Skip to main content
Healthcare Text 180,000 reports

Indic Radiology Report NLP Set (De-identified)

De-identified radiology reports with finding, impression and negation annotations — for clinical NLP, report generation and evaluation.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Indic Radiology Report NLP Set (De-identified) dataset
Languages
English, Hindi
Scale
180,000 reports · 180000 records
Inter-annotator agreement
α = 0.87
QA pass rate
98.40%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.87
QA pass rate
98.40%

Finding, impression and negation spans reach α = 0.87 with a 98.4% QA pass rate, reviewed with clinician sign-off. Annotations include anatomy, finding, severity and negation scope.

Provenance & consent

Reports are contributed with consent and rigorously de-identified to remove direct and quasi-identifiers before annotation. Collection, de-identification and storage occur entirely within India, aligned with the DPDP Act, 2023.

Methodology

Reports are de-identified under a documented SOP, validated by an independent reviewer, then annotated by medically-literate annotators against a curated ontology with clinician adjudication of contested cases.

Data preview

What's inside each record

A representative schema for Indic Radiology Report NLP Set (De-identified). The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
encounter_id string (uuid) "enc_b21d…"
audio_path string (wav, 16kHz) "clips/enc_b21d.wav"
transcript string (de-identified) "Patient reports two days of fever…"
speaker_role enum "clinician" | "patient"
intent enum "symptom_report"
phi_redacted bool true
language string (ISO 639) "eng"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Commercial Non-Exclusive

Contact us

Production licence for clinical-NLP and report-generation products with periodic refresh.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive licensing, additional specialties and on-prem / sovereign delivery.

Talk to sales

Related datasets

  • Healthcare
    Healthcare Multimodal 🇮🇳 Sovereign

    Indic Healthcare Patient–Clinician Dialogue Set

    De-identified, consented patient–clinician dialogues across four languages, with intent, symptom, and clinical-entity annotations for health AI.

    Languages:
    Hindi, Telugu, Bengali, Kannada
    Size:
    120 hrs audio · 42,000 turns
    IAA:
    0.86

    Licence from

    Contact us

    Rights-cleared

See all Healthcare datasets

Ready to license Indic Radiology Report NLP Set (De-identified)?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.