Skip to main content
Healthcare Multimodal 120 hrs audio · 42,000 turns

Indic Healthcare Patient–Clinician Dialogue Set

De-identified, consented patient–clinician dialogues across four languages, with intent, symptom, and clinical-entity annotations for health AI.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Indic Healthcare Patient–Clinician Dialogue Set dataset
Languages
Hindi, Telugu, Bengali, Kannada
Scale
120 hrs audio · 42,000 turns · 42000 records
Inter-annotator agreement
α = 0.86
QA pass rate
98.70%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.86
QA pass rate
98.70%
Word error rate
7.40%

Transcription word error rate is 7.4% after two-pass review; clinical-entity annotation reaches α = 0.86 with a 98.7% QA pass rate. Annotation covers symptoms, durations, medications, and follow-up intents.

Provenance & consent

Every dialogue is collected with explicit, written informed consent from both patient and clinician, then rigorously de-identified to remove direct and quasi-identifiers. Collection, de-identification, and storage are performed entirely within India in line with the DPDP Act, 2023.

Methodology

Recordings are captured in real consultation settings with consent, transcribed by medically-literate native speakers, then de-identified under a documented SOP and validated by an independent reviewer. Clinical entities are annotated against a curated ontology with clinician sign-off.

Data preview

What's inside each record

A representative schema for Indic Healthcare Patient–Clinician Dialogue Set. The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
encounter_id string (uuid) "enc_b21d…"
audio_path string (wav, 16kHz) "clips/enc_b21d.wav"
transcript string (de-identified) "मरीज़ को दो दिन से बुखार है…"
speaker_role enum "clinician" | "patient"
intent enum "symptom_report"
phi_redacted bool true
language string (ISO 639) "hin"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Commercial Non-Exclusive

Contact us

Production licence for health-AI products with audio, transcripts, and clinical annotations.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive licensing, additional specialties or languages, and on-prem / sovereign delivery.

Talk to sales

Related datasets

See all Healthcare datasets

Ready to license Indic Healthcare Patient–Clinician Dialogue Set?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.