Indic Healthcare Patient–Clinician Dialogue Set
De-identified, consented patient–clinician dialogues across four languages, with intent, symptom, and clinical-entity annotations for health AI.
- Rights-cleared
- Consent documented
- 🇮🇳 India data residency
- Languages
- Hindi, Telugu, Bengali, Kannada
- Scale
- 120 hrs audio · 42,000 turns · 42000 records
- Inter-annotator agreement
- α = 0.86
- QA pass rate
- 98.70%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.86
- QA pass rate
- 98.70%
- Word error rate
- 7.40%
Transcription word error rate is 7.4% after two-pass review; clinical-entity annotation reaches α = 0.86 with a 98.7% QA pass rate. Annotation covers symptoms, durations, medications, and follow-up intents.
Every dialogue is collected with explicit, written informed consent from both patient and clinician, then rigorously de-identified to remove direct and quasi-identifiers. Collection, de-identification, and storage are performed entirely within India in line with the DPDP Act, 2023.
Recordings are captured in real consultation settings with consent, transcribed by medically-literate native speakers, then de-identified under a documented SOP and validated by an independent reviewer. Clinical entities are annotated against a curated ontology with clinician sign-off.
What's inside each record
A representative schema for Indic Healthcare Patient–Clinician Dialogue Set. The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| encounter_id | string (uuid) | "enc_b21d…" |
| audio_path | string (wav, 16kHz) | "clips/enc_b21d.wav" |
| transcript | string (de-identified) | "मरीज़ को दो दिन से बुखार है…" |
| speaker_role | enum | "clinician" | "patient" |
| intent | enum | "symptom_report" |
| phi_redacted | bool | true |
| language | string (ISO 639) | "hin" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Commercial Non-Exclusive
Contact us
Production licence for health-AI products with audio, transcripts, and clinical annotations.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive licensing, additional specialties or languages, and on-prem / sovereign delivery.
Talk to salesRelated datasets
-
HealthcareHealthcare Text 🇮🇳 Sovereign
Indic Radiology Report NLP Set (De-identified)
De-identified radiology reports with finding, impression and negation annotations — for clinical NLP, report generation and evaluation.
- Languages:
- English, Hindi
- Size:
- 180,000 reports
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
Explore more Healthcare data
Ready to license Indic Healthcare Patient–Clinician Dialogue Set?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.