Child-Voice Expressive Speech Set (Major Indian Languages)
Consented child-voice speech in major Indian languages — expressive, read and spontaneous, with code-switching — for age-appropriate ASR, TTS and edtech voice models.
- Rights-cleared
- Consent documented
- 🇮🇳 India data residency
- Languages
- Hindi, Tamil, Telugu, Bengali, Marathi, Kannada
- Scale
- 180 hrs · 2,400 child speakers · 2400 records
- Inter-annotator agreement
- α = 0.86
- QA pass rate
- 98.10%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.86
- QA pass rate
- 98.10%
- Word error rate
- 8.60%
Transcripts and expressive-style labels (read, playful, storytelling, emphatic) are double-annotated at α = 0.86 with a 98.1% QA pass rate and a held-out WER of 8.6%. Balanced across age bands and gender.
Recorded from consenting minors under strict guardian-consent and child-safeguarding protocols, with fair-pay agreements granting commercial reuse. Every clip carries guardian-consent and age-band references; all data is stored and processed within India.
Guardian-supervised sessions capture read, spontaneous and expressive child speech — including natural code-switching — across quiet and everyday acoustic settings. Native annotators transcribe, tag expressive style and code-switch spans, and a second reviewer plus a safeguarding audit verify quality before each versioned release.
What's inside each record
A representative schema for Child-Voice Expressive Speech Set (Major Indian Languages). The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| clip_id | string (uuid) | "spk_4421…" |
| audio_path | string (wav, 16kHz) | "clips/00421.wav" |
| transcript | string (verbatim) | "हम सुबह बाज़ार गए…" |
| duration_sec | float | 7.42 |
| speaker_meta | object | {age_band, gender, region} |
| language | string (ISO 639) | "hin" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Commercial Non-Exclusive
Contact us
Production licence for child-directed ASR, TTS and edtech voice models, with balanced age/gender splits.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive licensing plus additional languages and expressive styles collected to spec.
Talk to salesRelated datasets
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Audiobook & Long-Form Narration — Sustained Expressive
Sustained, expressive long-form narration for audiobook-grade TTS across P0 and P1 languages.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Frontend & Conversational Voice — English + Hindi (Studio)
Natural, warm, neutral studio speech for assistant-style TTS in English and Hindi, including code-mixing — the P0 priority profile.
- Languages:
- English, Hindi (P0)
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Low-Resource Language ASR & TTS Corpus
Paired speech-and-text for low-resource languages, built for both ASR and TTS — read prompts plus spontaneous speech, with studio-grade single-speaker sets for voice building.
- Languages:
- Santali, Bodo, Dogri, Maithili + Global-South LRLs
- Size:
- 410 hrs · 1,300 speakers · 12 languages
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
Ready to license Child-Voice Expressive Speech Set (Major Indian Languages)?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.