Skip to main content
General Audio / Speech 180 hrs · 2,400 child speakers

Child-Voice Expressive Speech Set (Major Indian Languages)

Consented child-voice speech in major Indian languages — expressive, read and spontaneous, with code-switching — for age-appropriate ASR, TTS and edtech voice models.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Child-Voice Expressive Speech Set (Major Indian Languages) dataset
Languages
Hindi, Tamil, Telugu, Bengali, Marathi, Kannada
Scale
180 hrs · 2,400 child speakers · 2400 records
Inter-annotator agreement
α = 0.86
QA pass rate
98.10%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.86
QA pass rate
98.10%
Word error rate
8.60%

Transcripts and expressive-style labels (read, playful, storytelling, emphatic) are double-annotated at α = 0.86 with a 98.1% QA pass rate and a held-out WER of 8.6%. Balanced across age bands and gender.

Provenance & consent

Recorded from consenting minors under strict guardian-consent and child-safeguarding protocols, with fair-pay agreements granting commercial reuse. Every clip carries guardian-consent and age-band references; all data is stored and processed within India.

Methodology

Guardian-supervised sessions capture read, spontaneous and expressive child speech — including natural code-switching — across quiet and everyday acoustic settings. Native annotators transcribe, tag expressive style and code-switch spans, and a second reviewer plus a safeguarding audit verify quality before each versioned release.

Data preview

What's inside each record

A representative schema for Child-Voice Expressive Speech Set (Major Indian Languages). The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
clip_id string (uuid) "spk_4421…"
audio_path string (wav, 16kHz) "clips/00421.wav"
transcript string (verbatim) "हम सुबह बाज़ार गए…"
duration_sec float 7.42
speaker_meta object {age_band, gender, region}
language string (ISO 639) "hin"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Commercial Non-Exclusive

Contact us

Production licence for child-directed ASR, TTS and edtech voice models, with balanced age/gender splits.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive licensing plus additional languages and expressive styles collected to spec.

Talk to sales

Related datasets

See all General datasets

Ready to license Child-Voice Expressive Speech Set (Major Indian Languages)?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.