Low-Resource Language ASR & TTS Corpus
Paired speech-and-text for low-resource languages, built for both ASR and TTS — read prompts plus spontaneous speech, with studio-grade single-speaker sets for voice building.
- Rights-cleared
- Consent documented
- 🇮🇳 India data residency
- Languages
- Santali, Bodo, Dogri, Maithili + Global-South LRLs
- Scale
- 410 hrs · 1,300 speakers · 12 languages · 1300 records
- Inter-annotator agreement
- α = 0.87
- QA pass rate
- 98.00%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.87
- QA pass rate
- 98.00%
- Word error rate
- 9.70%
ASR splits reach a held-out WER of 9.7%; TTS voices are recorded as single-speaker studio sets with phonetic coverage checks and α = 0.87 transcript agreement at a 98.0% QA pass rate. Lexicons and grapheme-to-phoneme notes ship with each language.
Collected from consenting native speakers under fair-pay agreements granting commercial and synthetic-voice reuse. Each recording carries consent, speaker and language-variety references; Indic-language data is stored and processed within India, with in-region handling for others.
Each language pairs read prompts (for phonetic coverage and TTS) with spontaneous speech (for ASR robustness). Native transcribers and linguists produce verified transcripts, pronunciation lexicons and G2P notes; a second-pass audit verifies WER and TTS voice consistency per release.
What's inside each record
A representative schema for Low-Resource Language ASR & TTS Corpus. The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| clip_id | string (uuid) | "spk_4421…" |
| audio_path | string (wav, 16kHz) | "clips/00421.wav" |
| transcript | string (verbatim) | "ᱟᱞᱮ ᱥᱮᱛᱟᱜ ᱨᱮ ᱪᱟᱞᱟᱜ…" |
| duration_sec | float | 7.42 |
| speaker_meta | object | {age_band, gender, region} |
| language | string (ISO 639) | "san" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Commercial Non-Exclusive
Contact us
Production licence for low-resource ASR and TTS, with paired read + spontaneous splits.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive licensing plus additional low-resource languages and single-speaker TTS voices to spec.
Talk to salesRelated datasets
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Audiobook & Long-Form Narration — Sustained Expressive
Sustained, expressive long-form narration for audiobook-grade TTS across P0 and P1 languages.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Child-Voice Expressive Speech Set (Major Indian Languages)
Consented child-voice speech in major Indian languages — expressive, read and spontaneous, with code-switching — for age-appropriate ASR, TTS and edtech voice models.
- Languages:
- Hindi, Tamil, Telugu, Bengali, Marathi, Kannada
- Size:
- 180 hrs · 2,400 child speakers
- IAA:
- 0.86
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Frontend & Conversational Voice — English + Hindi (Studio)
Natural, warm, neutral studio speech for assistant-style TTS in English and Hindi, including code-mixing — the P0 priority profile.
- Languages:
- English, Hindi (P0)
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
Explore more General data
Ready to license Low-Resource Language ASR & TTS Corpus?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.