Skip to main content
General Audio / Speech 410 hrs · 1,300 speakers · 12 languages

Low-Resource Language ASR & TTS Corpus

Paired speech-and-text for low-resource languages, built for both ASR and TTS — read prompts plus spontaneous speech, with studio-grade single-speaker sets for voice building.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Low-Resource Language ASR & TTS Corpus dataset
Languages
Santali, Bodo, Dogri, Maithili + Global-South LRLs
Scale
410 hrs · 1,300 speakers · 12 languages · 1300 records
Inter-annotator agreement
α = 0.87
QA pass rate
98.00%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.87
QA pass rate
98.00%
Word error rate
9.70%

ASR splits reach a held-out WER of 9.7%; TTS voices are recorded as single-speaker studio sets with phonetic coverage checks and α = 0.87 transcript agreement at a 98.0% QA pass rate. Lexicons and grapheme-to-phoneme notes ship with each language.

Provenance & consent

Collected from consenting native speakers under fair-pay agreements granting commercial and synthetic-voice reuse. Each recording carries consent, speaker and language-variety references; Indic-language data is stored and processed within India, with in-region handling for others.

Methodology

Each language pairs read prompts (for phonetic coverage and TTS) with spontaneous speech (for ASR robustness). Native transcribers and linguists produce verified transcripts, pronunciation lexicons and G2P notes; a second-pass audit verifies WER and TTS voice consistency per release.

Data preview

What's inside each record

A representative schema for Low-Resource Language ASR & TTS Corpus. The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
clip_id string (uuid) "spk_4421…"
audio_path string (wav, 16kHz) "clips/00421.wav"
transcript string (verbatim) "ᱟᱞᱮ ᱥᱮᱛᱟᱜ ᱨᱮ ᱪᱟᱞᱟᱜ…"
duration_sec float 7.42
speaker_meta object {age_band, gender, region}
language string (ISO 639) "san"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Commercial Non-Exclusive

Contact us

Production licence for low-resource ASR and TTS, with paired read + spontaneous splits.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive licensing plus additional low-resource languages and single-speaker TTS voices to spec.

Talk to sales

Related datasets

See all General datasets

Ready to license Low-Resource Language ASR & TTS Corpus?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.