Studio-Quality Domain Audio Set (Broadcast-Grade)
Clean, studio-recorded speech for domain-specific use cases — voice assistants, IVR, audiobooks and broadcast — with professional voice talent and tightly controlled acoustics.
- Rights-cleared
- Consent documented
- 🇮🇳 India data residency
- Languages
- Hindi, English + 4 Indic languages
- Scale
- 320 hrs · 140 pro voice artists · 140 records
- Inter-annotator agreement
- α = 0.92
- QA pass rate
- 99.30%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.92
- QA pass rate
- 99.30%
- Word error rate
- 3.80%
Recorded at 48 kHz / 24-bit in treated studios with a controlled noise floor. Transcripts reach a held-out WER of 3.8% with α = 0.92 label agreement and a 99.3% QA pass rate, including per-clip SNR and clipping checks.
Recorded with professional voice artists under fair-pay agreements that grant commercial and synthetic-voice reuse rights. Each session carries a talent-consent and usage-rights reference; all recording, processing and storage take place within India.
Domain scripts (assistant, IVR, long-form narration) are recorded by vetted voice talent in acoustically treated rooms. Audio is mastered, loudness-normalised and QA'd for SNR, clipping and pronunciation, then aligned to verified transcripts before each versioned release.
What's inside each record
A representative schema for Studio-Quality Domain Audio Set (Broadcast-Grade). The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| clip_id | string (uuid) | "spk_4421…" |
| audio_path | string (wav, 16kHz) | "clips/00421.wav" |
| transcript | string (verbatim) | "हम सुबह बाज़ार गए…" |
| duration_sec | float | 7.42 |
| speaker_meta | object | {age_band, gender, region} |
| language | string (ISO 639) | "hin" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Commercial Non-Exclusive
Contact us
Production licence for TTS, voice-assistant and audiobook use cases across domains.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive voice packs, bespoke domains and dedicated talent recorded to spec.
Talk to salesRelated datasets
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Audiobook & Long-Form Narration — Sustained Expressive
Sustained, expressive long-form narration for audiobook-grade TTS across P0 and P1 languages.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Child-Voice Expressive Speech Set (Major Indian Languages)
Consented child-voice speech in major Indian languages — expressive, read and spontaneous, with code-switching — for age-appropriate ASR, TTS and edtech voice models.
- Languages:
- Hindi, Tamil, Telugu, Bengali, Marathi, Kannada
- Size:
- 180 hrs · 2,400 child speakers
- IAA:
- 0.86
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Frontend & Conversational Voice — English + Hindi (Studio)
Natural, warm, neutral studio speech for assistant-style TTS in English and Hindi, including code-mixing — the P0 priority profile.
- Languages:
- English, Hindi (P0)
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
Ready to license Studio-Quality Domain Audio Set (Broadcast-Grade)?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.