Skip to main content
General Audio / Speech 320 hrs · 140 pro voice artists

Studio-Quality Domain Audio Set (Broadcast-Grade)

Clean, studio-recorded speech for domain-specific use cases — voice assistants, IVR, audiobooks and broadcast — with professional voice talent and tightly controlled acoustics.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Studio-Quality Domain Audio Set (Broadcast-Grade) dataset
Languages
Hindi, English + 4 Indic languages
Scale
320 hrs · 140 pro voice artists · 140 records
Inter-annotator agreement
α = 0.92
QA pass rate
99.30%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.92
QA pass rate
99.30%
Word error rate
3.80%

Recorded at 48 kHz / 24-bit in treated studios with a controlled noise floor. Transcripts reach a held-out WER of 3.8% with α = 0.92 label agreement and a 99.3% QA pass rate, including per-clip SNR and clipping checks.

Provenance & consent

Recorded with professional voice artists under fair-pay agreements that grant commercial and synthetic-voice reuse rights. Each session carries a talent-consent and usage-rights reference; all recording, processing and storage take place within India.

Methodology

Domain scripts (assistant, IVR, long-form narration) are recorded by vetted voice talent in acoustically treated rooms. Audio is mastered, loudness-normalised and QA'd for SNR, clipping and pronunciation, then aligned to verified transcripts before each versioned release.

Data preview

What's inside each record

A representative schema for Studio-Quality Domain Audio Set (Broadcast-Grade). The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
clip_id string (uuid) "spk_4421…"
audio_path string (wav, 16kHz) "clips/00421.wav"
transcript string (verbatim) "हम सुबह बाज़ार गए…"
duration_sec float 7.42
speaker_meta object {age_band, gender, region}
language string (ISO 639) "hin"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Commercial Non-Exclusive

Contact us

Production licence for TTS, voice-assistant and audiobook use cases across domains.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive voice packs, bespoke domains and dedicated talent recorded to spec.

Talk to sales

Related datasets

See all General datasets

Ready to license Studio-Quality Domain Audio Set (Broadcast-Grade)?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.