Skip to main content
General Audio / Speech ≥10 hrs/speaker (up to 100)

Paralinguistic & Expressive Events — Fillers, Laughter, Physiological

Annotated non-lexical speech events — fillers, laughter, physiological sounds and expressive modifiers — for controllable, natural-sounding TTS.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Paralinguistic & Expressive Events — Fillers, Laughter, Physiological dataset
Languages
English, Hindi (P0) + 9 P1 Indic languages
Scale
≥10 hrs/speaker (up to 100)
Quality & methodology metrics

Measured, audited, reproducible

Audio: WAV (PCM), mono, 48 kHz (min. 44.1 kHz), 24-bit — studio-recorded (professional studio preferred; treated home studio with prior approval).

Transcripts: both verbatim and normalized, with documented handling of numbers, dates, currency, time, URLs, email, abbreviations, symbols and punctuation, and native code-mixing.

Annotation: emotion, paralinguistics, disfluencies and prosody/emphasis applied pre-delivery by native-speaker annotators. Inter-annotator agreement (Cohen's / Fleiss' κ) is reported per annotation type as our QA standard, with a two-pass QA and adjudication workflow.

Deliverables: audio (WAV); transcripts (TXT/CSV/JSONL); annotations (JSON/JSONL/CSV); metadata (JSON/CSV).

Exclusions warranted in writing: no music/singing, no synthetic or TTS-generated audio, no copyrighted material, no ASR-style/web-scraped data.

Provenance & consent

All audio originates from human speakers recorded in vendor-owned, controlled studio sessions under written contributor agreements that grant commercial reuse, including explicit consent for AI training and commercial voice-cloning / synthetic-voice generation. Never scraped, never repurposed from call-centre or telephony recordings.

Methodology

A specialized corpus of paralinguistic and expressive events: conversational fillers (um, uh, ah, er, hm, mm), backchannels, laughter, cheering, exclamations, physiological sounds (sigh, inhale, cough, throat-clearing) and expressive modifiers (whisper, shouting, emphasis) — the required taxonomy — plus disfluencies (filled pauses, false starts, repetitions, repairs, hesitations, elongations). A sample annotation file is included.

Data preview

What's inside each record

A representative schema for Paralinguistic & Expressive Events — Fillers, Laughter, Physiological. The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
clip_id string (uuid) "spk_4421…"
audio_path string (wav, 16kHz) "clips/00421.wav"
transcript string (verbatim) "We went to the market this morning…"
duration_sec float 7.42
speaker_meta object {age_band, gender, region}
language string (ISO 639) "eng"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Non-Exclusive License

Contact us

Worldwide, perpetual, transferable, sublicensable license for commercial TTS training, voice cloning and customer-facing API use. Other parties may also license the same set.

Request access

Category-Exclusive License

Contact us

Exclusive within a defined use-case category (e.g. TTS/voice-cloning) while remaining licensable for other categories. Scoped as an upgrade on the non-exclusive tier.

Talk to sales

Exclusive License

Contact us

Full single-buyer exclusivity — no other party holds or will receive the data — with worldwide, perpetual, transferable, sublicensable rights and provenance transfer.

Talk to sales

Related datasets

See all General datasets

Ready to license Paralinguistic & Expressive Events — Fillers, Laughter, Physiological?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.