Rights-cleared sovereign datasets
- Datasets live
- 41
- Language footprint
- 400
- Annotation IAA
- ≥ 0.85
- Residency
- India
Every dataset is rights-cleared with documented consent and a downloadable dataset card. Pricing is shown per licence tier on each catalog card below.
Showing 7 of 41 datasets
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Audiobook & Long-Form Narration — Sustained Expressive
Sustained, expressive long-form narration for audiobook-grade TTS across P0 and P1 languages.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Dramatic & Character Voice — Emotional Range & Personas
Character-driven, emotionally ranged dramatic studio speech for expressive and character TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Emotional & Acted Speech — Full Emotion Taxonomy
Emotion-labeled acted studio speech across eight categorical emotions for expressive TTS and emotion modeling.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
News & Educational Voice — Authoritative Neutral
Authoritative, neutral read studio speech for news-reader and educational TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Paralinguistic & Expressive Events — Fillers, Laughter, Physiological
Annotated non-lexical speech events — fillers, laughter, physiological sounds and expressive modifiers — for controllable, natural-sounding TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Prosody & Emphasis — Word/Phrase Stress & Contrastive
Prosody-annotated studio speech (word/phrase emphasis, contrastive stress) for emphasis-controllable TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Short Commands & Utterances — Sub-2s and 2–5s Clips
Very short commands and short utterances for command-style TTS and micro-utterance synthesis.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
Browse datasets by language
Dedicated pages for each language — including the low-resource Indic and African languages most vendors can't reach.
Browse datasets by domain
Vertical landing pages for each domain — built for the way each industry actually talks.
Languages we cover, across every continent
Our coverage spans a 400-language footprint — Indic and low-resource languages first, extending across Africa, Asia, the Americas and beyond. Select a continent to filter the list, or search by name, native script or ISO code. Multilingual data infrastructure built for language sovereignty, not just the major markets.
India — deepest Indic & low-resource coverage
South Asia flagged — our deepest Indic & low-resource coverage
India shown as a map; continents as a coverage grid — see the full searchable list (right) for every language we cover.
languages shown (filtered) · 400-language footprint in scope
Language coverage
Which languages does Cognegica's dataset registry cover?
Cognegica enumerates 392 languages across the registry within a 400-language collection footprint, with deep coverage of the low-resource long tail. Counts below are drawn directly from the registry, grouped by continent.
| Region / continent | Languages enumerated | Low-resource |
|---|---|---|
| South Asia | 110 | 97 |
| Asia | 43 | 35 |
| Middle East | 11 | 9 |
| Central Asia | 10 | 8 |
| Africa | 132 | 131 |
| Europe | 18 | 11 |
| Americas | 9 | 8 |
| Oceania | 6 | 6 |
| Global / Other | 53 | 37 |
| Total enumerated registry | 392 | 342 |
Enumerated registry counts by continent. The headline 400-language footprint includes additional collectable languages scoped per project.
FAQ
Dataset catalog — common questions
- What kind of datasets do you license?
Proprietary, rights-cleared, sovereign datasets for low-resource and Indic languages across a 400-language footprint — covering AI safety and evaluation, legal, healthcare, general multilingual, and Physical AI domains, in text, audio, speech, image, video and sensor modalities.
- Are the datasets rights-cleared and safe to train on?
Yes. Every dataset is authored or collected from consenting native speakers under agreements that grant commercial reuse, with documented provenance and consent per record — never scraped. Published quality metrics include inter-annotator agreement and QA pass rates.
- How does dataset licensing work?
Each dataset offers commercial non-exclusive and enterprise/sovereign exclusive licensing. Licensing is scoped and quoted per dataset against languages, modalities, volume and rights. Contact us for terms.
- Can the data stay in India (data residency)?
Yes. We offer an India data-residency / sovereign delivery option across collection, processing, annotation and storage, aligned with the DPDP Act, 2023. See Sovereign Data.
- Do you have Physical AI datasets?
Yes — Physical AI is a data service, not research. Browse multi-sensor and annotation samples at Physical AI datasets, or engage our collection and annotation services to build to spec.
- Can you build a custom dataset in a language not listed?
Yes. When the data you need does not exist, we collect it to spec across our 400-language footprint. Request a custom dataset.
Don't see the language or domain you need?
We build proprietary, rights-cleared datasets to spec — across 400 Indian languages and beyond. Tell us your target languages, modality, and quality bar.