Rights-cleared sovereign datasets
- Datasets live
- 41
- Language footprint
- 400
- Annotation IAA
- ≥ 0.85
- Residency
- India
Every dataset is rights-cleared with documented consent and a downloadable dataset card. Pricing is shown per licence tier on each catalog card below.
Showing 3 of 41 datasets
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Low-Resource Language ASR & TTS Corpus
Paired speech-and-text for low-resource languages, built for both ASR and TTS — read prompts plus spontaneous speech, with studio-grade single-speaker sets for voice building.
- Languages:
- Santali, Bodo, Dogri, Maithili + Global-South LRLs
- Size:
- 410 hrs · 1,300 speakers · 12 languages
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
South Asian Instruction-Following Eval (Rare Languages)
An instruction-following and reasoning evaluation suite for four under-served South Asian languages, with human reference answers and rubric-based scoring.
- Languages:
- Maithili, Bhojpuri, Santali, Sindhi
- Size:
- 9,600 instruction–response pairs
- IAA:
- 0.84
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Rare-Language Conversational Speech Corpus
Spontaneous, real-world conversational speech in four rare Indian languages — transcribed, segmented, and ready for ASR and speech-LLM training.
- Languages:
- Santali, Gondi, Tulu, Bhojpuri
- Size:
- 260 hrs · 1,850 speakers
- IAA:
- 0.85
Rights-clearedLicence from
Contact us
Browse datasets by language
Dedicated pages for each language — including the low-resource Indic and African languages most vendors can't reach.
Browse datasets by domain
Vertical landing pages for each domain — built for the way each industry actually talks.
Languages we cover, across every continent
Our coverage spans a 400-language footprint — Indic and low-resource languages first, extending across Africa, Asia, the Americas and beyond. Select a continent to filter the list, or search by name, native script or ISO code. Multilingual data infrastructure built for language sovereignty, not just the major markets.
India — deepest Indic & low-resource coverage
South Asia flagged — our deepest Indic & low-resource coverage
India shown as a map; continents as a coverage grid — see the full searchable list (right) for every language we cover.
languages shown (filtered) · 400-language footprint in scope
Language coverage
Which languages does Cognegica's dataset registry cover?
Cognegica enumerates 392 languages across the registry within a 400-language collection footprint, with deep coverage of the low-resource long tail. Counts below are drawn directly from the registry, grouped by continent.
| Region / continent | Languages enumerated | Low-resource |
|---|---|---|
| South Asia | 110 | 97 |
| Asia | 43 | 35 |
| Middle East | 11 | 9 |
| Central Asia | 10 | 8 |
| Africa | 132 | 131 |
| Europe | 18 | 11 |
| Americas | 9 | 8 |
| Oceania | 6 | 6 |
| Global / Other | 53 | 37 |
| Total enumerated registry | 392 | 342 |
Enumerated registry counts by continent. The headline 400-language footprint includes additional collectable languages scoped per project.
FAQ
Dataset catalog — common questions
- What kind of datasets do you license?
Proprietary, rights-cleared, sovereign datasets for low-resource and Indic languages across a 400-language footprint — covering AI safety and evaluation, legal, healthcare, general multilingual, and Physical AI domains, in text, audio, speech, image, video and sensor modalities.
- Are the datasets rights-cleared and safe to train on?
Yes. Every dataset is authored or collected from consenting native speakers under agreements that grant commercial reuse, with documented provenance and consent per record — never scraped. Published quality metrics include inter-annotator agreement and QA pass rates.
- How does dataset licensing work?
Each dataset offers commercial non-exclusive and enterprise/sovereign exclusive licensing. Licensing is scoped and quoted per dataset against languages, modalities, volume and rights. Contact us for terms.
- Can the data stay in India (data residency)?
Yes. We offer an India data-residency / sovereign delivery option across collection, processing, annotation and storage, aligned with the DPDP Act, 2023. See Sovereign Data.
- Do you have Physical AI datasets?
Yes — Physical AI is a data service, not research. Browse multi-sensor and annotation samples at Physical AI datasets, or engage our collection and annotation services to build to spec.
- Can you build a custom dataset in a language not listed?
Yes. When the data you need does not exist, we collect it to spec across our 400-language footprint. Request a custom dataset.
Don't see the language or domain you need?
We build proprietary, rights-cleared datasets to spec — across 400 Indian languages and beyond. Tell us your target languages, modality, and quality bar.