Rights-cleared sovereign datasets
- Datasets live
- 41
- Language footprint
- 400
- Annotation IAA
- ≥ 0.85
- Residency
- India
Every dataset is rights-cleared with documented consent and a downloadable dataset card. Pricing is shown per licence tier on each catalog card below.
Showing 20 of 41 datasets
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Audiobook & Long-Form Narration — Sustained Expressive
Sustained, expressive long-form narration for audiobook-grade TTS across P0 and P1 languages.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)
Native-authored red-team prompts and harm-labelled examples in code-mixed languages — surfacing caste, communal, gendered and code-switch harms English classifiers miss.
- Languages:
- Hinglish, Tanglish, English
- Size:
- 24,000 prompts · 12 harm categories
- IAA:
- 0.86
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Frontend & Conversational Voice — English + Hindi (Studio)
Natural, warm, neutral studio speech for assistant-style TTS in English and Hindi, including code-mixing — the P0 priority profile.
- Languages:
- English, Hindi (P0)
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
LegalLegal Text 🇮🇳 Sovereign
Indic Legal Judgments & Statutes Corpus
A structured, annotated corpus of Indian court judgments and statutes in four languages — with headnotes, citations, and clause-level entity tags.
- Languages:
- Hindi, Marathi, Tamil, English
- Size:
- 2.4M tokens · 31,000 documents
- IAA:
- 0.89
Rights-clearedLicence from
Contact us
-
Physical AIPhysical AI Video 🇮🇳 Sovereign
Multivision Egocentric Video Dataset (First-Person, Multi-View)
First-person (egocentric) video of everyday manual tasks, captured with synchronised multi-view rigs — for embodied AI, action recognition, hand-object interaction and imitation learning.
- Languages:
- Narration in English & Hindi
- Size:
- 220 hrs · 4-view sync · 1,600 sessions
- IAA:
- 0.89
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Studio-Quality Domain Audio Set (Broadcast-Grade)
Clean, studio-recorded speech for domain-specific use cases — voice assistants, IVR, audiobooks and broadcast — with professional voice talent and tightly controlled acoustics.
- Languages:
- Hindi, English + 4 Indic languages
- Size:
- 320 hrs · 140 pro voice artists
- IAA:
- 0.92
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Child Voices — Multiple Age Ranges (Studio)
Studio child speech across age ranges for child-voice TTS, offered as additional speakers.
- Languages:
- English, Hindi (P0)
- Size:
- <10 hrs/speaker (availability-based)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Text 🇮🇳 Sovereign
Code-Switched Hinglish & Tanglish Conversation Set
Real code-switched Hinglish and Tanglish conversations with language-ID, intent, and sentiment tags — the messy, mixed text production models actually face.
- Languages:
- Hinglish, Tanglish, English
- Size:
- 1.1M tokens · 55,000 messages
- IAA:
- 0.83
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Customer Support Voice — Empathetic & Measured
Clear, empathetic, measured support-agent studio speech for service-oriented TTS, with English terms embedded in Hindi.
- Languages:
- English, Hindi (P0)
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Dramatic & Character Voice — Emotional Range & Personas
Character-driven, emotionally ranged dramatic studio speech for expressive and character TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Elderly Voices — Senior Speakers (Studio)
Studio senior/elderly studio speech for age-diverse TTS, offered as additional speakers.
- Languages:
- English, Hindi (P0)
- Size:
- <10 hrs/speaker (availability-based)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Emotional & Acted Speech — Full Emotion Taxonomy
Emotion-labeled acted studio speech across eight categorical emotions for expressive TTS and emotion modeling.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
LegalLegal Text 🇮🇳 Sovereign
Indic Legal Contracts & Clause-Type Corpus
A rights-clean corpus of commercial contracts with clause-type, party and obligation annotations — for contract-analysis and legal-search AI.
- Languages:
- Hindi, English, Marathi
- Size:
- 1.1M tokens · 9,500 contracts
- IAA:
- 0.88
Rights-clearedLicence from
Contact us
-
HealthcareHealthcare Text 🇮🇳 Sovereign
Indic Radiology Report NLP Set (De-identified)
De-identified radiology reports with finding, impression and negation annotations — for clinical NLP, report generation and evaluation.
- Languages:
- English, Hindi
- Size:
- 180,000 reports
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
News & Educational Voice — Authoritative Neutral
Authoritative, neutral read studio speech for news-reader and educational TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Paralinguistic & Expressive Events — Fillers, Laughter, Physiological
Annotated non-lexical speech events — fillers, laughter, physiological sounds and expressive modifiers — for controllable, natural-sounding TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
Physical AIPhysical AI Sensor / Physical 🇮🇳 Sovereign
Physical AI Annotation Sample — Multi-Sensor Scenes
A sample of our Physical AI annotation service output: 3D boxes, point-cloud segmentation, and 6-DoF pose on multi-sensor scenes — a showcase of annotation capability, not a research release.
- Languages:
- Labels in English
- Size:
- 1,200 annotated frames
- IAA:
- 0.91
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Prosody & Emphasis — Word/Phrase Stress & Contrastive
Prosody-annotated studio speech (word/phrase emphasis, contrastive stress) for emphasis-controllable TTS.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Sales & Energetic Voice — High-Energy Persuasive
High-energy, persuasive, ad-style read studio speech for marketing and sales TTS.
- Languages:
- English, Hindi (P0)
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
-
GeneralGeneral Audio / Speech 🇮🇳 Sovereign
Short Commands & Utterances — Sub-2s and 2–5s Clips
Very short commands and short utterances for command-style TTS and micro-utterance synthesis.
- Languages:
- English, Hindi (P0) + 9 P1 Indic languages
- Size:
- ≥10 hrs/speaker (up to 100)
Rights-clearedLicence from
Contact us
Browse datasets by language
Dedicated pages for each language — including the low-resource Indic and African languages most vendors can't reach.
Browse datasets by domain
Vertical landing pages for each domain — built for the way each industry actually talks.
Languages we cover, across every continent
Our coverage spans a 400-language footprint — Indic and low-resource languages first, extending across Africa, Asia, the Americas and beyond. Select a continent to filter the list, or search by name, native script or ISO code. Multilingual data infrastructure built for language sovereignty, not just the major markets.
India — deepest Indic & low-resource coverage
South Asia flagged — our deepest Indic & low-resource coverage
India shown as a map; continents as a coverage grid — see the full searchable list (right) for every language we cover.
languages shown (filtered) · 400-language footprint in scope
Language coverage
Which languages does Cognegica's dataset registry cover?
Cognegica enumerates 392 languages across the registry within a 400-language collection footprint, with deep coverage of the low-resource long tail. Counts below are drawn directly from the registry, grouped by continent.
| Region / continent | Languages enumerated | Low-resource |
|---|---|---|
| South Asia | 110 | 97 |
| Asia | 43 | 35 |
| Middle East | 11 | 9 |
| Central Asia | 10 | 8 |
| Africa | 132 | 131 |
| Europe | 18 | 11 |
| Americas | 9 | 8 |
| Oceania | 6 | 6 |
| Global / Other | 53 | 37 |
| Total enumerated registry | 392 | 342 |
Enumerated registry counts by continent. The headline 400-language footprint includes additional collectable languages scoped per project.
FAQ
Dataset catalog — common questions
- What kind of datasets do you license?
Proprietary, rights-cleared, sovereign datasets for low-resource and Indic languages across a 400-language footprint — covering AI safety and evaluation, legal, healthcare, general multilingual, and Physical AI domains, in text, audio, speech, image, video and sensor modalities.
- Are the datasets rights-cleared and safe to train on?
Yes. Every dataset is authored or collected from consenting native speakers under agreements that grant commercial reuse, with documented provenance and consent per record — never scraped. Published quality metrics include inter-annotator agreement and QA pass rates.
- How does dataset licensing work?
Each dataset offers commercial non-exclusive and enterprise/sovereign exclusive licensing. Licensing is scoped and quoted per dataset against languages, modalities, volume and rights. Contact us for terms.
- Can the data stay in India (data residency)?
Yes. We offer an India data-residency / sovereign delivery option across collection, processing, annotation and storage, aligned with the DPDP Act, 2023. See Sovereign Data.
- Do you have Physical AI datasets?
Yes — Physical AI is a data service, not research. Browse multi-sensor and annotation samples at Physical AI datasets, or engage our collection and annotation services to build to spec.
- Can you build a custom dataset in a language not listed?
Yes. When the data you need does not exist, we collect it to spec across our 400-language footprint. Request a custom dataset.
Don't see the language or domain you need?
We build proprietary, rights-cleared datasets to spec — across 400 Indian languages and beyond. Tell us your target languages, modality, and quality bar.