Legal AI Training Datasets — Multilingual & Indic
Legal AI fails the moment it meets a Hindi cause-list, a Tamil judgment or a code-mixed affidavit. Our legal datasets are expert-annotated, rights-cleared and de-identified — court judgments, statutes and contracts across Indian languages, with clause-level entity tags and citation structure your model can actually learn from.
Legal datasets in the catalog
Showing 2 of 41 datasets.
-
LegalLegal Text 🇮🇳 Sovereign
Indic Legal Judgments & Statutes Corpus
A structured, annotated corpus of Indian court judgments and statutes in four languages — with headnotes, citations, and clause-level entity tags.
- Languages:
- Hindi, Marathi, Tamil, English
- Size:
- 2.4M tokens · 31,000 documents
- IAA:
- 0.89
Rights-clearedLicence from
Contact us
-
LegalLegal Text 🇮🇳 Sovereign
Indic Legal Contracts & Clause-Type Corpus
A rights-clean corpus of commercial contracts with clause-type, party and obligation annotations — for contract-analysis and legal-search AI.
- Languages:
- Hindi, English, Marathi
- Size:
- 1.1M tokens · 9,500 contracts
- IAA:
- 0.88
Rights-clearedLicence from
Contact us
Services & industry for Legal
When the data you need doesn't exist yet, we build it — collection, annotation, alignment and evaluation, all rights-cleared and documented.
Data Annotation
Speech, NLP, CV and multimodal annotation at IAA ≥ 0.85 with two-pass QA — built for foundation-model SLAs.
Explore the serviceSFT Gold-Standard Data
High-quality prompt/response curation for supervised fine-tuning — vetted by linguistic SMEs, not crowd-sourced noise.
Explore the serviceMultilingual Data Collection
Native-speaker audio, video, image and text collection across Indic, African and low-resource languages — field-grade, consented, documented.
Explore the serviceLegal industry
Multilingual legal-document, contract and judgment data — annotated by domain experts for legal-AI products that can't afford to hallucinate.
Explore the industryCommon questions
- Are the legal datasets rights-cleared and de-identified?
Yes. Source judgments come from publicly published court records and statutes from official gazettes; parties and PII are de-identified with documented SOPs, and processing happens within India.
- Which languages do the legal datasets cover?
Hindi, Marathi, Tamil and English today, with additional Indian languages and document types built to spec on request.
Need Legal data your model can train on?
License what's in the catalog, or tell us exactly what you need — we build proprietary, rights-cleared datasets to spec, with India data residency.