Skip to main content
Legal datasets

Legal AI Training Datasets — Multilingual & Indic

Legal AI fails the moment it meets a Hindi cause-list, a Tamil judgment or a code-mixed affidavit. Our legal datasets are expert-annotated, rights-cleared and de-identified — court judgments, statutes and contracts across Indian languages, with clause-level entity tags and citation structure your model can actually learn from.

Rights-cleared, multilingual dataset catalog for Legal AI Training Datasets — Multilingual & Indic
Matching datasets

Legal datasets in the catalog

Showing 2 of 41 datasets.

  • Legal
    Legal Text 🇮🇳 Sovereign

    Indic Legal Judgments & Statutes Corpus

    A structured, annotated corpus of Indian court judgments and statutes in four languages — with headnotes, citations, and clause-level entity tags.

    Languages:
    Hindi, Marathi, Tamil, English
    Size:
    2.4M tokens · 31,000 documents
    IAA:
    0.89

    Licence from

    Contact us

    Rights-cleared
  • Legal
    Legal Text 🇮🇳 Sovereign

    Indic Legal Contracts & Clause-Type Corpus

    A rights-clean corpus of commercial contracts with clause-type, party and obligation annotations — for contract-analysis and legal-search AI.

    Languages:
    Hindi, English, Marathi
    Size:
    1.1M tokens · 9,500 contracts
    IAA:
    0.88

    Licence from

    Contact us

    Rights-cleared
FAQ

Common questions

Are the legal datasets rights-cleared and de-identified?

Yes. Source judgments come from publicly published court records and statutes from official gazettes; parties and PII are de-identified with documented SOPs, and processing happens within India.

Which languages do the legal datasets cover?

Hindi, Marathi, Tamil and English today, with additional Indian languages and document types built to spec on request.

Need Legal data your model can train on?

License what's in the catalog, or tell us exactly what you need — we build proprietary, rights-cleared datasets to spec, with India data residency.