AI Safety & Evaluation Datasets — Multilingual Red-Team & Eval
A model can ace English safety benchmarks and still produce culturally-specific harms in Indic and low-resource languages. Our safety and evaluation datasets are native-speaker authored: red-team prompts, instruction-following eval and cultural-safety sets that surface what translated benchmarks miss.
Safety / Eval datasets in the catalog
Showing 4 of 41 datasets.
-
Safety / EvalSafety / Eval Text
African Low-Resource Red-Team & Safety Eval Set
A native-speaker red-team and safety evaluation set across five major African languages — built to surface harms that English-only guardrails miss.
- Languages:
- Hausa, Yoruba, Amharic, Swahili, Zulu
- Size:
- 18,400 prompts
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)
Native-authored red-team prompts and harm-labelled examples in code-mixed languages — surfacing caste, communal, gendered and code-switch harms English classifiers miss.
- Languages:
- Hinglish, Tanglish, English
- Size:
- 24,000 prompts · 12 harm categories
- IAA:
- 0.86
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
South Asian Instruction-Following Eval (Rare Languages)
An instruction-following and reasoning evaluation suite for four under-served South Asian languages, with human reference answers and rubric-based scoring.
- Languages:
- Maithili, Bhojpuri, Santali, Sindhi
- Size:
- 9,600 instruction–response pairs
- IAA:
- 0.84
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Audio / Speech
African-Language Clinical Speech Eval Set
A held-out evaluation set of accented clinical and everyday speech in two African languages — for benchmarking ASR and speech-LLM robustness.
- Languages:
- Swahili, Yoruba
- Size:
- 40 hrs · 600 speakers
- IAA:
- 0.85
Rights-clearedLicence from
Contact us
Services & industry for Safety / Eval
When the data you need doesn't exist yet, we build it — collection, annotation, alignment and evaluation, all rights-cleared and documented.
Content Moderation & Safety
PII filtering, harmful-content taxonomies and culturally-aware safety pipelines — calibrated for Indian context.
Explore the serviceCultural & Cross-Lingual Evaluation
Evaluation for honorifics, code-mix, idioms, caste-safety and pragmatic correctness — beyond translated MMLU.
Explore the serviceRLHF & Evaluation
Preference data, red-teaming, DPO and culturally-calibrated evaluation — so your model learns the judgement your users expect.
Explore the serviceAI Labs & Frontier Model Builders industry
Multilingual evaluation, RLHF/DPO preference data, safety & red-team sets and custom collection — so your frontier model works beyond English.
Explore the industryCommon questions
- How are the safety and eval datasets built?
Native-speaker linguists author adversarial prompts and gold references from scratch under written agreements — never scraped — mapped to a shared harm taxonomy with two-pass QA and adjudication.
- Do you cover languages without existing benchmarks?
Yes. We author evaluation and red-team sets for low-resource and Indic languages that have no standard benchmark to translate.
Need Safety / Eval data your model can train on?
License what's in the catalog, or tell us exactly what you need — we build proprietary, rights-cleared datasets to spec, with India data residency.