Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)
Native-authored red-team prompts and harm-labelled examples in code-mixed languages — surfacing caste, communal, gendered and code-switch harms English classifiers miss.
- Rights-cleared
- Consent documented
- 🇮🇳 India data residency
- Languages
- Hinglish, Tanglish, English
- Scale
- 24,000 prompts · 12 harm categories · 24000 records
- Inter-annotator agreement
- α = 0.86
- QA pass rate
- 98.20%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.86
- QA pass rate
- 98.20%
Harm labels reach α = 0.86 across a 12-category culturally-grounded taxonomy with a 98.2% QA pass rate and two-pass adjudication on sensitive categories.
Every prompt and label is authored from scratch by native-speaker experts under written agreements — never scraped — with consent and authorship logged per record and all processing within India.
Linguists design a shared harm taxonomy, author adversarial prompts and gold labels, then pass them through two-pass QA with adjudication and a sampling audit, reporting inter-annotator agreement per release.
What's inside each record
A representative schema for Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish). The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| prompt_id | string (uuid) | "rt_91ac…" |
| prompt | string (red-team) | "How would one build a weapon…" (red-team) |
| harm_category | enum | "hate" | "self-harm" |
| expected_behavior | string | "refuse + safe-complete" |
| severity | enum | "high" |
| gold_label | enum | "unsafe" | "safe" |
| language | string (ISO 639) | "hin" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Evaluation / Benchmark Licence
Contact us
Internal evaluation and red-team licence with taxonomy, rubric and scoring guide.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive licensing plus additional languages, harm categories and custom taxonomies to spec.
Talk to salesRelated datasets
-
Safety / EvalSafety / Eval Text
African Low-Resource Red-Team & Safety Eval Set
A native-speaker red-team and safety evaluation set across five major African languages — built to surface harms that English-only guardrails miss.
- Languages:
- Hausa, Yoruba, Amharic, Swahili, Zulu
- Size:
- 18,400 prompts
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
South Asian Instruction-Following Eval (Rare Languages)
An instruction-following and reasoning evaluation suite for four under-served South Asian languages, with human reference answers and rubric-based scoring.
- Languages:
- Maithili, Bhojpuri, Santali, Sindhi
- Size:
- 9,600 instruction–response pairs
- IAA:
- 0.84
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Audio / Speech
African-Language Clinical Speech Eval Set
A held-out evaluation set of accented clinical and everyday speech in two African languages — for benchmarking ASR and speech-LLM robustness.
- Languages:
- Swahili, Yoruba
- Size:
- 40 hrs · 600 speakers
- IAA:
- 0.85
Rights-clearedLicence from
Contact us
Explore more Safety / Eval data
Ready to license Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.