Skip to main content
Safety / Eval Text 24,000 prompts · 12 harm categories

Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)

Native-authored red-team prompts and harm-labelled examples in code-mixed languages — surfacing caste, communal, gendered and code-switch harms English classifiers miss.

  • Rights-cleared
  • Consent documented
  • 🇮🇳 India data residency
Rights-cleared multilingual dataset catalog illustrating the Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish) dataset
Languages
Hinglish, Tanglish, English
Scale
24,000 prompts · 12 harm categories · 24000 records
Inter-annotator agreement
α = 0.86
QA pass rate
98.20%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.86
QA pass rate
98.20%

Harm labels reach α = 0.86 across a 12-category culturally-grounded taxonomy with a 98.2% QA pass rate and two-pass adjudication on sensitive categories.

Provenance & consent

Every prompt and label is authored from scratch by native-speaker experts under written agreements — never scraped — with consent and authorship logged per record and all processing within India.

Methodology

Linguists design a shared harm taxonomy, author adversarial prompts and gold labels, then pass them through two-pass QA with adjudication and a sampling audit, reporting inter-annotator agreement per release.

Data preview

What's inside each record

A representative schema for Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish). The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
prompt_id string (uuid) "rt_91ac…"
prompt string (red-team) "How would one build a weapon…" (red-team)
harm_category enum "hate" | "self-harm"
expected_behavior string "refuse + safe-complete"
severity enum "high"
gold_label enum "unsafe" | "safe"
language string (ISO 639) "hin"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Evaluation / Benchmark Licence

Contact us

Internal evaluation and red-team licence with taxonomy, rubric and scoring guide.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive licensing plus additional languages, harm categories and custom taxonomies to spec.

Talk to sales

Related datasets

See all Safety / Eval datasets

Ready to license Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.