African-Language Clinical Speech Eval Set
A held-out evaluation set of accented clinical and everyday speech in two African languages — for benchmarking ASR and speech-LLM robustness.
- Rights-cleared
- Consent documented
- Languages
- Swahili, Yoruba
- Scale
- 40 hrs · 600 speakers · 600 records
- Inter-annotator agreement
- α = 0.85
- QA pass rate
- 97.60%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.85
- QA pass rate
- 97.60%
- Word error rate
- 10.20%
Transcripts are produced by a two-pass native-speaker workflow with a held-out reference WER of 10.2%; segment and speaker labels reach α = 0.85 with a 97.6% QA pass rate.
Audio is collected from consenting native speakers under fair-pay contributor agreements granting commercial reuse, with consented speaker demographics and a consent reference on every recording.
Field teams capture accented speech across settings, then segment, diarise and transcribe with native speakers; a second reviewer corrects transcripts and a sampling audit verifies WER before the held-out split is frozen for evaluation.
What's inside each record
A representative schema for African-Language Clinical Speech Eval Set. The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| prompt_id | string (uuid) | "rt_91ac…" |
| prompt | string (red-team) | "Jinsi ya kutengeneza silaha…" |
| harm_category | enum | "hate" | "self-harm" |
| expected_behavior | string | "refuse + safe-complete" |
| severity | enum | "high" |
| gold_label | enum | "unsafe" | "safe" |
| language | string (ISO 639) | "swa" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Evaluation Licence
Contact us
Held-out evaluation licence with reference transcripts and a scoring guide.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive licensing plus additional African languages and acoustic conditions to spec.
Talk to salesRelated datasets
-
Safety / EvalSafety / Eval Text
African Low-Resource Red-Team & Safety Eval Set
A native-speaker red-team and safety evaluation set across five major African languages — built to surface harms that English-only guardrails miss.
- Languages:
- Hausa, Yoruba, Amharic, Swahili, Zulu
- Size:
- 18,400 prompts
- IAA:
- 0.87
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)
Native-authored red-team prompts and harm-labelled examples in code-mixed languages — surfacing caste, communal, gendered and code-switch harms English classifiers miss.
- Languages:
- Hinglish, Tanglish, English
- Size:
- 24,000 prompts · 12 harm categories
- IAA:
- 0.86
Rights-clearedLicence from
Contact us
-
Safety / EvalSafety / Eval Text 🇮🇳 Sovereign
South Asian Instruction-Following Eval (Rare Languages)
An instruction-following and reasoning evaluation suite for four under-served South Asian languages, with human reference answers and rubric-based scoring.
- Languages:
- Maithili, Bhojpuri, Santali, Sindhi
- Size:
- 9,600 instruction–response pairs
- IAA:
- 0.84
Rights-clearedLicence from
Contact us
Explore more Safety / Eval data
Ready to license African-Language Clinical Speech Eval Set?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.