Skip to main content
Safety / Eval Audio / Speech 40 hrs · 600 speakers

African-Language Clinical Speech Eval Set

A held-out evaluation set of accented clinical and everyday speech in two African languages — for benchmarking ASR and speech-LLM robustness.

  • Rights-cleared
  • Consent documented
Rights-cleared multilingual dataset catalog illustrating the African-Language Clinical Speech Eval Set dataset
Languages
Swahili, Yoruba
Scale
40 hrs · 600 speakers · 600 records
Inter-annotator agreement
α = 0.85
QA pass rate
97.60%
Quality & methodology metrics

Measured, audited, reproducible

IAA / Krippendorff α
0.85
QA pass rate
97.60%
Word error rate
10.20%

Transcripts are produced by a two-pass native-speaker workflow with a held-out reference WER of 10.2%; segment and speaker labels reach α = 0.85 with a 97.6% QA pass rate.

Provenance & consent

Audio is collected from consenting native speakers under fair-pay contributor agreements granting commercial reuse, with consented speaker demographics and a consent reference on every recording.

Methodology

Field teams capture accented speech across settings, then segment, diarise and transcribe with native speakers; a second reviewer corrects transcripts and a sampling audit verifies WER before the held-out split is frozen for evaluation.

Data preview

What's inside each record

A representative schema for African-Language Clinical Speech Eval Set. The full data card ships the complete field dictionary, value ranges and annotation rubric.

Field Type Example
prompt_id string (uuid) "rt_91ac…"
prompt string (red-team) "Jinsi ya kutengeneza silaha…"
harm_category enum "hate" | "self-harm"
expected_behavior string "refuse + safe-complete"
severity enum "high"
gold_label enum "unsafe" | "safe"
language string (ISO 639) "swa"
annotator_id string (hashed) "anr_7f3…"
qa_status enum "passed"
consent_ref string "cns_2024_…"

Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.

Licensing

License tiers

Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.

Evaluation Licence

Contact us

Held-out evaluation licence with reference transcripts and a scoring guide.

Request access

Enterprise / Sovereign Exclusive

Contact us

Exclusive licensing plus additional African languages and acoustic conditions to spec.

Talk to sales

Related datasets

  • Safety / Eval
    Safety / Eval Text

    African Low-Resource Red-Team & Safety Eval Set

    A native-speaker red-team and safety evaluation set across five major African languages — built to surface harms that English-only guardrails miss.

    Languages:
    Hausa, Yoruba, Amharic, Swahili, Zulu
    Size:
    18,400 prompts
    IAA:
    0.87

    Licence from

    Contact us

    Rights-cleared
  • Safety / Eval
    Safety / Eval Text 🇮🇳 Sovereign

    Code-Mixed Safety & Red-Team Set (Hinglish, Tanglish)

    Native-authored red-team prompts and harm-labelled examples in code-mixed languages — surfacing caste, communal, gendered and code-switch harms English classifiers miss.

    Languages:
    Hinglish, Tanglish, English
    Size:
    24,000 prompts · 12 harm categories
    IAA:
    0.86

    Licence from

    Contact us

    Rights-cleared
  • Safety / Eval
    Safety / Eval Text 🇮🇳 Sovereign

    South Asian Instruction-Following Eval (Rare Languages)

    An instruction-following and reasoning evaluation suite for four under-served South Asian languages, with human reference answers and rubric-based scoring.

    Languages:
    Maithili, Bhojpuri, Santali, Sindhi
    Size:
    9,600 instruction–response pairs
    IAA:
    0.84

    Licence from

    Contact us

    Rights-cleared

See all Safety / Eval datasets

Ready to license African-Language Clinical Speech Eval Set?

A senior data PM will scope access, residency, and licensing terms and respond within one business day.