Skip to main content
Low-resource language Ge'ez Ethiopia (Africa)

Amharic AI Training Datasets

Amharic (አማርኛ) spoken in Ethiopia (Africa), written in Ge'ez, is one of the languages we build proprietary, rights-cleared datasets for. It is a low-resource language: usable training data barely exists on the open web, so it has to be collected from real speakers with consent. Every Amharic dataset ships with documented consent, provenance and published quality metrics, and is available with India data residency.

Global multilingual data network spanning scripts and regions, representing coverage for the Amharic AI Training Datasets
Matching datasets

Amharic datasets in the catalog

Showing 2 of 41 datasets.

  • Safety / Eval
    Safety / Eval Text

    African Low-Resource Red-Team & Safety Eval Set

    A native-speaker red-team and safety evaluation set across five major African languages — built to surface harms that English-only guardrails miss.

    Languages:
    Hausa, Yoruba, Amharic, Swahili, Zulu
    Size:
    18,400 prompts
    IAA:
    0.87

    Licence from

    Contact us

    Rights-cleared
  • General
    General Audio / Speech 🇮🇳 Sovereign

    Low-Resource Language ASR & TTS Corpus

    Paired speech-and-text for low-resource languages, built for both ASR and TTS — read prompts plus spontaneous speech, with studio-grade single-speaker sets for voice building.

    Languages:
    Santali, Bodo, Dogri, Maithili + Global-South LRLs
    Size:
    410 hrs · 1,300 speakers · 12 languages
    IAA:
    0.87

    Licence from

    Contact us

    Rights-cleared
FAQ

Common questions

Is the Amharic data rights-cleared and safe to train on?

Yes. Every Amharic record is created or sourced under written contributor agreements granting commercial reuse, with a consent reference and authorship log — never scraped.

Can you collect more Amharic data to spec?

Yes. When the Amharic data you need doesn't exist, we collect and annotate it to your specification across speech, text and multimodal modalities, with an India-residency option.

How is Amharic data quality measured?

Each dataset ships with a data card: inter-annotator agreement, QA pass rate, and (for speech) word error rate, with two-pass QA and adjudication.

Need Amharic data your model can train on?

License what's in the catalog, or tell us exactly what you need — we build proprietary, rights-cleared datasets to spec, with India data residency.