Inclusive AI

African languages are missing from AI

Large language models learn from what has been written down and digitised. African and indigenous languages are overwhelmingly absent from that record — so the systems now organising education, commerce and public life cannot speak them. This is a data problem, which means it is a fixable one.

of the world's languages sit in the lowest resource class — effectively invisible to language technology
88%of the world's languages sit in the lowest resource class — effectively invisible to language technologyJoshi et al. (2020), ACL
of the text in many large pretraining corpora is in African languages, combined
< 0.1%of the text in many large pretraining corpora is in African languages, combinedAfrican-language NLP surveys
of African languages have no basic digitised texts at all
92%of African languages have no basic digitised texts at allAfrican-language NLP surveys
lack any annotated dataset for fundamental language tasks
97%lack any annotated dataset for fundamental language tasksAfrican-language NLP surveys
The resource cliff

Nine in ten languages have almost nothing

Researchers grade languages by how much usable data and tooling exists for them. The distribution is not a gentle slope — it is a cliff, and almost every language is at the bottom of it.

The world's languages, by how much language technology exists for them

2,485 languages sorted into six resource classes, from class 0 — almost no data, almost no tools — up to class 5. Nearly nine in ten sit at the bottom.

Class 0 — The Left-Behinds

Exceptionally limited resources. Rarely considered in language technologies at all.

Where the text is

The web that models train on is a handful of languages

Measured consistently in Common Crawl tokens, the concentration is extreme — and it is the single biggest reason models perform the way they do.

English

532B tokens

Around 45–46% of all documents in the corpus.

Russian

101B tokens

Chinese

92B tokens

11languages clear 10 billion tokens — out of roughly 7,000 spoken in the world.
27languages clear 1 billion tokens — out of roughly 7,000 spoken in the world.

Source: Common Crawl corpus analyses. Shares vary by snapshot and by whether language is measured in documents or tokens.

The mechanism

Why it keeps getting worse, not better

Why the gap does not close on its own

Each step causes the next, and the fourth feeds the first. Left alone, the loop simply keeps turning.

…and step 4 feeds step 1. No new text keeps the language low-resource.

1. Scarce data

African languages are overwhelmingly absent from the written, digitised record that models learn from.

Breaking the loop takes data that is deliberately created, high quality and openly licensed. It does not appear on its own.

What it costs

This is not only a technology problem

When a language is missing from the digital world, its speakers are pushed to the margins of it. That is why language work is development work.

Exclusion from services AI now mediates

Translation, search, dictation, customer support, triage, benefits applications. As these move behind AI, speakers of unsupported languages are pushed to the margins of systems they are entitled to use.

Tools that fail quietly

A model that has barely seen a language does not refuse — it guesses. Confident, fluent, wrong output is more dangerous in a health or legal context than no output at all.

Data assembled elsewhere, on other terms

If the datasets that finally teach machines to speak African languages are built abroad, the terms of access to them are set abroad too.

Scraping without consent or credit

Where open, consented data does not exist, what gets used instead is whatever can be taken — without attribution to the communities whose knowledge it is.

What actually fixes this

Breaking the loop requires data that is deliberately created, high quality and openly licensed. It does not appear on its own, and it cannot be scraped into existence — there is nothing there to scrape.

The only way to create it is to go and get it, in person, from the people who still speak these languages: structured vocabulary and definitions on one side, natural speech on the other, both transcribed and reviewed by native speakers. Paired audio and transcript is the format speech technology actually needs, and it is the single biggest gap between what exists today and what a model can be trained on.

That is what CIALR builds — openly, with consent, with attribution, and free for anyone to use, including the companies that pay us for scale and curation.

A language that AI cannot speak is a language its speakers get locked out of

Sources

  • Joshi et al. (2020), ACL — Joshi, Santy, Budhiraja, Bali & Choudhury, “The State and Fate of Linguistic Diversity and Inclusion in the NLP World”, Proceedings of ACL 2020.
  • Common Crawl corpus analyses — Published analyses of the Common Crawl corpus; shares vary by snapshot and by whether language is measured in documents or tokens.
  • African-language NLP surveys — Surveys of digital text and annotated-dataset availability for African languages, and of dataset usage in NLP research papers 2015–2020.

Figures are reproduced from these published sources and are not estimates of our own. Where a source gives a range or a floor rather than a single number, we say so rather than picking a figure.

Donate