African languages are missing from AI
Large language models learn from what has been written down and digitised. African and indigenous languages are overwhelmingly absent from that record — so the systems now organising education, commerce and public life cannot speak them. This is a data problem, which means it is a fixable one.
- of the world's languages sit in the lowest resource class — effectively invisible to language technology
- 88%of the world's languages sit in the lowest resource class — effectively invisible to language technologyJoshi et al. (2020), ACL
- of the text in many large pretraining corpora is in African languages, combined
- < 0.1%of the text in many large pretraining corpora is in African languages, combinedAfrican-language NLP surveys
- of African languages have no basic digitised texts at all
- 92%of African languages have no basic digitised texts at allAfrican-language NLP surveys
- lack any annotated dataset for fundamental language tasks
- 97%lack any annotated dataset for fundamental language tasksAfrican-language NLP surveys
Nine in ten languages have almost nothing
Researchers grade languages by how much usable data and tooling exists for them. The distribution is not a gentle slope — it is a cliff, and almost every language is at the bottom of it.
2,485 languages sorted into six resource classes, from class 0 — almost no data, almost no tools — up to class 5. Nearly nine in ten sit at the bottom.
Class 0 — The Left-Behinds
Exceptionally limited resources. Rarely considered in language technologies at all.
The web that models train on is a handful of languages
Measured consistently in Common Crawl tokens, the concentration is extreme — and it is the single biggest reason models perform the way they do.
English
532B tokens
Around 45–46% of all documents in the corpus.
Russian
101B tokens
Chinese
92B tokens
Source: Common Crawl corpus analyses. Shares vary by snapshot and by whether language is measured in documents or tokens.
Why it keeps getting worse, not better
Each step causes the next, and the fourth feeds the first. Left alone, the loop simply keeps turning.
…and step 4 feeds step 1. No new text keeps the language low-resource.
1. Scarce data
African languages are overwhelmingly absent from the written, digitised record that models learn from.
Breaking the loop takes data that is deliberately created, high quality and openly licensed. It does not appear on its own.
This is not only a technology problem
When a language is missing from the digital world, its speakers are pushed to the margins of it. That is why language work is development work.
Exclusion from services AI now mediates
Translation, search, dictation, customer support, triage, benefits applications. As these move behind AI, speakers of unsupported languages are pushed to the margins of systems they are entitled to use.
Tools that fail quietly
A model that has barely seen a language does not refuse — it guesses. Confident, fluent, wrong output is more dangerous in a health or legal context than no output at all.
Data assembled elsewhere, on other terms
If the datasets that finally teach machines to speak African languages are built abroad, the terms of access to them are set abroad too.
Scraping without consent or credit
Where open, consented data does not exist, what gets used instead is whatever can be taken — without attribution to the communities whose knowledge it is.
What actually fixes this
Breaking the loop requires data that is deliberately created, high quality and openly licensed. It does not appear on its own, and it cannot be scraped into existence — there is nothing there to scrape.
The only way to create it is to go and get it, in person, from the people who still speak these languages: structured vocabulary and definitions on one side, natural speech on the other, both transcribed and reviewed by native speakers. Paired audio and transcript is the format speech technology actually needs, and it is the single biggest gap between what exists today and what a model can be trained on.
That is what CIALR builds — openly, with consent, with attribution, and free for anyone to use, including the companies that pay us for scale and curation.
A language that AI cannot speak is a language its speakers get locked out of
Sources
- Joshi et al. (2020), ACL — Joshi, Santy, Budhiraja, Bali & Choudhury, “The State and Fate of Linguistic Diversity and Inclusion in the NLP World”, Proceedings of ACL 2020.
- Common Crawl corpus analyses — Published analyses of the Common Crawl corpus; shares vary by snapshot and by whether language is measured in documents or tokens.
- African-language NLP surveys — Surveys of digital text and annotated-dataset availability for African languages, and of dataset usage in NLP research papers 2015–2020.
Figures are reproduced from these published sources and are not estimates of our own. Where a source gives a range or a floor rather than a single number, we say so rather than picking a figure.