Building the largest open dataset of African indigenous languages
Open infrastructure, built by the people who speak these languages, so that they survive both in daily life and in the technologies that shape it.
A third of the world's languages, almost none of them online
Africa is home to roughly a third of the world's languages. Almost none of them are meaningfully present in the digital systems that now organise education, commerce, administration and culture. A child today learns to read more easily in English or French than in the language their grandmother speaks. That asymmetry is not neutral — it is how languages die.
The Center of Indigenous & African Language Research (CIALR), the project Nzonza Foundation runs, is an open platform and open dataset built to close that gap. We collect, structure, verify and publish African language data — words, terms, proverbs, definitions, translations and native-speaker audio — under an open licence, through a network of volunteers, native speakers and language experts across more than 50 African countries.
The work is deliberately unglamorous: going to the communities where a language is still spoken, sitting with the people who speak it, recording them, and turning what they give us into data that a school, a researcher or a language model can actually use. Nobody else is going to do this. That is the whole premise of the project.
Why this is urgent
Several languages already in our dataset sit inside a one-generation window. These are precisely the languages no commercial actor will ever prioritise.
Languages are disappearing faster than they are being recorded
Urbanisation, migration and education systems built around colonial languages have broken transmission. Parents who speak an indigenous language raise children who understand it but do not speak it; those children raise children who do neither. Once a language stops being passed on at home, the window for documenting it is about one generation wide.
African languages are functionally invisible to AI
Models learn from what has been written down and digitised. Scarce data produces poor performance; poor performance means no products get built; no products means no new digital text; no new text keeps the language low-resource. Breaking that loop requires data that is deliberately created, high quality and openly licensed.
The tools for learning simply do not exist
For most African languages there is no usable dictionary, no pronunciation reference, no graded reader, no spellchecker, no keyboard layout. A diaspora child, a teacher, a nurse posted to a new region — none of them has anywhere to start.
The digital divide compounds all of it
Many communities lack the devices, connectivity and skills to take part in the digital economy at all. AI is arriving in the Global South as something built elsewhere, in other languages, for other users. Unless communities are equipped to use it — and to shape it — it widens inequality rather than closing it.
Why this has to be built from within Africa
Every major technology company now states that it wants African and indigenous languages represented in its models. The intention is real. The obstacle is that the data does not exist, and the only way to create it is to go and get it, in person, from the people who still speak these languages.
No large technology company is going to send a team into a village in the Niari valley to sit with an elder for three days. Nor should anyone expect it to. Documenting African languages is African work: ours to do, ours to own, and ours to benefit from.
Africa is not waiting for information. It is waiting for people who decide to take their destiny into their own hands and act on it.
If the datasets that finally teach machines to speak African languages are assembled abroad, the terms of access to them will be set abroad too. Built here instead, the same work becomes infrastructure the continent controls — and a source of revenue that flows back to the communities that hold the knowledge.
Three things working together
A dataset
Structured lexical data across 44 languages: terms, definitions, translations, usage, proverbs and idiomatic expressions, with linguistic metadata and native-speaker pronunciation audio.
A platform
Live software for contributing, reviewing, searching and retrieving language data. Not a static archive but a working system, used daily by contributors and open to developers.
A community
A distributed network of volunteers, native speakers, linguists and language experts across more than 50 African countries and the diaspora — the people who speak these languages, not outsiders extracting from them.
Why open matters here
Closed language datasets have been built before. They tend to die with the grant that funded them, or sit behind a licence no African developer can afford. CIALR is open by design — so a teacher in Malawi can build a classroom tool without asking permission, a researcher can verify and correct the data rather than take it on trust, an AI lab can train on African language data legitimately instead of scraping it, and the work outlives any single institution, funder or founder.
How the data gets made
The dataset is the visible part of the project. The method behind it is what makes the dataset possible — and it is where funding has the most direct effect.
- 01
Going to where the language lives
Regional coordinators travel into communities where a language is still in daily use and work through local structures — elders, teachers, associations, radio stations, churches. Elders are the priority: in fragile languages they are frequently the last speakers with full command of vocabulary, idiom and register. Participation is voluntary, informed and credited.
- 02
Two kinds of recording
Structured elicitation — vocabulary, translations, definitions, proverbs — term by term with a native speaker, which is what makes a dictionary or spellchecker possible. And natural speech: people talking freely about cooking, farming, trading, storytelling, which carries the tone, prosody, register and idiom a wordlist can never hold, and which is what speech technology actually requires.
- 03
From audio to text
Audio that has never been transcribed is an archive, not a dataset. We are building a transcription pipeline — tooling, conventions and trained local transcribers — so recordings become time-aligned text at scale. Paired audio and transcript is the single biggest gap between what exists today and what a model can train on.
- 04
Writing systems where none exist
Many languages in scope have never been written down in any settled way. Working with linguists, we adapt conventions from related languages — much of Bantu, for instance — rather than inventing from scratch, then test them with speakers and have the community ratify them.
- 05
Review before publication
Every entry is reviewed by native speakers and, where the language warrants it, by a linguist, before publication. Contributors and recordists are named. Communities can challenge and correct any entry about their own language.
44 languages across all five African regions
Deliberate representation of both major lingua francas and small, endangered languages — from Swahili and Hausa with over 100 million speakers, to languages with fewer than 30,000.
| Region | Languages | Examples |
|---|---|---|
| East Africa | 13 | Swahili, Amharic, Oromo, Somali, Tigrinya, Kinyarwanda, Kikuyu, Afar, Ge'ez |
| Central Africa | 12 | Lingala, Kikongo, Kituba, Tshiluba, Fang, Mboshi, Vili, Lari, Ghomala' |
| Southern Africa | 11 | Zulu, Xhosa, Shona, Sotho, Chichewa, Swazi, Afrikaans, Ndebele, Phuthi |
| West Africa | 10 | Hausa, Yoruba, Fula, Akan, Bambara, Wolof, Ewe, Nupe, Gungbe |
| North Africa | 2 | Coptic, Ancient Egyptian (hieroglyphic) |
One dot per language, on a logarithmic scale. The dataset deliberately spans five orders of magnitude — from lingua francas with hundreds of millions of speakers to languages with a few thousand.
Speaker figures are estimates including second-language speakers and vary by source. Hover a dot for the language; every figure is listed in the inventory below.
Several languages, including Fula and Wolof, span more than one region and are counted in each. Families represented: Niger-Congo (Bantu, Kwa, Benue-Congo, Atlantic), Afroasiatic (Cushitic, Semitic, Omotic, Egyptian), Indo-European, and Bantu-based creoles.
Technology + Language = Opportunity
Language is the medium through which people learn, access services, take part in public life and pass on what they know. When a language is missing from the digital world, its speakers are pushed to its margins. That is why language work is development work — and why a single project can move five Goals at once.
This work sits within the UN International Decade of Indigenous Languages (2022–2032), led by UNESCO.
SDG 4 — Quality Education
PrimaryTargets 4.4 · 4.5 · 4.7
Learning resources in local languages, and digital skills.
Open dictionaries, graded readers and pronunciation trainers make mother-tongue learning possible; training for youth, educators and local teams builds the digital and AI skills needed to use them.
SDG 10 — Reduced Inequalities
Target 10.2
More access to technology for underserved communities.
When AI systems understand African languages, speakers of those languages stop being excluded from the services, information and markets that AI increasingly mediates.
SDG 11 — Sustainable Cities & Communities
Target 11.4
Preserving cultural heritage and local knowledge.
Every verified recording of an elder, a proverb or a story is a permanent safeguard of intangible heritage — especially for fragile languages with fewer than 300,000 speakers.
SDG 16 — Peace, Justice & Strong Institutions
Targets 16.7 · 16.10
Building inclusive and transparent communities.
Open, community-verified, correctable data guarantees public access to information and gives communities a direct say over how their own language is recorded and used.
SDG 17 — Partnerships for the Goals
Targets 17.6 · 17.8 · 17.17
Working together for greater impact.
CIALR connects communities, universities, ministries, UN bodies and technology companies around shared open infrastructure — cooperation no single actor could achieve alone.
Where we go from here
Foundation
CompletePlatform built and deployed. More than 90,000 entries across 44 languages. Contributor network established in over 50 countries. Audio pipeline operational.
Depth and quality
0–12 monthsExtend audio coverage across all existing languages. Stand up the transcription pipeline and train transcribers. Launch training and capacity building for youth, educators and local teams. Standardise metadata and orthographies. Open a public API and structured exports.
Breadth
12–24 monthsScale from 44 languages toward 100 and beyond, prioritising endangered and undocumented languages in North Africa, the Sahel, the Horn and the smaller Central African clusters. Establish regional coordinators and field teams. Begin orthography development for unwritten languages.
Applications
18–36 monthsRelease reference learning tools built on the dataset: open dictionaries, pronunciation trainers, graded readers and the first version of the learning app. Publish benchmark datasets for African-language AI evaluation. Formal partnerships with universities, ministries and research groups.
Language data is community property, not a commodity
Community-sourced and community-verified
Contributions come from native speakers; entries are reviewed by speakers and experts before publication.
Attribution and consent
Contributors and recordists are credited and take part voluntarily, with a clear understanding of how the data will be used and licensed.
Open licensing
Data is published under terms that guarantee permanent public access and prevent enclosure by any single commercial actor.
African-led
Governance, direction and the majority of contribution come from the continent and its diaspora, not from external institutions studying it.
Correctability
Communities can challenge and correct entries about their own language. The dataset is a living record, not a colonial-era vocabulary list.
Return to source
Where the dataset generates commercial revenue, a defined share is committed back to fieldwork and to the communities whose knowledge made it possible.
How the work sustains itself
Fieldwork costs money — travel, equipment, stipends, transcription, review. Grants alone will not carry documentation at continental scale for the twenty years it needs. Nzonza therefore runs on two arms, deliberately separated.
The non-profit arm
Documentation, community fieldwork, training and capacity building, the open dataset and its permanent public archive. Funded by institutional support, grants and donations. Everything it produces is published openly and stays open. This is the part of the project that answers to future generations rather than to a market.
The commercial arm
Two revenue lines built on top of the open core rather than in place of it: data services for AI labs that need African language coverage with verified provenance and consent, and learning products — a language-learning app, dictionaries, phrasebooks, proverb collections.
The line between them
Commercial revenue exists to fund fieldwork, not to close the archive. The open dataset remains open and free to all users, including the companies that pay us. No exclusive rights are sold to anyone. What a commercial partner pays for is scale, speed, curation and legal certainty — never the right to keep the data from anybody else.
Three tracks for partners
A partner may engage with one track or with several.
Institutional endorsement
Formal recognition from UNESCO and comparable cultural, educational and linguistic bodies, and recognition of CIALR as an SDG action.
Requested: Letters of support; affiliation or observer status; inclusion in language-preservation programmes and the International Decade of Indigenous Languages; access to institutional networks of linguists and archives.
Funding
Sustained funding to move from volunteer capacity to professional throughput.
Requested: Regional field coordinators and audio collection stipends; transcription capacity; training for youth, educators and local teams; expert linguistic review; engineering for the platform, API and data infrastructure; recording equipment for low-connectivity regions; long-term archival hosting.
Technical and corporate partnership
Partnership with technology companies, AI labs and research institutions.
Requested: Compute and storage credits; engineering time; speech and NLP expertise; AI tools and platform access for community training; joint benchmark and evaluation work; commitments to include CIALR data in training and evaluation for African languages.
A UNESCO support application has been submitted. To discuss any of these tracks, email info@nzonza.org.
Language inventory
All 44 languages currently in the dataset. 8 are marked fragile — fewer than 300,000 speakers, and a documentation priority.
| Language | Native name | Family | Speakers | Region |
|---|---|---|---|---|
| Afar | Afaraf | Afroasiatic · Cushitic | ~2.6M | East |
| Afrikaans | Afrikaans | Indo-European | ~8.3M | Southern |
| Akan | Akan kasa | Niger-Congo · Kwa | ~20M | West |
| Amharic | Amarəñña | Afroasiatic · Semitic | ~60M | East |
| Ancient EgyptianFragile | r n km.t | Afroasiatic · Egyptian | < 1,000 | North |
| Bambara | Bamanankan | Niger-Congo | ~14M | West |
| BangiFragile | Bobangi | Niger-Congo · Bantu | ~120,000 | Central |
| Chichewa | Chichewa, Nyanja | Niger-Congo · Bantu | ~12M | Southern |
| Coptic | Coptic | Afroasiatic · Egyptian | ~10M | North |
| Ewe | Eʋegbe | Niger-Congo | ~6M | West |
| Fang | — | Niger-Congo · Bantu (A75) | ~1M | Central |
| Fula | Fulfulde, Pulaar | Niger-Congo | ~40M | West / Central / East |
| Ge'ezFragile | Gəʼəz | Afroasiatic · Semitic | < 1,000 | East |
| Ghomala'Fragile | Ghɔmáláʼ | Niger-Congo | ~260,000 | Central |
| Gungbe | Gungbe | Niger-Congo | ~1.5M | West |
| Hamer-BannaFragile | Hamer-Banna | Afroasiatic · Omotic | ~74,000 | East |
| Hausa | Hausa | Afroasiatic · Chadic | ~150M | West |
| Kikongo | — | Niger-Congo · Bantu (H10) | ~7M | Central |
| Kikuyu | Gĩkũyũ | Niger-Congo · Bantu | ~5M | East |
| Kinyarwanda | Kinyarwanda, Kirundi | Niger-Congo · Bantu | ~22M | East |
| Kituba | — | Bantu-based creole | ~6M | Central |
| Lari | — | Niger-Congo · Bantu (Kongo) | ~1M | Central |
| Lingala | — | Niger-Congo · Bantu (C30) | ~40M | Central |
| MboshiFragile | — | Niger-Congo · Bantu (C25) | ~150,000 | Central |
| Northern Ndebele | isiNdebele | Niger-Congo · Bantu (Nguni) | ~1.5M | Southern |
| Nupe | Nupe | Niger-Congo · Benue-Congo | ~4.5M | West |
| Oromo | Afaan Oromoo | Afroasiatic · Cushitic | ~45M | East |
| PhuthiFragile | Síphùthì | Niger-Congo · Bantu | ~20,000 | Southern |
| Shona | chiShona | Niger-Congo · Bantu | ~14M | Southern |
| Somali | Somali | Afroasiatic · Cushitic | ~28M | East |
| Sotho | Sesotho | Niger-Congo · Bantu | ~14M | Southern |
| Southern Ndebele | isiNdebele | Niger-Congo · Bantu | ~2.5M | Southern |
| Swahili | Kiswahili | Niger-Congo · Bantu (G40) | ~200M | East |
| Swazi | siSwati | Niger-Congo · Bantu | ~4M | Southern |
| Tigrinya | Tigrinya | Afroasiatic · Ethio-Semitic | ~9M | East |
| Tooro | Rutooro | Niger-Congo · Bantu | ~1.3M | East |
| Tshiluba | — | Niger-Congo · Bantu (L31) | ~6M | Central |
| ViliFragile | — | Niger-Congo · Bantu (H12) | ~30,000 | Central |
| Wolaytta | Wolayttatto Doonaa | Afroasiatic · Omotic | ~2.4M | East |
| Wolof | Wolof làkk | Niger-Congo | ~10M | West / Central |
| Xhosa | isiXhosa | Niger-Congo · Bantu | ~8M | Southern |
| Yao | chiYao | Niger-Congo · Bantu | ~3M | Southern |
| Yoruba | Yorùbá | Niger-Congo | ~50M | West |
| Zulu | isiZulu | Niger-Congo · Bantu | ~13.5M | Southern |
Speaker figures are estimates including second-language speakers and vary by source.
Leadership and contact
Amstrong Nouni
Co-founder & Managing Director
Data architect, bringing enterprise-grade experience in data modelling, pipeline design and system architecture to a domain where that rigour is rare. Member of the United Nations youth community, working at the intersection of technology, cultural preservation and development.
CIALR is an independent open-source initiative of Nzonza Foundation. Professional affiliations are listed as background and do not represent institutional endorsement.
The project is supported by a distributed network of volunteers, native speakers, linguists and language experts across more than 50 African countries and the diaspora.
A language is not a list of words. It is a way of seeing, and every one that disappears takes a way of seeing with it. CIALR exists so that fewer of them do.