Project Proposal 2026 · SDG Action

Building the largest open dataset of African indigenous languages

Open infrastructure, built by the people who speak these languages, so that they survive both in daily life and in the technologies that shape it.

90,000+language entries recorded and verifiedTerms, definitions, translations, proverbs and idiomatic expressions — each one reviewed by a native speaker before it is published.
44languages documentedAcross five African regions
44countries whose languages are represented
50+countries where our contributors work
5 of 17UN Sustainable Development Goals advanced
Summary

A third of the world's languages, almost none of them online

Africa is home to roughly a third of the world's languages. Almost none of them are meaningfully present in the digital systems that now organise education, commerce, administration and culture. A child today learns to read more easily in English or French than in the language their grandmother speaks. That asymmetry is not neutral — it is how languages die.

The Center of Indigenous & African Language Research (CIALR), the project Nzonza Foundation runs, is an open platform and open dataset built to close that gap. We collect, structure, verify and publish African language data — words, terms, proverbs, definitions, translations and native-speaker audio — under an open licence, through a network of volunteers, native speakers and language experts across more than 50 African countries.

The work is deliberately unglamorous: going to the communities where a language is still spoken, sitting with the people who speak it, recording them, and turning what they give us into data that a school, a researcher or a language model can actually use. Nobody else is going to do this. That is the whole premise of the project.

The problem

Why this is urgent

Several languages already in our dataset sit inside a one-generation window. These are precisely the languages no commercial actor will ever prioritise.

Languages are disappearing faster than they are being recorded

Urbanisation, migration and education systems built around colonial languages have broken transmission. Parents who speak an indigenous language raise children who understand it but do not speak it; those children raise children who do neither. Once a language stops being passed on at home, the window for documenting it is about one generation wide.

African languages are functionally invisible to AI

Models learn from what has been written down and digitised. Scarce data produces poor performance; poor performance means no products get built; no products means no new digital text; no new text keeps the language low-resource. Breaking that loop requires data that is deliberately created, high quality and openly licensed.

The tools for learning simply do not exist

For most African languages there is no usable dictionary, no pronunciation reference, no graded reader, no spellchecker, no keyboard layout. A diaspora child, a teacher, a nurse posted to a new region — none of them has anywhere to start.

The digital divide compounds all of it

Many communities lack the devices, connectivity and skills to take part in the digital economy at all. AI is arriving in the Global South as something built elsewhere, in other languages, for other users. Unless communities are equipped to use it — and to shape it — it widens inequality rather than closing it.

Why this has to be built from within Africa

Every major technology company now states that it wants African and indigenous languages represented in its models. The intention is real. The obstacle is that the data does not exist, and the only way to create it is to go and get it, in person, from the people who still speak these languages.

No large technology company is going to send a team into a village in the Niari valley to sit with an elder for three days. Nor should anyone expect it to. Documenting African languages is African work: ours to do, ours to own, and ours to benefit from.

Africa is not waiting for information. It is waiting for people who decide to take their destiny into their own hands and act on it.

If the datasets that finally teach machines to speak African languages are assembled abroad, the terms of access to them will be set abroad too. Built here instead, the same work becomes infrastructure the continent controls — and a source of revenue that flows back to the communities that hold the knowledge.

What CIALR is

Three things working together

A dataset

Structured lexical data across 44 languages: terms, definitions, translations, usage, proverbs and idiomatic expressions, with linguistic metadata and native-speaker pronunciation audio.

A platform

Live software for contributing, reviewing, searching and retrieving language data. Not a static archive but a working system, used daily by contributors and open to developers.

A community

A distributed network of volunteers, native speakers, linguists and language experts across more than 50 African countries and the diaspora — the people who speak these languages, not outsiders extracting from them.

Why open matters here

Closed language datasets have been built before. They tend to die with the grant that funded them, or sit behind a licence no African developer can afford. CIALR is open by design — so a teacher in Malawi can build a classroom tool without asking permission, a researcher can verify and correct the data rather than take it on trust, an AI lab can train on African language data legitimately instead of scraping it, and the work outlives any single institution, funder or founder.

Method

How the data gets made

The dataset is the visible part of the project. The method behind it is what makes the dataset possible — and it is where funding has the most direct effect.

  1. 01

    Going to where the language lives

    Regional coordinators travel into communities where a language is still in daily use and work through local structures — elders, teachers, associations, radio stations, churches. Elders are the priority: in fragile languages they are frequently the last speakers with full command of vocabulary, idiom and register. Participation is voluntary, informed and credited.

  2. 02

    Two kinds of recording

    Structured elicitation — vocabulary, translations, definitions, proverbs — term by term with a native speaker, which is what makes a dictionary or spellchecker possible. And natural speech: people talking freely about cooking, farming, trading, storytelling, which carries the tone, prosody, register and idiom a wordlist can never hold, and which is what speech technology actually requires.

  3. 03

    From audio to text

    Audio that has never been transcribed is an archive, not a dataset. We are building a transcription pipeline — tooling, conventions and trained local transcribers — so recordings become time-aligned text at scale. Paired audio and transcript is the single biggest gap between what exists today and what a model can train on.

  4. 04

    Writing systems where none exist

    Many languages in scope have never been written down in any settled way. Working with linguists, we adapt conventions from related languages — much of Bantu, for instance — rather than inventing from scratch, then test them with speakers and have the community ratify them.

  5. 05

    Review before publication

    Every entry is reviewed by native speakers and, where the language warrants it, by a linguist, before publication. Contributors and recordists are named. Communities can challenge and correct any entry about their own language.

Current coverage

44 languages across all five African regions

Deliberate representation of both major lingua francas and small, endangered languages — from Swahili and Hausa with over 100 million speakers, to languages with fewer than 30,000.

RegionLanguages
East Africa13
Central Africa12
Southern Africa11
West Africa10
North Africa2
Every language in the dataset, by speaker population

One dot per language, on a logarithmic scale. The dataset deliberately spans five orders of magnitude — from lingua francas with hundreds of millions of speakers to languages with a few thousand.

Fragile — under 300,000 speakers (8)Other languages (36)
1K10K100K1M10M100M300,000 — fragility thresholdAncient Egyptian — < 1,000 speakers (fragile)Ge'ez — < 1,000 speakers (fragile)Phuthi — ~20,000 speakers (fragile)Vili — ~30,000 speakers (fragile)Hamer-Banna — ~74,000 speakers (fragile)Bangi — ~120,000 speakers (fragile)Mboshi — ~150,000 speakers (fragile)Ghomala' — ~260,000 speakers (fragile)Fang — ~1M speakersLari — ~1M speakersTooro — ~1.3M speakersGungbe — ~1.5M speakersNorthern Ndebele — ~1.5M speakersWolaytta — ~2.4M speakersSouthern Ndebele — ~2.5M speakersAfar — ~2.6M speakersYao — ~3M speakersSwazi — ~4M speakersNupe — ~4.5M speakersKikuyu — ~5M speakersEwe — ~6M speakersKituba — ~6M speakersTshiluba — ~6M speakersKikongo — ~7M speakersXhosa — ~8M speakersAfrikaans — ~8.3M speakersTigrinya — ~9M speakersCoptic — ~10M speakersWolof — ~10M speakersChichewa — ~12M speakersZulu — ~13.5M speakersBambara — ~14M speakersShona — ~14M speakersSotho — ~14M speakersAkan — ~20M speakersKinyarwanda — ~22M speakersSomali — ~28M speakersFula — ~40M speakersLingala — ~40M speakersOromo — ~45M speakersYoruba — ~50M speakersAmharic — ~60M speakersHausa — ~150M speakersSwahili — ~200M speakersAncient EgyptianSwahili

Speaker figures are estimates including second-language speakers and vary by source. Hover a dot for the language; every figure is listed in the inventory below.

Several languages, including Fula and Wolof, span more than one region and are counted in each. Families represented: Niger-Congo (Bantu, Kwa, Benue-Congo, Atlantic), Afroasiatic (Cushitic, Semitic, Omotic, Egyptian), Indo-European, and Bantu-based creoles.

Our SDG action

Technology + Language = Opportunity

5 of 17GOALS ADVANCED

Language is the medium through which people learn, access services, take part in public life and pass on what they know. When a language is missing from the digital world, its speakers are pushed to its margins. That is why language work is development work — and why a single project can move five Goals at once.

This work sits within the UN International Decade of Indigenous Languages (2022–2032), led by UNESCO.

SDG 4 — Quality Education

Primary

Targets 4.4 · 4.5 · 4.7

Learning resources in local languages, and digital skills.

Open dictionaries, graded readers and pronunciation trainers make mother-tongue learning possible; training for youth, educators and local teams builds the digital and AI skills needed to use them.

SDG 10 — Reduced Inequalities

Target 10.2

More access to technology for underserved communities.

When AI systems understand African languages, speakers of those languages stop being excluded from the services, information and markets that AI increasingly mediates.

SDG 11 — Sustainable Cities & Communities

Target 11.4

Preserving cultural heritage and local knowledge.

Every verified recording of an elder, a proverb or a story is a permanent safeguard of intangible heritage — especially for fragile languages with fewer than 300,000 speakers.

SDG 16 — Peace, Justice & Strong Institutions

Targets 16.7 · 16.10

Building inclusive and transparent communities.

Open, community-verified, correctable data guarantees public access to information and gives communities a direct say over how their own language is recorded and used.

SDG 17 — Partnerships for the Goals

Targets 17.6 · 17.8 · 17.17

Working together for greater impact.

CIALR connects communities, universities, ministries, UN bodies and technology companies around shared open infrastructure — cooperation no single actor could achieve alone.

Roadmap

Where we go from here

Phase 1

Foundation

Complete

Platform built and deployed. More than 90,000 entries across 44 languages. Contributor network established in over 50 countries. Audio pipeline operational.

Phase 2

Depth and quality

0–12 months

Extend audio coverage across all existing languages. Stand up the transcription pipeline and train transcribers. Launch training and capacity building for youth, educators and local teams. Standardise metadata and orthographies. Open a public API and structured exports.

Phase 3

Breadth

12–24 months

Scale from 44 languages toward 100 and beyond, prioritising endangered and undocumented languages in North Africa, the Sahel, the Horn and the smaller Central African clusters. Establish regional coordinators and field teams. Begin orthography development for unwritten languages.

Phase 4

Applications

18–36 months

Release reference learning tools built on the dataset: open dictionaries, pronunciation trainers, graded readers and the first version of the learning app. Publish benchmark datasets for African-language AI evaluation. Formal partnerships with universities, ministries and research groups.

Governance and ethics

Language data is community property, not a commodity

Community-sourced and community-verified

Contributions come from native speakers; entries are reviewed by speakers and experts before publication.

Attribution and consent

Contributors and recordists are credited and take part voluntarily, with a clear understanding of how the data will be used and licensed.

Open licensing

Data is published under terms that guarantee permanent public access and prevent enclosure by any single commercial actor.

African-led

Governance, direction and the majority of contribution come from the continent and its diaspora, not from external institutions studying it.

Correctability

Communities can challenge and correct entries about their own language. The dataset is a living record, not a colonial-era vocabulary list.

Return to source

Where the dataset generates commercial revenue, a defined share is committed back to fieldwork and to the communities whose knowledge made it possible.

How the work sustains itself

Fieldwork costs money — travel, equipment, stipends, transcription, review. Grants alone will not carry documentation at continental scale for the twenty years it needs. Nzonza therefore runs on two arms, deliberately separated.

The non-profit arm

Documentation, community fieldwork, training and capacity building, the open dataset and its permanent public archive. Funded by institutional support, grants and donations. Everything it produces is published openly and stays open. This is the part of the project that answers to future generations rather than to a market.

The commercial arm

Two revenue lines built on top of the open core rather than in place of it: data services for AI labs that need African language coverage with verified provenance and consent, and learning products — a language-learning app, dictionaries, phrasebooks, proverb collections.

The line between them

Commercial revenue exists to fund fieldwork, not to close the archive. The open dataset remains open and free to all users, including the companies that pay us. No exclusive rights are sold to anyone. What a commercial partner pays for is scale, speed, curation and legal certainty — never the right to keep the data from anybody else.

What we are asking for

Three tracks for partners

A partner may engage with one track or with several.

Track A

Institutional endorsement

Formal recognition from UNESCO and comparable cultural, educational and linguistic bodies, and recognition of CIALR as an SDG action.

Requested: Letters of support; affiliation or observer status; inclusion in language-preservation programmes and the International Decade of Indigenous Languages; access to institutional networks of linguists and archives.

Track B

Funding

Sustained funding to move from volunteer capacity to professional throughput.

Requested: Regional field coordinators and audio collection stipends; transcription capacity; training for youth, educators and local teams; expert linguistic review; engineering for the platform, API and data infrastructure; recording equipment for low-connectivity regions; long-term archival hosting.

Track C

Technical and corporate partnership

Partnership with technology companies, AI labs and research institutions.

Requested: Compute and storage credits; engineering time; speech and NLP expertise; AI tools and platform access for community training; joint benchmark and evaluation work; commitments to include CIALR data in training and evaluation for African languages.

A UNESCO support application has been submitted. To discuss any of these tracks, email info@nzonza.org.

Appendix

Language inventory

All 44 languages currently in the dataset. 8 are marked fragile — fewer than 300,000 speakers, and a documentation priority.

LanguageNative nameFamilySpeakersRegion
AfarAfarafAfroasiatic · Cushitic~2.6MEast
AfrikaansAfrikaansIndo-European~8.3MSouthern
AkanAkan kasaNiger-Congo · Kwa~20MWest
AmharicAmarəññaAfroasiatic · Semitic~60MEast
Ancient EgyptianFragiler n km.tAfroasiatic · Egyptian< 1,000North
BambaraBamanankanNiger-Congo~14MWest
BangiFragileBobangiNiger-Congo · Bantu~120,000Central
ChichewaChichewa, NyanjaNiger-Congo · Bantu~12MSouthern
CopticCopticAfroasiatic · Egyptian~10MNorth
EweEʋegbeNiger-Congo~6MWest
Fang—Niger-Congo · Bantu (A75)~1MCentral
FulaFulfulde, PulaarNiger-Congo~40MWest / Central / East
Ge'ezFragileGəʼəzAfroasiatic · Semitic< 1,000East
Ghomala'FragileGhɔmáláʼNiger-Congo~260,000Central
GungbeGungbeNiger-Congo~1.5MWest
Hamer-BannaFragileHamer-BannaAfroasiatic · Omotic~74,000East
HausaHausaAfroasiatic · Chadic~150MWest
Kikongo—Niger-Congo · Bantu (H10)~7MCentral
KikuyuGĩkũyũNiger-Congo · Bantu~5MEast
KinyarwandaKinyarwanda, KirundiNiger-Congo · Bantu~22MEast
Kituba—Bantu-based creole~6MCentral
Lari—Niger-Congo · Bantu (Kongo)~1MCentral
Lingala—Niger-Congo · Bantu (C30)~40MCentral
MboshiFragile—Niger-Congo · Bantu (C25)~150,000Central
Northern NdebeleisiNdebeleNiger-Congo · Bantu (Nguni)~1.5MSouthern
NupeNupeNiger-Congo · Benue-Congo~4.5MWest
OromoAfaan OromooAfroasiatic · Cushitic~45MEast
PhuthiFragileSíphùthìNiger-Congo · Bantu~20,000Southern
ShonachiShonaNiger-Congo · Bantu~14MSouthern
SomaliSomaliAfroasiatic · Cushitic~28MEast
SothoSesothoNiger-Congo · Bantu~14MSouthern
Southern NdebeleisiNdebeleNiger-Congo · Bantu~2.5MSouthern
SwahiliKiswahiliNiger-Congo · Bantu (G40)~200MEast
SwazisiSwatiNiger-Congo · Bantu~4MSouthern
TigrinyaTigrinyaAfroasiatic · Ethio-Semitic~9MEast
TooroRutooroNiger-Congo · Bantu~1.3MEast
Tshiluba—Niger-Congo · Bantu (L31)~6MCentral
ViliFragile—Niger-Congo · Bantu (H12)~30,000Central
WolayttaWolayttatto DoonaaAfroasiatic · Omotic~2.4MEast
WolofWolof làkkNiger-Congo~10MWest / Central
XhosaisiXhosaNiger-Congo · Bantu~8MSouthern
YaochiYaoNiger-Congo · Bantu~3MSouthern
YorubaYorùbáNiger-Congo~50MWest
ZuluisiZuluNiger-Congo · Bantu~13.5MSouthern

Speaker figures are estimates including second-language speakers and vary by source.

Leadership and contact

Amstrong Nouni

Co-founder & Managing Director

Data architect, bringing enterprise-grade experience in data modelling, pipeline design and system architecture to a domain where that rigour is rare. Member of the United Nations youth community, working at the intersection of technology, cultural preservation and development.

CIALR is an independent open-source initiative of Nzonza Foundation. Professional affiliations are listed as background and do not represent institutional endorsement.

The project is supported by a distributed network of volunteers, native speakers, linguists and language experts across more than 50 African countries and the diaspora.

A language is not a list of words. It is a way of seeing, and every one that disappears takes a way of seeing with it. CIALR exists so that fewer of them do.

See exactly how donations are used

Donate