Help us reach 1,000,000 words — Get involved

Center of Indigenous & African Language Research

A project of Nzonza Foundation

Connecting people, preserving culture, and empowering communities through the power of language, rich traditions, shared roots, and collective wisdom.

90,000+language entries recorded and verifiedTerms, definitions, translations, proverbs and idiomatic expressions — each one reviewed by a native speaker before it is published.
44languages documentedAcross five African regions
44countries whose languages are represented
50+countries where our contributors work
5 of 17UN Sustainable Development Goals advanced
We're a non-profit

Help us save languages before they disappear

Nzonza Foundation relies entirely on donations to keep this site running and to fund the fieldwork behind CIALR. Hundreds of African and native languages are at risk of extinction — we can't preserve them without your help. If you'd like to support the mission, a donation goes a long way.

Donate Now

Secure checkout via Stripe

See exactly how your donation is used

In development

A learning app is on the way

Nzonza Learn turns the dataset into something you can learn from — lessons in African and indigenous languages, with the voices of the people who speak them. Here is how it is being designed.

Hear when it launches

Designs, not a released app — tap one to see it full size.

From the field

Our latest stories on social media

Recordings, fieldwork and the languages we are documenting — from @nzonza.connect.

  • See everythingAll our recordings and fieldwork on TikTok

Stories open in a player here; nothing is loaded from TikTok until you open one.

Common questions

Frequently asked questions

What the project is, how the data is made, and what happens to the money.

What is Nzonza Foundation, and what is CIALR?
Nzonza Foundation is the non-profit. The Center of Indigenous & African Language Research (CIALR) is the project it runs — an open platform and an open dataset of African indigenous languages. The Foundation is the organisation that receives donations and answers for the work; CIALR is the work itself.
What does CIALR do?
We collect, structure, verify and publish language data — words, definitions, translations, proverbs and native-speaker audio — through a network of volunteers, native speakers and language experts across more than 50 African countries and the diaspora. More than 90,000 entries across 44 languages are published today.
What is actually in the dataset?
Structured lexical data: terms, definitions, translations, usage, proverbs and idiomatic expressions, each with linguistic metadata — family, native name, speaker estimates, and the countries and regions where the language is used — alongside native-speaker pronunciation audio.
Which languages do you cover?
44 languages spanning all five African regions. The range is deliberate: Swahili with roughly 200 million speakers sits in the same dataset as Phuthi with about 20,000. Eight of the languages have fewer than 300,000 speakers and are a documentation priority. The full inventory is on the mission page.
How is the data collected?
Regional coordinators travel into communities where a language is still in daily use and work through local structures — elders, teachers, associations, radio stations. Elders are the priority: in fragile languages they are frequently the last speakers with full command of vocabulary, idiom and register. We record in two modes — structured elicitation of vocabulary and definitions, and natural speech about everyday life — because a wordlist alone cannot carry tone, register or idiom.
Is the data free to use?
Yes. It is published under terms that guarantee permanent public access and prevent enclosure by any single commercial actor. A teacher can build a classroom tool without asking permission, a researcher can verify and correct the data rather than take it on trust, and an AI lab can train on African language data legitimately instead of scraping it.
Do the speakers consent, and are they credited?
Participation is voluntary and informed, and recording only begins once the community understands how the material will be used and licensed. Contributors and recordists are named. Communities can challenge and correct any entry about their own language — the dataset is a living record, not a colonial-era vocabulary list.
How is the project funded, and where do donations go?
Nzonza Foundation runs on two deliberately separated arms. Donations go to the non-profit arm — documentation, community fieldwork, training, and the open dataset with its permanent public archive. A separate commercial arm sells data services and learning products built on top of the open core, to fund fieldwork rather than to close the archive, and a defined share of that revenue returns to the communities whose knowledge made it possible. See exactly what donations pay for.
How can I help?
Donate, contribute a recording in your mother tongue, introduce us to an elder or a speaker community, or partner with us on technology, funding or institutional support. For partnerships and anything not answered here, email info@nzonza.org.
Donate