Languages
Languages big models miss
Every language here has millions of speakers and too little of the right data. Tamil is live; tell us which one you need next.
- தமிழ்Live
Tamil · 80M+ speakers
Spoken Tamil differs sharply from written Tamil, and dialects from Jaffna to Chennai barely appear online.
Visit Tamil → - සිංහලNext
Sinhala · 17M+ speakers
Written and spoken Sinhala are almost two different registers, and models trained on formal text sound stiff in conversation.
Join the waitlist → - Èdè YorùbáPlanned
Yoruba · 45M+ speakers
A tonal language where meaning rides on diacritics that most web text leaves out.
Join the waitlist → - Harshen HausaPlanned
Hausa · 80M+ speakers
One of Africa's most spoken languages, yet with little web text, written in both Latin and Arabic (Ajami) script.
Join the waitlist → - አማርኛPlanned
Amharic · 55M+ speakers
Its own script (Ge'ez) and rich word forms get split into many pieces by tokenizers built for English.
Join the waitlist → - Asụsụ IgboPlanned
Igbo · 30M+ speakers
Tonal, with many dialects, and the little text online is often missing its tone marks.
Join the waitlist → - ភាសាខ្មែរPlanned
Khmer · 17M+ speakers
Written without spaces between words, so standard tokenizers and search struggle from the first step.
Join the waitlist → - پښتوPlanned
Pashto · 40M+ speakers
Strong regional dialects and very little high-quality digital text.
Join the waitlist → - Kreyòl ayisyenPlanned
Haitian Creole · 12M+ speakers
Spoken by nearly everyone in Haiti, while French dominates formal writing, so little text reflects how people talk.
Join the waitlist → - Runa SimiPlanned
Quechua · 8M+ speakers
A family of mostly oral varieties with little standardized text to learn from.
Join the waitlist →
Early access
Build with Language Brain.
We are opening Language Brain to a small group of teams building agents for these languages, and to speakers who want to help.