Uyir Lab · Language Brain
AI that speaks the languages big models miss.
Uyir Lab makes AI agents accurate in the languages big models get wrong, the way they are actually spoken.
The gap
Billions of people speak languages AI handles poorly.
Big models learn from the written web, and most of the world does not talk the way the web is written. People speak in dialects, slang, code-switching and cultural shorthand that never made it into training data. So agents misunderstand them, answer in stiff textbook language, or simply get it wrong.
Dialects
The same language sounds different from one city to the next. Models usually learn only the standard form.
Slang and code-switching
Real conversations mix languages and scripts mid-sentence. Training data rarely does.
Cultural context
Kinship terms, honorifics, food, faith and humour carry meaning a literal translation misses.
Our approach
Big labs go wide. We go deep.
One language at a time, we collect real conversations from native speakers, with their consent, and turn them into an understanding of how the language is really used. Depth beats breadth when your users are the ones being misunderstood.
Language Brain
A layer any AI agent plugs into.
Language Brain sits between your agent and the people it serves. It understands dialects, slang and cultural context, so the agent you already have becomes accurate in languages it used to get wrong.
- 01
Listen
Native speakers share real conversations, stories and voice recordings, with explicit consent and human review.
- 02
Learn
We build language models and evaluation sets from that data, starting with Nila, our Tamil model trained from scratch.
- 03
Plug in
Your agent calls Language Brain to understand people and reply the way they actually talk.
Input
நாளைக்கு கடைக்கு வாறியோ?
// brain.understand(text)
{
language: "ta",
dialect: "jaffna",
register: "casual",
meaning: "Are you coming to the shop tomorrow?"
}Jaffna Tamil, casual. A generic model often misreads the dialect verb.
The playbook
Then we repeat it, each language faster and cheaper than the last.
- 1
Community
Partner with native speakers, diaspora groups and local researchers.
- 2
Data
Collect real speech and text with consent, reviewed by people before training.
- 3
Models
Train and evaluate on how the language is really spoken, not just written.
- 4
Reuse
Pipelines, tokenizers and lessons carry over, so every next language costs less.
Languages
The languages big models miss.
We start where the gap is widest: languages with tens of millions of speakers and very little of the data big models learn from.
- தமிழ்Live
Tamil · 80M+ speakers
Spoken Tamil differs sharply from written Tamil, and dialects from Jaffna to Chennai barely appear online.
Join the waitlist → - සිංහලNext
Sinhala · 17M+ speakers
Written and spoken Sinhala are almost two different registers, and models trained on formal text sound stiff in conversation.
Join the waitlist → - Èdè YorùbáPlanned
Yoruba · 45M+ speakers
A tonal language where meaning rides on diacritics that most web text leaves out.
Join the waitlist → - Harshen HausaPlanned
Hausa · 80M+ speakers
One of Africa's most spoken languages, yet with little web text, written in both Latin and Arabic (Ajami) script.
Join the waitlist → - አማርኛPlanned
Amharic · 55M+ speakers
Its own script (Ge'ez) and rich word forms get split into many pieces by tokenizers built for English.
Join the waitlist → - Asụsụ IgboPlanned
Igbo · 30M+ speakers
Tonal, with many dialects, and the little text online is often missing its tone marks.
Join the waitlist → - ភាសាខ្មែរPlanned
Khmer · 17M+ speakers
Written without spaces between words, so standard tokenizers and search struggle from the first step.
Join the waitlist → - پښتوPlanned
Pashto · 40M+ speakers
Strong regional dialects and very little high-quality digital text.
Join the waitlist → - Kreyòl ayisyenPlanned
Haitian Creole · 12M+ speakers
Spoken by nearly everyone in Haiti, while French dominates formal writing, so little text reflects how people talk.
Join the waitlist → - Runa SimiPlanned
Quechua · 8M+ speakers
A family of mostly oral varieties with little standardized text to learn from.
Join the waitlist →
Early access
Build with Language Brain.
We are opening Language Brain to a small group of teams building agents for these languages, and to speakers who want to help.