Uyir LabLanguage BrainNilaLow-resource languages
Introducing Uyir Lab: AI for the Languages Big Models Miss
Uyir Lab is a Toronto AI lab building Language Brain, a layer that makes AI agents accurate in dialects, slang and cultural context. Tamil first, then 15 more languages and 854M+ speakers.
Uyir Lab · · 3 min
Uyir Lab is an independent AI lab in Toronto. We build Language Brain, a layer any AI agent can plug into so it understands people the way they actually speak: in dialects, in slang, switching between languages mid-sentence, and leaning on cultural context that never made it into training data. We are starting with Tamil, and 15 more languages are on our roadmap, together spoken by more than 854 million people.
Uyir (உயிர்) means life in Tamil. Our mission is to give every language a life in AI.
The problem: AI learned the internet, not the way people talk
Large language models learn from the written web. That works well for English and a handful of other languages, and for the formal, written register of everything else. It works badly for how most of the world talks.
Ask a general model to follow a Jaffna Tamil conversation, a Chennai chat written in Tanglish, or a customer mixing Punjabi and English on a support call, and it often misreads the meaning, answers in stiff textbook language, or gets it wrong. We wrote about why in Why AI struggles with Tamil and Spoken vs written Tamil.
This is about to matter much more. Companies are putting AI agents in front of customers for banking, health, telecom, government services and commerce. For hundreds of millions of people, those agents are becoming the front door, and today the front door does not understand them.
What we're building
Language Brain
Language Brain sits between an AI agent and the people it serves. It interprets dialect, slang and code-switching on the way in, and helps the agent reply in a way that sounds natural to the person on the other side. A business keeps the agent it already has and makes it accurate in languages it used to get wrong.
Nila, our Tamil model
Nila (நிலா, moon) is our first language model, trained from scratch on 9.6 billion tokens of Tamil, English and Tanglish. Training from scratch let us design the tokenizer around Tamil script and put Tamil at the centre of the training data instead of squeezing it into a model built for English. You can read about the approach in Building a Tamil large language model from scratch.
A community voice studio
Text alone is not enough. People speak differently from how they write, and every region sounds different. Our contribution studio lets Tamil speakers anywhere record short sentences or talk freely about everyday scenarios. Every recording draws one loop of a kolam, and five recordings complete it. Contributors give explicit consent, and every recording is reviewed by people before it is used. The Uyir app for iOS and Android brings the same experience to phones.
Why go deep instead of wide
The largest labs cover hundreds of languages thinly. We go the other way: one language at a time, every layer of it, from the written standard to regional dialects, slang, voice and accent. Depth is what makes an agent trustworthy for the people it serves.
The method also compounds. The pipelines, tokenizer work and evaluation methods we build for Tamil carry over, so each new language costs less than the last. After Tamil comes Sinhala, then Punjabi, Swahili, Hausa, Amharic, Javanese, Fula, Yoruba, Bhojpuri, Oromo, Pashto, Igbo, Khmer, Haitian Creole and Quechua. See them all on our languages page.
How we work
- Consent first. We train only on data people chose to give us, or data licensed for commercial use.
- People review everything. Every recording and text is checked by a person before it reaches a model.
- Communities shape the work. We partner with native speakers, diaspora groups and researchers.
Get involved
- Speak Tamil? Record your voice in the studio. It takes a few minutes and draws you a kolam.
- Speak another language on our roadmap? Join the list on the languages page so your language is next.
- Building AI agents for these markets? Get early access to Language Brain.
- Journalist or researcher? Our press page has facts, images and contacts.
Add your voice to Nila
Record a sentence, tell a story in your own dialect, or share something you have written. It takes a few minutes.
Open the contribution studio →