Uyir Lab · Language Brain

AI that speaks the languages big models miss.

Uyir Lab makes AI agents accurate in the languages big models get wrong, the way they are actually spoken.

The gap

Billions of people speak languages AI handles poorly.

Big models learn from the written web, and most of the world does not talk the way the web is written. People speak in dialects, slang, code-switching and cultural shorthand that never made it into training data. So agents misunderstand them, answer in stiff textbook language, or simply get it wrong.

Dialects

The same language sounds different from one city to the next. Models usually learn only the standard form.

Slang and code-switching

Real conversations mix languages and scripts mid-sentence. Training data rarely does.

Cultural context

Kinship terms, honorifics, food, faith and humour carry meaning a literal translation misses.

Our approach

Big labs go wide. We go deep.

One language at a time, we collect real conversations from native speakers, with their consent, and turn them into an understanding of how the language is really used. Depth beats breadth when your users are the ones being misunderstood.

Language Brain

A layer any AI agent plugs into.

Language Brain sits between your agent and the people it serves. It understands dialects, slang and cultural context, so the agent you already have becomes accurate in languages it used to get wrong.

  1. 01

    Listen

    Native speakers share real conversations, stories and voice recordings, with explicit consent and human review.

  2. 02

    Learn

    We build language models and evaluation sets from that data, starting with Nila, our Tamil model trained from scratch.

  3. 03

    Plug in

    Your agent calls Language Brain to understand people and reply the way they actually talk.

language-brain

Input

நாளைக்கு கடைக்கு வாறியோ?

// brain.understand(text)
{
  language: "ta",
  dialect: "jaffna",
  register: "casual",
  meaning: "Are you coming to the shop tomorrow?"
}

Jaffna Tamil, casual. A generic model often misreads the dialect verb.

⚑ Preview of the planned interface

The playbook

Then we repeat it, each language faster and cheaper than the last.

  1. 1

    Community

    Partner with native speakers, diaspora groups and local researchers.

  2. 2

    Data

    Collect real speech and text with consent, reviewed by people before training.

  3. 3

    Models

    Train and evaluate on how the language is really spoken, not just written.

  4. 4

    Reuse

    Pipelines, tokenizers and lessons carry over, so every next language costs less.

Languages

The languages big models miss.

We start where the gap is widest: languages with tens of millions of speakers and very little of the data big models learn from.

Early access

Build with Language Brain.

We are opening Language Brain to a small group of teams building agents for these languages, and to speakers who want to help.

I am
Languages you care about