Uyir Lab models

Models that speak the way people do.

Trained from scratch, one language at a time. Tamil first, with 15 more on the way. Move your cursor through the words.

The lineup

One model per language, built with its speakers

Live · research preview

நிலா

Nila

Tamil · 336M parameters · trained from scratch

Our first model. It reads and writes Tamil, English and Tanglish, and it is the start of Language Brain: one layer any agent can use to understand how people really speak.

Play with Nila →
  • සිංහල

    Sinhala

    20M+ speakers · next up

  • پنجابی · ਪੰਜਾਬੀ

    Punjabi

    156M+ speakers · planned

  • Kiswahili

    Swahili

    97M+ speakers · planned

  • Harshen Hausa

    Hausa

    94M+ speakers · planned

  • አማርኛ

    Amharic

    60M+ speakers · planned

  • Basa Jawa

    Javanese

    68M+ speakers · planned

  • Fulfulde

    Fula

    39M+ speakers · planned

  • Èdè Yorùbá

    Yoruba

    50M+ speakers · planned

  • भोजपुरी

    Bhojpuri

    52M+ speakers · planned

  • Afaan Oromoo

    Oromo

    45M+ speakers · planned

  • پښتو

    Pashto

    55M+ speakers · planned

  • Asụsụ Igbo

    Igbo

    34M+ speakers · planned

  • ភាសាខ្មែរ

    Khmer

    21M+ speakers · planned

  • Kreyòl ayisyen

    Haitian Creole

    13M+ speakers · planned

  • Runa Simi

    Quechua

    7M+ speakers · planned

Inside Nila

Trained from zero on 9.6 billion tokens

336M

parameters

9.6B

training tokens

24

transformer layers

64000

token vocabulary, built for Tamil

Watching it learn

19,000 steps over 51 hours on two GPUs

3.04.05.06.07.005k10k15k19k
Training loss Tamil held-out English held-out Tanglish held-outLower is better. Steps of about 524k tokens each; real numbers from Nila's training run, Oct 7 to 9, 2026.

What it read

  • Tamil50%
  • English43.6%
  • Tanglish6%
  • Classical Tamil0.4%

Under the hood

  • Llama-style decoder, grouped-query attention
  • 2,048-token context, rotary positions
  • Custom 64k tokenizer built for Tamil script
  • Instruction-tuned on 25,600 open, licensed chat examples
  • Commercially licensed data only

Playground

Talk to நிலா

A small research model: it will get things wrong, sometimes confidently. Every mistake shows us what to teach it next.

Nila Research preview

Nila is Uyir Lab's first language model, trained from scratch on Tamil, English and Tanglish. Type in Tamil or English, or speak in Tamil.

Nila is small and new. It makes mistakes, including confident wrong facts. Don't share personal information, and don't use it for medical, legal or immigration decisions. Privacy

The next model learns from your voice

Record a minute of how you really speak. It goes straight into teaching Nila the Tamil the internet never wrote down.

Record your voice