Uyir Bench · Tamil
How well does AI understand Tamil?
Not textbook Tamil: the way people actually speak it, from Jaffna to Chennai to Scarborough. Uyir Bench tests the leading AI models on dialects, Tanglish, culture, translation and reasoning, with every test written or checked by native speakers.
In progress · first results coming soon
We're building the first version now. Native speakers are reviewing the test questions, starting with a 50-question pilot and growing to 500. We only publish scores once the tests behind them are reviewed, so there are no numbers here yet.
Models we're testing: GPT, Claude, Gemini, open models such as Llama, Qwen and Gemma, Sarvam, and our own Nila.
What we test
Dialects
Jaffna, Batticaloa, Trincomalee and Upcountry Tamil from Sri Lanka; Chennai, Madurai, Kongu and Tirunelveli from India; and diaspora Tamil.
Code-mixing
Tanglish, Tamil typed in English letters, and sentences that switch script or language halfway through.
Spoken vs written
Tamil is written one way and spoken another. Can a model follow both, and answer in the right register and level of politeness?
Grammar
Word endings, sandhi, case markers and spelling, where small mistakes change the meaning.
Culture
Festivals, food, kinship terms, religion, Thirukkural and Sangam literature, and everyday life in Sri Lanka, Tamil Nadu and the diaspora.
Translation
Tamil to English and English to Tamil, including idioms that don't translate word for word.
Reasoning in Tamil
Maths word problems, logic and instructions written in Tamil, not translated from English.
Practical tasks
Forms, government and medical instructions, and customer support, where a misunderstanding has real cost.
Sensitive topics
Political and historical questions, which should get factual, neutral answers.
Speech (next)
Speech recognition on real accented and code-mixed recordings from our contributors and the AI challenge.
How we test
- Native speakers write and check every test
- AI can help draft a question, but nothing counts until two native speakers approve it.
- Same test for every model
- Identical prompts and settings, temperature 0, with model versions and run dates recorded.
- Fair scoring
- Exact or normalised matching where an answer is clear, translation metrics for translation, and a written rubric otherwise. When an AI judge is used, it is checked against human scores and never grades its own model family.
- A private test set
- About 10% of the tests are published. The rest stay private so no model can be trained on them, which keeps the scores honest.
- Outputs are for testing only
- Model answers are used to measure and report, never to train our own models.
For AI labs
Private evaluations and expert Tamil data
We run private, held-out Tamil evaluations for your models and supply native-speaker data that fixes what they get wrong: expert responses, preference rankings, cultural safety reviews and dialect speech, all with documented consent.
Email hello@uyirlab.com
For Tamil speakers
Test the AI yourself
Say a phrase in your dialect and see if the AI understands you. When it gets it wrong, you can share it with us, and real failures become candidate tests for the next version of the report. We're also looking for paid native-speaker reviewers, especially from Batticaloa, Trincomalee, the Upcountry, Madurai, Kongu, Tirunelveli, Malaysia and Singapore.