22 languages, score 0
Point a state-of-the-art language model at the names of the Northeast ~ Khasi surnames, Garo festivals, the titles a clan carries ~ and ask it to recognise them, and it returns a number that should stop a clearance meeting: an F1 score of zero. Not low. Zero. The finding is not a complaint on a blog; it sits in this year’s record of the Association for Computational Linguistics, where a researcher from Shillong, Badal Nyalang, showed two failures that compound ~ the model cannot see the culturally freighted name at all, and the tokeniser beneath it mangles the diacritics that carry meaning in Khasi and the morpheme mark that holds Garo together, corrupting them in up to half the cases across the models tested.
A purpose-built regional tagger, made on the community’s own terms, scored 0.964 on the same ground. The gap between zero and that is not compute. It is attention. This is the hole in India’s sovereign-AI story: it keeps a language count and no civilisation ledger. The country has decided, rightly, to build its own models rather than rent them from California, and put real money and hardware behind the decision. But a foundation model is only as sovereign as the meanings beneath it, and a model can be taught to produce fluent Khasi while remaining blind to the people who live inside the language. India is measuring breadth ~ how many tongues, how many datasets, how many chips ~ and calling it depth.
The build itself is serious, and its architects deserve to be heard fairly. Under the IndiaAI Mission, a common compute pool of tens of thousands of subsidised processors has been assembled where two years ago there was almost none, within reach of a start-up in Pune or a lab in Guwahati. From more than 500 applications the government has chosen twenty indigenous model proposals to fund. BharatGen, the IIT Bombay-led consortium at the centre of the effort, has released Param2, a seventeen-billion-parameter model it says now generates text across all twenty-two scheduled languages ~ the constitutional list it was asked to reach.
For years the world’s models treated Indian languages as a rounding error; a scheduled-language list on a state-backed model is a political fact before a technical........
