Artificial Intelligence

Public Compute, Indian Languages and the Question of Who Gets to Build the Models

Subsidised GPU access under the IndiaAI Mission is necessary but not sufficient if data and talent remain concentrated

By Meera Krishnan · 12 August 2026 · 5 min read
Public Compute, Indian Languages and the Question of Who Gets to Build the Models

When the government approved the IndiaAI Mission in March 2024 with an outlay exceeding ten thousand crore rupees, the component that drew the most immediate attention from the country's AI research and startup community was not any regulatory provision but a practical one: subsidised access to a pool of graphics processing units, the specialised chips required to train large AI models, procured through a public-private partnership structure and offered to Indian researchers, startups and academic institutions at rates well below prevailing market prices. The excitement was understandable. Compute access has been the single most binding constraint on Indian AI research for several years, with even well-funded Indian startups routinely priced out of the GPU clusters that American and Chinese labs treat as a baseline input, a scarcity driven by global supply constraints, export controls on the most advanced Nvidia chips, and the sheer capital intensity of building data centres at the scale frontier training runs require.

Why compute alone was never the whole bottleneck

The mission's logic, that removing the compute bottleneck would unlock a wave of Indian AI innovation, is only partially correct, and the partiality matters. Training a capable large language model requires three things in combination: compute, high-quality training data, and the specialised talent to design architectures and run training experiments efficiently. India has made real strides on compute access through this mission and has genuine talent, evidenced by the number of Indian researchers at OpenAI, Google DeepMind and Meta's AI research divisions. What India conspicuously lacks, particularly for its non-English languages, is data, and no amount of subsidised GPU time compensates for training corpora that simply do not exist at the scale and quality that model training requires.

The specific data problem in Indian languages

English-language AI models benefit from the accident of the internet's early development being disproportionately English-dominated, producing a vast, if messy, corpus of text that models like GPT and its successors could be trained on. Hindi has meaningfully less digitised text relative to its speaker population, and languages like Bhojpuri, Santhali, or Konkani, each spoken by millions, have digitised corpora that are smaller still, often limited to government documents, a handful of news outlets, and whatever devotional or literary texts have been digitised by university libraries, nowhere close to the scale needed to train a model with genuine fluency and cultural nuance. This is not a problem subsidised compute solves; it is a data curation and digitisation problem that requires sustained investment in transcription, translation and annotation work, much of it requiring native speaker involvement that cannot be crowdsourced cheaply without compromising quality.

What Bhashini has and has not achieved

The government's Bhashini initiative, launched under the broader Digital India Bhashini Division, has made a genuine effort here, building a growing repository of parallel corpora across Indian languages and offering APIs for translation and speech recognition that several state governments have integrated into citizen service portals. Its usefulness is real for narrower tasks like document translation and voice-based government service access, particularly for elderly or less literate users navigating schemes in their own language rather than in English or Hindi. But Bhashini's corpora remain considerably smaller and narrower in domain than what would be needed to train a foundation model with the fluency and reasoning capability of frontier systems, and conflating the initiative's genuine achievements with the much larger task of building India-specific foundation models risks both overselling Bhashini and understating the scale of work still required.

Who is actually queuing for the GPUs

There is also a distributional question embedded in how the subsidised compute pool gets allocated that the mission's early implementation has not fully resolved. Early allocation rounds have understandably favoured established AI startups and research institutions with the technical sophistication to submit compelling compute-access proposals, a reasonable criterion for ensuring the scarce resource is used productively, but one that risks reproducing the exact talent concentration, in Bengaluru, Hyderabad and a handful of elite technical institutions, that a more geographically distributed AI ecosystem would need to overcome. Researchers working on lower-resource language models at smaller state universities, who may have exactly the community relationships needed to source and validate quality training data in Santhali or Bodo, are precisely the applicants least likely to have the grant-writing experience or existing publication record that competitive compute allocation processes tend to reward.

What a more complete strategy would look like

A more complete approach would treat data curation as a public infrastructure investment in its own right, comparable in seriousness to the compute subsidy, funding sustained, well-paid work by native speaker linguists and community organisations to build genuinely large, high-quality corpora in India's lower-resource languages, rather than treating this as a residual task to be solved through crowdsourced volunteer contribution once compute access is sorted. It would also mean setting aside a meaningful share of subsidised compute specifically for this data-scarce, talent-thin segment of the research ecosystem, accepting that the return on this investment will look less impressive in the near term than funding another well-resourced startup building yet another English-fluent chatbot, because the payoff here is measured in linguistic and cultural preservation as much as in commercial output.

The IndiaAI Mission deserves credit for identifying and addressing a genuine, previously underserved bottleneck, and the compute subsidy will very likely produce a visible increase in the volume of Indian AI research output over the next several years. But if the mission is ultimately judged only by how much compute got deployed and how many startups it helped fund, it will have solved the easier half of the problem. The harder half, ensuring that India's AI capability actually reflects and serves its linguistic diversity rather than replicating an English-and-Hindi-dominant approximation of it, depends on investments in data and people that no GPU cluster, however generously subsidised, can substitute for.

#indiaai mission#gpu compute#indian languages ai#bhashini#low resource languages#public digital infrastructure

Related reading

IJP
News Brief
India's middle-income trap has a problem money may fix
Artificial Intelligence·The Economic Times·15 Sep 2026

India's middle-income trap has a problem money may fix

Read full story →
IJP
News Brief
Tamil Nadu CB-CID nabs Rajasthan scammer who circulated Al-generated deep fake video impersonating CM Vijay
Artificial Intelligence·The Hindu·14 Sep 2026

Tamil Nadu CB-CID nabs Rajasthan scammer who circulated Al-generated deep fake video impersonating CM Vijay

Read full story →
IJP
News Brief
Keralam chief minister VD Satheesan Never Danced. Can India’s Deepfake Laws Catch Up With AI?
Artificial Intelligence·Open Magazine·14 Sep 2026

Keralam chief minister VD Satheesan Never Danced. Can India’s Deepfake Laws Catch Up With AI?

Read full story →
IJP
News Brief
Data centre boom: India’s infrastructure opportunity could hit $90 billion
Artificial Intelligence·The Times of India·12 Sep 2026

Data centre boom: India’s infrastructure opportunity could hit $90 billion

Read full story →