Skip to content

On-Device AI Is Here. Which India Is It Arriving In?

Two Facts That Explain Everything

Drive north out of Chennai and keep going for eighteen hundred kilometres. You won't cross a border or change currency, but the alphabet on the road signs will change five times before you stop. Not the accent, the letters themselves. Five different writing systems, one country, one drive. And somewhere on that road, more than once, the phone in your pocket will drop to one bar and stay there.

How many languages there are, and how thin the network is: those two facts are the whole story of on-device AI in India. And the part that changes how you should be thinking about it isn't obvious until you look at the phones themselves. They aren't getting cheaper. They're getting more expensive, faster than at any point in the last decade, and AI itself is the reason.

A Market Getting More Expensive, Not Cheaper

According to TRAI's Quarterly Performance Indicator Report for January–March 2026, India's total internet subscriber base reached 1,092.79 million, roughly 64 million net additions in a single quarter. But the aggregate hides the real story: urban India sits at 126.80 internet subscribers per hundred people, which is saturation plus second devices, while rural India sits at 48.31. Same report, same quarter, same country.

Now the reversal. Smartphone shipments told a different story in the same window: IDC's tracker for Q2 2026 recorded shipments down 11.1% year-on-year to 33.2 million units, bringing first-half volumes to 64.2 million, the weakest first half in five years, while the average selling price surged 14.4% to roughly $315. The lower end of the market didn't just soften, it collapsed: phones under $100 dropped 74.3% year on year, falling from 15.6% of the market to 4.5%.

IDC's explanation is memory. Demand from AI data-centre construction has driven up the cost of DRAM and flash storage to the point where manufacturers either raised prices or stopped making entry-level phones altogether. Put plainly: the build-out of AI in the cloud is pricing Indian buyers out of the phones that would run AI locally.

What “On-Device” Actually Means

When a cloud assistant summarises a voice note, the audio leaves the phone, travels to a data centre, gets processed, and comes back, which needs a stable connection, and puts the recording on someone else's machine for the duration. On-device inference deletes that trip. The model runs on a neural processing unit inside the phone itself, silicon built to do the matrix arithmetic language models need without flattening the battery.

The model has to live in memory next to the operating system and everything else that's open, which is why RAM, not processor speed, turns out to be the real gate. And that has a consequence specific to localization: a cloud model is a service, and a language can be added on a Tuesday by pushing a new version to a server. A model on the device is a shipped artifact, quantized, size-constrained, frozen the moment the phone leaves the factory. Its language coverage is whatever the vendor chose to fit inside a file that has to share RAM with Instagram. On-device language support is a hardware roadmap decision, which is a very different negotiation than the one most of us are used to having with a client.

The Hardware Gap: Google vs. Apple

Google's Gemini Intelligence, the proactive layer built into Android 17, ships with a hard floor: 12GB of RAM or more, a qualified flagship chip, and Gemini Nano v3 or higher running through AI Core, plus five OS upgrades and six years of security patches from the manufacturer. The rollout starts with the newest Galaxy and Pixel phones. Last year's flagships, the Pixel 9 series and Galaxy S25 line included, don't clear the Nano v3 bar.

Now hold that 12GB floor next to the memory shortage above: Google's flagship AI experience requires a large allocation of the single most expensive component in the phone, during the worst memory-pricing environment in years.

Apple's hardware floor is lower, 8GB of unified memory on an iPhone 15 Pro or newer, but its language position is worse. Apple Intelligence reached India in localised English on 31 March 2025. As of WWDC in June 2026, none of India's languages, not Hindi, not Tamil, not any of the other twenty, are on Apple's supported list. Apple's roughly 8.5% of Indian shipments last quarter (about 27% by value) is mostly older hardware, so the actual audience for Apple's on-device AI in India is a slice of a slice, operating in a language most of the country doesn't speak at home.

Why Offline Isn't a Luxury Here

Rural India's 48.31 subscribers per hundred people counts subscriptions, it doesn't count the dead stretch on the highway, or what happens when an entire village shares one congested cell tower at eight in the evening. Think about who actually needs language AI in the field this week: a court interpreter in a district session, a health worker taking a case history in a language the clinic's system was never configured for, a journalist transcribing an interview on the road back. Every one of those jobs happens where the network is weakest, and every one involves material that shouldn't be going to a third-party server anyway.

That second point deserves more weight than it usually gets. If the material is medical records, legal filings, or witness statements, local inference is a confidentiality argument before it's a speed argument, the same conversation every translator has already had with a client about why a source file can't go through a free web tool.

The Three-Layer Stack the West Doesn't Have

In most Western markets, the state buys language technology and regulates it, the models come from private companies and government is a customer. India built something structurally different, in public, across three layers that talk to each other.

The first is Bhashini, the government's National Language Translation Mission: free translation, speech-recognition and text-to-speech APIs across the 22 scheduled languages, wired directly into government services, language capability treated the way roads are.

The second is AI4Bharat at IIT Madras, building the parts commercial vendors tend to skip. IndicTrans2 covers all 22 scheduled languages, including multiple scripts for lower-resource languages like Kashmiri, Manipuri and Sindhi, and the lab has published distilled translation models around 200 million parameters, small enough to actually run on a phone, which is the entire point of building them that size.

The third is the one that would cause an argument in Brussels. Sarvam AI, founded by AI4Bharat alumni, was selected under the IndiaAI Mission for subsidised access to 4,096 NVIDIA H100 GPUs for six months, a roughly $26M compute bill, with the government covering around $10.4M of it through compulsorily convertible debentures rather than a grant. In June 2026, Sarvam closed the first tranche of a Series B at a $1.5B valuation; the government's compute stake is reportedly set to convert into an equity holding. Weeks earlier, Sarvam open-sourced two models, roughly 30B and 105B parameters, under Apache 2.0, trained on that public compute. Those open models were trained on human-annotated corpora and measured against human-built benchmarks in languages like Odia, Assamese and Manipuri, the actual bottleneck for any Indian-language model, and not one that scraping more of the web solves.

Silicon Moves. Language Doesn't.

The silicon is coming down-market faster than expected. In late May 2026, MediaTek announced the Dimensity 8550, nearly identical to its predecessor except for one change that matters: an upgraded NPU with a dedicated LLM Booster, giving it native support for Gemini Nano v3, making upper-mid-range phones eligible for Gemini Intelligence for the first time. The catch: a phone with that chip still has to clear Google's 12GB RAM floor, so the chip reaches down market just as the memory specification pulls back up, in the same year memory got expensive.

The hardware curve does bend toward India, slowly, with something heavy pushing against it right now. The language curve doesn't bend on its own. No chip roadmap adds Kannada. That's a decision someone makes in a room, and it stays a decision until someone else changes it. Using “Hindi” as shorthand for Indian-language support is a habit worth dropping: a model that handles Hindi has solved a large market. It hasn't solved India.

Pressure-Test Your Own Rollout

The real test is narrow and specific: can a phone people actually buy, at a price people actually pay, take a voice note recorded in Meiteilon, in a place with one bar of signal, and do something useful with it without sending it anywhere? Not yet, but the pieces are on the table. Small translation models built for constrained devices. A public platform treating language as national infrastructure. Open weights across twenty-two languages. Mid-range silicon with a real NPU in it. Nobody has assembled them yet.

We built a one-page checklist, the pressure-test pack, for exactly this: six questions to ask before you promise a client on-device AI works for their Indian-language rollout, covering phone band, language, network condition, confidentiality, the public-infrastructure fit, and timing.

Download your pressure-test pack