The signal behind the AI headlines — ranked, not recappedGet the daily brief →
The AI Daily
Analysis · AI Economics · India

Fast-forward the AI ROI story: what enterprise AI costs 3 and 5 years out

It already works — the AI ROI you can’t justify today gets obvious about three years out. Add the two forces we left out of our first model — hardware depreciation and the relentless optimisation of inference — and today’s expensive pilot becomes tomorrow’s obvious yes. Provided you don’t let your own bill race the other way.

Fast-forward the AI ROI story: the same enterprise AI support desk costs about ₹2 crore a year now, ₹69 lakh in three years, and ₹34 lakh in five.

The same enterprise support desk, projected on our base case — a third of the cost in three years, a sixth in five. Set your own numbers in the calculator below.

~10×/yr
Fall in the price of a fixed unit of AI capability (a16z “LLMflation”)
₹65
Per subsidised GPU-hour under the IndiaAI Mission — about a third of global rates
$0.80
Per million tokens for Sarvam-105B, India-hosted — the sovereign price point
The short version

AI looks expensive because you are pricing it at today’s rates. The per-token price of a fixed capability is falling roughly 10× a year — so the real question is not “can we afford this?” but “what will it cost once we have scaled it?”

The levers pulling the LLM price down

  1. Algorithmic & serving efficiency — quantisation, distillation, mixture-of-experts, speculative decoding. The compute to hit a fixed quality bar roughly halves every eight months. The biggest lever, and it is not slowing.
  2. Smaller models that punch up — a 1-billion-parameter model today beats a 175-billion one from 2021. Move to the cheapest model that clears your quality bar and you ride the steep curve, not the gentle frontier one.
  3. Hardware price-performance — compute per rupee keeps improving ~35% a year, once today’s memory-price spike clears.
  4. Competition & open weights — open models chasing the frontier compress everyone’s margins.
  5. India’s sovereign stack — INR pricing, subsidised GPUs (~₹65/hr under the IndiaAI Mission), Indic-token efficiency, and India-hosted models like Sarvam. A second, steeper discount only India gets.

What that means right now

  • Prioritise use-cases that at least break even within two years, on a capex + opex basis. They may look marginal today; they compound into strong ROI over the following three years as costs fall.
  • Design for optionality — no overdependence on one model or one infrastructure layer. The jury is still out on who the winners will be; keep the freedom to switch models, providers and hosting as prices and leaders move.

The catalyst for this piece landed on 31 July 2026. At its first “Epoch” showcase, Bengaluru’s Sarvam AI unveiled Sarvam Code — a “made in India” coding agent that Sarvam claims edged out Anthropic’s Claude Code on the Terminal Bench evaluation (72 of 89), running on models it serves from Indian soil at a fraction of Western prices. Whatever the benchmark holds up to, the pricing is the real signal: Sarvam quotes its India-hosted inference at “up to 9× lower” than comparable offerings, and its flagship Sarvam-105B at about $0.80 per million tokens.

That is a direct challenge to a claim we made in our earlier primer, Enterprise AI Costs 101: that for Indian businesses, AI would stay stubbornly expensive because the bill is dollar-denominated and the API price is only the tip of the cost. That primer still holds for today. But it was a snapshot, and it left out the two forces that decide where costs go next: the depreciation of compute, and the compounding optimisation of how tokens are generated. This piece adds them back — and then hands you a calculator that projects your own use-case three and five years out.

1 · The per-token price is collapsing

Start with the number that sounds too good to be true. By a16z’s “LLMflation” measure, the cost of a fixed unit of capability — the price to hit a given quality bar — has fallen roughly 10× per year. Epoch AI’s more rigorous analysis puts the median at ~50×/year across benchmarks, and ~40×/year specifically for GPT-4-level reasoning. To reach MMLU performance that cost $60 per million tokens in late 2021, you now pay pennies.

The counter-intuitive part: this is mostly not a hardware story. Raw silicon price-performance (FLOP per dollar) improves only about 33–38% a year. The bulk of the collapse is algorithmic and systems work — quantisation to 8- and 4-bit, distillation into small capable models, mixture-of-experts, speculative decoding, continuous batching — plus brutal open-source competition compressing margins. A 1-billion-parameter model today beats a 175-billion-parameter model from 2021.

The catch: you only capture the full 10×/year if you keep down-shifting to the cheapest model that clears your quality bar. If you pin your product to the frontier flagship, you ride a much gentler curve — flagship output prices have fallen roughly 6× in two years (~55–60%/yr), not 10×. The decline rate you actually get is a choice, and the calculator below lets you set it.

2 · So why do the bills keep going up?

Because cost-per-token and cost-per-task are different animals. Two forces push the bill the other way:

So the honest model has three dials, not one: the per-token price falls, while tokens-per-task and total volume rise. Multiply them together and the same use-case can get dramatically cheaper — or dramatically more expensive.

3 · The memory wrinkle

There is one place the “hardware always gets cheaper” assumption is currently, badly wrong: memory. As we covered in The Memory Boom, AI demand has sent DRAM and high-bandwidth memory prices surging up to 90% from late 2025, and analysts call the squeeze structural — tight through 2027, with no firm normalisation date. Memory is now 40%+ of an AI server’s bill of materials, so for the next year or two it partly cancels the silicon deflation. Net for on-premise builds: costs still fall, but the curve flattens — and may briefly tick up in 2026–27 for memory-heavy configs before resuming its decline. The collapse is real; it is just not monotonic.

4 · India’s curve bends faster

Here is where the earlier “expensive in India” claim inverts. Several levers pull India’s effective cost below the dollar-priced Western stack:

None of this makes AI free. But it means the Indian curve has a second, steeper slope available — a sovereign track — that the dollar-priced frontier does not. The calculator lets you switch between them.

5 · The forward calculator

Take a large enterprise support operation — roughly 1.5 million queries a month. On today’s frontier prices that is about ₹2 crore a year. Hold the capability fixed and fast-forward on our base case, and it lands near ₹69 lakh in three years and ₹34 lakh in five — the same work at a third, then a sixth, of the cost. That is the ROI story; it just needed a calendar. The catch is that the calendar cuts both ways — bolt on agents and pin to the frontier, and the same desk can cost more.

Describe a use-case, set your assumptions, and see the monthly cost projected out to your horizon. The waterfall is the point: it decomposes the change from today’s bill to the future bill into its competing pieces — how much the price collapse takes off, and how much the reasoning tax and rising volume add back on. Defaults are our research-backed BASE case; every lever is yours to move.

📊 The AI Cost Curve calculator
Project a use-case 3 or 5 years out. Watch the deltas fight it out in the waterfall.
Per-token price decline / year
Reasoning-token tax / year (more tokens per task)
Usage growth / year (Jevons demand)
Today (Year 0)
Year 5
Today / future bill Price collapse (cheaper) Reasoning + volume (costlier) Sovereign discount
Monthly cost, year by year

6 · The leader’s verdict

Three things to take to a budget conversation:

Sources: a16z LLMflation; Epoch AI inference price trends; Menlo Ventures State of GenAI 2025; ARC Prize o3; Counterpoint; IndiaAI Mission; Analytics India Magazine (Sarvam Epoch). Projections are illustrative models built on these rates, not forecasts. Verify against your own workload before budgeting.