# Fast-forward the AI ROI story: what enterprise AI costs 3 and 5 years out

> Enterprise AI's per-token price is collapsing ~10x a year — but reasoning workloads and Jevons demand can make your bill rise anyway. A forward calculator with a waterfall that shows the deltas over 3 and 5 years, plus why sovereign models like Sarvam bend India's curve down.

Source: https://theaidaily.in/analysis/ai-cost-curve.html
Published: 2026-08-03

---

The signal behind the AI headlines — **ranked, not recapped**[Get the daily brief →](/subscribe.html)

[********The AI Daily](/)*[Subscribe →](/subscribe.html)☰

Analysis · AI Economics · India

Fast-forward the AI ROI story: what enterprise AI costs 3 and 5 years out

It already works — the AI ROI you can’t justify today gets obvious about three years out. Add the two forces we left out of our first model — hardware depreciation and the relentless optimisation of inference — and today’s expensive pilot becomes tomorrow’s obvious yes. Provided you don’t let your own bill race the other way.

**The AI Daily** · Analysis
~14 min read · Interactive forward calculator
August 3, 2026

The same enterprise support desk, projected on our base case — a third of the cost in three years, a sixth in five. Set your own numbers in the calculator below.

~10×/yr

Fall in the price of a fixed unit of AI capability (a16z “LLMflation”)

₹65

Per subsidised GPU-hour under the IndiaAI Mission — about a third of global rates

$0.80

Per million tokens for Sarvam-105B, India-hosted — the sovereign price point

The short version

AI looks expensive because you are pricing it at **today’s** rates. The per-token price of a fixed capability is falling roughly **10× a year** — so the real question is not “can we afford this?” but “what will it cost once we have scaled it?”

#### The levers pulling the LLM price down

- **Algorithmic & serving efficiency** — quantisation, distillation, mixture-of-experts, speculative decoding. The compute to hit a fixed quality bar roughly halves every eight months. The biggest lever, and it is not slowing.

- **Smaller models that punch up** — a 1-billion-parameter model today beats a 175-billion one from 2021. Move to the cheapest model that clears your quality bar and you ride the steep curve, not the gentle frontier one.

- **Hardware price-performance** — compute per rupee keeps improving ~35% a year, once today’s memory-price spike clears.

- **Competition & open weights** — open models chasing the frontier compress everyone’s margins.

- **India’s sovereign stack** — INR pricing, subsidised GPUs (~₹65/hr under the IndiaAI Mission), Indic-token efficiency, and India-hosted models like Sarvam. A second, steeper discount only India gets.

#### What that means right now

- **Prioritise use-cases that at least break even within two years**, on a capex + opex basis. They may look marginal today; they compound into strong ROI over the following three years as costs fall.

- **Design for optionality — no overdependence on one model or one infrastructure layer.** The jury is still out on who the winners will be; keep the freedom to switch models, providers and hosting as prices and leaders move.

The catalyst for this piece landed on **31 July 2026**. At its first “Epoch” showcase, Bengaluru’s [Sarvam AI unveiled **Sarvam Code**](https://analyticsindiamag.com/ai-features/sarvams-biggest-leap-yet-all-that-happened-at-epoch) — a “made in India” coding agent that Sarvam claims edged out Anthropic’s Claude Code on the Terminal Bench evaluation (72 of 89), running on models it serves from Indian soil at a fraction of Western prices. Whatever the benchmark holds up to, the pricing is the real signal: Sarvam quotes its India-hosted inference at **“up to 9× lower”** than comparable offerings, and its flagship **Sarvam-105B at about $0.80 per million tokens**.

That is a direct challenge to a claim we made in our earlier primer, [Enterprise AI Costs 101](/analysis/llm-cost.html): that for Indian businesses, AI would stay stubbornly expensive because the bill is dollar-denominated and the API price is only the tip of the cost. That primer still holds for today. But it was a snapshot, and it left out the two forces that decide where costs go next: the depreciation of compute, and the compounding optimisation of how tokens are generated. This piece adds them back — and then hands you a calculator that projects your own use-case three and five years out.

## 1 · The per-token price is collapsing

Start with the number that sounds too good to be true. By a16z’s [“LLMflation”](https://a16z.com/llmflation-llm-inference-cost/) measure, the cost of a fixed unit of capability — the price to hit a given quality bar — has fallen roughly **10× per year**. Epoch AI’s more rigorous [analysis](https://epoch.ai/data-insights/llm-inference-price-trends) puts the median at ~50×/year across benchmarks, and ~40×/year specifically for GPT-4-level reasoning. To reach MMLU performance that cost $60 per million tokens in late 2021, you now pay pennies.

The counter-intuitive part: **this is mostly not a hardware story.** Raw silicon price-performance (FLOP per dollar) improves only about [33–38% a year](https://epoch.ai/blog/trends-in-machine-learning-hardware). The bulk of the collapse is algorithmic and systems work — quantisation to 8- and 4-bit, distillation into small capable models, mixture-of-experts, speculative decoding, continuous batching — plus brutal open-source competition compressing margins. A 1-billion-parameter model today beats a 175-billion-parameter model from 2021.

**The catch:** you only capture the full 10×/year if you keep down-shifting to the cheapest model that clears your quality bar. If you pin your product to the frontier flagship, you ride a much gentler curve — flagship output prices have fallen roughly 6× in two years (~55–60%/yr), not 10×. The decline rate you actually get is a choice, and the calculator below lets you set it.

## 2 · So why do the bills keep going up?

Because cost-per-token and cost-per-task are different animals. Two forces push the bill the other way:

- **Jevons’ paradox.** As the unit price falls, usage explodes. Enterprise generative-AI spend went from [$1.7B in 2023 to $11.5B in 2024 to $37B in 2025](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) — up 3.2× year over year even as per-token prices fell. Menlo Ventures names this explicitly as the 2026 base case.

- **The reasoning tax.** Reasoning and agentic workloads burn far more tokens per task. On the ARC-AGI test, a single hard task on OpenAI’s o3 ranged from [about $20 to about $4,560](https://arcprize.org/blog/oai-o3-pub-breakthrough) in compute depending on how hard it “thought” — a ~170× spread on the same problem. As you move from single prompts to multi-step agents, tokens-per-task climbs fast.

So the honest model has three dials, not one: the per-token price falls, while tokens-per-task and total volume rise. Multiply them together and the same use-case can get dramatically cheaper — or dramatically more expensive.

## 3 · The memory wrinkle

There is one place the “hardware always gets cheaper” assumption is currently, badly wrong: memory. As we covered in [The Memory Boom](/analysis/memory-boom.html), AI demand has sent DRAM and high-bandwidth memory prices [surging up to 90%](https://www.counterpointresearch.com/en/insights/Memory-Prices-Surge-Up-to-90-From-Q4-2025) from late 2025, and analysts call the squeeze structural — tight through 2027, with no firm normalisation date. Memory is now 40%+ of an AI server’s bill of materials, so for the next year or two it partly cancels the silicon deflation. Net for on-premise builds: costs still fall, but the curve **flattens — and may briefly tick up in 2026–27** for memory-heavy configs before resuming its decline. The collapse is real; it is just not monotonic.

## 4 · India’s curve bends faster

Here is where the earlier “expensive in India” claim inverts. Several levers pull India’s effective cost below the dollar-priced Western stack:

- **The compute is basically duty-free.** Servers and GPUs carry ~0% basic customs duty under the ITA-1 schedule; the 18% IGST is fully creditable. India’s cost base is set by power (₹4.5–7.2/kWh) and utilisation, not tariffs.

- **Subsidised GPUs at scale.** The [IndiaAI Mission](https://www.digitalindia.gov.in/press_release/indiaai-mission-expands-ai-ecosystem-with-affordable-compute-and-startup-support/) has put 34,000+ GPUs on tap at roughly **₹65 per GPU-hour** after a ~40% subsidy — about a third of global on-demand rates.

- **INR-denominated, India-hosted models.** No dollar FX exposure, no cross-border egress, lower latency. Sarvam-105B at ~$0.80/M tokens and voice at ₹3.5/minute (versus an ₹8–12 industry norm) are priced in that world.

- **Token efficiency for Indian languages.** Models built with Indic-first tokenisers spend fewer tokens per word of Hindi, Tamil or Bengali — a direct discount on any Indian-language workload.

None of this makes AI free. But it means the Indian curve has a second, steeper slope available — a **sovereign track** — that the dollar-priced frontier does not. The calculator lets you switch between them.

## 5 · The forward calculator

Take a large enterprise support operation — roughly 1.5 million queries a month. On today’s frontier prices that is about **₹2 crore a year**. Hold the capability fixed and fast-forward on our base case, and it lands near **₹69 lakh in three years** and **₹34 lakh in five** — the same work at a third, then a sixth, of the cost. That is the ROI story; it just needed a calendar. The catch is that the calendar cuts both ways — bolt on agents and pin to the frontier, and the same desk can cost more.

Describe a use-case, set your assumptions, and see the monthly cost projected out to your horizon. The **waterfall** is the point: it decomposes the change from today’s bill to the future bill into its competing pieces — how much the price collapse takes off, and how much the reasoning tax and rising volume add back on. Defaults are our research-backed BASE case; every lever is yours to move.

📊 The AI Cost Curve calculator

Project a use-case 3 or 5 years out. Watch the deltas fight it out in the waterfall.

Use case

Queries / month

Horizon
3 years5 years

Scenario preset
ConservativeBaseAggressive

Per-token price decline / year

Reasoning-token tax / year (more tokens per task)

Usage growth / year (Jevons demand)

Frontier APIIndia-hosted (Sarvam-class)

₹ INR$ USD

Today (Year 0)

—

→

Year 5

—

—

*Today / future bill
**Price collapse (cheaper)
**Reasoning + volume (costlier)
**Sovereign discount

Monthly cost, year by year

## 6 · The leader’s verdict

Three things to take to a budget conversation:

- **Do not extrapolate today’s API bill in a straight line — in either direction.** Pinning to the frontier while bolting on agents can make a fixed capability cost *more* in five years, not less. Down-shifting models aggressively can cut it 80%. The gap between those two futures is a decision you control, not a market you wait on.

- **Budget for tokens-per-task, not just price-per-token.** The price line is falling; your architecture decides whether the task line rises faster.

- **For Indian workloads, price the sovereign track.** INR pricing, subsidised compute and Indic token efficiency are a real, compounding discount — and with Sarvam Code, that track now reaches all the way up to frontier coding agents. It is no longer only a cost story; it is a control story.

Sources: [a16z LLMflation](https://a16z.com/llmflation-llm-inference-cost/); [Epoch AI inference price trends](https://epoch.ai/data-insights/llm-inference-price-trends); [Menlo Ventures State of GenAI 2025](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/); [ARC Prize o3](https://arcprize.org/blog/oai-o3-pub-breakthrough); [Counterpoint](https://www.counterpointresearch.com/en/insights/Memory-Prices-Surge-Up-to-90-From-Q4-2025); [IndiaAI Mission](https://www.digitalindia.gov.in/press_release/indiaai-mission-expands-ai-ecosystem-with-affordable-compute-and-startup-support/); [Analytics India Magazine (Sarvam Epoch)](https://analyticsindiamag.com/ai-features/sarvams-biggest-leap-yet-all-that-happened-at-epoch). Projections are illustrative models built on these rates, not forecasts. Verify against your own workload before budgeting.
