Analysis
Nvidia's real moat isn't the chip — it's the physics of inference
Nvidia guided to a $108B quarter — a ~$430B run-rate. It's not a chip story, it's a cost-structure story: LLM inference is memory-bandwidth-bound, batching-driven, and exploding with the reasoning-model shift — and Nvidia owns exactly those layers. A technical read, with the honest bear case.
Technical deep dive · why Nvidia dominates AI compute now, read through the cost of a token · memory-bandwidth-bound inference, NVLink/NVL72, CUDA, the $500B capital moat · with the honest bear case (Jalapeño, 86% idle GPUs)