Jensen Huang spent two hours at GTC 2026 talking about watts.
Not FLOPS. Not parameters. Not training benchmarks. Watts — and the rate at which they convert into tokens. For anyone tracking the trajectory of enterprise AI, that inversion of Nvidia’s two-decade messaging hierarchy tells you more than the $1 trillion demand pipeline he projected through 2027. Let’s see why.
Focus On: Why Inference Speed Is the New Binding Constraint
The thesis Huang presented is very simple. Data centres are power-constrained systems. Within a fixed energy envelope, the only variable that matters is how many useful tokens you extract per watt consumed. He even offered a formula for the C-suite: Revenue = (Tokens per Watt) × (Available Gigawatts). It reads like a manufacturing equation — because that’s precisely what it is. Huang now describes data centres as “token factories,” and once you adopt that framing, every infrastructure decision becomes an optimisation problem with a single dependent variable: token velocity.
The numbers highlight the urgency. Nvidia’s forthcoming Vera Rubin platform generates 700 million tokens per second in the same power envelope where Blackwell manages 22 million. The Groq 3 LPX rack — the first product from Nvidia’s $20 billion acquisition of inference specialist Groq — delivers 35x more throughput per megawatt than its predecessor, targeting 1,500 tokens per second for agentic workloads. Nvidia’s own head of AI infrastructure said it clearly: CPUs are “becoming the bottleneck” in agentic workflows. Not GPUs. CPUs — the orchestration layer that coordinates agents, manages memory, and moves data between reasoning steps.
This matters to enterprise leaders beyond infrastructure teams because it redefines the economics of AI deployment. Training was a one-time capital expenditure. Inference is continuous, compounding with every agent deployed, every reasoning chain executed. Huang laid out a token pricing spectrum — from free tier through to $150 per million tokens at the premium end — and that spectrum maps directly to the quality of intelligence your organisation can afford to deploy. Slow inference becomes a revenue ceiling.
What This Means for Your Organisation
I’ve written previously about the transition from marketing automation to marketing autonomy — the shift from systems that execute predefined workflows to agents that reason, decide, and act within bounded parameters. Huang’s GTC presentation supplies the infrastructure economics that make that transition either feasible or prohibitively expensive. If your agentic AI deployments depend on commodity inference — and most enterprise pilots today do — you are building on an architecture that won’t support the workloads you’re planning for 2027.
The implication isn’t that every CMO needs to become a chip architect. It’s that the gap between organisations running agents on premium inference infrastructure and those running on the equivalent of economy class will produce qualitatively different capabilities. Real-time personalisation at scale, multi-agent orchestration across campaign and commercial workflows, autonomous research and competitive intelligence — these require token velocity that current commodity infrastructure cannot deliver within acceptable latency and cost parameters. As one technical analysis of GTC said, the information hierarchy has inverted: performance per watt got the stage, FLOPS got a passing mention.
Huang’s formula — Revenue = Tokens per Watt × Gigawatts — deserves a place in your next technology strategy review. Not as an abstraction, but as a forcing function. Your CFO will eventually ask what your AI agents cost per unit of output. When that conversation arrives, the answer had better be denominated in tokens.
If your AI agents’ response time were measured and priced like a utility — tokens per second, cost per million — would your current infrastructure strategy survive the audit?
Nvidia is building entirely separate chip architectures just for inference speed. What is your organisation doing to ensure its agentic investments don’t hit a latency wall before they hit an ROI?
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
