The most consequential shift in AI this year is not a new model or a founder prediction. It is a change in what the machines are actually doing. For years the money, the mythology, and the stock narratives were built around training: enormous one-time runs on GPU clusters to produce a model. That era is ending in plain sight. Gartner's August forecast has worldwide spending on AI-optimized cloud infrastructure reaching $42 billion in 2026, nearly double the $21.5 billion recorded in 2025, and of that total $23.3 billion will flow toward inference while $19 billion goes to training. The crossover, where the cost of running models overtakes the cost of building them, has now happened. If you hold any of the AI complex in your portfolio, this reshapes the thesis more than any headline about artificial general intelligence.
Why agentic workloads change the math
The driver behind the flip is agentic AI, and the mechanics are worth understanding because they explain the sheer scale of the demand. A chatbot answers a question and stops. An agent plans, calls tools, spawns sub-agents, and works through a task over many turns, consuming far more compute per job. One analysis found agentic workloads consume orders of magnitude more tokens, with consumption rates anywhere between four and fifteen times higher. Independent research houses that track this closely put the per-task multiplier even higher.
The important consequence for anyone modeling these companies is that inference is a cost that never switches off. A training run ends; a fleet of agents running inside an enterprise runs continuously, and the bill scales with usage rather than with a periodic capital decision. That is a structurally different demand curve, and it is only starting to show up in guidance. Goldman Sachs projects that global token usage will grow twenty-four-fold between 2026 and 2030, reaching 120 quadrillion tokens per month. Even discounting for the usual enthusiasm in these forecasts, the direction is not in doubt.
The CUDA moat looks different in an inference world
This is where the competitive picture gets more interesting than the standard "Nvidia wins everything" story. Nvidia's dominance was built on CUDA, two decades of software that developers know and trust, and that moat remains formidable. Nvidia still holds around 86% of data center GPU revenue in 2026, down from roughly 90% in 2024 as AMD gains ground in inference. The reason the moat matters less at the margins is that inference leans on open frameworks like vLLM and SGLang that abstract away a lot of the hardware-specific code, which is precisely where a challenger can slip in.
On the silicon itself, the gap has narrowed to the point of a genuine contest. AMD's MI355X matches Nvidia's B200 on compute and beats it on memory, at 288GB against 180 to 192GB, which is a real edge for serving large models. Memory is not a vanity number here: a model that needs two Nvidia cards to hold its weights can often fit on a single AMD card, removing a layer of complexity. Yet the detailed benchmarking tells you why this remains a knife fight rather than a rout. Recent testing on realistic agentic coding traffic showed Nvidia's higher-end configurations staying ahead, and after a round of software optimizations in late August, the performance-per-dollar of Nvidia's B200 pulled back ahead of the MI355X. AMD's remaining problem is less about chips than about execution: the software and the unglamorous work of continuous testing, where it still under-invests relative to Nvidia.
The investment read is that this is turning into a market where more than one vendor can win, because the pie is expanding fast enough to accommodate them. A credible second source of supply, plus the custom silicon that cloud providers are building for their own inference, chips away at the assumption that Nvidia captures every incremental dollar. That assumption is baked into a lot of valuations.
So what
Set against all this, Demis Hassabis putting a 50% chance on artificial general intelligence arriving by 2030 is the kind of statement that dominates coverage while changing very little about how to invest this year. His bar is deliberately high, matching the full range of human cognition rather than acing a benchmark, and even he expects one or two more fundamental breakthroughs before it arrives. Founders and allocators are better served watching where compute is actually being spent than debating a coin-flip four years out.
The practical takeaways are unglamorous but concrete. Inference economics now determine which AI businesses have durable margins and which are quietly subsidizing usage they cannot price. The hardware duopoly is loosening at the edges, which matters for anyone treating Nvidia's share as permanent. And the smartest venture money has already moved up the stack: the bet, made by investors like Conviction's Sarah Guo, is that the labs cannot build every application themselves, so the value accrues to whoever turns all this expensive inference into something a customer will pay for. That is the layer where the returns of this cycle will actually be decided.

