I’ve written previously about the evolution of AI pricing in the short to medium term, and my worry that many companies in a hurry to adopt the technology at today’s prices will find themselves in a bind for a couple of years: they will have encouraged their employees to rely on AI as much as possible, only to see the cost of that usage skyrocket.
I’ve been following this topic closely, so I enjoyed reading Dwarkesh Patel’s 29 July piece on the risk that the price of compute could increase tenfold in the coming years.
As I’ve highlighted in my previous posts, I agree with Dwarkesh’s take that the hard constraints on hyperscalers’ ability to add new compute capacity every year, despite massive investments, mean that demand will soon outstrip supply. Based on the latest earnings calls from Microsoft, Alphabet and others, this is already the case. These constraints are likely to persist for at least a few years: new fabs alone will not solve shortages in memory, power, data-centre capacity, advanced lithography equipment and leading-edge wafer allocation. (See the excellent work by SemiAnalysis on this.)
I struggle, however, to understand Ed Zitron, who argues that hyperscaler spending is fuelling a huge bubble ready to burst. I don’t dismiss the possibility that the amount of capital being deployed is excessive, and that investors’ expectations of future returns are unrealistic. But the investment in data centres and GPUs is justified. The demand is already here — I see it in my day-to-day work and among friends — and will continue to increase rapidly. This infrastructure takes years to plan and build, and the spending needs to happen now.
The returns might still be underwhelming relative to the sums invested, and much of this demand depends on Anthropic and OpenAI’s continued success. AI labs and hyperscalers are clearly betting that they will be able to charge much more in the future, which is one of the oldest business strategies in the world: get your customers hooked on your product and then jack up the prices. Customers may eventually be able to switch models, but those who own the compute infrastructure will have the market cornered.
As highlighted by Dwarkesh, the cost of renting an H100 GPU for a year is still around fifteen times cheaper than employing a software engineer. So there is still plenty of room for AI labs or hyperscalers to increase prices and find customers willing to pay, especially as models get closer to replacing some forms of human work one-for-one.
Patel then argues that the increased cost can be captured as profit either by the AI labs, which can use the lead of their frontier models over competitors as a differentiator, or by the hyperscalers and chipmakers, which will be able to charge more for compute. A combination of both is likely, but I’d argue that Anthropic and OpenAI are in the weaker position. We’ve seen significant competition from smaller labs, including open-weight models — especially out of China — that can come very close to their best.
Politics might be a factor here. It wouldn’t surprise me if the US government discouraged or prohibited US-based companies from relying too heavily on Chinese models, even if they are open-weight. However, Microsoft, Meta and others would still be happy to offer their own proprietary or open-weight models to customers looking for a cheaper alternative to OpenAI or Anthropic.
So where does this leave companies planning their AI strategy and adopting these tools across the board? I don’t think that the risk of price increases is a reason to stop or even slow down current adoption efforts. The biggest barrier to seeing a real ROI from AI deployment in most organisations is the time required to transform the culture, workflows, governance and systems they need to analyse information, make decisions and deliver to customers at the speed AI can enable. This will take years, especially for large organisations. Any business that achieves this transformation before its competitors will find itself in a great position to deploy and reap significant productivity boosts from future models, once the price of compute decreases.
In the meantime, companies should make sure they can precisely measure and compare the ROI of current use cases under consideration. Counterintuitively, this doesn’t necessarily mean relying only on cheaper models, as frontier models might use fewer tokens to achieve the same result, or cost more but produce a significantly better one. This is where the complexity lies: ROI will be measured differently for every company and use case, depending on factors such as the impact on the bottom line, customer experience or satisfaction, time saved and the cost of compute.
I would also focus on using AI to automate high-volume, low-value tasks, as those will rely most of all on the quality of the harness and internal setup. This should make it easier to switch to cheaper or even locally hosted models if necessary.
Finally, a few years of expensive compute might temporarily slow down some businesses’ eagerness to replace part of their headcount with AI. This might buy HR teams and workers a bit more time to reinvent certain jobs, understand where the human touch still matters, and work out how humans and machines should complement each other.