Qualcomm Is Raising Prices the Same Week Anthropic Cut Them
I want to like the “AI is deflationary” story. It’s clean, it’s optimistic, and every model launch this year has come with a chart showing cost per token falling off a cliff. I’ve written versions of that chart myself.
The chart is real. It’s just not the whole curve.
Qualcomm telling customers it’s raising prices by double digits, after exhausting its ability to absorb supplier costs, landed in the same news cycle as Anthropic pricing Opus 5 at roughly half of what Fable 5 costs for comparable capability. Two companies, one week, opposite directions — and the second story got almost all the attention, because it’s the one that flatters the industry’s preferred narrative.
Whose Costs Actually Fell?
Ask what Anthropic’s price cut is actually made of. Inference efficiency, better prompt-cache utilization, a model architecture tuned to do more per FLOP — real engineering, and I don’t doubt it. But inference runs on the same silicon Qualcomm is raising prices on. The compute layer got more expensive at the same moment the API layer got cheaper, which means somebody in the middle absorbed the difference, and it wasn’t Qualcomm.
Qualcomm’s letter is the tell. A supplier doesn’t announce it has exhausted its ability to absorb costs unless the costs have been rising for a while and the absorbing has already been happening quietly. The model layer’s price cuts are being subsidized, in part, by hardware margin compression that hasn’t shown up in a press release until now.
The Bill Nobody’s Circulating Yet
Every developer building on Opus 5’s token pricing is implicitly betting that this arrangement holds — that someone upstream keeps eating the silicon cost so the API price stays flat or falls further. Qualcomm’s letter says that bet is already being called in on at least one part of the stack. When a memory or logic supplier raises prices double digits, that cost has to land somewhere in twelve to eighteen months, and “somewhere” is usually the API price the buyer thought was structurally falling.
Where the Deflation Story Still Holds
None of this makes the AI deflation story false. Model efficiency gains are compounding in a way hardware cost gains aren’t, and a well-run lab really can outpace a supplier’s price hikes for a while through architecture alone. I’ll grant that Anthropic’s cut looks earned on its own terms.
I just don’t think the token price and the chip price can keep moving in opposite directions forever, and somebody is going to be surprised when they stop.