Apple's fiscal 2025 capex came in at $12.7B. Amazon, Microsoft, Google, and Meta are combining for roughly $725B in AI capex this year, up 77% year over year. Apple is spending under 2% of what the other four are spending combined. For most of the last two years, that gap has been read as a weakness — proof that Apple is behind on AI.
It isn't. It's the clearest signal yet of a different bet on where inference should happen and who should own the infrastructure underneath it.
Two different cost problems
The other four hyperscalers are solving the same problem twice. First they burn capex on training-scale GPU clusters. Then they have to fill those same clusters with inference demand to make the depreciation math work, because idle GPUs are just a write-down with a fan attached. Industry estimates now put inference at roughly 80% of AI GPU spend — which means most of that $725B is really an inference bill wearing a training-cluster costume.
Apple skipped that problem by never buying the cluster. The M5's neural accelerators are built directly into its GPU cores. The A19 Pro carries the same design into iPhone 17. On-device models are hitting roughly 30 tokens per second with sub-millisecond time-to-first-token, and the marginal cost of a query is close to zero because the silicon shipped a year ago, inside a device someone already paid for.
The GPU spend already happened
That's the real reframe. Apple's GPU spend for inference isn't a line item on next quarter's earnings call — it's already sunk, amortized across 1.5 billion devices that get replaced on a normal upgrade cycle regardless of what AI does. No cluster to keep full. No per-query bill to optimize. The economics of serving a token look completely different when the compute shipped inside a phone instead of a rack.
For anything heavier than what on-device silicon can handle, Apple isn't building a frontier foundation model either — it's licensing Gemini at roughly $1B a year, a fraction of what it costs competitors to train and serve their own frontier-scale systems.
The bet everyone else is making
Meanwhile everyone else is optimizing cost-per-token on infrastructure they own and have to keep utilized. Apple bet that the cost curve stops mattering once you own the endpoint instead of the datacenter. That's not a smaller ambition. It's a different one.