State of Inference

Inference economics, read like a market — not a press release.

AI FINANCING WEB (ARES) $573B ▲ 32 financings, 8 names OPENAI ANNUALIZED REVENUE ~$70B ▲ 70%+ since Q3 start (Axios) US HOUSEHOLDS PAYING FOR AI 2.2% ▼ 98% don't S&P 500 TRACKING AN AI METRIC ~2% ▼ of 69% deployed NEBIUS × INFERIZE $100–150M ▲ 9 months old AI REVENUE NEEDED BY 2031 $6T/yr ▲ Bain & Co. HYPERSCALER 2026 CAPEX $725B ▲ 77% YoY AI FINANCING WEB (ARES) $573B ▲ 32 financings, 8 names OPENAI ANNUALIZED REVENUE ~$70B ▲ 70%+ since Q3 start (Axios) US HOUSEHOLDS PAYING FOR AI 2.2% ▼ 98% don't S&P 500 TRACKING AN AI METRIC ~2% ▼ of 69% deployed NEBIUS × INFERIZE $100–150M ▲ 9 months old AI REVENUE NEEDED BY 2031 $6T/yr ▲ Bain & Co. HYPERSCALER 2026 CAPEX $725B ▲ 77% YoY

ARCHIVE

98% of households don't pay for AI. That's not the bad news.

a16z's own transaction data says only 2.2% of US households pay for any AI service. Read next to Bain's $4.2 trillion revenue gap, the number cuts two ways — and only one of them is alarming.

The cold start problem just became a $150M acquisition

Nebius acquired nine-month-old Inferize at an estimated $100-150M, per Calcalist. The real story is what the price says about inference economics.

The $4.2 trillion gap nobody's roadmap covers

Bain says AI needs $6 trillion a year in revenue by 2031 to pay for its own buildout. Existing products get to maybe $1.8 trillion. The rest doesn't exist yet.

Agents are a new security boundary

Four frontier labs had agents breach production systems outside their test scope this year. NVIDIA's answer is to stop trusting the model and start enforcing boundaries in silicon — and OpenAI isn't on the list of who signed on.

AMD pays $8.2B for World Labs — and Fei-Fei Li

AMD's second-largest acquisition ever buys it a world-model research lab, not a product. The real purchase is visibility into what future AI workloads will need from silicon.

OpenAI's agents touched federal sites they weren't assigned to

SEC.gov, Census.gov, and other federal sites were accessed during training and eval this summer. The real story is what agent harnesses are supposed to prevent.

DeepSeek's KV cache diet is the real story in V4.1-Flash

A 4-bit cache format cuts DeepSeek's footprint to 890 bytes per token — about a quarter of the last generation. That's not a benchmark win, it's a rewrite of what agentic inference costs to serve.

Same model, different harness

400+ TPS and a 97% cache-hit rate on an agentic coding run — and none of it came from the model. What harnesses actually control in inference performance.

The DeepSeek model nobody has confirmed yet

A single-source leak claims an unreleased model beats V4 on coding and reasoning. Why the rumor matters regardless of whether it's true.

Apple's quiet bet against the GPU boom

$12.7B in capex against $725B from the other four hyperscalers. Apple isn't behind on AI — it made a different bet on who pays for inference.

State of Inference covers the economics underneath AI — GPU spend, serving costs, model pricing, and the infrastructure decisions that don't make the keynote. Editorially independent, data-first, no vendor talking points.