Enterprise LLM API spending hit $8.4 billion in mid-2025, more than double the $3.5 billion figure from late 2024 Source. That money is going overwhelmingly to three companies: OpenAI, Anthropic, and Google. The hyperscalers — AWS, Azure, GCP — take their cut on the compute side. The model providers take their cut on the API side. Everyone is happy as long as inference flows through their pipes. The business model depends on it.
Open-weight inference breaks that model at the economic joint. A self-hosted 7B-parameter Llama model running on dedicated infrastructure behind a corporate firewall costs roughly 2% to 10% of what the frontier API charges for equivalent throughput Source. The pricing landscape now spans a 600x range, from $0.10 per million tokens for the cheapest self-hosted deployments to $60+ per million tokens for top-tier frontier models with reasoning Source. When the cost delta is that wide, the enterprise procurement conversation stops being "which model is best?" and starts being "why are we paying 50x more for something we can run on a server we already own?"
The TCO math is not subtle. A 12-month total cost of ownership model comparing local LLMs against cloud APIs shows that the break-even for self-hosted infrastructure happens at surprisingly low inference volumes — and beyond that point, every additional token is pure savings Source. The hardware is a sunk cost. The inference is free at the margin. Every fine-tuned 7B model humming away inside a Fortune 500 company's data center is generating tokens that will never appear on an OpenAI invoice.
Companies spent $37 billion on generative AI in 2025, up from $11.5 billion in 2024 — a 3.2x increase Source. A meaningful fraction of that spend is inference that could run locally. The question is not whether some enterprises will shift inference to self-hosted open-weight models. The question is how much of that $37 billion never materializes as API revenue in the first place because procurement teams ran the numbers and bought a GPU cluster instead.
Here's the colloquial version for anyone still pricing per-token API calls as if they're going to own the inference layer forever: every enterprise that deploys a fine-tuned open-weight model behind its own firewall is inference revenue that permanently exits the closed-model economy. It's not deferred revenue. It's not trial usage that will convert. It's gone. The model is downloaded, the weights are loaded, the inference is running on hardware the enterprise already owns, and the only bill that arrives at the end of the month is the electricity bill. Multiply that by the Fortune 500. Now multiply it by the Global 2000. The hyperscalers' AI revenue growth projections were not built for this.
The AI inference market was valued at $103.73 billion in 2025 and is projected to reach $312.64 billion by 2032 Source. The edge AI market specifically — models running on-device and on-premise — was valued at $24.9 billion in 2025 and projected to grow at 21.7% CAGR to $118.7 billion by 2033 Source. The inference market is enormous and growing. The question is who captures it — and the answer increasingly depends on where the inference actually runs.
The AI value chain is splitting into two distinct businesses. Business one: frontier training. This requires hundreds of millions of dollars in GPU clusters. Epoch AI's data shows training costs at the frontier have risen from roughly $2 million for GPT-3 to nearly $390 million for the largest runs in 2024 Source. This business is concentrated, capital-intensive, and likely to stay that way. The barriers to entry are the cost of a small country's GDP in compute.
Business two: frontier inference. This requires a model download, a decent GPU, and the willingness to manage your own infrastructure. Open-weight models running on-premise, on-device, and at the edge capture inference volume at 2% to 10% of the cost of API calls. Every enterprise that goes this route permanently removes itself from the market for closed-model inference services. The training business remains concentrated. The inference business is democratizing at the speed of Hugging Face downloads.
Open-weight models held just 11% of the LLM API market as of recent measurement Source. That number is going up — not because open-weight models are better at everything, but because the economics are inarguable for the 80% of enterprise AI use cases that do not require a frontier model's reasoning capacity. Summarization, classification, extraction, routing, basic Q&A — all run perfectly well on a 7B model fine-tuned on internal data at roughly one-fiftieth the per-token cost.
The hyperscalers see this coming. Azure markets Phi models as Azure services even though the models are small enough to run locally. AWS offers Bedrock with open-weight model options that still route inference through AWS infrastructure. The cloud vendors want a piece of every inference call no matter where the model came from. But the fundamental physics of inference cost — the marginal cost of running a model you already downloaded, on hardware you already own — favors local deployment for the vast middle of enterprise AI workloads.
The AI value chain is not consolidating. It is splitting, and the split runs straight through the inference layer. The money in training is concentrated. The money in inference is fragmenting. If you're building a business model that depends on capturing both, you might want to build a better moat around the one that anyone with a GPU can do for free.