The model layer is now a price war.
NVIDIA's data center gross margin held at 75.0% in the quarter ended July 26, 2026, even as Meta gave away an open model for $0.10 per million tokens and OpenAI cut its cheapest tier's price 80% (NVIDIA's Q2 FY27 results filed with the SEC; Martin Alderson's "The summer of open weights"). The model layer is in a price war. The infrastructure layer is not. That gap is the story.
Free weights were supposed to commoditize AI and starve the compute buildout. Instead, NVIDIA's Data Center revenue climbed from $51.2 billion to $89.0 billion across three quarters while margins never dropped below 73% (NVIDIA's Q3 FY26 results; NVIDIA's Q4 FY26 results; NVIDIA's Q2 FY27 results). Running an open model at scale still requires the memory, networking, and power that only a concentrated hardware supply chain provides.
What boards need to understand:
-
Open-weight adoption is not a hedge against infrastructure spend. A downloadable model still needs somewhere to run, and that bill has not shrunk. One industry analysis calls this "capex reallocation, not capex extinction" (Kingy.ai's analysis of open-weight economics).
-
Margin is migrating from model providers to infrastructure and distribution. NVIDIA's $12.9 billion purchase of Hugging Face, the platform hosting 3 million-plus models for 18 million developers, moves the company from selling chips to owning the open-model supply chain itself (CNBC's report on the Hugging Face acquisition).
-
NVIDIA is hedging which layer wins. Its equity stakes across labs, clouds, and infrastructure hit $99 billion in July 2026, up from roughly $7 billion a year earlier — a bet that pays out regardless of whether open or closed models dominate (CNBC's report on NVIDIA's equity investments).
Take to the next meeting: Before approving any open-model deployment as a cost-saving measure, finance should price the full infrastructure bill against the API alternative. The model may be free. The silicon underneath it is not.
For practitioners
A team modeling total cost of ownership against a vendor API needs to price GPU-hours, not tokens, since the GPU-hour is what self-hosting actually bills against, whatever happens to the sticker price on a vendor's token.
The token-price collapse is real and worth separating from the infrastructure question. Martin Alderson's account of pricing moves in mid-2026 tracks OpenAI cutting its cheapest tier 80% and its flagship 20%.
The kingy.ai analysis of open-weight capex effects frames the tradeoff correctly: open weights "make capable AI available to more developers, create reasons to self-host, and push spending into inference chips, cloud capacity, memory, networking, power, cooling, security and sovereign infrastructure." Teams evaluating self-hosting should price the full stack, not the model download. NVIDIA's own results back the mechanism: Jensen Huang told investors that "Grace Blackwell with NVLink is the king of inference today, delivering an order-of-magnitude lower cost per token," per NVIDIA's Q4 FY26 results filed with the SEC. The cost-per-token improvement is happening at the hardware layer. A free checkpoint does not change what silicon is required to serve it, and that silicon is the line item that still clears at 75.0% gross margin.
Deep dive
NVIDIA's data center segment did $89.0 billion in the quarter ended July 26, 2026, up 117% from a year earlier, with gross margin holding at 75.0% on both a GAAP and non-GAAP basis, according to NVIDIA's second-quarter fiscal 2027 results filed with the SEC. The timing captures the puzzle at the center of the open-weight boom: the free model and the record margin arrived together, not in sequence.
The stakes are straightforward. If open weights make frontier-quality models a commodity, the assumption embedded in trillions of dollars of AI infrastructure spending — that value concentrates wherever the smartest model lives — stops holding. Where the money then flows determines which companies in the stack are actually mispriced.
The evidence from NVIDIA's last three reported quarters argues against the simple version of the commoditization story. Data center revenue moved from $51.2 billion in the quarter ended October 26, 2025, to $62.3 billion in the quarter ended January 25, 2026, to $89.0 billion in the quarter ended July 26, 2026, according to NVIDIA's third-quarter fiscal 2026 results and fourth-quarter and full-year fiscal 2026 results.
The second-quarter fiscal 2027 filing cited above supplies the most recent of the three figures. Gross margin never dropped below 73.4% across those three quarters. That is not what a business looks like when its core input is becoming free.
The affirmative case is that open weights redirect capital rather than destroy it. Kingy.ai's analysis of the capex debate put it plainly: open-weight availability "will probably not kill the AI capex boom," and is "more likely to change what the money buys and who earns the return," according to Kingy.ai's analysis of open-weight economics and capex. The same analysis noted that Amazon, Alphabet, Microsoft and Meta all raised capital spending after the 2025 disruption caused by DeepSeek's release, and now run larger 2026 programs than before it. A free model still needs somewhere to run, and running a 2.8-trillion-parameter mixture-of-experts model at scale requires memory, networking, power and cooling that a research lab downloading weights from Hugging Face does not already own.
Jensen Huang has started saying this part out loud rather than leaving it implied. On the August earnings call he described "a thriving open-model ecosystem" in the same sentence as "multiple frontier labs scaling in parallel" and "physical AI coming online," framing all three as demand drivers rather than threats, according to the same fiscal 2027 filing. Five months earlier, announcing record fourth-quarter data center revenue, he credited Grace Blackwell with delivering "an order-of-magnitude lower cost per token" and called it "the king of inference today," according to NVIDIA's fourth-quarter and full-year fiscal 2026 results. The claim is specific: the binding constraint on cheap inference is the chip running the model, not the license attached to its weights.
The clearest evidence that margin is compressing somewhere is the pricing behavior of the model labs themselves. Anthropic's flagship model, priced at $10 and $50 per million tokens on its top tier, has struggled to attract users against cheaper alternatives, according to the Financial Times reporting cited in an independent analysis of summer 2026 open-weight pricing. OpenAI cut its cheapest tier 80% and its flagship tier 20% over the same stretch.
That same analysis also noted Anthropic's developer-facing social accounts signaling compute scarcity even as its prices fell relative to competitors. A lab that short on infrastructure is not positioned to be commoditizing anyone.
Nathan Lambert's framing complicates any tidy conclusion. Writing in April 2026, he argued the capability gap between open and closed models remains real, rejecting both the claim that open models "won't keep up" categorically and the opposite claim that they will simply converge, according to Nathan Lambert's analysis of open models at Interconnects.ai. He called the timing of any stable balance unclear, tying it to funding structures, distillation techniques, and regulatory risk to open releases, variables that could move the ledger in either direction and that no single quarter of NVIDIA earnings resolves.
NVIDIA's response to that uncertainty was not to wait it out. It agreed to acquire Hugging Face, the platform hosting more than 3 million models and used by over 18 million developers, for $12.9 billion, according to CNBC's report on the Hugging Face acquisition. Huang wrote that "NVIDIA compute will not be required to build on or deploy through Hugging Face," a disclaimer against lock-in issued by the company simultaneously buying the road everyone uses to reach it.