Open Weights and the New Model Economics
Inkling and DeepSeek made the weights free. The interesting question is who captures the value once the model itself stops being the scarce thing.
The received wisdom on open weights is that they are a cost story: download the model, skip the API meter, save money. That framing is comfortable and mostly wrong. My own read, having watched this for a couple of years now, is that open weights are quietly rearranging where profit sits in the AI stack, and the model layer is the part being hollowed out. When Mira Murati's Thinking Machines released Inkling in mid-July, it did not top any leaderboard, and the company said so openly. That candour is the tell. A well-capitalised US lab shipping a 975-billion-parameter model for free, under a permissive licence, while admitting it is not the strongest thing available, only makes sense if the model is no longer the product. The product is everything wrapped around it.
This piece is about what that shift does to money. Who pays, who gets paid, and what a business should actually do while the ground is still moving.
The model is becoming inventory
Start with the supply side, because it is where the economics break first. A foundation model has almost no durable moat at its own layer. Each release makes the previous one look dated, and the gap between the best closed model and the best open one keeps narrowing. The gap between the best open and best proprietary models has narrowed from roughly 20 to 30 percentage points in 2023 to 5 to 10 points on most evaluations by early 2026. On coding and mathematical reasoning, some open models already lead. When the thing you sell is replaceable every few months and a free version trails by single digits, you do not have pricing power. You have inventory.
Inkling illustrates the point precisely. It is a mixture-of-experts system with 975 billion total parameters, though it only draws on a fraction of that, about 41 billion, for any given task, a common design that keeps very large models faster and cheaper to run. The lab positions it not as a champion but as a base to be customised, released so that people can adapt it to their own work. The strategy is to give away the raw material and monetise the tooling around it, in this case the Tinker fine-tuning platform. That is a very different business from selling tokens.
The Chinese labs got there earlier and more aggressively. DeepSeek in particular reset expectations on price. Where developers route by cost, they route to open weights. At 90% capability parity, open models cost roughly 6 times less per call than closed alternatives. The result is a striking imbalance in actual usage. By mid-2026, Chinese-built open models route roughly 18 trillion tokens weekly, compared to 5.5 trillion for US-built models, a 3:1 ratio. When the cheap, good-enough option carries most of the traffic, the premium option has to justify its markup on something other than raw intelligence.
Owning the weights moves the cost, it does not remove it
The cost-savings pitch runs into a wall that most executives underestimate. Free weights do not mean free operation. As one governance consultant put it plainly, open-weight models do not remove enterprise cost, they move it. The bill simply shifts from a vendor invoice to your own payroll and infrastructure. Organizations inherit responsibility for inference infrastructure, GPU planning, fine-tuning pipelines, evaluation, security patching and perimeter defense, responsibilities typically handled by commercial AI providers. The scarcest input is people. The engineers who can run this well are rare and expensive, which is the part of the sum most spreadsheets leave out.
The economics only favour self-hosting under fairly narrow conditions. Hosted APIs are pure variable cost and remain the sensible default when volume is uncertain, while self-hosting wins on cost only when volume is high, steady, and your team can keep GPUs busy and the stack healthy. The silent killer is utilisation. A GPU you rent by the hour costs the same whether it's serving thousands of requests or sitting idle overnight. Real traffic is bursty, effective utilisation is usually far below the optimistic figure in the break-even model, and that pushes the crossover point much higher than a back-of-envelope estimate suggests.
The reason the math still tempts people is that the customisation cost has genuinely collapsed. Techniques like LoRA and quantisation mean a serious fine-tune of a large model now runs on a single high-end GPU in a night for tens of dollars, not a data centre for weeks. Set the recurring cost of an API meter against a one-off tuning run, and the case for owning a specialised model is stronger than it was even a year ago. But that lever pays off only for predictable, high-volume, well-defined workloads.
The value migrates to control and to the harness
If the model is inventory, value accrues to two things: control over your data, and the software that makes a model useful. Control is the honest reason most regulated buyers move to open weights, not price. For firms in finance, healthcare and the public sector, the ability to run a capable model inside their own perimeter, with data that never leaves, unblocks projects that compliance would otherwise veto. The buyer question through the first half of 2026 shifted from whether open weights could handle anything serious to which family for which workload and which hosting partner. That is the language of procurement, not experimentation, and it signals a market that has matured.
The more interesting money is in the orchestration layer, the scaffolding that turns a set of weights into a working agent. The frontier labs understand this, which is why they pulled the harness in-house. On agentic coding benchmarks, a lab-built harness tuned tightly to its own model beats independent harnesses on the identical model, and that tuning degrades on competitors' weights. It is a real moat, and it is the reason open models rarely top the official leaderboards even when the underlying weights are competitive. Put every model on one neutral scaffold and the ranking gap mostly disappears. So the value is not in the weights and not purely in the harness, but in the tight coupling between the two, which is exactly what a company like Thinking Machines is trying to sell with Tinker.
The public markets are about to price this
All of this lands just as the frontier labs seek public capital, which makes the timing worth watching. Anthropic confidentially filed for IPO on 1 June 2026, targeting an October Nasdaq listing at a $965B valuation. On roughly $47 billion of annualised revenue, that is around twenty times sales, a price that assumes the model layer stays lucrative for years. The number the market will actually fixate on is gross margin, and here the pressure is visible. Anthropic's gross margin is around 40%, with a target of 77% by 2028. Closing that gap requires token pricing to hold up. Open weights push in the opposite direction. When Microsoft reportedly cancelled most of its Claude Code licences after token billing chewed through annual budgets, it was a preview of the discipline enterprise buyers will start applying once agentic usage scales and the per-token meter runs hot.
What it means
My market read is that the closed frontier labs are excellent businesses attached to a fragile pricing model, and the open-weights wave is the mechanism that exposes the fragility. I do not think the frontier evaporates. There will always be a premium tier for the hardest reasoning and the most polished multimodal work. But the vast middle of enterprise workloads, the document processing, retrieval and routine agents, is drifting toward open weights on a neutral harness at a fraction of the cost, and that middle is where most of the token volume lives.
If you run a business deploying AI, the practical move is a layered one that the smarter buyers have already adopted: start on proprietary APIs for speed and validation, migrate the high-volume, well-defined workloads to a fine-tuned open model once usage justifies the operational burden, and keep an API relationship as a quality backstop for the hardest cases. Before self-hosting anything, cost it at your real, bursty utilisation rather than the idealised figure, and price in the engineers, not just the GPUs. The most expensive component is never the model file.
If you are looking at the labs as an investment, treat gross margin as the number that decides everything and treat the open-weights trend as a persistent headwind on it. The interesting equity story may not be the model makers at all but the layers that profit whoever wins: the managed-inference providers, the tuning and orchestration platforms, and the low-margin operating businesses that quietly bank the savings. Value in this market is moving away from renting intelligence and toward owning the workflow that intelligence runs inside. The weights were always going to be the commodity. The question worth your attention is what sits on top of them.

