What “open-weight” means
A model is its weights. That’s really all it is: a few billion numbers, locked in during training, and those numbers are where the capability lives. When people call a model “open-weight,” they mean that weight file is downloadable — you can pull it onto your own machines, run it, change it. Closed models like GPT or Claude keep those numbers behind an API where you never touch them. Fully open-source models go a step past open-weight and hand over the training data and code as well. Open-weight sits in between: the result is open, the recipe stays private. The food version, if it helps: a closed model is a restaurant — you eat what they bring you (Claude Opus 4.8, GPT-5.5). Open-weight is a meal kit — you cook it at home and season it however you like, but the recipe itself isn’t in the box (Kimi K3, Qwen, DeepSeek, Llama). Open source is the recipe, printed and handed over.
Companies started actually using the meal kits
Here’s what shifted this year, and it’s why I started paying attention. Open-weight stopped being a lab curiosity and became something companies will actually put into production. A few things pushed it over the line.
The quality gap got small. CAISI — the U.S. Center for AI Standards and Innovation — puts the best open-weight model, DeepSeek V4 Pro, roughly eight months behind the best closed one. Z.ai’s GLM-5.2 even lands second overall on one frontend-coding benchmark, ahead of every other open model (CSIS). When the gap was two or three years, waiting was the sensible move. Eight months is a different sum.
Price isn’t close, either. A cheap open-weight model like DeepSeek V4 Flash runs about $0.14 per million tokens in, $0.28 out. GPT-5.5 is $5 and $30. Claude Opus 4.8 is $5 and $25. On output that’s north of twenty to one (CSIS Futures Lab, Jul 2026). Vercel sees the same thing through its AI gateway: in July, open-weight models did 36% of all tokens but accounted for only 8.6% of the spending. If your workload is high-volume and can shrug off the occasional miss, that spread is hard to walk past.
And you keep your data. Run the weights on your own hardware and nothing sensitive leaves the building; fine-tune on your own data and what you end up with is yours, not something you rent back every month. Jellyfish, which builds developer tools, found that the top reason companies pick open-weight isn’t even the money — it’s wanting to decide how their own data and operations are handled (TechCrunch). It’s not a free lunch, obviously. Security, compliance, the cleanup when something breaks — all of that moves onto your desk.
So no one is switching over completely. One number tells the whole story. Take coding on that same Vercel gateway: cheap DeepSeek carried about a third of the tokens, but more than 80% of the actual money went to Anthropic’s models. Once you see it, the logic is plain. Big, forgiving, high-volume work goes to the cheap meal kit. Coding and the back-office agents, where a wrong answer turns into a real loss, still go to the model people trust, premium and all. Nobody flips the whole company. They flip one task at a time.
But the money isn’t in the model
Now the part I keep turning over. If the weights are free, where does the money actually come from? There’s a whole stack of answers.
The developers who build the models sell official hosted versions and sign big contracts — Moonshot, in China, is reportedly working out revenue-sharing with all three American hyperscalers at once (Azure, AWS, Google Cloud), per Reuters. For a company valued around $50 billion, the open questions are the split, who gets which data, and how anyone audits the token counts; if it closes, it’s the first revenue-share of that size between a Chinese lab and a big U.S. cloud. Brokers sit a layer out, putting hundreds of models behind one door and charging for the comparison, the calls, the routing, the billing. Cloud and chip vendors turn all that execution into infrastructure revenue — Together AI, which rents inference capacity for Chinese models, expected its $240 million cluster to sell out two or three months before it opened, and AMD and NVIDIA have each stood up 200-plus model repositories on Hugging Face so the open models run on their silicon. And at the top sit the application companies, wrapping generic weights in their own data and workflows. In one a16z survey, 27 of 50 enterprise buyers said they’d rather pay per task than per token — which is precisely the arrangement where quietly dropping in a cheaper open model that does the same job falls straight to the bottom line.
Put it plainly: the model is the bait, and every step between it and the customer is somewhere to put a cash register.
So the acquisition war is over the roads, not the models
Read this half-year’s two big deals with that in mind. Stripe, the payments company, agreed to buy OpenRouter — a model-routing platform — for north of $7 billion (Fortune). OpenRouter is really a switchboard: roughly 400 models from 80 providers, eight million developers plugged in. Stripe lays its own payments and cost tooling over the top, so choosing a model, calling it, and paying for it all collapse into one path — and Stripe keeps the transaction, usage, and spend data running through it. They bought a tollbooth that gets paid no matter which model wins, and they did it without owning a single model.
NVIDIA, for its part, is paying $12.93 billion for Hugging Face: three million models, 200,000 corporate users, the largest model hub there is — which is to say, where developers actually go and the pipes the models move through. One detail is easy to skip past. Back in February, Hugging Face absorbed the llama.cpp team, which stretched its reach all the way down to models running on a laptop or an in-house server. Distribution dropped a rung — from “where developers browse” to “where models actually run” — and that lower rung is exactly what NVIDIA was after.
Line the two deals up and they split cleanly. Stripe took the payment-and-transaction gate; NVIDIA took the developer-and-execution gate. Neither one bought a model. Everything gets more convenient while the gates quietly concentrate, and whoever holds the gate collects the data on what gets chosen, how heavily it’s used, and what it costs — the exact thing you’d want for catching the next model early and setting the price. The dependency slid off the models and onto the paths.
How open is “open,” really
Worth being precise about one thing: “open” means you can download it, not that you can do anything you want with it. Moonshot’s Kimi K3 license is a clean example. Use it, change it, sell it, fine-tune it — all free. But the moment you’re serving its inference to other people and your affiliated revenue clears $20 million a year, you need a separate commercial contract first; and once a service crosses roughly 100 million monthly users (or $20 million a month), you have to show “Kimi K3” on the screen. Stack national policy on top — the U.S. is weighing procurement, trade, and hosting limits on Chinese models — and “downloadable” and “usable” drift even further apart.
“Published” and “actually used” are two different things, too. Cross the 25 most-watched models on Hugging Face against the 25 most-downloaded and exactly one shows up on both lists; only 6.1% of U.S. enterprises had actually touched an open-model platform (HF, Ramp, Aug 2026). Heat on the leaderboard and adoption in the field are not the same picture.
So my own test for whether a thing is really open isn’t “can I download it.” It’s whether my right to keep running a business on it is actually secure.
My conclusion — and the patent question
Where this leaves me: the performance gap closes fast, but the distribution and execution paths only harden as users pile up. The advantage that lasts is holding the path, not owning this year’s model. (And if you’re a country or a company without a frontier model of your own, that’s quietly good news — verification, conversion, and execution are all still open ground.)
Which leaves the question a patent person can’t help asking. If the roads are worth this much, did anyone try to patent the road itself?
Someone did. Model routing — read the prompt, pick the right model, the whole reason OpenRouter exists — is the subject of US Patent 12,314,825 B2, “Prompt routing system and method,” granted May 2025 to a routing startup called Martian. Filed August 2024. So did Stripe just hand over $7 billion to step into a minefield?
You check it like this: set the application next to the granted patent and read the gap. A patent surfaces publicly twice — once as the claims filed (what the company wanted to own) and once as the claims that issued (what the examiner actually let them keep). The difference between the two is what they walked away with. Press releases quote the first version. Almost no one reads the second.
Claim 1 as filed (August 2024)
A method for routing machine learning model prompts, comprising: receiving a prompt; determining a candidate model response score for each of a set of candidate models based on the prompt, using a scoring model; selecting a runtime model from the set of candidate models based on the set of candidate model response scores; and facilitating runtime model response determination for the prompt using the runtime model.
Four lines. Strip the jargon and it says: score the models, pick one. That’s the whole thing. Had it issued as written, more or less anything that looks at a prompt and chooses a model would have been infringing.
Claim 1 as granted (May 2025)
A method for routing machine learning model prompts, comprising: training a scoring model comprising a neural network to predict candidate model response scores; extracting an encoder from the scoring model, wherein the encoder is a strict subset of layers of the scoring model; receiving a prompt; determining a candidate model response score for each of a set of candidate models by: determining a prompt encoding using the encoder extracted from the scoring model; and determining the candidate model scores based on the prompt encoding; selecting a runtime model based on the scores; and facilitating runtime model response determination using the runtime model.
Now read the difference. To push it through, they had to add three things: you have to train the scoring model yourself (borrow someone else’s and you’re clear); you have to pull an encoder out of it, and not just any slice — “a strict subset of layers”; and the scores have to come from the encodings that extracted encoder produces.
In the original application, that encoder business was tucked into dependent claim 3, an optional add-on. The file makes it obvious why it got promoted. The prior art here is thick: ten non-patent references on record — FrugalGPT (Chen et al., 2023), “Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing” (Ding et al., 2024), “Routing to the Expert” (Lu et al., 2023), and more — every one of them about choosing or routing between models on cost and quality. A bare “score and pick” claim was never going to survive that crowd, so the company hauled the specific architecture up out of the dependent claim and into claim 1 just to keep something standing.
What the diff tells you
So nobody owns model routing as an idea. What Martian owns is one particular recipe: carve an encoder out of a trained scoring network and route on its embeddings. Route some other way — hand-written rules, a classifier you trained on its own, simply asking an LLM to choose — and you never touch this claim.
Which is the real reason behind Stripe’s $7 billion. You can’t fence off the road with a patent, so you buy it. Where owning the idea fails, the moat is distribution — those eight million developers already standing at the gate — and that moat came with a $7 billion tag. The patent file quietly confirms, in its own dry language, the thing this whole piece keeps circling: the fight isn’t over who publishes models. It’s over who controls the way in.
Method note
You can run this yourself. Every U.S. patent goes public twice — the pre-grant version (A1) and the granted one (B2). Pull up claim 1 of each on Google Patents and look at what got added in between. Every phrase in the second that isn’t in the first is a place where the examiner pushed and the company gave ground. The references it was measured against — nine patents, ten papers — are printed right on the patent’s front page.
This is technical analysis of public documents, not legal advice, and makes no determination of infringement.









