Open-Weight Models Caught Up. Adoption Fell to 11%
Open-weight models matched proprietary ones on code, but enterprise open-source share fell from 19% to 11%. Why the real decision is not open versus closed but a governed selection problem, and the four-gate framework for licensing, cost, and provenance.
Open-Weight Models Caught Up. Adoption Fell to 11%
Open-weight models had their best year in 2026. Their best variants now rival proprietary systems on coding and much of everyday reasoning. So enterprise adoption should be climbing. It went the other way. Open-source models fell from 19% of enterprise usage in 2024 to 11% in 2025, according to Menlo Ventures. The technology got better and the buyers pulled back. That contradiction is the most useful signal in enterprise AI right now, because it tells you the decision leaders are struggling with is not the one they think it is.
New here? I write a weekly brief for enterprise leaders on the numbers, the frameworks, and the governance questions your roadmap is about to face. Read the latest at luizneto.ai.
Key takeaways
- Open-source enterprise share fell from 19% to 11% in a year.
- Capability is no longer the constraint. The license is.
- Self-hosting only pays off above a real token break-even.
- Open weight is not open source. Read the contract.
- Govern the choice in four gates before you deploy.
On this page
- The adoption paradox
- Open-weight models are not open source, and the difference is a contract
- You are not buying a model. You are signing a lease
- The economics have a break-even, and most workloads sit below it
- The framework, govern the choice in four gates
- Two leaders, two roadmaps
- What to do Monday
- FAQ
The adoption paradox
Start with the number that should not exist. Enterprise open-source model usage did not grow in 2025. It shrank. Menlo Ventures put the drop in plain terms: "Llama remains the most widely adopted open-weight model in the enterprise. But the model's stagnation has contributed to a decline in overall enterprise open-source share from 19% last year to 11% today."
Read that twice. The most adopted open-weight family stalled, and the whole category fell with it. That is not what a technology looks like when it is winning.
There is a quieter lesson underneath the headline. A category whose share moves this much when one vendor slows down was never as diversified as it looked. Enterprises had concentrated their open bets on a single family, so one vendor's release cadence became the whole segment's growth rate. That is a portfolio risk, not a capability problem, and it is the first hint that the real constraint here is structural rather than technical.
Meanwhile the capability story ran the opposite direction. Independent benchmark trackers through mid-2026 show the best open-weight models reaching parity with proprietary systems on coding tasks, and trailing only on the hardest composite reasoning and safety work. On the leaderboard, the distance closed. In the enterprise, the buyers stepped back.
So what happened. Most leaders framed this as a capability race. They watched benchmark scores, waited for open models to catch the frontier, and assumed adoption would follow the numbers. The numbers arrived. Adoption did not.
Here is the reframe. Picking the smartest model is the easy part. Living with the one you picked is the hard part, and that is where open-weight deployments break. A benchmark score tells you what a model can do in a lab. It tells you nothing about whether you can legally ship it, afford to run it at your volume, or prove where its weights came from when a regulator asks. Those three questions decide enterprise adoption. None of them appear on a leaderboard.
This is the same failure mode I described in why enterprise AI programs stall without a model portfolio strategy. The open-versus-closed question is one layer inside that portfolio decision, and it is the layer where the most expensive mistakes hide.
Open-weight models are not open source, and the difference is a contract
The word "open" is doing too much work. Most people use open-weight and open-source as if they mean the same thing. For your legal team, they do not, and the difference is the whole story.
Open source is a legal status you can verify. Open weight is a download you have to read the fine print on. Those are not the same thing. An open-source license, in the sense the Open Source Initiative defines it, grants broad rights with no field-of-use or scale restrictions. An open-weight release just means you can download the parameters. The license attached to them can say almost anything.
Take the most adopted example. Meta's Llama Community License permits commercial use, but it requires a separate license from Meta once the monthly active users of your product cross 700 million, bans using Llama to train a competing model, and mandates "Built with Llama" attribution. It also incorporates Meta's acceptable-use policy by reference, which Meta can update (TechTarget, 2026). A deployment that is compliant today can drift out of compliance as your product grows or as the policy changes underneath you.
The acceptable-use clause deserves its own line, because it turns a static license into a moving one. When a license incorporates a policy by reference and the vendor can revise that policy, you are not signing a fixed contract. You are signing a subscription to whatever the terms become. Legal has to monitor those updates the way they track a critical vendor's terms of service, not file the license once and forget it. Few AI teams have that habit yet, and it is exactly the kind of obligation that surfaces during an audit rather than a demo.
Permissive licenses behave differently. Apache 2.0 and MIT impose no user caps and no field-of-use limits on commercial use. Apache 2.0 adds an explicit patent grant and asks you to document your modifications (Recording Law, 2026). For an enterprise, that predictability is worth more than a few benchmark points.
| License | Commercial use | User / MAU cap | Bans training a competing model | Explicit patent grant |
|---|---|---|---|---|
| Llama Community | Allowed, with conditions | Yes, above 700M MAU | Yes | No |
| Apache 2.0 | Unrestricted | None | No | Yes |
| MIT | Unrestricted | None | No | No |
| Source: TechTarget and Recording Law, 2026 | ||||
There is a governance sting in the tail. Because most open-weight licenses carry use restrictions, the models are typically not OSI-approved open source, so they may not qualify for the open-source exemptions written into regimes like the EU AI Act (QubitTool, 2026). The label you assumed protected you may not.
Every open-weight model choice is a licensing decision wearing a benchmark's clothes. That is the sentence to bring to your next model review. Before anyone celebrates a score, someone in legal should have read the terms. If they have not, the score is not a decision. It is a liability waiting for a lawyer. For the regulatory frame around all of this, see my breakdown of the EU AI Act deadline that did not move to 2027.
You are not buying a model. You are signing a lease
An analogy makes the trade concrete. Think about how you occupy a building.
A closed model behind an API is a furnished rental. You pay predictable rent per token. The landlord maintains the plumbing, patches the roof, and upgrades the appliances. You can give notice and leave. You trade control for convenience, and for a lot of workloads that trade is correct.
An open-weight model is buying the building. You own it outright, which sounds better until you read the deed. You now hold the mortgage, the maintenance, and the security. And you are bound by a zoning code you did not write, the license, which the city can amend after you move in through the acceptable-use policy. Ownership is control and liability at the same time.
Neither is the smart choice in the abstract. A firm that runs enormous, steady volume and needs to control every byte of its data should own the building. A team shipping a feature next quarter should rent. The mistake is treating "own" as automatically more mature than "rent." It is not more mature. It is more responsibility, and responsibility has a cost you pay whether or not you planned for it.
Owning the building also means owning uptime. Production is where impressive models quietly fail, a pattern I traced in why AI agents in production succeed only 56.6% of the time. Most of the failed open-weight deployments I see made the same move. They bought the building because ownership felt like the serious choice, then discovered the mortgage was the token bill and the zoning code was the license. The next section puts a number on the mortgage.
The economics have a break-even, and most workloads sit below it
"Open models are cheaper" is the most repeated and least examined claim in this whole debate. The weights are free. Running them is not, and the total cost of ownership has a break-even that depends almost entirely on your volume.
The OECD modeled this across workload tiers in 2026, and the spread is stark. At roughly one billion tokens a month, self-hosting takes about 30 months to break even against an API. At ten billion tokens a month, the break-even arrives in about two months. At fifty billion, it lands in about one. Same model, same math, wildly different answer depending on how much you actually run.
Hardware follows the same curve. The OECD tiers scale from a single L4 GPU for small workloads under 100 million tokens, to one H100 around a billion, to a small cluster of eight H100s near 50 billion. Buying that hardware is only the start. A GPU sitting idle at low utilization inverts the economics completely, because you paid for capacity you did not use. The API's per-token markup is often cheaper than your own underused cluster.
Put a face on it. A support-automation workload running 300 million tokens a month sits well below the break-even, so a hosted API wins even though the weights are free. Move that same workload to three billion tokens a month across a fleet of agents, keep the hardware busy, and ownership starts to pay inside a year. The model did not change. The volume did, and volume is the variable that decides.
This is why preference and economics point different ways. McKinsey found 40% of enterprise leaders prefer models they can self-host for control over privacy and security. Preference is real. It is also not a budget. Wanting to own the building does not change the mortgage schedule.
The practical rule is simple. Below the break-even, rent. Above it, and only if you can keep utilization high, consider owning. And notice the hidden precondition: self-hosting assumes you already have the governed data and the MLOps discipline to run a model in production, which most organizations overrate in themselves. I made that case in detail in why only 7% of enterprises have AI-ready data.
The framework, govern the choice in four gates
You do not need another benchmark. You need a decision process that runs before the benchmark, so the score is the last input rather than the first. The analysts converge here. Forrester's AI Model Openness Framework, published in April 2026, scores any model, open or closed, on reproducibility, usage rights, and community momentum. Gartner frames the deployment choice as build, buy, or blend, with governance carried through its trust, risk, and security management discipline. Both are telling you the same thing: openness is a spectrum you assess, not a badge you accept.
Here is how I compress that into a decision a team can run in a single meeting. Four gates, in order. A model has to clear each one before it earns the next.
First, governance posture. Does this workload touch regulated data, sovereignty requirements, or a need to air-gap? If yes, you have a bias toward self-hosting an open-weight model, because control of the weights and the data path is the point. If no, this gate stays open and the others decide.
Second, volume and TCO. Run the token math from the section above. Below the break-even, the answer is an API almost regardless of preference. Above it, with utilization you can actually sustain, self-hosting earns a serious look.
Third, license fit. Before a line of integration code, legal reads the license. A monthly-active-user ceiling, a competitive-use ban, a field-of-use restriction, or an EU limitation is a design constraint, not a footnote. This gate has killed more good models than any benchmark.
Fourth, provenance and lifecycle. Can you prove the model's lineage, verify the weights you downloaded, and patch on a schedule you control? Public checkpoints can be altered after release, so treat a downloaded model like any other software dependency. Checksum it, record where it came from, and keep a bill of materials your security team can audit. Forrester calls this reproducibility. Your auditor calls it evidence. If you cannot answer it, you do not own a model. You own a risk with good benchmarks.
Score the candidates against four gates instead of one leaderboard and the field narrows fast. It also stops narrowing to the wrong answer. For the governance questions that sit above this, see the five questions every board should ask about AI governance.
Worth saving? The four-gate check is the part of this most teams skip. If you want the frameworks and numbers behind decisions like this every week, subscribe to the brief at luizneto.ai.
Two leaders, two roadmaps
Watch two leaders make this call and you can see the gates decide the outcome.
Leader A leads with the leaderboard. She picks the highest-scoring open-weight model of the quarter and mandates self-hosting, because owning the stack feels like the mature move. Three things arrive on a delay. The license has a user threshold her growth plan will cross. The GPU cluster she provisioned runs at low utilization, so her per-token cost is higher than the API she rejected. And when an auditor asks where the weights came from and how they were validated, she has a download link and no provenance trail. None of these showed up in the benchmark. All of them showed up in the budget and the risk register.
Leader B leads with the gates. She runs governance posture, TCO, license fit, and provenance first, and the score comes last. The result is not a single model. It is a blend. She self-hosts an open-weight model for the high-volume, sovereignty-sensitive workloads where ownership pays, and she calls a closed API for the low-volume, hard-reasoning tasks where renting is cheaper and better. McKinsey found this multimodel blend is how mature enterprises actually operate, and it is why their programs bend instead of breaking when one vendor or one license changes.
The difference between them was not intelligence or budget. It was sequence. Leader A let the benchmark pick and spent the next year managing what it picked. Leader B let governance pick and spent the year shipping. A blended portfolio needs an operating model to carry it, which I laid out in the enterprise agent control plane.
What to do Monday
The default that survives contact with reality is a blended, gate-governed portfolio. Open weights where volume and sovereignty justify ownership. Closed APIs where reasoning quality and operational simplicity matter more. The four gates decide each workload, and the benchmark is the tiebreaker, never the opener.
| If your workload signal is | Lean toward | Because |
|---|---|---|
| Regulated data, sovereignty, or air-gap | Self-hosted open weight | You control the weights and the data path |
| Below the token break-even, or low utilization | Closed API | The markup beats an underused GPU bill |
| Scale near a license ceiling (for example 700M MAU) | Legal review, then decide | The cap is a design constraint, not a footnote |
| Hardest reasoning, small volume | Closed API | Rent quality you do not run often |
| High steady volume you can keep utilized | Self-hosted open weight | Ownership pays above the break-even |
| A first-pass read. Run the four gates before committing. luizneto.ai analysis, 2026 | ||
Now the honesty an advisor owes you, since I just recommended the harder path. Self-hosting adds a real security and MLOps burden that most organizations underestimate, and a blended stack adds routing and evaluation complexity you will have to build and maintain. The framework does not remove that work. It moves the work to before the commitment instead of after, where it is cheaper to do and cheaper to change your mind.
Three moves for this week. Pull every model already in production and tag each one with its license, its monthly token volume, and whether anyone can prove its provenance. You will find at least one deployment sitting on the wrong side of a gate. Second, put a lawyer in the model-selection meeting, not the model-launch meeting. Third, write your break-even volume down before you price a single GPU, so the math leads the purchase instead of justifying it.
Open weights did not lose enterprise share because they got worse. They lost it because the industry finally started reading the contract. The teams that win the next year will not be the ones running the highest-scoring model. They will be the ones who can answer, for every model they run, three questions a benchmark never asks: can we legally ship it, can we afford it at our volume, and can we prove where it came from. Which of those three can your team answer today?
FAQ
What is an open-weight model?
An open-weight model is one whose trained parameters are available to download and run yourself. The architecture and weights are public. That does not make it open source, because the license attached can still restrict how you use it commercially.
What is the difference between open-weight and open-source?
Open source, in the OSI sense, grants broad rights with no field-of-use or scale limits. Open weight only means the parameters are downloadable. Most open-weight licenses add use restrictions, so they are typically not OSI-approved open source (QubitTool, 2026).
Is Llama actually open source?
No. Meta's Llama Community License permits commercial use but adds restrictions, including a separate license requirement above 700 million monthly active users and a ban on training competing models (TechTarget, 2026). It is open weight, not open source.
Read next. This decision sits inside a bigger one. Start with why enterprise AI programs stall without a model portfolio strategy, and if provenance is your concern, why transparency and explainability matter for enterprises. For the weekly brief on calls like this, subscribe at luizneto.ai.
Are open-weight models cheaper than closed APIs?
Only above a volume break-even. The OECD found self-hosting takes about 30 months to pay off at a billion tokens a month, but about two months at ten billion. Below that, and at low GPU utilization, APIs are usually cheaper.
Should our enterprise self-host its LLM?
Self-host when governance requires control of the data path, your token volume clears the break-even, the license fits your scale, and you can prove provenance. If any gate fails, a closed API or a blend is the safer call.
Do open-weight models qualify for the EU AI Act open-source exemption?
Often not. Because most open-weight licenses carry use restrictions, the models are usually not OSI-approved, so they may fall outside the open-source exemptions in regimes like the EU AI Act. Confirm each model's status with counsel before relying on it.