A number in Mozilla’s new State of Open Source AI report caught my attention this week: 4.4 months.
That is Mozilla’s estimate of the gap between the leading open-weight models and the closed frontier. Epoch AI arrives at roughly the same figure using a different measure.
I would be careful about treating 4.4 months as some new law of AI. Models improve unevenly, benchmarks have their limitations, and the best proprietary systems still hold a meaningful advantage on some difficult professional tasks. Still, four months is an awkward number if you are signing a three-year technology agreement.
It points to a problem that I think many companies are going to run into rather quickly. Our procurement cycles, architectures, operating models, and governance structures are designed to last for years, while the relative advantage of the model sitting inside them may now last for a fraction of that time.
Focus On: How Long Are You Actually Buying the Advantage For?
I have argued for some time that enterprises should avoid building their AI architecture too tightly around one model or provider. Until recently, you could reasonably call that good architectural hygiene: avoid unnecessary lock-in, keep your options open, build for change. The Mozilla numbers make the argument more concrete.
Suppose you select a model today because it is clearly better for an important workload. Six months from now, a cheaper alternative reaches the level of quality you actually require. Perhaps it can also run in an environment you control, offers lower latency, or better suits your privacy requirements. Can you switch?
If the answer involves months of engineering, reworking workflows, repeating governance approvals, retraining people, and renegotiating contracts, then the original decision cost rather more than the model bill suggests. You have attached a long-term dependency to what may have been a short-term advantage.
This is why I increasingly think of frontier intelligence as a depreciating asset. It is not a commodity yet, and for some tasks the frontier still matters enormously, but the useful life of that advantage is getting shorter. That should affect how much dependency we are prepared to build around it.
The Model Is Only Half the Story
Another finding in the report interested me even more. Among the professional developers Mozilla surveyed, 79 per cent use open models and 71 per cent use closed models, and half use both. Yet open models still reach production 12 percentage points less often. According to the report, the issue is increasingly the machinery around the model rather than the raw model itself.
Downloading model weights does not give you an enterprise AI capability. You still need context, permissions, memory, evaluation, security, observability, governance, workflow integration, and a safe way for the system to act. Mozilla uses the word harness for much of this surrounding layer.
I like the term because it describes where the work is moving. The model provides the intelligence; the harness determines what that intelligence can see, remember, and do, and whether anyone should trust the result. Model commoditisation, in other words, is arriving faster than operational commoditisation.
Regular readers may recognise the direction of travel. In The Deployment Layer Is The Moat, I argued that durable advantage was moving into the learning system around the model: your workflows, evaluations, permissions, feedback, and organisational memory. More recently, in Build for the Pattern, Not the Product, I made the case for architectures that assume individual products will change.
What the Mozilla report adds is a sense of speed. If relative model performance can shift materially within months, portability is an elegant design principle with a financial value attached.
Stop Buying “The Best Model”
This is where I think the conversation needs to become more practical for CMOs. Too many AI discussions still begin by asking which model is best, and that is usually the wrong question. I would rather ask what level of intelligence this particular task requires, and what the cheapest reliable way of supplying it looks like.
A three-point benchmark advantage may be hugely valuable if the system is analysing a major investment, interpreting complex regulation, or reasoning through a large body of contradictory evidence. The same three points may have almost no commercial value when the job is producing product copy, summarising customer comments, classifying content, or generating campaign variations.
So I would start measuring AI workloads differently. For every material use case, establish the quality level that is genuinely required, measure the volume of work, and understand what a successful output actually costs. Then calculate what it would cost to move that workload somewhere else.
That final number is usually missing. It matters because once a cheaper model crosses your required threshold for quality, reliability, and governance, you should at least have the option to move.
You may decide not to. Operational simplicity, vendor support, and having a contractual counterparty all have value, and frontier capability certainly does when you need it. But continuing to pay a premium should be a decision, not an architectural accident.
There Is Another Form of Lock-In Coming
There is a complication: the model companies understand these economics too. Mozilla highlights an example in which a third-party harness held a 21.8-point benchmark advantage in May. Roughly eight weeks later, after the model labs had incorporated more of that surrounding capability themselves, the difference had fallen to about three points.
That is revealing. As models become easier to substitute, vendors have a powerful incentive to own more of the layer above them: the agents, memory, evaluations, permissions, workflow state, and orchestration that make the model useful. It is an entirely rational move from where they sit, but it creates a subtler problem for the enterprise.
You may technically be able to replace the model, only to discover that everything which makes it valuable now lives inside the provider’s ecosystem. You will have changed the API and kept the dependency.
So portability has to mean more than keeping several model endpoints available. The durable assets should remain yours: proprietary context, evaluation data, workflow design, governance, integration with the rest of the business, and, perhaps most importantly, the organisational learning accumulated as people discover what these systems can and cannot do.
Four Questions I Would Ask Now
At the next serious AI investment review, I would ask:
Does this workload genuinely require frontier capability?
What measurable threshold would allow us to move it to a cheaper model?
What would changing models actually cost us?
What part of our competitive advantage would remain ours after the model changed?
The fourth is the one I would spend most time on. If the answer is “very little”, the company may be consuming a great deal of artificial intelligence without building much intelligence of its own.
There will be perfectly good reasons to keep paying for the frontier, and I expect that to remain true for a long time. But we should be clearer about what we are buying. The architecture may stay with us for five years. The advantage of the model inside it may not survive the next two quarters.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
