Thought Leaders
Open Models Keep Getting Better. The Economics Are Getting Messier.

AI models are obviously better than they were 18 months ago. But when I started looking deeper, the usual explanations for that progress didn’t add up.
They didn’t improve simply because they got bigger. “Open source is winning” is only half right. And the advice to look at benchmark scores turned out to be much less useful than I expected. That one took me a while to accept.
So I pulled 18 months of data across 46 models from eight labs. I wanted to understand whether intelligence is becoming a commodity, whether open models are catching the frontier, and what companies should measure when choosing the systems they’ll build on.
The key finding was straightforward: a benchmark score isn’t really a property of a model. It reflects the model, the software and context surrounding it, and how much time and computing power it was allowed to use. Publish one number, and you’ve hidden the other two.
The implications, however, are much broader. Companies are comparing individual models when they should be evaluating the complete system that has to perform the work.
Benchmarks Hide the System
When I stopped looking at composite scores and compared models test by test, the gap between leading open-weight and proprietary models wasn’t one number at all.
On SWE-bench Verified, an older test that has appeared in countless model announcements, the gap was 0.2 points. On SWE-bench Pro, a newer test designed around realistic, multilingual software work with a standardized setup, it was 18.2 points.
The apparent model gap depends on the benchmark
| Benchmark | Gap | Status |
|---|---|---|
| SWE-bench Verified | 0.2 pts | Saturated, contamination flagged |
| GPQA Diamond | 2.0 pts | Saturated, ceiling around 92–95% |
| Terminal-Bench 2.1 | 4.9 pts | Live, agentic |
| Humanity’s Last Exam | 9.8 pts | Live, frontier reasoning |
| SWE-bench Pro | 18.2 pts | Live, standardized scaffold |
Source: Jaspreet Singh’s analysis of leading open-weight and proprietary models, July 2026.
The same model can also receive different scores depending on who runs the test. Kimi K3 scored 80.9% on Terminal-Bench 2.1 in an independent evaluation and 88.3% when its maker reported the result. That 7.4-point swing is larger than the open-versus-frontier gap on three of the five benchmarks I analyzed.
A model receives instructions through surrounding software, draws on the context it’s given, uses the tools it can access, and operates within a reasoning budget set by the developer. Change those conditions and the result changes with them.
Public benchmarks are useful for understanding the direction of progress. They’re a poor basis for deciding what will work inside your company.
Agentic Reliability Changes the Math
This becomes more important as AI moves from answering questions to completing a sequence of actions.
Take a workflow with 20 steps, where the AI gets each step right 95% of the time. That sounds reliable. Across the full sequence, however, the workflow succeeds only 36% of the time.
A better model can improve the odds at each step, especially on difficult agentic work. It’s the surrounding engineering that determines whether one bad step ruins the job. Saving progress, verifying outputs, retrying safely, and recovering from failure often separate a convincing demonstration from a dependable product.
Model choice still matters. But the best-scoring model doesn’t automatically produce the most reliable system.
Choose Models by the Work
The data supports a narrower conclusion than either side of the open-versus-proprietary debate usually makes.
Open-weight models are competitive for well-defined, short-horizon work where the result can be measured directly. Classification, extraction, summarization, and structured generation increasingly fall into this category.
Frontier models still earn their premium on longer agentic tasks, instruction adherence under pressure, and work where failure is expensive to detect after the fact. That’s where newer benchmarks show a gap closer to 18 points than zero, and where the safety layer provided by a model vendor carries value.
Companies should select models by task, but they shouldn’t assume switching is free. Context, reasoning, and prompt caches don’t move cleanly between providers. A cheaper model can become more expensive if changing systems forces the work to restart or adds layers of engineering and governance.
Model choice is becoming less permanent. Switching is still an architectural decision.
As Intelligence Gets Cheaper, Judgment Stays Expensive
Competent general-purpose capability is becoming extraordinarily inexpensive. At the top of the curve, however, each additional point of performance costs more than the last. The final nine points of capability cost roughly ten times what everything below them costs.
Most companies don’t need the most powerful model for every request. They do need to know which work deserves it.
Summarizing, extracting, classifying, drafting, and translating are moving toward utility capability. What remains scarce is judgment: knowing which of five plausible answers is correct, noticing the question was wrong, and deciding when to stop and ask a human.
Judgment comes from proprietary context, business rules, workflows, historical knowledge, and the engineering that verifies work and contains failure.
Build the Evaluation No One Else Can Run
For an enterprise, the most valuable benchmark is built from its own work.
Create a private evaluation set with 50 to 200 real tasks from your business. Keep the surrounding system fixed, run the tests repeatedly, and score whether the job was completed correctly rather than how polished the answer sounds.
Apply the same discipline to cost. Cost per token can mislead when one model reasons longer, requires more retries, or creates more work for a human reviewer. Measure cost per accepted outcome instead.
Start with a small number of models. Route work at the beginning of a job rather than changing providers midway through it. Count the security, governance, and operational burden of each additional provider alongside its advertised price.
The market will keep producing new benchmark winners, and companies will keep being tempted to rebuild their AI strategy around each one. A better approach is to choose models for the work, then evaluate the entire system under the conditions where it has to perform.
The benchmark that should drive your architecture is the one only your company can run.












