Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced SemiAnalysis’s August 21 study, “Are Open Models Catching Up?”. Its answer is yes—but in a more interesting sense than a leaderboard snapshot suggests. Open models, many of them more precisely described as open-weight models, have repeatedly caught an earlier proprietary frontier, and SemiAnalysis finds that the delay has shortened with each major change in what language models can do.
The frontier moves in eras
The study rejects the idea that model progress can be measured on one timeless scoreboard. Benchmarks are built to separate the strongest systems of a particular moment. Once models saturate those tests, researchers invent harder ones that reflect a new capability frontier.
SemiAnalysis therefore divides modern language models into three eras. The early-scaling era emphasized knowledge, short math problems, and small coding tasks. The reasoning era shifted attention toward difficult mathematics and expert questions. The current agentic era asks whether a model can use a terminal, browse, write and repair software, operate tools, and sustain work across many steps.
This framing produces a recurring cycle. A proprietary lab discovers and productizes a new technique, opening a capability gap. Other labs study the result, reproduce the important advances, distill behavior from accessible models, and eventually close that gap. The leader then changes the game again.
In the early-scaling comparison, SemiAnalysis gives GPT-3.5 Turbo a normalized composite score of 75.7 and Llama 2 70B a score of 39.9. Llama 3.1 405B eventually exceeded that earlier GPT-3.5 reference, while DeepSeek V3 nearly matched GPT-4o by the end of 2024. In the reasoning era, the initial open-versus-closed gap was much smaller: 12.1 points rather than 35.8, and a later DeepSeek R1 checkpoint closed it in about 8.5 months.
The agentic era compressed the delay again. Using a composite of Terminal-Bench 2.1, BrowseComp-Plus, a banking-agent benchmark, and DeepSWE, the study says Kimi K2.6 surpassed its Opus 4.5 reference after 4.8 months and GLM-5.2 passed its GPT-5.2 reference after six months. The exact values depend on the chosen tests, but the direction is striking: open models are learning the frontier’s lessons faster.
What “caught up” actually means
SemiAnalysis built separate composites for each era, normalized results within them, and weighted their component benchmarks equally. It ran most evaluations using Prime Intellect’s stack, supplemented by Artificial Analysis and DeepSWE leaderboard results. Open models were served with the software, hardware, and sampling settings available near release; proprietary models used pinned API versions.
That is a serious attempt at historical comparison, but it does not establish one universal measure of intelligence. Choosing the representative models and benchmarks is partly subjective. Some selected benchmarks have known quality problems, and public tests invite model makers to train reinforcement-learning environments that resemble the evaluation. A high score can reflect real capability, deliberate hill-climbing, or both.
More importantly, catching a reference model is not the same as catching the current frontier. By the time an open model reproduces one generation’s advantage, the leading labs may have moved to a new interface, training method, or class of task. The study is best read as evidence for a shortening imitation cycle, not proof that the race has ended.
Release cadence makes that distinction sharper. SemiAnalysis calculates that leading proprietary labs released models about every 213 days in the early-scaling era, 120 days in the reasoning era, and 51 days in the agentic era. Open models are catching up faster while the closed frontier is also moving faster.
A model is not the whole product
The article supplies its own strongest warning against treating benchmarks as purchasing advice. Kimi K3 scores above Fable 5 on its curated composite, yet the SemiAnalysis team still prefers Fable for daily work. The difference comes partly from model behavior that benchmarks miss and partly from the surrounding product: the coding harness, tools, context management, and interaction design.
This explains why Anthropic could build a strong position without dominating the reasoning-era benchmark charts. Claude Code made the model useful as part of a working system. In agentic software, reliability across a long task, recovery from mistakes, tool orchestration, and the quality of the feedback loop can matter more than a few composite-score points.
The open-model advance is nevertheless strategically important. A shorter catch-up window gives developers more credible alternatives, puts pressure on inference prices, and weakens any moat based only on today’s weights. It also lets organizations trade some frontier performance for control over deployment, customization, and data handling.
For proprietary labs, the result is not necessarily a death sentence. It shifts the defensible layer upward—from a temporary benchmark lead toward products, distribution, safety work, developer ecosystems, and the ability to define the next capability era. The durable lesson is that model makers cannot assume a technical lead will remain scarce. The lead must be converted into a system people value before the rest of the field learns how it works.