The Frontier Is Leaking: How Open Models Caught Up in 12 Months
Why the 18-month lead became a 6-month lead — and what the last six weeks revealed about closed frontier valuations.
The Frontier Is Leaking: How Open Models Caught Up in 12 Months
Why the 18-month lead became a 6-month lead — and what the last six weeks revealed about closed frontier valuations.
Part 3 of The Intelligence Engine, a five-part series on how AI’s intelligence layer is being rebuilt — and what that means for the next decade of investment.

On February 23, 2026, Anthropic published a blog post that almost nobody outside the AI security community read carefully.
The post documented what Anthropic called an “industrial-scale distillation attack” against its Claude models. The numbers were extraordinary. Approximately 24,000 fraudulent accounts. Over 16 million exchanges harvested. Three Chinese AI laboratories named explicitly: DeepSeek with 150,000 exchanges, Moonshot with 3.4 million, and MiniMax — the biggest by far — with over 13 million. Anthropic said senior Moonshot staff were identified through request metadata matching public profiles.
The reason this should have made more noise: it is the closest thing we have to a public balance sheet for what closed frontier moats are actually worth in 2026.
And the answer on that balance sheet is — less than the market is pricing.
The disclosure was read two ways. In the US, the dominant frame was national security — proof of coordinated Chinese IP misappropriation, with calls for stronger export controls and access restrictions. In Asia, the dominant frame was market structure — evidence that the moat being defended was already thinner than headline valuations implied. Both frames capture something real. Neither captures what mattered most: what happened in the six weeks that followed.
In the six weeks that followed, frontier labs and open-weight labs took turns moving the goalposts on each other. The pace of release was unlike anything I’ve seen in this industry. By the time you finish this essay, the leaderboard will have rotated again.
This essay is about what those six weeks actually mean for the next phase of the AI investment cycle.
The signal from Part 2
In Part 2 of this series, I argued that the 2020–2024 one-axis scaling regime had split into four games — pre-training scale, post-training RL, inference-time compute, and frontier distillation. I called distillation the fourth axis, and noted that the open-weight Chinese models had caught up to closed US frontier models on coding and agentic benchmarks within a single year. I said that was not a sustainable equilibrium.
Six weeks of releases have now made that argument concrete. The fourth axis is not a theoretical lever. It is producing measurable, market-priced consequences in real time.
What changed in those six weeks is the structure of the moat itself.
What twelve months of catch-up actually looks like
Twelve months ago, the lead from a closed US frontier model over the best open-weight Chinese model — measured on the benchmarks that matter for paid enterprise workloads — was roughly eighteen months. By April 2026, that gap had collapsed to single digits on the headline coding benchmarks, and in some cases reversed.
Three data points tell the story.
On April 7, Z.ai (formerly Zhipu) released GLM-5.1, a 754-billion-parameter open-weight Mixture-of-Experts model under the MIT license. On SWE-Bench Pro — the harder, multi-language coding benchmark that the industry now considers the most credible measurement of real software engineering capability — GLM-5.1 scored 58.4%, ahead of GPT-5.4 (57.7%) and Claude Opus 4.6 (57.3%). It was the first time a Chinese open-weight model had taken the top spot on that benchmark. Z.ai has been on the US Entity List since January 2025, which means GLM-5.1 was trained entirely on Huawei Ascend chips — zero Nvidia silicon in the loop. The symbolism is hard to overstate.
On April 20, Moonshot released Kimi K2.6, a 1-trillion-parameter MoE under a modified MIT license. SWE-Bench Verified at 80.2%, sitting in a tight band with the top closed models. SWE-Bench Pro at 58.6%, edging GLM-5.1 by two-tenths of a point. API pricing of $0.60 input, $2.50 output per million tokens — roughly one-fifth to one-eighth the cost of equivalent closed frontier inference.
And in May, the US government weighed in. The Center for AI Standards and Innovation (CAISI) at NIST published an evaluation of DeepSeek V4 Pro, concluding that the most capable Chinese model they had assessed lagged the US frontier by approximately eight months. The Commerce Secretary framed this as proof of American dominance. Read carefully, it was the opposite. A year earlier, the comparable government framing was that Chinese models were structurally years behind. Eight months — measured on benchmarks that include non-public CAISI tests — is the same gap that exists between any two consecutive Claude releases.
The open-weight leaderboard now rotates every two weeks. Pricing has compressed by a factor of five to eight on equivalent intelligence. And the license is MIT — meaning any startup, any enterprise, any sovereign government can fine-tune and deploy without permission.
This is what the fourth axis produces. Not parity tomorrow, but a moat that compresses faster than annual planning cycles.

Four models, two camps. Open-weight Chinese models have matched the previous frontier generation on coding benchmarks at one-fifth the price. The closed frontier ships its next generation roughly four days before the open-weight equivalent — a release cadence that defines what’s left of the moat.
Two directions of distillation
What gets less attention is that the same technique is running in two directions at once.
The first direction — closed-frontier to open-weight — is what the Anthropic disclosure is about. The second direction is closed-frontier to small. Every major lab now routinely distills its own flagship models into smaller, cheaper, edge-deployable versions. OpenAI ships GPT-5.4 mini and nano. Google ships Gemma in multiple sizes. Microsoft ships Phi. Alibaba ships Qwen small variants. DeepSeek and Kimi both publish distilled versions of their flagships designed to run on consumer hardware.
The two flows look unrelated. They are the same phenomenon.
In both cases, the receiving model captures most of the capability of the originating model at a fraction of the cost. In both cases, the value of being the originator collapses faster than the value of being downstream of one. And in both cases, the practical effect on the application layer is the same: intelligence becomes a commodity that flows wherever it is allowed to flow.
For founders building application-layer products in 2026, this is the structural insight. You are no longer betting on which frontier model wins. You are betting on which intelligence you can deploy, where, at what unit economics, with what data control. The model is not the product anymore. The deployment is.
The six weeks that priced the moat
So how did frontier labs respond? They did not stop scaling. They did the only thing they could do — they accelerated on two axes the open-weight models cannot match easily: release cadence and product integration.
Look at what the closed frontier shipped in the same six-week window.
On approximately April 16, Anthropic released Claude Opus 4.7, scoring 87.6% on SWE-Bench Verified — a 6.8-point jump over Opus 4.6, and roughly seven points ahead of Kimi K2.6, which arrived four days later. On April 23, OpenAI released GPT-5.5 alongside a substantially upgraded Codex. The release notes did not lead with model benchmarks. They led with new capabilities — browser use, computer use, file handling, document workflows. GPT-5.5 in Codex is no longer pitched as “smartest model.” It is pitched as a teammate that completes multi-step work across applications.
And on May 19, Google used its I/O keynote to announce Gemini 3.5 Flash (already shipping), Gemini 3.5 Pro (one month out), Gemini Omni (any input, any output, starting with video), and Antigravity 2.0 — an agent-first development platform with workflow orchestration baked in. Workspace got Gemini Spark. The Gemini app got redesigned.
Now read those three releases together. The benchmark numbers are real but they are no longer the point. The product surface is the point.
This is the new moat. A frontier lab’s value is no longer a function of its model performance in isolation. It is a function of two levers that compound — lead time, and depth of product integration. Lead time is how far ahead the next-generation flagship stays of the best open-weight equivalent. Depth of integration is how embedded the model is in product surfaces the open-weight model cannot easily replicate — IDE plugins, computer-use agents, enterprise workflows, ecosystem distribution.
When both levers compound, the premium holds. Anthropic at $350B+ valuation, OpenAI at $500B+, both make sense if you believe Codex, Claude Code, Operator, Antigravity, and Gemini Enterprise are durable surfaces that capture switching costs the open-weight stack cannot match. When either lever fails, the premium evaporates.
The six weeks were a public stress test of both levers. Lead time held — barely. Product integration accelerated meaningfully. But the open-weight catch-up speed is the variable to watch over the next twelve months. If GLM-5.2 or Kimi K3 closes Opus 4.7’s seven-point gap in six months instead of twelve, the calculus changes.

A timeline of alternating releases. Each move from one camp triggered a response from the other within days. The pattern itself — not any single release — is the structural finding.
Where intelligence flows next
The other flow — frontier to edge — is moving even faster than the headlines suggest.
Twelve months ago, roughly 30% of frontier-class intelligence could be reproduced at meaningful quality on consumer hardware. Today the figure is 40 to 50%. By the end of 2027, I expect 60 to 70% — and that is the threshold that changes everything. Once edge models cross seventy percent of frontier capability, cloud dependency becomes a choice rather than a requirement for a wide class of applications.
The mechanics that drive this are mature now. Quantization to four bits. Knowledge distillation from teacher to student. Architectural pruning. Sparse mixture-of-experts that runs only the relevant experts per token. Combined, they routinely deliver eight to twelve times the parameter efficiency of two years ago.
The deployment tiers are stabilizing into two clear bands. Models in the one-to-three-billion-parameter range — Phi-4, Gemma 3, Qwen 3 small — now run on smartphones with acceptable latency for many agentic tasks. Models in the seven-to-thirty-billion range run on consumer laptops with unified memory architectures. Apple’s M-series, AMD’s Strix Halo, and a coming wave of NPU-equipped Windows laptops are turning the local-first deployment story from a niche into a default.
For founders, this is where the application-layer game gets interesting. You can build a product where the seventy percent of intelligence runs on the device, and only the hardest twenty to thirty percent hits the cloud. The unit economics are different by an order of magnitude. The data sovereignty story writes itself.’
Korea’s coordinates
Korea cannot win the frontier race directly. The capital gap is ten to twenty times. But Korea has spent the last year making a different bet, and the early returns are interesting.
In August 2025, the Korean government selected five national champions for its Sovereign AI Foundation Model Project — Naver Cloud, Upstage, SK Telecom, NC AI, and LG AI Research. The goal was explicit. Build foundation models trained from scratch with Korean data and Korean architectures, not fine-tunes of US or Chinese open weights. Roughly a year of competition followed, with substantial GPU compute provided by the government.
In January 2026, the first stage evaluation came back with a surprise. LG AI Research (EXAONE), SK Telecom (A.X), and Upstage (Solar) advanced to the next round. Naver Cloud was eliminated. The reason was technical: Naver’s submission had fine-tuned a foreign open-weight base, and the government’s “sovereign” definition required full from-scratch training with weight initialization. The Korean state had made a deliberate, unambiguous choice — Korean AI sovereignty means owning the weights, not just the surface.
In February, the government added a fourth slot. Motif Technologies — a one-year-old startup spun out of GPU infrastructure firm Moreh — won it. Motif’s 12.7-billion-parameter model, built from scratch in seven weeks, had topped the Korean ranking on the Artificial Analysis Intelligence Index in December 2025, outperforming Mistral Large 3 at 675 billion parameters. A model fifty-three times smaller, outranking on a global index. The Korean lesson is becoming clear: when you cannot win on scale, win on architecture and efficiency.
In April 2026, Krafton — the company behind PUBG — launched Raon, a four-model open-source release on Hugging Face. Raon-Speech, their 9-billion-parameter speech LLM, ranked first globally among open-source speech models under 10 billion parameters in both English and Korean according to their own evaluation. A gaming company shipping foundation models on day zero with global-class performance is a sentence that would not have been writable two years ago.
The Korean playbook now reads: domain depth plus Korean-language strength plus industrial vertical fit. Not frontier parity. Domain dominance. That is a defensible thesis for a country with ten to twenty times less capital — and the Sovereign AI program is the testbed.
Three lessons for founders and investors
After watching the last six weeks closely, three things are clear.
One. Closed frontier valuations are no longer priced primarily on model performance. They are priced on the joint function of lead time and product integration depth. If you are evaluating an exposure to OpenAI, Anthropic, or Google’s AI franchise, the question to underwrite is not “Is GPT-5.5 better than Kimi K2.6?” The question is “How long can the integration surface stay ahead of the open-weight stack?” If the answer is multiple product cycles, the premium holds. If the answer is one cycle, it does not.
Two. Inference unit cost has a floor that keeps falling. Open-weight models at one-fifth to one-eighth the price of closed frontier set a structural ceiling on what closed-model inference can charge for equivalent intelligence. This compresses gross margins at the application layer in a permanent way — and simultaneously opens unit economics for a class of vertical AI products that were not viable at 2024 inference prices.
Three. The application layer is the unblocked lane. Vertical AI, industry-specific copilots, regulated-data products, on-device intelligence — these are where the next decade of returns concentrate. The intelligence is now a commodity input. The differentiation is everything around it: data rights, workflow embedding, distribution, trust, regulatory positioning. Founders who understand this stop building general-purpose AI products. They build deeply specific ones.
What comes next
This essay has been about intelligence flowing horizontally — across borders, across model sizes, across pricing tiers. Part 4 of this series is about intelligence flowing into something else entirely. Bodies.
The same scaling logic that produced four axes for language models is now producing a fifth axis for action. Vision-Language-Action models. Robotic dexterity. A scaling law for hands. And the geography of that next axis — manufacturing, components, embodied data — runs straight through Korea, Japan, and China in a way the language model wars never did.
Part 4 of The Intelligence Engine — “The Fifth Scaling Law: When Robot Dexterity Became Predictable” — coming next.
The frontier is leaking. The intelligence is flowing. The next thing it picks up is a hand.
Jihoon Jeong (JJ) is Founding Partner at Asia2G Capital, with 150+ startup investments across AI and deep tech in Asia. He advises Samsung Electronics, SK Hynix, and Hyundai Motor Group on AI strategy. Follow him on X @hiconcep and LinkedIn.
메타데이터
- post_id
- efa04cde4fff
- slug
- the-frontier-is-leaking-how-open-models-caught-up-in-12-months-efa04cde4fff
- url
- https://medium.com/@hiconcep/the-frontier-is-leaking-how-open-models-caught-up-in-12-months-efa04cde4fff
- canonical_url
- https://medium.com/@hiconcep/the-frontier-is-leaking-how-open-models-caught-up-in-12-months-efa04cde4fff
- author_url
- https://medium.com/@hiconcep
- status
- ok
- fetched_at
- 2026-06-09 14:34:10