The illusion of parity: open AI models vs the frontier

Open Chinese LLMs lag frontier models on new benchmarks. In light of Kimi K3: where parity is paper-thin, where it is real, and how to choose a model for the task.

  • Paper Parity and the Real Gap: What DeepSeek V4 Pro Teaches
  • Precedent: Kimi K2 and the hype that "the number 1 model is now open"
  • 32 new benchmarks after release: zero wins
  • DeepSeek V4 Pro vs the frontier: three slices

Paper Parity and the Real Gap: What DeepSeek V4 Pro Teaches

Every release of a strong open Chinese model comes with the headline "caught up with the frontier," and with Kimi K3's release the story repeats. But the independent ikot.blog review shows a gap between paper parity and reliability in real work. Across 32 benchmarks published after release, DeepSeek V4 Pro did not win a single comparison against frontier models: median gap -14.8 pp, and -23 pp on agentic coding on average.

For mid-sized and large businesses, the takeaway is not that "Chinese models are bad"

, but an engineering one: where errors in production accumulate, the gap is real, and the model should be chosen based on data and the cost of mistakes, not hype. All numbers below come from the ikot.blog analysis; we unpack its findings and turn them into practical model selection.

Precedent: Kimi K2 and the hype that "the number 1 model is now open"

This has happened before

About 3.5 months before DeepSeek V4 Pro, Kimi K2 Thinking was released with the same narrative: the strongest model is now open and Chinese.

Three months later, according to the same author's dataset, on 16 shared benchmarks it lost to the frontier (GPT-5 and Sonnet 4.5) on 13, and on seven of them by more than 10 pp.

The pattern is consistent:

  • on release benchmarks, the picture looks like parity
  • and independent benchmarks
  • that pile up afterward
  • show systematic average lag

One striking chart in an announcement is not the same as the distribution of results across dozens of tasks gathered later and not tailored to the model.

32 new benchmarks after release: zero wins

DeepSeek V4 Pro (Preview) was released around April 24

To exclude tuning against known tests, the author took 32 benchmarks published after release and compared the model with the frontier at that same moment - GPT-5.4, GPT-5.5, Opus 4.6, and Opus 4.7. The result across the whole set was zero wins out of

The median gap to the frontier is -14.8 pp, the average is -18.6 pp, and on agentic coding benchmarks (SWE-bench-like) the gap is largest - -23 pp on average. Below are three comparison slices from this set; all numbers come from the ikot.blog review.

0 / 32wins by DeepSeek V4 Pro against all frontier models on 32 new benchmarks (ikot.blog review)
-14.8 ppmedian gap to the frontier; -18.6 pp on average
-23 ppaverage gap on agentic coding - SWE-bench-like benchmarks

DeepSeek V4 Pro vs the frontier: three slices

Comparison sliceBenchmarksWinsLossesAverage gap, pp
vs all frontier models (GPT-5.4/5.5, Opus 4.6/4.7)32032−18,6
vs Opus 4.61019−9,1
vs GPT-5.416313−6,9
Open AI Models vs Frontier: The Parity IllusionHorizontal bars show the average gap in percentage points: versus all frontier models, -18.6; versus Opus 4.6, -9.1; versus GPT-5.4, -6.9. The longer the bar, the greater the open model's lag. Data from the ikot.blog review.Average gap to the frontier, ppthe longer the bar, the greater the open model's lagvs all frontier models−18,6vs Opus 4.6−9,1vs GPT-5.4−6,9
Average gap of DeepSeek V4 Pro versus frontier models across three slices, in percentage points. The longer the bar, the greater the open model's lag. According to the ikot.blog review based on 32 benchmarks after release.

Assess where AI can deliver impact in your process

The honest nuance: where parity holds

It is important not to turn this into the claim that "Chinese models are trash."

The author of the analysis explicitly asks not to read it that way, and they are right

This is about systematic average lag, not every single task. In some subdomains, parity really holds: frontend layout, three.js and 3D in the web, web games, and probably presentation generation and Excel scenarios.

There, the average benchmark gap does not translate into a noticeable difference on the real task. The key formula is "paper parity -> real gap"

: the longer the task horizon and the costlier the mistake, the further frontier models pull ahead; the shorter and more visual the task, the closer the open model.

Parity by domain

Where parity really holds

  • Frontend layout and UI components.
  • three.js and 3D graphics on the web.
  • Web games and interactive graphics.
  • Probably presentation generation and Excel scenarios.

Where it systematically lags

  • Agentic coding (SWE-bench-like) - on average -23 pp.
  • Long autonomous tasks where errors accumulate.
  • Reliability in production - paper parity does not carry over.

Forecast: will Kimi K3 repeat the pattern

  1. In light of Kimi K3's release, the author cautiously predicts the pattern will repeat against the new frontier (GPT-5.6-Sol and Fable 5), but in a milder form - roughly 8 wins, 5 ties, and 17 losses out of

  2. In other words, the lag will remain on average but become smaller, and K3 may well lock in certain domains where open models no longer lag systematically.

  3. This is a forecast, not a fact, and it should be treated as a hypothesis.

  4. But the direction matches the two previous iterations (K2, DeepSeek V4 Pro), and for planning that matters more than the exact number: release hype about being "top-1" should be checked against independent benchmarks published later, not accepted from the announcement.

How to choose a model for the task

The practical business takeaway is not "use only frontier" and not "use open, it has caught up." Model choice is an engineering decision for a specific task, and it is made along two axes: the cost of error and the task horizon. Where errors accumulate and are expensive, frontier models work; where the task is short, visual, and cheap in terms of error cost, open and inexpensive models win.

This is exactly the logic we discuss in the piece about choosing an LLM for process and budget and in the review Claude in Mid-Market and Enterprise: do not chase a leaderboard rank; instead, calculate where higher accuracy justifies a more expensive model.

Three layers of model selection

Frontier - where mistakes are expensive

Agents that touch production; GDPR-bound environments; long tasks where errors accumulate; final synthesis. Here the gap is real - use Opus, GPT, or Gemini.

Open models - in suitable domains

Frontend, 3D web, helper code, drafts. There parity holds, and the cost of error is low - the open model saves money without losing quality.

Cheap layers - at scale

Ingest, summarization, classification, draft replies. Cheap on DeepSeek or Flash models, with verification by a deterministic gate or a frontier model at the output.

How we do it

  1. We do not choose models by hype, and we live by that logic ourselves.

  2. Our client concierge drafts replies on DeepSeek - V4 Flash for short confirmations, V4 Pro for commercial, technical, and long conversations.

  3. But sending is allowed by a deterministic policy engine, and lead qualification is handled by a separate classifier: the cheap model handles volume, while correctness comes from the predictable system, not the model itself.

  4. Our content machine is set up the other way around where mistakes can damage reputation: the final article synthesis in brand voice runs on Claude, while cheap signal collection and normalization run on simpler models. And our public AI calculator explicitly offers clients a hybrid "Frontier + DeepSeek" setup: about 60% of simple tasks go to DeepSeek, while complex ones stay on the frontier.

  5. The same engineering foundation - in our service AI-native development.

  6. This very site is maintained by an AI agent over the repository; we explain how such a setup works in our materials about AI-SDLC on an existing stack and developer and PM levels.

AI for the task

We will choose a model stack for your processes

We will break down your processes, calculate the cost of mistakes for each, and assemble a model stack: frontier where correctness is critical, open and inexpensive where it is cost-effective without losing quality.

Discuss implementing AI for your task

Sources

Sources checked on: 19.07.2026.

The analysis and all numerical comparisons come from the ikot.blog review. This article is our commentary on that review. Original analysis: "the illusion of parity"

and a set of 32 benchmarks released after DeepSeek V4 Pro's launch (zero wins, median -14.8 pp, average gap -18.6 pp, the Kimi K2 Thinking precedent, and the Kimi K3 forecast): ikot.blog/the-illusion-of-parity A detailed audit of DeepSeek V4 Pro with slice-by-slice comparisons (vs Opus 4.6, vs GPT-5.4, wins on TERMS-Bench / Long-Horizon Terminal-Bench / MLS-Bench, failure on DeepSWE): an interactive analysis with

with charts - deepseek-v4-pro-audit.seeyouall.chatgpt.site All benchmark numbers are taken from the ikot.blog review.

Discuss the article: The Illusion of Parity: Open AI Models...

Send via: