Every release of a strong open Chinese model comes with the headline "caught up with the frontier," and with Kimi K3's release the story repeats. But the independent ikot.blog review shows a gap between paper parity and reliability in real work. Across 32 benchmarks published after release, DeepSeek V4 Pro did not win a single comparison against frontier models: median gap -14.8 pp, and -23 pp on agentic coding on average.
Paper Parity and the Real Gap: What DeepSeek V4 Pro Teaches
For mid-sized and large businesses, the takeaway is not that "Chinese models are bad"
, but an engineering one: where errors in production accumulate, the gap is real, and the model should be chosen based on data and the cost of mistakes, not hype. All numbers below come from the ikot.blog analysis; we unpack its findings and turn them into practical model selection.



