DeepSeek V4-Pro closes the gap with Fable 5, Opus 4.8, and Kimi K3
DeepSeek released V4-Pro on August 12, and across its headline benchmarks the model now trades places with Fable 5, Claude Opus 4.8, Kimi K3, and GLM-5.2 — ahead on some tests, behind on others, with no single model sweeping the board.
How much DeepSeek V4-Pro improved on agentic benchmarks
The jump between this version and its preview is the real story here. DeepSWE, a benchmark built around real software engineering tasks, went from 12.8 to 62.7 — nearly a fivefold increase in a category DeepSeek had previously struggled with. Terminal Bench 2.1 climbed from 72.1 to 87.9, and two other benchmarks improved by similar margins: NL2Repo rose from 38.5 to 61.5, and Cybergym from 52.7 to 83.3, all inside a single release cycle.
Alongside those gains, DeepSeek added a reasoning-depth control with three settings: low for simple tasks, high as the default for everyday agent work, and max for the hardest multi-step problems that need more time and tokens to work through. V4-Pro is built on a 1.6-trillion-parameter mixture-of-experts architecture, the same general design DeepSeek has used across its flagship releases going back to earlier versions.
How close DeepSeek V4-Pro actually gets to the competition
On Terminal Bench 2.1, V4-Pro’s 87.9 sits ahead of Claude Opus 4.8 (85.0) and GLM-5.2 (81.0), trails Kimi K3 (88.3) by a fraction, and lands just 0.1 point behind Fable 5’s 88.0. DeepSWE follows a similar pattern: V4-Pro’s 62.7 beats Opus 4.8 (58.0) and GLM-5.2 (46.2), falls behind Kimi K3 (67.5), and sits 7.3 points behind Fable 5’s 70.0.
Cybergym is where V4-Pro actually edges ahead of Fable 5, 83.3 to 83.1, while also beating Opus 4.8 (78.3) and Kimi K3 (80.0). On NL2Repo, V4-Pro’s 61.5 clears GLM-5.2 (48.9) but falls short of Opus 4.8’s 69.7; neither Kimi K3 nor Fable 5 have a published score for that particular test.
The release also adds native support for OpenAI’s Responses API, though not the full feature set: function tools and server-side web search both work, but image and file inputs aren’t supported, and the endpoint is stateless, so it doesn’t retain conversation history or support background mode. V4-Pro is live now in DeepSeek’s app, on its website in Expert Mode, and through the API under the model ID deepseek-v4-pro.
Beating Opus 4.8 and GLM-5.2 on three separate benchmarks while still trailing Kimi K3 and Fable 5 on the hardest ones is a more honest picture than any single “DeepSeek catches up” headline — the model wins in some rooms and loses in others, same as everyone else at the frontier right now.
Share
SUBSCRIBE TO OUR PRIVATE CASES AND USEFUL TIPS
Subscribe to our newsletter, get only exclusive content and weekly digests, no any spam!
By providing my email, I accept the Privacy Policy.