DeepSeek V4 Pro 0813 costs 96 percent less than Claude Opus 4.8,
Hi everyone, I am Nitin Gavhane and In this blog I am looking at DeepSeek’s new V4 Pro 0813 release, because how quietly it landed says…
DeepSeek V4 Pro 0813 costs 96 percent less than Claude Opus 4.8, but the benchmarks tell a messier story
Hi everyone, I am Nitin Gavhane and In this blog I am looking at DeepSeek’s new V4 Pro 0813 release, because how quietly it landed says almost as much as what it can do.
Read this blog Free **here**

What shipped on August 12
DeepSeek’s flagship model spent close to four months in preview after its April 24 debut.

On August 12, the deepseek-v4-pro endpoint in DeepSeek’s own API documentation quietly switched to a new build called DeepSeek-V4-Pro-0813. No launch event, no blog post, and according to Simon Willison, who tracks these releases closely, not even an obvious announcement page. He ended up linking to OpenRouter’s listing instead of anything from DeepSeek itself.
South China Morning Post reported that DeepSeek did post a short statement on its own site that same day, claiming the update brought stronger agent capabilities. The statement was pulled by the following afternoon. Make of that what you will. The pattern echoes what DeepSeek did with V4 Flash, which left preview on July 31 after reportedly outscoring the Pro preview build on the company’s internal coding-agent tests.
The price tag everyone is talking about

Here is the number that will grab attention first. V4 Pro 0813 runs $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 per million output tokens, with a one million token context window and a max output of 384,000 tokens. Set that next to Claude Opus 4.8, priced at $5 per million input tokens and $25 per million output tokens on Anthropic’s standard tier, and V4 Pro 0813 comes out about 91 percent cheaper on input and 96 percent cheaper on output.

That gap is real, not a rounding error. But there is a catch worth flagging before anyone locks in a migration plan. DeepSeek’s own pricing page carries a notice that a significant increase to its API pricing is coming, with no date or number attached yet. Today’s rate is genuine. Next quarter’s is unknown.
What DeepSeek says its own numbers show
The benchmark figures DeepSeek is circulating come from its own model card and internal harness runs, tested at what the company calls V4-Pro-Max, its highest reasoning setting. The headline scores: 80.6 percent on SWE-bench Verified, 90.1 percent on GPQA Diamond, 93.5 percent on LiveCodeBench, 87.5 percent on MMLU-Pro, 67.9 percent on Terminal Bench 2.0, and a Codeforces rating of 3206. On SWE-bench Verified specifically, DeepSeek’s own comparison table places it level with Gemini 3.1 Pro and a hair behind Claude Opus 4.6’s 80.8 percent.

Those numbers come from DeepSeek testing DeepSeek. Plenty of vendors report their own results first, so that alone is not damning, but as of this writing no independent lab has reproduced any of these figures for the 0813 build specifically. Worth remembering before repeating them as settled fact.
What independent testers found
This is where the story gets more interesting. South China Morning Post cited two outside evaluations that paint a rougher picture. On the Artificial Analysis Intelligence Index, V4 Pro 0813 scored 53, putting it about even with Zhipu AI’s GLM-5.2 from June, four points behind the mid tier Terra model in OpenAI’s GPT-5.6 lineup, and seven points behind Moonshot AI’s Kimi K3. On the Vals Index, which scores models across a range of benchmarks, it ranked 12th, trailing even OpenAI’s older GPT-5.5 and landing well behind Kimi K3 and Claude Opus 5.
Vals AI pointed to two specific weak points: completing tasks inside a sandboxed terminal environment, and building complex financial models in Excel. Interesting side note, DeepSeek’s own Terminal Bench 2.0 score already trailed GPT-5.4 by more than seven points before Vals flagged terminal work as an outright soft spot. Two separate tests pointing the same direction carries more weight than either one alone.
Where it held up
None of this makes V4 Pro 0813 a dud. SCMP’s own headline calls out cybersecurity as a bright spot, and Chinese state outlet Global Times reported that DeepSeek’s benchmark charts show strong results on agent focused tests, covering terminal operation, code engineering, tool use, and security related work, with the company claiming it approached or beat some overseas rivals in specific cases. Treat that last claim with the same caution as the rest of DeepSeek’s self-reported numbers, but the cybersecurity strength keeps showing up across separate sources, which makes it easier to believe than a single vendor slide.
The weights are still missing
If you were hoping to run this one yourself, hold that thought. The Hugging Face repositories still host the April preview weights, not the 0813 build. Willison could not confirm whether open weights are coming for this version, though both April’s Pro preview and July’s Flash release shipped under an MIT license, so staying API only for good would break the pattern DeepSeek has followed so far. For what it is worth, the Pro repository has already logged more than 1.4 million downloads on Hugging Face in the past month, and that is on the older preview weights alone.
A pelican tells you more than a benchmark table
One detail from Willison’s writeup stuck with me longer than any of the numbers above. He runs an informal test where he asks each new model to draw a pelican riding a bicycle, and he noticed V4 Pro 0813 produced three visibly different pelicans across its low, medium, and high reasoning settings. He said he had not seen that kind of variation from any other model. Small detail, but it suggests the reasoning effort dial changes how the model approaches the task, not only how fast it answers. Worth testing on your own workload before assuming higher effort just means slower and more careful.
There is also something almost charming about how the actual benchmark numbers reached the public. Willison traced them to DeepSeek’s official WeChat group, from there to a Reddit post that got deleted by moderators for being low effort, and finally to an ASCII art table someone rebuilt on Hacker News. For a company shipping a model with 1.6 trillion parameters, that is an odd way to let the numbers out into the world.
Should you switch
If budget is your deciding factor and your workload sits in coding or agent work, V4 Pro 0813 is hard to ignore on price alone. A 96 percent discount on output tokens against Claude Opus 4.8 adds up fast at volume. But SCMP’s reporting suggests some developers walked away disappointed on both capability and pricing, which tells me the honeymoon period for this release did not last long. The benchmark claims are not yet independently verified, DeepSeek has warned of a price increase without giving a number, and the 0813 weights are still not public. Given all that, I would test this against my own evals before trusting the vendor slides.
Has anyone here run V4 Pro 0813 against real coding or agent workloads yet? I am curious whether the terminal and spreadsheet weak spots Vals AI flagged show up the same way in your testing.
메타데이터
- post_id
- c19e3c11eedb
- slug
- deepseek-v4-pro-0813-costs-96-percent-less-than-claude-opus-4-8-c19e3c11eedb
- url
- https://medium.com/@nitingavhane/deepseek-v4-pro-0813-costs-96-percent-less-than-claude-opus-4-8-c19e3c11eedb
- canonical_url
- https://medium.com/@nitingavhane/deepseek-v4-pro-0813-costs-96-percent-less-than-claude-opus-4-8-c19e3c11eedb
- author_url
- https://medium.com/@nitingavhane
- status
- ok
- fetched_at
- 2026-08-19 16:19:04