Chinese AI DeepSeek claims performance almost equivalent to Claude
Summarize this article with:

DeepSeek has just reminded the AI ​​industry that a successful model doesn’t have to cost a fortune. Without a press release or spectacular announcement, the Chinese laboratory has discreetly deployed the final version of its flagship model. Behind this update lies a rupture. The performance gap is narrowing, while the price gap is widening. For Web3, developers of autonomous agents and companies facing the increasing cost of AI, this new equation could reshuffle the cards. DeepSeek is no longer just looking to compete, but is attacking where its competitors are most vulnerable, price.

A duel between DeepSeek and Claude for control of the AI.

In brief

  • DeepSeek is rolling out the final commercial release DeepSeek-V4-Pro-0813 without an official announcement, replacing the April preview release.
  • Claude Fable 5 is only 5.3% ahead of the Chinese model on average across nine benchmarks, a superiority which drops to 2.8% when excluding a specific test.
  • Anthropic’s API costs 4,500% more than DeepSeek’s, with an average bill of $30 compared to $0.65 for an equivalent volume.
  • Running an agent task costs over $31 on Claude Fable 5, compared to around $0.04 on the DeepSeek architecture.

A discreet switch to the final version and tightened performance

This update carried out by DeepSeek, which is relaunching the global race for artificial intelligence, resulted in a discreet adjustment on the publisher’s price listwhere deepseek-v4-pro now identifies the final commercial version DeepSeek-V4-Pro-0813. So far, published independent tests have only relied on a preliminary version deployed in April. The Chinese firm recalled this on July 31 during the launch of its V4-Flash version, specifying that its Pro API remained unchanged and that the final model would follow shortly.

The data sheet on Hugging Face still shows the V4 series as a preliminary version, meaning that no external laboratory has yet independently evaluated this version 0813. Of the ten agent evaluation benchmarks published by the company, Anthropic’s competing Claude Fable 5 model only maintains an average lead of 5.3% across nine tests compared.

Detailed analysis of evaluations revealed determining nuances on the real level of performance. The overall gap of 5.3% is mainly explained by the test “Humanity’s Last Exam” without tools, where DeepSeek records a score of 42.7 compared to 53.3 for Claude Fable 5. By isolating this test showing a difference of 10.6%, the average superiority of its American rival collapses to 2.8% over the rest of the tests. It should be noted that DeepSeek measured these results on its own technical infrastructure, via the minimal mode of its framework. “DeepSeek Harness” configured at maximum effort level with high creativity. Two of the ten events, named “DSBench-FullStack” And “DSBench-Hard”are also based on internal test sets without verifiable public classification.

Here is the summary of the performance metrics published by the laboratory:

  • Claude Fable 5 maintains an average lead of 5.3% on nine benchmarks compared;
  • The gap drops to 2.8% excluding the test “Humanity’s Last Exam” without tools (42.7 versus 53.3);
  • DeepSeek comes out on top in two of the ten agent benchmarks presented;
  • The tests are based on the internal framework “DeepSeek Harness” minimum, including two events without public ranking.

A price gulf between the Chinese alternative DeepSeek and proprietary models

On the pricing front, the confrontation turns into a demonstration of industrial strength when examining the costs of access to the API. DeepSeek maintains its pricing at $0.435 per million input tokens and $0.87 per million output tokens, with a cost of $0.003625 for entry caching. For its part, Anthropic charges Claude Fable 5 at a price of 10 dollars per million input tokens and 50 dollars per million output tokens. By weighting these figures on standard usage ratios, the overall bill comes out to around $0.65 for the Chinese manufacturer compared to $30 for the American giant, i.e. a massive additional cost of 4,600%. This financial gap becomes even more pronounced on the scale of the task performed, with the American model thinking longer and generating more text.

The measurements carried out by Artificial Analysis thus estimate the average task at 3.15 dollars for Fable 5 compared to 3 cents for V4-Flash, i.e. a factor of 105. Clément Delangue, CEO of Hugging Face, estimates the overall bill at “over $31 per task” at Anthropic against “about $0.04” for its competitor.

Even within the Anthropic catalog, the presence of Claude Opus 5 complicates the situation, the latter outperforming Fable 5 on the majority of benchmarks at half its price. This pricing asymmetry is thus beginning to weaken the coherence of American proprietary ranges.

Join the ‘Read to Earn’ program
This link uses an affiliate program

Towards a redistribution of technological cards under the push of open source

This dynamic is part of a movement where Chinese open-weight laboratories, like Kimi whose K3 model had outclassed Fable 5 and GPT-5.6 Sol upon its launch, are drastically reducing inference costs.

The availability of DeepSeek on Hugging Face under the free MIT license offers developers the possibility of immediately verifying this performance on their own servers. This transparency contrasts sharply with the closed ecosystems of Silicon Valley giants, paving the way for permanent community auditing. Ultimately, this democratization of low-cost computing power could transform the global application ecosystem.

It would allow Web3 infrastructures, DeFi protocols and autonomous agent developers to integrate advanced reasoning capabilities without suffering the rent imposed by American giant landlords. However, reliance on internal benchmarks invites nuanced analysis. The future of the market will depend on the ability of open source players to maintain this level of efficiency while offering equivalent guarantees of security and reliability over the long term.

Maximize your Tremplin.io experience with our ‘Read to Earn’ program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.

Similar Posts