AI: Anthropic unveils Opus 4.8, just six weeks after Opus 4.7
Summarize this article with:

The market for advanced models is evolving at a steady pace, driven by frequent updates and ever more precise testing. Anthropic returns with Opus 4.8, a version that aims for better performance without increasing the standard price. In the AI ​​sector, this announcement stands out for its programming results, its new effort settings and its announced advances in security.

Illustration of Anthropic unveiling Opus 4.8 after Opus 4.7, with an AI robot, calendars 4.7 and 4.8 and the mention six weeks.

In brief

  • Anthropic is launching Opus 4.8 just six weeks after Opus 4.7, with better performance and unchanged standard pricing.
  • The model is progressing on several benchmarks, notably SWE-bench Pro, where it reached 69.2% compared to 64.3% for Opus 4.7.
  • Opus 4.8 offers a quick mode, more expensive, but advertised as less expensive than the quick modes of previous versions.
  • New effort settings allow the model to be adapted according to speed, precision and complexity of tasks.
  • Anthropic also highlights progress in security, with less deception and fewer bugs left unreported.

Anthropic launches Opus 4.8 shortly after Opus 4.7

Just six weeks after the launch of Opus 4.7, Anthropic presents Opus 4.8 with a striking promise: improving performance while keeping the same price. THE model remains offered at $5 per million input tokens and $25 per million output tokens. This stability provides a useful benchmark for users who compare costs between versions of Claude.

However, the actual cost may change depending on the tasks. The new tokenizer uses more tokens to execute certain requests. Thus, work carried out with Opus can cost more than with Claude Sonnet. The latter remains less powerful, but it may be sufficient for daily uses or complex problems that do not fall under advanced research.

At the same time, Anthropic offers a quick mode for Opus 4.8. This mode runs the same model at 2.5 times the speed. The price then increases to $10 per million input tokens and $50 per million output tokens. According to the company, this mode now costs three times less than previous models.

AI: increasing performance on benchmarks

The most observed result concerns SWE-bench Pro. This test measures an AI's ability to solve complex software engineering problems in real code bases. Opus 4.8 reached 69.2%, compared to 64.3% for Opus 4.7. It also outperforms GPT-5.5, rated at 58.6%, and Gemini 3.1 Pro, rated at 54.2%.

Comparative performance chart of Opus 4.8, Opus 4.7, GPT-5.5 and Gemini 3.1 Pro on several AI benchmarks, with Opus 4.8 leading on coding, reasoning, computing usage and financial analysis.Comparative performance chart of Opus 4.8, Opus 4.7, GPT-5.5 and Gemini 3.1 Pro on several AI benchmarks, with Opus 4.8 leading on coding, reasoning, computing usage and financial analysis.
Opus 4.8 displays better scores than several competing models on the majority of AI benchmarks presented. Source: Anthropic.

On Humanity's Last Exam, the model also posts high scores. This quiz covers several academic disciplines with expert-level questions.

  • Opus 4.8 scores 49.8% without tools;
  • Opus 4.8 reaches 57.9% with tools;
  • GPT-5.5 reaches 78.2% on Terminal-Bench 2.1;
  • Opus 4.8 obtains 74.6% on Terminal-Bench 2.1;
  • Opus 4.7 posted 66.1% on this same test;
  • Opus 4.8 reached 83.4% on OSWorld-Verified, compared to 82.8% for Opus 4.7.

These results put the model ahead of competitors cited in the data provided on Humanity's Last Exam. The results remain more nuanced on Terminal-Bench 2.1, which evaluates command line tasks carried out by an AI. Despite its second place, Anthropic significantly improves the score of Opus 4.7. On OSWorld-Verified, the model progresses more slightly compared to the previous version.

Cryptosteel: The best tools to stay safe
This link uses an affiliate program

Effort settings to better control Claude

Opus 4.8 gives more control to users. They can adjust the model's effort level depending on the difficulty of the task. The High level remains enabled by default and is suitable for most requests. The Extra level grants more resources to complex problems, while Max goes even further.

Conversely, Low and Medium levels reduce the resources used. They can save time, but they also reduce the expected accuracy. This choice therefore makes it possible to adapt the AI ​​according to the budget, the desired speed and the complexity of the work.

This control appears near the template picker in claude.ai and Cowork. It remains available for all subscriptions. Anthropic says that the High level consumes almost as many tokens as the standard Opus 4.7 setting, while still giving better results. The flow limits in Claude Code have also been raised to absorb uses of the Extra and Max levels.

Safety and comparison with Claude Mythos Preview

The Anthropic alignment team also highlights progress on model behavior. The data indicate less deception and less cooperation in cases of misuse. Opus 4.8 would also let four times fewer bugs pass in its own code without reporting them.

Chart comparing misaligned behavior scores of Sonnet 4.6, Mythos Preview, Opus 4.7, and Opus 4.8, with Opus 4.8 close to Mythos Preview's security level.Chart comparing misaligned behavior scores of Sonnet 4.6, Mythos Preview, Opus 4.7, and Opus 4.8, with Opus 4.8 close to Mythos Preview's security level.
Opus 4.8 shows a lower misaligned behavior score than Opus 4.7 and Sonnet 4.6. Source: Anthropic.

These results bring Opus 4.8 closer to Claude Mythos Preview on certain security criteria. However, Mythos remains presented as a larger and smarter model than Opus. It is only available in preview to a few selected cybersecurity organizations, including through the Glasswing project.

This limited framework is explained by its advanced capabilities. The UK AI Security Institute found that Mythos could conduct a 32-step network attack simulation on its own. This task usually takes 20 hours for expert security teams. For this reason, the model is not yet commercialized on a large scale.

In the short term, Opus 4.8 should above all strengthen competition between Claude, GPT and Gemini. The challenge will be the balance between cost, speed, precision and security. In this context, artificial intelligence no longer depends only on test scores, but also on the control offered to users and the limits set for deployment.

Maximize your Tremplin.io experience with our 'Read to Earn' program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.

Similar Posts