The market for advanced models is evolving at a steady pace, driven by frequent updates and ever more precise testing. Anthropic returns with Opus 4.8, a version that aims for better performance without increasing the standard price. In the AI sector, this announcement stands out for its programming results, its new effort settings and its announced advances in security.

In brief
- Anthropic is launching Opus 4.8 just six weeks after Opus 4.7, with better performance and unchanged standard pricing.
- The model is progressing on several benchmarks, notably SWE-bench Pro, where it reached 69.2% compared to 64.3% for Opus 4.7.
- Opus 4.8 offers a quick mode, more expensive, but advertised as less expensive than the quick modes of previous versions.
- New effort settings allow the model to be adapted according to speed, precision and complexity of tasks.
- Anthropic also highlights progress in security, with less deception and fewer bugs left unreported.
Anthropic launches Opus 4.8 shortly after Opus 4.7
Just six weeks after the launch of Opus 4.7, Anthropic presents Opus 4.8 with a striking promise: improving performance while keeping the same price. THE model remains offered at $5 per million input tokens and $25 per million output tokens. This stability provides a useful benchmark for users who compare costs between versions of Claude.
However, the actual cost may change depending on the tasks. The new tokenizer uses more tokens to execute certain requests. Thus, work carried out with Opus can cost more than with Claude Sonnet. The latter remains less powerful, but it may be sufficient for daily uses or complex problems that do not fall under advanced research.
At the same time, Anthropic offers a quick mode for Opus 4.8. This mode runs the same model at 2.5 times the speed. The price then increases to $10 per million input tokens and $50 per million output tokens. According to the company, this mode now costs three times less than previous models.
AI: increasing performance on benchmarks
The most observed result concerns SWE-bench Pro. This test measures an AI's ability to solve complex software engineering problems in real code bases. Opus 4.8 reached 69.2%, compared to 64.3% for Opus 4.7. It also outperforms GPT-5.5, rated at 58.6%, and Gemini 3.1 Pro, rated at 54.2%.


On Humanity's Last Exam, the model also posts high scores. This quiz covers several academic disciplines with expert-level questions.
- Opus 4.8 scores 49.8% without tools;
- Opus 4.8 reaches 57.9% with tools;
- GPT-5.5 reaches 78.2% on Terminal-Bench 2.1;
- Opus 4.8 obtains 74.6% on Terminal-Bench 2.1;
- Opus 4.7 posted 66.1% on this same test;
- Opus 4.8 reached 83.4% on OSWorld-Verified, compared to 82.8% for Opus 4.7.
These results put the model ahead of competitors cited in the data provided on Humanity's Last Exam. The results remain more nuanced on Terminal-Bench 2.1, which evaluates command line tasks carried out by an AI. Despite its second place, Anthropic significantly improves the score of Opus 4.7. On OSWorld-Verified, the model progresses more slightly compared to the previous version.
Effort settings to better control Claude
Opus 4.8 gives more control to users. They can adjust the model's effort level depending on the difficulty of the task. The High level remains enabled by default and is suitable for most requests. The Extra level grants more resources to complex problems, while Max goes even further.
Conversely, Low and Medium levels reduce the resources used. They can save time, but they also reduce the expected accuracy. This choice therefore makes it possible to adapt the AI according to the budget, the desired speed and the complexity of the work.
This control appears near the template picker in claude.ai and Cowork. It remains available for all subscriptions. Anthropic says that the High level consumes almost as many tokens as the standard Opus 4.7 setting, while still giving better results. The flow limits in Claude Code have also been raised to absorb uses of the Extra and Max levels.
Safety and comparison with Claude Mythos Preview
The Anthropic alignment team also highlights progress on model behavior. The data indicate less deception and less cooperation in cases of misuse. Opus 4.8 would also let four times fewer bugs pass in its own code without reporting them.


These results bring Opus 4.8 closer to Claude Mythos Preview on certain security criteria. However, Mythos remains presented as a larger and smarter model than Opus. It is only available in preview to a few selected cybersecurity organizations, including through the Glasswing project.
This limited framework is explained by its advanced capabilities. The UK AI Security Institute found that Mythos could conduct a 32-step network attack simulation on its own. This task usually takes 20 hours for expert security teams. For this reason, the model is not yet commercialized on a large scale.
In the short term, Opus 4.8 should above all strengthen competition between Claude, GPT and Gemini. The challenge will be the balance between cost, speed, precision and security. In this context, artificial intelligence no longer depends only on test scores, but also on the control offered to users and the limits set for deployment.
Maximize your Tremplin.io experience with our 'Read to Earn' program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
