Meta takes on Claude Code and Codex with a low-cost AI agent
Summarize this article with:

Meta launches Muse Code in beta on August 5, 2026, its first AI coding agent, in a market that Claude Code and Codex already control. The company promises a robust tool, designed to manage large codebases without interruption. Performance scores, however, paint a less flattering picture.

Robot inspired by Meta sprinting in a futuristic data center, pursued by two competitors, light chips in hand, dynamic 70s comics style.

In brief

  • Meta launches Muse Code (beta), an in-terminal coding agent powered by the Muse Spark 1.2 template.
  • On Terminal-Bench 2.1, the tool obtained 82.9%, behind Claude Code on Opus 5 (86.7%) but ahead of Codex on GPT-5.6 Terra (81.8%).
  • The central argument of Meta is about robustness: the agent logs every action and can resume exactly where it left off after a crash.

Meta arrives late, but with a new technical argument

The arrival of Meta recalls a well-known pattern: the latest entrant compensates for its delay by betting on reliability rather than raw performance, a logic already observed in the race for AI agents applied to crypto.

Your first cryptos with Bitpanda
This link uses an affiliate program

Muse Code fits into this sequence. In its official announcement, Meta presents the tool as “ a terminal coding agent powered by Muse Spark 1.2 » and promises “ bigger and better models coming soon “.

Designed for software engineering on large code repositories, the tool plans changes, writes code, and verifies the results. It also coordinates several persistent subagents for the same project, which accelerates the resolution of complex tasks with less human intervention.

The detail that really sets Muse Code apart is its internal workings. The agent logs every model call, tool execution, validation, and change in an event log that serves as a single source of truth. Meta explains that this architecture makes the system “identically replayable and capable of restarting without loss”. For tasks that run for hours, this argument matters more than raw speed.

The tool also includes ready-to-use commands:

  • “/plan” transforms a task into a plan subject to validation,
  • “/grill” tests its solidity before execution, and
  • “/goal” pushes the agent towards the completion of the set goal.

Meta specifies that it co-trained Muse Spark 1.2 with Muse Code to make the model and the agent work in tandem.

The scores remain behind against Claude Code and Codex

On the official benchmarks, Meta claims a clear progression from Muse Spark 1.2. The firm indicates that it has “considerably increased the computing power devoted to training on code tasks”. The published figures, however, tell a more nuanced story.

On Terminal-Bench 2.1, Muse Spark 1.2 associated with Muse Code obtains 82.9%, behind Claude Code on Opus 5 (86.7%) but ahead of GPT-5.6 Terra on Codex (81.8%) and Grok Build (81.6%). On DeepSWE 1.1, which measures agentic coding capabilities, the gap narrows: 59.3% for Muse compared to 65.0% for Opus 5 and 64.8% for Codex. On Meta’s internal benchmark, Muse tops out at 70.6% compared to 79.4% for Opus 5.

The progression curves over time partially reverse this classification. Over a thousand consecutive tool calls, Opus 5 shows the biggest gain over its base performance, around 74-75%, while Muse Spark 1.2 is in the middle of the table, between 61-69% depending on the tests. Meta highlights the fact that his agent continues to improve over the accumulated tool calls.

The most striking demonstrations concern this type of scenario. Meta states that Muse Code has “ iteratively optimized GPU cores over a thousand tool calls, up to 24 hours, on Nvidia Hopper GPUs “.

The company also features multimodal use: a user imports a video of a house filmed flying over it into the terminal, and Muse Code “interprets the video and produces a visually rich website with reservation functionalities”.

A market already saturated with AI coding agents

This quote, often used to describe the ambition of AI coding agents, summarizes the challenge that Meta is belatedly trying to address:

It’s not just about autocompletion, it’s unlocking creativity at scale.

OpenAI already runs cloud agents in parallel with Codex, DeepSeek has built its own competitor in Claude Code, and agentic tools like Hermes or OpenClaw already offer comparable or even superior capabilities for certain uses. The strength of Muse Code lies in its fault-resistant architecture and its management of sub-agents, not in numerical supremacy on the test benches.

The risk lies in the very nature of this autonomy. An agent that restarts after a crash and continues to run tools for twenty-four hours straight is still powerful, but unpredictable. Meta is betting that developers will value this autonomy more than pure performance, and the company has chosen to launch it now, rather than waiting for a more mature version. The tool is available for testing via a simple install command in the terminal.

In short, Muse Code arrives in a sector where competition is no longer limited to performance scores but now extends to the operational reliability of agents. Three catalysts will weigh on its adoption: the real solidity of the crash recovery system in production conditions, the ability of Meta to close the performance gap with Claude Code on future versions, and the appetite of developers for a multimodal tool that is still young. Meta’s bet remains risky, but it could redefine the evaluation criteria of these AI agents beyond the simple ranking of benchmarks.

Maximize your Tremplin.io experience with our ‘Read to Earn’ program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.

Similar Posts