Mark Zuckerberg launched Muse Code in beta on Wednesday, Meta’s first artificial intelligence (AI) coding agent. Anthropic’s Claude Opus 5 beats it in all four comparisons Meta published at launch.
Meta released those charts anyway. The company is selling a cheaper tool rather than a better one. Independent test data suggests the gap is wider than Meta showed.
Follow us on X to get the latest news as it happens
Meta’s Own Charts Hand Anthropic Every Round
Muse Spark 1.2 is the model inside Muse Code. It scored 82.9% on Terminal-Bench 2.1.
Claude Opus 5 scored 86.7% on the same test. Terminal-Bench comes from the Laude Institute and Stanford researchers. It sets 89 real jobs spanning system repair, data work, and security.
Second place is respectable. Muse Code beat OpenAI’s Codex at 81.8% and Grok Build at 81.6%.
The next chart was harsher. DeepSWE 1.1 sets 113 coding tasks with internet access switched off during grading. Muse Spark 1.2 dropped to third at 59.3%.
Meta then published a test it built itself, drawn from 440 real pull requests by its own engineers. Muse Spark 1.2 scored 70.6% there, roughly nine points behind Opus 5.
That score sits only 2.3 points above Muse Spark 1.1, the model Meta shipped in July.
The Model Meta Left Off Its Coding Charts
Meta measured itself against GPT-5.6 Terra. OpenAI sells a stronger model called Sol, and Meta left it out of all three coding charts.
Sol tops the independent Terminal-Bench 2.1 leaderboard at 89.5%. Opus 5 follows at 89.1%.
Both figures beat the 86.7% Meta reported for Opus 5. Meta picked a weaker setting of its strongest rival and still finished behind it.
Against Sol, the true leader, Muse Spark 1.2 trails by 6.6 points rather than 3.8.
Meta did include Sol in one place. On a graphics processing unit (GPU) kernel task running past 1,000 tool calls, Sol improved on the baseline by 71.2%. Muse Spark 1.2 managed 68.7% and placed fourth of six.
One caveat cuts the other way. Muse Spark 1.2 does not appear on that public leaderboard yet, where only 26 of 183 tracked models have been tested. Its 82.9% remains a Meta figure.
“Muse Spark 1.2 is our next step as we push toward frontier, with larger, more capable models on the way,” Zuckerberg said in a post.
Zuckerberg May Soon Host the Model Beating His Own
Meta is reportedly in talks to lease compute to Anthropic. The deal could reach $10 billion over two years. Meta data centers would then help run the Claude models Muse Code was built to unseat.
The leadership behind Muse Code was expensive. Zuckerberg paid $14.3 billion in June 2025 for Scale AI and its founder Alexandr Wang, who now heads Meta Superintelligence Labs.
Price is the lever Wang has left. Rates match the July launch of Meta’s first paid API at $1.25 per million input tokens and $4.25 per million output tokens.
A contributor tier costs more than 10 times less. Developers qualify by letting Meta train on their work. Wang declined to give adoption numbers for the Muse Spark line.
Meta’s accounts explain the discount. Revenue climbed 28% to $60.8 billion last quarter, yet operating profit fell 8% to $18.8 billion.
Operating margin slid to 31% from 43% a year earlier. Meta spent $31.08 billion on capital projects in the quarter alone, and guides to as much as $145 billion for the year.
Muse Code does offer engineering Claude Code lacks. Background agents hold context across a session. Sub-agents work in isolated copies of a repository.
Meta has built a solid second-best coder and priced it like a budget option. The beta will show whether developers trade a few points of accuracy for a bill roughly a tenth the size.
The post Zuckerberg’s Muse Code Loses to Anthropic on Meta’s Own Benchmark Charts appeared first on BeInCrypto.
