Meta's new coding agent is ready for action, yet it's already running into a familiar problem. Muse Code shipped Wednesday powered by Muse Spark 1.2, and while it beats OpenAI's Codex and Google's Antigravity on most tests, it consistently loses to Anthropic's Claude Opus 5 across every benchmark the company published.

The terminal-based agent comes with a practical advantage though. Its crash-safe event log records every model call, tool execution, approval, and edit in a single source of truth. If the system crashes mid-job, it can restart exactly where it stopped without replaying work. That durability matters when you're looking at the kinds of jobs Muse handles.

Speed over perfection

Meta scaled up compute on coding tasks during training for version 1.2, and the payoff shows in specific domains. The model handles code generation, debugging, and large codebase understanding better than before. In one case study, it rewrote GPU kernels for NVIDIA Hopper chips across more than 1,000 tool calls spanning 24 hours straight, working from a baseline without copying existing library code.

The company positioned this as a stepping stone. CEO Mark Zuckerberg noted bigger, more capable models are coming. For now, Meta is competing on cost and features rather than raw benchmark performance. The crash-safe log and session persistence for background subagents are designed for developers who need reliability in long-running tasks, not just accuracy on test suites.

The price play

This fits a broader pattern in AI development where companies target different customer needs instead of chasing the same performance metrics. Codex was the reference point for years. Antigravity filled a niche. Now Muse Code enters a market where teams might accept lower benchmark scores if the agent handles production crashes gracefully and keeps working through extended sessions.

This article is informational only and should not be considered financial or investment advice.