Meta released Muse Spark 1.2 and the agent Muse Code at 20 cents per million tokens, acknowledges shortcomings in its own benchmarks
Meta released the coding model Muse Spark 1.2 and the agent Muse Code (beta), with a plan starting at 20 cents per million output tokens in exchange for data sharing. The company itself acknowledges that its comparative benchmarks against competitors are not methodologically balanced.
Meta released the coding-focused model Muse Spark 1.2 along with its first in-house terminal coding agent, Muse Code, currently in beta. According to the company, the model invests more compute in training on programming tasks, plans steps ahead, tracks a fixed goal and compresses its previous context in long sessions instead of truncating it. Muse Code can resume work after a crash at the exact point of interruption instead of reloading the entire context, because it logs every model call, approval and change to a local log file. Meta also acknowledges methodological shortcomings in its comparative tests against competing models—the test environment was not optimized for competitors, the result achieved elsewhere by the model Claude Opus 5 is two percentage points higher according to other leaderboards, and the model Kimi K3, which the company tested according to its own methodology, is entirely absent from the published results.
According to Meta, Muse Spark 1.2 primarily improves coding capabilities over the model Muse Spark 1.1 released earlier this year, with advances in code generation, debugging and reasoning about large codebases. The model was trained mainly on long-running tasks, such as generating entire repositories or conducting independent research, and some of the training data was created by its predecessor Muse Spark 1.1 itself, which generated programming tasks and evaluated the quality of solutions.
According to Mark Zuckerberg, CEO of Meta, Muse Code can handle end-to-end software engineering across large repositories—planning changes, writing code and verifying results. It installs with a single command, similarly to Claude Code or Codex, and offers the commands “/plan” for plan approval, “/grill” for checking the plan for weaknesses and “/goal” for tracking a fixed goal. For larger tasks, the agent launches concurrent sub-agents working in isolated working copies without collisions; according to Zuckerberg, it built six game features in parallel without conflicts in a test using this approach.
The standard pricing plan remains at the level of version 1.1 (1.25 dollars per million input tokens and 4.25 dollars per million output tokens), with a new plan priced at 20 cents per million output tokens, conditional on sharing user data for further model training. According to the sources, Western competitors charge 10 to 30 dollars per million output tokens, Chinese providers start at 18 cents, and the model Kimi K3 costs 3 dollars per million input tokens and 15 dollars per million output tokens (on a cache hit, the input price drops to 30 cents). Alexandr Wang, head of the AI division at Meta, said the company competes primarily on price, not capabilities. The sources also state that shares in Meta fell by ten percent last week and that the company generates 98 percent of its revenue from advertising.
Why it matters
The cheaper plan expands the range of AI coding agents alongside Claude Code and Codex, but is conditional on sharing data with Meta, which is relevant for companies with sensitive code. The acknowledgment of methodological shortcomings in its own benchmarks calls into question how reliably comparisons with competitors (Claude Opus 5, Kimi K3, Gemini) can be taken, so decisions about choosing a tool should not rely solely on figures published by Meta.
Release card
Muse Spark 1.2
Meta
- Price
- Standard plan: 1.25 USD per million input tokens and 4.25 USD per million output tokens; discounted plan: 20 cents per million output tokens, conditional on sharing user data for further training
- Availability
- Available through the API from Meta; the model comes with the terminal coding agent Muse Code, currently in beta, which can be installed with a single command.
- code generation and debugging in large codebases
- long-running tasks such as generating entire repositories or conducting independent research
- cost-effective deployment on the plan priced at 20 cents per million output tokens in exchange for data sharing
- DeepSWE test results cannot be compared directly with the official leaderboard because each model ran in its own agent
- the discounted pricing plan requires sharing user data for further model training
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
What was added since the original report
Verified updates
-
Muse Spark 1.2 pricing: 20 cents per million output tokens; The model compresses context instead of truncating it in long sessions; Greater compute investment in training on programming tasks; Meta acknowledges methodological shortcomings in its own benchmark tests; Muse Code can restore context after a failure in long-running tasks
- Muse Spark 1.2 pricing: 20 cents per million output tokens
- The model compresses context instead of truncating it in long sessions
- Greater compute investment in training on programming tasks
- Meta acknowledges methodological shortcomings in its own benchmark tests
- Muse Code can restore context after a failure in long-running tasks
Two audiences, two different impacts
What this means
For individuals
Developers gain another terminal coding agent with a planning mode, concurrent sub-agents and crash recovery, with the cheapest plan priced at 20 cents per million output tokens in exchange for data sharing.
For a business
Companies considering deploying an AI coding agent gain another low-cost alternative to Claude Code and Codex, but the lowest price is tied to sharing company code with Meta for model training, and the comparative tests conducted by Meta are, by its own admission, methodologically unbalanced relative to competitors.
DevelopmentCheck the original
Event sources
confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.