Meta shipped Muse Spark 1.3 today, rolling out in Muse Code and the Meta Model API. Pricing carries over unchanged from 1.2: $1.25 per million input tokens and $4.25 per million output tokens on the standard tier, $0.10/$0.20 on the Contributor tier for developers who let Meta train on their prompts.
The scorecard has a stark split in it
Meta published an 11-benchmark comparison against Muse Spark 1.2, GPT-5.6 Sol (max), and Claude Opus 5 (max), grouped into three categories: Agent, Long Context, and Coding. The split across those categories is stark.
On the two Long Context rows, Muse Spark 1.3 wins outright: MRCR 256K–512K (98.5 against GPT-5.6 Sol's 91.5; Opus 5 isn't listed) and MRCR 512K–1M (98.1 against 73.8). On Coding, it wins two of three outright — DeepSWE v1.1 long-horizon agentic coding (75.4, edging out Opus 5's 74.0) and SWEAtlas CodeBase QnA (59.4 against Opus 5's 52.7) — and ties GPT-5.6 Sol for the top spot on the third, Terminal-Bench 2.1 (both at 88.8, ahead of Opus 5's 86.7).
On the six Agent-category benchmarks, it wins zero. Claude Opus 5 takes GDPval-AA v2 knowledge-work Elo (1824 against 1.3's 1754), JobBench professional tool use (65.7 against 64.9), OSWorld 2.0 agentic computer use (68.3 against 66.9), and AutomationBench end-to-end business workflows (50.3 against 49.4). GPT-5.6 Sol takes DeepSearchQA agentic browsing (93.0 against 89.4) and Meta's own internal Agentic IF Index for instruction following (60.5 against 57.8). Muse Spark 1.3 beats its own predecessor comfortably everywhere, but against the two frontier rivals, the categorical line holds: dominant on context and code, shut out on agentic task execution.
That's a useful correction to Meta's own framing. The research blog calls this a step toward "personal superintelligence"; Mark Zuckerberg's own launch tweet goes further, calling it "the biggest jump we've made so far on coding and agentic work." The benchmark table attached to that same tweet shows zero outright wins across all six Agent-category rows. On the numbers Meta itself chose to publish, this is a model that reads and writes code over huge context windows better than Opus 5 or GPT-5.6 Sol — not one that out-agents them, whatever the tweet claims.
Except on the composite Meta didn't publish
Artificial Analysis's independent Intelligence Index v4.1.1 — a 9-evaluation blend covering knowledge work, banking-agent tasks, terminal coding, science coding, and expert reasoning, among others — tells a less clean story. Muse Spark 1.3 (max) scores 62 on that index, tied for third place overall behind only Claude Fable 5.1 (66) and Claude Opus 5 (63), and ahead of GPT-5.6 Sol, Grok 4.6 (both 61), and every other model on the board, including Gemini 3.8 Flash (59). Muse Spark 1.3 (xhigh) scores 61, in the same top cluster. That's a genuinely strong showing on a composite Meta didn't choose and didn't publish itself — and it sits awkwardly next to "shut out on agent tasks," since three of the index's nine components are themselves agentic (τ³-Banking, Terminal-Bench v2.1, AA-LCR).
The catch is real, though: both Muse Spark 1.3 bars are marked "not currently available" on Artificial Analysis's own chart — a checkered pattern the site uses for scores it hasn't independently run yet. Whatever produced those numbers, they aren't yet Artificial Analysis's own verified result. Worth treating as provisional until that changes, not as confirmation that undoes the agent-benchmark shutout above.
What actually changed for users
Meta frames the qualitative gains around collaboration rather than raw capability. Muse Spark 1.3 asks clarifying questions when a prompt is ambiguous, asks for help when it's stuck, and confirms before taking actions it judges consequential — with an explicit toggle for whether it reports progress frequently or works silently until it has a deliverable. It's also trained to map incoming messages to the right task inside one messy, multi-turn thread, whether a new message is steering an existing request or interrupting it outright. Against Muse Spark 1.2, Meta's own engineering comparisons put the efficiency gain at roughly 20% fewer tool calls and 25% fewer tokens for comparable coding work.
The open-weights promise is still just a promise
Five days after Muse Spark 1.2 shipped closed, Meta said it would open the weights — no license, no date. That was August 10. Three weeks later, Muse Spark 1.3 ships closed too, and both Meta's research blog and Zuckerberg's own tweet repeat the same line — the blog's version reads "we also have bigger models, Muse Spark open weights, and more on the way soon"; Zuckerberg's is shorter: "Muse Spark open weights releases coming soon." Still no date, still no license, repeated by the CEO himself, on the second model generation since the reversal was first promised.
Safety
Meta says Muse Spark 1.3 has stronger adversarial robustness against prompt injection and better calibration on what counts as an irreversible action during long agentic runs. Max reasoning mode isn't shipping with today's release — Meta says it's coming "shortly" once additional safety testing finishes, leaving the model's stated highest capability tier unavailable at launch.
What to expect next
- Watch whether Artificial Analysis actually verifies the 62/61 scores. Until the "not currently available" marker clears, the strongest evidence against the agent-benchmark shutout is unconfirmed by the one party that didn't get to choose which numbers to show.
- Watch whether the agent-category gap closes in 1.4. A model that dominates long-context and coding but loses every benchmark in Meta's own Agent category reads like a lab that optimized for one axis this cycle — worth checking whether the next release rebalances toward agentic task execution specifically.
- Watch for an actual date on Muse Spark open weights. "Coming soon" has now been the answer for three weeks and two model generations; the next real signal is a license and a date, not another repetition of the same sentence.
- Watch for max reasoning mode's actual release. "Shortly after additional safety testing" is open-ended — worth checking how long the gap runs between today's launch and the promised capability tier actually shipping.