Muse Code and Muse Spark 1.2
94 points by paulkrush 3 hours ago | 55 comments

WhitneyLand 2 hours ago
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

They left Opus in and got beat in all but one benchmark.

Nothing wrong with trying to improve, but why the marketing games?

Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.

Then when your ready, come back and talk frontier without playing hide the model.

reply
jjice 21 minutes ago
While I won't take their limited benchmarks with much salt, if it actually is this close to opus, but at a third the cost, that's pretty solid. Now, Terra is pretty damn affordable too and you're right that it's suspicious that they don't put Sol in there at all.
reply
krm01 35 minutes ago
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
reply
lacker 9 minutes ago
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
reply
nrub 2 minutes ago
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
reply
bradfa 2 hours ago
If you got the $20 in free credits from Meta for signing up when muse-spark-1.1 was release, please note that there's now small print stating "While using free credits your content may be used for product improvement" which was not present at muse-spark-1.1 launch when the credits were given out.

If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.

reply
giancarlostoro 44 minutes ago
The API costs for the version of their model that feeds things back to meta is also drastically lower.
reply
tristanj 56 minutes ago
Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.

https://developer.meta.com/ai/models/muse-spark/

reply
ray_kay777 22 minutes ago
This makes it a very interesting alternative to Deepseek for personal work where I don't care about the training - judging by the AA benchmarks it seems like overall cost per task is similar to the new Deepseek Flash but with better benchmarks (and inbuilt vision capabilities).
reply
GodelNumbering 45 minutes ago
I think that's a fair offering tbh
reply
deno 37 minutes ago
I think it's limited to US or at least EU is excluded.
reply
mchusma 2 hours ago
This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.
reply
handzhiev 2 hours ago
If you are happy to share data for training, the contributor mode offers amazing price $0.10 / $0.20
reply
mchusma 7 minutes ago
Yes, that is the really compelling thing here IMO. Its a viable deepseek competitor for many people, and I missed that on the first pass.
reply
wxw 2 hours ago
Last I heard, everyone at Meta was using Claude Code.

Any insiders know how Muse Code is doing internally?

reply
GodelNumbering 57 minutes ago
If there were, do you believe it would be in their interest to answer this publicly?
reply
georgemcbay 44 minutes ago
> > Any insiders know how Muse Code is doing internally?

> If there were, do you believe it would be in their interest to answer this publicly?

If it were being adopted like gangbusters in their organization, sure!

So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...

reply
GodelNumbering 20 minutes ago
You should play Blood on the clocktower
reply
youre-wrong3 50 minutes ago
[dead]
reply
conradkay 2 hours ago
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg

Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?

Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data

reply
ipsum2 2 hours ago
I wonder why they didn't compare with GPT-5.6-sol, only Terra?
reply
wmf 2 hours ago
Clearly they're positioning it as a mid model.
reply
minimaxir 2 hours ago
Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive).

Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.

reply
redox99 2 hours ago
But why include Opus then?
reply
woadwarrior01 2 hours ago
Haven't you seen the kernel optimization case study at the bottom of the page? They compare against GPT-5.6 Sol and their model is worse.
reply
logicchains 2 hours ago
Presumably because it's worse than Sol, same reason they compared it to Opus 5 not Fable.
reply
Handy-Man 2 hours ago
Their bigger model is not ready - watermelon code name was still being prepared for release as of a month ago
reply
arjie 57 minutes ago
Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model.

The use traces must be crucial to functionality which is why they’re keeping prices so low.

reply
wmf 40 minutes ago
They rebooted less than one year ago so this is decent progress. Obviously users don't care about progress though.
reply
arjie 20 minutes ago
Yeah, progress is useful as an internal metric, but I'm going to measure against the present frontier unfortunately. Eager to see what they come up with in the future.
reply
daemonologist 2 hours ago
Pricing: https://dev.meta.ai/docs/pricing-rate-limits

Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.

reply
IceWreck 2 hours ago
Ive been poking with the muse code binary - seems to be written in rust, looks similar to codex but either its a very hard fork (i also see dissimilar things like config format is different, no acp, etc) or is just heavily inspired by it (more likely).
reply
sarjann 57 minutes ago
I do think some of features in their harness seem interesting (workers in separate worktrees at once), recovery from crashes seem interesting.
reply
liviux 2 hours ago
Does this muse code have any muse spark 1.2 usage included? Can't understand from the docs.
reply
Cappybara12 2 hours ago
Is this becoming a race where we have a usual flow of a company .. AI models, Coding agents, image generation tools, and more AI models ?
reply
minimaxir 2 hours ago
Muse Spark 1.1 was released July 16th, less than a month ago. A new version release this soon (particularly after Kimi K3's release drastically overshadowed it) is a bit sus and it appears that Meta is trying a first launch do-over.
reply
gaogao 60 minutes ago
Frequent minor version bumps are pretty common these days. Opus 4.7 -> 4.8 was 42 days.
reply
minimaxir 58 minutes ago
Which was in itself a do-over because Opus 4.7 received a lot of bad press on suspicion of being a regression from 4.6.
reply
Bolwin 2 hours ago
> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access

Wasn't the previous one us only? This is probably the biggest part of the post

Anyone know if muse code is open source?

reply
king_crimson 2 hours ago
Why does every AI lab feel the need to build their own coding agent…? Don’t we have more than enough already?
reply
kcb 2 hours ago
Open the weights.
reply
AtlanticThird 2 hours ago
I wish they would add a ZDR endpoint on OpenRouter
reply
Laurel1234 35 minutes ago
The only company less trustworthy than OpenAI and Anthropic is meta.
reply
fcoury 2 hours ago
Interesting, it seems like their muse code is built upon Codex CLI?
reply
paulkrush 3 hours ago
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.
reply
vcryan 43 minutes ago
It seems like one day, Google or Meta might produce a coding model worth discussing. That day is not today.
reply
qphe95 2 hours ago
Theres no actual evidence they didn't just distill Kimi K3
reply
toephu2 2 hours ago
At this point, it doesn't matter who is distilling from who.
reply
Jabrov 2 hours ago
Is there any actual evidence that they did?
reply
esafak 2 hours ago
If anyone from Meta is reading, please can you publish the cost and latency for each of your benchmarks, like OpenAI does? Show us how the reasoning effort level affects them in 2D charts. This needs to become standard practice.
reply
giancarlostoro 2 hours ago
Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary.

I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.

reply
greyb 60 minutes ago
I honestly think they're kinda banking on piggybacking off of Facebook account integrity systems to avoid the problems that other LLM providers are facing in trying to prevent mass free trial signups for token relays and so forth.

It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).

reply
hahahaa 60 minutes ago
China says hi.
reply
aanet 2 hours ago
+10000 to that
reply
Readerium 2 hours ago
Lol worse than DeepSeek
reply
rvz 2 hours ago
First of all, you have login to use it. Why?

After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.

Think twice before falling for this announcement and ask yourself what they are not telling you.

reply