Top
Best
New

Posted by paulkrush 1 hour ago

Muse Code and Muse Spark 1.2(research.meta.ai)
73 points | 44 comments
WhitneyLand 21 minutes ago|
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

They left Opus in and got beat in all but one benchmark.

Nothing wrong with trying to improve, but why the marketing games?

Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.

Then when your ready, come back and talk frontier without playing hide the model.

bradfa 30 minutes ago||
If you got the $20 in free credits from Meta for signing up when muse-spark-1.1 was release, please note that there's now small print stating "While using free credits your content may be used for product improvement" which was not present at muse-spark-1.1 launch when the credits were given out.

If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.

giancarlostoro 1 minute ago|
The API costs for the version of their model that feeds things back to meta is also drastically lower.
tristanj 13 minutes ago||
Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.

https://developer.meta.com/ai/models/muse-spark/

GodelNumbering 2 minutes ago|
I think that's a fair offering tbh
vcryan 22 seconds ago||
It seems like one day, Google or Meta might produce a coding model worth discussing. That day is not today.
mchusma 57 minutes ago||
This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.
handzhiev 35 minutes ago|
If you are happy to share data for training, the contributor mode offers amazing price $0.10 / $0.20
wxw 30 minutes ago||
Last I heard, everyone at Meta was using Claude Code.

Any insiders know how Muse Code is doing internally?

GodelNumbering 14 minutes ago||
If there were, do you believe it would be in their interest to answer this publicly?
georgemcbay 1 minute ago||
> > Any insiders know how Muse Code is doing internally?

> If there were, do you believe it would be in their interest to answer this publicly?

If it were being adopted like gangbusters in their organization, sure!

So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...

youre-wrong3 7 minutes ago||
[dead]
arjie 14 minutes ago||
Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model.

The use traces must be crucial to functionality which is why they’re keeping prices so low.

conradkay 44 minutes ago||
https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg

Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?

Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data

sarjann 14 minutes ago||
I do think some of features in their harness seem interesting (workers in separate worktrees at once), recovery from crashes seem interesting.
ipsum2 1 hour ago|
I wonder why they didn't compare with GPT-5.6-sol, only Terra?
wmf 1 hour ago||
Clearly they're positioning it as a mid model.
minimaxir 1 hour ago|||
Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive).

Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.

redox99 21 minutes ago|||
But why include Opus then?
woadwarrior01 1 hour ago|||
Haven't you seen the kernel optimization case study at the bottom of the page? They compare against GPT-5.6 Sol and their model is worse.
logicchains 1 hour ago|||
Presumably because it's worse than Sol, same reason they compared it to Opus 5 not Fable.
Handy-Man 59 minutes ago||
Their bigger model is not ready - watermelon code name was still being prepared for release as of a month ago
More comments...