Top
Best
New

Posted by gmays 5 hours ago

Ember-1(fireworks.ai)
256 points | 144 commentspage 2
intothemild 4 hours ago|
So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
netvarun 4 hours ago||
Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.
Evidlo 4 hours ago||
weight available
DonsDiscountGas 4 hours ago|||
It happens. Most open licenses aren't GPL style copyleft.
kingstnap 3 hours ago|||
There is little to no point reading the article as well. It's stripped of all alpha.

> task and environment feedback

> on-policy planning and learning

> feedback connects decisions to their consequences

These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.

Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.

intothemild 1 hour ago||
This is precisely my point.
makeramen 3 hours ago|||
Aren't Cursor Composer models like this too? At some point all the extra RL you do can be considered as proprietary information added.

Not suggesting this is right or wrong, but is sort of the nature of the technology.

swagatkonchada 3 hours ago|||
It happens with open source software all the time, why would we expect any different with open source weights.
reactordev 3 hours ago|||
Because we do. The GPL isn't a suggestion. If you can take open source code and make private software out of it then what are we all doing? No, license requirements and agreement are law for a reason.
bloggie 3 hours ago|||
Kimi K3 has its own license which is permissive, it isn't at all like GPL https://github.com/MoonshotAI/Kimi-K3/blob/main/LICENSE
dgellow 2 hours ago|||
GPL is a specific license, it’s not FLOSS as a whole
otterley 3 hours ago|||
Because the licenses that apply to software make no sense in the context of LLMs. With the latter, there is no source code to license.

The words of a license are what the license is.

dgellow 2 hours ago|||
Which is fine, that’s legal according to the license
spdustin 2 hours ago||
Been thinking about the feasibility of training a model using synthetic thinking traces that were reduced to caveman-speak prior to being used for training. Seems like it would be fairly easy to generate plenty of suitably lobotomized synthetic traces with a pair of cheap-ish models. Or even just using good old fashioned NLP to aggressively remove stop words and reduce trace words to lemmas.
dmkolobov 1 hour ago||
This is cool! But also: am I wrong for thinking “Pareto frontier” is some pretty silly/clever marketing jargon? Is this common phrasing for basically saying: test performance per spend on tokens is decent?
sebzim4500 1 hour ago|
I don't see why? It is a well defined term that existed prior to the recent AI bubble/revolution, and from what I can see they are using it appropriately.
dmkolobov 1 hour ago||
Fair enough!
tomrod 4 hours ago||
Well done, and great iteration.

The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).

drob518 4 hours ago||
Unfortunately, it’s hard to make a chart of that.
ttmoab 8 minutes ago||
[flagged]
nxtfari 2 hours ago||
The more I learn about Fireworks the more unsavory they seem as a company. I don’t care what the license says, Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it’s closed weights, “this is our own proprietary” nonsense? Where are we that China has better open source ethos than America?
peri-cl 1 hour ago|
Why is proprietary-licensed software unethical?

Kimi K3 itself isn't FOSS. Speaking of reciprocity: Fireworks is presumably paying Moonshot serious money for the right to do what they are doing here, since Kimi's license[0] excludes commercial inference providers (such as Fireworks) from gratis use. It requires them to: "...enter into a separate agreement with Moonshot AI before using the Software or its derivative works..."

[0] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE#...

srameshc 2 hours ago||
> The problem: thinking models think too much

I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there

riquito 3 hours ago||
Aside. I find the "cost per task" charts both useful and uncanny. Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars? Or a different model that too scores 90% in 1 dollar? How much will it cost me the last 10% or 5%? At the end of the day, cost to 100% is what matters and the half (90%) backed solution may require more to reach 100% (or not, who knows?)
swiftcoder 2 hours ago|
> Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars?

It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...

entrope 1 hour ago||
The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.

On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.

erichocean 4 hours ago||
Need this done for DeepSeek, ideally one of the Flash models.
drob518 3 hours ago||
And GLM. Both Deepseek 4.1 Flash and GLM 5.3 Flash are quote verbose when thinking.
atemerev 4 hours ago||
If you have the compute, I have the expertise.
blissofbeing 2 hours ago||
Would be nice to include in fire pass.
tdhz77 4 hours ago|
Does anybody know if this would be a good model for creative writing?
combobyte 3 hours ago|
> model

> creative

Choose one.

tdhz77 1 hour ago||
Are you a bot?
combobyte 52 minutes ago||
You really have no sense of irony, do you?
More comments...