Top
Best
New

Posted by seelos 19 hours ago

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra(cognition.com)
419 points | 175 commentspage 3
yipinwong 12 hours ago|
Incest in human biology causes mutations that's bad in the long term.

Same for AI models trained on Kimi-3 or other models like Chinese models do. They suffer from the same issue.

monkeydust 18 hours ago||
As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.
ccapitalK 17 hours ago||
Pareto frontiers are pretty commonly invoked to describe tradeoffs in computer science and have been for quite a while. I remember the term being used in one of my early algorithms courses to describe the tradeoff between data structures with fast writes, ones with fast reads and ones that tried to balance the two.

IIRC cognition boasted about hiring a lot of competitive programmers and algorithms experts back when they released Devin, so it tracks that they'd use the term.

walrus01 17 hours ago||
It's interesting watching people throw about pareto frontiers sort of like how RF nerds approach the shannon limit (in a practical real world sense of the term, like charting possible modulations/data rates on a two way satellite modem's manufacturer datasheet).
arrowleaf 17 hours ago||
Pareto leaked out of the sociology/econ bubble a long time ago :) Pareto principle, Pareto efficiency, Pareto distribution have been in the pop-sci buzzwords for quite awhile, I probably encountered it first in the 4-Hour Workweek. I don't think you can read a self-help book without the author introducing it as a groundbreaking principle to live your life by.
Tsarp 19 hours ago||
"SWE-2 is post-trained from Kimi K3"
airstrafer 18 hours ago||
Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.

Maybe still worth it if their "64% cheaper" figure holds.

Bolwin 18 hours ago|||
I don't think you know what distill means
airstrafer 17 hours ago||
I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?
FergusArgyll 16 hours ago||
Distilling you don't have the actual model weights of the teacher. All you have are the teachers answers to a lot of questions. You then teach your own smaller model to answer more similarly to the big teacher model.

Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way.

Bolwin 8 hours ago||
Distillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself.

What you're describing is just synthetic data.

Note Anthropic misused the term in their post about Chinese model distillation, deliberately I assume.

FergusArgyll 7 hours ago||
There are multiple kinds of distillation

https://arxiv.org/abs/2106.03310

https://arxiv.org/abs/2207.12106

Tsarp 18 hours ago|||
With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.
samyok 18 hours ago||
SWE-2 is free for all subscribers on the CLI to try out for the next month :)
teddyX 17 hours ago||
What about the gui/windsurf app? Same as cli?
xlbuttplug2 18 hours ago||
I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.

I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.

llmslave 18 hours ago||
At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

hnedeotes 18 hours ago||
that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way to high, days when I have diarrhoea my AI agents move even faster
llmslave 17 hours ago||
please keep thinking this so i can relax with my automated job
hnedeotes 17 hours ago||
It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu
mogwire 13 hours ago|||
If you are too dumb or too lazy to figure out what this guy has, then you are the problem and will be looked at as a relic.

Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently.

Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex command.

hnedeotes 4 hours ago||
Maybe it's you that needs to learn how to run `--help` or ask an AI how to not cry about burnout on open source instead? I don't get it, you should just be cruising on auto-pilot now.

The problem is retards that can only function on a cocktail of drugs, and as they were never good at anything other than anal retentive stuff built and continue to build these retarded systems. Those peddling RoR apps even when they couldn't serve more than 3 or 4 concurrent requests, JS backends to handle complex workflows that even after 2 years of dev. still have bugs and accrued a sprawl of crap to hide the issues of their own making, etc, and yet charge thousands of dollars, those that write shit software that's not even worth to clean your ass with, even though they have 20 years of experience, but then go give conferences and write books about their amazing architectural skills, those that write utils behind the "oh, it's open source, if you don't like it just fork it" and due to marketing get their crap everywhere, while making holes everywhere for their paycheques. Or the nepo babies that need their mexico border run to get their fix so they can have these "humanity changing" ideas? I bet they're the same that before would weasel a 2 week sprint to change the borders of a button. Or burn through 10k in meetings for irrelevant crap. Or get VC funding for a CSS styling company or a two prompt company. Or go on about the value of ideas, but then can't even get that going without outsourcing or an AI to help them have those same "ideas".

Ultimately, you just need to turn into a little pig and party in the pigsty, it's not that difficult either, they say pigs are very close anatomically to humans.

At least AI can help untangle the crap the anal retentive retards have built, and thank god, the pig-mor, this society can't even fuck to replacement levels (perhaps they'll manage now with AI).

ArsenDjukic 17 hours ago||
[flagged]
gigatexal 7 hours ago||
Not avail on open router?
ltsSmitty 18 hours ago||
Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at
microdrum 11 hours ago||
Do they have anything as good as Amp (which is able to use free models)? Amp has been in the lead for almost a year now and doesn't seem to be relinquishing it.
m3kw9 17 hours ago||
I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed.
_doctor_love 19 hours ago||
SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.
_doctor_love 3 hours ago|
Odd to get downvotes simply for sharing my experience. Like it or not, Cognition has a good frontier model and they are building serious products. They are working hard which is how you become successful. Sorry if that ruffles your feathers.
wqash71 17 hours ago|
The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all.
More comments...