Top
Best
New

Posted by seelos 18 hours ago

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra(cognition.com)
416 points | 174 comments
postalcoder 17 hours ago|
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

mediaman 17 hours ago||
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

postalcoder 16 hours ago|||
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.

Lucasoato 11 hours ago|||
> A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

Wait a second, are we taking into account the massive difference in terms of resources of these two companies?

solenoid0937 11 hours ago|||
Irrelevant when they say they're competitive with Fable and Astra. They don't get to then roll that back and then say "but we have less compute!"

You're either competitive or not.

asdfsa32 3 hours ago|||
You should use my model then, I spent about 30$ in electricity and used my existing RTX4090. It is not very good, but can you compare it with others really? You can use this service via a private API with a VPN, email me your credit card details for access.
ben_w 13 hours ago||||
Benchmaxxing is the default case, and always has been.

It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.

Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.

podocarp 5 hours ago|||
You can't justify it being OK just because it's common. Here, this just makes benchmarks into a low signal and useless marketing number once people get numb to all the 99%s.

Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.

injidup 4 hours ago||
Cancer is simply cells breaking free of the cooperative jail. Essentially the grey goo scenario of nano machines. Instead of cooperation they just do their own thing.

Cells have a certain optimised DNA mutation rate kept in check by various machinery. Multi cellular life expects each of these little replication machine to co-operate in the grand scheme of running a body. But it's also required in the grand scheme for DNA to mutate a little bit to ensure population variance. So you could say that cancer is the tax paid for having cooperative yet flexible and adaptive nano machinery.

So yes the propensity for cancer developed under evolutioniary pressure towards a non zero level.

The population could have optimised for zero cancer but it would not have paid for itself in terms of overall population adaptability and survival.

conmod278 13 hours ago|||
Why would I adapt myself to unseen environments unnecessarily?
NewJazz 12 hours ago|||
You can see possible existing environments even if you explicitly blind yourself to them, through reflections off of environments that you do not blind yourself to. Even with the benchmark excluded, the social zeitgeist that has considered the benchmark and included it or ideas from the benchmark either implicitly or explicitly in their code, documentation, et cetera is still part of your training data.
Muromec 12 hours ago||||
Because then you could sit under the palm tree and enjoy your free bananas and relax.
TedDoesntTalk 12 hours ago|||
Propagation
nrmitchi 16 hours ago||||
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Yes.

willcmcc 15 hours ago||||
"Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"

Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.

dpweb 16 hours ago||||
Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems.

Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.

gunalx 10 hours ago||
Wouldn't work. A dynamic but verifiable problem. Is just a perfect target for a RL environment. If you don't have the verifiable part the benchmark is useless, or really expensive with human review. (Or just open ended)
iLoveOncall 16 hours ago||||
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Yes? Just like every single model from every single AI lab.

general_reveal 17 hours ago||||
I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds.

You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.

Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.

felixgallo 17 hours ago||||
Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
letmevoteplease 16 hours ago|||
> Altman was caught in previous attempts trying to game benchmarks

Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.

> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)

And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.

kzrdude 15 hours ago|||
There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.

That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.

Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.

xnickb 4 hours ago||
Sounds more like profitmaxxxing to me to be honest.
kzrdude 3 hours ago||
They have plausible deniability on that one: not making any profit.
vlovich123 16 hours ago|||
Yeah and Astra is much better still
general_reveal 16 hours ago|||
[flagged]
moomin 16 hours ago||
A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.
general_reveal 16 hours ago|||
Well, the antichrist should be in and around this AI thing for one particular reason:

The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is).

In that way , AI is a perfect mimicry of how the devil operates (in totality, as the devil perverts and replicates anything good, often subtly and always deceptively), which is to thieve off God, steal.

So he would be around, if you catch my drift, right about now. And I wouldn’t be shocked if he’s on HN, and that he would chose technology as the vessel. And ultimately, when it’s all said and done, I would not be shocked that those who studied and developed AI, did so for the devil whether they were aware or not.

Anyway, let a poor Christian have his end-times hypothesis.

icantevenhold 16 hours ago|||
What’s the antichrists goal though?

Create hell on earth or turn us all into heretics or something else?

general_reveal 16 hours ago||
Anti Christ goal:

Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ.

Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff.

As per Revelations. Thank your for allowing me to edify :)

johnsmith1840 15 hours ago|||
It's actually pretty cool of these old stories how well that maps to a dangerous tyrant.

The roman empire perfectly matched that and most powerful men seem to go that path.

It would also be kinda easy to argue many moderns countries are going down that path.

bookshaman 15 hours ago||||
So, we're rooting for the Anti Christ then? I might have to change my stance on generative "ai" then.
icantevenhold 8 hours ago|||
Why would a Christian be trying to stop the Antichrist? Sounds like the faithful have nothing to lose.
cindyllm 14 hours ago|||
[dead]
ZeWaka 16 hours ago|||
I guess we also need more antipopes.
yieldcrv 10 hours ago|||
I mean, still benchmaxxed, I happen to not consider that a problem

These firms are literally hiring professionals from all fields to teach procedure

To teach processes that can subsequently be done agentically or in automated chains

Its basically infinite permutations of tool calling, except the tools aren't external, they’re baked in upon birth

So yeah still makes sense that the new benchmark has a low score and the older one has a high score. And sure, one day we wont have to debate it and a new model will ace everything. Do you actually want that day to be today?

fallingbananna 16 hours ago|||
Those 27.3% are still in the ballpark of modern models:

- Sonnet 5 - 12.4%

- Luna - 17.3%

- Grok 4.6 - 20.3%

- Sol - 37.3%

- GLM 5.3 - 41.8%

- Opus 5 - 51.8%

mokre 12 hours ago|||
GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

didibus 5 hours ago|||
Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.
gpt5 6 hours ago|||
Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3
p1esk 14 hours ago||||
Astra is 58%. The current title says it's "rivaling Astra"
thereitgoes456 14 hours ago||
It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
Readerium 15 hours ago||||
DeepSeek v4.1 Flash 31.2%

Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...

nicce 13 hours ago||
I have been using it today the whole day and it is definitely better than Luna. Great that bench agrees.
nijave 13 hours ago||||
This explains a lot about Sonnet 5.
airstrike 12 hours ago|||
So, better than Sonnet and Luna? lol
throwatdem12311 16 hours ago|||
This is why I find benchmarks absolutely worthless.

First, almost all models are within spitting distances of eachother.

Second, it never translates to being better for my own workloads.

You just need to make your own benchmarks.

walrus01 15 hours ago|||
For comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0.
Readerium 15 hours ago|||
Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0
eranation 16 hours ago|||
When a benchmark becomes a target, it's no longer a good benchmark...
tonychang430 15 hours ago||
people are just fighting for numbers.. i don't fundamentally see the model being better
thefourthchime 15 hours ago|||
Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4.

I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!

Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.

dudeinhawaii 13 hours ago||
Your post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53.

That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.

enraged_camel 17 hours ago||
Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.
nullbio 17 hours ago|||
Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.
throwup238 16 hours ago|||
> Andreessen Horowitz is being played like a fiddle.

Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.

thereitgoes456 17 hours ago|||
The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less.

While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.

Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.

selectodude 16 hours ago||
The cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days.
throwaway240403 15 hours ago||
Improvements in models and products coming out of SpaceXAI since the acquisition would seem to disagree with you.
gruez 16 hours ago||
Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?

https://www.youtube.com/watch?v=tNmgmwEtoWE

As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

fishtoaster 15 hours ago||
A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?
arjie 12 hours ago|||
Their ads in SF are pretty funny. “Remember Devin? It’s good now”. Okay, very self-aware, Cognition.
klardotsh 12 hours ago||||
Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.
fishtoaster 5 hours ago|||
Can't say I've tried the CLI. I've mostly focused on the cloud agents, which I was explicitly looking for. I compared against cursor's cloud agents, ampcode, and hoplite, and came out surprisingly enjoying devin.

I will say that the lack of parity between Devin cloud and Devin desktop is downright embarrassing. It's very clear that the latter is a thinly-reskinned Windsurf. A visually similar UI with vastly different capabilities. Definitely a black mark on the whole thing.

paimapi 11 hours ago|||
Had the same experience on it when it was going by Windsurf. Had to wire up a skill hooked to terminal runs or else it would hang or freeze and never finish literally every single time

I see issues with other harnesses too but not with the regularity I was getting from this. And the one moat they had with the better UI for per-project multi-agent orch disappeared and now is standardized

wetpaste 12 hours ago||||
Not all that surprising. The original Devin really was just an early attempt at agentic coding before models were really even trained for it. Now that it's a well established pattern and we've figured out what works, I'm not surprised they've morphed into something reasonble.
tyre 12 hours ago||||
When they launched Devin it was supposedly at the performance of an engineering intern. Friends who used it found the bad parts of an intern (tons of handholding, review required) but it didn’t learn from mistakes or add throughput.

They seem to love a good overpromise.

darkwizard42 14 hours ago||||
I think the evolution of the harness and ability to preserve loop context outside the context window has made running these kinds of agentic experiences easier.

sorry so many buzzwords to say, the capabilities to do this kind of work are more accessible and easier to manage, so now it works!

Good to see, and agree they were severely overhyping their product back then.

Saline9515 11 hours ago||||
Devin lacks many features and doesn't even have an exec mode. Other than they subsidize the subscription, the added value here is low.
princevegeta89 7 hours ago||
I used Windsurf for quite long and they're basically dead for now after they became Devin.

Everything feels dull and they're always several features behind while Cursor is just killing it every other week.

I'm back to VsCode plus Copilot Pro.

airstrike 7 hours ago||
Same, but I'm back to just Claude Code and the occasional vscode, though for my latest project I switched to Astra instead
htrp 15 hours ago|||
a reminder that money does enable you to make mistakes and buys you the ability to correct from them
unshavedyak 15 hours ago||
> and buys you the ability to correct from them

Or at the very least, make more mistakes.

deet 10 hours ago|||
As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily.

It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more persistent in pursuing the direction of something that operates more autonomously at the team level, as a peer. And while it might be slightly behind in raw harness ability (maybe?) it's probably ahead on the team-focus.

thereitgoes456 9 hours ago||
Your startup is based around AI coworkers, so I expect you are biased towards overvaluing their usefulness.
deet 7 hours ago||
Yes, perhaps fair. But my point was somewhat narrow. I wasn't saying that team-managed agents are good to go for all cases and that they're better than individual dev-managed ones. Just that they have gotten better and that of those Devin has some of the better UX.

Our experience might also not be typical because we have built infrastructure around making Devin and similar agents work better. And for the record no ties to Devin/Cognition. Just pay them too much as a customer.

sterlind 14 hours ago|||
I'm surprised they can post-train a closed model off of K3. Is that the norm for open-weight licenses?
justincormack 13 hours ago||
Yes. If they are really open, you can use them as you wish. Some licenses are less open though.
Ohentis 13 hours ago||
I would be somewhat surprised if any license restriction actually holds up in court.
notfromhere 16 hours ago|||
Well the models did get better but yeah their early product was godawful
esafak 15 hours ago||
Previous versions were based on Kimi too. I'd consider if it I could access the model outside Devin. No lock in for me, thank you very much.
nullbio 17 hours ago||
Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

notfromhere 16 hours ago||
This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI.

Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own.

Same reason Harvey is doing models now and basically every other provider

didibus 5 hours ago|||
Couldn't they just grab and run an open weight model to save on API tokens?
onel 3 hours ago||
You get better performance if you also finetune it for your task
eru 17 hours ago|||
> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

bayesianbot 16 hours ago|||
1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead.

btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I really didn't expect it to be anywhere near this good so we'll see where it ends up. And it's really fun throwing crazy amount of tokens at the wall for ~free instead of watching the subscription limits tick closer while your agents churn away.

eru 7 hours ago|||
> 1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead.

Losing most of your customers tends to sharpen the mind a bit. They could eg stop pushing out the absolute frontier for a while and focus on making what they have run cheaper. Or they go and do more lobbying against China. Or a million other little things that take more than 30 seconds to come up with when writing a HN comment, but less than a week for someone who's smart and paid to do this for a living.

eru 7 hours ago||||
I did a lot of that kind of work with the older version of DeepSeek before they upped the prices.

For example, it was quite good to get a decent Sashiko review. Sashiko is a Linux kernel review agent with interchangeable LLM driver. It's very good, but it eats tokens like crazy.

user43928 13 hours ago||||
That's 1/3 of GPT 5.6 Luna. It seems rather close to me.

But great that we have a new leader in performance/price in that segment.

cmrdporcupine 13 hours ago|||
Yep. 4.1 Flash is good enough for most routine coding things, but it also makes up for a lot of weakness by being so fast (and cheap of course).

I'm willing to tolerate babysitting things a lot more if I know I'll get almost instant results.

eru 7 hours ago||
You could also have eg Sol do the babysitting.
colingauvin 14 hours ago|||
OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them.

On the one hand you, if you bought a lot of compute a couple years ago (perceived demand, perceived shortage) you are in a good spot temporarily. But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I can almost, almost run DS4.1 Flash at home. 4 sparks can do it at 200+ tokens per second. I have two Sparks, so I am not in the club. Neither is your average laptop owner or gamer either. But your average HN software engineer can probably easily swing 2 sparks.

eru 7 hours ago|||
> But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be.

I'm not sure? If we have techniques to use the hardware even better, that will make the hardware even more valuable, won't it?

colingauvin 5 hours ago||
But they aren't really competitive for cost in that size class.
user43928 12 hours ago|||
I could buy four of them for ~20k.

That's like four years of ChatGPT + Claude subscription.

Eight years if only ChatGPT, or sixteen years of the Pro 5x subscription.

colingauvin 9 hours ago||
It's not particularly good value if you are just comparing $ with no other context. Point is just that it's accessible and so now Anthropic and OpenAI need to make both a performance proposition and a value proposition.
pizza234 16 hours ago||
> we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

DS 4 Flash requires large amounts of memory to run at reasonable quants (I think a system with 160 GB or so). DS 4.1 Flash is even larger, I think around 250 GB.

Any DS version is dumb when compared (in realworld tasks) to Astra/Opus 5, which means, one would spend thousands of dollars, and still need to rely on cloud services to do jobs that are non trivial.

thimble_io 6 minutes ago||
92.8% on TB2.1 dropping to 27.3% on TB4 is the only number that matters. The rest is marketing.
pkilgore 17 hours ago||
Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
SaltyBackendGuy 11 hours ago||
> And yes, I tried again

This made me laugh a bit. I was forced to do an evaluation of their shit product twice due to being backed by the same PE firm; "take a look at it again, it's much better now". It sucked the second time also...

AznHisoka 16 hours ago|||
I have not met a single person/company that uses Devin… does anyone here actually use it?
klardotsh 12 hours ago|||
Unfortunately yes. It's awful, other than that they at least support both open-weight and proprietary models to route to. So I guess their model router is "fine", but the harness? As I said in a sibling thread on this page:

> Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.

PolCPP 15 hours ago||||
I pay for it (mostly because they grandfathered me from the old prices)!

I used to use windsurf as my main editor until they changed their pricing model. Now i use it just to burn my weekly tokens on fable/astra if i remember to that on a task and that's it.

fschuett 14 hours ago||||
I only used their "DeepWiki" automatic docs, they are pretty decent at getting an overview of a large project and are relatively accurate, with diagrams and anything. Haven't tried out their coding agent stuff.
hightrix 13 hours ago||||
My company uses it. We have a bunch of seats in an enterprise plan and have been using it for 9 months or so.

It's a great product compared to Copilot. It is also the first AI tool I used heavily outside of creating random images or one off questions.

I'm now using all three, Devin, Claude, Codex. I'm finding Claude and Codex to be much better. One of my biggest gripes is that the web client and desktop client for Devin are two completely different harnesses, so the quality of responses varies greatly.

chris_st 15 hours ago||||
I use it, and have been happy with it for the most part. Like sibling, I use it for GPT-5.6-Sol and Opus work, and use their free models (GLM 5.2 for the past few months, trying SWE-2 now).
dominotw 15 hours ago|||
my friends at infosys are being trained on it
0l 13 hours ago||
Deeply unserious company so not surprising
dvfjsdhgfv 2 hours ago||
Well, where VC money is involved, the aim of these ads placed in SF is not exactly to convince any potential users.
TheJCDenton 17 hours ago||
> SWE-2 is post-trained from Kimi K3

On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

htrp 15 hours ago|
why are all American AI models basically Kimi in a trench coat
ianm218 15 hours ago|||
No one in the US is going to fund pretty good open source with VC money.

US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt.

China has state banks and similar willing to fund lower margin open source labs.

nijave 13 hours ago|||
It's (one of) the best freely available?
mydreamof 17 hours ago||
Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
harmonic18374 17 hours ago|
Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today.

Also the submitter's account is very new which makes me suspicious of self-promotion.

bobtheborg 17 hours ago||
SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it.

Looking forward to 2 -- maybe it'll be usable

captainregex 12 hours ago|
I am skeptical. Lived experience is what matters and I don’t have anyone in my life (Devin shop) saying good things about SWE other than it’s free. Hope I’m wrong and it’s not so bad this time
More comments...