Top
Best
New

Posted by albelfio 23 hours ago

Introducing System One Models and Jev(typesafe.ai)
1738 points | 464 commentspage 14
nightshift1 18 hours ago|
The whole page reads like it was vibe-written by an AI. If I'd built something as disruptive as this claims to be, I'd have spent at least fifteen minutes writing the announcement myself. Every time I see 'we' in an announcement like this, I picture one guy alone in his basement.
CompleteSkeptic 18 hours ago|
unfortunately all hand-written :( my chief-of-staff does unironically handwrite em dashes though
lwansbrough 18 hours ago|||
For what it’s worth, I didn’t get that impression, and even noticed a couple typos ;)
nojvek 16 hours ago|||
I really appreciated the hand-written release. Thank you.
pennomi 23 hours ago||
> Extraordinary claims require extraordinary evidence so see below for the receipts.

Yes, that’s the kind of attitude I want to see in these model releases

ramon156 22 hours ago|
But the evidence is not there...
pennomi 22 hours ago||
Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.
simianwords 22 hours ago||
They gesture at not using benchmarks for some reason...
meric_ 21 hours ago||
https://typesafe.ai/blog/antibenchmaxxing

But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do

Wazzymandias 17 hours ago||
This looks and feels a lot like productionized conformal prediction
elcomet 21 hours ago||
The technology and the results are very handwavy. What is RLCD exactly ? What are scores on benchmarks compared to LLMs ?

This website does not inspire confidence at all, it all sounds like a marketing piece. I wish it was true, some kind of text-prompted classifier with LLM performance would be cool, but I can't trust it with what we are given.

hi_hi 22 hours ago||
If I’m understanding correctly, this will work well for self driving cars?
darksaints 21 hours ago||
Okay, so it doesn't output text, that much is understood. What are the inputs like? I'm assuming maybe a text input? maybe an AST definition? Really hard to tell how this works at all from the demos, especially since we can't really try it out.
CompleteSkeptic 21 hours ago|
inputs are structured program state. there is an example at around second 30 of the doom demo

(though ideally everyone gets off the waitlist and can try it out for themselves )

andai 22 hours ago||
Why did they pick the name System One? It's not really explained what "System One tasks" and "System One shaped queries" are. Things that need a fast response?

Does this imply it's a very small model? I couldn't find anything about the model itself.

oblio 22 hours ago|
Maybe: https://thedecisionlab.com/reference-guide/philosophy/system...
hunterbrooks 22 hours ago||
Bingo. It's a Psychology term for the part of our brain that reacts instinctively rather than thoughtfully and logically
zmmmmm 21 hours ago||
The eval is baffling me

> we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. ... Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).

They assume there is a correct graph, but they don't compare to that, they compare to the average of the smarts models? So the smartest models are getting it wrong but you compare that anyway as a benchmark? So the outcome is "how much of a Fable am I getting" etc. Why not compare the actually correct thing?

But then even on this hand constructed eval, the first plot is showing Jev at less than Sonnet 5 accuracy. It is barely better than Luna. There are two Opus 5's and two Sonnet 5's without explanation. What is the plot showing?

I gave up.

_boffin_ 18 hours ago||
Any relation / inspiration to GLiClass?
bfeynman 22 hours ago|
Super intrigued by this - large scale automation using LLMs is quite annoying due to deprecation cycles of models from frontier labs and cost of running your own being prohibitive when you have a blend of them.
More comments...