Top
Best
New

Posted by nicowaltz 11 hours ago

Jeeves. Reasoning improves Jev-like decision models(github.com)
214 points | 85 commentspage 2
zerop 10 hours ago|
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
s2l 9 hours ago|
see: https://news.ycombinator.com/item?id=49883844
HarHarVeryFunny 5 hours ago||
The number of people who feel the need to try and argue that you don't need Jev can only be astroturfing by those with something to lose - Anthropic and OpenAI employees.

Like it or not, companies are going to use Jev unless you can offer something just as cheap and fast.

I wonder just how much of the business automation market, previously held by LLMs, is at risk here?

swingboy 8 hours ago||
Any good classifiers like this or Jev that support image input?
rgbrgb 4 hours ago||
openai's new Decisions API looks to be targeting that https://openai.com/index/devday-2026-recap/
quantized_state 8 hours ago||
I'd assume this would work with Qwen's image encoder probably better after a bit of tuning
druskacik 7 hours ago||
How's the performance compared to ordinary 9B LLM with structured outputs? Both accuracy and speed?
RamblingCTO 10 hours ago||
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.

But funny that jev is getting its lunch eaten apparently in under two weeks?

danieltanfh95 8 hours ago||
it just a classifier. I guess we have to thank typesafe for spending VC money on marketing classifiers as decision models instead.
santadays 8 hours ago||
Doesn't the fact that it's general purpose warrant a new term? It's partly that it doesn't need to be trained, but it's also able to play games based on game state, I'd imagine it would be hard to train a classifier to do something like this because you'd need to represent a good distribution of all the states. The general purpose llm world understanding underneath it allows for this.

I've used it to do web research where it follows the most appropriate links, decides what to record in state, etc. I struggle to see how you could implement something with a classifier. That said, I have no idea how deep the technology is and it might be replaced with open source pretty quickly since its drafting of the frontier models and the open source models seem almost as good.

I like the term decision model and I think it's warranted.

nico 7 hours ago||
> I'd imagine it would be hard to train a classifier to do something like this because you'd need to represent a good distribution of all the states

Yes, one general classifier would be very hard to train. However, you can create a sort of ensemble of classifiers, each trained in different tasks

I’m currently experimenting with this. So far I’ve combined classifiers for 13 different datasets, my target is 95 (the ones Laya used for training)

svachalek 6 hours ago|||
This isn't eating Jev's lunch. This is someone who doesn't understand the entire use case of Jev replacing it with something that doesn't handle it at all.
pavlov 10 hours ago||
It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs.
woadwarrior01 10 hours ago||
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
loclol101 8 hours ago||
How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
Naitik88 10 hours ago||
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
AnodicElegy 9 hours ago||
I'm surprised we haven't seen a "Jehovah" yet.
jadar 9 hours ago|
With the amount of talk about "inventing god", I'm surprised too.
mxkuzn 10 hours ago|
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
More comments...