Top
Best
New

Posted by jasondavies 5 hours ago

Clef: Open-source decision models, and new RL fine-tuning platform(blog.cloudflare.com)
357 points | 142 comments
manlymuppet 4 hours ago|
Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

And it's only been a few weeks.

segmondy 19 minutes ago||
A lot of people claim to have made better than jev, there's a jev benchmark, I have tried many of those models and they eventually end up failing, a non trivial task which doesn't seem like much but reminds me of the svg pelican bench is games, have one of these decision/classifier models play a game, hook it up to the input, most of the ones that are supposedly on jev level end up playing a terrible game, showing that they are very narrow. Cloudflare doesn't compare to the top open bench alternatives, I just finished downloading it and will compare it to jev for non trivial tasks tonight.
slopnt 4 hours ago|||
They have to have decision models already in production. Part of their business is detecting bots, DDoSers and spammers.
alightsoul 2 hours ago||
Yeah that's a decision tree, random Forest or some other machine learning classifier. They have existed for a long time
smallmancontrov 2 hours ago||
I'm all for rebranding discriminative models as decision models, though.

"Discriminative" always had pointlessly bad optics, but I knew it was over when I started seeing prominent machine learning researchers who p=100% knew better describe discriminative models as generative because that was the buzzword of the year. "Decision model" sells the value proposition much better and doesn't sound like an anti-woke crusade.

TeMPOraL 4 hours ago||
It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

seizethecheese 4 hours ago||
Name a few of these low hanging fruit left around.
TeMPOraL 4 hours ago|||
Jev is one.

Diffusion transformers are not "easy" but underfunded.

Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

Or imagine automated sliding doors that don't suck.

--

[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

mdp2021 2 hours ago|||
There are "low hanging fruits" - easier to achieve goals -, and there are super-fruits, milestone-fruits.

Among the most important ones:

-- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?

-- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.

-- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.

The long-term direction we got into must lead to this.

(You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)

patcon 1 hour ago|||
If anyone is interested, following Dr Michael Levin's Thoughtforms.life podcast is the cutting edge of where all this previously fuzzy stuff is becoming more concrete. So long as you can tolerate distinguished scientists flailing about as they discuss consciousness and life and developmental biology (and other less-obviously living things, like algorithms) as involving "free lunches" and "ingressing patterns from the platonic realm" :)
TeMPOraL 1 hour ago|||
Those are the absolutely fascinating parts, and I sincerely hope AI won't get out of control before we're able to tackle some of these.
blurbleblurble 3 hours ago||||
Diffusion models combined with these new looping techniques are gonna change the whole conversation about efficiency. Imagine control net but in one or more conceptual latent spaces.

But also harnesses and more generally new insights on "the control flow problem" could end up squeezing a ton of performance out of small models.

seizethecheese 2 hours ago||||
I commend you for actually answering, independent of what I think of the answers.
dominotw 1 hour ago||
not a very good answer though
flipping_beacon 3 hours ago||||
Definitely agree with edge computation, although inference extensively researched and funded if SOTA LLMs hit a dead end tomorrow,there is still a lot to explore and research in inference and edge computation
Amekedl 3 hours ago||||
yeah your reply, nobody can predict the future.

Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7

aeve890 3 hours ago||||
>Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

That's low hanging for you?

Twirrim 20 minutes ago|||
We already have examples of LLMs running 16k+ tokens a second using custom ASICs.

It's down at the moment (Not sure if it'll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens/sec. It was amazing to use, you'd no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.

I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we're currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.

We really don't always need newer better faster stronger models, there's quite a lot of room for "good enough" where getting 17kt/s at significantly lower power would be amazing.

[0] https://chatjimmy.ai/ [1] https://taalas.com/

AshamedBadger56 4 minutes ago||
>I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.

It would be interesting to pair the super fast model with a normal speed model. Have the super fast one do all the background research, code writing, etc. The normal model would just relay the needed info to you at a more reasonable pace.

guyomes 3 hours ago||||
If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document

mdp2021 1 hour ago||
> hardware dedicated to a specific LLM

That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...

msdz 3 hours ago||||
Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.
TeMPOraL 3 hours ago||
Yes. Point is, it's untapped only because everyone is running in the race (even if out of curiosity), and there's just not enough people with means to tap into these side threads. For the past few years, there's been many interesting papers that circulated the industry, got recognized as worthwhile pursuits, and then dropped because running behind the Big Vendors had massively better ROI.
mdp2021 2 hours ago||||
> That's low hanging for you

An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.

blurbleblurble 3 hours ago||||
It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.
TeMPOraL 3 hours ago||||
Yes. It's well within realm of possibility, but so far wasn't pursued because the Big Vendors went all-in into capability growth (rightfully testing "the bitter lesson" to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and/or starstruck with fascination, exploring what these models can do.

This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.

blurbleblurble 3 hours ago||
Just like renewable energy and so many other things. Hyperconcentration of capital is really tragic. I hope things turn around.
TeMPOraL 2 hours ago||
They will. That's the fallacy of the "S-curve" everyone likes to commit these days actually giving a positive outlook.

Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).

In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.

ekabod 3 hours ago|||
That's a high hanging fruit, not low.
alightsoul 2 hours ago|||
[dead]
sroussey 3 minutes ago||||
Just look at all the model type on hugging face. LLMs are a small percentage.
hobofan 4 hours ago||||
Closely connected to decision models: A good library to do ranking based on pairwise ranking on multiple attributes. By using a decision model (especially one that can make decisions on multiple fields at the same time) this becomes a lot faster and more powerful. Could make for a pretty nice search reranker as well as prioritizer for many problems.

Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.

sarkarghya 2 hours ago||||
I can imagine advancements on making smaller models work together better instead of a generalized core. Imagine a community or city having a https://pirateface.co/ so that the shard of the model that you need can be streamed in with minimal latency with your box only holding the minimal version (say deepseek v4 flash as orchestrator) of the model that you use on day to day basis.

We have overcome split brain problems before so this wont be our first

esseph 2 hours ago||
You're putting a lot of trust in uncorrupted, untrusted, unknown models (potentially).
murkt 4 hours ago||||
Easy to reach doesn’t automatically mean “easy to see”.
CamperBob2 2 hours ago||||
An example I like to use is: compare the quality and scope of games released with a brand-new console to the ones released for that console towards the end of its life, when everyone has learned how to take advantage of whatever weird, wacky hardware Sony invented for that console generation.

There is still a lot we don't know about how to get the most out of existing LLM components from a speed or cognitive-performance perspective. People could easily spend the next decade studying and refining what's been built so far, even if no new, original approaches ever arrive.

meander_water 8 minutes ago||
> This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason...

This seems misleading. Decision models do not produce deterministic output. Repeated calls can product different decisions just like an LLM with structured outputs.

vulture916 1 hour ago||
Jev = $0.042/m input, output free Clef = $0.24/m input, no output price listed

At 300 tokens per call, you'd get:

One million decisions on Jev cost about $12.60. One million decisions on Clef cost about $72.

Would probably make sense to self-host Clef, if you have the capability/resources. If not...

scronkfinkle 1 hour ago||
> no output price listed

It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input

jampekka 1 hour ago||
Output for decisions have so few output items (not really tokens here) that they are negligible anyway. Jev hyping "free output" is almost lying by omission.
buildbuildbuild 4 hours ago||
Open weights, not open source.

The weights have permissive licensing, but the data and training pipeline are not published to reproduce them from their proprietary Qwen starting points. Weights are not "source."

jMyles 4 hours ago|
Came directly to comments hoping not to see this one.

<sad trombone sound>

Surely someone will soon do what the title of this post makes it seem like cloudfare did. Truly modular open source training and inference logic, along with a totally open corpus and weights, will eventually out-compete the closed ecosystem.

ainch 1 hour ago||
There are some groups doing it for LLMs - like the Allen Institute for AI's Olmo models and Eluether AI's Pythia.
ssiddharth 4 hours ago||
Pricing is $0.24/million input tokens which is ~6x compared to Jev. Clef-flash is at $0.09 which is way more competitive.
CBLT 3 hours ago|
Yeah I also thought it was strange their pareto frontier didn't include cost.
bityard 4 hours ago||
Clef is based on Qwen3.8-27B and Clef-flash is based on Qwen3.8-9B (edit: actually Qwen3.5-9B). So, similar in spirit to Kev by my understanding, but based on a newer model.
NitpickLawyer 3 hours ago||
> and Clef-flash is based on Qwen3.8-9B

There is no official qwen 3.8 9b

From the model card:

> Clef-Flash is post-trained from Qwen/Qwen3.5-9B. See Clef for the larger variant.

bityard 2 hours ago|||
Thanks, I missed that. Fixed my comment.
ddarolfi 3 hours ago|||
It's based on Qwen3.5-9B, maybe a typo
okpatil 2 hours ago||
Atom is 60M Param (around 133x to 400x smaller).

16ms latency. And locally run.

https://at0m.pienomial.com/

Why go big when you can go small ?

kamranjon 2 hours ago||
Cause it's not open?
okpatil 2 hours ago||
Good point.

To counter, most of the AI is not open. So is none of Microsoft Products. As long as they work, we keep using them.

mrkn1 3 hours ago||
For smaller scale decision model that runs on CPU, check https://news.ycombinator.com/item?id=49923223
fastball 2 hours ago|
tbh saying all these dumb decision models are similar to Jev is like saying markov chains weren't far from GPT-2.

The value isn't really in the I/O shape, it is in the intelligence combined with the output shape. Every extra ounce of intelligence in these models unlocks additional use-cases. But the converse is also true: a dumb decision model is going to be less useful than using a more intelligent standard LLM.

That is the appeal of Jev: for certain usage it has more intelligence than some small SOTA LLMs. It is the first decision model that actually feels intelligent (to me).

mrkn1 4 minutes ago||
That makes sense, I agree. I still think there might be use cases when people don't want to use the Jev API, and end on a different trade-off.
amluto 2 hours ago||
I’ll go out on a limb and suggest that I don’t think a Jev-like model is particularly useful unless you can fine tune it. The Jev API has zero ability to pass in a prior [0], and, if you can neither pass in a prior nor fine tune for your system, you will get an output that may be almost meaningless.

I’d love to see someone build a model of this sort that can actually accept priors and do something intelligent with them.

[0] You can feed Jev a prior as text. I’ve tried it. It works poorly.

mikeocool 2 hours ago||
It seems like jev's major advantage over existing classifiers is that I dont have train it.

If I have to gather and tag data to fine-tune Jev, I can probably just train an "old school" classifier model and make it even cheaper, faster, and just as accurate.

jdthedisciple 22 minutes ago|||
I suppose a sort of prior-proxy can be encapsulated by a carefully written system prompt.
sheepscreek 2 hours ago|||
Also one of the more interesting features of Jev is the confidence rating that hardly any Jev-cc talks about.
amluto 2 hours ago||
It seems interesting to me only in the sense of being useless. From the horse’s mouth:

> Confidence is derived from the probabilities

https://docs.typesafe.ai/confidence

(Why is it much easier to find AI-slop websites quoting this than it is to find the actual documentation?)

My inner Bayesian would like for Jev to provide something resembling “evidence”, although I admit that one might ask Jev questions that are somewhat awkward to treat as typical Bayesian questions. If I ask “will this PR be merged”, it’s kind of strange to contemplate the probability of a PR conditioned in that PR being merged in the future. But I bet there is a way to formalize a prior-free classifier in a way that makes Bayesians and non-Bayesians happy, possibly involving actual learned probabilities and confidence levels. If you read the literature on scoring rules, you will find that classifier scores do somewhat naturally decompose into a few interpretable terms.

brokensegue 2 hours ago|||
I think better than priors would be a closed loop where you tell it what the right answer was (or some signal) and they monitor and fine-tune for you
okpatil 2 hours ago||
We were able to completely automate 20,100 token prompts with At0m[https://at0m.pienomial.com/].

We believe entire compliance workflows (even multilingual) could be automated.

Would you like to get a demo ?

sheepscreek 2 hours ago||
You’re coming on a bit strongly - a couple of comments with a link is sufficient. Before trying to sell, try to genuinely further the conversation, provide some useful knowledge in return for the reader’s attention.
okpatil 2 hours ago||
Point taken. Let me explain if you allow me.

It is possible with deterministic decision models, such as At0m, to gauge the probabilities at every decision. This behavior in addition to hard coded logic, it is possible to completely replicate a prompt's logic.

Using Fable 5.1, it is a matter of minutes.

I believe that most of the compliance check documents will be a solved problem, 3-6 months in future.

None of the LLMs can do it.

Hence I asked to the comment poster if he would want to demo, so that I can show it to him, how to do it step by step. By bad, if it came out too strongly.

fooker 3 hours ago||
This is awesome.

I bet the competition will result in research into how to make these decision models several more orders of magnitude faster and cheaper.

Here's a challenge problem - look at a 1M context window and produce N decisions (different queries) from it in 50-100ms.

yipinwong 4 hours ago|
A question someone not trainined in AI/ML field, Is a decision model that easy to crete that there are floods of these JEV alternatives already?

Or are companies/people already building this based on say an arXiv docs? n

---

The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding

TeMPOraL 4 hours ago||
Yes, it's easy. The thing people are missing (especially those believing AI is a "dead end" and "not transformative") is that the field has been advancing so fast in the past few years, that there's lots of such unexplored avenues, unpicked low-hanging fruits, that everyone just raced past. We've barely begun exploring the capabilities ML brought us - patterns, applications, and architectures.

Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.

nico 4 hours ago|||
The basics are pretty simple. And depending on what your specific need is, the model can be really really basic, fast and super effective (ie. run on a mobile device and process thousands of requests in <100ms)

I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests

Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window

calebkaiser 4 hours ago|||
There is a bunch of stuff to tease apart.

In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.

A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.

I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.

But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.

The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

tomrod 1 hour ago||
> The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

If CF's benchmark is representative and sufficient, Clef outperforms Jev!

Models by themselves don't guarantee market capture. Rather, its how they integrate. I think a lot of folks are burned by the closed nature of many models.

conmod278 4 hours ago|||
Live coding Jev from Scratch | Understanding Qwen architecture

https://www.youtube.com/watch?v=AzxoU7kxjig

orbital-decay 4 hours ago|||
Yes it's easy for an established shop, all they need to do is to tweak the post-training workflow. "Decision model" is the same kind of marketing as "LRM" attempted by OpenAI when RL CoT was new (to hyped up crowd). It's still fundamentally a classifier used for "decision making", games and RP were using generalist models and constrained outputs to do what the DOOM demo does for years.
janalsncm 3 hours ago|||
The interesting part is also the easy part. The model and architecture are not hard for an experienced machine learning engineer to build.

The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).

XCSme 4 hours ago|||
You can make a basic one in minutes based on existing open-source models.

Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.

Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.

It's not really a new technology, it's more like a new use-case.

sigbottle 4 hours ago||
What even are these new "decision models?" Take an existing LLM, feed it a prompt, force it to pick a choice; decode is 1 token (or rather, the whole logit set for only that last token; token implies selecting one logit) so you made a choice. That's it?
redox99 3 hours ago|||
Yes, although you probably want to calibrate your model if you want the probabilities to actually be meaningful.
orbital-decay 3 hours ago||||
Yes but optimized specifically for the purpose. Using that for "decision making" is also not a new use case, but turned out to be new to many people. Which is great, I hope they make something cool with it!
popinman322 3 hours ago|||
That's how Cygnet handles it.

https://github.com/blockbrain-ai/cygnet-recipe

redox99 4 hours ago||
Yes it's very easy if you have fairly basic ML knowledge.
More comments...