Posted by bashbjorn 17 hours ago
About Jev:
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities (even though they are not always correct).
Only 99% correctness! Borderline unusable!
About their model:
> It classifies: it gets a prompt with choices and outputs probabilities.
You want numbers, it gives you numbers! What more could you want?
Since most chat models want to answer with a human-readable message i think their logprobs are not as meaningful. It would be interesting to see if one choice is like "correct" and if the model wants to choose it more often, cause it might not answer the question but to prose to the user.
is a legal cya a la "Nathan For You" 's Dumb Starbucks
You can test Jev like model at 26B parameter count here (built few weeks ago): https://gambler-relay-us-west1.leo-fish.ts.net/demo (might not stay up for long)
Typesafe compatible API
This is just running on old hardware.
Speed and cost are obvious reasons, but isn’t this a tradeoff?
<think>\n\n</think>
but letting an LLM think would trade latency and performance for significant reliability above that of Jev.Why do I have to feed my e-mail into the model?
name = “qurren” print(f”hello {name}”)
ok