Top
Best
New

Posted by nandakishor_ml 14 hours ago

I built non-autoregressive decision models with RL a year ago(laya.convaiinnovations.com)
1060 points | 249 commentspage 5
pknerd 10 hours ago|
Correct me if I am wrong, can I use Jev and this tool for ticket classification? I mean, for instance, a level 1 ticket contains a screenshot of the login page that displays an error, LLM can do it perfectly, can Jev do it?
dgritsko 10 hours ago||
At least for now, Jev is not multimodal. So a screenshot alone wouldn't cut it.
adverbly 10 hours ago|||
I find that a bit interesting because the most system one part of the brain is probably the part used for visual processing.

It's trying to use a human analogy but the analogy breaks down if you try to apply it directly

Reubend 6 hours ago|||
Oh, I missed that! So the Doom demo was potentially just "harnessmaxxing"?
rgbrgb 10 hours ago|||
It can’t do images but it can do a pretty good job of triaging urgency or choosing when to escalate. So I’d guess yes but it depends how detailed your classification is.
someguy101010 8 hours ago||
you can use this for screenshots

https://huggingface.co/thaitea/laya-vision-smolvlm-256m

edot 10 hours ago||
I don’t understand Jev or this. I used this since it’s open source (good job btw!) with the following. State: “a 6 sided die rolled a 3”, question (noul): “Is the number odd?”

Answer: 9% chance, with 91% confidence.

Heh???

Ok, even worse. 75% chance a coin landed heads up?

State: I flipped a coin. Question:

{ "noul_result": { "type": "noul", "instructions": "Did the coin land heads up?" }, "choice_result": { "type": "choice", "instructions": "Determine if the coin landed heads or tails up.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }

Ran on: https://huggingface.co/spaces/convaiinnovations/laya-demo

Result: { "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.6839, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "heads", "probabilities": { "heads": 0.7407, "tails": 0.2593 }, "confidence": 0.1743, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 76, "output_tokens": 0 }, "latency_ms": 93.8 }

Trying to be even more good-faith:

State: "A fair coin was flipped once. The result was not observed. No other information about the outcome is available."

Questions: { "noul_result": { "type": "noul", "instructions": "Given only the supplied state, what is the probability that the coin landed heads up?" }, "choice_result": { "type": "choice", "instructions": "Given only the supplied state, determine which outcome occurred.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }

Result:

{ "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.1265, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "tails", "probabilities": { "heads": 0.2522, "tails": 0.7478 }, "confidence": 0.1853, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 123, "output_tokens": 0 }, "latency_ms": 154.5 }

jwpapi 8 hours ago||
Jev says you should restate state in the question and I tried it:

{ "decision": { "type": "noul", "instructions": "Is the rolled number in state odd?" }, "question": { "type": "noul", "instructions": "Is the number odd?" }, "question-3": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the number odd?" }, "question-4": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the rolled number odd?" } }

=>

decision,0.168,0.83 question,0.141,0.86 question-3,0.029,0.97 question-4,0.021,0.98

so im confused too..

A weakness with numbers?

bensyverson 7 hours ago||
Breaking news: small language models struggle with math
hbrn 2 hours ago||
But didn’t you hear?

> Jev is neither small nor an LLM

ksymph 8 hours ago|||
I don't think calculating mathematical odds from natural language is the sort of problem this is trying to solve. A typical LLM hooked up to a calculator would be more appropriate for that.

Jev (and similar) is more for data processing and sentiment analysis. Moderation, search engines, that sort of thing. Jev has a page of proposed use cases where you can get an idea of what they're going for: https://docs.typesafe.ai/concepts/use-case-map

prometheus1992 9 hours ago|||
try this model on HF - https://huggingface.co/MoritzLaurer/deberta-v3-large-zerosho...

a 6 sided die rolled a 3

possible class names - the number is odd, the number is even

result:

the number is odd 0.945 the number is even 0.055

as someone else said, that 0.055 is probably bc of 6 and 3 being there.

bensyverson 9 hours ago|||
This is not a good faith test of the system.
edot 9 hours ago|||
But it's hallucination-free, isn't it?
bensyverson 7 hours ago|||
I guess we’ve just reached the point where everyone has to state the obvious, and common sense is extremely uncommon.

So here goes: you should not use an AI model to validate a claim which is trivial to calculate deterministically. That is (obviously?) not what a model like Jev is for, thus it is not a good test of Jev.

usagisushi 9 hours ago|||
yeah, technically. (/s)

    python3 - <<'EOF'
    import json, urllib.request
    body = json.dumps({
        "state": "The car wash is only 100 meters away from my house.",
        "model": "jev-1.13-free",
        "questions": {"q": {"type": "choice",
            "instructions": "Should I drive or walk to the car wash?",
            "criteria": {"drive a car": None, "walk": None}}}
    }).encode()
    req = urllib.request.Request("https://opencode.ai/zen/v1/systemone", data=body,
        headers={"Content-Type": "application/json", "User-Agent": "opencode/1.18.31"})
    with urllib.request.urlopen(req, timeout=60) as r:
        print(json.dumps(json.load(r)["answers"]["q"], indent=2))
    EOF
    {
      "type": "choice",
      "choice": "walk",
      "confidence": 0.66,
      "probabilities": {
        "walk": 0.83,
        "drive a car": 0.17
      }
    }
hmokiguess 10 hours ago||
Maybe it’s the fact that “the number” could refer to both 6 and 3 to this model?
beeforpork 8 hours ago||
Is this as good as Laya 3? Unfortunately, it's production was moved from Bremen, Germany, to China, and it is not good anymore, in my opinion.
kburman 11 hours ago||
Loved the idea, but I don’t think it would be able to handle real-world data effectively. There are a lot of nuances that actually require a reasoning model to think through, connect the dots, and make sense of the broader context.
baobabKoodaa 8 hours ago|
If you need a reasoning model, then that is a System 2 decision, not a System 1 decision. This thread is about "Laya", a "Jev" competitor/precursor, which is a System 1 thing.
einpoklum 6 hours ago||
Pangram believes this text was authored with an LLM:

https://www.salahadawi.com/hacker-news-ai-detector/49765348

julianozen 3 hours ago||
Distribution > Creation
srameshc 11 hours ago||
from https://huggingface.co/convaiinnovations/laya > The policy reports a distribution; exploration adds zero-mean Gaussian noise to the logits; the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal questions). Expected reward is maximised only by reporting honest probabilities.
petesergeant 10 hours ago||
There are many, many, open-source versions of Jev, including three distinct projects sharing the name “openjev”

If you’re interested in the basic trick most are using (which is probably also what Jev does) then it’s here: https://sgnt.ai/p/jev/

skybrian 9 hours ago||
This sounds cool but it looks like it requires a GPU that I don't have. Is there an API to try it out?
jamienk 8 hours ago|
Why do we ("society") need the "frontier" companies at all? Their business goal has settled on trying to CONFUSE the shit out of us so that we don't understand the big pictures about various aspects of AI.

THANK YOU, Nandakishor Mukkunnoth, for putting in the work to help to clarify this stuff!

You are like a firefighter compared to their fire-insurance racket.

More comments...