Posted by nandakishor_ml 13 hours ago
“Claude, roast this noob, tell him that his model isn’t novel or frontier —”
both in unison “— and make no mistakes!”
It’s all so tiresome
The implosion of hype after the .com crash was actually kind of a ... relief.
It's trying to use a human analogy but the analogy breaks down if you try to apply it directly
Answer: 9% chance, with 91% confidence.
Heh???
Ok, even worse. 75% chance a coin landed heads up?
State: I flipped a coin. Question:
{ "noul_result": { "type": "noul", "instructions": "Did the coin land heads up?" }, "choice_result": { "type": "choice", "instructions": "Determine if the coin landed heads or tails up.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }
Ran on: https://huggingface.co/spaces/convaiinnovations/laya-demo
Result: { "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.6839, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "heads", "probabilities": { "heads": 0.7407, "tails": 0.2593 }, "confidence": 0.1743, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 76, "output_tokens": 0 }, "latency_ms": 93.8 }
Trying to be even more good-faith:
State: "A fair coin was flipped once. The result was not observed. No other information about the outcome is available."
Questions: { "noul_result": { "type": "noul", "instructions": "Given only the supplied state, what is the probability that the coin landed heads up?" }, "choice_result": { "type": "choice", "instructions": "Given only the supplied state, determine which outcome occurred.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }
Result:
{ "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.1265, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "tails", "probabilities": { "heads": 0.2522, "tails": 0.7478 }, "confidence": 0.1853, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 123, "output_tokens": 0 }, "latency_ms": 154.5 }
{ "decision": { "type": "noul", "instructions": "Is the rolled number in state odd?" }, "question": { "type": "noul", "instructions": "Is the number odd?" }, "question-3": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the number odd?" }, "question-4": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the rolled number odd?" } }
=>
decision,0.168,0.83 question,0.141,0.86 question-3,0.029,0.97 question-4,0.021,0.98
so im confused too..
A weakness with numbers?
> Jev is neither small nor an LLM
Jev (and similar) is more for data processing and sentiment analysis. Moderation, search engines, that sort of thing. Jev has a page of proposed use cases where you can get an idea of what they're going for: https://docs.typesafe.ai/concepts/use-case-map
a 6 sided die rolled a 3
possible class names - the number is odd, the number is even
result:
the number is odd 0.945 the number is even 0.055
as someone else said, that 0.055 is probably bc of 6 and 3 being there.
So here goes: you should not use an AI model to validate a claim which is trivial to calculate deterministically. That is (obviously?) not what a model like Jev is for, thus it is not a good test of Jev.
python3 - <<'EOF'
import json, urllib.request
body = json.dumps({
"state": "The car wash is only 100 meters away from my house.",
"model": "jev-1.13-free",
"questions": {"q": {"type": "choice",
"instructions": "Should I drive or walk to the car wash?",
"criteria": {"drive a car": None, "walk": None}}}
}).encode()
req = urllib.request.Request("https://opencode.ai/zen/v1/systemone", data=body,
headers={"Content-Type": "application/json", "User-Agent": "opencode/1.18.31"})
with urllib.request.urlopen(req, timeout=60) as r:
print(json.dumps(json.load(r)["answers"]["q"], indent=2))
EOF
{
"type": "choice",
"choice": "walk",
"confidence": 0.66,
"probabilities": {
"walk": 0.83,
"drive a car": 0.17
}
}