Posted by 1vuio0pswjnm7 1 day ago
> Since LLMs can give different answers to the same question, each question was run five times. That means, each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.
> All models were given the same zero-shot format. They were not given worked examples, previous conversations, hints or an opportunity to correct their answers. This is to make it as similar as possible to a response to a question from consumers.
As for the evaluation itself:
> Responses were checked against this (using an LLM-as-a-judge), and was only given a pass if every element was met; otherwise it was assessed as a fail. This all-pass approach was intentionally strict, so that the score measures whether an answer is complete enough to meet the expert legal standard, rather than how many individual points it gets right.
It's just AI slop and it should be taken with a mountain of salt.
Can't you see the irony. You are defeating the argument that LLMs are incorrect or weak with low effort with the term "AI slop" that itself is a narrative that AIs produce weak outputs with low effort.
https://news.ycombinator.com/item?id=49139102
I don't have the time to review the underlying research and decide which one is more correct. My personal biases make me want to believe the current one. Your personal biases may be pulling you in the other direction. How do we make the conversation more intelligent than that?
I mean, you absolutely should not trust any of those things on financial matters, bloody hell.
There's no reproducible set either. I'm not gonna trust this report.
[1]: not on HN obviously, but IRL, and probably among FT's readership as well.