Top
Best
New

Posted by 1vuio0pswjnm7 1 day ago

AI chatbots give wrong answers to financial queries 'most of the time'(www.ft.com)
154 points | 88 commentspage 2
sharts 8 hours ago|
So do a lot of financial planners / advisors
mizzao 22 hours ago||
Is this just because LLMs can't do math directly? If so, they can certainly write scripts to do math though and those will be a lot more reliable for queries.
sixtyj 1 day ago||
I would prefer to use agent-assisted python scripts that chatbot.
k7peak 1 day ago||
Agreed, this works really well for me. Double check the math/python, execute many times without a LLM that can change o
netless 1 day ago||
[dead]
includenotfound 1 day ago||
From the official report:

> Since LLMs can give different answers to the same question, each question was run five times. That means, each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.

> All models were given the same zero-shot format. They were not given worked examples, previous conversations, hints or an opportunity to correct their answers. This is to make it as similar as possible to a response to a question from consumers.

As for the evaluation itself:

> Responses were checked against this (using an LLM-as-a-judge), and was only given a pass if every element was met; otherwise it was assessed as a fail. This all-pass approach was intentionally strict, so that the score measures whether an answer is complete enough to meet the expert legal standard, rather than how many individual points it gets right.

It's just AI slop and it should be taken with a mountain of salt.

nicce 1 day ago|
> It's just AI slop and it should be taken with a mountain of salt.

Can't you see the irony. You are defeating the argument that LLMs are incorrect or weak with low effort with the term "AI slop" that itself is a narrative that AIs produce weak outputs with low effort.

johnnienaked 1 day ago||
They actually produce weak outputs with extremely high effort
in_absentia 1 day ago||
Now, compare this to a recent story that seemed to claim the opposite:

https://news.ycombinator.com/item?id=49139102

I don't have the time to review the underlying research and decide which one is more correct. My personal biases make me want to believe the current one. Your personal biases may be pulling you in the other direction. How do we make the conversation more intelligent than that?

sokoloff 1 day ago||
One way would be by “[reviewing] the underlying research and [deciding] which one is more correct.”
in_absentia 1 day ago||
And yet, looking at the thread the day after, no one has done that before giving their take.
Dilettante_ 1 day ago||
>"What's 1+1, and don't say 2"?
chilmers 1 day ago||
A company selling combined human + AI financial advice finds that AI advice alone is unreliable? Color me surprised.
rsynnott 1 day ago||
> Younger investors are also more likely to trust AI than financial influencers or TV shows

I mean, you absolutely should not trust any of those things on financial matters, bloody hell.

casey2 1 day ago||
This seems like undisclosed paid stealth-advertising for Thomson Reuters' new model.
simianwords 1 day ago||
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much.

There's no reproducible set either. I'm not gonna trust this report.

stymaar 1 day ago|
Most people[1] interacting with chatbots don't have a paid subscription and they do interact with the free-tier LLMs that are Luna and Haiku, so I still think it's relevant.

[1]: not on HN obviously, but IRL, and probably among FT's readership as well.

cillian64 1 day ago||
A free claude account with no subscription gets you access to sonnet and I believe uses it by default over haiku
74gee 1 day ago|
Well duh! If it's not using tools to look up the state of the market empirically it's not likely to be accurate financially.
More comments...