Posted by 1vuio0pswjnm7 1 day ago
very popular on Earth.
https://www.financialreporter.co.uk/ai-models-give-wrong-fin...
Much of the testing is on Haiku and Luna, and criticizing the quality of free AI (!). But they do claim Opus 5 with reasoning still failed 39% of their financial questions.
Getting results requires a harness like in coding and objective metrics, like tests.
Yet every time I open this website someone is trying to sell me that chatgpt solved abstract mathematics.