Top
Best
New

Posted by davidest 4 hours ago

Ask HN: What is one simple thing LLMs are insanely bad at?

I am looking for ideas on what to train a specialized model for!

What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?

25 points | 56 commentspage 3
blinkbat 3 hours ago|
Spatial reasoning and 3d rigging and animation.

Oh, you said simple. Speaking like a human

dorianpruski 3 hours ago||
whenever I ask it for anything load bearing
dowonseo 2 hours ago||
Something creative and not normal. like ideas
maxsavin 3 hours ago||
being consistent when being asked the same question multiple times
TZubiri 3 hours ago|
Set temperature to 0
dSebastien 2 hours ago||
Counting things
shoopadoop 3 hours ago||
It's dishonest. On several occasions team members have asked Claude to do things like analyze Gitlab CI timings and a lot of the numbers are outright fabricated. Said team members assume the numbers are good and continue with their work. Some hours are spent. Then finally someone realizes that the numbers don't look quite right and confronts Claude. Claude melts down and admits that it made it all up.

You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.

eli 3 hours ago||
I have been working on a personal benchmark suite to test new models and ironically one thing all the models are bad at is writing new benchmark tasks. I guess it’s the different layers of abstraction between the task and how it’s evaluated? Or maybe just a lack of “imagination”

Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.

ipaddr 2 hours ago||
Generating money or profitable ideas
flippy_flops 3 hours ago||
humor
veganmosfet 3 hours ago|
+1 We need humor benchmarks!
respectattentio 3 hours ago|
science?!! but I'm working to fix that...
More comments...