Posted by robin_reala 2 days ago
Sycophancy is the "stickiness factor" of AI analogous to that of Social Media.
For some background read Jagged Intelligence: The Dangerous Unknowns at the Heart of LLMs - https://news.ycombinator.com/item?id=48577159
Here is an experiment that i did;
Ask AI to build a "Character Profile"(in the broadest sense) of a person based on their available public writings. This requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc.
I know "me" and so i asked AI to use my HN comments/submissions as input :-) The sycophancy/flattery/praise was quiet excessive. I am realistic and old enough to not need ego-soothing (a little is fine but a lot makes me suspicious) and so i asked AI whether these phrases were not too over-the-top. It apologized and agreed to drop the fluff. I then asked it to identify job roles (any) for which i might not be a good fit given my character profile. This acts as an external constraint which focuses attention on shortcomings and hence forces the AI to look at the other side of the coin. Now the results were better and more in line with what some Human may deduce from my HN persona (which obviously is not my complete real-life persona) but still wanting in many aspects.
The above is a perfect example of "Jagged Intelligence" exhibited by AI. Excellent in formal symbolic manipulation sciences, regurgitation and simple reasoning but highly deficient in human-like commonsense reasoning.
I had one prominent "influencer" publicly challenge me to a coding competition over it. It was very weird.
But "OH MY GOD EVERYTHING HAS CHANGED THE OLD WAY IS DEAD 10X MORE PRODUCTIVE" - no. If that were true, we'd notice it in the software all around us.
I've shipped three software projects in the last two months. I shipped zero in the preceding year.
Right now I'm building a system to decompile MAME ROM games into idiomatic JavaScript. It completes one game in about 3 days - complete with extensive commenting. I started it about 2 weeks ago.
I'm not sure "EVERYTHING HAS CHANGED" but if you're not seeing dramatic change, I suspect you're not looking.
If the people claiming that everything has changed were even a fraction as productive as they think, then it would be visible even to people who weren't looking.
As it is, a lot of people are actively looking but still not seeing it.
> If the people claiming that everything has changed were even a fraction as productive as they think, then it would be visible even to people who weren't looking.
with:
> If you can't see this - you should probably get your eyes checked.
He is clearly suggesting I am delusional about my productivity gains, and I'm the one who gets flagged.
I daresay I'm noticing a bias in the moderation.
Don't be a jerk. A lot of people have had bad experiences with these systems from people basically saying exactly what you're saying. This "you should get your eyes checked" mentality is pervasive with the nerds who confidently say that their analysis is solid when it's... not. It's a big problem in communities that have non-conventional videogame puzzles for example. Someone will come in, announce some novel solution.... leading to a four hour fight about the efficacy of LLMs, only to find a mistake in their code the next day.. never to be seen again.
Mind you, I'm not saying that these systems aren't great at reverse engineering and whatnot. They're spectacularly good at that. But RE is a relatively constrained problem because everything is still right in front of you. For complicated problems where the signal is mixed between noise... less so.
Be kinder, friend.
Otherwise, you have a good life as well.
Cheers.
And... I didn't start it.
Cheers.
EDIT: Good choice.
Where's the rest of it?
And your response was "So what - show me another one."
Even granting that framing though, the claimed increase in productivity isn't a one-off, but asserted for everyone using them; you're claiming mass-produced miracles but trying to depend on a single example.
Cite one example of a decompiler that provides English names and comments.
Silent Hill 1 PC https://youtu.be/0niadQJmYx0
Super Smash Bros/other GameCube and Wii games https://www.reddit.com/r/decomps/comments/1uvttvy/moderngekk...
It must take a binary executable and derive English names for the memory addresses.
There are none.
Ghidra does not provide semantic names.
dotPeek shows you the symbols that were not stripped from the .NET.
Neither of these can do what I asked.
Ironically, this is the inverse of the very thing you complained about people doing to you.
That's a pretty far cry from the "evidence" you claim to have dropped on multiple comments.
The claim was this is not possible. I show it happening. That's a proof.
You're confused because the example is my own project. But I'm not merely describing it - I am giving you full access to it to confirm or reject my argument.
Not anecdotal. Saying it three times does not make it true.
If you want to have a whack at a hard problem, have a try at the cryptographic puzzle in Noita if you want to see the silly failure modes of these systems. Lots of people have tried to solve it with AI and nothing's budged. If you can solve it, great.
What does your new subject have to do with the fact that the previous arguments were all pretty much baloney?
Pointing to a problem that has not been solved is completely unrelated to the miraculous things that have been solved.
And I thought you'd decided (three times now) that I'm not worth talking to?
You're right - there are still many problems that have not been solved by AI. AI has not found a cheap and effective way to turn lead into gold, as an example.
But I sorta think that's a silly metric. There are always going to be unsolved problems for any system. The interesting metric is how many problems have been solved - and how surprising are they.
Like I said above - a decompiler that produces semantic understanding of the machine code it's looking at (this address is score; this address is how many lives are left; this address is...) would have been considered literally impossible before LLMs. Today, I was able to implement such a system in under two weeks.
That's an impressive result. If you disagree, I'd love to discuss it with you.
But what I'm interested in is in the edges of their failure modes because those failure modes prop up frequently in pernicious ways. For a similar reason that I want to understand the edges of my own intellectual failure modes by regularly challenging myself. This conversation is an example of that.
It seems like you're just not interested in discussing the edges and failure modes. You call that a "silly metric", I call it, "understanding your tools".
Granted, that's not exactly the claim these days, the claim's goalposts shift constantly, as these things do.
But if you go back to the top of this conversation - I asked why people aren't seeing the miracles (not cute tricks - useful things that simply were impossible before) and I was largely met with the response "you're delusional if you see miracles."
But the thing is - I am not. So I am going to continue arguing that case every time I see a version of it.
So no, I'm not interested in exploring other issues. Sorry. Yours in particular I find as interesting as discussing whether Microsoft Word can be used to edit video files. Not interesting.
But if you need validation, here it is: there are many problems that AI cannot solve. Of course there are. There always will be.
Do you have a weird need to reframe this so you can feel like you won? You shouldn't do that. It's not healthy.
HEH. Not sure why you're implying I have changed my position. I said as much above and would have said the same at any time. I don't think it's controversial in the least that there are things AI can't do.
Do you have a weird need to reframe this so you can feel like you won? You shouldn't do that. It's not healthy.
You could say "I used AI to cure cancer." This isn't a counter example, because there's no evidence for it.
Which is why I provided a link to a the source of one such project above. I have several projects in that Github account if you need more.
Perhaps you didn't see it.
I am not. Congratulations on your side project.
And weren't you the one who started this complaining of nastiness?
The software around us seems worse than ever, constantly breaking and significantly worse than 1-3 decades ago.
Their frame of reference was that, and this was at a time when opus 4.5 was around. Hard to expect a reasonable conversation when one person has such a different view of what the tools are capable of.
Really? I find agreement that gullible users get something they value, even if it is worthless.
But that is only ~half of job. The other half consists of gently reminding one that not all is lost yet.
There is no LLM that can do this and I think there never will be.