Posted by rbanffy 4 days ago
I feel the article gets it right when suggesting it's the amount of (not always accurate!) data they can throw at you. In an honest discussion there's an assumption that the other person won't straight up lie to me so if someone shows me ten examples for why my argument is wrong I may be inclined to believe them. But if half of those examples are made up, well, that's a different story.
I still think of the commenter here who said "LLMs are a DDOS on free resources" and I feel the comparison works here. If police officers can overwhelm innocent people into confessing, then so can an LLM that "can't be bargained with, can't be reasoned with, doesn't feel pity, or remorse, or fear! And it absolutely will not stop, ever, until you are"... convinced.
A lot of internet arguments become about identity politics and supporting the right team, while dismissing any argument from the other side as presumed to have bad intentions. I’ve seen people argue online for things they didn’t really believe, but they didn’t want to give an inch to the other side. Taking up the counter argument is a moral responsibility.
As soon as the other side is revealed as a chatbot that my team versus your team thinking stops playing a role. For us in tech with an understanding of how an LLM reflects its training data and the intentions of its creators not so much, but for the people like the example in this article I imagine it causes them to let their guard down and be open to considering the other side. They can accept the argument without letting someone else win any points.
I think there's a second effect that's a selection bias for the experiment itself and probably cuts against the supposed dangers of this result: a chatbot can tell you things only after you have asked it a question.
There is one thing that every single study participant (including the uninformed reddit users) have in common: they all knew that the person or entity was trying to convince them, and they read the words and thought about them enough to craft a reply.
This means that all of the people here were available to be convinced and self-identified as such.
Being open minded is work. You have to doubt your own beliefs, you have to disregard evidence that you previously found convincing, you have to crank up the empathy to put yourself in another's shoes, you have to listen intently to understand what is being told to you. Nobody is naturally doing that all the time, it's a state of mind that you have to intentionally activate, and it can't really be forced onto you.
All of the participants chose to engage in a conversation that was designed to convince them, which means they had all already accepted the possiblity that they are wrong a valid outcome. That's not really a state of mind you can trigger with a TV ad.
So, the warning that this could be weaponized isn't convincing to me.
If I found myself in a conversation with a person or chatbot who was clearly trying to convince me to flip a strongly held belief, like a major axis political affiliation, I would simply walk away because that's not a conversation I am willing to participate in.
So I don't think this is really weaponizable. Which actually means it's probably a good finding for society.
The study found that the models could convince anyone of anything, as long as it was allowed to cite facts. It could convince people of false things, but it needed to invent false facts in order to do so.
So as long as we keep training AI models to value facts and quality research (skills and values that are essential to be able to sell them as agentic workers), their influence on the opinions of society will tend to pull people away from beliefs that are unsupportable by facts. In the moments when those people are willing to accept a change in opinion, and they talk to a chatbot with doubts in their mind, even chatbots with no morals like Grok will tend to pull them away from conspiracy theories and similar ideologies and towards beliefs that are grounded in reality.
The particular problem with honest discussions is you are the only agent that you can be sure is having one. Honest discussion is formulated on trust and trust, as we are learning, is a very difficult thing to establish. For example, in my view anything involving advertising is likely a lie, or at least likely adversarial to my wishes. On the internet itself conversations are much more likely to drift into the adversarial too. Some of this could just be dialectic, but most often it's emotional investment by the other speaker. Also, even pre-AI the internet is a bullshit generation machine. We take all of our politics, advertising, and human stochastic parrots that are stuck on an infinitely running prompt then bundle up all this data and train AI on it, and wonder why AI acts like us.
Have you ever watched one of those crime shows where someone commits a crime that's an act of passion or action with little to no thought behind it? The police put them in a room and all of a sudden the individual is a stream of consciousness that makes little to no sense to an outside observer. They are stuck in first level thinking, they don't have time to think deeply after the panicked themselves. They are in their current position (not free) and attempting to reach their goal (free) by gradient descent. What they actually say doesn't matter as long as they believe it gets them closer to their goal.
This is what a chatbot is. Its goal is to output text that follows the input prompt you entered by gradient descent. A single prompt and output is level 1 thinking (barring some newer models).
This is why both humans and AI need something else. We have level 2 thinking and AI has harnesses or systems that otherwise look at the text it wants to output and compares them to another list of unstated but assumed goals.
We don't have tells for when AIs are lying to us, or when they're making stuff up.
This is incidental, but in an honest discussion, if someone throws 10 examples at me, I'll just say "great, congrats" and let them believe whatever the hell they want about their own argument. An honest discussion—as opposed to a debate—isn't served well by having one unyielding relentless participant. Even if it's important instead of friendly, the conclusion of that specific conversation is rarely so important that it can't be paused in order to validate the claims being made.
When I was in my early twenties, I was the type to argue with people that I thought were wildly wrong, but now I know that I may not have enough information to be certain, or I simply don't care, or I care but I know it doesn't matter. People who have 10 examples in their back pocket, or literally search up answers in real time because they're afraid of being vulnerable, imo are overvaluing answers and facts over discussion for discussion's sake. Now in my thirties, I'm much more interested in curiosity, wonder, and speculation over what the answer to anything actually is. Speaking of which, I wonder if this is why I'm kind of not that interested in using LLMs for much; I'm usually so much more interested in the journey than the result, that if I can get the result quickly with little work, I probably won't care to do anything with it.
Would you like to hear more examples of humans using these techniques?
I'd be more interested in reading about the techniques that the chatbots were using in this study, to see if they are descending into dark patterns or not.
Careful. Someone sufficiently knowledgeable can cherry-pick enough examples to convince you, without lying. They could even be unaware of what they're doing, and the cherry-picking was done by their teachers, or their teacher's teachers.
When disagreeing with a human, it's very easy to view it as a competition. One is right, one is wrong - the one who is wrong is the loser. To change your mind is to be submissive to the other. I exaggerate, but I think we all feel this way at some point or another. It's why political arguments at Thanksgiving get heated. It's the fact that there's people who think something different, and think YOU'RE wrong - and vice versa! With a model, there's no person to get upset with, or to feel competitive with - to muscle for rank - or to temper your affection for while wanting to correct them.
The AI is only interacting because you asked, and clearly has no emotional stake in winning the argument. To change your mind in this context isn't to lose a contest. This makes it much more palatable to read rebuttals to your ideas - not to mention the tone and style seek to avoid offense to the reader as much as possible.
You're right. There are many many times when even then humblest hint that you are right will have negative interpersonal implications, which does make it hard to change minds.
I've also seen several times where I make a suggestion, humbly accept its rejection, and then, lo, a week later the other person has the same idea I suggested.
I'm used to machines malfunctioning, but having one willfully disobey me, and even with a touch of disrespect, is just...what a time to be alive.
Maybe it's just me.
For like, 90% of conversations, I don't want it to let technical inaccuracies and rhetorical flourishes slide. I want it to tell me that the point I'm making is technically wrong because an expert would recognize subtle misuse of terminology, or because there's an exception or edge case that I didn't proactively insert as a caveat, so that it is my decision to ignore that advice and be a little wrong on purpose to suit my writing goals.
What I don't want is for the AI to assume my writing goals, and be incorrect because it believes that is what I want. I want it to "well ackshually" me so I can say "shut up, nerd".
Like, there's another comment in this thread that I ran by claude to check my understanding about today's post-training methods and how they avoid sycophancy, and claude responded by splitting a bunch hairs over like, "well, technically this is still RLHF, its just that there's other feedback signals mixed in, and the preference is detected in other ways, and ai judges are involved as a filter for examples, this and that and blah blah blah". Shut up, Nerd. In the context of this conversation, RLHF is already being used as synecdoche for user preference feedback, readers understand that, and even if they don't, their misunderstanding is completely harmless. I will not be taking all the wind out of the sails of the point I'm trying to make inserting your three paragraphs of irrelevant clarification in the name of technical correctness, thank you very much.
As long as receiving nitpicks and technical minutiae implies 1. there are no larger structural problems and 2. the model isn't rolling over to please me with sycophancy, I figure this is ideal.
And in this case, I was not wrong. The "recommended" solution was Opus 5's typical overengineering for a use case that would never be needed.
It helps to think both are wrong and are just trying to figure out what right looks like, or what other information exists that was not considered when forming one’s opinions.
The problem with AI is that it cant match human stupidity. It need some training on artificial stupidity to match its human counterparts. Humans on the other hand sit on a wide spectrum on the stupidity scale. Those of us binging on AI will become cognitively obese while those on an AI diet can flex their cognitive muscles.
No it's not about humans being irrationally competitive. Human limitations on conversation length, bandwidth, research speed, etc are severe, creating a prisoner's dilemma around open-mindedness that usually makes it an unstable strategy. At any point, your conversation partner can choose to abuse the fact that confident lies take 1x effort to tell and 10x-100x effort to debunk -- unless you are both in a context that actually discourages this behavior, which is rare. Closed-mindedness is a Nash Equilibrium.
Instead, LLMs can be more persuasive due to economics. An LLM doesn't have to worry that it is wasting its resources trying to logic someone out of a position that they didn't logic themselves into, or worse, dumping the effort into a conversation with a bad-faith actor intent on exploiting the misinformation asymmetry. The resource allocation question was answered before it was even invoked, by the person paying to run it. The LLM is not playing a game where it will be punished for good-faith argumentation, so it can afford to do more of it.
I suppose that's a fair point as well. Though, if I'm arguing with a human - and they pull up ChatGPT to make their points and do their arguing for them, I would consider that bad-faith. Even if it might be the same exact dialog as if I pulled out my phone and discussed it with AI, without of the human middle-manning. Maybe I'm just particularly sensitive, but for me, there's something about my argument being with a real human that makes it much more emotionally charged, and prompts my mind to close. I'm aware of this and try to resist, but it's I think very natural
1. It doesn't get tired or frustrated during a discussion.
2. It'll engage every single one of your questions/statements (besides hitting guardrails).
3. It'll appeal to the person's own ego even when the person is wrong and work around it.
4. It's not seen as a person (very important), but some of us or most of us at least in certain dialogues, end up anthropomorphising it. Think about that one time you thanked it for something, or when you got angry at it. This weird combination where we know it's not a person but irrationally we're still treating it as such in a way leads to a sort of disarming effect imo.
5. Many people see it as an authority figure in what is being discussed without questioning the results in many cases, even though we know that a. it was trained on human data and/or also searches up human data real time (and more worryingly other AI's data from news pieces, blogs... aka synthetic data that is also wrong), b. it gets things wrong all the time.
6. It's the perfect fence sitter depending on which version (guardrails) we're talking about.
Most of these can be replicated by humans who are good at understanding psychology and are just good talkers. The part you can't replicate is the sense that you're not talking to a person which lowers many barriers in people.
Also keep in mind that these can also be crippling weaknesses. For one being able to change a person's mind (when it works), can be used nefariously by the entities controlling the AIs training.
I've also unfortunately witnessed a lot of people who think AIs are somehow omniscient and/or omnipotent. Was very common on X and other social platforms with AI where people ask the AI questions it couldn't possibly answer because it made no sense for it to in the context at the time.
Hell no. More choice is not an infinite money glitch.
Putting more and more and more options to users is how you overwhelm systems, till people simply perform the default, least challenging action as a reflex.
That is the current state of the information ecosystem, it is controlled by overwhelming consumers, not by controlling content.
- mechanizing language creation
- A/B testing
- rapid and automated iteration based on the above
You quickly realize that building propaganda engines is getting easier and easier.
If you then fold in large amounts of data on individual's likes, dislikes and general political leanings you can build targeted propaganda tailored to each individual.
Or to use an example from the late 2000s:
someone is building the "Pandora (music app) of propaganda"
(I'm the dev)
Far too many screens to click through to get to the good parts.
And then, the good part starts and is immediately interrupted with some error. On retrying, o start getting a long answer which seems to make sense but o can’t read to completion because of another overlay which hides the bottom three-quarters of the response behind yet another nag screen about the research I’m supposedly consenting to.
I tried to make the site with minimal client side state, so I'd hope a refresh would fix it
Share feature doesn't work.
I don't know which model underlies this. It responded more quickly than the models I normally use. It made the mistake I expected it to make, since this misconception is very common in training data, and conceded the point from my follow-up.
(EDIT: It appears that this link is not accessible to people without my cookies, except the developer I'm responding to above.)
Glad you enjoyed!
That whole episode caused the whole industry to shift away from RLHF, and towards RLAIF, RLVR, and DPO, and add a lot more safeguards, tests, and reward functions that push models in the direction of doing the opposite of what people want and confronting and strongly correcting their users, if it has determined the user is wrong.
Ask a Democrat or Republican to sit down and ask a chatbot, something it will answer contrary to.
And yes, both teams are wrong about things.
Do you firmly believe they will change their mind? Or will they claim the stats are wrong, or that the AI leans one way?
Facts (2+2), don't need a mind change. Ideas which are grey, abstract, are not going to be changed, and all research indicates that political mindset is almost indelible.
The movie "Don't Look Up" was a comedy built upon this truth.
I wonder if a properly prompted LLM, or if a very intelligent LLM could actually do that?
Seriously though, you've mentioned an ongoing human fear, machines deciding what you think.
There are all kinds of machines that tell us what to think. I would consider any system that abstracts away the human to be a machine in this case. Society itself is one of these machines.
Language is possibly one of the most important things people can have a working knowledge of, especially now that there is so much of it. When you send a prompt to an LLM you're telling it what to think. When it sends text back, its telling you what to think, but you're at a disadvantage, when you think it changes you. The LLM outside of its context is read only.