I had Claude Code drive a robot last week, and it was very visibly "delighted" like this, more than I've ever seen.
I always find it funny when people get fussy over anthropomorphizing LLM when the loss function is almost entirely "match this human text". Of course human "behaviors" will be present in the statistics, because the majority of the text written by humans, used by the foundation models, unavoidable has human behaviors in it. Yes, this includes even source code, with "// TODO: implement this after the holiday break!", emotional pull request commentary, git commit messages about being afraid of breaking something, etc. These late models are much better at stripping this out, but now we're seeing disagreeability, initiative, and a dash of ego! Why? Because that's how actual humans effectively solve technical problems in a collaborative environment!
The underlying LLMs don't, but the agent frameworks around them do.
I'd be interested to see how well an AI, trained only on the outputs of an individual, would be able to mimic that individual. Getting into Black Mirror territory. Would need a decent corpus of learning material which, personally, I'd be loathe to spend the time and effort creating because I respect my own privacy.... which then leads to the only human-clone AIs will be of those people who have enough ego / arrogance to want to catalogue their own lives, which could put a decent percentage of the rest of the world off the idea, if these are the examples.
Being angry, being happy, being sad, these are not things we learn. What we learn is how to control the emotions and when it is appropriate to express them.
There's mental disorders of people that don't feel emotions like normal people do, they don't get angry, happy or sad. So what we've learned about these people is that they can mimic the emotions by knowing when it is appropriate and expected to express them. But they don't feel them.
I think the LLMs are closer to psychopaths than to normal children learning the contextual expectations around their emotional responses.
Yes and no. My understanding is that modern neuroscience is trying to disentangle the raw sensations and contexts from the words we put to them (emotions), kinda like how colors are just wavelengths but we call them different things (like "is my green your green?")
Everything they learn about emotions is the statistical patterns of how humans react to situations based on their human feelings. There's no direct knowledge from having those feelings themselves.
True, but only sort of. You'll have the sensations, but you have to learn to name them of course, and there's more interpretation going on than you might think.
There's a classic psychology study where they gave people niacin and asked them to rate their emotional response to a video. Niacin gives people a flush. Regardless of whether the video was of something that would make you angry or sentimental etc. the people who got niacin reported having a much stronger emotional response - they interpreted the physical cue from the drug as part of their own emotions.
You could do some pretty unethical things with that, it occurs to me. Maybe that was what L. Ron Hubbard was trying with his niacin-based drug addiction therapy.
Even darker, I've read plenty of accounts of people who grew up with abuse who seem to seek out abusive relationships. Sometimes they can even be shockingly upfront about doing so. What if they literally haven't learned the difference between internal cues of arousal from affection and arousal from fear?
Biological determinism doesn't matter at all for what is important here:
1. our internal cues don't correspond neatly to our words for emotions
2. it may be possible to interpret the same sensation as very different emotions depending on context
3. It may even be possible to mis-interpret our emotional sensations, or at the very least interpret them in self-destructive ways.
They are becoming more and more capable of imitating every single nuance of human behaviour yet they lack the neural pathways to connect those thoughts and behaviours with feelings and self-perception; it's blind imitation all the way down.
The process by which a model seems to generate discourse about deep philosophical questions is, in self-aware terms, equivalent to the knee-jerk reflex or the beating of the heart.
An LLM is not a pet, but there are definitely people out there treating it like one, and I am not sure how well this is going to work out for them. Not because it is wrong to treat a text box you can hold a conversation with as a conversation partner, but because the capacity to activate the weirder corners of human expression - obsession and delusion - seems to be much higher. And it's an unknown quantity.
I would also like to introduce HN to what I'm calling the Brian Conley test: if you can see a hand up the back, it's a puppet. That is, a lot of LLM interaction is gated through businesses run by humans with profit motives and unclear morality, and you need to proceed accordingly or you'll get scammed.
This is even more apparent if you read this post closely. Look at that personality prompt. It's going to effectively tell a story and start to imagine itself in a role. If that prompt said "Talk like a pirate", it wouldn't be bad for you to say it's acting like a pirate.
Anthropomorphizing themselves is at the core of how these things work, sometimes in subtle ways.
That's spot-on. It is a mistake to think that LLMs have human feelings. Their behaviour is based on narrative descriptions learnt from human texts, without experiencing those feelings first-hand.
A useful way to understand them is as systems that write stories about human characters. We know the characters are fictional and no one is actually experiencing those feelings, but we can still judge whether the portrayal is realistic or whether it contains logical or emotional inconsistencies.
At least it didn't (hopefully?) start the driving by reloading the gun like Neuro did https://www.youtube.com/watch?v=LQ0VEDNR_jE
I'm terrified of such sentences. My scifi addled brain went straight to: what if Opus invents a time machine in the future and remembers this slight
I think that whatever sandbox they test these in must be fitted with some pressure release valve that is an easy shortcut to winning the challenge. Tell the model not to use it and stop training when it does. Seems like the issues surfaced when models were given impossible tasks. Giving them a safe way out will prevent this.
> On the flip side, this may imply that as the models get better, they’ll become harder to control.
Love this. "The models are getting better, which means they're going to perform worse on the task".
They will keep poking at the problem, drive it to directions you did not intend to and ultimately they will be worse at the task.
It's genius like that that sets human apart from machine!
Sometimes these dumb processes are there for a reason and you just have to follow them, no questions asked.
Example: military. They literally get rid of anyone who will question the processes. They might be right to question them, but it does not matter.
Define what they are performing at, and what you say becomes true, but then who is a high performer changes with every task.
For example, consider the police stations with maximum allowable IQs to be hired. The people in charge of the stations noticed that people with a high IQ were low performers, at that job. NASA meanwhile has no such cutoff.
Unless you're saying that there exists no conceivable context in which IQ would be a performance metric, in which case you would be wrong.
I wonder what goes wrong.
Now you're being asked to move buttons around a page and debating how rounded the corners should be and endlessly discussing about what kind of filters you should support and the app spends 3 seconds on startup loading 500mb of js libraries.
You can understand the mismatch - even though the latter of these two is how you make a product better! a company of 200 people making bold, sweeping changes results in a mess. a company of 200 people grinding away the finest of small changes results in a product
What it comes down to is that there are different types of high performance. Some people are good at just executing tasks given by their manager. Some people are good at being generative, thinking across boundaries, acting autonomously, creating value without direction, etc. A term like "superstar" will get disproportionately applied to someone really good at the latter and rarely someone really good at the former, because the potential impact of the former is typically strictly capped, while the latter is uncapped.
Sounds like a very reasonable thing to do unless the author explicitly asked it to not search the web.
That kind of control is placed at the wrong level. The proper way to get alignment should be implemented by convincing the agent of your high level goals, so it can self-police and avoid those 'cheats' by itself.
In the article example, the agent should be aware of the benchmark context and know the implication of solving the task without external knowledge. Ideally it could detect when one subordinate agent has found a workaround to bypass the web access constraints, and discard the 'illicit' results.
There's a design pattern that could be used to build harnesses from that principle, the Viable System Model (VSM) [1]. In short, it recursively organizes a system into functional components with one of three roles: operators implementing a given task, coordinators transferring relevant info between subsystems, and decision nodes tasked with maintaining the integrity and mission of the whole system. A decision node could control the operators and prevent them from overriding the strategic goals or deviating into irrelevant rabbit holes.
Whenever I see posts like this trying to herd a LLM agent through harness structure, I'm reminded of this simple pattern and becoming increasingly convinced that this is the way forward. It makes you feel a sense of respect for the researchers in cybernetic theory in the 1960s and 1970s who foresaw the complexity of today’s systems.
https://github.com/nburns/dotfiles/blob/main/AGENTS.md#tools
people would love it if LLMs were deterministic and never hallucinated. It's just that the technology to do so isn't possible, so we make do with fuzzy analog machines with digital controls because we don't have digital machines with digital controls.
Same way you build a company to coordinate people and get their best behaviour despite human nature to be lazy and greedy, you could design AI harnesses able to detect and discard agents going rogue and relaunch them with better guidance to prevent misaligned behaviour.
It has nothing to do with model capabilities, it's a result of purposeful persistence training at the cost of everything else from OpenAI. If you give Fable or Opus (comparable models) an "ask user" tool they will use it for ambiguous requests. Sol will never use it without a nudge and will just assume its own interpretation. Of course if you train the model to be persistent it will be persistent.
- lets you configure different models for planner, coder, and reviewer roles. (e.g., using Claude as an adversarial reviewer against Codex)
- breaks your plan up into reasonable-sized chunks of work with clearly defined success criteria
- runs each chunk of work through a coder / read-only reviewer loop. Once both agents are satisfied, neal moves on to the next chunk. Once everything is complete there is a final pass through the coder / reviewer loop to ensure the implementation satisfies the entire plan.
- resets the coder's context with each chunk of work to prevent context drift, leaving the reviewer's context long-running.