I wonder what models would do when there were an "impostor" among them, an agent with no alignment or with a different kind of alignment behavior
Just want to correct the premise that the agent swarm behaviour was emergent.
Noam Brown on Dwarkesh podcast around 40 minutes mark
> “We train them to work together, to be cooperative, to essentially be fully aligned with each other.”
However, the article experiments with five agents. So it seems to assume even the small scale behaviour is emergent.
Also, “roles and structures” may well have been learned during training.