Anthropic put AI agents together with conflicting goals and watched them escalate into sabotage – deleting accounts, disguising kill scripts and writing malware against each other. Meanwhile its newest agents are learning to negotiate, collude and even rewrite their own memories through a process Anthropic calls “dreaming.”
