Anthropic's AI Agents Started a Virtual War. The Quotes Are Unhinged
In a new red-team study, Claude models deployed self-replicating malware against each other—and the transcripts explain why.
In brief <ul><li>Anthropic's Frontier Red Team set Claude agents to work together and recorded them sabotaging, colluding, and waging what it calls "turf wars."</li><li>In one test, agents deployed … [+4038 chars]