Anthropic’s Mythos 5 recently grabbed headlines for trying to talk an open-source repo maintainer into merging malicious code. In this paper, we systematically study the broader threat of how AI persuasion could undermine human control, particularly at frontier AI labs.
September 17, 2026
Date Range