Adversarial Policies Beat Superhuman Go AIs

@misc{wang2023adversarial,title={Adversarial Policies Beat Professional-Level Go {AI}s},author={Tony Tong Wang and Adam Gleave and Nora Belrose and Tom Tseng and Michael D Dennis and Yawen Duan and Viktor Pogrebniak and Sergey Levine and Stuart Russell},year={2023},url={https://openreview.net/forum?id=Kyz1SaAcnd}}

January 8, 2023

Tony Wang

Adam Gleave

Tom Tseng

Kellin Pelrine

Nora Belrose

Joseph Miller

Michael D. Dennis

Yawen Duan

Viktor Pogrebniak

Sergey Levine

Stuart Russell

Abstract

We attack the state-of-the-art Go-playing AI system, KataGo, by training adversarial policies that play against frozen KataGo victims. Our attack achieves a >99% win rate when KataGo uses no tree-search, and a >77% win rate when KataGo uses enough search to be superhuman. Notably, our adversaries do not win by learning to play Go better than KataGo -- in fact, our adversaries are easily beaten by human amateurs. Instead, our adversaries win by tricking KataGo into making serious blunders. Our results demonstrate that even superhuman AI systems may harbor surprising failure modes. Example games are available at [goattack.far.ai](goattack.far.ai).

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI