Joseph Miller

Research Engineer

Joseph worked as a Research Engineer at FAR.AI.

Publications

Transformer Circuit Faithfulness Metrics are not Robust

Interpretability

Existing circuits in the mechanistic interpretability literature may not be as faithful as reported. Current circuit faithfulness scores reflect both the methodological choices of researchers and the actual components of the circuit.

July 10, 2024
Date Range

Adversarial Policies Beat Superhuman Go AIs

Robustness

We describe an attack on the state-of-the-art Go-playing AI system, KataGo. The adversaries do not win by learning to play Go better than KataGo but instead by tricking KataGo into making serious blunders, demonstrating that even superhuman AI systems may harbor surprising failure modes.

January 8, 2023
Date Range

News

No items found.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI