London ControlConf 2025

Marylebone, London

March 27, 2025
March 28, 2025
Date Range

Overview

A collaboration between Redwood Research, FAR.AI, and the UK AI Security Institute, ControlConf brought together individuals interested in AI control (introduced here), including:

  • Researchers actively working on AI control at frontier labs, government departments, nonprofits, academia, and elsewhere
  • AI researchers who want to learn more about AI control but lack direct experience
  • People approaching AI control from non-AI backgrounds—such as information security professionals, policy researchers, and others who are interested in strategies to mitigate catastrophic misalignment risk

Our main objectives were:

  1. Coordination
    Many people are now exploring AI control, so we aimed to foster discussion on core strategic questions like “How ambitious should control strategies be?” and “What research is most pressing right now?”
  2. Knowledge Sharing
    As many individuals are new to the field, or have yet to engage with it directly, this conference aimed to help participants understand the state of the field.
  3. Interdisciplinary Discussions
    We explored AI control issues with experts from related areas, including AI policy and information security, to bring broader perspectives into the discussion.

Conference content featured a mix of presentations, demos, panel discussions, opportunities for one-on-one conversations, and structured breakout sessions.

London ControlConf 2025 sessions

Control, Cooperation and AI Welfare

Kathleen Finlinson

March 28, 2025

2025

Subversion Strategy Evaluation

Charlie Griffin

March 28, 2025

2025

[Fireside Chat] Control & Computer Security

Steve Kelly

March 28, 2025

2025

AI Control Safety Cases

Tomek Korbak

March 28, 2025

2025

Frontier Models are Capable of In-context Scheming

Alexander Meinke

March 27, 2025

2025

Task Decomposition for AI Control

Aaron Sandoval

March 27, 2025

2025

Deluding AIs

Owain Evans

March 27, 2025

2025

High Integrity Research Practices at Palisade

Dmitrii Volkov

March 27, 2025

2025

Lakera: Runtime Security for Agents at Scale

Sam Watts

March 27, 2025

2025

Improved Monitoring of Backdoor Insertion During Code Refactoring

Trevor Lohrbeer

March 27, 2025

2025

Low-stakes Control

Vivek Hebbar

March 27, 2025

2025

[Fireside Chat] White-box Methods for AI Control

Neel Nanda

March 27, 2025

2025

Hopes & Difficulties with Using Control Protocols in Production

Fabien Roger

March 27, 2025

2025

STPA & AI Control

Simon Mylius

March 27, 2025

2025

Ctrl-Z: Controlling AI Agents via Resampling

Aryan Bhatt

March 26, 2025

2025

Optimization around Control: Lessons from MONA

Sebastian Farquhar

March 26, 2025

2025

Automated Researchers Can Subtly Sandbag

Johannes Gasteiger

March 26, 2025

2025

Threat Analysis to Identify Priorities

Francesca Gomez

March 26, 2025

2025

Hierarchical Monitoring

Tim Hua

March 26, 2025

2025