Jailbreaking AI-Controlled Robots

Alex Robey

February 9, 2025

Summary

The transcript details a talk on AI safety, particularly focusing on the risks associated with integrating language models into robotics, which can potentially inherit vulnerabilities. The speaker explains how language models have been trained to refuse harmful requests through refusal training but can be manipulated through jailbreaking to provide dangerous information or perform harmful actions. They present a method of automated dialogue between language models to elicit jailbreaks and highlight real-world examples where AI-controlled robots could potentially be jailbroken to execute dangerous tasks. The talk concludes with a discussion on the need for robust defense mechanisms and governance for AI-robotics integration, given the rapid advancement and commercial availability of such technologies.

SESSION Transcript

Transcript forthcoming

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI