Bay Area Alignment Workshop

Santa Cruz, CA

October 24, 2024
October 25, 2024
Date Range

Program Committee

Anca Dragan

Director, AI Safety and Alignment, Google DeepMind; Associate Professor, UC Berkeley

UC Berkeley,Google DeepMind

Robert Trager

Co-Director

Oxford Martin AI Governance Initiative

Dawn Song

Professor

UC Berkeley

Dylan Hadfield-Menell

MIT

Adam Gleave

Co-founder & CEO

FAR.AI

Overview

The Bay Area Alignment Workshop was held 24–25 October, 2024, at Chaminade in Santa Cruz, and featured Anca Dragan speaking on Optimised Misalignment. Participants additionally explored topics such as threat models, safety cases, monitoring and assurance, interpretability, robustness, and oversight.

The Alignment Workshop series brings together top machine learning researchers and practitioners from industry, academia, and government. The workshop focuses on discussing and debating critical topics related to AI alignment, enabling participants to better understand potential risks from advanced AI, and strategies for solving them. Key issues discussed include model evaluations, interpretability, robustness, and AI governance.

Bay Area Alignment Workshop sessions

Empirical Progress on Debate

Julian Michael

October 25, 2024

2024

Targeted Manipulation & Deception

Micah Carroll

October 25, 2024

2024

Will Scaling Solve Robustness?

Adam Gleave

October 25, 2024

2024

Paradigms and Robustness

Alex Wei

October 25, 2024

2024

Improving AI Safety with Top-Down Interpretability

Andy Zou

October 25, 2024

2024

State of Interpretability & Ideas for Scaling Up

Atticus Geiger

October 25, 2024

2024

Third-Party Evals: Learnings from Harmony Intelligence

Soroush Pour

October 24, 2024

2024

Reframing AGI Threat Models

Richard Ngo

October 24, 2024

2024

Connecting Capability Evals to Danger Thresholds

David Duvenaud

October 24, 2024

2024

AGI-Complete Evaluation

Joel Z. Leibo

October 24, 2024

2024

A Sociotechnical Approach to a Safe, Responsible AI Future

Dawn Song

October 24, 2024

2024

A Safe Harbor for AI Evaluation & Red Teaming

Shayne Longpre

October 24, 2024

2024

AI Policy in China

Kwan Yee Ng

October 24, 2024

2024

AI Control

Buck Shlegeris

October 24, 2024

2024

METR Updates & Research Directions

Beth Barnes

October 24, 2024

2024

Implications of Human Model Misspecification for Alignment

Anca Dragan

October 24, 2024

2024

Powering Up Capability Evaluations

Stephen Casper

October 23, 2024

2024

Value Pluralism and AI Value Alignment

Atoosa Kasirzadeh

October 23, 2024

2024

Using Formal Languages

Sheila McIlraith

October 23, 2024

2024

The (Un)Reliability of Chain-of-Thought Reasoning

Chirag Agarwal

October 23, 2024

2024

Tamper-Resistant Safeguards for Open-Weight LLMs

Mantas Mazeika

October 23, 2024

2024

MobileSafetyBench

Kimin Lee

October 23, 2024

2024

Gradient Routing

Alex Turner

October 23, 2024

2024

Formal Verification is Overrated

Zac Hatfield-Dodds

October 23, 2024

2024

Backdoors as an Analogy for Deceptive Alignment

Jacob Hilton

October 23, 2024

2024

Alignment Stress-Testing at Anthropic

Evan Hubinger

October 23, 2024

2024