San Francisco Alignment Workshop

San Francisco, CA

February 27, 2023
February 28, 2023
Date Range

Overview

The inaugural Alignment Workshop was held 27–28 February 2023 in San Francisco, co-organized by researchers from top industry AI labs and universities, and was attended by 80 of the world’s leading machine learning researchers.

The Alignment Workshop series brings together top machine learning researchers and practitioners from industry, academia, and government. The workshop focuses on discussing and debating critical topics related to AI alignment, enabling participants to better understand potential risks from advanced AI, and strategies for solving them. Key issues discussed include model evaluations, interpretability, robustness, and AI governance.

San Francisco Alignment Workshop sessions

Lightning Talks (Day 2)

Multiple Speakers

February 27, 2023

2023

“Situational Awareness” Makes Measuring Safety Tricky

Ajeya Cotra

February 26, 2023

2023

Surveying Safety Research Directions

Dan Hendrycks

February 26, 2023

2023

Supervising AI on Hard Tasks

Jan Leike

February 26, 2023

2023

Opening Remarks: Confronting the Possibility of AGI

Ilya Sutskever

February 26, 2023

2023

Looking Inside Neural Networks with Mechanistic Interpretability

Chris Olah

February 26, 2023

2023

Lightning Talks (Day 1)

Multiple Speakers

February 26, 2023

2023

How Misalignment Could Lead to Takeover

Paul Christiano

February 26, 2023

2023

Aligning Massive Models: Current and Future Challenges

Jacob Steinhardt

February 26, 2023

2023