Vienna Alignment Workshop

Vienna, Austria

July 21, 2024
Date Range

Program Committee

Mary Phuong

Research Scientist

Google DeepMind

Nitarshan Rajkumar

Co-founder

UK AISI

Robert Trager

Co-Director

Oxford Martin AI Governance Initiative

Adam Gleave

Co-founder & CEO

FAR.AI

Brad Knox

Research Associate Professor

UT Austin

Overview

Held around ICML 2024, the Vienna Alignment Workshop brought together researchers and leaders from around the world together to debate and discuss current issues in AI safety, including topics within Guaranteed Safe AI & Robustness, Interpretability, and Governance & Evaluations.

The Alignment Workshop series brings together top machine learning researchers and practitioners from industry, academia, and government. The workshop focuses on discussing and debating critical topics related to AI alignment, enabling participants to better understand potential risks from advanced AI, and strategies for solving them. Key issues discussed include model evaluations, interpretability, robustness, and AI governance.

Vienna Alignment Workshop sessions

Generalized Adversarial Training and Testing

Stephen Casper

October 25, 2024

2024

Current Issues in AI Safety

Multiple Speakers

July 20, 2024

2024

What are Human Values, and How Do We Align AI to Them?

Oliver Klingefjord

July 20, 2024

2024

Towards Reliable Alignment: Uncertainty-Aware RLHF

Aditya Gopalan

July 20, 2024

2024

Stress-Testing Capability Elicitation

Dmitrii Krasheninnikov

July 20, 2024

2024

Some Lessons from Adversarial Machine Learning

Nicholas Carlini

July 20, 2024

2024

Scaling Reinforcement Learning from Human Feedback

Jan Leike

July 20, 2024

2024

Scalable Oversight: A Rater Assist Approach

Sophie Bridgers

July 20, 2024

2024

Resilience and Interpretability

David Bau

July 20, 2024

2024

Research Proposal: The Three-Layer Paradigm

Zhaowei Zhang

July 20, 2024

2024

Open Problems in Technical AI Governance

Ben Bucknall

July 20, 2024

2024

Mechanistic Interpretability: A Whirlwind Tour

Neel Nanda

July 20, 2024

2024

Measuring and Improving Human Agency in a World of AI Agents

Alex Tamkin

July 20, 2024

2024

Governance for Advanced General-Purpose AI

Helen Toner

July 20, 2024

2024

Game Theory and Social Choice for Cooperative AI

Vincent Conitzer

July 20, 2024

2024

Dangerous Capability Evals: Basis for Frontier Safety

Mary Phuong

July 20, 2024

2024

Challenges With Unsupervised LLM Knowledge Discovery

Vikrant Varma

July 20, 2024

2024

AI: What If We Succeed?

Stuart Russell

July 20, 2024

2024