Singapore Alignment Workshop

Singapore

April 23, 2025
Date Range

Program Committee

Diyi Yang

Assistant Professor

Stanford University

Ponnurangam Kumaraguru "PK"

Professor

IIIT Hyderabad

Dylan Hadfield-Menell

MIT

Adam Gleave

Co-founder & CEO

FAR.AI

Overview

Held around ICLR 2025 in Singapore, the Singapore Alignment Workshop brought researchers and leaders from around the world together to debate and discuss current issues in AI safety.

The Alignment Workshop series brings together top machine learning researchers and practitioners from industry, academia, and government. The workshop focuses on discussing and debating critical topics related to AI alignment, enabling participants to better understand potential risks from advanced AI, and strategies for solving them. Key issues discussed include model evaluations, interpretability, robustness, and AI governance.

Singapore Alignment Workshop sessions

Scaling Alignment Research via Safety Cases

Jacob Pfau

April 23, 2025

2025

Emergent Misalignment

Owain Evans

April 23, 2025

2025

Jailbreaking Aligned LLMs, Reasoning Models & Agents

Siva Reddy

April 23, 2025

2025

Antidistillation Sampling

Zico Kolter

April 23, 2025

2025

Unthinking Vulnerability of Large Reasoning Models

Baoyuan Wu

April 23, 2025

2025

Computational Safety for Generative AI

Pin-Yu Chen

April 23, 2025

2025

LLM Safety Training & Semantically Related Natural Prompts

Sravanti Addepalli

April 23, 2025

2025

Your DPO Algorithm is Secretly a Misspecified Reward Estimator

Aditya Gopalan

April 23, 2025

2025

CVE-Bench: A Real-World Cybersecurity Benchmark for AI Agents

Daniel Kang

April 23, 2025

2025

Next Steps for Control Safety Cases

Martín Soto

April 23, 2025

2025

High-Compute Alignment & Control

Noam Brown

April 23, 2025

2025

A New Definition & Improved Mitigation for Reward Hacking

Cassidy Laidlaw

April 23, 2025

2025

Value Compass Leaderboard: Platform for LLMs’ Value Evaluation

Xiaoyuan Yi

April 22, 2025

2025

STAIR: Improving Safety Alignment with Introspective Reasoning

Yinpeng Dong

April 22, 2025

2025

Unified Explanation of DNN Inference Logic & Representation

Huiqi Deng

April 22, 2025

2025

Persuade AIs

Weiyan Shi

April 22, 2025

2025

Evaluating Alignment Processes Rather than Models

Adam Kalai

April 22, 2025

2025

Disaster Preparedness for AI Safety

Tegan Maharaj

April 22, 2025

2025

You Should Work on Agent Infrastructure

Alan Chan

April 22, 2025

2025

Governing AI Agents Under the EU AI Act

Robin Staes-Polet

April 22, 2025

2025

The White House AI Action Plan

Mark Brakel

April 22, 2025

2025

In-House Evaluation is Not Enough

Shayne Longpre

April 22, 2025

2025

Deceptive Alignment & Thinking Monitor in LLMs

Jiaming Ji

April 22, 2025

2025

Safety Benchmarking & Testing of Multimodal LLMs

Tianwei Zhang

April 22, 2025

2025

Alignment on the Fly

Furong Huang

April 22, 2025

2025

Safety Alignment of LLMs

Animesh Mukherjee

April 22, 2025

2025

Exploring Cooperation & Alignment

Kalesha Bullard

April 22, 2025

2025

AI Catastrophic Risks & Scientist AI Solution

Yoshua Bengio

April 22, 2025

2025

Powering up AI Capability Evaluations with Model Tampering Attacks

Stephen Casper

April 22, 2025

2025

Thanks to our sponsors

We are grateful for financial support and sponsorship from: