Singapore Alignment Workshop

Singapore

•

April 23, 2025
Date Range

Program Committee

Diyi Yang

Assistant Professor

Stanford University

Ponnurangam Kumaraguru "PK"

Professor

IIIT Hyderabad

Dylan Hadfield-Menell

Associate Professor

MIT

Adam Gleave

Co-founder & CEO

FAR.AI

Overview

Held around ICLR 2025 in Singapore, the Singapore Alignment Workshop brought researchers and leaders from around the world together to debate and discuss current issues in AI safety.

The Alignment Workshop series brings together top machine learning researchers and practitioners from industry, academia, and government. The workshop focuses on discussing and debating critical topics related to AI alignment, enabling participants to better understand potential risks from advanced AI, and strategies for solving them. Key issues discussed include model evaluations, interpretability, robustness, and AI governance.

Singapore Alignment Workshop sessions

Scaling Alignment Research via Safety Cases

Jacob Pfau

•

April 23, 2025

2025

Emergent Misalignment

Owain Evans

•

April 23, 2025

2025

Jailbreaking Aligned LLMs, Reasoning Models & Agents

Siva Reddy

•

April 23, 2025

2025

Antidistillation Sampling

Zico Kolter

•

April 23, 2025

2025

Unthinking Vulnerability of Large Reasoning Models

Baoyuan Wu

•

April 23, 2025

2025

Computational Safety for Generative AI

Pin-Yu Chen

•

April 23, 2025

2025

LLM Safety Training & Semantically Related Natural Prompts

Sravanti Addepalli

•

April 23, 2025

2025

Your DPO Algorithm is Secretly a Misspecified Reward Estimator

Aditya Gopalan

•

April 23, 2025

2025

CVE-Bench: A Real-World Cybersecurity Benchmark for AI Agents

Daniel Kang

•

April 23, 2025

2025

Next Steps for Control Safety Cases

Martín Soto

•

April 23, 2025

2025

High-Compute Alignment & Control

Noam Brown

•

April 23, 2025

2025

A New Definition & Improved Mitigation for Reward Hacking

Cassidy Laidlaw

•

April 23, 2025

2025

Value Compass Leaderboard: Platform for LLMs’ Value Evaluation

Xiaoyuan Yi

•

April 22, 2025

2025

STAIR: Improving Safety Alignment with Introspective Reasoning

Yinpeng Dong

•

April 22, 2025

2025

Unified Explanation of DNN Inference Logic & Representation

Huiqi Deng

•

April 22, 2025

2025

Persuade AIs

Weiyan Shi

•

April 22, 2025

2025

Evaluating Alignment Processes Rather than Models

Adam Kalai

•

April 22, 2025

2025

Disaster Preparedness for AI Safety

Tegan Maharaj

•

April 22, 2025

2025

You Should Work on Agent Infrastructure

Alan Chan

•

April 22, 2025

2025

Governing AI Agents Under the EU AI Act

Robin Staes-Polet

•

April 22, 2025

2025

The White House AI Action Plan

Mark Brakel

•

April 22, 2025

2025

In-House Evaluation is Not Enough

Shayne Longpre

•

April 22, 2025

2025

Deceptive Alignment & Thinking Monitor in LLMs

Jiaming Ji

•

April 22, 2025

2025

Safety Benchmarking & Testing of Multimodal LLMs

Tianwei Zhang

•

April 22, 2025

2025

Alignment on the Fly

Furong Huang

•

April 22, 2025

2025

Safety Alignment of LLMs

Animesh Mukherjee

•

April 22, 2025

2025

Exploring Cooperation & Alignment

Kalesha Bullard

•

April 22, 2025

2025

AI Catastrophic Risks & Scientist AI Solution

Yoshua Bengio

•

April 22, 2025

2025

Powering up AI Capability Evaluations with Model Tampering Attacks

Stephen Casper

•

April 22, 2025

2025

Thanks to our sponsors

We are grateful for financial support and sponsorship from: