Results 1
Showing items 1 to3 of 6

Why Language Models Hallucinate
Santosh Vempala
•
December 1, 2025
2025

What Does It Mean for Agentic AI to Preserve Privacy?
Niloofar Mireshghallah
•
December 1, 2025
2025

Toy Models for Task-Horizon Scaling
Cozmin Ududec
•
December 1, 2025
2025

The State and the Science of AI-Bio Evals
Bryce Cai
•
December 1, 2025
2025

State of Jailbreaks
Xander Davies
•
December 1, 2025
2025

Privacy and Security Challenges in AI Agents
Kamalika Chaudhuri
•
December 1, 2025
2025

Peril and Potentials of Training with Lie Detectors
Chris Cundy
•
December 1, 2025
2025

Our Pivot To Pragmatic Interpretability
Neel Nanda
•
December 1, 2025
2025

Measuring AI Systems’ Ability to Influence Humans
Anna Gausen
•
December 1, 2025
2025

How do we know what AI can (and can't) do?
Anka Reuel
•
December 1, 2025
2025

Hidden Pitfalls of AI Scientist Agents
Atoosa Kasirzadeh
•
December 1, 2025
2025

Guarding the Age of Agents
2025

Current State of AI Agent Security
Andy Zou
•
December 1, 2025
2025

AI + Democracy
Divya Siddarth
•
December 1, 2025
2025

Toward Scalable and Actionable Interpretability
Yonatan Belinkov
•
November 30, 2025
2025

Scalable Oversight and Understanding
Sarah Schwettmann
•
November 30, 2025
2025

San Diego Alignment Workshop 2025 Opening Remarks
Adam Gleave
•
November 30, 2025
2025

Provably Safe AI
Max Tegmark
•
November 30, 2025
2025

Powerful Open-Weight AI Models: Wonderful, Terrible & Inevitable
Stephen Casper
•
November 30, 2025
2025

Polarity-Aware Probing for Quantifying Latent Alignment in LMs
Chirag Agarwal
•
November 30, 2025
2025

Multi-agent RL for Provably Robust LLM Safety
Natasha Jaques
•
November 30, 2025
2025

Lessons Learned from the First Misalignment Safety Case
Samuel Bowman
•
November 30, 2025
2025

Frontier AI in Cybersecurity: Risks, Challenges & Future Directions
Dawn Song
•
November 30, 2025
2025

Finding & Reactivating Safety Mechanisms of Post-Trained LLMs
Yisen Wang
•
November 30, 2025
2025

Eval Awareness is Becoming a Problem
Marius Hobbhahn
•
November 30, 2025
2025

Disentangling Agency & Predictive Power Without Solving ELK
Yoshua Bengio
•
November 30, 2025
2025

Consensus Sampling for Safer Generative AI
Adam Kalai
•
November 30, 2025
2025

Chain of Thought Monitorability for AI Safety
Tomek Korbak
•
November 30, 2025
2025

Automating Mechanistic Interpretability
Chenhao Tan
•
November 30, 2025
2025

AI Control Needs Redteaming
Asa Cooper-Stickland
•
November 30, 2025
2025

AI Agent Benchmarks Are Broken
Daniel Kang
•
November 30, 2025
2025

A Practical Approach to Verifying Code at Scale
Maja Trębacz
•
November 30, 2025
2025

Low probability estimation
Jacob Hilton
•
October 28, 2025
2025

Can you just train models not to scheme?
Marius Hobbhahn
•
October 7, 2025
2025

Singular Learning Theory and AI Safety
Jesse Hoogland
•
September 16, 2025
2025

What Would it Take to Stop the Development of Superintelligence? A Treaty Proposal
Aaron Scher
•
August 26, 2025
2025

The EU Code of Practice: Towards a Global Standard for Frontier AI Risk Management
Siméon Campos
•
August 19, 2025
2025

Alignment is social: lessons from human alignment for AI
Gillian Hadfield
•
August 5, 2025
2025

Realigning AI
Zhijing Jin
•
June 17, 2025
2025

Verifiable Compute & Building Trust in Agentic Supply Chains
Tina Morrison
•
May 31, 2025
2025

Sandpaper Socks: Operational Considerations in AI Evals
Olivia Shoemaker
•
May 31, 2025
2025

Rising Cost of Evaluations, Falling Cost of Intelligence
2025

Regulation of Frontier Models, Confidential Computing & AI Alignment
2025

Prioritizing Technical Governance Investments: Verification
Robert Trager
•
May 31, 2025
2025

Policy-Oriented AI Evaluations
2025

Normal Policy Tools for a Normal Technology
Asad Ramzanali
•
May 31, 2025
2025

National Security & Defense
2025

Hydropower: the Missing Piece
Charles Yang
•
May 31, 2025
2025

Hardware-Enabled Verifiability as a Data Center Security Property
2025

DoD & AGI Preparedness
2025

Day 2 Opening Remarks
Lennart Heim
•
May 31, 2025
2025

Compute in America: Playbook for Secure Clusters at Home
2025

Application of Cryptographic Primitives in AI Accelerators
Fatemeh Ganji
•
May 31, 2025
2025

Access Considerations for Generative AI Systems Beyond Release
Irene Solaiman
•
May 31, 2025
2025

AI-Military Integration Under President Trump
Hamza Chaudhry
•
May 31, 2025
2025

AI Supply Chains: An Emerging Ecosystem of AI Dependencies
2025

AI National Security Policy: Industry and Government
Sara McNaughton
•
May 31, 2025
2025

14 Ways of Looking at AI
2025

Unresolved Debates on the Future of AI
2025

Security: Datacenter & Model Weight
Lisa Thiergart
•
May 30, 2025
2025

Making Sense of the AI Auditing Ecosystem
Miranda Bogen
•
May 30, 2025
2025

Frontier Model Safety
Rep. Alex Bores
•
May 30, 2025
2025

Day 1 Opening Remarks
2025

An Overview of Technical AI Governance
Ben Bucknall
•
May 30, 2025
2025

AI Control: Addressing Risks from Agentic Internal Deployments
2025

Scaling Alignment Research via Safety Cases
Jacob Pfau
•
April 23, 2025
2025

Emergent Misalignment
Owain Evans
•
April 23, 2025
2025

Jailbreaking Aligned LLMs, Reasoning Models & Agents
Siva Reddy
•
April 23, 2025
2025

Antidistillation Sampling
Zico Kolter
•
April 23, 2025
2025

Unthinking Vulnerability of Large Reasoning Models
Baoyuan Wu
•
April 23, 2025
2025

Computational Safety for Generative AI
Pin-Yu Chen
•
April 23, 2025
2025

LLM Safety Training & Semantically Related Natural Prompts
Sravanti Addepalli
•
April 23, 2025
2025

Your DPO Algorithm is Secretly a Misspecified Reward Estimator
Aditya Gopalan
•
April 23, 2025
2025

CVE-Bench: A Real-World Cybersecurity Benchmark for AI Agents
Daniel Kang
•
April 23, 2025
2025

Next Steps for Control Safety Cases
Martín Soto
•
April 23, 2025
2025

High-Compute Alignment & Control
Noam Brown
•
April 23, 2025
2025

A New Definition & Improved Mitigation for Reward Hacking
Cassidy Laidlaw
•
April 23, 2025
2025

Value Compass Leaderboard: Platform for LLMs’ Value Evaluation
Xiaoyuan Yi
•
April 22, 2025
2025

STAIR: Improving Safety Alignment with Introspective Reasoning
Yinpeng Dong
•
April 22, 2025
2025

Unified Explanation of DNN Inference Logic & Representation
Huiqi Deng
•
April 22, 2025
2025

Persuade AIs
Weiyan Shi
•
April 22, 2025
2025

Evaluating Alignment Processes Rather than Models
Adam Kalai
•
April 22, 2025
2025

Disaster Preparedness for AI Safety
Tegan Maharaj
•
April 22, 2025
2025

You Should Work on Agent Infrastructure
2025

Governing AI Agents Under the EU AI Act
Robin Staes-Polet
•
April 22, 2025
2025

The White House AI Action Plan
Mark Brakel
•
April 22, 2025
2025

In-House Evaluation is Not Enough
Shayne Longpre
•
April 22, 2025
2025

Deceptive Alignment & Thinking Monitor in LLMs
Jiaming Ji
•
April 22, 2025
2025

Safety Benchmarking & Testing of Multimodal LLMs
Tianwei Zhang
•
April 22, 2025
2025

Alignment on the Fly
Furong Huang
•
April 22, 2025
2025

Safety Alignment of LLMs
Animesh Mukherjee
•
April 22, 2025
2025

Exploring Cooperation & Alignment
Kalesha Bullard
•
April 22, 2025
2025

AI Catastrophic Risks & Scientist AI Solution
Yoshua Bengio
•
April 22, 2025
2025

Powering up AI Capability Evaluations with Model Tampering Attacks
Stephen Casper
•
April 22, 2025
2025

Control, Cooperation and AI Welfare
Kathleen Finlinson
•
March 28, 2025
2025

Subversion Strategy Evaluation
Charlie Griffin
•
March 28, 2025
2025

[Fireside Chat] Control & Computer Security
Steve Kelly
•
March 28, 2025
2025

AI Control Safety Cases
Tomek Korbak
•
March 28, 2025
2025

Frontier Models are Capable of In-context Scheming
Alexander Meinke
•
March 27, 2025
2025
There are no available cars matching the current filters.