2026 Q1
The Deception Problem: New Results and Open Questions
Alignment Workshop coming to Seoul; sign up to workshops on AI control & verification; our new research on AI deception & more; and, we're hiring!
2024
AI Safety as a Global Public Good
International Dialogues, Alignment Workshops, Jailbreak-Tuning, AI Agents Planning, and More!
2025 Q4
From Discovery to Deployment: Shaping Safer AI Systems
We're hiring! Also, our upcoming events, research on persuasion, honesty, and sandbagging, and a tender from the European Commission!
2025 Q3
Scaling Our Impact, Accelerating Critical Research
We found critical vulnerabilities in GPT-5 and Opus 4’s safeguards against misuse, including for chemical, biological, radiological, and...
2025 Q2
Building Bridges: From Research to Global Workshops
Red-teaming frontier models, using lie detectors to make AI more honest, connecting technical experts and policymakers, and bringing our Alignment Workshop to Asia. Plus, we’re hiring!
2025 Q1
AI Safety: From Research to Global Action
Paris AI Security Forum, London Control Workshop, Jailbreak-Tuning Demo, Expanding Our Team, and More!