Joar Max Viktor Skalse

University of Oxford

Publications

Open Problems in Mechanistic Interpretability

Interpretability

This review discusses the current frontier of mechanistic interpretability, which aims to understand the computational mechanisms underlying neural networks. While the field has made progress, many open problems remain, including the need for improved methods, better applications to specific goals, and engagement with socio-technical challenges.

January 26, 2025
Date Range

Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems

Alignment

This paper introduces Guaranteed Safe (GS) AI, an approach to AI safety that ensures high-assurance quantitative safety guarantees. It relies on three core components—a world model, a safety specification, and a verifier—to mathematically verify that AI systems meet safety requirements.

May 9, 2024
Date Range

STARC: A General Framework For Quantifying Differences Between Reward Functions

Alignment

STARC (STAndardised Reward Comparison) metrics, a class of pseudometrics, quantify differences between reward functions, providing theoretical and empirical tools to improve the analysis and safety of reward learning algorithms in reinforcement learning.

April 7, 2024
Date Range

News

No items found.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI