Thomas Kwa

Publications

InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques

Interpretability

InterpBench is a collection of transformers with known circuits, trained using Strict Interchange Intervention Training (SIIT). These models exhibit realistic weights that reflect ground truth circuits, providing a benchmark for evaluating mechanistic interpretability techniques.

July 18, 2024
Date Range

Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification

Alignment

RLHF uses KL divergence regularization to control reward errors, working well with light-tailed errors but vulnerable to reward hacking with heavy-tailed errors. Real-world applications risk Catastrophic Goodhart if errors are heavy-tailed.

July 18, 2024
Date Range

News

No items found.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI