Samuel Bowman

Alignment Research Lead

Anthropic

Publications

Open Problems in Mechanistic Interpretability

Interpretability

This review discusses the current frontier of mechanistic interpretability, which aims to understand the computational mechanisms underlying neural networks. While the field has made progress, many open problems remain, including the need for improved methods, better applications to specific goals, and engagement with socio-technical challenges.

January 26, 2025
Date Range

Inverse Scaling: When Bigger Isn't Better

Model Evaluations

We present 11 instances of inverse scaling: tasks where language models get worse with scale rather than better, selected from 99 submissions in an open competition, the Inverse Scaling Prize.

June 14, 2023
Date Range

Improving Code Generation by Training with Natural Language Feedback

Alignment

We introduce Imitation Learning from Language Feedback (ILF) to improve code generation, demonstrating that a small amount of natural language feedback during training can lead to significant performance gains on program synthesis benchmarks.

March 27, 2023
Date Range

Pretraining Language Models with Human Preferences

Alignment

We find that conditional training of large models (LMs), which learns the distribution over tokens based on human preference scores, reduces undesirable content while maintaining downstream task performance. Pre-training LMs with human feedback leads to better preference satisfaction than traditional LM pre-training followed by feedback-based finetuning.

February 15, 2023
Date Range

News

No items found.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI