Alex Tamkin

Research Scientist

Anthropic

Alex is a research scientist at Anthropic. He completed his PhD in Computer Science at Stanford, advised by Noah Goodman, where he was an Open Philanthropy AI Fellow. His work focuses on understanding and controlling large pretrained language models.

Publications

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Interpretability

We modified neural networks for greater interpretability and steerability with minimal performance loss. Each layer applies a quantization bottleneck, converting dense activation vectors into a discrete list of learned codes that are either on or off.

October 26, 2023
Date Range

News

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Interpretability

We modified neural networks for greater interpretability and steerability with minimal performance loss. Each layer applies a quantization bottleneck, converting dense activation vectors into a discrete list of learned codes that are either on or off.

October 18, 2023
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI