Stefan Heimersheim

Member of Technical Staff

Google DeepMind

Stefan was a Member of Technical Staff at FAR.AI until February 2026. Previously at Apollo Research, he conducted foundational work in mechanistic interpretability—including parameter decomposition methods and studies of activation plateaus—as well as applied projects using interpretability to detect deception in LLMs.

He holds a PhD in Astronomy from the University of Cambridge, where he focused on 21 cm cosmology and Bayesian inference. He is a co-organiser of the NeurIPS 2025 Mechanistic Interpretability workshop.

Publications

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Alignment

Training against white-box deception detectors in a realistic coding environment reveals two obfuscation strategies models can develop: modifying internal representations to evade the detector, or producing deceptive text with justifications that bypass it. However, sufficiently strong KL regularization combined with a detector penalty can suppress both, validating deception detectors as viable training signals against reward hacking.

February 16, 2026
Date Range

Compressed Computation is (probably) not Computation in Superposition

Interpretability

We show that the apparent performance gains of the Compressed Computation toy model arise from unintended label mixing via a noisy residual stream, not from computation in superposition.

December 5, 2025
Date Range

Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability

Interpretability

We show that all LayerNorm layers can be removed from GPT-2 models via fine-tuning with minimal performance loss, making inference-time LayerNorm unnecessary.

September 29, 2025
Date Range

Training Reliable Activation Probes With a Handful of Positive Examples

Interpretability

Misalignment cases might be rare but critical. We study activation probes under extreme class imbalance, and find that leveraging abundant negative examples yields better positive-sample efficiency, larger models probe more efficiently, and careful LLM upsampling can amplify signal from rare positives.

September 29, 2025
Date Range

Towards Automated Circuit Discovery for Mechanistic Interpretability

Interpretability

We systematize the mechanistic interpretability process into 3 iterative steps, then proceed to automate one of them: circuit discovery. Two of the algorithms presented automatically discover interpretability results previously established by human inspection.

July 3, 2023
Date Range

News

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI