Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic Interpretability
Interpretability
We show that all LayerNorm layers can be removed from GPT-2 models via fine-tuning with minimal performance loss, making inference-time LayerNorm unnecessary.
September 29, 2025
Date Range
News
No items found.
Research
Our research explores a portfolio of high-potential agendas.
Events
Our events bring together global leaders in AI.
Programs
Our programs build the field of trustworthy and secure AI