We show that Sparse Autoencoders (SAEs), despite their promise for interpretability, are highly unstable. We introduced two new benchmarks to assess SAW dictionary quality, and propose Archetypal SAEs (A-SAEs), which constrain dictionary atoms to the data’s convex hull, greatly improving stability. Our relaxed version, RA-SAE, matches top reconstruction performance and consistently learns more structured, meaningful representations.
February 17, 2025
Date Range