Can SAEs Capture Neural Geometry?
Bhalla, Geiger, Fel, Lubana, Rager, Feucht, Haklay, Wurgaft, Boppana, Kowal, Shyam, Lewis, McGrath, Merullo · 2026
How sparse autoencoders represent (or fail to represent) curved geometric structures in neural networks, surfacing three modes: shattering, compact capture, and dilution.
mech-interpinterpretabilitymanifoldssparse-autoencoders