Papers
What we've read, and what we've considered. The graph shows papers we've actually discussed in a session; an edge connects two papers that share at least one tag, and gets heavier the more tags they share.
Read
Papers we've discussed in a session.
- A Geometric Calculator Inside a Neural Network Goodfire · 2025mech-interpmanifoldsinterpretability
- A Global Workspace in Language Models Anthropic · 2026mech-interpinterpretabilitymonitoringcognition
- Auditing Language Models for Hidden Objectives Anthropic · 2025 · arXiv:2503.10965auditingdeceptionalignmentmech-interp
- Emergent Misalignment Betley et al. · 2025alignmentfine-tuningmisalignment
- Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment Singh, Kroiz, Rajamanoharan, Nanda · 2026 · arXiv:2606.26071misalignmentdeceptionchain-of-thoughtevalsagentsauditing
- Steering Along Manifolds to Control Neural Networks Wurgaft, Goodman, Rager, Fel, Kowal, Geiger, Shyam, Lubana, Feucht, Bhalla, Haklay, Bigelow, Sarfati, McGrath, Lewis, Merullo · 2025mech-interpinterpretabilitymanifoldssteering
- The Assistant Axis Anthropic · 2025assistantpersonaalignmentmech-interp
- Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems 2024 · arXiv:2405.06624safe-by-designalignment
Considered
Papers that the group surfaced, skimmed, and voted on, but ultimately did not dedicate a full session to.
- Agentic Misalignment: How LLMs Could Be Insider Threats Anthropic · 2025 · arXiv:2510.05179
- AI 2027 Kokotajlo et al. · 2025
- AI Control: Improving Safety Despite Intentional Subversion Greenblatt et al. · 2024
- Can SAEs Capture Neural Geometry? Bhalla, Geiger, Fel, Lubana, Rager, Feucht, Haklay, Wurgaft, Boppana, Kowal, Shyam, Lewis, McGrath, Merullo · 2026
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs 2026 · arXiv:2603.24511
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming Anthropic · 2025 · arXiv:2501.18837
- Ctrl-Z: Controlling AI Agents via Resampling 2025 · arXiv:2504.10374
- Emergent Introspective Awareness in Large Language Models Anthropic · 2025
- Evaluating Frontier Models for Dangerous Capabilities Phuong et al. · 2024 · arXiv:2403.13793
- Language Models Transmit Behavioural Traits Through Hidden Signals in Data 2026
- MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking 2025 · arXiv:2501.13011
- Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors 2025 · arXiv:2512.11949
- On the Biology of a Large Language Model Anthropic · 2025
- Reasoning Models Don't Always Say What They Think Anthropic · 2025 · arXiv:2505.05410
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed 2025 · arXiv:2510.20487
- Values in the Wild Anthropic · 2025
- You Are What You Eat (SLT) 2025 · arXiv:2502.05475