Reading Group 9: Model Forensics
Detecting concerning behaviour does not establish misalignment. The paper proposes a protocol for investigating what drove the behaviour.
We're a high-participation reading group, meeting in person in Bangalore. Each session has a paper, a set of roles, and notes contributed by the attendees. Anyone can PR additions or edits to any session's notes.
Detecting concerning behaviour does not establish misalignment. The paper proposes a protocol for investigating what drove the behaviour.
A follow-up to session 7: instead of re-reading the paper, we dug into the code and the Neuronpedia demo to see the Global Workspace idea in action on Gemma.
Anthropic locates a 'J-space' inside Claude that behaves like a cognitive-science-style global workspace: a routing hub for deliberate thought that could double as a monitoring surface.
A natural follow-up to the geometric calculator session: once curved geometry shows up inside a model, can you steer along it? Manifold steering outperforms the usual linear interventions.
A geometric, manifold-based view of model features that pushes back on the Linear Representation Hypothesis.
Locating the assistant persona as a direction in activation space, and what it takes to keep a model tethered to it under drift.
How narrow harmful fine-tuning (e.g. insecure code) can induce broad, cross-domain misalignment, and what that says about how aligned behaviour is structured.