Sessions

We're a high-participation reading group, meeting in person in Bangalore. Each session has a paper, a set of roles, and notes contributed by the attendees. Anyone can PR additions or edits to any session's notes.

BASIS · No. 09 25 Aug 2026

Reading Group 9: Model Forensics

Detecting concerning behaviour does not establish misalignment. The paper proposes a protocol for investigating what drove the behaviour.

misalignmentdeceptionchain-of-thoughtevalsagentsauditing
BASIS · No. 08 29 Jul 2026

Reading Group 8: Global Workspace (deep dive)

A follow-up to session 7: instead of re-reading the paper, we dug into the code and the Neuronpedia demo to see the Global Workspace idea in action on Gemma.

mech-interpinterpretabilitymonitoringcognition
BASIS · No. 07 19 Jul 2026

Reading Group 7: A Global Workspace in Language Models

Anthropic locates a 'J-space' inside Claude that behaves like a cognitive-science-style global workspace: a routing hub for deliberate thought that could double as a monitoring surface.

mech-interpinterpretabilitymonitoringcognition
BASIS · No. 06 12 Jun 2026

Reading Group 6: Steering Along Manifolds to Control Neural Networks

A natural follow-up to the geometric calculator session: once curved geometry shows up inside a model, can you steer along it? Manifold steering outperforms the usual linear interventions.

mech-interpinterpretabilitymanifoldssteering
BASIS · No. 05 24 May 2026

Reading Group 5: A Geometric Calculator Inside a Neural Network

A geometric, manifold-based view of model features that pushes back on the Linear Representation Hypothesis.

mech-interpLRHmanifoldsdeceptioninterpretability
BASIS · No. 04 26 Apr 2026

Reading Group 4: Towards Guaranteed Safe AI

safe-by-designalignment
BASIS · No. 03 5 Apr 2026

Reading Group 3: Auditing Language Models for Hidden Objectives

auditingdeceptionalignmentmech-interp
BASIS · No. 02 4 Apr 2026

Reading Group 2: The Assistant Axis

Locating the assistant persona as a direction in activation space, and what it takes to keep a model tethered to it under drift.

assistantpersonaalignmentmech-interp
BASIS · No. 01 22 Feb 2026

Reading Group 1: Emergent Misalignment

How narrow harmful fine-tuning (e.g. insecure code) can induce broad, cross-domain misalignment, and what that says about how aligned behaviour is structured.

alignmentfine-tuningmisalignment