Paper

AI Control: Improving Safety Despite Intentional Subversion

Greenblatt et al. · 2024

Read the paper →

controlalignment