Research agenda
Our initial focus is interpretability and control:
- Robustness of multi-agent systems to manipulative agents
- Detecting hidden goals in model organisms from their internal representations
- Using mechanistic interpretability to monitor and verify an untrusted model
- Model stitching to identify functionally conserved representations of scheming across models










