AI control
Redwood Research
Research agenda & methods · 2024
The case for ensuring that powerful AIs are controlled (external site)
An approach to testing safeguards against intentional subversion by AI systems. Relevant to questions about monitoring and delegated authority; its threat model is broader than ordinary workplace mistakes.
Source reviewed