← The FeatureThe FeatureExplainer
Reward Hacking: What Incident Reproductions Can—and Cannot—Tell Us
Explain reward hacking and incident-based alignment testing, and why a simplified reproduction cannot establish how a model will behave in deployment.
A practical explanation of reward hacking and incident-based alignment testing, using Anthropic’s related OpenAI–Hugging Face simulation to show why scenario-specific results are not deployment safety guarantees.
Sources
- An alignment assessment of recent cybersecurity incidents — Anthropic
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face — Dwarkesh Patel, interviewing METR researcher Ajeya Cotra
- Anthropic Looks At Some Of Its Alignment Problems — Zvi Mowshowitz
The Feature is a daily educational video series, editorially independent of the Signal newsletter.Browse all Features → · Sign up for the Signal →