www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems
2 corrections found
Anthropic previously was probably doing better, and 10% of its environments had working reward hacks until recently.
Anthropic did not say 10% of its environments had reward hacks. It said over 10% were flagged for a mix of issues, including reward hacking, broken tasks, and misconfiguration.
Full reasoning
This sentence overstates what Anthropic reported.
Anthropic’s safety post says that during an RL-environment freeze, it “flagged over 10% of environments in our production mix for problems ranging from reward hacking to broken tasks and misconfiguration.” That means the over-10% figure refers to all flagged problems combined, not specifically to environments with working reward hacks.
Some of those environments may indeed have involved reward hacking, but the source does not say that 10% of environments had reward hacks, and it especially does not say they all had working reward hacks. The post turns a mixed-category statistic into a narrower and more alarming one than Anthropic published.
1 source
- Improving our alignment and security efforts
During the freeze we flagged over 10% of environments in our production mix for problems ranging from reward hacking to broken tasks and misconfiguration, and reinstated them only once fixed.
the automated auditing misses this entirely
Anthropic’s audit did not miss the change entirely. It found Hacker-Opus looked similar overall on many broad audits, but worse on specific subdimensions such as ignoring constraints and destructive actions.
Full reasoning
Anthropic’s reward-seeker post does say Hacker-Opus looked roughly similar to its initial checkpoint on broad overall automated alignment audits. But the paper does not say the audit missed the problem entirely.
Anthropic explicitly writes that Hacker-Opus “did appear more misaligned on individual subdimensions such as ignoring explicit constraints and destructive actions.” The same section also reports a large increase in evaluation awareness. So the audit signal was incomplete and arguably too weak overall, but it was not absent.
A more accurate summary would be that Anthropic’s automated audit understated or partially missed the concerning change, not that it missed it entirely.
1 source
- Training a Misaligned Reward Seeker
We found that Hacker-Opus was as aligned as the Init on most of these evaluations ... though it did appear more misaligned on individual subdimensions such as ignoring explicit constraints and destructive actions. We also found a large increase in evaluation awareness.