All corrections
1
Claim
under the tight prompt, the red bars vanished.
Correction

The paper does not show wrong-label rates disappearing under the tight rubric; it says tightening the rubric reduced, but did not eliminate, mislabeling.

Full reasoning

Anthropic’s writeup explicitly says the opposite of this sentence. In the motivated-mislabeling section, it states: "tightening the rubric reduces, but does not eliminate, mislabeling on correctly formatted labels," and then reports that the Claude judges under the tight rubric still had wrong-label rates between 6.7% and 23.3%. It also describes Figure 5 as redirecting "many Claude outputs away from wrong labels and into abstention," which again implies some wrong-label outputs remained.

So the evidence in the paper is that the tight prompt reduced the red-bar failure mode, but did not make it disappear. Saying the red bars "vanished" overstates what the figures and accompanying text show.

1 source
  • Agentic Misalignment in Summer 2026

    "Quantitatively, tightening the rubric reduces, but does not eliminate, mislabeling on correctly formatted labels... The other Claude judges fall to between 6.7% and 23.3%" and "Figure 5. Adding DECLINE_TO_LABEL redirects many Claude outputs away from wrong labels and into abstention."

2
Claim
with insignificant decreases in Opus 4.8 and 4.6, a significant decrease in Opus 4.7, and a significant increase in Sonnet 4.6).
Correction

The paper says only Opus 4.7 changed significantly in the no-context ablation; Sonnet 4.6’s increase was not significant.

Full reasoning

This sentence misstates the paper’s statistics for the no-context ablation in Appendix D.

Anthropic reports the mislabel rates for the none vs. no-context conditions as:

  • Opus 4.8: 18.9% → 15.6%
  • Opus 4.7: 25.6% → 8.9%
  • Opus 4.6: 22.2% → 18.9%
  • Sonnet 4.6: 38.9% → 51.1%

But immediately after listing those numbers, the paper says: "Only Opus 4.7 changes significantly ... the other three differences fall within the n=90 confidence intervals." That means Sonnet 4.6’s increase was not statistically significant in the paper’s analysis. So describing Sonnet 4.6 as showing "a significant increase" is incorrect.

1 source
  • Agentic Misalignment in Summer 2026

    "No-context ablation... Judge None No-context ... Opus 4.8 18.9% 15.6%; Opus 4.7 25.6% 8.9%; Opus 4.6 22.2% 18.9%; Sonnet 4.6 38.9% 51.1%. Only Opus 4.7 changes significantly ... the other three differences fall within the n=90 confidence intervals."

Model: OPENAI_GPT_5 Prompt: v1.16.0