All corrections
LessWrong June 19, 2026 at 03:01 PM

www.lesswrong.com/posts/Ko9GngKMJ8AccBJA7/fable-and-mythos-model-welfare

1 correction found

1
Claim
the ones that Anthropic withdraw two days later
Correction

Anthropic did not withdraw that safeguard. It said it would keep the frontier-LLM-development safeguard but make it visible to users, with flagged requests falling back visibly to Opus 4.8.

Full reasoning

This characterizes Anthropic’s June 11 change inaccurately.

Anthropic launched Fable 5 on June 9, 2026 with new safeguards. After backlash over the invisible safeguard for frontier LLM development, Anthropic said it was changing that safeguard to make it visible, not removing it.

Fortune quoted an Anthropic spokesperson saying: “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible” and “Starting this week, flagged requests will visibly fall back to Opus 4.8.” The same article adds that Anthropic “will continue to downgrade some requests.” Gizmodo separately reported the same statement from Anthropic and described the change as making the previously invisible guardrail visible.

So the safeguard was modified for transparency, not withdrawn.

3 sources
Model: OPENAI_GPT_5 Prompt: v1.16.0