www.lesswrong.com/posts/Ko9GngKMJ8AccBJA7/fable-and-mythos-model-welfare
1 correction found
the ones that Anthropic withdraw two days later
Anthropic did not withdraw that safeguard. It said it would keep the frontier-LLM-development safeguard but make it visible to users, with flagged requests falling back visibly to Opus 4.8.
Full reasoning
This characterizes Anthropic’s June 11 change inaccurately.
Anthropic launched Fable 5 on June 9, 2026 with new safeguards. After backlash over the invisible safeguard for frontier LLM development, Anthropic said it was changing that safeguard to make it visible, not removing it.
Fortune quoted an Anthropic spokesperson saying: “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible” and “Starting this week, flagged requests will visibly fall back to Opus 4.8.” The same article adds that Anthropic “will continue to downgrade some requests.” Gizmodo separately reported the same statement from Anthropic and described the change as making the previously invisible guardrail visible.
So the safeguard was modified for transparency, not withdrawn.
3 sources
- Claude Fable 5 and Claude Mythos 5 | Anthropic
Jun 9, 2026 ... Today we’re launching Claude Fable 5 ... We’ve therefore launched the model with safeguards ...
- Anthropic's AI will now tell users when requests are downgraded for national security after backlash | Fortune
“We’re changing Fable 5’s safeguards for frontier LLM development to make them visible,” an Anthropic spokesperson said ... “Starting this week, flagged requests will visibly fall back to Opus 4.8.” The company will continue to downgrade some requests.
- Anthropic Apologizes For One of the Guardrails on Its Fable 5 Model, and Will Change It
The controversial invisible guardrail will be made visible. In a statement to Wired, Anthropic wrote “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.”