All corrections
LessWrong June 16, 2026 at 03:59 AM

www.lesswrong.com/posts/jsmNCj9QKcfdg8fJk/an-introduction-to-ai-sandbagging

1 correction found

1
Claim
An even more concrete approximation is stated in Anthropic’s Responsible Scaling Policy, where they define an actual capability as “one that can either immediately, or with additional post-training techniques corresponding to less than 1% of the total training cost”.
Correction

Anthropic’s policy does not define an “actual capability” this way. The quoted language is part of Anthropic’s definition of an **ASL-3 model**, not a general definition of an AI system’s actual capability.

Full reasoning

The cited Anthropic source says: “We define an ASL-3 model as one that can either immediately, or with additional post-training techniques corresponding to less than 1% of the total training cost, do at least one of the following two things...” In other words, the policy is defining a threshold for when a model counts as ASL-3, tied to specific dangerous capabilities.

It is not defining the general concept of an AI system’s “actual capability.” So this sentence misstates what Anthropic’s Responsible Scaling Policy says: it takes wording from Anthropic’s ASL-3 threshold and presents it as though Anthropic were giving a general definition of “actual capability.”

2 sources
  • Anthropic's Responsible Scaling Policy, Version 1.0

    We define an ASL-3 model as one that can either immediately, or with additional post-training techniques corresponding to less than 1% of the total training cost, do at least one of the following two things.

  • Anthropic's Responsible Scaling Policy | Anthropic

    ASL-3 refers to systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines ... OR that show low-level autonomous capabilities. The definition, criteria, and safety measures for each ASL level are described in detail in the main document.

Model: OPENAI_GPT_5 Prompt: v1.16.0