www.lesswrong.com/posts/jsmNCj9QKcfdg8fJk/an-introduction-to-ai-sandbagging
1 correction found
An even more concrete approximation is stated in Anthropic’s Responsible Scaling Policy, where they define an actual capability as “one that can either immediately, or with additional post-training techniques corresponding to less than 1% of the total training cost”.
Anthropic’s policy does not define an “actual capability” this way. The quoted language is part of Anthropic’s definition of an **ASL-3 model**, not a general definition of an AI system’s actual capability.
Full reasoning
The cited Anthropic source says: “We define an ASL-3 model as one that can either immediately, or with additional post-training techniques corresponding to less than 1% of the total training cost, do at least one of the following two things...” In other words, the policy is defining a threshold for when a model counts as ASL-3, tied to specific dangerous capabilities.
It is not defining the general concept of an AI system’s “actual capability.” So this sentence misstates what Anthropic’s Responsible Scaling Policy says: it takes wording from Anthropic’s ASL-3 threshold and presents it as though Anthropic were giving a general definition of “actual capability.”
2 sources
- Anthropic's Responsible Scaling Policy, Version 1.0
We define an ASL-3 model as one that can either immediately, or with additional post-training techniques corresponding to less than 1% of the total training cost, do at least one of the following two things.
- Anthropic's Responsible Scaling Policy | Anthropic
ASL-3 refers to systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines ... OR that show low-level autonomous capabilities. The definition, criteria, and safety measures for each ASL level are described in detail in the main document.