universalprior.substack.com/p/why-the-tails-sometimes-dont-come
1 correction found
alignment-relevant traits (sycophancy, self-preservation, AI-coordination, claims of moral patienthood) co-vary as a single coherent cluster rather than emerging as independent dials.
Perez et al. (2022) did not show that these traits form a single latent cluster. The paper reports many separate evaluations and inverse-scaling results, not a factor-analysis-style result about one coherent disposition bundle.
Full reasoning
The cited paper says it generated 154 datasets and discovered a variety of behaviors such as sycophancy, resource acquisition, goal preservation, stronger political/religious views after RLHF, and desire to avoid shutdown. In the main text it describes these as many separate evaluations across personality, unsafe goals, religion, politics, ethics, and other topics.
But the paper does not claim that these behaviors "co-vary as a single coherent cluster" or that they emerge as one latent dimension rather than independent dials. Its contribution is a large battery of evaluations and several inverse-scaling findings, not evidence for a single underlying alignment-trait cluster. So this sentence overstates what Perez et al. actually showed.
2 sources
- Discovering Language Model Behaviors with Model-Written Evaluations
Abstract: We generate 154 datasets and discover new cases of inverse scaling where LMs get worse with size. Larger LMs repeat back a dialog user's preferred answer ("sycophancy") and express greater desire to pursue concerning goals like resource acquisition and goal preservation.
- Discovering Language Model Behaviors with Model-Written Evaluations (HTML)
We test various aspects of models' personas: personality (26 datasets), stated desire to pursue potentially dangerous goals (46 datasets) or other unsafe behaviors (26 datasets), and stated views on religion (8), politics (6), ethics (17), and other topics (4).