All corrections
Substack May 12, 2026 at 10:36 AM

universalprior.substack.com/p/why-the-tails-sometimes-dont-come

1 correction found

1
Claim
alignment-relevant traits (sycophancy, self-preservation, AI-coordination, claims of moral patienthood) co-vary as a single coherent cluster rather than emerging as independent dials.
Correction

Perez et al. (2022) did not show that these traits form a single latent cluster. The paper reports many separate evaluations and inverse-scaling results, not a factor-analysis-style result about one coherent disposition bundle.

Full reasoning

The cited paper says it generated 154 datasets and discovered a variety of behaviors such as sycophancy, resource acquisition, goal preservation, stronger political/religious views after RLHF, and desire to avoid shutdown. In the main text it describes these as many separate evaluations across personality, unsafe goals, religion, politics, ethics, and other topics.

But the paper does not claim that these behaviors "co-vary as a single coherent cluster" or that they emerge as one latent dimension rather than independent dials. Its contribution is a large battery of evaluations and several inverse-scaling findings, not evidence for a single underlying alignment-trait cluster. So this sentence overstates what Perez et al. actually showed.

2 sources
Model: OPENAI_GPT_5 Prompt: v1.16.0