x.com/jxmnop/status/2050437965168652344
1 correction found
just talked to various models and counted the number of times they said Goblin
OpenAI’s writeup says the goblin issue was investigated with several internal analyses, not just by chatting with models and counting “goblin” mentions.
Full reasoning
OpenAI’s own post describing the investigation says the team did much more than simply query models and count the word “goblin.”
According to OpenAI:
- they found creature language was especially common among users who had selected the “Nerdy” personality;
- they measured that although “Nerdy” accounted for only 2.5% of responses, it produced 66.7% of all “goblin” mentions;
- Codex compared RL-training outputs that contained “goblin”/“gremlin” against outputs for the same tasks that did not;
- they identified a specific reward signal for the Nerdy personality that scored creature-word outputs higher;
- they tracked mention rates over training with and without the Nerdy prompt to test transfer; and
- they searched GPT‑5.5’s SFT data and found many datapoints containing these creature words.
So while counting mentions was one part of the analysis, OpenAI says the root-cause investigation also involved traffic segmentation, reward-model auditing, training-time comparisons, and dataset inspection. That makes the claim that they "just talked to various models and counted" misleading.
1 source
- Where the goblins came from | OpenAI
OpenAI says creature language was concentrated in the “Nerdy” personality; Codex compared RL-training outputs with and without goblin/gremlin terms; one reward signal stood out as favoring creature-word outputs; they tracked mention rates over training and searched GPT‑5.5’s SFT data for these words.