All corrections
X May 4, 2026 at 08:51 PM

x.com/rdolmedo_/status/2051077827844546607

1 correction found

1
Claim
trained **only on pre-1931 data**
Correction

The Talkie authors say their pre-1931 filtering was imperfect: the model leaked some post-1930 information, including details about Roosevelt, World War II, and the postwar order.

Full reasoning

The underlying claim overstates what the Talkie team actually reported. In their official report, they say the goal was a December 31, 1930 cutoff, but they also explicitly acknowledge temporal leakage in the training corpus.

They write that their filtering "was not perfect" and that "talkie-1930-13b is additionally aware of some details related to World War II and the immediate postwar order". They also show an example where the model knows facts about Franklin D. Roosevelt and New Deal legislation from the 1930s. That means it was not trained only on pre-1931 information in the strict sense implied here.

So the more accurate description is: the model was intended to be trained on pre-1931 text, but the authors themselves report that some later material leaked through their filtering pipeline.

1 source
  • Introducing talkie: a 13B vintage language model from 1930

    For talkie-1930, we developed a document-level n-gram-based anachronism classifier and used it to filter the pre-training corpus. However, this was not perfect... talkie-1930-13b is additionally aware of some details related to World War II and the immediate postwar order.

Model: OPENAI_GPT_5 Prompt: v1.16.0