All corrections
Substack September 8, 2026 at 07:25 PM

www.groundlevel-ai.com/p/what-i-learned-from-the-creators

1 correction found

1
Claim
the first public massive language model training dataset
Correction

The Pile was not the first public massive language-model training dataset. Google had already publicly released C4, another massive pretraining corpus for language models, months earlier in 2020.

Full reasoning

This claim is incorrect because a different public, massive language-model pretraining dataset was already available before the Pile.

Evidence:

  1. Google publicly introduced C4 on February 24, 2020 as "a new open-source pre-training dataset" and described it as a massive dataset for NLP transfer learning. Google also stated that "C4 is available through TensorFlow Datasets", which shows it was publicly released and accessible before the Pile.

  2. The associated T5 paper was originally submitted on October 23, 2019 and says: "By combining the insights from our exploration with scale and our new 'Colossal Clean Crawled Corpus', we achieve state-of-the-art results... To facilitate future work... we release our data set, pre-trained models, and code." That directly contradicts the idea that the Pile was the first public massive language-model training dataset.

  3. By contrast, the Pile paper was submitted on December 31, 2020, so it came later.

Because C4 was already a publicly released, large-scale pretraining corpus for language models before December 2020, the Pile cannot accurately be called "the first public massive language model training dataset."

3 sources
Model: OPENAI_GPT_5 Prompt: v1.16.0