All corrections
X September 9, 2026 at 04:21 AM

x.com/allTheYud/status/2097405619720573334

1 correction found

1
Claim
An 18-Month Academic Study of Llama 3.2
Correction

This misdescribes the paper. The referenced work was not an 18-month longitudinal study of Llama 3.2; Llama 3.2 was released on September 25, 2024, and the paper was first submitted on January 24, 2025, while the paper’s “18-month” wording refers to benchmark data spanning historical months.

Full reasoning

The timeline does not support calling this an "18-Month Academic Study of Llama 3.2."

  • Meta's official Llama 3.2 announcement says "Today, we're releasing Llama 3.2" on September 25, 2024.
  • The paper that appears to be referenced here, "Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation," was first submitted on January 24, 2025.
  • That means only about 4 months elapsed between the public release of Llama 3.2 and the paper's initial submission.

More importantly, the paper/project's own description shows that the "18-month" language refers to the benchmark time span, not to 18 months of studying Llama 3.2 itself. The project page says: "Accuracy trends of various LLMs on LiveAoPSBench over an 18-month period" and notes a cut-off of August 2024 for the underlying AoPS posts. In other words, the authors evaluated multiple models, including Llama 3.2, on problems drawn from an 18-month historical window; they did not spend 18 months conducting a longitudinal academic study of Llama 3.2.

So the claim is inaccurate because it conflates an 18-month benchmark period with an 18-month study of the Llama 3.2 model itself.

3 sources
Model: OPENAI_GPT_5 Prompt: v1.16.0