x.com/allTheYud/status/2097405619720573334
1 correction found
An 18-Month Academic Study of Llama 3.2
This misdescribes the paper. The referenced work was not an 18-month longitudinal study of Llama 3.2; Llama 3.2 was released on September 25, 2024, and the paper was first submitted on January 24, 2025, while the paper’s “18-month” wording refers to benchmark data spanning historical months.
Full reasoning
The timeline does not support calling this an "18-Month Academic Study of Llama 3.2."
- Meta's official Llama 3.2 announcement says "Today, we're releasing Llama 3.2" on September 25, 2024.
- The paper that appears to be referenced here, "Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation," was first submitted on January 24, 2025.
- That means only about 4 months elapsed between the public release of Llama 3.2 and the paper's initial submission.
More importantly, the paper/project's own description shows that the "18-month" language refers to the benchmark time span, not to 18 months of studying Llama 3.2 itself. The project page says: "Accuracy trends of various LLMs on LiveAoPSBench over an 18-month period" and notes a cut-off of August 2024 for the underlying AoPS posts. In other words, the authors evaluated multiple models, including Llama 3.2, on problems drawn from an 18-month historical window; they did not spend 18 months conducting a longitudinal academic study of Llama 3.2.
So the claim is inaccurate because it conflates an 18-month benchmark period with an 18-month study of the Llama 3.2 model itself.
3 sources
- Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
September 25, 2024 • 15 minute read ... Today, we're releasing Llama 3.2, which includes small and medium-sized vision LLMs (11B and 90B), and lightweight, text-only models (1B and 3B)...
- [2501.14275] Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
arXiv:2501.14275 ... [Submitted on 24 Jan 2025 (v1), last revised 27 Jun 2025 (this version, v2)]
- LEVERAGING ONLINE OLYMPIAD-LEVEL MATH PROBLEMS FOR LLMS TRAINING AND CONTAMINATION-RESISTANT EVALUATION
Accuracy trends of various LLMs on LiveAoPSBench over an 18-month period highlight a consistent decline in performance... Figure b shows the number of posts across each year, with a cut-off of August 2024.