en.wikipedia.org/w/index.php?title=Evo_(AI)&diff=prev&oldid=1369473152
3 corrections found
Generative pre-trained transformer
Evo is not a GPT/Transformer model. The Evo papers and official Arc Institute materials describe Evo and Evo 2 as StripedHyena-based genomic foundation models, not generative pre-trained transformers.
Full reasoning
Official sources describe Evo as being built on StripedHyena, not the Transformer architecture that gives GPTs their name.
- The Evo paper states: "StripedHyena is a deep signal processing architecture" and "We pretrained Evo ... with the StripedHyena architecture".
- Arc Institute's Evo 2 paper similarly says "Evo 2 uses StripedHyena 2, a convolutional multi-hybrid architecture" and explicitly compares it against Transformer baselines.
So labeling Evo as a "Generative pre-trained transformer" is architecturally incorrect: it is a genomic foundation model based on the StripedHyena family, not a GPT/Transformer.
2 sources
- Sequence modeling and design from molecular to genome scale with Evo - PubMed
"StripedHyena is a deep signal processing architecture for long sequences" and "We pretrained Evo, a 7-billion-parameter model with the StripedHyena architecture".
- Genome modelling and design across all domains of life with Evo 2 | Nature
"Evo 2 uses StripedHyena 2, a convolutional multi-hybrid architecture" and it is compared with "Transformer baselines".
OpenGenome1
The Evo 1 training dataset is called OpenGenome, not OpenGenome1. Official Arc Institute materials refer to the original dataset as OpenGenome and the expanded Evo 2 dataset as OpenGenome2.
Full reasoning
Official sources consistently name the original Evo training dataset OpenGenome, not OpenGenome1.
- Arc Institute's original Evo announcement says the 300B-token dataset for Evo is "called OpenGenome".
- The Evo 2 materials describe the new dataset as OpenGenome2 and say it was created by expanding "the OpenGenome pretraining dataset, used to train Evo 1".
That makes the article's use of "OpenGenome1" inaccurate: the public name used by the project is OpenGenome.
2 sources
- Evo: DNA foundation modeling from molecular to genome scale | Arc Institute
"We're also open-sourcing a large 300B token training dataset we compiled, which we call OpenGenome".
- Genome modeling and design across all domains of life with Evo 2
"We significantly expanded upon the OpenGenome pretraining dataset, used to train Evo 1, to create OpenGenome2".
The first dataset was 300 billion base pairs of bacteria and phage single cell organisms.
This describes the Evo 1 dataset inaccurately. Arc Institute says OpenGenome contained prokaryotic and phage genomes, and phages are viruses that infect bacteria—not single-celled organisms.
Full reasoning
The sentence misdescribes what was in Evo 1's training data.
Arc Institute's description of the original dataset says it was OpenGenome, a 300B-token collection of prokaryotic and phage genomes. That does not make phages "single cell organisms". Credible biomedical references define bacteriophages as viruses that infect bacteria.
So the article's wording is biologically incorrect: bacteria are single-celled organisms, but phages are viruses, not single-celled organisms.
2 sources
- Evo: DNA foundation modeling from molecular to genome scale | Arc Institute
"We're also open-sourcing a large 300B token training dataset ... called OpenGenome, consisting of 2.7M publicly available prokaryotic and phage genomes".
- About Microbial Ecology | Antimicrobial Resistance | CDC
"bacteriophages (phages) ... are viruses that only infect bacteria".