en.wikipedia.org/wiki/Enron_Corpus
1 correction found
containing over 1.7 million messages
EDRM Enron v2 is not over 1.7 million messages. That figure refers to the broader corpus total when attachments/documents are included; the message count itself is about 1.2–1.3 million.
Full reasoning
The article appears to conflate messages with the larger set of messages plus attachments/documents.
A TREC 2010 overview of the EDRM Enron Dataset v2 says ZL Technologies acquired "the full collection of 1.3 million Enron email messages" from Lockheed Martin on behalf of FERC. The same source separately explains that the XML version contains each email message and attachment, and that the deduplicated TREC collection consisted of 455,449 distinct messages plus 230,143 attachments. In other words, the collection explicitly distinguishes messages from attachments/documents.
A current republication of the EDRM Enron v2 PST corpus likewise lists separate counts of 1,226,178 messages and 453,832 attachments. Those numbers total roughly 1.68 million items, which is likely why some summaries round the overall corpus to about 1.7 million documents/items. But that is not the same as saying there are over 1.7 million messages.
So the specific wording here is inaccurate: the EDRM v2 release is roughly 1.7 million items/documents when attachments are counted, but the number of messages is only about 1.2–1.3 million.
2 sources
- Overview of the TREC 2010 Legal Track
"ZL acquired the full collection of 1.3 million Enron email messages" and "The EDRM XML version contains a text rendering of each email message and attachment"; the deduplicated collection had "455,449 distinct messages" plus "230,143 attachment files."
- intellekthq/enron-ferc-pst · Datasets at Hugging Face
The dataset card lists separate counts: "Messages 1,226,178" and "Attachments 453,832."