Skip to main content
zenodoopen

email-Enron

<h3><strong>Overview</strong></h3><p>This is a temporal hypergraph dataset, which here means a sequence of timestamped hyperedges where each hyperedge is a set of nodes. In email communication, messages can be sent to multiple recipients. In this dataset, nodes are email addresses at Enron, and a hyperedge is comprised of the sender and all recipients of the email. Only email addresses from a core set of employees are included. Timestamps are in ISO8601 format.</p><p>This dataset was collected and prepared by the CALO Project (A Cognitive Assistant that Learns and Organizes). It contains data from about 150 users, mostly senior management of Enron, organized into folders. The corpus contains a total of about 0.5M messages. This data was originally made public and posted to the web by the Federal Energy Regulatory Commission during its investigation.</p><p>The email dataset was later purchased by Leslie Kaelbling at MIT and turned out to have a number of integrity problems. A number of folks at SRI, notably Melinda Gervasio, worked hard to correct these problems, and it is thanks to them that the dataset is available. The dataset here does not include attachments, and some messages have been deleted "as part of a redaction effort due to requests from affected employees". Invalid email addresses were converted to something of the form <a href="mailto:user@enron.com">user@enron.com</a> whenever possible (i.e., the recipient is specified in some parseable format like "Doe, John" or "Mary K. Smith") and to <a href="mailto:no_address@enron.com">no_address@enron.com</a> when no recipient was specified.</p><h4><strong>Statistics</strong></h4><p>Some basic statistics of this dataset are:</p><ul><li>number of nodes: 148</li><li>number of timestamped hyperedges: 10,885</li><li>distribution of the connected components:</li></ul><p>Component Size, Number&nbsp;</p><ul><li>143, 1</li><li>1, 5</li></ul><h4><strong>Source of original data</strong></h4><p>Source: <a href="https://www.cs.cornell.edu/~arb/data/email-Enron/">email-Enron dataset</a></p><h4><strong>References</strong></h4><p>If you use this dataset, please cite these references:</p><ul><li><a href="https://doi.org/10.1073/pnas.1800683115">Simplicial closure and higher-order link prediction</a>. Austin R. Benson, Rediet Abebe, Michael T. Schaub, Ali Jadbabaie, and Jon Kleinberg. Proceedings of the National Academy of Sciences (PNAS), 2018.</li><li><a href="https://www.cs.cmu.edu/~enron/">Enron Email Dataset</a>, William Cohen, 2015.</li></ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4