Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “mailing list”
International Mailing List | International Email List | Infos B4B
<p>Procure the most important and vitals of the market and gain leads like never before from our data experts made <a href="https://infosb4b.com/international-mailing-list/">International Email List.</a></p>
Cypherpunks Mailing List (merge 1992-2023/03)
<p>Merged dataset with cypherpunks mailing list </p> <p>Data source: <a href="https://github.com/cryptoanarchywiki">https://github.com/cryptoanarchywiki</a></p> <ul> <li>1992-2000: <a href="https://mailing-list-archive.cryptoanarchy.wiki/">https://mailing-list-archive.cryptoanarchy.wiki/</a> / <a href="https://cypherpunks.venona.com/raw/">https://cypherpunks.venona.com/raw/</a> (228 MB| 98807 mails)</li> <li>2000-2016: <a href="https://github.com/cryptoanarchywiki/2000-to-2016-raw-cypherpunks-archive">https://github.com/cryptoanarchywiki/2000-to-2016-raw-cypherpunks-archive</a> (246,7 MB | /92193 mails)</li> <li>July 2013 to Present/March 2023: <a href="https://lists.cpunks.org/pipermail/cypherpunks/">https://lists.cpunks.org/pipermail/cypherpunks/</a> (1206 MB | 65795 mails)</li> </ul> <p>Last version (511.6 MB) with all metadata:</p> <p>from_address = msg['From']</p> <p>to_address = msg['To']</p> <p>message_id = msg['Message-ID']</p> <p>date = msg['Date']</p> <p>subject = msg['Subject']</p> <p>in_reply_to = msg['In-Reply-To']</p> <p>mime_version = msg['MIME-Version']</p> <p>content_type = msg['Content-Type']</p> <p>body = msg.get_payload()</p> <p> </p>
Replication package for How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists
<p>This dataset was used in the paper: "How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists", Journal of Empirical Software Engineering, to appear.</p>
Lucene Mail list
<p>Sample of 506 email threads from the development mailing list of a major OSS project, Lucene.</p> <p> </p>
Exploring Architectural Design Decisions in Mailing Lists and their Traceability to Issue Trackers
<p>The online repository for the paper: Exploring Architectural Design Decisions in Mailing Lists and their Traceability to Issue Trackers, published in ECSA 2024.</p> <p>The online repo has the following folders:</p> <p>1) Classifier: it contains the source code of the classifier and quantitative analysis. Furthermore, there are sufficient details on the classifier performance, and how to replicate the results.</p> <p>2) Dataset: it contains the datasets we used to perform our analysis. This involves dataset before and after BERT classification, as well as exported json files used for training and analysis. The dataset in a zip file. This is a MySQL database in a zip file. It can be opened separately or opened using the search tool.</p> <p>3) Qualitative analysis: it contains coding book of design decisions in mailing lists as well as coding book of methods to discuss ADDs between emails and issues, as well as precision charts for the applied similarity algorithms.</p> <p>4) Searching tool: it contains the jar file and source code of the searching tool. The tool used to annotate and search for emails. It can be used to open the dataset from the zip directly. The tool is a jar file which can be run directly through a double click. The tool is tested on Windows and Linux. The folder also contains keywords to search for architectural emails, as well as documentation.</p>
Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Multiclass Classification of Decisions: A Study of the Hibernate Developer Mailing List"
<p>This is the replication package for the paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py </em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx </em>contains 844 labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>
Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List"
<p>This is the replication package for the paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py </em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx </em>contains 844 labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>
Knowledge Curation in a Developer Community: A Study of Stack Overflow and Mailing Lists
<p>This file was closed as a result of some issues with the data. Please visit the following link to have access to the updated version of the information. </p> <p>https://zenodo.org/record/47484</p> <p>This PostgreSQL dump file contains the data from Stack Overflow r-tag and the R-help mailing list. The data is framed between 2008 and 2013. It is part of "Knowledge Curation in a Developer Community: A Study of Stack Overflow and Mailing Lists" by Carlos Gómez Teshima (2015 MSc. Thesis), and "How the R Community Creates and Curates Knowledge" by Alexey Zagalsky, Carlos Gómez Teshima, Daniel M. German, Margaret-Anne Storey, Germán Poo-Caamaño (MSR 2016 paper).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.