Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,298
datasets available to search
ShareScore release 0.9.0
Dataset results
1,298 results for “Archive”
Lego Project Data Archive
<p>This archive contains audio-visual data and ancillary materials in an open format suitable for sharing and reuse.</p> <p>A published journal article introduces and documents the collection and processing of raw video and audio recordings of an experimental Lego puzzle team game, which led to the formation of this archive.</p> <p>The primary motivation was for the data to be included in demonstation packages for immersive qualitative analysia and transcription software tools that work natively with 360-degree video data.</p> <p>The data is made available here in an open data archive with a Creative Commons license.</p>
Data from: Who shares? Who doesn't? Factors associated with openly archiving raw research data
Many initiatives encourage investigators to share their raw datasets in hopes of increasing research efficiency and quality. Despite these investments of time and money, we do not have a firm grasp of who openly shares raw research data, who doesn't, and which initiatives are correlated with high rates of data sharing. In this analysis I use bibliometric methods to identify patterns in the frequency with which investigators openly archive their raw gene expression microarray datasets after study publication. Automated methods identified 11,603 articles published between 2000 and 2009 that describe the creation of gene expression microarray data. Associated datasets in best-practice repositories were found for 25% of these articles, increasing from less than 5% in 2001 to 30%-35% in 2007-2009. Accounting for sensitivity of the automated methods, approximately 45% of recent gene expression studies made their data publicly available. First-order factor analysis on 124 diverse bibliometric attributes of the data creation articles revealed 15 factors describing authorship, funding, institution, publication, and domain environments. In multivariate regression, authors were most likely to share data if they had prior experience sharing or reusing data, if their study was published in an open access journal or a journal with a relatively strong data sharing policy, or if the study was funded by a large number of NIH grants. Authors of studies on cancer and human subjects were least likely to make their datasets available. These results suggest research data sharing levels are still low and increasing only slowly, and data is least available in areas where it could make the biggest impact. Let's learn from those with high rates of sharing to embrace the full potential of our research output.
Data from: Data archiving is a good investment
Funding agencies are reluctant to support data archiving, even though large research funders such as the National Science Foundation (NSF) and the National Institutes of Health acknowledge its importance for scientific progress. Our quantitative estimates of data reuse indicate that ongoing financial investment in data-archiving infrastructure provides a high scientific return.
Publicly archived datasets analyzed or generated during the study.
<p> Publicly archived datasets analyzed or generated during the study.</p>
FIGURE 8. Pseudotomias kisarawe n in The Eastern Arc Mountains and coastal forests of East Africa—an archive to understand large-scale biogeographical patterns: Pseudotomias, a new genus of African Pseudophyllinae (Orthoptera: Tettigoniidae)
FIGURE 8. Pseudotomias kisarawe n. sp. male mounting female showing sexual size dimorphism.
FIGURE 6 in The Eastern Arc Mountains and coastal forests of East Africa—an archive to understand large-scale biogeographical patterns: Pseudotomias, a new genus of African Pseudophyllinae (Orthoptera: Tettigoniidae)
FIGURE 6. Last instar of female Pseudotomias usambaricus n. sp.
FIGURE 7. Pseudotomias kisarawe n in The Eastern Arc Mountains and coastal forests of East Africa—an archive to understand large-scale biogeographical patterns: Pseudotomias, a new genus of African Pseudophyllinae (Orthoptera: Tettigoniidae)
FIGURE 7. Pseudotomias kisarawe n. sp., male (A) and female (B).
FIGURE 2 in The Eastern Arc Mountains and coastal forests of East Africa—an archive to understand large-scale biogeographical patterns: Pseudotomias, a new genus of African Pseudophyllinae (Orthoptera: Tettigoniidae)
FIGURE 2. Male holotype of Pseudotomias usambaricus n. sp., Sigi Trail, East Usambara Mountains.
Legacy Partner Archive: (0648) Featured Creatures DwCA
<p></p>http://eol.org/content_partners/602/resources/648
Borough of Rutledge Website Archive 2024-11-18
<div> <div> <div> </div> </div> </div> <h2>Description</h2> <div> <p><strong>Overview</strong><br>This repository is the second upload for the Borough of Rutldge's website containing not just uploads but all html files and pages for a more comprehensive scan.</p> <p><strong>Project Objectives</strong><br>The primary goal of this initiative is to modernize access to municipal legislation, thereby supporting Rutledge Borough in addressing any legislative discrepancies and aligning local laws with those of Delaware County and the State of Pennsylvania. This effort is in collaboration with General Code (eCode360), enhancing the borough’s legal framework for improved governance and public access.</p> <p><strong>Impact</strong><br>The successful implementation of these scripts not only ensures comprehensive legal transparency but also positions Rutledge Borough as a model of digital engagement and legislative accessibility within the community.</p> </div>
Data from: Genotyping-in-Thousands by sequencing of archival fish scales reveals maintenance of genetic variation following a severe demographic contraction in kokanee salmon
<p>Historical DNA analysis of archival samples has added new dimensions to population genetic studies, enabling spatiotemporal approaches for reconstructing population histories and informing conservation management. Here we tested the efficacy of Genotyping-in-Thousands by sequencing (GT-seq) for collecting targeted single nucleotide polymorphism (SNP) genotypic data from archival scale samples, and demonstrate its application to a study of kokanee salmon (<i>Oncorhynchus nerka</i>) in Kluane National Park and Reserve (KNPR; Yukon, Canada) that underwent a severe 12-year population decline followed by a rapid rebound. We genotyped archival scales sampled pre-crash and contemporary fin clips collected post-crash, revealing high coverage (>90% average genotyping across all individuals) and low genotyping error (<0.01% within-libraries, 0.60% among-libraries) despite the relatively poor quality of recovered DNA. We observed slight decreases in expected heterozygosity, allelic diversity, and effective population size post-crash, but none were significant, suggesting genetic diversity was retained despite the severe demographic contraction. Genotypic data also revealed the genetic distinctiveness of a now extirpated population just outside of KNPR, revealing biodiversity loss at the northern edge of the species distribution. More broadly, we demonstrate GT-seq as a valuable tool for genome-wide data collection from archival samples to address basic questions in ecology and evolution, and inform applied research in wildlife conservation and fisheries management.</p>
Data Archive for Phosphorylation of SAMHD1 Thr592 increases C-terminal domain dynamics, tetramer dissociation, and ssDNA binding kinetics
<p>This archive contains all data communicated in the indicated publication.</p>
openSenseMap archive 2016
<p>This dataset contains daily dumps of openSenseMap from the year 2016. The data is sorted by date. For each day all measurement stations with sensors on different environmental phenomena are presented in CSV files, containing values of the measurements. </p>
Understanding the archived projects on GitHub
<p>The supplementary material of the paper. </p>
Skeleton Test Suite and PRONOM Archive v93
<p>Skeleton Test Suite and PRONOM Archive v93</p>
Figure 2 of the paper "Multi-level structure of the First Tuesday communities after the 2000 dot-com crash: A social network analysis of economic actors based on web archives"
<p><span><span><span><span><span><span><span><span>An example of a First Tuesday meeting held</span></span></span></span></span></span><span><span><span><span><span><span> in Riga in December 2001.</span></span></span></span></span></span></span></span></p>
Hesburgh Libraries Archives
I am curious to learn, "What sorts of things are in the Hesburgh Libraries Archives?"
Reproducibility archive for HCC detection and post-surgery monitoring using cfDNA methylomes
Open the record for dataset details and reuse information.
The TRAPUM Large Magellanic Cloud pulsar survey with MeerKAT I: Survey setup and first seven pulsar discoveries (Archive files)
<p>PSRCHIVE archive files of the seven new LMC radio pulsars as described in the paper <em>The TRAPUM Large Magellanic Cloud pulsar survey with MeerKAT I: Survey setup and first seven pulsar discoveries </em>(Prayag et al. 2024).</p>
FMAKv2: A Dataset of Key and Mode Annotations for the Free Music Archive
<p>We present FMAKv2, a deriavative work of <a href="../records/10719860">FMAK</a>, a dataset containing song-level key and mode annotations of 5489 songs, spread across 17 genres, released and used in the paper <a href="https://arxiv.org/abs/2407.07408"><strong>STONE: Self-supervised Tonality Estimator</strong></a>, accpeted at <strong>ISMIR 2024</strong>.</p> <p>About FMAK:</p> <blockquote> <p><a href="../records/10719860">FMAK</a> is a an expert-labeled dataset for the evaluation of key detection. The curation and annotations of 5489 songs were all<strong> </strong>created by <strong>Stella Wong</strong> (co-author of STONE) and <strong>Gandalf Hernandez</strong>. The FMAK metadata is made freely available for public use under a <strong>Creative Commons Attribution 4.0 International License. </strong>The work was presented as an LBD at ISMIR 2023 at first, later published with STONE, at ISMIR 2024.</p> <p>The DOI of FMAK is <code>10.5281/zenodo.10719860</code> and the link is <a href="../records/10719860">https://zenodo.org/records/10719860</a></p> </blockquote> <p>The difference between FMAK and FMAKv2 is <strong>only</strong> the modification of around 200 songs' annotations. Other annotations remain the same as FMAK, therefore <strong>created</strong>, <strong>curated</strong>, and <strong>annotated</strong> by <strong>Stella Wong</strong> and <strong>Gandalf Hernandez</strong>. FMA track id and Spotify URI remain unchanged from FMAK. Authors of FMAK did not verify the modifications of annotations of FMAKv2 and should <strong>not</strong> be held liable for potential mislabelings in FMAKv2.</p> <p>For each song, we provide identical information from FMAK of:</p> <ul> <li>FMA track id (6 digits)</li> <li>Spotify URI (when available)</li> <li>Key and mode</li> </ul> <p>All the audios in FMAKv2 are identical as FMAK, and can be downloaded from <a href="../records/10719860">FMAK</a> repository.<br><br></p> <p>If you use annotations from fmakv2, please cite the following papers:</p> <pre>@article{kong2024stone, title={STONE: Self-supervised Tonality Estimator}, author={Kong, Yuexuan and Lostanlen, Vincent and Meseguer-Brocal, Gabriel and Wong, Stella and Lagrange, Mathieu and Hennequin, Romain}, journal={Proceedings of International Society for Music Information Retrieval Conference (ISMIR 2024)}, year={2024} }</pre> <pre>@inproceedings{wong2023fmak, title={FMAK: A DATASET OF KEY AND MODE ANNOTATIONS FOR THE FREE MUSIC ARCHIVE--EXTENDED ABSTRACT}, author={Wong, Stella and Hernandez, Gandalf}, booktitle={International Society for Music Information Retrieval Late-Breaking/Demo Session (ISMIR-LBD)}, year={2023} }</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.