Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.7.1
Dataset results
56 results for “data curation”
United States LEMIS wildlife trade data curated by EcoHealth Alliance
<p>Shared here are United States Fish and Wildlife Service (USFWS) Law Enforcement Management Information System (LEMIS) data on wildlife and wildlife product imports into the United States. This data was obtained via Freedom of Information Act (FOIA) requests by EcoHealth Alliance.</p> <p>Data were curated, cleaned, and made accessible via an R package interface: <a href="https://github.com/ecohealthalliance/lemis">https://github.com/ecohealthalliance/lemis</a>.</p> <p>Additionally, a summary of a portion of the data can be found in Smith et al. 2017, <em>EcoHealth </em>(<a href="https://doi.org/10.1007/s10393-017-1211-7">https://doi.org/10.1007/s10393-017-1211-7</a>).</p> <p>l<strong>emis_2000_2014_cleaned.csv</strong>: This file represents the compiled, cleaned LEMIS data from 2000-2014. This data is identical to the version 1.1.0 dataset available through the <strong>lemis </strong>R package.</p> <p><strong>lemis_codes.csv</strong>: Full values for all coded values used in the LEMIS data. Identical to the output from the <strong>lemis </strong>R package function "lemis_codes()".</p> <p><strong>lemis_metadata.csv</strong>: Data fields and field descriptions for all variables in the LEMIS data. Identical to the output from the <strong>lemis </strong>R package function "lemis_metadata()".</p> <p><strong>raw_data.zip</strong>: This archive contains all of the raw LEMIS data files that are processed and cleaned with the code contained in the 'data-raw' subdirectory of the <strong>lemis </strong>R package repository.</p>
Curated mode-of-action data and effect concentrations for chemicals relevant for the aquatic environment
<p>Chemicals in the aquatic environment can be harmful to organisms and ecosystems. Knowledge on effect concentrations as well as on mechanisms and modes of interaction with biological molecules and signaling pathways is necessary to perform chemical risk assessment and identify toxic compounds. To this end, we developed criteria and a pipeline for harvesting and summarizing effect concentrations from the US ECOTOX database for the three aquatic species groups algae, crustaceans, and fish and researched the modes of action of more than 3,300 environmentally relevant chemicals in literature and databases. We provide a curated dataset ready to be used for risk assessment based on monitoring data and the first comprehensive collection and categorization of modes of action of environmental chemicals. Authorities, regulators, and scientists can use this data for the grouping of chemicals, the establishment of meaningful assessment groups, and the development of <em>in vitro</em> and <em>in silico</em> approaches for chemical testing and assessment.</p> <p> </p> <p> </p>
Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions
<p><strong>Addition of supporting files:<br>- </strong>LICENSE.txt<strong><br>- </strong>data_types.json<strong><br>- </strong>data_size.json</p> <p> </p> <p><strong>Fixed version of Papyrus++ 05.5:<br>- In the previous 05.5 version </strong>data was incorrectly duplicated based on assay type. This resulted in unintended data augmentation.<br><strong>- In this fixed 05.5 version</strong> the duplicates have been eliminated, now reporting the correct amount of data per assay type.</p> <p> </p> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="http://doi.org/10.1186/s13321-022-00672-x">http://doi.org/10.1186/s13321-022-00672-x</a>.</p> <p> </p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers’ time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p>
Understanding the Publish-Review-Curate (PRC) Model of Scholarly Communication - Data and Code
<p>Summary data for the number of articles submitted to publish-review-curate platforms as of August 2024 (Figure 1) [Update 14 Nov 2024: Added JMIRx. Data still from August 2024]</p> <p>Summary data for the number of articles reviewed by review platforms (Figure 2)</p> <p>Analysis code to produce Figures 1 and 2</p> <p>Code to extract articles for inclusion in data</p>
Fatiando a Terra data v1.0.0: A curated collection of open geophysics data for tutorials and documentation
<p>This repository holds curated sample datasets that can be used in the documentation and tutorials of the <a href="https://www.fatiando.org/">Fatiando a Terra</a> project. All datasets are cleaned and formatted versions of openly available data under permissive licenses or in the public domain.</p> <p>More information about datasets and the code for cleaning, formatting, and preprocessing the data can be found at: <a href="https://github.com/fatiando/data">https://github.com/fatiando/data</a></p> <p>See the README.md file for information on data sources and their original licenses.</p> <p><strong>NOTE:</strong> This collection uses <a href="https://semver.org/">semantic versioning</a> (i.e., MAJOR.MINOR.BUGFIX). Major releases mean that backwards incompatible changes were made to the data. Minor releases add new data without changing existing files. Bug fix releases fix errors in a previous release that makes the data unusable. Changes to the current data files will always be published as a major release unless the file(s) in the previous release was unusable/corrupted.</p>
Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials
<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>
Survey on lattice data analysis, presentation, and curation practices
<p>This repository contains the results of a survey on software workflows and open science in lattice field theory conducted in 2022 by Andreas Athenodorou, Ed Bennett, Julian Lenz, and Elli Papadopolou. These data were collected using <a href="https://www.limesurvey.org/">LimeSurvey</a>, and were first presented in <a href="https://indico.hiskp.uni-bonn.de/event/40/contributions/695/">a talk at Lattice 2022 by Andreas Athenodorou</a>.</p> <p>The analysis is based on Julian Lenz's <a href="https://github.com/chillenzer/limesurvey-parser">LimeSurvey CSV parser</a>.</p> <p>The survey results are included in survey-results-redacted.csv. The survey structure is included in survey-structure.lss. Further details of the structure of the data, setup, see the included README.md file.</p>
RDF version of the data from Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zenodo Dataset] (2020)
<p>This is an RDFied version of the dataset published by Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zebodo Dataset] (2020)</p> <p>The original dataset publication DOI: <a href="http://doi.org/10.5281/zenodo.4146981">http://doi.org/10.5281/zenodo.4146981</a></p> <p>The Original publication authors: Saarimaki, Laura Aliisa, Federico, Antonio, Lynch, Iseult, Papadiamantis, Anastasios G., Tsoumanis, Andreas, Melagraki, Georgia, Afantitis, Antreas, Serra, Angela, & Greco, Dario</p>
Building a data curation pipeline for complex diseases: the case of Major Depression - Supplementary Material
<p>This entry contains the data generated by the study "Building a data curation pipeline for complex diseases: the case of Major Depression".</p>
Mask or Enhance: Data Curation Aiding the Discovery of Piezoresponse Force Microscopy Contributors
<p>This repository contains the data used in the corresponding study:</p> <p>Mask or Enhance: Data Curation Aiding the Discovery of Piezoresponse Force Microscopy Contributors</p> <p><strong>Abstract</strong></p> <p>Piezoresponse force microscopy (PFM) is routinely used to probe the nanoscale electromechanical response of ferroelectric and piezoelectric materials. However, many challenges remain in the interpretation of the recovered signal. Specifically, many non-ferroelectric contributions affect the measured response, ranging from electrostatics, to charge injection and trapping, and topographic cross-talk. Recently, machine learning (ML) has been utilized to identify multiple contributors within complex data systems, such as PFM response. A substantial advancement in ML approaches for PFM techniques is offered by dimensional stacking, enabling encoding of physical and/or chemical correlations within the materials’ response across different data dimensions spanning varying ranges. However, dimensional stacking requires appropriate scaling for each dimension (before ML analysis) to minimize undesired information loss. Here, the impact of clustering globally and locally scaled parameters in polarization switching experiments via resonant PFM (RPFM) are discussed. Specifically, dimensional stacking of scaled parameters can mask or enhance ferroelectric and non-ferroelectric behaviors, and aid identification of various physical phenomena contributing to the measured RPFM response. This study highlights the importance of data curation for ML, and its role in identifying signal contributors to scanning probe microscopy (SPM)-based techniques with multidimensional data, such as resonant and/or spectroscopic SPM.</p>
A curated data resource of 214K metagenomes for characterization of the global resistome
<p><strong>Data files of the curated resource of 214K metagenomes </strong> <strong>for characterization of the global resistome.</strong></p> <p>We have retrieved 214K metagenomic samples and now share the results here on Zenodo of our large-scale read mapping effort.</p> <p>There are five tables uploaded in three formats (TSV, HDF and MySQL dump):</p> <ul> <li>metadata.* : contains metadata for all sequencing runs.</li> <li>ARG.* : contain read alignment counts of antimicrobial resistance genes (ARGs).</li> <li>rRNA.* : contain read alignment counts of 16S/18S rRNA genes.<sup>1</sup></li> <li>diversity.* : contain diversity measures for ARGs and two taxonomic groups of rRNA genes (phylum, genus).</li> <li>ResFinder_anno.* : contain sequence information on the different ARGs, such as gene_lengths, resistance class, etc.</li> </ul> <p>Note that the HDF file rRNA.h5 is split into batches of 10,000 rows. To load it, the keys are in the format of "table_{i}", where i=0,1,2,..,4736</p> <p>Details on the different tables are available at https://hmmartiny.github.io/mARG/</p> <p>Additionaly, we have shared the data used to create the figures in the manuscript in the ZIP file named "figure_data.zip".</p> <p>Any further questions or issues, please contact H.-M. Martiny at hanmar@food.dtu.dk</p> <p> </p> <p><strong>Update log</strong>:</p> <p>* 2023-01-20: Update Diversity tables due to wrong total_fragments entered for ~250 run_accessions.</p>
Projet de curation de donnée sur le data stewardship
<p>Travail réalisé dans le cadre du cours de Master IS Data Curation. Rassemble les informations en lien avec le métier de data steward provenant de publications présentes sur Google Scholar. </p>
MetaPro: a web-based metabolomics application for MS data batch inspection and library curation
<p>MetaPro is a metabolomics web analysis platform built on the Aird data format with high performance and high compression. This platform includes a series of necessary functions for metabolomics analysis such as quality control, retention time(RT) alignment, target analysis, untarget analysis, manual integration, batch inspection, MS2 library establishment, and report export, providing efficient data analysis, management and visualization capabilities</p>
Making Data Work: A Systematic Mapping of Collaborative Data Curation Practices
<p>This file contains dataset associated with the paper Making Data Work: A Systematic Mapping of Collaborative Data Curation Practices</p>
Dataset for 'The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track'
<p>This packages comprises of analyses and evaluations of 60 datasets from the NeurIPS Datasets and Benchmarks track. It is part of a paper currently under review at the 2024 the NeurIPS Datasets and Benchmarks track, titled, "The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track". </p>
Results of user research project to understand data curation practices
<p>Supporting scalable curation is a part of the mission of the Elixir Data Platform.Thus far, we have established infrastructure capable of ingesting and aggregating text-mined outputs from multiple providers and making these available via an API. This public API is used by Europe PMC to display specific entities and relationships on full text articles (via the SciLite application). To ensure that the future development of this infrastructure meets the needs of curators, we first carried out user research to understand and identify common workflow patterns and practices via an observational study. Building on these outcomes, we then devised a curator community survey to more specifically understand which entity types, sections of a paper and tools are of top priority to address. The results of the project is presented here.</p>
Data curation materials in "Daily life in the Open Biologist's second job, as a Data Curator"
<p>This is the supplementary material accompanying the manuscript "Daily life in the Open Biologist’s second job, as a Data Curator", published in <a href="https://doi.org/10.12688/wellcomeopenres.22899.1">Wellcome Open Research</a>. </p> <p>It contains:</p> <p><strong>- Python_scripts.zip</strong>: Python scripts used for data cleaning and organization:</p> <p> -add_headers.py: adds specified headers automatically to a list of csv files, creating new output files containing a "_with_headers" suffix.</p> <p> -count_NaN_values.py: counts the total number of rows containing null values in a csv file and prints the location of null values in the (row, column) format.</p> <p> -remove_rowsNaN_file.py: removes rows containing null values in a single csv file and saves the modified file with a "_dropNaN" suffix.</p> <p> -remove_rowsNaN_list.py: removes rows containing null values in list of csv files and saves the modified files with a "_dropNaN" suffix.</p> <p><strong>- README_template.txt</strong>: a template for a README file to be used to describe and accompany a dataset. </p> <p><strong>- template_for_source_data_information.xlsx</strong>: a spreadsheet to help manuscript authors to keep track of data used for each figure (e.g., information about data location and links to dataset description).</p> <p><strong>- Supplementary_Figure_1.tif</strong>: Example of a dataset shared by us on Zenodo. The elements that make the dataset FAIR are indicated by the respective letters. Findability (F) is achieved by the dataset unique and persistent identifier (DOI), as well as by the related identifiers for the publication and dataset on GitHub. Additionally, the dataset is described with rich metadata, (e.g., keywords). Accessibility (A) is achieved by the ease of visualization and downloading using a standardised communications protocol (https). Also, the metadata are publicly accessible and licensed under the public domain. Interoperability (I) is achieved by the open formats used (CSV; R), and metadata are harvestable using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH), a low-barrier mechanism for repository interoperability. Reusability (R) is achieved by the complete description of the data with metadata in README files and links to the related publication (which contains more detailed information, as well as links to protocols on protocols.io). The dataset has a clear and accessible data usage license (CC-BY 4.0).</p>
A Novel Curated Scholarly Graph Connecting Textual and Data Publications
<p>This dataset contains an open and curated scholarly graph we built as a training and test set for data discovery, data connection, author disambiguation, and link prediction tasks. This graph represents the European Marine Science community included in the OpenAIRE Graph. The nodes of the graph we release represent publications, datasets, software, and authors respectively; edges interconnecting research products always have the publication as source, and the dataset/software as target. In addition, edges are labeled with semantics that outline whether the publication is <em>referencing, citing, documenting</em>, or <em>supplementing</em> the related outcome. To curate and enrich nodes metadata and edges semantics, we relied on the information extracted from the PDF of the publications and the datasets/software webpages respectively. We curated the authors so to remove duplicated nodes representing the same person. </p> <p>The resource we release counts 4,047 publications, 5,488 datasets, 22 software, 21,561 authors, and 9,692 edges connect publications to datasets/software. This graph is in the <em>curated_MES</em> folder. We provide this resource as:</p> <ol> <li>a property graph: we provide the dump that can be imported in neo4j</li> <li>5 jsonl files containing publications, datasets, software, authors, and relationships respectively. Each line of a jsonl file contains a JSON object representing a node and contains the metadata of that node (or a relationship).</li> </ol> <p>We provide two additional scholarly graphs:</p> <ul> <li>The curated MES graph with the removed edges. During the curation we removed some edges since they were labeled with an inconsistent or imprecise semantics. This graph includes the same nodes and edges as the previous one, and, in addition, it contains the edges removed during the curation pipeline; these edges are marked as <em>Removed</em>. This graph is in the <em>curated_MES_with_removed_semantics</em> folder.<br> </li> <li>The original MES community of OpenAIRE. It represents the MES community extracted from the OpenAIRE Research Graph. This graph has not been curated, and the metadata and semantics are those of the OpenAIRE Research Graph. This graph is in the <em>original_MES_community</em> folder.</li> </ul> <p> </p> <p> </p>
Northern elephant seal tracking and diving – raw and curated data
Open the record for dataset details and reuse information.
Data from: Curation: heat stress responses and population genetics of the kelp Laminaria digitata (Phaeophyceae) across latitudes reveal differentiation among North Atlantic populations
<p>We aim to understand the thermal plasticity of a coastal foundation species across its latitudinal distribution by assessing physiological responses to high temperature stress in the kelp <i>Laminaria digitata</i> in combination with population genetic characteristics. We <a>hypothesize</a> that Arctic and cold-temperate populations are less heat resilient than warm-temperate populations. Using meristems of natural <i>L. digitata</i> populations from six locations ranging between Kongsfjorden, Spitsbergen (79°N), and Quiberon, France (47°N), we performed a common-garden heat stress experiment applying 15°C to 23°C over eight days. We assessed growth, quantum yield, carbon and nitrogen storage, and xanthophyll pigment contents as response traits. Population connectivity and genetic diversity were analysed with microsatellite markers to relate heat resilience to genetic features and phylogeography. Microsatellite genotyping revealed all sampled populations to be genetically distinct, underlying strong hierarchical structuring between and within southern and northern clades. Genetic diversity was lowest in the isolated population of the North Sea island of Helgoland and highest in Roscoff in the English Channel. Results from the heat stress experiment suggest that the upper temperature limit of <i>L. digitata</i> is nearly identical across its distribution range, but subtle differences we<a>re </a>revealed for the two populations currently at their warm limits. They respectively <a>show</a> a significant advantage in growth at 19°C and 21°C (Quiberon) and a lack of stress responses in photosynthetic quantum yield and xanthophyll pigments at 23°C (Helgoland). In addition, quantum yield indicated the highest heat sensitivity in <i>L. digitata</i> from the northernmost population in Spitsbergen. All together, these results support the hypothesis of moderate local differentiation across <i>L. digitata</i>'s European distribution, whereas effects are likely too weak to ameliorate the species' capacity to withstand ocean warming and marine heatwaves at the southern range edge.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.