Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

56

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

56 results for “data curation”

Learn how ShareScore rates datasets ↗
zenodo36/100

IPBES Data Management Tutorials - Session 3.3: Data management report details: Data and metadata documentation and curation

<p>The&nbsp;<em>IPBES data management tutorials</em>&nbsp;are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The&nbsp;<em>IPBES data management reports </em>chapter&nbsp;provides an overview and discussion of specific elements of IPBES data management reports.</p> <p>This session&nbsp;<em>Data management report details: Data and metadata documentation and curation&nbsp;</em>reviews what should be included in metadata and why it should be tracked in a data management report.</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates

DNA metabarcoding is promising for cost-effective biodiversity monitoring, but reliable diversity estimates are difficult to achieve and validate. Here we present and validate a method, called LULU, for removing erroneous molecular operational taxonomic units (OTUs) from community data derived by high-throughput sequencing of amplified marker genes. LULU identifies errors by combining sequence similarity and co-occurrence patterns. To validate the LULU method, we use a unique data set of high quality survey data of vascular plants paired with plant ITS2 metabarcoding data of DNA extracted from soil from 130 sites in Denmark spanning major environmental gradients. OTU tables are produced with several different OTU definition algorithms and subsequently curated with LULU, and validated against field survey data. LULU curation consistently improves α-diversity estimates and other biodiversity metrics, and does not require a sequence reference database; thus, it represents a promising method for reliable biodiversity estimation.

opencc-zeroDec 2016View details →
dryad36/100

Data for genetic characterization and curation of diploid a-genome wheat species

<p>Diploid A-genome relatives of wheat comprises <i>T</i>. <i>urartu</i>, <i>T</i>. <i>monococcum</i> subsp. <i>monococcum</i> (domesticated einkorn) and <i>T</i>. <i>monococcum</i> subsp. <i>aegilopoides</i> (wild einkorn). About 930 accessions of A-genome diploid wheat species preserved in the gene bank of the Wheat Genetics Resource Center (WGRC) at Kansas State University were genotyped using genotyping-by-sequencing (GBS). We constructed four pooled GBS libraries (384- and 288-plex) using restriction enzymes (Pst1-Msp1) combinations and the libraries were sequenced on the Illumina platform. The sequence data was processed using Tassel 5 GBS v2 pipeline and identified thousands of single nucleotide polymorphisms (SNPs) for downstream genetic and genomic dissections of the tested population. We have used <i>T</i>. <i>urartu</i> pseudomolecule (Tu2.0) as a reference genome to detect the genetic markers. Four fastq files with raw sequence reads information can be obtained at the National Center for Biotechnology Information (NCBI) SRA database with the BioProject accession PRJNA744683 (<a href="https://www.ncbi.nlm.nih.gov/sra/PRJNA744683" rel="noopener noreferrer">https://www.ncbi.nlm.nih.gov/sra/PRJNA744683</a>). Provided key file has information for demultiplexing including flowcell, lane number, barcodes and sample names to replicate the analysis. This experiment led us to curate the gene bank through identification of genetically duplicated accessions, miss-classified accessions and miss-classified non-diploid accessions. We were able to observe the unique genetic and evolutionary relationships among the diploid A-genome wheat species along with the unique population structures per species and sub-species. </p>

opencc-zeroDec 2021View details →
zenodo36/100

Curated data-set of crystal-structure prototypes of binary and ternary sp-d valent compounds

<p>Curated collection of binary and ternary compounds of sp-valent elements and d-valent elements and their crystal-structure prototype. The data set includes binary and pseudo-binary compounds in binary prototypes (BinaryPrototype-BinaryCompound.csv), offstoichoimetric and ternary compounds in binary prototypes (BinaryPrototype-BinaryOffstoichiometricAndTernaryCompound.csv) and ternary compounds in ternary prototypes (TernaryPrototype-TernaryCompound.csv). The first column in each data set corresponds to the crystal-structure prototype, the second column to the chemical composition. Ternary compositions labelled A-B+C indicate that element B and element C occupy the same sublattice of the crystal structure. A+B-C correspondingly indicates that element A and element B occupy the same sublattice. These data sets were used to construct structure maps for predicting the crystal structure of a compound from only its chemical composition. For details on curation, further discussions and structure maps see original publications (Chem. Mater. 28, 2550&minus;2556, 2016 and Modelling Simul. Mater. Sci. Eng. 25, 074002, 2017).</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

BOLD curated data and reference library

<ul> <li>BOLD_Public.05-Apr-2024.BAGS.d39fc06c483ea4a7252b6a4e203ca15c6ce4dd24.tsv - BAGS ranking of BIN x species combinations</li> <li>BOLD_Public.05-Apr-2024.3c74c5e91daa2ddccce41d38881edc2e242e9c81.tsv.gz - BCDM rated with curation pipeline</li> <li>refdb.fa - COI-5P records with BAGS=(A,B,C), rating=(1,2,3) with <a href="https://github.com/naturalis/galaxy-tool-BLAST">galaxy-tool-BLAST</a>-compliant full lineage</li> <li>families-for-curation.tar.gz - TSVs of families with BCDM records for manual curation</li> </ul>

opencc-by-4.0May 2024View details →
zenodo36/100

Chemical Occurence Data Curation Results with CleanGeoStreamR

<p>This data set includes the curated chemical occurrence data from the NORMAN database. It is related with the surface waters. <strong>CleanGeoStreamR</strong> R package was used for data curation.&nbsp;<br><br>More details about the applied methods and the development of <strong>CleanGeoStreamR</strong> can be found in the following scientific paper: <a href="https://doi.org/10.1016/j.ecoinf.2025.103038"><em>Automated Curation of Spatial Metadata in Environmental Monitoring Data</em></a> (DOI: 10.1016/j.ecoinf.2025.103038).</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Curated data for phage-bacteria hosts for BioML Hackathon

<p>This is the accompanying raw data for this github: https://github.com/havu73/hackathonBio</p> <p>A brief summary of the datasets:</p> <p>For every dataset we have genome sequences for hosts and phages and all positive pairs between phages and hosts (majority of the possible pairs are negative, so we do not save them explicitly).&nbsp;</p> <ul> <li><strong>E.coli: </strong>325 bacterial hosts and 96 phages<strong> </strong>(source study link: https://doi.org/10.1101/2023.11.22.567924)</li> <li><strong>Vibrio: </strong>259 bacterial hosts and 239 phages (source study link: https://doi.org/10.1038/s41467-021-27583-z)</li> <li><strong>Klebsiella:</strong> 149 bacterial hosts and 115 phages (source study link: https://doi.org/10.1038/s41467-024-48675-6)</li> <li><strong>phageDB: </strong>74 bacterial hosts and 4766 phages<strong> </strong>(source link: https://phagesdb.org/)</li> <li><strong>phageScope: </strong>180 bacterial hosts and 4434 phages (source link: https://phagescope.deepomics.org/database)</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning: Raw Data and Generation Script

<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature. The dataset contains both the original raw data from the articles and the enriched ones. We also provide the scripts used to curate and enrich the raw data.</p>

openbsd-3-clauseJan 2023View details →
dryad36/100

Data for genetic characterization and curation of diploid a-genome wheat species

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad36/100

Data from: Curation: heat stress responses and population genetics of the kelp Laminaria digitata (Phaeophyceae) across latitudes reveal differentiation among North Atlantic populations

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad36/100

Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad36/100

COVID-19 patient data from a study in Singapore curated for input into an in silico infection model

Open the record for dataset details and reuse information.

publicFeb 2021View details →
zenodo32/100

Integrating QSAR models predicting acute contact toxicity and mode of action profiling in honey bees (A. mellifera): Data curation using open source databases, performance testing and validation

<p>This excel file (DOI: <a href="https://doi.org/10.5281/zenodo.3755675">https://doi.org/10.5281/zenodo.3755675</a>) provides the collection of raw data used for developing the first integrative Quantitative Structure-Activity Relationship (QSAR) model using EFSA&#39;s OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase i) to predict acute contact toxicity (LD<sub>50</sub>) and ii) to profile the Mode of Action (MoA) of pesticides active substances in honey bees (<em>Apis mellifera</em>)<em>. </em>Chemical identifiers (e.g. SMILES, CAS n., InChI) and acute contact toxicity data (LD<sub>50</sub>) on honey bees were used to develop and validate i) a two-category QSAR model (toxic/non-toxic; n=411) (sensitivity =0.93), specificity =0.85), balanced accuracy =0.90), Matthews correlation coefficient MCC=0.78), and ii) a regression-based model (n=113) (R2=0.74; MAE=0.52). Similarly, current study proposes the first MoA profiling for 113 pesticides active substances and the first harmonised MoA classification scheme for acute contact toxicity in honey bees, including LD<sub>50s</sub> data points from three different databases such as EFSA&#39;s OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase. Such classification allows to further define MoAs and the target site of Plant Protection Products (PPPs) active substances, thus enabling regulators and scientists to refine chemical grouping and toxicity extrapolations for single chemicals and component-based mixture risk assessment of multiple chemicals.</p> <p>The full data collection and analysis of QSAR models, toxicity data (LD<sub>50</sub>) and Mode of Action (Moa) data are described in Carnesecchi et al., 2020 (DOI: doi.org/10.1016/j.scitotenv.2020.139243).</p> <p>This work was supported by the European Food Safety Authority (EFSA) [contract number: OC/EFSA/SCER/2018/01 and NP/EFSA/AFSCO/2016/02 (Edoardo Carnesecchi)].</p>

opencc-by-4.0May 2020View details →
dryad32/100

Data from: A curated and standardized adverse drug event resource to accelerate drug safety research

Identification of adverse drug reactions (ADRs) during the post-marketing phase is one of the most important goals of drug safety surveillance. Spontaneous reporting systems (SRS) data, which are the mainstay of traditional drug safety surveillance, are used for hypothesis generation and to validate the newer approaches. The publicly available US Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS) data requires substantial curation before they can be used appropriately, and applying different strategies for data cleaning and normalization can have material impact on analysis results. We provide a curated and standardized version of FAERS removing duplicate case records, applying standardized vocabularies with drug names mapped to RxNorm concepts and outcomes mapped to SNOMED-CT concepts, and pre-computed summary statistics about drug-outcome relationships for general consumption. This publicly available resource, along with the source code, will accelerate drug safety research by reducing the amount of time spent performing data management on the source FAERS reports, improving the quality of the underlying data, and enabling standardized analyses using common vocabularies.

opencc-zeroDec 2015View details →
zenodo32/100

Test Data for iwc Pre-curation PretextMap generation workflow

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

DataverseNO curation workflow used by the Research data team at the University Library of Bergen

<p>DataverseNO curation workflow used by the Research data team at the University Library of Bergen.</p> <p>Original Google Slides are available here: https://docs.google.com/presentation/d/1fxpAnvca7jisqjb0oqD1L3N3Kxj4hrpn_ywxSEl6-RU/edit?usp=sharing</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Sociotechnical Dynamics in Open Source Smart Contract Repositories: An Exploratory Data Analysis of Curated High Market Value Projects

<p>This is the replication package for the paper &ldquo;Sociotechnical Dynamics in Open Source Smart Contract Repositories: An Exploratory Data Analysis of Curated High Market Value Projects&rdquo;.</p> <p>In project_curation_selection, there is the curation process of the 100 selected projects including the identification of GitHub repositories and classification of evolution scenarios.&nbsp;</p> <p>In distribution_commits_issues_contributors_market_value_before_after_deploy, data collection from GitHub projects includes the distribution of total commits, contributors, and issues before and after deployment of each investigated project.&nbsp;</p> <p>In analysis_commit_messages, there is qualitative analysis of commit message content from all investigated projects.&nbsp;</p> <p>In the analysis_contributors section, the data focuses on analyzing the profiles of each GitHub contributor involved in the investigated projects.</p> <p>In analysis_market_value_by_project, data refers to the market value and volume of each investigated project.&nbsp;</p> <p>In codes, there are scripts used to obtain the analyzed data.</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
dryad32/100

Data from: Opening the door to the past: accessing phylogenetic, pathogen, and population data from museum curated bees

Open the record for dataset details and reuse information.

publicOct 2018View details →
dryad32/100

Data from: A curated and standardized adverse drug event resource to accelerate drug safety research

Open the record for dataset details and reuse information.

publicApr 2017View details →
dryad32/100

Data from: A reference set of curated biomedical data and metadata from clinical case reports

Open the record for dataset details and reuse information.

publicNov 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record