Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
RefAHL: A curated quorum sensing reference linking diverse LuxI-type signal synthases with their acyl-homoserine lactone products
Open the record for dataset details and reuse information.
Data for genetic characterization and curation of diploid a-genome wheat species
Open the record for dataset details and reuse information.
Dataset from: A curated DNA barcode reference library for parasitoids of northern European cyclically outbreaking geometrid moths
Open the record for dataset details and reuse information.
Data from: Curation: heat stress responses and population genetics of the kelp Laminaria digitata (Phaeophyceae) across latitudes reveal differentiation among North Atlantic populations
Open the record for dataset details and reuse information.
Aligned and curated mtDNA sequences from: Ancient DNA of narrow-headed voles reveals common features of the Late Pleistocene population dynamics in cold-adapted small mammals
Open the record for dataset details and reuse information.
A new taxonomist-curated reference library of DNA barcodes for Neotropical electric fishes (Teleostei: Gymnotiformes)
Open the record for dataset details and reuse information.
Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates
Open the record for dataset details and reuse information.
Assessment of hepatitis B viral lymphotrophism using deep curation
Open the record for dataset details and reuse information.
COVID-19 patient data from a study in Singapore curated for input into an in silico infection model
Open the record for dataset details and reuse information.
Taranaki Basin Curated Well Logs
<p>Machine learning (ML) models are being widely used in the geosciences for various tasks involving well log data, including prediction of missing well log curves, picking of stratigraphic surfaces, facies classification, and segmentation of different rock types. Even though various ML applications have been proposed in the literature, it is difficult to reproduce and advance the prior art without having access to the data and preprocessing steps used. In fact, there is an increasing need for benchmark cases to assess past and future solutions. The present dataset integrates well log data curated from the 2016 New Zealand Petroleum Exploration Public Data Pack.</p>
Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)
<p><span><span><span><span><span><span><span><span><span><span><span>Litter-decomposing Agaricales play key role in terrestrial carbon cycling, but little is known about their decomposition mechanisms. We assembled datasets of 42 gene families involved in plant-cell-wall decomposition from seven newly sequenced litter decomposers and 35 other Agaricomycotina members, mostly white-rot and brown-rot species. Using sequence similarity and phylogenetics, we split the families into phylogroups and compared their gene composition across nutritional strategies. Subsequently, we used Raman spectroscopy to examine the ability of litter decomposers, white-rot fungi, and brown-rot fungi to decompose crystalline cellulose. Both litter decomposers and white-rot fungi share the enzymatic cellulose decomposition, whereas brown-rot fungi possess a distinct mechanism that disrupts cellulose crystallinity. However, litter decomposers and white-rot fungi differ with respect to hemicellulose and lignin degradation phylogroups, suggesting adaptation of the former group to the litter environment. Litter decomposers show high phylogroup diversity, which is indicative of high functional versatility within the group, whereas a set of white-rot species shows adaptation to bulk-wood decomposition. In both groups, we detected species that have unique characteristics associated with hitherto unknown adaptations to diverse wood and litter substrates. Our results suggest that the terms white-rot fungi and litter decomposers mask a much larger functional diversity.</span></span></span></span></span></span></span></span></span></span></span></p>
Integrating QSAR models predicting acute contact toxicity and mode of action profiling in honey bees (A. mellifera): Data curation using open source databases, performance testing and validation
<p>This excel file (DOI: <a href="https://doi.org/10.5281/zenodo.3755675">https://doi.org/10.5281/zenodo.3755675</a>) provides the collection of raw data used for developing the first integrative Quantitative Structure-Activity Relationship (QSAR) model using EFSA's OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase i) to predict acute contact toxicity (LD<sub>50</sub>) and ii) to profile the Mode of Action (MoA) of pesticides active substances in honey bees (<em>Apis mellifera</em>)<em>. </em>Chemical identifiers (e.g. SMILES, CAS n., InChI) and acute contact toxicity data (LD<sub>50</sub>) on honey bees were used to develop and validate i) a two-category QSAR model (toxic/non-toxic; n=411) (sensitivity =0.93), specificity =0.85), balanced accuracy =0.90), Matthews correlation coefficient MCC=0.78), and ii) a regression-based model (n=113) (R2=0.74; MAE=0.52). Similarly, current study proposes the first MoA profiling for 113 pesticides active substances and the first harmonised MoA classification scheme for acute contact toxicity in honey bees, including LD<sub>50s</sub> data points from three different databases such as EFSA's OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase. Such classification allows to further define MoAs and the target site of Plant Protection Products (PPPs) active substances, thus enabling regulators and scientists to refine chemical grouping and toxicity extrapolations for single chemicals and component-based mixture risk assessment of multiple chemicals.</p> <p>The full data collection and analysis of QSAR models, toxicity data (LD<sub>50</sub>) and Mode of Action (Moa) data are described in Carnesecchi et al., 2020 (DOI: doi.org/10.1016/j.scitotenv.2020.139243).</p> <p>This work was supported by the European Food Safety Authority (EFSA) [contract number: OC/EFSA/SCER/2018/01 and NP/EFSA/AFSCO/2016/02 (Edoardo Carnesecchi)].</p>
Reactome COVID-19: the literature curation strategy
<p>ABSTRACT</p> <p>In response to the deluge of COVID-19-related publications Reactome developed a computational triaging strategy to review and identify publications appropriate for manual curation (66,100 SARS-Cov-2 articles on PUBMED, tallied on 30/October/2020 https://www.ncbi.nlm.nih.gov/research/coronavirus/; Chen et al., 2020). </p> <p>The literature triaging approach consisted of 4 main steps: 1) Literature screening; 2) Literature selection; 3) Reference tagging; and 4) Reference database construction. Two primary reference databases were downloaded and automatically text-mined: CDC COVID-19 downloadable database and bioRxiv database. Other collections of SARS-Cov-2 literature were manually screened with a focus on molecular interactions. These included: a Zotero Library built and updated by members of COVID-19 Disease Map (Ostaszewski et al., 2020); CORD-19 (Wang et al., 2020); LitCOVID (Chen et al., 2020); Johns Hopkins literature summary; Cell Press Coronavirus Resource Hub; Nature Coronavirus and COVID-19 updates; and Science’s Latest Coronavirus research. </p> <p>About 5% of the articles made it through reference screening focused on Reactome SARS-CoV-2 map construction to the literature selection step. If relevance was confirmed, the reference was then tagged in step 3. SARS-CoV-2 selected references were tagged regarding multiple features: (i) type of publication (e.g., article, review, pre-print, comment); (ii) Virus and host species (e.g., SARS-CoV-2, SARS-CoV-1; MERS, ACE2); (iii) Entity (specific molecules studied); (iv) Methods (e.g., Cryo-EM, ELISA, IC50); (v) Cell line and/or Tissue (e.g., vero-E6, lung tissue); (vi) subcellular localization (e.g., plasma membrane, ER); (vii) Molecular event (e.g., virus cycle step, pathway, host response); and (viii) Phenotype (e.g., immune, coagulation). Features i, ii, iii, vii and viii were mandatory. This Reference Database is stored in a shared spreadsheet, in which Reactome team members can edit and refine the Library (e.g., inclusion of tags). </p> <p> As Reactome is an evidence-based database built on reliable experimental published data, the process of SARS-CoV-2 reference selection is stringent and prioritizes peer-reviewed references. Nevertheless, the final decision on the reliability of the scientific evidence to support a molecular interaction was made by Reactome curators. In this case the focused literature triaging provides curators with a deeply researched trove of articles.</p> <p> The Reactome strategy of first creating an individual map for SARS-CoV-1 supported and guided the construction of a refined SARS-CoV-2 map. The SARS-CoV-2 Reactome map is an ongoing task, built on the foundation of a strong literature curation strategy. Together the literature triage and curatorial groups have built open-source COVID-19 viral infection pathways incorporating rapid scientific development and literature availability.</p> <p> </p> <p>References.</p> <p> </p> <p>Chen, Q., Allot A., Lu Z. Keep up with the latest coronavirus research. Nature 579, 193 (2020). doi: 10.1038/d41586-020-00694-1 </p> <p> </p> <p>Ostaszewski, M., Mazein, A., Gillespie, M.E. et al. COVID-19 Disease Map, building a computational repository of SARS-CoV-2 virus-host interaction mechanisms. Sci Data 7, 136 (2020).<a href="https://doi.org/10.1038/s41597-020-0477-8"> https://doi.org/10.1038/s41597-020-0477-8</a></p> <p> </p> <p>Wang, L., Lo K, Chandrasekhar, Y., et al. CORD-19: The Covid-19 Open Research Dataset. Preprint. ArXiv. 2020;arXiv:2004.10706v2. Published 2020 Apr 22.CORD-19.<a href="https://covidsearch.sinequa.com/app/covid-search/#/home"> https://covidsearch.sinequa.com/app/covid-search/#/home</a></p> <p> </p> <p>Reference Databases: CDC Covid-19 Database (<a href="https://www.cdc.gov/library/researchguides/2019novelcoronavirus/researcharticles.html">https://www.cdc.gov/library/researchguides/2019novelcoronavirus/researcharticles.html</a>); bioRxiv database (<a href="https://www.biorxiv.org/about-biorxiv">https://www.biorxiv.org/about-biorxiv</a>); Johns Hopkins literature summary (<a href="https://ncrc.jhsph.edu/topics/">https://ncrc.jhsph.edu/topics/</a>); Cell Press Coronavirus Resource Hub (<a href="https://www.cell.com/COVID-19">https://www.cell.com/COVID-19</a>); Nature Coronavirus and COVID-19 updates (<a href="https://www.nature.com/collections/aijdgieecb">https://www.nature.com/collections/aijdgieecb</a>); Science’s Latest Coronavirus research (<a href="https://www.sciencemag.org/collections/coronavirus?IntCmp=coronavirussiderail-128">https://www.sciencemag.org/collections/coronavirus?IntCmp=coronavirussiderail-128</a>).</p> <p> </p>
Data from: A curated and standardized adverse drug event resource to accelerate drug safety research
Identification of adverse drug reactions (ADRs) during the post-marketing phase is one of the most important goals of drug safety surveillance. Spontaneous reporting systems (SRS) data, which are the mainstay of traditional drug safety surveillance, are used for hypothesis generation and to validate the newer approaches. The publicly available US Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS) data requires substantial curation before they can be used appropriately, and applying different strategies for data cleaning and normalization can have material impact on analysis results. We provide a curated and standardized version of FAERS removing duplicate case records, applying standardized vocabularies with drug names mapped to RxNorm concepts and outcomes mapped to SNOMED-CT concepts, and pre-computed summary statistics about drug-outcome relationships for general consumption. This publicly available resource, along with the source code, will accelerate drug safety research by reducing the amount of time spent performing data management on the source FAERS reports, improving the quality of the underlying data, and enabling standardized analyses using common vocabularies.
Manually Curated Library of Transposable Elements (TEs) and TE Annotations for Drosophila amaguana
<p>This data collection provides a manually curated library of transposable elements (TEs) for <em>Drosophila amaguana</em>, including consensus sequences, genome-wide TE annotations, and individual TE copy sequences. The library was built using <em>de novo</em> generated by EDTA (Extensive <em>de novo</em> TE Annotator) and RepeatModeler, curated with MCHelper, and further processed for genome annotation using RepeatMasker and OneCodeToFindThemAll. </p> <p>Below is a description of the included files:</p> <ul> <li><strong><code>Dama_curated_TE_library.fasta</code>:</strong> Contains 737 consensus TE sequences manually curated for <em>D. amaguana</em>. Sequence identifiers include classification and origin (e.g., new families or similarity to known elements). The identifier for each sequence in the FASTA file includes information about its classification:<br><br> <ul> <li><strong>For sequences corresponding to potentially new families</strong>: The identifier consists of a three-letter abbreviation for <em>D. amaguana</em> (Dama), followed by the new family identifier and the superfamily name, all separated by underscores. <em>Example: </em>Dama_NF_BELPAO_1.</li> <li><strong>For consensus sequences that show similarity to TE sequences previously reported in other species</strong>: The identifier includes the abbreviation Dama, followed by the superfamily name and an abbreviation for the species in which the TE was previously reported, all separated by underscores. Example: Dama_Helitron-1_DVir.<br><br></li> </ul> </li> <li><strong><code>Dama_TE_annotations.out</code>:</strong> Genome-wide annotation file of TE insertions produced with RepeatMasker and post-processed using OneCodeToFindThemAll to merge fragmented elements.</li> <li><strong><code>Dama_TE_copies.fasta</code>: </strong>FASTA file containing the extracted sequences of all annotated TE copies from the <em>D. amaguana</em> genome.</li> <li><strong><code>TEcopies_sequences.sh</code>:</strong> Shell script used to extract TE copy sequences from the genome using the annotation coordinates.</li> <li><strong><code>Dynamics_Dama.ipynb</code>: </strong>Jupyter Notebook for the analysis of transposable element (TE) dynamics in<strong> </strong><em>D. amaguana.</em></li> </ul>
Curated Dataset for Analysis of the Impact of COVID-19 on Florida's Housing Market
Open the record for dataset details and reuse information.
FIGURE 5 in The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium)
FIGURE 5. The entomologists behind the collection of Megachilidae specimen records at RMCA (Tervuren, Belgium).
FIGURE 6 in The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium)
FIGURE 6. The diversity of entomologists behind the collection of megachilid specimen records at RMCA (Tervuren, Belgium). We list here the inventory number corresponding to each photograph, as well as the names of the photographers, for the images kindly provided by the RMCA archives department. 1. Mrs. Lebrun and Vrydagh (left) and Mrs. Corbisier and Staner (right) (AP.0.0.29742; photo by P. Staner 1930); 2. Reverend Father Hyacinthe Vanderyst in Kisantu (AP.0.0.23368; photo by "Jesuit Order", 1910-1925); 3. Reverend Father Hyacinthe Vanderyst (R.P. Vanderyst) in Ipamu (AP.0.2.368; photo by R.P. Delaere, 1922); 4. Frans Guillaume Overlaet (AP.0.0.25895; photo G. Fr. De Witte, 1931); 5. Maurice Bequaert (HP.1950.7.37, photographer not identified, 1950); 6. Professor Theodore D. Cockerell and the pygmies chief Kasulo, who lived in the forest near Tshibinda in the Sud-Kivu Province (photo by Alice Mackie); 7. Gaston-François De Witte (left) and Hans Joseph Brédo (right) (HP.2011.62.20-221; photo G. Fr. De Witte, 1947); 8. Prince Léopold III of Belgium (left) and Mr Ghesquière (right) in front of the laboratory in Stanleyville (now Kisangani) (AP.0.0.25895; photo Ghesquière, 1925); 9. Charles Seydel alias "Bwana Bilulu" (HP.1953.21.8, photo E. Devroey, 1922); 10. Charles Seydel alias "Bwana Bilulu" (HP.2011.62.13-226, photo GastonFrançois De Witte, 1931); 11. Charles Seydel alias "Bwana Bilulu" (HP. 2011.62.5-145; photo Gaston-François De Witte, 1925); 12. Sheffield Airey Neave; 13. Hans Joseph Brédo on the right, Director of the International Locust Laboratory in Abercorn (Northern Rhodesia), welcomes Belgian personalities on an official visit to the Laboratory (2017.24.30; photo Infocongo, 1952). All rights reserved for photographs 1, 2, 3, 5, 8 and 9; photographs 4, 7, 10, 11 and 13 under Creative Commons license (CC-BY 4.0); photograph 6 after Cockerell (1932) and 12 ©National Portrait Gallery.
FIGURE 3. Top 4 in The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium)
FIGURE 3. Top 4 of the Megachilidae species represented by the highest number of specimens in the RMCA collection (Tervuren, Belgium). A. Gronoceras cinctum (Fabricius, 1781) (female at nest entrance) (n=1,270 specimens); B. Euaspis abdominalis (Fabricius, 1793) (female on Duranta erecta (Verbenaceae)) (n=394 specimens); C. Megachile rufipes (Fabricius, 1781) (male on Antigonon leptopus (Polygonaceae)) (n=390 specimens); D. Megachile bituberculata Ritsema, 1880 (female on Pueraria javanica (Fabaceae)) (n=281). All photographs by NJ Vereecken.
FIGURE 2 in The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium)
FIGURE 2. Linear models between collection abundance and associated species richness in (i) the six historical provinces in DRC (top row), (ii) the 26 present-day provinces (middle row), and (iii) the 10 known phytogeographical districts (bottom row) defined within the country. Some administrative/natural units deviate from the overall attributable values. The Province Equateur deviates from the number of collections for species richness close to the mean.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.