Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
Survey on lattice data analysis, presentation, and curation practices
<p>This repository contains the results of a survey on software workflows and open science in lattice field theory conducted in 2022 by Andreas Athenodorou, Ed Bennett, Julian Lenz, and Elli Papadopolou. These data were collected using <a href="https://www.limesurvey.org/">LimeSurvey</a>, and were first presented in <a href="https://indico.hiskp.uni-bonn.de/event/40/contributions/695/">a talk at Lattice 2022 by Andreas Athenodorou</a>.</p> <p>The analysis is based on Julian Lenz's <a href="https://github.com/chillenzer/limesurvey-parser">LimeSurvey CSV parser</a>.</p> <p>The survey results are included in survey-results-redacted.csv. The survey structure is included in survey-structure.lss. Further details of the structure of the data, setup, see the included README.md file.</p>
Curated reference files for GCAP (WES)
<p>Provides big reference files or extra datasets/models for (in) GCAP project. </p> <p>The reference files are adapted from https://github.com/Wedge-lab/battenberg, more specifically, https://ora.ox.ac.uk/objects/uuid:08e24957-7e76-438a-bd38-66c48008cf52.</p> <p>News:</p> <ul> <li>Removed the correction files marked with 'update', which does not work for the latest version of ASCAT v3.</li> </ul>
S103 | NORMANUVCB | NORMAN Dataset of Curated UVCB Mappings
<p>This is the collection associated with list S103 | NORMANUVCB | NORMAN Dataset of Curated UVCB Mappings on the NORMAN Suspect List Exchange. <a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>A curated NORMAN Network effort to provide information for substances of Unknown or Variable Composition, Complex Reaction Products, or Biological Materials (UVCBs).</p> <p>v0.1.3: added BioHackathon UVCBs. v0.1.4: added EAWAGSURF. v0.1.5: added molecular formula to chemicals mapping for better table sorting. v0.1.6: added HESI UVCBs. v0.1.7: adjusted source display for annotation content of UVCBs. v0.1.8 added NIST polymers. v0.1.9 updated NIST polymers. </p>
A Simulated Heterozygous Diploid Genome for Third-gen Sequencing, Assembly, and Curation
<p>A simulated heterozygous diploid genome based on <em>Saccharomyces</em> <em>cerevisiae</em>, and <em>S. paradoxus</em> homologous chromosomes.</p> <p>Simulated PacBio subreads were generated from both parent haplomes and mixed together. A phased assembly was produced using FALCON assembler and FALCON Unzip (doi:10.1038/nmeth.4035). This dataset and assembly were then used to validate the Purge Haplotigs pipeline (https://bitbucket.org/mroachawri/purge_haplotigs). See workflow.sh for commands, comments and file descriptions.</p>
BacSPaD: A robust bacterial strains' pathogenicity resource based on integrated and curated genomic metadata
<p>The vast array of omics data in microbiology presents significant opportunities for studying bacterial pathogenesis and creating computational tools for predicting pathogenic potential. However, the field lacks a comprehensive, curated resource that catalogs bacterial strains and their ability to cause human infections. Current methods for identifying pathogenicity determinants often introduce biases and miss critical aspects of bacterial pathogenesis.<br>In response to this gap, we introduce BacSPaD (Bacterial Strains’ Pathogenicity Database), a thoroughly curated database focusing on pathogenicity annotations for a wide range of high-quality, complete bacterial genomes. Our rule-based annotation workflow combines metadata from trusted sources with automated keyword matching, extensive manual curation, and detailed literature review. Our analysis classified 5,502 genomes as pathogenic to humans (HP) and 490 as non-pathogenic to humans (NHP), encompassing 532 species, 193 genera, and 96 families. Statistical analysis demonstrated a significant but moderate correlation between virulence factors and HP classification, highlighting the complexity of bacterial pathogenicity and the need for ongoing research. This resource is poised to enhance our understanding of bacterial pathogenicity mechanisms and aid in the development of predictive models. To improve accessibility and provide key visualization statistics, we developed a user-friendly web interface, accessible at<a href="https://bacspad.altrabio.com/"> </a><a href="https://bacspad.altrabio.com/"><u>https://bacspad.altrabio.com</u></a>.</p>
Literature Curated PPIs from "Accurate and Sensitive Interactome Profiling Using a Quantitative Protein-Fragment Complementation Assay"
<p>This dataset includes the literature curated protein-protein interactions supporting the manuscript titled 'Accurate and Sensitive Interactome Profiling Using a Quantitative Protein-Fragment Complementation Assay'.</p>
RDF version of the data from Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zenodo Dataset] (2020)
<p>This is an RDFied version of the dataset published by Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zebodo Dataset] (2020)</p> <p>The original dataset publication DOI: <a href="http://doi.org/10.5281/zenodo.4146981">http://doi.org/10.5281/zenodo.4146981</a></p> <p>The Original publication authors: Saarimaki, Laura Aliisa, Federico, Antonio, Lynch, Iseult, Papadiamantis, Anastasios G., Tsoumanis, Andreas, Melagraki, Georgia, Afantitis, Antreas, Serra, Angela, & Greco, Dario</p>
Dataset - Papyrus 2024 - A large scale curated dataset aimed at bioactivity predictions
<p><strong>This update of release 2024.1 fixes the following:</strong></p> <ul> <li>Metadata in the columns <em>type_IC50</em>, <em>type_EC50</em>, <em>type_KD</em>, <em>type_Ki</em>, and <em>type_other</em> did not contain multiple values when multiple pChEMBL values where available but reported only a single value. This fix ensures all values are reported.</li> <li>Molecules were incorrectly standardized and mixtures were included in the dataset. Standardization (using the <a href="https://github.com/OlivierBeq/papyrus_structure_pipeline" target="_blank" rel="noopener">papyrus_structure_pipeline</a>) is now correctly enforced and mixtures have been removed.</li> </ul> <p><strong>Changes since version 05.6</strong></p> <ul> <li>ChEMBL data was updated to ChEMBL version 34</li> <li>data from the IUPHAR/BPS Guide to PHARMACOLOGY has been included</li> <li>data from Pickett et al.'s publication on MMP-12 has been included (<a href="https://doi.org/10.1021/ml100191f">ACS Med Chem Lett. 2011 Jan 13; 2(1): 28–33. DOI: 10.1021/ml100191f</a>)</li> </ul> <p><strong>Papyrus++:</strong></p> <p>Previous versions mistakenly considered a deviation of 2 log units around compound-target pairs to determine the reproducibility of assays (see published article for more details). This has been fixed to 0.5 log units to ensure data points fall within a maximum range of 1 log unit. As a result, the number of entries in the Papyrus++ set from this release has drastically reduced compared to previous releases.</p>
TF-Marker: A comprehensive manually curated database for transcription factors and related markers in specific cell and tissue types in human.
<p>Here, we developed the TF-Marker database (TF-Marker, http://bio.liclab.net/TF-Marker/) which is committed to a comprehensive manual curation of TFs and related markers with experimental evidence in specific cell and tissue types in human. Currently, through reviewing <strong>2,091</strong> published literature, we have manually classified TFs and related markers into five types according to their functions: 1) <strong>TF</strong>: TFs, which regulate the expression of markers; 2) <strong>T Marker</strong>: markers, which are regulated by TFs (TF and T Marker pairs can identify cell types more specifically); 3) <strong>I Marker</strong>: markers, which influence the activity of TFs (I Markers can also influence the development of specific cells and tissues); 4) <strong>TFMarker</strong>: TFs, which play roles as markers (TFMarkers are cell/tissue-specific TFs used as cell markers in biology experiments); and 5) <strong>TF Pmarker</strong>: TFs, which play roles as potential markers. By curating thousands of published literature, <strong>5,905</strong> entries including <strong>1,316</strong> TFs, <strong>1,092</strong> T Markers, <strong>473</strong> I Markers, <strong>1,600</strong> TFMarkers and <strong>1,424</strong> TF Pmarkers, were annotated in <strong>383</strong> cell types and <strong>95</strong> tissue types in human. Moreover, TF-Marker divided markers into disease markers and tissue/cell-specific markers. TF-Marker is an elaborate database, which provides TFs and related markers supported by experimental evidence. We believe TF-Marker will provide strong support for research into cell/tissue-specific TFs and related markers.</p>
Building a data curation pipeline for complex diseases: the case of Major Depression - Supplementary Material
<p>This entry contains the data generated by the study "Building a data curation pipeline for complex diseases: the case of Major Depression".</p>
A curated database of fungal pathogens and their host range
<p>This database contains a manually curated set of human, animal and plant pathogens, annotated with their confirmed host range and relevant sources. In addition to that, we include additional sets of plant-associated fungi (which may include non-pathogens), as well as fungi with an automatically assigned, putative human, animal or plant host. The labelled fungal species are linked to their representative GenBank genomes wherever possible. Genomes that were screened, but no label was found, are also included.</p> <p><strong>[Last update on: 11 Dec 2022]</strong><br> [Home page: <a href="https://dacs-hpi.gitlab.io/pathogenic-fungi/">https://dacs-hpi.gitlab.io/pathogenic-fungi/</a>]<br> <br> The database is stored in a flat-file format. All metadata are stored in all_data_[date].csv, and all_data_[date].rds contains the same data in a compressed format that can be easily loaded in R. The database was first compiled on 9 Oct 2021 (v1.0), and then updated on 2 Jan 2022 (v1.1) and 11 Dec 2022 (v1.2).</p> <p>The core database is limited to manually confirmed human, animal and plant pathogens with available genomes as of 9 Oct 2021. Those data are a subset of all_data, and are stored in core_fungal_pathogens.csv and core_fungal_pathogens.rds.</p> <p>The temporal-test subset contains confirmed pathogens with genomes added to GenBank between 9 Oct 2021 and 2 Jan 2022.</p> <p>You may also be interested in trained neural network models predicting pathogenic potentials of novel fungi from DNA sequences (<a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>) and simulated Illumina read sets used to train them (<a href="https://zenodo.org/record/5846397">https://zenodo.org/record/5846397</a>).<br> <br> See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a> and <strong>the paper</strong> presented at ECCB '22 and published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btac495">https://doi.org/10.1093/bioinformatics/btac495.</a></p>
Mask or Enhance: Data Curation Aiding the Discovery of Piezoresponse Force Microscopy Contributors
<p>This repository contains the data used in the corresponding study:</p> <p>Mask or Enhance: Data Curation Aiding the Discovery of Piezoresponse Force Microscopy Contributors</p> <p><strong>Abstract</strong></p> <p>Piezoresponse force microscopy (PFM) is routinely used to probe the nanoscale electromechanical response of ferroelectric and piezoelectric materials. However, many challenges remain in the interpretation of the recovered signal. Specifically, many non-ferroelectric contributions affect the measured response, ranging from electrostatics, to charge injection and trapping, and topographic cross-talk. Recently, machine learning (ML) has been utilized to identify multiple contributors within complex data systems, such as PFM response. A substantial advancement in ML approaches for PFM techniques is offered by dimensional stacking, enabling encoding of physical and/or chemical correlations within the materials’ response across different data dimensions spanning varying ranges. However, dimensional stacking requires appropriate scaling for each dimension (before ML analysis) to minimize undesired information loss. Here, the impact of clustering globally and locally scaled parameters in polarization switching experiments via resonant PFM (RPFM) are discussed. Specifically, dimensional stacking of scaled parameters can mask or enhance ferroelectric and non-ferroelectric behaviors, and aid identification of various physical phenomena contributing to the measured RPFM response. This study highlights the importance of data curation for ML, and its role in identifying signal contributors to scanning probe microscopy (SPM)-based techniques with multidimensional data, such as resonant and/or spectroscopic SPM.</p>
A curated data resource of 214K metagenomes for characterization of the global resistome
<p><strong>Data files of the curated resource of 214K metagenomes </strong> <strong>for characterization of the global resistome.</strong></p> <p>We have retrieved 214K metagenomic samples and now share the results here on Zenodo of our large-scale read mapping effort.</p> <p>There are five tables uploaded in three formats (TSV, HDF and MySQL dump):</p> <ul> <li>metadata.* : contains metadata for all sequencing runs.</li> <li>ARG.* : contain read alignment counts of antimicrobial resistance genes (ARGs).</li> <li>rRNA.* : contain read alignment counts of 16S/18S rRNA genes.<sup>1</sup></li> <li>diversity.* : contain diversity measures for ARGs and two taxonomic groups of rRNA genes (phylum, genus).</li> <li>ResFinder_anno.* : contain sequence information on the different ARGs, such as gene_lengths, resistance class, etc.</li> </ul> <p>Note that the HDF file rRNA.h5 is split into batches of 10,000 rows. To load it, the keys are in the format of "table_{i}", where i=0,1,2,..,4736</p> <p>Details on the different tables are available at https://hmmartiny.github.io/mARG/</p> <p>Additionaly, we have shared the data used to create the figures in the manuscript in the ZIP file named "figure_data.zip".</p> <p>Any further questions or issues, please contact H.-M. Martiny at hanmar@food.dtu.dk</p> <p> </p> <p><strong>Update log</strong>:</p> <p>* 2023-01-20: Update Diversity tables due to wrong total_fragments entered for ~250 run_accessions.</p>
Projet de curation de donnée sur le data stewardship
<p>Travail réalisé dans le cadre du cours de Master IS Data Curation. Rassemble les informations en lien avec le métier de data steward provenant de publications présentes sur Google Scholar. </p>
Curated GWAS summary statistics on African ancestry on 19 blood count traits and glycemic traits (hg38)
<p>Genome wide curated summary statistics on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>). GWAS of glycemic traits come from the <a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a> study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p> <p> </p>
Curated GWAS summary statistics on East Asian ancestry on 19 blood count traits and glycemic traits
<p>Genome wide curated summary statistics on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>). GWAS of glycemic traits come from the <a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a> study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p> <p>Full description of the method used to derive this dataset can be found in </p>
HawaiiCoast_GT: Curated AIS for Hawaii's coast correlated with ground truth incidents
<p>Because of the high-risk nature of emergencies and illegal activities at sea, it is critical that algorithms designed to detect anomalies from maritime traffic data be robust. However, there exist no publicly available maritime traffic datasets with real-world labelled anomalies. As a result, most anomaly detection algorithms for maritime traffic are validated without ground truth. We introduce the HawaiiCoast_GT dataset, the first ever publicly available automatic identification system dataset with a large corresponding set of true anomalous incidents. This dataset—cleaned and curated from Bureau of Ocean Energy Management (BOEM) and National Oceanic and Atmospheric Administration (NOAA) automatic identification system (AIS) data--covers Hawaii’s coastal waters for four years (2017-2020) and contains 88,749,176 AIS points for a total of 2,622 unique vessels. 208 tracks are labelled corresponding to 154 labelled real-world incidents. The codebase used to curate the original AIS data is being made openly available on GitHub.</p>
Curated dataset on protein's properties and post-translational modification protein properties
<p>Proteins perform essential cellular functions, which range from cell division and metabolism to DNA replication. Thus, decoding the mechanism of action of cells, requires understanding of the functioning and physicochemical properties of proteins [1]. While the genetic code encodes the primary structure of proteins, they undergo various modifications as part of their normal functioning including addition of modifying groups, such as acetyl, phosphoryl, glycosyl, and methyl, to one or more amino acids after translation, which is known as post-translational modification (PTM) [2, 3]. PTMs play an essential role in regulating protein functions by altering their physicochemical properties and understanding these reactions provides valuable insights regarding cell function. Advances in proteomics research have significantly deepened our understanding of PTMs and their impact on cellular functions and disease mechanisms. The study of PTMs is now at the forefront of research in molecular biology and biochemistry.</p> <p>Many databases, software, and tools have been developed to enhance our understanding of the various PTMs that affect human plasma proteins and help to simplify the analysis of complex PTM data [4]. These PTM databases and tools contain significant information and are a valuable resource for the research community. Key databases include dbPTM, UniProt, and PubChem. Utilising these databases, protein-related information like substrate peptides, amino acid sequence numbers, and experimentally validated PTM sites can be identified and curated.</p> <p>This dataset presents curated information regarding PTM-related changes in the physicochemical properties of the 16 most abundant plasma proteins [5], i.e., Serum Albumin, Serotransferrin, Antithrombin-III, Apolipoprotein A-I, Apolipoprotein A-IV, Apolipoprotein B-100, Apolipoprotein C-II, Apolipoprotein C-III, Apolipoprotein E, Clusterin, Complement C3, Haptoglobin, Histidine-rich glycoprotein, Mannose-binding protein C, Hemoglobin, and Fibrinogen alpha chain. The physicochemical properties studied, and the impact of different PTMs on the properties, include the protein molecular weight, isoelectric point, surface hydrophobicity, and solubility. The PTMs explored include phosphorylation, acetylation, glycosylation, methylation, ubiquitination, SUMOylation, lipidation, glutathionylation, nitrosylation, sulfoxidation, succinylation, neddylation, malonylation, hydroxylation, oxidation, and palmitoylation.</p> <p>References</p> <ol> <li>Alberts B, Johnson A, Lewis J, et al. Molecular Biology of the Cell. 4th edition. New York: Garland Science; 2002. Analyzing Protein Structure and Function.</li> <li>Chen, H.; Venkat, S.; McGuire, P.; Gan, Q.; Fan, C. Recent Development of Genetic Code Expansion for Posttranslational Modification Studies. Molecules 2018, 23, 1662.</li> <li>Marc Oeller, Ryan Kang, Hannah Bolt, Ana Gomes dos Santos, Annika Langborg Weinmann, Antonios Nikitidis, Pavol Zlatoidsky, Wu Su, Werngard Czechtizky, Leonardo De Maria,Pietro Sormanni, Michele Vendruscolo: Sequence-based prediction of the solubility of peptides containing non-natural amino acids [bioRiv].</li> <li>Ramazi S, Zahiri J. Posttranslational modifications in proteins: resources, tools and prediction methods. Database (Oxford). 2021 Apr 7;2021:baab012.</li> </ol>
Curated Research on Network Traffic Analysis
<p>With the NTA Database we aim to collect relevant information about the research in network traffic analysis conducted during the last years. To this end, we have curated related papers from journals and conferences and stored the extracted data in JSON files. </p>
ActDES – a Curated Actinobacterial Database for Evolutionary Studies
<p>ActDES constitutes a novel resource for the community of Actinobacterial researchers that will be useful primarily for two types of analyses: (i) comparative genomic studies - facilitated by reliable orthologs identification across a set of defined, phylogenetically representative genomes, and (ii) phylogenomic studies which will be improved by identification of gene subsets at specified taxonomic level. These studies can then act as a springboard for the study of the evolution of virulence genes, studying the evolution of metabolism and metabolic engineering target identification.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.