Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
SeuratExtend Tutorial: Curated Example Datasets for Single-Cell Analysis
<p>This repository contains example datasets specifically curated for the SeuratExtend tutorial, aimed at facilitating advanced analyses and visualization techniques in single-cell genomics. The datasets have been derived from publicly available data obtained from the 10X Genomics website and have undergone careful preprocessing to serve specific tutorial goals.</p> <p>The collection includes the following datasets:</p> <ol> <li> <p><strong>Myeloid Subset from PBMC 10k Dataset:</strong> This subset focuses on myeloid cells extracted from the larger PBMC 10k dataset, showcasing a preprocessed SeuratObject stored as an RDS file. The data serve as a primary example for demonstrating the capabilities of SeuratExtend differentiation trajectory analysis.</p> </li> <li> <p><strong>Velocyto LOOM File of Myeloid Subset from PBMC 10k Dataset:</strong> Accompanying the first dataset, this Velocyto-generated LOOM file represents a subset of the same myeloid cells, focusing on RNA velocity analyses. It provides a dynamic perspective on gene expression changes over time, enriching the tutorial with advanced single-cell transcriptomics insights.</p> </li> <li> <p><strong>SCENIC-Processed PBMC 3k Dataset:</strong> An outcome of running the SCENIC workflow on the PBMC 3k dataset, this LOOM file represents a refined dataset highlighting gene regulation networks. It serves as an advanced example for users interested in exploring gene regulatory mechanisms using SeuratExtend.</p> </li> </ol> <p>Each dataset has been subsetted and processed, making them ideal for users ranging from beginners to advanced researchers in the field of single-cell genomics. The provided data are intended for educational and tutorial purposes, allowing users to gain hands-on experience with real-world single-cell analysis scenarios.</p> <p> </p>
IRIS Carbon Mapping Project: Curated Dataset
<p>Dataset to support IRIS Carbon Mapping Project final report.</p>
Data for genetic characterization and curation of diploid a-genome wheat species
<p>Diploid A-genome relatives of wheat comprises <i>T</i>. <i>urartu</i>, <i>T</i>. <i>monococcum</i> subsp. <i>monococcum</i> (domesticated einkorn) and <i>T</i>. <i>monococcum</i> subsp. <i>aegilopoides</i> (wild einkorn). About 930 accessions of A-genome diploid wheat species preserved in the gene bank of the Wheat Genetics Resource Center (WGRC) at Kansas State University were genotyped using genotyping-by-sequencing (GBS). We constructed four pooled GBS libraries (384- and 288-plex) using restriction enzymes (Pst1-Msp1) combinations and the libraries were sequenced on the Illumina platform. The sequence data was processed using Tassel 5 GBS v2 pipeline and identified thousands of single nucleotide polymorphisms (SNPs) for downstream genetic and genomic dissections of the tested population. We have used <i>T</i>. <i>urartu</i> pseudomolecule (Tu2.0) as a reference genome to detect the genetic markers. Four fastq files with raw sequence reads information can be obtained at the National Center for Biotechnology Information (NCBI) SRA database with the BioProject accession PRJNA744683 (<a href="https://www.ncbi.nlm.nih.gov/sra/PRJNA744683" rel="noopener noreferrer">https://www.ncbi.nlm.nih.gov/sra/PRJNA744683</a>). Provided key file has information for demultiplexing including flowcell, lane number, barcodes and sample names to replicate the analysis. This experiment led us to curate the gene bank through identification of genetically duplicated accessions, miss-classified accessions and miss-classified non-diploid accessions. We were able to observe the unique genetic and evolutionary relationships among the diploid A-genome wheat species along with the unique population structures per species and sub-species. </p>
ariercole/UNITE-COVID: Version 3.1.0 release curation pipeline
<p>Release of ESICM UNITE-COVID data curation pipeline and data dictionary after main analysis completed.</p>
GeMo : a web-based platform for the visualization and curation of mosaic genomes
<p>Dataset that will be used by TraceAncestor or by VCFHunter.</p> <p><strong>Reference for dataset</strong></p> <ul> <li> <p><a href="https://doi.org/10.1093/molbev/msy199">Baurens,F.-C. et al.(2019) Recombination and Large Structural Variations Shape Interspecific Edible Bananas Genomes. Mol Biol Evol, 36, 97–111.</a></p> </li> <li> <p><a href="https://doi.org/10.1111/tpj.14683">Martin et al., 2020a. Martin G, Cardi C, Sarah G, Ricci S, Jenny C, Fondi E, Perrier X, Glaszmann J-C, D’Hont A, Yahiaoui N. 2020. Genome ancestry mosaics reveal multiple and cryptic contributors to cultivated banana. Plant J. 102:1008–1025.</a></p> </li> <li> <p><a href="https://doi.org/10.1093/aob/mcz029">Ahmed,D. et al. (2019) Genotyping by sequencing can reveal the complex mosaic genomes in gene pools resulting from reticulate evolution: a case study in diploid and polyploid citrus. Annals of Botany, 123, 1231–1251.</a></p> </li> </ul>
Assessment of hepatitis B viral lymphotrophism using deep curation
<p><strong><span>Background</span></strong><span><strong>:</strong> The replicative forms of the hepatitis B virus (HBV) is found in several types of white blood cells within the host defense system. To determine the dimensionality of the extrahepatic manifestation of HBV in host white blood cells, it is important to understand the complete biology of its pathogenesis and lymphotropic nature.</span></p> <p><strong><span>Methods</span></strong><span><strong>: </strong>Deep curation of the literature from the PubMed database pertaining to the HBV manifestation in the human host white blood cells was conducted and then manually filtered to determine the behavioral trend of the virus within the human white blood cells.</span></p> <p><strong><span>Results</span></strong><span><strong>:</strong> The curation of 198 research articles identified 28 genes, 92 proteins, and 20 Peripheral Blood Mononuclear cells involved in HBV pathogenesis, while 20 immune cells were found to be permissive for the viral penetration and replication. The presence of the replicative forms of HBV in the host immune cells led to the further elucidation of 28 genes and 92 proteins that interact with one or more viral genes and proteins.</span></p> <p><strong><span>Conclusions</span></strong><span><strong>:</strong> A multi-dimensional analysis using deep curation identified a possible lymphotropic character of HBV. Moreover, there are certain pathways that could aid in the propagation of viral infection by using immune cells to its advantage. Thus, instead of eliminating HBV, the immune system may contribute to the population expansion of the virus.</span></p>
Curated data-set of crystal-structure prototypes of binary and ternary sp-d valent compounds
<p>Curated collection of binary and ternary compounds of sp-valent elements and d-valent elements and their crystal-structure prototype. The data set includes binary and pseudo-binary compounds in binary prototypes (BinaryPrototype-BinaryCompound.csv), offstoichoimetric and ternary compounds in binary prototypes (BinaryPrototype-BinaryOffstoichiometricAndTernaryCompound.csv) and ternary compounds in ternary prototypes (TernaryPrototype-TernaryCompound.csv). The first column in each data set corresponds to the crystal-structure prototype, the second column to the chemical composition. Ternary compositions labelled A-B+C indicate that element B and element C occupy the same sublattice of the crystal structure. A+B-C correspondingly indicates that element A and element B occupy the same sublattice. These data sets were used to construct structure maps for predicting the crystal structure of a compound from only its chemical composition. For details on curation, further discussions and structure maps see original publications (Chem. Mater. 28, 2550−2556, 2016 and Modelling Simul. Mater. Sci. Eng. 25, 074002, 2017).</p>
A new taxonomist-curated reference library of DNA barcodes for Neotropical electric fishes (Teleostei: Gymnotiformes)
<p>DNA barcoding is a useful tool for identifying species; however, successful barcode-based identification requires a reference library of barcode sequences from accurately identified specimens. Here we present a reference library of <em>co1</em> barcode sequences for the Neotropical electric knifefish order Gymnotiformes (Teleostei: Ostariophysi), a model taxon for studies of tropical diversification and biogeography, genomics, behaviour, and neurobiology. Our library contains barcodes for 167 of the ca. 270 valid species of gymnotiforms derived from geo-referenced museum voucher specimens, and includes sequences from 26 type specimens and 21 specimens from type localities, most of which we collected. To assess the state of gymnotiform barcodes in two main public barcode repositories, GenBank and BOLD, we compared the barcodes in these databases to our reference library. Our analysis shows that a considerable proportion of gymnotiform barcodes in GenBank and BOLD are mis- or unidentified. We encourage taxonomists to develop and publish barcode reference libraries composed of carefully curated barcode sequences.</p>
ARF / YA16Sdb collection of curated 16S rRNA alleles
<p>A collection of 16s rRNA alleles filtered by read quality and annotation quality.</p> <p>This repository contains 376,934 full-length 16s rRNA alleles with validated (by majority-rules) taxonomic annotations. These are a subset of 727,361 full-length 16s rRNA alleles.</p> <p>The pipeline used to create this repo is also available at: github.com/jgolob/arf</p>
Curation of blade parts
<p><strong>Curation of the blade parts at the Swiderian occupation in Lubrza 10, Poland </strong></p> <p> </p> <p>The data is used in conference presentation A. Diachenko and I. Sobkowiak-Tabaka, "Smallest-scale movement: Tracing the intra-site relocation of hunter-gatherers " (ArcheoFOSS 2022)</p>
Curated Phage Database (CPD) fasta file
<p>This is the fasta file including phage genomes used to generate the Curated Phage Database (CPD) utilized in our in review manuscript "The circulating phageome reflects bacterial infections". The corresponding phage characteristic data will be present in the manuscript as a supplemental file, and can be used to connect a Genbank ID to identified bacteriophage host and phage taxonomic information if known.</p> <p>Please note that this database is built from phage sequences in the NCBI nucleotide repository. Due to field bias towards sequencing human disease-related bacteria and their phage, this database is reflective of this bias and is most representative of bacteriophage associated with human pathogens and as such underrepresents environmental phages in comparison - a limitation to keep in mind when utilizing to interpret potential phage sequences.</p>
InSAR Interferograms and Products from 2021, M7.2 Nippes, Haiti earthquake (Curated Dataset)
<p>This dataset contains curated interferograms and derived products from the August 14, 2021 M7.2 Nippes, Haiti earthquake and accompanies the BSSA publication: <a href="https://doi.org/10.1785/0120220109" target="_blank" rel="noopener">https://doi.org/10.1785/0120220109</a>. Data products are also available via https://topex.ucsd.edu/haiti_7.2/index.html</p> <p>The data is organized in directories by satellite, relative orbit, and pair:</p> <p>├── ALOS1_A138<br>│ └── 20100116_20100603<br>├── ALOS2_A042<br>│ ├── 20210101_20210827<br>│ ├── 20210101_20211231<br>│ ├── 20210827_20211231<br>│ └── phasegrad<br>├── ALOS2_A043<br>│ ├── 20201223_20210818<br>│ └── phasegrad<br>├── ALOS2_D138<br>│ └── 20191210_20210817<br>├── S1_A004<br>│ ├── 20210805_20210817<br>│ ├── 20210817_20210823<br>│ ├── 20210823_20210829<br>│ └── 20210829_20210904<br>└── S1_D142<br> ├── 20210803_20210815<br> ├── 20210815_20210821<br> ├── 20210821_20210827<br> └── 20210827_20210902</p> <p>Each pair directory contains the relevant .grd files and HDF5 file, named according to the standard:</p> <p><SAT>_<SW>_<RELORB>_<FRAME>_<DATE1>-<DATE2>_<TBASE>_<BPERP>.h5</p> <p>We also include phase gradient .grd files, calculated in the range (xphase_mask_ll.grd) and azimuth (yphase_mask_ll.grd) directions.</p> <p>Please contact Zoe Yin (hyin@ucsd.edu) or Jennifer S. Haase (jhaase@ucsd.edu) with any questions. </p>
Curated dataset of bug fix commits from "An Empirical Study on Real Bug Fixes"
<p>To cite it:</p> <p><code>@misc{bfdataset,<br> author = {Martin Monperrus},<br> title = {Curated dataset of bug fix commits from "An Empirical Study on Real Bug Fixes"},<br> year = 2017,<br> doi = {10.5281/zenodo.1004734},<br> url = {https://doi.org/10.5281/zenodo.1004734}<br> }</code></p> <pre> </pre> <p> </p>
Database of 16S sequences from SILVA (r114), filtered, curated and annotated to be used easily by programs of taxonomic assignments
<p>The database used for the taxonomic assignment of reads generally comes from the SILVA database (http://www.arb-silva.de/). The logic behind this database is to use the information from the best one to the worst one. This is why the curated database was splitted in two parts : the [C] sequences for Complete sequences in terms of taxonomy, and the [I] and [E] sequences, for Incomplete and Environmental sequences.</p> <p>Each sequence included into the database must have a specific format summarizing all needed information (example below):<br> >[I]AACY020336309;Archaea(superkingdom);Euryarchaeota(phylum);Thermoplasmata(class);Thermoplasmatales(order);Marine_Group_II(no_rank);;marine_metagenome</p> <p>This sequence is an incomplete one ([I]), with a specific accession number from NCBI or SILVA, or another database (AACY020336309). Then, all taxonomic data is separated using ';' characters, for each considered level (superkingdom, phylum, <br> class, order, family, and genus). The species name is the last one and separated by two ';' characters from the rest of the descriptive line. Finally, the descriptive line must not contain specific characters like spaces. If one or several levels are unknown, this is indicated by 'no_rank'.</p> <p>Another example here for [C] sequences:<br> >[C]AAAK03000010;Bacteria(superkingdom);Firmicutes(phylum);Bacilli(class);Lactobacillales(order);Enterococcaceae(family);Enterococcus(genus);;Enterococcus_faecium_DO<br> This sequence is a complete one ([C]), with a specific accession number from NCBI or SILVA, or another database (AACY020187844). Then, all taxonomic data is separated using ';' characters, for each considered level (superkingdom, phylum, <br> class, order, family, and genus). The species is the last one and separated by two ';' characters from the rest of the descriptive line. Complete sequences must have six levels of information (superkingdom, phylum, class, order, family, and genus). If it is not the case, the sequence will be considered as Incomplete ([I]) (between three and five levels), or Environmental ([E]) (with only the superkingdom and the phylum levels).</p> <p>Another example here for [E] sequences:<br> >[E]U59968;Archaea(superkingdom);Thaumarchaeota(phylum);Soil_Crenarchaeotic_Group(SCG)(no_rank);;uncultured_crenarchaeote<br> This sequence is a environmental one ([E]), with a specific accession number from NCBI or SILVA, or another database (U59968). Then, all taxonomic data is separated using ';' characters, for each considered level (superkingdom, phylum, class, order, family, and genus). The species is the last one and separated by two ';' characters from the rest of the descriptive line. Complete sequences <br> must have six levels of information (superkingdom, phylum, class, order, family, and genus). If it is not the case, the sequence will be considered as Incomplete ([I]) (between three and five levels), or Environmental ([E]) (with only the superkingdom and the phylum levels).</p> <p>More details on the steps defined to clean and define this new database can be available on demand (sebastien.terrat@inra.fr).</p>
BOLD curated data and reference library
<ul> <li>BOLD_Public.05-Apr-2024.BAGS.d39fc06c483ea4a7252b6a4e203ca15c6ce4dd24.tsv - BAGS ranking of BIN x species combinations</li> <li>BOLD_Public.05-Apr-2024.3c74c5e91daa2ddccce41d38881edc2e242e9c81.tsv.gz - BCDM rated with curation pipeline</li> <li>refdb.fa - COI-5P records with BAGS=(A,B,C), rating=(1,2,3) with <a href="https://github.com/naturalis/galaxy-tool-BLAST">galaxy-tool-BLAST</a>-compliant full lineage</li> <li>families-for-curation.tar.gz - TSVs of families with BCDM records for manual curation</li> </ul>
User Generated Content (EOL v2): User Added Text, curated
For questions or use cases calling for large, multi-use aggregate data files, please visit the EOL Services forum at <p></p>http://discuss.eol.org/c/eol-services
User Generated Content (EOL v2): curation of media objects
trust, untrust, hide, set as exemplar <p></p>https://eol-jira.bibalex.org/browse/DATA-1731 Oct 24, 2018 For questions or use cases calling for large, multi-use aggregate data files, please visit the EOL Services forum at <p></p>http://discuss.eol.org/c/eol-services
Chemical Occurence Data Curation Results with CleanGeoStreamR
<p>This data set includes the curated chemical occurrence data from the NORMAN database. It is related with the surface waters. <strong>CleanGeoStreamR</strong> R package was used for data curation. <br><br>More details about the applied methods and the development of <strong>CleanGeoStreamR</strong> can be found in the following scientific paper: <a href="https://doi.org/10.1016/j.ecoinf.2025.103038"><em>Automated Curation of Spatial Metadata in Environmental Monitoring Data</em></a> (DOI: 10.1016/j.ecoinf.2025.103038).</p>
Curated data for phage-bacteria hosts for BioML Hackathon
<p>This is the accompanying raw data for this github: https://github.com/havu73/hackathonBio</p> <p>A brief summary of the datasets:</p> <p>For every dataset we have genome sequences for hosts and phages and all positive pairs between phages and hosts (majority of the possible pairs are negative, so we do not save them explicitly). </p> <ul> <li><strong>E.coli: </strong>325 bacterial hosts and 96 phages<strong> </strong>(source study link: https://doi.org/10.1101/2023.11.22.567924)</li> <li><strong>Vibrio: </strong>259 bacterial hosts and 239 phages (source study link: https://doi.org/10.1038/s41467-021-27583-z)</li> <li><strong>Klebsiella:</strong> 149 bacterial hosts and 115 phages (source study link: https://doi.org/10.1038/s41467-024-48675-6)</li> <li><strong>phageDB: </strong>74 bacterial hosts and 4766 phages<strong> </strong>(source link: https://phagesdb.org/)</li> <li><strong>phageScope: </strong>180 bacterial hosts and 4434 phages (source link: https://phagescope.deepomics.org/database)</li> </ul>
MFCC Chroma features from "A New Curated Corpus of Historical Electronic Music"
<p>Audio Feature data to accompany the paper </p> <p>(2018) Nick Collins, Peter Manning and Simone Tarsitani "A New Curated Corpus of Historical Electronic Music: Collation, Data and Research Findings" Transactions of the International Society for Music Information Retrieval 1(1): 34-43</p> <p>https://transactions.ismir.net/articles/10.5334/tismir.5</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.