Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data files: Single-cell RNA sequencing of Plasmodium vivax sporozoites reveals stage- and species-specific transcriptomic signatures

<p>Scripts, preprocessed count matrices, single-cell data objects, and generated data (tables and .rds files)&nbsp;from the scRNA-seq analyses performed in&nbsp;<strong>&ldquo;Single-cell RNA sequencing of Plasmodium vivax sporozoites reveals stage- and species-specific transcriptomic signatures&quot;.</strong></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Data and analysis script to support "Multi-site analysis of sequence in leaf-out and flowering reveals evidence of local adaptation"

<p>Data files and R script used to download and analyze plant leaf-out and flowering observations maintained by the USA National Phenology Network to evaluate the consistency of leaf-out and flowering among species pairs over multiple years.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Data from : Genomic Sequence of Klebsiella pneumoniae IIEMP-3, a Vitamin B12-Producing Strain from Indonesian Tempeh

<p>Klebsiella pneumoniae&nbsp;strain IIEMP-3, isolated from Indonesian tempeh, is a vitamin B<sub>12</sub>-producing strain that exhibited a different genetic profile from pathogenic isolates. Here we report the draft genome sequence of strain IIEMP-3, which may provide insights on the nature of fermentation, nutrition, and immunological function of Indonesian tempeh.</p>

opencc-by-4.0Feb 2016View details →
zenodo40/100

Figure 2. A in Mitochondrial Dna Sequence Data Indicate Evidence For Multiple Species Within Peromyscus Maniculatus

Figure 2. A) Phylogenetic tree generated using Bayesian (MrBayes; Huelsenbeck and Ronquist 2001), maximum likelihood (RAxML; Version 8.1.17, Stamatakis 2006), and parsimony methods (PAUP* v. 4.0a165, Swofford 2002) and DNA sequence data from the mitochondrial cytochrome-b gene. The topology depicted is from the Bayesian analysis. Clade probability values (≥ 0.95) for the Bayesian analysis are indicated by an asterisk (*) and are to the left of the first slash, bootstrap values for the maximum likelihood analysis are shown between the two slashes, and bootstrap values obtained from the parsimony analysis are to the right of the last slash. Line at bottom of figure depicts the nucleotide substitution rate per site per million years. B) Same phylogenetic tree as depicted in Figure 2A except unsupported nodes (C, G, and H) were collapsed.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Figure 4. Approximate distributions and associated divergence times for A in Mitochondrial Dna Sequence Data Indicate Evidence For Multiple Species Within Peromyscus Maniculatus

Figure 4. Approximate distributions and associated divergence times for A) Peromyscus maniculatus-like ancestor; B) P. melanotis-like ancestor; C) P. gambelii/keeni/sejugis/sp.-like ancestor; D) P. polionotus-like ancestor; E) P. sonoriensis-like ancestor; F) P. labecula and P. maniculatus - like ancestor; G) P. keeni/sp.-like ancestor; and H) P. keeni-like, P. gambelii-like, P. sejugis-like, and P. sp.-like ancestors. Divergence times were estimated from the BEAST analysis (Version 2.4, Bouckaert et al. 2014) of the mitochondrial cytochrome-b gene dataset (see Fig. 3). Shading schemes that correspond to species distributions are shown in the inset.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Figure 1 in Mitochondrial Dna Sequence Data Indicate Evidence For Multiple Species Within Peromyscus Maniculatus

Figure 1. Distribution of selected populations and species of the Peromyscus maniculatus species group from Canada, Mexico, and the United States. Shaded areas represent distributions of taxa (defined in figure insert) as originally defined by Hall (1981) and modified based on the results of this study. Closed circles represent collecting localities listed in the Appendix; note that multiple individuals may be represented by a single closed circle. White boxes with black stars indicate type localities for each taxon and triangles indicate localities where haplotypes representing P. sonoriensis were found to be in sympatry with samples of P. gambelii and P. labecula, respectively.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Figure 3 in Mitochondrial Dna Sequence Data Indicate Evidence For Multiple Species Within Peromyscus Maniculatus

Figure 3. Time-calibrated ultrametric tree obtained from the BEAST analysis (Version 2.4, Bouckaert et al. 2014) of the mitochondrial cytochrome-b gene dataset. Scale bars at nodes represent the 95% highest posterior densities and numbers associated to each node are the estimated divergence times in million years ago.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Phertilizer: growing a clonal tree from single-cell DNA sequencing data of tumors

<p>The is the supplementary data repository for the simulation input data for&nbsp;Phertilizer: growing a clonal tree from single-cell DNA sequencing data of tumors.</p>

opencc-by-4.0Apr 2022View details →
dryad40/100

Supplementary data for: Comparison of optical flow derivation techniques for retrieving tropospheric winds from satellite image sequences

<p>This study introduces a validation technique for quantitative comparison of algorithms which retrieve winds from passive detection of cloud- and water vapor-drift motions, also known as Atmospheric Motion Vectors (AMVs).  The technique leverages airborne wind-profiling lidar data collected in tandem with 1-min refresh rate geostationary satellite imagery.  AMVs derived with different approaches are used with accompanying numerical weather prediction model data to estimate the full profiles of lidar-sampled winds which enables ranking of feature tracking, quality control, and height-assignment accuracy and encourages meso-scale, multi-layer, multi-band wind retrieval solutions.  The technique is used to compare the performance of two brightness motion, or "optical flow," retrieval algorithms used within AMVs, 1) Patch Matching (PM; used within operational AMVs) and 2) an advanced Variational Optical Flow (VOF) method enabled for most atmospheric motions by new-generation imagers.  The VOF AMVs produce more accurate wind retrievals than the PM method within the benchmark in all imager bands explored.  It is further shown that image regions with low texture and multi-layer-cloud scenes in visible and infrared bands are tracked significantly better with the VOF approach, implying VOF produces representative AMVs where PM typically breaks down.  It is also demonstrated that VOF AMVs have reduced accuracy where the brightness texture does not advect with the mean wind (e.g. gravity waves), where the image temporal noise exceeds the natural variability, and when the height-assignment is poor.  Finally, it is found that VOF AMVs have improved performance when using fine-temporal refresh rate imagery, such as 1-min versus 10-min data.</p>

opencc-zeroOct 2022View details →
zenodo40/100

Data from: Bratzel et al. (2022) Target-enrichment sequencing reveals for the first time a well-resolved phylogeny of the core Bromelioideae (Bromeliaceae). Taxon

<p>DNA sequence alignments used for phylogenetic analyses in Bratzel et al. (2022) Target-enrichment sequencing reveals for the first time a well-resolved phylogeny of the core Bromelioideae (Bromeliaceae). Taxon.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Data and code for "Differential methylation analysis of reduced representation bisulfite sequencing experiments using edgeR"

<p>This data set provides data files and R code to accompany the article <em>Differential methylation analysis of reduced representation bisulfite sequencing experiments using edgeR</em> published by F1000Research.</p> <p>The data consists of Reduced Representation BS-seq methylation profiles of epithelial populations from the mouse mammary gland, with n=2 biological replicates for each of three cell populations.</p> <p>RNA-seq expression profiles of luminal and basal mammary epithelial populations are also provided.</p> <p>The R code undertakes an differential methylation analysis of the BS-seq profiles and demonstrates a strong negative correlation between the differential methylation and differential expression results.</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

Targeted Gene Panel Sequencing Data of RELN paper - Meyer Children's Hospital IRCCS

<h3>Dataset description</h3> <p>&nbsp;</p> <p>The dataset has been prepared according the Minimal Information about a high throughput SEQuencing Experiment (MINSEQE) as reported in: <a href="https://doi.org/10.5281/zenodo.5706412">https://doi.org/10.5281/zenodo.5706412</a></p> <p>This dataset includes:</p> <ul> <li>The Targeted Gene Panel Sequencing Raw Data (FASTQ files) from two individuals harbouring RELN variants</li> <li>The &lsquo;final&rsquo; processed data,&nbsp;submitted both as VCF and TXT files, and obtained from the ANNOVAR annotations of the two patients</li> </ul> <p>The gene panel list used in the targeted capture and the essential experimental and data processing protocols has been reported in the RELN paper.</p> <h3>Identifiers</h3> <p>The 444D indentifier correspond to&nbsp;<strong>DN1 patient</strong> in the RELN paper.</p> <p>Tissue: peripheral blood sample</p> <p>Sex: female</p> <p>Age at sequencing: 21 years</p> <p>The 528T indentifier correspond to <strong>DN2 patient </strong>in the RELN paper.</p> <p>Tissue: peripheral blood sample</p> <p>Sex: female</p> <p>Age at sequencing: 1.5 years</p>

opencc-by-4.0May 2024View details →
zenodo40/100

FedscGen: privacy-aware federated batch effect correction of single-cell RNA sequencing data -- Preprocessed datasets

<div> <div> <div> <div> <p>This dataset accompanies the publication "FedscGen: Privacy-Aware Federated Batch Effect Correction of Single-Cell RNA Sequencing Data" and includes eight single-cell RNA sequencing (scRNA-seq) datasets used to benchmark the FedscGen and scGen methods. The datasets are provided in <code>.h5ad</code> format and include comprehensive metadata necessary for replication and further analysis.</p> <h3>Datasets</h3> <p>We analyze various datasets to compare FedscGen against scGen (centralized) in terms of batch correction. For simplicity, we refer to the dataset by abbreviations:</p> <ol> <li> <p><strong>Cell Line (CL)</strong>:</p> <ul> <li>Derived from the 293t_jurkat experiment with three batches: Zheng et al., 2017.</li> </ul> </li> <li> <p><strong>Human Dendritic Cells (HDC)</strong>:</p> <ul> <li>scRNA-seq data of human dendritic cells across two batches: Villani et al., 2017.</li> </ul> </li> <li> <p><strong>Human Pancreas (HP)</strong>:</p> <ul> <li>Consolidated data from five sources with 14,767 cells each: Baron et al., 2016; Muraro et al., 2016; Segerstolpe et al., 2016; Wang et al., 2016; Xin et al., 2016.</li> </ul> </li> <li> <p><strong>Mouse Brain (MB)</strong>:</p> <ul> <li>Merged datasets with 691,600 and 141,606 cells: Saunders et al., 2018; Rosenberg et al., 2018.</li> </ul> </li> <li> <p><strong>Mouse Cell Atlas (MCA)</strong>:</p> <ul> <li>Data focusing on 11 cell types from various organs: Han et al., 2018; The Tabula Muris Consortium, 2018.</li> </ul> </li> <li> <p><strong>Mouse Hematopoietic Stem and Progenitor Cells (MHSPC)</strong>:</p> <ul> <li>Data from SMART-seq2 and MARS-seq protocols: Nestorowa et al., 2016; Paul et al., 2015.</li> </ul> </li> <li> <p><strong>Mouse Retina (MR)</strong>:</p> <ul> <li>Data from two unassociated laboratories with 26,830 and 44,808 cells: Macosko et al., 2015; Shekhar et al., 2016.</li> </ul> </li> <li> <p><strong>PBMC (human Peripheral Blood Mononuclear Cell)</strong>:</p> <ul> <li>scRNA-seq data with two batches: Zheng et al., 2017.</li> </ul> </li> </ol> <p><strong>Usage Notes</strong>: Each dataset is provided in <code>.h5ad</code> format, compatible with common single-cell analysis tools such as Scanpy. Detailed metadata is included within each file.</p> <p><strong>Keywords</strong>: Single-cell RNA sequencing, scRNA-seq, Batch effect correction, Privacy-aware, Federated learning, scGen, FedscGen, Clinical multi-center studies, Genomics, Bioinformatics</p> <p><strong>Contact</strong>: For questions or further information, please contact Mohammad Bakhtiari at <a href="mailto:mohammad.bakhtiari@uni-hamburg.de.">mohammad.bakhtiari@uni-hamburg.de.</a></p> <p><strong>License</strong>: Creative Commons Attribution 4.0 International (CC BY 4.0)</p> </div> </div> </div> </div> <div> <div> <div>&nbsp;</div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Paired-end sequencing data toy example (measles)

<p>700,000 paired-end reads essentially Morbillivirus hominis to be used for testing purpose in the Sequana project</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Fig. 2 in Morphology and Sequence Data of Mexican Populations of the Ciliate Parasite of Marine Fishes Trichodina rectuncinata (Ciliophora: Trichodinidae)

Fig. 2. Photomicrographs of silver-impregnated adhesive discs and diagrammatic drawings of the denticles of respective morphotypes studied in the present paper; a and a'. From Enneanectes reticulatus, San Carlos, Sonora. b and b'. From Enneanectes reticulatus, San Carlos, Sonora. c and c'. From Tomicodon zebra, Zihuatanejo, Guerrero. d and d'. From Tomicodon zebra, Cuatunalco, Oaxaca.

opencc-by-4.0Dec 2018View details →
zenodo40/100

Fig. 3 in Morphology and Sequence Data of Mexican Populations of the Ciliate Parasite of Marine Fishes Trichodina rectuncinata (Ciliophora: Trichodinidae)

Fig. 3. Bayesian inference tree of sequences of the 18S gene of trichodinid species of the genus Trichodina and Trichodinella, emphasizing on Trichodina rectuncinata. Numbers near internal nodes show the support value. Codes: ♦ Cuatunalco; * Zihuatanejo; ● San Carlos.

opencc-by-4.0Dec 2018View details →
zenodo40/100

Fig. 1 in Morphology and Sequence Data of Mexican Populations of the Ciliate Parasite of Marine Fishes Trichodina rectuncinata (Ciliophora: Trichodinidae)

Fig. 1. Map showing the location of Mexico, and localities where populations of Trichodina rectuncinata were obtained.

opencc-by-4.0Dec 2018View details →
zenodo40/100

Protein haplotype sequences obtained by ProHap from the 1000 Genomes Project data set

<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the 1000 Genomes Project, aligned with the GRCh38 genome build (<a href="https://www.internationalgenome.org/data-portal/data-collection/grch38">https://www.internationalgenome.org/data-portal/data-collection/grch38</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts. The complete configuration file for each ProHap run is attached to this repository.</p> <p>This data set contains six compressed directories, five representing the superpopulations included in the 1000 Genomes Project (<a href="https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project">https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project</a>), and one created using all the samples included in the 1000 Genomes data set:</p> <ul> <li>AFR - African</li> <li>AMR - American</li> <li>EUR - European</li> <li>SAS - South Asian</li> <li>EAS - East Asian</li> <li>ALL - all participants in the 1000 Genomes Project</li> </ul> <p>Each of the directories contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap, using alleles with at least 1 % frequency within the selected population</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the simplified fasta file.&nbsp;</li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to&nbsp;<a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Va&scaron;&iacute;ček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families, ascertained to contain at least four members affected with lymphoid cancer.&nbsp; Please note that previous versions of this repository omitted a key file linking the genotypes of individuals to their family and individual IDs; this file, geno_key.txt, is now included. All other files remain the same as in previous versions.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>geno_key.txt &ndash; a plain-text file that links the genotyped individuals to their family and individual IDs.</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at <a href="../records/12694914">https://zenodo.org/records/12694914</a></p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →
dryad40/100

Data from: Sequence-based detection of emerging antigenically novel influenza A viruses

<p>The detection of evolutionary transitions in influenza A (H3N2) viruses' antigenicity is a major obstacle to effective vaccine design and development. In this study, we describe NIAViD, an unsupervised machine learning tool, adept at identifying these transitions, using HA1 sequence and associated physicochemical properties. NIAViD, performed with 88.9% (95% CI, 56.5%–98.0%) and 72.7% (95% CI,43.4%– 90.3%) sensitivity in training and validation respectively, outperforming the uncalibrated null model – 33.3% (95% CI,12.1%–64.6%) and does not require the need for potentially biased, time-consuming and costly laboratory assays. The pivotal role of Boman's index, indicative of the virus's cell surface binding potential, is underscored, enhancing the precision of detecting antigenic transitions. NIAViD's efficacy is not only in identifying influenza isolates that belong to novel antigenic clusters, but also in pinpointing potential sites driving significant antigenic changes, without the reliance on explicit modeling of hemagglutinin inhibition titers. Our approach holds immense promise to augment existing surveillance networks, offering timely insights for the development of updated, effective influenza vaccines. Consequently, NIAViD, in conjunction with other resources, could be used to support surveillance efforts and inform the development of updated influenza vaccines.</p>

opencc-zeroJul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record