Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
dryad32/100

Genotype likelihoods for low-coverage whole-genome sequencing data of yellow warblers

<p>The following datasets include the required input files used to empirically test population assignment in WGSassign on Yellow Warbler data. The file "yewa.known.ind105.ds_2x.beagle.gz" includes the filtered variants of 105 Yellow Warbler individuals output as genotype likelihoods and stored in a Beagle-formatted file. The ID file, "yewa.known.ind105.reference.IDs.txt", is a tab-delimited file with 2 columns, the first being the sample ID, and the second being the known reference population. The sample order in the ID file should match that of the input beagle file. To measure the assignment accuracy of WGSassign, we used leave-one-out cross validation using the input beagle file and our ID file.</p>

opencc-zeroJan 2024View details →
zenodo32/100

Supplementary material 1 from: Duan Y-B, Wang Y-J, Zhu D-H, Zeng Y, Wang X-D (2024) Description and mitochondrial genome sequencing of a new species of inquiline gall wasp, Synergus nanlingensis (Hymenoptera, Cynipidae, Synergini), from China. Journal of Hymenoptera Research 97: 105-126. https://doi.org/10.3897/jhr.97.119433

List of universal insect mitochondrial short fragments of the cox1, cob, rrnL and D2 genes primers used for long PCR primer developments

opencc-zeroMar 2024View details →
zenodo32/100

Figure 2 in Genomic survey sequencing and complete mitochondrial genome of the elkhorn coral crab Domecia acanthophora (Desbonne in Desbonne & Schramm, 1867) (Decapoda: Brachyura: Domeciidae)

Figure 2. Visualisation of assembled mitochondrial genome of Domecia acanthophora. Photo by Yun Scholten.

opennotspecifiedAug 2023View details →
zenodo32/100

Figure 1 in Genomic survey sequencing and complete mitochondrial genome of the elkhorn coral crab Domecia acanthophora (Desbonne in Desbonne & Schramm, 1867) (Decapoda: Brachyura: Domeciidae)

Figure 1. Repetitive elements in the genome of Domecia acanthophora. Each bar corresponds to a different type of annotated cluster. Numbers between parentheses in the legend represent the total number of annotated clusters in that category.

opennotspecifiedAug 2023View details →
zenodo32/100

Figure 4 in Genomic survey sequencing and complete mitochondrial genome of the elkhorn coral crab Domecia acanthophora (Desbonne in Desbonne & Schramm, 1867) (Decapoda: Brachyura: Domeciidae)

Figure 4. Mitochondrial gene order (MGO) of Domecia acanthophora compared to that of the brachyuran basic gene order, with the two transposition events highlighted.

opennotspecifiedAug 2023View details →
dryad32/100

Data from: Demographic inference from whole-genome and RAD sequencing data suggests alternating human impacts on goose populations since the last ice age

We investigated how population changes and fluctuations in the pink-footed goose might have been affected by climatic and anthropogenic factors. First, genomic data confirmed the existence of two separate populations: western (Iceland) and eastern (Svalbard/Denmark). Second, emographic inference suggests that the species survived the last glacial period as a single ancestral population with a low population size (100-1,000 individuals) that split into the current populations at the end of the Last Glacial Maximum with Iceland being the most plausible glacial refuge. While population changes during the last glaciation were clearly environmental, we hypothesize that more recent demographic changes are human-related: (1) the inferred population increase in the Neolithic is due to deforestation to establish new lands for agriculture, increasing available habitat for pink-footed geese (2) the decline inferred during the Middle Ages is due to human persecution and (3) improved protection explains the increasing demographic trends during the 20th century. Our results suggest both environmental (during glacial cycles) and anthropogenic effects (more recent) can be a threat to species survival.

opencc-zeroDec 2016View details →
zenodo32/100

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing

<p>These are the VCF files of structural variant (SV) calls for two sibling patients (II:2 [GMPB009_1] and II:3 [GMPB009_4]) generated by PacBio HiFi long-read genome sequencing.</p> <p>Sequence reads were processed using the <a href="https://github.com/PacificBiosciences/pb-human-wgs-workflow-snakemake">PacBio Human WGS workflow</a> with the human reference genome (hg38), and SVs were identified using '<a href="https://github.com/PacificBiosciences/svpack">svpack</a>'.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Sequence and functional analyses of native plasmids from plant pathogenic Gammaproteobacteria: comparative genomics, conjugative mobilization and fitness effects

<p>These data tables are part of the Supplementary Material for Chapter I of the thesis titled <em>"Sequence and Functional Analyses of Native Plasmids from Plant-Pathogenic Gammaproteobacteria: Comparative Genomics, Conjugative Mobilization, and Fitness Effects."</em></p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Draft Genome Sequences of 38 Aspergillus parasiticus isolates

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".

<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Whole genome re-sequencing workshop data: fastq files and reference genomes

<p>&nbsp;</p> <p>Whole genome re-sequencing data analysis workshop datasets. The files are necessary inputs for the workshop in https://github.com/PoODL-CES/Genomics_learning_workshop</p> <p>Tools and scripts listed in the https://github.com/PoODL-CES/Genomics_learning_workshop repository.</p> <p>&nbsp;</p> <p>These are subsampled fastq files from:<br><br>Khan, A., Patel, K., Shukla, H., Viswanathan, A., van der Valk, T., Borthakur, U., Nigam, P., Zachariah, A., Jhala, Y.V., Kardos, M. and Ramakrishnan, U., 2021. Genomic evidence for inbreeding depression and purging of deleterious genetic variation in Indian tigers. <em>Proceedings of the National Academy of Sciences</em>, <em>118</em>(49), p.e2023018118.</p> <p>The reference genome is from :</p> <p>Shukla, H., Suryamohan, K., Khan, A., Mohan, K., Perumal, R.C., Mathew, O.K., Menon, R., Dixon, M.D., Muraleedharan, M., Kuriakose, B. and Michael, S., 2023. Near-chromosomal de novo assembly of Bengal tiger genome reveals genetic hallmarks of apex predation. <em>GigaScience</em>, <em>12</em>, p.giac112.</p> <p>&nbsp;</p> <p>The reference has been indexed using:</p> <p>bwa index <a target="_blank" rel="noopener noreferrer">GCA_021130815.1_PanTigT.MC.v3_genomic.fna</a></p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

BactPrep: A user-friendly whole-genome sequencing analysis platform for the detection of homologous recombination and horizontal gene transfer in bacteria - Sample Dataset

<p>This is the dataset used as the sample dataset for the pipeline BactPrep. This&nbsp;dataset consists of 218&nbsp;<em>Streptococcus pneumoniae</em>&nbsp; PMEN1 WGS assemblies collected from the year 1984&nbsp;- 2008 from 22 unique countries globally. The raw sequencing data was originally published in the work:&nbsp;Rapid pneumococcal evolution in response to clinical interventions (doi: 10.1371/journal.ppat.1002745) under the bioproject&nbsp;PRJEB2085.</p> <p>We have assembled the raw sequences records with the following steps: 1)&nbsp;raw reads were&nbsp;first quality checked using fastQC 0.11.9;&nbsp;2) adapters and low quality reads were removed using Trimmomatic 0.39&nbsp;with parameter &ldquo;ILLUMINACLIP:TruSeq2-PE.fa:2:30:10:2:keepBothReads LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:36&rdquo;; 3)&nbsp;trimmed reads were error-corrected and assembled into WGS assemblies using SPAdes 3.15.0 with parameters &quot;--careful --mismatch-correction&rdquo;.</p>

opencc-by-4.0Oct 2021View details →
dryad32/100

Valenzuela phylogenomic dataset from: Illumina whole genome sequencing indicates ploidy level differences within the Valenzuela flavidus (Psocodea: Psocomorpha: Caeciliusidae) species complex

<p>This contains data for the manuscript: "Illumina Whole Genome Sequencing indicates Ploidy Level Differences within the <i>Valenzuela flavidus </i>(Psocodea: Psocomorpha: Caeciliusidae) Species Complex".</p> <p><i>Valenzuela flavidus</i> is a species of bark louse which is known to have asexual parthenogenetic populations in Europe but is believed to have sexual and asexual populations in North America as well. Historically, <i>Valenzuela aurantiacus</i> was the species epithet recognized for North American members until reports of asexual reproduction surfaced in certain North American populations. Cytogenetic studies have demonstrated European all-female populations are triploid. However, males are often reported in North America suggesting diploidy for sexual populations. With the use of Illumina whole genome sequencing, genetic diversity among North American and European populations was explored with phylogenomic methods. Ploidy level was estimated by examining allele frequencies of read-mapped homologous gene regions. Results indicate divergent populations between Europe and North America. North American populations containing males are estimated to be diploid suggesting a different mechanism of genomic reproduction. These results suggest divergent population structure among European asexual and North American sexual members of <i>V. flavidus</i> providing insight for future studies to understand patterns of asexuality reported within the complex.</p> <p>The following file contains all gene alignments, concatenated supermatrix, and mitochondrial alignment for this manuscript. In addition, the BAM files used to estimate allele frequencies. Also, gene trees for coalescent analysis, resultant treefiles from IQ-tree searches, and MCMCtree result.</p>

opencc-zeroNov 2021View details →
zenodo32/100

Supplementary File to "Both binding strength and evolutionary accessibility affect the population frequency of transcription factor binding sequences in Arabidopsis thaliana" (Genome Biology and Evolution)

<p>This data is supplementary file 1 of the following publication:</p> <p>Schweizer G, Wagner A. &quot;Both binding strength and evolutionary accessibility affect the population frequency of transcription factor binding sequences in Arabidopsis thaliana&quot; (Genome Biology and Evolution)</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of Ms4a3Ai14, BM chimeric mice (CD45.2 Csf2rb-/-: CD45.1 Csf2rb+/+ and CD45.2 Ifngr1-/-: CD45.1 Ifngr1+/+) using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.

<p><strong>Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE,&nbsp; BM chimeric mice (CD45.2 <em>Csf2rb</em><sup>-/-</sup>: CD45.1 <em>Csf2rb</em><sup>+/+</sup> and CD45.2 <em>Ifngr1<sup>-/-</sup></em>: CD45.1 <em>Ifngr1<sup>+/+</sup></em>) using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer&#39;s protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger&#39;s in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Single-cell RNA sequencing of Lymph node-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.

<p><strong>Single-cell RNA sequencing of Lymph node-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer&#39;s protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger&#39;s in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Single-cell RNA sequencing of Bone Marrow-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.

<p><strong>Single-cell RNA sequencing of Bone Marrow-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer&#39;s protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger&#39;s in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Single-cell RNA sequencing of Blood-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.

<p><strong>Single-cell RNA sequencing of Blood-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer&#39;s protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger&#39;s in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>

opencc-by-4.0Nov 2021View details →
dryad32/100

Mitochondrial genome sequencing of marine leukemias reveals cancer contagion between clam species in the Seas of Southern Europe

<p>Clonally transmissible cancers are tumour lineages that are transmitted between individuals via the transfer of living cancer cells. In marine bivalves, leukemia-like transmissible cancers, called hemic neoplasias, have demonstrated the ability to infect individuals from different species. We performed whole-genome sequencing in eight <i>V. verrucosa</i> clams that were diagnosed with hemic neoplasia, from two sampling points located more than 1,000 nautical miles away in the Atlantic Ocean and the Mediterranean Sea Coasts of Spain. Mitochondrial genome sequencing of tumour tissues from neoplastic animals revealed the coexistence of haplotypes from two different clam species. Phylogenies estimated from mitochondrial and nuclear markers confirmed this leukemia originated in <i>C. gallina </i>(or a closely related taxa) and was later transmitted to <i>V. verrucosa</i>, in which it survived as a contagious cancer. The analysis of mitochondrial and nuclear gene sequences supports all the studied tumours belonging to a single neoplastic <i>C. gallina </i>lineage that spread in the Seas of Southern Europe.</p>

opencc-zeroDec 2021View details →
dryad32/100

Long-read genome sequencing of bread wheat facilitates disease resistance gene cloning

<p>Cloning agronomically important genes from large, complex crop genomes remains challenging. Here, we generate a 14.7-gigabase chromosome-scale<i> </i>assembly of the South African bread wheat (<i>Triticum aestivum</i>) cultivar Kariega by combining high-fidelity long reads, optical mapping, and chromosome conformation capture. The resulting assembly is an order of magnitude more contiguous than previous wheat assemblies. Kariega shows durable resistance against the devastating fungal stripe rust disease. We identified the race-specific disease resistance gene <i>Yr27</i>, encoding an intracellular immune receptor, as a major contributor to this resistance. <i>Yr27</i> is allelic to the leaf rust resistance gene <i>Lr13,</i> with the Yr27 and Lr13 proteins sharing 97% sequence identity. Our results thus demonstrate the feasibility of generating chromosome-scale wheat assemblies to clone genes and also exemplify that highly similar alleles of a single-copy gene can confer resistance to different pathogens, which might provide a basis for engineering <i>Yr27</i> alleles with multiple recognition specificities in future.</p>

opencc-zeroDec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record