Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

138

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

138 results for “Exome sequencing”

Learn how ShareScore rates datasets ↗
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"

<p>These are the burden test results (summary statistics) for the publication:</p> <p>&quot;Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer&rsquo;s Disease&quot;,</p> <p>Nature Genetics, 2022.</p> <p>&nbsp;</p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on&nbsp;6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL&gt;=75, LOF+REVEL&gt;=50, LOF+REVEL&gt;=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Twist Whole-Exome Sequencing Dataset - High Coverage - WGGC SIG4 Benchmarking

<p>GIAB Reference Genome for Benchmarking Initiatives in the West German Genome Center (WGGC) - SIG4.&nbsp;</p> <p>Twist Whole-Exome Sequencing Dataset - High Coverage - 200M Reads.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)

<p>The data used in this tutorial are a subset of the data&nbsp;published previously in&nbsp;<a href="https://zenodo.org/record/3243160">Training material for the course &quot;Exome analysis with GALAXY&quot;</a>. Credit for uploading the original data goes to&nbsp;Paolo Uva and Gianmauro&nbsp;Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p>&nbsp;</p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you&#39;ll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load&nbsp;and, in fact, without having to install GEMINI&#39;s bundled annotation data.</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families, ascertained to contain at least four members affected with lymphoid cancer.&nbsp; Please note that previous versions of this repository omitted a key file linking the genotypes of individuals to their family and individual IDs; this file, geno_key.txt, is now included. All other files remain the same as in previous versions.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>geno_key.txt &ndash; a plain-text file that links the genotyped individuals to their family and individual IDs.</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at <a href="../records/12694914">https://zenodo.org/records/12694914</a></p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →
dryad40/100

Data from: Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra

Open the record for dataset details and reuse information.

publicDec 2018View details →
zenodo36/100

Training material for exome sequencing

<p>Exome sequencing means that all protein-coding genes in a genome are sequenced.</p> <p>In Humans, there are ~180,000 exons that makes up 1% of the human genome which&nbsp;contain ~30 million base pairs. Mutations in the exome have usually a higher&nbsp;impact and more severe consequences, than in the remaining 99% of the genome.</p> <p>With exome sequencing, one can identify genetic variation that is responsible&nbsp;for&nbsp;both Mendelian and common diseases without the high costs&nbsp;associated with&nbsp;whole-genome sequencing. Indeed, exome sequencing is the&nbsp;most efficient way to&nbsp;identify the genetic variants in all of an individual&#39;s genes.&nbsp;Exome sequencing&nbsp;is cheaper also than whole-genome sequencing.&nbsp;</p> <p>&nbsp;</p> <p>For training on exome sequencing data analysis, the Galaxy community proposes two tutorials (https://github.com/bgruening/training-material/tree/master/Exome-Seq). Here, you can find the needed datasets for these tutorials.</p>

opencc-by-4.0Aug 2016View details →
zenodo36/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains&nbsp;all the data files for a simulated exome-sequencing study of 150 families ascertained to contain at least four members affected with lymphoid cancer.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at&nbsp;<a href="https://zenodo.org/record/6505385">https://zenodo.org/record/6505385</a>&nbsp;.</p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →
zenodo36/100

Exome sequence analysis identifies rare coding variants associated with a machine learning-based marker for coronary artery disease.

<p>*.sh and *.R are codes to test rare coding variants for association with ISCAD.</p> <p>Petrazzini_etal_2024_*_level_meta_analysis.txt.gz are summary statistics of variant- and gene-level associations of rare coding variants in the exome sequences of 604,914 individuals with an in-silico score for coronary artery disease (ISCAD).</p> <p>Chromosomal positions are mapped to the GRCh38 (hg38) human genome reference.</p> <p>Directions of effect correspond to associations in the UK Biobank, the All of Us Research Program, the BioMe Biobank sample 1 and the BioMe Biobank sample 2, in that order.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Analysis results for association study of long-term kidney transplant rejection using whole-exome sequencing

<p>Association study results for long-term kidney transplant rejection. Single-variant association results are provided as Plink output files. Meta-analysis results are provided as METAL output files. FDR results are sorted lists of the top association result from random sample label permutations and are included with the plink and meta-analysis results. SKAT and GSEA results are provided for gene and pathway level analyses, respectively.</p>

opencc-by-4.0Oct 2018View details →
dryad36/100

Exome sequencing of a hybrid pine species complex on the Qinghai-Tibetan Plateau

<p>This study investigates the evolutionary history of <em>Pinus</em> <em>densata</em> on the Qinghai-Tibetan Plateau (QTP) and genomic heterogeneity across a zone of species transition to understand contemporary dynamics of selection and evolution of species barriers. We analyzed the genetic diversity in a range-wide collection of <em>P. densata</em> and representative populations of its progenitors <em>P. tabuliformis</em> and <em>P. yunnanensis</em> using 40,000 exome probe capture sequencing.</p>

opencc-zeroJan 2023View details →
ClinicalTrials.gov36/100

Enabling Personalized Medicine Through Exome Sequencing in the U.S. Air Force

ClinicalTrials.gov study NCT03276637. IPD Sharing: UNDECIDED. Countries: 1. Publications: 18.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

North Carolina Newborn Exome Sequencing for Universal Screening

ClinicalTrials.gov study NCT02826694. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

NCGENES: North Carolina Clinical Genomic Evaluation by NextGen Exome Sequencing

ClinicalTrials.gov study NCT01969370. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Clinical Utility of Prenatal Whole Exome Sequencing

ClinicalTrials.gov study NCT03482141. IPD Sharing: NO. Countries: 1. Publications: 6.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

North Carolina Genomic Evaluation by Next-generation Exome Sequencing, 2

ClinicalTrials.gov study NCT03548779. IPD Sharing: YES. Countries: 1. Publications: 61.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Clinical Utility of Pediatric Whole Exome Sequencing

ClinicalTrials.gov study NCT03525431. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
dryad36/100

Whole exome sequencing reveals a long-term decline in effective population size of red spruce (Picea rubens)

Open the record for dataset details and reuse information.

publicApr 2020View details →
dryad36/100

Exome sequencing of a hybrid pine species complex on the Qinghai-Tibetan Plateau

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad32/100

Data from: A high-density exome capture genotype-by-sequencing panel for forestry breeding in Pinus radiata

Development of genome-wide resources for application in genomic selection or genome-wide association studies, in the absences of full reference genomes, present a challenge to the forestry industry, where longer breeding cycles could benefit from the accelerated selection possible through marker-based breeding value predictions. In particular, large conifer megagenomes require a strategy to reduce complexity, whilst ensuring genome-wide coverage is achieved. Using a transcriptome-based reference template, we have successfully developed a high density exome capture genotype-by-sequencing panel for radiata pine (Pinus radiata D.Don), capable of capturing in excess of 80,000 single nucleotide polymorphism (SNP) markers with a minor allele frequency above 0.03 in the population tested. This represents approximately 29,000 gene models from a core set of 48,914 probes. A set of 704 SMP markers capable of pedigree reconstruction and differentiating individual genotypes were tested within two full-sib mapping populations. While as few as 70 markers could reconstruct parentage in almost all cases, the impact of missing genotypes was noticeable in several offspring. Therefore, sets of 60 sets of 110 randomly selected SNP markers were compared for both parentage reconstruction and clone differentiation. The performance in parentage reconstruction showed little variation over 60 iterations. However, there was notable variation in discriminatory power between closely related individuals, indicating a higher density SNP marker panel may be required to elucidate hidden relationships in complex pedigrees.

opencc-zeroOct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record