Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
138
datasets available to search
ShareScore release 0.9.0
Dataset results
138 results for “Exome sequencing”
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"
<p>These are the burden test results (summary statistics) for the publication:</p> <p>"Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer’s Disease",</p> <p>Nature Genetics, 2022.</p> <p> </p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on 6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL>=75, LOF+REVEL>=50, LOF+REVEL>=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p> </p>
Twist Whole-Exome Sequencing Dataset - High Coverage - WGGC SIG4 Benchmarking
<p>GIAB Reference Genome for Benchmarking Initiatives in the West German Genome Center (WGGC) - SIG4. </p> <p>Twist Whole-Exome Sequencing Dataset - High Coverage - 200M Reads.</p> <p> </p> <p> </p>
Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)
<p>The data used in this tutorial are a subset of the data published previously in <a href="https://zenodo.org/record/3243160">Training material for the course "Exome analysis with GALAXY"</a>. Credit for uploading the original data goes to Paolo Uva and Gianmauro Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p> </p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you'll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load and, in fact, without having to install GEMINI's bundled annotation data.</p>
Simulated exome-sequencing data for a family study of lymphoid cancer
<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families, ascertained to contain at least four members affected with lymphoid cancer. Please note that previous versions of this repository omitted a key file linking the genotypes of individuals to their family and individual IDs; this file, geno_key.txt, is now included. All other files remain the same as in previous versions.</p> <p>The simulated data can be found in the files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the SLiM-simulated, exome-wide, SNV data generated under an American-admixture demographic model, for the American-admixed sub-population only.</li> <li>SLiM_output_chr8&9.txt - contains the SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but only for chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained pedigrees.</li> <li>Genotypes.zip - a zipfile that contains 22 text files of genotypes for each chromosome. The genotypes are for simulated single-nucleotide variants on the exome and are in gene-dosage format. </li> <li>geno_key.txt – a plain-text file that links the genotyped individuals to their family and individual IDs.</li> <li>SNVmaps.zip - a zipfile that contains 22 text files giving the single-nucleotide variant information for each chromosome. </li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip - a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data can be found in the GitHub repository archived at <a href="../records/12694914">https://zenodo.org/records/12694914</a></p> <p>We have also uploaded one intermediate .Rdata file, Chromwide.Rdata, to save the user substantial time when running the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>
Data from: Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra
Open the record for dataset details and reuse information.
Training material for exome sequencing
<p>Exome sequencing means that all protein-coding genes in a genome are sequenced.</p> <p>In Humans, there are ~180,000 exons that makes up 1% of the human genome which contain ~30 million base pairs. Mutations in the exome have usually a higher impact and more severe consequences, than in the remaining 99% of the genome.</p> <p>With exome sequencing, one can identify genetic variation that is responsible for both Mendelian and common diseases without the high costs associated with whole-genome sequencing. Indeed, exome sequencing is the most efficient way to identify the genetic variants in all of an individual's genes. Exome sequencing is cheaper also than whole-genome sequencing. </p> <p> </p> <p>For training on exome sequencing data analysis, the Galaxy community proposes two tutorials (https://github.com/bgruening/training-material/tree/master/Exome-Seq). Here, you can find the needed datasets for these tutorials.</p>
Simulated exome-sequencing data for a family study of lymphoid cancer
<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families ascertained to contain at least four members affected with lymphoid cancer.</p> <p>The simulated data can be found in the files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the SLiM-simulated, exome-wide, SNV data generated under an American-admixture demographic model, for the American-admixed sub-population only.</li> <li>SLiM_output_chr8&9.txt - contains the SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but only for chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained pedigrees.</li> <li>Genotypes.zip - a zipfile that contains 22 text files of genotypes for each chromosome. The genotypes are for simulated single-nucleotide variants on the exome and are in gene-dosage format. </li> <li>SNVmaps.zip - a zipfile that contains 22 text files giving the single-nucleotide variant information for each chromosome. </li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip - a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data can be found in the GitHub repository archived at <a href="https://zenodo.org/record/6505385">https://zenodo.org/record/6505385</a> .</p> <p>We have also uploaded one intermediate .Rdata file, Chromwide.Rdata, to save the user substantial time when running the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>
Exome sequence analysis identifies rare coding variants associated with a machine learning-based marker for coronary artery disease.
<p>*.sh and *.R are codes to test rare coding variants for association with ISCAD.</p> <p>Petrazzini_etal_2024_*_level_meta_analysis.txt.gz are summary statistics of variant- and gene-level associations of rare coding variants in the exome sequences of 604,914 individuals with an in-silico score for coronary artery disease (ISCAD).</p> <p>Chromosomal positions are mapped to the GRCh38 (hg38) human genome reference.</p> <p>Directions of effect correspond to associations in the UK Biobank, the All of Us Research Program, the BioMe Biobank sample 1 and the BioMe Biobank sample 2, in that order.</p>
Analysis results for association study of long-term kidney transplant rejection using whole-exome sequencing
<p>Association study results for long-term kidney transplant rejection. Single-variant association results are provided as Plink output files. Meta-analysis results are provided as METAL output files. FDR results are sorted lists of the top association result from random sample label permutations and are included with the plink and meta-analysis results. SKAT and GSEA results are provided for gene and pathway level analyses, respectively.</p>
Exome sequencing of a hybrid pine species complex on the Qinghai-Tibetan Plateau
<p>This study investigates the evolutionary history of <em>Pinus</em> <em>densata</em> on the Qinghai-Tibetan Plateau (QTP) and genomic heterogeneity across a zone of species transition to understand contemporary dynamics of selection and evolution of species barriers. We analyzed the genetic diversity in a range-wide collection of <em>P. densata</em> and representative populations of its progenitors <em>P. tabuliformis</em> and <em>P. yunnanensis</em> using 40,000 exome probe capture sequencing.</p>
Enabling Personalized Medicine Through Exome Sequencing in the U.S. Air Force
ClinicalTrials.gov study NCT03276637. IPD Sharing: UNDECIDED. Countries: 1. Publications: 18.
North Carolina Newborn Exome Sequencing for Universal Screening
ClinicalTrials.gov study NCT02826694. IPD Sharing: YES. Countries: 1. Publications: 1.
NCGENES: North Carolina Clinical Genomic Evaluation by NextGen Exome Sequencing
ClinicalTrials.gov study NCT01969370. IPD Sharing: YES. Countries: 1. Publications: 1.
Clinical Utility of Prenatal Whole Exome Sequencing
ClinicalTrials.gov study NCT03482141. IPD Sharing: NO. Countries: 1. Publications: 6.
North Carolina Genomic Evaluation by Next-generation Exome Sequencing, 2
ClinicalTrials.gov study NCT03548779. IPD Sharing: YES. Countries: 1. Publications: 61.
Clinical Utility of Pediatric Whole Exome Sequencing
ClinicalTrials.gov study NCT03525431. IPD Sharing: NO. Countries: 1. Publications: 3.
Whole exome sequencing reveals a long-term decline in effective population size of red spruce (Picea rubens)
Open the record for dataset details and reuse information.
Exome sequencing of a hybrid pine species complex on the Qinghai-Tibetan Plateau
Open the record for dataset details and reuse information.
Data from: A high-density exome capture genotype-by-sequencing panel for forestry breeding in Pinus radiata
Development of genome-wide resources for application in genomic selection or genome-wide association studies, in the absences of full reference genomes, present a challenge to the forestry industry, where longer breeding cycles could benefit from the accelerated selection possible through marker-based breeding value predictions. In particular, large conifer megagenomes require a strategy to reduce complexity, whilst ensuring genome-wide coverage is achieved. Using a transcriptome-based reference template, we have successfully developed a high density exome capture genotype-by-sequencing panel for radiata pine (Pinus radiata D.Don), capable of capturing in excess of 80,000 single nucleotide polymorphism (SNP) markers with a minor allele frequency above 0.03 in the population tested. This represents approximately 29,000 gene models from a core set of 48,914 probes. A set of 704 SMP markers capable of pedigree reconstruction and differentiating individual genotypes were tested within two full-sib mapping populations. While as few as 70 markers could reconstruct parentage in almost all cases, the impact of missing genotypes was noticeable in several offspring. Therefore, sets of 60 sets of 110 randomly selected SNP markers were compared for both parentage reconstruction and clone differentiation. The performance in parentage reconstruction showed little variation over 60 iterations. However, there was notable variation in discriminatory power between closely related individuals, indicating a higher density SNP marker panel may be required to elucidate hidden relationships in complex pedigrees.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.