Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
87
datasets available to search
ShareScore release 0.9.0
Dataset results
87 results for “WES”
Curated reference files for GCAP (WES)
<p>Provides big reference files or extra datasets/models for (in) GCAP project. </p> <p>The reference files are adapted from https://github.com/Wedge-lab/battenberg, more specifically, https://ora.ox.ac.uk/objects/uuid:08e24957-7e76-438a-bd38-66c48008cf52.</p> <p>News:</p> <ul> <li>Removed the correction files marked with 'update', which does not work for the latest version of ASCAT v3.</li> </ul>
TCGA WES processed by Demichelis Lab
<p>Here we provide processed data of the TCGA WES collection related to the study by Ciani Y. et al, "Allele-specific informed genomic analysis identifies loss of heterozygosity as a common trait of impaired tumor-suppressive processes" that implements the SPICE pipeline. Processed data includes SNV/indels, genomic calls as allele specific copy number, and LOH/gene expression linear models results. For more information visit <a href="https://github.com/demichelislab/SPICE-pipeline">https://github.com/demichelislab/SPICE-pipeline</a></p>
WES cropbioBonn WGGC CN 200M vcf results
<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Agilent v6 2x101bp PE, NovaSeq 6000, 200M reads</p>
WES cropbioBonn WGGC CN Twist Exome vcf results
<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Twist Exome, NovaSeq 6000</p>
WES cropbioBonn WGGC CN 75M vcf results
<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Agilent v6 2x101bp PE, NovaSeq 6000, 75M reads</p>
NA12878 WES Benchmark dataset
<p>This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session <strong>NA12878 WES Benchmark </strong>files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. </p> <p>The <a href="https://usegalaxy.org/u/erinija/p/omim-genes-in-na12878-wes-benchmark">"Procedure and datasets to cross-reference OMIM genes with the genomic regions of interest"</a> Galaxy page on <strong>usegalaxy.org</strong> server's <em><strong>Shared Data Pages</strong></em> describes practical procedure and several possible use cases for this data set. This page can be accessed freely by users logged into their accounts on usegalaxy.org. Please register if you don't have an account on usegalaxy.org Galaxy server. </p> <p>All genomic variant calls in all VCF files of this data set were decomposed and normalized with vt. This dataset contains: </p> <ol> <li>Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : <ol> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz</li> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi</li> <li>GIAB_v3.3.2_NA12878_HC_regions.bed</li> </ol> </li> <li>HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : <ul> <li>ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser <ol> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz</li> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi</li> <li> ARUP_SeqCap_EZ_Exome.bed</li> </ol> </li> <li>UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser <ol> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz</li> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi</li> <li>UCSF_WES_Agilent_V4_Custom.bed</li> </ol> </li> <li>Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory <ol> <li>CHEO_NA12878_WES_S1dataset.vcf.gz</li> <li>CHEO_NA12878_WES_S1dataset.vcf.gz.tbi</li> <li>Agilent_CRE_v2.bed</li> </ol> </li> </ul> </li> <li>Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : <ul> <li>Omim_Genes.bed </li> </ul> </li> </ol>
gnomAD SQLite database WES v4.0
<p>This package scales the huge gnomAD files (on average ~120G/chrom) to a SQLite database with a size of <100G and allows scientists to look for various variant annotations present in gnomAD (i.e. Allele Count, Depth, Minor Allele Frequency, etc.). (A query containing 300.000 variants takes ~40s.)</p><p>Find more information on <a href="https://github.com/KalinNonchev/gnomAD_DB">here</a>.</p><p>gnomAD SQLite database WES v4.0</p>
GHGA_sarek3_variant_calling_Agilent200M_WES
<p>Using the sarek pipeline default values. Aligned to HG38 using bwa, and dragmap. Variants are called using either haplotypecaller, strelka, deepvariant, or freebayes. Data input was Agilent 200M WES reads. Sample A006850052 (GIAB: HG001)</p>
gnomAD SQLite database WES v4.1
<div> <div> <div> </div> </div> </div> <div> <p>This package scales the huge gnomAD files (on average ~120G/chrom) to a SQLite database with a size of <100G and allows scientists to look for various variant annotations present in gnomAD (i.e. Allele Count, Depth, Minor Allele Frequency, etc.). (A query containing 300.000 variants takes ~40s.)</p> <p>Find more information on <a href="https://github.com/KalinNonchev/gnomAD_DB">here</a>.</p> <p>gnomAD SQLite database WES v4.1</p> </div>
DeepRVAT gene-trait association testing results on the 470k UK Biobank WES dataset
<p>Association testing results from DeepRVAT on the 470k UK Biobank WES dataset, covering all tested genes and traits. Tests were performed on the full dataset and on Caucasian individuals ("cohort" column). "Significant" indicates significance after multiple testing correction (FWER < 5%).</p>
WES benchmark results nf-core/sarek v3.1.1
<p>Variant calling results on benchmarking datasets produced with the nf-core/sarek v3.1.1 pipeline.</p>
WES CH UKBB 200K Summary statistics
<p>Bi-allelic summary statistics for "Exome-wide evidence of compound heterozygous effects".</p>
Crohn's Disease WES meta results
<p>Synopsis</p> <p>IBD exome-wide assocaition statistics from the meta-analysis of two callsets: Nextera and Twist. We used the fixed-effect meta-analysis in <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2922887/">METAL</a>. Column headers are self-explanatory. Allele2 is the tested/effect allele.</p> <p>Code availability</p> <p>Computer code used in this study:</p> <ul> <li>Hail (quality control, variants effect annotation; <a href="https://hail.is/">https://hail.is</a>)</li> <li>SAIGE (variant-based and gene-based association test; <a href="https://github.com/weizhouUMICH/SAIGE">https://github.com/weizhouUMICH/SAIGE</a>)</li> <li>METAL (meta-analysis; <a href="https://genome.sph.umich.edu/wiki/METAL_Documentation">https://genome.sph.umich.edu/wiki/METAL_Documentation</a>)</li> <li>PLINK (ancestry assignment, IBD relatedness QC and sample heterozygosity QC; plink1.9, <a href="https://www.cog-genomics.org/plink/">https://www.cog-genomics.org/plink/</a>; sample level QC pipeline, <a href="https://github.com/Annefeng/PBK-QC-pipeline">https://github.com/Annefeng/PBK-QC-pipeline</a>)</li> </ul> <p>Citation</p> <p><a href="https://www.medrxiv.org/content/10.1101/2021.06.15.21258641v2">Sazonovs, Stevens, Venkataraman, Yuan et al., MedRxiv, 2021</a></p>
Crohn-s-Disease-WES-meta
<p>Sequencing of over 100,000 individuals identifies multiple genes and rare variants associated with Crohn’s disease susceptibility</p> <p>Synopsis</p> <p>IBD exome-wide assocaition statistics from the meta-analysis of two callsets: Nextera and Twist. We used the fixed-effect meta-analysis in <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2922887/">METAL</a>. Column headers are self-explanatory. Allele2 is the tested/effect allele.</p> <p>Code availability</p> <p>Computer code used in this study:</p> <ul> <li>Hail (quality control, variants effect annotation; <a href="https://hail.is/">https://hail.is</a>) Analysis scripts are available in folder "Hail-scripts" of this repository.</li> <li>SAIGE (variant-based and gene-based association test; <a href="https://github.com/weizhouUMICH/SAIGE">https://github.com/weizhouUMICH/SAIGE</a>)</li> <li>METAL (meta-analysis; <a href="https://genome.sph.umich.edu/wiki/METAL_Documentation">https://genome.sph.umich.edu/wiki/METAL_Documentation</a>)</li> <li>PLINK (ancestry assignment, IBD relatedness QC and sample heterozygosity QC; plink1.9, <a href="https://www.cog-genomics.org/plink/">https://www.cog-genomics.org/plink/</a>; sample level QC pipeline, <a href="https://github.com/Annefeng/PBK-QC-pipeline">https://github.com/Annefeng/PBK-QC-pipeline</a>)</li> </ul> <p>Citation</p> <p><a href="https://www.medrxiv.org/content/10.1101/2021.06.15.21258641v2">Sazonovs, Stevens, Venkataraman, Yuan et al., MedRxiv, 2021</a></p>
WES benchmark results nf-core/sarek v3.4.3
Variant calling results on benchmarking datasets produced with nf-core/sarek
WES benchmark results nf-core/sarek v3.4.3
Variant calling results on benchmarking datasets produced with nf-core/sarek
WES benchmark results nf-core/sarek v3.4.3
Variant calling results on benchmarking datasets produced with nf-core/sarek
WES benchmark results nf-core/sarek v3.4.4
Variant calling results on benchmarking datasets produced with nf-core/sarek
NEXMIF encephalopathy: DeNovogear output of WES data of the family
<p>The developmental and epileptic encephalopathies (DEE) are the most severe group of epilepsies. Recently, <i>NEXMIF</i> mutations have been shown to cause a DEE in females, characterized by myoclonic–atonic epilepsy and recurrent nonconvulsive status. Here we used advanced neuroimaging techniques in a patient with a novel <i>NEXMIF</i> de novo mutation presenting with recurrent absence status with eyelid myoclonia, to reveal brain structural and functional changes that can bring the clinical phenotype to alteration within specific brain networks. Indeed, the alterations found in the patient involved the visual pericalcarine cortex and the middle frontal gyrus, regions that have been demonstrated to be a core feature in epilepsy phenotypes with visual sensitivity and eyelid myoclonia with absences.</p>
Whole-Exome Sequencing (WES) of Cancer Patients
ClinicalTrials.gov study NCT02127359. IPD Sharing: Not stated. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.