Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

87

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

87 results for “WES”

Learn how ShareScore rates datasets ↗
zenodo44/100

Curated reference files for GCAP (WES)

<p>Provides big reference files or extra datasets/models for (in) GCAP project.&nbsp;</p> <p>The reference files are adapted from https://github.com/Wedge-lab/battenberg, more specifically, https://ora.ox.ac.uk/objects/uuid:08e24957-7e76-438a-bd38-66c48008cf52.</p> <p>News:</p> <ul> <li>Removed the correction files marked with 'update', which does not work for the latest version of ASCAT v3.</li> </ul>

opencc-by-4.0Sep 2021View details →
zenodo44/100

TCGA WES processed by Demichelis Lab

<p>Here we provide processed data of the TCGA WES collection related to the study by Ciani Y. et al, &quot;Allele-specific informed genomic analysis identifies loss of heterozygosity as a common trait of impaired tumor-suppressive processes&quot; that implements the SPICE pipeline. Processed data includes SNV/indels, genomic calls as allele specific copy number, and LOH/gene expression linear models results. For more information visit <a href="https://github.com/demichelislab/SPICE-pipeline">https://github.com/demichelislab/SPICE-pipeline</a></p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

WES cropbioBonn WGGC CN 200M vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Agilent v6 2x101bp PE, NovaSeq 6000, 200M reads</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

WES cropbioBonn WGGC CN Twist Exome vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Twist Exome, NovaSeq 6000</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

WES cropbioBonn WGGC CN 75M vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Agilent v6 2x101bp PE, NovaSeq 6000, 75M reads</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

NA12878 WES Benchmark dataset

<p>This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session <strong>NA12878 WES Benchmark </strong>files in a single dataset so that these files can be used in other applications or genome browsers such as IGV.&nbsp;</p> <p>The <a href="https://usegalaxy.org/u/erinija/p/omim-genes-in-na12878-wes-benchmark">&quot;Procedure and datasets to cross-reference OMIM genes with the genomic regions of interest&quot;</a>&nbsp; Galaxy page&nbsp; on&nbsp; <strong>usegalaxy.org</strong> server&#39;s&nbsp;<em><strong>Shared Data Pages</strong></em> describes&nbsp;practical procedure and several possible use cases for this data set. This page can be accessed freely by users logged into their accounts on usegalaxy.org.&nbsp; Please register if you don&#39;t have an account on usegalaxy.org Galaxy server. &nbsp;</p> <p>All&nbsp; genomic variant calls in&nbsp; all VCF files of this data set were decomposed and normalized with vt. This dataset contains:&nbsp;</p> <ol> <li>Genome in a bottle (GIAB)&nbsp;version 3.3.2 high confidence (HC)&nbsp; variant calls and genomic regions for HapMap individual NA12878 : <ol> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz</li> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi</li> <li>GIAB_v3.3.2_NA12878_HC_regions.bed</li> </ol> </li> <li>HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : <ul> <li>ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser <ol> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz</li> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi</li> <li>&nbsp;ARUP_SeqCap_EZ_Exome.bed</li> </ol> </li> <li>UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser <ol> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz</li> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi</li> <li>UCSF_WES_Agilent_V4_Custom.bed</li> </ol> </li> <li>Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory <ol> <li>CHEO_NA12878_WES_S1dataset.vcf.gz</li> <li>CHEO_NA12878_WES_S1dataset.vcf.gz.tbi</li> <li>Agilent_CRE_v2.bed</li> </ol> </li> </ul> </li> <li>Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : <ul> <li>Omim_Genes.bed&nbsp;</li> </ul> </li> </ol>

opencc-by-4.0Jan 2020View details →
zenodo36/100

gnomAD SQLite database WES v4.0

<p>This package scales the huge gnomAD files (on average ~120G/chrom) to a SQLite database with a size of &lt;100G &nbsp;and allows scientists to look for various variant annotations present in gnomAD (i.e. Allele Count, Depth, Minor Allele Frequency, etc.). (A query containing 300.000 variants takes ~40s.)</p><p>Find more information on <a href="https://github.com/KalinNonchev/gnomAD_DB">here</a>.</p><p>gnomAD SQLite database WES v4.0</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

GHGA_sarek3_variant_calling_Agilent200M_WES

<p>Using the sarek pipeline default values. Aligned to HG38 using bwa, and dragmap. Variants are called using either haplotypecaller, strelka, deepvariant, or freebayes. Data input was Agilent 200M WES reads. Sample A006850052 (GIAB: HG001)</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

gnomAD SQLite database WES v4.1

<div> <div> <div>&nbsp;</div> </div> </div> <div> <p>This package scales the huge gnomAD files (on average ~120G/chrom) to a SQLite database with a size of &lt;100G &nbsp;and allows scientists to look for various variant annotations present in gnomAD (i.e. Allele Count, Depth, Minor Allele Frequency, etc.). (A query containing 300.000 variants takes ~40s.)</p> <p>Find more information on <a href="https://github.com/KalinNonchev/gnomAD_DB">here</a>.</p> <p>gnomAD SQLite database WES v4.1</p> </div>

opencc-by-4.0Apr 2024View details →
zenodo36/100

DeepRVAT gene-trait association testing results on the 470k UK Biobank WES dataset

<p>Association testing results from DeepRVAT on the 470k UK Biobank WES dataset, covering all tested genes and traits. Tests were performed on the full dataset and on Caucasian individuals ("cohort" column). "Significant" indicates significance after multiple testing correction (FWER &lt; 5%).</p>

openmit-licenseAug 2024View details →
zenodo36/100

WES benchmark results nf-core/sarek v3.1.1

<p>Variant calling results on benchmarking datasets produced with the nf-core/sarek v3.1.1 pipeline.</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

WES CH UKBB 200K Summary statistics

<p>Bi-allelic summary statistics for "Exome-wide evidence of compound heterozygous effects".</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Crohn's Disease WES meta results

<p>Synopsis</p> <p>IBD exome-wide assocaition statistics from the meta-analysis of two callsets: Nextera and Twist. We used the fixed-effect meta-analysis in&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2922887/">METAL</a>. Column headers are self-explanatory. Allele2 is the tested/effect allele.</p> <p>Code availability</p> <p>Computer code used in this study:</p> <ul> <li>Hail (quality control, variants effect annotation;&nbsp;<a href="https://hail.is/">https://hail.is</a>)</li> <li>SAIGE (variant-based and gene-based association test;&nbsp;<a href="https://github.com/weizhouUMICH/SAIGE">https://github.com/weizhouUMICH/SAIGE</a>)</li> <li>METAL (meta-analysis;&nbsp;<a href="https://genome.sph.umich.edu/wiki/METAL_Documentation">https://genome.sph.umich.edu/wiki/METAL_Documentation</a>)</li> <li>PLINK (ancestry assignment, IBD relatedness QC and sample heterozygosity QC; plink1.9,&nbsp;<a href="https://www.cog-genomics.org/plink/">https://www.cog-genomics.org/plink/</a>; sample level QC pipeline,&nbsp;<a href="https://github.com/Annefeng/PBK-QC-pipeline">https://github.com/Annefeng/PBK-QC-pipeline</a>)</li> </ul> <p>Citation</p> <p><a href="https://www.medrxiv.org/content/10.1101/2021.06.15.21258641v2">Sazonovs, Stevens, Venkataraman, Yuan et al., MedRxiv, 2021</a></p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Crohn-s-Disease-WES-meta

<p>Sequencing of over 100,000 individuals identifies multiple genes and rare variants associated with Crohn&rsquo;s disease susceptibility</p> <p>Synopsis</p> <p>IBD exome-wide assocaition statistics from the meta-analysis of two callsets: Nextera and Twist. We used the fixed-effect meta-analysis in&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2922887/">METAL</a>. Column headers are self-explanatory. Allele2 is the tested/effect allele.</p> <p>Code availability</p> <p>Computer code used in this study:</p> <ul> <li>Hail (quality control, variants effect annotation;&nbsp;<a href="https://hail.is/">https://hail.is</a>) Analysis scripts are available in folder &quot;Hail-scripts&quot; of this repository.</li> <li>SAIGE (variant-based and gene-based association test;&nbsp;<a href="https://github.com/weizhouUMICH/SAIGE">https://github.com/weizhouUMICH/SAIGE</a>)</li> <li>METAL (meta-analysis;&nbsp;<a href="https://genome.sph.umich.edu/wiki/METAL_Documentation">https://genome.sph.umich.edu/wiki/METAL_Documentation</a>)</li> <li>PLINK (ancestry assignment, IBD relatedness QC and sample heterozygosity QC; plink1.9,&nbsp;<a href="https://www.cog-genomics.org/plink/">https://www.cog-genomics.org/plink/</a>; sample level QC pipeline,&nbsp;<a href="https://github.com/Annefeng/PBK-QC-pipeline">https://github.com/Annefeng/PBK-QC-pipeline</a>)</li> </ul> <p>Citation</p> <p><a href="https://www.medrxiv.org/content/10.1101/2021.06.15.21258641v2">Sazonovs, Stevens, Venkataraman, Yuan et al., MedRxiv, 2021</a></p>

opencc-by-4.0May 2022View details →
zenodo32/100

WES benchmark results nf-core/sarek v3.4.3

Variant calling results on benchmarking datasets produced with nf-core/sarek

opencc-zeroAug 2024View details →
zenodo32/100

WES benchmark results nf-core/sarek v3.4.3

Variant calling results on benchmarking datasets produced with nf-core/sarek

opencc-zeroAug 2024View details →
zenodo32/100

WES benchmark results nf-core/sarek v3.4.3

Variant calling results on benchmarking datasets produced with nf-core/sarek

opencc-zeroAug 2024View details →
zenodo32/100

WES benchmark results nf-core/sarek v3.4.4

Variant calling results on benchmarking datasets produced with nf-core/sarek

opencc-zeroSep 2024View details →
dryad32/100

NEXMIF encephalopathy: DeNovogear output of WES data of the family

<p>The developmental and epileptic encephalopathies (DEE) are the most severe group of epilepsies. Recently, <i>NEXMIF</i> mutations have been shown to cause a DEE in females, characterized by myoclonic–atonic epilepsy and recurrent nonconvulsive status. Here we used advanced neuroimaging techniques in a patient with a novel <i>NEXMIF</i> de novo mutation presenting with recurrent absence status with eyelid myoclonia, to reveal brain structural and functional changes that can bring the clinical phenotype to alteration within specific brain networks. Indeed, the alterations found in the patient involved the visual pericalcarine cortex and the middle frontal gyrus, regions that have been demonstrated to be a core feature in epilepsy phenotypes with visual sensitivity and eyelid myoclonia with absences.</p>

opencc-zeroAug 2021View details →
ClinicalTrials.gov32/100

Whole-Exome Sequencing (WES) of Cancer Patients

ClinicalTrials.gov study NCT02127359. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record