Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

257

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

257 results for “exome”

Learn how ShareScore rates datasets ↗
zenodo48/100

Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants

<p>This dataset&nbsp;consists of the reference data files, metadata and processed results files for the paper &quot;Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants,&quot; which&nbsp;investigates clonality in normal human dermal fibroblast cell populations in 32 cell lines from distinct donors, using bulk whole-exome sequencing and single-cell RNA-sequencing data.</p> <p>This dataset contains everything required to reproduce the results presented in the paper from&nbsp;processed data and results of our data processing workflows. Our analyses can be reproduced using the <a href="https://github.com/davismcc/fibroblast-clonality">source code</a>&nbsp;and instructions available at our <a href="https://davismcc.github.io/fibroblast-clonality/">project website</a>.</p> <p>The <em>entire</em> analysis workflow from raw data to final results is also reproducible but&nbsp;is substantially more complicated and computationally intensive.&nbsp;It also requires large datasets to be obtained from other repositories. Specifically, single-cell RNA-seq data have been deposited in the ArrayExpress database at EMBL-EBI under accession number E-MTAB-7167. Whole-exome sequencing data is available through the HipSci portal (www.hipsci.org). Combined with the dataset in this repository and following the instructions on the project website, it is possible to run our entire analysis pipeline.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"

<p>These are the burden test results (summary statistics) for the publication:</p> <p>&quot;Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer&rsquo;s Disease&quot;,</p> <p>Nature Genetics, 2022.</p> <p>&nbsp;</p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on&nbsp;6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL&gt;=75, LOF+REVEL&gt;=50, LOF+REVEL&gt;=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Twist Whole-Exome Sequencing Dataset - High Coverage - WGGC SIG4 Benchmarking

<p>GIAB Reference Genome for Benchmarking Initiatives in the West German Genome Center (WGGC) - SIG4.&nbsp;</p> <p>Twist Whole-Exome Sequencing Dataset - High Coverage - 200M Reads.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Exome data: Phoxinellus alepidotus and Telestes ukliva

<p>Reference exome and variant files for Phoxinellus alepidotus and Telestes ukliva</p> <p>Published in :&nbsp;Daane, J.M., Rohner, N., Konstantinidis, P., Djuranovic, S. and Harris M.P. (2016).&nbsp; Parallelism and epistasis in skeletal evolution identified through use of phylogenomic mapping strategies.&nbsp; <em>Mol Biol. Evol. </em>33:162-173.</p>

opencc-by-4.0Oct 2015View details →
zenodo44/100

Genetic variants (chr. 6) from Old World Schistosoma mansoni exomes

<p>Variant calling file (VCF) produced from exome libraries of <em>Schistosoma mansoni</em> (bloodfluke) samples from the Old Wold (West Africa (Senegal, Niger), East Africa (Tanzania), and Middle East (Oman)). One sample form the New World (Caribbean (HR9)) was added for comparison. The variants were called on the 3 Mb of chromosome 6 centered on the <em>SmSULT-OR</em> gene. This gene is involved in resistance to the drug oxamniquine&nbsp; (OXA). The aim of the related article was to investigate the origin of OXA resistant mutations in the New Wolrd by identifying sequence variation in <em>SmSULT-OR</em> in <em>S. mansoni</em> from the Old World, where OXA has seen minimal usage.</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)

<p>The data used in this tutorial are a subset of the data&nbsp;published previously in&nbsp;<a href="https://zenodo.org/record/3243160">Training material for the course &quot;Exome analysis with GALAXY&quot;</a>. Credit for uploading the original data goes to&nbsp;Paolo Uva and Gianmauro&nbsp;Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p>&nbsp;</p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you&#39;ll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load&nbsp;and, in fact, without having to install GEMINI&#39;s bundled annotation data.</p>

opencc-by-4.0May 2019View details →
dryad40/100

Data from: Association genetics of growth and adaptive traits in loblolly pine (Pinus taeda L.) using whole-exome-discovered polymorphisms

In the United States, forest genetics research began over 100 years ago and loblolly pine breeding programs were established in the 1950s. However, the genetics underlying complex traits of loblolly pine remains to be discovered. To address this, adaptive and growth traits were measured and analyzed in a clonally tested loblolly pine (Pinus taeda L.) population. Over 2.8 million single nucleotide polymorphism (SNP) markers detected from exome sequencing were used to test for single locus associations, SNP-SNP interactions and correlation of individual heterozygosity with phenotypic traits. A total of 36 SNP-trait associations were found for specific leaf area (5 SNPs), branch angle (2), crown width (3), stem diameter (4), total height (9), carbon isotope discrimination (4), nitrogen concentration (2), and pitch canker resistance traits (7). Eleven SNP-SNP interactions were found to be associated with branch angle (1 SNP-SNP interaction), crown width (2), total height (2), carbon isotope discrimination (2), nitrogen concentration (1), and pitch canker resistance (3). Non-additive effects imposed by dominance and epistasis account for a large fraction of the genetic variance for the quantitative traits. Genes that contain the identified SNPs have a wide spectrum of functions. Individual heterozygosity positively correlated with water use efficiency and nitrogen concentration. In conclusion, multiple effects identified in this study influence the performance of loblolly pines, provide resources for understanding the genetic control of complex traits, and have potential value for assessing with breeding through marker assisted selection and genomic selection.

opencc-zeroDec 2018View details →
zenodo40/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families, ascertained to contain at least four members affected with lymphoid cancer.&nbsp; Please note that previous versions of this repository omitted a key file linking the genotypes of individuals to their family and individual IDs; this file, geno_key.txt, is now included. All other files remain the same as in previous versions.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>geno_key.txt &ndash; a plain-text file that links the genotyped individuals to their family and individual IDs.</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at <a href="../records/12694914">https://zenodo.org/records/12694914</a></p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →
zenodo40/100

Training material for the course "Exome analysis with GALAXY"

<p>Galaxy is an open source, web-based platform for data intensive biomedical research. It makes accessible bioinformatics applications to users lacking programming skills, enabling them to easily build analysis workflows for NGS data.<br /> &nbsp;<br /> The course &quot;<strong>Exome analysis using Galaxy</strong>&quot; is aimed at PhD student, biologists, clinicians and researchers who are analysing, or need to analyse in the near future, high throughput exome sequencing data. The aim of the course is to make participants familiarise with the Galaxy platform and prepare them to work independently, using state-of-the art tools for the analysis of exome sequencing data.</p> <p>The course will be delivered using a mixture of lectures and computer based hands-on practical sessions. Lectures will provide an up-to-date overview of the strategies for the analysis of exome next-generation experiments, starting from the raw sequence data. Analyses include sequence quality control, alignment to a reference genome, refinement of aligned sequences, variant calling, annotation and interpretation, and tools for visual inspection of results. Participants will apply the knowledge gained during the course to the analysis of Illumina&rsquo;s real exome datasets, and implement workflows to reproduce the complete analysis. After the course, participants will be able to create pipeline for their individual analyses.</p> <p>Those are the needed datasets for this course.</p>

opencc-zeroSep 2016View details →
zenodo40/100

A clinical exome study on a family segregating pontocerebellar hyploplasia

<p><span>Pontocerebellar hypoplasia type 2D (PCH2D) is caused by mutations in the SEPSECS gene (chr. 4p15.2), encoding O-Phosphoseryl-tRNA:selenocysteinyl-tRNA synthase. This is a key enzyme in the biosynthesis of selenoproteins, which act in maintaining antioxidant systems. . We describe a novel patient with compound heterozygosity in the SEPSECS gene including a novel missense variant.&nbsp;</span><span> This study broadens the genetic background and associated PCH2D phenotype, supporting the causal link with mitochondrial disorders in selenoproteins biosynthesis deficiency</span></p> <p>&nbsp;</p> <p>22M1764&nbsp; vcf files are related to the proband</p> <p>22M1764M vcf&nbsp; files are related to his mother</p> <p>22M1764P files are related to his father</p> <p>&nbsp;</p> <p>The proband suffers with pontocerebellar hyplosia while his parents are healthy</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

WES cropbioBonn WGGC CN Twist Exome vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Twist Exome, NovaSeq 6000</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Data files for manuscript "Re-evaluation and Re-analysis of 152 research exomes five years after the initial report reveals clinically relevant changes in 18%"

<p>#2023-06-16<br> #Summary<br> This ZIP-file contains the data files used for all analyses for the manuscript &quot;Re-evaluation and Re-analysis of 152 research exomes five years after the initial report reveals clinically relevant changes in 18%&quot;.</p> <p><br> #File structure<br> README.txt&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;This README file.<br> File S02 (&quot;FileS2_conNDD-cohort.xlsx&quot;)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;All variants identified by Reuter et al. previously with reevaluated variants and addition variants identified in this&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;project togetehr with information about the families, individuals, samplesand the BAM files assessed in this project.<br> File S03 (&quot;FileS3_conNDD-variants.xlsx&quot;)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;All variant data analyzed from the cohort. Including a sheet with thresholdes for in silico predictions tools used to predict effect of variants,&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;a table with exome wide homozygous variants in 4 categories (A45, LGD, Missense, Splice), a table with exome wide variants in 4 categories (A45, LGD, Missense, Splice)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;filtered for domiant genes associated with neurodevelopmental disorders in SysID (Prime and Candidate list), a table with exome wide variants in 4 categories (A45, LGD, Missense, Splice) filtered for recessive genes associated with neurodevelopmental disorders in SysID (Prime and Candidate list), a table withcopy number (CN) calls for the cohort and a table withcalls for runs of homozygosity (RoH) regions.</p> <p>#Files and checksums<br> 29c4b2f3dd8985d268f50dd3e0265798&nbsp;&nbsp; &nbsp;./FileS2_conNDD-cohort.xlsx<br> a054334637b8b22a9bf743db1e348663&nbsp;&nbsp; &nbsp;./FileS3_conNDD-variants.xlsx<br> &nbsp;</p>

opencc-by-4.0Sep 2022View details →
dryad40/100

Data from: Association genetics of growth and adaptive traits in loblolly pine (Pinus taeda L.) using whole-exome-discovered polymorphisms

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad40/100

Data from: Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra

Open the record for dataset details and reuse information.

publicDec 2018View details →
dryad40/100

Data from: Optimizing exome captures in species with large genomes using species-specific repetitive DNA blocker

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad36/100

Data from: Exome resequencing reveals signatures of demographic and adaptive processes across the genome and range of black cottonwood (Populus trichocarpa)

Extant variation in temperate and boreal plant species has been influenced by both demographic histories associated with Pleistocene glacial cycles and adaptation to local climate. We used sequence capture to investigate the role of these neutral and adaptive processes in shaping diversity in black cottonwood (Populus trichocarpa). Nucleotide diversity and Tajima's D were lowest at replacement sites and highest at intergenic sites, while LD showed the opposite pattern. With samples grouped into three populations arrayed latitudinally, effective population size was highest in the north, followed by south and centre, and LD was highest in the south followed by the north and centre, suggesting a possible northern glacial refuge. FST outlier analysis revealed that promoter, 5′-UTR and intronic sites were enriched for outliers compared with coding regions, while no outliers were found among intergenic sites. Codon usage bias was evident, and genes with synonymous outliers had 30% higher average expression compared with genes containing replacement outliers. These results suggest divergent selection related to regulation of gene expression is important to local adaptation in P. trichocarpa. Finally, within-population selective sweeps were much more pronounced in the central population than in putative northern and southern refugia, which may reflect the different demographic histories of the populations and concomitant effects on signatures of genetic hitchhiking from standing variation.

opencc-zeroDec 2013View details →
zenodo36/100

Training material for exome sequencing

<p>Exome sequencing means that all protein-coding genes in a genome are sequenced.</p> <p>In Humans, there are ~180,000 exons that makes up 1% of the human genome which&nbsp;contain ~30 million base pairs. Mutations in the exome have usually a higher&nbsp;impact and more severe consequences, than in the remaining 99% of the genome.</p> <p>With exome sequencing, one can identify genetic variation that is responsible&nbsp;for&nbsp;both Mendelian and common diseases without the high costs&nbsp;associated with&nbsp;whole-genome sequencing. Indeed, exome sequencing is the&nbsp;most efficient way to&nbsp;identify the genetic variants in all of an individual&#39;s genes.&nbsp;Exome sequencing&nbsp;is cheaper also than whole-genome sequencing.&nbsp;</p> <p>&nbsp;</p> <p>For training on exome sequencing data analysis, the Galaxy community proposes two tutorials (https://github.com/bgruening/training-material/tree/master/Exome-Seq). Here, you can find the needed datasets for these tutorials.</p>

opencc-by-4.0Aug 2016View details →
zenodo36/100

Exome-wide association analysis (ExWAS) of clonal haematopoiesis in 136,401 Admixed Americans and 416,118 Europeans

<p>We performed exome-wide association analysis (ExWAS) of germline genetic variants identified from whole-exome sequencing (WES) to identify novel inherited genetic determinants of clonal haematopoiesis (CH). Here, we provide the summary statistics from ExWAS CH performed on Admixed Americans recruited to the Mexico City Prospective Study (MCPS), Europeans recruited to the United Kingdom Biobank (UKB), and cross-ancestry meta-analysis of Admixed Americans and Europeans. Analyses was performed with REGENIE software (Firth's logistic regression), Fisher's exact test, and METAL software (inverse variance-weighted average method to derive effect size and <em>P</em>-value method to derive P value), respectively.</p> <p>&nbsp;</p> <p>In version 1 of this repository, UKB variants (*_UKB.tsv) with minor allele frequency (MAF) 1% or more were uploaded. In version 2, this is now rectified so that rare variants with MAF of 0.1% or more were uploaded. This threshold now matches the MCPS (*_MCPS.tsv) and UKB-MCPS meta-analysis summary statistics (*_MCPS-UKB_meta-analysis.tsv)</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains&nbsp;all the data files for a simulated exome-sequencing study of 150 families ascertained to contain at least four members affected with lymphoid cancer.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at&nbsp;<a href="https://zenodo.org/record/6505385">https://zenodo.org/record/6505385</a>&nbsp;.</p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record