Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

257

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

257 results for “exome”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Genomic prediction accuracies in space and time for height and wood density of Douglas-fir using exome capture as the genotyping platform

Background Genomic selection (GS) can offer unprecedented gains, in terms of cost efficiency and generation turnover, to forest tree selective breeding; especially for late expressing and low heritability traits. Here, we used: 1) exome capture as a genotyping platform for 1372 Douglas-fir trees representing 37 full-sib families growing on three sites in British Columbia, Canada and 2) height growth and wood density (EBVs), and deregressed estimated breeding values (DEBVs) as phenotypes. Representing models with (EBVs) and without (DEBVs) pedigree structure. Ridge regression best linear unbiased predictor (RR-BLUP) and generalized ridge regression (GRR) were used to assess their predictive accuracies over space (within site, cross-sites, multi-site, and multi-site to single site) and time (age-age/ trait-trait). Results The RR-BLUP and GRR models produced similar predictive accuracies across the studied traits. Within-site GS prediction accuracies with models trained on EBVs were high (RR-BLUP: 0.79–0.91 and GRR: 0.80–0.91), and were generally similar to the multi-site (RR-BLUP: 0.83–0.91, GRR: 0.83–0.91) and multi-site to single-site predictive accuracies (RR-BLUP: 0.79–0.92, GRR: 0.79–0.92). Cross-site predictions were surprisingly high, with predictive accuracies within a similar range (RR-BLUP: 0.79–0.92, GRR: 0.78–0.91). Height at 12 years was deemed the earliest acceptable age at which accurate predictions can be made concerning future height (age-age) and wood density (trait-trait). Using DEBVs reduced the accuracies of all cross-validation procedures dramatically, indicating that the models were tracking pedigree (family means), rather than marker-QTL LD. Conclusions While GS models' prediction accuracies were high, the main driving force was the pedigree tracking rather than LD. It is likely that many more markers are needed to increase the chance of capturing the LD between causal genes and markers.

opencc-zeroDec 2016View details →
dryad32/100

Data from: A high-density exome capture genotype-by-sequencing panel for forestry breeding in Pinus radiata

Development of genome-wide resources for application in genomic selection or genome-wide association studies, in the absences of full reference genomes, present a challenge to the forestry industry, where longer breeding cycles could benefit from the accelerated selection possible through marker-based breeding value predictions. In particular, large conifer megagenomes require a strategy to reduce complexity, whilst ensuring genome-wide coverage is achieved. Using a transcriptome-based reference template, we have successfully developed a high density exome capture genotype-by-sequencing panel for radiata pine (Pinus radiata D.Don), capable of capturing in excess of 80,000 single nucleotide polymorphism (SNP) markers with a minor allele frequency above 0.03 in the population tested. This represents approximately 29,000 gene models from a core set of 48,914 probes. A set of 704 SMP markers capable of pedigree reconstruction and differentiating individual genotypes were tested within two full-sib mapping populations. While as few as 70 markers could reconstruct parentage in almost all cases, the impact of missing genotypes was noticeable in several offspring. Therefore, sets of 60 sets of 110 randomly selected SNP markers were compared for both parentage reconstruction and clone differentiation. The performance in parentage reconstruction showed little variation over 60 iterations. However, there was notable variation in discriminatory power between closely related individuals, indicating a higher density SNP marker panel may be required to elucidate hidden relationships in complex pedigrees.

opencc-zeroOct 2019View details →
zenodo32/100

Easy BAM files from practising RNA-seq and exome-seq visualization

<p>Two sets of small BAM files for practising RNA-seq and WES visualization (human)&nbsp;</p> <p>All data was processed from open access SRA datasets.&nbsp;</p> <p>Description:</p> <table> <tbody> <tr> <td>NAME</td> <td>RNA-EMT</td> <td>WES-LUNG</td> </tr> <tr> <td>SOURCE</td> <td>Yang et al. https://doi.org/10.1128/mcb.00019-16, SRA: PRJNA304419</td> <td>Ju et al. https://doi.org/10.1101/gr.133645.111; SRA: ERP001071)</td> </tr> <tr> <td>DESCRIPTION</td> <td>Human RNA-seq BAM files&nbsp;</td> <td>Human WES BAM files: Lung cancer</td> </tr> <tr> <td>SEQUENCING TYPE</td> <td>Illumina paired-end, stranded</td> <td>Illumina paired-end</td> </tr> <tr> <td>METHODS</td> <td>Original data are from cell lines subjected to epthithelium-mesenchymal transition. We took replicates corresponding to Day 0 (uninduced cells) and Day 7 (7 days after induction). Reads were aligned with STAR. BAM files are restricted to chromosome 18. Gene of interest: CDH2, up in mesenchymal state.</td> <td>WES were obtained from tumor and blood tissues from a non small cell lung cancer patient. Reads were aligned with BWA. A fraction of sequences corresponding to a panel of genes were retained. Genes of interest: DUSP27, KRAS</td> </tr> <tr> <td>MAPPED_TO</td> <td>HG19</td> <td>HG19</td> </tr> <tr> <td>SAMPLING RATE</td> <td>0.50%</td> <td>0.10%</td> </tr> <tr> <td>FILE SIZE</td> <td>2*40Mb</td> <td>2*7Mb</td> </tr> </tbody> </table>

opencc-by-4.0Nov 2024View details →
dryad32/100

Schistosoma mansoni raw genotype calls from exome data

<p><em>Schistosoma mansoni, </em>a snail-vectored, blood fluke that infects humans, was introduced into the Americas from Africa during the Trans-Atlantic slave trade. As this parasite shows strong specificity to the snail intermediate host, we expected that adaptation to S. American <em>Biomphalaria</em> spp. snails would result in population bottlenecks and strong signatures of selection. We scored 475,081 single nucleotide variants (SNVs) in 143 <em>S. mansoni</em> from the Americas (Brazil, Guadeloupe, and Puerto Rico) and Africa (Cameroon, Niger, Senegal, Tanzania, and Uganda), and used these data to ask: (i) Was there a population bottleneck during colonization? (ii) Can we identify signatures of selection associated with colonization? And (iii) what were the source populations for colonizing parasites? We found a 2.4-2.9-fold reduction in diversity and much slower decay in linkage disequilibrium (LD) in parasites from East to West Africa. However, we observed similar nuclear diversity and LD in West Africa and Brazil, suggesting no strong bottlenecks and limited barriers to colonization. We identified five genome regions showing selection in the Americas, compared with three in West Africa and none in East Africa, which we speculate may reflect adaptation during colonization. Finally, we infer that unsampled African populations from central African regions between Benin and Angola, with contributions from Niger, are likely the major source(s) for Brazilian <em>S. mansoni</em>. The absence of a bottleneck suggests that this is a rare case of a serendipitous invasion, where <em>S. mansoni </em>parasites were preadapted to the Americas and were able to establish with relative ease.</p>

opencc-zeroMar 2022View details →
zenodo32/100

Broad exome calling interval list

<p>Exome calling interval list (v1, GRCh38)&nbsp;used by the Broad Institute</p>

opencc-by-4.0Jun 2022View details →
dryad32/100

HyRAD-X Exome Capture Museomics Unravels Giant Ground Beetle Evolution

<p>Abstract Advances in phylogenomics contribute toward resolving long-standing evolutionary questions. Notwithstanding, genetic diversity contained within more than a billion biological specimens deposited in natural history museums remains recalcitrant to analysis owing to challenges posed by its intrinsically degraded nature. Yet that tantalizing resource could be critical in overcoming taxon sampling constraints hindering our ability to address major evolutionary questions. We addressed this impediment by developing phyloHyRAD, a new bioinformatic pipeline enabling locus recovery at a broad evolutionary scale from HyRAD-X exome capture of museum specimens of low DNA integrity using a benchtop RAD-derived exome-complexity-reduction probe set developed from high DNA integrity specimens. Our new pipeline can also successfully align raw RNAseq transcriptomic and ultraconserved element reads with the RAD-derived probe catalog. Using this method, we generated a robust timetree for Carabinae beetles, the lack of which had precluded study of macroevolutionary trends pertaining to their biogeography and wing-morphology evolution. We successfully recovered up to 2,945 loci with a mean of 1,788 loci across the exome of specimens of varying age. Coverage was not significantly linked to specimen age, demonstrating the wide exploitability of museum specimens. We also recovered fragmentary mitogenomes compatible with Sanger-sequenced mtDNA. Our phylogenomic timetree revealed a Lower Cretaceous origin for crown group Carabinae, with the extinct Aplothorax (Waterhouse, 1841) nested within the genus Calosoma (Weber, 1801) demonstrating the junior synonymy of Aplothorax syn. nov., resulting in the new combination Calosoma burchellii (Waterhouse, 1841) comb. nov. This study compellingly illustrates that HyRAD-X and phyloHyRAD efficiently provide genomic-level data sets informative at deep evolutionary scales.</p>

opencc-zeroSep 2022View details →
zenodo32/100

Scalable mixed model approaches for set-based association studies on large-scale categorical data analysis and its application to 450k exome sequencing data in UK Biobank

<p>The ongoing release of large-scale sequencing data in the UK Biobank allows for identifying associations between rare variants and complex traits. SAIGE-GENE+ is a valid approach to conducting set-based association tests for quantitative and binary traits. However, for ordinal categorical phenotypes, applying SAIGE-GENE+ with treating the trait as quantitative or binarizing the trait can cause inflated type I error rates or power loss. In this study, we propose a novel method for rare-variant association tests, POLMM-GENE, in which a proportional odds logistic mixed model was used to characterize ordinal categorical phenotypes while adjusting for sample relatedness. POLMM-GENE fully utilizes the categorical nature of phenotypes and thus can well control type I error rates while remaining powerful. In the analyses of UK Biobank 450k whole exome-sequencing data for 5 ordinal categorical traits, POLMM-GENE identified 54 gene-phenotype associations.</p>

opencc-by-4.0Sep 2022View details →
dryad32/100

Data from: HyRAD-X, a versatile method combining exome capture and RAD sequencing to extract genomic information from ancient DNA

Over the last decade, protocols aimed at reproducibly sequencing reduced-genome subsets in non-model organisms have been widely developed. Their use is however limited to DNA of relatively high molecular weight. During the last year, several methods exploiting hybridization capture using probes based on RAD-sequencing loci have circumvented this limitation and opened avenues to the study of samples characterized by degraded DNA, such as historical specimens. Here, we present a major update to those methods, namely Hybridization capture from RAD-derived probes obtained from a reduced eXome template (hyRAD-X), a technique applying RAD-sequencing to messenger RNA from one or few fresh specimens to elaborate bench-top produced probes, i.e., a reduced representation of the exome, further used to capture homologous DNA from a samples set. In contrast to previous hybridization-capture methods, the reference catalog on which reads are aligned does not rely on de novo assembly of anonymous RAD-sequencing loci, but on an assembled transcriptome obtained from RNAseq data, thus increasing the accuracy of loci definition and Single-Nucleotide-Polmorphisms (SNP) call, and targeting, specifically, expressed genes. Finally, the capture step of hyRAD-X relies on RNA probes, increasing stringency of hybridization, making it well suited for low-content DNA samples. As a proof of concept, we applied hyRAD-X to subfossil needles from the coniferous tree Abies alba, collected in lake sediments (Origlio, Switzerland) and dating back from 7200-5800 years before present (BP). More specifically we investigated genetic variation before, during, and after an anthropogenic perturbation that caused an abrupt decrease in Abies alba population size, 6500-6200 years BP. HyRAD-X produced a matrix encompassing 524 exome-derived SNPs. Despite a lower observed heterozygosity was observed during the 6.500-6.200 years BP time slice, genetic composition was nearly identical before and after the perturbation, indicating that re-expansion of the population after the decline was driven by autochthonous specimens. To the best of our knowledge, this is the first time a population genomic study incorporating ancient DNA samples of tree subfossils is conducted at a moderate cost using reproducible exome-reduced complexity.

opencc-zeroDec 2016View details →
dryad32/100

Haploid, diploid, and pooled exome capture recapitulate features of biology and paralogy in two non-model tree species

<p>Despite their suitability for studying evolution, many conifer species have large and repetitive giga-genomes (16-31Gbp) that create hurdles to producing high coverage SNP datasets that capture diversity from across the entirety of the genome. Due in part to multiple ancient whole genome duplication events, gene family expansion and subsequent evolution within <i>Pinaceae</i>, false diversity from the misalignment of paralog copies creates further challenges in accurately and reproducibly inferring evolutionary history from sequence data. Here, we leverage the cost-saving benefits of pool-seq and exome-capture to discover SNPs in two conifer species, Douglas-fir (<i>Pseudotsuga menziesii</i> var. <i>menziesii </i>(Mirb.) Franco, <i>Pinaceae</i>) and jack pine (<i>Pinus banksiana</i> Lamb., <i>Pinaceae</i>). We show, using minimal baseline filtering, that allele frequencies estimated from pooled individuals show a strong positive correlation with those estimated by sequencing the same population as individuals (r &gt; 0.948), on par with such comparisons made in model organisms. Further, we highlight the utility of haploid megagametophyte tissue for identifying sites that are likely due to misaligned paralogs. Together with additional minor filtering, we show that it is possible to remove many of the loci with large frequency estimate discrepancies between individual and pooled sequencing approaches, improving the correlation further (r &gt; 0.973). Our work addresses bioinformatic challenges in non-model organisms with large and complex genomes, highlights the use of megagametophyte tissue for the identification of paralog sites, and suggests the combination of pool-seq and exome capture to be robust for further evolutionary hypothesis testing in these systems.</p>

opencc-zeroDec 2020View details →
ClinicalTrials.gov32/100

NIPD of CFTC by WGA Coupled to Mini-exome Sequencing

ClinicalTrials.gov study NCT03743948. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Exome Sequencing Study in Cardiomyopathy to Identify New Risk Variants

ClinicalTrials.gov study NCT03754101. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

A Study of Consent Forms for Whole Exome and Whole Genome Sequencing

ClinicalTrials.gov study NCT01927770. IPD Sharing: Not stated. Countries: 1. Publications: 4.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Combining Exome and Transcriptome Data to Unravel the Genetic Basis of the Lissencephalies

ClinicalTrials.gov study NCT05185414. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Whole Exome Sequencing in CKD Hypertension

ClinicalTrials.gov study NCT03718585. IPD Sharing: NO. Countries: 1. Publications: 25.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Exome and Genome Analysis to Elucidate Genetic Etiologies and Population Characteristics in the Plain Community

ClinicalTrials.gov study NCT02927158. IPD Sharing: YES. Countries: 1. Publications: 18.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov32/100

Contribution of High-throughput Exome Sequencing in the Diagnosis of the Cause Fetal Polymalformation Syndromes

ClinicalTrials.gov study NCT02512354. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Comparison of the Non-invasive Approach and Fetal Exome Sequencing in Prenatal Diagnosis When Fetal Ultrasound Signs Are Discovered

ClinicalTrials.gov study NCT05182242. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Exome Sequencing in Autistic Spectrum Disorder

ClinicalTrials.gov study NCT01059201. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Whole Exome Sequencing and Whole Genome Sequencing for Nonimmune Fetal/Neonatal Hydrops

ClinicalTrials.gov study NCT03911531. IPD Sharing: YES. Countries: 1. Publications: 6.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov32/100

Adult Patients With Undiagnosed Conditions and Their Responses to Clinically Uncertain Results From Exome Sequencing

ClinicalTrials.gov study NCT03605004. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record