Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

244

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

244 results for “genomic variants”

Learn how ShareScore rates datasets ↗
zenodo36/100

Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.

<p>Ebola genome is manually mutated to contain structural variants. There are seven mutated ebola genomes each one for two large deletions, insertions, duplications, translocations, inversions, one complex variant1 (consecutive three mutations - insertion, duplication and deletion) and one complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion. Paired-end illumine RNASEQ reads are simulated in fastq format.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.

<p>Ebola genome is manually mutated to contain structural variants. There are seven mutated ebola genomes each one for large deletion, insertion, duplication, translocation, inversion, complex variant1 (consecutive three mutations - insertion, duplication and deletion) and complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion. Paired-end illumine RNASEQ reads are simulated in fastq format.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Rare variant replaced Korea reference genome fasta

<p>The rare variants of Korea reference genome fasta&nbsp;were&nbsp;replaced with using 396 Korean vcf information</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Genomic variant data for the Jurkat cell line

<p>Variant calling data from whole-genome sequencing of the Jurkat cell line. The data set includes output files from four variant calling tools (jurkat_raw_variant_caller_output.tar.gz), final variant calls after filtering and merging the calls from the separate tools (jurkat_final_variant_calls.tar.gz), and files with variant effect information (jurkat_variant_effects.tar.gz).</p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

Genomic Resources for BIC@MSKCC Mouse Variant Pipeline

<p>Custom genomic resource files for Mouse Variant pipeline. Details and code are here: https://github.com/soccin/MusVar</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

A phased genome of the highly heterozygous 'Texas' almond uncovers patterns of allele-specific expression linked to heterozygous structural variants

<h2># Genomic datasets associated to the publication:&nbsp;</h2> <h3># Gene-ID conversion with previous genome version</h3> <p>Texasv3_vs_Texasv2_GeneID.txt -- gene ID conversion between Texasv3 and Texasv2 (https://www.rosaceae.org/analysis/295)</p> <p>pdulcis26_to_F1_liftoff_polished.gff3 -- Texasv2 gene annotation liftoff on Phase-1 assembly (Phase-1 coordinates)</p> <h3># Phase-1</h3> <p>Texas_F1_K80_chr.fasta &nbsp;-- genome asssembly, phase-1&nbsp;<br>Texas_F1_gene_models.gff3 -- &nbsp;phase-1 &nbsp;gene annotation (de novo annotation)<br>Functional_annotation_TexasF1.csv -- &nbsp;phase-1 gene functions &nbsp;<br>Texas_F1_ref_SV.vcf --- Structural variations relative to phase-0 (This file uses Phase-1 as reference)</p> <h3># Phase-0</h3> <p>Texas_F0_K80_chr.fasta -- genome asssembly, phase-0&nbsp;<br>Texas_F0_gene_models.gff3 -- &nbsp;phase-0 &nbsp;gene annotation (liftoff from Phase-1)<br>Functional_annotation_TexasF0.csv -- &nbsp;phase-0 &nbsp;gene functions &nbsp;<br>Texas_F0_ref_SV.vcf --- Structural variations relative to phase-1 (This file uses Phase-0 as reference)</p> <h3># Transposable element annotation</h3> <p>Texas_F0_HiConf_TE_v3.gff3 --- TE annotation in Phase-0<br>Texas_F1_HiConf_TE_v3.gff3 --- TE annotation in Phase-1<br>Texasv3_TElib.fa --- TE library of TexasV3 (non-redundant repeat consensuses taking the account the two genome phases)</p> <h3># Gene sequences in fasta</h3> <p>Transcript, CDS and protein sequences in fasta format for phase-0 (F0) and phase-1 (F1)&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Data and code for: A naturally occurring mitochondrial genome variant confers broad protection from infection in Drosophila (doi: https://doi.org/10.1101/2024.03.28.587162)

<p><strong>&nbsp;</strong>Raw data and R code for:</p> <p>Tiina S. Salminen, Laura Vesala, &nbsp;Yuliya Basikhina, Megan Kutzer, Tea Tuomela, Ryan Lucas, Katy Monteith, Arun Prakash, Tilman Tietz<sup> </sup>and Pedro F. Vale.&nbsp;A naturally occurring mitochondrial genome variant confers broad protection from infection in <em>Drosophila &nbsp;</em>(<strong>doi:</strong>&nbsp;https://doi.org/10.1101/2024.03.28.587162)</p> <p><em>&nbsp;</em></p> <p><strong>Data files:</strong></p> <p>CFU_stats.R: statistics for CFU data</p> <p>Prettgeri_CFU.csv: colony forming units (CFUs) in <em>P. rettgeri</em>-infected flies</p> <p>Saureus_CFU.csv: colony forming units (CFUs) in <em>S. aureus</em>-infected flies</p> <p>&nbsp;</p> <p>gene exp_mitotypes-stats.R: statistics for the qPCR data</p> <p>gene exp_mitotypes.csv: data on the relative gene expression of&nbsp;<em>FucTC</em>, <em>AANATL3</em>, <em>Acp1</em> and <em>CG3397</em> among mitotypes</p> <p>&nbsp;</p> <p>copynumber_stats.R: statistics for the copynumber data</p> <p>copynumber.csv: Mitochondrial copy number as copies of the mtDNA target gene <em>16S</em> relative to nuclear target gene <em>RpL32 </em></p> <p>&nbsp;</p> <p>respirometry_stats.R: statistics for the respirometry data</p> <p>respirometry.csv: Mitochondrial respiration measured in uninfected male flies by measuring oxygen consumption of the OXPHOS complexes I, III and IV</p> <p>&nbsp;</p> <p>ROS_stats.R: statistics for ROS</p> <p>ROS.csv: hydrogen peroxide (H<sub>2</sub>O<sub>2</sub>) quantification (whole flies)</p> <p>&nbsp;</p> <p>wasp_stats.R: Statistics on the proportion of melanised and unmelanised wasp larvae</p> <p>Wasps_data.csv: Data on proportion of melanised (killed) and unmelanised (living) <em>L. boulardi</em> parasitoid wasp larvae found in <em>Drosophila</em> larvae</p> <p>&nbsp;</p> <p>CC_stats.R: statistics on crystal cell counts</p> <p>CC.csv: amount of hemocytes called crystal cell in the larvae</p> <p>&nbsp;</p> <p>HC_stats.R: statistics on hemocyte counts</p> <p>hc_uninf.csv: total and activated hemocyte (blood cell) counts in uninfected larvae</p> <p>hc_inf.csv: total and activated hemocyte (blood cell) counts in infected larvae (48 h after infection by <em>L. boulardi</em> parasitoid wasps)</p> <p>&nbsp;</p> <p>PPO3.R: statistics on PPO3 expression levels</p> <p>PPO3.csv: data on gene expression level of prophenoloxidase (PPO3)</p> <p>&nbsp;</p> <p>MMP.R: statistics on MMP</p> <p>cybridMMP.csv: mitochondrial membrane potential (MMP) measured from larval hemocytes using the TMRM dye</p> <p>&nbsp;</p> <p>CellROX.R: statistics on CellROX<sup>TM </sup>signal</p> <p>CellROX.csv: Reactive oxygen species (ROS) measured from larval hemocytes using CellROX<sup>TM</sup> green reagent</p> <p>&nbsp;</p> <p>Hazard.Rmd: hazard ratios</p> <p>Combined.xlsx: data on hazard ratios Hazard ratios of survival post&nbsp;<em>P. rettgeri, S. aureus, </em>Kallithea and DCVinfections</p> <p>&nbsp;</p> <p>Mito_priming.R: Statistics on the effect of priming</p> <p>mito_priming.csv: data on the effect of immune priming with heat-killed bacteria prior to infection</p> <p>&nbsp;</p> <p>hc_stats_supplementary.R: statistics for sex-specific differences in hemocyte counts (data: hc_uninf.csv &amp; hc_inf.csv )</p> <p>PI_hc_stats.R: statistics on dead (Propidium iodide, PI-positive) vs. living (PI-negative) larval hemocytes (data: hc_uninf.csv &amp; hc_inf.csv )</p> <p>&nbsp;</p> <p>melanisation_stats.R: statistics on melanisation response</p> <p>melanisation_males.csv: melanisation response in hemolymph</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

LYCEUM: Learning to call copy number variants on low coverage ancient genomes

<p>Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduce additional noise into sequencing data; and finally, (iii) the typically low coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high coverage read-depth signals, underperform under such conditions. To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high- confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection.</p>

opencc-by-4.0Oct 2024View details →
dryad36/100

Large scale across-breed genome-wide association study reveals a variant in HMGA2 associated with inguinal cryptorchidism risk in dogs

<p class="MsoNormal"><span>Cryptorchidism is the most common congenital sex development disorder in dogs. Despite this, little progress has been made in understanding its genetic background. Extensive genetic testing of dogs through consumer and veterinary channels using a high-density SNP genotyping microarray coupled with links to clinical records presents the opportunity for a large-scale genome-wide association study to elucidate the molecular risk factors associated with cryptorchidism in dogs. Using an inter-breed genome-wide association study approach, a significant statistical association on canine chromosome 10 was identified, with the top SNP pinpointing a variant of <em>HMGA2 </em>previously associated with adult weight variance. In further analysis we show that incidence of cryptorchidism is skewed towards smaller dogs in concordance with the identified variant's previous association with adult weight. This study represents the first putative variant to be associated with cryptorchidism in dogs.</span></p>

opencc-zeroMay 2022View details →
zenodo36/100

Whole genome sequence variant discovery from Ethiopian Boran, N'Dama and Holstein cattle

<p>Whole genome sequence variants (SNPs) from forty samples of each Ethiopian Boran, N&#39;Dama and Holstein cattle breeds which were utilised in the&nbsp;Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle study by Riggio et al., 2022, (in review).&nbsp;The NDama sequences included samples from Guinea (n =21), Nigeria (n =10) and Senegal (n = 9); and the Boran samples from Ethiopia (n = 30) and Kenya (n = 10). Variants discovery followed the protocol&nbsp;described in the Material and Method section of the paper.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Project Adotto Whole-Genome Variants

<p>Squared-off project-level VCF on GRCh38 of 172 haplotype-resolved long-read assemblies from 86 samples of 78 individuals. Variants stored in compressed BCF and split by chromosome.&nbsp;Created as part of a&nbsp;project attempting to catalog tandem-repeat regions in Human genomes. Details on the project can be found on the <a href="https://github.com/ACEnglish/adotto">github</a>.</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024

<p>Expression quantitative trait loci (eQTLs) provide a key bridge between noncoding DNA sequence variants and organismal traits. The effects of eQTLs can differ among tissues, cell types, and cellular states, but these differences are obscured by gene expression measurements in bulk populations. We developed a one-pot approach to map eQTLs in <em>Saccharomyces cerevisiae</em> by single-cell RNA sequencing (scRNA-seq) and applied it to over 100,000 single cells from three crosses. We used scRNA-seq data to genotype each cell, measure gene expression, and classify the cells by cell-cycle stage. We mapped thousands of local and distant eQTLs and identified interactions between eQTL effects and cell-cycle stages. We took advantage of single-cell expression information to identify hundreds of genes with allele-specific effects on expression noise. We used cell-cycle stage classification to map 20 loci that influence cell-cycle progression. One of these loci influenced the expression of genes involved in the mating response. We showed that the effects of this locus arise from a common variant (W82R) in the gene <em>GPA1</em>, which encodes a signaling protein that negatively regulates the mating pathway. The 82R allele increases mating efficiency at the cost of slower cell-cycle progression and is associated with a higher rate of outcrossing in nature. Our results provide a more granular picture of the effects of genetic variants on gene expression and downstream traits.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Calibration of variant effect predictors on genome-wide data masks heterogeneous performance across genes

<p><strong>Supplemental files contain all necessary datasets to reproduce the analysis and figures in "Calibration of variant effect predictors on genome-wide data masks heterogeneous performance across genes" Code and additional files can be found at https://github.com/FowlerLab/VEP-calibrations</strong></p>

opencc-by-4.0Mar 2024View details →
dryad36/100

The role of structural variants in pest adaptation and genome evolution of the Colorado potato beetle, Leptinotarsa decemlineata (Say)

<p>Structural variation has been associated with genetic diversity and adaptation in diverse taxa. Despite these observations, it is not yet clear what their relative importance is for microevolution, especially with respect to known drivers of diversity, e.g., nucleotide substitutions, in rapidly adapting species. Here we examine the significance of structural variants (SVs) in pesticide resistance evolution of the agricultural super-pest, the Colorado potato beetle,<em> Leptinotarsa decemlineata</em>. By employing a parent offspring trio sequencing procedure, we develop highly contiguous reference genomes to characterize structural variation within this species. These updated assemblies represent &gt;100-fold improvement of contiguity and include derived pest and ancestral non-pest individuals. We identify &gt;200,000 SVs, which appear to be non-randomly distributed across the genome as they co-occur with transposable elements and genes. SVs intersect exons for a large proportion of gene annotations (~20%) and are associated with insecticide resistance, development, and transcription, most notably cytochrome P450 (CYP) genes. To understand the role that SVs might play in adaptation we measure allele frequencies of SVs for an additional 57 individuals, using whole genome resequencing data, representing pest and non-pest populations of North America. Incorporating multiple independent tests of significance using SNP data, we identify 14<strong> </strong>positively selected genes that include SVs and SNPs of elevated frequency within the sampled pest lineages. Among these, four are associated with insecticide resistance. One of these genes, glycosyltransferase-13, is a duplicated gene enclosed within a structural variant that resides inside the <em>CYP4g15</em> genic region. Both gene products have been observed to be co-induced during insecticide exposure. These results demonstrate the significance of structural variations as a genomic feature to describe species history, genetic diversity, and adaptation.</p>

opencc-zeroJun 2024View details →
zenodo36/100

Identifying genetic variants associated with chromatin looping and genome function

<p><span>Here<span> we present a comprehensive HiChIP dataset on na&iuml;ve CD4 T cells (nCD4) from 30 donors and identify QTLs that associate with genotype-dependent and/or allele-specific variation of HiChIP contacts defining loops between active regulatory regions (iQTLs). We observe a substantial overlap between iQTLs and previously defined eQTLs and histone QTLs, and an enrichment for fine-mapped QTLs and GWAS variants. Furthermore, we describe a distinct subset of nCD4 iQTLs, for which the significant variation of chromatin contacts in nCD4 are translated into significant eQTL trends in CD4 T cell memory subsets. Finally, we define connectivity-QTLs as iQTLs that are significantly associated with concordant genotype-dependent changes in chromatin contacts over a broad genomic region (e.g., GWAS SNP in the <em>RNASET2</em> locus). Our results demonstrate the importance of chromatin contacts as a complementary modality for QTL mapping and their power in identifying novel classes of QTLs linked to cell-specific gene expression and connectivity.</span></span></p> <p>&nbsp;</p> <p><span><span>This repository contains the source code, supplementary datasets for the manuscript (Nature Communications 2024).</span></span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

NGSAP-VC : Genomic Variant Calling as an Installable GALAXY Workflow Using NGS data.

<p>Implementation of genomic variants calling as an installable GALAXY workflows using NGS data. Repository contains two separate sets of simulated ebola test data. One for SNPs and INDELs calling and another for Structural Variants calling.</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

Canadian Arctic narwhal and bowhead whale genomic variants

<p>This dataset contains resequencing data used in our Arctic whale genomics research exaiming population structure and demographic history in bowhead whales and narwhal, including unfiltered genomic variants and filtered SNPs. Source code for genomic analyses is available at&nbsp;<a href="http://github.com/edegreef/NBW-resequencing">github.com/edegreef/arctic-whales-resequencing</a>. Data uploaded here contain:</p> <ul> <li><strong>bowhead_sample_info.csv</strong>&nbsp;- metadata for bowhead whale samples</li> <li><strong>bowhead_RWmap_allvariants.vcf.gz</strong>&nbsp;- all variant calls in bowhead whales, including indels and SNPs (n = 21)</li> <li><strong>bowhead_RWmap_snps.filtered.autosomes.vcf.gz&nbsp;</strong>- bowhead whale SNPs filtered for quality, bi-allelic sites, and autosomes (n = 21)</li> <li><strong>bowhead_RWmap_snps.filtered.autosomes.hwe_maf_LDprunedr08_n20.vcf.gz</strong>&nbsp;- bowhead whale SNPs filtered for quality, bi-allelic sites, out of HWE, MAF &gt; 0.05, LD-pruned, removal of close kin (n = 20).<br><br></li> <li><strong>narwhal_sample_info.csv</strong>&nbsp;- metadata for narwhal samples</li> <li><strong>narwhal_allvariants.vcf.gz</strong>&nbsp;- all variant calls in narwhals, including indels and SNPs for bowhead whales (n = 62*)</li> <li><strong>narwhal_snps.filtered.autosomes.vcf.gz&nbsp;</strong>- narwhal SNPs filtered for quality, bi-allelic sites, and autosomes (n = 62*)</li> <li><strong>narwhal_snps.filtered.autosomes.hwe_maf_LDprunedr08_n57.vcf.gz</strong>&nbsp;- narwhal SNPs filtered for quality, bi-allelic sites, out of HWE, MAF &gt; 0.05, LD-pruned, removal of duplicates, close kin, and samples with high missingness (n = 57)</li> </ul> <p>The asterisk* is to note this includes 2 duplicate samples that were removed from the analyses and respective manuscript.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Genome-wide comparison reveals large structural variants in the cassava landraces Authors

<p><span>Structural variants (SVs) are critical for plant genomic diversity and phenotypic variation. This study investigates a large, 9.7 Mbp highly repetitive segment on chromosome 12 of <em><span>TMEB117</span></em>, a region not previously characterized in cassava. We aim to explore its presence and variability across multiple cassava landraces, providing insights into its genomic significance and potential implications.</span></p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Mode-of-inheritance predictions for all possible missense variants in the human genome (hg38)

<p><span>Ensemble and consensus approaches to prediction of recessive inheritance for missense variants in human disease.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Assessing reproduciblity of Inherited Variants Detected with Short-read Whole Genome Sequencing

<p>This dataset is part of the following study:</p> <p><a href="https://www.fda.gov/science-research/bioinformatics-tools/microarraysequencing-quality-control-maqcseqc">https://www.fda.gov/science-research/bioinformatics-tools/microarraysequencing-quality-control-maqcseqc</a></p> <p>The raw sequencing data can be downloaded from SRA:</p> <p><a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA723125">https://www.ncbi.nlm.nih.gov/bioproject/PRJNA723125</a></p> <p>and NODE:</p> <p><a href="https://www.biosino.org/node/project/detail/OEP0018966">https://www.biosino.org/node/project/detail/OEP0018966</a></p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record