Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

244

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

244 results for “genomic variants”

Learn how ShareScore rates datasets ↗
zenodo48/100

PopDel identifies medium-size deletions jointly in tens of thousands of genomes - Variant call sets

<p>This data set contains the variant calls sets generated by different tools for the benchmarks in the paper <a href="https://www.nature.com/articles/s41467-020-20850-5">PopDel identifies medium-size deletions simultaneously in tens of thousands of genomes</a>. It includes the VCFs/BCFs for the following test cases:</p> <ul> <li>Random deletion simulation on up to 1000 chromosome 21 samples</li> <li>1000 Genomes Project deletions inserted into simulated chromosomes 17 to 22 of up to 500 samples</li> <li>HG001 (NA12878)</li> <li>Trio of <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/">HG002</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG003_NA24149_father/NIST_HiSeq_HG003_Homogeneity-12389378/">HG003</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG004_NA24143_mother/NIST_HiSeq_HG004_Homogeneity-14572558/">HG004</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Diversity-Cohort">Polaris Diversity cohort</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Kids-Cohort">Polaris Kids cohort</a></li> </ul> <p>Further, the long and short read reference call sets for HG001 are provided. For HG002 the reference call set and the high confidence regions by the Genome in a Bottle consortium are provided.</p> <p>For details on how the files have been created, please refer to the paper and the script repository on <a href="https://github.com/kehrlab/PopDel-scripts">GitHub</a>.</p>

opencc-by-4.0Aug 2020View details →
zenodo48/100

Variant, Metabolite and Source Data for: Population genomics uncover loci for trait improvement in the indigenous African cereal tef (Eragrostis tef)

<p>These files contain the variant and metabolome for a collection of 220 tef (<em>Eragrsotis tef)</em> accessions from an ethiopian diversity panel. The accessions were assembled and managed by the Ethiopian Institute of Agricultural Research (EIAR, Ethiopia). The variant data was produced at the John Innes Centre (UK). The metabolome data was produced at Aberystwyth University (UK). These dataset are described in Jones et al. (2024), <em>bioRxiv</em>, https://doi.org/10.1101/2024.09.30.615331. The source data for main figures in the publication are also included.</p> <p>The submission contains</p> <ol> <li>EIAR_filtered.vcf.gz: This is the variant data obtained from alignment of Illumina reads from all 220 teff accessions to the reference assembly of tef (Dabbi). &nbsp;Low quality variants were filtered out. This variant data was used for constructing the phylogenetic relationship between the accessions. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>pooled_EIAR_filtered.vcf.gz: After the phylogentic analysis described above, reads from accessions that were found to be genetically redundant were pooled before variant calling. This file was used for the SNP GWAS analysis. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>&nbsp;Metabolite_Profile.xlxs (source data for Figure 5): This file contains m/z feature intensities from untargeted metabolite fingerprinting using Flow Infusion Electrospray High-resolution Mass Spectrometry (FIE-HRMS). The sample names contains a combination of Location code and Plot number in Supplementary Table S10 e.g AT plot 1, CD plot 1, DZ plot 1, where AT, CD and DZ represent Alem Tena, Chefe Donsa and Debre Zeit, respectively. The data was used for the partial least squares discriminant analysis and differentially accumulated metabolites analysis presented in Figure 5.</li> <li>Source data: Numerical source data for graphs and charts in Figures 3 - 7.</li> <li>Tsedey TT2 Sequence from Improved Assembly: The 4A and 4B sequences around the TT2 orthologue in tef from the improved PacBio-based chromosome-scale assembly of tef. These sequences were used for plotting the LTR Copia alignments presented in Supplementary Figure 9. We thank Corteva for pre-publication access to this improved Tsedey genome assembly.</li> </ol>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Analysis of variant-dependent m6A modifications within the Human genome

<p>Interactive and machine-readable results produced by the&nbsp;<a href="https://github.com/cumbof/m6Ad-SNVs" target="_blank" rel="noopener">m6Ad-SNVs</a> tool to asses if m6A-distal SNVs affect DRACH site accessibility, specifically by evaluating the alteration of base-pairing of nucleotides within segments of the DRACH motif.</p> <p>These results contain the predicted m6Ad-SNV candidates with the length of the reference and m6Ad-SNV-containing alternate sequences limited to 250 base pairs. This constraint has been applied to maintain the reliability of the results predicted by RNAFold (<a href="https://www.tbi.univie.ac.at/RNA/">ViennaRNA</a> package). The sequence composition contains up to 100 base pairs from 3'UTRs, with the remaining base pairs limited to the last two exons.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.

openmit-licenseApr 2024View details →
zenodo44/100

Functional genomics analysis to disentangle the role of genetic variants in major depression - Supplementary Tables

<p>This entry contains the data generated by the study &quot;Functional genomics analysis to disentangle the role of genetic variants in major depression&quot; that are part of the Supplementary information of&nbsp;the article describing the study.</p> <p>The entry contains the following data:</p> <p><strong>Supplementary Tables S1-S7</strong></p> <p>Supplementary Table S1. Summary of resources.</p> <p>Supplementary Table S2. Causal GVs for MD.</p> <p>Supplementary Table S3. pGenes functional and disease enrichment analysis.</p> <p>Supplementary Table S4. Fine-mapped MD causal GVs disease enrichment analysis.</p> <p>Supplementary Table S5. Colocalizing GWAS-eQTLs association to disease.</p> <p>Supplementary Table S6. TFBS analysis.</p> <p>Supplementary Table S7. GVs state annotation.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Variant dataset and code for "Population-level whole genome sequencing of Ascochyta rabiei identifies genomic loci associated with isolate aggressiveness"

<p>This dataset contains genetic variants (SNPs) of <em>Ascochyta rabiei</em> isolates and the R code used in their analysis to generate the results and figures described in the manuscript "<strong>Population-level whole genome sequencing of <em>Ascochyta rabiei</em> identifies genomic loci associated with isolate aggressiveness</strong>".</p> <div> <div>&nbsp;</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Data For: Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data

<p>Simulation output and Genome-wide scan for nIBD variants in UK10K data as reported in:</p> <p>Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data</p> <p>Johnson KE, Adams CJ, Voight BF. Methods Ecol Evol 2022 Nov;13(11):&nbsp;2429&ndash;2442.</p> <p>Code available at:&nbsp;https://github.com/kelsj/EVICORD</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Variant calls for 'Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration'

<p>A VCF of SNP calls used for input to GWAS in https://elifesciences.org/articles/26255</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance

<p>The below data are associated with our paper entitled "Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance."</p> <p>1) SNP-gene link predictions generated by pgBoost and existing methods SCENT (Sakaue et al. 2024 <em>Nat Genet</em>), Signac (Stuart et al. 2021 <em>Nat Methods</em>), ArchR (Granja et al. 2021 <em>Nat Genet</em>), and Cicero (Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><strong>pgBoost_scores.tsv.gz </strong>contains linking predictions made by pgBoost.</p> <p><strong>constituent_method_scores.tsv.gz</strong> contains linking predictions made by constituent methods.</p> <p><em><span>**NOTE: promoters (+/- 1kb from TSS) and candidate links &gt;500kb are excluded from linking predictions (see manuscript)**</span></em></p> <p>Linking scores and percentiles are reported for each method (pgBoost score, SCENT FDR, Signac correlation, ArchR correlation, Cicero co-accessibility). Rank percentiles are computed as: 1 - (rank / n). When multiple links receive the same score, they are assigned the percentile of the top rank. Links unscored by each method (denoted by zeros* in the linking score column) are assigned a percentile equivalent to the percent of links unscored by the focal method. See the Methods section of the paper for further details on computing linking scores and summarizing scores across cell types and data sets.</p> <p>*Candidate links tested and assigned a co-accessibility of zero by the Cicero method are given a score of 1e-100 in the "Cicero" column to distinguish between unscored candidate links and candidate links assigned a partial correlation of zero (see Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><em>NOTE: The predictions associated with this release (version 2) were generated using an expanded set of data sets, an expanded training set, and corrected TSS coordinates.</em></p> <p>2) GWAS-derived evaluation SNP-gene link evaluation set.</p> <p><strong>gwas_evaluation.tsv</strong>: GWAS-derived evaluation SNP-gene link evaluation set. Column 1 provides SNP coordinates in the format &lt;chr-start-end&gt;. This evaluation framework was proposed by Weeks et al. 2024 <em>Nature Genetics</em> based on fine-mapping results from Kanai et al. <em>medrxiv</em>&nbsp;(see Methods: <em>Evaluation data sets</em> of Dorans et al.). "True" links (gold = 1) are non-coding variants fine-mapped to a focal trait (PIP &gt; 0.1) with a coding variant for exactly one candidate gene within 1 Mb&nbsp;attaining PIP &gt; 0.5 for the same trait. "False" links (gold = 0) are candidate SNP-gene pairs involving a SNP with a "true" link. This file of SNP-gene links was adapted from credible set-gene links <a href="https://github.com/Deylab999/GWAS_benchmark_IGVF/blob/bb91d08cc02d59cdd829eb1430569057ac26c5fe/V2G/ENCODE_E2G_2023/UKBiobank.ABCGene.anyabc.tsv">here</a> (the "truth" column defines true/false links) by identifying SNPs with PIP &gt; 0.1 within each credible set-gene link.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Genome-wide analysis identified candidate variants and genes associated with heat stress adaptation in Egyptian sheep breeds

<p>The current study was conducted from 2009 to 2019 in three hot and dry agroecological zones in Egypt: Western Desert coastal zone, New Valley desert oasis, and hot-dry Upper Egypt. Within these zones, three local sheep breeds were studied: Barki (83 ewes), Wahati (55 ewes) and Saidi (68 ewes). During the study period, the animals exercised under natural heat stress (simulating summer grazing on poor pasture). Meteorological and physiological parameters were measured and recorded. The heat tolerance index of the animals was calculated to identify animals with high and low heat tolerance based on the animals&#39; response to the five main physiological parameters (scale from 0 to 5). DNA samples were extracted for genomic analysis. The genetic diversity measurements showed a significant influence of breed and location on the populations. The influence of breed is more significant than that of location. The inbreeding analysis shows that the desert breeds (Wahati and Barki) have lower values than the urban breed (Saidi). The high rate of sub-clustering indicates the process of sub-population through inbreeding pressure. Wahati and Barki are very distinct breeds with strong identification, while Saidi breed has crosses with other breeds. The most significant SNPs associated with heat tolerance were found in MYO5A, PRKG1, GSTCD, and RTN1 genes (P &lt; 0.0001). MYO5A had an effect of 0.74 on the trait heat tolerance in the studied population. It produces a protein that is widely distributed in the melanin-producing neural crest of the skin. Genetic association between genetic and phenotypic variations showed that OAR1 18300122.1, located in ST3GAL3, had the greatest positive effect on heat tolerance. GWAS analysis identified SNPs associated with heat tolerance in the PLCB1, STEAP3, KSR2, UNC13C , PEBP4, and GPAT2 genes.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data for publication: A pipeline for in-depth analysis of DNA virus populations by profiling the low abundant virus variants and partial genomic components

<p>Raw and processed sequence data from Oxford Nanopore and BGI short read sequencing platforms used in the publication: "A pipeline for in-depth analysis of DNA virus populations by profiling the low abundant virus variants and partial genomic components".</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Canadian Arctic killer whale genomic variants

<p>This dataset contains resequencing data used in our killer whale genomics research exaiming population structure and demographic history, including unfiltered genomic variants and filtered SNPs. Source code for genomic analyses is available at <a href="http://github.com/edegreef/NBW-resequencing">github.com/edegreef/killerwhale-resequencing</a>. Data uploaded here contain:</p> <ul> <li><strong>orca_sample_info.csv</strong>&nbsp;- metadata for the killer whale samples</li> <li><strong>orca_unfiltered_diploid.vcf.gz</strong> - all variant calls, including indels and SNPs (n = 29)</li> <li><strong>orca_snps_q30_biallelic.vcf.gz </strong>- SNPs filtered for quality and bi-allelic sites (n = 29)</li> <li><strong>orca_snps_q30_biallelic_HWE0.005_miss0.4.vcf.gz</strong> - SNPs filtered for quality, bi-allelic sites, out of HWE, and missingness &gt; 0.4 (n = 29).</li> <li><strong>orca_snps_q30_biallelic_HWE0.005_miss0.4_maf0.05_LDprunedr08_n24.vcf.gz</strong> - SNPs further filtered for MAF, LD-pruned, and removal of kin &amp; duplicates (n = 24).</li> </ul>

opencc-by-4.0May 2024View details →
zenodo40/100

Rare Genomic Copy Number Variants Implicate New Candidate Genes for Bicuspid Aortic Valve

<p>Whole genome genotyping data in dbGAP format and copy number variant calls in dbVar format.</p> <p>dbVar data includes CNV calls from cases with early onset bicuspid aortic valve disease (EBAV), cases from the International BAV Consortium (BAVCon), and controls from the dbGAP Wisconsin Longitudinal Study on Aging dataset (WLS).</p> <p>dbGAP files are divided into 12 batches of genotypes from EBAV subjects (EBAV1-12):</p> <p>1) PLINK output files (.map and .ped)</p> <p>2) GenomeStudio Final Report files</p> <p>3) One master pedigree file</p> <p>4) dbGAP subject mapping files</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.

<p>Ebola genome is manually mutated to contain non-structural as well as structural variants. One ebola genome contains non-structural variants - 10 SNPs, 10 INDELs, 05 TRANSLOCATIONs, 05 INSERSIONs and their reverse complements. Similarly, other two set of mutated genomes contain structural variants. Each set contains seven mutated ebola genome each one for large deletion, insertion, duplication, translocation, inversion, complex variant1 (consecutive three mutations - insertion, duplication and deletion) and complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Alliance of Genome Resources Sequence Variants

<p>Variant Call Format (VCF) formatted spreadsheets of sequence variants associated with phenotypic alleles from the Alliance of Genome Resources. Variants are in Human Genome Variation Society (HGVS) nomenclature syntax.</p> <p>Files include variants in VCF format for the following organisms:</p> <ul> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> </ul>

opencc-by-4.0Jul 2023View details →
dryad40/100

Variant calling in the Goldilocks Zone: how reference genome choice and read mapping stringency impact heterozygosity estimates and phylogenetic analyses

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad40/100

Negative linkage disequilibrium between amino acid changing variants reveals interference among deleterious mutations in the human genome

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad40/100

Data and source code for: ClinVar and HGMD genomic variant classification accuracy has improved over time, as measured by implied disease burden

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad40/100

Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants

Open the record for dataset details and reuse information.

publicOct 2021View details →
zenodo36/100

Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.

<p>Ebola genome is manually mutated to contain non-structural variants - 10 SNPs, 10 INDELs, 05 TRANSLOCATIONs, 05 INSERSIONs and their reverse complements. Paired-end illumina RNASEQ reads are simulated in fastq format.&nbsp;</p>

opencc-by-4.0Sep 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record