Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

91

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

91 results for “variant identification”

Learn how ShareScore rates datasets ↗
zenodo44/100

Variant calls for 'Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration'

<p>A VCF of SNP calls used for input to GWAS in https://elifesciences.org/articles/26255</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Identification and characterization of novel splice variants of human farnesoid X receptor

<p>Dataset related to publication:</p> <blockquote> <p>Mustonen E-K, Lee SML, Nie&szlig; H, Schwab M, Pantsar T, Burk O:&nbsp;&quot;Identification and characterization of novel splice variants of human farnesoid X receptor&quot;.&nbsp;<em>Archives of Biochemistry and Biophysics</em>&nbsp;<a href="https://doi.org/10.1016/j.abb.2021.108893">https://doi.org/10.1016/j.abb.2021.108893</a></p> </blockquote> <p>&nbsp;</p> <p>Including:</p> <p>I. Full-length raw-trajectories of the 1 microsecond Desmond simulations (-out.cms, trj)</p>

opencc-by-4.0Aug 2020View details →
dryad36/100

Identification of genetic variants associated with anterior cruciate ligament rupture and AKC standard coat color in the Labrador Retriever

<p>Canine anterior cruciate ligament (ACL) rupture is a common complex disease. Prevalence of ACL rupture is breed-dependent. In an epidemiological study, yellow coat color was associated with increased risk of ACL rupture in the Labrador Retriever. ACL rupture risk variants may be linked to coat color through genetic selection or through linkage with coat color genes. To investigate these associations, Labrador Retrievers were phenotyped as ACL rupture cases or controls and for coat color and were single nucleotide polymorphism (SNP) genotyped. After filtering, ~697K SNPs were analyzed using GEMMA and mvBIMBAM for multivariate association. Functional annotation clustering analysis with DAVID was performed on candidate genes. A large 8Mb region on chromosome 5 that included <em>ACSF3</em>, as well as 32 additional SNPs, met genome-wide significance at <em>P</em>&lt;6.07E-7 or Log<sub>10</sub>(BF) = 3.0 for GEMMA and mvBIMBAM, respectively. On chromosome 23, SNPs were located within or near <em>PCCB</em> and <em>MSL2</em>. On chromosome 30, a SNP was located within <em>IGDCC3</em>. SNPs associated with coat color were also located within <em>ADAM9</em>,<em> FAM109B</em>,<em> SULT1C4</em>,<em>RTDR1</em>,<em> BCR</em>, and <em>RGS7</em>. <em>DZIP1L</em> was associated with ACL rupture. Several significant SNPs on chromosomes 2, 3, 7, 24, and 26 were located within uncharacterized regions or long non-coding RNA sequences. This study validates associations with the previous ACL rupture candidate genes <em>ACSF3</em> and <em>DZIP1L</em> and identifies novel candidate genes. These variants could act as targets for treatment or as factors in disease prediction modeling. The study highlighted the importance of regulatory SNPs in the disease, as several significant SNPs were located within non-coding regions.</p>

opencc-zeroOct 2023View details →
zenodo36/100

Comprehensive identification of pathogenic variants in retinoblastoma by long- and short-read sequencing

<p>Retinoblastoma (RB) is the most common intraocular malignancy in childhood. The causal variants in RB are mostly characterized by previously used short-read sequencing (SRS) analysis, which has technical limitations in identifying structural variants (SVs) and&nbsp;phasing information. Long-read sequencing (LRS) technology has significant advantages over SRS in detecting SVs, phased genetic variants, and methylation. In this study, we comprehensively characterized the genetic landscape of RB using combinatorial LRS and SRS of 16 RB tumors and 16 matched blood samples. We detected a total of 232 somatic SVs, with an average of 14.5 SVs per sample across the cohort.&nbsp;We identified 20 distinct pathogenic variants, including <span>three</span>&nbsp;novel small variants and <span>five</span>&nbsp;<span>somatic</span>&nbsp;SVs. Furthermore, our analysis shows that the vast majority (93.8%) of <em>RB1</em>-disrupted patients fit the two-hit hypothesis through diverse types, including the biallelic hypermethylated promoter as well as small and large compound heterozygous mutations which were missing in SRS analysis. By constructing the evolutionary history of all genetic variants, we&nbsp;reveal the evolution trends that <em>RB1</em> disruption early and followed by copy number changes, including amplifications of Chr2p, and deletions of Chr16q, during RB tumorigenesis. Altogether, we characterize the comprehensive genetic landscape of RB, providing novel insights into the genetic alterations and mechanisms contributing to RB initiation and development. Our work also establishes a framework to analyze genomic landscape of cancers based on LRS data.&nbsp;<br><br>We took IGV screenshots of SV and SNV/indels. The IGV naming conventions for SV is: sample name, chromosome 1, start, chromosome 2, end, SV type, SV length, SV breakpoint. The IGV naming convention for SNV/indels is: sample name, chromosome, position, reference allele, alternate allele.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Identification of genetic variants associated with clinical features of sickle cell disease

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo36/100

Training data for "Identification of allelic variants in SARS-CoV-2 from deep sequencing reads"

<p>Effectively monitoring global infectious disease crises, such as the COVID-19 pandemic, requires capacity to generate and analyze large volumes of sequencing data in near real time. These data have proven essential for monitoring the emergence and spread of new variants, and for understanding the evolutionary dynamics of the virus.</p> <p>Two sequencing platforms in combination with several established library preparation strategies are predominantly used to generate SARS-CoV-2 sequence data. However, data alone do not equal knowledge: they need to be analyzed. The Galaxy community developed analysis workflows to support the <strong>identification of allelic variants (AVs) in SARS-CoV-2 from deep sequencing reads</strong>.</p> <p>These workflows allow one to identify AVs and lineages in SARS-CoV-2 genomes with variant allele frequencies ranging from 5% to 100% (i.e., they detect variants with intermediate frequencies as well.</p> <p>In this tutorial we will see how to run these workflows for the different types of input data:</p> <ul> <li>Single end data derived from Illumina-based RNAseq experiments</li> <li>Paired end data derived from Illumina-based RNAseq experiments</li> <li>Paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols</li> <li>ONT fastq files generated with Oxford nanopore (ONT)-based Ampliconic (ARTIC) protocols</li> </ul> <p>To illustrate the tutorial, we took some example datasets (paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols) from COG-UK, the COVID-19 Genomics UK Consortium.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Supplementarty Data: Retention time and fragmentation predictors increase confidence in identification of common variant peptides

<p>Supplementary data related to the paper &quot;Retention time and fragmentation predictors increase confidence in identification of common variant peptides&quot;</p> <p>The database directory contains the FASTA files of the four protein sequence databases used in the analysis of the paper.</p> <p>The data directory contains exports with lists of all peptide-to-spectrum matches obtained from the analysis.</p> <p>More information and scripts to reproduce the post-processing steps are available at: https://github.com/ProGenNo/VariantPeptideIdentification</p>

opencc-by-4.0Aug 2023View details →
dryad36/100

Identification of genetic variants associated with anterior cruciate ligament rupture and AKC standard coat color in the Labrador Retriever

Open the record for dataset details and reuse information.

publicNov 2023View details →
ClinicalTrials.gov32/100

Validation of Optical Genome Mapping for the Identification of Constitutional Genomic Variants in a Postnatal Cohort

ClinicalTrials.gov study NCT05295277. IPD Sharing: NO. Countries: 1. Publications: 7.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Identification of Genetic Variants Associated With Unexpected Infant Death Syndrome

ClinicalTrials.gov study NCT06244433. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

SVXplorer: three-tier approach to identification of structural variants via sequential recombination of discordant cluster signatures

<p>This repository contains the bgzipped variants calls in&nbsp;VCF format for CHM1, NA12878 and AJ trio dataset that are used in the SVXplorer manuscript. The names of the files contain the name of the sample (CHM1/NA12878/HG002/HG003/HG004), the name of the method (SVXplorer/DELLY/LUMPY/TIDDIT/TARDIS/MANTA) used to call the variants. There are three separate files for the DELLY calls which have the deletions, duplications and the inversion calls made by DELLY for each of the samples. For NA12878, there are two sets of calls, one for each of the libraries (ERR194147/SRR505885)</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
dryad28/100

Data from: Identification of novel variants in LTBP2 and PXDN using whole-exome sequencing in developmental and congenital glaucoma

Open the record for dataset details and reuse information.

publicDec 2016View details →
geo24/100

Identification and functional impact of genomic copy number variants in zebrafish, an important human disease model (Zebrafish Strain CNVs) (CGH ZV81M 2)

GEO Series GSE28278. Danio rerio. 76 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenDec 2011View details →
geo24/100

Massively parallel identification of cis-regulatory variants in yeast promoters - Experimental measurements

GEO Series GSE155943. Saccharomyces cerevisiae. 36 samples. Type: Other.

openGEO-OpenAug 2020View details →
geo24/100

Massively parallel identification of cis-regulatory variants in yeast promoters

GEO Series GSE155944. Escherichia coli; Saccharomyces cerevisiae. 38 samples. Type: Other.

openGEO-OpenAug 2020View details →
geo24/100

Identification of host dependency factors shared by multiple SARS-CoV-2 variants of concern

GEO Series GSE207981. Homo sapiens. 48 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2024View details →
geo24/100

Systematic identification of genotype-dependent enhancer variants in eosinophilic esophagitis and atopic dermatitis

GEO Series GSE232341. Homo sapiens. 39 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenNov 2023View details →
geo24/100

Identification of functional non-coding variants associated with orofacial cleft [MPRA]

GEO Series GSE297145. Homo sapiens. 8 samples. Type: Other.

openGEO-OpenJun 2025View details →
geo24/100

Identification of Breast Cancer Associated Variants That Modulate Transcription Factor Binding

GEO Series GSE89013. Homo sapiens. 5 samples. Type: Other.

openGEO-OpenDec 2016View details →
geo24/100

Direct identification of hundreds of expression-modulating variants using a multiplexed reporter assay

GEO Series GSE75661. Homo sapiens; synthetic construct. 31 samples. Type: Other.

openGEO-OpenJun 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record