Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
91
datasets available to search
ShareScore release 0.9.0
Dataset results
91 results for “variant identification”
Variant calls for 'Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration'
<p>A VCF of SNP calls used for input to GWAS in https://elifesciences.org/articles/26255</p>
Identification and characterization of novel splice variants of human farnesoid X receptor
<p>Dataset related to publication:</p> <blockquote> <p>Mustonen E-K, Lee SML, Nieß H, Schwab M, Pantsar T, Burk O: "Identification and characterization of novel splice variants of human farnesoid X receptor". <em>Archives of Biochemistry and Biophysics</em> <a href="https://doi.org/10.1016/j.abb.2021.108893">https://doi.org/10.1016/j.abb.2021.108893</a></p> </blockquote> <p> </p> <p>Including:</p> <p>I. Full-length raw-trajectories of the 1 microsecond Desmond simulations (-out.cms, trj)</p>
Identification of genetic variants associated with anterior cruciate ligament rupture and AKC standard coat color in the Labrador Retriever
<p>Canine anterior cruciate ligament (ACL) rupture is a common complex disease. Prevalence of ACL rupture is breed-dependent. In an epidemiological study, yellow coat color was associated with increased risk of ACL rupture in the Labrador Retriever. ACL rupture risk variants may be linked to coat color through genetic selection or through linkage with coat color genes. To investigate these associations, Labrador Retrievers were phenotyped as ACL rupture cases or controls and for coat color and were single nucleotide polymorphism (SNP) genotyped. After filtering, ~697K SNPs were analyzed using GEMMA and mvBIMBAM for multivariate association. Functional annotation clustering analysis with DAVID was performed on candidate genes. A large 8Mb region on chromosome 5 that included <em>ACSF3</em>, as well as 32 additional SNPs, met genome-wide significance at <em>P</em><6.07E-7 or Log<sub>10</sub>(BF) = 3.0 for GEMMA and mvBIMBAM, respectively. On chromosome 23, SNPs were located within or near <em>PCCB</em> and <em>MSL2</em>. On chromosome 30, a SNP was located within <em>IGDCC3</em>. SNPs associated with coat color were also located within <em>ADAM9</em>,<em> FAM109B</em>,<em> SULT1C4</em>,<em>RTDR1</em>,<em> BCR</em>, and <em>RGS7</em>. <em>DZIP1L</em> was associated with ACL rupture. Several significant SNPs on chromosomes 2, 3, 7, 24, and 26 were located within uncharacterized regions or long non-coding RNA sequences. This study validates associations with the previous ACL rupture candidate genes <em>ACSF3</em> and <em>DZIP1L</em> and identifies novel candidate genes. These variants could act as targets for treatment or as factors in disease prediction modeling. The study highlighted the importance of regulatory SNPs in the disease, as several significant SNPs were located within non-coding regions.</p>
Comprehensive identification of pathogenic variants in retinoblastoma by long- and short-read sequencing
<p>Retinoblastoma (RB) is the most common intraocular malignancy in childhood. The causal variants in RB are mostly characterized by previously used short-read sequencing (SRS) analysis, which has technical limitations in identifying structural variants (SVs) and phasing information. Long-read sequencing (LRS) technology has significant advantages over SRS in detecting SVs, phased genetic variants, and methylation. In this study, we comprehensively characterized the genetic landscape of RB using combinatorial LRS and SRS of 16 RB tumors and 16 matched blood samples. We detected a total of 232 somatic SVs, with an average of 14.5 SVs per sample across the cohort. We identified 20 distinct pathogenic variants, including <span>three</span> novel small variants and <span>five</span> <span>somatic</span> SVs. Furthermore, our analysis shows that the vast majority (93.8%) of <em>RB1</em>-disrupted patients fit the two-hit hypothesis through diverse types, including the biallelic hypermethylated promoter as well as small and large compound heterozygous mutations which were missing in SRS analysis. By constructing the evolutionary history of all genetic variants, we reveal the evolution trends that <em>RB1</em> disruption early and followed by copy number changes, including amplifications of Chr2p, and deletions of Chr16q, during RB tumorigenesis. Altogether, we characterize the comprehensive genetic landscape of RB, providing novel insights into the genetic alterations and mechanisms contributing to RB initiation and development. Our work also establishes a framework to analyze genomic landscape of cancers based on LRS data. <br><br>We took IGV screenshots of SV and SNV/indels. The IGV naming conventions for SV is: sample name, chromosome 1, start, chromosome 2, end, SV type, SV length, SV breakpoint. The IGV naming convention for SNV/indels is: sample name, chromosome, position, reference allele, alternate allele.</p>
Identification of genetic variants associated with clinical features of sickle cell disease
Open the record for dataset details and reuse information.
Training data for "Identification of allelic variants in SARS-CoV-2 from deep sequencing reads"
<p>Effectively monitoring global infectious disease crises, such as the COVID-19 pandemic, requires capacity to generate and analyze large volumes of sequencing data in near real time. These data have proven essential for monitoring the emergence and spread of new variants, and for understanding the evolutionary dynamics of the virus.</p> <p>Two sequencing platforms in combination with several established library preparation strategies are predominantly used to generate SARS-CoV-2 sequence data. However, data alone do not equal knowledge: they need to be analyzed. The Galaxy community developed analysis workflows to support the <strong>identification of allelic variants (AVs) in SARS-CoV-2 from deep sequencing reads</strong>.</p> <p>These workflows allow one to identify AVs and lineages in SARS-CoV-2 genomes with variant allele frequencies ranging from 5% to 100% (i.e., they detect variants with intermediate frequencies as well.</p> <p>In this tutorial we will see how to run these workflows for the different types of input data:</p> <ul> <li>Single end data derived from Illumina-based RNAseq experiments</li> <li>Paired end data derived from Illumina-based RNAseq experiments</li> <li>Paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols</li> <li>ONT fastq files generated with Oxford nanopore (ONT)-based Ampliconic (ARTIC) protocols</li> </ul> <p>To illustrate the tutorial, we took some example datasets (paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols) from COG-UK, the COVID-19 Genomics UK Consortium.</p>
Supplementarty Data: Retention time and fragmentation predictors increase confidence in identification of common variant peptides
<p>Supplementary data related to the paper "Retention time and fragmentation predictors increase confidence in identification of common variant peptides"</p> <p>The database directory contains the FASTA files of the four protein sequence databases used in the analysis of the paper.</p> <p>The data directory contains exports with lists of all peptide-to-spectrum matches obtained from the analysis.</p> <p>More information and scripts to reproduce the post-processing steps are available at: https://github.com/ProGenNo/VariantPeptideIdentification</p>
Identification of genetic variants associated with anterior cruciate ligament rupture and AKC standard coat color in the Labrador Retriever
Open the record for dataset details and reuse information.
Validation of Optical Genome Mapping for the Identification of Constitutional Genomic Variants in a Postnatal Cohort
ClinicalTrials.gov study NCT05295277. IPD Sharing: NO. Countries: 1. Publications: 7.
Identification of Genetic Variants Associated With Unexpected Infant Death Syndrome
ClinicalTrials.gov study NCT06244433. IPD Sharing: Not stated. Countries: 1. Publications: 1.
SVXplorer: three-tier approach to identification of structural variants via sequential recombination of discordant cluster signatures
<p>This repository contains the bgzipped variants calls in VCF format for CHM1, NA12878 and AJ trio dataset that are used in the SVXplorer manuscript. The names of the files contain the name of the sample (CHM1/NA12878/HG002/HG003/HG004), the name of the method (SVXplorer/DELLY/LUMPY/TIDDIT/TARDIS/MANTA) used to call the variants. There are three separate files for the DELLY calls which have the deletions, duplications and the inversion calls made by DELLY for each of the samples. For NA12878, there are two sets of calls, one for each of the libraries (ERR194147/SRR505885)</p> <p> </p>
Data from: Identification of novel variants in LTBP2 and PXDN using whole-exome sequencing in developmental and congenital glaucoma
Open the record for dataset details and reuse information.
Identification and functional impact of genomic copy number variants in zebrafish, an important human disease model (Zebrafish Strain CNVs) (CGH ZV81M 2)
GEO Series GSE28278. Danio rerio. 76 samples. Type: Genome variation profiling by genome tiling array.
Massively parallel identification of cis-regulatory variants in yeast promoters - Experimental measurements
GEO Series GSE155943. Saccharomyces cerevisiae. 36 samples. Type: Other.
Massively parallel identification of cis-regulatory variants in yeast promoters
GEO Series GSE155944. Escherichia coli; Saccharomyces cerevisiae. 38 samples. Type: Other.
Identification of host dependency factors shared by multiple SARS-CoV-2 variants of concern
GEO Series GSE207981. Homo sapiens. 48 samples. Type: Expression profiling by high throughput sequencing.
Systematic identification of genotype-dependent enhancer variants in eosinophilic esophagitis and atopic dermatitis
GEO Series GSE232341. Homo sapiens. 39 samples. Type: Expression profiling by high throughput sequencing; Other.
Identification of functional non-coding variants associated with orofacial cleft [MPRA]
GEO Series GSE297145. Homo sapiens. 8 samples. Type: Other.
Identification of Breast Cancer Associated Variants That Modulate Transcription Factor Binding
GEO Series GSE89013. Homo sapiens. 5 samples. Type: Other.
Direct identification of hundreds of expression-modulating variants using a multiplexed reporter assay
GEO Series GSE75661. Homo sapiens; synthetic construct. 31 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.