Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,880
datasets available to search
ShareScore release 0.9.0
Dataset results
2,880 results for “Variant.”
Adsorption patterns of the WT, delta and omicron variants onto hydrophobic and hydrophilic surfaces
<p>Adsorption of the WT, delta and omicron variants onto hydrophobic and hydrophilic surfaces. This contains 3 replicas of each.</p>
Maverick Variant Pathogenicity Data Resources
<p>MAVERICK is a Mendelian Approach to Variant Effect pRedICtion built in Keras. It classifies protein-altering variants as either dominant disease-causing, recessive disease-causing, or benign. Here, we provide the pre-computed scores for all missense and nonsense SNVs in Gencode Basic V33 on GRCh37 and lifted over to GRCh38 as well as the datasets on which MAVERICK was trained and primarily evaluated. </p>
Perturbations in fatty acid metabolism and collagen production infer pathogenicity of a novel MBTPS2 variant in Osteogenesis imperfecta
<p>(i) PCR-sequencing: Chromatograms were generated by PCR-sequencing of a region within exon 4 of MBTPS2 using DNA extracted from a healthy control and the proband's fibroblasts</p> <p>(ii) Gene expression was quantified by qRT-PCR using RNA extracted from fibroblasts. Transcript levels of each gene of interest was calculated using the 2^-deltaCt method with normalization to the average Ct values of endogenous control genes GAPDH, IPO8 and TBP.</p> <p>(iii) Cellular fatty acid content was quantified by GC-MS/MS. Each table represents one technical replicate. Absolute values of each fatty acid are listed in the tables; relative ratios of various fatty acids are calculated at the bottom of each table. </p> <p>(iv) Immunocytochemistry images of ECM proteins (COL1 = collagen type I; COL4 = collagen type IV; COL5 = collagen type V; a2b1 = integrin a2b1) and binding of collagen-hybridising peptide (R-CHP).</p>
BRCA1-specific machine learning model predicts variant pathogenicity with high accuracy - Supplementary material
<p>Figure S1: Distribution of the reviewed 141 <em>BRCA1</em> missense variants; Figure S2: The Shapely values for the <em>BRCA1</em> XGBoost models; Figure S3: The Shapely values of the <em>BRCA1</em> XGBoost model used to predict the functional assays’ results for variants of uncertain significance; Table S1: The receiver operating characteristic (ROC) curve analysis for the different in silico predictions; Table S2: Cross validation of the BRCA1 model in 5 different random training and test samples; Table S3: Pathogenicity prediction and prioritization of the 31,058 unreviewed BRCA1 variants from the BRCA Exchange database.</p>
Naturally segregating variants contributing to thermal tolerance in a D. melanogaster model system.
<p>Main_Incapacitation.zip and Incapacitation_founders.zip contain raw thermal tolerance scores for individuals measured within the heat box. Each folder is labeled with the RIL or founder ID and replicates within each file are labeled with group numbers. </p> <p>RNAi_files_to_tar.txt contains the metadata for the Combined_tracks_RNAi_1.Rds.zip and Combined_tracks_RNAi_2.Rds.zip.</p> <p>Combined_tracks_RNAi_1.Rds.zip and Combined_tracks_RNAi_2.Rds.zip. contains raw data for RNAi lines measured on the heat plate. </p> <p>plate_finder-kinglab-2021-05-02.zip contains the DeepLabCut model used for finding the corners of aluminum mounting plate used to hold the fly vials for thermal sensitivity testing. This directory contains the training data as well as the trained and evaluated model. No retraining should be necessary for use.</p> <p>fly_tracker_2-king-2021-09-27.zip contains the DeepLabCut model used for tracking individual flies during thermal sensitivity testing. This directory contains the training data as well as the trained and evaluated model. No retraining should be necessary for use.</p> <p>fly_tracker_batch.py is a python (>= 3.0) script that processes the raw movie files collected via the Raspberry Pi. This script uses the plate finder DeepLabCut model to find the corners of the plate, rotate and crop the images, and output movie files for individual flies. It then uses the fly tracker DeepLabCut model to track the flies and output the data for subsequent processing in R.</p>
CPT-1 pre-computed whole-proteome variant effect predictions and model source code
<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p><p>This repository contains the variant effect predictions of CPT-1 for 18,602 human proteins, initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects". The proteins are split into three files.</p><p><i>CPT1_score_EVE_set.zip</i>: Proteins in the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>)</p><p><i>CPT1_score_no_EVE_set_1.zip</i> & <i>CPT1_score_no_EVE_set_2.zip</i>: Proteins not in the EVE set. Predictions for these proteins use imputed values for features depending on the EVE MSA.</p><p>The protein names are UniProt gene names.</p><p>We also provide source code to train CPT-1 model and reproduce results in the manuscript :</p><p><i>source_code.zip </i>(corresponds to GitHub repository songlab-cal/CPT version as of Jul 12, 2023)</p><p> </p><p><strong>Citation</strong></p><p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br>"Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p><p>*These authors contributed equally to this work.<br>†To whom correspondence should be addressed: <a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p><p>DOI: <a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p><p> </p>
Rare Variants Prioritized by the Multio-omic Watershed Model
<p>These datasets include genomic annotations included in the Multi-omic Watershed model, trained using data from 1,319 individuals from the Multi-Ethnic Study of Atherosclerosis (MESA) cohort. The resulting rare genetic variants prioritized by the model were provided with their posteriors in each omic dimension (RNA expression, methylation, splicing, and protein expression). </p> <p>For details, please refer to https://www.biorxiv.org/content/10.1101/2022.09.07.507008v1.abstract </p>
Alliance of Genome Resources Sequence Variants
<p>Variant Call Format (VCF) formatted spreadsheets of sequence variants associated with phenotypic alleles from the Alliance of Genome Resources. Variants are in Human Genome Variation Society (HGVS) nomenclature syntax.</p> <p>Files include variants in VCF format for the following organisms:</p> <ul> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> </ul>
Patient-specific induced pluripotent stem cell properties implicate Ca2+-homeostasis in clinical arrhythmia associated with combined heterozygous RYR2 and SCN10A variants
Open the record for dataset details and reuse information.
Quantifying CpG variants in a pied flycatcher population
Open the record for dataset details and reuse information.
Pathogenic and low frequency variants in children with central precocious puberty
Open the record for dataset details and reuse information.
Variant calling in the Goldilocks Zone: how reference genome choice and read mapping stringency impact heterozygosity estimates and phylogenetic analyses
Open the record for dataset details and reuse information.
Variant discovery in full-sibling families of Pinus taeda L
Open the record for dataset details and reuse information.
Data from: Introgressed variants obscure phylogenetic relationships but are not subject to positive selection in Australasian long-tailed parrots
Open the record for dataset details and reuse information.
In search of the genetic variants of human sex ratio at birth: Was Fisher wrong about sex ratio evolution?
Open the record for dataset details and reuse information.
Negative linkage disequilibrium between amino acid changing variants reveals interference among deleterious mutations in the human genome
Open the record for dataset details and reuse information.
Data and source code for: ClinVar and HGMD genomic variant classification accuracy has improved over time, as measured by implied disease burden
Open the record for dataset details and reuse information.
Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants
Open the record for dataset details and reuse information.
Variant call file for mountain yellow-legged frog (MYLF) selection analysis
Open the record for dataset details and reuse information.
ZENODO TEST: ALE Variant Analysis data and scripts
<p>ALE Variant Analysis data and scripts</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.