Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,363
datasets available to search
ShareScore release 0.7.1
Dataset results
6,363 results for “Mutations”
Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Data From: PD-linked LRRK2 G2019S mutation impairs astrocyte morphology and synapse maintenance via ERM hyperphosphorylation
<p>Source data files for Wang et al 'PD-linked LRRK2 G2019S mutation impairs astrocyte morphology and synapse maintenance via ERM hyperphosphorylation'</p>
Joint host-pathogen genomic analysis identifies hepatitis B virus mutations associated with human NTCP and HLA class I variation
<p>Summary statistics for "Joint host-pathogen genomic analysis identifies hepatitis B virus mutations associated with human NTCP and HLA class I variation" </p><p>Files are organized in the following directory structure:</p><p><strong>G2G/</strong> - Summary statistics of G2G associations (SNPs, HLA, and gene-level analysis). </p><p><strong>HLA/ </strong>- Peptide binding prediction results</p><p><strong>preS1_haplotypes/ - </strong>Resolved intra-host haplotypes of the preS1 binding region. </p><p><strong>DnDs/</strong> - Calculation of intra-host positive selection, within the preS1 binding region. </p><p> </p>
Rate-enhancing PETase mutations determined through DFT/MM molecular dynamics simulations†
<p>Raw data for classical MD simulations ran with Gromacs 2018.3 for the two mutants Asp83Asn and Asp89Asn.</p><p>Raw data for quantum mechanics/molecular mechanics simulations ran with CP2K 6.1 for the two mutants Asp83Asn and Asp89Asn.</p><p>Distance and free energy analysis from the QM/MM MD simulations for the wild-type and the two mutants Asp83Asn and Asp89Asn.</p>
Binding of Cholesterol to the N-terminal Domain of the NPC1L1 Transporter: Analysis of the Epimerisation-Related Binding Selectivity and Loop Mutations
<p>Input files, topologies and trajectories of the work "Binding of Cholesterol to the N-terminal Domain of the NPC1L1 Transporter: Analysis of the Epimerisation-Related Binding Selectivity and Loop Mutations". </p>
ESCOTT mutational effect predictions for ProteinGym Substitutions Dataset with Colabfold MSAs
<p>This dataset includes all escott calculations for ProteinGym dataset and the calculations were performed with v1.6.0 of PRESCOTT package. </p> <p>ProteinGym dataset contains 72 proteins and there are 87 experiments performed on these 72 proteins.</p> <p>This dataset is organized into three folders:<br>1-predictions<br>2-analyses<br>3-figures-and-csv-files</p> <p>'predictions' folder has 72 subdirectories, one for each protein.<br>Let's use BLAT as an example to show the most essential files in each protein folder inside 'predictions'.<br>You can find the following files in BLAT_ECOLX_full_11 directory:<br> *aliBLAT.fasta (input file): This is the MSA file used for the calculations and it was obtained with Colabfold.<br> *ranked_0.pdb (input file): This is the pdb model used to deduce structural parameters by escott algorithm.<br> *BLAT_ECOLX_Stiffler_2015.mut (input file): This file is a list of mutations in simple text format. <br> *BLAT_jet.res (data file): This file contains trace(tjet), cv and pc parameters for each amino acid in a protein. This data is produced and used by escott.<br> *BLAT_ECOLX_Stiffler_2015_normPred_evolCombi.txt (output file): This file contains raw (unprocessed) escott predictions. This file is the most important output of escott algorithm.<br> *BLAT_ECOLX_Stiffler_2015_singleline.csv (experimental data file): We compare our escott predictions to the experimental measurements given in this file.<br> <br> <br>'analyses' folder contains 3 types of analyses that we used in our study:<br>1-secondary-structure-analysis<br>2-escott-averages-analysis<br>3-amino-acid-type-analysis</p> <p>'figures-and-csv-files' folder contains png files and their data in csv format.</p>
Generative AI in the Advancement of Viral Therapeutics for Predicting and Targeting Immune-Evasive SARS-CoV-2 Mutations
<p>This dataset <strong>encompasses</strong> and describes the following features:</p> <ul> <li>Mutations in viruses like SARS-CoV-2 can make them escape vaccines and treatments.</li> <li>Accurately predicting these mutations is crucial for developing effective countermeasures.</li> <li>The study uses a type of AI called a Generative Adversarial Network (GAN) to analyze the virus's spike protein, which plays a key role in infection.</li> <li>The GAN generates protein sequences similar to natural ones, but which are also likely to evade immune responses.</li> <li>By analyzing these generated sequences, the researchers improve their AI model's ability to predict real-world escape mutations.</li> <li>This improved prediction could help design better vaccines and treatments, and prepare for future viral threats.</li> </ul>
Artifact for the paper "Mutation-based Lifted Repair of Software Product Lines"
<p>This work presents a novel lifted repair algorithm for program families (Software Product Lines - SPLs) based on code mutations.<br>The inputs of our tool are an erroneous SPL and a specification given in the form of assertions. We use variability encoding to transform the given SPL into a single program, called family simulator, which is translated into a set of SMT formulas whose conjunction is satisfiable iff the simulator (i.e. the input SPL) violates an assertion. We use a predefined set of mutations applied to feature and program expressions of the given SPL. The tool repeatedly mutates the erroneous family simulator and checks if it becomes (bounded) correct. The outputs of our tool are all minimal repairs in the form of minimal number of (feature and program) expression replacements such that the repaired SPL is (bounded) correct with respect to a given set of assertions.</p> <p>We present a prototype tool for repairing #ifdef-based C programs (i.e., annotative SPLs). The experimental results show that our approach is able to successfully repair various interesting SPLs.</p>
Data and code for: Assessing the Reliability of Point Mutation as Data Augmentation for Deep Learning with Genomic Data
<p>Data and code for the paper "Assessing the Reliability of Point Mutation as Data Augmentation for Deep Learning with Genomic Data".</p>
Lack of shared neoantigens in prevalent mutations in cancer
<p><span>Tumors are mostly characterized by genetic instability, as result of mutations in surveillance mechanisms, such as DNA damage checkpoint, DNA repair machinery and mitotic checkpoint. Defect in one or more of these mechanisms causes additive accumulation of mutations. Some of these mutations are drivers of transformation and are positively selected during the evolution of the cancer, giving a growth advantage on the cancer cells. If such mutations would result in mutated neoantigens, these could be actionable targets for cancer vaccines and/or adoptive cell therapies. However, the results of the present analysis show, for the first time, that the most prevalent mutations identified in human cancers do not express mutated neoantigens. The hypothesis is that this is the result of the selection operated by the immune system in the very early stages of tumor development. At that stage, the tumor cells characterized by mutations giving rise to highly antigenic non-self mutated neoantigens would be efficiently targeted and eliminated. Consequently, the outgrowing tumor cells cannot be controlled by the immune system, with an ultimate growth advantage to form large tumors embedded in an immunosuppressive tumor microenvironment (TME). The outcome of such a negative selection operated by the immune system is that the development of off-the-shelf vaccines, based on shared mutated neoantigens, does not seem to be at hand.</span></p>
Data for: A new threshold selection method for species distribution models with presence-only data: extracting the mutation point of the P/E curve by threshold regression
<p>Selecting thresholds to convert continuous predictions of species distribution models proves critical for many real-world applications and model assessments. Prevalent threshold selection methods for presence-only data require unproven pseudo-absence data or subjective researchers' decisions. This study proposes a new method, Boyce-Threshold Quantile Regression (BTQR), to determine thresholds objectively without pseudo-absence data. We summarize that the mutation point is a typical shape feature of the predicted-to-expected (P/E) curve after reviewing relevant articles. Analysis based on source-sink theory suggests that this mutation point may represent a transition in habitat types and serve as an appropriate threshold. Threshold regression is introduced to accurately locate the mutation point.</p> <p>To validate the effectiveness of BTQR, we used four virtual species of varying prevalence and a real species with reliable distribution data. Six different species distribution models were employed to generate continuous suitability predictions. BTQR and nine other traditional methods transformed these continuous outputs into binary results. Comparative experiments show that BTQR has advantages in terms of accuracy, applicability, and consistency over the existing methods.</p>
Data required for "Low mutation rate of spontaneous mutants enables detection of causative genes by comparing whole genome sequences"
<p>In the early 1900s,mutation breeding to select varieties with desirable traits using spontaneous mutation was actively conducted around the world, including Japan. In rice, the number of fixed mutations per generation was estimated to be 1.38-2.25. Although this low mutation rate was a major problem for breeding in those days, in the modern era with the development of NGS technology, it was conversely considered to be an advantage for efficient gene identification. In this paper, we proposed an in silico approach using next-generation sequencing (NGS) to compare the whole genome sequence of a spontaneous mutant with that of a closely related strain with a nearly identical genome, to find polymorphisms that differ between them, and to identify the causal gene by predicting the functional variation of the gene caused by the polymorphism. Using this approach, we found four causal genes for the dwarf mutation, the round shape grain mutation and the awnless mutation. Three of these genes were the same as those previously reported, but one was a novel gene involved in awn formation. The novel gene was isolated from Bozu-Aikoku, a mutant of Aikoku with the awnless trait, in which nine polymorphisms were predicted to alter gene function by their whole-genome comparison. Based on the information on gene function and tissue-specific expression patterns of these candidate genes, Os03g0115700/LOC_Os03g02460, annotated as a shortchain dehydrogenase/reductase SDR family protein, is most likely to be involved in the awnless mutation. Indeed, complementation tests by transformation showed that it is involved in awn formation. Thus, this method is an effective way to accelerate genome breeding of various crop species by enabling the identification of useful genes that can be used for crop breeding with minimal effort for NGS analysis.</p>
Data for: Viral receptor-binding protein evolves new function through mutations that cause trimer instability and functional heterogeneity
<p>When proteins evolve new activity, a concomitant decrease in stability is often observed because the mutations that confer new activity can destabilize the native fold. In the conventional model of protein evolution, reduced stability is considered a purely deleterious cost of molecular innovation because unstable proteins are prone to aggregation and are sensitive to environmental stressors. However, recent work has revealed that non-native, often unstable protein conformations play an important role in mediating evolutionary transitions, raising the question of whether instability can itself potentiate the evolution of new activity. We explored this question in a bacteriophage receptor binding protein (RBP) during host-range evolution. We studied the properties of the RBP of bacteriophage before and after host-range evolution and demonstrated that the evolved protein is relatively unstable and may exist in multiple conformations with unique receptor preferences. Through a combination of structural modeling and in vitro oligomeric state analysis, we found that the instability arises from mutations that interfere with trimer formation. This study raises the intriguing possibility that protein instability might play a previously unrecognized role in mediating host-range expansions in viruses.</p>
Project Section: Searching for Mutations Leading to Isoniazid Resistance in TB (HackBio)
<p><strong>Introduction</strong></p> <p><span>Tuberculosis (TB) remains a global health concern, and the emergence of drug-resistant strains poses a significant challenge to effective treatment. Isoniazid is a key first-line drug used in the treatment of TB, and resistance to this drug can compromise the success of therapy. </span></p> <p><span>This project aims to identify and characterise genetic mutations associated with isoniazid resistance in </span><em>Mycobacterium tuberculosis</em><span>, the causative agent of TB. </span></p> <p><strong>Datasets</strong><span>: Here</span></p> <p><strong>Task:</strong></p> <ul> <li><span>Given the datasets above, perform the complete variant calling and annotation pipeline to specifically identify the following known mutations in </span><em>Mycobacterium tuberculosis</em><span> (against Isoniazid). </span></li> <li><span>Provide pretty IGV visualisation of mutations in </span><em>katG</em><span> sequences (Position: 2156111 - 2153889) across all the samples. N/B: </span><em>katG</em><span> mutations are prevalent in isoniazid resistance.</span></li> <li><span>Also, explore associated biological pathways to understand the molecular mechanisms of isoniazid resistance.</span></li> </ul> <p><strong><span>Expected Output:</span></strong></p> <p><span>Kindly submit a comprehensive project report, including methods, results, and interpretations.</span></p>
Data from: Prophage maintenance is determined by environment-dependent selective sweeps rather than mutational availability
<p>Prophages, viral sequences integrated into bacterial genomes, can confer fitness benefits and costs. Despite the risk of prophage activation and subsequent bacterial death, active prophages are present in most bacterial genomes. However, our understanding about the selective forces that maintain prophages in bacterial populations is scarce. Combining experimental evolution with stochastic modelling we found that prophage maintenance and loss are primarily determined by environmental conditions that amplify the net fitness effect of the prophage. Whole genome sequencing revealed that prophage loss occurs through environment-specific sequences of selective sweeps, leading to rapid prophage loss when prophages are costly. However, conflicting selection pressures that select against the prophage but for a prophage-encoded accessory gene can prolong prophage maintenance. The extent of prolonged maintenance depends on the sociality of this accessory gene. Selection for non-cooperative genes is more effective for prophage maintenance as cooperative genes allow for protection of phage-free 'cheaters' that may emerge if prophage costs outweigh their benefits. Our mathematical model suggests that environmental variation plays a larger role than mutation rates in determining prophage maintenance. This challenges our understanding of the role of random chance events relative to environmental factors in shaping the evolutionary trajectory of bacterial populations.</p>
Fitness landscape of substrate-adaptive mutations in evolved APC transporters
<p>Growth rate calculations:</p> <p>Single colonies of <em>S. cerevisiae</em> Δ10AA pADHXC3GH-<em>GOI</em> were inoculated in YB media supplemented with 4 mm NH<sub>4</sub><sup>+</sup> and 0.1 mg/ml ampicillin, and grown until late logarithmic phase. The cultures were pelleted at 750 × <em>g</em> for 10 min at 30 °C and washed with YB media. The wells in the microplate were filled with the amino acids of interest to a final concentration of 2 mM and with culture cells to a final OD<sub>600</sub> of 0.04, to a final total well volume of 200 µl. Sterile water was added in the space between the wells to avoid evaporation. The prepared microplates included three biological replicates of the strains with the plasmid containing the gene of interest (GOI) and one biological replicate of the strain with the empty vector. The absorbance in each well was measured at 600 nm in 30 min intervals without shaking of the microplate, at 30 °C for 72 h in a SpectraMax ABS Plus plate reader. The data sets (CSV files) are the raw optical density readings from the growth assays, along with plate layout metadata (CSV files). The growth rates were derived based on the Baranyi growth model, using the <em>growthrates</em> package in <em>R</em>.</p> <p>Included transporter genes:</p> <p>Mutated transporters (script growth_rates_mutants_Baranyi2.r)</p> <p><em>AGP1</em>, <em>AGP1</em>-N, <em>AGP1</em>-V,<em> AGP1</em>-NV, <em>AGP1</em>-G, <em>AGP1</em>-T, <em>PUT4</em>, <em>PUT4</em>-S</p> <p>Wild-type transporters (script growth_rates_wild_types_Baranyi2.r)</p> <p><em>AGP1</em>, <em>BAP2</em>, <em>CAN1</em>, <em>HIP1</em>, <em>LYP1</em>, <em>MMP1</em>, <em>PUT4</em></p> <p> </p> <p>Plasmid sequences (Genbank files):</p> <p>pADHXC3GH-AGP1, pADHXC3GH-BAP2, pADHXC3GH-CAN1, pADHXC3GH-HIP1, pADHXC3GH-LYP1, pADHXC3GH-MMP1, pADHXC3GH-PUT4-S L207S, pADHXC3GH-PUT4, pADHXC3GH</p> <p> </p> <p>Measureing relative membrane fluorescence (all transporters in this study have C-terminal GFP tags):</p> <p>Micrographs were analyzed with ImageJ using the script analyse_cell_perimeter.ijm</p>
Transcriptomic Analysis Data for MSTN Mutations and Mechanisms of Muscle Hypertrophy in a New Guinea Pig Breed
<p>This dataset contains raw RNA-seq data from six guinea pig muscle samples, split into two groups:</p> <ul> <li><strong>Native guinea pigs (B1 to B3):</strong> Control group with no selective breeding.</li> <li><strong>Kuri breed guinea pigs (B4 to B6):</strong> Synthetic hybrid group selectively bred for increased muscle mass.<br>Each sample has paired-end FASTQ files (e.g., B1_1.fq.gz and B1_2.fq.gz).</li> </ul>
The Cellular and Extra-Cellular Proteomic Signature of Human Dopaminergic Neurons Carrying the LRRK2 G2019S Mutation
<p>The provided data sets were generated during the work described in "The Cellular and Extra-Cellular Proteomic Signature of Human Dopaminergic Neurons Carrying the LRRK2 G2019S Mutation" published in Frontiers in Neuroscience, 2024.</p> <p>They contain the results from our differential expression analysis of our raw DIA Protemic Data as well as a list of input data for GO analyses.</p>
Mitochondrial oxidant stress promotes alpha-synuclein aggregation and spreading in mice with mutated glucocerebrosidase
<p>These are the data sets and plasmid sequences of the viral vectors used in the study "Mitochondrial oxidant stress promotes alpha-synuclein aggregation and spreading in mice with mutated glucocerebrosidase"</p>
PRESCOTT/ESCOTT/iGEMME mutational effect predictions of all single point mutations for ~3000 proteins
<p>This dataset contains all necessary data to reproduce our analyses on ~3000 proteins:</p> <p>It is made up of 4 compressed datasets. </p> <ol> <li>All colabfold MSAs and structures for ~3000 proteins: colabfold-sequences-structures-3000-proteins.tar.bz2</li> <li>All escott prediction: escott-v-1-6-0-max-two-components-colabfold-msas-entire-single-point-mutations-cvRC7.tgz</li> <li>All igemme predictions: escott-v-1-6-0-tjet-only-colabfold-msas-entire-single-point-mutations.tgz</li> <li>All gnomad v4.0.0 csv files used for prescott predictions: gnomadv4-0-0-csv-files.tgz</li> </ol> <p>The dataset contains list of 500 proteins used in determining PRESCOTT coefficients in TRAINING-all_gene_names_v4_set1_no_acmg.txt. </p> <p>Furthermore, the list of 1883 proteins used to measure for testing purpose is given in TESTING-all_gene_names_v4_set1_with_acmg.txt. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.