Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,363

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,363 results for “Mutations”

Learn how ShareScore rates datasets ↗
zenodo44/100

Data for: Regularized sequence-context mutational trees capture variation in mutation rates across the human genome

<p>Additional data on output models from Bayer as reported in:</p> <p>Regularized sequence-context mutational trees capture variation in mutation rates across the human genome</p> <p>Adams CJ, Conery M, Auerbach BJ, Jensen ST, Mathieson I, Voight BF. BioRxiv&nbsp;https://doi.org/10.1101/2022.10.14.512160</p> <p>Accepted, PLoS Genetics.&nbsp;</p> <p>Code Available at:&nbsp;https://github.com/bvoightlab/Baymer</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Genome-Wide Mutational Signatures of Aristolochic Acid and Its Application as a Screening Tool

<p>The .zip files contain the fastq files for the cell line data published in DOI: 10.1126/scitranslmed.3006086 plus results of analyzing these data. HK2_AA and HK2b-8d2 were exposed to AA (aristolochic acid I), and HK2_ctrl was an unexposed control. The .tsv files are VCF-like files with the mutations found in the exposed cells. spectrum_counts.txt are what are now called &quot;SBS96&quot; spectra -- counts of single base substitutions in the contexts of preceding and following bases. The PDF file has plots of these.&nbsp; (The HUC1 cells that were exposed did not show AA-related mutations.)</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2013View details →
zenodo44/100

Processed snRNA-seq data from "Divergent single cell transcriptome and epigenome alterations in ALS and FTD patients with C9orf72 mutation"

<p>Processed snRNA-seq data from &quot;Divergent single cell transcriptome and epigenome alterations in ALS and FTD patients with C9orf72 mutation&quot;. All nuclei passed QC and were corrected for background noise using cellBender.&nbsp;Files are in R objects saved in RDS (R Data Serialization) format. This repo contains one Seurat v4 object and one gene-by-cell&nbsp;raw RNA count matrix in sparse matrix format (dgCMatrix).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

AlphaFold structures reported in "AlphaFold2 Can Predict Single-Mutation Effects"

<p>This contains AlphaFold predictions for X proteins that are found in the Protein Data Bank (PDB), that were used to evalluate AlphaFold's predictions of mutation effects. This includes one set of structures predicted by AlphaFold2.0, using default settings, and one structure for each of 5 models. This also includes structures predicted by the ColabFold version of AlphaFold (6 recycles, 5 models, no template, amber minimization, 4 repeats).</p><p>There are also additional predicted structures that are found in the PDB that were not analyzed in the paper.</p><p>There are AlphaFold predictions for three proteins (BFP / RFP, GFP, and PafA), covering either all (BFP/RFP, PafA) or a subset (GFP) of the sequences in three datasets of phenotype measurements from high-throughput experiments.</p><p>Results are separated into tar files based on whether DeepMind &nbsp;(AF2.0) or ColabFold implementation was used.</p><p>Folders under "ColabFold/PDB" are labelled according to a sequence ID, since multiple PDB structures can exist for a single sequence. These sequence IDs can be mapped back to PDB IDs using the information in "seq_id_pdb_id.json".</p><p>All PDB files have been compressed using Foldcomp (<a href="https://github.com/steineggerlab/foldcomp">https://github.com/steineggerlab/foldcomp</a>). Foldcomp is required to decompress the ".fcz" files in order to recover the ".pdb" files.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Mutation testing of smart contracts at scale

<p>Replication package for the paper&nbsp;<a href="https://arxiv.org/abs/1909.12563">Mutation testing of smart contracts at scale</a>.</p> <p>It is crucial that smart contracts are tested thoroughly due to their immutable nature. Even small bugs in smart contracts can lead to huge monetary losses. However, testing is not enough; it is also im- portant to ensure the quality and completeness of the tests. There are already several approaches that tackle this challenge with mutation test- ing, but their effectiveness is questionable since they only considered small contract samples. Hence, we evaluate the quality of smart contract mutation testing at scale. We choose the most promising of the existing (smart contract specific) mutation operators, analyse their effectiveness in terms of killability and highlight severe vulnerabilities that can be in- jected with the mutations. Moreover, we improve the existing mutation methods by introducing a novel killing condition that is able to detect a deviation in the gas consumption, i.e., in the monetary value that is required to perform transactions.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Dataset related to the article "Clinical and Molecular Data Define a Diagnosis of Arrhythmogenic Cardiomyopathy in a Carrier of a Brugada-Syndrome-Associated PKP2 Mutation"

<p>This record contains raw data related to the article &ldquo;&nbsp;Molecular Data Define a Diagnosis of Arrhythmogenic Cardiomyopathy in a Carrier of a Brugada-Syndrome-Associated PKP2 Mutation&rdquo;.&nbsp;</p> <p>Plakophilin-2 (<em>PKP2</em>) is the most frequently mutated desmosomal gene in arrhythmogenic cardiomyopathy (ACM), a disease characterized by structural and electrical alterations predominantly affecting the right ventricular myocardium. Notably, ACM cases without overt structural alterations are frequently reported, mainly in the early phases of the disease. Recently, the&nbsp;<em>PKP2</em>&nbsp;p.S183N mutation was found in a patient affected by Brugada syndrome (BS), an inherited arrhythmic channelopathy most commonly caused by sodium channel gene mutations. We here describe a case of a patient carrier of the same BS-related&nbsp;<em>PKP2</em>&nbsp;p.S183N mutation but with a clear diagnosis of ACM. Specifically, we report how clinical and molecular investigations can be integrated for diagnostic purposes, distinguishing between ACM and BS, which are increasingly recognized as syndromes with clinical and genetic overlaps. This observation is fundamentally relevant in redefining the role of genetics in the approach to the arrhythmic patient, progressing beyond the concept of &quot;one mutation, one disease&quot;, and raising concerns about the most appropriate approach to patients affected by structural/electrical cardiomyopathy. The merging of genetics, electroanatomical mapping, and tissue and cell characterization summarized in our patient seems to be the most complete diagnostic algorithm, favoring a reliable diagnosis.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Data and analysis scripts associated with the paper 'Long-term experimental evolution of HIV-1 reveals effects of environment and mutational history''

<p><em>Eva Bons, Christine Leemann,&nbsp; Karin J. Metzner, Roland R. Regoes</em></p> <p>This repository contains all the data and analysis scripts associated with the paper &#39;Long-term experimental evolution of HIV-1 reveals effects of environment and mutational history&#39;</p> <p>See the readme after unpacking the .zip for a description of the files</p>

opencc-by-4.0Oct 2020View details →
dryad40/100

Allopatric divergence of cooperators confers cheating resistance and limits the effects of a defector mutation

<p>Studies of microbial social defectors that 'cheat' on cooperative genotypes generally focus on interactions with their cooperative parents, yet in nature defectors may meet diverse cooperators. Genotype-by-genotype interactions may constrain the ranges of cooperators upon which particular defectors can cheat, limiting the cheaters' spread and potentially the overall equilibrium frequency of cheaters. The bacterium Myxococcus xanthus undergoes cooperative multicellular development upon starvation, but some developmental defectors can cheat on cooperators, outcompeting them within mixed groups. We show that a defector disrupted at the signaling gene csgA has a narrow cheating range among diverse natural cooperators owing to antagonisms not specifically targeted at defectors. More strikingly, lab-evolved cooperators only slightly differentiated from the defector have allopatrically evolved beyond its cheating range by accumulating fewer than 20 mutations when development was not directly under selection. Cooperators might diversify not only with respect to which defectors cheat on them, but also in the potential for a particular mutation to reduce expression of cooperative trait or generate a cheating phenotype. We tested this by constructing a new csgA mutation in several highly diverged cooperators. The mutation generated very different sporulation phenotypes – from a complete defect to no defect – indicating that genetic background effects can limit the set of genomes for which a given mutation creates a defector and potentiates cheating. Our results suggest that natural populations feature geographic mosaics of cooperators diversified in susceptibility to cheating by any given defector and in the social phenotypes generated by any given mutation in a cooperation gene.</p>

opencc-zeroDec 2020View details →
zenodo40/100

Genome wide Illumina 450K array in AML patients with or without CEBPA mutation

<ul> <li>Num: Row Number</li> <li>Hybridization REF: CpG Reference ID</li> <li>Position: Position in hg19</li> <li>Gene: Entrez Gene Symbol</li> <li>Chrome: Chromosome Number</li> <li>p_vals: P value for t-test between patients with CEBPA and without CEBPA mutation</li> <li>mean_cebpa_mut: mean methylation score for patients with CEBPA mutation</li> <li>mean_no_cebpa_mut:&nbsp;mean methylation score&nbsp;for patients with CEBPA mutation&nbsp;</li> <li> <p>site_status:&nbsp;prediction status for CEBPA sites&nbsp;</p> </li> <li> <p>fdr:&nbsp;fdr p_value&nbsp;</p> </li> </ul>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Evolution enhances mutational robustness and suppresses the emergence of a new phenotype

<p>Source codes and data sets for&nbsp;arXiv:2012.03030. For details, see readme.txt file.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Data and material for the manuscript "Mutation testing and self/peer assessment: analyzing their effect on students in a software testing course"

<p><strong>This repository is composed of two different parts: </strong></p> <ul> <li><a href="https://zenodo.org/record/4464300/files/Assessment%20data%20and%20Mutation%20Scores.xlsx?download=1">Assessment data and Mutation Scores</a> file contains the student-generated data used in the experience.</li> <li><a href="https://zenodo.org/record/4464300/files/experience-material.zip?download=1">Experience-material</a>&nbsp;file contains the files to be able to reproduce the experience.</li> </ul> <p>&nbsp;</p> <p><strong>The </strong><strong> <a href="https://zenodo.org/record/4464300/files/experience-material.zip?download=1">Experience-material</a> file for the lab is used in two sessions:</strong></p> <p>Session 1: Development and assessment of test suites</p> <p>In this session, the student has to develop a test suite for a program under test. At the end of the session, the test suite will be evaluated against a set of assessment criteria regarding the quality of the developed test suite.</p> <p>Files for this session:</p> <ul> <li>VVS-Lab6-S1 pdf file , with the description of this session.</li> <li>Material-S1 zip file, with the files required to complete this session.</li> </ul> <p>Session 2: Evaluation applying mutation testing with MuCPP</p> <p>In this session, the test cases designed in the first part of this lab will be evaluated based on the mutation adequacy criterion. This will be done by using the&nbsp;<a href="https://ucase.uca.es/mucpp/">MuCPP mutation tool</a>.</p> <p>Files for this session:</p> <ul> <li>VVS-Lab6-S2 pdf file, with the description of this session.</li> <li>Material-S2 zip file, with the files required to complete this session.</li> </ul> <p><em>The source code files family.[cpp|hpp] have been adapted from a listing in [1]. Note that, while considered to be fault free in this lab, these source files are used in other sessions where students are expected to detect some defects in them.</em></p> <p>[1] S. Wiener and L. J. Pinson, The C++ Workbook. USA: Addison-Wesley Longman Publishing Co., Inc., 1990.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Early postzygotic mutations contribute to de novo variation in a healthy monozygotic twin pair.

<p>Human de novo single-nucleotide variation (SNV) rate is estimated to range between 0.82-1.70&times;10(-8) mutations per base per generation. However, contribution of early postzygotic mutations to the overall human de novo SNV rate is unknown.</p> <p>METHODS:</p> <p>We performed deep whole-genome sequencing (more than 30-fold coverage per individual) of the whole-blood-derived DNA samples of a healthy monozygotic twin pair and their parents. We examined the genotypes of each individual simultaneously for each of the SNVs and discovered de novo SNVs regarding the timing of mutagenesis. Putative de novo SNVs were validated using Sanger-based capillary sequencing.</p> <p>RESULTS:</p> <p>We conservatively characterised 23 de novo SNVs shared by the twin pair, 8 de novo SNVs specific to twin I and 1 de novo SNV specific to twin II. Based on the number of de novo SNVs validated by Sanger sequencing and the number of callable bases of each twin, we calculated the overall de novo SNV rate of 1.31&times;10(-8) and 1.01&times;10(-8) for twin I and twin II, respectively. Of these, rates of the early postzygotic de novo SNVs were estimated to be 0.34&times;10(-8) for twin I and 0.04&times;10(-8) for twin II.</p> <p>CONCLUSIONS:</p> <p>Early postzygotic mutations constitute a substantial proportion of de novo mutations in humans. Therefore, genome mosaicism resulting from early mitotic events during embryogenesis is common and could substantially contribute to the development of diseases.</p>

opencc-by-4.0Oct 2015View details →
zenodo40/100

Secretin-like class B G protein-coupled receptor (GPCR) mutation data set

<p>Curated set of&nbsp;2463 quantitative mutation data points&nbsp;covering 13 secretin-like class B&nbsp;G protein-coupled receptors.</p> <p>Annotated according to GPCRdb standards (http://gpcrdb.org/), complemented by GPCRdb Ballesteros-Weinstein numbers assigned&nbsp;using GPCRdb KNIME nodes (https://github.com/3D-e-Chem/knime-gpcrdb).</p> <p>The data set is build from data sets published in:</p> <p>- Siu et al. Nature 2013,&nbsp;499: 444-449. doi:10.1038/nature12393</p> <p>- Hollenstein, de Graaf&nbsp;et al. Tr Pharmacol Sci&nbsp;2014, 35: 12-22. doi: 10.1016/j.tips.2013.11.001</p> <p>- Yang, de Graaf et al. J Biol Chem 2016, 291: 12991-3004. doi: 10.1074/jbc.M116.721977</p>

opencc-zeroAug 2016View details →
zenodo40/100

Putative mutations associated with tetracycline resistance detected in Treponema spp.- an analysis of 4,355 Spirochaetales genomes

<p>Supplementary Table 1&ndash;4, , as well as multiple sequence alignment files for&nbsp;16S rRNA, rpsC, and rpsJ loci, have been deposited.</p> <p>Additionally, the archive contains a <code>scripts.txt</code> file that includes all custom Bash scripts used for:</p> <ul> <li> <p>rRNA gene detection and quantification using <em>Barrnap</em></p> </li> <li> <p>Genome quality assessment using <em>CheckM</em></p> </li> <li> <p>Gene-by-gene schema creation, allele calling, and cgMLST extraction using <em>chewBBACA</em></p> </li> <li> <p>GFF merging and wide-format transformation for summarizing rRNA copy number per genome</p> </li> </ul> <p>These materials are provided to ensure reproducibility of the analyses and to support further investigations into antimicrobial resistance in <em>Spirochaetales</em>.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Low-frequency somatic mutations are heritable in tropical trees Dicorynia guianensis and Sextonia rubra

<p>Somatic mutations potentially play a role in plant evolution, but common expectations pertaining to plant somatic mutation remain insufficiently tested. Unlike in most animals, the plant germline is assumed to be set aside late in development, leading to the expectation that plants accumulate somatic mutations along growth. Therefore, several predictions were made on the fate of somatic mutations: mutations have generally low frequency in plant tissues; mutations at high frequency have a higher chance of intergenerational transmission; branching topology of the tree dictates mutation distribution; and, exposure to UV radiation increases mutagenesis. To provide new insights into mutation accumulation and transmission in plants, we produced two high-quality reference genomes and a unique dataset of 60 high-coverage whole-genome sequences of two tropical tree species, <i>Dicorynia guianensis</i> (Fabaceae) and <i>Sextonia rubra </i>(Lauraceae). We identified 15,066 <i>de novo</i> somatic mutations in <i>D. guianensis</i> and&nbsp; 3,208 in <i>S. rubra</i>, surprisingly almost all found at low frequency. We demonstrate that: 1) low-frequency mutations can be transmitted to the next generation; 2) mutation phylogenies deviate from the branching topology of the tree; and 3) mutation rates and mutation spectra are not demonstrably affected by differences in UV exposure. Altogether, our results suggest far more complex links between plant growth, ageing, UV exposure, and mutation rates than commonly thought.</p>

opencc-by-4.0Nov 2023View details →
dryad40/100

Data from: The role of mutation bias in adaptive molecular evolution: insights from convergent changes in protein function

<p>An underexplored question in evolutionary genetics concerns the extent to which mutational bias in the production of genetic variation influences outcomes and pathways of adaptive molecular evolution. In the genomes of at least some vertebrate taxa, an important form of mutation bias involves changes at CpG dinucleotides: If the DNA nucleotide cytosine (C) is immediately 5' to guanine (G) on the same coding strand, and if the C is methylated, then C→T and G→A mutations occur at an elevated rate relative to mutations at non-CpG sites. Here we examine experimental data from case studies in which it has been possible to identify the causative substitutions that are responsible for adaptive changes in the functional properties of vertebrate hemoglobin (Hb). Specifically, we examine the molecular basis of convergent increases in Hb-O<sub>2</sub> affinity in high-altitude birds. Using a data set of experimentally verified, affinity-enhancing mutations in the Hbs of highland avian taxa, we tested whether causative changes are enriched for mutations at CpG dinucleotides relative to the frequency of CpG mutations among all possible missense mutations. The tests revealed that a disproportionate number of causative amino acid replacements were attributable to CpG mutations, demonstrating that mutation bias can influence outcomes of molecular adaptation.</p>

opencc-zeroNov 2023View details →
zenodo40/100

Characterization of a loss-of-function NAPB mutation in monozygotic triplets affected with epilepsy and autism using cortical neurons from proband-derived and CRISPR-corrected iPSC lines. Author names and affiliations

<p>RNA-seq data of matured cortical neurons (8-weeks old) derived from induced pluripoent stem cells (iPSC). There are three replicates (Rep1, Rep2, Rep3) for each sample&nbsp; with Forwad read (R1_001.fastq.gz) and reverse read (R2_001.fastq.gz).</p> <p>CtrlF:&nbsp; Control Father sample</p> <p>CtrlM: Control mother sample</p> <p>NDD_01: Proband sample</p> <p>NDD_04: Proband sample</p> <p>NDD_05: Proband sample</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

iGEMME Missense Mutational Effect Predictions for Entire Human Proteome

<p>This dataset contains iGEMME single point mutation predictions of about ~19000 human proteins. In iGEMME predictions, only evolutionary data coming from multiple sequence alignment files is used.&nbsp;</p> <h2>Description of the data and file structure</h2> <p>This dataset contains iGEMME predictions for all human proteins.&nbsp;</p> <p>Data of each human protein is in a folder named after its uniprotID. Inside uniprotID folder, there is a subfolder called results that contain all input and output. An example results folder for uniprotID A0A0B4J245 will contain the following files:</p> <ol> <li> <p><strong>Raw igemme predictions (output file):</strong> A0A0B4J245_normPred_evolCombi_igemme.txt</p> </li> <li> <p><strong>Ranksorted (between 0-1) igemme predictions in csv format (output file):</strong> A0A0B4J245_normPred_evolCombiTransposedRanksorted_igemme.csv</p> </li> <li> <p><strong>Colabfold MSA file (input file):</strong> aliA0A0B4J245.fasta</p> </li> <li> <p><strong>JET2 file containing JET scores for each amino acid (output file) :</strong> A0A0B4J245_jet_igemme.res</p> </li> <li> <p><strong>Configuration file containing default parameters (output file):</strong> default.conf</p> </li> <li> <p><strong>Log file (output file):</strong> igemme.log</p> </li> </ol>

opencc-by-nc-4.0Dec 2023View details →
dryad40/100

Data from: Deep mutational scanning of HBV reveals a mechanism for cis preferential reverse transcription

<p>Hepatitis B virus (HBV) is a small double-stranded DNA virus that chronically infects 296 million people. Over half of its compact genome encodes protein in two overlapping reading frames, and during evolution, multiple selective pressures can act on shared nucleotides. This study combines an RNA-based HBV cell culture system with deep mutational scanning to uncouple <em>cis-</em> and <em>trans</em>-acting sequence requirements in the HBV genome. The results support a leaky ribosome scanning model for polymerase translation, provide a fitness map of the HBV polymerase at single nucleotide resolution, and identify conserved prolines adjacent to the HBV polymerase termination codon that stall ribosomes. Further experiments indicated that stalled ribosomes tether the nascent polymerase to its template RNA, ensuring <em>cis</em>-preferential RNA packaging and reverse transcription of the HBV genome.</p>

opencc-zeroFeb 2024View details →
zenodo40/100

Dataset for Gas-centered Mutation Testing of Ethereum Smart Contracts

<p>This dataset contains additional material for the paper entitled "<em>Gas-centered Mutation Testing of Ethereum Smart Contracts</em>", published by&nbsp;Journal of Software: Evolution and Process. This material includes:</p> <ul> <li><strong>Mutation tool</strong>, with the implementation of our gas-centered mutation operators.</li> <li><strong>Mutation operators - Example</strong>, with the code excerpts used to illustrate the gas perturbation caused by the mutation operators.</li> <li><strong>Smart contract - Accelerator_0b58e2.dir</strong>, with the mutants generated in a smart contract and their execution results.</li> <li><strong>Smart contracts_features.xlsx,</strong> with the complete list of smart contracts and their characteristics.</li> </ul>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record