Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,363

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,363 results for “Mutations”

Learn how ShareScore rates datasets ↗
zenodo48/100

Overcoming Limitation of AlphaFold2 by Deep-mutational Scanning and Stability-Selection of Protein Sequences

<p>This repository contains the processed datasets and corresponding code used in our study. While AlphaFold2 revolutionizes protein structure prediction, its accuracy critically depends on evolutionary information from natural homologs&mdash;limiting applications for proteins with sparse sequence families. Here, we bypass this bottleneck by employing deep mutational scanning and stability-guided selection to generate artificial homologs. Fed into AlphaFold2, these synthetic sequences match the accuracy achieved on well-predicted proteins with rich natural homology, while providing highly accurate predictions for difficult targets&mdash;including orphan proteins previously deemed "unpredictable." Our approach achieves high accuracy (&lt;3 &Aring; RMSD for 5/8 and &lt;2 &Aring; RMSD for 7/8 targets after excluding intrinsically flexible regions). Thus, integrating simple, scalable molecular biology (mutagenesis/selection) with high-throughput sequencing can deliver the accuracy similar to but at a fraction of the cost and time of traditional experimental structure-determination methods. This hybrid framework could democratize high-resolution structural biology, opening avenues to determine structures of protein complexes, modified proteins, and condition-dependent conformations.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo48/100

Characterization of a loss-offunction NSF attachment protein beta mutation in monozygotic triplets affected with epilepsy and autism using cortical neurons from proband-derived and CRISPR-corrected induced pluripotent stem cell lines

<p>RNA-seq data of matured cortical neurons (8-weeks old) derived from the induced pluripoent stem cells (iPSC) of control parents (CtrlF and CtrlM) and corrected proband. There are three replicates (Rep1, Rep2, Rep3) for each sample&nbsp; with Forwad read (R1_001.fastq.gz)</p> <p>CtrlF:&nbsp; Control Father sample</p> <p>CtrlM: Control mother sample</p> <p>NDD_01_Corr_Het: Heterozygous correction of NAPB mutation (c.354+2T&gt;G) in NDD_01 proband</p> <p>NDD_05_Corr_Hom: Homozygous correction of NAPB mutation (c.354+2T&gt;G) in NDD_05 proband</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

An estimate of fitness reduction from mutation accumulation in a mammal allows assessment of the consequences of relaxed selection: Dataset

<p>Supplementary files (data and analysis) for "An estimate of fitness reduction from mutation accumulation in a mammal allows assessment of the consequences of relaxed selection"</p> <p>Supplementary File 1: C3H_pheno_fix_Jun7_2023_nolowmut.csv</p> <p>Data for all mice in MA experiment including: mouse ID, sire, dam, generation, mating ID, sex, weight at 3 weeks, weight at 6 weeks, tail length, litter size, litter ID, line ID</p> <p>&nbsp;</p> <p>Supplementary File 2: C3H_pheno_Kontrol_June2023.csv</p> <p>Data for all control mice including: mouse ID, sire, dam, generation, mating ID, sex, weight at 3 weeks, weight at 6 weeks, tail length, litter size, litter ID, line ID</p> <p>&nbsp;</p> <p>Supplementary File 3: C3H_birthdates.csv</p> <p>Data for all C3H mice including: mouse ID, birthdate</p> <p>&nbsp;</p> <p>Supplementary File 4: MA_pheno.R</p> <p>R code for visualising trait data, running linear regressions, and comparing control and MA experiment data</p> <p>&nbsp;</p> <p>Supplementary File 5: C3H_pheno_burnin20_Jun7_2023_nolowmut.csv</p> <p>Data for all mice in MA experiment including a 20 generation burn-in to simulate mutation-drift balance for Animal model analyses: mouse ID, sire, dam, generation, mating ID, sex, weight at 3 weeks, weight at 6 weeks, tail length, litter size, litter ID, line ID</p> <p>&nbsp;</p> <p>Supplementary File 6: asreml_C3H_ALL.R</p> <p>R code for estimating mutational heritabilities using mixed model analysis</p> <p>&nbsp;</p> <p>Supplementary File 7: C3H_ped_rekey_Jun2023.csv</p> <p>Pedigree data for all mice in MA experiment</p> <p>&nbsp;</p> <p>Supplementary File 8: C3H_ped_rekey_KEY.csv</p> <p>Key for pedigree data file</p> <p>&nbsp;</p> <p>Supplementary File 9: plot_pedigree_tree_MS_final.R</p> <p>R code for visualising pedigree of mice in MA experiment</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages

<p>Supplementary information relating to the manuscript titled "SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages" that has been published on the preprint server arXiv.</p> <ul> <li>PersistentInfectionScore.nb: Mathematica code used to process the data and generate Figure 2</li> <li>PersistentInfectionScore.pdf: pdf version of the above file</li> <li>Supplementary_tables_Harari_et_al_2022.xlsx: raw data from (<a href="https://www.nature.com/articles/s41591-022-01882-4#Sec19">Harari et al. 2022</a>) that was used to generate mutation distributions</li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Bedrock radioactivity influences the rate and spectrum of mutation - Orthologous genes

<p>Alignments of the 2490 orthologous genes used in the article &quot;Natural Bedrock radioactivity influences the rate and spectrum of mutation&quot; to estimate the mutational spectrum and synonymous substitution rate.</p> <p>To compute accurate synonymous substitution rate, we removed genes with short sequences (&lt;half of the alignment) and genes strongly supporting another phylogeny using ProfileNJ <a href="https://paperpile.com/c/Klqlpb/W5sS">(Noutahi et al. 2016)</a> with a bootstrap threshold of 90%, resulting in a subset of 769 genes listed in the file &quot;List_769_1-to-1_orthologs_EvolutionRate.txt&quot;.</p> <p>Transcriptome paired-end reads used to define these orthologous genes have been deposited to the European Nucleotide Archive and are available under the study ID PRJEB14193.</p> <p>Sequences were aligned with Prank<a href="https://paperpile.com/c/Klqlpb/pilh"> (L&ouml;ytynoja &amp; Goldman 2008)</a> using a codon model and sites ambiguously aligned were removed with Gblocks <a href="https://paperpile.com/c/Klqlpb/c5kb">(Castresana 2000)</a>.</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'

<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Dynamics of SARS-CoV-2 spike protein in open and closed states and identification of key structural perturbations upon mutations

<p>The SARS-Cov-2 spike protein resides on the exterior surface of the coronavirus, and therefore, acts as the first point of contact that mediates cell attachment and fusion. &nbsp;During this process, it undergoes dramatic conformational changes upon host receptor binding. We are leveraging high-performance computing to identify these structural perturbations in wildtype and mutant spike protein models. The files contain structures from molecular dynamics simulations of closed SARS-Cov-2 spike protein embedded in POPC membrane.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

SIRAH-CoV2 initiative: Spike D614G mutation (introduced on PDB id:6XR8)

<p>This dataset contains the raw data of a coarse-grained molecular dynamics simulation of the D614G mutant of&nbsp;Spike&nbsp;protein&nbsp;from&nbsp;SARS-CoV-2. The initial coordinates correspond to the PDB structure 6XR8, where the coordinates of the side chain of Aspartate&nbsp;614 were deleted to produce a&nbsp;Glycine. Missing loops were introduced with Swiss-Model (https://swissmodel.expasy.org).&nbsp;Simulations were performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in&nbsp;<a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to&nbsp;<a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado &amp; Pantano JCTC 2020</a>.&nbsp;Glycan parameters correspond to those reported by&nbsp;<a href="http://doi.org/10.1101/2020.12.18.423446">Garay et at.</a>2020.&nbsp;</p> <p>The files spike_D614G_SIRAHcg_rawdata_0-2us.tar, spike_D614G_SIRAHcg_rawdata_2-4us.tar,&nbsp;spike_D614G_SIRAHcg_rawdata_4-6us.tar, contain&nbsp;all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing&nbsp;CG trajectories using&nbsp;<a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a>&nbsp;can be found at www.sirahff.com.</p> <p>Additionally, the&nbsp;file&nbsp;spike_D614G_SIRAHcg_prot_10us.tar&nbsp;contains only the protein coordinates, while&nbsp;spike_D614G_SIRAHcg_prot_10us_skip10ns.tar contains one frame every 10ns. Notice that these two files contain a 10-microsecond trajectory, while the raw data is limited to 6 microseconds because of size restrictions on the database. Raw data from microseconds 6 to 10 is available upon request (see below).</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar&nbsp;the file&nbsp;spike_D614G_SIRAHcg_prot_10us_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd spike_D614G_SIRAHcg_prot.prmtop&nbsp;spike_D614G_SIRAHcg_prot.ncrst&nbsp;spike_D614G_SIRAHcg_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc.,&nbsp;and coloring by&nbsp;restype, element, name, etc.&nbsp;</p> <p>This dataset is part of the SIRAH-CoV-2&nbsp;initiative.</p> <p>For further details, please contact Pablo Garay (pgaray@pasteur.edu.uy) or Sergio Pantano (spantano@pasteur.edu.uy).</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Hotspot propensity across mutational processes

<p>Mutational hotspots identified by HotspotFinder from 49 cancer types analysed in: Arnedo-Pac C, Mui&ntilde;os F, Gonzalez-Perez A, Lopez-Bigas N. Hotspot propensity across mutational processes. Mol Syst Biol. 2023:1-22. doi: <a href="https://doi.org/10.1038/s44320-023-00001-w">doi.org/10.1038/s44320-023-00001-w</a></p> <p>Hotspots have been computed using somatic mutations in each cancer type, after excluding those overlapping cancer driver elements. HotspotFinder has been run with a threshold of 2 mutated samples per hotspot and using alternate specific mutations. For additional details, please check Materials and Methods and Appendix Note 1 in our manuscript.</p> <p>HotspotFinder code is available at: <a href="https://bitbucket.org/bbglab/hotspotfinder">bitbucket.org/bbglab/hotspotfinder</a></p> <p>Source code to reproduce this data can be found at: <a href="https://github.com/bbglab/hotspot_propensity">github.com/bbglab/hotspot_propensity</a></p> <p>Cancer type identifiers are listed as follows:</p> <ul> <li>Acute Lymphoblastic Leukemia (ALL)</li> <li>Acute Myeloid Leukemia (AML)</li> <li>Adrenocortical Carcinoma (ACC)</li> <li>Anal Cancer (AN)</li> <li>Basal Cell Carcinoma (BCC)</li> <li>Biliary Tract (BILIARY_TRACT)</li> <li>Bladder/Urinary Tract (BLADDER_URI)</li> <li>Bone/Soft Tissue (BONE_SOFT_TISSUE)</li> <li>Bowel (BOWEL)</li> <li>CNS/Brain (BRAIN)</li> <li>Cervix (CERVIX)</li> <li>Colorectal Adenocarcinoma (COADREAD)</li> <li>Cutaneous Melanoma (SKCM)</li> <li>Cutaneous Squamous Cell Carcinoma (CSCC)</li> <li>Endometrial Carcinoma (UCEC)</li> <li>Ependymoma (EPM)</li> <li>Esophageal cancer (ES)</li> <li>Esophagus/Stomach cancers (ESOPHA_STOMACH)</li> <li>Glioblastoma Multiforme (GBM)</li> <li>Head and Neck (HEAD_NECK)</li> <li>High-Grade Glioma NOS (HGGNOS)</li> <li>Invasive Breast Carcinoma (BRCA)</li> <li>Kidney (KIDNEY)</li> <li>Liver (LIVER)</li> <li>Low-Grade Glioma NOS (LGGNOS)</li> <li>Lung (LUNG)</li> <li>Lung Neuroendocrine Tumor (LNET)</li> <li>Lymphoid Neoplasm (LNM)</li> <li>Medulloblastoma (MBL)</li> <li>Myeloid Neoplasm (MNM)</li> <li>Myeloproliferative Neoplasms (MPN)</li> <li>Neuroblastoma (NBL)</li> <li>Non-Hodgkin Lymphoma (NHL)</li> <li>Non-Small Cell Lung Cancer (NSCLC)</li> <li>Oligodendroglioma (ODG)</li> <li>Ovarian Cancer (OV)</li> <li>Pan-cancer (PANCANCER)</li> <li>Pancreas (PANCREAS)</li> <li>Pilocytic Astrocytoma (PAST)</li> <li>Pleural Mesothelioma (PLMESO)</li> <li>Prostate (PROSTATE)</li> <li>Retinoblastoma (RBL)</li> <li>Skin (SKIN)</li> <li>Small Bowel Cancer (SBC)</li> <li>Small Bowel Neuroendocrine Tumor (SBNET)</li> <li>Small Cell Lung Cancer (SCLC)</li> <li>Stomach cancer (ST)</li> <li>Thyroid (THYROID)</li> <li>Vulva (VULVA)</li> </ul>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Characterization of cefiderocol resistant spontaneous mutant variants of Klebsiella pneumoniae producing NDM-5 with single mutation in cirA

<p>Cefiderocol (CFDC) is a siderophore-cephalosporin antibiotic designed to combat highly resistant Gram-negative bacterial infections. Its mechanism involves a strong affinity for iron and active transport into bacterial cells, providing an alternative against strains resistant to common antibiotics. However, the emergence of CFDC resistance in Klebsiella is a growing concern. Recent reports highlight increasing CFDC resistance in K. pneumoniae, particularly associated with mutations in the cirA gene, responsible for encoding a siderophore receptor. Co-localization of blaNDM-like gene and cirA mutations correlates with higher CFDC resistance. The study focuses on a carbapenem-resistant K. pneumoniae strain (Kp-1) with carbapenemases blaNDM-5 and blaOXA-181, recovered from a post-surgery patient. The strain exhibited resistance to all tested antibiotics but susceptibility to CFDC. Heteroresistant populations with the halo on inhibition of CFDC were observed. Genomic analysis identified a novel mutation (W123*) in the cirA gene associated with CFDC resistance. Additionally, increased blaNDM-5 expression in Kp-1 IHC (intra-halo colony) compared to Kp-1 was noted. The coexistence of blaNDM-like and cirA variants, along with high blaNDM-5 expression, explains the observed 21-fold increase in Minimum Inhibition Concentration (MIC) in Kp-1 IHC. The study contributes to understanding the molecular mechanisms driving the emergence of cefiderocol resistance, emphasizing the significance of coexisting mutations in cirA and blaNDM-like genes.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Early deficits in an in vitro striatal microcircuit model carrying the Parkinson's GBA-N370S mutation

<p><strong>ABSTRACT</strong></p> <p>Understanding medium spiny neuron (MSN) physiology is essential to understand motor impairments in Parkinson&rsquo;s disease (PD) given the architecture of the basal ganglia. Here, we developed a custom three-chambered microfluidic platform and established a cortico-striato-nigral microcircuit partially recapitulating the striatal presynaptic landscape in vitro using induced pluripotent stem cell (iPSC)-derived neurons. We found that, cortical glutamatergic projections facilitated MSN synaptic activity, and dopaminergic transmission enhanced maturation of MSNs in vitro. Replacement of wild-type iPSC-derived dopamine neurons (iPSC-DaNs) in the striatal microcircuit with those carrying the PD-related GBA-N370S mutation led to a depolarisation of resting membrane potential and an increase in rheobase in iPSC-MSNs, as well as a reduction in both voltage-gated sodium and potassium currents. Such deficits were resolved in late microcircuit cultures, and could be reversed in younger cultures with antagonism of protein kinase A activity in iPSC-MSNs. Taken together, our results highlight the unique utility of modelling striatal neurons in a modular physiological circuit to reveal mechanistic insights into GBA1 mutations in PD.</p> <p><strong>FILE DESCRIPTIONS</strong></p> <p>Source Data.xlsx: Tabular datasets plotted on main figures 1, 3, 4, 5 and 6.</p> <p>Supplementary Data.xlsx: Tabular datasets plotted on supplementary figures 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, and 12.</p> <p>Key Resources Table.xlsx: Table containing key resources (primary and secondary antibodies, cell lines and software) used in this study.</p> <p>List of Primers.xlsx: Primers used in&nbsp;RT-qPCR.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Neither alpha-synuclein-preformed fibrils derived from patients with GBA1 mutations nor the host murine genotype significantly influence seeding efficacy in the mouse olfactory bulb

<p>Data sets for;</p> <p>Neither alpha-synuclein-preformed fibrils derived from patients with&nbsp;<em>GBA1</em> mutations nor the host murine genotype significantly influence seeding efficacy in the mouse olfactory bulb</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Characterizing and explaining impact of disease-associated mutations in proteins without known structures or structural homologues

<p>AlphaFold and RoseTTAFold models of domains of disease associated human proteins without structures/known homologues.</p> <p>Tables containing the model quality, model region, sequence&nbsp;alignment statistics, matched FunFam, associated GO terms for the FunFam, ddG of mutation, pathogenicity of mutation,&nbsp;if mutation is near a predicted functional site (conserved residue/ligand binding site/protein-protein interface)</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Assessment of mutation probabilities of KRAS G12 missense mutants and their long-time scale dynamics by atomistic molecular simulations and Markov state modeling: Datasets.

<p>Datasets related to the publication [1].<br> Including:</p> <ul> <li>KRAS G12X mutations derived from COSMIC v.79 [http://cancer.sanger.ac.uk/cosmic/] (KRAS_G12X_mut_COSMICv79..xlsx)</li> <li>RMSFs (300-2000ns) of GDP-systems (300_2000rmsf_GDP_systems_RAW_AVG_SE.xlsx)</li> <li>RMSFs (300-2000ns) of GTP-systems (300_2000RMSF_GTP_systems_RAW_AVG_SE.xlsx)</li> <li>PyInteraph analysis data for salt-bridges and hydrophobic clusters (.dat files for each system in the PyInteraph_data.zip-file)</li> <li>Backbone&nbsp;trajectories for each system (residues 4-164; frames for every 1ns). Last number (e.g. _1) refers to the replica of the&nbsp;simulated system.</li> <li>backbone_4-164.gro/.pdb/.tpr -files (resid 4-164)&nbsp;&nbsp;</li> </ul> <p><br> [1] Pantsar T et al.&nbsp;Assessment of mutation probabilities of KRAS G12 missense mutants and their long-time scale dynamics by atomistic molecular simulations and Markov state modeling. <em>PLoS Comput Biol Submitted</em>&nbsp;(2018)</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

LRRK2 G2019S mutation suppresses differentiation of Th9 and Treg cells via JAK/STAT3

<p>Flow cytometry (Immune profiling) Tidy Data: Figure 1d,e, 2d</p> <p>ELISA Tidy Data: Figure 2c, 4b</p> <p>qRT-PCR Tidy Data: Figure 1b, 3a, 4c,d</p> <p>Cell count Tidy Data: Figure 2b</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Ultrasensitive detection of cancer-associated nucleic acids and mutations by primer exchange reaction-based signal amplification and flow cytometry

<p>This dataset contains the raw data that were used for the publication entitled, "Ultrasensitive detection of cancer-associated nucleic acids and mutations by primer exchange reaction-based signal amplificaiton and flow cytometry" published in Biosensors and Bioelectronics on 5 October 2024.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Dataset for Efficient, robust, and versatile fluctuation data analysis using MLE MUtation Rate calculator (mlemur)

<p>This file contains the R and C++ code used for simulating experiments, simulated fluctuation data, and the results of estimations used in the paper &quot;Efficient, robust, and versatile fluctuation data analysis using MLE MUtation Rate calculator (mlemur)&quot;.</p>

opengpl-2.0Jan 2023View details →
zenodo44/100

gnomAD polymorphism and de novo mutation data for analysis of mutation rates in highly mutable gene classes

<p>We analyze the human mutation rate in three gene classes (IGK, RNU, and tRNA)&nbsp;which deviate from the expectations of a mutation rate model. We examine the distribution of allele frequencies for SNVs within these genes and we analyze the counts of de novo mutations stratified by whether the SNV was observed or not.&nbsp;</p> <p>{CHR}_IGK_SFS_v2_denovo.gz: allele frequencies, mutation rate estimates, and whether the de novo mutation was observed for IGK, RNU, and tRNA genes. Based on gnomAD v3.&nbsp;</p> <p>{CHR}_indiv_mu.csv: quality information for variants in these gene classes from the 1kg subset of gnomAD.</p> <p>&quot;CHR&quot;, &quot;POS&quot;, &quot;REF&quot;, &quot;ALT&quot;, &quot;FILTER&quot;, &quot;AC&quot;, &quot;AN&quot;, &quot;MQRankSum&quot;, &quot;pab_max&quot;, &quot;VQSLOD&quot;, &quot;AB&quot;, &quot;PN&quot;, &quot;MR&quot;, &quot;AR&quot;, &quot;MG&quot;, &quot;MC&quot;, &quot;QUAL&quot;</p> <p>all_variants_chr21_mu_h.csv.gz: all variants from chromosome 21 to use for comparing allele frequencies to those in our gene classes.</p> <p>21_indiv_mu_all.csv.gz: quality information from all variants on chromosome 21 from the&nbsp;1kg subset of gnomAD to use for comparison with gene classes.</p> <p>&quot;CHR&quot;, &quot;POS&quot;, &quot;REF&quot;, &quot;ALT&quot;, &quot;FILTER&quot;, &quot;AC&quot;, &quot;AN&quot;, &quot;MQRankSum&quot;, &quot;pab_max&quot;, &quot;VQSLOD&quot;, &quot;AB&quot;, &quot;PN&quot;, &quot;MR&quot;, &quot;AR&quot;, &quot;MG&quot;, &quot;MC&quot;, &quot;QUAL&quot;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Supplemental data for: Increased mutation and gene conversion within human segmental duplications

<p>Data used for figure generation and analysis in: <strong>Increased mutation and gene conversion within human segmental duplications</strong></p> <ol> <li>new-assemblies.zip contains all the new assemblies added in this work beyond the HPRC assemblies (Clint PTR, CHM1, HG00514, NA12878, HG03125). All other assemblies used in this analysis are available&nbsp;through the HPRC:&nbsp;<a href="https://github.com/human-pangenomics/HPP_Year1_Assemblies/blob/main/assembly_index/Year1_assemblies_v2_genbank.index">assembly_index/Year1_assemblies_v2_genbank.index</a>.</li> <li>all-sample.vcf is a vcf file with all the variant calls used in this analysis.&nbsp;</li> <li>alignments.zip contains all the syntenic&nbsp;alignments used for analysis.&nbsp;</li> <li>data.zip contains annotation data and other information used in analysis and figure making.&nbsp;</li> <li>Online tables 1-4 (Online-tables.xlsx)</li> </ol> <p>Code used in figure making and analysis is&nbsp;on <a href="https://github.com/mrvollger/sd-divergence-and-igc-figures">GitHub</a>.</p> <p>Snakemake pipelines used in the analysis are also on GitHub:</p> <ul> <li>Assembly alignment and IGC calling: https://github.com/mrvollger/asm-to-reference-alignment</li> <li>Variant calling from assembly alignments: https://github.com/mrvollger/sd-divergence</li> <li>Analysis of the triplet content of SNVs: https://github.com/mrvollger/mutyper_workflow</li> </ul>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Supporting data for "CoVEffect: Interactive System for Mining the Effects of SARS-CoV-2 Mutations and Variants Based on Deep Learning"

<p>This repository contains the datasets created and extracted for the paper:</p> <p>Giuseppe Serna Garc&iacute;a, Ruba Al Khalaf, Francesco Invernici, Stefano Ceri, and Anna Bernasconi. 2022.<br> &quot;<strong>CoVEffect</strong>: Interactive System for Mining the <strong>Effects of SARS-CoV-2 Mutations and Variants</strong> Based on Deep Learning&quot;. (Available online at http://gmql.eu/coveffect)</p> <p>--------------------------------------------------------------------------------<br> LIST OF FILES WITH DESCRIPTION:<br> --------------------------------------------------------------------------------</p> <p>AdditionalFile1-effects-taxonomy:<br> Descriptions of legal values for the &#39;Effect&#39; field, based on a categorized taxonomy.</p> <p>AdditionalFile2-levels-taxonomy:<br> Descriptions of legal values for the &#39;Level&#39; field.</p> <p>AdditionalFile3-training_dataset_target:<br> List of target tuples (manually annotated) of 221 abstracts considered for training the model. For each abstract, target tuples&nbsp; follow the schema ID, DOI, title, entity, effect, level, type (mutation or variant), tuples_count (&gt;1 when an effect/level is shared by multiple entities, #abstracts containing the same effect described in the tuple).</p> <p>AdditionalFile4-validation_dataset_target:<br> List of target tuples (manually annotated) of 50 abstracts considered for validating the prepared prediction model.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile5-validation_dataset_highlighted:<br> Textual abstracts of the 50 manuscripts considered for validation; the text used to support the manual target annotations has been highlighted in yellow.</p> <p>AdditionalFile6-validation_dataset_prediction:<br> List of predicted annotations of 50 abstracts considered for validating the prepared prediction model. The file is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p> <p>AdditionalFile7-keywords_query_list:<br> Keyword-based search run on the CORD-19 dataset to extract a relevant subset of abstracts regarding the scope of interest of CoVEffect. The Boolean logic used to combine keywords is explained in the section &#39;Annotations of the biology-related CORD-19 cluster&#39;.</p> <p>AdditionalFile8-CORD-19_batch_dataset_metadata:<br> Metadata of the 7,230 papers extracted by the keyword-based query in AdditionalFile7.<br> These abstracts have been annotated by the prediction framework.</p> <p>AdditionalFile9-CORD-19_batch_dataset_prediction:<br> List of predicted annotations of 7,230 abstracts extracted from the biology-related cluster of CORD-19.</p> <p>AdditionalFile10-test_dataset_target:<br> List of target tuples (manually annotated) of 100 abstracts randomly selected from the 7,230 extracted as in AdditionalFile8.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile11-test_dataset_prediction:<br> List of predicted annotations of 100 abstracts considered for testing the prediction model on a subset of the CORD-19 biology-related cluster. As AdditionalFile6, it is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p>

opencc-zeroDec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record