Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

100

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

100 results for “protein evolution”

Learn how ShareScore rates datasets ↗
zenodo48/100

The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma

<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup>&nbsp;(IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution

<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook&nbsp;</li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

LukProt - an animal evolution-centric eukaryotic protein database

<p>LukProt is the EukProt database with additional species added, mostly the undersampled animal and some holozoan taxa.&nbsp;The database is composed of sequences translated from annotated genomes, transcriptomes or ESTs. <strong>The main purposes of the database are to consolidate sequences from undersampled animal taxa</strong> and provide usable search tools. The publication associated with LukProt can be found here: <a href="https://doi.org/10.1093/gbe/evae231">https://doi.org/10.1093/gbe/evae231</a>.</p> <p>The current version of the database (v1.5.1) is based on <a href="https://doi.org/10.24072/pcjournal.173">EukProt v3</a>. The home of all public versions of LukProt is this page (Zenodo).</p> <p>Proteomes that are novel in LukProt are denoted as LPXXXXX and those coming from AniProtDB are called APXXXXX. The sequence IDs from EukProt are conserved in LukProt. This means that each sequence is assigned an ID in the following format:</p> <pre><code>(A/E/L)PXXXXX_Species_epithet_(strain)_PYYYYYY</code></pre> <p>where XXXXX is a number from 00001 to 99999 and YYYYYY is a number from 000001 to 999999. Each sequence is assigned a unique number YYYYYY, and each taxon XXXXXX. All the IDs are compatible with BLAST v5 "-parse_seqids" option and the database can be readily deployed, for example on a server running <a href="https://doi.org/10.1093/molbev/msz185">SequenceServer</a>. Within each of the source fasta files, the source sequence identifier was kept after a blank space, so that it can still be retrieved if needed.</p> <p>A publicly available BLAST server providing LukProt search is available at: <a title="LukProt BLAST server" href="https://lukprot.hirszfeld.pl/" target="_blank" rel="noopener">https://lukprot.hirszfeld.pl/</a>.</p> <p>Comparison of EukProt v2/v3, LukProt 1.4.1 and LukProt v1.5.1 in their main areas of difference:</p> <table> <tbody> <tr> <th>Taxogroup</th> <th>EukProt v2</th> <th>EukProt v3</th> <th>LukProt v1.4.1</th> <th>LukProt v1.5.1</th> </tr> <tr> <th> <p>Holozoa</p> <p>(excluding Metazoa)</p> </th> <td>31</td> <td>40</td> <td>39</td> <td>43</td> </tr> <tr> <th>Ctenophora</th> <td>2</td> <td>2</td> <td>35</td> <td>38</td> </tr> <tr> <th>Porifera</th> <td>4</td> <td>5</td> <td>30</td> <td>47</td> </tr> <tr> <th>Placozoa</th> <td>2</td> <td>2</td> <td>3</td> <td>6</td> </tr> <tr> <th>Cnidaria</th> <td>3</td> <td>5</td> <td>65</td> <td>88</td> </tr> <tr> <th>Bilateria</th> <td>51</td> <td>51</td> <td>94</td> <td>142</td> </tr> </tbody> </table> <p>Included with the database are:</p> <ul> <li>ready to use main database files: <ul> <li><em>LukProt_v1.5.1_single_species_FASTA.7z</em> &ndash; a FASTA file with the sequences - <a href="https://en.wikipedia.org/wiki/7z">7-zipped</a>, <strong>uncompressed size: 17.6 GB</strong><br> <ul> <li>to concatenate all into one file, run this in the parent directory: <code>for file in $(find . -type f -name "*.fasta"); do awk 'FNR==1{print ""}1' $file &gt;&gt; LukProt_v1.5.1.fa; done</code>. This will create single FASTA file with all the sequences in the parent directory. <code>awk</code> is used to insert a new line after every file because&nbsp;<code>cat</code> would sometimes merge the last sequence with the header of the first sequence.</li> </ul> </li> <li><em>LukProt_v1.5.1_full_BLAST_db.7z</em> &ndash; a preformatted, full BLAST database (NCBI BLAST database format version: v5, masked with segmasker), <strong>uncompressed size: 28.3 GB</strong></li> <li><em>LukProt_v1.5.1_taxogroup_BLAST_db.7z</em> &ndash; a collection of BLAST databases where each proteome is one taxogroup and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.3 GB</strong></li> <li><em>LukProt_v1.5.1_single_species_BLAST_db.7z</em> &ndash; a collection of BLAST databases where each proteome is one BLAST database and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.4 GB</strong></li> </ul> </li> <li>auxiliary database files: <ul> <li><em>LukProt_v1.5.1.cdhit70.7z</em> &ndash; the full database clustered at 70% identity using CD-HIT with the following command: <code>cd-hit -g 1 -d 0 -T 20 -M 90000 -c 0.7 -uL 0.2 -uS 0.9 -s 0.2</code>,&nbsp;<strong>uncompressed sizes: fasta file - 11 GB, clstr file - 2.5 GB</strong></li> <li><em>LukProt_IDs_mapped.txt.gz</em> &ndash; a text file mapping the LukProt IDs to the AniProtDB IDs and EukProt IDs that are different</li> <li><em>BUSCO_tables.ods</em> &ndash; a spreadsheet with full result tables generated by BUSCO analysis</li> <li><em>OMAmer_output.zip</em> &ndash; a folder with full results of OMAmer analyses (includes per-sequence taxonomy classification)</li> <li><em>OMArk_output.zip</em> &ndash; a folder with the results of all OMArk analyses</li> </ul> </li> <li>metadata: <ul> <li><em>README.md</em> &ndash; a README file describing the metadata</li> <li><strong><em>LukProt_metadata_sheet.ods</em> &ndash; main metadata file. A spreadsheet with information about each proteome (in an open .ods format, most compatible with <a href="https://www.libreoffice.org/">LibreOffice</a>)</strong></li> <li><em>LukProt_metadata_other.zip</em> &ndash; an archive with other metadata files, documented in the README. Contents include:<br> <ul> <li>the LukProt taxonomy in various formats</li> <li>supporting scripts for data manipulation and visualization</li> </ul> </li> <li>a recoloring script (modified by LFS, originally by Dr. Celine Petitjean). The script is in&nbsp;<a title="formatFigtree2" href="https://doi.org/10.5281/zenodo.10654583">public domain</a> and reuploaded here only for convenience.&nbsp;</li> <li>other files - see README</li> </ul> </li> <li><em>changelog.md</em> &ndash; database changelog</li> </ul> <p>Words of caution:</p> <ul> <li>The database has been synchronized to EukProt v3 in version v1.5.1. This means that identifiers were modified in comparison to LukProt v1.4.1. The convention is not expected to change any more in future updates.</li> <li>Many proteomes, especially those transcriptome-based, may contain contamination from different species. In addition, the translation algorithms often introduce errors (e.g. the transcript may not represent a full length protein). For this reason, to get accurate sequences from each organism, users are directed to source data and to the included OMAmer, OMArk and BUSCO data for details.</li> <li>The taxonomy is different to UniEuk/EukMap, but UniEuk data were integrated where possible.</li> <li>A few NCBI taxids are missing and will be added in due course.</li> <li>Proteomes from NCBI and UniProt will be updated to current versions.</li> <li>A number of proteomes present in some metadata, are unpublished and were held back.</li> <li>While the database contains metadata that present a particular phylogeny of animals, holozoans and other eukaryotes, no particular claims or hypotheses are made by the author(s). However, in the future efforts will be made to name clades officially, once they are more firmly established.</li> </ul> <p><strong>Please report any problems or suggestions to Lukasz Sobala: lukasz.sobala (at) hirszfeld.pl.</strong></p> <p>&nbsp;</p> <p>Acknowledgements:</p> <ul> <li> <p>Andrew E. Allen Lab for creating the original <a href="https://allenlab.ucsd.edu/data/" target="_blank" rel="noopener">PhyloDB</a>.</p> </li> <li> <p>Daniel Richter <em>et al.</em> for creating <a href="https://doi.org/10.6084/m9.figshare.12417881">EukProt</a> and keeping it updated.</p> </li> <li> <p>Members of <a href="https://multicellgenome.com/">the Multicellgenome Lab</a>, especially Michelle Leger (for donating her database), for the bioinformatics support and for doing great science.</p> </li> <li> <p>All the authors of the original data.</p> </li> <li> <p>National Science Centre of Poland for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this database.</p> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo44/100

PhasAGE Training School 1 -Overview of bioinformatics tools for the life sciences & Classification and evolution of non-globular proteins- LECTUREs

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
dryad40/100

Data from: Convergent rates of protein evolution identify novel targets of sexual selection in primates

<p>Sexual selection is the differential reproductive success of individuals, resulting from competition for mates, mate choice, or success in fertilization. In primates, this selective pressure often leads to the development of exaggerated traits which play a role in sexual competition and successful reproduction. In order to gain insight into the mechanisms driving the development of sexually selected traits, we used an unbiased genome-wide approach across 21 primate species to correlate individual rates of protein evolution to relative testes size and sexual dimorphism in body size, two anatomical hallmarks of sexual selection in mammals. Among species with presumed high levels of sperm competition, we detected strong conservation of testes-specific proteins responsible for spermatogenesis and ciliary form and function. In contrast, we identified accelerated evolution of female reproductive proteins expressed in the vagina, cervix, and fallopian tubes in these same species. Additionally, we found accelerated protein evolution in lymphoid tissue, indicating that adaptive immune functions may also be influenced by sexual selection. This study demonstrates the distinct complexity of sexual selection in primates revealing contrasting patterns of protein evolution between male and female reproductive tissues.</p>

opencc-zeroOct 2023View details →
dryad40/100

Data from: The role of mutation bias in adaptive molecular evolution: insights from convergent changes in protein function

<p>An underexplored question in evolutionary genetics concerns the extent to which mutational bias in the production of genetic variation influences outcomes and pathways of adaptive molecular evolution. In the genomes of at least some vertebrate taxa, an important form of mutation bias involves changes at CpG dinucleotides: If the DNA nucleotide cytosine (C) is immediately 5' to guanine (G) on the same coding strand, and if the C is methylated, then C→T and G→A mutations occur at an elevated rate relative to mutations at non-CpG sites. Here we examine experimental data from case studies in which it has been possible to identify the causative substitutions that are responsible for adaptive changes in the functional properties of vertebrate hemoglobin (Hb). Specifically, we examine the molecular basis of convergent increases in Hb-O<sub>2</sub> affinity in high-altitude birds. Using a data set of experimentally verified, affinity-enhancing mutations in the Hbs of highland avian taxa, we tested whether causative changes are enriched for mutations at CpG dinucleotides relative to the frequency of CpG mutations among all possible missense mutations. The tests revealed that a disproportionate number of causative amino acid replacements were attributable to CpG mutations, demonstrating that mutation bias can influence outcomes of molecular adaptation.</p>

opencc-zeroNov 2023View details →
zenodo40/100

Borg tandem repeats undergo rapid evolution and are under strong selection to create new intrinsically disordered regions in proteins

<p>This repository contains files that accompany the Schoelmerich <em>et al.&nbsp;</em>(2022) bioRxiv preprint.</p> <p>These files include</p> <p>- all Borg proteins used for protein family clustering (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/all_Borg_proteins.fasta">all_Borg_proteins.fasta</a>)</p> <p>- 37 additional aaTR-proteins from manually curated Borg contigs (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/37_Borg_aaTR-proteins.fasta?versionId=bfdfe7b6-e3b5-48b8-b7e6-a20304097e7d">37_Borg_aaTR-proteins.fasta</a>)</p> <p>- IQ-TREE&nbsp;of Borg DNA polymerases and reference sequences from doi: 10.1093/nar/gkaa760 (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/DNAPolB_iqtree.treefile?versionId=c588a98c-822d-4a2f-a92d-7dcb3b1675f0">DNAPolB_iqtree.treefile</a>)</p> <p>- Borg Sm ribonucleoprotein sequences (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/21_Borg_Sm_ribonucleoproteins.fasta">21_Borg_Sm_ribonucleoproteins.fasta</a>)</p> <p>- Borg MHC sequences (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/14_Borg_MHC_proteins.fasta">14_Borg_MHC_proteins.fasta</a>)</p> <p>- AlphaFold2 predicted structural models&nbsp;of Borg Sm ribonucleoproteins and MHCs with aaTRs</p>

opencc-by-4.0May 2022View details →
dryad40/100

Data from: The structure of an ancient genotype-phenotype map shaped the functional evolution of a protein family

Open the record for dataset details and reuse information.

publicMay 2025View details →
dryad40/100

Data from: The role of mutation bias in adaptive molecular evolution: insights from convergent changes in protein function

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad40/100

Data from: Convergent rates of protein evolution identify novel targets of sexual selection in primates

Open the record for dataset details and reuse information.

publicOct 2023View details →
zenodo36/100

the supplementary of Novel plastid genome characteristics in Fugacium kawagutii and accelerated evolution of plastid proteins in dinoflagellates

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo36/100

Data from: Protein Conformational Space at the Edge of Allostery: Turning a Non-allosteric Malate Dehydrogenase into an "Allosterized" Enzyme using Evolution Guided Punctual Mutations

<p>This data&nbsp;accompanies the paper&nbsp;entitled <em>Protein Conformational Space at the Edge of Allostery: Turning a Non-allosteric Malate Dehydrogenase into an &ldquo;Allosterized&rdquo; Enzyme using Evolution Guided Punctual Mutations</em></p> <p>The zip archive contains the results of molecular dynamics simulations of the 4 systems investigated in the paper: wt of A. ful MalDH and three mutants. Each system has been simulated at two temperatures, 300 K and 340 K. Starting configurations of the proteins after equilibration are provided for all the systems in GRO Gromos87 format. Trajectories with the positions of the proteins every 100 ps are provided for all the systems in XTC gromacs format.</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

Evolution of a novel female reproductive strategy in Drosophila melanogaster populations subjected to long term protein restriction

<p>Reproductive output is often constrained by availability of macronutrients, especially protein. Long term protein restriction, therefore, is expected to select for traits maximizing reproduction even under nutritional challenge. We subjected four replicate populations of <em>Drosophila melanogaster</em> to a complete deprivation of yeast supplement, thereby mimicking a protein restricted ecology. Following 24 generations, compared to their matched controls, females from experimental populations showed increased reproductive output early in life, both in presence and absence of yeast supplement. The observed increase in reproductive output was without associated alterations in egg size, development time, pre-adult survivorship, body mass at eclosion, and lifespan of the females. Further, selection was ineffective on lifelong cumulative fecundity. However, females from experiment regime were found to have a significantly faster rate of reproductive senescence following the attainment of the reproductive peak early in life. Therefore, adaptation to yeast deprivation ecology in our study involved a novel reproductive strategy whereby females attained higher reproductive output early in life followed by faster reproductive aging. To the best of our knowledge, this is one of the cleanest demonstration of optimization of fitness by fine tuning of reproductive schedule during adaptation to a prolonged nutritional deprivation.</p>

opencc-zeroMay 2022View details →
zenodo36/100

Protein interaction networks are substantially rewired across evolution

<p>Complete data and source code deposit for Protein interaction networks are substantially rewired across evolution.</p>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Docking data for "The evolution of the SARS-CoV-2 spike protein for differential usage of the host transmembrane serine proteases entry pathway"

<p><br>The dataset includes predicted complexes of the SARS-CoV-2 Spike protein (specifically at the S2' cleavage site) with Hepsin and TMPRSS2 proteins. It contains data on three variants: Wuhan, Delta, and Omicron BA.1.</p> <p><strong>Compressed folders:</strong></p> <p>-357596-DeltaHepsin.tgz</p> <p>-357597-DeltaTMPRSS2.tgz</p> <p>-360039-WuhanHepsin.tgz</p> <p>-360042-TMPRSSWuhan.tgz</p> <p>-392981-TMPRSS-BA1_all.tgz</p> <p>-392982-Hepsin-BA-all.tgz</p> <p><strong>Each compressed folder contains the following:</strong></p> <p>-Initial structures in pdb format</p> <p>-Output complexes in pdb format</p> <p>-Clusters in pdb format</p> <p>-Protocols</p> <p>-Parameters</p> <p>-Scoring files</p> <p>&nbsp;</p> <p><strong>Protein-protein docking&nbsp;</strong><br>Molecular docking between the SARS-CoV-2 S protein of Wuhan, Delta (PDB: 7W92, [DOI: 10.1038/s41467-022-28528-w]), and BA.1 (PDB: 7XO5, [DOI: 10.1038/s41422-022-00672-4]) and the human proteases TMPRSS2 (PDB: 8HD8, [DOI: 10.1038/s41467-023-42527-5]) and Hepsin (PDB: 1Z8G, [DOI: 10.1042/BJ20041955]) was performed using the HADDOCK v2.5-2024.03 webserver ([DOI: 10.1021/ja026939x], [DOI: 10.1016/j.jmb.2015.09.014]). Missing loops in the protein structures were reconstructed using Modeller v10.5 ([DOI: 10.1006/jmbi.1993.1626]). Every heteroatom was removed from the reference structures. The relaxed atomistic coordinates for each S protein variant were derived via all-atom molecular dynamics (MD) simulations. These simulations were performed using AMBER22 with the FF19SB force fields and the pmemd.cuda module for enhanced performance ([DOI: 10.1021/acs.jcim.3c01153], [DOI: 10.1021/jz501780a], [DOI:10.1021/ct400314y]). For the Wuhan variant the S protein was retrieved from our previous modeling study [DOI: 10.1039/D0NR03969A] where for Delta and BA.1, ecah S protein was placed in a dodecahedral box, extending 20 &Aring; beyond the solute in every cartesian direction, and solvated with the four-site OPC water model ([DOI: 10.1021/jz501780a]). The systems were neutralized with counterions, specifically one Cl&minus; ion for the Delta variant and three Cl- ions for the BA.1 variant. To remove local clashes, a geometric optimization was performed using the steepest descent algorithm for 5000 cycles. The MD equilibration process consisted of several stages. First, temperature equilibration in the NVT ensemble was performed by gradually increasing the temperature through steps of 150, 200, 250, 300, and finally 310 K, each lasting 200 ps. During this phase, position restraints were applied to the heavy atoms of the proteins, with progressively decreasing spring constants of 5.0, 4.0, 3.0, and 1.0 kcal mol&minus;1 &Aring;&minus;2, facilitating gradual relaxation. This was followed by a 1 ns equilibration at 310 K in the NPT ensemble without restraints. For production MD, the simulations were run in the NPT ensemble with periodic boundary conditions and Particle Mesh Ewald (PME) method ([DOI: 10.1063/5.0040966], [DOI: 10.1021/ct9001015]) using a grid spacing of 1.0 &Aring; for long-range electrostatics. Non-bonded interactions were modeled with a Lennard-Jones potential using a 9&Aring; cutoff. Temperature control was maintained using Langevin dynamics ([DOI: 10.1021/ct800573m]) with a collision frequency of 4.0 ps&minus;1, and pressure control was managed by the Monte Carlo barostat ([DOI: 10.1016/j.cplett.2003.12.039]) with a 2.0 ps relaxation time at 1 bar. Bond constraints on hydrogen atoms were applied using the SHAKE algorithm ([DOI: 10.1016/0021-9991(77)90098-5]), and the hydrogen mass repartitioning scheme was applied via ParmEd ([DOI: 10.1371/journal.pcbi.1005659]), enabling a 4 fs integration time step ([DOI: 10.1021/ct5010406]). Each protein complex was simulated for a total of 20 ns. For the Wuhan variant, the 3D coordinates were retrieved from [DOI: 10.5281/zenodo.3817446].<br>The active interaction region on the spike protein was defined as the cleavage site (residues P809-R815). For TMPRSS2 and Hepsin, the active sites were defined based on their catalytic residues: H296, D345, D435, S441, S460, and G462 for TMPRSS2, and H203, D257, D347, A348, and S353 for Hepsin. These specific regions were selected to guide the docking process and maximize biologically relevant interactions. Docking clusters were analyzed by selecting those with the lowest interaction energies for further structural analysis. To evaluate binding accuracy, native contacts between the S protein and proteases were computed using the contact map analysis based on the OV+rCSU method ([DOI: 10.12693/APhysPolA.145.S9, 10.1021/acs.jctc.6b00986]), which allows for a precise identification of critical stabilizing interactions, both specific and non-specifics. High-frequency contacts, defined as those appearing in over 70% of the generated models, were highlighted as key determinants of protein-protein recognition, providing insight into the most stable and consistent interactions across docking configurations.</p>

opencc-by-4.0Nov 2024View details →
dryad36/100

Data from: Detailed characterization of the UMAMITs proteins provides insight into their evolution, amino acid transport properties, and role in the plant

<p>Amino acid transporters play a critical role in distributing amino acids within the cell compartments and between the plant organs. Despite this importance, relatively few amino acid transporter genes have been characterized and their role elucidated with certainty. Two main families of proteins encode amino acid transporters in plants: the Amino Acid-Polyamine-Organocation superfamily, containing mostly importers, and the Usually Multiple Acids Move In and out Transporter family, apparently encoding exporters, totaling 63 and 44 genes in Arabidopsis, respectively. Knowledge on UMAMITs is scarce, based on six Arabidopsis genes and a handful of genes from other species. To get insight into the role of the members of this family and provide data to be used for future characterization, we studied the evolution of the UMAMITs in plants, and determined the functional properties, the structure, and the localization of the 47 Arabidopsis UMAMITs. Our analysis showed that the AtUMAMITs are essentially localized at the tonoplast or the plasma membrane, and that most of them are able to export amino acids from the cytosol, confirming a role in intra- and inter-cellular amino acid transport. As an example, this set of data was used to hypothesize the role of a few AtUMAMITs in the plant and the cell.</p>

opencc-zeroAug 2021View details →
zenodo36/100

Supplementary Dataset for Trenner et al. 2022: Evolution and Functions of Plant U-box proteins (PUBs): From protein quality control to signalling

<p>This record contains additional information and a supplementary dataset of sequence alignment files and phylograms that support the publication:</p> <p>Trenner J, Monaghan J, Saeed B, Quint M, Shabek N, Trujillo M. 2022. Evolution and Functions of Plant U-box proteins (PUBs): From protein quality control to signalling. (submitted to Annual Review of Plant Biology)</p> <p>Included files:</p> <p><em>Additional information.pdf</em></p> <p>This file contains Supplementary Methods and References.</p> <p><br> <em>Supp_Fig1_Ubox_full_protein_maximum_likelihood_phylogram_linear.pdf</em></p> <p>A multiple protein sequence alignment by MAFFT of 1121 U-box protein sequences from 21 species and four outgroup sequences containing a RING finger domain (<em>A. thaliana</em> RBX1, <em>A. thaliana</em> RMA1, <em>S. cerevisiae</em> RAD18, <em>Homo sapiens</em> TRAF6 ) was used to infer a phylogenetic tree by maximum likelihood with IQ-TREE. The implemented ultrafast bootstrap approximation, set to 2000 bootstrap samples, was used for branch support. The consensus tree was annotated using iTOL.</p> <p>&nbsp;</p> <p><em>Supp_Fig2_Ubox_domain_maximum_likelihood_phylogram_linear.pdf</em></p> <p>In order to analyse specifically the evolution of the U-box domain, the U-box domain of the final 1121 U-box protein sequences as well as the RING finger domain of the four outgroup sequences were isolated and aligned by MAFFT. The U-box domain alignment was then used to infer a phylogenetic tree by maximum likelihood with IQ-TREE. The implemented ultrafast bootstrap approximation, set to 2000 bootstrap samples, was used for branch support. The consensus tree was annotated using iTOL. This tree was used to interpret the evolutionary history of U-box proteins in the main text.</p> <p>&nbsp;</p> <p><em>MAFFT_alignment _1121seqs_Ubox_full_protein_plus_4seqs_RING_outgroup.fasta</em></p> <p>Multiple protein sequences alignment by MAFFT of 1121 U-box full length protein sequences and four RING finger protein sequences, FASTA formatted.</p> <p>&nbsp;</p> <p><em>MAFFT_alignment _1121seqs_Ubox_domain_plus_4seqs_RING_outgroup.fasta</em></p> <p>Multiple protein sequences alignment by MAFFT of U-box domain sequences of 1121 U-box protein and four RING finger domain sequences of outgroup sequences, FASTA formatted.</p> <p>&nbsp;</p> <p><em>Ubox_protein_sequences_information.xlsx</em></p> <p>Spreadsheet with sequence ID information and source databases.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Directed Evolution of a Surface-Displayed Artificial Allylic Deallylase Relying on a GFP Reporter Protein

<p>Data underlying the figures in the publication &ldquo;Directed Evolution of a Surface-Displayed Artificial Allylic Deallylase Relying on a GFP Reporter Protein&rdquo;, published in <em>ACS Catal</em>. <strong>2021</strong>, 11, 17, 10705&ndash;10712.</p> <p>https://pubs.acs.org/doi/10.1021/acscatal.1c02405</p> <p>Table of contents:</p> <p><strong>1. Dataset</strong>; Excel file with the experimental data for <em>figures 2a, 2b, 3b</em> and <em>5</em>, plus protocol for in vivo catalysis, sequencing data and the selection for rescreening.</p>

opencc-by-4.0Sep 2021View details →
dryad36/100

Tudor genes of Holozoa: Early evolution and within Metazoa diversification of a multifaceted protein family

<p>Early metazoan evolution was characterized by the expansion of many gene families involved in novel multicellularity-related functions, like the Tudor family. In eukaryotes, Tudor genes are numerous and heterogeneous, mostly associated with gene expression regulation. However, the family underwent a lineage-specific expansion in animals, with novel elements almost exclusively involved in the germline-specific regulation of retrotransposons through piRNAs (as spatiotemporal regulators of the key-element Piwi, another previously supposedly animal-specific gene). In the present analysis, we used online-available proteomes for a total of 25 major taxonomic groups to characterize the Tudor gene family at a holozoan-wide level, and we confirmed the apomorphic expansion of piRNA-related Tudor genes in animals. However, we could also interestingly observe the presence of elements of the piRNA pathway, both Tudor and Piwi genes, in some Ichthyosporea species, suggesting that some elements of the pathway were already present in the last common ancestor of Holozoa. Moreover, we observed an outstanding variability (34-fold) of Tudor gene number both between and within metazoan phyla, that could be associated with convergent genomic and phenotypic evolutions. Expansions were usually sided by whole genome duplications and/or life history traits such as parthenogenesis, possibly leading to the expansion of retrotransposon silencing pathways. Reductions were instead mostly associated with overall phenotypic and genomic simplifications, like almost all endoparasites of our dataset. Lastly, we phylogenetically tested a previously proposed model for the evolution of the three possible secondary structures of the Tudor domains and we could mostly (but not completely) confirm the model.</p>

opencc-zeroSep 2023View details →
dryad36/100

Data from: Dissection of the role of a SH3 domain in the evolution of binding preference of paralogous proteins

<p><span>Protein-protein interactions drive many cellular processes. Some protein interactions are directed by Src homology 3 (SH3) domains that bind proline-rich motifs on other proteins. The evolution of the binding specificity of SH3 domains is not completely understood, particularly following gene duplication. Paralogous genes accumulate mutations that can modify protein functions and, for SH3 domains, their binding preferences. Here, we examined how the binding of the SH3 domains of two paralogous yeast type I myosins, Myo3 and Myo5, evolved following duplication. We found that the paralogs have subtly different SH3-dependent interaction profiles. However, by swapping SH3 domains between the paralogs and characterizing the SH3 domains freed from their protein context, we find that few of the differences in interactions, if any, depend on the SH3 domains themselves. We used ancestral sequence reconstruction to resurrect the pre-duplication SH3 domains and examined, moving back in time, how the binding preference changed. Although the closest ancestor of the two domains had a very similar binding preference as the extant ones, older ancestral domains displayed a gradual loss of interaction with the modern interaction partners when inserted in the extant paralogs. Molecular docking and experimental characterization of the free ancestral domains showed that their affinity with the proline motifs is likely not the cause for this loss of binding. Taken together, our results suggest that the SH3 and its host protein could create intramolecular or allosteric interactions essential for the SH3-dependent PPIs, making domains not functionally equivalent even when they have the same binding specificity. </span></p>

opencc-zeroSep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record