Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

261

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

261 results for “Protein prediction”

Learn how ShareScore rates datasets ↗
zenodo56/100

Cross-phyla protein annotation by structural prediction and alignment

<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is&nbsp;typically limited to a few model organisms. In non-model species, the sequence-based prediction of&nbsp;gene orthology can be used to infer protein identity, however this approach loses predictive power&nbsp;at longer evolutionary distances. Here we propose a workflow for protein annotation using structural&nbsp;similarity, exploiting the fact that similar protein structures often reflect homology and are more&nbsp;conserved than protein sequences.</p> <p><strong>Results:</strong>&nbsp;&nbsp;We propose a workflow of openly available tools for the functional annotation of proteins via&nbsp;structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete&nbsp;proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet&nbsp;their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with&nbsp;known homology in &gt;90%&nbsp;cases, and annotates an additional 50%&nbsp;of the proteome beyond&nbsp;standard sequence-based methods. We uncover new functions for sponge cell types, including extensive&nbsp;FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in&nbsp;myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes,&nbsp;proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Discoba protein sequences for protein structure predictions

<p>Comprehensive database of Discoba protein sequences, gathered for the purpose of improving protein structure predictions of Discoba species (including <em>Trypanosoma </em>and <em>Leishmania</em>) by AlphaFold and RoseTTAFold. Originally gathered for use with:&nbsp;https://github.com/zephyris/discoba_alphafold</p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

Prediction of Conformational Variability for RRM proteins in inter3m data base

<p>Predictions for protein Conformational Variability for the entries in&nbsp;InteR3M (<a href="https://inter3mdb.loria.fr/">https://inter3mdb.loria.fr/</a>), performed with the software ConforMine (in preparation).</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Associated Data: RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features

<p>Additional digital data to &quot;RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features&quot; (ChemRxiv preprint:<a href="https://doi.org/10.26434/chemrxiv.12636704.v1">https://doi.org/10.26434/chemrxiv.12636704</a>).</p> <p>Associated code can be found at:&nbsp;<a href="https://github.com/HITS-MCM/RASPDplus">https://github.com/HITS-MCM/RASPDplus</a></p> <p>Files:</p> <ul> <li>weights.tar.gz: contains the model weights of one random dataset split and its associated crossvalidation folds. Used for standard RASPD+ evaluation.</li> <li>additional_model_replicates.tar.gz: contains the remaining models trained on the full set of descriptors.</li> <li>external_test_sets.tar.gz: contains the descriptor tables for all external test sets used</li> <li>dude.tar.gz: contains the descriptor tables for and several identifier lists for evaluation on the Directory of Useful Decoys - Enhanced (DUD-E)</li> <li>run_outputs.tar.gz: Performance metric data and predicted values created during the model training and evaluation runs. Basis for the figures and metrics in the manuscript.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

RNA-Protein Interaction Prediction Using Network-Guided Deep Learning

<p>RNA-protein interactions are critical to various life processes, including fundamental translation and gene regulation. Identifying these interactions is vital for understanding the mechanisms underlying life processes. Then, ZHMolGraph is an advanced pipeline that integrates graph neural network sampling strategy and unsupervised large language models to enhance binding predictions for novel RNAs and proteins.</p> <div>&nbsp;</div>

openmit-licenseJul 2024View details →
zenodo44/100

Datasets for Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes

<p>All used Datasets to pair with &quot;Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes&quot; by Nicholas Ouassil*, Rebecca L. Pinals*, Jackson Travis Del Bonis-O&#39;Donnell, Jeffrey W. Wang, and Markita P. Landry</p> <p>*Co-authors</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Predicting Exon Criticality from Protein Sequence

<p>Exon ByPASS (predicting Exon-skipping Based in Protein amino acid SequenceS), predictions on test exons from Human and Mouse transcripts. The exons in the test set from the two genomes are those that are not predicted to be skippable in hg38 and mm10 annotation and are also exons that in-frame when skipped. The preprocessed data includes the ensemble transcript id and exon rank as well the amino acid sequence for the upstream, downstream, and exon of interest. Additionally, the table contains the output probability from the model in the last column. The input data is&nbsp;transformed data of the amino acid sequence that is the Exon ByPASS model can use as an input.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "

<p>Datasets for&nbsp;preprint&nbsp;entitled &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi&quot;</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of&nbsp;<em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi&quot;.&nbsp;&nbsp;&nbsp;</p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict.&nbsp;For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22.&nbsp;In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using&nbsp;Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes.&nbsp;</p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset&nbsp;is made up of predicted protein tertiary structures representing the main member of each up-regulated&nbsp;<em>V. inaequalis</em> effector candidate family. Structures were predicted using&nbsp;Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>).&nbsp;In cases where&nbsp;the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2.&nbsp;Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold&nbsp;(<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>)&nbsp;open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input.&nbsp;</p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence&nbsp;proteins from other fungi&quot; study. These structures were predicted using&nbsp;Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input.&nbsp;</p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>

opencc-by-2.0Feb 2022View details →
zenodo44/100

Replication Data for: Geometric Transformers for Protein Interface Contact Prediction

<p>This dataset contains replication data for the paper titled &quot;Geometric Transformers for Protein Interface Contact Prediction&quot;. The dataset consists of pickled Python dictionaries containing pairs of DGLGraphs&nbsp;that can be used to train and validate&nbsp;protein interface contact prediction models. It also contains our best model checkpoints saved as&nbsp;PyTorch LightningModules.&nbsp;Our GitHub repository, DeepInteract, linked in the &quot;Additional notes&quot; metadata section below provides more details on how we use&nbsp;these files as&nbsp;examples for cross-validation.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Input features of E. coli proteome for predicting and modeling protein-protein interactions with AF2Complex

<p>Input features to be used with AF2Complex for predicting protein-protein interactions among ~4400 E. coli proteins. A pickled feature file was generated by the feature data pipeline of AF2Complex for each E. coli protein. To reduce storage size, we limited up to 10,000 MSA sequences and up to 10 structural templates from the Protein Data Bank. The cutoff date for sequence libraries and the Protein Data Bank releases used for feature generation is no later than 11-30-2021.</p> <ul> <li>ecoli_af2c_fea.txt -- A list of all E coli protein with pre-generated input features</li> <li>af2c_fea_ecoli_220331_msa10ktem10.tar&nbsp;-- Input features named after the UniProt ID of each proteins. Note that after untar the tarball, you may use the gzipped feature pickle files directly with AF2Complex w/o gunzip.</li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Supplementary Data for NIPS Publication: Protein Interface Prediction using Graph Convolutional Networks.

<p>These data sets can be used to re-run the experiments from our paper, Protein Interface Prediction using Graph Convolutional Networks. The data are derived from protein complexes in the docking benchmark dataset v. 5.0. Each file is a python tuple&nbsp;that has been saved using cPickle and compressed using gzip.</p> <p>Links:</p> <p>Paper: https://papers.nips.cc/paper/7231-protein-interface-prediction-using-graph-convolutional-networks</p> <p>Poster:&nbsp;https://zenodo.org/record/1134154</p> <p>Code:&nbsp;https://github.com/fouticus/pipgcn</p> <p>&nbsp;</p> <p><strong>File Descriptions:</strong></p> <p>train.cpkl.gz and test.cpkl.gz have the data formatted for neighborhood based graph convolutions. The diffc_ files are the same data formatted for the diffusion convolutional neural networks that we compare against.&nbsp;</p> <p>train.cpkl.gz is a tuple of length 2:</p> <ul> <li>element 0 is a list of length 175 containing the PDB codes from the docking benchmark dataset</li> <li>element 1 is a list of length 175 containing features for each protein. Each element is a dictionary containing the following keys: <ul> <li>r_vertex: vertex (residue) features for the receptor. numpy array of shape (x, 70) where x is the number of residues in the receptor and 70 is the number of features.</li> <li>l_vertex: vertex (residue) features for the ligand. analogous to above, with shape (y, 70) where y is the number of residues in the ligand.</li> <li>complex_code: PDB code&nbsp;of the complex. matches the list of codes described above.</li> <li>l_edge: edge features for the neighborhood around each residue in the ligand. numpy array of shape (y, 20, 2) where y&nbsp;is defined as above. the second dimension is the edges to the 20 nearest neighboring residues,&nbsp;ordered by decreasing distance. The third dimension allows for two features per edge.&nbsp;</li> <li>r_edge: edge features for the neighborhood around each residue in the receptor. numpy array of shape (x, 20, 2) where x&nbsp;is as above.&nbsp;</li> <li>l_hood_indices: the index of the 20 closest residues to each residue, ordered by decreasing distance. numpy array of shape (y, 20, 1). &quot;Index&quot; means which row in l_vertex gives the vertex features for the closest neighbor, second closest neighbor, etc.&nbsp;</li> <li>r_hood_indices: analogous to above, shape (x, 20, 1).</li> <li>label: 1 or -1 label for each residue pair. numpy array of shape (x*y, 3). Each row looks like (i, j, k) where i is the index of the ligand&nbsp;residue, j is the index of the receptor residue, and k is either -1 (negative example) or 1 (positive example).</li> </ul> </li> </ul> <p>test.cpkl.gz matches the structure of train.cpkl.gz except it has the test set of 55 complexes.&nbsp;</p> <p>Descriptions of the vertex and edge features can be found in Appendix A of &nbsp;<a href="https://mountainscholar.org/handle/10217/185661">this.</a></p> <p>diffc_g2_p2_train.cpkl.gz is a tuple of length 2:</p> <ul> <li>element 0 is a list of the same 175 PDB codes as above.&nbsp;</li> <li>element 1 is a list of features for the 175 complexes. Each element is a dictionary of features with these keys: <ul> <li>r_vertex, l_vertex, complex_code, label: these are the same as described above.&nbsp;</li> <li>&#39;r_power_series&#39;: Stacked diffusion matrices which are powers of the similarity matrix used in the DCNN method. numpy array of shape (x, 2, x) where x&nbsp;is the number of receptor residues. the middle dimension 2 indicates how many &quot;hops&quot; is used for that diffusion (1 vs. 2). In other words, element (i, 0, j) is the similarity after 1 hops between residues i and j. element (i, 1, j) is the similarity after 2 hops.&nbsp;See DCNN paper for details.</li> <li>&#39;l_power_series&#39;: same as above but for the ligand. shape is (y, 2, y).</li> </ul> </li> </ul> <p>diffc_g2_p2_test.cpkl.gz is the same as diffc_g2_p2_train.cpkl.gz but for the 55 test complexes.</p> <p>diff_g2_p5_train.cpkl.gz and diff_g2_p5_test.cpkl.gz are the same as the p2 version above, except that the diffusion matrices have shape (x, 5, x) and (y, 5, y) because one of our comparisons against&nbsp;the DCNN model uses 5 hops instead of just 2.&nbsp;</p> <p>&nbsp;</p> <p>Note: these files were pickled with Python 2.7. If you&#39;re unpickling with Python 3.x you might have to specify encoding as &#39;latin1&#39;.&nbsp;</p> <p>&nbsp;</p> <p>Please direct any questions to:</p> <ul> <li>Alex Fout (fout@colostate.edu)</li> <li>Jonathon Byrd (jonbyrd@colostate.edu)</li> <li>Basir Shariat (basir@cs.colostate.edu</li> <li>Asa Ben-Hur (asa@cs.colostate.edu)</li> </ul>

opencc-by-sa-4.0Dec 2017View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction and databases of protein phase separation - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction and databases of protein phase separation Overview- LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1-Protein aggregation prediction-PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction of intrinsic disorder in proteins-DisProt - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction of intrinsic disorder in proteins-MobiDB - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction

<p>This dataset contains replication data for the paper titled &quot;DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction&quot;. The dataset consists of pickled Pandas DataFrame files, along with training, validation, and (for DB5-Plus) test&nbsp;filename lists for cross-validation, that can be used to develop and evaluate&nbsp;protein interface prediction models. This dataset also contains the externally generated residue-level PSAIA and HH-suite3 features for users&#39; convenience (e.g. raw MSAs and profile HMMs for each protein complex).&nbsp;Our GitHub repository linked in the &quot;Additional notes&quot; metadata section below provides more details on how we parsed through these files to create our cross-validation&nbsp;datasets. The GitHub repository&nbsp;for DIPS-Plus&nbsp;also includes scripts that can be used&nbsp;to impute missing feature values and convert the&nbsp;final &quot;raw&quot; complexes into DGL-compatible graph objects. Since our final DGL graph representation for each complex uses PyTorch tensors in its construction of&nbsp;residue&nbsp;embeddings, the final representation of each complex can easily be adapted to fit the users&#39; needs (e.g. feeding a complex&#39;s&nbsp;2D residue feature tensors into a convolutional neural network).</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Embeddings from protein language models predict conservation and variant effects

<p>For this work, we used protein language model representations (embeddings) to predict sequence conservation without multiple sequence alignments (MSAs). Embeddings alone predicted residue conservation almost as accurately from single sequences as ConSeq using MSAs (two-state Matthew Correlation Coefficient &ndash; MCC - for ProtT5 embeddings of 0.596&plusmn;0.006 vs. 0.608&plusmn;0.006 for ConSeq).</p> <p><strong><em>ConSurf10k</em>- Dataset for the development of ProtT5cons:</strong> The method (ProtT5cons) predicting residue conservation used <em>ConSurf-DB </em>(Ben Chorin et al. 2020). This resource provided sequences and conservation for 89,673 proteins. For all, experimental high-resolution three-dimensional (3D) structures were available in the Protein Data Bank (PDB) (Berman et al. 2000). As standard-of-truth for the conservation prediction, we used the values from ConSurf-DB generated using HMMER (Mistry et al. 2013), CD-HIT (Fu et al. 2012), and MAFFT-LINSi (Katoh and Standley 2013) to align proteins in the PDB (Burley et al. 2019). For proteins from families with over 50 proteins in the resulting MSA, an evolutionary rate at each residue position is computed and used along with the MSA to reconstruct a phylogenetic tree. The ConSurf-DB conservation scores ranged from 1 (most variable) to 9 (most conserved). The PISCES server (Wang and Dunbrack 2003) was used to redundancy reduce the data set such that no pair of proteins had more than 25% pairwise sequence identity. We removed proteins with resolutions &gt;2.5&Aring;, those shorter than 40 residues, and those longer than 10,000 residues. The resulting data set (ConSurf10k) with 10,507 proteins (or domains) was randomly partitioned into training (9,392 sequences), cross-training/validation (555) and test (519) sets.</p> <p>Uploaded data:</p> <ul> <li>ConSuf10k_PDBid_seq_cons.fasta: fasta file with PDBid, sequence and conservation annotation</li> <li>consurf10k_test_ids.txt: txt file with id&#39;s of test set</li> <li>consurf10k_train_ids.txt: txt file with id&#39;s of train set</li> <li>consurf10k_val_ids.txt: txt file with id&#39;s of cross-validation set</li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo44/100

DATASET: Predicting Protein Function and Orientation on a Gold Nanoparticle Surface Using a Residue-Based Affinity Scale

<p>This upload contains data for the manuscript &quot;<strong>Predicting Protein Function and Orientation on a Gold Nanoparticle Surface Using a Residue-Based Affinity Scale</strong>.&quot; It contains kinetics data, UV-vis data, surface calculations, and activity assays for the systems described in the manuscript.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Predictive modeling of moonlighting DNA-binding proteins

<p>This repository contains the codes used for the prediction of moonlighting proteins&nbsp;in the paper &quot;Predictive modeling of moonlighting DNA binding proteins&quot;.</p> <p>The repository is organized as the following:</p> <p>1. The DNA binding protein identifiers&nbsp;and their features that were used to train the models for the prediction of DNA binding Moonlighting proteins.</p> <p>2. Five feature sets were used to create Catboost models that make predictions. The source code for generating predictions based on all the features and predictions based on particular features is supplied. In addition, the source code for generating maximum and average ensemble predictions has been made available. A detailed explanation is given in README file.</p>

opencc-by-4.0Nov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record