Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22
datasets available to search
ShareScore release 0.9.0
Dataset results
22 results for “protein function prediction”
DATASET: Predicting Protein Function and Orientation on a Gold Nanoparticle Surface Using a Residue-Based Affinity Scale
<p>This upload contains data for the manuscript "<strong>Predicting Protein Function and Orientation on a Gold Nanoparticle Surface Using a Residue-Based Affinity Scale</strong>." It contains kinetics data, UV-vis data, surface calculations, and activity assays for the systems described in the manuscript.</p>
Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "
Open the record for dataset details and reuse information.
Supporting data for: A close look at protein function prediction evaluation protocols
<p>Data files and predictions associated with the paper: A close look at protein function prediction evaluation protocols.</p>
A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2
<p>Data produced and analyzed in the manuscript "A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2" by Pasquadibisceglie et al.</p> <p><br> If you include these data in your manuscript, please cite: Pasquadibisceglie A, Leccese A and Polticelli F (2022) A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2. <em>Front. Chem.</em> 10:1004815. doi: 10.3389/fchem.2022.1004815</p>
Multimodal dataset: Protein Function Prediction using STRING data & COVID19 Mortality Model by EI
<p>The PFP.zip file contains 1. 5 well-formated GO terms dataset for EI, 2. STRING data 3. GO term annotation. The last two could be merged by the 'generate_data.py' script in https://github.com/GauravPandeyLab/ensemble_integration</p> <p>The covid19_model_built.zip contained the EI model built based on the COVID-19 Mortality dataset, the detail of usage are here:.</p>
How can we biochemically validate protein function predictions with the Ras GTPase family? - Associated data
<p>This is the data that accompanies the pub "<a href="https://doi.org/10.57844/arcadia-74ad-345f">How can we biochemically validate ProteinCartography with the Ras GTPase family?</a>" It's part of a group of pubs focused on validating ProtienCartography that begins with "<a href="https://doi.org/10.57844/arcadia-cae9-96c4">A strategy to validate protein functions <em>in vitro</em></a><a href="https://doi.org/10.57844/arcadia-cae9-96c4">." </a></p> <p>For this repository, we ran ProteinCartography <a href="https://github.com/Arcadia-Science/ProteinCartography/releases/tag/v0.5.0">v0.5.0</a> using human HRas and KRas as our inputs for a single run (UniProt ID: <a href="https://www.uniprot.org/uniprotkb/P01112/entry">P01112</a> and <a href="https://www.uniprot.org/uniprotkb/P01116/entry">P01116</a>). We asked for 3,000 Foldseek hits and 7,000 BLAST hits for a total of 10,000 structures. The updated configuration file is in the zipped folder in this repository. Also included in the zipped folder are the inputs, structures of all hits, and all ProteinCartography results. </p> <p>Finally, we created a custom overlay for the protein map using this <a href="https://github.com/Arcadia-Science/2023-actin-embedding/blob/main/notebooks/3_plotting_overlays.ipynb">notebook</a> and the manually annotated TSV file in this repository, where we denoted which group of substrates a protein is predicted to act on based on its annotation from UniProt.</p>
How can we biochemically validate protein function predictions with the deoxycytidine kinase family? - Associated data
<p>This is the data that accompanies the pub "<a href="https://doi.org/10.57844/arcadia-1e5d-e272">How can we biochemically validate ProteinCartography with the deoxycytydine kinase family?</a>" It's part of a group of pubs focused on validating ProtienCartography that begins with "<a href="https://doi.org/10.57844/arcadia-cae9-96c4">A strategy to validate protein functions <em>in vitro</em></a><a href="https://doi.org/10.57844/arcadia-cae9-96c4">." </a></p> <p>For this repository, we ran ProteinCartography <a href="https://github.com/Arcadia-Science/ProteinCartography/releases/tag/v0.5.0">v0.5.0</a> on the deoxycytidine kinase (dCK) using human dCK as our input (UniProt ID: <a href="https://www.uniprot.org/uniprotkb/P27707/entry">P27707</a>). We asked for 3,000 Foldseek hits and 7,000 BLAST hits for a total of 10,000 structures. The updated configuration file is in the zipped folder in this repository. Also included in the zipped folder are the inputs, structures of all hits, and all ProteinCartography results. </p> <p>Finally, we created a custom overlay for the protein map using this <a href="https://github.com/Arcadia-Science/2023-actin-embedding/blob/main/notebooks/3_plotting_overlays.ipynb">notebook</a> and the manually annotated TSV file in this repository, where we denoted which group of substrates a protein is predicted to act on based on its annotation from UniProt.</p>
Input Data for "Protein Function Prediction for newly sequenced organisms"
<p>The input sequence files in FASTA format and the detailed list of all organisms excluded when testing each specific bacterium.</p>
Benchmark dataset for protein function prediction that integrates various protein information
<p>This research was funded by the National Science Centre in Poland (grant number 2021/41/N/ST6/01919)</p>
Data for DualNetGO: A Dual Network Model for Protein Function Prediction via Effective Feature Selection
<p>Data used in the paper, including annotation files, graph embeddings from TransformerAE, and protein attributes for both human and mouse, and for cafa3 data. Extract and place them in the <em>data </em>folder.</p>
A strategy to validate protein function predictions in vitro - accompanying data
<p>This repository contains the ProteinCartography analyses that accompany the pub "<a href="https://doi.org/10.57844/arcadia-cae9-96c4">A strategy to validate protein function predictions <em>in vitro</em></a><a href="https://doi.org/10.57844/arcadia-cae9-96c4">."</a> These proteins were selected from a list of the 200 most studied human proteins in the Protein Data Bank (PDB) which was published in <a href="https://onlinelibrary.wiley.com/doi/10.1002/pro.4038">Liu and Buck, 2021</a>. </p> <p>Each analysis was done using <a href="https://github.com/Arcadia-Science/ProteinCartography/releases/tag/v0.5.0">v0.5.0</a> of ProteinCartography using the standard configuration parameters. Each analysis is in its own zipped file corresponding to the protein name. The zipped files contain inputs, configuration files, all the structures, and all ProteinCartography results. Additionally, for a handful of the proteins, we ran a scaled-up version of ProteinCartography asking for 10,000 hit proteins instead of the standard. These are also included with the file prefix "Scaled-up_". </p> <p>Finally, we've decided to move forward with 2 protein families for validation of our tool, ProteinCartography. These protein families are <a href="https://doi.org/10.57844/arcadia-1e5d-e272">deoxycytidine kinase</a> and <a href="https://doi.org/10.57844/arcadia-74ad-345f">Ras GTPase</a>. For more about these two protein families, visit the mentioned pubs! </p>
Fueling ab initio folding with oceanic metagenomics enables structure and function predictions of new protein families
<p>Code and protein sequence database to construct multiple sequence alignment from Tara Ocean data.</p>
Predicted protein functions for 2031 Saccharomyces cerevisiae genome assemblies
<p>Predicted protein functions for 2031 Saccharomyces cerevisiae genome assemblies</p>
Monte Carlo Arithmetic Instrumented DeepGOPlus Protein Function Predictions
<p>This dataset contains the perturbed protein function predictions by the DeepGOPlus model excluding the Diamond tool component. The model was perturbed with Verrou, an implementation of Monte Carlo Arithmetic (MCA), a stochastic arithmetic technique that injects noise into a program that simulates changes in a user's execution environment. The folders contain pkl files that can be read with the Pandas python library to load up dataframes containing the predictions and original values. Each file is one MCA sample run across the entire DeepGOPlus test set.</p> <p>The folder "Verrou_All" contains predictions where the entirety of the model was instrumented with MCA.</p> <p>The folder "Verrou_TF" contains predictions where only the Tensorflow library was instrumented with MCA.</p> <p>The folder "Fuzzy_Python" contains predictions where only the Python interpreter was instrumented with MCA.</p> <p>The folder "VPREC_Outbound_Mode" contains predictions where the virtual precision of the floating point operations was reduced witht the VPREC precision simulator tool in outbound mode.</p> <p>The folder "VPREC_Inbound_Mode" contains predictions where the virtual precision of the floating point operations was reduced witht the VPREC precision simulator tool in inbound mode.</p> <p>More information can be found by consulting <a href="https://arxiv.org/abs/2212.06361">this paper</a> or <a href="https://github.com/big-data-lab-team/deepgoplus-stability">this Github repository</a>.</p>
DeepGO-SE protein function prediction model data
<p>Data for training and running DeepGO-SE protein function prediction model</p>
Common protein interactions explain variant oncofusion function and predict a novel oncofusion in rhabdomyosarcoma
<p>Original Westerns from publications</p>
Dataset for the paper: Predicting protein functions using positive-unlabeled ranking with ontology-based priors
Open the record for dataset details and reuse information.
Sequence data and structural data utilized in the study and analysis of grain protein function prediction.
Open the record for dataset details and reuse information.
Functional prediction of select proteins from the gut archaeome
<p>Synteny plots of archaeal gut-specific unique and homologous gene clusters. </p>
DualNetGO: A Dual Network Model for Protein Function Prediction via Effective Feature Selection
<p>Data used in the paper, including annotation files, graph embeddings from TransformerAE, and protein attributes for both human and mouse. Extract and place them in the <em>data </em>folder, and there will be two two folders <em>human </em>and <em>mouse </em>containing necessary data for training and testing.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.