Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “protein structure prediction”
Cross-phyla protein annotation by structural prediction and alignment
<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is typically limited to a few model organisms. In non-model species, the sequence-based prediction of gene orthology can be used to infer protein identity, however this approach loses predictive power at longer evolutionary distances. Here we propose a workflow for protein annotation using structural similarity, exploiting the fact that similar protein structures often reflect homology and are more conserved than protein sequences.</p> <p><strong>Results:</strong> We propose a workflow of openly available tools for the functional annotation of proteins via structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with known homology in >90% cases, and annotates an additional 50% of the proteome beyond standard sequence-based methods. We uncover new functions for sponge cell types, including extensive FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes, proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>
Discoba protein sequences for protein structure predictions
<p>Comprehensive database of Discoba protein sequences, gathered for the purpose of improving protein structure predictions of Discoba species (including <em>Trypanosoma </em>and <em>Leishmania</em>) by AlphaFold and RoseTTAFold. Originally gathered for use with: https://github.com/zephyris/discoba_alphafold</p>
Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "
<p>Datasets for preprint entitled "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi"</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of <em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi". </p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict. For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22. In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes. </p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset is made up of predicted protein tertiary structures representing the main member of each up-regulated <em>V. inaequalis</em> effector candidate family. Structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). In cases where the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2. Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold (<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>) open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input. </p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi" study. These structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input. </p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>
DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction
<p>This dataset contains replication data for the paper titled "DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction". The dataset consists of pickled Pandas DataFrame files, along with training, validation, and (for DB5-Plus) test filename lists for cross-validation, that can be used to develop and evaluate protein interface prediction models. This dataset also contains the externally generated residue-level PSAIA and HH-suite3 features for users' convenience (e.g. raw MSAs and profile HMMs for each protein complex). Our GitHub repository linked in the "Additional notes" metadata section below provides more details on how we parsed through these files to create our cross-validation datasets. The GitHub repository for DIPS-Plus also includes scripts that can be used to impute missing feature values and convert the final "raw" complexes into DGL-compatible graph objects. Since our final DGL graph representation for each complex uses PyTorch tensors in its construction of residue embeddings, the final representation of each complex can easily be adapted to fit the users' needs (e.g. feeding a complex's 2D residue feature tensors into a convolutional neural network).</p>
ESM Atlas v0 representative random sample of predicted protein structures
<p>A representative random sample of the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> Sample size: 997,405.</p>
ESM Atlas v0 random sample of high confidence predicted protein structures
<p>A random sample out of the 225M high confidence predictions in the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> High confidence is defined as mean pLDDT > 0.7 and pTM > 0.7 and corresponds to ∼36% of the total 617M proteins folded.<br> This is the random sample used for analysis in the paper as well as visualization on the <a href="http://esmatlas.com/">esmatlas.com</a> Explore page.<br> Sample size: 999,520 based on 999,996 unique randomly sampled IDs and 0.05% missing data in the processing pipeline.</p>
DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction (Supplementary Data)
<p>This dataset contains supplementary replication data for the paper titled "DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction". In particular, it contains a new version of our `final_raw_dips.tar.gz` protein pair representations which now contain (1) residue-level annotations for intrinsic disorder regions (IDRs) as well as (2) a copy of each protein pair representation in the HDF5 file format for programming language-agnostic read capabilities. In addition, this record also contains (3) raw MSAs (in HDF5 file format) generated for each protein pair using Jackhmmer and AlphaFold's small version of the Big Fantastic Database (BFD). Lastly, this record contains (4) PDB metadata derived for each DIPS-Plus complex using Graphein's PDBManager API as well as (5) structure-based (i.e., FoldSeek-based) training and validation splits of the dataset's complexes in the form of respective text files containing the file paths of complexes assigned to each split.</p>
Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "
Open the record for dataset details and reuse information.
Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions
<p>Using AlphaFold2 and HADDOCK, we have generated a predicted structure for the SARS-CoV-2 B.1.1.529 variant's Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in the receptor binding domain of B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>
[Accompanying Dataset for PHIStruct] ColabFold-Predicted Structures of Receptor-Binding Proteins
<p><strong>This dataset contains protein structures, computationally predicted via <a href="https://doi.org/10.1038/s41592-022-01488-1">ColabFold</a>, of 19,081 non-redundant (i.e., with duplicates removed) receptor-binding proteins from 8,525 phages across 238 host genera</strong>. We identified these receptor-binding proteins based on GenBank annotations. For phage sequences without GenBank annotations, we employed a pipeline that uses the viral protein library <a href="https://doi.org/10.1093/nargab/lqab067">PHROG</a> and the machine learning model <a href="https://doi.org/10.3390/v14061329">PhageRBPdetect</a>. </p> <p>More details can be found in our paper <strong>"PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings."</strong> The project page is <a href="https://github.com/bioinfodlsu/PHIStruct">https://github.com/bioinfodlsu/PHIStruct</a>. Our paper is published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btaf016" rel="nofollow">https://doi.org/10.1093/bioinformatics/btaf016</a></p> <p>Our research was supported with Cloud TPUs from <a href="https://sites.research.google/trc/about/" rel="nofollow">Google's TPU Research Cloud (TRC)</a> and with computing resources from the <a href="https://docs.mlerp.cloud.edu.au/" rel="nofollow">Machine Learning eResearch Platform (MLeRP)</a> of Monash University, University of Queensland, and Queensland Cyber Infrastructure Foundation Ltd.</p>
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain
Figure 3. Distribution of Q3 values-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The estimated accuracy for the α- helices (QH), β- strands (QE), C-coil states (QC), and three<br> state together (Q3) for the system is shown in Figure 3.</p>
Figure 2. PAM250 matrix for the encoded sequence-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The PAM matrix (Dayhoff et al., 1978) describes the probability that original amino acid<br> will be replaced by another amino acid over a defined evolutionary interval. The unit of<br> evolutionary divergence is defined as the interval in which 1% of the amino acids have been<br> changed between two sequences. The work uses PAM250, which assumes the occurrence of 250-<br> point mutations per 100 amino acids.<br> So, for the given the protein sequence GIVEQCCASVCSLYQLENYCN, A will be replaced<br> by 1 -3 0 1 -3 -1 0 5 -2 -3 -4 -2 -3 -5 0 1 0 -7 -5 -1 as shown in Figure 2.</p>
Figure 1: Snapshot of the CB396 dataset-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning AlgorithmSecondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The dataset used for this work is CB396. This dataset contains 396 non-redundant sequences<br> derived from the 3Dee database created by Cuff and Barton (Cuff & Barton, 1999). It contains 396<br> proteins with their respective secondary structure as shown in Figure 1.</p>
Data to accompany the paper "Improved fragment-based protein structure prediction by redesign of search heuristics"
<p>This repository contains the older and newer input fragment sets and other data used for the analyses in our paper. The filenames for each tarball contain the PDB identifier of each protein along with a chain ID if applicable, followed by 'old' or 'new' for old and new fragments, respectively. Each tarball contains: a .fasta file of the input sequence, a matching PDB structure file, the relevant PSIPRED secondary structure prediction file, and the 9mer and 3mer fragment files. <br> <br> An additional tarball, ScoreRMSDplots_3protocols.tgz, contains extended versions of Figure 3 which show score and RMSD distributions clearly. Additionally, the same data is shown for equivalent experiments using the older fragment set.</p>
Extended data for the paper "Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction"
<p>Extended data for the paper:<br> Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction</p> <p>Authors:<br> Shaun M Kandathil, Mario Garza-Fabre, Simon C Lovell and Julia Handl</p> <p>--------------------------------</p> <p>Contents of the zip file:</p> <p> </p> <p>Directory 'ECDFplots':<br> ----------------------<br> Data corresponding to Figure 3 for all targets, for the bilevel and ILS protocols. Data are available following stages 3 and 4 of the low-resolution protocol.</p> <p>Directory 'ScoreRMSDplots_3archivers':<br> --------------------------------------<br> Data corresponding to Figures 6 and 9 for all targets. Data corresponding to decoys obtained after low-resolution stages 3 and 4 can be found in subdirectories 'Stage3' and 'Stage4', respectively.<br> </p>
DPCstruct Classification of AlphaFold2-Predicted Protein Structures
<p>This dataset contains DPCstruct domain classifications for protein structures predicted by AlphaFold2, as presented in the paper "Unsupervised Domain Classification of AlphaFold2-Predicted Protein Structures."</p> <p>DPCstruct was applied to a non-redundant set of the AlphaFold Database v4.0, known as Foldseek Clusters, which includes approximately 15 million representative proteins, as described in the work by <a href="https://doi.org/10.1038/s41586-023-06510-w">Barrio-Hernandez et al.</a></p> <p>This repository provides the results of our classification, along with all the data related to the analyses presented in our study. DPCstruct algorithm can be found at <a href="https://github.com/RitAreaSciencePark/DPCstruct">https://github.com/RitAreaSciencePark/DPCstruct</a> together with examples on how to use it.</p> <p><strong>FILES DESCRIPTION:</strong></p> <ul> <li><strong>dpcstruct_classification.tsv: </strong>List of domains identified by DPCstruct and their corresponding metacluster. Columns: Metacluster ID, Protein Uniprot ID, domain start, domain end.</li> <li><strong>mcs_reps.fasta:</strong> For each metacluster, two representative domains were selected: one representing the center of the cluster and the other being the domain with the highest pLDDT score. If these are the same, only one domain is included as the representative. This file contains the list of representative domains and their sequences in FASTA format.</li> <li><strong><span>mcs_reps_pdbs.zip: </span></strong>Contains a PDB file for each representative domain. The filename is structured as 'proteinID_metacluster.pdb'.</li> <li><strong>mcs_properties.tsv:</strong> Set of properties per metacluster, including: <ul> <li><strong>mcID:</strong> Metacluster ID.</li> <li><strong>size:</strong> Number of domains.</li> <li><strong>len_aa:</strong> Average length of domains (number of amino acids).</li> <li><strong>len_std:</strong> Standard deviation of domain lengths.</li> <li><strong>len_ratio:</strong> Ratio of len_std to len_aa.</li> <li><strong>plddt:</strong> Average predicted LDDT as reported by AlphaFold2.</li> <li><strong>disorder:</strong> Average intrinsic disorder score calculated with AIUPred.</li> <li><strong>alntmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs.</li> <li><strong>tmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs, using the maximum between TM-score normalized by query or target.</li> <li><strong>lddt:</strong> Pairwise LDDT score, averaged over all pairs.</li> <li><strong>prob:</strong> Pairwise probability of homology according to SCOPe, as reported by Foldseek.</li> <li><strong>pident:</strong> Pairwise percentage identity, averaged over all pairs.</li> </ul> </li> <li><span><strong>annotated_[cath|scop]_qc[x]_t[x]_l[x].tsv:</strong> </span>For each fold in [CATH|SCOP], we provide the best matching DPCstruct domain, if available, along with the structural alignment information as reported by Foldseek. A fold is considered annotated if its alignment values meet or exceed the following thresholds: <ul> <li>qc: query coverage.</li> <li>t: template modelling score of the alignment.</li> <li>l: lddt score of the alignment.</li> </ul> </li> <li><strong>dpcstruct_consistency.tsv:</strong> Consistency of DPCstruct metaclusters with respect to Pfam 36.0 labels. Note that we consider a Pfam label to overlap with a DPCstruct domain even if it shares just one amino acid, which is why some metaclusters have many labels. In such cases, we only display 5 representative labels.</li> <li><strong>pfam_consistency.tsv:</strong> Consistency of Pfam Clans with respecto to DPCstruct labels.</li> </ul> <p><strong>Note:</strong> All 'tsv' files contain a header as the first row.</p> <p>If there is any doubt regarding the data or there is something missing please contact us: </p> <p>federico.barone@areasciencepark.it</p>
Protein structure model predictions for secreted fungal proteins
<p><strong>Dataset A - Alphafold2 prediction output data for 753 secreted proteins of <em>Rhizophagus irregularis </em>DAOM197198</strong>. Gene IDs are taken from the annotation by Yildirir et al. 2021, <a href="https://doi.org/10.1111/nph.17842">doi.org/10.1111/nph.17842</a></p> <p><strong>Dataset B - Alphafold2 prediction output data for 10 fungal effectors.</strong><strong> </strong>These are nine effectors from <em>Fusarium oxysporum</em> f. sp.<em> lycopersici</em> and RiSLM from <em>Rhizophagus irregularis</em> as well as their amino acid sequences. Signal peptides and sequences preceding a predicted Kex2 processing site were removed.</p> <p><strong>Dataset C - Alphafold2 prediction output data for 454 matches of a MycFOLD-HMM search</strong> across the Mycocosm genome database (<a href="https://mycocosm.jgi.doe.gov/mycocosm/home">https://mycocosm.jgi.doe.gov/mycocosm/home</a>) and 36 Glomeromycotina fungal genomes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.