Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

347

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

347 results for “protein structure”

Learn how ShareScore rates datasets ↗
zenodo56/100

Homologous membrane protein structures (HOMEP) version 1

<p><strong>Table 1</strong> = List of membrane protein structures in the <strong>HOMEP</strong>&nbsp;data set (version 1).<br> From Forrest, Tang&nbsp;&amp; Honig&nbsp;2006 Biophysical Journal (Supplementary Table 1)<br> <a href="https://www.ncbi.nlm.nih.gov/pubmed/16648166">https://www.ncbi.nlm.nih.gov/pubmed/16648166</a></p> <p>Contains the following columns:<br> PDB-Code Protein-Name &nbsp; &nbsp;Source &nbsp;Res-(&Aring;) Length (Num-TM) Number-of-TM-domains &nbsp; &nbsp;Family</p> <p><strong>Table 2</strong> =&nbsp;List of pairs of membrane protein structures in the <strong>HOMEP</strong>&nbsp;data set (version 1).<br> From Forrest, Tang&nbsp;&amp; Honig&nbsp;2006 Biophysical Journal (Supplementary Table 2)</p> <p>Contains the following columns:<br> Model &nbsp; Family &nbsp;Query &nbsp; Template &nbsp; ID(%) &nbsp;RMS(&Aring;) GDT_TS(%) &nbsp;TM-ID(%) &nbsp; TM-RMS(&Aring;) &nbsp;TM GDT_TS(%)</p> <p><strong>Table 3 </strong>= Manually-defined transmembrane regions in the <strong>HOMEP</strong>&nbsp;data set (version 1),&nbsp;listed for each family by transmembrane segment number. From Forrest, Tang&nbsp;&amp; Honig&nbsp;2006 Biophysical Journal (Supplementary Table 3).</p> <p>Contains&nbsp;the columns defined as follows:<br> Protein chain identifier, start (-s) and end (-e) residues for each PDB structure in the family</p>

opencc-by-4.0Apr 2006View details →
zenodo56/100

Cross-phyla protein annotation by structural prediction and alignment

<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is&nbsp;typically limited to a few model organisms. In non-model species, the sequence-based prediction of&nbsp;gene orthology can be used to infer protein identity, however this approach loses predictive power&nbsp;at longer evolutionary distances. Here we propose a workflow for protein annotation using structural&nbsp;similarity, exploiting the fact that similar protein structures often reflect homology and are more&nbsp;conserved than protein sequences.</p> <p><strong>Results:</strong>&nbsp;&nbsp;We propose a workflow of openly available tools for the functional annotation of proteins via&nbsp;structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete&nbsp;proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet&nbsp;their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with&nbsp;known homology in &gt;90%&nbsp;cases, and annotates an additional 50%&nbsp;of the proteome beyond&nbsp;standard sequence-based methods. We uncover new functions for sponge cell types, including extensive&nbsp;FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in&nbsp;myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes,&nbsp;proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>

opencc-by-4.0Mar 2023View details →
zenodo52/100

Structures of S-protein in complex with ligands deposited in the PDB between the 1st January 2021 and the 13th May 2021

<p>All 174 structures of SARS-CoV-2 S-protein in complex with a ligand released on the PDB between the 1<sup>st</sup> January 2021 and the 13<sup>th</sup> May 2021 were categorised by ligand type: hACE2, antibody Fab fragments, VHH antibody fragments or <em>de novo</em> designed peptide scaffolds. Information concerning the method by which the structures were determined and their resolution were retrieved from the PDB. The categorisation of ligands by S-protein binding site were achieved by visual analysis of all the structures using molecular visualisation software PyMOL, in which no new binding sites were found beyond those already categorised for the structures released on the PDB until the 1<sup>st</sup> January 2021 (10.5281/zenodo.5503855).</p> <p>The Pure project is funded by the European Union&rsquo;s Horizon 2020 program under grant agreement No. 899732.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo52/100

List of the structures of S-protein in complex with ligands deposited in the Protein Data Bank until the 1st January 2021.

<p>All 131 structures of SARS-CoV-2 S-protein in complex with a ligand released on the PDB until the 1<sup>st</sup> January 2021 were categorised by ligand type: hACE2, antibody Fab fragments, VHH antibody fragments or <em>de novo</em> designed peptide scaffolds. The ligands&rsquo; amino acid sequences, the method by which the structures were determined and their resolution were retrieved from the PDB. Information regarding the ligands&#39; production method, dissociation constants (K<sub>D</sub>), S-protein segment against which the K<sub>D</sub> were measured and the determination methods were retrieved from the respective references. The categorisation of ligands by S-protein binding site and listing of S-protein conformation in each structure were achieved by visual analysis of all the structures using molecular visualisation software PyMOL.</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

Bioactivity deep learning for structure-free compound-protein interaction

<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p>&nbsp;</p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p>&nbsp;</p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Protein structure files for the paper "Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth" in Nature Cell Biology by Tang et al.

<p>This archive contains models of HRAS, KRAS, and NRAS homo- and heterodimers with various mutations discussed in the paper,&nbsp; &quot;Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth&quot; in Nature Cell Biology by Tang et al.<br> as well as crystallographic dimers of these proteins as identified by the ProtCAD database, http://dunbrack2.fccc.edu/ProtCAD/Results/PfamArchClusterInfo.aspx?GroupId=8 (cluster 5). Several of the models are shown in Supp. Figure 11b and the crystallographic dimers of RAS that provide evidence for the possible biological relevance of these models are shown in Supp. Figure 11a.</p> <p>The crystallographic dimers were identified by clustering all possible interfaces generated by symmetry operators in crystals of HRAS, KRAS, and NRAS as described in the paper: Xu, Q., Dunbrack, R.L. ProtCID: a data resource for structural information on protein interactions. <em>Nat Commun</em> <strong>11</strong>, 711 (2020). https://doi.org/10.1038/s41467-020-14301-4.</p> <p>The models were created by superposing monomers of HRAS, KRAS, or NRAS onto the alpha4-alpha5 dimer present in the crystal of PDB entry 3k8y. Mutations were made in PyMOL. The structures were relaxed with the FastRelax protocol and the Ref2015 scoring function in the program Rosetta, which uses the backbone-dependent rotamer library of Shapovalov and Dunbrack to repack side chains.</p> <p>The crystallographic dimers are contained in a zipped PyMOL session. The mmCIF format for all the structures is present in a zip file, Tang_et_al_crystallographic_and_modeled_RAS_dimer_ciffiles.zip. The PyMOL session and zip file contains 87 HRAS dimers, 14 KRAS dimers, and 1 NRAS dimer, all having the interface consisting of the alpha4 and alpha5 helices. The PyMOL session also contains the modeled structures. Only Mg ions and GTP/GNP/GDP ligands are shown. Others are present but hidden and may be displayed by PyMOL (&quot;show sticks, het&quot;).</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Homologous membrane protein structures (HOMEP) dataset version v2

<p><strong>Protein structures from the dataset of&nbsp;Homologous MEmbrane Protein structures (HOMEP)</strong> version v2 created in 2010, published in 2013. A more automated version of HOMEP v1:&nbsp;<a href="https://doi.org/10.5281/zenodo.2646534">10.5281/zenodo.2646534</a><br> &nbsp;</p> <p><strong>Table 1</strong> = List of protein databank&nbsp;structure entries<br> From Stamm et al, PLOS One 2013,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/23469223">https://www.ncbi.nlm.nih.gov/pubmed/23469223</a>, Supplementary Table 1, with the following entries:<br> Family grouping, Protein databank identifier, Name, Source organism, Resolution (&Aring;)</p> <p>&nbsp;</p> <p><strong>Table 2</strong> = List of pairs of structures<br> From Stamm et al, PLOS One 2013,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/23469223">https://www.ncbi.nlm.nih.gov/pubmed/23469223</a>, Supplementary Table 2, with the following entries:<br> Family grouping, PDB code for first structure, Chain ID from PDB1, PDB for second structure, Chain ID from PDB2, protein structural difference (PSD), % sequence identity</p> <p>&nbsp;</p> <p><strong>File S2 HOMEP2 Dataset.tar.gz</strong> = Protein databank format files (PDB) are attached in the Dataset tar zipped file,&nbsp;organized by family.&nbsp;From Stamm et al, PLOS One 2013,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/23469223">https://www.ncbi.nlm.nih.gov/pubmed/23469223</a>, Supplementary dataset.</p>

openother-openMar 2013View details →
zenodo48/100

Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution

<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook&nbsp;</li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

Discoba protein sequences for protein structure predictions

<p>Comprehensive database of Discoba protein sequences, gathered for the purpose of improving protein structure predictions of Discoba species (including <em>Trypanosoma </em>and <em>Leishmania</em>) by AlphaFold and RoseTTAFold. Originally gathered for use with:&nbsp;https://github.com/zephyris/discoba_alphafold</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Dynamics of SARS-CoV-2 spike protein in open and closed states and identification of key structural perturbations upon mutations

<p>The SARS-Cov-2 spike protein resides on the exterior surface of the coronavirus, and therefore, acts as the first point of contact that mediates cell attachment and fusion. &nbsp;During this process, it undergoes dramatic conformational changes upon host receptor binding. We are leveraging high-performance computing to identify these structural perturbations in wildtype and mutant spike protein models. The files contain structures from molecular dynamics simulations of closed SARS-Cov-2 spike protein embedded in POPC membrane.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Global comparative structural analysis of responses to protein phosphorylation

<p>This contains the structures and data used for the structural analysis presented in <em>Global comparative structural analysis of responses to protein phosphorylation</em> (Correa Marrero et al.,&nbsp;https://doi.org/10.1101/2024.10.18.617420 ). To summarize:</p> <ul> <li>filtered_df.xlsx: dataset of paired phosphorylated structures and their non-phosphorylated counterparts. Each row contains one such pair.</li> <li>chains_by_protein.zip: each directory (named with a UniProt ID) contains the used structures that form the basis for the analysis. The structures are in PDB format, in a separate directory for each protein in the dataset. The exception is the annotation_per_psite directory, which contains annotation as a csv file for each phosphosite.</li> <li>extracted_domains.zip: contains structures of Pfam domains (extracted from the previous dataset) in PDB format. Each filename follows the format {PDB ID}_{Chain ID}_{Pfam domain ID}. The domain_coverage.csv file lists the domain coverage of the structure, as well as its length compared to the whole sequence and the whole structure it was extracted from. These are the structures used for the analysis shown in Fig. 1 f-h.</li> <li>extracted_pfam_domains.zip: contains structures of a broader set Pfam domain structures (the whole set of Pfam domains found to contain a phosphosite in filtered_df.csv) in PDB format. Each directory (named with the Pfam ID) contains the structures. merged_pfam_data.tsv contains&nbsp;metadata about the structures (structure quality, coverage of the domain structure, phosphosite location...). These are the structures used for the analysis shown in Fig. 2.</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo44/100

PSSH2 - database of protein sequence-to-structure homologies (including Sars-CoV-2 structures)

<p><strong>Protein sequence and structure data</strong></p> <p>This data set contains data from Uniprot (in the files called protein_sequence, protein_synonyms, protein_names, organism_synonyms) and PDB (in the files called PDB and PDB_chain) as used by the <a href="https://github.com/ODonoghueLab/Aquaria">Aquaria web resource</a> at the time of download (2022-02-08).</p> <p>&nbsp;</p> <p><strong>The&nbsp;PSSH2 data set</strong><br> <br> PSSH2 is a database of protein sequence-to-structure homologies based on HHblits, an alignment method employing iterative comparisons of hidden Markov models (HMMs). To ensure the highest possible final alignment quality for matches in Aquaria using HHblits, we first calculate HMM profiles for each unique PDB sequence (PDB_full) and also for each unique Swiss-Prot sequence. We generated PSSH2 using HHblits to find similarities between HMMs from PDB and HMMs from UniProt sequences.</p> <p>&nbsp;</p> <p><strong>Calculating PSSH2</strong></p> <p>The&nbsp;Swissprot and PDB data was downloaded in November 2021.<br> Generating PSSH2: We used <a href="https://gwdu111.gwdg.de/~compbiol/uniclust/2021_03/UniRef30_2021_03.tar.gz">UniRef30_2021_03</a> (originally called UniRef30_2021_06)&nbsp;from HH-suite, a database of non-redundant UniProt sequence clusters in which the highest pairwise sequence identity between clusters was 30%. The HHblits code and the code for running the calculations&nbsp;was retrieved from git (https://github.com/soedinglab/hh-suite.git and https://github.com/aschafu/PSSH2.git respectively)&nbsp;at the respective&nbsp;time of calculation in the timeframe until December&nbsp;2021.&nbsp;<br> &nbsp;</p> <p><strong>PDB based sequence-to-structure alignments</strong></p> <p>In addition to the PSSH2 data, new PDB structures were retrieved based on the primary accession of the proteins, by querying for all chains in all PDB entries with exact matches using the sequence cross references records given in PDB. Sequence-to-structure alignments were then created, again based on information provided in each PDB entry. These are contained in the PDBchain data.</p> <p>This data covers sequences and PDB structures in the timeframe until February 2022.&nbsp;</p> <p>&nbsp;</p> <p><strong>Evaluating PSSH2</strong></p> <p>The resulting alignment data was analysed using CATH domain assignments downloaded from&nbsp;/cath/releases/all-releases/v4_2_0/cath-classification-data/ to define correct hits and false hits:&nbsp;</p> <ul> <li>The set of query sequences is defined by the CATH non-redundant S40_overlap_60 dataset (ftp://orengoftp.biochem.ucl.ac.uk/cath/releases/all-releases/v4_2_0/non-redundant-data-sets/)</li> <li>The set of all expected hits are all pdb structures containing a domain with the same CATH code if contained in the set of processed sequences (-&gt; all) or&nbsp;only if also contained in the set of non redundant sequences (-&gt; nr40).</li> <li>The set of true positives is defined by sharing the same CATH code up to the level of homology (&quot;CATH&quot;) or up to the level of topology (&quot;CAT&quot;).</li> </ul> <p>The data was evaluated with respect to false discovery rate (FDR) and recall (true positive rate TPR) by cumulatively considering all hits with an E-value below the threshold (&quot;C&quot;) or in bins with an E-value between the threshold and one tenth of the threshold (&quot;B&quot;). This evaluation was carried out for the data obtained in November 2021 (202111)&nbsp;as well as previous data from October 2020 (202010), February 2020 (202002) and&nbsp;September 2017 (201709). The results are&nbsp;&nbsp;collected in&nbsp;<a href="https://zenodo.org/api/files/445add84-fcf1-4dfe-b8a1-63dc55f378ee/PSSH%20CATH%20validation.csv?versionId=a5df7473-6efd-442b-b422-3944e9452003">PSSH CATH validation.csv</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Known errors</strong></p> <p>Due to processing error, the profile of pdb structure 5fia A / B (sequence md5 052667679fc644184f40063c7602c9e1) is incomplete in the pdb_full hhblits database which led to further errors in generating sequence based alignments for sequences for 1vtm P (sequence md5 c844aff103449363cb8489c78c58ebf1) and 434t A / B (sequence md5 d67aa1c3a36492c719cb48b5e7ecc624).<br> <br> &nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "

<p>Datasets for&nbsp;preprint&nbsp;entitled &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi&quot;</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of&nbsp;<em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi&quot;.&nbsp;&nbsp;&nbsp;</p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict.&nbsp;For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22.&nbsp;In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using&nbsp;Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes.&nbsp;</p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset&nbsp;is made up of predicted protein tertiary structures representing the main member of each up-regulated&nbsp;<em>V. inaequalis</em> effector candidate family. Structures were predicted using&nbsp;Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>).&nbsp;In cases where&nbsp;the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2.&nbsp;Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold&nbsp;(<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>)&nbsp;open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input.&nbsp;</p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence&nbsp;proteins from other fungi&quot; study. These structures were predicted using&nbsp;Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input.&nbsp;</p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>

opencc-by-2.0Feb 2022View details →
zenodo44/100

Characterizing and explaining impact of disease-associated mutations in proteins without known structures or structural homologues

<p>AlphaFold and RoseTTAFold models of domains of disease associated human proteins without structures/known homologues.</p> <p>Tables containing the model quality, model region, sequence&nbsp;alignment statistics, matched FunFam, associated GO terms for the FunFam, ddG of mutation, pathogenicity of mutation,&nbsp;if mutation is near a predicted functional site (conserved residue/ligand binding site/protein-protein interface)</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Identifying and profiling structural similarities between Spike of SARS-CoV-2 and other viral or host proteins with Machaon - Pre-computed features for replication

<p>Machaon&#39;s computed features that were used in the structural comparisons with Spike protein.</p> <p>DATA_PDBS_vir_whole_1-3.zip files are parts of a single folder.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

X-Ray Diffraction data from Membrane transport protein AcrB, V612F mutant with bound minocycline, source of 9FHC structure

<p>Crystals were grown of the membrane transport protein AcrB, V612F mutant, with bound minocycline.&nbsp;</p> <p>X-ray diffraction data of this upload: 400 frames of 0.5&deg; width were collected on 2007-04-30 at the X06SA beamline of Swiss Light Source at Paul-Scherrer-Institute (Switzerland).</p> <p>The data can be processed with XDS; XDS.INP is provided as part of the upload.</p> <p>The data are the basis of the PDB 9FHC structure.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Structural basis of actin monomer re-charging by cyclase-associated protein

<p>1) table_of_simulations.pdf:&nbsp; table of simulations</p> <p>2) toppar_HIC.str: methylhistidine (HIC) topologies and parameters</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; -prepared based on analogy</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; -to be used with top_all36_prot.rtf and par_all36_prot.prm</p> <p>3) simulation_archive.tar.gz</p> <p>&nbsp;&nbsp;&nbsp; The Contents:</p> <p>1_ADP-Actin--CARP, 2_ADP-Actin--CAP1, 3_ATP-Actin--WH2, 4_ADP-Actin<br> All systems presented in the paper; see table_of_simulations.pdf<br> Each directory contains<br> 000README&nbsp; gromacs_topologies&nbsp; gromacs_tpr_files&nbsp; index.ndx&nbsp; processed_trajectories&nbsp; prod.mdp&nbsp; systems_at_t=0</p> <p>*** The rosetta models for WH2 domain and the proline-rich loop that connects it to the CARP domain can be found in&nbsp; 2_ADP-Actin--CAP1/rosetta_models</p> <p><br> _Topologies:<br> &nbsp;&nbsp;&nbsp; toppar_c36_jul16:<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The charmm force field version used to generate topologies before conversion to gromacs; see 000README in the systems directory<br> &nbsp; &nbsp;<br> &nbsp;&nbsp;&nbsp; ***toppar_c36_jul16/toppar_HIC.str: The topology and parameters for methylated histidine used in the simulations.</p> <p>&nbsp;&nbsp;&nbsp; gromacs_topologies:<br> &nbsp;&nbsp;&nbsp; Contains all itp files (converted from&nbsp; psf file using PyTopol&#39;s psf2top utility) and parameters.<br> &nbsp;&nbsp;&nbsp; Note that relevant files can also be found in directories corresponding to each system ( 1_ADP-Actin--CARP&nbsp; 2_ADP-Actin--CAP1&nbsp; 3_ATP-Actin--WH2&nbsp; 4_ADP-Actin)</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo44/100

Homologous membrane protein structures (HOMEP) version 3

<p><strong>Homologous membrane protein structures (HOMEP)</strong> version 3 (created 2013)<br> An updated version of v2&nbsp;<a href="https://doi.org/10.5281/zenodo.2646539">10.5281/zenodo.2646539</a>&nbsp;and&nbsp;v1:&nbsp;<a href="https://doi.org/10.5281/zenodo.2646534">10.5281/zenodo.2646534</a></p> <p><strong>Table 1</strong> =&nbsp;Alpha-helical membrane protein structures&nbsp;in the HOMEP3 data set (2013), listed by family<br> From Stamm M, Forrest LR, Proteins 2015 (Supplementary Table 1):<a href="http://https://www.ncbi.nlm.nih.gov/pubmed/26178143">&nbsp;https://www.ncbi.nlm.nih.gov/pubmed/26178143</a>&nbsp;</p> <p>Contains the following columns:<br> Protein family, Protein databank identifier, Chain identifier, Name, Source organism, Resolution (&Aring;)</p> <p>&nbsp;</p> <p><strong>Table 2</strong> =&nbsp;Beta-barrel&nbsp;membrane protein structures&nbsp;in the HOMEP3 data set (2013), listed by family<br> From Stamm M, Forrest LR, Proteins 2015 (Supplementary&nbsp;Table 2)</p> <p>Contains the following columns:<br> Protein family, Protein databank identifier, Chain identifier, Name, Source organism, Resolution (&Aring;)</p> <p>&nbsp;</p> <p><strong>HOMEP3_pairs_alpha.txt</strong><br> List of all pairs of protein structure chains of&nbsp;alpha-helical proteins</p> <p>&nbsp;</p> <p><strong>HOMEP3_pairs_beta.txt</strong><br> List of all pairs of protein structure chains of&nbsp;beta-barrel proteins</p> <p>&nbsp;</p> <p><strong>HOMEP3_pdbs.tar.gz</strong><br> All pdb files for individual chains&nbsp;in both alpha-helical and beta-barrel subsets</p>

openother-atJul 2015View details →
zenodo44/100

Critical Assessment of automated Structure Determination of Proteins by NMR

<p>The community-wide initiative &quot;Critical Assessment of Automated Structure Determination of Proteins by NMR (<strong>CASD-NMR</strong>)&quot; was launched in 2009 to to evaluate the ability of automated methods to produce 3D protein structures from NMR data that closely match structures manually determined by experts.</p> <p>This dataset includes all the experimental data made available to the participants of CASD-NMR in the two completed rounds of the initiative.</p> <p>Also refer to http://www-nmr.cabm.rutgers.edu/blindtest/blind.html for additional details, including first release date and link to each final PDB entry</p>

opencc-by-4.0Aug 2015View details →
zenodo44/100

Project files provided as supporting information to the manuscript "A deep learning approach to the structural analysis of proteins"

<p><strong>README file to the project files provided as supporting information to the manuscript &ldquo;A deep learning approach to the structural analysis of proteins&rdquo;</strong></p> <p>Dec. 30, 2018</p> <p>Authors: Marco Giulini and Raffaello Potestio</p> <p>==================================</p> <p>The dataset contains the following files:</p> <p>&nbsp;</p> <p>- datasets.zip: archive containing five .csv files, namely:</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - decoys_cm.csv : all the data for 10728 protein decoys, training set</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - evaluation_cm.csv : all data for 146 proteins in the evaluation set</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - random_CG.csv : 1200 Coulomb matrices. 100 CG models for each protein with 120 amino acids</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - 1e5g_centered_sphere.csv : 100 CG models in which the central atoms in 1e5g are not removed</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - 1e5g_random_sphere.csv : 10 CG models for 10 different (random) locations for the sphere that includes atoms that have to be retained. 100 CG models in total</p> <p>&nbsp;</p> <p>- decoys_labels.lab containing the labels associated to the 10728 decoys present in the training set</p> <p>- evaluation_labels.lab containing the labels associated to the 146 pdb files in the evaluation set</p> <p>- random_CG_labels.lab containing the labels associated to the 6 proteins with 120 amino acids</p> <p>- network_development_training: a python script that performs cross validation and full training of the model</p> <p>- saved_networks.zip FOLDER containing 10 networks: the architecture is included in .json files while weight parameters are inside .hs files</p> <p>&nbsp;</p> <p>- pdb_files.zip&nbsp;FOLDER containing the PDB files that have been employed in the project, namely:</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - pdb_files_len100 : pdb files with 100 amino acids</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - pdb_files_len101-110 : pdb files with a number of amino acids between 101 and 110</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - decoys : decoys of length 100 extracted from the above folder: name syntax == PDBNAME_decoy_STARTRES_ENDRES.pdb</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; EXAMPLE 6gsp.pdb will give rise to 6gsp_decoy_0_100.pdb , 6gsp_decoy_1_101.pdb , 6gsp_decoy_2_102.pdb , 6gsp_decoy_3_103.pdb&nbsp; , 6gsp_decoy_4_104.pdb</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - pdb_files_len100 : 6 pdb files with 120 amino acids</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record