Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
347
datasets available to search
ShareScore release 0.9.0
Dataset results
347 results for “structural proteins”
Fig. 4 in Structural characterization of the Pet c 1.0201 PR-10 protein isolated from roots of Petroselinum crispum (Mill.) Fuss
Fig. 4. Superpositions of optimized tertiary structures of the Pet c 1.0201 3D homology model and the Api g 1.0101 crystal structure (PDB accession 2BK0). A: blue – the template structure of Api g 1, red – the 3D model of Pet c 1.0201, both with annotated secondary structure elements: α – α-helix, β – β-sheet, L – loop, and N- and C-terminals. B: monomers of Api g 1.0101 and C: monomers of Pet c 1.0201 after 200 ns MD simulations. Red indicates the position of Glu45 responsible for IgE binding. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 2 in Structural characterization of the Pet c 1.0201 PR-10 protein isolated from roots of Petroselinum crispum (Mill.) Fuss
Fig. 2. Peptide sequences of the Pet c 1.021 protein purified from the Petroselinum crispum roots, where identified tryptic (red types) and chymotryptic (blue types) peptides are highlighted. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
A3D database: structure-based predictions of protein aggregation for the human proteome
<p>A3D database: structure-based predictions of protein aggregation for the human proteome</p>
Understanding structural and functional diversity of ATP-PPases using protein domains and functional families in CATH database
<p>The dataset of AF2-predicted HUP domains with overall pLDDT > 90, culled at 90% identity.</p>
Peak lists, assignments, and structures of 100 proteins determined by the ARTINA algorithm
<h2> </h2><p>The time-consuming and complex data analysis process constitutes a limitation for biomolecular NMR. It typically requires weeks or months of manual work of a trained expert to identify and interpret thousands of signals, recorded in a series of multidimensional NMR spectra, in order to determine the sequence-specific resonance assignments or the three-dimensional structure of a single protein. </p><p>To solve this problem, we developed ARTINA, a machine learning-based method that uses as input only NMR spectra and the protein sequence, and delivers signal positions, resonance assignments, and structures strictly without any human intervention. Tested on a 100-protein benchmark comprising 1329 multidimensional NMR spectra, ARTINA demonstrated its ability to solve structures with 1.44 Å median RMSD to the PDB reference and to identify 91.36% correct NMR resonance assignments. </p><p>ARTINA is freely accessible for non-commercial users at NMRtist (https://nmrtist.org), an online platform that combines deep learning, large-scale optimization, and cloud computing to offer full automation of NMR spectra analysis. Our website provides virtual storage for NMR spectra deposition together with a set of applications designed for automated peak picking, chemical shift assignment, and protein structure determination. The system can be used by non-experts and without IT infrastructure on the user's side, allowing protein NMR spectra interpretation within hours after completion of the measurements. With NMRtist, the effort for a protein assignment or structure determination by NMR essentially reduced to the preparation of the sample and the spectrum measurements. </p>
Structure analysis of a p53 fusion protein
Open the record for dataset details and reuse information.
Structure-guided discovery of potent antifungals that prevent Ras signaling by inhibiting protein farnesyltransferase
Open the record for dataset details and reuse information.
Protein structure alignments and structural similarity score code for: Benchmarking methods of protein structure alignment
Open the record for dataset details and reuse information.
A hybrid structure determination approach to investigate the druggability of the nucleocapsid protein of SARS-CoV-2
Open the record for dataset details and reuse information.
Data from: Spider silk colour co-varies with thermal properties but not protein structure
Open the record for dataset details and reuse information.
Deep-time structural evolution of retroviral and filoviral surface envelope proteins
Open the record for dataset details and reuse information.
Data from: The evolution of protein-coding gene structure in eukaryotes
Open the record for dataset details and reuse information.
Annotation and analysis of the secondary structure elements in the Cytochrome P450 protein family
<p>We collected all currently available structures for proteins in the Cytochrome P450 family and annotated their secondary structure elements using SecStrAnnotator software (https://webchem.ncbr.muni.cz/Wiki/SecStrAnnotator). We used 2nnjA as the template domain for the annotation. Based on these annotations, we analysed the occurrence, length distribution, amino acid sequence, and presence of structural irregularities (β-bulges, 3<sub>10</sub>-helices, π-helices) of each secondary structure element class. We also statistically compared the bacterial vs eukaryotic structures. For the secondary structure element classes with sufficient sequence conservation, the most conserved residue is annotated as the reference residue.</p> <p>Main files:</p> <ul> <li><strong>set_ALL.json</strong> - Set-ALL: list of 1012 protein domains belonging to the Cytochrome P450 family (CATH accession 1.10.630.10 + Pfam accession PF00067, accessed on 7 July 2020, one domain per PDB entry)</li> <li><strong>set_NR.json</strong> - Set-NR: non-redundant list of 183 domains (one domain per UniProt ID)</li> <li><strong>domain_lists_table.tsv</strong> - Overview of Set-ALL and Set-NR and separation into subsets Set-NR-Bact (bacterial), Set-NR-Euka (eukaryotic), Set-NR-Arch (archaeal), Set-NR-Viru (viral)</li> <li><strong>structures/template_2NNJ-template.sses.json</strong> - Manually prepared annotation template (domain 2nnjA)</li> <li><strong>structures/template_2NNJ.cif</strong> - Structure of the template domain (2nnjA)</li> <li><strong>annotations_with_reference_residues_ALL.json, annotations_with_reference_residues_ALL.tsv</strong> - Annotation of secondary structure elements for Set-ALL</li> <li><strong>annotations_with_reference_residues_NR.json, annotations_with_reference_residues_NR.tsv</strong> - Annotation of secondary structure elements for Set-NR</li> <li><strong>aligments_NR</strong> - Multiple sequence alignments for each SSE class (Set-NR)</li> <li><strong>logos_NR</strong> - Sequence logos for each SSE class (Set-NR)</li> <li><strong>plots</strong> - Plots of SSE occurrence, length distribution, contained helix types and beta-bulge occurrence (Set-NR), some plots show the comparison between Set-NR-Bact and Set-NR-Euka</li> <li><strong>statistical_tests.ods</strong> - Comparison of SSE occurrence between Set-NR-Bact and Set-NR-Euka by the test of equal proportions and the Fisher test, comparision of the SSE length by the Kolmogorov-Smirnov test and the two-sample Wilcoxon test</li> </ul>
Supplementary Structural Models (SARS-CoV-2 Spike-RBD:ACE2 complex and TMPRSS2) - SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals
<p>Structural Models (PDB) of SARS-CoV-2 Spike RBD bound to ACE2 receptors of 215 animals.</p> <p>Structural model of Human TMPRSS2.</p> <p>Modelled using the FunMod pipeline and referenced in the preprint</p> <p><a href="https://www.biorxiv.org/content/10.1101/2020.05.01.072371v5">SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals</a></p> <p> </p>
Host plant defense produces species-specific alterations to flight muscle protein structure and flight-related fitness traits of two armyworms
<p>Insects manifest phenotypic plasticity in their development and behavior in response to plant defenses, via molecular mechanisms that produce tissue-specific changes. Phenotypic changes might vary between species that differ in their preferred hosts and these effects could extend beyond larval stages. To test this, we manipulated the diet of southern armyworm (SAW; Spodoptera eridania) and fall armyworm (FAW; Spodoptera frugiperda) using a tomatomutant for jasmonic acid plant defense pathway (def1), and wild-type plants, and then quantified gene expression of Troponin t (Tnt) and flight muscle metabolism of the<br> adult insects. Differences in Tnt spliceform ratios in insect flight muscles correlate with changes to flight muscle metabolism and flight<br> muscle output. We found that SAW adults reared on induced def1 plants had a higher relative abundance (RA) of the A isoform of Troponin t (Tnt A) in their flight muscles; in contrast, FAW adults reared on induced def1 plants had a lower RA of Tnt A in their flight muscles compared with adults reared on def1 and controls. Although massadjusted flightmetabolic rate showed no independent host plant effects in either species, higher flight metabolic rates in SAW correlated with increased RA of Tnt A. Flight muscle metabolism also showed an interaction of host plants with Tnt A in both species, suggesting that host plants might be influencing flight muscle metabolic output by altering Tnt. This study illustrates how insects respond to variation in host plant chemical defense by phenotypic modifications to their flight muscle proteins, with possible implications for dispersal.</p>
Data from: Exploring the universe of protein structures beyond the Protein Data Bank
It is currently believed that the atlas of existing protein structures is faithfully represented in the Protein Data Bank. However, whether this atlas covers the full universe of all possible protein structures is still a highly debated issue. By using a sophisticated numerical approach, we performed an exhaustive exploration of the conformational space of a 60 amino acid polypeptide chain described with an accurate all-atom interaction potential. We generated a database of around 30,000 compact folds with at least 30% of secondary structure corresponding to local minima of the potential energy. This ensemble plausibly represents the universe of protein folds of similar length; indeed, all the known folds are represented in the set with good accuracy. However, we discover that the known folds form a rather small subset, which cannot be reproduced by choosing random structures in the database. Rather, natural and possible folds differ by the contact order, on average significantly smaller in the former. This suggests the presence of an evolutionary bias, possibly related to kinetic accessibility, towards structures with shorter loops between contacting residues. Beside their conceptual relevance, the new structures open a range of practical applications such as the development of accurate structure prediction strategies, the optimization of force fields, and the identification and design of novel folds.
Data from: Knowledge-based prediction of protein backbone conformation using a structural alphabet
Libraries of structural prototypes that abstract protein local structures are known as structural alphabets and have proven to be very useful in various aspects of protein structure analyses and predictions. One such library, Protein Blocks, is composed of 16 standard 5-residues long structural prototypes. This form of analyzing proteins involves drafting its structure as a string of Protein Blocks. Predicting the local structure of a protein in terms of protein blocks is the general objective of this work. A new approach, PB-kPRED is proposed towards this aim. It involves (i) organizing the structural knowledge in the form of a database of pentapeptide fragments extracted from all protein structures in the PDB and (ii) applying a knowledge-based algorithm that does not rely on any secondary structure predictions and/or sequence alignment profiles, to scan this database and predict most probable backbone conformations for the protein local structures. Though PB-kPRED uses the structural information from homologues in preference, if available. The predictions were evaluated rigorously on 15,544 query proteins representing a non-redundant subset of the PDB filtered at 30% sequence identity cut-off. We have shown that the kPRED method was able to achieve mean accuracies ranging from 40.8% to 66.3% depending on the availability of homologues. The impact of the different strategies for scanning the database on the prediction was evaluated and is discussed. Our results highlights the usefulness of the method in the context of proteins without any known structural homologues. A scoring function that gives a good estimate of the accuracy of prediction was further developed. This score estimates very well the accuracy of the algorithm (R2 of 0.82). An online version of the tool is provided freely for non-commercial usage at http://www.bo-protscience.fr/kpred/.
The 1.1 Å Structure of the Periplasmic Phosphate-Binding Protein from Stenotrophomonas maltophilia - a crystallisation contaminant identified by molecular replacement using the entire protein database (X-ray diffraction images).
<p>During efforts to crystallise the enzyme 2,4-dihydroxyacetophenone dioxygenase (DAD) from <em>Alcaligenes</em> sp. 4HAP, a small number of strongly diffracting protein crystals were obtained after two years of crystal growth in one condition. The crystals diffracted synchrotron radiation to almost 1.0 Å resolution and were, until recently, assumed to be formed by the DAD protein. However, when another crystal form of this enzyme was eventually solved at lower resolution, molecular replacement using this structure as the search model did not give a convincing solution with the original atomic resolution dataset. Hence we considered that these crystals might be due to a protein impurity, although molecular replacement using the structures of common crystallisation contaminants as search models again failed. A script to perform molecular replacement using MOLREP (Vagin, A. & Teplyakov, A. (2010). Acta Crystallogr. D 66, 22-25.) in which the first chain of every structure in the PDB was used as a search model was run on a multi-core cluster. This identified a number of prokaryotic phosphate binding proteins as scoring highly in the MOLREP peak lists. Calculation of an electron density map at 1.1 Å resolution allowed most of the amino acids to be identified visually and built into the model. A BLAST search then indicated that the molecule was most probably a phosphate binding protein from <em>Stenotrophomonas maltophilia</em> (UniProt ID: B4SL31; gene ID: Smal_2208) and fitting of the corresponding sequence to the atomic resolution map fully corroborated this. Proteins in this family have been linked with the virulence of antibiotic resistant strains of pathogenic bacteria and with biofilm formation. The structure has been refined to an R-factor of 10.15 % and an R-free of 12.46 % at 1.1 Å resolution. The molecule adopts the type-II periplasmic binding protein fold with a number of extensively elaborated loop regions. A fully-dehydrated phosphate anion is bound tightly between the two domains of the protein and interacts with conserved residues and a number of helix dipoles. </p>
Expansion of the RNAStructuromeDB to include secondary structural data spanning the human protein-coding transcriptome
<p>This dataset includes the -2, -1, and no filter z-score dot bracket files from ScanFold for all protein coding transcript isoforms.</p>
Ubiquity and evolution of structural maintenance of chromosomes (SMC) proteins in eukaryotes
<p>Structural maintenance of chromosomes (SMC) protein complexes are common in Bacteria, Archaea, and Eukaryota. SMC proteins, together with the proteins related to SMC (SMC-related proteins), constitute a superfamily of ATPases. Bacteria/Archaea and Eukaryotes are distinctive from one another in terms of the repertory of SMC proteins. A single type of SMC protein is dimerized in the bacterial and archaeal complexes, whereas eukaryotes possess six distinct SMC subfamilies (SMC1-6), constituting three heterodimeric complexes, namely cohesin, condensin, and SMC5/6 complex. Thus, to bridge the homodimeric SMC complexes in Bacteria and Archaea to the heterodimeric SMC complexes in Eukaryota, we need to invoke multiple duplications of an SMC gene followed by functional divergence. However, to our knowledge, the evolution of the SMC proteins in Eukaryota had not been examined for more than a decade. In this study, we reexamined the ubiquity of SMC1-6 in phylogenetically diverse eukaryotes that cover the major eukaryotic taxonomic groups recognized to date and provide two novel insights into the SMC evolution in eukaryotes. First, multiple secondary losses of SMC5 and SMC6 occurred in the eukaryotic evolution. Second, the SMC proteins constituting cohesin and condensin (i.e., SMC1-4), and SMC5 and SMC6 were derived from closely related but distinct ancestral proteins. Based on the above-mentioned findings, we discuss how SMC1-6 have diverged from the archaeal homologs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.