Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
347
datasets available to search
ShareScore release 0.9.0
Dataset results
347 results for “structural proteins”
Nanoscale chemical characterization of secondary protein structure of F-Actin using mid-infrared photoinduced force microscopy (PiF-IR)
<p>Raw data for manuscript for special issue in Spectrochimica Acta related to ECSBM 2022</p> <p> </p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Flavivirus Helicases -- Output Files
<p>These are the output files generated using the input files and Galaxy workflows for flavivirus helicase simulations, from: </p> <pre>https://doi.org/10.5281/zenodo.7493015</pre>
Artificial intelligence method to design and fold alpha-helical structural proteins from the primary amino acid sequence
<p>Dataset for paper: Z. Qin, L. Wu, H. Sun, S. Huo, T. Ma, E. Lim, P.-Y. Chen, B. Marelli, M.J. Buehler, Artificial intelligence method to design and fold alpha-helical structural proteins from the primary amino acid sequence, Extreme Mechanics Letters, Vol. 36, 100652, 2020. <a href="https://doi.org/10.1016/j.eml.2020.100652">https://doi.org/10.1016/j.eml.2020.100652</a>.</p> <p>Code: https://github.com/lamm-mit/MNNN/ </p>
Optimized structures for Optical control of ultrafast structural motion in a fluorescent protein
<p>QM-MM Optimized structures of the<strong> </strong>hydrogen bonding configuration in states A1, A2 and Transition State (TS) between them for rsKiiro protein on ground (s0) and first excited (s1) states. Structures were optimized at PBE0-D3/cc-pVDZ//Amber03 level.</p>
Trypanosoma brucei predicted protein structures, part 2 of 2
<p>AlphaFold2-predicted protein structures for the <em>Trypanosoma brucei</em> (TREU927) proteome, predicted using input multiple sequence alignments optimised for the Discoba lineage in which <em>T. brucei </em>sits. The structure prediction methodology was exactly as described in <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p> <p>This deposition contains part 2 of 2. To get the full dataset, also download TbruceiTREU927_part1.zip from <a href="https://zenodo.org/record/7940748">doi:10.5281/zenodo.7940748</a>.</p> <p>Data are organised with one directory per <em>T. brucei </em>TREU927 gene ID (eg. Tb927.1.3600). Within each directory you will find:</p> <p><strong><gene id>_predmap.png</strong> A left to right representation of the linear protein sequence with one pixel per amino acid. Each horizontal bar represents one structure prediction, colour coded by pLDDT. For small proteins, there will likely be one prediction of the full-length protein. For large proteins, there may be many overlapping predictions.</p> <p><strong><gene id>_<start aa>-<end_aa> </strong>A directory containing structure prediction of that gene ID between the start and end amino acid. Within this directory you will find:</p> <p><strong><gene id>_<start aa>-<end_aa>.json</strong> The full data in a JSON format, including linear protein sequence and metadata, along with 5 predicted protein structures ranked from best to worst overall pAE. For each predicted protein structure, the structure (PDB format), its pLDDT per residue and pairwise pAE.</p> <p><strong><gene id>_<start aa>-<end_aa>_1.pdb</strong> The PDB file of the highest ranked structure.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-pae.png</strong> A plot of pAE, for the highest ranked structure, at one pixel per amino acid. Shade of green represents pAE for that amino acid pair, see below.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-plddt.png</strong> A plot of pLDDT, for the highest ranked structure, at one horizontal pixel per amino acid. Bar height and colour both represent pLDDT for that amino acid, see below.</p> <p>PDB structure and pAE/pLDDT of lower ranked models are embedded in the JSON file.</p> <p>All pLDDT and pAE plots use the colour scales used by https://alphafold.ebi.ac.uk/: For pLDDT: > 90 (dark blue), 90 > pLDDT > 70 (light blue), 70 > pLDDT > 50 (orange), < 50 (yellow) discontinuous. For pAE: 0 (white angstrom) to dark green (32 angstrom) continuous.</p> <p>If you use this resource, please cite this Zenodo deposition and <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p>
Trypanosoma brucei predicted protein structures, part 1 of 2
<p>AlphaFold2-predicted protein structures for the <em>Trypanosoma brucei</em> (TREU927) proteome, predicted using input multiple sequence alignments optimised for the Discoba lineage in which <em>T. brucei </em>sits. The structure prediction methodology was exactly as described in <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p> <p>This deposition contains part 1 of 2. To get the full dataset, also download TbruceiTREU927_part2.zip from <a href="http://zenodo.org/record/7948119">10.5281/zenodo.7948119</a>.</p> <p>Data are organised with one directory per <em>T. brucei </em>TREU927 gene ID (eg. Tb927.1.3600). Within each directory you will find:</p> <p><strong><gene id>_predmap.png</strong> A left to right representation of the linear protein sequence with one pixel per amino acid. Each horizontal bar represents one structure prediction, colour coded by pLDDT. For small proteins, there will likely be one prediction of the full-length protein. For large proteins, there may be many overlapping predictions.</p> <p><strong><gene id>_<start aa>-<end_aa> </strong>A directory containing structure prediction of that gene ID between the start and end amino acid. Within this directory you will find:</p> <p><strong><gene id>_<start aa>-<end_aa>.json</strong> The full data in a JSON format, including linear protein sequence and metadata, along with 5 predicted protein structures ranked from best to worst overall pAE. For each predicted protein structure, the structure (PDB format), its pLDDT per residue and pairwise pAE.</p> <p><strong><gene id>_<start aa>-<end_aa>_1.pdb</strong> The PDB file of the highest ranked structure.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-pae.png</strong> A plot of pAE, for the highest ranked structure, at one pixel per amino acid. Shade of green represents pAE for that amino acid pair, see below.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-plddt.png</strong> A plot of pLDDT, for the highest ranked structure, at one horizontal pixel per amino acid. Bar height and colour both represent pLDDT for that amino acid, see below.</p> <p>PDB structure and pAE/pLDDT of lower ranked models are embedded in the JSON file.</p> <p>All pLDDT and pAE plots use the colour scales used by https://alphafold.ebi.ac.uk/: For pLDDT: > 90 (dark blue), 90 > pLDDT > 70 (light blue), 70 > pLDDT > 50 (orange), < 50 (yellow) discontinuous. For pAE: 0 (white angstrom) to dark green (32 angstrom) continuous.</p> <p>If you use this resource, please cite this Zenodo deposition and <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p>
Subset of non-redundant, high resolution multi-pass membrane protein PDB structure files from OPM
<p>A subset of OPM PDB structure files with membrane predictions curated for the MSc thesis project '<strong>Structural characterization of triplets of transmembrane α-helical segments consecutive in sequence in integral membrane proteins: a fragment-based computational pipeline'. </strong>The dataset contains PDB files downloaded from OPM on 22/02/2023. The files are alpha-helical integral membrane proteins with at least three transmembrane regions at better than 2.7 Angtrom resolution. They are non-redundant, with a sequence identity of less than 40% based on clustering with the cd-hit algorithm. </p>
Alphafold predicted structures of VPS13 proteins from model organisms
<p>This upload contains AlphaFold-predicted structures of VPS13 proteins from a variety of organisms. Given the large size of these proteins, only partial sequences were predicted with AlphaFold(1) and the resulting structures were aligned in PyMOL(2). A summary of the structures uploaded here is presented as a collection of domain cartoons in the "VPS13 domain organization across eukaryotic evolution.pdf" file. </p> <p>The structures were generated with AlphaFold v2.029 on the Yale High Performance Cluster. Each *.zip file contains the best ranked predictions (out of five) for each sequence (*.pdb files) and the PyMOL assembled full structure (*.pse file). In a few cases, where a good alignment was not possible due to long disordered regions in the C-terminal portions (mostly in proteins from <em>D. discoideum</em> and <em>A. thaliana</em>), the full structures were aligned manually in PyMOL based on the continuity of the lipid transfer groove. The structures in PyMOL can be colour-coded by the confidence value of AlphaFold predictions using the following prompt:</p> <p>set_color n0, [0.051, 0.341, 0.827]<br> set_color n1, [0.416, 0.796, 0.945]<br> set_color n2, [0.996, 0.851, 0.212]<br> set_color n3, [0.992, 0.490, 0.302]<br> color n0, b < 100; color n1, b < 90<br> color n2, b < 70; color n3, b < 50</p> <p>Considering that full length structures were assembled by aligning different protein fragments and in view of the presence of flexible loops with low prediction confidence scores, the relative positions of different folded domains are not necessarily correct.</p> <p> </p> <p><strong>References</strong></p> <p>1. J. Jumper, <em>et al.</em>, Highly accurate protein structure prediction with AlphaFold. <em>Nature</em> 596, 583–589 (2021).</p> <p>2. The PyMOL Molecular Graphics System, Version 2.0. Schrödinger LLC.</p>
The impact of genetically controlled splicing on exon inclusion and protein structure
<p><strong>This repository contains raw and processed files used in Einson et. al 2022. </strong></p> <p>Code used to generate these files can be found here: https://github.com/jeinson/sqtl_manuscript</p> <p><strong><em>Descriptions of files contained within each sub directory</em></strong></p> <p><strong>01_raw_psi</strong></p> <ul> <li><em>{GTEx_tissue_id}_v8.psi.tsv.gz: </em>Unfiltered PSI output from IPSA-nf, per tissue. See methods for details about how files were created. </li> <li><em>gtex_v8_exon_id_map.tsv: </em>Mapping file between exon coordinates and Ensembl gene IDs, with suffix used in GTEx v8 gencode annotation. </li> </ul> <p><strong>02_qtl_results</strong></p> <ul> <li><strong>cross_tissue</strong> <ul> <li><em>top_sQTLs_MAF05.tsv: </em>List of top GTEx v8 sQTLs across tissues, with one exon and top variant per tissue. See methods for details. See matching file for column descriptions. </li> <li><em>top_sQTLs_median_psi.tsv: </em>The median, mean, and standard deviation of PSI of each significant exon from the previous file, taken across all individuals from GTEx with data available.</li> <li><em>top_sQTLs_MAF05_w_anc_allele.tsv: </em>List of top sQTLs across tissues, with additional columns for the top ψQTL ancestral and derived alleles, where available. </li> </ul> </li> <li><strong>per_tissue</strong> <ul> <li><em>{GTEx_tissue_id}_combined_sQTLs.tsv.gz: </em>Raw output of ψQTL calling using QTLtools in grouped permutational mode per tissue, with groups specified by gene. See methods for more details, and https://qtltools.github.io/qtltools/ for column descriptions. </li> </ul> </li> </ul> <p><strong>03_qtl_credible_sets</strong></p> <ul> <li> <em>GTEx_psi_{GTEx_tissue_id}.collapsed.txt.gz: </em>Output of the QTL catalog fine mapping pipeline (https://github.com/eQTL-Catalogue/qtlmap), run on all exons and tissues, and collapsed using the procedure described in Methods. </li> </ul> <p><strong>04_qtl_coloc</strong></p> <ul> <li><em>combined_coloc_results_full.tsv.gz: </em>Combined output of running coloc on ψQTLs from the 18 GTEx tissues against 87 sets of GWAS summary statistics. This file contains all results, including non-significant associations. A nominal QTLtools pass was used as input. We do not include these files in this repository due to size limitations, but contact the authors if you need access to nominal QTL calls. </li> <li><em>top_sQTLs_with_top_coloc_event.tsv: </em>The QTLs in <em>top_sQTLs_MAF05.tsv</em> with additional columns for the GWAS with the highest posterior probability of a colocalization event. Importantly, the tissue and top variant may not match the main <em>top_sQTLs_MAF05.tsv </em>file for every gene. </li> </ul> <p><strong>05_exon_features: </strong>See matching files for description of each column. </p> <ul> <li><em>cross_tissue_constitutive_exons_with_AF.tsv: </em>Detailed features of cross tissue constitutive exons. See methods for definition of constitutive exons. </li> <li><em>cross_tissue_nonsignificant_genes_with_AF.tsv: </em>Detailed features of sufficiently variable exons with no significant variant across tissues. See methods for more details. </li> <li><em>top_sQTLs_MAF05_with_AF.tsv: </em>Detailed features of top sQTLs. </li> <li><em>top_sQTLs_with_top_coloc_with_AF.tsv: </em>Detailed features of sQTLs that colocalize with at least one GWAS trait. Contains columns for Euclidean distances between PAE matrices and RMSD between isoforms, among genes with a significant GWAS colocalization event. </li> </ul> <p><strong>06_predicted_structures: </strong>Each prediction was run 5 times, and we report the best model in the manuscript. </p> <ul> <li><strong>{protein.id}[_mutant].result</strong> <ul> <li><em>{protein.id}[_mutant]{_run.id}_coverage.png.gz: </em>Plot of the number of sequences per position in MSA</li> <li><em>{protein.id}[_mutant]{_run.id}_PAE.png.gz: </em>PAE matrix plots for each model</li> <li><em>{protein.id}[_mutant]{_run.id}_plddt.png.gz: </em>pLDDT plots for each model</li> <li><em>{protein.id}[_mutant]{_run.id}_predicted_aligned_error_v1.json.gz</em>: A PAE matrix for the best model using <a href="https://alphafold.ebi.ac.uk/faq#faq-7">AlphaFold-DB's format</a></li> <li><em>{protein.id}[_mutant]{_run.id}_unrelaxed_rank_{rank.num}_model_{model.num}_scores.json.gz</em>: Per model array (list of lists) with PAE, a list of the average pLDDT and the pTM score. </li> <li><em>{protein.id}[_mutant]{_run.id}_unrelaxed_rank_{rank.num}_model_{model.num}_pdb.gz: </em>Per model predicted structure in pd format</li> <li><em>{protein.id}[_mutant]{_run.id}.a3m.gz</em>: A3M formatted input MSA</li> <li><em>cite.bibtex: </em>BibTex file with citations for all used tools and databases</li> <li><em>config.json</em>: Model input parameters</li> </ul> </li> </ul> <p><strong>07_other_data</strong></p> <ul> <li><em>cross_tissue_constitutive_exons.tsv: </em>List of exons that are constitutively spliced across multiple tissues. See methods for details. </li> <li><em>cross_tissue_nonsignificant_genes.tsv</em>: List variably spliced exons with no significant sVariant in any tissue. See methods for details. </li> <li><em>gtex_v8_exon_id_map.rds: </em>rds representation of a map between exon IDs, as used in the modified version of gencode v26, and exon hg38 coordinates. </li> <li><em>gtex_v8_n_exons_per_gene.tsv: </em>Number of exons per gene, as annotated in the modified version of gencode v26 used in GTEx v8. </li> </ul> <p><strong>08_geuvadis</strong></p> <ul> <li><em>geuvadis_psi.tsv.gz: </em>Unfiltered PSI output from IPSA-nf, run on Geuvadis BAM files. See methods for details. (Raw data was downloaded from ftp://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/GEUV)</li> <li><em>geuvadis_sQTLs.tsv.gz: </em>Raw output of ψQTL calling using QTLtools in grouped permutational mode for geuvadis data, with groups specified by gene. See methods for more details, and https://qtltools.github.io/qtltools/ for column descriptions. </li> <li><em>remapped_gencode.v26.GRCh37.GTEx_v8.nochr.genes.gtf.gz:</em> Lifted over version of the gencode v26 gtf file, used to define exons for PSI and qtl mapping in the geuvadis analysis. The original version that was used in the GTEx analysis is based on GRCh38, and is available here: https://storage.googleapis.com/gtex_analysis_v8/reference/gencode.v26.GRCh38.genes.gtf</li> </ul>
Modeling flexible protein structure with AlphaFold2 and cross-linking mass spectrometry
<p>Ensembles of models predicted by AlphaFold for the proteins C3 (Complement component 3), luciferase and QBP (glutamine-binding periplasmic protein). Models interpolated between two conformations of C3, and luciferase are also included. This dataset is cited in the following paper: https://www.biorxiv.org/content/10.1101/2023.09.11.557128v1</p>
Dataset for Peptide binder design with inverse folding and protein structure prediction
<p>Dataset for a paper on peptide design</p> <p> </p> <p><br> mutated_peptides - results for randomly intriduced mutations in protein-peptide complexes that can be predicted at 2 Å (Figure 1)<br> pdb_peptide - variation in the number of recycles (1-10) for 96 peptides (Figure 1)<br> minibinder - results for the minibinder set (Figure 2)<br> Pfam - results for the Pfam set (Figures 4+5)<br> protein_mpnn - results on protein_mpnn test set (Figure 6)</p> <p> </p> <p> </p>
Dataset for Protein Structure Refinement using PCS Restrains of the dArmRP A4M4C
<p>The dataset contains all the experimental data and information used to refine the A4M4C dArmRP in the free and (KR)4 bound state from pseudocontact shifts. Individual modules are restraints during refinement.</p>
Files for training purposes - Protein structure manipulation training session @BIOI2
<p>AlphaFold2 prediction of ORF1 Nter dimer (using ColabFold v1.5)</p>
extHomFam v37.0: structural benchmark for protein multiple sequence alignments
<p>extHomFam v37.0 was constructed by combining Homstrad reference alignments (2 December 2023 release) with Pfam 37.0 (UniProt release) families containing at least 200 sequences. Homstrad entries with less than 3 reference sequences and those pointing to dead Pfam families were discarded.</p> <p> </p>
Data from: Cohesin protein Smc3 influences kinocilial structure and function
Open the record for dataset details and reuse information.
Structural ontogeny of protein-protein interactions
Open the record for dataset details and reuse information.
MD simulations data for: Role of the αC-β4 loop in protein kinase structure and dynamics
Open the record for dataset details and reuse information.
Conserved structural elements specialize ATAD1 as a membrane protein extraction machine
Open the record for dataset details and reuse information.
Supplementary data from: Assessing the structural boundaries of broadly reactive antibody interactions with diverse H3 influenza hemagglutinin proteins
Open the record for dataset details and reuse information.
DeeplyTough: Learning Structural Comparison of Protein Binding Sites (pre-processed datasets)
<p>This deposit includes the processed versions of TOUGH-M1 (Govindaraj and Brylinski, 2018), Vertex (Chen et al., 2016), and ProSPECCTs (Ehrt et al., 2018) datasets as used in our JCIM paper (Simonovsky and Meyers, 2020) - https://pubs.acs.org/doi/abs/10.1021/acs.jcim.9b00554. Specifically, this may include PDB files with associated pockets, pre-processed HTMD features, or UniProt accession number and cluster mappings. Note that these datasets have to be combined with the original dataset files released by the respective authors.</p> <p>Liability: We do not represent and/or warrant that no third party rights exist which might prevent the use of the database or that no third party rights would be infringed by said use.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.