Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
347
datasets available to search
ShareScore release 0.9.0
Dataset results
347 results for “Protein structures”
DProQ: A Gated-Graph Transformer for Protein Complex Structure Assessment
<p>The file contains two benchmark sets: Heterodimer-AF2 and Docking benchmark 5.5-AF2 test. Each set includes (1) doecy folder, (2) native folder, and (3) label_info.csv. </p>
Benchmark Datasets for: EGR: Equivariant Graph Refinement and Assessment of 3D Protein Complex Structures
<p>This archive contains three benchmark datasets associated with the Equivariant Graph Refiner (EGR), two for protein complex structure refinement (PSR Test and Benchmark 2) and the other for protein complex structure assessment (M4S Test). The refinement datasets contain (1) a `pred` directory that contains decoy structure PDB files and (2) a `true` directory that contains native structure PDB files. The quality assessment dataset contains (1) `target_name` directories that each contain decoy structure PDB files for a given protein target and (2) a `label_info.csv` file listing each decoy structure's DockQ score and CAPRI class label.</p>
DProQ: A Gated-Graph Transformer for Protein Complex Structure Assessment: MAF2 set
<p>Multimer AF2 set. For more information, please read our paper on biorxiv:</p> <p><a href="https://www.biorxiv.org/content/10.1101/2022.05.19.492741v2">https://www.biorxiv.org/content/10.1101/2022.05.19.492741v2</a> </p>
Full Datasets for: EGR: Equivariant Graph Refinement and Assessment of 3D Protein Complex Structures
<p>This archive contains three datasets associated with the Equivariant Graph Refiner (EGR), two for protein complex structure refinement (PSR Test and Benchmark 2) and the other for protein complex structure assessment (M4S Test). The refinement datasets contain (1) a `pred` directory that contains decoy structure PDB files and (2) a `true` directory that contains native structure PDB files. The quality assessment dataset contains (1) `target_name` directories that each contain decoy structure PDB files for a given protein target and (2) a `label_info.csv` file listing each decoy structure's DockQ score and CAPRI class label.</p>
Sequence-structure-function relationships in the microbial protein universe
<p>The Microbiome Immunity Project (MIP) dataset contains models predicted with both Rosetta and DMPFold (folder `dataset/`). It also contains DeepFRI function predictions for all models. </p> <p>The `metadata` folder contains additional data which may be useful for searching the MIP database (FASTA files, BLAST databases and useful scripts for structure/function search) as well as retrieving the sequence/structural annotations.</p> <p>The `intermediate_data` folder contains preprocessed output for reproducing many of the figures in our manuscript in conjunction with scripts and Juypter notebooks found in our git repository: https://github.com/microbiome-immunity-project/protein_universe .</p> <p>More information about the dataset and associated metadata is provided in the `README.md` file).</p> <p>We are also providing workflows to search the MIP database against a protein sequence or structure or function of interest (see `SEARCHING.md` for more details).</p>
A structural database of chain-chain and domain-domain interfaces of proteins
<p>Library of protein-protein and domain-domain interfaces from the protein data bank. The data also contains the structural clusters of protein-protein and domain-domain interfaces.</p>
Source codes and datasets for the paper "DRLComplex: Reconstruction of protein quaternary structures using deep reinforcement learning"
<p>This contains the<strong> reproducible source code and dataset </strong>for the paper "DRLComplex : Reconstruction of protein quaternary structures using deep reinforcement learning paper"</p>
Profiling phage-host interactions between Skunavirus receptor binding proteins and lactococcal cell wall polysaccharide structures
Open the record for dataset details and reuse information.
[Accompanying Dataset for PHIStruct] ColabFold-Predicted Structures of Receptor-Binding Proteins
<p><strong>This dataset contains protein structures, computationally predicted via <a href="https://doi.org/10.1038/s41592-022-01488-1">ColabFold</a>, of 19,081 non-redundant (i.e., with duplicates removed) receptor-binding proteins from 8,525 phages across 238 host genera</strong>. We identified these receptor-binding proteins based on GenBank annotations. For phage sequences without GenBank annotations, we employed a pipeline that uses the viral protein library <a href="https://doi.org/10.1093/nargab/lqab067">PHROG</a> and the machine learning model <a href="https://doi.org/10.3390/v14061329">PhageRBPdetect</a>. </p> <p>More details can be found in our paper <strong>"PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings."</strong> The project page is <a href="https://github.com/bioinfodlsu/PHIStruct">https://github.com/bioinfodlsu/PHIStruct</a>. Our paper is published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btaf016" rel="nofollow">https://doi.org/10.1093/bioinformatics/btaf016</a></p> <p>Our research was supported with Cloud TPUs from <a href="https://sites.research.google/trc/about/" rel="nofollow">Google's TPU Research Cloud (TRC)</a> and with computing resources from the <a href="https://docs.mlerp.cloud.edu.au/" rel="nofollow">Machine Learning eResearch Platform (MLeRP)</a> of Monash University, University of Queensland, and Queensland Cyber Infrastructure Foundation Ltd.</p>
FAPM: Functional annotation of proteins using multi-modal models beyond structural modeling
<p>Assigning accurate property labels to proteins, like functional terms and catalytic activity, is challenging, especially for proteins without homologs and "tail labels" with few known examples. Unlike previous methods that mainly focused on protein sequence features, we use a pretrained large natural language model to understand the semantic meaning of protein labels. Specifically, we introduce FAPM, a contrastive multi-modal model that links natural language with protein sequence language. This model combines a pretrained protein sequence model with a pretrained large language model to generate labels, such as Gene Ontology (GO) functional terms and catalytic activity predictions, in natural language. Our results show that FAPM excels in understanding protein properties, outperforming models based solely on protein sequences or structures. It achieves state-of-the-art performance on public benchmarks and in-house experimentally annotated phage proteins, which often have few known homologs. Additionally, FAPM's flexibility allows it to incorporate extra text prompts, like taxonomy information, enhancing both its predictive performance and explainability. This novel approach offers a promising alternative to current methods that rely on multiple sequence alignment for protein annotation.</p>
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain
Figure 3. Distribution of Q3 values-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The estimated accuracy for the α- helices (QH), β- strands (QE), C-coil states (QC), and three<br> state together (Q3) for the system is shown in Figure 3.</p>
Figure 2. PAM250 matrix for the encoded sequence-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The PAM matrix (Dayhoff et al., 1978) describes the probability that original amino acid<br> will be replaced by another amino acid over a defined evolutionary interval. The unit of<br> evolutionary divergence is defined as the interval in which 1% of the amino acids have been<br> changed between two sequences. The work uses PAM250, which assumes the occurrence of 250-<br> point mutations per 100 amino acids.<br> So, for the given the protein sequence GIVEQCCASVCSLYQLENYCN, A will be replaced<br> by 1 -3 0 1 -3 -1 0 5 -2 -3 -4 -2 -3 -5 0 1 0 -7 -5 -1 as shown in Figure 2.</p>
Figure 1: Snapshot of the CB396 dataset-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning AlgorithmSecondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The dataset used for this work is CB396. This dataset contains 396 non-redundant sequences<br> derived from the 3Dee database created by Cuff and Barton (Cuff & Barton, 1999). It contains 396<br> proteins with their respective secondary structure as shown in Figure 1.</p>
Data to accompany the paper "Improved fragment-based protein structure prediction by redesign of search heuristics"
<p>This repository contains the older and newer input fragment sets and other data used for the analyses in our paper. The filenames for each tarball contain the PDB identifier of each protein along with a chain ID if applicable, followed by 'old' or 'new' for old and new fragments, respectively. Each tarball contains: a .fasta file of the input sequence, a matching PDB structure file, the relevant PSIPRED secondary structure prediction file, and the 9mer and 3mer fragment files. <br> <br> An additional tarball, ScoreRMSDplots_3protocols.tgz, contains extended versions of Figure 3 which show score and RMSD distributions clearly. Additionally, the same data is shown for equivalent experiments using the older fragment set.</p>
A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage - Fig. 4e
<p>The single-molecule FRET dataset underlying Fig. 4e of "A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage", DOI: 10.1039/C8SC00681D</p>
A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage - Fig. 4f
<p>The single-molecule FRET dataset underlying Fig. 4f of "A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage", DOI: 10.1039/C8SC00681D</p>
A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage - Fig. 4d
<p>The single-molecule FRET dataset underlying Fig. 4d of "A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage", DOI: 10.1039/C8SC00681D</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.