Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

347

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

347 results for “Protein structures”

Learn how ShareScore rates datasets ↗
zenodo40/100

DProQ: A Gated-Graph Transformer for Protein Complex Structure Assessment

<p>The file contains two benchmark sets: Heterodimer-AF2 and Docking benchmark 5.5-AF2 test. Each set includes (1) doecy folder, (2) native folder, and (3) label_info.csv.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Benchmark Datasets for: EGR: Equivariant Graph Refinement and Assessment of 3D Protein Complex Structures

<p>This archive&nbsp;contains three benchmark datasets associated with the Equivariant Graph Refiner (EGR), two for protein complex structure refinement&nbsp;(PSR Test and Benchmark 2) and the other for protein complex structure assessment (M4S Test). The refinement datasets contain&nbsp;(1) a&nbsp;`pred` directory that contains decoy structure PDB files and (2) a `true` directory that contains native structure PDB files. The quality assessment dataset contains (1) `target_name` directories that each contain decoy structure PDB files for a given protein target and (2) a `label_info.csv` file listing each decoy structure&#39;s DockQ score and CAPRI class label.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

DProQ: A Gated-Graph Transformer for Protein Complex Structure Assessment: MAF2 set

<p>Multimer AF2&nbsp; set. For more information, please read our paper on biorxiv:</p> <p><a href="https://www.biorxiv.org/content/10.1101/2022.05.19.492741v2">https://www.biorxiv.org/content/10.1101/2022.05.19.492741v2</a>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Full Datasets for: EGR: Equivariant Graph Refinement and Assessment of 3D Protein Complex Structures

<p>This archive&nbsp;contains three&nbsp;datasets associated with the Equivariant Graph Refiner (EGR), two for protein complex structure refinement&nbsp;(PSR Test and Benchmark 2) and the other for protein complex structure assessment (M4S Test). The refinement datasets contain&nbsp;(1) a&nbsp;`pred` directory that contains decoy structure PDB files and (2) a `true` directory that contains native structure PDB files. The quality assessment dataset contains (1) `target_name` directories that each contain decoy structure PDB files for a given protein target and (2) a `label_info.csv` file listing each decoy structure&#39;s DockQ score and CAPRI class label.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Sequence-structure-function relationships in the microbial protein universe

<p>The Microbiome Immunity Project (MIP) dataset contains models predicted with both Rosetta and DMPFold (folder `dataset/`). It also contains DeepFRI function predictions for all models.&nbsp;</p> <p>The `metadata` folder contains additional data which may be useful for searching the MIP database (FASTA files, BLAST databases and useful scripts for structure/function search) as well as retrieving the sequence/structural annotations.</p> <p>The `intermediate_data` folder contains preprocessed output for reproducing many of the figures in our manuscript in conjunction with scripts and Juypter notebooks found in our git repository: https://github.com/microbiome-immunity-project/protein_universe .</p> <p>More information about the dataset and associated metadata is provided in the `README.md` file).</p> <p>We are also providing workflows to search the MIP database against a protein sequence or structure or function of interest (see `SEARCHING.md` for more details).</p>

opencc-by-4.0May 2022View details →
zenodo40/100

A structural database of chain-chain and domain-domain interfaces of proteins

<p>Library of protein-protein and domain-domain interfaces from the protein data bank. The data also contains the structural clusters of protein-protein and domain-domain interfaces.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Source codes and datasets for the paper "DRLComplex: Reconstruction of protein quaternary structures using deep reinforcement learning"

<p>This contains the<strong> reproducible&nbsp;source code and dataset </strong>for the paper &quot;DRLComplex&nbsp;: Reconstruction of protein quaternary structures using deep reinforcement learning paper&quot;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Profiling phage-host interactions between Skunavirus receptor binding proteins and lactococcal cell wall polysaccharide structures

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo40/100

[Accompanying Dataset for PHIStruct] ColabFold-Predicted Structures of Receptor-Binding Proteins

<p><strong>This dataset contains protein structures, computationally predicted via <a href="https://doi.org/10.1038/s41592-022-01488-1">ColabFold</a>, of 19,081 non-redundant (i.e., with duplicates removed) receptor-binding proteins from 8,525 phages across 238 host genera</strong>. We identified these receptor-binding proteins based on GenBank annotations. For phage sequences without GenBank annotations, we employed a pipeline that uses the viral protein library&nbsp;<a href="https://doi.org/10.1093/nargab/lqab067">PHROG</a> and the machine learning model <a href="https://doi.org/10.3390/v14061329">PhageRBPdetect</a>.&nbsp;</p> <p>More details can be found in our paper <strong>"PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings."</strong> The project page is <a href="https://github.com/bioinfodlsu/PHIStruct">https://github.com/bioinfodlsu/PHIStruct</a>. Our paper is published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btaf016" rel="nofollow">https://doi.org/10.1093/bioinformatics/btaf016</a></p> <p>Our research was supported with Cloud TPUs from&nbsp;<a href="https://sites.research.google/trc/about/" rel="nofollow">Google's TPU Research Cloud (TRC)</a>&nbsp;and with computing resources from the&nbsp;<a href="https://docs.mlerp.cloud.edu.au/" rel="nofollow">Machine Learning eResearch Platform (MLeRP)</a> of Monash University, University of Queensland, and Queensland Cyber Infrastructure Foundation Ltd.</p>

openmit-licenseMay 2024View details →
dryad40/100

FAPM: Functional annotation of proteins using multi-modal models beyond structural modeling

<p>Assigning accurate property labels to proteins, like functional terms and catalytic activity, is challenging, especially for proteins without homologs and "tail labels" with few known examples. Unlike previous methods that mainly focused on protein sequence features, we use a pretrained large natural language model to understand the semantic meaning of protein labels. Specifically, we introduce FAPM, a contrastive multi-modal model that links natural language with protein sequence language. This model combines a pretrained protein sequence model with a pretrained large language model to generate labels, such as Gene Ontology (GO) functional terms and catalytic activity predictions, in natural language. Our results show that FAPM excels in understanding protein properties, outperforming models based solely on protein sequences or structures. It achieves state-of-the-art performance on public benchmarks and in-house experimentally annotated phage proteins, which often have few known homologs. Additionally, FAPM's flexibility allows it to incorporate extra text prompts, like taxonomy information, enhancing both its predictive performance and explainability. This novel approach offers a promising alternative to current methods that rely on multiple sequence alignment for protein annotation.</p>

opencc-zeroJul 2024View details →
zenodo40/100

Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain

opencc-by-4.0Jul 2024View details →
zenodo40/100

Figure 3. Distribution of Q3 values-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm

<p>The estimated accuracy for the &alpha;- helices (QH), &beta;- strands (QE), C-coil states (QC), and three<br> state together (Q3) for the system is shown in Figure 3.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Figure 2. PAM250 matrix for the encoded sequence-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm

<p>The PAM matrix (Dayhoff et al., 1978) describes the probability that original amino acid<br> will be replaced by another amino acid over a defined evolutionary interval. The unit of<br> evolutionary divergence is defined as the interval in which 1% of the amino acids have been<br> changed between two sequences. The work uses PAM250, which assumes the occurrence of 250-<br> point mutations per 100 amino acids.<br> So, for the given the protein sequence GIVEQCCASVCSLYQLENYCN, A will be replaced<br> by 1 -3 0 1 -3 -1 0 5 -2 -3 -4 -2 -3 -5 0 1 0 -7 -5 -1 as shown in Figure 2.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Figure 1: Snapshot of the CB396 dataset-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning AlgorithmSecondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm

<p>The dataset used for this work is CB396. This dataset contains 396 non-redundant sequences<br> derived from the 3Dee database created by Cuff and Barton (Cuff &amp; Barton, 1999). It contains 396<br> proteins with their respective secondary structure as shown in Figure 1.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Data to accompany the paper "Improved fragment-based protein structure prediction by redesign of search heuristics"

<p>This repository contains the older and newer input fragment&nbsp;sets&nbsp;and other data used for the analyses in our paper. The filenames for each tarball contain the PDB identifier of each protein along with a chain ID if applicable, followed by &#39;old&#39; or &#39;new&#39; for old and new fragments, respectively. Each tarball contains: a .fasta file of the input sequence, a matching PDB structure file, the relevant PSIPRED secondary structure prediction file, and the 9mer and 3mer fragment files.&nbsp;<br> <br> An additional tarball, ScoreRMSDplots_3protocols.tgz, contains extended versions of Figure 3 which show score and RMSD distributions clearly. Additionally, the same data is shown for equivalent experiments using the older fragment&nbsp;set.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage - Fig. 4e

<p>The single-molecule FRET dataset underlying Fig. 4e of &quot;A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage&quot;, DOI: 10.1039/C8SC00681D</p>

opencc-by-nc-4.0Mar 2018View details →
zenodo40/100

A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage - Fig. 4f

<p>The single-molecule FRET dataset underlying Fig. 4f&nbsp;of &quot;A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage&quot;, DOI: 10.1039/C8SC00681D</p>

opencc-by-nc-4.0Mar 2018View details →
zenodo40/100

A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage - Fig. 4d

<p>The single-molecule FRET dataset underlying Fig. 4d of &quot;A bi-terminal protein ligation strategy to probe chromatin structure during DNA damage&quot;, DOI: 10.1039/C8SC00681D</p>

opencc-by-nc-4.0Mar 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record