Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

295

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

295 results for “structure prediction”

Learn how ShareScore rates datasets ↗
zenodo40/100

DPCstruct Classification of AlphaFold2-Predicted Protein Structures

<p>This dataset contains DPCstruct domain classifications for protein structures predicted by AlphaFold2, as presented in the paper "Unsupervised Domain Classification of AlphaFold2-Predicted Protein Structures."</p> <p>DPCstruct was applied to a non-redundant set of the AlphaFold Database v4.0, known as Foldseek Clusters, which includes approximately 15 million representative proteins, as described in the work by <a href="https://doi.org/10.1038/s41586-023-06510-w">Barrio-Hernandez et al.</a></p> <p>This repository provides the results of our classification, along with all the data related to the analyses presented in our study. DPCstruct algorithm can be found at <a href="https://github.com/RitAreaSciencePark/DPCstruct">https://github.com/RitAreaSciencePark/DPCstruct</a> together with examples on how to use it.</p> <p><strong>FILES DESCRIPTION:</strong></p> <ul> <li><strong>dpcstruct_classification.tsv: </strong>List of domains identified by DPCstruct and their corresponding metacluster. Columns: Metacluster ID, Protein Uniprot ID, domain start, domain end.</li> <li><strong>mcs_reps.fasta:</strong> For each metacluster, two representative domains were selected: one representing the center of the cluster and the other being the domain with the highest pLDDT score. If these are the same, only one domain is included as the representative. This file contains the list of representative domains and their sequences in FASTA format.</li> <li><strong><span>mcs_reps_pdbs.zip: </span></strong>Contains a PDB file for each representative domain. The filename is structured as 'proteinID_metacluster.pdb'.</li> <li><strong>mcs_properties.tsv:</strong> Set of properties per metacluster, including: <ul> <li><strong>mcID:</strong> Metacluster ID.</li> <li><strong>size:</strong> Number of domains.</li> <li><strong>len_aa:</strong> Average length of domains (number of amino acids).</li> <li><strong>len_std:</strong> Standard deviation of domain lengths.</li> <li><strong>len_ratio:</strong> Ratio of len_std to len_aa.</li> <li><strong>plddt:</strong> Average predicted LDDT as reported by AlphaFold2.</li> <li><strong>disorder:</strong> Average intrinsic disorder score calculated with AIUPred.</li> <li><strong>alntmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs.</li> <li><strong>tmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs, using the maximum between TM-score normalized by query or target.</li> <li><strong>lddt:</strong> Pairwise LDDT score, averaged over all pairs.</li> <li><strong>prob:</strong> Pairwise probability of homology according to SCOPe, as reported by Foldseek.</li> <li><strong>pident:</strong> Pairwise percentage identity, averaged over all pairs.</li> </ul> </li> <li><span><strong>annotated_[cath|scop]_qc[x]_t[x]_l[x].tsv:</strong>&nbsp;</span>For each fold in [CATH|SCOP], we provide the best matching DPCstruct domain, if available, along with the structural alignment information as reported by Foldseek. A fold is considered annotated if its alignment values meet or exceed the following thresholds: <ul> <li>qc: query coverage.</li> <li>t: template modelling score of the alignment.</li> <li>l: lddt score of the alignment.</li> </ul> </li> <li><strong>dpcstruct_consistency.tsv:</strong> Consistency of DPCstruct metaclusters with respect to Pfam 36.0 labels. Note that we consider a Pfam label to overlap with a DPCstruct domain even if it shares just one amino acid, which is why some metaclusters have many labels. In such cases, we only display 5 representative labels.</li> <li><strong>pfam_consistency.tsv:</strong> Consistency of Pfam Clans with respecto to DPCstruct labels.</li> </ul> <p><strong>Note:</strong> All 'tsv' files contain a header as the first row.</p> <p>If there is any doubt regarding the data or there is something missing please contact us:&nbsp;</p> <p>federico.barone@areasciencepark.it</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Development and Comparison of Model-Based and Data-Driven Approaches for the Prediction of the Mechanical Properties of Lattice Structures

<p>This dataset comes from the following paper:</p> <p>Chiara Pasini, Oscar Ramponi, Stefano Pandini, Luciana Sartore, Giulia Scalet, Development and Comparison of Model-Based and Data-Driven Approaches for the Prediction of the Mechanical Properties of Lattice Structures, J. of Materi Eng and Perform, 2024. <a href="https://doi.org/10.1007/s11665-024-10199-x">https://doi.org/10.1007/s11665-024-10199-x</a></p> <p>It contains:</p> <ul> <li>"Notes.pdf" describing all the files uploaded</li> <li>. m of the neural network</li> <li>. inp of the Abaqus finite element simulations</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Fig. 3 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 3. Scatterplot of canonical correspondence analysis (CCA) for the fish communities of the Jogui and Iguatemi Rivers.

opencc-by-4.0Mar 2007View details →
zenodo40/100

Fig. 2 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 2. Similarity dendrogram of fish communities in Jogui River (above) and Iguatemi Rivers (below).

opencc-by-4.0Mar 2007View details →
zenodo40/100

Fig. 4 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 4. Altitudinal distributions of the main fish species in the Jogui (A) and Iguatemi (B) rivers. Black dots represent sampling sites. Horizontal black lines represent species distribution range.

opencc-by-4.0Mar 2007View details →
zenodo40/100

Neural Networks for Structure-Informed Prediction of Formation Energy (employed in SIPFENN)

<p>pySIPFENN Documentation:&nbsp;<a href="https://pysipfenn.org">pysipfenn.org</a></p> <p>pySIPFENN GitHub:&nbsp;<a href="https://github.com/PhasesResearchLab/pySIPFENN">git.pysipfenn.org</a></p> <p>Original SIPFENN Paper:&nbsp;<a href="https://doi.org/10.1016/j.commatsci.2022.111254">10.1016/j.commatsci.2022.111254</a></p> <p>&nbsp;</p> <p>Network Changelog:</p> <p>V 0.10 - All models moved to the open ONNX format for improved interchangeability; NN30 neural network similar to NN20 but accepting the new KS2022 feature vector; Python code migrated to public GitHub repository.</p> <p>V 0.9 - Python code updated to the release version; paper published</p> <p>V 0.8 - Python code (beta)&nbsp;to run models included</p> <p>V 0.7 - Original upload of development models&nbsp;</p> <p>&nbsp;</p> <p>Selected works with SIPFENN alongside DFT and experiments:</p> <p>-&nbsp;<a href="https://doi.org/10.1016/j.actamat.2021.117448">10.1016/j.actamat.2021.117448</a></p> <p>-&nbsp;<a href="https://doi.org/10.1038/s41598-021-03578-0">10.1038/s41598-021-03578-0</a></p> <p>&nbsp;</p> <p>SIPFENN Abstract (original publication, 2021):</p> <p>In recent years, numerous studies have employed machine learning (ML) techniques to enable orders of magnitude faster high-throughput materials discovery by augmentation of existing methods or as standalone tools. In this paper, we introduce a new neural network-based tool for the prediction of formation energies based on elemental and structural features of Voronoi-tessellated materials. We provide a self-contained overview of the ML techniques used. Of particular importance is the connection between the ML and the true material-property relationship, how to improve the generalization accuracy by reducing overfitting, and how new data can be incorporated into the model to tune it to a specific material system.<br> &nbsp; &nbsp;&nbsp;<br> &nbsp; &nbsp; In the course of this work, over 30 novel neural network architectures were designed and tested. This lead to three final models optimized for (1) highest test accuracy on the Open Quantum Materials Database (OQMD), (2) performance in the discovery of new materials, and (3) performance at a low computational cost. On a test set of 21,800 compounds randomly selected from OQMD, they achieve mean average error (MAE) of 28, 40, and 42 meV/atom respectively. The second model provides better predictions on materials far from ones reported in OQMD, while the third reduces the computational cost by a factor of 8.<br> &nbsp; &nbsp;&nbsp;<br> &nbsp; &nbsp; We collect our results in a new open-source tool called SIPFENN (Structure-Informed Prediction of Formation Energy using Neural Networks). SIPFENN not only improves the accuracy beyond existing models but also ships in a ready-to-use form with pre-trained neural networks and a user interface.&nbsp;</p> <p>&nbsp;</p> <p>Contacts:</p> <p>- Adam Krajewski: ak@psu.edu</p> <p>- Prof. Zi-Kui Liu: zxl15@psu.edu</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Protein structure model predictions for secreted fungal proteins

<p><strong>Dataset A - Alphafold2 prediction output data for 753 secreted proteins of <em>Rhizophagus irregularis </em>DAOM197198</strong>.&nbsp;Gene IDs are taken from the annotation by Yildirir et al. 2021,&nbsp;<a href="https://doi.org/10.1111/nph.17842">doi.org/10.1111/nph.17842</a></p> <p><strong>Dataset B - Alphafold2 prediction output data for 10 fungal effectors.</strong><strong>&nbsp;</strong>These are nine effectors from&nbsp;<em>Fusarium oxysporum</em>&nbsp;f. sp.<em>&nbsp;lycopersici</em>&nbsp;and RiSLM from&nbsp;<em>Rhizophagus irregularis</em>&nbsp;as well as their amino acid sequences. Signal peptides and sequences preceding a predicted Kex2 processing site were removed.</p> <p><strong>Dataset C - Alphafold2 prediction output data for 454 matches of a MycFOLD-HMM search</strong>&nbsp;across the Mycocosm genome database (<a href="https://mycocosm.jgi.doe.gov/mycocosm/home">https://mycocosm.jgi.doe.gov/mycocosm/home</a>) and 36 Glomeromycotina fungal genomes.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Primary data for: "Remotely sensed localised primary production anomalies predict the burden and community structure of infection in long-term rodent datasets"

<p>Datasets</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15

<p>Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Automated benchmarking of combined protein structure and ligand conformation prediction

<p>The prediction of protein-ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein-Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three-dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time-split test-set is inappropriate for comprehensive PLC evaluation, with state-of-the-art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC-prediction-specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein-ligand interactions.</p> <p>This repository contains:</p> <ol> <li> <p>all_validation_clustering_data.tsv - X-ray validation data and MMSeqs cluster identifiers at different sequence identities for over a million small molecule and ion-binding pockets in the PDB.&nbsp;</p> </li> <li> <p>hqr_dataset.tsv - PDB IDs and ligand information for the high quality representative (HQR) dataset described in the manuscript</p> </li> <li> <p>score_files.tar.gz - Full docking results for all detected pockets for the PDBBind time-split test-set, the HQR dataset, and the subsets of AF models created for both datasets. One file per tool benchmarked with the following columns: Tool, Complex, Pocket, Rank, lDDT-PLI, lDDT-LP, BiSyRMSD, Reference_Ligand, Tool-generated Score</p> </li> <li> <p>errors_all_sets.csv - Report of failures running the pipeline with the following columns: Process, Complex/Ligand/Receptor, Problem</p> </li> </ol>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Data and analytical codes for: Learning beyond-pairwise interactions enables the bottom-up prediction of microbial community structure

<p>Data and analytical codes for: Ishizawa et al. (2023) Learning beyond-pairwise interactions enables the bottom-up prediction of microbial community structure, bioRxiv, 2023.07.04.546222</p> <p>&nbsp;https://www.biorxiv.org/content/10.1101/2023.07.04.546222v1</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
dryad40/100

Early insight into social network structure predicts climbing the social ladder

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad40/100

Data from: Habitat edge responses of generalist predators are predicted by prey and structural resources

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad40/100

Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad40/100

UltraScan Solution Modeler (US-SOMO) hydrodynamic parameter, structural small angle scattering and SESCA circular dichroism (CD) calculations on AlphaFold predicted structures

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad40/100

Fast-slow traits predict competition network structure and its response to resources and enemies

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Male song structure predicts offspring recruitment to the breeding population in a migratory bird

Open the record for dataset details and reuse information.

publicFeb 2024View details →
dryad40/100

Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad36/100

Data from: Chemical structure predicts the effect of plant-derived low-molecular weight compounds on soil microbiome structure and pathogen suppression

<p>1. Plant-derived low molecular weight compounds play a crucial role in shaping soil microbiome functionality. While various compounds have been demonstrated to affect soil microbes, most data are case-specific and do not provide generalizable predictions on their effects. Here we show that the chemical structural affiliation of low molecular weight compounds typically secreted by plant roots – sugars, amino acids, organic acids and phenolic acids – can predictably affect microbiome diversity, composition and functioning in terms of plant disease suppression.</p> <p>2. We amended soil with single or mixtures of representative compounds, mimicking carbon deposition by plants. We then assessed how different classes of compounds, or their combinations, affected microbiome composition and the protection of tomato plants from the soil-borne Ralstonia solanacearum bacterial pathogen.</p> <p>3. We found that chemical class predicted well the changes in microbiome composition and diversity. Organic and amino acids generally decreased the microbiome diversity compared to sugars and phenolic acids. These changes were also linked to disease incidence, with amino acids and nitrogen-containing compound mixtures inducing more severe disease symptoms connected with a reduction in bacterial community diversity.</p> <p>4. Together, our results demonstrate that low molecular weight compounds can predictably steer rhizosphere microbiome functioning providing guidelines to engineer microbiomes based on root exudation patterns by specific plant cultivars or crop regimes.</p>

opencc-zeroDec 2019View details →
dryad36/100

Data from: Long-term mechanistic hindcasts predict the structure of experimentally-warmed intertidal communities

Increases in global temperatures are expected to have dramatic effects on the abundance and distribution of species in the coming years. Intertidal organisms, which already experience temperatures at or beyond their thermal limits, provide a model system in which to investigate these effects. We took advantage of a previous study in which experimental plates were deployed in the intertidal zone and passively warmed for 12 years to a daily maximum temperature on average 2.7°C higher than control plots on the adjacent bedrock. We compared the composition of the biological communities on each experimental plate with its neighboring bedrock control. Plate communities showed decreased richness of taxa and percent cover of filamentous algae, mussels and mobile grazers relative to bedrock, and increased percent cover of biofilm. We then used short-term time-series measurements of plate and bedrock temperatures and a mechanistic heat-budget model to hindcast those temperatures back 12 years. Greater differences in long-term average temperature between the experimental plates and bedrock controls were correlated with lower similarity in community composition. Additionally, years with higher average differences between plate and bedrock temperatures were more predictive of current compositional similarity between plate and bedrock communities, even though they occurred farther in the past than did more recent, but cooler, years. We conclude that current intertidal communities reflect their long-term, rather than short-term, thermal histories. Mechanistic heat-budget models based on short-term measurements can provide this valuable, long-term information.

opencc-zeroAug 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record