Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

295

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

295 results for “structure prediction”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
dryad40/100

Fast-slow traits predict competition network structure and its response to resources and enemies

<p>Plants interact in complex networks but how network structure depends on resources, natural enemies, and species resource-use strategy remains poorly understood. Here, we quantified competition networks among 18 plants varying in fast-slow strategy, by testing how increased nutrient availability and reduced foliar pathogens affected intra- and inter-specific interactions. Our results show that nitrogen and pathogens altered several aspects of network structure, often in unexpected ways due to fast and slow-growing species responding differently. Nitrogen addition increased competition asymmetry in slow-growing networks, as expected, but decreased it in fast-growing networks. Pathogen reduction made networks more even and less skewed because pathogens targeted weaker competitors. Surprisingly, pathogens and nitrogen dampened each other's effect. Our results show that plant growth strategy is key to understanding how competition responds to resources and enemies, a prediction from classic theories that has rarely been tested by linking functional traits to competition networks.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Research data for "Predicting Dynamics from Structure in a Sodium Silicate Glass"

<p>This dataset supports the paper "Predicting Dynamics from Structure in a Sodium Silicate Glass".</p> <p>The following files are provided.</p> <p>File: dataset_800.zip</p> <p>- Pickle files for:</p> <ul> <li>400 Sodium silicate glass structures of 3000 atoms</li> <li>30 Trajectories sampled eight times at various timescales up to 1 ns for each if the 400 glass structures</li> </ul> <p>File: in.comb</p> <p>- Lammps inputfile used to generate simulation from with the data in dataset was sampled</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Predicting glycan structure from tandem mass spectrometry via deep learning

<p>Curated set of LC-MS/MS data from glycomics studies. Used for training and applying CandyCrunch, a deep learning model to predict glycan structure from LC-MS/MS data, described in Urban et al., Nat Methods, 2024 and https://github.com/BojarLab/CandyCrunch.</p> <p>Files:</p> <p>full_dataset.xlsx: Full dataset with all annotated LC-MS/MS glycan spectra</p> <p>X_train.pkl: spectra and metadata from our training set</p> <p>y_train.pkl: labels from our training set</p> <p>X_test.pkl: spectra and metadata from our independent test set</p> <p>y_test.pkl: labels from our independent test set</p> <p>glycans.pkl: glycans in IUPAC-condensed nomenclature in the same order as the label-encoding</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Progress Toward SHAPE Constrained Computational Prediction of Tertiary Interactions in RNA Structure

<p>Supplementary&nbsp;repository for the &quot;Progress Toward SHAPE Constrained Computational &nbsp;Prediction of Tertiary Interactions in RNA Structure&quot; article.&nbsp;Contains the simulation on&nbsp;the&nbsp;<em>Didymium iridis</em>&nbsp;lariat-capping ribozyme (DiLCrz, PDB ID: 4P8Z).</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions

<p>Using AlphaFold2 and HADDOCK, we have generated a predicted&nbsp;structure for the SARS-CoV-2 B.1.1.529 variant&#39;s Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in&nbsp;the receptor binding domain&nbsp;of&nbsp;B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Mining folded proteomes in the era of accurate structure prediction

<p>Supplementary data to accompany the manuscript &ldquo;Mining folded proteomes in the era of accurate structure prediction&rdquo;. Contains three zip files with fold matching search results to support results in the main text.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

[Accompanying Dataset for PHIStruct] ColabFold-Predicted Structures of Receptor-Binding Proteins

<p><strong>This dataset contains protein structures, computationally predicted via <a href="https://doi.org/10.1038/s41592-022-01488-1">ColabFold</a>, of 19,081 non-redundant (i.e., with duplicates removed) receptor-binding proteins from 8,525 phages across 238 host genera</strong>. We identified these receptor-binding proteins based on GenBank annotations. For phage sequences without GenBank annotations, we employed a pipeline that uses the viral protein library&nbsp;<a href="https://doi.org/10.1093/nargab/lqab067">PHROG</a> and the machine learning model <a href="https://doi.org/10.3390/v14061329">PhageRBPdetect</a>.&nbsp;</p> <p>More details can be found in our paper <strong>"PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings."</strong> The project page is <a href="https://github.com/bioinfodlsu/PHIStruct">https://github.com/bioinfodlsu/PHIStruct</a>. Our paper is published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btaf016" rel="nofollow">https://doi.org/10.1093/bioinformatics/btaf016</a></p> <p>Our research was supported with Cloud TPUs from&nbsp;<a href="https://sites.research.google/trc/about/" rel="nofollow">Google's TPU Research Cloud (TRC)</a>&nbsp;and with computing resources from the&nbsp;<a href="https://docs.mlerp.cloud.edu.au/" rel="nofollow">Machine Learning eResearch Platform (MLeRP)</a> of Monash University, University of Queensland, and Queensland Cyber Infrastructure Foundation Ltd.</p>

openmit-licenseMay 2024View details →
dryad40/100

Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure

<p>Predicting how tree populations will respond to climate change is an urgent societal concern. An increasingly popular way to make such predictions is the genomic offset (GO) approach, which aims to use genomic and climate data to identify populations that may experience climate maladaptation in the near future. More precisely, GO tries to represent the change in allele frequencies required to maintain the current gene-climate relationships under climate change. However, the GO approach has major limitations and, despite promising validation of its predictions using height data from common gardens, it still lacks broad empirical testing. In the present study, we evaluated the consistency and empirical validity of GO predictions in maritime pine (<em>Pinus pinaster</em> Ait.), a tree species from southwestern Europe and North Africa with a marked population genetic structure. First, gene-climate relationships were estimated using 9,817 SNPs genotyped in 454 trees from 34 populations; and candidate SNPs potentially involved in climate adaptation were identified. Second, GO was predicted using four methods, namely Gradient Forest (GF), Redundancy Analysis (RDA), latent factor mixed model (LFMM) and Generalised Dissimilarity Modeling (GDM), two sets of SNPs (candidate and control SNPs) and five climate general circulation models (GCMs) to account for uncertainty in future climate predictions. Last, the empirical validity of GO predictions was evaluated within a Bayesian framework by estimating the associations between GO predictions and two independent data sources: mortality data from National Forest Inventories (NFI), and mortality and height data from five common gardens in contrasting environments. We found high variability in GO predictions across methods, SNP sets and GCMs. Regarding validation, GO predictions with GDM and GF (and to a lesser extent RDA) based on the candidate SNPs showed the strongest and most consistent associations with mortality rates in common gardens and NFI plots. We found almost no association between GO predictions and tree height in common gardens, most likely due to the overwhelming effect of population genetic structure on tree height in this species. Our study demonstrates the imperative to validate GO predictions with a range of independent data sources before they can be used as informative and reliable metrics in conservation or management strategies.</p>

opencc-zeroMay 2024View details →
zenodo40/100

In situ conductometry for studying the homogenization of Al-Mg-Si alloys and predicting extrudate grain structure through machine learning

<p>This dataset includes the <em>in situ</em> impedance and time/temperature data from [1], grain structure data created by extrusion simulation coupled with physically-based microstructural simulation [2], and the predictions of the feed-forward neural network GRAINN-1/2 [1].</p> <p>[1] &Ouml;sterreicher, J. A., Zivanovic, D., Walenta, W., Maimone, S.,Hofbauer, M., Hovden, S., T&uuml;k&ouml;r, Z., Arnoldt, A., Cerny, A. Kronsteiner, A., Antic, M., Zickler, G., Ehmeier, F., Mikulovic, M., Kunschert, G. (2024) . In situ conductometry for studying the homogenization of Al-Mg-Si alloys and predicting extrudate grain structure through machine learning. <em>Materials &amp; Design</em>, 113070.</p> <p>[2] Hovden, S., Kronsteiner, J., Arnoldt, A., Horwatitsch, D., Kunschert, G., &amp; &Ouml;sterreicher, J. A. (2024). Parameter study of extrusion simulation and grain structure prediction for 6xxx alloys with varied Fe content. <em>Materials Today Communications</em>, <em>38</em>, 108128.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain

opencc-by-4.0Jul 2024View details →
zenodo40/100

Supplementary material for publication "Multi-Echelon Inventory Optimization in Supply Chain Networks: Exploring Network Structures and Predictive Modeling"

<div> <div> <div> <p>This dataset collects different supply chain network structures generated artificially. We present four types of networks: Serial, Convergent, Divergent, and General, each type consisting of 20,000 individual instances. All 80,000 network instances generated are available to researchers and practitioners in Excel. The repository consists of separate files for each network instance consisting of each network inventory data, node connections, and a visual representation.</p> </div> </div> </div>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Figure 3. Distribution of Q3 values-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm

<p>The estimated accuracy for the &alpha;- helices (QH), &beta;- strands (QE), C-coil states (QC), and three<br> state together (Q3) for the system is shown in Figure 3.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Figure 2. PAM250 matrix for the encoded sequence-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm

<p>The PAM matrix (Dayhoff et al., 1978) describes the probability that original amino acid<br> will be replaced by another amino acid over a defined evolutionary interval. The unit of<br> evolutionary divergence is defined as the interval in which 1% of the amino acids have been<br> changed between two sequences. The work uses PAM250, which assumes the occurrence of 250-<br> point mutations per 100 amino acids.<br> So, for the given the protein sequence GIVEQCCASVCSLYQLENYCN, A will be replaced<br> by 1 -3 0 1 -3 -1 0 5 -2 -3 -4 -2 -3 -5 0 1 0 -7 -5 -1 as shown in Figure 2.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Figure 1: Snapshot of the CB396 dataset-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning AlgorithmSecondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm

<p>The dataset used for this work is CB396. This dataset contains 396 non-redundant sequences<br> derived from the 3Dee database created by Cuff and Barton (Cuff &amp; Barton, 1999). It contains 396<br> proteins with their respective secondary structure as shown in Figure 1.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Data to accompany the paper "Improved fragment-based protein structure prediction by redesign of search heuristics"

<p>This repository contains the older and newer input fragment&nbsp;sets&nbsp;and other data used for the analyses in our paper. The filenames for each tarball contain the PDB identifier of each protein along with a chain ID if applicable, followed by &#39;old&#39; or &#39;new&#39; for old and new fragments, respectively. Each tarball contains: a .fasta file of the input sequence, a matching PDB structure file, the relevant PSIPRED secondary structure prediction file, and the 9mer and 3mer fragment files.&nbsp;<br> <br> An additional tarball, ScoreRMSDplots_3protocols.tgz, contains extended versions of Figure 3 which show score and RMSD distributions clearly. Additionally, the same data is shown for equivalent experiments using the older fragment&nbsp;set.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

AbDb processed and pickled for use in deep learning CDR-H3 Structure prediction

<p>This is a pickle file, ready for training by the neural network described in &quot;Improving CDR-H3 modelling in Antibodies&quot; found at the following URL:</p> <p><a href="https://github.com/OniDaito/MRes">https://github.com/OniDaito/MRes</a></p> <p>The data is derived from the AbDb dataset found at:</p> <p><a href="http://www.bioinf.org.uk/abs/abdb/">http://www.bioinf.org.uk/abs/abdb/</a></p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Extended data for the paper "Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction"

<p>Extended data for the paper:<br> Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction</p> <p>Authors:<br> Shaun M Kandathil, Mario Garza-Fabre, Simon C Lovell and Julia Handl</p> <p>--------------------------------</p> <p>Contents of the zip file:</p> <p>&nbsp;</p> <p>Directory &#39;ECDFplots&#39;:<br> ----------------------<br> &nbsp;&nbsp; &nbsp;Data corresponding to Figure 3 for all targets, for the bilevel and ILS protocols. Data are available following stages 3 and 4 of the low-resolution protocol.</p> <p>Directory &#39;ScoreRMSDplots_3archivers&#39;:<br> --------------------------------------<br> &nbsp;&nbsp; &nbsp;Data corresponding to Figures 6 and 9 for all targets. Data corresponding to decoys obtained after low-resolution stages 3 and 4 can be found in subdirectories &#39;Stage3&#39; and &#39;Stage4&#39;, respectively.<br> &nbsp;</p>

opencc-by-4.0Jul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record