Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
295
datasets available to search
ShareScore release 0.9.0
Dataset results
295 results for “structure prediction”
Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "
Open the record for dataset details and reuse information.
Fast-slow traits predict competition network structure and its response to resources and enemies
<p>Plants interact in complex networks but how network structure depends on resources, natural enemies, and species resource-use strategy remains poorly understood. Here, we quantified competition networks among 18 plants varying in fast-slow strategy, by testing how increased nutrient availability and reduced foliar pathogens affected intra- and inter-specific interactions. Our results show that nitrogen and pathogens altered several aspects of network structure, often in unexpected ways due to fast and slow-growing species responding differently. Nitrogen addition increased competition asymmetry in slow-growing networks, as expected, but decreased it in fast-growing networks. Pathogen reduction made networks more even and less skewed because pathogens targeted weaker competitors. Surprisingly, pathogens and nitrogen dampened each other's effect. Our results show that plant growth strategy is key to understanding how competition responds to resources and enemies, a prediction from classic theories that has rarely been tested by linking functional traits to competition networks.</p>
Research data for "Predicting Dynamics from Structure in a Sodium Silicate Glass"
<p>This dataset supports the paper "Predicting Dynamics from Structure in a Sodium Silicate Glass".</p> <p>The following files are provided.</p> <p>File: dataset_800.zip</p> <p>- Pickle files for:</p> <ul> <li>400 Sodium silicate glass structures of 3000 atoms</li> <li>30 Trajectories sampled eight times at various timescales up to 1 ns for each if the 400 glass structures</li> </ul> <p>File: in.comb</p> <p>- Lammps inputfile used to generate simulation from with the data in dataset was sampled</p>
Predicting glycan structure from tandem mass spectrometry via deep learning
<p>Curated set of LC-MS/MS data from glycomics studies. Used for training and applying CandyCrunch, a deep learning model to predict glycan structure from LC-MS/MS data, described in Urban et al., Nat Methods, 2024 and https://github.com/BojarLab/CandyCrunch.</p> <p>Files:</p> <p>full_dataset.xlsx: Full dataset with all annotated LC-MS/MS glycan spectra</p> <p>X_train.pkl: spectra and metadata from our training set</p> <p>y_train.pkl: labels from our training set</p> <p>X_test.pkl: spectra and metadata from our independent test set</p> <p>y_test.pkl: labels from our independent test set</p> <p>glycans.pkl: glycans in IUPAC-condensed nomenclature in the same order as the label-encoding</p>
Progress Toward SHAPE Constrained Computational Prediction of Tertiary Interactions in RNA Structure
<p>Supplementary repository for the "Progress Toward SHAPE Constrained Computational Prediction of Tertiary Interactions in RNA Structure" article. Contains the simulation on the <em>Didymium iridis</em> lariat-capping ribozyme (DiLCrz, PDB ID: 4P8Z).</p>
Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions
<p>Using AlphaFold2 and HADDOCK, we have generated a predicted structure for the SARS-CoV-2 B.1.1.529 variant's Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in the receptor binding domain of B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>
Mining folded proteomes in the era of accurate structure prediction
<p>Supplementary data to accompany the manuscript “Mining folded proteomes in the era of accurate structure prediction”. Contains three zip files with fold matching search results to support results in the main text.</p>
[Accompanying Dataset for PHIStruct] ColabFold-Predicted Structures of Receptor-Binding Proteins
<p><strong>This dataset contains protein structures, computationally predicted via <a href="https://doi.org/10.1038/s41592-022-01488-1">ColabFold</a>, of 19,081 non-redundant (i.e., with duplicates removed) receptor-binding proteins from 8,525 phages across 238 host genera</strong>. We identified these receptor-binding proteins based on GenBank annotations. For phage sequences without GenBank annotations, we employed a pipeline that uses the viral protein library <a href="https://doi.org/10.1093/nargab/lqab067">PHROG</a> and the machine learning model <a href="https://doi.org/10.3390/v14061329">PhageRBPdetect</a>. </p> <p>More details can be found in our paper <strong>"PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings."</strong> The project page is <a href="https://github.com/bioinfodlsu/PHIStruct">https://github.com/bioinfodlsu/PHIStruct</a>. Our paper is published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btaf016" rel="nofollow">https://doi.org/10.1093/bioinformatics/btaf016</a></p> <p>Our research was supported with Cloud TPUs from <a href="https://sites.research.google/trc/about/" rel="nofollow">Google's TPU Research Cloud (TRC)</a> and with computing resources from the <a href="https://docs.mlerp.cloud.edu.au/" rel="nofollow">Machine Learning eResearch Platform (MLeRP)</a> of Monash University, University of Queensland, and Queensland Cyber Infrastructure Foundation Ltd.</p>
Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure
<p>Predicting how tree populations will respond to climate change is an urgent societal concern. An increasingly popular way to make such predictions is the genomic offset (GO) approach, which aims to use genomic and climate data to identify populations that may experience climate maladaptation in the near future. More precisely, GO tries to represent the change in allele frequencies required to maintain the current gene-climate relationships under climate change. However, the GO approach has major limitations and, despite promising validation of its predictions using height data from common gardens, it still lacks broad empirical testing. In the present study, we evaluated the consistency and empirical validity of GO predictions in maritime pine (<em>Pinus pinaster</em> Ait.), a tree species from southwestern Europe and North Africa with a marked population genetic structure. First, gene-climate relationships were estimated using 9,817 SNPs genotyped in 454 trees from 34 populations; and candidate SNPs potentially involved in climate adaptation were identified. Second, GO was predicted using four methods, namely Gradient Forest (GF), Redundancy Analysis (RDA), latent factor mixed model (LFMM) and Generalised Dissimilarity Modeling (GDM), two sets of SNPs (candidate and control SNPs) and five climate general circulation models (GCMs) to account for uncertainty in future climate predictions. Last, the empirical validity of GO predictions was evaluated within a Bayesian framework by estimating the associations between GO predictions and two independent data sources: mortality data from National Forest Inventories (NFI), and mortality and height data from five common gardens in contrasting environments. We found high variability in GO predictions across methods, SNP sets and GCMs. Regarding validation, GO predictions with GDM and GF (and to a lesser extent RDA) based on the candidate SNPs showed the strongest and most consistent associations with mortality rates in common gardens and NFI plots. We found almost no association between GO predictions and tree height in common gardens, most likely due to the overwhelming effect of population genetic structure on tree height in this species. Our study demonstrates the imperative to validate GO predictions with a range of independent data sources before they can be used as informative and reliable metrics in conservation or management strategies.</p>
In situ conductometry for studying the homogenization of Al-Mg-Si alloys and predicting extrudate grain structure through machine learning
<p>This dataset includes the <em>in situ</em> impedance and time/temperature data from [1], grain structure data created by extrusion simulation coupled with physically-based microstructural simulation [2], and the predictions of the feed-forward neural network GRAINN-1/2 [1].</p> <p>[1] Österreicher, J. A., Zivanovic, D., Walenta, W., Maimone, S.,Hofbauer, M., Hovden, S., Tükör, Z., Arnoldt, A., Cerny, A. Kronsteiner, A., Antic, M., Zickler, G., Ehmeier, F., Mikulovic, M., Kunschert, G. (2024) . In situ conductometry for studying the homogenization of Al-Mg-Si alloys and predicting extrudate grain structure through machine learning. <em>Materials & Design</em>, 113070.</p> <p>[2] Hovden, S., Kronsteiner, J., Arnoldt, A., Horwatitsch, D., Kunschert, G., & Österreicher, J. A. (2024). Parameter study of extrusion simulation and grain structure prediction for 6xxx alloys with varied Fe content. <em>Materials Today Communications</em>, <em>38</em>, 108128.</p>
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain
Supplementary material for publication "Multi-Echelon Inventory Optimization in Supply Chain Networks: Exploring Network Structures and Predictive Modeling"
<div> <div> <div> <p>This dataset collects different supply chain network structures generated artificially. We present four types of networks: Serial, Convergent, Divergent, and General, each type consisting of 20,000 individual instances. All 80,000 network instances generated are available to researchers and practitioners in Excel. The repository consists of separate files for each network instance consisting of each network inventory data, node connections, and a visual representation.</p> </div> </div> </div>
Figure 3. Distribution of Q3 values-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The estimated accuracy for the α- helices (QH), β- strands (QE), C-coil states (QC), and three<br> state together (Q3) for the system is shown in Figure 3.</p>
Figure 2. PAM250 matrix for the encoded sequence-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The PAM matrix (Dayhoff et al., 1978) describes the probability that original amino acid<br> will be replaced by another amino acid over a defined evolutionary interval. The unit of<br> evolutionary divergence is defined as the interval in which 1% of the amino acids have been<br> changed between two sequences. The work uses PAM250, which assumes the occurrence of 250-<br> point mutations per 100 amino acids.<br> So, for the given the protein sequence GIVEQCCASVCSLYQLENYCN, A will be replaced<br> by 1 -3 0 1 -3 -1 0 5 -2 -3 -4 -2 -3 -5 0 1 0 -7 -5 -1 as shown in Figure 2.</p>
Figure 1: Snapshot of the CB396 dataset-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning AlgorithmSecondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The dataset used for this work is CB396. This dataset contains 396 non-redundant sequences<br> derived from the 3Dee database created by Cuff and Barton (Cuff & Barton, 1999). It contains 396<br> proteins with their respective secondary structure as shown in Figure 1.</p>
Data to accompany the paper "Improved fragment-based protein structure prediction by redesign of search heuristics"
<p>This repository contains the older and newer input fragment sets and other data used for the analyses in our paper. The filenames for each tarball contain the PDB identifier of each protein along with a chain ID if applicable, followed by 'old' or 'new' for old and new fragments, respectively. Each tarball contains: a .fasta file of the input sequence, a matching PDB structure file, the relevant PSIPRED secondary structure prediction file, and the 9mer and 3mer fragment files. <br> <br> An additional tarball, ScoreRMSDplots_3protocols.tgz, contains extended versions of Figure 3 which show score and RMSD distributions clearly. Additionally, the same data is shown for equivalent experiments using the older fragment set.</p>
AbDb processed and pickled for use in deep learning CDR-H3 Structure prediction
<p>This is a pickle file, ready for training by the neural network described in "Improving CDR-H3 modelling in Antibodies" found at the following URL:</p> <p><a href="https://github.com/OniDaito/MRes">https://github.com/OniDaito/MRes</a></p> <p>The data is derived from the AbDb dataset found at:</p> <p><a href="http://www.bioinf.org.uk/abs/abdb/">http://www.bioinf.org.uk/abs/abdb/</a></p>
Extended data for the paper "Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction"
<p>Extended data for the paper:<br> Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction</p> <p>Authors:<br> Shaun M Kandathil, Mario Garza-Fabre, Simon C Lovell and Julia Handl</p> <p>--------------------------------</p> <p>Contents of the zip file:</p> <p> </p> <p>Directory 'ECDFplots':<br> ----------------------<br> Data corresponding to Figure 3 for all targets, for the bilevel and ILS protocols. Data are available following stages 3 and 4 of the low-resolution protocol.</p> <p>Directory 'ScoreRMSDplots_3archivers':<br> --------------------------------------<br> Data corresponding to Figures 6 and 9 for all targets. Data corresponding to decoys obtained after low-resolution stages 3 and 4 can be found in subdirectories 'Stage3' and 'Stage4', respectively.<br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.