Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21
datasets available to search
ShareScore release 0.9.0
Dataset results
21 results for “conformal prediction”
Prediction of Conformational Variability for RRM proteins in inter3m data base
<p>Predictions for protein Conformational Variability for the entries in InteR3M (<a href="https://inter3mdb.loria.fr/">https://inter3mdb.loria.fr/</a>), performed with the software ConforMine (in preparation).</p>
Conformer datasets for "Equivariant Graph Neural Networks for Toxicity Prediction"
<p>Predictive modeling of toxicity is a crucial step in the drug discovery pipeline. It can help filter out molecules with a high probability of failing in the early stages of de novo drug design. Thus, several machine learning (ML) models have been developed to predict the toxicity of molecules by combining classical ML techniques or deep neural networks with well-known molecular representations such as fingerprints or 2D graphs. But the more natural, accurate representation of molecules is expected to be defined in physical 3D space like in ab initio methods. Recent studies successfully used equivariant graph neural networks (EGNNs) for representation learning based on 3D structures to predict quantum-mechanical properties of molecules. Inspired by this, we investigated the performance of EGNNs to construct reliable ML models for toxicity prediction. We used the equivariant transformer (ET) model in TorchMD-NET for this. Eleven toxicity data sets taken from MoleculeNet, TDCommons, and ToxBenchmark have been considered to evaluate the capability of ET for toxicity prediction. Our results show that ET adequately learns 3D representations of molecules that can successfully correlate with toxicity activity, achieving good accuracies on most data sets comparable to state-of-the-art models. We also test a physicochemical property, namely, the total energy of a molecule, to inform the toxicity prediction with a physical prior. However, our work suggests that these two properties can not be related. We also provide an attention weight analysis for helping to understand the toxicity prediction in 3D space and thus increase the explainability of the ML model. In summary, our findings offer promising insights considering 3D geometry information via EGNNs and provide a straightforward way to integrate molecular conformers into ML-based pipelines for predicting and investigating toxicity prediction in physical space. We expect that in the future, especially for larger, more diverse data sets, EGNNs will be an essential tool in this domain.</p> <p>PAPER</p> <p>https://pubs.acs.org/doi/full/10.1021/acs.chemrestox.3c00032</p> <p>CODE and MODELS:</p> <p>The conformer data sets and trained toxicity models will be published upon acceptance of this work. The code has been made available at <a href="https://github.com/jule-c/ET-Tox">https://github.com/jule-c/ET-Tox</a>, and the processed data as well as pretrained models for training and testing can be downloaded from <a href="../record/7942946">https://zenodo.org/record/7942946</a>. We can provide the full list of conformers as XYZ files upon request.</p>
Data - AlphaFold2 Predicts Alternative Conformation Populations in Green Fluorescent Protein Variants
<p><strong>MSAs.zip </strong>Multiple sequences alignments generated by AlphaFold2 structure prediction of 7 engineered GFPs.</p> <p><strong>AF2_models_column_masking.zip </strong>AlphaFold2 models of the alternative conformations of 7 engineered GFPs.</p> <p><strong>MD_trajectories_PyMOL.zip</strong> Molecular dynamics trajectories (PyMOL sessions) of the alternative conformations of 7 engineered GFPs.</p> <p><strong>MD_analysis.zip </strong>Root mean square deviation and per-residue root mean square fluctuations along molecular dynamics simulations of 7 engineered GFPs.</p> <p><strong>rmsd_values.zip</strong> Root mean square deviation relative to crystallographic GFP structure for AlphaFold2 models and molecular dynamics frames (global and central alpha-helix)</p>
Automated benchmarking of combined protein structure and ligand conformation prediction
<p>The prediction of protein-ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein-Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three-dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time-split test-set is inappropriate for comprehensive PLC evaluation, with state-of-the-art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC-prediction-specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein-ligand interactions.</p> <p>This repository contains:</p> <ol> <li> <p>all_validation_clustering_data.tsv - X-ray validation data and MMSeqs cluster identifiers at different sequence identities for over a million small molecule and ion-binding pockets in the PDB. </p> </li> <li> <p>hqr_dataset.tsv - PDB IDs and ligand information for the high quality representative (HQR) dataset described in the manuscript</p> </li> <li> <p>score_files.tar.gz - Full docking results for all detected pockets for the PDBBind time-split test-set, the HQR dataset, and the subsets of AF models created for both datasets. One file per tool benchmarked with the following columns: Tool, Complex, Pocket, Rank, lDDT-PLI, lDDT-LP, BiSyRMSD, Reference_Ligand, Tool-generated Score</p> </li> <li> <p>errors_all_sets.csv - Report of failures running the pipeline with the following columns: Process, Complex/Ligand/Receptor, Problem</p> </li> </ol>
Dataset for "ConfSolv: Prediction of solute conformer free energies across a range of solvents"
<p>This dataset contains three archives. The first archive, full_dataset.zip, contains geometries and free energies for nearly 44,000 solute molecules with almost 9 million conformers, in 42 different solvents. The geometries and gas phase free energies are computed using density functional theory (DFT). The solvation free energy for each conformer is computed using COSMO-RS and the solution free energies are computed using the sum of the gas phase free energies and the solvation free energies. The geometries for each solute conformer are provided as ASE_atoms_objects within a pandas DataFrame, found in the compressed file dft coords.pkl.gz within full_dataset.zip. The gas-phase energies, solvation free energies, and solution free energies are also provided as a pandas DataFrame in the compressed file free_energy.pkl.gz within full_dataset.zip. Ten example data splits for both random and scaffold split types are also provided in the ZIP archive for training models. Scaffold split index 0 is used to generate results in the corresponding publication. </p><p>The second archive, refined_conf_search.zip, contains geometries and free energies for a representative sample of 28 solute molecules from the full dataset that were subject to a refined conformer search and thus had more conformers located. The format of the data is identical to full_dataset.zip.</p><p>The third archive contains one folder for each solvent for which we have provided free energies in full_dataset.zip. Each folder contains the .cosmo file for every solvent conformer used in the COSMOtherm calculations, a dummy input file for the COSMOtherm calculations, and a CSV file that contains the electronic energy of each solvent conformer that needs to be substituted for "EH_Line" in the dummy input file.</p>
Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors - apo and validation MD
<p>Supplementary data of "Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors" paper.</p> <p>This dataset consists of molecular dynamics simulations trajectories and topology of Cruzain, Cathepsin K and Cathepsin L enzymes in it apo form, together with validation simulations. We ran five replicate 100ns simulations on each complex, with randomized initial velocities.</p>
Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors - ICL and IKR complexes MD
<p>Supplementary data of "Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors" paper.</p> <p>This dataset consists of the molecular dynamics simulations trajectory and topology of Cruzain, Cathepsin K and Cathepsin L enzymes in noncovalent and covalent complexes with selective ligand against Cruzain (ICL* and CCL) and Cathepsin K (IKR and CKR). We ran five replicate 100ns simulations on each complex, with randomized initial velocities.</p> <p> </p> <p> </p>
Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors - ICR and ICK complexes MD
<p>Supplementary data of "Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors" paper.</p> <p>This dataset consists of molecular dynamics simulations trajectory and topology of Cruzain, Cathepsin K and Cathepsin L enzymes in noncovalent and covalent complexes with selective ligand against Cruzain (ICR and CCR) and Cathepsin K (ICK and CCK). We ran five replicate 100ns simulations on each complex, with randomized initial velocities.</p>
Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors - Inputs and Analysis
<p>Supplementary data of "Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors" paper.</p> <p>This dataset consists of the parametrized ligand (covalent and noncovalent form of ICR, ICK, ICL, IKR) and complexes files, sample of input files used for Molecular dynamics simulations and analysis procedures, and the raw data of results. </p>
Data for: Redshift Prediction with Images for Cosmology using a Bayesian Convolutional Neural Network with Conformal Predictions
<p>These files contain the predictions from the CNN and BCNN model from the paper titled: "Redshift Prediction with Images for Cosmology using a Bayesian Convolutional Neural Network with Conformal Predictions" (Jones et al. 2024). These files will allow reproduction of the performance metrics described in the paper.</p> <p> </p> <p>full_prediction_set_CNN.csv - predictions for the redshift using the CNN model of the entire dataset<br>cnn_evaluation.csv - predictions from just the evaluation dataset that was not used in training</p> <p>Columns are:</p> <p>photoz - predicted photoz from the model<br>specz - spectroscopic redshift<br>objectid - object ID from HSC PDR2 data release (Aihara et al. 2019)</p> <p><br>full_prediction_set_BCNN.csv - predictions for the redshift using the BCNN model of the entire dataset<br>bcnn_evaluation.csv - predictions from just the evaluation dataset that was not used in training</p> <p>Columns are:</p> <p>photoz - predicted photoz from the model<br>specz - spectroscopic redshift<br>objectid - object ID from HSC PDR2 data release (Aihara et al. 2019)<br>photoz_uncertainty - uncertainty in the predicted photoz</p>
Validation of Time Series Technique for Prediction of Conformational States of Amino Acids
<p>Validation of Time Series Technique for Prediction of Conformational States of Amino Acids</p> <p>- a project for fulfillment of M.Sc (Master of Science) in Bioinformatics.</p>
Sequence/simulation data for Direct Prediction of Intrinsically Disordered Protein Conformational Properties From Sequence
<p>This is a DOI-linked deposition of sequence/biophysical properties pairs used in the associated paper by Lotthammer et al:</p><p>Lotthammer, J. M.<strong>*</strong>, Ginell, G. M.<strong>*</strong>, Griffith, D.<strong>*</strong>, Emenecker, R. J. & Holehouse, A. S. <br>Direct Prediction of Intrinsically Disordered Protein Conformational Properties From Sequence.<br><i><strong>Nature Methods</strong></i> (<i>in press</i>), (2023).</p><p> </p>
Datasets for understanding the importance of conformation in property prediction models
<p>Descriptor and conformer data sets for molecular property and reaction selectivity prediction tasks. The PQC data set was created based on a part of the PubChemQC PM6 dataset (J. Chem. Inf. Model. 2020, 60, 12, 5891–5899), which contains two- and three-dimensional descriptors and conformers. The APTC data sets are based on the data sets for asymmetric phase transfer catalysts with enantio-selectivity (<a href="https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets" target="_blank" rel="noopener">https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets</a>). The melting point data set was created from the Jean-Claude Bradley Double Plus Good (Highly Curated and Validated) Melting Points Dataset (<a href="https://doi.org/10.6084/m9.figshare.1031638.v1">https://doi.org/10.6084/m9.figshare.1031638.v1</a>).</p> <p>They contained descriptors and conformers to train and validate machine learning models.</p> <p>Detailed explanations on how to use these datasets are found in the Github repository: <a href="https://github.com/YuHamakawa/Conformation-Importance-ML-Models">https://github.com/YuHamakawa/Conformation-Importance-ML-Models</a>. </p> <p> </p> <p> </p>
Raw data for "Ancestral structure prediction reveals the conformational impact of the RuBisCO small subunit across time".
<p>The repository contains raw data for the article "Ancestral structure prediction reveals the conformational impact of the RuBisCO small subunit across time".</p> <p>The repository contains:</p> <ol> <li>Dataset for the RbcL and RbcS sequences along with the inferred sequences for the ancestors of interest</li> <li>Phylogenetic tree for the concatenated and separate RbcL-RbcS sequences.</li> <li>Structures for the extant and ancestral RuBisCO complexes used in the study.</li> <li>Solvated pdb files for creating the topology files required for MD-simulations.</li> </ol>
Data from: Patterns of male fitness conform to predictions of evolutionary models of late-life
We studied lifetime male virility, a male fitness component, in five populations of Drosophila melanogaster. Virility was measured as the number of females, out of eight total, that a male could fertilize in 24 hours. Individual males were measured at weekly intervals until they died. Virility declined in an approximately linear fashion for the first three weeks of adult life. It then stayed low but relatively constant for another three weeks, exhibiting a clear plateau. These observations are consistent with the evolutionary theories of late-life.The results were not consistent with a simple heterogeneity theory of late-life. This is the first demonstration of a late-life plateau for a male fitness component. We also found that the virility of males that were within seven days of death was significantly lower than similarly aged males that were not about to die. This rapid deterioration of virility prior to death, or death spiral, is similar to a decline in fecundity that we had previously documented.
Data from: Knowledge-based prediction of protein backbone conformation using a structural alphabet
Libraries of structural prototypes that abstract protein local structures are known as structural alphabets and have proven to be very useful in various aspects of protein structure analyses and predictions. One such library, Protein Blocks, is composed of 16 standard 5-residues long structural prototypes. This form of analyzing proteins involves drafting its structure as a string of Protein Blocks. Predicting the local structure of a protein in terms of protein blocks is the general objective of this work. A new approach, PB-kPRED is proposed towards this aim. It involves (i) organizing the structural knowledge in the form of a database of pentapeptide fragments extracted from all protein structures in the PDB and (ii) applying a knowledge-based algorithm that does not rely on any secondary structure predictions and/or sequence alignment profiles, to scan this database and predict most probable backbone conformations for the protein local structures. Though PB-kPRED uses the structural information from homologues in preference, if available. The predictions were evaluated rigorously on 15,544 query proteins representing a non-redundant subset of the PDB filtered at 30% sequence identity cut-off. We have shown that the kPRED method was able to achieve mean accuracies ranging from 40.8% to 66.3% depending on the availability of homologues. The impact of the different strategies for scanning the database on the prediction was evaluated and is discussed. Our results highlights the usefulness of the method in the context of proteins without any known structural homologues. A scoring function that gives a good estimate of the accuracy of prediction was further developed. This score estimates very well the accuracy of the algorithm (R2 of 0.82). An online version of the tool is provided freely for non-commercial usage at http://www.bo-protscience.fr/kpred/.
Atomistic Predictions and Network-Based Allosteric Analysis of Conformational Ensembles for the State-Switching ABL Kinase Mutants Using Combination of Alanine Sequence Scanning and Shallow Subsampling in AlphaFold2
Open the record for dataset details and reuse information.
Conformation Database for Publication: Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction
<p><strong>Conformation database</strong> for 2022 Publication "Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction"</p> <ul> <li>DOI of Physica A publication: <a href="https://doi.org/10.1016/j.physa.2022.128395">https://doi.org/10.1016/j.physa.2022.128395</a></li> <li>GitHub source code: <a href="https://github.com/CompSoftMatterBiophysics-CityU-HK/Applying-DRL-to-HP-Model-for-Protein-Structure-Prediction">https://github.com/CompSoftMatterBiophysics-CityU-HK/Applying-DRL-to-HP-Model-for-Protein-Structure-Prediction</a></li> </ul> <p>This conformation database shows the distinct conformations of best-known and next best energies:</p> <p>├── <strong>20merA</strong><br> │ ├── <strong>20merA_E8_set</strong><br> │ ├── <strong>20merA_E9_set</strong><br> │ ├── confs_20merA_E8.txt<br> │ └── confs_20merA_E9.txt<br> ├── <strong>20merB</strong><br> │ ├── <strong>20merB_E10_set</strong><br> │ ├── <strong>20merB_E9_set</strong><br> │ ├── confs_20merB_E10.txt<br> │ └── confs_20merB_E9.txt<br> ├── <strong>24mer</strong><br> │ ├── <strong>24mer_E8_set</strong><br> │ ├── <strong>24mer_E9_set</strong><br> │ ├── confs_24mer_E8.txt<br> │ └── confs_24mer_E9.txt<br> ├── <strong>25mer</strong><br> │ ├── <strong>25mer_E7_set</strong><br> │ ├── <strong>25mer_E8_set</strong><br> │ ├── confs_25mer_E7.txt<br> │ └── confs_25mer_E8.txt<br> ├── <strong>36mer</strong><br> │ ├── <strong>36mer_E13_set</strong><br> │ ├── <strong>36mer_E14_set</strong><br> │ ├── confs_36mer_E13.txt<br> │ └── confs_36mer_E14.txt<br> ├── <strong>48mer</strong><br> │ ├── <strong>48mer_E22_set</strong><br> │ ├── <strong>48mer_E23_set</strong><br> │ ├── confs_48mer_E22.txt<br> │ └── confs_48mer_E23.txt<br> └── <strong>50mer</strong><br> ├── <strong>50mer_E20_set</strong><br> ├── <strong>50mer_E21_set</strong><br> ├── confs_50mer_E20.txt<br> └── confs_50mer_E21.txt</p>
Data from: Knowledge-based prediction of protein backbone conformation using a structural alphabet
Open the record for dataset details and reuse information.
Data from: Patterns of male fitness conform to predictions of evolutionary models of late-life
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.