Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30
datasets available to search
ShareScore release 0.9.0
Dataset results
30 results for “in silico prediction”
Dataset related to article "New in silico models to predict in vitro micronucleus induction as marker of genotoxicity"
<p>The .txt file contains the dataset of the in silico model for genotoxicity as induction of micronuclei.</p> <p>The .doc file contains the descriptors of the models and the structural alerts.</p>
In silico prediction of ARB resistance: A first step in creating personalized ARB therapy
<p><strong>AT1R Model preparation</strong><br> The crystal structure of human AT1R bound to olmesartan (PDB: 4ZUD) was downloaded from the RCSB Protein Data Bank. 4ZUD contains apocytochrome b562RIL fused to the amino terminus, and many of the flexible regions, as well as helix 8, are not resolved. In order to generate an appropriate starting structure, olmesartan and the apocytochrome b562RIL fusion were removed from 4ZUD, and the missing regions were added to the protein with MOE software (Chemical Computing Group ULC, Montreal, Canada). Specifically, the N-Terminus (residues 1 to 25), intracellular loop 2 (residues 134 to 140), extracellular loop 2 (residues 186 to 188), intracellular loop 3 (residues 223 to 234), and helix 8 (residues 305 to 316) were added to the AT1R in accordance to the human AT1R sequence and PDB:4YAY. The remaining carboxyl-tail of the AT1R (residues 317 to 359) was not modeled. The AT1R model then underwent an energy minimization within MOE using the Amber10:Extended Huckel Theory (EHT) force field.</p> <p><strong>Molecular dynamic (MD) simulations and analysis</strong><br> The MOE minimized AT1R was loaded into CHARMM-GUI. An 80 Å by 80 Å lipid bi-layer composed of 13% cholesterol and 87% Phosphatidylcholine (POPC) was generated around the receptor. Water was packed 17.5 Å above and below the lipid bi-layer, and 150 mM Na+ and Cl- ions were added to the system via Monte-Carlo ion placing. The all-atom CHARMM C36 force field for proteins and ions, and the CHARMM TIP3P force field for water were selected. A hard non-bonded cutoff of 8.0 angstroms was utilized. All molecular dynamics simulations were performed using the PMEMD module of the AMBER16 package with support for MPI multi-process control and GPU acceleration code. Orthorhombic periodic boundary conditions with a constant pressure of 1 atm was set via the NPT ensemble and temperature was set to 310.15°K (37°C) using Langevin dynamics. The SHAKE algorithm was used to constrain bonds containing hydrogens. The dynamics were propagated using Langevin dynamics with Langevin damping coefficient of 1 ps-1 and a time step of 2 fs. Before the production run, the AT1R model was minimized for 5000 steps using the steepest descent method and then equilibrated for 600 ps. The protein coordinates were saved in 10 ps intervals. The production run lasted 150 ns, at which point all three replicas were stable for at least the last 20 ns.</p>
Overview of available toxicity data for calystegines - results of the in silico genotoxicity predictions
<p>Results of the<em> in silico</em> genotoxicity predictions complementing the EFSA scientific report on calystegines: https://doi.org/10.2903/j.efsa.2019.5574</p>
In silico design, docking simulation, and ANN-QSAR model for predicting the anticoagulant activity of thiourea isosteviol compounds as FXa inhibitors
<p>The present work combined molecular modeling and docking approach for searching and designing novel thiourea isosteviol-based compounds as potential FXa inhibitors. Elaborated regression model establishes the relationships between experimentally determined anticoagulant activity and molecular descriptors and enables the prediction of FXa inhibitory activity for novel compounds. The obtained results proved that the Artificial Neural Network algorithm facilitates the search for the most promising isosteviol derivatives incorporating thiourea fragments as FXa inhibitors. Moreover, docking simulation confirms the prominent binding of the newly in silico designed molecules with the active sites of the protein, which may be the lead molecules and can be further optimized for the efficient pharmacodynamic and pharmacokinetic profiles. The enclosed files are representations of molecular structures of thiourea isosteviol compounds with experimentally tested FXa inhibitory activity (i20-i39) geometrically optimized in hyperchem, newly in silico designed thiourea isosteviol compounds geometrically optimized in hyperchem (e1-e11), one file contains molecular descriptors for optimized structures calculated in Dragon and there is also a code for ANN QSAR model for predicting activity of novel thiourea isosteviol compounds. </p>
In silico subcellular targeting predictions for cytosolic aminoacyl tRNA-synthetases (aaRS) in parasitic plants
<p>Eukaryotic nuclear genomes often encode distinct sets of protein translation machinery for function in the cytosol vs. organelles (mitochondria and plastids). This phenomenon raises questions about why multiple translation systems are maintained even though they are capable of comparable functions, and whether they evolve differently depending on the compartment where they operate. These questions are particularly interesting in land plants because translation machinery, including aminoacyl-tRNA synthetases (aaRS), is often dual-targeted to both the plastids and mitochondria. These two organelles have quite different metabolisms, with much higher rates of translation in plastids to supply the abundant, rapid-turnover proteins required for photosynthesis. Previous studies have indicated that plant organellar aaRS evolve more slowly compared to mitochondrial aaRS in other eukaryotes that lack plastids. Thus, we investigated the evolution of nuclear-encoded organellar and cytosolic translation machinery across a broad sampling of angiosperms, including non-photosynthetic (heterotrophic) plant species with reduced rates of plastid gene expression to test the hypothesis that translational demands associated with photosynthesis constrain the evolution of bacterial-like enzymes involved in organellar tRNA metabolism. Remarkably, heterotrophic plants exhibited wholesale loss of many organelle-targeted aaRS and other enzymes, even though translation still occurs in their mitochondria and plastids. These losses were often accompanied by apparent retargeting of cytosolic enzymes and tRNAs to the organelles, sometimes preserving aaRS-tRNA charging relationships but other times creating surprising mismatches between cytosolic aaRS and mitochondrial tRNA substrates. Our findings indicate that the presence of a photosynthetic plastid drives the retention of specialized systems for organellar tRNA metabolism.</p>
In silico subcellular targeting predictions for cytosolic aminoacyl tRNA-synthetases (aaRS) in parasitic plants
Open the record for dataset details and reuse information.
Raw data and metadata of SiO2 NP physicochemical characterisation, in vitro investigations and in silico predictions on protein corona formation
<p>Raw data and metadata of SiO<sub>2</sub> NP physicochemical characterisation, <em>in vitro</em> investigations and<em> in silico</em> predictions on protein corona formation. Data repository for Hasenkopf I. et al., 2022. Please note that “pristine” NPs in metadata are referred to as “bare” NPs in the paper.</p> <p>1. NP phys-chem raw data (xlsx-formatted) + corresponding metadata (xlsx-formatted)</p> <p>2. NP-protein binding raw data (xlsx-formatted) + corresponding metadata (xlsx-/pdf-formatted, ImageLab-formatted analyses, pdf-formatted ImageLab reports)</p> <p>3. huRBL mediator release raw data (xlsx-formatted) + corresponding metadata (pdf-formatted)</p> <p>4. <em>in silico</em> modelling raw data (xlsx-formatted) + corresponding metadata (config-/map-formatted)</p> <p> </p>
In silico prediction of high-resolution Hi-C interaction matrices (part III)
<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>). There are a total of six files in this dataset: Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz and Nhek.tgz. The Data.tgz include predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz contain trained models, predictions, feature files for two chromosomes for in each cell line.</p> <p>This is part III of the dataset which contains Nhek.tgz and Data.tgz.</p>
In silico prediction of high-resolution Hi-C interaction matrices (part I)
<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>). There are a total of six files in this dataset: Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz and Nhek.tgz. The Data.tgz include predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz contain trained models, predictions, feature files for two chromosomes for in each cell line.</p> <p>This is part I of the dataset which contains Gm12878.tgz and Hmec.tgz.</p>
In silico prediction of high-resolution Hi-C interaction matrices (part II)
<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>). There are a total of six files in this dataset: Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz and Nhek.tgz. The Data.tgz include predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz contain trained models, predictions, feature files for two chromosomes for in each cell line.</p> <p>This is part II of the dataset which contains K562.tgz and Huvec.tgz.</p>
In silico prediction and biophysical validation of novel 14-3-3σ homodimer stabilizers
<p>this dataset is related to "In silico prediction and biophysical validation of novel 14-3-3σ homodimer stabilizers" Aljabal G., Teh, A.-H., Yap B.K.</p>
Fig. 4 in Characterization of undescribed melanoma inhibitors from Euphorbia mauritanica L. cultivated in Egypt targeting BRAF and MEK 1 kinases via in-silico study and ADME prediction
Fig. 4. Experimental and TDDFT-simulated electronic circular dichroism spctra of (A) euphomauritanol A (1), (B) euphomauritanol B (2), and euphomauritanophane A (3).
Fig. 6 in Characterization of undescribed melanoma inhibitors from Euphorbia mauritanica L. cultivated in Egypt targeting BRAF and MEK 1 kinases via in-silico study and ADME prediction
Fig. 6. Bioavailability radar chart from Swiss ADME online web tool for compounds (A) 1, (B) 2 and (C) 3. The pink area represents the range of the optimal property values for oral bioavailability and the red line is compounds (A) 1, (B) 2 and (C) 3 predicted properties. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 5 in Characterization of undescribed melanoma inhibitors from Euphorbia mauritanica L. cultivated in Egypt targeting BRAF and MEK 1 kinases via in-silico study and ADME prediction
Fig. 5. (A) 2D binding mode, (B) 3D binding mode of Euphomauritanol A (1) in the active site of BRaf Kinase V600E (PDB ID: 4XV2) (C) 2D binding mode, and (D) 3D binding mode of euphomauritanophane A (3) in the active site of MEK 1 kinase (PDB ID: 4LMN).
Supporting data for "The benefit of in silico predicted spectral libraries in data-independent acquisition data analysis workflows"
Open the record for dataset details and reuse information.
Galaxy Histories with in silico mass spectra of Mirex, Ethylene, Benzophenone and Enilconazole predicted via QCxMS
<p>Galaxy histories containing spectra predicted using the QCxMS software. The molecules are Mirex, Ethylene, Benzophenone and Enilconazole. Calculations have been carried our using Galaxy and histories exported as ROCrates.</p>
C.Origami: Cell type-specific prediction of 3D chromatin organization enables high-throughput in silico genetic screening
<p>This data repository includes dataset generated from IMR-90, GM12878 cell lines and C.Origami pre-trained model weights.</p> <p>GitHub page: https://github.com/tanjimin/C.Origami</p>
In silico Antibody-Peptide Epitope prediction for Personalized cancer therapy publication data
<p>In silico Antibody-Peptide Epitope prediction for Personalized cancer therapy publication data</p>
Data from: In silico site-directed mutagenesis informs species-specific predictions of chemical susceptibility derived from the Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool
Chemical hazard assessment requires extrapolation of information from model organisms to all species of concern. The Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool was developed as a rapid, cost effective method to aid cross-species extrapolation of susceptibility to chemicals acting on specific protein targets through evaluation of protein structural similarities and differences. The greatest resolution for extrapolation of chemical susceptibility across species involves comparisons of individual amino acid residues at key positions involved in protein-chemical interactions. However, a lack of understanding of whether specific amino acid substitutions among species at key positions in proteins affect interaction with chemicals made manual interpretation of alignments time consuming and potentially inconsistent. Therefore, this study used in silico site-directed mutagenesis coupled with docking simulations of computational models for acetylcholinesterase (AChE) and ecdysone receptor (EcR) to investigate how specific amino acid substitutions impact protein-chemical interaction. This study found that computationally derived substitutions in identities of key amino acids caused no change in protein-chemical interaction if residues share the same side chain functional properties and have comparable molecular dimensions, while differences in these characteristics can change protein-chemical interaction. These findings were considered in the development of capabilities for automatically generated species-specific predictions of chemical susceptibility in SeqAPASS. These predictions for AChE and EcR were shown to agree with SeqAPASS predictions comparing the primary sequence and functional domain sequence of proteins for more than 90 % of the investigated species, but also identified dramatic species-specific differences in chemical susceptibility that align with results from standard toxicity tests. These results provide a compelling line-of-evidence for use of SeqAPASS in deriving screening level, species-specific, susceptibility predictions across broad taxonomic groups for application to human and ecological hazard assessment.
Fig. 7. Predicted Boiled-Egg plot from Swiss ADME online web tool for compounds 1–3 in Characterization of undescribed melanoma inhibitors from Euphorbia mauritanica L. cultivated in Egypt targeting BRAF and MEK 1 kinases via in-silico study and ADME prediction
Fig. 7. Predicted Boiled-Egg plot from Swiss ADME online web tool for compounds 1–3.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.