Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.7.1
Dataset results
56 results for “virtual screening”
Datasets for practical model selection for prospective virtual screening
<p>This repository contains datasets for the manuscript "Practical model selection for prospective virtual screening":</p> <ul> <li><strong>pria_rmi_cv.tar.gz</strong>: A compressed directory containing chemical screening data for the <strong>PriA-SSB AS</strong>, <strong>PriA-SSB FP</strong>, and <strong>RMI-FANCM FP</strong> binary datasets. The files also contain the associated continuous % inhibition values and chemical features represented as SMILES and Morgan fingerprints. The dataset has been split into five folds for cross validation.</li> <li><strong>pria_rmi_pcba_cv.tar.gz</strong>: A compressed directory containing chemical screening data for the <strong>PriA-SSB AS</strong>, <strong>PriA-SSB FP</strong>, and <strong>RMI-FANCM FP</strong> binary datasets as well as public PubChem BioAssay datasets. The files also contain the PriA-SSB and RMI-FANCM continuous % inhibition values and chemical features represented as SMILES and Morgan fingerprints. The dataset has been split into five folds for cross validation. Missing values are left blank.</li> <li><strong>pria_prospective.csv.gz</strong>: A compressed file containing chemical screening data for the binary dataset <strong>PriA-SSB prospective</strong>. The file also contains the continuous % inhibition values and chemical features represented as SMILES and Morgan fingerprints.</li> </ul> <p>If you use these data in a publication, please cite:</p> <p>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter. Practical Model Selection for Prospective Virtual Screening. Journal of Chemical Information and Modeling. 2018 <a href="https://doi.org/10.1021/acs.jcim.8b00363">doi:10.1021/acs.jcim.8b00363</a></p> <p>PubChem data were provided by the <a href="https://pubchem.ncbi.nlm.nih.gov/">PubChem database</a>. Follow the <a href="https://pubchemdocs.ncbi.nlm.nih.gov/citation-guidelines">PubChem citation guidelines</a> if you use the PubChem data. See <a href="https://doi.org/10.1177/2472555217712001">Voter et al. 2017</a> (PubChem AID <a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>) for the PriA-SSB screening data and <a href="https://doi.org/10.1177/1087057116635503">Voter et al. 2016</a> (PubChem AID <a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1159607">1159607</a>) for RMI-FANCM.</p> <p>Version 1.1.0 updates all of the data files. We standardized the SMILES in all files by generating canonical SMILES with RDKit version 2016.03.4. In addition, we removed 2845 chemicals from pria_prospective.csv.gz that were duplicates of compounds in pria_rmi_cv.tar.gz.</p>
Virtual screening on Nsp16: screening of 1084 compounds in VeroE6-eGFP cells
<p>This report describes the most relevant results of virtually screening the Janssen Pharmaceutica compound collection for potential activity against SARS-CoV-2 Nsp16 and confirmation of potential hits in a VeroE6 cell-based anti-SARS-CoV-2 assay.</p>
Screening of 2694 RdRP virtual screening hits in RdRP/Nsp7/Nsp8 biochemical assay and confirmation in cellular SARS-CoV-2 assay
<p>This report describes the most relevant results of virtually screening the Janssen Pharmaceutica compound collection for potential activity against SARS-CoV-2 RNA polymerase and confirmation of potential hits in a biochemical SARS-CoV RTC assay and A549-hACE2 cell-based anti-SARS-CoV-2 assay.</p>
Identification of potential modulators of IFITM3 by in-silico modeling and virtual screening
<p>Modeled structure of IFITM3 and Desmond MD trajectory files for IFITM3 and IFITM3-ligand complexes. Please see README file.</p>
PSnpBind: A database of mutated binding site protein-ligand complexes constructed using a multithreaded virtual screening workflow
<p>A key concept in drug design is how natural variants, especially the ones occurring in the binding site of drug targets, affect the inter-individual drug response and efficacy by altering binding affinity. These effects have been studied on very limited and small datasets while, ideally, a large dataset of binding affinity changes due to binding site single-nucleotide polymorphisms (SNPs) is needed for evaluation. However, to the best of our knowledge, such a dataset does not exist. Thus, a reference dataset of ligands binding affinities to proteins with all their reported binding sites’ variants was constructed using a molecular docking approach. Having a large database of protein-ligand complexes covering a wide range of binding pocket mutations and a large small molecules’ landscape is of great importance for several types of studies. For example, developing machine learning algorithms to predict protein-ligand affinity or a SNP effect on it requires an extensive amount of data. In this work, we present PSnpBind: A large database of mutated binding site protein-ligand complexes constructed using a multithreaded virtual screening workflow. It provides a web interface to explore and visualize the protein-ligand complexes and a REST API to programmatically access the different aspects of the database contents. PSnpBind is freely available at <a href="https://psnpbind.org">https://psnpbind.org</a>.<strong> </strong>The source code of the tools used in constructing PSnpBind is available on <a href="https://github.com/ammar257ammar/PSnpBind-Build">GitHub</a>.</p>
Figure 5: Screen shot of the generated collaborative virtual environment
<p>The output of the design is an xml file. The technical user has to customize<br> the designed session in order to put in the virtual environment engine in order<br> to generate the session. The customization regards technical parameters (for<br> example spatial coordinate, intensity of the light and so on). A screen shot of<br> the generated session is in figure 5.</p>
The Pan-Canadian Chemical Library: A Mechanism to Open Academic Chemistry to High-Throughput Virtual Screening
<h1>Pan-Canadian Chemical Library</h1> <p>This Zenodo repository contains the cheap and druglike subset of the Pan-Canadian Chemical Library (PCCL) project. For more information, visit <a href="https://pccl.thesgc.org/" rel="nofollow">https://pccl.thesgc.org</a>.</p> <h2>PCCL library</h2> <p>The PCCL library is splitted by reaction, then by number of heavy atoms. Two types of files are available in zip archives:</p> <ul> <li>The SMILES format files, with the SMILES string and their product name,</li> <li>The CSV format file, with all the information generated during their enumeration: reagents, druglike properties, etc.</li> </ul> <p>Note: Purchasability is defined according to two integers: 1 for products only composed of BB-50 reagents, and 2 for products composed of BB-40 or BB-50 reagents. Read more about the meaning of these reagents groups in the article below.</p> <h2>Citation</h2> <p>If you find the PCCL useful or if you use it, please cite our paper:</p> <p>Bedart, C. <em>et al.</em> The Pan-Canadian Chemical Library: A mechanism to open academic chemistry to high-throughput virtual screening. Scientific Data 11, (2024).<br>doi: <a title="10.1038/s41597-024-03443-5" href="https://www.nature.com/articles/s41597-024-03443-5">10.1038/s41597-024-03443-5</a></p> <p> </p> <p> </p>
Prototyping 3D Virtual Learning Environments with X3D-based Content and Visualization Tools-Figure 10. Room model generated with Autodesk 123D Catch - the 3D model (screen capture from GLC Player)
<p>Structure from motion was used for rapid modeling of a small room with all its objects. Two files were generated, a Wavefront obj and mtl (corresponding to the texture). The 3D model was post-processed with MeshLab, during which several filters were applied to clean up the model. The mesh model was also connected with the scanned model, by choosing at least 4 connection points. The 2D and 3D results are shown in Figures 9, 10. A post-processing could also be performed using the Autodesk 123D Catch web application.</p>
Hit Expansion using Substructure Search, Virtual Screening & Free Energy Perturbation
<p>Identification of commercially available chemical analogs of primary hits previously crystallized in complex with the zinc finger ubiquitin binding domain (Zf-UBD) of USP5 and prioritization of chemical analogues by free energy perturbation (FEP). </p>
Virtual Screening with Molecular Forecaster
<p>Commercially available compounds for USP5 zinc-finger ubiquitin binding domain (ZnF-UBD) were identified with Molecular Forecasters (MFI) FITTED docking platform. Preliminary assessment of docking for USP5 ZnF-UBD with FITTED can be found <a href="https://zenodo.org/record/2620208#.XQO2pYhKjIV">here</a>.</p>
Deep Reinforcement Learning Enables Better Bias Control in Benchmark for Virtual Screening
<p>This compressed file contains all datasets made for the validation of MUBDsyn.</p><ul><li>datasets_int_val: 17 cases in this folder are derived from <a href="https://github.com/jwxia2014/ULS-UDS">MUBD for GPCRs</a>. MUBDreal was made by <a href="https://github.com/jwxia2014/MUBD-DecoyMaker2.0">MUBD-DecoyMaker2.0</a> and MUBDsyn was made by <a href="https://github.com/taoshen99/MUBDsyn">MUBD-DecoyMakersyn</a>.</li><li>datasets_ext_val_classical_VS: Five cases in this folder are derived from the shared cases of MUV and DUD-E. The active sets of MUV were taken as the input to make corresponding MUBD datasets. Files in SBVS are raw molecular docking results by smina.</li><li>datasets_ext_val_SI_classical_VS: DeepCoy and TocoDecoy were used to make the datasets corresponding to the same five cases above. The data of DeepCoy was directly retrieved from <a href="https://opig.stats.ox.ac.uk/resources">DeepCoy resources at OPIG</a> while topology decoys of TocoDecoy_9W were made based on the scripts provided at <a href="https://github.com/5AGE-zhang/TocoDecoy">TocoDecoy GitHub Repository</a>. Files in SBVS are raw molecular docking results by smina.</li><li>datasets_ext_val_ML_VS: Ten cases in this folder are derived from <a href="http://nrlist.drugdesign.fr/">NRLiSt-BDB</a>. Corresponding MUBD datasets were made as described above.</li></ul><p>All these datasets can be used for the reproduction of validation performed in the manuscript or to benchmark various virtual screening methods.</p>
Optimal design of virtual screening benchmarks from in vitro screening data
<p>Input Datasets and output data for the creation of a new Benchmark dataset using PubChem Bioassay data. </p>
Benchmarking of PROTAC docking and virtual screening tools - dataset
<p>This repository includes all the input files for the PROTAC docking and virtual screening benchmark. The raw data files are available on request (all raw data files combined is ~70GB and compressed ~47GB). Researchers can access and download the data to reproduce the results. Details about files/folders is included in the README.txt.</p>
Large scale virtual screening for finding inhibitor against the main protease from herbal medicine for SARS-Cov2 therapy
<p>The pandemic COVID-19 caused by SARS-CoV-2 has raised global health concerns. However, there is still no targeted medicine available for treatment of this disease. It is reported that 3CLpro plays an important role for the life cycle of virus and this protein has also been proved as an effective drug target in the case of severe acute respiratory syndrome coronavirus (SARS-CoV) and Middle East respiratory syndrome coronavirus (MERS-CoV). Medicinal plantsare important resource for drug discovery. Therefore, we executed large scale virtual screening on the herbal medicine library and hoped to find a potential drug againstthe main protease. As a result, we obtained three available compounds derived from Chinese herbal medicines through the docking and binding free energy calculation.Then 100 ns molecular dynamic (MD) were employed to uncover the potential mechanisms which is helpful for further drug optimization.</p>
Performance of virtual screening against GPCR homology models: Impact of template selection and treatment of binding site plasticity
<p>Rational drug design for G protein-coupled receptors (GPCRs) is limited by the small number of available atomic resolution structures. We assessed the use of homology modeling to predict the structures of two therapeutically relevant GPCRs and strategies to improve the performance of virtual screening against modeled binding sites. Homology models of the D<sub>2</sub> dopamine (D<sub>2</sub>R) and serotonin 5-HT<sub>2A</sub> receptors (5-HT<sub>2A</sub>R) were generated based on crystal structures of 16 different GPCRs. Comparison of the homology models to D<sub>2</sub>R and 5-HT<sub>2A</sub>R crystal structures showed that accurate predictions could be obtained, but not necessarily using the most closely related template. Assessment of virtual screening performance was based on molecular docking of ligands and decoys. The results demonstrated that several templates and multiple models based on each of these must be evaluated to identify the optimal binding site structure. Models based on aminergic GPCRs displayed ligand enrichment and there was a trend toward improved virtual screening performance with increasing binding site accuracy. The best models even displayed ligand enrichment better than that of the D<sub>2</sub>R and 5-HT<sub>2A</sub>R crystal structures. Methods to consider binding site plasticity were explored to further improve predictions. Molecular docking to ensembles of structures did not outperform the best individual binding site models, but could increase the diversity of hits from virtual screens and be advantageous for GPCR targets with few known ligands. Molecular dynamics refinement resulted in moderate improvements of structural accuracy and the virtual screening performance of snapshots was either comparable to or worse than that of the raw homology models. These results provide guidelines for successful application of structure-based ligand discovery using GPCR homology models.</p>
Screen Capture & Audior Recordings - Empirical Study "Business Process Model Validation Through Virtual Enactment"
<p>Screen captures of task completion and audio recordings of interviews of the empirical study which has been conducted within the context of the master thesis "Business Process Model Validation Through Virtual Enactment"</p>
ESSENCE-Dock: A Consensus-Based Approach to Enhance Virtual Screening Enrichment in Drug Discovery
<p>All of the individual docking data and ESSENCE-Dock consensus results for 21 diverse DUD-E targets as presented in the paper "ESSENCE-Dock: A Consensus-Based Approach to Enhance Virtual Screening Enrichment in Drug Discovery".</p> <p>The data is sorted per DUD-E target. It contains the prepared data that was used for the docking calculations (in the Undocked directory), as well as our docking results. Finally, our ESSENCE-Dock Consensus results are included as well</p> <p>Docking calculations were performed using:</p> <ul> <li><a href="https://github.com/bio-hpc/metascreener">Metascreener (V1.1)</a> (Gnina and LeadFinder Calculations; prefix VS_GN_ and VS_LF_ respectively)</li> <li><a href="https://github.com/Jnelen/DiffDockHPC/tree/DiffDockHPCv1.0">DiffDockHPC (v1.0)</a> (DiffDock calculations; prefix VS_DD_ )</li> </ul> <p>The consensus calculations were performed using ESSENCE-Dock, available via <a href="https://github.com/bio-hpc/metascreener">Metascreener </a>as well.</p> <p>The whole methodology and all of the details are described in the ESSENCE-Dock paper: <a href="https://doi.org/10.1021/acs.jcim.3c01982">https://doi.org/10.1021/acs.jcim.3c01982</a></p> <p><strong>Paper Abstract</strong></p> <p>Drug development is a complex, costly, and time-consuming endeavor. While high-throughput screening (HTS) plays a critical role in the discovery stage, it is one of many factors contributing to these challenges. In certain contexts, virtual screening can complement HTS, potentially offering a more streamlined approach in the initial stages of drug discovery. Molecular docking is an example of a popular virtual screening technique that is often used for this purpose, however, its effectiveness can vary greatly. This has led to the use of consensus docking approaches, which combine results from different docking methods to improve the identification of active compounds and reduce the occurrence of false positives. However, many of these methods do not fully leverage the latest advancements in molecular docking.<br>In response, we present ESSENCE-Dock (Effective Structural Screening ENrichment ConsEnsus Dock), a new consensus docking workflow aimed at decreasing false positives and increasing the discovery of active compounds. By utilizing a combination of novel docking algorithms, we improve the selection process for potential active compounds. ESSENCE-Dock has been made to be user-friendly, requiring only a few simple commands to perform a complete screening, while also being designed for use in high-performance computing (HPC) environments.</p>
"iDCNNPred: An interpretable deep learning model for virtual screening and identification of PI3Ka inhibitors against triple-negative breast cancer"
<p>In this study, we proposed a novel interpretable deep convolutional neural network prediction (iDCNNPred) system for classifying molecular bioactivity and identifying predictive potential inhibitors for the PI3Ka isoform protein. This system utilizes 2D molecular image representation as input features, instead of traditional molecular fingerprints or descriptors.</p> <p><strong>The datasets used for model construction, prediction and screening of chemical library are provided in this uploaded data in <a href="../api/records/10947610/draft/files/Molecular_image_Custom_DCNN_datasets.zip/content" target="_blank" rel="noopener noreferrer">Molecular_image_Custom_DCNN_datasets.zip</a> file for Custom-DCNN models and <a href="../api/records/10947610/draft/files/Molecular_image_pre_trained_datasets.zip/content" target="_blank" rel="noopener noreferrer">Molecular_image_pre_trained_datasets.zip</a> file for Pre-trained fine-tuned models. </strong><strong>The final run of models results given in file <a href="../api/records/10947610/draft/files/Custom_DCNN_Pre_trained_models.zip/content" target="_blank" rel="noopener noreferrer">Custom_DCNN_Pre_trained_models.zip</a></strong></p>
CryoXKit virtual screening set
<p>Virtual screening dataset used in <a href="https://chemrxiv.org/engage/chemrxiv/article-details/6723b0a0f9980725cfb49751"><em>Docking guidance </em><em>with experimental ligand structural density </em><em>improves docking pose prediction and virtual </em><em>screening performance</em></a>. Modification of <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155">LIT-PCBA</a> dataset. The dataset used for TRPV1 is available from <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c00312">Llanos et al.</a></p> <p> </p> <p>Also included are </p> <pre><code>run_bias.template.qsub</code></pre> <pre><code>submit.sh</code></pre> <p>Together, these scripts provide examples of how to set-up and run screening density-guided simulations on a SLURM-based HPC cluster. Please note that this dataset does not include the structural density files associated with each receptor. These must be downloaded separately from the RCSB PDB (hyperlinks for each receptor may be found in SI Table 2 of the associated publication for this dataset).</p> <p> </p> <p><strong>Targets included:</strong></p> <p>ADRB2</p> <p>ALDH1</p> <p>FEN1</p> <p>GBA</p> <p>HSP90a</p> <p>MAPK1</p> <p>MTORC1</p> <p>PKM2</p> <p>VDR</p>
PanDDA files from a ligand screen against the NSP3 macrodomain of SARS-CoV-2 - ligands from fragment merging/linking and virtual screening
<p>This deposition contains the X-ray diffraction data used to run PanDDA in the ligand screen against the NSP3 macrodomain of SARS-CoV-2 described in Gahbauer et al. 2022 (doi: https://doi.org/10.1101/2022.06.27.497816).</p> <p>mac1_pandda.zip contains the structure factor intensities, PanDDA input/ouput and refined models/maps. A description of the files can be found in the README file. </p> <p>mac1_ligand-bound_states.zip contains the ligand-bound states extracted from the multi-state PDB files. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.