Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

56

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

56 results for “Virtual screening”

Learn how ShareScore rates datasets ↗
zenodo52/100

Datasets for practical model selection for prospective virtual screening

<p>This repository contains datasets for the manuscript &quot;Practical model selection for prospective virtual screening&quot;:</p> <ul> <li><strong>pria_rmi_cv.tar.gz</strong>: A compressed directory containing chemical screening data for the&nbsp;<strong>PriA-SSB AS</strong>,&nbsp;<strong>PriA-SSB FP</strong>, and <strong>RMI-FANCM FP</strong> binary datasets.&nbsp; The files also contain the associated continuous % inhibition values and chemical features represented as SMILES and Morgan fingerprints.&nbsp; The dataset has been split into five folds for cross validation.</li> <li><strong>pria_rmi_pcba_cv.tar.gz</strong>: A compressed directory containing chemical screening data for the&nbsp;<strong>PriA-SSB AS</strong>,&nbsp;<strong>PriA-SSB FP</strong>, and <strong>RMI-FANCM FP</strong> binary datasets as well as public PubChem BioAssay datasets.&nbsp; The files also contain the&nbsp;PriA-SSB and&nbsp;RMI-FANCM&nbsp;continuous % inhibition values and chemical features represented as SMILES and Morgan fingerprints.&nbsp; The dataset has been split into five folds for cross validation.&nbsp; Missing values are left blank.</li> <li><strong>pria_prospective.csv.gz</strong>: A compressed file containing chemical screening data for the binary&nbsp;dataset&nbsp;<strong>PriA-SSB prospective</strong>.&nbsp;&nbsp;The file&nbsp;also contains the continuous % inhibition values and chemical features represented as SMILES and Morgan fingerprints.</li> </ul> <p>If you use&nbsp;these&nbsp;data in a publication, please cite:</p> <p>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter. Practical Model Selection for Prospective Virtual Screening. Journal of Chemical Information and Modeling. 2018 <a href="https://doi.org/10.1021/acs.jcim.8b00363">doi:10.1021/acs.jcim.8b00363</a></p> <p>PubChem data were provided by the&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/">PubChem database</a>.&nbsp; Follow the <a href="https://pubchemdocs.ncbi.nlm.nih.gov/citation-guidelines">PubChem citation guidelines</a> if you use the PubChem data.&nbsp; See <a href="https://doi.org/10.1177/2472555217712001">Voter et al. 2017</a>&nbsp;(PubChem AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>) for the PriA-SSB screening data and <a href="https://doi.org/10.1177/1087057116635503">Voter et al. 2016</a> (PubChem AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1159607">1159607</a>) for RMI-FANCM.</p> <p>Version 1.1.0 updates&nbsp;all of the data files.&nbsp; We standardized the SMILES in all files by generating canonical SMILES with RDKit&nbsp;version 2016.03.4.&nbsp; In addition, we removed 2845 chemicals from&nbsp;pria_prospective.csv.gz that were duplicates of compounds in&nbsp;pria_rmi_cv.tar.gz.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Virtual screening on Nsp16: screening of 1084 compounds in VeroE6-eGFP cells

<p>This report describes the most relevant results of virtually screening the Janssen Pharmaceutica compound collection for potential activity against SARS-CoV-2 Nsp16 and confirmation of potential hits in a VeroE6 cell-based anti-SARS-CoV-2 assay.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Screening of 2694 RdRP virtual screening hits in RdRP/Nsp7/Nsp8 biochemical assay and confirmation in cellular SARS-CoV-2 assay

<p>This report describes the most relevant results of virtually screening the Janssen Pharmaceutica compound collection for potential activity against SARS-CoV-2 RNA polymerase and confirmation of potential hits in a biochemical SARS-CoV RTC assay and A549-hACE2 cell-based anti-SARS-CoV-2 assay.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Identification of potential modulators of IFITM3 by in-silico modeling and virtual screening

<p>Modeled structure of IFITM3 and Desmond MD trajectory files for IFITM3 and IFITM3-ligand complexes. Please see README file.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

PSnpBind: A database of mutated binding site protein-ligand complexes constructed using a multithreaded virtual screening workflow

<p>A key concept in drug design is how natural variants, especially the ones occurring in the binding site of drug targets, affect the inter-individual drug response and efficacy by altering binding affinity. These effects have been studied on very limited and small datasets while, ideally, a large dataset of binding affinity changes due to binding site single-nucleotide polymorphisms (SNPs) is needed for evaluation. However, to the best of our knowledge, such a dataset does not exist. Thus, a reference dataset of ligands binding affinities to proteins with all their reported binding sites&rsquo; variants was constructed using a molecular docking approach. Having a large database of protein-ligand complexes covering a wide range of binding pocket mutations and a large small molecules&rsquo; landscape is of great importance for several types of studies. For example, developing machine learning algorithms to predict protein-ligand affinity or a SNP effect on it requires an extensive amount of data. In this work, we present PSnpBind: A large database of mutated binding site protein-ligand complexes constructed using a multithreaded virtual screening workflow. It provides a web interface to explore and visualize the protein-ligand complexes and a REST API to programmatically access the different aspects of the database contents. PSnpBind is freely available at <a href="https://psnpbind.org">https://psnpbind.org</a>.<strong> </strong>The source code of the tools used in constructing PSnpBind is available on <a href="https://github.com/ammar257ammar/PSnpBind-Build">GitHub</a>.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Figure 5: Screen shot of the generated collaborative virtual environment

<p>The output of the design is an xml file. The technical user has to customize<br> the designed session in order to put in the virtual environment engine in order<br> to generate the session. The customization regards technical parameters (for<br> example spatial coordinate, intensity of the light and so on). A screen shot of<br> the generated session is in figure 5.</p>

opencc-by-4.0Oct 2010View details →
zenodo40/100

The Pan-Canadian Chemical Library: A Mechanism to Open Academic Chemistry to High-Throughput Virtual Screening

<h1>Pan-Canadian Chemical Library</h1> <p>This Zenodo repository contains the cheap and druglike subset of the Pan-Canadian Chemical Library (PCCL) project. For more information, visit&nbsp;<a href="https://pccl.thesgc.org/" rel="nofollow">https://pccl.thesgc.org</a>.</p> <h2>PCCL library</h2> <p>The PCCL library is splitted by reaction, then by number of heavy atoms. Two types of files are available in zip archives:</p> <ul> <li>The SMILES format files, with the SMILES string and their product name,</li> <li>The CSV format file, with all the information generated during their enumeration: reagents, druglike properties, etc.</li> </ul> <p>Note: Purchasability is defined according to two integers: 1 for products only composed of BB-50 reagents, and 2 for products composed of BB-40 or BB-50 reagents. Read more about the meaning of these reagents groups in the article below.</p> <h2>Citation</h2> <p>If you find the PCCL useful or if you use it, please cite our paper:</p> <p>Bedart, C. <em>et al.</em> The Pan-Canadian Chemical Library: A mechanism to open academic chemistry to high-throughput virtual screening. Scientific Data 11, (2024).<br>doi: <a title="10.1038/s41597-024-03443-5" href="https://www.nature.com/articles/s41597-024-03443-5">10.1038/s41597-024-03443-5</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Prototyping 3D Virtual Learning Environments with X3D-based Content and Visualization Tools-Figure 10. Room model generated with Autodesk 123D Catch - the 3D model (screen capture from GLC Player)

<p>Structure from motion was used for rapid modeling of a small room with all its objects. Two files were generated, a Wavefront obj and mtl (corresponding to the texture). The 3D model was post-processed with MeshLab, during which several filters were applied to clean up the model. The mesh model was also connected with the scanned model, by choosing at least 4 connection points. The 2D and 3D results are shown in Figures 9, 10. A post-processing could also be performed using the Autodesk 123D Catch web application.</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Hit Expansion using Substructure Search, Virtual Screening & Free Energy Perturbation

<p>Identification of&nbsp;commercially available chemical analogs of primary hits previously crystallized in complex with the zinc finger ubiquitin binding domain (Zf-UBD) of USP5 and prioritization of chemical analogues by free energy perturbation (FEP).&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Virtual Screening with Molecular Forecaster

<p>Commercially available compounds for USP5 zinc-finger ubiquitin binding domain (ZnF-UBD) were identified with Molecular Forecasters (MFI) FITTED docking platform. Preliminary assessment of docking for USP5 ZnF-UBD with FITTED can be found <a href="https://zenodo.org/record/2620208#.XQO2pYhKjIV">here</a>.</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

Deep Reinforcement Learning Enables Better Bias Control in Benchmark for Virtual Screening

<p>This compressed file contains all datasets made for the validation of MUBDsyn.</p><ul><li>datasets_int_val: 17 cases in this folder are derived from&nbsp;<a href="https://github.com/jwxia2014/ULS-UDS">MUBD for GPCRs</a>. MUBDreal was made by <a href="https://github.com/jwxia2014/MUBD-DecoyMaker2.0">MUBD-DecoyMaker2.0</a> and MUBDsyn&nbsp;was made by <a href="https://github.com/taoshen99/MUBDsyn">MUBD-DecoyMakersyn</a>.</li><li>datasets_ext_val_classical_VS: Five&nbsp;cases in this folder are derived from the shared cases of MUV and DUD-E. The active sets of MUV were taken as the input to make corresponding MUBD datasets. Files in SBVS are raw molecular docking results by smina.</li><li>datasets_ext_val_SI_classical_VS: DeepCoy and TocoDecoy were used to make the datasets corresponding to the same five cases above. The data of&nbsp;DeepCoy was directly&nbsp;retrieved from&nbsp;<a href="https://opig.stats.ox.ac.uk/resources">DeepCoy resources at OPIG</a>&nbsp;while topology decoys of TocoDecoy_9W were&nbsp;made based on the scripts provided at&nbsp;<a href="https://github.com/5AGE-zhang/TocoDecoy">TocoDecoy GitHub Repository</a>. Files in SBVS are raw molecular docking results by smina.</li><li>datasets_ext_val_ML_VS: Ten&nbsp;cases in this folder are derived from <a href="http://nrlist.drugdesign.fr/">NRLiSt-BDB</a>. Corresponding MUBD datasets were made as described above.</li></ul><p>All these datasets can be used for the reproduction of validation performed in the manuscript or to benchmark various virtual screening methods.</p>

openapache2.0May 2023View details →
zenodo40/100

Optimal design of virtual screening benchmarks from in vitro screening data

<p>Input Datasets and output data for the creation of a new Benchmark dataset using PubChem Bioassay data.&nbsp;</p>

opencc-byAug 2023View details →
zenodo40/100

Benchmarking of PROTAC docking and virtual screening tools - dataset

<p>This&nbsp;repository includes all the input files&nbsp;for the PROTAC docking and virtual screening benchmark. The raw data files are available on request (all raw data files combined is ~70GB and compressed ~47GB).&nbsp;Researchers can access and download the data to reproduce&nbsp;the results. Details about files/folders is included in the README.txt.</p>

opencc-by-4.0Aug 2023View details →
dryad36/100

Large scale virtual screening for finding inhibitor against the main protease from herbal medicine for SARS-Cov2 therapy

<p>The pandemic COVID-19 caused by SARS-CoV-2 has raised global health concerns. However, there is still no targeted medicine available for treatment of this disease. It is reported that 3CLpro plays an important role for the life cycle of virus and this protein has also been proved as an effective drug target in the case of severe acute respiratory syndrome coronavirus (SARS-CoV) and Middle East respiratory syndrome coronavirus (MERS-CoV). Medicinal plantsare important resource for drug discovery. Therefore, we executed large scale virtual screening on the herbal medicine library and hoped to find a potential drug againstthe main protease. As a result, we obtained three available compounds derived from Chinese herbal medicines through the docking and binding free energy calculation.Then 100 ns molecular dynamic (MD) were employed to uncover the potential mechanisms which is helpful for further drug optimization.</p>

opencc-zeroJan 2021View details →
dryad36/100

Performance of virtual screening against GPCR homology models: Impact of template selection and treatment of binding site plasticity

<p>Rational drug design for G protein-coupled receptors (GPCRs) is limited by the small number of available atomic resolution structures. We assessed the use of homology modeling to predict the structures of two therapeutically relevant GPCRs and strategies to improve the performance of virtual screening against modeled binding sites. Homology models of the D<sub>2</sub> dopamine (D<sub>2</sub>R) and serotonin 5-HT<sub>2A</sub> receptors (5-HT<sub>2A</sub>R) were generated based on crystal structures of 16 different GPCRs. Comparison of the homology models to D<sub>2</sub>R and 5-HT<sub>2A</sub>R crystal structures showed that accurate predictions could be obtained, but not necessarily using the most closely related template. Assessment of virtual screening performance was based on molecular docking of ligands and decoys. The results demonstrated that several templates and multiple models based on each of these must be evaluated to identify the optimal binding site structure. Models based on aminergic GPCRs displayed ligand enrichment and there was a trend toward improved virtual screening performance with increasing binding site accuracy. The best models even displayed ligand enrichment better than that of the D<sub>2</sub>R and 5-HT<sub>2A</sub>R crystal structures. Methods to consider binding site plasticity were explored to further improve predictions. Molecular docking to ensembles of structures did not outperform the best individual binding site models, but could increase the diversity of hits from virtual screens and be advantageous for GPCR targets with few known ligands. Molecular dynamics refinement resulted in moderate improvements of structural accuracy and the virtual screening performance of snapshots was either comparable to or worse than that of the raw homology models. These results provide guidelines for successful application of structure-based ligand discovery using GPCR homology models.</p>

opencc-zeroMar 2020View details →
zenodo36/100

Screen Capture & Audior Recordings - Empirical Study "Business Process Model Validation Through Virtual Enactment"

<p>Screen captures of task completion and audio recordings of interviews of the empirical study which has been conducted within the context of the master thesis "Business Process Model Validation Through Virtual Enactment"</p>

opencc-by-4.0Aug 2017View details →
zenodo36/100

ESSENCE-Dock: A Consensus-Based Approach to Enhance Virtual Screening Enrichment in Drug Discovery

<p>All of the individual docking data and ESSENCE-Dock consensus results for 21 diverse DUD-E targets as presented in the paper "ESSENCE-Dock: A Consensus-Based Approach to Enhance Virtual Screening Enrichment in Drug Discovery".</p> <p>The data is sorted per DUD-E target. It contains the prepared data that was used for the docking calculations (in the Undocked directory), as well as our docking results. Finally, our ESSENCE-Dock Consensus results are included as well</p> <p>Docking calculations were performed using:</p> <ul> <li><a href="https://github.com/bio-hpc/metascreener">Metascreener (V1.1)</a> (Gnina and LeadFinder Calculations; prefix VS_GN_ and VS_LF_ respectively)</li> <li><a href="https://github.com/Jnelen/DiffDockHPC/tree/DiffDockHPCv1.0">DiffDockHPC (v1.0)</a> (DiffDock calculations; prefix VS_DD_ )</li> </ul> <p>The consensus calculations were performed using ESSENCE-Dock, available via <a href="https://github.com/bio-hpc/metascreener">Metascreener </a>as well.</p> <p>The whole methodology and all of the details are described in the ESSENCE-Dock paper: <a href="https://doi.org/10.1021/acs.jcim.3c01982">https://doi.org/10.1021/acs.jcim.3c01982</a></p> <p><strong>Paper Abstract</strong></p> <p>Drug development is a complex, costly, and time-consuming endeavor. While high-throughput screening (HTS) plays a critical role in the discovery stage, it is one of many factors contributing to these challenges. In certain contexts, virtual screening can complement HTS, potentially offering a more streamlined approach in the initial stages of drug discovery. Molecular docking is an example of a popular virtual screening technique that is often used for this purpose, however, its effectiveness can vary greatly. This has led to the use of consensus docking approaches, which combine results from different docking methods to improve the identification of active compounds and reduce the occurrence of false positives. However, many of these methods do not fully leverage the latest advancements in molecular docking.<br>In response, we present ESSENCE-Dock (Effective Structural Screening ENrichment ConsEnsus Dock), a new consensus docking workflow aimed at decreasing false positives and increasing the discovery of active compounds. By utilizing a combination of novel docking algorithms, we improve the selection process for potential active compounds. ESSENCE-Dock has been made to be user-friendly, requiring only a few simple commands to perform a complete screening, while also being designed for use in high-performance computing (HPC) environments.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

"iDCNNPred: An interpretable deep learning model for virtual screening and identification of PI3Ka inhibitors against triple-negative breast cancer"

<p>In this study, we proposed a novel interpretable deep convolutional neural network prediction (iDCNNPred) system for classifying molecular bioactivity and identifying predictive potential inhibitors for the PI3Ka isoform protein. This system utilizes 2D molecular image representation as input features, instead of traditional molecular fingerprints or descriptors.</p> <p><strong>The datasets used for model construction, prediction and screening of chemical library are provided in this uploaded data in <a href="../api/records/10947610/draft/files/Molecular_image_Custom_DCNN_datasets.zip/content" target="_blank" rel="noopener noreferrer">Molecular_image_Custom_DCNN_datasets.zip</a> file for Custom-DCNN models and <a href="../api/records/10947610/draft/files/Molecular_image_pre_trained_datasets.zip/content" target="_blank" rel="noopener noreferrer">Molecular_image_pre_trained_datasets.zip</a> file for Pre-trained fine-tuned models. </strong><strong>The final run of models results given in file <a href="../api/records/10947610/draft/files/Custom_DCNN_Pre_trained_models.zip/content" target="_blank" rel="noopener noreferrer">Custom_DCNN_Pre_trained_models.zip</a></strong></p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

CryoXKit virtual screening set

<p>Virtual screening dataset used in&nbsp;<a href="https://chemrxiv.org/engage/chemrxiv/article-details/6723b0a0f9980725cfb49751"><em>Docking guidance </em><em>with experimental ligand structural density </em><em>improves docking pose prediction and virtual </em><em>screening performance</em></a>. Modification of <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155">LIT-PCBA</a> dataset. The dataset used for TRPV1 is available from&nbsp;<a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c00312">Llanos et al.</a></p> <p>&nbsp;</p> <p>Also included are&nbsp;</p> <pre><code>run_bias.template.qsub</code></pre> <pre><code>submit.sh</code></pre> <p>Together, these scripts provide examples of how to set-up and run screening density-guided simulations on a SLURM-based HPC cluster. Please note that this dataset does not include the structural density files associated with each receptor. These must be downloaded separately from the RCSB PDB (hyperlinks for each receptor may be found in SI Table 2 of the associated publication for this dataset).</p> <p>&nbsp;</p> <p><strong>Targets included:</strong></p> <p>ADRB2</p> <p>ALDH1</p> <p>FEN1</p> <p>GBA</p> <p>HSP90a</p> <p>MAPK1</p> <p>MTORC1</p> <p>PKM2</p> <p>VDR</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

PanDDA files from a ligand screen against the NSP3 macrodomain of SARS-CoV-2 - ligands from fragment merging/linking and virtual screening

<p>This deposition contains the X-ray diffraction data&nbsp;used to run PanDDA&nbsp;in the&nbsp;ligand screen against the NSP3 macrodomain of SARS-CoV-2 described in Gahbauer et al. 2022 (doi: https://doi.org/10.1101/2022.06.27.497816).</p> <p>mac1_pandda.zip contains the structure factor intensities,&nbsp;PanDDA input/ouput and&nbsp;refined models/maps.&nbsp;A description of the files can be found in the README&nbsp;file.&nbsp;</p> <p>mac1_ligand-bound_states.zip contains the ligand-bound states extracted from the multi-state PDB files.&nbsp;</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record