Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

873

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

873 results for “ligands”

Learn how ShareScore rates datasets ↗
zenodo40/100

CheckMyBlob ligand data set (CMB)

<p>Ligand data set prepared for the CheckMyBlob study, described&nbsp;in&nbsp;<em>&quot;Automatic recognition of ligands in electron density by machine learning methods&quot;</em>&nbsp;by Kowiel, M.&nbsp;<em>et al.</em>&nbsp;It contains only structures from X-ray diffraction experiments determined to at least 4.0 &Aring; resolution. Entries with R factor above 0.3 or ligands below 0.3 occupancy (according to wwPDB validation reports) were rejected. Only ligands with at least 2 non-H atoms were considered and structures with low ligand map correlation coefficients (RSCC &lt; 0.6, RSZO &lt;= 1, RSZD &gt; 6.0) were removed. Apart from taking into account quality factors, we removed from the experimental data set all moieties that are not considered proper ligands. These included: unknown species, water molecules, standard amino acids, and selected nucleotides. Moreover, connected ligands (as per the naming convention in the PDB) were labeled as alphabetically ordered strings of hetero-compound codes (e.g., NAG-NAG-NAG-NAG). Finally, the data set was limited to 200 most popular ligands. The resulting data set consisted of 219,986 examples with individual ligand counts ranging from 48,490 examples for SO4 (sulfate ion) to 106 for A2G (n-acetyl-2-deoxy-2-amino-galactose). More details concerning data selection can be found in the paper of Kowiel&nbsp;<em>et al.</em></p> <p>For machine learning (classification) purposes, the target attribute is: <strong>res_name</strong>.</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

SAXS-based titrations of periplasmic binding proteins against known ligands

<p>This dataset contains the small-angle X-ray scattering intensities of a series of titrations between proteins HisBP, GlnBP, and DEBP against a number of amino acids. A total of five synchrotron beamline sessions are included to test for reproducibility across setups and total exposure time per sample:</p> <ul> <li>DESY P12 (4s)</li> <li>ESRF BM29 (5s)</li> <li>Australian Synchrotron SAXS/WAXS (20s)</li> <li>Diamond B21 (40s)</li> </ul> <p>The accompanying manuscripts can be found on bioRxiv MS 715193, entitled: &quot;Structure-based screening of binding affinities via small-angle X-ray scattering&quot;. It is currently under consideration with te Biophysical Journal.</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

The data for: Fucose as a nutrient ligand for Dikarya and a building block of early diverging lineages

<div><br><strong>Supplementary Figure&nbsp;</strong> <div> <div> <div></div> <div> <p>&nbsp;</p> <p>CLANS clustering of GT-A clan peptidyltransferases present in Fungi with animal representatives of transferases acting on glucosylated, fucosylated and galactosylated peptides. After clustering of fungal Fringe homologs with 42 animal representatives of both LFNG and C1GLT transferases, we obtained four clusters with 4946 fungal sequences. However, only four fungal sequences (Rozella allomycis RKP19558.1,RKP21212.1, EPZ36320.1 and Fusarium oxysporum EXK84998.1) group together with the human LFNG (Q8NES3) protein sequence. Two major fungal clusters are equally distant from LFNG and C1GLT clans. Basidiobolus meristosporus ORX91553.1 and Batrachochytrium salamandrivorans KAH6567989.1, KAH6579126.1 and B. dendrobatidis EGF83417.1 sequences group together with the animal C1GLT sequences. We provide all Fringe-like accessions in Supplementary Table S1 despite unresolved specificity questions.</p> <p>&nbsp;</p> </div> <strong>Supplementary File</strong></div> <div>Phylogenetic trees for fungal fucosyltrasferases as a newick file format.<br> <div> <div> <p><strong>&nbsp;Supplementary Table S1.</strong></p> <p>All protein identifiers, protein names and aliases, expression profiles</p> <p>&nbsp;</p> <p>&nbsp;</p> </div> </div> </div> </div> </div>

opencc-by-4.0Sep 2023View details →
zenodo40/100

EBEC-MicroED: Static electron diffraction movies collected at different incident flux on a direct electron detector (DE Apollo) on crystals of (S,S) Jacobsen's salen ligand and Co(II) porphyrin, and diffraction tilt series recorded on the DE Apollo and CetaD detector for the same crystals of Jacobsen's Ligand

<p>This record contains static diffraction movies recorded from crystals of (S,S) Jacobsen's salen ligand, and crystals of Co(II) meso-tetraphenyl porphyrin, using a direct electron detector (DE Apollo) in counting mode. Data were acquired at varying different incident flux settings, referred to as "spotsize11" or "spot11" (0.01 electrons per square Angstoms per second),&nbsp; "spotsize10" or "spot10" (0.03 electrons per square Angstrom per second),&nbsp; "spotsize9" or "spot9" (0.045 electrons per square angstrom per second)", and "spotsize8" or "spot8" (0.084 electrons per sqaure Angstrom per second. For each compound these trials, the same crystal ("crystal1", "crystal2", etc.) was conserved across a dose series, and illuminated at each incident flux from lowest to highest in sequence.</p> <p>Additionally, this record contains diffraction tilt series acquired from crystals of (S,S) Jacobsen's ligand, first on the Ceta D and next on the DE Apollo, rotating at 2 degrees per second with an incident flux of either 0.01 or 0.045 electrons per square Angstrom per second.</p> <p>All data is saved in mrc file format, with the exception of movies from the Ceta D, which are saved in ser file format.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

EBEC-MicroED: Electron diffraction tilt series recorded on crystals of (S,S) Jacobsen's salen ligand at 293 K using a direct electron detector (DE Apollo)

<p>This record contains diffraction tilt series recorded from crystals of (S,S) Jacobsen's salen ligand using a direct electron detector (DE Apollo) in counting mode. Incident flux and stage rotation rate are varied, for a total of three data collection settings giving variable total electron beam fluence. The keywords "fast" and "slow" in the titles of the dataset file indicates that the stage was rotated at either 2 degrees/second, or 0.33 degrees/second, respectively. Data were additionally acquired at two different incident flu settings, referred to as "spotsize11" or "spot11" (0.01 electrons per square Angstoms per second) and "spotsize9" or "spot9" (0.045 electrons per square angstrom per second)"</p> <p>All data is saved in mrc file format.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Data for manuscript: The Conformational Space of the SARS-CoV-2 Main Protease Active Site Loops is Determined by Ligand Binding and Interprotomer Allostery

<div>The data is provided as a part of the manuscript "<strong>The Conformational Space of the SARS-CoV-2 Main Protease Active Site Loops is Determined by Ligand Binding and Interprotomer Allostery</strong>".&nbsp; This repository includes an archive with folders:</div> <div>&nbsp;</div> <div><strong>md_data&nbsp;</strong></div> <div> <ul> <li>a directory with MD data for all simulation systems considered in the manuscript. Initial and final conformations are provided.</li> </ul> </div> <div>&nbsp;</div> <div><strong>fig_data</strong></div> <div> <ul> <li>a directory with the data underlying all the main text in the manuscript.&nbsp;</li> </ul> </div> <div>&nbsp;</div> <div>Videos S1-S3 are also included.</div>

opencc-by-4.0Sep 2024View details →
zenodo40/100

AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes

<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span>&nbsp;</span>alanine<span>&nbsp; </span>sequence masking<span>&nbsp; </span>with shallow multiple sequence alignment<span>&nbsp; </span>subsampling to expand the conformational diversity of the predicted structural<span>&nbsp; </span>ensembles and<span>&nbsp;&nbsp; </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span>&nbsp; </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span>&nbsp; </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span>&nbsp; </span>the view </span><span>&nbsp;</span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span>&nbsp; </span>a significant <span>correlation </span>between conformational flexibility and <span>&nbsp;</span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span>&nbsp; </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span>&nbsp;</span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span>&nbsp; </span></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Dataset for publication Electron-induced ligand loss from iron tetracarbonyl methyl acrylate

<p>Experimental dataset for the publication including readme files with description and all metadata.</p> <p><a href="https://doi.org/10.3762/bjnano.15.66">https://doi.org/10.3762/bjnano.15.66</a></p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Binding constants of clinical drugs and other organic ligands with human and mammalian serum albumins

<p>The dataset contains literature values of the experimental equilibrium binding constants of drugs and some other organic ligands with human and mammalian (predominantly bovine) serum albumins. There are 1755 records gathered from 346 original literature sources describing the albumin affinity of 324 different substances. The data were extracted from both articles and existing protein binding databases applying strict data selection rules in order to exclude the values influenced by the third-party compounds. For each experiment with a particular ligand&ndash;albumin system found in literature we provide (if possible) the following details: albumin source organism, ligand chemical name, canonical SMILES, InChIKey, the binding (either association <em>K<sub>a</sub></em> or dissociation <em>K<sub>d</sub></em>) constant value in molarity-based scale, temperature in K, albumin and ligand concentrations, buffer pH, composition, and concentration, experimental method, model used for the binding constant calculation, DOI or link to the source paper. For the results obtained using independent binding sites model, the average number of binding sites <em>n</em> is also provided. If two different types of binding sites were suggested, we put the second site-specific binding constant <em>K<sub>a2</sub></em> or <em>K<sub>d2</sub></em> and the number of the second-type sites<em> </em> <em>n</em><sub>2</sub> into separate columns. For some systems, the enthalpies of binding have also been determined either from the temperature dependence of the binding constant or using direct calorimetric measurements. Their values are also given in the respective column. The dataset can be used as the reference one, for the development of predictive models to calculate the binding constants, and for the choice of the experimental setup in the future albumin binding studies.</p>

opencc-by-4.0Apr 2021View details →
zenodo40/100

Data for article "Assembling diuranium complexes in different states of charge with a bridging redox-active ligand"

<p>This upload contains raw data (NMR, X-Ray, Elemental Analysis, AC &amp; DC SQUID, EPR) files for the article.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

FRidge GA13 Organic Ligands CSV data

<p>Competitive Ligand Exchange-Adsorptive Cathodic Stripping Voltammetry samples from 11 sites along the Mid-Atlantic Ridge. Samples were collected as part of the 2017-18 FRidge (GA 13) research expedition on board the RSS James Cook.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

scPDB BO1 subset (protein ligand-binding sites)

<p>The BO1 subset of the scPDB database.</p> <p>BO1 consists in 766 pairs of non-redundant binding-sites<br> (383 similar pairs, 383 dissimilar pairs).</p> <p>BO1 was describbed in:<br> ---<br> Eguida, M., &amp; Rognan, D. (2020).<br> A computer vision approach to align and compare protein cavities:<br> application to fragment-based drug design.<br> Journal of Medicinal Chemistry, 63(13), 7127-7142.<br> https://doi.org/10.1021/acs.jmedchem.0c00422<br> ---</p> <p>The scPDB was recently describbed in:<br> ---<br> Desaphy, J., Bret, G., Rognan, D., &amp; Kellenberger, E. (2015).<br> sc-PDB: a 3D-database of ligandable binding sites&mdash;10 years on.<br> Nucleic acids research, 43(D1), D399-D404.<br> https://doi.org/10.1093/nar/gku928<br> ---</p> <p>The scPDB is available at:<br> http://bioinfo-pharma.u-strasbg.fr/scPDB/<br> &nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Vertex balanced dataset (protein ligand-binding sites)

<p>Vertex balanced dataset.</p> <p>This dataset consists in 676 pairs of binding-sites<br> (338 similar pairs, 338 dissimilar pairs).</p> <p>The original Vertex dataset was describbed in (but is not downloadable from):<br> ---<br> Chen, Y. C., Tolbert, R., Aronov, A. M., McGaughey, G., Walters, W. P., &amp; Meireles, L. (2016).<br> Prediction of protein pairs sharing common active ligands using protein sequence,<br> structure, and ligand similarity.<br> Journal of chemical information and modeling, 56(9), 1734-1745.<br> https://doi.org/10.1021/acs.jcim.6b00118<br> ---</p> <p>Its balanced version was recently describbed in (but is not downloadable from):<br> ---<br> Eguida, M., &amp; Rognan, D. (2020).<br> A computer vision approach to align and compare protein cavities:<br> application to fragment-based drug design.<br> Journal of Medicinal Chemistry, 63(13), 7127-7142.<br> https://doi.org/10.1021/acs.jmedchem.0c00422<br> ---<br> &nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

PLBD (Protein Ligand Binding Database) table description XML file

<p>PLBD (Protein Ligand Binding Database) table description XML file<br> =================================================================</p> <p>General<br> -------</p> <p>The provided ZIP archive contains an XML file &quot;main-database-description.xml&quot; with the description of all tables (VIEWS) that are exposed publicly at the PLBD server (https://plbd.org/). In the XML file, all columns of the visible tables are described, specifying their SQL types, measurement units, semantics, calculation formulae, SQL statements that can be used to generate values in these columns, and publications of the formulae derivations.</p> <p>The XML file conforms to the published XSD schema created for descriptions of relational databases for specifications of scientific measurement data. The XSD schema (&quot;relational-database_v2.0.0-rc.18.xsd&quot;) and all included sub-schemas are provided in the same archive for convenience. All XSD schemas are validated against the &quot;XMLSchema.xsd&quot; schema from the W3C consortium.</p> <p>The ZIP file contains the excerpt from the files hosted in the https://plbd.org/ at the moment of submission of the PLBD database in the Scientific Data journal, and is provided to conform the journal policies. The current data and schemas should be fetched from the published URIs:</p> <p>&nbsp;&nbsp; &nbsp;https://plbd.org/<br> &nbsp;&nbsp; &nbsp;https://plbd.org/doc/db/schemas<br> &nbsp;&nbsp; &nbsp;https://plbd.org/doc/xml/schemas</p> <p>Software that is used to generate SQL schemas, RestfulDB metadata and the RestfulDB middleware that allows to publish the databases generated from the XML description on the Web are available at public Subversion repositories:</p> <p>&nbsp;&nbsp; &nbsp;svn://www.crystallography.net/solsa-database-scripts<br> &nbsp;&nbsp; &nbsp;svn://saulius-grazulis.lt/restfuldb</p> <p>Usage<br> -----</p> <p>The unpacked ZIP file will create the &quot;db/&quot; directory with the tree layout given below. In addition to the database description file &quot;main-database-description.xml&quot;, all XSD schemas necessary for validation of the XML file are provided. On a GNU/Linux operating system with a GNU Make package installed, the XML file validity can be checked by unpacking the ZIP file, entering the unpacked directory, and running &#39;make distclean; make&#39;. For example, on a Linux Mint distribution, the following commands should work:</p> <p>&nbsp;&nbsp; &nbsp;unzip main-database-description.zip<br> &nbsp;&nbsp; &nbsp;cd db/release/v0.10.0/tables/<br> &nbsp;&nbsp; &nbsp;sh -x dependencies/Linuxmint-20.1/install.sh<br> &nbsp;&nbsp; &nbsp;make distclean<br> &nbsp;&nbsp; &nbsp;make</p> <p>If necessary, additional packages can be installed using the &#39;install.sh&#39; script in the &#39;dependencies/&#39; subdirectory corresponding to your operating system. As of the moment of writing, Debian-10 and Linuxmint-20.1 OSes are supported out of the box; similar OSes might work with the same &#39;install.sh&#39; scripts. The installation scripts require to run package installation command under system administrator privileges, but they use *only* the standard system package manager, thus they should not put your system at risk. For validation and syntax checking, the &#39;rxp&#39; and &#39;xmllint&#39; programs are used.</p> <p>The log files provided in the &quot;outputs/validation&quot; subdirectory contain validation logs obtained on the system where the XML files were last checked and should indicate validity of the provided XML file against the references schemas.</p> <p>Layout of the archived file tree<br> --------------------------------</p> <p>&nbsp;&nbsp; &nbsp;db/<br> &nbsp;&nbsp; &nbsp;└── release<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; └── v0.10.0<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── tables<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── Makeconfig-validate-xml<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── Makefile<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── Makelocal-validate-xml<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── dependencies<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── main-database-description.xml<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── outputs<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── schema</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Example run script for ligand dissociation in wepy

<p>This provides an example script for initializing a wepy script with files from CHARMM-GUI.&nbsp; The wepy software and installation instructions can be found here:&nbsp;https://github.com/ADicksonLab/wepy .</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Speciation and Structures in Pt Surface Sites Stabilized by N-Heterocyclic Carbene Ligands Revealed by DNP Enhanced Indirect-ly Detected 195Pt NMR Spectroscopic Signatures and Fingerprint Analysis

<p>Raw NMR data for the paper published under DOI: 10.1021/jacs.2c08300</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Protein-ligand interactions between mAMCase and chitin.

<p>This directory contains all&nbsp;files to analyze mouse AMCase-ligand interactions as presented in <strong>Supplemental&nbsp;Figure 6</strong>&nbsp;in&nbsp;the manuscript <a href="https://www.biorxiv.org/content/10.1101/2023.06.03.542675">D&iacute;az et al.<em>&nbsp;</em>(2023)</a>.</p> <p>Structure models were analyzed in PyMOL. Figures were compiled using Adobe Illustrator.</p> <p>&nbsp;</p> <p>Contact:</p> <p>Roberto Efra&iacute;n D&iacute;az, robertoefrain.diaz@ucsf.edu</p> <p>James Fraser, jfraser@fraserlab.com</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Asp138 orientation correlates with ligand subsite occupancy.

<p>This directory contains all files required to plot Asp138 orientation against ligand subsite occupancy across all&nbsp;mouse AMCase&nbsp;structures presented in <strong>Figure 3</strong> of&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2023.06.03.542675">D&iacute;az et al.<em> </em>(2023)</a>.</p> <p>Data was analyzed using Graphpad Prism. Figures were compiled using Adobe Illustrator.</p> <p>&nbsp;</p> <p><strong>Figures</strong></p> <p>- contains PDFs&nbsp;of the occupancy at sugar-binding subsite across all mouse AMCase models presented in D&iacute;az et al., (2023).</p> <p>&nbsp;</p> <p><strong>Occupancy.pzfx&nbsp;</strong></p> <p>- contains raw data and plots of ligand occupancy, and Asp138 <em>active</em>&nbsp;conformation occupancy.</p> <p>&nbsp;</p> <p>Contact:<br> Roberto Efra&iacute;n D&iacute;az, robertoefrain.diaz@ucsf.edu</p> <p>James Fraser, jfraser@fraserlab.com</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Computational data on binding of a pyrene-based fluorescent amyloid ligand (Py1SA) to transthyretin (TTR)

<p>This repository contains the data and files for the computational study on binding of a pyrene-based fluorescent amyloid ligand (Py1SA) to transthyretin (TTR).&nbsp;<br> The repository is organized into different folders as described below:</p> <p>====================================================================================================<br> 1_Starting_structure<br> This folder contains the starting structures of the four binding modes obtained from the crystallographic study. These structures serve as the initial configurations for the Molecular Dynamics (MD) simulations.</p> <p>====================================================================================================<br> 2_mdp_files<br> This folder contains two subfolders:</p> <p>1_MD<br> The &quot;1_MD&quot; subfolder includes the MD simulation files in the mdp format. These files define the parameters and settings for running the MD simulations with Gromacs version 2019.3.</p> <p>2_US<br> The &quot;2_US&quot; subfolder includes the files required for performing Umbrella Sampling (US) simulations for each of the binding modes using Gromacs version 2021.3. Within each mode folder, you will find the following files:</p> <p>constraint: Position restraints files used in the Umbrella Sampling simulations.<br> pull: mdp files containing the settings for pulling in the US simulations.<br> us: mdp files used for the US simulations.</p> <p>====================================================================================================<br> 3_MD_results<br> This folder contains the results of the MD simulations. It includes the structure files (.gro) and trajectory files (.xtc) for each simulation. Due to the large size of the files, the solvent has been excluded, and the results are provided at every 1 nanosecond (ns) interval.</p> <p>====================================================================================================<br> 4_US_results<br> The &quot;4_US_results&quot; folder includes the trajectories of the US simulations for each of the binding modes. For each mode, two trajectories are provided.</p> <p>====================================================================================================<br> 5_pdb<br> This folder includes the pdb files of the simulated structures of the two binding modes after the equilibration step.</p> <p>====================================================================================================</p> <p>We acknowledge funding by the German Research Foundation (DFG) through the Emmy Noether Young Group Leader Programme (CK, project KO 5423/1-1), the Swedish e-Science Research Centre (SeRC, ML, PN), the Swedish Research Council (PN, Grant No. 2018-4343). Computing resources were provided by the Swedish National Infrastructure for Computing (SNIC).</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Automated benchmarking of combined protein structure and ligand conformation prediction

<p>The prediction of protein-ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein-Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three-dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time-split test-set is inappropriate for comprehensive PLC evaluation, with state-of-the-art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC-prediction-specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein-ligand interactions.</p> <p>This repository contains:</p> <ol> <li> <p>all_validation_clustering_data.tsv - X-ray validation data and MMSeqs cluster identifiers at different sequence identities for over a million small molecule and ion-binding pockets in the PDB.&nbsp;</p> </li> <li> <p>hqr_dataset.tsv - PDB IDs and ligand information for the high quality representative (HQR) dataset described in the manuscript</p> </li> <li> <p>score_files.tar.gz - Full docking results for all detected pockets for the PDBBind time-split test-set, the HQR dataset, and the subsets of AF models created for both datasets. One file per tool benchmarked with the following columns: Tool, Complex, Pocket, Rank, lDDT-PLI, lDDT-LP, BiSyRMSD, Reference_Ligand, Tool-generated Score</p> </li> <li> <p>errors_all_sets.csv - Report of failures running the pipeline with the following columns: Process, Complex/Ligand/Receptor, Problem</p> </li> </ol>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record