Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.9.0
Dataset results
56 results for “pKa*”
pKaDatabase for Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge
<p>A curated a database of small molecules with experimentally measured pKa values. </p> <p>This pickle file can be loaded into memory using Pandas. In the code block below we will print out the columns of the DataFrame:</p> <pre><code class="language-python">import pandas as pd df = pd.load("pKaDatabase.pkl") print(df.keys()). # print the columns</code></pre> <blockquote> <p>['deprotonated microstate ID', 'protonated microstate ID', 'deprotonated microstate smiles', 'protonated microstate smiles', 'AM1BCC partial charge (prot. atom)', 'AM1BCC partial charge (deprot. atom)', 'AM1BCC partial charge (prot. atoms 1 bond away)', 'AM1BCC partial charge (deprot. atoms 1 bond away)', 'AM1BCC partial charge (prot. atoms 2 bond away)', 'AM1BCC partial charge (deprot. atoms 2 bond away)', 'Gasteiger partial charge (prot. atom)', 'Gasteiger partial charge (deprot. atom)', 'Gasteiger partial charge (prot. atoms 1 bond away)', 'Gasteiger partial charge (deprot. atoms 1 bond away)', 'Gasteiger partial charge (prot. atoms 2 bond away)', 'Gasteiger partial charge (deprot. atoms 2 bond away)', 'Extented Hückel partial charge (prot. atom)', 'Extented Hückel partial charge (deprot. atom)', 'Extented Hückel partial charge (prot. atoms 1 bond away)', 'Extented Hückel partial charge (deprot. atoms 1 bond away)', 'Extented Hückel partial charge (prot. atoms 2 bond away)', 'Extented Hückel partial charge (deprot. atoms 2 bond away)', '∆G_solv (kJ/mol) (prot-deprot)', 'SASA (Shrake)', 'SASA (Lee)', 'Bond Order', 'Change in Enthalpy (kJ/mol) (prot-deprot)', 'pKa','href', 'num ionizable groups', 'Weight', 'pKa source']</p> </blockquote> <p> </p> <p>For more information regarding feature calculations, please read the following paper:</p> <blockquote> <p>Raddi, Robert, and Vincent Voelz. "Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge." (2021). <a href="https://doi.org/10.26434/chemrxiv.14650302.v1">10.26434/chemrxiv.14650302.v1</a></p> </blockquote>
Raw Data to "Prediction of acid pKa values in the solvent acetone based on COSMO-RS"
<p>This data is a supplement to the publication entitled "Prediction of Acid pKa Values in the Solvent Acetone based on COSMO-RS" in the Journal of Computational Chemistry (DOI:10.1002/jcc.26864). The data includes initial starting structures as inputs for conformer searches using COSMOconf (versions 2020 and 2021) in combination with TURBOMOLE (version 7.3). The corresponding output files serve as inputs for the calculation of Gibbs free energies using COSMO-RS as provided by COSMOtherm.</p> <p>Additional information on the file structure is given in the README file.</p>
pKa of GH18 chitinases in the inactive and active conformation
<p>This directory contains all files required to plot the theoretical pKa of D1, D2, and E from mouse AMCase, human AMCase, and other GH18 chitinases presented in <strong>Figure 4</strong> and <strong>Supplemental Figure 4 </strong>of <a href="https://www.biorxiv.org/content/10.1101/2023.06.03.542675">Díaz et al.<em> </em>(2023)</a>.</p> <p>Structure models were analyzed using PyMOL. Data was analyzed using Graphpad Prism. Figures were compiled using Adobe Illustrator.</p> <p> </p> <p>Files included in this directory:</p> <p><strong>Figures</strong></p> <p>- contains PDFs of Graphpad plots for the pKa of human AMCase, mouse AMCase, and GH18 chitinases D2 in the <em>active </em>or <em>inactive </em>conformation.</p> <p> </p> <p><strong>PDB2PQR</strong></p> <p>- contains all log files from PDB2PQR sorted by data type (mouse AMCase <em>active</em> and <em>inactive </em>D2, human AMCase <em>active</em> and <em>inactive </em>D2, and GH18 chitinase <em>active</em> and <em>inactive </em>D2).</p> <p> </p> <p><strong>PyMOL</strong></p> <p>- contains PyMOL script and session file for each data type (mouse AMCase <em>active</em> and <em>inactive </em>D2, human AMCase <em>active</em> and <em>inactive </em>D2, and GH18 chitinase <em>active</em> and <em>inactive </em>D2) shown in <strong>Figure 4</strong>.</p> <p> </p> <p><strong>Reference Models</strong></p> <p>- contains all structure models shown in <strong>Figure 4</strong>.</p> <p> </p> <p>Contact:<br> Roberto Efraín Díaz, robertoefrain.diaz@ucsf.edu</p> <p>James Fraser, jfraser@fraserlab.com</p>
Supporting materials for: pKa Prediction in Non-Aqueous Solvents
<p>This repository includes datasets and supplementary materials for the manuscript: "pKa Prediction in Non-Aqueous Solvents" by Jonathan W. Zheng, Emad Al Ibrahim, and William H. Green. <strong>Citations should refer directly to the manuscript:</strong></p> <blockquote> <p>Zheng, J. W., Al Ibrahim, E., Kaljurand, I., Leito, I., & Green, W. H. (2024). pKa Prediction in Non-Aqueous Solvents. <em>Journal of Computational Chemistry (2024), </em>doi:10.1002/jcc.27517</p> </blockquote> <p>This compilation includes the predicted and experimental pKa values for all compounds studied in the corresponding work, as well as .xyz files corresponding to all conformers used in the COSMO-RS calculations. </p> <p>For the <strong>.csv </strong>files, column <code>pKa_exp</code> corresponds to the originally-reported experimental value whereas <code>pKa_OK</code> corresponds to the corrected value.</p> <p>The data from <strong>low_error_solvent_preds.csv</strong> and <strong>high_error_solvent_preds.csv</strong> and <strong>unreliable_solvent_preds.csv</strong> are formatted in part based on their compilation in the source manuscript: <em>Busch, M., Ahlberg, E., Ahlberg, E., & Laasonen, K. (2022). How to Predict the pKa of Any Compound in Any Solvent. ACS Omega, 7(20), 17369-17383.</em></p> <p> </p> <ul> <li>Note on version Aug. 28, 2024: fixed erroneous SMILES for several of the benzoic acids in the test set.</li> <li>Note on version Nov. 18, 2024: updated many of the values in the training and test sets, excluding some values and correcting others. SMILES (and corresponding .xyz files) for two values in the test set were updated. </li> </ul>
Simulations of PKA RIα Homodimer Reveal cAMP-coupled Conformational Dynamics of Each Protomer and the Dimer Interface with Functional Implications
<p>Protein kinase A (PKA) is a ubiquitous cAMP-dependent enzyme in mammalian tissues. The inactive PKA holoenzyme disassociates into a homodimer of regulatory (R) subunits and two active catalytic (C) subunits upon cAMP binding to two tandem domains (termed <span>CBD-A</span> <span>and</span> <span>CBD</span>-<span>B</span>) in R subunits. The release of cAMP facilitates reassociation of R and C subunits<span>,</span> resetting PKA to its basal state. The cAMP-mediated structural changes in the activation-termination cycle <span>remain</span> <span>partially</span> <span>understood</span>. The multimeric states of PKA complicate the issue and are particularly <span>less</span> <span>studied</span>. Therefore, we computationally investigate the conformational dynamics of PKA <span>RI</span>a homodimer in different cAMP-bound states. The absence of cAMP in two CBDs affect differently the <span>conformational</span> <span>dynamics</span> of protomers. Moreover, <span>such</span> disparate <span>responses</span> are extended to the dimer interface <span>constituted</span> <span>by</span> <span>the</span> <span>N</span>-<span>terminal</span> <span>helical</span> <span>sub-domains</span> termed N3A motifs. T<span>he removal of cAMP from CBD-A induces large-scale structure</span> changes <span>of individual </span>R subunits <span>towards the holoenzyme state,</span> consist with previous simulations of a single R subunit. <span>Meanwhile</span> <span>it</span> <span>keeps the structural heterogeneity of the </span>N3A-N3A' <span>dimer interface observed in the fully bound state.</span> By contrast, the removal of cAMP from CBD-B does not affect <span>individual </span>R subunits but alters the conformational space of the N3A-N3A' dimer interface. The cAMP-coupled s<span>tructural</span> <span>changes</span> <span>of</span> <span>each</span> <span>protomer</span> <span>and</span> conserved conformational space of <span>the</span> N3A-N3A' <span>dimer</span> <span>interface</span> are essential for t<span>he</span> <span>transition</span> <span>between</span> the fully cAMP-bound R<sub>2</sub> homodimer and the R<sub>2</sub>C<sub>2</sub> holoenzyme as suggested by their crystal structures. Our work provides structural insights into <span>the</span> regulatory mechanism of cAMP in PKA signaling. </p>
pKa dataset. Molecule are coded by ECFP4 fingeprint singularly exploited
<p>This dataset has been used for ML modelling (see <a href="https://github.com/Fraunhofer-ITMP/pKa_model_shiny">Fraunhofer-ITMP/pKa_model_shiny: Repository of the FhG-ITMP ADME models for pKa (github.com)</a>)</p>
Evaluation of the pKa's of Quinazoline Derivatives : Usage of Quantum Mechanical Based Descriptors
<p>In this study, several quantum mechanical-based computational approaches have been used in order to propose accurate protocols for predicting the p<em>K<sub>a</sub></em>’s of quinazoline derivatives, which constitute a very important class of natural and synthetic compounds in organic, pharmaceutical, agricultural and medicinal chemistry areas. Linear relationships between the experimental p<em>K<sub>a</sub></em>’s and nine different DFT descriptors (atomic charge on nitrogen atoms (<em>Q</em>(N), ionization energy (<em>I</em>), electron affinity (<em>A</em>), chemical potential (m), hardness (h), electrophilicity index (w), fukui functions (<em>f <sup>+</sup></em>, <em>f <sup>-</sup></em>), condensed dual descriptor (D<em>f</em>) and local hypersoftness (s<sup>(2)</sup>)) were considered. Several DFT methods (a combination of five DFT functionals and two basis sets) in conjunction with two different implicit solvent models were tested, and among them, M06L/6-311++G(d,p) level of theory employing the CPCM solvation model was found to give the strongest correlations between the DFT descriptors and the experimental p<em>K<sub>a</sub></em>’s of the quinazoline derivatives. The calculated atomic charge on N<sub>1</sub> atom (<em>Q</em>(N<sub>1</sub>)) was shown to be the best descriptor to reproduce the experimental p<em>K<sub>a</sub></em>’s (R<sup>2</sup>=0.927), whereas strong correlations were also derived for <em>A</em>, w, m, and Δ<em>f</em>. In the last part, the applicability of isodesmic reaction scheme to the p<em>K<sub>a</sub></em> prediction of quinazoline derivatives was tested, and the calculated <em>A</em> was shown to be a well-established method for the classification of molecules, and thus, for the identification of a suitable reference molecule for the calculations. The QM-based protocols presented in this study will enable fast and accurate high-throughput p<em>K<sub>a</sub></em> predictions of quinazoline derivatives and the relationships derived can be effectively used in data generation for successful machine learning models for p<em>K<sub>a</sub></em> predictions.</p>
Machine learning meets pKa - Datasets
<p>Datesets belonging to the publication:</p> <p>Baltruschat M and Czodrowski P. Machine learning meets pK<sub>a</sub> [version 2; peer review: 2 approved]. <em>F1000Research</em> 2020, <strong>9</strong>(Chem Inf Sci):113 (<a href="https://doi.org/10.12688/f1000research.22090.2">https://doi.org/10.12688/f1000research.22090.2</a>)</p> <p>Corresponding GitHub repository: <a href="https://github.com/czodrowskilab/Machine-learning-meets-pKa">https://github.com/czodrowskilab/Machine-learning-meets-pKa</a></p>
QIZILO'NGACH NUQSONI BILAN TUG'ILGAN CHAQALOQLARDA O'PKA TUZILMALARINING SHIKASTLANISHI
Open the record for dataset details and reuse information.
pKa data for 5000+ compounds from large pharma companies
<p>Three dataset with pKa for 5000+ drug discovery compounds</p>
Spatially compartmentalized phase regulation of a Ca2+-cAMP-PKA oscillatory circuit
Open the record for dataset details and reuse information.
Elevated PKA activity at synapses and broad molecular disturbances in the striatum of Akap11 mutant mice, a genetic model of schizophrenia and bipolar disorder [snRNA-Seq]
GEO Series GSE272098. Mus musculus. 61 samples. Type: Expression profiling by high throughput sequencing.
Bipolar and schizophrenia risk gene AKAP11 encodes an autophagy receptor coupling the regulation of PKA kinase network homeostasis to synaptic transmission
GEO Series GSE304457. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.
Hepatic PKA Mediates the Liver and Pancreatic Alpha-Cell Crosstalk
GEO Series GSE268864. Mus musculus. 90 samples. Type: Expression profiling by high throughput sequencing.
Identification of PKA-dependent signaling network using CRISPR-Cas9 coupled with quantitative transcriptomics, proteomics and phosphoproteomics
GEO Series GSE95008. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.
Activation of P-TEFb by cAMP-PKA signaling in autosomal dominant polycystic kidney disease
GEO Series GSE128524. Mus musculus. 14 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Protein Kinase A (PKA) inhibition in Saccharomyces cerevisiae and Kluyveromyces lactis
GEO Series GSE163741. Saccharomyces cerevisiae; Kluyveromyces lactis. 95 samples. Type: Expression profiling by high throughput sequencing.
CUT&Tag analysis of HDAC1 binding profile with or without PKA inhibitor H89 treatment
GEO Series GSE285017. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
A time-gated PKA–CREB signaling circuit licenses IL-12 responsiveness and Th1 fate in CD4⁺ T cells
GEO Series GSE303408. Mus musculus. 30 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Newly identified Aspergillus nidulans PKA targets play a role in carbon-catabolite repression
GEO Series GSE116579. Aspergillus nidulans. 4 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.