Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

46

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

46 results for “Scoring function”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data for "A learned score function improves the power of mass spectrometry database search"

<div> <h1>DATA for "A learned score function improves the power of mass spectrometry database search"</h1> <br> <div>These data files are associated with the following publication:</div> <br> <div> <ul> <li>Varun Ananth, Justin Sanders, Melih Yilmaz, Sewoong Oh and William Stafford Noble. "<a title="biorXiv Preprint Link" href="https://www.biorxiv.org/content/10.1101/2024.01.26.577425v2" target="_blank" rel="noopener">A learned score function improves the power of mass spectrometry database search</a>". Bioinformatics (Proceedings of the ISMB). &nbsp;2024.</li> </ul> </div> <br> <div>For the benchmarking data, we used a dataset that is publicly available on ProteomeXchange (PXD028735). The paper that introduced this dataset is:</div> <br> <div> <ul> <li>Van Puyvelde, B., Daled, S., Willems, S., Gabriels, R., Gonzalez de Peredo, A., Chaoui, K., Mouton-Barbosa, E., Bouyssi&eacute;, D., Boonen, K., Hughes, C. J., Gethings, L. A., Perez-Riverol, Y., Bloomfield, N., Tate, S., Schiltz, O., Martens, L., Deforce, D., &amp; Dhaenens, M. (2022). A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. In Scientific Data (Vol. 9, Issue 1). Springer Science and Business Media LLC. https://doi.org/10.1038/s41597-022-01216-6</li> </ul> </div> <br> <div>More specifically, the following `.raw` files were downloaded:</div> <br> <ul> <li><code>LFQ_Orbitrap_DDA_Ecoli_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Human_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Yeast_01.raw</code></li> </ul> <br> <div>Those files can be accessed via FTP&nbsp;<a title="Link to ProteomeXchange: PXD028735" href="https://ftp.pride.ebi.ac.uk/pride/data/archive/2022/02/PXD028735/" target="_blank" rel="noopener">here</a>.</div> <br> <div>We upload here the annotated <code>.mgf</code> files created from these <code>.raw</code> files, as described in our paper.</div> <br> <div>The human, yeast, and E. coli .fasta files used in all database searches were downloaded from UniProt on 11/6/23, 4:30 PM.</div> <br> <div> <ul> <li>Bateman, A., Martin, M.-J., Orchard, S., Magrane, M., Ahmad, S., Alpi, E., Bowler-Barnett, E. H., Britto, R., Bye-A-Jee, H., Cukura, A., Denny, P., Dogan, T., Ebenezer, T., Fan, J., Garmiri, P., da Costa Gonzales, L. J., Hatton-Ellis, E., Hussein, A., &hellip; Zhang, J. (2022). UniProt: the Universal Protein Knowledgebase in 2023. In Nucleic Acids Research (Vol. 51, Issue D1, pp. D523&ndash;D531). Oxford University Press (OUP). https://doi.org/10.1093/nar/gkac1052</li> </ul> </div> <br> <div>We include these files here, with only minor modifications to replace `U` amino acids with `X` so that all amino acids fall into Casanovo-DB's vocabulary.</div> </div>

opencc-by-4.0Mar 2024View details →
zenodo40/100

AA-Score: a New Scoring Function Based on Amino Acid Specific Interaction for Molecular Docking

<p>The protein-ligand scoring function plays an important role in computer-aided drug discovery, which is heavily used in virtual screening and lead optimization. In this study, we developed a new empirical protein-ligand scoring function,&nbsp;which is a linear combination of empirical energy components, including hydrogen bond, van der Waals, electrostatic, hydrophobic, &pi;-stacking, &pi;-cation, and metal-ligand interaction. Different from previous empirical scoring functions, AA-Score uses several amino acid-specific empirical interaction components. We tested AA-Score on several test sets. The resulting performance shows AA-Score performs well on scoring, docking, and ranking compared with other widely used traditional scoring functions. Our results suggest that AA-Score gains substantial improvements from using detailed protein-ligand interaction components. Besides, we developed an easy-to-use tool to analyze protein-ligand interaction fingerprint and predict binding affinity using AA-Score.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions

<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schr&ouml;dinger.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Visualization of the numerical pose optimization with the JAMDA scoring function using the BFGS and the LSL-BFGS algorithm

<p>These videos demonstrate the behavior of two different optimization algorithms (BFGS and LSL-BFGS) during pose optimization using the JAMDA protein-ligand scoring function.</p> <p>Flachsenberg et al. (2020) (<a href="http://doi.org/10.1021/acs.jcim.0c01095" target="_blank" rel="noopener">10.1021/acs.jcim.0c01095</a>) describes the JAMDA protein-ligand scoring function and the LSL-BFGS algorithm.<br>The data for these videos stems from Experiment 5 in Flachsenberg et al. (2020). In this experiment, the crystal structure of a ligand was numerically optimized in the binding site with respect to the JAMDA scoring function to create the JAMDA-minimized crystal structure. The JAMDA-minimized crystal structure was randomly deflected to generate various starting poses for the numerical optimization.</p> <p>These videos demonstrate the behavior of two optimization algorithms (BFGS and LSL-BFGS) when optimizing one of the generated starting poses. The chosen example for the videos is a structure of ribonuclease A with a 5'-deoxy-5'-N-piperidinouridine inhibitor (PDB code 3d6q, <a href="https://doi.org/10.2210/pdb3D6Q/pdb" target="_blank" rel="noopener">10.2210/pdb3D6Q/pdb</a>, <a href="https://doi.org/10.1021/jm800724t" target="_blank" rel="noopener">10.1021/jm800724t</a>). Each of the videos shows all the intermediate steps the optimization algorithm takes until convergence.</p> <p>The main observation (that is discussed in detail in Flachsenberg et al. (2020)) is that the BFGS algorithm tends to take inappropriately large steps when clashes are present in the structure, resulting in unwanted binding mode changes. This is <em>not</em> the case for the LSL-BFGS algorithm.</p> <h3>Legend</h3> <p>For each iteration, the JAMDA score value, the RMSD to the JAMDA-minimized crystal structure (yellow), and the RMSD to the optimization's starting point (blue) are given. Furthermore, the gradient's norm (representing the main convergence criterion) is shown. In addition to the optimized ligand, also the JAMDA-minimized crystal structure (yellow) and the optimization's starting structure (blue) are shown.</p> <p><br>Each optimization algorithm is shown in two videos: In one video, the optimized ligand is colored by elements. Here, atoms with clashes (positive JAMDA scores) are marked with orange balls. In the other video, the atoms and bonds of the optimized ligand are colored by their individual JAMDA score.</p> <h3>Used Software</h3> <p>The snapshots of the optimization algorithms were rendered using PyMOL 3.0 (<a href="https://pymol.org/" target="_blank" rel="noopener">https://pymol.org/</a>) and further processed using the Pillow 10.4 Python library (<a href="https://doi.org/10.5281/zenodo.12606429" target="_blank" rel="noopener">10.5281/zenodo.12606429</a>). Videos were created from the individual snapshots using FFmpeg 7.0 (<a href="https://www.ffmpeg.org/" target="_blank" rel="noopener">https://www.ffmpeg.org/</a>).</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Challenges in adjusting scoring matrices when comparing functional motifs with non-standard compositions

<p>This research was funded by the National Science Centre in Poland (grant number 2021/41/N/ST6/01919)</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Functional near-infrared spectroscopy visually evoked measurements and Autism Questionnaire score in adults and children

<p><strong>Functional near-infrared spectroscopy (fNIRS) visually-evoked data collected from healthy adults&nbsp;and children. We recruited a total of 40 adult participants (20 women, age: 31.05 &plusmn; 3.94 (SD) years) and 19 children (5 girls, age: 7.20 &plusmn; 3.01 (SD) years).&nbsp;Adult participants filled in the Autistic-traits Quotient (AQ) questionnaire, a 50-items self-administered report validated for the Italian version.&nbsp;The items consist of descriptive statements assessing personal preferences and typical behavior.&nbsp;Since child self-report might be affected by reading and comprehension difficulties, the children&#39;s version of Autism Spectrum Quotient (Italian version of AQ-child) was completed by parents.&nbsp;To measure changes in total Hb (THb) concentration and relative oxygenation levels (OHb and DHb) in the occipital cortex during the task, we used a continuous-wave NIRS system with&nbsp;8 red light-sources operating at 760 nm and 850 nm, and 7 detectors, forming an array of 22 multi-distant channels.&nbsp;</strong></p> <p>&nbsp;</p> <p><strong>The dataset includes 2 &#39;.zip&#39; files containing artifact-free visually-evoked transients HDF data, exported from Homer 3 (https://github.com/BUNPC/Homer3) in &#39;.txt&#39; format.&nbsp;Subfolders correspond to different stimulation conditions, for details refer to this paper [preprint coming soon].</strong></p> <p><strong>Each txt file is named with the corresponding subject code and&nbsp;contains two events [tagged as 1 (baseline) and 2 (stimulus)].&nbsp;</strong></p> <p><strong>The file &#39;annotations.csv&#39; contains&nbsp;the following fields:</strong></p> <p><strong>Code [Subjects id name, that corresponds to the filename in the recording folders]</strong></p> <p><strong>Sex [M or F]</strong></p> <p><strong>Valid&nbsp; [1 or 0, 0 means excluded subject]</strong></p> <p><strong>Age [int]</strong></p> <p><strong>CH_exclude [list of excluded channels]</strong></p> <p><strong>CH_Red [ list of channels that did not passed the calibration step]</strong></p> <p><strong>Category&nbsp; [adult or child]&nbsp; &nbsp;</strong></p> <p><strong>Cartoon_fixed&nbsp; [cartoon name]</strong></p> <p><strong>Cartoon_chosen&nbsp;&nbsp; &nbsp;[cartoon name]</strong></p> <p><strong>AQ&nbsp; [global AQ score]&nbsp;&nbsp;</strong></p> <p><strong>AQ_S&nbsp; [&nbsp;AQ subscale social ]&nbsp; &nbsp;</strong></p> <p><strong>AQ_C&nbsp;&nbsp;[&nbsp;AQ subscale communication ]&nbsp; &nbsp;</strong></p> <p><strong>AQ_A&nbsp;&nbsp;[&nbsp;AQ subscale attention]&nbsp; &nbsp;</strong></p> <p><strong>AQ_D&nbsp;&nbsp;[&nbsp;AQ subscale detail]&nbsp;</strong></p> <p><strong>AQ_I&nbsp;[&nbsp;AQ subscale imagination]&nbsp;</strong><br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
ClinicalTrials.gov36/100

Trial Evaluating New Strategy in the Functional Assessment of 3-vessel Disease Using SYNTAXII Score in Patients With PCI

ClinicalTrials.gov study NCT02015832. IPD Sharing: Not stated. Countries: 4. Publications: 4.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

The Effect of NAC on Lung Function and CT Mucus Score

ClinicalTrials.gov study NCT03822637. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
zenodo32/100

ToolBoxSF: Robustly interrogating machine learning-based scoring functions: what are they learning?

<p>This Zenodo repository provides comprehensive resources for the pre-print research paper titled "Robustly interrogating machine learning-based scoring functions: what are they learning?" Our collection includes Singularity containers containing pre-trained models, benchmark datasets, and training/test CSV files, offering valuable insights into the inner workings of machine learning-based scoring functions.</p><p>Key Components:</p><p>Singularity Containers:</p><ul><li>Machine Learning Models: Explore state-of-the-art scoring models used in the study, enabling reproducibility and in-depth analysis.</li><li>Environment Setup: Simplify model deployment and experimentation by utilizing our pre-configured environments.</li></ul><p>Benchmark Datasets:</p><ul><li>Curated benchmark datasets used in the pre-print, facilitating validation and evaluation of scoring functions.</li></ul><p>Training and Test CSV Files:</p><ul><li>Training and test data in CSV format, along with associated metadata.</li><li>Facilitate model testing and comparison using the provided data.</li></ul><p>This Zenodo collection is a valuable resource for researchers, data scientists, and machine learning enthusiasts seeking to replicate the study's findings, explore model behaviors, and conduct further investigations into machine learning-based scoring functions. Detailed documentation and usage instructions are included to support your research efforts at <a href="https://github.com/guydurant/toolboxsf">https://github.com/guydurant/toolboxs</a>f.</p><p>Citation Information: Please cite this Zenodo repository when using our resources in your work, and consider acknowledging the original pre-print when publishing research based on these materials.</p>

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov32/100

Assessment of Nebulized Tranexamic Acid in Functional Endoscopic Sinus Surgery Using Modena Bleeding Score

ClinicalTrials.gov study NCT06725732. IPD Sharing: Not stated. Countries: 1. Publications: 4.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Validation of a Kinematic Functional Shoulder Score Including Only Essential Movements

ClinicalTrials.gov study NCT01431417. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Validation of a Score for Shoulder Function Evaluation Based on Movement Analysis

ClinicalTrials.gov study NCT01281085. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Association of Functional Dyspepsia Symptom Diary Score and Other Scores Related to Functional Dyspepsia and Its Severity

ClinicalTrials.gov study NCT04953975. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Pain, Range of Motion, Edema, Sensibility, Strength (PRESS) & Self-reported Function Create a Comprehensive Score

ClinicalTrials.gov study NCT06155617. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
zenodo28/100

Left atrial polygenic scores for "Deep Learning of Left Atrial Structure and Function Provides Link to Atrial Fibrillation Risk"

Open the record for dataset details and reuse information.

opencc-zeroMar 2024View details →
zenodo28/100

Data for "A learned score function improves the power of mass spectrometry database search"

<div>These data files are associated with the following publication:</div> <div> <ul> <li>Varun Ananth, Justin Sanders, Melih Yilmaz, Bo Wen, Sewoong Oh and William Stafford Noble. "<a title="biorXiv Preprint Link" href="https://www.biorxiv.org/content/10.1101/2024.01.26.577425v2" target="_blank" rel="noopener">A learned score function improves the power of mass spectrometry database search</a>". Bioinformatics (Proceedings of the ISMB). &nbsp;2024.</li> </ul> </div> <div>For the benchmarking data, we used a dataset that is publicly available on ProteomeXchange (PXD028735). The paper that introduced this dataset is:</div> <div> <ul> <li>Van Puyvelde, B., Daled, S., Willems, S., Gabriels, R., Gonzalez de Peredo, A., Chaoui, K., Mouton-Barbosa, E., Bouyssi&eacute;, D., Boonen, K., Hughes, C. J., Gethings, L. A., Perez-Riverol, Y., Bloomfield, N., Tate, S., Schiltz, O., Martens, L., Deforce, D., &amp; Dhaenens, M. (2022). A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. In Scientific Data (Vol. 9, Issue 1). Springer Science and Business Media LLC. https://doi.org/10.1038/s41597-022-01216-6</li> </ul> </div> <div>More specifically, the following `.raw` files were downloaded:</div> <ul> <li><code>LFQ_Orbitrap_DDA_Ecoli_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Human_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Yeast_01.raw</code></li> </ul> <div>Those files can be accessed via FTP&nbsp;<a title="Link to ProteomeXchange: PXD028735" href="https://ftp.pride.ebi.ac.uk/pride/data/archive/2022/02/PXD028735/" target="_blank" rel="noopener">here</a>.</div> <div>We upload here the annotated&nbsp;<code>.mgf</code>&nbsp;files created from these&nbsp;<code>.raw</code> files, as described in our paper.</div> <div>The human, yeast, and E. coli .fasta files used in all database searches were downloaded from UniProt on 11/6/23, 4:30 PM.</div> <div> <ul> <li>Bateman, A., Martin, M.-J., Orchard, S., Magrane, M., Ahmad, S., Alpi, E., Bowler-Barnett, E. H., Britto, R., Bye-A-Jee, H., Cukura, A., Denny, P., Dogan, T., Ebenezer, T., Fan, J., Garmiri, P., da Costa Gonzales, L. J., Hatton-Ellis, E., Hussein, A., &hellip; Zhang, J. (2022). UniProt: the Universal Protein Knowledgebase in 2023. In Nucleic Acids Research (Vol. 51, Issue D1, pp. D523&ndash;D531). Oxford University Press (OUP). https://doi.org/10.1093/nar/gkac1052</li> </ul> </div> <div>We include these files here, with only minor modifications to replace <code>U</code> amino acids with <code>X</code> so that all amino acids fall into Casanovo-DB's vocabulary.</div>

openMar 2024View details →
geo24/100

CaSSiDI: Novel single-cell “Cluster Similarity Scoring and Distinction Index” reveals critical functions for PirB and context-dependent Cebpb repression

GEO Series GSE252466. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2024View details →
geo24/100

Functional activity scores of a DMS library representing coding single residue substitution variants in the transcription factor CRX measured in an engineered reporter cell line

GEO Series GSE262060. synthetic construct; Homo sapiens. 18 samples. Type: Other.

openGEO-OpenMar 2024View details →
geo24/100

DNA Repair Function Scores for 2172 Variants in the BRCA1 Amino-Terminus

GEO Series GSE234482. Homo sapiens. 16 samples. Type: Other.

openGEO-OpenJun 2023View details →
zenodo24/100

Beware of the generic machine learning-based scoring functions in structure-based virtual screening

<p>Data sets and the rescoing scores utilized in the paper &quot;Beware of the generic machine learning-based scoring functions in structure-based virtual screening&quot; (DOI: 10.1093/bib/bbaa070)</p>

opencc-by-4.0Jun 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record