Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,129
datasets available to search
ShareScore release 0.9.0
Dataset results
2,129 results for “scores”
Nikolai Medtner – Tales (A corpus of annotated scores)
https://dcmlab.github.io/medtner_tales/
The Annotated Mozart Sonatas: Score, Harmony, and Cadence
Harmony, cadence, and phrase annotations for Mozart's 18 piano sonatas
Pyotr Tchaikovsky – The Seasons (A corpus of annotated scores)
No description provided.
Franz Liszt – Années de Pèlerinage (A corpus of annotated scores)
https://dcmlab.github.io/liszt_pelerinage/
Edvard Grieg – Lyric Pieces (A corpus of annotated scores)
https://dcmlab.github.io/grieg_lyric_pieces/
Arcangelo Corelli – Trio Sonatas (A corpus of annotated scores)
<p>mvt 2 ready for review. See two issues at mm. 12 and 14, details in commits.</p>
FunMap prediction scores for all gene pairs
<p>This gzipped TSV (Tab-Separated Values) file contains prediction scores generated by a predictive model trained on different types of feature set. The file structure is as follows:</p><p><strong>Column 1 (Gene Pair):</strong> This column represents pairs of gene ids. Each entry in this column consists of two gene ids, potentially denoting associations between genes.</p><p><strong>Column 2 (RNA Data Prediction):</strong> The values in this column represent prediction scores generated by the model using RNA features as input. </p><p><strong>Column 3 (Protein Data Prediction):</strong> This column contains prediction scores produced by the model when trained on protein features. </p><p><strong>Column 4 (Combined RNA and Protein Data Prediction):</strong> The values in this column represent prediction scores resulting from the model trained with both RNA and protein features.</p>
Host biomarkers and combinatorial scores for the detection of serious and invasive bacterial infection in pediatric patients with fever without source
<h3><span>Background </span></h3> <p><span>Improved tools are required to detect bacterial infection in children with fever without source (FWS), especially when younger than 3 years old. The aim of the present study was to investigate the diagnostic accuracy of a host signature combining for the first time two viral-induced biomarkers, tumor necrosis factor-related apoptosis-inducing ligand (TRAIL) and interferon </span><span>γ</span><span>-induced protein-10 (IP-10), with a bacterial-induced one, C-reactive protein (CRP), to reliably predict bacterial infection in children with fever without source (FWS) and to compare its performance to routine individual biomarkers (CRP, procalcitonin (PCT), white blood cell and absolute neutrophil counts, TRAIL, and IP-10) and to the Labscore.</span></p> <h3><span>Methods</span></h3> <p><span>This was a prospective diagnostic accuracy study conducted in a single tertiary center in children aged less than 3 years old presenting with FWS. Reference standard etiology (bacterial or viral) was assigned by a panel of three independent experts. Diagnostic accuracy (AUC, sensitivity, specificity) of host individual biomarkers and combinatorial scores was evaluated in comparison to reference standard outcomes (expert panel adjudication and microbiological diagnosis). </span></p> <h3><span>Results </span></h3> <p><span>241 patients were included. 68 of them (28%) were diagnosed with a bacterial infection and 5 (2%) with invasive bacterial infection (IBI). Labscore, ImmunoXpert, and CRP attained the highest AUC values for the detection of bacterial infection, respectively 0.854 (0.804–0.905), 0.827 (0.764–0.890), and 0.807 (0.744–0.869). Labscore and ImmunoXpert outperformed the other single biomarkers with higher sensitivity and/or specificity and showed comparable performance to one another although slightly reduced sensitivity in children < 90 days of age.</span></p> <h3><span>Conclusion </span></h3> <p><span>Labscore and ImmunoXpert demonstrate high diagnostic accuracy for safely discriminating bacterial infection in children with FWS aged under and over 90 days, </span><span>supporting their adoption in the assessment of febrile patients. </span></p>
Appendix_verb_scores_UD
<p>A dataset that plots verb scores (1 to 3) in main and adverbial clauses in the sample languages covered by UD.</p>
Forecasts, score summary files, target observational data and meteorological driver files to accompany the manuscript "Skill of process-based forecasts relative to multiple null models varies across time and depth for water temperature and dissolved oxygen"
<p>This data publication includes raw ensemble forecast output (forecasts.zip), as well as summary score files (scores.zip) for process-based forecasts produced with the Forecasting Lake and Reservoir Ecosystems (FLARE) framework. In addition, it includes scores for climatology (climatology_scores.csv) and random walk (RW_scores.csv) null forecasts, formatted observational data of target variables (sunp-targets-insitu.csv), and meteorological driver files required for analysis to accompany the manuscript "Skill of process-based forecasts relative to multiple null models varies across time and depth for water temperature and dissolved oxygen". Forecasts were made of water temperature and dissolved oxygen at Lake Sunapee, NH in 2021 and 2022.</p>
Gene prioritization scores and drug target information
<p>This data folder accompanies the article Sadler MC, Auwerx C, Deelen P, Kutalik Z. Multi-layered genetic approaches to identify approved drug targets. Cell Genomics 3, no. 7 (July 2023): 100341(https://doi.org/10.1016/j.xgen.2023.100341).</p><p>It contains disease-drug-target links as well as gene prioritisation scores of all the assessed methods.</p><p> </p>
ClinVar Annovar annotated file with predictor scores
<p>ClinVar vcf file with annovar annotated predictor scores. Used for downstream parsing and filtering to create relevant datasets for analysis in "<strong>Calibration of variant effect predictors on genome-wide data masks heterogeneous performance across genes". </strong></p>
AlphaFold structures with AlphaMissense scores
<p>These repository provides:</p> <ol> <li>NEW: AFwAM-pdb-qb.tar file including pdb.gz files for human protein structures from the AlphaFoldDb with occupancy column set to residue-wise mean of all a.a. variations and temperature factor column set to residue-wise mean of single nucleotde variations; a PyMOL plugin file (coloram-qb.py) for coloring these structures (<code>coloram column=b</code> or <code>coloram column=q</code>; b is the default)</li> <li>AFwAM-pdb.tar file including pdb.gz files for human protein structures from the <a href="https://alphafold.ebi.ac.uk/">AlphaFoldDb</a> with occupancy and temperature factor columns set to residue-wise mean of <a href="../records/8208688">AlphaMissense</a> scores;</li> <li>A PyMOL plugin file (coloram.py) for coloring these structures;</li> <li>For data, Python scripts, and notebooks, please refer to the pub.zip file; detailed instructions are provided in the README.md within this archive and further explained in our manuscript.</li> </ol> <p><br>For alternative data access, please visit <a href="https://alphamissense.hegelab.org/">https://alphamissense.hegelab.org</a>.</p> <p> </p> <p>Disclaimer: The AlphaMissense Database and other information provided on or linked to this site is for theoretical modelling only, caution should be exercised in use. It is provided "as-is" without any warranty of any kind, whether express or implied. For clarity, no warranty is given that use of the information shall not infringe the rights of any third party (and this disclaimer takes precedence over any contrary provisions in the Google Cloud Platform Terms of Service). The information provided is not intended to be a substitute for professional medical advice, diagnosis, or treatment, and does not constitute medical or other professional advice.</p> <p>Data contained within the AlphaMissense Database is provided for non-commercial research use only under CC BY-NC-SA 4.0 license.</p> <p>DeepMind - AlphaMissense: <a href="https://doi.org/10.1126/science.adg7492">https://doi.org/10.1126/science.adg7492</a></p>
The abilitator summary scale items and summary scores
<p><strong>Objectives</strong></p> <p>According to the Consensus-based Standards for the selection of health Measurement Instruments (COSMIN) panel, structural validity describes how well Patient-Reported Outcome Measures' (PROM) scores reflect the dimensions of the measured construct. Reaching structural validity is important for PROMs that reflect unidimensional effect indicators, but not for formative PROMs in which the items are not necessarily correlated. The main purpose of this study was to examine the structural components of the Abilitator, a co-developed self-report questionnaire on work ability and functioning for the population in a weak labour market position.</p> <p><strong>Methods</strong></p> <p>We examined to what extent the Abilitator has reflective and formative elements in its five summary scales: "C. Inclusion", "D. Mind", "E. Everyday life", "F. Skills", and "G. Body". The Abilitator data sample (n=4555, men 51%, mean age 37 years) was collected in 2017–2022 by the Finnish Institute of Occupational Health in cooperation with the European Social Fund Priority 5 projects in which the participants have multiple challenges to gain employment. For the structural components and validity analysis we implemented both Confirmatory Factor Analysis (CFA) and Exploratory Factor Analysis (EFA).</p> <p><strong>Results</strong></p> <p>Based on the COSMIN criteria for structural validity, the Abilitator reached approximate model fit with CFA when we analysed the different concepts of the questionnaire separately rather than in one unified model. An exception was "E. Everyday life" which was a formative summary scale, and it did not reach approximate fit. EFA showed that the items in the Abilitator's summary scales loaded on ten factors.</p> <p><strong>Conclusions </strong></p> <p>The Abilitator had both reflective and formative elements in its structure. It reached structural validity in those separate concepts that were based on a reflective model. This study revealed interesting connections between different aspects of the Abilitator and produced valuable information for further modification of the questionnaire.</p>
Dataset for Interactive Profiling Narrative (IPN) with Toxicity Tolerance Score of each individual user to different categories of toxicity
<p>he dataset is the collection of the results of the Interactive Profiling Narrative that was developed as a part of the project Listener Aware Content Detoxification.<br>It contains the toxicity tolerance scores of each individual user to different categories.<br>High scores indicate that the user is extremely sensitive to the particular category where as low scores indicate that the user isn't that triggered by the category.Medium scores show moderate tolerance.</p> <p>Columns:<br>1.UserId: The unique id given to each user (helps preserve anonymity).<br>2.Race: The user's toxicity tolerance to the category race. High score means racist comment trigger them. <br>3.Sex: The user's toxicity tolerance to the category sexuality. High scores indicate sensitivity to comments on sexual identity.<br>4.Body_image: The user's sensitivity to comments regarding body image. High scores indicate that remarks on body appearance strongly affect the user.<br>5.Disability:The user's sensitivity to comments about disabilities. A high score means the user is highly sensitive to potentially ableist remarks.<br>6.Religion_culture: The user's sensitivity to content involving religion or culture. High scores show the user is easily triggered by insensitive comments about religious or cultural aspects.<br>7.Physical_abuse: User's sensitivity to comments regarding physical abuse. High scores indicate the user is strongly affected by remarks on physical abuse.<br>8.Mental_health: This measures the user's sensitivity to comments on mental health. High scores suggest that the user is particularly affected by comments stigmatizing mental health issues.<br>9.Politics:The user's sensitivity to political content. High scores indicate that political discussions or comments are likely to evoke a strong reaction in the user</p>
Challenges in adjusting scoring matrices when comparing functional motifs with non-standard compositions
<p>This research was funded by the National Science Centre in Poland (grant number 2021/41/N/ST6/01919)</p>
Dataset: 'Protected Designations of Origin' (PDO) per NUTS-3 regions plus PDO-scores and social-ecological indicators (v2)
<p>The dataset contains the original research data belonging to the research article "EU-wide mapping of ‘Protected Designations of Origin’ food products (PDOs) reveals correlations with social-ecological landscape values" (Flinzberger et al. 2022: <a href="https://doi.org/10.1007/s13593-022-00778-4">https://doi.org/10.1007/s13593-022-00778-4</a>).</p> <p>Based on the mapping of 638 food products registered within the EU as 'Protected Designation of Origin' (PDO), the first excel file <em>"EU28_all_PDO_products_per_NUTS3.xlsx"</em> shows the distribution of PDO-labeled products across the EU28 NUTS-3 regions. All NUTS-3 regions are listed once for each PDO product they contain (as by 30 June 2020).</p> <p>The second excel file <em>"PDO_scores_and_social_ecological_indicator_data.xlsx"</em> includes the accumulated numbers of PDOs per NUTS-3 region (called PDO score) as registered by 30 June 2020. The PDO score is presented for all food products combined, and for four major food categories separately. Further, the excel file contains all aggregated values of social-ecological indicators (including their sources), used for the correlation analysis between PDO scores and social-ecological characteristics in the above-mentioned article.</p> <p>Code explanations (meta data) can be found on the second sheet of each excel file.</p> <p>- - - </p> <p>second file updated to v2</p>
Expanding drug targets for 112 chronic diseases using a machine learning-assisted genetic priority score
<h2>ML-GPS: Machine Learning-Assisted Genetic Priority Score</h2> <p>This Zenodo repository contains data and code associated with the publication:</p> <p>Chen R, Duffy Á, Petrazzini BO, Vy HM, Stein D, Mort M, Park JK, Schlessinger A, Itan Y, Cooper DN, Jordan DM, Rocheleau G, Do R. Expanding drug targets for 112 chronic diseases using a machine learning-assisted genetic priority score. Nat Commun. 2024 Oct 15;15(1):8891. doi: <a href="https://doi.org/10.1038/s41467-024-53333-y">10.1038/s41467-024-53333-y</a>.</p> <h3>Important notes</h3> <ul> <li>You can interactively view the top 10% of ML-GPS predictions without download at <a href="https://rstudio-connect.hpc.mssm.edu/mlgps/">https://rstudio-connect.hpc.mssm.edu/mlgps/</a>.</li> <li>For running Jupyter notebooks, please follow the instructions in the README of the GitHub repository at <a href="https://github.com/robchiral/ML-GPS">https://github.com/robchiral/ML-GPS</a>.</li> </ul> <h3>Repository contents</h3> <p>Files needed to train ML-GPS and ML-GPS DOE:</p> <ul> <li><strong>Files needed for Jupyter notebooks.zip</strong>: Data files required for preprocessing and training.</li> <li><strong>Jupyter notebooks.zip</strong>: Notebooks for cleaning data, training models, and generating predictions.</li> </ul> <h3>Other files:</h3> <ul> <li><strong>Predictions for all gene-phecode pairs.zip</strong>: ML-GPS and ML-GPS DOE scores for all analyzed gene-phecode pairs.</li> <li><strong>Summary statistics.zip</strong>: Genetic association summary statistics for all tested gene-phecode pairs.</li> </ul> <h3>Updated performance metrics</h3> <table> <tbody> <tr> <td><strong>Model</strong></td> <td><strong>Open Targets AUPRC</strong></td> <td><strong>SIDER AUPRC</strong></td> </tr> <tr> <td>ML-GPS (non-DOE)</td> <td>0.074</td> <td>0.080</td> </tr> <tr> <td>ML-GPS DOE (activator predictions)</td> <td>0.029</td> <td>0.042</td> </tr> <tr> <td>ML-GPS DOE (inhibitor predictions)</td> <td>0.067</td> <td>0.064</td> </tr> </tbody> </table> <h3>Zenodo versions</h3> <ul> <li><strong>Version 4: </strong>Updated notebooks and external data to use Open Targets 2024.9; summary statistics are unchanged</li> <li><strong>Version 3: </strong>Corrected error where DOE for rare and ultrarare variants was incorrectly incorporated</li> <li><strong>Version 2: </strong>Original release accompanying the publication</li> </ul>
Big Five Inventory scores for University Students with relevant item scores reversed.
<ul> <li>First Year (n = 97, cohort 6). Sample details, 55% Male, 45% Female; age profile, typically 18-19*; 94% UK, 6% international. </li> <li>Final Year (n = 44, cohort 7): Sample details: 41% Male, 59% Female; age profile 20-35, m = 22.3; 84% UK, 16% international. </li> <li>Masters (n = 51, cohort 8): Sample details: 43% Male, 57% Female; age profile 20-36, m = 24.3; 49% UK, 51% international. </li> </ul>
MACIE scores for human genome assembly GRCh37 Part 1 (Chr1 - Chr3)
<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not damaging protein functional and evolutionarily conserved” (MACIE01); “damaging protein functional and not evolutionarily conserved” (MACIE10); “not damaging protein functional and not evolutionarily conserved” (MACIE00); “both damaging protein functional and evolutionarily conserved” (MACIE11). MACIE_protein is the estimated posterior probability of “damaging protein functional”, which is the sum of MACIE10 and MACIE11; MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “damaging protein functional” or “evolutionarily conserved”, which is the sum of MACIE01, MACIE10, and MACIE11.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not evolutionarily conserved and regulatory functional” (MACIE01); “evolutionarily conserved and not regulatory functional” (MACIE10); “not evolutionarily conserved and not regulatory functional” (MACIE00); “both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of “regulatory functional”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “evolutionarily conserved” or “regulatory functional”, which is the sum of MACIE01, MACIE10, and MACIE11.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.