Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,863
datasets available to search
ShareScore release 0.7.1
Dataset results
2,863 results for “Challenges”
Cadenza Challenge (CAD2): databases for rebalancing classical music task
<h1>Cadenza</h1> <p>Please, cite CadenzaWoodwind as</p> <blockquote> <p><strong>Gerardo Roa-Dabike , Trevor J. Cox , Alex J. Miller , Bruno M. Fazenda , Simone Graetzer , Rebecca R. Vos , Michael A. Akeroyd , Jennifer Firth , William M. Whitmer , Scott Bannister , Alinka Greasley , Jon P. Barker , The Cadenza Woodwind Dataset: Synthesised Quartets for Music Information Retrieval and Machine Learning, Data in Brief (2024), doi: https://doi.org/10.1016/j.dib.2024.111199</strong></p> </blockquote> <p>This is the training and validation data for the rebalancing classic music task from the <a href="https://cadenzachallenge.org/">Second Cadenza Machine Learning Challenge (CAD2).</a></p> <p>The Cadenza Challenges are improving music production and processing for people with a hearing loss. According to The World Health Organization, 430 million people worldwide have a disabling hearing loss. Hearing aid users report several issues when listening to music, including distortion in the bass, difficulties in perceiving the full range of the music, especially high-frequency pitches, and a tendency to miss the impact of quieter parts of compositions [1]. In a pilot study, we found giving listeners sliders to allow them to rebalance different instruments in a classical music ensemble was desirable.</p> <p>Overview of files:</p> <ol> <li>CadenzaWoodwind. Synthesized dataset of small ensembles of woodwind instruments for training and validation.</li> <li>EnsembleSet_Mix_1. A subset of the synthesised <a href="../records/6519024">EnsembleSet [7]</a> for training and validation (Mix_1 render).</li> <li>Real Data for Tuning: <a href="../api/records/12664932/draft/files/Stereo_Reverb_Real_Data_For_Tuning.zip/content" target="_blank" rel="noopener noreferrer">Stereo_Reverb_Real_Data_For_Tuning.zip</a>.</li> <li>metadata.zip contains audiograms, scene details, target gains and compressor settings.</li> </ol> <p>The audio files are in FLAC format in the .zip archives. The json files contain metadata.</p> <p>More details below.</p> <p> </p>
Chemical data accompanying the manuscript "Chromium cycling in redox-stratified basins challenges δ53Cr paleoredox proxy applications" in Geophysical research Letters
<p>Water column and sediment chromium concentration and stable isotope data and ancillary metal data from Lake Cadagno, Switzerland. These data accompany a manuscript by the same authors in Geophysical Research Letters (doi: 10.1029/2022GL099154).</p> <p> </p> <p>The associated CTD data are available in the following Zenodo dataset: Sepúlveda Steiner, O., Carlino, C., Haizmann, E., Roman, S., Wüest, A., & Bouffard, D. (2022). Lake Cadagno 2017 CTD and water quality monitoring [Data set]. Zenodo. <a href="http://doi.org/10.5281/zenodo.7127882">http://doi.org/10.5281/zenodo.7127882</a></p>
CryoEM Maps and Associated Data Submitted to the 2015/2016 EMDataBank Map Challenge
<p>Files and metadata associated with the EMDataBank/Unified Data Resource for 3DEM 2015/2016 Map Challenge hosted at challenges.emdatabank.org are deposited.</p> <p>All members of the Scientific Community--at all levels of experience--were invited to participate as Challengers, and/or as Assessors.</p> <p>Seven benchmark raw image datasets were selected for the challenge. Six are selected from recently described single particle structure determinations with image data collected as multi-frame movies; one is based on simulated (in silico) images. All of the raw image datasets are archived at pdbe.org/empiar.</p> <p>27 Challengers created 66 single particle reconstructions from the targets, and then uploaded their results with associated details. 15 of the reconstructions were calculated using the SDSC Gordon supercomputer.</p> <p>This map challenge was one of two community-wide challenges sponsored by EMDataBank in 2015/2016 to critically evaluate 3DEM methods that are coming into use, with the ultimate goal of developing validation criteria associated with every 3DEM map and map-derived model.</p> <p> </p>
Dataset of "Anomaly Detection in Industrial Networks: Current State, Classification, and Key Challenges"
<p>Industrial networks are adapted to their specific requirements, especially in terms of industrial processes. To ensure sufficient security in these networks, it is necessary to set and use security policies that complement government regulations, recommendations, and relevant security standards. This paper aims to provide an in-depth analysis of the anomalies occurring within the networks and propose a structure for collecting valuable data from the experimental site based on dividing anomalies into three main categories:<br>security, operational, and service anomalies (and regular traffic recognition). We present a proof-of-concept solution/design aggregating data in industrial networks for advanced anomaly classification. Multiple data sources such as industrial communication, sensor data (additional sensors controlling device behavior), and HW status data are used as data sources. A total of three scenarios (using a physical testbed) were implemented, where we achieved an accuracy of 0.8540/0.9972 in advanced anomaly classification.</p>
Supplementary data for the article: Future environmental impacts of metals: a systematic review of impact trends, modelling approaches, and challenges
<p>This repository provides the supplementary data to the paper titled <a href="https://doi.org/10.1016/j.resconrec.2024.107572" target="_blank" rel="noopener"><em>"Future environmental impacts of metals: a systematic review of impact trends, modelling approaches, and challenges"</em></a>, published 2024 in <em>Resources, Conservation and Recycling</em>.</p> <h4><strong>Contents</strong></h4> <p>The repository is split in 3 parts and comprises the following files (more details are provided in the <em>README.md</em>):</p> <p><strong>A_Database of reviewed studies:</strong></p> <ul> <li>contains the detailed review data, meant for readers to use as an overview file to gather studies relevant to them. It also includes an overview of all data sources that the reviewed studies used.</li> </ul> <p><strong>B_Scientific supplement to paper:</strong></p> <ul> <li>Contains all data relevant to the related publication Harpprecht et al. (2024), such as studies screened , FAIR data analysis, or analyzed impact trends.</li> </ul> <p><strong>C_Data for figures in paper:</strong></p> <ul> <li>This file contains all the data for Figures 3, 4 and 5 in tabular form, representing impact trends, scenario variables, scenario modelling approaches and data sources used.</li> </ul> <h4><strong>Summary</strong></h4> <p>These files allow to reproduce the results of our study. In this work, we systematically reviewed studies which assessed future environmental impacts of metal supply chains. Our review yielded 40 publications covering 15 metals: copper, iron, aluminium, nickel, zinc, lead, cobalt, lithium, gold, manganese, neodymium, dysprosium, praseodymium, terbium, and titanium. We evaluated their results regarding future impact trends, and their methods, i.e., modelling approaches, scenario variables, and data sources of scenario variables. We identified 15 scenario variables. The most common variables are background electricity mix, ore grade, recycling shares, demand, and energy efficiency. We identified 229 unique data sources for the reviewed scenario variables.</p> <h4><strong>Related publication</strong></h4> <p>More details on the data and its interpretation as well as the scientific context are provided in the publication itself:</p> <p><a href="https://doi.org/10.1016/j.resconrec.2024.107572" target="_blank" rel="noopener">Harpprecht, C., Miranda Xicotencatl, B., van Nielen, S., van der Meide, M., Li, C. , Li, Z., Tukker, A., Steubing, B. (2024). <em>Future environmental impacts of metals: a systematic review of impact trends, modelling approaches, and challenges.</em> Resources, Conservation and Recycling.</a></p> <h4><strong>Funding </strong></h4> <p>Carina Harpprecht received funding from the Energy Program of the German Aerospace Center in 2022. Zhijie Li received funding from the European Institute of Innovation and Technology (EIT) under the project Valomag (Project No. 14049).</p> <h4><strong>License</strong></h4> <p>CC-BY 4.0 license for DLR (German Aerospace Center)</p>
Bridge2AI Grand Challenge AI-Readiness Evaluation Data Year 2 of 4
<p>This excel workbook and set of radar plots contains current and projected AI-readiness evaluation datasets of four NIH Bridge2AI Program Grand Challenges in Functional Genomics, Clinical Care Informatics, Precision Public Health, and Return to Health (Salutogenesis). These evaluations were collected in late 2024, at the conclusion of Year 2 of the 4-year Bridge2AI program, by Grand Challenge (GC) representatives on the Bridge2AI Standards Working Group, in consultation with their GC leadership team, and will be updated in subsequent years and the program progresses. They assess biomedical AI readiness of existing collected data only. </p>
Supplementary Datasets for the publication "Rousettus aegyptiacus Fruit Bats Do Not Support Productive Replication of Cedar Virus upon Experimental Challenge"
<p>Cedar henipavirus (CedV), which was isolated from the urine of pteropodid bats in Australia, belongs to the genus Henipavirus in the family of Paramyxoviridae. It is closely related to the Hendra virus (HeV) and Nipah virus (NiV), which have been classified at the highest biosafety level (BSL4) due to their high pathogenicity for humans. Meanwhile, CedV is apathogenic for humans and animals. As such, it is often used as a model virus for the highly pathogenic henipaviruses HeV and NiV. In this study, we challenged eight Rousettus aegyptiacus fruit bats of different age groups with CedV in order to assess their age-dependent susceptibility to a CedV infection. Upon intranasal inoculation, none of the animals developed clinical signs, and only trace amounts of viral RNA were detectable at 2 days post-inoculation in the upper respiratory tract and the kidney as well as in oral and anal swab samples. Continuous monitoring of the body temperature and locomotion activity of four animals, however, indicated minor alterations in the challenged animals, which would have remained unnoticed otherwise.</p>
Supplementary Datasets for the publication "Increased Susceptibility of Rousettus aegyptiacus Bats to Respiratory SARS-CoV-2 Challenge Despite Its Distinct Tropism for Gut Epithelia in Bats"
<p>Increasing evidence suggests bats are the ancestral hosts of the majority of coronaviruses. In gen-eral, coronaviruses primarily target the gastrointestinal system, while some strains, especially Be-tacoronaviruses with the most relevant representatives SARS-CoV, MERS-CoV, and SARS-CoV-2, also cause severe respiratory disease in humans and other mammals. We previously reported the susceptibility of Rousettus aegyptiacus (Egyptian fruit bats) to intranasal SARS-CoV-2 infection. Here, we compared their permissiveness to an oral infection versus respiratory challenge (in-tranasal or orotracheal) by assessing virus shedding, host immune responses, tissue-specific pa-thology, and physiological parameters. While respiratory challenge with a moderate infection dose of 1 × 104 TCID50 caused a systemic infection with oral and nasal shedding of replica-tion-competent virus, the oral challenge only induced nasal shedding of low levels of viral RNA. Even after a challenge with a higher infection dose of 1 × 106 TCID50, no replication-competent vi-rus was detectable in any of the samples of the orally challenged bats. We postulate that SARS-CoV-2 is inactivated by HCl and digested by pepsin in the stomach of R. aegyptiacus, thereby decreasing the efficiency of an oral infection. Therefore, fecal shedding of RNA seems to depend on systemic dissemination upon respiratory infection. These findings may influence our general understanding of the pathophysiology of coronavirus infections in bats.</p>
New Challenges in Point Cloud Visual Quality Assessment: A Systematic Review (Dataset)
<p>This dataset is a collection of annotated information on the scientific papers screened and analyzed for the systematic review of the literature in Point Cloud Visual Quality Assessment. </p> <p>The data is structured as follows:</p> <ul> <li>General information <ul> <li>Document title</li> <li>Authors</li> <li>Year of publication</li> <li>Venue (Conference or Journal title)</li> <li>Citations (number)</li> <li>URL/DOI</li> </ul> </li> </ul> <ul> <li>About the content <br> <ul> <li>Content Type: Point clouds (PC), Colored Point clouds (CPC), Meshes, Dynamic Point Clouds (DPC)</li> <li>Content source: Source of the content used in a subjective QA test or the evaluation of one or more QA metrics</li> </ul> </li> </ul> <ul> <li>About metric benchmarks <ul> <li>Subjective Ground-truth Data: Dataset(s) Source of the subjective scores used as ground-truth in a QA metric benchmark</li> <li>Assessed Metrics: Types of metrics assessed in a benchmark (JPEG standards, IQM, NR, State-of-the-art, others)</li> <li>Performance Measures: PLCC, SROCC, KRCC, RMSE, OR, others</li> </ul> </li> </ul> <ul> <li>About Objective QA metrics <ul> <li>Metric: Name given to the metric introduced in this paper</li> <li>Base: 3D-based or Projection-based</li> <li>Categories: Categories that characterize the approach of the proposed metric (Feature-based, Learning-Based, Perceptual-based, IQM, others) </li> <li>Reference: Full-Reference (FR), Reduced-Reference (RR) or No-Reference (NR)</li> </ul> </li> </ul> <ul> <li>About Subjective QA experiments <ul> <li>Display: Type of display (2D, 3D, AR, MR, VR) and interaction approach (passive, interactive, 3DoF, 6DoF) used in the described experiment.</li> <li>Rendering: Type of rendering used to display the stimuli (Points, Squares, Cubes, Surface)</li> <li>Lab/Remote: The experiment was run in one or more lab environments, or remotely (Lab, Cross-Lab, Remote)</li> <li>Rating: Subjective rating methodology used in the experiment (ACR, DSIS, PWC, others)</li> <li>Dataset: Name of the new subjective dataset if the experiment's results were published.</li> <li>Observers: Number of observers </li> <li>Distortion type: Types of distortions applied to the stimuli and assessed in the experiment</li> </ul> </li> </ul>
Cadenza Challenge (CAD2): databases for lyric intelligibility task
<h2>Cadenza</h2> <p>This is the training and validation data for the lyric intelligibility task from the <a href="https://cadenzachallenge.org/">Second Cadenza Machine Learning Challenge (CAD2).</a></p> <p>The Cadenza Challenges are improving music production and processing for people with a hearing loss. According to The World Health Organization, 430 million people worldwide have a disabling hearing loss. Studies show that not being able to understand lyrics is an important problem to tackle for those with hearing loss. Consequently, this task is about improving the intelligibility of lyrics when listening to pop/rock over headphones. But this needs to be done without losing too much audio quality - you can't improve intelligibility just by turning off the rest of the band! We will be using one metric for intelligibility and another metric for audio quality, and giving you different targets to explore the balance between these metrics.</p> <p>Please see the <a href="https://cadenzachallenge.org/">Cadenza website</a> for a full description of the data</p>
Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study
<p>Efficiency is essential to support responsiveness w.r.t. ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code that supports symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development tends to produce DL code that is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, less error-prone imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. While hybrid approaches aim for the "best of both worlds," the challenges in applying them in the real world are largely unknown. We conduct a data-driven analysis of challenges—and resultant bugs—involved in writing reliable yet performant imperative DL code by studying 250 open-source projects, consisting of 19.7 MLOC, along with 470 and 446 manually examined code patches and bug reports, respectively. The results indicate that hybridization: (i) is prone to API misuse, (ii) can result in performance degradation—the opposite of its intention, and (iii) has limited application due to execution mode incompatibility. We put forth several recommendations, best practices, and anti-patterns for effectively hybridizing imperative DL code, potentially benefiting DL practitioners, API designers, tool developers, and educators.</p>
Tractography Challenge ISMRM 2015 b=3000s/mm² Data.
<p>This archive contains a version of the ISMRM 2015 Tractography Challenge data simulated with a b-value of 3000s/mm², 64 non-zero gradient directions and 1 baseline volume. The other simulation parameters are the same as for the original phantom.</p>
Unblinded Data for PLAsTiCC Classification Challenge
<p>For classification challenge (https://plasticc.org), the unblinded data files are included here. See PDF note above for more information. The original challenge (Sep 28, 2018 - Dec 17, 2018) was hosted at https://www.kaggle.com/c/PLAsTiCC-2018 .</p>
pKaDatabase for Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge
<p>A curated a database of small molecules with experimentally measured pKa values. </p> <p>This pickle file can be loaded into memory using Pandas. In the code block below we will print out the columns of the DataFrame:</p> <pre><code class="language-python">import pandas as pd df = pd.load("pKaDatabase.pkl") print(df.keys()). # print the columns</code></pre> <blockquote> <p>['deprotonated microstate ID', 'protonated microstate ID', 'deprotonated microstate smiles', 'protonated microstate smiles', 'AM1BCC partial charge (prot. atom)', 'AM1BCC partial charge (deprot. atom)', 'AM1BCC partial charge (prot. atoms 1 bond away)', 'AM1BCC partial charge (deprot. atoms 1 bond away)', 'AM1BCC partial charge (prot. atoms 2 bond away)', 'AM1BCC partial charge (deprot. atoms 2 bond away)', 'Gasteiger partial charge (prot. atom)', 'Gasteiger partial charge (deprot. atom)', 'Gasteiger partial charge (prot. atoms 1 bond away)', 'Gasteiger partial charge (deprot. atoms 1 bond away)', 'Gasteiger partial charge (prot. atoms 2 bond away)', 'Gasteiger partial charge (deprot. atoms 2 bond away)', 'Extented Hückel partial charge (prot. atom)', 'Extented Hückel partial charge (deprot. atom)', 'Extented Hückel partial charge (prot. atoms 1 bond away)', 'Extented Hückel partial charge (deprot. atoms 1 bond away)', 'Extented Hückel partial charge (prot. atoms 2 bond away)', 'Extented Hückel partial charge (deprot. atoms 2 bond away)', '∆G_solv (kJ/mol) (prot-deprot)', 'SASA (Shrake)', 'SASA (Lee)', 'Bond Order', 'Change in Enthalpy (kJ/mol) (prot-deprot)', 'pKa','href', 'num ionizable groups', 'Weight', 'pKa source']</p> </blockquote> <p> </p> <p>For more information regarding feature calculations, please read the following paper:</p> <blockquote> <p>Raddi, Robert, and Vincent Voelz. "Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge." (2021). <a href="https://doi.org/10.26434/chemrxiv.14650302.v1">10.26434/chemrxiv.14650302.v1</a></p> </blockquote>
Cultural heritage adaptive reuse in Salerno: challenges and solutions. Dataset
<p>Dataset analysed in Pintossi, N., Ikiz Kaya, D., Pereira Roders, A. (2023). Cultural heritage adaptive reuse in Salerno: Challenges and solutions. City, Culture and Society, 100505. https://doi.org/10.1016/j.ccs.2023.100505</p> <ul> <li>Date of data collection: 27/11/2018</li> <li>Geographic location of data collection: Salerno, Italy. The venue of the data collection is <em>Salone dei marmi, Palazzo di Città</em>, via Roma, 84121 Salerno, Italy </li> <li>Activity of data collection: Historic Urban Landscape workshop 2 - Salerno. Held in Salerno, Italy, on 26-27/11/2018</li> <li>Aim of data collection: Multi-scale, participatory identification of challenges entailed in the adaptive reuse of cultural heritage and solutions </li> <li>Methods for collection/generation of data: see the methodology section in Pintossi, N., Ikiz Kaya, D., Pereira Roders, A. (2023). Cultural heritage adaptive reuse in Salerno: Challenges and solutions. City, Culture and Society, 100505. https://doi.org/10.1016/j.ccs.2023.100505</li> <li>Researchers facilitating roundtable discussion and writing down paper version of data: Marco Acri, Gaia Daldanise, Gamze Dane, Cristina Garzillo, Antonia Gravagnuolo, Lu Lu, Nadia Pintossi, and Ruba Saleh</li> <li>Researcher translating to English, transcribing data in the digital tabular dataset, and cleaning the data: Nadia Pintossi</li> <li>Original language of the data: English, Italian, and mix of English and Italian</li> </ul>
Arts and humanities in the LTER Network: understanding extent, values, and challenges by assessing the relevance of empathy in the LTER Network, 2013-2014
The Long-term Ecological Research (LTER) Network is a collection of 25 National Science Foundation-funded sites committed to long-term, place-based investigation of the natural world. While activities primarily focus on ecological research, arts and humanities inquiry emerged in 2002 and since then a substantial body of creative work has been produced at LTER-affiliated sites. These art-humanities-science collaborations parallel a wider trend in universities and nonprofits. However, there is little empirical work on the value and effectiveness of this work. After launching a survey in 2013 to assess the values and challenges associated with arts and humanities in the LTER Network, which identified empathy as a meaningful potential outcome of this creative work, we conducted a follow-up analysis to understand: the relevance of empathy in the LTER Network; the role of empathy in bridging arts, humanities, and science collaborations; and the capacity of empathy to connect wider audiences both to LTER science and to the natural world. Our research included phone interviews with representatives from 15 LTER sites and an audience perception survey at an LTER-hosted art show. We found that arts-humanities-science collaborations have great potential to catalyze relationships between scholars, the public, and the natural world; cultivate inspiration and empathy for the natural world; and spark awareness shifts that can enable pro-environmental behavior. Our research demonstrates the potential for art-humanities-science collaborations to facilitate conservation attitudes and action in the Network and beyond.
NRM2018 PET Grand Challenge Dataset
Open the record for dataset details and reuse information.
EEG: Simon Conflict w/ Reinforcement + Cabergoline Challenge
Open the record for dataset details and reuse information.
Compound database and subsets generated by the fragment network for stage 3 of the PHIP2 SAMPL7 Challenge
<p>The fragment network provides a convenient way to filter-out compounds that are dissimilar to the input hit(s). Overall, this search algorithm requires a compound input and 3 parameters: 1- the number of graph traversals (hops), 2- number of changes in heavy atom count (hac), 3- number of changes in ring atoms counts (rac). Please, read the reference (Hall, Murray and Verdonk, 2017) for the specifics of the methods.</p>
The Hurricane Challenge Dataset
<p>The Hurricane Challenge was an international evaluation of intelligibility-enhancing speech modifications that took place in 2013 at Interspeech. The dataset consists of stereo files containing clean speech (channel 1) and noise (channel 2). The purpose of the Challenge is to modify the clean speech only in order to make it more intelligibility in the presence of the corresponding noise sample, without changing root-mean-square level, and within certain constraints on changes in duration. Two different types of noise, stationary speech-shaped noise, and competing speech, are provided, each at three signal-to-noise ratios. The Challenge is described in this <a href="https://www.isca-speech.org/archive/archive_papers/interspeech_2013/i13_3552.pdf">Interspeech article</a>. </p> <p>The dataset consists of two zip files, one for each masker (cs.zip for competing speech, ssn.zip for speech-shaped noise). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.