Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
96
datasets available to search
ShareScore release 0.7.1
Dataset results
96 results for “Reinforcement Learning”
Offline replay supports planning in human reinforcement learning
Open the record for dataset details and reuse information.
Deep Reinforcement Learning for END-To-END Local Motion Planning of Autonomous Aerial Robots in Unknown Outdoor Environments: Real-Time Flight Experiments
<p> </p> <p>Videos for the real flight tests and the simulation experiments </p>
Reaching the limit in autonomous racing: Optimal control versus reinforcement learning
<p>A central question in robotics is how to design a control system for an agile mobile robot. This paper studies this question systematically, focusing on a challenging setting: autonomous drone racing. We show that a neural network controller trained with reinforcement learning (RL) outperformed optimal control (OC) methods in this setting. We then investigated which fundamental factors have contributed to the success of RL or have limited OC. Our study indicates that the fundamental advantage of RL over OC is not that it optimizes its objective better but that it optimizes a better objective. OC decomposes the problem into planning and control with an explicit intermediate representation, such as a trajectory, that serves as an interface. This decomposition limits the range of behaviors that can be expressed by the controller, leading to inferior control performance when facing unmodeled effects. In contrast, RL can directly optimize a task-level objective and can leverage domain randomization to cope with model uncertainty, allowing the discovery of more robust control responses. Our findings allowed us to push an agile drone to its maximum performance, achieving a peak acceleration greater than 12 times the gravitational acceleration and a peak velocity of 108 kilometers per hour. Our policy achieved superhuman control within minutes of training on a standard workstation. This work presents a milestone in agile robotics and sheds light on the role of RL and OC in robot control.</p>
Supplemental data to AALE 2024 publication "Adaptive manufacturing: dynamic resource allocation using multi-agent reinforcement learning"
<p>Release as supplementary material for our contribution at AALE 2024: "Adaptive manufacturing: dynamic resource allocation using multi-agent reinforcement learning"<br><br>The evaluation datasets stored in this collection are used to compare the performance of multi-agent reinforcement learning. In addition, the performance of other methods such as (meta-) heuristic algorithms or single agent reinforcement learning algorithms or novel methods of search space reduction can also be compared.</p>
Keynote: Bringing Reinforcement learning Into Radio Light Network for Massive Connections
<blockquote><p>3GPP standardization has been progressing at an astonishingly rapid phase, where Release 15 and Release 16 have set the foundations of the 5G system, while Release 17 provides enhancements and optimizations to enable support for further use cases. In parallel to 5G standardization efforts, several initiatives worldwide endeavour to drive and support the evolution of smart networks and services. Among others Europe is establishing the <i>Joint Undertaking on Smart Networks and Services</i> in the frame of the Horizon Europe programme for research and innovation. Other initiatives are complementing the European initiative, such as <i>Secure 5G & Beyond Act</i> in the U.S., <i>roadmap towards 6G </i>in Japan, <i>MSIT 6G programme</i> in S. Korea, and <i>MIIT 6G programme</i> in China.</p></blockquote><blockquote><p>The workshop will provide an opportunity for reflection and discussion about requirements and architectural considerations for future generations of mobile systems. The focus will be on presenting version 4.0 of the Architecture white paper developed by the 5G PPP architecture working group. It will allow to move from 5G and beyond towards a fully-fledged 6G architecture.</p></blockquote>
TensorBoard runs for reinforcement learning on automated conjectured bounds on Laplacian spectral radius of graphs
<p>The zip archive contains separate folders with TensorBoard event files for the runs of our reinforcement learning implementation on the conjectured upper bounds for the Laplacian spectral radius, which are listed in Appendix B of the forthcoming paper: S. Al-Yakoob, M. Ghebleh, A. Kanso, D. Stevanović, Reinforcement learning for graph theory, I. Reimplementation of Wagner’s approach, Art Discrete Appl. Math. (2024).</p> <p>To view the contained graphs and the evolution of rewards, unzip the archive in a folder of your choice, and run in the terminal the command "tensorboard --logdir runs" from the parent folder of the unzipped "runs" folder.</p>
Dataset for "Reinforcement Learning reveals fundamental limits on the mixing of active particles"
<p>Open access data set for manuscript "Reinforcement Learning reveals fundamental limits on the mixing of active particles" currently (2021) in preparation.</p>
The AI Economist: Taxation policy design via two-level deep reinforcement learning
<p>This dataset contains all raw experimental data for the paper "The AI Economist: Taxation Policy Design via Two-level Deep Multi-Agent Reinforcement Learning". </p> <p>The accompanying simulation, reinforcement learning, and data visualization code can be found at https://github.com/salesforce/ai-economist.</p> <p>For the one-step economy experiments, we provide:</p> <ul> <li> <p>training histories,</p> </li> <li> <p>configuration files (these experiments do not use phases), and</p> </li> <li> <p>final agent and planner models.</p> </li> </ul> <p>For the Gather-Trade-Build scenario, the data covers 6 spatial layouts: two Open-Quadrant (with 4 and 10 agents), and four Split-World maps with different configurations of the high-skilled and low-skilled agents. It also covers 4 tax policies (the AI Economist, Saez, free-market, and US federal). In addition, the AI Economist has been optimized for two social welfare functions: the product of equality and productivity, and inverse-income weighted utility. The Saez tax policy also uses estimated elasticities. </p> <p>Each experiment was repeated with different random seeds: 10 seeds for the Open-Quadrant scenarios, and 5 seeds for the Split-World scenarios. For each individual experiment, we provide: </p> <ul> <li> <p>Training histories (e.g. equality and productivity throughout training)</p> </li> <li> <p>the phase 1 and phase 2 configuration files, </p> </li> <li> <p>40 episode dense logs (the final 10 simulation logs across 4 environment replicas),</p> </li> <li> <p>phase 1 final agent models, and</p> </li> <li> <p>phase 2 final agent and planner models.</p> </li> </ul> <p>Finally, we include all data and results used to calibrate the Saez elasticity estimates and to estimate elasticity directly from a sweep over flat-rate tax policies:</p> <ul> <li> <p>training histories,</p> </li> <li> <p>the phase 1 and phase 2 configuration files, </p> </li> <li> <p>phase 1 final agent models, and</p> </li> <li> <p>phase 2 final agent and planner models.</p> </li> </ul>
Deep Reinforcement Learning-based Project Prioritization for Rapid Post-Disaster Recovery of Transportation Infrastructure Systems
<p>Among various natural hazards that threaten transportation infrastructure, flooding represents a major hazard in Region 6's states to roadways as it challenges their design, operation, efficiency, and safety. The catastrophic flooding disaster event generally leads to massive obstruction of traffic, direct damage to highway/bridge structures/pavement, and indirect damages to economic activities and regional communities that may cause loss of many lives. After disasters strike, reconstruction and maintenance of an enormous number of damaged transportation infrastructure systems require each DOT to take extremely expensive and long-term processes. In addition, planning and organizing post-disaster reconstruction and maintenance projects of transportation infrastructures are extremely challenging for each DOT because they entail a massive number and the broad areas of the projects with various considerable factors and multi-objective issues including social, economic, political, and technical factors. Yet, amazingly, a comprehensive, integrated, data-driven approach for organizing and prioritizing post-disaster transportation reconstruction projects remains elusive. In addition, DOTs in Region 6 still need to improve the current practice and systems to robustly identify and accurately predict the detailed factors and their impacts affecting post-disaster transportation recovery. The main objective of this proposed research is to develop a deep reinforcement learning-based project prioritization system for rapid post-disaster reconstruction and recovery of damaged transportation infrastructure systems. This project also aims to provide a means to facilitate the systematic optimization and prioritization of the post-disaster reconstruction and maintenance plan of transportation infrastructure by focusing on social, economic, and technical aspects. The outcomes from this project would help engineers and decision-makers in Region 6's State DOTs optimize and sequence transportation recovery processes at a regional network level with necessary recovery factors and evaluating its long-term impacts after disasters.</p>
RL-MLZerD: Multimeric protein docking using reinforcement learning
<p>This is the dataset used in the publication:<strong> RL-MLZerD: Multimeric protein docking using reinforcement learning by Tunde Aderinwale, Charles Christoffer and Daisuke Kihara</strong>, in submission 2022.</p> <p>The dataset included 30 protein complexes targets from 3 to 5 chains used for docking. Each target in the dataset are randomly rotated and shifted three times to remove bias to the starting conformation.</p> <p>We also include the docking results for each target in the corresponding directory for all the methods reported in the manuscript.</p> <p>For each target, 1500 to 12,000 models were provided that were generated by RL-MLZerD, 1700 to 12,000 models for Multi-LZerD, about 100 to 300 models by Combdock, and 5 models each by Alphafold-Multimer, Alphafold-Multimer (no template), and ColabFold respectively.</p> <p>Directory description:</p> <p>Target: contains all the files related to all experiment for that target e.g 1A0R<br> Target/*_mod_run_? contain files for all random rotation experiments and combdock results. e.g 1A0R/1A0R_mod_run_1/<br> Target/*_mod_combine_athird contains files for combined rotation experiment for RL-MLZerD and Multilzerd. e.g 1A0R/1A0R_mod_combine_athird/<br> Target/alphafold_multimer contains prediction files for AlphaFold-Multimer, both regular and no template run. e.g alphafold_multimer/1a0r and alphafold_multimer/1a0r_nodb<br> Target/colabfold contains prediction files for ColabFold results. e.g colabfold/prediction_1a0r</p>
Data for: Evolution Reinforces Cooperation with the Emergence of Self-Recognition Mechanisms: an empirical study of the Moran process for the iterated Prisoner's dilemma using reinforcement learning
<p>This contains data used for a paper titled: Evolution Reinforces Cooperation with the Emergence of Self-Recognition Mechanisms: an empirical study of the Moran process for the iterated Prisoner's dilemma using reinforcement learning.</p> <p>Numerous data sets are included, the main being `main.csv` which includes the fixation counts for a number of Moran processes between pairs of players from the Axelrod library.</p> <p>The source code and explanation of the data is here: https://github.com/Axelrod-Python/axelrod-moran. </p> <p> </p>
Deep Gradient Reinforcement learning for Music Improvisation in cloud computing framework
<p><span>The improvised music is further rendered in the MIDI format. The Bach Chorales dataset with six different attributes relevant to musical compositions is employed in implementing the present research. The model was set up in a containerised cloud environment and controlled for smooth load distribution. Five different parameters, such as pitch frequency (PF), standard pitch delay (SPD), average distance between peaks (ADP), note duration gradient (NDG) and pitch class gradient (PCG) are leveraged to assess the quality of the improvised music.</span></p>
Benchmarking Study of Deep Generative Models for Inverse Polymer Design: Reinforcement Learning
<p>Well-trained models and generation results for reinforcement learning part of <a href="https://github.com/ytl0410/Polymer-Generative-Models-Benchmark">ytl0410/Polymer-Generative-Models-Benchmark: Well-trained models and generative outcomes for the paper "Benchmarking Study of Deep Generative Models for Inverse Polymer Design" (github.com)</a></p>
FreeCiv games for the experiment on comparing Knowledge-Based Reinforcement Learning and Neural Networks in Strategic Games
<p>Dataset provides played FreeCiv games. The Tournament subset of them was played fully by two Artificial Intelligence (AI) agents against each other and one more computer player. The Human games subset was played by humans for demonstration/teaching purposes. The rest of the games were played 120 turns and stopped.</p>
Reinforcement learning theory reveals the cognitive requirements for solving the cleaner fish market task
<p>Learning is an adaptation that allows individuals to respond to environmental stimuli in ways that improve their reproductive outcomes. The degree of sophistication in learning mechanisms potentially explains variation in behavioural responses. Here, we present a model of learning that is inspired by documented intra- and interspecific variation in the performance in a simultaneous two-choice task, the 'biological market task'. The task presents a problem that cleaner fish often face in nature: the decision of choosing between two client types; one that is willing to wait for inspection and one that may leave if ignored. The cleaners' choice hence influences the future availability of clients, i.e. it influences food availability. We show that learning the preference that maximizes food intake requires subjects to represent in their memory different combinations of pairs of client types rather than just individual client types. In addition, subjects need to account for future consequences of actions, either by estimating expected long-term reward or by experiencing a client leaving as a penalty (negative reward). Finally, learning is influenced by the absolute and relative abundance of client types. Thus, cognitive mechanisms and ecological conditions jointly explain intra and interspecific variation in the ability to learn the adaptive response.</p>
Michalis Panayides thesis: Reinforcement learning dataset
<p>A dataset that consists of the reinforcement learning experiments used in my thesis. There are 5 main subdirectories. Each subdirectory consists of scripts, data, jupyter notebooks and plots for the corresponding utility function of the reinforcement learning algorithm. See README.md for more information.</p>
Policies used in "Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark"
<p>Policies used in the paper "Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark". They can be tested using the software "OTTO-benchmark" available at https://github.com/auroreloisy/otto-benchmark.</p>
YudengLin/memristorBDNN: Uncertainty quantification via a memristor Bayesian deep neural network for risk-sensitive reinforcement learning
<p>This code repository is partly to support risk-sensitive reinforcement learning experiment in the manuscript "Uncertainty quantification via a memristor Bayesian deep neural network for risk-sensitive reinforcement learning" submitted to Nature Machine Intelligence.</p>
Dataset for the paper: Emergent Resource Exchange and Tolerated Theft Behavior using Multi-Agent Reinforcement Learning
<p>Data generated from the experiments for the paper "Emergent Resource Exchange and Tolerated Theft Behavior using Multi-Agent Reinforcement Learning"</p>
Reinforcement Learning in Diabetes Mellitus Trial
ClinicalTrials.gov study NCT04473326. IPD Sharing: NO. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.