Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,654
datasets available to search
ShareScore release 0.7.1
Dataset results
1,654 results for “Automation”
Student's logs and perceptions of an automated assessment tool in a software engineering MOOC specialization
<p>Our dataset contains students' perceptions and usage of an automated assessment tool (MOOCauto) for obtaining formative feedback in software engineering assignments that are part of a MOOC specialization at Universidad Politécnica de Madrid (Spain), delivered by the MiriadaX platform. The dataset has previously been used in a study to evaluate students' perceptions of the tool and to analyze their usage patterns using Growth Mixture Models <a href="https://www.computer.org/csdl/magazine/so/5555/01/10196480/1P9AhkBLYXK">(López-Pernas et al., 2023)</a>. The code of each of the assignments is available on Github: <a href="https://github.com/ging-moocs">https://github.com/ging-moocs</a>.</p> <p>Our dataset contains two files:</p> <h2>MOOCauto usage logs</h2> <p>The first file is called<strong> moocauto_logs.csv </strong>and it contains 9,108 anonymized logs of students' use of the automated assessment tool in the MOOC specialization assignments. The columns of the dataset are as follows:</p> <ul> <li><strong>MOOCid</strong>: Unique numeric identifier for the MOOC (1-4)</li> <li><strong>MOOC: </strong>Name of the MOOC: Frontend Development, Backend Development, Git & Github, Fullstack Development</li> <li><strong>AssignmentName</strong>: Name of the assignment.</li> <li><strong>AssignmentId</strong>: Unique identifier for each assignment (1-17)</li> <li><strong>user: </strong>Unique identifier of the student (it varies per assignment)</li> <li><strong>timestamp: </strong>Time in which the assessment was performed</li> <li><strong>score</strong>: Score obtained (0-10)</li> </ul> <h2>Students' perceptions of MOOCauto</h2> <p>The second file is called <strong>moocauto_questionnaire.csv</strong> and it contains 213 students' responses to the questionnaire conducted at the end of each MOOC in order to evaluate their opinion of the tool and perception on usefulness, ease of use, and other aspects related to the Technology Acceptance Model (TAM). The questions were as follows:</p> <ul> <li><strong>What is your general opinion of MOOCauto?</strong> (1 Horrible - 5 Excellent)</li> <li><strong>Indicate your level of agreement with the following statements </strong>(1 Strongly disagree - 5 Strongly agree) <ul> <li>MOOCauto has been easy to install</li> <li>MOOCauto has been easy to use</li> <li>The feedback provided by MOOCauto was easy to understand</li> <li>The feedback provided by MOOCauto was useful</li> <li>The feedback provided by MOOCauto helped me improve my assignments</li> <li>The documentation Of MOOCauto was useful</li> <li>MOOCauto has increased my motivation to work on the assignments</li> <li>I prefer the feedback from MOOCauto than from peer assessment</li> <li>I would like to have a bot like MOOCauto in other MOOCs</li> </ul> </li> <li><strong>How useful do you perceive the following features of MOOCauto?</strong> (1 Useless - 5 Very useful) <ul> <li>It works locally on my computer</li> <li>It allows to run the test suite as many times as I want</li> <li>It provides instantaneous feedback every time the test suite is executed</li> <li>It has documentation that explains its use and available options</li> </ul> </li> </ul>
Evaluating Automated Seismic Event Detection Approaches: An Application to Victoria Land, East Antarctica
<p>This repository contains the waveform data used by Ho et al. (2024), along with all generated fine-tuned models and event catalogs. See the README file for a summary. The corresponding software packages are available on GitHub at <a href="https://github.com/jakewalter/easyQuake.git">https://github.com/jakewalter/easyQuake.git</a>, <a href="https://github.com/seisbench/seisbench">https://github.com/seisbench/seisbench</a>, and <a href="https://github.com/longmho/Transfer_Learning_and_Seisbench">https://github.com/longmho/Transfer_Learning_and_Seisbench</a>. See the README.md file at <a href="https://github.com/longmho/Transfer_Learning_and_Seisbench">https://github.com/longmho/Transfer_Learning_and_Seisbench</a> for additional details.</p>
Final model for "Automated Large-Scale Full Seismic Waveform Inversion for North America and the North Atlantic" by Krischer et al. (2018)
<p>The HDF5 file contains the final model of the paper "Automated Large-Scale Full Seismic Waveform Inversion for North America and the North Atlantic" by Krischer et al. (2018), soon to be published in the Journal of Geophysical Research - Solid Earth.</p> <p>The "coordinates_0", "coordinates_1", and "coordinates_2" data sets are the coordinates along each dimension, here colatitude in degree, longitude in degree, and radius in meter, respectively. The regularly sampled data is available in five 3D-arrays in the "data" group: "vp", "vsv", "vsh", "rho", and "Q". Velocities are defined at 1 Hertz and are given in km/s, the density in kg/m^3. Q is Q_mu.</p> <p>The coordinates have to be rotated to yield true spherical Earth coordinates. They have to be rotated around on axis vector of 0.766044443118978/0.6427876096865393/0.0 in cartesian x/y/z coordinates by -30.0 degrees. Conversion of spherical to cartesian coordinates happens with the standard convention:</p> <p>x = r sin(theta) cos(phi)<br> y = r sin(theta) sin(phi)<br> z = r cos(theta)</p>
Dataset for ICFHR2018 Competition on Automated Text Recognition on a READ Dataset
<p>The main idea of this dataset is to analyse the impact of training data. How many training data specific to the document, you are transcribing, is necessary? </p> <p><strong>general data: </strong>This is a collection of heterogeneous documents to train an initial system. For each text line there is an image file of that line, a file with the ground truth text and an information file containing an automatically generated surrounding polygon.</p> <p><strong>specific data: </strong>The specific data contains documents related to the test data. For the specific systems only the images of the train list may be used. The file are of the same type as the general data.</p> <p><strong>test data: </strong>The test data contains only the images and the information files.</p> <p>More Information, some published results and an evaluation procedure at https://scriptnet.iit.demokritos.gr/competitions/10/</p>
Critical Assessment of automated Structure Determination of Proteins by NMR
<p>The community-wide initiative "Critical Assessment of Automated Structure Determination of Proteins by NMR (<strong>CASD-NMR</strong>)" was launched in 2009 to to evaluate the ability of automated methods to produce 3D protein structures from NMR data that closely match structures manually determined by experts.</p> <p>This dataset includes all the experimental data made available to the participants of CASD-NMR in the two completed rounds of the initiative.</p> <p>Also refer to http://www-nmr.cabm.rutgers.edu/blindtest/blind.html for additional details, including first release date and link to each final PDB entry</p>
Automated Classification of Conversation Valence and Arousal using Autonomic Nervous System Responses
<p>This repository contains the supplementary file for our study "Automated Classification of Conversation Valence and Arousal using Autonomic Nervous System Responses". The MS Excel file contains all physiological features (individual features and synchrony features) for all valid dyads and all intervals together with self-report ratings of the conversation (Self-Assessment Manikin) and personality trait data (CES-D, BFNES, QCAE). Synchrony features were calculated using code from a previous Zenodo submission (https://zenodo.org/record/7140829).</p>
Vul4J+: A Dataset of Vulnerabilities for Automated Vulnerability Repair
<div> <div><strong>Vul4J+</strong> is a dataset of vulnerability fixes for automated vulnerability repair (AVR) in Java. Each entry of the dataset represents a <strong>vulnerability</strong> affecting an open-source Java project, having reference to the commit (revision) containing the code affected by the vulnerability and its version fixed by a human developer (the "left" and "right" parts of the commit). Each vulnerability is equipped with at least one <strong>"oracle"</strong> that shows the presence of the vulnerability, and that can be used to validate the correctness of patches generated by AVR tools. This *"oracle"* might have the form of a:</div> <div>- <strong>Vulnerability-witnessing test</strong>, i.e., a JUnit test case that fails on the vulnerable version of the code but passes on the patched version.</div> <div>- <strong>Warning/report</strong> raised by a vulnerability static analyzer, i.e., SpotBugs, that is presented in the vulnerable version of the code but not in the patched version.</div> <br> <div>In essence, Vul4J+ is a cleaned up and extended version of Vul4J containing:</div> <div>- 106 known vulnerabilities with executable vulnerability-witnessing test cases in Docker containers and warnings (reports) from SpotBugs static analyzer (if found);</div> <div>- 79 come from the original Vul4J;</div> <div>- 27 result from the replication of the same protocol used in the original Vul4J;</div> <div>- 50 vulnerabilities stored in Docker containers with the warnings (reports) from SpotBugs static analyzer ;</div> <div>- 35 known vulnerabilities matched with vulnerability-witnessing test cases retrieved from projects in the wild.</div> <br> <div>In total, Vul4J+ points to <strong>191 vulnerabilities</strong>, each with at least one vulnerability oracle.</div> </div>
Scores for calculating automated FAIR assessments in the low carbon energy domain
<p>Results for an automated FAIR assessment of 80 databases from the low carbon energy domain. The assessment was performed with the help of the FAIR maturity evaluation service of Wilkinson et al. The FAIR status with respect to 16 FAIR criteria is listed. The scores are defined to be consistent with the FAIR assessment tool of the Australian Research Data Commons. More details can be found in an additional publication on Zenodo as well as in an upcoming publication by Schwanitz et al.</p>
The datasets for "Automated Recovery of Issue-Commit Links Leveraging Both Textual and Non-textual Data" paper
<p>Paper title: Automated Recovery of Issue-Commit Links Leveraging Both Textual and Non-textual Data</p> <p>Conference: ICSME 2021</p>
Replication package for "Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation"
<p>This repository contains the replication package for the paper "Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation" by Fernando Vallecillos Ruiz, Anastasiia Grishina, Max Hort and Leon Moonen, accepted for publication in ACM Transactions on Software Engineering and Methodology on 2025-10-09.</p> <p>A preprint is deposited on arXiv with DOI: <a href="https://doi.org/10.48550/arXiv.2401.07994">10.48550/arXiv.2401.07994</a>.</p> <p>The replication package is archived on Zenodo with DOI: <a href="https://doi.org/10.5281/zenodo.10500593">10.5281/zenodo.10500593</a>. It is maintained on GitHub at <a href="https://github.com/secureIT-project/RTT_for_APR">https://github.com/secureIT-project/RTT_for_APR</a>.</p> <p>This project builds on code from the <a href="https://github.com/lin-tan/clm/">clm</a> project, which is (c) 2023, The ASSET research group led by Lin Tan, Purdue University, licensed under the BSD 3-Clause License (see jasper/LICENSE.BSD). All modifications and new contributions are (c) 2025 by the authors of this replication package and distributed under the MIT License (see LICENSE.MIT). The data, models and preprint are distributed under the CC BY 4.0 license.</p> <h2>Citation<code> </code></h2> <p>If you build on this data or code, please cite this work by referring to the paper:</p> <div> <pre><code>@article{ruiz2025:rtt, title = {Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation}, author = {Vallecillos Ruiz, Fernando and Anastasiia Grishina and Max Hort and Leon Moonen}, journal = {ACM Transactions on Software Engineering and Methodology (TOSEM)}, year = {2025}, publisher = {{ACM}} }</code></pre> </div> <h2>Organization</h2> <p>The replication package is organized as follows:</p> <ul> <li>clm-apr <ul> <li>plbart: code to generate patches with PLBART models.</li> <li>codet5: code to generate patches with CodeT5 models.</li> <li>transcoder: code to generate patches with the TransCoder model.</li> <li>incoder: code to generate patches with InCoder models.</li> <li>santacoder: code to generate patches with the SantaCoder model.</li> <li>starcoder: code to generate patches with the StarCoderBase model.</li> <li>quixbugs: code to validate patches generated for the QuixBugs benchmark.</li> <li>defects4j: code to validate patches generated for any of the Defects4J benchmarks.</li> <li>humaneval: code to validate patches generated for the HumanEval-Java benchmark.</li> </ul> </li> <li>humaneval-java: the HumanEval-Java benchmark proposed by Jiang et al. 2023</li> <li>jasper: a Java tool to parse Java programs needed to preprocess input.</li> <li>model: folder to download the language models.</li> <li>analysis_wandb: data from WandB and Jupyter notebook to create graphs.</li> <li>tmp_benchmarks: folder for temporary files used in patch validation. The folder may contain pairs of `paralell’ folders src and src_org for each benchmark, used to replace buggy code with candidate patches.</li> </ul> <h2>Replication</h2> <h3>Prerequisites</h3> <ul> <li>Python version: 3.8—3.10.</li> <li><a href="https://git-lfs.com/">Git LFS</a> is required for model downloading.</li> </ul> <h4>Weight and Biases (WandB)</h4> <ol> <li>Create an account on <a href="https://wandb.ai/">Weights and Biases</a></li> <li>Install the <a href="https://docs.wandb.ai/ref/python">Weights and Biases</a> library</li> <li>Run <code>wandb login</code> and follow the instructions</li> </ol> <h4>Set up OpenAI access</h4> <p>OpenAI account is needed with access to <code>gpt-3.5-turbo</code> and <code>gpt-4</code> . The <code>OPENAI_API_KEY</code> environment variable should be set to your OpenAI API access token.</p> <h3>Dependencies</h3> <ul> <li><a href="https://github.com/rjust/defects4j">Defects4J</a> - To generate inputs for the Defects4J datasets or to validate them, you need to have installed <a href="https://github.com/rjust/defects4j">their tool</a>.</li> <li>Java 8</li> <li>Apache Maven</li> </ul> <h3>Setup</h3> <p>We recommend the use of the setup script:</p> <pre><code>setup.sh </code></pre> <p>which performs the following:</p> <ol> <li>Creates a virtual environment for Python and activate it.</li> <li>Install the packages in <code>requirements.txt</code>.</li> <li>Compiles Jasper.</li> <li>Downloads parsers.</li> <li>Check if the Defects4J installation is correct.</li> </ol> <h3>Download models</h3> <p>The following bash script contains the code to download all of the models used:</p> <pre><code>models/download_models.sh </code></pre> <p>We recommend downloading only the models you are going to use due to their size</p> <pre><code>cd models chmod +x download_models.sh ./download_models.sh </code></pre> <p>To run one specific model, for example, PLBART (C#), use the following commands:</p> <pre><code>cd models git lfs install git clone https://huggingface.co/uclanlp/plbart-java-cs git clone https://huggingface.co/uclanlp/plbart-cs-java cd ../.. </code></pre> <h3>Step 1: Preprocessing and Prompting:</h3> <p>Each script in each <code>clm-apr/[model]</code> folder connects one or more models with<br>one dataset. These scripts follow the template: [benchmark]_[model]_[technique].py.<br>The scripts first create an <code>[model]_input.json</code> file with the preprocessed<br>input. Then generate outputs based on that file with one or more models.<br>For example:</p> <pre><code>cd clm-apr/plbart python quixbugs_plbart_round.py # Generates input for QuixBugs and generate patches using Java<->C# RTT. python quixbugs_plbart_round_nl.py # Generates input for QuixBugs and generate patches using Java<->NL RTT. </code></pre> <p>Optionally, use argument <code>--device_map cpu</code> if you wish to run the script on<br>CPU, for example:</p> <pre><code>python quixbugs_plbart_round.py --device_map cpu </code></pre> <p>Otherwise, the script will be run on all available CUDA GPU’s.</p> <p>We have commented the generation of inputs in the scripts. Users are free to<br>uncomment this method and try for themselves. It is easily recognizable by<br>their name template <code>[model]_[benchmark]_input()</code>. In the previous case:</p> <pre><code>quixbugs_plbart_input() </code></pre> <h3>Step 2 and 3: Round Trip Translation and Postprocessing</h3> <p>These steps are also included in the [benchmark]_[model]_[technique].py<br>script mentioned above. They are modularized in the method recognizable by<br>their name template [model]_[benchmark]_output().<br>For example:</p> <pre><code>quixbugs_incoder_output() </code></pre> <p>This method:</p> <ol> <li>Reads the input json file.</li> <li>Generates outputs through the LLM.</li> <li>Postprocess the output (extract the patch, clean up extra token, etc.).</li> <li>Creates [model]_output_[technique]_[extra].json.</li> </ol> <p>The last 3 steps are repeated according to the number of runs set to performed<br>(10 in our experiments). Each run will produce a different file with the seed<br>used in its generation. For example, <code>quixbugs\_plbart\_round.py</code> and<br><code>quixbugs\_plbart\_round_nl.py</code> scripts create:</p> <pre><code>clm-apr/quixbugs/plbart_results/run_0/plbart_java_cs_java_output_round_csharp_batch.json clm-apr/quixbugs/plbart_results/run_0/plbart_java_nl_java_output_round_nl_batch.json </code></pre> <h3>Step 4: Evaluation of RTT Results:</h3> <p>The last step evaluates the generated outputs against the test-suites of each<br>benchmark. This script reads the previous outputs files and generates a new one<br>with the results of the test for one model. Furthermore, it connects with the<br><em>WandB</em> tool to calculate metrics and send them to analyze.</p> <p>Following the previous examples, to validate the results previously obtained,<br>we execute the following:</p> <pre><code>cd clm-apr/quixbugs python validate_quixbugs_parallel.py </code></pre> <p>Given the included JSON, this script would create:</p> <pre><code>clm-apr/quixbugs/plbart_results/run_0/plbart_java_cs_java_validate_round_csharp_batch.json </code></pre> <p>We have disabled <em>WandB</em> in the script to allow users to try the script first.<br>However, it can be easily activated by changing the parameter <code>mode="disabled"</code><br>to <code>mode="online"</code>.<br>We have set the variable <code>total_runs = 1</code>, as well as <code>input_file</code> and <code>output_file</code><br>to the results included. They should be modified accordingly to validate more runs<br>or to validate other files/models.</p> <h3>Included Results</h3> <p>We include two CSV files obtained through WandB.</p> <pre><code>'data_cleaned_grouped.csv': Aggregated metrics of the 25 outputs for all runs. 'full_data_all_runs.csv': All metrics for all outputs on all runs. </code></pre> <h2>Changelog</h2> <ul> <li>v1.0 - updates corresponding to the accepted version of the manuscript in TOSEM</li> <li>v0.1 - initial replication package corresponding to v1 of arXiv deposit: includes raw data, code, and example outputs.</li> </ul> <h2>References</h2> <p>Jiang, N.; Liu, K.; Lutellier, T.; and Tan, L. 2023. Impact of Code Language<br>Models on Automated Program Repair. In 45th International Conference on<br>Software Engineering (ICSE), 1430–1442. IEEE. ISBN 978-1-66545-701-9.</p> <div> </div>
Automated Minirhizotron Validation Data
<p>This dataset is contains the validation data (raw images and binary masks generated by hand) for the automated minirhizotrons in our paper 'HIGH FREQUENCY ROOT DYNAMICS: SAMPLING AND INTERPRETATION USING REPLICATE ROBOTIC MINIRHIZOTRONS' published in Journal of Experimental Botany. <a href="https://doi.org/10.1093/jxb/erac427">https://doi.org/10.1093/jxb/erac427</a></p> <p>The dataset consists of one readme, four .7z files and one python script all contained in one .7z file.</p> <p>The files are described in the readme file.</p> <p>Some of these data were collected and all processed as part of the Marie Sklodowska-Curie project 748893 'MrPARTS' awarded to Richard Nair. We also acknowledge the generous support of Markus Reichstein at MPI-BGC Jena including to the MaNiP project through the Max Planck research prize 2013</p>
Supplementary Material for 'Leveraging the GIDAS Database for the Criticality Analysis of Automated Driving Systems'
<p>This repository contains the supplementary material for the publication 'Leveraging the GIDAS Database for the Criticality Analysis of Automated Driving Systems'.<br> It consists of four files:</p> <ol> <li>Criticality-Phenomena-Catalog.CSV: The catalog of criticality phenomena (CP)</li> <li>Criticality-Phenomena-Phi-Coefficient.CSV: The calculation of the Phi coefficient between all pairs of CP</li> <li>Criticality-Phenomena-Risk-Calculation.CSV: The case-phenomenon relation matrix, including the calculated values for the risk of each CP for all three severity levels</li> <li>Criticality-Phenomena-Sorted-By-Risk.CSV: A list of the CP from the CP catalog sorted by risk for all three severity classes</li> </ol>
Automated monitoring of biodiversity in the tropics: A pilot study at Barro Colorado Island
<p>Automated monitoring of biodiversity using camera systems and acoustic devises is becoming increasingly common practise. The development of camera traps for insects, however, is relatively new, and has not been tested in tropical environments. To understand both the challenges and opportunities for automated monitoring in the tropics we deployed three newly developed cameras systems for monitoring insects, with a focus on night-flying insects. In addition we deployed audible and ultrasound recording equipment. The locations of each device was changed over the 5 day pilot study, and the configuration of each system was changed to allow an assessment of the relative importance of position, light, etc. The study was undertaken on Barro Colorado Island, Panama, in January 2023.</p> <p>Data are aggregated into <em>.zip</em> files, and details of their contents, and metadata are given in <em>read_me.txt</em> and <em>metadata.csv</em>.</p>
Lithium-Ion Batteries in Automated Guided Vehicles (AGVs) dataset for article "Automated Battery Power Fade Estimation for Fast Charge and Discharge Operations"
<p>Dataset of aggregated information related to discharge-only cycles of lithium-ion battery packs employed in Automated Guided Vehicle systems.</p> <p>The dataset supports the study in conference article "Automated Battery Power Fade Estimation for Fast Charge and Discharge Operations"</p>
Catalytic Rules and Validation Results for "EzMechanism: An Automated Tool to Propose Catalytic Mechanisms of Enzyme Reactions"
<p>Dataset containing the "Rules of Enzyme Catalysis" as created during the development of EzMechanism and the validation results of the software. For more information see https://www.biorxiv.org/content/10.1101/2022.09.05.506575v1, and the M-CSA website in https://www.ebi.ac.uk/thornton-srv/m-csa/</p>
Dataset: all TROPOMI detected plumes for 2021. [Schuit et al. 2023: Automated detection and monitoring of methane super-emitters using satellite data]
<p>Dataset of all TROPOMI detected methane plumes in 2021, including estimates for the source location, emission quantification and source type. Corresponds to Figure 6 of Schuit et al. 2023 [Automated detection and monitoring of methane super-emitters using satellite data, https://doi.org/10.5194/acp-23-9071-2023]. Additional details and context are provided in Section 3 of the paper.</p> <p> </p> <p><em>Contents and data formats</em></p> <p><strong>date</strong>, date of the TROPOMI observation. format: YYYYMMDD</p> <p><strong>time_UTC</strong>, time of the TROPOMI observation in UTC. format: HH:MM:SS</p> <p><strong>lat</strong>, latitude of the center of the TROPOMI pixel at the estimated source location. format: float</p> <p><strong>lon</strong>, longitude of the center of the TROPOMI pixel at the estimated source location. format: float</p> <p><strong>source_rate_t/h</strong>, estimated emission source rate in tonnes per hour, the methodology is described in Section 2.5.1 of the paper. format: int</p> <p><strong>uncertainty_t/h</strong>, the uncertainty of the emission source rate in tonnes per hour, the methodology is described in Section 2.5.1 of the paper. format: int</p> <p><strong>estimated_source_type</strong>, the locally dominant anthropogenic source sector based on bottom-up inventories, the methodology is described in Section 2.5.3 of the paper. format: str</p> <p> </p> <p>Full citation of the paper:</p> <p>Schuit, B. J., Maasakkers, J. D., Bijl, P., Mahapatra, G., van den Berg, A.-W., Pandey, S., Lorente, A., Borsdorff, T., Houweling, S., Varon, D. J., McKeever, J., Jervis, D., Girard, M., Irakulis-Loitxate, I., Gorroño, J., Guanter, L., Cusworth, D. H., and Aben, I.: Automated detection and monitoring of methane super-emitters using satellite data, Atmos. Chem. Phys., 23, 9071–9098, https://doi.org/10.5194/acp-23-9071-2023, 2023.</p>
Convolutional neural network for automated surface crack detection using inductive thermography
<p>Two phase images of the samples AIT_01 and AIT_08, analysed in the publication "Convolutional neural network for automated surface crack detection using inductive thermography", submitted to the Journal of Electronic Imaging.</p>
Evaluation of an adapted semi-automated DNA extraction for human salivary shotgun metagenomics
<p>This deposit contains :</p> <p>- a RMarkdown filte containing the codes for the mcirobial analysis of saliva samples</p> <p>- the html report with codes, results and figures</p> <p>- a RData containing microbial datasets (MSp species abundance table, genus, family and phylum abundance tables, matrix of genes correlations, taxonomy)</p> <p>- a RData containing associated metadata </p>
Versailles, Urban driving, Automated Driving and VRU detection
<p>Scenario description: cyclist and pedestrian detection during automated driving.</p> <p>Session description: Automated Driving part of the trip only. Two rounds: the first one without IoT (VRUs not connected) then with IoT (pedestrian with smartphone/watch and cyclist with connected bike).</p> <p>Datasets descriptions:</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_Vehicle: </strong>Data generated from the vehicle sensors</p> <p>Vehicle datasets generated by the vehicle sensors during urban driving at Versailles. This includes the data coming from the CAN bus and GPS. It includes following kind of datasets: Vehicle: general data (speed, battery), PositioningSystem: data from GPS, VehicleDynamics: data about dynamic (acceleration...), Accel: acceleration data, EnvironmentSensorsAbsolute: environment sensors in absolute coordinates</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_V2X: </strong>V2X messages during uraban driving sessions</p> <p>Data exchanged with other vehicles and pedestrian during urban driving. SortOfCam: messages sent or received from bicycles</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_IoT: </strong>Data extracted from IoT oneM2M platform</p> <p>This dataset refers to messages exchanged by urban driving and car sharing application with vehicle, across oneM2M platform. oneM2M: car sharing status data</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_CAM: </strong>CAM messages</p> <p>This dataset refers to messages captured inside the vehicle during car sharing.</p>
Brainport, Automated valet parking, dropoff scenario TNO vehicle
<p><strong>Scenario description</strong>:</p> <p>The AD-vehicle receives parking command message (AutoPilot.VehicleCommand) containing the destination parking spot and the free obstacle route and drives from the drop-off position and parks to the destination parking spot. During the parking process the vehicle send two type of messages (AutoPilot.PositionEstimate and AutoPilot.VehicleAVPStatus)</p> <p><strong>Session description</strong>:</p> <p>The dropoff scenario with TNO vehicle. TNO vehicle parks autonomously from the dropoff location to the selected parking spot at the parking area on the automotive campus.</p> <p><strong>Datasets descriptions</strong>:</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_DriverVehicleInteraction</strong>: Data extracted from the CAN of the vehicle</p> <p>Dataset Description This dataset contains e.g. throttlestatus, clutchstatus, brakestatus, brakeforce, wipersstatus, steeringwheel for the vehicle</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_DroneAvpCommand</strong>: Data sent from drone</p> <p>Dataset Description This dataset contains route information for a vehicle to a designated parking spot</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_EnvironmentSensorsAbsolute</strong>: Data extracted from the vehicle environment sensors</p> <p>Dataset Description This dataset contains information about detected object, with absolute coordinates</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_EnvironmentSensorsRelative</strong>: Data extracted from the vehicle environment sensors</p> <p>Dataset Description This dataset contains information about detected object, with relative coordinates</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_IotVehicleMessage</strong>: Data sent between all devices, vehicles and services</p> <p>Dataset Description Each sensor data submission is a Message. A Message has an Envelope, a Path, and optionally (but likely) Path Events and optionally Path Media. The envelope bears fundamental information about the individual sender (the vehicle) but not to a level that owner of the vehicle can be identified or different messages can be identified that originate from a single vehicle.</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_ParkingSpotDetection</strong>: Data sent from drone to parkingService</p> <p>Dataset Description This dataset contains informaton about detected parking spots</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_PositioningSystem</strong>: Data from GPS on the vehicle</p> <p>Dataset Description This dataset contains speed, longitude, latitude, heading from the GPS</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_PositioningSystemResampled</strong>: Data from GPS on the vehicle</p> <p>Dataset Description This dataset contains speed,longitude,latitude,heading from the GPS, resampled to 100 milliseconds</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_Vehicle</strong>: Data from the CAN and sensors about the state of the vehicle</p> <p>Dataset Description This dataset contains a.o temperature and battery state of the vehicles</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_VehicleAvpCommand</strong>: Data sent from ParkingService to vehicle</p> <p>Dataset Description This dataset contains route to parkingspot, and some other environmental information</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_VehicleAvpStatus</strong>: Data sent from vehicle to ParkingService</p> <p>Dataset Description This dataset contains information about the current status and parkingstatus of the vehicle</p> <p><strong>AUTOPILOT_BrainPort_AutomatedValetParking_VehicleDynamics</strong>: Data from the CAN and sensors about the state of the vehicle</p> <p>Dataset Description This dataset contains a.o accelerations and speedlimit of the vehicle, as observed from the CAN and the external sensors</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.