Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
73
datasets available to search
ShareScore release 0.9.0
Dataset results
73 results for “Unit Test”
Catalan United Nations v1.0 test set
<p>Catalan version [1] of the test set from the United Nations v1.0 [2]. The translation was performed in two steps: we did a first automatic translation from the Spanish test set version into Catalan and then a professional translator post-edited the output.</p> <p><br> [1] Marta R. Costa-Jussà, Noé Casas, Carlos Escolano, and José A. R. Fonollosa. 2019. Chinese-Catalan: A Neural Machine Translation Approach Based on Pivoting and Attention Mechanisms. <em>ACM Trans. Asian Low-Resour. Lang. Inf. Process.</em> 18, 4, Article 43 (August 2019), 8 pages. DOI:https://doi.org/10.1145/3312575</p> <p>[2] Michal Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The United Nations parallel corpus v1.0. In<br> Proceedings of the LREC, 2016</p>
Block-oriented description of a fictional CHP generation plant for unit commitment testing
<p>Block-oriented description of a fictional combined heat&power generation plant for unit commitment testing.</p> <p>Example input data is available at: https://doi.org/10.5281/zenodo.10875471</p>
Dataset for testing unit commitment models of CHCP plants
<p>Zipped set of 365 JSON files, one for each day of year 2022, containing fictional but realistic data for a combined heat, power and cooling generation plant:</p> <ul> <li>heat_demand: heat demand in kWh</li> <li>cooling_demand: cooling demand in kWh</li> <li>electricity_purchase: price for purchasing electricity in €/MWh</li> <li>electricity_selling: price for selling electricity in €/MWh</li> <li>gas_purchase: price for natural gas in €/MWh</li> </ul> <p> </p>
Dataset for Automated Unit Test Generation via Chain of Thought Prompt and Reinforcement Learning
<p>This is the replication package including three types datasets: training dataset with CoT prompts, reward dataset for training reward model, rl dataset for optimizing policy model. The training dataset includes filter_test_cot_rule_50k.csv, filter_train_cot_rule_50k.csv, and filter_valid_cot_rule_50k.csv. These three datasets includes multiple fields (i.e., src_fm, intention, plan, elaboration, gpt_test, src_fm_cot_gpt, target, src_fm_fc_ms_ff,src_fm_intention,src_fm_plan,src_fm_elaboration,idx,rule_cot,rule_cot_nlp,combine_cot,src_fm_rule_cot_nlp,src_fm_cot_nlp_gpt,gpt_cot_filter,src_fm_plan_intention). The reward dataset includes test_athena.json, train_athena.json, and valid_athena.json three files. The rl dataset includes three files: filter_test_cot_gpt_rl.csv, filter_train_cot_gpt_rl.csv, filter_valid_cot_gpt_rl.csv. These files include mulitple fields: src_fm,intention,plan,elaboration,gpt_test,src_fm_cot_gpt,target,src_fm_fc_ms_ff,src_fm_intention,src_fm_plan,src_fm_elaboration,gpt_cot_filter.</p>
Unit Test of GalapagosSenDT128
<p>A unit test dataset.</p>
Artifact of An LLM-based Readability Measurement for Unit Tests' Context-aware Inputs
<div> <p><strong>Artifact of An LLM-based Readability Measurement for Unit Tests' Context-aware Inputs</strong></p> <p> </p> <p>A README.md file can be found in the base folder after unzipping the artifact.</p> <p> </p> </div>
Data for Effective Unit Test Generation for Java Null Pointer Exceptions
<p>This is to provide all contents of our work, NPETest, including all raw experimental data used in our paper which will be published at ASE`24.</p> <p> </p> <p>NPETest is an unit test generation tool for Java projects, which utilizes both static and dynamic analysis techniques for effective NPE detection. This tool is implemented on the top of EvoSuite, a publicly available unit test generation tool for Java.</p> <p>For more technical details, please read our paper which will be published at ASE`24.</p> <p> </p> <p>The descriptions for the uploaded files are as follows:</p> <p>npetest_result.zip: results of NPETest for all benchmarks, containing the generated test-cases</p> <p>evosuite_opt_result.zip: results of EvoSuite with fine-tuned options for all benchmarks, containing the generated test-cases</p> <p>evosuite_def_result.zip: results of EvoSuite with default options for all benchmarks, containing the generated test-cases</p> <p>randoop_NPEX.tar.gz: results of Randoop for NPEX benchmarks, containing only the log-files. </p> <p>randoop_other.tar.gz: results of Randoop for Bears, BugSwarm, Defects4J, Genesis benchmarks, containing only the log-files.</p> <p>subject_gits.tar.gz: information of the buggy version for each benchmark. </p> <p>NPETestArtifact-main.zip: all contents of NPETest from the public Github respository <a href="https://github.com/kupl/NPETestArtifact" target="_blank" rel="noopener">NPETestArtifact</a>.</p> <p> </p> <p> </p> <p>The detailed description for NPETest (e.g., Install, Usage) of the tool is available on the public repository: <a href="https://github.com/kupl/NPETestArtifact" target="_blank" rel="noopener">NPETestArtifact</a>.<br>You can also download VM from the following link: <a href="https://doi.org/10.5281/zenodo.13371823" target="_blank" rel="noopener">Zenodo</a></p>
Artifacts for paper "CITYWALK: Enhancing LLM-Based C++ Unit Test Generation via Project-Dependency Awareness and Language-Specific Knowledge" submitted to TOSEM
<p>The project includes the data and code used in the submitted TOSEM paper titled "CITYWALK: Enhancing LLM-Based C++ Unit Test Generation via Project-Dependency Awareness and Language-Specific Knowledge"</p>
Daily United States COVID-19 Testing and Outcomes Data By State, March 7, 2020 to March 7, 2021
<p>The COVID Tracking Project was a volunteer organization launched from The Atlantic and dedicated to collecting and publishing the data required to understand the COVID-19 outbreak in the United States. Our dataset was in use by national and local news organizations across the United States and by research projects and agencies worldwide.</p> <p>Every day, we collected data on COVID-19 testing and patient outcomes from all 50 states, 5 territories, and the District of Columbia by visiting official public health websites for those jurisdictions and entering reported values in a spreadsheet. The files in this dataset represent the entirety of our COVID-19 testing and outcomes data collection from March 7, 2020 to March 7, 2021. This dataset includes official values reported by each state on each day of antigen, antibody, and PCR test result totals; the total number of probable and confirmed cases of COVID-19; the number of people currently hospitalized, in intensive care, and on a ventilator; the total number of confirmed and probable COVID-19 deaths; and more.</p>
Replication Kit: "Are Unit and Integration Test Definitions Still Valid for Modern Java Projects? An Empirical Study on Open-Source Projects"
<p><strong>Replication Kit for the Paper "Are Unit and Integration Test Definitions Still Valid for Modern Java Projects? An Empirical Study on Open-Source Projects"</strong><br> This additional material shall provide other researchers with the ability to replicate our results. Furthermore, we want to facilitate further insights that might be generated based on our data sets.</p> <p><strong>Structure</strong><br> The structure of the replication kit is as follows:</p> <ul> <li><strong>additional_visualizations</strong>: contains additional visualizations (Venn-Diagrams) for each projects for each of the data sets that we used</li> <li><strong>data_analysis</strong>: contains python scripts that we used to analyze our raw data</li> <li><strong>data_collection_tools</strong>: contains all source code used for the data collection, including the used versions of the <a href="https://github.com/comfort-framework">COMFORT framework</a>, the <a href="https://github.com/ftrautsch/BugFixClassifier">BugFixClassifier</a>, and the used tools of the <a href="https://github.com/smartshark">SmartSHARK environment</a>;</li> <li><strong>mongodb_no_authors</strong>: Archived dump of our MongoDB that we created by executing our data collection tools. The "comfort" database can be restored via the mongorestore command.</li> </ul> <p><br> <strong>Additional Visualizations</strong><br> We provide two additional visualizations for each project:<br> 1) <project_name>\_disj\_ieee\_venn (visualizations for the DISJ data set)<br> 2) <project_name>\_all\_ieee\_venn (visualizations for the ALL data set)</p> <p>For each of these data sets there exist one visualization for each project that shows four Venn-Diagrams for each of the different defect types. These Venn-Diagrams show the number of defects that were detected by either unit, or integration tests (or both).</p> <p>Furthermore, we added boxplots for each of the data sets (i.e., ALL and DISJ) showing the scores of unit and integration tests for each defect type.</p> <p><br> <strong>Analysis scripts</strong><br> Requirements:<br> - python3.5<br> - tabulate<br> - scipy<br> - seaborn<br> - mongoengine<br> - pycoshark<br> - pandas<br> - matplotlib</p> <p>Both python files contain all code for the statistical analysis we performed.</p> <p><strong>Data Collection Tools</strong><br> We provide all data collection tools that we have implemented and used throughout our paper:</p> <ul> <li><strong>BugFixClassifier</strong>: Used to classify our defects.</li> <li><strong>comfort-core</strong>: Core of the comfort framework. Used to classify our tests into unit and integration tests and calculate different metrics for these tests.</li> <li><strong>comfort-jacoco-listner</strong>: Used to intercept the coverage collection process as we were executing the tests of our case study projects.</li> <li><strong>jSHARK</strong>: Library that contains models for the used ORM mapper that is used inside the SmartSHARK environment (for Java).<strong> </strong></li> <li><strong>pycoSHARK</strong>: Library that contains models for the used ORM mapper that is used inside the SmartSHARK environment (for Python).</li> <li><strong>tools-changedistiller</strong>: Version of ChangeDistiller that we used within our comfort-core framework.</li> <li><strong>vcsSHARK</strong>: Used to collect data from the VCSs of the projects.</li> </ul> <p> </p> <p> </p>
Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation
<p><strong>Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation</strong></p> <p> </p> <p>This artifact contains:</p> <p>1. The binary folder. Its instruction can be found in [./binary/README.md](./binary/README.md).</p> <p>2. All classes used in the experiments [./classes.csv](./classes.csv).</p> <p>3. The experimental data zip file ([experimental_data.zip](./experimental_data.zip)).</p> <p> </p>
Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation
<p><strong>Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation</strong></p> <p> </p> <p>This zip file is the artifact that contains the following:</p> <p>1. The binary. </p> <p>2. All classes used in the experiments.</p> <p>3. The experimental data.</p> <p> </p>
Artefact for "Combining Type Inference and Automated Unit Test Generation for Python"
<p>Contains the artefact for our ASE 2023 submission “Combining Type Inference and Automated Unit Test Generation for Python”.</p>
Artifact of Enhancing Search-Based Unit Test Generation with Method-Level Java Class Splitter
<p># Artifact of Enhancing Search-Based Unit Test Generation with Method-Level Java Class Splitter</p> <p> </p> <p>This artifact contains:</p> <p>1. The binary folder. Its instruction can be found in [./binary/README.md](./binary/README.md).</p> <p>2. All classes used in the experiments [./classes.csv](./classes.csv).</p> <p>3. The experimental data zip file ([experimental_data.zip](./experimental_data.zip)).</p>
Determination of Blomia Tropicalis Allergen Extract in Prick Test Units
ClinicalTrials.gov study NCT07195929. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Evaluation of Obstructive Sleep Apnea (OSA) Using Portable Sleep Testing (PST) Devices on an Inpatient Stroke Unit
ClinicalTrials.gov study NCT06516354. IPD Sharing: NO. Countries: 1. Publications: 7.
The Effect of Blood Tests Performed Within the Routine Protocol in the Surgical Intensive Care Unit on Hemoglobin Levels
ClinicalTrials.gov study NCT05833178. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Assessment of Transcutaneous Oxygen Tension/Oxygen Challenge Test in Intensive Care Unit (ICU) Patients
ClinicalTrials.gov study NCT01174966. IPD Sharing: Not stated. Countries: 1. Publications: 5.
Preoperative Exercise Test and Postoperative Intensive Care Unit Need in Bariatric Surgery
ClinicalTrials.gov study NCT03419104. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Air Leak Test In Pediatric Intensive Care Unit
ClinicalTrials.gov study NCT05328206. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.