Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
708
datasets available to search
ShareScore release 0.7.1
Dataset results
708 results for “Java”
Replication Kit: "Are Unit and Integration Test Definitions Still Valid for Modern Java Projects? An Empirical Study on Open-Source Projects"
<p><strong>Replication Kit for the Paper "Are Unit and Integration Test Definitions Still Valid for Modern Java Projects? An Empirical Study on Open-Source Projects"</strong><br> This additional material shall provide other researchers with the ability to replicate our results. Furthermore, we want to facilitate further insights that might be generated based on our data sets.</p> <p><strong>Structure</strong><br> The structure of the replication kit is as follows:</p> <ul> <li><strong>additional_visualizations</strong>: contains additional visualizations (Venn-Diagrams) for each projects for each of the data sets that we used</li> <li><strong>data_analysis</strong>: contains python scripts that we used to analyze our raw data</li> <li><strong>data_collection_tools</strong>: contains all source code used for the data collection, including the used versions of the <a href="https://github.com/comfort-framework">COMFORT framework</a>, the <a href="https://github.com/ftrautsch/BugFixClassifier">BugFixClassifier</a>, and the used tools of the <a href="https://github.com/smartshark">SmartSHARK environment</a>;</li> <li><strong>mongodb_no_authors</strong>: Archived dump of our MongoDB that we created by executing our data collection tools. The "comfort" database can be restored via the mongorestore command.</li> </ul> <p><br> <strong>Additional Visualizations</strong><br> We provide two additional visualizations for each project:<br> 1) <project_name>\_disj\_ieee\_venn (visualizations for the DISJ data set)<br> 2) <project_name>\_all\_ieee\_venn (visualizations for the ALL data set)</p> <p>For each of these data sets there exist one visualization for each project that shows four Venn-Diagrams for each of the different defect types. These Venn-Diagrams show the number of defects that were detected by either unit, or integration tests (or both).</p> <p>Furthermore, we added boxplots for each of the data sets (i.e., ALL and DISJ) showing the scores of unit and integration tests for each defect type.</p> <p><br> <strong>Analysis scripts</strong><br> Requirements:<br> - python3.5<br> - tabulate<br> - scipy<br> - seaborn<br> - mongoengine<br> - pycoshark<br> - pandas<br> - matplotlib</p> <p>Both python files contain all code for the statistical analysis we performed.</p> <p><strong>Data Collection Tools</strong><br> We provide all data collection tools that we have implemented and used throughout our paper:</p> <ul> <li><strong>BugFixClassifier</strong>: Used to classify our defects.</li> <li><strong>comfort-core</strong>: Core of the comfort framework. Used to classify our tests into unit and integration tests and calculate different metrics for these tests.</li> <li><strong>comfort-jacoco-listner</strong>: Used to intercept the coverage collection process as we were executing the tests of our case study projects.</li> <li><strong>jSHARK</strong>: Library that contains models for the used ORM mapper that is used inside the SmartSHARK environment (for Java).<strong> </strong></li> <li><strong>pycoSHARK</strong>: Library that contains models for the used ORM mapper that is used inside the SmartSHARK environment (for Python).</li> <li><strong>tools-changedistiller</strong>: Version of ChangeDistiller that we used within our comfort-core framework.</li> <li><strong>vcsSHARK</strong>: Used to collect data from the VCSs of the projects.</li> </ul> <p> </p> <p> </p>
ConPlag: a Dataset of Programming Contest Plagiarism in Java
<p>This package presents ConPlag, the first dataset of programming contest plagiarism in Java. To find the details about how to use the dataset and how to run tools on it, please refer to the README.</p>
Data Analysis of Adolescent Unintended Pregnancy and Pre-marital Sex in East Java (BKKBN, 2019)
<p>The data used is secondary data in the form of raw data from SKAP (Program Performance and Accountability Survey) for East Java in 2019 conducted by BKKBN (National Population and Family Planning Agency). SKAP is a survey conducted by the BKKBN in order to measure the success of population development programs, family planning, and family development.<br> The data consist of women aged 15-18 years were selected, totaling 1815 women of childbearing age and adolescents. The adolescents included in this study were residents of East Java.</p>
The relocated catalog of the 2022 Cianjur (West Java) earthquake
<p>The relocated catalog of the 2022 Cianjur (West Java) earthquake</p> <p>Please refer to:</p> <p>Supendi, P., Winder, T., Rawlinson, N., Bacon, C. A., Palgunadi, K. H., Simanjuntak, A., Kurniawan, A., Widiyantoro, S., Nugraha, A. D., Shiddiqi, H. A., Ardianto, Daryono, Adi, S. P., Karnawati, D., Priyobudi, Marliyani, G. I., Imran, I., and Jatnika, J. (2023). A conjugate fault revealed by the destructive Mw 5.6 (November 21, 2022) Cianjur earthquake, West Java, Indonesia, <em>Journal of Asian Earth Sciences</em>,<strong> </strong>257, 105830. https://doi.org/10.1016/j.jseaes.2023.105830</p>
Supplementary data from: The Holocene fossil record of the slow loris (Nycticebus sp.) in Java (Indonesia)
<p>Supplementary data from: The Holocene fossil record of the slow loris (Nycticebus sp.) in Java (Indonesia)</p> <p>Abstract:</p> <p>Fossil lorises are rare in Southeast Asia. Their taxonomic relationship with extant populations, and the extent to which their distribution and morphology are influenced by changing environmental conditions, remain poorly understood. This study provides a synthesis of <em>Nycticebus</em> occurrences in Holocene Java (Indonesia). A morphometric analysis of a sample of craniodental remains aims to improve our understanding of their taxonomic status. Morphometrics were also used to explore potential size changes during the Holocene.</p> <p>Based on the literature and a review of museum catalogs, a synthesis was compiled of fossil slow loris occurrences in Java. Morphometric data on the mandible and maxilla of 11 fossil lorises were compared with a dataset of extant specimens to assess variation in size and shape.</p> <p>Five Holocene <em>Nycticebus</em> occurrences were identified in eastern Java. All specimens fall in the range of <em>N. javanicus</em> and <em>N. coucang</em>. The specimens from Hoekgrot, Gua Jimbe and Sampung suggest a closer affinity to <em>N. javanicus</em>. The fossils from Gua Jimbe and Hoekgrot gave values close to the largest <em>N. javanicus</em> specimens, but the (presumably older) Song Terus fossil was of average size.</p> <p>The distribution of <em>Nycticebus</em> suggests that it originally occurred throughout the island. The fossils are probably best identified as <em>N. javanicus</em> or <em>N. coucang</em>, but the Neolithic finds from Hoekgrot and Gua Jimbe are presumably <em>N. javanicus</em>. Size variation in <em>Nycticebus</em> was clinal, but although some large specimens were present, no evidence was found for size diminution during the Holocene.</p>
Dataset of the publication: Large-scale Characterization of Java Streams
<p>This repository contains all the data supporting the findings of the research article: <em>Large-scale Characterization of Java Streams</em>, including the list of all analyzed projects and their metadata. The repository contains two files as follows:</p> <p>1. database.tar: This archive contains all the data that supports the findings reported in the manuscript entitled “Large-scale Characterization of Java Streams” which is under revision.</p> <p>The data is provided as a database created via PostgreSQL and can be restored to an installation of PostgreSQL (version 12 or later) by following the next procedure:</p> <p>A. Using pgAdmin (https://www.pgadmin.org/):</p> <p>- Enter to the "Restore Dialog" (https://www.pgadmin.org/docs/pgadmin4/development/restore_dialog.html).</p> <p>- Select as "Format" the "Custom or tar" option.</p> <p>- When selecting the "Filename", browse to the folder where this archive was decompressed and select the file "database.tar".</p> <p>2. ListOfProjects.csv: This comma-separated value file contains the list of all analyzed projects and their metadata.</p>
Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation
<p><strong>Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation</strong></p> <p> </p> <p>This artifact contains:</p> <p>1. The binary folder. Its instruction can be found in [./binary/README.md](./binary/README.md).</p> <p>2. All classes used in the experiments [./classes.csv](./classes.csv).</p> <p>3. The experimental data zip file ([experimental_data.zip](./experimental_data.zip)).</p> <p> </p>
Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation
<p><strong>Artifact of Method-Level Java Class Splitter for Search-Based Unit Test Generation</strong></p> <p> </p> <p>This zip file is the artifact that contains the following:</p> <p>1. The binary. </p> <p>2. All classes used in the experiments.</p> <p>3. The experimental data.</p> <p> </p>
Artifact of Enhancing Search-Based Unit Test Generation with Method-Level Java Class Splitter
<p># Artifact of Enhancing Search-Based Unit Test Generation with Method-Level Java Class Splitter</p> <p> </p> <p>This artifact contains:</p> <p>1. The binary folder. Its instruction can be found in [./binary/README.md](./binary/README.md).</p> <p>2. All classes used in the experiments [./classes.csv](./classes.csv).</p> <p>3. The experimental data zip file ([experimental_data.zip](./experimental_data.zip)).</p>
The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach"
<p>The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach".<br> Please refer to README.md for detailed instructions.</p>
Artifact of Method-Level Java Class Splitter
<p><strong>Artifact of Method-Level Java Class Splitter</strong></p> <p> </p> <p>It contains:</p> <p>1. The binary folder. The instructions are in [./binary/README.md](./binary/README.md).</p> <p>2. All subjects used in the experiments [./classes.csv](./classes.csv).</p> <p>3. The experimental result zip file ([experimental_data.zip](./experimental_data.zip)).</p>
FIGURES 2a–d in Arthrostoma supriatnai sp. nov. (Nematoda: Ancylostomatidae) parasitic in Mydaus javanensis (Carnivora: Mephitidae) from Mount Ciremai, Java, Indonesia
FIGURES 2a–d. Scanning electron microscopy images of Arthrostoma supriatnai sp. nov. collected from Mydaus javanensis (Carnivora: Mephitidae) caught on Mount Ciremai, Java, Indonesia. a. Cephalic end, apical view; b. Cervical papilla; c. Vulva, ventral view; d. Female posterior end, lateral view. Scale bars: a, b. 10 µm; c. 5 µm; d. 100 µm. Abbreviations: am. amphid; cp. cephalic papillae.
FIGURES 1a– l in Arthrostoma supriatnai sp. nov. (Nematoda: Ancylostomatidae) parasitic in Mydaus javanensis (Carnivora: Mephitidae) from Mount Ciremai, Java, Indonesia
FIGURES 1a– l. Arthrostoma supriatnai sp. nov. collected from Mydaus javanensis (Carnivora: Mephitidae) caught on Mount Ciremai, Java, Indonesia. a. Anterior portion, lateral view; b. Cephalic end, apical view; c. Lateral views of buccal capsule; d. ventral views of buccal capsule; e. Female posterior end, lateral view; f. Vulva, lateral view; g. Egg; h. Bursa, ventral view; i. Genital cone; j. Dorsal ray, ventral view; k. Spicule, lateral view; l. Gubernaculum, ventral view. Scale bars: a, c, h, k. 100 µm; b, i. 10 µm; d, e, f. 50 µm; g, j, l. 25 µm. Abbreviations: de. duct of dorsal esophageal gland; dg. dorsal gutter; la. lancet.
Dataset and Tool for the Paper: Cross-Language Dependencies: An Empirical Study of Kotlin-Java
<ul> <li>Folder <strong>tool-to-extract-kotlin-java-dependencies</strong> contains a runnable jar of our tool to extract Kotlin-Java dependencies, please follow README in this folder to execute this tool.</li> <li>Folder <strong>accuracy-verification </strong>contains the source code of the ground-truth project, along with manual dependency inspections from file to file and from line to line. The extracted dependencies result of this project from our tool is also included in this folder.</li> <li>Folder <strong>rq1-dependencies-graph-for-each-project </strong>contains all dependency graphs of our 23 subjects in Json format.</li> <li>Folder<strong> rq1-circular-layout-for-each-project</strong> contains the circular layout of dependencies among source files of our 23 subjects.</li> <li>Folder <strong>rq2-maintenance-cost-for-file-with-out-kotlin-java-interactions </strong>contains <ul> <li><strong>details</strong> lists the #commits and loc for each type of source files in details <ul> <li>_java_in_cross_language.csv represents Java source files participating in Kotlin-Java interactions</li> <li>_java_not_in_cross_language.csv represents Java source files not participating in Kotlin-Java interactions </li> <li>_kt_in_cross_language.csv represents Kotlin source files participating in Kotlin-Java interactions</li> <li>_kt_not_in_cross_language.csv represents Kotlin source files not participating in Kotlin-Java interactions </li> </ul> </li> </ul> </li> <li>Folder <strong>rq3-common-mistakes </strong>contains 103 cases in the manual inspections.</li> </ul>
Data from: Carving out turf in a biodiversity hotspot: multiple, previously unrecognized shrew species co-occur on Java Island, Indonesia
Open the record for dataset details and reuse information.
Data from: Wood and non wood forest products of Central Java, Indonesia
Open the record for dataset details and reuse information.
Data from: Characterising bird-keeping user-groups on Java reveals distinct behaviours, profiles and potential for change
Open the record for dataset details and reuse information.
Genetic diversity and the origin of commercial plantation of Indonesian teak on Java Island
Open the record for dataset details and reuse information.
Data from: Revisiting the ichthyodiversity of Java and Bali through DNA barcodes: taxonomic coverage, identification accuracy, cryptic diversity and identification of exotic species
Open the record for dataset details and reuse information.
List of licenses discovered in the study of Potential Code Borrowing and License Violations in Java Projects on GitHub
<p>This is the list of licenses discovered in the study of Potential Code Borrowing and License Violations in Java Projects on GitHub. The licenses are ranged by the amount of files that they cover, there are a total of 95 different licenses. Where possible, the names are presented as identifiers at https://spdx.org/licenses/. "GitHub" stands for no license in the file or the project.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.