Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

33

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

33 results for “information retrieval”

Learn how ShareScore rates datasets ↗
zenodo44/100

Supporting Data for: Information Retrieval Interfaces in Virtual Reality - A Scoping Review Focused on Current Generation Technology

<p>This is the full data set of all reviewed research items obtained from Google Scholar, Web of Science and Scopus for the Scoping Literature Review&nbsp;<em><a href="https://doi.org/10.1371/journal.pone.0246398">Information Retrieval Interfaces in Virtual Reality - A Scoping Review Focused on Current Generation VR technology</a>.</em></p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Bangla Information Retrieval Test Collection | Revisiting Anwesha

<p>There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, no Gold standard dataset existed for Bangla IR until recently (https://zenodo.org/record/6583149). Our work expands the existing Gold standard dataset by creating 100 query document relevance pairs across a new test collection of 1000 documents. The corpus contains news articles&nbsp;from&nbsp;Ebela, Zee News and Anandabazar&nbsp;Patrika, Vikaspedia and various Bangla travel blogs.&nbsp;The definition of&nbsp;the complexity level of a query is described below:</p> <p>Complexity Level 1:&nbsp;The query contains exact words, phrases or sentence from the document.</p> <p>Complexity Level 2:&nbsp;The query is not present as it is in the document. There is a slight deviation.</p> <p>Complexity Level 3:&nbsp;The query is a generalised phrase capturing the overall story or the document&rsquo;s theme.</p> <p>Complexity Level 4: It is a general query not related to any specific document.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Supplementary Information: Exoplanet atmosphere retrievals in 3D using phase curve data with ARCiS: application to WASP-43b. Chubb and Min, A&A (2022).

<p>Supplementary information containing additional figures of the journal article &#39;Exoplanet atmosphere retrievals in 3D using phase curve data with ARCiS: application to WASP-43b&#39; by K. L. Chubb and&nbsp;M. Min, published in Astronomy &amp; Astrophysics (2022).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Bangla Information Retrieval Test Collection

<p>There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, there is no Gold standard dataset available to test the effectiveness of Bangla IR. So, we have created a document collection containing 182 short stories, novels, and essays written by Rabindranath Tagore11 and 1000 newspaper articles published in 2013 crawled from the Bangla newspaper Prothom Alo12. The collection contains 100 newspaper articles each from one of the ten categories: বাংলাদেশ/ Bānlādēśa(EN: `Bangladesh&#39;), খেলা/ khēlā(EN: `sports&#39;), বিজ্ঞান ও প্রযুক্তি/ bij&ntilde;āna ō prayukti(EN: `technology&#39;), বিনোদন/ binōdana(EN: `entertainment&#39;), আন্তর্জাতিক/ āntarjātika(EN: `international&#39;), অর্থনীতি/ arthanīti(EN: `economy&#39;), জীবনযাপন/ jībanayāpana(EN: `life-style&#39;), মতামত/ &nbsp;matāmata(EN: `opinion&#39;), শিক্ষা/ śikṣā(EN: `education&#39;) and আমরা/ āmarā(EN:`we-are&#39;). There are 94 queries in the dataset, 26 queries belonging to complexity levels 1 and 2, 19 queries&nbsp;in complexity level 3 and 23 queries in complexity level 4.&nbsp;The definition of&nbsp;the complexity level of a query is described below:</p> <p>Complexity Level 1:&nbsp;The query contains exact words, phrases or sentence from the document.</p> <p>Complexity Level 2:&nbsp;The query is not present as it is in the document. There is a slight deviation.</p> <p>Complexity Level 3:&nbsp;The query is a generalised phrase capturing the overall story or the document&rsquo;s theme.</p> <p>Complexity Level 4: It is a general query not related to any specific document.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Data archive for the peer-reviewed journal article "Information content and aerosol property retrieval potential for different types of in situ polar nephelometer data"

<p>Data archive accompanying the peer-reviewed journal article &quot;Information content and aerosol property retrieval potential for different types of in situ polar nephelometer data&quot;. This article was accepted for publication in the journal <em>Atmospheric Measurement Techniques</em> in 2022. The original contributions presented in the study are included in the article and its supplementary information. The GRASP-OPEN model was used to perform forward calculations: this model is publicly available on the official GRASP website (https://www.grasp-open.com/; last access: 14 September, 2022). The specific GRASP-OPEN model outputs that were used for the study are contained in this data archive.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

UnmixDB: A Dataset for DJ-Mix Information Retrieval

<p>A collection of automatically generated DJ mixes with ground truth, based on creative-commons-licensed freely available and redistributable electronic dance tracks.</p> <p>In order to evaluate the DJ mix analysis and reverse engineering methods, we created a dataset of excerpts of open licensed dance tracks and automatically generated mixes based on these.</p> <p>Each mix is based on a playlist that mixes 3 track excerpts beat-synchronously, such that the middle track is embedded in a realistic context of beat-aligned linear cross fading to the other tracks.<br> The first track&#39;s BPM is used as the seed tempo onto which the other tracks are adapted.</p> <p>Each playlist of 3 tracks is mixed 12 times with combinations of 4 variants of effects and 3 variants of time scaling using the treatments of the sox open source command-line program [http://sox.sourceforge.net].</p> <p>Each track excerpt contains about 20s of the beginning and 20s of the end of the source track. However, the exact choice is made taking into account the metric structure of the track. The cue-in region, where the fade-in will happen, is placed on the second beat marker starting a new measure, and lasts for 4 measures.&nbsp; The cue-out region ends with the 2nd to last measure marker. We assure at least 20s for the beginning and end parts. The cut points where they are spliced together is again placed on the start of a measure, such that no artefacts due to beat discontinuity are introduced.</p> <p>The UnmixDB dataset contains the ground truth for the source tracks and mixes in ASCII label format with tab-separated columns starttime, endtime, label.<br> For each mix, the start, end, and cue points of the constituent tracks are given, along with their BPM&nbsp; and speed factors.<br> We use the convention that the label starts with a number indicating which of the 3 source tracks the label refers to.</p> <p>The song excerpts are accompanied by their cue region and tempo information in .txt files in table format.</p> <p>Additionally, we provide the .beat.xml files containing the beat tracking results for the full tracks available from Sonnleitner et. al. 2016.</p> <p>Our DJ mix dataset is based on the curatorial work of Sonnleitner et. al. (ISMIR 2016), who collected Creative-Commons licensed source tracks of 10 free dance music mixes from Mixotic. We used their collected tracks to produce our track excerpts, but regenerated artificial mixes with perfectly accurate ground truth.</p> <p>The code used to create the dataset from the above is published at https://github.com/Ircam-RnD/unmixdb-creation, such that other researchers can create test data from other track collections or in other variants.</p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Sep 2018View details →
zenodo40/100

Fig. 1 Retrieved data distribution. A in Same information, new applications: revisiting primers for the avian COI gene and improving DNA barcoding identification

Fig. 1 Retrieved data distribution. A Distribution of published primers for the barcode region of the avian COI gene throughout the years. B Number of complete COI sequences available for each bird order

opencc-by-4.0Aug 2021View details →
zenodo36/100

Learning Unsupervised Knowledge-Enhanced Representations to Reduce the Semantic Gap in Information Retrieval (Evaluation datasets)

<p>This dataset contains all the runs, pools, plots and analyses to reproduce the results presented in the paper: &quot;Learning Unsupervised Knowledge-Enhanced Representations to Reduce the Semantic Gap in Information Retrieval&nbsp;&quot;, 2020.</p>

opencc-by-4.0Jun 2020View details →
zenodo36/100

Bilingual Dataset for Information Retrieval and Question Answering over the Spanish Workers Statute

<p>A bilingual dataset of questions and answers over a key document in Spanish labor law legislation is presented. The document contains 150 questions and their respective answers in the form of one part number from the 130 parts in which the Workers Statute is divided (articles and other provisions), and with the most relevant excerpt of information for the answer.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

ir_metadata: An Extensible Metadata Schema for Information Retrieval Experiments

<p>This dataset accompanies our work that introduces a metadata schema for TREC run files based on the PRIMAD model. PRIMAD considers essential components of computational experiments that possibly can affect reproducibility on a conceptual level. We propose to align the metadata annotations to the PRIMAD components. In order to demonstrate the potential of metadata annotations, we curated a dataset with run files derived from experiments with different instantiations of PRIMAD components and annotated these with the corresponding metadata. With this work, we hope to stimulate IR researchers to annotate run files and improve the reuse value of experimental artifacts even further.</p> <p>&nbsp;</p> <p>This archive contains the following data:</p> <ul> <li> <p><strong>demo.tar.xz</strong> : Selected annotated runs files that are used in the Colab demonstration.</p> </li> <li> <p><strong>metadata.zip</strong> : YAML files containing only the metadata annotations for each run.</p> </li> <li> <p><strong>runs.zip</strong> : The entire set of run files with annotations.</p> </li> </ul> <p>&nbsp;</p> <p>The annotated runs result from the following experiments:</p> <ul> <li> <p>Grossman and Cormack @ TREC Common Core 2017 <a href="https://trec.nist.gov/pubs/trec26/papers/MRG_UWaterloo-CC.pdf">Paper</a> |&nbsp;<a href="https://trec.nist.gov/">Source</a></p> </li> <li> <p>Grossman and Cormack @ TREC Common Core 2018 <a href="https://trec.nist.gov/pubs/trec27/papers/MRG_UWaterloo-CC.pdf">Paper</a> | <a href="https://trec.nist.gov/">Source</a></p> </li> <li> <p>Yu et al. @ TREC Common Core 2018 <a href="https://trec.nist.gov/pubs/trec27/papers/h2oloo-CC.pdf">Paper</a> | <a href="https://github.com/castorini/Anserini/blob/master/docs/runbook-trec2018-h2oloo.md">Source</a></p> </li> <li> <p>Yu et al. @ ECIR 2019 <a href="https://link.springer.com/chapter/10.1007/978-3-030-15712-8_26">Paper</a> | <a href="https://github.com/castorini/anserini/blob/master/docs/runbook-ecir2019-ccrf.md">Source</a></p> </li> <li> <p>Breuer et al. @ SIGIR 2020 <a href="https://dl.acm.org/doi/10.1145/3397271.3401036">Paper</a> | <a href="https://zenodo.org/record/3856042">Source</a></p> </li> <li> <p>Breuer et al. @ CLEF 2021 <a href="https://link.springer.com/chapter/10.1007/978-3-030-85251-1_5">Paper</a> | <a href="https://zenodo.org/record/4105885">Source</a></p> </li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Knowledge organization systems and their consequences for information retrieval

<p>Traditionally, research on knowledge organization systems (KOS) and information retrieval discussed the relative advantages or disadvantages of using controlled vocabularies versus free-text or intellectual indexing versus automatic indexing methods for indexing and search. Experiments and case studies variously showed the superiority of either approach without reaching a final conclusion on this seemingly basic question. As full-text indexing has become more possible and now prevalent, the discussion of the relative merits of KOS &ndash; not only&nbsp;as substitute but in combination with full-text &ndash; was not settled but continued with new challenges. With the advent of the Semantic Web, KOS (now appearing as ontologies) became important tools in new information retrieval applications and were pushed once again to the research forefront. With different disciplines working in the field, the terminology around KOS has become more and more ambiguous up to the point that tracing research in the literature is difficult &ndash; ironically something that traditional KOS have always tried to mitigate.<br> This paper summarizes recent discussions of the impact of KOS on information retrieval and attempts to show and unify different research strands from library science research on subject indexing, information retrieval and the Semantic Web. Whereas earlier impact studies on retrieval resulted in clearly measurable outcomes (for example changes in precision/recall), recent use of KOS in Semantic Web applications or other information systems has switched from pure search scenarios to exploration (browse) and contextualization, for which clear (and calculable) evaluation or quality standards and&nbsp;</p>

opencc-by-4.0Jul 2011View details →
zenodo36/100

Supplementary material for 'The MAP metric in Information Retrieval Fault Localization'

<pre># map_bench4bl This is the supplementary material, data, and evaluation source code for the paper &quot;The MAP metric in Information Retrieval Fault Localization&quot; by Thomas Hirsch and Birgit Hofer. ## Preliminaries ### Python environment - Python 3.8 - pandas - numpy - matplotlib ## Datasets The [Bench4BL](<em>https://github.com/exatoa/Bench4BL</em>) dataset has been used in this evaluation, with the addition of intermediate files taken from the [SABL](<em>http://dx.doi.org/10.5281/zenodo.4681242</em>) experiment performed on this Bench4BL dataset. All data used in our evaluation is included in this repository. However, if the data is to be re-imported directly from these benchmark and datasets they have to be downloaded first and their local paths have to be set in [paths.py](<em>paths.py</em>). ### Bench4BL The Bench4BL dataset was published with the paper &quot;Bench4BL: Reproducibility study on the performance of IR-based bug localization&quot; by Lee, J., Kim, D., Bissyand&eacute;, T.F., Jung, W. and Le Traon, Y.. The dataset can be obtained [here](<em>https://github.com/exatoa/Bench4BL</em>). Follow the steps described in the corresponding [README](<em>https://github.com/exatoa/Bench4BL/blob/master/README.md</em>) to set up the dataset. The Bench4BL dataset contains the _old subjects_ subdataset, containing 558 bugs from AspectJ, JDT, PDE, SWT, and ZXing that have been widely used in older IRFL studies. This _old subjects_ subdataset was used in answering our RQ1, as discussed below, the corresponding scripts use _old subjects_ in their name to highlight this. #### SABL The SABL dataset is the online appendix of the paper &quot;An Extensive Study of Smell-Aware Bug Localization&quot; by TTakahashi, A., Sae-Lim, N., Hayashi, S. and Saeki, M.. The dataset can be downloaded [here](<em>http://dx.doi.org/10.5281/zenodo.4681242</em>). The experiments in this dataset build on top of Bench4BL and intermediate files are provided in the datapackage. #### Rankings Rankings for BLIA, BRTracer, and BugLocator were produced by running these tools on Bench4BL locally. Rankings for AmaLgam and BLUiR were taken from the SABL experiment dataset. ## Structure ### Folders Bench4BL ground truths: - bench4bl_old_subjects_summary - bench4bl_summary Localization results of the included tools in Bench4BL: - bench4bl_localization_results - bench4bl_localization_results_sabl Target projects size metrics: - cloc_results - cloc_results_old_subjects Utility functions: - utils Output folders containing results, generated figures and tables: - results - results_old_subjects ### Scripts Scripts for re-importing data from Bench4BL and SABL datasets: - data_preparation_step_1_cloc_bench4bl.py - data_preparation_step_1_cloc_old_subjects_bench4bl.py - data_preparation_step_2_import_ground_truth_from_bench4bl.py - data_preparation_step_2_import_ground_truth_from_old_subjects_bench4bl.py - data_preparation_step_3_import_bench4bl_ranking_results.py - data_preparation_step_3_import_sabl_ranking_results.py Utilities: - paths.py - utils/bench4bl_utils.py - utils/Logger.py ### Evaluation scripts for the corresponding research questions: **Dataset analysis:** - rq_0_dataset_analysis_bench4bl_issues.py **RQ1: How big is the average ground truth in Bench4BL datasets, and what proportion of bugs have a ground truth containing multiple files?** - rq_1_bench4bl_ground_truth_size.py - rq_1_old_subjects_bench4bl_ground_truth_size.py RQ2: Do the IRFL tools included in Bench4BL truncate their results? - rq_2_ranking_lengths.py **RQ3: How strong is $AP_{asrd}$ overestimating $AP_{mb}$ for truncated BugLocator retrieval results on the Bench4BL dataset? RQ3a: How strong is $AP_{asrd}$ overestimating $AP_{mb}$ for truncated BugLocator retrieval results when considering the bloated ground truth issue found in Bench4BL?** - rq_3_truncating_BugLocator_rankings_bench4bl.py **RQ3b: How strong is $AP_{asrd}$ overestimating $AP_{mb}$ for truncated BugLocator retrieval results when undefined $AP$ values are simply ignored?** - rq_3b_undefined_ap_BugLocator_rankings_bench4bl.py ## Licence All code and results are licensed under [CCA v4](<em>https://creativecommons.org/licenses/by/4.0/</em>), according to LICENSE file. Other licences may apply for some tools and datasets contained in this repo: [cloc-1.92.pl](<em>https://github.com/AlDanial/cloc</em>) under GPL v2, [Bench4BL](<em>https://github.com/exatoa/Bench4BL</em>) and [SABL](<em>http://dx.doi.org/10.5281/zenodo.4681242</em>) under CCA 4.0.</pre>

opencc-by-4.0Apr 2023View details →
zenodo36/100

How doctors apply semantic components to specify search in work-related information retrieval

<p>Workplace searching is often context-specific and targets a &lsquo;right answer&rsquo; within some<br> domain-specific aspect of the search topic. We have developed the semantic component<br> (SC) model that allows searchers to specify a search within context-specific aspects of the<br> main topic of documents. The goal of our study was to gain insight into how family practice<br> physicians at sundhed.dk, a national healthcare portal in Denmark, applied the SC model<br> to formulate queries to solve work-related search tasks. The results showed that doctors<br> used the model purposively when choosing search facets and search concepts. They were<br> relatively consistent in their use. The findings provide promising evidence of the model&rsquo;s<br> potential usefulness.</p>

opencc-by-4.0Jul 2011View details →
zenodo36/100

Scientific information retrieval systems

<p>Analysis of four scientific information retrieval systems: Google Scholar, Semantic Scholar, Internet Archive Scholar and BASE</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

Supplementary Information to: Increasing the spatial resolution of cloud property retrievals from Meteosat SEVIRI by use of its high–resolution visible channel: Evaluation of candidate approaches with MODIS observations

<p>This repository contains the Python programmes and datasets used for</p> <p>the research described in the following paper:<br> https://doi.org/10.5194/amt-2019-334</p> <p>It has been prepared as supplementary information to the final paper.<br> <br> Note that the actual Cloud Property Retrieval (CPP) which would be<br> required for full reproducability of the results cannot be made<br> available by the authors, that the code included here is based on Python2, and that<br> paths to the dataset have to be adapted in the code.<br> <br> The repository consists of three separate parts/directories:<br> <br> * method: Python routines for generating the cloud property retrieval<br> &nbsp; input based on the Meteosat and MODIS data<br> <br> * analysis: Python routines for analysing the different experiments<br> &nbsp; from the CPP outputs, producing the figures and calculating the<br> &nbsp; comparison statistics<br> <br> * datasets: the various input and output datasets used for the<br> &nbsp; analyses of the paper</p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Retrieved studies for Impacts of the Adoption of Hybrid Work on Collaboration in Information Technology Teams

<p>Retrieved studies for Impacts of the Adoption of Hybrid Work on Collaboration in Information Technology Teams</p>

opencc-by-4.0Aug 2024View details →
ClinicalTrials.gov32/100

Retrieval of Economic Incentives and Information on Quality-of-care Indicators in Primary Care

ClinicalTrials.gov study NCT06829589. IPD Sharing: NO. Countries: 1. Publications: 10.

closedIPD-NOFeb 2026View details →
dryad28/100

Data from: Test collections for EHR-based clinical information retrieval

Objectives: To create test collections for evaluating clinical Information Retrieval (IR) systems and advancing clinical IR research. Materials and Methods: Electronic Health Records (EHR) data, including structured and free text data, from 45,000 patients who are a part of the Mayo Clinic Biobank cohort was retrieved from the clinical data warehouse. The clinical IR system indexed 42 million free-text EHR documents. The search queries consisted of 56 topics developed through a collaboration between Mayo Clinic and Oregon Health &amp; Science University. We described the creation of test collections, including a to-be-evaluated document pool using five retrieval models, and human assessment guidelines. We analyzed the relevance judgment results in terms of human agreement and time spent, and results of three levels of relevance, and reported performance of five retrieval models. Results: The two judges had a moderate overall agreement with a Kappa value of 0.49, spent a consistent amount of time judging the relevance, and were able to identify easy and difficult topics. The conventional retrieval model performed best overall on most topics while a concept-based retrieval model had better performance on the topics requiring conceptual level retrieval. Discussion: Information Retrieval can provide an alternate approach to leveraging clinical narratives for patient information discovery as it is less dependent on semantics. Our study showed the feasibility of test collections as well as challenges. Conclusion: The conventional test collections for evaluating the IR system show potential for successfully evaluating clinical IR systems with a few challenges to be investigated.

opencc-zeroJun 2019View details →
zenodo28/100

Biodiversity information retrieval across networked data sets

<p>Recording of presentation. An approach to integrate data through wrapping of various datasets stored in relational databases located on networked platforms. It is designed to overcome copyright problems in data sharing.</p>

opencc-ncJun 2009View details →
zenodo28/100

Supplemental Material for Predictive Reranking using Code Smells for Information Retrieval Fault Localization

<pre># Predictive Reranking using Code Smells for Information Retrieval Fault Localization This repository constitutes the supplementary material, data, and source code for the paper &quot;Predictive Reranking using Code Smells for Information Retrieval Fault Localization&quot;, by Thomas Hirsch and Birgit Hofer, 2023. Source code and results are also made available on GitHub: https://github.com/AmadeusBugProject/PredictiveRerankingUsingCodeSmellsForIRFL </pre> <pre>## Preliminaries ### Python environment - python=3.8 - pandas - numpy - joblib - scikit-learn==1.0.2 - keras - tensorflow - nltk - sentence-transformers - matplotlib - seaborn Conda files are located in the root directory of the repository, [conda_from_history.yml](<em>conda_from_history.yml</em>). ### Datasets The [Bench4BL](<em>https://github.com/exatoa/Bench4BL</em>) dataset was used in our evaluation. All data necessary for our machine learning and localization experiments is included in this repository. However, if the data is to be re-imported and recalculated directly from Bench4BL: Bench4BL has to be downloaded and paths to the benchmark root set accordingly in [paths.py](<em>paths.py</em>). BugLocator, BRTracer, and BLIA have to be run on the Bench4BL dataset using the scripting provided by the benchmark. PMD has to be installed in version 6.45.0 and path to PMD set accordingly in [paths.py](<em>paths.py</em>). # Structure of this repository ## General utility functions - [constants.py](<em>constants.py</em>) Contains parameters for the NN model, and other parameters. - [paths.py](<em>paths.py</em>) Contains paths to external datasources and tools, e.g. Bench4BL and PMD. - [utils/bench4bl_utils.py](<em>utils/bench4bl_utils.py</em>) Helper functions for resolving paths to Bench4BL benchmark. - [utils/dataset_utils.py](<em>utils/dataset_utils.py</em>) Helper functions for loading datasets and performing dataset splits. - [utils/Logger.py](<em>utils/Logger.py</em>) Logging. - [utils/nn_classifier.py](<em>utils/nn_classifier.py</em>) NN classifier model. - [utils/scoring_utils.py](<em>utils/scoring_utils.py</em>) Metrics. - [utils/stats_utils.py](<em>utils/stats_utils.py</em>) Wrapper methods for statistical tests. ## Experiment ### Dataset setup and preparation The following scripts are responsible to create, import, and set up data that is used in our experiments. The produced data is already part of this repository, the execution of these scripts is therefore only necessary when data is to be re-imported from the Bench4BL repository. - [a00_pmd_bench4bl.py](<em>a00_pmd_bench4bl.py</em>) Runs PMD on all projects and versions contained in the Bench4BL dataset. The utilized ruleset is defined in [all_java_ruleset.xml](<em>all_java_ruleset.xml</em>). Results are stored in [pmd_results](<em>pmd_results</em>). - [a01_vectorize_pmd.py](<em>a01_vectorize_pmd.py</em>) Creates csv vectors from PMD outputs. - [a02_cloc_bench4bl.py](<em>a02_cloc_bench4bl.py</em>) Runs cloc on all projects and versions contained in the Bench4BL dataset. Only Java files are considered. Results are stored in [cloc_results](<em>cloc_results</em>). - [a02_pmd_usage_in_bench4bl_projects.py](<em>a02_pmd_usage_in_bench4bl_projects.py</em>) Searches for occurrence of PMD in the build files of all projects and versions contained in the Bench4BL dataset. Results are stored in [pmd_usage_in_bench4bl_projects](<em>pmd_usage_in_bench4bl_projects</em>). - [a03_import_bugs_from_bench4bl.py](<em>a03_import_bugs_from_bench4bl.py</em>) Imports textual bug reports and corresponding fixed files ground truth from Bench4BL. Results are stored in [bench4bl_summary](<em>bench4bl_summary</em>). - [a04_normalize_smells_by_loc.py](<em>a04_normalize_smells_by_loc.py</em>) Normalizes the smell vectors for each file with its LOC count. Results are stored in [pmd_results](<em>pmd_results</em>). - [a05_bench4bl_file_features.py](<em>a05_bench4bl_file_features.py</em>) Creates feature vectors for each bug report from PMD smells. Results are stored in [bug_smell_vectors](<em>bug_smell_vectors</em>). - [a05_bench4bl_ranking_results.py](<em>a05_bench4bl_ranking_results.py</em>) Imports the results of BugLocator, BRTracer, and BLIA from the Bench4BL benchmark. Results are stored in [bench4bl_localization_results](<em>bench4bl_localization_results</em>). - [a09_bench4bl_stackoverflow_mpnet.py](<em>a09_bench4bl_stackoverflow_mpnet.py</em>) Creates document embeddings for all textual bug reports using the [stackoverflow_mpnet-base](<em>https://huggingface.co/flax-sentence-embeddings/stackoverflow_mpnet-base</em>) model. Results are stored in [stackoverflow_mpnet_embeddings](<em>stackoverflow_mpnet_embeddings</em>). ### Preliminary experiments and dataset splitting The following scripts perform data set splitting, and the preliminary experiments used for feature selection as discussed in Section VI of the paper. - [b00_analyze_most_promising_smells.py](<em>b00_analyze_most_promising_smells.py</em>) Assumes a perfect smell oracle (by using the known smells of the ground truth files) and applies it to rerank the IRFL tools outputs on the Classification Training Set (the older half of versions in the dataset). Then evaluates the achievable localization performance increase for each smell group. Results are stored in [h_analyze_most_promising_smells](<em>h_analyze_most_promising_smells</em>). - [c00_make_dataset_splits_bootstrap.py](<em>c00_make_dataset_splits_bootstrap.py</em>) Performs dataset splitting. Splits are performed on a temporal ordering of versions of each contained software project. Data is greedily split into 50/25/25, resulting in a Classification training set (used for NN model training), a ranking training set (used to estimate weights for linear combination of smell distance and IRFL suspicousness scores), and a test set (used for evaluating the localization performance achievable by our pipeline). Bootstrapping is applied by resampling fractions of 0.8, 20 times, resulting in 20 sets of the three splits to be used in the following eperiments. Datasets are stored in [p_FINAL_Bench4BL/p_model_for_smell_classification_bootstrap](<em>p_FINAL_Bench4BL/p_model_for_smell_classification_bootstrap</em>) for the full Bench4BL dataset, for the single project experiments please refer to p_FINAL_CAMEL, p_FINAL_HBASE, and p_FINAL_ROO accordingly. - [c01_full_dataset_stats.py](<em>c01_full_dataset_stats.py</em>) Calculates various statistics on the created datasets. Results are stored in [px_summary_dataset](<em>px_summary_dataset</em>). - [d00_model_for_smell_classification_performance_all_groups_bootstrap.py](<em>d00_model_for_smell_classification_performance_all_groups_bootstrap.py</em>) Trains a NN model on the Classification data set and evaluates its classification performance on the Ranking training set. This is performed for all 20 bootstrap iterations, the resulting data is stored in [p_FINAL_Bench4BL/p_model_for_smell_classification_bootstrap](<em>p_FINAL_Bench4BL/p_model_for_smell_classification_bootstrap</em>) for the full Bench4BL dataset, for the single project experiments please refer to p_FINAL_CAMEL, p_FINAL_HBASE, and p_FINAL_ROO accordingly. - [d01_classification_performance_evaluation_for_feature_selection_preliminariy_all_smell_groups.py](<em>d01_classification_performance_evaluation_for_feature_selection_preliminariy_all_smell_groups.py</em>) Creates summary and bootstrap statistics from the previous step. Results are stored in [p_FINAL_Bench4BL/p_summary_classification](<em>p_FINAL_Bench4BL/p_summary_classification</em>) for the full Bench4BL dataset. ### Localization experiments The following scripts perform our actual localization experiments. These scripts are applied to bootstrapped dataset splits. Results are stored in [p_FINAL_Bench4BL](<em>p_FINAL_Bench4BL</em>) for the full Bench4BL dataset, for the single project experiments please refer to p_FINAL_CAMEL, p_FINAL_HBASE, and p_FINAL_ROO accordingly. For a detailed experiment setup we refer to our paper. - [e01_model_for_localization_bootstrap.py](<em>e01_model_for_localization_bootstrap.py</em>) Trains NN models for smell classification and performs predictions on the corresponding test sets. Results are stored in [p_FINAL_Bench4BL/p_model_for_smell_classification_bootstrap](<em>p_FINAL_Bench4BL/p_model_for_smell_classification_bootstrap</em>). - [e02_ranking_training_bootstrap.py](<em>e02_ranking_training_bootstrap.py</em>) Performs reranking of the Ranking training set based on predicted smells by the model created in the previous step. Score combination is performed by linear combination of IRFL tools suspicousness scores and smell distances calculated based on our predictions. Results are stored in [p_FINAL_Bench4BL/p_ranking_training_bootstrap](<em>p_FINAL_Bench4BL/p_ranking_training_bootstrap</em>). - [e03_get_best_weigths_per_project_bootstrap.py](<em>e03_get_best_weigths_per_project_bootstrap.py</em>) Evaluates the outputs of the previous steps in order to pick the best weights for each project and IRFL tool. Results are stored in [p_FINAL_Bench4BL/p_ranking_training_bootstrap](<em>p_FINAL_Bench4BL/p_ranking_training_bootstrap</em>). - [e04_ranking_test_project_wise_bootstrap.py](<em>e04_ranking_test_project_wise_bootstrap.py</em>) Uses predictions of the final NN model and the weights obtained from the previous step to perform rerankings on the Test set. Results are stored in [p_FINAL_Bench4BL/p_ranking_test_proejct_wise_bootstrap](<em>p_FINAL_Bench4BL/p_ranking_test_proejct_wise_bootstrap</em>). ### Result collection and evaluation The following scripts calculate final scores and statistics from the 20 bootstrap iterations of the previous block of scripts. - [f00_bootstrap_summary_compare_map_and_ttest.py](<em>f00_bootstrap_summary_compare_map_and_ttest.py</em>) Calculates localization performance using the MAP metric and performs statistical tests. Results are stored in [p_FINAL_Bench4BL/p_summary_bootstrap](<em>p_FINAL_Bench4BL/p_summary_bootstrap</em>). - [f01_boostrap_summary_compare_classification_performance.py](<em>f01_boostrap_summary_compare_classification_performance.py</em>) Calculates classifier performance of our final model. Results are stored in [p_FINAL_Bench4BL/p_summary_bootstrap](<em>p_FINAL_Bench4BL/p_summary_bootstrap</em>). - [f02_bootstrap_summary_model_classification_performance_eval_test_set_for_all_projects.py](<em>f02_bootstrap_summary_model_classification_performance_eval_test_set_for_all_projects.py</em>) Calculates classifier performance and project wise classifier performance of our final model. Results are stored in [p_FINAL_Bench4BL/p_summary_classification_test_set](<em>p_FINAL_Bench4BL/p_summary_classification_test_set</em>). ### Further analysis The following scripts collect statistics and results to create latex tables and additional analysis used in our paper. - [g00_project_wise_perf_stats.py](<em>g00_project_wise_perf_stats.py</em>) Creates overview latex table comparing the Bench4BL and single project trained pipelines. Results are stored in [px_summary_performance](<em>px_summary_performance</em>). - [g01_project_multiple_file_smell_distances.py](<em>g01_project_multiple_file_smell_distances.py</em>) Analyses smell distances within each bug&#39;s ground truth files. Results are stored in [px_summary_dataset](<em>px_summary_dataset</em>). - [g02_performance_correlation_analysis.py](<em>g02_performance_correlation_analysis.py</em>) Performs correlation analysis of our pipeline&#39;s MAP localization performance, classification performance, and smell distance measures from previous step. Results are stored in [p_FINAL_Bench4BL/p_summary_correlations](<em>p_FINAL_Bench4BL/p_summary_correlations</em>) for the full Bench4BL dataset. ## Results - [pmd_catalogue/all_smells.json](<em>pmd_catalogue/all_smells.json</em>) lists all PMD smells and associated groups that occur in the dataset. - [px_summary_dataset](<em>px_summary_dataset</em>) contains statistics and information about the utilized dataset. Results for preliminary experiments for feature selection: - [h_analyze_most_promising_smells](<em>h_analyze_most_promising_smells</em>) contains the results of our preliminary experiment into each smell group&#39;s information content towards localization. - [p_FINAL_Bench4BL/p_summary_classification](<em>p_FINAL_Bench4BL/p_summary_classification</em>) contains the results of our preliminary experiments into the classifiability of smell groups from textual bug reports. Results of our localization eperiments: - [p_FINAL_Bench4BL/p_summary_bootstrap](<em>p_FINAL_Bench4BL/p_summary_bootstrap</em>) contains the results of our localization experiments, project wise MAP performance summary can be found in [p_FINAL_Bench4BL/p_summary_bootstrap/project_scores_tool_wise_others.tex](<em>p_FINAL_Bench4BL/p_summary_bootstrap/project_scores_tool_wise_others.tex</em>). - [p_FINAL_Bench4BL/p_summary_classification_test_set](<em>p_FINAL_Bench4BL/p_summary_classification_test_set</em>) contains the results of our classification performance analysis of the predictions used in localization. A project wise classification performance summary can be found in [p_FINAL_Bench4BL/p_summary_classification_test_set/macro_average_classification_performances_per_projectother.tex](<em>p_FINAL_Bench4BL/p_summary_classification_test_set/macro_average_classification_performances_per_projectother.tex</em>) ## Licence All code and results are licensed under [AGPL v3](<em>https://www.gnu.org/licenses/agpl-3.0.html.en</em>), according to LICENSE file. Other licences may apply for some tools and datasets contained in this repo: [cloc-1.92.pl](<em>https://github.com/AlDanial/cloc</em>) under [GPL v2](<em>https://www.gnu.org/licenses/old-licenses/gpl-2.0.en.html</em>), and data originating from [Bench4BL](<em>https://github.com/exatoa/Bench4BL</em>) under [CCA 4.0](<em>https://creativecommons.org/licenses/by/4.0/</em>). </pre>

openapgl-v3Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record