Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.9.0
Dataset results
26 results for “App Reviews”
AWARE: Dataset for Aspect-Based Sentiment Analysis of Apps Reviews
<p> </p> <p><em><strong>The peer-reviewed paper of AWARE dataset is published in ASEW 2021, and can be accessed through: <a href="http://doi.org/10.1109/ASEW52652.2021.00049">http://doi.org/10.1109/ASEW52652.2021.00049</a>. Kindly cite this paper when using AWARE dataset.</strong></em></p> <p> </p> <p>Aspect-Based Sentiment Analysis (ABSA) aims to identify the opinion (sentiment) with respect to a specific aspect. Since there is a lack of <em>smartphone apps reviews</em> dataset that is annotated to support the ABSA task, we present AWARE: <strong>A</strong>BSA <strong>W</strong>arehouse of <strong>A</strong>pps <strong>RE</strong>views.</p> <p>AWARE contains apps reviews from three different domains (Productivity, Social Networking, and Games), as each domain has its distinct functionalities and audience. Each sentence is annotated with three labels, as follows: </p> <ul> <li><strong>Aspect Term: </strong>a term that exists in the sentence and describes an aspect of the app that is expressed by the sentiment. A term value of “N/A” means that the term is not explicitly mentioned in the sentence.</li> <li><strong>Aspect Category:</strong> one of the pre-defined set of domain-specific categories that represent an aspect of the app (e.g., security, usability, etc.).</li> <li><strong>Sentiment:</strong> positive or negative.</li> </ul> <p><em>Note: games domain does not contain aspect terms.</em></p> <p>We provide a comprehensive dataset of 11323 sentences from the three domains, where each sentence is additionally annotated with a Boolean value indicating whether the sentence expresses a positive/negative opinion. In addition, we provide three separate datasets, one for each domain, containing only sentences that express opinions. The file named “AWARE_metadata.csv” contains a description of the dataset’s columns.</p> <p><strong>How AWARE can be used?</strong></p> <p>We designed AWARE such that it can be used to serve various tasks. The tasks can be, but are not limited to:</p> <ul> <li>Sentiment Analysis.</li> <li>Aspect Term Extraction.</li> <li>Aspect Category Classification.</li> <li>Aspect Sentiment Analysis.</li> <li>Explicit/Implicit Aspect Term Classification.</li> <li>Opinion/Not-Opinion Classification.</li> </ul> <p>Furthermore, researchers can experiment with and investigate the effects of different domains on users' feedback.</p>
A Replication Package of Learning Features that Predict Developer Responses for iOS App Store Reviews
<p>This replication package contains the dataset and script used in our paper "<em>Learning Features that Predict Developer Responses for iOS App Store Reviews.</em>" The paper has been accepted at the ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), 2020. For further modification and versioning of the dataset (as well as the preprint) please go to <a href="https://github.com/Kamonphop/ESEM20-Replication">https://github.com/Kamonphop/ESEM20-Replication</a></p>
SURF: Replication Package for: "What Would Users Change in My App? Summarizing App Reviews for Recommending Software Changes"
<p>Description of the content of folder "SURF_replication_package": 1) "Experiment I" contains: a) the folder "summaries" which contains all the html summaries generated through SURF and browsed by study participants involved in the Experiment I. b) the folder "XMLreviews" which contains, for each of the apps involved in the Experiment I, the corresponding XML file containing all the collected reviews for that app. These xml files have been used as input files for the SURF tool for generating the summaries contained in the "summaries" folder c) "Experiment_I_results.xlsx" which contains all the answers to our survey collected from the Experiment I participants.</p> <p>2) "Experiment II" contains: a) the folder "summaries" which contains the two html summaries generated through SURF and browsed by study participants in the Experiment II. b) the folder "XMLreviews" which contains, for each of the two apps involved in the Experiment II, the corresponding XML file containing all the collected reviews for that app. These xml files have been used as input of the SURF tool for generating the summaries contained in the "summaries" folder. c) "Experiment_II_results.xlsx" which contains all the user feedbacks extracted/validated by survey participants in the two sub-experiments. d) "Experiment_II_survey_answers.xlsx" which contains all the answers to our survey collected in the Experiment II participants.</p> <p>3) "Survey.pdf" which contains the pdf version of the survey performed by the participants</p> <p>4) "SURF_tool.zip" contains: a) "SURF.jar", which contains the class files of a prototypical implementation of SURF b) "README.txt" which contains the instructions to run the SURF tool c) the "lib" folder, which contains all the java libraries needed for running SURF.</p>
Unveiling Competition Dynamics in Mobile App Markets through User Reviews
<p>This replication package contains the datasets and evaluation results for the research titled <i>"<strong>Unveiling Competition Dynamics in Mobile App Markets through User Reviews"</strong>, </i>by Quim Motger, Xavier Franch, Vincenzo Gervasi and Jordi Marco.</p><p>Latest version of the full code is available at: <a href="https://github.com/quim-motger/app-market-analysis">https://github.com/quim-motger/app-market-analysis</a></p>
App-Review-Multiclass-Classification
<p>This is an app review classification dataset comprises 1679 manually labeled reviews and it is designed for multi-class classification tasks. Each review is categorized into one or more categories: Bug Report, Feature Request, and User Experience.</p>
Automatic Classification of Non-functional Requirements in App User Reviews Based on System Model and Artificial Intelligence
<p>This is the replication package for the paper: "Automatic Classification of Non-functional Requirements in App User Reviews Based on System Model and Artificial Intelligence". It contains the dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package in the following.</p> <p><strong>1. dataset folder</strong></p> <ul> <li>dataset_user_reviews.xlsx contains 1278 labelled non-requirement user reviews.</li> <li>readme.txt describes the meaning of the data in dataset_user_reviews.xlsx in detail.</li> </ul>
Empirical data for project: Mining user reviews of COVID contact-tracing apps
<p>Empirical data for project: Mining user reviews of COVID contact-tracing apps</p>
AH-CID: A Tool to Automatically Detect Human-Centric Issues in App Reviews
<p>The keyword list used in the paper to pre-filter the app reviews.</p>
Age-Inclusive Mobile App Reviews
Open the record for dataset details and reuse information.
dataset and script for app reviews classification
Open the record for dataset details and reuse information.
Exploring Transparency Concerns in Social Media App Reviews: A Human-LLM Perspective
<p>This is a dataset for the "Exploring Transparency Concerns in Social Media App Reviews: A Human-LLM Perspective." paper submitted to EASE2025.</p>
GLARE: Google Apps Arabic Reviews Dataset
<p>This paper introduces GLARE an Arabic Apps Reviews dataset collected from Saudi Google PlayStore. It consists of 76M reviews, 69M of which are Arabic reviews of 9,980 Android Applications. We present the data collection methodology, along with a detailed Exploratory Data Analysis (EDA) and Feature Engineering on the gathered reviews. We also highlight possible use cases and benefits of the dataset.</p>
Dataset: Gold standard dataset for explainability need detection in app reviews.
<p>We crawled 90,000 app reviews from both Google Play Store and Apple App Store, including reviews from both free and paid apps. These reviews were filtered for explainability needs, and after this process, 4,495 reviews remained. Among them, 2,185 reviews indicated an explanation need, while 2,310 did not. This resulting gold standard dataset was used to train and evaluate several machine learning models and rule-based approaches for detecting explanation needs in app reviews.</p> <p>The dataset includes both balanced and unbalanced evaluation sets, as well as the original crawled data from October 2023. In addition to machine learning approaches, rule-based methods optimized for F1 score, precision, and recall are also included.</p> <p>We provide several pre-trained machine learning models (including BERT, SetFit, AdaBoost, K-Nearest Neighbor, Logistic Regression, Naive Bayes, Random Forest, and SVM) along with training scripts and evaluation notebooks. These models can be applied directly or retrained using the included datasets.</p> <p>For further details on the structure and usage of the dataset, please refer to the README.md file within the provided ZIP archive.</p>
RARE((Repository for App review REfinement)
<p># This directory contains the Benchmark RARE Dataset and Code File as described in the paper:</p> <p><br>- RARE_Dataset: In this folder, we introduce RARE, a benchmark for App Review Refinement. This folder contains two subfolders named Gold_Corpus and Silver_Corpus.</p> <p>1. Gold_Corpus: In this folder, a corpus of 10,000 annotated reviews, collaboratively refined by software engineers and a large language model (LLM) sourced from 10 different application domains, is provided.</p> <p>2. Silver_Corpus: This folder includes a set of 10,000 automatically refined reviews using the best-performing model, Flan-T5, which was trained on 10,000 reviews from the gold corpus, forming the silver corpus.</p> <p> </p> <p><br>- Code_File: In this folder, all the code files used in the entire experiment and research are provided. This folder contains four subfolders named Data_Extraction, Refined_Review_Generation_through_Prompting, Model_Finetuning_and_Inferences, and Result_Evaluation.</p> <p>1. Data_Extraction: This folder contains 2 Python files named 'Google_Play_Store_Reviews_Extraction_from_10_different_App.py', which was used for extracting 10,000 raw reviews from the Google Play Store, and 'Apple_App_Store_Reviews_Extraction_from_10_different_App.py', which was used for extracting 10,000 raw reviews from the Apple App Store.</p> <p>2. Refined_Review_Generation_through_Prompting: This folder contain a Python file named 'Prompting_GPT_3.5_TURBO_For_Refined_Review_Generation.py', which was used to guide GPT-3.5-Turbo in generating refined versions of the raw reviews.</p> <p>3. Model_Finetuning_and_Inferences: This folder contains 16 Python files: one for fine-tuning and another for inference, each for eight models, including BART, Flan-T5, Pegasus, Llama-2, Falcon, Mistral, Orca-2, and Gemma.</p> <p>4. Result_Evaluation: This folder contains 2 Python files: 'Reference_free_Automatic_Metrics_Evaluation.py' for evaluating reference-free metrics such as FKGL, FKRE, LEN, and SS, and 'Reference_Based_Automatic_Metrics_Evaluation.py' for evaluating reference-based metrics such as SARI and BERTScore Precision.</p>
App Store Reviews
<p><strong>App Store Reviews</strong></p> <ul> <li>Apple App Store</li> <li>Google Play Store</li> </ul>
Explanation Needs in App Reviews: Taxonomy and Automated Detection
<p><strong>Replication package for our paper submission to RE 2023</strong></p> <p>It contains the following files:</p> <ol> <li>Our dataset of 5,564 app reviews that we manually labeled with respect to the tags "explanation need present" and "explanation need not present".</li> <li>Code/notebooks that we used for the training and evaluation of our explanation need detection approaches</li> <li>.bin file of our best performing model</li> </ol>
Systematic Review of Health App Gamification for Lifestyle Intervention Adherence
ClinicalTrials.gov study NCT04633070. IPD Sharing: NO. Countries: 1. Publications: 11.
Data from: Evidence assessing the diagnostic performance of medical smartphone apps: a systematic review and exploratory meta-analysis
Objective: The number of mobile applications addressing health topics is increasing. Whether these apps underwent scientific evaluation is unclear. We comprehensively assessed papers investigating the diagnostic value of available diagnostic health applications using in-built smartphone-sensors. Methods: Systematic Review - Medline, Scopus, Web of Science inclusive Medical Informatics and Business Source Premier (by citation of reference) were searched from inception until December 15th, 2016. Checking of reference lists of review articles and of included articles complemented electronic searches. We included all studies investigating a health application that used in-built sensors of a smartphone for diagnosis of disease. The methodological quality of 11 studies used in an exploratory meta-analysis was assessed with the QUADAS-2 tool and the reporting quality with the STARD statement. Sensitivity and specificity of studies reporting two-by-two tables were calculated and summarized. Results We screened 3'296 references for eligibility. Eleven studies, most of them assessing melanoma screening apps, reported 17 two-by-two tables. Quality assessment revealed high risk of bias in all studies. Included papers studied 1'048 subjects (758 with the target conditions and 290 healthy volunteers). Overall, the summary estimate for sensitivity was 0.82 (95 % confidence interval (CI); 0.56 to 0.94) and 0.89 (95 %CI; 0.70 to 0.97) for specificity. Conclusions The diagnostic evidence of available health apps on Apple's and Google's app stores is scarce. Consumers and healthcare professionals should be aware of this when using or recommending them.
Supplementary material 4 from: Howard L, van Rees CB, Dahlquist Z, Luikart G, Hand BK (2022) A review of invasive species reporting apps for citizen science and opportunities for innovation. NeoBiota 71: 165-188. https://doi.org/10.3897/neobiota.71.79597
Table S4. Mean dimension scores by app
Supplementary material 2 from: Howard L, van Rees CB, Dahlquist Z, Luikart G, Hand BK (2022) A review of invasive species reporting apps for citizen science and opportunities for innovation. NeoBiota 71: 165-188. https://doi.org/10.3897/neobiota.71.79597
Table S2. Reviewer correlation
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.