Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
725
datasets available to search
ShareScore release 0.9.0
Dataset results
725 results for “recommendation”
movies recommendations dataset using the "BestSimilar" database
<p>movies recommendations dataset using the "BestSimilar" database</p>
Automated Pattern-Based Recommendation for Improving API Operation Performance and Reliability in Cloud-Based Architectures
<p>The extensive use of APIs as the entry point to many Cloud-based applications has created challenging problems, especially concerning API quality properties such as performance and reliability. API best practices and patterns, such as bundling requests, rate limiting, or load balancing, have been proposed to solve these challenges. Unfortunately, no study investigating the impact of existing API practices and patterns on such quality properties exists beyond informal recommendations. In this paper, we fill this gap by proposing a pattern-based, automated recommendation approach to improve the performance and reliability of API operations. We provide a benchmark suite based on a realistic open-source microservice application to enable the automatic generation of comprehensive decision tree models. These models are then processed to generate API design recommendation algorithms to improve API operations regarding performance and reliability stored in catalogs for reuse. We validate our algorithms using extensive data sets generated by running the benchmark on a private cloud and AWS. For both environments, based on the decision tree models automatically generated from the measured data, API design recommendation algorithms have been calculated using our approach.</p>
Replication Package for ASE 2023 Paper "Personalized First Issue Recommender for Newcomers in Open Source Projects"
<p>This replication package contains a replication package for ASE 2023 paper titled "Personalized First Issue Recommender for Newcomers in Open Source Projects." This package includes a dataset of 68,858 issues from 100 GitHub projects, records of 123 manually labeled issue samples, and Python scripts for analyzing the data and evaluating models. The package is also stored in the GitHub repository <a href="https://github.com/mcxwx123/PFIRec">https://github.com/mcxwx123/PFIRec</a>.</p> <p>Required Environment</p> <p>We recommend setting up the required environment on a commodity Linux machine with at least 1 CPU Core, 8GB Memory, and 100GB empty storage space. Our experiments were conducted on an Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>Files and Replicating Results</p> <p>We used the GFI-bot database and the GitHub GraphQL API to collect features of 68,858 candidate issues and restore historical states of resolvers of 11,615 FIs (first issues).</p> <p>The followings are the files and replicating results:</p> <p>Dataset:</p> <p>The raw data of newcomer-issue pairs' features are stored in <code>ReplicationPackage/data/dataset_{bertmodel}_{num}.pkl</code>, where {bertmodel} is one of the four BERT-based language models: SIMCSE, RoBERTa, CodeBERT, and BERTOverflow, corresponding to the dataset whose textual features are extracted by one of the four language models. And {num} is 0 to 19, corresponding to the 20 chronological folds. The training sets of the GFI-Bot approach are contained in <code>ReplicationPackage/data/training_set_recgfi_simcse_{num}.pkl</code>. <code>ReplicationPackage/data/newcomerdata.json</code> contains first issues' title and description and their resolvers' total commit number and number of commits in the latest month, and <code>ReplicationPackage/data/processeddata.pkl</code> contains the 37 developers' features for the empirical study. <code>ReplicationPackage/data/isstexts.json</code> contains issues titles and descriptions for Stanik et al.'s approach.</p> <p>Python scripts:</p> <p><code>ReplicationPackage/empirical.py</code> is the script for reproducing all the results in Section III of the paper. <code>ReplicationPackage/model.py</code> is the script for reproducing all the results in Section IV of the paper.</p> <p>Records:</p> <p><code>ReplicationPackage/PFIs.csv</code> records the manually labeled issues for the empirical study.</p> <p>Figures:</p> <p>By running <code>ReplicationPackage/empirical.py</code> and <code>ReplicationPackage/model.py</code>, you can get all the figures in the fold <code>ReplicationPackage/figures/</code>. Besides the figures in the paper, <code>ReplicationPackage/figures/</code> also contains <code>typedis_{num}.png</code>, and <code>domaindis_{num}.png</code>, {num} is 1 to 4, representing additional results of newcomer features for Figure 4 in the paper.</p>
Deep Learning-Based Recommendation System: Systematic Review and Classification - Outputs
<p>The datasets provided are the outputs of the paper titled "Deep Learning-Based Recommendation System: Systematic Review and Classification." They encompass multiple outputs, including primary articles, domain-focused articles, technique mapping, and domain mapping for each category.</p>
KuaiSAR: A Unified Search And Recommendation Dataset
<p>The confluence of Search and Recommendation (S&R) services is a vital aspect of online content platforms like Kuaishou and TikTok. The integration of S&R modeling is a highly intuitive approach adopted by industry practitioners. However, there is a noticeable lack of research conducted in this area within the academia, primarily due to the absence of publicly available datasets. Consequently, a substantial gap has emerged between academia and industry regarding research endeavors in this field. To bridge this gap, we introduce the first large-scale, real-world dataset KuaiSAR of integrated Search And Recommendation behaviors collected from Kuaishou, a leading short-video app in China with over 300 million daily active users. Previous research in this field has predominantly employed publicly available datasets that are semi-synthetic and simulated, with artificially fabricated search behaviors. Distinct from previous datasets, KuaiSAR records genuine user behaviors, the occurrence of each interaction within either search or recommendation service, and the users’ transitions between the two services. This work aids in joint modeling of S&R, and the utilization of search data for recommenders (and recommendation data for search engines). Additionally, due to the diverse feedback labels of user-video interactions, KuaiSAR also supports a wide range of other tasks, including intent recommendation, multi-task learning, and long sequential multi-behavior modeling etc. We believe this dataset will facilitate innovative research and enrich our understanding of S&R services integration in real-world applications.</p>
Office Posture Analysis for Tailored Exercise Recommendations
<p>Reproducibility code of the paper with info about the dataset: <a href="https://github.com/GaetanoDibenedetto/healthrecsys24.git">Github</a><br><br>Zip strutures:<br>archives_data<br>│ <br>├───frames<br>│<br>├───keypoints<br>│ ap_1_250.json<br>│ ...<br>│ ...<br>│ ms_3_51581.json<br>│ <br>├───keypoints_augmented<br>│ augmented_ap_1_250.json<br>│ ...<br>│ ...<br>│ augmented_ms_3_51581.json<br>│ <br>│───labels<br>│ └───result<br>│ labels_for_train.csv<br>│<br>└───model_checkpoint<br> └───keypoint<br> yyyyMMddHHmmss.pkl <br><br>Note: The dataset contains only the keypoints extracted with the relative Human Pose Estimation Model used in the research paper</p>
2-Year Study of Vagus Nerve Stimulation for Higher-Grade Treatment-Resistant Depression: Clinical Outcomes and Policy Recommendation
ClinicalTrials.gov study NCT07097025. IPD Sharing: UNDECIDED. Countries: 1. Publications: 9.
The Primary Objective of This Study to Evaluate the Safety and Tolerability of IBI334 and Determine the Maximum Tolerated Dose (MTD) and the Recommended Phase 2 Dose (RP2D)and Anti Tumor Activity of I
ClinicalTrials.gov study NCT05774873. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Making Effective Human Papillomavirus (HPV) Vaccine Recommendations
ClinicalTrials.gov study NCT02377843. IPD Sharing: NO. Countries: 1. Publications: 2.
Effectiveness of Food-Based Recommendations for Minangkabau Women of Reproductive Age With Dyslipidemia
ClinicalTrials.gov study NCT04085874. IPD Sharing: NO. Countries: 1. Publications: 1.
Specialist Recommendation on FBC (Familial Breast Cancer) Chemoprevention Prescribing
ClinicalTrials.gov study NCT04058418. IPD Sharing: NO. Countries: 1. Publications: 2.
An Iodine Balance Study to Investigate the Recommended Iodine Intake in an Elderly Population
ClinicalTrials.gov study NCT06218043. IPD Sharing: NO. Countries: 1. Publications: 7.
Referral Recommendations for Axial Spondyloarthritis
ClinicalTrials.gov study NCT00383617. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Study on the Effects of a Recommendation Based Supply of Vitamin D3 in Healthy Volunteers
ClinicalTrials.gov study NCT01711905. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Remaja ASIK and Optimized Food-based Recommendations
ClinicalTrials.gov study NCT03946475. IPD Sharing: YES. Countries: 1. Publications: 7.
National Evaluation of the Adherence to Recommendations of Venous Thrombo Embolism Treatment in Cancer Patients
ClinicalTrials.gov study NCT01362933. IPD Sharing: Not stated. Countries: 1. Publications: 7.
Study to Evaluate the Safety and Tolerability of ABL501, and to Determine the Maximum Tolerated Dose (MTD) and Recommended Phase 2 Dose (RP2D) of ABL501 in Subjects With Any Progressive, Locally Advan
ClinicalTrials.gov study NCT05101109. IPD Sharing: NO. Countries: 1. Publications: 1.
Comparison Between Low Carbohydrate Diet and Traditionally Recommended Diabetic Diet in the Treatment of Diabetes Mellitus Type 2.
ClinicalTrials.gov study NCT01005498. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Gamified Family-based Health Exercise Intervention to Improve Adherence to 24-h Movement Behaviors Recommendations in Children.
ClinicalTrials.gov study NCT05741879. IPD Sharing: NO. Countries: 1. Publications: 2.
Impact of Currently Recommended Postnatal Nutrition on Neonatal Body Composition
ClinicalTrials.gov study NCT02622373. IPD Sharing: Not stated. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.