Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
100
datasets available to search
ShareScore release 0.9.0
Dataset results
100 results for “request”
Dependabot and Security Pull Requests
<p>This deposit contains four (4) main datasets that were used in the study "<em>Dependabot and Security Pull Requests: Large Empirical Study</em>" (<a href="https://doi.org/10.1007/s10664-024-10523-y" target="_blank" rel="noopener">Link</a>). Each dataset is described as follows :</p> <ol> <li><strong>Dataset (1) - Dependency Update</strong> : This dataset concerns issues related to pull requests (PRs) that were created by both users and bots to manage dependency updates in GitHub projects. The search was based on the keywords "Dependency, Update" in the title, body or comment of a PR created in the time period between 26/05/2017 and 15/06/2021 for the 1st partition, and between 01/01/2023 and 30/09/2023 for the 2nd partition. We obtained a total of 6,573,489 PR-related issues belonging to a total of 927,007 repositories for partition (1); and for partition (2), we obtained a total of 3,342,829 PR-related issues belonging to a total of 816,028 repositories.</li> <li><strong>Dataset (2) - Dependabot Security PRs</strong> : The second dataset is related to PRs created by Dependabot to handle security vulnerabilities in project dependencies. In our search, we look for PR-related issues created by "Dependabot-preview" or "Dependabot" and with the label "security", also created during the time period between 26/05/2017 and 30/09/2023. With these parameters, our results consist of 422,388 issues from 47,987 repositories.</li> <li><strong>Dataset (3) - Manual Security PRs</strong> : For this dataset, we were interested in PRs created only by users to handle security vulnerabilities. The search consists of finding the keywords "Dependency, Vulnerable" in the title, body or comment of a PR created in the time period between 26/05/2017 and 30/09/2023. We only consider pull requests created by authors with the type "user". The final results include a total of 186,186 issues for 60,758 repositories.</li> <li><strong>Dataset (4) - Bots' Security PRs</strong> : This dataset is related to PRs created by several bots to handle security vulnerabilities in project dependencies. In the search query, we look for PR-related issues where the keywords "Dependency", and "Security", and "Vulnerability" are mentioned in the title, body, or comment of the PR. These PRs are created by one of the following bots: "Snyk", "Renovate", "Greenkeeper", or "Depfu", also created during the time period between 26/05/2017 and 30/09/2023. The obtained results for the 4 bots consists of a collection of 628,495 PR-related issues in a total of 105,342 repositories.</li> </ol> <p>We also included :</p> <ul> <li><strong>Derived Sample</strong> : This sample contains the data that was selected and extracted to conduct our manual qualitative analysis, and the manual feature extraction.</li> </ul>
"@alex, this fixes #9": Analysis of Referencing Patterns in Pull Request Discussions
<p>This publication consists of a dataset of 7k references manually identified in 450 Pull request (PR) discussion threads sampled from GitHub in CSV format. In addition to the dataset, it also contains R code files which were written to analyze this dataset statistically. This dataset is released under the research, which is accepted for publication at CSCW 2021 conference, titled "@alex, this fixes #9": Analysis of Referencing Patterns in Pull Request Discussions".</p> <p><strong>Paper Abstract</strong></p> <p>Pull Requests (PRs) are a frequently used method for proposing changes to source code repositories. When discussing proposed changes in a PR discussion, stakeholders often reference a wide variety of information objects for establishing shared awareness and common ground. Previous work has not considered how referential behavior impacts collaborative software development via PRs. This knowledge gap is the major barrier in evaluating the current support for referencing in PRs and improving them. We conducted an explorative analysis of ~7K references, collected from 450 public PRs on GitHub, and constructed taxonomies of referent types and expressions. Using our annotated dataset, we identified several patterns in the use of references. Referencing source code elements was prevalent but the authoring interface lacks support for it. Three classes of contextual factors influence referencing behaviors: referent type, discussion thread, and project attributes. Referencing patterns may indicate PR outcomes (e.g., merged PRs frequently reference issues, users, and tests). We conclude with design implications to support more effective referencing in PR discussion interfaces.</p>
French trainset for chatbots dealing with usual requests on bank cards
<p><strong>[EN] French training dataset for chatbots dealing with usual requests on bank cards.</strong></p> <ul> <li><strong>Description</strong>: This dataset represents examples of common customer requests relating to bank cards management. It can be used as a training set for a small chatbot intended to process these usual requests.</li> <li><strong>Content</strong>: The questions are asked in French. The dataset is divided into 10 intents of 100 questions each, for a total of 1 000 questions.</li> <li><strong>Intents scope</strong>: Intents are constructed in such a way that all questions arising from the same intention have the same response or action. The scope covered concerns: loss or theft of cards; the swallowed card; the card order; consultation of the bank balance; insurance provided by a card; card unlocking; virtual card management; management of bank overdraft; management of payment limits; management of contactless mode.</li> <li><strong>Origin</strong>: Intents scope is inspired by a chatbot currently in production, and the wording of the questions are inspired by the usual customers requests.</li> </ul> <p><br> <strong>[FR] Jeu d'entraînement en français d'assistants conversationnels traitant des demandes courantes sur les cartes bancaires.</strong></p> <ul> <li><strong>Description </strong>: Cet ensemble de données représente des exemples de demandes usuelles des clients concernant la gestion des cartes bancaires. Il peut être utilisé comme jeu d'entraînement pour un assistant conversationnel destiné à traiter ces demandes courantes.</li> <li><strong>Contenu </strong>: Les questions sont formulées en français. L'ensemble de données est divisé en 10 intentions de 100 questions chacune, pour un total de 1 000 questions.</li> <li><strong>Périmètre des intentions</strong> : Les intentions sont construites de telle manière que toutes les questions issues d'une même intention ont la même réponse ou action. Le périmètre couvert concerne : la perte ou le vol de cartes ; la carte avalée ; la commande des cartes ; la consultation du solde bancaire ; l'assurance fournie par une carte ; le déverrouillage de la carte ; la gestion de cartes virtuelles ; la gestion du découvert bancaire ; la gestion des plafonds de paiement ; la gestion du mode sans contact.</li> <li><strong>Origine </strong>: Le périmètre des intentions est inspiré par un chatbot actuellement en production, et la formulation des questions est inspirée de demandes courantes de clients.</li> </ul>
ChatGPT's answers to requests about the librarian profession
<p>Dataset containing questions and answers about the library profession requested to ChatGPT.</p>
US Army Aviation air movement request problem instances
<p>Although lacking the same preeminent status of air assault planning, air movement operations comprise a majority of Army utility and cargo helicopter combat aviation operations in terms of volume of customers and the endless appetite for rapid movement of troops across the battlespace. The data provided enabled research and development of a US Army Aviation air movement mission planning model to assist the mission planner by rapidly providing courses of action based on the commander's priorities. Features of the problem and the data provided include priority demand, multi-node refueling, aircraft and passenger time windows, maximum passenger transportation time, and the minimization of unsupported demand, aircraft utilization, and total flight time. The mathematical model provided is an extension of the dial-a-ride problem (DARP) that will coordinate air mission requests (AMRs) at the aviation task force-level or lower to generate courses of action that optimize helicopter fleet resourcing and routing decisions against mission variables, while supporting the optimal number of AMRs that sustain combat power over time. </p>
US Army Aviation air movement request problem instances
Open the record for dataset details and reuse information.
Dataset - How do you propose your code changes? Empirical Analysis of Affect Metrics of Pull Requests on GitHub
<p>This package contains the raw open data for the study </p> <p>Marco Ortu, Giuseppe Destefanis, Daniel Graziotin, Michele Marchesi, Roberto Tonelli. 2020. How do you propose your code changes? Empirical Analysis of Affect Metrics of Pull Requests on GitHub. Under Review.</p> <p>The dataset is based on GHTorrent dataset:</p> <p>Georgios Gousios. 2013. The GHTorent dataset and tool suite. In Proceedings of the 10th Working Conference on Mining Software Repositories (MSR ’13). IEEE Press, 233–236</p> <p>And released with the same license (CC BY-SA 4.0).</p>
Data supplementing the conference paper "Who you gonna call? Analyzing web requests in Android applications", 14th International Conference on Mining Software Repositories 2017.
<p>This repository contains the data supplementing the paper:</p> <p>M. Rapoport, P. Suter, E. Wittern, O. Lhótak, J. Dolby, "Who you gonna call? Analyzing web requests in Android applications", MSR 2017.</p> <p>A detailed description of the data is included in the archive in README.md.</p>
Parmalat Scheduling Production Requests
<div> <div> <div> <p>The dataset provided in this repository is an anonymized collection of Scheduling Production Requests data, gathered by Parmalat from October to December 2023. Each entry in the dataset is a JSON object with the following keys: <code>prodQty</code> (production quantity), <code>prodPriority</code> (production priority), <code>prodArea</code> (anonymized production area), <code>lastChangeAccepted</code> (timestamp of the last accepted change), <code>prodUM</code> (anonymized unit of measure of quantity), and <code>itemFormat</code> (anonymized item format). The dataset provides valuable insights into the production scheduling process, while ensuring the confidentiality of sensitive information through anonymization performed by the knowlEdge Data Quality Assurance component. The data has been originally collected by the knowlEdge Data Collection Platform and stored within the knowlEdge Historical Data Storage. This dataset can be a valuable resource for researchers and analysts studying production scheduling patterns and strategies. Please note that all identifiers in the dataset have been anonymized for privacy reasons.</p> </div> </div> </div>
Terms from the SPARC Term Request Pipeline (Oct 2019 - Aug 2021)
<p>Terms analyzed for the paper "Extending and using anatomical vocabularies in the Stimulating Peripheral Activity to Relieve Conditions (SPARC) project". These terms were submitted by SPARC investigators to the SPARC term request pipeline. The SPARC anatomical term request pipeline is an iterative curation process that consists of three main steps: term request, term review, and term engineering. <strong> </strong>This work was supported by NIH grant 3OT2OD030541 from the Office of the Director through the Stimulating Peripheral Activity to Relieve Conditions (SPARC) program.</p>
Towards Just-In-Time Feature Request Approval Prediction
<p>This is the dataset of our publication "Towards Just-In-Time Feature Request Approval Prediction", which consists of feature requests collected from Sourceforge.net. </p>
Energy consumption, execution time and fail requests rate of a proactive energy-aware auto-scaling solution for edge-based infrastructures applied to real-world workload.
<p>Spreadsheet of the results obtained with our horizontal auto-scaling proposal presented in "A proactive energy-aware auto-scaling solution for edge-based infrastructures". In that research, we present a proactive horizontal auto-scaling framework for edge infrastructures, which considers both the base (idle) and dynamic (due to application execution) energy consumption of edge nodes and the node scaling mechanism. Simulations were performed with the EdgeCloudSim simulator with a workload provided by Shanghai Telecom and the results show up to a 92.5% decrease in energy consumption, a failed request rate of up to 0%, and reasonable execution times of the auto-scaling process for different problem sizes.</p> <p>Proactive auto-scaling mechanisms in edge-based infrastructures can anticipate user service requests by allocating computing resources while supporting the quality of service needed by a vast range of applications requiring, e.g., a low latency or response time. </p> <p>This work is supported by the European Union's H2020 research and innovation program under grant agreement DAEMON 101017109 and by the projects co-financed by FEDER funds LEIA UMA18-FEDERJA-15, MEDEA RTI2018-099213-B-I00 (MCI/AEI) and RHEA P18-FR-1081.</p>
GitHub Pull Request Demonstration by Alexandr Smagin
<p>In this demonstration, NLU student Alexandr Smagin walks us through pull requests for the TOPS SCHOOL GitHub repository. You can watch the video below or find a link in the 'Additional details' section.</p>
Dataset for ESE submission "Pull Request Latency Explained: An Empirical Overview"
<p>This is the dataset for ESE submission "Pull Request Latency Explained: An Empirical Overview".</p> <p>For research purpose, if you need `pull request id`, please request <a href="https://zenodo.org/record/7299639#.Y2mQ-HpBwUE">the column</a>.</p>
Pull Request Classification - Replication Package
<p><strong>Pull Request Classification - Replication Package:</strong> Contains two .java files (Heuristics.java and a helper class, namely Formatter.java) for the classification of pull request into the eight categories (i.e., bug, feature, test, resources, refactoring, merge, deprecate and others). Additionally, the results of the classification are presented in .csv files.</p>
A Randomized Trial of the Early Referral and Request Approach (ERRA) Intervention to Increase Consent to Organ Donation
ClinicalTrials.gov study NCT02138227. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Measuring Unique Changes: How do Distinct Changes Affect the Size and Lifetime of Pull Requests?
<p>Size metrics are commonly cited features in studies that analyze influencing factors on pull request lifetime. These metrics are also important for integrators, as based on them, they may prefer to prioritize pull requests easy to assess. However, code changes that form pull requests may not be unique, and repetitive changes may represent less complexity than expected by considering size metrics like lines of code or source files. The goal of this study is to analyze the influence of unique changes over pull requests relative to its size and lifetime. We collected data from 83,000+ pull requests of 26 projects hosted on GitHub. Also, we proposed a metric called unique changes rate to measure the proportion of unique changes over the total changes made by a pull request. We conducted experiments with Random Forest regression models and association rules to examine the influence of unique changes rate. Results show that unique changes have more influence over the lifetime of large pull requests, is determined mainly by the number of source files, and low levels of unique changes rate affect more the relationship between pull request size and lifetime than high levels. We conclude that unique changes can figure as an interesting feature in the context of the pull request lifetime. Results indicate that unique changes may increase or decrease the influence of pull request size on its lifetime. Our work has implications for researchers and core team members in software projects since unique changes represent helpful information.</p>
On the Use of Dependabot Security Pull Requests
<p>This dataset contains the data files used to analyze our RQs in the manuscript "On the Use of Dependabot Security Pull Requests."</p> <p>For more information on how to understand the folder structure and dataset, please read the README.md.</p>
GitHub Public Pull Request Comments
<p>Over 13 MILLION pull request comments</p><p>Dataset used for the master's thesis "LLMs for Code Comment Consistency." Covers the languages Go, Java, JavaScript, TypeScripp, and Python. All data is mined from permissively-licensed GitHub public projects with at least 25 stars and 25 pull requests submitted at the time of access.</p><p> </p><p>This dataset pertains specifically to **pull request comments that are made on files.** In other words, every comment in this dataset is linked to a specific file in a pull request.</p><p> </p><p>### What can I do with this data?</p><p>Anything you want, of course, but here are some starter ideas:</p><p>- Sentiment analysis of comments, is there a correlation between number of contributions and positivity of reviews?</p><p>- Pull request comment generation: can we automatically make code review comments?</p><p>- PR text mining: can we mine out examples of a specific type of comment? (in my project, this was comments about function documentation)</p><p> </p><p>The mining code is publicly accessible.</p><p> </p><p>Each file is a JSON object where each key is a Github repository, and each value is a pull request comment in that repository.</p>
Characterizing Support for a Third-party Library: A Study of External Pull Requests for npm packages
<p>Third-party libraries play a key role in building contemporary software applications. Despite this, most libraries are open source that often rely on volunteer (usually unpaid and overworked) contributions for their sustainability. Our motivation is to understand the extent to which third-party libraries are supported by contributions in the form of Pull Requests (PR) from outside the project team (i.e., External PR). Concretely, we analyze 1,076,123 PRs to investigate the External PR prevalence, bots, and the PR characteristics. Our results show that external contributions are prevalent, with packages receiving a high rate of (median of 73.45%) External PR . Furthermore, contributors are also submitting a high proportion of External PR (median of 87.62%). Results indicate a statistical difference in the acceptance of PRs submitted by bots compared to abandoned or open PRs. Furthermore, comparing external and internal PR, we find that Internal PR are more likely to be accepted. Statistically, we find that submitted patches (i.e., commit and code metrics) submitted by Internal PR are higher than patches submitted by External PR. We find that the External PR and Internal PR both have the same content (i.e., introducing new features and fixing bugs). Differently, External PR have more PRs that relate to documentation content, while Internal PR relates to refactoring-related changes</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.