Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,947
datasets available to search
ShareScore release 0.7.1
Dataset results
3,947 results for “Requirements”
Requirements Quality Factor Ontology
<p>We investigated a previously published set of research articles concerned with requirements quality, extracted quality factors and other relevant elements from eligible publications, and iteratively constructed an ontology of quality factors for natural language requirements. The documentation of the process, the resulting data set, and a web application to visualize the results are contained in this artifact bundle.</p>
Net irrigation requirement under different climate scenarios using AquaCrop over Europe
<p>This repository contains the setup and data related to the peer-reviewed article "Net irrigation requirement under different climate scenarios using AquaCrop over Europe" accepted for HESS (https://hess.copernicus.org/preprints/hess-2021-631/).</p> <p>The README.txt file contains all information about the repository. Please contact Louise Busschaert (louise.busschaert@kuleuven.be) or Gabrielle De Lannoy (gabrielle.delannoy@kuleuven.be) for any further questions.</p>
Supplementary Material for Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study
<p>This repository delivers the supplementary material for the paper: <em>Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study.</em></p> <p>In the following, we present the abstract of the study:</p> <p><strong>Context:</strong> Agile Software Development (ASD) is a disruptive process compared to traditional software development. Therefore, traditional Requirements Engineering (RE) forms may not be the best way to do RE for ASD (RE-ASD). <strong>Objective:</strong> Working with ASD using traditional RE ways could limit ASD's potential. Thus, it is necessary to investigate what academia and industry have done in RE to take full advantage of all of the capabilities of ASD beyond traditional RE. <strong>Method: </strong>We conducted a Tertiary Study looking for solutions for RE-ASD using the Systematic Literature Review (SLR) protocol described by Kitchenham and Charters. We then categorized the solutions into families using Targeted Coding and Constant Comparison, tools from Socio-Technical Grounded Theory (STGT). Afterward, we classified the solutions as disruptive using our model based on the Hype Level Curve concept, assessing their hype (popularity) in the software engineering community using Google Trends and Google Colab tools. <strong>Results:</strong> After executing the SLR protocol, we accepted 37 studies and encountered 136 solutions used by academia and industry for RE-ASD. We categorized these solutions into 21 solution families, six of which we classified as disruptive. Design Thinking (DT) and Artificial Intelligence (AI) were the two families of solutions that stood out the most. We also identified the type of solution (e.g., process, method, technique, tool, model, framework) and domain (academia or industry). Furthermore, we cataloged the challenges presented by the solutions. <strong>Conclusion:</strong> We concluded that only a few solutions that have been used for RE-ASD have the power to successfully challenge the mainstream Agile Software Development process by using innovation (26 out of 106). There is a gap between academia and industry regarding these disruptive solutions, and some challenges still need to be addressed in using these solutions.</p> <p>The repository contains the following:</p> <ul> <li>Dataset from the Tertiary Study: <ul> <li>Data of the retrieved studies. It presents the classifications of the documents as 'Accepted,' 'Rejected' (with the indication of the step of the protocol the authors rejected the study), or 'Duplicated.'</li> <li>Data of all solutions retrieved from the accepted studies</li> </ul> </li> <li>Socio-Technical Grounded Theory (STGT) tools <ul> <li>Result of the use of Targeted Coding and Constant Comparison</li> </ul> </li> <li>The Google Colab Notebook <ul> <li>Code in python</li> <li>Results</li> </ul> </li> </ul> <p> </p>
Strategies, Benefits and Challenges App Store-inspired Requirements Elicitation - Supplementary Material
<p>This is the supplementary material for the paper "Strategies, Benefits and Challenges App Store-inspired Requirements Elicitation". </p> <p>Abstract: App store-inspired elicitation is the practice of exploring competitors’ apps, to get inspiration for requirements. This activity is common among developers, but little insight is available on its practical use, advantages and possible issues. This paper aims to study strategies, benefits and challenges of app store-inspired elicitation, and compare this technique with more traditional requirements elicitation interviews. We conduct an experimental simulation with 58 analysts, and collect qualitative data. Our results show that specific guidelines and procedures are required to better conduct app store-inspired elicitation. Furthermore, current search features made available by app stores are not suitable for this practice, and more tool support is required to help analysts in the retrieval and<br> evaluation of competing products. While interviews focus on the why dimension of requirements engineering (i.e., goals), app store-inspired elicitation focuses on how (i.e., solutions), offering indications for implementation and improved usability. Our study provides a framework for researchers to address existing challenges, and suggests possible benefits to foster app store-inspired elicitation among practitioners.</p> <p>The package contains the following files:</p> <p>1.Protocol.pdf - it describes in details the steps of the protocol and the intermediate results obtained during the execution.</p> <p>2. Codebooks:<br> 2.a. Codebook Strategies: codebook of the strategies to select apps<br> 2.b Codebook Benefits: codebook of the benefits of use IBE (sheet 1) and ASE (sheet 2)<br> 2.c Codebook Challenges: codebook of the challenged of use IBE (sheet 1) and ASE (sheet 2)<br> 2.d Differences IBE-ASE: table of the identified (categorized) differences between IBE and ASE</p> <p>3. Labelled Data <br> 3.a Strategies - labelled data: the file contains the name of the selected apps, the motivation behind the selection, and the themes assigned to them (refer to 2.a for explanation of the themes).<br> 3.b Benefits IBE - labelled data: the file contains the extract of the raw data about IBE benefits and the themes assigned (refer to 2.b for explanation of the themes).<br> 3.c Challenges IBE - labelled data: the file contains the extract of the raw data about IBE challenges and the themes assigned (refer to 2.c for explanation of the themes).<br> 3.d Benefits ASE - labelled data: the file contains the extract of the raw data about ASE benefits and the themes assigned (refer to 2.b for explanation of the themes).<br> 3.e Challenges ASE - labelled data: the file contains the extract of the raw data about IBE benefits and the themes assigned (refer to 2.c for explanation of the themes).</p> <p>4. Raw data.xls: it contains the raw data used in the work (two sheets, one for strategies and one for reflections).</p> <p>5. SLR data: data related to the lightweight systematic literature review <br> 5.a Codebook Scopus.xlsx: codebook for the themes elicited from the SLR. The themes are also present in the files in the folder Codebooks.<br> 5.b SLR-scopus-results-and-selected.xlsx: results of the search string, and, in green, the selected papers. </p> <p>6. Readme.txt: summary file. </p> <p>Note that some of the row in Raw data.xls (and in the corresponding "Labelled Data" files) are substituted with N/A. This corresponds to those participants who asked to not publicly share their responses.</p>
PURE: a Dataset of Public Requirements Documents
<p>Please cite this dataset as <strong>Ferrari, A., Spagnolo, G. O., & Gnesi, S. (2017, September). PURE: A dataset of public requirements documents. In <em>2017 IEEE 25th International Requirements Engineering Conference (RE) </em>(pp. 502-505). IEEE.</strong></p> <p><a href="https://ieeexplore.ieee.org/abstract/document/8049173">https://ieeexplore.ieee.org/abstract/document/8049173</a></p> <p>This dataset presents PURE (PUblic REquirements dataset), a dataset of 79 publicly available natural language requirements documents collected from the Web. The dataset includes 34,268 sentences and can be used for natural language processing tasks that are typical in requirements engineering, such as model synthesis, abstraction identification and document structure assessment. It can be further annotated to work as a benchmark for other tasks, such as ambiguity detection, requirements categorisation and identification of equivalent re-quirements. In the associated paper, we present the dataset and we compare its language with generic English texts, showing the peculiarities of the requirements jargon, made of a restricted vocabulary of domain-specific acronyms and words, and long sentences. We also present the common XML format to which we have manually ported a subset of the documents, with the goal of facilitating replication of NLP experiments. The XML documents are also available for download.</p> <p>The paper associated to the dataset can be found here: </p> <p>https://ieeexplore.ieee.org/document/8049173/</p> <p>More info about the dataset is available here: </p> <p>http://nlreqdataset.isti.cnr.it</p> <p>Preprint of the paper available at ResearchGate:</p> <p>https://goo.gl/HxJD7X</p> <p>The dataset includes:</p> <p>- all the documents in PDF format</p> <p>- a subset of 19 documents in XML format</p> <p>- the .xsd schema of the XML files</p> <p>The dataset has been created by gathering data from web sources and we are not aware of license agreements or intellectual property rights on the requirements. The curator took utmost diligence in minimizing the risks of copyright infringement by using non-recent data that is less likely to be critical, by sampling a subset of the original requirements collection, and by qualitatively analyzing the requirements. In case of copyright infringement, please contact the dataset curator (Alessio Ferrari, alessio.ferrari@cnr.it, alessio.ferrari@ucd.ie) to discuss the possibility of removal of that dataset [see <a href="https://support.zenodo.org/help/en-gb/13-policies/140-what-is-your-take-down-procedure" target="_blank" rel="noopener">Zenodo's policies</a>].</p>
Text-fig. 3. Geology of the Cheringoma Plateau, Mozambique. Sections and geological map adapted from Tinley (1977). The star symbols close to Mhengere Hill represent fossil wood and stem sites. Note that the fault relationships proposed in the northernmost Inhaminga section require re-examination. The Nguere Hills were called Gadjiua by Tinley (1977). in Stratigraphy, Chronology And Palaeontology Of The Tertiary Rocks Of The Cheringoma Plateau, Mozambique
Text-fig. 3. Geology of the Cheringoma Plateau, Mozambique. Sections and geological map adapted from Tinley (1977). The star symbols close to Mhengere Hill represent fossil wood and stem sites. Note that the fault relationships proposed in the northernmost Inhaminga section require re-examination. The Nguere Hills were called Gadjiua by Tinley (1977).
Supplementary Material for Documentation artifacts for conversation-related requirements specification in chatbots: a systematic review and a meta-model proposal
<p>This is a supplementary data of the tertiary systematic literature review conducted in the paper "Conversation-related requirements specification in chatbots: a systematic review and a meta-model proposal".</p> <p>Context: Chatbots are complex applications due to their capacity to engage and maintain a conversation with humans. However, the conversational-related requirements of chatbots are hard to elicit, document, and test. Another challenge is the documentation since there are not so many directions on how to register and test subjective requirements.</p> <p>Methods: We followed systematic literature review (SLR) guidelines and identified 42 relevant papers that address the artifacts used by practitioners to document conversational-related requirements in literature. We also investigated what conversational requirements are addressed in requirements documentation.</p> <p>Results: The main results indicate that UML diagrams, prototypes, tables of requirements, conversational flows, and scenarios are present in most chatbot documentation. Except for UML diagrams, those artifacts are used to document standard requirements or conversational requirements. In those artifacts, context-dependent behavior, assertivity, error handling, and human-like attitude are the most approached conversational requirements in the studies. In sequence, based on our findings, we propose the conversational integrated map, a meta-model solution as documentation of conversational requirements.</p>
Recovery of unavailable Requirements Quality Artifacts
<p>We extracted a <a href="http://www.reqfactoront.com/">requirements quality factors ontology</a> from existing requirements quality literature in previous research. This ontology revealed that several artifacts (data sets and tools/implementations) are unavailable, hindering progress in the research domain. In the project based on this replication package, we attempted to recover lost artifacts by requesting authors to disclose their artifacts according to open science principles. This repository contains both the process description, tools for conduction of the recovery, the results, and the evaluation thereof.</p>
Data from: Plant ammonium sensitivity is associated with the external pH adaptation, repertoire of nitrogen transporters, and nitrogen requirement
<p>Modern crops exhibit diverse sensitivities to ammonium as the primary nitrogen source, influenced by environmental factors such as external pH and nutrient availability. Despite its significance, there is currently no systematic classification of plant species based on their ammonium sensitivity. This study conducts a meta-analysis of 50 plant species and presents a new classification method based on the comparison of fresh biomass obtained under ammonium and nitrate nutrition. The classification uses the natural logarithm of biomass ratio as the size effect indicator of ammonium sensitivity. This numerical parameter is associated with critical factors for nitrogen demand and form preference, such as Ellenberg indicators and the repertoire of nitrogen transporters for ammonium and nitrate uptake. Finally, a comparative analysis of the developmental and metabolic responses, including hormonal balance, is conducted in two species with divergent ammonium sensitivity values in the classification. Results indicate that nitrate has a key counteracting role of ammonium toxicity in species with a higher abundance of genes encoding NRT2-type proteins and fewer of the AMT2-type proteins. Additionally, the study confirms the reliability of the phytohormone balance and methylglyoxal content as indicators for anticipating ammonium toxicity.</p>
Figure 1 in Character mapping and cladogram comparison versus the requirement of total evidence: does it matter for polychaete systematics?
Figure 1. Example of the error of cladogram comparisons. A, phylogenetic hypotheses inferred from separate sets of premises. Letters on cladogram 'nodes' indicate population-splitting events relevant to the various hypotheses of character origin/fixation within ancestral populations. The requirement of total evidence precludes such a comparison of cladogram topologies because explanations of characters 1(1)–5(1) by population-splitting events A–C (left cladogram) contradict explanations of 6(1)–8(1) by population-splitting events D–F. See text for further discussion. B, explaining observations in accordance with the requirement of total evidence, correcting the problem in 'A'.
Figure 2 in Character mapping and cladogram comparison versus the requirement of total evidence: does it matter for polychaete systematics?
Figure 2. Example of the error of character mapping. A, phylogenetic hypotheses are inferred for a set of characters. Numbers on cladogram 'nodes' indicate population-splitting events relevant to the various hypotheses of character origin/fixation within ancestral populations (not shown; cf. fig. 1). B, a different set of characters are 'mapped' onto the branches of the cladogram in 'A'. C, the 'mapped' characters in 'B' actually refer to phylogenetic hypotheses inferred separately from the hypotheses implied by the cladogram in 'A' and 'B'. D, explaining observations in accordance with the requirement of total evidence, correcting the problem in 'B' and 'C'. See text for further discussion.
Fig. 4 in Habitat requirements and occurrence of Crematogaster pilosa (Hymenoptera: Formicidae) ants within intertidal salt marshes
Fig. 4. Logistic regression model (P = 0.03) of the probability of Crematogaster pilosa as a function of brown leaf density between 0.61 and 1.20 m. Stars indicate plots containing ants, and open symbols indicate plots not containing ants. Vertical dashed line represents a 50% probability of ants and occurs at a brown leaf density of 2.5 m−1, which equals 1.5 brown leaves between 0.61 and 1.20 m above the marsh surface.
Fig. 2 in Habitat requirements and occurrence of Crematogaster pilosa (Hymenoptera: Formicidae) ants within intertidal salt marshes
Fig. 2. Mean vegetation heights for marsh plots containing ants (n = 8) and plots not containing ants classified by their dominant vegetation type: short (n = 7) and tall (n = 2). All plots were from Dean Creek and Odum's Marsh. Mean heights are the weighted average of all vegetation counts within plots. Letters above whiskers signify significant difference using Tukey's HSD with P <0.05.
Fig. 1 in Habitat requirements and occurrence of Crematogaster pilosa (Hymenoptera: Formicidae) ants within intertidal salt marshes
Fig. 1. Southern tip of Sapelo Island, Georgia (USA). Location of Crematogaster pilosa observations and vegetation assessments in Odum's Marsh (A) and Dean Creek (C). Presence/absence of ants along Lighthouse Creek (B) from canoe and baited trap survey. Sites containing C. pilosa were labeled "ants", those not containing ants were labeled by their vegetation (i.e., short or tall) based on maximum vegetation height.
Fig. 3 in Habitat requirements and occurrence of Crematogaster pilosa (Hymenoptera: Formicidae) ants within intertidal salt marshes
Fig. 3. Height-specific vegetation density for marsh plots with and without ants in Dean Creek and Odum's Marsh. Vegetation density is the number of vegetation features (i.e., stems and leaves) per vertical meter above an average point on the marsh surface. Integrating vertically produces the average number of vegetation features above a single point. All plots containing Crematogaster pilosa were grouped (ants); plots not containing ants were classified by the maximum vegetation height of Spartina alterniflora (i.e., tall or short). Vegetation density distributions are the means for tall (n = 2), ants (n = 8), and short (n = 7) plots.
Datasets for Crowd-based Requirements Engineering and aspect-based detection of learning-centered emotion from the text in Serbian language
<h1>Datasets for the paper "Enhancing Software and Learning with Serbian Student Feedback Corpora"</h1> <p>These datasets include student feedback on an Intelligent Tutoring System written in Serbian, annotated with categories for Crowd-based Requirements Engineering (CrowdRE) and aspect-based detection of learning-centered emotions. Four annotators manually annotated each sentence. </p> <p>The CrowdRE dataset includes two JSON files:</p> <ul> <li><strong>crowdre_english.json</strong> - annotated text with columns and classes written in English. Columns are: <ul> <li><em>Comment</em> - the entire student feedback</li> <li><em>Sentence</em> - sentence extracted from the feedback that was annotated</li> <li><em>Intention </em>- class representing the intention of the sentence</li> <li><em>Topic </em>- class representing the topic of the sentence</li> </ul> </li> <li><strong>crowdre_srpski.json</strong> - annotated text with columns and classes written in Serbian. Columns are:<br> <ul> <li><em>Komentar</em> - the entire student feedback</li> <li><em>Recenica </em>- sentence extracted from the entire feedback that was annotated</li> <li><em>Namera </em>- class representing the intention of the sentence</li> <li><em>Tema </em>- class representing the topic of the sentence.</li> </ul> </li> </ul> <p>The dataset for the aspect-based detection of learning-centered emotions includes two JSON files:</p> <ul> <li><strong>emotions_english.json </strong>- annotated text with columns and classes written in English. Columns are: <ul> <li><em>Comment </em>- the entire student feedback</li> <li><em>Sentence </em>- sentence extracted from the feedback that was annotated</li> <li><em>Aspect </em>- class representing the aspect of the sentence</li> <li><em>Emotion </em>- class representing the learning-centered emotion of the sentence</li> </ul> </li> <li><strong>emocije_srpski.json </strong>- annotated text with columns and classes written in Serbian. Columns are: <ul> <li><em>Komentar </em>- the entire student feedback</li> <li><em>Recenica</em> - sentence extracted from the entire feedback that was annotated</li> <li><em>Aspekt</em> - class representing the aspect of the sentence</li> <li><em>Emotion </em>- class representing the learning-centered emotion of the sentence.</li> </ul> </li> </ul> <p>Annotators annotated the dataset based on the annotation procedure and guidelines available <a href="https://github.com/Clean-CaDET/student-feedback-mining">here</a>. </p> <h2>Citation</h2> <p>If you use this in your research, please cite:</p> <blockquote> <p>Vidaković, D., Luburić, N., Kovačević, A., & Slivka, J. Enhancing software and learning with Serbian student feedback corpora. Language Resources & Evaluation (2025). https://doi.org/10.1007/s10579-025-09855-y</p> </blockquote> <p> </p>
Figure 6. Adding a Global Goal with its required type-Designing a Growing Functional Modules "Artificial Brain"
<p>The next step consists of adding a Global Goal expressing a motivation required by the<br> controller. The goal is to keep the vehicle's front free of obstacles, thus the Sensation “free” should<br> stay equal to “1”. After adding a new Global Goal, its assigned type should be “Cst” corresponding<br> to a constant output request (see figure 6). In the parameter field, its specified value is “1”.</p>
Microsat Data for 'Simulated Disperser Analysis: determining the number of loci required to genetically identify dispersers'
<p>Microsattelite data from 94 samples (<em>Stunus vulgaris</em>) from 3 populations and including 29 loci. Used in the paper 'Simulated Disperser Analysis: determining the number of loci required to genetically identify dispersers'. </p>
Software Requirement Risk Prediction Dataset
<p>@attribute Requirements {'The system shall display all the products that can be configured.','The system shall allow user to select the product to configure.','The system shall display all the available components of the product to configure','The system @attribute 'project Category ' {'Transaction Processing System','Management Information System','Enterprise System','Safety Critical System'}@attribute 'Requirement Category' {Functional,Usability,'Reliability & Availability',Performance,Security,Supportability,Constraints,Interfaces,Standards,Safety}@attribute 'Risk Target Category' {Budget,Quality,Schedule,Personal,Performance,FunctionalValidity,People,'Project complexity','Planning & Control',Team,'Resource availability',User,Requirement,'Time Dimension','Organizational Environment',Cost,Design,Bus@attribute Probability numeric@attribute 'Magnitude of Risk' {Negligible,'Very Low',Low,Medium,High,'Very High',Extreme}@attribute Impact {high,catastrophic,moderate,Low,insignificant}@attribute 'Dimension of Risk' {Requirements,User,'Project complexity','planning and control',Team,'Organizational Environment',Estimations,'Software Requirement','Planning and Control',Schedule,Cost,'Project Complexity','Organizational Requirements'}@attribute 'Afftecting No of Modules' numeric@attribute 'Fixing Duration (Days)' numeric@attribute 'Fix Cost (\% of Project)' numeric@attribute Priority numeric@attribute 'Risk Level' {1,2,3,4,5}</p>
The REquirements TRacing On target (RETRO).NET Dataset
<p> For this dataset, we took the RETRO Requirement Specification (version 1.0, written to document the features in our original RETRO tool) and used RETRO.NET to trace it to the code used to implement RETRO.NET (copyright Jody Larsen). There are 66 requirements (functional requirements only) that have been extracted from the document. RETRO.NET expects all source elements to be in a folder and all target elements to be in a folder. Therefore, in our dataset, each requirement has been stored in its own file with the identifier as the file name and the file containing its text (RETRONET Requirements folder). The original requirements specification as well as the document subset have been provided in the dataset (as .docx and as data.txt, respectively) along with the python script (parser.py, copyright to Jared Payne) that was used to parse the text only version of the document subset. There are 118 code files in the dataset, primarily C# files. Each code file is listed in the code directory (RETRONET Trunk folder). The answer set has been provided in xml format (result.xml) as well as in traditional answer set format of our research group (results.txt). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.