Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
32
datasets available to search
ShareScore release 0.9.0
Dataset results
32 results for “Research Software Engineering”
Replication data [The who, what and how of the current research at the Brazilian Symposium on Software Engineering]
<p>Replication Data for the SBES paper <em>"The who, what and how of the current research at the Brazilian Symposium on Software Engineering"</em></p> <p>Dataset containing analysis (who, what and how) of 90 SBES papers: 27 from SBES’19, 43 from SBES’20, and 20 from SBES’21.</p>
Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis
<p>Dataset of the research paper: <strong>Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis</strong></p> <p>Existing work on the practical impact of software engineering (SE) research examines industrial relevance rather than adoption of study results, hence the question of how results have been practically applied remains open. To answer this and investigate the outcomes of impactful research, we performed a quantitative and qualitative analysis of 4 354 SE patents citing 1 690 SE papers published in four leading SE venues between 1975–2017. Moreover, we conducted a survey on 475 authors of 593 top-cited and awarded publications, achieving 26% response rate. Overall, researchers have equipped practitioners with various tools, processes, and methods, and improved many existing products. SE practice values knowledge-seeking research and is impacted by diverse cross-disciplinary SE areas. Practitioner-oriented publication venues appear more impactful than researcher-oriented ones, while industry-related tracks in conferences could enhance their impact. Some research works did not reach a wide footprint due to limited funding resources or unfavorable cost-benefit trade-off of the proposed solutions. The need for higher SE research funding could be corroborated through a dedicated empirical study. In general, the assessment of impact is subject to its definition. Therefore, academia and industry could jointly agree on a formal description to set a common ground for subsequent research on the topic.</p> <p>The following data files are included.</p> <ul> <li><em>./fields</em>: <ul> <li><strong>engi-fields.csv</strong>: Publication and PhD dissertation counts of main engineering branches</li> <li><strong>engi-fields-queries.txt</strong>: Queries applied to Elsevier's Scopus and Open Access Theses and Dissertations databases to retrieve the publication and dissertation counts</li> </ul> </li> <li><em>./patents</em>: <ul> <li><strong>sample-se-references-verified.csv</strong>: Manual verification of a random sample of references by software engineering (SE) patents to SE papers</li> <li><strong>se-cpc.tsv</strong>: Manually-identified SE-related Cooperative Patent Classification (CPC) categories</li> <li><strong>se-references-in-patents.csv</strong>: SE references made by SE patents to SE papers</li> <li><em>./patents/litigation</em>: <ul> <li><strong>case-values.csv</strong>: Manually-retrieved litigation damages of citing SE patents</li> <li><strong>lit-per-paper.csv</strong>: Litigation cases of citing SE patents</li> </ul> </li> <li><em>./patents/maintenance</em>: <ul> <li><strong>maint-code-fee-mapping.csv</strong>: Mapping of patent maintenance fee codes to their fee values</li> <li><strong>maint-fees.csv</strong>: Fee values of maintenance fee codes</li> <li><strong>maint-per-paper.csv</strong>: Maintenance fee events of citing SE patents</li> </ul> </li> <li><em>./patents/reports</em>: <ul> <li><strong>lit-sum-per-paper.csv</strong>: Counts and total damages of litigation cases of patent-cited SE papers</li> <li><strong>maint-sum-per-paper.csv</strong>: Counts and total values of maintenance fee events of patent-cited SE papers</li> <li><strong>patent-ref-counts.csv</strong>: SE patent citation counts of patent-cited SE papers</li> </ul> </li> </ul> </li> <li><em>./survey</em>: <ul> <li><strong>emse-top.csv</strong>: Most-cited papers of the Empirical Software Engineering (EMSE) journal</li> <li><strong>icse-bp.csv</strong>: Distinguished papers of the International Conference of Software Engineering (ICSE)</li> <li><strong>icse-mip.csv</strong>: Most influential ICSE papers</li> <li><strong>icse-top.csv</strong>: Most-cited ICSE papers</li> <li><strong>survey-questionnaire-emse.pdf</strong>: The EMSE survey questionnaire</li> <li><strong>survey-questionnaire.pdf</strong>: The ICSE, TSE, and TOSEM survey questionnaire</li> <li><strong>survey-responses.csv</strong>: The anonymized survey responses</li> <li><strong>tosem-top.csv</strong>: Most-cited papers of the ACM Transactions on Software Engineering and Methodology (TOSEM)</li> <li><strong>tse-top.csv</strong>: Most-cited papers of the IEEE Transactions on Software Engineering (TSE)</li> <li><em>./survey/manual-coding</em>: <ul> <li><strong>feedback.txt</strong>: Manual coding of survey feedback</li> <li><strong>practical-impact.csv</strong>: Manual coding of responses about practical impact of work</li> <li><strong>practical-impact-lack.csv</strong>: Manual coding of responses about lack of practical impact</li> <li><strong>research-methods.csv</strong>: Manual coding of additional research methods of surveyed papers</li> <li><strong>state-of-practice.csv</strong>: Manual coding of responses about changes in state of practice</li> </ul> </li> </ul> </li> <li><em>./venues</em>: <ul> <li><strong>se-venues.csv</strong>: Top SE venues according to Google Scholar Metrics</li> <li><strong>se-venues-impact.csv</strong>: SE patent citations and patent-based impact factors of SE venues</li> <li><strong>se-venues-scopus-queries.txt</strong>: Queries applied to Scopus to retrieve the publication counts of the SE venues</li> </ul> </li> </ul>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>This is a course project and I collect the data in a rush.</p> <p>If you want to use this dataset and find any error, please contact me ;-)</p> <p>My email: echo.xiangchen@gmail.com</p>
Diversity Awareness in Software Engineering Participant Research
<p>This dataset contains the result of a classification of three ICSE venues namely, ICSE 2019, 2020, and 2021 technical tracks, as stated in the methodology of the paper “Diversity awareness in software engineering participant studies” by Dutta et al. (2023).</p>
Rapid Review Dataset for Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team
<p>A collection of evidence briefings produced through a rapid literature review protocol the Department of Software Engineering and Research at Sandia National Laboratories. These briefings are described in our research paper, "Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team", which was accepted for publication at the 1st Annual Conference of the United States Research Software Engineer Association (US-RSE'23).</p> <p>Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525. SAND2023-06549O.</p>
Research Software Engineers Supporting Science: Survey Responses
<p>Raw survey data for "Not everyone can use git: Research Software Engineers’ recommendations for scientist-centred software support (and what researchers think of them)", a talk given by Caroline Jay at RSE16, Manchester, UK.</p>
Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal
<p>This Research Artifact contains supplemental material from the study: "Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal". The supplementary material contains:</p> <ol> <li>The list of the 106 primary studies classified by journals.</li> <li>The list of the 12 research artifacts classified by conferences.</li> <li>The .xlsx file of the dataset used to analyze the RQs.</li> <li>The .xlsx file of the dataset used to analyze the research artifacts problems (Table 3). </li> <li>The list of the figures published in the scientific article.</li> </ol>
Replication package for "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Levels of a Research Software Engineer
<p><strong>Levels of a Research Software Engineer: </strong>The diverse role of the RSE can be captured by the degree or level to which they work with researchers, and in what scope. Level 1 of RSE "domain" are closest to researchers, working directly on their behalf. Level 2 "generalist" RSE work on core technologies needed across the scientific community, and level 3 "researcher" take this a step further, researching the space or models underlying the software itself.</p>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>We collect the data in a rush.</p> <p>If you want to use this dataset and find any errors, please contact us ;-)</p> <p> </p> <p>Our emails:</p> <ul> <li>echo.xiangchen@gmail.com</li> <li>zhifengyao731@gmail.com</li> </ul>
Research on Cognition in Software Engineering
<p>This dataset includes the primary studies selected for literature review on cognition in software engineering. </p>
Using Open Citation Databases for Snowballing in Software Engineering Research
<p>Dataset for our study on the coverage of software engineering articles in open citation databases:</p> <ul> <li>a list of the 23 sampled venues with their respective CORE ranks and publishers, <ul> <li>01-venues.csv,</li> </ul> </li> <li>a list of the 204 sampled articles with their respective number of references/citations per citation database, <ul> <li>02-articles.csv (articles with publication information),</li> <li>03-references-absolute.csv (number of references in published PDF & absolute numbers for reference coverage in databases),</li> <li>04-references-relative.csv (relative numbers for reference coverage in databases),</li> <li>05-citations-absolute.csv (absolute numbers for citation coverage in databases),</li> <li>06-citations relative.csv (relative numbers for citation coverage in databases),</li> </ul> </li> <li>a list of the 8 articles analyzed in more detail with complete references data from the citation databases, <ul> <li>07-selected-articles.csv (articles with publication information),</li> <li>08A–08H (comparison of references found in databases for each article),</li> </ul> </li> <li>and additional statistical measures and plots <ul> <li>09-Statistics.{pdf,xlsx} (statistical measures – i.e., minimum, maximum, median, average, variance – for the whole dataset and for subsets by publisher, CORE rank, or year of publication),</li> <li>10-Figures.zip (figures for references as shown in the study and additional figures for citations – each in EPS and PNG format).</li> </ul> </li> </ul>
Experimental Data for: Research Perspective on Supporting Software Engineering via Physical 3D Models
<p>Experimental data for the experiment presented in the technical report 1507: "Research Perspective on Supporting Software Engineering via Physical 3D Models"</p>
Dataset for "Beyond Self-Promotion: How Software Engineering Research Is Discussed on LinkedIn"
<p>This repository contains the artifacts of our study on how software engineering research papers are shared and interacted with on LinkedIn, a professional social network. This includes:</p> <ul> <li><em>included-papers.csv</em>: the list of the 79 ICSE and FSE papers we found on LinkedIn</li> <li><em>linkedin-post-data.csv</em>: the final data of the 98 LinkedIn posts we collected and synthesized</li> <li><em>linkedin-post-scraping.zip</em>: the scripts used to automatically collect several attributes of the LinkedIn posts</li> <li><em>analysis.zip</em>: the Jupyter notebook used to analyze and visualize <em>linkedin-post-data.csv</em></li> </ul>
Social Science Theories in Software Engineering Research - Replication Package
<p>Replication package for the article "Social Science Theories in Software Engineering Research".</p> <p>See README file for additional details.</p>
On a Microservice System Benchmark with Multiple Architected Variants for Software Engineering Research
Open the record for dataset details and reuse information.
Dataset about Reproducibility in Software Engineering Research: A Systematic Mapping Study
<p>This artifact contains the results collected in one Systematic Mapping Study (SMS), about Reproducibility in Software Engineering Research. The results are associated with the selected studies and with the research questions considered.</p>
Survey on how do Brazilian Software Engineering Researchers Perceive and Practice Open Science
<p>This document provides our survey questionnaire and the responses of 31 researchers.</p>
Topic modeling in software engineering research
<p>Raw data collected from 111 papers applying topic modeling techniques in software engineering studies.</p>
Psychometric Instruments in Software Engineering Research on Personality: Status Quo After Fifty Years
<p>Step files.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.