Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
113
datasets available to search
ShareScore release 0.9.0
Dataset results
113 results for “research software”
Replication Data for "Mapping the Structure and Evolution of Software Testing Research Over the Past Three Decades"
<p>In this research (publication included in the package), we have used author-assigned keywords as a quantitative data source for understanding the connections between keywords and research topics in software testing research, based on a large sample of studies from Scopus.</p> <p>We apply co-word analysis to map the topology of testing research as a network where author-assigned keywords are connected by edges indicating co-occurrence in publications. Keywords are clustered based on edge density and frequency of connection. We examine the most popular keywords, summarize clusters into high-level research topics, examine how topics connect, and examine how the field is changing. This package contains the map and network files used to perform our analyses, as well as the publication sample.</p>
Research Software Funders Global South
<p>The Research Software Alliance's (ReSA) mission is to bring research software communities together to collaborate on the advancement of research software. Given the ReSA mission, it is important to understand the landscape of communities involved with research software. In 2020, ReSA completed an initial exercise to scope the international research software community landscape. This work was reported by ReSA's Software Landscape Analysis task force via a <a href="https://www.researchsoft.org/blog/2020-03/">blog post</a>. The majority of the communities in the previous analysis represented the global north. To improve the extent of this landscape analysis, ReSA announced a paid opportunity for short-term contractors located in <a href="https://www.researchsoft.org/2022-mapping/">the global south</a> to collect data on communities and funders in their region in early 2022. This document describes how the work was undertaken, a summary of findings, the gaps and opportunities perceived by the data collectors and some highlights. This work identified 126 organisations and communities and 62 funder bodies that support research software in the global south. Their main activities are connecting people, training, and networking, and support through research grants.</p> <p> </p> <p>To add to this funders list please fill in the following form: <a href="https://forms.gle/CJWo24MUCjhWKh9U8">https://forms.gle/CJWo24MUCjhWKh9U8</a></p>
Research Software Communities Global South
<p>The Research Software Alliance's (ReSA) mission is to bring research software communities together to collaborate on the advancement of research software. Given the ReSA mission, it is important to understand the landscape of communities involved with research software. In 2020, ReSA completed an initial exercise to scope the international research software community landscape. This work was reported by ReSA's Software Landscape Analysis task force via a <a href="https://www.researchsoft.org/blog/2020-03/">blog post</a>. The majority of the communities in the previous analysis represented the global north. To improve the extent of this landscape analysis, ReSA announced a paid opportunity for short-term contractors located in <a href="https://www.researchsoft.org/2022-mapping/">the global south</a> to collect data on communities and funders in their region in early 2022. This document describes how the work was undertaken, a summary of findings, the gaps and opportunities perceived by the data collectors and some highlights. This work identified 126 organisations and communities and 62 funder bodies that support research software in the global south. Their main activities are connecting people, training, and networking, and support through research grants.</p> <p> </p> <p>To add to this communities list please fill in the following form <a href="https://forms.gle/KJE9vkBnM6vhh7cEA">https://forms.gle/KJE9vkBnM6vhh7cEA</a></p>
Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal
<p>This Research Artifact contains supplemental material from the study: "Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal". The supplementary material contains:</p> <ol> <li>The list of the 106 primary studies classified by journals.</li> <li>The list of the 12 research artifacts classified by conferences.</li> <li>The .xlsx file of the dataset used to analyze the RQs.</li> <li>The .xlsx file of the dataset used to analyze the research artifacts problems (Table 3). </li> <li>The list of the figures published in the scientific article.</li> </ol>
Figure 2. Some screenshots from the software system-Design and Development of a Software System for Swarm Intelligence Based Research Studies
<p>All of the mentioned operations can be performed easily by using the provided controls over<br> the related interfaces – windows of each algorithm. It is also important that each algorithm interface<br> is supported by visual controls to view obtained results with typical iteration-based graphics or<br> problem oriented visual elements. For instance, resulting graph structures are automatically shown<br> by the algorithm interfaces after solving some specific, popular problems like Travelling Salesman<br> Problem (TSP), Vehicle Routing Problem (VCP)…etc. Visually improved using features and<br> functions of the software system are critical aspects to provide more effective and useful platform to<br> perform SI based research studies better.<br> Related to the designed and developed software system, some screenshots from the software<br> system [interfaces of two algorithms (IWDs and ABC)] are represented in Fig. 2.</p>
Results of a research software programming and development survey at the University of Reading
<p>In 2017 an online survey of University of Reading staff active in or supporting research and registered PhD students was undertaken to assess the nature and extent of research programming and software development activities in the University, and to understand how the University might provide guidance, training and support. The survey was a administered by the Research Data Manager on behalf of the University's Research Data Management Steering Group. The survey ran from 1st November to 15th December 2017 and collected a total of 170 responses.</p> <p>The survey sought responses from anyone in the University who was involved in any of the following activities:</p> <ul> <li>writing code and using software for numerical and statistical analysis;</li> <li>creating and contributing to computational models or simulations;</li> <li>conducting Text and Data Mining (TDM) and content analysis;</li> <li>creating and contributing to software distributed as a product or implemented as a service;</li> <li>creating data visualisations;</li> <li>using markup languages to structure and render content.</li> </ul> <p>The survey was distributed using the Bristol Online Survey. A dataset of anonymised survey responses and a PDF of the survey questions are here included.</p>
What makes research software sustainable? Anonymized interview transcripts
<p>Anonymized transcripts of a series of interviews with the developers of research software.</p> <p>We also include the Participant Information Sheet provided to all participants before the interview commences.</p>
Replication package for "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Levels of a Research Software Engineer
<p><strong>Levels of a Research Software Engineer: </strong>The diverse role of the RSE can be captured by the degree or level to which they work with researchers, and in what scope. Level 1 of RSE "domain" are closest to researchers, working directly on their behalf. Level 2 "generalist" RSE work on core technologies needed across the scientific community, and level 3 "researcher" take this a step further, researching the space or models underlying the software itself.</p>
Contemporary Software Modernization: Strategies, Driving Forces, and Research Opportunities - Supplementary Material
<p>This dataset contains the bibs, data extraction (raw) and the list of selected studies that are part of our paper entitled: "Contemporary Software Modernization: Strategies, Driving Forces, and Research Opportunities"</p>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>We collect the data in a rush.</p> <p>If you want to use this dataset and find any errors, please contact us ;-)</p> <p> </p> <p>Our emails:</p> <ul> <li>echo.xiangchen@gmail.com</li> <li>zhifengyao731@gmail.com</li> </ul>
Evaluating an instrument of the research software related to software use and disclosure - Dimension 1 - Dataset of Focus Groups
<p>Artifacts used for data collection and analysis of the focus groups sessions during the evaluation of an instrument for research software related to software use and disclosure - dimension 1.</p>
BEE-STEWARD: a research and decision support software for effective land management to promote bumblebee populations
<p><span><span>The demand for agent-based models to explore the effects of environmental change on pollinator population dynamics is growing. However, models need a simple yet flexible interface to enable adoption by a wide range of stakeholders. </span></span><span><span>We introduce BEE-STEWARD: a research and decision-support software tool, enabling researchers, policy-makers, land management advisors, and practitioners to predict and compare the effects of bee-friendly management interventions on bumblebee populations over several years. </span></span><span><span>BEE-STEWARD integrates the BEESCOUT and <i>Bumble</i>-BEEHAVE agent-based models of bumblebee behaviour, colony growth and landscape exploration into a user-friendly interface, with reconstructed code, and expanded functionality. Bespoke automatic reports can be created to illustrate how different land management interventions can affect the densities of bumblebees and their colonies over time. </span></span><span><span>BEE-STEWARD could be an important virtual test-bed for scientists exploring the impacts of different stressors on bumblebees and used by those with little or no modelling experience, enabling a shared methodology between research, policy, and practice.</span></span></p>
How can software containers help your research?
<p>This video explains software containers to a research audience. It is an introduction to why containers are beneficial for research. These benefits are standardisation, portability, reliability and reproducibility. </p> <p>Software Containers in research are a solution that addresses the challenge of a replicable computational environment and supports reproducibility of research results. Understanding the concept of software containers enables researchers to better communicate their research needs with their colleagues and other researchers using and developing containers.</p> <p><strong>Watch the video here: <a href="https://www.youtube.com/watch?v=HelrQnm3v4g">https://www.youtube.com/watch?v=HelrQnm3v4g</a></strong></p> <p>If you want to share this video please use this:</p> <p>Australian Research Data Commons, 2021. <em>How can software containers help your research?</em>. [video] Available at: <a href="https://www.youtube.com/watch?v=HelrQnm3v4g">https://www.youtube.com/watch?v=HelrQnm3v4g</a> DOI: <a href="http://doi.org/10.5281/zenodo.5091260">http://doi.org/10.5281/zenodo.5091260</a> [Accessed dd Month YYYY].</p>
Research Data and Software (re)Use Indications in High Energy Physics related Scholarly Works (Full dataset)
<p>This dataset contains research data and software (re)use indications (formal citations, informal mentions) in scholarly works related to High Energy Physics. 1,411 research and software indications were identified by a mix of approaches: use of citation discovery services and multiple search approaches in Google Scholar. The dataset contains indications by what approach the (re)use indications were found. All identified research data and software (re)use indications were classified according to their purpose, location, and elements.</p> <p>The data was collected in 2018 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>
Identification of Research Data and Software (re)Use Indications in High Energy Physics related Scholarly Works (1951 - 2018)
<p>This dataset contains research data and software (re)use indications in High Energy Physics related scholarly works. A minimal random sample of scholarly works that contain highly processed research data was taken from HEPData and manually read for research data and software (re)use indications. The sample contains 368 works, which were published between 1951 and 2018.</p> <p>The data was collected in 2019 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>
Research on Cognition in Software Engineering
<p>This dataset includes the primary studies selected for literature review on cognition in software engineering. </p>
ICITS'23 - Understanding the Success Factors of Research Software: Interviews with Brazilian Computer Science Academic Researchers
<p>Artifacts used for data collection and analysis of the article accepted for publication in ICITS'23.</p> <p>Mourão, E., Trevisan, D., Viterbo, J. (2022).Understanding the Success Factors of Research Software: Interviews with Brazilian Computer Science Academic Researchers. In: ICITS'23 - 6th International Conference on Information Technology & Systems. Advances in Intelligent Systems and Computing, Springer, Cham.</p>
Using Open Citation Databases for Snowballing in Software Engineering Research
<p>Dataset for our study on the coverage of software engineering articles in open citation databases:</p> <ul> <li>a list of the 23 sampled venues with their respective CORE ranks and publishers, <ul> <li>01-venues.csv,</li> </ul> </li> <li>a list of the 204 sampled articles with their respective number of references/citations per citation database, <ul> <li>02-articles.csv (articles with publication information),</li> <li>03-references-absolute.csv (number of references in published PDF & absolute numbers for reference coverage in databases),</li> <li>04-references-relative.csv (relative numbers for reference coverage in databases),</li> <li>05-citations-absolute.csv (absolute numbers for citation coverage in databases),</li> <li>06-citations relative.csv (relative numbers for citation coverage in databases),</li> </ul> </li> <li>a list of the 8 articles analyzed in more detail with complete references data from the citation databases, <ul> <li>07-selected-articles.csv (articles with publication information),</li> <li>08A–08H (comparison of references found in databases for each article),</li> </ul> </li> <li>and additional statistical measures and plots <ul> <li>09-Statistics.{pdf,xlsx} (statistical measures – i.e., minimum, maximum, median, average, variance – for the whole dataset and for subsets by publisher, CORE rank, or year of publication),</li> <li>10-Figures.zip (figures for references as shown in the study and additional figures for citations – each in EPS and PNG format).</li> </ul> </li> </ul>
Dataset of article: Investigating Developers' Perception on Success Factors for Research Software Development
<p>This dataset is an addendum to the article "Investigating Developers' Perception on Success Factors for Research Software Development" to provide information regarding the anonymously collected data.</p> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.