Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
113
datasets available to search
ShareScore release 0.9.0
Dataset results
113 results for “research software”
Empirical research in software architecture - perceptions of the community (supplementary material)
<p>This repository contains supplementary material for the manuscript "Empirical research in software architecture - perceptions of the community" published in the Journal of Systems and Software (<a href="https://doi.org/10.1016/j.jss.2023.111684">https://doi.org/10.1016/j.jss.2023.111684</a>):</p> <ul> <li>ICSA_papers.xlsx: This file contains the papers published at ICSA 2017-2021.</li> <li>invited_PC_members.xlsx: This file contains all names of PC members of ECSA, ICSA, QoSA, CBSE and WICSA for all instances of these conferences. This information was collected in June 2017 from the publicly available websites of these conferences as well as the publicly available online profiles of PC members.</li> <li>protocol.pdf: This files contains a summary of the survey protocol.</li> <li>responses.xls: This file contains all the responses of the survey from all respondents for all questions.</li> </ul>
Supplemental package of a study on the use of qualitative surveys in software engineering researchers
<p>This package contains the list of analyzed papers in a meta-study on the use of qualitative surveys in software engineering research. It contains all the files obtained in queries to scientific databases, and the list of analyzed papers containing how they were classified according to the aspects analyzed.</p>
Data from: Research and exploratory analysis driven - time-data visualization (read-tv) software
Open the record for dataset details and reuse information.
Data: Researcher Perspectives on the Use and Sharing of Software
Open the record for dataset details and reuse information.
Role and practice of research software development at DLR
<p>The deposit contains the results of a survey concerning the role and practice of research software development at the German Aerospace Center (DLR). The survey started end of November 2018 and ended mid-February 2019. We received 773 answers which gave us interesting insights concerning developer demographics, tool usage, documentation, testing as well as software citation at DLR.</p>
Public Money? Public Code: What 'Free' Software Really Means in Research
<p><strong>Episode Summary:</strong></p> <p>Our guest Dr Christian Busse spoke to us about the Free Software Foundation Europe and the challenges and opportunities connected to Open and Free Software...and what the differences between those two things might be. Christian has very kindly supplied some notes for us to add this week.</p> <p><strong>Links: </strong> </p> <p>First, the main website of the "Public Money, Public Code" campaign is: <a href="https://publiccode.eu/">https://publiccode.eu</a> </p> <p>Second, the "Legal activities" section of the FSFE website, which includes further links for licensing questions, workshops and the Legal Network, can<br> be accessed via <a href="https://fsfe.org/activities/ftf/activities.en.html">https://fsfe.org/activities/ftf/activities.en.html</a> </p> <p>Finally, there is an position paper by the FSFE on Free Software in Horizon2020: <a href="https://fsfe.org/activities/ftf/activities.en.html">https://fsfe.org/activities/ftf/activities.en.html</a> </p> <p><strong>Quotes:</strong></p> <p>'Free as in freedom and free as in beer'</p>
Software used in research based on combined surveys
<p>The combined results of five surveys run by the Software Sustainability Institute, which were run between 2014 to 2016. The data relate to 1261 survey participants who were asked “What software do you use in your research?”.</p> <p>The data are described here:</p> <p>https://www.software.ac.uk/blog/2016-08-13-quick-and-dirty-analysis-software-being-used-research-python-matlab-and-r</p>
Dataset for paper: Research Artifacts in Secondary Studies: A Systematic Mapping in Software Engineering
Open the record for dataset details and reuse information.
ICITS'24 - A Systematic Mapping Study on the Use and Development of Research Software
<p>Artifacts used for data collection and analysis of the article accepted for publication in ICITS'24.</p><p>Mourão, E., Trevisan, D., Viterbo, J. and Pantoja, C.E. (2024). A Systematic Mapping Study on the Use and Development of Research Software. In: ICITS'24 - 7th International Conference on Information Technology & Systems. Lecture Notes in Networks and Systems. Springer, Cham.</p>
Supplementary Material for: Starting Collaborations Between SMEs and Researchers in Software Engineering
<p>Supplementary Material for: Starting Collaborations Between SMEs and Researchers in Software Engineering</p>
Supplementary Material for Paper Titled "Challenges, Adaptations, and Fringe Benefits of Conducting Software Engineering Research with Human Participants during the COVID-19 Pandemic"
<p>This archive contains the supplementary material associated with our study titled "Characterizing Human Aspects in Reviews of COVID-19 Apps".</p> <p>Specifically, the archive contains:</p> <ol> <li>Codebooks for interviews</li> <li>Codebooks for qualitative questions in our survey</li> <li>The email template</li> <li>Survey questions as a PDF</li> </ol>
Research software funding policies and programs: Results from an international survey (Dataset)
<p>Research software is increasingly recognized as critical infrastructure in contemporary science. Research software spans a broad spectrum, including source code files, algorithms, scripts, computational workflows, and executables, all created for or during research. Research funders have developed programs, initiatives and policies to bolster research software’s role. However, there has been no empirical study of how research funders prioritize support for research software. This information is needed to clarify where current funder support is concentrated and where strategic gaps may exist. Here, we present data from a survey of research software funders (n=36) from around the world. The survey explored these funders’ priorities, finding a strong emphasis on developing skills, software sustainability, embedding open science, building community and collaboration, advancing research software funding, increasing software visibility and use, innovation and security. </p> <h1>Methods</h1> <p>This research was carried out using a survey combining qualitative and quantitative items. The survey was designed to investigate how research software funders support research software’s sustainability and impact.</p> <p>The study was reviewed and given an exempt determination by the University of Illinois Urbana-Champaign Institutional Review Board (no. 24374).</p> <h2>Survey design</h2> <p>The survey designed for this study began by collecting profile information, including institutional affiliation and job title. The survey gathered information about respondents’ organization’s initiatives, policies, or programs to support research software. The range of questions yielded too much data for one article. In this article, we focus exclusively on the results generated via an open-ended question asking about the top priorities for the respondents’ organizations’ support for research software: “What are your organization's top priorities related to research software?”. Four open-response text boxes were provided for respondents to indicate and list these priorities.</p> <h2>Sampling</h2> <p>This survey was aimed at international research funders, including governmental and non-governmental (e.g., philanthropic) funders. A list of contacts to invite to participate in this survey was created based on participation in the Research Software Association (ReSA) and responsibility for research software funding known to the authors. This initial list of people was refined, with removals based on individuals having moved to unrelated professional roles or being unavailable long-term, for example, due to personal issues.</p> <p>The final, refined contact list comprised 71 people. After removing individuals when a member of their organization already provided a complete answer or when the person turned out to no longer be working on a relevant topic or to be otherwise unavailable (total of n=30), 41 people remained. Five of these individuals did not complete the survey, while 36 people (representing 30 research funding organizations) did, yielding a response rate of 87.8%. Fully completed survey responses were not required for individuals to be retained in the sample, resulting in varied sample bases across survey questions.</p> <p>The sample includes research funders in North and South America, Europe, Oceania and Asia, but over-represents North America and European funder representatives. Some participating funders cover a broad spectrum of disciplines, while others focus on a particular domain such as social science, health, environment, physical sciences or humanities.</p> <table> <tbody> <tr> <td> <p><strong>Continent</strong></p> </td> <td> <p><strong>Count</strong></p> </td> </tr> <tr> <td> <p>North America</p> </td> <td> <p>15</p> </td> </tr> <tr> <td> <p>South America</p> </td> <td> <p>4</p> </td> </tr> <tr> <td> <p>Europe</p> </td> <td> <p>12</p> </td> </tr> <tr> <td> <p>Oceania</p> </td> <td> <p>3</p> </td> </tr> <tr> <td> <p>Asia</p> </td> <td> <p>1</p> </td> </tr> </tbody> </table> <p>The respondents represented research funders supported by governmental (n=26), philanthropic (n=6) and corporate (n=1) resources.</p> <p>Respondents’ job titles span the following categories: <em>Senior Leadership and Executive</em>, such as a Vice President of Strategy; <em>Program and Project Management</em>, such as Senior Program Manager; <em>Planning and Business Development</em>; <em>Scientific, Technical and IT</em>, such as Scientific Information Lead.</p> <p>Most respondents 72.7% (n=24) answered ‘Yes’ to the question, “Has your organization established any policies, initiatives or programs aimed at supporting research software?”, while 18.2% (n=6) said ‘No’ and 9.1% (n=3) ‘Unsure’.</p> <h2>Data collection, management and analysis</h2> <p>Data collection took place from December 2023 to May 2024. The mean completion time for the detailed survey was 28 minutes and 13 seconds.</p> <p>The data were cleaned and prepared for analysis by removing any identifiable respondent details. The data analysis process followed a standard thematic qualitative analysis approach (e.g., Jensen & Laurie, 2016). This involved first identifying themes and organizing the data accordingly. Dimensions of each theme were identified where relevant. Then data extracts were selected from the survey responses associated with each theme and theme dimension. </p> <h1>Additional data: Evolving funding strategies for research software: Insights from an international survey of research funders</h1> <p>Data were uploaded in December 2024 to support another paper drawing on the same overall survey data. This one is entitled: 'Evolving funding strategies for research software: Insights from an international survey of research funders'. The survey data for this upload were generated using the following survey items.</p> <table> <tbody><tr> <td> <p><strong><span>Variable</span></strong></p> </td> <td> <p><strong><span>Survey Item</span></strong></p> </td> <td> <p><strong><span>Response Options</span></strong></p> </td> </tr> </tbody><tbody> <tr> <td> <p><span>Policies, initiatives, or programs aimed at supporting research software</span></p> </td> <td> <p><span>“Has your organization established any policies, initiatives or programs aimed at supporting research software?”<br>(This could include grants, fellowships, funding policies, conference funding, or other kinds of support aimed at bolstering the sustainability or impact of research software)</span></p> </td> <td> <p><span>Yes, No, Unsure</span></p> <p><span>(If ‘Yes’, then the next question was asked)</span></p> </td> </tr> <tr> <td> <p><span>Number of policies or programs to be reported</span></p> </td> <td> <p><span>“How many of your organization’s policies, initiatives or programs to support research software are you familiar with?”</span></p> </td> <td> <p><span>1, 2, 3, 4, 5+</span></p> </td> </tr> <tr> <td> <p><em><span>The following questions were asked for each policy, initiative, or program</span></em></p> </td> </tr> <tr> <td> <p><span>Name of policy or program</span></p> </td> <td> <p><span>“Please name the policy, initiative or program (starting with the one you are most familiar with):”</span></p> </td> <td> <p><span>[Text line]</span></p> </td> </tr> <tr> <td> <p><span>Status of policy or program</span></p> </td> <td> <p><span>“What is the status of this policy, initiative or program?”</span></p> </td> <td> <p><span>Completed/closed, In progress/open, Other (please specify)</span></p> </td> </tr> <tr> <td> <p><span>Link(s)/description</span></p> </td> <td> <p><span>“Please provide link(s) to the policy, initiative or program, upload or email to [the researcher’s contact details].”<br>“Link(s)/Description:”<br>(If there is no documentation available, please describe it here:)</span></p> </td> <td> <p><span>[Textarea], [File upload]</span></p> </td> </tr> <tr> <td> <p><span>Type of policy or program</span></p> </td> <td> <p><span>“Which of the following best describes the policy, initiative or program you named above?”</span></p> </td> <td> <p><a name="_Hlk180534652"></a><span>Funding program, Policy that affects funding decision-making or outcomes (funder side), Policy that affects funding applicants or recipients (applicant/awardee side), Other (please specify)</span></p> </td> </tr> <tr> <td> <p><em><span>If ‘Funding program’ was selected in the previous question, then the next question was asked</span></em></p> </td> </tr> <tr> <td> <p><span>Type of funding</span></p> </td> <td> <p><span>“Which of the following best describes the available funding?”</span></p> </td> <td> <p><span>Funding that <strong>includes</strong> research software, Dedicated funding <strong>only</strong> for research software, Other (please specify)</span></p> </td> </tr> <tr> <td> <p><em><span>For all categories of policy, initiative or program, the following questions were asked.</span></em></p> </td> </tr> <tr> <td> <p><span>Problem(s) addressed</span></p> </td> <td> <p><span>“Please summarize the problem(s) this policy, initiative or program is aiming to address from your organization’s perspective:”</span></p> </td> <td> <p><span>[Text Area]</span></p> </td> </tr> <tr> <td> <p><span>Perceived level of program success</span></p> </td> <td> <p><span>“What factors have contributed to its success or lack of success?”</span></p> </td> <td> <p><span>Very successful, Successful, Neutral, Unsuccessful, Very unsuccessful, Not applicable / No opinion</span></p> </td> </tr> </tbody> </table>
Research Software focus area Maturity Model (RSMM) dataset
<p><span>The Research Software project focus area Maturity Model (RSMM) dataset provides a comprehensive description of 79 practices of RSMM. </span><span>This description includes </span><span>when the practices are implemented and is organized based</span><span> on the MoSCoW prioritization (Must have, Should </span><span>have, Could have, Won’t have). Additionally, it details </span><span>the resources required for execution, dependencies among </span><span>neighboring practices, and references.</span></p>
Dataset for "Evaluation Methods and Replicability of Software Architecture Research Objects"
<p># Content<br> In this package, please find the following content:</p> <p>* Investigated Papers.bib<br> A BibTeX file with all papers investigated in the paper "Evaluation Methods and Replicability of Software Architecture Research Objects"<br> * Raw-Data-Table Content-Data.html and Raw-Data-Table Meta-Data.html<br> Tables with the raw data as extracted during the systematic literature review<br> * Colection of Data Visualizations.pdf<br> Multiple visualizations of the raw data for analysis. A copy of summary.pdf as described below.<br> * Data and Visualization<br> Contains:<br> - The data as CSV files,<br> - scripts for creating visualizations<br> - *.awk -- Awk scripts are used to create the corresponding of the *.csv files in data<br> - *.rb -- Ruby scripts to build the respective figures in figs as *.tex files<br> - make-all.sh -- A script to call all other scripts for creating diagrams and the summary<br> - make-paper-figures.sh -- A script to build "paper-figures.pdf" with all diagrams used in the accompanying paper<br> - A documentation of the contained scripts (Data and Visualization/README.md)<br> - summary.pdf -- A collection of diagrams (as built by make-all.sh)<br> - paper-figures.pdf -- A collection of all diagrams as used in the accompanying paper (as built by make-paper-figures.sh and make-all.sh)<br> * Wiki/<br> A copy of the wiki used during data extraction.<br> Constains:<br> - descriptions of all data items<br> - the process description<br> - the taxonomy used for data extraction</p> <p><br> # Reproduction<br> You can reproduce the visualizations with the following commands in a UNIX command line environment.</p> <p>> cd "Data and Visualization"<br> > ./make-paper-figures.sh<br> > ./make-all.sh</p> <p>The requirements are:<br> * A UNIX command line environment (e.g., bash) with awk installed<br> * Ruby (>2.5)<br> * latex (e.g., tex-live)</p> <p>The command "./make-paper-figures.sh" produces the file “paper-figures.pdf”, which contains all diagrams that are used in the paper.<br> The command "./make-all.sh" produces the file "summary.pdf", which contains diagrams used for data analysis, and the file "paper-figures.pdf". All figures describing the results in the paper are also in "summary.pdf".<br> These commands each take about 2 ("/make-paper-figures.sh") / 8 ("./make-all.sh") minutes to run on current standard laptop (Intel i5-8250U, 16 GB memory).<br> Calling the commands produces many log statements (information and warnings), which show the progress and can be ignored.</p>
User Base of rOpenGov's Music Research Software
<p>The Digital Music Observatory has already created open-source software for the music industry that had been tested in real-life policy advocacy and business cases and scientific uses related to piracy research. While developed with a clear music industry focus, they have found thousands of users in the open research community worldwide for other purposes, too. We aim to further improve them to work as as a software ecosystem, and whenever possible, add web-based application interfaces with the ambition to make them useable for music organization that do not possess in-house R&D, IT, or data science capacities.<br> <br> The datasets contains the download statistics of these packages from CRAN.</p>
On a Microservice System Benchmark with Multiple Architected Variants for Software Engineering Research
Open the record for dataset details and reuse information.
Replication package for: An Analysis of Positionality Statements in Software Engineering Research
Open the record for dataset details and reuse information.
Research Results on SATD in Quantum Software
<p>This is a replication package for research on quantum-specific SATDs (QSATDs). It consists of three directories.</p> <ul> <li>`src/`: Contains source code for collecting data from GitHub, detecting Python file comments, and detecting SATD from them.</li> <li>`RQs_result/`: The results of manual coding to answer each RQ of this study are recorded as Excel files.</li> <li>`data/`: Archive of data that we have actually analyzed. </li> </ul> <p>Please check README.md for details of the package.</p>
The Research Software Stack
<p>A description of the differnet layers that underlying research computing in bioinformatics (from top down: research computing, bioinformatics tools & libraries, scientific libraries and general purpose computing) with the people involved in creating them (posgraduates and faculty, research software engineers and systems engineers).</p>
Common in Research Software, Hardware and Data: the experience of the SEEKCommons project (recording)
<p>Luis Felipe R. Murillo, professor of anthropology at University of Notre Dame, delivered this keynote address on August 1st, 2024, at the Ethical Open Science for Past Global Change Data 2024 Symposium, in Keshena, Wisconsin, on the lands of the Menominee Nation. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.