Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,916
datasets available to search
ShareScore release 0.7.1
Dataset results
1,916 results for “software,”
The Stanford Software Survey, 2020
<p>This is the Stanford Software Survey, interface and results for the year 2020.</p> <p>https://github.com/stanford-rc/stanford-software-survey/releases/tag/2020</p>
LabelGit: A dataset for software repositories classification using attributed dependency graphs
<p>A dataset for software repositories classification using attributed dependency graphs</p>
Supplemental Material: What we talk about when we talk about software test flakiness
<p>This supplemental material details the definitions of the concepts that have been found by conducting the scoping review of both the white and grey literature introduced in Section 2 of the manuscript titled: <strong>"What we talk about when we talk about software test flakiness</strong>".</p> <p> </p>
UK Research Software Survey 2014
<p>This spreadsheet contains the anonymised data collected as part of a survey of UK researchers in their use of research software.</p> <p>We asked people specifically about “research software” which we defined as:</p> <blockquote> <p>“Software that is used to generate, process or analyse results that you intend to appear in a publication (either in a journal, conference paper, monograph, book or thesis). Research software can be anything from a few lines of code written by yourself, to a professionally developed software package. Software that does not generate, process or analyse results - such as word processing software, or the use of a web search - does not count as ‘research software’ for the purposes of this survey.”</p> </blockquote> <p>We contacted 1,000 randomly selected researchers at each of 15 Russell Group universities. From the 15,000 invitations to complete the survey, we received 417 responses – a rate of 3% which is fairly normal for a blind survey. We used Google Forms to collect responses.</p> <p>The responses have good representation from across the disciplines, seniorities and genders. This is a statistically significant number of responses that can be used to represent the views of people in research-intensive universities in the UK.</p> <p>An overview of the data is available on the worksheet "Summary data". Responses to questions are ordered by unique respondent ID. Please read the "README" worksheet for additional information about the collection and processing of this data.</p> <p>This survey data is licensed under a Creative Commons by Attribution licence. Copyright resides with The University of Edinburgh on behalf of the Software Sustainability Institute.</p> <p>Please cite as:</p> <p><strong>APA</strong></p> <p>Hettrick. S. J., et al. (2014). UK Research Software Survey 2014 [Data set]. doi:10.5281/zenodo.14809</p> <p><strong>Chicago</strong></p> <p>S.J. Hettrick et al, UK Research Software Survey 2014 (accessed December 4, 2014), 10.5281/zenodo.14809.</p> <p><strong>MLA</strong></p> <p>Hettrick S.J., et al. “UK Research Software Survey 2014” ZENODO, 2014. Web. 4 December 2014. .</p>
The Software Sustainability Institute's Collaborations Workshop 2015 (CW15) attendees computational tools dataset
<p>Contains the question, raw data, and cleaned data for producing the most used software word cloud for those who attended the Software Sustainability Institute's Collaborations Workshop 2015 (CW15) held at the Oxford e-Research Institute, Oxford, UK from 25-27 March 2015</p>
SURF: Replication Package for: "What Would Users Change in My App? Summarizing App Reviews for Recommending Software Changes"
<p>Description of the content of folder "SURF_replication_package": 1) "Experiment I" contains: a) the folder "summaries" which contains all the html summaries generated through SURF and browsed by study participants involved in the Experiment I. b) the folder "XMLreviews" which contains, for each of the apps involved in the Experiment I, the corresponding XML file containing all the collected reviews for that app. These xml files have been used as input files for the SURF tool for generating the summaries contained in the "summaries" folder c) "Experiment_I_results.xlsx" which contains all the answers to our survey collected from the Experiment I participants.</p> <p>2) "Experiment II" contains: a) the folder "summaries" which contains the two html summaries generated through SURF and browsed by study participants in the Experiment II. b) the folder "XMLreviews" which contains, for each of the two apps involved in the Experiment II, the corresponding XML file containing all the collected reviews for that app. These xml files have been used as input of the SURF tool for generating the summaries contained in the "summaries" folder. c) "Experiment_II_results.xlsx" which contains all the user feedbacks extracted/validated by survey participants in the two sub-experiments. d) "Experiment_II_survey_answers.xlsx" which contains all the answers to our survey collected in the Experiment II participants.</p> <p>3) "Survey.pdf" which contains the pdf version of the survey performed by the participants</p> <p>4) "SURF_tool.zip" contains: a) "SURF.jar", which contains the class files of a prototypical implementation of SURF b) "README.txt" which contains the instructions to run the SURF tool c) the "lib" folder, which contains all the java libraries needed for running SURF.</p>
Research Software Engineers Supporting Science: Survey Responses
<p>Raw survey data for "Not everyone can use git: Research Software Engineers’ recommendations for scientist-centred software support (and what researchers think of them)", a talk given by Caroline Jay at RSE16, Manchester, UK.</p>
Simulated benchmark metagenome used to demonstrate and evaluate MGLEX software
<p>This is a mock dataset of 120 000 artificial contigs of 1 kb length derived by simulating reads from 295 unique genomes and 44 species with each two or three strain genomes using the ART read simulator (Huang et al., 2012) and a lognormal abundance distribution. Genomes were chosen according to the CAMI2015 (www.cami-challenge.org) medium complexity toy dataset. The dataset contains four replicate samples with varied abundances and corresponding sequence feature files in MGLEX v0.1.1 format to use for genome reconstruction. Our aim was to create a benchmark dataset under controlled settings, minimizing potential biases introduced by specific software. This package also includes MGLEX benchmark scripts.</p>
On the Understandability of Semantic Constraints for Behavioral Software Architecture Compliance: A Controlled Experiment
<p>Software architecture compliance is concerned with the alignment of implementation with its desired architecture and detecting potential inconsistencies. The study is specifically concerned with behavioral architecture compliance. That is, the focus is on semantic alignment of implementation and architecture. In particular, the study evaluates three representative approaches for describing semantic constraints in terms of their understandability, namely natural language descriptions as used in many architecture documentations today, a structured language based on specification patterns that abstract underlying temporal logic formulas, and a structured cause-effect language that is based on Complex Event Processing. We conducted a controlled experiment with 190 participants using a simple randomized design with one alternative per experimental unit.</p>
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Matrix multiplication software and results bundle for paper "Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library" for P^3MA submission
<p>This is the archive containing the matrix multiplication software and the results of the publication "<em>Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library</em>" submitted to the P^3MA workshop 2017.</p> <p><strong>The archive has the following content:</strong></p> <ul> <li>Source code for the (tiled) matrix multiplication in "src": <ul> <li>regular version in "src/matmul": <ul> <li>Remote: https://github.com/theZiz/matmul.git (copy will be removed)</li> <li>Branch: topic-compatible-alpaka-0-1-0</li> <li>Commit: a63ba4810d6bfcca62c68dd57408af15028e78a3</li> </ul> </li> <li>forked version for XL in "src/matmul": <ul> <li>Remote: https://github.com/theZiz/matmul.git (copy will be removed)</li> <li>Branch: topic-xl-workaround</li> <li>Commit: 1fee028eccb8cf7b677e8071233e08aa9f81846a</li> </ul> </li> </ul> </li> <li>The compiled binaries and the results of the tuning and scaling runs are in "runs" in sub folders for each type of run and architectures.</li> </ul>
Data for: Drivers and Barriers for Microservice Adoption in the German Software Industry
<p>Microservices are an architectural style for software which currently receives a lot of attention in both industry and academia. Several companies employ microservice architectures with great success, and there is a wealth of blog posts praising their advantages. Especially so-called Internet-scale systems use them to satisfy their enormous scalability requirements and to rapidly deliver new features to their users.<br> However, microservices are not only popular with large, Internet-scale systems. Many traditional companies are also considering whether microservices are a viable option for their applications. However, these companies may have other motivations to employ microservices, and see other barriers which may prevent them from adopting microservices. Furthermore, these drivers and barriers may differ among industry sectors.<br> This dataset contains the questions and results of a survey on drivers and barriers for microservice adoption among professionals in the German software industry. In addition to overall drivers and barriers, we particularly focused on the use of microservices to modernize existing software, with special emphasis on implications for runtime performance and transactionality.</p>
NeuroGentoo Presentation - Bringing the Power of Gentoo Package Management to (Neuro)Scientific Software Environments
<p>Neuroscience, one of the most computation-reliant fields in the natural sciences, is dependent upon dozens of highly complex software suites, which scientists are often forced to manage manually. Upstream developers often ship bundled dependencies to better support this flawed workflow, and in the resulting mess documenting and reproducing analysis pipelines is neigh-impossible. We seek to correct these shortcoming of both modern software distribution as well as modern data science, by integrating high-quality ebuilds for neuroscientific software into the Gentoo Science Overlay. We also seek to publish a simple NuroGentoo world file (along with appropriate usage instructions) to allow scientists with access to OpenStack, Amazon Elastic Computing, or Docker to launch up and build a system fit for reproducing data analysis run on other NeuroGentoo systems with minimum effort.</p>
1151 commits with software maintenance activity labels (corrective,perfective,adaptive)
<p>Data format: CSV</p> <p>Separator character: '#'</p> <p><strong>This dataset contains 1151 commits manually labeled with maintenance activities ("c" for corrective, "p" for perfective, "a" for adaptive)</strong> according to the definition by Mockus et al. in <em>"Mockus, A. and Votta, L.G., 2000, October. Identifying Reasons for Software Changes using Historic Databases. In icsm (pp. 120-130)"</em>.</p> <p>In addition, this dataset also contains <strong>further information (features) extracted from the commits</strong>:</p> <ol> <li>The <strong>source code changes</strong> performed by the commit author as part of a given commit (statement added, statement removed, etc.) <ul> <li>The source code change taxonomy is detailed in <em>"Fluri, B. and Gall, H.C., 2006, June. Classifying change types for qualifying change couplings. In Program Comprehension, 2006. ICPC 2006. 14th IEEE International Conference on (pp. 35-45). IEEE."</em></li> </ul> </li> <li>A binary indication (1/0) whether a given commit contains any of the <strong>keywords from a pre-computed </strong>(according to a word frequency analysis)<strong> set of keywords</strong> <strong>indicative of each maintenance activity</strong>.</li> </ol> <p>The dataset consists of commits sampled from the following open source projects:</p> <ol> <li>RxJava</li> <li>hbase</li> <li>elasticsearch</li> <li>intellij-community</li> <li>hadoop</li> <li>drools</li> <li>kotlin</li> <li>restlet-framework-java</li> <li>orientdb</li> <li>camel</li> <li>spring-framework </li> </ol> <p>This dataset is a supporting material for the paper <strong>"Boosting Automatic Commit Classification Into Maintenance Activities By Utilizing Source Code Changes", to appear in PROMISE 2017.</strong></p>
SleepEEGpy: a Python-based software integration package to organize preprocessing, analysis, and visualization of sleep EEG data
<p>This dataset includes three high-density sleep EEG recordings of healthy participants, downsampled to 250 Hz and stored in FIF format:</p> <ol> <li>Nap recording of a young adult participant</li> <li>Overnight recording of a young adult participant</li> <li>Overnight recording of an older adult participant</li> </ol> <p>Additionally, the dataset includes three text files for each recording:</p> <ul> <li>bad_channels.txt: Indexes of noisy channels</li> <li>annotations.txt: Onset and duration of noisy temporal intervals</li> <li>staging.txt: Sleep staging vector</li> </ul> <p>The corresponding package can be found on <a href="https://github.com/NirLab-TAU/sleepeegpy">GitHub.</a></p> <p>For citation, please use:<br>Falach, R., G. Belonosov, J. F. Schmidig, M. Aderka, V. Zhelezniakov, R. Shani-Hershkovich, E. Bar, and Y. Nir. "SleepEEGpy: a Python-based software integration package to organize preprocessing, analysis, and visualization of sleep EEG data." Computers in Biology and Medicine 192 (2025): 110232.<br><a href="https://doi.org/10.1016/j.compbiomed.2025.110232" rel="nofollow">https://doi.org/10.1016/j.compbiomed.2025.110232</a></p>
Impact of Antipatterns on Software systems- Snowballing Process
<p> This study presents a systematic literature review that accumulates, summarizes, and reports the results of 97 relevant Primary Studies (PSs) concerning the impact of APs on Object Oriented (OO), Service Oriented (SO), and mobile Oriented (MO) software applications from 2005 to 2024 while considering several internal and external quality attributes. The PSs are classified based on the techniques used to find the impact of APs, type of datasets, evaluation measures, and tool support. </p>
Package and Dependency Metadata for CZI Hackathon: Mapping the Impact of Research Software in Science
<p>A collection of useful datasets extracted from <a href="https://packages.ecosyste.ms">https://packages.ecosyste.ms</a> and <a href="https://repos.ecosyste.ms/">https://repos.ecosyste.ms</a> for use at the CZI Hackathon: Mapping the Impact of Research Software in Science.</p><p>All data is provided as NDJSON (new line delimited JSON), each line represents a valid JSON object, and they are separated by newline characters. There are <a href="https://pypi.org/project/ndjson/">python</a> and <a href="https://www.rdocumentation.org/packages/ndjson/versions/0.9.0/topics/stream_in">R</a> libraries for reading these files, or you can maually read each line and parse each line as a single JSON object.</p><p>Each ndjson file has been compressed with gzip (actual command: `tar -czvf`) to reduce download size, they expand to significantly bigger files after extraction.</p><h4>Package Data</h4><p>Package names from cran, bioconductor and pypi that have been parsed by the <a href="https://github.com/chanzuckerberg/software-mentions">software-mentions</a> project (data: <a href="https://datadryad.org/stash/dataset/doi:10.5061/dryad.6wwpzgn2c">https://datadryad.org/stash/dataset/doi:10.5061/dryad.6wwpzgn2c</a>) are collected together with their latest release at time of publishing along with the names of their dependencies, those dependency names have then also been recursively fetched with latest release and dependencies until the full list of transitive dependencies is included. </p><p>Note: This approach uses a simplified method of dependency resolution, always picking the latest version of each package rather than taking into account each dependencies specific version range requirements, this is primarily due to time constraints and allows all software ecosystems to be processed in the same way. A future improvement would be to use each package ecosystem's specific dependency resolution algorithm to compute the full transitive dependency tree for each mentioned software package.</p><h4>GitHub Data</h4><p>Two different approaches were taken for collecting data for referenced GitHub mentions:</p><p>1. `github.ndjson` is metadata for each repository from GitHub, including "manifest" files which are known files that contain dependency information for a project such as requirements.txt, DESCRIPTION and package.json, parsed using <a href="https://github.com/ecosyste-ms/bibliothecary">https://github.com/ecosyste-ms/bibliothecary</a>, which may include transitive dependencies that have been discovered in a `lockfile` within the repository.</p><p>2. `github_packages.ndjson` is metadata for each package that was found on any package manager that references the GitHub url as it's repository url/source/homepage, these packages, like the cran and pypi data above, include the latest release and their direct dependencies. There may be more than one package for each GitHub URL as it is a one to many relationship. `github_packages_with_transitive.ndjson` follows the same format but also includes the extra resolved transitive dependencies of all packages using the same approach as with cran and pypi data above with the same caveats. </p><p>There are also many more ecosystems referenced in these files than just cran, bioconductor and pypi, https://packages.ecosyste.ms provides a standardized metadata format for all of them to enable comparison and simplification of automation.</p><h4>Contact</h4><p>If you would like any help, support or more data from Ecosyste.ms please do get in touch via email: hello@ecosyste.ms or open an issue on GitHub: https://github.com/ecosyste-ms/packages/issues</p>
Dataset - Understanding the software and data used in the social sciences
<p>This is a repository for a UKRI Economic and Social Research Council (ESRC) funded project to understand the software used to analyse social sciences data.</p><p>Any software produced has been made available under a BSD 2-Clause license and any data and other non-software derivative is made available under a CC-BY 4.0 International License. Note that the software that analysed the survey is provided for illustrative purposes - it will not work on the decoupled anonymised data set.</p><p>Exceptions to this are:</p><ul><li>Data from the UKRI ESRC is mostly made available under a <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">CC BY-NC-SA 4.0</a> Licence.</li><li>Data from Gateway to Research is made available under an <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/">Open Government Licence</a> (Version 3.0).</li></ul><h2>Contents</h2><ul><li>Survey data & analysis: esrc_data-survey-analysis-data.zip</li><li>Other data: esrc_data-other-data.zip</li><li>Transcripts: esrc_data-transcripts.zip</li><li>Data Management Plan: esrc_data-dmp.zip</li></ul><h3>Survey data & analysis</h3><p>The survey ran from 3rd February 2022 to 6th March 2023 during which 168 responses were received. Of these responses, three were removed because they were supplied by people from outside the UK without a clear indication of involvement with the UK or associated infrastructure. A fourth response was removed as both came from the same person which leaves us with 164 responses in the data.</p><p>The survey responses, Question (Q) Q1-Q16, have been decoupled from the demographic data, Q17-Q23. Questions Q24-Q28 are for follow-up and have been removed from the data. The institutions (Q17) and funding sources (Q18) have been provided in a separate file as this could be used to identify respondents. Q17, Q18 and Q19-Q23 have all been independently shuffled.</p><p>The data has been made available as Comma Separated Values (CSV) with the question number as the header of each column and the encoded responses in the column below. To see what the question and the responses correspond to you will have to consult the survey-results-key.csv which decodes the question and responses accordingly. </p><p><strong>A pdf copy of the survey questions is </strong><a href="https://github.com/softwaresaved/esrc-software-study/blob/main/Docs/esrc-survey.pdf"><strong>available</strong></a><strong> on GitHub.</strong></p><p>The survey data has been decoupled into:</p><ul><li>survey-results-key.csv - maps a question number and the responses to the actual question values.</li><li>q1-16-survey-results.csv- the non-demographic component of the survey responses (Q1-Q16).</li><li>q19-23-demographics.csv - the demographic part of the survey (Q19-Q21, Q23).</li><li>q17-institutions.csv - the institution/location of the respondent (Q17).</li><li>q18-funding.csv - funding sources within the last 5 years (Q18).</li></ul><p>Please note the code that has been used to do the analysis will not run with the decoupled survey data. </p><h3>Other data files included</h3><ul><li>CleanedLocations.csv - normalised version of the institutions that the survey respondents volunteered.</li><li>DTPs.csv - information on the UKRI Doctoral Training Partnerships (DTPs) scaped from the UKRI <a href="https://esrc.ukri.org/skills-and-careers/doctoral-training/doctoral-training-partnerships/doctoral-training-partnership-dtp-contacts/">DTP contacts</a> web page in October 2021.</li><li>projectsearch-1646403729132.csv.gz - data snapshot from the <a href="https://gtr.ukri.org/">UKRI Gateway to Research</a> released on the 24th February 2022 made available under an <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/">Open Government Licence</a>.</li><li>locations.csv - latitude and longitude for the institutions in the cleaned locations.</li><li>subjects.csv - research classifications for the ESRC projects for the 24th February data snapshot.</li><li>topics.csv - topic classification for the ESRC projects for the 24th February data snapshot.</li></ul><h3>Interview transcripts</h3><p>The interview transcripts have been anonymised and converted to markdown so that it's easier to process in general. List of interview transcripts:</p><ul><li>1269794877.md</li><li>1578450175.md</li><li>1792505583.md</li><li>2964377624.md</li><li>3270614512.md</li><li>40983347262.md</li><li>4288358080.md</li><li>4561769548.md</li><li>4938919540.md</li><li>5037840428.md</li><li>5766299900.md</li><li>5996360861.md</li><li>6422621713.md</li><li>6776362537.md</li><li>7183719943.md</li><li>7227322280.md</li><li>7336263536.md</li><li>75909371872.md</li><li>7869268779.md</li><li>8031500357.md</li><li>9253010492.md</li></ul><h3>Data Management Plan</h3><p>The study's Data Management Plan is provided in PDF format and shows the different data sets used throughout the duration of the study and where they have been deposited, as well as how long the SSI will keep these records. </p>
Understanding the Building Blocks of Accountability in Software Engineering
<p>In the social and organizational sciences, accountability has been linked to the efficient operation of organizations. However, it has received limited attention in software engineering (SE) research, in spite of its central role in the most popular software development methods (i.e., Scrum). In this article, we seek to explore the mechanisms of accountability in SE environments and investigate the factors that foster software engineers' individual accountability within their teams through an interview study with 12 people. Our findings recognize two primary forms of accountability shaping software engineers individual senses of accountability: \emph{institutionalized} and \emph{grassroots}. While the former is directed by formal processes and mechanisms, like performance reviews, grassroots accountability arises organically within teams, driven by factors such as peers' expectations and intrinsic motivation. This organic form cultivates a shared sense of collective responsibility, emanating from shared team standards and individual engineers' inner commitment to their personal, professional values, and self-set standards. While institutionalized accountability relies on traditional "carrot and stick" approaches, such as financial incentives or denial of promotions, grassroots accountability operates on reciprocity with peers and intrinsic motivations, like maintaining one's reputation in the team.</p>
Demonstration of semantic and inter-input constraints on software in OWL 2 and SPARQL for fulfilling the M1 Machine FAIR Use Case
<p>This video demonstrates using hypothetical examples how to (1) find a valid dataset for input into a software using OWL 2 classification inference, (2) validly combine two software using OWL 2 subsumption inference to infer that the output of software 1 is valid input to software 2, and (3) combine OWL 2 inference with a SPARQL query to find two datasets that satisfy a software's inter-input constraints.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.