Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,916

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,916 results for “software,”

Learn how ShareScore rates datasets ↗
zenodo44/100

Fuzzy modelling and mapping soil moisture in Germany, link to research data and scientific software

<p>Research data and scientific software related to spatio-temporal estimations of ecological soil moisture with available data covering the whole territory of Germany and the Kellerwald National Park (Hesse). Temporal trends of modelled soil moisture for the time period 1961&ndash;2070 were statistically analyzed. Soil moisture changes (drying-out) at both national and regional levels were mapped.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Prädiktive Kartierung und Analyse klimawandelbedingter Veränderungen von Wäldern in Deutschland, Link zu Forschungsdaten und wissenschaftlicher Software

<p>Forschungsdaten zu einer Projektion von klimawandelbedingten Ver&auml;nderungen von Wald&ouml;kosystemen in Deutschland (1961-90, 1991-2010, 2011-40, 2041-70) und m&ouml;glicher &Auml;nderungen von ausgew&auml;hlten Lebensraumtypen nach Anhang I der Fauna-Flora-Habitat-Richtlinie.</p>

opencc-by-4.0Feb 2016View details →
zenodo44/100

research-software/resosuma-data: 0.4.1

<p><strong><em>resosuma-data</em></strong> represents activities in the research software sustainability space in the CSV format, where column 1 contains actors in the space, column 2 contains activities, and column 3 contains actees.</p> <p>resosuma-0.4.1.csv fixes issues <a href="https://github.com/research-software/resosuma-data/issues/3">#3</a> and <a href="https://github.com/research-software/resosuma-data/issues/4">#4</a>.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo44/100

The Software Heritage Graph Dataset

<p>Software Heritage is the largest existing public archive of software source<br> code and accompanying development history: it currently spans more than five<br> billion unique source code files and one billion unique commits, coming from<br> more than 80 million software projects.</p> <p>This is the Software Heritage graph dataset: a fully-deduplicated<br> Merkle DAG representation of the Software Heritage archive. The dataset links<br> together file content identifiers, source code directories, Version Control<br> System (VCS) commits tracking evolution over time, up to the full states of VCS<br> repositories as observed by Software Heritage during periodic crawls. The<br> dataset&rsquo;s contents come from major development forges (including GitHub and<br> GitLab), FOSS distributions (e.g., Debian), and language-specific package<br> managers (e.g., PyPI). &nbsp;Crawling information is also included, providing<br> timestamps about when and where all archived source code artifacts have been<br> observed in the wild.</p> <p>The Software Heritage graph dataset is available in multiple formats, including<br> downloadable CSV dumps and Apache Parquet files for local use, as well as a<br> public instance on Amazon Athena interactive query service for ready-to-use<br> powerful analytical processing.</p> <p>By accessing the dataset, you agree with the Software Heritage&nbsp;<a href="https://www.softwareheritage.org/legal/users-ethical-charter/">Ethical Charter<br> for using the archive data</a>, and the&nbsp;<a href="https://www.softwareheritage.org/legal/bulk-access-terms-of-use/">terms of use for bulk access</a>.</p> <p>If you use this dataset for research purposes, please cite the following paper:</p> <ul> <li>Antoine Pietri, Diomidis Spinellis, Stefano Zacchiroli.&nbsp;<br> <em>The Software Heritage Graph Dataset: Public software development under one roof</em>.&nbsp;<br> In proceedings of&nbsp;<a href="http://2019.msrconf.org/">MSR 2019</a>: The 16th International Conference on Mining Software Repositories, May 2019, Montreal, Canada. Co-located with&nbsp;<a href="https://2019.icse-conferences.org/">ICSE 2019</a>.&nbsp;<br> <a href="https://upsilon.cc/~zack/research/publications/msr-2019-swh.pdf">preprint</a>,&nbsp;<a href="https://upsilon.cc/~zack/research/publications/msr-2019-swh.bib">bibtex</a></li> </ul> <p>You can also refer to the above paper for more information the dataset and sample queries.</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Compound annual growth rate for software: replication package

<p>This repository contains the reproducibility package (software and data) for the following paper.</p> <p>Les Hatton, Diomidis Spinellis, and Michiel van Genuchten. The long-term growth rate of evolving software: Empirical results and implications. <em>Journal of Software: Evolution and Process</em>, 29(5), May 2017. <a href="http://dx.doi.org/10.1002/smr.1847">doi:10.1002/smr.1847</a></p> <p>The amount of code in evolving software-intensive systems appears to be growing relentlessly, affecting products and entire businesses. Objective figures quantifying the software code growth rate bounds in systems over a large time scale can be used as a reliable predictive basis for the size of software assets. We analyze a reference base of over 404 million lines of open source and closed software systems to provide accurate bounds on source code growth rates. We find that software source code in systems doubles about every 42 months on average, corresponding to a median compound annual growth rate (CAGR) of 1.21&plusmn;0.01. Software product and development managers can use our findings to bound estimates, to assess the trustworthiness of road maps, to recognise unsustainable growth, to judge the health of a software development project, and to predict a system&rsquo;s hardware footprint.</p> <p>&nbsp;</p>

openapache2.0Jan 2017View details →
zenodo44/100

[Dataset] Software Process Line as an Approach to Support Software Process Reuse: a Systematic Literature Review

<p>Dataset of a&nbsp;Systematic Literature Review on Software Process Line as an Approach to Support Software Process Reuse</p> <p>Please, read README.txt file before go through dataset.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Dataset and replication package for Temporal Discounting in Software Engineering: A Replication Study

<p>Dataset and replication package for the paper Temporal Discounting in Software Engineering: A Replication Study (Fagerholm, F., Becker, C., Chatzigeorgiou, A., Betz, S., Duboc, L., Penzenstadler, B., Mohanani, R., Venters, C. (2019). Temporal Discounting in Software Engineering: A Replication Study. 13th ACM/IEEE International Symposium of Empirical Software Engineering and Measurement (ESEM 2019)). The dataset consists of answers to a questionnaire on temporal discounting in a technical debt context. Two questionnaire templates illustrate how to gather the data for professional and student participants. An analysis script is provided which shows the details of the calculations and analyses performed for the paper. More information is given in the description file.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Citations to software and data in Zenodo via open sources

<p>In January 2019, the Asclepias Broker harvested citation links to Zenodo objects from three discovery systems: the NASA Astrophysics Datasystem (ADS), Crossref Event Data and Europe PMC. Each row of our dataset represents one unique link between a citing publication and a Zenodo DOI. Both endpoints are described by basic metadata. The second dataset contains usage metrics for every cited Zenodo DOI of our data sample.&nbsp;</p> <p>&nbsp;</p>

opencc-zeroOct 2019View details →
zenodo44/100

A Modified Doyle-Fuller-Newman Model Enables the Macroscale Physical Simulation of Dual-ion Batteries - Dataset and Software

<p>This dataset contains:</p> <p>- all the raw cycling data of the three-electrode cell used to gather the experimental data for the model validation (VMP data, exported with EC-LAB);<br>- the specific, processed data used in the model validation step (0.2C discharge, 5C discharge, EIS data);<br>- the COMSOL dual-ion battery model (version 6.0). IMPORTANT: activate the "Electric potential at the positive electrode current collector (only for EIS)" boundary condition when simulating impedance spectroscopy, and deactivate it when simulating charge/discharge curves; the charge-discharge profile can be modified by changing the duration of the test, the C-rate, and the conditions set in the "Events" section.</p> <p>Update: Fixed the model to work also in the 6.2 version of COMSOL (Substituted Dleff with Dleffxx in the modified weak expression of the cathode mass conservation equation). Download the new version!</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Replication package for paper: Insights on the Use of Software Design Principles in Machine Learning Pipelines

<p>This is the replication package of the paper "Insights on the Use of Software Design Principles in Machine Learning Pipelines".</p> <p>This replication package contains two files:</p> <ul> <li><a href="../api/records/13828806/draft/files/Data%20extraction.xlsx/content" target="_blank" rel="noopener noreferrer">Data extraction.xlsx</a>: file containing the details of the extracted data for each single ML project.&nbsp;</li> <li><a href="../api/records/13828806/draft/files/Source%20Code%20and%20Metadata.zip/content" target="_blank" rel="noopener noreferrer">Source Code and Metadata.zip</a>: zip file including the source code local copy analyzed and the repository metadata (.json) provided by GitHub API for each ML project .repository&nbsp;</li> </ul> <p>Reference: [1] Lidia L&oacute;pez, Cristina G&oacute;mez, and Claudia Ayala. Insights on the Use of Software Design Principles in Machine Learning Pipelines. <em>Accepted </em>in the 2024 edition of the International Conference on Product-Focused Software Process Improvement (PROFES 2024).</p> <p><strong>Note</strong>: The licence is applicable to the excel file. "Source Code and Metadata.zip" file contains source code repositories downloaded from GitHub, the license for each repository is defined in the corresponding GitHub repository by their authors.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Data and Software for: 'A NICER View of PSR J1231−1411: A Complex Case'

<p>Posterior sample files associated with the publication "A NICER View of PSR J1231&minus;1411: A Complex Case" by Salmi et al. (2024b;&nbsp;<a href="https://doi.org/10.48550/arXiv.2409.14923">arXiv.2409.14923</a>; <a href="https://doi.org/10.3847/1538-4357/ad81d2">https://doi.org/10.3847/1538-4357/ad81d2</a>).</p> <p>Also included are: the data products; the numeric model files including the telescope calibration products; model modules in the Python language using the X-PSI framework; and Jupyter analysis notebooks.</p> <p>Please refer to the README for detailed information.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Replication Package for the Paper Titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction"

<p>This is a replication package for the paper titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction".</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Software, Dataset, and Techreport: Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration

<p>This upload contains a techreport titled "Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration" together with the software (with documentation) and dataset generating the results. The software is also available on GitHub at https://github.com/croci/mpfem-paper-experiments-2024/ . The GitHub version may be updated in the future. This upload corresponds to commit number 8506dd368b84655201c8c72b1307239b9b4e43fd . See README.md file for installation instructions. The manuscript is also available on the arXiv: https://arxiv.org/abs/2410.12614.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Open dataset for publication "Systematic Mapping Study on Requirements Engineering for Regulatory Compliance of Software Systems"

<p>This publication contains open dataset for the journal publication "Systematic Mapping Study on Requirements Engineering for Regulatory Compliance of Software Systems".</p> <p>The dataset contains the data extracted from 280 selected primary studies.</p> <p>The dataset includes the following data:</p> <ul> <li>study metadata (title, venue, publication year, authors, authors&rsquo; affiliation, abstract);</li> <li>challenges to regulatory compliance (direct excerpts from studies);</li> <li>categories of challenges to compliance;</li> <li>principles and practices (direct excerpts from text);</li> <li>categories of principles and practices;</li> <li>types of automation of principles and practices;</li> <li>involved stakeholders (direct excerpts from studies);</li> <li>categories of involved stakeholders;</li> <li>phase of the principle and practice life cycle for which involvement of stakeholders was considered;</li> <li>SDLC process areas covered by the study;</li> <li>regulations considered in the study;</li> <li>fields of regulations that were considered;</li> <li>domains of application that were considered;</li> <li>assessment of rigor and relevance of the study.</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Software Engineering Education Knowledge versus Industrial Needs

<p>Dataset of the research paper:&nbsp;<strong>Software Engineering Education Knowledge versus Industrial&nbsp;Needs</strong></p> <p><em>Contribution</em>: Determine and analyze the gap between software practitioners&rsquo; education outlined in the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and industrial needs pointed by Wikipedia articles referenced in Stack Overflow (SO) posts.<br> <em>Background</em>: Previous work has uncovered deficiencies in the coverage of computer fundamentals, people skills, software processes, and human-computer interaction, suggesting rebalancing.<br> <em>Research Questions</em>: 1) To what extent are developers&rsquo; needs, in terms of Wikipedia articles referenced in SO posts, covered by the SEEK knowledge units? 2) How does the popularity of Wikipedia articles relate to their SEEK coverage? 3) What areas of computing knowledge can be better covered by the SEEK knowledge units? 4) Why are Wikipedia articles covered by the SEEK knowledge units cited on SO?<br> <em>Methodology</em>: Wikipedia articles were systematically collected from SO posts. The most cited were manually mapped to the SEEK knowledge units, assessed according to their degree of coverage. Articles insufficiently covered by the SEEK were classified by hand using the 2012 ACM Computing Classification System. A sample of posts referencing sufficiently covered articles was manually analyzed. A survey was conducted on software practitioners to validate the study findings.<br> <em>Findings</em>: SEEK appears to cover sufficiently computer science fundamentals, software design and mathematical concepts, but less so areas like the World Wide Web, software engineering components, and computer graphics. Developers seek advice, best practices and explanations about software topics, and code review assistance. Future SEEK models and the computing education could dive deeper in information systems, design, testing, security, and soft skills.</p> <p>The following data files are included.</p> <ul> <li><strong>wikipedia_articles.csv</strong>: Wikipedia articles mapped to the knowledge units of the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and the first and second level categories of the 2012 ACM Computing Classification System (CCS).</li> <li> <p><strong>posts_analysis.csv</strong>: Stack Overflow post data and metadata.</p> </li> <li> <p><strong>posts_aggregated_codes.csv</strong>: The aggregated codes that resulted from the manual analysis of the Stack Overflow posts by grouping individual keywords assigned to the posts.</p> </li> <li> <p><strong>survey_questionnaire.csv</strong>:&nbsp;The final survey questionnaire.</p> </li> <li> <p><strong>survey_responses.csv</strong>:&nbsp;Anonymized responses of the final survey questionnaire. (E-mail addresses have been excluded for privacy reasons.)</p> </li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Replication Package "Applying Test Case Prioritization to Software Microbenchmarks"

<p>Replication package for the paper &quot;Applying Test Case Prioritization to Software Microbenchmarks&quot; accepted for publication in Empirical Software Engineering.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Softcite software mention extraction from the CORD-19 publications

<p><strong>Softcite software mention extraction from the CORD-19 publications </strong></p> <p>This dataset is the result of the extraction of software mentions from the set of publications of the CORD-19 corpus (<a href="https://allenai.org/data/cord-19">https://allenai.org/data/cord-19</a>) by the Softcite software recognizer (SciBERT-CRF fine-tuned model), see <a href="https://github.com/ourresearch/software-mentions">https://github.com/ourresearch/software-mentions</a>.</p> <p>The CORD-19 version used for this dataset is the one dated <strong>2021-07-26,</strong> using the <em>metadata.csv</em> file only. We re-harvested the PDF with <a href="https://github.com/kermitt2/article-dataset-builder">https://github.com/kermitt2/article-dataset-builder</a> in order to also extract coordinates of software mentions in the PDF and to take advantage of the latest version of GROBID to produce better full text extraction from PDF. We also harvested 61,230 full-texts more than the standard CORD-19 distribution.</p> <p>Note that this is the third version of this dataset (version 0.3.0). The previous Softcite software mention extraction from the CORD-19 was based on 2020-09-11 and 2021-03-22 versions. The new version cover a larger set of documents and is using an improved version of the extraction tools.</p> <p><strong>Data format </strong></p> <p>The extraction consists of 3 JSON files:</p> <p><strong>annotations.jsonl</strong> contains the individual software annotations including <em>software name</em> and possible attached attributes (<em>publisher</em>, <em>URL</em> and <em>version</em>). Each annotation is associated with coordinates expressed as bounding boxes in the original PDF. See <a href="https://grobid.readthedocs.io/en/latest/Coordinates-in-PDF/">Coordinates of structures in the original PDF</a>&nbsp; for more details on the coordinate format.</p> <p>The context of citation is the sentence where the software name and its attributes are extracted. It is added to the JSON structure (field <em>context</em>), as well as the identifier of the document where the annotation belongs (field <em>document</em>, pointing to entries available in <em>documents.json</em>) and a list of bibliographical references attached to the software name (field <em>references</em>, pointing to entries available in <em>references.json</em>, with the used reference marker string). See <a href="https://github.com/ourresearch/software-mentions">https://github.com/ourresearch/software-mentions</a> for more details on the extracted attributes.</p> <p>If the software name was sucessfully disambiguated against WikiData (&quot;entity linking&quot;), it appears in the field <em>wikidataId</em> as Wikidata entity identifier and in the field <em>wikipediaExternalRef</em> as a Wikipedia PageID from the English Wikipedia. Entity linking is realized with <a href="https://github.com/kermitt2/entity-fishing">entity-fishing</a>.</p> <p><strong>documents.jsonl</strong> contains the metadata of the all the CORD-19 documents containing at least one software annotation. The metadata are given as a CrossRef JSON structure. The abstract should be included in the metadata most of the time, as well as some complements extracted by GROBID directly from the PDF. In addition, the size of the pages and the unique file path to the PDF can be found to allow annotations directly on the PDF (see <a href="https://grobid.readthedocs.io/en/latest/Coordinates-in-PDF/">Coordinates of structures in the original PDF</a> for more details on the PDF annotation display mechanism).</p> <p><strong>references.jsonl </strong>contains the parsed reference entries associated to software mentions. These references are given in the field <em>tei</em> encoded in the XML TEI format of GROBID extraction. The extracted raw references have been matched against CrossRef to get a DOI and more complete metadata with <a href="https://github.com/kermitt2/biblio-glutton">biblio-glutton</a>.</p> <p><strong>Statistics</strong></p> <p>CORD-19 version: 2021-07-26</p> <p>- total Open Access full texts: 296,686<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - with at least one software mention: 115,073</p> <p>- total software name annotations: 652,518<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - with linked Wikidata ID: 231,599</p> <p>- associated field&nbsp;&nbsp;&nbsp;<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - publisher: 107,421<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - version: 188,724<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; - URL: 59,366<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; - references: 230,145</p> <p>- associated bibliographical references: 92,573<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - references with matched DOI: 49,350<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; - references with matched PMID: 32,895<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - references with matched PMC ID: 18,741</p> <p><strong>License and acknowledgements</strong></p> <p>This dataset is licensed under a Creative Commons Attribution 4.0 International License.</p> <p>We thank the Alfred P. Sloan Foundation and of the Gordon and Betty Moore Foundation for supporting this work.</p>

opencc-by-4.0May 2021View details →
zenodo44/100

Relação de artigos científicos sobre Humanidades Digitais obtidos pelo software Google Scholar Crawler

<p>Resultado do processamento do Google Scholar Crawler para o termo &quot;Humanidades Digitais&quot;. Neste arquivo j&aacute; foram retirados os itens duplicados e os falsos positivos (artigos que n&atilde;o consideramos como sendo de Humanidades Digitais.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Software repository for data-driven reconstruction of doping profiles in semiconductors

<p>Datasets and code used described in paper: &quot;Data-driven solutions of ill-posed inverse problems arising from doping reconstruction in semiconductors&quot; [arXiv:2208.00742]</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

OpenEBench Software Observatory data dump

<p>Data dump of the OpenEBench Software Observatory, consisting in&nbsp;metadata of software in the Life Sciences domain extracted from several sources. All collections contain the same information at different points in the processing pipeline:</p> <p>- <em>alambique</em> collection contains&nbsp;metadata as extracted from sources.</p> <p>- <em>pretools</em>&nbsp;is the harmonized version of such metadata.</p> <p>- <em>tools&nbsp;</em>is the final integrated collection of software metadata used to identify trends and perform&nbsp;FAIRness evaluations.&nbsp;</p> <p><strong>Repository</strong> containing the code used to generate the datasets:&nbsp;<a href="https://gitlab.bsc.es/inb/elixir/software-observatory/FAIRsoft_ETL">https://gitlab.bsc.es/inb/elixir/software-observatory/FAIRsoft_ETL&nbsp;</a></p>

opencc-by-4.0Nov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record