Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

166

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

166 results for “open-source”

Learn how ShareScore rates datasets ↗
dryad36/100

Quantitative analysis of subcellular distributions with an open-source, object-based tool

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad36/100

Mapping built infrastructure in semi-arid systems using data integration and open-source approaches for image classification

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad36/100

Data from: The BumbleBox: An open-source platform for quantifying behavior in bumblebee colonies

Open the record for dataset details and reuse information.

publicApr 2025View details →
zenodo32/100

Security vulnerabilities in open-source reused systems

<p>This dataset comprise 2017 Java projects. It contains information related to their external dependencies and its&nbsp; potential and disclosed security vulnerabilities.</p> <p>The potential vulnerabilities were detected with the use of the SpotBugs static analyzer tool, while the disclosed ones with the use of OWASP Dependency Check tool..</p> <p>This dataset was generated during a research effort to correlate software reuse to security vulnerabilities.</p> <p>The scripts for reproducing the dataset and analyzing it are available on GitHub under this link [https://github.com/AntonisGkortzis/Vulnerabilities-in-Reused-Software].</p>

opencc-by-4.0Nov 2019View details →
zenodo32/100

[Dataset] Expanding the Number of Reviewers in Open-Source Projects by Recommending Appropriate Developers

<p>A rich collection of review and development data, including information about reviewers, developers and their<br> commits within five large ASF projects and four Gerrit communities.</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

An open-source nnU-net algorithm for automatic segmentation of MRI scans in the male pelvis for adaptive radiotherapy

<p>Data related to the article:</p> <p>Front. Oncol.</p> <p>Sec. Radiation Oncology</p> <p>Volume 13 - 2023 | doi: 10.3389/fonc.2023.1285725</p> <p>&nbsp;</p> <p>An open-source nnU-net algorithm for automatic segmentation of MRI scans in the male pelvis for adaptive radiotherapy</p> <p>Ebbe Laugaard Lorenzen 1,2*, Bahar Celik 1, Nis Sarup1, Lars Dysager3, Rasmus L&uuml;beck Christiansen1, Anders Smedegaard Bertelsen1, Uffe Bernchou1,2, S&oslash;ren Nielsen Agergaard1, Maximilian Lukas Konrad1, Carsten Brink1,2*, Faisal Mahmood1,2, Tine Schytte2,3,&nbsp;Christina Junker Nyborg3</p> <p>1 Laboratory of Radiation Physics, Department of Oncology, Odense University Hospital, J. B. Winsl&oslash;ws Vej 4, 5000 Odense C, Denmark&nbsp;</p> <p>2 Department of Clinical Research, University of Southern Denmark, J.B. Winsl&oslash;ws Vej 19 3., 5000 Odense C, Denmark</p> <p>3 Department of Oncology, Odense University Hospital, J. B. Winsl&oslash;ws Vej 4, 5000 Odense C, Denmark</p> <p>* Correspondence:&nbsp;</p> <p>Ebbe Laugaard Lorenzen</p> <p>ebbe.lorenzen@rsyd.dk</p> <p>Carsten Brink&nbsp;</p> <p>carsten.brink@rsyd.dk</p> <p>&nbsp;</p> <p>NOTE: Version 2 of this repository contains the nnU-Net v2 model, while version 1 contains the original, nnU-Net v1 model</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Image distortion data from "An open-source MRI compatible frame for multimodal presurgical mapping in macaque and capuchin monkeys"

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo32/100

X-COBOL: A Dataset of Open-Source COBOL repositories

<p>A dataset of 5195 open source COBOL files across 168 repositories</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

SCoV2-VAR: A light-weighted, customizable, and open-source database of 12 million SARS-CoV-2 genomes

<pre>Explosive accumulation of SARS-CoV-2 variants is posing a challenge to monitoring virus mutation and other data dealing, particularly based on centralized databases. The present study aimed to establish a light-weighted, customizable, and open-source database for SARS-CoV-2 genomes and annotations, without any access limit. The database, named SCoV2-VAR, was constructed, based on the variations (VAR) of the full-length SARS-CoV-2 (SCoV2) data uploaded on websites. All sequence samples were subject to quality control, single nucleotide polymorphism (SNP) annotation, format conversion, and final compression before appending to SCoV2-VAR. The final version of SCoV2-VAR (up to Feb 2024) contained more than 12 million SARS-CoV-2 records, with full genome and annotations. SCoV2-VAR was extremely light-weighted, with a storage size of 937 Mb for all 12 million sequences, post a 1: 596 compression. SCoV2-VAR is capable of timely updating, quickly querying, and customizable outputting SARS-CoV-2 sequences and their annotations. Additionally, the present study provided an overview of all 12 million SARS-CoV-2 samples, for both sequences and annotations.<br> <br><br></pre>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Understanding the Adoption of Modern JavaScript Features: An Empirical Study on Open-Source Systems

<p>This repository contains the data and analysis from an empirical study investigating the adoption trends of modern JavaScript features introduced with ECMAScript 6 (ES6) and beyond. By mining the source code history of 158 open-source JavaScript projects, the study identifies efforts to rejuvenate legacy code by replacing outdated constructs with modern ones. The findings highlight the extensive use of modern features, their widespread adoption within one to two years after ES6's release, and ongoing trends in the rejuvenation of JavaScript codebases.<br><br></p> <ul> <li> <p><strong>scripts.zip</strong>: Contains Python scripts used to analyze data and generate the graphs presented in the study's results.</p> </li> <li><strong>scripts-threats-analysis.zip</strong>:&nbsp; Contains the Python scripts used to analyze the projects without applying the study's filtering criteria and to generate the table presented in the Threats to Validity section.</li> <li> <p><strong>jsminer-tool.zip</strong>: Includes the tool developed to analyze GitHub repository history and collect metrics on the adoption of modern JavaScript features.</p> </li> <li> <p><strong>jsminer_database_backup.zip</strong>: Provides a PostgreSQL database dump containing all code review comments from the repositories analyzed in the study.</p> </li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo32/100

WildfireDB: An Open-Source Dataset Connecting Wildfire Spread with Relevant Determinants

<p>Modeling fire spread is critical in fire risk management. Creating data-driven models to forecast spread remains challenging due to the lack of comprehensive data sources that relate fires with relevant covariates. We present the first comprehensive and open-source dataset that relates historical fire data with relevant covariates such as weather, vegetation, and topography. Our dataset, named <em>WildfireDB</em>, contains over 17 million data points that capture how fires spread in the continental USA in the last decade. The paper accompanying this dataset is part of the 2021 Neural Information Processing Systems (NeurIPS) Dataset and Benchmark Track. The paper&nbsp;describes the algorithmic approach used to create and integrate the data, describe the dataset, and present benchmark results regarding data-driven models that can be learned to forecast the spread of wildfires.</p> <p>&nbsp;</p> <p>Please see&nbsp;https://colab.research.google.com/drive/1cm2Z4E0HzXMAcuUrE26wHXL2FS_pIj3t?usp=sharing for an introduction about how to load the database using python (pandas).</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Open-Source Low-Cost Cardiac Optical Mapping System

<p>Optical fluorescence imaging data for&nbsp;<strong>Open-Source Low-Cost Cardiac Optical Mapping System&nbsp;</strong>PLOS ONE paper submission</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Dataset for Mistic: an open-source multiplexed image t-SNE viewer

<p>This link consists of 10 anonymized non-small cell lung cancer (NSCLC)&nbsp;field&nbsp;of Views (FoVs) to test Mistic.</p> <p><strong>Mistic</strong></p> <p>Understanding the complex ecology of a tumor tissue and the spatio-temporal relationships between its cellular and microenvironment components is becoming a key component of translational research, especially in immune-oncology. The generation and analysis of multiplexed images from patient samples is of paramount importance to facilitate this understanding. In this work, we present Mistic, an open-source multiplexed image t-SNE viewer that enables the simultaneous viewing of multiple 2D images rendered using multiple layout options to provide an overall visual preview of the entire dataset. In particular, the positions of the images can be taken from t-SNE or UMAP coordinates. This grouped view of all the images further aids an exploratory understanding of the specific expression pattern of a given biomarker or collection of biomarkers across all images, helps to identify images expressing a particular phenotype or to select images for subsequent downstream analysis. Currently there is no freely available tool to generate such image t-SNEs.</p> <p><strong>Links</strong></p> <p><br> <a href="https://github.com/MathOnco/Mistic">Mistic code</a></p> <p><a href="https://mistic-rtd.readthedocs.io/">Mistic documentation</a></p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.10.08.463728v1">Paper</a></p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Replication package for "Blended Modeling in Commercial and Open-source Model-Driven Software Engineering Tools: A Systematic Study"

<p>Replication package for the paper&nbsp;<em>Blended Modeling in Commercial and Open-source Model-Driven Software Engineering Tools: A Systematic Study</em>.</p> <p>Protocol</p> <ul> <li><code>/01-protocol/protocol.pdf</code></li> </ul> <p>Data &amp; analysis scripts</p> <p>This replication package is structured as follows:</p> <ul> <li><code>/02-search</code>&nbsp;- Detailed data on the&nbsp;<code>/academic</code>&nbsp;and&nbsp;<code>/grey literature</code>&nbsp;search.</li> <li><code>/03-tools</code>&nbsp;- Identified tools and inclusion/exclusion decisions.</li> <li><code>/04-classification_schema</code>&nbsp;- Classification framework and the corresponding data extraction form.</li> <li><code>/05-data</code>&nbsp;- Clean data in a processable form.</li> <li><code>/06-analysis</code>&nbsp;- Analysis scripts and results.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Data Supplement: Open-source model-based reconstruction in Julia (ISMRM 2022)

<p>Dataset to accompany the abstract presented at ISMRM 2022:</p> <p>Open-source model-based reconstruction in Julia: A pipeline for spiral diffusion imaging</p> <p>Instructions after download are found at the github link for the project.</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

Open-source release of tensor-network software

<p>This is a package<sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-SBmodel-85d3b1b42ae4f41239f8975ec68008b2">1</a></sup>&nbsp;for calculating FLUCTUATIONS of heat transfer in the Spin-Boson model<sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-PRX2020-85d3b1b42ae4f41239f8975ec68008b2">2</a></sup>&nbsp;using the&nbsp;<strong>Time Evolving Density matrices using Orthogonal Polynomial Algorithm (<em>TEDOPA</em>)</strong><sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-Prior2010-85d3b1b42ae4f41239f8975ec68008b2">3</a></sup><sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-Chin2010-85d3b1b42ae4f41239f8975ec68008b2">4</a></sup>.</p> <p>We employ the&nbsp;<strong>Thermofield-based chain-mapping approach for open quantum systems</strong><sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-PRA2015-85d3b1b42ae4f41239f8975ec68008b2">5</a></sup>&nbsp;that enables us to use a vacuum initial matrix product state (pure) for the environment instead of a thermal state (mixed), thereby speeding up the computation greatly.</p> <p>In this package, we use the ITensor library<sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-Itensor-85d3b1b42ae4f41239f8975ec68008b2">6</a></sup>&nbsp;in Julia for tensor network manipulations. &nbsp;</p> <p>This package uses&nbsp;<strong>julia = "1.8.2"</strong> version.</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Covert Communication Channels based on Hardware Trojans: open-source dataset

<h1>Dataset of Hardware-Trojan (HT) based Covert Channels (HT-CCs) for the IEEE 802.11 (WiFi) standard</h1> <div> <h2>Datasets arhitecture</h2> <ul> <li>The dataset has 80 dataset elements organized in 5 classes, namely CC-free (HT0-CC) and CC-infected (HTX-CC, X={1, &middot; &middot; &middot; , 4}), each class has 16 elements.</li> <li>Those 16 elements are 8 acquisitions for different SNR values ranging from 1dB to 29dB with a step of 4dB, 2 acquisitions for each SNR value. The second acquisition was obtained immediately after the first one, with no other parameters changing between one acquisition and the other. The second acquisition is marked with an underscore in the naming of the file.</li> <li>Each dataset element has 2000 frames and it is the concatenation of ten 200-frames sub-acquisitions.</li> </ul> </div> <div> <h3>Naming convention</h3> </div> <p>Files are named according to the following convention:</p> <p>First acquisition:</p> <div> <pre><code>rxSig_&lt;SNR value&gt;dB_HT&lt;HT-CC attack&gt;.mat </code></pre> <div>&nbsp;</div> </div> <p>Second acquisition:</p> <div> <pre><code>rxSig_&lt;SNR value&gt;dB_HT&lt;HT-CC attack&gt;_.mat </code></pre> <div>&nbsp;</div> </div> <ul> <li>SNR value: ranging from 1dB to 29dB with a step of 4dB, i.e., &lt;1&gt;=SNR of 1db, &lt;2&gt;=SNR of 5db, &lt;3&gt;=SNR of 9dB, &lt;4&gt;=SNR of 13dB, &lt;5&gt;=SNR of 17dB, &lt;6&gt;=SNR of 21dB, &lt;7&gt;=SNR of 25dB, &lt;8&gt;=SNR of 29dB.</li> <li>HT-CC attack: CC-free (HT0-CC) and CC-infected (HTX-CC, X={1, &middot; &middot; &middot; , 4}).</li> </ul> <div> <h3>HT0-CC</h3> </div> <p>The HT0-CC is the benchmark CC-free transmission according to the WiFi standard.</p> <div> <h3>HT1-CC</h3> </div> <p>The HT1-CC is the Amplitude Modulation (AM) Short Training Sequence (STS) HT attak based on the paper</p> <blockquote> <p>A. R. D&iacute;az-Rizo, H. Aboushady and H.-G. Stratigopoulos, "Leaking Wireless ICs via Hardware Trojan-Infected Synchronization," in IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 5, pp. 3845-3859, 1 Sept.-Oct. 2023, doi: 10.1109/TDSC.2022.3218507.</p> </blockquote> <p>The amplitude modulation value alpha is set to 10%.</p> <div> <h3>HT2-CC</h3> </div> <p>The HT2-CC is the PSK STF HT attack based on the paper</p> <blockquote> <p>J. Classen, M. Schulz and M. Hollick, "Practical covert channels for WiFi systems," 2015 IEEE Conference on Communications and Network Security (CNS), Florence, Italy, 2015, pp. 209-217, doi: 10.1109/CNS.2015.7346830.</p> </blockquote> <p>The implemented version is leaking 8 bits per OFDM PPDU.</p> <div> <h3>HT3-CC</h3> </div> <p>The HT3-CC is the Dirty Constellation attack based on paper</p> <blockquote> <p>A. Dutta, D. Saha, D. Grunwald and D. Sicker, "Secret Agent Radio: Covert Communication through Dirty Constellations," in Proceedings of the 14th international conference on Information Hiding (IH'12). Springer-Verlag, Berlin, Heidelberg, 160&ndash;175, doi: 10.1007/978-3-642-36373-3_11.</p> </blockquote> <p>The implemented version of this attack has an embedding frequency of 5 subcarriers per OFDM symbol, where each OFDM symbol comprises 48 data subcarriers. Each dirty subcarrier is leaking 2 bits; therefore, we are leaking 10 bits per OFDM symbol,</p> <div> <h3>HT4-CC</h3> </div> <p>The HT4-CC is the AM analog/RF attack based on paper</p> <blockquote> <p>K. S. Subramani, N. Helal, A. Antonopoulos, A. Nosratinia and Y. Makris, "Amplitude-Modulating Analog/RF Hardware Trojans in Wireless Networks: Risks and Remedies," in IEEE Transactions on Information Forensics and Security, vol. 15, pp. 3497-3510, 2020, doi: 10.1109/TIFS.2020.2990792.</p> </blockquote> <p>Due to SPI limitations, the implemented version of the attack is a digitally-emulated version of the AM analog/RF attack. It leaks 8 bits per transmitted frame.</p> <div> <h2>Hardware Platform</h2> </div> <p>Acquisitions are received downconverted IQ samples using a Software Defined Radio (SDR) board bladeRF xA9.</p>

opencc-by-nc-4.0Mar 2024View details →
zenodo32/100

Code Smells and their Collocations : A Large-scale Experiment on Open-source Systems

<p>This dataset includes classes with code smells, acquired from Qualitas Corpus (QC).<br> Folder &#39;all&#39; contains data coming from the QC rev.20130901 (92 systems).<br> Folder &#39;domains&#39; contains data coming from QC rev.20111026 (76 systems updated to their most recent releases from rev.20130901).&nbsp;<br> Folder &#39;pca&#39; includes results of the PCA analysis, generated with the R prcomp() function for regular PCA, and logisticPCA() function for the binary data.</p> <p>Filenames include information about the base release of the QC, and a number (25, 50 or 75) that specifies the minimum number of detectors that identified a specific smell instance (25%, 50%, and 75%, respectively). For example, if a given code smell in a class X has been identified by 1 out of 4 available detecting tools, then the smell for the class X will be reported in the respective file 25, but not in 50 or 75. Please note, that for smells detected with only one tool, the values would be equal in all datasets (in that case, the smell was detected by 0% or 100% of tools)</p> <p>In all files, &quot;1&quot; denotes that the smell was identified (subject to the limitations with the number of detectors, described above), and &ldquo;0&rdquo; that the smell was not found in a given class.</p> <p>The filename also includes the domain abbreviation (app, css, dev, dgdv) or a keyword ALL, which indicates that the dataset includes data from all domains.</p> <p>The smells have been detected by 11 tools. Most of the tools detect more than one smell.&nbsp;<br> Information about the tool used to detect a given smell is given in headers of each file. Additionally, in &#39;smell detectors.csv&#39; file we present the information about smells detected by a specific tool.</p>

opencc-by-nc-4.0May 2018View details →
zenodo32/100

All the open-source SRT datasets used in STMGAC

<p>All the open-source SRT datasets used in STMGAC</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Reproduction package for the paper "The open-source sunbather code: Modeling escaping planetary atmospheres and their transit spectra"

<p>This is a reproduction package for the paper "The open-source sunbather code: modeling escaping planetary atmospheres and their transit spectra" by Dion Linssen, Jim Shih, Morgan MacLeod &amp; &nbsp;Antonija Oklopčić (2024). It provides a front-to-end reproduction script to reproduce the results and Figures 1-5 of the paper. Figures 6&amp;7 can be reproduced with the example notebook found in the sunbather installation.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record