Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
311
datasets available to search
ShareScore release 0.9.0
Dataset results
311 results for “Open source”
ThermoCyte: an inexpensive open-source temperature control system for in vitro live cell imaging
Open the record for dataset details and reuse information.
Cardio PyMEA: A user-friendly, open-source Python application for cardiomyocyte microelectrode array analysis
Open the record for dataset details and reuse information.
SOils DAta Harmonization database (SoDaH): an open-source synthesis of soil data from research networks
This SOils DAta Harmonization (SoDaH) database is designed to bring together soil carbon data from diverse research networks into a harmonized dataset that can be used for synthesis activities and model development. The research network sources for SoDaH span different biomes and climates, encompass multiple ecosystem types, and have collected data across a range of spatial, temporal, and depth gradients. The rich data sets assembled in SoDaH consist of observations from monitoring efforts and long-term ecological experiments. The SoDaH database also incorporates related environmental covariate data pertaining to climate, vegetation, soil chemistry, and soil physical properties. The data are harmonized and aggregated using open-source code that enables a scripted, repeatable approach for soil data synthesis.
Key input and output data for the multi-model analysis "Open Source Energiewende"
<p>This repository contains key input and output data of the multi-model analysis carried out in the project "Open Source Energiewende", financed by the German Federal Ministry for Economic Affairs and Energy.</p> <p>The results are presented and discussed in the paper "Power sector effects of cheaper stationary batteries: insights from an open multi-model analysis".</p> <p>The model codes are availabe in individual repositories, which are provided in the paper.</p>
Refactoring Test Smells: A Perspective from Open-Source Developers
<p>Presentation video for the <strong>5th Brazilian Symposium on Systematic and Automated Software Testing (SAST)</strong>, during the <strong>11th Brazilian Conference on Software: Practice and Theory (CBSoft 2020)</strong></p>
Quantitative analysis of subcellular distributions with an open-source, object-based tool
<p>The subcellular localization of objects, such as organelles, proteins, or other molecules, instructs cellular form and function. Understanding the underlying spatial relationships between objects through colocalization analysis of microscopy images is a fundamental approach used to inform biological mechanisms. We generated an automated and customizable computational tool, the SubcellularDistribution pipeline, to facilitate object-based image analysis from 3D fluorescence microcopy images. To test the utility of the SubcellularDistribution pipeline, we examined the subcellular distribution of mRNA relative to centrosomes within <i>Drosophila</i> embryos. Centrosomes are microtubule-organizing centers, and RNA enrichments at centrosomes are of emerging importance. Our open-source and freely available software detected RNA distributions comparably to commercially available image analysis software. The SubcellularDistribution pipeline is designed to guide the user through the complete process of preparing image analysis data for publication, from image segmentation and data processing to visualization.</p>
Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Article Pool)
<p>This pdf includes all of the articles that analyzed in the study: "Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review".</p>
Open-source software collaboration network mining dataset
<p>The resulting dataset of the <a href="https://github.com/gotec/git2net">git2net </a>and <a href="https://github.com/wschuell/repo_tools">repo_tools </a>mining process for randomly selected large open-source repositories.</p>
South African Open Data in Higher Education: Sources, resources and providers
<p>Spreadsheet of data sourced on South African sources, resources and providers of higher education open data. Composed through desk review as principle component of the situational analysis conducted for the 'Use of open data in the governance of South African higher education' research project, in the IDRC/WWWF 'Exploring Emerging Impacts of Open Data in the South' initiative.</p>
Mandelbugs in Open-Source Software
<p>This dataset contains a list bugs from four open-source projects (the Linux kernel, the MySQL DBMS, the Apache HTTPD web server, and the Apache AXIS WS framework). The bugs have been classified into Mandelbugs, Bohrbugs, or Aging-Related Bugs, by analyzing the conditions that exercise the bug (i.e., the "fault trigger"). This classification is useful to get insights into bugs and failures that can occur in OSS projects, and to tune testing and fault-tolerance strategies according to the distribution of bug types in a project.</p> <p>The dataset contains an ARFF file for each subsystem of the four open-source projects. Each row of the ARFF file contains:</p> <p>- An IDs of the bug, which can be used to retrieve more information about the bug from the issue tracker of the project;</p> <p>- A string that represents the class of the bug (BOH = Bohrbug; NAM = Mandelbug; ARB = Aging-Related Bug; UNK = Unknown class);</p> <p>- A string that represents the sub-class of the bug (for Bohrbugs, the sub-class is not available; the subclasses for Mandelbugs are LAG, ENV, TIM, SEQ; the subclasses for Aging-Related bugs are MEM, STO, LOG, NUM, TOT).<br> </p>
Komplementäres Datenmodell für das Open-Source-Tool OptIES
<p>Der hier verfügbare Datensatz stellt das komplementäre Datenmodell zum Open-Source-Tool <a href="https://github.com/znes/OptIES" target="_blank" rel="noopener">OptIES</a> dar. </p> <p>Das Tool <code>OptIES</code> dient der Optimierung regionaler Energiesysteme und wurde im Rahmen des Forschungsprojekts "OptIES Dörpum - Offene Optimierung sektorgekoppelter regionaler Energiesysteme am Beispiel des IES Dörpum" entwickelt. Es basiert auf der offenen Software-Toolbox <a href="https://github.com/PyPSA/PyPSA" target="_blank" rel="noopener">PyPSA</a>.</p> <p>Das Forschungsprojekt <a href="https://www.uni-flensburg.de/eum/forschung/laufende-projekte/opties-doerpum" target="_blank" rel="noopener">OptIES Dörpum</a> wurde durch die <a href="https://www.uni-flensburg.de/" target="_blank" rel="noopener">Europa-Universität Flensburg</a> und die <a href="https://www.ecowert360.com" target="_blank" rel="noopener">EcoWert360° GmbH</a> bearbeitet und durch das Förderprogramm HWT Energie und Klimaschutz der <a href="https://www.eksh.org/" target="_blank" rel="noopener">Gesellschaft für Energie und Klimaschutz Schleswig-Holstein (EKSH)</a> finanziert. Es stellte die wissenschaftliche Begleitung des Praxisprojekts <a href="https://www.aktivregion-nf-nord.de/fileadmin/user_upload/KT_Klimawandel_Energie/Projekte/IES_D%C3%B6rpum/07.51_-_Beschreibung_-_Projekt_57_IES_D%C3%B6rpum.pdf" target="_blank" rel="noopener">IES Dörpum</a> dar.</p> <p>Ziel des Forschungsprojekts war es, lokale sowie nationale Herausforderungen der Energiewende durch Einordnung der kommunalen Bestrebungen in das nationale Energiesystem zu beleuchten und analysieren. Das vorliegende Tool dient der Optimierung des regionalen Energiesystems zur Untersuchung und Bewertung von Betriebsstrategien und Entwicklungspfaden des isolierten Systems. Im weiteren Projektverlauf wurde das regionale System unter Berücksichtigung des übergelagerten Gesamtsystems optimiert, um integrierte Handlungsempfehlungen für die erfolgreiche Energiewende auf kommunaler und nationaler Ebene abzuleiten.</p> <p>Unter den hochgeladenen Dateien findet sich der Projektabschlussbericht mit einer ausführlichen Darstellung des Projekts, der entwickelten Modelle und Tools, der durchgeführten Berechnungen und Analysen sowie der Ergebnisse. Zudem werden die im Rahmen des Projekts angefertigten Masterarbeiten von Matthias Winschu und Mohsen Mansouri bereitgestellt.</p> <p>Die hier veröffentlichten Daten enthalten Informationen zu Verbrauchs- und Erzeugungseinheiten innerhalb des IES Dörpum, welche für Modellrechnungen mit der Software-Toolbox notwendig sind. Die Daten enthalten sowohl synthetische als auch reale Daten in anonymisierter Form. Es werden die innerhalb von <code>PyPSA</code> als Standard definierten Einheiten verwendet (also z. B. MW). Für die Erstellung der synthetische Zeitreihen ist das Jahr 2019 als Referenzjahr herangezogen. </p> <p>Zur Erstellung der synthetischen Lastzeitreihen (Strom und Wärme) wurden Jahresgesamtverbräuche der letzten Jahre herangezogen und basierend auf der Arbeit von <a href="https://doi.org/10.1186/s42162-022-00201-y">Büttner et al.</a> zeitlich verteilt. Zur Erstellung der synthetischen Lastzeitreihen aus Ladevorgängen von Elektrofahrzeugen wurden Daten der Studie <a href="https://www.mobilitaet-in-deutschland.de/">Mobilität in Deutschland</a> innerhalb des Tools <a href="https://github.com/RAMP-project/RAMP-mobility">RAMP-mobility</a> verwendet. Die synthetische Zeitreihe der potentiellen Erzeugung durch PV ist mithilfe <a href="https://www.renewables.ninja/">renewables_ninja</a> erzeugt. Alle weiteren Daten zur Parametrisierung der Komponenten sowie die synthetische Gaslast und Eigenverbräuche der enthaltenen Biogasanlage sind aufgrund von Angaben der Biogas Dörpum GmbH festgelegt und mithilfe der Daten der <a href="https://github.com/PyPSA/technology-data">PyPSA technology data</a> ergänzt (insbesondere Kostenannahmen).</p> <p>Die gemessenen Zeitreihen ergeben sich aus dem Zeitraum Mitte Februar 2023 bis Mitte Febraur 2024. Aufgrund von Hard- und Softwarefehlern stehen keine lückenlosen Messungen zur Verfügung, insbesondere die Messungen der Ladevorgänge durch E-Fahrzeuge sind beeinträchtigt. Eine Übersicht über die Verfügbarkeit der gemessenen Daten ist der Abbildung <strong>Verfuegbarkeit_Messungen.png</strong> zu entnehmen.</p> <p><strong>buses.csv - </strong>enthält Informationen, wie Name, Energieträger und wenn anwendbar, die Spannungsebene, der verschiedenen im Modell abgebildeten Knoten. </p> <p><strong>generators.csv - </strong>stellt alle Erzeugungseinheiten mit ihren zugehörigen Knoten, installierten Kapazitäten und Kostenparametern dar. </p> <p><strong>lines.csv - </strong>beschreibt Verlauf, Länge und Art der verbauten Kabel und die zugehörigen elektrischen Parameter. </p> <p><strong>links.csv - </strong>liefert Informationen zu allen kontrollierbaren Komponenten, die zwei Knoten zweier beliebiger Energieträger verbinden. Im vorliegenden Beispiel können das KWK-Anlagen sowie Be- und Entladeeinheiten für Wärmespeicher sein. </p> <p><strong>loads.csv - </strong>beschreibt alle Verbrauchseinheiten mit deren Energieträger und zugehörigem Knoten<strong>. </strong></p> <p><strong>storage_unit.csv - </strong>stellt die vorhandenen Batteriespeicher mit Kapazitäten, Kosten- und Effizienzparametern dar. </p> <p><strong>stores.csv - </strong>enthält Informationen zu anderen Speichereinheiten für Wärme oder Biogas mit Kapazitäten, Kosten- und Effizienzparametern.</p> <p><strong>el_load_synth.csv - </strong>stellt synthetisch erzeugte stündliche elektrischen Lastzeitreihen für alle im Model abgebildeten Anschlussnehmer bereit. Unter 'BGA_AC' ist der elektrische Eigenverbrauch der Biogasanlage angenommen. 'LS1' und 'LS2' beinhalten synthetische elektrische Lastzeitreihen zum Verbrauch durch Ladevorgänge von Elektrofahrzeugen unter der Annahme von zehn ('LS1') und fünf ('LS2') Nutzer*innen.</p> <p><strong>ev_load_synth.csv - </strong>stellt zusätzliche, synthetische elektrische Lastzeitreihen zum Verbrauch durch Ladevorgänge von Elektrofahrzeugen bereit. Die verschiedenen Zeitreihen variieren in der Annahme der Anzahl der Nutzer*innen, diese ist in dem Spaltentitel vermerkt.</p> <p><strong>heat_load_synth.csv -</strong> hält synthetische Daten zum Wärmeverbrauch des angeschlossenen Nahwärmenetzes und zum Eigenverbrauch der Biogasanlage in stündlicher Auflösung bereit. Weiterhin ist der Wärmeeigenbedarf der Biogasanlage angenommen.</p> <p><strong>gas_load_synth.csv -</strong> liefert stündliche synthetische Daten zum Gasverbrauch eines Satelliten-BHKWs. </p> <p><strong>pot_pv_timeseries_synth.csv - </strong>enthält Daten zu potenziellen Stromeinspeisezeitreihen aus Photovoltaik in stündlicher Auflösung. </p> <p><strong>el_load_real.csv - </strong>stellt real gemessene elektrische Lastzeitreihen für alle im Model abgebildeten Anschlussnehmer in einer zeitlichen Auflösung von fünf Minuten bereit. Bei Messlücken größer als fünf Minuten wurde interpoliert. Unter 'BGA_AC' ist der elektrische Eigenverbrauch der Biogasanlage angenommen. 'LS1' und 'LS2' beinhaltet die selbe gemessene elektrische Lastzeitreihe zum Verbrauch durch Ladevorgänge von Elektrofahrzeugen. Diese wurde bei Messlücken >1h interpoliert. </p>
Open-Source Cardiac MR Fingerprinting
<p>Magnetic Resonance (MR) raw data acquired with an open-source cardiac MR Fingerprinting (cMRF) sequence of a phantom at four different MR scanners. More details can be found here: https://github.com/PTB-MR/cMRF. The colormaps are taken from https://zenodo.org/records/11185704 because zenodo_get failed on trying to download this record in a jupyter notebook.</p> <p>Additionally cMRF data was acquired in three volunteers who were scanned at two different scanners. Cartesian and golden radial cine data was acquired to verify the anatomical features seen in the quantitative maps.</p>
Dataset: A continuous open source data collection platform for architectural technical debt assessment
<p>The dataset and replication package of the study "A continuous open source data collection platform for architectural technical debt assessment".</p> <p> </p> <p>Abstract</p> <p>Architectural decisions are the most important source of technical debt. In recent years, researchers spent an increasing amount of effort investigating this specific category of technical debt, with quantitative methods, and in particular static analysis, being the most common approach to investigate such a topic.</p> <p> </p> <p>However, quantitative studies are susceptible, to varying degrees, to external validity threats, which hinder the generalisation of their findings.</p> <p>In response to this concern, researchers strive to expand the scope of their study by incorporating a larger number of projects into their analyses. This practice is typically executed on a case-by-case basis, necessitating substantial data collection efforts that have to be repeated for each new study.</p> <p> </p> <p>To address this issue, this paper presents our initial attempt at tackling this problem and enabling researchers to study architectural smells at large scale, a well-known indicator of architectural technical debt. Specifically, we introduce a novel approach to data collection pipeline that leverages Apache Airflow to continuously generate up-to-date, large-scale datasets using Arcan, a tool for architectural smells detection (or any other tool).</p> <p>Finally, we present the publicly-available dataset resulting from the first three months of execution of the pipeline, that includes over 30,000 analysed commits and releases from over 10,000 open source GitHub projects written in 5 different programming languages and amounting to over a billion of lines of code analysed.</p>
Sample Stripped Pre-supernova Progenitors for open-source code CHIPS (Complete History for Interaction-Powered Supernovae)
<p>Inlists, mainly based on the test suite "example_make_pre_ccsn" in r12778, with slight amendments for removal of hydrogen (and helium, for Ic progenitors) envelope at core hydrogen (helium) exhausion.</p><p>For details: https://ui.adsabs.harvard.edu/abs/2023arXiv230810785T/abstract</p>
Artifact for "Inside Bug Report Templates: An Empirical Study on Bug Report Templates in Open-Source Software"
<p>This is the artifact for the paper "Inside Bug Report Templates: An Empirical Study on Bug Report Templates in Open-Source Software".</p> <p><strong>What the artifact does:</strong><br>1) a questionnaire that we used for our online survey (PDF);<br>2) the valid responses of our online survey (CSV).</p> <p>3) the code of preprocessing (.py).</p> <p>4) the dataset of preprocessing and labeling (CSV).</p>
Estimating Usage Of Open Source Projects - Flutter Telemetry Case Study - MSR' 24
<p>This dataset (CSV) was assembled to support analysis within a case study that will be published in the proceedings of the <a href="https://conf.researchr.org/home/msr-2024">Mining Software Repositories</a> conference (MSR ‘24) April 14-15 2024: <a title="Estimating Usage Of Open Source Projects" href="https://doi.org/10.1145/3643991.3645066">Estimating Usage Of Open Source Projects.</a></p> <p>This case study explored whether publicly available metrics could serve as proxies to estimate usage of an open source project. Using the <a href="https://flutter.dev/">Flutter</a> project as our case study, we collected monthly proxy metrics from GitHub, StackOverflow and Slack to compare with Flutter’s monthly active user count over the same time period: January 2018 through February 2021.</p> <p>All metrics correspond to the last day of the month and/or represent aggregate activity in that month. Data from GitHub shows aggregate activity counts across the entire <a href="https://github.com/flutter">Flutter GitHub organization</a> (up to 33 repositories). </p> <p>Our specific metrics include:</p> <div> <table> <tbody> <tr> <td> <p>Source</p> </td> <td> <p>Metric</p> </td> <td> <p>Aggregation method</p> </td> <td> <p>Details</p> </td> </tr> <tr> <td> <p>Flutter</p> </td> <td> <p>Monthly Active Users (MAU)</p> </td> <td> <p>Google internal tooling</p> </td> <td> <p>Flutter users active in the last 30 days, collected on the last day of each month</p> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>PullRequest Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> <p>As Google governs changes to this code base, we excluded known Google employees in code change related metrics </p> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Issues Created in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Issue Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Fork events in month</p> </td> <td> <p><a href="http://gharchive.org">gharchive.org</a></p> </td> <td> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Fork cumulative count at the end of month</p> </td> <td> <p><a href="http://gharchive.org">gharchive.org</a></p> </td> <td> </td> </tr> <tr> <td> <p>StackOverflow</p> </td> <td> <p>Question Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> </td> </tr> <tr> <td> <p>StackOverflow</p> </td> <td> <p>Questions in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> </td> </tr> <tr> <td> <p>Slack</p> </td> <td> <p>Claimed (cumulative) members at the end of the month</p> </td> <td> <p>fluttercommunity.slack</p> </td> <td> </td> </tr> <tr> <td> <p>Slack</p> </td> <td> <p>Cumulative messages at the end of the month</p> </td> <td> <p>fluttercommunity.slack</p> </td> <td> </td> </tr> </tbody> </table> </div> <p>Tools: </p> <ul> <li> <p>We used an instance of <a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a> hosted by <a href="https://bitergia.com/">Bitergia</a> to aggregate GitHub PullRequest Authors, GitHub Issue Authors, GitHub Issues, across all repositories under the Flutter Organization, and StackOverflow Question Authors, and StackOverflow Questions for questions that mention Flutter. We used the <a href="http://github.com/chaoss/grimoirelab-sortinghat">Sorting Hat</a> of feature GrimoireLab to identify Google employees in this sample.</p> </li> </ul> <ul> <li> <p><a href="http://gharchive.org">GHArchive</a> via <a href="https://cloud.google.com/blog/topics/public-datasets/github-on-bigquery-analyze-all-the-open-source-code">BigQuery</a> was used to count GitHub Star and Fork events across all repositories under the Flutter Organization</p> </li> <li> <p>We pulled Slack activity directly from <a href="http://fluttercommunity.slack.com">Flutter's slack channel </a>dashboard</p> </li> </ul>
Chinese Reference Population: open-source age-dependent computational phantoms of reference Chinese population
<p>The<strong> Chinese Reference Population (CRP)</strong> phantoms dataset encompass <strong>30 phantoms</strong> available in both voxel and NURBS formats, with age in 0, 1, 2, 3, 4, 5, 6, 8, 10, 12, 15, 18 years and adult male and female, as well as 4 pregnant women and fetus in early pregnancy, first trimester, second trimester and third trimester.</p> <ul> <li><strong>Voxelized phantoms</strong> are accessible in NII format :<strong> <em>"XXX.nii", which could be opened in AMIDE software.</em></strong></li> <li>Excel file<strong> </strong>containing<strong> organ masses and other descriptive information </strong>:<strong> <em>"CRP_descriptive_Info.xlsx"</em></strong></li> <li>In the application of F18−FDG dose calculation, <strong>organ absorbed doses per unit activity administered </strong>is provided in an Excel file :<strong> <em>"Application_F18-FDG.xlsx"</em></strong></li> </ul> <p>All data are stored on Zenodo and can be publicly accessed.</p>
GloHydroRes - a global dataset combining open-source hydropower plant and reservoir data
<div> <div> <div> <div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>Analyzing the impacts of drought and climate change on hydropower requires detailed data not only on hydropower attributes such as plant type, head, and installed capacity, but also on reservoir characteristics like area, depth, and volume. Current open-source hydropower datasets typically lack information on reservoirs, while reservoir datasets often omit hydropower details. GloHydroRes is a global dataset that integrates open-source hydropower and reservoir data, offering 29 attributes, including key information such as installed capacity, plant type, dam height, reservoir depth, area, volume, and river name. Overall, GloHydroRes provides data on 7,775 hydropower plants across 128 countries.</p> </div> </div> </div> </div> <div> <div> <div> </div> </div> </div> </div> </div>
E2EGit: A Dataset of End-to-End Web Tests in Open Source Projects
<p><strong>ABSTRACT </strong><br>End-to-end (E2E) testing is a software validation approach that simulates realistic user scenarios throughout the entire workflow of an application. In the context of web<br>applications, E2E testing involves two activities: Graphic User Interface (GUI) testing, which simulates user interactions with the web app’s GUI through web browsers, and performance testing, which evaluates system workload handling. Despite its recognized importance in delivering high-quality web applications, the availability of large-scale datasets featuring real-world E2E web tests remains limited, hindering research in the field.<br>To address this gap, we present E2EGit, a comprehensive dataset of non-trivial open-source web projects collected on GitHub that adopt E2E testing. By analyzing over 5,000 web repositories across popular programming languages (JAVA, JAVASCRIPT, TYPESCRIPT, and PYTHON), we identified 472 repositories implementing 43,670 automated Web GUI tests with popular browser automation frameworks (SELENIUM, PLAYWRIGHT, CYPRESS, PUPPETEER), and 84 repositories that featured 271 automated performance tests implemented leveraging the most popular open-source tools (JMETER, LOCUST). Among these, 13 repositories implemented both types of testing for a total of 786 Web GUI tests and 61 performance tests.</p> <p><br><strong>DATASET DESCRIPTION </strong><br>The dataset is provided as an SQLite database, whose structure is illustrated in Figure 3 (in the paper), which consists of five tables, each serving a specific purpose.<br>The <em>repository </em>table contains information on 1.5 million repositories collected using the SEART tool on May 4. It includes 34 fields detailing repository characteristics. The<br><em>non_trivial_repository</em> table is a subset of the previous one, listing repositories that passed the two filtering stages described in the pipeline. For each repository, it specifies whether it is a web repository using JAVA, JAVASCRIPT, TYPESCRIPT, or PYTHON frameworks. A repository may use multiple frameworks, with corresponding fields (e.g., is web java) set to true, and the field web dependencies listing the detected web frameworks. For Web GUI testing, the dataset includes two additional tables; <em>gui_testing_test _details</em>, where each row represents a test file, providing the file path, the browser automation framework used, the test engine employed, and the number of tests implemented in the file. <em>gui_testing_repo_details</em>, aggregating data from the previous table at the repository level. Each of the 472 repositories has a row summarizing<br>the number of test files using frameworks like SELENIUM or PLAYWRIGHT, test engines like JUNIT, and the total number of tests identified. For performance testing, the <em>performance_testing_test_details</em> table contains 410 rows, one for each test identified. Each row includes the file path, whether the test uses JMETER or LOCUST, and extracted details such as the number of thread groups, concurrent users, and requests. Notably, some fields may be absent—for instance, if external files (e.g., CSVs defining workloads) were unavailable, or in the case of Locust tests, where parameters like duration and concurrent users are specified via the command line.</p> <p><strong>To cite this article refer to this citation:</strong></p> <p>@inproceedings{di2025e2egit,<br> title={E2EGit: A Dataset of End-to-End Web Tests in Open Source Projects},<br> author={Di Meglio, Sergio and Starace, Luigi Libero Lucio and Pontillo, Valeria and Opdebeeck, Ruben and De Roover, Coen and Di Martino, Sergio},<br> booktitle={2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR)},<br> pages={10--15},<br> year={2025},<br> organization={IEEE/ACM}<br>}<br><br>This work has been partially supported by the Italian PNRR MUR project PE0000013-FAIR. </p>
Current and future sources of open citations
<p>Video recording of the presentation done in the context of the Austrian DataCite Consortium on 19 November 2021. The video finishes after the last question.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.