Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

311

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

311 results for “Open source”

Learn how ShareScore rates datasets ↗
dryad40/100

ThermoCyte: an inexpensive open-source temperature control system for in vitro live cell imaging

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad40/100

Cardio PyMEA: A user-friendly, open-source Python application for cardiomyocyte microelectrode array analysis

Open the record for dataset details and reuse information.

publicMay 2022View details →
edi40/100

SOils DAta Harmonization database (SoDaH): an open-source synthesis of soil data from research networks

This SOils DAta Harmonization (SoDaH) database is designed to bring together soil carbon data from diverse research networks into a harmonized dataset that can be used for synthesis activities and model development. The research network sources for SoDaH span different biomes and climates, encompass multiple ecosystem types, and have collected data across a range of spatial, temporal, and depth gradients. The rich data sets assembled in SoDaH consist of observations from monitoring efforts and long-term ecological experiments. The SoDaH database also incorporates related environmental covariate data pertaining to climate, vegetation, soil chemistry, and soil physical properties. The data are harmonized and aggregated using open-source code that enables a scripted, repeatable approach for soil data synthesis.

openCC0Jul 2020View details →
zenodo36/100

Key input and output data for the multi-model analysis "Open Source Energiewende"

<p>This repository contains key input and output data of the multi-model analysis carried out in the project &quot;Open Source Energiewende&quot;, financed by the German Federal Ministry for Economic Affairs and Energy.</p> <p>The results are presented and discussed in the paper &quot;Power sector effects of cheaper stationary batteries: insights from an open multi-model analysis&quot;.</p> <p>The model codes are availabe in individual repositories, which are provided in the paper.</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Refactoring Test Smells: A Perspective from Open-Source Developers

<p>Presentation video for the <strong>5th Brazilian Symposium on Systematic and Automated Software Testing (SAST)</strong>, during the&nbsp;<strong>11th Brazilian Conference on Software: Practice and Theory (CBSoft 2020)</strong></p>

opencc-by-4.0Oct 2020View details →
dryad36/100

Quantitative analysis of subcellular distributions with an open-source, object-based tool

<p>The subcellular localization of objects, such as organelles, proteins, or other molecules, instructs cellular form and function. Understanding the underlying spatial relationships between objects through colocalization analysis of microscopy images is a fundamental approach used to inform biological mechanisms. We generated an automated and customizable computational tool, the SubcellularDistribution pipeline, to facilitate object-based image analysis from 3D fluorescence microcopy images. To test the utility of the SubcellularDistribution pipeline, we examined the subcellular distribution of mRNA relative to centrosomes within <i>Drosophila</i> embryos. Centrosomes are microtubule-organizing centers, and RNA enrichments at centrosomes are of emerging importance. Our open-source and freely available software detected RNA distributions comparably to commercially available image analysis software. The SubcellularDistribution pipeline is designed to guide the user through the complete process of preparing image analysis data for publication, from image segmentation and data processing to visualization.</p>

opencc-zeroJan 2021View details →
zenodo36/100

Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Article Pool)

<p>This pdf includes all of the articles that analyzed in the study:&nbsp;&quot;Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review&quot;.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Open-source software collaboration network mining dataset

<p>The resulting dataset of the <a href="https://github.com/gotec/git2net">git2net </a>and <a href="https://github.com/wschuell/repo_tools">repo_tools </a>mining process for randomly selected large open-source repositories.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

South African Open Data in Higher Education: Sources, resources and providers

<p>Spreadsheet of data sourced on South African sources, resources and providers of higher education open data. Composed through desk review as principle component of the situational analysis conducted for the &#39;Use of open data in the governance of South African higher education&#39; research project, in the IDRC/WWWF &#39;Exploring Emerging Impacts of Open Data in the South&#39; initiative.</p>

opencc-by-sa-4.0May 2014View details →
zenodo36/100

Mandelbugs in Open-Source Software

<p>This dataset contains a list bugs from four open-source projects (the Linux kernel, the MySQL DBMS, the Apache HTTPD web server, and the Apache AXIS WS framework). The bugs have been classified into Mandelbugs, Bohrbugs, or Aging-Related Bugs, by analyzing the conditions that exercise the bug (i.e., the "fault trigger"). This classification is useful to get insights into bugs and failures that can occur in OSS projects, and to tune testing and fault-tolerance strategies according to the distribution of bug types in a project.</p> <p>The dataset contains an ARFF file for each subsystem of the four open-source projects. Each row of the ARFF file contains:</p> <p>- An IDs of the bug, which can be used to retrieve more information about the bug from the issue tracker of the project;</p> <p>- A string that represents the class of the bug (BOH = Bohrbug; NAM = Mandelbug; ARB = Aging-Related Bug; UNK = Unknown class);</p> <p>- A string that represents the sub-class of the bug (for Bohrbugs, the sub-class is not available; the subclasses for Mandelbugs are LAG, ENV, TIM, SEQ; the subclasses for Aging-Related bugs are MEM, STO, LOG, NUM, TOT).<br>  </p>

opencc-by-4.0Aug 2017View details →
zenodo36/100

Komplementäres Datenmodell für das Open-Source-Tool OptIES

<p>Der hier verf&uuml;gbare Datensatz stellt das komplement&auml;re Datenmodell zum Open-Source-Tool <a href="https://github.com/znes/OptIES" target="_blank" rel="noopener">OptIES</a> dar.&nbsp;</p> <p>Das Tool <code>OptIES</code> dient der Optimierung regionaler Energiesysteme und wurde im Rahmen des Forschungsprojekts "OptIES D&ouml;rpum &nbsp;- Offene Optimierung sektorgekoppelter regionaler Energiesysteme am Beispiel des IES D&ouml;rpum" entwickelt. Es basiert auf der offenen Software-Toolbox <a href="https://github.com/PyPSA/PyPSA" target="_blank" rel="noopener">PyPSA</a>.</p> <p>Das Forschungsprojekt <a href="https://www.uni-flensburg.de/eum/forschung/laufende-projekte/opties-doerpum" target="_blank" rel="noopener">OptIES D&ouml;rpum</a> wurde durch die <a href="https://www.uni-flensburg.de/" target="_blank" rel="noopener">Europa-Universit&auml;t Flensburg</a> und die <a href="https://www.ecowert360.com" target="_blank" rel="noopener">EcoWert360&deg; GmbH</a> bearbeitet und durch das F&ouml;rderprogramm HWT Energie und Klimaschutz der <a href="https://www.eksh.org/" target="_blank" rel="noopener">Gesellschaft f&uuml;r Energie und Klimaschutz Schleswig-Holstein (EKSH)</a> finanziert. Es stellte die wissenschaftliche Begleitung des Praxisprojekts <a href="https://www.aktivregion-nf-nord.de/fileadmin/user_upload/KT_Klimawandel_Energie/Projekte/IES_D%C3%B6rpum/07.51_-_Beschreibung_-_Projekt_57_IES_D%C3%B6rpum.pdf" target="_blank" rel="noopener">IES D&ouml;rpum</a> dar.</p> <p>Ziel des Forschungsprojekts war es, lokale sowie nationale Herausforderungen der Energiewende durch Einordnung der kommunalen Bestrebungen in das nationale Energiesystem zu beleuchten und analysieren. Das vorliegende Tool dient der Optimierung des regionalen Energiesystems zur Untersuchung und Bewertung von Betriebsstrategien und Entwicklungspfaden des isolierten Systems. Im weiteren Projektverlauf wurde das regionale System unter Ber&uuml;cksichtigung des &uuml;bergelagerten Gesamtsystems optimiert, um integrierte Handlungsempfehlungen f&uuml;r die erfolgreiche Energiewende auf kommunaler und nationaler Ebene abzuleiten.</p> <p>Unter den hochgeladenen Dateien findet sich der Projektabschlussbericht mit einer ausf&uuml;hrlichen Darstellung des Projekts, der entwickelten Modelle und Tools, der durchgef&uuml;hrten Berechnungen und Analysen sowie der Ergebnisse. Zudem werden die im Rahmen des Projekts angefertigten Masterarbeiten von Matthias Winschu und Mohsen Mansouri bereitgestellt.</p> <p>Die hier ver&ouml;ffentlichten Daten enthalten Informationen zu Verbrauchs- und Erzeugungseinheiten innerhalb des IES D&ouml;rpum, welche f&uuml;r Modellrechnungen mit der Software-Toolbox notwendig sind. Die Daten enthalten sowohl synthetische als auch reale Daten in anonymisierter Form. Es werden die innerhalb von <code>PyPSA</code> als Standard definierten Einheiten verwendet (also z. B. MW). F&uuml;r die Erstellung der synthetische Zeitreihen ist das Jahr 2019 als Referenzjahr herangezogen.&nbsp;</p> <p>Zur Erstellung der synthetischen Lastzeitreihen (Strom und W&auml;rme) wurden Jahresgesamtverbr&auml;uche der letzten Jahre herangezogen und basierend auf der Arbeit von <a href="https://doi.org/10.1186/s42162-022-00201-y">B&uuml;ttner et al.</a> zeitlich verteilt. Zur Erstellung der synthetischen Lastzeitreihen aus Ladevorg&auml;ngen von Elektrofahrzeugen wurden Daten der Studie <a href="https://www.mobilitaet-in-deutschland.de/">Mobilit&auml;t in Deutschland</a> innerhalb des Tools <a href="https://github.com/RAMP-project/RAMP-mobility">RAMP-mobility</a> verwendet. Die synthetische Zeitreihe der potentiellen Erzeugung durch PV ist mithilfe <a href="https://www.renewables.ninja/">renewables_ninja</a> erzeugt. Alle weiteren Daten zur Parametrisierung der Komponenten sowie die synthetische Gaslast und Eigenverbr&auml;uche der enthaltenen Biogasanlage sind aufgrund von Angaben der Biogas D&ouml;rpum GmbH festgelegt und mithilfe der Daten der <a href="https://github.com/PyPSA/technology-data">PyPSA technology data</a> erg&auml;nzt (insbesondere Kostenannahmen).</p> <p>Die gemessenen Zeitreihen ergeben sich aus dem Zeitraum Mitte Februar 2023 bis Mitte Febraur 2024. Aufgrund von Hard- und Softwarefehlern stehen keine l&uuml;ckenlosen Messungen zur Verf&uuml;gung, insbesondere die Messungen der Ladevorg&auml;nge durch E-Fahrzeuge sind beeintr&auml;chtigt. Eine &Uuml;bersicht &uuml;ber die Verf&uuml;gbarkeit der gemessenen Daten ist der Abbildung <strong>Verfuegbarkeit_Messungen.png</strong> zu entnehmen.</p> <p><strong>buses.csv - </strong>enth&auml;lt Informationen, wie Name, Energietr&auml;ger und wenn anwendbar, die Spannungsebene, der verschiedenen im Modell abgebildeten Knoten.&nbsp;</p> <p><strong>generators.csv - </strong>stellt alle Erzeugungseinheiten mit ihren zugeh&ouml;rigen Knoten, installierten Kapazit&auml;ten und Kostenparametern dar.&nbsp;</p> <p><strong>lines.csv -&nbsp; </strong>beschreibt Verlauf, L&auml;nge und Art der verbauten Kabel und die zugeh&ouml;rigen elektrischen Parameter.&nbsp;</p> <p><strong>links.csv - </strong>liefert Informationen zu allen kontrollierbaren Komponenten, die zwei Knoten zweier beliebiger Energietr&auml;ger verbinden. Im vorliegenden Beispiel k&ouml;nnen das KWK-Anlagen sowie Be- und Entladeeinheiten f&uuml;r W&auml;rmespeicher sein.&nbsp;</p> <p><strong>loads.csv - </strong>beschreibt alle Verbrauchseinheiten mit deren Energietr&auml;ger und zugeh&ouml;rigem Knoten<strong>.&nbsp;</strong></p> <p><strong>storage_unit.csv - </strong>stellt die vorhandenen Batteriespeicher mit Kapazit&auml;ten, Kosten- und Effizienzparametern dar.&nbsp;</p> <p><strong>stores.csv - </strong>enth&auml;lt Informationen zu anderen Speichereinheiten f&uuml;r W&auml;rme oder Biogas mit Kapazit&auml;ten, Kosten- und Effizienzparametern.</p> <p><strong>el_load_synth.csv -&nbsp;</strong>stellt synthetisch erzeugte st&uuml;ndliche elektrischen Lastzeitreihen f&uuml;r alle im Model abgebildeten Anschlussnehmer bereit. Unter 'BGA_AC' ist der elektrische Eigenverbrauch der Biogasanlage angenommen. 'LS1' und 'LS2' beinhalten synthetische elektrische Lastzeitreihen zum Verbrauch durch Ladevorg&auml;nge von Elektrofahrzeugen unter der Annahme von zehn ('LS1') und f&uuml;nf ('LS2') Nutzer*innen.</p> <p><strong>ev_load_synth.csv -&nbsp;</strong>stellt zus&auml;tzliche, synthetische elektrische Lastzeitreihen zum Verbrauch durch Ladevorg&auml;nge von Elektrofahrzeugen bereit. Die verschiedenen Zeitreihen variieren in der Annahme der Anzahl der Nutzer*innen, diese ist in dem Spaltentitel vermerkt.</p> <p><strong>heat_load_synth.csv -</strong> h&auml;lt synthetische Daten zum W&auml;rmeverbrauch des angeschlossenen Nahw&auml;rmenetzes und zum Eigenverbrauch der Biogasanlage in st&uuml;ndlicher Aufl&ouml;sung bereit. Weiterhin ist der W&auml;rmeeigenbedarf der Biogasanlage angenommen.</p> <p><strong>gas_load_synth.csv -</strong> liefert st&uuml;ndliche synthetische Daten zum Gasverbrauch eines Satelliten-BHKWs.&nbsp;</p> <p><strong>pot_pv_timeseries_synth.csv -&nbsp;</strong>enth&auml;lt Daten zu potenziellen Stromeinspeisezeitreihen aus Photovoltaik in st&uuml;ndlicher Aufl&ouml;sung.&nbsp;</p> <p><strong>el_load_real.csv -&nbsp;</strong>stellt real gemessene elektrische Lastzeitreihen f&uuml;r alle im Model abgebildeten Anschlussnehmer in einer zeitlichen Aufl&ouml;sung von f&uuml;nf Minuten bereit. Bei Messl&uuml;cken gr&ouml;&szlig;er als f&uuml;nf Minuten wurde interpoliert. Unter 'BGA_AC' ist der elektrische Eigenverbrauch der Biogasanlage angenommen. 'LS1' und 'LS2' beinhaltet die selbe gemessene elektrische Lastzeitreihe zum Verbrauch durch Ladevorg&auml;nge von Elektrofahrzeugen. Diese wurde bei Messl&uuml;cken &gt;1h interpoliert.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Open-Source Cardiac MR Fingerprinting

<p>Magnetic Resonance (MR) raw data acquired with an open-source cardiac MR Fingerprinting (cMRF) sequence of a phantom at four different MR scanners. More details can be found here: https://github.com/PTB-MR/cMRF. The colormaps are taken from https://zenodo.org/records/11185704 because zenodo_get failed on trying to download this record in a jupyter notebook.</p> <p>Additionally cMRF data was acquired in three volunteers who were scanned at two different scanners. Cartesian and golden radial cine data was acquired to verify the anatomical features seen in the quantitative maps.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Dataset: A continuous open source data collection platform for architectural technical debt assessment

<p>The dataset and replication package of the study &quot;A continuous open source data collection platform for architectural technical debt assessment&quot;.</p> <p>&nbsp;</p> <p>Abstract</p> <p>Architectural decisions are the most important source of technical debt.&nbsp; In recent years, researchers spent an increasing amount of effort investigating this specific category of technical debt, with quantitative methods, and in particular static analysis, being the most common approach to investigate such a topic.</p> <p>&nbsp;</p> <p>However, quantitative studies are susceptible, to varying degrees, to external validity threats, which hinder the generalisation of their findings.</p> <p>In response to this concern, researchers strive to expand the scope of their study by incorporating a larger number of projects into their analyses. This practice is typically executed on a case-by-case basis, necessitating substantial data collection efforts that have to be repeated for each new study.</p> <p>&nbsp;</p> <p>To address this issue, this paper presents our initial attempt at tackling this problem and enabling researchers to study architectural smells at large scale, a well-known indicator of architectural technical debt. Specifically, we introduce a novel approach to data collection pipeline that leverages Apache Airflow to continuously generate up-to-date, large-scale datasets using Arcan, a tool for architectural smells detection (or any other tool).</p> <p>Finally, we present the publicly-available dataset resulting from the first three months of execution of the pipeline, that includes over 30,000 analysed commits and releases from over 10,000 open source GitHub projects written in 5 different programming languages and amounting to over a billion of lines of code analysed.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Sample Stripped Pre-supernova Progenitors for open-source code CHIPS (Complete History for Interaction-Powered Supernovae)

<p>Inlists, mainly based on the test suite "example_make_pre_ccsn" in r12778, with slight amendments for removal of hydrogen (and helium, for Ic progenitors) envelope at core hydrogen (helium) exhausion.</p><p>For details: https://ui.adsabs.harvard.edu/abs/2023arXiv230810785T/abstract</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Artifact for "Inside Bug Report Templates: An Empirical Study on Bug Report Templates in Open-Source Software"

<p>This is the artifact for the&nbsp;paper "Inside Bug Report Templates: An Empirical Study on Bug Report Templates in Open-Source Software".</p> <p><strong>What the artifact&nbsp;does:</strong><br>1) a questionnaire that we used for our online survey (PDF);<br>2) the valid responses of our online survey (CSV).</p> <p>3) the code of preprocessing (.py).</p> <p>4) the dataset of preprocessing and labeling (CSV).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Estimating Usage Of Open Source Projects - Flutter Telemetry Case Study - MSR' 24

<p>This dataset (CSV) was assembled to support analysis within a case study that will be published in the proceedings of the <a href="https://conf.researchr.org/home/msr-2024">Mining Software Repositories</a> conference (MSR &lsquo;24) April 14-15 2024: <a title="Estimating Usage Of Open Source Projects" href="https://doi.org/10.1145/3643991.3645066">Estimating Usage Of Open Source Projects.</a></p> <p>This case study explored whether publicly available metrics could serve as proxies to estimate usage of an open source project. Using the <a href="https://flutter.dev/">Flutter</a> project as our case study, we collected monthly proxy metrics from GitHub, StackOverflow and Slack to compare with Flutter&rsquo;s monthly active user count over the same time period: January 2018 through February 2021.</p> <p>All metrics correspond to the last day of the month and/or represent aggregate activity in that month. Data from GitHub shows aggregate activity counts across the entire <a href="https://github.com/flutter">Flutter GitHub organization</a> (up to 33 repositories).&nbsp;</p> <p>Our specific metrics include:</p> <div> <table> <tbody> <tr> <td> <p>Source</p> </td> <td> <p>Metric</p> </td> <td> <p>Aggregation method</p> </td> <td> <p>Details</p> </td> </tr> <tr> <td> <p>Flutter</p> </td> <td> <p>Monthly Active Users (MAU)</p> </td> <td> <p>Google internal tooling</p> </td> <td> <p>Flutter users active in the last 30 days, collected on the last day of each month</p> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>PullRequest Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> <p>As Google governs changes to this code base, we excluded known Google employees in code change related metrics&nbsp;</p> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Issues Created in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Issue Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Fork events in month</p> </td> <td> <p><a href="http://gharchive.org">gharchive.org</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Fork cumulative count at the end of month</p> </td> <td> <p><a href="http://gharchive.org">gharchive.org</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>StackOverflow</p> </td> <td> <p>Question Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>StackOverflow</p> </td> <td> <p>Questions in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>Slack</p> </td> <td> <p>Claimed (cumulative) members at the end of the month</p> </td> <td> <p>fluttercommunity.slack</p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>Slack</p> </td> <td> <p>Cumulative messages at the end of the month</p> </td> <td> <p>fluttercommunity.slack</p> </td> <td>&nbsp;</td> </tr> </tbody> </table> </div> <p>Tools:&nbsp;</p> <ul> <li> <p>We used an instance of <a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a> hosted by <a href="https://bitergia.com/">Bitergia</a> to aggregate GitHub PullRequest Authors, GitHub Issue Authors, GitHub Issues, across all repositories under the Flutter Organization, and StackOverflow Question Authors, and StackOverflow Questions for questions that mention Flutter. We used the <a href="http://github.com/chaoss/grimoirelab-sortinghat">Sorting Hat</a> of feature GrimoireLab to identify Google employees in this sample.</p> </li> </ul> <ul> <li> <p><a href="http://gharchive.org">GHArchive</a> via <a href="https://cloud.google.com/blog/topics/public-datasets/github-on-bigquery-analyze-all-the-open-source-code">BigQuery</a> was used to count GitHub Star and Fork events across all repositories under the Flutter Organization</p> </li> <li> <p>We pulled Slack activity directly from&nbsp;<a href="http://fluttercommunity.slack.com">Flutter's slack channel </a>dashboard</p> </li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Chinese Reference Population: open-source age-dependent computational phantoms of reference Chinese population

<p>The<strong> Chinese Reference Population (CRP)</strong> phantoms dataset encompass <strong>30 phantoms</strong> available in both voxel and NURBS formats, with age in 0, 1, 2, 3, 4, 5, 6, 8, 10, 12, 15, 18 years and adult male and female, as well as 4&nbsp;pregnant women and fetus in early pregnancy, first trimester, second trimester and third trimester.</p> <ul> <li><strong>Voxelized phantoms</strong> are accessible in NII format :<strong> <em>"XXX.nii", which could be opened in AMIDE software.</em></strong></li> <li>Excel file<strong> </strong>containing<strong> organ masses and other descriptive information </strong>:<strong> <em>"CRP_descriptive_Info.xlsx"</em></strong></li> <li>In the application of F18&minus;FDG dose calculation,&nbsp;<strong>organ absorbed doses per unit activity administered </strong>is provided in an Excel file :<strong> <em>"Application_F18-FDG.xlsx"</em></strong></li> </ul> <p>All data are stored on Zenodo and can be publicly accessed.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

GloHydroRes - a global dataset combining open-source hydropower plant and reservoir data

<div> <div> <div> <div>&nbsp;</div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>Analyzing the impacts of drought and climate change on hydropower requires detailed data not only on hydropower attributes such as plant type, head, and installed capacity, but also on reservoir characteristics like area, depth, and volume. Current open-source hydropower datasets typically lack information on reservoirs, while reservoir datasets often omit hydropower details. GloHydroRes is a global dataset that integrates open-source hydropower and reservoir data, offering 29 attributes, including key information such as installed capacity, plant type, dam height, reservoir depth, area, volume, and river name. Overall, GloHydroRes provides data on 7,775 hydropower plants across 128 countries.</p> </div> </div> </div> </div> <div> <div> <div>&nbsp;</div> </div> </div> </div> </div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

E2EGit: A Dataset of End-to-End Web Tests in Open Source Projects

<p><strong>ABSTRACT&nbsp;</strong><br>End-to-end (E2E) testing is a software validation&nbsp;approach that simulates realistic user scenarios throughout&nbsp;the entire workflow of an application. In the context of web<br>applications, E2E testing involves two activities: Graphic User&nbsp;Interface (GUI) testing, which simulates user interactions with&nbsp;the web app&rsquo;s GUI through web browsers, and performance&nbsp;testing, which evaluates system workload handling. Despite its&nbsp;recognized importance in delivering high-quality web applications, the availability of large-scale datasets featuring real-world&nbsp;E2E web tests remains limited, hindering research in the field.<br>To address this gap, we present E2EGit, a comprehensive&nbsp;dataset of non-trivial open-source web projects collected on&nbsp;GitHub that adopt E2E testing. By analyzing over 5,000 web repositories across popular programming languages (JAVA,&nbsp;JAVASCRIPT, TYPESCRIPT, and PYTHON), we identified 472 repositories implementing 43,670 automated Web GUI tests with&nbsp;popular browser automation frameworks (SELENIUM, PLAYWRIGHT, CYPRESS, PUPPETEER), and 84 repositories that featured 271 automated performance tests implemented leveraging&nbsp;the most popular open-source tools (JMETER, LOCUST). Among&nbsp;these, 13 repositories implemented both types of testing for a total&nbsp;of 786 Web GUI tests and 61 performance tests.</p> <p><br><strong>DATASET DESCRIPTION&nbsp;</strong><br>The dataset is provided as an SQLite database,&nbsp;whose structure is illustrated in Figure 3 (in the paper), which consists of&nbsp;five tables, each serving a specific purpose.<br>The <em>repository </em>table contains information on 1.5 million&nbsp;repositories collected using the SEART tool on May 4. It&nbsp;includes 34 fields detailing repository characteristics. The<br><em>non_trivial_repository</em> table is a subset of the previous one, listing repositories that passed the two filtering stages described in the pipeline. For each repository, it specifies whether it is a web repository using JAVA, JAVASCRIPT, TYPESCRIPT, or PYTHON frameworks. A repository may use multiple frameworks, with corresponding fields (e.g., is web java) set to true, and the field web dependencies listing the detected web&nbsp;frameworks. For Web GUI testing, the dataset includes two&nbsp;additional tables; <em>gui_testing_test _details</em>, where each row represents a test file, providing the file path, the browser&nbsp;automation framework used, the test engine employed, and the&nbsp;number of tests implemented in the file. <em>gui_testing_repo_details</em>, aggregating data from the previous table at the repository&nbsp;level. Each of the 472 repositories has a row summarizing<br>the number of test files using frameworks like SELENIUM or&nbsp;PLAYWRIGHT, test engines like JUNIT, and the total number&nbsp;of tests identified. For performance testing, the <em>performance_testing_test_details</em> table contains 410 rows, one for each test identified. Each row includes the file path, whether the test uses JMETER or LOCUST, and extracted details such as the number of thread groups, concurrent users, and requests. Notably, some fields may be absent&mdash;for instance, if external files (e.g., CSVs defining workloads) were unavailable, or in the case of Locust tests, where parameters like duration and concurrent users are specified via the command line.</p> <p><strong>To cite this article refer to this citation:</strong></p> <p>@inproceedings{di2025e2egit,<br>&nbsp; title={E2EGit: A Dataset of End-to-End Web Tests in Open Source Projects},<br>&nbsp; author={Di Meglio, Sergio and Starace, Luigi Libero Lucio and Pontillo, Valeria and Opdebeeck, Ruben and De Roover, Coen and Di Martino, Sergio},<br>&nbsp; booktitle={2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR)},<br>&nbsp; pages={10--15},<br>&nbsp; year={2025},<br>&nbsp; organization={IEEE/ACM}<br>}<br><br>This work has been partially supported by the Italian PNRR MUR project PE0000013-FAIR.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Current and future sources of open citations

<p>Video recording of the presentation done in the context of the Austrian DataCite Consortium on 19 November 2021. The video finishes after the last question.</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record