Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

374

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

374 results for “Power Data”

Learn how ShareScore rates datasets ↗
zenodo44/100

Simulation data for "Characteristics of Wave-Particle Power Transfer as a Function of Electron Pitch Angle in Nonlinear Frequency Chirping" which will be submitted to Journal of Geophysical Research: Space Physics

<p>Simulation data for "Characteristics of Wave-Particle Power Transfer as a Function of Electron Pitch Angle in Nonlinear Frequency Chirping" which will be submitted to Journal of Geophysical Research: Space Physics.</p> <p>Including the simulation input parameter file and the necessary output data to plot each figure in the article.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Codes and Data for 'Cost-effective Planning of Decarbonized Power-Gas Infrastructure to Meet the Challenges of Heating Electrification'

<p>The codes and data used in the followng paper</p> <p>''Khorramfar, R., Santoni-Calvin, M., Mallapragada, D., Amin, S., Botterud, A.,<br>Norfork L., (2025) Cost-effective Planning of Power-Gas Infrastructure to Meet the Challenges<br>of Heating Electrification, Cell Reports Sustainability</p> <p>&nbsp;</p> <p>Link (open source): https://www.cell.com/cell-reports-sustainability/fulltext/S2949-7906(25)00003-5</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Open data and power dynamics

<p>Podcast with an expert -Mor Rubenstein- about the management of open data and the power dynamics that are at play in society. In particular we are referring mainly to western Europe as the societal context.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

BOREALIS Power Analysis Code and Data

<p>This contains the code and data necessary to rerun the power analysis used in testing BOREALIS.</p> <p>Borealis is an R library performing outlier analysis for count-based bisulfite sequencing data. It detects outlier methylated CpG sites from bisulfite sequencing (BS-seq). The core of Borealis is modeling Beta-Binomial distributions. This can be useful for rare disease diagnoses.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Data set of German coal power stations for coal phase out auctions

<p>This repository contains the coal power station data set used in the paper &quot;Auctions to phase out coal power: Lessons learned from Germany&quot; by Silvana Tiedemann and Finn M&uuml;ller-Hansen, published in Energy Policy in 2023. All details about the data set are provided in the readme.md.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Data to reproduce the results: Statistical power of spatial earthquake forecast tests

<p>We provide data needed to reproduce the figures from the publication titled &quot;Statistical power of spatial earthquake forecast tests&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Reference data set for a Norwegian medium voltage power distribution system

<p>This reference data set describes a representative Norwegian radial, medium voltage (MV) electric power distribution system operated at 22 kV. The data set is developed in the Norwegian research centre CINELDI and will in brief be referred to as the CINELDI MV reference system.</p> <p>Data for a real Norwegian distribution system were provided by a distribution grid company. The data have been anonymized and processed to obtain a simplified but still realistic grid model with 124 nodes. The data set consists of the following three parts:<br> 1. Grid data files: describe the base version of the reference system that represents the present-day state of the grid, including information about topology, electrical parameters, and existing load points.<br> 2. Load data files: comprise load demand time series for a year with hourly resolution and scenarios for the possible long-term development of peak load. These data describe an extended version of the reference system with information about possible new load points being added to the system in the future.<br> 3. Reliability data files: contain data necessary for carrying out reliability of supply analyses for the system.</p> <p>The data set is described in detail in the following data article:<br> I. B. Sperstad, O. B. Fosso, S. H. Jakobsen, A. O. Eggen, J. H. Evenstuen, and G. Kj&oslash;lle, &ldquo;Reference data set for a Norwegian medium voltage power distribution system,&rdquo; Data in Brief, 109025, 2023, doi: 10.1016/j.dib.2023.109025.</p>

opencc-by-4.0Oct 2022View details →
edi44/100

Optimizing sampling across methods improves the power of ecological monitoring data

Transect-based monitoring has long been a valuable tool in ecosystem monitoring. These transects are often used to measure multiple ecosystem attributes. The line-point intercept (LPI), vegetation height, and canopy gap intercept methods comprise a set of core methods, which provide indicators of ecosystem condition. However, users struggle to design a sampling strategy that optimizes the ability to detect ecological change using transect-based methods. We assessed the sensitivity of these core methods on a one-hectare plot to transect length, number, and sampling interval to determine: 1) minimum sampling required to describe ecosystem characteristics and detect change for each method and 2) optimal transect length and number for all three methods to make recommendations for future analyses and monitoring efforts. We used data from 13 National Wind Erosion Research Network locations spanning the western US, which included 151 measurements over time across five biomes. We found that longer and increased numbers of transects were more important for reducing sampling error than increased sample intensity along transects. For all methods and indicators across plots, three 100-m transects reduced sampling error so that indicator estimates fall within an 95% confidence interval of +/- 5% for canopy gap intercept and LPI-total foliar cover, +/- 5 cm for height and +/- two species for LPI-species counts. For the same criteria at 80% confidence intervals, two 100-m transects are needed. Site-scale inference was strongly affected by sample design, consequently our understanding of ecological dynamics may be influenced by sampling decisions.

openCC (other)Sep 2024View details →
zenodo40/100

Raw data for: Sex and Power: Sexual dimorphism in trait variability and its eco-evolutionary and statistical implications

<p>This is a dataset obtained from the EBI in August 2018. The full code, analyses and processed data that are associated with the paper that is based on this large dataset can be found here:&nbsp;<a href="https://github.com/itchyshin/mice_sex_diff">https://github.com/itchyshin/mice_sex_diff</a></p> <p>&nbsp;</p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Data From: Powerful detection of polygenic selection and environmental adaptation in US beef cattle

<p>GEMMA output containing summary statistics for generation proxy selection mapping (GPSM) and environmental GWAS (envGWAS) selection analyses from&nbsp;<br> Rowan et al. &quot;Powerful detection of polygenic selection and environmental adaptation in US beef cattle&quot; 2021<br> https://doi.org/10.1101/2020.03.11.988121&nbsp; &nbsp;&nbsp;</p> <p>File names identify the analysis run, for example<br> &quot;Gelbvieh_envgwas_desert_summary_stats.txt.gz&quot;<br> Is the Gelbvieh dataset analyzed using the Desert ecoregion as the dependent variable&nbsp;<br> in a univariate envGWAS model.&nbsp;</p> <p>Files are formated according to GEMMA output.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Data for Solar Field Output Temperature Optimization Using a MILP Algorithm and a 0D Model in the Case of a Hybrid Concentrated Solar Thermal Power Plant for SHIP Applications

<p>These data were generated for the Open-Acces Article :</p> <p>Kamerling, S.; Vuillerme, V.; Rodat, S. Solar Field Output Temperature Optimization Using a MILP Algorithm and a 0D Model in the Case of a Hybrid Concentrated Solar Thermal Power Plant for SHIP Applications.&nbsp;<em>Energies</em>&nbsp;<strong>2021</strong>,&nbsp;<em>14</em>, 3731. https://doi.org/10.3390/en14133731</p> <p>In these dataset, the data for the Case Study and the Sensitivity Analysis are available. Jupyter Notebooks for further process of these data are also available. The NoteBooks AnalyseHourlyValues,&nbsp;AnalyseDailyValues and&nbsp;AnalyseMonthlyValues allow for easy change of variable, whereas CaseStudyAnalysis is for one specific set of data. The AnalyseSets were created in order to analyse the influence of the optimization on the solar fraction of the different datasets.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Data and plot scripts for "Rising complexity and falling explanatory power in ecology"

<p>Analyses of published research can provide a realistic perspective on the progress of science. By analyzing more than 18 000 articles published by the preeminent ecological societies, we found that (1) ecological research is becoming increasingly statistically complex, reporting a growing number of&nbsp;<em>P</em>&nbsp;values per article and (2) the value of reported coefficient of determination (<em>R</em>2) has been falling steadily, suggesting a decrease in the marginal explanatory power of ecology. These trends may be due to changes in the way ecology is studied or in the way the findings of investigations are reported. Determining the reason for increasing complexity and declining marginal explanatory power would require a critical review of the scientific process in ecology, from research design to dissemination, and could influence the public interpretation and policy implications of ecological findings.<br /> <br /> <br /> Read More:&nbsp;http://www.esajournals.org/doi/abs/10.1890/130230</p>

openother-openSep 2014View details →
zenodo40/100

Illustrative dataset for the article: Vieira, R., McDonald, S., Araujo-Soares, V., Sniehotta, F., Henderson, R. (2017) "Dynamic modelling of n-of-1 data: Powerful and flexible data analytics applied to individualised studies"

<p>This dataset is supplementary material of the manuscript "Dynamic modelling of n-of-1 data: Powerful and flexible data analytics applied to individualised studies. McDonald et al. (2016) presents a series of novel n-of-1 studies that intended to explore the relationship between physical activity change during the retirement transition. The file contains the data of one participant. The column names correspond to the following variables:</p> <p>time: duration of follow-up (minutes);<br> minute: time of day (hours and minutes);<br> day_num: day since beginning of follow-up (the first two days were considered as adaptation phase and therefore removed); <br> PAscore: accelerometer raw score; <br> startBout: 1 (a bout of PA was initiated in this minute) or 0 (a bout of PA wasn't <br> initiated in this minute); <br> nPAbouts_day: number of PA bouts per day; <br> nPAbouts_day.l1: number of PA bouts in previous day (lag 1); <br> nPAbouts_day.l2: number of PA bouts two day before (lag 2); <br> nBoutsLast2hours: number of PA bouts in previous 2 hours; <br> retirement: 0 (before retirement) or 1 (after retirement)<br> weekday: 0 (workday) or 1 (weekend)<br> sleepLength: number of hours of sleep last night<br> sleepLength.l1: number of hours of sleep the night before<br> sleepLength.l2: number of hours of sleep two nights before<br> pers: personalised measure of partner's influence (scale 0-1)<br> periodDay: morning, evening or afternoon</p> <p>McDonald, S., Vieira, R., O'Brien, N., White, M., &amp; Sniehotta, F. F. (2016). Does physical activity and sedentary behavior change during the retirement transition? Findings from a series of novel n-of-1 natural experiments. <em>International Journal of Behavioral Medicine, 23</em>, S261-S261.</p> <p> </p>

opencc-by-4.0May 2017View details →
zenodo40/100

Data for: Downscaled gridded global dataset for Gross Domestic Product (GDP) per capita at purchasing power parity (PPP) over 1990-2022

<p>This dataset provides a gridded dataset for GDP per capita at purchasing power parity (PPP) downscaled to an admin 2 level (43,501 admin units). The dataset is based on reported subnational admin data (from 89 countries and 2,708 subnational units) and spans three decades from 1990 to 2022.&nbsp;</p> <p>The dataset is presented in details in the following publication.&nbsp;<strong><em>Please cite this paper when using data.&nbsp;</em></strong></p> <p>Kummu, M., Kosonen, M. &amp; Masoumzadeh Sayyar, S. 2025. Downscaled gridded global dataset for gross domestic product (GDP) per capita PPP over 1990&ndash;2022. Scientific Data 12: 178. <a href="https://doi.org/10.1038/s41597-025-04487-x" target="_blank" rel="noopener">https://doi.org/10.1038/s41597-025-04487-x</a></p> <p><strong>Code is available</strong> at: <a href="https://github.com/mattikummu/griddedGDPpc" target="_blank" rel="noopener">https://github.com/mattikummu/griddedGDPpc&nbsp;</a></p> <p>&nbsp;</p> <p><strong>The following data is given (formats in brackets)</strong></p> <ul> <li>GDP per capita (PPP) at admin 0 level (national) (GeoTIFF, gpkg, csv)</li> <li>GDP per capita (PPP) at admin 1 level (at the level of reporting, either admin 1 level or admin 0 level) (GeoTIFF, gpkg, csv)</li> <li>GDP per capita (PPP) at admin 2 level (downscaled from admin 1 level) (GeoTIFF, gpkg, csv)</li> <li>Total GDP (PPP), downscaled admin 2 level GDP per capita (PPP) multiplied by gridded population count, with three resolutions: 30 arc-sec, 5 arc-min, and 30 arc-min (GeoTIFF)&nbsp;</li> <li>Input data for the script that was used to generate the data above (code_input_data.zip). Code available at https://github.com/mattikummu/griddedGDPpc&nbsp;</li> </ul> <p><strong>Files are named as follows</strong><br><em>Format</em>: raster data (GeoTIFF) starts with rast_*, polygon data (gpkg) with polyg_*, and tabulated with tabulated_*.&nbsp;<br><em>Admin levels:</em> adm0 for admin 0 level, adm1 for admin 1 level, and adm2 for admin 2 level<br><em>Product type:</em> GDP per capita at purchasing power parity (PPP): _gdp_perCapita_; and total GDP at purchasing power parity (PPP): _gdp_tot_</p> <p>&nbsp;</p> <p><strong>Metadata&nbsp;</strong></p> <p><em>Grids for GDP per capita data:</em></p> <p>Resolution: 5 arc-min (0.083333333 degrees) &nbsp;(for admin 2 level also 30 arc-min, 0.5 degree, resolution is provided)</p> <p>Spatial extent: Lon: -180, 180; -90, 90 (xmin, xmax, ymin, ymax)&nbsp;</p> <p>Coordinate ref system: EPSG:4326 - WGS 84&nbsp;</p> <p>Format: Multiband geotiff; each band for each year over 1990-2022&nbsp;</p> <p>Unit: USD in 2017 international dollars</p> <p>&nbsp;</p> <p><em>Grids for total GDP:</em></p> <p>Resolution: 30 arc-sec, 5 arc-min or 30 arc-min</p> <p>Spatial extent: Lon: -180, 180; -90, 90 (xmin, xmax, ymin, ymax)&nbsp;</p> <p>Coordinate ref system: EPSG:4326 - WGS 84&nbsp;</p> <p>Format: Multiband geotiff; each band for each year over 1990-2022 (5 arc-min, 30 arc-min) or for each five years 1990, 1995, ... 2015, 2020 (30 arc-sec)</p> <p>Unit: USD in 2017 international dollars</p> <p>&nbsp;</p> <p><em>Geospatial polygon (gpkg) files:&nbsp;</em></p> <p>Spatial extent:&nbsp;-180, 180; -90, 83.67 (xmin, xmax, ymin, ymax)&nbsp;</p> <p>Temporal extent: annual over 1990-2022</p> <p>Coordinate ref system: EPSG:4326 - WGS 84&nbsp;</p> <p>Format: gkpk&nbsp;</p> <p>Unit: &nbsp;USD in 2017 international dollars</p>

opencc-by-4.0Oct 2022View details →
dryad40/100

Data from: harnessing the power of regional baselines for broad-scale genetic stock identification: a multistage, integrated, and cost-effective approach

<p>In mixed-stock fishery analyses, genetic stock identification (GSI) estimates the contribution of each population to a mixture and is typically conducted at a regional scale using genetic baselines specific to the stocks expected in that region. Often these regional baselines cannot be combined to produce broader geographical baselines due to non-overlapping populations and genetic markers. In cases where the mixture contains stocks spanning across a wide area, a broad-scale baseline is created, but often at the cost of resolution. Here, we introduce a new GSI method to harness the resolution capabilities of baselines developed for regional applications in the analysis of mixtures containing individuals from a broad geographic range. This method employs a multistage framework that allows disparate baselines to be used in a single integrated process that produces estimates along with the propagated errors from each stage. All individuals in the mixture sample are required to be genotyped for all genetic markers in the baselines used by this model, but the baselines do not require overlap in genetic markers or populations representing the broad-scale or regional baselines.</p> <p>We demonstrate our integrated multistage GSI model using a synthesized data set made up of Chinook salmon, <em>Oncorhynchus tshawytscha</em>, from the North Bering Sea of Alaska. The data set is designed to be run using R package, Ms.GSI, and it does not represent the composition of the real fishery. The results show an improved accuracy for estimates using an integrated multistage framework, compared to the conventional framework of using separate hierarchical steps. The integrated multistage framework allows GSI of a wide geographic area without first developing a large scale, high-resolution genetic baseline or dividing a mixture sample into smaller regions beforehand. This approach is more cost-effective than updating range-wide baselines with all regionally important markers.</p>

opencc-zeroDec 2023View details →
zenodo40/100

MaStR power unit registry for eGon-data (legacy)

<p>The data set contains selected data from the <a href="https://www.marktstammdatenregister.de/MaStR">Marktstammdatenregister</a> for eGon-data pipeline.<br><strong>Note:</strong> This dataset is legacy and should be only used for compatibility reasons in eGon-data.</p> <p><strong>Dump version: 2021-04-30</strong></p> <p>The data set has been downloaded and processed with a very early version of the software&nbsp;<a href="https://github.com/OpenEnergyPlatform/open-MaStR">open-MaStR</a>.</p> <p>License information:<br>Marktstammdatenregister - &copy; Bundesnetzagentur f&uuml;r Elektrizit&auml;t, Gas, Telekommunikation, Post und Eisenbahnen |&nbsp;<a href="https://www.govdata.de/dl-de/by-2-0">DL-DE-BY-2.0</a></p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

MaStR power unit registry for eGon-data

<p>The data set contains selected data from the <a href="https://www.marktstammdatenregister.de/MaStR">Marktstammdatenregister</a> for eGon-data pipeline.</p> <p><strong>Dump version: 2022-11-17</strong></p> <p>The data set has been downloaded and processed with the software <a href="https://github.com/OpenEnergyPlatform/open-MaStR">open-MaStR</a>.<br>Compared to the <a href="../doi/10.5281/zenodo.6807425">official open-MaStR releases</a> a custom format postprocessing has been applied.</p> <p>License information:<br>Marktstammdatenregister - &copy; Bundesnetzagentur f&uuml;r Elektrizit&auml;t, Gas, Telekommunikation, Post und Eisenbahnen |&nbsp;<a href="https://www.govdata.de/dl-de/by-2-0">DL-DE-BY-2.0</a></p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Increase of active-power-based flexibility (data for KPI evaluation)

<p>There are stored data collected from new EV charging stations &ndash; this will be used as an aggregated source of active power-based flexibility procured for system operator (DSO) and managed through non frequency platform.&nbsp; Data originated from part of the CZ DEMO run directly by ČEZ distribuce called "e &ndash; fleet". At Zenodo there are data from all sites (EV charging poles while the first excel sheet contains calculation of the KPI &ldquo;Increase of active-power-based flexibility&rdquo;. Detailed explanation on KPI evaluation, data format and main findings from the tests are included in the <a title="link" href="https://www.onenet-project.eu/wp-content/uploads/2024/03/OneNet_D10.5_V1.0.pdf">Deliverable 10.5 (section 2.1.2).&nbsp;</a></p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Dragon_Pi: IoT Side-Channel Power Data Intrusion Detection Dataset and Unsupervised Convolutional Autoencoder for Intrusion Detection

<h2><strong>Dragon_Pi</strong></h2> <div> <div>For a more in depth description of the Dragon_Pi dataset, please consult the journal article of the same name:</div> <div>Lightbody <em>et al.</em>, Future Internet, 2024, <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a> - specifically Section 3.2: Dataset Overview.</div> <div>&nbsp;</div> </div> <p>Dragon_Pi is an intrusion detection dataset for IoT devices. In the field of IoT security there are few datasets, and those which do exist tend to focus solely on network traffic. The Dragon_Pi dataset seeks to provide not only more data for the field of IoT security, but also, data of a somewhat under-published type: linear time series power consumption data.</p> <p>Dragon_Pi is a fully labelled Intrusion Detection dataset for IoT devices. It is composed of both normal and under-attack power consumption data obtained from two separate testbeds - one using a DragonBoard 410c and the other a Raspberry Pi Model 3 - Hence the moniker&nbsp;<em>Dragon_Pi</em>.&nbsp;</p> <p>These testbeds were set up with predefined normal behavour as described in the attached publications. The normal linear time series power consumption&nbsp; was sampled from the testbed under these normal conditions. Both testbeds were then attacked using some common attacks on IoT - the linear time series power consumption captured under these condtions as well.&nbsp;</p> <p>Specifically, the testbeds were subjected to the Port Scan (using Nmap), SSH Brute Force (using Hydra) and SYNFlood Denial of Service (using Hping3) attacks. These attacks were repeated to gain insight to what their signatures looked like and also how varying the tool settings effected the resultant signature.&nbsp; A fourth type of scenario was also conducted on the testbeds - the "Capture the Flag" scenarios. In these files multiple attack types were used with a more specific target - to exfiltrate a hidden file from the testbeds.</p> <p>Each file has three hierarchical levels of annotation for <strong>each sample</strong> within:</p> <ol> <li>A simple "Normal or Anomaly" label for the specific sample</li> <li>A specifc attack type label e.g. "SSH Bruteforce", for the specific sample</li> <li>A specific tool setting for that attack e.g. "Hydra_T16", for the specific sample</li> </ol> <p>Users can decide for themselves what level of annotation they require for their specific task.&nbsp;</p> <p>Each file in the Dragon_Pi dataset is accompanied by its own legend file. This file explains the contents of the specific .csv file and the specific indexes of the events within.</p> <p>The Dragon_Pi dataset consists of approximately 67 files, as shown in Table 1. Compressed, the datset totals approximately 13GB. Completely decompressed the dataset is approximately 80GB ( 30GB Pi data, 50 GB Dragon data).&nbsp;</p> <div>&nbsp;</div> <div> <table> <tbody> <tr> <td>Label Type</td> <td>Specific Label&nbsp;</td> <td>Number of Files DragonBoard 410c</td> <td>Number of Files Raspberry Pi</td> </tr> <tr> <td>Normal&nbsp;</td> <td>Normal&nbsp;</td> <td>3&nbsp;</td> <td>2</td> </tr> <tr> <td>Port Scan Attack&nbsp;</td> <td>Nmap_T5</td> <td>2</td> <td>1</td> </tr> <tr> <td>&nbsp;</td> <td>Nmap_T4</td> <td>1</td> <td>1</td> </tr> <tr> <td>&nbsp;</td> <td>Nmap_T3</td> <td>1</td> <td>1</td> </tr> <tr> <td>&nbsp;</td> <td>Nmap_T2</td> <td>1</td> <td>1</td> </tr> <tr> <td>SSH Brute Force</td> <td>Hydra_T32</td> <td>4</td> <td>2</td> </tr> <tr> <td>&nbsp;</td> <td>Hydra_T16</td> <td>16</td> <td>2</td> </tr> <tr> <td>&nbsp;</td> <td>Hydra_T3</td> <td>8</td> <td>2</td> </tr> <tr> <td>&nbsp;</td> <td>Hydra_T1</td> <td>5</td> <td>2</td> </tr> <tr> <td>SYNFlood DOS</td> <td>SYNFlood DOS</td> <td>1</td> <td>1</td> </tr> <tr> <td>Capture the Flag</td> <td>Misc Attacks</td> <td>3</td> <td>5</td> </tr> </tbody> </table> </div> <div>Table 1. Enumeration of the in the Dragon_Pi dataset.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>For a more in depth description of the Dragon_Pi dataset, please consult the journal article of the same name:</div> <div>Lightbody <em>et al.</em>, Future Internet, 2024, <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a> - specifically Section 3.2: Dataset Overview.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div><strong>Publication of this dataset:</strong></div> <div>&nbsp;</div> <div>This dataset was published in Lightbody&nbsp;<em>et al.</em>, Future Internet, 2024, <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a>. Consult and cite this article for a more in depth dataset description, as well as an in depth review of first AI Intrusion Detection model trained on this dataset.&nbsp;</div> <div>&nbsp;</div> <div>See article Lightbody <em>et al.</em>, Future Internet, 2023, <a href="https://doi.org/10.3390/fi15050187">https://doi.org/10.3390/fi15050187</a> for a detailed investigation on&nbsp; the attack signatures discovered while creating this dataset. This work was an inital investigation of the dataset and can serve as a part 1 to the Dragon_Pi paper.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div><strong>How to cite this dataset in your work:&nbsp;</strong></div> <div>&nbsp;</div> <div>Please cite these two DOIs when publishing using this dataset:</div> <div> <ol> <li>Dragon_Pi release publication: <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a> (most important)</li> <li>Zenodo Dataset DOI: https://doi.org/10.5281/zenodo.10784947</li> </ol> </div> <div> <div>&nbsp;</div> </div> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Data for "A learned score function improves the power of mass spectrometry database search"

<div> <h1>DATA for "A learned score function improves the power of mass spectrometry database search"</h1> <br> <div>These data files are associated with the following publication:</div> <br> <div> <ul> <li>Varun Ananth, Justin Sanders, Melih Yilmaz, Sewoong Oh and William Stafford Noble. "<a title="biorXiv Preprint Link" href="https://www.biorxiv.org/content/10.1101/2024.01.26.577425v2" target="_blank" rel="noopener">A learned score function improves the power of mass spectrometry database search</a>". Bioinformatics (Proceedings of the ISMB). &nbsp;2024.</li> </ul> </div> <br> <div>For the benchmarking data, we used a dataset that is publicly available on ProteomeXchange (PXD028735). The paper that introduced this dataset is:</div> <br> <div> <ul> <li>Van Puyvelde, B., Daled, S., Willems, S., Gabriels, R., Gonzalez de Peredo, A., Chaoui, K., Mouton-Barbosa, E., Bouyssi&eacute;, D., Boonen, K., Hughes, C. J., Gethings, L. A., Perez-Riverol, Y., Bloomfield, N., Tate, S., Schiltz, O., Martens, L., Deforce, D., &amp; Dhaenens, M. (2022). A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. In Scientific Data (Vol. 9, Issue 1). Springer Science and Business Media LLC. https://doi.org/10.1038/s41597-022-01216-6</li> </ul> </div> <br> <div>More specifically, the following `.raw` files were downloaded:</div> <br> <ul> <li><code>LFQ_Orbitrap_DDA_Ecoli_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Human_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Yeast_01.raw</code></li> </ul> <br> <div>Those files can be accessed via FTP&nbsp;<a title="Link to ProteomeXchange: PXD028735" href="https://ftp.pride.ebi.ac.uk/pride/data/archive/2022/02/PXD028735/" target="_blank" rel="noopener">here</a>.</div> <br> <div>We upload here the annotated <code>.mgf</code> files created from these <code>.raw</code> files, as described in our paper.</div> <br> <div>The human, yeast, and E. coli .fasta files used in all database searches were downloaded from UniProt on 11/6/23, 4:30 PM.</div> <br> <div> <ul> <li>Bateman, A., Martin, M.-J., Orchard, S., Magrane, M., Ahmad, S., Alpi, E., Bowler-Barnett, E. H., Britto, R., Bye-A-Jee, H., Cukura, A., Denny, P., Dogan, T., Ebenezer, T., Fan, J., Garmiri, P., da Costa Gonzales, L. J., Hatton-Ellis, E., Hussein, A., &hellip; Zhang, J. (2022). UniProt: the Universal Protein Knowledgebase in 2023. In Nucleic Acids Research (Vol. 51, Issue D1, pp. D523&ndash;D531). Oxford University Press (OUP). https://doi.org/10.1093/nar/gkac1052</li> </ul> </div> <br> <div>We include these files here, with only minor modifications to replace `U` amino acids with `X` so that all amino acids fall into Casanovo-DB's vocabulary.</div> </div>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record