Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,505
datasets available to search
ShareScore release 0.7.1
Dataset results
7,505 results for “Generation”
Data associated with the following publication: "Giant thermoelectric response of confined electrolytes with thermally activated charge carrier generation"
<p>Data associated with the following publication: "Giant thermoelectric response of confined electrolytes with thermally activated charge carrier generation" (DOI: <a title="" href="https://doi.org/10.48328/tudatalib-1376">https://doi.org/10.48328/tudatalib-1376</a>)</p>
Network Digital Twin-Generated Dataset for Machine Learning-based Detection of Benign and Malicious Heavy Hitter Flows
<h3>Overview</h3> <p>This record provides a dataset created as part of the study presented in the following publication and is made <strong>publicly available for research purposes</strong>. The associated article provides a comprehensive description of the dataset, its structure, and the methodology used in its creation. If you use this dataset, please <strong>cite the following article </strong>published in the journal <strong>IEEE Communications Magazine</strong>:</p> <blockquote> <p><strong>A. Karamchandani, J. Nunez, L. de-la-Cal, Y. Moreno, A. Mozo, and A. Pastor, “On the Applicability of Network Digital Twins in Generating Synthetic Data for Heavy Hitter Discrimination,” IEEE Communications Magazine, pp. 2–8, 2025, DOI: 10.1109/MCOM.003.2400648.</strong></p> </blockquote> <p>More specifically, the record contains several synthetic datasets generated to differentiate between benign and malicious heavy hitter flows within a realistic virtualized network environment. Heavy Hitter flows, which include high-volume data transfers, can significantly impact network performance, leading to congestion and degraded quality of service. Distinguishing legitimate heavy hitter activity from malicious Distributed Denial-of-Service traffic is critical for network management and security, yet existing datasets lack the granularity needed for training machine learning models to effectively make this distinction.</p> <p>To address this, a Network Digital Twin (NDT) approach was utilized to emulate realistic network conditions and traffic patterns, enabling automated generation of labeled data for both benign and malicious HH flows alongside regular traffic.</p> <h3>Feature Set:</h3> <p>The feature set includes the following flow statistics commonly used in the literature on network traffic classification:</p> <ul> <li>The protocol used for the connection, identifying whether it is TCP, UDP, ICMP, or OSPF.</li> <li>The time (relative to the connection start) of the most recent packet sent from source to destination at the time of each snapshot.</li> <li>The time (relative to the connection start) of the most recent packet sent from destination to source at the time of each snapshot.</li> <li>The cumulative count of data packets sent from source to destination at the time of each snapshot.</li> <li>The cumulative count of data packets sent from destination to source at the time of each snapshot.</li> <li>The cumulative bytes sent from source to destination at the time of each snapshot.</li> <li>The cumulative bytes sent from destination to source at the time of each snapshot.</li> <li>The time difference between the first packet sent from source to destination and the first packet sent from destination to source.</li> </ul> <h3>Dataset Variations:</h3> <p>To accommodate diverse research needs and scenarios, the dataset is provided in the following variations:</p> <ol> <li> <p><strong><code>All at Once</code></strong>:</p> <ol> <li>Contains a synthetic dataset where all traffic types, including benign, normal, and malicious DDoS heavy hitter (HH) flows, are combined into a single dataset.</li> <li>This version represents a holistic view of the traffic environment, simulating real-world scenarios where all traffic occurs simultaneously.</li> </ol> </li> <li> <p><strong><code>Balanced Traffic Generation</code></strong>:</p> <ol> <li>Represents a balanced traffic dataset with an equal proportion of benign, normal, and malicious DDoS traffic.</li> <li>Designed for scenarios where a balanced dataset is needed for fair training and evaluation of machine learning models.</li> </ol> </li> <li> <p><strong><code>DDoS at Intervals</code></strong>:</p> <ol> <li>Contains traffic data where malicious DDoS HH traffic occurs at specific time intervals, mimicking real-world attack patterns.</li> <li>Useful for studying the impact and detection of intermittent malicious activities.</li> </ol> </li> <li> <p><strong><code>Only Benign HH Traffic</code></strong>:</p> <ol> <li>Includes only benign HH traffic flows.</li> <li>Suitable for training and evaluating models to identify and differentiate benign heavy hitter traffic patterns.</li> </ol> </li> <li> <p><strong><code>Only DDoS Traffic</code></strong>:</p> <ol> <li>Contains only malicious DDoS HH traffic.</li> <li>Helps in isolating and analyzing attack characteristics for targeted threat detection.</li> </ol> </li> <li> <p><strong><code>Only Normal Traffic</code></strong>:</p> <ol> <li>Comprises only regular, non-HH traffic flows.</li> <li>Useful for understanding baseline network behavior in the absence of heavy hitters.</li> </ol> </li> <li> <p><strong><code>Unbalanced Traffic Generation</code></strong>:</p> <ol> <li>Features an unbalanced dataset with varying proportions of benign, normal, and malicious traffic.</li> <li>Simulates real-world scenarios where certain types of traffic dominate, providing insights into model performance in unbalanced conditions.</li> </ol> </li> </ol> <p>For each variation, the output of the different packet aggregators is provided separated in its respective folder.</p> <p>Each variation was generated using the NDT approach to demonstrate its flexibility and ensure the reproducibility of our study's experiments, while also contributing to future research on network traffic patterns and the detection and classification of heavy hitter traffic flows. The dataset is designed to support research in network security, machine learning model development, and applications of digital twin technology.</p>
Synthetic time series data generation for edge analytics
<p>In this research, we create synthetic data with features that are like data from IoT devices. We use an existing air quality dataset that includes temperature and gas sensor measurements. This real-time dataset includes component values for the Air Quality Index (AQI) and ppm concentrations for various polluting gas concentrations. We build a JavaScript Object Notation (JSON) model to capture the distribution of variables and structure of this real dataset to generate the synthetic data. Based on the synthetic dataset and original dataset, we create a comparative predictive model. Analysis of synthetic dataset predictive model shows that it can be successfully used for edge analytics purposes, replacing real-world datasets. There is no significant difference between the real-world dataset compared the synthetic dataset. The generated synthetic data requires no modification to suit the edge computing requirements. The framework can generate correct synthetic datasets based on JSON schema attributes. The accuracy, precision, and recall values for the real and synthetic datasets indicate that the logistic regression model is capable of successfully classifying data</p>
What does it take to generate new growth - Survey data on company perceptions on innovative behavior
<p>This data includes raw survey data, a codebook and the survey form for the survey <em>what does it take to generate new growth? </em>The survey focused on comprehensively mapping the Finnish companies growth outlooks and their underlying management practices and principles. The study creates an overview of top managers’ views on Finnish companies’ growth, innovativeness, and the ability for renewal. It allows us to identify what sets high-growing companies apart from others. The Codebook is associated with an SPSS and CSV file including the data.</p>
MUHAI Benchmark : Task 2 (Credibility of knowledge-based generated gossip stories)
<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 2 (Credibility of knowledge-based generated gossip stories)</strong></p> <p> </p> <p>This dataset aims at investigating whether the use of Knowledge Graphs has an impact on the credibililty of automatically-generated stories.</p> <p><br> The submission includes the following data:</p> <ol> <li>Generated stories (.txt)</li> <li>Story generation template </li> <li>A tsv file with entities and triples (to be used for generating stories)</li> <li>Evaluation description : Questions and metrics submitted to the users</li> </ol> <p>The "gossip stories" are generated with the T5 languge model fine-tuned on the WebNLG challenge. The model takes the triples (file 3) as input and generates one sentence each. A link prediction algorithm based on Jaccard's similarity learns the likelihood of two entities to be related (3). Then, the narrative continues with automatically generated celebrity background descriptions.</p> <p>The credibility of the story is evaluated using a questionnaire based on Gaziano et. al. The questionnaire was filled in by the test subjects after reading each generated article. One for a KG-generated text where links were predicted using the link prediction and one for text that was generated using triples of random entities (celebrities). </p> <p>Full code available at : https://github.com/kmitd/muhai-credibility-KR</p>
Data and code used in manuscript: Basal freeze-on generates complex ice-sheet stratigraphy
<p>Mapped plumes location obtained from ice-sheet radio echo sounding data of North Greenland (https://data.cresis.ku.edu/data/rds/ for 2010-2014_Greenland files) and map of calculated freeze-on index are found in 'FreezeOnIndex_MappedPlume_Data.nc'. Model code of the three models used to obtain the findings shown in the manuscript 'Basal freeze-on generates complex ice-sheet stratigraphy'. As well as code to calculate the freeze-on index.</p>
Testing absolute plate reference frames and the implications for the generation of geodynamic mantle heterogeneity structure
<div>Description of Resources - Shephard et al. (2012)</div> <div> </div> <div>This file provides a detailed description of all of the files that make up the data collection associated with the publication: Shephard, G. E., Bunge, H. P., Schuberth, B. S., Müller, R. D., Talsma, A. S., Moder, C., & Landgrebe, T. C. W. (2012). Testing absolute plate reference frames and the implications for the generation of geodynamic mantle heterogeneity structure. Earth and Planetary Science Letters, 317, 204-217. doi: <a href="https://doi.org/10.1016/j.epsl.2011.11.027" target="_blank" rel="noopener">10.1016/j.epsl.2011.11.027</a></div> <div> </div> <div>Note: For information on file formats and what programs to use to interact with various file formats, see "File Formats and Recommended Programs”.</div> <div> </div> <div>This data collection includes both the rotations and topologically closed polygons* for each of the 5 absolute reference frames that were tested in the publication. They are to be loaded in GPlates (<a href="http://www.gplates.org" target="_blank" rel="noopener">http://www.gplates.org</a>).</div> <div> </div> <div>*Topologically closed plate polygons are constructed from the intersection of ridges, transforms, subduction zones and other plate boundary geometries. These 'resolved topologies' are valid at 1 Myr intervals. The plate boundary geometries and plate polygons have been assigned plate reconstruction IDs to allow them to be reconstructed using the supplied rotation files. </div> <div> </div> <div>The files associated with this data collection include:</div> <div>• <strong>Hybrid hotspot model (Moving and Fixed hotspots) (HHS)</strong></div> <div>* Caltech_Global_20110311HHS.gpml (37 MB) - topologically closed plate polygons and plate boundary geometries</div> <div>* Caltech_Global_20110412HHS.rot (287 KB)- global rotation model</div> <div> </div> <div>• <strong>Fixed hotspot model (FHS)</strong></div> <div>* Caltech_Global_20110311FHS.gpml (36.8 MB) - topologically closed plate polygons and plate boundary geometries</div> <div>* Caltech_Global_20110412FHS.rot (291 KB) - global rotation model</div> <div> </div> <div>•<strong> Hybrid hotspot and palaeomagnetic model (PMG)</strong></div> <div>* Caltech_Global_20110311PMG.gpml (35.9 MB) - topologically closed plate polygons and plate boundary geometries</div> <div>* Caltech_Global_20110412PMG.rot (287 KB) - global rotation model</div> <div> </div> <div>• <strong>Subduction reference frame model (SUB)</strong></div> <div>* Caltech_Global_20110311SUB.gpml (36 MB) - topologically closed plate polygons and plate boundary geometries</div> <div>* Caltech_Global_20110412SUB.rot (287 KB) - global rotation model</div> <div> </div> <div>• <strong>Hybrid hotspot and TPW-corrected palaeomagnetic model (TPW)</strong></div> <div>* Caltech_Global_20110311TPW.gpml (36.8 MB) - topologically closed plate polygons and plate boundary geometries</div> <div>* Caltech_Global_20110412TPW.rot (287 KB)- global rotation model</div> <div> </div> <div>Project files (.gproj) are included for each .gpml/.rot pair.</div> <div> </div> <div>This article has additional supplementary data available with the online publication.</div> <div> </div> <div> </div> <div>Additional notes:</div> <div>*.rot contains the rotations for all plates and topological polygons.</div> <div>Each model is specific according to the African Plate (Plate ID 701) rotations. The rotations for all other plates are the same across each of the five models with the exception of cross-overs involving Pacific/Panthalassa plates for times earlier than 83.5Ma; these must be absolute reference frame specific and were re-calculated for each model. Programs used to calculate the new finite rotations include "adder" and "seaflow" </div> <div> </div> <div>*.gpml and .shp files contain continuously closing plate polygons i.e. from plate boundaries, from 140 Ma to present-day in 1 million year increments. </div> <div>These files differ slightly from those used in the paper, but are the most up-to-date version (as at May 2011) and are based on an updated model, Seton et al. (2012).</div> <div>They are specific to each of the five absolute reference frames. </div> <div> </div> <div>Note on velocity calculations in GPlates:</div> <div>GPlates calculates the velocity within each plate based on the stage rotation for that time period and averages for that respective period. For this reason, the velocities of a plate do not change incrementally within the time period and then abruptly change according to the next time period/stage rotation. </div> <div>This is also why there appears to be a "jump" in velocity magnitude and direction between 140 and 139 Ma.</div>
Experimental data generated on the stability of hydrophobic porous materials
<div>/* **********</div> <div>/* This work is licensed under a Creative Commons Attribution 4.0 International License.</div> <div>/* **********</div> <div> </div> <div>Open access to experimental data generated by the project Electro-Intrusion (101017858, Horizon 2020, European Union, https://www.electro-intrusion.eu/en) along with the research to be used in intrusion-extrusion applications. Research pertaining to Task 2.1 (WP2). </div> <div>Underlying data for the publication Amayuelas, E. et al. Bimetallic Zeolitic Imidazole Frameworks for Improved Stability and Performance of Intrusion-Extrusion Energy Applications. The Journal of Physical Chemistry 2023, 127, 18310-18315. https://doi.org/10.1021/acs.jpcc.3c04368. Data related to Figures 2, 3 and 4 in the article.</div> <div> </div> <div>Dataset Identifier: 10.5281/zenodo.11273904</div> <div> </div> <div>Contact person: Eder Amayuelas (CIC energiGUNE). ORCID: </div> <div> </div> <div> </div> <div>The archive 'JPCC_3c04368.zip' contains 25 files:</div> <p> </p>
RESOLUTE atlas for brain PET/MR pseudo-CT generation
<p>Template and mask images for performing the <em>Region specific optimization of continuous linear attenuation coefficients based on UTE</em> (RESOLUTE) pseudo-CT generation approach. This dataset can be used in conjunction with an open-source C++ implementation of RESOLUTE (<a href="https://github.com/UCL/petmr-RESOLUTE">https://github.com/UCL/petmr-RESOLUTE</a>) for the Siemens mMR scanner.</p>
Third harmonic generation images of the lacuno-canalicular network in bone femoral diaphysis of mice from the BionM1 project (space flight)
<p>Data set for 11 samples in 3 groups of Control, Space Flight and Synchro (ground control with space flight housing and feeding conditions). Contains THG images in tif format of 2D mosaic of selected samples and 3D stacks in selected anatomical regions of interest. See readme file for more information.</p>
Generated WSP: Validation of a water-sensitive paper-based method for the characterization of agricultural spray droplets
<p>Synthetic images were generated in a Python environment using the OpenCV library to replicate the distribution of droplets in WSP. The images display droplet stains represented by blue circles (255,0,0) on a yellow background (0,255,255) to enhance contrast and enable more precise analysis. The synthetic images were created in two distinct resolutions, namely 640x480 and 2560x1440 pixels, with the aim of reproducing the output of two specific digital microscopes: the Jiusion 640x480 and the Jiusion HD 2560x1440 (Shenzen, China). The resolution is chosen based on the expected practical application, ensuring that any image analysis algorithm developed can effectively process images with similar characteristics to those obtained under real conditions by these microscopes. Each pixel in this configuration corresponds to a physical size of 18.125 µm in images with a resolution of 640x480, and a size of 6.875 µm in images with a resolution of 2560x1440. Multiple patterns were created to simulate various configurations of droplet stains in WSP. The sizes of single droplet stains varied between 100 and 600 µm, with spacings of either 1000 µm or 2000 µm between drops (see attached figure). Furthermore, the same size range was utilised to generate patterns with double and overlaid droplet stains, with a consistent spacing of 2800 µm between each stain (see attached figure). The implementation of this systematic method guarantees the accurate calibration and application of image analysis algorithms in real-world situations. This allows for the representation of precise measurements and spacing that would be encountered in actual experimental conditions.</p>
LCZ-Generator Training Areas
<p>This dataset contains all training areas (TA) submitted to the <a href="https://lcz-generator.rub.de/">LCZ Generator</a> (<a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a>) since 2021-04-14. The LCZ Generator follows a crowdsourcing approach, making fast and easy LCZ-mapping available to the public, while collecting LCZ maps and TAs in a centralized, easy to access, location. The crowdsoucing approach overcomes the limitations of previous approaches where a manual review was mandatory before publication. While this improved the quality of individual LCZ-maps, the number of cities mapped during this period remained low. The LCZ Generator removed the manual review process, allowing for faster collection of LCZ maps and TAs, however, sacrificing some quality since any person can submit to the LCZ Generator without prior training or review.</p> <p>This dataset is based on crowdsourcing, hence LCZs may be mislabelled, polygon shapes may not be perfect etc. also city names may not be correct. Some contributors chose to name their city e.g. "..", "....amsa" etc. also some author names may not be correct. The automated quality control (see table below and section 2.3 in <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a>) may help filter out some of the incorrect TAs. We intentionally included all available TAs to allow for (the development of) custom filtering.</p> <p>The data was extracted from the LCZ-Generator database taking into account:</p> <ol> <li>Whether or not the submitting author agreed to show their name (if not, it is also left blank in this dataset)</li> <li>The license the TA was submitted under (a license change happened with version 2.0.0 of the LCZ Generator)</li> <li>Duplicate geometries were dropped, since multiple (re-)submission may have the same geometries. Only the <strong>first</strong> submitted version is kept and attributed to the <code>submission_id</code> of the first submission.</li> </ol> <p>The data, up to December 2021, was used during creation of the global LCZ Map (<a href="https://doi.org/10.5194/essd-14-3835-2022">Demuzere et al. 2022</a>).</p> <p>Additional TAs were extracted from the <a href="https://wudapt.cs.purdue.edu">WUDAPT Portal</a> and processed using the LCZ-Generator.</p> <p><strong>Note</strong>: The data is updated periodically, but not on a fixed schedule.</p> <h2>Data Description</h2> <p>The data is provided as GeoPackage (<code>.gkpg</code>) which can be used with most GIS.</p> <table> <tbody> <tr> <th>Column Name</th> <th>Description</th> </tr> </tbody> <tbody> <tr> <td><code>geometry</code></td> <td>The polygon geometry of the training area (TA) in EPSG:4326</td> </tr> <tr> <td><code>submission_id</code></td> <td>The ID of the corresponding submission in the <a href="https://lcz-generator.rub.de/">LCZ Generator</a></td> </tr> <tr> <td><code>submission_date</code></td> <td>The date and time in <strong>UTC</strong> the TA was submitted to the LCZ Generator</td> </tr> <tr> <td><code>city</code></td> <td>The city the TAs are for. Note: This is sometimes incorrect due to users entering incorrect information and the LCZ Generator following a crowdsourcing approach.</td> </tr> <tr> <td><code>reference</code></td> <td>The submitting author may have provided (additional) references via this field. This can be a scientific paper or a citation of the original creator of this TA</td> </tr> <tr> <td><code>remarks</code></td> <td>General information: e.g. co-authors, information about the study/framework the TAs were generated for</td> </tr> <tr> <td><code>representative_date</code></td> <td>The date the TAs are representative for (i.e. the date the aerial image was taken)</td> </tr> <tr> <td><code>firstname</code></td> <td>First name of the submitting author (if the author did not agree to publish their name, this is left blank and the submission is treated anonymously)</td> </tr> <tr> <td><code>lastname</code></td> <td>Last name of the submitting author (if the author did not agree to publish their name, this is left blank and the submission is treated anonymously)</td> </tr> <tr> <td><code>license</code></td> <td>The license this specific polygon is licensed under. With version 2.0.0 of the LCZ Generator the license was changed from CC BY-SA to CC BY-NC-SA 4.0</td> </tr> <tr> <td><code>cite_as</code></td> <td>A suggestion how to cite the TA (-set) based on the name, year, and city information. If the author submitted anonymously, this is left blank</td> </tr> <tr> <td><code>version</code></td> <td>The version of the LCZ Generator the polygon was submitted to. Detailed information can be found in the <a href="https://github.com/RUBclim/LCZ-Generator-Issues?tab=readme-ov-file#changelog">Changelog</a></td> </tr> <tr> <td><code>class</code></td> <td>The LCZ Class the TA-polygon has been labelled (1 - 17)</td> </tr> <tr> <td><code>area</code></td> <td>The area of the TA-polygon in km<sup>2</sup></td> </tr> <tr> <td><code>perimeter</code></td> <td>The perimeter of the TA-polygon km</td> </tr> <tr> <td><code>shape</code></td> <td>The shape of the TA-polygon calculated as: (perimeter<sup>2</sup>) / (4 π · area)</td> </tr> <tr> <td><code>vertices</code></td> <td>The number of vertices of the TA-polygon</td> </tr> <tr> <td><code>qc_step1</code></td> <td>Whether the TA-polygon passed the automated quality control (QC) step 1: Surface area below 0.04 km<sup>2</sup> (too small) or a shape ratio 3 (too complex shape) are flagged. More information about the QC can be found in the <a href="https://lcz-generator.rub.de/faq#why-suspicious-tas">FAQ</a> and the corresponding paper <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a></td> </tr> <tr> <td><code>qc_step2</code></td> <td>Whether the TA-polygon passed the automated quality control (QC) step 2: Average spectral value of a polygon of LCZ class is considered as an outlier compared to the average spectral values of all other polygons of that class. Note that this is done on a per-submission basis More information about the QC can be found in the <a href="https://lcz-generator.rub.de/faq#why-suspicious-tas">FAQ</a> and the corresponding paper <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a></td> </tr> <tr> <td><code>qc_step3</code></td> <td>Whether the TA-polygon passed the automated quality control (QC) step 3: Considers all individual pixel values of all polygons in each LCZ class compared to the polygon average approach from QC Step 2. More information about the QC can be found in the <a href="https://lcz-generator.rub.de/faq#why-suspicious-tas">FAQ</a> and the corresponding paper <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a></td> </tr> <tr> <td><code>oa</code></td> <td>The overall accuracy of the submission based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>oau</code></td> <td>The overall accuracy for the urban LCZ classes only of the submission based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>oabu</code></td> <td>The overall accuracy of the built versus natural LCZ classes only of the submission based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>oaw</code></td> <td>A weighted accuracy taking the similarities of LCZs into account based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>f1_1</code></td> <td>Class-wise metric F1 for LCZ class 1</td> </tr> <tr> <td><code>f1_2</code></td> <td>Class-wise metric F1 for LCZ class 2</td> </tr> <tr> <td><code>f1_3</code></td> <td>Class-wise metric F1 for LCZ class 3</td> </tr> <tr> <td><code>f1_4</code></td> <td>Class-wise metric F1 for LCZ class 4</td> </tr> <tr> <td><code>f1_5</code></td> <td>Class-wise metric F1 for LCZ class 5</td> </tr> <tr> <td><code>f1_6</code></td> <td>Class-wise metric F1 for LCZ class 6</td> </tr> <tr> <td><code>f1_7</code></td> <td>Class-wise metric F1 for LCZ class 7</td> </tr> <tr> <td><code>f1_8</code></td> <td>Class-wise metric F1 for LCZ class 8</td> </tr> <tr> <td><code>f1_9</code></td> <td>Class-wise metric F1 for LCZ class 9</td> </tr> <tr> <td><code>f1_10</code></td> <td>Class-wise metric F1 for LCZ class 10</td> </tr> <tr> <td><code>f1_11</code></td> <td>Class-wise metric F1 for LCZ class 11</td> </tr> <tr> <td><code>f1_12</code></td> <td>Class-wise metric F1 for LCZ class 12</td> </tr> <tr> <td><code>f1_13</code></td> <td>Class-wise metric F1 for LCZ class 13</td> </tr> <tr> <td><code>f1_14</code></td> <td>Class-wise metric F1 for LCZ class 14</td> </tr> <tr> <td><code>f1_15</code></td> <td>Class-wise metric F1 for LCZ class 15</td> </tr> <tr> <td><code>f1_16</code></td> <td>Class-wise metric F1 for LCZ class 16</td> </tr> <tr> <td><code>f1_17</code></td> <td>Class-wise metric F1 for LCZ class 17</td> </tr> </tbody> </table> <h2>Acknowledgements</h2> <p>We acknowledge all WUDAPT contributors and community members for providing the training areas via the LCZ Generator.</p>
Dataset for the publication "Implementation of an exact completing method of generation for face-milled spiral bevel gears with uniform depth taper"
<p>This dataset contains geometric and graphics data associated with the referenced paper, enabling the reproduction of the conducted research. </p>
SignAture_Electricity_generation_data_compare_Latvia_2020_2022
<p>This dataset, related to the article 'Power System Modelling in the Baltic Countries: Data Accessibility and Consistency Aspects' (2023), compares electricity generation data for 2020 and 2022 from various sources in Latvia, providing both input and output values and associated metadata.</p>
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
An analysis of meiofauna knowledge generated by Latin American researchers
<p>Bibliographic databases used to analyse the document production of benthic meiofauna in Latin American countries. To be opened on R, bibliometrix package.</p> <p> </p>
Dataset generated to evaluate in situ sampling strategies to reconstruct fine-scale ocean currents in the context of SWOT satellite mission (H2020 EuroSea project)
<p><strong>Dataset generated in Subtask 2.3.1 of the H2020 EuroSea project.</strong></p> <ul> <li> <p><em>H2020 EuroSea project:</em><br> The H2020 EuroSea project aims at improving and integrating the European Ocean Observing and Forecasting System (see official website: <a href="https://eurosea.eu/">https://eurosea.eu/</a>). It has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 862626).</p> </li> <li> <p><em>Task 2.3:</em><br> Task 2.3 has the objective to improve the design of multi-platform experiments aimed to validate the Surface Water and Ocean Topography (SWOT) satellite observations with the goal to optimize the utility of these observing platforms. Observing System Simulation Experiments (OSSEs) have been conducted to evaluate different configurations of the in situ observing system, including rosette and underway CTD, gliders, conventional satellite nadir altimetry and velocities from drifters. High-resolution models have been used to simulate the observations and to represent the “ocean truth”. Several methods of reconstruction have been tested: spatio-temporal optimal interpolation, machine-learning techniques, model data assimilation and the MIOST tool. The planned OSSEs are detailed in this public report <a href="https://doi.org/10.3289/eurosea_d2.1">Barceló-Llull et al. (2020)</a> and the complete analysis is available here <a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al. (2022)</a>. Contributors to Task 2.3 are CSIC (Spain), CLS (France), SOCIB (Spain), IMT-Atlantique (France) and Ocean-Next (France).</p> </li> <li> <p><em>Subtask 2.3.1:</em><br> Subtask 2.3.1 aims to evaluate different in situ sampling strategies to reconstruct fine-scale ocean currents (~20 km) in the context of SWOT. An advanced version of the classic optimal interpolation used in field experiments, which considers the spatial and temporal variability of the observations, has been applied to reconstruct different configurations with the objective to evaluate the best sampling strategy to validate SWOT.</p> </li> <li> <p><em>Where?</em><br> The analysis focuses on two regions of interest: (i) the western Mediterranean Sea and (ii) the Subpolar North West Atlantic. In the western Mediterranean Sea, the target area is located within a swath of SWOT, while in the North West Atlantic the region of study includes a crossover of SWOT during the fast-sampling phase.</p> </li> </ul> <p><strong>Report with the full analysis</strong></p> <p>The complete analysis can be found in this report: <a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al. (2022)</a>.</p> <p><strong>Codes for the analysis</strong></p> <p>The codes generated to develop Subtask 2.3.1 can be found on GitHub: <a href="https://github.com/bbarcelollull/EuroSea_subTask_2.3.1">https://github.com/bbarcelollull/EuroSea_subTask_2.3.1</a></p> <p><strong>The dataset</strong></p> <p>The dataset includes:</p> <p>1) Model outputs used to simulate the observations in different configurations in both regions of study. The folder "2D_model_outputs" contains 2D data used to simulate SSH observations for the analysis of the temporal correlation scale (<a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al., 2022</a>, p. 28-42). The folder "3D_model_outputs" contains 3D model outputs used to simulate observations of temperature and salinity. Note that eNATL60 outputs have been interpolated onto a new regular grid. </p> <p>2) Simulated configurations (or sampling strategies) in each region (PKL file format).</p> <p>3) Observations simulated in each configuration in both regions of study. The observations simulated are temperature and salinity. ADCP horizontal velocities are also simulated, however for eNATL60 they will be corrected in the future to account for the rotated original axes. File format: region_configuration_period_model.nc. The folder "SSH" includes the simulated SSH observations for the analysis of the temporal correlation scale (<a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al., 2022</a>, p. 28-42).</p> <p>4) Reconstructed fields with the spatio-temporal optimal interpolation. File format: region_configuration_period_model_stOI_Lx_Lt_cd_YYYYMMDDhhmm_var.nc (stOI = spatio-temporal optimal interpolation, Lx = spatial correlation scale, Lt = temporal correlation scale, cd = map on the central date of the sampling, YYYYMMDDhhmm = date and time of the map, var = variable interpolated (temperature and salinity) or the derived variables (dynamic height, geostrophic velocities and the Rossby number)).</p> <p>5) Compared fields (ocean truth from model outputs vs. reconstructed fields) for each region and model (PKL file format).</p> <p> </p>
Dataset for Accessing Cosmic Radiation as an Entropy Source for a Non-Deterministic Random Number Generator
<p>The dataset contains all gathered data from the experiment from Wednesday, March 16, 2022 11:58:41.929 AM UTC+0 (1647431921929) until Sunday, April 3, 2022 1:08:35.353 PM UTC+0 (1648991315353). The experiment was executed during physical presence within the Arctic Circle in Tromsø, Norway 69° 40' 53.117'' N 18° 58' 36.027'' E at 35m elevation above sea level. The dataset was gathered with a prototype [1] based on the CREDO android application [2]. The main research is to use Ultra High Energy Cosmic Rays (UHECR) as an entropy source for a Random Bit Generator (RBG). </p> <p>The associated publication will probably have the title "Accessing Cosmic Radiation as an Entropy Source for a Non-Deterministic Random Number Generator"</p> <p>In order to reproduce the results the SQLite3 database "mrng_arctic_experiment_2022.db" is needed. To get the visual representations of the detections use "image_decoding_and_codesnippets.py" to generate the cleaned (414 detections / ~15MB) or the uncleaned (5567 detections / ~195 MB) dataset. The compressed folder "raw_data_incl_space_weather.7z" contains all raw data as gathered with the MRNG prototype, unprocessed, uncleaned, and unmerged. </p> <p> </p> <p>[1] https://github.com/StefanKutschera/mrng-prototype, visited on 27.03.2023</p> <p>[2] https://github.com/credo-science/credo-detector-android, visited on 27.03.2023</p>
Hybrid quantum-classical machine learning for generative chemistry and drug design: Generated molecules
<p>Deep generative chemistry models emerge as powerful tools to expedite drug discovery. How- ever, the immense size and complexity of the structural space of all possible drug-like molecules pose significant obstacles, which could be overcome with hybrid architectures combining quantum computers with deep classical networks. As the first step toward this goal, we built a compact discrete variational autoencoder (DVAE) with a Restricted Boltzmann Machine (RBM) of reduced size in its latent layer. The size of the proposed model was small enough to fit on a state-of-the-art D-Wave quantum annealer and allowed training on a subset of the ChEMBL dataset of biologically active compounds. Finally, we generated 2331 novel chemical structures with medicinal chemistry and synthetic accessibility properties in the ranges typical for molecules from ChEMBL. The pre- sented results demonstrate the feasibility of using already existing or soon-to-be-available quantum computing devices as testbeds for future drug discovery applications.</p>
Data for Figures 4, A-F and Table J of Publication "Blue skies over China: The effect of pollution-control on solar power generation and revenues"
<p>This repository contains the data to produce Figures 4, A-F and Table J and emission data in the paper:</p> <p>"Labordena M, Neubauer D, Folini D, Patt A, Lilliestam J (2018) Blue skies over China: The effect of pollution-control on solar power generation and revenues. PLoS ONE 13(11): e0207028. https://doi.org/10.1371/journal.pone.0207028"</p> <p>Note that the scripts are to be found in the accompanying package (https://doi.org/10.5281/zenodo.8130726)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.