Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Mars surface image (Curiosity rover) labeled data set
<p>This data set consists of 6691 images spanning 24 classes that were collected by the Mars Science Laboratory (MSL, Curosity) rover by three instruments (Mastcam Right eye, Mastcam Left eye, and MAHLI). These images are the "browse" version of each original data product, not full resolution. They are roughly 256x256 pixels each.</p> <p>We divided the MSL images into train, validation, and test data sets according to their sol (Martian day) of acquisition. This strategy was chosen to model how the system will be used operationally with an image archive that grows over time. The images were collected from sols 3 to 1060 (August 2012 to July 2015). The exact train/validation/test splits are given in individual files. Full-size images can be obtained from the PDS at https://pds-imaging.jpl.nasa.gov/search/ .</p> <p><strong>Contents</strong>:</p> <ul> <li>calibrated/: Directory containing calibrated MSL images</li> <li>train-calibrated-shuffled.txt: Training labels (images in shuffled order)</li> <li>val-calibrated-shuffled.txt: Validation labels</li> <li>test-calibrated-shuffled.txt: Test labels</li> <li>msl_synset_words-indexed.txt: Mapping from class IDs to class names</li> </ul> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1049137</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, You Lu, Alice Stanboli, Kevin Grimes, Thamme Gowda, and Jordan Padams. "Deep Mars: CNN Classification of Mars Imagery for the PDS Imaging Atlas." <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>
Data set for the paper: Intercomparison of ocean colour algorithms for picophytoplankton carbon in the ocean
<p>This dataset contains the phytoplankton carbon,Cphy, obtained from in situ counts of phytoplankton cells using ow cytometry presented in the paper [13]. The location and time of the samples have been matched with the satelllite data in the Ocean Colour Climate Change Initiative (OCCCI) dataset. This dataset is the match between the in situ Cphy and the products from using the OCCCI inputs (i.e. chlorophyll concentration, backscattering coecient, phytoplankton absorption) with 6 different algorithms. This document describes the dataset details: data sources, computation of Cphy, selected data.</p>
Data set for "Reward-based learning drives rapid sensory signals in medial prefrontal cortex and dorsal hippocampus necessary for goal-directed behavior"
<p>Data set for: Le Merre P, Esmaeili V, Charrière E, Galan K, Salin P-A, Petersen CCH, Crochet S (2018) Reward-based learning drives rapid sensory signals in medial prefrontal cortex and dorsal hippocampus necessary for goal-directed behavior. Neuron, https://doi.org/10.1016/j.neuron.2017.11.031</p> <p>There are 44 files in this data upload:<br> 1. '2018_LeMerre_Neuron.pdf' - this is a pdf version of the online publication.<br> 2. 'Chronic_LFP_data.mat' - this is a Matlab data structure, which contains all the chronic LFP data for the publication.<br> 3. 'Silicon_Probe_data.mat' - this is a Matlab data structure, which contains all the mPFC silicon probe recording data for the publication.<br> 4. 'Opto_Inactivation_data.mat' - this is a Matlab data structure, which contains all the optogenetic inactivation data for the publication.<br> 5. 'Mus_Inactivation_data.mat' - this is a Matlab data structure, which contains all the pharmacological (Muscimol) inactivation data for the publication.<br> 6. 'Learning_Days_Mtrx.mat' - this is a Matlab data file, which contains the selected training days analyzed for the Trained condition in the Detection Task.<br> 7. 'Exposed_Days_Mtrx.mat' - this is a Matlab data file, which contains the selected days analyzed for the Exposed condition in the Neutral Exposure.<br> 8. 'p_value_colormap.mat' - this is a Matlab data file, which contains the color map used to display the p value in the Matlab codes 'plot_fig2A_SEP_D1_vs_Trained.m'; 'plot_fig2B_Amplitude_D1_vs_Trained.m'; 'plot_fig3A_SEP_D1_vs_Exposed.m'; 'plot_fig4A_SEP_H_vs_M.m’.<br> 9. 'p_value_colormap2.mat' - this is a Matlab data file, which contains the color map used to display the p value in the Matlab code 'plot_figS3B_Stim_vs_Catch_for_significantly_inc_dec_units.m’; ’plot_figS4A_H_vs_M_for_inc_dec_units_and_zscored_PSTH.m’.<br> 10. 'scatterplot_colormap.mat' - this is a Matlab data file, which contains the color map used to display the p value in the Matlab code 'plot_fig2C_Scatterplot_Amplitude_vs_dprime.m'.<br> 11. 'SEP_colormtrx.mat' - this is a Matlab data file, which contains the color map used to display the p value in the Matlab code 'plot_fig1B_Sensory_Evoked_Potentials.m'; 'plot_figS3A_SEP_EMG_amplitude_ReactionTime.m'.<br> 12. 'zscore_colormap.mat' - this is a Matlab data file, which contains the color map used to display the p value in the Matlab code 'plot_fig3D_mPFC_PSTH_and_zscore_DT_vs_NE.m'; 'plot_figS4A_H_vs_M_for_inc_dec_units_and_zscored_PSTH.m’.<br> 13. 'Chronic_LFP_dataViewer.fig' - this is a Matlab Figure file, which is the GUI layout for 'Chronic_LFP_dataViewer.m'.<br> 14. 'Chronic_LFP_dataViewer.m' - this is a Matlab code, which displays the data contained in 'Chronic_LFP_data.mat'.<br> 15. 'Silicon_Probe_dataViewer.fig' - this is a Matlab Figure file, which is the GUI layout for 'Silicon_Probe_dataViewer.m'.<br> 16. 'Silicon_Probe_dataViewer.m' - this is a Matlab code, which displays the data contained in 'Silicon_Probe_data.mat'.<br> 17. 'plot_fig1B_Sensory_Evoked_Potentials.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results published in figure 1, panel B (Le Merre et al., 2018).<br> 18. 'plot_fig1C_Silicon_Probe_Hit_trials.m' - this is a Matlab code, which analyses the data in 'Silicon_Probe_data.mat', and displays the results in the same way as the published figure 1, panel C (Le Merre et al., 2018).<br> 19. 'plot_fig2A_SEP_D1_vs_Trained.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 2, panel A (Le Merre et al., 2018).<br> 20. 'plot_fig2B_Amplitude_D1_vs_Trained.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 2, panel B (Le Merre et al., 2018).<br> 21. 'plot_fig2C_Scatterplot_Amplitude_vs_dprime.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 2, panel C (Le Merre et al., 2018).<br> 22. 'plot_fig3A_SEP_D1_vs_Exposed.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 3, panel A (Le Merre et al., 2018).<br> 23. 'plot_fig3B_Amplitude_D1_vs_Exposed.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 3, panel B (Le Merre et al., 2018).<br> 24. 'plot_fig3C_ROC_Trained_vs_Exposed.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the ROCs in the same way as the published figure 3, panel C (Le Merre et al., 2018).<br> 25. 'plot_fig3C_ROC_Randomization.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the label shuffled ROCs in the same way as the published figure 3, panel C (Le Merre et al., 2018).<br> 26. 'plot_fig3D_mPFC_PSTH_and_zscore_DT_vs_NE.m' - this is a Matlab code, which analyses the data in 'Silicon_Probe_data.mat', and displays the results in the same way as the published figure 3, panel D (Le Merre et al., 2018).<br> 27. 'plot_fig4A_SEP_H_vs_M.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 4, panel A (Le Merre et al., 2018).<br> 28. 'plot_fig4B_Amplitude_ H_vs_M.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure 4, panel B (Le Merre et al., 2018).<br> 29. 'plot_fig4C_mPFC_PSTH_Hit_vs_Miss.m' - this is a Matlab code, which analyses the data in 'Silicon_Probe_data.mat', and displays the results in the same way as the published figure 4, panel C, left panel (Le Merre et al., 2018).<br> 30. 'plot_fig4C_Scatterplot_modulation_Hit_vs_Miss.m' - this is a Matlab code, which analyses the data in 'Silicon_Probe_data.mat', and displays the results in the same way as the published figure 4, panel C, right panel (Le Merre et al., 2018).<br> 31. 'plot_fig4D_Photoinhibitions.m' - this is a Matlab code, which analyses the data in 'Opto_Inactivation_data.mat', and displays the results published in figure 4, panel D (Le Merre et al., 2018).<br> 32. 'plot_figS2D_Performance_DetectionTask_NeutralExposition.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure S2, panel D (Le Merre et al., 2018).<br> 33. 'plot_figS3A_SEP_EMG_amplitude_ReactionTime.m' - this is a Matlab code, which analyses the data in 'Chronic_LFP_data.mat', and displays the results in the same way as the published figure S3, panel A (Le Merre et al., 2018).<br> 34. 'plot_figS3B_Stim_vs_Catch_for_significantly_inc_dec_units.m' - this is a Matlab code, which analyses the data in 'Silicon_Probe_data.mat', and displays the results in the same way as the published figure S3, panel B (Le Merre et al., 2018).<br> 35. 'plot_figS4A_H_vs_M_for_inc_dec_units_and_zscored_PSTH.m' - this is a Matlab code, which analyses the data in 'Silicon_Probe_data.mat', and displays the results in the same way as the published figure S4, panel A (Le Merre et al., 2018).<br> 36. 'plot_figS4B_Pharmacological_Inactivations.m' - this is a Matlab code, which analyses the data in 'Mus_Inactivation_data.mat', and displays the results published in figure S4 (Le Merre et al., 2018).<br> 37. 'Load_LFP_Multisite_database.m' - this is a Matlab code, which is called in the Matlab codes that analyze the data in 'Chronic_LFP_data.mat'.<br> 38. 'Load_Silicon_Probe_database.m' - this is a Matlab code, which is called in the Matlab codes that analyze the data in 'Silicon_Probe_data.mat'.<br> 39. 'Load_Optogenetic_Inactivation_database.m' - this is a Matlab code, which is called in the Matlab code that analyzes the data in 'Opto_Inactivation_data.mat'.<br> 40. 'Load_Pharmacological_Inactivation_database.m' - this is a Matlab code, which is called in the Matlab code that analyzes the data in 'Mus_Inactivation_data.mat'.<br> 41. 'bonf_holm.m' - this is a Matlab code developed by D. M. Groppe, which is called in the Matlab code 'plot_figS4B_Pharmacological_Inactivations.m':<br> https://ch.mathworks.com/matlabcentral/fileexchange/28303-bonferroni-holm-correction-for-multiple-comparisons<br> 42. 'boundedline.m' - this is a Matlab code developed by K. Kearney, which is called in the Matlab codes 'plot_fig1C_Silicon_Probe_Hit_trials.m'; 'plot_fig2A_SEP_D1_vs_Trained.m'; 'plot_fig3A_SEP_D1_vs_Exposed.m’; 'plot_fig3C_ROC_Trained_vs_Exposed.m'; 'plot_fig3D_mPFC_PSTH_and_zscore_DT_vs_NE.m'; 'plot_fig4A_SEP_H_vs_M.m'; 'plot_fig4C_mPFC_PSTH_Hit_vs_Miss.m'; 'plot_figS3B_Stim_vs_Catch_for_significantly_inc_dec_units.m'; 'plot_figS4A_H_vs_M_for_inc_dec_units_and_zscored_PSTH.m':<br> https://ch.mathworks.com/matlabcentral/fileexchange/27485-boundedline-m<br> 43. 'inpaint_nans.m' - this is a Matlab code, which is called in the Matlab code 'boundedline.m'.<br> 44. 'PSTH_Simple.m' - this is a Matlab code developed by V. Esmaeili, which is called in the Matlab codes 'plot_fig1C_Silicon_Probe_Hit_trials.m'; 'plot_fig3D_mPFC_PSTH_and_zscore_DT_vs_NE.m'; 'plot_fig4C_mPFC_PSTH_Hit_vs_Miss.m'; 'plot_figS3B_Stim_vs_Catch_for_significantly_inc_dec_units.m'; 'plot_figS4A_H_vs_M_for_inc_dec_units_and_zscored_PSTH.m’.</p>
Reference Data Set: Electricity, Heat, and Gas Sector Data for Modeling the German System
<p>This reference data set representing the status quo of the German electricity, heat, and natural gas sectors was compiled within the research project ‘LKD-EU’ (Long-term planning and short-term optimization of the German electricity system within the European framework: Further development of methods and models to analyze the electricity system including the heat and gas sector).</p> <p>While the focus is on the electricity sector, the heat and natural gas sectors are covered as well. With this reference data set, we aim to increase the transparency of energy infrastructure data in Germany. Where not otherwise stated, the data included in this report is given with reference to the year 2015 for Germany. The data set is documented in DIW Data Documentation 92 (see references).</p> <p>The project is a joined effort by the German Institute for Economic Research (DIW Berlin), the Workgroup for Infrastructure Policy (WIP) at Technische Universität Berlin (TUB), the Chair of Energy Economics (EE2) at Technische Universität Dresden (TUD), and the House of Energy Markets & Finance at University of Duisburg-Essen. The project was funded by the German Federal Ministry for Economic Affairs and Energy through the grant ‘LKD-EU’, FKZ 03ET4028A-D.</p>
A data set on "Utilizing Constant Energy Difference between sp-Peak and C 1s Core Level in Photoelectron Spectra for Unambiguous Identification and Quantification of Diamond Phase in Nanodiamonds"
<p>The data set to paper: </p> <p>Utilizing Constant Energy Difference between sp-Peak and C 1s Core Level in Photoelectron Spectra for Unambiguous Identification and Quantification of Diamond Phase in Nanodiamonds</p> <p>Oleksandr Romanyuk1,*, Štěpán Stehlík1,2, Josef Zemek1, Kateřina Aubrechtová Dragounová1,3 and Alexander Kromka1</p> <p>1 Institute of Physics of the Czech Academy of Sciences, Cukrovarnická 10, 162 00 Prague, Czech Republic<br>2 New Technologies—Research Centre, University of West Bohemia, Univerzitní 8, 306 14 Pilsen, Czech Republic<br>3 Faculty of Nuclear Sciences and Physical Engineering, Czech Technical University in Prague, Břehová 7, 115 19 Prague, Czech Republic</p> <p>* corresponding author: romanyuk@fzu.cz</p> <p>Data manager: Kristýna Dostálová: dostalovak@fzu.cz</p> <p>Date of data collection: 1. 1. 2024 - 15. 03. 2024</p> <p>All the data showed in the pictures are provided in X-Y format with described sample. Always, the respective Figure to which the data belong is provided in high resolution. <br>The data are in the following formats: <br>Figure 1: tiff, csv<br>Figure 2: tiff, csv<br>Figure 3: tiff, csv<br>Figure 4: tiff, csv<br>Figure 5: tiff, csv</p> <p>Data acquistion and processing is provided in the Experimental part in the publication: DOI:10.3390/nano14070590</p>
Data set and scripts - Influence of Festive Periods on Road Safety: Multidimensional Analysis (Road Accidents in Colombia 2017-2021)
<p>This dataset comprises historical information about road accidents in Colombia from 2017 to 2021, titled 'Road Accidents 2017-2021', containing 18,600 records of accident events on roads managed by the National Roads Institute (INVÍAS, 2021). The dataset includes 41 descriptors and was last updated on July 15, 2022. It has been published under the Open Data initiative (Law 1712 of 2014 on Transparency and Access to National Public Information).</p> <p>In addition to accident information, the dataset integrates a database with holiday dates and road identifiers, ensuring data coherence and quality for data analysis purposes. Statistical analysis is conducted through exploratory data analysis focusing on the years 2017 to 2021, utilizing Python (version 3.10) within the Jupyter Notebooks execution environment and specialized libraries (Pandas, NumPy, Matplotlib, and Seaborn), due to their ease of application for this dataset. After data normalization, the dataset comprises 18,554 records, with 46 excluded due to inconsistent data formats.</p>
Data set: Variations in water economy traits in two Sphagnum species across their distribution boundaries
<p><em>Sphagnum</em> trait data collected (2016-2017) across a climatic gradient in Sweden. Trait data for both shoot and canopy traits. Data for <em>Sphagnum cuspidatum</em> and <em>Sphagnum lindbergii</em>. Also contains data on species occurrence records in Sweden and output from speceis distribution modelling. See published paper for more information.</p> <p>Files contain (i) processed data ("calculated_trait_data...cvs"), (ii) raw data ("Campbell_etal_clim_traits_...cvs"), (iii) their readme files, and (iv) R-scripts to run the analyses. Note that you need the files in the zip-file to run the analyses in the R-script. The zip-file contains all raw data (climate, traits, species occurences), MaxEnt output, and raster files from photogrammetry.</p> <p>More info in paper: <a href="https://doi.org/10.1002/ajb2.16347" target="_blank" rel="noopener">https://doi.org/10.1002/ajb2.16347</a></p>
Data Set for "Alteration's control on frictional behavior and the depth of the ductile shear zone in geothermal reservoirs in volcanic arcs" I: Lesser Antilles Volcanic Arc
<p>Data set for the 48 friction experiments performed for gouge samples (altered andesitic rocks) from the Lesser Antilles used in the manuscript, "Alteration's control on frictional behavior and the depth of the ductile shear zone in geothermal reservoirs in volcanic arcs". This data set can be used in combination with the data set for the Cascades used in the same manuscript (doi:10.5281/zenodo.10964936). This large combined data set (of 108 frictional experiments) represents a unique opportunity to systematically study frictional behaviour in the framework of rate and state. All samples are tested in wet and dry conditions at 10, 30, and 50 MPa with velocity steps and slide-hold-slides. These two data sets have the further advantage of being performed with exactly the same protocol (same run in, same initial gouge thickness, same velocity steps, same hold periods), in the same machine, by the same operator (or by an operator who was trained and supervised by the original operator). </p>
Minimal data set for: Air-liquid interface exposure of A549 human lung cells to characterize the hazard potential of a gaseous bio-hybrid fuel blend
<p>This minimal data set presents the values behind the means and standard deviation for the publication entitled: "Air-liquid interface exposure of A549 human lung cells to characterize the hazard potential of a gaseous bio-hybrid fuel blend"</p>
Data set of pup retrieval test in control and V1b vasopressin receptor knockout mice
<div> <div> <div> <p>This repository contains the images and code used in the paper "Computer vision analysis of mother-infant interaction identified efficient pup retrieval in V1b receptor knockout mice".</p> <p>For effective nursing, close contact between lactating mothers and their infants is necessary. However, evaluation of maternal motivation to retrieve pups is challenging, because multiple infants were randomly accessed multiple times in changing background. We applied computer vision and deep learning analysis in this process to precisely calculate maternal behavior. Object detection in an open filed test identified less entry into the center area and less moving distance in virgin female mice lacking V1b vasopressin receptor (V1bKO) than wild-type (WT) mice. Although this character was replicated in a V1bKO mother, a pup retrieval test showed that total distances among a V1bKO mother and infants come closer in a shorter time than with a WT mother. In the medial preoptic area, V1b receptor message was partly detected in galanin- and c-fos-positive neurons after the mother was stimulated by infants. Our deep learning analysis effectively evaluated mother-infant relationship in V1bKO mice.</p> </div> </div> </div>
TCOM-COF2: TOMCAT CTM and Occultation measurement-based Stratospheric COF2 profile data set
<p>TCOM-COF2: TOMCAT CTM and Occultation measurement-based Stratospheric TCOM-COF2 profile data set </p> <p>Sandip S. Dhomse </p> <p>School of Earth and Enviro, University of Leeds, Leeds, UK</p> <p>National Centre for Earth Observations, University of Leeds, Leeds, UK</p> <p> email: s.s.dhomse@leeds.ac.uk</p> <p> </p> <p>Methodology: TOMCAT simulation is performed at T64L32 resolution for the 2000-2023 time period. Collocated COF2 profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement differences are calculated for each zonal bins (51 height levels, 10km to 60km). Separate XGBoost regression models are trained for the differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids. Estimated corrections for a given model grid that are added to the original TOMCAT simulated daily (at 1.30 local time) COF2 profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values. For more details see attached presentation.</p> <p>Dataset also includes two files containing daily mean zonal mean COF2 profiles on height (10-60 km) and pressure (300-0.1 hPa) levels (8766 days/64 latitudes):</p> <p>zmcof2_TCOM_hlev_T2Dz_2000_2023.nc – height level data (10 to 60 km)</p> <p>zmcof2_TCOM_plev_T2Dz_2020_2023.nc – pressure level data (300 to 0.1 hPa)</p> <p>Daily 3D profiles on height and pressure levels would be made available on request. Xarrays “resample” can be used to get monthly means.</p>
Data set: UAS-based optical- and thermal infrared remote sensing of the fumarole field of La Fossa cone, Vulcano Island (Italy), reveals the degassing and hydrothermal alteration structure
<p>This is the data set supporting the paper "Anatomy of a fumarole field; drone remote sensing and petrological approaches reveal the degassing and alteration structure at La Fossa cone, Vulcano Island, Italy" (DOI: <a href="https://doi.org/10.5194/egusphere-2023-1692" target="_blank" rel="noopener noreferrer">10.5194/egusphere-2023-1692</a>).</p> <p> </p> <p><strong>Short description of the study:</strong> Hydrothermal alteration is common on actively degassing volcanoes and can lead to significant changes in the physical and chemical properties of the volcanic rocks, such as changes in permeability or rock strength. Despite the potentially far-reaching consequences of hydrothermal alteration for volcano stability, less is known about the detailed structures and dynamics of degassing and alteration systems. In this study, we use UAS-derived high-resolution data to analyze the fumarole field at La Fossa cone, Vulcano Island (Italy), aiming to better understand the structures and dynamics of volcanic degassing and alteration systems. By combining Principal Component Analysis, image analysis, and classification applied to high-resolution optical data and analysis of thermal infrared data, we resolve the detailed structure of the surficial degassing and alteration system based on optical and thermal anomalies. We identified characteristic anomaly patterns that indicate local degassing and alteration variability, and larger units of diffuse activity that, next to high-temperature fumaroles, contribute significantly to the total activity. We compared the observed anomaly patterns with the mineralogical and geochemical composition of representative rock samples, and with the surface degassing activity, and are able to provide the anatomy of the La Fossa fumarole field at great resolution. We show local alteration gradients, the presence of larger diffuse active complexes, and evidence for dynamic processes associated with the hydrothermal alteration. For more details, please read on: "<em>Müller, D., Walter, T. R., Troll, V. R., Stammeier, J., Karlsson, A., De Paolo, E., ... & De Jarnatt, B. (2023). Anatomy of a fumarole field; drone remote sensing and petrological approaches reveal the degassing and alteration structure at La Fossa cone, Vulcano Island, Italy. EGUsphere, 2023, 1-45. </em> https://doi.org/10.5194/egusphere-2023-1692".</p> <p> </p> <p> </p> <p><strong>Data set:</strong> We provide a UAS-based high-resolution dataset covering the whole La Fossa cone, including aerial Orthomosaic, Digital Elevation Model, and a Temperature Map derived from an airborne optical- and thermal infrared sensor (acquired in 2018 and 2019). </p> <p>The dataset is organized in 1) photogrammetric data, and 2) relevant processing results and related data. <strong>Filenames</strong> are written in bold letters and are a composite of the file type and the date (YYYYMMDD). </p> <p> </p> <p> </p> <p><strong>1) Photogrammetric data: </strong></p> <ul> <li><strong>Orthomosaic_20191114.tif</strong> is the in Agisoft Metashape processed orthomosaic of a 150 m (above fumarole field) optical overflight (DJI Phantom 4 Pro camera). </li> <li><strong>DigitalElevationModel_20191114.tif</strong> is the in Agisoft Metashape processed Digital Elevation Model (DEM) from the above-mentioned 150 m overflight. </li> <li><strong>Hillshade_20191114.tif</strong> is the 2.5-D representation of the DigitalElevationModel_20191114. Note, for viewing use a stretched (black to white) color scale.</li> <li><strong>TemperatureMap_20181115.tif</strong> is showing the apparent surface temperature for the La Fossa cone, acquired by a Flir Tau 2 thermal infrared camera at ~150 m (above fumarole field) flight altitude in the early morning hours (before sunrise) of 15 November 2018. Note that apparent temperatures shown may underestimate real in situ fumarole temperatures due to pixel-to-vent size ratios and atmospheric- or gas-plume distortion effects. Note further that the data has some processing artifacts, due to blind pixels of our IR camera system. For more detailed information or an updated data set please contact dmueller@gfz-potsdam.de.</li> <li><strong>T_20to40C.tif</strong> shows the diffuse thermally active surface at the fumarole field of the La Fossa cone (units a-g, see Fig. 4 in "Anatomy of a fumarole field...", https://doi.org/10.5194/egusphere-2023-1692). This raster shows the extracted pixels from TemperatureMap_20181115 in the range of 22 - 40 °C.</li> <li><strong>T_higher40C.tif</strong> outlines the high-temperature fumarole locations of the La Fossa fumarole field (HTF, see Fig. 4 in "Anatomy of a fumarole field...", https://doi.org/10.5194/egusphere-2023-1692), based on the extracted pixels with temperatures > 40 °C from TemperatureMap_20181115.</li> </ul> <p>Shapefiles for temperatures > 40 °C representing the high-temperature fumarole locations (HTF) and for temperatures of 20 - 40 °C representing diffuse active units, are attached at the end of the upload list and named <strong>T_higher40C_polygon</strong> and <strong>T_20_40C_polygon</strong> and consist of multiple files per shapefile with the file extensions .CPG, .dbf, .prj, .sbn, .sbx, .shp, .shp.xml, .shx. </p> <p>The coordinate system of the data sets is WGS84 EPSG:4326. For nadir projection use WGS 84 / UTM zone 33N - EPSG:32633. Note that the data might have horizontal and vertical offsets in the typical range of SfM-derived products with single-band GPS accuracy.</p> <p> </p> <p> </p> <p><strong>2) Relevant processing steps and related data:</strong></p> <ul> <li>Step 1) Principal Component Analysis applied to Orthomosaic_20191114 results in the following 3 Principal Components (decorrelated variance representations of the initial RGB bands): <ul> <li><strong>1_PCA_PC1.tif </strong>1st principal component </li> <li><strong>1_PCA_PC2.tif</strong> 2nd principal component</li> <li><strong>1_PCA_PC3.tif</strong> 3rd principal component - highlights well the effects of concentrated and diffuse degassing, resulting in different alteration effects from a simple shift from reddish oxidized surface to gray, up to strong silicic alteration effects. This can be used to extract the data of interest, the hydrothermally altered surface, and to create a new alteration sub-dataset. </li> </ul> </li> <li>Step 2) Extraction of hydrothermally altered surface / alteration sub-dataset <ul> <li><strong>2_alteration_subdata_RGB.tif</strong> The alteration sub-data set was extracted from the original Orthomosaic_20191114 based on a mask obtained from Principal Component 3 (1_PCA_PC3) for values > 85. The resulting raster data set is an extract of the original RGB data.</li> </ul> </li> <li>Step 3) PCA applied to 2_alteration_subdata_RGB will adjust to the reduced spectral range of the alteration sub-data set, provide a more sensitive variance representation, and highlight variability within the hydrothermally altered surface. <ul> <li><strong>3_PCA_PC1.tif</strong> 1st principal component of 2_alteration_subdata_RGB</li> <li><strong>3_PCA_PC2.tif</strong> 2nd principal component of 2_alteration_subdata_RGB</li> <li><strong>3_PCA_PC3.tif</strong> 3rd principal component of 2_alteration_subdata_RGB</li> </ul> </li> <li>Step 4) Unsupervised classification <ul> <li><strong>4_classification.tif</strong> is the unsupervised classification result of 3_PCA (all Principal Components), classified into 32 classes to achieve a high class resolution. When combining different classes, they form larger spatial units / surface types with similar spectral characteristics. This way, we divide the alteration surface into 3 surface types (see Fig. 4B in "Anatomy of a fumarole field..." DOI: 10.5194/egusphere-2023-1692) representing different alteration gradients and important structural units. To achieve the same results, combine classes 1 -19 (surface type 3), 20 - 25 (surface type 2), 26 - 30 (surface type 1), and 31 - 32 for sulfur/fumarole plume. See Image <strong>optical_structure.jpg</strong> for comparison. </li> </ul> </li> </ul> <p>Note that Principal Components and Classification of Principal Components highlight data variability along the axes of highest data variance. Results have to be evaluated carefully and may be valid only locally. They are efficient for identifying variability in degassing and alteration areas, but at the same time may also highlight certain fractions of vegetation or settlements for instance. We evaluated the structure defined by our classification results by analyzing the thermal structure (<strong>thermal_structure.jpg</strong>) of the fumarole field and additional geochemical- and mineralogical investigations (XRD and XRF) of rock samples and by measuring the diffuse degassing from surface (see "Anatomy of a fumarole field..." DOI: 10.5194/egusphere-2023-1692) to prove that the observed degassing/alteration units are true.</p> <p>To highlight alteration effects throughout the entire La Fossa cone, including the southern inner and outer crater rim, the alteration zones of La Forgia, or alteration on the outer flanks of La Fossa e.g. the 1988 Landslide, we provide the raster <strong>La_Fossa_alteration.tif </strong>and image <strong>La_Fossa_alteration.jpg (</strong>Note that the color scale for strong alteration (classes 31 - 32) was changed from white to purple for highlighting purpose).</p> <p> </p> <p>In case of further questions about the dataset, please contact dmueller@gfz-potsdam.de.</p> <p> </p> <p> </p> <p> </p> <p> </p>
Data set of "Capacitive and Inductive Characteristics of Volatile Perovskite Resistive Switching Devices with Analog Memory"
<p>The dataset of all data presented in the article published in the virtual special issue of the Journal of Physical Chemistry Letters:</p> <p>"Capacitive and Inductive Characteristics of Volatile Perovskite Resistive Switching Devices with Analog Memory"</p> <p>DOI: <a title="DOI URL" href="https://doi.org/10.1021/acs.jpclett.4c00945">https://doi.org/10.1021/acs.jpclett.4c00945</a></p> <p> </p> <p>The dataset contains the following raw data:</p> <p>## FILE DESCRIPTION<br>--------------<br>### Figure 2<br>- Fig2a.txt : Representative characteristic _I-V_ response of memristor (5 cycles)<br>- Fig2b.txt : Upper vertex-dependent multilevel/multistate analog resistive switching<br>- Fig2c.txt : Characteristic _I-V_ response of 20 distinct devices<br>- Fig2d.txt : Endurance measurements for 1000 cycles of the LRS (ON state) and HRS (OFF state)</p> <p>### Figure 3<br>- Fig3a.txt : Characteristic _I-V_ response with an upper vertex of 0.25 V<br>- Fig3b.txt : Characteristic _I-V_ response with an upper vertex of 0.75 V<br>- Fig3c.txt : Characteristic _I-V_ response with an upper vertex of 1.25 V</p> <p>### Figure 4<br>- Fig4a.txt : IS spectrum under dark conditions at 0 V<br>- Fig4b.txt : IS spectrum under dark conditions at 0.2 V<br>- Fig4c.txt : IS spectrum under dark conditions at 0.3 V<br>- Fig4d.txt : IS spectrum under dark conditions at 0.4 V<br>- Fig4e.txt : IS spectrum under dark conditions at 0.6 V<br>- Fig4f.txt : IS spectrum under dark conditions at 1.0 V</p> <p>### Figure 5<br>- Fig5a.txt : Voltage-dependent transient current response of the perovskite memristor<br>- Fig5b.txt : Magnified view of the transient current response of a single voltage pulse at representative applied voltages<br>- Fig5c.txt : Pulse width-dependent transient current response<br>- Fig5d.txt : Corresponding magnified view of the first and last transient responses<br>- Fig5e.txt : Synaptic potentiation and depression characteristic response of the memristor</p> <p>### Figure 6<br>- Fig6a.txt : Transient current response of the volatile perovskite memristor with a single long pulse vs. a train of short pulses at 0.4 V<br>- Fig6b.txt : Corresponding magnified view of the transient current response of a first voltage pulse at 0.4 V<br>- Fig6c.txt : Transient current response of the volatile perovskite memristor with a single long pulse vs. a train of short pulses at 0.8 V<br>- Fig6d.txt : Corresponding magnified view of the transient current response of a first voltage pulse at 0.8 V<br>- Fig6e.txt : Transient current response of the volatile perovskite memristor with a single long pulse vs. a train of short pulses at 1.2 V<br>- Fig6f.txt : Corresponding magnified view of the transient current response of a first voltage pulse at 1.2 V</p>
Data Set of Industrial Metaverse Use Cases
<p>Data Set of Industrial Metaverse Use Cases</p> <p>Potential of the Industrial Metaverse – A Taxonomic Approach<br>IFIP 21st International Conference on Product Lifecycle Management (2024)</p> <p>This data set comprises the following components:</p> <ul> <li>Use Cases & Classification</li> <li>Dimensions & Characteristics</li> <li>Review Documentation</li> </ul>
Mud and organic content are strongly correlated with microplastic contamination in a meandering riverbed - Data sets
<p>This is the dataset relating to publication "Microplastics distribution in a meandering riverbed reveasl mud content as a universal normalizer for microplastic contamination in aquatic environments" by Van Daele, M., Van Bastelaere, B., de Clercq, J., Meyer, I., Vercauteren, M. and Asselman, J.</p> <p>It contains the microplastic and sedimentological data that support the findings of that publication, with a seperate file for data obtained from riberbed sediments and the water column. It further contains a file with all source data for the graphs in the figures of the main manuscript and the Supplementary Information<span><span>.</span></span></p>
Mass mortality among colony-breeding seabirds in the German Wadden Sea in 2022 due to distinct genotypes of HPAIV H5N1 clade 2.3.4.4b: data sets on phylogeographic analyses
<p>Highly pathogenic avian influenza viruses (HPAIV) of clade 2.3.4.4b of the H5 goose/Guangdong (gs/GD) lineage have repeatedly emerged in Germany since 2016. Both poultry holdings and wild birds have been heavily hit but the 2020-2021 and 2021-2022 HPAI winter seasons exceeded all previously recorded epizootics in Germany in terms of number of wild bird cases recorded, genetic diversity of viruses, and duration of virus activity. In past seasons regional massing of wild bird cases were seen at the German coasts of the Baltic and North Sea, but species mainly affected varied from season to season. In 2022 a new and, in Europe, unprecedented aspect was observed when several cormorant and seabird breeding colonies became affected since May at the Baltic Sea coast and in the Wadden Sea, respectively by HPAI H5N1 viruses.</p> <p> </p>
AIT Netflow Data Set
<p><strong>AIT Netflow Data Sets</strong></p> <p>This repository contains labeled synthetic netflows suitable for evaluation of intrusion detection systems, federated learning, and alert aggregation. The netflows are generated from the packet captures contained in the <a href="http://doi.org/10.5281/zenodo.5789064">AIT-LDS-v2.0</a>. A detailed description of that dataset is available in [1]. The packet captures were collected from eight testbeds that were built at the Austrian Institute of Technology (AIT) following the approach by [2]. Please cite these papers if the data is used for academic publications.</p> <p>In brief, each of the datasets corresponds to a testbed representing a small enterprise network including mail server, file share, WordPress server, VPN, firewall, etc. Normal user behavior is simulated to generate background noise over a time span of 4-6 days. At some point, a sequence of attack steps is launched against the network. The following attacks are launched in the network:</p> <ul> <li>Scans (nmap, WPScan, dirb)</li> <li>Webshell upload (CVE-2020-24186)</li> <li>Password cracking (John the Ripper)</li> <li>Privilege escalation</li> <li>Remote command execution</li> <li>Data exfiltration (DNSteal)</li> </ul> <p>This repository contains the following files:</p> <ul> <li><em><testbed>_netflows.zip</em>: CSV files of labeled TCP and UDP netflows for each testbed.</li> <li><em>label_info.txt</em>: File describing which labels in TCP and UDP are benign and which ones are malicious.</li> <li><em>README.md</em>: Instructions on how to reproduce the generation and labeling of the netflows from the <a href="http://doi.org/10.5281/zenodo.5789064">AIT-LDS-v2.0</a>. Note that it is only necessary to run the python scripts if you want to extend or change the labeling procedure.</li> <li><em>1_format_dataset_info.ipynb</em>: Generates the tables necessary for labeling (see README.md).</li> <li><em>2_label_logs.ipynb</em>: Labels the netflows (see README.md).</li> </ul> <p>Acknowledgements: Partially funded by the FFG projects INDICAETING (868306) and DECEPT (873980), and the EU projects GUARD (833456) and PANDORA (SI2.835928).</p> <p><strong>If you use the dataset, please cite the following publications:</strong></p> <p>[1] M. Landauer, F. Skopik, M. Frank, W. Hotwagner, M. Wurzenberger, and A. Rauber. <a href="https://ieeexplore.ieee.org/abstract/document/9866880">"Maintainable Log Datasets for Evaluation of Intrusion Detection Systems"</a>. IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 4, pp. 3466-3482. [<a href="https://arxiv.org/pdf/2203.08580.pdf">PDF</a>]</p> <p>[2] M. Landauer, F. Skopik, M. Wurzenberger, W. Hotwagner and A. Rauber, <a href="https://ieeexplore.ieee.org/document/9262078">"Have it Your Way: Generating Customized Log Datasets With a Model-Driven Simulation Testbed,"</a> in IEEE Transactions on Reliability, vol. 70, no. 1, pp. 402-415, March 2021, doi: 10.1109/TR.2020.3031317. [<a href="https://www.skopik.at/ait/2020_trel.pdf">PDF</a>]</p>
Data sets for FASTREAD
<p>Three datasets created with information in existing systematic literature review publications.</p> <p>Hall.csv is created according to Hall, et al., "A systematic literature review on fault prediction performance in software engineering", 2012.</p> <p>Wahono.csv is created according to Wahono, et al., "A systematic literature review of software defect prediction: research trends, datasets, methods and frameworks", 2015.</p> <p>Radjenovic.csv is created according to Radjenovic, et al., "Software fault prediction metrics: A systematic literature review", 2013.</p> <p>One dataset (Kitchenham.zip) provided by Barbara Kitchenham (cleaned and summarized by Zhe Yu) from Kitchenhm et al. "Systematic literature reviews in software engineering- A tertiary study", 2010.</p> <p>These datasets are used in our submission to Empirical Software Engineering: "How to Read Less: On the Benefit of Active Learning for Primary Study Selection in Systematic Literature Reviews".</p>
MPS Data set with images of medieval charters for handwriting-style based dating of manuscripts
<pre>The MPS benchmark data set for handwritten manuscript dating ____________________________________________________________ This data set is collected for the Dutch NWO project: Medieval Paleographical Scale (MPS) by Petros Samara Project website: http://application02.target.rug.nl/monk/Projects/MPS/ Copyright (c) Huygensinstituut, Den Haag, 2016 University of Groningen, 2016. All rights reserved. Organisation of the data: Each .tar.gz file contains a number of NetPBM images. The format is chosen because of its simplicity. Also, there is no doubt about lossy compression in the processing chain. The file names are of the format 'MPS<year>_<seqnr>.ppm', for example, 'MPS1300_0056.ppm'. Note: the files are not in a separate directory, they will be extracted in place. However, due to the unique naming, there is no problem extracting them in one single current (destination) directory. The actual type of the image can be gray scale (.pgm) or color (.ppm), in '8-bit DirectClass' according to ImageMagick's 'identify' tool. The images were cropped out of larger photographs because of irrelevant elements such as a Kodak color calibrator and non-text content such as supporting surface (table) backgrounds, seals (emblems), ribbons, etc. No effort has been made to obtain a balanced set of samples over years: the given frequencies of occurrence in archives are used. There is evidently less data in years before 1375 A.D. while some periods provides us with ample data for historical reasons (e.g, 1450 A.D.). It would have been a pity if the scarce years had determined and limited the size of this data set. Selection criteria for data reduction, whether random or systematic, would have been arbitrary. In any case, these images were used in our publications, such that the performance results of future attempts on manuscript dating can be compared with earlier results. The performances that have been reached using our algorithms are in the order of an MAE (mean average error) of 10 years. If you have any questions, please contact us: Sheng He (heshengxgd@gmail.com) Petros Samara (petros.samara@huygens.knaw.nl) Jan Burgers (jan.burgers@huygens.knaw.nl) Lambert Schomaker (L.Schomaker@ai.rug.nl) Please cite our papers if you use this data set: [1] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Image-based historical manuscript dating using contour and stroke fragments. Pattern Recognition(PR), Vol. 59, pp. 159-171, 2016 [2] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Towards style-based dating of historical documents. International Conference on Frontiers in Handwriting Recognition(ICFHR), Crete, Greece, 2014 [3] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Multiple-Label Guided Clustering Algorithm for Historical Document Dating and Localization IEEE Trans. on Image Processing, Vol. 25(11), Nov. 2016. http://ieeexplore.ieee.org/document/7551181/</pre> <p>Data are collected thanks to Dutch NWO grant project 380-50-006</p>
Data set and benchmarks from Eriksson et al., ICAPS 2018
<p>These are the benchmarks and the experiment data used in the paper "A Proof System for Unsolvable Planning Tasks" by Eriksson et al. (ICAPS 2018). The raw data contains the logs from all runs, while the eval-directories contain a json file with all parsed attributes. The benchmarks directory contains all benchmarks used in the experiments. Finally, the file eriksson-et-al-icaps2018.html provides an overview over the most interesting attributes across all runs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.