Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,185

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,185 results for “Learning”

Learn how ShareScore rates datasets ↗
zenodo48/100

Learning Dynamics of Electrophysiological Brain Signals During Human Fear Conditioning (Open Data and Open Materials)

<p><strong>Open Data and Open Materials of: Sperl, M. F. J.,&nbsp;Wroblewski, A., Mueller, M., Straube, B., &amp; Mueller, E. M. (2021).&nbsp;Learning Dynamics of Electrophysiological Brain Signals During Human Fear Conditioning.&nbsp;<em>NeuroImage</em>,&nbsp;<em>226</em>, 117569.</strong></p> <p>Electrophysiological studies in rodents allow recording neural activity during threats with high temporal and spatial precision. Although fMRI has helped translate insights about the anatomy of underlying brain circuits to humans, the temporal dynamics of neural fear processes remain opaque and require EEG. To date, studies on electrophysiological brain signals in humans have helped to elucidate underlying perceptual and attentional processes, but have widely ignored how fear memory traces&nbsp;<em>evolve</em>&nbsp;over time. The low signal-to-noise ratio of EEG demands aggregations across high numbers of trials, which will wash out transient neurobiological processes that are induced by learning and prone to habituation. Here, our goal was to unravel the plasticity and temporal emergence of EEG responses during fear conditioning. To this end, we developed a new sequential-set fear conditioning paradigm that comprises three successive acquisition and extinction phases, each with a novel CS+/CS- set. Each set consists of two different neutral faces on different background colors which serve as CS+ and CS-, respectively. Thereby, this design provides sufficient trials for EEG analyses while tripling the relative amount of trials that tap into more transient neurobiological processes. Consistent with prior studies on ERP components, data-driven topographic EEG analyses revealed that ERP amplitudes were potentiated during time periods from 33&ndash;60 ms, 108&ndash;200 ms, and 468&ndash;820 ms indicating that fear conditioning prioritizes early sensory processing in the brain, but also facilitates neural responding during later attentional and evaluative stages. Importantly, averaging across the three CS+/CS- sets allowed us to probe the temporal evolution of neural processes: Responses during each of the three time windows gradually increased from early to late fear conditioning, while long-latency (460&ndash;730 ms) electrocortical responses diminished throughout fear extinction. Our novel paradigm demonstrates how short-, mid-, and long-latency EEG responses change during fear conditioning and extinction, findings that enlighten the learning curve of neurophysiological responses to threat in humans.</p>

opencc-by-4.0Nov 2020View details →
zenodo48/100

Experimental data for the motor learning study performed: "Promoting Motor Variability During Robotic Assistance Enhances Motor Learning of Dynamic Tasks"

<p>The dataset contains the kinematic data and the questionnaire responses for a robot-assisted motor learning study performed in the Motor Learning and Neurorehabilitation Laboratory at University of Bern. The details of the study are described in [doi: 10.3389/fnins.2020.600059]. The kinematic data for each participant is stored as a data frame inside a &ldquo;pickle&rdquo; (serialized python object) file. The questionnaire responses are stored as a &ldquo;csv&rdquo; file. The variables inside the files are explained in &ldquo;DataframeVariableDescription.rtf&rdquo;. For questions, please contact oezhan.oezen@artorg.unibe.ch or L.MarchalCrespo@tudelft.nl.</p>

opencc-by-4.0Dec 2020View details →
zenodo48/100

Written and spoken digits database for multimodal learning

<p><strong>Database description:</strong></p> <p>The written and spoken digits database is not a new database but a constructed database from existing ones, in order to provide a ready-to-use database for multimodal fusion [1].</p> <p>The written digits database is the original MNIST handwritten digits database [2] with no additional processing. It consists of 70000 images&nbsp;(60000 for training and 10000 for test) of 28 x 28 = 784 dimensions.</p> <p>The spoken digits database was extracted from Google Speech Commands [3], an audio dataset of spoken words that was proposed to train and evaluate keyword spotting systems. It consists of 105829 utterances of 35 words, amongst which 38908 utterances of the ten digits (34801 for training and 4107 for test). A pre-processing was done via the extraction of the Mel Frequency Cepstral Coefficients (MFCC) with a framing window size of 50 ms and frame shift size of 25 ms. Since the speech samples are approximately 1 s long, we end up with 39 time slots. For each one, we extract 12 MFCC coefficients with an additional energy coefficient. Thus, we have a final vector of 39 x 13 = 507 dimensions. Standardization and normalization were&nbsp;applied on the MFCC features.</p> <p>To construct the multimodal digits dataset, we associated written and spoken digits of the same class respecting the initial partitioning in [2] and [3] for the training and test subsets. Since we have less samples for the spoken digits, we duplicated some random samples to match the number of written digits and have a multimodal digits database of 70000 samples&nbsp;(60000 for training and 10000 for test).</p> <p>The dataset is provided in six files as described below. Therefore, if a shuffle is performed on the training or test subsets, it must be performed in unison with the same order for the written digits, spoken digits and labels.</p> <p>&nbsp;</p> <p><strong>Files:</strong></p> <ul> <li>data_wr_train.npy: 60000 samples of 784-dimentional written digits for training;</li> <li>data_sp_train.npy: 60000 samples of 507-dimentional spoken digits for training;</li> <li>labels_train.npy: 60000 labels for the training subset;</li> <li>data_wr_test.npy: 10000 samples of 784-dimentional written digits for test;</li> <li>data_sp_test.npy: 10000 samples of 507-dimentional spoken digits for test;</li> <li>labels_test.npy: 10000 labels for the test subset.</li> </ul> <p>&nbsp;</p> <p><strong>References:</strong></p> <ol> <li>Khacef, L. et al. (2020), &quot;Brain-Inspired Self-Organization with Cellular Neuromorphic Computing for Multimodal Unsupervised Learning&quot;.</li> <li>LeCun, Y. &amp; Cortes, C. (1998), &ldquo;MNIST handwritten digit database&rdquo;.</li> <li>Warden, P. (2018), &ldquo;Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition&rdquo;.</li> </ol>

opencc-by-4.0Oct 2019View details →
zenodo48/100

Dataset Nucleation Patterns of Polymer Crystals Analyzed by Machine Learning Models

<p>This dataset contains the raw data (01_raw_data), processed data (02_processed_data), and plotting scripts (03_figures) related to the paper:</p> <p>"Nucleation Patterns of Polymer Crystals Analyzed by Machine Learning Models"<br>Atmika Bhardwaj, Jens-Uwe Sommer, Marco Werner</p> <p>Macromolecules <strong>2024</strong>; DOI: <a href="10.1021/acs.macromol.4c00920">10.1021/acs.macromol.4c00920</a></p> <p>Please refer to the README.md files in their respective folders.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

Dataset of "Advanced machine learning techniques for State-of-Health estimation in lithium-ion batteries: A comparative study"

This research focuses on State-of-Health (SOH) estimation of lithium-ion (Li-ion) batteries to enhance lifespan and reliability. Using Samsung INR18650-35E cells, 600 cycles were analyzed with machine learning (ML) techniques, including Gaussian Process Regression (GPR), Support Vector Regression (SVR), Feed-Forward Neural Network (FFNN) and Adaptive Neuro-Fuzzy Inference System (ANFIS). Input features from charging and discharging cycles were selected with Pearson Correlation Analysis (PCA) and Exhaustive Search (ES) to optimize inputs for each ML method. Models were tested on datasets of varying sizes to evaluate performance and overfitting, including an experiment where SOH estimation of one battery was performed using training data from another. The findings highlight each model's strengths and limitations, guiding their application in battery health prediction.

opencc-by-4.0Nov 2024View details →
zenodo48/100

Bioactivity deep learning for structure-free compound-protein interaction

<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p>&nbsp;</p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p>&nbsp;</p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

RADIT: A Machine Learning-Reconstructed Dataset of River Discharge, Temperature, and Heat Flux into the Arctic Ocean

<p>The Reconstructed Arctic-draining river DIscharge and Temperature (RADIT) dataset provides daily records of river discharge, temperature, and heat flux for 25 major Arctic-draining rivers from 1950 to 2023. Using machine learning methods and ERA5-Land reanalysis data, we reconstructed these key hydrological variables with high accuracy (most NSEs &gt; 0.8).</p> <p>Due to licensing restrictions and to encourage adherence to the stated licenses of the original input data, this dataset only provides the reconstructed (filled) values. Users can obtain the complete historical observational data from their original publicly available sources as detailed in our documentation. By combining these original observations with our reconstructed data, a comprehensive and continuous daily dataset from 1950 to 2023 can be assembled. Clear instructions and links for downloading the original observational data used in this study can be found at: <a href="https://github.com/zhwang24/RADIT-Reconstructed-Arctic-River-Data" target="_blank" rel="noopener">https://github.com/zhwang24/RADIT-Reconstructed-Arctic-River-Data</a>. Should you encounter any issues or have questions, please feel free to contact the first author, Zihan Wang (zhwang2018@163.com).</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

ML-TOMCAT V2.0: Machine-Learning-Based Satellite-Corrected Global Stratospheric Ozone Profile Dataset

<p>MLTOMCAT V2 is 46 years (1979-2024) of gap free ozone profile data sets that is created by correcting biases in a TOMCAT Chemical Transport Model (CTM) simulated ozone profiles. We use Random Forest regression model to correct model biases.&nbsp;</p> <p>Each file contain monthly mean zonal mean ozone profiles. There are 6 data files.</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_vmr_V2.nc</a>&nbsp;contains ozone profiles on&nbsp;geometric height levels (1 to 60 km) &nbsp;in&nbsp;mixing ratio units, whereas&nbsp;&nbsp;<a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_nd_V2.nc</a>&nbsp;contains ozone profile in number density units.</p> <p>Similarly,&nbsp;</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_vmr_V2.nc</a>&nbsp;contains ozone profiles on 43 MLS pressure levels&nbsp;&nbsp;(1000 to 0.1&nbsp;hPa) &nbsp;in&nbsp;mixing ratio units, whereas&nbsp;&nbsp;<a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_nd_V2.nc</a>&nbsp;contains ozone profile in number density units.</p> <p>Please note that data below 300 hPa (~8km) and 1 hPa (~50 km) should be used with caution.</p> <p>There are two straospheric column files</p> <p>ML-TOMCAT-SCO_120ppb_boundary_V2_197901-202412.nc and</p> <p>ML-TOMCAT-SCO_150ppb_boundary_V2_197901-202412.nc</p> <p>Stratospheric column files calculated using 120 ppb and 150 ppb as a chemical ozone boundaries.</p> <p>A manuscript describing MLTOMCAT would be published in EESD (Dhomse et al., 2021).</p>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Learning Elements in Learning Management Systems (LMSs)

<p>Results of a survey in the higher education area. Participants are professors, lecturers, and tutors.</p> <p>&nbsp;</p> <p>The final definitions for the elements are:</p> <ul> <li>Brief Overview (BO): Short summary or recap without details of the actual learning material</li> <li>Quiz (QU): Quiz questions related to the content taught</li> <li>Learning Goal (LG): Description of the competences, skills or abilities that the learners should acquire in relation to a specific learning content</li> <li>Manuscript (MS): Complete or brief elaboration of a speech, a lecture, a course, or similar</li> <li>Exercise (EX): Opportunity to apply and deepen the learned. Varied tasks are possible beside the classic exercise sheet</li> <li>Summary (SU): Elementalization (reduction to the essentials) of the actual content with details</li> <li>Auditory additional material (AAM): Material with the aim of applying and deepening the learned with audio files</li> <li>Textual additional material (TAM): Material with the aim of applying and deepening the learned with textual further information (also named additional literature)</li> <li>Visual additional material (VAM): Material with the aim of applying and deepening the learned with videos or similar</li> <li>Collaboration Tool (CT): Cooperative and interactive communication medium with the aim of knowledge sharing between learners and learners and/or lecturers, and is used for collaborative work</li> </ul> <p>The corresponding scientific paper can be found via ORCID as of December 2023.</p> <p>&nbsp;</p> <p>The presented work is supported by the &lsquo;German Federal Ministry of Research, Technology and Space&rsquo; (BMFTR) through the granting of the funding project HASKI (FKZ: 16DHBKI035).</p>

opencc-by-4.0Oct 2023View details →
zenodo48/100

augMENTOR: Simulated Student Learning Profiles and their Engagement Metrics in TryHackMe Platform_V1

<p>The dataset provides simulated insights into student engagement and performance within the THM platform. It outlines mathematical representations of student learning profiles, detailing behaviors ranging from high achievers to inconsistent performers. Additionally, the dataset includes key performance indicators, offering metrics like room completion, points earned, and time spent to gauge student progress and interaction within the platform's modules.</p><p>Here are definitions of the learning profiles, along with mathematical representations of their behaviors:</p><ul><li>High Achiever: These are students who consistently perform well across all modules. Their performance can be described as a normal distribution centered at a high mean value. Their performance P in a given module can be modelled as: P = N(90, 5) where N is the normal distribution function, 90 is the mean, and 5 is the standard deviation.</li><li>Average Performer: These are students who typically perform at the average level across all modules. Their performance can be described as a normal distribution centered at a medium mean value: P = N(70, 10), where 70 is the mean, and 10 is the standard deviation.</li><li>Late Bloomer: These are students whose performance improves as they progress through the modules. Their performance can be modelled as: P = N(50 + i*10, 10), where i is the module index and shows an increasing trend.</li><li>Specialized Talent: These are students who have average performance in most modules but excel in a particular module (e.g., module5). Their performance can be described as: P = N(90, 5) if the module is module 5, else P = N(70, 10).</li><li>Inconsistent Performer: These are students whose performance varies significantly across modules. Their performance can be described as a normal distribution with a high standard deviation: P = N(70, 30), where 70 is the mean, and 30 is the high standard deviation, reflecting inconsistency.</li></ul><p>Note that the actual performances are bounded between 0 and 100 using the function max(0, min(100, performance)) to ensure valid percentages.</p><p>In these formulas, the <i>np.random.normal</i> function is used to simulate the variability in student performance around the mean values. The first argument to this function is the mean, and the second argument is the standard deviation, reflecting the level of variability around the mean. The function returns a number drawn from the normal distribution described by these parameters. Note that the proposed method is experimental and has not been validated.&nbsp;</p><p>&nbsp;</p><p>List of Key Performance Indicators (KPIs) for Student Engagement and Progress within the Platform:</p><ul><li>Room Name: This represents the unique identifier or name of a specific room (or module). Think of each room as a separate module or lesson within an educational platform. For example, Room1, Room2, etc.</li><li>Total rooms completed: Indicates the cumulative number of rooms that a student has fully completed. Completion is typically determined by meeting certain criteria, like answering all questions or achieving a certain score.</li><li>Rooms registered in: Represents the number of rooms a student has registered or enrolled in. This could be different from the total number of rooms they've completed.</li><li>Ratio of Questions completed per room: This gives an insight into a student's progress in a particular room. For instance, a ratio of 7/10 suggests the student has completed 7 out of 10 available questions in that room.</li><li>Room Completed (yes no): Indicates whether a student has fully completed a specific room or not. This could be determined by the percentage of material covered, questions answered, or a certain score achieved.</li><li>Room Last deploy (count of days): Refers to the number of days since the last update or deployment was made to that room. It can give an idea about the effort of the student.</li><li>Points in room used for the leaderboard (range 0-560): Each room assigns points based on student performance, and these points contribute to leaderboards. The range suggests that a student can earn anywhere from 0 to 560 points in a particular room.</li><li>Last answered question in a room (27th Jan 2023): This indicates the date when a student last answered a question in a specific room. It can provide insights into a student's recent activity and engagement.</li><li>Total points in all rooms (range 0-560): The cumulative score a student has achieved across all rooms.</li><li>Path Percentage completed (range 0-100): Indicates the percentage of the overall learning path that the student has completed. A path could consist of multiple modules or rooms.</li><li>Module Percentage completed (range 0-100): Represents how much of a specific module (which could have multiple lessons or topics) a student has completed.</li><li>Room Percentage completed (range 0-100): Shows the percentage of a specific room that has been completed by a student.</li><li>Time Spent on the platform (seconds): This provides an aggregate of the total time a student has spent on the entire educational platform.</li><li>Time spent on each room (seconds): Represents the amount of time a student has dedicated to a specific room. This can give insights into which rooms or modules are the most time-consuming or engaging for students.</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo48/100

Potential forest conservation value rasters for Denmark from Assmann et al. "LiDAR data fusion and machine learning identify temperate forests of high conservation value"

<p>Potential forest conservation value (high / low) rasters for Denmark based on a remote sensing data fusion approach. Please see manuscript (below) for a detailed description of the methods and data products.&nbsp;</p> <p><br>Jakob J. Assmann, Pil B. M. Pedersen, Jesper E. Moeslund, Cornelius Senf, Urs A. Treier, Derek Corcoran, Zs&oacute;fia Koma, Thomas Nord-Larsen, Signe Normand. In prep. LiDAR data fusion and machine learning identify temperate forests of high conservation value.</p> <p><br>When using the data, please cite the above manuscript.&nbsp;</p> <p><br>Files description:</p> <ul> <li>Compressed and cloud optimised rasters of potential forest conservation value projections for Denmark (10 m res.) in EPSG:3857 <ul> <li>forest_quality_ranger_biowide_10m_cog_epsg3857.tif &nbsp; &nbsp; RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_10m_cog_epsg3857.tif RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_10m_cog_epsg3857.tif GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_10m_cog_epsg3857.tif &nbsp; &nbsp; GBM model projections based on SustainScapes stratification</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li>Aggregated rasters of potential forest conservation value projections for Denmark (100 m res.) in EPSG:25832 <ul> <li>forest_quality_ranger_biowide_100m.tif RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_100m.tif RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_100m.tif GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_100m.tif GBM model projections based on SustainScapes stratification&nbsp;</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li>Uncompressed and tiled rasters of potential forest conservation value projections for Denmark (10 m res.) in EPSG:25832<br>Please note: the archives contain approx. 42k tiles, each 10 x 10 km, as well as a VRT file for covenient loading.&nbsp; <ul> <li>forest_quality_ranger_biowide_10m.zip RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_10m.zip RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_10m.zip GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_10m.zip GBM model projections based on SustainScapes stratification</li> </ul> </li> </ul>

opencc-by-4.0Dec 2023View details →
zenodo48/100

A Bayesian Machine Learning Framework for Animal Telemetry Data

<p>The data and tutorial in this repository are intended to be used in conjunction with the tutorial with our manuscript titled "A Bayesian Machine Learning Framework for Animal Telemetry Data." Telemetry data for three lesser prairie-chickens are provided here as .csv files. For more information about the data, please refer to our manuscript or contact Andrew Whetten or David Haukos for more information.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Dataset for "Machine learning predictions on an extensive geotechnical dataset of laboratory tests in Austria"

<p>This dataset comprises over 20 years of geotechnical laboratory testing data collected primarily from Vienna, Lower Austria, and Burgenland. It includes 24 features documenting critical soil properties derived from particle size distributions, Atterberg limits, Proctor tests, permeability tests, and direct shear tests. Locations for a subset of samples are provided, enabling spatial analysis.</p> <p>The dataset is a valuable resource for geotechnical research and education, allowing users to explore correlations among soil parameters and develop predictive models. Examples of such correlations include liquidity index with undrained shear strength, particle size distribution with friction angle, and liquid limit and plasticity index with residual friction angle.</p> <p>Python-based exploratory data analysis and machine learning applications have demonstrated the dataset's potential for predictive modeling, achieving moderate accuracy for parameters such as cohesion and friction angle. Its temporal and spatial breadth, combined with repeated testing, enhances its reliability and applicability for benchmarking and validating analytical and computational geotechnical methods.</p> <p>This dataset is intended for researchers, educators, and practitioners in geotechnical engineering. Potential use cases include refining empirical correlations, training machine learning models, and advancing soil mechanics understanding. Users should note that preprocessing steps, such as imputation for missing values and outlier detection, may be necessary for specific applications.</p> <p><strong>Key Features</strong>:</p> <ul> <li><strong>Temporal Coverage</strong>: Over 20 years of data.</li> <li><strong>Geographical Coverage</strong>: Vienna, Lower Austria, and Burgenland.</li> <li><strong>Tests Included</strong>: <ul> <li>Particle Size Distribution</li> <li>Atterberg Limits</li> <li>Proctor Tests</li> <li>Permeability Tests</li> <li>Direct Shear Tests</li> </ul> </li> <li><strong>Number of Variables</strong>: 24</li> <li><strong>Potential Applications</strong>: Correlation analysis, predictive modeling, and geotechnical design.</li> </ul> <p><strong>Technical Details</strong>:</p> <ul> <li>Missing values have been addressed using K-Nearest Neighbors (KNN) imputation, and anomalies identified using Local Outlier Factor (LOF) methods in previous studies.</li> <li>Data normalization and standardization steps are recommended for specific analyses.</li> </ul> <p><strong>Acknowledgments</strong>:<br>The dataset was compiled with support from the European Union's MSCA Staff Exchanges project 101182689 Geotechnical Resilience through Intelligent Design (GRID).</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Network Digital Twin-Generated Dataset for Machine Learning-based Detection of Benign and Malicious Heavy Hitter Flows

<h3>Overview</h3> <p>This record provides a dataset created as part of the study presented in the following publication and is made <strong>publicly available for research purposes</strong>. The associated article provides a comprehensive description of the dataset, its structure, and the methodology used in its creation. If you use this dataset, please <strong>cite the following article </strong>published in the journal <strong>IEEE Communications Magazine</strong>:</p> <blockquote> <p><strong>A. Karamchandani, J. Nunez, L. de-la-Cal, Y. Moreno, A. Mozo, and A. Pastor, &ldquo;On the Applicability of Network Digital Twins in Generating Synthetic Data for Heavy Hitter Discrimination,&rdquo; IEEE Communications Magazine, pp. 2&ndash;8, 2025, DOI: 10.1109/MCOM.003.2400648.</strong></p> </blockquote> <p>More specifically, the record contains several synthetic datasets generated to differentiate between benign and malicious heavy hitter flows within a realistic virtualized network environment. Heavy Hitter flows, which include high-volume data transfers, can significantly impact network performance, leading to congestion and degraded quality of service. Distinguishing legitimate heavy hitter activity from malicious Distributed Denial-of-Service traffic is critical for network management and security, yet existing datasets lack the granularity needed for training machine learning models to effectively make this distinction.</p> <p>To address this, a Network Digital Twin (NDT) approach was utilized to emulate realistic network conditions and traffic patterns, enabling automated generation of labeled data for both benign and malicious HH flows alongside regular traffic.</p> <h3>Feature Set:</h3> <p>The feature set includes the following flow statistics commonly used in the literature on network traffic classification:</p> <ul> <li>The protocol used for the connection, identifying whether it is TCP, UDP, ICMP, or OSPF.</li> <li>The time (relative to the connection start) of the most recent packet sent from source to destination at the time of each snapshot.</li> <li>The time (relative to the connection start) of the most recent packet sent from destination to source at the time of each snapshot.</li> <li>The cumulative count of data packets sent from source to destination at the time of each snapshot.</li> <li>The cumulative count of data packets sent from destination to source at the time of each snapshot.</li> <li>The cumulative bytes sent from source to destination at the time of each snapshot.</li> <li>The cumulative bytes sent from destination to source at the time of each snapshot.</li> <li>The time difference between the first packet sent from source to destination and the first packet sent from destination to source.</li> </ul> <h3>Dataset Variations:</h3> <p>To accommodate diverse research needs and scenarios, the dataset is provided in the following variations:</p> <ol> <li> <p><strong><code>All at Once</code></strong>:</p> <ol> <li>Contains a synthetic dataset where all traffic types, including benign, normal, and malicious DDoS heavy hitter (HH) flows, are combined into a single dataset.</li> <li>This version represents a holistic view of the traffic environment, simulating real-world scenarios where all traffic occurs simultaneously.</li> </ol> </li> <li> <p><strong><code>Balanced Traffic Generation</code></strong>:</p> <ol> <li>Represents a balanced traffic dataset with an equal proportion of benign, normal, and malicious DDoS traffic.</li> <li>Designed for scenarios where a balanced dataset is needed for fair training and evaluation of machine learning models.</li> </ol> </li> <li> <p><strong><code>DDoS at Intervals</code></strong>:</p> <ol> <li>Contains traffic data where malicious DDoS HH traffic occurs at specific time intervals, mimicking real-world attack patterns.</li> <li>Useful for studying the impact and detection of intermittent malicious activities.</li> </ol> </li> <li> <p><strong><code>Only Benign HH Traffic</code></strong>:</p> <ol> <li>Includes only benign HH traffic flows.</li> <li>Suitable for training and evaluating models to identify and differentiate benign heavy hitter traffic patterns.</li> </ol> </li> <li> <p><strong><code>Only DDoS Traffic</code></strong>:</p> <ol> <li>Contains only malicious DDoS HH traffic.</li> <li>Helps in isolating and analyzing attack characteristics for targeted threat detection.</li> </ol> </li> <li> <p><strong><code>Only Normal Traffic</code></strong>:</p> <ol> <li>Comprises only regular, non-HH traffic flows.</li> <li>Useful for understanding baseline network behavior in the absence of heavy hitters.</li> </ol> </li> <li> <p><strong><code>Unbalanced Traffic Generation</code></strong>:</p> <ol> <li>Features an unbalanced dataset with varying proportions of benign, normal, and malicious traffic.</li> <li>Simulates real-world scenarios where certain types of traffic dominate, providing insights into model performance in unbalanced conditions.</li> </ol> </li> </ol> <p>For each variation, the output of the different packet aggregators is provided separated in its respective folder.</p> <p>Each variation was generated using the NDT approach to demonstrate its flexibility and ensure the reproducibility of our study's experiments, while also contributing to future research on network traffic patterns and the detection and classification of heavy hitter traffic flows. The dataset is designed to support research in network security, machine learning model development, and applications of digital twin technology.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior

<p>The datasets used in the paper "Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior". A detailed description of these datasets is available at https://github.com/jaydu1/VITAE/tree/master/data.</p>

opencc-by-4.0Dec 2020View details →
zenodo48/100

LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models

<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p>&nbsp;</p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 &Aring;. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 &Aring; gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines.&nbsp;</p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP&sup3; Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Experimental data for the motor learning study performed: "Towards functional robotic training: Motor learning of dynamic tasks is enhanced by haptic rendering but hampered by robotic assistance"

<p>The dataset contains the kinematic data and the questionnaire responses for a robot-assisted motor learning study performed in the Motor Learning and Neurorehabilitation Laboratory at the University of Bern. The details of the study are&nbsp;described in [doi: ]. The kinematic data for each participant is stored as a data frame inside a &ldquo;pickle&rdquo; (serialized python object) file. The questionnaire responses and population metrics&nbsp;are stored as&nbsp;&ldquo;CSV&rdquo; files. The variables inside the files are explained in &ldquo;DataframeVariableDescription.rtf&rdquo;. For questions, please contact oezhan.oezen@artorg.unibe.ch or L.MarchalCrespo@tudelft.nl.</p>

opencc-by-4.0Jul 2021View details →
zenodo48/100

Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study

<p>Efficiency is essential to support responsiveness w.r.t. ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code that supports symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development tends to produce DL code that is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, less error-prone imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. While hybrid approaches aim for the &quot;best of both worlds,&quot; the challenges in applying them in the real world are largely unknown. We conduct a data-driven analysis of challenges&mdash;and resultant bugs&mdash;involved in writing reliable yet performant imperative DL code by studying 250 open-source projects, consisting of 19.7 MLOC, along with 470 and 446 manually examined code patches and bug reports, respectively. The results indicate that hybridization: (i) is prone to API misuse, (ii) can result in performance degradation&mdash;the opposite of its intention, and (iii) has limited application due to execution mode incompatibility. We put forth several recommendations, best practices, and anti-patterns for effectively hybridizing imperative DL code, potentially benefiting DL practitioners, API designers, tool developers, and educators.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

On the Effectiveness of Transfer Learning for Code Search - Replication Package

<p>This repository represents the replication package for the paper <em>On the Effectiveness of Transfer Learning for Code Search</em>.</p> <p>The paper is published in&nbsp;the journal&nbsp;<em>IEEE Transactions on Software Engineering (TSE)</em>.</p> <p>In this replication package, we provide all the data and scripts we used in our study.</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)

<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1&nbsp;contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE).&nbsp;</p> <p>&nbsp;</p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called &quot;Sentinel2LULC_GeoTiff.zip&quot;&nbsp;contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called &quot;Sentinel2LULC_JPEG.zip&quot; contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called &quot;Sentinel2LULC_CSV.zip&quot; includes 29 zip-compressed CSV files with as many rows as provided images and with 12&nbsp;columns containing the following metadata (this same metadata is provided in the image filenames):&nbsp;</p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class&nbsp;</li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products&nbsp;</li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is&nbsp;the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing&nbsp;and exporting&nbsp;its corresponding&nbsp;image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class&nbsp;</li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with &quot;including_non_downloaded_images&quot;.</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset,&nbsp; we&nbsp;included in the version v2.1, &nbsp;a compressed folder called &quot;Geographic_Representativeness.zip&quot;. This zip-compressed folder&nbsp;contains a csv file&nbsp;for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>&copy;&nbsp;<a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset&nbsp;</a>by&nbsp;&nbsp;Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujem&acirc;a Achchab, Francisco Herrera &amp; Siham Tabik&nbsp;is marked with Attribution 4.0 International&nbsp;(CC-BY 4.0)</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record