Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

672

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

672 results for “Logs”

Learn how ShareScore rates datasets ↗
zenodo40/100

The Resolution of Keller's Conjecture - Computation Logs

<p>Logs of the computations performed to settle Keller&#39;s conjecture on cube tilings. For every value of s∊{3,4,6} there are files:</p> <ul> <li>s?.cnf encoding the problem for that value of s as explained in the paper, with additional symmetry breaking clauses.</li> <li>s?.dnf which is a tautology of assignments that need to be verified by SAT solvers.</li> </ul> <p>Inside the Keller-logs.zip archive, for every value of s∊{3,4,6} you will find:</p> <ul> <li>A folder sym-s? containing all the verification logs of the symmetry breaking clauses added to the original encoding of the problem for that value of s as a SAT instance in order to obtain s?.cnf.</li> <li>A folder unsat-s? containing one log file for reach assignment in s?.dnf.</li> </ul>

opencc-by-4.0Apr 2020View details →
zenodo40/100

AIT Log Data Set V1.1

<p><strong>AIT Log Data Sets</strong></p> <p>This repository contains synthetic log data suitable for evaluation of intrusion detection systems. The logs were collected from four independent testbeds that were built at the Austrian Institute of Technology (AIT) following the approach by Landauer et al. (2020) [1]. Please refer to the paper for more detailed information on automatic testbed generation and cite it if the data is used for academic publications. In brief, each testbed simulates user accesses to a webserver that runs Horde Webmail and OkayCMS. The duration of the simulation is six days. On the fifth day (2020-03-04) two attacks are launched against each web server.</p> <p>The archive AIT-LDS-v1_0.zip contains the directories &quot;data&quot; and &quot;labels&quot;.</p> <p>The data directory is structured as follows. Each directory mail.&lt;name&gt;.com contains the logs of one web server. Each directory user-&lt;ID&gt; contains the logs of one user host machine, where one or more users are simulated. Each file log&lt;UID&gt;.log in the user-&lt;ID&gt; directories contains the activity logs of one particular user.</p> <p>Setup details of the web servers:</p> <ul> <li>OS: Debian Stretch 9.11.6</li> <li>Services: <ul> <li>Apache2</li> <li>PHP7</li> <li>Exim 4.89</li> <li>Horde 5.2.22</li> <li>OkayCMS 2.3.4</li> <li>Suricata</li> <li>ClamAV</li> <li>MariaDB</li> </ul> </li> </ul> <p>Setup details of user machines:</p> <ul> <li>OS: Ubuntu Bionic</li> <li>Services: <ul> <li>Chromium</li> <li>Firefox</li> </ul> </li> </ul> <p>User host machines are assigned to web servers in the following way:</p> <ul> <li>mail.cup.com is accessed by users from host machines user-{0, 1, 2, 6}</li> <li>mail.spiral.com is accessed by users from host machines user-{3, 5, 8}</li> <li>mail.insect.com is accessed by users from host machines user-{4, 9}</li> <li>mail.onion.com is accessed by users from host machines user-{7, 10}</li> </ul> <p>The following attacks are launched against the web servers (different starting times for each web server, please check the labels for exact attack times):</p> <ul> <li>Attack 1: multi-step attack with sequential execution of the following attacks: <ul> <li>nmap scan</li> <li>nikto scan</li> <li>smtp-user-enum tool for account enumeration</li> <li>hydra brute force login</li> <li>webshell upload through Horde exploit (CVE-2019-9858)</li> <li>privilege escalation through Exim exploit (CVE-2019-10149)</li> </ul> </li> <li>Attack 2: webshell injection through malicious cookie (CVE-2019-16885)</li> </ul> <p>Attacks are launched from the following user host machines. In each of the corresponding directories user-&lt;ID&gt;, logs of the attack execution are found in the file attackLog.txt:</p> <ul> <li>user-6 attacks mail.cup.com</li> <li>user-5 attacks mail.spiral.com</li> <li>user-4 attacks mail.insect.com</li> <li>user-7 attacks mail.onion.com</li> </ul> <p>The log data collected from the web servers includes</p> <ul> <li>&nbsp;Apache access and error logs</li> <li>&nbsp;syscall logs collected with the Linux audit daemon</li> <li>&nbsp;suricata logs</li> <li>&nbsp;exim logs</li> <li>&nbsp;auth logs</li> <li>&nbsp;daemon logs</li> <li>&nbsp;mail logs</li> <li>&nbsp;syslogs</li> <li>&nbsp;user logs</li> </ul> <p>&nbsp;<br> Note that due to their large size, the audit/audit.log files of each server were compressed in a .zip-archive. In case that these logs are needed for analysis, they must first be unzipped.<br> &nbsp;<br> Labels are organized in the same directory structure as logs. Each file contains two labels for each log line separated by a comma, the first one based on the occurrence time, the second one based on similarity and ordering. Note that this does not guarantee correct labeling for all lines and that no manual corrections were conducted.</p> <p>Version history and related data sets:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.3723083">AIT-LDS-v1.0</a>: Four datasets, logs from single host, fine-granular audit logs, mail/CMS. <ul> <li><a href="https://doi.org/10.5281/zenodo.4264796">AIT-LDS-v1.1</a>: Removed carriage return of line endings in audit.log files.</li> </ul> </li> <li><a href="http://doi.org/10.5281/zenodo.5789064">AIT-LDS-v2.0</a>: Eight datasets, logs from all hosts, system logs and network traffic, mail/CMS/cloud/web.</li> </ul> <p>Acknowledgements: Partially funded by the FFG projects INDICAETING (868306) and DECEPT (873980), and the EU project GUARD (833456).</p> <p><strong>If you use the dataset, please cite the following publication:</strong></p> <p>[1] M. Landauer, F. Skopik, M. Wurzenberger, W. Hotwagner and A. Rauber, <a href="https://ieeexplore.ieee.org/document/9262078">&quot;Have it Your Way: Generating Customized Log Datasets With a Model-Driven Simulation Testbed,&quot;</a> in IEEE Transactions on Reliability, vol. 70, no. 1, pp. 402-415, March 2021, doi: 10.1109/TR.2020.3031317. [<a href="https://www.skopik.at/ait/2020_trel.pdf">PDF</a>]</p>

opencc-by-nc-sa-4.0Nov 2020View details →
zenodo40/100

Techmaster event log

<p>Techmaster log for the replication of experiments described in&nbsp;Context-Aware Process Performance Indicator Prediction in IEEE Access 2020.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

IODP Expedition 368X Piece log

<p>Dataset includes length data for every whole-round piece: bin length, whole-round piece length (measured by curation staff), and both the archive- and working-half piece lengths (optionally measured by scientists).</p>

opencc-zeroJan 2021View details →
zenodo40/100

IODP Expedition 366 Piece log

<p>Dataset includes length data for every whole-round piece: bin length, whole-round piece length (measured by curation staff), and both the archive- and working-half piece lengths (optionally measured by scientists).</p>

opencc-zeroMar 2020View details →
zenodo40/100

solar_home_system_data_log: Initial release of data sets and script

<p>In this first release, this repository includes three sets of data (date/time, temperature, current, and voltage) of over 6 months of electricity consumption of three households in an off-grid area in the state of Jharkhand, India. The goal of this data set is to be made open so that the community working in off-grid electricity access can get a sense of electricity consumption patterns and apply various analytical and visualization techniques.</p>

openother-openAug 2016View details →
zenodo40/100

System Call Logs with Natural Random Faults: Experimental Design and Application

<p>In this work we present a large dataset for use in embedded systems research which has been gathered from a realistic development environment operating in the path of an accelerated neutron beam. The dataset contains traces with events from a real-time operating system as as user events from a safety critical application. All data is carefully timestamped and in human understandable form.</p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

Digitized Copies of the Zurich Overnight Visitor Logs ("Nachtzedel") after 1780

<div> <div>This dataset contains digital copies - image files and metadata - of the "Z&uuml;rcher Nachtzedel" (overnight visitor logs = Fremdenliste, 1780 to 1784).</div> </div>

opencc-by-4.0Aug 2024View details →
dryad40/100

Data for: Temporal variations in female moose responses to roads and logging in the absence of wolves

<p>Animal movements, needed to acquire food resources, avoid predation risk, and find breeding partners, are influenced by annual and circadian cycles. Decisions related to movement reflect a quest to maximize benefits while limiting costs, especially in heterogeneous landscapes. Predation by wolves (<em>Canis lupus</em>) has been identified as the major driver of moose (<em>Alces alces</em>) habitat selection patterns, and linear features have been shown to increase wolf efficiency to travel, hunt and kill prey. However, few studies have described moose behavioral response to roads and logging in Canada in the absence of wolves. We thus characterized temporal changes (i.e., day phases and biological periods) in eastern moose (<em>Alces alces americana</em>) habitat selection and space use patterns near a road network in a wolf-free area located south of the St. Lawrence River (eastern Canada). We used telemetry data collected on 18 females between 2017 and 2019 to build resource selection functions and mixed linear regressions to explain variations in habitat selection patterns, home-range size and movement rates. Female moose selected forest stands providing forage when movement was not impeded by snow cover (i.e., spring/green-up, summer/rearing, fall/rut) and stands offering protection against incidental predation during calving. In winter, home-range size decreased with an increasing proportion of stands providing food and shelter against harsh weather, limiting the energetic costs associated with movement. Our results reaffirmed the year-round aversive effect of roads, even in the absence of wolves, but the magnitude of this avoidance differed between day phases, being lower during the "dusk-night-dawn" phase, perhaps due to a lower level of human activity on and near roads. Female moose behavior in our study area was similar to what was observed in landscapes where moose and wolves cohabit, suggesting that the risk associated with humans, perceived as another type of predator, and with incidental predators (coyote <em>Canis latrans</em>,<em> </em>black bear <em>Ursus americanus</em>), equates that of wolf predation in heavily managed landscapes.</p>

opencc-zeroJan 2024View details →
zenodo40/100

Processed Datasets - Imputation in Well Log Data: A Benchmark

<p>Imputation of well log data is a common task in the field. However a quick review of the literature reveals a lack of padronization when evaluating methods for the problem.&nbsp;The goal of the benchmark is to introduce a standard evaluation protocol to any imputation method for well log data.&nbsp;</p> <p>In the proposed benchmark, three public datasets are used:</p> <ul> <li><strong>Geolink:</strong> The Geolink Dataset is another public dataset of wells in the Norwegian offshore. The data is provided by the company of the same name,&nbsp;<a href="https://www.geolink-s2.com/" target="_blank" rel="noopener">GEOLINK</a> and follows the NOLD 2.0 license. <br>This dataset contains a total of 223 wells. It also has lithology labels for the wells with a total of 36 lithology classes. [<a href="https://drive.google.com/drive/folders/1EgDN57LDuvlZAwr5-eHWB5CTJ7K9HpDP" target="_blank" rel="noopener">download original</a>]</li> <li><strong>Taranaki Basin:</strong> The Taranaki Basin Dataset is a curated set of wells and a convenient option for experimentation especially due to it is ease of accessibility and use.<br>This collection, under the CDLA-Sharing-1.0 license, contains well logs extracted from the <a href="https://geodata.nzpam.govt.nz/" target="_blank" rel="noopener">New Zealand Petroleum &amp; Minerals Online Exploration Database</a> and&nbsp;<a href="http://pet.gns.cri.nz/" target="_blank" rel="noopener">Petlab</a>.<br>There are a total of 407 wells, of which 289 are onshore and 118 are offshore exploration and production wells. [<a href="https://developer.ibm.com/exchanges/data/all/taranaki-basin-curated-well-logs/" target="_blank" rel="noopener">download original</a>]</li> <li><strong>Teapot Dome:</strong> The Teapot Dome dataset is provided by the Rocky Mountain Oilfield Testing Center (RMOTC) and the US Department of Energy.<br>It contains different types of data related to the Teapot Dome oil field, such as 2D and 3D seismic data, well logs, and GIS data. The data is licensed under the Creative Commons 4.0 license. <br>In total, the dataset has 1,179 wells with available logs. The number of available logs varies across wells. There are only 91 wells with the gamma ray, bulk density, and neutron porosity logs, while only three wells have the complete basic suite. [<a href="http://s3.amazonaws.com/open.source.geoscience/open_data/teapot/rmotc.tar" target="_blank" rel="noopener">direct download</a>]</li> </ul> <p>Here you can download all three datasets already preprocessed to be used with our implementation, found <a href="https://github.com/uai-ufmg/well-log-imputation" target="_blank" rel="noopener">here</a>.</p> <p>&nbsp;</p> <h3>File Description:</h3> <p>There are six files for each fold partition for each dataset.</p> <ul> <li><code><em>datasetname_fold_k_well_log_metadata_train.json </em></code>: JSON file with general information of the slices of <strong>training </strong>partition of the fold <strong>k</strong>. Contains total number of slices and the number of slices per well.<em>&nbsp;&nbsp;</em></li> <li><em><code>datasetname_fold_k_well_log_metadata_val.json</code> </em>: JSON file with general information of the slices of <strong>validation </strong>partition of the fold&nbsp;<strong>k</strong>. Contains total number of slices and the number of slices per well.&nbsp;</li> <li><em><code>datasetname_fold_k_well_log_slices_train.npy</code>: </em>.npy (numpy) file ready to be loaded with the slices for <strong>training </strong>of the fold&nbsp;<strong>k </strong>already processed. When loaded<em> </em>should have shape of<em> (total_slices, 256, number_of_logs)</em></li> <li><em><code>datasetname_fold_k_well_log_slices_val.npy</code>&nbsp;</em>:&nbsp; .npy (numpy) file ready to be loaded with the slices for <strong>validation </strong>of the fold&nbsp;<strong>k </strong>already processed.</li> <li><em><code>datasetname_fold_k_well_log_slices_meta_train.json</code> :&nbsp;</em>JSON file with the slices info for all slices in the <strong>training </strong>partition of the fold <strong>k</strong>. For each slice, 7 data points are provided, the last four are discarded (it would contain other information that was not used). The first three are in order the: origin well name, the starting position in that well, and the end position of the slice in that well.</li> <li><em><code>datasetname_fold_k_well_log_slices_meta_val.json</code> </em>: JSON file with the slices info for all slices in the <strong>validation </strong>partition of the fold <strong>k</strong>.</li> </ul>

opencc-by-4.0Apr 2024View details →
dryad40/100

Investigating cooccurrence patterns and dynamics for many imperfectly detected species, using a log-linear modelling parameterisation

<p>1. Patterns in, and the underlying dynamics of, species cooccurrence is of interest in many ecological applications. Unaccounted for, imperfect detection of the species can lead to misleading inferences about the nature and magnitude of any interaction. A range of different parameterisations have been published that could be used with the same fundamental modelling framework that accounts for imperfect detection, although each parameterisation has different advantages and disadvantages.</p> <p>2. We propose a parameterisation based on log-linear modelling that does not require a species hierarchy to be defined (in terms of dominance), and enables a numerically robust approach for estimating covariate effects.</p> <p>3. Conceptually the parameterisation is equivalent to using the presence of species in the current, or a previous, time period as predictor variables for the current occurrence of other species. This leads to natural, 'symmetric', interpretations of parameter estimates.</p> <p>4. The parameterisation can be applied to many species, in either a maximum-likelihood or Bayesian estimation framework. We illustrate the method using camera trapping data collected on three mesocarnivore species in South Texas.</p>

opencc-zeroApr 2022View details →
zenodo40/100

UIS Log: Synthetic User Interface with Screenshots Log

<p>These data correspond to the set of problems for evaluating the proposal detailed in Mart&iacute;nez-Rojas et al. 2022.&nbsp;The evaluation utilizes a set of synthetic problems that simulate realistic administrative use cases. Each problem includes a UI Log with a synthetic screenshot corresponding to each event, capturing 3 distinct processes (<em>P</em>) marked by varying complexity levels. These levels are defined by the number of activities, the process execution variants, and the visual features influencing decisions between these variants.</p> <p>The implementation of this proposal can be found in the tool available at&nbsp;<a href="https://github.com/RPA-US/screenrpa" target="_new">this GitHub repository</a>, which utilizes the logs of these 3 processes for validation. Here they are described:<br><br></p> <ul> <li><em>P1 Client creation</em>. A process with&nbsp;<strong>5 activities and 2 variants</strong>. The single decision in this process is made based on the existence of an attachment in the reception email.</li> <li><em>P2 Client validation</em>. A process with&nbsp;<strong>7 activities and 2 variants</strong>. The decision is made based on the user&rsquo;s response to a query.</li> <li><em>P3 Client deletion</em>. A process with&nbsp;<strong>7 activities and 4 variants</strong>. The decisions are made based on two conditions: (1) the existence of pending invoices and (2) the existence of an attachment to justify the payment of the invoices.</li> </ul> <p>These processes all contain a single decision point, although the one in P3 is complex. All processes include</p> <ol> <li>synthetic screen captures for their activities and</li> <li>a sample event log with a single instance for each variant.</li> </ol> <p>To generate the objects for the valuation, we generate event logs of different sizes (|<em>L</em>|) for each of these processes by deriving events from the sample event log. We consider log sizes in the range of {10, 25, 50, 100} events. Note that we consider complete instances in the log and thus, we remove the last instance if it goes beyond |<em>L</em>|.<br>Some of these logs are generated with a balanced number of instances, while others are unbalanced (<em>B</em>?) which present more than 20% of different frequency between the most frequent and less frequent variants. To average the result over a collection of problems, 30 instances are randomly generated for each tuple &lt;&nbsp;<em>P</em>, |<em>L</em>|,&nbsp;<em>B</em>? &gt;.<br>In this dataset there are 3 zips, one for each family. Each family corresponds to a process:</p> <ul> <li><em>Basic </em>corresponds to P1</li> <li><em>Intermediate </em>corresponds to P12</li> <li><em>Advanced </em>corresponds to P3</li> </ul> <p>Within these folders, we find 30 different scenarios (folder), in which the look and feel of the applications present in the screenshots have suffered little variations. Within each of these scenarios, variations are carried out respecting to the data entered in the forms and the images or attachments present in the user interface to generate log instances depending on the characteristics of each process.<br>For each scenario, we find 8 folders with the concrete problem which is defined by Log_size (in {10,25,50,100}) and Balanced (in {Balanced, Unbalanced}). The name of these folders have this format<em>: Family_LogSize_Balanced.</em><br>Inside each problem folder the UI Log and the screen captures can be found.<br><br><strong>References</strong><br><br>Mart&iacute;nez-Rojas, A., Jim&eacute;nez-Ram&iacute;rez, A., Enr&iacute;quez, J. G., &amp; Reijers, H. A. (2022, September). Analyzing variable human actions for robotic process automation. In <em>International Conference on Business Process Management</em> (pp. 75-90). Cham: Springer International Publishing.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Kyoushi Log Data Set

<p>This repository contains synthetic log data suitable for evaluation of intrusion detection systems. The logs were collected from a testbed that was built at the Austrian Institute of Technology (AIT) following the approaches by [1], [2], and [3]. Please refer to these papers for more detailed information on the dataset and cite them if the data is used for academic publications. Other than the related <a href="https://zenodo.org/record/4264796">AIT-LDSv1.1</a>, this dataset involves a more complex network structure, makes use of a different attack scenario, and collects log data from multiple hosts in the network. In brief, the testbed simulates a small enterprise network including mail server, file share, WordPress server, VPN, firewall, etc. Normal user behavior is simulated to generate background noise. After some days, two attack scenarios are launched against the network. Note that the <a href="https://zenodo.org/record/5789064">AIT-LDSv2.0</a> extends this dataset with additional attack cases and variations of attack parameters.</p> <p>The archives have the following structure. The <em>gather </em>directory contains the raw log data from each host in the network, as well as their system configurations. The <em>labels </em>directory contains the ground truth for those log files that are labeled. The <em>processing</em> directory contains configurations for the labeling procedure and the <em>rules </em>directory contains the labeling rules. Labeling of events that are related to the attacks is carried out with the <a href="https://github.com/ait-aecid/kyoushi-dataset">Kyoushi Labeling Framework</a>.</p> <p>Each dataset contains traces of a specific attack scenario:</p> <ul> <li>Scenario 1 (see <em>gather/attacker_0/logs/sm.log</em> for detailed attack log): <ul> <li>nmap scan</li> <li>WPScan</li> <li>dirb scan</li> <li>webshell upload through wpDiscuz exploit (CVE-2020-24186)</li> <li>privilege escalation</li> </ul> </li> <li>Scenario 2 (see <em>gather/attacker_0/logs/dnsteal.log</em> for detailed attack log): <ul> <li>DNSteal data exfiltration</li> </ul> </li> </ul> <p>The log data collected from the servers includes</p> <ul> <li>Apache access and error logs (labeled)</li> <li>audit logs (labeled)</li> <li>auth logs (labeled)</li> <li>VPN logs (labeled)</li> <li>DNS logs (labeled)</li> <li>syslog</li> <li>suricata logs</li> <li>exim logs</li> <li>horde logs</li> <li>mail logs</li> </ul> <p>Note that only log files from affected servers are labeled. Label files and the directories in which they are located have the same name as their corresponding log file in the <em>gather </em>directory. Labels are in JSON format and comprise the following attributes: line (number of line in corresponding log file), labels (list of labels assigned to that log line), rules (names of labeling rules matching that log line). Note that not all attack traces are labeled in all log files; please refer to the labeling rules in case that some labels are not clear.</p> <p>Acknowledgements: Partially funded by the FFG projects INDICAETING (868306) and DECEPT (873980), and the EU project GUARD (833456).</p> <p><strong>If you use the dataset, please cite the following publications:</strong></p> <p>[1] <a href="https://ieeexplore.ieee.org/document/9262078">M. Landauer, F. Skopik, M. Wurzenberger, W. Hotwagner and A. Rauber, &quot;Have it Your Way: Generating Customized Log Datasets With a Model-Driven Simulation Testbed,&quot; in IEEE Transactions on Reliability, vol. 70, no. 1, pp. 402-415, March 2021, doi: 10.1109/TR.2020.3031317.</a></p> <p>[2] <a href="https://dl.acm.org/doi/10.1145/3510547.3517924">M. Landauer, M. Frank, F. Skopik, W. Hotwagner, M. Wurzenberger, and A. Rauber, &quot;A Framework for Automatic Labeling of Log Datasets from Model-driven Testbeds for HIDS Evaluation&quot;. ACM Workshop on Secure and Trustworthy Cyber-Physical Systems (ACM SaT-CPS 2022), April 27, 2022, Baltimore, MD, USA. ACM.</a></p> <p>[3] <a href="https://repositum.tuwien.at/handle/20.500.12708/17804">M. Frank, &quot;Quality improvement of labels for model-driven benchmark data generation for intrusion detection systems&quot;, Master&#39;s Thesis, Vienna University of Technology, 2021. </a></p>

opencc-by-nc-sa-4.0Dec 2021View details →
zenodo40/100

Compositional discovery of architecture-aware and sound process models from event logs of multi-agent systems: experimental data.

<p>This repository contains the experimental data used for the evaluation of the compositional approach to the discovery of process models from event logs of multi-agent systems, where agents interact according to specific patterns of synchronous and asynchronous interactions.</p> <p>According to the experiment plan, there is the folder for each interface pattern containing:</p> <ol> <li>The reference model (Petri net encoded in PNML-file)</li> <li>The event log obtained by simulating the behavior of the reference model (XES-file)</li> <li>The model discovered directly from the generated event log (Petri net encoded in PNML-file)</li> <li>The model discovered by composing the agent model w.r.t. the interface pattern (Petri net encoded in&nbsp;PNML-file)</li> </ol>

opencc-by-4.0May 2021View details →
zenodo40/100

SHP event log

<p>SHP log for the replication of experiments described in&nbsp;Context-Aware Process Performance Indicator Prediction in IEEE Access 2020.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

AIT Log Data Set V2.0

<p><strong>AIT Log Data Sets</strong></p> <p>This repository contains synthetic log data suitable for evaluation of intrusion detection systems, federated learning, and alert aggregation. A detailed description of the dataset is available in [1]. The logs were collected from eight testbeds that were built at the Austrian Institute of Technology (AIT) following the approach by [2]. Please cite these papers if the data is used for academic publications.</p> <p>In brief, each of the datasets corresponds to a testbed representing a small enterprise network including mail server, file share, WordPress server, VPN, firewall, etc. Normal user behavior is simulated to generate background noise over a time span of 4-6 days. At some point, a sequence of attack steps is launched against the network. Log data is collected from all hosts and includes Apache access and error logs, authentication logs, DNS logs, VPN logs, audit logs, Suricata logs, network traffic packet captures, horde logs, exim logs, syslog, and system monitoring logs. Separate ground truth files are used to label events that are related to the attacks. Compared to the <a href="../record/4264796">AIT-LDSv1.1</a>, a more complex network and diverse user behavior is simulated, and logs are collected from all hosts in the network. If you are only interested in network traffic analysis, we also provide the <a href="../record/6610489">AIT-NDS</a> containing the labeled netflows of the testbed networks. We also provide the <a href="../records/8263181">AIT-ADS</a>, an alert data set derived by forensically applying open-source intrusion detection systems on the log data.</p> <p>The datasets in this repository have the following structure:</p> <ul> <li>The <em>gather </em>directory contains all logs collected from the testbed. Logs collected from each host are located in <em>gather/&lt;host_name&gt;/logs/</em>.</li> <li>The <em>labels </em>directory contains the ground truth of the dataset that indicates which events are related to attacks. The directory mirrors the structure of the gather directory so that each label files is located at the same path and has the same name as the corresponding log file. Each line in the label files references the log event corresponding to an attack by the line number counted from the beginning of the file ("line"), the labels assigned to the line that state the respective attack step ("labels"), and the labeling rules that assigned the labels ("rules"). An example is provided below.</li> <li>The <em>processing </em>directory contains the source code that was used to generate the labels.</li> <li>The <em>rules </em>directory contains the labeling rules.</li> <li>The <em>environment </em>directory contains the source code that was used to deploy the testbed and run the simulation using the <a href="https://github.com/ait-aecid/kyoushi-environment">Kyoushi Testbed Environment</a>.</li> <li>The <em>dataset.yml</em> file specifies the start and end time of the simulation.</li> </ul> <p>The following table summarizes relevant properties of the datasets:</p> <ul> <li>fox <ul> <li>Simulation time: 2022-01-15 00:00 - 2022-01-20 00:00</li> <li>Attack time: 2022-01-18 11:59 - 2022-01-18 13:15</li> <li>Scan volume: High</li> <li>Unpacked size: 26 GB</li> </ul> </li> <li>harrison <ul> <li>Simulation time: 2022-02-04 00:00 - 2022-02-09 00:00</li> <li>Attack time: 2022-02-08 07:07 - 2022-02-08 08:38</li> <li>Scan volume: High</li> <li>Unpacked size: 27 GB</li> </ul> </li> <li>russellmitchell <ul> <li>Simulation time: 2022-01-21 00:00 - 2022-01-25 00:00</li> <li>Attack time: 2022-01-24 03:01 - 2022-01-24 04:39</li> <li>Scan volume: Low</li> <li>Unpacked size: 14 GB</li> </ul> </li> <li>santos <ul> <li>Simulation time: 2022-01-14 00:00 - 2022-01-18 00:00</li> <li>Attack time: 2022-01-17 11:15 - 2022-01-17 11:59</li> <li>Scan volume: Low</li> <li>Unpacked size: 17 GB</li> </ul> </li> <li>shaw <ul> <li>Simulation time: 2022-01-25 00:00 - 2022-01-31 00:00</li> <li>Attack time: 2022-01-29 14:37 - 2022-01-29 15:21</li> <li>Scan volume: Low</li> <li>Data exfiltration is not visible in DNS logs</li> <li>Unpacked size: 27 GB</li> </ul> </li> <li>wardbeck <ul> <li>Simulation time: 2022-01-19 00:00 - 2022-01-24 00:00</li> <li>Attack time: 2022-01-23 12:10 - 2022-01-23 12:56</li> <li>Scan volume: Low</li> <li>Unpacked size: 26 GB</li> </ul> </li> <li>wheeler <ul> <li>Simulation time: 2022-01-26 00:00 - 2022-01-31 00:00</li> <li>Attack time: 2022-01-30 07:35 - 2022-01-30 17:53</li> <li>Scan volume: High</li> <li>No password cracking in attack chain</li> <li>Unpacked size: 30 GB</li> </ul> </li> <li>wilson <ul> <li>Simulation time: 2022-02-03 00:00 - 2022-02-09 00:00</li> <li>Attack time: 2022-02-07 10:57 - 2022-02-07 11:49</li> <li>Scan volume: High</li> <li>Unpacked size: 39 GB</li> </ul> </li> </ul> <p>The following attacks are launched in the network:</p> <ul> <li>Scans (nmap, WPScan, dirb)</li> <li>Webshell upload (CVE-2020-24186)</li> <li>Password cracking (John the Ripper)</li> <li>Privilege escalation</li> <li>Remote command execution</li> <li>Data exfiltration (DNSteal)</li> </ul> <p>Note that attack parameters and their execution orders vary in each dataset. Labeled log files are trimmed to the simulation time to ensure that their labels (which reference the related event by the line number in the file) are not misleading. Other log files, however, also contain log events generated before or after the simulation time and may therefore be affected by testbed setup or data collection. It is therefore recommended to only consider logs with timestamps within the simulation time for analysis.</p> <p>The structure of labels is explained using the audit logs from the intranet server in the russellmitchell data set as an example in the following. The first four labels in the <em>labels/intranet_server/logs/audit/audit.log</em> file are as follows:</p> <blockquote> <p>{"line": 1860, "labels": ["attacker_change_user", "escalate"], "rules": {"attacker_change_user": ["attacker.escalate.audit.su.login"], "escalate": ["attacker.escalate.audit.su.login"]}}</p> <p>{"line": 1861, "labels": ["attacker_change_user", "escalate"], "rules": {"attacker_change_user": ["attacker.escalate.audit.su.login"], "escalate": ["attacker.escalate.audit.su.login"]}}</p> <p>{"line": 1862, "labels": ["attacker_change_user", "escalate"], "rules": {"attacker_change_user": ["attacker.escalate.audit.su.login"], "escalate": ["attacker.escalate.audit.su.login"]}}</p> <p>{"line": 1863, "labels": ["attacker_change_user", "escalate"], "rules": {"attacker_change_user": ["attacker.escalate.audit.su.login"], "escalate": ["attacker.escalate.audit.su.login"]}}</p> </blockquote> <p>Each JSON object in this file assigns a label to one specific log line in the corresponding log file located at <em>gather/intranet_server/logs/audit/audit.log</em>. The field "line" in the JSON objects specify the line number of the respective event in the original log file, while the field "labels" comprise the corresponding labels. For example, the lines in the sample above provide the information that lines 1860-1863 in the <em>gather/intranet_server/logs/audit/audit.log</em> file are labeled with "attacker_change_user" and "escalate" corresponding to the attack step where the attacker receives escalated privileges. Inspecting these lines shows that they indeed correspond to the user authenticating as root:</p> <blockquote> <p>type=USER_AUTH msg=audit(1642999060.603:2226): pid=27950 uid=33 auid=4294967295 ses=4294967295 msg='op=PAM:authentication acct="jhall" exe="/bin/su" hostname=? addr=? terminal=/dev/pts/1 res=success'</p> <p>type=USER_ACCT msg=audit(1642999060.603:2227): pid=27950 uid=33 auid=4294967295 ses=4294967295 msg='op=PAM:accounting acct="jhall" exe="/bin/su" hostname=? addr=? terminal=/dev/pts/1 res=success'</p> <p>type=CRED_ACQ msg=audit(1642999060.615:2228): pid=27950 uid=33 auid=4294967295 ses=4294967295 msg='op=PAM:setcred acct="jhall" exe="/bin/su" hostname=? addr=? terminal=/dev/pts/1 res=success'</p> <p>type=USER_START msg=audit(1642999060.627:2229): pid=27950 uid=33 auid=4294967295 ses=4294967295 msg='op=PAM:session_open acct="jhall" exe="/bin/su" hostname=? addr=? terminal=/dev/pts/1 res=success'</p> </blockquote> <p>The same applies to all other labels for this log file and all other log files. There are no labels for logs generated by "normal" (i.e., non-attack) behavior; instead, all log events that have no corresponding JSON object in one of the files from the <em>labels </em>directory, such as the lines 1-1859 in the example above, can be considered to be labeled as "normal". This means that in order to figure out the labels for the log data it is necessary to store the line numbers when processing the original logs from the <em>gather </em>directory and see if these line numbers also appear in the corresponding file in the <em>labels </em>directory.</p> <p>Beside the attack labels, a general overview of the exact times when specific attack steps are launched are available in <em>gather/attacker_0/logs/attacks.log</em>. An enumeration of all hosts and their IP addresses is stated in processing/config/servers.yml. Moreover, configurations of each host are provided in <em>gather/&lt;host_name&gt;/configs/</em> and <em>gather/&lt;host_name&gt;/facts.json</em>.</p> <p>Version history:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.3723082">AIT-LDS-v1.x</a>: Four datasets, logs from single host, fine-granular audit logs, mail/CMS.</li> <li><a href="http://doi.org/10.5281/zenodo.5789064">AIT-LDS-v2.0</a>: Eight datasets, logs from all hosts, system logs and network traffic, mail/CMS/cloud/web.</li> </ul> <p>Acknowledgements: Partially funded by the FFG projects INDICAETING (868306) and DECEPT (873980), and the EU projects GUARD (833456) and PANDORA (SI2.835928).</p> <p><strong>If you use the dataset, please cite the following publications:</strong></p> <p>[1] M. Landauer, F. Skopik, M. Frank, W. Hotwagner, M. Wurzenberger, and A. Rauber. <a href="https://ieeexplore.ieee.org/abstract/document/9866880">"Maintainable Log Datasets for Evaluation of Intrusion Detection Systems"</a>. IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 4, pp. 3466-3482, doi: 10.1109/TDSC.2022.3201582. [<a href="https://arxiv.org/pdf/2203.08580.pdf">PDF</a>]</p> <p>[2]&nbsp;M. Landauer, F. Skopik, M. Wurzenberger, W. Hotwagner and A. Rauber, <a href="https://ieeexplore.ieee.org/document/9262078">"Have it Your Way: Generating Customized Log Datasets With a Model-Driven Simulation Testbed,"</a> in IEEE Transactions on Reliability, vol. 70, no. 1, pp. 402-415, March 2021, doi: 10.1109/TR.2020.3031317. [<a href="https://www.skopik.at/ait/2020_trel.pdf">PDF</a>]</p>

opencc-by-nc-sa-4.0Feb 2022View details →
dryad40/100

Data from: eDNA metabarcoding of log hollow sediments and soils highlights the importance of substrate type, frequency of sampling and animal size, for vertebrate species detection

<p>Fauna monitoring often relies on visual monitoring techniques such as camera trappings, which have biases leading to underestimates of vertebrate species diversity. Environmental DNA (eDNA) has emerged as a new source of biodiversity data that may improve biomonitoring; however, eDNA based assessments of species richness remain relatively untested in terrestrial environments. We investigated the suitability of fallen log hollow sediment as a source of vertebrate eDNA, across two sites in south-western Australia - one with a Mediterranean climate and the other semi-arid. We compared two different approaches (camera trapping and eDNA metabarcoding) for monitoring of vertebrate species, and investigated the effect of other factors (frequency of species, timing of visits, frequency of sampling, body size) on vertebrate species detectability. Metabarcoding of hollow sediments resulted in the detection of higher species richness in comparison Hollow sediment detected higher species richness (29 taxa: six birds, three reptiles and 20 mammals) to metabarcoding of soil at the entrance of the hollow (13 taxa: three birds, two reptiles and eight mammals). We detected 31 taxa in total with eDNA metabarcoding and 47 with camera traps, with 14 taxa detected by both (12 mammals and two birds). By comparing camera trap data with eDNA read abundance, we were able to detect vertebrates through eDNA metabarcoding that had visited the area up to two months prior to sample collection. Larger animals were more likely to be detected, and so were vertebrates that were identified multiple times in the camera traps. These findings demonstrate the importance of substrate selection, frequency of sampling, and animal size, on eDNA based monitoring. Future eDNA experimental design should consider all these factors as they affect detection of target taxa. </p>

opencc-zeroMay 2022View details →
zenodo40/100

Logs and QC information for the northern Borneo Orogeny Seismic Survey seismic network

<p>Instrument log files (mass positions + GPS offset/syncs and system information) for the northern Borneo Orogeny Seismic Survey (nBOSS) seismic network, which operated 2018&ndash;2020.<br> <br> FDSN network code: YC (<a href="https://doi.org/10.7914/SN/YC_2018">https://doi.org/10.7914/SN/YC_2018</a>).</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Dataset and Logs from Crime in Medellin Forecasting

<p>Datasets and logs that exced the maximun capacity of githyb repository&nbsp;https://github.com/BioAITeam/Crimes-in-Medellin-Forecasting/</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

IDE Action Log Dataset from a CS1 MOOC

<p>This is a a dataset containing Integrated Development Environment (IDE) logs from an introductory programming MOOC. The dataset contains information on when actions in the IDE were performed in relation to deadlines over the different parts of the course. One exceptional aspect of the dataset is that part of the logs have been gathered at the keystroke level, allowing for fine-grained insight into the learning process. In addition to the IDE logs themselves, the dataset has information on whether students included in the data passed the course. This can facilitate further research that analyzes how time-related behavior relates to performance in introductory programming courses.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record