Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.7.1
Dataset results
6 results for “Malicious Traffic”
MAD (MAlicious Traffic Dataset) in home and commercial environments - Home environment
<p>For the home environment we have: 01 Wifi Modem Router, 03 Smartphones, 01 server, 01 desktop, 01 Multifunction Printer, 01 network extender, 01 SmartTV, 01 Cable TV decoder and 01 firewall. This environment is a local network. The server has the Monitoring Environment and a network card, which provides connectivity and receives all network traffic for analysis.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz - Telegraf collections from the TICK stack</p> <p>netstat.csv.gz - Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the 'timestamp' column events.csv.gz with the 'time' column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2018-09-15 to 2019-02-04</p>
MAD (MAlicious Traffic Dataset) in home and commercial environments - Internal environment
<p>In this environment we have: 01 Wifi Router, 01 Smartphone, 01 server and 01 desktop with virtual machines. This environment, called Internal, is a local network. One of the servers has the Security and Performance Monitoring Environment installed. In addition, 05 virtual machines were instantiated via QEMU on the same network. In this server, a network card provides connectivity to the environment and the other network card receives all network traffic for analysis by the Monitoring Environment. Getting traffic to Suricata is done by Ettercap. The desktop has two virtual machines instantiated via Oracle VirtualBox, on the same network and acts on the network as a client as well.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz - Telegraf collections from the TICK stack</p> <p>netstat.csv.gz - Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the 'timestamp' column events.csv.gz with the 'time' column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2018-06-06 to 2019-01-31</p> <p> </p>
MAD (MAlicious Traffic Dataset) in home and commercial environments - Internet environment
<p>We have for the Internet environment: 01 Switch, 01 IP camera, 01 server for monitoring, 01 server for honeypot and no firewall. This environment is directly connected to the Internet. We installed a server, functioning as a Monitoring Environment. The network traffic was obtained via Port Mirroring on the switch to the Monitoring Environment server.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz - Telegraf collections from the TICK stack</p> <p>netstat.csv.gz - Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the 'timestamp' column events.csv.gz with the 'time' column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2018-08-28 to 2019-11-14</p>
MAD (MAlicious Traffic Dataset) in home and commercial environments - Environment with scalability
<p>We have used the Internet environment: 01 Switch, 01 IP camera, 01 server for monitoring, 01 server for honeypot and no firewall. This environment is directly connected to the Internet. We installed a server, functioning as a Monitoring Environment. The network traffic was obtained via Port Mirroring on the switch to the Monitoring Environment server.</p> <p>We added 08 virtual machines and performed the following test with a denial of service DoS attack:</p> <p>01 virtual machine from 04:00 pm to 23:55 pm on 2019-12-04 with an interval every 01 hour;<br> 02 virtual machines from 23:55 am on 2019-12-04 to 08:50 am on 2019-12-05 with an interval every 01 hour;<br> 04 virtual machines as of 08:55 am on 2019-12-05 to 05:25 pm on 2019-12-06 with an interval every 5 minutes;<br> 08 virtual machines from 05:30 pm on 2019-12-06 to 23:59 on 2019-12-06 with an interval every 5 minutes;<br> End of tests with shutdown of virtual machines at 23:59 on 2019-12-06.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz - Telegraf collections from the TICK stack</p> <p>netstat.csv.gz - Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the 'timestamp' column events.csv.gz with the 'time' column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2019-12-04 to 2019-12-06</p>
CTU-SME-11: a labeled dataset with real benign and malicious network traffic mimicking a small medium-size enterprise environment
<p>As technology advances, the number and complexity of cyber-attacks increase, forcing defense techniques to be updated and improved. To help develop effective tools for detecting security threats it is essential to have reliable and representative security datasets. Many existing security datasets have limitations that make them unsuitable for research, including lack of labels, unbalanced traffic, and outdated threats.</p> <p>CTU-SME-11 is a labeled network dataset designed to address the limitations of previous datasets. The dataset was captured in a real network that mimics a small-medium enterprise setting. Raw network traffic (packets) was captured from 11 devices using tcpdump for a duration of 7 days, from 20th to 26th of February, 2023 in Prague, Czech Republic. The devices were chosen based on the enterprise setting and consists of IoT, desktop and mobile devices, both bare metal and virtualized. The devices were infected with malware or exposed to Internet attacks, and factory reset to restore benign behavior. </p> <p>The raw data was processed to generate network flows (Zeek logs) which were analyzed and labeled. The dataset contains two types of levels, a high level label and a descriptive label, which were put by experts. The former can take three values, benign, malicious or background. The latter contains detailed information about the specific behavior observed in the network flows. The dataset contains 99 million labeled network flows. The overall compressed size of the dataset is 80GB and the uncompressed size is 170GB.</p>
IoT Emulated Dataset for ICMP/Ping Normal and Malicious Traffic
<p>These datasets are related to Intrusion Detection System, Computer Network Traffic and IoT.</p> <p>These datasets are generated for the purpose of differentiating <strong>ICMP/Ping normal and malicious traffic </strong>that are generated from an embedded device (IoT). The differentiation analysis is done using machine learning.</p> <p>There are three types of files that depend on each module of our research framework. The data generation sequence is as follows:</p> <p>The <strong>pcap files </strong>(network traffic) are generated first, the device used to generate the data is an ESP-01s. Afterwards, the pcap files are transformed into <strong>log files </strong>using the Zeek tool, the log files are then extracted and placed into <strong>CSV files</strong>.</p> <p>The CSV files are labeled and ready for the Machine Learning process.</p> <p> </p> <p>The publication reference for this work is here : https://doi.org/10.1109/ACCESS.2023.3327061</p> <p>The code link: <a href="../badge/latestdoi/619245496">https://zenodo.org/badge/latestdoi/619245496</a></p> <p><strong>This version of the release (0.2.0) is for ping flood with spoofed IPs, however, the previous version (</strong>0.1.0)<strong> is for static IP</strong></p> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.