Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

374

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

374 results for “Traffic”

Learn how ShareScore rates datasets ↗
zenodo44/100

Simulated application load on a Kubernetes system based on Human traffic pattern

<p>Contained here is the dataset generated using the tool based on the report 'A Dynamic Kubernetes Load Generation Solution Mimicking Human Traffic Pattern'. The tool was created to fill in the gap of having to generate artificial load within Kubernetes infrastructure. This tool could then be used as a stable testing framework to determine whether the Autoscaling solutions in place can handle the expected load pattern or not. It can also be used as a standard benchmarking tool to test various algorithms that aim to improve the already existing autoscaling solutions. The tool is primarily for IaaS (Infrastructure As A Service) providers, who have to deal with applications as a black box and are unaware of the scaling needs of the application. Some of the data generated in this dataset were made in reference to existing datasets obtained from Google Dataset of different containers. The reference dataset will also be uploaded in the near future.&nbsp;</p> <p>The dataset contains CSV files, which have the name of the pod, average CPU usage, average memory usage, and the timestamp of when the data was collected. These measurements were taken directly from the Prometheus service which uses using Kubernetes metrics server to scrape the data from the pods.&nbsp;<br>The `avg_cpu_usage` is the average CPU percent of the total CPU available the application utilized over a 1-minute interval window. While the `avg_memory_usage` is the average memory utilized in bytes over 1 1-minute interval window. The CSV file names also mention the period of time from when the measurements were taken. Any date and time within the dataset are in the CEST timezone. All of the CSV files in the dataset have the same layout.&nbsp;</p> <p>The datasets were collected from a Kubernetes node in a VM, with the following specifications,<br>- Intel core (Broadwell, no TSX, IBRS), 2 cores @ 2.594 GHzkb<br>- 4GiB of memory<br>- Ubuntu 22.04.4 LTS x86_64<br>- Kernel: 5.15.0-105-generic</p> <p>The names of the file itself specify the type of application load generated. The file names can be divided into segments separated by `__`. The last two segments signify the start time of the measurement data, and the end time of the measurement data respectively. The optional segment before the mentioned two segments signifies the container from the original input dataset that was used as a reference to produce the given data.</p> <p>Similarly, there is an optional tag of `CPU`, `mem`, or `test` at the beginning of the file name. This signifies the type of application load that was used to generate the data. `CPU` would mean the application was created by using options on the tool that used CPU stressing algorithms from `ng-stress`. This was used to better mimic a CPU-intensive application. Similarly, `mem` primarily makes use of a test algorithm that stresses memory. `test` is a special option making use of a combination of tests to have a balanced outcome.&nbsp;</p> <p>In the real world, the applications are never static, and between each run, with the same parameters, the application will always show some variations. To account for these variations, we added some inherent randomness in the number of concurrent connections and the amount of time a virtual user waits to make a request. There are some measurements where these measurements were frozen, which makes the class,<br>- fixed-load_fixed-sleep&nbsp;<br>- random-load_fixed-sleep<br>- fixed-load_random-sleep</p> <p>measurements.<br>The default class of (random-load_random-sleep) is every other measurement not marked by these tags in their name.</p> <p>The dataset itself was collected over a period of 3 days, from 25-03-2024 to 28-03-2024. Each dataset generally contains 34 min of simulation data, with the first and last 2 minutes without any traffic being directed to stabilize the system. The actual 30 min of data corresponds to the daily behavioral pattern as shown in a 30-day period.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Extended Malaysian Traffic Sign Dataset (EMTD)

<p>An extension of the existing Malaysian Traffic Sign Dataset for traffic sign (TS) detection. This contains 66 TS categories, and contains an additional&nbsp;814 new TS instances than the original dataset. Containing 1,413 images in total. In particular, there has been a great increase in classes that previously had fewer than 75 examples</p>

opencc-by-4.0Apr 2018View details →
zenodo44/100

Reference Workloads for Traffic Generation

<p>Collection of traffic generator profiles, implemented for the needs of Superfluidity 5G project.</p> <p>Used in the context of the performance characterization&nbsp;(WP4) and validation&nbsp;(WP7) activities.</p>

opencc-by-4.0Apr 2018View details →
zenodo44/100

CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting

<h2><strong>CESNET-TimeSeries24: The dataset for network traffic forecasting and anomaly detection</strong></h2> <p>The dataset called CESNET-TimeSeries24 was collected by long-term monitoring of selected statistical metrics for 40 weeks for each IP address on the ISP network CESNET3 (Czech Education and Science Network). The dataset encompasses network traffic from more than 275,000 active IP addresses, assigned to a wide variety of devices, including office computers, NATs, servers, WiFi routers, honeypots, and video-game consoles found in dormitories. Moreover, the dataset is also rich in network anomaly types since it contains all types of anomalies, ensuring a comprehensive evaluation of anomaly detection methods.<br><br>Last but not least, the CESNET-TimeSeries24 dataset provides traffic time series on institutional and IP subnet levels to cover all possible anomaly detection or forecasting scopes. Overall, the time series dataset was created from the 66 billion IP flows that contain 4 trillion packets that carry approximately 3.7 petabytes of data. The CESNET-TimeSeries24 dataset is a complex real-world dataset that will finally bring insights into the evaluation of forecasting models in real-world environments.<br><br></p> <p>Please cite the usage of our dataset as:</p> <blockquote> <p>Koumar, J., Hynek, K., Čejka, T. <em>et al.</em> CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting. <em>Sci Data</em> <strong>12</strong>, 338 (2025). https://doi.org/10.1038/s41597-025-04603-x<br><br>@Article{cesnettimeseries24,<br>&nbsp;&nbsp;&nbsp; author={Koumar, Josef and Hynek, Karel and {\v{C}}ejka, Tom{\'a}{\v{s}} and {\v{S}}i{\v{s}}ka, Pavel},<br>&nbsp;&nbsp;&nbsp; title={CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting},<br>&nbsp;&nbsp;&nbsp; journal={Scientific Data},<br>&nbsp;&nbsp;&nbsp; year={2025},<br>&nbsp;&nbsp;&nbsp; month={Feb},<br>&nbsp;&nbsp;&nbsp; day={26},<br>&nbsp;&nbsp;&nbsp; volume={12},<br>&nbsp;&nbsp;&nbsp; number={1},<br>&nbsp;&nbsp;&nbsp; pages={338},<br>&nbsp;&nbsp;&nbsp; issn={2052-4463},<br>&nbsp;&nbsp;&nbsp; doi={10.1038/s41597-025-04603-x},<br>&nbsp;&nbsp;&nbsp; url={https://doi.org/10.1038/s41597-025-04603-x}<br>}<br><br></p> </blockquote> <p>&nbsp;</p> <h3>Time series</h3> <p>We create evenly spaced time series for each IP address by aggregating IP flow records into time series datapoints. The created datapoints represent the behavior of IP addresses within a defined time window of 10 minutes. The vector of time-series metrics v_{ip, i} describes the IP address ip in the i-th time window. Thus, IP flows for vector v_{ip, i} are captured in time windows starting at t_i and ending at t_{i+1}. The&nbsp;time series are built from these datapoints.&nbsp;&nbsp;</p> <p>Datapoints created by the aggregation of IP flows contain the following time-series metrics:</p> <ul> <li><strong><em>Simple volumetric metrics:</em></strong> the number of IP flows, the number of packets, and the transmitted data size (i.e. number of bytes)</li> <li><strong><em>Unique volumetric metrics:</em></strong> the number of unique destination IP addresses, the number of unique destination Autonomous System Numbers (ASNs), and the number of unique destination transport layer ports. The aggregation of \textit{Unique volumetric metrics} is memory intensive since all unique values must be stored in an array. We used a server with 41 GB of RAM, which was enough for 10-minute aggregation on the ISP network. &nbsp;&nbsp;</li> <li><strong><em>Ratios metrics:</em></strong> the ratio of UDP/TCP packets, the ratio of UDP/TCP transmitted data size, the direction ratio of packets, and the direction ratio of transmitted data size</li> <li><em><strong>Average metrics:</strong></em> the average flow duration, and the average Time To Live (TTL)</li> </ul> <p>&nbsp;</p> <p><strong>Multiple time aggregation:&nbsp;</strong> The original datapoints in the dataset are aggregated by 10 minutes of network traffic. The size of the aggregation interval influences anomaly detection procedures, mainly the training speed of the detection model. However, the 10-minute intervals can be too short for longitudinal anomaly detection methods. Therefore, we added two more aggregation intervals to the datasets--1 hour and 1 day.</p> <p><strong>Time series of institutions:</strong>&nbsp; We identify 283 institutions inside the CESNET3 network. These time series aggregated per each institution ID provide a view of the institution's data.&nbsp;</p> <p><strong>Time series of institutional subnets:</strong> We identify 548 institution subnets inside the CESNET3 network. These time series aggregated per each institution ID provide a view of the institution subnet's data.&nbsp;</p> <p>&nbsp;</p> <h3>Data Records</h3> <p>The file hierarchy is described below:</p> <blockquote> <p>cesnet-timeseries24/</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; |- institution_subnets/</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_institution&gt;.csv</p> <p>&nbsp; &nbsp;&nbsp; | &nbsp; &nbsp; |- agg_1_hour/&lt;id_institution&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_day/&lt;id_institution&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- institutions/</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_institution_subnet&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_hour/&lt;id_institution_subnet&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_day/&lt;id_institution_subnet&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- ip_addresses_full/</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_ip_folder&gt;/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_hour/&lt;id_ip_folder&gt;/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_day/&lt;id_ip_folder&gt;/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- ip_addresses_sample/</p> <p>&nbsp; &nbsp; &nbsp;| &nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- agg_1_hour/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- agg_1_day/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- times/</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp;&nbsp; |- times_10_minutes.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- times_1_hour.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- times_1_day.csv</p> <p>&nbsp; &nbsp; &nbsp;|- ids_relationship.csv<br>&nbsp; &nbsp; &nbsp;|- weekends_and_holidays.csv</p> </blockquote> <p>The following list describes time series data fields in CSV files:</p> <ul> <li><strong>id_time: &nbsp;</strong>Unique identifier for each aggregation interval within the time series, used to segment the dataset into specific time periods for analysis.</li> <li><strong>n_flows: </strong>Total number of flows observed in the aggregation interval, indicating the volume of distinct sessions or connections for the IP address.</li> <li><strong>n_packets:&nbsp;</strong>Total number of packets transmitted during the aggregation interval, reflecting the packet-level traffic volume for the IP address.</li> <li><strong>n_bytes: </strong>Total number of bytes transmitted during the aggregation interval, representing the data volume for the IP address.</li> <li><strong>n_dest_ip: </strong>Number of unique destination IP addresses contacted by the IP address during the aggregation interval, showing the diversity of endpoints reached.</li> <li><strong>n_dest_asn: </strong>Number of unique destination Autonomous System Numbers (ASNs) contacted by the IP address during the aggregation interval, indicating the diversity of networks reached.</li> <li><strong>n_dest_port: </strong>Number of unique destination transport layer ports contacted by the IP address during the aggregation interval, representing the variety of services accessed.</li> <li><strong>tcp_udp_ratio_packets: </strong>Ratio of packets sent using TCP versus UDP by the IP address during the aggregation interval, providing insight into the transport protocol usage pattern. This metric belongs to the interval &lt;0, 1&gt; where 1 is when all packets are sent over TCP, and 0 is when all packets are sent over UDP.</li> <li><strong>tcp_udp_ratio_bytes:</strong> Ratio of bytes sent using TCP versus UDP by the IP address during the aggregation interval, highlighting the data volume distribution between protocols. This metric belongs to the interval &lt;0, 1&gt; &nbsp;with same rule as <em>tcp_udp_ratio_packets</em>.</li> <li><strong>dir_ratio_packets: </strong>Ratio of packet directions (inbound versus outbound) for the IP address during the aggregation interval, indicating the balance of traffic flow directions. This metric belongs to the interval &lt;0, 1&gt;, where 1 is when all packets are sent in the outgoing direction from the monitored IP address, and 0 is when all packets are sent in the incoming direction to the monitored IP address.</li> <li><strong>dir_ratio_bytes: </strong>Ratio of byte directions (inbound versus outbound) for the IP address during the aggregation interval, showing the data volume distribution in traffic flows. This metric belongs to the interval &lt;0, 1&gt; with the same rule as <em>dir_ratio_packets</em>.</li> <li><strong>avg_duration: </strong>Average duration of IP flows for the IP address during the aggregation interval, measuring the typical session length.</li> <li><strong>avg_ttl: </strong>Average Time To Live (TTL) of IP flows for the IP address during the aggregation interval, providing insight into the lifespan of packets.</li> </ul> <p>Moreover, the time series created by re-aggregation contains following time series metrics instead of <strong>n_dest_ip</strong>,&nbsp;<strong>n_dest_asn</strong>, and&nbsp;<strong>n_dest_port</strong>:</p> <ul> <li><strong>sum_n_dest_ip:&nbsp;</strong>Sum of numbers of unique destination IP addresses.</li> <li><strong>avg_n_dest_ip:&nbsp;</strong>The average number of unique destination IP addresses.</li> <li><strong>std_n_dest_ip: </strong>Standard deviation of numbers of unique destination IP addresses.</li> <li><strong>sum_n_dest_asn:&nbsp;</strong>Sum of numbers of unique destination ASNs.</li> <li><strong>avg_n_dest_asn:&nbsp;</strong>The average number of unique destination ASNs.</li> <li><strong>std_n_dest_asn: </strong>Standard deviation of numbers of unique destination ASNs)</li> <li><strong>sum_n_dest_port: </strong>Sum of numbers of unique destination transport layer ports.</li> <li><strong>avg_n_dest_port:&nbsp;</strong>&nbsp;The average number of unique destination transport layer ports.</li> <li><strong>std_n_dest_port: </strong>Standard deviation of numbers of unique destination transport layer ports.</li> </ul> <p>&nbsp;</p> <p>Moreover, files &nbsp;<em>identifiers.csv</em> in each dataset type contain IDs of time series that are present in the dataset. Furthermore, the <em>ids_relationship.csv</em> file contains a relationship between IP addresses, Institutions, and institution subnets. The <em>weekends_and_holidays.csv</em> contains information about the non-working days in the Czech Republic.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

MAD (MAlicious Traffic Dataset) in home and commercial environments - Home environment

<p>For the home environment we have: 01&nbsp;Wifi Modem Router, 03 Smartphones, 01&nbsp;server, 01&nbsp;desktop, 01&nbsp;Multifunction Printer, 01&nbsp;network extender, 01&nbsp;SmartTV, 01&nbsp;Cable TV decoder and 01&nbsp;firewall. This environment is a local network. The server has the Monitoring Environment and a network card, which provides connectivity and receives all network traffic for analysis.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>netstat.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the &#39;timestamp&#39; column events.csv.gz with the &#39;time&#39; column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2018-09-15 to 2019-02-04</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

MAD (MAlicious Traffic Dataset) in home and commercial environments - Internal environment

<p>In this environment we have: 01 Wifi Router, 01&nbsp;Smartphone, 01&nbsp;server and 01&nbsp;desktop with virtual machines. This environment, called Internal, is a local network. One of the servers has the Security and Performance Monitoring Environment installed. In addition, 05 virtual machines were instantiated via QEMU on the same network. In this server, a network card provides connectivity to the environment and the other network card receives all network traffic for analysis by the Monitoring Environment. Getting traffic to Suricata is done by Ettercap. The desktop has two virtual machines instantiated via Oracle VirtualBox, on the same network and acts on the network as a client as well.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>netstat.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the &#39;timestamp&#39; column events.csv.gz with the &#39;time&#39; column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from&nbsp;2018-06-06 to 2019-01-31</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

MAD (MAlicious Traffic Dataset) in home and commercial environments - Internet environment

<p>We have for the Internet environment: 01 Switch, 01 IP camera, 01 server for monitoring, 01&nbsp;server for honeypot and no firewall. This environment is directly connected to the Internet. We installed a server, functioning as a Monitoring Environment. The network traffic was obtained via&nbsp;Port Mirroring on the switch to the Monitoring Environment server.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>netstat.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the &#39;timestamp&#39; column events.csv.gz with the &#39;time&#39; column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2018-08-28 to 2019-11-14</p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

MAD (MAlicious Traffic Dataset) in home and commercial environments - Environment with scalability

<p>We have used the Internet environment: 01 Switch, 01 IP camera, 01 server for monitoring, 01&nbsp;server for honeypot and no firewall. This environment is directly connected to the Internet. We installed a server, functioning as a Monitoring Environment. The network traffic was obtained via&nbsp;Port Mirroring on the switch to the Monitoring Environment server.</p> <p>We added 08 virtual machines and performed the following test with a denial of service DoS attack:</p> <p>01 virtual machine from 04:00 pm to&nbsp;23:55 pm on 2019-12-04&nbsp;with an interval every 01 hour;<br> 02 virtual machines from 23:55 am on 2019-12-04&nbsp;to 08:50 am&nbsp;on 2019-12-05&nbsp;with an interval every 01 hour;<br> 04 virtual machines as of 08:55 am on 2019-12-05 to 05:25&nbsp;pm on 2019-12-06 with an interval every 5 minutes;<br> 08 virtual machines from 05:30 pm on 2019-12-06 to 23:59&nbsp;on 2019-12-06 with an interval every 5 minutes;<br> End of tests with shutdown of virtual machines at 23:59&nbsp;on 2019-12-06.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>netstat.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the &#39;timestamp&#39; column events.csv.gz with the &#39;time&#39; column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2019-12-04 to 2019-12-06</p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Dataset of traffic accidents reported on Twitter Bogotá Colombia

<p><strong>1 Classification Dataset</strong></p> <p>This dataset for the classification model contains 3,804 tweets, where 1,902 are related to traffic accident reports (TA, positive class) and 1,902 are unrelated (NTA, negative class).</p> <p>For training the tweet classification model, a collaborative labeling strategy was designed. Here, 30 people labeled data according to the instructions given. Each participant had to evaluate a tweet to manually classify it into one of three categories defined as: traffic accident related, unrelated and don&acute;t know/no response. Each tweet was evaluated by 3 participants. The correct label was selected by voting; the 3 people must agree on the selected label, otherwise the tweet was excluded from training. This process took a month and required the development and deployment of a web application.</p> <p><strong>2 NER Dataset (Named Entity Recognition)</strong></p> <p>For the entity recognition model training, a sample of the filtered tweets resulting from the previous classification phase was taken. 1,340 tweets were extracted, where 800 are from &ldquo;unofficial&rdquo; users, almost 60% of the sample. These tweets were user reports on traffic incident occurred in Bogota from October 2018 to July 2019, including other tweets that contained some location references such as reports on the state of road infrastructure; some tweets from the years 2016 and 2017 were also included. Although these posts were not related to accidents per se, they were selected because they contained location information. The purpose was to train a model that would recognize these entities, because a classifier of accident-related tweets was previously created. Additionally, the dataset was split, reserving 1,072 tweets for training and 268 for evaluation.</p> <p>This dataset was manually labeled using the IOB (Inside-outside- beginning) format. The labeling tool called Brat Annotation Tools was used for this task. The labels defined are Location, which refers to the location of the report; and Time, which refers to the time or date of the incident. Accordingly, 5 labels were generated: B-loc, I-loc, B-time, I-time and O. The O label refers to Others.</p> <p><strong>3 Traffic accident Twitter geolocation</strong></p> <p>A dataset with 26362 traffic accident tweets with the coordinates of the incident and the date of publication.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Traffic network routing index for the Czech Republic

<p><strong>Graph representation of road network of the Czech Republic</strong></p> <p>This dataset contains graph of the entire Czech road network. The data are derived from the Open Street Map project and stored in a HDF5 file. The file contains graph topology and metadata for edges and vertices.&nbsp;</p> <p>Spatial index is also included for the purpose of point snapping and other spatial queries. The spatial index is stored in a SQLite file with SpatiaLite extension.</p> <p>Creation of this dataset has been supported by the <a href="http://antarex-project.eu">Antarex project</a>.</p> <p><strong>Routing Index in HDF5</strong></p> <p><strong>File:&nbsp;</strong><a href="/api/files/f1fb6db8-a2e2-425b-b15e-8e401eb22d44/CZE-1528295206-proc-20180724134457.hdf?versionId=3294ae0b-c3b2-4155-ad36-311d2bb06ed9">CZE-1528295206-proc-20180724134457.hdf</a></p> <p>The index is divided into parts according to geographical boundaries defined by country borders. In this case it contains only single part - CZE. This information along with the creation time is stored as an attribute in the root group of the file.</p> <p>The graph topology is stored in the following way. Nodes have assigned a row index in the edges dataset which points to an outbound edge plus a number of the subsequent edges which are also output to this node. The edge metadata are stored in the EdgeData dataset to avoid redundancy. NodeMap dataset provides a convenient way to query nodes based on their uniqe identifiers.</p> <p><strong>File structure:</strong></p> <pre><code class="language-javascript">HDF5 "CZE-1528295206-proc-20180724134457.hdf" { GROUP "/" { GROUP "Index" { ATTRIBUTE "CreationTime" { DATATYPE H5T_STD_I64LE DATASPACE SCALAR DATA { (0): 1535451896 } } ATTRIBUTE "PartsCount" { DATATYPE H5T_STD_I32LE DATASPACE SCALAR DATA { (0): 1 } } ATTRIBUTE "PartsInfo" { DATATYPE H5T_COMPOUND { H5T_STRING { STRSIZE 4; STRPAD H5T_STR_NULLPAD; CSET H5T_CSET_ASCII; CTYPE H5T_C_S1; } "id"; H5T_STD_I64LE "nodeCount"; H5T_STD_I64LE "edgeCount"; } DATASPACE SIMPLE { ( 1 ) / ( 1 ) } DATA { (0): { "CZE\000", 904085, 2223222 } } } GROUP "CZE" { ATTRIBUTE "PartInfo" { DATATYPE H5T_STRING { STRSIZE 4; STRPAD H5T_STR_NULLPAD; CSET H5T_CSET_ASCII; CTYPE H5T_C_S1; } DATASPACE SCALAR DATA { (0): "CZE\000" } } DATASET "Edges" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "edgeId"; H5T_STD_I32LE "nodeIndex"; H5T_STD_I32LE "computed_speed"; H5T_STD_I32LE "length"; H5T_STD_I32LE "edgeDataIndex"; } DATASPACE SIMPLE { ( 2223222, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } DATASET "Nodes" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "id"; H5T_STD_I32LE "latitudeInt"; H5T_STD_I32LE "longtitudeInt"; H5T_STD_U8LE "edgeOutCount"; H5T_STD_I32LE "edgeOutIndex"; H5T_STD_U8LE "edgeInCount"; } DATASPACE SIMPLE { ( 904085, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } } DATASET "EdgeData" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "id"; H5T_STD_U8LE "speed"; H5T_STD_U8LE "funcClass"; H5T_STD_U8LE "lanes"; H5T_STD_U8LE "vehicleAccess"; H5T_STD_U16LE "specificInfo"; H5T_STD_U16LE "maxWeight"; H5T_STD_U16LE "maxHeight"; H5T_STD_U8LE "maxAxleLoad"; H5T_STD_U8LE "maxWidth"; H5T_STD_U8LE "maxLength"; H5T_STD_I8LE "incline"; } DATASPACE SIMPLE { ( 2838, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } DATASET "NodeMap" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "nodeId"; H5T_STRING { STRSIZE 4; STRPAD H5T_STR_NULLPAD; CSET H5T_CSET_ASCII; CTYPE H5T_C_S1; } "partId"; H5T_STD_I32LE "nodeIndex"; } DATASPACE SIMPLE { ( 904085, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } } } }</code></pre> <p><strong>Spatial index</strong></p> <p><strong>File:&nbsp;</strong><a href="https://zenodo.org/api/files/f1fb6db8-a2e2-425b-b15e-8e401eb22d44/CZE-1528295206-proc-20180829135251.sqlite?versionId=56bb1e60-ff7c-407f-8041-3c628ef78db0">CZE-1528295206-proc-20180829135251.sqlite</a><br> &nbsp;</p> <p>SQLite database contains two primary tables. Table&nbsp;<em>nodes</em>&nbsp;and table&nbsp;<em>segments</em>, that&nbsp;resluts&nbsp;of selection and projection from the tables of the same name in primary db. Both tables are suitable for&nbsp;<em>searching nearest lines</em>&nbsp;task and&nbsp;<em>searching closest node of nearest line</em>&nbsp;task.&nbsp;</p> <pre><code class="language-sql">CREATE TABLE nodes  (      gid INTEGER,      part_gid INTEGER,      node_type INTEGER,      geom POINT  );  CREATE TABLE segments  (      gid INTEGER,      node_gid_from INTEGER,       node_gid_to INTEGER,      frc TEXT,      transition_time DOUBLE,      computed_speed REAL,      geom_length DOUBLE,      geom LINESTRING  ); </code></pre> <p>There are three more tables. Table&nbsp;<em>segments_rt</em>&nbsp;is a copy of table&nbsp;<em>segments</em>, that exclude loops in segments, hence it supports network routing task in&nbsp;SpatiaLite. Next table is&nbsp;<em>rt_network</em>&nbsp;(static graph generated from table&nbsp;<em>segments_rt</em>&nbsp;suitable for routing). The last one is table&nbsp;<em>virtual_rt_network</em>, that is an interface for routing query.&nbsp;&nbsp;</p> <p>&nbsp;</p>

openodc-byDec 2018View details →
zenodo44/100

Extended Wikipedia Web Traffic Daily Dataset (without Missing Values)

<p>This dataset contains 145063 time series representing the number of hits or web traffic for a set of Wikipedia pages from 2015-07-01 to 2022-06-30. This is an extended version of the dataset that was used in the Kaggle Wikipedia Web Traffic forecasting competition. For consistency, the same Wikipedia pages that were used in the competition have been used in this dataset as well.&nbsp;The colons (:) in article names have been replaced by dashes (-) to make the .tsf file readable using our&nbsp;<a href="https://github.com/rakshitha123/TSForecasting/tree/master/utils">data loaders</a>.</p> <p>The original dataset contains missing values. They have been simply replaced by zeros.</p> <p>The data were downloaded from the&nbsp;<a href="https://wikimedia.org/api/rest_v1/#/Pageviews%20data/get_metrics_pageviews_per_article__project___access___agent___article___granularity___start___end_">Wikimedia REST API</a>. According to the conditions of the API, this dataset is licensed under&nbsp;<a href="https://creativecommons.org/licenses/by-sa/3.0/">CC-BY-SA 3.0</a>&nbsp;and&nbsp;<a href="https://www.gnu.org/licenses/fdl-1.3.html">GFDL</a>&nbsp;licenses.</p>

opencc-by-3.0Nov 2022View details →
zenodo44/100

Extended Wikipedia Web Traffic Daily Dataset (with Missing Values)

<p>This dataset contains 145063 time series representing the number of hits or web traffic for a set of Wikipedia pages from 2015-07-01 to 2022-06-30. This is an extended version of the dataset that was used in the Kaggle Wikipedia Web Traffic forecasting competition. For consistency, the same Wikipedia pages that were used in the competition have been used in this dataset as well.&nbsp;The colons (:) in article names have been replaced by dashes (-) to make the .tsf file readable using our <a href="https://github.com/rakshitha123/TSForecasting/tree/master/utils">data loaders</a>.</p> <p><br> The data were downloaded from the <a href="https://wikimedia.org/api/rest_v1/#/Pageviews%20data/get_metrics_pageviews_per_article__project___access___agent___article___granularity___start___end_">Wikimedia REST API</a>. According to the conditions of the API, this dataset is licensed under <a href="https://creativecommons.org/licenses/by-sa/3.0/">CC-BY-SA 3.0</a> and <a href="https://www.gnu.org/licenses/fdl-1.3.html">GFDL</a> licenses.</p>

opencc-by-3.0Nov 2022View details →
zenodo44/100

Dataset of traffic dynamics during the 2019 Kincade Wildfire Evacuation

<p>This dataset has been sourced from the Performance Measurement System of the California Department of Transportation. The data has been processed, analysed, presented and summarized in the paper:&nbsp;<em>Rohaert et al., &lsquo;Traffic dynamics during the 2019 Kincade wildfire evacuation&rsquo;, [Submitted for peer-review to an international journal.], 2022.</em></p> <p><strong>CRediT author statement</strong></p> <p><strong>Arthur Rohaert:&nbsp;</strong>Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Software, Validation, Visualization, Writing - original draft, Writing - review &amp; editing.&nbsp;<strong>Erica D. Kuligowski:&nbsp;</strong>Conceptualization, Funding acquisition, Investigation, Methodology, Resources, Supervision, Validation, Writing - review &amp; editing.&nbsp;<strong>Adam Ardinge:&nbsp;</strong>Conceptualization, Formal analysis, Investigation, Methodology, Resources, Validation, Writing - review &amp; editing.&nbsp;<strong>Jonathan Wahlqvist:</strong>&nbsp;Conceptualization, Formal analysis, Investigation, Methodology, Resources, Supervision, Validation, Writing - review &amp; editing.&nbsp;<strong>Steven M.V. Gwynne:&nbsp;</strong>Conceptualization, Funding acquisition, Investigation, Methodology, Resources, Supervision, Validation, Writing - review &amp; editing.&nbsp;<strong>Amanda Kimball:&nbsp;</strong>Conceptualization, Funding acquisition, Investigation, Methodology, Project Administration, Resources, Supervision, Validation, Writing - review &amp; editing.&nbsp;<strong>Noureddine B&eacute;nichou:&nbsp;</strong>Conceptualization, Funding acquisition, Investigation, Methodology, Resources, Supervision, Validation, Writing - review &amp; editing.&nbsp;<strong>Enrico Ronchi:&nbsp;</strong>Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Resources, Supervision, Validation, Writing - original draft, Writing - review &amp; editing</p> <p><strong>Acknowledgements</strong></p> <p>This work has been funded under award 60NANB20D191 from the National Institute of Standards and Technology (NIST), U.S. Department of Commerce. The authors would like to thank the WUI-NITY team (Guillermo Rein, Nikolaos Kalogeropoulos, Harry Mitchell, Max Kinateder, Maxime Berthiaume). The authors also acknowledge the technical panel of the project for their support and guidance: Carole Adam, Amy Christianson, Tom Cova, Lauren Folk, Abishek Gaur, Paolo Intini, Justice Jones, Bryan Klein, Chris Lautenberger, Ruggiero Lovreglio, Jerry McAdams, Ruddy Mell, Elise Miller-Hooks, Cathy Stephens, Steve Taylor, Sandra Vaiciulyte, Xilei Zhao, Rita Fahy, Lucian Deaton, and Michele Steinberg.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Dataset: "Traffic Noise at Moderate Levels Affects Cognitive Performance: Do Distance-Induced Temporal Changes Matter?"

<p>This repository contains the dataset presented in&nbsp;&quot;Traffic Noise at Moderate Levels Affects Cognitive Performance: Do Distance-Induced Temporal Changes Matter?&quot; (https://doi.org/10.3390/ijerph20053798) as well as&nbsp;the SPSS syntax used for the statistical evaluation. Additionally, calibrated binaural recordings of the evaluated stimuli are provided as 32 bit .wav files, the values stored in those files correspond to pascals.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Dataset from VR Streaming Server (Emulated) and Radio Access Network for Streaming Traffic

<p>The dataset contains an experiment in&nbsp;&nbsp;a site where UEs attach to a gNodeB that provides access to a streaming server that is stressed with high demanding transcoding workloads to emulate VR/AR processes. The UEs are realized through the Remote UE mode enabled by Amarisoft Simbox emulator, and the gNodeB is realized through the Amarisoft Callbox, which also provides the user plane function. The emulated VR streaming server is deployed as a Nginx pod in a Kubernetes cluster.</p> <p>We rely on MonB5G sampling functions that feed monitoring data (CPU and RAN parameters) to the monitoring system.&nbsp;A streaming video server has been deployed with the help of a NGINX server. It provides video-on-demand and video streaming, which can be accessed by any user (or UE) for real-time reproduction. This VR video streaming emulation aids to assess the performance of the network and therefore the benefits that each solution has brought. The video &ldquo;Big Buck Bunny&rdquo; with h.264 encoding and a resolution of 1920x1080p has been used for the experiments.The description of dataset features&nbsp;are:<br> 1-) Index Number,<br> 2-) Time: Time of the experiment,<br> 3-) N:&nbsp;number of VR streaming clients,<br> 4-) C: Average CPU of VR streaming server [mc]<br> 5-) O: Outbound traffic at the server average outbound traffic (O) flowing from the data interface of the video server.&nbsp;<br> 6-) R: Instantaneous downlink bit rate [Mbps],</p> <p>The original video file information:</p> <table> <tbody> <tr> <td> <p>Video codec&nbsp;</p> </td> <td> <p>Advanced Video Codec (AVC)&nbsp;</p> </td> </tr> <tr> <td> <p>Width&nbsp;</p> </td> <td> <p>1920 pixels&nbsp;</p> </td> </tr> <tr> <td> <p>Height&nbsp;</p> </td> <td> <p>1080 pixels&nbsp;</p> </td> </tr> <tr> <td> <p>Display aspect radio&nbsp;</p> </td> <td> <p>16:9&nbsp;</p> </td> </tr> <tr> <td> <p>Duration&nbsp;</p> </td> <td> <p>10 min 34 s&nbsp;</p> </td> </tr> <tr> <td> <p>Max Bitrate&nbsp;</p> </td> <td> <p>16.7 Mb/s&nbsp;</p> </td> </tr> <tr> <td> <p>Frame rate&nbsp;</p> </td> <td> <p>30 FPS&nbsp;</p> </td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Data set and classification method for low quality web traffic identification in video marketing campaigns

<p>Final outcomes of the InPreVi (AI4Media) project developed in 2022.&nbsp;</p> <p>1. Data set describing the statistics of the video ad marketing campaigns</p> <p>2. Script for web traffic classification</p>

opencc-by-4.0May 2023View details →
zenodo44/100

CTU-SME-11: a labeled dataset with real benign and malicious network traffic mimicking a small medium-size enterprise environment

<p>As technology advances, the number and complexity of cyber-attacks increase, forcing defense techniques to be updated and improved. To help develop effective tools for detecting security threats it is essential to have reliable and representative security datasets. Many existing security datasets have limitations that make them unsuitable for research, including lack of labels, unbalanced traffic, and outdated threats.</p> <p>CTU-SME-11 is a labeled network dataset designed to address the limitations of previous datasets. The dataset was captured in a real network that mimics a small-medium enterprise setting. Raw network traffic (packets) was captured from 11 devices using tcpdump for a duration of 7 days, from 20th to 26th of February, 2023 in Prague, Czech Republic. The devices were chosen based on the enterprise setting and consists of IoT, desktop and mobile devices, both bare metal and virtualized. The devices were infected with malware or exposed to Internet attacks, and factory reset to restore benign behavior.&nbsp;</p> <p>The raw data was processed to generate network flows (Zeek logs) which were analyzed and labeled. The dataset contains two types of levels, a high level label and a descriptive label, which were put by experts. The former can take three values, benign, malicious or background. The latter contains detailed information about the specific behavior observed in the network flows. The dataset contains 99 million labeled network flows. The overall compressed size of the dataset is 80GB and the uncompressed size is 170GB.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Urban Traffic Simulation Data from Real Fusion Estimates

<p>The datasets contain vehicle data from the simulation of the city center of Guimar&atilde;es. The simulation was created using data collected from real sensor data, thus, providing an accurate view of the traffic flows. The data contains route information, fuel consumption, emissions, driving distance, and the amount of time each vehicle is stopped.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Network traffic datasets created by Single Flow Time Series Analysis

<p><strong>Network traffic datasets created by Single Flow Time Series Analysis</strong></p> <p>Datasets were created for the paper: Network Traffic Classification based on Single Flow Time Series Analysis -- Josef Koumar, Karel Hynek, Tom&aacute;&scaron; Čejka -- which was published at The 19th International Conference on Network and Service Management (CNSM) 2023. Please cite usage of our datasets as:<br>&nbsp;</p> <blockquote> <p>J. Koumar, K. Hynek and T. Čejka, "Network Traffic Classification Based on Single Flow Time Series Analysis," <em>2023 19th International Conference on Network and Service Management (CNSM)</em>, Niagara Falls, ON, Canada, 2023, pp. 1-7, doi: 10.23919/CNSM59352.2023.10327876.</p> </blockquote> <p>This Zenodo repository contains 23 datasets created from 15 well-known published datasets which are cited in the table below. Each dataset contains 69 features created by Time Series Analysis of Single Flow Time Series. The detailed description of features from datasets is in the file: <em>feature_description.pdf</em></p> <p>&nbsp;</p> <p>In the following table is a description of each dataset file:</p> <table> <tbody> <tr> <td><strong>File name</strong></td> <td><strong>Detection problem</strong></td> <td><strong>Citation of original raw dataset</strong></td> </tr> <tr> <td>botnet_binary.csv&nbsp;</td> <td>Binary detection of botnet&nbsp;</td> <td>S. Garc&iacute;a et al. An Empirical Comparison of Botnet Detection Methods. Computers &amp; Security, 45:100&ndash;123, 2014.&nbsp;</td> </tr> <tr> <td>botnet_multiclass.csv&nbsp;</td> <td>Multi-class classification of botnet&nbsp;</td> <td>S. Garc&iacute;a et al. An Empirical Comparison of Botnet Detection Methods. Computers &amp; Security, 45:100&ndash;123, 2014.&nbsp;</td> </tr> <tr> <td>cryptomining_design.csv</td> <td>Binary detection of cryptomining; the design part&nbsp;</td> <td>Richard Pln&yacute; et al. Datasets of Cryptomining Communication. Zenodo, October 2022&nbsp;</td> </tr> <tr> <td>cryptomining_evaluation.csv&nbsp;</td> <td>Binary detection of cryptomining; the evaluation part&nbsp;</td> <td>Richard Pln&yacute; et al. Datasets of Cryptomining Communication. Zenodo, October 2022&nbsp;</td> </tr> <tr> <td>dns_malware.csv&nbsp;</td> <td>Binary detection of malware DNS&nbsp;</td> <td>Samaneh Mahdavifar et al. Classifying Malicious Domains using DNS Traffic Analysis. In DASC/PiCom/CBDCom/CyberSciTech 2021, pages 60&ndash;67. IEEE, 2021.&nbsp;</td> </tr> <tr> <td>doh_cic.csv&nbsp;</td> <td>Binary detection of DoH&nbsp;</td> <td> <p>Mohammadreza MontazeriShatoori et al. Detection of doh tunnels using time-series classification of encrypted traffic. In DASC/PiCom/CBDCom/CyberSciTech 2020, pages 63&ndash;70. IEEE, 2020&nbsp;</p> </td> </tr> <tr> <td>doh_real_world.csv&nbsp;</td> <td>Binary detection of DoH&nbsp;</td> <td>Kamil Jeř&aacute;bek et al. Collection of datasets with DNS over HTTPS traffic. Data in Brief, 42:108310, 2022&nbsp;</td> </tr> <tr> <td>dos.csv&nbsp;</td> <td>Binary detection of DoS&nbsp;</td> <td>Nickolaos Koroniotis et al. Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset. Future Gener. Comput. Syst., 100:779&ndash;796, 2019.</td> </tr> <tr> <td>edge_iiot_binary.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.</td> </tr> <tr> <td>edge_iiot_multiclass.csv</td> <td>Multi-class classification of IoT malware</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.</td> </tr> <tr> <td>https_brute_force.csv</td> <td>Binary detection of HTTPS Brute Force</td> <td>Jan Luxemburk et al. HTTPS Brute-force dataset with extended network flows, November 2020</td> </tr> <tr> <td>ids_cic_binary.csv</td> <td>Binary detection of intrusion in IDS</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108&ndash;116, 2018.</td> </tr> <tr> <td>ids_cic_multiclass.csv&nbsp;</td> <td>Multi-class classification of intrusion in IDS&nbsp;</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108&ndash;116, 2018.&nbsp;</td> </tr> <tr> <td>ids_unsw_nb_15_binary.csv&nbsp;</td> <td>Binary detection of intrusion in IDS&nbsp;</td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1&ndash;6. IEEE, 2015.</td> </tr> <tr> <td>ids_unsw_nb_15_multiclass.csv&nbsp;</td> <td>Multi-class classification of intrusion in IDS&nbsp;</td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1&ndash;6. IEEE, 2015.</td> </tr> <tr> <td>iot_23.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Sebastian Garcia et al. IoT-23: A labeled dataset with malicious and benign IoT network traffic, January 2020. More details here https://www.stratosphereips.org /datasets-iot23</td> </tr> <tr> <td>ton_iot_binary.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021</td> </tr> <tr> <td>ton_iot_multiclass.csv&nbsp;</td> <td>Multi-class classification of IoT malware&nbsp;</td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021</td> </tr> <tr> <td>tor_binary.csv&nbsp;</td> <td>Binary detection of TOR&nbsp;</td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253&ndash;262. SciTePress, 2017.&nbsp;</td> </tr> <tr> <td>tor_multiclass.csv&nbsp;</td> <td>Multi-class classification of TOR&nbsp;</td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253&ndash;262. SciTePress, 2017.&nbsp;</td> </tr> <tr> <td>vpn_iscx_binary.csv&nbsp;</td> <td>Binary detection of VPN&nbsp;</td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407&ndash;414, 2016.&nbsp;</td> </tr> <tr> <td>vpn_iscx_multiclass.csv&nbsp;</td> <td>Multi-class classification of VPN&nbsp;</td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407&ndash;414, 2016.&nbsp;</td> </tr> <tr> <td>vpn_vnat_binary.csv&nbsp;</td> <td>Binary detection of VPN&nbsp;</td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022</td> </tr> <tr> <td>vpn_vnat_multiclass.csv</td> <td>Multi-class classification of VPN&nbsp;</td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Dynamic traffic noise levels (LAeq) in the city of Tartu for two times of the day: 04:00 and 16:00

<p>Animations of the dynamic traffic noise levels (LAeq) in the city of Tartu, Estonia, for two times of the day: (a) 04:00 (sparse traffic) and (b) 16:00 (dense traffic). Blue circles represent the location of&nbsp;vehicles.&nbsp;Colours on receiver points represent noise levels according to legend (c).</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record