Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

37 results for “Network Traffic”

Learn how ShareScore rates datasets ↗
zenodo48/100

ADS-C Air Traffic Data Collected by the OpenSky Network

<p>ADS-C data collected by the OpenSky Network since 7th July 2023.&nbsp;</p> <p>Data underlying (Version 1.1)</p> <h1>A First Look at Exploiting the Automatic Dependent Surveillance-Contract Protocol for Open Aviation Research</h1> <p>https://journals.open.tudelft.nl/joas/article/view/7229</p>

opencc-by-4.0Oct 2023View details →
zenodo48/100

Short-term traffic flow prediction based on secondary hybrid decomposition and deep echo state networks

<p>The publication titled "Short-term traffic flow prediction based on secondary hybrid decomposition and deep echo state networks" is supported by the STRIDE K3 project. The dataset used in the publication is uploaded here.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Host Network Traffic 2019

<p><strong><em>Dataset Summary</em></strong></p> <ul> <li><strong>Timespan</strong>: 2019-01-01 : 2019-12-31</li> <li><strong>Granularity:&nbsp;</strong>1-hour disjoint time windows</li> <li><strong># of&nbsp;characteristics observed:&nbsp;</strong>9</li> <li><strong>Hosts observed: </strong>65536</li> <li><strong>Labels:&nbsp;</strong>included</li> <li><strong>Unzipped volume:&nbsp;</strong>approx. 10 GB</li> </ul> <p><strong><em>Dataset Origins</em></strong></p> <p>Dataset&nbsp;was collected over the <strong>whole year</strong>&nbsp;<strong>&nbsp;2019</strong>. The observation points for the collection of IP flows were located at the borders of the university campus network. The campus university network has /16 CIDR IPv4 network range at disposal and contains various network segments from segments connecting dormitories, over server segments, to a segment containing working stations of university administrative workers.&nbsp;<strong>A host in our dataset is identified by its source IPv4 address. &nbsp;</strong></p> <p><em><strong>Variables</strong></em></p> <p>The dataset contains the following variables:</p> <ul> <li><strong>Aggregations</strong>&nbsp;- created sums of the individual variables over a one-hour interval: <ul> </ul> <ul> <li><strong># of flows &nbsp;</strong>- number of flows for a given source IP&nbsp;</li> <li><strong># of packets </strong>&nbsp;-&nbsp;number of packets for a given source IP</li> <li><strong># of bytes </strong>&nbsp;-&nbsp;number of packets for a given source IP</li> <li><strong>flow duration </strong>&nbsp;- average flow duration in seconds</li> </ul> </li> <li><strong>Distinct Counts&nbsp;</strong>- count of distinct values for each variable over a one-hour window <ul> <li><strong># of peers </strong>&nbsp;- number of distinct communication peers for a given source IP</li> <li><strong># of ports </strong>&nbsp;- number of distinct destination ports&nbsp;for a given source IP</li> <li><strong># of protocols</strong>&nbsp;- number of distinct communication protocols&nbsp;for a given source IP</li> <li><strong># of AS numbers</strong>&nbsp;- number of distinct destination AS numbers for a given source IP</li> <li><strong># of countries </strong>&nbsp;- number of distinct destination countries&nbsp;for a given source&nbsp;</li> </ul> </li> </ul> <p><em><strong>Dataset Structure</strong></em></p> <ul> <li><strong>Dataset Files</strong> - each variable is contained in one <strong>Comma-Separated File (.csv)&nbsp;</strong>file <ul> <li><strong>Row index&nbsp;-&nbsp;</strong>&nbsp;timestamp of the observation window (8760 rows)</li> <li><strong>Columns index -&nbsp;</strong>&nbsp;anonymized IP addresses (65536&nbsp;columns)</li> </ul> </li> <li><strong>Label File -&nbsp;</strong>contains labels of the individual IP addresses from the Dataset Files <ul> <li><strong>Row index </strong>- anonymized IP addresses (65536 rows)</li> <li><strong>Columns index </strong>- labels for the IP addresses <ul> <li><strong>Subnet </strong>- ID&nbsp;of a subnet - hosts belonging to the same subnet have the same Id.</li> <li><strong>Subnet_range&nbsp;</strong>- CIDR range of a&nbsp;subnet</li> <li><strong>Unit -&nbsp;</strong>an ID of&nbsp;&nbsp;administrative unit owning the network range</li> <li><strong>Sub-unit </strong>&nbsp;- an ID of&nbsp;&nbsp;administrative sub-unit owning the network range</li> <li><strong>Subnet_label -&nbsp;&nbsp;</strong>subnet label <ul> <li><strong>Servers - </strong>selected subnets containing mostly servers (133.250.178.0/24, 133.250.163.0/24)</li> <li><strong>Workstations - </strong>selected subnets containing mostly workstations&nbsp;(133.250.146.0/24,&nbsp;133.250.157.128/25)</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p><strong><em>Further notes</em></strong></p> <ul> <li><strong>N/A values </strong> <ul> <li><strong>Variables&nbsp;</strong>- means that in a given observation window, the host did not communicate</li> <li><strong>Labels -&nbsp;</strong>no additional information on this IP is available</li> </ul> </li> <li><strong>Dataset load&nbsp;</strong> <ul> <li> <pre><code class="language-python">df = pd.read_csv(&lt;filename&gt;,header=[0], index_col=[0])</code></pre> </li> </ul> </li> </ul>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Crowdsourced air traffic data from The OpenSky Network 2020 [CC-BY]

<p><strong>WARNING! </strong>This dataset is no longer updated after January 2022. Refer to the <a href="https://doi.org/10.5281/zenodo.3737101">original dataset</a> with different license terms for an up to date version.</p> <p><strong>Motivation</strong></p> <p>The data in this dataset is derived and cleaned from the full OpenSky dataset to illustrate the development of air traffic during the COVID-19 pandemic. It spans all flights seen by the network&#39;s more than 2500 members since 1 January 2019. More data will be periodically included in the dataset until the end of the COVID-19 pandemic.</p> <p><strong>License</strong></p> <p>Creative Commons CC-BY</p> <p>The only difference with the <a href="https://doi.org/10.5281/zenodo.3737101">original dataset</a> comes from anonymised aircraft information.</p> <p><strong>WARNING:</strong>This dataset is now longer updated after January 2022. The original dataset is still updated.</p> <p><strong>Disclaimer</strong></p> <p>The data provided in the files is provided as is. Despite our best efforts at filtering out potential issues, some information could be erroneous.</p> <ul> <li>Origin and destination airports are computed online based on the ADS-B trajectories on approach/takeoff: no crosschecking with external sources of data has been conducted.<br> Fields <strong>origin</strong> or <strong>destination</strong> are empty when no airport could be found.</li> <li>Aircraft information come from the OpenSky aircraft database. Fields <strong>typecode</strong> and <strong>registration</strong> are empty when the aircraft is not present in the database.</li> </ul> <p><strong>Description of the dataset</strong></p> <p>One file per month is provided as a csv file with the following features:</p> <ul> <li><strong>callsign</strong>: the identifier of the flight displayed on ATC screens (usually the first three letters are reserved for an airline: AFR for Air France, DLH for Lufthansa, etc.)</li> <li><strong>number</strong>: the commercial number of the flight, when available (the matching with the callsign comes from public open API)</li> <li><strong>aircraft_uid</strong>: a unique anonymised identifier for aircraft;</li> <li><strong>typecode</strong>: the aircraft model type (when available);</li> <li><strong>origin</strong>: a four letter code for the origin airport of the flight (when available);</li> <li><strong>destination</strong>: a four letter code for the destination airport of the flight (when available);</li> <li><strong>firstseen</strong>: the UTC timestamp of the first message received by the OpenSky Network;</li> <li><strong>lastseen</strong>: the UTC timestamp of the last message received by the OpenSky Network;</li> <li><strong>day</strong>: the UTC day of the last message received by the OpenSky Network;</li> <li><strong>latitude_1</strong>, <strong>longitude_1</strong>, <strong>altitude_1</strong>: the first detected position of the aircraft;</li> <li><strong>latitude_2</strong>, <strong>longitude_2</strong>, <strong>altitude_2</strong>: the last detected position of the aircraft.</li> </ul> <p><strong>Examples</strong></p> <p>Possible visualisations and a more detailed description of the data are available at the following page:<br> &lt;<a href="https://traffic-viz.github.io/scenarios/covid19.html">https://traffic-viz.github.io/scenarios/covid19.html</a>&gt;</p> <p><strong>Credit</strong></p> <p>Martin Strohmeier, Xavier Olive, Jannis L&uuml;bbe, Matthias Sch&auml;fer, and Vincent Lenders<br> <strong>&quot;</strong>Crowdsourced air traffic data from the OpenSky Network 2019&ndash;2020<strong>&quot;</strong><br> <em>Earth System Science Data</em> 13(2), 2021<br> <a href="https://doi.org/10.5194/essd-13-357-2021">https://doi.org/10.5194/essd-13-357-2021</a></p> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

HIKARI-2021: Generating Network Intrusion Detection Dataset Based on Real and Encrypted Synthetic Attack Traffic

<p>Available datasets from the paper&nbsp;Generating Encrypted Network Traffic for Intrusion Detection Datasets.</p> <p>To produce the dataset follow the technical detail in <a href="https://github.com/andreysfc/generating-encrypted-network">github</a></p>

opencc-by-4.0May 2021View details →
zenodo44/100

CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting

<h2><strong>CESNET-TimeSeries24: The dataset for network traffic forecasting and anomaly detection</strong></h2> <p>The dataset called CESNET-TimeSeries24 was collected by long-term monitoring of selected statistical metrics for 40 weeks for each IP address on the ISP network CESNET3 (Czech Education and Science Network). The dataset encompasses network traffic from more than 275,000 active IP addresses, assigned to a wide variety of devices, including office computers, NATs, servers, WiFi routers, honeypots, and video-game consoles found in dormitories. Moreover, the dataset is also rich in network anomaly types since it contains all types of anomalies, ensuring a comprehensive evaluation of anomaly detection methods.<br><br>Last but not least, the CESNET-TimeSeries24 dataset provides traffic time series on institutional and IP subnet levels to cover all possible anomaly detection or forecasting scopes. Overall, the time series dataset was created from the 66 billion IP flows that contain 4 trillion packets that carry approximately 3.7 petabytes of data. The CESNET-TimeSeries24 dataset is a complex real-world dataset that will finally bring insights into the evaluation of forecasting models in real-world environments.<br><br></p> <p>Please cite the usage of our dataset as:</p> <blockquote> <p>Koumar, J., Hynek, K., Čejka, T. <em>et al.</em> CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting. <em>Sci Data</em> <strong>12</strong>, 338 (2025). https://doi.org/10.1038/s41597-025-04603-x<br><br>@Article{cesnettimeseries24,<br>&nbsp;&nbsp;&nbsp; author={Koumar, Josef and Hynek, Karel and {\v{C}}ejka, Tom{\'a}{\v{s}} and {\v{S}}i{\v{s}}ka, Pavel},<br>&nbsp;&nbsp;&nbsp; title={CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting},<br>&nbsp;&nbsp;&nbsp; journal={Scientific Data},<br>&nbsp;&nbsp;&nbsp; year={2025},<br>&nbsp;&nbsp;&nbsp; month={Feb},<br>&nbsp;&nbsp;&nbsp; day={26},<br>&nbsp;&nbsp;&nbsp; volume={12},<br>&nbsp;&nbsp;&nbsp; number={1},<br>&nbsp;&nbsp;&nbsp; pages={338},<br>&nbsp;&nbsp;&nbsp; issn={2052-4463},<br>&nbsp;&nbsp;&nbsp; doi={10.1038/s41597-025-04603-x},<br>&nbsp;&nbsp;&nbsp; url={https://doi.org/10.1038/s41597-025-04603-x}<br>}<br><br></p> </blockquote> <p>&nbsp;</p> <h3>Time series</h3> <p>We create evenly spaced time series for each IP address by aggregating IP flow records into time series datapoints. The created datapoints represent the behavior of IP addresses within a defined time window of 10 minutes. The vector of time-series metrics v_{ip, i} describes the IP address ip in the i-th time window. Thus, IP flows for vector v_{ip, i} are captured in time windows starting at t_i and ending at t_{i+1}. The&nbsp;time series are built from these datapoints.&nbsp;&nbsp;</p> <p>Datapoints created by the aggregation of IP flows contain the following time-series metrics:</p> <ul> <li><strong><em>Simple volumetric metrics:</em></strong> the number of IP flows, the number of packets, and the transmitted data size (i.e. number of bytes)</li> <li><strong><em>Unique volumetric metrics:</em></strong> the number of unique destination IP addresses, the number of unique destination Autonomous System Numbers (ASNs), and the number of unique destination transport layer ports. The aggregation of \textit{Unique volumetric metrics} is memory intensive since all unique values must be stored in an array. We used a server with 41 GB of RAM, which was enough for 10-minute aggregation on the ISP network. &nbsp;&nbsp;</li> <li><strong><em>Ratios metrics:</em></strong> the ratio of UDP/TCP packets, the ratio of UDP/TCP transmitted data size, the direction ratio of packets, and the direction ratio of transmitted data size</li> <li><em><strong>Average metrics:</strong></em> the average flow duration, and the average Time To Live (TTL)</li> </ul> <p>&nbsp;</p> <p><strong>Multiple time aggregation:&nbsp;</strong> The original datapoints in the dataset are aggregated by 10 minutes of network traffic. The size of the aggregation interval influences anomaly detection procedures, mainly the training speed of the detection model. However, the 10-minute intervals can be too short for longitudinal anomaly detection methods. Therefore, we added two more aggregation intervals to the datasets--1 hour and 1 day.</p> <p><strong>Time series of institutions:</strong>&nbsp; We identify 283 institutions inside the CESNET3 network. These time series aggregated per each institution ID provide a view of the institution's data.&nbsp;</p> <p><strong>Time series of institutional subnets:</strong> We identify 548 institution subnets inside the CESNET3 network. These time series aggregated per each institution ID provide a view of the institution subnet's data.&nbsp;</p> <p>&nbsp;</p> <h3>Data Records</h3> <p>The file hierarchy is described below:</p> <blockquote> <p>cesnet-timeseries24/</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; |- institution_subnets/</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_institution&gt;.csv</p> <p>&nbsp; &nbsp;&nbsp; | &nbsp; &nbsp; |- agg_1_hour/&lt;id_institution&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_day/&lt;id_institution&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- institutions/</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_institution_subnet&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_hour/&lt;id_institution_subnet&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_day/&lt;id_institution_subnet&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- ip_addresses_full/</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_ip_folder&gt;/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_hour/&lt;id_ip_folder&gt;/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- agg_1_day/&lt;id_ip_folder&gt;/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- ip_addresses_sample/</p> <p>&nbsp; &nbsp; &nbsp;| &nbsp; &nbsp;&nbsp; |- agg_10_minutes/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- agg_1_hour/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- agg_1_day/&lt;id_ip&gt;.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- identifiers.csv</p> <p>&nbsp; &nbsp; &nbsp;|- times/</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp;&nbsp;&nbsp; |- times_10_minutes.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- times_1_hour.csv</p> <p>&nbsp; &nbsp; &nbsp;|&nbsp; &nbsp; &nbsp; |- times_1_day.csv</p> <p>&nbsp; &nbsp; &nbsp;|- ids_relationship.csv<br>&nbsp; &nbsp; &nbsp;|- weekends_and_holidays.csv</p> </blockquote> <p>The following list describes time series data fields in CSV files:</p> <ul> <li><strong>id_time: &nbsp;</strong>Unique identifier for each aggregation interval within the time series, used to segment the dataset into specific time periods for analysis.</li> <li><strong>n_flows: </strong>Total number of flows observed in the aggregation interval, indicating the volume of distinct sessions or connections for the IP address.</li> <li><strong>n_packets:&nbsp;</strong>Total number of packets transmitted during the aggregation interval, reflecting the packet-level traffic volume for the IP address.</li> <li><strong>n_bytes: </strong>Total number of bytes transmitted during the aggregation interval, representing the data volume for the IP address.</li> <li><strong>n_dest_ip: </strong>Number of unique destination IP addresses contacted by the IP address during the aggregation interval, showing the diversity of endpoints reached.</li> <li><strong>n_dest_asn: </strong>Number of unique destination Autonomous System Numbers (ASNs) contacted by the IP address during the aggregation interval, indicating the diversity of networks reached.</li> <li><strong>n_dest_port: </strong>Number of unique destination transport layer ports contacted by the IP address during the aggregation interval, representing the variety of services accessed.</li> <li><strong>tcp_udp_ratio_packets: </strong>Ratio of packets sent using TCP versus UDP by the IP address during the aggregation interval, providing insight into the transport protocol usage pattern. This metric belongs to the interval &lt;0, 1&gt; where 1 is when all packets are sent over TCP, and 0 is when all packets are sent over UDP.</li> <li><strong>tcp_udp_ratio_bytes:</strong> Ratio of bytes sent using TCP versus UDP by the IP address during the aggregation interval, highlighting the data volume distribution between protocols. This metric belongs to the interval &lt;0, 1&gt; &nbsp;with same rule as <em>tcp_udp_ratio_packets</em>.</li> <li><strong>dir_ratio_packets: </strong>Ratio of packet directions (inbound versus outbound) for the IP address during the aggregation interval, indicating the balance of traffic flow directions. This metric belongs to the interval &lt;0, 1&gt;, where 1 is when all packets are sent in the outgoing direction from the monitored IP address, and 0 is when all packets are sent in the incoming direction to the monitored IP address.</li> <li><strong>dir_ratio_bytes: </strong>Ratio of byte directions (inbound versus outbound) for the IP address during the aggregation interval, showing the data volume distribution in traffic flows. This metric belongs to the interval &lt;0, 1&gt; with the same rule as <em>dir_ratio_packets</em>.</li> <li><strong>avg_duration: </strong>Average duration of IP flows for the IP address during the aggregation interval, measuring the typical session length.</li> <li><strong>avg_ttl: </strong>Average Time To Live (TTL) of IP flows for the IP address during the aggregation interval, providing insight into the lifespan of packets.</li> </ul> <p>Moreover, the time series created by re-aggregation contains following time series metrics instead of <strong>n_dest_ip</strong>,&nbsp;<strong>n_dest_asn</strong>, and&nbsp;<strong>n_dest_port</strong>:</p> <ul> <li><strong>sum_n_dest_ip:&nbsp;</strong>Sum of numbers of unique destination IP addresses.</li> <li><strong>avg_n_dest_ip:&nbsp;</strong>The average number of unique destination IP addresses.</li> <li><strong>std_n_dest_ip: </strong>Standard deviation of numbers of unique destination IP addresses.</li> <li><strong>sum_n_dest_asn:&nbsp;</strong>Sum of numbers of unique destination ASNs.</li> <li><strong>avg_n_dest_asn:&nbsp;</strong>The average number of unique destination ASNs.</li> <li><strong>std_n_dest_asn: </strong>Standard deviation of numbers of unique destination ASNs)</li> <li><strong>sum_n_dest_port: </strong>Sum of numbers of unique destination transport layer ports.</li> <li><strong>avg_n_dest_port:&nbsp;</strong>&nbsp;The average number of unique destination transport layer ports.</li> <li><strong>std_n_dest_port: </strong>Standard deviation of numbers of unique destination transport layer ports.</li> </ul> <p>&nbsp;</p> <p>Moreover, files &nbsp;<em>identifiers.csv</em> in each dataset type contain IDs of time series that are present in the dataset. Furthermore, the <em>ids_relationship.csv</em> file contains a relationship between IP addresses, Institutions, and institution subnets. The <em>weekends_and_holidays.csv</em> contains information about the non-working days in the Czech Republic.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Traffic network routing index for the Czech Republic

<p><strong>Graph representation of road network of the Czech Republic</strong></p> <p>This dataset contains graph of the entire Czech road network. The data are derived from the Open Street Map project and stored in a HDF5 file. The file contains graph topology and metadata for edges and vertices.&nbsp;</p> <p>Spatial index is also included for the purpose of point snapping and other spatial queries. The spatial index is stored in a SQLite file with SpatiaLite extension.</p> <p>Creation of this dataset has been supported by the <a href="http://antarex-project.eu">Antarex project</a>.</p> <p><strong>Routing Index in HDF5</strong></p> <p><strong>File:&nbsp;</strong><a href="/api/files/f1fb6db8-a2e2-425b-b15e-8e401eb22d44/CZE-1528295206-proc-20180724134457.hdf?versionId=3294ae0b-c3b2-4155-ad36-311d2bb06ed9">CZE-1528295206-proc-20180724134457.hdf</a></p> <p>The index is divided into parts according to geographical boundaries defined by country borders. In this case it contains only single part - CZE. This information along with the creation time is stored as an attribute in the root group of the file.</p> <p>The graph topology is stored in the following way. Nodes have assigned a row index in the edges dataset which points to an outbound edge plus a number of the subsequent edges which are also output to this node. The edge metadata are stored in the EdgeData dataset to avoid redundancy. NodeMap dataset provides a convenient way to query nodes based on their uniqe identifiers.</p> <p><strong>File structure:</strong></p> <pre><code class="language-javascript">HDF5 "CZE-1528295206-proc-20180724134457.hdf" { GROUP "/" { GROUP "Index" { ATTRIBUTE "CreationTime" { DATATYPE H5T_STD_I64LE DATASPACE SCALAR DATA { (0): 1535451896 } } ATTRIBUTE "PartsCount" { DATATYPE H5T_STD_I32LE DATASPACE SCALAR DATA { (0): 1 } } ATTRIBUTE "PartsInfo" { DATATYPE H5T_COMPOUND { H5T_STRING { STRSIZE 4; STRPAD H5T_STR_NULLPAD; CSET H5T_CSET_ASCII; CTYPE H5T_C_S1; } "id"; H5T_STD_I64LE "nodeCount"; H5T_STD_I64LE "edgeCount"; } DATASPACE SIMPLE { ( 1 ) / ( 1 ) } DATA { (0): { "CZE\000", 904085, 2223222 } } } GROUP "CZE" { ATTRIBUTE "PartInfo" { DATATYPE H5T_STRING { STRSIZE 4; STRPAD H5T_STR_NULLPAD; CSET H5T_CSET_ASCII; CTYPE H5T_C_S1; } DATASPACE SCALAR DATA { (0): "CZE\000" } } DATASET "Edges" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "edgeId"; H5T_STD_I32LE "nodeIndex"; H5T_STD_I32LE "computed_speed"; H5T_STD_I32LE "length"; H5T_STD_I32LE "edgeDataIndex"; } DATASPACE SIMPLE { ( 2223222, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } DATASET "Nodes" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "id"; H5T_STD_I32LE "latitudeInt"; H5T_STD_I32LE "longtitudeInt"; H5T_STD_U8LE "edgeOutCount"; H5T_STD_I32LE "edgeOutIndex"; H5T_STD_U8LE "edgeInCount"; } DATASPACE SIMPLE { ( 904085, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } } DATASET "EdgeData" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "id"; H5T_STD_U8LE "speed"; H5T_STD_U8LE "funcClass"; H5T_STD_U8LE "lanes"; H5T_STD_U8LE "vehicleAccess"; H5T_STD_U16LE "specificInfo"; H5T_STD_U16LE "maxWeight"; H5T_STD_U16LE "maxHeight"; H5T_STD_U8LE "maxAxleLoad"; H5T_STD_U8LE "maxWidth"; H5T_STD_U8LE "maxLength"; H5T_STD_I8LE "incline"; } DATASPACE SIMPLE { ( 2838, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } DATASET "NodeMap" { DATATYPE H5T_COMPOUND { H5T_STD_I32LE "nodeId"; H5T_STRING { STRSIZE 4; STRPAD H5T_STR_NULLPAD; CSET H5T_CSET_ASCII; CTYPE H5T_C_S1; } "partId"; H5T_STD_I32LE "nodeIndex"; } DATASPACE SIMPLE { ( 904085, 1 ) / ( H5S_UNLIMITED, H5S_UNLIMITED ) } } } } }</code></pre> <p><strong>Spatial index</strong></p> <p><strong>File:&nbsp;</strong><a href="https://zenodo.org/api/files/f1fb6db8-a2e2-425b-b15e-8e401eb22d44/CZE-1528295206-proc-20180829135251.sqlite?versionId=56bb1e60-ff7c-407f-8041-3c628ef78db0">CZE-1528295206-proc-20180829135251.sqlite</a><br> &nbsp;</p> <p>SQLite database contains two primary tables. Table&nbsp;<em>nodes</em>&nbsp;and table&nbsp;<em>segments</em>, that&nbsp;resluts&nbsp;of selection and projection from the tables of the same name in primary db. Both tables are suitable for&nbsp;<em>searching nearest lines</em>&nbsp;task and&nbsp;<em>searching closest node of nearest line</em>&nbsp;task.&nbsp;</p> <pre><code class="language-sql">CREATE TABLE nodes  (      gid INTEGER,      part_gid INTEGER,      node_type INTEGER,      geom POINT  );  CREATE TABLE segments  (      gid INTEGER,      node_gid_from INTEGER,       node_gid_to INTEGER,      frc TEXT,      transition_time DOUBLE,      computed_speed REAL,      geom_length DOUBLE,      geom LINESTRING  ); </code></pre> <p>There are three more tables. Table&nbsp;<em>segments_rt</em>&nbsp;is a copy of table&nbsp;<em>segments</em>, that exclude loops in segments, hence it supports network routing task in&nbsp;SpatiaLite. Next table is&nbsp;<em>rt_network</em>&nbsp;(static graph generated from table&nbsp;<em>segments_rt</em>&nbsp;suitable for routing). The last one is table&nbsp;<em>virtual_rt_network</em>, that is an interface for routing query.&nbsp;&nbsp;</p> <p>&nbsp;</p>

openodc-byDec 2018View details →
zenodo44/100

Dataset from VR Streaming Server (Emulated) and Radio Access Network for Streaming Traffic

<p>The dataset contains an experiment in&nbsp;&nbsp;a site where UEs attach to a gNodeB that provides access to a streaming server that is stressed with high demanding transcoding workloads to emulate VR/AR processes. The UEs are realized through the Remote UE mode enabled by Amarisoft Simbox emulator, and the gNodeB is realized through the Amarisoft Callbox, which also provides the user plane function. The emulated VR streaming server is deployed as a Nginx pod in a Kubernetes cluster.</p> <p>We rely on MonB5G sampling functions that feed monitoring data (CPU and RAN parameters) to the monitoring system.&nbsp;A streaming video server has been deployed with the help of a NGINX server. It provides video-on-demand and video streaming, which can be accessed by any user (or UE) for real-time reproduction. This VR video streaming emulation aids to assess the performance of the network and therefore the benefits that each solution has brought. The video &ldquo;Big Buck Bunny&rdquo; with h.264 encoding and a resolution of 1920x1080p has been used for the experiments.The description of dataset features&nbsp;are:<br> 1-) Index Number,<br> 2-) Time: Time of the experiment,<br> 3-) N:&nbsp;number of VR streaming clients,<br> 4-) C: Average CPU of VR streaming server [mc]<br> 5-) O: Outbound traffic at the server average outbound traffic (O) flowing from the data interface of the video server.&nbsp;<br> 6-) R: Instantaneous downlink bit rate [Mbps],</p> <p>The original video file information:</p> <table> <tbody> <tr> <td> <p>Video codec&nbsp;</p> </td> <td> <p>Advanced Video Codec (AVC)&nbsp;</p> </td> </tr> <tr> <td> <p>Width&nbsp;</p> </td> <td> <p>1920 pixels&nbsp;</p> </td> </tr> <tr> <td> <p>Height&nbsp;</p> </td> <td> <p>1080 pixels&nbsp;</p> </td> </tr> <tr> <td> <p>Display aspect radio&nbsp;</p> </td> <td> <p>16:9&nbsp;</p> </td> </tr> <tr> <td> <p>Duration&nbsp;</p> </td> <td> <p>10 min 34 s&nbsp;</p> </td> </tr> <tr> <td> <p>Max Bitrate&nbsp;</p> </td> <td> <p>16.7 Mb/s&nbsp;</p> </td> </tr> <tr> <td> <p>Frame rate&nbsp;</p> </td> <td> <p>30 FPS&nbsp;</p> </td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo44/100

CTU-SME-11: a labeled dataset with real benign and malicious network traffic mimicking a small medium-size enterprise environment

<p>As technology advances, the number and complexity of cyber-attacks increase, forcing defense techniques to be updated and improved. To help develop effective tools for detecting security threats it is essential to have reliable and representative security datasets. Many existing security datasets have limitations that make them unsuitable for research, including lack of labels, unbalanced traffic, and outdated threats.</p> <p>CTU-SME-11 is a labeled network dataset designed to address the limitations of previous datasets. The dataset was captured in a real network that mimics a small-medium enterprise setting. Raw network traffic (packets) was captured from 11 devices using tcpdump for a duration of 7 days, from 20th to 26th of February, 2023 in Prague, Czech Republic. The devices were chosen based on the enterprise setting and consists of IoT, desktop and mobile devices, both bare metal and virtualized. The devices were infected with malware or exposed to Internet attacks, and factory reset to restore benign behavior.&nbsp;</p> <p>The raw data was processed to generate network flows (Zeek logs) which were analyzed and labeled. The dataset contains two types of levels, a high level label and a descriptive label, which were put by experts. The former can take three values, benign, malicious or background. The latter contains detailed information about the specific behavior observed in the network flows. The dataset contains 99 million labeled network flows. The overall compressed size of the dataset is 80GB and the uncompressed size is 170GB.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Network traffic datasets created by Single Flow Time Series Analysis

<p><strong>Network traffic datasets created by Single Flow Time Series Analysis</strong></p> <p>Datasets were created for the paper: Network Traffic Classification based on Single Flow Time Series Analysis -- Josef Koumar, Karel Hynek, Tom&aacute;&scaron; Čejka -- which was published at The 19th International Conference on Network and Service Management (CNSM) 2023. Please cite usage of our datasets as:<br>&nbsp;</p> <blockquote> <p>J. Koumar, K. Hynek and T. Čejka, "Network Traffic Classification Based on Single Flow Time Series Analysis," <em>2023 19th International Conference on Network and Service Management (CNSM)</em>, Niagara Falls, ON, Canada, 2023, pp. 1-7, doi: 10.23919/CNSM59352.2023.10327876.</p> </blockquote> <p>This Zenodo repository contains 23 datasets created from 15 well-known published datasets which are cited in the table below. Each dataset contains 69 features created by Time Series Analysis of Single Flow Time Series. The detailed description of features from datasets is in the file: <em>feature_description.pdf</em></p> <p>&nbsp;</p> <p>In the following table is a description of each dataset file:</p> <table> <tbody> <tr> <td><strong>File name</strong></td> <td><strong>Detection problem</strong></td> <td><strong>Citation of original raw dataset</strong></td> </tr> <tr> <td>botnet_binary.csv&nbsp;</td> <td>Binary detection of botnet&nbsp;</td> <td>S. Garc&iacute;a et al. An Empirical Comparison of Botnet Detection Methods. Computers &amp; Security, 45:100&ndash;123, 2014.&nbsp;</td> </tr> <tr> <td>botnet_multiclass.csv&nbsp;</td> <td>Multi-class classification of botnet&nbsp;</td> <td>S. Garc&iacute;a et al. An Empirical Comparison of Botnet Detection Methods. Computers &amp; Security, 45:100&ndash;123, 2014.&nbsp;</td> </tr> <tr> <td>cryptomining_design.csv</td> <td>Binary detection of cryptomining; the design part&nbsp;</td> <td>Richard Pln&yacute; et al. Datasets of Cryptomining Communication. Zenodo, October 2022&nbsp;</td> </tr> <tr> <td>cryptomining_evaluation.csv&nbsp;</td> <td>Binary detection of cryptomining; the evaluation part&nbsp;</td> <td>Richard Pln&yacute; et al. Datasets of Cryptomining Communication. Zenodo, October 2022&nbsp;</td> </tr> <tr> <td>dns_malware.csv&nbsp;</td> <td>Binary detection of malware DNS&nbsp;</td> <td>Samaneh Mahdavifar et al. Classifying Malicious Domains using DNS Traffic Analysis. In DASC/PiCom/CBDCom/CyberSciTech 2021, pages 60&ndash;67. IEEE, 2021.&nbsp;</td> </tr> <tr> <td>doh_cic.csv&nbsp;</td> <td>Binary detection of DoH&nbsp;</td> <td> <p>Mohammadreza MontazeriShatoori et al. Detection of doh tunnels using time-series classification of encrypted traffic. In DASC/PiCom/CBDCom/CyberSciTech 2020, pages 63&ndash;70. IEEE, 2020&nbsp;</p> </td> </tr> <tr> <td>doh_real_world.csv&nbsp;</td> <td>Binary detection of DoH&nbsp;</td> <td>Kamil Jeř&aacute;bek et al. Collection of datasets with DNS over HTTPS traffic. Data in Brief, 42:108310, 2022&nbsp;</td> </tr> <tr> <td>dos.csv&nbsp;</td> <td>Binary detection of DoS&nbsp;</td> <td>Nickolaos Koroniotis et al. Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset. Future Gener. Comput. Syst., 100:779&ndash;796, 2019.</td> </tr> <tr> <td>edge_iiot_binary.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.</td> </tr> <tr> <td>edge_iiot_multiclass.csv</td> <td>Multi-class classification of IoT malware</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.</td> </tr> <tr> <td>https_brute_force.csv</td> <td>Binary detection of HTTPS Brute Force</td> <td>Jan Luxemburk et al. HTTPS Brute-force dataset with extended network flows, November 2020</td> </tr> <tr> <td>ids_cic_binary.csv</td> <td>Binary detection of intrusion in IDS</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108&ndash;116, 2018.</td> </tr> <tr> <td>ids_cic_multiclass.csv&nbsp;</td> <td>Multi-class classification of intrusion in IDS&nbsp;</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108&ndash;116, 2018.&nbsp;</td> </tr> <tr> <td>ids_unsw_nb_15_binary.csv&nbsp;</td> <td>Binary detection of intrusion in IDS&nbsp;</td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1&ndash;6. IEEE, 2015.</td> </tr> <tr> <td>ids_unsw_nb_15_multiclass.csv&nbsp;</td> <td>Multi-class classification of intrusion in IDS&nbsp;</td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1&ndash;6. IEEE, 2015.</td> </tr> <tr> <td>iot_23.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Sebastian Garcia et al. IoT-23: A labeled dataset with malicious and benign IoT network traffic, January 2020. More details here https://www.stratosphereips.org /datasets-iot23</td> </tr> <tr> <td>ton_iot_binary.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021</td> </tr> <tr> <td>ton_iot_multiclass.csv&nbsp;</td> <td>Multi-class classification of IoT malware&nbsp;</td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021</td> </tr> <tr> <td>tor_binary.csv&nbsp;</td> <td>Binary detection of TOR&nbsp;</td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253&ndash;262. SciTePress, 2017.&nbsp;</td> </tr> <tr> <td>tor_multiclass.csv&nbsp;</td> <td>Multi-class classification of TOR&nbsp;</td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253&ndash;262. SciTePress, 2017.&nbsp;</td> </tr> <tr> <td>vpn_iscx_binary.csv&nbsp;</td> <td>Binary detection of VPN&nbsp;</td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407&ndash;414, 2016.&nbsp;</td> </tr> <tr> <td>vpn_iscx_multiclass.csv&nbsp;</td> <td>Multi-class classification of VPN&nbsp;</td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407&ndash;414, 2016.&nbsp;</td> </tr> <tr> <td>vpn_vnat_binary.csv&nbsp;</td> <td>Binary detection of VPN&nbsp;</td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022</td> </tr> <tr> <td>vpn_vnat_multiclass.csv</td> <td>Multi-class classification of VPN&nbsp;</td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Network traffic datasets with novel extended IP flow called NetTiSA flow

<p><strong>Network traffic datasets with novel extended IP flow called NetTiSA flow</strong></p> <p>Datasets were created for the paper: NetTiSA: Extended IP Flow with Time-series Features for Universal Bandwidth-constrained High-speed Network Traffic Classification -- Josef Koumar, Karel Hynek, Jaroslav Pe&scaron;ek, Tom&aacute;&scaron; Čejka -- which is published in The International Journal of Computer and Telecommunications Networking&nbsp;<a href="https://doi.org/10.1016/j.comnet.2023.110147" rel="nofollow">https://doi.org/10.1016/j.comnet.2023.110147</a><br><br>Please cite the usage of our datasets as:</p> <blockquote> <p>Josef Koumar, Karel Hynek, Jaroslav Pe&scaron;ek, Tom&aacute;&scaron; Čejka, "NetTiSA: Extended IP flow with time-series features for universal bandwidth-constrained high-speed network traffic classification", Computer Networks, Volume 240, 2024, 110147, ISSN 1389-1286<br><br></p> <pre><code>@article{KOUMAR2024110147, title = {NetTiSA: Extended IP flow with time-series features for universal bandwidth-constrained high-speed network traffic classification}, journal = {Computer Networks}, volume = {240}, pages = {110147}, year = {2024}, issn = {1389-1286}, doi = {https://doi.org/10.1016/j.comnet.2023.110147}, url = {https://www.sciencedirect.com/science/article/pii/S1389128623005923}, author = {Josef Koumar and Karel Hynek and Jaroslav Pe&scaron;ek and Tom&aacute;&scaron; Čejka} } </code></pre> </blockquote> <p>This Zenodo repository contains 23 datasets created from 15 well-known published datasets, which are cited in the table below. Each dataset contains the NetTiSA flow feature vector.<br><br>&nbsp;</p> <p><strong>NetTiSA flow feature vector</strong></p> <p><br>The novel extended IP flow called NetTiSA (Network Time Series Analysed) flow contains a universal bandwidth-constrained feature vector consisting of 20 features. We divide the NetTiSA flow classification features into three groups by computation. The first group of features is based on classical bidirectional flow information---a number of transferred bytes, and packets.&nbsp; The second group contains statistical and time-based features calculated using the time-series analysis of the packet sequences. The third type of features can be computed from the previous groups (i.e., on the flow collector) and improve the classification performance without any impact on the telemetry bandwidth.</p> <p>&nbsp;</p> <p><strong>Flow features</strong></p> <p>The flow features are:</p> <ul> <li><strong><em>Packets</em></strong> is the number of packets in the direction from the source to the destination IP address.</li> <li><em><strong>Packets in reverse order</strong></em> is the number of packets in the direction from the destination to the source IP address.</li> <li><strong><em>Bytes</em> </strong>is the size of the payload in bytes transferred in the direction from the source to the destination IP address.</li> <li><strong><em>Bytes in reverse order</em></strong> is the size of the payload in bytes transferred in the direction from the destination to the source IP address.</li> </ul> <p>&nbsp;</p> <p><strong>Statistical and Time-based features</strong></p> <p>The features that are exported in the extended part of the flow. All of them can be computed (exactly or in approximative) by stream-wise computation, which is necessary for keeping memory requirements low. The second type of feature set contains the following features:</p> <ul> <li><strong><em>Mean</em></strong> represents mean of the payload lengths of packets</li> <li><strong><em>Min</em></strong> is the minimal value from payload lengths of all packets in a flow</li> <li><strong><em>Max</em></strong> is the maximum value from payload lengths of all packets in a flow</li> <li><strong><em>Standard deviation</em></strong> is a measure of the variation of payload lengths from the mean payload length</li> <li><strong><em>Root mean square</em></strong> is the measure of the magnitude of payload lengths of packets</li> <li><strong><em>Average dispersion</em></strong> is the average absolute difference between each payload length of the packet and the mean value</li> <li><strong><em>Kurtosis</em></strong> is the measure describing the extent to which the tails of a distribution differ from the tails of a normal distribution</li> <li><em><strong>Mean of relative times</strong></em> is the mean of the relative times which is a sequence defined as <span>\(st = \{t_1 - t_1, t_2 - t_1, ..., t_n - t_1\} \)</span></li> <li><em><strong>Mean of time differences</strong></em> is the mean of the time differences which is a sequence defined as <span>\(dt = \{ t_j - t_i | j = i + 1, i \in \{1, 2, \dots, n - 1\} \}.\)</span></li> <li><em><strong>Min from time differences</strong></em> is the minimal value from all time differences, i.e., min space between packets.</li> <li><em><strong>Max from time differences</strong></em> is the maximum value from all time differences, i.e., max space between packets.</li> <li><em><strong>Time distribution</strong></em> describes the deviation of time differences between individual packets within the time series. The feature is computed by the following equation:<br><span>\(tdist = \frac{ \frac{1}{n-1} \sum_{i=1}^{n-1} \left| \mu_{\{dt_{n-1}\}} - dt_i \right| }{ \frac{1}{2} \left(max\left(\{dt_{n-1}\}\right) - min\left(\{dt_{n-1}\}\right) \right) }\)</span></li> <li><em><strong>Switching ratio</strong></em> represents a value change ratio (switching) between payload lengths. The switching ratio is computed by equation:<br><span>\(sr = \frac{s_n}{\frac{1}{2} (n - 1)}\)</span></li> </ul> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; where&nbsp;<span>\(s_n\)</span> is number of switches.</p> <p>&nbsp;&nbsp;</p> <p><strong>Features computed at the collector</strong><br>The third set contains features that are computed from the previous two groups prior to classification. Therefore, they do not influence the network telemetry size and their computation does not put additional load to resource-constrained flow monitoring probes. The NetTiSA flow combined with this feature set is called the Enhanced NetTiSA flow and contains the following features:</p> <ul> <li><em><strong>Max minus min</strong></em>&nbsp; is the difference between minimum and maximum payload lengths</li> <li><em><strong>Percent deviation</strong></em> is the dispersion of the average absolute difference to the mean value</li> <li><em><strong>Variance</strong></em> is the spread measure of the data from its mean</li> <li><em><strong>Burstiness</strong></em> is the degree of peakedness in the central part of the distribution</li> <li><em><strong>Coefficient of variation</strong></em> is a dimensionless quantity that compares the dispersion of a time series to its mean value and is often used to compare the variability of different time series that have different units of measurement</li> <li><em><strong>Directions</strong></em> describe a percentage ratio of packet direction computed as <span>\(\frac{d_1}{ d_1 + d_0}\)</span>, where&nbsp;<span>\(d_1\)</span> is a number of packets in a direction from source to destination IP address and&nbsp;<span>\(d_0\)</span> the opposite direction. Both&nbsp;&nbsp;<span>\(d_1\)</span> and&nbsp;<span>\(d_0\)</span> are inside the classical bidirectional flow.</li> <li><em><strong>Duration</strong></em> is the duration of the flow</li> </ul> <p>&nbsp;</p> <p>The NetTiSA flow is implemented into IP flow exporter <a href="https://github.com/CESNET/ipfixprobe">ipfixprobe</a>.</p> <p>&nbsp;</p> <p><strong>Description of dataset files</strong></p> <p>In the following table is a description of each dataset file:</p> <table> <tbody> <tr> <td> <p><strong>File name</strong></p> </td> <td> <p><strong>Detection problem</strong></p> </td> <td> <p><strong>Citation of the original raw dataset</strong></p> </td> </tr> <tr> <td>botnet_binary.csv&nbsp;</td> <td>Binary detection of botnet&nbsp;</td> <td>S. Garc&iacute;a et al. An Empirical Comparison of Botnet Detection Methods. Computers &amp; Security, 45:100&ndash;123, 2014.&nbsp;</td> </tr> <tr> <td>botnet_multiclass.csv&nbsp;</td> <td>Multi-class classification of botnet&nbsp;</td> <td>S. Garc&iacute;a et al. An Empirical Comparison of Botnet Detection Methods. Computers &amp; Security, 45:100&ndash;123, 2014.&nbsp;</td> </tr> <tr> <td>cryptomining_design.csv&nbsp;</td> <td>Binary detection of cryptomining; the design part&nbsp;</td> <td>Richard Pln&yacute; et al. Datasets of Cryptomining Communication. Zenodo, October 2022&nbsp;</td> </tr> <tr> <td>cryptomining_evaluation.csv&nbsp;</td> <td>Binary detection of cryptomining; the evaluation part&nbsp;</td> <td>Richard Pln&yacute; et al. Datasets of Cryptomining Communication. Zenodo, October 2022&nbsp;</td> </tr> <tr> <td>dns_malware.csv&nbsp;</td> <td>Binary detection of malware DNS&nbsp;</td> <td>Samaneh Mahdavifar et al. Classifying Malicious Domains using DNS Traffic Analysis. In DASC/PiCom/CBDCom/CyberSciTech 2021, pages 60&ndash;67. IEEE, 2021.&nbsp;</td> </tr> <tr> <td>doh_cic.csv&nbsp;</td> <td>Binary detection of DoH&nbsp;</td> <td>Mohammadreza MontazeriShatoori et al. Detection of doh tunnels using time-series classification of encrypted traffic. In DASC/PiCom/CBDCom/CyberSciTech 2020, pages 63&ndash;70. IEEE, 2020&nbsp;</td> </tr> <tr> <td>doh_real_world.csv&nbsp;</td> <td>Binary detection of DoH&nbsp;</td> <td>Kamil Jeř&aacute;bek et al. Collection of datasets with DNS over HTTPS traffic. Data in Brief, 42:108310, 2022&nbsp;</td> </tr> <tr> <td>dos.csv&nbsp;</td> <td>Binary detection of DoS&nbsp;</td> <td>Nickolaos Koroniotis et al. Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset. Future Gener. Comput. Syst., 100:779&ndash;796, 2019.&nbsp;</td> </tr> <tr> <td>edge_iiot_binary.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.&nbsp;</td> </tr> <tr> <td>edge_iiot_multiclass.csv&nbsp;</td> <td>Multi-class classification of IoT malware&nbsp;</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.&nbsp;</td> </tr> <tr> <td>https_brute_force.csv&nbsp;</td> <td>Binary detection of HTTPS Brute Force&nbsp;</td> <td>Jan Luxemburk et al. HTTPS Brute-force dataset with extended network flows, November 2020&nbsp;</td> </tr> <tr> <td>ids_cic_binary.csv&nbsp;</td> <td>Binary detection of intrusion in IDS&nbsp;</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108&ndash;116, 2018.&nbsp;</td> </tr> <tr> <td>ids_cic_multiclass.csv&nbsp;</td> <td>Multi-class classification of intrusion in IDS&nbsp;</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108&ndash;116, 2018.&nbsp;</td> </tr> <tr> <td>unsw_binary.csv&nbsp;</td> <td>Binary detection of intrusion in IDS&nbsp;</td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1&ndash;6. IEEE, 2015.&nbsp;</td> </tr> <tr> <td>unsw_multiclass.csv&nbsp;</td> <td>Multi-class classification of intrusion in IDS&nbsp;</td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1&ndash;6. IEEE, 2015.&nbsp;</td> </tr> <tr> <td>iot_23.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Sebastian Garcia et al. IoT-23: A labeled dataset with malicious and benign IoT network traffic, January 2020. More details here https://www.stratosphereips.org /datasets-iot23&nbsp;</td> </tr> <tr> <td>ton_iot_binary.csv&nbsp;</td> <td>Binary detection of IoT malware&nbsp;</td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021&nbsp;</td> </tr> <tr> <td>ton_iot_multiclass.csv&nbsp;</td> <td>Multi-class classification of IoT malware&nbsp;</td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021&nbsp;</td> </tr> <tr> <td>tor_binary.csv&nbsp;</td> <td>Binary detection of TOR&nbsp;</td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253&ndash;262. SciTePress, 2017.&nbsp;</td> </tr> <tr> <td>tor_multiclass.csv&nbsp;</td> <td>Multi-class classification of TOR&nbsp;</td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253&ndash;262. SciTePress, 2017.&nbsp;</td> </tr> <tr> <td>vpn_iscx_binary.csv&nbsp;</td> <td>Binary detection of VPN&nbsp;</td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407&ndash;414, 2016.&nbsp;</td> </tr> <tr> <td>vpn_iscx_multiclass.csv&nbsp;</td> <td>Multi-class classification of VPN&nbsp;</td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407&ndash;414, 2016.&nbsp;</td> </tr> <tr> <td>vpn_vnat_binary.csv&nbsp;</td> <td>Binary detection of VPN&nbsp;</td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022&nbsp;</td> </tr> <tr> <td>vpn_vnat_multiclass.csv&nbsp;</td> <td>Multi-class classification of VPN&nbsp;</td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Host network traffic time series 2019/01

<p><em><strong>General info</strong></em></p> <p>Dataset&nbsp;was collected over one <strong>month period in January 2019</strong>. The observation points for the collection of IP flows were located at the borders of the university campus network. The campus university network has /16 CIDR IPv4 network range at disposal and contains various network segments from segments connecting dormitories, over server segments, to a segment containing working stations of university administrative workers. The size of the raw IP flows used to create the dataset was over 860GB. <strong>A host in our dataset is identified by its source IPv4 address. &nbsp;</strong><br> &nbsp;</p> <p><em><strong>Variables</strong></em></p> <p>The dataset contains the following variables:</p> <ul> <li><strong>Aggregations</strong> - created from five-minute total volumes aggregated&nbsp;over&nbsp;one-hour disjoint windows using&nbsp;mean/max/min aggregation functions <ul> <li><strong># of flows (FL) </strong>- number of flows for a given source IP&nbsp;</li> <li><strong># of packets (PKT)</strong> -&nbsp;number of packets for a given source IP</li> <li><strong># of bytes (BYT)</strong> -&nbsp;number of packets for a given source IP</li> <li><strong>flow duration (DUR)</strong> - average flow duration in seconds</li> </ul> </li> <li><strong>Distinct Counts&nbsp;</strong>- count of distinct values for each variable in five-minute window aggregated&nbsp;over&nbsp;one-hour disjoint windows using&nbsp;mean/max/min aggregation functions <ul> <li><strong># of peers (PEER)</strong> - number of distinct communication peers for a given source IP</li> <li><strong># of ports (PORTS)</strong> - number of distinct destination ports&nbsp;for a given source IP</li> <li><strong># of protocols (PROTO)</strong> - number of distinct communication protocols&nbsp;for a given source IP</li> <li><strong># of AS numbers (AS)</strong> - number of distinct destination AS numbers for a given source IP</li> <li><strong># of countries (CTRY)</strong> - number of distinct destination countries&nbsp;for a given source IP</li> </ul> </li> <li><strong>Labels</strong> <ul> <li><strong>Range (RNG)</strong> - a network range a host belongs to (anonymized)</li> <li><strong>Unit (UNT) </strong>- an administrative unit owning the network range</li> <li><strong>Sub-unit (SUB-UNT)</strong> - a sub-unit of the unit</li> </ul> </li> </ul> <p>&nbsp;</p> <p><em><strong>Dataset format</strong></em></p> <ul> <li>The dataset is in <strong>comma-separated values (CSV)</strong> format.&nbsp;</li> <li><strong>Header</strong> - multilevel, first 3 lines <ul> <li>1 level - aggregation type {mean|min|max}</li> <li>2 level - variable {see above}</li> <li>3 level - hour of a day {00,01,02,03,...,22,23}</li> </ul> </li> <li><strong>Lablels</strong> - last 4 columns</li> <li><strong>Dataset size&nbsp;</strong> <ul> <li>rows: 65536 host records&nbsp;+ 3 headers</li> <li>columns: 648 variables + 4 labels</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Curated Research on Network Traffic Analysis

<p>With the NTA Database we aim to collect relevant information about the research in network traffic analysis conducted during the last years. To this end, we have curated related papers from journals and conferences and stored the extracted data in JSON files.&nbsp;</p>

opencc-by-4.0May 2018View details →
zenodo40/100

CESNET-QUIC22: A large one-month QUIC network traffic dataset from backbone lines

<p><strong>Please refer to the original data article for further data description:&nbsp;</strong>Jan Luxemburk et al. CESNET-QUIC22: A&nbsp;large one-month QUIC network traffic dataset from backbone lines, Data in Brief, 2023, 108888, ISSN 2352-3409, <a href="https://doi.org/10.1016/j.dib.2023.108888">https://doi.org/10.1016/j.dib.2023.108888</a>.&nbsp;</p><p><strong>We recommend using the</strong> <strong>CESNET DataZoo python library, which facilitates the work with large network traffic datasets. </strong>More information about the DataZoo project can be found in the GitHub repository <a href="https://github.com/CESNET/cesnet-datazoo">https://github.com/CESNET/cesnet-datazoo</a>.</p><p>The QUIC (Quick UDP Internet Connection) protocol has the potential to replace TLS over TCP, which is the standard choice for reliable and secure Internet communication. Due to its design that makes the inspection of QUIC handshakes challenging and its usage in HTTP/3, there is an increasing demand for research in QUIC traffic analysis. This dataset contains one month of QUIC traffic collected in an ISP backbone network, which connects 500 large institutions and serves around half a million people. The data are delivered as enriched flows that can be useful for various network monitoring tasks. The provided server names and packet-level information allow research in the encrypted traffic classification area. Moreover, included QUIC versions and user agents (smartphone, web browser, and operating system identifiers) provide information for large-scale QUIC deployment studies.</p><p><strong>Data capture</strong>&nbsp;The data was captured in the flow monitoring infrastructure of the <a href="https://www.cesnet.cz">CESNET2</a>&nbsp;network. The capturing was done for four weeks between 31.10.2022 and 27.11.2022. The following list provides per-week flow count, capture period, and uncompressed size:</p><ul><li><strong>W-2022-44</strong><ul><li>Uncompressed Size: 19 GB</li><li>Capture Period: 31.10.2022 - 6.11.2022</li><li>Number of flows: 32.6M</li></ul></li><li><strong>W-2022-45</strong><ul><li>Uncompressed Size: 25 GB</li><li>Capture Period: 7.11.2022 - 13.11.2022</li><li>Number of flows: 42.6M</li></ul></li><li><strong>W-2022-46</strong><ul><li>Uncompressed Size: 20 GB</li><li>Capture Period: 14.11.2022 - 20.11.2022</li><li>Number of flows: 33.7M</li></ul></li><li><strong>W-2022-47</strong><ul><li>Uncompressed Size: 25 GB</li><li>Capture Period: 21.11.2022 - 27.11.2022</li><li>Number of flows: 44.1M</li></ul></li><li><strong>CESNET-QUIC22&nbsp;</strong><ul><li>Uncompressed Size: 89 GB</li><li>Capture Period: 31.10.2022 - 27.11.2022</li><li>Number of flows: 153M</li></ul></li></ul><p>&nbsp;</p><p><strong>Data description</strong> The dataset consists of network flows describing encrypted QUIC communications. Flows were created using <a href="https://github.com/CESNET/ipfixprobe">ipfixprobe</a> flow exporter and are extended with packet metadata sequences, packet histograms, and with fields extracted from the QUIC Initial Packet, which is the first packet of the QUIC connection handshake. The extracted handshake fields are the Server Name Indication (SNI) domain, the used version of the QUIC protocol, and the user agent string that is available in a subset of QUIC communications.</p><p><strong>Packet Sequences</strong> Flows in the dataset are extended with sequences of packet sizes, directions, and inter-packet times. For the packet sizes, we consider payload size after transport headers (UDP headers for the QUIC case). Packet directions are encoded as ±1, <i>+1</i>&nbsp;meaning a packet sent from client to server, and <i>-1</i>&nbsp;a packet from server to client. Inter-packet times depend on the location of communicating hosts, their distance, and on the network conditions on the path. However, it is still possible to extract relevant information that correlates with user interactions and, for example, with the time required for an API/server/database to process the received data and generate the response to be sent in the next packet. Packet metadata sequences have a length of 30, which is the default setting of the used flow exporter. We also derive three fields from each packet sequence: its length, time duration, and the number of roundtrips. The roundtrips are counted as the number of changes in the communication direction (from packet directions data); in other words, each client request and server response pair counts as one roundtrip.</p><p><strong>Flow statistics</strong> Flows also include standard flow statistics, which represent aggregated information about the entire bidirectional flow. The fields are: the number of transmitted bytes and packets in both directions, the duration of flow, and packet histograms. Packet histograms include binned counts of packet sizes and inter-packet times of the entire flow in both directions (more information in the <a href="https://github.com/CESNET/ipfixprobe/tree/master#phists">PHISTS plugin documentation</a>&nbsp;There are eight bins with a logarithmic scale; the intervals are 0-15, 16-31, 32-63, 64-127, 128-255, 256-511, 512-1024, &gt;1024&nbsp;[ms or B]. The units are milliseconds for inter-packet times and bytes for packet sizes. Moreover, each flow has its&nbsp;end reason - either it was idle, reached the active timeout, or ended due to other reasons. This corresponds with the official <a href="https://www.iana.org/assignments/ipfix/ipfix.xhtml#ipfix-flow-end-reason">IANA IPFIX-specified values</a>. The <i>FLOW_ENDREASON_OTHER</i>&nbsp;field represents the <i>forced end</i>&nbsp;and <i>lack of resources</i>&nbsp;reasons. The <i>end of flow detected</i>&nbsp;reason is not considered because it is not relevant for UDP connections.</p><p><strong>Dataset structure</strong> The dataset flows&nbsp;are delivered in compressed CSV files. CSV files contain one flow per row; data columns are summarized in the provided list below. For each flow data file, there is a JSON file with the number of saved and seen (before sampling) flows per service and total counts of all received (observed on the CESNET2 network), service (belonging to one of the dataset's services), and saved (provided in the dataset) flows. There is also the <i>stats-week.json</i> file aggregating flow counts of a whole week and the <i>stats-dataset.json</i> file aggregating flow counts for the entire dataset. Flow counts before sampling can be used to compute sampling ratios of individual services and to resample the dataset back to the original service distribution. Moreover, various dataset statistics, such as feature distributions and value counts of QUIC versions and user agents, are provided in the&nbsp;<i>dataset-statistics</i>&nbsp;folder. The mapping between services and service providers is provided in the&nbsp;<i>servicemap.csv</i>&nbsp;file, which also includes SNI domains used for ground truth labeling.&nbsp;The following list describes flow data fields in CSV files:</p><ul><li><strong>ID:</strong> Unique identifier</li><li><strong>SRC_IP:</strong> Source IP address</li><li><strong>DST_IP:</strong> Destination IP address</li><li><strong>DST_ASN:</strong> Destination Autonomous System number</li><li><strong>SRC_PORT:</strong> Source port</li><li><strong>DST_PORT:</strong> Destination port</li><li><strong>PROTOCOL:</strong> Transport protocol</li><li><strong>QUIC_VERSION QUIC:</strong> protocol version</li><li><strong>QUIC_SNI:</strong> Server Name Indication domain</li><li><strong>QUIC_USER_AGENT:</strong> User agent string, if available in the QUIC Initial Packet</li><li><strong>TIME_FIRST:</strong> Timestamp of the first packet in format YYYY-MM-DDTHH-MM-SS.ffffff</li><li><strong>TIME_LAST:</strong> Timestamp of the last packet in format YYYY-MM-DDTHH-MM-SS.ffffff</li><li><strong>DURATION:</strong> Duration of the flow in seconds</li><li><strong>BYTES:</strong> Number of transmitted bytes from client to server</li><li><strong>BYTES_REV:</strong> Number of transmitted bytes from server to client</li><li><strong>PACKETS:</strong> Number of packets transmitted from client to server</li><li><strong>PACKETS_REV:</strong> Number of packets transmitted from server to client</li><li><strong>PPI:</strong> Packet metadata sequence in the format: [[inter-packet times], [packet directions], [packet sizes]]</li><li><strong>PPI_LEN:</strong> Number of packets in the PPI sequence</li><li><strong>PPI_DURATION:</strong> Duration of the PPI sequence in seconds</li><li><strong>PPI_ROUNDTRIPS:</strong> Number of roundtrips in the PPI sequence</li><li><strong>PHIST_SRC_SIZES:</strong> Histogram of packet sizes from client to server</li><li><strong>PHIST_DST_SIZES: </strong>Histogram of packet sizes from server to client</li><li><strong>PHIST_SRC_IPT: </strong>Histogram of inter-packet times from client to server</li><li><strong>PHIST_DST_IPT:</strong> Histogram of inter-packet times from server to client</li><li><strong>APP:</strong> Web service label</li><li><strong>CATEGORY:</strong> Service category</li><li><strong>FLOW_ENDREASON_IDLE:</strong> Flow was terminated because it was idle</li><li><strong>FLOW_ENDREASON_ACTIVE:</strong> Flow was terminated because it reached the active timeout</li><li><strong>FLOW_ENDREASON_OTHER:</strong> Flow was terminated for other reasons</li></ul><p>&nbsp;</p><p><strong>Link&nbsp;to other CESNET datasets</strong></p><ul><li><a href="https://www.liberouter.org/technology-v2/tools-services-datasets/datasets/">https://www.liberouter.org/technology-v2/tools-services-datasets/datasets/</a></li><li><a href="https://github.com/CESNET/cesnet-datazoo">https://github.com/CESNET/cesnet-datazoo</a></li></ul><p><strong>Please cite the original data article:</strong></p><blockquote><p>@article{CESNETQUIC22, author = {Jan Luxemburk and Karel Hynek and Tomáš Čejka and Andrej Lukačovič and Pavel Šiška}, title = {CESNET-QUIC22: a large one-month QUIC network traffic dataset from backbone lines}, journal = {Data in Brief}, pages = {108888}, year = {2023}, issn = {2352-3409}, doi = {https://doi.org/10.1016/j.dib.2023.108888}, url = {https://www.sciencedirect.com/science/article/pii/S2352340923000069} }</p></blockquote>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Impacts of centralized control on mixed traffic network performance: A strategic games analysis

<p>This dataset contains the data that were used to assess the proposed framework within the context of the case study in the paper entitled "Impacts of centralized control on mdaixed traffic network per-formance: A strategic games analysis".</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Freeway Inductive Loop Detector Dataset for Network-wide Traffic Speed Prediction

<p>The data is collected by the inductive loop detectors deployed on freeways in Seattle area. The freeways contain&nbsp;I-5, I-405, I-90, and SR-520. This data set contains spatiotemporal speed information of the freeway system. At each milepost, the speed information collected from main lane loop detectors in the&nbsp;same direction are averaged and integrated into 5 minutes interval speed data. The raw&nbsp;data is provided by Washington Start Department of Transportation (WSDOT) and processed by the <a href="http://www.uwstarlab.org/">STAR Lab</a> in the University of Washington according to data quality control and data imputation procedures [1][2].&nbsp;&nbsp;</p> <p>The data file is a pickle file that can be easily read using the read_pickle() function in the Pandas package. The data forms as a matrix and each cell of the matrix is speed value for the specific milepost and time period. The&nbsp;horizontal header of the data set denotes the milepost and the vertical header indicates the timestamps. For more information on the definition of milepost, please refer to this <a href="http://data.wsdot.wa.gov/traffic/">website</a>.</p> <p>This data set been used for traffic prediction tasks in several research studies [3][4]. For more detailed information about the data set, you can also refer to this <a href="https://github.com/zhiyongc/Seattle-Loop-Data">link</a>.</p> <p><strong>References</strong>:</p> <p>[1].&nbsp;Henrickson, K., Zou, Y., &amp; Wang, Y. (2015). Flexible and robust method for missing loop detector data imputation.&nbsp;<em>Transportation Research Record</em>,&nbsp;<em>2527</em>(1), 29-36.</p> <p>[2]. Wang, Y., Zhang, W., Henrickson, K., Ke, R., &amp; Cui, Z. (2016).&nbsp;<em>Digital roadway interactive visualization and evaluation network applications to WSDOT operational data usage</em>&nbsp;(No. WA-RD 854.1). Washington (State). Dept. of Transportation.</p> <p>[3].&nbsp;Cui, Z., Ke, R., &amp; Wang, Y. (2018). Deep bidirectional and unidirectional LSTM recurrent neural network for network-wide traffic speed prediction.&nbsp;<em>arXiv preprint arXiv:1801.02143</em>.</p> <p>[4]. Cui, Z., Henrickson, K., Ke, R., &amp; Wang, Y. (2018). Traffic Graph Convolutional Recurrent Neural Network: A Deep Learning Framework for Network-Scale Traffic Learning and Forecasting.&nbsp;<em>arXiv preprint arXiv:1802.07007</em>.</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

Required data for simulating a typical large-scale urban traffic network

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo36/100

Cyber4OT: ICS network traces containing normal activity and full attack traffic

<p>The <em><strong>Cyber4OT</strong></em> dataset contains prepared in the test-bed environment packet traces from normal activity of OT network, as well as, full network attack. During recorded activity, the attacker performs full network reconnaissance, later disconnects legal Modbus TCP connection and performs PLC device hijacking.</p> <p>The dataset contains 96 files with more than 4,25 millions of packets.</p> <p><em><strong>ReadMe.txt</strong></em> file contains short description of each trace file content.</p> <p>Detailed description of the test bed, where data was prepared, is provided in the <em><strong>Cyber4OT_testbed_description.pdf</strong></em> file.</p>

opencc-by-nc-4.0Nov 2024View details →
zenodo36/100

Population estimation from mobile network traffic metadata

<p><em><strong>Please cite our paper if you publish material based on those datasets</strong></em></p> <blockquote> <p>G. Khodabandelou, V. Gauthier, M. El-Yacoubi, M. Fiore, &quot;Estimation of Static and Dynamic Urban Populations with Mobile Network Metadata&quot;, in IEEE Trans. on Mobile Computing, 2018 (in Press). <a href="http://dx.doi.org/10.1109/TMC.2018.2871156">10.1109/TMC.2018.2871156</a></p> </blockquote> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>Communication-enabled devices that are physically carried by individuals are today pervasive,<br> which opens unprecedented opportunities for collecting digital metadata about the mobility of large populations. In this paper, we propose a novel methodology for the estimation of people density at metropolitan scales, using subscriber presence metadata collected by a mobile operator. We show that our approach suits the estimation of static population densities, i.e., of the distribution of dwelling units per urban area contained in traditional censuses. Specifically, it achieves higher accuracy than that granted by previous equivalent solutions. In addition, our approach enables the estimation of dynamic population densities, i.e., the time-varying distributions of people in a conurbation. Our results build on significant real-world mobile network metadata and relevant ground-truth information in multiple urban scenarios.</p> <p><strong>Dataset Columns</strong></p> <p>This dataset cover one month of data taken during the month of April 2015 for three Italian cities: Rome, Milan, Turin. The raw data has been provided during the Telecom Italia Big Data Challenge (http://www.telecomitalia.com/tit/en/innovazione/archivio/big-data-challenge-2015.html)</p> <p>1. <strong>grid_id</strong>: the coordinate of the grid can be retrieved with the shapefile of a given city<br> 2. <strong>date</strong>: format Y-M-D H:M:S<br> 4. <strong>landuse_label</strong>: the land use label has been computed by through method described in [2]<br> 5. <strong>population</strong>: Census population of a given grid block as defined by the Istituto nazionale di statistica (ISTAT https://www.istat.it/en/censuses) in 2011<br> 6. <strong>estimation</strong>: Dynamics density population estimation (in person) as the result of the method described in [1]<br> 7. <strong>area</strong>: surface of the &quot;grid id&quot; considered in km^2<br> 8. <strong>geometry</strong>: the shape of the area considered with the EPSG:3003 coordinate system (only with quilt)</p> <p><strong>Note</strong></p> <p>Due to legal constraints, we cannot share directly the original data from the Telecom Italia Big Data Challenge we used to build this dataset.</p> <p><strong>Easy access to this dataset with&nbsp;quilt</strong></p> <p>Install the dataset repository:</p> <p>$ quilt install vgauthier/DynamicPopEstimate</p> <p>Use the dataset with a Panda Dataframe</p> <p>&gt;&gt;&gt; from quilt.data.vgauthier import DynamicPopEstimate<br> &gt;&gt;&gt; import pandas as pd<br> &gt;&gt;&gt; df = pd.DataFrame(DynamicPopEstimate.rome())<br> <br> Use the dataset with a GeoPanda Dataframe<br> <br> &gt;&gt;&gt; from quilt.data.vgauthier import DynamicPopEstimate<br> &gt;&gt;&gt; import geopandas as gpd<br> &gt;&gt;&gt; df = gpd.DataFrame(DynamicPopEstimate.rome())</p> <p><strong>References</strong></p> <p>[1] G. Khodabandelou, V. Gauthier, M. El-Yacoubi, M. Fiore, &quot;Population estimation from mobile network traffic metadata&quot;, in proc of the 17th International Symposium on A World of Wireless, Mobile and Multimedia Networks (WoWMoM), pp. 1 - 9, 2016.&nbsp;</p> <p>[2] A. Furno, M. Fiore, R. Stanica, C. Ziemlicki, and Z. Smoreda, &quot;A tale of ten cities: Characterizing signatures of mobile traffic in urban areas,&quot; IEEE Transactions on Mobile Computing, Volume: 16, Issue: 10, 2017.<br> &nbsp;</p>

openodc-odblOct 2017View details →
zenodo36/100

Relationship between Road Network Characteristics and Traffic Safety

<p>Corresponding data set for Tran-SET Project No. 17ITSTSA01. Abstract of the final report is stated below for reference:</p> <p>&quot;The Transportation and Capital Improvement of the City of San Antonio, Texas Department of Transportation (TxDOT) and other related agencies often make several efforts based on traffic data to improve safety at intersections, but the number of intersection crashes is still on the high side. There is no one size fits all solution for intersections and the City is often usually confronted with doing best value option analysis on different solutions to choose the least expensive yet more advancements. The goal of this project was to obtain the relationship between road network characteristics and public safety with a focus on intersections; perform a thorough analysis of critical intersections with high crash incidents and crash rates within the city of San Antonio, Texas, and analyze key factors that lead to crashes and recommend effective safety countermeasures. Researchers conducted the following tasks: literature review, crash data analysis, factors affecting crashes at intersections, and the development of possible solutions to some of the identified challenges. Several variables and factors were analyzed, including driver characteristics, like age and gender, road-related factors and environmental factors such as weather conditions and time of day ArcGIS was used to analyze crash frequency at different intersections, and hotspot analysis was carried out to identify high-risk intersections. The crash rates were also calculated for some intersections. The research outcome shows that there are more male drivers than female drivers involved in crashes, even though we have more licensed female drivers than male drivers. The highest number of crashes involved drivers within the age range of 15 &ndash; 34 years; this is an indication that intersection crash is one of the top threats to the young generation. The study also shows that the most common crash type is the angle crash which represents over 23% of the intersection crashes. Driver&rsquo;s inattention ranked first among all the contributing factors recorded. The highrisk intersections based on crash frequency and crash rate show that the intersection along the Bandera Road and Loop 1604 is the worst in the city, with 399 crashes and 8.5 crashes per million entering vehicles. The research concluded with some suggested countermeasures, which include public enlightenment and road safety audit as a proactive means of identifying high-risk intersections.&quot;</p>

opencc-by-4.0Nov 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record