Skip to main content
zenodoopen

Host Network Traffic 2019

<p><strong><em>Dataset Summary</em></strong></p> <ul> <li><strong>Timespan</strong>: 2019-01-01 : 2019-12-31</li> <li><strong>Granularity:&nbsp;</strong>1-hour disjoint time windows</li> <li><strong># of&nbsp;characteristics observed:&nbsp;</strong>9</li> <li><strong>Hosts observed: </strong>65536</li> <li><strong>Labels:&nbsp;</strong>included</li> <li><strong>Unzipped volume:&nbsp;</strong>approx. 10 GB</li> </ul> <p><strong><em>Dataset Origins</em></strong></p> <p>Dataset&nbsp;was collected over the <strong>whole year</strong>&nbsp;<strong>&nbsp;2019</strong>. The observation points for the collection of IP flows were located at the borders of the university campus network. The campus university network has /16 CIDR IPv4 network range at disposal and contains various network segments from segments connecting dormitories, over server segments, to a segment containing working stations of university administrative workers.&nbsp;<strong>A host in our dataset is identified by its source IPv4 address. &nbsp;</strong></p> <p><em><strong>Variables</strong></em></p> <p>The dataset contains the following variables:</p> <ul> <li><strong>Aggregations</strong>&nbsp;- created sums of the individual variables over a one-hour interval: <ul> </ul> <ul> <li><strong># of flows &nbsp;</strong>- number of flows for a given source IP&nbsp;</li> <li><strong># of packets </strong>&nbsp;-&nbsp;number of packets for a given source IP</li> <li><strong># of bytes </strong>&nbsp;-&nbsp;number of packets for a given source IP</li> <li><strong>flow duration </strong>&nbsp;- average flow duration in seconds</li> </ul> </li> <li><strong>Distinct Counts&nbsp;</strong>- count of distinct values for each variable over a one-hour window <ul> <li><strong># of peers </strong>&nbsp;- number of distinct communication peers for a given source IP</li> <li><strong># of ports </strong>&nbsp;- number of distinct destination ports&nbsp;for a given source IP</li> <li><strong># of protocols</strong>&nbsp;- number of distinct communication protocols&nbsp;for a given source IP</li> <li><strong># of AS numbers</strong>&nbsp;- number of distinct destination AS numbers for a given source IP</li> <li><strong># of countries </strong>&nbsp;- number of distinct destination countries&nbsp;for a given source&nbsp;</li> </ul> </li> </ul> <p><em><strong>Dataset Structure</strong></em></p> <ul> <li><strong>Dataset Files</strong> - each variable is contained in one <strong>Comma-Separated File (.csv)&nbsp;</strong>file <ul> <li><strong>Row index&nbsp;-&nbsp;</strong>&nbsp;timestamp of the observation window (8760 rows)</li> <li><strong>Columns index -&nbsp;</strong>&nbsp;anonymized IP addresses (65536&nbsp;columns)</li> </ul> </li> <li><strong>Label File -&nbsp;</strong>contains labels of the individual IP addresses from the Dataset Files <ul> <li><strong>Row index </strong>- anonymized IP addresses (65536 rows)</li> <li><strong>Columns index </strong>- labels for the IP addresses <ul> <li><strong>Subnet </strong>- ID&nbsp;of a subnet - hosts belonging to the same subnet have the same Id.</li> <li><strong>Subnet_range&nbsp;</strong>- CIDR range of a&nbsp;subnet</li> <li><strong>Unit -&nbsp;</strong>an ID of&nbsp;&nbsp;administrative unit owning the network range</li> <li><strong>Sub-unit </strong>&nbsp;- an ID of&nbsp;&nbsp;administrative sub-unit owning the network range</li> <li><strong>Subnet_label -&nbsp;&nbsp;</strong>subnet label <ul> <li><strong>Servers - </strong>selected subnets containing mostly servers (133.250.178.0/24, 133.250.163.0/24)</li> <li><strong>Workstations - </strong>selected subnets containing mostly workstations&nbsp;(133.250.146.0/24,&nbsp;133.250.157.128/25)</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p><strong><em>Further notes</em></strong></p> <ul> <li><strong>N/A values </strong> <ul> <li><strong>Variables&nbsp;</strong>- means that in a given observation window, the host did not communicate</li> <li><strong>Labels -&nbsp;</strong>no additional information on this IP is available</li> </ul> </li> <li><strong>Dataset load&nbsp;</strong> <ul> <li> <pre><code class="language-python">df = pd.read_csv(&lt;filename&gt;,header=[0], index_col=[0])</code></pre> </li> </ul> </li> </ul>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
8

Topics