Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
37
datasets available to search
ShareScore release 0.9.0
Dataset results
37 results for “Network Traffic”
Speed Prediction in Large and Dynamic Traffic Sensor Networks
<p>Aggregated traffic sensor data from Fortaleza (Brazil) in 2014.</p> <p>Please cite the following paper when using the dataset:</p> <p>R.P. Magalhaes, F. Lettich, J.A. Macedo, F.M. Nardini, R. Perego, C. Renso, R. Trani., <strong>Speed prediction in large and dynamic traffic sensor networks</strong>, Information Systems (2019) 101444, <a href="https://doi.org/10.1016/j.is.2019.101444">https://doi.org/10.1016/j.is.2019.101444</a></p> <p>You can also check details regarding the dataset in the paper.</p>
COVID-19 and Traffic Safety: Exploring Exposure, Crash Frequency and Severity, and Roadway and Network Design
<p>Early COVID-19 lockdowns in the first half of 2020 largely kept people at home, thereby reducing motor vehicle traffic levels. Theoretically, reduced traffic exposure should have resulted in reduced motor vehicle crashes. However, a variety of factors may have complicated this relationship. In order to better understand the impact of COVID-19 lockdowns on traffic safety outcomes, we explore fatalities, injuries, and total crashes before and during the lockdowns on both the national and state levels. We provide descriptive statistics and create negative binomial regressions exploring the role of vehicle, user, and built environment factors on traffic safety outcomes. Findings suggest that crash counts in Region 6 were 35%-50% lower in 2020 during the COVID-19 lockdowns. Crashes that occurred during the COVID-19 lockdowns were more likely to be more severe. Fatal pedestrian crashes across the U.S. decreased during COVID-19 (although not as much as overall fatal crashes) and fatal bicyclist crash counts increased. Drunk drivers were less prevalent in nationwide fatalities but more prevalent in overall Region 6 crashes. Overall, crashes were more likely single-vehicle fixed-object or rollover crashes involving unsafe speeds. In Texas, suburban areas saw the most crashes before and during COVID-19, although they also saw the greatest decrease. Rural Texas crashes were most likely to result in a fatality or serious injury, and that likelihood got worse during the COVID-19 lockdowns. While Texas freeways and arterials saw the largest decreases in crash counts, these functional classifications still had the most crashes. Urban interstates and rural local roads in Texas were notable because these two functional classifications actually saw increases in the number of fatal and serious injury crashes during COVID-19 lockdowns.</p>
Understanding the impact of host networking elements on traffic bursts: Raw measurement data
<p>This record contains the raw trace files gathered by the Valinor network traffic burst measurement framework in Redis dump (rdb) format. Please refer to the artifact repository for instructions on how to parse and use the datasets:</p> <p><a href="https://github.com/hopnets/valinor-rawdata">hopnets/valinor-rawdata: Raw Redis datasets containing the measurement results of Valinor NSDI '23 paper (github.com)</a></p>
Data from: Climate change and pathways of vessel traffic create marine protected area networks of invasion
Open the record for dataset details and reuse information.
IoT network traffic dataset using the custom flow representation
Open the record for dataset details and reuse information.
Crowdsourced air traffic data from The OpenSky Network 2020 [CC-BY]
<p><strong>Motivation</strong></p> <p>The data in this dataset is derived and cleaned from the full OpenSky dataset to illustrate the development of air traffic during the COVID-19 pandemic. It spans all flights seen by the network's more than 2500 members since 1 January 2020. More data will be periodically included in the dataset until the end of the COVID-19 pandemic.</p> <p><strong>License</strong></p> <p>Creative Commons CC-BY</p> <p>The only difference with the <a href="https://zenodo.org/record/3928550">original dataset</a> comes from anonymised aircraft information.</p> <p><strong>Disclaimer</strong></p> <p>The data provided in the files is provided as is. Despite our best efforts at filtering out potential issues, some information could be erroneous.</p> <ul> <li>Origin and destination airports are computed online based on the ADS-B trajectories on approach/takeoff: no crosschecking with external sources of data has been conducted.<br> Fields <strong>origin</strong> or <strong>destination</strong> are empty when no airport could be found.</li> <li>Aircraft information come from the OpenSky aircraft database. Fields <strong>typecode</strong> and <strong>registration</strong> are empty when the aircraft is not present in the database.</li> </ul> <p><strong>Description of the dataset</strong></p> <p>One file per month is provided as a csv file with the following features:</p> <ul> <li><strong>callsign</strong>: the identifier of the flight displayed on ATC screens (usually the first three letters are reserved for an airline: AFR for Air France, DLH for Lufthansa, etc.)</li> <li><strong>number</strong>: the commercial number of the flight, when available (the matching with the callsign comes from public open API)</li> <li><strong>aircraft_uid</strong>: a unique anonymised identifier for aircraft;</li> <li><strong>typecode</strong>: the aircraft model type (when available);</li> <li><strong>origin</strong>: a four letter code for the origin airport of the flight (when available);</li> <li><strong>destination</strong>: a four letter code for the destination airport of the flight (when available);</li> <li><strong>firstseen</strong>: the UTC timestamp of the first message received by the OpenSky Network;</li> <li><strong>lastseen</strong>: the UTC timestamp of the last message received by the OpenSky Network;</li> <li><strong>day</strong>: the UTC day of the last message received by the OpenSky Network;</li> <li><strong>latitude_1</strong>, <strong>longitude_1</strong>, <strong>altitude_1</strong>: the first detected position of the aircraft;</li> <li><strong>latitude_2</strong>, <strong>longitude_2</strong>, <strong>altitude_2</strong>: the last detected position of the aircraft.</li> </ul> <p><strong>Examples</strong></p> <p>Possible visualisations and a more detailed description of the data are available at the following page:<br> <<a href="https://traffic-viz.github.io/scenarios/covid19.html">https://traffic-viz.github.io/scenarios/covid19.html</a>></p> <p><strong>Credit</strong></p> <p>If you use this dataset, please cite the original OpenSky paper:</p> <p>Matthias Schäfer, Martin Strohmeier, Vincent Lenders, Ivan Martinovic and Matthias Wilhelm.<br> "Bringing Up OpenSky: A Large-scale ADS-B Sensor Network for Research".<br> In<em> Proceedings of the 13th IEEE/ACM International Symposium on Information Processing in Sensor Networks (IPSN)</em>, pages 83-94, April 2014.</p> <p>and the traffic library used to derive the data:</p> <p>Xavier Olive.<br> "traffic, a toolbox for processing and analysing air traffic data."<br> <em>Journal of Open Source Software</em> 4(39), July 2019.</p>
Application specific network traffic with specified activities
<p>This dataset contains network traffic of 12 different containerized, isolated applications. It includes headers of network traffic data in the initial two minutes of execution of the application.<br> We have specified each network traffic with an activity that the application was involved in during traffic capturing.<br> </p> <p>------------------------------------<br> Application name, No. of samples, Engaged activity while capturing<br> ------------------------------------<br> Curl, 60, downloading<br> Firefox, 47, browsing<br> mpg123, 99, audio streaming<br> mplayer, 82, audio streaming<br> mpv, 48, video streaming<br> Slack, 58, no chatting (just opening the app and menus)<br> streamlink, 57, video streaming<br> Trello, 46, just opening the app (just opening the app and menus)<br> Vlc, 102, audio streaming<br> w3m, 60, downloading<br> wget, 20, file downloading<br> wgetaudio, 20, file downloading(audio)<br> wgetvideo, 20, file downloading(video)<br> ydm (mplayer + youtube-dl),61, video streaming</p>
The dataset of "Characteristics of Urban Interaction Network in Shandong Province from the Perspective of Traffic Flow and Information Flow"
<p>There are the datasets of the research named "Characteristics of Urban Interaction Network in Shandong Province from the Perspective of Traffic Flow and Information Flow".</p>
Simulated performance data: MILC, LAMMPS, and uniform random traffic patterns on 72-ndoe dragonfly network
<p>Data generated from an old, private fork of the CODES simulation toolkit: https://github.com/codes-org/codes</p> <p>Data dictionary: https://github.com/kevinabrown/codes/wiki/Dragonfly-Dally-DEBUG-Metrics</p>
CESNET-USTS23: a benchmark dataset of Unevenly spaced time series from network traffic
<p>This dataset was created to evaluate characteristics of <em>Unevenly sampled time series from network traffic (USTS)</em> for the paper <em>Unevenly Spaced Time Series from Network Traffic</em>.</p> <p>The file named <code>time_series.tar.gz</code> contains a folder with time series CSV files as raw data of the experiment. In the folder are the following files:</p> <ul> <li><code>fts.csv</code> -- contains 2.6 million <em>Flow time series (FTS)</em> created from 259 million IP flows,</li> <li><code>pts.csv</code> -- contains 19 million <em>Packet time series (PTS)</em> created from 110 million network packets,</li> <li><code>sfts.csv</code> -- contains 15 million <em>Single flow time series (SFTS)</em> created from 160 million network packets.</li> </ul> <p>Traffic was captured on the national CESNET2 network from February 2023 to April 2023. All IP addresses in the dataset were anonymized.</p> <p>The <code>fts.csv</code> has the following format:</p> <ul> <li>ID_DEPENDENCY -- Identification of a network dependency observed as a Flow time series. (real IP address was anonimized by replacing with a random IP address)</li> <li>N_FLOWS -- Number of flows in time series, i.e., number of data points.</li> <li>N_PACKETS -- Number of packets in time series, i.e., the sum of metric PACKETS.</li> <li>N_BYTES -- Number of bytes in time series, i.e., the sum of metric PACKETS.</li> <li>PACKETS -- The array containing the time series metric number of packets in the IP flow.</li> <li>BYTES -- The array containing the time series metric number of bytes in the IP flow.</li> <li>START_TIMES -- The array containing the time series time axis of the flows starts.</li> <li>END_TIMES -- The array containing the time series time axis of the flows ends.</li> </ul> <p>The <code>pts.csv</code> has the following format:</p> <ul> <li>ID_DEPENDENCY -- Identification of a network dependency observed as a Packet time series. (real IP address was anonymized by replacing with a random IP address)</li> <li>BYTES -- The array containing the time series metric payload length of the network packet.</li> <li>TIMES -- The array containing the time series time axis of the transmission of network packets.</li> </ul> <p>The <code>sfts.csv</code> has the following format:</p> <ul> <li>SRC_IP -- Source IP address. (real IP address was anonimized by replacing with a random IP address)</li> <li>SRC_PORT -- Source port.</li> <li>DST_IP -- Destination IP address (real IP address was anonymized by replacing with a random IP address)</li> <li>DST_PORT -- Destination port.</li> <li>bytes -- The array containing the time series metric payload length of the network packet.</li> <li>time -- The array containing the time series time axis of the transmission of network packets.</li> </ul> <p>The file named <code>characteristics.tar.gz</code> contains a folder with characteristics gained by experiments from time series files. In the folder are the following files:</p> <ul> <li><code>fts.characteristics.csv</code> -- Characteristics about Flow time series from the fts.csv.</li> <li><code>pts.characteristics.csv</code> -- Characteristics about Packet time series from the pts.csv.</li> <li><code>sfts.characteristics.csv</code> -- Characteristics about Single flow time series from the sfts.csv.</li> </ul> <p>The <code>fts.characteristics.csv</code> has the following format:</p> <ul> <li>LENGTH -- Number of data points in the source time series.</li> <li>DURATION -- Duration of the source time series.</li> <li>H_BYTES -- Hurst exponent of the source time series metric BYTES.</li> <li>STATIONARITY_PACKETS -- Stationarity of the source time series metric PACKETS.</li> <li>STATIONARITY_BYTES -- Stationarity of the source time series metric BYTES.</li> <li>OVERALL_STATIONARITY -- Overal stationarity created by merging STATIONARITY_PACKETS and STATIONARITY_BYTES.</li> </ul> <p>The <code>pts.characteristics.csv</code> and <code>sfts.characteristics.csv</code> have the following format:</p> <ul> <li>LENGTH -- Number of data points in the source time series.</li> <li>DURATION -- Duration of the source time series.</li> <li>H -- Hurst exponent of the source time series.</li> <li>STATIONARITY -- Stationarity of the source time series.</li> </ul> <p>We provide the samples of all zipped files for a quick lookup: <code>fts.characteristics.sample.csv</code>, <code>fts.sample.csv</code>, <code>pts.characteristics.sample.csv</code>, <code>pts.sample.csv</code>, <code>sfts.characteristics.sample.csv</code>, <code>sfts.sample.csv</code></p> <p> </p>
Data from: A model to identify urban traffic congestion hotspots in complex networks
The rapid growth of population in urban areas is jeopardizing the mobility and air quality worldwide. One of the most notable problems arising is that of traffic congestion. With the advent of technologies able to sense real-time data about cities, and its public distribution for analysis, we are in place to forecast scenarios valuable for improvement and control. Here, we propose an idealized model, based on the critical phenomena arising in complex networks, that allows to analytically predict congestion hotspots in urban environments. Results on real cities' road networks, considering, in some experiments, real- traffic data, show that the proposed model is capable of identifying susceptible junctions that might becomes hotspots if mobility demand increases.
Smartphone-Based Incentive Framework for Dynamic Network-Level Traffic Congestion Management Project H3
<p>Task 2 numerical experiment results</p>
SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events
<p><strong>SUTD-TrafficQA</strong> is a dataset that takes the form of <strong>Video</strong> <strong>QA</strong> based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive capability of causal inference and event understanding models in complex traffic scenarios. Specifically, the dataset proposes 6 challenging reasoning tasks corresponding to various traffic scenarios, so as to evaluate the reasoning capability over different kinds of complex yet practical traffic events.</p>
Data from: A model to identify urban traffic congestion hotspots in complex networks
Open the record for dataset details and reuse information.
IoT-deNAT: Outbound flow-based network traffic data of IoT and non-IoT devices behind a home NAT
<p>This dataset is comprised of NetFlow records, which capture the outbound network traffic of 8 commercial IoT devices and 5 non-IoT devices, collected during a period of 37 days in a lab at Ben-Gurion University of The Negev. The dataset was collected in order to develop a method for telecommunication providers to detect vulnerable IoT models behind home NATs. Each NetFlow record is labeled with the device model which produced it; for research reproducibilty, each NetFlow is also allocated to either the "training" or "test" set, in accordance with the partitioning described in:</p> <p>Y. Meidan, V. Sachidananda, H. Peng, R. Sagron, Y. Elovici, and A. Shabtai, A novel approach for detecting vulnerable IoT devices connected behind a home NAT, Computers & Security, Volume 97, 2020, 101968, ISSN 0167-4048, https://doi.org/10.1016/j.cose.2020.101968. (http://www.sciencedirect.com/science/article/pii/S0167404820302418)</p> <p> </p> <p>Please note:</p> <ul> <li>The dataset itself is free to use, however users are requested to cite the above-mentioned paper, which describes in detail the research objectives as well as the data collection, preparation and analysis.</li> <li>Following is a brief description of the features used in this dataset.</li> </ul> <p> </p> <p># NetFlow features, used in the related paper for analysis</p> <p>'FIRST_SWITCHED': System uptime at which the first packet of this flow was switched<br> 'IN_BYTES': Incoming counter for the number of bytes associated with an IP Flow<br> 'IN_PKTS': Incoming counter for the number of packets associated with an IP Flow<br> 'IPV4_DST_ADDR': IPv4 destination address<br> 'L4_DST_PORT': TCP/UDP destination port number<br> 'L4_SRC_PORT': TCP/UDP source port number<br> 'LAST_SWITCHED': System uptime at which the last packet of this flow was switched<br> 'PROTOCOL': IP protocol byte (6: TCP, 17: UDP)<br> 'SRC_TOS': Type of Service byte setting when there is an incoming interface<br> 'TCP_FLAGS': Cumulative of all the TCP flags seen for this flow</p> <p> </p> <p># Features added by the authors</p> <p>'IP': Prefix of the destination IP address, representing the network (without the host)<br> 'DURATION': Time (seconds) between first/last packet switching</p> <p> </p> <p># Label<br> 'device_model': <type>.<manufacturer>.<model number></p> <p> </p> <p># Partition<br> 'partition': Training or test</p> <p> </p> <p># Additional NetFlow features (mostly zero-variance)<br> 'SRC_AS': Source BGP autonomous system number<br> 'DST_AS': Destination BGP autonomous system number<br> 'INPUT_SNMP': Input interface index<br> 'OUTPUT_SNMP': Output interface index<br> 'IPV4_SRC_ADDR': IPv4 source address<br> 'MAC': MAC address of the source</p> <p> </p> <p># Additional data<br> 'category': IoT or non-IoT<br> 'type': IoT, access_point, smartphone, laptop<br> 'date': Datepart of FIRST_SWITCHED<br> 'inter_arrival_time': Time (seconds) between successive flows of the same device (identified by its MAC address)</p>
Project dataset: predicting traffic volumes and congestion in road networks equipped with ANPR cameras
<p><a href="https://github.com/ppintosilva/anpr-predict">https://github.com/ppintosilva/anpr-predict</a></p>
OLSR traffic from Funkfeuer Graz network
<p>pcap files containing all OLSR traffic captured from May 20, 2020 at 12:49:15 to May 22, 2020 at 22:18:27 in a virtual machine conected via VPN to the network.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.