Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “honeypots”
PANDAcap SSH Honeypot Dataset
<p>This is a dataset of <strong>63 <a href="https://github.com/panda-re/panda">PANDA</a> traces</strong>, collected using the <a href="https://github.com/vusec/pandacap">PANDAcap</a> framework. The dataset aims to offer a starting point for the analysis of <em>ssh brute force attacks</em>. The traces were collected through the course of approximately 3 days from 21 to 23 February 2020. A VM was configured using PANDAcap so that it accepts all passwords for user <code>root</code>. When an ssh session starts for the user, PANDA is signaled by the <a href="https://github.com/panda-re/panda/tree/master/panda/plugins/recctrl">recctrl plugin</a> to start recording for 30'.</p> <p>You can read more details about the experimental setup and an overview of the dataset <strong>EuroSec 2020</strong> publication:</p> <ul> <li> <p>Manolis Stamatogiannakis, Herbert Bos, and Paul Groth. PANDAcap: A Framework for Streamlining Collection of Full-System Traces. In <em>Proceedings of the 13th European Workshop on Systems Security</em>, <a href="https://www.concordia-h2020.eu/eurosec-2020/">EuroSec '20</a>, Heraklion, Greece, April 2020. doi: <a href="https://doi.org/10.1145/3380786.3391396">10.1145/3380786.3391396</a>, preprint: <a href="https://www.vusec.net/publications/#stamatogiannakis-bos-groth-pandacapaframeworkforstreamliningcollectionoffullsystemtraces-2020">vusec.net</a></p> </li> </ul> <p>The dataset is split in 3 zip files/directories:</p> <ul> <li><strong>rr</strong>: Contains the 63 PANDA traces of the dataset. The traces are in the upcoming RRArchive format. Note that PANDA support for the format is still wip at the time of writing (April 2020). If you need to downgrade to the traditional PANDA trace format, you can use the snippet in <a href="https://github.com/vusec/pandacap/blob/master/docs/xxx">foo</a>.</li> <li><strong>qcow</strong>: Contains the QCOW base image (<code>ubuntu16-planb.qcow2</code>) used to create the dataset, as well as the disk deltas for the 63 traces. These can be mounted to inspect the contents of the filesystem before and after each session. and disk deltas for the 63 traces. Quick instructions on how to mount and inspect a QCOW image can be found below.</li> <li><strong>pcap</strong>: Contains the pcap network traces for the sessions in the PANDA traces. These have been extracted using the PANDA <a href="https://github.com/panda-re/panda/tree/master/panda/plugins/network">network plugin</a>. We decided to also include them in the dataset as standalone files for convenience.</li> </ul> <p>Additionally, we provide the PANDA linux kernel profile <code>ubuntu16-planb-kernelinfo.conf</code>, which can be used to analyze the traces using the PANDA <a href="https://github.com/panda-re/panda/tree/master/panda/plugins/osi_linux">osi_linux plugin</a>.</p> <p>Additional information:</p> <ul> <li>To convert RRArchive traces to the traditional PANDA format, run the following snippet inside the <code>rr</code> directory: <pre><code class="language-bash">for f in *.tar.gz; do tar -zxvf "$f" --exclude=PANDArr --xform='s%/%-%' --xform='s%-metadata%%' rm -f "$f" done</code></pre> </li> <li>If you wish to reuse the VM image in your project, it is available as a standalone download through <a href="https://academictorrents.com/details/39df3904460e909e175434cbd87764b8c487891d">academictorrents.com</a>, along with more detailed information on its contents.</li> <li>If you wish to download individual samples rather than the whole dataset, you can use the dataset torrent file available through <a href="https://academictorrents.com/details/4a3eadf47425cb60111ec224de272997294eec93">academictorrents.com</a>. Unlike this Zenodo deposit, the files in the torrent have not been zipped.</li> <li>A better formatted (and possibly more up-to-date) version of this information can be found <a href="https://github.com/vusec/pandacap/blob/master/docs/eurosec20-dataset.md">here</a>.</li> </ul>
CTU Hornet 65 Niner: A Network Dataset of Geographically Distributed Low-Interaction Honeypots
<p>CTU Hornet 65 Niner is a dataset of 65 days of network traffic attacks captured in cloud servers used as honeypots to help understand how geography may impact the inflow of network attacks. The honeypots were placed in nine different geographical locations: Amsterdam, London, Frankfurt, San Francisco, New York, Singapore, Toronto, Bangalore, and Sydney. The data was captured from April 28th to July 1st, 2024.</p> <p>The nine cloud servers were created and configured following identical instructions using Ansible [1] in DigitalOcean [2] cloud provider. The network capture was performed using the Zeek [3] network monitoring tool, which was installed on each cloud server. The cloud servers had only one service running (SSH on a non-standard port) and were fully dedicated to being used as a honeypot. No honeypot software was used in this dataset.</p> <p>The dataset is composed of nine scenarios:</p> <ul> <li>Honeypot-Cloud-DigitalOcean-Geo-1: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-2: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-3: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-4: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-5: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-6: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-7: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-8: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> <li>Honeypot-Cloud-DigitalOcean-Geo-9: has 65 folders (YYYY-MM-DD), each containing 24 Zeek conn.log files and other Zeek files</li> </ul> <p><strong>References:</strong></p> <p>[1] Ansible IT Automation Engine, https://www.ansible.com/. Accessed on 08/28/2024.</p> <p>[2] DigitalOcean, https://www.digitalocean.com/. Accessed on 08/28/2024.</p> <p>[3] Zeek Documentation, https://docs.zeek.org/en/master/index.html. Accessed on 08/28/2024.</p> <p><strong>Funding:</strong></p> <p>The authors acknowledge support by the Strategic Support for the Development of Security Research in the Czech Republic 2019--2025 (IMPAKT 1) program, by the Ministry of the Interior of the Czech Republic under No. VJ02010020 -- AI-Dojo: Multi-agent testbed for the research and testing of AI-driven cyber security technologies.</p>
Hornet 40: Network Dataset of Geographically Placed Honeypots
<p>Hornet 40 is a dataset of 40 days of network traffic attacks captured in cloud servers used as honeypots to help understand how geography may impact the inflow of network attacks. The honeypots are located in eight different cities: Amsterdam, London, Frankfurt, San Francisco, New York, Singapore, Toronto, Bangalore. The data was captured in April, May, and June 2021.</p> <p>The eight cloud servers were created and configured simultaneously following identical instructions. The network capture was performed using the Argus network monitoring tool in each cloud server. The cloud servers had only one service running (SSH on a non-standard port) and were fully dedicated as a honeypot. No honeypot software was used in this dataset.</p> <p>The dataset consists of eight scenarios, one for each geographically located cloud server. Each scenario contains bidirectional NetFlow files in the following format:</p> <ul> <li>hornet40-biargus.tar.gz: all scenarios with bidirectional NetFlow files in Argus binary format;</li> <li>hornet40-netflow-v5.tar.gz: all scenarios with bidirectional NetFlow v5 files in CSV format;</li> <li>hornet40-netflow-extended.tar.gz: all scenarios with bidirectional NetFlows files in CSV format containing all features provided by Argus.</li> <li>hornet40-full.tar.gz: download all the data (biargus, NetFlow v5, and extended NetFlows)</li> </ul>
Dataset of Thesis - Attacks on the Cloud: Unveiling Cyber Assaults on Cloud Infrastructure Through Honeypot Analysis
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.