Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3 results for “message board”

Learn how ShareScore rates datasets ↗
zenodo52/100

Cebulka (Polish dark web cryptomarket and image board) messages data

<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Cebulka (Polish dark web cryptomarket and image board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Patrycja Cheba (Jagiellonian University); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: &ldquo;Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade&rdquo; (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish dark web cryptomarket and image board called Cebulka (<a href="http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php">http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php</a>). &nbsp;&nbsp;</p> <p>5. <strong>Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project focuses on studying internet behavior concerning disruptive actions, particularly emphasizing the online narcotics market in Poland. The research seeks to (1) investigate how the open internet, including social media, is used in the drug trade; (2) outline the significance of darknet platforms in the distribution of drugs; and (3) explore the complex exchange of content related to the drug trade between the surface web and the darknet, along with understanding meanings constructed within the drug subculture.</p> <p>Within this context, Cebulka is identified as a critical digital venue in Poland&rsquo;s dark web illicit substances scene. Besides serving as a marketplace, it plays a crucial role in shaping the narratives and discussions prevalent in the drug subculture. The dataset has proved to be a valuable tool for performing the analyses needed to achieve the project&rsquo;s objectives.</p> <h3><strong>Data Content</strong></h3> <p>6. <strong>Data Description</strong></p> <p>The data was collected in three periods, i.e., in January 2023, June 2023, and January 2024.</p> <p>The dataset comprises a sample of messages posted on Cebulka from its inception until January 2024 (including all the messages with drug advertisements). These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories. The &ldquo;cebulka_adverts&rdquo; directory contains posts related to drug advertisements (both advertisements and comments). In contrast, the &ldquo;cebulka_community&rdquo; directory holds a sample of posts from other parts of the cryptomarket, i.e., those not related directly to trading drugs but rather focusing on discussing illicit substances. The dataset consists of 16,842 posts.</p> <p>7. <strong>Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>8. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (&ldquo;cebulka_adverts.zip&rdquo; and &ldquo;cebulka_community.zip&rdquo;) containing all messages. These files are organized into individual directories that mirror the folder structure found on Cebulka.</li> <li>Two .csv files that list all the messages, including file names and the content of each post. The first .csv lists messages from &ldquo;cebulka_adverts.zip,&rdquo; and the second .csv lists messages from &ldquo;cebulka_community.zip.&rdquo;</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>9.&nbsp;<strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Hyperreal Talk (Polish clear web message board) messages data

<h3><strong>General Information</strong></h3> <p>1.<strong> Title of Dataset</strong></p> <p>Hyperreal Talk (Polish clear web message board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: &ldquo;Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade&rdquo; (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish clear web message board called Hyperreal Talk (<a href="https://hyperreal.info/talk/">https://hyperreal.info/talk/</a>).</p> <p>5.<strong> Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project delves into internet dynamics within disruptive activities, specifically focusing on the online drug trade in Poland. It aims to (1) examine the utilization of the open internet, including social media, in the drug trade; (2) delineate the role of darknet environments in narcotics distribution; and (3) uncover the intricate flow of drug trade-related content and its meanings between the open web and the darknet, and how these meanings are shaped within the so-called drug subculture.</p> <p>The Hyperreal Talk forum emerges as a pivotal online space on the Polish internet, serving as a hub for discussions and the exchange of knowledge and experiences concerning drug use. It plays a crucial role in investigating the narratives and discourses that shape the drug subculture and the broader societal perceptions of drug consumption. The dataset has been instrumental in conducting analyses pertinent to the earlier project goals.</p> <p>6. <strong>Collection Method</strong></p> <p>The dataset was compiled using the Scrapy framework, a web crawling and scraping library for Python. This tool facilitated systematic content extraction from the targeted message board.</p> <p>7.<strong> Collection Date</strong></p> <p>The data was collected in two periods, i.e., in September 2023 and November 2023.</p> <h3><strong>Data Content</strong></h3> <p>8.&nbsp;<strong>Data Description</strong></p> <p>The dataset comprises all messages posted on the Polish-language Hyperreal Talk message board from its inception until November 2023. These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories: &ldquo;hyperreal&rdquo; and &ldquo;hyperreal_hidden.&rdquo; The &ldquo;hyperreal&rdquo; directory contains accessible posts without needing to log in to Hyperreal Talk, while the &ldquo;hyperreal_hidden&rdquo; directory holds posts that can only be viewed by logged-in users. For each directory, a .txt file has been prepared detailing the structure of the message board folders from which the posts were extracted. The dataset includes 6,248,842 posts.</p> <p>9.<strong> Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>10. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (hyperreal.zip) containing messages that are visible without logging into Hyperreal Talk. These files are organized into individual directories that mirror the folder structure found on the Hyperreal Talk message board.</li> <li>Zipped .txt files (hyperreal_hidden.zip) containing messages that are visible only after logging into Hyperreal Talk. Similar to the first type, these files are organized into directories corresponding to the website&rsquo;s folder structure.</li> <li>A .csv file that lists all the messages, including file names and the content of each post.</li> </ul> <h3><strong>Accessibility and Usage</strong></h3> <p>11.<strong> Access Conditions</strong></p> <p>The data can be accessed without any restrictions.</p> <p>12. <strong>Related Documentation</strong></p> <p>Attached are .txt files detailing the tree of folders for &ldquo;hyperreal.zip&rdquo; and &ldquo;hyperreal_hidden.zip.&rdquo;</p> <p>Documentation on the Python regular expressions used for scraping, cleaning, processing, and anonymizing the data can be found on GitHub at the following URLs:</p> <ul> <li><a href="https://github.com/LeszekSwieca/Project_2021-43-B-HS6-00710">https://github.com/LeszekSwieca/Project_2021-43-B-HS6-00710</a></li> <li><a href="https://github.com/HaitaoShi/Scrapy_hyperreal">https://github.com/HaitaoShi/Scrapy_hyperreal</a>"</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>13. <strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that scraping and automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Dopek.eu (Polish clear web and dark web message board) messages data

<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Dopek.eu (Polish clear web and dark web message board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: &ldquo;Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade&rdquo; (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4. <strong>Data Source</strong></p> <p>Clear web and dark web message board called dopek.eu (<a href="https://dopek.eu/">https://dopek.eu/</a>). &nbsp;&nbsp;</p> <p>5.<strong> Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project delves into internet dynamics within disruptive activities, specifically focusing on the online drug trade in Poland. It aims to (1) examine the utilization of the open internet, including social media, in the drug trade; (2) delineate the role of darknet environments in narcotics distribution; and (3) uncover the intricate flow of drug trade-related content and its meanings between the open web and the darknet, and how these meanings are shaped within the so-called drug subculture.</p> <p>The dopek.eu forum emerges as a pivotal online space on the Polish internet, serving as a hub for trading, discussions, and the exchange of knowledge and experiences concerning the use of the so-called new psychoactive substances (designer drugs). The dataset has been instrumental in conducting analyses pertinent to the earlier project goals.</p> <p>6.<strong> Collection Method</strong></p> <p>The dataset was compiled using the Scrapy framework, a web crawling and scraping library for Python. This tool facilitated systematic content extraction from the targeted message board.</p> <p>7.<strong> Collection Date</strong></p> <p>The data was collected in October 2023.</p> <h3><strong>Data Content</strong></h3> <p>8.<strong> Data Description</strong></p> <p>The dataset comprises all messages posted on dopek.eu from its inception until October 2023. These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. A .txt file has been prepared detailing the structure of the message board folders from which the posts were extracted. The dataset includes 171,121 posts.</p> <p>9.<strong> Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>10. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following types of files:</p> <ul> <li>Zipped .txt files (dopek.zip) containing all messages (posts).</li> <li>A .csv file that lists all the messages, including file names and the content of each post.</li> </ul> <h3><strong>Accessibility and Usage</strong></h3> <p><strong>11. Access Conditions</strong></p> <p>The data can be accessed without any restrictions.</p> <p><strong>12. Related Documentation</strong></p> <p>Attached are .txt files detailing the tree of folders for &ldquo;dopek.zip&rdquo;.</p> <h3><strong>Ethical Considerations</strong></h3> <p><strong>13. Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the posts, utilizing automated systems for irreversible hashing. Recognizing that scraping and automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record