Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

354

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

354 results for “data accessibility”

Learn how ShareScore rates datasets ↗
zenodo44/100

Data from: The interaction of ice and law in Arctic marine accessibility

<p>Sea ice levies an impost on maritime navigability in the Arctic. But ice cover diminution due to anthropogenic climate change is generating expectations for improved accessibility in coming decades. Projections of sea ice cover retreating preferentially from the eastern Arctic suggest key provisions of international law of the sea will require revision. Specifically, protections against marine pollution in ice covered seas enshrined in Article 234 of the United Nations Convention on the Law of the Sea have been used in recent decades to extend jurisdictional competence over the Northern Sea Route only loosely associated with environmental outcomes. Projections show that plausible open water routes through international waters may be accessible by mid-century under all but the most aggressive of emissions control scenarios. While inter- and intra-annual variability places the economic viability of these routes in question for some time, the inevitability of a seasonally ice-free Arctic will be attended by a reduction of regulatory friction and a recalibration of associated legal frameworks.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Modular control of human movement during running: an open access data set

<p>The human body is an outstandingly complex machine including around 1000 muscles and joints acting synergistically. Yet, the coordination of the enormous amount of degrees of freedom needed for movement is mastered by our one brain and spinal cord. The idea that some synergistic neural components of movement exist was already suggested at the beginning of the XX century. Since then, it has been widely accepted that the central nervous system might simplify the production of movement by avoiding the control of each muscle individually. Instead, it might be controlling muscles in common patterns that have been called muscle synergies. Only with the advent of modern computational methods and hardware it has been possible to numerically extract synergies from electromyography (EMG) signals. However, typical experimental setups do not include a big number of individuals, with common sample sizes of five to 20 participants. With this study, we make publicly available a set of EMG activities recorded during treadmill running from the right lower limb of 135 healthy and young adults (78 males, 57 females). Moreover, we include in this open access data set the code used to extract synergies from EMG data using non-negative matrix factorization and the relative outcomes. Muscle synergies, containing the time-invariant muscle weightings (motor modules) and the time-dependent activation coefficients (motor primitives), were extracted from 13 ipsilateral EMG activities using non-negative matrix factorization. Four synergies were enough to describe as many gait cycle phases during running: weight acceptance, propulsion, early swing and late swing. We foresee many possible applications of our data, that we can summarize in three key points. First, it can be a prime source for broadening the representation of human motor control due to the big sample size. Second, it could serve as a benchmark for scientists from multiple disciplines such as musculoskeletal modelling, robotics, clinical neuroscience, sport science, etc. Third, the data set could be used both to train students or to support established scientists in the perfection of current muscle synergies extraction methods.</p> <p>The &quot;RAW_DATA.RData&quot;&nbsp;R list consists of elements of S3 class &quot;EMG&quot;, each of which is a human locomotion trial containing cycle segmentation timings and raw electromyographic (EMG) data from 13 muscles of the right-side leg. Cycle times are structured as data frames containing two columns that&nbsp;correspond to touchdown (first column) and lift-off (second column).&nbsp;Raw EMG data sets are also structured as data frames with one row for each recorded data point&nbsp;and 14 columns. The first column contains the incremental time in seconds. The remaining 13 columns contain the raw EMG data, named with the following muscle abbreviations:&nbsp;ME = gluteus medius, MA = gluteus maximus, FL = tensor fasci&aelig; lat&aelig;, RF = rectus femoris, VM = vastus medialis, VL = vastus lateralis, ST = semitendinosus, BF = biceps femoris, TA = tibialis anterior, PL = peroneus longus, GM = gastrocnemius medialis, GL = gastrocnemius lateralis, SO = soleus.</p> <p>The file &quot;dataset.rar&quot; contains data in older format,&nbsp;not compatible with the R package <a href="https://CRAN.R-project.org/package=musclesyneRgies">musclesyneRgies</a>.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

bulk-tumour-api: a programmatically accessible dataset of pre-processed bulk tumour sequencing data

<p><strong>This repository, including the API,&nbsp;are&nbsp;currently under development.</strong></p> <p><strong>bulk-tumour-api</strong>: A programmatically accessible dataset of pre-processed bulk tumour sequencing data. The python API can be found at&nbsp;https://github.com/tomouellette/bulk-tumour-api. All data stored in this repository have&nbsp;been collected from&nbsp;open access&nbsp;online sources. Original references and sources are provided in database.tsv (for empirical patient data) and synthetic.tsv (for simulated data).</p> <p><strong>A note on datasets: </strong></p> <ul> <li>All <em>empirical patient sequencing </em>samples&nbsp;have&nbsp;been processed into&nbsp;pseudo-VCF files&nbsp;which at minimum contain the following columns:&nbsp; sample identifier (sample), patient identifier (patient), chromosome (chr), position (pos), variant allele frequency (VAF), alternate read counts (t_alt_count), depth (DP), and total copy number (total_cn). However, if more data is required, unprocessed data including copy number segments or gene-level calls, clinical, and/or biopsy level information can be found in the /raw/.&nbsp;</li> <li>All <em>synthetic datasets </em>have also been processed in pseudo-VCF files. In some cases, all ground truth information (e.g. subclone frequency) is contained within the pseudo-VCF. In other cases, additional meta/ground-truth information are in separate files; any simulated sample with a column marked has_meta&nbsp;= True will have multiple files that will be downloaded together.</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Data and codes: Changing the Academic Gender Narrative through Open Access

<p>This Zenodo entry includes data and R codes used to generate the figures included in the manuscript &quot;Changing the Academic Gender Narrative through Open Access&quot;, authored by members of the Curtin Open Knowledge Initiative (COKI). These include data that are either publicly available or&nbsp;derived through the COKI data infrastructure.</p> <p>The R file includes codes used to generate Figures 1, 2, 3, 4, 1A, 2A and 3A. It uses data contained in the files &quot;au_data_all.csv&quot;, &quot;au_groupings.csv&quot;, &quot;uk_data_all.csv&quot; and &quot;uk_groupings.csv&quot;.</p> <p>This entry also includes the full data files (.csv and .xlsx) for Figures 5 and 6 included in the manuscript:</p> <ul> <li>Figure 5: &lsquo;Percentages of women academic staff (headcount) compared to the total number of academics in the institution for 43 Australian universities by grouping, 2020&rsquo;. The analysis is of publicly available data sourced from the Australian Department of Education, Skills and Employment.</li> <li>Figure 6: &lsquo;Percentages of women academic staff (headcount) compared to the total number of academics in the institution for a subset of 165 United Kingdom higher education institutions by grouping, 2020&rsquo;. The analysis is of publicly available data sourced from the United Kingdom Higher Education Statistics Agency (HESA).</li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Chromatin accessibility data for the CRISPRai prediction algorithm implemented in crisprScore

<p>Chromatin accessibility data for the CRISPRai prediction algorithm implemented in&nbsp;crisprScore; see&nbsp;https://github.com/crisprVerse/crisprScore for more detail.</p> <p>&nbsp;</p>

openmit-licenseJun 2022View details →
zenodo44/100

Data from "Fast acquisition of propagating waves in humans with low-field MRI: towards accessible MR elastography"

<p>Data presented in the Science Advances manuscript &quot;<em>Fast acquisition of propagating waves in humans with low-field MRI: towards accessible MR elastography</em>&quot; by Yushchenko M., Sarracanie M., Salameh N.</p> <p>See further details in <em>Description.txt.</em></p> <p>The 3D wave datasets acquired in humans at 0.1 T can be used for elastogram reconstruction with appropriate methods.<br> <br> &nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Monitoring open access publishing of NWO funded research (2015-2021) data set

<p>This is the dataset underlying the report &quot;Monitoring open access publishing of NWO funded research&quot;&nbsp; (<a href="https://doi.org/10.5281/zenodo.7041897">https://doi.org/10.5281/zenodo.7041897</a>)</p> <p>The report presents statistics on the extent to which publications from the period 2015&ndash;2021&nbsp;funded by NWO are available in Open Access. The analyses presented in this report also cover publications funded by the Netherlands Organisation for Health Research and Development ZonMw.&nbsp;This report builds on two earlier reports, published in <a href="https://zenodo.org/record/4446042">2020</a> and <a href="https://zenodo.org/record/5056043">2021</a>, covering publications from the period 2015&ndash;2018 and 2015-2020, respectively.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

ACCESS-AM2 Southern Ocean cloud and radiation data for k-means clustering and analysis

<p>The ACCESS-AM2&nbsp;(Australian Community Climate and Earth-System Simulator - Atmospheric Model Version 2) data and k-means analysis used for the&nbsp;study described in Fiddes et al. 2022 &#39;<em>Southern Ocean cloud and shortwave radiation biases in a nudged climate model simulation: does the model ever get it right?&#39; .</em>&nbsp;</p> <p>Included files:&nbsp;</p> <ul> <li>modis_cluster_centres_2015-2019.nc&nbsp; - kmeans derived cluster centres for MODIS</li> <li>modis_cluster_labels_2015-2019.nc&nbsp; -&nbsp; kmeans derived cluster labels for MODIS&nbsp;</li> <li>bx400_cluster_labels_2015-2019.nc&nbsp; -&nbsp; kmeans fitted cluster label for model&nbsp;</li> <li>COSP_vars_bx400_2015-2019.nc&nbsp; -&nbsp; model data for analysis&nbsp;</li> </ul> <p>The code that performs the analysis/generates this data and has instructions for where to download MODIS data&nbsp;can be found here:&nbsp;https://github.com/sfiddes/code_for_publications_2022/tree/main/ACCESS_cloud_radiation_eval</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Supplementary data to `Do science maps from open access literature capture the overall topic structure of an academic field?`

<p>The dataset contains the 8,528 academic articles records related to Sustainable Food research sourced with the query `TS=("sustainab*" NEAR/2 "food*")` .</p> <p>They are the records present in the largest component of the citation network, as specified in the manuscript. &nbsp;</p> <p>The dataset was sourced from OpenAlex based on the original data used in the manuscript and it is composed of the following columns:</p> <table> <tbody> <tr> <td><em><strong>Column</strong></em></td> <td><em><strong>Description</strong></em></td> </tr> <tr> <td>Id</td> <td>OpenAlex ID</td> </tr> <tr> <td>DOI</td> <td>Document Object Identifier</td> </tr> <tr> <td>display_name</td> <td>The article title</td> </tr> <tr> <td>publication_year</td> <td>The publication year of the article</td> </tr> <tr> <td>open_access</td> <td>An object with details of the open access status of the article</td> </tr> </tbody> </table> <p>We choose the `.rdata` format for easy loading in R. Use the function `load()` to add the data frame to the enviroment.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Improving access to and reuse of research results, publications and data for scientific purposes - Stakeholders' consultations results

<p>The data sets were created via data collection effort for the Horizon Europe-funded "study to evaluate&nbsp;the effects of the EU copyright framework on research and the effects of potential interventions and to identify and present relevant provisions for research in EU data and digital legislation, with a focus on rights and obligations". The study was contacted by DG RTD.&nbsp;</p> <p>This research project supports Action 2 objectives of the European Research Area (ERA) Policy Agenda 2022-2024, which aims to propose&nbsp;an EU legislative and regulatory framework for copyright and data that is fit for research. The report provides a comprehensive analysis of barriers to the access and reuse of publicly funded research, including scientific publications and data. It assesses existing EU copyright legislation and EU data and digital legislation. It also assesses regulatory frameworks and national initiatives and identifies potential areas for improvement.</p> <p>Using a methodological, evidence-based approach (including the survey results posted in this repository), the study presents possible&nbsp;legislative and non-legislative measures to improve the current EU copyright and data framework and align it with the needs of scientific research and open research data principles.&nbsp;</p> <p>The data sets include the raw data of the three surveys (survey 1 targeted at researchers, survey 2 targeted at research-performing organisations, and survey 3 targeted at publishers). All surveys have two major parts: one concerning copyright legislation and another concerning data and digital legislation. In addition, we provide interview notes, they are also organised into two parts: one concerning copyright legislation and another concerning data and digital legislation.&nbsp;</p> <p>The data collection effort was partially supported by our colleagues from the Institute for Information Law (IVIR) and KU Leuven CiTIP.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Data and Statistical analysis for: "Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research"

<p>Data and Statistical analysis for: &quot;Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research&quot; published in&nbsp;<em>Frontiers in Marine Science</em></p>

openmit-licenseMar 2018View details →
zenodo44/100

Method Classification of Open Access INTACT Molecular Interaction data.

<p>Simple&nbsp;classification data derived from open access papers indexed in&nbsp;the INTACT database (https://www.ebi.ac.uk/intact/downloads) based on PSI-MI25 codes for interaction detection methods&nbsp;or participant detection methods based on the subfigure caption text.&nbsp;<br> <br> intact_records_and_captions_complete.tsv - This file links available text of subfigure captions to PSI-MI25 codes for the interaction detection method and participant detection method.&nbsp;&nbsp;</p> <p>evidx_run_file.txt - This file provides execution codes for the &#39;EvidX&#39; machine learning text&nbsp;classifier (https://github.com/SciKnowEngine/evidX/releases/tag/v0.1.0)</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

Monitoring and evaluation of UKRI's Open Access Policy: Exploring the use of open data sources to inform baseline values - Dataset

<p>This dataset accompanies the report <em>"Monitoring and evaluation of UKRI's Open Access Policy: Exploring the use of open data sources to inform baseline values"</em>, which is available via Zenodo.<br><br>It provides record-level data of UKRI-funded and UK-affiliated research output (limited to journal articles with Crossref DOIs) published between 2012 and 2022 - including bibliographic metadata as well as data on open access availability, publisher, national and international collaborations, citations, views and downloads, altmetrics and subjects (fields).&nbsp;All variables are documented in the data dictionary included in this Zenodo record.</p> <p>The code used to generate the dataset from open data sources is available on GitHub.&nbsp;</p> <p>The following data sources were used:</p> <ul> <li> <p>Gateway to Research (records downloaded between 2023-11-05 and 2023-11-13)</p> </li> <li> <p>Crossref (Metadata Plus snaphot 2023-10-31, Crossref member route API 2024-01-23)</p> </li> <li> <p>OpenAlex (data snapshot 2023-10-18)</p> </li> <li> <p>Unpaywall (data snapshot 2023-11-27)</p> </li> <li> <p>IRUS UK (2024-04-03)</p> </li> <li> <p>Crossref Event Data (2023-04-01)</p> </li> </ul> <p><strong></strong><br><br>The project made use of Curtin Open Knowledge Initiative (COKI) infrastructure, which is documented on GitHub: <a href="https://github.com/The-Academic-Observatory">https://github.com/The-Academic-Observatory</a>.&nbsp;</p>

opencc-zeroSep 2024View details →
zenodo44/100

Data accessibility in the chemical sciences: an analysis of recent practice in organic chemistry journals

<div> <p>Data is the analysis of the data outputs of 240 randomly selected research papers from 12 top-ranked journals published in early 2023. We investigate author compliance with recommended (but not compulsory) data policies, whether there is evidence to suggest that authors apply FAIR data guidance in their data publishing, and if the existence of specific recommendations for publishing NMR data by some journals encourages compliance. Files in the data package have been provided in both human and machine-readable forms. The main dataset is available in the Excel file Data worksheet.XLSX, the contents of which can also be found in Main_dataset.CSV, Data_types.CSV, and Article_selection.CSV with explanations of the variable coding used in the studies in Variable_names.CSV, Codes.CSV, and FAIR_variable_coding.CSV. The R code used for the article selection can be found in Article_selection.R. Data about article types from the journals that contain original research data is in Article_types.CSV. Data collected for analysis in our sister paper[4] can be found in Extended_Adherence.CSV, Extended_Crystallography.CSV, Extended_DAS.CSV, Extended_File_Types.CSV, and Extended_Submission_Process.CSV. A full list of files in the data package and a short description for each is given in README.TXT.</p> </div>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Environmental and AIS data collected during the EUMarineRobots Trans-National Access activities experiments using the NATO STO-CMRE Littoral Ocean Observatory Network testbed

<p>Environmental and AIS data collected during the H2020 project EUMarineRobots&nbsp;Trans-National Access activities&nbsp;experiments using the NATO STO-CMRE Littoral Ocean Observatory Network (LOON) testbed. Environmental data consists of temperature measured across the water column; sound velocity measured close to the surface and close to the sea bottom; meteorological data at the surface (i.e., pressure, temperature, wind speed and direction, humidity and rain). The environmental dataset is complemented with Automatic Identification System (AIS) data for the ships transiting close to &nbsp;the LOON area (Gulf of La Spezia, Italy)</p> <p>Temperature measured across the water column in the LOON area (Gulf of La Spezia, Italy). The dataset includes measurements for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021</p> <p><br> Meteorological data at the surface (i.e., pressure, temperature, wind speed and direction, humidity and rain) in the LOON area (Gulf of La Spezia, Italy). The dataset includes measurements for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021</p> <p><br> Sound velocity measured close to the surface (SVP1) and close to the sea bottom (SVP2) in the LOON area (Gulf of La Spezia, Italy). The dataset includes measurements for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021</p> <p>SVP2 data &nbsp;missing for &nbsp;Dec 14-20 (2020) and Jan 24, 27-28 (2021).</p> <p>Automatic Identification System (AIS) data for the ships transiting close to &nbsp;the LOON area (Gulf of La Spezia, Italy). The dataset includes AIS data for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021<br> &nbsp;</p> <p>For reference, see: &quot;Environmental data collected on the CMRE LOON tested during the EUMR project: dataset description&quot;,&nbsp;&nbsp;Petroccia, Roberto; Zappa, Giovanni; Cimino, Giampaolo; Grati, Alberto; Alves, Jo&atilde;o. CMRE-DA-2021-001. July 2021, available&nbsp; at&nbsp;https://www.cmre.nato.int/research/publications/latest-techreports/1638-cmre-da-2021-001</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Open-Access Data for "Received SignalStrength Measurements with BLE Signals for Contact Tracing and Proximity Detection"

<p>This archive contains three folders which are supplementary material for the paper accepted for publishing in IEEE Sensors Journal.</p> <p><strong>Contents:</strong></p> <ul> <li>&nbsp;The folder `open-access-data/upb/` contains the measurements acquired at UPB. The subfolders are named as `upb_ble_*`, where an asterisk masks&nbsp;the directory number. Whenever UPB is specified, use the data sets from the corresponding directory.</li> <li>The folder `open-access-data/tau/` contains the measurements acquired at TAU. The subfolders are named as `tau_ble_*`, where an asterisk masks the directory number. Whenever TAU is specified, use the data sets from the corresponding directory.</li> <li>The folder `open-access-data/wifi-on-off/` contains a sample code to read the files and plot the data from Fig. 14 in `open-access-data/wifi-on-off/wifi_on_off_read_plot.py` and Fig. 15 in `open-access-data/wifi-on-off/wifi_on_off_read_plot.ipynb`.</li> </ul> <p><strong>Results based on the data have been presented in the paper:</strong><br> Flueratoru, L., Shubina, V., Niculescu, D., Lohan, E.S. (2021). On the High Fluctuations of Received Signal Strength Measurements with BLE Signals for Contact Tracing and Proximity Detection, IEEE Sensors, Special Issue on Advanced Sensors and Sensing Technologies for Indoor Positioning and Navigation</p> <p><strong>To cite these data sets please use the following:</strong><br> Laura Flueratoru, Viktoriia Shubina, Dragoș Niculescu, &amp; Elena Simona Lohan. (2021). Open Access Data for &quot;Received SignalStrength Measurements with BLE Signals for Contact Tracing and Proximity Detection&quot; [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4643668</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

yangclaraliu/armslist_scraping: provide access to data

<p>This is the data used in the research letter:</p> <p>Drake, Coleman, Ashley M. Hernandez, Yang Liu, Adam H. Schwartz, and Maria E. Sundaram. &quot;Evidence of Background Checks in an Online Firearms Marketplace.&quot;&nbsp;<em>American journal of preventive medicine</em>&nbsp;57, no. 5 (2019): 718.</p> <p>https://www.ajpmonline.org/article/S0749-3797(19)30273-9/fulltext#%20</p>

opencc-byAug 2021View details →
zenodo44/100

Improving the data access control using blockchain for healthcare domain

<p>This research reviews the importance of blockchains in healthcare&nbsp;as they provide infinite possibilities to individuals, companies, and governments.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Data from the OPERAS business models survey on open access books

<p>OPERAS (the European Research Infrastructure for the development of open scholarly communication in the social sciences and humanities) has conducted a survey of publishing organisations throughout Europe to identify and better understand existing and potential business models to support the Open Access publication of research monographs. The results of the survey are used to inform the formulation of recommendations about how to create a sustainable open access book publishing ecosystem within Europe.</p> <p>The survey was designed to serve two core aims:&nbsp;<br> 1. To further, better or improve our understanding of the scholarly publishing landscape and of the challenges that publishers face in the context of publishing OA monographs;<br> 2. To identify main trends (including opportunities and challenges) and the knowledge of collaborative funding and infrastructure models in OA publishing in SSH.&nbsp;</p> <p>The survey was open&nbsp;between 16 February and 14 April 2021.</p> <p>The results are presented in two versions of the white paper&nbsp;of the Open Access Business Models Special Interest Group:&nbsp;&nbsp;Stone, Graham, Błaszczyńska, Marta, Lebon, Chlo&eacute;, Morka, Agata, Mosterd, Tom, Mounier, Pierre, Proudman, Vanessa, Speicher, Lara, &amp; Melin&scaron;čak Zlodi, Iva. (2021). Collaborative models for OA book publishers (1.0). Zenodo. https://doi.org/10.5281/zenodo.5494731 and the second version to be published in Spring 2023.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Malware Finances and Operations: a Data-Driven Study of the Value Chain for Infections and Compromised Access

<p><strong>Description</strong><br> <br> The datasets demonstrate the malware economy and the value chain published in our paper,&nbsp;<a href="https://doi.org/10.1145/3600160.3605047"><em>Malware Finances and Operations: a Data-Driven Study of the Value Chain for Infections and Compromised Access</em></a>,&nbsp;at the 12th International Workshop on Cyber Crime (IWCC 2023), part of the ARES Conference, published by the International Conference Proceedings Series of the ACM ICPS.</p> <p>Using the well-documented scripts, it is straightforward to reproduce our findings. It takes an estimated 1 hour of human time and 3 hours of computing time to duplicate our key findings from MalwareInfectionSet; around one hour with VictimAccessSet; and minutes to replicate the price calculations using AccountAccessSet. See the included README.md files and Python scripts.</p> <p>We choose to represent each victim by a single JavaScript Object Notation (JSON) data file. Data sources provide sets of victim JSON data files from which we&#39;ve extracted the essential information and omitted Personally Identifiable Information (PII). We collected, curated, and modelled three datasets, which we publish under the&nbsp;<a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>1. MalwareInfectionSet</strong><br> We discover (and, to the best of our knowledge, document scientifically for the first time) that malware networks appear to dump their data collections online. We collected these infostealer malware logs available for free. We utilise 245 malware log dumps from 2019 and 2020 originating from 14 malware networks. The dataset contains 1.8 million victim files, with a dataset size of 15 GB.</p> <p><strong>2. VictimAccessSet</strong><br> We demonstrate how Infostealer malware networks sell access to infected victims. Genesis Market focuses on user-friendliness and continuous supply of compromised data. Marketplace listings include everything necessary to gain access to the victim&#39;s online accounts, including passwords and usernames, but also detailed collection of information which provides a clone of the victim&#39;s browser session. Indeed, Genesis Market simplifies the import of compromised victim authentication data into a web browser session. We measure the prices on Genesis Market and how compromised device prices are determined. We crawled the website between April 2019 and May 2022, collecting the web pages offering the resources for sale. The dataset contains 0.5 million victim files, with a dataset size of 3.5 GB.</p> <p><strong>3. AccountAccessSet</strong><br> The Database marketplace operates inside the anonymous Tor network. Vendors offer their goods for sale, and customers can purchase them with Bitcoins. The marketplace sells online accounts, such as PayPal and Spotify, as well as private datasets, such as driver&#39;s licence photographs and tax forms. We then collect data from Database Market, where vendors sell online credentials, and investigate similarly. To build our dataset, we crawled the website between November 2021 and June 2022, collecting the web pages offering the credentials for sale. The dataset contains 33,896 victim files, with a dataset size of 400 MB.</p> <p><strong>Credits Authors</strong></p> <ul> <li>Billy Bob Brumley (Tampere University, Tampere, Finland)</li> <li>Juha Nurmi (Tampere University, Tampere, Finland)</li> <li>Mikko Niemel&auml; (Cyber Intelligence House, Singapore)</li> </ul> <p><strong>Funding</strong></p> <p>This project has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation programme under project numbers 804476 (SCARE) and 952622 (SPIRS).<br> <br> <strong>Alternative links to download:</strong>&nbsp;<a href="https://mega.nz/folder/aJwVyIYJ#9SWh-Z3-TpPfjHZeFxbeew">AccountAccessSet</a>,&nbsp;<a href="https://mega.nz/folder/iUQ3RaKB#48ZkXnFYSR0qXLcbkrZLqw">MalwareInfectionSet</a>, and <a href="https://mega.nz/folder/aNYCFCrK#pbDkJL-PNWjn1ABXbtdR4w">VictimAccessSet</a>.</p>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record