Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

431

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

431 results for “Training Datasets”

Learn how ShareScore rates datasets ↗
zenodo44/100

Phononic crystals dataset for supervised training of surrogate deep learning model

<p>The dataset contains shapes of unit cells of phononic crystals (inputs) in the form of images and corresponding dispersion diagrams (outputs). The dataset is used for deep learning (DL) model training.<br> Outputs are in the form of .mat files which contain vectors of reduced wavevector and corresponding frequencies, and also displacements u, v, w which can be used for polarization calculation.</p> <p>The dataset contains 11000 cases.</p> <p>Note: Ignore names &quot;labels&quot; as these are actually inputs to the DL model, not labels.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Dataset: Shell Commands Used by Participants of Hands-on Cybersecurity Training

<p>This repository contains supplementary materials for the following journal paper:</p> <p>Valdemar &Scaron;v&aacute;bensk&yacute;, Jan Vykopal, Pavel Seda, Pavel Čeleda.<br> <em>Dataset of Shell Commands Used by Participants of Hands-on Cybersecurity Training.</em><br> In Elsevier Data in Brief. 2021.<br> <a href="https://doi.org/10.1016/j.dib.2021.107398">https://doi.org/10.1016/j.dib.2021.107398</a></p> <ul> </ul> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <pre><code>@article{Svabensky2021dataset, author = {\v{S}v\'{a}bensk\'{y}, Valdemar and Vykopal, Jan and Seda, Pavel and \v{C}eleda, Pavel}, title = {{Dataset of Shell Commands Used by Participants of Hands-on Cybersecurity Training}}, journal = {Data in Brief}, publisher = {Elsevier}, volume = {38}, year = {2021}, issn = {2352-3409}, url = {https://doi.org/10.1016/j.dib.2021.107398}, doi = {10.1016/j.dib.2021.107398}, }</code></pre> <p>The data were collected using a logging toolset referenced <a href="https://zenodo.org/record/5126693">here</a>.</p> <p><strong>Attached content</strong></p> <ol> <li><strong>Dataset (data.zip).</strong> The collected data are attached here on Zenodo. A&nbsp;copy is also available in&nbsp;<a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/datasets/commands">this repository</a>.</li> <li><strong>Analytical tools (toolset.zip).</strong> To analyze the data, you can instantiate the toolset or&nbsp;<a href="https://gitlab.ics.muni.cz/muni-kypo/tools/commands-elk">this project for ELK</a>.</li> </ol> <p><strong>Version history</strong></p> <ul> <li>Version 1 (<a href="https://zenodo.org/record/5137355">https://zenodo.org/record/5137355</a>)&nbsp;contains&nbsp;13446 log records from 175 trainees. These data are precisely those that are described in the associated journal paper. Version 1 provides a snapshot of the state when the article was published.</li> <li>Version 2 (<a href="https://zenodo.org/record/5517479">https://zenodo.org/record/5517479</a>)&nbsp;contains&nbsp;13446 log records from 175 trainees. The data are unchanged from Version 1, but the analytical toolset includes a minor fix.</li> <li>Version 3 (<a href="https://zenodo.org/record/6670113">https://zenodo.org/record/6670113</a>)&nbsp;contains&nbsp;21762 log records from 275 trainees. It is a superset of Version 2, with newly collected data added to the dataset.</li> <li>The current Version 4 (<a href="https://zenodo.org/record/8136017">https://zenodo.org/record/8136017</a>)&nbsp;contains&nbsp;21459 log records from 275 trainees. Compared to Version 3, we cleaned 303 invalid/duplicate command records.</li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Datasets and trained diffusion models for "Diffusion Models for Interferometric Satellite Aperture Radar"

<p>A set of trained Probabilistic Diffusion Models (PDMs) and corresponding training datasets for the paper &quot;<a href="https://doi.org/10.48550/arXiv.2308.16847">Diffusion Models for Interferometric Satellite Aperture Radar</a>&quot;, by Tuel, Kerdreux et al. The code for this paper can be found at <a href="https://github.com/thomaskerdreux/PDM_SAR_InSAR_generation">this link</a>.</p> <p><strong>Training datasets</strong></p> <p>- &quot;InSAR_noise_32x32.zip&quot;: a dataset of 32x32 ground deformation scenes obtained from InSAR interferograms over New Mexico with the small baseline subset (SBAS) algorithm. Images were normalised to [0, 1].</p> <p>- &quot;insar_unwrapped_phase_normalised.zip&quot;: a dataset of 128x128 InSAR interferograms obtained from Sentinel-1 acquisitions over Nex Mexico. Images were normalised to [0, 1].</p> <p><strong>Trained Models</strong></p> <p>We provide 6 trained PDMs in separate .zip files. Each .zip file contains the model weights (in *.pt format) and the model metadata file (in *.json format).</p> <p>- &quot;mnist_32_cond_sigma_100.zip&quot;: a class-conditional model trained with 100 diffusion time steps on 32x32 MNIST images;</p> <p>- &quot;mnist_32_no_cond_sigma_100.zip&quot;: an unconditional model trained with 100 diffusion time steps on 32x32 MNIST images;</p> <p>- &quot;SAR_lowres_128_cond_sigma_2000.zip&quot;: a low-resolution (256 to 128) model trained with 2000 diffusion time steps on 128x128 TenGeoP-SARwv images;</p> <p>- &quot;SAR_superres_128_to_256_cond_sigma_2000.zip&quot;: a super-resolution (128 to 256) model trained with 2000 diffusion time steps on TenGeoP-SARwv images;</p> <p>- &quot;insar_phase_128_sigma_2000.zip&quot;: an unconditional model trained with 2000 diffusion time steps on 128x128 Sentinel-1 InSAR interferograms over New Mexico;</p> <p>- &quot;insar_noise_32_sigma_1000.zip&quot;: an unconditional model trained with 1000 diffusion time steps on 32x32 Sentinel-1 InSAR ground deformation scenes over New Mexico.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

RnR-ExM Training Dataset

<p>This dataset was released as part of the 2023 ISBI challenge,&nbsp;<a href="https://rnr-exm.grand-challenge.org">RnR-ExM</a>.&nbsp;The organizers thank Ruihan Zhang (MIT), Margaret Elizabeth Schroeder (MIT) and Chi Zhang (MIT) for contributing data to this competition.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Forest fire assessement training dataset (2022-07-18 fire at Maclas - France)

<p>This dataset has been created to train Univ. Eiffel personnels on raster data handling with QGIS.</p><p>It provides the following elements:</p><ul><li>Geopackage database with the following layers:<ul><li>QGIS project</li><li>Extract from the SENTINEL-2 2022-06-11 B8A band</li><li>Extract from the SENTINEL-2 2022-06-11 B12 band</li><li>Extract from the SENTINEL-2 2022-07-21 B8A band</li><li>Extract from the SENTINEL-2 2022-07-21 B12 band</li><li>Reclassified delta NBR raster layer</li><li>Delta NBR vector layer</li><li>Studied area bounding box</li></ul></li><li>Intermediate results:<ul><li>pre-event NBR raster file</li><li>post-event NBR raster file</li><li>Delta NBR raster file</li><li>Delta NBR raster file multiplied by 1000 (for easier reclassification)</li></ul></li></ul><p>Data sources IDs from opensearch-theia.cnes.fr-sentinel2-l2a catalogue :</p><ul><li>SENTINEL2B_20220721-104826-811_L2A_T31TFL_D</li><li>SENTINEL2B_20220611-104824-395_L2A_T31TFL_D</li></ul><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Training dataset used in the magazine paper entitled "A Flexible Machine Learning-Aware Architecture for Future WLANs"

<p><a href="https://arxiv.org/pdf/1910.03510.pdf"><strong>A Flexible Machine Learning-Aware Architecture for Future WLANs</strong></a></p> <p><strong>Authors: </strong>Francesc Wilhelmi, Sergio Barrachina-Mu&ntilde;oz, Boris Bellalta, Cristina Cano, Anders Jonsson &amp; Vishnu Ram.</p> <p><strong>Abstract:&nbsp;</strong>Lots of hopes have been placed in Machine Learning (ML) as a key enabler of future wireless networks. By taking advantage of the large volumes of data generated by networks, ML is expected to deal with the ever-increasing complexity of networking problems. Unfortunately, current networking systems are not yet prepared for supporting the ensuing requirements of ML-based applications, especially for enabling procedures related to data collection, processing, and output distribution. This article points out the architectural requirements that are needed to pervasively include ML as part of future wireless networks operation. To this aim, we propose to adopt the International Telecommunications Union (ITU) unified architecture for 5G and beyond. Specifically, we look into Wireless Local Area Networks (WLANs), which, due to their nature, can be found in multiple forms, ranging from cloud-based to edge-computing-like deployments. Based on ITU&#39;s architecture, we provide insights on the main requirements and the major challenges of introducing ML to the multiple modalities of WLANs.</p> <p><strong>Dataset description:&nbsp;</strong>This is the dataset generated for training a Neural Network (NN) in the Access Point (AP) (re)association problem in IEEE 802.11 Wireless Local Area Networks (WLANs).&nbsp;</p> <p>In particular, the NN is meant to output a prediction function of the throughput that a given station (STA) can obtain from a given Access Point (AP) after association. The features included in the dataset are:</p> <ol> <li>Identifier of the AP to which the STA has been associated.</li> <li>RSSI obtained from the AP to which the STA has been associated.</li> <li>Data rate in bits per second (bps) that the STA is allowed to use for the selected AP.</li> <li>Load in packets per second (pkt/s)&nbsp;that the STA generates.</li> <li>Percentage of data that the AP is able to serve before the user association is done.</li> <li>Amount of traffic load in pkt/s handled by the AP before the user association is done.</li> <li>Airtime in % that the AP enjoys before the user association is done.</li> <li>Throughput in pkt/s that the STA receives after the user association is done.</li> </ol> <p>The dataset has been generated through random simulations, based on the model provided in <a href="https://github.com/toniadame/WiFi_AP_Selection_Framework">https://github.com/toniadame/WiFi_AP_Selection_Framework</a>. More details regarding the dataset generation have been provided in&nbsp;<a href="https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans">https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans</a>.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

ZeroCostDL4Mic - CARE (2D) example training and test dataset

<p><strong>Name</strong>: ZeroCostDL4Mic - CARE (2D) example training and test dataset</p> <p>(see <a href="https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki">our Wiki</a> for details)</p> <p>&nbsp;</p> <p><strong>Data type</strong>: Paired microscopy images (fluorescence) of low and high signal-to-noise ratio</p> <p><strong>Microscopy data type</strong>: Fluorescence microscopy (Lifeact-RFP)</p> <p><strong>Microscope</strong>: Structured Illumination Microscopy (SIM) with a 60x 1.42 NA objective&nbsp;&nbsp;&nbsp;</p> <p><strong>Cell type</strong>: DCIS.COM Lifeact-RFP</p> <p><strong>File format</strong>: .tif (32-bit)</p> <p><strong>Image size</strong>: 1024x1024 (Pixel size: 40 nm)</p> <p>&nbsp;</p> <p><strong>Author(s)</strong>: Guillaume Jacquemet<sup>1,2</sup></p> <p><strong>Contact email</strong>: guillaume.jacquemet@abo.fi</p> <p><strong>Affiliation</strong>:&nbsp;</p> <p>1) Faculty of Science and Engineering, Cell Biology, &Aring;bo Akademi University, 20520 Turku, Finland</p> <p>2) Turku Bioscience Centre, University of Turku and &Aring;bo Akademi University, FI-20520 Turku, Finland</p> <p>&nbsp;</p> <p><strong>Associated publications</strong>: Unpublished</p> <p><strong>Funding bodies</strong>: G.J. was supported by grants awarded by the Academy of Finland, the Sigrid Juselius Foundation and &Aring;bo Akademi University Research Foundation (CoE CellMech) and by Drug Discovery and Diagnostics strategic funding to &Aring;bo Akademi University.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

ZeroCostDL4Mic - Noise2Void (2D) example training and test dataset

<p><strong>Name</strong>: ZeroCostDL4Mic - Noise2Void (2D) example training and test dataset</p> <p>(see <a href="https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki">our Wiki</a> for details)</p> <p>&nbsp;</p> <p><strong>Data type</strong>: Microscopy images (fluorescence)</p> <p><strong>Microscopy data type</strong>: Fluorescence microscopy (paxillin-GFP)&nbsp;</p> <p><strong>Microscope</strong>: Spinning disk confocal microscope with a 63x 1.4 NA objective&nbsp;</p> <p><strong>Cell type</strong>: U-251 glioma cells, endogenously expressing paxillin-GFP</p> <p><strong>File format</strong>: .tif (16-bit)</p> <p><strong>Image size</strong>: 512x512 (Pixel size: 248 nm)</p> <p>&nbsp;</p> <p><strong>Author(s)</strong>: Aki Stubb<sup>1</sup>, Guillaume Jacquemet<sup>1,2</sup> and Johanna Ivaska<sup>1</sup></p> <p><strong>Contact email</strong>: guillaume.jacquemet@abo.fi</p> <p><strong>Affiliation</strong>:&nbsp;</p> <p>1) Turku Bioscience Centre, University of Turku and &Aring;bo Akademi University, FI-20520 Turku, Finland</p> <p>2) Faculty of Science and Engineering, Cell Biology, &Aring;bo Akademi University, 20520 Turku, Finland</p> <p><br> &nbsp;</p> <p><strong>Associated publication</strong>: Stubb <em>et al.</em> 2020, Nano letters DOI: 10.1021/acs.nanolett.9b04083</p> <p><strong>Funding bodies</strong>: A.S. has been supported by the University of Turku Doctoral programme for Molecular Medicine (TuDMM).</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

DCASE 2020 Challenge Task 2 Additional Training Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;additional training&nbsp;dataset&quot; for the&nbsp;<strong>DCASE 2020 Challenge Task 2 &quot;Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring&quot; </strong><a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">[task description]</a>.&nbsp;</p> <p>In the task, three datasets have been or will be released:&nbsp;&quot;<a href="http://zenodo.org/record/3678171">development dataset</a>&quot;, &quot;additional training&nbsp;dataset&quot;,&nbsp;and &quot;<a href="https://zenodo.org/record/3841772">evaluation dataset</a>&quot;.&nbsp;This additional training&nbsp;dataset was released before the &quot;<a href="https://zenodo.org/record/3841772">evaluation dataset</a>&quot;.&nbsp;This&nbsp;dataset&nbsp;includes around 1,000 normal samples for each Machine Type and Machine ID used in the <a href="https://zenodo.org/record/3841772">evaluation dataset</a> and can be used for model training in advance.</p> <p>The recording procedure and data format are the same as&nbsp;the <a href="http://zenodo.org/record/3678171">development dataset</a>.&nbsp;The Machine IDs in this dataset are different from those in the <a href="http://zenodo.org/record/3678171">development dataset</a>.&nbsp;For more information, please see the pages of the&nbsp;<a href="http://zenodo.org/record/3678171">development dataset</a> and the <a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">task description</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Directory structure</strong></p> <p>Once&nbsp;you unzip the downloaded files from&nbsp;Zenodo, you can see the following directory structure. Machine Type information is given by directory name, and Machine ID and condition information are given by file name, as:</p> <ul> </ul> <p>/eval_data</p> <ul> <li>/ToyCar <ul> <li>/train (Only normal data for all Machine IDs are included.) <ul> <li>/normal_id_05_00000000.wav</li> <li>...</li> <li>/normal_id_05_00000999.wav</li> <li>/normal_id_06_00000000.wav</li> <li>...</li> <li>/normal_id_07_00000999.wav</li> </ul> </li> </ul> </li> <li>/ToyConveyor (The other Machine Types have the same directory structure as ToyCar.)</li> <li>/fan</li> <li>/pump</li> <li>/slider</li> <li>/valve</li> </ul> <p>&nbsp;</p> <p>The paths of audio files are:</p> <ul> <li>&quot;/eval_data/&lt;Machine_Type&gt;/train/normal_id_&lt;Machine_ID&gt;_[0-9]+.wav&quot;</li> </ul> <p>For example, the Machine Type and Machine ID of&nbsp;&quot;/ToyCar/train/normal_id_05_00000000.wav&quot; are &quot;ToyCar&quot; and &quot;05&quot;, respectively, and&nbsp;its condition is normal (This dataset includes only normal samples).&nbsp;</p> <p>&nbsp;</p> <p><strong>Baseline system</strong></p> <p>A simple baseline system is available&nbsp;on the Github repository <a href="https://github.com/y-kawagu/dcase2020_task2_baseline">[URL]</a>. The baseline system provides a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. It is a good starting point, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p>&nbsp;</p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>NTT Corporation</strong> and <strong>Hitachi, Ltd.</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three papers</strong>:</p> <p>Yuma Koizumi, Shoichiro Saito, Noboru Harada, Hisashi Uematsu, and Keisuke Imoto, &quot;ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection,&quot; in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019.&nbsp;<a href="https://ieeexplore.ieee.org/document/8937164">[pdf]</a></p> <p>Harsh Purohit, Ryo Tanabe, Kenji Ichige, Takashi Endo, Yuki Nikaido, Kaori Suefusa, and Yohei Kawaguchi, &ldquo;MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection,&rdquo; in Proc. 4th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019.&nbsp;<a href="http://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf">[pdf]</a></p> <p>Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto, Toshiki Nakamura, Yuki Nikaido, Ryo Tanabe, Harsh Purohit, Kaori Suefusa, Takashi Endo, Masahiro Yasuda, and Noboru Harada,&nbsp;&quot;Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring<em>,&quot;</em>&nbsp;in Proc. 5th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE),&nbsp;2020.&nbsp;<a href="https://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Koizumi_3.pdf">[pdf]</a></p> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yuma Koizumi, <a href="mailto:koizumi.yuma@ieee.org">koizumi.yuma@ieee.org</a></li> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>

opencc-by-nc-sa-4.0Mar 2020View details →
zenodo40/100

ZeroCostDL4Mic - Label-free prediction (fnet) example training and test dataset

<p><strong>Name</strong>: ZeroCostDL4Mic - Label-free prediction (fnet) example training and test dataset</p> <p>(see <a href="https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki">our Wiki</a>&nbsp;for details)</p> <p>&nbsp;</p> <p><strong>Data type</strong>: 3D paired microscopy images (fluorescence and transmitted light)</p> <p><strong>Microscopy data type</strong>:&nbsp;Confocal microscopy data (TOM20 labeled with Alexa Fluor 594)</p> <p><strong>Microscope</strong>:&nbsp;Leica SP8, HC PL APO 63x 1.40 NA oil objective&nbsp;</p> <p><strong>Cell type</strong>:&nbsp;HeLa (fixed using an organelle-preserving protocol)</p> <p><strong>File format</strong>:&nbsp;.tif (8-bit)</p> <p><strong>Image size</strong>:&nbsp;512 x 512 x 32&nbsp;(Pixel size: x and y: 90 nm pixel size, z: 150 nm)</p> <p>&nbsp;</p> <p><strong>Author(s)</strong>: Christoph Spahn</p> <p><strong>Contact email</strong>:&nbsp;c.spahn@chemie.uni-frankfurt.de, heilemann@chemie.uni-frankfurt.de</p> <p><strong>Affiliation</strong>:&nbsp;Institute of Physical and Theoretical Chemistry, Goethe-University Frankfurt, Frankfurt, Germany</p> <p>&nbsp;</p> <p><strong>Associated publications</strong>: Unpublished</p> <p><br> <strong>Funding body(ies)</strong>:&nbsp;M.H. and C.S.:&nbsp;German Science Foundation (grant nr. SFB1177).&nbsp;C.S.:&nbsp;European Molecular Biology Organization (short term fellowship 8589)</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Training dataset for semantic segmentation (U-Net) of structural conservation practices

<p>In this research, the best management practices include vegetative/structural conservation practices (SCP) across crop fields, such as grassed waterways&nbsp;and terraces. This reference dataset includes 500,000 pair patches (false-color image (B1: NIR, B2: Red, B3: Green)&nbsp;and binary label (SCP: yes[1] or no[0]).&nbsp;These training samples were randomly extracted from Iowa BMP project (<a href="https://www.gis.iastate.edu/gisf/projects/conservation-practices">https://www.gis.iastate.edu/gisf/projects/conservation-practices</a>) and present 90% of patches with SCP areas and 10% of patches non-SCP area. The patch dimension is 256 x&nbsp; 256 pixels at 2-m resolution. Due to the file size, the images were upload in different *.rar files (imagem_0_200k.rar, imagem_200_400k.rar, imagem_400_500k.rar), and the user should download all and merge them in the same folder. The corresponding labels are all in &quot;class_bin.rar&quot; file.</p> <p>Application: These pair images are useful for conservation practitioners interested in the classification of vegetative/structural SCPs using deep-learning semantic segmentation methods.</p> <p>Further information will be available in future.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

MESINESP: Medical Semantic Indexing in Spanish - Train dataset

<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>INTRODUCTION</strong>:</p> <p>The Mesinesp (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) training set has a total of 369,368 records.&nbsp;</p> <p>The training dataset contains all records from LILACS and IBECS databases at the Virtual Health Library (VHL) with a non-empty abstract written in Spanish. The URL used to retrieve records is as follows:<br> http://pesquisa.bvsalud.org/portal/?output=xml&amp;lang=es&amp;sort=YEAR_DESC&amp;format=abstract&amp;filter[db][]=LILACS&amp;filter[db][]=IBECS&amp;q=&amp;index=tw&amp;</p> <p>We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;</p> <p>The training dataset was crawled on 10/22/2019. This means that the data is a snapshot of that moment and that may change over time. In fact, it is very likely that the data will undergo minor changes as the different databases that make up LILACS and IBECS may add or modify the indexes.</p> <p>&nbsp;</p> <p><strong>ZIP STRUCTURE:</strong></p> <p>The training data sets contain 369,368 records from 26,609 different journals. Two different data sets are distributed as described below:</p> <p>&nbsp;- <em>Original Train set</em> with 369,368 records that also include the qualifiers, as retrieved from VHL.&nbsp;<br> &nbsp;- <em>Pre-processed Train set</em><strong> </strong>with the 318,658 records with at least one DeCS code and with no qualifiers.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>STATISTICS</strong>:</p> <p>Abstracts&rsquo; length (measured in characters)<br> Min: 12<br> Avg: 1140.41<br> Median: 1094<br> Max: 9428</p> <p>Number of DeCS codes per file<br> Min: 1<br> Avg: 8.12<br> Median: 7<br> Max: 53</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>CORPUS FORMAT</strong>:</p> <p>The training data sets are distributed as a JSON file with the following format:</p> <pre><code>{   "articles": [     {       "id": "Id of the article",       "title": "Title of the article",       "abstractText": "Content of the abstract",       "journal": "Name of the journal",       "year": 2018,       "db": "Name of the database",       "decsCodes": [         "code1",         "code2",         "code3"       ]     }   ] } </code></pre> <p>Note that the decsCodes field lists the DeCs Ids assigned to a record in the source data. Since the original XML data contain descriptors (no codes), we provide a DeCs conversion table (https://temu.bsc.es/mesinesp/wp-content/uploads/2019/12/DeCS.2019.v5.tsv.zip) with:</p> <p>&nbsp;- DeCs codes<br> &nbsp;- Preferred descriptor (the label used in the European DeCs 2019 set)<br> &nbsp;- List of synonyms (the descriptors and synonyms from both European and Latin Spanish DeCs 2019 data sets, separated by pipes)</p> <p>&nbsp;</p> <p>For more details on the Latin and European Spanish DeCs codes see: http://decs.bvs.br and http://decses.bvsalud.org/ respectively.</p> <p>Please, cite: Krallinger M, Krithara A, Nentidis A, Paliouras G, Villegas M. BioASQ at CLEF2020: Large-Scale Biomedical Semantic Indexing and Question Answering. InEuropean Conference on Information Retrieval 2020 Apr 14 (pp. 550-556). Springer, Cham.</p> <p>&nbsp;</p> <p>Copyright (c) 2020 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0May 2020View details →
zenodo40/100

A Dataset of Pull Requests and A Trained Random Forest Model for predicting Pull Request Acceptance

<p>A Curated Dataset of 470,925 pull requests for 3349 popular NPM packages, description of the variables, code snippet for creating a Random Forest model for predicting pull request acceptance, and a pre-trained&nbsp;&nbsp;Random Forest model (in R). The dataset is for the ESEM-2020 paper: &quot;Impact of Technical and Social Factors on Pull Request Quality for the NPM Ecosystem&quot; (<a href="https://arxiv.org/abs/2007.04816">https://arxiv.org/abs/2007.04816</a>).&nbsp;</p> <p>Citation:</p> <pre>@inproceedings{dey2020effect, title={Effect of technical and social factors on pull request quality for the npm ecosystem}, author={Dey, Tapajit and Mockus, Audris}, booktitle={Proceedings of the 14th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM)}, pages={1--11}, year={2020} }</pre>

opencc-by-4.0May 2020View details →
zenodo40/100

Popularity Dataset for Online Stats Training

<p>This is a dataset&nbsp;used for the online stats training website (<a href="https://www.rensvandeschoot.com/tutorials/">https://www.rensvandeschoot.com/tutorials/</a>) and is based on the data used by&nbsp;&nbsp;<a href="https://doi.org/10.1016/j.adolescence.2009.12.004">van de Schoot, van der Velden, Boom, and Brugman (2010)</a>.</p> <p>The dataset is based on a study that investigates an association between popularity status and antisocial behavior from at-risk adolescents (n = 1491), where gender and ethnic background are moderators under the association. The study distinguished subgroups within the popular status group in terms of overt and covert antisocial behavior.For more information on the sample, instruments, methodology, and research context, we refer the interested readers to <a href="https://doi.org/10.1016/j.adolescence.2009.12.004">van de Schoot, van der Velden, Boom, and Brugman (2010)</a>.</p> <p>&nbsp;</p> <p>Variable name&nbsp;&nbsp; Description</p> <p>Respnr =&nbsp; Respondents&rsquo; number</p> <p>Dutch =&nbsp; Respondents&rsquo; ethnic background (0 = Dutch origin, 1 = non-Dutch origin)</p> <p>gender&nbsp; = Respondents&rsquo; gender (0 = boys, 1 = girls)</p> <p>sd =&nbsp;&nbsp;Adolescents&rsquo; socially desirable answering patterns</p> <p>covert =&nbsp;Covert antisocial behavior</p> <p>overt =&nbsp; Overt antisocial behavior</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Endless Foram, MD022508 and MD9712138 training datasets

<p>Training datasets for the paper &quot;Automated analysis of foraminifera fossil records by image classification using a convolutional neural network&quot;</p> <p>The Endless Forams dataset is a derivative work, and was created from the original dataset (endlessforams.org,&nbsp;<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2019PA003612">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2019PA003612</a>) by removing the text and border from each image.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Dataset and trained models belonging to the article 'Distant reading patterns of iconicity in 940.000 online circulations of 26 iconic photographs'

<p>Quantifying Iconicity - Zenodo</p> <p><br> ## The Dataset<br> This dataset contains the material collected for the article &quot;Distant reading 940,000 online circulations of 26 iconic photographs&quot; (to be) published in New Media &amp; Society (DOI: 10.1177/14614448211049459). We identified 26 iconic photographs based on earlier work (Van der Hoeven, 2019). The Google Cloud Vision (GCV) API was subsequently used to identify webpages that host a reproduction of the iconic image. The GCV API uses computer vision methods and the Google index to retrieve these reproductions. The code for calling the API and parsing the data can be found on GitHub: https://github.com/rubenros1795/ReACT_GCV.</p> <p>The core dataset consists of .tsv-files with the URLs that refer to the webpages. Other metadata provided by the GCV API is also found in the file and manually generated metadata. This includes:<br> - the URL that refers specifically to the image. This can be an URL that refers to a full match or a partial match<br> - the title of the page<br> - the iteration number. Because the GCV API puts a limit on its output, we had to reupload the identified images to the API to extend our search. We continued these iterations until no more new unique URLs were found<br> - the language found by the ``langid`` Python module [link](https://github.com/saffsd/langid.py), along with the normalized score.<br> - the labels associated with the image by Google<br> - the scrape date</p> <p>Alongside the .tsv-files, there are several other elements in the following folder structure:</p> <p>```<br> ├── data<br> │&nbsp;&nbsp; ├── embeddings<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── doc2vec<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── input-text<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── metadata<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── umap<br> │&nbsp;&nbsp; └── evaluation<br> │&nbsp;&nbsp; └── results<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── diachronic-plots<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── top-words<br> │&nbsp;&nbsp; └── tsv<br> ```</p> <p>1. The ```/embeddings``` folder contains the doc2vec models, the training input for the models, the metadata (id, URL, date) and the UMAP embeddings used in the GMM clustering. Please note that the date parser was not able to find dates for all webpages and for this reason not all training texts have associated metadata.<br> 2. The ```/evaluation``` folder contains the AIC and BIC scores for GMM clustering with different numbers of clusters.<br> 3. The ```/results``` folder contains the top words associated with the clusters and the diachronic cluster prominence plots.</p> <p>## Data Cleaning and Curation<br> Our pipeline contained several interventions to prevent noise in the data. First, in between the iterations we manually checked the scraped photos for relevance. We did so because reuploading an iconic image that is paired with another, irrelevant, one results in reproductions of the irrelevant one in the next iteration. Because we did not catch all noise, we used Scale Invariant Feature Transform (SIFT), a basic computer vision algorithm, to remove images that did not meet a threshold of ten keypoints. By doing so we removed completely unrelated photographs, but left room for variations of the original (such as painted versions of Che Guevara, or cropped versions of the Napalm Girl image). Another issue was the parsing of webpage texts. After experimenting with different webpage parsers that aim to extract &#39;relevant&#39; text it proved too difficult to use one solution for all our webpages. Therefore we simply parsed all the text contained in commonly used html-tags, such as ```&lt;p&gt;```, ```&lt;h1&gt;``` etc.</p>

openNov 2020View details →
zenodo40/100

Training dataset: Mass spectrometry based proteomics of healthy human serum samples

<p>The two raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>Serum of a healthy person was obtained by centrifugation of full blood in a serum-gelmonovette. One serum sample was depleted for high abundant proteins, the other not.<br> For the non-depleted sample: 5&micro;l of serum was diluted with 0.1% Rapigest, resulting in a concentration of 1mg/ml.<br> Depletion was performed with the Seppro IgY14 Spin columns which are able to deplete 14 high abundent blood proteins by immunoaffinity. For the depleted sample 9&micro;l of serum was diluted with TBS/HCl/NaCl buffer and added to the Seppro IgY14 spin column. After depletion the sample was buffered with Hepes pH 8.0 and Rapigest was added to a final 0.1% Rapigest concentration. From here on, both samples were reduced by adding TCEP, alkylated by IAA and quenched with DTT in solution. Digestion was performed by adding trypsin in a ratio of 1:50 to the samples. After incubation at 37&deg;C, 600rpm, over night, the sample clean-up was performed with the PreOmics desalting columns. iRT peptides were added and the sample was measured with a Q-Exactive Plus mass spectrometer. Besides the two raw files, we uploaded a fasta file that serves as human protein sequence database and the Galaxy MaxQuant training result files: protein groups, peptides, mqpar and PTXQC.</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Enhanced Bug Prediction in JavaScript Programs with Hybrid Call-Graph Based Invocation Metrics (Training Dataset)

<p>This dataset consists of multiple files which contain bug prediction training data.</p> <p>The entries in the dataset are JavaScript functions either being buggy or non-buggy. Bug related information was obtained from the project EsLint contained in BugsJS (https://github.com/BugsJS/eslint). The buggy instances were collected throughout the lifetime of the project, however we added non-buggy entries from the latest version which is tagged as fix (entries which were previously included as buggy were not included as non-buggy later on).</p> <p>The dataset is based on hybrid call graphs&nbsp;which are constructed by&nbsp;https://github.com/sed-szeged/hcg-js-framework. The result of this tool is a call graph where the edges are associated with a confidence level which shows how likely the given edge is a valid call edge.</p> <p>We used different threshold values from which we considered the edges to be valid. The following threshold values were used:</p> <ul> <li>0.00</li> <li>0.05</li> <li>0.20</li> <li>0.30</li> </ul> <p>The prefix in the dataset file names are coming from the used threshold. The the datasets include coupling metrics NII (Nubmer of Incoming Invocations) and NOI (Number of Outgoing Invocations) which were calculated by a static source code analyzer called SourceMeter. Hybrid counterparts of these metrics (HNII and HNOI) are based on the given threshold values.</p> <p>There are four variants for all of these datasets:</p> <ul> <li>Both static (NII, NOi) and hybrid (HNII, HNOI) coupling metrics are included&nbsp;with additional static source code metrics and information about the entries (file without any&nbsp;postfix). Column contained only in this dataset are: <ul> <li>ID</li> <li>Name</li> <li>Longname</li> <li>Parent ID</li> <li>Component ID</li> <li>Path</li> <li>Line</li> <li>Column</li> <li>EndLine</li> <li>EndColumn</li> </ul> </li> <li>Both static (NII, NOi) and hybrid (HNII, HNOI) coupling metrics are included&nbsp;with additional&nbsp;static source code metrics&nbsp;(file with &#39;_h+s&#39; postfix)</li> <li>Only static (NII, NOI) coupling metrics are included with additional static source code metrics&nbsp;(file with &#39;_s&#39; postfix)</li> <li>Only hybrid (HNII, HNOI) coupling metrics are included with additional static source code metrics (file with &#39;_h&#39; postfix)</li> </ul> <p>Static source code metrics which are contained in all dataset are the following:</p> <ul> <li>McCC - McCabe Cyclomatic Complexity</li> <li>NL - Nesting Level</li> <li>NLE - Nesting Level&nbsp;Else If</li> <li>CD - Comment Density</li> <li>CLOC - Comment Lines of Code</li> <li>DLOC - Documentation Lines of Code</li> <li>TCD - Total Comment Density (Comment Lines in an emedded function will be also considered)</li> <li>TCLOC - Total Comment Lines of Code&nbsp;(Comment Lines in an emedded function will be also considered)</li> <li>LLOC - Logical Lines of Code (Comment and empty lines not counted)</li> <li>LOC - Lines of Code (Comment and empty lines are counted)</li> <li>NOS - Number of Statements</li> <li>NUMPAR - Number of Parameters</li> <li>TLLOC -&nbsp;Logical Lines of Code (Lines in embedded functions are also counted)</li> <li>TLOC -&nbsp;Lines of Code (Lines in embedded functions are also counted)</li> <li>TNOS - Total Number of Statements (Statements in embedded functions are also counted)</li> </ul>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Training dataset: Generation of a spectral library from HEK-Ecoli Spike-in mass spectrometry data

<p>The five raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 &deg;C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in the following ratios (amount in &micro;g):</p> <p>Sample&nbsp;&nbsp; &nbsp;HEK&nbsp;&nbsp; &nbsp;E.coli&nbsp;&nbsp; &nbsp;MS method<br> Sample1&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.00&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample2&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.05&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample3&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample4&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.40&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample5&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DDA</p> <p>Additionally, iRT peptides were added and 1&micro;g of each samples&nbsp;was measured with a Q-Exactive Plus mass spectrometer. Besides the five&nbsp;raw files, we uploaded two&nbsp;fasta files that serve&nbsp;as human and ecoli protein sequence databases, an transition list for the iRT peptides as well as an experimental design for the MaxQuant search.<br> Additionally, we uploaded&nbsp;the Galaxy MaxQuant training result files: protein groups, peptides, mqpar, msms, evidence&nbsp;and PTXQC.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

NewsEye / READ OCR training dataset from French Newspapers (18th, 19th, early 20th C.)

<p>The dataset comprises French newspaper pages from 18th, 19th and early 20th century with carefully corrected text. The page images were provided by the&nbsp;<a href="https://www.bnf.fr/en">French National Library</a> and comprise 127 pages (training set) and 8 pages (validation set). The data are formed according to the PAGE format (cf.&nbsp;Cf.&nbsp;<a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a>&nbsp;and the&nbsp;<a href="http://read.transkribus.eu/">READ </a>project.</p>

opencc-by-4.0Nov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record