Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11,687
datasets available to search
ShareScore release 0.7.1
Dataset results
11,687 results for “training”
replicAnt - Plum2023 - Pose-Estimation Datasets and Trained Models
<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for semantic and instance segmentation experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data have been generated with the open-source photgrammetry platform scAnt <a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>. All synthetic data has been generated with the associated replicAnt project available from <a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Two pose-estimation datasets were procured. Both datasets used first instar <em>Sungaya nexpectata</em> (Zompro 1996) stick insects as a model species. Recordings from an evenly lit platform served as representative for controlled laboratory conditions; recordings from a hand-held phone camera served as approximate example for serendipitous recordings in the field. </p> <p>For the platform experiments, walking <em>S. inexpectata</em> were recorded using a calibrated array of five FLIR blackfly colour cameras (Blackfly S USB3, Teledyne FLIR LLC, Wilsonville, Oregon, U.S.), each equipped with 8 mm c-mount lenses (M0828-MPW3 8MM 6MP F2.8-16 C-MOUNT, CBC Co., Ltd., Tokyo, Japan). All videos were recorded with 55 fps, and at the sensors’ native resolution of 2048 px by 1536 px. The cameras were synchronised for simultaneous capture from five perspectives (top, front right and left, back right and left), allowing for time-resolved, 3D reconstruction of animal pose.<br> <br> The handheld footage was recorded in landscape orientation with a Huawei P20 (Huawei Technologies Co., Ltd., Shenzhen, China) in stabilised video mode: <em>S. inexpectata </em>were recorded walking across cluttered environments (hands, lab benches, PhD desks etc), resulting in frequent partial occlusions, magnification changes, and uneven lighting, so creating a more varied pose-estimation dataset.<br> <br> Representative frames were extracted from videos using DeepLabCut (DLC)-internal k-means clustering. 46 key points in 805 and 200 frames for the platform and handheld case, respectively, were subsequently hand-annotated using the DLC annotation GUI.</p> <p><strong>Synthetic data</strong></p> <p>We generated a synthetic dataset of 10,000 images at a resolution of 1500 by 1500 px, based on a 3D model of a first instar <em>S. inexpectata </em>specimen, generated with the <a href="https://peerj.com/articles/11155/"><em>scAnt</em> photogrammetry workflow</a>. Generating 10,000 samples took about three hours on a consumer-grade laptop (6 Core 4 GHz CPU, 16 GB RAM, RTX 2070 Super). We applied 70\% scale variation, and enforced hue, brightness, contrast, and saturation shifts, to generate 10 separate sub-datasets containing 1000 samples each, which were combined to form the full dataset.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College’s President’s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>
replicAnt - Plum2023 - Segmentation Datasets and Trained Models
<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for semantic and instance segmentation experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data have been generated with the open-source photgrammetry platform scAnt <a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>. All synthetic data has been generated with the associated replicAnt project available from <a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Semantic and instance segmentation is used only rarely in non-human animals, partially due to the laborious process of curating sufficiently large annotated datasets. <em>replicAnt </em>can produce pixel-perfect segmentation maps with minimal manual effort. In order to assess the quality of the segmentations inferred by networks trained with these maps, semi-quantitative verification was conducted using a set of macro-photographs of <em>Leptoglossus zonatus</em> (Dallas, 1852) and <em>Leptoglossus phyllopus</em> (Linnaeus, 1767), provided by Prof. Christine Miller (University of Florida), and Royal Tyler (Bugwood.org. For further qualitative assessment of instance segmentation, we used laboratory footage, and field photographs of <em>Atta vollenweideri</em> provided by Prof. Flavio Roces. More extensive quantitative validation was infeasible, due to the considerable effort involved in hand-annotating larger datasets on a per-pixel basis.</p> <p><strong>Synthetic data</strong></p> <p>We generated two synthetic datasets from a single 3D scanned <em>Leptoglossus zonatus</em> (Dallas, 1852) specimen: one using the default pipeline, and one with additional plant assets, spawned by three dedicated scatterers. The plant assets were taken from the Quixel library and include 20 grass and 11 fern and shrub assets. Two dedicated grass scatterers were configured to spawn between 10,000 and 100,000 instances; the fern and shrub scatterer spawned between 500 to 10,000 instances. A total of 10,000 samples were generated for each sub dataset, leading to a combined dataset comprising 20,000 image render and ID passes. The addition of plant assets was necessary, as many of the macro-photographs also contained truncated plant stems or similar fragments, which networks trained on the default data struggled to distinguish from insect body segments. The ability to simply supplement the asset library underlines one of the main strengths of <em>replicAnt</em>: training data can be tailored to specific use cases with minimal effort.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College’s President’s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>
Galaxy Training Material for Mass spectrometry: GC-MS data processing (with XCMS, RAMClustR, RIAssigner, and matchms)
<p>This dataset contains the training data for the <strong>Mass spectrometry: GC-MS data processing (with XCMS, RAMClustR, RIAssigner, and matchms)</strong> GTN tutorial. It includes 3 GC-[EI+]-HRMS files from seminal plasma samples, the RECETOX Metabolome HR-[EI+]-MS library collected from mostly endogoenous compounds from MetaSci Human Metabolite Library, reference alkanes, sample metadata table, and preprocessed XCMS object.</p>
Physiological Signals During Motor Imagery Brain-Computer Interface Training Using Virtual Reality and Haptics
<p><strong>Participant demographics:</strong></p> <p>The sample is consisted by 20 healthy volunteers with a mean age of 24.79 years (SD = 3.54 years). The cohort was 68% male and 32% female. In terms of education, 16% had attended only high school, while 32% had a bachelor's degree, 42% a master's degree, and 11% a doctorate. All participants signed an informed consent before participating in the study in accordance with the 1964 Declaration of Helsinki.</p> <p><strong>Experiment Description:</strong></p> <p>The experiment consisted in having the subjects perform motor imagery of a bimanual rowing task with two individual paddles, one in each hand, under five experimental conditions. Four of these conditions used NeuRow (<a href="https://link.springer.com/chapter/10.1007/978-3-030-27950-9_1"><strong>Vourvopoulos et al. (2016-2019</strong>))</a>—a VR environment that renders virtual arms from a first-person perspective—while the other conditions used abstract feedback based on the BCI-Graz paradigm<a href="https://ieeexplore.ieee.org/abstract/document/1214714"> (<strong>Pfurtscheller et al. (2003))</strong></a>. All six conditions and their acronyms are described below:</p> <ol> <li><strong>Motor Imagery(MI)</strong>: The standard motor imagery training, with a fixation cross and directional arrows on a black background guiding the subjects through the experiment.</li> <li><strong>Motor Imagery/Motor Observation (MIMO):</strong> A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a monitor.</li> <li><strong>Motor Imagery/Motor Observation with Haptics (MIMOHP): </strong>A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a monitor. Hand controllers also provided haptic feedback through vibrotactile stimulation.</li> <li><strong>Motor Imagery/Motor Observation with VR HMD (MIMOVR):</strong> A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a VR HMD.</li> <li><strong>Motor Imagery/Motor Observation with VR HMD and Haptics (MIMOVRHP):</strong> A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a VR HMD. Hand controllers also provided haptic feedback through vibrotactile stimulation.</li> <li><strong>Motor Execution (ME):</strong> A fixation cross and directional arrows were displayed on a black background through a monitor (same as in MI), and guided the subjects through the experiment by having them tap their fingers accordingly. Data from this condition was available only after S07, so only 10 subjects<br> have performed ME.</li> </ol> <p>Finally, this experiment followed a within-subject design, in a randomized order of the conditions to minimize any order effects, while MI and ME conditions acted as control.</p> <p><strong>Equipment:</strong></p> <p>A wireless EEG amplifier (LiveAmp; Brain Products GmbH, Gilching, Germany) was used, with 32 active electrodes(+3 ACC) with a sampling rate of 500Hz. In addition, <strong>ECG, PPG</strong> and <strong>Respiration</strong> signals have been recorded synchronously in a bipolar montage, and connected to the EEG amplifier’s AUX input through the Brain Products BIP2AUX adapter.</p> <p>Visual feedback was provided through a monitor in all conditions except in MIMOVR and MIMOVRHP, in which an Oculus Rift CV1 headset (Reality Labs, formerly Facebook, Inc., CA, USA) was used instead. Haptic feedback was provided through the Oculus Rift hand controllers.<br> </p> <p><strong>Channel Indices:</strong></p> <p><strong>EEG</strong>: 1-32<br> <strong>PPG</strong> (AUX1): 33<br> <strong>Resp</strong>. (AUX2): 34<br> <strong>ECG</strong> (AUX3): 35<br> <strong>ACC</strong>: 36-38</p> <p> </p> <p><strong>Event codes:</strong></p> <table> <tbody> <tr> <td><strong>Code</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>S01</td> <td>Experiment Start</td> </tr> <tr> <td>S02</td> <td>Baseline Start</td> </tr> <tr> <td>S03</td> <td>Baseline Stop</td> </tr> <tr> <td>S04</td> <td>Start Of Trial</td> </tr> <tr> <td>S05</td> <td>Cross On Screen</td> </tr> <tr> <td>S07</td> <td>class1, Left hand </td> </tr> <tr> <td>S08</td> <td>class2, Right hand </td> </tr> <tr> <td>S09</td> <td>Feedback Continuous</td> </tr> <tr> <td>S10</td> <td>End of Trial</td> </tr> <tr> <td>S11</td> <td>End Of Session</td> </tr> <tr> <td>S12</td> <td>Experiment Stop</td> </tr> </tbody> </table> <p> </p> <p><strong>Directory tree:</strong></p> <p>ROOT<br> |<br> +--- USER #<br> | +---SESSION #<br> | | +---TASK #<br> | | | +---MI<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMO<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMOHP<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMOVR<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMOHPVR<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---ME<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk</p> <p> </p> <p><strong>Note: </strong>The first three datasets are from pilot sessions: sub-p01 to p03. From sub-01 to 19, subjects 10 and 11 have been removed due to the lack of markers. Subject sub-13, task MIMOVRHP is missing.</p> <p> </p>
Alterations in RNA editing in skeletal muscle following exercise training in individuals with Parkinson's disease
<p>Parkinson’s Disease (PD) is the second most common neurodegenerative disease behind Alzheimer’s Disease, currently affecting more than 10 million people worldwide. The progression of PD results in the loss of function due to neurodegeneration and neuroinflammation. The etiology of PD is multifactorial, including both genetic and environmental origins. We explored changes in RNA editing, specifically editing through the actions of the Adenosine Deaminases Acting on RNA (ADARs), in the progression of PD. Analysis of ADAR editing of skeletal muscle transcriptomes from PD patients and controls, including those that engaged in a rehabilitative exercise training program revealed significant differences in ADAR editing patterns based on age, disease status, and following rehabilitative exercise. Further, deleterious editing events in protein coding regions were identified in multiple genes with known associations to PD pathogenesis. Our findings of differential ADAR editing complement findings of changes in transcriptional network identified by a recent Lavin et al. 2020 (<a href="https://doi.org/10.3389/fphys.2020.00653">https://doi.org/10.3389/fphys.2020.00653)</a> study and offer insights into dynamic ADAR editing changes associated with PD pathogenesis. VCF files were generated using AIDD (Plonski et al., 2020) (<a href="https://doi.org/10.1186/s12859-020-03888-6">https://doi.org/10.1186/s12859-020-03888-6</a>).</p>
Image and label patches used to train GLASS-AI
<p>This archive contains the paired image and label patches used to train our machine learning pipeline, Grading of Lung Adenocarcinoma with Simultaneous Segmentation by Artificial Intelligence (GLASS-AI). </p> <p>Image patches were generated from whole slide images of H&E-stained sections using an Aperio ScanScope AT2 Slide Scanner (Leica) at 20x magnification with a 0.5022 microns/pixel resolution. The individual tumors and airways were annotated by an expert human before being divided into 224x224 pixel patches of the H&E image and paired annotation layer.</p> <p>For more details regarding how these data were used to train GLASS-AI, please see our forthcoming manuscript. </p>
Molecular adaptations in response to exercise training are associated with tissue-specific transcriptomic and epigenomic signatures
<p>Processed data associated with the manuscript DOI: <a href="https://doi.org/10.1016/j.xgen.2023.100421" target="_blank" rel="noopener">10.1016/j.xgen.2023.100421 </a></p> <p>Analysis code on GitHub: <a href="../doi/10.5281/zenodo.8253917" target="_blank" rel="noopener">10.5281/zenodo.8253917</a></p> <p> </p>
Phononic crystals dataset for supervised training of surrogate deep learning model
<p>The dataset contains shapes of unit cells of phononic crystals (inputs) in the form of images and corresponding dispersion diagrams (outputs). The dataset is used for deep learning (DL) model training.<br> Outputs are in the form of .mat files which contain vectors of reduced wavevector and corresponding frequencies, and also displacements u, v, w which can be used for polarization calculation.</p> <p>The dataset contains 11000 cases.</p> <p>Note: Ignore names "labels" as these are actually inputs to the DL model, not labels.</p>
QLKNN7D-edge training set
<p><strong>QLKNN7D-edge training set</strong></p> <p>This dataset contains a large-scale run of ~15 million flux calculations of the quasilinear gyrokinetic transport model QuaLiKiz. The dataset is in a parameter regime typical of the L-mode near edge (pedestal forming region). QuaLiKiz is applied in numerous tokamak integrated modelling suites, and is openly available at <a href="https://gitlab.com/qualikiz-group/QuaLiKiz/">https://gitlab.com/qualikiz-group/QuaLiKiz/</a>. This dataset was generated with QuaLiKiz 2.8.4, which includes numerical improvements increasing the robustness of strongly driven (high gradient) calculations typical of the L-mode near-edge. See <a href="https://gitlab.com/qualikiz-group/QuaLiKiz/-/tags/2.8.4">https://gitlab.com/qualikiz-group/QuaLiKiz/-/tags/2.8.4</a> for the in-repository tag.</p> <p>The dataset is appropriate for the training of learned surrogates of QuaLiKiz, e.g. with neural networks. See <a href="https://doi.org/10.1063/1.5134126">https://doi.org/10.1063/1.5134126</a> for a Physics of Plasmas publication illustrating the development of a learned surrogate (QLKNN10D-hyper) of an older version of QuaLiKiz (2.4.0) with a 300 million point 10D dataset. The paper is also available on <a href="https://arxiv.org/abs/1911.05617">arXiv</a> and the older dataset on <a href="https://doi.org/10.5281/zenodo.3497066">Zenodo</a>. For an application example, see <a href="http://https://doi.org/10.1088/1741-4326/ac0d12">Van Mulders et al 2021</a>, where QLKNN10D-hyper was applied for ITER hybrid scenario optimization. An additional, larger, QuaLiKiz dataset is found at <a href="https://zenodo.org/record/8017522">https://zenodo.org/record/8017522</a>. Neither the QLKNN10D or QLKNN11D datasets include L-mode near-edge parameters. For any learned surrogates developed for QLKNN7D-edge, the effective addition of the alphaMHD input dimension through rescaling the input magnetic shear (s) by s = s - alpha_MHD/2, as carried out in Van Mulders et al., is recommended.</p> <p>Related repositories:</p> <ul> <li><a href="https://qualikiz.com">General QuaLiKiz documentation </a></li> <li><a href="https://qualikiz.com/QuaLiKiz/Input-and-output-variables">QuaLiKiz/QLKNN input/output variables naming scheme </a></li> <li><a href="https://gitlab.com/Karel-van-de-Plassche/QLKNN-develop">Training, plotting, filtering, and auxiliary tools </a></li> <li><a href="https://gitlab.com/qualikiz-group/QuaLiKiz-pythontools">QuaLiKiz related tools </a></li> <li><a href="https://gitlab.com/qualikiz-group/QLKNN-fortran">FORTRAN QLKNN implementation with wrapper for Python and MATLAB </a></li> <li><a href="https://gitlab.com/qualikiz-group/qlknn-hyper">Weights and biases of 'hyperrectangle style' QLKNN </a></li> </ul> <p> </p>
Dataset: Shell Commands Used by Participants of Hands-on Cybersecurity Training
<p>This repository contains supplementary materials for the following journal paper:</p> <p>Valdemar Švábenský, Jan Vykopal, Pavel Seda, Pavel Čeleda.<br> <em>Dataset of Shell Commands Used by Participants of Hands-on Cybersecurity Training.</em><br> In Elsevier Data in Brief. 2021.<br> <a href="https://doi.org/10.1016/j.dib.2021.107398">https://doi.org/10.1016/j.dib.2021.107398</a></p> <ul> </ul> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <pre><code>@article{Svabensky2021dataset, author = {\v{S}v\'{a}bensk\'{y}, Valdemar and Vykopal, Jan and Seda, Pavel and \v{C}eleda, Pavel}, title = {{Dataset of Shell Commands Used by Participants of Hands-on Cybersecurity Training}}, journal = {Data in Brief}, publisher = {Elsevier}, volume = {38}, year = {2021}, issn = {2352-3409}, url = {https://doi.org/10.1016/j.dib.2021.107398}, doi = {10.1016/j.dib.2021.107398}, }</code></pre> <p>The data were collected using a logging toolset referenced <a href="https://zenodo.org/record/5126693">here</a>.</p> <p><strong>Attached content</strong></p> <ol> <li><strong>Dataset (data.zip).</strong> The collected data are attached here on Zenodo. A copy is also available in <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/datasets/commands">this repository</a>.</li> <li><strong>Analytical tools (toolset.zip).</strong> To analyze the data, you can instantiate the toolset or <a href="https://gitlab.ics.muni.cz/muni-kypo/tools/commands-elk">this project for ELK</a>.</li> </ol> <p><strong>Version history</strong></p> <ul> <li>Version 1 (<a href="https://zenodo.org/record/5137355">https://zenodo.org/record/5137355</a>) contains 13446 log records from 175 trainees. These data are precisely those that are described in the associated journal paper. Version 1 provides a snapshot of the state when the article was published.</li> <li>Version 2 (<a href="https://zenodo.org/record/5517479">https://zenodo.org/record/5517479</a>) contains 13446 log records from 175 trainees. The data are unchanged from Version 1, but the analytical toolset includes a minor fix.</li> <li>Version 3 (<a href="https://zenodo.org/record/6670113">https://zenodo.org/record/6670113</a>) contains 21762 log records from 275 trainees. It is a superset of Version 2, with newly collected data added to the dataset.</li> <li>The current Version 4 (<a href="https://zenodo.org/record/8136017">https://zenodo.org/record/8136017</a>) contains 21459 log records from 275 trainees. Compared to Version 3, we cleaned 303 invalid/duplicate command records.</li> </ul>
Data deposit accompanying Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields
<p>Dataset accompanying the paper: <em>"Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields"</em>. Contains the training sets curated during active learning as well as .xyz files used for creating the Figures. <br> <br> The paper highlights that the computational efficiency of ML force fields not only results in decreased computational costs for routine catalytic investigations but also facilitates more comprehensive exploration of catalytic pathways.</p> <p><strong>Published in NPJ Computational Materials</strong>: <a href="https://www.nature.com/articles/s41524-023-01124-2">https://www.nature.com/articles/s41524-023-01124-2</a><br> Formerly on Arxiv: <a href="https://arxiv.org/abs/2301.09931">https://arxiv.org/abs/2301.09931</a></p>
Datasets and trained diffusion models for "Diffusion Models for Interferometric Satellite Aperture Radar"
<p>A set of trained Probabilistic Diffusion Models (PDMs) and corresponding training datasets for the paper "<a href="https://doi.org/10.48550/arXiv.2308.16847">Diffusion Models for Interferometric Satellite Aperture Radar</a>", by Tuel, Kerdreux et al. The code for this paper can be found at <a href="https://github.com/thomaskerdreux/PDM_SAR_InSAR_generation">this link</a>.</p> <p><strong>Training datasets</strong></p> <p>- "InSAR_noise_32x32.zip": a dataset of 32x32 ground deformation scenes obtained from InSAR interferograms over New Mexico with the small baseline subset (SBAS) algorithm. Images were normalised to [0, 1].</p> <p>- "insar_unwrapped_phase_normalised.zip": a dataset of 128x128 InSAR interferograms obtained from Sentinel-1 acquisitions over Nex Mexico. Images were normalised to [0, 1].</p> <p><strong>Trained Models</strong></p> <p>We provide 6 trained PDMs in separate .zip files. Each .zip file contains the model weights (in *.pt format) and the model metadata file (in *.json format).</p> <p>- "mnist_32_cond_sigma_100.zip": a class-conditional model trained with 100 diffusion time steps on 32x32 MNIST images;</p> <p>- "mnist_32_no_cond_sigma_100.zip": an unconditional model trained with 100 diffusion time steps on 32x32 MNIST images;</p> <p>- "SAR_lowres_128_cond_sigma_2000.zip": a low-resolution (256 to 128) model trained with 2000 diffusion time steps on 128x128 TenGeoP-SARwv images;</p> <p>- "SAR_superres_128_to_256_cond_sigma_2000.zip": a super-resolution (128 to 256) model trained with 2000 diffusion time steps on TenGeoP-SARwv images;</p> <p>- "insar_phase_128_sigma_2000.zip": an unconditional model trained with 2000 diffusion time steps on 128x128 Sentinel-1 InSAR interferograms over New Mexico;</p> <p>- "insar_noise_32_sigma_1000.zip": an unconditional model trained with 1000 diffusion time steps on 32x32 Sentinel-1 InSAR ground deformation scenes over New Mexico.</p>
RnR-ExM Training Dataset
<p>This dataset was released as part of the 2023 ISBI challenge, <a href="https://rnr-exm.grand-challenge.org">RnR-ExM</a>. The organizers thank Ruihan Zhang (MIT), Margaret Elizabeth Schroeder (MIT) and Chi Zhang (MIT) for contributing data to this competition.</p>
Forest fire assessement training dataset (2022-07-18 fire at Maclas - France)
<p>This dataset has been created to train Univ. Eiffel personnels on raster data handling with QGIS.</p><p>It provides the following elements:</p><ul><li>Geopackage database with the following layers:<ul><li>QGIS project</li><li>Extract from the SENTINEL-2 2022-06-11 B8A band</li><li>Extract from the SENTINEL-2 2022-06-11 B12 band</li><li>Extract from the SENTINEL-2 2022-07-21 B8A band</li><li>Extract from the SENTINEL-2 2022-07-21 B12 band</li><li>Reclassified delta NBR raster layer</li><li>Delta NBR vector layer</li><li>Studied area bounding box</li></ul></li><li>Intermediate results:<ul><li>pre-event NBR raster file</li><li>post-event NBR raster file</li><li>Delta NBR raster file</li><li>Delta NBR raster file multiplied by 1000 (for easier reclassification)</li></ul></li></ul><p>Data sources IDs from opensearch-theia.cnes.fr-sentinel2-l2a catalogue :</p><ul><li>SENTINEL2B_20220721-104826-811_L2A_T31TFL_D</li><li>SENTINEL2B_20220611-104824-395_L2A_T31TFL_D</li></ul><p> </p>
Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)
<p>The data used in this tutorial are a subset of the data published previously in <a href="https://zenodo.org/record/3243160">Training material for the course "Exome analysis with GALAXY"</a>. Credit for uploading the original data goes to Paolo Uva and Gianmauro Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p> </p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you'll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load and, in fact, without having to install GEMINI's bundled annotation data.</p>
Training Data for 'ewas_suite' Analysis
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes EWAS data from a study published by Hugo, Willy, et al., 2015 (DOI: <a href="https://doi.org/10.1016/j.cell.2015.07.061">10.1016/j.cell.2015.07.061</a>) to identify differentially methylated regions and positions associated with melanoma MAPKi resistance.</p>
Training dataset used in the magazine paper entitled "A Flexible Machine Learning-Aware Architecture for Future WLANs"
<p><a href="https://arxiv.org/pdf/1910.03510.pdf"><strong>A Flexible Machine Learning-Aware Architecture for Future WLANs</strong></a></p> <p><strong>Authors: </strong>Francesc Wilhelmi, Sergio Barrachina-Muñoz, Boris Bellalta, Cristina Cano, Anders Jonsson & Vishnu Ram.</p> <p><strong>Abstract: </strong>Lots of hopes have been placed in Machine Learning (ML) as a key enabler of future wireless networks. By taking advantage of the large volumes of data generated by networks, ML is expected to deal with the ever-increasing complexity of networking problems. Unfortunately, current networking systems are not yet prepared for supporting the ensuing requirements of ML-based applications, especially for enabling procedures related to data collection, processing, and output distribution. This article points out the architectural requirements that are needed to pervasively include ML as part of future wireless networks operation. To this aim, we propose to adopt the International Telecommunications Union (ITU) unified architecture for 5G and beyond. Specifically, we look into Wireless Local Area Networks (WLANs), which, due to their nature, can be found in multiple forms, ranging from cloud-based to edge-computing-like deployments. Based on ITU's architecture, we provide insights on the main requirements and the major challenges of introducing ML to the multiple modalities of WLANs.</p> <p><strong>Dataset description: </strong>This is the dataset generated for training a Neural Network (NN) in the Access Point (AP) (re)association problem in IEEE 802.11 Wireless Local Area Networks (WLANs). </p> <p>In particular, the NN is meant to output a prediction function of the throughput that a given station (STA) can obtain from a given Access Point (AP) after association. The features included in the dataset are:</p> <ol> <li>Identifier of the AP to which the STA has been associated.</li> <li>RSSI obtained from the AP to which the STA has been associated.</li> <li>Data rate in bits per second (bps) that the STA is allowed to use for the selected AP.</li> <li>Load in packets per second (pkt/s) that the STA generates.</li> <li>Percentage of data that the AP is able to serve before the user association is done.</li> <li>Amount of traffic load in pkt/s handled by the AP before the user association is done.</li> <li>Airtime in % that the AP enjoys before the user association is done.</li> <li>Throughput in pkt/s that the STA receives after the user association is done.</li> </ol> <p>The dataset has been generated through random simulations, based on the model provided in <a href="https://github.com/toniadame/WiFi_AP_Selection_Framework">https://github.com/toniadame/WiFi_AP_Selection_Framework</a>. More details regarding the dataset generation have been provided in <a href="https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans">https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans</a>.</p>
ZeroCostDL4Mic - CARE (2D) example training and test dataset
<p><strong>Name</strong>: ZeroCostDL4Mic - CARE (2D) example training and test dataset</p> <p>(see <a href="https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki">our Wiki</a> for details)</p> <p> </p> <p><strong>Data type</strong>: Paired microscopy images (fluorescence) of low and high signal-to-noise ratio</p> <p><strong>Microscopy data type</strong>: Fluorescence microscopy (Lifeact-RFP)</p> <p><strong>Microscope</strong>: Structured Illumination Microscopy (SIM) with a 60x 1.42 NA objective </p> <p><strong>Cell type</strong>: DCIS.COM Lifeact-RFP</p> <p><strong>File format</strong>: .tif (32-bit)</p> <p><strong>Image size</strong>: 1024x1024 (Pixel size: 40 nm)</p> <p> </p> <p><strong>Author(s)</strong>: Guillaume Jacquemet<sup>1,2</sup></p> <p><strong>Contact email</strong>: guillaume.jacquemet@abo.fi</p> <p><strong>Affiliation</strong>: </p> <p>1) Faculty of Science and Engineering, Cell Biology, Åbo Akademi University, 20520 Turku, Finland</p> <p>2) Turku Bioscience Centre, University of Turku and Åbo Akademi University, FI-20520 Turku, Finland</p> <p> </p> <p><strong>Associated publications</strong>: Unpublished</p> <p><strong>Funding bodies</strong>: G.J. was supported by grants awarded by the Academy of Finland, the Sigrid Juselius Foundation and Åbo Akademi University Research Foundation (CoE CellMech) and by Drug Discovery and Diagnostics strategic funding to Åbo Akademi University.</p>
ZeroCostDL4Mic - Noise2Void (2D) example training and test dataset
<p><strong>Name</strong>: ZeroCostDL4Mic - Noise2Void (2D) example training and test dataset</p> <p>(see <a href="https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki">our Wiki</a> for details)</p> <p> </p> <p><strong>Data type</strong>: Microscopy images (fluorescence)</p> <p><strong>Microscopy data type</strong>: Fluorescence microscopy (paxillin-GFP) </p> <p><strong>Microscope</strong>: Spinning disk confocal microscope with a 63x 1.4 NA objective </p> <p><strong>Cell type</strong>: U-251 glioma cells, endogenously expressing paxillin-GFP</p> <p><strong>File format</strong>: .tif (16-bit)</p> <p><strong>Image size</strong>: 512x512 (Pixel size: 248 nm)</p> <p> </p> <p><strong>Author(s)</strong>: Aki Stubb<sup>1</sup>, Guillaume Jacquemet<sup>1,2</sup> and Johanna Ivaska<sup>1</sup></p> <p><strong>Contact email</strong>: guillaume.jacquemet@abo.fi</p> <p><strong>Affiliation</strong>: </p> <p>1) Turku Bioscience Centre, University of Turku and Åbo Akademi University, FI-20520 Turku, Finland</p> <p>2) Faculty of Science and Engineering, Cell Biology, Åbo Akademi University, 20520 Turku, Finland</p> <p><br> </p> <p><strong>Associated publication</strong>: Stubb <em>et al.</em> 2020, Nano letters DOI: 10.1021/acs.nanolett.9b04083</p> <p><strong>Funding bodies</strong>: A.S. has been supported by the University of Turku Doctoral programme for Molecular Medicine (TuDMM).</p>
DCASE 2020 Challenge Task 2 Additional Training Dataset
<p><strong>Description</strong></p> <p>This dataset is the "additional training dataset" for the <strong>DCASE 2020 Challenge Task 2 "Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring" </strong><a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">[task description]</a>. </p> <p>In the task, three datasets have been or will be released: "<a href="http://zenodo.org/record/3678171">development dataset</a>", "additional training dataset", and "<a href="https://zenodo.org/record/3841772">evaluation dataset</a>". This additional training dataset was released before the "<a href="https://zenodo.org/record/3841772">evaluation dataset</a>". This dataset includes around 1,000 normal samples for each Machine Type and Machine ID used in the <a href="https://zenodo.org/record/3841772">evaluation dataset</a> and can be used for model training in advance.</p> <p>The recording procedure and data format are the same as the <a href="http://zenodo.org/record/3678171">development dataset</a>. The Machine IDs in this dataset are different from those in the <a href="http://zenodo.org/record/3678171">development dataset</a>. For more information, please see the pages of the <a href="http://zenodo.org/record/3678171">development dataset</a> and the <a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">task description</a>. </p> <p> </p> <p><strong>Directory structure</strong></p> <p>Once you unzip the downloaded files from Zenodo, you can see the following directory structure. Machine Type information is given by directory name, and Machine ID and condition information are given by file name, as:</p> <ul> </ul> <p>/eval_data</p> <ul> <li>/ToyCar <ul> <li>/train (Only normal data for all Machine IDs are included.) <ul> <li>/normal_id_05_00000000.wav</li> <li>...</li> <li>/normal_id_05_00000999.wav</li> <li>/normal_id_06_00000000.wav</li> <li>...</li> <li>/normal_id_07_00000999.wav</li> </ul> </li> </ul> </li> <li>/ToyConveyor (The other Machine Types have the same directory structure as ToyCar.)</li> <li>/fan</li> <li>/pump</li> <li>/slider</li> <li>/valve</li> </ul> <p> </p> <p>The paths of audio files are:</p> <ul> <li>"/eval_data/<Machine_Type>/train/normal_id_<Machine_ID>_[0-9]+.wav"</li> </ul> <p>For example, the Machine Type and Machine ID of "/ToyCar/train/normal_id_05_00000000.wav" are "ToyCar" and "05", respectively, and its condition is normal (This dataset includes only normal samples). </p> <p> </p> <p><strong>Baseline system</strong></p> <p>A simple baseline system is available on the Github repository <a href="https://github.com/y-kawagu/dcase2020_task2_baseline">[URL]</a>. The baseline system provides a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. It is a good starting point, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p> </p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>NTT Corporation</strong> and <strong>Hitachi, Ltd.</strong> and is available under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p> </p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three papers</strong>:</p> <p>Yuma Koizumi, Shoichiro Saito, Noboru Harada, Hisashi Uematsu, and Keisuke Imoto, "ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection," in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019. <a href="https://ieeexplore.ieee.org/document/8937164">[pdf]</a></p> <p>Harsh Purohit, Ryo Tanabe, Kenji Ichige, Takashi Endo, Yuki Nikaido, Kaori Suefusa, and Yohei Kawaguchi, “MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection,” in Proc. 4th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019. <a href="http://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf">[pdf]</a></p> <p>Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto, Toshiki Nakamura, Yuki Nikaido, Ryo Tanabe, Harsh Purohit, Kaori Suefusa, Takashi Endo, Masahiro Yasuda, and Noboru Harada, "Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring<em>,"</em> in Proc. 5th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2020. <a href="https://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Koizumi_3.pdf">[pdf]</a></p> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yuma Koizumi, <a href="mailto:koizumi.yuma@ieee.org">koizumi.yuma@ieee.org</a></li> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.