Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
120
datasets available to search
ShareScore release 0.7.1
Dataset results
120 results for “data augmentation”
Data augmentation for Wasm
<p>We used wasm-mutate to augment the original dataset of MINOS. We partially recreated the dataset from MINOS, and in this record, it is composed by 134 benign WebAssembly programs and 36 malign programs. After gathering the binaries we pass each program to wasm-mutate generating up to 1k variants. For the programs and its variants, we generate their grayscale representation and that is what is saved in the datasets.</p>
Variable Misuse tool: Dataset for data augmentation
<p>Dataset used for data augmentation in the training phase of the Variable Misuse tool. It contains some source code files extracted from third-party repositories.</p>
Variable Misuse tool: Dataset for data augmentation (5)
<p>Dataset used for data augmentation in the training phase of the Variable Misuse tool. It contains some source code files extracted from third-party repositories.</p>
LOFAs data of NEREA Augmented Observatory
<p>LOFAs samples were collected at the selected depth using Niskin bottle, pre-filtered onto 200 μm mesh nylon net, filtered onto polycarbonate filters (2 μm mesh size) and then frozen in liquid nitrogen. Filters were sonicated in water, incubated for 30 min at room temperature and then LOFAs extracted and quantified through liquid chromatography-mass spectrometry (LC-MS) using an internal standard. LOFAs amount are normalized by the volume of the filtered seawater sample.</p>
Flow cytometry data of NEREA Augmented Observatory
<p>FCM sample were processed using a Becton-Dickinson FACSVerse flow cytometer, equipped with standard laser (488 nm) and filter set. Volumes were estimated using TruCount Beads run before every batch run or estimated from the volumetric count of the instrument.</p> <p> </p>
Biogeochemical data of NEREA Augmented Observatory
<p>Analyses are performed using a Flow-Sys III Systea auto-analyzer using the methods reported by Hansen and Grasshoff (1983), slightly modified, for inorganic nutrients and after digestion at 120 °C for 30 minutes, using the adjustments proposed by Valderrama, 1981 to the original methods by Koroleff (1976 a, b). DOC analysis is carried out by high-temperature catalytic oxidation with a Shimadzu TOC analyzer (TOC-Vcsn) following the methods described in Santinelli 2015. POC analysis is conducted using a variable volume (1-3 litres) of seawater filtered on GF/F filters. After filtration, filters are rinsed with deionized water and stored at -20 °C. The analyses are performed with a Thermo Scientific FlashEA 1112 elemental analyzer (Thermo Fisher Scientific) following (Hedges and Stern 1984) and using cyclohexanone-2,4-dinitrophenylhydrazone as standard.</p>
Assemblies of metagenomic data of NEREA Augmented Observatory
<p>The NEREA_metaG_assemblies project contains key initial analyses of NEREA metagenomic data that are stored in the project NEREA_metaG.</p> <p><strong>Primary_analysis</strong>: This directory consists of key initial analyses including the assembly of metagenomic data and subsequent gene prediction. </p>
Prokaryotic gene catalog, prokaryotic Metagenome-Assembled Genomes (MAGs) and taxonomic profiling of metagenomic data of NEREA Augmented Observatory
<p>The NEREA_metaG directory is dedicated to the in-depth analysis of NEREA microbial communities using metagenomic sequencing data. </p> <p><strong>Gene catalog:</strong> This directory contains the gene catalog compiled from metagenomic data, which includes: Protein and nucleotide sequence files for genes; Cluster files grouping similar genes; Annotation files mapping genes to KEGG pathways; Normalized gene abundance profiles.</p> <div><strong>MAGs:</strong> Directory for Metagenome-Assembled Genomes (MAGs). It contains comprehensive annotation files for the MAGs, providing insights into gene functions, metabolic pathways, and other genomic features. It also contains the individual MAGs categorized by sample origin. Each MAG is stored in a compressed FASTA format.</div> <p><strong>mOTUs</strong>: Contains files related to microbial taxonomic units identified and quantified using the mOTUs profiler. </p>
Phytoplankton data of NEREA Augmented Observatory
<p>Samples were collected at the selected depth using Niskin bottles and immediately fixed with Lugol's solution (1%). Quantitative analysis is carried out using an inverted microscope after the settling of a variable volume of sample (Utermhöl, 1958). Volumes vary in relation to the number of cells present in the sample, which is estimated on the basis of the chlorophyll a concentration or of the Secchi Disk depth. Counting is done over varying fractions of the sedimentation chamber (transects, random fields, or entire chamber), depending on the characteristics of the sample.</p>
HPLC data of NEREA Augmented Observatory
<p>The water used for the determination of the pigment spectrum by HPLC accounts for 2-3 litres for each depth. The seawater is filtered on GF/F filters (diameter 47mm). The filters are stored in liquid nitrogen until pigments HPLC analysis. HPLC pigments separations were made on an Agilent 1100 HPLC (Agilent technologies, United States). The system was equipped with an HP 1050 photodiode array detector and a HP 1046A fluorescence detector for the determination of chlorophyll degradation products. Instrument calibration was carried out with external standard pigments provided by the International Agency for 14C determination-VKI Water Quality Institute. </p>
Microzooplankton data of NEREA Augmented Observatory
<p>Samples were collected at the selected depth using Niskin bottles and immediately fixed with Lugol's solution (1%). Quantitative analysis is carried out using an inverted microscope after the settling of a variable volume of sample (Utermhöl, 1958). Volumes vary in relation to the number of cells present in the sample. At first, 100mL are settled form each sample and cells enumerated. The volume of sample is considered sufficient if it is possible to count at least 100 cells on a transect. Otherwise additional volume is settled on top of the same chamber up to a maximum of 250mL. Counting is done over the whole sedimentation chamber or half of it, depending on the characteristics of the sample.</p>
Mesozooplankton data of NEREA Augmented Observatory
<p>Mesozooplankton was collected with double WP2, which has a mouth area of 0.25 m2 and mesh aperture width of 200 μm. The net, ballasted with a 3 kg weight, was towed vertically to the surface at low speed (0.7-1.0 m s-1). </p>
Dataset: Data augmentation experiments with style-based quantum generative adversarial networks on trapped-ion and superconducting-qubit technologies
<p>Dataset for the following paper: <a href="https://arxiv.org/abs/2405.04401">"Data augmentation experiments with style-based quantum generative adversarial networks on trapped-ion and superconducting-qubit technologies", Julien Baglio, arXiv:2405.04401</a></p> <p>It contains:</p> <ul> <li>one folder named "data_for_all_plots" containing the raw data for the s, t, and y distributions for all the figures of the paper as well as a Jupyter notebook to generate the figures.</li> <li>one file named "variance_calculations_qGAN.txt" containing the data to calculate the errors for the KL divergences.</li> </ul>
Stable isotope ratios data of NEREA Augmented Observatory
<p>Samples were collected at 1m depth or along the whole water column using Niskin bottles and plankton nets, transferred into bins or jars and transported to the laboratory, where the individual size fractions were isolated. For each size fraction, samples were frozen and freeze-dried individually. Carbon and nitrogen stable isotope values (‰) were determined at the Stable Isotope Facility at the University of california in Davis (USA).</p>
16S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory
<p><strong>Metabarcoding - 16S ASV generation and taxonomic assignment</strong>. Raw 16S paired-end sequences were subjected to a data quality control step and subsequently imported into the QIIME2 pipeline v.2022.2.0. Leftover primers and adapter sequences were removed through cutadapt. The amplicon sequence variants (ASV) table, which represent true biological sequences within each sample, was generated using the denoised-paired method including truncation, denoising, dereplication, merging, and chimera filtering of the DADA2 (Divisive Amplicon Denoising Algorithm 2) plugin inside QIIME2. Default parameters were used with the exception of the forward and reverse sequence length (--p-trunc-len-f and --p-trunc-len-r), that were set to 220 and 180, respectively. Processed reads that passed all these filters were used for taxonomy classification. The V4-V5 regions were extracted from the pre-formatted reference sequences and taxonomy file built on the SILVA 138 99% OTUs database and the vsearch v.2.6.2 global alignment implemented in QIIME2 was used.</p>
Data for the study "'Look at the Trees': A Verbal Nudge to Reduce Screen Time When Learning Biodiversity with Augmented Reality"
Open the record for dataset details and reuse information.
FreeHi-C simulates high fidelity Hi-C data for benchmarking and data augmentation (.hic files)
<p>FreeHi-C simulated replicates for GM12878 and A549 on chromosome 1, and whole-genome of P.falciparum using different sequencing depth. Data are save in .hic format for direct visualization on Juicebox.</p>
Augmented-BDD: Diverse Adverse Weather and Lighting Conditions through Data Augmentations
<h1>The Augmented-BDD (A-BDD) dataset</h1>
Class-specific data augmentation for plant stress classification
<p>This is a companion dataset for the paper titled "<strong>Class-specific data augmentation for plant stress classification</strong>" by Nasla Saleem, Aditya Balu, Talukder Zaki Jubery, Arti Singh, Asheesh K. Singh, Soumik Sarkar, and Baskar Ganapathysubramanian published in <em>The Plant Phenome Journal</em>, https://doi.org/10.1002/ppj2.20112<br><br><br><strong>Abstract:</strong><br><br>Data augmentation is a powerful tool for improving deep learning-based image classifiers for plant stress identification and classification. However, selecting an effective set of augmentations from a large pool of candidates remains a key challenge, particularly in imbalanced and confounding datasets. We propose an approach for automated class-specific data augmentation using a genetic algorithm. We demonstrate the utility of our approach on soybean [<em>Glycine max (L.) Merr</em>] stress classification where symptoms are observed on leaves; a particularly challenging problem due to confounding classes in the dataset. Our approach yields substantial performance, achieving a mean-per-class accuracy of 97.61% and an overall accuracy of 98% on the soybean leaf stress dataset. Our method significantly improves the accuracy of the most challenging classes, with notable enhancements from 83.01% to 88.89% and from 85.71% to 94.05%, respectively. A key observation we make in this study is that high-performing augmentation strategies can be identified in a computationally efficient manner. We fine-tune only the linear layer of the baseline model with different augmentations, thereby reducing the computational burden associated with training classifiers from scratch for each augmentation policy while achieving exceptional performance. This research represents an advancement in automated data augmentation strategies for plant stress classification, particularly in the context of confounding datasets. Our findings contribute to the growing body of research in tailored augmentation techniques and their potential impact on disease management strategies, crop yields, and global food security. The proposed approach holds the potential to enhance the accuracy and efficiency of deep learning-based tools for managing plant stresses in agriculture.</p>
Image-based automated species identification: Can virtual data augmentation overcome problems of insufficient sampling?
<p></p><p>Automated species identification and delimitation is challenging, particularly in rare and thus often scarcely sampled species, which do not allow sufficient discrimination of infraspecific versus interspecific variation. Typical problems arising from either low or exaggerated interspecific morphological differentiation are best met by automated methods of machine learning that learn efficient and effective species identification from training samples. However, limited infraspecific sampling remains a key challenge also in machine learning.</p> <p>In this study, we assessed whether a data augmentation approach may help to overcome the problem of scarce training data in automated visual species identification. The stepwise augmentation of data comprised image rotation as well as visual and virtual augmentation. The visual data augmentation applies classic approaches of data augmentation and generation of artificial images using a Generative Adversarial Networks (GAN) approach. Descriptive feature vectors are derived from bottleneck features of a VGG-16 convolutional neural network (CNN) that are then stepwise reduced in dimensionality using Global Average Pooling and PCA to prevent overfitting. Finally, data augmentation employs synthetic additional sampling in feature space by an oversampling algorithm in vector space (SMOTE). Applied on four different image datasets, which include scarab beetle genitalia (Pleophylla, Schizonycha) as well as wing patterns of bees (Osmia) and cattleheart butterflies (Parides), our augmentation approach outperformed a deep learning baseline approach by means of resulting identification accuracy with non-augmented data as well as a traditional 2D morphometric approach (Procrustes analysis of scarab beetle genitalia).</p><p></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.