Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
120
datasets available to search
ShareScore release 0.7.1
Dataset results
120 results for “data augmentation”
Data from: Intestinal Ralstonia pickettii augments glucose intolerance in obesity
Open the record for dataset details and reuse information.
Data from: Outcomes in patients with infections and augmented renal clearance: a multicenter retrospective study
Recently, augmented renal clearance (ARC), which accelerates glomerular filtration of renally eliminated drugs thereby reducing the systemic exposure to these drugs, has started to receive attention. However, the clinical features associated with ARC are still not well understood, especially in the Japanese population. This study aimed to evaluate the clinical characteristics and outcomes of ARC patients with infections in Japanese intensive care unit (ICU) settings. We conducted a retrospective observational study from April 2013 to May 2017 at two tertiary level ICUs in Japan, which included 280 patients with infections (median age 74 years; interquartile range, 64–83 years). We evaluated the estimated glomerular filtration rate (eGFR) at ICU admission using the Japanese equation, and ARC was defined as eGFR >130 mL/min/1.73 m2. Multivariable logistic regression analysis was performed to identify the independent risk factors for ARC and to determine if it was a predictor of ICU mortality. In addition, a receiver operating curve (ROC) analysis was performed, and the area under the ROC (AUROC) was determined to examine the significant variables that predict ARC. In total, 19 patients (6.8%) manifested ARC. Multivariable logistic regression analysis identified younger age as an independent risk factor for ARC (odds ratio [OR], 0.94; 95% confidence interval [CI], 0.91–0.96). However, ARC was not found to be a predictor of ICU mortality (OR, 0.57; 95% CI, 0.11–2.92). In addition, the AUROC of age was 0.79 (95% CI, 0.68–0.91), and the optimal cut off age for ARC was ≤63 years (sensitivity, 68.4%; specificity, 78.9%). The incidence of ARC was, therefore, low among patients with infections in the Japanese ICUs. Although younger age was associated with the incidence of ARC, it was not an independent predictor of ICU mortality.
Augment Single-cell RNA-seq data with Surface Protein Levels using Gene set-based Deep Learning and Transfer Learning Methods
<p><span>Necessary data, scripts and saved models for "Augment Single-cell RNA-seq data with Surface Protein Levels using Gene set-based Deep Learning and Transfer Learning Methods" manuscript.</span></p>
18S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory
<p>Illumina paired-end V9-18S raw reads (FASTQ format, 2 X 150 PE) were pre-processed with cutadapt and vsearch to remove primer sequences, trim low quality bases and unify mixed orientation reads produced in the ligation-based library preparation; the procedure was implemented in a custom bash script. Processed reads were then used to generate amplicon sequence variants (ASVs) using the DADA2 R library; the pipeline was adapted from the one described on the program website (<span><span><a href="https://benjjneb.github.io/dada2/tutorial.html" target="_blank" rel="noopener">https://benjjneb.github.io/dada2/tutorial.html</a></span></span>); no further quality filtering was implemented at this stage, except for discarding all reads with ambiguities (parameter maxN = 0 of function filterAndTrim). Filtered forward (F) and reverse (R) reads were used to train the error model and then denoised by applying the trained error model to generate ASVs. Finally, F and R reads were merged and checked for chimeras; allowing no mismatches in read merging (default parameter maxMismatch = 0 of function mergePairs). ASVs were then classified with BLAST against the PR2 v5.01 reference database, integrated with 1,293 sequences from GoN protist strains and fungi environmental sequences. Highest bit score matching with the best taxonomic resolution were then selected among the returned results.</p>
Data augmentation with Generative AI for DoW attack detection in serverless architectures
<p>Serverless computing is one of the latest paradigms in cloud computing. It offers a framework for the development of event-driven applications whose functions are executed in a scalable environment provided by the corresponding cloud platform. In this way, resources are obtained on demand, paying only for the time the function is running. This new model has new vulnerabilities and, therefore, new types of cybersecurity attacks. However, there are still not enough transaction datasets for serverless systems with a sufficient amount of data to develop advanced detection methods for this type of threat. Therefore, we present this dataset that has been built with generative AI to advance the development of models that can effectively deal with these threats.</p>
Enhancing melanoma skin cancer classification through data augmentation
Open the record for dataset details and reuse information.
Medical Data Collection for the Evaluation of Radiofrequency Ablation and Cement Augmentation for the Treatment of Secondary Metastases to the Spine
ClinicalTrials.gov study NCT04751422. IPD Sharing: Not stated. Countries: 1. Publications: 0.
PMCF Study on the Safety, Performance and Clinical Benefits Data of the NexGen TM Augmentation Patella
ClinicalTrials.gov study NCT05253976. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Design of an Augmented Reality System by Integration of CT Scan or MRI Data With Endoscopic Images for Video-assisted Endonasal Endoscopic Surgery
ClinicalTrials.gov study NCT04968561. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Data from: Outcomes in patients with infections and augmented renal clearance: a multicenter retrospective study
Open the record for dataset details and reuse information.
Data from: Partially incorrect fossil data augment analyses of discrete trait evolution in living species
Open the record for dataset details and reuse information.
Ground-Based Global Navigation Satellite System (GNSS) Satellite-Based Augmentation System (SBAS) Broadcast Ephemeris Data (30-second sampling, hourly files) from NASA CDDIS
This dataset consists of ground-based Global Navigation Satellite System (GNSS) Satellite-Based Augmentation System (SBAS) Broadcast Ephemeris Data (hourly files) from the NASA Crustal Dynamics Data Information System (CDDIS). GNSS provide autonomous geo-spatial positioning with global coverage. GNSS data sets from ground receivers at the CDDIS consist primarily of the data from the U.S. Global Positioning System (GPS) and the Russian GLONASS. Since 2011, the CDDIS GNSS archive includes data from other GNSS (Europe’s Galileo, China’s Beidou, Japan’s Quasi-Zenith Satellite System/QZSS, the Indian Regional Navigation Satellite System/IRNSS, and worldwide Satellite Based Augmentation Systems/SBASs), which are similar to the U.S. GPS in terms of the satellite constellation, orbits, and signal structure. The hourly SBAS broadcast ephemeris files contain one day of SBAS broadcast navigation data in RINEX format from a global permanent network of ground-based receivers, one file per site. More information about these data is available on the CDDIS website at https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/hourly_30second_data.html.
Ground-Based Global Navigation Satellite System (GNSS) Satellite-Based Augmentation System (SBAS) Broadcast Ephemeris Data (sub-hourly files) from NASA CDDIS
This dataset consists of ground-based Global Navigation Satellite System (GNSS) Satellite-Based Augmentation System (SBAS) Broadcast Ephemeris Data (sub-hourly files) from the NASA Crustal Dynamics Data Information System (CDDIS). GNSS provide autonomous geo-spatial positioning with global coverage. GNSS data sets from ground receivers at the CDDIS consist primarily of the data from the U.S. Global Positioning System (GPS) and the Russian GLONASS. Since 2011, the CDDIS GNSS archive includes data from other GNSS (Europe’s Galileo, China’s Beidou, Japan’s Quasi-Zenith Satellite System/QZSS, the Indian Regional Navigation Satellite System/IRNSS, and worldwide Satellite Based Augmentation Systems/SBASs), which are similar to the U.S. GPS in terms of the satellite constellation, orbits, and signal structure. The sub-hourly SBAS broadcast ephemeris files contain 15 minutes of SBAS broadcast navigation data in RINEX format from a global permanent network of ground-based receivers, one file per 15 minutes per site. More information about these data is available on the CDDIS website at https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/high-rate_data.html.
Ground-Based Global Navigation Satellite System (GNSS) Satellite-Based Augmentation System (SBAS) Broadcast Ephemeris Data (daily files) from NASA CDDIS
This dataset consists of ground-based Global Navigation Satellite System (GNSS) Satellite-Based Augmentation System (SBAS) Broadcast Ephemeris Data (daily files) from the NASA Crustal Dynamics Data Information System (CDDIS). GNSS provide autonomous geo-spatial positioning with global coverage. GNSS data sets from ground receivers at the CDDIS consist primarily of the data from the U.S. Global Positioning System (GPS) and the Russian GLONASS. Since 2011, the CDDIS GNSS archive includes data from other GNSS (Europe’s Galileo, China’s Beidou, Japan’s Quasi-Zenith Satellite System/QZSS, the Indian Regional Navigation Satellite System/IRNSS, and worldwide Satellite Based Augmentation Systems/SBASs), which are similar to the U.S. GPS in terms of the satellite constellation, orbits, and signal structure. The daily SBAS broadcast ephemeris files contain one day of SBAS broadcast navigation data in RINEX format from a global permanent network of ground-based receivers, one file per site. More information about these data is available on the CDDIS website at https://cddis.nasa.gov/Data_and_Derived_Products/GNSS/daily_30second_data.html.
Source data for "Fast multi-source nanophotonic simulations using augmented partial factorization"
<p>Source data for Fig. 3b-c and Fig. 5a-b</p>
Data augmentation
<p>Data augmentation</p>
Data Augmentation Dataset
Open the record for dataset details and reuse information.
Augmenting Existing Normative Data for QbMobile in Healthy Children Aged 6-11 Years
ClinicalTrials.gov study NCT07330960. IPD Sharing: NO. Countries: 0. Publications: 0.
16S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory
<p>Cleaned raw 16S paired-end sequences were imported into the QIIME2 pipeline v. 2022.2.0. Leftover primers and adapters’ sequences were removed through cutadapt. The amplicon sequence variants (ASV) table, which represent true biological sequences within each sample, was generated using the denoised-paired method including truncation, denoising, dereplication, and chimera filtering of the DADA2 (Divisive Amplicon Denoising Algorithm 2) plugin inside QIIME2. Default parameters were used with the exception of the forward and reverse sequence length (--p-trunc-len-f and --p-trunc-len-r), that were set to 220 and 180, respectively. For taxonomy classification, the V4-V5 region were extracted from the pre-formatted reference sequences and taxonomy file build on the SILVA 138 99% OTUS database and the vsearch v. 2.6.2 global alignment implemented in QIIME2 was used.</p>
18S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory
<p>The quality of the Illumina paired-end V9-18S raw reads (FASTQ format, 2 X 150 PE) was checked using vsearch (vsearch --fastq_stats), then pre-processed with cutadapt and vsearch to remove primer sequences, trim low quality bases and unify mixed orientation reads produced in the ligation-based library preparation. Processed reads were then used to generate amplicon sequence variants (ASVs) using the DADA2 R library; the pipeline was adapted from the one described on the program website (https://benjjneb.github.io/dada2/tutorial.html); no further quality filtering was implemented at this stage, except for discarding all reads with ambiguities (parameter maxN=0 of function filterAndTrim). Filtered F and R reads were used to train the error model and then denoised by applying the trained error model to generate ASVs. Finally, F and R reads were merged and checked for chimeras; up to 9 mismatches were allowed for read merging (parameter maxMismatch=9 of function mergePairs). ASVs were then classified with BLAST against the PR2 v5.01 reference database, integrated with 1293 sequences from Gulf of Naples protist strains and fungi environmental sequences. Highest bit score matches with the best taxonomic resolution were then selected among the returned results.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.