Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
120
datasets available to search
ShareScore release 0.7.1
Dataset results
120 results for “data augmentation”
Generalizable cone beam CT esophagus segmentation using physics-based data augmentation
<p>This upload contains open source AAPM thoracic auto-segmentation data (http://aapmchallenges.cloudapp.net/competitions/3) augmented with physics-based data augmentation technique introduced in this paper:https://iopscience.iop.org/article/10.1088/1361-6560/abe2eb .</p> <p>The original data contained thoracic planning CT along with organs-at-risk segmentation masks for Esophagus, heart, lungs and spinal cord. The physics-based augmentation pipeline was used to convert planning CT images to pseudoCBCT (psCBCT) images which are routinely used in weekly radiotherapy treatment sessions for cancer patients. Further geometric data augmentations are also applied to convert one planning CT/OAR dataset into 23 perfectly paired CT/psCBCT/OAR pairs.</p> <p>This dataset has been used to train deep learning models for organs-at-risk segmentation from psCBCT images and multitask simultaneous CBCT to CT translation and segmentation tasks.</p>
[Data S2] Microbiota dictate T cell clonal selection to augment graft-vs-host disease after stem cell transplantation
<p><strong>Data</strong> <strong>S2. GLIPH2 hits for all recipient pairs amongst B6 to B6D2F1 transplants with or without antibiotic exposure and 900cGy vs. 1300cGy TBI conditioning. </strong>For all TCRs within each recipient pair (12 mice total, 66 pairs for spleen; 15 pairs for SILP analysis), GLIPH2 specificity groups are generated as described in methods.</p>
[Data S1] Microbiota dictate T cell clonal selection to augment graft-vs-host disease after stem cell transplantation
<p><strong>Data S1. ALICE hits for all recipient pairs amongst B6 to B6D2F1 transplants with or without antibiotic exposure and 900cGy vs. 1300cGy TBI conditioning. </strong>For each recipient pair (12 mice total, 66 pairs for spleen; 15 pairs for SILP analysis), all ALICE hits generated as described in methods.</p>
Data for manuscript "Adaptive Ensemble Refinement of Protein Structures in High Resolution Electron Microscopy Density Maps with Radical Augmented Molecular Dynamics Flexible Fitting"
<p>The tar file contains the input files for RADICAL augmented MDFF implementation (R-MDFF) for two protein systems, Adenylate Kinase (ADK) and Carbon Monoxide Dehydrogenase (CODH). These examples demonstrate the implementation of R-MDFF using RADICAL-Cybertools to flexibly fit biomolecules in cryo-EM density maps with on-the-fly decision making.</p> <p>All molecular simulations were performed using CUDA enabled NAMD 2.14 installed on OLCF Summit HPC resource. The CHARMM36 force field parameters were used for the proteins. Synthetic density maps were prepared at 1.8, 3 and 5 Å for ADK and 1.8 and 3 Å for CODH using VMD 1.9.3 software installed on OLCF Summit HPC resource. During the analysis stage, the cross correlation coefficients between density maps and atomic model were computed using VMD 1.9.3 on Summit HPC as part of the R-MDFF workflow.</p> <p>The source code is publicly available on GitHub: <a href="https://github.com/radical-collaboration/MDFF-EnTK">https://github.com/radical-collaboration/MDFF-EnTK </a></p> <p>The preprint of this research is submitted on bioRxiv, doi: <a href="https://doi.org/10.1101/2021.12.07.471672">https://doi.org/10.1101/2021.12.07.471672 </a></p> <p>To obtain maximum compression of the data, the tar command used to generate this tarball was:</p> <pre><code class="language-bash">GZIP=-9 tar --exclude='last.pdb' --exclude='*last_from_prev_iter.pdb' --exclude='*old' --exclude='*log' --exclude='*coor' --exclude='*vel' --exclude='*xsc' --exclude='*dcd' --exclude='lastframepdbs_fix' --exclude='*out' --exclude='*sl' --exclude='*rs' --exclude='*prof' --exclude='*err' --exclude='*dx' --exclude='*grid.pdb' --exclude='*txt' -cvzf rmdffv2.tar.gz rmdff-zenodo/</code></pre> <p> </p>
User Unfairness Mitigation by Graph Data Augmentation in Recommendation
<p>Dataset for the paper submission `User Unfairness Mitigation by Graph Data Augmentation in Recommendation`.</p> <p>The included datasets are: MovieLens-1M, Last.FM 1K.</p> <p> </p>
Assessing the Performance of Artificial Intelligence (AI)-Augmented Electronic Health Record (EHR) Data Abstraction for Clinical Trial Patient Screening
ClinicalTrials.gov study NCT06561217. IPD Sharing: NO. Countries: 1. Publications: 1.
Data From: Tactile Echoes: Multisensory Augmented Reality for the Hand
Open the record for dataset details and reuse information.
Data from: Evaluating the target‐tracking performance of scanning avian radars by augmenting data with simulated echoes
Open the record for dataset details and reuse information.
Data from: Trends in plant cover derived from vegetation-plot data using ordinal zero-augmented beta regression
Open the record for dataset details and reuse information.
Data from: In situ foliar augmentation of multiple species for optical phenotyping and bioengineering using soft robotics
Open the record for dataset details and reuse information.
Data from: Group augmentation, collective action, and territorial boundary patrols by male chimpanzees
Open the record for dataset details and reuse information.
Multispectral and augmented Landsat data with land cover labels
<p>Benchmark set at 77.1% O.A at: https://doi.org/10.1117/1.JRS.14.048503</p> <p>The dataset consists of 60,000 images, corresponding to Landsat patches of 33x33 pixels with 102 bands. Randomly selected from Mexico (country). Each patch is labeled with one of 12 Land Use and Vegetation classes according to the classification described at https://doi.org/10.3390/rs6053923.</p> <p>The zip file contains 12 folders numbered 1-12 and each contains 5,000 .npy python files (can be loaded with the NumPy library).</p> <p>The labeled classes correspond to the following identifier.</p> <p>1, Temperate Coniferous forest<br> 2, Temperate Decidius Forest<br> 3, Temperate Mixed Forest<br> 4, Tropical Evergreen Forest<br> 5, Tropical Deciduous Forest<br> 6, Scrubland<br> 7, Wetland Vegetation<br> 8, Agriculture<br> 9, Grassland<br> 10, Water body<br> 11, Barren Land<br> 12, Urban Area</p> <p>To build that dataset, we take the information of the National Continuum of Land Use and Vegetation series number 5 generated by the National Institute of Statistics and Geography from Mexico (INEGI) from The National Commission for the Knowledge and Use of Biodiversity (CONABIO) web page (http://geoportal.conabio.gob.mx/metadatos/doc/html/usv250s5ugw.html).</p> <p>The file used for this dataset construction is the shape format file with geographic coordinates located in http://www.conabio.gob.mx/informacion/gis/maps/geo/usv250s5ugw.zip.<br> Later, a transformation to Albers equal-area conic projection was done with the followings parameters:</p> <p>Fake east: 2500000.0<br> Fake North: 0.0<br> Origin longitude: -102.0º<br> Origin latitude: 12.0º<br> First standard parallel: 17.5º<br> Second standard parallel: 29.5º<br> Linear unit: Meter (1.0)<br> Reference ellipsoid: GRS80</p> <p><br> Once the data was projected, using the classes identified in the National Continuum of Land Use and Vegetation, correspondence was applied to the classes identified in https://doi.org/10.3390/rs6053923, these classes being: Agriculture, Barren land, Grassland, Scrubland, Temperate coniferous forest, Temperate deciduous forest, Temperate mixed forest, Tropical deciduous forest, Tropical evergreen forest, Urban area, Waterbody and Wetland vegetation.</p> <p>Once the information layer was generated with the 12 classes indicated above, the reference layer was rasterized.<br> Thus, a national grid of 1,975,940 regions of 1 x 1 kilometers was generated and the percentage of pixels of the dominant class in each corresponding 1 km region was associated.</p> <p>A total of cells with 70% or more pixels from one dominant class corresponds to 1,640,827 which represents a total of 83% of the Mexican territory. That means, only 17% of cells have less than 70% of their pixels from one dominant class.<br> Then, 5000 regions were randomly selected from each land cover class at the national level. For this random selection only were selected the regions in which cells have 70% or more of their pixels from one dominant class. The above, for looking to have consistent and reliable data for the automatic classification task. This random selection generates a total of 60,000 regions selected.</p> <p>Image patches were extracted from the selected regions in the sample.</p> <p>The image used is the result of the application of multiple time series analysis algorithms on a cube of image data with mainly Tier 1 (T1) quality and a few Tier 2 (T2) as described in https: // www. usgs.gov/land-resources/nli/landsat/landsat-collection-1. An Open Data Cube (ODC, https://www.opendatacube.org/) was constructed from 3,515 Landsat 5 and 7 images corresponding to the year 2011, which is the same reference year of the National Continuum of Land Use and Vegetation Series 5.</p> <p>From the analysis of the ODC images, the Geomedian (https://doi.org/10.1109/TGRS.2017.2723896) was calculated, which generated a national cloud-free mosaic from 2011, pixels at 30 meters resolution and 6 spectral bands (blue, green, red, nir, swir 1, swir 2). Finally, 15 spectral indices were calculated for each pixel in the image. This resulted in 15 national mosaics from the analysis of the time series of each pixel available for the year 2011 using all the combinations of normalized difference indices, which were possible with the 6 bands that were incorporated into the data cube, with which resulted in 102 information channels. Since Landsat images have a resolution of 30 meters, we have images of 33 pixels x 33 pixels for each region of 1 km x 1 km.</p> <p>The 102 channels in the patches correspond to:</p> <p>Geomedian Bands (6): blue, green, red, nir, swir 1, swir 2<br> Geomedian Based Indexes (15): evi, bu, sr, arvi, ui, ndbi, ibi, ndvi, ndwi, mndwi, nbi, brba, nbai, baei, bi<br> Geomedian Based Tasseled cap transformation (6): brightness, greenness, wetness, fourth, fifth, sixth</p> <p>2011 Landsat Time Analysis Series by Pixel</p> <p>(red-swir 1)/(red+swir 1); (5): min, mean, max, std, median<br> (red-nir)/( red+nir); (5): min, mean, max, std, median<br> (swir 1-swir 2)/( swir 1+swir 2); (5): min, mean, max, std, median<br> (nir-swir 2)/(nir+swir 2); (5): min, mean, max, std, median<br> (nir-swir 1)/( nir+swir 1); (5): min, mean, max, std, median<br> (red-swir 2)/( red+swir 2); (5): min, mean, max, std, median<br> (green-swir 2)/(green+swir 2); (5): min, mean, max, std, median<br> (green-swir 1)/(green+swir 1); (5): min, mean, max, std, median<br> (green-red)/(green+red); (5): min, mean, max, std, median<br> (green-nir)/(green+nir); (5): min, mean, max, std, median<br> (blue-swir 2)/(blue+swir 2); (5): min, mean, max, std, median<br> (blue-swir 1)/(blue+swir 1); (5): min, mean, max, std, median<br> (blue-red)/(blue+red); (5): min, mean, max, std, median<br> (blue-nir)/(blue+nir); (5): min, mean, max, std, median<br> (blue-green)/( blue+green); (5): min, mean, max, std, median</p>
Fast and accurate estimation of species-specific diversification rates using data augmentation
Diversification rates vary across species as a response to various factors, including environmental conditions and species-specific features. Phylogenetic models that allow accounting for and quantifying this heterogeneity in diversification rates have proven particularly useful for understanding clades diversification. Recently, we introduced the cladogenetic diversification rate shift model (ClaDS), which allows inferring subtle rate variations across lineages. Here we present a new inference technique for this model that considerably reduces computation time through the use of data augmentation and provide an implementation of this method in Julia. In addition to drastically reducing computation time, this new inference approach provides a posterior distribution of the augmented data, that is the tree with extinct and unsampled lineages as well as associated diversification rates. In particular, this allows extracting the distribution through time of both the mean rate and the number of lineages. We assess the statistical performances of our approach using simulations and illustrate its application on the entire bird radiation.
Augmentation Via Registration: AutoImplant 2020 Augmented Data Set
<p>Augmented data derived from the AutoImplant 2020 Challenge. Data has been augmented via non-linear SyN registrations using Advanced Normalization Tools (ANTs). This data was used to obtain first place in the AutoImplant 2020 challenge. The AutoImplant 2020 Challenge data was derived from the QC500 dataset from qure.ai. The original dataset is licensed under the CC BY-NC-SA 4.0 license (Attribution-NonCommercial-ShareAlike) and complies with the End User License Agreement (EULA) which are both detailed in the "LICENSE" file in this folder. Please refer to the "LICENSE" file for terms of use.</p> <p>If you use this data, please cite the following paper:</p> <p>Ellis D.G., Aizenberg M.R. (2020) Deep Learning Using Augmentation via Registration: 1st Place Solution to the AutoImplant 2020 Challenge. In: Li J., Egger J. (eds) Towards the Automatization of Cranial Implant Design in Cranioplasty. AutoImplant 2020. Lecture Notes in Computer Science, vol 12439. Springer, Cham. <a href="https://doi.org/10.1007/978-3-030-64327-0_6">https://doi.org/10.1007/978-3-030-64327-0_6</a></p>
Data from: Kin selection, not group augmentation, predicts helping in an obligate cooperatively breeding bird
Kin selection theory has been the central model for understanding the evolution of cooperative breeding, where non-breeders help bear the cost of rearing young. Recently the dominance of this idea has been questioned; particularly in obligate cooperative breeders where breeding without help is uncommon and seldom successful. In such systems, the direct benefits gained through augmenting current group size have been hypothesised to provide a tractable alternative (or addition) to kin selection. However, clear empirical tests of the opposing predictions are lacking. Here, we provide convincing evidence to suggest that kin selection and not group augmentation accounts for decisions of whether, where and how often to help in an obligate cooperative breeder, the chestnut-crowned babbler (Pomatostomus ruficeps). We found no evidence that group members base helping decisions on the size of breeding units available in their social group, despite both correlational and experimental data showing substantial variation in the degree to which helpers affect productivity in units of different size. By contrast, 98% of group members with kin present helped, 100% directed their care towards the most related brood in the social group and those rearing half/full-sibs helped approximately 3 times harder than those rearing less/non-related broods. We conclude that kin selection plays a central role in the maintenance of cooperative breeding in this species, despite the apparent importance of living in large groups.
Physical-oceanographic data (CTD and thermosalinograph) of NEREA Augmented Observatory
<p>CTD data are obtained using a SeaBird Electronic SBE 911 plus v2 multi-parameter profiler. This probe acquires 24 data variables per second with an accuracy of 0.002 °C for temperature and 0.001 S/m for conductivity, and it is bore by a ROSETTE-type sampler. </p> <p>The thermosalinograph was used at 0 m depth during the sampling campaigns, allowing to collect continuous data for the superficial temperature, conductivity and salinity. </p>
Code and Data from: An Imputation-Based Approach for Augmenting Sparse Epidemiological Signals
<p>This directory contains R code and required data to run the full data augmentation described in, "An Imputation-Based Approach for Augmenting Sparse Epidemiological Signals." This is the updated code corresponding to the updated medRxiv manuscript. It now includes ILINet data as a predictor in the imputation.</p> <p> </p> <p>"aug_pipeline.R" runs through all component steps and calls individual functions and data files within the directory. "plots_for_pipeline.R" uses data created during the aug_pipeline script to visualize individual steps in the augmentation process.</p>
Machine Learning Augmented DA Code and Data
Open the record for dataset details and reuse information.
Data For: Retrieval Augmented Docking using Hierarchical Navigable Small Worlds
<p>These are the DOCK scores for the DUDE-Z "goldilocks" molecules docked to each of the 43 DUDE-Z proteins used in the paper: Retrieval Augmented Docking using Hierarchical Navigable Small Worlds.</p> <p>The data is saved as a pickle of a python dictionary. The keys are the ZINC IDs of the molecules, and the values are lists where the first entry is the corresponding SMILES string, and the second is a dictionary of DOCK scores for each DUDE-Z receptor. If a receptor does not appear in a particularly molecule's dictionary, it means that the molecule failed to dock to the receptor.</p> <p>{</p> <p>zinc_id1: [SMILES, {receptor1:score, receptor2:score,...} ],</p> <p>zinc_id2: [SMILES, {receptor1:score, receptor2:score,...} ],</p> <p>....</p> <p>}</p>
Upscaling Tower-Based Net Ecosystem Productivity to global 250m using the Data Augmentation Method by Considering their Spatial Distribution
<p>Terrestrial ecosystems have emerged as critical carbon sinks, holding a crucial role in the carbon cycle. Net ecosystem productivity (NEP) is a highly significant parameter in terrestrial ecosystems, representing the net ecosystem exchange (NEE) between ecosystems and the atmosphere, without considering other carbon fluxes from disturbances. In this NEP product, we harmonized various sets of tower-based NEP from flux sites as target variable, remote sensing product and meteorological data as traning variables. We further optimizied these smaple sets to address the problems in spatial distribution, culminating in a global NEP product spanning the years 2001-2022, achieved through the application of the random forest method. This dataset contains NEP data for global terrestrial ecosystems for the period 2001-2022 in MgC with a temporal resolution of 1 year. The spatial resolution of the product is 250m and the data format is TIFF.</p> <p><strong>For detailed instructions on how to use the dataset, see User Guides.doc!</strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.