Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,805

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,805 results for “Data model”

Learn how ShareScore rates datasets ↗
zenodo36/100

SDUST2021GRA: Global marine gravity anomaly model recovered from Ka-band and Ku-band satellite altimeter data

<p>SDUST2021GRA is the global marine gravity anomaly model on&nbsp;&nbsp;a grid of 1&prime;&times;1&prime;, which is established from the altimeter data of&nbsp;<strong>&nbsp;</strong>Ka-band and Ku-band&nbsp; altimetry satellite including HY-2A.&nbsp;Its spatial coverage is&nbsp;80&deg;S-80&deg;N.&nbsp;Assessed by the shipborne gravity data, the accuracy of SDUST2021GRA in the global is 2.37 mGal, and that in the open ocean is about 1.5 mGal.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Modelling snowpack bulk density using snow depth, cumulative degree-days and climatological predictor variables -- data set

<p>This file constitutes the data set containing the snow course survey, North American Regional Reanalysis (NARR)-derived degree-day indices, and climatological variables data used to conduct the analysis, and generate the figures and tables in the manuscript titled &quot;Modelling snowpack bulk density using snow depth, cumulative degree-days and climatological predictor variables&quot; by Andras J. Szeitz and R. Dan Moore. The manuscript was submitted for publication in the journal &#39;Hydrological Processes&#39;.</p> <p>Due to the size of the NARR data files used to derive the air temperature time series for each snow course location, we recommend acquiring them from the National Oceanic and Atmospheric Administration&#39;s data portal directly (<a href="https://psl.noaa.gov/data/gridded/data.narr.html">https://psl.noaa.gov/data/gridded/data.narr.html</a>).</p> <p>Likewise, the ClimateNA software application used to extract the climatological variables for each snow course location can be obtained from the Centre for Forest Conservation Genetics, Department of Forest and Conservation Sciences, UBC, directly (<a href="https://climatena.ca/">https://climatena.ca/</a>).</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Data availability: Random encounter model is a reliable method for estimating population density of multiple species using camera traps

<p>Data of the paper entitled &quot;Random encounter model is a reliable method for estimating population density of multiple species using camera traps&quot; published on Remote Sensing in Ecology and Conservation</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Source Data files for: Primary cilia and SHH signaling impairments in human and mouse models of Parkinson's disease

<p>Parkinson&rsquo;s disease (PD) as a progressive neurodegenerative disorder arises from multiple genetic and environmental factors. However, underlying pathological mechanisms remain poorly understood. Using multiplexed single-cell transcriptomics, we analyze human neural precursor cells (hNPCs) from sporadic PD (sPD) patients. Alterations in gene expression appear in pathways related to primary cilia (PC). Accordingly, in these hiPSC-derived hNPCs and neurons, we observe a shortening of PC. Additionally, we detect a shortening of PC in <em>PINK1</em>-deficient human cellular and mouse models of familial PD. Furthermore, in sPD models, the shortening of PC is accompanied by an increased SHH signal transduction. Inhibition of this pathway rescues the alterations in PC morphology and mitochondrial dysfunction. Thus, increased SHH activity due to ciliary dysfunction is needed for the development of pathoetiological phenotypes observed in sPD, like mitochondrial dysfunction. In sum, altered PC function is part of early PD pathoetiology and inhibiting the overactive SHH signaling is a potential neuroprotective therapy.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

CESM2 data for "Ocean complexity shapes sea surface temperature variability in a CESM2 coupled model hierarchy" - submitted to JCLI

<p><strong>CESM2 Experiment names:</strong></p> <ul> <li>FC = fully coupled model, CESM2 (variables freely available on https://esgf-node.llnl.gov/search/cmip6/)</li> <li>MD&nbsp;= mechanically decoupled model, CESM2</li> <li>SOM = slab ocean model, CESM2</li> </ul> <p>All datasets are for pre-industrial forcing (e.g., piControl), nominal 1-degree horizontal resolution&nbsp;</p> <p>---</p> <p>Decoding the files names:</p> <ul> <li><strong>climatology_monthly </strong>= 12 month&nbsp;climatology&nbsp;</li> <li><strong>climatology_annual</strong> = time mean climatology</li> <li><strong>variance</strong> = anomaly variance computed over time</li> </ul> <p>---</p> <p>Variables:</p> <ul> <li><strong>PRECL</strong> = large-scale convective precipitation</li> <li><strong>PRECC</strong> = convective precipitation</li> <li><strong>total precipitation (not provided but can be calculated)</strong> = PRECC + PRECL</li> <li><strong>HMXL</strong> = mixed layer depth</li> <li><strong>SST</strong> = sea surface temperature&nbsp;</li> </ul> <p><strong>Files for the CESM2 MD piControl run:</strong></p> <ol> <li>forcing_coupled.F90: POP2 (ocean) source code changes for cesm2.1.4-rc08 (search for &quot;slarson&quot; throughout code to find our changes</li> <li>cesm2.1.4-exp03-CTRL_B1850_f09_g17_hourlyclim_TAUX.nc: 6 hourly climatology for TAUX, from a FC run of CESM2. This file and the TAUY climatology&nbsp;are opened and read in the &quot;rotate wind stress&quot; subroutine in forcing_coupled.F90. This file is&nbsp;named &quot;x2oavg_Foxx_taux_6hourly.nc&quot;&nbsp;in forcing_coupled (we wanted a shorter file name in the code)</li> <li>cesm2.1.4-exp03-CTRL_B1850_f09_g17_hourlyclim_TAUY.nc: 6 hourly climatology for TAUY.&nbsp;This file is&nbsp;named &quot;x2oavg_Foxx_tauy_6hourly.nc&quot;&nbsp;in forcing_coupled&nbsp;(we wanted a shorter file name in the code)&nbsp;</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
dryad36/100

Data from: Integrating 3D models with morphometric measurements to improve volumetric estimates in marine mammals

<p>1. Studies of body condition are key to understanding the health, bioenergetics, and ecological roles of marine mammals. Due to challenges in studying marine mammals at sea, body condition is often approximated using metrics representing the size of the dorsal surface visible from aerial imagery, but quantifying variability in body volume would enable a more holistic understanding of bioenergetics. Further, the number and location of measurements needed to accurately quantify body condition has received little attention. Three-dimensional (3D) models provide a promising tool for representing morphology and providing holistic estimates of marine mammal body condition when combined with field-based morphometric measurements.</p> <p>2. We use humpback whales (Megaptera novaeangliae) to demonstrate the utility of 3D models for estimating body condition in marine mammals. We integrate morphometric measurements taken from Unoccupied Aerial Vehicles (UAVs) with scalable 3D models to generate estimates of humpback whale body volume. We assess which and how many morphometric measurements are required to accurately estimate body volume and compare the error between volume estimates derived from 3D models and previously developed models representing volume as a series of ellipses. Using UAV measurements, we assess the contribution of each morphometric measurement to volumetric estimates, and quantify the error produced by all combinations and numbers of morphometric measurements (131,072 combinations).</p> <p>3. Error in volume estimates from 3D models generated with as few as five width measurements was &lt;5% compared to the full models and was lower than the error produced when using five width measurements with the elliptical approach. We suggest that by conserving the external morphology of marine mammals, 3D models allow body volume and body condition to be estimated accurately with few measurements.</p> <p>4. We provide code and guidelines for creating 3D models using the open-source software Blender and for assessing which measurements are needed to accurately capture the morphology of cetaceans. The 3D modeling approach we present will facilitate studies of intra- and interannual changes in body volume in marine mammals, which is vital to providing a more holistic understanding of bioenergetics and to assessing responses to environmental change and anthropogenic stressors.</p>

opencc-zeroJul 2022View details →
zenodo36/100

Supplementary data - "Learning Reduced Models for Large-Scale Agent-Based Systems"

<p>This repository contains supplementary data on my PhD thesis &quot;Learning Reduced Models for Large-Scale Agent-Based Systems&quot;.&nbsp;Chapters 1-3, 7 and&nbsp;A&nbsp;do not have supplementary data.</p> <p><strong>Chapter 4</strong></p> <ul> <li><em>Large_deviation_example.zip</em> contains the trajectory for Figure 4.8.</li> <li><em>mean_exit_time*</em> contains the raw data to compute the mean exit time and standard deviation for the ABM process (JP) and SDE process (CLE). It contains additionally a precomputed mean and standard deviation as well as the corresponding numbers of agents.</li> <li><em>transition_matrix*</em> contain the computed box discretizations as MATLAB and Numpy files as used for Figures 4.2-4.4, 4.6 and Tables 4.1&nbsp;and 4.2.</li> </ul> <p><strong>Chapter 5</strong></p> <ul> <li><em>CVM_2021-07-09-15-53_training_data.npz</em>&nbsp;contains the training&nbsp;data for Figure 5.7 a and b.</li> <li><em>CVM_2021-09-29-07-13_distribution.npz&nbsp;</em>contains the raw data for Figure 5.7 c.</li> <li>The remaining data for Chapter 5 can be found in the related dataset&nbsp;<a href="https://doi.org/10.5281/zenodo.4522119">doi.org/10.5281/zenodo.4522119</a>.</li> </ul> <p><strong>Chapter 6</strong></p> <ul> <li><em>CVM_pareto_estimate</em> contains trajectory data required for Figure 6.6&nbsp;b to estimate points in the Pareto Front using the civil violence model.&nbsp;</li> <li><em>CVM_training_data</em> contains the training data to construct the surrogate model. Each data set consists of&nbsp;<em>CVM_*_cops_train.npz</em> as training set,&nbsp;<em>CVM_*_cops_trajectory.npz</em> as sample trajectory and <em>CVM_*_cops.pkl</em>&nbsp;to compute the training data.</li> <li><em>CVM_covering_iterations_8.mat</em> Pareto set covering after 8 iterations for the&nbsp;civil violence model.&nbsp;Required for Figure 6.6&nbsp;a.</li> <li><em>CVM_pareto_set+front.npz</em> is required for Figure 6.6&nbsp;b.&nbsp;</li> <li><em>CVM_surrogate_model.mat</em> contains the surrogate model for the civil violence model</li> <li><em>Expl_iterations_*</em> contains Pareto set coverings after 8 and 12 iterations for Example 6.1.4 and Figure 6.1.</li> <li><em>VM_covering_iterations_12.mat</em> contains the Pareto set covering depicted in Figure 6.4 a.</li> <li><em>VM_ODE_covering_iterations_12_subset_front.mat</em>&nbsp;contains the Pareto set covering depicted in Figure 6.5&nbsp;and 6.5&nbsp;c.</li> <li><em>VM_ODE_covering_iterations_12_subset.mat</em>&nbsp;contains the Pareto set covering depicted in Figure 6.5&nbsp;and 6.5&nbsp;d.</li> <li><em>VM_ODE_covering_iterations_12.mat</em> contains the Pareto set covering depicted in Figure 6.4 b.</li> <li><em>VM_surrogate_model.mat</em>&nbsp;contains the surrogate model for the extended voter model.</li> <li><em>VM_test_points_non_pareto.npz</em> contains Non-Pareto points in Figure 6.5 and 6.5&nbsp;d.</li> <li><em>VM_test_points_pareto.npz</em>&nbsp;contains Pareto points in Figure 6.5 and 6.5&nbsp;c.</li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Model data for "Surface heating over the Tibetan Plateau associated with the Antarctic Oscillation"

<p>In HYSPLIT.rar, the ????06.backjectory.10day.sh.p.pnum.nc data are the hysplit results.</p> <p>In CESM.rar, the pres_f.inc6hr.????.cam.h0.????-05_06.nc and pres_f.ins6hr.????.cam.h0.????-05_06.nc are the CTL and EXP experiment results of AGCM.</p> <p>resp_Amundv7.t42l20.nc is the response of the LBM model.</p> <p>Detailed description is shown in the paper &quot;Surface heating over the Tibetan Plateau associated with the Antarctic Oscillation&quot;.&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Discovering molecular regulators of ageing using mixture models with RNA-sequencing data

<p>Identifying the molecular regulators that control ageing is challenging because the ageing process is influenced by a combination of genetic and environmental factors which makes it difficult to source the contribution of a single gene. Multiple studies have demonstrated that as humans age, increased gene expression heterogeneity results in the dysregulation of key regulators and pathways. Given the dynamic nature of gene expression, it is vital that this data be modelled by statistical approaches that can appropriately account for changes in variability to understand the contribution of heterogeneity during the aging process and properly identify its regulators. This study demonstrates the utility of using mixture models to model biological variability of gene expression occurring during ageing and how novel potential regulators of ageing can be identified.</p> <p>Our mixture modelling approach was applied to gene expression data from the Genotype-Tissue Expression (GTEx) cohort. For every gene, the expression profile was modelled using a mixture model across the cohort where the subset of donors corresponding to each mode was tested for a significant change in age group. The multi-tissue aspect of GTEx was leveraged to find ageing regulators based on this mixture model approach genes that were common across multiple tissues, suggesting that the regulation of ageing may also be controlled through a set of genes that have non-tissue-specific activity.</p> <p>Our approach identified well-documented ageing regulators <em>mTOR </em>and <em>RICTOR</em> and other potential ageing regulators such as <em>IL4</em> and <em>GPR4</em> which were detected only by our approach. Genes identified by edgeR, DESeq2 and the mixture model-based approach were enriched for similar biological pathways. This suggests that while the specific ageing regulators identified from our approach may be distinct, they generally belong in the same pathways as the genes identified by standard approaches. Overall, these results indicate that modelling gene expression variability using mixture models in conjunction with standard differential gene expression can help uncover new regulators that have a potential role for understanding human ageing.</p> <p>I</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Spatz neutron reflectometer commissioning data and corresponding models to fits

<p>This is the data to the hot commissioning of the Spatz neutron reflectometer at the OPAL Research Reactor, ANSTO. The data also contains the Jupyter&nbsp;notebook to reduce the data to give complete reflectivity profile and contains the models used to fit the data within refnx using the GUI option.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2022View details →
dryad36/100

Data from: Influence of different data cleaning solutions of point-occurrence records on downstream macroecological diversity models

<p><span>Digital point-occurrence records from the Global Biodiversity Information Facility (GBIF) and other data providers enable a wide range of research in macroecology and biogeography. However, data errors may hamper immediate use. Manual data cleaning is time-consuming and often unfeasible, given that the databases may contain thousands or millions of records. Automated data cleaning pipelines are therefore of high importance. This study examined the extent to which cleaned data from six pipelines using data cleaning tools (e.g., the GBIF web application, different R packages) affect downstream species distribution models. In addition, we assessed how the pipeline data differ from expert data. From 13,889 North American <i>Ephedra</i> observations in GBIF, the pipelines removed 31.7% to 62.7% false-positives, invalid coordinates, and duplicates, leading to data sets that included between 9,484 (GBIF application) and 5,196 records (manual-guided filtering). The expert data consisted of 703 thoroughly handpicked records, comparable to data from field studies. Although differences in the record numbers were relatively large, stacked species distribution models (sSDM) from the pipelines and the expert data were strongly related (mean Pearson's <i>r</i> across the pipelines: 0.9986, versus the expert data: 0.9173). The ever-stronger correlations resulted from occurrence information that became increasingly condensed in the course of the workflow (from individual occurrences to collectivized occurrences in grid cells to predicted probabilities in the sSDMs). In sum, our results suggest that the <i>R</i> package-based pipelines reliably identified invalid coordinates. In contrast, the GBIF-filtered data still contained both spatial and taxonomic errors. However, major drawbacks emerge from the fact that no pipeline fully discovered misidentified specimens without the assistance of expert taxonomic knowledge. We conclude that application-filtered GBIF data will still need additional review to achieve higher spatial data quality. Achieving high-quality taxonomic data will require extra effort, probably by thoroughly analyzing the data for misidentified taxa, supported by experts.</span></p>

opencc-zeroJul 2022View details →
zenodo36/100

Data from: Filtering ground noise from LiDAR returns produces inferior models of forest aboveground biomass in heterogenous landscapes

<p>Airborne LiDAR has become an essential data source for large-scale, high-resolution modeling of forest aboveground biomass and carbon stocks, enabling predictions with much higher resolution and accuracy than can be achieved using optical imagery alone. Ground noise filtering -- that is, excluding returns from LiDAR point clouds based on simple height thresholds -- is a common practice meant to improve the &#39;signal&#39; content of LiDAR returns by preventing ground returns from masking useful information about tree size and condition contained within canopy returns. However, ground returns may be helpful for making accurate aboveground biomass predictions in heterogeneous landscapes that include a patchy mosaic of vegetation heights and land cover types.<br> &nbsp;<br> &nbsp; In this paper, we applied several ground noise filtering thresholds while mapping forest AGB across New York State (USA), a heterogenous landscape composed of both contiguously forested and highly fragmented areas with mixed land cover types. We fit random forest models to predictor sets derived from each filtering intensity threshold and compared model accuracies, paying attention to how changes in accuracy correlated with landscape structure. We observed that removing ground noise via any height threshold systematically biases many of the LiDAR-derived variables used in AGB modeling, with mean correlation (Spearman&#39;s $\rho$) between variables increasing from 0.183 to 0.266. We found that that ground noise filtering yields models of forest AGB with lower accuracy than models trained using predictors derived from unfiltered point clouds, with RMSE increasing by up to 2.2 Mg ha^-1^ statewide. Although we only modeled AGB for forest cover types, models fit to predictors derived from filtered point clouds performed worse as landscape heterogeneity (as measured by patch density and edge density) increased, suggesting ground returns are particularly useful when modeling edge forests. Our results suggest that ground filtering should be a carefully considered decision when mapping forest AGB, particularly when mapping heterogeneous and highly fragmented landscapes, as ground returns are more likely to represent useful &#39;signal&#39; than extraneous &#39;noise&#39; in these cases.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Land-use fluxes: data from global models and national inventories

<p>This file includes the&nbsp;data&nbsp;from Supplementary Table 1 of&nbsp;Grassi et al. (ESSD, submitted), for the 42 countries having a managed forest area greater than 10 Million ha. The data includes:&nbsp;</p> <p>(i)&nbsp;areas&nbsp;of managed forest, used in this study and based on country data</p> <p>(ii) The&nbsp;CO2 fluxes (2001-2020 average) from global models - i.e.&nbsp;bookkeeping models (BMs) and Dynamic Global Vegetation Models (DGVMs) -, and from a collection of National GHG inventories (NGHGIs) for LULUCF, forest land, deforestation, and other fluxes (organic soils, cropland, grassland etc.).&nbsp;</p> <p>BM values are averages of three models and DGVM values are averages of 17 models, consistent with the Global Carbon Budget 2021 (<a href="https://priv-bx-myremote.tech.ec.europa.eu/articles/14/1917/2022/,DanaInfo=.aetugDhuwm0xto76O48y,SSL+">https://essd.copernicus.org/articles/14/1917/2022/</a>).&nbsp;Values for NGHGIs are from&nbsp;<a href="https://priv-bx-myremote.tech.ec.europa.eu/preprints/essd-2022-104/,DanaInfo=.aetugDhuwm0xto76O48y,SSL+">https://essd.copernicus.org/preprints/essd-2022-104/</a>&nbsp;&nbsp;</p> <p>For further methodological details, see Grassi et al. (ESSD, submitted):</p> <p>Giacomo Grassi, Clemens Schwingshackl, Thomas Gasser, Richard A. Houghton, Stephen Sitch, Josep G. Canadell, Alessandro Cescatti, Philippe Ciais, San1, Etsushi Kato, Daniel Kennedy, J&uuml;rgen Knauer, Anu Korosuo, Matthew J. McGrath, Julia Nabel, Benjamin Poulter, Simone Rossi, Anthony P. Walker, Wenping Yuan, Xu YueJulia Pongratz.&nbsp;Mapping land-use fluxes for 2001-2020 from global models to national inventories. ESSD (submitted)</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Supplementary data and code of the manuscript: "Strong evidence for the adaptive walk model of gene evolution in Drosophila and Arabidopsis"

<p>This repository contains all data tables ad code to reproduce the analysis performed in &quot;Strong evidence for the adaptive walk model of gene evolution in Drosophila and Arabidopsis&quot;.&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

National Data Files for Pre-built Sector-coupled Euro-Calliope Model

<p>National time series data derived from the <a href="https://zenodo.org/record/5774988#.YtUQ9-zP3Ph">Sector-coupled Euro-Calliope Pre-built Model</a></p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Dynamics of the euphotic zone in the Black Sea: The synergy of data from profiling floats, machine learning and numerical modeling

<p>The datasets contain input data and data emulated by Neural networks (NN) used in the study &#39;Dynamics of the euphotic zone in the Black Sea: The synergy of data from profiling floats, machine learning and numerical modeling&#39;</p> <p>- <strong>NN2018_CMEMS_ARGO.tar.gz</strong>: archive contains Matlab binary files consisting of NN-derived BGC variables (Chlorophyll-a, Oxygen and backscatter at 700nm) along ARGO float paths in 2018 with vertical resolution taken from CMEMS (13 depth levels in the depth range studied here); NN was applied either on CMEMS physics (&#39;C&#39;) or on ARGO physics (&#39;A&#39;); additionally, CMEMS BGC model data (Chlorophyll-a and Oxygen) along these paths are included; (filenames follow the names of floats given in Table 1:<em> floatname</em>_2018_NNARGOCMEMS.mat)</p> <p>- <strong>NNalongARGO.tar.gz</strong>: archive contains Matlab binary files consisting of NN-derived BGC variables (Chlorophyll-a, Oxygen and backscatter at 700 nm) along ARGO float paths together with input ARGO data (time, latitude, longitude, salinity, temperature, sigma_T and BGC variables); all variables&nbsp;are mapped with 1m vertical resolution; depth range is [1m 150m] ; additionally, float ogs7 data include NO3, float hzg1 data does not contain Chlorophyll-a and backscatter at 700 nm; (filenames follow the names of floats given in Table 1:<em> floatname</em>_euph_1mRes_150mALLINCLNN.mat)</p> <p>- <strong>NNReconBlackSea.tar.gz</strong>: archive contains Matlab binary files consisting of basin wide NN derived BGC variables (Chlorophyll-a, Oxygen and backscatter at 700nm) for the years 2015-19 and 2010/11 (weekly mean data); additionally CMEMS data (time, latitude, longitude, salinity, temperature, sea surface height) are provided&nbsp;for the photic zone; (filenames are reconNNCMEMS_2015_2019.mat and reconNNCMEMS_2010_2011.mat, respectively)</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Data for: Deep Learning of Model- and Reanalysis- Based Precipitation and Pressure Mismatches over Europe

<p>This study focuses on using UNet Convolutional Neural Networks to predict the spatiotemporal mismatches (errors) between TSMP-G2A model-based and COSMO-REA6 reanalysis-based precipitation and surface pressure over Europe.</p> <p>The following data are provided in this dataset:</p> <p>1) The remapped and NetCDF-merged TSMP-G2A and COSMO-REA6 precipitation and surface pressure over the study area (EU-11 EUROCORDEX, ~0.11 degrees) for the years 1995-2017. Files: COSMO-REA6_PREPROCESSED.zip and TSMP_PREPROCESSED.zip</p> <p>2) The actual and predicted spatiotemporal mismatch data for training, validation, and testing periods (1995-2017). Files: MISMATCH_ACTUAL.zip and MISMATCH_PREDICTED.zip<br> &nbsp;</p> <p>References for original TSMP-G2A and COSMO-REA6 data:<br> TSMP-G2A: http://doi.org/10.17616/R31NJMGR<br> COSMO-REA6: doi:10.1002/qj.2486, 2015</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

FOD CT Data: air pockets in avocado and stone in modelling clay

<p><strong>Summary</strong></p> <p>This submission contains X-ray CT data of avocado fruits and pieces of modelling clay containing pebble stones.<br> Data for every object include binned pre-processed projections and volume segmentations.<br> These datasets can be used for training and testing deep learning methods for foreign object detection.</p> <p>The data is made available as a part of the paper &quot;CT-based data generation for foreign object detection on a single X-ray projection&quot;.</p> <p><strong>Data acquisition</strong></p> <p>A majority of raw data for modeling clay (excluding 10 samples without pebble stones in the Test subset) is taken from the dataset<br> &quot;A collection of 131 CT datasets of pieces of modeling clay containing stones&quot;<br> [![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.5866228.svg)](https://doi.org/10.5281/zenodo.5866228)</p> <p>The remaining pieces of modeling clay and all avocado fruits were scanned at the FleX-ray laboratory<br> of the Centrum Wiskunde &amp; Informatica (CWI) in Amsterdam, the Netherlands (details can be found in [Coban 2020]).<br> For every fruit, we made scans with significantly different amounts of air pockets by waiting for a few days between experimental acquisitions.<br> The measurements were performed with the voltage of 90 kV, power of 45 W, exposure time of 300 ms per projection, and magnification factor of 1.3.<br> The original X-ray image size was 1912 px x 1520 px with a pixel size of 75 &mu;m, 1440 images were acquired for every sample.<br> For faster deep learning model training, images and reconstructions were downsampled with a factor of 4, leading to the effective pixel size of 300 &mu;m and voxel size of 230 &mu;m.<br> Additional scans of the pieces of modeling clay were acquired with settings similar to the main collection.</p> <p><strong>Data Description</strong></p> <p>The submission is split into &quot;Avocado&quot; and &quot;Playdoh&quot; (pieces of modeling clay) datasets. Each dataset is further split into Training and Test subsets.</p> <p>The folder for every scanned object contains<br> - ./log/ - subfolder with logarithmed X-ray projections after darkfield and flatfield correction.<br> - ./segm/ - subfolder with slices of the segmented volume.<br> - ./scan settings.txt - a file with scanner metadata containing scan geometry<br> - ./volume_info.csv - a file with a voxel count for every class in the segmentation.</p> <p>For playdoh objects, the segmentation classes are modeling clay (Class 1) and pebble stone (Class 2). In this case, pebble stones are foreign objects.</p> <p>For avocado objects, the segmentation classes are peel (Class 1), avocado meat (Class 2), seed (Class 3) and air pockets (Class 4). Air pockets are considered a foreign object.</p> <p><strong>Additional Links</strong></p> <p>These datasets are produced by the Computational Imaging group at Centrum Wiskunde &amp; Informatica (CI-CWI). For any relevant Python/MATLAB scripts for the FleX-ray datasets, we refer the reader to our group&#39;s GitHub page.</p> <p><strong>Contact Details</strong></p> <p>For more information or guidance in using these datasets, please get in touch with<br> - vladyslav.andriiashen [at] cwi.nl</p> <p><strong>References</strong></p> <p>[Coban 2020] S. B. Coban, F. Lucka, W. J. Palenstijn, D. Van Loo, and K. J. Batenburg, &ldquo;Explorative imaging and its implementation at the FleX-ray Laboratory,&rdquo; J. Imaging, vol. 6, no. 18, 2020, doi: 10.3390/jimaging6040018.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

data sets from "Updated trends of the stratospheric ozone vertical distribution in the 60S–60N latitude range based on the LOTUS regression model"

<p>Monthly means data sets from satellite, ground-based and model records used in the article entitled: &quot;Updated trends of the stratospheric ozone vertical distribution in the 60 S&ndash;60 N latitude range based on the LOTUS regression model&quot;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Data for "Properties of the lateral mesoscale eddy-induced transport in a high-resolution ocean model: Beyond the flux-gradient relation" (Lu et al. JPO)

<p>Preprocessed data to reproduce results and figures in &quot;Properties of the lateral mesoscale eddy-induced transport in a high-resolution ocean model: Beyond the flux-gradient relation&quot; (Lu et al., In Review of&nbsp;<em>Journal of Physical Oceanography</em>).&nbsp;</p> <p>Feel free to contact Yueyang Lu via&nbsp;<strong>yxl1496@miami.edu</strong>&nbsp;if you have any questions.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record