Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

50

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

50 results for “Dataset construction”

Learn how ShareScore rates datasets ↗
zenodo48/100

Construction Industry Steel Ordering Lists (CISOL) Dataset

<p>The Construction Industry Steel Ordering Lists (CISOL) dataset comprises table-centric, real-world documents from the construction industry, annotated to facilitate the testing and training of deep learning models for table detection (TD) and table structure recognition (TSR).&nbsp;</p> <p>CISOL Key Features:</p> <ul> <li>Steel ordering lists from 24 construction projects carried out between 2015-2023, contributed by 10 distinct German structural engineering firms.</li> <li>Anonymized images to ensure the unrecognizability of specific project or creator information.</li> <li>A total of 3280 images, with 844 annotated following the CISOL annotation guidelines.</li> </ul> <p>CISOL is structured into two tracks:</p> <ul> <li><strong>Track A: TD-TSR&nbsp;</strong>version for end-to-end table detection and table structure recognition tasks.</li> <li><strong>Track B: TSR-</strong>only version for table structure recognition tasks, featuring images cropped to the actual table areas with accordingly adjusted annotations.</li> </ul> <p>The dataset is developed in accordance with the FAIR Principles, ensuring that it is Findable, Accessible, Interoperable, and Reusable. The CISOL dataset permits expansion following the established annotation guideline.</p> <p>Access to the CISOL Leaderboard will be provided at <a href="https://eval.ai/web/challenges/challenge-page/2257" target="_blank" rel="noopener">EvalAI.</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Dataset supporting the paper "Doublet-Singlet-Doublet Transition in a Single Organic Molecule Magnet On-Surface Constructed with up to 3 Aluminum Atoms. Nano Letters 21, 8317 (2021)"

<p>Dataset corresponding to theoretical calculations in the paper &quot;Doublet-Singlet-Doublet Transition in a Single Organic Molecule Magnet On-Surface Constructed with up to 3 Aluminum Atoms&quot; Nano Letters 21, 8317 (2021), <a href="https://doi.org/10.1021/acs.nanolett.1c02881">https://doi.org/10.1021/acs.nanolett.1c02881</a></p> <p>List of files:</p> <p>Several folders corresponding to the figures of the paper. They contain:</p> <ul> <li>.siesta files: STM images in WsXM format (http://www.wsxm.eu/) simulated using STMpw (<a href="https://doi.org/10.5281/zenodo.3581159">https://doi.org/10.5281/zenodo.3581159</a>).</li> <li>CONTCAR and POSCAR files: relaxed structures in VASP format. They can be visualized with VESTA (<a href="https://jp-minerals.org/vesta/en/">https://jp-minerals.org/vesta/en/</a>).</li> <li>.agr: grace files (<a href="https://plasma-gate.weizmann.ac.il/Grace/">https://plasma-gate.weizmann.ac.il/Grace/</a>).<br> &nbsp;</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo44/100

TCOM-H2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric H2O profile dataset [1991-2021] constructed using machine-learning.

<p>Methodology: &nbsp;</p> <p>The <strong>TOMCAT simulation</strong> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized <strong>ERA-5 reanalysis data</strong>.</p> <h3>H2O Profile Processing and Bias Correction</h3> <p><strong>Collocated H2O profiles</strong> are organized into five distinct latitude bins:</p> <ul> <li> <p><strong>NH polar</strong>: 90∘N - 50∘N</p> </li> <li> <p><strong>NH mid-lat</strong>: 20∘N - 70∘N</p> </li> <li> <p><strong>Tropics</strong>: 40∘S - 40∘N</p> </li> <li> <p><strong>SH mid-lat</strong>: 70∘S - 20∘S</p> </li> <li> <p><strong>SH polar</strong>: 90∘S - 50∘S</p> </li> </ul> <p>Initially, <strong>differences between TOMCAT and satellite measurements</strong> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from 10,km to 60,km). Note that TOMCAT may not accurately capture H2O evolution post-HTHH eruption due to the sparse spatial coverage of ACE measurements, which limits training data.</p> <p><strong>Separate XGBoost regression models</strong> are then trained for these H2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate <strong>H2O bias corrections</strong> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</p> <p><strong>Height-resolved H2O profile data</strong> are then interpolated onto 28 standard pressure levels (from 300,hPa to 0.1,hPa), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</p> <p>We acknowledge the inherent <strong>dry biases in the original TOMCAT H2O profiles</strong>, largely because the TTL entry mixing ratios are based on a simplistic sinusoidal seasonal cycle, which omits the H2O enhancement contributed by tropical convective clouds.</p> <h3>Data Files</h3> <p>The dataset includes two files containing daily mean zonal mean H2O profiles:</p> <ul> <li> <p><code>zmh2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>height level data</strong> (10,km to 60,km).</p> </li> <li> <p><code>zmh2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>pressure level data</strong> (300,hPa to 0.1,hPa).</p> </li> </ul> <h3>Reference Publication</h3> <p>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</p> <p>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, <a title="null" href="https://doi.org/10.5194/essd-15-5105-2023">https://doi.org/10.5194/essd-15-5105-2023</a>, 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

TCOM-O3: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric ozone profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;TOMCAT simulation is performed at T64L32 resolution for the 2000-2024 time period. Collocated Ozone (O3) profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement &nbsp;differences are calculated for each zonal bins (51 height levels, 10km to 60km). Note that if enough ACE measurements are not avaliable for a particular level then data is purely based on TOMCAT simulated output field. Separate XGBoost regression models are trained for the &nbsp;differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids. &nbsp;TOMCAT output sampled at 1.30 pm local time at the equator. Estimated corrections for a given model grid that are added to the original TOMCAT simulated day and night time ozone profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values. &nbsp;For more details see attached presentation. Previous version use both HALOE and ACE data. Here only ACE data is used (hence starting date is 01 January 2000). PDF file shows comparison between v1.0 and v1.1 as well as TOMCAT data.</p> <p>Dataset also includes two files containing daily mean zonal mean hydrogen fluoride &nbsp;profiles on height (10-60 km) and pressure (300-0.1 hPa) levels:</p> <p>zmo3_TCOM_hlev_T2Dz_2000_2024.nc &ndash; height level data (10 to 60 km)</p> <p>zmo3_TCOM_plev_T2Dz_2000_2024.nc &ndash; pressure level data (300 to 0.1 hPa)</p> <p>Daily 3D profiles on height and pressure levels would be made available on request.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Inertial Dataset for Posture Recognition in Agriculture and Construction Tasks

<p>This <strong>dataset, manually labeled</strong>, contains <strong>10 hours and 40 minutes</strong> (60 Hz) of <strong>8 typical working posture classes</strong> (standing, reaching, stooping, squatting, kneeling, lifting/lowering, carrying, and others), acquired with 16 subjects in three distinct scenarios in a lab environment:</p> <ol> <li>Isolated postures or short sequences without any associated task;</li> <li>Agriculture task (bricklaying) circuit;</li> <li>Construction task (harvesting) circuit.</li> </ol> <p>Two full-body inertial motion caption systems (<strong>17</strong> Xsens MTw Awinda <strong>IMUs</strong> each, from Xsens Technologies, B.V., The Netherlands) were used, connected to, respectively:</p> <ol> <li>Xsens MT Manager, providing raw inertial data (acceleration, angular velocity, and magnetic field data - csv files);</li> <li>Xsens MVN Analyze, providing processed data (quaternions, Euler angles, position, linear velocity, acceleration, angular velocity, angular acceleration, joint angles, ergonomic angles, center of mass, and magnetic field - xlsx files).</li> </ol> <div> <p>More information about the dataset acquisition and organization is detailed in readme.pdf file.&nbsp;For any questions, please contact Diogo R. Martins at&nbsp;<a href="mailto:diogo-martins-9@live.com.pt">diogo-martins-9@live.com.pt</a>&nbsp;or Sara M. Cerqueira at <a href="mailto:saracerqueira1996@gmail.com">saracerqueira1996@gmail.com</a>.</p> </div>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A Comprehensive Self-Consolidating Concrete Dataset for Advanced Construction Practices

<ul> <li><span>Size: over 2500 Self-consolidating concrete mixtures from 176 published papers.</span></li> <li><span>Material type: Self-consolidating concrete (SCC).</span></li> <li><span>Features:</span> <ul> <li><span>Identification features (5 features): References, number of the mixture, the authors, year of publication, &amp; the mixture code.</span></li> <li><span>Powders type, content, &amp; density (76 features): Cement, various supplementary cementitious materials, &amp; other mineral additions.</span></li> <li><span>Paste properties (8 features): The total amount of powder used, the water content, the calculated volume of the paste, the water-to-cement ratio, the water-to-binder ratio, the water-to-powder ratio, the volume of water to the volume of powder ratio, &amp; the volume of water to the volume of cement ratio.</span></li> <li><span>Aggregate properties (7 features): Content and density of fine and coarse aggregates, the total aggregate, the maximum size of the aggregate, &amp; the fine-to-total-aggregate ratio.</span></li> <li><span>Admixture properties (3 features): Quantity of admixture used, its proportion relative to the cement &amp; the total binder content.</span></li> </ul> </li> <li>Properties: <ul> <li>Fresh properties (13 features): Including filling ability properties, i.e., slump flow spread, V-funnel flow time, &amp; the T50 time; Passing ability properties, i.e., J-Ring flow spread, L-box H1/H2 ratio, &amp; U-box flow; Segregation resistance i.e., sieve segregation index, column segregation index, dynamic segregation index, segregation factor, &amp; sieve GTM stability test. Additionally, the percentage of air content is also documented.</li> <li><span>Rheological properties (3 features): yield stress &amp; plastic viscosity values alongside with the used rheometer. The instruments employed in these measurements include the ICAR Rheometer, R/S Plus Rheometer, ConTec5 Viscometer, ConTec4SCC, Concrete Shear Box, &amp; TR-CRI Concrete Rheometer.</span></li> </ul> </li> <li><span>Application: Essential in choosing Self-Compacting Concrete (SCC) mixtures for different uses, considering the importance of both fresh &amp; rheological properties. Intended to support the creation of sustainable &amp; eco-friendly building materials.</span></li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo44/100

A Construction Waste Landfill Dataset of Two Districts in Beijing, China from High Resolution Satellite Images

<p>CWLD_model project shows scripts and instructions on how to use this dataset to train a segmentation model. requirements.txt files provide the libraries you need to run your project. The README.md document details the deployment process and features of each module.</p> <p>You can also visit the GitHub page for scripts and instructions on how to use this dataset for visualizing and plotting basic statistics. The models and the code to execute them are released on&nbsp;<a href="https://github.com/huangleinxidimejd/CWLD_Model">https://github.com/huangleinxidimejd/CWLD_Model</a>.</p> <h2>Training details</h2> <p>The model was trained with two GPUs, an Nvidia GeForce RTX 2080Ti, and the following parameters:</p> <ul> <li>'train_batch_size': 4,</li> <li>'val_batch_size': 4,</li> <li>'train_crop_size': 512,</li> <li>'val_crop_size': 512,</li> <li>'lr': 0.001, # the learning rate used during training. It determines how quickly the model learns from the data</li> <li>'Epoch Times': 200,</li> <li>'gpu': correct,</li> <li>'weight_decay': 5E-4,</li> <li>'Momentum': 0.9,</li> <li>'print_freq': 100,</li> <li>'predict_step': 5,</li> </ul> <h2>usage</h2> <ul> <li>After downloading the dataset from Zenodo, place the train and val files from the Deep Learning Datasets file into the data folder of the CWLD semantic segmentation model.</li> <li>Open: CWLD_ Open the root directory in CWLD_model/dataset/ and start training with the WasteSeg_Train.py file. The modelss module provides five convolutional networks, Improved_DeeplabV3_plus, PSPNet, ResNet, SegNet, and UNet, which can be selected and modified accordingly.</li> <li>The utils package provides a large number of data processing tools to use.</li> <li>The trained model can be predicted from a EvalSeg.py file.</li> </ul>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Large-scale 3D building and tree datasets constructed from airborne LiDAR point clouds in Glasgow, UK

<p>This is the updated version of building 3D model data. The revision includes appending attributes to the lod1 and lod2 shapefile and creating cityjson file for each 3D building model. All 3D building models are available in mesh (.obj), multipath shapefile, and cityjson (.json) now.</p> <p><strong>IMPORTANT NOTE: We suggest using the building footprint, lod1, and lod2 data of this version (Version v4).</strong></p> <p>Urban Big Data Centre of the University of Glasgow generates 3D city models via the airborne LiDAR point clouds acquired between 2020-2021 on behalf of Glasgow City Council. It is a large-scale 3D city model containing 3D information on terrain, trees, and buildings in Glasgow City. This dataset comprises terrain, tree canopy, and building products derived from high-density airborne LiDAR point clouds.&nbsp;</p> <p>The terrain products include Digital Terrain Model (DTM), Digital Surface Model (DSM), and normalized Digital Surface Model (nDSM) in 0.5 m spatial resolution. The DTM and DSM rasters were provided by the vendor and nDSM rasters were obtained by subtracting DTM from DSM. Terrain products are provided in 5 km by 5 km GeoTIF format raster.</p> <p>The tree canopy products are composed of canopy height models (CHM) and tree top locations. Classified tree point clouds were applied with pit-free algorithm to generate CHM in 0.5 m grid raster in GeoTIF format [1]-[2]. Treetop locations were identified by using Local Maximum Filter based on CHM and are recorded as points in Shapefile format. The tree canopy products are provided in 5 km by 5 km tiles.</p> <p>Building 3D model products include footprint polygons with building height attributes and 3D mesh of building models in LoD1 and LoD2 levels. A series of processes such as converting building point clouds to building height models (BHM), converting BHM to polygons, and polygon regularization were conducted to obtain the building footprint polygons. Building height attributes were calculated from BHM for each footprint. The building footprint data are provided in Shapefile format. LoD1 models were generated based on the footprint and average height of the building. LoD2 models were constructed based on footprint and building point cloud with City3D tool[3]. LoD1 and LoD2 models are provided in OBJ and shapefile format. Building 3D model products are provided in 5 km by 5 km tiles. The RMSE of Euclidean distances between each point in the point cloud to the reconstructed model was calculated to evaluate the LoD2 model construction. A table of RMSE and a note for a few problematic models are provided.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Integrated Harmonized Dataset Adolescent Substance Use, Psychosocial Constructs, and Demographics

<p>This dataset contains final analysis cases used in our paper&nbsp;Psychosocial Constructs Related to Alcohol, Cigarette, and Marijuana Use:&nbsp;An Integrated and Harmonized Analysis.&nbsp;We assembled raw data from 25 longitudinal research projects. We collected data from our own research projects (7 projects) as well data provided by 18 researchers. Datasets included epidemiological studies and prevention studies. For the latter, only control group and pretest data were included. All data, including surveys and projects have been de-identified.</p>

opencc-by-3.0-usAug 2021View details →
zenodo44/100

Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: datasets collected using light microscopy and DNA metabarcoding

<p>This dataset contains all data required to reproduce the analyses conducted in Macgregor&nbsp;<em>et al.&nbsp;</em>(2018), using the R Notebook archived at doi: <a href="https://dx.doi.org/10.5281/zenodo.1322712">10.5281/zenodo.1322712</a>.</p> <p>Specifically, the dataset contains details of pollen transport detected on two matched samples, each containing 311 moths of 41 species, using two methods: a traditional light microscopy approach and a novel DNA metabarcoding approach. Both raw and manually-curated versions of each dataset are archived for full clarity.&nbsp;The dataset additionally contains all metadata required to fully interpret these data, including the RGB tables used to prepare Fig 4 in Macgregor <em>et al. </em>(2018).</p> <p>Macgregor&nbsp;<em>et al.&nbsp;</em>(2018) Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: a comparison using light microscopy and DNA metabarcoding.&nbsp;<em>Ecological Entomology</em>,&nbsp;doi: <a href="https://dx.doi.org/10.1111/een.12674">10.1111/een.12674</a>.</p>

opencc-by-4.0Sep 2018View details →
zenodo44/100

Dataset for a typological study of adpossessive constructions

<p>This material contains the dataset and the scripts from the <a href="https://version.helsinki.fi/gramadapt/udw2020-adpossessive-constructions">gitlab repository</a> of the following article. Please cite the article when using the data.</p> <p>Sinnem&auml;ki, Kaius &amp; Viljami Haakana 2020. Variation in Universal Dependencies annotation: A token-based typological case study on adpossessive constructions. In Marie-Catherine de Marneffe, Miryam de Lhoneux, Joakim Nivre &amp; Sebastian Schuster (eds.), <em>Proceedings of the Fourth Workshop on Universal Dependencies (UDW 2020)</em>, 158&ndash;167. Barcelona (online): The Association for Computational Linguistics. Available at&nbsp;<a href="https://www.aclweb.org/anthology/2020.udw-1.0">https://www.aclweb.org/anthology/2020.udw-1.0</a>.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

TCOM-CH4: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric methane profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;</p> <p><span>he </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>CH4 Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated CH4 profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>). It is important to note that unlike previous versions that might have used both HALOE and ACE measurements, this version exclusively utilizes </span><strong><span>ACE-FTS data</span></strong><span>, which is why the dataset starts from 2000.</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these CH4 differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>CH4 bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved CH4 profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean CH4 profiles:</span></p> <ul> <li> <p><code><span>zmch4_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmch4_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023.</span></p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

TCOM-N2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric nitrous oxide profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;</p> <p><span>The </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>N2O Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated N2O profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these N2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>N2O bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved N2O profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean N2O profiles:</span></p> <ul> <li> <p><code><span>zmn2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmn2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023.</span></p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Dataset for the paper "A framework for robotic excavation and dry stone construction using on-site materials"

<p>Stone data from the <i>Science Robotics</i> paper "A framework for robotic excavation and dry stone construction using on-site materials" containing:</p><ul><li>Mesh files of 1,100 stones (quarried boulders, erratics, and concrete debris) that were digitized by the autonomous excavator HEAP<ul><li><a href="https://zenodo.org/api/records/10038881/draft/files/1100%20Unprocessed%20Stone%20Meshes.zip/content">1100 Unprocessed Stone Meshes.zip: </a>Raw mesh files directly from the poisson reconstruction of accumulated LiDAR points, containing some artifacts and floating geometries</li><li><a href="https://zenodo.org/api/records/10038881/draft/files/1100%20Closed%20Stone%20Meshes.zip/content">1100 Closed Stone Meshes.zip: </a>Clean, closed, downsampled meshes</li><li><a href="https://zenodo.org/api/records/10038881/draft/files/Stone_Shape_Properties.csv/content">Stone_Shape_Properties.csv: </a>Properties file with a list of the stone IDs (IDs in the 1xxx and 3xxx range typically correspond to concrete elements) and select shape properties</li></ul></li><li>A dataset of candidate placements from automatically generated stone walls. &nbsp;The candidate placement data zip files contain:<ul><li><a href="https://zenodo.org/api/records/10038881/draft/files/Candidate_Placement_Data-npy.zip/content">Candidate_Placement_Data-npy.zip: </a>SDF (.npy) representation of each candidate, with three channels of 32x32x32 for distances to the stone, the already-placed stones, and the target wall</li><li><a href="https://zenodo.org/api/records/10038881/draft/files/Candidate_Placement_Data-pcd.zip/content">Candidate_Placement_Data-pcd.zip: </a>Point cloud (.pcd) representations of each candidate, with separate files for the placed stone, target wall (search volume), and already-placed stones (where they exist)</li></ul></li><li><a href="https://zenodo.org/api/records/10038881/draft/files/sdf_classifier.zip/content">sdf_classifier.zip: </a>Python examples:<ul><li>Rendering the three channel SDF data to mesh geometry using marching cubes and libigl</li><li>Candidate SDF classification using the pretrained model</li></ul></li><li>Candidate attributes and labels<ul><li><a href="https://zenodo.org/api/records/10038881/draft/files/candidate_attributes_labels.csv/content">candidate_attributes_labels.csv: </a>CSV file containing a list of UUID's corresponding to each candidate placement in the dataset. &nbsp;For each candidate, additional information is included about the dimensions and location of the solution, together with the (subjectively) hand-labelled binary value for placement viability.&nbsp;</li></ul></li><li><a href="https://zenodo.org/api/records/10038881/draft/files/README.md/content">README.md: </a>An additional readme with some details on the&nbsp;attributes file</li></ul><p>If you use this data in your research, please cite the <a href="https://www.science.org/doi/10.1126/scirobotics.abp9758">journal article</a>.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Generic Constructions Datasets

<p>Add Basic Results from ProSenseAIR Algorithm (https://github.com/JanChristianRedlich/ProSenseAIR) on NOW Corpus inclusive Visualizations.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Dataset: Construction Partners, Inc. (ROAD) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

D-PLACE dataset derived from Binford 2001 'Constructing Frames of Reference'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Binford, L. 2001. Constructing Frames of Reference: An Analytical Method for Archaeological Theory Building Using Hunter-gatherer and Environmental Data Sets. University of California Press</p> </blockquote>

opencc-by-nc-4.0Nov 2023View details →
zenodo40/100

Image Dataset for Object Detection of Small Size Construction Tools

<p>&nbsp;This is an image dataset established as input data for object detection model of small-sized construction tools. In the dataset, there are 12 classes of target tools&nbsp; (bucket, cutter, drill, grinder, hammer, knife, saw, shovel, spanner, tacker, trowel, and wrench) which are typically used at indoor construction sites. 25,084 sets of image and the corresponding label data have been established and shared.&nbsp;</p> <p>&nbsp;The diversity of objects in the images of the 12 small tools was considered by photographing tools of various shapes, sizes, and colors. In addition, to improve the model performance, images were also captured with various changes (e.g., image resolution, occlusion, lighting, and background). Among the 25,084 images in the dataset, 6,258 (25%) were obtained from the actual construction site.&nbsp;</p> <p>&nbsp;Object annotations in each image were done by bounding boxes and were saved into a text file. The coordinates of the bounding box have the form of (Class, Center X, Center Y, Width, Height). Class refers to one of 12 construction tool types. Center X and Center Y are the center coordinates of the bounding box for an object from an image when the resolution of the image has min-max normalized. Width and Height are the width and height of the bounding box for an object, respectively, also from the image with the min-max normalized resolution.</p> <p>&nbsp;</p> <p>The peer-reviewed publication for this dataset has now been published in &quot; KSCE Journal of Civil Engineering&quot; a Springer journal as follows:</p> <p><strong>* Lee, K., Jeon, C., and Shin, D. (2023, In press) &quot;Small Tool Image Database and Object Detection Approach for Indoor Construction Site Safety&quot; <em>KSCE Journal of Civil Engineering</em>. DOI: https://doi.org/10.1007/s12205-022-1011-7</strong></p> <p>Please cite this reference when using the dataset.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

ReCon Soil Project: Dataset for Identification of Soil Health Indicators in Construction

<p>This dataset represents experimental results from Work Package 4 of the ReCon Soil Project: https://www.plymouth.ac.uk/research/institutes/sustainable-earth/the-recon-soil-project&nbsp;</p> <p>Soil health indicators for maintenance in a construction context, are less about soil fertility and more about maintaining those indicators identified, to minimise loss of carbon&nbsp;(C) and nitrogen (N), maintain soil microbial activity and to ensure stockpiled soils are maintained to a point where they can be reused effectively in the future. To this end, an experiment was designed in order to identify and monitor key soil health indicators for construction, and to investigate methods to mitigate against carbon losses in stockpiled soils.</p> <p>The main questions that were&nbsp;addressed were:</p> <p>1. Will sowing grass and other plant species in stockpiled soils reduce losses of C and N from the soil?&nbsp;</p> <p>2. What effect does soil stockpiling have on microbial communities?&nbsp;</p> <p>3. What effect does wood chip application on stockpiled soils have on C and N stocks and microbial communities? &nbsp;</p> <p>4. How do rapid in-field methods for carbon and microbial analysis compare with established lab methods?</p> <p>The experiment was established at the end of August 2022. Any vegetation was cleared from the study site by using a digger to scrape away the top 5cm of soil. Diggers were used to take soil from an area adjacent to the study site which had been undisturbed for 3-4 years, vegetation was cleared from this area and the topsoil was formed into 12 stockpiles (6m length x 4m width x 2m height). The stockpiles were then randomly assigned one of three treatments; amenity grass seed mix, herbal ley seed mix&nbsp;or control. For the seeded treatments, each stockpile was sown by hand to an application rate of 70g/m<sup>2</sup>. After seeds were sown, wood chip was applied by hand to half of each stockpile (randomly assigned), to a depth of covering of approximately 2cm. This resulted in 4 replicates of each combination producing 24 sample plots.</p> <p>Soil samples were collected using a soil sampling auger to collect samples at 0-30 cm and 90-100cm. One sample was taken from each replicate at both depths, these were taken at baseline (T<sub>0</sub>), and intervals of 1 week (T<sub>1</sub>), 2 weeks (T<sub>2</sub>), 4 weeks (T<sub>4</sub>) and every 4 weeks after that. The soil cores were sealed in sterile plastic bags and transported to the laboratory for processing.</p> <p>At each time point, soil was analysed in the laboratory at Eden Project Learning for the following parameters: non-purgeable organic carbon (NPOC), total nitrogen (TN), soil organic matter (SOM), gravimetric water content, pH, and microbial activity using a fluorescein diacetate assay. All remaining bagged soil samples were then transferred to a freezer and stored at -18&deg;C. These were later defrosted and analysed for total organic carbon (TOC), available iron, ammonia, nitrate and nitrite at a subset of time points. Using in-field equipment, soils were analysed for CO<sub>2</sub> emissions at every time point, fungi:bacteria ratio and for soil carbon every 4 weeks. Samples for soil eDNA were taken from each replicate at T<sub>0</sub>, T<sub>8</sub> and T<sub>20</sub>.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Dataset for the collected responses for the items measuring constructs affecting eHS non-acceptance behavior in Nigeria

<p>This dataset is a collection of responses from the questionnaire distributed to study the e-health service non-acceptance behavior prominent in Nigeria. A total of 543 valid responses were collected. This research used an integration model based on TPB and SOR theory The dataset were analysed using PLS-SEM. Refer to the article for the results of this study.<br><br></p> <p><strong>Note:</strong> CO = Communication overload; CHO = Choice overload; PR = Perceived oisk; HL = Health literacy; NA = Negative attitude; SN = Subjective norms; PBC = Perceived behavioral control; INTU = Intention not to use eHS; NAB = Non-acceptance behavior</p>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record