Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,038
datasets available to search
ShareScore release 0.7.1
Dataset results
8,038 results for “validation”
Initial Sample of HYPERNETS Hyperspectral Surface Reflectance Measurements for Satellite Validation from the Wytham Woods site in the United Kingdom
<p>The HYPERNETS project (www.hypernets.eu) has the overall aim to ensure that high quality in situ measurements are available to support the (VNIR/SWIR) optical Copernicus products. Therefore, it established a new autonomous hyperspectral spectroradiometer (HYPSTAR® - www.hypstar.eu) dedicated to land and water surface reflectance validation with instrument pointing capabilities. In the prototype phase, the instrument is being deployed at 24 sites covering a range of water and land types and a range of climatic and logistic conditions. This dataset provides the first published data for the Wytham Woods HYPERNETS site in the United Kingdom (WWUK). It is a subset of the complete data record which consists of the best quality WWUK measurements which could be used for satellite validation. </p> <p>The provided NetCDF files are the L2A hypernets products with surface reflectances, their associated uncertainties and error-correlation information. The reflectance in the L2A products is the Hemispherical-directional Reflectance Factor (HDRF) defined as: HDRF = π L / E where L is the directional upwelling radiance (with field of view of 5 degrees) and E is the (hemispherical) downwelling irradiance (i.e. including both direct solar and diffuse sky irradiance). These reflectances have dimensions of wavelength and series, where each series is a set of measurements for a given geometry (combination of viewing zenith and azimuth angle). In addition to variables for wavelength and bandwidth, the files also contain variables that provide for each series the acquisition time, viewing and solar angles, number of valid scans used, and quality flags (typically no flags are set in the data provided in this dataset). These NetCDF files also contain further relevant metadata as attributes. See https://hypernets-processor.readthedocs.io/ for further info.</p> <p>The WWUK site is a deciduous broadleaf forest comprised primarily of Oak, Hazel, Ash, Sycamore and Beech. It is located approximately 5 km North-West of Oxford, UK and has an extensive history of scientific research. The site follows the typical seasonal dynamics of a temperate forest with distinctive periods of leaf-off, green up and senescence across the growing season. The HYPERNETS site itself (51.777206 degrees N, 1.338494 W), is located at a height of 28 m upon a flux tower in the centre of the forest. The HYPSTAR®-XR sensor was installed in October 2021. Data are collected e very 30 minutes between 9am and 6pm local time between viewing zenith angles of 0 and 30 degrees.</p> <p>The HYPSTAR®-XR (eXtended Range) instruments deployed at each land HYPERNETS site consist of a VNIR and a SWIR sensor and autonomously collect data between 380-1700 nm at various viewing geometries and send it to a central server for quality control and processing. The VNIR sensor spans 1330 channels between 380 and 1000 nm with a FWHM of 3 nm and the SWIR sensor has 220 channels between 1000 and 1700 nm with a FWHM of 10 nm. The hypernets_processor (Goyens et al. 2021; De Vis et al. in prep.) automatically processes all this data into various products, including the L2A surface reflectance product provided here. All of the products have associated uncertainties (divided into random and systematic uncertainties, including error-correlation information) which were propagated using the CoMet toolkit (www.comet-toolkit.org). </p> <p>To obtain this dataset, we start from the full WWUK data record and omit all the data that do not pass all of the quality checks performed as part of the hypernets_processor. In addition, two additional screening procedures are developed to remove outliers and only supply the best quality data suitable for satellite validation. For Wytham wood, sequences are only supplied that match a typical vegetation spectrum. As such, data is only provided between April and October during the leaf-on period. Reflectances are then tested against three parameters to check that they are vegetation spectrum. Firstly, that there is a peak in the green portion of the visible wavebands (560 nm). Secondly, that a red edge is detected. Finally, the Normalized Difference Vegetation Index (NDVI) is calculated. Spectra with an NDVI of less than 0.42 are removed from the final data set.</p> <p>After the vegetation quality flag are applied, a sigma-clipping method is used to remove outliers. First reflectances are extracted in separate 2 hour windows throughout the day (to account for BRDF differences due to different solar position) for 4 different wavelengths (500, 900, 1100 and 1600 nm). Outliers in these reflectances are then identified by iteratively calculating the mean reflectance trend with time (by binning the data per maximum 30 data points), calculating the standard deviation from this trend, and masking any data that is more than 3 standard deviations away from the trend. This process is repeated on the unmasked data until the standard deviation does not vary by more than 5% between two iterations. The masks for the 4 different wavelengths are then combined (keeping only measurements for which none of the 4 wavelengths is an outlier). The reflectances and associated uncertainties for any masked series (i.e. a geometry that is masked either by the sigma-clipping procedure or from the masks of the hypernets_processor) are replaced by NaNs. Any sequence that has more than half of its series masked is removed entirely. </p>
Datasets and results from: "Random Forest Classification and Solar Flares Data: Analysis and Validation"
<p><strong>Instructions for the data and code repository</strong></p> <p>Results, post-processing workflow, and datasets for the research paper titled "Random Forest Classification and Solar Flares Data: Analysis and Validation".</p> <p>The folder contains three .csv files: the complete dataset (dataset.csv), the balanced training dataset (train_dataset.csv), and the testing dataset (test_dataset.csv).</p> <p>The folder also contains the result files from the research (.csv output files with predictions and .html files with evaluation metrics, etc.) exported from the JASP software. The number in each file name corresponds to the number of trees utilized in Random Forest modelling.</p> <p>In addition, the Python script for the post-processing workflow is provided, with comments located in the script.</p> <p>The soft range X-ray irradiance and VLF amplitude data were obtained from:<br> National Centers for Environmental Information (NCEI) Available online: https://www.ncei.noaa.gov/. Accessed on: 24th June 2023. <br> Worldwide archive of low-frequency data and observations (WALDO) Available online: https://waldo.world/. Accessed on: 24th June 2023.</p>
Dataset of experiment of wire-harness manufacturing framework validation
<p>Dataset created during the validation experiments of the wire harness manufacturing framework developed within the REMODEL European project.</p>
Plant regeneration in leaf culture of Centaurium erythraea Rafn. Part 3: de novo transcriptome assembly and validation of housekeeping genes for studies of in vitro morphogenesis
<p>Six centaury transcriptomes (embryogenic calli, globular somatic embryos, cotyledonary somatic embryos, adventitious buds, leaves and roots of <em>in vitro</em> grown plants) were sequenced and <em>de novo</em> assembled using <a href="https://github.com/trinityrnaseq/trinityrnaseq/wiki">Trinity</a> .</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly.tar.gz">CE_Assembly.tar.gz</a> - Centaury referent transcriptome comprises of 160.839 Trinity transcripts grouped in 105.726 Trinity genes.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly_fpkm.tar.gz">CE_Assembly_fpkm.tar.gz</a> - fpkm normalized read counts of the assembled transcripts in the six sequenced centaury tissues.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/nt.db_CE_assembly.tar.gz">nt.db_CE_assembly.tar.gz</a> - annotation of assembled transcripts by mapping them against NCBI nucleotide (NT) database using BLASTn . The obtained results were filtered with E-value E ≤ 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">swissprot.db_CE_assembly.tar.gz</a> - annotation of assembled transcripts by mapping them against NCBI nucleotide (<a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">s</a>wissprot) database using BLASTx . The obtained results were filtered with E-value E ≤ 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/pfam30.db_CE_assembly.tar.gz">pfam30.db_CE_assembly.tar.gz</a> - annotation of assembled transcripts by mapping them against Pfam30 domain database using hmmer3. The obtained results were filtered with independent E-value E ≤ 10<sup>-3</sup>.</p>
A Dataset of European Union Land Cover Validation Samples
<p>A dataset of European Union land cover validation samples in 2015 and 2010 based on the LUCAS micro dataset ( publicly available at <a href="https://ec.europa.eu/eurostat/web/lucas/data/lucas-grid">https://ec.europa.eu/eurostat/web/lucas/data/lucas-grid</a> ) . The dataset provides 9 land cover types of land cover including cropland, forest, grassland, shrubland, wetland, water, bareland, impervious surface and permanent snow/ice. The dataset is provided in .csv format.</p>
LAI_TS_Val: LAI time-series validation datasets in the 1-km pixel grid at global scale from 2001 to 2011
<p>Leaf area index (LAI), which is defined as one half of the total green leaf area per unit ground surface area, is a critical structural variable for quantifying the exchange processes of energy and matter between the land surface and atmosphere, it is thus identified as a key parameter in most terrestrial ecosystem models. To acquire long-term LAI records at the global scale, several remote sensing LAI products have been generated from various satellite sensors. However, assessing the uncertainties associated with these LAI products through comparisons with independent ground-truth measurements is pivotal for an effective application of products. Many sites from global networks have collected and provided invaluable ground LAI measurements covering a wide range of biome types and spatial variabilities. These site-based LAI measurements have been obtained about 30 years (1990-now). However, the spatial scale mismatch between site and pixel observations restricts the utilization of LAI measurements for product time-series validation. This datasets were generated from site-based LAI measurements of FLUXET and Chinese Ecosystem Research Network (CERN), using the proposed GUGM (Grading and Upscaling of Ground Measurements) method to resolve the scale-mismatch issue between site and sensor observations and maximize the utility of time-series of site-based LAI measurements, which can achieve the goal of product time-series validation. This GUGM approach first ingests both high-resolution images and site-based LAI measurements to capture the spatiotemporal variability in the product pixel grid. Then, a strategy was employed to grade the spatial representativeness of LAI measurements in the product pixel grid. For those LAI measurements which cannot be directly used in the validation of products, a strategy was adopted to calculate the spatial upscaling coefficient based on site-based LAI measurements and aggregated high-resolution reference maps to derive reliable LAI time-series validation datasets. The GUGM method has been applied to the site-based LAI measurements to generate global time-series LAI validation datasets from 2001 to 2011 in the 1 km pixel grid. The datasets include 28 sites which are mainly located in North America and Asia, providing 924 validation data in total. Among these sites, 16 sites with 508 (55.0%) validation data were obtained for forest, while 11 sites with 341 (36.9%) validation data and one site with 75 (8.1%) were obtained for crops and grasses, respectively. This datasets were saved in two formats: *.xls and *.kmz and each format was zipped for 63 KB and 31 KB, respectively.</p>
A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)
<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie "Forrest Gump" (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation) and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>
D3.3. Experimental Wave-Tank Validation Database
<p>This database is a deliverable of the FLOATECH project, funded under the European Union's Horizon 2020 research and innovation program under grant agreement No 101007142.</p><p>The aim of the accompanying document is to describe the experimental testing campaign C2 at the LHEEA wave-tank facility. The campaign took place between May and June 2023. The objective of the campaign is to test several FOWT control strategy, including a feed forward wave-based control, using the software-in-the-loop SOFTWIND system. The accompanying report on the E.U. portal describes the database created from these experiments and aimed to be shared for model validation. </p>
Validation dataset
<p>This dataset presents reported patient home communities and health facilities which they attended for malaria testing.</p>
Dataset of publication "Derivation and validation of a reference data-based real gas model for hydrogen"
<p>In this repository, a new real gas model for hydrogen based on the Reference Fluid Thermodynamic and Transport Properties Database (REFPROP) v10.0 is provided for the use in the simulation software OpenFOAM v2012. The model is valid in a temperature and pressure range of 150-400 K and 0.1-1000 bar, respectively. Usage beyond this range is not recommended as it may lead to unrealistic results.</p>
Auxiliary files and data to generate eddy flux and validate 2D model for MALTA
<p>This repository contains the following directories to accompany the manuscript 'A Zonally-Averaged Global Atmospheric Transport Model for Long-lived Trace Gases', submitted to JAMES:</p><p>1) <strong>GEOSChem </strong>This directory contains the run directory template and (slurm) runscript to generate the tracer fields used to generate the eddy fluxes. The GEOSChem model will have to be installed locally to run this, and the run directory built to your local area. It may be easiest to just copy the relevant bits in /Tracer_2D_template/ (i.e., the .rc files, /RestartFiles/, input.geos, reset_restart.py and species_database.yml) into a GEOSChem Transport run directory and change the directories in the copied files. If using slurm on an HPC, just change the directories in the runtracers_inputs.sh script to match that of your own HPC. Else, a different script will have to be written copying the slurm functionality.</p><p>2) <strong>GEOSChem_SF6 </strong>This directory contains the monthly mean SF6 mole fractions generated using GEOSChem used to validate the 2D model MALTA. Emissions come from the EDGAR v4.2 emissions inventory. Emissions after 2008 continue to use 2008 as the emissions value.</p><p>3) <strong>CFC11_inversion</strong> This directory contains the relevant script and files to quantify emissions of CFC-11 using an output mole fraction from the TOMCAT 3D model using MALTA, and compare these to the TOMCAT emissions used to generate the mole fractions. The directory paths at the beginning of the main script in CFC11_inversion.py must be changed to point to the remaining files in the /CFC11_inversion/ directory, and a save directory must be specified, before running locally. MALTA must be installed to run this.</p><p>4) <strong>singapore.dat </strong>This file contains the QBO winds above Singapore, taken from https://www.geo.fu-berlin.de/en/met/ag/strat/produkte/qbo/index.html</p><p> </p>
Spectrophotometric Assay for the Detection of 2,5-Diformylfuran and Its Validation through Laccase-Mediated Oxidation of 5-Hydroxymethylfurfural
<p>Modern biocatalysis requires fast, sensitive, and efficient high-throughput screening methods to screen enzyme libraries in order to seek out novel biocatalysts or enhanced variants for the production of chemicals. For instance, the synthesis of bio-based furan compounds like 2,5-diformylfuran (DFF) from 5-hydroxymethylfurfural (HMF) via aerobic oxidation is a crucial process in industrial chemistry. Laccases, known for their mild operating conditions, independence from cofactors, and versatility with various substrates, thanks to the use of chemical mediators, are appealing candidates for catalyzing HMF oxidation. Herein, Schiff-based polymers based on the coupling of DFF and 1,4-phenylenediamine (PPD) have been used in the set-up of a novel colorimetric assay for detecting the presence of DFF in different reaction mixtures. This method may be employed for the fast screening of enzymes (Z' values ranging from 0.68 to 0.72). The sensitivity of the method has been proved, and detection (8.4 μM) and quantification (25.5 μM) limits have been calculated. Notably, the assay displayed selectivity for DFF and enabled the measurement of kinetics in DFF production from HMF using three distinct laccase–mediator systems.</p>
Qiime2 classifiers (rbcl, Mollusc 18s) for testing the validity of using eDNA for carbon origin analysis from sediment cores
<p>Qiime2 formatted classifiers that were created for a Natural England funded project by researchers at the James Hutton Institute. The pilot project aims to test the validity of using eDNA for carbon origin analysis from sediment cores. These classifiers for the rbcl and 18 Mollusc genes were made using RESCRIPt and Qiime2. </p> <p>The scripts used to created these classifiers are available at the James Hutton ICS GitHub <a href="https://github.com/HuttonICS/blue-carbon-db">blue-carbon-db</a> . The files are as follows:</p> <p><a href="../api/records/10046481/draft/files/mollusc-espineira-classifier.qza/content" target="_blank" rel="noopener noreferrer">mollusc-espineira-classifier.qza</a> is a classifer built from ncbi 18s Mollusc sequences, trained on the primer set from Espiñeira et al (2009).</p> <div>rbcl-vasselon-zimmerman-F3-R1-classifier.qza is a classifer built from ncbi rbcl sequences, trained on the F3 and R1 primer set fromVasselon et al (2017).</div> <p> </p> <p> </p> <p><strong>Important: </strong>If you use these classifiers please be aware of the process used to create them and be sure to review the methods. These databases were created by downloading data from the NCBI in October 2023, sequence data available at the NCBI changes over time. To create the most up to date database a fresh download and re-evaluations of the databases would be preferable. All method and scripts can be found at <a href="https://github.com/HuttonICS/blue-carbon-db">blue-carbon-db </a></p> <p>If you use these database please reference this repository along with RESCRIPt and Qiime2 </p> <p> </p> <p>Espiñeira, M., González-Lavín, N., Vieites, J. M. and Santaclara, F. J. 2009 Development of a method for the genetic identification of commercial bivalve species based on mitochondrial 18S rRNA sequences. J Agric Food Chem, 28, 495-502 https://doi.org/10.1021/jf802787d</p> <p> </p> <p>Vasselon, V., Rimet, F., Tapolczai, K. and Bouchez, A. 2017. Assessing ecological status with diatoms DNA metabarcoding: Scaling-up on a WFD monitoring network (Mayotte island, France). Ecological Indicators, 82, 1-12 <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.ecolind.2017.06.024" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.ecolind.2017.06.024</a></p>
global interpreted planted forest, natural forests validation samples
<p>This dataset provided the global validation dataset including planted forest, natural forest, and non-forest in 2015. This dataset was visually interpreted using the high spatial resolution (<1 m) images from Google Earth, and combining the spatial distribution of planted forets, and forest gain map.</p> <p>1 denotes planted forest, 2 denotes the natural forest, and 3 denotes the nonforest.</p> <p>Detailed information about how the global validation dataset was visually interpreted can be seen in the following reference:</p> <p>Xu, H., He, B., Guo, L., Yan, X., Zeng, Y., Yuan, W., et al. (2024). Global forest plantations mapping and biomass carbon estimation. Journal of Geophysical Research: Biogeosciences, 129, e2023JG007441.</p>
Data from: Validating a practical methodology for thatch - mat - soil distinction in turfgrass soils
<p><span>We described a <a name="_Hlk163223211"></a>practical method for the distinction of thatch, mat and soil layers in turfgrass soils, being a combination of visually and manually observable characteristics. For two widely different turfgrass species, we analyzed total organic matter (TOM) and dry bulk density (ρd) in thin slices of 6 mm (between 0 and 10 cm soil depth), resulting in clear patterns for both soil properties with increasing depth. Statistical analysis of TOM patterns resulted in similar boundary depths between calculated and observed layers, validating our practical method for the distinction of thatch, mat and soil layers as a reliable method. </span><span>Furthermore, we characterized thatch, mat and soil layer by different TOM fractions. <span>TOM was fractionalized into three distinctive and functional pools of organic matter: (1) visible organic matter (VOM), consisting of mainly non-decomposed plant structures, (2) decomposed organic matter (DOM), consisting of mainly decomposed plant structures with its associated microbial biomass, and (3) soil organic matter (SOM), being the background value or recalcitrant native organic matter in a soil and its local microbial biomass.</span></span></p> <p><span><span>We distinguished thatch, mat and soil layer based on visual and manual observable characteristics of the layers and a protocol as described in Evers et al. (2024) https://doi.org/10.1002/its2.148. This study was conducted on a well-established turfgrass demonstration field with monoculture plots of turfgrass varieties (turfgrass seed company DLF; Moerstraten, the Netherlands; 51°32'27'' N 4°20'54" E) 3.5 years after sowing, reflecting the result of organic matter accumulation over these initial years of turfgrass establishment. The climate regime was marine with cool to medium summer temperatures and mild winters (Cfb/Cfa according to the Köppen-Geiger climate classification system (Peel, et al., 2007)). The field was built on a sandy soil (Hortic Anthrasol as described in the FAO/UNESCO soil map of the world (2006)). Sampling of the soil took place in June 2016 before the field was sown. We tested our methodology with two turfgrass species, slender creeping red fescue (<em>Festuca rubra trichophylla </em>(Frt), variety Beudin of DLF) <span>as an example of a spreading turfgrass</span><em> </em><span>and perennial ryegrass (<em>Lolium perenne </em></span>(Lp)<em>,</em> variety Duparc of DLF) as an example of a bunch-type grass, as these two species were expected to differ widely in thatch and mat depth. The individual plot size for each variety was approximately 1 m<sup>2</sup> (0.8 x 1.2 m). Plots of each species were sampled at the end of February 2020. A subplot of 0.25 m<sup>2</sup> (0.5 m x 0.5 m) in the center of one plot per species was selected, to avoid contamination with other varieties (at least 90% pure monoculture), and it was marked with a metal frame. From the 25 cells of 5 x 5 cm in this frame, nine evenly dispersed cells were chosen to take a set of soil samples of 10 cm depth, using a core sampler with 2.8 cm diameter, for thatch-mat and mat-deeper soil boundary observation and TOM and <em>ρ</em><sub>d</sub> analyses. A<span>nother set of nine samples was taken next to the previous cells for VOM analyses. </span>Every fresh 10 cm soil core was first photographed (Canon Powershot S5 camera, 24 megapixels; Canon Europe, Amstelveen, the Netherlands), judged on thatch-mat and mat-deeper soil boundaries following the method described in supplementary information, and then sliced into 14 subsamples of 6 mm each plus a 16 mm subsample at the bottom, starting to measure just below the green canopy with an accurate ruler (Sola HK ¼ W12, EU-accuracy class 3), for analyzing TOM and <em>ρ</em><sub>d</sub>.in every subsample. TOM and <em>ρ</em><sub>d</sub> were analyzed after samples were dried at 105°C for 24 h. VOM was analyzed after a<span>ll sediment per slice was carefully washed off with tap water in a fine sieve (approximately 600 µm (27 mesh)), after which the remaining (dead and living) plant biomass, mainly roots and rhizomes, was dried at 65 °C for at least 48 h. SOM was determined separately in the bulk soil of the study site and is 2.6% of the dry matter. DOM was calculated via subtraction of VOM and SOM from TOM</span></span></span></p> <p><span>Based on the analyzed TOM content of the 14 slices per nine replicates, the boundaries of distinctive layers were calculated. For this, we rescaled the TOM results per replicate via the normalized function <em>F </em>(χ) = (χ-χ<sub>min</sub>)/(χ<sub>max</sub>-χ<sub>min</sub>) into values between 0 and 1 to overcome scale differences between replicates but keeping distributions the same. Per turfgrass species, the best fitted line, i.e., the smallest <em>rse</em>, and its 95%-confidence interval through all 135 points was iteratively calculated by non-linear least square regression with the nls function of R (version 3.5.2; 2018-12-20). For the fitting of the statistical models, either a logistic function (<em>F </em>(χ) = α/(1+e<sup>-(βχ+γ)</sup>) or a bell-shaped Gaussian function (<em>G </em>(χ) = αe<sup>-((χ-β)^2/2γ^2)</sup>) was used, based on the best fit for the respective turfgrass species. The characteristics of these mathematical functions were used to explain the TOM-dynamics in the soil. To this end, the turning points of the curves were calculated by finding where the first derivative of both functions equals zero, i.e., solving <em>F’</em> (χ) = 0 and <em>G’ </em>(χ) = 0 respectively. The inflection points of each curve were determined by analyzing the second derivative, i.e., identifying the points where the second derivative <em>F’’</em> (χ) or <em>G’’</em> (χ) = 0, indicating changes in concavity. The inflection points and turning points indicated changes in TOM content in the turfgrass soil profile as likely boundaries between distinctive soil layers. </span><span>Differences in parameters between distinctive soil layers and turfgrass species were determined based on the calculated means of parameters per nine replicates of soil slices Normality of residuals and the equality of variances was checked with diagnostic plots and Levene’s test, respectively. Normally distributed means were compared with one-way ANOVA for a 3-layered soil system, followed by either Tukey post hoc tests in case of equality of variance, or by the Games-Howell post hoc test in case of no equality of variance, or with a one sample t-test for a 2-layered soil system. </span></p>
Extended dataset for the validation the competent Computational Thinking test in grades 3-6
<p>Extended dataset for the validation the competent Computational Thinking test in grades 3-6<br>=======================================================</p> <p>• If you publish material based on this dataset, please cite the following :</p> <p> • The Zenodo repository : Laila El-Hamamsy, Barbara Bruno, Jessica Dehler Zufferey, & Francesco Mondada (2023). Extended dataset for the validation of the competent Computational Thinking test in grades 3-6 [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7983525 </p> <p> • The article on the validation of the computational thinking test for grades 3-6 : El-Hamamsy, L., Zapata-Cáceres, M., Martín-Barroso, E., Mondada, F., Zufferey, J. D., Bruno, B., & Román-González, M. (2025). The competent Computational Thinking test (cCTt): A valid, reliable and gender-fair test for longitudinal CT studies in grades 3–6. <em>Technology, Knowledge and Learning</em>, 1-55. https://doi.org/10.1007/s10758-024-09777-8 </p> <p>• License : This work is licensed under a Creative Commons Attribution 4.0 International license (CC-BY-4.0)</p> <p>• Creators : El-Hamamsy, L., Bruno, B., Dehler Zufferey, J., and Mondada, F.</p> <p>• Date May 30th 2023</p> <p>• Subject : Computational Thinking (CT), Assessment, Primary education, Psychometric validation</p> <p>• Dataset format : CSV. The dataset contains four files (one per grade, see detailed description below). Please note that the spreadsheets may contain missing values due to students not being present for a part of the data collection. To have access to the specific cCTt questions please refer to the original publication [1] and Zenodo repository [2] which provide the full set of questions and correct responses.</p> <p>• Dataset size < 500 kB</p> <p>• Data collection period : January and November 2021</p> <p>• Abbreviations :<br> - CT : Computational Thinking<br> - cCTt: competent CT test</p> <p>• Funding : This work was funded by the the NCCR Robotics, a National Centre of Competence in Research, funded by the Swiss National Science Foundation (grant number 51NF40_185543)</p> <p># References</p> <p>[1] El-Hamamsy, L., Zapata-Cáceres, M., Barroso, E. M., Mondada, F., Zufferey, J. D., & Bruno, B. (2022). The Competent Computational Thinking Test: Development and Validation of an Unplugged Computational Thinking Test for Upper Primary School. Journal of Educational Computing Research, 60(7), 1818–1866. https://doi.org/10.1177/07356331221081753 </p> <p>[2] El-Hamamsy, L., Zapata-Cáceres, M., Marcelino, P., Dehler Zufferey, J., Bruno, B., Martín Barroso, E., & Román-González, M. (2022). Dataset for the comparison of two Computational Thinking (CT) test for upper primary school (grades 3-4) : the Beginners' CT test (BCTt) and the competent CT test (cCTt) (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5885034 </p> <p>[3] El-Hamamsy, L., Zapata-Cáceres, M., Martín-Barroso, E. <em>et al.</em> The Competent Computational Thinking Test (cCTt): A Valid, Reliable and Gender-Fair Test for Longitudinal CT Studies in Grades 3–6. <em>Tech Know Learn</em> (2025). https://doi.org/10.1007/s10758-024-09777-8</p> <p>[4] Brennan, K. and Resnick, M. (2012). New frameworks for studying and assessing the development of computational thinking. page 25</p> <p>[5] El-Hamamsy, L., Zapata-Cáceres, M., Marcelino, P., Bruno, B., Dehler Zufferey, J., Martín-Barroso, E., & Román-González, M. (2022). Comparing the psychometric properties of two primary school Computational Thinking (CT) assessments for grades 3 and 4: The Beginners’ CT test (BCTt) and the competent CT test (cCTt). Frontiers in Psychology, 13. https://www.frontiersin.org/articles/10.3389/fpsyg.2022.1082659</p>
Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper
<p>Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper (complement for a tutorial at https://github.com/datngu/LmTag)</p> <p>This repo includes chromosome 10 reference panel data constructed for 3 populations:</p> <p>- EAS</p> <p>- EUR</p> <p>- SAS</p>
Database for the validation of the French COHIP-SF-19 among 12-year-olds in New Caledonia
<p>Extract from the New Caledonian epidemiological survey database for the validation of the French COHIP-SF-19 among 12-year-olds </p>
Simulated X-ray micro-computed tomography based particle tracking velocimetry dataset for validation purposes
<p>Authors: Tom Bultreys, Stefanie Van Offenwert, Wannes Goethals, Matthieu N. Boone, Jan Aelterman and Veerle Cnudde; Ghent University (Belgium)<br> Date: 8th February 2022<br> For any usage, please cite the accompanying publication: T. Bultreys, S. Van Offenwert, W. Goethals, M. N. Boone, J. Aelterman and V. Cnudde, "X-ray Tomographic Micro-Particle Velocimetry in Porous Media", Physics of Fluids, 34, 042008 (2022).<br> https://doi.org/10.1063/5.0088000<br> -----------------------------</p> <p>Validation dataset for micro-computed tomography based particle tracking velocimetry: a simulated micro-CT based velocimetry experiment with associated ground-truth particle trajectories</p> <p>- The ground truth trajectories were based on randomly dropping virtual particles in the pore space, and tracking their movement through a CFD-based velocity field (see below). The positions were calculated for the time corresponding to each radiograph of a micro-CT experiment. The folder "GroundTruthData" contains the locations of all particles at the central time of each micro-CT scan, as well as their radii. Check the associated readme file to read the data file.</p> <p>- The main data is contained in the directory "TimeFrames", containing the reconstructed 3D images at 7 time steps (70 seconds interval), with a voxel size of 11.8 µm, in 3D .tif format. This can be opened in for example Fiji/ImageJ.</p> <p>- The directory "clearFrame" contains an image of the pore space without particles, matching with the time frame images, in the same format and with the same voxel size as the time frame images.</p> <p>- The directory "SegmentedImage" contains two binary 3D images (same format as images before) which was created by segmenting the clearFrame image. There are two versions: the original segmentation, and a version where pores were eroded. The eroded segmentation was used to mask the pore space during particle detection (this avoids spurious detections near pore walls, caused by minor mis-alignments of the clearImage).</p> <p>- The original segmentation was used as input to simulate the velocity fields in the directory "simulatedVelocityFields", which contains 3D .tif images that represent the three components of the velocity vector field (the X-direction was the axis of the sample, equaling the flow direction). There is also an input text file and an output text file. The simulation was performed with the code from single-phase OpenFOAM implementation from Ali Raeini and others at Imperial College London: http://www.imperial.ac.uk/earth-science/research/research-groups/perm/research/pore-scale-modelling/</p> <p>- The trackingOutput folder contains the experimentally determined velocity points (.csv, only particles that could be tracked at least 6 time frames) and the experimentally determined velocity magnitude field (.tif, voxel size 23.6 µm)</p>
Copper Rings Insertion Validation Dataset
<p>The dataset consists of color images of the fixture with inserted copper sliding rings, which is used to evaluate and validate the process of assembly of an object with low tolerances supported by multi modal exception strategy learning and ergodic control. This dataset is used to classify the insertion process into three states: OK, NotOK or NoPart. The dataset consists of two main classes: Valid Insertion, Invalid Insertion in each of the four insertion slots.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.