Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
447
datasets available to search
ShareScore release 0.9.0
Dataset results
447 results for “Model validation”
Community science validates climate suitability projections from ecological niche modeling
<p><span>Climate change poses an intensifying threat to many bird species, and projections of future climate suitability provide insight into how species may shift their distributions in response. Climate suitability is characterized using ecological niche models (ENMs), which correlate species occurrence data with current environmental covariates and project future distributions using the modeled relationships together with climate predictions. Despite their widespread adoption, ENMs rely on several assumptions that are rarely validated <i>in situ </i>and can be highly sensitive to modeling decisions, precluding their reliability in conservation decision-making. Using data from a novel, large-scale community science program, we developed dynamic occupancy models to validate near-term climate suitability projections for bluebirds and nuthatches in summer and winter. We estimated occupancy, colonization, and extinction dynamics across species' ranges in the United States in relation to projected climate suitability in the 2020s, and used a Gibbs variable selection approach to quantify evidence of species-climate relationships. We also included a Bird Conservation Region strata-level random effect to examine among-strata variation in occupancy that may be attributable to land-use and ecoregional differences. Across species and seasons, we found strong evidence that initial occupancy and colonization were positively related to 2020 climate suitability, illustrating an independent validation of projections from ENMs across a large geographic area. </span><span>Random strata effects revealed that occupancy probabilities were generally higher than average in core areas and lower than average in peripheral areas of species' ranges, and served as a first step in identifying spatial patterns of occupancy from these community science data. </span><span>Our findings lend much-needed support to the use of ENM projections for addressing questions about potential climate-induced changes in species' occupancy dynamics. More broadly, </span>our work highlights the value of community scientist observations for ground-truthing projections from statistical models and for refining our understanding of the processes shaping species' distributions under a changing climate.</p>
Data from: Development and validation of an expert-based habitat suitability model to support boreal caribou conservation
The long-term persistence of boreal caribou (Rangifer tarandus caribou) is threatened by the negative impacts of human activities, including industrial development. The vast geographical distribution and behavioural variability of caribou warranted the development and validation of an appropriate conservation tool for managers in eastern Canada. We developed a habitat suitability model for boreal caribou in Québec by integrating expert knowledge into a hierarchical analysis. We elicited responses from 14 experts on caribou ecology to determine the best spatiotemporal scales of application, the relative importance of habitat variables, the zone of influence of human infrastructure, and the parameterization of the model. Based on this input, we built the model using 8 habitat categories and 3 human infrastructure variables. Experts identified mature conifer-dominated forests and open lichen woodlands as the most important vegetation categories for boreal caribou, whereas density of and proximity to paved roads, forest roads, and mines decreased habitat quality. We mapped the resulting model over the entire province of Québec (up to the northern forest allocation limit), and validated it using independent GPS telemetry datasets acquired in 3 distinct regions. Our model predicted that habitat suitability was highest in the northeastern part of our study area, where timber harvest activities and roads were virtually absent. Conversely, southern parts of Québec were generally unsuitable for boreal caribou. Our habitat suitability model is among the first tools for boreal caribou conservation available to wildlife and land managers in Québec.
Data from: Testing the validity of functional response models using molecular gut content analysis for prey choice in soil predators
Analysis of predator - prey interactions is a core concept of animal ecology, explaining structure and dynamics of animal food webs. Measuring the functional response, i.e. the intake rate of a consumer as a function of prey density, is a powerful method to predict the strength of trophic links and assess motives of prey choice, particularly in arthropod communities. However, due to their reductionist set-up, functional responses, which are based on laboratory feeding experiments, may not display field conditions, possibly leading to skewed results. Here, we tested the validity of functional responses of centipede predators and their prey by comparing them with empirical gut content data from field-collected predators. Our predator - prey system included lithobiid and geophilomorph centipedes, abundant and widespread predators of forest soils and their soil-dwelling prey. First, we calculated the body size-dependent functional responses of centipedes using a published functional response model in which we included natural prey abundances and animal body masses. This allowed us to calculate relative proportions of specific prey taxa in the centipede diet. In a second step, we screened field-collected centipedes for DNA of eight abundant soil-living prey taxa and estimated their body size-dependent proportion of feeding events. We subsequently compared empirical data for each of the eight prey taxa, on proportional feeding events with functional response-derived data on prey proportions expected in the gut, showing that both approaches significantly correlate in five out of eight predator - prey links for lithobiid centipedes but only in one case for geophilomorph centipedes. Our findings suggest that purely allometric functional response models, which are based on predator-prey body size ratios are too simple to explain predator - prey interactions in a complex system such as soil. We therefore stress that specific prey traits, such as defence mechanisms, must be considered for accurate predictions.
Observed monthly N2O emission dataset used for model carlibaration and validation
<p>The N<sub>2</sub>O emission dataset for calibration sites was extracted from the published figures and tables using GetData Graph Digitizer version 2.24; the other information such as biome, geographic location, experimental period, soil organic carbon content, soil pH and soil texture was selected from corresponding literature. If the data related to soil was not avaliable, we extracted them from the soil database(IGBP-DIS)</p>
A field-validated ensemble species distribution model of Eriogonum pelinophilum, an endangered subshrub in Colorado, USA
<p>Understanding the suitable habitat of endangered species is crucial for agencies such as the Bureau of Land Management to plan management and conservation. However, few species distribution models are directly validated, potentially limiting their application. In preparation for a Species Status Assessment of clay‐loving wild buckwheat (<em>Eriogonum pelinophilum</em>), an endangered subshrub found in southwest Colorado, we ran a series of species distribution models to estimate the species' potential occupied habitat and validated these models in the field. A 1‐meter resolution digital elevation model derived from LiDAR and a high‐resolution geology mapping helped identify biologically relevant characteristics of the species' habitat. We employed a weighted ensemble model based on two Random Forest and one Boosted Regression Tree model, and the discrimination performance of the ensemble model was high (AUC-PR = 0.793). We then conducted a systematic field survey of model habitat suitability predictions, during which we discovered 55 new subpopulations of the species and demonstrated that new species observations were strongly associated with model predictions (p < .0001, Cliff's delta = 0.575). We then further refined our original models by incorporating the additional species occurrences collected in the field survey, a new explanatory variable, and a more diverse set of models. These iterative changes to the model marginally improved performance (AUC‐PR = 0.825). Direct validation of species distribution models is extremely rare, and our field survey provides strong validation of our model results. This helps increase confidence in utilizing predictions in planning. The final model predictions greatly improve the Bureau of Land Management's understanding of the species' habitat and increase our ability to consider potential habitat in planning land use activities such as road development and travel management.</p>
Model Output and Validation Data for Online Determination of GNSS Differential Code Biases using Rao-Blackwellized Particle Filtering
<p>This dataset contains the outputs of 10 test runs of the A-CHAIM data assimilation model used to evaluate the performance of the bias estimation procedure used by the model. The 5-minute output files for each test run are included in a separate folder.</p> <p>The IGS DCBs in SINEX format and the various global ionospheric maps used in the study are included.</p> <p>This dataset also includes all of the processed GNSS data used in the study, which were used to generate the DCBs used in each test run.<br> <br> The in-situ electron density measurements from DMSP as well as the ionosonde data from GIRO which were used to measure the performance are also included.</p>
Model Development for State-of-Power Estimation of Large-Capacity Nickel-Manganese-Cobalt Oxide-Based Lithium-Ion Cell Validated Using a Real-Life Profile
Open the record for dataset details and reuse information.
Validation of the predictive accuracy of health-state utility values based on the Lloyd model for metastatic or recurrent breast cancer in Japan
<p>Although there is a lack of data on health-state utility values (HSUVs) for calculating quality-adjusted life years in Japan, Cost-utility analysis has been introduced by the Japanese government to inform decision-making in the medical field since 2016. This study aimed to determine whether the Lloyd model which was a predictive model of HSUVs for metastatic breast cancer (MBC) patients in the United Kingdom can accurately predict actual HSUVs for Japanese patients with MBC. The prospective observational study, followed by the validation study of the clinical predictive model.<b> </b>Forty-four Japanese patients with MBC were studied at 336 survey points. This study consisted of two phases. In the first phase, we constructed a database of clinical data prospectively and HSUVs for Japanese patients with MBC to evaluate the predictive accuracy of HSUVs calculated using the Lloyd model. In the second phase, Bland-Altman analysis was used to determine how accurately predicted HSUVs (based on the Lloyd model) correlated with actual HSUVs obtained using the EuroQol 5-Dimension 5-Level questionnaire, a preference-based measure of HSUVs in patients with MBC. In the Bland-Altman analysis, the mean difference between HSUVs estimated by the Lloyd model and actual HSUVs, or systematic error, was -0.106. The precision was 0.165. The 95% limits of agreement ranged from -0.436 to 0.225. The t value was 4.6972, which was greater than the t value with 2 degrees of freedom at the 5% significance level (p=0.425). There were acceptable degrees of fixed and proportional errors associated with the prediction of HSUVs based on the Lloyd model for Japanese patients with MBC. We recommend that sensitivity analysis be performed when conducting cost-effectiveness analyses with HSUVs calculated using the Lloyd model.</p>
Artifacts for the SPLC 2020 Paper "A Conceptual Model for Unifying Variability in Space and Time" and the ESE journal extension "A Conceptual Model for Unifying Variability in Space and Time: Rationale, Validation, and Illustrative Applications"
<p>These artifacts relate to the SPLC'20 research paper "A Conceptual Model for Unifying Variability in Space and Time" and the Empirical Software Engineering journal extension "A Conceptual Model for Unifying Variability in Space and Time: Rationale, Validation, and Illustrative Applications."</p>
Validating x-ray line-profile defect analysis using atomistic models of deformed material
<p>Data and analysis (scripts and Jupyter notebooks) associated with publications:</p> <p>Validating x-ray line-profile defect analysis using atomistic models of deformed material (submitted to Physical Review Materials)</p> <p>Validation of x-ray line-profile analysis of extended defects in deformed crystals (submitted to Physical Review Letters)</p> <p>Preprints of these papers are included here</p>
Validation Data used for manuscript "Climate Projections over the Great Lakes Region: Using Two-way Coupling of a Regional Climate Model with a 3-D Lake Model"
<p>those are the processed data that used for model-data comparison in the manuscript "Climate Projections over the Great Lakes Region: Using Two-way Coupling of a Regional Climate Model with a 3-D Lake Model", including Lake Surface Temperature and Lake Surface Ice Cover from Great Lakes Surface Environmental Analysis (GLSEA), Surface Air temperature and Precipitation from Climatic Research Unit (CRU). </p>
Datasets and model outputs used for validating WASP for SURFEX v8.0
<p>These datasets contain the model outputs used for validating the WASP SURFEX v8.0 parameterization of turbulent fluxes at sea. The file LION_BUOY.tar contains (ascii format) the tubrulent fluxes computed from the surface parameters recorded at the Météo-France LION buoy (in the centre of the Gulf of Lion, NW Mediterranean Sea) at an hourly frequency between 2001 and 2014, using the bulk softwares COARE 3.0 with various wave representation and WASP with waves. The file AROME_MED.tar contains the outputs of the AROME operational model with ECUME or WASP representation of the turbulent fluxes, on the case study of October 2016, that were used to validate the parameterization as presented in the paper. The file AROME_OI.tar contains the outputs of AROME OI on the case of the tropical cyclone Batsirai in the Indian Ocean, in February 2022. The file ARPEGE_CLIM_AMIP contains the outputs of the low-resolution, AMIP runs performed using the ECUME and WASP parameterization.</p>
Validation of Peak Ground Velocities Recorded on Very-high rate GNSS Against NGA-West2 Ground Motion Models: Dataset
<p>Contains all the data files for the manuscript of the same name.</p>
Experimental data for validation of a single-storey flexible double-skin façade system model
<p>Double skin façades are adaptive envelopes aiming at improving building energy use and comfort performance. Their adaptive principle relies on the dynamic management of the cavity’s ventilation flow and the shading device (when available). They can also be integrated with the environmental systems for heating, cooling, and ventilation. In most cases, though, the possible exploitation of the ventilation airflow is not fully enabled, as the adoption of only one or two possible airpath limits the possibility that this façade architecture offers and flexible interaction with the environmental systems is not planned. This work aims to develop, using an existing software tool for building energy simulation, a numerical model of a flexible double-skin façade module capable of fully exploiting the adaptive features of such envelope concept by switching between different cavity ventilation strategies. Leveraging on the Double Glass Facade component available in IDA ICE, a new model for a flexible double-skin façade module was developed, and its performance in replicating the thermophysical behaviours of such a dynamic system has been assessed by comparison with experimental data collected through a dedicated experimental activity using one the outdoor test cells of the TWINS facility in Torino (Italy). The accuracy of the predictions resulted in line with the performance obtained by the Double Glass Facade component to simulate conventional double-skin facades. By establishing a new archetype model to study the performance and optimal integration of a large class of double-skin façade modules, including fully flexible ones, this works demonstrated the possibility of modifying existing models in building energy simulation tools to study unconventional building envelope model solutions such as adaptive façade systems.</p>
Results and Validation of Global-Regional Nested model for polycyclic aromatic hydrocarbons.
<p>The repository is used to store data from the results and collected observations in the submitted paper "Modeling PAHs from Global to Regional Scales: Model Development and Investigation of Health risks in China 2013-2018," which was published in the Earth Science Model Development (GMD) journal.</p>
Four psychometrically validated datasets for benchmarking large language models, based on the TIMSS 2008 and 2011 released items.
<p>Four datasets validated according to psychometric principles that can be used to benchmark large language models in terms of achievements in advanced school math, advanced school physics, 8th grade math and 8th grade science.</p> <p>These four datasets are derived from items released by Trends in International Mathematics and Science Study Advanced 2008 and Trends in International Mathematics and Science Study 2011. See <a href="https://nces.ed.gov/timss/released-questions.asp">link</a>.</p> <p>For more information, see our paper <a href="https://arxiv.org/abs/2404.01799">PATCH! Psychometrics-AssisTed benCHmarking of Large Language Models: A Case Study of Mathematics Proficiency</a>.</p>
Outputs from fitted models across the cross-validation scenarios for 'Space-time species distribution modeling with opportunistic presence-only data: a case study of passerines in a protected area'
<p>Three Zenodo repositories are linked to the preprint <em>Space-time Species Distribution Modeling for Opportunistic Presence-Only Data: A Case Study of Passerines in a Protected Area </em>(Lasgorceux et al., unpublished, <a href="https://hal.science/hal-04616332">https://hal.science/hal-04616332</a>):</p> <ul> <li>Data, scripts and, code (Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12545052">https://doi.org/10.5281/zenodo.12545052</a>)</li> <li>Outputs from fitted models across the cross-validation scenarios (Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12544212">https://doi.org/10.5281/zenodo.12544212</a>)</li> <li>Supplementary information at (Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12541412">https://doi.org/10.5281/zenodo.12541412</a>)</li> </ul> <p>This repository contains the outputs from fitted models across the cross-validation scenarios.</p> <p>In the folder <em>Ouputs_cross_validation</em>, each species is represented by a .RData file, numbered from 1 to 77 (excluding 7, which corresponds to <em>Bombycilla garrulus</em>; see the preprint for details). This dataset is specifically used to generate Figure 1, which shows the AUC of various cross-validation scenarios. To reproduce this figure in R, place all the files in the <em>Results/Fitted_models</em> folder and run the <em>Models_Outputs.R</em> script located in the <em>Results</em> folder of Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12545052">https://doi.org/10.5281/zenodo.12545052.</a></p> <p>Note: These data have been separated due to memory requirements (23.14GB).</p>
RNA 3D structural models used to train, test and validate lociPARSE
<p>This repository contains all the training, validation and test decoy sets to train and evalaute lociPARSE. It also contains training and benchmarks set-2 decoys from ARES.</p>
Validation and Stabilization of a Prophage Lysin of Clostridium perfringens by Using Yeast Surface Display and Coevolutionary Models
<p>Bacteriophage lysins are compelling antimicrobial proteins whose biotechnological utility and evolvability would be aided by elevated stability. Lysin catalytic domains, which evolved as modular entities distinct from cell wall binding domains, can be classified into one of several families with highly conserved structure and function, many of which contain thousands of annotated homologous sequences. Motivated by the quality of these evolutionary data, the performance of generative protein models incorporating coevolutionary information was analyzed to predict the stability of variants in a collection of 9,749 multimutants across 10 libraries diversified at different regions of a putative lysin from a prophage region of a Clostridium perfringens genome. Protein stability was assessed via a yeast surface display assay with accompanying highthroughput sequencing. Statistical fitness of mutant sequences, derived from secondorder Potts models inferred with different levels of sequence homolog information, was predictive of experimental stability with areas under the curve (AUCs) ranging from 0.78 to 0.85. To extract an experimentally derived model of stability, a logistic model with site-wise score contributions was regressed on the collection of multimutants. This achieved a cross-validated classification performance of 0.95. Using this experimentally derived model, 5 designs incorporating 5 or 6 mutations from multiple libraries were constructed. All designs retained enzymatic activity, with 4 of 5 increasing the melting temperature and with the highest-performing design achieving an improvement of 4°C.</p>
A multi-species benchmark for training and validating large scale mass spectrometry proteomics machine learning models
<p>This is a de novo sequencing benchmark dataset derived from nine<br>publicly available mass spectrometry datasets. There are two versions<br>of the benchmark: main and balanced. The balanced version randomly<br>eliminates some spectra associated with some species in order to<br>create a smaller, more evenly balanced dataset. Also provided are two<br>zip files containing the raw data as well as intermediate results.<br>Details about how the benchmark was created are provided in an<br><a href="../records/13653420">associated zenodo release</a>, which contains the source code as well as a<br>manuscript describing the benchmark.</p> <p>This release fixes a bug that incorrectly detected shared peptides <br>between different species. It also includes the annotated spectra in <br>mzSpecLib format.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.