Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data from: a physics-based digital twin for model predictive control of autonomous unmanned aerial vehicle landing
<p>This paper proposes a two-level, data-driven, digital twin concept for the autonomous landing of aircraft, under some assumptions. It features a digital twin instance for model predictive control; and an innovative, real-time, digital twin prototype for fluid-structure interaction and flight dynamics to inform it. The latter digital twin is based on the linearization about a pre-designed glideslope trajectory of a high-fidelity, viscous, nonlinear computational model for flight dynamics; and its projection onto a low-dimensional approximation subspace to achieve real-time performance, while maintaining accuracy. Its main purpose is to predict in real-time, during flight, the state of an aircraft and the aerodynamic forces and moments acting on it. Unlike static lookup tables or regression-based surrogate models based on steady-state wind tunnel data, the aforementioned real-time digital twin prototype allows the digital twin instance for model predictive control to be informed by a truly dynamic flight model, rather than a less accurate set of steady-state aerodynamic force and moment data points. The paper describes in detail the construction of the proposed two-level digital twin concept and its verification by numerical simulation. It also reports on its preliminary flight validation in autonomous mode for an off-the-shelf unmanned aerial vehicle instrumented at Stanford University.</p>
Economic modeling data of Egypt
<p>The data that support the findings of this study are available from the corresponding author upon reasonable request.</p>
GAIA model simulate data of doubled CO2 (Forces, advections, and winds)
<p>This dataset contains forces, advections, and winds output from the GAIA model, that are related to the Figures in the paper. The forces and advections are divided by the Coriolis parameter or a zonal mean absolute vorticity. </p>
Forecasting suppression of invasive Sea Lamprey in Lake Superior: data and code for Bayesian forecast model
<p>Resource managers frequently are tasked with mitigating or reversing adverse effects of invasive species through management policies and actions. In Lake Superior, of the Laurentian Great Lakes, invasive sea lamprey populations are suppressed to protect valuable fish stocks. However, the relationship between choice of long-term control strategy and the future chance of achieving the suppression target is unclear.</p> <p>Using a 60+ year time-series of suppression effort and monitoring data from 50 assessment sites located on Lake Superior tributaries, we developed a Bayesian state-space model to forecast the probability of suppressing lamprey below the suppression target.</p> <p>With annual application of lampricide (i.e., lamprey-specific pesticide) at historical mean levels, we forecasted a 15% chance of achieving the Lake Superior sea lamprey suppression target in 2040.</p> <p>Increasing lampricide effort and/or supplementing lampricide control with age-1 recruitment reduction increased suppression chance. Annual application of the maximum historical lampricide effort resulted in a 50% predicted chance of achieving the target, annual application of the mean historic lampricide effort plus a 40% reduction in recruitment resulted in a 54% chance, and the maximum amount of effort considered (maximum historic lampricide and 60% reduction in recruitment) resulted in a 94% chance.</p> <p><em><a>Policy </a>implications</em>. <a>We</a> developed a simulation model from a robust, long-term monitoring dataset that improves understanding of why long-term sea lamprey suppression objectives have been difficult to achieve in Lake Superior. Furthermore, the model provides a means to gauge efficacy of sea lamprey control policy and action scenarios based on forecasted chance of achieving the suppression target. Creating processes for iteratively refining our forecasting model with stakeholder and technical-expert input and integration with a decision analysis framework could strengthen the link between ecological knowledge obtained from long-term monitoring and invasive sea lamprey management.</p>
Supplementary data to Analyzing and Modeling the Spread of SARS-CoV-2 Omicron Lineages BA.1 and BA.2, France, September 2021–February 2022
<p>ZIP folder containing supplementary files to the article entitled <em>Analyzing and Modeling the Spread of SARS-CoV-2 Omicron Lineages BA.1 and BA.2, France, September 2021–February 2022</em> and published in Emerging Infectious Diseases with doi <a href="https://dx.doi.org/10.3201/eid2807.220033">10.3201/eid2807.220033</a></p> <ul> <li>Script_EID_1.R is the R script analysing the data_EID1.csv screening test data file (Figure 1 and Suppl Figure F1, and Tables 1 and 3).</li> <li>Script_EID_2.R is the R script analysing the data_EID2.csv screening test data file (Figures 2, 3, 5, and Suppl Figure F2, and Table 2).</li> <li>Script_EID_sequencing_raw.R is the R script analysing the data_EID_sequencing.csv sequencing data file (Figures 4, 5, and Suppl Figure F3).</li> </ul>
Data from: Occurrence-habitat mismatching and niche truncation when modelling distributions affected by anthropogenic range contractions
<p><strong>Aims: </strong>Human-induced pressures such as deforestation cause anthropogenic range contractions (ARCs). Such contractions present dynamic distributions that may engender data misrepresentations within species distribution models. The temporal bias of occurrence data—where occurrences represent distributions before (past bias) or after (recent bias) ARCs—underpins these data misrepresentations. Occurrence-habitat mismatching results when occurrences sampled before contractions are modelled with contemporary anthropogenic variables; niche truncation results when occurrences sampled after contractions are modelled without anthropogenic variables. Our understanding of their independent and interactive effects on model performance remains incomplete but is vital for developing good modelling protocols. Through a virtual ecologist approach, we demonstrate how these data misrepresentations manifest and investigate their effects on model performance.</p> <p><strong>Location:</strong> Virtual Southeast Asia</p> <p><strong>Methods:</strong> Using 100 virtual species, we simulated ARCs with 100-year land-use data and generated temporally biased (past, recent) occurrence datasets. We modelled datasets with and without a contemporary land-use variable (conventional modelling protocols) and with a temporally dynamic land-use variable. We evaluated each model's ability to predict historical and contemporary distributions.</p> <p><strong>Results:</strong> Greater ARC resulted in greater occurrence-habitat mismatching for datasets with past bias and greater niche truncation for datasets with recent bias. Occurrence-habitat mismatching prevented models with the contemporary land-use variable from predicting anthropogenic-related absences, causing overpredictions of contemporary distributions. Although niche truncation caused underpredictions of historical distributions (environmentally suitable habitats), incorporating the contemporary land-use variable resolved these underpredictions, even when mismatching occurred. Models with the temporally dynamic land-use variable consistently outperformed models without.</p> <p><strong>Main conclusions:</strong> We showed how these data misrepresentations can degrade model performance, undermining their use for empirical research and conservation science. Given the ubiquity of anthropogenic range contractions, these data misrepresentations are likely inherent to most datasets. Therefore, we present a three-step strategy for handling data misrepresentations: maximise the temporal range of anthropogenic predictors, exclude mismatched occurrences, and test for residual data misrepresentations.</p>
Data and models in Support of "Joint and Constrained Inversion as Hypothesis Testing Tools"
<p>The model and data files as well as the plotting and run scripts to reproduce the examples in "Joint and Constrained Inversion as Hypothesis Testing Tools".</p>
Research data supporting "Ephemeral Ice-Like Local Environments in Classical Rigid Models of Liquid Water"
<p>This repository contains the set of data shown in the paper <strong>"</strong><em>Ephemeral Ice-Like Local Environments in Classical Rigid Models of Liquid Water</em>", published on the Journal of Chemical Physics (DOI:10.1063/5.0088599).</p>
Exploring Hierarchy and Dependency of Rules for Consistency Checking Between Code and Model (Evaluation Data)
<p>This repository contains the data related to the protocol and results of the evaluation conducted on the HiDeoCR approach. </p>
Data from: When can model-based estimates replace surveys of wildlife populations that span many discrete management units?
<p>Monitoring widely distributed species on a budget presents challenges for the spatio-temporal allocation of survey effort. When there are multiple discrete units to monitor, survey alternatives such as model-based estimates can be useful to fill information-gaps but may not reliably reflect biological complexity and change. The spatio-temporal allocation of survey effort that minimizes uncertainty for the greatest number of units within a budget can help to ensure monitoring efforts are optimized.</p> <p>We used aerial survey-based population estimates of moose (Alces alces) across 30 Wildlife Management Units (WMUs) in Ontario, Canada to parameterize simulated populations and test the performance of different monitoring scenarios in capturing WMU-specific annual variation and trends. Firstly, we tested scenarios that prioritized conducting a survey for a unit based on one of three management criteria: population state, population uncertainty, or number of years between surveys. Also incorporated in the decision framework were WMU-specific costs and annual budget constraints. Secondly, we tested how using model-based estimates to fill information-gaps improved population and trend estimates. Lastly, we assessed how the utility (based on minimizing population uncertainty) of using a model-based estimate rather than conducting a survey was impacted by population density, severity of environmental stressors, and years since the last survey.</p> <p>Interval-based monitoring that minimized the number of years between surveys captured accurate trends for the highest number of WMUs, but annual variation was poorly captured regardless of management criteria prioritized. Using model-based estimates to fill information gaps improved trend estimation. Further, the utility of conducting a survey increased with time since the last survey and was greater for populations with low densities when the severity of environmental stressors was high, while being greater for populations with high densities when environmental severity was low.</p> <p>Overall, the utility of aerial survey monitoring was strongly associated with WMU-specific monitoring precision and the predictive power of model-based estimates. If long-term trends are evident then there is greater value in using alternatives such as model-based predictions to replace surveys, but model-based estimates may be a poor substitute when there is strong annual variation and when using a simple model.</p>
Data supporting "Stems Matter: Xylem Physiological Limits Are an Accessible and Critical Improvement to Models of Plant Gas Exchange in Deep Time"
<p>Model outputs from from <em>Paleo</em>-BGC and <em>Paleo</em>-BGC+. Code detailing data structure and allowing reproduction of analysis can be found at github.com/wjmatthaeus</p>
Auxiliary Euro-Calliope datasets: Spatial data to represent a European energy system model at several spatial resolutions
<p>Main output generated with the <a href="https://github.com/brynpickering/possibility-for-electricity-autarky/tree/custom-regions">custom-region possibility-for-electricity-autarky</a> workflow.</p> <p>This output provides similar data to <a href="https://doi.org/10.5281/zenodo.3246302">https://doi.org/10.5281/zenodo.3246302</a> (technically eligible land area for renewables and other spatially disaggregated energy system data), but with two key differences:</p> <ol> <li>The spatial extent has been expanded to include Iceland.</li> <li>Two new spatial resolutions have been added: `ehighways` and `ehighways_disaggregated`.</li> </ol> <p>`ehighways` defines 98 regions based on the result of work undertaken in the European Commission Seventh Framework Programme project e-HIGHWAY 2050 [1]. The regions cover 35 European countries; 19 are described at a national resolution and the rest at a subnational resolution. Those at a subnational resolution are aggregated from NUTS3-2006 statistical units. `ehighways_disaggregated` provides the data at the resolution of statistical units in Europe, which is then aggregated to produce the data at the `ehighways` resolution. The mapping from statistical units to ehighways regions is defined in `./ehighways/statistical_units_to_ehighways_regions.csv`. `./ehighways/units.png` shows a map of the resulting 98 `ehighways` regions. The region colours are used to help differentiate regions and have no other meaning.</p> <p>This dataset is used as an input to the <a href="https://github.com/calliope-project/sector-coupled-euro-calliope">Sector-Coupled Euro-Calliope workflow</a>.</p> <p>[1] Anderski, T., Surmann, Y., Stemmer, S., Grisey, N., Momot, E., Leger, A.-C., Betraoui, B., and van Roy, P. (2014). European cluster model of the Pan-European transmission grid (e-HIGHWAY 2050)</p>
Data and modeling input and parameter files for the 2021 Acapulco, Mexico earthquake and tsunami
<p>Raw and processed strong motion, GNSS, InSAR and tide gauge data for the event. Also includes MudPy slip inversion parameter files as well as GeoClaw input files.</p>
Documenting And Assessing Open Innovation: Co-creation Of An Open Data Model For Surgical Training (Additional materials, tables 2 & 3)
<p>Challenge competitions have recently resurged for promoting open innovation in areas where markets fail to provide incentives, such as the Sustainable Development Goals (SDGs). Challenges call for the general public to contribute novel solutions to a well-defined problem, in exchange for prizes, credentials and the promise of further development of selected solutions. The aim of this paper is to report on the development of an open and collaborative data model to document and evaluate innovations in the context of a challenge competition, while also being compatible with the work of other open source communities to validate and improve them. By reusing open documentation standards and embedding them into a semantic collaborative platform, the model aimed to be flexible enough to respond to the evaluation needs of the project organisers and self-assessment for participants. We expect our experience provides insights on the potential of semantic, collaborative platforms and standards for increasing the impact of innovations towards the SDGs.</p> <p>The developer team defined the goal and scope of the ontology in collaboration with the GSTC organisers. This was done by agreeing on scenarios where the ontology will be used and establishing competency questions that the ontology has to be able to respond to. Table 2 describes the four motivating scenarios, including actors involved, requirements, sequence of actions and main problems identified. Table 3 details the competency questions for each scenario.</p>
Random forest modelling of multi-scale, multi-species habitat associations within KAZA transfrontier conservation area using spoor data
<p>As landscape-scale conservation models grow in prominence, assessments of how wildlife utilise multiple-use landscapes are required to inform effective conservation and management planning. Such efforts should strive to incorporate multi-species perspectives to maximise value for conservation, and should account for scale to accurately capture species-environment relationships. We show that the random forest machine learning algorithm can be used to model large-scale sign-based data in a multi-scale framework. We used this method to investigate scale-dependent habitat associations for 16 mammal species of high conservation importance across the southern Kavango Zambezi (KAZA) Transfrontier Conservation Area in Botswana and Zimbabwe. Our findings revealed substantial variation in the factors shaping habitat use across species, and illustrate that different species often have divergent responses to the same environmental and anthropogenic factors, and differ in the scales at which they respond to them. For all variables across all species, scale optimisation most often selected our largest scale. Precipitation, soil nutrients, and vegetation appeared to be the most important factors determining mammal distributions, likely through their associations with food resources for herbivores and, in turn, prey availability for carnivores. Anthropogenic pressures also had an important influence on habitat use, with many species selecting against areas with high cattle density. The variety of relationships with human density indicated that species vary in their tolerance of humans. We found a consistent positive relationship with areas under high protection, and negative relationship with unprotected and less-strictly protected areas. Policy implications: This study highlights the importance of adopting a multi-scale, multi-species approach for critical decision-making processes that depend on understanding wildlife distributions and habitat associations, such as protected area, corridor, and buffer zone prioritisation. We use our findings to identify changing rainfall patterns and increasing livestock numbers as two emerging trends that may impact wildlife distributions, both within sub-Saharan Africa and on a global scale.</p>
Dynamic data of body weight and feed intake in fattening pigs and the determination of energetic allocation factors using a dynamic linear model
<p>This is the R script (DLM_script.R) to characterize the evolution of the energetic allocation factor (α<sub>t</sub>) which represents the link between the cumulative net energy available (estimated from feed intake) and cumulative weight gain during fattening period. The data for the 100 fattening pigs are stored in the csv file (DataAxiom.csv) and structured as follows:</p> <ul> <li>ID: pig identification number;</li> <li>Fattening_group.Pen : fattening group and pen number for a given ID;</li> <li>t (day): time in days since the transfer to fattening room;</li> <li>Wt (kg) : median weight in kg at day t for a given ID;</li> <li>FIt (kg day-1) : total feed intake in kg at day t for a given ID.</li> </ul> <p>For detail description of the procedure please see article.</p>
Data for "Physically Based Deep Learning Framework to Model Intense Precipitation Events at Engineering Scales"
<p>The dataset consists of high resolution (250 m) and low resolution (0.025 degree) climate model outputs in netCDF format. Each file contains data for one variable and one month.</p> <p>Low resolution files follow the naming scheme: montrealC_0025deg_200x200_ERA5_1m_YYYYMM_VAR.nc</p> <p>High resolution files follow the naming scheme: montrealC_250m_324x324_ERA5_TEB_100_noconv_YYYYMM_VAR.nc</p> <p>YYYYMM stands for the year (first 4 digits) and month (last 2 digits).</p> <p>_VAR indicates the variable contained in the file:</p> <ul> <li>_UU700 stands for the east-west component of wind at a pressure level of 700 hPa (hourly frequency)</li> <li>_VV700 stands for the north-south component of wind at a pressure level of 700 hPa (hourly frequency)</li> <li>When _VAR is omitted, the variable is precipitation at 1-minute temporal resolution</li> </ul>
Training and test data, plus saved models for the paper "Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" submitted to the SVRHM 2022 Workshop @ NeurIPS
<p>Each .pkl file contains a training or test dataset in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images used for model training. 20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in 'train_images'. All natural images are labeled with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0, according to their texture family.</li><li>'test_images': 64,000 float32 images used for model testing. 20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in 'test_images'. All natural images are labeled with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0, according to their texture family.</li></ul><p>Each .zip file contains a saved model. Details on these are coming soon.</p><p>For more details, see the paper "Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" published at the SVRHM 2022 Workshop @ NeurIPS (<a href="https://openreview.net/forum?id=8dfboOQfYt3">link</a>).</p>
Sample Input Data and Supporting Files for the SELECT Model of Urbanization
<p>Sample Input Data and Supporting Files for the SELECT Model of Urbanization</p> <p>Code available at: https://github.com/IMMM-SFA/select</p>
Deep cross-omics cycle attention model for joint analysis of single-cell multi-omics data
<p>We proposed DCCA for accurately dissecting the cellular heterogeneity on joint-profiling multi-omics data from the same individual cell by transferring representation between each other.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.