Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Code and data for publication "pyGRETA, pyCLARA, pyPRIMA: A pre-processing suite to generate flexible model regions for energy system models"
<p>This dataset contains the code of the three pre-processing tools <a href="https://github.com/tum-ens/pyGRETA">pyGRETA</a>, <a href="https://github.com/tum-ens/pyPRIMA">pyPRIMA</a> and <a href="https://github.com/tum-ens/pyCLARA">pyCLARA</a> and an examplary database for the scope of Austria.</p> <p>To run the code with full functionality additional data is needed. Check the documentation of the tools for further information.</p> <p> </p> <p>Sources for data can be found here: </p> <p>pyGRETA: https://pygreta.readthedocs.io/en/stable/user_manual.html#recommended-input-sources</p> <p>pyPRIMA: https://pyprima.readthedocs.io/en/stable/user_manual.html#recommended-input-sources</p> <p>pyCLARA: https://pyclara.readthedocs.io/en/stable/user_manual.html#recommended-input-sources</p>
GHG data from inverse models and UNFCCC national inventories v0.1
<p><strong>GHG (CO<sub>2</sub>, CH<sub>4</sub>, N<sub>2</sub>O) data from inverse models and UNFCCC national inventories</strong></p> <p>This dataset contains 5 datasets, including GHG data from inverse models and UNFCCC national inventories in the top emitter countries:</p> <p>- <strong>CO2_inversion_1990-2019</strong>: annual CO<sub>2</sub> flux from from 6 inversion models in three sectors:</p> <ul> <li>'land flux (all land)' -> land flux from all land </li> <li>'land flux (managed land)' -> land flux from managed land</li> <li>'land flux (managed land + lateral adjustment)' -> land flux from managed land by adjusting the lateral flux</li> </ul> <p>- <strong>CH4_inversion_2000-2017</strong>: CH<sub>4</sub> flux from from 10 in-situ inversion (2000-2017) and 11 satellite inversion (2010-2017) models from four sectors:</p> <ul> <li>'anthropogenic (method x)' -> anthropogenic emissions from managed land. x could be 1, 2, 3.1 and 3.2, representing different methods to calculate the emissions in this sector:</li> <li>'fossil' -> emissions from the fossil sector</li> <li>'agriculture & waste' -> emissions from the agriculture and waste sector combined</li> <li>'biomass burning' -> emissions from biomass burning</li> </ul> <p>- <strong>N2O_inversion_1997-2016</strong>: anthropogenic N<sub>2</sub>O emissions from from 3 models.</p> <p>- <strong>Inventory_1990-2019</strong>: inventory data collecting from UNFCCC national inventories. The classification of sectors is corresponding with the inversion data files for each gas specie.</p> <p>- <strong>Inventory_1990-2019_IPCC</strong>: inventory data collecting from UNFCCC national inventories in IPCC category.</p> <ul> </ul>
Data from: Transcriptional remodeling upon light removal in a model cnidarian: losses and gains in gene expression
Organismal responses to light:dark cycles can result from two general processes: (i) direct response to light or (ii) a free-running rhythm (i.e., a circadian clock). Previous research in cnidarians has shown that candidate circadian clock genes have rhythmic expression in the presence of diel lighting, but these oscillations appear to be lost quickly after removal of the light cue. Here, we measure whole-organism gene expression changes in 136 transcriptomes of the sea anemone Nematostella vectensis, entrained to a light:dark environment and immediately following light cue removal to distinguish two broadly defined responses in cnidarians: light entrainment and circadian regulation. Direct light exposure resulted in significant differences in expression for hundreds of genes, including more than 200 genes with rhythmic, 24-hour periodicity. Removal of the lighting cue resulted in the loss of significant expression for 80% of these genes after one day, including most of the hypothesized cnidarian circadian genes. Further, 70% of these candidate genes were phase shifted. Most surprisingly, thousands of genes, some of which are involved in oxidative stress, DNA damage response, and chromatin modification, had significant differences in expression in the 24 hours following light removal, suggesting that loss of the entraining cue may induce a cellular stress response. Together, our findings suggest that a majority of genes with significant differences in expression for anemones cultured under diel lighting are largely driven by the primary photoresponse rather than a circadian clock when measured at the whole animal level. These results provide context for the evolution of cnidarian circadian biology and help to disassociate two commonly confounded factors driving oscillating phenotypes.
Comparing and Combining Existing and Emerging Data Collection and Modeling Strategies in Support of Signal Control Optimization and Management (Project M2)
<p>For decades, traffic signal management agencies have used signal timing optimization tools combined with fine-tuning of signal timing based on field observations in their updates of time-of-day signal timing plans. These traditional signal optimization methods and tools use very limited amount of data and depend on default values in the signal timing optimization/simulation tools to estimate network performance under different signal optimization strategies. In recent years, new data collection technologies are emerging including high resolution controller data, more advanced detection technologies such as video image detection that are based on vehicle tracking and possible integration with microwave detectors, automatic-vehicle based identification technologies, third party crowdsourcing data, connected vehicles, and connected automated vehicles data. The objective of the proposed study is to propose methods and algorithms to combine data collected from existing and emerging sources with enhanced models and optimization algorithms to optimize and manage signal operations. The results from applying the developed methods and algorithms will be compared with traditional signal timing and optimization methods currently used by transportation agencies. </p>
Training data for the Sei framework sequence model
<p>Training data for the Sei framework deep learning model. The data contains chromatin profiles from the Cistrome Project: <strong>please agree to the terms of usage at the Cistrome Project (http://cistrome.org/db/#/bdown) before downloading. </strong></p>
Modelling data in the South China Sea
<p>Model simulated daily current, temperature, salinity, chlorophyll and nitrate concentration for three experiments (control, case1 and case2) from our coupled physical-biological model of the South China Sea . For nitrate concentration, the dataset also includes the 16-day averaged budget terms along section S1. The modeled lagrangian particl tracking data is also avaiable. And all data used to draw Fig. 1 to Fig. 10 are in Figure_data.rar file.</p>
Data from: Can conflicting selection from pollinators and nectar-robbing antagonists drive adaptive pollen limitation? A conceptual model and empirical test
<p><span><span><span><span><span><span><span><span><span><span><span>Pollen limitation is widespread, despite predictions that it shouldn't be. We propose a novel mechanism generating pollen limitation: conflicting selection by pollinators and antagonists on pollinator attraction traits. We introduce a heuristic model demonstrating antagonist-induced adaptive pollen limitation, and present a field study illustrating its occurrence in a wild population. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>For antagonist-induced adaptive pollen limitation to occur, four criteria must be met: 1) correlated attraction of pollinators and antagonists, 2) greater response by antagonists than pollinators to altered investment in attraction traits, 3) reduced investment in pollinator attraction leading to pollen limitation, and 4) higher fitness for plants with reduced investment in pollinator attraction.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>We surveyed nectar robbery and reproductive output for 109 <i>Odontonema cuspidatum </i>(Acanthaceae) plants in a pollen-limited population<i> </i>over two years and used experimental floral arrays to evaluate how flower number affects pollination and nectar robbery. Both pollinators and nectar robbers preferred larger floral displays, and nectar robbery reduced reproductive output, suggesting conflicting selection. Survey and experimental data agreed closely on the optimum flower number under antagonist-induced pollen limitation; this number was substantially overrepresented in the population. While criteria for antagonist-induced adaptive pollen limitation are restrictive, the necessary conditions may often be realized. Considering interactions beyond the plant-pollinator dyad illuminates previously overlooked mechanisms generating pollen limitation. </span></span></span></span></span></span></span></span></span></span></span></p>
Using machine learning to model nontraditional spatial dependence in occupancy data
<p>Spatial models for occupancy data are used to estimate and map the true presence of a species, which may depend on biotic and abiotic factors as well as spatial autocorrelation. Traditionally researchers have accounted for spatial autocorrelation in occupancy data by using a correlated normally distributed site-level random effect, which might be incapable of modeling nontraditional spatial dependence such as discontinuities and abrupt transitions. Machine learning approaches have the potential to model nontraditional spatial dependence, but these approaches do not account for observer errors such as false absences. By combining the flexibility of Bayesian hierarchal modeling and machine learning approaches, we present a general framework to model occupancy data that accounts for both traditional and nontraditional spatial dependence as well as false absences. We demonstrate our framework using six synthetic occupancy data sets and two real data sets. Our results demonstrate how to model both traditional and nontraditional spatial dependence in occupancy data which enables a broader class of spatial occupancy models that can be used to improve predictive accuracy and model adequacy.</p>
Data from: Assessing the usefulness of Citizen Science Data for habitat suitability modelling: opportunistic reporting versus sampling based on a systematic protocol
<p><strong>Aim:</strong> To evaluate the potential of models based on opportunistic reporting (OR) compared to models based on data from a systematic protocol (SP) for modelling species distributions. We compared model performance for eight forest bird species with contrasting spatial distributions, habitat requirements, and rarity. Differences in the reporting of species were also assessed. Finally, we tested potential improvement of models when inferring high quality absences from OR based on questionnaires sent to observers.</p> <p><strong>Location:</strong> Both datasets cover the same large area (Sweden) and time period (2000 -2013).</p> <p><strong>Methods:</strong> Species distributions were modelled using logistic regression. Predictive performance of OR models to predict SP data were assessed based on AUC. We quantified the congruence in spatial predictions using Spearman's rank correlation coefficient. We related these results to species characteristics and reporting behaviour of observers. We also assessed the gain in predictive performance of OR models by adding inferred absences. Finally, we investigated the potential impact of sampling bias in OR.</p> <p><strong>Results:</strong> For all species, and despite the sampling biases, results from OR overall agreed well with those of SP, for the nationwide spatial congruence of habitat suitability maps and the selection and directions of species-environment relationships. The OR models also performed well in predicting the SP data. The predictive performance of the OR models increased with species rarity and even outperformed the SP model for the rarest species. No significant impact of observer behaviour was found.</p> <p><strong>Main Conclusions:</strong> Relatively simple analyses with inferred absences could produce reliable spatial predictions of habitat suitability. This was especially true for rare species. OR data should be seen as a complement to SP, as the weakness of one is the strength of the other, and OR may be especially useful at large spatial scales or where no systematic data collection protocols exist.</p>
Data supplement for: Agreement of analytical and simulation-based estimates of the required land depth in climate models
<p>Many current-generation climate models have land components that are too shallow. Under climate change conditions, the long-term warming trend at the surface propagates deeper into the ground than the commonly used 3-10m. Shallow models alter the terrestrial heat storage and distribution of temperatures in the subsurface, influencing the simulated land-atmosphere interactions. Previous studies focusing on annual timescales suggest that deeper models are required to match subsurface-temperature observations and the classic analytical heat conduction solution. However, for a systematic investigation of land-model deepening in the frame of anthropogenic climate change, the classic analytical solution is inaccurate because it does not mimic the timescale and amplitude of the simulated warming trend. This study intends to bridge the gap between analytical and simulation-based estimates of the subsurface thermodynamic state by adapting the classic analytical framework to mimic long-term anthropogenic warming. The analysis shows that a land-model depth of at least 170m is recommended for a proper simulation of the post-1850 ground climate, which differs up to 30% from the estimate of the classic approach. Compared to previous studies, this provides an accurate estimate of the required land model depth for long-term climate-change simulations and indicates the relative bias in insufficiently deep land models.</p>
Models and data for "Modelling and controlling a Segway using bond graphs"
<p>The files here contain models and data for the technical report "Modelling and controlling a Segway using bond graphs".</p>
Model output data for compressible EULAG dynamical core in COSMO: convective-scale Alpine weather forecasts
<p>The archive contains model output data for article "Compressible EULAG dynamical core in COSMO: convective-scale Alpine weather forecasts" to be published in Monthly Weather Review.</p> <p>The article presents the semi-implicit compressible EULAGas a newdynamical core for convective-scale<br> numerical weather prediction. The core is implemented within the infrastructure of the<br> operational model of the Consortium for Small Scale Modeling (COSMO), forming the NWP<br> COSMO-EULAG model (CE). This regional high-resolution implementation of the dynamical<br> core complements its global implementation in the Finite-Volume Module of ECMWF’s Integrated<br> Forecasting System. The paper documents the first operational-like application of the dynamical<br> core for realistic weather forecasts. After discussing the formulation of the core and its coupling<br> with the host model, the paper considers several high-resolution prognostic experiments over<br> complex Alpine orography. Standard verification experiments examine the sensitivity of the CE<br> forecast to the choice of the advection routine and assess the forecast skills against those of the<br> default COSMO Runge-Kutta dynamical core at the 2.2 km grid showing a general improvement.<br> The skills are also compared using satellite observations for a weak-flow convective Alpine weather<br> case-study, showing favorable results. Additional validation of the new CE framework for partly<br> convection-resolving forecasts using 1.1 km, 0.55 km, 0.22 km, and 0.1 km grids, designed to<br> challenge its numerics and test the dynamics-physics coupling, demonstrates its high robustness in<br> simulating multi-phase flows over complex mountain terrain, with slopes reaching 85 degrees, and<br> the flow’s realistic representation.</p>
Data from: Genomic data and multi-species demographic modelling uncover past hybridization between currently allopatric freshwater species
<p>Evidence for ancient interspecific gene flow through hybridization has been reported in many animal and plant taxa based on genetic markers. The study of genomic patterns of closely related species with allopatric distributions allows the assessment of the relative importance of vicariant isolating events and past gene flow. Here, we investigated the role of gene flow in the evolutionary history of four closely related freshwater fish species with currently allopatric distributions in western Iberian rivers - Squalius carolitertii, S. pyrenaicus, S. torgalensis and S. aradensis - using a population genomics dataset of 23 562 SNPs from 48 individuals, obtained through genotyping by sequencing (GBS). We uncovered a species tree with two well differentiated clades: (i) S. carolitertii and S. pyrenaicus; and (ii) S. torgalensis and S. aradensis. By using D-statistics and demographic modelling based on the site frequency spectrum, comparing alternative demographic scenarios of hybrid origin, secondary contact and isolation, we found that the S. pyrenaicus North lineage is likely the result of an ancient hybridization event between S. carolitertii (contributing ~84%) and S. pyrenaicus South lineage (contributing ~16%), consistent with a hybrid speciation scenario. Furthermore, in the hybrid lineage we identify outlier loci potentially affected by selection favouring genes from each parental lineage at different genomic regions. Our results suggest that ancient hybridization can affect speciation and that freshwater fish species currently in allopatry are useful to study these processes.</p>
Complementary dataset of the paper "The viewing angle in AGN SED models, a data-driven analysis"
<p>In this repository, you can find the SED data from the X-CIGALE estimates from the article: "<a href="https://academic.oup.com/mnras/article/510/1/687/6448487">The viewing angle in AGN SED models: a data-driven analysis</a>"</p>
Veillon et al. (2020), Geoscientific Model Development : data and development
<p>F.Veillon, M.Dumont, C.Amory, M.Fructus : A versatile method for computing optimized snow albedo from spectrally fixed radiative variables : VALHALLA v1.0, Geoscientific Model Development, in review, 2020. </p> <p>See README for a full description of the dataset content</p> <p>Please contact me at marie.dumont@meteo.fr if you need more details on the dataset</p>
Data for "Unsupervised learning of sequence-specific aggregation behavior for a model copolymer"
<p>These are the data associated with the paper, "Unsupervised learning of sequence-specific aggregation behavior for a model copolymer" (DOI 10.1039/D1SM01012C). Each of the directories contains subdirectories with `GSD` files dumped from HOOMD. Each subdirectory roughly corresponds to one or two of the figures in the paper.</p>
Noise data for the study of consonant-in-noise discrimination using an auditory model with different speech-based decision devices
<p>The data stored in this repository correspond to two sets of 5000 speech-shaped noises (SSN) that were used in the conference paper titled "Consonant-in-noise discrimination using an auditory model with different speech-based decision devices" by the same authors, presented in the DAGA conference in Vienna, Austria, on 17/08/2021. </p> <p>The two zip files (<strong>osses2021c_S01</strong> and <strong>osses2021c_S02</strong> for participants S01 and S02, respectively) have following structure:</p> <ul> <li><strong>NoiseStim-SSN</strong>: Folder containing the 5000 noises</li> <li><strong>Results</strong>: Results of the listening experiment collected using the fastACI toolbox (https://github.com/aosses-tue/fastACI).</li> </ul> <p>To obtain similar results for other participants the same experiment has to be run using the fastACI toolbox. For instance, to collect new data for participant 'S03', you have to input the following command in MATLAB:</p> <pre><code class="language-bash">fastACI_experiment('speechACI_varnet2013','S03','SSN');</code></pre>
Measured and modeled data of the thermo-radiative and hydrological exhanges of a lawn in a residential area (FluxSAP 2012 campaign)
<p>The dataset includes wind, soil water content and soil temperature data measured in a private home garden collected during the FluxSAP campaign in Nantes in 2012. It also includes equivalent data modeled by the surface model for natural soils and vegetation called ISBA for four sensitivity experiments corresponding to different model configurations.</p>
Limitations of using surrogates for behaviour classification of accelerometer data: refining methods using random forest models in Caprids
<p>Animal-attached devices can be used on cryptic species to measure their movement and behaviour, enabling unprecedented insights into fundamental aspects of animal ecology and behaviour. However, direct observations of subjects are often still necessary to translate biologging data accurately into meaningful behaviours. As many elusive species cannot easily be observed in the wild, captive or domestic surrogates are typically used to calibrate data from devices. However, the utility of this approach remains equivocal. </p> <p>Here, we assess the validity of using captive conspecifics, and phylogenetically-similar domesticated counterparts (surrogate species) for calibrating behaviour classification. Tri-axial accelerometers and tri-axial magnetometers were used with behavioural observations to build random forest models to predict the behaviours. We applied these methods using captive Alpine ibex (Capra ibex) and a domestic counterpart, pygmy goats (Capra aegagrus hircus), to predict the behaviour including terrain slope for locomotion behaviours of captive Alpine ibex. </p> <p>Behavioural classification of captive Alpine ibex and domestic pygmy goats was highly accurate (> 98%). Model performance was reduced when using data split per individual, i.e., classifying behaviour of individuals not used to train models (mean ± sd = 56.1 ± 11%). Behavioural classifications using domestic counterparts, i.e., pygmy goat observations to predict ibex behaviour, however, were not sufficient to predict all behaviours of a phylogenetically similar species accurately (> 55%).</p> <p>We demonstrate methods to refine the use of random forest models to classify behaviours of both captive and free-living animal species. We suggest there are two main reasons for reduced accuracy when using a domestic counterpart to predict the behaviour of a wild species in captivity; domestication leading to morphological differences and the terrain of the environment in which the animals were observed. We also identify limitations when behaviour is predicted in individuals that are not used to train models. Our results demonstrate that biologging device calibration needs to be conducted using: (i) with similar conspecifics, and (ii) in an area where they can perform behaviours on terrain that reflects that of species in the wild.</p>
Simulated CO2 time series data based on Jena CO2 inversion and TM3 transport model, and MIROC-ACTM
<p>Each file contains simulated CO2 time series at each surface station. The model, simulation type, and station are specified in the file name. These simulations are driven by either varying winds alone (e.g., Jena_W) or varying winds and fluxes (e.g., Jena_WF). The only MIROC-ACTM run is named ACTM_W_MLO.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.