Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data from: Model parameterization of four species distribution models
<p>Species Distribution Models (SDMs) are practical tools to assess the habitat suitability of species with numerous applications in environmental management and conservation planning. The manipulation of the input data to deal with their spatial bias is one of the advantageous methods to enhance the performance of SDMs. However, the development of a model parameterization approach covering different SDMs to achieve well-performing models has rarely been implemented. We integrated input data manipulation and model tuning for four commonly-used SDMs: generalized linear model (GLM), gradient boosted model (GBM), random forest (RF), and maximum entropy (MaxEnt), and compared their predictive performance to model geographically imbalanced biased data of a rare species complex of mountain vipers. Models were tuned up based on a range of model-specific parameters considering two background selection methods: random and background weighting schemes. The performance of the fine-tuned models was assessed based on a recently identified localities of the species. The results indicated that although the fine-tuned version of all models shows great performance in predicting training data (AUC > 0.9 and TSS > 0.5), they produce different results in classifying out-of-bag data. The GBM and RF with higher sensitivity of training data showed more different performances. The GLM, despite having high predictive performance for test data, showed lower specificity. It was only the MaxEnt model that showed high predictive performance and comparable results for identifying test data in both random and background weighting procedures. Our results highlight that while GBM and RF are prone to overfitting training data and GLM over-predict non-sampled areas MaxEnt is capable of producing results that are both predictable (extrapolative) and complex (interpolative). We discuss the assumptions of each model and conclude that MaxEnt could be considered as a practical method to cope with imbalanced-biased data in species distribution modeling approaches.</p>
Raw Data for the article: Human Amnion-Derived Mesenchymal Stromal/Stem Cells Pre-Conditioning Inhibits Inflammation and Apoptosis of Immune and Parenchymal Cells in an In Vitro Model of Liver Ischemia/Reperfusion
<p>Ischemia/reperfusion injury (IRI) represents one of the leading causes of primary non-function acute liver transplantation failure. IRI, generated by an interruption of organ blood flow and the subsequent restoration upon transplant, i.e., reperfusion, generates the activation of an inflammatory cascade from the resident Kupffer cells, leading first to neutrophils recruitment and second to apoptosis of the parenchyma. Recently, human mesenchymal stromal/stem cells (hMSCs) and derivatives have been implemented for reducing the damage induced by IRI. Interestingly, sparse data in the literature have described the use of human amnion-derived MSCs (hAMSCs) and, more importantly, no evidence regarding hMSCs priming on liver IRI have been described yet. Thus, our study focused on the definition of an in vitro model of liver IRI to test the effect of primed hAMSCs to reduce IRI damage on immune and hepatic cells. We found that the IFNγ pre-treatment and 3D culture of hAMSCs strongly reduced inflammation induced by M1-differentiated macrophages. Furthermore, primed hAMSCs significantly inhibited parenchymal apoptosis at early timepoints of reperfusion by blocking the activation of caspase 3/7. All together, these data demonstrate that hAMSCs priming significantly overcomes IRI effects in vitro by engaging the possibility of defining the molecular pathways involved in this process.</p>
Air temperature and thermal comfort data measured and modelled for 121 workplaces in the Upper Rhine Valley
<p>The uploaded files contain the measured and modelled indoor data at different workplaces in the Upper Rhine Valley between August 1, 2021, and July 31, 2022 presented in the article "Predicting Indoor Air Temperature and Thermal Comfort in Occupational Settings Using Weather Forecasts, Indoor Sensors, and Artificial Neural Networks" by Sulzer et al. (2023), <a href="http://doi.org/10.1016/j.buildenv.2023.110077">doi.org/10.1016/j.buildenv.2023.110077</a>. Information about the characteristics of each workplace can be found in the appendix of the article. For every workplace two files are uploaded, one for the data of the indoor air temperature (Ta) and one for indoor physiological equivalent temperature (PET). The workplace ID and variable are stated in the filename. The columns in the files contain the following data:</p> <ul> <li>"Datetime (UTC)": This column contains the timestamp in UTC of the starting point of the interval used for the one-hour mean values .</li> <li>"MoBiMet data": This column contains the one-hour mean values of Ta or PET of the data derived every five minutes by the low-cost Mobile Biometeorology System (MoBiMet) at the corresponding workplace in °C. The MoBiMet are presented in detail in <a href="http://doi.org/10.3390/s22051828">doi.org/10.3390/s22051828</a>.</li> <li>"Used for": This column contains the information if the data point was used for training of the models (t), evaluation of the models (e), or not used for model training or evaluation due to missing indoor observation data (n).</li> <li>"Product 0 (ICON_outdoor)": This column, in the files for the indoor air temperature, contains the ICON-D2 data of the air temperature 2m a.g.l. in °C of the grid cell in which the associated work station is located.</li> <li>"Product 0 (ICON_outdoor) air temperatur 2m (C) input for PET calculation using RayMan","Product 1 (ICON_outdoor) vapor pressure 2m (hPa) input for PET calculation using RayMan","Product 1 (ICON_outdoor) wind speed 10m (m/s) input for PET calculation using RayMan", and "Product 1 (ICON_outdoor) global radiation surface (W/m²) input for PET calculation using RayMan": This columns contain the data of the outdoor air temperature 2m a.g.l. in °C, the vapor pressure 2m a.g.l., derived from the ICON-D2 weather forecast data of the grid cell in which the associated work station is located, which were used in RayMan Pro to calculate the PET for outdoors.</li> <li>"Product 2 (ANN_Generic)": This column contains the indoor data of PET or air temperature in °C modelled by an artificial neural network using generic data as input. The generic data contain hourly solar altitude and azimuth at each location, the weekday, and a sine and cosine function of the daily and yearly cycle.</li> <li>"Product 3 (ANN_AWS) without past data": This column contains the indoor data of PET or air temperature in °C modelled by an artificial neural network using generic data and the meteorological data of air temperature, vapor pressure, mean sea level pressure, global radiation, longwave downwelling radiation, and wind speed of an automated weather station in Freiburg (Station FRCHEM; 48°00’04’’ N; 7°50’55’’ E).</li> <li>"Product 3 (ANN_AWS) with past data": This column contains similar data than the column before but the artificial neural network models used "past data" of the automated weather station as additional input variables to model indoor air temperature and PET in °C. Additional to the hourly average of the meteorological data for each actual time (t), hourly averages for t-1 h, t-3 h, t-6 h, t-12 h, and t-24 h of air temperature, global radiation, and Longwave downwelling radiation are used as so called "past data".</li> <li>"Product 4 (ANN_ICON) without past data": This column contains the indoor data of PET or air temperature in °C modelled by an artificial neural network using generic data and the meteorological data of air temperature, vapor pressure, mean sea level pressure, global radiation, longwave downwelling radiation, and wind speed derived from the ICON-D2 weather forecast data of the grid cell in which the associated work station is located.</li> <li>"Product 4 (ANN_ICON) with past data": This column contains similar data than the column before but the artificial neural network models used "past data" of the ICON-D2 weather forecast data as additional input variables to model indoor air temperature and PET in °C. Additional to the hourly average of the meteorological data for each actual time (t), hourly averages for t-1 h, t-3 h, t-6 h, t-12 h, and t-24 h of air temperature, global radiation, and Longwave downwelling radiation are used as so called "past data".</li> <li>"Product 5 (ANN_Mixed) without past data": This column contains the indoor data of PET or air temperature in °C modelled by the same artificial neural network models as in Product 3 but applied for the same input data of Product 4.</li> <li>"Product 5 (ANN_Mixed) with past data": This column contains similar data than the column before but also takes into account the "past data".</li> </ul> <p>The data of Product 3 (ANN_AWS) and Product 5 (ANN_Mixed) are only available for locations in Freiburg, because the data of an automated weather station in Freiburg was used.<br> Data which is not available is stated as NA.</p>
Processed data for RCPANN model
<p>Processed RBSP/RBSPICE data, satellite coordinates and geomagnetic indices</p>
Toward minimal composite Higgs models from regular geometries in bottom-up holography—data release
<p>This dataset contains the data points in the plots of the preprint <a href="https://arxiv.org/abs/2303.00541">Towards composite Higgs: minimal coset from a regular bottom-up holographic model</a>.</p> <p>If you use this data release in the context of your research, please cite the aforementioned paper.</p> <p>Further details are given in the file ReadMe.md.</p> <p> </p>
Data from: When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers
<p>This dataset includes metadata of the newspaper articles used for the paper "When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers". The dataset is in JSON format. The metadata includes: "uuid" (unique identifier we associated to an article), "URLs" (the URLs where the article was published), "sources" (newspaper and feed/section where the article was published), "datesPublished" (dates when the article was published/updated).</p> <p>License: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">https://creativecommons.org/licenses/by-sa/4.0/legalcode</a>)</p> <p> </p>
Data and Code for Modeling Lake Cahuilla
<p>Data and Code for Modeling Lake Cahuilla</p>
Data for Port-Hamiltonian Heat and Wave models
<p>Data files for port-Hamiltonian heat and wave equation models</p>
The data and code for "Gap-Filling of Turbulent Heat Fluxes over Rice–Wheat-Rotation Croplands Using the Random Forest Model""
<p>This file contains the dataset and code for the paper "Gap-Filling of Turbulent Heat Fluxes over Rice–Wheat-Rotation Croplands Using the Random Forest Model".</p>
Data Set Generated by the Fuzzy Model Constructed to Describe Execution Tracing Quality
<p>The uploaded data set was generated by the fuzzy model published in T. Galli, F. Chiclana, and F. Siewe. Genetic algorithm-based fuzzy inference system for describing execution tracing quality. Mathematics, 9(21), 2021. ISSN 2227-7390. doi: https://doi.org/10.3390/ma th9212822. URL https://www.mdpi.com/2571-5577/4/1/20.</p> <p>The goal of the data generation is to make the published model available in the form of data points in a 5D space, which facilitates the construction of simpler models to approximate the original model. The names of the columns in the .csv file constitute the quality properties of execution tracing: (1) accuracy, (2) legibility, (3) implementation, and (4) security, while column (5) contains execution tracing quality derived from the fuzzy model. The indices in brackets show the column indices in the .csv file.</p> <p>All variables lie in the continuous range [0, 100], where 100 means the best possible quality value and 0 the complete lack of quality or the lack of the given quality property. While generating the data, the inputs were increased by a step-size 5 and the model's output was collected, i.e. 4 inputs, from including 0 to 100 with 21 data points (21^4 = 194481).</p> <p> </p>
Calculated data from Thermodynamic modelling of the nature of speciation and phase behaviour of binary and ternary mixtures of formaldehyde, water and methanol. MolPhys 2023
<p>Calculated data in the figures of the publication. </p>
Data from Real-time, model-based magnetic field correction for moving, wearable MEG
<p>OP-MEG data used to generate figures from manuscript titled "Real-time, model-based magnetic field correction for moving, wearable MEG". Each zipped folder relates to a different experiment: EnvironmentalNoise relates to the environmental noise experiments, AEF is the auditory evoked response experiment and ExternalCoils relates to the recordings using a set of external coils to produce interference presented in the supplementary material of the paper.</p>
Colon Data - Lipid MSI from Chronic Pain Model (Rat)
<p>Rat model of comorbid visceral pain. Colon tissues swiss-rolled in 2% gelatin. Cryosections at 12 um thickness. Norharmane matrix, negative ion mode lipid scans by MALDI-MSI at 50um spatial resolution. Converted from native file structure to the common file format, .imzML. Compressed zip folders contain both the .ibd and .imzML files.</p> <p>Data used in: Kasun Pathirage, Aman Virmani, Alison Scott, Richard Traub, Robert Ernst, Reza Ghodssi, Behtash Babadi, and Pamela Abshire. (2023). “Interpretable dimensionality reduction and classification of mass spectrometry imaging data in a visceral pain model via non-negative matrix factorization,” bioRxiv 2023.04.24.538180. doi: <a href="https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdoi.org%2F10.1101%2F2023.04.24.538180&data=05%7C01%7CAscott1%40umaryland.edu%7Cc2a3f9b67a964db86c8a08db484f3b69%7C3dcdbc4a7e4c407b80f77fb6757182f2%7C0%7C0%7C638183279460394978%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=0yzFTOcsD9FiNnZK8R28EVKqKiW1MwydsF0LW0KzJgE%3D&reserved=0">https://doi.org/10.1101/2023.04.24.538180</a></p> <p> </p>
The data from CropPol which are not included in my linear mixed models and the reasons.
<p>The data from CropPol which are not included in my linear mixed models and the reasons. Note_for_not_include column refers to the reason why these data was excluded. </p>
Open Research Data for "A New Framework for Evaluating Model Simulated Inland Tropical Cyclone Wind Fields"
<p>The (1) NOAA GFDL T-SHiELD outputs, (2) processed ASOS data, and (3) observation-based, theory-driven wind profiles data used in the manuscript "A New Framework for Evaluating Model Simulated Inland Tropical Cyclone Wind Fields". </p>
the data for WPV model for budburst
<p>The budburst data in five populations for three species</p>
Supplementary data: "Physics-informed machine learning for power grid frequency modelling"
<p>This repository contains result files for the paper "Physics-informed machine learning for power grid frequency modelling" <a href="https://doi.org/10.48550/arXiv.2211.01481">(Preprint)</a>. The code for producing the processed data and the results is <a href="https://github.com/johkruse/PIML-for-grid-frequency-modelling">available at github</a>.</p> <p><strong>Results</strong></p> <p>The result folder comprises the results of hyper-parameter optimisation, scaling variation and interpretation via SHAP. In particular, it contains these sub-folders and files:</p> <ul> <li><em>tuning </em>: Results of hyper-parameter tuning.</li> <li><em>best_model </em>: Weights of the trained model with best hyper-parameters.</li> <li><em>best_model_<scaling-variation> </em>: Weights of the trained models with best hyper-parameters but with a variation of the parameter scaling.</li> <li><em>fixed_model_hps.pkl </em>: Hyper-parameters that are not optimised.</li> <li><em>shap_values_<parameter>_long.h5</em> : SHAP values for the prediction of the system parameters.</li> </ul>
Model simulations utilizing the latest urban underlying surface and anthropogenic heat data
<p>Based on numerical simulations utilizing the latest urban underlying surface and anthropogenic heat data over the Yangtze River Delta urban agglomeration, we find that LU change and AH emission can result in opposite effects on summer precipitation. The related model simulations are included in this dataset.</p>
Data supplement for "Stationary broken parity states in active matter models''
<p>Data sets and python codes that produce all figures in "Stationary broken parity states in active matter models''. Additionally it contains example Matlab codes that perform the numerical path continuation with pde2path.</p>
Data from: Analysing ecological dynamics with relational event models: the case of biological invasions
<p>Dynamic species invasions network: sender nodes are species and receiver nodes are regions as defined in the FirstRecords database. The dataset covers four taxonomic groups (i.e., mammals, birds, plants, insects) and a fixed time-frame [1880–2005] over which the invasion process of the species occurs into the regional set of the 272 pre-specified regions. The dataset includes 16,403 invasion events recorded for 4,835 species (615 birds, 186 mammals, 3,920 plants, and 114 insect species). The average number of invasion events per taxonomic group is ~5 records per insect species, ~4 per bird and mammal species, and ~3 recorded events per plant species. On average, ~46 invasion events are recorded per region, ranging from 1 to 1,685 events. 157 regions had less than 15 observations, with 33% of these regions located in Africa; 36 regions had more than 100 records.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.