Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
310
datasets available to search
ShareScore release 0.9.0
Dataset results
310 results for “State Model”
Bayesian analysis of the equation of state of quantum chromodynamics from a holographic model
<p>Prior and posterior samples obtained from a Bayesian analysis of the equation of state of quantum chromodynamics (QCD) within a holographic Einstein-Maxwell-Dilaton model, constrained by state-of-the art lattice QCD results at a vanishing net density of baryons.</p> <p>Samples contain metadata, model parameters, and model predictions for the location of the QCD critical point.</p> <p>Supplement to <a title="Bayesian location of the QCD critical point from a holographic perspective" href="https://arxiv.org/abs/2309.00579">arXiv:2309.00579</a>.</p>
Bridging data silos to holistically model plant macrophenology data, Contiguous United States, 2013-2021
Phenological responses to climate change can have dire implications for ecosystem functions. Despite the availability of diverse datasets (e.g., herbarium specimens, community science initiatives, observatory networks, and remote sensing), holistic modeling of plant events across scales remains limited due to fragmented data and disciplinary silos. This is an important topic that has been overdue for attention. Here we use two different plant phenological datasets, herbarium and USA-NPN (includes NEON), to look at the overall flowering period of Acer rubrum between 2013-2021, distributed across the Contiguous United States. We harmonize the data to demonstrate its use to leverage the spatial and biological organizational scales at which these data are captured. Both datasets include phenophase status (presence or absence) across the flowering season (day of year). These harmonized data exemplify their usefulness to holistically model plant phenology using an integrated species distribution model framework, while accounting for the heterogeneity across data types (presence-only, presence-absence). These data can be used to explore general questions about intraspecific synchrony of Acer rubrum flowering phenology across populations, or questions with coarser scales of interest (e.g., community level, global scales).
Modeling an auditory stimulated brain under altered states of consciousness using the generalized ising model
Open the record for dataset details and reuse information.
Dataset for "Modeling Dipolar Nonprotogenic Solvents with PC-SAFT-Type Equations of State: Pure Substance Properties"
<p>Dipolar nonprotogenic solvents (DNS) are important chemical substances used across a wide range of applications, including renewable green solvent media, sustainable energy sources, and as efficient solvents for fabricating and processing semiconductive materials used in organic photovoltaics. Therefore, for efficient solvent screening or process design, a description and prediction of the thermodynamic properties of DNS using thermodynamic models is essential. This dataset contains calculation results of four different modeling strategies within the PC-SAFT equation of state for pure-substance properties of six DNSs: gamma-valerolactone, propylene carbonate, acetonitrile, dihydrolevoglucosenone, 1-methyl-2-pyrrolidone, and sulfolane. The modeling strategies differ in the treatment of the strong dipolar interactions of DNSs. The pure-substance properties include liquid density, vapor pressure, enthalpy of vaporization, and residual isobaric liquid heat capacity. The PC-SAFT performance was analyzed and evaluated based on the calculated data. Additionally, the dataset includes input files for quantum mechanical calculations of optimal molecular geometries and dipole moments of the considered DNSs using Gaussian 16 software.</p>
A Finite State Tranducer that models Chinese Historical Phonology
<p>This is a finite state transducer that tries to model Chinese historical phonology from Old Chinese as reconstructed by Baxter and Sagart in <em>Old Chinese: a New Reconstruction</em> (Oxford, 2014) to Middle Chinese as presented in the system of Baxter in <em>A Handbook of Old Chinese Phonology </em>(Mouton, 1992).</p>
Beyond the Four-Level Model: Dark and Hot States in Quantum Dots Degrade Photonic Entanglement
<p><strong>Dataset for "Beyond the four-level model: Dark and hot states in quantum dots degrade photonic entanglement"</strong></p> <p><em>Nano Lett.</em> 2023, 23, 4, 1409–1415<br> Publication Date: February 6, 2023<br> <a href="https://doi.org/10.1021/acs.nanolett.2c04734">https://doi.org/10.1021/acs.nanolett.2c04734</a></p> <p>A description of the dataset is found in the <strong>readme.md</strong> file (markdown markup language).</p> <p><strong>Data reuse</strong><br> Please cite B.U. Lehner et al., <em>Nano Lett.</em> 2023, 23, 4, 1409–1415 (2023) in publications that reuse this data and if possible inform the corresponding authors.</p>
Lake chloride concentrations and model predictions for 49,432 lakes in the Midwest and Northeast United States.
Lakes in the Midwest and Northeast United States are at risk of anthropogenic chloride contamination, but we have little knowledge of the prevalence and spatial distribution of the problem. The majority of salt pollution in north temperate regions stems from road salt application but other chloride sources include water softeners, synthetic fertilizers, and livestock excretion. Although chloride contamination of lakes is well documented, it is unknown how many lakes are at risk of long-term salinization. We used a quantile regression forest to leverage information from 2,773 lakes to predict the chloride concentration of all 49,432 lakes greater than 4 ha in a 17-state area. The QRF used 22 predictor variables, which included lake morphometry characteristics, watershed land use, and distance to the nearest interstate and road. Model predictions had an r2 of 0.94 for all chloride observations, and 0.87 for predictions of the mean chloride concentration observed at each lake.
Model-based aversive learning in humans is supported by preferential task state reactivation
Open the record for dataset details and reuse information.
Arctic Ocean state estimates for 2013 using the GECCO model
<p>The dataset contains the 2013 data of a 10-year ocean synthesis (2007-2016) obtained by assimilating available observations of sea ice and ocean parameters into the GECCO model. Data from, among others, several satellite programs such as AMSRE, SSMI, AMSR2, Envisat, Jason, Cryosat., AVHRR, and SMOS, and available moorings in the Davis Strait, the Bering Strait, the Fram Strait, the Barents Sea Opening, and by the Nansen and Amundsen Basins Observational System (NABOS), the North Pole Environmental Observatory (NPEO), and the Beaufort Gyre Exploration Project (BGEP) project. A detailed description can be found in Lyu et al., 2020.</p> <p>Guokun Lyu, Nuna Serra, Armin Koehl and Detlef Stammer, 2020. INTAROS Deliverable 6.4 Ice-ocean statistics and state estimation V1. https://intaros.nersc.no/sites/intaros.nersc.no/files/D6.4_INTAROS_Data_assimilation_v1.3.pdf </p>
Arctic Ocean state estimates for 2012 using the GECCO model
<p>The dataset contains the 2012 data of a 10-year ocean synthesis (2007-2016) obtained by assimilating available observations of sea ice and ocean parameters into the GECCO model. Data from, among others, several satellite programs such as AMSRE, SSMI, AMSR2, Envisat, Jason, Cryosat., AVHRR, and SMOS, and available moorings in the Davis Strait, the Bering Strait, the Fram Strait, the Barents Sea Opening, and by the Nansen and Amundsen Basins Observational System (NABOS), the North Pole Environmental Observatory (NPEO), and the Beaufort Gyre Exploration Project (BGEP) project. A detailed description can be found in Lyu et al., 2020.</p> <p>Guokun Lyu, Nuna Serra, Armin Koehl and Detlef Stammer, 2020. INTAROS Deliverable 6.4 Ice-ocean statistics and state estimation V1. https://intaros.nersc.no/sites/intaros.nersc.no/files/D6.4_INTAROS_Data_assimilation_v1.3.pdf </p>
Additional steady-state simulations of Miocene Antarctic ice-sheet variability using 3D thermodynamical ice-sheet model IMAU-ICE
<div> </div> <div> <div> <div>We supplement our previous dataset (<a href="https://doi.pangaea.de/10.1594/PANGAEA.939114">doi:10.1594/PANGAEA.939114</a>), with six additional steady-state simulations of the Miocene Antarctic ice sheet using the reference Miocene settings.</div> <div> </div> <div>IMAU-ICE was run using a 40x40km grid covering the Antarctic continent. Initial conditions were obtained from reconstructions of the Antarctic bathymetry and bedrock topography pertaining to 23 to 24 million years (Myr) ago (dataset <a href="https://doi.pangaea.de/10.1594/PANGAEA.923109" target="_self">doi:10.1594/PANGAEA.923109</a>). The simulations were forced by climate input data obtained from GENESIS simulations with varying CO2 levels (280 to 840 ppm) and Antarctic ice sheet cover (no ice to a large East-Antarctic ice sheet), and with present-day insolation. We utilized a matrix interpolation method to construct the time-varying climate forcing, based on the prescribed CO2 levels and ice cover simulated by IMAU-ICE.</div> <div> </div> <div>For each simulation, we provide the run script, 1D output variables including CO2 level and the sea level contribution of the Antarctic ice sheet, and 3D output variables including ice thickness, bedrock and surface height, surface mass balance, basal mass balance, ice velocities, and ice temperatures. For more information, please contact L.B. Stap at l.b.stap@uu.nl.</div> </div> </div>
Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities
<p>Numerical data used to make Figures in the manuscript entitled "Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities". We provide the data to create Figures 1 to 4 from the main text and Figures S1 to S9 from the supplementary information. We also provide Python scripts to plot them.</p>
Code and measurement data - State of charge and state of health diagnosis of batteries with voltage-controlled models
<p><strong>This dataset contains the research data (code and measurement data) of the journal article: <a href="https://doi.org/10.1016/j.jpowsour.2022.231828">J. A. Braun, R. Behmann, D. Schmider, W. G. Bessler, "State of charge and state of health diagnosis of batteries with voltage-controlled models", Journal of Power Sources 544 (2022), 231828</a>.</strong></p> <p> </p> <p><strong>Abstract:</strong><br> The accurate diagnosis of state of charge (SOC) and state of health (SOH) is of utmost importance for battery users and for battery manufacturers. State diagnosis is commonly based on measuring battery current and using it in Coulomb counters or as input for a current-controlled model. Here we introduce a new algorithm based on measuring battery voltage and using it as input for a voltage-controlled model. We demonstrate the algorithm using fresh and pre-aged lithium-ion battery single cells operated under well-defined laboratory conditions on full cycles, shallow cycles, and a dynamic battery electric vehicle load profile. We show that both SOC and SOH are accurately estimated using a simple equivalent circuit model. The new algorithm is self-calibrating, is robust with respect to cell aging, allows to estimate SOH from arbitrary load profiles, and is numerically simpler than state-of-the-art model-based methods.</p> <p> </p> <p><strong>Intellectual property information:</strong><br> The Matlab codes and the research data provided here are under <strong><a href="https://creativecommons.org/licenses/by-nc/4.0/legalcode">CC-BY-NC-4.0</a></strong> license. Please note that the algorithms themselves are subject to industrial property rights, including, but not necessarily limited to, German patent <strong><a href="https://patents.google.com/patent/DE102019127828B4/en">DE102019127828B4</a></strong> and international patent application <strong><a href="https://patents.google.com/patent/WO2021073690A2/en">WO2021073690A2</a></strong>. Any use of the codes and algorithms presented here is subject to these property rights.</p> <p> </p> <p><strong>Overview of files:</strong><br> <strong>SOC_SOH_simple_model.m:</strong> Matlab script performing SOC and SOH diagnosis with the voltage-controlled "simple" equivalent circuit model. The script also reproduces the figures shown in the manuscript.</p> <p><strong>SOC_SOH_simple_extended.m:</strong> Matlab script performing SOC and SOH diagnosis with the voltage-controlled "extended" equivalent circuit model. The script also creates figures of additional data not shown in the manuscript.</p> <p><strong>Experimental_data_fresh_cell.csv:</strong> Tabulated experimental data (time, current, voltage, temperature) of the long-term experiment (99 h total with 1 s resolution) of a fresh lithium-ion cell. The cell is initally completely discharged. The data consist of full cycling, shallow cycling, and WLTP cycling.</p> <p><strong>Experimental_data_aged_cell.csv:</strong> Tabulated experimental data (time, current, voltage, temperature) of the long-term experiment (85 h total with 1 s resolution) of a pre-aged lithium-ion cell. The cell is initally completely discharged. The data consist of full cycling, shallow cycling, and WLTP cycling.</p> <p><strong>OCV_vs_SOC_curve.csv:</strong> Tabulated experimentally-derived open-circuit voltage (OCV) as function of state of charge (SOC). 1001 data points between SOC = 0 and SOC = 1 in increments of 0.001.</p> <p><strong>readme.txt:</strong> Overview of files with a short description.</p>
Exploring the Relationship Between Upper Ocean States and the Falling Ice Radiative Effects using ECCO Product and Global Climate Models
<p><strong><span>Sensitivity test using CESM1-CAM5 following CMIP5 protocool from 1980-2005</span></strong></p> <p><strong><span>NOS: no falling ice radiative effects (FIREs), four data sets</span></strong></p> <p><strong><span>SON: with FIREs, for data sets</span></strong></p> <p><strong><span> Xsize = 362 Ysize = 182 Zsize = 18</span></strong></p> <p><strong><span>Format: netcdf</span></strong></p> <p><strong><span>Upper 200 meter ocean variables</span></strong></p> <p><strong><span>Annual mean (ANN)</span></strong></p> <p><strong><span>CESM2-var-NOS (or SON)-ANN.nc, var = (UO, VO, WO, TO) = (zonal velocity, meridional velocity, ascending velocity, potential temperature) : (cm/s, cm/s, cm/s, K)</span></strong></p>
Assessment of mutation probabilities of KRAS G12 missense mutants and their long-time scale dynamics by atomistic molecular simulations and Markov state modeling: Datasets.
<p>Datasets related to the publication [1].<br> Including:</p> <ul> <li>KRAS G12X mutations derived from COSMIC v.79 [http://cancer.sanger.ac.uk/cosmic/] (KRAS_G12X_mut_COSMICv79..xlsx)</li> <li>RMSFs (300-2000ns) of GDP-systems (300_2000rmsf_GDP_systems_RAW_AVG_SE.xlsx)</li> <li>RMSFs (300-2000ns) of GTP-systems (300_2000RMSF_GTP_systems_RAW_AVG_SE.xlsx)</li> <li>PyInteraph analysis data for salt-bridges and hydrophobic clusters (.dat files for each system in the PyInteraph_data.zip-file)</li> <li>Backbone trajectories for each system (residues 4-164; frames for every 1ns). Last number (e.g. _1) refers to the replica of the simulated system.</li> <li>backbone_4-164.gro/.pdb/.tpr -files (resid 4-164) </li> </ul> <p><br> [1] Pantsar T et al. Assessment of mutation probabilities of KRAS G12 missense mutants and their long-time scale dynamics by atomistic molecular simulations and Markov state modeling. <em>PLoS Comput Biol Submitted</em> (2018)</p>
Dataset of "Comprehensive Machine Learning Approaches for Modelling the State of Charge of Lithium-ion Batteries"
<p>This paper evaluates three ML approaches for SOC modeling in LIBs: the multilayer perceptron (MLP), long short-term memory (LSTM), and the nonlinear autoregressive with exogenous input (NARX) neural network architectures. These models were tested using an experimental dataset with multiple input variables, including electrochemical impedance spectroscopy (EIS) data, voltage, and capacity readings for commercial LIB cells. Results indicate that MLP and LSTM are more adaptable with a smaller training dataset (14 samples), while the NARX model required more than 34 out of 67 samples to achieve reasonable accuracy. Additionally, the NARX model is more sensitive to changes in the learning rate (α) and exhibits larger output error deviations. The MLP and LSTM models consistently performed well across various hidden layer sizes, showing no upper bound constraints, whereas the NARX model’s performance deteriorated with certain hidden layer configurations.</p>
Integrative in situ mapping of single-cell transcriptional states and tissue histopathology in an Alzheimer disease model
<p>Amyloid-β plaques and neurofibrillary tau tangles are the neuropathologic hallmarks of Alzheimer’s disease (AD), but the spatiotemporal cellular responses and molecular mechanisms underlying AD pathophysiology remain poorly understood. Here we introduce STARmap PLUS to simultaneously map single-cell transcriptional states and disease marker proteins in brain tissues of AD mouse models at a voxel size of 95 95 350 nm. This high-resolution spatial transcriptomics map revealed a core-shell structure where disease-associated microglia (DAM) closely contact amyloid-β plaques, whereas disease-associated astrocyte-like cells (DAA-like) and oligodendrocyte precursor cells (OPC) are enriched in the outer shells surrounding the plaque-DAM complex. Hyperphosphorylated tau emerged mainly in excitatory neurons in the CA1 region accompanied by infiltration of oligodendrocyte subtypes into the axon bundles of hippocampal alveus. The integrative STARmap PLUS method bridges single-cell gene expression profiles with tissue histopathology at subcellular resolution, providing an unprecedented roadmap to pinpoint the molecular and cellular mechanisms of AD pathology and neurodegeneration.</p>
Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 2
<p>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al.</p> <p>The code to use with these data and reproduce the manuscript results is available at https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. param_fixing.zip - self-explanatory (Figure 4 & 5); contains an explanatory note for this part (experiment_details.txt), and the file containing Km values fetched from the BRENDA database (Km_database.csv).</p> <p>2. scripts.zip - scripts to generate figure 2-5 on toy data</p>
Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 1
<p><strong>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al (https://doi.org/10.1101/2023.02.21.529387).</strong></p> <p>The code to use with these data and reproduce the manuscript results is available at https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. models.zip - contains thermodynamically curated steady-state and nonlinear kinetic models of <em>E. coli </em>metabolism used in this study. Also contains the samples of steady-state metabolite concentrations and metabolic fluxes used in the study presented in Figure 3 (steady-state samples used for preparing Figures 2 and 4).</p> <p>2. renaissance_incidence_results.zip - self-explanatory (Figure 2a and 2b)</p> <p>3. ODE_solutions.zip - self-explanatory (Figure 2c)</p> <p>4. bioreactor_simulations1-3.zip - self-explanatory (Figure 2d)</p> <p>5. steady_state_analysis.zip - RENAISSANCE results obtained for each of the steady states (Figure 3a)</p> <p>6. subspace_analysis.zip - RENAISSANCE results presented in Figure 3b-g</p> <p><strong>The remaining datasets are published in the following links</strong></p> <p><em> - https://doi.org/10.5281/zenodo.7930084</em></p> <p><em> - https://doi.org/10.5281/zenodo.10391802</em></p>
Asymmetry of AMOC Hysteresis in a State-of-the-Art Global Climate Model
<p>These directories contain Python (v3) scripts for plotting/analysing model output.</p> <p>Python scripts can be found in the directory 'Program'. Model output can be found in the directory 'Data'.</p> <p>The processed model output are stored as NETCDF files and using the relevant scripts one can regenerate all the figures. We provided the original model output (native grid) and is only converted to yearly-averaged data (due to storage limitations). Some scripts (e.g., FOV_index.py and AMOC_transport.py) use the original model output and running these script generates in the time series, which are presented in the manuscript.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.