Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

33

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

33 results for “state machine”

Learn how ShareScore rates datasets ↗
zenodo48/100

Dataset of "Advanced machine learning techniques for State-of-Health estimation in lithium-ion batteries: A comparative study"

This research focuses on State-of-Health (SOH) estimation of lithium-ion (Li-ion) batteries to enhance lifespan and reliability. Using Samsung INR18650-35E cells, 600 cycles were analyzed with machine learning (ML) techniques, including Gaussian Process Regression (GPR), Support Vector Regression (SVR), Feed-Forward Neural Network (FFNN) and Adaptive Neuro-Fuzzy Inference System (ANFIS). Input features from charging and discharging cycles were selected with Pearson Correlation Analysis (PCA) and Exhaustive Search (ES) to optimize inputs for each ML method. Models were tested on datasets of varying sizes to evaluate performance and overfitting, including an experiment where SOH estimation of one battery was performed using training data from another. The findings highlight each model's strengths and limitations, guiding their application in battery health prediction.

opencc-by-4.0Nov 2024View details →
zenodo44/100

Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities

<p>Numerical data used to make Figures in the manuscript entitled "Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities". We provide the data to create Figures 1 to 4 from the main text and Figures S1 to S9 from the supplementary information. We also provide Python scripts to plot them.</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Machine-Readable Vocabulary Files of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)

<p>This dataset contains two versions of vocabulary files of the&nbsp;<a href="https://ark.staatsbibliothek-berlin.de/">ARK (Alter Realkatalog)</a> in .tsv and .ttl format used for training models for automatic subject indexing with the modular <a href="https://github.com/NatLibFi/Annif">Annif</a> tool. As the ARK is a historical classification system which has been used to describe historical works in the Staatsbibliothek zu Berlin &ndash; Berlin State Library&rsquo;s collections up to 1955, this dataset has been created for generating automatic indexing suggestions for historical texts which have not yet been manually classified with the help of the ARK (for a detailed description of the ARK, see also <a href="../doi/10.5281/zenodo.12783813">Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)</a>. Together with specific corpus training data, these vocabulary files serve as input to Annif, with which the corresponding models on <a href="https://huggingface.co/SBB">Hugging Face at the Staatsbibliothek zu Berlin &ndash; Preu&szlig;ischer Kulturbesitz</a> community have been created. Associated corpus training data have been extracted from the <a href="../doi/10.5281/zenodo.12783813" target="_blank" rel="noopener">Metadata of the "Alter Realkatalog" (ARK)</a> (title data).</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Dataset of "Comprehensive Machine Learning Approaches for Modelling the State of Charge of Lithium-ion Batteries"

<p>This paper evaluates three ML approaches for SOC modeling in LIBs: the multilayer perceptron (MLP), long short-term memory (LSTM), and the nonlinear autoregressive with exogenous input (NARX) neural network architectures. These models were tested using an experimental dataset with multiple input variables, including electrochemical impedance spectroscopy (EIS) data, voltage, and capacity readings for commercial LIB cells. Results indicate that MLP and LSTM are more adaptable with a smaller training dataset (14 samples), while the NARX model required more than 34 out of 67 samples to achieve reasonable accuracy. Additionally, the NARX model is more sensitive to changes in the learning rate (&alpha;) and exhibits larger output error deviations. The MLP and LSTM models consistently performed well across various hidden layer sizes, showing no upper bound constraints, whereas the NARX model&rsquo;s performance deteriorated with certain hidden layer configurations.</p>

embargoedcc-by-4.0Aug 2024View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 2

<p>Supplementary files containing datasets needed to reproduce the results of the manuscript &quot;Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states&quot; by S. Choudhury et al.</p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. param_fixing.zip - self-explanatory (Figure 4 &amp; 5); contains an explanatory note for this part (experiment_details.txt), and the file containing Km values fetched from the BRENDA database (Km_database.csv).</p> <p>2. scripts.zip - scripts to generate figure 2-5 on toy data</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 1

<p><strong>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al (https://doi.org/10.1101/2023.02.21.529387).</strong></p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. models.zip - contains thermodynamically curated steady-state and nonlinear kinetic models of <em>E. coli </em>metabolism used in this study. Also contains the samples of steady-state metabolite concentrations and metabolic fluxes used in the study presented in Figure 3 (steady-state samples used for preparing Figures 2 and 4).</p> <p>2. renaissance_incidence_results.zip - self-explanatory (Figure 2a and 2b)</p> <p>3. ODE_solutions.zip - self-explanatory (Figure 2c)</p> <p>4. bioreactor_simulations1-3.zip - self-explanatory (Figure 2d)</p> <p>5. steady_state_analysis.zip - RENAISSANCE results obtained for each of the steady states (Figure 3a)</p> <p>6. subspace_analysis.zip - RENAISSANCE results presented in Figure 3b-g</p> <p><strong>The remaining datasets are published in the following links</strong></p> <p><em>&nbsp;- https://doi.org/10.5281/zenodo.7930084</em></p> <p><em>&nbsp;- https://doi.org/10.5281/zenodo.10391802</em></p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

XIS: A daily spatiotemporal machine-learning model for environmental exposures in the contiguous United States

<p>These Parquet files contain the outputs used for many analyses and plots in the linked papers. For temperature and humidity, the full sets of observations for cross-validation aren't included because we used restricted-use MADIS data.</p>

opencc-by-sa-4.0Jun 2023View details →
zenodo40/100

Supplementary Material for "To Do or Not to Do: Semantics and Patterns for Do Activities in UML PSSM State Machines"

<p>This dataset provides artifacts about the semantics of doActivity in the Precise Semantics of UML State Machines (PSSM) specification. It collects:</p> <ul> <li>execution traces and screenshots from two simulators (Cameo, Papyrus Moka),</li> <li>analysis about doActivity features present in PSSM test suite, focusing on concurrency,</li> <li>collection of state machine models and doActivity patterns used in the Thirty Meter Telescope (TMT) SysML model.</li> </ul> <p>Find the related paper at <a href="https://arxiv.org/abs/2309.14884" target="_blank" rel="noopener">arXiv:2309.14884</a>.</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Supplementary data: A machine learning approach for dynamical modelling of Al distributions in zeolites via 23Na/27Al solid-state NMR

<p><strong>Content:</strong></p> <p>This dataset provides supplementary data to "A machine learning approach for dynamical modelling of Al distributions in zeolites via 23Na/27Al solid-state NMR". It contains trained Neural Network Potentials (NNP), energy and force data used for accuracy evaluation of the NNPs. Energy and forces are stored as ASE trajectory files (traj), readable by the&nbsp;<a href="https://wiki.fysik.dtu.dk/ase/index.html">Atomic Simulation Environment </a>(ASE). In addition, this repository contains the generated training database with DFT (SCAN+D3(BJ)) energies and forces as SchNetPack1.0 database (SiAlOHNa.db) file readable by ASE and&nbsp;<a href="https://github.com/atomistic-machine-learning/schnetpack/tree/schnetpack1.0">SchNetPack version 1.0</a>.&nbsp; Also, the structure files used to calculate NMR properties are involved.</p> <ul> <li>"nnps.zip" - (pytorch) NNP model files (compatible with <a href="https://github.com/atomistic-machine-learning/schnetpack/tree/schnetpack1.0">SchNetPack version 1.0</a>)</li> <li>"SiAlOHNa.db" - DFT (SCAN+D3(BJ)) training database as SchNetPack1.0 database file readable by ASE and <a href="https://github.com/atomistic-machine-learning/schnetpack/tree/schnetpack1.0">SchNetPack version 1.0</a></li> <li>"error_stats.zip" - traj files storing energies/forces at the DFT (SCAN+D3(BJ)) and NNP level for all test simulations to calcuate energy/force errors</li> <li>"Structures_CHA17.zip" - the structures files of CHA(17).&nbsp;</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo40/100

The state-of-the-art machine learning model for Plasma Protein Binding Prediction: computational modeling with OCHEM and experimental validation

<p><span>Institute of Materia Medica,&nbsp;Chinese Academy of Medical Sciences purchased 10,000 ChemDiv databases.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

State-of-the-Art Review on the Aspects of Martensitic Alloys Studied via Machine Learning

<p>Description</p> <p>The dataset for the&nbsp; review paper titled &quot;State-of-the-Art Review on the Aspects of Martensitic Alloys Studied via Machine Learning&quot; consists of the four files with the names (i) alloy_names.csv, (ii)&nbsp;machine_learning_methods.csv, (iii)&nbsp;nomenclature.csv, and (iv) ptmc_terminologies.csv.</p> <p><strong>(i)&nbsp;alloy_names.csv </strong>: This file presents the summarized list&nbsp;of alloys&#39; names which have been discussed in the review paper.&nbsp;The list thus provides the names of the alloys for which data-driven studies have been attempted to explore one of the effects - martensitic transformation, phase transformation or shape memory effect.&nbsp;</p> <p><strong>(ii)&nbsp;machine_learning_methods.csv</strong> : The machine learning methods that have been discussed in the review paper in relation to the simulation, modeling or prediction tasks in martensitic alloys are listed in this file. This csv file conssits of three columns. The first column &quot;Methods&quot; lists the names of the machine learning methods whereas the second column &quot;Purpose&quot; briefly reveals the objective of the use of the named machine learning method. The final column &quot;Reference&quot; provides the information about the original work (source) from which the data is obtained.&nbsp;</p> <p><strong>(iii)&nbsp;nomenclature.csv </strong>: This file lists all of the acronyms utilized in the review paper, and provides their corresponding full forms.&nbsp;</p> <p><strong>&nbsp;(iv) ptmc_terminologies.csv</strong> : One of the major theories considered significant in the study of martensitic alloys and shape memory effects is&nbsp;phenomenological theory of martensite crystallography (PTMC). The review paper discusses this theory. The different concepts that might be helpful in understanding PTMC , have been assembled in the form of terminologies.&nbsp;</p>

opencc-zeroJun 2023View details →
zenodo36/100

Data for simultaneous inference of sea ice state and surface emissivity model using machine learning and data assimilation

<h2>Overview</h2> <p>This dataset supports the draft manuscript "Simultaneous inference of sea ice state and surface emissivity model using machine learning and data assimilation" which describes a way to infer the daily maps of the sea ice concentration and empirical properties of the sea ice (relating to its snow cover and its physical properties, such as air inclusions) along with the creation of a new empirical model for the sea ice surface emissivity. This is done using knowledge of the atmosphere state, skin temperature and ocean water emissivity from the European Centre for Medium-range Weather Forecasts (ECMWF) weather forecasting model and the observed radiances at microwave frequencies from the Advanced Microwave Scanning Radiometer 2 (AMSR2). The inverse modelling and state estimation is achieved by combining empirical machine learning elements in a Bayesian-inspired network along with a number of physical components. The work also introduces the idea of an "empirical state", in this case describing the aspects of the sea ice physical state which affect the observations, and which is defined by the inputs to the new empirical model component (in machine learning terms, it is defined by the latent input state of a neural network). This dataset includes the &nbsp;data used in training the model and inferring the sea ice parameters, as well as the outputs from that training process. The software used to perform the training is in Python and uses the Keras and Tensorflow software. See the draft manuscript for full details of this data.</p> <p>The code used in the draft manuscript is archived at <a href="https://doi.org/10.5281/zenodo.10013542">https://doi.org/10.5281/zenodo.10013542</a></p> <p>The data used in the draft manuscript is archived at <a href="https://doi.org/10.5281/zenodo.10033377">https://doi.org/10.5281/zenodo.10033377</a></p> <h2>Training data&nbsp;</h2> <h3>Observation space training and ancillary data</h3> <p>Training is done at the location of AMSR2 superobservations (superobs) over ocean with less than 1% land contamination and polewards of 45 degrees latitude, between 1st July 2020 and 30th June 2021. There are 64,184,021 superobs used. A superob is the average of all raw JAXA level 1B observations from one orbit falling into a grid box on an approximately constant area (reduced Gaussian) grid at approximately 40 km by 40 km resolution (noting that polar regions can thus have up to around 7 superobs per day). The superobs have been computed using the field of view central locations for each channel as derived from the JAXA level 1B data. A subset of 10 of the AMSR2 channels is used, from 10 GHz, V polarised, to 89 GHz, H polarised.</p> <p>At each superob location, the relevant fields from the ECMWF 12 hour 'background' forecast are interpolated to the observation time and location. The atmosphere is represented indirectly by the relevant radiative transfer terms from a scattering radiative transfer model. The sea ice concentration from the ECMWF OCEAN5 analysis is included as a validation reference but is not used in the training itself, except to provide a monthly mean first guess to speed up the training. Each field is provided in a separate netCDF file:</p> <ul> <li>field_v2_JULIAN_DAY.nc - superob time in days since 12 UTC on Nov 24th 4714 BC on the proleptic Gregorian calendar</li> <li>field_v2_LAT.nc - superob central latitude in degrees</li> <li>field_v2_LON.nc - superob central longitude in degrees</li> <li>field_v2_IGRID.nc - corresponding grid number on the map grid used in this work (see below)</li> <li>field_v2_OBSVALUE.nc - observed superob brightness temperature at each of 10 AMSR2 channels.</li> <li>field_v2_TSFC.nc - skin temperature computed by the ECMWF forecast model</li> <li>field_v2_WINDSPEED10M.nc - 10m wind speed computed by the ECMWF forecast model</li> <li>field_v2_EMIS_WATER.nc - Ocean water surface emissivity at 10 AMSR2 channels, simulated from the ECMWF forecast fields using the FASTEM-6 model</li> <li>field_v2_CLOUD_FRACTION.nc - Effective cloud fraction used in the atmospheric radiative transfer model at each of 10 AMSR2 channels</li> <li>field_v2_TAUSFC_CLD.nc - Surface to space transmittance in the cloudy column at each of 10 AMSR2 channels</li> <li>field_v2_TUP_CLD.nc - Upwelling brightness temperature from the atmosphere in the cloudy column at each of 10 AMSR2 channels</li> <li>field_v2_TDOWN_CLD.nc - Downwelling brightness temperature from the atmosphere in the cloudy column at each of 10 AMSR2 channels</li> <li>field_v2_TAUSFC.nc - Equivalently for the clear column</li> <li>field_v2_TUP.nc - Equivalently for the clear column</li> <li>field_v2_TDOWN.nc - Equivalently for the clear column</li> <li>field_v2_SEAICE.nc - Sea ice concentration from the ECMWF OCEAN5 analysis, for validation only (not used in training)</li> </ul> <h3>Grid space data: initial data for training; validation sea ice data</h3> <p>A number of properties are provided to the hybrid physical-empirical model that is being trained, on a special map grid defined in this project, including all 62,499 of the reduced Gaussian 40km grid points that have at least one superob at some point during the year of training data. These are:</p> <ul> <li>ifs_seaice_initials_year.nc - sea ice concentration from OCEAN5, monthly averaged on the grid, and then provided on all days of the relevant month as initial conditions (technically, first guess) for the training. This includes an additional day before the beginning of the training, used for time-lagging (see draft paper).</li> <li>ifs_tsfc_year_dailyx.nc - skin temperature from ECMWF forecast fields at observation locations, averaged onto the daily grid, to help provide constraints on the likelihood of sea ice as part of a sea ice loss function.</li> </ul> <p>For diagnostic and validation purposes, the ECMWF OCEAN5 analysis is also provided on the grid:</p> <ul> <li>ifs_seaice_year.nc - sea ice concentration from OCEAN5 at observation locations, averaged onto the daily grid</li> </ul> <p>All these fields are provided on the following dimensions:</p> <ul> <li>LON - the longitude of the grid point in degrees</li> <li>DAY - the day through the training year (0-364, 1st July 2020 to 30th June 2021) or through the training year extended forward by one day (30th June 2020) for the sea ice (0-365). In practice the days are offset by 3 hours from the UTC day to match the ECMWF data assimilation windows, which start at 21 UTC the day before.</li> </ul> <p>The latitude is also provided</p> <ul> <li>LAT - the latitude of the grid point in degrees</li> </ul> <p>Note that the observation location IGRID is on the custom grid of the ML model that is defined implicitly in these gridded files. The LON and LAT vectors in these files are the longitude and latitude points of the grid and are of 62499 in length. The IGRID number for an observation is the index into these arrays from 0-62498.</p> <h2>Outputs from training</h2> <p>The following files are the output and diagnostics from the year-long training. The python code and the draft paper are the primary documentation for these:</p> <ul> <li>models_year.nc - settings of the model are recorded here, along with the trained values of the smaller empirical components/layers within the hybrid model. For example, the layer weights of the wind speed bias correction, the observation space bias correction, and the empirical surface emissivity model are recorded here. The values of the loss function at each epoch are also recorded here.</li> <li>properties_year.nc - trained values of each of 3 empirical properties of sea ice on the map grid (3 properties by 62499 locations by 365 days from 1st July 2020)</li> <li>seaice_year.nc - inferred values of sea ice fraction on the map grid (62499 locations by 365 days from 1st July 2020, discarding the additional day at the start)</li> <li>tbsim_year.nc - simulated AMSR2 brightness temperatures from the trained network</li> <li>tbsim_initial_year.nc - simulated AMSR2 brightness temperatures using the untrained network</li> </ul> <p>The longitude and latitude of the map grid is found in any of the initial data files described in the previous section. The days are 0-364 corresponding to 1st July 2020 to 30th June 2021.</p> <h3>Sea ice surface emissivity at grid locations</h3> <p>A packaged version of the sea ice surface emissivity is provided at grid locations, alongside the surface emissivity model, the sea ice concentration and the four inputs to the model, i.e. the normalised skin temperature and the three empirical variables:</p> <ul> <li>emissivity_grid_year.nc</li> </ul> <p>Note that in the training, the surface emissivity is computed at observation locations and has not been stored due to memory limitations. For easier comparison to other datasets, the surface emissivity has been recomputed on grid locations in this package, using the year-long trained emissivity model and its trained inputs. The sea ice surface emissivity is only physically meaningful for sea ice concentrations above around 0.25. Also be aware of the "hole at the pole" which is the small region of the Arctic ocean that is sometimes not covered by an AMSR2 overpass, and which is found from 88 degrees N. On days where the hole or part of the hole exists, the sea ice emissivity on the grid is not valid at these locations. These locations can be identified by having all values of the empirical properties zero (because the empirical properties were never constrained by any observations on that day, and remain at their initial values before training).</p> <h2>Sensitivity tests</h2> <p>Extensive sensitivity tests were carried out, as described in the appendices of the draft paper and as documented in the Python code, using the month of August 2020 as an example. These required equivalent month-long training and initial data similar to those described above, but all observation space fields are contained within the same file in this case. Output files follow similar principles to those described above. The full package is provided as a tar file:</p> <ul> <li>sensitivity.tar</li> </ul> <p>This contains the training and initial files:</p> <ul> <li>amsr2_v2_202008.nc</li> <li>ifs_tsfc_dailyx_202008.nc</li> <li>ifs_seaice_202008.nc</li> </ul> <p>as well as directories containing the trained model outputs and diagnostics at each of the sensitivity tests, using the same formats as described for the yearly training, with these names:</p> <ul> <li>nprop - number of empirical properties</li> <li>epoch - number of epochs</li> <li>deep - configuration of the empirical sea ice emissivity model, including multiple layers of nonlinear dense neural network</li> <li>bseaice - background error for the sea ice physical bounds background error (loss) term</li> <li>bemis - background error for the sea ice emissivity background error (loss) term</li> <li>bbias - background error for the bias correction background error (loss) term</li> <li>batchsize - batch size used in training</li> <li>bbatchsize - extended epochs testing of batch size used in training</li> </ul> <h2>Licensing</h2> <p>This data product is published under a Creative Commons Attribution 4.0 International (CC BY&nbsp;4.0). To view a copy of this licence, visit <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>You are free to:</p> <ul> <li>Share &mdash; copy and redistribute the material in any medium or format</li> <li>Adapt &mdash; remix, transform, and build upon the material&nbsp;for any purpose, even commercially.</li> </ul> <p>Under the following terms:</p> <ul> <li>You must give appropriate credit (attribution) to ECMWF as outlined below, provide a link to the licence, and indicate if changes were made.</li> <li>No additional restrictions &mdash; You may not apply legal terms or technological measures that legally restrict others from doing anything the licence permits.</li> </ul> <p>The following wording shall be attached to the use of this ECMWF data product:&nbsp;</p> <ol> <li>Copyright statement: Copyright "&copy; 2023 European Centre for Medium-Range Weather&nbsp;Forecasts (ECMWF)".</li> <li>Source <a href="http://www.ecmwf.int/">www.ecmwf.int </a>and <a href="https://doi.org/10.5281/zenodo.10009497">https://doi.org/10.5281/zenodo.10009497</a></li> <li>Licence Statement: This data is published under a Creative Commons Attribution 4.0&nbsp;International (CC BY 4.0). <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></li> <li>Disclaimer: ECMWF does not accept any liability whatsoever for any error or omission in&nbsp;the data, their availability, or for any loss or damage arising from their use.</li> <li>Where applicable, an indication if the material has been modified and an indication of previous modifications.</li> <li>DOI: 10.5281/zenodo.10009498</li> </ol> <p>Original data for this value-added product was provided by Japan Aerospace Exploitation Agency (JAXA). Specifically, this dataset builds on the Advanced Microwave Scanning Radiometer 2 (AMSR2) level 1B data available from the JAXA G-Portal, https://gportal.jaxa.jp/gpr/, which has the following attribution and licensing:</p> <ol> <li>Give credit for the original data to JAXA, i.e. "Original data for this value added data product was provided by Japan Aerospace Exploration Agency"</li> <li>DOI for original JAXA data is L1B-Brightness temperature (TB) GCOM-W/AMSR2 L1B Brightness Temperature: <a href="https://doi.org/10.57746/EO.01gs73ans548qghaknzdjyxd2h">https://doi.org/10.57746/EO.01gs73ans548qghaknzdjyxd2h</a></li> <li>Original terms of data service from JAXA, with highlighted extracts: <ul> <li><a href="https://gportal.jaxa.jp/gpr/index/eula?lang=en">https://gportal.jaxa.jp/gpr/index/eula</a> <ol> <li>The user is entitled to use G-Portal data free of charge without any restrictions (including commercial use) except for the condition about acknowledgement of data credit as stipulated in Article 7.(2). (see above)</li> <li>JAXA is collecting results (papers, theses, reports, etc.) using G-Portal data. If you have any results using G-Portal data, please mail/e-mail a copy of the result to G-Portal Support Desk (Contact Information written at the end of the Terms of Use). We appreciate your cooperation very much.</li> </ol> </li> </ul> </li> </ol>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Sensing Performance of Artificially Intelligent Nanopores Developed by Integrating Solid-State Nanopores with Machine Learning Methods

<p>Ionic current-time data obtained from measuring nanoparticles with the diameters of 90, 100, 150, 200, 220, 270, and 300 nm, using nanopores with a diameter of 300nm.</p> <p>Test_100nm_1 means a test data of nanoparticles with a diameter of 100 nm.</p> <p>Train_100nm_1 means a training data of&nbsp;nanoparticles with a diameter of 100 nm.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Low-cost prediction of molecular and transition state partition functions via machine learning

<p>This dataset contains the vibrational, rotational, translational, and electronic partition functions for 35,883 organic chemistry molecular structures taken from the Grambow et. al dataset [1]. It was used to train ML deep neural networks to predict unknown transition state partition functions as well as partition functions for known molecular structures [2]</p> <p>The partition functions were computed at temperatures in the range T= [50, 2000] K with the rigid rotor, rigid body, harmonic oscillator approximations. Reactions involve no more than 7 C, N, or O atoms.</p> <p>Frequencies for the vibrational partition functions were taken from [1] where they were computed with DFT at the &omega;B97X-D3/def2-TZVP level of theory.</p> <p>For the rotational partition function, symmetry numbers were obtained by evaluating proper and improper invariant rotations of the structures. We note that structures involving two molecules were not separated: vibrational frequencies and symmetry numbers were computed for the aggregate structure.</p> <p>For each reaction, partition functions were calculated at 50 temperatures sampled uniformly from the inverse temperature range 1/T = [1/2000, 1/50] K<sup>-1</sup>. This corresponds to 11,961 reactions, 35,883 total structures, and 1,794,150 total partition function examples.</p> <p>The file Partition_Functions.tar.gz contains directories entitled &ldquo;rxnXXXXXX&rdquo; where XXXXXX is a reaction number identifier. Each contain three files &ldquo;rXXXXXX.csv&rdquo;, &ldquo;pXXXXXX.csv&rdquo;, &ldquo;tsXXXXXX.csv&rdquo; corresponding to data from the reactant (&ldquo;r&rdquo;), product (&ldquo;p&rdquo;), and transition state (&ldquo;ts&rdquo;) for reaction XXXXXX. Note that the directory structure and the reaction identifiers are the same as used in the original structure dataset by Grambow et al. and the corresponding structures can easily be extracted from that dataset. Each comma separated value (csv) file contains 50 rows and the following columns:</p> <table> <tbody> <tr> <td> <p><strong>&nbsp;Column label</strong></p> </td> <td> <p><strong>&nbsp;Values</strong></p> </td> </tr> <tr> <td> <p>&nbsp;T [K]</p> </td> <td> <p>&nbsp;Temperature</p> </td> </tr> <tr> <td> <p>&nbsp;qpart_ele [unitless]</p> </td> <td> <p>&nbsp;Electronic partition function</p> </td> </tr> <tr> <td> <p>&nbsp;qpart_trans [unitless]</p> </td> <td> <p>&nbsp;Translational partition function</p> </td> </tr> <tr> <td> <p>&nbsp;qpart_vib [unitless]</p> </td> <td> <p>&nbsp;Vibrational partition function</p> </td> </tr> <tr> <td> <p>&nbsp;qpart_rot [unitless]</p> </td> <td> <p>&nbsp;Rotational partition function</p> </td> </tr> <tr> <td> <p>&nbsp;qpart [unitless]</p> </td> <td> <p>&nbsp;Partition function</p> </td> </tr> <tr> <td> <p>&nbsp;log_qpart_trans [unitless]&nbsp;</p> </td> <td> <p>&nbsp;Natural logarithm of translational partition function</p> </td> </tr> <tr> <td> <p>&nbsp;log_qpart_rot [unitless]</p> </td> <td> <p>&nbsp;Natural logarithm of rotational partition function</p> </td> </tr> <tr> <td> <p>&nbsp;log_qpart_vib [unitless]</p> </td> <td> <p>&nbsp;Natural logarithm of vibration partition function</p> </td> </tr> <tr> <td> <p>&nbsp;log_qpart [unitless]</p> </td> <td> <p>&nbsp;Natural logarithm of partition function</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>[1] &nbsp;&nbsp;&nbsp;&nbsp; C. A. Grambow, L. Pattanaik, and W. H. Green, &ldquo;Reactants, products, and transition states of elementary chemical reactions based on quantum chemistry,&rdquo; <em>Sci. Data</em>, <strong>7</strong>:1&ndash;8, 2020.</p> <p>[2]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Komp, E. Valleau, S. &ldquo;Low-cost prediction of molecular and transition state partition functions via machine learning&rdquo;, arXiv:, 2022.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Machine Learning the Hohenberg-Kohn Map to Molecular Excited States

<p>The dataset and code used in paper&rdquo;Machine Learning the Hohenberg-Kohn Map to Molecular Excited States&rdquo; For detailed information of each file, see Readme.txt</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Machine readable code lists for an algorithm to identify incident non-small cell lung cancer (NSCLC) in United States healthcare claims data

<p>Machine readable code lists for an algorithm to identify incident non-small cell lung cancer (NSCLC) in United States healthcare claims data</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Machine Learning-Assisted Discovery of Hidden States in Expanded Free Energy Space

<p>Collective variables (CVs) are crucial parameters in enhanced sampling calculations and strongly impact the quality of the obtained free energy surface. However, many existing CVs are unique to and dependent on the system they are constructed with, making the developed CV non-transferable to other systems. Herein, we develop a non-instructor-led deep autoencoder neural network (DAENN) for discovering general-purpose CVs. The DAENN is used to train a model by learning molecular representations upon unbiased trajectories that contain only the reactant conformers. The prior knowledge of nonconstraint reactants coupled with the here-introduced topology variable and loss-like penalty function are only required to make the biasing method able to expand its configurational (phase) space to unexplored energy basins. Our developed autoencoder is efficient and relatively inexpensive to use in terms of <em>a priori</em> knowledge, enabling one to automatically search for hidden CVs of the reaction of interest.</p>

opencc-by-4.0Feb 2022View details →
dryad36/100

Estimated roadway segment traffic data by vehicle class for the United States: A machine learning approach

Open the record for dataset details and reuse information.

publicApr 2025View details →
zenodo32/100

Dataset for "The State of the ML-universe: 10 Years of Artificial Intelligence & Machine Learning Software Development on GitHub"

<p>Supplementary data to &quot;The State of the ML-universe: 10 Years of Artificial Intelligence &amp; Machine Learning Software Development on GitHub&quot; accepted for publication at MSR 2020.</p> <p>The data included in this package were used to conduct analyses to characterize the AI &amp; ML software development community hosted on GitHub. Please read the paper for a full understanding of what data was collected and how it was used.</p> <p>Questions and comments can be directed to Danielle Gonzalez dng2551@rit.edu</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

ACE: Abstract Consensus Encapsulation for Liveness Boosting of State Machine Replication (video)

Full video presentation of the paper: ACE: Abstract Consensus Encapsulation for Liveness Boosting of State Machine Replication.<br><br>Appears in Session 3 of the 24th International Conference on Principles of Distributed Systems OPODIS 2020<br><a href="https://opodis2020.unistra.fr">https://opodis2020.unistra.fr</a>

opencc-by-4.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record