Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

Supplementary data for "Machine learning-based prediction of activity and substrate specificity for OleA enzymes in the thiolase superfamily"

<p>Supplementary data for &quot;Machine learning-based prediction of activity and substrate specificity for OleA enzymes in the thiolase superfamily&quot;</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Data bundle for "Spherical-angular dark field imaging and sensitive microstructural phase clustering with unsupervised machine learning"

<p>Prepared by Tom McAuliffe (t.mcauliffe17@imperial.ac.uk)</p> <p>This repository is a release of the raw data and analysis results for: &#39;Spherical-angular dark field imaging and sensitive microstructural phase clustering with unsupervised machine learning&#39;&nbsp;</p> <p>The raw data is given as &#39;yprime.h5&#39; - this contains patterns&nbsp;and metadata in the Bruker-exported format.</p> <p>Scripts for dataset decomposition into latent factors are given in &#39;Scripts&#39;.</p> <p>Our spherical analysis code is included in &#39;SphericalAngleDF&#39;.</p> <p>Outputs of our analysis code&nbsp;are contained in &#39;Analysis&#39;.</p> <p>Figures for the paper are included in &#39;Figures&#39;.</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Dataset for "Machine Learning Stability and Bandgaps of Lead-Free Perovskites for Photovoltaics"

<p>Datasets used in the publication &quot;Machine Learning Stability and Bandgaps of Lead-Free Perovskites for Photovoltaics&quot;&nbsp; [doi:10.1002/adts.201900178].</p> <p>All structures were relaxed with the following parameters using Quantumwise QATK 2017:</p> <p>- SG15-GGA norm-conserving (Vanderbilt) pseudopotentials employed in a LCAO-approach (200 Hartree cutoff)<br> - 2x1x2-cubic-perovskite-supercells, relaxed from cubic 11.4&Aring;x5.7&Aring;x11.4&Aring;-structures (forces &lt; 0.01eV/&Aring;)<br> - 300K Fermi-Dirac-smearing<br> - a 6x12x6 k-point grid (Monkhorst-Pack)</p> <p><br> Specifically, the included files are:</p> <p><strong>db_2.data: </strong>the actual database used for model building (json-format)<br> <strong>lead_set.data:</strong> the &quot;external&quot; test set used to test predictive power with out of sample compounds (json-format)<br> <strong>load_stanley_c.py:</strong> a python script to parse the .json-files to a python-dictionary including the structures (relaxed and unrelaxed) as <a href="https://gitlab.com/ase/ase">ASE</a>-atoms</p> <p>The format of the datafiles is as follows (-1 generally denote values not parsed from the raw data):<br> {<br> &nbsp;&nbsp;&nbsp;&nbsp;&quot;&lt;idstring&gt;&quot; : {<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;trajectory&quot; : n/a,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;energy&quot; : total DFT energy in eV,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;rstruc&quot; : relaxed structure, 3-tuple: (cell-vectors, scaled_positions, elements),<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;gaps&quot; : { &quot;opt_gap&quot;, &quot;ind_gap } - both direct and indirect gap,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;effective_mass&quot; : n/a,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;iterations&quot; : number of relaxation steps,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;calc&quot; : some calculation metadata,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;ustruc&quot; : unrelaxed input structure,<br> &nbsp;&nbsp;&nbsp;&nbsp;<br> &nbsp;&nbsp;&nbsp;&nbsp;}<br> }<br> Missing ids relate to structures filtered out, because the calculation didn&#39;t converge.</p> <p>Some code which works with a different representation of this data can be found at&nbsp; https://github.com/jstanai/Machine-Learning-Perovskite-Properties-for-Photovoltaics</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

GAP-20 machine learning force field for phosphorus

<p>This dataset contains the force-field parameter files and reference database described in the manuscript &quot;A general-purpose machine-learning force field for bulk and nanostructured phosphorus&quot; (to be published).</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Data for: Machine learning identifies robust matrisome markers and regulatory mechanisms in cancer

<p>The expression and regulation of matrisome genes - the ensemble of extracellular matrix, ECM, ECM-associated proteins and regulators as well as cytokines, chemokines and growth factors - is of paramount importance for the many biological processes and signals within the tumor microenvironment. The availability of large and diverse multi-omics data enables mapping and understanding the regulatory circuitry governing the tumor matrisome to an unprecedented level, though such a volume of information requires robust approaches to data analysis and integration. In this study, we show that combining Pan-Cancer expression data from The Cancer Genome Atlas (TCGA) with genomics, epigenomics and microenvironmental features from TCGA and other sources enables the identification of &ldquo;landmark&rdquo; matrisome genes and machine learning-based reconstruction of their regulatory networks in 74 clinical and molecular subtypes of human cancers and approx. 6700 patients. These results, enriched for prognostic genes and cross-validated markers at the protein level, unravel the role of genetic and epigenetic programs in governing the tumor matrisome and allow the prioritization of tumor-specific matrisome genes (and their regulators) for the development of novel therapeutic approaches.</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Systematic Data Analysis and Diagnostic Machine Learning Reveal Differences between Compounds with Single- and Multitarget Activity

<p>The deposited files contain balanced data sets of multi-target (MT) and single-target (ST) compounds (CPDs) used for machine learning studies (https://dx.doi.org/10.1021/acs.molpharmaceut.0c00901).&nbsp; The first file (st_mt_data.tsv) contains 15,142 MT- and 15,081 ST-CPDs and the second (st_dt_data.tsv)&nbsp; 1828 DT- and 1776 ST-CPDs. For each CPD, a nonstereo_aromatic_SMILES representation, the original ChEMBL_cid, UniProt (target) IDs, and CPD category (CPD_CAT) (i.e. DT/MT/ST) is provided. DT stands for &#39;diverse-target&#39; and denotes a subset of MT-CPDs (as detailed in the publication). In addition, a CPD is tagged &ldquo;Y&rdquo; if it continued to be present in the data set after removal of 50% randomly selected CPDs or 50%&nbsp; CPD nearest neighbors (NN), respectively.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Adele 3D seismic survey segy format used in the FORCE 2020 machine learning competition for fault identification

<p>Adele seismic 3D&nbsp; survey segy format used in the FORCE 2020 machine learning competition for fault identification.</p> <p>Dataset is courtesy of GEOSCIENCE Australia who need to be acknowledged in each publication</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Airborne Radar Quality Control with Machine Learning

<p>This repository contains the radar data collected by ELDORA required to train and test the random forest model discussed in&nbsp;&quot;Airborne radar quality control with machine learning&quot; by Alexander DesRosiers and Michael M. Bell at the Colorado State University Department of Atmospheric Science. The model used in the manuscript&nbsp;is also contained in a &#39;.pkl&#39; file.&nbsp;Upon publication, a link to the paper will be provided here.&nbsp;Finer points of the methodology were discussed in the manuscript and the python script (make_radarQC_rf_model.py)&nbsp;is commented to guide users through the process of creating the model.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Combinatorial and machine learning approaches for the analysis of Cu2ZnGeSe4: influence of the off-stoichiometry on defect formation and solar cell performance

<p>Dataset of the results published in the&nbsp;<a href="https://zenodo.org/record/4742379#.YMzExOgzYmJ">J. Mater. Chem. A, 2021, 9, 10466</a>. The files represent: i)&nbsp;the measured compositional and optoelectronic data of each solar cell, as well as the data generated from the Raman spectra analysis; ii) Raman spectra of the representative cells; iii) Machine Learning discriminants.</p> <p>The elemental composition of the different cells of the combinatorial sample was determined by X-ray fluorescence (XRF) using a Fischerscope XDV system with a 1 mm spot diameter, a 50 kV acceleration voltage, a Ni10 lter and a 45 s acquisition time. Raman analysis with blue (442 nm) and green (532 nm) excitation wavelengths were performed on the bare absorber, while measurements with NIR (785 nm) were performed in complete devices using Horiba Jobin Yvon FHR640 and iHR320 monochromators coupled with CCD detectors. The first monochromator is optimized for the UV and visible spectral ranges and was used with 442 nm (He&ndash;Cd gas laser) and 532 nm (solid state laser) excitation wavelengths. The second monochromator is optimized for the NIR range and was used with a 785 nm (solid state laser) excitation wavelength. The power&nbsp;density of the lasers was kept below 150 W cm<sup>2</sup> and the spot size was ~70 <span class="math-tex">\(\mu\)</span>m. The measurements were performed in a backscattering configuration through a specific probe designed at IREC. The J&ndash;V characteristics of the devices were obtained under simulated AM1.5 illumination (1000 W m2 intensity at room temperature) using a pre-calibrated Class AAA solar simulator (Abet Technologies Sun 3000).</p>

opencc-by-3.0Apr 2021View details →
zenodo44/100

Dataset of experimental measurements for "Demonstration of quantum advantage in machine learning"

<p>Dataset of experimental measurements for "Demonstration of quantum advantage in machine learning",  <em>npj Quantum Information</em><strong> 3</strong>, Article number: 16 (2017).</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

Dataset supporting "Using Machine Learning to decide when to Precondition Cylindrical Algebraic Decomposition with Groebner Bases"

<p>Dataset supporting the paper:</p> <p>Z. Huang, M. England, J.H. Davenport and L.C. Paulson<br> Using Machine Learning to decide when to Precondition Cylindrical Algebraic Decomposition with Groebner Bases.<br> Proceedings of the 18th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC '16), pp. 45--52. IEEE, 2016. Digital Object Identifier: 10.1109/SYNASC.2016.020 </p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

TCOM-H2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric H2O profile dataset [1991-2021] constructed using machine-learning.

<p>Methodology: &nbsp;</p> <p>The <strong>TOMCAT simulation</strong> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized <strong>ERA-5 reanalysis data</strong>.</p> <h3>H2O Profile Processing and Bias Correction</h3> <p><strong>Collocated H2O profiles</strong> are organized into five distinct latitude bins:</p> <ul> <li> <p><strong>NH polar</strong>: 90∘N - 50∘N</p> </li> <li> <p><strong>NH mid-lat</strong>: 20∘N - 70∘N</p> </li> <li> <p><strong>Tropics</strong>: 40∘S - 40∘N</p> </li> <li> <p><strong>SH mid-lat</strong>: 70∘S - 20∘S</p> </li> <li> <p><strong>SH polar</strong>: 90∘S - 50∘S</p> </li> </ul> <p>Initially, <strong>differences between TOMCAT and satellite measurements</strong> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from 10,km to 60,km). Note that TOMCAT may not accurately capture H2O evolution post-HTHH eruption due to the sparse spatial coverage of ACE measurements, which limits training data.</p> <p><strong>Separate XGBoost regression models</strong> are then trained for these H2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate <strong>H2O bias corrections</strong> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</p> <p><strong>Height-resolved H2O profile data</strong> are then interpolated onto 28 standard pressure levels (from 300,hPa to 0.1,hPa), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</p> <p>We acknowledge the inherent <strong>dry biases in the original TOMCAT H2O profiles</strong>, largely because the TTL entry mixing ratios are based on a simplistic sinusoidal seasonal cycle, which omits the H2O enhancement contributed by tropical convective clouds.</p> <h3>Data Files</h3> <p>The dataset includes two files containing daily mean zonal mean H2O profiles:</p> <ul> <li> <p><code>zmh2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>height level data</strong> (10,km to 60,km).</p> </li> <li> <p><code>zmh2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>pressure level data</strong> (300,hPa to 0.1,hPa).</p> </li> </ul> <h3>Reference Publication</h3> <p>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</p> <p>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, <a title="null" href="https://doi.org/10.5194/essd-15-5105-2023">https://doi.org/10.5194/essd-15-5105-2023</a>, 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

TCOM-O3: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric ozone profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;TOMCAT simulation is performed at T64L32 resolution for the 2000-2024 time period. Collocated Ozone (O3) profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement &nbsp;differences are calculated for each zonal bins (51 height levels, 10km to 60km). Note that if enough ACE measurements are not avaliable for a particular level then data is purely based on TOMCAT simulated output field. Separate XGBoost regression models are trained for the &nbsp;differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids. &nbsp;TOMCAT output sampled at 1.30 pm local time at the equator. Estimated corrections for a given model grid that are added to the original TOMCAT simulated day and night time ozone profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values. &nbsp;For more details see attached presentation. Previous version use both HALOE and ACE data. Here only ACE data is used (hence starting date is 01 January 2000). PDF file shows comparison between v1.0 and v1.1 as well as TOMCAT data.</p> <p>Dataset also includes two files containing daily mean zonal mean hydrogen fluoride &nbsp;profiles on height (10-60 km) and pressure (300-0.1 hPa) levels:</p> <p>zmo3_TCOM_hlev_T2Dz_2000_2024.nc &ndash; height level data (10 to 60 km)</p> <p>zmo3_TCOM_plev_T2Dz_2000_2024.nc &ndash; pressure level data (300 to 0.1 hPa)</p> <p>Daily 3D profiles on height and pressure levels would be made available on request.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

CLRD-GLPS: A Long-term Seasonal Dataset of Ruminant Livestock Distribution in China's Grazing Production Systems (2000-2021) Using Stacking-based Interpretable Machine Learning

<p>Advanced computational methods integrating ensemble learning with interpretable machine learning are essential for precision livestock management under increasing environmental constraints and food security pressures. This study develops a novel stacking-based interpretable machine learning (IML) framework that combines multiple algorithms with SHAP analysis techniques to generate the China's Long-term Ruminant Livestock Distribution in Grazing Livestock Production Systems (CLRD-GLPS) dataset. Our computational approach addresses critical challenges in livestock distribution modelling: livestock segmentation and spatial prediction accuracy. The framework integrates Random Forest, XGBoost, CatBoost, LightGBM, and Extra Trees through a two-layer stacking architecture, enhanced with SHAP (Shapley Additive Explanations) analysis for model interpretability. We also implemented interpretable machine learning for livestock production system segmentation to distinguish grazing from total livestock populations. The stacking ensemble demonstrated superior performance over individual algorithms, achieving R&sup2; values of 0.954-0.961 for cattle and 0.896-0.901 for sheep and goats, with improvements of up to 8.3% compared to best performance single-model approaches. Multi-scale validation confirmed computational robustness: livestock segmentation achieved R&sup2; = 0.80 at county level, while independent city-level validation of CLRD-GLPS datasets yielded R&sup2; = 0.76-0.80. SHAP interpretability analysis revealed distinct environmental drivers, with vegetation indices and topography primarily influencing cattle distribution, while snow conditions and elevation dominated sheep and goat patterns. This computational framework advances livestock distribution modelling through enhanced prediction accuracy, model stability, and interpretability, while the CLRD-GLPS dataset provides essential spatial-temporal information for rangeland sustainability assessments and evidence-based livestock management policies. This dataset is supported by the Second Tibetan Plateau Scientific Expedition and Research Program (STEP, grant no. 2019QZKK0906).</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

GTWS-MLrec: Global terrestrial water storage reconstruction by machine learning from 1940 to present

<p>Terrestrial water storage (TWS) includes all forms of water stored on and below the land surface, and is a key determinant of global water and energy budgets. However, TWS data from measurements by the Gravity Recovery and Climate Experiment (GRACE) satellite mission are only available from 2002, limiting global and regional investigation of the long-term trends and variabilities in the terrestrial water cycle under climate change. This study presents long-term (i.e., 1940-2022) and high-resolution (i.e., 0.25°) monthly time series of TWS anomalies over the global land surface. The reconstruction is achieved by using a set of machine learning models with a large number of predictors, including climatic and hydrological variables, land use/land cover data, and vegetation indicators (e.g., leaf area index). The outcome, machine learning-reconstructed TWS estimates (i.e., GTWS-MLrec), fits well with the GRACE/GRACE-FO measurements, showing high correlation coefficients and low biases in the GRACE era. We also evaluate GTWS-MLrec with other independent datasets such as the land-ocean mass budget, large-scale water balance in 341 large river basins, and streamflow measurements at 10,168 gauges. We find that the proposed approach performs overall as well as or is more reliable than previous TWS datasets. Moreover, our reconstructions successfully reproduce the impact of climate variability, such as strong El Niño events. GTWS-MLrec dataset consists of three reconstructions based on JPL, CSR and GSFC mascons, three detrended and de-seasonalized reconstructions, and six global average TWS series over land areas, both with and without Greenland and Antarctica. Along with its extensive attributes, GTWS_MLrec can support a broad range of applications such as better understanding the global water budget, constraining and evaluating hydrological models, climate-carbon coupling, and water resources management.</p><p>Please cite the reference: <strong>Yin J, Slater L, Khouakhi A, et al. GTWS-MLrec: Global terrestrial water storage reconstruction by machine learning from 1940 to present. Earth System Science Data. 2023.</strong></p><p>For any inquiry about the dataset, welcome to contact Dr. Jiabo Yin (jboyn@whu.edu.cn).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

MLFMF: Data Sets for Machine Learning for Mathematical Formalization

<h3>MLFMF</h3><p><strong>MLFMF (Machine Learning for Mathematical Formalization) </strong>is a collection of data sets for benchmarking recommendation systems used to support formalization of mathematics with proof assistants. These systems help humans identify which previous entries (theorems, constructions, datatypes, and postulates) are relevant in proving a new theorem or carrying out a new construction.&nbsp;</p><p>The MLFMF data sets provide solid benchmarking support for further investigation of the numerous machine learning approaches to formalized mathematics. With more than 250,000 entries in total, this is currently the largest collection of formalized mathematical knowledge in machine learnable format.&nbsp;</p><p>In addition to benchmarking the recommendation systems, the data sets can also be used for benchmarking <strong>node classification</strong> and <strong>link prediction</strong> algorithms.&nbsp;</p><h3>The four data sets</h3><p>Each data set is derived from a library of formalized mathematics written in proof assistants <a href="https://agda.readthedocs.io/en/v2.6.4/"><i>Agda</i></a> or <a href="https://lean-lang.org/"><i>Lean</i></a>. The collection includes &nbsp;</p><ol><li>the largest Lean 4 library <a href="https://github.com/leanprover-community/mathlib4"><strong>Mathlib</strong></a>,</li><li>the three largest Agda libraries:<ul><li>the <a href="https://github.com/agda/agda-stdlib"><strong>standard library</strong></a></li><li>the library of univalent mathematics <a href="https://github.com/UniMath/agda-unimath"><strong>Agda-unimath</strong></a>, and</li><li>the <a href="https://github.com/martinescardo/TypeTopology"><strong>TypeTopology</strong></a> library.</li></ul></li></ol><p>Each data set represents the corresponding library in two ways: as a heterogeneous network, and as a list of syntax trees of all the entries in the library. The network contains the (modular) structure of the library and the references between entries, while the syntax trees give complete and easily parsed information about each entry.</p><p>The Lean library data set was obtained by converting <strong>.olean</strong> files into s-expressions (see the <a href="https://github.com/andrejbauer/lean2sexp"><strong>lean2sexp</strong></a> tool).</p><p>The Agda data sets were obtained with an <a href="https://github.com/andrejbauer/agda/tree/master-sexp">s-expression extension</a> of the official Agda repository (use either master-sexp or release-2.6.3-sexp branch).</p><p>For more details, see our <a href="https://arxiv.org/abs/2310.16005"><strong>arXiv copy</strong></a><strong> </strong>of the paper.</p><h3>Directory structure</h3><p>First, the <strong>mlfmf.zip</strong> archive needs to be unzipped. It contains a separate directory for every library (for example, the standard library of Agda can be found in the stdlib directory) and some auxiliary files. Every library directory contains</p><ul><li>the <strong>network file</strong> from which the heterogeneous network can be loaded,</li><li>a zip of the <strong>entries directory</strong> that contains (many) files with abstract syntax trees. Each of those files describes a single entry of the library.</li></ul><p>In addition to the auxiliary files which are used for loading the data (and described below), the zipped sources of lean2sexp and Agda s-expression extension are present.</p><h4>Loading the data</h4><p>In addition to the data files, there is also a simple python script <strong>main.py</strong> for loading the data. To run it, you will have to install the packages listed in the file <strong>requirements.txt</strong>: <strong>tqdm</strong> and <strong>networkx</strong>. The easiest way to do so is calling <i><strong>pip install -r requirements.txt</strong></i>.</p><p>When running <strong>main.py </strong>for the first time, the script will unzip the entry files into the directory named <strong>entries</strong>. After that, the script loads the syntax trees of the entries (see the <strong>Entry</strong> class) and the network (as <i>networkx.MultiDiGraph</i> object).</p><p><i>Note. The entry files have extension <strong>.dag </strong>(directed acyclic graph), since Lean uses node sharing, which breaks the tree structure (a shared node has more than one parent node).</i></p><h3>More information</h3><p>For more information about the <strong>data collection process</strong>, <strong>detailed data (and data format) description</strong>, and <strong>baseline experiments</strong> that were already performed with these data, see our <a href="https://arxiv.org/abs/2310.16005"><strong>arXiv copy</strong></a><strong> of the paper</strong>.</p><p>For the code that was used to perform the experiments and data format description, visit our github repository <a href="https://github.com/ul-fmf/mlfmf-data"><strong>https://github.com/ul-fmf/mlfmf-data.</strong></a></p><h3>Funding</h3><p>Since not all the funders are available in the Zenodo's database, we list them here:</p><ol><li>This material is based upon work supported by the Air Force Office of Scientific Research under award number FA9550-21-1-0024.</li><li>The authors also acknowledge the financial support of the Slovenian Research Agency via the research core funding No. P2-0103 and No. P1-0294.</li></ol><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

A machine learning-based high-precision density functional method for drug-like molecules

<h2><strong>Models</strong></h2><p>The repo contains the models and test datasets for our aticles. The energy unit is in <strong>Hartree,</strong> The coordinate unit is in<strong> Bohr.</strong></p><p><strong>## DeePHF</strong></p><p>you need first prepare the `dm_eig.npy` in data_test and do predict `l_e_delta.npy`, you can use</p><p>```</p><p>deepks test -m model.pth -o test/test -d data_test/* -D dm_eig -G</p><p>```</p><p><strong>## DeePKS</strong></p><p>first you should prepare the `atom.npy`, and `energy.npy` in data_test. you can test the datasets by command.&nbsp;</p><p>```</p><p>deepks scf scf_input.yaml -m model.pth -s data_test -d test_out</p><p>```</p><p><strong># Datasets</strong></p><p>All datasets only have `atom.npy` and `energy.npy`. The coordinate unit is `bohr`, and energy unit is `Hartree`.</p><p><strong>## small molecules torsion</strong></p><p>Contains 62 small molecules with 36 conformation for each under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] B. D. Sellers, N. C. James, A. Gobbi, A comparison of quantum and molecular mechanical methods to estimate strain energy in druglike fragments, Journal of chemical information and modeling 57 (6) (2017) 1265–127</p><p><br>&nbsp;</p><p><strong>## MPCONF91</strong></p><p>Contains 6 molecules with 91 conformations under LNO-CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] J. Rezac, D. Bím, O. Gutten, L. Rulisek, Toward accurate conformational energies of smaller peptides and medium-sized macrocycles: Mpconf196 benchmark energy data set, Journal of chemical theory and computation 14 (3) (2018) 1254–1</p><p><br>&nbsp;</p><p><strong>## torsionNet206</strong></p><p>Contains 206 molecules with 4494 conformations under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] B. K. Rai, V. Sresht, Q. Yang, R. Unwalla, M. Tu, A. M. Mathiowetz,G. A. Bakken, Torsionnet: A deep neural network to rapidly predict small-molecule torsional energy profiles with the accuracy of quantum mechanics, Journal of Chemical Information and Modeling 62 (4) (2022) 785–80</p><p><br>&nbsp;</p><p><strong>## Out-of-plane bending</strong></p><p>Contains 242 molecules with 3315 conformations under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] X. Yang, C. Liu, P. Ren, High order ab initio valence force field with chemical pattern based parameter assignment., Journal of Computational Biophysics and Chemistry 21 (4) (2021) 43</p><p><br><br>&nbsp;</p><p><strong>## DrugBank-T</strong></p><p>Contains 165 molecules with 1155 conformations under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] V. Law, C. Knox, Y. Djoumbou, T. Jewison, A. C. Guo, Y. Liu, A. Maciejewski, D. Arndt, M. Wilson, V. Neveu, et al., Drugbank</p><p>4.0: shedding new light on drug metabolism, Nucleic acids research 42 (D1) (2014) D1091–D1097 &nbsp;</p><p>[2] Z. Qiao, M. Welborn, A. Anandkumar, F. R. Manby, T. F. Miller III, Orbnet: Deep learning for quantum chemistry using symmetry adapted atomic-orbital features, The Journal of chemical physics 153 (12) (2020) 124111</p><p><strong>## Notice</strong></p><p>if you use above datasets, please cite the original articals too</p>

opencc-byAug 2023View details →
zenodo44/100

Data for: Machine-learning-accelerated simulations enable heuristic-free surface reconstruction

<p>This is the dataset for the publication "Machine-learning-accelerated simulations to enable automatic surface reconstruction", by X. Du, J.K. Damewood, J.R. Lunger, R. Millan, B.&nbsp;Yildiz, L. Li, and R. Gómez-Bombarelli. The repository contains the density-functional theory (DFT) data used to train the neural network force fields (NFF), selected results from our GaN(0001), Si(111), and SrTiO3(001) Virtual Surface Site Relaxation-Monte Carlo (VSSR-MC) runs, and Jupyter notebooks used for analysis and plots. To run the .ipynb's, you will need to install <a href="https://github.com/learningmatter-mit/surface-sampling">surface-sampling</a> (tested up to commit 02820d339eed6291b6af6ccb809f154ad6244110 on master) and <a href="https://github.com/learningmatter-mit/NeuralForceField">NeuralForceField</a>&nbsp;(tested up to commit 72d1f32f43f202c1a466116beeed15845a6456e7 on master) from the <a href="https://github.com/learningmatter-mit">Rafael Gómez-Bombarelli Group @ MIT</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

GLAB-VOD: Global L-band AI-Based Vegetation Optical Depth Dataset Based on Machine Learning and Remote Sensing

<p>GLAB VOD is a Global L-band Ai-Based vegetation optical depth dataset with 18-day temporal and 25 km spatial resolution, covering 2002 to 2020. The dataset is created using a neural network with SMOS-SMAP-INRAE-BORDEAUX (SMOSMAP-IB) VOD product as a target (over 2015-2020) and brightness temperatures (TB) from the SMOS, AMSR-E, and AMSR-2 spaceborne missions alongside with a novel soil moisture dataset (CASM) as inputs. The GLAB-VOD dataset was created using a recently developed methodology previously used to create a long-term consistent soil moisture dataset CASM, adapted to the&nbsp; VOD retrievals. First, the TB and VOD signals were divided into fixed seasonal cycle and residuals, where the residual part of the signal contains sub-seasonal periodic signals, trends, extremes, and noise. Then, a multi-staged neural network training scheme was used to achieve internally consistent predictions by merging data from different sources without introducing biases or compromising data distribution. A side-product of this project is GLAB TB - a global long-term brightness temperature dataset that matches SMOS TB quality and spawns back to 2002.&nbsp;GLAB TB has daily temporal resolution and 25 km spatial resolution.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

CPAZMAL: Cryosphere PAZ satellite MAchine Learning

<p>CPAZMAL:<strong> C</strong>ryosphere <strong>PAZ</strong> satellite <strong>MA</strong>chine <strong>L</strong>earning</p> <p>The aim of this dataset is to serve as a foundation for machine learning in multi-class classification, specifically in mountainous regions. It comprises descending images acquired by the PAZ X-band satellite, focusing on the Mont Blanc region during the period from January 2020 to November 2021, totaling 56 acquisitions.</p> <p>The time series is divided into two sub-sections:</p> <ol> <li>From January 2020 to 8th January 2021 included: dual polarisation HH and HV,</li> <li>After 8th January 2021: single polarisation HH.</li> </ol> <div> <div>The datas are divided into 8 classes:</div> <div> <ul> <li>Hanging Glacier (HAG)</li> <li>Ice Aperon (ICA)</li> <li>Ablation area</li> <li>Accumulation area</li> <li>Rock</li> <li>Plain</li> <li>Forest</li> <li>City</li> </ul> <p>In each classe, between 4 to 10 groups or distinct areas, where their complete description (position, aspect, elevation, ...) can be found in the&nbsp;<em>desc_topo_areas.png&nbsp;</em>file</p> </div> <div>We provide code that directly extracts temporal or spatial datasets, consisting of homogeneous windows paired with respective labels.</div> <div> <pre><code># Request and save data into hdf5 file rqtemp = "classe in ['ICA','HAG','ABL','ACC','FOR','CIT','ROC','PLA'] &amp; date &lt; '2021-01-01'" cdlf = Dataset_tiff2hdf5 ( path_to_folder_extracted, different_group=True, n_jobs=1, outpath="path_to_dataset.h5", extension="temporal" ) cdlf.extract_data(rqtemp, polarisation="HH", winsize=7, save=True) # Load the previously extracted data set ( x, y, gr, org, _, ) = load_h5(path_to_dataset.h5)</code></pre> </div> <div>An example of how to use it can be found at <a href="https://github.com/Matthieu-Gallet/PAZ_DTW_classification" target="_blank" rel="noopener">Github</a>.</div> <div>&nbsp;</div> <blockquote> <div>The authors would like to thank the <em>Spanish Instituto Nacional de Tecnica Aerospacial</em> (INTA) for the PAZ images (Project AO-001-051)&nbsp;</div> </blockquote> </div>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record