Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,185

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7,185 results for “learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset for: "Incremental Semiparametric Inverse Dynamics Learning"

<p>Dataset used in the experimental section of the paper:</p> <blockquote> <p>R. Camoriano, S. Traversaro, L. Rosasco, G. Metta and F. Nori, "<strong>Incremental semiparametric inverse dynamics learning,</strong>" <em>2016 IEEE International Conference on Robotics and Automation (ICRA)</em>, Stockholm, 2016, pp. 544-550.<br> <br> doi: 10.1109/ICRA.2016.7487177<br> <br> Abstract: This paper presents a novel approach for incremental semiparametric inverse dynamics learning. In particular, we consider the mixture of two approaches: Parametric modeling based on rigid body dynamics equations and nonparametric modeling based on incremental kernel methods, with no prior information on the mechanical properties of the system. The result is an incremental semiparametric approach, leveraging the advantages of both the parametric and nonparametric models. We validate the proposed technique learning the dynamics of one arm of the iCub humanoid robot.<br> <br> URL: http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&amp;arnumber=7487177&amp;isnumber=7487087<br>  </p> </blockquote> <p> </p> <p><strong>Description</strong></p> <p>The file "iCubDyn_2.0.mat" contains data collected from the right arm of the iCub humanoid robot, considering as input the positions, velocities and accelerations of the 3 shoulder joints and of the elbow joint, and as outputs the 3 force and 3 torque components measured by the six-axis F/T sensor in-built in the upper arm.</p> <p>The dataset is collected at 10Hz at as the end-effector tracks circumferences with 10cm radius on the transverse (XY) and sagittal (XZ) planes (For more information on the iCub reference frames, see [4]) at approximately 0.6 m/s. The total number of points for each dataset is 10000, corresponding to approximately 17 minutes of continuous operation. Trajectories are generated by means of the Cartesian Controller presented in [5].</p> <p>    Input (X)</p> <p>        columns 1-4: Joint (3 shoulder joints + 1 elbow joint) positions<br>         columns 5-8: Joint (3 shoulder joints + 1 elbow joint) velocities<br>         columns 9-12: Joint (3 shoulder joints + 1 elbow joint) accelerations</p> <p>    Output (Y)</p> <p>        Columns 1-3: Measured forces (N) along the X, Y, Z axes by the force-torque (F/T) sensor placed in the upper arm<br>         Columns 4-6: Measured torques (N*m) along the X, Y, Z axes by the force-torque (F/T) sensor placed in the upper arm</p> <p> </p> <p><strong>Preprocessing</strong><br> <br> - Velocities and accelerations are computed by an Adaptive Window Polynomial Fitting Estimator, implemented through a least-squares based algorithm on a adpative window (see [2], [3]). Velocity estimation max window size: 16. Acceleration estimation max window size: 25.<br> - Positions, velocities and accelerations are recorded at 9Hz and oversampled to 20 Hz via cubic spline interpolation.<br> - Forces and torques are directly recorded at 20Hz.</p> <p>This dataset was used in [1] for experimental purposes. See section IV therein for further details.</p> <p>For more information, please contact:<br> Raffaello Camoriano - raffaello.camoriano@iit.it<br> Silvio Traversaro - silvio.traversaro@iit.it</p> <p> </p> <p><strong>References</strong><br> <br> [1] Camoriano, Raffaello; Traversaro, Silvio; Rosasco, Lorenzo; Metta, Giorgio; Nori, Francesco, "Incremental Semiparametric Inverse Dynamics Learning", eprint arXiv:1601.04549, 01/2016<br> [2] F. Janabi-Sharifi ; Dept. of Mech. Eng., Ryerson Polytech. Univ., Toronto, Ont., Canada ; V. Hayward ; C. -S. J. Chen, "Discrete-time adaptive windowing for velocity estimation",  IEEE Transactions on Control Systems Technology, 1003 - 1009, Vol. 8,  Issue 6, Nov 2000<br> [3] https://github.com/robotology/icub-main/blob/master/src/libraries/ctrlLib/include/iCub/ctrl/adaptWinPolyEstimator.h<br> [4] http://wiki.icub.org/wiki/ICubForwardKinematics<br> [5] U. Pattacini; F. Nori; L. Natale; G. Metta; and G. Sandini; “An experimental evaluation of a novel minimum-jerk cartesian controller for humanoid robots,” in Intelligent Robots and Systems (IROS), 2010 IEEE/RSJ International Conference on, Oct 2010, pp. 1668–1674.</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

Final Results from the RDM Survey - LEARN project (June 2017)

<p> </p> <p>Data obtained from the open survey developed by the LEARN project (http://www.learn-rdm.eu/) as a self-assessment tool to assist institutions discover how ready they are for managing research data. This dataset replaces the previous ones published at http://doi.org/10.5281/zenodo.61903 and http://doi.org/10.5281/zenodo.290635. The survey is based on the issues posed to institutions by the LERU Roadmap for Research Data published at the end of 2013, and available at: http://www.learn-rdm.eu/material/leru_roadmap_for_research_data<br> The survey has thirteen questions addressing the main elements to be taken into account in developing an institutional strategy for research data management. Each question has three possible answers representing green, yellow or red light. The more ‘green light’ responses recorded, the readier an institution probably is for managing its research data.</p> <p>The survey is available in English at http://learn-rdm.eu/en/rdm-readiness-survey/ and in Spanish at http://learn-rdm.eu/encuesta-rdm/</p>

opencc-by-4.0Jun 2017View details →
zenodo44/100

Deep learning based automatic grounding line delineation in DInSAR interferograms

<p>This dataset contains a small subset of the AIS_cci GLL product, which covers several key glaciers and the corresponding HED-delineated grounding lines generated from our automatic delineation pipeline. A description of the attributes of the AIS_cci GLL product is provided in the&nbsp;<a href="https://climate.esa.int/media/documents/ST-UL-ESA-AISCCI-PUG-0001.pdf" target="_blank" rel="noopener">Product User Guide</a>.&nbsp; We do not indicate the split of the interferograms into training, validation and test sets as the complete AIS_cci dataset is not open-access.</p> <p>We also provide eight double difference interferograms at 100 m pixel size to demonstrate the generation of the features stack. Please note, eight samples are not sufficient to train the neural network to achieve the delineation capability described in our work.</p> <p>The "UUID" attribute in both GeoJSON files is an identifier that links the vector geometries to the interferogram TiFF files.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

TCOM-H2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric H2O profile dataset [1991-2021] constructed using machine-learning.

<p>Methodology: &nbsp;</p> <p>The <strong>TOMCAT simulation</strong> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized <strong>ERA-5 reanalysis data</strong>.</p> <h3>H2O Profile Processing and Bias Correction</h3> <p><strong>Collocated H2O profiles</strong> are organized into five distinct latitude bins:</p> <ul> <li> <p><strong>NH polar</strong>: 90∘N - 50∘N</p> </li> <li> <p><strong>NH mid-lat</strong>: 20∘N - 70∘N</p> </li> <li> <p><strong>Tropics</strong>: 40∘S - 40∘N</p> </li> <li> <p><strong>SH mid-lat</strong>: 70∘S - 20∘S</p> </li> <li> <p><strong>SH polar</strong>: 90∘S - 50∘S</p> </li> </ul> <p>Initially, <strong>differences between TOMCAT and satellite measurements</strong> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from 10,km to 60,km). Note that TOMCAT may not accurately capture H2O evolution post-HTHH eruption due to the sparse spatial coverage of ACE measurements, which limits training data.</p> <p><strong>Separate XGBoost regression models</strong> are then trained for these H2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate <strong>H2O bias corrections</strong> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</p> <p><strong>Height-resolved H2O profile data</strong> are then interpolated onto 28 standard pressure levels (from 300,hPa to 0.1,hPa), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</p> <p>We acknowledge the inherent <strong>dry biases in the original TOMCAT H2O profiles</strong>, largely because the TTL entry mixing ratios are based on a simplistic sinusoidal seasonal cycle, which omits the H2O enhancement contributed by tropical convective clouds.</p> <h3>Data Files</h3> <p>The dataset includes two files containing daily mean zonal mean H2O profiles:</p> <ul> <li> <p><code>zmh2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>height level data</strong> (10,km to 60,km).</p> </li> <li> <p><code>zmh2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>pressure level data</strong> (300,hPa to 0.1,hPa).</p> </li> </ul> <h3>Reference Publication</h3> <p>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</p> <p>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, <a title="null" href="https://doi.org/10.5194/essd-15-5105-2023">https://doi.org/10.5194/essd-15-5105-2023</a>, 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

TCOM-O3: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric ozone profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;TOMCAT simulation is performed at T64L32 resolution for the 2000-2024 time period. Collocated Ozone (O3) profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement &nbsp;differences are calculated for each zonal bins (51 height levels, 10km to 60km). Note that if enough ACE measurements are not avaliable for a particular level then data is purely based on TOMCAT simulated output field. Separate XGBoost regression models are trained for the &nbsp;differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids. &nbsp;TOMCAT output sampled at 1.30 pm local time at the equator. Estimated corrections for a given model grid that are added to the original TOMCAT simulated day and night time ozone profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values. &nbsp;For more details see attached presentation. Previous version use both HALOE and ACE data. Here only ACE data is used (hence starting date is 01 January 2000). PDF file shows comparison between v1.0 and v1.1 as well as TOMCAT data.</p> <p>Dataset also includes two files containing daily mean zonal mean hydrogen fluoride &nbsp;profiles on height (10-60 km) and pressure (300-0.1 hPa) levels:</p> <p>zmo3_TCOM_hlev_T2Dz_2000_2024.nc &ndash; height level data (10 to 60 km)</p> <p>zmo3_TCOM_plev_T2Dz_2000_2024.nc &ndash; pressure level data (300 to 0.1 hPa)</p> <p>Daily 3D profiles on height and pressure levels would be made available on request.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Automated MESSENGER Plasma Region Classifications via Unsupervised Transfer Learning

<p>This file contains the 1-minute resolution dataset (&ldquo;labeled_sunside_data_3labels.csv&rdquo;) for Toy-Edens et al.&rsquo;s Automated Classification of MESSENGER Plasma Observations via Unsupervised Transfer Learning. The 1-minute resolution file contains the rolled up 1-minute epoch, features that go into clustering and post-cleaning methods, spacecraft positions (in MSO), total magnetic field, raw and cleaned clustering labels, and raw and cleaned transition name.</p> <p>We ask that if you use any parts of the dataset that you cite Toy-Edens et al.&rsquo;s Automated Classification of MESSENGER Plasma Observations via Unsupervised Transfer Learning (DOI: 10.3389/fspas.2025.1608091).</p> <p>This work was supported by NASA grants 80NSSC19K0789 and 80NSSC22K0993.</p> <p>&nbsp;</p> <p>The following tables detail the contents of the described files:</p> <p><strong>labeled_sunside_data_3labels.csv description</strong></p> <table style="width: 100.063%; height: 851.2px;"> <tbody> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p><strong>Column Name</strong></p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p><strong>Description</strong></p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;Epoch</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Epoch in datetime (YYYY-MM-DD HH:MM:SS)</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;x_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>x position of the spacecraft in MSO [km]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;y_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>y position of the spacecraft in MSO [km]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;z_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>z position of the spacecraft in MSO [km]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;btot_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Total magnetic field [nT]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;norm_Btot</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Magnitude of the total magnetic field normalized to 150nT. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;ratio_max_width</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Ratio of the width of the most prominent ion spectra peak (in number of energy channels) to max number of energy channels. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;ratio_high_low</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Ratio of the mean of the log intensity of high energies in the ion spectra to the mean of the log intensity of low energies in the ion spectra. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;high_intensity</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Boolean if there is a peak with a higher minimum intensity threshold. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;spectra_counts</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>A ratio of spectra bins with non-zero counts to all possible spectra bins (i.e. way to determine if too much missing spectra data). See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;raw_named_label</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Raw cluster assigned plasma region label (allowed values: magnetosheath, magnetosphere, solar wind)</p> </td> </tr> <tr> <td style="width: 17.3792%;"> <p>intermediate_named_label</p> </td> <td style="width: 78.9512%;"> <p>Cleaned cluster assigned plasma region label with only relabeling rules applied. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;named_label</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Cleaned cluster assigned plasma region label with relabeling rules and post-processing applied (use these unless have a specific reason to use raw labels). See paper for more information</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p>&nbsp;raw_transition_name</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Raw transition names (e.g. bow shock, magnetopause) based on "raw_named_label" cluster labels. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p>&nbsp;transition_name</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Cleaned transition names (e.g. bow shock, magnetopause) after removing likely transient transitions based on "named_label" cluster labels. See paper for more information</p> </td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

CLRD-GLPS: A Long-term Seasonal Dataset of Ruminant Livestock Distribution in China's Grazing Production Systems (2000-2021) Using Stacking-based Interpretable Machine Learning

<p>Advanced computational methods integrating ensemble learning with interpretable machine learning are essential for precision livestock management under increasing environmental constraints and food security pressures. This study develops a novel stacking-based interpretable machine learning (IML) framework that combines multiple algorithms with SHAP analysis techniques to generate the China's Long-term Ruminant Livestock Distribution in Grazing Livestock Production Systems (CLRD-GLPS) dataset. Our computational approach addresses critical challenges in livestock distribution modelling: livestock segmentation and spatial prediction accuracy. The framework integrates Random Forest, XGBoost, CatBoost, LightGBM, and Extra Trees through a two-layer stacking architecture, enhanced with SHAP (Shapley Additive Explanations) analysis for model interpretability. We also implemented interpretable machine learning for livestock production system segmentation to distinguish grazing from total livestock populations. The stacking ensemble demonstrated superior performance over individual algorithms, achieving R&sup2; values of 0.954-0.961 for cattle and 0.896-0.901 for sheep and goats, with improvements of up to 8.3% compared to best performance single-model approaches. Multi-scale validation confirmed computational robustness: livestock segmentation achieved R&sup2; = 0.80 at county level, while independent city-level validation of CLRD-GLPS datasets yielded R&sup2; = 0.76-0.80. SHAP interpretability analysis revealed distinct environmental drivers, with vegetation indices and topography primarily influencing cattle distribution, while snow conditions and elevation dominated sheep and goat patterns. This computational framework advances livestock distribution modelling through enhanced prediction accuracy, model stability, and interpretability, while the CLRD-GLPS dataset provides essential spatial-temporal information for rangeland sustainability assessments and evidence-based livestock management policies. This dataset is supported by the Second Tibetan Plateau Scientific Expedition and Research Program (STEP, grant no. 2019QZKK0906).</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

GTWS-MLrec: Global terrestrial water storage reconstruction by machine learning from 1940 to present

<p>Terrestrial water storage (TWS) includes all forms of water stored on and below the land surface, and is a key determinant of global water and energy budgets. However, TWS data from measurements by the Gravity Recovery and Climate Experiment (GRACE) satellite mission are only available from 2002, limiting global and regional investigation of the long-term trends and variabilities in the terrestrial water cycle under climate change. This study presents long-term (i.e., 1940-2022) and high-resolution (i.e., 0.25°) monthly time series of TWS anomalies over the global land surface. The reconstruction is achieved by using a set of machine learning models with a large number of predictors, including climatic and hydrological variables, land use/land cover data, and vegetation indicators (e.g., leaf area index). The outcome, machine learning-reconstructed TWS estimates (i.e., GTWS-MLrec), fits well with the GRACE/GRACE-FO measurements, showing high correlation coefficients and low biases in the GRACE era. We also evaluate GTWS-MLrec with other independent datasets such as the land-ocean mass budget, large-scale water balance in 341 large river basins, and streamflow measurements at 10,168 gauges. We find that the proposed approach performs overall as well as or is more reliable than previous TWS datasets. Moreover, our reconstructions successfully reproduce the impact of climate variability, such as strong El Niño events. GTWS-MLrec dataset consists of three reconstructions based on JPL, CSR and GSFC mascons, three detrended and de-seasonalized reconstructions, and six global average TWS series over land areas, both with and without Greenland and Antarctica. Along with its extensive attributes, GTWS_MLrec can support a broad range of applications such as better understanding the global water budget, constraining and evaluating hydrological models, climate-carbon coupling, and water resources management.</p><p>Please cite the reference: <strong>Yin J, Slater L, Khouakhi A, et al. GTWS-MLrec: Global terrestrial water storage reconstruction by machine learning from 1940 to present. Earth System Science Data. 2023.</strong></p><p>For any inquiry about the dataset, welcome to contact Dr. Jiabo Yin (jboyn@whu.edu.cn).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

MLFMF: Data Sets for Machine Learning for Mathematical Formalization

<h3>MLFMF</h3><p><strong>MLFMF (Machine Learning for Mathematical Formalization) </strong>is a collection of data sets for benchmarking recommendation systems used to support formalization of mathematics with proof assistants. These systems help humans identify which previous entries (theorems, constructions, datatypes, and postulates) are relevant in proving a new theorem or carrying out a new construction.&nbsp;</p><p>The MLFMF data sets provide solid benchmarking support for further investigation of the numerous machine learning approaches to formalized mathematics. With more than 250,000 entries in total, this is currently the largest collection of formalized mathematical knowledge in machine learnable format.&nbsp;</p><p>In addition to benchmarking the recommendation systems, the data sets can also be used for benchmarking <strong>node classification</strong> and <strong>link prediction</strong> algorithms.&nbsp;</p><h3>The four data sets</h3><p>Each data set is derived from a library of formalized mathematics written in proof assistants <a href="https://agda.readthedocs.io/en/v2.6.4/"><i>Agda</i></a> or <a href="https://lean-lang.org/"><i>Lean</i></a>. The collection includes &nbsp;</p><ol><li>the largest Lean 4 library <a href="https://github.com/leanprover-community/mathlib4"><strong>Mathlib</strong></a>,</li><li>the three largest Agda libraries:<ul><li>the <a href="https://github.com/agda/agda-stdlib"><strong>standard library</strong></a></li><li>the library of univalent mathematics <a href="https://github.com/UniMath/agda-unimath"><strong>Agda-unimath</strong></a>, and</li><li>the <a href="https://github.com/martinescardo/TypeTopology"><strong>TypeTopology</strong></a> library.</li></ul></li></ol><p>Each data set represents the corresponding library in two ways: as a heterogeneous network, and as a list of syntax trees of all the entries in the library. The network contains the (modular) structure of the library and the references between entries, while the syntax trees give complete and easily parsed information about each entry.</p><p>The Lean library data set was obtained by converting <strong>.olean</strong> files into s-expressions (see the <a href="https://github.com/andrejbauer/lean2sexp"><strong>lean2sexp</strong></a> tool).</p><p>The Agda data sets were obtained with an <a href="https://github.com/andrejbauer/agda/tree/master-sexp">s-expression extension</a> of the official Agda repository (use either master-sexp or release-2.6.3-sexp branch).</p><p>For more details, see our <a href="https://arxiv.org/abs/2310.16005"><strong>arXiv copy</strong></a><strong> </strong>of the paper.</p><h3>Directory structure</h3><p>First, the <strong>mlfmf.zip</strong> archive needs to be unzipped. It contains a separate directory for every library (for example, the standard library of Agda can be found in the stdlib directory) and some auxiliary files. Every library directory contains</p><ul><li>the <strong>network file</strong> from which the heterogeneous network can be loaded,</li><li>a zip of the <strong>entries directory</strong> that contains (many) files with abstract syntax trees. Each of those files describes a single entry of the library.</li></ul><p>In addition to the auxiliary files which are used for loading the data (and described below), the zipped sources of lean2sexp and Agda s-expression extension are present.</p><h4>Loading the data</h4><p>In addition to the data files, there is also a simple python script <strong>main.py</strong> for loading the data. To run it, you will have to install the packages listed in the file <strong>requirements.txt</strong>: <strong>tqdm</strong> and <strong>networkx</strong>. The easiest way to do so is calling <i><strong>pip install -r requirements.txt</strong></i>.</p><p>When running <strong>main.py </strong>for the first time, the script will unzip the entry files into the directory named <strong>entries</strong>. After that, the script loads the syntax trees of the entries (see the <strong>Entry</strong> class) and the network (as <i>networkx.MultiDiGraph</i> object).</p><p><i>Note. The entry files have extension <strong>.dag </strong>(directed acyclic graph), since Lean uses node sharing, which breaks the tree structure (a shared node has more than one parent node).</i></p><h3>More information</h3><p>For more information about the <strong>data collection process</strong>, <strong>detailed data (and data format) description</strong>, and <strong>baseline experiments</strong> that were already performed with these data, see our <a href="https://arxiv.org/abs/2310.16005"><strong>arXiv copy</strong></a><strong> of the paper</strong>.</p><p>For the code that was used to perform the experiments and data format description, visit our github repository <a href="https://github.com/ul-fmf/mlfmf-data"><strong>https://github.com/ul-fmf/mlfmf-data.</strong></a></p><h3>Funding</h3><p>Since not all the funders are available in the Zenodo's database, we list them here:</p><ol><li>This material is based upon work supported by the Air Force Office of Scientific Research under award number FA9550-21-1-0024.</li><li>The authors also acknowledge the financial support of the Slovenian Research Agency via the research core funding No. P2-0103 and No. P1-0294.</li></ol><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

A machine learning-based high-precision density functional method for drug-like molecules

<h2><strong>Models</strong></h2><p>The repo contains the models and test datasets for our aticles. The energy unit is in <strong>Hartree,</strong> The coordinate unit is in<strong> Bohr.</strong></p><p><strong>## DeePHF</strong></p><p>you need first prepare the `dm_eig.npy` in data_test and do predict `l_e_delta.npy`, you can use</p><p>```</p><p>deepks test -m model.pth -o test/test -d data_test/* -D dm_eig -G</p><p>```</p><p><strong>## DeePKS</strong></p><p>first you should prepare the `atom.npy`, and `energy.npy` in data_test. you can test the datasets by command.&nbsp;</p><p>```</p><p>deepks scf scf_input.yaml -m model.pth -s data_test -d test_out</p><p>```</p><p><strong># Datasets</strong></p><p>All datasets only have `atom.npy` and `energy.npy`. The coordinate unit is `bohr`, and energy unit is `Hartree`.</p><p><strong>## small molecules torsion</strong></p><p>Contains 62 small molecules with 36 conformation for each under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] B. D. Sellers, N. C. James, A. Gobbi, A comparison of quantum and molecular mechanical methods to estimate strain energy in druglike fragments, Journal of chemical information and modeling 57 (6) (2017) 1265–127</p><p><br>&nbsp;</p><p><strong>## MPCONF91</strong></p><p>Contains 6 molecules with 91 conformations under LNO-CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] J. Rezac, D. Bím, O. Gutten, L. Rulisek, Toward accurate conformational energies of smaller peptides and medium-sized macrocycles: Mpconf196 benchmark energy data set, Journal of chemical theory and computation 14 (3) (2018) 1254–1</p><p><br>&nbsp;</p><p><strong>## torsionNet206</strong></p><p>Contains 206 molecules with 4494 conformations under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] B. K. Rai, V. Sresht, Q. Yang, R. Unwalla, M. Tu, A. M. Mathiowetz,G. A. Bakken, Torsionnet: A deep neural network to rapidly predict small-molecule torsional energy profiles with the accuracy of quantum mechanics, Journal of Chemical Information and Modeling 62 (4) (2022) 785–80</p><p><br>&nbsp;</p><p><strong>## Out-of-plane bending</strong></p><p>Contains 242 molecules with 3315 conformations under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] X. Yang, C. Liu, P. Ren, High order ab initio valence force field with chemical pattern based parameter assignment., Journal of Computational Biophysics and Chemistry 21 (4) (2021) 43</p><p><br><br>&nbsp;</p><p><strong>## DrugBank-T</strong></p><p>Contains 165 molecules with 1155 conformations under CCSD(T)/def2-TZVP.</p><p><br>&nbsp;</p><p>[1] V. Law, C. Knox, Y. Djoumbou, T. Jewison, A. C. Guo, Y. Liu, A. Maciejewski, D. Arndt, M. Wilson, V. Neveu, et al., Drugbank</p><p>4.0: shedding new light on drug metabolism, Nucleic acids research 42 (D1) (2014) D1091–D1097 &nbsp;</p><p>[2] Z. Qiao, M. Welborn, A. Anandkumar, F. R. Manby, T. F. Miller III, Orbnet: Deep learning for quantum chemistry using symmetry adapted atomic-orbital features, The Journal of chemical physics 153 (12) (2020) 124111</p><p><strong>## Notice</strong></p><p>if you use above datasets, please cite the original articals too</p>

opencc-byAug 2023View details →
zenodo44/100

Developing Learning Paths - The Learning Paths Protocol

<p>Learning Paths (LPs) are pathways that guide learners through a set of learning courses or materials to be undertaken progressively to acquire the desired knowledge and skills on a subject of interest. &nbsp;In this video, we explain how to develop learning paths step-by-step using the Learning Paths protocol. &nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Deep learning to extract the meteorological by-catch of wildlife cameras: Supporting data, models and code

<p>This repository contains the data, models and code to train and deploy deep learning models related to the paper "Deep learning to extract the meteorological by-catch of wildlife cameras" published in the journal Global Change Biology (<a href="https://doi.org/10.1111/gcb.17078"><strong>https://doi.org/10.1111/gcb.17078</strong></a>).</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Deep learning segmentation projects of FIB-SEM dataset of U2-OS cell

<p>This submission includes ground truth datasets that were used to segment the nuclear envelope (NE), mitochondria, endoplasmic reticulum (ER) and Golgi from a human bone osteosarcoma epithelial cell (U2-OS) imaged using focused-ion beam scanning electron microscopy (FIB-SEM).</p><p>The full FIB-SEM dataset is deposited to EMPIAR (<a href="https://www.ebi.ac.uk/empiar">https://www.ebi.ac.uk/empiar</a>, EMPIAR-11746).&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Data for: Machine-learning-accelerated simulations enable heuristic-free surface reconstruction

<p>This is the dataset for the publication "Machine-learning-accelerated simulations to enable automatic surface reconstruction", by X. Du, J.K. Damewood, J.R. Lunger, R. Millan, B.&nbsp;Yildiz, L. Li, and R. Gómez-Bombarelli. The repository contains the density-functional theory (DFT) data used to train the neural network force fields (NFF), selected results from our GaN(0001), Si(111), and SrTiO3(001) Virtual Surface Site Relaxation-Monte Carlo (VSSR-MC) runs, and Jupyter notebooks used for analysis and plots. To run the .ipynb's, you will need to install <a href="https://github.com/learningmatter-mit/surface-sampling">surface-sampling</a> (tested up to commit 02820d339eed6291b6af6ccb809f154ad6244110 on master) and <a href="https://github.com/learningmatter-mit/NeuralForceField">NeuralForceField</a>&nbsp;(tested up to commit 72d1f32f43f202c1a466116beeed15845a6456e7 on master) from the <a href="https://github.com/learningmatter-mit">Rafael Gómez-Bombarelli Group @ MIT</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

GLAB-VOD: Global L-band AI-Based Vegetation Optical Depth Dataset Based on Machine Learning and Remote Sensing

<p>GLAB VOD is a Global L-band Ai-Based vegetation optical depth dataset with 18-day temporal and 25 km spatial resolution, covering 2002 to 2020. The dataset is created using a neural network with SMOS-SMAP-INRAE-BORDEAUX (SMOSMAP-IB) VOD product as a target (over 2015-2020) and brightness temperatures (TB) from the SMOS, AMSR-E, and AMSR-2 spaceborne missions alongside with a novel soil moisture dataset (CASM) as inputs. The GLAB-VOD dataset was created using a recently developed methodology previously used to create a long-term consistent soil moisture dataset CASM, adapted to the&nbsp; VOD retrievals. First, the TB and VOD signals were divided into fixed seasonal cycle and residuals, where the residual part of the signal contains sub-seasonal periodic signals, trends, extremes, and noise. Then, a multi-staged neural network training scheme was used to achieve internally consistent predictions by merging data from different sources without introducing biases or compromising data distribution. A side-product of this project is GLAB TB - a global long-term brightness temperature dataset that matches SMOS TB quality and spawns back to 2002.&nbsp;GLAB TB has daily temporal resolution and 25 km spatial resolution.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

University Learning Management System xAPI data for the ILEDA project

<p>Anonymized data in xAPI format collected from learning management systems (Moodle and LAMS) for the ILEDA (2021-1-BG01-KA220-HED-000031121) project https://ileda.eu/ileda-project. The dataset contains 306,741 records of 829 individuals participating in eight different blended learning courses following either a flipped classroom methodology or a project-based learning methodology from four universities: University of Eastern Finland (Finland), University of León (Spain), Belgrade Metropolitan University (Serbia) and Sofia University (Bulgaria).</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo44/100

CPAZMAL: Cryosphere PAZ satellite MAchine Learning

<p>CPAZMAL:<strong> C</strong>ryosphere <strong>PAZ</strong> satellite <strong>MA</strong>chine <strong>L</strong>earning</p> <p>The aim of this dataset is to serve as a foundation for machine learning in multi-class classification, specifically in mountainous regions. It comprises descending images acquired by the PAZ X-band satellite, focusing on the Mont Blanc region during the period from January 2020 to November 2021, totaling 56 acquisitions.</p> <p>The time series is divided into two sub-sections:</p> <ol> <li>From January 2020 to 8th January 2021 included: dual polarisation HH and HV,</li> <li>After 8th January 2021: single polarisation HH.</li> </ol> <div> <div>The datas are divided into 8 classes:</div> <div> <ul> <li>Hanging Glacier (HAG)</li> <li>Ice Aperon (ICA)</li> <li>Ablation area</li> <li>Accumulation area</li> <li>Rock</li> <li>Plain</li> <li>Forest</li> <li>City</li> </ul> <p>In each classe, between 4 to 10 groups or distinct areas, where their complete description (position, aspect, elevation, ...) can be found in the&nbsp;<em>desc_topo_areas.png&nbsp;</em>file</p> </div> <div>We provide code that directly extracts temporal or spatial datasets, consisting of homogeneous windows paired with respective labels.</div> <div> <pre><code># Request and save data into hdf5 file rqtemp = "classe in ['ICA','HAG','ABL','ACC','FOR','CIT','ROC','PLA'] &amp; date &lt; '2021-01-01'" cdlf = Dataset_tiff2hdf5 ( path_to_folder_extracted, different_group=True, n_jobs=1, outpath="path_to_dataset.h5", extension="temporal" ) cdlf.extract_data(rqtemp, polarisation="HH", winsize=7, save=True) # Load the previously extracted data set ( x, y, gr, org, _, ) = load_h5(path_to_dataset.h5)</code></pre> </div> <div>An example of how to use it can be found at <a href="https://github.com/Matthieu-Gallet/PAZ_DTW_classification" target="_blank" rel="noopener">Github</a>.</div> <div>&nbsp;</div> <blockquote> <div>The authors would like to thank the <em>Spanish Instituto Nacional de Tecnica Aerospacial</em> (INTA) for the PAZ images (Project AO-001-051)&nbsp;</div> </blockquote> </div>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Dataset for Machine Learning Assisted Citation Screening for Systematic Reviews

<p>The work "Machine Learning Assisted Citation Screening for Systematic Reviews" explored the problem of citation screening automation using machine-learning (ML) with an aim to accelerate the process of generating <a href="https://en.wikipedia.org/wiki/Systematic_review#:~:text=Systematic%20reviews%20are%20a%20type,synthesize%20findings%20qualitatively%20or%20quantitatively." rel="nofollow">systematic reviews</a>. Manual process of citation screening involve two reviewers manually screening the searched studies using a predefined inclusion criteria. If the study passes the "inclusion" criteria, it is included for further analysis or is excluded. As apparant through manual screening process, the work considered citation screening as a binary classification problem whereby any ML classifier could be trained to separate the searched studies into these two classes (include&nbsp;and&nbsp;exclude).</p> <p>&nbsp;</p> <p>A physiotherapy citation screening dataset was used to test automation approaches and the dataset includes the studies identified for citation screening in an update to the systematic review by Hilfiker <em>et al.</em> The dataset included titles and abstracts (citations) from 31,279 (deduplicated: 25,540) studies identified during the search phase of this SR. These studies were already manually assessed for relevance and labelled by two reviewers into two mutually exclusive labels. The uploaded file consists of 25,540 data samples, with each data sample separated by a new line. It is a tab separated file and the data in it is structured as shown below. This dataset was manually labelled into include and exclude by Hilfiker&nbsp;<em>et al.</em></p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Title</strong></td> <td><strong>PMID</strong></td> <td><strong>Abstract&nbsp;</strong></td> <td><strong>Class</strong></td> <td><strong>MeSH terms (separated by a pipe)</strong></td> </tr> <tr> <td>Structured exercise improves physical functioning in women with stages I and II breast cancer: results of a randomized controlled trial. &nbsp;</td> <td>11157015</td> <td>Abstract PURPOSE: Self-directed and supervised exercise were compared with usual care in a clinical trial designed to evaluate the effect of structured exercise on physical functioning and other dimensions of health-related quality of life in women with stages I and II breast cancer. PATIENTS AND METHODS: One hundred twenty-three women with stages I and II breast cancer completed baseline evaluations of generic and disease- and site-specific health-related quality of life, aerobic capacity, and body weight. Participants were randomly allocated to one of three intervention groups: usual care (control group), self-directed exercise, or supervised exercise. Quality of life, aerobic capacity, and body weight measures were repeated at 26 weeks...</td> <td>include or exclude</td> <td>Clinical Trial | Comparative Study | Randomized Controlled Trial | Research Support, Non-U.S. Gov't | Antineoplastic Combined Chemotherapy Protocols | Breast Neoplasms | Breast Neoplasms | Breast Neoplasms | Chemotherapy, Adjuvant | Exercise | Female | Humans | Middle Aged | Neoplasm Staging | Quality of Life | Radiotherapy, Adjuvant</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>If you use this dataset in your research, please cite our papers.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

e-DIPLOMA - Dataset: European remote e-learning ecosystem survey data

<p>This is the dataset "European remote e-learning ecosystem survey data" of the e-DIPLOMA project.</p> <p>In 2022 the evaluation of the European tertiary training ecosystem capacity for using disruptive technologies in practice based e-learning was explored. It was done in the eDiploma project WP2. The research problem was: What are the main gaps in tertiary education in the institutional capacity to perform practice based e-learning with disruptive technologies? Three survey instruments were developed for three target groups in institutions: technology specialists, educators and students. The survey was composed of four blocks of capacity elements:&nbsp;</p> <ul> <li> <p>infrastructural capacities,&nbsp;</p> </li> <li> <p>normative and regulatory capacities (institutional level),&nbsp;</p> </li> <li> <p>teaching cultures (community level),&nbsp;</p> </li> <li> <p>competences, attitudes and values (personal level).&nbsp;</p> </li> </ul> <p>The data were collected with the anonymous web based survey approach in countries: Spain, Estonia, Hungary, Bulgaria, Italy, Cyprus.&nbsp;</p> <p>In each HEI or VET institution the respondents were:</p> <ul> <li> <p>Technical and didactical support staff: educational technologist, IT or technical support specialists, lecturers responsible for technology training, Digital policy administrative specialist</p> </li> <li> <p>Lecturers or researchers who have experiences with some forms of group-learning or practice based learning</p> </li> <li> <p>Students from the institution who have experiences with some forms of group-learning or practice based learning / to be spread among each institution, so that different areas students respond, these should not be one group from one class only)</p> </li> </ul> <p>The answers were collected totally from the following number of the technology specialists-experts (N=96), the educators (N=351), and the students (N=516). The generalizability of the data is limited due to the sampling structure: it was not attempted to reach regional coverage because countries in our sample differ greatly in size. In Estonia responses were collected from 9 institutions (3 vocational schools and 6 HEIs). In Bulgaria responses were from 3 institutions (all HEIs). In Cyprus responses were from 3 institutions (all HEIs). In Hungary responses were from 6 institutions (1 vocational school and 5 HEIs). In Spain responses were from 116 institutions (28 high schools, 41 vocational schools, 47 HEIs). In Italy responses were from 9 institutions (4 HEIs and 5 social enterprises).&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Stiffness Moduli Modelling and Prediction in Four-Point Bending of Asphalt Mixtures: A Machine Learning-Based Framework within Weave-UNISONO 2021 project, NCN project No 2021/03/Y/ST8/00079, and GACR project GA22-04047K

<div><strong>Summary:</strong></div> <div>Two selected mixtures were thoroughly investigated in an experimental trial carried out by means of a four-point bending test (4PBT) apparatus. The mixtures were prepared using spilite aggregate, a conventional 50/70 penetration grade bitumen, and limestone filler. Their stiffness moduli (SM) were determined while samples were exposed to 11 loading frequencies (from 0.1 to 50 Hz) and 4 testing temperatures (from 0 to 30 &deg;C). Observations were recorded and used to develop a machine learning (ML) model. The main scope was the prediction of the stiffness moduli based on the volumetric properties and testing conditions of the corresponding mixtures, which would provide the advantage of reducing the laboratory efforts required to determine them.</div> <div>&nbsp;</div> <div><strong>The dataset includes:</strong></div> <div>Characteristics of bituminous binder, CSV raw data</div> <div> <ul> <li>bituminous binder.csv</li> </ul> </div> <div>Grading curves of tested asphalt mixtures</div> <ul> <li>AML16 Grading curves.csv</li> <li>AMP22 Grading curves.csv</li> </ul> <div>Volumetric characterizations of AML16 and AMP22 mixtures</div> <ul> <li>AML16 Volumetric characterizations.csv</li> <li>AMP22 Volumetric characterizations.csv</li> </ul> <div>Outcomes of the 4PBT experimental trial carried out on AML16 and AMP22 mixtures</div> <ul> <li>AML16 Stiffness Modulus 4PB.csv</li> <li>AMP22 Stiffness Modulus 4PB.csv</li> </ul>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record