Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

285

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

285 results for “Machine learning dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset of experimental measurements for "Demonstration of quantum advantage in machine learning"

<p>Dataset of experimental measurements for "Demonstration of quantum advantage in machine learning",  <em>npj Quantum Information</em><strong> 3</strong>, Article number: 16 (2017).</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

Dataset supporting "Using Machine Learning to decide when to Precondition Cylindrical Algebraic Decomposition with Groebner Bases"

<p>Dataset supporting the paper:</p> <p>Z. Huang, M. England, J.H. Davenport and L.C. Paulson<br> Using Machine Learning to decide when to Precondition Cylindrical Algebraic Decomposition with Groebner Bases.<br> Proceedings of the 18th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC '16), pp. 45--52. IEEE, 2016. Digital Object Identifier: 10.1109/SYNASC.2016.020 </p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

TCOM-H2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric H2O profile dataset [1991-2021] constructed using machine-learning.

<p>Methodology: &nbsp;</p> <p>The <strong>TOMCAT simulation</strong> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized <strong>ERA-5 reanalysis data</strong>.</p> <h3>H2O Profile Processing and Bias Correction</h3> <p><strong>Collocated H2O profiles</strong> are organized into five distinct latitude bins:</p> <ul> <li> <p><strong>NH polar</strong>: 90∘N - 50∘N</p> </li> <li> <p><strong>NH mid-lat</strong>: 20∘N - 70∘N</p> </li> <li> <p><strong>Tropics</strong>: 40∘S - 40∘N</p> </li> <li> <p><strong>SH mid-lat</strong>: 70∘S - 20∘S</p> </li> <li> <p><strong>SH polar</strong>: 90∘S - 50∘S</p> </li> </ul> <p>Initially, <strong>differences between TOMCAT and satellite measurements</strong> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from 10,km to 60,km). Note that TOMCAT may not accurately capture H2O evolution post-HTHH eruption due to the sparse spatial coverage of ACE measurements, which limits training data.</p> <p><strong>Separate XGBoost regression models</strong> are then trained for these H2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate <strong>H2O bias corrections</strong> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</p> <p><strong>Height-resolved H2O profile data</strong> are then interpolated onto 28 standard pressure levels (from 300,hPa to 0.1,hPa), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</p> <p>We acknowledge the inherent <strong>dry biases in the original TOMCAT H2O profiles</strong>, largely because the TTL entry mixing ratios are based on a simplistic sinusoidal seasonal cycle, which omits the H2O enhancement contributed by tropical convective clouds.</p> <h3>Data Files</h3> <p>The dataset includes two files containing daily mean zonal mean H2O profiles:</p> <ul> <li> <p><code>zmh2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>height level data</strong> (10,km to 60,km).</p> </li> <li> <p><code>zmh2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</code>: Contains <strong>pressure level data</strong> (300,hPa to 0.1,hPa).</p> </li> </ul> <h3>Reference Publication</h3> <p>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</p> <p>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, <a title="null" href="https://doi.org/10.5194/essd-15-5105-2023">https://doi.org/10.5194/essd-15-5105-2023</a>, 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

TCOM-O3: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric ozone profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;TOMCAT simulation is performed at T64L32 resolution for the 2000-2024 time period. Collocated Ozone (O3) profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement &nbsp;differences are calculated for each zonal bins (51 height levels, 10km to 60km). Note that if enough ACE measurements are not avaliable for a particular level then data is purely based on TOMCAT simulated output field. Separate XGBoost regression models are trained for the &nbsp;differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids. &nbsp;TOMCAT output sampled at 1.30 pm local time at the equator. Estimated corrections for a given model grid that are added to the original TOMCAT simulated day and night time ozone profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values. &nbsp;For more details see attached presentation. Previous version use both HALOE and ACE data. Here only ACE data is used (hence starting date is 01 January 2000). PDF file shows comparison between v1.0 and v1.1 as well as TOMCAT data.</p> <p>Dataset also includes two files containing daily mean zonal mean hydrogen fluoride &nbsp;profiles on height (10-60 km) and pressure (300-0.1 hPa) levels:</p> <p>zmo3_TCOM_hlev_T2Dz_2000_2024.nc &ndash; height level data (10 to 60 km)</p> <p>zmo3_TCOM_plev_T2Dz_2000_2024.nc &ndash; pressure level data (300 to 0.1 hPa)</p> <p>Daily 3D profiles on height and pressure levels would be made available on request.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

CLRD-GLPS: A Long-term Seasonal Dataset of Ruminant Livestock Distribution in China's Grazing Production Systems (2000-2021) Using Stacking-based Interpretable Machine Learning

<p>Advanced computational methods integrating ensemble learning with interpretable machine learning are essential for precision livestock management under increasing environmental constraints and food security pressures. This study develops a novel stacking-based interpretable machine learning (IML) framework that combines multiple algorithms with SHAP analysis techniques to generate the China's Long-term Ruminant Livestock Distribution in Grazing Livestock Production Systems (CLRD-GLPS) dataset. Our computational approach addresses critical challenges in livestock distribution modelling: livestock segmentation and spatial prediction accuracy. The framework integrates Random Forest, XGBoost, CatBoost, LightGBM, and Extra Trees through a two-layer stacking architecture, enhanced with SHAP (Shapley Additive Explanations) analysis for model interpretability. We also implemented interpretable machine learning for livestock production system segmentation to distinguish grazing from total livestock populations. The stacking ensemble demonstrated superior performance over individual algorithms, achieving R&sup2; values of 0.954-0.961 for cattle and 0.896-0.901 for sheep and goats, with improvements of up to 8.3% compared to best performance single-model approaches. Multi-scale validation confirmed computational robustness: livestock segmentation achieved R&sup2; = 0.80 at county level, while independent city-level validation of CLRD-GLPS datasets yielded R&sup2; = 0.76-0.80. SHAP interpretability analysis revealed distinct environmental drivers, with vegetation indices and topography primarily influencing cattle distribution, while snow conditions and elevation dominated sheep and goat patterns. This computational framework advances livestock distribution modelling through enhanced prediction accuracy, model stability, and interpretability, while the CLRD-GLPS dataset provides essential spatial-temporal information for rangeland sustainability assessments and evidence-based livestock management policies. This dataset is supported by the Second Tibetan Plateau Scientific Expedition and Research Program (STEP, grant no. 2019QZKK0906).</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

GLAB-VOD: Global L-band AI-Based Vegetation Optical Depth Dataset Based on Machine Learning and Remote Sensing

<p>GLAB VOD is a Global L-band Ai-Based vegetation optical depth dataset with 18-day temporal and 25 km spatial resolution, covering 2002 to 2020. The dataset is created using a neural network with SMOS-SMAP-INRAE-BORDEAUX (SMOSMAP-IB) VOD product as a target (over 2015-2020) and brightness temperatures (TB) from the SMOS, AMSR-E, and AMSR-2 spaceborne missions alongside with a novel soil moisture dataset (CASM) as inputs. The GLAB-VOD dataset was created using a recently developed methodology previously used to create a long-term consistent soil moisture dataset CASM, adapted to the&nbsp; VOD retrievals. First, the TB and VOD signals were divided into fixed seasonal cycle and residuals, where the residual part of the signal contains sub-seasonal periodic signals, trends, extremes, and noise. Then, a multi-staged neural network training scheme was used to achieve internally consistent predictions by merging data from different sources without introducing biases or compromising data distribution. A side-product of this project is GLAB TB - a global long-term brightness temperature dataset that matches SMOS TB quality and spawns back to 2002.&nbsp;GLAB TB has daily temporal resolution and 25 km spatial resolution.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Dataset for Machine Learning Assisted Citation Screening for Systematic Reviews

<p>The work "Machine Learning Assisted Citation Screening for Systematic Reviews" explored the problem of citation screening automation using machine-learning (ML) with an aim to accelerate the process of generating <a href="https://en.wikipedia.org/wiki/Systematic_review#:~:text=Systematic%20reviews%20are%20a%20type,synthesize%20findings%20qualitatively%20or%20quantitatively." rel="nofollow">systematic reviews</a>. Manual process of citation screening involve two reviewers manually screening the searched studies using a predefined inclusion criteria. If the study passes the "inclusion" criteria, it is included for further analysis or is excluded. As apparant through manual screening process, the work considered citation screening as a binary classification problem whereby any ML classifier could be trained to separate the searched studies into these two classes (include&nbsp;and&nbsp;exclude).</p> <p>&nbsp;</p> <p>A physiotherapy citation screening dataset was used to test automation approaches and the dataset includes the studies identified for citation screening in an update to the systematic review by Hilfiker <em>et al.</em> The dataset included titles and abstracts (citations) from 31,279 (deduplicated: 25,540) studies identified during the search phase of this SR. These studies were already manually assessed for relevance and labelled by two reviewers into two mutually exclusive labels. The uploaded file consists of 25,540 data samples, with each data sample separated by a new line. It is a tab separated file and the data in it is structured as shown below. This dataset was manually labelled into include and exclude by Hilfiker&nbsp;<em>et al.</em></p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Title</strong></td> <td><strong>PMID</strong></td> <td><strong>Abstract&nbsp;</strong></td> <td><strong>Class</strong></td> <td><strong>MeSH terms (separated by a pipe)</strong></td> </tr> <tr> <td>Structured exercise improves physical functioning in women with stages I and II breast cancer: results of a randomized controlled trial. &nbsp;</td> <td>11157015</td> <td>Abstract PURPOSE: Self-directed and supervised exercise were compared with usual care in a clinical trial designed to evaluate the effect of structured exercise on physical functioning and other dimensions of health-related quality of life in women with stages I and II breast cancer. PATIENTS AND METHODS: One hundred twenty-three women with stages I and II breast cancer completed baseline evaluations of generic and disease- and site-specific health-related quality of life, aerobic capacity, and body weight. Participants were randomly allocated to one of three intervention groups: usual care (control group), self-directed exercise, or supervised exercise. Quality of life, aerobic capacity, and body weight measures were repeated at 26 weeks...</td> <td>include or exclude</td> <td>Clinical Trial | Comparative Study | Randomized Controlled Trial | Research Support, Non-U.S. Gov't | Antineoplastic Combined Chemotherapy Protocols | Breast Neoplasms | Breast Neoplasms | Breast Neoplasms | Chemotherapy, Adjuvant | Exercise | Female | Humans | Middle Aged | Neoplasm Staging | Quality of Life | Radiotherapy, Adjuvant</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>If you use this dataset in your research, please cite our papers.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

3d Transition Metal K-edge XANES Dataset for Machine Learning Models

<p><strong>Data</strong><br><br>This dataset contains machine learning data for K-edge X-ray Absorption Near-Edge Structure (XANES) prediction models for eight 3d transition metals (Ti -Cu).</p> <ul> <li><strong>features_and_spectra:</strong> Material features (X) and corresponding XAS spectra (y) for each dataset split: training (train), validation (val), and test.</li> <li><strong>&nbsp;material_id_and_site:</strong> Material identifiers and site indices (according to <a href="https://github.com/AI-multimodal/Lightshow">Lightshow</a>) for each dataset split.&nbsp;</li> </ul> <p><strong>Funding</strong><br><br>This research is based upon work supported by the U.S. Department of Energy, Office of Science, Office Basic Energy Sciences, under Award Number FWP PS-030. This research also used theory and computational resources of the Center for Functional Nanomaterials, which is a U.S. Department of Energy Office of Science User Facility, and the Scientific Data and Computing Center, at Brookhaven National Laboratory under Contract No. DE-SC0012704 and by Brookhaven National Laboratory (BNL), Laboratory Directed Research and Development (LDRD) grant no. 24-004.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics

<h1><strong><span><span>Dataset from the Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University:&nbsp; </span></span></strong></h1> <p>&nbsp;</p> <div> <div> <p><span><span>The </span><span>Social Media</span><span> &amp; Hate research lab at the Institute for the Study of Contemporary Antisemitism compiled this dataset using an annotation portal (Jikeli, Soemer, and Karali 2024), which was used to label tweets as either antisemitic or non-antisemitic, among other labels. Note that annotation was done on live data, including images and context, such as threads. All data was annotated by two experts, and all discrepancies were discussed</span><span> (Jikeli et al. 2023)</span><span>.</span></span><span>&nbsp;</span></p> </div> </div> <p><br><strong>Content: </strong></p> <p><span><span>This dataset&nbsp;</span><span>contains</span> <span>1</span><span>1</span><span>311</span><span> tweets </span><span>covering</span><span>&nbsp;a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and </span><span>April 2023</span><span>.&nbsp;</span><span>The dataset consists of random samples of relevant keywords during this </span><span>time period</span><span>.</span><span>&nbsp;1,</span><span>953</span><span> tweets (1</span><span>7</span><span>%) </span><span>are antisemitic </span><span>according to </span><span>the IHRA definition of antisemitism.</span><span>&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> <div> <p><span><span>The distribution of tweets by year is as follows:</span><span>&nbsp;1499 (</span><span>13</span><span>%) from 2019, 371</span><span>2</span><span> (</span><span>33</span><span>%) from 2020, </span><span>2591</span><span> (2</span><span>3</span><span>%) from 2021</span><span>, 2644 from 2022 </span><span>(23%)</span> <span>and 865 </span><span>(8%)</span> <span>f</span><span>rom 2023</span><span>. </span><span>6365</span><span> (</span><span>56</span><span>%) </span><span>contain</span><span> the keyword "Jews,"</span><span> 4134 </span><span>(</span><span>3</span><span>7</span><span>%) include "Israel," 529 (</span><span>5</span><span>%) feature the derogatory term "</span><span>ZioNazi</span><span>*," and 283 (</span><span>3</span><span>%) use the slur "K---s." Some tweets may </span><span>contain</span><span> multiple keywords.&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>725</span><span> out of the </span><span>6365</span><span> tweets with the keyword "Jews" (11%) and </span><span>664</span><span> out of the </span><span>4134</span><span> tweets with the keyword "Israel" (1</span><span>6</span><span>%) were classified as antisemitic. 97 out of the 283 tweets using the antisemitic slur "K---s" (34%) are antisemitic.</span> <span>Interestingly, many tweets featuring the slur "K---s" actually </span><span>call out</span><span> its u</span><span>s</span><span>e.</span><span> In contrast, </span><span>the majority of</span><span> tweets </span><span>using</span><span> the derogatory term "</span><span>ZioNazi</span><span>*" are antisemitic, with 467 out of 529 (88%) being classified as such.&nbsp;</span></span><span>&nbsp;</span></p> </div> <p>&nbsp;</p> <p><strong>File Description:&nbsp;</strong></p> <div> <div> <p><span><span>The dataset is provided in a csv file format, with each row </span><span>representing</span><span> a single message, including replies, quotes, and retweets. The file </span><span>contains</span><span> the following columns:&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&nbsp;</span></span><span><span>&nbsp;</span><br></span><span><span>&lsquo;ID&rsquo;:</span> <span>Represents</span><span> the tweet ID.&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Username&rsquo;: </span><span>Represents</span><span> the username </span><span>that posted </span><span>the tweet</span><span>.&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Text&rsquo;: </span><span>Represents</span><span> the full text of the tweet (not pre-processed).</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;</span><span>CreateDate</span><span>&rsquo;: </span><span>Represents</span><span> the date </span><span>on which </span><span>the tweet was created</span><span>.&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Biased&rsquo;: </span><span>Represents</span><span> the label given by our annotations as to whether the tweet </span><span>is antisemitic or no</span><span>t</span><span>.</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Keyword&rsquo;: </span><span>Represents</span><span> the keyword that was used in the query. The keyword can be in the text, including </span><span>hashtags, </span><span>mentioned </span><span>users</span><span>, or the username</span><span> itself.</span><span>&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> </div> <p>&nbsp;</p> <p>Licences&nbsp;</p> <p>Data is published under the terms of the "Creative Commons Attribution 4.0 International" licence (https://creativecommons.org/licenses/by/4.0)&nbsp;</p> <p>&nbsp;</p> <p>Acknowledgements&nbsp;</p> <p>We are grateful for the support of Indiana University&rsquo;s Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media &amp; Hate Research Lab at Indiana University&rsquo;s Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenh&auml;usler, Sophie von M&aacute;ri&aacute;ssy, Mabel Poindexter, Jenna Solomon, Clara Schilling, and Victor Tschiskale.&nbsp;</p> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Simulated datasets for detector and particle flow reconstruction: CLIC detector, machine learning format

<p><strong>Synopsis</strong></p> <p>Machine-learning friendly format of tracks, clusters and target particles in electron-positron events, simulated with the CLIC detector.&nbsp;Ready to be used with <a href="https://zenodo.org/records/14930299">jpata/particleflow:v2.3.0</a>. Derived from the EDM4HEP ROOT files in <a href="https://zenodo.org/record/8260741">https://zenodo.org/record/8260741</a>.</p> <ul> <li>clic_edm_ttbar_pf.zip: e+e- -&gt; ttbar, center of mass energy at 380 GeV</li> <li>clic_edm_qq_pf.zip: e+e- -&gt; Z* -&gt; qqbar, center of mass energy at 380 GeV</li> <li>clic_edm_ww_fullhad_pf.zip: e+e- -&gt; WW -&gt; W decaying hadronically, center of mass energy at 380 GeV</li> <li>clic-tfds.ipynb: an example notebook on how to load the files</li> </ul> <p><strong>Contents</strong></p> <p>Each .zip file contains the dataset in the <a href="https://github.com/tensorflow/datasets">tensorflow-datasets</a>,&nbsp;<a href="https://github.com/google/array_record">array_record</a> format. We have split the full datasets into 10 subsets, due to space considerations on zenodo, two subsets from each dataset are uploaded. Each dataset contains a train and test split of events.</p> <p><strong>Dataset semantics (to be updated)</strong></p> <p>Each dataset consists of events that can be iterated over using the tensorflow-datasets library and used in either tensorflow or pytorch. Each event has the following information available:</p> <ul> <li>X: the reconstruction input features, i.e. tracks and clusters</li> <li>ytarget: the ground truth particles with the features ["PDG", "charge", "pt", "eta", "sin_phi", "cos_phi", "energy", "jet_idx"], with "jet_idx" corresponding to the gen-jet assignment of this particle</li> <li>ycand: the baseline Pandora PF particles with the features ["PDG", "charge", "pt", "eta", "sin_phi", "cos_phi", "energy", "jet_idx"], with "jet_idx" corresponding to the gen-jet assignment of this particle</li> </ul> <p>The full semantics, including the list of features for X, are available at https://github.com/jpata/particleflow/blob/v2.3.0/mlpf/heptfds/clic_pf_edm4hep/utils_edm.py and https://github.com/jpata/particleflow/blob/v2.3.0/mlpf/data/key4hep/postprocessing.py.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Machine learning code and dataset for "Nowcasting thunderstorm hazards using machine learning: the impact of data sources on performance"

<p>This repository contains the code and dataset for the paper:</p> <p>Nowcasting&nbsp;thunderstorm&nbsp;hazards&nbsp;using&nbsp;machine&nbsp;learning:&nbsp;the&nbsp;impact&nbsp;of&nbsp;data&nbsp;sources&nbsp;on&nbsp;performance,&nbsp;Natural&nbsp;Hazards&nbsp;and&nbsp;Earth System&nbsp;Sciences,&nbsp;2022,&nbsp;<a href="https://doi.org/10.5194/nhess-2021-171">https://doi.org/10.5194/nhess-2021-171</a></p> <p>The GitHub code repository at <a href="https://github.com/meteoswiss-mdr/ts-nowcast-datasources">https://github.com/meteoswiss-mdr/ts-nowcast-datasources</a> may contain a more up-to-date version of the code if bug fixes etc. have been necessary. The file <a href="https://zenodo.org/api/files/41faa1b7-17f6-4a75-be09-7743426ef13c/ts-nowcast-datasources-publication.zip">ts-nowcast-datasources-publication.zip</a> in this Zenodo release contains the status of the GitHub repository at the time of the publication of the paper.</p> <p>For instructions for using the data, please see the <a href="https://github.com/meteoswiss-mdr/ts-nowcast-datasources">code repository</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Dataset for: Application of Machine Learning for the Spatial Analysis of Binaural Room Impulse Responses

<p>This repository contains supplementary material for the paper titled `Application of Machine Learning for the Spatial<br> Analysis of Binaural Room Impulse Responses&#39; Available at: <a href="http://dx.doi.org/10.3390/app8010105">dx.doi.org/10.3390/app8010105</a>&nbsp;. These programs and audio files are distributed in the hopes that they will prove useful under the Creative Commons Attribution 4.0, with no warranty; or the implied warranty of merchantability or fitness for a particular problem. Please give appropriate credit for use of the material provided in this repository back to the author.&nbsp;</p> <p>In order to use the MatLab code the Auditory Toolbox by Malcolm Slaney [1] and the Cochleagram function distributed by Bin Gao [2] are required.</p> <p>The python scrips require the following Python libraries to be installed: Numpy[3], SciPy[4] and Tensorflow [5].</p> <p>The MatLab code was tested using MatLab R2017a on a Computer running windows 7.</p> <p>The python code was tested using Python 3.2.5, using an anaconda Python environment - in windows command line.</p> <p>--</p> <p>The repository contains:</p> <p>Folders:</p> <p><br> 1.) neg90 - This folder contains the gaussian normalisation parameters stored as text files and the weights and biases for the trained neural network - these are all for the -90&deg; rotation neural network.</p> <p>2.) pos90 - This folder contains the gaussian normalisation parameters stored as text files and the weights and biases for the trained neural network - these are all for the +90&deg; rotation neural network.</p> <p>3.) testData - this folder contains pre-generated test data for the different binaural dummy head microphones, speaker, and signal type combinations.</p> <p>Python Scripts:</p> <p><br> 1.) AnalyseDoA.py - A python script that can be run to test the neural network using the pre-generated test data - running the script will allow the user to input the binaural dummy head, speaker, and signal type. The important variables generated by this script are DoA - the direction of arrival for each signal in the feature vector, and yDiff - the difference between the predicted DoA and the expected direction of arrival</p> <p>2.) DirectionAnalysis.py - This python file contains a set of function that are used to define the neural network, and run it. The function called DoAPrediction takes the feature vector generated by the MatLab code as its input argument, these features will then be passed to the neural network, and the output of this function is the direction of arrival predicted by the neural network for each signal. The functions: DoAAnalysis_neg90 and DoAAnalysis_pos90 are called by the DoAPrediction function, these functions create the neural network using the NN function, import the weights and biases, and passes the feature matrix (provided as input) through the neural network - the output of these functions are the predicted direction of arrival.</p> <p>MatLab files:</p> <p><br> 1.) runAnalysis.m - This&nbsp;MatLab script&nbsp;analyses the dataset provided as part of this repository. Users can change the variables head (&#39;KEMAR&#39; or &#39;KU100&#39;), signalType (&#39;directSound&#39; or &#39;reflection&#39;), and speaker (&#39;EquatorD5&#39; or &#39;Genelec8030&#39;). This script will produce the gaussian normalised feature vector and expected direction of arrival for all signals with the defined head, signal type, and speaker combination. These variables are then saved in .mat files so they can be imported by the python scripts.</p> <p>2.) BinauralModelCochlea.m - This MatLab function analyses a given binaural signal and outputs the interaural cross-correlation, interaural level difference, interaural time difference, the cochlea output for the left and right channel and the centre frequencies of the gammatone filter band. The input variables are: IR - the signal to be analysed, N - the number of gammatone filters, freqLow - the lowest centre frequency of the gammatone filter bank (centre frequency of the first gammatone filter), and freqHigh - the highest centre frequency of the gammatone filter bank (the centre frequency of the Nth gammatone filter). This function requires Malcolm Slaney&#39;s Auditory Toolbox [1] and Bin Gao&#39;s Cochleagram function [2] in order to work.</p> <p>3.) generateFeatureVector.m - This MatLab function generates a feature vector from an input binaural signal x, and a version of the signal captured after the binaural dummy head has been rotated by either +90&deg; or -90&deg; degree (variables xPos90 and xNeg90 respectively). If the sampling frequency (Fs) isn&#39;t 44100, the signals are resampled to be at 44100. This file also contains a function &#39;gaussianNormalisationTestData&#39; which gaussian normalises the data using the mean and standard deviation calculated from the data used to train the neural networks - the mean and standard deviation values are stored in the folder GMParams in the pos90 and neg90 folders.</p> <p>4.) generateTestData.m - This&nbsp;MatLab function analyses the included binaural dataset, it takes the input variables: head - the binaural dummy head used for the measurements either &#39;KEMAR&#39; or &#39;KU100&#39;, speaker - the speaker used for the measurements either &#39;EquatorD5&#39; or &#39;Genelec8030&#39;, and signalType - the type of signal being analysed either &#39;directSound&#39; or &#39;reflection&#39;.</p> <p>Text files:</p> <p><br> 1.) noLayers.txt - a text file containing the number of layers used when training the neural network - with the current version of the code the neural network contains only 1 layer.</p> <p>2.) README.txt - Read me file containing information about the repository.</p> <p>Audio files:</p> <p><br> This repository contains 1152 binaural signals half of which are direct sounds segmented from a binaural room impulse responses and the other half are reflections segmented from binaural room impulse responses (detailed in the paper this material supports) the direct sounds are recorded at angles from 0&deg; to 357.5&deg; in steps of 2.5&deg; and the reflections are recorded at angles of 1&deg; to 358.5&deg; in steps of 2.5&deg;. In the paper only recordings relating to signals recorded with the Equator D5 are analysed.</p> <p>The combination of audio files include:</p> <p>1.) 144 direct sound recordings captured with the KEMAR 45BC binaural dummy head microphone and the Equator D5 speaker<br> 2.) 144 reflection recordings captured with the KEMAR 45BC binaural dummy head microphone and the Equator D5 speaker<br> 3.) 144 direct sound recordings captured with the KU100 binaural dummy head microphone and the Equator D5 speaker<br> 4.) 144 reflection recordings captured with the KU100 binaural dummy head microphone and the Equator D5 speaker<br> 5.) 144 direct sound recordings captured with the KEMAR 45BC binaural dummy head microphone and the Genelec 8030 speaker<br> 6.) 144 reflection recordings captured with the KEMAR 45BC binaural dummy head microphone and the Genelec 8030 speaker<br> 7.) 144 direct sound recordings captured with the KU100 binaural dummy head microphone and the Genelec 8030 speaker<br> 8.) 144 reflection recordings captured with the KU100 binaural dummy head microphone and the Genelec 8030 speaker</p> <p>The files are stored using the following file naming convention:<br> head_Test3_speaker_signalType_000_0_Degrees.wav - where _000_0 defines the azimuth direction of arrival so for example for a direct sound measured with the KEMAR unit and the Genelec8030 at 5 degrees would be &#39;KEMAR_Test3_Genelec8030_directSound_005_0Degrees.wav&#39; and for a reflection measured with the KU100 and the Equator D5 at 298.5 degrees would be &#39;KU100_Test3_EquatorD5_reflection_298_5Degrees.wav&#39;</p> <p>--</p> <p>Bibliography:<br> [1]&nbsp;Slaney, M. (1998). Auditory Toolbox. Palo Alto, CA. [Online]. Available: https://engineering.purdue.edu/~malcolm/interval/1998-010/ [Accessed: Oct. 27, 2017]</p> <p>[2]&nbsp;Gao, B. (2014). Cochleagram and IS-NMF2D for Blind Source Separation. [Online] Available:&nbsp;http://uk.mathworks.com/matlabcentral/fileexchange/48622-cochleagram-and-is-nmf2d-for-blind-source-separation?focused=3855900&amp;tab=function&nbsp;[Accessed: Oct. 27, 2017]</p> <p>[3]&nbsp;NumFocus. (n.d.). NumPy. [Online]. Available: http://www.numpy.org/ [Accessed: Oct. 27, 2017]</p> <p>[4]&nbsp;SciPy. (n.d.). SciPy. [Online]. Available:&nbsp;https://www.scipy.org/&nbsp;[Accessed: Oct. 27, 2017]</p> <p>[5]&nbsp;Google. (n.d.). TensorFlow. [Online] Available:&nbsp;https://www.tensorflow.org/&nbsp;[Accessed: Oct. 27, 2017]</p> <p>--</p> <p>All code and audio produced by: Michael Lovedee-Turner, PhD candidate in Music Technology at the Audio Lab, Department of Electronic Engineering, University of York</p> <p>Contact: mjlt500@york.ac.uk</p>

opencc-by-4.0Oct 2017View details →
zenodo44/100

Global derived datasets for use in k-NN machine learning prediction of global seafloor total organic carbon

<p>This&nbsp;dataset includes 663 predictor grids used for k-NN global prediction of seafloor total organic carbon.</p> <p>663 predictor grids available in netCDF4 HDF5 file format. Grids are cell-centered sized 4320 x 2160. File names adhere to the naming conventions discussed below. The naming structure is partioned by underscores and periods in the following order: interface to which the gridded values refer to, quantity of values contained within the grid, units and reference values/units (e.g. meters below sea level), data source, statistic calculated (if applicable), grid pitch, and file extension.</p> <p>Possible interfaces from the top &ndash; down:</p> <p>SS &ndash; Sea surface &ndash; atmosphere interface (may also be average of the entire water column)</p> <p>SF &ndash; Seafloor &ndash; water interface (may also be denoted by GL)</p> <p>GL&nbsp;&nbsp; &ndash; Ground level (e.g. bottom of pure liquid, top of dirt)</p> <p>SC &ndash; Sediment &ndash; crust interface (e.g. sediment above, igneous/metamorphic below)</p> <p>CM &ndash; Crust &ndash; mantle interface (e.g. Mohorovicic discontinuity)</p> <p>Appropriate reference naming marker (bold), original data source, and date of last access:</p> <p><strong>Becker</strong></p> <p>Becker, J. J., Wood, W. T., &amp; Martin, K. M. (2014). <em>Global crustal heat flow using random decision forest prediction</em>, Abstract NG31A-3788 presented at 2014 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 06/23/2015.</p> <p><strong>CRUST1</strong>&nbsp;</p> <p>Pasyanos, M.E., Masters, G., Laske, G. &amp; Ma, Z. (2012). <em>LITHO1.0 - An Updated Crust and Lithospheric Model of the Earth Developed Using Multiple Data Constraints</em>, Abstract T11D-09 presented at 2012 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 07/01/2014.</p> <p><strong>CRUST1_NOAA</strong></p> <p>&nbsp;As the NOAA sediment thickness database is globally not complete, data gaps in the NOAA grid with this have been supplemented by the CRUST1 sediment thickness (see above citation).</p> <p>Whittaker, J., Goncharov, A., Williams, S., M&uuml;ller, R. D., &amp; Leitchenkov, G. (2013) Global sediment thickness dataset updated for the Australian-Antarctic Southern Ocean, <em>Geochemistry, Geophysics, Geosystems. </em>https://doi.org/10.1002/ggge.2018.<em> </em>Last access:&nbsp; 09/02/2018.</p> <p><strong>GVP</strong></p> <p>Global Volcanism Program (2013) Volcanoes of the World. In E. Venzke (ed.). (Vol. 4.7.3).&nbsp; Smithsonian Institution. https://doi.org/10.5479/si.GVP.VOTW4-2013. Last access: 09/22/2014.</p> <p><strong>ETOPO2v2</strong></p> <p>National Geophysical Data Center (2006). 2-minute Gridded Global Relief Data (ETOPO2) v2. National Geophysical Data Center, NOAA. DOI: 10.7289/V5J1012Q. Last access: 02/06/2013.</p> <p><strong>PLATES</strong></p> <p>Coffin, M.F., Gahagan, L.M., &amp; Lawver, L.A. (1998). Present-day Plate Boundary Digital Data Compilation. University of Texas Institute for Geophysics Technical Report (No. 174, pp. 5). Last access: 09/15/2014.</p> <p><strong>ONRL</strong></p> <p>Ludwig,W., Amiotte-Suchet, P., &amp; Probst, J. L. (2011). ISLSCP II Global River Fluxes of Carbon and Sediments to the Oceans. In F. G. Hall, G. Collatz, B. Meeson, S. Los, E. Brown de Colstoun, and D. Landis (Eds.), <em>ISLSCP Initiative II Collection</em>. Oak Ridge National Laboratory Distributed Active Archive Center, Oak Ridge, Tennessee, U.S.A. http://dx.doi.org/10.3334/ORNLDAAC/1028. Last Access: 02/15/2015.</p> <p><strong>Muller</strong></p> <p>M&uuml;ller, R. D., Sdrolias, M., Gaina, C., &amp; Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world&rsquo;s ocean crust, <em>Geochemistry, Geophysics, Geosystems</em>, 9(4), Q04006. https://doi.org/10.1029/2007GC001743. Last accessed: 07/19/2011.</p> <p><strong>Woa13x</strong></p> <p>Boyer, T.P., Antonov, J. I., Baranova, O. K., Coleman, C., Garcia, H. E., Grodsky, A., et al. (2013) World Ocean Database 2013. In &nbsp;S. Levitus, A. Mishonov (Ed.), <em>NOAA Atlas NESDIS 72, Technical Ed</em>. Silver Spring, MD. http://doi.org/10.7289/V5NZ85MT. Last Access: 09/18/2014.</p> <p><strong>KIM</strong></p> <p>Kim, S.S. &amp; Wessel, P. (2011). New global seamount census from the altimetry-derived gravity data, <em>Geophysical Journal International</em>, 186, 615-631. https://doi.org/10.1111/j.1365-246X.2011.05076.x.&nbsp; Last access: 09/22/2014.</p> <p><strong>HYCOM</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data.Last access: 03/19/2014.</p> <p><strong>NCEDC</strong></p> <p>NCEDC (2016). Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC. Last access: 09/21/2014.</p> <p><strong>Wei2010</strong></p> <p>Wei, C.-L., Rowe, G. T., Escobar-Briones, E., Boetius, A., Soltwedel, T., Caley, M. J., et al.(2010). Global patterns and predictions of seafloor biomass using random forests. <em>PLoS ONE</em>,5(12), e15323. https://doi.org/10.1371/journal.pone.0015323 Last access: 06/20/2016.</p> <p><strong>NGA_egm2008</strong></p> <p>Pavlis, N.K., Holmes, S. A., Kenyon, S. C., &amp; Factor, J. K. (2008). <em>The</em> <em>EGM2008 Global Gravitational Model</em>, Abstract 2008AGUFM.G22A..01P presented at the 2008 General Assembly of the European Geosciences Union, Vienna, Austria. Last access: 07/10/2014.</p> <p><strong>WAVEWATCH3</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data. Last access: 03/19/2014.</p> <p>Updated global seafloor porosity grid using our k-nearest neighbors algorithm using 5 nearest neighbors. Observed data used for prediction from Martin et al. (2015).&nbsp;</p> <p>Martin, K. M., Wood, W. T., &amp; Becker, J. J. (2015). A global prediction of seafloor sediment porosity using machine learning. <em>Geophysical Research Letters</em>, 42(24), 10640. https://doi.org/10.1002/2015GL065279</p> <p>Other grids which have been generated by empirical means are latitude (and derivatives), longitude (and derivatives), Coriolis, coast_is_1.0, and the random noise grids.&nbsp;</p> <p>Units referenced are as follows:</p> <p>KGM3 - kilogram per cubic meter<br> MS - meters per second<br> KM - kilometer<br> M_ASL - meters above sea level (i.e. meters referenced to sea level)<br> MWM2 - milliwatt per square meter<br> TGCYR - terragram of carbon per year<br> TGYR - terragram per year<br> MA - megaannum<br> M - meters<br> MGCM2 - milligram of carbon per square meter<br> DEG - degree<br> S - seconds</p> <p>Statistics grids are calculated within a given radius (e.g. 10km, 50km, 125km, 250km, 500km, 1000km) of the respective cell-centered value. The statistics grids include mean (.men), average absolute deviation from the mean (.aad), and the common logarithm (.log) of the absolute value of the mean (.mlg). Additionally, some grids are a weighted count for given radii (e.g. seamounts) where weight is a cosine taper from the center of the grid cell.&nbsp;</p> <p>The grid pitch for this dataset is uniformly at 5-arc minute denoted by &ldquo;.5m&rdquo;. Additionally, the extension used (netCDF4) is denoted by &ldquo;.nc&rdquo;.</p>

opencc-by-4.0Oct 2018View details →
zenodo44/100

Dataset and code for "Classification of Solar Wind With Machine Learning"

<p>Matlab software and data from http://www.mlspaceweather.org/ for the paper</p> <p>https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017JA024383</p>

opencc-by-4.0Oct 2017View details →
zenodo44/100

Modeling Autonomic Pupillary Responses from External Stimuli Using Machine Learning - Dataset

<p>This page contains the data collected for the paper:&nbsp;<em><strong>Modeling Autonomic Pupillary Responses from External Stimuli Using Machine Learning&nbsp;</strong></em>(<a href="https://doi.org/10.26717/BJSTR.2019.20.003446">DOI:10.26717/BJSTR.2019.20.003446</a>).&nbsp;The&nbsp;dataset consists of spectral and pupillometric data collected during three&nbsp;outdoor/indoor walks. The folders &ldquo;raw&rdquo;, &ldquo;merged&rdquo;, and &ldquo;cleaned&rdquo; contain data collected by the Konica Minolta CL-500A Illuminance Spectrophotometer and Tobii Pro Glasses 2 at three different stages in the data preparation process. The &ldquo;raw&rdquo; folder contains uncleaned and unsynchronized .csv/.json files. The &ldquo;merged&rdquo; folder contains uncleaned, but synchronized light and ocular data in .csv format. The &ldquo;cleaned&rdquo; folder contains a single .csv of cleaned and synchronized data with the derived variables: average pupil diameter and pupil diameter difference.&nbsp;</p> <p>The best choice of data files will depend on desired analysis. More guidance on how to handle this data can be found in the readMe files located in each subsequent folder. More information on the sensing devices used here can be found at the Minolta and Tobii information links below.&nbsp;</p> <p><strong>Minolta Information</strong>: <a href="https://sensing.konicaminolta.us/uploads/cl-500a_instruction217a_eng-250cl60686.pdf">https://sensing.konicaminolta.us/uploads/cl-500a_instruction217a_eng-250cl60686.pdf</a></p> <p><strong>Tobii Information</strong>: <a href="https://www.tobiipro.com/siteassets/tobii-pro/user-manuals/tobii-pro-glasses-2-user-manual.pdf/?v=1.1.3">https://www.tobiipro.com/siteassets/tobii-pro/user-manuals/tobii-pro-glasses-2-user-manual.pdf/?v=1.1.3</a></p> <p>The codes used to prepare, analyze, and visualize this data is available in the LightOcular GitHub Repository linked below.&nbsp;</p> <p><strong>LightOcular GitHub Repo</strong>: <a href="https://github.com/mi3nts/LightOcular">https://github.com/mi3nts/LightOcular</a></p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Dataset of "Comprehensive Machine Learning Approaches for Modelling the State of Charge of Lithium-ion Batteries"

<p>This paper evaluates three ML approaches for SOC modeling in LIBs: the multilayer perceptron (MLP), long short-term memory (LSTM), and the nonlinear autoregressive with exogenous input (NARX) neural network architectures. These models were tested using an experimental dataset with multiple input variables, including electrochemical impedance spectroscopy (EIS) data, voltage, and capacity readings for commercial LIB cells. Results indicate that MLP and LSTM are more adaptable with a smaller training dataset (14 samples), while the NARX model required more than 34 out of 67 samples to achieve reasonable accuracy. Additionally, the NARX model is more sensitive to changes in the learning rate (&alpha;) and exhibits larger output error deviations. The MLP and LSTM models consistently performed well across various hidden layer sizes, showing no upper bound constraints, whereas the NARX model&rsquo;s performance deteriorated with certain hidden layer configurations.</p>

embargoedcc-by-4.0Aug 2024View details →
zenodo44/100

Structured Power Grid Simulation Dataset for Machine Learning: Failure and Survival Events in Grid2Op's L2RPN WCCI 2022 Environment

<p>This dataset was developed for and used in the paper titled <em>"Fault Detection for Agents in Power Grid Topology Optimization: A Comprehensive Analysis"</em> by Malte Lehna, Mohamed Hassouna, Dmitry Degtyar, Sven Tomforde, and Christoph Scholz, presented at the <em>Workshop on Machine Learning for Sustainable Power Systems (ML4SPS)</em>, part of <em>ECML PKDD 2024</em>. While the paper is pending formal publication, a preprint version is available on arXiv.</p> <p>The dataset contains structured training, validation, and test data comprising failure and survival events observed in transmission power grid simulations. These were generated using Grid2Op with the WCCI 2022 L2RPN environment. Each data instance is labeled with one of four classes, representing survival or impending failure in 1, 3, and 5 timesteps. This dataset was used to train, validate and test machine learning models that predict grid agent failures in topology optimization tasks.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

PSML: A Multi-scale Time-series Dataset for Machine Learning in Decarbonized Energy Grids (Dataset)

<p><strong>Abstract</strong></p> <p>The electric grid is a key enabling infrastructure for the ambitious transition towards carbon neutrality as we grapple with climate change. With deepening penetration of renewable energy resources and electrified transportation, the reliable and secure operation of the electric grid becomes increasingly challenging. In this paper, we present PSML, a first-of-its-kind open-access multi-scale time-series dataset, to aid in the development of data-driven machine learning (ML) based approaches towards reliable operation of future electric grids. The dataset is generated through a novel transmission + distribution (T+D) co-simulation designed to capture the increasingly important interactions and uncertainties of the grid dynamics, containing electric load, renewable generation, weather, voltage and current measurements at multiple spatio-temporal scales. Using PSML, we provide state-of-the-art ML baselines on three challenging use cases of critical importance to achieve: (i) early detection, accurate classification and localization of dynamic disturbance events; (ii) robust hierarchical forecasting of load and renewable energy with the presence of uncertainties and extreme events; and (iii) realistic synthetic generation of physical-law-constrained measurement time series. We envision that this dataset will enable advances for ML in dynamic systems, while simultaneously allowing ML researchers to contribute towards carbon-neutral electricity and mobility.&nbsp;</p> <p><strong>Data Navigation</strong></p> <p>Please download, unzip and put somewhere for later benchmark results reproduction and data loading and performance evaluation for proposed methods.</p> <pre><code>wget https://zenodo.org/record/5130612/files/PSML.zip?download=1 7z x 'PSML.zip?download=1' -o./ </code></pre> <p><strong>Minute-level Load and Renewable</strong></p> <ul> <li>File Name <ul> <li>ISO_zone_#.csv: `CAISO_zone_1.csv` contains minute-level load, renewable and weather data from 2018 to 2020 in the zone 1 of CAISO.</li> </ul> </li> <li>- Field Description <ul> <li>Field `<em>time</em>`: Time of minute resolution.</li> <li>Field `<em>load_power</em>`: Normalized load power.</li> <li>Field `<em>wind_power</em>`: Normalized wind turbine power.</li> <li>Field `<em>solar_power</em>`: Normalized solar PV power.</li> <li>Field `<em>DHI</em>`: Direct normal irradiance.</li> <li>Field `<em>DNI</em>`: Diffuse horizontal irradiance.</li> <li>Field `<em>GHI</em>`: Global horizontal irradiance.</li> <li>Field <em>`Dew Point</em>`: Dew point in degree Celsius.</li> <li>Field `<em>Solar Zeinth Angle</em>`: The angle between the sun&#39;s rays and the vertical direction in degree.</li> <li>Field `<em>Wind Speed</em>`: Wind speed (m/s).</li> <li>Field `<em>Relative Humidity</em>`: Relative humidity (%).</li> <li>Field `<em>Temperature</em>`: Temperature in degree Celsius.</li> </ul> </li> </ul> <p><strong>Minute-level PMU Measurements</strong></p> <ul> <li>File Name <ul> <li>case #: The `case 0` folder contains all data of scenario setting #0. <ul> <li>pf_input_#.txt: Selected load, renewable and solar generation for the simulation.</li> <li>pf_result_#.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> <li>Filed Description <ul> <li>Field <em>`time`</em>: Time of minute resolution.</li> <li>Field <em>`Vm_###`</em>: Voltage magnitude (p.u.) at the bus ### in the simulated model.</li> <li>Field <em>`Va_###`</em>: Voltage angle (rad) at the bus ### in the simulated model.</li> <li>Field <em>`P_#_#_#`</em>: `P_3_4_1` means the active power transferring in the #1 branch from the bus 3 to 4.</li> <li>Field <em>`Q_#_#_#`</em>: `Q_5_20_1` means the reactive power transferring in the #1 branch from the bus 5 to 20.</li> </ul> </li> </ul> <p><strong>Millisecond-level PMU Measurements</strong></p> <ul> <li>File Name <ul> <li>Forced Oscillation: The folder contains all forced oscillation cases. <ul> <li>row_#: The folder contains all data of the disturbance scenario #. <ul> <li>dist.csv: Three-phased voltage at nodes in the distribution system via T+D simualtion.</li> <li>&nbsp;info.csv: This file contains the start time, end time, location and type of the disturbance</li> <li>trans.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> <li>Natural Oscillation: The folder contains all natural oscillation cases. <ul> <li>row_#: The folder contains all data of the disturbance scenario #. <ul> <li>dist.csv: Three-phased voltage at nodes in the distribution system via T+D simualtion.</li> <li>info.csv: This file contains the start time, end time, location and type of the disturbance.</li> <li>trans.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> </ul> </li> <li>Filed Description <ul> <li>trans.csv <ul> <li>&nbsp; - Field <em>`Time(s)`</em>: Time of millisecond resolution.</li> <li>&nbsp; - Field <em>`VOLT ###`</em>: Voltage magnitude (p.u.) at the bus ### in the transmission model.</li> <li>&nbsp; - Field <em>`POWR ### TO ### CKT #`</em>: `POWR 151 TO 152 CKT &#39;1 &#39;` means the active power transferring in the #1 branch from the bus 151 to 152.</li> <li>&nbsp; - Field <em>`VARS ### TO ### CKT #`</em>: `VARS 151 TO 152 CKT &#39;1 &#39;` means the reactive power transferring in the #1 branch from the bus 151 to 152.</li> </ul> </li> <li>dist.csv <ul> <li>Field <em>`Time(s)`</em>: Time of millisecond resolution.</li> <li>Field <em>`####.###.#`</em>: `3005.633.1` means per-unit voltage magnitude of the phase A at the bus 633 of the distribution grid, the one connecting to the bus 3005 in the transmission system.</li> </ul> </li> </ul> </li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Dataset of Machine Learning forecasted VTEC from paper: Uncertainty Quantification for Machine Learning-based Ionosphere and Space Weather Forecasting

<p>The *csv files contain forecasted one-day-ahead Vertical Total Electron Content (VTEC), consisting of the mean/median VTEC values and the upper and lower VTEC bounds of the 95% confidence intervals of 4 models based on machine learning for test data.</p> <p>The first part of the *csv file name corresponds to the type of model: SE stands for the super-ensemble VTEC model, QGB stands for the quantile gradient boosting VTEC model, BNN1 stands for the Bayesian neural network VTEC model, and BNN2 stands for the Bayesian neural network with negative log-likelihood (NLL) loss VTEC model. The second part of the file name refers to the geographic location of the VTEC points for which the forecast is performed, i.e., 10E70N for 10 degree of longitude and 70 degree of latitude, 10E40N for 10 degree of longitude and 40 degree of latitude, and 10E10N for 10 degree of longitude and 10 degree of latitude. The last part of the file name corresponds to the test year, i.e., year 2017.</p> <p>The SE_*_2017.csv file consists of 14 columns. The index column (&quot;Date-time&quot;) is expressed in Coordinated Universal Time (UTC) as YYYY-MM-DD. Columns 1-3 contain the VTEC forecast results of Random Forest (RF) trained on three data subsets; columns 4-6 contain the VTEC forecast results of Adaptive Boosting (AB) trained on three data subsets; columns 7-9 contain the VTEC forecast results&nbsp; of Gradient Boosting (XGBoost) trained on three data subsets. Column 10 (&quot;Mean&quot;) represents the mean of columns 1-9, i.e., the ensemble mean; column 11 (&quot;Std&quot;) represents the standard deviation of columns 1-9, i.e., the ensemble spread; columns 12 (&quot;UB&quot;) and 13 (&quot;LB&quot;) contain the upper and lower bounds of the 95% confidence interval of VTEC, respectively; and column 14 contains the&nbsp;Global Ionosphere Maps (GIM) values of CODE, i.e., the ground-truth in this study.</p> <p>The QGB_*_2017.csv file consists of 4 columns. The index column (&quot;Date-time&quot;) is expressed in UTC as YYYY-MM-DD. Column 1 (&quot;Median&quot;) contains the median VTEC forecast, column 2 (&quot;LB&quot;) contains the lower VTEC bound of the 95% confidence interval, column 3 (&quot;UB&quot;) contains the upper VTEC bound of the 95% confidence interval, and column 4 contains the GIM values of CODE, i.e., the ground-truth in this study.</p> <p>The BNN*_2017.csv file consists of 5 columns. The index column (&quot;Date-time&quot;) is expressed in UTC as YYYY-MM-DD. Column 1 (&quot;Mean&quot;) contains the mean VTEC forecast, column 2 (&quot;Std&quot;) contains the standard deviation, column 3 contains GIM values of CODE, i.e., ground-truth in this study; column 4 (&quot;UB&quot;) contains the upper VTEC bound of the 95% confidence interval, and column 5 (&quot;LB&quot;) contains the lower VTEC bound of the 95% confidence interval.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>Contact</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>If you have any questions regarding these data, please contact:</p> <p>Randa Natras</p> <p>Deutsches Geod&auml;tisches Forschungsinstitut (DGFI-TUM)</p> <p>Technical University of Munich</p> <p>Arcisstra&szlig;e 21</p> <p>80333 M&uuml;nchen</p> <p>randa.natras@tum.de</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset

<p><b>Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset.</b></p><p>Ribonucleic acids (RNA) play crucial roles in living organisms as they are involved in key processes necessary for proper cell functioning. Some RNA molecules, such as bacterial ribosomes and precursor messenger RNA, are targets of small molecule drugs, while others, e.g., bacterial riboswitches or viral RNA motifs are considered as potential therapeutic targets. Thus, the continuous discovery of new functional RNA increases the demand for developing compounds targeting them and for methods for analyzing RNA—small molecule interactions. We recently developed fingeRNAt - a software for detecting non-covalent bonds formed within complexes of nucleic acids with different types of ligands. The program detects several non-covalent interactions, such as hydrogen and halogen bonds, ionic, Pi, inorganic ion- and water-mediated, lipophilic interactions, and encodes them as computational-friendly Structural Interaction Fingerprint (SIFt). Here we present the application of SIFts accompanied by machine learning methods for binding prediction of small molecules to RNA targets. We show that SIFt-based models outperform the classic, general-purpose scoring functions in virtual screening. We discuss the aid offered by Explainable Artificial Intelligence in the analysis of the binding prediction models, elucidating the decision-making process, and deciphering molecular recognition processes.</p>

opencc-zeroDec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record