Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

Supplementary material for "Towards a metagenomics machine learning interpretable model for understanding the transition from adenoma to colorectal cancer"

<p>Supplementary files for&nbsp;&quot;Towards a metagenomics machine learning interpretable model for understanding the transition from adenoma to colorectal cancer&quot;.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Evaluation of Machine learning algortihms for classification

<p>The purpose of this report is to compare three different classifiers through supervised machine learning on two diverse datasets. The whole machine learning process was applied and conducted in different experiments. The exploration of the datasets as well as the preprocessing strategies are outlined in the following. Furthermore, the modelling processes and the performance measures on which their results are evaluated will be explained. Finally, different parameter adjustments and settings are compared and discussed which leads to a conclusion.</p>

opencc-byApr 2022View details →
zenodo44/100

A generalized machine learning framework to predict the space-time yield of methanol from thermocatalytic CO2 hydrogenation

<p>Thermocatalytic CO<sub>2</sub> hydrogenation to methanol is an attractive decarbonization technology to combat climate change while producing a valuable platform chemical and energy carrier. However, predicting the performance of catalytic systems for this process remains a challenge. Herein, we present a machine learning framework to predict catalyst performance from experimental descriptors. A database of Cu-, Pd-, In<sub>2</sub>O<sub>3</sub>-, and ZnO-ZrO<sub>2</sub>-based catalysts with 1425 datapoints is compiled from literature and subjected to data mining. Accurate ensemble-tree models (<em>R</em><sup>2</sup> &gt; 0.85) are developed to predict the methanol space-time yield (<em>STY</em>) from 12 descriptors, where the significance of space velocity, pressure, and metal content is revealed. The model prediction and its insights are experimentally validated, with a root mean squared error of 0.11&nbsp;g<sub>MeOH</sub>&nbsp;h<sup>&minus;1</sup>&nbsp;g<sub>cat</sub><sup>&minus;1 </sup>between the actual and predicted methanol<em> STY</em>. The framework is purely data-driven, interpretable, cross-deployable to other catalytic processes, and serves as an invaluable tool for guided experiments and optimization.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Gaia EDR3 Catalogs of Machine-Learned Radial Velocities

<p><strong>Gaia EDR3 Catalogs of Machine-Learned Radial Velocities</strong></p> <p>Spatially complete&nbsp;Test-Set and Machine-Learned Radial Velocity (ML-RV)&nbsp;Catalogs described in Dropulic et al., arXiv:<a href="https://arxiv.org/abs/2205.12278">2205.12278</a>. The spatially complete&nbsp;Test-Set Catalog contains a total of 4,332,657&nbsp;stars, while the &nbsp;spatially complete ML-RV Catalog contains 91,840,346 stars. We provide Gaia EDR3 Source IDs, the network-predicted line-of-sight velocity in km/s, and the network-predicted uncertainty in km/s.&nbsp;</p> <p>We have included a simple Jupyter notebook demonstrating how to import the data, and make a simple histogram with it.</p> <p>If you find this catalog useful in your work, please cite Dropulic et al.&nbsp;arXiv:<a href="https://arxiv.org/abs/2205.12278">2205.12278</a>, as well as Dropulic et al. <a href="https://doi.org/10.3847/2041-8213/ac09ef">ApJL 915, L14 (2021)</a>&nbsp;arXiv:<a href="https://arxiv.org/abs/2103.14039">2103.14039</a>.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Multi-omic machine learning predictor of breast cancer therapy response

<p>H&amp;E slides used in the training dataset described in&nbsp;&quot;Multi-omic machine learning predictor of breast cancer therapy response&quot;&nbsp;published in&nbsp;<em>Nature</em>:&nbsp;<a href="https://www.nature.com/articles/s41586-021-04278-5">https://www.nature.com/articles/s41586-021-04278-5</a></p> <p>Metadata associated with these images also included in file Slide metadata.xlsx</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Reflectometry curves (XRR and NR) and corresponding fits for machine learning

<p>This is a compiled dataset of raw X-ray reflectivity (XRR, reflectometry) measurements together with corresponding fit parameters, intentionally published to use as training or test data for machine learning models. (The authors aim to include NR data in further versions of this dataset and plan to include other substrates and materials for XRR. Contributions welcome!)</p> <p><br> <strong>An interactive documentation can be found in <em>&quot;README.html&quot;</em> or at <a href="https://schreiber-lab.github.io/reflectometry-dataset">https://schreiber-lab.github.io/reflectometry-dataset</a>.</strong></p> <ul> <li>Data structure</li> </ul> <p>All data is provided in an hdf5 file, following <a href="https://www.nexusformat.org/">NeXus</a> convention with respect to the provided metadata in the hdf5 attributes. Some datesets have been measured in-situ and therefore there are stacks of curves that correspond to the different layer thicknesses of the same material on top of SiOx. The measured data is provided under experimental and the corresponding fit parameters under fit. Additional information is collected in metadata.</p> <ul> <li>Where to find the dataset and how to contribute</li> </ul> <p>Have a look at <a href="https://github.com/schreiber-lab/reflectometry-dataset">github</a> and <a href="https://doi.org/10.5281/zenodo.6497438">zenodo</a>. In case you wish to contribute further curves to this dataset or have ideas how to improve the dataset or where else to deposit it, please contact the authors at softmatter AT ifap.uni-tuebingen.de.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Machine learning classifiers for species classification of fungi using error-prone long-reads on extended metabarcodes

<p>Machine learning models used in the decision tree of linked machine learning models (<a href="https://github.com/teenjes/fungal_ML">https://github.com/teenjes/fungal_ML</a>)</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Tomato Classification using Mass Spectrometry-Machine Learning Technique: a Food Safety-enhancing Platform

<p>Food safety and quality assessment mechanisms are unmet needs that industries and countries have been continuously facing in recent years. Our study aimed at developing a platform using Machine Learning algorithms to analyze Mass Spectrometry data for classification of tomatoes on organic and non-organic. Tomato samples were analyzed using silica gel plates and direct-infusion electrospray-ionization mass spectrometry technique. Decision Tree algorithm was tailored for data analysis. This model achieved 92% accuracy, 94% sensitivity and 90% precision in determining to which group each fruit belonged. Potential biomarkers evidenced differences in treatment and production for each group.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Surrogate Machine Learning for Parmec Advanced Gas-cooled Reactor (AGR) Analysis

<p>Data associated with research towards a surrogate machine learning model for the Advanced Gas-cooled Reactor (AGR). This data was generated using the Parmec software package [<a href="https://parmes.org/parmec/index.html">1</a>] and can be used to train machine learning models using the Surrogate Machine Optimisation and Learning (SMOL) framework [<a href="https://gitlab.cs.man.ac.uk/q59494hj/parmec_agr_ml_surrogate">2</a>]. Visit the aforementioned repository, clone the code, then download the files into repository folder.</p> <p>If you are not planning on working with data augmentation, exclude the files with flip and rotate in the title, e.g.&nbsp;dataset__flip_13_rotate_123_cases.pkl.<br> <br> 1. Koziara, T., 2019. Parmec documentation. URL: <a href="https://parmes.org/parmec/index.html">https://parmes.org/parmec/index.html</a>&nbsp;[Online; accessed 05-August-2022].</p> <p>2. Github Repository. URL:&nbsp;<a href="https://gitlab.cs.man.ac.uk/q59494hj/parmec_agr_ml_surrogate">https://gitlab.cs.man.ac.uk/q59494hj/parmec_agr_ml_surrogate</a>&nbsp;[Online; accessed 05-August-2022].</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Machine Learning Snowfall LPM

<p>This dataset contains a 65-year snowfall climatology for stations in Lower Michigan Penisular and associated environment variables.</p> <p>We also have several R codes examples for the machine learning model training.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Machine Learning for Software Engineering: A Tertiary Study

<p>Dataset of the research paper:&nbsp;<strong>Machine Learning for Software Engineering: A Tertiary Study</strong></p> <p>Machine learning (ML) techniques increase the effectiveness of software engineering (SE) lifecycle activities. We systematically collected, quality-assessed, summarized, and categorized 83 reviews in ML for SE published between 2009&ndash;2022, covering 6,117 primary&nbsp;studies. The SE areas most tackled with ML are software quality and testing, while human-centered areas appear more challenging for ML. We propose a number of ML for SE research challenges and actions including: conducting further empirical validation and industrial studies on ML; reconsidering deficient SE methods; documenting and automating data collection and pipeline processes; reexamining how industrial practitioners distribute their proprietary data; and implementing incremental ML approaches.</p> <p>The following data and source&nbsp;files&nbsp;are included.</p> <ul> <li><strong>review-protocol.md</strong>: The protocol employed in this tertiary study</li> </ul> <p><strong>data/</strong></p> <p><strong>&nbsp; dl-search/</strong></p> <p><strong>&nbsp; &nbsp; input/</strong></p> <ul> <li><strong>acm_comput_surveys_overviews.bib</strong>: Surveys of ACM Computing Surveys journal</li> <li><strong>acm_comput_surveys_overviews_titles.txt</strong>: Titles of surveys</li> <li><strong>acm_comput_ml_surveys.bib</strong>: Machine learning (ML)-related surveys of ACM Computing Surveys journal</li> <li><strong>acm_comput_ml_surveys_titles.txt</strong>: Titles of ML-related surveys</li> <li><strong>dl_search_queries.txt</strong>: Search queries applied to IEEE Xplore, ACM Digital Library, and Elsevier Scopus</li> <li><strong>ml_keywords.txt</strong>: ML-related keywords extracted from ML-related survey titles and used in the search queries</li> <li><strong>se_keywords.txt</strong>: Software Engineering (SE)-related keywords derived from the 15 SWEBOK Knowledge Areas (KAs&mdash;except for Computing Foundations, Mathematical Foundations, and Engineering Foundations) and used in the search queries</li> <li><strong>secondary_studies_keywords.txt</strong>: Survey-related keywords composed of the 15 keywords introduced in the tertiary study on SLRs in SE by Kitchenham <em>et al.</em> (2010), and the survey titles, and used in the search queries</li> </ul> <p><strong>&nbsp; &nbsp; output/</strong></p> <ul> <li><strong>acm/</strong> <ul> <li><strong>acm{1&ndash;9}.bib</strong>: Search results from ACM Digital Library</li> </ul> </li> <li><strong>ieee.csv</strong>: Search results from IEEE Xplore</li> <li><strong>scopus_analyze_year.csv</strong>: Yearly distribution of ML and SE documents extracted from Scopus&#39;s <em>Analyze search results</em> page</li> <li><strong>scopus.csv</strong>: Search results from Scopus</li> </ul> <p><strong>&nbsp; study-selection/</strong></p> <ul> <li><strong>backward_snowballing.csv</strong>: Additional secondary studies found through the backward snowballing process</li> <li><strong>backward_snowballing_references.csv</strong>: References of quality-accepted secondary studies</li> <li><strong>cohen_kappa_agreement.csv</strong>: Inter-rater reliability of reviewers in study selection</li> <li><strong>dl_search_results.csv</strong>: Aggregated search results of all three digital libraries</li> <li><strong>forward_snowballing_reviewer_{1,2}.csv</strong>: Divided forward snowballing citations of quality-accepted studies assessed by reviewer 1 and 2, correspondingly, based on IC/EC</li> <li><strong>study_selection_reviewer_{1,2}.csv</strong>: Divided search results assessed by reviewer 1 and 2, correspondingly, based on IC/EC</li> </ul> <p><strong>&nbsp; quality-assessment/</strong></p> <ul> <li><strong>dare_assessment.csv</strong>: Quality assessment (QA) of selected secondary studies based on the Database of Abstracts of Reviews of Effects (DARE) criteria by York University, Centre for Reviews and Dissemination</li> <li><strong>quality_accepted_studies.csv</strong>: Details of quality-accepted studies</li> <li><strong>studies_for_review.bib</strong>: Bibliography details and QA scores of selected secondary studies</li> </ul> <p><strong>&nbsp; data-extraction/</strong></p> <ul> <li><strong>further_research.csv</strong>: Recommendations for further research of quality-accepted studies</li> <li><strong>further_research_general.csv</strong>: The complete list of associated studies for each general recommendation</li> <li><strong>knowledge_areas.csv</strong>: Classification of quality-accepted studies using the SWEBOK KAs and subareas</li> <li><strong>ml_techniques.csv</strong>: Classification of the quality-accepted studies based on a four-axis ML classification scheme, along with extracted ML techniques employed in the studies</li> <li><strong>primary_studies.csv</strong>: Details of reviewed primary studies by the quality-accepted secondary</li> <li><strong>research_methods.csv</strong>: Citations of the research methods employed by the quality-accepted studies</li> <li><strong>research_types_methods.csv</strong>: Research types and methods employed by the quality-accepted studies</li> </ul> <p><strong>src/</strong></p> <ul> <li><strong>data-analysis.ipynb</strong>: Analysis of data extraction results (data preprocessing, top authors and institutions, study types, yearly distribution of publishers, QA scores, and SWEBOK KAs) and creation of all figures included in the study</li> <li><strong>scopus-year-analysis.ipynb</strong>: Yearly distribution of ML and SE publications retrieved from Elsevier Scopus</li> <li><strong>study-selection-preprocessing.ipynb</strong>: Processing of digital library search results to conduct the inter-rater reliability estimation and study selection process</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Assessing the Influence of Zeolite Composition on Oxygen-Bridged Diamino Dicopper(II) Complexes in Cu-CHA DeNOx Catalysts by Machine Learning-Assisted X‑ray Absorption Spectroscopy

<ul> <li><strong>Data type</strong>: Experimental spectroscopic measurements and related elaboration from Figures 1-4 of the corresponding article</li> <li>Files are with filename extensions: <strong>txt</strong></li> <li>Information on <strong>origin of the data</strong>:</li> </ul> <p>In situ XANES and EXAFS data were collected at the BM23 beamline of the European Synchrotron Radiation Facility (ESRF, Grenoble, France) in a Microtomo reactor cell; measured Cu-CHA samples are indicated in the following with &ldquo;Cu/Al&rdquo;-&ldquo;Si/Al&rdquo; labels</p> <ul> <li><strong>fig_01_XANES:</strong> Normalized Cu K-edge XANES for Cu-CHA samples 0.1-5; 0.5-15; 0.6-29, collected at 200 &deg;C after pretreatment in O<sub>2</sub>, reduction in NO+NH<sub>3</sub> and subsequent oxidation in O<sub>2</sub>.</li> <li><strong>fig_02_Conversion:</strong> NOx conversion in the 150&minus;500 &deg;C temperature range for Cu-CHA samples 0.1-5, 0.5-15, 0.6-29; TOF at 200 &deg;C versus fraction of Cu(I) from XANES LCF after oxidation and fraction of Cu(I) from XANES LCF after oxidation versus Cu density for the same catalysts.</li> <li><strong>fig_03_EXAFS_FT_WT:</strong> Magnitude of experimental EXAFS spectra, obtained by Fourier transforming k<sup>2</sup>&chi;(k) spectra in the 2.4&minus;12.0 &Aring;<sup>&minus;1</sup> range for Cu-CHA samples 0.1-5, 0.5-15, 0.6-29 after reduction in NO+NH<sub>3</sub> and subsequent oxidation in O<sub>2</sub>; corresponding EXAFS WT maps magnified in high-R range (2-4 &Aring;), obtained using a Morlet WT with parameters (&sigma;=1, &eta;=7).</li> <li><strong>fig_04_EXAFS_MLfit:</strong> Magnitude of experimental and best fit EXAFS spectra, obtained by Fourier transforming k<sup>2</sup>&chi;(k) spectra in the 2.4&minus;12.0 &Aring;<sup>&minus;1</sup> range for Cu-CHA samples 0.1-5, 0.5-15, 0.6-29 after oxidation in O<sub>2</sub>. Scaled components 1 ([Cu<sup>I</sup>(NH<sub>3</sub>)<sup>2</sup>]<sup>+</sup>), 2 and 3 (planar and bent &mu;-&eta;<sup>2</sup>,&eta;<sup>2</sup>-peroxo diamino dicopper(II)) isolated by ML-assisted EXAFS fitting are also reported, vertically translated.</li> <li><strong>Information on</strong>:</li> <li>specialized abbreviations: <strong>CHA</strong>&ndash; chabazite; <strong>XANES</strong>&ndash; X-ray absorption near edge structure, <strong>EXAFS</strong> &ndash; Extended X-ray absorption fine structure; <strong>LCF</strong> &ndash; Linear Combination Fit;<strong> FT</strong>: Fourier Transform; <strong>WT</strong> &ndash; Wavelet Transform; <strong>ML</strong> &ndash; Machine Learning; <strong>TOF</strong> &ndash; Turn Over Frequency;</li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Dataset for: Application of Machine Learning for the Spatial Analysis of Binaural Room Impulse Responses

<p>This repository contains supplementary material for the paper titled `Application of Machine Learning for the Spatial<br> Analysis of Binaural Room Impulse Responses&#39; Available at: <a href="http://dx.doi.org/10.3390/app8010105">dx.doi.org/10.3390/app8010105</a>&nbsp;. These programs and audio files are distributed in the hopes that they will prove useful under the Creative Commons Attribution 4.0, with no warranty; or the implied warranty of merchantability or fitness for a particular problem. Please give appropriate credit for use of the material provided in this repository back to the author.&nbsp;</p> <p>In order to use the MatLab code the Auditory Toolbox by Malcolm Slaney [1] and the Cochleagram function distributed by Bin Gao [2] are required.</p> <p>The python scrips require the following Python libraries to be installed: Numpy[3], SciPy[4] and Tensorflow [5].</p> <p>The MatLab code was tested using MatLab R2017a on a Computer running windows 7.</p> <p>The python code was tested using Python 3.2.5, using an anaconda Python environment - in windows command line.</p> <p>--</p> <p>The repository contains:</p> <p>Folders:</p> <p><br> 1.) neg90 - This folder contains the gaussian normalisation parameters stored as text files and the weights and biases for the trained neural network - these are all for the -90&deg; rotation neural network.</p> <p>2.) pos90 - This folder contains the gaussian normalisation parameters stored as text files and the weights and biases for the trained neural network - these are all for the +90&deg; rotation neural network.</p> <p>3.) testData - this folder contains pre-generated test data for the different binaural dummy head microphones, speaker, and signal type combinations.</p> <p>Python Scripts:</p> <p><br> 1.) AnalyseDoA.py - A python script that can be run to test the neural network using the pre-generated test data - running the script will allow the user to input the binaural dummy head, speaker, and signal type. The important variables generated by this script are DoA - the direction of arrival for each signal in the feature vector, and yDiff - the difference between the predicted DoA and the expected direction of arrival</p> <p>2.) DirectionAnalysis.py - This python file contains a set of function that are used to define the neural network, and run it. The function called DoAPrediction takes the feature vector generated by the MatLab code as its input argument, these features will then be passed to the neural network, and the output of this function is the direction of arrival predicted by the neural network for each signal. The functions: DoAAnalysis_neg90 and DoAAnalysis_pos90 are called by the DoAPrediction function, these functions create the neural network using the NN function, import the weights and biases, and passes the feature matrix (provided as input) through the neural network - the output of these functions are the predicted direction of arrival.</p> <p>MatLab files:</p> <p><br> 1.) runAnalysis.m - This&nbsp;MatLab script&nbsp;analyses the dataset provided as part of this repository. Users can change the variables head (&#39;KEMAR&#39; or &#39;KU100&#39;), signalType (&#39;directSound&#39; or &#39;reflection&#39;), and speaker (&#39;EquatorD5&#39; or &#39;Genelec8030&#39;). This script will produce the gaussian normalised feature vector and expected direction of arrival for all signals with the defined head, signal type, and speaker combination. These variables are then saved in .mat files so they can be imported by the python scripts.</p> <p>2.) BinauralModelCochlea.m - This MatLab function analyses a given binaural signal and outputs the interaural cross-correlation, interaural level difference, interaural time difference, the cochlea output for the left and right channel and the centre frequencies of the gammatone filter band. The input variables are: IR - the signal to be analysed, N - the number of gammatone filters, freqLow - the lowest centre frequency of the gammatone filter bank (centre frequency of the first gammatone filter), and freqHigh - the highest centre frequency of the gammatone filter bank (the centre frequency of the Nth gammatone filter). This function requires Malcolm Slaney&#39;s Auditory Toolbox [1] and Bin Gao&#39;s Cochleagram function [2] in order to work.</p> <p>3.) generateFeatureVector.m - This MatLab function generates a feature vector from an input binaural signal x, and a version of the signal captured after the binaural dummy head has been rotated by either +90&deg; or -90&deg; degree (variables xPos90 and xNeg90 respectively). If the sampling frequency (Fs) isn&#39;t 44100, the signals are resampled to be at 44100. This file also contains a function &#39;gaussianNormalisationTestData&#39; which gaussian normalises the data using the mean and standard deviation calculated from the data used to train the neural networks - the mean and standard deviation values are stored in the folder GMParams in the pos90 and neg90 folders.</p> <p>4.) generateTestData.m - This&nbsp;MatLab function analyses the included binaural dataset, it takes the input variables: head - the binaural dummy head used for the measurements either &#39;KEMAR&#39; or &#39;KU100&#39;, speaker - the speaker used for the measurements either &#39;EquatorD5&#39; or &#39;Genelec8030&#39;, and signalType - the type of signal being analysed either &#39;directSound&#39; or &#39;reflection&#39;.</p> <p>Text files:</p> <p><br> 1.) noLayers.txt - a text file containing the number of layers used when training the neural network - with the current version of the code the neural network contains only 1 layer.</p> <p>2.) README.txt - Read me file containing information about the repository.</p> <p>Audio files:</p> <p><br> This repository contains 1152 binaural signals half of which are direct sounds segmented from a binaural room impulse responses and the other half are reflections segmented from binaural room impulse responses (detailed in the paper this material supports) the direct sounds are recorded at angles from 0&deg; to 357.5&deg; in steps of 2.5&deg; and the reflections are recorded at angles of 1&deg; to 358.5&deg; in steps of 2.5&deg;. In the paper only recordings relating to signals recorded with the Equator D5 are analysed.</p> <p>The combination of audio files include:</p> <p>1.) 144 direct sound recordings captured with the KEMAR 45BC binaural dummy head microphone and the Equator D5 speaker<br> 2.) 144 reflection recordings captured with the KEMAR 45BC binaural dummy head microphone and the Equator D5 speaker<br> 3.) 144 direct sound recordings captured with the KU100 binaural dummy head microphone and the Equator D5 speaker<br> 4.) 144 reflection recordings captured with the KU100 binaural dummy head microphone and the Equator D5 speaker<br> 5.) 144 direct sound recordings captured with the KEMAR 45BC binaural dummy head microphone and the Genelec 8030 speaker<br> 6.) 144 reflection recordings captured with the KEMAR 45BC binaural dummy head microphone and the Genelec 8030 speaker<br> 7.) 144 direct sound recordings captured with the KU100 binaural dummy head microphone and the Genelec 8030 speaker<br> 8.) 144 reflection recordings captured with the KU100 binaural dummy head microphone and the Genelec 8030 speaker</p> <p>The files are stored using the following file naming convention:<br> head_Test3_speaker_signalType_000_0_Degrees.wav - where _000_0 defines the azimuth direction of arrival so for example for a direct sound measured with the KEMAR unit and the Genelec8030 at 5 degrees would be &#39;KEMAR_Test3_Genelec8030_directSound_005_0Degrees.wav&#39; and for a reflection measured with the KU100 and the Equator D5 at 298.5 degrees would be &#39;KU100_Test3_EquatorD5_reflection_298_5Degrees.wav&#39;</p> <p>--</p> <p>Bibliography:<br> [1]&nbsp;Slaney, M. (1998). Auditory Toolbox. Palo Alto, CA. [Online]. Available: https://engineering.purdue.edu/~malcolm/interval/1998-010/ [Accessed: Oct. 27, 2017]</p> <p>[2]&nbsp;Gao, B. (2014). Cochleagram and IS-NMF2D for Blind Source Separation. [Online] Available:&nbsp;http://uk.mathworks.com/matlabcentral/fileexchange/48622-cochleagram-and-is-nmf2d-for-blind-source-separation?focused=3855900&amp;tab=function&nbsp;[Accessed: Oct. 27, 2017]</p> <p>[3]&nbsp;NumFocus. (n.d.). NumPy. [Online]. Available: http://www.numpy.org/ [Accessed: Oct. 27, 2017]</p> <p>[4]&nbsp;SciPy. (n.d.). SciPy. [Online]. Available:&nbsp;https://www.scipy.org/&nbsp;[Accessed: Oct. 27, 2017]</p> <p>[5]&nbsp;Google. (n.d.). TensorFlow. [Online] Available:&nbsp;https://www.tensorflow.org/&nbsp;[Accessed: Oct. 27, 2017]</p> <p>--</p> <p>All code and audio produced by: Michael Lovedee-Turner, PhD candidate in Music Technology at the Audio Lab, Department of Electronic Engineering, University of York</p> <p>Contact: mjlt500@york.ac.uk</p>

opencc-by-4.0Oct 2017View details →
zenodo44/100

Experiment on the performance of different machine learning algorithms for classification - Results

<h2>Results of a short performance study of machine learning algorithms</h2> <h3>Context and methodology</h3> <ul> <li>This data was produced while performing a university project to examine the performance of various machine learning algorithms on different prediction datasets</li> <li>The data serves the purpose of comparing the metrics of performing the different tasks</li> <li>The dataset contains a number of matrices for every classifier and every dataset</li> <li>The data was produced with python scripts provided further down and with the usage of the external datasets: <ul> <li>Membership Woes Dataset (OpenML): <a href="https://api.openml.org/d/44224">https://api.openml.org/d/44224</a></li> <li>Zoo dataset (UCI): <a href="https://doi.org/10.24432/C5R59V">https://doi.org/10.24432/C5R59V</a></li> <li>Breast Cancer Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> <li>Loan Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> </ul> </li> </ul> <h3>Technical details</h3> <ul> <li>The data consists of one JSON file</li> <li>The source code for producing this data is available at&nbsp;<a href="https://doi.org/10.5281/zenodo.11085222">https://doi.org/10.5281/zenodo.11085222</a></li> </ul> <h3>Structure of the data</h3> <p>[ {"classifier": ...,<br>"dataset": ...,<br>"hyper_parameters": ...,<br>"cross_validation_results": {<br>&nbsp; &nbsp; "fit_time": {} ,<br>&nbsp; &nbsp; "score_time": ...,<br>&nbsp; &nbsp; "metrics": {}<br>},&nbsp;<br>"holdout_test_results": ...}, ]</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

DFT Calculated xyz and log Files as well as csv Files for Machine Learning in Support of "Tailoring Phosphine Ligands for Improved C H Activation: Insights from Δ-Machine Learning"

<p>Transition metal complexes have played crucial roles in various homogeneous catalytic processes due to their exceptional versatility. This adaptability stems not only from the central metal ions but also from the vast array of choices of the ligand spheres, which form an enormously large chemical space. For example, Rh complexes, with a well-designed ligand sphere, are known to be efficient in catalyzing the C-H activation process in alkanes. To investigate the structure-property relation of the Rh complex and identify the optimal ligand that minimizes the calculated reaction energy &Delta;E&nbsp;of an alkane C-H activation, we have applied a &Delta;-Machine Learning method trained on various features to study 1,743 pairs of reactants (Rh(PLP)(Cl)(CO)) and intermediates (Rh(PLP)(Cl)(CO)(H)(propyl)). Our findings demonstrate that the models exhibit robust predictive performance when trained on features derived from electron density (R<sup>2 </sup>= 0.816), and SOAPs (R<sup>2 </sup>= 0.819), a set of position-based descriptors. Leveraging the model trained on xTB-SOAPs that only depend on the xTB-equilibrium structures, we propose an efficient and accurate screening procedure to explore the extensive chemical space of bisphosphine ligands. By applying this screening procedure, <a>we identify ten newly selected reactant-intermediate pairs with an average &Delta;E&nbsp;</a>of 33.2 kJ mol<sup>-1</sup>, remarkably lower than the average &Delta;E of the original data set of 68.0 kJ mol<sup>-1</sup>. This underscores the efficacy of our screening procedure in pinpointing structures with significantly lower energy levels.</p> <p>_______________________________________________________________________</p> <p>The dataset contains three file types:</p> <p>Version 1.0:</p> <ol> <li>xyz files of the final optimized Rh-phosphine complexes; one set for the starting materials denoted as "molecule-XXXX_4-times" and one set for the intermediates after C-H activation denoted as "molecule-XXXX_6-times"</li> <li>Gaussian16 log files for the optimization process; one set for the starting materials denoted as "molecule-XXXX_4-times" and one set for the intermediates after C-H activation denoted as "molecule-XXXX_6-times"</li> <li>csv files containing the per molecule features used for training the different machine learning models. The name of the csv files indicates which property was predicted and which model was used</li> </ol> <p>New in version 1.1 (other data is unchanged):</p> <ol> <li>Gaussian16 log files for the ten newly identified bisphosphine ligands; one set for the product material denoted as "LXX_6-times-axial" and one set for the transition state for the C-H activation denoted as "LXX_C-H-activation_TS"</li> </ol>

opencc-by-4.0Jan 2024View details →
zenodo44/100

ΔvapHm-VOC: Standard Molar Vaporization Enthalpy Database for Machine Learning Prediction Models

<p>We present the full database of the article "Data-Driven, Explainable Machine Learning Model for Predicting Volatile Organic Compounds&rsquo; Standard Vaporization Enthalpy".</p> <p>This is the database used for data driven, explainable supervised ML model to predict &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; of VOCs. The model was built on an established experimental database of 2410 unique molecules and 223 VOCs categorized by chemical groups. Using supervised ML regression algorithms, the Random Forest successfully predicted VOCs&rsquo; &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; with a mean absolute error of 3.02 kJ mol<sup>-1</sup> and a 94% test score. The model was successfully validated through the prediction of &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; for a known database of VOCs and through molecular group hold-out tests.</p> <div> <div> <div> <div> <p>The model's database was built with a variety of molecules from diverse chemical families with known experimental &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; values. Entries were collected from <a href="https://doi.org/10.1063/1.3309507" target="_blank" rel="noopener">Acree and Chickos&rsquo; 2010 compilation</a>, curated by <a href="https://doi.org/10.1016/j.fluid.2013.09.021" target="_blank" rel="noopener">Gharagheizi (2013)</a>, with experimental vaporization enthalpy at the standard temperature of 298.15 K. This database was selected as it is an open-access repository, generally presenting experimental values with low uncertainties and corrected for the real-to-ideal behavior of the gas phase. We introduced a routine to convert and present each chemical entry into a SMILES string, along with chemical family categorization. For VOCs, we built a specific database of compounds documented in a VOC regulatory environmental guideline (<a href="https://www.s-t-a.org/Files%20Public%20Area/Documents/The%20Categorisation%20of%20Volatile%20Organic%20Compounds%20HMIP%20(1996).pdf">Marlowe <em>et al</em>., 1995</a>), and we used our web-scrapping routine to gather experimental &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; values. The external dataset for validation studies was also collected from <a href="https://doi.org/10.1016/j.fluid.2013.09.021" target="_blank" rel="noopener">Gharagheizi (2013)</a>.</p> <p>Along with &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; experimental values, each molecule is represented by its CAS number, SMILES string and InChlKey. We generated 106 chemical descriptors for every molecule in the database, using <a href="http://http//www.rdkit.org/">RDKit</a> software version 2022.09.4, running on top of Python 3.9. Descriptors were calculated from the &ldquo;MolFromSmiles&rdquo; function in &ldquo;RDKIT.Chem&rdquo; as descriptors with non-numerical values were removed. The descriptors encode significant chemical information and are used to present physicochemical characteristics of compounds, building a relationship between structure and &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg;.</p> </div> </div> </div> </div> <p>Through chemical feature importance analysis, the explainable model revealed that VOC polarizability, connectivity indexes and electrotopological state are key for the model&rsquo;s prediction accuracy. We thus present a replicable and explainable model, which can be further expanded towards the prediction of other thermodynamic properties of VOCs.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Results: Towards Realistic SATD Identification Through Machine Learning Models: Ongoing Research and Preliminary Results

<p>Automated identification of self-admitted technical debt (SATD) has been crucial for advancements in managing such debt.&nbsp;<br>However, state-of-the-arts studies often overlook chronological factors, leading to experiments that do not faithfully replicate the conditions developers face in their daily routines.<br>This study initiates a chronological analysis of SATD identification through machine learning models, emphasizing the significance of temporal factors in automated SATD detection.&nbsp;<br>The research is in its preliminary phase, divided into two stages: evaluating model performance trained on historical data and tested in prospective contexts, and examining model generalization across various projects. Preliminary results reveal that the chronological factor can positively or negatively influence model performance and that some models are not sufficiently general when trained and tested on different projects.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Quantifying the basic reproduction number and the under-estimated fraction of mpox cases around the world at the onset of the outbreak: a mathematical modeling and machine learning- based study

<p><span>In 2022, there was a global resurgence of mpox, with different clinico-epidemiological features compared with previous</span><br><span>outbreaks. During this resurgence, sexual contact was hypothesized as the primary transmission route, with the community</span><br><span>of men having sex with men (MSM) being disproportionately affected. Because of the stigma associated with sexually</span><br><span>transmitted infections, especially those impacting MSM, the real burden of mpox could be masked.</span><br><span>We quantified the basic reproduction number (R</span><span>0</span><span>) and the under-estimated fraction of mpox cases in 16 countries, from the</span><br><span>onset of the outbreak until early September 2022, using Bayesian inference and a compartmentalized, risk-structured (high-</span><br><span>and low-risk populations), two-route (sexual and non-sexual transmission) mathematical model. Machine learning (ML) was</span><br><span>leveraged to identify under-estimation determinants.</span><br><span>Estimated R</span><span>0</span><span> </span><span>ranged between 1&middot;37 (Canada) and 3&middot;68 (Germany). The under-estimation rates for the high- and low-risk</span><br><span>populations varied between 25-93% and 65-85%, respectively. The estimated total number of mpox cases, relative to the</span><br><span>reported cases, is highest in Colombia (3&middot;60) and lowest in Canada (1&middot;08). In the ML analysis, two clusters of countries could</span><br><span>be identified, differing in terms of attitudes towards the 2SLGBTQIAP+ community and importance of religion.</span><br><span>Given the substantial mpox under-estimation, surveillance should be enhanced and campaigns against the stigmatization of</span><br><span>MSM should be organized. Countries have different social characteristics, potentially explaining the various degrees of under-</span><br><span>reporting in mpox cases, which should be considered by studies assessing the effectiveness of community-based</span><br><span>interventions.</span></p>

opencc-by-4.0May 2024View details →
zenodo44/100

Solar Asset Mapper: A continuously-updated global inventory of solar energy facilities built with satellite data and machine learning

<p><strong>TransitionZero&rsquo;s Solar Asset Mapper is a global, satellite-derived dataset of utility-scale solar farms generated with a combination of machine learning and human annotation. Our Q1 2024 dataset contains the location and shape of 63,616 assets, along with estimated capacities. We estimate the construction date for over 80% of these assets. The dataset contains over 19,100 square kilometres of solar farms across 183 countries, with a total estimated capacity of 711 GW.</strong></p> <p>Download the dataset, read the explainer and explore our polygon browser UI at&nbsp;<a href="https://www.transitionzero.org/products/solar-asset-mapper" target="_blank" rel="noopener">TransitionZero.org.</a></p> <p><a href="https://blog.transitionzero.org/hubfs/Data%20Products/TZ-SAM/tz-sam-scientific-methodology-Q12024.pdf" target="_blank" rel="noopener">Download our methodology paper here&nbsp;</a></p> <h1><strong>1. Dataset Description</strong></h1> <p>We publish six files.</p> <ul> <li><em>analysis_polygons.gpkg:</em> our &ldquo;analysis-ready&rdquo; dataset containing geometries, capacity estimates and construction date estimates.</li> <li><em>analysis_polygons.csv:</em> a version of analysis_polygons.gpkg containing a central latitude and longitude in place of a geometry, to allow parsing without geospatial software.</li> <li><em>sources.csv</em>: a table mapping the IDs of our analysis-ready dataset to the raw geometries that make them up.</li> <li><em>raw_polygons.gpkg:</em> the raw geometries used to compose analysis_polygons.gpkg.</li> <li><em>TZ Solar Asset Mapper Q1 2024.xlsx</em>: an Excel formatted version of the analysis_polygons.csv file.</li> <li><em>tz-sam_scientific_data.pdf</em>: A pre-print aricle that explains the methodology in detail.</li> </ul> <h2><strong>1.1 Analysis-level datasets</strong></h2> <p>Our analysis-level dataset comprises our most complete view of global asset-level solar installations, incorporating our own detections as well as known solar farm geometries from other datasets.</p> <p>The geospatial dataset contains the following fields:</p> <ul> <li>id: unique ID for the asset</li> <li>geometry: Polygon or MultiPolygon defining the asset</li> <li>capacity_mw: estimated capacity of the asset in megawatts</li> <li>constructed_before: upper bound for construction date (estimated date of the image in which the solar plant was first seen in a constructed state)</li> <li>constructed_after: lower bound for construction date (estimated date of the image in which construction began for the solar plant)</li> </ul> <p>The CSV version replaces the Geometry column with:</p> <ul> <li>latitude: the latitude of the centroid of the asset</li> <li>longitude: the longitude of the centroid of the asset</li> <li>country: administrative country name</li> </ul> <h2><strong>1.2 Raw datasets and sources</strong></h2> <p>The analysis-level datasets hide some complexity in the underlying data that we expose in the <em>raw_polygons</em> and <em>sources</em> file.</p> <ul> <li>We produce new sets of polygons for each run. Often these overlap, sometimes in complicated ways.</li> <li>We cluster together overlapping and nearby geometries from both our detections and external sources. Currently these sources are:</li> <li>Large solar farms scraped from OpenStreetMap (OSM)</li> <li>Validated geometries from <a href="../records/5005868">Kruitwagen et. al., A global inventory of solar photovoltaic generating units</a>.</li> </ul> <p>Each cluster comprises one row in the analysis-level dataset. In order to enable tracking raw detections from run to run, as well as to provide detailed sourcing information, we provide all of these raw polygons, along with a source file that lists all of the raw polygons contained in each analysis-level polygon.</p> <p>raw_polygons.gpkg contains the following fields:</p> <ul> <li>id: ID of the raw source polygon</li> <li>geometry<strong>:&nbsp;</strong>Polygon or MultiPolygon defining the asset</li> <li>source: either &ldquo;solar asset mapper&rdquo;, &ldquo;osm&rdquo; or &ldquo;2019_global_pv&rdquo;.</li> <li>acquisition_date: for solar asset mapper polygons, this is the date of the inference run that produced the polygon; for OSM polygons it is the date that the polygon was scraped from OSM; for 2019_global_pv it is 2019-01-01, the approximate detection date of that dataset.</li> </ul> <p>Sources.csv contains the following fields:</p> <ul> <li>cluster_id: ID of the corresponding item in the analysis-level dataset</li> <li>source_id: ID of the raw source polygon</li> <li>source: either &ldquo;solar asset mapper&rdquo;, &ldquo;osm&rdquo; or &ldquo;2019_global_pv&rdquo;.</li> <li>acquisition_date: for solar asset mapper polygons, this is the date of the inference run that produced the polygon; for OSM polygons it is the date that the polygon was scraped from OSM; for 2019_global_pv it is 2019-01-01, the approximate detection date of that dataset.</li> </ul> <h2><strong>1.3 Caveats and limitations</strong></h2> <h3><strong>1.3.1 Capacity Estimates</strong></h3> <p>While we have made every effort to remove false positives from the published dataset, some will remain due to the difficulty of manually validating detections in 10-metre satellite imagery. To estimate false positive prevalence throughout the data a subset of approximately 2000 detections were selected at random from our positively labelled solar assets. Each of these were validated through a higher degree of scrutiny utilising high-resolution imagery. This analysis yielded an expected rate of false positives of around 1%.</p> <h3>1.3.2 Plant Shapes</h3> <p>Our plant outlines are not perfect. They will occasionally be much smaller or larger than the underlying plant. Our tests show that on average, these effects average out.</p> <h3>1.3.3 Capacity Updates</h3> <p>Our capacity estimation model should produce relatively unbiased country-level aggregates, since it is trained to learn the typical ground coverage ratio of plants by country. The model has no way to distinguish between a very dense and a very sparse (e.g. dual-axis-tracking) plant in the same country. Plants with unusually high or low ground coverage ratios will not have accurate capacity estimates.</p> <h3>1.3.4 Construction Date Estimates</h3> <p>We are not able to directly estimate the construction date of a plant. We estimate an upper bound (the date of the image in which the plant was first seen in constructed state) and a lower bound (the date of the image in which the plant was last seen in an unconstructed state). For plants that were constructed before the launch date of Sentinel-2 in 2017, we produce only an upper bound.</p> <p>We leave it to consumers of the data to interpret these bounds and/or estimate likely grid connection dates.</p> <p><strong>2. Attribution</strong></p> <p>TZ-SAM is made available under a Creative Commons Attribution Non-Commercial 4.0 International License (CC-BY-NC-4.0). Attribution to TransitionZero is required. You must also clearly indicate if you have made any changes to the TZ-SAM dataset and what these are. Please refer to the suggested citation formats:</p> <ul> <li>&ldquo;TransitionZero Solar Asset Mapper, TransitionZero, May 2024 release.&rdquo;</li> <li>&ldquo;TZ-SAM, TransitionZero, May 2024 release.&rdquo;</li> <li>&ldquo;TransitionZero (2024) Solar Asset Mapper.&rdquo;</li> </ul>

opencc-by-nc-4.0May 2024View details →
Figshare44/100

Unstable Crystallographic & Molecular Structures for Machine Learning of System Energies

<div> <div> <div> <p>Extended QM9 (E-QM9) includes diverse sizes (i.e. number of atoms) and compositions of OoE molecules, through extending a subset of QM9 with OoE versions of 10k of its molecules.</p> <p>Periodic crystals (PC) allows learning regular bonding patterns that arise in periodic structures by repeating the base crystal lattice. We use the Face-Centred Cubic (fcc) Bravais lattice for aluminium (Al) and copper (Cu) crystals.</p> <p>Crystal Growth (CG) contains growing crystals of increasing size and complexity. Starting from a basic fcc crystal seed of 14 atoms, new systems are generated by iteratively placing atoms at a random location on the surface of the growing crystal following its lattice pattern, with sizes ranging from 15 to 114 atoms. We use 20 random seeds for each atom type, thus creating 40 varied Al and Cu crystal growths and 4,000 stable systems. As a result, for a given crystal size and composition (atom type), there are 20 samples with differently located atoms. CG enables experi- menting with large scale atomic interactions in non-regular sys- tems, and enables evaluation of an ML method&rsquo;s ability to learn how each atom contributes to the final potential energy.</p> <p>In all datasets, OoE systems are obtained by compressing/dilating all interatomic distances (i.e. isometrically) at regular intervals within 90-150% of stable geometry, which we refer to as &lsquo;scaling&rsquo;. In other words, scaling is applied to the coordinates of all atoms within the system. At each geometry, the ground-truth potential energy is calculated using CP2K7&rsquo;s DFT.</p> </div> </div> </div>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record