Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
WaivOps WRLD-SMB: Open Audio Resources for Machine Learning in Music
<p><strong>WRLD-SMB Dataset</strong></p> <p>WRLD-SMB is an open audio dataset featuring a collection of synthetic drum recordings in the style of Brazilian samba music. It includes 1,100 audio loops recorded in uncompressed stereo WAV format, along with paired JSON files intended for the supervised training of generative AI audio models.</p> <p><strong>Overview</strong></p> <p>This dataset was developed using multi-velocity audio samples and a paired MIDI dataset. The intended use of this dataset is to train or fine-tune AI models in learning high-performance drum notations, aiming to replicate the live sound of a small drum ensemble. To facilitate augmentation and supervised training with labeled audio data, a dropout technique was employed on the rendered audio files to generate variational mixes of the drum tracks.</p> <p>The primary purpose of this dataset is to provide accessible content for machine learning applications in music and audio. Potential use cases include generative music, feature extraction, tempo detection, audio classification, rhythm analysis, drum synthesis, music information retrieval (MIR), sound design and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>1,100 audio loops (approximately 5.5 hours)</li> <li>16-bit 44.1kHz WAV format</li> <li>Tempo range: 90–120 BPM</li> <li>Paired label data (WAV + JSON)</li> <li>Variational drum patterns</li> <li>Subgenre styles (Traditional and modern samba, bossa nova, fusion)</li> </ul> <p>A JSON file is provided for referencing and converting MIDI note numbers to text labels. You can update the text labels to suit your preferences.</p> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The WRLD-SMB dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-WRLD-SMB">GitHub repository</a>.</p>
Structured Power Grid Simulation Dataset for Machine Learning: Failure and Survival Events in Grid2Op's L2RPN WCCI 2022 Environment
<p>This dataset was developed for and used in the paper titled <em>"Fault Detection for Agents in Power Grid Topology Optimization: A Comprehensive Analysis"</em> by Malte Lehna, Mohamed Hassouna, Dmitry Degtyar, Sven Tomforde, and Christoph Scholz, presented at the <em>Workshop on Machine Learning for Sustainable Power Systems (ML4SPS)</em>, part of <em>ECML PKDD 2024</em>. While the paper is pending formal publication, a preprint version is available on arXiv.</p> <p>The dataset contains structured training, validation, and test data comprising failure and survival events observed in transmission power grid simulations. These were generated using Grid2Op with the WCCI 2022 L2RPN environment. Each data instance is labeled with one of four classes, representing survival or impending failure in 1, 3, and 5 timesteps. This dataset was used to train, validate and test machine learning models that predict grid agent failures in topology optimization tasks. </p>
Supporting data for "Raising awareness of potential biases in medical machine learning: Experience from a Datathon"
<p>This archive contains files from a Datathon held virtually in February<br>2024 to introduce clinicians and data scientists to the challenge of<br>reviewing a clinical dataset for potential biases.</p>
WaivOps SYN-SE1: Open Audio Resources for Machine Learning in Music
<div> <p><strong>SYN-SE1 Dataset</strong></p> <p>SYN-SE1 is an open audio dataset containing archived recordings of a Studio Electronics SE1 analog synthesizer. It includes 1,000 one-shot audio samples recorded in uncompressed stereo WAV format, labeled by note key across a two-octave range. The presets encompass a variety of distinct synth bass and lower-pitched lead sounds, featuring filter modulations and spatial stereo imaging, providing a valuable resource for soundfont design, audio production, and training data for generative AI models.</p> <p>The primary purpose of this dataset is to provide accessible content for machine learning applications in music and audio. Potential use cases include pitch detection, musical note classification, audio synthesis, music information retrieval (MIR), sound design, and signal processing.</p> <br><strong>Specifications</strong> <ul> <li>1,000 audio samples</li> <li>16-bit WAV format</li> <li>Key note labeled samples</li> <li>JSON reference</li> <li>Analog synth bass and leads</li> </ul> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed and published by Patchbanks. All recordings have been obtained from verified sources to ensure copyright clearance.</p> <p>The SYN-SE1 dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="noopener">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-SYN-SE1" target="_blank" rel="noopener">GitHub repository</a>.</p> </div>
WaivOps POP-ROK: Open Audio Resources for Machine Learning in Music
<div> <p><strong>POP-ROK Dataset</strong></p> <p>POP-ROK is an open audio dataset featuring an uncurated collection of synthetic drum recordings in the style of pop rock music. It includes 5,378 audio loops recorded in uncompressed stereo WAV format, along with paired JSON files intended for the supervised training of generative AI audio models.</p> <p><strong>Overview</strong></p> <p>The POP-ROK Dataset was developed by sonifying a collection of approximately 30 acoustic drum kits with a paired MIDI dataset covering basic rhythm patterns, excluding toms. Data augmentation included a random drum-swapping method to generate unique drum kits and reverb simulations to represent various room sizes. This dataset is intended for training or fine-tuning AI models in rhythm notation with paired drum note labels, aiming to replicate the sound of live drumming.</p> <p>The primary purpose of this dataset is to provide accessible content for machine learning applications in music and audio. Potential use cases include generative music, feature extraction, tempo detection, audio classification, rhythm analysis, drum synthesis, music information retrieval (MIR), sound design and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>5,378 audio loops (approximately 24 hours)</li> <li>16-bit WAV format</li> <li>Tempo range: 100-130 BPM</li> <li>Paired label data (WAV + JSON)</li> <li>Variational drum patterns</li> <li>Subgenre styles (Pop, classic rock, soft rock, country)</li> </ul> <p>A JSON file is provided for referencing and converting MIDI note numbers to text labels. You can update the text labels to suit your preferences.</p> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The POP-ROK dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="noopener">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-POP-ROK" target="_blank" rel="noopener">GitHub repository</a>.</p> </div>
Machine Learning Articles Extracted from Google Scholar
<p>This dataset was created as part of a web scraping practice aimed at capturing academic information from Google Scholar. It contains data on <em>Machine Learning</em> research articles, including the article's title, authors, summary, direct link, citation count, and APA reference. This dataset was collected using Python and Selenium to develop skills in web scraping tools for extracting data from websites with dynamic content.</p> <p>The dataset was generated specifically as part of an academic exercise to learn and apply web scraping techniques, without a deep analysis intent for the data obtained. This dataset is intended as a resource for learning and evaluating the methods used in web data collection.</p> <p><strong>Included Fields</strong>:</p> <ul> <li><strong>title</strong>: Title of the research article.</li> <li><strong>link</strong>: Direct link to the article.</li> <li><strong>authors</strong>: Names of the article’s authors.</li> <li><strong>description</strong>: Summary or brief description of the article.</li> <li><strong>citations</strong>: Number of times the article has been cited on Google Scholar.</li> <li><strong>APA_citation</strong>: APA-formatted citation of the article.</li> </ul> <p>This dataset was created solely for educational purposes and to demonstrate the application of web scraping techniques in a controlled environment.</p>
Unveiling Web Fingerprinting in the Wild Via Code Mining and Machine Learning
<p>Dataset of Javascripts used for training and testing the fingerprinting algorithms described in </p> <p>Rizzo, Valentino, Stefano Traverso, and Marco Mellia. "Unveiling Web Fingerprinting in the Wild Via Code Mining and Machine Learning." <em>Proceedings on Privacy Enhancing Technologies</em> 2021.1 (2021): 43-63.</p>
InChI to IUPAC name machine learning model
<p>This is a machine learning model that predicts IUPAC names from InChI. It was trained on a dump of PubChem's database, and has a transformer encoder-decoder architecture.</p> <p><strong>Instructions</strong></p> <p>Requires:</p> <ul> <li>Python >= 3.6</li> <li>PyTorch == 1.6.0</li> </ul> <p>1. Install <a href="https://github.com/OpenNMT/OpenNMT-py/tree/2.0.0">OpenNMT-py</a> version 2.0.0:</p> <pre><code class="language-bash">pip install OpenNMT-py==2.0.0</code></pre> <p>2. Prepare InChI to be translated by splitting into individual characters separated by whitespace and saving in a text file. You can predict multiple IUPAC names by having one InChI per line (see example.inchi for reference).</p> <p>3. Perform the prediction with the supplied model file:</p> <pre><code class="language-bash">onmt_translate --beam_size 10 --length_penalty wu --alpha 1.0 --model inchi2iupac_step_259200.pt --src <infile> --max_length 300 --output <outfile></code></pre> <p> </p>
Data Set for 'Self-Supervised Machine Learning for Live Cell Imagery Segmentation'
<p><strong>Self-supervised machine learning code and data for segmenting live cell imagery (Matlab)</strong></p> <p><em>Running the Code</em></p> <p>SSL_Demo_2.m : main program for self-supervised machine learning segmentation</p> <p>SSL_Declumping_2.m : main program for declumping application (applied to output of SSL_Demo_2.m)</p> <p>This Matlab code is designed to be used with time-resolved live cell microscopy images (tiffs) for the automated segmentation of cells from background.</p> <p>It is recommended you first run this code with its accompanying demo data (included in this package), keeping the current directory structure.</p> <p>Simply open SSL_Demo_2.m or SSL_Declumping_2.m in Matlab and hit Run.</p> <p><em>Code Methodology</em></p> <p>The principle of self-supervised machine learning is that you simply load your images and Run - no parameter tuning needed, no training imagery required.</p> <p>Run from start to finish, the SSL_Demo_2.m code uses consecutive pairs of images to generate training data of 'cells' and 'background' via dynamic feature vectors based on optical flow (unsupervised). These self-labeled pixels are then used to generate static feature vectors (entropy, gradient), which in turn are used to train a classifier model. The training data is updated every image in order to automatically adapt to temporal changes in cell morphologies or background illumination.</p> <p>The code was tested for high fidelity segmentation using five different modes of light microscopy: transmitted light, DIC, phase contrast, fluorescence and interference reflection microscopy.</p> <p>Six different cell lines were imaged to cover a range of morphologies and phenotypic dynamics using three cameras of differing resolutions.</p> <p>The associated manuscript for this work can be found here (although the latest version is under peer review as of this writing): </p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1">https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1</a></p> <p>This code was tested on Matlab v2020a and v2021a using commercially available laptop computers running the Windows 10 operating system.</p>
PSML: A Multi-scale Time-series Dataset for Machine Learning in Decarbonized Energy Grids (Dataset)
<p><strong>Abstract</strong></p> <p>The electric grid is a key enabling infrastructure for the ambitious transition towards carbon neutrality as we grapple with climate change. With deepening penetration of renewable energy resources and electrified transportation, the reliable and secure operation of the electric grid becomes increasingly challenging. In this paper, we present PSML, a first-of-its-kind open-access multi-scale time-series dataset, to aid in the development of data-driven machine learning (ML) based approaches towards reliable operation of future electric grids. The dataset is generated through a novel transmission + distribution (T+D) co-simulation designed to capture the increasingly important interactions and uncertainties of the grid dynamics, containing electric load, renewable generation, weather, voltage and current measurements at multiple spatio-temporal scales. Using PSML, we provide state-of-the-art ML baselines on three challenging use cases of critical importance to achieve: (i) early detection, accurate classification and localization of dynamic disturbance events; (ii) robust hierarchical forecasting of load and renewable energy with the presence of uncertainties and extreme events; and (iii) realistic synthetic generation of physical-law-constrained measurement time series. We envision that this dataset will enable advances for ML in dynamic systems, while simultaneously allowing ML researchers to contribute towards carbon-neutral electricity and mobility. </p> <p><strong>Data Navigation</strong></p> <p>Please download, unzip and put somewhere for later benchmark results reproduction and data loading and performance evaluation for proposed methods.</p> <pre><code>wget https://zenodo.org/record/5130612/files/PSML.zip?download=1 7z x 'PSML.zip?download=1' -o./ </code></pre> <p><strong>Minute-level Load and Renewable</strong></p> <ul> <li>File Name <ul> <li>ISO_zone_#.csv: `CAISO_zone_1.csv` contains minute-level load, renewable and weather data from 2018 to 2020 in the zone 1 of CAISO.</li> </ul> </li> <li>- Field Description <ul> <li>Field `<em>time</em>`: Time of minute resolution.</li> <li>Field `<em>load_power</em>`: Normalized load power.</li> <li>Field `<em>wind_power</em>`: Normalized wind turbine power.</li> <li>Field `<em>solar_power</em>`: Normalized solar PV power.</li> <li>Field `<em>DHI</em>`: Direct normal irradiance.</li> <li>Field `<em>DNI</em>`: Diffuse horizontal irradiance.</li> <li>Field `<em>GHI</em>`: Global horizontal irradiance.</li> <li>Field <em>`Dew Point</em>`: Dew point in degree Celsius.</li> <li>Field `<em>Solar Zeinth Angle</em>`: The angle between the sun's rays and the vertical direction in degree.</li> <li>Field `<em>Wind Speed</em>`: Wind speed (m/s).</li> <li>Field `<em>Relative Humidity</em>`: Relative humidity (%).</li> <li>Field `<em>Temperature</em>`: Temperature in degree Celsius.</li> </ul> </li> </ul> <p><strong>Minute-level PMU Measurements</strong></p> <ul> <li>File Name <ul> <li>case #: The `case 0` folder contains all data of scenario setting #0. <ul> <li>pf_input_#.txt: Selected load, renewable and solar generation for the simulation.</li> <li>pf_result_#.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> <li>Filed Description <ul> <li>Field <em>`time`</em>: Time of minute resolution.</li> <li>Field <em>`Vm_###`</em>: Voltage magnitude (p.u.) at the bus ### in the simulated model.</li> <li>Field <em>`Va_###`</em>: Voltage angle (rad) at the bus ### in the simulated model.</li> <li>Field <em>`P_#_#_#`</em>: `P_3_4_1` means the active power transferring in the #1 branch from the bus 3 to 4.</li> <li>Field <em>`Q_#_#_#`</em>: `Q_5_20_1` means the reactive power transferring in the #1 branch from the bus 5 to 20.</li> </ul> </li> </ul> <p><strong>Millisecond-level PMU Measurements</strong></p> <ul> <li>File Name <ul> <li>Forced Oscillation: The folder contains all forced oscillation cases. <ul> <li>row_#: The folder contains all data of the disturbance scenario #. <ul> <li>dist.csv: Three-phased voltage at nodes in the distribution system via T+D simualtion.</li> <li> info.csv: This file contains the start time, end time, location and type of the disturbance</li> <li>trans.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> <li>Natural Oscillation: The folder contains all natural oscillation cases. <ul> <li>row_#: The folder contains all data of the disturbance scenario #. <ul> <li>dist.csv: Three-phased voltage at nodes in the distribution system via T+D simualtion.</li> <li>info.csv: This file contains the start time, end time, location and type of the disturbance.</li> <li>trans.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> </ul> </li> <li>Filed Description <ul> <li>trans.csv <ul> <li> - Field <em>`Time(s)`</em>: Time of millisecond resolution.</li> <li> - Field <em>`VOLT ###`</em>: Voltage magnitude (p.u.) at the bus ### in the transmission model.</li> <li> - Field <em>`POWR ### TO ### CKT #`</em>: `POWR 151 TO 152 CKT '1 '` means the active power transferring in the #1 branch from the bus 151 to 152.</li> <li> - Field <em>`VARS ### TO ### CKT #`</em>: `VARS 151 TO 152 CKT '1 '` means the reactive power transferring in the #1 branch from the bus 151 to 152.</li> </ul> </li> <li>dist.csv <ul> <li>Field <em>`Time(s)`</em>: Time of millisecond resolution.</li> <li>Field <em>`####.###.#`</em>: `3005.633.1` means per-unit voltage magnitude of the phase A at the bus 633 of the distribution grid, the one connecting to the bus 3005 in the transmission system.</li> </ul> </li> </ul> </li> </ul>
Code and data for "KGML-ag: A Modeling Framework of Knowledge-Guided Machine Learning to Simulate Agroecosystems: A Case Study of Estimating N2O Emission using Data from Mesocosm Experiments "
<p>This is code and data for manuscript: <br> "KGML-ag: A Modeling Framework of Knowledge-Guided Machine Learning to Simulate Agroecosystems: <br> A Case Study of Estimating N<sub>2</sub>O Emission using Data from Mesocosm Experiments"<br> Licheng Liu, Shaoming Xu, Zhenong Jin*, Jinyun Tang, Kaiyu Guan, Timothy J. Griffis, <br> Matt D. Erickson, Alexander L. Frie, Xiaowei Jia, Taegon Kim, Lee T. Miller, Bin Peng, Shaowei Wu, Yufeng Yang, Wang Zhou, Vipin Kumar</p> <p>All the files belong to Prof. Zhenong Jin, University of Minnesota, UA. jinzn@umn.edu<br> "code" foler includes code for data processing, model training, and results plotting.<br> "trained_model_saved" includes all trained model so you can use to reproduce the results showed in the study;<br> "data" includes all data presented in the study. Finetuning data is refering to Miller, L.T. , Griffis, T. J., Erickson, M. D., Turner, P. A., Deventer, M. J., Chen, Z., Yu, Z., Venterea, R.T., Baker, J. M., and Frie, A. L. (2021). Response of nitrous oxide emissions to future changes in precipitation and individual rain events. Journal of Environmental Quality, In review</p>
Machine Learning Quantum Reaction Rate Constants
<p>Dataset of 1,517,419 quantum reaction rate constant products <span class="math-tex">\(k^{\text {QM}}(T)Q_{\text R}(T)\)</span> computed from the transmission coefficient for model single and double barrier minimum energy paths. Here <span class="math-tex">\(k^{QM}(T) \)</span> is the quantum reaction rate constant at temperature <span class="math-tex">\(T\)</span> and <span class="math-tex">\(Q_\text{R}(T)\)</span> is the reactant partition function computed with the rigid rotor and harmonic oscillator approximations.This dataset was created for Ref [1] where it was used to train and test a DNN to predict <span class="math-tex">\(\log{k^{\text{QM}}(T)Q_\text{R}(T)}\)</span>.</p> <p><strong>Cite as</strong></p> <p>Please cite the following references when using this dataset:</p> <p>[1] E. Komp and S. Valleau, Machine Learning Quantum Reaction Rate Constants, <em>J. Phys. Chem. A</em>, 124:8607–8613, 2020, <a href="https://pubs.acs.org/doi/abs/10.1021/acs.jpca.0c05992">doi: 10.1021/acs.jpca.0c05992</a>.</p> <p>[2] E. Komp and S.Valleau, Machine Learning Quantum Reaction Rate Constants (1.0.0) [Data set], 2020, <em>Zenodo</em>, <a href="https://doi.org/10.5281/zenodo.5510392">https://doi.org/10.5281/zenodo.5510392 </a></p> <p><strong>Contents</strong></p> <p>Descriptions of entries in the tabular dataset file `QM_kQ.csv`. Please refer to the publication [1] for details.</p> <ul> <li>`mass_au`: Mass of the reactants in atomic units. </li> <li>`width_1_au`: Width of first potential energy barrier in atomic units.</li> <li>`width_2_au`: Width of second (if present) potential barrier in atomic units. For single barriers width_2_au = 0.0.</li> <li>`height_1_au`: Activation energy of first potential energy barrier in atomic units.</li> <li>`height_2_au`: Activation energy of second (if present) potential energy barrier in atomic units. For single barriers height_2_au = 0.0.</li> <li>`dist_au`: For double barriers, absolute value of the difference between the position of the two potential energy maxima along the reaction coordinate in atomic units. For single barriers dist_au = 0.0.</li> <li>`alpha_symm`: Symmetry constant for single barriers, defined as the difference between product and reactant energies in atomic units.</li> <li>`alpha_double`: Symmetry constant for double barriers, defined as the sum of the normalized difference between barrier heights and the normalized difference between barrier widths, unitless.</li> <li>`slope`: Slope of the first reaction barrier along the reaction coordinate in atomic units. Slope values were evaluated numerically from the type of potential energy barrier see Ref [1] SI.</li> <li>`temp_K`: Temperature in Kelvin.</li> <li>`kQ_rate`: Quantum reaction rate constant times reactant partition function in units of 1/ps. </li> <li>`log_kQ_rate`: Natural logarithm of the quantum reaction rate constant times the reactant partition function in units of log(1/ps).</li> </ul>
Data for: Machine learning for predicting environmental mobility based on retention behaviour
<p>This repository contains the data and supplementary information for the paper: "Machine learning for predicting environmental mobility based on retention behaviour".</p>
A map of global peatland extent created using machine learning (Peat-ML)
<p>Map of global peatland extent estimated by machine learning. The download includes both a netcdf file version and a GeoTIFF (as a zip archive)</p> <p>Abstract from associated paper:</p> <p>Peatlands store large amounts of soil carbon and freshwater, constituting an important component of the global carbon and<br> hydrologic cycles. Accurate information on the global extent and distribution of peatlands is presently lacking but is needed<br> by Earth System Models (ESMs) to simulate the effects of climate change on the global carbon and hydrologic balance. Here,<br> we present Peat-ML, a spatially continuous global map of peatland fractional coverage generated using machine learning<br> techniques suitable for use as a prescribed geophysical field in an ESM. Inputs to our statistical model follow drivers of<br> peatland formation and include spatially distributed climate, geomorphological and soil data, along with remotely-sensed<br> vegetation indices. Available maps of peatland fractional coverage for 14 relatively extensive regions were used along with<br> mapped ecoregions of non-peatland areas to train the statistical model. In addition to qualititative comparisons to other maps<br> in the literature, we estimated model error in two ways. The first estimate used the training data in a blocked leave-one-out<br> cross-validation strategy designed to minimize the influence of spatial autocorrelation. That approach yielded an average r<sup>2</sup><br> of 0.73 with a root mean squared error and mean bias error of 9.11% and -0.36%, respectively. Our second error estimate<br> was generated by comparing Peat-ML against a high-quality, extensively ground-truthed map generated by Ducks Unlimited<br> Canada for the Canadian Boreal Plains region. This comparison suggests our map to be of comparable quality to mapping<br> products generated through more traditional approaches, at least for boreal peatlands.</p>
Predicting Survival of Tongue Cancer Patients by Machine Learning Models
<p>This repository contains the dataset used in the paper "Predicting Survival of Tongue Cancer Patients by Machine Learning Models." The dataset contains information on 1712 tongue cancer curative surgery recipients. Each row represents one patient. The meaning of each variable is summarized here:</p> <ul> <li>id: patient identifier</li> <li>gender: patient sex</li> <li>survival: patient survival status at follow-up; 0: survival, 1: death</li> <li>follow_time: length of follow-up period in days</li> <li>part: site of operation</li> <li>stage: tumor stage; 0: very small, no spreading, 1: small, no spreading, 2: some growth, spreading, 3: large, spreading to surrounding tissue or lymph nodes, 4A/4B/4C: larger, metastasis to at least one other organ</li> <li>op: operation status; 1: complete</li> <li>rt: radiation therapy status; 0: not received, 1: received</li> <li>ct: chemotherapy status; 0: not received, 1: received</li> <li>t_stage: tumor size; 1 (small) to 4 (large)</li> <li>n_stage: metastasis to lymph nodes; 0 (no metastasis) to 3 (metastasis to multiple lymph nodes)</li> <li>grade: tumor grade; 1 (no proliferation) to 3 (aggressive proliferation)</li> </ul>
Data for: Explainable Quantum Machine Learning
<p>Data used in the numerical experiments for the publication "Explainable Quantum Machine Learning" (<a href="https://arxiv.org/abs/2301.09138">arXiv:2301.09138</a>).</p>
A Systematic Literature Review of Machine Learning for Uncovering Software Faults and Failures
<p>This data set contains the results of an extensive, systematic literature review on the use of machine learning (ML) for uncovering software faults and failures. Covering the period of 2019 to 2022, this literature review identifies 874 relevant publications, classified into six distinct quality assurance tasks. Results show a compound annual growth rate (CAGR) of relevant publications of 38% over the last five years.</p> <p>This literature review particularly analyzed in how far these relevant papers leverage synergies between different quality assurance tasks. Results show that only 3% of all relevant papers leverage such synergies, indicating ample opportunities for future research. For example, a single type of quality assurance activity may not suffice to deliver the expected software quality. Ideally, one would use a suitable combination of different types of activities – such as combining dynamic testing with static code analysis. Also, leveraging the synergies between different quality assurance activities can increase the effectiveness of the individual activities. For example, having a good estimate of the fault density of a software component (e.g., using deep learning-driven fault prediction techniques) could help optimize and prioritize testing effort and budget.</p>
Development of a machine learning model for river bedload - Data, Model, and Scripts
<p>This repository for “Development of a machine learning model for river bedload” Hosseiny et. Al (in review at Earth Surface Dynamics) contains the following assets. These assets may need to be modified for your purposes. You are responsible for inspecting these assets and making adjustments as necessary.</p> <p>Assets:</p> <p>1) The trained ANN model described in Hosseiny et al. (in review) as a .zip file named ‘Hosseiny_et_al_trained_ANN.zip’. This contains a folder (BEDLOAD_MODEL_FINAL) which contains the trained ANN Model (saved_model.pb) and associated information related to the input variables and variable weights.</p> <p>2) An accompanying Jupyter notebook named “bedload_ann_example.ipynb” that provides a step-by-step guide for implemented the trained ANN model (Asset 1).</p> <p>3) A .xls file named “Hosseiny et al_Supplemental_Data_Tables.xlsx” which provides the original observations that the model was trained and tested on, the summary statistics of the original input data as a compilation and for individual sites, the model errors associated with training and validation steps, the bedload calculations from the four uncalibrated existing bedload transport models described in the original study and for the ANN for the test data population, associated summary statistics with model output, and additional site-specific calculations of model error.</p>
Dataset of Machine Learning forecasted VTEC from paper: Uncertainty Quantification for Machine Learning-based Ionosphere and Space Weather Forecasting
<p>The *csv files contain forecasted one-day-ahead Vertical Total Electron Content (VTEC), consisting of the mean/median VTEC values and the upper and lower VTEC bounds of the 95% confidence intervals of 4 models based on machine learning for test data.</p> <p>The first part of the *csv file name corresponds to the type of model: SE stands for the super-ensemble VTEC model, QGB stands for the quantile gradient boosting VTEC model, BNN1 stands for the Bayesian neural network VTEC model, and BNN2 stands for the Bayesian neural network with negative log-likelihood (NLL) loss VTEC model. The second part of the file name refers to the geographic location of the VTEC points for which the forecast is performed, i.e., 10E70N for 10 degree of longitude and 70 degree of latitude, 10E40N for 10 degree of longitude and 40 degree of latitude, and 10E10N for 10 degree of longitude and 10 degree of latitude. The last part of the file name corresponds to the test year, i.e., year 2017.</p> <p>The SE_*_2017.csv file consists of 14 columns. The index column ("Date-time") is expressed in Coordinated Universal Time (UTC) as YYYY-MM-DD. Columns 1-3 contain the VTEC forecast results of Random Forest (RF) trained on three data subsets; columns 4-6 contain the VTEC forecast results of Adaptive Boosting (AB) trained on three data subsets; columns 7-9 contain the VTEC forecast results of Gradient Boosting (XGBoost) trained on three data subsets. Column 10 ("Mean") represents the mean of columns 1-9, i.e., the ensemble mean; column 11 ("Std") represents the standard deviation of columns 1-9, i.e., the ensemble spread; columns 12 ("UB") and 13 ("LB") contain the upper and lower bounds of the 95% confidence interval of VTEC, respectively; and column 14 contains the Global Ionosphere Maps (GIM) values of CODE, i.e., the ground-truth in this study.</p> <p>The QGB_*_2017.csv file consists of 4 columns. The index column ("Date-time") is expressed in UTC as YYYY-MM-DD. Column 1 ("Median") contains the median VTEC forecast, column 2 ("LB") contains the lower VTEC bound of the 95% confidence interval, column 3 ("UB") contains the upper VTEC bound of the 95% confidence interval, and column 4 contains the GIM values of CODE, i.e., the ground-truth in this study.</p> <p>The BNN*_2017.csv file consists of 5 columns. The index column ("Date-time") is expressed in UTC as YYYY-MM-DD. Column 1 ("Mean") contains the mean VTEC forecast, column 2 ("Std") contains the standard deviation, column 3 contains GIM values of CODE, i.e., ground-truth in this study; column 4 ("UB") contains the upper VTEC bound of the 95% confidence interval, and column 5 ("LB") contains the lower VTEC bound of the 95% confidence interval.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>Contact</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>If you have any questions regarding these data, please contact:</p> <p>Randa Natras</p> <p>Deutsches Geodätisches Forschungsinstitut (DGFI-TUM)</p> <p>Technical University of Munich</p> <p>Arcisstraße 21</p> <p>80333 München</p> <p>randa.natras@tum.de</p>
Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset
<p><b>Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset.</b></p><p>Ribonucleic acids (RNA) play crucial roles in living organisms as they are involved in key processes necessary for proper cell functioning. Some RNA molecules, such as bacterial ribosomes and precursor messenger RNA, are targets of small molecule drugs, while others, e.g., bacterial riboswitches or viral RNA motifs are considered as potential therapeutic targets. Thus, the continuous discovery of new functional RNA increases the demand for developing compounds targeting them and for methods for analyzing RNA—small molecule interactions. We recently developed fingeRNAt - a software for detecting non-covalent bonds formed within complexes of nucleic acids with different types of ligands. The program detects several non-covalent interactions, such as hydrogen and halogen bonds, ionic, Pi, inorganic ion- and water-mediated, lipophilic interactions, and encodes them as computational-friendly Structural Interaction Fingerprint (SIFt). Here we present the application of SIFts accompanied by machine learning methods for binding prediction of small molecules to RNA targets. We show that SIFt-based models outperform the classic, general-purpose scoring functions in virtual screening. We discuss the aid offered by Explainable Artificial Intelligence in the analysis of the binding prediction models, elucidating the decision-making process, and deciphering molecular recognition processes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.