Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

285

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

285 results for “Machine learning dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

EO4WildFires: An Earth Observation multi-sensor, time-series machine-learning-ready benchmark dataset for wildfire impact prediction

<p>This paper presents a benchmark dataset called EO4WildFires; a multi-sensor (multi spectral; Sentinel-2, Synthetic-Aperture Radar - SAR; Sentinel-1, meteorological parameters; NASA Power) time-series dataset that spans 45 countries, which can be used for developing machine learning and deep learning methods targeted for the estimation of the area that a forest wildfire might cover.</p> <p>This novel EO4WildFires dataset is annotated using EFFIS (European Forest Fire Information System) as forest fire detection and size estimation data source. A total of 31,742 wildfire events are gathered from 2018 to 2022. For each event, Sentinel-2 (multispectral), Sentinel-1 (SAR) and meteorological data are assembled into a single data cube. The meteorological parameters that are included in the data cube are: ratio of actual partial pressure of water vapor to the partial pressure at saturation, average temperature, bias corrected average total precipitation, average wind speed, fraction of land covered by snowfall, percent of root zone soil wetness, snow depth, snow precipitation, as well as percent of soil moisture.</p> <p>The main problem that this dataset is designed to address, is the severity forecasting before wildfires occur. The dataset is not used to predict wildfire events, but rather to predict the severity (size of area damaged by fire) of a wildfire event, if that happens in a specific place under the current and historical forest status, as recorded from multispectral and SAR images, and meteorological data.</p> <p>Using the data cube for the collected wildfire events, the EO4WildFires dataset is used to realize three (3) different preliminary experiments, in order to evaluate the contributing factors for wildfire severity prediction. The first experiment evaluates wildfire size using only the meteorological parameters, the second one utilizes both the multispectral and SAR parts of the dataset, while the third exploits all dataset parts. In each experiment, machine learning models are developed, and their accuracy is evaluated.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

A dataset of Earth Observation Data for Lithological Mapping using Machine Learning

<p><strong>Dataset Information</strong></p> <p>Machine Learning (ML) algorithms had successfully contributed in the creation of automated methods of recognizing patterns in high-dimensional data. Remote sensing data&nbsp; covers&nbsp; wide&nbsp; geographical areas and could be used to solve the problem of the demand of various&nbsp; in-situ data.&nbsp; Lithologicall mapping using remotely sensed data&nbsp; is one of the most challenging&nbsp; applications of ML algorithms. In the framework of the &ldquo;AI for Geoapplications&rdquo; project , ML and especially Deep Learning (DL) methodologies are investigated&nbsp; for&nbsp; the identification and characterization of the lithology based on remote sensing data in various&nbsp; pilot areas&nbsp; in Greece.&nbsp; In order to train and test the various ML algorithms, a dataset consisting of&nbsp; 30 ROIs selected&nbsp; mainly&nbsp; from low -vegetated areas,&nbsp; that cover 2% of the total&nbsp; area of Greece was created</p> <p><strong>Dataset Preprocessing</strong></p> <p>Dataset preprocessing was executed using a combination of SNAP, QGIS and ENVI tools.</p> <p>Preprocessing steps:</p> <p>Defining areas with the following properties:</p> <ul> <li> <p>Zero cloud and snow coverage</p> </li> <li> <p>No water bodies</p> </li> <li> <p>Minimum vegetation</p> </li> </ul> <p>For the Aster Images:</p> <ul> <li> <p>Subset on defined areas</p> </li> <li> <p>Mosaic images when needed</p> </li> <li> <p>Digitising clouds</p> </li> </ul> <p>For the Labels:</p> <ul> <li> <p>We got the Soil map from YPEN (<a href="https://ypen.gov.gr/">https://ypen.gov.gr/</a>)</p> </li> <li> <p>Subset on defined areas</p> </li> <li> <p>All categories are represented with good analogies</p> </li> <li> <p>Clip label files with digitised clouds</p> </li> <li> <p>Rasterize</p> </li> </ul> <p>&nbsp;</p> <p>For the Labels we have eighteen categories for the twenty-eight areas that we collected data.&nbsp;We use the following coding&nbsp;for the&nbsp;Labels of our <strong>Dataset</strong>:</p> <table> <tbody> <tr> <td> <p><strong>Alluvial deposits</strong></p> </td> <td> <p><strong>0</strong></p> </td> </tr> <tr> <td> <p><strong>Limestone colluvial deposits</strong></p> </td> <td> <p><strong>1</strong></p> </td> </tr> <tr> <td> <p><strong>Limestones</strong></p> </td> <td> <p><strong>2</strong></p> </td> </tr> <tr> <td> <p><strong>Schists</strong></p> </td> <td> <p><strong>3</strong></p> </td> </tr> <tr> <td> <p><strong>Quaternary sediments</strong></p> </td> <td> <p><strong>4</strong></p> </td> </tr> <tr> <td> <p><strong>Gneiss</strong></p> </td> <td> <p><strong>5</strong></p> </td> </tr> <tr> <td> <p><strong>Slope fan debris</strong></p> </td> <td> <p><strong>6</strong></p> </td> </tr> <tr> <td> <p><strong>Mixed flysch</strong></p> </td> <td> <p><strong>7</strong></p> </td> </tr> <tr> <td> <p><strong>Flysch shale and cherts</strong></p> </td> <td> <p><strong>8</strong></p> </td> </tr> <tr> <td> <p><strong>Dolomites</strong></p> </td> <td> <p><strong>9</strong></p> </td> </tr> <tr> <td> <p><strong>Granite</strong></p> </td> <td> <p><strong>10</strong></p> </td> </tr> <tr> <td> <p><strong>Sandstone flysch</strong></p> </td> <td> <p><strong>11</strong></p> </td> </tr> <tr> <td> <p><strong>Flysch colluvial deposits</strong></p> </td> <td> <p><strong>12</strong></p> </td> </tr> <tr> <td> <p><strong>Peridotite and Gabbro</strong></p> </td> <td> <p><strong>13</strong></p> </td> </tr> <tr> <td> <p><strong>River bed deposits</strong></p> </td> <td> <p><strong>14</strong></p> </td> </tr> <tr> <td> <p><strong>Gneiss colluvial deposits</strong></p> </td> <td> <p><strong>15</strong></p> </td> </tr> <tr> <td> <p><strong>Not available</strong></p> </td> <td> <p><strong>-100</strong></p> </td> </tr> <tr> <td> <p><strong>cloud coverage</strong></p> </td> <td> <p><strong>-999</strong></p> </td> </tr> </tbody> </table> <p>The following table lists the available <strong>areas </strong>and the <strong>categories </strong>that each contains<strong>:&nbsp;<a href="https://docs.google.com/spreadsheets/d/17q0L5Ltz7V4uBY9i6DhULJsJCtf7BOY1nbblB-hf3Pw/edit?usp=share_link">Lithology_Dataset</a> </strong></p> <p>&nbsp;</p> <p>For the <strong>Sentinel-2 images</strong>, we made the following process:</p> <ul> <li> <p><strong>Resampling 10m</strong></p> </li> <li> <p><strong>Subset on defined areas</strong></p> </li> </ul> <p>The Sentinel-2 map contains: Sentinel 2 false colour composite 11/8/4 with OSM background</p> <p>The Final step is the collocation of the previous into a datacube i.e a multidimensional array with 25 bands (datacube dimensions differentiate for every area) using the Aster image as base (15m spatial resolution).&nbsp;</p> <ul> <li> <p>Bands 1-14: Aster</p> </li> <li> <p>Bands 15-24: S2</p> </li> <li> <p>Band 25: Label</p> </li> </ul> <p>The code for preprocessing the dataset in order to be used for machine learning algorithms can be found in the following link:&nbsp;&nbsp;</p> <p><a href="https://github.com/georgegiannop/Lithology">https://github.com/georgegiannop/Lithology</a></p> <p><strong>Citation</strong></p> <p>If you use this dataset in your work, please cite our paper:</p> <p>Vernikos, I., Giannopoulos, G., Christopoulou, A., Begaj, A., Stefouli, M., Bratsolis, E., and Charou, E.: A dataset of Earth Observation Data for Lithological Mapping using Machine Learning, EGU General Assembly 2023, Vienna, Austria, 24&ndash;28 Apr 2023, EGU23-17570,&nbsp;<a href="https://doi.org/10.5194/egusphere-egu23-17570">https://doi.org/10.5194/egusphere-egu23-17570</a>, 2023.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 2

<p>Supplementary files containing datasets needed to reproduce the results of the manuscript &quot;Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states&quot; by S. Choudhury et al.</p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. param_fixing.zip - self-explanatory (Figure 4 &amp; 5); contains an explanatory note for this part (experiment_details.txt), and the file containing Km values fetched from the BRENDA database (Km_database.csv).</p> <p>2. scripts.zip - scripts to generate figure 2-5 on toy data</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 1

<p><strong>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al (https://doi.org/10.1101/2023.02.21.529387).</strong></p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. models.zip - contains thermodynamically curated steady-state and nonlinear kinetic models of <em>E. coli </em>metabolism used in this study. Also contains the samples of steady-state metabolite concentrations and metabolic fluxes used in the study presented in Figure 3 (steady-state samples used for preparing Figures 2 and 4).</p> <p>2. renaissance_incidence_results.zip - self-explanatory (Figure 2a and 2b)</p> <p>3. ODE_solutions.zip - self-explanatory (Figure 2c)</p> <p>4. bioreactor_simulations1-3.zip - self-explanatory (Figure 2d)</p> <p>5. steady_state_analysis.zip - RENAISSANCE results obtained for each of the steady states (Figure 3a)</p> <p>6. subspace_analysis.zip - RENAISSANCE results presented in Figure 3b-g</p> <p><strong>The remaining datasets are published in the following links</strong></p> <p><em>&nbsp;- https://doi.org/10.5281/zenodo.7930084</em></p> <p><em>&nbsp;- https://doi.org/10.5281/zenodo.10391802</em></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

MISATO - Machine learning dataset for structure-based drug discovery

<p>Developments in Artificial Intelligence (AI) have had an enormous impact on scientific research in recent years. Yet, relatively few robust methods have been reported in the field of structure-based drug discovery. To train AI models to abstract from structural data, highly curated and precise biomolecule-ligand interaction datasets are urgently needed. We present MISATO, a curated dataset of almost 20000 experimental structures of protein-ligand complexes, associated molecular dynamics traces, and electronic properties. Semi-empirical quantum mechanics was used to systematically refine protonation states of proteins and small molecule ligands. Molecular dynamics traces for protein-ligand complexes were obtained in explicit water. The dataset is made readily available to the scientific community via simple python data-loaders. AI baseline models are provided for dynamical and electronic properties. This highly curated dataset is expected to enable the next-generation of AI models for structure-based drug discovery. Our vision is to make MISATO the first step of a vibrant community project for the development of powerful AI-based drug discovery tools.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Dataset - What are the Machine Learning best practices reported by practitioners on Stack Exchange?

<p>The data correspond to the posts (questions and answers) retrieved by querying for posts related to the tag &#39;machine learning&#39; and the phrase &#39;best practice(s).&#39; The data were used as the basis for a study currently under review on discussing machine learning best practices as discussed by practitioners in question-and-answer communities such as Stack Exchange. The information from each type of post (i.e., questions and answers) is presented in multiple formats (i.e., .txt, .csv, and .xlsx).</p> <p>&nbsp;</p> <p><strong>Answers - Variables</strong></p> <ul> <li><strong>AID</strong>:<strong>&nbsp;</strong>&nbsp;Unique identification of the answer in the Q&amp;A website.</li> <li><strong>ParentId</strong>: Unique identification of the question associated with the answer in the Q&amp;A website&nbsp;</li> <li><strong>AcceptedAnswerId</strong>&nbsp;: In the case in which an answer is the most voted question associated with the&nbsp;<em>ParentId</em>, and it is different from the accepted answer, a different identifier from the&nbsp;<em>AID</em>&nbsp;is available. In the case in which the accepted question had a&nbsp;<em>score</em>&nbsp;lower than 1, a -1 is assigned.&nbsp;</li> <li><strong>ABody:</strong>&nbsp;&nbsp;HTML text of the answer.</li> <li><strong>Score:</strong>&nbsp;Upvotes - downvotes of the answer.</li> <li><strong>url_Answer:</strong>&nbsp;&nbsp;URL of the answer. The question URL can be from different websites.&nbsp;&nbsp;</li> <li><strong>type:</strong>&nbsp;best or accepted. Accepted in the case that the information belongs to the accepted answer of the&nbsp;<em>ParentId&nbsp;</em>question and best in the case in which it is the most voted question of the&nbsp;<em>ParentId&nbsp;</em>question.</li> <li><strong>Date:&nbsp;</strong>Creation date of the answer.</li> </ul> <p><strong>Questions - Variables</strong></p> <ul> <li><strong>QID</strong>: Unique identification of the question in the Q&amp;A website.&nbsp;</li> <li><strong>AcceptedAnswerId</strong>: Unique identification of the accepted answer for a specific question in the Q&amp;A website. In the case in which a question had a most-voted answer different from the accepted one, and the accepted one had a negative score, a -1 was assigned to the&nbsp;&nbsp;<em>AcceptedAnswerId</em><strong>.&nbsp;</strong></li> <li><strong>BestAnswerId</strong>: Unique identification of the most voted answer for a specific question in the Q&amp;A website. In the case in which the most voted and accepted questions were the same, then a -1 was assigned to the&nbsp;<em>BestAnswerId</em>.&nbsp;&nbsp;</li> <li><strong>Qtitle</strong>: Title of the question.</li> <li><strong>QBody</strong>: HTML text of the question.</li> <li><strong>Score</strong>: Upvotes - downvotes of the questions.</li> <li><strong>QTags</strong>: Tags that are associated with each question.</li> <li><strong>url_question</strong>: URL of the question. The question URL can be from different websites. &nbsp;</li> <li><strong>Date</strong>: Creation date of the question</li> </ul> <p>This dataset is a subset of the Stack Exchange dump of 03.2021 (<a href="https://archive.org/details/stackexchange_20210301">https://archive.org/details/stackexchange_20210301</a>) in which a series of filters were applied to obtain the data used in the study.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Dataset - Downscaling ERA5 Wind Speed Data: A Machine Learning approach considering Topographic Influences

<p>This dataset provides three products:</p> <p><strong>1. The topographic data.&nbsp;&nbsp;</strong></p> <p>These data are provided as GeoTIFF files for Europe with 1km x 1km spatial resolution. These maps include:</p> <ul> <li>Digital Elevation Model (DEM) map: Europe_DEM.tif</li> <li>Slope map: Europe_slope.tif</li> <li>Aspect map: Europe_aspect.tif</li> <li>Topographic Position Index (TPI) with a 5 km radius map: Europe_TPI_5.tif</li> <li>Topographic Position Index (TPI) with a 75 km radius map: Europe_TPI_75.tif</li> <li>Terrain Diversity Index (TDI) map: Europe_TDI.tif</li> </ul> <p>These data can be used as input maps for the preprocessing step. In addition, the two TPI maps can also be used in the regression process.</p> <p><strong>2. The resulting map of the preprocessing step.</strong>&nbsp;</p> <p>This map offers predictions on the quality of ERA5 data across Europe and is also provided as a GeoTIFF file with 1km x 1km spatial resolution under the name:</p> <ul> <li>&nbsp;Europe_classification.tif</li> </ul> <p>In this map, Class1 represents a good ERA5 quality with an RMSE of less than 1.5 m/s, Class2 represents a moderate ERA5 quality with an RMSE bigger than 1.5 m/s but less than 3 m/s, while Class 3 indicates a poor ERA5 quality with an RMSE greater than 3 m/s.</p> <p><strong>3. The downscaled wind speed time series data. </strong></p> <p>Europe has been divided into 64 equal area blocks to accommodate the large data size. Each downscaled dataset is provided as a NetCDF file, offering hourly wind speed time series for a year (8760 hours) at approximately 1km x 1km spatial resolution. Each NetCDF file has three dimensions: 'lon' representing longitude, 'lat' representing latitude, and 'time' representing the hour. The variable name for wind speed in the NetCDF file is 'WindSpeed'. The 'WindSpeed' variable is stored as an Int32 data type in the NetCDF file, with values multiplied by 10000 in order to significantly reduce the data size. To utilize this variable, please divide it by 10000.</p> <p>For regions identified as Class1 and Class2, the downscaled wind speed is obtained through a simple nearest neighbour spatial interpolation of ERA5 due to the good quality of ERA5 in these regions. However, for the regions identified as Class3, the downscaled wind speed is derived using the machine learning-based regression approach described in the relevant publication. The geographic extent and the visual representation for each block are provided&nbsp;in 'Readme.pdf' document.</p> <p>&nbsp;</p> <p>To cite this dataset, please cite our published paper in Environmental Research Letters (<strong>DOI:</strong> 10.1088/1748-9326/aceb0a)</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Training dataset used in the magazine paper entitled "A Flexible Machine Learning-Aware Architecture for Future WLANs"

<p><a href="https://arxiv.org/pdf/1910.03510.pdf"><strong>A Flexible Machine Learning-Aware Architecture for Future WLANs</strong></a></p> <p><strong>Authors: </strong>Francesc Wilhelmi, Sergio Barrachina-Mu&ntilde;oz, Boris Bellalta, Cristina Cano, Anders Jonsson &amp; Vishnu Ram.</p> <p><strong>Abstract:&nbsp;</strong>Lots of hopes have been placed in Machine Learning (ML) as a key enabler of future wireless networks. By taking advantage of the large volumes of data generated by networks, ML is expected to deal with the ever-increasing complexity of networking problems. Unfortunately, current networking systems are not yet prepared for supporting the ensuing requirements of ML-based applications, especially for enabling procedures related to data collection, processing, and output distribution. This article points out the architectural requirements that are needed to pervasively include ML as part of future wireless networks operation. To this aim, we propose to adopt the International Telecommunications Union (ITU) unified architecture for 5G and beyond. Specifically, we look into Wireless Local Area Networks (WLANs), which, due to their nature, can be found in multiple forms, ranging from cloud-based to edge-computing-like deployments. Based on ITU&#39;s architecture, we provide insights on the main requirements and the major challenges of introducing ML to the multiple modalities of WLANs.</p> <p><strong>Dataset description:&nbsp;</strong>This is the dataset generated for training a Neural Network (NN) in the Access Point (AP) (re)association problem in IEEE 802.11 Wireless Local Area Networks (WLANs).&nbsp;</p> <p>In particular, the NN is meant to output a prediction function of the throughput that a given station (STA) can obtain from a given Access Point (AP) after association. The features included in the dataset are:</p> <ol> <li>Identifier of the AP to which the STA has been associated.</li> <li>RSSI obtained from the AP to which the STA has been associated.</li> <li>Data rate in bits per second (bps) that the STA is allowed to use for the selected AP.</li> <li>Load in packets per second (pkt/s)&nbsp;that the STA generates.</li> <li>Percentage of data that the AP is able to serve before the user association is done.</li> <li>Amount of traffic load in pkt/s handled by the AP before the user association is done.</li> <li>Airtime in % that the AP enjoys before the user association is done.</li> <li>Throughput in pkt/s that the STA receives after the user association is done.</li> </ol> <p>The dataset has been generated through random simulations, based on the model provided in <a href="https://github.com/toniadame/WiFi_AP_Selection_Framework">https://github.com/toniadame/WiFi_AP_Selection_Framework</a>. More details regarding the dataset generation have been provided in&nbsp;<a href="https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans">https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans</a>.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Released Experimental Dataset for Sampled Automated Machine Learning

<p>Released Experimental Dataset of &quot;Doing More with Less: Characterizing Dataset Downsampling for AutoML&quot;</p> <p>&nbsp;</p> <p>Experiments were run for 5 and 60 minutes on 16 datasets:<br> 4 small: &lt; 10.000<br> 5 medium: &lt; 100.000<br> 7 large: &gt; 100.000<br> &nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Comprehensive Datasets for RNA Design, Machine Learning and Beyond

<p>This repository contains a comprehensive collection of RNA multi-loops extracted from major RNA databases, along with benchmark results for various RNA design algorithms. The resource is intended to facilitate research and development in RNA design, particularly for multi-loop structures.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

TCOM-CH4: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric methane profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;</p> <p><span>he </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>CH4 Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated CH4 profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>). It is important to note that unlike previous versions that might have used both HALOE and ACE measurements, this version exclusively utilizes </span><strong><span>ACE-FTS data</span></strong><span>, which is why the dataset starts from 2000.</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these CH4 differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>CH4 bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved CH4 profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean CH4 profiles:</span></p> <ul> <li> <p><code><span>zmch4_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmch4_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023.</span></p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

TCOM-N2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric nitrous oxide profile dataset [1991-2021] constructed using machine-learning

<p>Methodology: &nbsp;</p> <p><span>The </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>N2O Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated N2O profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these N2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>N2O bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved N2O profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean N2O profiles:</span></p> <ul> <li> <p><code><span>zmn2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmn2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105&ndash;5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023.</span></p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Dataset - FetMRQC: an open-source machine learning framework for multi-centric fetal brain MRI quality control

<p>This dataset contains the data and model used in the paper</p> <blockquote> <p>Thomas Sanchez, Oscar Esteban, Yvan Gomez, Alexandre Pron, M&eacute;riam Koob, Vincent Dunet, Nadine Girard, Andras Jakab, Elisenda Eixarch, Guillaume Auzias, and Meritxell Bach Cuadra. "FetMRQC: an open-source machine learning framework for multi-centric fetal brain MRI quality control." <a href="https://arxiv.org/abs/2311.04780"><em>arXiv preprint arXiv:2311.04780</em></a> (2023).</p> </blockquote> <p>If you found this dataset useful or used it in your research, please cite this reference.</p> <p>This dataset contains manual quality annotations and image quality metrics (IQMs) obtained from 1647 stacks of T2-weighted (T2w) slices of fetal brain magnetic resonance (MR) images collected from 233 subjects at four different institutions Lausanne University Hospital (CHUV) in Switzerland, BCNatal at Hospital Sant Joan de D&eacute;u in Barcelona (Spain), University Children's Hospital Z&uuml;rich (KISPI) in Switzerland and La Timone University Hospital in Marseille, France. The data were acquired on scanners from different vendors (Siemens at CHUV, BCNatal and Marseille, General Electrics at KISPI), MR sequences (Half Fourier Single-shot Turbo spin-Echo &ndash;HASTE&ndash; for Siemens scanners and Single-Short Fast Spin Echo &ndash;SS-FSE&ndash; for GE scanners), magnetic field strengths (1.5 T and 3 T), image resolutions, fields of view, repetition times and echo times, with both neurotypical and pathological cases.</p> <p>These data and the derived IQMs were used to train and evaluate models for quality assessment and quality control of fetal brain MR images. The code to reproduce the experiments is available on <a href="https://github.com/Medical-Image-Analysis-Laboratory/fetal_brain_qc">GitHub.</a></p> <p>Each entry describe the information for a single stack of T2w slices. It contains information regarding which subject it belongs to, its manual quality rating, scanner-related information and 332 IQMs, starting at the `centroid` column in the file. Further description of the data is available in the materials and methods section of the <a href="https://arxiv.org/abs/2311.04780">paper</a>.</p> <p>The model is a 2D nnUNet [1] segmentation network trained on the super-resolution reconstructed data and manual segmentations available as part of the<a href="https://www.synapse.org/#!Synapse:syn25649159/wiki/610007"> Fetal Tissue Annotation Challenge</a> (FeTA).</p> <p>Copyright (c) - All rights reserved. Medical Image Analysis Laboratory - Department of Radiology, Lausanne University Hospital (CHUV) and University of Lausanne (UNIL), Lausanne, Switzerland &amp; CIBM Center for Biomedical Imaging. 2023.</p> <p>[1] Isensee, Fabian, et al. "nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation."&nbsp;<em>Nature methods</em> 18.2 (2021): 203-211.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics

<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims&nbsp;</h1> <div> <h2>&nbsp;</h2> <h2>Description&nbsp;</h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias.&nbsp;&nbsp;</p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias".&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>The label was determined by a 75% majority vote. We classified &ldquo;probably biased&rdquo; and &ldquo;confident biased&rdquo; as biased, and &ldquo;confident not biased,&rdquo; &ldquo;probably not biased,&rdquo; and &ldquo;don't know&rdquo; as not biased.&nbsp;</p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India.&nbsp;</p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022.&nbsp;&nbsp; The percentage of biased tweets with the keyword 'Blacks' remained nearly the same.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews".&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Content</h2> <h3>Cohort 2022&nbsp;</h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias.&nbsp;&nbsp;</p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and &ndash;&ndash;as mentioned above&ndash;&ndash;146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims.&nbsp;&nbsp;</p> </div> <div> <h3>Cohort 2023&nbsp;</h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords &ldquo;Asians, Blacks, Jews, Latinos and Muslims&rdquo; from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias.&nbsp;</p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims.&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns:&nbsp;&nbsp;</p> <p>'TweetID': Represents the tweet ID.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed).&nbsp;&nbsp;</p> </div> <div> <p>'CreateDate': Represents the date the tweet was created.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>&nbsp;&lsquo;Cohort&rsquo;: Represents the year the data was annotated (class of 2022 or class of 2023)&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Acknowledgements&nbsp; &nbsp;</h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy.&nbsp;</p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;</p> </div> <div> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Kinodata-3D: an in silico kinase-ligand complex dataset for kinase-focused machine learning.

<p><strong>Project Description</strong></p> <p>Drug discovery pipelines nowadays rely on machine learning models to explore and evaluate large chemical spaces. While the inclusion of 3D complex information is considered to be beneficial, structural ML for affinity prediction suffers from data scarcity.&nbsp;<br>We provide kinodata-3D, a dataset of <strong>~138 000</strong> docked complexes to enable more robust training of 3D-based ML models for kinase activity prediction (see <a href="https://github.com/volkamerlab/kinodata-3D-affinity-prediction">github.com/volkamerlab/kinodata-3D-affinity-prediction</a>).</p> <h2>Dataset</h2> <h3>1. Data</h3> <p>This data set consists of three-dimensional protein-ligand complexes that were generated using computational docking from the OpenEye toolkit. The modeled proteins cover the kinase family for which a fair amount of structural data, i.e. co-crystallized protein-ligand complexes in the PDB, enriched through KLIFS annotations, is available. This enables us to use template docking (OpenEye&rsquo;s POSIT functionality) in which the ligand placement is guided according to a similar co-crystallized ligand pose. The kinase-ligand pairs to dock are sourced from binding assay data via the public ChEMBL archive, version 33. In particular, we use kinase activity data as curated through the&nbsp;<a href="https://github.com/openkinome/kinodata">OpenKinome kinodata</a> project. The final protein-ligand complexes are annotated with a predicted RMSD of the docked poses. The RMSD model is a simple neural network trained on a <a href="https://github.com/openkinome/kinase-docking-benchmark">kinase-docking benchmark</a> data set using ligand (fingerprint) similarity, docking score (ChemGauss 4), and Posit probability (see <a href="https://github.com/volkamerlab/kinodata-3D" target="_blank" rel="noopener">kinodata-3D repository</a>).</p> <p>The final data set contains in total&nbsp;<strong>138 286</strong> deduplicated kinase-ligand pairs, covering <strong>~98 000</strong> distinct compounds and ~<strong>271</strong> distinct kinase structures.</p> <h3>2. File structure</h3> <p>The archive <strong>kinodata_3d.zip&nbsp;</strong>uses the following file structure</p> <blockquote> <p>data/raw<br>&nbsp;|&nbsp; kinodata_docked_with_rmsd.sdf.gz<br>&nbsp;|&nbsp; pocket_sequences.csv<br>&nbsp;|&nbsp; mol2/pocket<br>&nbsp;&nbsp;&nbsp;&nbsp; | 1_pocket.mol2<br>&nbsp;&nbsp;&nbsp;&nbsp; | ...</p> </blockquote> <p>The file <strong>kinodata_docked_with_rmsd.sdf.gz</strong> contains the docked ligand poses and the information on the protein-ligand pair inherited from <em>kinodata</em>. The protein pockets located in <strong>mol2/pocket</strong> are stored according to the MOL2 file format.</p> <p>The pocket structures were sourced from KLIFS (<a href="https://klifs.net" target="_blank" rel="noopener">klifs.net)</a> and complete the poses in the aforementioned SDF file. The files are named <strong>{klifs_structure_id}_pocket.mol2</strong>. The structure ID is given in the SDF file along with the ligand poses.</p> <p>The file <strong>pocket_sequences.csv&nbsp;</strong>contains all KLIFS pocket sequences relevant to the kinodata-3D dataset.</p> <h3>3. Related code</h3> <p>The code used to create the poses can be found in the <a href="https://github.com/volkamerlab/kinodata-3D" target="_blank" rel="noopener">kinodata-3D repository</a>. The docking pipeline makes heavy use of the <a href="https://github.com/openkinome/kinoml" target="_blank" rel="noopener">kinoml</a> framework, which in turn uses <a href="https://www.eyesopen.com" target="_blank" rel="noopener">OpenEye's</a> Posit template docking implementation. The details of the original pipeline can also be found in the manuscript by <a href="https://www.biorxiv.org/content/10.1101/2023.09.11.557138v1">Schaller et al. (<strong>2023</strong>). Benchmarking Cross-Docking Strategies for Structure-Informed Machine Learning in Kinase Drug Discovery. <em>bioRxiv</em>.</a></p>

openmit-licenseMar 2024View details →
zenodo40/100

Datasets for "Machine-Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks"

<p>This collection contains the datasets and associated files used in the research presented in the manuscript titled "Machine Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks". The datasets are critical for the development and validation of machine learning and symbolic regression models aiming to predict methane storage capacities in covalent organic frameworks (COFs).</p> <p><strong>Included Datasets:</strong></p> <ol> <li><code>COF_Data_for_ML.csv</code>: This dataset was utilized for the development of machine learning models.</li> <li><code>COF_Data_for_SISSO.csv</code>: This dataset was employed for the development of SISSO-based symbolic regression models.</li> <li><code>ML_vs_GCMC.xlsx</code>: This comparative dataset features GCMC-calculated results alongside machine learning predictions.</li> <li><code>Feature_Combination.xlsx</code>: This file contains data detailing all the feature combinations explored in the study.</li> <li><code>ML_SISSO_GCMC.xlsx</code>: This comparative dataset includes GCMC calculations, SISSO-based symbolic regression model predictions, and ML predictions.</li> <li><code>Crystallographic_Properties_of_535k_COFs.xlsx</code>: This consolidated dataset presents the crystallographic properties of 535,293 COFs.</li> </ol> <p><strong>Software Used:</strong></p> <ul> <li>Machine Learning Computations: Scikit-Learn (<a href="https://scikit-learn.org/stable/" target="_new">https://scikit-learn.org/stable/</a>)</li> <li>GCMC Simulations: RASPA2 (<a href="https://github.com/iRASPA/RASPA2" target="_new">https://github.com/iRASPA/RASPA2</a>)</li> <li>SISSO Calculations: SISSO toolkit (<a href="https://github.com/rouyang2017/SISSO" target="_new">https://github.com/rouyang2017/SISSO</a>)</li> <li>Crystallographic property calculations: Zeo++ (<a href="https://www.zeoplusplus.org/" target="_new">https://www.zeoplusplus.org/</a>)</li> </ul> <p>The datasets are provided to enable replication of the study's findings, encourage further research in the field, and facilitate the development of advanced predictive models by the scientific community. Researchers who use these datasets are requested to cite this Zenodo entry as well as the associated paper upon its publication.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Dataset related to the article titled "Prediction of elastic modulus of basaltic rocks using machine learning methods."

<p>Dataset related to the article titled "Prediction of elastic modulus of basaltic rocks using machine learning methods."</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Small dataset machine-learning approach for efficient design space exploration: engineering ZnTe-based high-entropy alloys for water splitting

<p>Atomic structure data used in the research article entitled "Small Dataset Machine-Learning Approaches to Explore the Design Space of High-Entropy Alloys: Engineering ZnTe-based Multicomponent Alloys for the Photo-Splitting of Water"</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Architectural Design Decisions for the Machine Learning Workflow: Dataset and Code

<p><strong>Title:</strong> Architectural Design Decisions for the Machine Learning Workflow: Dataset and Code</p> <p><strong>Authors:</strong> Stephen John Warnett; Uwe Zdun</p> <p><strong>About:</strong> This is the dataset and code artifact for the article entitled &quot;Architectural Design Decisions for the Machine Learning Workflow&quot;.</p> <p><strong>Contents:</strong> The &quot;_generated&quot; directory contains the generated results, including latex files with tables for use in publications and the Architectural Design Decision model in textual and graphical form. &quot;Generators&quot; contains Python applications that can be run to generate the above. &quot;Metamodels&quot; contains a Python file with type definitions. &quot;Sources_coding&quot; contains our source codings and audit trail. &quot;Add_models&quot; contains the Python implementation of our model and source codings. Finally, &quot;appendix&quot; contains a detailed description of our research method.</p> <p><strong>Article Abstract:&nbsp;</strong>Bringing machine learning models to production is challenging as it is often fraught with uncertainty and confusion, partially due to the disparity between software engineering and machine learning practices, but also due to knowledge gaps on the level of the individual practitioner. We conducted a qualitative investigation into the architectural decisions faced by practitioners as documented in gray literature based on Straussian Grounded Theory and modeled current practices in machine learning. Our novel Architectural Design Decision model is based on current practitioner understanding of the topic and helps bridge the gap between science and practice, foster scientific understanding of the subject, and support practitioners via the integration and consolidation of the myriad decisions they face. We describe a subset of the Architectural Design Decisions that were modeled, discuss uses for the model, and outline areas in which further research may be pursued.</p> <p><strong>Objective:</strong> This article aims to study current practitioner understanding of architectural concepts associated with data processing, model building, and Automated Machine Learning (AutoML) within the context of the machine learning workflow.</p> <p><strong>Method:</strong> Applying Straussian Grounded Theory to gray literature sources containing practitioner views on machine learning practices, we studied methods and techniques currently applied by practitioners in the context of machine learning solution development and gained valuable insights into the software engineering and architectural state of the art as applied to ML.</p> <p><strong>Results:</strong> Our study resulted in a model of Architectural Design Decisions, practitioner practices, and decision drivers in the field of software engineering and software architecture for machine learning.</p> <p><strong>Conclusions:</strong> The resulting Architectural Design Decisions model can help researchers better understand practitioners&#39; needs and the challenges they face, and guide their decisions based on existing practices. The study also opens new avenues for further research in the field, and the design guidance provided by our model can also help reduce design effort and risk. In future work, we plan on using our findings to provide automated design advice to machine learning engineers.</p>

openapache2.0Nov 2021View details →
zenodo40/100

Datasets for "Unexplored Antarctic meteorite collection sites revealed through machine learning"

<p>This archive provides datasets related to the following publication:</p> <p>V. Tollenaar, H. Zekollari, S. Lhermitte, D. Tax, V. Debaille, S. Goderis, P. Claeys, F. Pattyn, Unexplored Antarctic meteorite collection sites revealed through machine learning. Science Advances 8, eabj8138 (2022). <a href="https://doi.org/10.1126/sciadv.abj8138">DOI: 10.1126/sciadv.abj8138</a></p> <p>Contact: Veronica Tollenaar, Veronica.Tollenaar@ulb.be</p> <p>Users should cite the original publication when using all or part of the data.&nbsp;</p> <p>About the datasets: it includes a shapefile with the outline of the 613 Meteorite Stranding Zones (Fig. 7, &quot;613MSZs.zip&quot;), the observations used for classification, and the continent-wide probability to find meteorites (at 450-meter resolution, Fig. 5, &quot;positive_classified.nc&quot;). References to the literature are provided in the corresponding publication. Meteorite locations are based on the Meteoritical Bulletin Database (available at https://www.lpi.usra.edu/meteor/).</p> <p>- bias_above200m1kmbuff_expanded_dissolved: shapefile of polygons of unlabelled observations<br> - meteorite_locations_raw.csv: contains locations of meteorite finds as defined in the meteoritical bulletin consulted on 05/07/2019<br> - meteorite_types.csv: contains meteorite names and types as defined in the meteoritical bulletin consulted on 05/07/2019<br> - validation_neg.csv: contains locations of negative observations used for validation<br> - TEST_neg.csv: contains locations of negative test observations<br> - TEST_pos.csv: contains locations of positive test obesrvations<br> - MSZs_ranked: shapefile of ranked meteorite stranding zones<br> - Test_neg4326: shapefile of locations used as negative test data<br> - Cal_neg4326: shapefile of locations used as negative calibration/validation data<br> - TestMSZs_pos4326: shapefile of locations used as positive test data in MSZ-level assesment<br> - 613MSZs: shapefile of outlines of meteorite stranding zones<br> - positive_classified.nc: netcdf of positive classified observations with their estimated a posteriori probabilities</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record