Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
16
datasets available to search
ShareScore release 0.9.0
Dataset results
16 results for “Multi-Task Learning”
MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification
<p>Recent research in disaster informatics demonstrates a practical and important use case of artificial intelligence to save human lives and suffering during natural disasters based on social media contents (text and images). While notable progress has been made using texts, research on exploiting the images remains relatively under-explored. To advance image-based approaches, we propose MEDIC\footnote{Available~at: \url{https://crisisnlp.qcri.org/medic/index.html}}, which is the largest social media image classification dataset for humanitarian response consisting of 71,198 images to address four different tasks in a multi-task learning setup. This is the first dataset of its kind: social media images, disaster response, and multi-task learning research. An important property of this dataset is its high potential to facilitate research on \textit{multi-task learning}, which recently receives much interest from the machine learning community and has shown remarkable results in terms of memory, inference speed, performance, and generalization capability. Therefore, the proposed dataset is an important resource for advancing image-based disaster management and multi-task machine learning research. <br> </p>
Multi-Task Regression-based Learning for Autonomous Unmanned Aerial Vehicle Flight Control within Unstructured Outdoor Environments [dataset]
<p>This dataset is related to "Multi-Task Regression-based Learning for Autonomous Unmanned Aerial Vehicle Flight Control within Unstructured Outdoor Environments" in IEEE RA-L,2019.</p> <p> </p> <p>Data Capture<br> ========================<br> Data is obtained by manually flying the UAV through the redwood forest environment using a FrSky Taranis (Plus) Digital Telemetry Radio System. In total, 81,674 frames were captured together with the flight behaviour that comprehends flights under and above the forest canopy, navigation inside caves and on river beds, lakes and mountains.</p> <p> </p> <p>Folder Structure<br> ========================<br> |-manual_0 - manual_5: sequences containing training data</p> <p>|-test_0 - sequences containing testing data</p> <p> </p> <p>Data Protection<br> ========================<br> Gathered by simulated flight using Microsoft AirSim (2019) and released in accordance with MSR Aerial Information and Robotics Simulator (AirSim) lisence, which is described in details bellow:</p> <p> </p> <blockquote> <p>The MIT License (MIT)</p> <p>MSR Aerial Informatics and Robotics Platform<br> MSR Aerial Informatics and Robotics Simulator (AirSim)<br> Copyright (c) Microsoft Corporation<br> All rights reserved.<br> MIT License</p> <p>Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the ""Software""), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:<br> The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.<br> THE SOFTWARE IS PROVIDED *AS IS*, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p> </blockquote>
MTL-QA : A dataset and multi-task learning approach for knowledge graph and natural language question answering
<p>The dataset used for this project is created by enhancing the publicly available MetaQA (Movie Text Audio QA), which is primarily a KGQA dataset pertaining to movies, an extension of WikiMovies. This involves questions requiring 1, 2, and 3 hops which can be answered by using a MetaQA Knowledge Graph. The questions are available in text and audio format. The text has vanilla (original) and its paraphrased version, and is called ntm. </p> <p>In order to develop a dataset to support NLQA, a series of dataset augmentation steps has been performed.</p> <p>The dataset consists of natural language questions and a tagged topic entity as ground truth. This topic entity is used to retrieve textual information related to the question from Wikipedia. The introduction section of the entity's page is used as the context that is required for NLQA. Hence, this dataset has information related to both KGQA and NLQA. Certain preliminary checks and validations are done to only retain those data samples whose context can be used to answer a given question.</p>
scPretrain: Multi-task self-supervised learning for cell type classification
<p>The dataset and code for paper, scPretrain: Multi-task self-supervised learning for cell type classification.</p>
Multi-task self-supervised learning for wearables - human activity recognition
<p>Datasets used to train and evaluated the self-supervised-learning model</p>
DCP-MTL: Vectorization of Agricultural Cultivation Field Parcels via Boundary-Parcel Multi-Task Learning Network in Ultra-High-Resolution Remote Sensing Images
<p><span>This paper introduces the first UHR UAV dataset specifically for CFP, designed to evaluate the performance of the proposed model in identifying these parcels. </span><span>The dataset offers ultra-high spatial resolution, various field parcel types, and broad geographic coverage. </span><span>Figure 8 </span><span>shows </span><span>the spatial distribution of the study data. Jilin Province is the primary region for training and evaluating the model, while Hebei, Henan, Anhui, Zhejiang, and Hainan are auxiliary regions for testing the model's transferability. </span></p>
Data sets and machine learning models for: Predicting critical properties and acentric factor of fluids using multi-task machine learning
<p>The experimental data sets, data splits, additional features, QM calculations, model predictions, and final machine learning models for the manuscript "Predicting Critical Properties and Acentric Factor of Fluids Using Multi-Task Machine Learning". <strong>Citation should refer directly to the manuscript:</strong></p> <ul> <li> <p>Biswas, S.; Chung, Y.; Ramirez, J.; Wu, H.; Green, W. H. Predicting Critical Properties and Acentric Factors of Fluids Using Multitask Machine Learning. <em>Journal of Chemical Information and Modeling.</em> <strong>2023</strong> <em>63</em> (15), 4574-4588. DOI: <a href="https://doi.org/10.1021/acs.jcim.3c00546">10.1021/acs.jcim.3c00546</a></p> </li> </ul> <p>To use the machine learning models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a>. </p> <p>Detailed information can be found in README.md file.</p> <p> </p> <p><strong>Details on the properties considered</strong></p> <p>The data set includes the following 8 properties:</p> <ul> <li>Tc: critical temperature, in K</li> <li>Pc: critical pressure, in bar</li> <li>rhoc: critical density, in mol/L</li> <li>omega: acentric factor, unitless</li> <li>Tb: boiling point, in K</li> <li>Tm: melting point, in K</li> <li>dHvap: enthalpy of vaporization at boiling point, in kJ/mol</li> <li>dHfus: enthalpy of fusion at melting point, in kJ/mol</li> </ul> <p><strong>Details on the files</strong></p> <p>1. Data sets under CritProp_v1.1.0:</p> <ul> <li>all_data: includes the data sets used in this work. All data points are listed for each chemical compound as well as its corresponding data source. The details of the data sources can be found in the README.md file. The distribution of the data set is included in each folder. <ul> <li>estimated_data_for_pretraining: contains the estimated data from Yaws' handbook that are used to pre-train our machine learning (ML) model.</li> <li>experimental_data: contains the experimental data (references 1 - 15) used to fine-tune our final ML model.</li> </ul> </li> <li>additional_features: includes the additional features tested for the ML model. The Abraham features are generated for all data (references 1 - 15) while the acsf, qm, and rdkit features are only generated for the data from references 1 - 9. <ul> <li>abraham: Abraham solute parameters (E, S, A, B, L). Molecular features.</li> <li>acsf: ACSF (atom-centered symmetry functions). Atomic features that are coverted from the 3D coordinates of the compound</li> <li>qm_atom: QM (quantum chemical) atomic feature. </li> <li>qm_mol: QM molecular feature.</li> <li>rdkit: Selected RDKit 2D molecular features.</li> </ul> </li> <li>data_splits_and_model_predictions: contains the training and test sets used to evaluate the model. It also contains the predicted values from our final ML model for each test set. <ul> <li>random and scaffold splits: training and test sets that include the data from references 1 - 9.</li> <li>external test set: a test set that includes the data from only references 10 - 15.</li> </ul> </li> </ul> <p>2. Machine learning (ML) model files:</p> <ul> <li>CritProp_ML_model_files_with_abraham_feat.zip: contains the Chemprop ML model files that are trained using Abraham features as additional molecular features. This gives the best results.</li> <li>CritProp_ML_model_files_without_additional_feat.zip: contains the Chemprop ML model files that are trained without any additional features. This gives the second best results.</li> </ul> <p>To use these ML models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a></p> <p>3. QM (quantum chemical) calculations:</p> <ul> <li>QM_calculations.zip: contains the results of the QM calculations that are performed to compute QM features.</li> </ul> <p> </p> <p> </p>
Processed data for the manuscript: Promoting Multi-Task Learning as a General Approach for Deep-Learning-based Hydrological Models
<div> <div>Below is a brief overview of the processed data in this repository:</div> <br> <div>- camels_streamflow: This directory contains streamflow data for CAMELS basins covering the period from January 1, 2015, to December 31, 2021. We have not included the original CAMELS dataset, which contains attributes, meteorological forcing, and streamflow data from January 1, 1980, to December 31, 2014, as it can be easily downloaded from the CAMELS website (https://gdex.ucar.edu/dataset/camels.html) and is too large for us to upload to Zenodo.</div> <div>- modiset4camels: This directory includes multiple versions of basin-mean Evapotranspiration (ET) data retrieved from the MOD16A2 data product. The dataset spans from January 1, 2001, to December 31, 2021, with an 8-day temporal resolution.</div> <div>- nldas4camels: This directory contains basin-mean daily meteorological forcing data from the NLDAS-2 dataset, obtained via Google Earth Engine (GEE). The dataset covers the period from January 1, 2001, to December 31, 2021.</div> <div>- smap4camels: This directory features basin-mean Soil Moisture (SSM) data from the NASA-USDA Enhanced SMAP Global Soil Moisture dataset, covering the period from April 2, 2015, to October 3, 2021. The dataset provides SSM measurements at a 5 cm depth. Additionally, we provide basin-mean daily SMAP L4 data spanning from April 1, 2015, to December 31, 2023.</div> </div>
DrugBLIP: Exploring the Protein-Molecule Interaction Mechanisms with a Multi-task Learning Graph Transformer
Open the record for dataset details and reuse information.
Dataset for Multi-Task Learning for Simultaneous Retrievals of Passive Microwave Precipitation Estimates and Rain/No-Rain Classification
<p>Dataset for Multi-Task Learning for "Simultaneous Retrievals of Passive Microwave Precipitation Estimates and Rain/No-Rain Classification".</p> <p>Global Precipitation Measurement (GPM) dual-frequency precipitation radar (DPR) and Goddard Profiling Algorithm (GPROF) were provided from NASA Global Precipitation Measurement Precipitation Data Directory (https://gpm.nasa.gov/data/directory). The original data used for this study have been supplied by JAXA’s GSMaP.</p> <p>The source code for preprocessing and model training is available at https://doi.org/10.5281/zenodo.7627112</p>
A_Dataset for Multi-Task Learning for Simultaneous Retrievals of Passive Microwave Precipitation Estimates and Rain/No-Rain Classification
<p>Dataset for Multi-Task Learning for "Simultaneous Retrievals of Passive Microwave Precipitation Estimates and Rain/No-Rain Classification".</p> <p>Global Precipitation Measurement (GPM) dual-frequency precipitation radar (DPR) and Goddard Profiling Algorithm (GPROF) were provided from NASA Global Precipitation Measurement Precipitation Data Directory (https://gpm.nasa.gov/data/directory). The original data used for this study have been supplied by JAXA’s GSMaP.</p> <p>The source code for preprocessing and model training is available at https://doi.org/10.5281/zenodo.7627112.</p>
SOUND-BASED DRONE FAULT CLASSIFICATION USING MULTI-TASK LEARNING
<p>arxiv : https://arxiv.org/abs/2304.11708</p> <p>Accepted at 29th International Congress on Sound and Vibration (ICSV29). </p> <p>The drone has been used for various purposes including military applications, aerial photography, and pesticide spraying. However, the drone is vulnerable to external disturbances, and malfunction in propellers and motors can easily occur. To improve the safety of drone operations, early detection of mechanical faults should be made in real-time. In this paper, we propose a sound-based deep neural network (DNN) fault classifier and drone sound dataset. The dataset was constructed by collecting the operating sounds of drones from microphones mounted on three different drones in an anechoic chamber. The dataset includes various operating conditions of drones, such as flight directions (front, back, right, left, clockwise, counter clockwise) and faults on propellers and motors. The drone sounds were then mixed with noises recorded in five different spots on the university campus, with a signal-to-noise ratio (SNR) varying from 10 dB to 15 dB. Using the acquired dataset, we train a DNN classifier, 1DCNN-ResNet, that classifies the types of mechanical faults and their locations from short-time input waveforms. We employ multitask learning (MTL) and incorporate the direction classification task as an auxiliary task to make the classifier learn more general audio features. The test over unseen data reveals that the proposed multitask model can successfully classify faults in drones and outperforms single-task models even with less training data. </p> <p> </p> <p>please reorganize the file directory like below</p> <p>drone</p> <p>ㄴA</p> <p>ㄴB</p> <p>ㄴC</p> <p> </p> <p>For each drone type A, B, and C have 54000*2 files. (Here, *2 means stereo channel, you can find mic1 and mic2 in subdirectory) They are divided into train, valid, and test by a 6:2:2 ratio. For each file, recording information is labeled below.</p> <p>{model_type}_{maneuvering_direction}_{fault}_{drone_file_index}_{background}_{background_file_index}_{SNR}</p> <p>model_type: A, B, C</p> <p>maneuvering_direction: F(Front), B(Back), R(Right), L(Left), C(Clockwise), CC(Counter-clockwise)</p> <p>fault: N (Normal), MF1~4 (Moter Failure), PC1~4 (Propeller Cut) -> 1~4 means each motor/propeller of the quadcopter.</p> <p> </p>
MultiTune: Multiple-Environment Configuration Tuning via Multi-Task Learning and Propensity Score Matching
Open the record for dataset details and reuse information.
The data of Multi-task learning aids in assessing microbial genome quality
Open the record for dataset details and reuse information.
Multi-task learning uncovers robust translation cis-regulatory features
GEO Series GSE201766. Homo sapiens. 1 samples. Type: Other.
Skillful bias correction of offshore near-surface wind speed and wind direction forecasting based on a multi-task machine learning model
<h3>Dataset</h3> <p>1. observation data over 14 weather stations</p> <p>Variables: hourly near-surface 2-min average wind speed, wind direction </p> <p>2. ECMWF-IFS forecast data over 14 weather stations</p> <p>Variables: hourly predictors at surface level and upper level in next 48 hours (shown in Table 1. and Table 2.)</p> <p>Table 1. ECMWF-IFS forecast data at surface level</p> <div> <table> <tbody> <tr> <td> <p>Predictors</p> </td> <td> <p>Abbreviation</p> </td> <td> <p>Unit</p> </td> </tr> <tr> <td> <p>Temperature at 2 m</p> </td> <td> <p>2t</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Sea surface temperature</p> </td> <td> <p>sst</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Dewpoint temperature at 2 m</p> </td> <td> <p>2d</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Convective precipitation in the past hour</p> </td> <td> <p>cp</p> </td> <td> <p>mm</p> </td> </tr> <tr> <td> <p>Mean sea level pressure</p> </td> <td> <p>msl</p> </td> <td> <p>hPa</p> </td> </tr> <tr> <td> <p>Zonal component of wind speed at 10 m</p> </td> <td> <p>10u</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Meridional component of wind speed at 10 m</p> </td> <td> <p>10v</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind speed at 10 m</p> </td> <td> <p>10ws</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind direction at 10 m</p> </td> <td> <p>10wd</p> </td> <td> <p>°</p> </td> </tr> <tr> <td> <p>Zonal component of wind speed at 100 m</p> </td> <td> <p>100u</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Meridional component of wind speed at 100 m</p> </td> <td> <p>100v</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind speed at 100 m</p> </td> <td> <p>100ws</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind direction at 100 m</p> </td> <td> <p>100wd</p> </td> <td> <p>°</p> </td> </tr> </tbody> </table> </div> <div> </div> <p>Table 2. ECMWF-IFS forecast data at upper level</p> <table> <tbody> <tr> <td> <p>Predictors</p> </td> <td> <p>Abbreviation</p> </td> <td> <p>Unit</p> </td> </tr> <tr> <td> <p>Relative humidity at xxx hPa</p> </td> <td> <p>r_Lxxx</p> </td> <td> <p>%</p> </td> </tr> <tr> <td> <p>Temperature at xxx hPa</p> </td> <td> <p>t_Lxxx</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Vertical velocity of wind at xxx hPa</p> </td> <td> <p>w_Lxxx</p> </td> <td> <p>Pa s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Zonal component of wind at xxx hPa</p> </td> <td> <p>u_Lxxx</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Meridional component of wind at xxx hPa</p> </td> <td> <p>v_Lxxx</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind speed at xxx hPa</p> </td> <td> <p>ws_Lxxx</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind direction at xxx hPa</p> </td> <td> <p>wd_Lxxx</p> </td> <td> <p>°</p> </td> </tr> </tbody> </table> <div> </div> <p>3. key variables constructed by feature engineering</p> <p>(1) sort-term statistics, including <em>maximum, minimum, mean </em>and <em>variance</em> of key variables (<em>2t</em>,<em> 10u</em>, <em>10v </em>and <em>10ws</em>) from ECMWF-IFS model during the next 48 hours,</p> <p> (2) long-term statistics, including <em>mean </em>and <em>deviation</em> of key variables (<em>2t</em>,<em> 10u</em>, <em>10v </em>and <em>10ws</em>) from ECMWF-IFS model during history 3-yr period (January 2020–December 2022),</p> <p> (3) thermodynamic factors, including the low-level wind shear between <em>10ws</em> and <em>100ws</em>, vertical wind shear between 200 hPa and 850 hPa<em>, </em>the differences between <em>sst</em><em> </em>and <em>2t</em><em>.</em></p> <h3>Scripts</h3> <p>1. Random Forest model training code</p> <p>2. LightGBM model training code</p> <p>3. XGBoost model training code</p> <p>4. TabNet-MTL model training code</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.