Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
66
datasets available to search
ShareScore release 0.9.0
Dataset results
66 results for “Multi-source”
Data from: Long-term trends in the occupancy of ants revealed through use of multi-sourced datasets
Open the record for dataset details and reuse information.
A STP-HSI index method for urban built-up area extraction based on multi-source remote sensing data
Open the record for dataset details and reuse information.
MGP: a new 1-hourly 0.25° global precipitation product (2000-2020) based on multi-source precipitation data fusion
<p>The multi-source merged global precipitation product (MGP; 0.25°/ hourly; 2000-2020; 60°N/S), which takes advantage of the complementary strengths of satellite, reanalysis, and gauge data to obtain reliable precipitation estimates over the global land surface, provides a new higher-quality precipitation dataset for data users to realize their respective research purposes and the social and economic activities. The data developers hope that MGP will play an important role in a variety of science communities (e.g., hydrology, meteorology, climatology, ecology, and agriculture).</p>
Typhoon track tracking and forecasting algorithm based on multi-source information
<p>Typhoon track tracking and forecasting algorithm based on multi-source information</p>
MosqVision-3K: A Balanced Multi-Source Dataset of 3,000 Annotated Images for Culex, Anopheles, and Aedes Mosquito Species Classification
<p><strong>Comprehensive Mosquito Species Image Dataset for Machine Learning</strong><br><strong>Description:</strong><br>This dataset is a meticulously curated collection of high-quality images featuring three major mosquito species: <strong>Culex</strong>, <strong>Anopheles</strong>, and <strong>Aedes</strong>. These species are significant vectors for transmitting vector-borne diseases such as malaria, dengue, and Zika. The dataset has been compiled to support research and development in entomology, vector-borne disease control, and image recognition.<br>With <strong>3,000 images in total</strong>, the dataset is structured to ensure a balanced representation of the three species, each having <strong>1,000 images</strong>. Images were sourced from four reputable platforms, including <strong>MosquitoAlert.com</strong>, <strong>Mendeley Data</strong>, <strong>IEEE DataPort</strong>, and the <strong>Dryad Digital Repository</strong>. These sources ensure a comprehensive and diverse representation of mosquito appearances, including variations in morphology, lighting conditions, and orientations.</p> <h2>The dataset is organized into directories for each species, making it easy to integrate into machine learning workflows for tasks like species identification and classification. The collection also includes metadata and annotations to enhance usability.</h2> <p><strong>Key Features:</strong></p> <ul> <li><strong>Species Represented:</strong> <ul> <li><em>Culex</em></li> <li><em>Anopheles</em></li> <li><em>Aedes</em></li> </ul> </li> <li><strong>Total Images:</strong> 3,000 (1,000 images per species)</li> <li><strong>Image Sources:</strong> <ul> <li><strong>MosquitoAlert.com</strong> (1,234 images)</li> <li><strong>Mendeley Data</strong> (876 images)</li> <li><strong>IEEE DataPort</strong> (748 images)</li> <li><strong>Dryad Digital Repository</strong> (600 images)</li> </ul> </li> </ul> <h2>- <strong>Image Annotations:</strong> Metadata and species labels are included for enhanced usability.</h2> <p><strong>Applications:</strong><br>This dataset is ideal for a variety of applications, including:</p> <ul> <li>Training machine learning models for mosquito species identification.</li> <li>Developing computer vision algorithms for pest control and public health.</li> </ul> <h2>- Enhancing vector control strategies to mitigate disease spread.</h2> <p><strong>Data Structure:</strong><br>The dataset is organized as follows:</p> <pre><code>Mosquito_Dataset/ ├── Anopheles/ │ ├── img_001<span>.jpg</span> │ ├── img_002<span>.jpg</span> │ └── ... ├── Aedes/ │ ├── img_001<span>.jpg</span> │ ├── img_002<span>.jpg</span> │ └── ... └── Culex/ ├── img_001<span>.jpg</span> ├── img_002<span>.jpg</span> └── ... </code></pre> <h2> </h2> <h2><strong>Acknowledgments:</strong><br>We acknowledge the following data sources for their contributions:</h2> <blockquote> <ul> <li>MosquitoAlert.com</li> <li>Mendeley Data</li> <li>IEEE DataPort</li> <li> Dryad Digital Repository</li> </ul> </blockquote>
Dynamic implicit modeling of tunnel unfavorable geology based on multi-source data fusion using support vector machine
<p>This is the relevant data of the article "Dynamic implicit modeling of tunnel unfavorable geology based on multi-source data fusion using support vector machine"</p>
Water level of Qinghai Lake based on multi-source satellite altimetry data
Open the record for dataset details and reuse information.
FIGURE 3 in Synonymy of two important crop pests of burrower bugs, Cyrtomenus mirabilis and C. bergi (Hemiptera: Cydnidae), based in a multi-source approach
FIGURE 3. Genital structures; (A) external male genitalia C. mirabilis, (B) C. bergi, (C) phallus lateral view C. mirabilis, (D) C. bergi, (E) pygophore dorsal view C. mirabilis, (F) C. bergi, (G) left paramere C. mirabilis, (H) C. bergi, (I) Female external genitalia C. mirabilis, (J) C. bergi. (K) spermatheca C. mirabilis, (L) C. bergi. sX, segment X; am, apical margin; as, apical surface; gcVIII, gonocoxite VIII; gcIX, gonocoxite IX; laVIII, laterotergite VIII; laIX, laterotergite IX; stVII, sternum VII; so, spermathecal opening; rs, ring sclerites; pd, proximal duct; vp, vaginal pouche; di, dilation; in, invagination; dd, distal duct; pf, proximal flange; ip, intermediate part; df, distal flange; nd, "neck" duct; sr, seminal receptacle; dr, dorsal rim; pa, paramere; v, vesica; pv, processus vesicae; 2pc, second conjunctival appendage. Scale bar: 0.5 mm.
FIGURE 5 in Synonymy of two important crop pests of burrower bugs, Cyrtomenus mirabilis and C. bergi (Hemiptera: Cydnidae), based in a multi-source approach
FIGURE 5. Discriminant function analysis plot after cross-validation procedure, based on the discriminant variate of shape analysis and correspondent mean shapes for C. mirabilis and C. bergi of the head (A), pronotum (C), scutellum (E) and hemelytron (G). Multiple-group discriminant function analysis plot based on canonical variates, CV1 and CV 2 for species in each group of latitude for the head (B), pronotum (D), scutellum (F) and hemelytron (H).
GDCLD:A globally distributed dataset of coseismic landslide mapping via multi-source high-resolution remote sensing images
<p>GDCLD : A globally distributed dataset of coseismic landslide mapping via multi-source high-resolution remote sensing images</p> <p>Fang, C., Fan, X., Wang, X., Nava, L., Zhong, H., Dong, X., Qi, J., and Catani, F.: A globally distributed dataset of coseismic landslide mapping via multi-source high-resolution remote sensing images, Earth Syst. Sci. Data Discuss. [preprint], https://doi.org/10.5194/essd-2024-239, in review, 2024.</p> <p> </p> <p>Data description:</p> <p> </p> <p>The training dataset and the validation dataset are composed of UAV, PlanetScope, Gaofen-6 and Map World images of the 5 earthquake regions of Luding, Nippes, Hokkaido, Jiuzhaigou and Mainling. There is no overlapping area in each TIFF. The training dataset and the validation dataset are randomly divided at a ratio of approximately 0.75:0.25.</p> <p> </p> <p>train_dataset:</p> <p>train_data: The train dataset part of the GDCLD data set contains 11162 data matrices (TIFF), and the shape of the matrix is (1024, 1024, 3) (TIFF).</p> <p>train_label: The train dataset part of the GDCLD data set contains 11162 data matrices (TIFF), and the shape of the matrix is (1024, 1024, 1) (TIFF).</p> <p> </p> <p>Validation_dataset</p> <p>val_data: The validation dataset part of the GDCLD data set contains 4459 data matrices (TIFF), and the shape of the matrix is (1024, 1024, 3) (TIFF).</p> <p>val_label: The validation dataset part of the GDCLD data set contains 4459 data matrices (TIFF), and the shape of the matrix is (1024, 1024, 1) (TIFF).</p> <p> </p> <p>Test_dataset (Lushan, Sumatra, Mesetas and Palu dataset)</p> <p>This package contains the original files of remote sensing images from three sources: UAV, Map World, and PlaneScope belonging to the Lushan, Sumatra, Mesetas and Palu earthquake regiones, which are used to display the test area.<br><br>Future work:<br>The future work includes additional landslide data that the authors will continue to upload. In this 2.0 version update, we have added UAV imagery and interpreted data for loess landslides triggered by the December 2023 M6.2 earthquake in Gansu, China, with a resolution of 0.1 m. Due to authorization constraints, we can only provide PNG files without geographic coordinates. Additionally, this update includes PlanetScope imagery of landslides induced by heavy rainfall in Guangdong, China, in 2024, as well as PlanetScope imagery and landslide labels for events triggered by the Hualien earthquake in Taiwan.<br><br>Please note that this landslide dataset is publicly available exclusively for scientific research purposes and must not be used for commercial purposes.</p>
Dataset accompanying "Integrated nowcasting of convective precipitation with Transformer-based models using multi-source data"
<p>Dataset accompanying the article <a href="https://arxiv.org/abs/2409.10367" target="_blank" rel="noopener">Integrated nowcasting of convective precipitation with Transformer-based models using multi-source data</a>. </p> <p>Contains almost 8000 events that are sampled from the summer months (where convective precipitation events are most likely to occur) of 2019-2023, centred over Austria.</p> <p>Each sample has a temporal span of 4 hours with a spatial extent of 400 x 700 km, with following data streams:</p> <ul> <li>4 MSG infrared channels with central wavelengths of 6.2, 7.3, 8.7, and 10.8 μm</li> <li>Rain rates mosaicked from ground-based radar observations</li> <li>Lightning data from ground-based observations</li> <li>INCA precipitation analysis</li> <li>INCA convective available potential energy (CAPE) estimates</li> </ul> <p>The dataset is accompanied by elevation and coordinate information. </p> <p>Please refer to the manuscript and the <a href="https://github.com/caglarkucuk/earthformer-multisource-to-inca">GitHub repository</a> for further information and helper code for reading the data files.</p>
An Integrated Approach for enhanced SMAP Soil Moisture Retrieval: Multi-Source Data Fusion and Data-Driven Machine Learning
<p><span>Accurate satellite-based soil moisture (SM) retrieval is essential for hydrometeorological and agroecological applications, yet traditional physical models for L-band SM retrieval are hindered by uncertainties stemming from inaccuracies in prior parameters. This work combines multi-source data fusion and a physically-guided machine learning framework to develop a Soil Moisture Active Passive (SMAP) SM retrieval model (Fusion-LightGBM, F-LGB) that bypasses the need for static prior parameters, resulting in a new SM product. The retrieval benchmark is a new seamless SM data constructed by combining Triple Collection correlation coefficients (TC-R) and the Maximized-R method, which demonstrates superior temporal correlation on 20 International Soil Moisture Network (ISMN) <em>in-situ</em> networks compared to existing SM data, including ECMWF Reanalysis v5-Land (ERA5-Land), SMAP Level 4 (SMAP L4), and Global Land Data Assimilation System (GLDAS) Noah. The machine learning model incorporates input variables that represent the Tau-Omega model’s radiative transfer process, including brightness temperature, vegetation optical depth, soil temperature, and an external variable for precipitation. In the 2015-2020 validation set, F-LGB demonstrated the highest correlation (mean R = 0.72, significantly surpassing the second-best SMAP-INRAE-BORDEAUX (SMAP-IB) SM and deep neural network (DNN) SM at 0.67) and the lowest ubRMSE (mean value of 0.052 m<sup>3</sup>/m<sup>3</sup>, better than 0.055 m<sup>3</sup>/m<sup>3</sup> for both DNN and SMAP-IB). F-LGB performed well across diverse land covers, vegetation densities, and climates, with SHAP analysis showing H-polarized brightness temperature as crucial, especially in areas with low to moderate vegetation. This new machine learning-based SMAP SM product may improve global satellite-based SM estimation capabilities.</span></p>
Data and Python codes used in "A New Approach to Estimate Total Nitrogen Concentration in a Seasonal Lake Based on Multi-Source Data Methodology"
<p>These are the dataset and python codes used in a manuscript titled "A New Approach to Estimate Total Nitrogen Concentration in a Seasonal Lake Based on Multi-Source Data Methodology"</p>
Randomized Controlled Trial of Multi-Source Feedback to Pediatric Residents
ClinicalTrials.gov study NCT00302783. IPD Sharing: Not stated. Countries: 1. Publications: 1.
B-slim, a Multi-source Digital Super Coach for Sustainable Weight Loss
ClinicalTrials.gov study NCT02595671. IPD Sharing: UNDECIDED. Countries: 1. Publications: 2.
REBECCA-3 Study, Research on Breast Cancer Induced Chronic Conditions Supported by Causal Analysis of Multi-source Data
ClinicalTrials.gov study NCT06435091. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: Improving species distribution models for stream networks by incorporating spatial autocorrelation in multi-sourced datasets: An assessment of Idaho giant salamander status and future risk
Open the record for dataset details and reuse information.
Data from: Live, dead, and fossil mollusks in Florida freshwater springs and spring-fed rivers: taphonomic pathways and the formation of multi-sourced, time-averaged death assemblages
Taphonomic processes are informative about the magnitude and timing of paleoecological changes but remain poorly understood with respect to freshwater invertebrates in spring-fed rivers and streams. We compared taphonomic alteration among freshwater gastropods in live, dead (surficial shell accumulations), and fossil (late Pleistocene-early Holocene in situ sediments) assemblages from two Florida spring-fed systems, the Wakulla and Silver/Ocklawaha Rivers. We assessed taphonomy of two gastropod species: the native <i>Elimia floridensis</i> (n=2504) and introduced <i>Melanoides tuberculata</i> (n=168). We quantified seven taphonomic attributes (aperture condition, color, fragmentation, abrasion, juvenile spire condition, dissolution, and exterior luster) and combined those attributes into a total taphonomic score (TT). Fossil <i>E. floridensis</i> specimens exhibited the greatest degradation (highest TT scores), whereas live specimens of both species were least degraded. Specimens of <i>E. floridensis</i> from death assemblages were less altered than fossil specimens of the same species. Within death assemblages, specimens of <i>M. tuberculata</i> were significantly less altered than specimens of <i>E. floridensis</i>, but highly degraded specimens dominated in both species. Radiocarbon dates on fossils clustered between 9792 and 7087 cal. BP, whereas death assemblage ages ranged from 10,692 to 1173 cal. BP. Possible explanations for the observed taphonomic patterns include: (1) rapid taphonomic shell alteration, (2) prolonged near-surface exposure to moderate alteration rates, and/or (3) introduction of reworked fossil shells into surficial assemblages. Combined radiocarbon dates and taphonomic analyses suggest that all these processes may have played a role in death assemblage formation. In these fluvial settings, shell accumulations develop as a complex mixture of specimens derived from multiple sources and characterized by multi-millennial time averaging. These findings suggest that, when available, fossil assemblages may be more appropriate than death assemblages for assessing pre-industrial faunal associations and recent anthropogenic changes in freshwater ecosystems.
Benchmark Dataset of Cropland Parcel Boundaries from Multi-Source Remote Sensing Imagery for AI application
<p><span>We provide a standardized dataset of diversified parcel labels based on multi-source remote sensing imagery. The parcels included in this dataset are sourced from four countries: the Netherlands, Denmark, Spain, and China, encompassing over 1</span><span>4</span><span>0,000 </span><span>cropland parcel</span><span>. The average parcel size across different study areas ranges from 0.05 hectares to 10 hectares. The remote sensing imagery is derived from more than 10 data sources, including satellites, UAV (unmanned aerial vehicle), and public mapping platforms, with spatial resolutions varying from 0.05 meters to 10 meters.</span></p>
Data from: Assessing urban-scale spatiotemporal heterogeneous metro station coverage using multi-source mobility data
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.