Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
35
datasets available to search
ShareScore release 0.7.1
Dataset results
35 results for “foundation model”
Modeling Foundation Species in Food Webs
Foundation species are basal species that play an important role in determining community composition by physically structuring ecosystems and modulating ecosystem processes. Foundation species largely operate via non-trophic interactions, presenting a challenge to incorporating them into food-web models. Here, we used non-linear, bioenergetic predator-prey models to explore the role of foundation species and their non-trophic effects. We explored four types of models in which the foundation species reduced the metabolic rates of species in a specific trophic position. We examined the outcomes of each of these models for six metabolic rate “treatments” in which the foundation species altered the metabolic rates of associated species by one-tenth to ten times their allometric baseline metabolic rates. For each model simulation, we looked at how foundation species influenced food-web structure during community assembly and the subsequent change in food-web structure when the foundation species was removed. When a foundation species lowered the metabolic rate of only basal species the resultant webs were complex, species-rich, and robust to foundation species removals. On the other hand, when a foundation species lowered the metabolic rate of only consumer species, all species, or no species the resultant webs were species poor and the subsequent removal of the foundation species webs resulted in the further loss of species and complexity. This suggests that in nature we should look for foundation species to predominantly facilitate basal species.
A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys
<p>A multi-type geobody dataset for training SAG model, including channel, paloekarst, salt body, and so on.</p> <p>A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys (<a href="https://arxiv.org/abs/2409.04962">[2409.04962] A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys (arxiv.org)</a>)</p> <p> </p>
Gene embeddings used in GenePT: A Simple But Hard-to-Beat Foundation Model for Genes and Cells Built From ChatGPT
<p>These are the pulled NCBI (and UniProt, when applicable) summaries of genes, as well as the corresponding OpenAI text embeddings (text-embedding-ada-002 and text-embedding-3-large) computed on the summaries. See methods details in Chen and Zou (2024+).</p> <p>The unzipped folder contains four different files: </p> <ol> <li>NCBI_summary_of_genes.json (NCBI gene card summary of human genes)</li> <li>NCBI_UniProt_summary_of_genes.json (NCBI gene card and UniProt protein (when applicable) summary of human genes)</li> <li>GenePT_gene_embedding_ada_text.pickle (a dictionary of numpy array where gene names (upper case) are keys and text-embedding-ada-002 embeddings of the summary in 1. are the values)</li> <li>GenePT_gene_protein_embedding_model_3_text.pickle (a dictionary of numpy array where gene names (upper case) are keys and text-embedding-3-large embeddings of the summary in 1. are the values)</li> </ol> <p>Reference:</p> <p>Chen YT, Zou J. (2024+) GenePT: A Simple But Effective Foundation Model for Genes and Cells Built From ChatGPT. bioRxiv preprint: <a href="https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1">https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1</a>.</p>
Tuning-less Object Naming with a Foundation Model - Data recorded during testing
<p>We implement a real-time object naming system that enables learning a set of named entities never seen. Our approach employs an existing foundation model that we consider ready to see anything before starting. It turns seen images into relatively small feature vectors that we associate with index to a gradually built vocabulary without any training of fine-tuning of the model. Our contribution is using the association mechanism known from transformers as attention. It has features that support generalization from irrelevant information for distinguishing the entities and potentially enable associating with much more than indices to vocabulary. As a result, the system can work in a one-shot manner and correctly name objects named in different contents. We also outline implementation details of the system modules integrated by a blackboard architecture. Finally, we investigate the<br>system's quality, mainly how many objects it can handle in this way.</p>
High-Resolution Canopy Fuel Maps Based on GEDI: A Foundation for Wildfire Modeling in Germany
<p>Open access publication under review.</p> <p>Visit <a href="https://ee-forestfuels-ger.projects.earthengine.app/view/gedi-fuels"><strong>this Earth Engine app</strong></a> to explore the data interactively.</p> <p> </p> <p>Abstract:</p> <p>Forest fuels are essential for wildfire behavior modeling and risk assessments but difficult to quantify accurately. An increase in fire frequency in recent years, particularly in regions traditionally not prone to fire, such as central Europe, has increased demands for large-scale remote sensing fuel information. This study develops a methodology for mapping canopy fuels over large areas (Germany) at high spatial resolution, exclusively relying on open remote sensing data.</p> <p><br>We propose a two-step approach where we first use measurements from NASA’s GEDI instrument to estimate canopy fuel variables at the footprint level, before predicting high-resolution raster maps. Instead of using field measurements, we generate (GEDI-) footprint-level estimates for Canopy (Base) Height (CH, CBH),<br>Cover (CC), Bulk Density (CBD), and Fuel Load (CFL) by segmenting airborne LiDAR point clouds and processing tree-level metrics with allometric crown biomass<br>models. To predict footprint-level canopy fuels we fit and tune Random Forest models, which are cross-validated using k-fold Nearest Neighbor Distance Matching.<br>Predictions at >1.6 M GEDI footprints and biophysical raster covariates are combined with a Universal Kriging method to produce countrywide maps at 20-meter resolution.</p> <p><br>Agreement (RMSE/R²) with validation data (from the same population) was strong for footprint-level predictions and moderate for map predictions. A validation<br>with estimates based on National Forest Inventory data revealed low to modest agreement. Better accuracy was achieved for variables related to height (CH, CBH)<br>rather than to cover or biomass (CBD, CFL). Error analysis pointed towards a mixture of biases in model predictions and validation data, as well as underestimation of<br>model prediction standard errors. Contributing factors may be simplification through allometric equations and spatial and temporal mismatch of data inputs.<br>The proposed workflow has the potential to support regions where wildfire is an emerging issue, and fuel and field information is scarce or unavailable.</p> <p> </p> <p>Data:</p> <p>This repository contains modeling data, model objects (R), and predicted maps. The TIFF-files each have six bands, which includes (1) the final Universal Kriging result, (2) the linear model prediction (3) the prediction of residual Kriging, (4) the Kriging variance, (5) the linear model prediction standard error, and (6) Universal Kriging standard error.</p> <p> </p> <p>Disclaimer:<br>Maps in this repository are predicted using canopy fuel estimates from GEDI measurements. These are limited the region between 51.6° North and South. Map predictions exceeding this range should be considered an extrapolation of the model to an unknown biophysical domain. Error maps (6) can aid in utilizing our canopy fuel maps.</p>
First Street Foundation Flood Model Hazard Layers V1.3
<p>Up to 15 different hazard layers are available, representing 3 different time periods (2021, 2036, 2051) and 4-5 different return periods from the 2-year (coastal only) to the 500-year intervals.</p> <p>Data is delivered in GeoTIFF format and at a 3 meter resolution with each pixel representing depth of flooding in centimeters. This high resolution dataset allows you to visualize flood extents at multiple return periods both today and in the future.</p> <p>The hazard inundation layers are emailed through a clickable link that automatically starts the download of the datasets. The Version 1.3 hazards are available for the contiguous United States.</p> <p>You can download a sample of the hazard layers generated from First Street's Flood Model on this page. You can request access to the hazard layers for areas within the contiguous United States on the First Street website<a href="https://firststreet.org/data-access/paid-access/?utm_source=Hazard_Layers&utm_medium=Purchase_Data&utm_campaign=Zenodo#pricing-component"> here</a>. You can find the data dictionary which breaks down the data that is available with each hazard layer purchase<a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-hazard-dictionary/?utm_source=Hazard_Layers&utm_medium=Hazard_Dictionary&utm_campaign=Zenodo"> here</a>. If you are also interested in the flood risk statistics, you can find more information <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/data-dictionary/?utm_source=Hazard_Layers&utm_medium=Data_Dictionary&utm_campaign=Zenodo">here</a>.</p>
Semi-Supervised Pre-trained Foundation Model for 3D Structural Feature Analysis of Seismic Images
<p>Codes, trained model, and datasets for the paper "Semi-Supervised Pre-trained Foundation Model for 3D Structural Feature Analysis of Seismic Images".</p>
UniFMIR: Pre-training a Foundation Model for Universal Fluorescence Microscopy Image Restoration
<p>This repository contains the preprocessed dataset for [UniFMIR](https://github.com/cxm12/UNiFMIR/). All training and test data involved in the experiments are publicly available datasets. Licenses of the original dataset are applied. You can refer to the Github repository for details.</p> <p>* The 3D denoising/isotropic reconstruction/projection datasets can be downloaded from [Content Aware Image Restoration dataset](https://publications.mpi-cbg.de/publications-sites/7207/). `Projection_Flywing/train_data/my_training_data.npz` are generated according to the [CSBDeep](http://csbdeep.bioimagecomputing.com/doc/).</p> <p>* The SR dataset can be downloaded from [BioSR dataset](https://doi.org/10.6084/m9.figshare.13264793). The dataset is augmented according to the instructions in [DFCAN](https://github.com/qc17-THU/DL-SR/tree/main#train-a-new-model) and `my_training_data.npz` files are generated following [CSBDeep](http://csbdeep.bioimagecomputing.com/doc/datagen.html). </p> <p>* The Volumetric reconstruction dataset are from [VCD-LFM dataset](https://doi.org/10.5281/zenodo.4390067). The dataset is prepared according to the instructions in [VCD-Net](https://github.com/feilab-hust/VCD-Net).</p> <p>* DeepBacs dataset can be downloaded from [DeepBacs dataset](https://zenodo.org/record/6460867). We split the dataset into 5 folds for cross-validation. Shareloc dataset can be downloaded from [Shareloc dataset](https://zenodo.org/record/7234161).</p> <p> </p> <p>The data paths should be as follows:</p> <p>```</p> <p>VCD/vcdnet/</p> <p>CSB/DataSet/</p> <p> Denoising_Planaria/</p> <p> Denoising_Tribolium/</p> <p> Isotropic/Isotropic_Liver/</p> <p> Projection_Flywing/</p> <p> BioSR_WF_to_SIM/DL-SR-main/dataset/</p> <p> Synthetic_tubulin_gfp/</p> <p> Synthetic_tubulin_granules/</p> <p>DeepBacs/</p> <p>Shareloc/</p> <p>```</p>
stFormer: a foundation model for spatial transcriptomics
<p>stFormer incorporates ligand genes within the spatial niche into transformer encoder of single-cell transcriptomics, and outputs gene embeddings specific to the intracellular context and spatial niche. These gene representations can serve as input of various downstream applications, including cell clustering, cell type prediction, gene function prediction, and <em>in silico</em> perturbation analysis of ligand-receptor interaction.</p> <p>The model architecture is designed for ST data resolved at the single-cell level. We propose a biased cross-attention method to enable the model to do learning with single-cell resolution on low-resolution, whole-transcriptome Visium data, which is a widely available spatial resource.</p> <p>We assembled a pretraining corpus comprising ~4.1 million spatial samples from public human Visium datasets, spanning diverse tissues, development stages, and disease states. After pretraining, stFormer is compatible with both single-cell and spot resolution ST data.</p>
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
<p>We introduce a new image-text dataset, providing high-quality natural language descriptions for global-scale satellite data. Specifically, we utilize Sentinel-2 data for its global coverage as the foundational image source, employing semantic segmentation labels from the European Space Agency's WorldCover project to enrich the descriptions of land covers. By conducting in-depth semantic analysis, we formulate detailed prompts to elicit rich descriptions from ChatGPT. We then include a manual verification process to enhance the dataset's quality further. This step involves manual inspection and correction to refine the dataset. Finally, we offer the community ChatEarthNet, a large-scale image-text dataset characterized by global coverage, high quality, wide-ranging diversity, and detailed descriptions. ChatEarthNet consists of 163,488 image-text pairs with captions generated by ChatGPT-3.5 and an additional 10,000 image-text pairs with captions generated by ChatGPT-4V(ision). This dataset has significant potential for both training and evaluating vision-language geo-foundation models for remote sensing.</p>
Scupa: Single-cell unified polarization assessment of immune cells using the single-cell foundation model
<p>This dataset contains the processed scRNA-seq data and code for single-cell unified polarization assessment. Please refer to the article "Scupa: Single-cell unified polarization assessment of immune cells using the single-cell foundation model" for the detailed data and method description.</p> <p>*.rds: the processed scRNA-seq datasets with Universal Cell Embeddings saved as a Seurat object. The Universal Cell Embeddings are saved in the assay 'uce'.</p> <ul> <li>immune_dict_uce.rds: Immune Dictionary.</li> <li>ifnb_treatment_uce.rds: human PBMCs treated with IFN-beta.</li> <li>macrophage_stimulation_uce.rds: human macrophages treated with one or two of IFN-beta, IFN-gamma, TNF-alpha and IL-4.</li> <li>il2_treatment_uce.rds: mouse spleen CD8+ T cells treated with IL-2 or anti-PD-L1.</li> <li>pan_cancer_myeloid_uce.rds: tumor-infiltrating myeloid cells across seven cancer types.</li> </ul> <p>notebooks.zip: Jupyter notebooks containing code for training models and applications to multiple datasets.</p>
Dataset for the study "Foundation Models as Assistive Tools in Hydrometeorology: Opportunities, Challenges, and Perspectives"
<p>This dataset contains the materials, files, and codes used in the study "Foundation Models as Assistive Tools in Hydrometeorology: Opportunities, Challenges, and Perspectives". </p>
Data for Developing Machine Learning Models to Predict Base Resistance of Pile Foundation
<p>This data was collected from 86 static pile load tests across 37 different high-rise buildings in Vietnam, especially soft soil region in Mekong Delta (Ho Chi Minh City). The data was used to develop machine learning models to predict base resistance of piles. Further details can be found in publication: "<strong>Influence of Settlement on Base Resistance of Long Piles in Soft Soil—Field and Machine Learning Assessments</strong>", Link: https://www.mdpi.com/2673-7094/4/2/25.</p> <p>Recommended citation: Nguyen, Thanh T., Viet D. Le, Thien Q. Huynh, and Nhu H.T. Nguyen. 2024. "Influence of Settlement on Base Resistance of Long Piles in Soft Soil—Field and Machine Learning Assessments" <em>Geotechnics</em> 4, no. 2: 447-469. https://doi.org/10.3390/geotechnics4020025</p> <p> </p>
Towards Neural Scaling Laws for Foundation Models on Temporal Graphs
<p>Datasets provided in this storage are introduced in the paper: <em>Towards Neural Scaling Laws for Foundation Models on Temporal Graphs</em></p> <ul> <li>Each .csv file represents all transactions of the token network that has the same name as the file name (<em><tokenname.csv>)</em></li> <li>Each transaction corresponds to a row in each file.</li> <li>Each transaction has: <ul> <li> blockNumber : is the block ID of Ethereum that includes this transaction</li> <li>timestamp: time that the transaction is made in UNIX timestamp format</li> <li>tokenAddress : the address that specifies a unique ERC20 token</li> <li>from: address of sender</li> <li>to: address of receiver</li> <li>value: the amount the transaction</li> <li>fileBlock: we split the whole number of blocks count to 35 buckets and assigned the bucket ID to the transaction to trace the blocks </li> </ul> </li> <li>To use the same setting as described in the papers, we include edge list and label that contain node interactions and labels for each snapshot in each token network. <ul> <li>Each transaction in the edge list also has "from","to" and "amount" fields, but with an additional "snapshot" field to indicate the index of the snapshot that the transaction below to</li> <li>Each row in label file indicates the ground truth label of the snapshot having an index corresponding to the index of the row (e.g first row indicates the label of the first snapshot)</li> <li>We provided the way to generate edge lists and label files in the following Github repository: https://github.com/benjaminnNgo/ScalingTGNs/blob/main/script/utils/TGS.py</li> </ul> </li> <li>However, we also provide raw <em>.csv </em> to divide into generate <em>edgeslist </em>and <em>label with a different setting.</em></li> </ul> <p><br><br></p>
Foundation Detection Model
<p>Foundation Detection Model</p>
Reproducibility Test Data of Foundation Model for Cancer Imaging x Mhub
<p>This dataset provides test data for the FMCIB model integrated within the MHub platform, a robust solution for deploying, managing, and testing deep learning models tailored for medical imaging. </p> <h3>Dataset Composition:</h3> <ul> <li><strong>Sample Folder:</strong> Contains the input data utilized for testing the model’s functionality.</li> <li><strong>Reference Folder:</strong> Contains the corresponding output provided by the original model contributor.</li> <li><strong>Test.yml File:</strong> This file includes the original contributor’s test setup, which has been accepted by the MHub team.</li> </ul> <h3>Sample Data Source:</h3> <p>The sample images used in this dataset are sourced from public datasets available through the <strong><a href="https://datacommons.cancer.gov/repository/imaging-data-commons" target="_blank" rel="noopener">Imaging Data Commons (IDC)</a></strong>, a repository that provides access to a wide range of medical imaging data. This ensures that the test cases reflect real-world clinical scenarios, facilitating robust validation of model performance.</p> <h3>Purpose and Utility:</h3> <p>The primary objective of this dataset is to enable the rigorous testing and validation of model performance within MHub workflows. To assess the performance of a model, users can process the sample data and compare the resulting output to the reference data. Additionally, users may inspect the sample and reference data independently to better understand the input-output structure that defines each model’s workflow.</p> <p>This dataset streamlines the process of model validation. By providing a standardized testing framework, the dataset facilitates reproducible results and accelerates the development of reliable AI models for medical imaging.</p> <h3>About MHub:</h3> <p>MHub (<a href="https://mhub.ai/" target="_new" rel="noopener">mhub.ai</a>) is an innovative platform designed to simplify the deployment, management, and testing of deep learning models for medical imaging. It enables researchers and clinicians to integrate AI-based solutions into clinical workflows while ensuring reproducibility and scalability. The platform provides a modular framework where users can execute complex workflows, such as image segmentation, classification, and registration, leveraging state-of-the-art AI models. MHub's goal is to accelerate the development and clinical adoption of medical imaging models by providing a streamlined, user-friendly environment for testing and validating new algorithms.</p> <p>For more information on the platform and its capabilities, visit <a href="https://mhub.ai/" target="_blank" rel="noopener">mhub.ai</a>.</p>
SpectralGPT: The first remote sensing foundation model customized for spectral data
<p>SpectralGPT is the first purpose-built foundation model designed explicitly for spectral RS data. It considers unique characteristics of spectral data, i.e., spatial-spectral coupling and spectral sequentiality, in the MAE framework with a simple yet effective 3D GPT network.</p> <p>We will gradually release the trained models (SpectralGPT, SpectralGPT+), the new benchmark dataset (SegMunich) for the downstream task of semantic segmentation, original code, and implementation instructions.</p>
NextVir: Enabling Classification of Tumor-Causing Viruses with Genomic Foundation Models
<p>These are a collection of 150bp reads synthesized using ART. Viral reference genomes were downloaded using iCAV, and the primary assemblies of GRCh38.p14 were used for the human reference.</p>
FloodCastBench: A Large-Scale Dataset and Foundation Models for Flood Modeling and Forecasting
<div>Effective flood forecasting is crucial for informed decision-making and emergency response. Existing flood datasets mainly describe flood events but lack dynamic process data suitable for machine learning (ML). This work introduces the FloodCastBench dataset, designed for ML-based flood modeling and forecasting, featuring four major flood events: Pakistan 2022, UK 2015, Australia 2022, and Mozambique 2019. FloodCastBench provides comprehensive low-fidelity and high-fidelity flood forecasting datasets specifically for ML.</div> <div> </div> <div>This dataset comprises three folders: the low-fidelity flood forecasting folder, the high-fidelity flood forecasting folder, and the relevant data folder. The low-fidelity flood forecasting folder includes data on the 2022 Pakistan flood and the 2019 Mozambique flood, both with a spatial resolution of 480 m. The high-fidelity flood forecasting folder contains two subfolders: one for the 2022 Australia flood and the 2015 UK flood with a spatial resolution of 30 m, and another for the same floods with a spatial resolution of 60 m. All data files are stored in TIFF format, with a temporal resolution of 300 seconds, and file names are numbered sequentially, incremented every 300 seconds until the simulation endpoint. The relevant data folder includes five subfiles: DEM, land use and land cover, rainfall data, georeferenced files, and initial condition files. The DEM, land use and land cover, rainfall, and initial condition data are all provided in TIFF format. The rainfall data is organized in a format of year-month-day-hour-minute-second. Georeferenced files provide geographic extent and spatial reference to support viewing and analysis of the associated TIFF files in GIS.</div> <div> </div> <div>FloodCastBench details the process of flood dynamics data acquisition, starting with input data preparation (e.g., topography, land use, rainfall) and flood measurement data collection (e.g., SAR-based maps, surveyed outlines) for hydrodynamic modeling. We deploy a widely recognized finite difference numerical solution to construct high-resolution spatiotemporal dynamic processes with 30-m spatial and 300-second temporal resolutions. Flood measurement data are used to calibrate the hydrodynamic model parameters and validate the flood inundation maps. Furthermore, we establish a benchmark of foundational models for neural flood forecasting using FloodCastBench, validating its effectiveness in supporting ML models for spatiotemporal, cross-regional, and downscaled flood forecasting.</div>
Dataset for Foundation Model for Composite Materials and Microstructural Analysis
<p>This repository contains datasets generated for the publication <strong>"Foundation Model for Composite Materials and Microstructural Analysis". </strong></p> <p>The datasets are provided to facilitate replication of results, further research, and development in the field of composite materials and machine learning.</p> <h1>Contents:</h1> <ol> <li><strong>downstream_circular_inclusion.zip</strong> <ul> <li> <p>Description: This dataset consists of microstructure images and associated properties for circular inclusion composites used in the downstream tasks of the study.</p> </li> <li> <p>Structure:</p> <ul> <li> <p>train/: Contains training data directories and files.</p> </li> <li>valid/: Contains training data directories and files.</li> <li> <p>descriptors.csv: A CSV file with descriptors and homogenized stiffness components for validation data.</p> </li> </ul> </li> </ul> </li> <li><strong>downstream_short_fiber.zip</strong> <ul> <li> <p>Description: This dataset consists of microstructure images and associated properties for short-fiber composites used in the downstream tasks of the study.</p> </li> <li> <p>Structure:</p> <ul> <li> <p>train/: Contains training data directories and files.</p> </li> <li>valid/: Contains training data directories and files.</li> <li> <p>descriptors.csv: A CSV file with descriptors and homogenized stiffness components for validation data.</p> </li> </ul> </li> </ul> </li> <li><strong>inclusion_train_100k.zip</strong> <ul> <li>Description: This dataset contains 100,000 microstructure images of short-fiber composites used for self-supervised pre-training of the Masked Material Autoencoder (MMAE).</li> <li> <p>Structure:</p> <ul> <li> <p>mesh0dir/, mesh1dir/, ... mesh99999dir/: : Folders containing microstructure images used for pre-training.</p> </li> <li>descriptors.csv: A CSV file with descriptors related to the microstructures.</li> </ul> </li> </ul> </li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.