Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

708

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

708 results for “Global dataset”

Learn how ShareScore rates datasets ↗
nasa24/100

Last of the Wild Project, Version 1, 2002 (LWP-1): Global Human Footprint Dataset (Geographic)

The Global Human Footprint Dataset of the Last of the Wild Project, Version 1, 2002 (LWP-1) is the Human Influence Index (HII) normalized by biome and realm. The HII is a global dataset of 1-kilometer grid cells, created from nine global data layers covering human population pressure (population density, population settlements), human land use and infrastructure (built up areas, nighttime lights, land use/land cover), and human access (coastlines, roads, railroads, navigable rivers). The dataset in Clarke 1866 Geographic Coordinate System is produced by the Wildlife Conservation Society (WCS) and the Columbia University Center for International Earth Science Information Network (CIESIN).

restrictednotspecifiedApr 2025View details →
nasa24/100

Global Man-made Impervious Surface (GMIS) Dataset From Landsat

The Global Man-made Impervious Surface (GMIS) Dataset From Landsat consists of global estimates of fractional impervious cover derived from the Global Land Survey (GLS) Landsat dataset for the target year 2010. The GMIS dataset consists of two components: 1) global percent of impervious cover; and 2) per-pixel associated uncertainty for the global impervious cover. These layers are co-registered to the same spatial extent at a common 30m spatial resolution. The spatial extent covers the entire globe except Antarctica and some small islands. This dataset is one of the first global, 30m datasets of man-made impervious cover to be derived from the GLS data for 2010 and is a companion dataset to the Global Human Built-up And Settlement Extent (HBASE) dataset. The dataset is expected to have a rather broad spectrum of users, from those wishing to examine/study the fine details of urban land cover over the globe at full 30m resolution to global modelers trying to understand the climate/environmental impacts of man-made surfaces at continental to global scales. For example, the data are applicable to local modeling studies of urban impacts on the energy, water, and carbon cycles, as well as analyses at the individual country level.

restrictednotspecifiedMar 2025View details →
nasa24/100

Randolph Glacier Inventory - A Dataset of Global Glacier Outlines, Version 5

The Randolph Glacier Inventory (RGI) is a global set of glacier outlines; it is intended as a snapshot of the world’s glaciers. This data set provides a single outline for each glacier and is produced in coordination with the Global Land Ice Measurements from Space (GLIMS) initiative. The RGI is not suitable for measuring glacier-by-glacier rates of area change, but can be used to estimate glacier volumes, rates of elevation change at regional and global scales, and cryospheric responses to climatic forcing.Glacier mapping data are contributed to both GLIMS and the RGI from the glaciological community. RGI is produced by the Working Group on the Randolph Glacier Inventory and Infrastructure for Glacier Monitoring, a body of the International Association of Cryospheric Sciences (IACS). This data set is updated approximately annually.Glacier outlines are distributed as Shapefiles. Hypsometric data (CSV files) and gridded auxiliary data (GeoTIFFs) are also available. All RGI data are packaged globally and by region, with regions based upon the Global Terrestrial Network for Glaciers.

restrictednotspecifiedApr 2025View details →
zenodo20/100

Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict

<p>We present a dataset that collects tweets from news media channels worldwide that pertain to the Russo-Ukrainian war. This dataset spans a period of February 2022-May 2023. The dataset is unique in its global scope, encompassing tweets in various languages and from different parts of the world. Additionally, we extracted information about the stance, sentiment, prominent entities &amp; concepts that occur in tweets to be able to answer questions about the discourse: who says what (prominent entities), who stands (stance) where on what aspect (prominent concepts), how are the aspects portrayed (sentiment). We also downloaded the images attached to the post and classified them to extract image tags for each image. The dataset includes 1,524,826 tweets, out of which 306,295 tweets have images, for 60 languages.<br><br>The source code for the collection and processing of tweets can be found on here:&nbsp;<a href="https://github.com/sherzod-hakimov/ru-ua-news-discourse-twitter"><em>https://github.com/sherzod-hakimov/ru-ua-news-discourse-twitter</em></a></p> <p>Each entry in the dataset is a single JSON line and has the following entries:</p> <pre><code>{ 'tweet_id': 'lang': 'stanza_output': 'stanza_named_entities': 'sentiment': 'stance': 'channel': 'country': 'verified':<br>'image_tags': }</code></pre> <pre>&nbsp;</pre> <p><em><strong>If you need access to the full text of the dataset, please</strong> <strong>contact us via an email: <a href="mailto:sherzodhakimov@gmail.com">sherzodhakimov (at sign) gmail.com</a></strong></em><br><br>If you find the resources useful, please cite us:<br><br>```</p> <p>@inproceedings{hakimov2023unveiling,<br>&nbsp; &nbsp; &nbsp; title={Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict},&nbsp;<br>&nbsp; &nbsp; &nbsp; author={Sherzod Hakimov and Gullal S. Cheema},<br>&nbsp; &nbsp; &nbsp; booktitle={Proceedings of the 2024 {ACM} International Conference on Multimedia Retrieval, {ICMR} 2024},<br>&nbsp; &nbsp; &nbsp; year={2024}<br>}<br>```</p>

openJun 2023View details →
zenodo20/100

Understanding Global Brain Network Alterations in Glioma Patients [dataset]

<p>This dataset was published to support the reproducibility of the paper&nbsp;<a href="https://pubmed.ncbi.nlm.nih.gov/33947274/">Understanding Global Brain Network Alterations in Glioma Patients</a>. They should be used with the MATLAB scripts posted at&nbsp;<a href="https://github.com/multinetlab-amsterdam/projects/tree/master/clustering_paper_2021">https://github.com/multinetlab-amsterdam/projects/tree/master/clustering_paper_2021</a>.</p> <p><strong>Important information: the files were organized in random order and subsequently used indices/identifiers do not reflect any patient/subject codes. Data is fully anonymous.</strong></p>

openDec 2021View details →
zenodo20/100

Global Datasets of carbon isotope composition (δ13C) for Ecological and Earth System Research

<p>This is dataset for paper: global Datasets of carbon isotope composition (&delta;13C) for Ecological and Earth System Research, which&nbsp;include&nbsp;soil&nbsp;&delta;13C samples from 2500 research sites across global.</p>

restrictedMay 2022View details →
zenodo20/100

A monthly full-coverage satellite-based global atmospheric CO2 dataset at 0.05° resolution from 2015 to 2021

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo20/100

Subset of CPTAC CCRCC Global DIA Dataset

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo20/100

Dataset for "Global dominance of seasonality in shaping lake surface extent dynamics" (in review)

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo20/100

A spatially-explicit dataset of global wind erosion based on distributed RWEQ model

<p>This is a global wind erosion estimation dataset obtained based on the distributed RWEQ model with a resolution of 0.05&deg;. It's the raw data of the article &ldquo;Global Wind Erosion Reduction Driven by Changing Climate and Land Use&rdquo; (https://doi.org/10.1029/2024EF004930). The dataset is NetCDF (Network Common Data Format) and contains 38 years between 1982 and 2019. Dimension is the year information for wind erosion, beginning in 1982 and ending in 2019.</p> <p>Coordinate system: WGS_84 (EPSG:4326)</p> <p>NoData Value=-10000</p> <p>NETCDF_VARNAME=SoilLoss</p> <p>&nbsp;</p> <h3><strong>Please cite this dataset via the Earth's Future journal article</strong><strong>&ldquo;Global Wind Erosion Reduction Driven by Changing Climate and Land Use&rdquo; </strong></h3> <h3><strong>https://doi.org/10.1029/2024EF004930</strong></h3> <p>Citation:<br>Sun, R., He, H., Jing, Y., Leng, S., Yang,G., L&uuml;, Y., et al. (2024). Global winderosion reduction driven by changingclimate and land use. Earth's Future, 12,e2024EF004930. https://doi.org/10.1029/2024EF004930</p> <p>&nbsp;</p> <p>For access to the data, please contact the corresponding author.</p>

restrictedcc-by-sa-4.0Oct 2024View details →
zenodo16/100

Datasets and R code for Liu et al. A global meta-analysis on the drivers of salt marsh planting success and implications for ecosystem services

<p>Planting has been widely adopted to battle the loss of salt marshes and to establish living shorelines. However, the drivers of success in salt marsh planting and their ecological effects are poorly understood at the global scale. Here, we assemble a global database, encompassing 22,074 observations reported in 210 studies, to examine the drivers and impacts of salt marsh planting. We show that, on average, 53% of plantings survived globally, and plant survival and growth can be enhanced by careful design of sites, species selection, and novel planted technologies. Planting enhances shoreline protection, primary productivity, soil carbon storage, biodiversity conservation and fishery production (effect sizes = 0.61, 1.55, 0.21, 0.10 and 1.01, respectively), compared with degraded wetlands. However, the ecosystem services of planted marshes, except for shoreline protection, have not yet fully recovered compared with natural wetlands (effect size = -0.25, 95% CI&thinsp;-0.29, -0.22). Fortunately, the levels of most ecological functions related to climate change mitigation and biodiversity increase with plantation age when compared with natural wetlands, and achieve equivalence to natural wetlands after 5-25 years. Overall, our results suggest that salt marsh planting could be used as a strategy to enhance shoreline protection, biodiversity conservation and carbon sequestration.</p>

restrictedcc-by-4.0Dec 2023View details →
zenodo16/100

Dataset of global gridded monthly crop coefficient, yearly and monthly blue-to-total water footprint ratio, and national unit blue and green water footprints of maize (2000-2021)

<p>The data includes monthly <span><span>crop coefficient</span></span>, yearly and monthly blue-to-total water footprint ratio at a 5 arcminute spatial scale, and the unit water footprint at an annual national (regional) scale of global maize.</p>

restrictedcc-by-4.0Jul 2024View details →
zenodo16/100

CloudSEN12 - a global dataset for semantic understanding of cloud and cloud shadow in Sentinel-2

<p><strong>Description</strong></p> <p>CloudSEN12 is a large dataset for cloud semantic understanding that consists of 9880 regions of interest (ROIs). Each ROI has five 5090x5090 meters image patches (IPs) collected on different dates; we manually choose the images to guarantee that each IP inside an ROI matches one of the following cloud cover groups:</p> <p>- clear (0%)</p> <p>- low-cloudy (1% - 25%)&nbsp;</p> <p>- almost clear (25% - 45%)</p> <p>- mid-cloudy (45% - 65%)</p> <p>- cloudy (65% &gt;)</p> <p>An IP is the core unit in CloudSEN12. Each IP contains data from Sentinel-2 optical levels 1C and 2A, Sentinel-1 Synthetic Aperture Radar (SAR), digital elevation model, surface water occurrence, land cover classes, and cloud mask results from eight cutting-edge cloud detection algorithms. Besides, in order to support standard, weakly, and self-/semi-supervised learning procedures, cloudSEN12 includes three distinct forms of hand-crafted labelling data: high-quality, scribble, and no annotation. Consequently, each ROI is randomly assigned to a different annotation group:</p> <ul> <li> <p>2000 ROIs with pixel-level annotation, where the average annotation time is 150 minutes (high-quality group).</p> </li> <li> <p>2000 ROIs with scribble level annotation, where the annotation time is 15 minutes (scribble group).</p> </li> <li> <p>5880 ROIs with annotation only in the cloud-free (0\%) image (no annotation group).</p> </li> </ul> <p>For high-quality labels, we use the Intelligence foR Image Segmentation\cite{iris2019} (IRIS) active learning technology, a system that combines human photo-interpretation and machine learning. For scribble, ground truth pixels were drawn using IRIS but without ML support. Finally, the no annotation dataset is generated automatically, with manual annotation only in the clear image patch. The dataset is already available here: <strong><a href="https://shorturl.at/cgjtz">https://shorturl.at/cgjtz</a></strong>. Check out our website <strong><a href="https://cloudsen12.github.io/">https://cloudsen12.github.io/</a></strong> for examples of how to download the dataset via STAC.</p>

restrictedAug 2022View details →
zenodo16/100

Global River Discharge Reanalysis dataset (GRDR)

<p>This repository contains the&nbsp;Global River Discharge Reanalysis dataset (GRDR) generated from a work that is currently under review. GRDR is&nbsp;a global river discharge product that assimilated remotely sensed river discharge and hydrologic model simulations. It contains&nbsp;daily flows in ~2.9 million vectorized river reaches for 1984-2018.</p> <p>The data format is netCDF.</p> <p>&nbsp;</p>

restrictedJul 2023View details →
zenodo16/100

Predictor variables of the Global Tree-Canopy Cover Change dataset (GTCCC)

<p>Compilation of variables used to generate the Global Tree-Canopy Cover Change dataset (GTCCC). The contents are described in the file &quot;variable_list.xlsx&quot;, and their use is documented in a complementary manuscript. The GTCC data and the code used to generate it can be found in a <a href="https://doi.org/10.5281/zenodo.7901290">separate Zenodo repository.</a></p>

restrictedAug 2023View details →
zenodo16/100

Social Media Big Dataset for Research, Analytics, Prediction, and Understanding the Global Climate Change Trends

<p>Yuriy Syerov, October 6, 2023, &quot;Social Media Big Dataset for Research, Analytics, Prediction, and Understanding the Global Climate Change Trends&quot;, IEEE Dataport, doi: https://dx.doi.org/10.21227/71ms-8v86</p> <p>https://ieee-dataport.org/documents/social-media-big-dataset-research-analytics-prediction-and-understanding-global-climate</p>

restrictedDec 2022View details →
zenodo12/100

MEG dataset adult local-global paradigm

<p>MEG data of 16 healthy adults with an auditory local-global paradigm. For closer description of data see metadata and data description files. </p> <p>If unziping the dataset on linux  -FF option should be used.</p>

restrictedJun 2017View details →
zenodo12/100

Dataset for "Global dominance of seasonality in shaping lake surface extent dynamics" (in review)

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Sep 2024View details →
zenodo12/100

Global dataset of fallout radionuclides in cryoconite

<p>These data describe the activity concentrations of fallout radionuclides (137Cs, 241Am, 210Pb) in cryoconite and proglacial sediment samples collected on glaciers&nbsp;around the global cryosphere.</p>

restrictedOct 2021View details →
zenodo12/100

Global crop-specific nitrogen fertilization dataset in 1961-2020

<p>This is a dataset for crop-specific N fertilization with 5-arc-min resolution during 1961 to 2020, including N fertilizer inputs, types and placements. The N fertilization data was classified into 21 crop types, 13 fertilizer types and 2 fertilization placements.The datasets extend from 180&deg;E to 180&deg;W longtitude and 90&deg;S to 90&deg;N latitude with a resolution of 5 arc-min for the temporal period 1961&ndash;2020 in standard WGS84 coordinate system, with a temporal resolution. The data are provided in Hierarchical Data Format (HDF5) format.</p>

restrictedDec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record