Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
708
datasets available to search
ShareScore release 0.9.0
Dataset results
708 results for “Global dataset”
Global derived datasets for use in k-NN machine learning prediction of global seafloor total organic carbon
<p>This dataset includes 663 predictor grids used for k-NN global prediction of seafloor total organic carbon.</p> <p>663 predictor grids available in netCDF4 HDF5 file format. Grids are cell-centered sized 4320 x 2160. File names adhere to the naming conventions discussed below. The naming structure is partioned by underscores and periods in the following order: interface to which the gridded values refer to, quantity of values contained within the grid, units and reference values/units (e.g. meters below sea level), data source, statistic calculated (if applicable), grid pitch, and file extension.</p> <p>Possible interfaces from the top – down:</p> <p>SS – Sea surface – atmosphere interface (may also be average of the entire water column)</p> <p>SF – Seafloor – water interface (may also be denoted by GL)</p> <p>GL – Ground level (e.g. bottom of pure liquid, top of dirt)</p> <p>SC – Sediment – crust interface (e.g. sediment above, igneous/metamorphic below)</p> <p>CM – Crust – mantle interface (e.g. Mohorovicic discontinuity)</p> <p>Appropriate reference naming marker (bold), original data source, and date of last access:</p> <p><strong>Becker</strong></p> <p>Becker, J. J., Wood, W. T., & Martin, K. M. (2014). <em>Global crustal heat flow using random decision forest prediction</em>, Abstract NG31A-3788 presented at 2014 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 06/23/2015.</p> <p><strong>CRUST1</strong> </p> <p>Pasyanos, M.E., Masters, G., Laske, G. & Ma, Z. (2012). <em>LITHO1.0 - An Updated Crust and Lithospheric Model of the Earth Developed Using Multiple Data Constraints</em>, Abstract T11D-09 presented at 2012 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 07/01/2014.</p> <p><strong>CRUST1_NOAA</strong></p> <p> As the NOAA sediment thickness database is globally not complete, data gaps in the NOAA grid with this have been supplemented by the CRUST1 sediment thickness (see above citation).</p> <p>Whittaker, J., Goncharov, A., Williams, S., Müller, R. D., & Leitchenkov, G. (2013) Global sediment thickness dataset updated for the Australian-Antarctic Southern Ocean, <em>Geochemistry, Geophysics, Geosystems. </em>https://doi.org/10.1002/ggge.2018.<em> </em>Last access: 09/02/2018.</p> <p><strong>GVP</strong></p> <p>Global Volcanism Program (2013) Volcanoes of the World. In E. Venzke (ed.). (Vol. 4.7.3). Smithsonian Institution. https://doi.org/10.5479/si.GVP.VOTW4-2013. Last access: 09/22/2014.</p> <p><strong>ETOPO2v2</strong></p> <p>National Geophysical Data Center (2006). 2-minute Gridded Global Relief Data (ETOPO2) v2. National Geophysical Data Center, NOAA. DOI: 10.7289/V5J1012Q. Last access: 02/06/2013.</p> <p><strong>PLATES</strong></p> <p>Coffin, M.F., Gahagan, L.M., & Lawver, L.A. (1998). Present-day Plate Boundary Digital Data Compilation. University of Texas Institute for Geophysics Technical Report (No. 174, pp. 5). Last access: 09/15/2014.</p> <p><strong>ONRL</strong></p> <p>Ludwig,W., Amiotte-Suchet, P., & Probst, J. L. (2011). ISLSCP II Global River Fluxes of Carbon and Sediments to the Oceans. In F. G. Hall, G. Collatz, B. Meeson, S. Los, E. Brown de Colstoun, and D. Landis (Eds.), <em>ISLSCP Initiative II Collection</em>. Oak Ridge National Laboratory Distributed Active Archive Center, Oak Ridge, Tennessee, U.S.A. http://dx.doi.org/10.3334/ORNLDAAC/1028. Last Access: 02/15/2015.</p> <p><strong>Muller</strong></p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust, <em>Geochemistry, Geophysics, Geosystems</em>, 9(4), Q04006. https://doi.org/10.1029/2007GC001743. Last accessed: 07/19/2011.</p> <p><strong>Woa13x</strong></p> <p>Boyer, T.P., Antonov, J. I., Baranova, O. K., Coleman, C., Garcia, H. E., Grodsky, A., et al. (2013) World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), <em>NOAA Atlas NESDIS 72, Technical Ed</em>. Silver Spring, MD. http://doi.org/10.7289/V5NZ85MT. Last Access: 09/18/2014.</p> <p><strong>KIM</strong></p> <p>Kim, S.S. & Wessel, P. (2011). New global seamount census from the altimetry-derived gravity data, <em>Geophysical Journal International</em>, 186, 615-631. https://doi.org/10.1111/j.1365-246X.2011.05076.x. Last access: 09/22/2014.</p> <p><strong>HYCOM</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data.Last access: 03/19/2014.</p> <p><strong>NCEDC</strong></p> <p>NCEDC (2016). Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC. Last access: 09/21/2014.</p> <p><strong>Wei2010</strong></p> <p>Wei, C.-L., Rowe, G. T., Escobar-Briones, E., Boetius, A., Soltwedel, T., Caley, M. J., et al.(2010). Global patterns and predictions of seafloor biomass using random forests. <em>PLoS ONE</em>,5(12), e15323. https://doi.org/10.1371/journal.pone.0015323 Last access: 06/20/2016.</p> <p><strong>NGA_egm2008</strong></p> <p>Pavlis, N.K., Holmes, S. A., Kenyon, S. C., & Factor, J. K. (2008). <em>The</em> <em>EGM2008 Global Gravitational Model</em>, Abstract 2008AGUFM.G22A..01P presented at the 2008 General Assembly of the European Geosciences Union, Vienna, Austria. Last access: 07/10/2014.</p> <p><strong>WAVEWATCH3</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data. Last access: 03/19/2014.</p> <p>Updated global seafloor porosity grid using our k-nearest neighbors algorithm using 5 nearest neighbors. Observed data used for prediction from Martin et al. (2015). </p> <p>Martin, K. M., Wood, W. T., & Becker, J. J. (2015). A global prediction of seafloor sediment porosity using machine learning. <em>Geophysical Research Letters</em>, 42(24), 10640. https://doi.org/10.1002/2015GL065279</p> <p>Other grids which have been generated by empirical means are latitude (and derivatives), longitude (and derivatives), Coriolis, coast_is_1.0, and the random noise grids. </p> <p>Units referenced are as follows:</p> <p>KGM3 - kilogram per cubic meter<br> MS - meters per second<br> KM - kilometer<br> M_ASL - meters above sea level (i.e. meters referenced to sea level)<br> MWM2 - milliwatt per square meter<br> TGCYR - terragram of carbon per year<br> TGYR - terragram per year<br> MA - megaannum<br> M - meters<br> MGCM2 - milligram of carbon per square meter<br> DEG - degree<br> S - seconds</p> <p>Statistics grids are calculated within a given radius (e.g. 10km, 50km, 125km, 250km, 500km, 1000km) of the respective cell-centered value. The statistics grids include mean (.men), average absolute deviation from the mean (.aad), and the common logarithm (.log) of the absolute value of the mean (.mlg). Additionally, some grids are a weighted count for given radii (e.g. seamounts) where weight is a cosine taper from the center of the grid cell. </p> <p>The grid pitch for this dataset is uniformly at 5-arc minute denoted by “.5m”. Additionally, the extension used (netCDF4) is denoted by “.nc”.</p>
Datasets for Watset: Local-Global Graph Clustering with Applications in Sense and Frame Induction
<p>This dataset supplements the article “<a href="https://doi.org/10.1162/COLI_a_00354">Watset: Local-Global Graph Clustering with Applications in Sense and Frame Induction</a>” published in the Computational Linguistics journal:</p> <ul> <li> <p><code>watset-coli-lcc-performance.tsv</code>: runtime analysis</p> </li> <li> <p><code>watset-coli-synsets.zip</code>: synset induction experiment (note that <code>pairwise-{en-babelnet,ru-rwn}.pkl</code> files are excluded due to the licensing issues)</p> </li> <li> <p><code>watset-coli-triframes.zip</code>: semantic frame induction experiment</p> </li> <li> <p><code>watset-coli-classes.zip</code>: semantic class induction experiment</p> </li> </ul>
Global monthly water temperature dataset, derived from dynamical 1-D water-energy routing model (DynWat) at 10 km spatial resolution
<pre>Global 10km spatial resolution water temperature dataset at the global scale for all major rivers, lakes and reservoirs. Data are provided at a monthly temporal resolution.</pre> <p>V1.1 update includes a improved version of the model removing some initial spikes related to rapid ice melt and streams that fall dry. The record has been reduced from 1981 tot 2014 to remove potential spinup impacts.</p> <p>The 1960-2010 data from v1.0 can be used for the earlier years.</p> <p>Consistent forcing is used for both time periods to remove potential biases that might occur otherwise.</p>
The global renewable power support policy dataset
<p>The global renewable power support policy dataset was compiled by Sarah Hafner (Anglia Ruskin University, United Kingdom) and Johan Lilliestam (Institute for Advanced Sustainability Studies (IASS), Germany) in February-July 2017 and completed during 2017. The work was led by Johan Lilliestam but each author gathered half of the data. The data was formatted and checked for internal consistency by Tim Tröndle, IASS.</p> <p>All non-commercial users are allowed to use and manipulate our data, but are required to give appropriate attribution. Hence, <strong>please cite this data as</strong>:</p> <p>Hafner, S. & Lilliestam, J. (2019): <em>The global renewable power support dataset</em>. Institute for Advanced Sustainability Studies (IASS) & Anglia Ruskin University, Potsdam & Cambridge. Doi: https://doi.org/ 10.5281/zenodo.3371375.</p> <p><strong>If you are interested in contributing</strong> to and further developing the dataset: please contact Johan Lilliestam (IASS Potsdam).</p> <p>The search was done in publically available sources, including but not limited to the IEA renewables policy database, res-legal.eu, Worldbank data, as well as data from the responsible national ministries.</p> <p>Our data holds information on 10 specific policy instruments explicitly dedicated to the support for expansion of renewable electricity generation 1990-2016; some instruments, including taxation of non-renewables or emission trading, affect other sectors than renewable power, but are mentioned in their original policy description to also be dedicated to increasing renewable power. Our data concerns national policy measures, but ignores policies enacted on higher (e.g. EU-level in Europe) or lower (e.g. state-level policies in Canada, USA) political levels. For example, the “no support” entry for the United Arab Emirates indicates that there were no national-level policies: all policies were, in this case, emirate-specific.</p> <p>The data exists in two versions: one version readable for humans (RE_policies_fullglobal.xlsx) and for each instrument type as .csv. The information in the two versions is identical and differs only in the way it is displayed.</p> <p><strong>Please refer to the metadata file for a detailed description of the dataset and the data categories.</strong></p>
A Dataset of Global Land Cover Validation Samples
<p>A dataset of global land cover validation samples in 2015. In order to guarantee the confidence and objective of the validation samples, several existing reference datasets such as GLCNMO2008 training dataset, VIIRS reference dataset, STEP reference dataset, Global cropland reference data and so on, high resolution imagery in the Google earth and time-series NDVI,NDSI values of each related point are integrated to derive the global validation datasets. The dataset is provided in .shp format.</p>
Global trends and collaborations in electrochemical methods: a dataset on etching and deposition research
<p><span>This dataset supports the study "Electrochemical Etching vs. Electrochemical Deposition: A Comparative Bibliometric Analysis," which examines scientific publications on electrochemical etching and electrochemical deposition from 1970 to 2023. The dataset is derived from the Science Citation Index Expanded (SCIE) database and includes bibliometric information on publication trends, leading contributors, research areas, and keyword co-occurrences in both fields.</span></p>
D-PLACE dataset derived from Wessel and Smith 2015 'Global Self-consistent, Hierarchical, High-resolution Geography Database'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Wessel, P., and W. H. F. Smith (1996), A global, self-consistent, hierarchical, high-resolution shoreline database, J. Geophys. Res., 101(B4), 8741–8743, doi:10.1029/96JB00104. Wessel P, Smith, W. H. F. Global Self-consistent, Hierarchical, High-resolution Geography Database (GSHHS) v2.3.4 [Internet]. 2015. Available: https://www.ngdc.noaa.gov/mgg/shorelines/gshhs.html</p> </blockquote>
D-PLACE dataset derived from Jenkins et al. 2013 'Global patterns of terrestrial vertebrate diversity and conservation'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Jenkins CN, Pimm SL, Joppa LN. Global patterns of terrestrial vertebrate diversity and conservation. Proc Natl Acad Sci. 2013;110: E2602–E2610.</p> </blockquote>
Global CHIRPS MCWD (Maximum Cumulative Water Deficit) Dataset
<p><strong>Global CHIRPS MCWD Dataset</strong></p> <p>The MCWD (Maximum Cumulative Water Deficit) is a measure of drought severity, which corresponds to the maximum value of the monthly accumulated water deficit reached for each pixel within the year. The MCWD is a useful indicator of meteorologically induced water stress without taking into account local soil conditions and plant adaptations, which are poorly understood in Amazonia. The full method of MCWD is described in Aragão et al. (2007; <a href="https://doi.org/10.1029/2006GL028946">https://doi.org/10.1029/2006GL028946</a>). Detail about CHIRPS (Rainfall Estimates from Rain Gauge and Satellite Observations) can be found in Funk et al. (2015; <a href="https://doi.org/10.1038/sdata.2015.66">https://doi.org/10.1038/sdata.2015.66</a>).</p> <p> </p> <p><strong>Coverage:</strong> Spanning 50°S-50°N (and all longitudes/land areas)</p> <p><strong>Period:</strong> 1981 to 2020</p> <p><strong>Spatial resolution:</strong> 0.05-degree</p> <p><strong>Temporal resolution:</strong> Annual</p> <p><strong>Coordinate reference system:</strong> Geographic Coordinate System (Datum WGS84)</p> <p><strong>File format:</strong> The zip file containing 40 files (one per year) in compressed TIFF format.</p> <p><strong>Code:</strong> <a href="https://zenodo.org/record/5034650">https://zenodo.org/record/5034650</a></p> <p><strong>Dataset usage</strong>: It is free to use, but if you use this dataset in your work, please make sure to cite the repository and our paper properly. We also welcome users to invite us for collaboration.</p> <p><strong>For the use of this dataset, please cite:</strong></p> <p>Silva Junior, C.H.L. et al. Global CHIRPS MCWD (Maximum Cumulative Water Deficit) Dataset. <em>Zenodo</em> (2021). DOI: 10.5281/zenodo.4903340. <a href="https://doi.org/10.1038/s41597-020-00600-4">https://doi.org/10.5281/zenodo.4903340</a></p> <p>Silva Junior, C.H.L. et al. Fire Responses to the 2010 and 2015/2016 Amazonian Droughts. <em>Front. Earth Sci</em>. (2019). DOI: 10.3389/feart.2019.00097. <a href="https://doi.org/10.3389/feart.2019.00097">https://doi.org/10.3389/feart.2019.00097</a></p> <p>Funk, C. et al. The climate hazards infrared precipitation with stations—a new environmental record for monitoring extremes. <em>Scientific Data</em> (2015). DOI: 10.1038/sdata.2015.66. <a href="https://doi.org/10.1038/sdata.2015.66">https://doi.org/10.1038/sdata.2015.66</a></p>
Global Wheat Head Dataset 2021
<p>This is the full Global Wheat Head Dataset 2021. Labels are included in csv.</p> <p>Tutorials available here: https://www.aicrowd.com/challenges/global-wheat-challenge-2021</p> <p> </p> <p>🕵️ Introduction</p> <p>Wheat is the basis of the diet of a large part of humanity. Therefore, this cereal is widely studied by scientists to ensure food security. A tedious, yet important part of this research is the measurement of different characteristics of the plants, also known as Plant Phenotyping. Monitoring plant architectural characteristics allow the breeders to grow better varieties and the farmers to make better decisions, but this critical step is still done manually. The emergence of UAV, camera and smartphone makes in-field RGB images more available and could be a solution to manual measurement. For instance, the counting of the wheat head can be done with Deep Learning. However, this task can be visually challenging. There is often an overlap of dense wheat plants, and the wind can blur the photographs, making identify single heads difficult. Additionally, appearances vary due to maturity, colour, genotype, and head orientation. Finally, because wheat is grown worldwide, different varieties, planting densities, patterns, and field conditions must be considered. To end manual counting, a robust algorithm must be created to address all these issues. </p> <p>💾 Dataset</p> <p>The dataset is composed of more than 6000 images of 1024x1024 pixels containing 300k+ unique wheat heads, with the corresponding bounding boxes. The images come from 11 countries and covers 44 unique measurement sessions. A measurement session is a set of images acquired at the same location, during a coherent timestamp (usually a few hours), with a specific sensor. In comparison to the 2020 competition on Kaggle, it represents 4 new countries, 22 new measurements sessions, 1200 new images and 120k new wheat heads. This amount of new situations will help to reinforce the quality of the test dataset. The 2020 dataset was labelled by researchers and students from 9 institutions across 7 countries. The additional data have been labelled by Human in the Loop, an ethical AI labelling company. We hope these changes will help in finding the most robust algorithms possible!</p> <p>The task is to localize the wheat head contained in each image. The goal is to obtain a model which is robust to variation in shape, illumination, sensor and locations. A set of boxes coordinates is provided for each image.</p> <p>The training dataset will be the images acquired in Europe and Canada, which cover approximately 4000 images and the test dataset will be composed of the images from North America (except Canada), Asia, Oceania and Africa and covers approximately 2000 images. It represents 7 new measurements sessions available for training but 17 new measurements sessions for the test!</p> <p>📁 Files</p> <p>Following files are available in the <code>resources</code> section:</p> <ul> <li> <p><code>images: the folder contains all images</code></p> </li> <li> <p><code>competition_train.csv , competition_val.csv, competition_test.csv : contains the splits used for the 2021 Global Wheat Challenge</code></p> <ul> <li> <p><code>Val contains the "public test", which is the test set of Global Wheat Head 2020</code></p> </li> <li> <p><code>Test contains the "private test".</code></p> </li> </ul> </li> <li> <p><code>Metadata.csv : contains additional metadatas for each domain</code></p> </li> </ul> <p>💻 Labels</p> <ul> <li>All boxes are contained in a csv with three columns <code>image_name</code>, BoxesString and domain</li> <li><code>image_name</code> is the name of the image, without the suffix. All images have a .png extension</li> <li>BoxesString is a string containing all predicted boxes with the format [x_min,y_min, x_max,y_max]. To concatenate a list of boxes into a PredString, please concatenate all list of coordinates with one space (" ") and all boxes with one semi-column ";". If there is no box, BoxesString is equal to "no_box".</li> <li>domain give the domain for each image</li> </ul> <p> </p> <p>If you use the dataset for your research, please do not forget to quote:</p> <pre>@article{david2020global, title={Global Wheat Head Detection (GWHD) dataset: a large and diverse dataset of high-resolution RGB-labelled images to develop and benchmark wheat head detection methods}, author={David, Etienne and Madec, Simon and Sadeghi-Tehran, Pouria and Aasen, Helge and Zheng, Bangyou and Liu, Shouyang and Kirchgessner, Norbert and Ishikawa, Goro and Nagasawa, Koichi and Badhon, Minhajul A and others}, journal={Plant Phenomics}, volume={2020}, year={2020}, publisher={Science Partner Journal} } </pre> <p>@misc{david2021global,<br> title={Global Wheat Head Dataset 2021: more diversity to improve the benchmarking of wheat head localization methods},<br> author={Etienne David and Mario Serouart and Daniel Smith and Simon Madec and Kaaviya Velumani and Shouyang Liu and Xu Wang and Francisco Pinto Espinosa and Shahameh Shafiee and Izzat S. A. Tahir and Hisashi Tsujimoto and Shuhei Nasuda and Bangyou Zheng and Norbert Kichgessner and Helge Aasen and Andreas Hund and Pouria Sadhegi-Tehran and Koichi Nagasawa and Goro Ishikawa and Sébastien Dandrifosse and Alexis Carlier and Benoit Mercatoris and Ken Kuroki and Haozhou Wang and Masanori Ishii and Minhajul A. Badhon and Curtis Pozniak and David Shaner LeBauer and Morten Lilimo and Jesse Poland and Scott Chapman and Benoit de Solan and Frédéric Baret and Ian Stavness and Wei Guo},<br> year={2021},<br> eprint={2105.07660},<br> archivePrefix={arXiv},<br> primaryClass={cs.CV}<br> }</p>
A global inventory of solar photovoltaic generating units - dataset
<p>This is the data repository accompanying Kruitwagen, L., Story, K., Friedrich, J., Byers, L., Skillman, S., & Hepburn, C. (2021) A global inventory of photovoltaic solar generating units, <strong>Nature, </strong><em>forthcoming</em>. This repository contains the training, cross-validation, test, and predicted data set as described in the publication. The contents of this repository are briefly summarised here, see the publication for further details.</p> <p><strong>Repository contents:</strong></p> <p><em>trn_tiles.geojson: </em>18,570 rectangular areas-of-interest used for sampling training patch data.</p> <p><em>trn_polygons.geojson: </em>36,882 polygons obtained from OSM in 2017 used to label training patches.</p> <p><em>cv_tiles.geojson: </em>560 rectangular areas-of-interest used for sampling cross-validation data seeded from <a href="https://www.wri.org/research/global-database-power-plants">WRI GPPDB</a></p> <p><em>cv_polygons.geojson: </em>6,281 polygons corresponding to all PV solar generating units present in cv_tiles.geojson at the end of 2018.</p> <p><em>test_tiles.geojson: </em>122 rectangular regions-of-interest used for building the test set.</p> <p><em>test_polygons.geojson: </em>7,263 polygons corresponding to all utility-scale (>10kW) solar generating units present in test_tiles.geojson at the end of 2018.</p> <p><em>predicted_polygons.geojson: </em>68,661 polygons corresponding to predicted polygons in global deployment, capturing the status of deployed photovoltaic solar energy generating capacity at the end of 2018.</p>
The dataset of the Global Collections survey of natural history collections
<p>From 2016 to 2018, we surveyed the world’s largest natural history museum collections to begin mapping this globally distributed scientific infrastructure. The resulting dataset includes 73 institutions across the globe. It has:</p> <ul> <li> <p>Basic institution data for the 73 contributing institutions, including estimated total collection sizes, geographic locations (to the city) and latitude/longitude, and Research Organization Registry (ROR) identifiers where available.</p> </li> <li> <p>Resourcing information, covering the numbers of research, collections and volunteer staff in each institution.</p> </li> <li> <p>Indicators of the presence and size of collections within each institution broken down into a grid of 19 collection disciplines and 16 geographic regions.</p> </li> <li> <p>Measures of the depth and breadth of individual researcher experience across the same disciplines and geographic regions.</p> </li> </ul> <p>This dataset contains the data (raw and processed) collected for the survey, and specifications for the schema used to store the data. It includes:</p> <ol> <li>A diagram of the MySQL database schema.</li> <li>A SQL dump of the MySQL database schema, excluding the data.</li> <li>A SQL dump of the MySQL database schema with all data. This may be imported into an instance of MySQL Server to create a complete reconstruction of the database.</li> <li>Raw data from each database table in CSV format.</li> <li>A set of more human-readable views of the data in CSV format. These correspond to the database tables, but foreign keys are substituted for values from the linked tables to make the data easier to read and analyse.</li> <li>A text file containing the definitions of the size categories used in the collection_unit table.</li> </ol> <p>The global collections data may also be accessed at<a href="https://rebrand.ly/global-collections"> https://rebrand.ly/global-collections</a>. This is a preliminary dashboard, constructed and published using Microsoft Power BI, that enables the exploration of the data through a set of visualisations and filters. The dashboard consists of three pages:</p> <p><strong>Institutional profile:</strong> Enables the selection of a specific institution and provides summary information on the institution and its location, staffing, total collection size, collection breakdown and researcher expertise.</p> <p><strong>Overall heatmap:</strong> Supports an interactive exploration of the global picture, including a heatmap of collection distribution across the discipline and geographic categories, and visualisations that demonstrate the relative breadth of collections across institutions and correlations between collection size and breadth. Various filters allow the focus to be refined to specific regions and collection sizes.</p> <p><strong>Browse:</strong> Provides some alternative methods of filtering and visualising the global dataset to look at patterns in the distribution and size of different types of collections across the global view.</p>
A High-Resolution Dataset of Global Urban Fraction for Mesoscale Urban Modelling
<p>Coupled urban-atmospheric models are extensively used to understand the urban environment and its impact on atmospheric processes. A common requirement of these models is information about the “urban fraction” (fraction of model grid covered by impervious surface area (ISA)). The European Space Agency (ESA) WorldCover product provides a global land cover map for the base year of 2020 and 2021 at a spatial resolution of 10 m. The dataset is based on Sentinel-1 and Sentinel-2 data with an overall accuracy of 74.4% (2020) and 76.7% (2021). In this study we process the WorldCover dataset and provide a ready-to-use “urban fraction” that can be incorporated in urban modelling systems. The dataset contains GeoTIFF and Weather Research and Forecasting Pre-processing System (WRF-WPS) format files for 1, 0.5, 0.25, 0.009 (~1 km), 0.0027 (~300 m), and 0.0009 (~100 m) degree spatial resolutions. The GeoTIFF files can be converted to other urban mesoscale modelling systems. Please check the README.txt for more information on using the dataset.</p> <p>Note: version 2.0.0 uses WorldCover 2021 v200 dataset for processing of urban fractions, while version 1.0.0 uses WorldCover 2020 v100 dataset.</p> <p>For more information please see here: <a href="https://1drv.ms/w/s!Ai5IcIuv5U4DioElr0E8CEq0DE3CSw?e=Wdp9i4">https://1drv.ms/w/s!Ai5IcIuv5U4DioElr0E8CEq0DE3CSw?e=Wdp9i4</a></p>
A global Lagrangian eddy dataset based on satellite altimetry (GLED v1.0)
<p>Mesoscale eddies, defined as rotating structures ranging typically from tens to hundreds of kilometers and lasting for several weeks to months, are ubiquitous in the global ocean. Isolated mesoscale eddies are generally considered as coherent structures with a material barrier that can trap the fluid within the eddy interior. Methods employed to identify coherent eddies can be classified into Eulerian and Lagrangian frameworks. Eddy datasets based on Eulerian methods, especially the eddy census of Chelton et al. (2011), have been used in a huge range of applications, from physics to biology. However, recent works have shown that Eulerian eddies are not necessarily coherent because there is strong and persistent water exchange across the Eulerian eddy boundary. In this study, millions of Lagrangian particles are advected by satellite-derived surface geostrophic velocities over a period of 1993-2019. Using the method of Lagrangian-averaged vorticity deviation by Haller et al. (2016), we present a global Lagrangian eddies dataset (GLED v1.0). This open-source dataset contains not only general features (eddy center position, equivalent radius, rotation property, etc.) of eddies with lifespans of 30, 90, and 180 days, but also the trajectory of particles trapped by coherent eddy boundaries over the lifetime. The greatest strength of GLED v1.0 is that the identified eddies are all material objects by construction. Our eddy dataset provides an additional option for oceanographers in studying the interaction between coherent eddies and other physical or biochemical processes in the Earth system.</p>
Dataset linking to the publication "An assessment of data sources, data quality and changes in national forest monitoring capacities in the Global Forest Resources Assessment 2005–2020"
<p>This dataset links to the study “An assessment of data sources, data quality and changes in national forest monitoring capacities in the Global Forest Resources Assessment 2005–2020”. This study is published in the journal “Environmental Research Letters” which can be found at <a href="https://iopscience.iop.org/article/10.1088/1748-9326/abd81b">https://iopscience.iop.org/article/10.1088/1748-9326/abd81b</a>. The dataset contains two files, one csv file, and one shape file. The two files contain the same data to meet the different users' needs. The dataset contains variables for assessing national forest monitoring data sources i.e., RS and/or NFI. Separate indicators namely 'Use of RS', and 'Use of NFI' were used to analyze the two data sources (RS and NFI). The description of each variable for these two indicators contained in the dataset is given in the Table below.</p> <table> <caption><strong>The description of the variables in the datase</strong>t <strong>for country capacity assessment</strong></caption> <tbody> <tr> <td><strong>Variables Name</strong></td> <td><strong>Description of the variables</strong></td> </tr> <tr> <td>Country</td> <td>Country</td> </tr> <tr> <td>ISO_A3_CODE</td> <td>ISO A3 Code for country</td> </tr> <tr> <td>ADM0_CODE</td> <td>ADMO Code for country</td> </tr> <tr> <td>CONTINENT</td> <td>Continent</td> </tr> <tr> <td>Region</td> <td>Region</td> </tr> <tr> <td>RSInd_05</td> <td>Use of remote sensing (RS) for forest area (change) monitoring 2005 Indicator</td> </tr> <tr> <td>RSSc_05</td> <td>Use of RS for forest area (change) monitoring 2005 Score</td> </tr> <tr> <td>RSInd _10</td> <td>Use of RS for forest area (change) monitoring 2010 Indicator</td> </tr> <tr> <td>RSSc _10</td> <td>Use of RS for forest area (change) monitoring 2010 Score</td> </tr> <tr> <td>RSInd_15</td> <td>Use of RS for forest area (change) monitoring 2015 Indicator</td> </tr> <tr> <td>RSSc _15</td> <td>Use of RS for forest area (change) monitoring 2015 Score</td> </tr> <tr> <td>RSInd_20</td> <td>Use of RS for forest area (change) monitoring 2020 Indicator</td> </tr> <tr> <td>RSSc _20</td> <td>Use of RS for forest area (change) monitoring 2020 Score</td> </tr> <tr> <td>DRS05_20</td> <td>Difference ‘use of RS’ 2005-2020</td> </tr> <tr> <td>NFIInd_05</td> <td>Use of national forest inventories (NFI) for forest monitoring 2005 Indicator</td> </tr> <tr> <td>NFISc_05</td> <td>Use of NFI for forest monitoring 2005 Score</td> </tr> <tr> <td>NFIInd _10</td> <td>Use of NFI for forest monitoring 2010 Indicator</td> </tr> <tr> <td>NFISc _10</td> <td>Use of NFI for forest monitoring 2010 Score</td> </tr> <tr> <td>NFIInd_15</td> <td>Use of NFI for forest monitoring 2015 Indicator</td> </tr> <tr> <td>NFISc _15</td> <td>Use of NFI for forest monitoring 2015 Score</td> </tr> <tr> <td>NFIInd_20</td> <td>Use of NFI for forest monitoring 2020 Indicator</td> </tr> <tr> <td>NFISc _20</td> <td>Use of NFI for forest monitoring 2020 Score</td> </tr> <tr> <td>DNFI05_20</td> <td>Difference ‘Use of NFI’ 2005-2020</td> </tr> </tbody> </table> <p>Indicators and Scores in the above Table for showing the use of RS and NFI data for forest monitoring in Figure 1 (1a, 1b, and 2a, 2b) are related in the following way.</p> <table> <caption><strong>The indicator values and scores of the country capacity assessment</strong></caption> <tbody> <tr> <td><strong>Indicator</strong></td> <td><strong>Score</strong></td> </tr> <tr> <td>Low</td> <td>0</td> </tr> <tr> <td>Limited</td> <td>1</td> </tr> <tr> <td>Intermediate</td> <td>2</td> </tr> <tr> <td>Good</td> <td>3</td> </tr> <tr> <td>Very Good</td> <td>4</td> </tr> </tbody> </table> <p>The capacity changes from 2005 to 2020 in Figure 1 (1c & 2c) are related in the following way.</p> <table> <caption><strong>The indicator values and levels for country capacity changes</strong></caption> <tbody> <tr> <td><strong>Capacity change values</strong></td> <td><strong>Capacity change levels</strong></td> </tr> <tr> <td>1,2,3,4</td> <td>Increase</td> </tr> <tr> <td>0</td> <td>No change</td> </tr> <tr> <td>-1,-2,-3,-4</td> <td>Decrease</td> </tr> </tbody> </table> <p> </p>
A hybrid 100-m global land cover dataset with Local Climate Zones for WRF
<p>This hybrid 100-m CGLC-MODIS-LCZ global land cover dataset is produced for the Weather Research and Forecasting (WRF) model starting from version 4.5. It is based on 1) the Copernicus Global Land Service Land Cover (CGLC, Buchhorn et al., 2021) product resampled to MODIS IGBP classes (CGLC-MODIS), and 2) the global map of Local Climate Zones (LCZ, Demuzere et al., 2022a, b) that describes the urban and built-up land surface. Both the CGLC and LCZ products are available at a 100-m spatial resolution, are representative for the year 2018, and cover -180°W to 180°E and -60°S to 78°N. Remaining areas are filled with the MODIS land cover classes. This dataset has been implemented into the WRF Preprocessing System (WPS) as <a href="https://www2.mmm.ucar.edu/wrf/users/download/get_sources_wps_geog.html">tiled binary data files</a> with <a href="https://github.com/wrf-model/WPS/blob/develop/geogrid/GEOGRID.TBL.ARW_LCZ">a new GEOGRID table entry</a> to allow WRF/WPS users to flexibly use this dataset in their studies particularly for urban modeling applications.</p> <p>To display the dataset in QGIS, <em>cmap_Qgis_CGLC_MOD_LCZ.txt </em>can be used as a color scheme.</p> <p>For more details, please read the technical documentation: <a href="https://doi.org/10.5281/zenodo.7670792">https://doi.org/10.5281/zenodo.7670792</a>.<br> <br> References:</p> <p><em>Buchhorn, M., Smets, B., Bertels, L., De Roo, B., Lesiv, M., Tsendbazar, N.-E., Li, L., Tarko, A. Copernicus Global Land Service: Land Cover 100m: version 3 Globe 2015-2019: Product User Manual (Dataset v3.0, doc issue 3.4). Product User Manual; Zenodo, Geneve, Switzerland, September 2020; doi: 10.5281/zenodo.3938963<br> <br> Demuzere M, Kittner J, Martilli A, et al. A global map of local climate zones to support earth system modelling and urban-scale environmental science. Earth Syst Sci Data. 2022a;14(8):3835-3873. doi:10.5194/essd-14-3835-2022</em></p> <p><em>Demuzere M, Kittner J, Martilli A, et al. (2022). Global map of Local Climate Zones (2.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6364593</em></p>
Global monthly and 0.1° gap-free XCO2 dataset
<p>This dataset consists of the global continental continuous column-averaged dry-air mole fraction of carbon dioxide (XCO2, unit: parts per million i.e., ppm) covering the period from September 2014 to December 2020. </p> <p>The dataset is derived from the official OCO-2 XCO2 product and incorporates data from multiple sources, including CO2 concentration data, vegetation index data, and meteorological data. We utilized a machine learning approach to generate a monthly-scale gapless CO2 product with a spatial resolution of 0.1°, stored in GeoTiff format.</p> <p> </p> <p><strong>Please cite the following article when using the dataset:</strong></p> <p>L. Zhang, T. Li, J. Wu, H. Yang, Global estimates of gap-free and fine-scale CO2 concentrations during 2014–2020 from satellite and reanalysis data, Environment International (2023), doi: https://doi.org/10.1016/j.envint.2023.108057</p>
Global Hydrogen Production during high-pressure Serpentinisation of Subducting Slabs—Dataset
<p>Data-set and code for recreating results of:</p> <p><strong>Global Hydrogen Production during high-pressure Serpentinisation of Subducting Slabs </strong></p> <p>A manuscript submitted to G-cubed.</p> <p> </p>
Global Biodiversity Information Facility (GBIF): an exhaustive list of gbif record ids, dataset keys, and their associated Occurrence IDs, Institution Code, Collection Codes and Catalog Numbers. hash://sha256/ea88f03a7bfd1ba853fdbea3203d54ab81ac3cdc8e8da7c96bbbba9c4b05d933 hash://md5/c49fe34785354847b37ea4509261e130
<p>The Global Biodiversity Information Facility (GBIF) indexes thousands of biodiversity datasets from Natural History Collections, citizen science initiatives (e.g., iNaturalist, eBird), and other sources. As part of the index process, GBIF associates at least two identifiers with indexed records: a record id (aka gbifID) and a dataset id (aka dataset key). These ids are central to do lookup, reference data, and package interpreted data products.</p> <p>This publication contains an exhaustive list of GBIF IDs and ids associated by their data providers as derived from:</p> <p>GBIF.org (01 March 2023) GBIF Occurrence Download https://doi.org/10.15468/dl.pk3trq</p> <p>The resource (size: ~260GB) provided by GBIF had content id hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 and was used to generate the resource included in this publication using</p> <pre><code class="language-bash">preston cat 'zip:hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97!/0015281-230224095556074.csv'\ | cut -f 1,2,3,37,38,39\ | gzip\ > gbifid.tsv.gz </code></pre> <p>with the content id of gbifid.tsv.gz (size: ~35GB) being hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8 .</p> <p>the first 10 lines of gbifid.tsv.gz as extracted via</p> <pre><code>preston cat --remote https://zenodo.org/record/7789866/files,https://linker.bio hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8\ | gunzip\ | head</code></pre> <p>are:</p> <pre><code>gbifID datasetKey occurrenceID institutionCode collectionCode catalogNumber 2997162320 c71c8000-9fc7-422c-804a-ce6abe751771 3399442 CEPEC CEPEC CEPEC00109669 2997162309 c71c8000-9fc7-422c-804a-ce6abe751771 2733085 CEPEC CEPEC CEPEC00000818 2997162317 c71c8000-9fc7-422c-804a-ce6abe751771 2733086 CEPEC CEPEC CEPEC00000888 2997162313 c71c8000-9fc7-422c-804a-ce6abe751771 3399443 CEPEC CEPEC CEPEC00109744 2997162306 c71c8000-9fc7-422c-804a-ce6abe751771 2733087 CEPEC CEPEC CEPEC00000889 2997162316 c71c8000-9fc7-422c-804a-ce6abe751771 3399440 CEPEC CEPEC CEPEC00109605 2997162324 c71c8000-9fc7-422c-804a-ce6abe751771 2733088 CEPEC CEPEC CEPEC00000890 2997162308 c71c8000-9fc7-422c-804a-ce6abe751771 3399441 CEPEC CEPEC CEPEC00109615 2997162303 c71c8000-9fc7-422c-804a-ce6abe751771 2733089 CEPEC CEPEC CEPEC00000891</code></pre> <p>Note that at time of writing, the html resource associated with the occurrence id 2997162320, and data set key c71c8000-9fc7-422c-804a-ce6abe751771 (extracted from of the first data row example above) are available via:</p> <p>https://gbif.org/occurrence/2997162320</p> <p>and</p> <p>https://gbif.org/dataset/c71c8000-9fc7-422c-804a-ce6abe751771</p> <p>respectively.</p> <p>This resource was initially created to help integrate with Bionomia (https://bionomia.net) to help associate people identifiers provided by bionomia to their original records via their GBIF ids. Bionomia re-uses GBIF records ids as a way to define links between records and the people (e.g., curators, collectors, identifiers) that worked on them. </p> <p>In other words, this resource provides a versioned translation table from the GBIF data universe (as defined by GBIF record ids, and dataset keys) to the data collections that exist (and evolve) independent of it. </p> <p>Note that the resource identified by hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 was not included in this publication it was too big (260GB) to fit. You may be able to retrieve the resource from its original location at https://api.gbif.org/v1/occurrence/download/request/0015281-230224095556074.zip .</p>
Datasets for Wohlfarth et al. (2023) An advanced thermal roughness model for airless planetary bodies - Implications for global variations of lunar hydration and mineralogical mapping of Mercury with the MERTIS spectrometer
<p>This document describes the datasets and modeling results presented and discussed in our full research article.<br> <br> Wohlfarth, K., Wöhler, C., Hiesinger, H., Helbert, J. 2023, An advanced thermal roughness model for airless planetary bodies - Implications for global variations of lunar hydration and mineralogical mapping of Mercury with the MERTIS spectrometer, Astronomy and Astrophysics, 672<br> <br> <a href="https://doi.org/10.1051/0004-6361/202245343">https://doi.org/10.1051/0004-6361/202245343</a><br> <br> We provide several visualization scripts that read and display the results for convenience. Access to the original MATLAB® code for the thermal model implementation is available upon request (<a href="mailto:kay.wohlfarth@tu-dortmund.de">kay.wohlfarth@tu-dortmund.de</a>).<br> <br> More info in Dataproducts.pdf</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.