Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
679
datasets available to search
ShareScore release 0.7.1
Dataset results
679 results for “retrieval”
A ten-year (2010-2019) global terrestrial NEE inferred from the GOSAT v9 XCO2 retrievals (GCAS2021)
<p>Top-down atmospheric inversion is one of the major methods to estimate the global NEE. Here, by assimilating the GOSAT ACOS v9 XCO<sub>2</sub> product, we generate a ten-year (2010–2019) global monthly terrestrial NEE dataset using the Global Carbon Assimilation System, version 2 (GCASv2), which is named as GCAS2021. It includes(1) monthly and annual prior and posterior NEE and OCN fluxes, and prescribed FIRE and FFC emissions in a global spatial resolution of 1°×1°; (2) globally, latitudinally, and regionally aggregated monthly and annual posterior NEE and NBE, and their uncertainties; and (3) weekly gridded ensemble members of posterior NEE and OCN fluxes. The regional fluxes are aggregated both in the TRANSCOM and RECCAP2 regions. The latitudinal fluxes are aggregated in northern mid-high latitudes (> 30° N, NL), low latitudes (30° S ~ 30° N, TL), and southern middle latitudes (<30° N, SL). The weekly grided ensemble members could be used for calculating the posterior uncertainties of the user's area and time scale. </p> <p>Combining the OCN flux, and FIRE and FFC emissions, the net biosphere flux (NBE) and atmospheric growth rate (AGR) as well as their inter-annual variabilities (IAVs) agree well with the estimates of Global Carbon Budget 2020. Regionally, GCAS2021 shows that eastern North America, Amazon, Congo Basin, Europe, boreal forests, southern China and Southeast Asia are carbon sinks, while western US, African grasslands, Brazilian plateaus and parts of South Asia are carbon sources. In the TRANSCOM land regions, the NBEs of temperate N. America, northern Africa and boreal Asia are between the estimates of CMS-Flux NBE 2020 and CT2019B, and those in temperate Asia, Europe, and Southeast Asia are consistent with CMS-Flux NBE 2020 but significantly different from CT2019B. In the RECCAP2 regions, except for Africa and South Asia, the NBEs are comparable with the latest bottom-up estimate of Ciais et al. (2021). Compared with previous studies, the IAVs and seasonal cycles of NEE of this dateset could clearly reflect the impacts of extreme climates and large-scale climate anomalies on the carbon flux. The evaluations also show that the posterior CO<sub>2</sub> concentrations at remote sites and in regional scale, as well as on vertical CO<sub>2</sub> profiles in the Asia-Pacific region and the Amazon basin, are all consistent with independent CO<sub>2</sub> measurements from surface flask and aircraft CO<sub>2</sub> observations, indicating that this dataset captures surface carbon fluxes well. We believe that this data set will contribute to regional or national-scale carbon cycle and carbon neutrality assessment, and carbon dynamics research.</p>
Tropospheric delays and precipitable water vapor retrieved from global radiosonde observations from 2014 to 2019
<p>The data consists of a set of meteorological quantities including tropospheric delays (zenith wet delay, zenith hydrostatic delay, and zenith total delay), precipitable water vapor, and surface temperature and pressure. The data is retrieved from the observations of 414 globally distributed radiosonde stations from 2014 to 2019. In addition to the geographic information of radiosonde stations, the profiles of tropospheric delays and precipitable water vapor are contained in the data file. This data has a wide range of applications, e.g., validating the tropospheric delays and precipitable water vapor derived from other techniques, investigating the spatial-temporal variations of water vapor, and acting as training data of machine learning to build tropospheric delay models.</p> <p>In the manuscript "Machine Learning-based Model for Real-time GNSS Precipitable Water Vapor Sensing", this data is used to train a machine learning model to map the zenith total delays to precipitable water vapor. The data is split into training data and test data, where the data from 2014 to 2018 are employed for model training, and the data of 2019 are used for testing. The developed models and the results for the manuscript are saved in the directories of Models and Results, respectively.</p>
Data from: Clathrin-independent endocytic retrieval of SV proteins mediated by the clathrin adaptor AP-2 at mammalian central synapses
<p><span>Neurotransmission is based on the exocytic fusion of synaptic vesicles (SVs) followed by endocytic membrane retrieval and the reformation of SVs. Conflicting models have been proposed regarding the mechanisms of SV endocytosis, most notably clathrin/ AP-2-mediated endocytosis and clathrin-independent ultrafast endocytosis. Partitioning between these pathways has been suggested to be controlled by temperature and stimulus paradigm. We report on the comprehensive survey of six major SV proteins to show that SV endocytosis in mouse hippocampal neurons at physiological temperature occurs independent of clathrin while the endocytic retrieval of a subset of SV proteins including the vesicular transporters for glutamate and GABA depend on sorting by the clathrin adaptor AP-2. Our findings highlight a clathrin-independent role of the clathrin adaptor AP-2 in the endocytic retrieval of select SV cargos from the presynaptic cell surface and suggest a revised model for the endocytosis of SV membranes at mammalian central synapses.</span></p>
ir_metadata: An Extensible Metadata Schema for Information Retrieval Experiments
<p>This dataset accompanies our work that introduces a metadata schema for TREC run files based on the PRIMAD model. PRIMAD considers essential components of computational experiments that possibly can affect reproducibility on a conceptual level. We propose to align the metadata annotations to the PRIMAD components. In order to demonstrate the potential of metadata annotations, we curated a dataset with run files derived from experiments with different instantiations of PRIMAD components and annotated these with the corresponding metadata. With this work, we hope to stimulate IR researchers to annotate run files and improve the reuse value of experimental artifacts even further.</p> <p> </p> <p>This archive contains the following data:</p> <ul> <li> <p><strong>demo.tar.xz</strong> : Selected annotated runs files that are used in the Colab demonstration.</p> </li> <li> <p><strong>metadata.zip</strong> : YAML files containing only the metadata annotations for each run.</p> </li> <li> <p><strong>runs.zip</strong> : The entire set of run files with annotations.</p> </li> </ul> <p> </p> <p>The annotated runs result from the following experiments:</p> <ul> <li> <p>Grossman and Cormack @ TREC Common Core 2017 <a href="https://trec.nist.gov/pubs/trec26/papers/MRG_UWaterloo-CC.pdf">Paper</a> | <a href="https://trec.nist.gov/">Source</a></p> </li> <li> <p>Grossman and Cormack @ TREC Common Core 2018 <a href="https://trec.nist.gov/pubs/trec27/papers/MRG_UWaterloo-CC.pdf">Paper</a> | <a href="https://trec.nist.gov/">Source</a></p> </li> <li> <p>Yu et al. @ TREC Common Core 2018 <a href="https://trec.nist.gov/pubs/trec27/papers/h2oloo-CC.pdf">Paper</a> | <a href="https://github.com/castorini/Anserini/blob/master/docs/runbook-trec2018-h2oloo.md">Source</a></p> </li> <li> <p>Yu et al. @ ECIR 2019 <a href="https://link.springer.com/chapter/10.1007/978-3-030-15712-8_26">Paper</a> | <a href="https://github.com/castorini/anserini/blob/master/docs/runbook-ecir2019-ccrf.md">Source</a></p> </li> <li> <p>Breuer et al. @ SIGIR 2020 <a href="https://dl.acm.org/doi/10.1145/3397271.3401036">Paper</a> | <a href="https://zenodo.org/record/3856042">Source</a></p> </li> <li> <p>Breuer et al. @ CLEF 2021 <a href="https://link.springer.com/chapter/10.1007/978-3-030-85251-1_5">Paper</a> | <a href="https://zenodo.org/record/4105885">Source</a></p> </li> </ul>
Dataset of synergistically retrieved water vapor from SPICAM and PFS nadir observations
<p>Compilation of processed data files used for the creation of the figures presented in the manuscript "Constraining near-surface water vapor on Mars: a spectral synergy climatological survey applied to PFS and SPICAM nadir observations".</p>
Mesoscale stereo retrievals from Hunga Tonga-Hunga Ha'apai Eruption of 15 January 2022
<p>Stereo methods using GOES-17 and Himawari-8 applied to the Hunga Tonga-Hunga Ha'apai volcanic plume on 15 January 2022 show overshooting tops reaching 50-55 km altitude, a record in the satellite era. Plume height is important to understand dispersal and transport in the stratosphere and climate impacts. Stereo methods, using geostationary satellite pairs, offer the ability to accurately capture the evolution of plume top morphology quasi-continuously over long periods. Manual photogrammetry estimates plume height during the most dynamic early phase of the eruption and a fully automated algorithm retrieves both plume height and advection every 10 minutes during a more frequently sampled and stable phase beginning three hours after the eruption. Stereo heights are confirmed with Global Navigation Satellite System Radio Occultation (GNSS-RO) bending angles, showing that much of the plume was lofted 30–40 km into the atmosphere. Cold bubbles are observed in the stratosphere with brightness temperature of ~173K.</p>
Fig. 1 in The Relationship Between Fish Length And Otolith Size And Weight Of The Australian Anchovy, Engraulis Australis (Clupeiformes, Engraulidae), Retrieved From The Food Of The Australasian Gannet, Morus Serrator (Suliformes, Sulidae), Hauraki Gulf, New Zealand
Fig. 1. Map showing the location of the gannet's colonies in
IAA/CSIC Temperature and CO2 density profiles in Mars Year 34 retrieved from the 1st year of NOMAD/TGO solar occultation observations [dataset]
<p>Dataset associated to manuscript 2022JE007278, "Martian atmospheric temperature and density profiles during the 1st year of NOMAD/TGO solar occultation measurements", submitted on Feb 28th, 2022, to Journal of Geophysical Research - Planets (AGU) for the Special Issue "ExoMars Trace Gas Orbiter: One Martian Year of Science".</p> <p>Format and Number of datafiles: 1 README.txt and 1 tar file containing 325 ASCII files, one for each retrieved NOMAD scan. Each of these files contains one vertical profile of 7 parameters: altitude (km), Temperature in the last iteration (K), atmospheric pressure (mb), Temperature of First Guess (K), pressure of First Guess (mb), Temperature retrieval error (K) and diagonal element fo the Averaging Kernel matrix at that tangent altitude (normalized to 0-1). Also, the header of each file contains extra info, like the NOMAD "internal" data filename, and the ranges in latitude, longitude, Local time, Solar Longitude, tangent heights observed, and Line-of-Sight shift needed to obtain a good fit, together with the number of altitudes in the retrieval vector (usually around 100 km). See README file for more details.<br> </p>
Bumblebees retrieve only the ordinal ranking of foraging options when comparing memories obtained in distinct settings
<p><span>Are animals' preferences determined by absolute memories for options (e.g. reward sizes) or by their remembered ranking (better/worse)? The only studies examining this question suggest humans and starlings utilise memories for both absolute and relative information. We show that bumblebees' learned preferences are based only on memories of ordinal comparisons. A series of experiments showed that after learning to discriminate pairs of different flowers by sucrose concentration, bumblebees preferred flowers (in novel pairings) with 1) higher ranking over equal absolute reward, 2) higher ranking over higher absolute reward, and 3) identical qualitative ranking but different quantitative ranking equally. Bumblebees used absolute information in order to rank different flowers. However, additional experiments revealed that, even when ranking information was absent (i.e. bees learned one flower at a time), memories for absolute information were lost or could no longer be retrieved after at most one hour. Our results illuminate a divergent mechanism for bees (compared to starlings and humans) of learned preferences that may have arisen from different adaptations to their natural environment.</span></p>
Large-scale grid computing for content-based image retrieval
<p>The author presents an approach in which a large distributed processing Grid has been used to apply a range of content-based image retrieval methods to a substantial number of images. By massively distributing the required computational task across thousands of Grid nodes, we have achieved very high throughput at relatively low overheads.</p> <p> </p>
TEAMx-PC22 (TEAMx pre-campaign 2022) - Animations of radial velocity and coplanar-retrieved horizontal wind fields from KITcube Leosphere/Vaisala Windcube WLS200s-124 and WLS200s-159
<p><strong>Abstract</strong></p> <p>This data set was collected during the TEAMx pre-campaign in summer 2022 (TEAMx-PC22) in the Inn Valley Target Area, Austria.</p> <p><strong>Data Description</strong></p> <p>This data set is comprised of 64 daily .mp4 files showing:</p> <ul> <li>post-processed radial velocity fields sampled by KITcube Leosphere/Vaisala Windcube WLS200s-124 and WLS200s-159</li> <li>their coplanar-retrieved horizontal wind field output</li> <li>additionally, horizontal wind speed and direction sampled by the KITcube Vaisala Windcube v2.1 (WLS7-1489) at 60-m above ground level is added as an independent measurement for subjective validation of the coplanar-retrieved wind</li> </ul> <p>The wind fields shown in the animations originate from an accompanying Zenodo data set (DOI: 10.5281/zenodo.7212801), where complete information concerning lidar locations, scan parameters, and post-processing may be found.</p>
Planetary Boundary Layer Height Retrievals from Micropulse-lidar at four Multiple ARM Sites Around the World
<p><span>Planetary Boundary Layer Height Retrievals from Micropulse-lidar at the following ARM campaigns: GOAMAZON (MAO), COPS (FKB), CACTI (COR), and BAECC (TMP).</span><span> These retrievals </span><span>were computed</span><span> using the Different Thermo-Dynamic Stabilities (DTDS) algorithm </span><span>(Su et al., 2020; Su et al., 2022)</span><span>. The quality-control flag is provided in the dataset file, where zero (0) indicates a high-quality flag. </span></p> <p>Su, T., Zheng, Y. and Li, Z., 2022. Methodology to determine the coupling of continental clouds with surface and boundary layer height under cloudy conditions from lidar and meteorological data. Atmospheric Chemistry and Physics, 22(2), pp.1453-1466.</p> <p>Su, T., Li, Z. and Kahn, R., 2020. A new method to retrieve the diurnal variability of planetary boundary layer height from lidar under different thermodynamic stability conditions. Remote Sensing of Environment, 237, p.111519.</p> <p><span>Roldán-Henao, N., Su, T., and Li, Z. (2024, under review). Refining Planetary Boundary Layer Height Retrievals from Micropulse-lidar at Multiple ARM Sites Around the World. Submitted to <em><span>Journal of Geophysical Research: Atmospheres. </span></em></span></p>
Synthetic temporal dataset for temporal trend analysis and retrieval
<p>This repository contains a synthetic, temporal data set that was generated by the authors by sampling values from the Gaussian distribution. The dataset contains eight nontemporal dimensions, a temporal dimension, and a numerical measure attribute. The data set was generated according to the scheme and procedure detailed in this source paper: Kaufmann, M., Fischer, P.M., May, N., Tonder, A., Kossmann, D. (2014). TPC-BiH: A Benchmark for Bitemporal Databases. In: Performance Characterization and Benchmarking. TPCTC 2013. Lecture Notes in Computer Science, vol 8391. Springer, Cham. https://doi.org/10.1007/978-3-319-04936-6_2. The data set can be used for analyzing and locating temporal trends of interest, where a temporal trend is generated by selecting the desired values of the nontemporal dimensions, and then selecting the corresponding values of the temporal dimension and the numerical measure attribute. Locating temporal trends of interest, e.g., unusual trends, is a common task in many applications and domains. It can also be of interest to understand which nontemporal dimensions are associated with the temporal trends of interest. To this end, the data set can be used for analyzing and locating temporal trends in the data cube induced by the data set.</p>
Retrieve, Merge, Predict: Augmenting Tables with Data Lakes
<p>Files composing the YADL data lake, for the paper "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes (Experiment, Analysis & Benchmark Paper)"</p> <p>We present an in-depth analysis of data discovery for analytics in data lakes, focusing on table augmentation for given machine learning tasks. We analyze alternative methods used in the three key steps: retrieving joinable tables, merging information, and predicting with the resultant table. As data lakes, the paper uses YADL (Yet Another Data Lake) -- a novel dataset developed as a tool for benchmarking this data discovery task -- and Open Data US, a well-referenced real data lake. Through systematic exploration on both lakes, our study outlines the importance of accurately retrieving join candidates, and the efficiency of simple aggregation methods. We report new insights on the benefits of existing solutions and on the their limitations, aiming at guiding future research in this space.</p> <p>Archives provided here follow the notation used for the experiments, which is different from what is reported in the paper. The four YADL versions available here are:</p> <ul> <li>"binary_update" (YADL Binary)</li> <li>"wordnet_full" (YADL Base)</li> <li>"wordnet_vldb_10" (YADL 10k)</li> <li>"wordnet_vldb_50" (YADL 50k)</li> </ul>
The role of OCO-3 XCO2 retrievals in estimating global terrestrial net ecosystem exchanges
<p>1.<strong>Exp_OCO3.zip </strong> includes regional posterior carbon fluxes from the assimilation using OCO-3 observations.</p> <p>2.<strong>Exp_OCO2.zip </strong> includes regional posterior carbon fluxes from the assimilation using OCO-2 observations.</p> <p>3.<strong>Exp_OCO3&2.zip </strong> includes regional posterior carbon fluxes from the joint assimilation using OCO-3 and OCO-2 observations together.</p> <p>4.<strong>posterior.fluxes.Exp_OCO3.nc </strong>includes information on the spatial distribution of annual as well as monthly posterior carbon fluxes from the assimilation using OCO-3 observations during August 2019 to December 2022.</p> <p>5.<strong>posterior.fluxes.Exp_OCO2.nc </strong>includes information on the spatial distribution of annual as well as monthly posterior carbon fluxes from the assimilation using OCO-2 observations during August 2019 to December 2022.</p> <p>6.<strong>posterior.fluxes.Exp_OCO3&2.nc </strong>includes information on the spatial distribution of annual as well as monthly posterior carbon fluxes from the joint assimilation using OCO-3 and OCO-2 observations together during August 2019 to December 2022.</p> <p>7.<strong>evaluation_result.txt</strong> includes the results of the evaluation of posterior carbon fluxes using independent CO2 observations from 66 surface flask sites.</p> <p> </p> <p> </p> <p> </p>
Data: Soil moisture modeling with ERA5-Land retrievals, topographic indices, and in situ measurements and its use for predicting ruts
<p>Data for: <br><br>Soil moisture modeling with ERA5-Land retrievals, topographic indices, and in situ measurements and its use for predicting ruts</p> <p>Marian Schönauer<sup>1</sup>, Anneli M. Ågren<sup>2</sup>, Klaus Katzensteiner<sup>3</sup>, Florian Hartsch<sup>1</sup>, Paul Arp<sup>4</sup>, Simon Drollinger<sup>5</sup>, Dirk Jaeger<sup>1</sup></p> <p><sup>1</sup>Department of Forest Work Science and Engineering, University of Göttingen, Göttingen, Germany</p> <p><sup>2</sup>Department of Forest Ecology and Management, Swedish University of Agricultural Sciences, Umeå, Sweden</p> <p><sup>3</sup>Institute of Forest Ecology, University of Natural Resources and Life Sciences, Vienna, Vienna, Austria</p> <p><sup>4</sup>Forestry and Environmental Management, University of New Brunswick, New Brunswick, Canada</p> <p><sup>5</sup>Department of Physical Geography, University of Göttingen, Göttingen, Germany</p>
Dataset Retrieval Benchmark
<p>The sample of dataset retrieval benchmark corpus:</p> <ul> <li><strong>docs.tsv</strong> contains 9802 dataset metadata records</li> <li><strong>topics.tsv, keyword_queries.tsv</strong> contain 51 original post and keywords query for each post</li> <li><strong>qrels.tsv</strong> contains 57 relevance judgements for provided sample of queries</li> </ul>
Fig. 1 in Retrieving climate change dependent Sea Surface Temperature (SST) in Southern Turkey by using Landsat thermal imagery
Fig. 1 — Map of the study area, Bay of Goköva
Fig. 4 in Retrieving climate change dependent Sea Surface Temperature (SST) in Southern Turkey by using Landsat thermal imagery
Fig. 4 — Average annual temperature in Gökova Bay
Research Compendium for Himes et al. (2024): "Using neural networks for near-real-time aerosol retrievals from OMPS Limb Profiler measurements"
<p>This archive is the Reproducible Research Compendium for</p> <p>Using neural networks for near-real-time aerosol retrievals from OMPS Limb Profiler measurements</p> <p>by Himes et al. (2024), submitted to Atmospheric Measurement Techniques.</p> <p>This compendium includes all files related to MARGE associated with the manuscript.</p> <p>NN model files are split into smaller files for convenience, given their sizes. To recombine the files, do, e.g., <br> cat cnn_weights_NH-LW.h5* > cnn_weights_NH-LW.h5</p> <p>User interested in running MARGE will need to clone the GitHub repo (https://github.com/exosports/MARGE), apply the patch file to checksum fc95b3c, organize the relevant files into directories as listed in the configuration files (Zenodo does not support organizing files into directory structures) and calculate the number of training, validation, and test cases to be stored in the relevant input file specified in the configuration file. MARGE is under the Reproducible Research Software License (https://planets.ucf.edu/resources/reproducible-research/software-license/). For more details on MARGE, see the User Manual on GitHub.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.