Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

679

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

679 results for “retrieval”

Learn how ShareScore rates datasets ↗
zenodo36/100

Supplement to "Retrieval of an ice water path over the ocean from ISMAR and MARSS millimeter and submillimeter brightness temperatures"

<p>This data set is was used for the retrieval of ice water path of a precipitating frontal system west of the coast of Iceland using the millimeter/submillimeter radiometer International SubMillimetre Airborne Radiometer (ISMAR) and Microwave Airborne Radiometer Scanning System (MARSS) on board the (Facility for Airborne Atmospheric Measurements) FAAM BAE-146.</p> <p>It is a supplement to the journal article &quot;Retrieval of an ice water path over the ocean from ISMAR and MARSS millimeter and submillimeter brightness temperatures&quot;, which will be published in Atmospheric Measurement Techniques (AMT).</p> <p>The data set consists of</p> <ul> <li>the observed brightness temperatures of ISMAR and MARSS (FAAM_Flight_B897_radiometer-data.zip)</li> <li>the simulated brightness temperatures of ISMAR and MARSS and the corresponding atmospheric states (FAAM_Flight_B897_simulation.zip)</li> <li>the retrieval training database, which has been used to train the neural network retrieval.</li> </ul>

opencc-by-4.0Jan 2018View details →
zenodo36/100

CVL Database - An Off-line Database for Writer Retrieval, Writer Identification and Word Spotting

<p>The CVL Database is a public database for writer retrieval, writer identification and word spotting. The database consists of 7 different handwritten texts (1 German and 6 Englisch Texts). In total 310 writers participated in the dataset. 27 of which wrote 7 texts and 283 writers had to write 5 texts. For each text a rgb color image (300 dpi) comprising the handwritten text and the printed text sample is available as well as a cropped version (only handwritten). An unique id identifies the writer, whereas the Bounding Boxes for each single word are stored in an XML file.</p> <p>The CVL-database consists of images with cursively handwritten german and english texts which has been choosen from literary works. All pages have a unique writer id and the text number (separated by a dash) at the upper right corner, followed by the printed sample text. The text is placed between two horizontal separatores. Beneath the printed text individuals have been asked to write the text using a ruled undersheet to prevent curled text lines. The layout follows the style of the IAM database. The database was updated on 12/09/2013 since one writer ID (265/266) was wrong. The version number was changed to 1.1.</p> <p>Samples of the following texts have been used:</p> <ul> <li>Edwin A. Abbot &ndash; Flatland: A Romance of Many Dimension (92 words).</li> <li>William Shakespeare &ndash; Mac Beth (49 words).</li> <li>Wikipedia &ndash; Mail&uuml;fterl (73 words, under CC Attribution-ShareALike License).</li> <li>Charles Darwin &ndash; Origin of Species (52 words).</li> <li>Johann Wolfgang von Goethe &ndash; Faust. Eine Trag&ouml;die (50 words).</li> <li>Oscar Wilde &ndash; The Picture of Dorian Gray (66 words).</li> <li>Edgar Allan Poe &ndash; The Fall of the House of Usher (78 words).</li> </ul> <p>This database may be used for non-commercial research purpose only. If you publish material based on this database, we request you to include a reference to:</p> <p>Florian Kleber, Stefan Fiel, Markus Diem and Robert Sablatnig, <em>CVL-Database: An Off-line Database for Writer Retrieval, Writer Identification and Word Spotting</em>, In Proc. of the 12th Int. Conference on Document Analysis and Recognition (ICDAR) 2013, pp. 560-564, 2013.</p>

opencc-by-nc-4.0Nov 2018View details →
zenodo36/100

Towards improving short-term predictions of fine particulate matter over the United States via assimilation of satellite aerosol optical depth retrievals

<p>This dataset contains paired modeled and observed values of different trace gases and aerosol species over the CONUS for the period of 15 July to 14 August 2014 for the three different data assimilation experiments (BKG, MET_BE and MET+EMIS_BE) described in the paper.&nbsp;</p>

opencc-by-4.0Feb 2019View details →
zenodo36/100

Retrieval and summarization of microblogs posted after a disaster event: SMERP 2017 dataset

<p>This is the dataset used in the Data Challenge track of the ECIR 2017 Workshop on Exploitation of Social Media for Emergency Relief and Preparedness (<a href="https://www.computing.dcu.ie/~dganguly/smerp2017/">SMERP 2017</a>).</p> <p>The Data Challenge track was about extracting and summarizing information relevant to a set of practical information needs (topics) that are critical for post-disaster relief operations, such as need and availability of resources, infrastructure damage and restoration, etc. The track used a dataset of tweets / microblogs posted during the August 2016 earthquake in central Italy. Specifically, the data challenge consisted of two tasks:<br> (1) Retrieve the microblogs that are relevant to the given set of topics, and<br> (2) Summarizing the microblogs that are relevant to the given set of topics. &nbsp;</p> <p><br> This dataset can be used to develop algorithms for retrieval and summarization of microblogs that are useful for post-disaster relief operations, in the aftermath of a disaster.</p> <p>For more details, refer to the <a href="https://dl.acm.org/citation.cfm?id=3130338">workshop report.</a></p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Retrieving monthly and interannual pHT in the East China Sea shelf using an artificial neural network: ANN-pHT-v1

<p><br> The reliability of the artificial neural network model was&nbsp;evaluated by independent sampled data from 3 cruises in 2018.</p> <p>Monthly water column pHT for the period 2000-2016 was obtained passing T, S, DO, N, P, and Si from the Finite-Volume Coastal Ocean Model with the European Regional Sea Ecosystem Model through the artificial neural network. The spatiotemporal resolution of monthly pHT is 1-10 km in the horizontal, 10 depth levels in the vertical, and 12 months. Seasonal pHT dynamics in the East China Sea shelf can be primarily attributed to temperature changes and the shifting balance of production and respiration processes.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Prognostic value of a modified pathological staging system for gastric cancer based on the number of retrieved lymph nodes and metastatic lymph node ratio raw data

<p><span>Clinical data from the US Surveillance, Epidemiology, and End Results (SEER) Program from 2010-2015 (https://seer.cancer.gov/) was extracted and analyzed as training set, data from 2016-2017 was adopted as internal validation set. Data from The Cancer Genome Atlas Program (TCGA) (https://portal.gdc.cancer.gov/) and prognosis data from Gastrointestinal surgery Department, Third Affiliated Hospital of Sun Yat-sen University were applied as external validation sets. </span></p> <p><span>Screening criteria for gastric cancer cases were as follow: exclusion of cases with only autopsy or death certificate, cases where initial tumor location was not stomach, patients with stage 0 and stage IV, cases without radical surgery, non-adenocarcinoma cases, death cases within one month after operation, and cases with unknown lymph node information and AJCC TNM stage.</span></p> <p><span>The study analyzed various factors such as age of diagnosis (&lt;50 years, 50-69 years, &gt;69 years), gender, race (white, black, other), AJCC T stage (T1-T4b), AJCC TNM stage (I-III), primary tumor location (stomach body, antrum/pylorus, cardia/fundus, greater gastric recurve, lesser gastric recurve, overlapping area, NOS), Clinical features such as tumor size (&ge;5cm,&lt;5cm, unknown), tumor grade (I-IV), chemotherapy, radiotherapy, number of lymph nodes retrieved and number of metastases, and lymph node positive rate. The populations of American Indian/Alaskan and Asian/Pacific Islander were classified as "other" due to small sample sizes. Tumor grade was also analyzed, with grades I-IV representing highly differentiated, moderately differentiated, poorly differentiated, and signed-ring cell carcinoma, respectively. Overall survival (OS) is the time from cancer diagnosis to death from any cause, while disease-specific survival (DSS) is the time from cancer diagnosis to death specifically due to the disease.</span></p> <p><strong><span>&nbsp;</span></strong></p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Investigation of satellite vertical sensitivity on long-term retrieved lower tropospheric ozone trends

<p>Regional time-series (monthly mean) of lower tropospheric column ozone (LTCO3; 0-6 km or surface to 450 hPa) between 2008 and 2017 from three satellite products and an Earth System Model (UKESM1.0 - https://ukesm.ac.uk/). The regions of focus are North America, Europe and East Asia based on the HTAP-2 land mask (https://htap.org/). The three satellite products are from the Ozone Monitoring Instrument (OMI) (RAL Space - https://www.ralspace.stfc.ac.uk/Pages/Remote-Sensing.aspx), the Infrared Atmospheric Sounding Interferometer (IASI) FORLI (Fast Optimal Retrievals on Layers for IASI) scheme (https://iasi.aeris-data.fr/cos_iasi_b_arch/) and the IASI SOFRID (SOftware for Fast Retrievals of IASI Data) scheme (https://iasi-sofrid.sedoo.fr/). These data have been used to investigate long-term trends and investigation of satellite long-term discrepancies in retrieved LTCO3. The pre-print of the relevant manuscript can be found at https://doi.org/10.5194/egusphere-2023-3109.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana (dataset)

<p>The dataset contains all the data required to reproduce the experiments done in the paper &quot;Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana&quot;, published in the 25th International Conference on Theory and Practice of Digital Libraries (<a href="http://www.tpdl.eu/tpdl2021/">TPDL&#39;21</a>). In that work we&nbsp;run an experiment using the Europeana CH digital library as a use case, and we evaluated the effectiveness of a multilingual information retrieval strategy using machine translations to English as pivot language. We used the&nbsp;CEF translation service (eTranslation) for the translation&nbsp;of queries and content to English (<a href="https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation">https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation</a>).</p> <p>The dataset is&nbsp;also available at&nbsp;<a href="https://rnd-2.eanadev.org/share/crosslingual-search/">https://rnd-2.eanadev.org/share/crosslingual-search/</a>, and it is&nbsp;organized in four main folders:</p> <ul> <li><strong>queries</strong>: sample of 68 queries and their translations to English. The queries were&nbsp;issued in languages other than English from the Europeana Portal,&nbsp;using the Europeana&rsquo;s 1914-1918 thematic collection, between January and August 2019.</li> <li><strong>transcriptions</strong>:&nbsp;sample of 18,257 handwriting transcriptions&nbsp; and its translations to English. The transcriptions are taken&nbsp; from the Europeana 1914-1918 thematic collection, and obtained from the Transcribathon crowdsourcing platform (https://europeana.transcribathon.eu/).</li> <li><strong>solr_configuration</strong>: Apache Solr search engine configuration used in the experiments (which replicates the one used in Europeana).</li> <li><strong>results</strong>: manual evaluation of the query translations, and automatic evaluation of the multilingual&nbsp;retrieval.</li> </ul> <p>&nbsp;</p>

opencc-by-sa-4.0Jun 2021View details →
zenodo36/100

Retrieving 2D laterally varying structures from multi-station surface wave dispersion curves using multiscale window analysis

<p>Here are the waveform data used in the Geophysical Journal International paper entitled &quot;Retrieving 2D laterally varying structures from multi-station surface wave dispersion curves using multiscale window analysis&quot;.&nbsp;The dataset is used for the reader who wants to reproduce the result in the paper.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Belgian Statutory Article Retrieval Dataset (BSARD)

<p>The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval. BSARD consists of more than 22,600 statutory articles from Belgian law and about 1,100 legal questions posed by Belgian citizens and labeled by experienced jurists with relevant articles from the corpus.</p>

opencc-by-nc-sa-4.0Aug 2021View details →
zenodo36/100

Data Archive for: Hurricane Laura (2020): A Comparison of Drop Size Distribution Moments Using Ground and Radar Remote Sensing Retrieval Methods

<p>This archive corresponds to the data described in Brauer&nbsp;et al. (2021) to be published in&nbsp;<em>Journal of Geophysical Research: Atmospheres.</em>&nbsp;Please see the included readme.txt file for details about each data file.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Impact of 3D Cloud Structures on the Atmospheric Trace Gas Products from UV-VIS Sounders: Synthetic dataset for validation of trace gas retrieval algorithms

<p>This data set is described in detail in a paper submitted to AMTD:</p> <p><strong>Impact of 3D Cloud Structures on the Atmospheric Trace Gas Products from UV-VIS Sounders - Part I: Synthetic dataset for validation of trace gas retrieval algorithms</strong></p> <p>by Claudia Emde, Huan Yu, Arve Kylling, Michel van Roozendael, Kerstin Stebel, Ben Veihelmann, and<br> Bernhard Mayer</p> <p>&nbsp;</p> <p>The subdirectory <em>boxcloud</em> includes synthetic reflectances for clearsky, 1D cloud and box cloud.</p> <p>The subdirectory <em>les_cloud</em> includes synthetic reflectances for the LES cloud scenario for low earth orbit (<em>leo</em>) and geostationary orbit (<em>geo</em>).</p> <p>All data are provided in <em>netcdf</em> format.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Data and code for the manuscript Retrieving Water Vapor From an E-band Microwave Link With an Empirical Model Not Requiring In-situ Calibration

<p>Data and code for the manuscript <em>Retrieving Water Vapor From an E-band Microwave Link With an Empirical Model Not Requiring In-situ Calibration</em> accepted for publication to<em> </em> <em>Earth and Space Science</em> in October 2021.</p> <p>The dataset contains 7 month of total losses (transmitted - received power levels) and retrieved water vapor density from a 4.87 km long full-duplex E-band commercial microwave link (CML) operating at 73.5 and 83.5 GHz in Prague, CZ. The CML was operated as a part of a mobile phone backhaul. Furthermore, observations of air temperature, and air relative humidity from sites close to the CML end nodes are provided. Finally, theoretical gaseous attenuation calculated from the air temperature and relative humidity is included as a part of the dataset.</p> <p>Data are stored in semicolon-delimited csv files. Time stamps are in UTC time in the format yyyy-mm-dd HH:MM:SS. All time series are regular and have 5-min temporal resolution. Metadata are stored in text files.</p> <p>The code is in a form of R Markdown files and html notebooks. Results presented in the manuscript Retrieving Water Vapor From an E-band Microwave Link With an Empirical Model Not Requiring In-situ Calibration and in its Supporting information are fully reproducible using this dataset.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Gauge data used in 'An improved near-real-time precipitation retrieval for Brazil' by Pfreundschuh et al.

<p>The rain gauge measurements were compiled by the National Institute of Meteorology of Brazil 40 and consist of hourly gauge measurements covering the time range May 2000 until May 2020. They are used as reference data &nbsp;&#39;An improved near-real-time precipitation retrieval for Brazil&#39; by Pfreundschuh et al. and published here in order to ensure reproducibility.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Collected weakness description retrieved from the CWE database.

<p>Meta data of 464&nbsp;weaknesses introduced in the requirements, architecture and design, and policy phases retrieved from the CWE database.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Lidar size retrieval of soot aerosol

<p>The original data of figures shown in the paper: &quot;Size retrieval errors of soot particles observed using triple-wavelength lidar and the effects on recalculated optical parameters: A numerical investigation on fractal models&quot;.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

PAReTT: a Python package for the Automated Retrieval and management of divergence time data from the TimeTree resource for downstream analyses (Dataset)

<p>Dataset for article by the same title submitted the the <em>Journal of Molecular Evolution</em>.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Cross-spectra used in "Retrieval and precise phase-velocity estimation of Rayleigh waves by the spatial autocorrelation method between distributed acoustic sensing and seismometer data"

<p>Cross-spectra used in "Retrieval and precise phase-velocity estimation of Rayleigh waves by the spatial autocorrelation method between distributed acoustic sensing and seismometer data</p> <p>", by Shun Fukushima, Masanao Shinohara, Kiwamu Nishida, Akiko Takeo, Tomoaki Yamada, and Kiyoshi Yomogida&nbsp;</p> <p>For more information, please contact Shun Fukushima (s-fuku@eri.u-tokyo.ac.jp)</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

data set for Design of architectural environment integration of cyber-physical systems based on image segmentation and retrieval technology

<p>This data set is used to implement the project&nbsp;&nbsp;Design of architectural environment integration of cyber-physical systems based on image segmentation and retrieval technology</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

On retrieval system theory

<p>This paper re-reviews Vickery&rsquo;s book On retrieval system theory, first published 50 years ago,<br> and discusses the changing nature of theoretical work on information retrieval and the possibility of developing a general theory of IR. Stephen Robertson writes:<br> &lsquo;What kinds of theory or theories do we need for the field of information retrieval?&rsquo;<br> Brian Vickery&rsquo;s book whose title I have purloined was first published in 1961; in the preface he makes the following disclaimer: &lsquo;There is as yet no unified theory of retrieval systems&rsquo;. I have made the same statement many times myself, and it is as true now as it was a half-century ago. The number of papers published in the field of information retrieval, in every year of the first decade of the third millennium, would astonish the Brian Vickery of 1961, and many of these papers appeal to theoretical arguments of various more-or-less formal kinds. But can we expect, and do we need or want, a unified theory? In this talk I will attempt some discussion of these issues.</p>

opencc-by-4.0Jul 2011View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record