Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

90

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

90 results for “data retrieval”

Learn how ShareScore rates datasets ↗
zenodo32/100

WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models

<p>#############</p> <h1>WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models</h1> <p>#############</p> <p>Authors: Valentin Gabeff, Marc Russwurm, Devis Tuia &amp; Alexander Mathis</p> <p>Affiliation: EPFL</p> <p>Date: January, 2024</p> <p>Link to the article: <a href="https://link.springer.com/article/10.1007/s11263-024-02026-6">https://link.springer.com/article/10.1007/s11263-024-02026-6</a></p> <p>--------------------------------</p> <p>WildCLIP is a fine-tuned CLIP model that allows to retrieve camera-trap events with natural language from the Snapshot Serengeti dataset. This project intends to demonstrate how vision-language models may assist the annotation process of camera-trap datasets.</p> <p>Here we provide the processed Snapshot Serengeti data used to train and evaluate WildCLIP, along with two versions of WildCLIP (model weights).</p> <p>Details on how to run these models can be found in the project <a href="https://github.com/amathislab/wildclip">github repository</a>.</p> <h2>Provided data (images and attribute annotations):&nbsp;</h2> <p>The data consists of 380 x 380 image crops corresponding to the MegaDetector output of Snapshot Serengeti with a confidence threshold above 0.7. We considered only camera trap images containing single individuals.</p> <p>A description of the original data can be found on LILA <a href="https://lila.science/datasets/snapshot-serengeti">here</a>, released under the <a href="https://cdla.dev/permissive-1-0/" rel="nofollow">Community Data License Agreement (permissive variant)</a>.</p> <p>We warmly thank the authors of LILA for making the MegaDetector outputs publicly available, as well as for structuring the dataset and facilitating its access.</p> <h2>Adapted CLIP model (model weights):&nbsp;</h2> <p>WildCLIP models provided:</p> <ul> <li><strong>[New] WildCLIP_vitb16_t1.pth:&nbsp;</strong>CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1. Trained on both base and novel vocabulary (see paper for details).</li> <li><strong>[New] WildCLIP_vitb16_t1_lwf.pth:&nbsp;</strong>CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1, and with the additional VR-LwF loss. Trained on both base and novel vocabulary (see paper for details).</li> <li><strong>WildCLIP_vitb16_t1_base.pth:</strong> CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1. Model used for evaluation and trained on base vocabulary only. (previously named <em>WildCLIP_vitb16_t1.pth</em>)</li> <li><strong>WildCLIP_vitb16_t1t7_lwf_base.pth</strong>: CLIP model with the ViT-B/16 visual backbone trained on data with captions following templates 1 to 7, and with the additional VR-LwF loss. Model used for evaluation and trained on base vocabulary only.&nbsp;(previously named <em>WildCLIP_vitb16_t1t7_lwf.pth</em>)</li> </ul> <p>We also provide the CSV files containing the train / val / test splits. The train / test splits follow camera split from LILA (https://lila.science/datasets/snapshot-serengeti). The validation split is custom, and also at the camera level.</p> <ul> <li><strong>train_dataset_crops_single_animal_template_captions_T1T7_ID.csv</strong>: Train set with captions from templates 1 through 7 (column "all captions") or template 1 only (column "template 1")</li> <li><strong>val_dataset_crops_single_animal_template_captions_T1T7_ID.csv</strong>: Validation set with captions from templates 1 through 7 (column "all captions") or template 1 only (column "template 1")</li> <li><strong>test_dataset_crops_single_animal_template_captions_T1T8T10.csv</strong>: Test set with captions from templates 1, 8, 9 and 10 (columns "all captions")</li> </ul> <p>Details on how the models were trained can be found in the associated&nbsp;<a href="https://link.springer.com/article/10.1007/s11263-024-02026-6" target="_blank" rel="noopener">publication</a>.</p> <h2>References:&nbsp;</h2> <p>If you find our code, or weights, please cite:</p> <pre>@article{gabeff2024wildclip, title={WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models}, author={Gabeff, Valentin and Ru{\ss}wurm, Marc and Tuia, Devis and Mathis, Alexander}, journal={International Journal of Computer Vision}, pages={1--17}, year={2024}, publisher={Springer} }</pre> <p>If you use the adapted Snapshot Serengeti data please also cite their article:</p> <pre>@article{swanson2015snapshot, title={Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna}, author={Swanson, Alexandra and Kosmala, Margaret and Lintott, Chris and Simpson, Robert and Smith, Arfon and Packer, Craig}, journal={Scientific data}, volume={2}, number={1}, pages={1--14}, year={2015}, publisher={Nature Publishing Group} }</pre>

opencdla-permissive-1.0Dec 2023View details →
zenodo32/100

Instruction about codes and data produced by the study entitled: Diurnal carbon monoxide retrieval from FY-4B/GIIRS using a novel machine learning method

<p>Instruction about codes and data produced by the study entitled: Diurnal carbon monoxide retrieval from FY-4B/GIIRS using a novel machine learning method.</p> <p>Please note that <strong>the manuscript is under review</strong> in a peer-reviewed journal.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Data for "Improving semantic video retrieval models by training with a relevance-aware online mining strategy"

<p>This repository contains all the data available for the publication:</p> <p><a href="https://doi.org/10.1016/j.cviu.2024.104035">Alex Falcon, Giuseppe Serra, and Oswald Lanz.&nbsp;<em>Improving semantic video retrieval models by training with a relevance-aware online mining strategy</em>. <strong>Computer Vision and Image Understanding</strong>. 2024.</a></p> <p>Code is available at: <a href="https://github.com/aranciokov/ranp/">https://github.com/aranciokov/ranp/</a></p> <p>The data includes:</p> <ul> <li>pre-extracted features (ordered_feature_*.zip files)</li> <li>annotations, such as pre-extracted semantic graphs, glove checkpoints, class annotations, etc (annotations_*.zip files)</li> <li>train/val/test, when available, split information (public_split_*.zip) files</li> <li>pretrained models for HGR and EAO (details in the github repo)</li> </ul>

opencc-by-4.0May 2024View details →
zenodo32/100

Test and Train data for the retrieval experiment in "A molecule generation-oriented lead compound optimization architecture: discovery of potent, selective, oral NLRP3 inflammasome inhibitors"

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Data for "Using Simulated Radiances to Understand the Limitations of Satellite Retrieved Volcanic Ash Data and the Implications for Volcanic Ash Cloud Forecasting"

<p>This location contains all of the data used in the analysis for the paper "Using Simulated Radiances to Understand the Limitations of Satellite Retrieved Volcanic Ash Data and the Implications for Volcanic Ash Cloud Forecasting" which is currently in prep.</p> <p>All of the retrieved satellite data can be seen in the retrieved_satellite_data.zip folder. the data is organised by the input ash cloud properties being simulated and the hdf files contain all of the retrieved variables where ash has been successfully detected.</p> <p>All of the output dispersion model data is available in NAME_output_data.zip.&nbsp;</p> <p>All of the input source data used in the dispersion model simulations (including the data from REFIR) is available in NAME_source_data_REFIR.zip.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Volcanic ash source inversion data for paper "A near-real-time method for estimating volcanic ash emissions using satellite retrievals"

<p>This dataset consists of volcanic ash source inversion data for the paper &quot;A near-real-time method for estimating volcanic ash emissions using satellite retrievals&quot; by Rachel E. Pelley, David J. Thomson, Helen N. Webster, Michael C. Cooke, Alistair J. Manning, Claire S. Witham and Matthew C. Hort, Atmosphere, 2021, 12, 1573, https://doi.org/10.3390/atmos12121573. Satellite retrievals, dispersion model simulations and inversion calculations are included for the eruptions of Eyjafjallajokull in 2010 and Grimsvotn in 2011.</p>

opencc-by-4.0Jun 2019View details →
zenodo32/100

An Integrated Approach for enhanced SMAP Soil Moisture Retrieval: Multi-Source Data Fusion and Data-Driven Machine Learning

<p><span>Accurate satellite-based soil moisture (SM) retrieval is essential for hydrometeorological and agroecological applications, yet traditional physical models for L-band SM retrieval are hindered by uncertainties stemming from inaccuracies in prior parameters. This work combines multi-source data fusion and a physically-guided machine learning framework to develop a Soil Moisture Active Passive (SMAP) SM retrieval model (Fusion-LightGBM, F-LGB) that bypasses the need for static prior parameters, resulting in a new SM product. The retrieval benchmark is a new seamless SM data constructed by combining Triple Collection correlation coefficients (TC-R) and the Maximized-R method, which demonstrates superior temporal correlation on 20 International Soil Moisture Network (ISMN)&nbsp;<em>in-situ</em> networks compared to existing SM data, including ECMWF Reanalysis v5-Land (ERA5-Land), SMAP Level 4 (SMAP L4), and Global Land Data Assimilation System (GLDAS) Noah. The machine learning model incorporates input variables that represent the Tau-Omega model&rsquo;s radiative transfer process, including brightness temperature, vegetation optical depth, soil temperature, and an external variable for precipitation. In the 2015-2020 validation set, F-LGB demonstrated the highest correlation (mean R = 0.72, significantly surpassing the second-best SMAP-INRAE-BORDEAUX (SMAP-IB) SM and deep neural network (DNN) SM at 0.67) and the lowest ubRMSE (mean value of 0.052 m<sup>3</sup>/m<sup>3</sup>, better than 0.055 m<sup>3</sup>/m<sup>3</sup> for both DNN and SMAP-IB). F-LGB performed well across diverse land covers, vegetation densities, and climates, with SHAP analysis showing H-polarized brightness temperature as crucial, especially in areas with low to moderate vegetation. This new machine learning-based SMAP SM product may improve global satellite-based SM estimation capabilities.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Satellite-retrieved cloud top radiative cooling data in 2014 over global ocean

<p>Satellite-retrieved cloud-top radiative cooling data in 2014 over global ocean, used in a manuscript submitted to GRL (Zheng et al., 2021, GRL).</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Data for "Evaluating vertical velocity retrievals from vertical vorticity equation constrained dual-Doppler analysis of real, rapid-scan radar data"

<p>This archive contains data from the Rapid Scanning X-Band Polarimetric (RaXPol) radar, Atmospheric Imaging Radar (AIR), and Shared Mobile Atmospheric Research and Teaching radar (SMART-R) for 4 September 2018 in central Oklahoma. This is a rapid-scan dual-Doppler dataset of a convective storm. RaXPol and AIR were the two radars that can be used for dual-Doppler retrievals and the SMART-R data can be used for verification of vertical velocity.</p> <p>The AIR and RaXPol data have been quality controlled and are provided in cfRadial format. The SMART-R data has not been quality controlled and are available in its raw data format. All data can be read using the Python ARM Radar Toolkit.</p> <p>This dataset was used for the manuscript:</p> <p>Gebauer, J. G., A. Shapiro, C. K. Potvin, N. A. Dahl, M. I. Biggerstaff, and A. A. Alford, 2021: Evaluating vertical velocity retrievals from vertical vorticity equation constrained dual-Doppler analysis of rapid-scan radar data. <em>J. Atmos. Meas. Tech.,&nbsp;</em>in review.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Input data for RemoTeC synthetic measurements and retrieval

<p>RemoTeC is a retrieval algorithm developed for the retrieval of trace gas column-averaged dry air mole fractions from measured level 1b radiance spectra in the near-infrared (NIR) and shortwave-infrared (SWIR) bands. It is open access software developed by The Netherlands Institute for Space Research (SRON) and Karlsruhe Institute for Technology (KIT). &nbsp;</p> <p>The dataset available here is the input data for the RemoTeC synthetic measurement and retrieval code. The source code for these algorithms can be found at:</p> <p><a href="https://bitbucket.org/sron_earth/remotec_synthetic_measurements/src/master/">https://bitbucket.org/sron_earth/remotec_multi_purpose/src/main/</a></p> <p>Here we provide the data for both of these algorithms. The retrieval code should be used in conjunction with the synthetic measurement generator, but the synthetic measurement generator can be used as stand alone code. Any publications from using either of these codes should reference DOI for this dataset.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Data for the paper Evaluation of error components in rainfall retrieval from collocated commercial microwave links

<p>The published data for the paper includes rain-induced attenuation and rainfall intensities for commercial microwave links.</p> <p>The data are stored in text files. Timestamps are in UTC time in format yyyy-mm-dd HH:MM:SS. The data are at 1-min temporal resolution.</p> <p>Metadata can be found in the paper (Appendix A: Metadata table of CMLs): &Scaron;pačkov&aacute;, A., Fencl, M., and Bare&scaron;, V.: Evaluation of error components in rainfall retrieval from collocated commercial microwave links, Atmos. Meas. Tech. Discuss. [preprint], https://doi.org/10.5194/amt-2022-340, in review, 2023.</p> <p>The repository contains 2 folders (rain-induced attenuation and rainfall intensities) and Read_me file.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

CO2 retrievals at global power plants using PRISMA satellite data

<p>PRISMA (prisma.asi.it) data for a set of global power plants tasked between 2021-2022. Each scene includes a NETCDF file containing raw radiance data (1e-4 * W/(str &micro;m m<sup>2</sup>)), retrieved XCO2 (using an IMAP-DOAS algorithm), and retrieval precision. Also included is a spreadsheet &quot;PRISMA tracking-2023-06-23.xlsx&quot; that details the result of an analyst&#39;s QC of each scene regarding retrieval quality and CO2 plume detection. Another tracking sheet &quot;PRISMA emissions-2023-06-23.xlsx&quot; lists derived emission rates (via Integrated Mass Enhancement approach) with uncertainties and ERA5 wind speeds.</p> <p>Also included for each scene is an RGB and XCO2 PNG file that allows the user to quickly scan retrieval results.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
ClinicalTrials.gov32/100

Contacting Authors to Retrieve Individual Patient Data

ClinicalTrials.gov study NCT02569411. IPD Sharing: Not stated. Countries: 0. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Data from: Reservoir in-situ stress state determined by retrieved granite cores from the Gonghe enhanced geothermal system and its implications, northeastern Tibetan Plateau, China

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad32/100

Data from: A causal role for the precuneus in network-wide theta and gamma oscillatory activity during complex memory retrieval

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad32/100

Data from: Avian mitochondrial genomes retrieved from museum eggshell

Open the record for dataset details and reuse information.

publicFeb 2019View details →
zenodo28/100

Data Publication accompanying the paper "How FAIR can you get? Image Retrieval as a Use Case to calculate FAIR Metrics"

<pre>This dataset is the result of a benchmark run for a use-case-centric FAIR metric. The applied tech stack uses OAI-PMH and DataCite. The use case central to this benchmark is the retrieval of temporally and spatially annotated images. The zipped archives includes the data created during the first test run in June 2018. </pre>

opencc-by-4.0Oct 2018View details →
dryad28/100

Data from: Test collections for EHR-based clinical information retrieval

Objectives: To create test collections for evaluating clinical Information Retrieval (IR) systems and advancing clinical IR research. Materials and Methods: Electronic Health Records (EHR) data, including structured and free text data, from 45,000 patients who are a part of the Mayo Clinic Biobank cohort was retrieved from the clinical data warehouse. The clinical IR system indexed 42 million free-text EHR documents. The search queries consisted of 56 topics developed through a collaboration between Mayo Clinic and Oregon Health &amp; Science University. We described the creation of test collections, including a to-be-evaluated document pool using five retrieval models, and human assessment guidelines. We analyzed the relevance judgment results in terms of human agreement and time spent, and results of three levels of relevance, and reported performance of five retrieval models. Results: The two judges had a moderate overall agreement with a Kappa value of 0.49, spent a consistent amount of time judging the relevance, and were able to identify easy and difficult topics. The conventional retrieval model performed best overall on most topics while a concept-based retrieval model had better performance on the topics requiring conceptual level retrieval. Discussion: Information Retrieval can provide an alternate approach to leveraging clinical narratives for patient information discovery as it is less dependent on semantics. Our study showed the feasibility of test collections as well as challenges. Conclusion: The conventional test collections for evaluating the IR system show potential for successfully evaluating clinical IR systems with a few challenges to be investigated.

opencc-zeroJun 2019View details →
dryad28/100

Data from: Learning relevance models for patient cohort retrieval

OBJECTIVE We explored how judgements provided by physicians can be used to learn relevance models that enhance the quality of patient cohorts retrieved from Electronic Health Records (EHR) collections. METHODS A very large number of features were extracted from patient cohort descriptions as well as electronic health record collections. Specifically, we investigated retrieving (1) neurology-specific patient cohorts from the Temple University Hospital EEG Corpus as well as (2) the more general cohorts evaluated in the TREC Medical Records Track (TRECMed) from the de-identified hospital records provided by the University of Pittsburgh Medical Center. The features informed a Learning Relevance Model (LRM) that took advantage of relevance judgements provided by physicians. The LRM implements a pairwise learning-to-rank framework, which enables our learning patient cohort retrieval (L-PCR) system to learn from physicians' feedback. RESULTS AND DISCUSSION We evaluated the L-PCR system against state-of-the-art traditional patient cohort retrieval systems, and observed a 27% improvement when operating on EEGs and a 53% improvement when operating on TRECMed EHRs, showing the promise of the L-PCR system. We also performed extensive feature analyses to reveal the most effective strategies for representing cohort descriptions as queries, encoding EHRs, and measuring relevance. CONCLUSION The learning patient cohort retrieval system has significant promise for reliably retrieving patient cohorts from EHRs in multiple settings when trained with relevance judgments. When provided with additional cohort descriptions, the L-PCR will continue to learn, thus offering a potential solution to the performance barriers of current cohort retrieval systems.

opencc-zeroDec 2017View details →
zenodo28/100

Data for cloud retrieval algorithm

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record