Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,655
datasets available to search
ShareScore release 0.9.0
Dataset results
1,655 results for “Subset”
Figure 6. The occurrence with data from subset 1 in Modelling hot spot areas for the invasive alien plant Elodea nuttallii in the EU
Figure 6. The occurrence with data from subset 1 (A) and subset 3 (B). The combined output of both is also presented (C). Areas with low uncertainty and above the habitat suitability threshold (i.e. priority areas) are marked in red while areas with high uncertainty are marked in yellow. The grid size of all 3 maps is 1 km2.
Subset of Opensky Dataset for aircraft trajectories
<p>To reduce the size complexity we have filtered the OpenSky dataset such that it encompasses flight trajectory information from the Nordrhein Westfalen region, with latitude ranging from 50 to 52 degrees north and longitude ranging from 5 to 9 degrees east, capturing the movement of aircraft's from July 1st 2019 to July 31st 2019.</p>
MADFORWATER: WP1: Water and water-related vulnerabilities in Egypt, Morocco and Tunisia: Task1.2: Analysis and mapping of water stress, water vulnerability and potential for water reuse in Egypt, Morocco and Tunisia: Subtask1.2.b: Data collection on water stress and vulnerability: Souss-Massa Region Subset
<p>This folder contains the dataset that I used to write my conference paper "Groundwater Resources Scarcity in Souss-Massa Region and Alternative Solutions for Sustainable Agricultural Development"</p>
Tracking-NEMO-movies_subset
<p>This is a subset of the image data used to perform MSD analysis on NEMO dots, for the publiction: </p> <ul> <li><a href="http://jcb.rupress.org/content/204/2/231.long">"TNF and IL-1 exhibit distinct ubiqui-tin requirements for inducing NEMO-IKK supramolecular structures", Tarantino et al. (2014)</a>.</li> </ul> <p>This dataset is used in educational resources to introduce single-particle tracking and mean-square displacement analysis.</p> <p> </p>
RSYD-BASIC results for AMR benchmarking dataset subset (MiSeq data)
<p><strong>Input data:</strong></p> <ul> <li>20240905_test_config.yaml: original config file used to run the pipeline</li> <li>20241004_rsyd_largeset_reads.zip: renamed Illumina MiSeq reads</li> <li>20241011-sample-overview.xlsx: overview of SRR accession numbers to internal sample numbers</li> <li>ILM_Run0001_Y20240904_kts_new.xlsx: runsheet </li> <li>input_en.yaml: column name configuration for the run</li> <li>lis_data.zip: LIS report and bacteria list used for LIS-specific results</li> </ul> <p><strong>Expected results:</strong></p> <ul> <li>20240910_test_illumina_largeset.zip: Results of the RSYD-BASIC pipeline, version 1.15.1, with the reads used</li> </ul>
CLIP Features and Selected Relevance Judgments Subset for TRECVID Ad-hoc Search (2019-2023)
<div> <div> <div> <div> <div> </div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>This repository contains CLIP features and annotations for a subset of V3C images, based on their relevance to selected queries from the TREC Video Retrieval Evaluation (TRECVID) Ad-hoc Video Search (AVS) task. The data includes annotations for AVS queries and judgments conducted in TRECVID from 2019 to 2023 [1], using the V3C1 and V3C2 collections [2]. Specifically, the TRECVID-AVS collection covers 89 queries, with video shots manually labeled as relevant (1), non-relevant (0), or not annotated (-1).</p> <p>We used approximately 2.6 million keyframes extracted from these video shots, mapping the annotations to the corresponding keyframes (note that there may not be a one-to-one correspondence between TRECVID shotID since multiple frames might be extracted from a single shot). Image representations are based on CLIP ViT-H/14 - LAION-2B features [3]. The timestamps of the keyframes and their CLIP features are sourced from the VISIONE repository [4].</p> <p>Given the incomplete nature of the TRECVID ground truth (where only a subset of video segments were judged per query), we focused on queries with at least 200 positive and 1400 negative annotations. This resulted in 80 datasets, each containing 1500 images—10% labeled as relevant and 90% as non-relevant.</p> <h3>Contents of the Repository:</h3> <ol> <li> <p><strong>Query-specific CSV Files:</strong> For each of the 80 selected AVS query (e.g., <code>1591</code>), the corresponding CSV file (e.g., <code>1591.csv</code>) contains a column for each image, where:</p> <ul> <li><strong>VISIONE image ID</strong> is in the first row.</li> <li><strong>CLIP features</strong> are in the subsequent rows.</li> <li><strong>Relevance annotations</strong> are in the last row: <code>1</code> for relevant, <code>0</code> for non-relevant.</li> </ul> </li> <li> <p><strong>Post-processed Datasets:</strong></p> <ul> <li><code>dataset_normalized.zip</code>: L2-normalized CLIP features.</li> <li><code>dataset_softmax.zip</code>: CLIP features converted into probabilities using a softmax function.</li> <li><code>dataset_logistic.zip</code>: CLIP features converted into probabilities using a logistic function followed by L1 normalization.</li> </ul> </li> <li> <p><strong>Text Feature Data:</strong> <code>clip_laion_text_features.csv</code> contains additional details for each query, including the query ID, query text, and L2 normalized CLIP features extracted from the query text.</p> </li> </ol> <h3>Citation and Usage:</h3> <p>This data was used in the experiments described in:</p> <p>Lucia Vadicamo, Francesca Scotti, Alan Dearle, Richard Connor, <em>Comparative Analysis of Relevance Feedback Techniques for Image Retrieval</em>, in Proceedings of the 31st International Conference on Multimedia Modeling (MMM 2025).</p> <p>The data is released under a Creative Commons Attribution license. If you use it in your research, please cite the above work. </p> <h3>References:</h3> <p>[1]TRECVID Data: <a href="https://www-nlpir.nist.gov/projects/trecvid/trecvid.data.html">https://www-nlpir.nist.gov/projects/trecvid/trecvid.data.html</a><br>[2] Rossetto, L., Schuldt, H., Awad, G., Butt, A.A.: V3C - A research video collection. <em>In: International Conference on Multimedia Modeling</em>, pp. 349–360. Springer (2019).<br>[3] https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K<br>[4] VISIONE Repository: <a href="https://zenodo.org/records/8188570">https://zenodo.org/records/8188570</a></p> </div> </div> </div> </div> </div> </div> <p> </p> <p> </p> <p> </p>
Linked collectors and determiners for: ENBI WP 13 Martius C. F. P. von Collection subset (Muenchen).
Natural history specimen data linked to collectors and determiners held within, "ENBI WP 13 Martius C. F. P. von Collection subset (Muenchen)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/f6c83992-27ff-11e2-85e3-00145eb45e9a">https://bionomia.net/dataset/f6c83992-27ff-11e2-85e3-00145eb45e9a</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/f6c83992-27ff-11e2-85e3-00145eb45e9a">https://gbif.org/dataset/f6c83992-27ff-11e2-85e3-00145eb45e9a</a>. Formatted as a Frictionless Data package.
Example variant files and corresponding annotations for GnomAD v3.1.1 on a subset of chromosome 22
<p>Example variant files and corresponding hg38 annotations for `chr22:15518158-20127355`.</p> <p>Sources:</p> <ul> <li><a href="http://dx.doi.org/10.1093/nar/gky955">Gencode v34 (hg38)</a></li> <li><a href="https://doi.org/10.1038/s41586-020-2308-7">GnomAD v3.1.1</a></li> </ul>
RHARM (Radiosounding HARMonization) dataset - subset
<p>In the context of the Copernicus Climate Change Service (C3S), a novel approach, named RHARM (Radiosounding HARMonization), has been developed to provide a harmonized dataset of temperature, humidity and wind profiles along with an estimation of the measurement uncertainties for about 650 radiosounding stations globally. </p> <p>The RHARM method is applied to the Integrated Global Radiosonde Archive (IGRA) Version 2 which is the most comprehensive collection of historical and near-real-time radiosonde and pilot balloon observations from around the globe, maintained and distributed by the National Oceanic and Atmospheric Administration’s National Centers for Environmental Information (NCEI). The daily (0000 and 1200 UTC) radiosonde data holdings on 16 standard pressure levels (from 1000 to 10 hPa) from 1978 to present are post-processed or statistically homogenized using the RHARM approach. The applied adjustments are interpolated to all reported significant levels to retain information content contained within each individual ascent profile. </p> <p>The RHARM algorithm is the first to provide homogenized time series of temperature, relative humidity and wind profiles alongside an estimation of the observational uncertainty for each single observation at each pressure level. </p> <p>A copy of the RHARM dataset is stored in the Copernicus Climate Data Store (CDS) although not publicly available yet. For review purposes only, a subset has been made available here.</p>
Main-sequence pulsar: the peculiar subset of hot magnetic stars
<p>In my talk, I will describe some peculiar phenomena observed from a subset of hot magnetic stars, also known as 'main-sequnce pulsar's and discuss their usefulness. This class of stars is distinguished from the rest of the population by their ability to produce highly directed coherent radio emission via electron cyclotron maser emission (ECME) in their magnetospheres. The fundamental difference between the two classes is however still vague. I will describe our efforts to understand this difference; this includes a survey being carried out with the Giant Metrewave Radio telescope (GMRT) in order to find more hot magnetic stars that show coherent radio emission, and also wideband observations taken with two telescopes: the GMRT and the Very Large Array (VLA), of the already known coherent radio emitters to find out the bandwidth of the emission process. In this process, we have also come across certain new phenomena like the inversion of circular polarization and that of the sequence of ECME pulse arrival with frequency. I will discuss the implications of these new revelations and their potential to become tools for estimating plasma density close to the stellar surface, as well as constraining the inhomogeneity and asymmetry in the magnetosphere.</p>
Wikidata 3 Topical Subsets (Gene Wiki, Music, Ships) and 4 Random Subsets
<p>This dataset contains the N-Triples files of 3 Wikidata topical subsets corresponding to 3 Wikidata WikiProject: Gene Wiki, Music, and Ships along with 4 random subsets in different sizes: two of 100K items, one 500K items, and one 1M items. Subsets are extracted from the <a href="https://academictorrents.com/details/229cfeb2331ad43d4706efd435f6d78f40a3c438">3 January 2022 dump</a>. All subsets have been extracted with <a href="https://github.com/seyedahbr/wdumper">WDumper</a> using these <a href="https://github.com/seyedahbr/RQSS_Evaluation/tree/main/WDumper%20Specification%20Files">JSON specification files</a>. The files are:</p> <ul> <li>GeneWiki.zip: contains 25 `.nt.gz` RDF files each of which corresponds to one of the main Gene Wiki WikiProject classes, e.g. protein, gene, chemical compound, etc.</li> <li>music.nt.gz: the RDF file corresponding to the Music WikiProject.</li> <li>ships.nt.gz: the RDF file corresponding to the Ships WikiProject.</li> <li>Random100K_1.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random100K_2.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random500K.zip: contains 10 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 500,000 items in total.</li> <li>Random1M.zip: contains 20 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 1,000,000 items in total.</li> </ul> <p> </p>
Output files of the RQSS extractor and framework on 3 Topical (Gene Wiki, Music, Ships) subsets and 4 Random Subsets
<p>This dataset contains the `.csv` output files of the Referencing Quality Scoring System - RQSS performed on 3 topical subsets (Gene Wiki, Music, Ships) and 4 random subsets. Subset RDF files are in <a href="https://doi.org/10.5281/zenodo.7332161">this dataset</a>.</p>
Spectral-index maps for the MeerKAT-2019 subset of the G4Jy Sample
<p>Spectral-index maps for sources in the MeerKAT-2019 subset of the G4Jy Sample. The component MeerKAT spectral-index maps can be accessed as both .png and .fits files via the SARAO archive: https://archive-gw-1.kat.ac.za/public/repository/10.48479/wyab-t838/index.html . See Sejake et al. (2022) for further details.</p>
Wikidata subset from 2018 dumps created with WDSub at Biohackathon 2022 - Turtle format
<p>Subset of Wikidata obtained using this Shape Expression: https://github.com/kg-subsetting/datasets-biohackathon2022/blob/main/GeneWiki/GeneWiki.shex</p> <p>And the wdsub tool version 0.0.28: https://github.com/weso/wdsub</p> <p>The input dump is: wikidata-20180115-all</p> <p>And the dumpformat is TURTLE</p>
GeneWiki subset created with WDSub at Biohackathon 2022 - Turtle format - Only labels en english
<p>Subset of Wikidata obtained using this Shape Expression: https://github.com/kg-subsetting/datasets-biohackathon2022/blob/main/GeneWiki/GeneWiki.shex</p> <p>And the wdsub tool version 0.0.31: https://github.com/weso/wdsub which is available at docker</p> <p>The input dump is: wikidata-20220630-all.json.gz</p> <p>And the dumpformat is Turtle/RDF</p>
scPDB BO1 subset (protein ligand-binding sites)
<p>The BO1 subset of the scPDB database.</p> <p>BO1 consists in 766 pairs of non-redundant binding-sites<br> (383 similar pairs, 383 dissimilar pairs).</p> <p>BO1 was describbed in:<br> ---<br> Eguida, M., & Rognan, D. (2020).<br> A computer vision approach to align and compare protein cavities:<br> application to fragment-based drug design.<br> Journal of Medicinal Chemistry, 63(13), 7127-7142.<br> https://doi.org/10.1021/acs.jmedchem.0c00422<br> ---</p> <p>The scPDB was recently describbed in:<br> ---<br> Desaphy, J., Bret, G., Rognan, D., & Kellenberger, E. (2015).<br> sc-PDB: a 3D-database of ligandable binding sites—10 years on.<br> Nucleic acids research, 43(D1), D399-D404.<br> https://doi.org/10.1093/nar/gku928<br> ---</p> <p>The scPDB is available at:<br> http://bioinfo-pharma.u-strasbg.fr/scPDB/<br> </p>
Simple dataset obtained as a Wikidata subset from 2022 dump using entity schema about taxon
<p>The subset has been obtained using wdsub version 0.0.33 and the schema:</p> <p> </p> <pre>PREFIX p: <http://www.wikidata.org/prop/> PREFIX ps: <http://www.wikidata.org/prop/statement/> PREFIX prov: <http://www.w3.org/ns/prov#> PREFIX wd: <http://www.wikidata.org/entity/> PREFIX wdt: <http://www.wikidata.org/prop/direct/> start = @<taxon_by_wd_ontology> OR @<taxon_by_identifier> <taxon_by_wd_ontology> { wdt:P31 [wd:Q16521] ; } <taxon_by_identifier> {wdt:P685 . +;} OR # NCBI taxonomy ID {wdt:P846 . +;} OR # GBIF taxon ID {wdt:P3151 . +;} OR # iNaturalist taxon ID {wdt:P3444 . +;} # eBird taxon ID</pre>
GTEx subset
<p>This is a subset of GTEx created for teaching purposes. It contains 3382 samples with expression counts from 1000 genes for 7 different tissues. </p>
PatchCamelyon (PCAM) 5% subset
<p>This subset was created for teaching purposes. It is derived from the original data on drive (https://drive.google.com/drive/folders/1gHou49cA1s5vua2V5L98Lt8TiWA3FrKB) and contains every 20th image in each of the splits. References for this dataset are:</p> <p>- B. S. Veeling, J. Linmans, J. Winkens, T. Cohen, M. Welling. "Rotation Equivariant CNNs for Digital Pathology". <a href="http://arxiv.org/abs/1806.03962">arXiv:1806.03962</a>.</p> <p>- Ehteshami Bejnordi et al. Diagnostic Assessment of Deep Learning Algorithms for Detection of Lymph Node Metastases in Women With Breast Cancer. JAMA: The Journal of the American Medical Association, 318(22), 2199–2210. <a href="https://doi.org/10.1001/jama.2017.14585">doi:jama.2017.14585</a>.</p>
Datalog subsetting input files (Wikidata 2015 NTriple-to-CSV dump)
<p>This is a Wikidata 2015 NTriple dump in which the delimiter is changed to ','. The file is used in subsetting experiment via <a href="https://github.com/seyedahbr/radlog">Radlog</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.