Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “random subset”

Learn how ShareScore rates datasets ↗
zenodo40/100

Wikidata 3 Topical Subsets (Gene Wiki, Music, Ships) and 4 Random Subsets

<p>This dataset contains the N-Triples files of 3 Wikidata topical subsets corresponding to 3 Wikidata WikiProject: Gene Wiki, Music, and Ships along with 4 random subsets in different sizes: two of 100K items, one 500K items, and one 1M items. Subsets are extracted from the <a href="https://academictorrents.com/details/229cfeb2331ad43d4706efd435f6d78f40a3c438">3 January 2022 dump</a>. All subsets have been extracted with <a href="https://github.com/seyedahbr/wdumper">WDumper</a> using these <a href="https://github.com/seyedahbr/RQSS_Evaluation/tree/main/WDumper%20Specification%20Files">JSON specification files</a>. The files are:</p> <ul> <li>GeneWiki.zip: contains 25 `.nt.gz` RDF files each of which corresponds to one of the main Gene Wiki WikiProject classes, e.g. protein, gene, chemical compound, etc.</li> <li>music.nt.gz: the RDF file corresponding to the Music WikiProject.</li> <li>ships.nt.gz: the RDF file corresponding to the Ships WikiProject.</li> <li>Random100K_1.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random100K_2.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random500K.zip: contains 10 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 500,000 items in total.</li> <li>Random1M.zip: contains 20 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 1,000,000 items in total.</li> </ul> <p>&nbsp;</p>

opencc-byNov 2022View details →
zenodo40/100

Output files of the RQSS extractor and framework on 3 Topical (Gene Wiki, Music, Ships) subsets and 4 Random Subsets

<p>This dataset contains the `.csv` output files of the Referencing Quality Scoring System - RQSS performed on 3 topical subsets (Gene Wiki, Music, Ships) and 4 random subsets. Subset RDF files are in <a href="https://doi.org/10.5281/zenodo.7332161">this dataset</a>.</p>

opencc-byNov 2022View details →
nasa28/100

Sounder SIPS: Aqua AIRS Level-1C Calibration Subset: Random Full Spectra V2 (SNDRAQIML1CCALSUBRND) at GES DISC

The Atmospheric Infrared Sounder (AIRS) is a grating spectrometer (R = 1200) aboard the second Earth Observing System (EOS) polar-orbiting platform, EOS Aqua. AIRS/Aqua Level-1C calibration subset including clear cases, special calibration sites, random nadir spots, and high clouds. Infrared temperature sounders generate a large amount of Level-1B spectral data. For example, the AIRS instrument with 2378 channels, its visible light component and AMSU with 15 channels create 3x240 files each day, for a total of over 500 MB of data. The purpose of the Calibration Data Subsets is extract key information from these data into a few daily files to: 1. Facilitate a quick evaluation of the absolute calibration of the instruments. 2. Facilitate an assessment of the instrument performance under clear, cloudy, and extreme hot and cold conditions. 3. Facilitate the evaluation of instrument trends and their significance relative to climate trends. 4. Facilitate the comparison of AIRS with CrIS using their equivalent data subsets.The output files are constructed from Level-1B or Level-1C IR and MW brightness or antenna temperatures. Each file contains selected observations taken from a nominal 24-hour period.

restrictednotspecifiedApr 2025View details →
nasa28/100

Sounder SIPS: Suomi NPP CrIS Level-1B NSR Calibration Subset: Random full spectra V2 (SNDRSNIL1BCALSUBRNDN) at GES DISC

The CrIS/ATMS instruments used for this product are on board the Suomi National Polar-orbiting Partnership (SNPP) platform and use the Normal Spectral Resolution (NSR) data. The CrIS instrument is a Fourier transform spectrometer with a total of 1305 NSR infrared sounding channels covering the longwave (655-1095 cm-1), midwave (1210-1750 cm-1), and shortwave (2155-2550 cm-1) spectral regions. The ATMS instrument is a cross-track scanner with 22 channels in spectral bands from 23 GHz through 183 GHz. Infrared temperature sounders generate a large amount of Level-1B spectral data. The purpose of the Calibration Data Subsets is extract key information from these data into a few daily files to: 1. Facilitate a quick evaluation of the absolute calibration of the instruments. 2. Facilitate an assessment of the instrument performance under clear, cloudy, and extreme hot and cold conditions. 3. Facilitate the evaluation of instrument trends and their significance relative to climate trends. 4. Facilitate the comparison of AIRS with CrIS using their equivalent data subsets.The output files are constructed from Level-1B and MW brightness or antenna temperatures. Each file contains selected observations taken from a nominal 24-hour period. The S-NPP CrIS summary subset contains Level-1B BTs for all selection types but only for selected channels, while the S-NPP CrIS random calibration subset contains full Level-1B spectra for only the randomly selected observations.

restrictednotspecifiedApr 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record