Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “random subset”
Wikidata 3 Topical Subsets (Gene Wiki, Music, Ships) and 4 Random Subsets
<p>This dataset contains the N-Triples files of 3 Wikidata topical subsets corresponding to 3 Wikidata WikiProject: Gene Wiki, Music, and Ships along with 4 random subsets in different sizes: two of 100K items, one 500K items, and one 1M items. Subsets are extracted from the <a href="https://academictorrents.com/details/229cfeb2331ad43d4706efd435f6d78f40a3c438">3 January 2022 dump</a>. All subsets have been extracted with <a href="https://github.com/seyedahbr/wdumper">WDumper</a> using these <a href="https://github.com/seyedahbr/RQSS_Evaluation/tree/main/WDumper%20Specification%20Files">JSON specification files</a>. The files are:</p> <ul> <li>GeneWiki.zip: contains 25 `.nt.gz` RDF files each of which corresponds to one of the main Gene Wiki WikiProject classes, e.g. protein, gene, chemical compound, etc.</li> <li>music.nt.gz: the RDF file corresponding to the Music WikiProject.</li> <li>ships.nt.gz: the RDF file corresponding to the Ships WikiProject.</li> <li>Random100K_1.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random100K_2.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random500K.zip: contains 10 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 500,000 items in total.</li> <li>Random1M.zip: contains 20 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 1,000,000 items in total.</li> </ul> <p> </p>
Output files of the RQSS extractor and framework on 3 Topical (Gene Wiki, Music, Ships) subsets and 4 Random Subsets
<p>This dataset contains the `.csv` output files of the Referencing Quality Scoring System - RQSS performed on 3 topical subsets (Gene Wiki, Music, Ships) and 4 random subsets. Subset RDF files are in <a href="https://doi.org/10.5281/zenodo.7332161">this dataset</a>.</p>
Sounder SIPS: Aqua AIRS Level-1C Calibration Subset: Random Full Spectra V2 (SNDRAQIML1CCALSUBRND) at GES DISC
The Atmospheric Infrared Sounder (AIRS) is a grating spectrometer (R = 1200) aboard the second Earth Observing System (EOS) polar-orbiting platform, EOS Aqua. AIRS/Aqua Level-1C calibration subset including clear cases, special calibration sites, random nadir spots, and high clouds. Infrared temperature sounders generate a large amount of Level-1B spectral data. For example, the AIRS instrument with 2378 channels, its visible light component and AMSU with 15 channels create 3x240 files each day, for a total of over 500 MB of data. The purpose of the Calibration Data Subsets is extract key information from these data into a few daily files to: 1. Facilitate a quick evaluation of the absolute calibration of the instruments. 2. Facilitate an assessment of the instrument performance under clear, cloudy, and extreme hot and cold conditions. 3. Facilitate the evaluation of instrument trends and their significance relative to climate trends. 4. Facilitate the comparison of AIRS with CrIS using their equivalent data subsets.The output files are constructed from Level-1B or Level-1C IR and MW brightness or antenna temperatures. Each file contains selected observations taken from a nominal 24-hour period.
Sounder SIPS: Suomi NPP CrIS Level-1B NSR Calibration Subset: Random full spectra V2 (SNDRSNIL1BCALSUBRNDN) at GES DISC
The CrIS/ATMS instruments used for this product are on board the Suomi National Polar-orbiting Partnership (SNPP) platform and use the Normal Spectral Resolution (NSR) data. The CrIS instrument is a Fourier transform spectrometer with a total of 1305 NSR infrared sounding channels covering the longwave (655-1095 cm-1), midwave (1210-1750 cm-1), and shortwave (2155-2550 cm-1) spectral regions. The ATMS instrument is a cross-track scanner with 22 channels in spectral bands from 23 GHz through 183 GHz. Infrared temperature sounders generate a large amount of Level-1B spectral data. The purpose of the Calibration Data Subsets is extract key information from these data into a few daily files to: 1. Facilitate a quick evaluation of the absolute calibration of the instruments. 2. Facilitate an assessment of the instrument performance under clear, cloudy, and extreme hot and cold conditions. 3. Facilitate the evaluation of instrument trends and their significance relative to climate trends. 4. Facilitate the comparison of AIRS with CrIS using their equivalent data subsets.The output files are constructed from Level-1B and MW brightness or antenna temperatures. Each file contains selected observations taken from a nominal 24-hour period. The S-NPP CrIS summary subset contains Level-1B BTs for all selection types but only for selected channels, while the S-NPP CrIS random calibration subset contains full Level-1B spectra for only the randomly selected observations.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.