Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,655
datasets available to search
ShareScore release 0.9.0
Dataset results
1,655 results for “subset”
subset of the data
<p>The subset of the data that used in the article</p>
Data for: Nanopore R10.4.1 LSK114 HG002: subset of 20000 reads in BLOW5 format
<p>HG002 (NA24385) is a reference human genome sample used for benchmarking and comparing bioinformatics applications. This dataset contains a subset of 20,000 reads from the HG002 human reference sample, sequenced using an Oxford Nanopore Technologies PromethION sequencer on an R10.4.1 flowcell. Sheared DNA libraries (~17Kb) were prepared using the ONT LSK114 ligation library prep and an R10.4.1 flow cell was used to generate ~30X genome coverage. The original data in the FAST5 format was converted to BLOW5 format using slow5tools v0.8.0. This is a downsampled subset containing 20,000 reads in BLOW5 format.</p>
Hydrographs from validation subsets of PCR-GLOBWB-RF (30 arcmin)
<p>Plots showing hydrographs and flow duration curves (observed discharge, uncalibrated PCR-GLOBWB discharge and hybrid modelled discharge, 1979-2019) and respective residuals at all stations used for cross-validation (5 subsamples).</p> <p>A selected number of these plots were used in figure 7 in the paper, to show that the hybrid approach can improve (most of the time) or worsen the performance of uncalibrated PCR-GLOBWB. Here we store all the plots produced during cross-validation, to satisfy any potential curiosity of the reader. Plot numbers reflect stations from the Global Runoff Data Centre (GRDC). Note that the hybrid approach can produce a continuous synthetic streamflow time-series, irrespective of the availability of data at a particular station.</p> <p>A description of the hybrid methodology used to produce these time-series can be found in the paper 'Global streamflow modelling using process-informed machine learning', by Michele Magni, Edwin H. Sutanudjaja, Youchen Shen and Derek Karssenberg.</p>
A subset of the Scientific Literature Comparison Tables Dataset
<p>The Scientific Literature Comparison Table (SLCT) dataset was collected using the arXiv and Semantic Scholar APIs. It underwent a series of processing steps, including manual inspection and editing.</p> <p>The processing steps are summarized as follows:</p> <p>1) Downloading Survey Papers’ LaTeX files using the Arxiv API.</p> <p>2) Preprocessing LaTeX files to HTML format.</p> <p>3) Extracting tables from the HTML files.</p> <p>4) Creating a Golden Table as a reference.</p> <p>5) Generating descriptions for column headers.</p> <p>6) Acquiring citation data.</p> <p>7) Finalizing the dataset.</p>
FER-2013 subset ES-MAPs
<p>A modified subset of the FER-2013 dataset generated as ES-MAPs (emotional state heatmaps)</p>
UniProt subset about proteins and annotations generated using Shape Expressions
<p>Subset of Uniprot obtained from Shape Expression</p> <p>Link to Shape expression: https://github.com/shex-consolidator/subsetting-examples/blob/master/protein/protein.shex</p> <p>Dumps from Uniprot downloaded on 26-June-2023</p> <p>Tool employed in the creation of the subset: Pschea-rs (https://github.com/angelip2303/pschema-rs)</p>
RDF Subset From PDBj for Evaluating Oxigraph Server
<p>This repository contains the code and data files for the <a href="https://2023.biohackathon.org">DBCLS BioHackathon 2023</a> project "Evaluating Oxigraph Server as a Triple Store for Small and Medium-Sized Datasets".</p> <p>The evaluation focused on the use of [Oxigraph](https://oxigraph.org), a modern graph database implemented in the Rust programming language, as a triple store for handling small to medium-sized (sub-billion) datasets in the context of RDF and SPARQL technologies.</p>
Subset of MARIDA used at JuliaEO23
<p>Subset of MARIDA provided by E. Castanho</p> <p>Kikaki K, Kakogeorgiou I, Mikeli P, Raitsos DE, Karantzalos K (2022) MARIDA: A benchmark for Marine Debris detection from Sentinel-2 remote sensing data. PLoS ONE 17(1): e0262247. https://doi.org/10.1371/journal.pone.0262247</p>
Schema.org Characteristic Sets computed from the JSON-LD subset of Web Data Commons dataset (October 2021 release)
<p>This dataset reports the computation of Characteristic Sets from the JSON-LD subset of Web Data Commons dataset (October 2021 release). Each row consists in a combination of Schema.org properties and its cardinality. </p>
CCBB congenital CMV cohort immune cell subset RNA seq analysis
<p>CCBB congenital CMV cohort immune cell subset RNA seq analysis includes:</p> <p>1) metadata file with sample cell type and CMV infection status</p> <p>2) read counts for FAC-sorted cord blood immune cell subsets</p> <p>3) sample code for bulk RNA seq analysis and data visualization</p>
The Role of CD4+ T Cell Subsets in the Mechanism of Action of Vedolizumab in Ulcerative Colitis
ClinicalTrials.gov study NCT02721719. IPD Sharing: NO. Countries: 1. Publications: 31.
Safety and Immunogenicity Trial of an Oral SARS-CoV-2 Vaccine (VXA-CoV2-1) for Prevention of COVID-19 in Healthy Adults and Boost (VXA-CoV2-1.1-S) at 1 Year Post Initial Vaccination in Subset of Subje
ClinicalTrials.gov study NCT04563702. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Monocyte Subsets Altered by Anesthesia
ClinicalTrials.gov study NCT03431532. IPD Sharing: NO. Countries: 1. Publications: 5.
A Study to Evaluate the Accuracy of a Subset of the Length-109 Probe Set Panel (a Genetic Test) in Predicting Response to Golimumab in Participants With Moderately to Severely Active Ulcerative Coliti
ClinicalTrials.gov study NCT01988961. IPD Sharing: Not stated. Countries: 12. Publications: 2.
The Monocyte Subsets in Obese Patients With and Without Metabolic Syndrome
ClinicalTrials.gov study NCT03241394. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Identifocation the B Cell Subsets Responsible for Anti-pneumococcal Response
ClinicalTrials.gov study NCT02126384. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Postoperative REcurrence and DynamICs of T Cell Subsets in Crohn's Disease
ClinicalTrials.gov study NCT02770495. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Role of ARMA in Selective Subset of Refractory GERD Patients.
ClinicalTrials.gov study NCT05899491. IPD Sharing: NO. Countries: 1. Publications: 3.
Data from: A nonrandom subset of olfactory genes is associated with host preference in the fruit fly Drosophila orena
Open the record for dataset details and reuse information.
Data from: Generalist haemosporidian parasites are better adapted to a subset of host species in a multiple host community
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.