Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,655
datasets available to search
ShareScore release 0.9.0
Dataset results
1,655 results for “subset”
Subset de imagenes para experimentar con las rutinas de cálculo de disparidad
<p>Subset de imagenes para experimentar con las rutinas de cálculo de disparidad, extraido de https://vision.middlebury.edu/stereo/data/ para uso docente</p>
Seurat object subset of mouse liver scRNAseq data (Guilliams et al., Cell 2022)
<p>Seurat object containing a subset of the mouse liver scRNAseq data (Guilliams et al., Cell 2022)</p> <p>Data used only for demonstration purpose. Namely, to demonstrate the Differential NicheNet pipeline: https://github.com/saeyslab/nichenetr/blob/master/vignettes/differential_nichenet.md</p>
Regridded subset of MODIS chlorophyll, OC-CCI chlorophyll, MODIS sea surface temperature; basin bathymetric depth
<p>The North Atlantic phytoplankton bloom depends on a confluence of environmental factors that drive transient periods of exponential phytoplankton growth and interannual variability in bloom magnitude. I analyze interannual bloom variability in the North Atlantic via extreme value theory where the Generalized Extreme Value Distribution (GEVD) is fitted spatially to annual maxima of satellite-measured surface chlorophyll. I find excellent agreement between the observed distribution of interannual bloom maxima and those predicted from the GEVD. The spatial distribution of fitted GEVD parameters closely follows basin bathymetry where the largest extremes and heaviest distribution tails are found on the continental shelves and slopes. Trend analyses suggest weak evidence for changes in GEVD parameters, despite regional trends in mean chlorophyll levels and sea surface temperature. These results provide a framework to quantify interannual bloom variability and call for further work examining how extreme blooms propagate through food webs and contribute to carbon export.</p>
Results of Simulation Study II for "New weighting methods when cases are only a subset of events in a nested case-control study"
<p>This file includes the result of Simulation Study II in Section 4.3 of the manuscript.</p>
The 6-class subset of the MS1MV3 dataset for visualization
<pre>The 6-class subset of the [MS1MV3](https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_) dataset for visualization used in the paper: Hyperbolic Softmax Loss: Balancing Penalty on Intra-Class and Inter-Class Distances</pre> <p> </p>
Wikidata subset with revision history information [JSON]
<p>This dataset consists the complete revision history of every instance of the 100 most important classes in Wikidata. It contains 9.3 million classes and around 450 million revisions made to those classes. This dataset was exported from a MongoDB database. After decompressing the files, the resulting JSON files can be imported into MongoDB using the following commands:</p> <pre><code>mongoimport --db=db_name --collection=wd_entities --file=wd_entities.json mongoimport --db=db_name --collection=wd_revisions --file=wd_revisions.json </code></pre> <p>Make sure that <em>db_name</em> is replaced by the database where this data will be imported.</p> <p>Documents within the <em>wd_entities</em> collection have the following schema:</p> <ul> <li><strong>id</strong>: Internal id of the entity used by Wikidata (e.g. 8195238).</li> <li><strong>entity_id</strong>: Public id of the entity in Wikidata (e.g. 'Q42')</li> <li><strong>class_ids</strong>: List of classes that the entity belongs to (e.g. ['Q5', 'Q100'])</li> <li><strong>entity_json</strong>: JSON contents of the entity, following Wikidata's JSON data model (https://doc.wikimedia.org/Wikibase/master/php/md_docs_topics_json.html).</li> </ul> <p>Documents within the <em>wd_revisions</em> collection have the following schema:</p> <ul> <li><strong>id</strong>: Identifier of the revision (e.g. 15921539)</li> <li><strong>entity_id</strong>: Public id of the entity in Wikidata affected by this revision (e.g. 'Q42')</li> <li><strong>class_ids</strong>: List of classes that the entity affected by this revision belongs to (e.g. ['Q5', 'Q100'])</li> <li><strong>parent_id</strong>: Identifier of the previous revision to this one, if it exists (e.g. 15921214)</li> <li><strong>timestamp</strong>: Date where the revision was made, following the ISO 8601 format (e.g. +2019-05-27T09:31:10Z)</li> <li><strong>username</strong>: Username of the user that made the revision.</li> <li><strong>comment</strong>: Comments made by the user in the revision, if any.</li> <li><strong>entity_diff</strong>: List of operations made in this revision, following the JSON Patch format.</li> </ul>
Data subset from Sorbie, Jimenez and Benakis, iScience 2022
<p>This repository contains a subset of the raw fastq files from Sorbie, Jimenez and Benakis, iScience, 2022, as well as the processed datasets. This data is provided to test our pipeline and provide a citable link to the processed dataset. </p>
blase II: PHOENIX Subset Clone Archive
<p>As part of our study, we cloned 1,314 individual PHOENIX spectra (Husser et al. 2013) with blase and optimized their line shapes with ML. These interpretable clones have been saved as state dictionaries in .pt files, which we upload here to be readily accessible to others.</p> <p> </p> <p>You can download and unzip the file in order to access the clone state dictionaries. They can be loaded from disk using torch.load() (Paszke et al. 2019), and input into blase's SparseLinearEmulator (Gully-Santiago & Morley 2022) with its init_state_dict constructor argument.</p>
FERMO 1.0.0 example dataset: metabolomics data from Planomonospora - data subset
<p>A subset of the <em>Planomonospora</em> dataset published under <a href="https://doi.org/10.1021/acs.jnatprod.0c00807" target="_blank" rel="noopener">Zdouc et al 2021</a> and available under <a href="https://massive.ucsd.edu/ProteoSAFe/dataset.jsp?accession=MSV000085376">MSV000085376</a>. This subset encompasses strains from different phylogroups, all grown in the same medium. A medium blank is provided as well. This data is referenced in the <em>FERMO</em> 1.0.0 Documentation and made available for convenient access.</p>
Data for "The 20-year highest tropical cyclone-generated waves associated with the maximum energy of seismic noises" (Subset 4)
<p>This is dataset of ocean wave simulations used in the paper "The 20-year highest tropical cyclone-generated waves associated with the maximum energy of seismic noises" by Shimura et al. (submitted).</p> <p>The dataset contains </p> <ul> <li>significant wave height (Subset 1),</li> <li>wave induced surface pressure (Subset 2)</li> <li>long-period componet of wave heights (Subset 3)</li> <li>long-period component of surface pressure (Subset 4)</li> </ul> <p>during 2004 from 2023 summer. </p>
Data for "The 20-year highest tropical cyclone-generated waves associated with the maximum energy of seismic noises" (Subset 1)
<p>This is dataset of ocean wave simulations used in the paper "The 20-year highest tropical cyclone-generated waves associated with the maximum energy of seismic noises" by Shimura et al. (submitted).</p> <p>The dataset contains </p> <ul> <li>significant wave height (Subset 1),</li> <li>wave induced surface pressure (Subset 2)</li> <li>long-period componet of wave heights (Subset 3)</li> <li>long-period component of surface pressure (Subset 4)</li> </ul> <p>during 2004 from 2023 summer. </p>
Data for "The 20-year highest tropical cyclone-generated waves associated with the maximum energy of seismic noises" (Subset 2)
<p>This is dataset of ocean wave simulations used in the paper "The 20-year highest tropical cyclone-generated waves associated with the maximum energy of seismic noises" by Shimura et al. (submitted).</p> <p>The dataset contains </p> <ul> <li>significant wave height (Subset 1),</li> <li>wave induced surface pressure (Subset 2)</li> <li>long-period componet of wave heights (Subset 3)</li> <li>long-period component of surface pressure (Subset 4)</li> </ul> <p>during 2004 from 2023 summer. </p>
Fictive reactions validated by TTL with USPTO as source of molecules - equilibrated subset (1M)
Open the record for dataset details and reuse information.
MAST rhythm re-annotated subset
<p>In the work described in the paper <em><a href="http://archives.ismir.net/ismir2019/paper/000052.pdf">A dataset of rhythmic pattern reproductions and baseline automatic assessment system</a></em><a href="https://ismir2019.ewi.tudelft.nl/?q=accepted-papers"><em> (ISMIR, 2019)</em></a>, Falcão et al. applied efforts on annotating a subset of the full <a href="https://zenodo.org/record/2620357#.XKLB39szZuQ">MAST rhythm dataset</a> with grades for a subset of performances.</p> <p>Out of a total of 1040 student submissions comprised by the original dataset, 80 performances (equally distributed between 20 references) were graded by annotators using a custom annotation tool. During this process annotators could play each <em>(reference, performance)</em> pair as many times as desired before deciding on one of the available assessments: <em>4-Perfect, 3-Minor errors, 2-Major errors, 1-Completely off.</em></p> <p><strong>Naming convention</strong></p> <p>This dataset is a collection of audio files, onset features and user-annotated data. Audio files are stored into the <em>Only Performances</em> and <em>Only References </em>folders, together with some extra files containing their original onsets, quantized onsets and scaling parameters (for more information regarding this auxiliary data, please check the paper). An indexed list of files (<em>listreferences </em>and<em> listperformances</em>) is also stored in each of these folders in order to allow for mapping between audio files and auxiliary data. Example:</p> <ul> <li>The 10th line in <em>Only Performances/MAST Onsets [Performances] </em>contains the pure onsets for the 10th audio file listed in the <em>Only Performances/listperformances </em>file ('55_rhy2_per138459_pass.wav')</li> </ul> <p>Also, in order to check the reference which relates to a specific performance, one must also look for audio file names in the list of files: the <em>i-eth</em> audio file listed in <em>Only Performances/listperformances </em>is a specific performance whose reference is defined by the<em> i-eth</em> line of <em>Only References/listreferences. </em>Example:</p> <ul> <li>The 10th audio file listed in <em>Only Performances/listperformances </em>file ('55_rhy2_per138459_pass.wav') is a performance whose reference is stored in the 10th line of <em>Only References/listreferences </em>(55_rhy2_ref162559.wav)</li> </ul> <p>Such convention is also applied to the folder that contains the annotations (<em>Performances Annotations</em>). Each one of the <em>annotator_<strong>X</strong>.txt </em>file lists the grades assigned by the <em><strong>X</strong>-ith</em> annotator, and the mapping between the assessments and their referred performances is indexed in <em>listfiles.txt</em></p>
Experimental data from the paper "Subset-Saturated Cost Partitioning for Optimal Classical Planning"
<p>The data set contains the raw experiment data, parsed values and basic reports for the experiments in the paper. For each experiment there are two directories. The first directory contains the raw data of all experiment runs. The code directories and benchmark files have been removed to avoid duplication and save space. The second directory (*-eval) contains a "properties" file with all parsed values and an HTML report.</p>
Subset of BioID dataset (https://www.bioid.com/facedb/) used for image denoising benchmark as used in DivNoising paper (https://arxiv.org/abs/2006.06072)
<p>The original BioID dataset comes from https://www.bioid.com/facedb/. </p> <p>A subset of original BioID dataset was used for image denoising benchmark (corrupted with zero mean Gaussian noise of std 15) as in DivNoising paper (https://arxiv.org/abs/2006.06072)</p>
Subset of AmazonQA annotated with answerability and semantic & syntactic embedding
<p>3755 Instances taken from the review-question dataset AmazonQA. The file data_answerability contain the 3755 instances annotated with answerability, answer tag (aligned with the passage) and answer text. The file data_annotated contains 1818 answerable instances, annotated with embedding constructions in the answer text. Inventory_phrases and inventory_simple_words contain the expressions of logical operator, implicative and factive predicates that can be used in embedding annotation. annotate.py is the script for embedding annotation. </p>
Refseq Test Subsets for Frame Classification with and without Errors
<p>These test files extend the<a href="https://zenodo.org/record/4306248"> 'Refseq datasets for training frame classification</a>' dataset. It provided the original test file and three variations containing erroneous sequences to simulate realistic data.</p> <p>The data is based on randomly selected viral and bacterial genomes and the human193(GRCh38.p13) reference genome which was downloaded from GenBank. From each original nucleic acid sequence, we created multiple patches of length 300 in all possible reading frames using a sliding window on the initial sequence and its reversed complement. The data is stored in the FASTA format according to the following convention:<br> <br> >{ID}_subsequence{patch index}_frame{frame index}|{class marker}|{frame index}<br> sequence<br> <br> with<br> <br> ID - denotes the ReSeq accession of the original sequence in the Refseq dataset.<br> sequence - nucleic acid sequence patch of length 300 or 250<br> patch_index - denotes the starting triplet of the given patch within the original sequence or reverse complemented sequence (i.e. 3*patch_index is the starting index of frame 0 in the original sequence)<br> class marker - indicates the taxonomic domain<br> 0 - virus<br> 1 - bacteria<br> 2 - human / mammal<br> frame index - indicates the reading frame<br> 0 - on-frame<br> 1 - shifted by one<br> 2 - shifted by two<br> 3 - reverse complemented<br> 4 - shifted by one and reverse complemented<br> 5 - shifted by two and reverse complemented</p> <p>Each file contains 212.618 patches per frame.</p>
FIGURE. Median network analyses (MNA) of a subset of the C. trilobus aggregate (i.e. those in the clade A from Fig. 11) based on concatenated DNA sequence data from ITS, trnL-trnF and psbJ-petA. Stars and arrow indicate accessions discussed in the text. NI: North Island, SI: South Island. in Five new species of Corybas (Diurideae, Orchidaceae) endemic to New Zealand and phylogeny of the Nematoceras clade
FIGURE. Median network analyses (MNA) of a subset of the C. trilobus aggregate (i.e. those in the clade A from Fig. 11) based on concatenated DNA sequence data from ITS, trnL-trnF and psbJ-petA. Stars and arrow indicate accessions discussed in the text. NI: North Island, SI: South Island.
10-arcmin overlays for the MeerKAT-2019 subset of the G4Jy Sample
<p>Overlays created for 140 sources in the G4Jy Sample, referred to as the MeerKAT-2019 subset. The radio contours are from MeerKAT (purple), GLEAM (red), NVSS/SUMSS (blue), and TGSS (yellow), with the background image being from WISE and having a 10-arcmin across field-of-view. The number of sigma in the filename refers to the lowest contour-level of the MeerKAT contours, and the component MeerKAT images can be accessed via the SARAO archive: https://archive-gw-1.kat.ac.za/public/repository/10.48479/wyab-t838/index.html . See Sejake et al. (2022) for further details.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.