Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,655
datasets available to search
ShareScore release 0.9.0
Dataset results
1,655 results for “Subset”
Subset of Global Sub-Daily Rainfall (GSDR) dataset
<p>A subset of the Global Sub-Daily Rainfall (GSDR) dataset, including the countries marked as 'open' in Table A2 of the accompanying paper (DIO: 10.1175/JCLI-D-18-0143.1).</p> <p>The dataset includes archive data, raw data, quality control flags and quality controlled data and raw data. Example quality control files (with headers), raw data processing scripts and INTENSE Python code are also included.</p>
Study of Purified Vero Rabies Vaccine Compared With Two Reference Rabies Vaccines, Given in a Pre-exposure Regimen to Children and Adults and as Single Booster Dose to a Subset of Adults
ClinicalTrials.gov study NCT04127786. IPD Sharing: YES. Countries: 1. Publications: 1.
Reproducible data, example subsets, and analysis pipeline for the extended TAaCGH study of breast cancer genomic and transcriptomic profiles
Open the record for dataset details and reuse information.
MERRA-2 subset for evaluation of renewables with merra2ools R-package: 1980-2020 hourly, 0.5° lat x 0.625° lon global grid
Open the record for dataset details and reuse information.
ASTRAL-SCOPe subset 2.04 in ActivePapers format
<p>This ActivePaper contains the structures in version 2.04 of the ASTRAL SCOPe subset with less than 40% sequence identity. For more information about ASTRAL and SCOPe, see</p> <p> http://scop.berkeley.edu/astral/</p> <p>Each ASTRAL entry describes a domain from a protein structure in the PDB. This ActivePaper contains these domains in the MOSAIC HDF5 format. For more information about MOSAIC, see</p> <p> http://mosaic-data-model.github.io/</p> <p>The structures are arranged by its SCOPe classification. For example, ASTRAL entry d1v0aa1 is found under /data/b/18/1/30/d1v0aa1, because SCOPe classifies it as</p> <p> b: all-beta proteins<br /> b.18: Galactose-binding domain-like<br /> b.18.1: Galactose-binding domain-like<br /> b.18.1.30: CBM11</p> <p>For each entry, the ASTRAL database provides a reference to the PDB entry with chain and residue identifiers plus a sequence. The importlet in /code/import_structures reads this information, downloads the corresponding PDB entry in mmCIF format, extracts the domain, checks that its sequence matches the one given by ASTRAL, and stores the domain in MOSAIC format.</p> <p>For a small number of ASTRAL entries, this process failed for various reasons: mismatch between the ASTRAL reference and the PDB data, mismatch in the sequences, a mistake in the PDB mmCIF file, or unjustified hypotheses in the conversion script. The number of failures was deemed sufficiently small (28 failures out of 13042 entries) for not attempting a time-consuming in-detail analysis of each failure. The list of missing entries (generated automatically during the import process) can be found in this ActivePaper under /documentation/missing-entries.</p>
Dendritic cell subsets in oral mucosa of allergic and healthy subjects
<p><strong>Abstract</strong></p> <p>Immunohistochemistry was used to identify, enumerate, and describe the tissue distribution of Langerhans type (CD1a and CD207), myeloid (CD1c and CD141), and plasmacytoid (CD303 and CD304) dendritic cell subsets in oral mucosa of allergic and non-allergic individuals. Allergic individuals have more CD141+ myeloid cells in epithelium and more CD1a+ Langerhans cells in the lamina propria compared to healthy controls, but similar numbers for the other DC subtypes. Our data are the first to describe the presence of CD303+ plasmacytoid DCs in human oral mucosa and a dense intraepithelial network of CD141+ DCs. The number of Langerhans type DCs (CD1a and CD207) and myeloid DCs (CD1c), was higher in the oral mucosa than in the nasal mucosa of the same individual independent of the atopic status.</p>
Grib and ASCII data, subset ERA-I for shallow water waves Ocean Science study
<p>Specific output from ERA-I reanalysis (wave model component) containing interated parameters, see https://doi.org/10.5194/os-13-1-2017</p>
The winter subset of the Saildrone 2021-2022 Mission to the Gulf Stream used for the publication "The importance of contemporaneous measurements for regional air-sea CO2 flux estimates"
<p>Data from the Saildrone 2021-2022 observational mission to the Gulf Stream. These data are published to accompany the publication "The importance of contemporaneous measurements for regional air-sea CO<sub>2</sub> flux estimates." Included in this dataset are the primary and processed variables used throughout the paper. The data associated with each saildrone is named by the drone number. Additionally, included in the structure for each drone are the gas transfer velocities for each scenario, MBL atmospheric CO2 interpolated to the time and location of the drone, and ERA-5 wind speed, sea level pressure, significant wave height, and drag coefficient interpolated to the time and location of the drone. These variables are used to calculate CO<sub>2</sub> fluxes for each scenario and are named as follows: "F" + gas transfer velocity equation used (DM18 or W14) + drone ID + scenario. Scenario A-D correspond to those outlined in the paper. Scenarios E and F correspond to the calculation of air-sea fluxes using all saildrone observed variables except for atmospheric CO<sub>2</sub> (from MBL product) and significant wave height (from ERA-5), respectively. </p>
Subset of Project FeederWatch data demonstrating observer shift
<p>This is a subset of data from the Project FeederWatch dataset, <a href="https://feederwatch.org/explore/raw-dataset-requests/">available on the their website</a>. This dataset is restricted to a 15 km radius around Washington, D.C., and includes observations from the most consistent observers. My GitHub repository, titled <a href="https://github.com/GatesDupont/observer_shift">observer_shift</a>, includes R scripts that processes the raw Project FeederWatch dataset and generates this data subset.</p>
Subset of GSE192456
<p>This repository contains the code and materials for the paper "Teaching Biomedical Students to Use Complex Objects for Omics Data Storage and Analysis: A Classroom Implementation Strategy Using Jupyter Notebooks"</p> <ul> <li>Leonardo D. Garma*, Breast Cancer Clinical Research Unit, Centro Nacional de Investigaciones Oncológicas – CNIO, Madrid, Spain;</li> <li>Nuno S. Osório*, Life and Health Sciences Research Institute (ICVS), School of Medicine, University of Minho, Portugal and ICVS/3B’s –PT Government Associate Laboratory, Braga, Portugal</li> <li>Corresponding authors: <a href="mailto:lgarma@cnio.es">lgarma@cnio.es</a>, <a href="mailto:nosorio@med.uminho.pt">nosorio@med.uminho.pt</a></li> </ul>
E3SM simulations of Hurricane Irene (Delaware River basin subset)
<p>This repository consists of the E3SM simulation outputs of Hurricane Irene (subsetted within Delaware River basin) associated with the manuscript: "Simulation of Compound Flooding using River-Ocean Two-way Coupled E3SM Ensemble on Variable-resolution Meshes". Below are the descriptions for each file:</p> <p>EAM_ensemble.nc - 25 EAM ensemble simulations.<br>MOSART_1way_baseline.nc - 25 MOSART ensemble simulations for Experiment 1way_baseline<br>MOSART_1way_datm_GSWP.nc - MOSART simulation for Experiment 1way_datm using GSWP forcing<br>MOSART_1way_datm_JRA.nc - MOSART simulation for Experiment 1way_datm using JRA forcing<br>MOSART_1way_r0125.nc - 25 MOSART ensemble simulations for Experiment 1way_r0125<br>MOSART_2way.nc - 25 MOSART ensemble simulations for Experiment 2way<br>MOSART_1way_SSH.nc - 25 MOSART ensemble simulations for Experiment 1way_SSH<br>MOSART_1way_MSL.nc - 25 MOSART ensemble simulations for Experiment 1way_MSL<br>MPASO_subset_1way_ens001.nc ~ MPASO_subset_1way_ens025.nc - 25 MPAS-O ensemble simulations for Experiment 1way_baseline<br>MPASO_subset_2way_ens001.nc ~ MPASO_subset_2way_ens001.nc - 25 MPAS-O ensemble simulations for Experiment 2way</p>
COAMPS-TC atmospheric data subset for Hurricane Michael
<p>Hurricane Michael was a Category 5 hurricane when it made landfall on the Florida Panhandle on 10 October 2018. The data subset of the Hurricane Michael Coupled Ocean/Atmosphere Mesoscale Prediction System for Tropical Cyclones (COAMPS-TC) was utilized for the manuscript "In situ observations at the air-sea interface by expendable air-deployed drifters under Hurricane Michael (2018)". Zonal and meridional 10-m winds were subset for designated time steps and interpolated to the time and location of the drifter observations. The wind fields were used in calculations of wave age and to compare to drifter observed winds and European Centre for Medium-Range Weather Forecasts (ECMWF) Earth Reanalysis version 5 reanalysis (ERA5) wind.</p>
10K-Cell Subset of PBMC CITE-Seq Dataset for CITEViz
<p>This repository contains an example CITE-Seq data (10K peripheral blood mononuclear cells) to test the CITEViz program. The CITEViz preprint is available <a href="https://www.biorxiv.org/content/10.1101/2022.05.15.491411v1">here</a>, and the documentation website is located <a href="https://maxsonbraunlab.github.io/CITEViz/">here</a>. The original data underlying this article are available in GEO (Gene Expression Omnibus) at <a href="https://www.ncbi.nlm.nih.gov/geo/">https://www.ncbi.nlm.nih.gov/geo/</a>, and can be accessed with GSE164378. </p>
LIT-PCBA nine targets subset
<p>This dataset is the accompanying data for the submitted manuscript:<br>"An ANI-2 Enabled Open-Source Protocol To Estimate Ligand Strain After Docking"</p>
Gene Wiki subsets
<p>Gene Wiki Subset of Wikidata created with wdsub (https://github.com/weso/wdsub).</p> <p>Subsets are described using Shape Expressions: https://github.com/weso/genewikisub/tree/master/shex/directP31</p> <p>Input: https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.gz downloaded between 2021/12/06 10:26:10 pm (CET) and 2021/12/07 08:40:42 am (CET)</p>
Coloc summary results for "Dissection of multiple sclerosis genetics identifies B and CD4+ T cells as driver cell subsets"
<p>Text files containing coloc results between MS GWAS loci and CD4 T and B cell cis-eQTLs from DICE. These results accompany the paper "<strong>Dissection of multiple sclerosis genetics identifies B and CD4+ T cells as driver cell subsets"</strong></p>
TIGER training dataset (ROI-level annotations of WSIROIS subset)
<p>This dataset contains data and ROI-level annotations of the WSIROIS subset of the TIGER training dataset, released in conjunction with the <a href="https://tiger.grand-challenge.org/">TIGER challenge</a>. Note that the WSIROIS dataset with whole-slide image-level annotations can be downloaded via the <a href="https://tiger.grand-challenge.org/Data/">Data </a>section of the TIGER challenge, together with the two additional subsets released with the challenge, namely the WSIBULK and the WSITILS subsets. </p> <p>The data is derived from digital pathology images of Her2 positive (Her2+) and Triple Negative (TNBC) breast cancer whole-slide images, together with manual annotations. Data comes from multiple sources. A subset of Her2+ and TNBC cases is provided by the Radboud University Medical Center (RUMC) (Nijmegen, Netherlands). A subset of Her2+ and TNBC cases is provided by the Jules Bordet Institut (JB) (Bruxelles, Belgium). A third subset of TNBC cases only is derived from the TCGA-BRCA archive obtained from the Genomic Data Commons Data Portal.</p> <p>This dataset of ROI-level annotation of WSIROIS is released in a format that is fully compatible with segmentation and detection pipelines used in the computer vision community. For this reason, we release regions of interest and manual annotations in PNG format and cell locations as bounding boxes in COCO format. In this way, we hope to make TIGER accessible to people that do not have experience with whole-slide images but still want to participate and contribute to this project.</p> <p>In this set, we release regions of interest from n=195 whole-slide images of breast cancer, both (core-needle) biopsies and surgical resections, with regions of interest (ROI) selected and manually annotated. All data (both images and manual annotations) are released at 0.5 um/px magnification. This dataset contains images and annotations from multiple sources:</p> <ul> <li><strong>TCGA: </strong>regions of interest cropped from<strong> </strong>n=151<strong> </strong>WSIs of TNBC cases from the TGCA-BRCA archive (the original slides can also be downloaded from the <a href="https://portal.gdc.cancer.gov/">GDC Data Portal</a>). Annotations are extracted and adapted from the publicly available <a href="https://bcsegmentation.grand-challenge.org/">BCSS</a> and <a href="https://nucls.grand-challenge.org/">NuCLS</a> datasets. </li> <li><strong>RUMC: </strong>regions of interest cropped from n=26 WSIs of TNBC and Her2+ cases from Radboud University Medical Center (Netherlands). Annotations were made by a panel of board-certified breast pathologists.</li> <li><strong>JB: </strong>regions of interest cropped from n=18 WSIs of TNBC and Her2+ cases from Jules Bordet Institute (Belgium). Annotations were made by a panel of board-certified breast pathologists.</li> </ul> <p>In this dataset, we release ROI-level annotations of both tissue compartments and cells. ROI images are released in PNG format; cell annotations are released as bounding boxes in the standard COCO format for object detection; tissue compartment annotations are released as PNG images containing pixel-wise class labels. In each image file, the coordinates of the region of interest in the WSI are indicated in the filename as imagefilename_[x1,y1,x2,y2].png, where (x1,y1) are the coordinates of the top-left corner and (x2,y2) are the coordinates of the bottom-right corner of each ROI.</p> <p>Check the <a href="https://tiger.grand-challenge.org/Data/">Data</a> section of the TIGER challenge for additional information about this dataset.</p> <ul> </ul>
Segmentation Zoo UNet models for Landsat-8 satellite imagery, Coast Train v1 Landsat-8 4-class subset.
<p><strong>Doodleverse/Segmentation Zoo UNet models for Landsat-8 satellite imagery, Coast Train v1 Landsat-8 4-class subset.</strong></p> <p>These UNet model data are based on the Coast Train v1 Landsat-8 labeled imagery subset. Models have been fitted to 4 different types of data</p> <p>1. NDWI (1 band): (g-nir)/(g+nir)</p> <p>2. MNDWI (1 band): (swir-g)/(swir+g)</p> <p>3. RGB (3 band): red, green, blue</p> <p>4. RGB-NIR-SWIR (5 band): red, green, blue, nir, swir</p> <p>Classes are: {0: water, 1: whitewater, 2:sediment, 3:other}. These classes have been remapped from the original 11 classes<br> </p> <p>These files are used in conjunction with Segmentation Zoo*</p> <p>For each model, there are 3 files with the same root name:</p> <p>1. <strong>'.json' </strong>config file: this is the file that was used by Segmentation Gym** to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p> </p> <p>2.<strong> '.h5'</strong> weights file: this is the file that was created by the Segmentation Gym** function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym** function `seg_images_in_folder.py` or the Segmentation Zoo* function `select_model_and_batch_process_folder.py` to segment a folder of images</p> <p> </p> <p>3.<strong> '_modelcard.json'</strong> model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p> </p> <p>References</p> <p>* https://github.com/Doodleverse/segmentation_zoo</p> <p>** https://github.com/Doodleverse/segmentation_gym</p>
Segmentation Zoo Res-UNet models for Landsat-8 satellite imagery, Coast Train v1 Landsat-8 4-class subset.
<p><strong>Doodleverse/Segmentation Zoo models for Landsat-8 satellite imagery, Coast Train v1 Landsat-8 4-class subset.</strong></p> <p>These model data are based on the Coast Train v1 Landsat-8 labeled imagery subset. Models have been fitted to 4 different types of data</p> <p>1. NDWI (1 band): (g-nir)/(g+nir)</p> <p>2. MNDWI (1 band): (swir-g)/(swir+g)</p> <p>3. RGB (3 band): red, green, blue</p> <p>4. RGB-NIR-SWIR (5 band): red, green, blue, nir, swir</p> <p>Classes are: {0: water, 1: whitewater, 2:sediment, 3:other}. These classes have been remapped from the original 11 classes<br> </p> <p>These files are used in conjunction with Segmentation Zoo*</p> <p>For each model, there are 3 files:</p> <p>1. config file: this is the file that was used by Segmentation Gym** to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p> </p> <p>2. weights file: this is the file that was created by the Segmentation Gym** function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym** function `seg_images_in_folder.py` or the Segmentation Zoo* function `select_model_and_batch_process_folder.py` to segment a folder of images</p> <p> </p> <p>3. model card file: this is a json file containing the following fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata</p> <p> </p> <p>References</p> <p>* https://github.com/Doodleverse/segmentation_zoo</p> <p>** https://github.com/Doodleverse/segmentation_gym</p> <p> </p>
Wikidata Subsets of 4 Gene Wiki Classes
<p>Chemical compound (Q11173), disease (Q12136), gene (Q7187), and protein (Q8054) are some of the main classes containing the Gene Wiki WikiProjects. In this repository, we put the corresponding subsets of each class. The subsets contain all instances of the four main classes (no sub-classes). All subsets are extracted from the Wikidata JSON dump of 3 January 2022 (<a href="https://t.co/vdmJc8V1v2">https://t.co/vdmJc8V1v2</a>)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.