Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
151
datasets available to search
ShareScore release 0.7.1
Dataset results
151 results for “Reference dataset”
Protein haplotype sequences obtained by ProHap from the Haplotype Reference Consortium Release 1.1 dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Haplotype Reference Consortium, Release 1.1 (<a href="https://ega-archive.org/datasets/EGAD00001002729" target="_blank" rel="noopener">https://ega-archive.org/datasets/EGAD00001002729</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>Release 1.1 of the HRC is provided aligned with the GRCh37 reference genome. We have performed a liftover to the GRCh38 reference using GeneBe (https://genebe.net/tools/liftover). Variants for which the reported alternative allele is considered as reference in GRCh38 were removed. A threshold of 1% minor allele frequency was applied to filter the remaining variants. After translation, a frequency threshold of 0.5% was applied to filter the resulting unique non-canonical sequences. The complete configuration file for the ProHap run is attached to this repository.</p> <p>This dataset contains one compressed directory, contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Mar2025 reference datasets for Mitohelper
<p><strong>Reference datasets (Mar 2025 update) for Mitohelper (</strong><a href="https://github.com/aomlomics/mitohelper"><em><strong>https://github.com/aomlomics/mitohelper</strong></em></a><strong>)</strong></p> <p><a href="http://github.com/aomlomics/mitohelper">Mitohelper</a> is a repository built to facilitate experimental design, alignment visualization, and reference sequence analysis in fish eDNA studies. Refer to our <a href="https://doi.org/10.1002/edn3.187">paper</a> and Mitohelper's <a href="https://github.com/aomlomics/mitohelper/wiki">wiki</a> for database construction pipeline.</p> <p>I. Reference database files in tab-separated format, containing gene, taxonomy, and sequence information:</p> <ul> <li>mitofish.all.Mar2025.tsv (883,519 records)</li> <li>mitofish.12S.Mar2025.tsv (61,379 records of 12S rRNA gene >50 bp long)</li> <li>mitofish.12S.Mar2025_NR.fasta (FASTA file of 12S rRNA gene records)</li> <li>mitofish.COI.Mar2025.tsv (323,337 records)</li> </ul> <p>II. De-replicated QIIME 2-compatible 12S/12S+16S+18S rRNA reference datasets:</p> <ul> <li>12S-seqs-derep-uniq.qza</li> <li>12S-tax-derep-uniq.qza</li> <li>12S-16S-18S-seqs.qza</li> <li>12S-16S-18S-tax.qza</li> </ul> <p>If you use Mitohelper, please cite:<br>Jean Lim, S, Thompson, LR. Mitohelper: A mitochondrial reference sequence analysis tool for fish eDNA studies. Environmental DNA. 2021; 00: 1– 10. <a href="https://doi.org/10.1002/edn3.187">https://doi.org/10.1002/edn3.187</a></p>
Epidemiological and clinical characteristics predictive of ICU mortality of traumatic brain injury patients treated at a trauma reference hospital – A cohort study - Dataset
<p><strong>Dataset of a cohort whose summary is described below.</strong></p> <p><strong>ABSTRACT</strong></p> <p><strong>Background</strong>: Traumatic brain injury (TBI) has substantial physical, psychological, social and economic impacts, with high rates of morbidity and mortality. Considering its high incidence, the aim of this study was to identify epidemiological and clinical characteristics that predict mortality in patients hospitalized for TBI in intensive care units (ICUs). <strong>Methods</strong>: A retrospective cohort study was carried out with patients over 18 years old with TBI admitted to an ICU of a Brazilian trauma referral hospital between January 2012 and August 2019. TBI was compared with other traumas in terms of clinical characteristics of ICU admission and outcome. Univariate and multivariate analyses were used to estimate the odds ratio for mortality. <strong>Results</strong>: Of the 4816 patients included, 1114 had TBI, with a predominance of males (85.1%). Compared with patients with other traumas, patients with TBI had a lower mean age (45.3 ± 19.1 versus 57.1 ± 24.1 years, p < 0.001), higher median APACHE II (19 versus 15, p <0.001) and SOFA (6 versus 3, p < 0.001) scores, lower median Glasgow Coma Scale (GCS) score (10 versus 15, p < 0.001), higher median length of stay (7 days versus 4 days, p < 0.001) and higher mortality (27.6% versus 13.3%, p < 0.001). In the multivariate analysis, the predictors of mortality were older age (OR: 1.008 [1.002-1.015], p = 0.016), higher APACHE II score (OR: 1.180 [1.155-1.204], p < 0.001), lower GCS score for the first 24 hours (OR: 0.730 [0.700-0.760], p < 0.001), and greater number of brain injuries and presence of associated chest trauma (OR: 1.727 [1.192-2.501], p < 0.001). <strong>Conclusion</strong>: Patients admitted to the ICU for TBI were younger and had worse prognostic scores, longer hospital stays and higher mortality than those admitted to the ICU for other traumas. The independent predictors of mortality were advanced age, APACHE II score, first 24-hour GCS score, number of brain injuries and chest trauma.</p>
The Object Detection for Olfactory References (ODOR) Dataset
<p><strong>The Object Detection for Olfactory References (ODOR) Dataset</strong></p> <p>Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. </p> <p>Existing datasets provide instance-level annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The ODOR dataset fills this gap, offering 38,116 object-level annotations across 4,712 images, spanning an extensive set of 139 fine-grained categories. </p> <p>It has challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. </p> <p>Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.</p> <p><strong>How to use</strong></p> <p>The annotations are provided in COCO JSON format. To represent the two-level hierarchy of the object classes, we make use of the supercategory field in the categories array as defined by COCO. In addition to the object-level annotations, we provide an additional CSV file with image-level metadata, which includes content-related fields, such as Iconclass codes or image descriptions, as well as formal annotations, such as artist, license, or creation year. </p> <p>In addition to a zip containing the dataset images, we provide links to their source collections in the metadata file and a Python script to conveniently download the artwork images (`download_imgs.py`).</p> <p>The mapping between the `images` array of the `annotations.json` and the `metadata.csv` file can be accomplished via the `file_name` attribute of the elements of the `images` array and the unique `File Name` column of the `metadata.csv` file, respectively.</p>
Protein haplotype sequences obtained by ProHap from the Human Pangenome Reference Consotruim dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Human Pangenome Reference Consotruim (HPRC), first release (<a href="https://github.com/human-pangenomics/hpp_pangenome_resources">https://github.com/human-pangenomics/hpp_pangenome_resources</a>), 44 samples. We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>This repository contains one database created using all 43 samples of the HPRC release (the haplotypes of the sample NA21309 did not encode any non-canonical sequences), and then a database for each of the 43 samples separately. No filtering on allele frequency or haplotype frequency was applied in any of the databases. The complete configuration file for the ProHap run is attached to this repository.</p> <p>There is one compressed directory for each of the databases, containing the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>For this dataset, only the simplified format is provided. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
A Fully-Parameterized Object-Side Light Field Dataset and Theory for Using Entrance and Exit Pupils as Natural Light Field Reference Planes for an Unfocused Plenoptic Camera
<p>We describe a dataset of light fields with full object-side parameterizations. The dataset contains PNG and ESLF files for all 32 images. 12 of them additionally contain depth maps and point clouds.</p>
Datasets of "Differences in the stool metabolome between vegans and omnivores: analyzing the NIST stool reference material" publication
<p>To gain confidence in results of omic-data acquisitions, methods must be benchmarked by validated quality control materials. We here report data combining both untargeted and targeted metabolomics assays for the analysis of four new human fecal reference materials developed by the U.S. National Institute of Standards and Technologies (NIST) for metagenomics and metabolomics measurements. These reference grade test materials (RGTM) were established by NIST based on two different diets and two different samples treatments: homogenized fecal matter from subjects eating vegan diets, stored and submitted in either lyophilized (RGTM 10162) or aqueous form (RGTM 10171); secondly, homogenized fecal matter from subjects eating omnivore diets, stored and submitted in either lyophilized (RGTM 10172) or aqueous form (RGTM 10173). We used four untargeted metabolomics assays (lipidomics, primary metabolites, biogenic amines and polyphenols) and one targeted assay on bile acids.</p>
Reference Dataset for Benchmarking Organ Doses Derived from Monte Carlo Simulations of CT Exams
<p>This reference dataset contains CT scanner x-ray source characteristics, filtration profile, de-identified patient image data and size characteristics, voxelized patient models, exam characteristics, x-ray tube current data, and organ dose results in tabular form from Monte Carlo (MC) simulations of abdominal/pelvis CT exams of pregnant patients. This dataset can be used for benchmarking MC simulation codes for CT dosimetry.</p>
Global Pasture Watch - Grassland reference samples based on visual interpretation of VHR imagery and harmonized datasets (2000–2024)
<p>Reference point samples used in the production of the <a href="https://doi.org/10.5281/zenodo.13890401">global maps of annual grassland class and extent for 2000—2022</a><strong> </strong>within the scope of the <a href="https://landcarbonlab.org/data/global-grassland-and-livestock-monitoring/">Global Pasture Wath</a> initiative. </p> <p>The reference samples (estabilished by Feature Space Coverage Sampling-FSCS) comprises <strong>2.3M points</strong> visually classified (<em>using Very High Resolution imagery</em>) in:</p> <ol> <li><strong>Cultivated grassland,</strong></li> <li><strong>Natural/semi-natural grassland</strong></li> <li><strong>Other land cover</strong></li> </ol> <p>The file <code>gpw_grassland_fscs.vi.vhr_tile.samples_20000101_20241231_go_epsg.4326_v2.gpkg</code> aggregates the samples by visual interpretation units ( 1x1 km) and includes the follow collumns:</p> <ul> <li>cluster_id: Cluster id defined by k-means (FSCS),</li> <li>cluster_distance: Distance from the sample tile to center of the cluster (FSCS),</li> <li>cluster_size: Size of cluster (strata) defined by the FSCS,</li> <li>priority: Priority used by the visual interpretation,</li> <li>tile_id: Sample tile id,</li> <li>imagery: VHR reference images used by the visual interpretation,</li> <li>min_year: Minimum of year covered by the reference samples,</li> <li>max_year: Maximum of year covered by the reference samples,</li> <li>n_years: Number of years covered by the reference samples,</li> <li>n_samples_c1: Number of reference samples for "Cultivated grass" (1),</li> <li>n_samples_c2: Number of reference samples for "Natural / Semi-natural grass" (2),</li> <li>n_samples_c3: Number of reference samples for "Open Shrubland" (2),</li> <li>n_samples_c4: Number of reference samples for "Not grass" (3),</li> <li>n_samples_all: Total number of reference samples,</li> </ul> <p>The file <code>gpw_grassland_fscs.vi.vhr_point.samples_20000101_20241231_go_epsg.4326_v2.gpkg</code> provides individual points (with 60-m spatial support) and include the follow collumns:</p> <ul> <li>sample_id: Sample id deribed by MD5 Hash of columns x, y, imagery and year,</li> <li>x: Longitude in WGS84 (EPSG:4326),</li> <li>y: Latitude in WGS84 (EPSG:4326),</li> <li>vi_tile_id: 1-km tile id,</li> <li>tile_id: GLAD tile id (1x1 degree)</li> <li>imagery: VHR Reference image used by the visual interpretation (Google; Bing; Interpolated),</li> <li>ref_date: Reference date of GPW samples (based on VHR image) and of other existing datasets,</li> <li>year: Reference year of GPW samples (based on VHR image) and of other existing datasets,</li> <li>class: Class id (1: Cultivated grassland; 2: Natural/semi-natural grassland; 3: Open shrubland; 4: Other land cover) ,</li> <li>class_label: Class labels (Cultivated grassland; Natural/semi-natural grassland; Open shrubland; Other land cover) ,</li> <li>dataset_name: Existing dataset names (CGLS-LC, EuroCrops, GeoWiki, GeoWiki-feedback, LCMap-Conus, LUCAS, MapBiomas, WorldCereal, GPW) <br>dataset_class: Original land cover class provided by the maintainer of existing dataset</li> <li>esa_worldcover_2020: Land cover class labels extracted from ESA WorldCover 2020,</li> <li>glad_glcluc_yyyy: Land cover class labels extracted from UMD GLAD GLCLUC for the reference date,</li> <li>glc_fcs30d_yyyy: Land cover class labels extracted from GLC_FCS30D for the reference date,</li> <li>gpw_fscs_cluster: K-Means output ranging from 0—9999 according to Feature Space Coverage Sampling (FSCS),</li> <li>ml_cv_group: spatial block CV group (based on vi_tile_id),</li> <li>ml_type: specify if the sample was used for (1) training or (2) calibration.</li> </ul> <p>The file <code>gpw_grassland_fscs.vi.vhr_grid.samples_20000101_20241231_go_epsg.4326_v2.gpkg</code> provides the grid samples (with 10-m spatial support) and include the follow collumns:</p> <ul> <li>tile_id: 1-km tile id,</li> <li>bing_class: Class labels (Cultivated grassland; Natural/semi-natural grassland; Other land cover) defined using as reference Bing Maps Images,</li> <li>bing_image_start_date: Start date of the Bing Maps Images used in the visual interpretation,</li> <li>bing_image_end_date: End date of the Bing Maps Images used in the visual interpretation,</li> <li>google_class: Class labels (Cultivated grassland; Natural/semi-natural grassland; Other land cover) defined using as reference Google Maps Images,</li> <li>google_image_start_date: Start date of the Google Maps Images used in the visual interpretation,</li> <li>google_image_end_date: End date of the Google Maps Images used in the visual interpretation,</li> <li>missing_image_date: No images available,</li> <li>same_image_bing_google: Images from the same date available in Google and Bing Maps.</li> </ul> <p>The dataset was produced through the <a href="https://plugins.qgis.org/plugins/qgis-fgi-plugin/">QGIS plugin Fast Grid Inspection</a>.</p> <h3>Related resources</h3> <ul> <li><strong>Maps of dominant grassland:</strong><br><a href="https://zenodo.org/records/13890400">2000-2002</a> <a href="https://zenodo.org/records/13890402">2003-2005</a> <a href="https://zenodo.org/records/13890404">2006-2008</a> <a href="https://zenodo.org/records/13890408">2009-2011</a> <a href="https://zenodo.org/records/13890410">2012-2014</a> <a href="https://zenodo.org/records/13890412">2015-2017</a> <a href="https://zenodo.org/records/13890414">2018-2020</a> <a href="https://zenodo.org/records/13890416">2021-2022</a></li> <li><strong>Probability maps of cultivated grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Probability maps of natural/semi-natural grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Grassland reference samples based on VHR imagery (2000–2022):</strong><br><a href="https://doi.org/10.5281/zenodo.11281157">GeoPackage files</a></li> <li><strong>Global machine learning models (Random Forest):</strong><br><a href="https://doi.org/10.5281/zenodo.13952806">Parquet and joblib python files</a></li> <li><strong>Reference sampling design derived by FSCV:</strong><br><a href="https://doi.org/10.5281/zenodo.11391517">GeoPackage and raster files</a></li> <li><strong>Harmonized reference samples based on existing LULC dataset:</strong><br><a href="https://doi.org/10.5281/zenodo.13951976">GeoPackage and raster files</a></li> <li><strong>Source code for reproducibility:<br></strong><a href="https://doi.org/10.5281/zenodo.13952867">GitHub release</a><strong><br></strong></li> <li><strong>Mapping feedback tool:</strong><br><a href="https://geo-wiki.org">GeoWiki</a></li> <li><strong>Data catalogues:</strong><br><a href="https://stac.openlandmap.org/gpw_ggc-30m/collection.json?.language=en">OpenLandMap STAC</a> <a href="https://global-pasture-watch.projects.earthengine.app/view/ggc-30m">Google Earth Engine</a></li> </ul> <h3>Support</h3> <p>For questions of bugs/inconsistencies related to the dataset raise a GitHub issue in <a href="https://github.com/wri/global-pasture-watch">https://github.com/wri/global-pasture-watch</a></p>
Dataset of publication "Derivation and validation of a reference data-based real gas model for hydrogen"
<p>In this repository, a new real gas model for hydrogen based on the Reference Fluid Thermodynamic and Transport Properties Database (REFPROP) v10.0 is provided for the use in the simulation software OpenFOAM v2012. The model is valid in a temperature and pressure range of 150-400 K and 0.1-1000 bar, respectively. Usage beyond this range is not recommended as it may lead to unrealistic results.</p>
Reference dataset for comparison of cloud detection algorithms for Sentinel-2 imagery
<p>Sentinel-2 cloud mask reference dataset generated and analyzed as part of Tarrio, K., Tang, X., Masek, J.G., Claverie, M., Ju, J., Qiu, S., Zhu, Z. and Woodcock, C.E., 2020. Comparison of cloud detection algorithms for Sentinel-2 imagery. Science of Remote Sensing, 2, p.100010. [https://www.sciencedirect.com/science/article/pii/S2666017220300092](https://www.sciencedirect.com/science/article/pii/S2666017220300092)</p> <p><strong>1. Reference masks</strong></p> <p><strong>Algorithms:</strong></p> <ul> <li>Fmask 1.x</li> <li>Fmask 2.x</li> <li>Fmask 4.x</li> <li>Tmask</li> <li>Sen2Cor</li> <li>MAJA</li> <li>LaSRC</li> </ul> <p><strong>Locations:</strong></p> <ul> <li>South Africa (35JPM)</li> <li>Senegal (28PDC)</li> <li>Switzerland (32TLT)</li> <li>France (31TCJ, 31TFJ)</li> <li>Morocco (29RNQ)</li> </ul> <p><strong>Standardized legend:</strong></p> <p>Original algorithm outputs were standardized to the same categorical legend.</p> <ul> <li>0 = clear land</li> <li>1 = clear water</li> <li>2 = cloud shadow</li> <li>3 = snow/ice</li> <li>4 = cloud</li> </ul> <p>All reference masks processed to both 10m and 30m resolution, with the exception of Tmask, which is available only at a 30m resolution.</p> <p><strong>Mask naming convention:</strong></p> <p>All processed masks are named according to the following convention:<br> M<*resolution*><*S2 MGRS tile ID*><*YYYY*><*DOY*><*algorithm*><br> e.g. **M30T28PDC2016351TMASK**</p> <p><br> <strong>2. Interpreted sample points</strong></p> <p>Sample points were selected based on agreement among different map products. This record includes a shapefile with the final interpretations for each of the sampled sites. (See publication for additional information.)</p>
Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper
<p>Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper (complement for a tutorial at https://github.com/datngu/LmTag)</p> <p>This repo includes chromosome 10 reference panel data constructed for 3 populations:</p> <p>- EAS</p> <p>- EUR</p> <p>- SAS</p>
Dataset: A systematic multi-technique comparison of luminescence characteristics of two reference quartz samples
<p>Original dataset corresponding to the manuscript Schmidt et al. submitted to the Journal of Luminescence. The dataset is structured as follows: </p> <ul> <li>01_original_data <ul> <li>This folder contains the original measurement data for each experiment (e.g., LM-OSL, TL), organised following the manuscript figures, e.g., Fig_4. These subfolders contain the original measurement data and the measurement sequence. For Risø readers, the output format is *.binx and *.xsyg for results from the Freiberg Instruments lexsyg readers (e.g., TR-OSL).</li> </ul> </li> <li>02_processed_data <ul> <li>Comma-separated value (*.csv) to reproduce each figure in the manuscript. </li> </ul> </li> </ul> <p> </p>
Human kidney cortex CODEX reference dataset 1
<p>CODEX image stack of human kidney cortex stained with markers as indicated in CODEX_antibody_list_010621.csv and imaged in the order given in CODEX_channel_index_010621.csv. Tissue preparation and analysis as described <a href="https://www.biorxiv.org/content/10.1101/2021.12.27.474025v1">here.</a></p>
Diagnostic accuracy of a set of clinical and radiological criteria for screening of COVID-19 using RT-PCR as the reference standard - Dataset
<p>Dataset of a cohort whose summary is described below.</p> <p>Abstract</p> <p><strong>Objective:</strong> To evaluate the accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) of a set of clinical-radiological criteria for COVID-19 screening in patients with severe acute respiratory failure (SARF) admitted to intensive care units (ICUs), using reverse-transcriptase polymerase chain reaction (RT-PCR) as the reference standard. <strong>Method: </strong>Diagnostic accuracy study including a historical cohort of 1009 patients consecutively admitted to ICUs across six hospitals in Curitiba (Brazil) from March to September, 2020. The sample was stratified into groups by the strength of suspicion for COVID-19 (strong <em>versus</em> weak) using parameters based on three clinical and radiological (chest computed tomography) criteria. The diagnosis of COVID-19 was confirmed by RT-PCR (referent). <strong>Results:</strong> With respect to RT-PCR, the proposed criteria had 98.5% (95% confidence interval [95% CI] 97.5–99.5%) sensitivity, 70% (95% CI 65.8–74.2%) specificity, 85.5% (95% CI 83.4–87.7%) accuracy, PPV of 79.7% (95% CI 76.6–82.7%) and NPV of 97.6% (95% CI 95.9–99.2%). <strong>Conclusion: </strong>The proposed set of clinical-radiological criteria were accurate in identifying patients with strong <em>versus</em> weak suspicion for COVID-19 and had high sensitivity and considerable specificity with respect to RT-PCR. These criteria may be useful for screening COVID-19 in patients presenting with SARF.</p>
Characterisation and calibration of low-cost PM sensors at high temporal resolution to reference grade performances - dataset
<p>This repository contains the data used for the analysis of the paper "Characterisation and calibration of PM sensors at high temporal resolution to reference grade performances" submitted to Heliyon and available as a pre-print:</p> <p>Bulot, Florentin M. J. and Ossont, Steven J. and Morris, Andrew and Basford, Philip J. and Easton, Natasha H. C. and Mitchell, Hazel L. and Foster, Gavin L. and Cox, Simon J. and Loxham, Matthew, Characterisation and Calibration of Low-Cost Pm Sensors at High Temporal Resolution to Reference-Grade Performance. Available at SSRN: <a href="https://ssrn.com/abstract=4360707">https://ssrn.com/abstract=4360707</a> or <a href="http://dx.doi.org/10.2139/ssrn.4360707">http://dx.doi.org/10.2139/ssrn.4360707</a></p> <p> </p> <p>The code used to conduct the data analysis is available at <a href="https://doi.org/10.5281/zenodo.7261417">https://doi.org/10.5281/zenodo.7261417</a></p> <p> </p> <p>.</p> <p> </p> <p>The files are available in .csv and in .rds (for R) formats. For details about the measurement equipment used<br> during this study, please refer to the methods section of the paper.</p> <p>Description of the files.</p> <p>202007_to_202107_nocs - contains the data from the low-cost sensors</p> <p>It contains the following headers:<br> - "sensor" - sensor id<br> - "site" - name of the air quality monitor hosting the sensor<br> - "median_PM1" - PM1 mass concentration (ug/m3)<br> - "median_PM10" - PM10 mass concentration (ug/m3)<br> - "median_PM25" - PM25 mass concentration (ug/m3)<br> - "median_PM4" - PM4 mass concentration (ug/m3) (only available for SPS30)<br> - "median_n05" - particle number concentration (SPS30) of particles between 0.3um and 0.5um<br> - "median_n1" - particle number concentration (SPS30) of particles between 0.3um and 1um<br> - "median_n10" - particle number concentration (SPS30) of particles between 0.3um and 10um<br> - "median_n25" - particle number concentration (SPS30) of particles between 0.3um and 2.5um<br> - "median_n4" - particle number concentration (SPS30) of particles between 0.3um and 4um<br> - "median_gr03um" - particle number concentration (PMS5003) of particles >0.3um<br> - "median_gr05um" - particle number concentration (PMS5003) of particles >0.5um<br> - "median_gr100um" - particle number concentration (PMS5003) of particles >10um<br> - "median_gr10um" - particle number concentration (PMS5003) of particles >1um<br> - "median_gr25um" - particle number concentration (PMS5003) of particles >2.5um<br> - "median_gr50um" - particle number concentration (PMS5003) of particles >5um<br> - "median_pm100_cf1" - PM10 mass concentration with cf1 calibration for PMS5003<br> - "median_pm10_cf1" - PM1 mass concentration with cf1 calibration for PMS5003<br> - "median_pm25_cf1" - PM25 mass concentration with cf1 calibration for PMS5003<br> - "date" - date, format "yyyy-mm-dd HH:MM:SS GMT" </p> <p> </p> <p>df_pm_2min - contains the PM mass concentration data from the Fidas 200S.</p> <p>It contains the following headers:<br> - "PM2.5" - PM2.5 mass concentration (ug/m3) Fidas 200S<br> - "PM10" - PM10 mass concentration (ug/m3) Fidas 200S<br> - "PMtot" - PM total mass concentration (ug/m3) Fidas 200S<br> - "PM1" - PM1 mass concentration (ug/m3) Fidas 200S<br> - "date" - date, format "yyyy-mm-dd HH:MM:SS GMT" </p> <p> </p> <p>df_weather_2min - contains the weather data from the Fidas 200S</p> <p>It contains the following headers:<br> - "rh" - relative humidity (%)<br> - "dew_point_temperature" - dew point temperature (Celsius)<br> - "air_pressure" - Air pressure (hPa)<br> - "temperature" - temperature (Celsius)<br> - "date" - date, format "yyyy-mm-dd HH:MM:SS GMT"</p>
Reference site conditions for floating wind arrays: dataset of reference sites
<h1>IEA Task 49: Integrated Design of Floating Wind Arrays</h1> <h3>This dataset contains atmospheric and oceanographic data of 11 locations around the globe of future floating offshore wind farms.<br>Each dataset was used to perform a metocean analysis for preliminary design, published in the Work Package 1 Report of IEA Task 49.</h3> <div> <div> <div><a name="_msocom_1"></a></div> </div> </div> <table> <tbody> <tr> <td><strong>4COffshore ID</strong></td> <td><strong>Name</strong></td> <td><strong>Latitude [deg]</strong></td> <td><strong>Longitude [deg]</strong></td> <td><strong>Water Depth [m] (GEBCO)</strong></td> <td><strong>Distance from shore [m]</strong></td> <td><strong>Country</strong></td> <td><strong>Dataset curated by</strong></td> </tr> <tr> <td>IT95 </td> <td>Hannibal </td> <td> <div>37.842</div> </td> <td> <div>12.0722</div> </td> <td> <div> -353</div> </td> <td> <div> 35</div> </td> <td>Italy </td> <td>RSE</td> </tr> <tr> <td>US0W </td> <td> Humboldt</td> <td> <div> 40.928</div> </td> <td> <div>-124.708</div> </td> <td> <div> -707</div> </td> <td> <div> 43.8</div> </td> <td>USA</td> <td>NREL</td> </tr> <tr> <td>KR0R </td> <td>Ulsan </td> <td> <div> 35.449</div> </td> <td> <div> 129.949</div> </td> <td> <div> -188</div> </td> <td> <div> 32</div> </td> <td>South Korea </td> <td>UOU</td> </tr> <tr> <td>IE34 </td> <td>Moneypoint One</td> <td> <div> 52.519</div> </td> <td> <div>-10.276</div> </td> <td> <div> -1<a>02</a> </div> </td> <td> <div> 23.4</div> </td> <td>Ireland </td> <td>GDG</td> </tr> <tr> <td>UK6L </td> <td>Havbredey </td> <td> <div>58.862</div> </td> <td> <div> <div> <p>-5.54</p> </div> </div> </td> <td> <div> <div> <p> -91</p> </div> </div> </td> <td> <div> <div> <p> 41.6</p> </div> </div> </td> <td>Scotland </td> <td>DHI</td> </tr> <tr> <td>JP06 </td> <td>Fukushima</td> <td> <div>37.311</div> </td> <td> <div>141.251</div> </td> <td> <div> -90</div> </td> <td> <div> 19.4</div> </td> <td>Japan </td> <td>AIT</td> </tr> <tr> <td>NO44 </td> <td>Utsira nord<strong>*</strong></td> <td> <div>59.276</div> </td> <td> <div>4.541</div> </td> <td> <div> -273</div> </td> <td> <div> 42.4</div> </td> <td>Norway </td> <td>4subsea / UiS,UiB*</td> </tr> <tr> <td>USZ3 </td> <td>Gulf of Maine</td> <td> <div>43.25</div> </td> <td> <div>-69.5</div> </td> <td> <div> -148</div> </td> <td> <div> 138</div> </td> <td>USA </td> <td>NREL</td> </tr> <tr> <td>KR88 </td> <td>Geomundo<strong>**</strong></td> <td> <div>34.039</div> </td> <td> <div>126.901</div> </td> <td> <div> -70</div> </td> <td> <div> 47</div> </td> <td>South Korea </td> <td>IAE**</td> </tr> <tr> <td>FR87 </td> <td>Sud de la Bretagne II</td> <td> <div>47.325</div> </td> <td> <div>-3.659</div> </td> <td> <div> -94</div> </td> <td> <div> 30.7</div> </td> <td>France </td> <td>UiB</td> </tr> <tr> <td>NO66</td> <td>Sørlige Nordsjø II<strong>***</strong></td> <td> <div> 56.78 </div> </td> <td> <div> 4.92 </div> </td> <td> <div> -60</div> </td> <td> <div> 180</div> </td> <td>Norway</td> <td>4subsea / UiS,UiB***</td> </tr> </tbody> </table> <p>* Suplementary dataset published at: https://doi.org/10.5281/zenodo.10048048</p> <p>** Dataset is confidential, for details of usage reach out to the contact person.</p> <p>*** Suplementary dataset published at: https://doi.org/10.5281/zenodo.7057407</p>
Dataset for the comparison of performance of two leak detectors using hydrogen reference leaks (supplement to paper "Advancing Hydrogen Leak Detection: Design and Calibration of Reference Leaks")
<p>Excel file containing some measurements made in December 2023, using three hydrogen reference leaks, to assess the performance of two distinct leak detectors, one portable and made specifically for hydrogen and one MSLD in hydrogen-mode.</p>
InTeReC: In-text Reference Corpus - Single References Dataset
<p>This dataset contains a set of sentences extracted from articles published by the Public Library of Science (PLOS) up to September 2013. Information is given on the position of the sentences relative to the article and the section in which they appear, the section type with respect to the four main types of the IMRaD structure, as well as verb phrases that occur in the sentence. Each sentence contains one single in-text reference.</p> <p>The dataset is in the CSV format. Size: 314023 sentences.</p> <p>Column list:</p> <ul> <li><em>journal</em>: journal title</li> <li><em>doi</em>: DOI of the article from which the sentence was extracted</li> <li><em>article-length</em>: size of the article, as number of sentences</li> <li><em>article-pos</em>: position of the sentence in the article, as number of sentences from the beginning of the article</li> <li><em>section-length</em>: size of the section, as number of sentences</li> <li><em>section-pos</em>: position of the sentence in the section, as number of sentences from the beginning of the section</li> <li><em>section-type</em>: section type (see below)</li> <li><em>sentence-text</em>: full text of the sentence</li> <li><em>verb-phrases</em>: a list of verb phrases that occur in the sentence, comma separated</li> </ul> <p>Possible section types are:</p> <ul> <li>I: Introduction</li> <li>M: Methods</li> <li>R: Results</li> <li>D: Discussion</li> <li>MR: Methods and Results</li> <li>RD: Results and Discussion</li> </ul> <p> </p> <p>Full description of the construction of the dataset is published in:</p> <p>Marc Bertin and Iana Atanassova (2018) InTeReC : an In-text Reference corpus for applying Natural Language Processing to Bibliometrics. Bibliometric-enhanced Information Retrieval: 7th International BIR workshop (7th BIR workshop) at the 40th European Conference on Information Retrieval (ECIR).</p>
Top Quark Tagging Reference Dataset
<p>A set of MC simulated training/testing events for the evaluation of top quark tagging architectures.</p> <p>In total 1.2M training events, 400k validation events and 400k test events. Use “train” for training, “val” for validation during the training and “test” for final testing and reporting results.</p> <p><strong>Description</strong></p> <ul> <li> <p>14 TeV, hadronic tops for signal, qcd diets background, Delphes ATLAS detector card with Pythia8</p> </li> <li> <p>No MPI/pile-up included</p> </li> <li> <p>Clustering of particle-flow entries (produced by Delphes E-flow) into anti-kT 0.8 jets in the pT range [550,650] GeV</p> </li> <li> <p>All top jets are matched to a parton-level top within ∆R = 0.8, and to all top decay partons within 0.8</p> </li> <li> <p>Jets are required to have |eta| < 2</p> </li> <li> <p>The leading 200 jet constituent four-momenta are stored, with zero-padding for jets with fewer than 200</p> </li> <li> <p>Constituents are sorted by pT, with the highest pT one first</p> </li> <li> <p>The truth top four-momentum is stored as truth_px etc.</p> </li> <li> <p>A flag (1 for top, 0 for QCD) is kept for each jet. It is called is_signal_new</p> </li> <li> <p>The variable "ttv" (= test/train/validation) is kept for each jet. It indicates to which dataset the jet belongs. It is redundant as the different sets are already distributed as different files.</p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.