Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,129
datasets available to search
ShareScore release 0.9.0
Dataset results
2,129 results for “scores”
PMD hypomethylation human (hg19) neural network scores
<p>Global loss of DNA methylation in mammalian genomes occurs cumulatively as a mitotic process during aging and cancer, primarily in Partially Methylated Domains (PMDs). It has been shown that local sequence context (100bp) has a strong effect on the rate of demethylation of individual CpG dinucleotides within PMDs. Here, we train a deep learning model to characterize this sequence dependence further, finding that methylation loss can be predicted from a CpG’s 150bp sequence context alone with an AUC of 0.95. We use re-methylation rates of newly synthesized DNA to show that CpGs with fast-loss sequence context are inefficiently re-methylated. Interestingly, we find that the 10% of CpGs predicted to have the “slowest” rate of loss lose almost no DNA methylation in healthy cell types. These same slow-loss CpGs lose a significant amount of DNA methylation in cancer, suggesting that they could be responsible for deregulation of genes and transposable elements that are associated with DNA hypomethylation in cancer.</p> <p>This directory contains the Nov. 18, 2020 version of the human (hg19) CpG hypomethylation Neural network scores in a single tab-delimited (bedgraph) file:<br> <strong>multitissue-nn-scores.allCGs.0based.hg19.bedgraph.gz</strong><br> with the following columns:<br> 1: chromosome (hg19)<br> 2: start coord (hg19, 0-based)<br> 3: end coord (hg19, 0-based)<br> 4: multi-tissue NN score (0-1). Close to 0 is classified as slow-loss CpG, close to 1 is classified as fast loss CpG5: Num CpGs in 150 bp window (including central CpG, so minimum is 1).</p> <p> </p> <p>The full version of the NN scores with additional details are in the file <strong>zhou-bian.allCGs.1based.hg19.tsv.gz</strong></p> <p>Each row is a CG which provides (1) chromosome, (2) the corresponding C coordinate on the forward (watson) strand of the reference genome in one-based coordinates, (3) Neural network score, (4) number of CpGs within the 150bp sequence centered on this CpG, including the center CpG, (5) CpG is within a CpG island (0, no; 1, yes), CpG is within ENCODE blacklist (0, no; 1, yes)</p> <p> Here the CpG islands are the union set of Irizarry (Irizarry et al. 2009, Nat Genet), Takai-Jones (Takai et al. 2002, PNAS), Gardner-Gardin CGIs (Gardner-Gardin et al. 1987, J Mol Biol.). The blacklist was downloaded from https://github.com/Boyle-Lab/Blacklist/tree/master/lists.</p> <p>Additional files are included here:<br> <strong>zhou_pmds.0based.hg19.bed.gz</strong>: Input PMD CpGs from the Zhou (multi-tissue) dataset<br> <strong>bian_pmds.crc01.0based.hg19.bed.gz</strong>: Input PMD CpGs from the Bian (intra-tumor) dataset<br> <strong>zhou_bian_train_test_data.tar.gz</strong>: All training and test CpGs, including labels and sequence windows.</p> <p> </p> <p> </p>
UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits
<p>Summary-level GWAS data for 53 traits generated by <a href="https://www.genomicsplc.com/">Genomics plc</a> as presented in:</p> <p>Thompson D. et al. UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits (<a href="https://doi.org/10.1101/2022.06.16.22276246">https://doi.org/10.1101/2022.06.16.22276246</a>)</p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at <a href="mailto:research@genomicsplc.com">research@genomicsplc.com</a></p> <p><strong>NOTES</strong></p> <p>These analyses were carried out using the full UK Biobank (UKB) imputation data release (v3b). After removal of exclusions and withdrawals, a subset of 337,151 UKB individuals, the White British Unrelated (WBU) subgroup, was defined as the intersection of two sample groups created by Bycroft et al 2018 (Nature 562, 203-209): the ‘White British ancestry’ group (UKB Data Field 22006) and the ‘used in genetic principal components’ group (UKB Data Field 22020), the latter being high quality samples that were filtered to avoid closely related individuals. All GWAS analyses were performed on the WBU subgroup.</p> <p>Phenotypes were defined as described in Supplementary Table 1 ‘Phenotype definitions’ using a combination of Hospital Episode Statistics, Cancer Registry reports (where applicable) and self-report responses, with the exception of coronary artery disease (CAD). GWAS data was generated for both a “narrow” and a “broad” definition of CAD. The former was used as part of the training data for the Enhanced CAD PRS, the latter was used as part of the training data for the Enhanced CVD PRS. The phenotype definitions for “narrow” and a “broad” CAD are as follows:</p> <table> <tbody> <tr> <td>Narrow CAD<br> (includes angina)</td> <td>ICD10 codes (where .X indicates all subcodes) from both hospital and death records: I21, I22, I23, I24.1, I25.2, I20.X. ICD9 codes: 410-412, 42979, 413.X. OPCS-4 codes (K40.1–40.4, K41.1–41.4, K45.1–45.5,K49.1–49.2, K49.8–49.9, K50.2, K75.1–75.4, K75.8–75.9), self-reported heart attack (UKB codes 1075 in field 20002; code 1 in field 6150), self-reported coronary angioplasty (ptca) or coronary artery bypass graft (UKB codes 1070 and 1095 in field 20004), self-reported angina.</td> </tr> <tr> <td>Broad CAD<br> (includes angina and all ischaemic heart disease)</td> <td>As for Narrow CAD, plus ICD10 codes I24.X, I25X, and ICD9 codes 414.X (where .X indicates all subcodes).</td> </tr> </tbody> </table> <p>Note that there is no GWAS for cardiovascular disease (CVD) per se. This is because the UKB training data for the Enhanced CVD PRS consisted of separate GWASs for “narrow” CAD and ischaemic stroke.</p> <p>All analyses included Age at assessment, sex (for non-sex specific traits), genotyping chip, and 10 principal components as covariates.</p> <p>GWAS summary statistics for each trait were generated by applying PLINK 2.0 to the WBU subgroup, using a logistic regression for disease traits, and a linear regression model for quantitative traits. For chromosome X variants males were treated as having 0 or 2 alternative alleles.</p> <p>The results are not adjusted for genomic control.</p> <p><strong>DATA FILE CONTENT DESCRIPTION (DISEASE TRAITS)</strong></p> <table> <tbody> <tr> <td>cpra</td> <td>Variant ID in ‘CPRA’ format. Position reflects position in b37</td> </tr> <tr> <td>chrom</td> <td>Chromosome</td> </tr> <tr> <td>pos</td> <td>Position in base pairs (b37, 1-based)</td> </tr> <tr> <td>alt</td> <td>Alternative allele (effect allele)</td> </tr> <tr> <td>beta</td> <td>Effect size (log odds ratio)</td> </tr> <tr> <td>standard_error</td> <td>Standard error of beta</td> </tr> <tr> <td>minus_log10_p</td> <td>Minus log(base 10) of P-value</td> </tr> <tr> <td>ref</td> <td>Reference allele (non-effect allele)</td> </tr> <tr> <td>ncase</td> <td>Number of cases</td> </tr> <tr> <td>ncontrol</td> <td>Number of controls</td> </tr> </tbody> </table> <p><strong>DATA FILE CONTENT DESCRIPTION (QUANTITATIVE TRAITS)</strong></p> <table> <tbody> <tr> <td>cpra</td> <td>Variant ID in ‘CPRA’ format. Position reflects position in b37</td> </tr> <tr> <td>chrom</td> <td>Chromosome</td> </tr> <tr> <td>pos</td> <td>Position in base pairs (b37, 1-based)</td> </tr> <tr> <td>alt</td> <td>Alternative allele (effect allele)</td> </tr> <tr> <td>beta</td> <td>Effect size</td> </tr> <tr> <td>standard_error</td> <td>Standard error of beta</td> </tr> <tr> <td>minus_log10_p</td> <td>Minus log(base 10) of P-value</td> </tr> <tr> <td>ref</td> <td>Reference allele (non-effect allele)</td> </tr> <tr> <td>ntotal</td> <td>Total sample size</td> </tr> </tbody> </table> <p><strong>FILE NAMES</strong></p> <p>The following is a list of traits and their corresponding file names.</p> <p><em><strong>DISEASE TRAITS</strong></em></p> <table> <tbody> <tr> <td>Age-related macular degeneration</td> <td>amd_strict_UKB_WBU.csv.gz</td> </tr> <tr> <td>Alzheimer's disease</td> <td>alzheimers_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Asthma</td> <td>asthma_UKB_WBU.csv.gz</td> </tr> <tr> <td>Atrial fibrillation</td> <td>atrial_fibrillation_UKB_WBU.csv.gz</td> </tr> <tr> <td>Bipolar disorder</td> <td>bipolar_disorder_UKB_WBU.csv.gz</td> </tr> <tr> <td>Bowel cancer</td> <td>CRC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Breast cancer</td> <td>BC_UKB_WBU_women.csv.gz</td> </tr> <tr> <td>Coeliac disease</td> <td>celiac_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Narrow coronary artery disease</td> <td>NARROW_CAD_UKB_WBU.csv.gz</td> </tr> <tr> <td>Broad coronary artery disease</td> <td>BROAD_CAD_UKB_WBU.csv.gz</td> </tr> <tr> <td>Crohn's disease</td> <td>crohns_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Epithelial ovarian cancer</td> <td>OC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Hypertension</td> <td>HT_UKB_WBU.csv.gz</td> </tr> <tr> <td>Ischaemic stroke</td> <td>IS_stroke_UKB_WBU.csv.gz</td> </tr> <tr> <td>Melanoma</td> <td>melanoma_UKB_WBU.csv.gz</td> </tr> <tr> <td>Multiple sclerosis</td> <td>multiple_sclerosis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Osteoporosis</td> <td>OP_WBU_training.csv.gz</td> </tr> <tr> <td>Prostate cancer</td> <td>PC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Parkinson's disease</td> <td>parkinsons_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Primary open angle glaucoma</td> <td>POAG_WBU_training.csv.gz</td> </tr> <tr> <td>Psoriasis</td> <td>psoriasis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Rheumatoid arthritis</td> <td>rheumatoid_arthritis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Schizophrenia</td> <td>schizophrenia_UKB_WBU.csv.gz</td> </tr> <tr> <td>Systemic lupus erythematosus</td> <td>lupus_UKB_WBU.csv.gz</td> </tr> <tr> <td>Type 1 diabetes</td> <td>t1d_UKB_WBU.csv.gz</td> </tr> <tr> <td>Type 2 diabetes</td> <td>T2D_UKB_WBU.csv.gz</td> </tr> <tr> <td>Ulcerative colitis</td> <td>ulcerative_colitis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Venous thromboembolic disease</td> <td>VTE_UKB_WBU.csv.gz</td> </tr> </tbody> </table> <p><em><strong>QUANTITATIVE TRAITS</strong></em></p> <table> <tbody> <tr> <td>Age at menopause</td> <td>age_at_menopause_UKB_WBU.csv.gz</td> </tr> <tr> <td>Apolipoprotein A1</td> <td>apolipoprotein_a1_UKB_WBU.csv.gz</td> </tr> <tr> <td>Apolipoprotein B</td> <td>apolipoprotein_b_UKB_WBU.csv.gz</td> </tr> <tr> <td>Body mass index</td> <td>bmi_UKB_WBU.csv.gz</td> </tr> <tr> <td>Calcium</td> <td>calcium_UKB_WBU.csv.gz</td> </tr> <tr> <td>Docosahexaenoic acid</td> <td>docosahexaenoic_acid_UKB_WBU.csv.gz</td> </tr> <tr> <td>Estimated bone mineral density T-score</td> <td>BMD_WBU_training.csv.gz</td> </tr> <tr> <td>Estimated glomerular filtration rate (creatinine based)</td> <td>egfr_UKB_WBU.csv.gz</td> </tr> <tr> <td>Estimated glomerular filtration rate (cystatin based)</td> <td>egfr_cys_UKB_WBU.csv.gz</td> </tr> <tr> <td>Glycated haemoglobin</td> <td>hba1c_UKB_WBU_nodiabetes.csv.gz</td> </tr> <tr> <td>High density lipoprotein cholesterol</td> <td>hdl_cholesterol_UKB_WBU.csv.gz</td> </tr> <tr> <td>Height</td> <td>height_UKB_WBU.csv.gz</td> </tr> <tr> <td>Intraocular pressure</td> <td>iop_WBU_training.csv.gz</td> </tr> <tr> <td>Low density lipoprotein cholesterol</td> <td>ldl_UKB_WBU_nostatins.csv.gz</td> </tr> <tr> <td>Omega-6 fatty acids</td> <td>omega_6_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Omega-3 fatty acids</td> <td>omega_3_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Phosphatidylcholines</td> <td>phosphatidylcholines_UKB_WBU.csv.gz</td> </tr> <tr> <td>Phosphoglycerides</td> <td>phosphoglycerides_UKB_WBU.csv.gz</td> </tr> <tr> <td>Polyunsaturated fatty acids</td> <td>polyunsaturated_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Resting heart rate</td> <td>resting_heart_rate_UKB_WBU.csv.gz</td> </tr> <tr> <td>Remnant cholesterol (Non-HDL, Non-LDL cholesterol)</td> <td>remnant_cholesterol__UKB_WBU.csv.gz</td> </tr> <tr> <td>Sphingomyelins</td> <td>sphingomyelins_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total cholesterol</td> <td>total_cholesterol_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total fatty acids</td> <td>total_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total triglycerides</td> <td>total_triglycerides_UKB_WBU.csv.gz</td> </tr> </tbody> </table>
Philosphical papers reviewed and scored for their compliance with a vegan ethic 1975-2020
<p>This is the updated dataset to complement the submitted manuscript "<strong>Has the philosophical case for animal liberation been proved? A systematic and narrative review of the philosophical literature 1975-2020."</strong></p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
DEM Intercomparison eXercise (DEMIX) - Maps of completeness criteria scores for global DEMs
<h2>Introduction</h2> <p>This introduction gives a brief overview of the context in which the dataset has been produced. Readers curious about the detailed standards and procedures described in this section are encouraged to open the resources linked to this dataset.</p> <h3>The Digital Elevation Model Intercomparison eXercise (DEMIX)</h3> <p>This work is part of the Digital Elevation Model Intercomparison eXercise (DEMIX), initiated by the <a href="https://ceos.org/ourwork/workinggroups/wgcv/current-activites/#:~:text=DEMIX%3A%20Digital%20Elevation%20Model%20Intercomparison,elevation%20model%20for%20their%20application.">Committee on Earth Observation Satellites (CEOS)</a>. This initiative aims at "<a href="https://isprs-archives.copernicus.org/articles/XLIII-B4-2021/395/2021/">providing harmonised terminology and methods, as well as practical guidelines and results allowing the intercomparison of continental or global Digital Elevation Models (DEM)</a>" (Strobl et al., 2021). Several publications have defined the framework of DEMIX, from <a href="https://doi.org/10.3390/rs13183581">the terminology and definitions</a> (Guth et al., 2021) to the <a href="https://doi.org/10.1109/TGRS.2024.3368015">DEM ranking methods</a> (Bielski et al., 2024). An additional methodology paper has been publicated regarding the assessment of <a href="https://doi.org/10.3390/ijgi13030096">planimetric displacements between DEMs</a> (Riazanoff et al., 2024), which are a common source of biases in DEM comparisons.</p> <h3>The DEMIX grid</h3> <p>Studies performed within the DEMIX framework rely on the <a href="../records/7504791">DEMIX grid</a> (Guth et al., 2023), a geodetic grid (EPSG:4326) dividing the world in areas of approximately 10x10km. These standard areas are called DEMIX tiles, and can be precisely located thanks to their identifier.</p> <h3>Criteria and scores</h3> <p>Within DEMIX, several criteria have been defined to assess the quality of DEMs. These criteria take as input a DEM and a DEMIX tile, and provide as output the score of the DEM for this specific tile. Repeating this process over several DEMs and DEMIX tiles of interest allow for a comparison of scores, leading to a ranking of DEMs. <a href="https://doi.org/10.1109/TGRS.2024.3368015">DEMIX rankings are based on the Randomized Complete Block Design (RCBD)</a> (Bielski et al., 2024).</p> <h2>This dataset</h2> <p>This dataset is composed of global maps of one map per (DEM, criterion) tuple. Each GeoTIFF map can be superimposed with the <a href="../records/7504791">DEMIX grid</a> (Guth et al., 2023) in a GIS (tested in QGIS 3.16).</p> <h3>Completeness criteria</h3> <p>The completeness criteria have originally been defined by Peter Strobl. A brief description of each criterion is given in the next table. Please see the column "Original document" and files of this repository for the complete definitions.</p> <table> <tbody> <tr> <td><strong>Criterion</strong></td> <td><strong>Description</strong></td> <td><strong>Requirements</strong></td> <td><strong> Original document</strong></td> </tr> <tr> <td>A01 - Product fractional cover</td> <td>Fraction of a DEMIX tile <strong>covered</strong> by the DEM product</td> <td>None</td> <td>See document "DEMIX_CDD-A01_20211103.docx"</td> </tr> <tr> <td>A02 - Valid data fraction</td> <td>Fraction of a DEMIX tile <strong>covered</strong> by <strong>valid </strong>pixels of the DEM product</td> <td>"No data" or "void" value in metadata</td> <td>See document "DEMIX_CDD-A02_20211103.docx"</td> </tr> <tr> <td>A03 - Primary data fraction</td> <td>Fraction of a DEMIX tile <strong>covered </strong>by <strong>valid </strong>pixels generated from the <strong>main source of data</strong> of the DEM product</td> <td>"No data" or "void" value in metadata + source data/editing mask</td> <td>See document "DEMIX_CDD-A03_20211103.docx"</td> </tr> <tr> <td>A04 - Valid land fraction</td> <td>Fraction of a DEMIX tile <strong>covered </strong>by <strong>valid </strong>pixels of <strong>land </strong>of the DEM product</td> <td>"No data" or "void" value in metadata + water body mask</td> <td>See document "DEMIX_CDD-A04_20211103.docx"</td> </tr> <tr> <td>A05 - Primary land fraction</td> <td>Fraction of a DEMIX tile <strong>covered </strong>by <strong>valid </strong>pixels of <strong>land </strong>generated from the <strong>main source of data </strong>of the DEM product</td> <td>"No data" or "void" value in metadata + water body mask + source data/editing mask</td> <td>See document "DEMIX_CDD-A05_20211103.docx"</td> </tr> </tbody> </table> <h3>DEMs and ancillary data</h3> <p>The following DEM products and ancillary layers have been used to generate the dataset.</p> <table> <tbody> <tr> <td><strong>Identifier</strong></td> <td><strong>Used layers</strong></td> <td><strong>Data access</strong></td> </tr> <tr> <td> <p>ASTGTM v003</p> </td> <td>ASTER GDEM elevations (dem.tif) + editing / source masks (num.tif)</td> <td><a href="https://lpdaac.usgs.gov/products/astgtmv003/">https://lpdaac.usgs.gov/products/astgtmv003/</a></td> </tr> <tr> <td> <p>ASTWBD v001</p> </td> <td>ASTER GDEM water body mask (att.tif)</td> <td><a href="https://lpdaac.usgs.gov/products/astwbdv001/">https://lpdaac.usgs.gov/products/astwbdv001/</a></td> </tr> <tr> <td> <p>AW3D30 v2003</p> </td> <td>ALOS World 3D elevations (DSM.tif) + editing / source / water body masks (MSK.tif)</td> <td><a href="https://www.eorc.jaxa.jp/ALOS/en/dataset/aw3d30/aw3d30_e.htm">https://www.eorc.jaxa.jp/ALOS/en/dataset/aw3d30/aw3d30_e.htm</a></td> </tr> <tr> <td>COP-DEM_GLO-30-DGED v2019_1</td> <td>Copernicus DEM GLO-30 elevations (DEM.tif) + editing (EDM.tif) + source (SRC.tif) + water body (WBM.tif) masks</td> <td><a href="https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model">https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model</a></td> </tr> <tr> <td>COP-DEM_GLO-90-DGED v2019_1</td> <td>Copernicus DEM GLO-90 elevations (DEM.tif) + editing (EDM.tif) + source (SRC.tif) + water body (WBM.tif) masks</td> <td><a href="https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model">https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model</a></td> </tr> <tr> <td> <p>NASADEM_HGT v001</p> </td> <td>NASADEM elevations (.hgt) + editing / source (.num) + water body (.swb) masks</td> <td><a href="https://lpdaac.usgs.gov/products/nasadem_hgtv001/">https://lpdaac.usgs.gov/products/nasadem_hgtv001/</a></td> </tr> <tr> <td> <p>SRTMGL1 v003</p> </td> <td>SRTMGL1 elevations (.hgt)</td> <td><a href="https://lpdaac.usgs.gov/products/srtmgl1v003/">https://lpdaac.usgs.gov/products/srtmgl1v003/</a></td> </tr> <tr> <td> <p>SRTMGL1N v003</p> </td> <td>SRTMGL1 editing / source / water body masks (.num)</td> <td><a href="https://lpdaac.usgs.gov/products/srtmgl1nv003/">https://lpdaac.usgs.gov/products/srtmgl1nv003/</a></td> </tr> </tbody> </table> <h3>Computation of scores</h3> <p>For each DEMIX tile and DEM, each "fractional cover" has been computed using the following procedure:</p> <ol> <li><strong>Crop DEMIX tile layers</strong> - The tiles of each DEM layer (elevations, editing, sources and water bodies) are cropped to the extent of the DEMIX tile.</li> <li><strong>Compute standardized layers</strong><strong> </strong>- Given the cropped DEM layers, four standardized layers are produced, which are: <ul> <li>Heights layer - Containing the heights of the DEM</li> <li>Land/water mask layer - Indicating whether DEM pixels are land or water: <ul> <li>0 = NO_DATA</li> <li>1 = BACKGROUND</li> <li>2 = INVALID</li> <li>3 = WATER</li> <li>4 = LAND</li> </ul> </li> <li>Source mask layer - Indicating the source data of DEM heights (or "edited" value): <ul> <li>0 = NO_DATA</li> <li>1 = BACKGROUND</li> <li>2 = INVALID</li> <li>3 = PRIMARY_DATA</li> <li>4 = EXTERNAL_DATA</li> <li>5 = EDITED</li> </ul> </li> <li>Valid mask layer - Indicating if the DEM pixels are valid or not: <ul> <li>0 = NO_DATA</li> <li>1 = BACKGROUND</li> <li>2 = INVALID</li> <li>3 = VALID</li> </ul> </li> </ul> </li> <li><strong>Retrieve pixel number N</strong><em><strong> </strong>-<strong> </strong></em>The total pixel number N is computed for one of the layers (all layers have the same number of pixels).</li> <li><strong>Retrieve criterion pixel number C </strong>-<strong> </strong>The criterion pixel number C is computed based on the standard layers, more precisely: <ul> <li>A01 - Product fractional cover - Number of pixels of <strong>valid mask layer equal to 1, 2 or 3</strong></li> <li>A02 - Valid data fraction - Number of pixels of <strong>valid mask layer equal to 3</strong></li> <li>A03 - Primary data fraction - Number of pixels of <strong>source mask layer equal to 3</strong></li> <li>A04 - Valid land fraction - Number of pixels of <strong>land/water mask layer equal to 4</strong></li> <li>A05 - Primary land fraction - Number of pixels of <strong>source mask layer equal to 3</strong> and<strong> land/water mask layer equal to 4</strong></li> </ul> </li> <li><strong>Compute the final score S</strong><strong> </strong>- The final score S is expressed as the following percentage: <strong>S = ceil(C/N*100)</strong></li> </ol> <h2>Known issues</h2> <p>The "SRTMGL1N v003" is known to have "tile repeating issues", where part of the data is wrongly flagged as water. This issue has been reported with no particular response from the providers of the DEM (see <a href="https://forum.earthdata.nasa.gov/viewtopic.php?t=2752">https://forum.earthdata.nasa.gov/viewtopic.php?t=2752</a>).</p> <p><strong>References:</strong></p> <ul> <li>Guth, P.L.; Strobl, P.; Gross, K.; Riazanoff, S. <em>DEMIX 10k Tile Data Set (1.0)</em> [Data set]. Zenodo 2023. <a href="https://doi.org/10.5281/zenodo.7504791" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.7504791</a></li> <li>Guth, P.L.; Van Niekerk, A.; Grohmann, C.H.; Muller, J.-P.; Hawker, L.; Florinsky, I.V.; Gesch, D.; Reuter, H.I.; Herrera-Cruz, V.; Riazanoff, S.; López-Vázquez, C.; Carabajal, C.C.; Albinet, C.; Strobl, P. <em>Digital Elevation Models: Terminology and Definitions</em>. Remote Sens. 2021, 13, 3581. <a href="https://doi.org/10.3390/rs13183581">https://doi.org/10.3390/rs13183581</a></li> <li>Riazanoff, S.; Corseaux, A.; Albinet, C.; Strobl, P.A.; López-Vázquez, C.; Guth, P.L.; Tadono, T. <em>Best BiCubic Method to Compute the Planimetric Misregistration between Images with Sub-Pixel Accuracy: Application to Digital Elevation Models</em>. <em>ISPRS Int. J. Geo-Inf.</em> 2024, <em>13</em>, 96. <a href="https://doi.org/10.3390/ijgi13030096">https://doi.org/10.3390/ijgi13030096</a></li> <li>Bielski, C.; López-Vázquez, C.; Grohmann, C.H.; Guth, P.L.; Hawker, L.; Gesch, D.; Trevisani, S.; Herrera-Cruz, V.; Riazanoff, S.; Corseaux, A.; Reuter, H.I.; Strobl, P.A.; <em>Novel Approach for Ranking DEMs: Copernicus DEM Improves One Arc Second Open Global Topography</em> in <em>IEEE Transactions on Geoscience and Remote Sensing</em>, vol. 62, pp. 1-22, 2024, Art no. 4503922. <a href="https://doi.org/10.1109/TGRS.2024.3368015">https://doi.org/10.1109/TGRS.2024.3368015</a></li> <li>Strobl, P.A.; Bielski, C.; Guth, P.L.; Grohmann, C.H.; Muller, J.P.; López-Vázquez, C.; Gesch, D.B.; Amatulli, G.; Riazanoff, S.; Carabajal, C. The Digital Elevation Model Intercomparison eXperiment DEMIX, a community based approach at global DEM benchmarking. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2021, XLIII-B4-2021, 395–400. <a href="https://doi.org/10.5194/isprs-archives-XLIII-B4-2021-395-2021">https://doi.org/10.5194/isprs-archives-XLIII-B4-2021-395-2021</a></li> </ul>
Air Quality Index Scores by CBSA with Population
<p>The AQI describes the five main types of air pollution regulated by the Clean Air Act: sulfur dioxide, nitrogen dioxide, carbon monoxide, ground-level ozone, and particle pollution. The EPA and its partners take regular readings of these pollutants and converts the results into a number ranging from 0 to 500, along with a specific color corresponding to a level of health concern. Generally, if the air quality is good, the air quality index is low (0 to 50) or moderate (51-100), and the color associated with it is green or yellow. As the air quality gets worse, the numbers go up, and the color linked with it goes from orange, to red, to purple, all the way to a dark shade of maroon for hazardous (300+).</p> <p>This dataset contains the AQI scores by metropolitant area (CBSA) during 2017. I've enhanced some publically available data from the EPA's <a href="https://airnow.gov/">airnow</a> website with census data, to be able to provide context about the number of people who are actually impacted when an <a href="https://www.cleanairresources.com/resources/how-do-i-read-the-air-quality-index">AQI score</a> is high or low in a given area.</p> <p>Related datasets on <a href="https://www.cleanairresources.com/data#sources">relative composition of air pollution by source: typical distribution and during wildfire season</a> are available here.</p>
GenoNet scores for human genome assembly GRCh37
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
Integration of expression datasets to identify biomarkers for accurate Gleason scoring in Prostate Cancer -- Supplementary data
<p>This dataset contains expression data from multiple sources used to identify biomarker candidates for prostate cancer aggressiveness. The data includes transcriptional expression levels, patient metadata, and other relevant features utilized in our machine-learning models. The training dataset was extracted from the repository described at Matos-Filipe, et al. (2022) [1].</p> <p>Raw ML metrics from models gaussian Naïve Bayes classifiers using this dataset are available in </p> <p> </p> <p>[1] <span><span><span>Matos-Filipe, P, et al. "</span></span></span>The usage of transcriptomics datasets as sources of Real-World Data for clinical trialling". <span>bioRxiv (</span><span>2022). </span><span><span>doi:</span> https://doi.org/10.1101/2022.11.10.515995</span></p>
Raw data for analysis of archaeological collections using scoring
<p>The file includes data in the form of attributes identifying the cultural or natural origin of archaeological collections composed of flakes and blades made of flint. The collections attributed to the Lower and Middle Palaeolithic in Poland and Germany. The last collection consists of experimentally produced samples. These collections are held at the Muzeum Śląska Opolskiego in Opole, University of Wrocław, University of Silesia in Katowice, Sosnowiec and Landesmuseum für Vorgeschichte in Halle (Saale). The collected data was used to improve the method hitherto applied to distinguish sets composed of artefacts and pseudo-artefacts using so-called scoring. This work was financially supported by the National Science Centre,<br>Poland (Grant 2020/39/B/HS3/02277).</p>
The Brief Symptom Inventory in the Swiss general population: Presentation of norm scores and predictors of psychological distress: Data supporting the publication
This is the dataset on which the following publication is based: • Michel G, Baenziger J, Brodbeck J, Mader L, Kuehni CE, Roser K (2024). The Brief Symptom Inventory in the Swiss general population: Presentation of norm scores and predictors of psychological distress. PLOS One. 19(7), e0305192. Doi: 10.1371/journal.pone.0305192, https://doi.org/10.1371/journal.pone.0305192 A description of the sample and the data collection procedure is available in the publication. The dataset contains the following variables: • Socio-demographic characteristics of the sample - Weights according to representative general population sample - Sex from Swiss Federal Statistical Office (SFSO) - Age at study (rounded to integer) - Age categories (10-year age groups) - Language questionnaire (German/Rumantsch, French, Italian) - Nationality from SFSO - Migration background - Education - Employment status • Original and prepared data on the Brief Symptom Inventory A detailed data dictionary is available in a separate excel file. Version • 1.0 (15 August 2024)
Shakespeare: Julius Caesar 1.2.30-187, sentences with Maximum Similarty Score. Appendix Table for "Innovation and Repetiton in Dramatic Texts"
<p>This is an illustrative table for the study "Innovation and Repetition in Dramatic Texts", published in the journal JCLS by Botond Szemes and Mihály Nagy. The table contains a scene from <em>Julius Caesar</em> 1.2.30-187' , assigning to each utterance the most similar sentence and the degree of similarity based on an S-BERT model.</p>
Meta analysis of prognostic scoring systems for pancreatitis
Open the record for dataset details and reuse information.
Consistency test scores for aftershock+mainshock RELM forecasts
<p><strong>Summary</strong></p> <p>Files are transcribed from Zechar et al. (2013) into comma separated values (csv) files. The consistency test scores are shown in the electronic supplement table S4 and the catalog is found in Table 1 of the main text.</p> <p><strong>Reference</strong></p> <p>Zechar, J. D., D. Schorlemmer, M. J. Werner, M. C. Gerstenberger, D. A. Rhoades, and T. H. Jordan (2013). Regional Earthquake Likelihood Models I: First-Order Results, Bulletin of the Seismological Society of America 103 787-798.</p> <p> </p>
Partitioned linkage disequilibrium scores for active regulatory elements in ROADMAP datasets
<p>Partitioned linkage disequilibrium scores for active regulatory elements in ROADMAP epigenomics datasets, to accompany paper Lynall et al 2021</p> <p>Accompanying code available at https://github.com/maryellenlynall/psychimmgen2021</p> <p>Active regulatory elements annotations are a union of the following IDEAS annotations, representing enhancers and active promoters (see http://bx.psu.edu/~yuzhang/Roadmap_ideas/trackDb_test.txt for IDEAS track hubs): </p> <p>4_Enh<br> 6_EnhG<br> 8_TssAFlnk<br> 10_TssA<br> 14_TssWk<br> 17_EnhGA</p> <p>tissues.txt provides the list of ROADMAP tissues </p> <p>The partitioned_LD_scores folder contains partitioned LD scores in a format suitable for stratified LDSC analysis for European participants</p>
Scores for calculating automated FAIR assessments in the low carbon energy domain
<p>Results for an automated FAIR assessment of 80 databases from the low carbon energy domain. The assessment was performed with the help of the FAIR maturity evaluation service of Wilkinson et al. The FAIR status with respect to 16 FAIR criteria is listed. The scores are defined to be consistent with the FAIR assessment tool of the Australian Research Data Commons. More details can be found in an additional publication on Zenodo as well as in an upcoming publication by Schwanitz et al.</p>
SCoRe- Prototyp 2 - Erprobung des Forschungsszenarios "Urbane Grünflächen" – UGF-1
<p>Dieses Datenset enthält Materialien (Videos, Protokolle und Fallbeschreibungen) aus der ersten prototypischen Durchführung des Forschungsszenarios "Urbane Grünflächen" im Teilprojekt <a href="http://www.360total.de/score/">SCoRe-VideoLearning</a> des <a href="https://scoreforschung.com/ueber/">Score-Projektes</a> ..</p> <p>Hierin finden sich drei exemplarische Fälle von Studierenden, welche sich videografisch forschend mit urbanen Grünflächen auseinandersetzten und dabei die Merkmale der Grünfläche hinsichtlich urbanen Nutzungsmöglichkeiten und der biologischen Vielfalt untersuchten. Dazu wurden die Grünflächen zunächst ausgewählt und in Bezug auf verschiedene vorgegebene Ordnungskriterien beschrieben und bewertet (Fallbeschreibung). Zur Produktion der Videoforschungsdaten - als Basismaterial der empirischen Untersuchung - waren die Studierenden angehalten ein Produktionsprotokoll während aller drei Produktionsphasen der Videografie (Vorproduktion, Produktion im Feld sowie Nachproduktion) auszufüllen und somit für sich sowie andere analysierende Studierende die Entscheidungsprozesse zur Gestaltung der Videoforschungsdaten zu explizieren und zu dokumentieren.. Diese Protokolle bilden entsprechend die Grundlagen für Gütekriterien qualitativer Forschungsdaten: Transparenz und intersubjektive Nachvollziehbarkeit. </p>
Metrics As Scores Dataset: Price, Weight, and Other Properties of Over 1,200 Ideal-Cut and Best-Clarity Diamonds
<p>This dataset is a subset of the original diamonds dataset with more than 54,000 diamonds. It was reduced to only contain diamonds of the best cut (ideal) and clarity (IF). The group is now given by the colors from J (worst) to D (best). This dataset comes from the R-package ggplot2 (Wickham 2016). For each color, we can examine the following attributes (<strong>features</strong>) of each diamond:</p> <ul> <li><em>Carat</em>: Weight of the diamond</li> <li><em>Depth</em>: Total depth percentage</li> <li><em>Price</em>: Price in US dollars [discrete]</li> <li><em>Table</em>: Width of top of diamond relative to widest point</li> <li><em>X</em>: Length in mm</li> <li><em>Y</em>: Width in mm</li> <li><em>Z</em>: Depth in mm</li> </ul> <p>It has a total of 7 Colors (<strong>groups</strong>): <em>D</em>, <em>E</em>, <em>F</em>, <em>G</em>, <em>H</em>, <em>I</em>, and <em>J</em>. The best color is <em>D</em> and the worst color is <em>J</em>. This dataset was created to analyze whether there are differences between the colors.</p>
Metrics As Scores Dataset: Elisa Spectrophotometer Positive Samples
<p>The ELISA dataset contains data from a spectrophotometer that determined the optical density of positive control samples, from five different lots, across five different runs. This dataset was introduced by (Schmid 1991).</p> <p>This dataset has the following <strong>Features</strong>:</p> <ul> <li><em>Lot1</em>: The first lot</li> <li><em>Lot2</em>: The second lot</li> <li><em>Lot3</em>: The third lot</li> <li><em>Lot4</em>: The fourth lot</li> <li><em>Lot5</em>: The fifth lot</li> </ul> <p>It has a total of 5 <strong>Groups</strong>: <em>Run1</em>, <em>Run2</em>, <em>Run3</em>, <em>Run4</em>, and <em>Run5</em>.</p>
Metrics As Scores Dataset: The Iris Flower Data Set
<p>The Iris flower data set or Fisher’s Iris data set is a multivariate data set used and made famous by the British statistician and biologist Ronald Fisher. The dataset was introduced in his 1936 paper "The Use of Multiple Measurements in Taxonomic Problems" (Fisher 1936) as an example of linear discriminant analysis.</p> <p>This dataset has the following Features:</p> <ul> <li><em>Petal.Length</em>: Length of the petal</li> <li><em>Petal.Width</em>: Width of the petal</li> <li><em>Sepal.Length</em>: Length of the sepal</li> <li><em>Sepal.Width</em>: Width of the sepal</li> </ul> <p>It has a total of 3 <strong>Groups</strong>: <em>setosa</em>, <em>versicolor</em>, and <em>virginica</em>.</p>
Metrics As Scores Dataset: Metrics and Domains From the Qualitas.class Corpus
<p>This dataset was created by extracting software metrics data from the Qualitas.class corpus (Terra et al. 2013; Tempero et al. 2010). Therefore, the principal quantity type is Metric, and the context is given by a system’s Domain (e.g., "Game", "Middleware", etc.). Some metrics were obtained on program-level, while others are package- or method-level metrics. Most of the metrics in the corpus are of discrete/integral nature. The corpus holds 23 types of pre-computed software metrics for a total of 111 systems which are spread across eleven different domains.</p> <p>This dataset has the following discrete <strong>Features</strong> (Metrics):</p> <ul> <li>CA: Afferent Coupling</li> <li>CE: Efferent Coupling</li> <li>DIT: Depth of Inheritance Tree</li> <li>MLOC: Method Lines of Code</li> <li>NBD: Nested Block Depth</li> <li>NOC: Number of Classes</li> <li>NOF: Number of Attributes</li> <li>NOI: Number of Interfaces</li> <li>NOM: Number of Methods</li> <li>NOP: Number of Packages</li> <li>NORM: Number of Overridden Methods</li> <li>NSC: Number of Children</li> <li>NSF: Number of Static Attributes</li> <li>NSM: Number of Static Methods</li> <li>PAR: Number of Parameters</li> <li>TLOC: Total Lines of Code</li> <li>VG: McCabe Cyclomatic Complexity</li> <li>WMC: Weighted Methods per Class</li> </ul> <p>The following features are continuous:</p> <ul> <li>LCOM: Lack of Cohesion in Methods</li> <li>RMA: Abstractness</li> <li>RMD: Normalized Distance</li> <li>RMI: Instability</li> <li>SIX: Specialization Index</li> </ul> <p>It has a total of 11 <strong>Groups</strong> (Domains): 3D; Graphics; Media, Databases, Diagrams; Visualiz., Games, IDE, Middleware, Parsers; Generators, Progr. Language, SDK, Testing, and Tool.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.