Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.9.0
Dataset results
18 results for “quantile”
A Pan-European, Quantile Machine learning (QML) based, Total, Fine-Mode and Coarse-Mode Aerosol Optical Depth dataset (QML AOD))
<p>The V 1.1.0 product is an improved Aerosol Optical Depth (AOD) product based on Gap-filled MAIAC AOD, which provide first full-coverage, high-resolution monitoring of fine-mode and coarse-mode aerosols in Europe from 2003-20. This dataset has successfully rectified the previously identified issue of weak associations between satellite AOD and PM2.5 in Europe, which was primarily attributable to current limitations of AOD data. Our innovative approach has yielded stronger correlations with PM10, PM2.5, and PMcoarse than previous AOD product, laying a critical groundwork for improving PM10, PM2.5, and PMcoarse predictions in further epidemiological studies or environmental monitoring.</p> <p>We have uploaded three QML AOD datasets in Geotiff format, covering the region from -27° to 72° latitude and from -25° to 45° longitude. These datasets will be useful for researchers and policymakers to better understand the impacts of aerosols on the environment and human health.</p> <p> Note: v1.0.0 product do not include MAIAC AOD in their models.</p> <p>Please read more details in our paper </p> <h1><span>Estimation of pan-European, daily total, fine-mode and coarse-mode Aerosol Optical Depth at 0.1° resolution to facilitate air quality assessments</span></h1> <p><a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.scitotenv.2024.170593" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.scitotenv.2024.170593</span></a></p>
Code and Data to "Quantile regression for temporal streamflow modeling"
<p>This is the accompanying code to "Quantile regression for temporal streamflow modeling", which is part of the manuscript "The Role of Process Heterogeneity in Statistical Modeling", which was submitted to the Austrian Journal of Statistics. </p> <p> </p> <p>The data used in this publication is fully accessible through the <a href="https://doi.org/10.5194/essd-13-4529-2021" target="_blank" rel="noopener">LamaH-CE</a> dataset. The two scripts "functions_create_data.R" and "create_data.R" will create the final dataset used for modelling. </p> <p>"functions_modelling.R" provide the functions for tuning the XGBoost model and computing the SHAP values. An example script is also attached (calc_predictions_shap.R). "analyzing_results.R" and "error_metrics.R" will produce the final output used in the manuscript. Finally, two plots produced in the script are added as pdf. </p> <p>All data analysis was performed in R, and we want to acknowledge the following packages: <a href="https://dplyr.tidyverse.org/">dplyr</a>, <a href="https://tidyr.tidyverse.org/">tidyr</a>, <a href="https://www.jstatsoft.org/v40/i03/">lubridate</a>, <a href="https://purrr.tidyverse.org/">purrr</a>, <a href="https://doi.org/10.18637/jss.v033.i01">glmnet</a>, <a href="https://cran.r-project.org/web/packages/xgboost/index.html">xgboost</a>, <a href="https://CRAN.R-project.org/package=shapr">shapr</a>, <a href="https://CRAN.R-project.org/package=Metrics">Metrics</a>, <a href="https://CRAN.R-project.org/package=gridExtra" target="_blank" rel="noopener">gridExtra</a>, <a href="https://doi.org/10.18637/jss.v014.i06">zoo</a> and <a href="https://CRAN.R-project.org/package=wesanderson">wesanderson</a>. </p> <p> </p>
Data and code for "Uncertainty Displays Using Quantile Dotplots or CDFs Improve Transit Decision-Making" (CHI 2017)
<p>This repository contains data and analysis code for the following paper:</p> <p>Michael Fernandes, Logan Walls, Sean Munson, Jessica Hullman, and Matthew Kay. "Uncertainty Displays Using Quantile Dotplots or CDFs Improve Transit Decision-Making", Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems - CHI 2018. DOI: 10.1145/3173574.3173718<br> </p>
UKCP18 RCM precipitation and temperature bias corrected using ISIMIP3BA change-preserving quantile mapping.
<p>We present bias-corrected UK Climate Projections 2018 (UKCP18; Met Office Hadley Centre, 2018) regional datasets for temperature, precipitation, and potential evapotranspiration (1981-2080). All 12 members of the 12 km ensemble were corrected using quantile mapping and a change-preserving variant (Lange, 2019; Lange, 2020). Both methods effectively reduce biases in multiple statistics, while maintaining projected climatic changes. We provide guidance on using the bias-corrected datasets for climate change impact assessment. Please find a detailed description and evaluation in the metadata and accompanying data paper (Reyniers et al., 2025).</p> <p>---</p> <p>Met Office Hadley Centre (2018): UKCP18 Regional Projections on a 12km grid over the UK for 1980-2080. CEDA, <em>8 March 2022</em>. <a href="https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604">https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604</a></p> <p>Lange, S. (2019). Trend-preserving bias adjustment and statistical downscaling with ISIMIP3BASD (v1. 0). <em>GMD,</em> <em>12</em>(7), 3055-3070.</p> <p>Lange, S. (2020). ISIMIP3BASD (2.4.1). Zenodo. https://doi.org/10.5281/zenodo.3898426</p> <p>Reyniers, N., Zha, Q., Addor, N., Osborn, T. J., Forstenhäusler, N., & He, Y. (2025). Two sets of bias-corrected regional UK Climate Projections 2018 (UKCP18) of temperature, precipitation and potential evapotranspiration for Great Britain. <em>Earth System Science Data</em>, <em>2025, 17(5)</em>, 2113–2133.</p>
Data from: On the use of double quantile regression and visual assessment to estimate performance constraints
Open the record for dataset details and reuse information.
Quantile regression in genomic selection for oligogenic traits in autogamous plants: a simulation study
<p>This study assessed the efficiency of Genomic selection (GS) or genome‐wide selection (GWS), based on Regularized Quantile Regression (RQR), in the selection of genotypes to breed autogamous plant populations with oligogenic traits. To this end, simulated data of an F<sub>2</sub> population were used, with traits with different heritability levels (0.10, 0.20 and 0.40), controlled by four genes. The generations were advanced (up to F<sub>6</sub>) at two selection intensities (10% and 20%). The genomic genetic value was computed by RQR for different quantiles (0.10,0.50 and 0.90), and by the traditional GWS methods, specifically RR-BLUP and BLASSO. A second objective was to find the statistical methodology that allows the fastest fixation of favorable alleles. In general, the results of the RQR model were better than or equal to those of traditional GWS methodologies, achieving the fixation of favorable alleles in most of the evaluated scenarios. At a heritability level of 0.40 and a selection intensity of 10%, RQR (0.50) was the only methodology that fixed the alleles quickly, i.e., in the fourth generation. Thus, it was concluded that the application of RQR in plant breeding, to simulated autogamous plant populations with oligogenic traits, could reduce time and consequently costs, due to the reduction of selfing generations to fix alleles in the evaluated scenarios.</p>
Gaussian Quantiles Datasets - DiRo2C
<p>The two datasets are used to simulate two different black boxes that are supposed to predict different results in certain data areas.</p> <p>Dataset diro2c_gaussian_dataset.csv:</p> <p>Two-dimensional dataset with the continuous features x1 and x2. It was created with µ = 0 and σ 2 = 0.8. It contains 300 instances (data points) with the following properties for feature x1: min = − 275.71, max = 255.90, µ = 0.04, and σ = 88.54 and with the following properties for feature x2: min = − 252.57, max = 201.03, µ = − 9.44, and σ = 86.20. The instances of the datasets are classified into two classes “0” and “1”. 150 are assigned to the class “0” and 150 instances are assigned to the class “1”. The instances of the dataset are generated by the sklearn “make_gaussian_quantiles” function with the following parameters: make_gaussian_quantiles(n_samples = 300, n_classes = 2, shuffle = False, cov = 0.8, random_state = 7). Afterward, we scale the instances by the factor 100.</p> <p>Dataset diro2c_gaussian_manipulated_dataset.csv:</p> <p>Two-dimensional dataset with the continuous features x1 and x2. The manipulated dataset is generated with µ = 0 and σ 2 = 1.3. It contains 300 instances (datapoints) with the following properties for feature x1: min = − 351.46, max = 326.21, µ = 0.04, and σ = 112.87 and with the following properties for feature x2: min = − 321.97, max = 256.27, µ = − 12.04, and σ = 109.89. The instances of the datasets are classified into two classes “0” and “1”. 150 instances are assigned to the class “0”, and 150 instances are assigned to the class “1”.</p>
Review data for: SnowQM 1.0: A fast R Package for bias-correcting spatial fields of snow water equivalent using quantile mapping
<p>Climatology of snow water equivalent of Switzerland between winters 1962 and 2021. Obtained using quartile mapping between a model using data assimilation since 1998 and a model without data assimilation. This version of the dataset corresponds to the publication revision time. The publication has been submitted to GMD Copernicus journal as: <em>SnowQM 1.0: A fast R Package for bias-correcting spatial fields of snow water equivalent using quantile mapping</em></p>
UKCP18 RCM precipitation and temperature bias corrected using non-parametric quantile mapping method
<p>The UKCP18 RCM PPE (Met Office Hadley Centre, 2018) projections of precipitation and daily average temperature were bias adjusted using a non-parametric quantile mapping method based on empirical quantiles (Boe et al, 2007, Gudmundsson et al, 2012). The datasets cover the period from December 1980 to November 2080 and are intended for use in climate change impact assessments, where the bias correction helps reduce biases in multiple statistics while <span>maintaining projected climatic changes</span>.</p> <p>-------------------------------------------------</p> <p>Met Office Hadley Centre (2018): UKCP18 Regional Projections on a 12km grid over the UK for 1980-2080. CEDA, <em>8 March 2022</em>. <a href="https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604">https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604</a></p> <p>Boe, J.; Terray, L.; Habets, F. & Martin, E. Statistical and dynamical downscaling of the Seine basin climate for hydro-meteorological studies. International Journal of Climatology, 2007, 27, 1643-1655, doi: 10.1002/joc.1602.</p> <p>Gudmundsson, L.; Bremnes, J. B.; Haugen, J. E. & Engen-Skaugen, T. Technical Note: Downscaling RCM precipitation to the station scale using statistical transformations - a comparison of methods. Hydrology and Earth System Sciences, 2012, 16, 3383-3390, doi:10.5194/hess-16-3383-2012.</p> <p><strong>Paper Citation:</strong><br>We kindly ask users of this dataset to cite the paper that describes the dataset. The paper is published and can be accessed via the following link: <a href="https://doi.org/10.5194/essd-17-2113-2025" target="_new" rel="noopener">https://doi.org/10.5194/essd-17-2113-2025</a>.</p> <p><br>Please reference the paper as:<br>Reyniers, N., Zha, Q., Addor, N., Osborn, T. J., Forstenhäusler, N., and He, Y.: Two sets of bias-corrected regional UK Climate Projections 2018 (UKCP18) of temperature, precipitation and potential evapotranspiration for Great Britain, Earth Syst. Sci. Data, 17, 2113–2133, https://doi.org/10.5194/essd-17-2113-2025, 2025.</p>
Data from: Managing more than the mean: using quantile regression to identify factors related to large elk groups
1. Animal group size distributions are often right-skewed, whereby most groups are small, but most individuals occur in larger groups that may also disproportionately affect ecology and policy. In this case, examining covariates associated with upper quantiles of the group size distribution could facilitate better understanding and management of large animal groups. 2. We studied wintering elk groups in Wyoming, where group sizes span several orders of magnitude, and issues of disease, predation and property damage are affected by larger group sizes. We used quantile regression to evaluate relationships between the group size distribution and variables of land use, habitat, elk density and wolf abundance to identify conditions important to larger elk groups. 3. We recorded 1263 groups ranging from 1 to 1952 elk and found that across all quantiles of group size, group sizes were larger in open habitat and on private land, but the largest effect occurred between irrigated and non-irrigated land [e.g. the 90th quantile group size increased by 135 elk (95% CI = 42, 227) on irrigation]. 4. Only upper quantile group sizes were positively related to broad-scale measures of elk density and wolf abundance. For wolf abundance, this effect was greater on elk groups found in open habitats and private land than those in closed habitats or public land. If we had limited our analysis to mean or median group sizes, we would not have detected these effects. 5. Synthesis and applications. Our analysis of elk group size distributions using quantile regression suggests that private land, irrigation, open habitat, elk density and wolf abundance can affect large elk group sizes. Thus, to manage larger groups by removal or dispersal of individuals, we recommend incentivizing hunting on private land (particularly if irrigated) during the regular and late hunting seasons, promoting tolerance of wolves on private land (if elk aggregate in these areas to avoid wolves) and creating more winter range and varied habitats. Relationships to the variables of interest also differed by quantile, highlighting the importance of using quantile regression to examine response variables more completely to uncover relationships important to conservation and management.
Does Investor Sentiment Predict Bitcoin Return and Volatility? - A Quantile Regression Approach
<p>This dataset was used in generating findings for the paper titled "<strong>Does Investor Sentiment Predict Bitcoin Return and Volatility? - A Quantile Regression Approach".</strong></p>
Ice water path retrievals from Meteosat-9 with quantile regression neural networks: video supplement
<p>Supplementary videos used from in A. Amell, P. Eriksson, S. Pfreundschuh: Ice water path retrievals form Meteosat-9 with quantile regression neural networks.</p>
Joint quantile disease mapping with application to Malaria and G6PD deficiency
Statistical analysis based on quantile regression methods is more comprehensive, flexible, and less sensitive to outliers when compared to mean regression methods. When the link between different diseases are of interest, joint disease mapping is useful for inferring correlation between them. Most studies study this link through multiple correlated mean regressions. In this paper we propose a joint quantile regression framework for multiple diseases where different quantile levels can be considered. We are motivated by the theorized link between the presence of Malaria and the gene deficiency G6PD, where medical scientist have anecdotally discovered a possible link between high levels of G6PD and lower than expected levels of Malaria initially pointing towards the occurrence of G6PD inhibiting the occurrence of Malaria. This link cannot be investigated with mean regressions and thus the need for flexible joint quantile regression in a disease mapping framework arise. Our joint quantile disease mapping model can be used for linear and non-linear effects of covariates by stochastic splines, since we define it as a latent Gaussian model. We perform Bayesian inference of this model using the INLA framework embedded in the R software package INLA, resulting in a very efficient model even for large datasets. Finally, we illustrate the applicability of the model by analyzing the malaria and G6PD deficiency incidences, jointly, in 21 countries.
The Quantile-Quantile Plots of Analyzed Variables in a Neuroscience Study
<p>Enclosed with this submission are the Quantile-Quantile Plots pertinent to the variables involved in our study. These plots serve as graphical evaluation methods for assessing the normality of our variables. </p>
Data from: Managing more than the mean: using quantile regression to identify factors related to large elk groups
Open the record for dataset details and reuse information.
Joint quantile disease mapping with application to Malaria and G6PD deficiency
Open the record for dataset details and reuse information.
Dataset for the paper 'Debris-flow return times and pre-event terrains: a novel framework using copula functions and quantile-based approaches'
<p>here attached you can find the dataset we used for our paper. While in the first two columns, you have X, and Y coordinates of the domain's pixels, the other five columns are DEMs' elevations (2006, 2010, 2011, 2015, and 2021). See in the paper's 'Methodology' Section how to employ these data. For any questions, please reach us at massimiliano.schiavo@unipd.it. Thanks!</p>
Does urbanization drive up housing prices? Novel evidence from remote sensing and dynamic panel quantile regression
<p><strong>Purpose:</strong> This study aims to quantify the influence of urbanization on housing prices at the districtbased level, while also investigating the heterogeneous impacts across different quantiles of housing prices.</p> <p><br><strong>Design/methodology/approach:</strong> The study uses remote-sensed spectral images from the Landsat 7 ETM+ satellite to measure urbanization, replacing prior reliance solely on urban population metrics. Subsequently, the two-step system Generalized Method of Moments is employed to evaluate how urbanization influences district-based housing prices through three spectrometrics: Urban Index (𝑈𝐼), Normalized Difference Built-up Index (𝑁𝐷𝐵𝐼), and Built-Up Index (𝐵𝑈𝐼). Finally, this study examines the heterogeneous impacts across various housing price quantiles through Dynamic Panel Quantile Regression with non-additive fixed effects under Markov Chain Monte Carlo Simulation.</p> <p><br><strong>Findings:</strong> The study demonstrates that urbanization leads to an increase in regional housing prices. However, these impact magnitudes vary across housing price quantiles. Specifically, the impact exhibits an inverse V-shaped curve, with urbanization exerting a more pronounced influence on the 60𝑡ℎ percentile of housing prices, while its effect on the 10𝑡ℎ and 90𝑡ℎ percentile is comparatively weaker.</p> <p><br><strong>Originality/value:</strong> This study employs a novel method of utilizing remote sensing to measure urbanization and investigates its effects on housing prices. Furthermore, it provides an empirical application of non-additive fixed effect quantile regression for analyzing heterogeneity.<br><br></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.