Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,243
datasets available to search
ShareScore release 0.7.1
Dataset results
1,243 results for “statistics”
Quantifying changes in fish population stability using statistical early warnings of regime shifts
This data package describes long-term trends in metrics describing population stability and used as statistical early warnings of regime shifts in 29 fish species that inhabit the San Francisco Bay-Delta in central California, USA. Metrics used in this study include spatial synchrony, temporal coefficient of variation (CV), and lag-1 temporal autocorrelation. Trends were measured using ordinary least squares linear regression. These derived data were developed from abundance (as CPUE) time series based on three long-term fish monitoring studies included in https://doi.org/10.6073/pasta/a29a6e674b0f8797e13fbc4b08b92e5b; the Fall Midwater Trawl Survey, Delta Juvenile Monitoring Program, and Bay Study. Selected data were from fall months (September to December) in 1980-2023, from midwater trawl and beach seine surveys for which sampling effort (e.g., tow volume) was recorded. Data on fish exceeding maximum length thresholds for age-0 fish were discarded, except for white sturgeon, where the maximum length threshold corresponded to approximately 10 years of age, the onset of reproductive maturity. Observations from different sampling stations were aggregated into 10 sub-regions (South San Francisco Bay, Central San Francisco Bay, San Pablo Bay, Napa River, Suisun Bay, Delta Confluence, South Delta, North Delta, San Joaquin River, Sacramento River, and midwater trawl samples and beach seine samples were considered separately because the methods sample distinct habitat types. Combinations of sub-region and sampling method were considered distinct spatial units. EWI metrics were measured in 5-year rolling windows to permit assessment of changes over time. The temporal CV and lag-1 autocorrelation were measured on individual spatial unit time series, ignoring windows with >1 year of missing data. The coefficient of variation divides the standard deviation by the mean. Lag-1 autocorrelation was measured as Pearson correlation. Spatial synchrony was measured across spatial un
OLVSL_ Object-location visual statistical learning
Open the record for dataset details and reuse information.
ICARIA: climate projections from statistical downscaling outputs
<p><strong>ICARIA </strong>project had as one of its main purposes to develop coherent, reliable and usable downscaled climate projections from the last CMIP6 in order to construct the basis for efficient support to climate adaptation and decision-making of the related stakeholders, supporting the adaptation of critical assets within the project. These projections were obtained with also the purpose to be freely available for further use in subsequent studies and, hence, foster adaptation to climate change in more areas. Therefore, ICARIA’s climate information is already based on CMIP6 models and incorporating in its workflow the current SSPs. The presented high-resolution future climate projections display a unique dataset. These models will provide the scenarios to be considered within the Risk Assessment and the design and development of all adaptation measures coming as ICARIA outcomes.</p> <p>For further details, find here a brief of the <strong>methodology </strong>followed:<strong> </strong></p> <p><strong>----- </strong></p> <p><em>The statistical downscaling methodology applied in ICARIA by FIC, named FICLIMA (Ribalaygua et al. 2013), consists of a two-step analogue/regression statistical method which has been used in national and international projects with good verification results (i.e.: Monjo et al. 2016). The first step is common for all simulated climate variables and it is based on an analogue stratification (Zorita et al. 1993). An analogue method was applied based on the hypothesis that ‘analogue’ atmospheric patterns (predictors) should cause analogue local effects (predictands), which means that the number of days that were most similar to the day to be downscaled was selected. The similarity between any two days was measured according to three nested synoptic windows (with different weights) and four large-scale fields using a pseudo-Euclidean distance between the large-scale fields used as predictors. For each predictor, the weighted Euclidean distance was calculated and standardised by substituting it with the closest percentile of a reference population of weighted Euclidean distances for that predictor. This method is a good method for reproducing nonlinear relationships between predictors and the predictands, but it could not be used to simulate values outside of the range of observed values. In order to overcome this problem and obtain a better simulation, a second step was required.</em></p> <p><em>For this second step, the procedures applied depend on the variable of interest. To determine the temperature, multiple linear regression analysis for the selected number of most analogous days was performed for each station and for each problem day. From a group of potential predictors, the linear regression selected those with the highest correlation, using a forward and backward stepwise approach.</em></p> <p><em>For precipitation, a group of m problem days (we use the whole days of a month) is downscaled. For each problem day we obtain a “preliminary precipitation amount” averaging the rain amount of its n most analogous days, so we can sort the m problem days from the highest to the lowest “preliminary precipitation amount”. For assigning the final precipitation amount, all amounts of the m×n analogous days are sorted and clustered in m groups. Every quantity is finally assigned, orderly, to the m days previously sorted by the “preliminary precipitation amount”.</em></p> <p><em>For wind or relative humidity, the second step is a transfer function between the observed probability distribution and the simulated one using the averaged values from the n = 30 analogous days. Particularly, a parametric bias correction was performed to the time series obtained from the analogue stratification (first step). In order to estimate the improvement of this procedure, the bias correction was also applied to the direct model outputs.</em></p> <p><em>This second step done at a daily scale with an inner thorough verification procedure is essential and the main differentiating process of FICLIMA method. It extends beyond mean values to include extremes and covers all time scales, including daily intervals. With the verification it can be proven If the method correctly simulates changes from one day to the next, indicating an effective capture of the underlying physical connections between predictors and predictands. These physical links remain relatively consistent, even in the face of climate change (as opposed to purely empirical relationships that might shift). In essence, this approach theoretically addresses the primary challenge in statistical downscaling known as the non-stationarity problem. This problem questions the stability of predictor/predictand relationships established in the past, probing whether these relationships will persist in the future.</em></p> <p>-----</p> <p>The dataset shared here includes information for the three case studies tackled in ICARIA: <strong>Barcelona Metropolitan Area (AMB), Salzburg Region (SLZ), and South Aegean Region (SAR)</strong>. The information provided covers data and outcomes by 10 models belonging to CMIP6. Each model has a historical archive, from 01/01/1950 to 31/12/2014 and 4 future scenarios (ssp126, ssp245, ssp370 and ssp585) ranging from 01/01/2015 to 31/12/2100. The relation of the selected models is detailed in the next Table:</p> <p><strong>Table 1</strong>.<em> Information about the 10 climate models belonging to the 6 Coupled Model Intercomparison Project (CMIP6) corresponding to the IPCC AR6. Models were retrieved from the Earth System Grid Federation (ESGF) portal in support of the Program for Climate Model Diagnosis and Intercomparison (PCMDI).</em></p> <div> <div> <table> <tbody> <tr> <td> <p><strong>CMIP6 MODELS</strong></p> </td> <td> <p><strong>Resolution</strong></p> </td> <td> <p><strong>Responsible Centre</strong></p> </td> <td> <p><strong>References</strong></p> </td> </tr> <tr> <td> <p>ACCESS-CM2</p> </td> <td> <p>1,875º x 1,250º</p> </td> <td> <p>Australian Community Climate and Earth System Simulator (ACCESS), Australia</p> </td> <td> <p>Bi, D. et al (2020)</p> </td> </tr> <tr> <td> <p>BCC-CSM2-MR</p> </td> <td> <p>1,125º x 1,121º</p> </td> <td> <p>Beijing Climate Center (BCC), China Meteorological Administration, China.</p> </td> <td> <p>Wu T. et al. (2019)</p> </td> </tr> <tr> <td> <p>CanESM5</p> </td> <td> <p>2,812º x 2,790º</p> </td> <td> <p>Canadian Centre for Climate Modeling and Analysis (CC-CMA), Canadá.</p> </td> <td> <p>Swart, N.C. et al. (2019)</p> </td> </tr> <tr> <td> <p>CMCC-ESM2</p> </td> <td> <p>1,000º x 1,000º</p> </td> <td> <p>Centro Mediterraneo sui Cambiamenti Climatici (CMCC).</p> </td> <td> <p>Cherchi et al, 2018</p> </td> </tr> <tr> <td> <p>CNRM-ESM2-1</p> </td> <td> <p>1,406º x 1,401º</p> </td> <td> <p>CNRM (Centre National de Recherches Meteorologiques), Meteo-France, Francia.</p> </td> <td> <p>Seferian, R. (2019)</p> </td> </tr> <tr> <td> <p>EC-EARTH3</p> </td> <td> <p>0,703º x 0,702º</p> </td> <td> <p>EC-EARTH Consortium</p> </td> <td> <p>EC-Earth Consortium. (2019)</p> </td> </tr> <tr> <td> <p>MPI-ESM1-2-HR</p> </td> <td> <p>0,938º x 0,935º</p> </td> <td> <p>Max-Planck Institute for Meteorology (MPI-M), Germany.</p> </td> <td> <p>Müller et al., (2018)</p> </td> </tr> <tr> <td> <p>MRI-ESM2-0</p> </td> <td> <p>1,125º x 1,121º</p> </td> <td> <p>Meteorological Research Institute (MRI), Japan.</p> </td> <td> <p>Yukimoto, S. et al. (2019)</p> </td> </tr> <tr> <td> <p>NorESM2-MM</p> </td> <td> <p>1,250º x 0,942º</p> </td> <td> <p>Norwegian Climate Centre (NCC), Norway.</p> </td> <td> <p>Bentsen, M. et al. (2019)</p> </td> </tr> <tr> <td> <p>UKESM1-0-LL</p> </td> <td> <p>1,875º x 1,250º</p> </td> <td> <p>UK Met Office, Hadley Centre, United Kingdom</p> </td> <td> <p>Good, P. et al. (2019)</p> </td> </tr> </tbody> </table> <p> </p> <p>The results shared here are developed over each of the observational locations that were retrieved to run the statistical downscaling. Both the observational datasets and the future climate change projections can be found here in a TXT format for each of the locations where they were developed. Observations include the main variables retrieved after a quality and homogeneity control, and climate projections together with extreme indicators include each of the 10 models, the 4 Tier 1 SSPs and data until the year 2100. The variables treated belong to the main climate variables and their related extreme indicators as they were defined during the ICARIA project. You can find here a summary table of all the variables and indicators that were used to develop the projections.</p> <strong>Table 2.</strong> <em>Summary of selected thermal and precipitation indicators, grouped aligned with the main hazards they feed. “nd” = number of days; “ne” = number of events.</em> <div> <table> <tbody> <tr> <td> <p><strong>Index/name</strong></p> </td> <td> <p><strong>Short description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Variable</strong></p> </td> <td> <p><strong>Units</strong></p> </td> <td> <p><strong>Threshold</strong></p> </td> </tr> <tr> <td> <p><strong>Thermal indicators</strong></p> </td> </tr> <tr> <td> <p>TX90 / TX10</p> </td> <td> <p>Warm/cold days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>90 / 10%</p> </td> </tr> <tr> <td> <p>HD</p> </td> <td> <p>Heat day</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>> 30 °C</p> </td> </tr> <tr> <td> <p>EHD</p> </td> <td> <p>Extreme heat day</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>> 35 °C</p> </td> </tr> <tr> <td> <p>TR</p> </td> <td> <p>Tropical nights</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>> 20 °C</p> </td> </tr> <tr> <td> <p>EQ</p> </td> <td> <p>Equatorial nights</p> </td> <td> <p>AEMet 2020, ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>> 25 °C</p> </td> </tr> <tr> <td> <p>IN</p> </td> <td> <p>Infernal nights</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>> 30 °C</p> </td> </tr> <tr> <td> <p>FD</p> </td> <td> <p>Frost days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>< 0 °C</p> </td> </tr> <tr> <td> <p>Max consec</p> </td> <td> <p>Max spell length for above thermal indicators</p> </td> <td> <p>ICARIA</p> </td> <td> <p>-</p> </td> <td> <p>nd</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>Nº events</p> </td> <td> <p>Number of above thermal indicators events</p> </td> <td> <p>ICARIA</p> </td> <td> <p>-</p> </td> <td> <p>ne</p> </td> <td> <p>> 3 days</p> </td> </tr> <tr> <td> <p>TXm</p> </td> <td> <p>Mean maximum temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>°C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>TNm</p> </td> <td> <p>Mean minimum temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>°C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>TM</p> </td> <td> <p>Mean temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TA</p> </td> <td> <p>°C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>HWle</p> </td> <td> <p>Heatwave length</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HWim/HWix</p> </td> <td> <p>Mean and maximum heatwave intensity</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>°C</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HWf</p> </td> <td> <p>Heatwave frequency</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>ne</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HWd</p> </td> <td> <p>Heatwave days</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HI - P90</p> </td> <td> <p>Heat Index (percentile 90)</p> </td> <td> <p>NWS (1994)</p> </td> <td> <p>TX, RH</p> </td> <td> <p>°C</p> </td> <td> <p>TX>27 °C, HR> 40%</p> </td> </tr> <tr> <td> <p>UTCI</p> </td> <td> <p>Universal Thermal Climate Index</p> </td> <td> <p>Bröde et al. (2012)</p> </td> <td> <p>TA<br>RH, W</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>UHI</p> </td> <td> <p>Isla de calor (BCN) anual y estacional</p> </td> <td> <p>AMB, Metrobs 2015</p> </td> <td> <p>T</p> </td> <td> <p>°C</p> </td> <td> <p>TM1-TM2 > 0 °C</p> </td> </tr> <tr> <td> <p><strong>Precipitation indicators</strong></p> </td> </tr> <tr> <td> <p>R20</p> </td> <td> <p>Number of heavy precipitation days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>>20 mm</p> </td> </tr> <tr> <td> <p>R50, R100</p> </td> <td> <p>Days with extreme heavy rain</p> </td> <td> <p>AMB et al. (2017)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>>50mm</p> <p>>100mm</p> </td> </tr> <tr> <td> <p>Ra</p> </td> <td> <p>Yearly and seasonal rainfall relative change</p> </td> <td> <p>ICARIA</p> </td> <td> <p>P</p> </td> <td> <p>mm</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p>IDF - CCF</p> </td> <td> <p>IDF Curves - Climate Change Factor</p> </td> <td> <p>Arnbjerg-Nielsen (2012)</p> </td> <td> <p>P</p> </td> <td> <p>-</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p><strong>Forest fire indicators</strong></p> </td> </tr> <tr> <td> <p>Mean FWI</p> </td> <td> <p>Mean Canadian FWI in fire season</p> </td> <td> <p>Stock, B.J. et al. (1989)</p> </td> <td> <p>RHn, TX, P, W</p> </td> <td> <p>.</p> </td> <td> <p>June-<br>September</p> </td> </tr> <tr> <td> <p>Very High FWI</p> </td> <td> <p>Very High Canadian FWI</p> </td> <td> <p>Stock, B.J. et al. (1989)</p> </td> <td> <p>RHn, TX, P, W</p> </td> <td> <p>nd</p> </td> <td> <p>FWI > 38</p> </td> </tr> </tbody> </table> </div> <p> </p> <p><strong>Table 3</strong>. <em>Summary of selected drought, oceanic and wind indicators, grouped aligned with the main hazards they feed. “nd” = number of days; “ne” = number of events.</em></p> <div> <table> <tbody> <tr> <td> <p><strong>Index/name</strong></p> </td> <td> <p><strong>Short description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Variable</strong></p> </td> <td> <p><strong>Units</strong></p> </td> <td> <p><strong>Threshold</strong></p> </td> </tr> <tr> <td> <p><strong>Drought indicators</strong></p> </td> </tr> <tr> <td> <p>CDDx</p> </td> <td> <p>Maximum dry spell duration</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>< 1 mm</p> </td> </tr> <tr> <td> <p>CDDm</p> </td> <td> <p>Mean dry spell duration</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>< 1 mm</p> </td> </tr> <tr> <td> <p>SPI</p> </td> <td> <p>SPI </p> <p>of 1, 3, 6, 12, 24 & 36 months</p> </td> <td> <p>McKee et al. (1993) </p> </td> <td> <p>P, TA</p> </td> <td> <p>mm</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p>SPEI</p> </td> <td> <p>SPEI </p> <p>of 1, 3, 6, 12, 24 & 36 months</p> </td> <td> <p>Vicente-Serrano et al. (2010)</p> </td> <td> <p>P, TA</p> </td> <td> <p>mm</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p><strong>Oceanic indicators</strong></p> </td> </tr> <tr> <td> <p>SS</p> </td> <td> <p>Storm surge</p> </td> <td> <p>Bryant et al. (2016)</p> </td> <td> <p>MT</p> </td> <td> <p>cm</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>OW</p> </td> <td> <p>Significant/maximum wave height</p> </td> <td> <p>ICARIA</p> </td> <td> <p>WH</p> </td> <td> <p>m</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>Wind indicators</p> </td> </tr> <tr> <td> <p>EWG</p> </td> <td> <p>Extreme wind gusts</p> </td> <td> <p>ICARIA</p> </td> <td> <p>W</p> </td> <td> <p>km/h</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> </div> <p> </p> </div> </div>
Appendix - Potential COVID-19 test fraud detection: Findings from a pilot study comparing conventional and statistical approaches
<p>The methods and results of the publication "COVID-19 test fraud detection: Findings from a pilot study comparing conventional and statistical approaches" are described in more detail in this appendix. The R-syntax for the calculation is provided, as well as a pseudo data set with which the syntax can also be tested.</p>
ICARIA: spatially distributed climate projections from statistical downscaling
<p><strong>ICARIA </strong>project had as one of its main purposes to develop coherent, reliable and usable downscaled climate projections from the last CMIP6 in order to construct the basis for efficient support to climate adaptation and decision-making of the related stakeholders, supporting the adaptation of critical assets within the project. These projections were obtained with also the purpose of being freely available for further use in subsequent studies and, hence, foster adaptation to climate change in more areas. Therefore, ICARIA’s climate information is already based on CMIP6 models and incorporating in its workflow the current SSPs. The presented high-resolution future climate projections display a unique dataset, being obtained from a high-quality and high-density set of weather observations that are then interpolated to the case studies of interest in a <strong>100x100m resolution grid, </strong>which is the main outcome offered in this publication. These models will provide the scenarios to be considered within the Risk Assessment and the design and development of all adaptation measures coming as ICARIA outcomes.</p> <p>For further details, find here a brief of the <strong>methodology </strong>followed:<strong> </strong></p> <p><strong>----- </strong></p> <p><em>The statistical downscaling methodology applied in ICARIA by FIC, named FICLIMA (Ribalaygua et al. 2013), consists of a two-step analogue/regression statistical method which has been used in national and international projects with good verification results (i.e.: Monjo et al. 2016). The first step is common for all simulated climate variables and it is based on an analogue stratification (Zorita et al. 1993). An analogue method was applied based on the hypothesis that ‘analogue’ atmospheric patterns (predictors) should cause analogue local effects (predictands), which means that the number of days that were most similar to the day to be downscaled was selected. The similarity between any two days was measured according to three nested synoptic windows (with different weights) and four large-scale fields using a pseudo-Euclidean distance between the large-scale fields used as predictors. For each predictor, the weighted Euclidean distance was calculated and standardised by substituting it with the closest percentile of a reference population of weighted Euclidean distances for that predictor. This method is a good method for reproducing nonlinear relationships between predictors and the predictands, but it could not be used to simulate values outside of the range of observed values. In order to overcome this problem and obtain a better simulation, a second step was required.</em></p> <p><em>For this second step, the procedures applied depend on the variable of interest. To determine the temperature, multiple linear regression analysis for the selected number of most analogous days was performed for each station and for each problem day. From a group of potential predictors, the linear regression selected those with the highest correlation, using a forward and backward stepwise approach.</em></p> <p><em>For precipitation, a group of m problem days (we use the whole days of a month) is downscaled. For each problem day we obtain a “preliminary precipitation amount” averaging the rain amount of its n most analogous days, so we can sort the m problem days from the highest to the lowest “preliminary precipitation amount”. For assigning the final precipitation amount, all amounts of the m×n analogous days are sorted and clustered in m groups. Every quantity is finally assigned, orderly, to the m days previously sorted by the “preliminary precipitation amount”.</em></p> <p><em>For wind or relative humidity, the second step is a transfer function between the observed probability distribution and the simulated one using the averaged values from the n = 30 analogous days. Particularly, a parametric bias correction was performed to the time series obtained from the analogue stratification (first step). In order to estimate the improvement of this procedure, the bias correction was also applied to the direct model outputs.</em></p> <p><em>This second step done at a daily scale with an inner thorough verification procedure is essential and the main differentiating process of FICLIMA method. It extends beyond mean values to include extremes and covers all time scales, including daily intervals. With the verification it can be proven If the method correctly simulates changes from one day to the next, indicating an effective capture of the underlying physical connections between predictors and predictands. These physical links remain relatively consistent, even in the face of climate change (as opposed to purely empirical relationships that might shift). In essence, this approach theoretically addresses the primary challenge in statistical downscaling known as the non-stationarity problem. This problem questions the stability of predictor/predictand relationships established in the past, probing whether these relationships will persist in the future.</em></p> <p>-----</p> <p>The dataset shared here includes information for the three case studies tackled in ICARIA: <strong>Barcelona Metropolitan Area (AMB), Salzburg Region (SLZ), and South Aegean Region (SAR)</strong>. The information provided covers data and outcomes by 10 models belonging to CMIP6. Each model has a historical archive, from 01/01/1950 to 31/12/2014 and 4 future scenarios (ssp126, ssp245, ssp370 and ssp585) ranging from 01/01/2015 to 31/12/2100. The relation of the selected models is detailed in the next Table:</p> <p><strong>Table 1</strong>.<em> Information about the 10 climate models belonging to the 6 Coupled Model Intercomparison Project (CMIP6) corresponding to the IPCC AR6. Models were retrieved from the Earth System Grid Federation (ESGF) portal in support of the Program for Climate Model Diagnosis and Intercomparison (PCMDI).</em></p> <div> <div> <table> <tbody> <tr> <td> <p><strong>CMIP6 MODELS</strong></p> </td> <td> <p><strong>Resolution</strong></p> </td> <td> <p><strong>Responsible Centre</strong></p> </td> <td> <p><strong>References</strong></p> </td> </tr> <tr> <td> <p>ACCESS-CM2</p> </td> <td> <p>1,875º x 1,250º</p> </td> <td> <p>Australian Community Climate and Earth System Simulator (ACCESS), Australia</p> </td> <td> <p>Bi, D. et al (2020)</p> </td> </tr> <tr> <td> <p>BCC-CSM2-MR</p> </td> <td> <p>1,125º x 1,121º</p> </td> <td> <p>Beijing Climate Center (BCC), China Meteorological Administration, China.</p> </td> <td> <p>Wu T. et al. (2019)</p> </td> </tr> <tr> <td> <p>CanESM5</p> </td> <td> <p>2,812º x 2,790º</p> </td> <td> <p>Canadian Centre for Climate Modeling and Analysis (CC-CMA), Canadá.</p> </td> <td> <p>Swart, N.C. et al. (2019)</p> </td> </tr> <tr> <td> <p>CMCC-ESM2</p> </td> <td> <p>1,000º x 1,000º</p> </td> <td> <p>Centro Mediterraneo sui Cambiamenti Climatici (CMCC).</p> </td> <td> <p>Cherchi et al, 2018</p> </td> </tr> <tr> <td> <p>CNRM-ESM2-1</p> </td> <td> <p>1,406º x 1,401º</p> </td> <td> <p>CNRM (Centre National de Recherches Meteorologiques), Meteo-France, Francia.</p> </td> <td> <p>Seferian, R. (2019)</p> </td> </tr> <tr> <td> <p>EC-EARTH3</p> </td> <td> <p>0,703º x 0,702º</p> </td> <td> <p>EC-EARTH Consortium</p> </td> <td> <p>EC-Earth Consortium. (2019)</p> </td> </tr> <tr> <td> <p>MPI-ESM1-2-HR</p> </td> <td> <p>0,938º x 0,935º</p> </td> <td> <p>Max-Planck Institute for Meteorology (MPI-M), Germany.</p> </td> <td> <p>Müller et al., (2018)</p> </td> </tr> <tr> <td> <p>MRI-ESM2-0</p> </td> <td> <p>1,125º x 1,121º</p> </td> <td> <p>Meteorological Research Institute (MRI), Japan.</p> </td> <td> <p>Yukimoto, S. et al. (2019)</p> </td> </tr> <tr> <td> <p>NorESM2-MM</p> </td> <td> <p>1,250º x 0,942º</p> </td> <td> <p>Norwegian Climate Centre (NCC), Norway.</p> </td> <td> <p>Bentsen, M. et al. (2019)</p> </td> </tr> <tr> <td> <p>UKESM1-0-LL</p> </td> <td> <p>1,875º x 1,250º</p> </td> <td> <p>UK Met Office, Hadley Centre, United Kingdom</p> </td> <td> <p>Good, P. et al. (2019)</p> </td> </tr> </tbody> </table> <p>The climate projections have been developed over each of the observational locations that were retrieved to run the statistical downscaling. The results from these projections have been <strong>spatially interpolated into a 100x100m grid with a Multi-lineal Regression Model</strong> considering diverse adjustments and topographic corrections. The results presented here are the<strong> median of the 10 models used, obtained for each of the 4 SSP</strong>s and each of the time periods considered in ICARIA until the year 2100. The variables treated belong to the main climate variables and their related extreme indicators as they were defined during the ICARIA project. You can find here a summary table of all the variables and indicators that were used to develop the projections.</p> <strong>Table 2.</strong> <em>Summary of selected thermal and precipitation indicators, grouped aligned with the main hazards they feed. “nd” = number of days; “ne” = number of events.</em> <div> <table> <tbody> <tr> <td> <p><strong>Index/name</strong></p> </td> <td> <p><strong>Short description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Variable</strong></p> </td> <td> <p><strong>Units</strong></p> </td> <td> <p><strong>Threshold</strong></p> </td> </tr> <tr> <td> <p><strong>Thermal indicators</strong></p> </td> </tr> <tr> <td> <p>TX90 / TX10</p> </td> <td> <p>Warm/cold days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>90 / 10%</p> </td> </tr> <tr> <td> <p>HD</p> </td> <td> <p>Heat day</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>> 30 °C</p> </td> </tr> <tr> <td> <p>EHD</p> </td> <td> <p>Extreme heat day</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>> 35 °C</p> </td> </tr> <tr> <td> <p>TR</p> </td> <td> <p>Tropical nights</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>> 20 °C</p> </td> </tr> <tr> <td> <p>EQ</p> </td> <td> <p>Equatorial nights</p> </td> <td> <p>AEMet 2020, ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>> 25 °C</p> </td> </tr> <tr> <td> <p>IN</p> </td> <td> <p>Infernal nights</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>> 30 °C</p> </td> </tr> <tr> <td> <p>FD</p> </td> <td> <p>Frost days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>< 0 °C</p> </td> </tr> <tr> <td> <p>Max consec</p> </td> <td> <p>Max spell length for above thermal indicators</p> </td> <td> <p>ICARIA</p> </td> <td> <p>-</p> </td> <td> <p>nd</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>Nº events</p> </td> <td> <p>Number of above thermal indicators events</p> </td> <td> <p>ICARIA</p> </td> <td> <p>-</p> </td> <td> <p>ne</p> </td> <td> <p>> 3 days</p> </td> </tr> <tr> <td> <p>TXm</p> </td> <td> <p>Mean maximum temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>°C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>TNm</p> </td> <td> <p>Mean minimum temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>°C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>TM</p> </td> <td> <p>Mean temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TA</p> </td> <td> <p>°C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>HWle</p> </td> <td> <p>Heatwave length</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HWim/HWix</p> </td> <td> <p>Mean and maximum heatwave intensity</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>°C</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HWf</p> </td> <td> <p>Heatwave frequency</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>ne</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HWd</p> </td> <td> <p>Heatwave days</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>3d > 95% TX</p> </td> </tr> <tr> <td> <p>HI - P90</p> </td> <td> <p>Heat Index (percentile 90)</p> </td> <td> <p>NWS (1994)</p> </td> <td> <p>TX, RH</p> </td> <td> <p>°C</p> </td> <td> <p>TX>27 °C, HR> 40%</p> </td> </tr> <tr> <td> <p>UTCI</p> </td> <td> <p>Universal Thermal Climate Index</p> </td> <td> <p>Bröde et al. (2012)</p> </td> <td> <p>TA<br>RH, W</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>UHI</p> </td> <td> <p>Isla de calor (BCN) anual y estacional</p> </td> <td> <p>AMB, Metrobs 2015</p> </td> <td> <p>T</p> </td> <td> <p>°C</p> </td> <td> <p>TM1-TM2 > 0 °C</p> </td> </tr> <tr> <td> <p><strong>Precipitation indicators</strong></p> </td> </tr> <tr> <td> <p>R20</p> </td> <td> <p>Number of heavy precipitation days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>>20 mm</p> </td> </tr> <tr> <td> <p>R50, R100</p> </td> <td> <p>Days with extreme heavy rain</p> </td> <td> <p>AMB et al. (2017)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>>50mm</p> <p>>100mm</p> </td> </tr> <tr> <td> <p>Ra</p> </td> <td> <p>Yearly and seasonal rainfall relative change</p> </td> <td> <p>ICARIA</p> </td> <td> <p>P</p> </td> <td> <p>mm</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p>IDF - CCF</p> </td> <td> <p>IDF Curves - Climate Change Factor</p> </td> <td> <p>Arnbjerg-Nielsen (2012)</p> </td> <td> <p>P</p> </td> <td> <p>-</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p><strong>Forest fire indicators</strong></p> </td> </tr> <tr> <td> <p>Mean FWI</p> </td> <td> <p>Mean Canadian FWI in fire season</p> </td> <td> <p>Stock, B.J. et al. (1989)</p> </td> <td> <p>RHn, TX, P, W</p> </td> <td> <p>.</p> </td> <td> <p>June-<br>September</p> </td> </tr> <tr> <td> <p>Very High FWI</p> </td> <td> <p>Very High Canadian FWI</p> </td> <td> <p>Stock, B.J. et al. (1989)</p> </td> <td> <p>RHn, TX, P, W</p> </td> <td> <p>nd</p> </td> <td> <p>FWI > 38</p> </td> </tr> </tbody> </table> </div> <p> </p> <p><strong>Table 3</strong>. <em>Summary of selected drought, oceanic and wind indicators, grouped aligned with the main hazards they feed. “nd” = number of days; “ne” = number of events.</em></p> <div> <table> <tbody> <tr> <td> <p><strong>Index/name</strong></p> </td> <td> <p><strong>Short description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Variable</strong></p> </td> <td> <p><strong>Units</strong></p> </td> <td> <p><strong>Threshold</strong></p> </td> </tr> <tr> <td> <p><strong>Drought indicators</strong></p> </td> </tr> <tr> <td> <p>CDDx</p> </td> <td> <p>Maximum dry spell duration</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>< 1 mm</p> </td> </tr> <tr> <td> <p>CDDm</p> </td> <td> <p>Mean dry spell duration</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>< 1 mm</p> </td> </tr> <tr> <td> <p>SPI</p> </td> <td> <p>SPI </p> <p>of 1, 3, 6, 12, 24 & 36 months</p> </td> <td> <p>McKee et al. (1993) </p> </td> <td> <p>P, TA</p> </td> <td> <p>mm</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p>SPEI</p> </td> <td> <p>SPEI </p> <p>of 1, 3, 6, 12, 24 & 36 months</p> </td> <td> <p>Vicente-Serrano et al. (2010)</p> </td> <td> <p>P, TA</p> </td> <td> <p>mm</p> </td> <td> <p>≥ 0.1mm</p> </td> </tr> <tr> <td> <p><strong>Oceanic indicators</strong></p> </td> </tr> <tr> <td> <p>SS</p> </td> <td> <p>Storm surge</p> </td> <td> <p>Bryant et al. (2016)</p> </td> <td> <p>MT</p> </td> <td> <p>cm</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>OW</p> </td> <td> <p>Significant/maximum wave height</p> </td> <td> <p>ICARIA</p> </td> <td> <p>WH</p> </td> <td> <p>m</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>Wind indicators</p> </td> </tr> <tr> <td> <p>EWG</p> </td> <td> <p>Extreme wind gusts</p> </td> <td> <p>ICARIA</p> </td> <td> <p>W</p> </td> <td> <p>km/h</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> </div> <p> </p> </div> </div>
Genome-wide association summary statistics for human blood plasma glycome
<p>The dataset contains results of genome-wide association study of human blood plasma glycome. The 113 files contain association summary statistics for 113 glycome traits, of which 36 were directly measured by UPLC technology and 77 were derived glycome traits. Description of each glycome trait can be found in the <strong>Additional notes</strong> section. This dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org">http://gwasarchive.org</a>. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Sharapov, S. Z., Tsepilov, Y. A., Klaric, L., Mangino, M., Thareja, G., Shadrina, A. S., … Aulchenko, Y. (2019). Defining the genetic control of human blood plasma N-glycome using genome-wide association study. <em>Human Molecular Genetics</em>. http://doi.org/10.1093/hmg/ddz054</li> <li>Sodbo Sharapov, Yakov Tsepilov, Lucija Klaric, Massimo Mangino, Gaurav Thareja, Mirna Simurina, Concetta Dagostino, Julia Dmitrieva, Marija Vilaj, FranoVuckovic, Tamara Pavic, Jerko Stambuk, Irena Trbojevic-Akmacic, Jasminka Kristic, Jelena Simunovic, Ana Momcilovic, Harry Campbell, Malcolm Dunlop, Susan Farrington, Maria Pucic-Bakovic, Christian Gieger, Massimo Allegri, Edouard Louis, Michel Georges, Karsten Suhre, Tim Spector, Frances MK Williams, Gordan Lauc, Yurii Aulchenko. (2018). Genome-wide association summary statistics for human blood plasma glycome (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1298406</li> </ol> <p><strong>Funding</strong></p> <p>This work was supported by the European Community’s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736) and by the European Structural and Investments funding for the "Croatian National Centre of Research Excellence in Personalized Healthcare" (contract #KK.01.1.1.01.0010).</p> <p>The work of SSh was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme.</p> <p>The work of YT was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).</p> <p>Karsten Suhre and Gaurav Thareja are supported by ‘Biomedical Research Program’ funds at Weill Cornell Medicine - Qatar, a program funded by the Qatar Foundation. We thank all staff at Weill Cornell Medicine - Qatar and Hamad Medical Corporation, and especially all study participants who made the QMDiab study possible.</p> <p>The SOCCS study was supported by grants from Cancer Research UK (C348/A3758, C348/A8896, C348/ A18927); Scottish Government Chief Scientist Office (K/OPR/2/2/D333, CZB/4/94); Medical Research Council (G0000657-53203, MR/K018647/1); Centre Grant from CORE as part of the Digestive Cancer Campaign (<a href="http://www.corecharity.org.uk">http://www.corecharity.org.uk</a>).</p> <p>TwinsUK is funded by the Wellcome Trust, Medical Research Council, European Union, the National Institute for Health Research (NIHR)-funded BioResource, Clinical Research Facility and Biomedical Research Centre based at Guy’s and St Thomas’ NHS Foundation Trust in partnership with King’s College London.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNP: SNP rsID</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>OTHER_ALLELE: reference allele (coded as "0")</li> <li>EFFECT_ALLELE: effective allele (coded as "1")</li> <li>EAF: effective allele frequency </li> <li>N: sample size</li> <li>BETA: effect size of effective allele</li> <li>SE: standard error of effect size</li> <li>PVAL: P-value of association (without GC correction)</li> <li>IMPUTATION: imputation quality</li> </ol>
Genome-wide association summary statistics for human healthspan
<p>The dataset contains genome-wide association summary statistics computed for heathspan. The UKB sub-population of 300,447 genetically Caucasian, British individuals were analyzed. For more details see [1].</p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Zenin, A., Tsepilov, Y., Sharapov, S., Getmantsev, E., Menshikov, L. I., Fedichev, P. O., & Aulchenko, Y. (2019). Identification of 12 genetic loci associated with human healthspan. <em>Communications Biology</em>, <em>2</em>(1), 41. http://doi.org/10.1038/s42003-019-0290-0</li> <li>Aleksandr Zenin, Yakov Tsepilov, Sodbo Sharapov, Evgeny Getmantsev, Leonid Menshikov, Peter Fedichev, & Yurii Aulchenko. (2018). Genome-wide association summary statistics for human healthspan (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1302861</li> </ol> <p><strong>Funding</strong></p> <p>The work was supported by Russian Ministry of Science and Education under 5-100 Excellence Programme. <br> The work was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017). <br> This research has been conducted using the UK Biobank Resource. <br> The study has been funded by Gero LLC.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNPID - SNP rsID</li> <li>chr - chromosome</li> <li>pos - position (GRCh37 build / hg19)</li> <li>EA - effective allele (coded as "1")</li> <li>RA - reference allele (coded as "0")</li> <li>EAF - effective allele frequency</li> <li>beta - effect size of effective allele</li> <li>se - standard error of effect size</li> <li>Z - Z-value of association</li> <li>-log10(p-value) - minus log10(P-value) of association</li> </ol>
Genome-wide association summary statistics for back pain
<p>The dataset contains results of a genome-wide association study of back pain. Two files contain association summary statistics for discovery GWAS based on the analysis of 350,000 white British individuals from the UK Biobank and meta-analysis GWAS based on the meta-analysis of the same 350,000 individuals and additional 103,862 individuals of European Ancestry from the UK biobank (total N = 453,862). The phenotype of back pain was defined by the answer provided by the UK biobank participants to the following question: "Pain type(s) experienced in last month". Those who reported “Back pain”, were considered as cases, all the rest were considered as controls. Individuals who did not reply or replied: "Prefer not to answer" or "Pain all over the body" were excluded. This dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org/">http://gwasarchive.org</a>. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Insight into the genetic architecture of back pain and its risk factors from a study of 509,000 individuals. Freidin, Maxim; Tsepilov, Yakov; Palmer, Melody; Karssen, Lennart; Suri, Pradeep; Aulchenko, Yurii; Williams, Frances MK,# CHARGE Musculoskeletal Working Group. PAIN: February 06, 2019 - Volume Articles in Press - Issue - p<br> doi: 10.1097/j.pain.0000000000001514</li> <li>Maxim B Freidin, Yakov A Tsepilov, Melody Palmer, Lennart Karssen, CHARGE Musculoskeletal Working Group, Pradeep Suri, … Frances MK Williams. (2018). Genome-wide association summary statistics for back pain (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1319332</li> </ol> <p><strong>Funding:</strong></p> <p>This study was supported by the European Community’s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736). <br> The research has been conducted using the UK Biobank Resource (project # 18219).</p> <p>The development of software implementing SMR/HEIDI test and database for GWAS results was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Program”.</p> <p>Dr. Suri’s time for this work was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development Service. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p>Dr. Tsepilov’s time for this work was supported in part by the Russian Ministry of Science and Education under the 5-100 Excellence Program.</p> <p><strong>Column headers - discovery (350K)</strong></p> <ol> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>ID: SNP rsID</li> <li>REF: reference allele (coded as "0")</li> <li>ALT: effect allele (coded as "1")</li> <li>CASE_ALLELE_CT: allele observation count in cases</li> <li>CTRL_ALLELE_CT: allele observation count in controls</li> <li>ALT_FREQ: effect allele frequency </li> <li>MACH_R2: imputation quality</li> <li>TEST: model of association test (additive)</li> <li>OBS_CT: sample size</li> <li>BETA: effect size of effect allele</li> <li>SE: standard error of effect size</li> <li>T_STAT: Z-value of effect allele</li> <li>P: P-value of association (without GC correction)</li> <li>MAF: minor allele frequency</li> </ol> <p><strong>Column headers - meta-analysis (450K)</strong></p> <ol> <li>MarkerName: SNP rsID</li> <li>Allele1: effect allele (coded as "1")</li> <li>Allele2: reference allele (coded as "0")</li> <li>Freq1: effect allele frequency</li> <li>FreqSE: standard error of effect allele frequency</li> <li>Effect: effect size of effect allele</li> <li>StdErr: standard error of effect size</li> <li>P-value: P-value of association (without GC correction)</li> <li>Direction: sign of effect in discovery and replication samples</li> <li>n_total: Total sample size</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>MACH_R2_discovery: imputation quality in discovery sample</li> </ol>
SQLite database to accompany the paper, "Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring"
<p>This dataset is a SQLite database that accompanies methods and analysis described in the paper, "Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring" (Balantic & Donovan 2019, Bioacoustics, https://www.tandfonline.com/doi/full/10.1080/09524622.2019.1605309). </p> <p>A Github repository containing code for using the SQLite database also accompanies this paper at: <a href="https://github.com/cbalantic/false-positive-mitigation">http://github.com/cbalantic/false-positive-mitigation</a></p>
Dataset for paper "Ejecta cloud distributions for the statistical analysis of impact cratering events onto asteroids' surfaces: a sensitivity analysis"
<p>Dataset for the paper "Ejecta cloud distributions for the statistical analysis of impact cratering events onto asteroids' surfaces: a sensitivity analysis" published in Icarus.</p>
NEON distributed initial soil characterization dataset (DP1.10047.001) modified for statistical analysis of organic carbon and extractable metals in Hall and Thompson (2021)
We compiled National Ecological Observatory Network (NEON) datasets related to the initial distributed soil sampling effort and subsetted them (removed samples with missing values for certain variables, and several samples with extreme values) for use in statistical analyses to describe relationships between soil organic carbon (SOC) and metals measured in several soil chemical extractions. The NEON provisional data products we used were DP1.10047.001 and DP1.10008.001, which were subsequently combined by NEON as a single data product DP1.10047.001, “Soil physical and chemical properties, distributed initial characterization”. These datasets were used for the analyses reported in a manuscript by Hall and Thompson (2021) in the Soil Science Society of America Journal.
Seasonal and annual summary statistics of urbanization, vegetation, land surface temperature, and bioclimatic variables derived from remotely-sensed imagery in areas surrounding long-term bird monitoring locations in the greater Phoenix, Arizona, USA metropolitan area (1997-2023)
This data package consists of 26 years (1998-2023) of environmental data and 22 years (2000-2022) years of bioclimatic data associated with CAP-LTER long-term point-count bird censusing sites (https://doi.org/10.6073/pasta/4777d7f0a899f506d6d4f9b5d535ba09), temporally aggregated by year and by four meteorological seasons (Winter, Spring, Summer, Fall). The environmental variables include land surface temperature (LST), three spectral indices of vegetation and water – the normalized difference vegetation index (NDVI), the soil adjusted vegetation index (SAVI), and modified normalized difference water index (MNDWI) – and four spectral indices of impervious surface/urbanization. Impervious surface indices include the normalized difference built-up index (NDBI), the normalized difference impervious surface index (NDISI), the enhanced normalized differences impervious surface index (ENDISI), and the normalized impervious surface index (NISI). LST and all spectral indices were derived from annual and seasonal composites of 30-m resolution Landsat 5-9 Level-2 Surface Reflectance imagery. The seven bioclimatic variables (e.g., air temperature, precipitation) were sourced from 1-km resolution gridded estimates of daily climatic data from NASA Daymet V4. We created temporally-aggregated Daymet raster images by calculating mean pixel-values for each season and year, as well as seasonally and annually summed precipitation. We summarized the values of each environmental variable by generating variously-sized (100-m, 500-m, 1000-m) buffers around each bird point count location and extracting weighted mean values of each environmental variable, with each pixel's values weighted by the proportion of its area falling within the buffer. All imagery retrieval and data processing were completed with Google Earth Engine (Gorelick et al. 2017) and program R. A complete description of data processing methods, including the aggregation of imagery by year and season and the calculation of s
Summary statistics and annual trends in chloride concentration in urban Minnesota lakes and streams
These data tables describe statistical summaries of chloride concentration, temporal trends in annual chloride, and projected risk of future chloride pollution in lakes and streams of the 17 most urban counties in Minnesota. Data were summarized separately for lakes and for streams, and include statistics (mean, median, standard deviation, upper and lower confidence intervals, and maximum) over the entire data record, over the warm season (May - October), and over the most recent 5 years (i.e., since 2018). Trends were computed on annual means and medians. For lakes, data were aggregated by lake basin (MN DNR Lake ID, or DOW) as well as by depth of sample (surface and deep). For streams, data were aggregated by individual site level as well as by stream reach (per MN Pollution Control Agency assessment units). Risk of chloride pollution was also determined for sites with longer records (10+ years) based on current concentration, number of exceedances of chronic standards, and projected chloride concentration based on current trends. Raw data were extracted from two sources: (1) the National Water Quality Portal (USGS & EPA) and (2) the Metropolitan Council Environmental Information Management System. The data were originally collected by a large number of entities, including watershed management authorities in the state of Minnesota, tribal groups, the Minnesota Pollution Control Agency, municipalities, university researchers, private consultants, and the Metropolitan Council. Some data records begin as early as the 1950's or 1960's, with many sites still including active data collection. A total of approximately 45,000 observations of chloride were included for lakes and wetlands, and approximately 70,000 observations for streams. Nearly 1600 stream sites and 700 lake/wetland sites were represented in the raw data, with 356 stream sites and 600 lakes represented in the summaries (after filtering out sites with less than 1 year of data collection). Primary data retr
ReFAB 10,000 Year Statistical Estimate of Aboveground Woody Biomass, Midwest US, Level 2
How terrestrial biomass changed before the advent of industrial society is a major gap in our understanding of the Earth's carbon cycle. Here, we archive data used to reconstruct 10,000 years of aboveground woody biomass across the US Upper Midwest using statistical models based on historical forest surveys and fossil pollen assemblages. From our analyses we document a 5,000 year long carbon sink into vegetation, primarily caused by the range expansion of two late-successional species into the region during the late Holocene. The importance of such large slow-growing tree species in storing carbon during the pre-industrial past argues for protecting similar species in wild forests today. This material is based upon work supported by the National Science Foundation under grants #DEB-1241874, 1241868, 1241870, 1241851, 1241891, 1241846, 1241856, 1241930.
Models and Infrastructure used in "Deep Statistical Model Checking"
<p>This repository contains the models and all other infrastructure (learning procedure, NNs, Jani generator, maps, modes & mcsta binaries) used in the FORTE 2020 paper "Deep Statistical Model Checking".</p>
Supplementary data for Nested sampling cross-checks using order statistics
<p>This is the raw data for tables 1 and 2 in <a href="https://arxiv.org/abs/2006.03371">Nested sampling cross-checks using order statistics</a>. The file names for the MultiNest results in table 1 are:</p> <pre><code>MN_{PROBLEM}_{number of dimensions}d_efr_{MultiNest efr parameter}.txt</code></pre> <p>The file names for the PolyChord results in table 2 are:</p> <pre><code>PC_{PROBLEM}_{number of dimensions}d_nr_{PolyChord number of repeats}.txt</code></pre> <p>where PROBLEM specifies one of four test functions described in Appendix C</p> <ul> <li>gaussian = Gaussian</li> <li>mixture = Gaussian-log-gamma mixture</li> <li>rosenbrock = Rosenbrock function</li> <li>shells = Gaussian shells</li> </ul> <p>Each file contains 100 rows (corresponding to 100 runs) and 8 columns</p> <ol> <li>log evidence</li> <li>error log evidence</li> <li>KS statistic from test on all iterations</li> <li>number of iterations</li> <li>p-value from all iterations</li> <li>Greatest KS statistic from tests on chunks of iterations</li> <li>Iteration of start of chunk at which greatest KS statistic occurred</li> <li>Bonferroni corrected p-value from greatest KS statistic from tests on chunks of iterations</li> </ol>
Statistics of the sonnendach.ch dataset
<p>This dataset contains statistics of the sonnendach.ch dataset at the national level. See README.md for more information.</p>
UShER performance statistics, SARS-CoV-2 daily builds 2021-2023
<p>For each day from 2021-01-07 through 2023-08-01 on which the daily build update of the UShER tree of SARS-CoV-2 genomes completed, the number of new sequences added to the tree, the number of sequences in the updated tree, the number of parallel usher jobs (original usher through 2022-04-27, usher-sampled starting 2022-04-29), the number of CPU cores per usher job, and approximate runtime of the usher batch in hours are listed. The number of sequences in the updated tree is "n/a" for most dates prior to 2021-03-11 because before that point, daily updates were for the public-sequence-only tree and the comprehensive GISAID and public sequence tree was updated only occasionally. On and after 2021-03-11, the comprehensive tree and public tree were updated daily. The runtime figures are approximate because they are calculated by subtracting the file modification date of the VCF input to usher from the file modification date of the MAT output of usher. On most days, that was a good proxy for usher runtime, but occasionally there was a crash that required debugging and/or restart, and the "runtime" includes those delays.</p>
A Data-driven Analysis of a Cloud Data Center: Statistical Characterization of Workload, Energy and Temperature
<p>A characterization of cloud data center logs, analyzing its workload, energy and thermal characteristics. For more details of the dataset, please read the following paper: <a href="http://hpc.ec.tuwien.ac.at/files/UCC_23_data_center_analysis.pdf">http://hpc.ec.tuwien.ac.at/files/UCC_23_data_center_analysis.pdf.</a></p><p> </p><p>If you use the dataset, please cite the following work:</p><p>Shashikant Ilager, Adel N. Toosi, Mayank Raj Jha, Ivona Brandic, Rajkumar Buyya, "A Data-driven Analysis of a Cloud Data Center: Statistical Characterization of Workload, Energy and Temperature", In Proceedings of the 16th IEEE/ACM International Conference on Utility and Cloud Computing (UCC2023), Messina, Italy, December 4-7, 2023.</p>
Towards standardising the collection of game statistics in Europe: a dataset
<p>This dataset is part of the article:</p><p>Title : <strong>Towards standardising the collection of game statistics in Europe: a case study</strong></p><p>Journal:<i><strong> European Journal of Wildlife Research.</strong></i><br><strong>DOI : 10.1007/s10344-023-01746-3</strong></p><p>Two different data sets have been incorporated : <br>1) Data collected from a questionnaire to regional governmental hunting agencies (mainland Spain)</p><p><a href="https://zenodo.org/api/records/10080464/draft/files/QuestionnaireData_DOI_10.1007_s10344-023-01746-3.csv/content">QuestionnaireData_DOI_10.1007_s10344-023-01746-3.csv</a></p><p>Metadata with variable vocabulary and descriptors has been included</p><p><a href="https://zenodo.org/api/records/10080464/draft/files/QuestionnaireMetadata_DOI_10.1007_s10344-023-01746-3.docx/content">QuestionnaireMetadata_DOI_10.1007_s10344-023-01746-3.docx</a></p><p>2) Characterisation of each of the Autonomous Communities, including information on economic and human resources and the volume or coverage of hunting resources available in each of the regional administrations. </p><p><a href="https://zenodo.org/api/records/10080464/draft/files/SocioEconomicData_DOI_10.1007_s10344-023-01746-3.csv/content">SocioEconomicData_DOI_10.1007_s10344-023-01746-3.csv</a></p><p>Metadata with variable vocabulary and descriptors has been included</p><p><a href="https://zenodo.org/api/records/10080464/draft/files/SocioEconomicMetadata_DOI_10.1007_s10344-023-01746-3.docx/content">SocioEconomicMetadata_DOI_10.1007_s10344-023-01746-3.docx</a></p><p> </p><p>Please remember to cite correctly the doi associated to this databases <strong>10.5281/zenodo.10080464</strong></p><p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.