Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
144
datasets available to search
ShareScore release 0.9.0
Dataset results
144 results for “statistical modeling”
Models and Infrastructure used in "Deep Statistical Model Checking"
<p>This repository contains the models and all other infrastructure (learning procedure, NNs, Jani generator, maps, modes & mcsta binaries) used in the FORTE 2020 paper "Deep Statistical Model Checking".</p>
Regional Estimates of Chemical Composition of Fine Particulate Matter Using a Combined Geoscience-Statistical Method with Information from Satellites, Models, and Monitors: V4.NA.02.MAPLE
<p>We estimate ground-level fine particulate matter (PM<sub>2.5</sub>) total and compositional mass concentrations over North America by combining Aerosol Optical Depth (AOD) retrievals from the NASA MODIS, MISR, and SeaWIFS instruments with the GEOS-Chem chemical transport model, and subsequently calibrated to regional ground-based observations of both total and compositional mass using Geographically Weighted Regression (GWR) as detailed in the provided reference for V4.NA.02. V4.NA.02.MAPLE further modified the V4.NA.02 GWR method with additional developments as part of the MAPLE (Mortality–Air Pollution Associations in Low-Exposure Environments) project. This adjustment was of particular value over low concentrations. The GWR method of individual components remains unchanged from V4.NA.02, but are provided are percentages to ensure mass closure and recommended to be applied to the V4.NA.02.MAPLE total PM<sub>2.5</sub>.</p> <p>Annual datasets are provided in NetCDF [.nc]. Gridded files use the WGS84 projection. Compositional estimates are provided for sulfate (SO4), nitrate (NO3), ammonium (NH4), organic matter (OM), black carbon (BC), mineral dust (DUST), and sea-salt (SS). Percentages are denoted with a ‘p’ after component identifiers within filenames. A slight change in file name has been included for 2017, corresponding to minor internal changes compared to earlier years. Overall, however, the dataset is consistent throughout its entire time period and can be appropriately used for trend analysis.</p> <p><strong>Reference:</strong><br> van Donkelaar, A., R. V. Martin, et al. (2019). <strong>Regional Estimates of Chemical Composition of Fine Particulate Matter using a Combined Geoscience-Statistical Method with Information from Satellites, Models, and Monitors.</strong> Environmental Science & Technology, 2019, doi:10.1021/acs.est.8b06392.</p>
Data for publication "Statistical characteristics of extreme daily precipitation during 1501 BCE - 1849 CE in the Community Earth System Model".
<p>Here, the data used in Kim, W. M., Blender, R., Sigl, M., Messmer, M., & Raible, C. C. (2021). "Statistical characteristics of extreme daily precipitation during 1501 BCE–1849 CE in the Community Earth System Model" in <em>Climate of the Past </em>(<a href="https://doi.org/10.5194/cp-2021-61">https://doi.org/10.5194/cp-2021-6</a>) are provided.</p> <p>Two simulations covering the period 1501 BCE - 2008 CE are performed with CESM 1.2.2: the orbital-only and the full-forcing simulations. The full-forcing transient simulation includes the new long record of volcanic eruptions (<a href="https://doi.org/10.1594/PANGAEA.928646">https://doi.org/10.1594/PANGAEA.928646</a>) that covers the last 3500 years. The output from the simulations is used to examine the long-term variability and characteristics of daily extreme precipitation during 1501BCE-1849 CE.</p> <p>The following files are provided:</p> <ul> <li> <strong>CESM122.transient.PRECT.anom.above99th.1501BCE-1849CE_I and II</strong>: Daily precipitation anomalies above the 99th percentiles relative to 1501BCE-1849CE from the full forcing simulation. The file is split into two parts, with the first file containing the first 50% of extremes (I) and the second file containing the rest 50% (II).</li> <li><strong>CESM122.orbital.PRECT.anom.above99th.1501BCE-1849CE I and II:</strong> Daily precipitation anomalies above the 99th percentiles relative to 1501BCE-1849CE from the orbital-only simulation.</li> <li> <strong>CESM122.transient.variables.mon.1979-2008CE:</strong> monthly precipitation, temperature, and geopotential height at 500 hPa for 1979-2008CE from the full-forcing simulation.</li> <li><strong>CESM122.trans.variable_names.years:</strong> Monthly variables from the full-forcing simulation. The simulation starts from the model year 1, which corresponds to the actual year 1501BCE. The variables are solar insolation (SOLIN), clear-sky net surface shortwave radiation (FSNSC), geopotential height at 500hPa (Z500), and surface temperature (TS).</li> <li><strong>CESM122.orbital.variable_names.years:</strong> Monthly variables from the orbital-only simulation. The simulation starts from the model year 1, which corresponds to the actual year 1501BCE.</li> <li><strong>CESM122.*.log-likelihood-GPDmodel-ExtForcing</strong>: Negative log-likelihood for the stationary and non-stationary Generalized Pareto Distribution models for external forcings.</li> <li> <strong>CESM122.*.log-likelihood-GPDmodel-ModesVar</strong>: Negative log-likelihood for the stationary and non-stationary Generalized Pareto Distribution models for modes of variability.</li> <li><strong>Evolk_EVA_distribution_1501BCE-2015CE</strong>: Distribution of volcanic aerosol for CAM5, produced based on Kim et al. (2021).</li> </ul> <p>If you use this dataset, please cite:</p> <p><em>Kim, W. M., Blender, R., Sigl, M., Messmer, M., & Raible, C. C. (2021). Statistical characteristics of extreme daily precipitation during 1501 BCE–1849 CE in the Community Earth System Model. Climate of the Past Discussions, 1-38. <a href="https://doi.org/10.5194/cp-2021-61">https://doi.org/10.5194/cp-2021-61</a></em></p>
Ensemble statistics for modelled Eddy Kinetic Energy in the Southern Ocean
<p>This dataset contains surface eddy kinetic energy over the Southern Ocean region, sourced from a 50-member ensemble of 0.25° ocean model simulations. It is used in the paper "Circumpolar variations in the chaotic nature of Southern Ocean eddy dynamics" published in Journal of Geophysical Research - Oceans.</p> <p>This dataset has been computed from the OceaniC Chaos – ImPacts, strUcture, predicTability (OCCIPUT) global ocean/sea-ice ensemble simulation. It is composed of 50 members with a horizontal resolution of 1/4° and 75 geopotential levels (<a href="http://doi.org/10.5194/gmd-10-1091-2017">Bessières et al., 2017</a>, Penduff et al., 2014). The numerical configuration is based on the version 3.5 of the NEMO model (<a href="https://www.nemo-ocean.eu/doc">Madec, 2008</a>). The 50 members were started on January 1st 1960 from a common 21-year spinup. A small stochastic perturbation is applied to the equation of state of sea water (as in <a href="https://doi.org/10.1016/j.ocemod.2013.02.004">Brankart, 2013</a>) within each member during 1960, then switched off during the rest of the simulation. This 1-year perturbation generates an ensemble spread which grows and saturates after a few months up to a few years depending on the region. The 50 members are driven through bulk formulae during the whole 1960-2015 simulation by the same realistic 6-hourly atmospheric forcing (Drakkar Forcing Set DFS5.2, Dussin et al., 2016) derived from ERA interim atmospheric reanalysis. Data is for the period 1979-2015.</p> <p>The sea level anomaly is found according to <a href="http://doi.org/10.1016/j.pocean.2020.102314">Close et al (2020)</a> and converted into surface geostrophic velocity anomaly using the geostrophic relation. This velocity field is then used to calculate the eddy kinetic energy (EKE). Data is averaged over calendar month, and restricted to the latitude range 40°-60°S. A full description of this process is included in the companion paper.</p> <p>The dataset includes EKE files (eke_0??.nc), with monthy EKE saved for the period 1979-2015 for each ensemble member, and a single file (tau.nc) for the monthly-averaged wind stress over the same period.</p>
Statistical characterization of Andalusian wave climate for several combinations of Global Climate Models and Regional Climate Models and periods 2026 - 2045 and 2081 - 2100.
<p>The following text is an extract of the extended abstract entitled "<strong>Parametric Characterization of Wave Climate along the Andalusian Coast for Non-Stationary Stochastic Simulation</strong>" whose authors are Manuel Cobos, Pedro Magaña, Pedro Otiñar and Asunción Baquerizo, and that was included in proceedings of <em>39th IAHR World Congress</em> where this dataset is included.</p> <p><em>Processed data comes from PIMA Adapta Costas project (Ramírez et al., 2019), in particular, from projections of maritime climate for 2026-2045 and 2081-2100. Sea climate contains, among other information, time series of the significant wave height (H<sub>s</sub>) obtained for several combinations of GCM-RCM projections of EUR-11 for the RCP 8.5. GCM-RCM combinations ACCE, CMCC, CNRM, GFDL, HADG, IPSL, MIRO with a 0.1 degrees grid were used for the Atlantic facade while CNRM, HADG, IPSL, MIRO, MEDC, MPIE, ESM2, EART models with 1/11 degrees were used for the Mediterranean one. A total of 210 locations were analyzed, 54 at the Atlantic facade and 156 at the Mediterranean one (Figure 1). The data was bias adjusted using the Empirical Quantile Mapping (Déqué et al., 2007; Michelangeli et al., 2009). Information of the significant wave height and the dependence between the values at a given time with previous values with a VAR(q) model is already available. </em></p> <p><em>At each location, the methodology of Lira-Loarca et al. (2021) was applied, using the software described in Cobos et al. (2022a). More precisely, for every GCM-RCM (hereinafter, model n for n = 1, .., N where N = 7 for Atlantic data and N = 8 for the Mediterranean data), a non-stationary marginal distribution of H<sub>s</sub>, , assuming that the year was the largest periodicity of the climate, was fitted to data using a lognormal model for the central part and two generalized Pareto distribution for the lower and upper tails, as in Solari and Losada (2011). The non- stationarity is considered by assuming a decomposition of the parameters of the distribution and of the percentiles of the common end points of the interval into a trigonometric truncated expansion.</em></p> <p><em>In addition, the coefficients of the matrix, C<sub>n</sub>, of a VAR(q) model with q up to 92 hours were estimated. The ensemble multi-model characteristics of the data were obtained from the compound distributions and the weighted averaged matrix coefficients. </em></p> <p><em>Soon, the results of the peak period (T<sub>p</sub>) and mean incoming wave direction (ϑ<sub>m</sub>) and the coefficients of the multivariate VAR model will also be included.</em></p> <p> </p> <p> </p>
Multiscale continuum figures from Tratnyek et al. (2017) "In silico environmental chemical science: Properties and processes from statistical and computational modelling"
<p>Accessible versions of selected figures from Tratnyek et al. (2017) "In silico environmental chemical science: Properties and processes from statistical and computational modelling" Environ. Sci. Processes Impacts 19(3): 188-202. DOI: 10.1039/C7EM00053G.</p> <p>The Abstract Art figure shows a classification of variables for predictive/diagnostic models used in silico environmental chemical science, in terms of system scales and variable types. Figure 3 shows a continuum of system scales encompassing the whole scope of predictive/diagnostic modelling for in silico environmental chemical sciences, juxtaposing earth and biological scales.</p> <p>The published version of Figure 3 is tall, for two-column page-layouts, but a wide version of Figure 3 is provided for landscape oriented formats. The 300 dpi versions of each figure should be adequate resolution for most purposes, and therefore are recommended. The large versions of the figures may take significant time to download, but may be useful for high resolution applications.</p> <p>This work is from the perspectives/review paper at the beginning of a themed issue on "Quantitative Structure-Activity Relationships (QSARs) and Computational Chemistry Methods in the Environmental Chemical Sciences", published in the March 2017 issue of the Royal Society of Chemistry journal Environmental Sciences: Process and Impacts. The whole collection of papers can be accessed at rsc.li/qsars.</p>
Development and validation of statistical shape models of the primary functional bone segments of the foot.
<p>This dataset comprises manually segmented three-dimensional point clouds (.STL) of magnetic resonance images of the primary functional segments of the foot - first metatarsal, midfoot (second-to-fifth metatarsals, cuneiforms, cuboid, and navicular), calcaneus, and talus. These data were used to create statistical shape models of the foot bones, utilising the GIAS2 toolbox (https://pypi.org/project/gias2/).</p>
Combining statistical and mechanistic models to unravel the drivers of mortality within a rear-edge beech population - Supporting Material
<p>Supporting material for the study:</p> <p><strong>"Combining statistical and mechanistic models to unravel the drivers of mortality within a rear-edge beech population."</strong></p> <p><strong>Authors:</strong></p> <p>Cathleen Petit-Cailleux1, Hendrik Davi1, François Lefèvre1, Joseph Garrigue<strong>2</strong>, Jean-André Magdalou<strong>2</strong>, Christophe Hurson<strong>2,3</strong><strong>, </strong>Elodie Magnanou<strong>2,4</strong>, and Sylvie Oddou-Muratorio1.</p> <p> </p> <p>Adresses</p> <p>1INRA, UR 629 Ecologie des Forêts Méditerranéennes, URFM, Avignon, France</p> <p><strong>2</strong>Réserve Naturelle Nationale de la Forêt de la Massane, France</p> <p><strong>3</strong>Fédération des Réserves Naturelles Catalanes, Prades, France</p> <p><strong>4</strong>Sorbonne Université, CNRS, Biologie Intégrative des Organismes Marins, BIOM, F-66650 Banyuls-sur-Mer, France</p> <p><strong>ORCID:</strong></p> <p>Cathleen Petit-Cailleux: <a href="https://orcid.org/0000-0001-7714-6583">https://orcid.org/0000-0001-7714-6583</a></p> <p>François Lefèvre : <a href="https://orcid.org/0000-0003-2242-7251">https://orcid.org/0000-0003-2242-7251</a></p> <p>Sylvie Oddou-Muratorio <a href="https://orcid.org/0000-0003-2374-8313">https://orcid.org/0000-0003-2374-8313</a></p> <p> </p> <p>-------------</p> <p>Raw data of the Table_Massane_moratlity_trees.csv and climate can be obtained from Joseph Garrigue, Jean-André Magdalou and Christophe Hurson.</p> <p>The inventories files and daily climate are the input dataset to run CASTANEA models.</p> <p>All details are provided in the article.</p>
A Bi-atrial Statistical Shape Model and 100 Volumetric Anatomical Models of the Atria
<p>This dataset is part of the publication "A bi-atrial statistical shape model for large-scale in silico studies of human atria: Model development and application to ECG simulations" by Nagel et al. (<a href="https://doi.org/10.1016/j.media.2021.102210">https://doi.org/10.1016/j.media.2021.102210</a>). It includes a bi-atrial statistical shape model built based on 47 MR and CT images (Left atrium segmentation challenge (Tobon-Gomez, 2015), Left atrium fibrosis and scar segmentation challenge (Karim, 2013), Left atrial wall thickness challenge (Karim, 2018)). ScalismoLab (https://scalismo.org) was used for parts of the model generation. Further Details are explained in the paper. The SSM is available as an h5 file including information about the mean shape's vertex locations and their triangulation as well as the eigenvectors and -values. </p> <p>100 random instances derived from the model are available. Each zip file contains the volumetric bi-atrial geometry as vtk file, which was augmented in a post-processing step with a homogeneous wall thickness, fiber orientation, intra-atrial bridges and material tags so that they are ready to use for electrophysiological simulations of atrial signals. Furthermore, the scalar field resulting from computing the gradient of the Laplace equation with the boundary conditions described by Piersanti et al. (Modeling cardiac muscle fibers in ventricular and atrial electrophysiology simulations, Computer Methods in Applied Mechanics and Engineering, 2020, <a href="https://doi.org/10.1016/j.cma.2020.113468">https://doi.org/10.1016/j.cma.2020.113468</a>) are available on the left and the right atrial instances. </p> <p>Furthermore, 95 geometries with uniformly distributed left atrial volumes are available in LAE_geometries.zip. </p>
SUMMA/mizuRoute model configurations, parameters, and ensemble statistics for representative cryosphere basins
<p>Meteorological forcing is a major source of uncertainty in hydrological modeling. The recent development of probabilistic large-domain meteorological datasets enables convenient uncertainty characterization, which however is rarely explored in large-domain research. Tang et al. (2023) analyze how uncertainties in meteorological forcing data affect hydrological modeling in 289 representative cryosphere basins by forcing the Structure for Unifying Multiple Modeling Alternatives (SUMMA) and mizuRoute models with precipitation and air temperature ensembles from the Ensemble Meteorological Dataset for Planet Earth (EM-Earth). EM-Earth probabilistic estimates are used in ensemble simulation for uncertainty analysis. The results reveal the magnitude, spatial distribution, and scale effect of uncertainties in meteorological, snow, runoff, soil water, and energy variables.</p>
A computationally efficient statistically downscaled 100 m resolution Greenland product from the regional climate model MAR: accompanying dataset
<p>Dataset containing surface temperature and surface mass balance datasets generated from the MAR regional climate model over Greenland over two test areas using statistical downscaling tools from 6 km to 100m. The abstract of the accompanying submitted paper follows: </p> <p> </p> <p>The Greenland Ice Sheet (GrIS) has been contributing directly to sea level rise and this contribution is projected to accelerate over next decades. A crucial tool for studying the evolution surface mass loss (e.g., surface mass balance, SMB) consists of regional climate models (RCMs) which can provide current estimates and future projections of sea level rise associated with such losses. However, one of the main limitations of RCMs is the relatively coarse horizontal spatial resolution at which outputs are currently generated. Here, we report results concerning the statistical downscaling of the SMB modeled by the Modèle Atmosphérique Régional (MAR) RCM from the original spatial resolution of 6 km to 100 m building on the relationship between elevation and mass losses in Greenland. To this goal, we developed a geospatial framework that allows the parallelization of the downscaling process, a crucial aspect to increase the computational efficiency of the algorithm. The results obtained in the case of the SMB, assessed through the comparison of the modeled outputs with in-situ SMB measurements, show a considerable improvement in the case of the downscaled product with respect to the original, coarse output. In the case of the downscaled MAR product, the coefficient of determination (R<sup>2</sup>) increases from 0.868 for the original MAR output to 0.935 for the downscaled product. Moreover, the value of the slope and intercept of the linear regression fitting modeled and measured SMB values shifts from 0.865 for the original MAR to 1.015 for the downscaled product in the case of the intercept and from the value -235mm (original) to -57 mm (downscaled) in the case of the slope, considerably improving upon results previously published in the literature.</p>
Statistical blending of global-gridded climatological products: an approach to inverse hydrological model
<p>The growing use of global-scale environmental products in hydro-climatic modeling (with different assumptions, resolutions, and precisions) has increased the variety of their applications and the complications of their uncertainties and evaluations. Researchers have recently turned to statistical blending (fusion) of these products to achieve optimal modeling while avoiding difficulties. The proposed statistical blending in this study includes five large-scale and satellite precipitation (Climate Hazards Group Infrared Precipitation with Stations (CHIRPS), ERA5-Land of ECMWF (ERA), Integrated Multi-Satellite Retrievals for GPM (IMERG), Tropical Rainfall Measuring Mission (TRMM), and Terra) and evapotranspiration (Global Land Evaporation Amsterdam Model (GLEAM), SSEBop, Moderate Resolution Imaging Spectroradiometer (MODIS), Terra, and ERA) products committed in three modeling scenarios. The blending procedures are organized using a conceptual water balance model to achieve the best precipitation and evapotranspiration results for the conceptual production of streamflow using hydrological inverse modeling. Based on the results, the proposed blending procedures of precipitation and evapotranspiration improved the performance of the model using different statistical metrics. In addition, the results show the conformity of the pattern and behavior of the blended precipitation calculated using the moving least square method in the study area. This happened by changing the estimation based on <em>in situ</em> values, particularly in cold months considering the orographic/snow effects. The combining method provides a good fusion procedure to improve the realistic estimation of precipitation and evapotranspiration in ungagged watersheds as well<strong>.</strong></p>
Dataset for "Topography-based statistical modelling reveals high spatial variability and seasonal emission patches in forest floor methane flux"
<p>This dataset provides measured and upscaled forest floor methane (CH4) fluxes and soil moisture.</p> <p>This dataset is related to the following manuscript:</p> <p>Vainio et al., Topography-based statistical modelling reveals high spatial variability and seasonal emission patches in forest floor methane flux, Biogeosciences, in review. (The discussion preprint is available at https://doi.org/10.5194/bg-2020-263.)</p>
A statistical shape model of craniosynostosis patients and 100 model instances of each pathology
<p>This dataset is part of the publication "A statistical shape model for radiation-free assessment and classification of craniosynostosis" (M. Schaufelberger et al.). It includes several 3D head models constructed of surface scans of craniosynostosis patients: The full shape model, a texture model, and submodels of four classes: sagittal suture fusion (scaphocephaly), metopic suture fusion (trigonocephaly), coronal suture fusion (brachycephaly and anterior plagiocephaly), and a control model (normocephaly and positional plagiocephaly). Each of the models is available in an .h5 file. We also include 100 mesh instances as a .ply file in a zip file. The model's statistical information can be incorporated into the [Liverpool-York child head model (Dai et al. 2019)](https://doi.org/10.1007/s11263-019-01260-7) as it uses the same vertex order and IDs (starting from index 0). If you want to synthesize new models, take a look a the demo.py file. For information about the hierarchy in the h5-file, take a look at documentation.md.</p>
Data for fitting a statistical global burned area model for seamless integration into Dynamic Global Vegetation Models
<p>The dataset is a large R data.table object saved in RDS format. It contains global, monthly data spanning the period from 2002 to 2018, with a 0.5 degrees spatial resolution. The dataset is utilized to develop and validate statistical models for predicting global burnt areas resulting from wildfires.</p>
Capturing features of hourly-resolution energy models through statistical annual indicators
<p>Dear colleagues,</p> <p>This is the official repository of the Task 7.4 of H2020 Locomotion project. Feel free to use our data by citing this work and comment about our work by referencing the main authors of it. The article explaining this work is under revision. it will be referenced as soon as posible.</p> <p><strong>Python scripts </strong>("create_inputs.txt" and "run_simulations.txt") creates the input files for EnergyPLAN. The second one runs iteratively EnergyPLAN to generate the outputs of combinations (which are saved in the "EU_Iterate_case.xlsx" file). Hourly distributions of demands and supply technologies are contained in the RAR file ("EUdist.rar") and "EU_start_v2_noFlex.txt" initialize the starting configuration of the European energy system. Those files are required to run EnergyPLAN. The <strong>PowerPoint file</strong> ("EnergyPLAN_instructions.pptx") explains the procedure to carry out the runs of combinations in Python/Excel.</p> <p>In case you couldn't properly do the combinations, the<strong> Excel file</strong> ("EU.xlsx") saves this information, so the steps of the approach could be followed from this point with the Excel file. We have used Power Query (Excel) to prepare the data for the next step of building the regression models.</p> <p>The <strong>Matlab file (</strong>"CreateRegressionModels.m"<strong>)</strong> automatically generates the regression models for the European region of WILIAM (official model of the Locomotion project).</p> <p>Best regards,</p> <p>Gonzalo.</p>
Chronogram or phylogram for ancestral state estimation? Model-fit statistics indicate the branch lengths underlying a binary character's evolution: R scripts and simulated trees
<p>All R scripts used in this study, and the set of simulated phylogenetic trees used in the study.</p> <p>1. Modern methods of ancestral state estimation (ASE) incorporate branch length information, and it has been demonstrated that ASEs are more accurate when conducted on the branch lengths most correlated with a character's evolution; however, a reliable method for choosing between alternate branch length sets for discrete characters has not yet been proposed.<br><br>2. In this study, we simulate paired chronograms and phylograms, and generate binary characters that evolve in correlation with one of these. We then investigate (1) the effect of alternate branch lengths on ASE error, and (2) whether phylogenetic signal statistics and/or model-fit statistic can be used to select the branch lengths most correlated with a binary character.<br><br>3. In agreement with previous studies, we find that ASEs are more accurate when conducted on the branch lengths most correlated with the character. Phylogenetic signal statistics show limited utility for selecting the correct branch lengths, but model-fit statistics are found to be more accurate, with the correct branch lengths generally returning greater model-fit (lower AICc and BIC values). Using this method to choose between alternate branch length sets is more accurate when tree and character properties are more favorable for model optimization, and when shape differences between alternate phylogenies are greater.<br><br>4. Our results indicate that researchers conducting ASEs on discrete characters should carefully consider which branch lengths are appropriate, and, in the absence of other evidence, we suggest estimating model-fit values over alternate branch length sets and evolutionary models and choosing the branch length/model combination that returns better model fit.</p>
CONTENT -- Multi-context genetic modeling TWAS and eAssociation summary statistics
<p>We provide the summary statistics of running CONTENT, the context-by-context approach, and UTMOST on over 22 phenotypes. The phenotypes are listed in the manuscript, and their respective studies and sample size can be found in a table under the supplementary section of the manuscript. All 3 methods were trained on GTEx v7 as well as CLUES, a single-cell RNA sequencing dataset of PBMCs. The data include the gene name, model, cross-validated R^2, prediction pvalue, TWAS p value, TWAS Z score, and a column titled "hFDR" indicating whether the association was statistically significant while employing hierarchical FDR. The benefits of employing such an approach for all methods can be found in the manuscript.</p> <p> </p> <p>We also include the eAssociations that we obtain by training prediction models on GTEx and CLUES alone. For the CxC and UTMOST approaches, these files contain the gene, context, pvalue and adjusted R^2. For CONTENT, these include the gene, context and pvalue and adjusted R^2 for each CONTENT model--the column names are described like a regression of y~x, rsq_y_x, so rsq_observed_full is the adjusted R^2 from regressing the observed expression onto the cross-validated full model predictions. In cases where the R^2 is higher from the specific or shared models, it's best to use either of those rather than the full model for out of sample prediction.</p>
Data associated with the manuscript "Simple statistical models can be sufficient for testing hypotheses with population time series data"
<p>This is a revised version of the archive of R code and data used in the manuscript, <em>Simple statistical models can be sufficient for testing hypotheses with population time series data. </em>The data are in three files. <em>etodata1.csv</em> and <em>etodata2.csv</em> contain two versions of the same data for shoal-dwelling fishes in the Etowah River and associated environmental covariates. <em>knz_dat</em> contains data for small mammals collected in the Konza Prairie Biological Station and associated environmental covariates. The R code consists of four primary files that call nine auxiliary files. CaseStudy1-main_code and CaseStudy2-main_code are the primary files for running the two case studies. Simulations1 and Simulations2 are the files for running the two batteries of simulations. We thank the Konza Prairie Biological Station and Konza Prairie Long-Term Ecological Research Program supported by the National Science Foundation (DEB-1440484) for collecting and providing access to mammal community data. More details are in the manuscript and supporting information. </p>
Yearly CDIP:MOP-alongshore modeled wave statistics for California, January 2000 - July 2022
<p><strong>Overview</strong></p> <ul> <li>Yearly wave averages for all 11,594 CDIPS-MOPS alongshore sites in California.</li> <li>Sites are defined in the files "CDIP_Transects.csv" and "CDIP_Transects.geojson". The bounds of each site are listed in the file "CA_region_bounds.csv"</li> <li>Data are described here: https://cdip.ucsd.edu/documents/index/product_docs/mops/mop_intro.html</li> <li>Data are obtained from here: https://thredds.cdip.ucsd.edu/thredds/catalog.html</li> </ul> <p><strong>Methods</strong></p> <ul> <li>Data are computed from hourly inshore wave hindcasts and nowcasts. Data download script is the file "CDIP_MassDownloader.ipynb"</li> <li>wave summary statistics have been computed using the file "Create_stats.ipynb". All yearly data are simple averages (i.e. mean values) of the hourly data</li> </ul> <p><strong>Data files</strong><br> Data have been split into 25 regions, defined in "CA_regions.json"</p> <p>Data are provided in geoJSON format, in the form of one file per region, and one file for all regions</p> <p><strong>Data fields</strong></p> <ul> <li>Hs: significant wave height [meters]</li> <li>Tp: peak wave period [seconds]</li> <li>Ta: average wave period [seconds]</li> <li>Dp: peak wave direction [degrees]</li> <li>Da: average wave direction [degrees]</li> <li>Ea: wave energy density, averaged over wave frequencies</li> <li>Es: wave energy density, summed over wave frequencies</li> <li>QC: quality flag</li> <li>waveTime: UTC time string</li> <li>metaWaterDepth: water depth of modeled wave data (range is 10-15m)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.