Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.9.0
Dataset results
33 results for “level statistics”
ReFAB 10,000 Year Statistical Estimate of Aboveground Woody Biomass, Midwest US, Level 2
How terrestrial biomass changed before the advent of industrial society is a major gap in our understanding of the Earth's carbon cycle. Here, we archive data used to reconstruct 10,000 years of aboveground woody biomass across the US Upper Midwest using statistical models based on historical forest surveys and fossil pollen assemblages. From our analyses we document a 5,000 year long carbon sink into vegetation, primarily caused by the range expansion of two late-successional species into the region during the late Holocene. The importance of such large slow-growing tree species in storing carbon during the pre-industrial past argues for protecting similar species in wild forests today. This material is based upon work supported by the National Science Foundation under grants #DEB-1241874, 1241868, 1241870, 1241851, 1241891, 1241846, 1241856, 1241930.
Voxel-level summary statistics of hippocampus shape, white matter microstructure, and cortical surface curvature in UK Biobank (n=33,324)
<p>This deposit hosts GWAS summary statistics of hippocampus shape (n=33,324), white matter microstructure (n=33,324), and cortical surface curvature (n=15,752) using UKB unrelated white subjects. The data was generated by using the highly efficient imaging genetics (<a href="https://github.com/Zhiwen-Owen-Jiang/heig">HEIG v1.1.0</a>) framework where only the triplets - summary statistics of low-dimensional representations (LDRs), the functional bases, and the variance-covariance matrix LDRs - are shared, which is sufficient to recover all voxel-variant pairs as well as to conduct voxel-level heritability and (cross-trait) genetic correlation analysis. Check the <a href="https://github.com/Zhiwen-Owen-Jiang/heig/wiki">tutorial</a> and the <a href="../records/13770930">example data</a> used in the tutorial. </p> <p>The shared data includes:</p> <p>1. Triplets for hippocampus shape measured by the radial distance from the medial model for each vertex. The original images contain 30,000 vertices while the shared data contains 49 LDRs. Left and right hemispheres were analyzed separately, each with 15,000 vertices.</p> <p>2. Triplets for 21 white matter tracts measured by fractional anisotropy. The original images contain 32,217 voxels and each tract contains 88 ~ 3503 voxels while the shared data contains 1,034 LDRs. Tracts were analyzed separately.</p> <p>3. Triplets for cortical surface curvature. The original images contain 59,412 vertices while the shared data contains 1,750 LDRs. The entire brain was analyzed as a whole.</p> <p>4. LD matrix and its inverse for 22 chromosomes including 460k genotyped SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 8.4k white unrelated subjects in UKB. Two regularization levels are provided: {85%, 80%} for heritability and genetic correlations within images and {75%, 70%} for cross-trait genetic correlations.</p> <p>5. LD matrix and its inverse for 22 chromosomes including 1.2 million imputed HapMap3 SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 42k white unrelated subjects in UKB. Two regularization levels are provided: {98%, 95%} for heritability and genetic correlations within images and {90%, 85%} for cross-trait genetic correlations.</p>
First Street Foundation Property Level Flood Risk Statistics V1.3
<p>The property level flood risk statistics generated by the First Street Foundation Flood Model Version 1.3 come in CSV format. The data that is included in the CSV includes:</p> <ul> <li> <p>An FSID; a First Street ID (FSID) is a unique identifier assigned to each location.</p> </li> <li> <p>The latitude and longitude of a parcel as well as the zip code, census block group, census tract, county, congressional district, and state of a given parcel.</p> </li> <li> <p>The property’s Flood Factor as well as data on economic loss.</p> </li> <li> <p>The flood depth in centimeters at the low, medium, and high CMIP 4.5 climate scenarios for the 2, 5, 20, 100, and 500 year storms in 2021, 2036, and 2051.</p> </li> <li> <p>Data on the cumulative probability of a flood event exceeding the 0cm, 15cm, and 30cm threshold depth is provided at the low, medium, and high climate scenarios for years 2021, 2036, and 2051.</p> </li> <li> <p>Information on historical events and flood adaptation, such as ID and name.</p> </li> </ul> <p>You can download a sample of the property level flood risk statistics generated by First Street's Flood Model on this page. You can purchase the property level data for areas within the contiguous United States on the First Street website <a href="https://firststreet.org/data-access/paid-access/?utm_source=Property_Statistics&utm_medium=Purchase_Data&utm_campaign=Zenodo#pricing-component">here</a>. You can find the data dictionary which breaks down the data that is available with each property-level data purchase <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/data-dictionary/?utm_source=Property_Statistics&utm_medium=Data_Dictionary&utm_campaign=Zenodo">here</a>. If you are also interested in the hazard layers, you can find more information <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-hazard-dictionary/?utm_source=Property_Statistics&utm_medium=Hazard_Dictionary&utm_campaign=Zenodo">here</a>.</p>
First Street Foundation Property Level Flood Risk Statistics V2.0
<p>The property level flood risk statistics generated by the First Street Foundation Flood Model Version 2.0 come in CSV format. </p> <p>The data that is included in the CSV includes:</p> <ul> <li> <p>An FSID; a First Street ID (FSID) is a unique identifier assigned to each location.</p> </li> <li> <p>The latitude and longitude of a parcel as well as the zip code, census block group, census tract, county, congressional district, and state of a given parcel.</p> </li> <li> <p>The property’s Flood Factor as well as data on economic loss.</p> </li> <li> <p>The flood depth in centimeters at the low, medium, and high CMIP 4.5 climate scenarios for the 2, 5, 20, 100, and 500 year storms this year and in 30 years.</p> </li> <li> <p>Data on the cumulative probability of a flood event exceeding the 0cm, 15cm, and 30cm threshold depth is provided at the low, medium, and high climate scenarios for this year and in 30 years.</p> </li> <li> <p>Information on historical events and flood adaptation, such as ID and name.</p> </li> </ul> <p> </p> <p>This dataset includes <a href="https://firststreet.org/">First Street</a>'s aggregated flood risk summary statistics. The data is available in CSV format and is aggregated at the congressional district, county, and zip code level. The data allows you to compare FSF data with FEMA data. You can also view aggregated flood risk statistics for various modeled return periods (5-, 100-, and 500-year) and see how risk changes due to climate change (compare FSF 2020 and 2050 data). There are various <a href="https://floodfactor.com/">Flood Factor</a> risk score aggregations available including the average risk score for all properties (flood factor risk scores 1-10) and the average risk score for properties with risk (i.e. flood factor risk scores of 2 or greater). This is version 2.0 of the data and it covers the 50 United States and Puerto Rico. There will be updated versions to follow.</p> <p>If you are interested in acquiring First Street flood data, you can request to access the data <a href="https://firststreet.org/data-access/paid-access/?utm_source=Summary_Statistics_v1.3&utm_medium=Purchase_Data&utm_campaign=Zenodo#pricing-component">here</a>. More information on First Street's flood risk statistics can be found <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-data-dictionaryv2/">here</a> and information on First Street's hazards can be found <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-hazard-dictionary/?utm_source=Summary_Statistics_v1.3&utm_medium=Hazard_Dictionary&utm_campaign=Zenodo">here</a>.</p> <p>The data dictionary for the parcel-level data is below.</p> <table> <tbody> <tr> <td> <p><strong>Field Name</strong></p> </td> <td> <p><strong>Type</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>fsid</p> </td> <td> <p>int</p> </td> <td> <p>First Street ID (FSID) is a unique identifier assigned to each location</p> </td> </tr> <tr> <td> <p>long</p> </td> <td> <p>float</p> </td> <td> <p>Longitude</p> </td> </tr> <tr> <td> <p>lat</p> </td> <td> <p>float</p> </td> <td> <p>Latitude</p> </td> </tr> <tr> <td> <p>zcta</p> </td> <td> <p>int</p> </td> <td> <p>ZIP code tabulation area as provided by the US Census Bureau</p> </td> </tr> <tr> <td> <p>blkgrp_fips</p> </td> <td> <p>int</p> </td> <td> <p>US Census Block Group FIPS Code</p> </td> </tr> <tr> <td> <p>tract_fips</p> </td> <td> <p>int</p> </td> <td> <p>US Census Tract FIPS Code</p> </td> </tr> <tr> <td> <p>county_fips</p> </td> <td> <p>int</p> </td> <td> <p>County FIPS Code</p> </td> </tr> <tr> <td> <p>cd_fips</p> </td> <td> <p>int</p> </td> <td> <p>Congressional District FIPS Code for the 116th Congress</p> </td> </tr> <tr> <td> <p>state_fips</p> </td> <td> <p>int</p> </td> <td> <p>State FIPS Code</p> </td> </tr> <tr> <td> <p>floodfactor</p> </td> <td> <p>int</p> </td> <td> <p>The property's Flood Factor, a numeric integer from 1-10 (where 1 = minimal and 10 = extreme) based on flooding risk to the building footprint. Flood risk is defined as a combination of cumulative risk over 30 years and flood depth. Flood depth is calculated at the lowest elevation of the building footprint (largest if more than 1 exists, or property centroid where footprint does not exist)</p> </td> </tr> <tr> <td> <p>CS_depth_RP_YY</p> </td> <td> <p>int</p> </td> <td> <p>Climate Scenario (low, medium or high) by Flood depth (in cm) for the Return Period (2, 5, 20, 100 or 500) and Year (today or 30 years in the future). Today as year00 and 30 years as year30. ex: low_depth_002_year00</p> </td> </tr> <tr> <td> <p>CS_chance_flood_YY</p> </td> <td> <p>float</p> </td> <td> <p>Climate Scenario (low, medium or high) by Cumulative probability (percent) of at least one flooding event that exceeds the threshold at a threshold flooding depth in cm (0, 15, 30) for the year (today or 30 years in the future). Today as year00 and 30 years as year30. ex: low_chance_00_year00</p> </td> </tr> <tr> <td> <p>aal_YY_CS</p> </td> <td> <p>int</p> </td> <td> <p>The annualized economic damage estimate to the building structure from flooding by Year (today or 30 years in the future) by Climate Scenario (low, medium, high). Today as year00 and 30 years as year30. ex: aal_year00_low</p> </td> </tr> <tr> <td> <p>hist1_id</p> </td> <td> <p>int</p> </td> <td> <p>A unique First Street identifier assigned to a historic storm event modeled by First Street</p> </td> </tr> <tr> <td> <p>hist1_event</p> </td> <td> <p>string</p> </td> <td> <p>Short name of the modeled historic event</p> </td> </tr> <tr> <td> <p>hist1_year</p> </td> <td> <p>int</p> </td> <td> <p>Year the modeled historic event occurred</p> </td> </tr> <tr> <td> <p>hist1_depth</p> </td> <td> <p>int</p> </td> <td> <p>Depth (in cm) of flooding to the building from this historic event</p> </td> </tr> <tr> <td> <p>hist2_id</p> </td> <td> <p>int</p> </td> <td> <p>A unique First Street identifier assigned to a historic storm event modeled by First Street</p> </td> </tr> <tr> <td> <p>hist2_event</p> </td> <td> <p>string</p> </td> <td> <p>Short name of the modeled historic event</p> </td> </tr> <tr> <td> <p>hist2_year</p> </td> <td> <p>int</p> </td> <td> <p>Year the modeled historic event occurred</p> </td> </tr> <tr> <td> <p>hist2_depth</p> </td> <td> <p>int</p> </td> <td> <p>Depth (in cm) of flooding to the building from this historic event</p> </td> </tr> <tr> <td> <p>adapt_id</p> </td> <td> <p>int</p> </td> <td> <p>A unique First Street identifier assigned to each adaptation project</p> </td> </tr> <tr> <td> <p>adapt_name</p> </td> <td> <p>string</p> </td> <td> <p>Name of adaptation project</p> </td> </tr> <tr> <td> <p>adapt_rp</p> </td> <td> <p>int</p> </td> <td> <p>Return period of flood event structure provides protection for when applicable</p> </td> </tr> <tr> <td> <p>adapt_type</p> </td> <td> <p>string</p> </td> <td> <p>Specific flood adaptation structure type (can be one of many structures associated with a project)</p> </td> </tr> <tr> <td> <p>fema_zone</p> </td> <td> <p>string</p> </td> <td> <p>Specific FEMA zone categorization of the property ex: A, AE, V. Zones beginning with "A" or "V" are inside the Special Flood Hazard Area which indicates high risk and flood insurance is required for structures with mortgages from federally regulated or insured lenders</p> </td> </tr> <tr> <td> <p>footprint_flag</p> </td> <td> <p>int</p> </td> <td> <p>Statistics for the property are calculated at the centroid of the building footprint (1) or at the centroid of the parcel (0)</p> </td> </tr> </tbody> </table> <p> </p>
Summary Statistics from "Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2"
<p>GWAMA summary statistics of PCSK9 levels using fixed-effect model. Genome-wide data is given for Europeans with statin adjustment and Europeans without statin treatment only (subset of the population). In addition, locus-wide data of the PCSK9 gene locus for African-Americans without statin treatment is listed.</p> <p>When using this data, please cite: Pott J, Gadin J, Theusch E, et al.. Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2. Hum Mol Genet. 2021 Sep 30:ddab279. doi: 10.1093/hmg/ddab279. PMID: 34590679</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>
Australian Statistical-Area (SA) Level Regions and Census Income Data (2011)
<p>The Australian Statistical Geography Standard (ASGS) defines a series of nested geographical areas in Australia known as Statistical Area (SA) Levels. SA3 regions are aggregations of SA2 regions, and SA2 regions are aggregations of SA1 regions. This data set contains the shapefiles of all SA1, SA2, and SA3 regions across Australia at the time of the 2011 census, originally downloaded from the Australian Bureau of Statistics (<a href="https://www.abs.gov.au/AUSSTATS/abs@.nsf/DetailsPage/1270.0.55.001July\%20201.">ABS</a>).</p><p>This data set also contains income information from the 2011 census, at the SA1 and SA2 level in New South Wales (NSW). Specifically, it contains the number of families of various types within a range of weekly income brackets.</p><p>Sainsbury-Dale et al. (2023) used a subset of this data set in a study on poverty levels in an area of (NSW) surrounding Sydney. </p><p> </p><p><strong>References</strong></p><p>Sainsbury-Dale, M., Zammit-Mangion, A., and Cressie, N. (2023) "Modelling Big, Heterogeneous, Non-Gaussian Spatial and Spatio-Temporal Data using FRK", <i>Journal of Statistical Software</i>, to appear.</p>
Supporting data for the AI education publication statistics in "An Experience Report of Executive-Level Artificial Intelligence Education in the United Arab Emirates"
<p>Supporting data for the AI education publication statistics presented in the paper "An Experience Report of Executive-Level Artificial Intelligence Education in the United Arab Emirates" to be published at the Twelfth AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-22). The data was used to plot the figure showing the cumulative number of publications from 1976 to 2020 relating to AI education.</p>
Data, stimuli, and analyses for "High-level aftereffects reveal the role of statistical features in visual shape encoding"
<h2><strong>Data and code share.</strong></h2><p>This record contains data and code (written in MATLAB) to reproduce the results shown in:</p><p>Morgenstern, Y. , Storrs, K., R., Schmidt, F., Hartmann, F., Tiedemann, H., Tiedemann, H, ., Wagemans, J., & Fleming, R. (<i>in press</i>) . High-level aftereffects reveal the role of statistical features in visual shape coding. Current Biology</p><p>Below is a summary of shared scripts that load data, run the analysis (including options for fitting model parameters or loading pre-computed fitted parameters), and plotting the results.</p><h3><strong>Figure 1</strong></h3><ul><li><i>Fig1C_shapespace.m</i>: draw shape space (as in Figure 1C)</li><li><i>Fig1EFG_plotpsychometricdata.m</i>: fit psychometric model to pooled data and plot (as in Figure 1EFG)</li><li><i>getExptShapesHbias.m</i>: saves a data structure (which we call 'package') with adaptor, test, and human biases from experiment 1. (Used to fit models; e.g., see <i>fitGabPyr2Hbais.m</i> or <i>fitTAEGANfit2Hbais.m</i>)</li></ul><h3><strong>Figure 2 and 3A</strong></h3><ul><li><i>fig3A_modeval_expt1.m</i>: generate figure that evaluates models in Figure 3A on how well they predict aftereffects in Experiment 1. (script located in the 'Figures 2 and 3A/models' directory).</li></ul><p>Code to fit the models, and figures that show examples of model predictions are in the model directories, and summarized below:</p><h4>Model: <strong>GabPyrAE</strong> </h4><ul><li><i><strong>note: </strong></i>To run GabPyr, you will likely need to recompile the .mex files in 'matlabPyrTools/mex'. Then move the recompiled files into the 'matlabPyrToos' directory</li><li><i>fig2B_GabPyrAEVisFigs.m</i>: produce GabPyrAE model images for example adaptor and test image ( as in Figure 2B )</li><li><i>figS2BC_GabPyrAEExp</i>.m: get GabPyrAE model responses to simulated tilt aftereffect experiment using anisotropic noise, and plot model responses as in Figures S2BC.</li><li><i>getGabPyrAEMod.m: </i>given an adaptor and test image, this function produces the unfit GabPyrAE prediction</li><li><i>fitGabPyr2Hbias</i>.m: fit GabPyrAE model to best predict human baises in Experiment 1 using data structure from <i>getExptShapesHbias.m</i>. This is the general function that calls on MATLAB's GA algorithm to minimize the error function in <i>fitGabPyrNormMod2Stims</i>.m</li><li><i>evalGabPyrAE_fitmod_aic.m</i>: evaluate GabPyrAE fitted model on how well it predicts human biases from experiment 1.</li><li><i>evalGabPyrAE_unfitmod_aic.m</i>: evaluate GabPyrAE fitted model on how well it predicts human biases from experiment 1</li></ul><h4>Model: <strong>TAE</strong></h4><ul><li><i>fig2CD_TAEmod.m</i>: produce TAE model images for example adaptor and test image (as in Figure 2CD)</li><li><i>getTAEModonShape.m: </i>given an adaptor and test image, this function produces TAE Original prediction. Input to function is adaptor and test shapes, and TAE model parameters alpha and sigma.</li><li><i>getTAEGAN_spwt_onShape.m: </i>given an adaptor and test image, this function produces TAEGAN prediction. Input to function is adaptor and test shapes, and TAEGAN model parameters which include a constant term, alpha and sigma, as well as Gaussian pooling parameter that determines how much TAE to incorporate on the test or mean shapes from neighbouring adaptor line segments.</li><li><i>getTAE_spwt_onShape_nn.m: </i>given an adaptor and test image, this function produces TAE nearest neighbour prediction. Input to function is adaptor and test shapes, and TAE nearest neighbour model parameters which include a constant term, alpha and sigma, as well as Gaussian pooling parameter that determines how much TAE to incorporate on the test or mean shapes from neighbouring adaptor line segments.</li><li><i>fitTAEGAN2Hbias</i>.m: fit TAEGAN model to best predict human biases in Experiment 1 using data structure from <i>getExptShapesHbias.m</i>. This is the general function that calls on MATLAB's GA algorithm to minimize the error function in <i>fitTAEGANMod2Stims</i>.m</li><li><i>fitTAENN2Hbias</i>.m: fit TAE nearest neighbour model to best predict human biases in Experiment 1 using data structure from <i>getExptShapesHbias.m</i>. This is the general function that calls on MATLAB's GA algorithm to minimize the error function in <i>fitTAENNMod2Stims</i>.m.</li><li><i>evalTAEGAN_fitmod_aic.m</i>: evaluate TAEGAN fitted model on how well it predicts human biases from experiment 1.</li><li><i>evalTAENN_fitmod_aic.m</i>: evaluate TAE nearest neighbour fitted model on how well it predicts human biases from experiment 1.</li><li><i>evalTAE_unfitmod_aic.m</i>: evaluate TAE Original model on how well it predicts human biases from experiment 1</li></ul><h4>Model: <strong>PSAE</strong></h4><ul><li><i>fig2EF_PSAEmod.m</i>: produce PSAE model images for example adaptor and test image</li><li><i>getPos_spwt_ShiftononShape_io.m: </i>given an adaptor and test image, this function produces PSAEGAN prediction. Input to function is adaptor and test shapes, and PSAEGAN model parameters which include a constant term, alpha and sigma, as well as Gaussian pooling parameter that determines how much PSAE to incorporate on the test or mean shapes from neighbouring adaptor line segments.</li><li><i>getPos_spwt_ShiftonShape_io_nn.m: </i>given an adaptor and test image, this function produces PSAE nearest neighbour prediction. Input to function is adaptor and test shapes, and PSAE nearest neighbour model parameters which include a constant term, alpha and sigma, as well as Gaussian pooling parameter that determines how much PSAE to incorporate on the test or mean shapes from neighbouring adaptor line segments.</li><li><i>fitPSAEGAN2Hbias</i>.m: fit PSAEGAN model to best predict human biases in Experiment 1 using data structure from <i>getExptShapesHbias.m</i>. This is the general function that calls on MATLAB's GA algorithm to minimize the error function in <i>fitPSAEGANMod2Stims</i>.m</li><li><i>fitPSAENN2Hbias</i>.m: fit PSAE nearest neighbour model to best predict human biases in Experiment 1 using data structure from <i>getExptShapesHbias.m</i>. This is the general function that calls on MATLAB's GA algorithm to minimize the error function in <i>fitPSAENNMod2Stims</i>.m.</li><li><i>evalPSAEGAN_fitmod_aic.m</i>: evaluate PSAEGAN fitted model on how well it predicts human biases from experiment 1.</li><li><i>evalPSAENN_fitmod_aic.m</i>: evaluate PSAE nearest neighbour fitted model on how well it predicts human biases from experiment 1.</li></ul><h4>Model: <strong>ShapeComp </strong>and <strong>No Adaptation</strong></h4><ul><li><i>eval_ShapeComp _aic.m</i>: evaluate ShapeComp 1 parameter fitted model on how well it predicts human biases from experiment 1.</li><li><i>eval_NoAdaptation _aic.m</i>: evaluate model that predicts no adaptation on how well it predicts human biases from experiment 1.</li></ul><h3><strong>Figure 3BC</strong></h3><ul><li><i>fig3BC_Experiment2.m</i>: load, analyze, and plot experiment 2 data (as in Figure 3B and C).</li><li><i>figS4_Expt2_stimuli.m</i>: show adaptors (in black) and test shapes for ShapeComp (purple), PSAE fit GAN (green), and no adaptation model (white) (as in Figure S4)</li></ul><p> </p>
Data and code for: Fixed effects or random effects in statistical models? fewer than five levels of a grouping factor
<p>This is code (and simulated data from that code) to assess how sample size and the numbers of levels of random effects influence parameter estimates of fixed effects in linear mixed-effects models. </p>
Statistical and power calculations animal trial persistant antimicrobials and resistance levels
<p>This repository is containing the datasets with statistical analysis of the paper: "Extended period of selection pressure due to persistancy of antimicrobials in broilers."</p> <p><em>Power calculation</em></p> <ul> <li>Two CSV-files: Power calculation resistant isolates and Power calculation antibiotics</li> <li>One R-file: Power calculation antimicrobial persistent residues</li> </ul> <p><em>The phenotypic resistance</em></p> <ul> <li>R-file: Phenotypic resistance analysis trial, CSV-file: E. coli resistance per treatment</li> </ul> <p><em>The resistome analysis </em></p> <ul> <li>one CSV-files: Resistome workfile </li> <li>one R-file: Resistome statistical analysis</li> </ul> <p><em>Microbiome analysis</em><strong><em><br></em></strong></p> <ul> <li>Phyloseq object <ul> <li>R-file: Rarfied phyloseq object from biom file, CSV-file: metedata, biom-file: reads_trail </li> </ul> </li> <li>Alpha- and beta-diversity analysis<br> <ul> <li>R-file: Beta en alpha diversity animal trail, Rdata-file: rarefy.animaltrial</li> </ul> </li> <li>Microbial abundance analysis <ul> <li>R-file: relative abundance animal trial, Rdata-file: rarefy.animaltrial</li> </ul> </li> </ul>
Table 11. Results of statistical calculation of hydroxyproline levels analysis of variance (ANOVA) two way spss 23.00
<p>Table 11. Results of statistical calculation of hydroxyproline levels analysis of variance (ANOVA) two way spss 23.00</p> <p> </p>
Level statistics and entanglement entropy of Rydberg dressed bosons in a triple-well potential
<p>We study the signatures of quantum chaos in Rydberg dressed bosonic atoms held in a 1 triple-well potential. Dynamics of the bosons are governed by an extended Bose-Hubbard model (EBHM) where long-range nearest-neighbor and next-nearest-neighbor interactions are induced by laser coupling the ground state to Rydberg state. We analyze the level statistics of the EBHM for finite number N of atoms through numerical diagonalization. In the presence of a tilting potential, the level statistics are Poissonian distribution for weak dressed interaction. It becomes a Wigner-Dyson distribution for strong interaction, signifying the emergence of quantum chaos. A hybrid distribution is obtained when the dressed interaction is much stronger than the hopping rate. Using the Fock basis, we further calculate dynamical evolution of the entanglement entropy. The maximal value (upper bound) of the entanglement entropy is proved to depend on particle numbers in the form ln(N + 1). It is found that the maximum of the time-averaged entanglement entropy appears when the chaos is strong. The location of the maximum as a function of the dressed interaction and tilting potential is independent of atom number N.</p>
Statistical evaluation of character support reveals the instability of higher-level dinosaur phylogeny
<p>The interrelationships of the three major dinosaur clades (Theropoda, Sauropodomorpha, and Ornithischia) have come under increased scrutiny following the recovery of conflicting phylogenies by a large new character matrix and its extensively modified revision. Here, we use tools derived from recent phylogenomic studies to investigate the strength and causes of this conflict. Using maximum likelihood as an overarching framework, we examine the global support for alternative hypotheses as well as the distribution of phylogenetic signal among individual characters in both the original and rescored dataset. We find the three possible ways of resolving the relationships among the main dinosaur lineages (Saurischia, Ornithischiformes, and Ornithoscelida) to be statistically indistinguishable and supported by nearly equal numbers of characters in both matrices. While the changes made to the revised matrix increased the mean phylogenetic signal of individual characters, this amplified rather than reduced their conflict, resulting in greater sensitivity to character removal or coding changes and little overall improvement in the ability to discriminate between alternative topologies. We conclude that early dinosaur relationships are unlikely to be resolved without fundamental changes to both the quality of available datasets and the techniques used to analyze them.</p>
Data for: Intracranial entrainment reveals statistical learning across levels of abstraction
<div class="t-landing__text-wall "> <p>The following submission contains the data reported from the manuscript "Intracranial entrainment reveals statistical learning across levels of abstraction". The dataset was obtained from 8 neurosurgical patients who had intracranially implanted electrodes for seizure monitoring. Intracranial EEG (iEEG) data were recorded while the patients viewed a rapid stream of scene images. In the Category-Level Structured condition, patients viewed a series of trial-unique scene images, in which the categories of scenes were paired across repetitions (e.g., images of beach always followed by images of canyons). In the Exemplar-level Structured condition, participants viewed a sequence of 6 repeating scene images (from non-overlapping scene categories), which were paired across repetitions (e.g., image A always followed by image B). In the baseline Random condition, participants again viewed a sequence of 6 repeating scene images, but the images were presented in a random temporal order. Each image was presented for 250 ms, followed by a 250 ms inter-stimulus-interval period.</p> <p>The data presented here contain the raw iEEG data from the 8 patients during these task conditions, as well as information about the anatomical placement of their electrodes.</p> </div>
Dataset for: Estimating pregnancy rate from blubber progesterone levels of a blindly biopsied beluga population poses methodological, analytical and statistical challenges
<p class="MsoBodyText"><span>Beluga (<em>Delphinapterus leucas</em>) from the St. Lawrence Estuary, Canada, have been declining since the early 2000s, suggesting recruitment issues as a result of low fecundity, abnormal abortion rates or poor calf or juvenile survival. Pregnancy is difficult to observe in cetaceans, making the ground-truthing of pregnancy estimates in wild individuals challenging. Blubber progesterone concentrations were contrasted among 62 SLE beluga with a known reproductive state (i.e., pregnant, resting, parturient, and lactating females), that were found dead in 1997–2019. The suitability of a threshold obtained from decaying carcasses to assess reproductive state and pregnancy rate of freshly-dead or free-ranging and blindly-sampled beluga was examined using three statistical approaches and two datasets (135 freshly-harvested carcasses in Nunavik, and 65 biopsy-sampled SLE beluga). Progesterone concentrations in decaying carcasses were considerably higher in known-pregnant (mean </span><span>±</span><span> sd: </span><span>365 </span><span>±</span><span> 244 ng g<sup>-1</sup> of tissue) </span><span>than resting (</span><span>3.1 </span><span>±</span><span> 4.5 ng g<sup>-1</sup> of tissue</span><span>) or lactating (</span><span>38.4 </span><span>±</span><span> 100 ng g<sup>-1</sup> of tissue</span><span>) females. An approach based on statistical mixtures of distributions and a logistic regression was compared to the commonly-used, fixed threshold approach (here, 100 ng g<sup>-1</sup>) for discriminating pregnant from non-pregnant females. The error rate for classifying individuals of known reproductive status was the lowest for the fixed threshold and logistic regression approaches, but the mixture approach required limited <em>a priori</em> knowledge for clustering individuals of unknown pregnancy status. Mismatches in assignations occurred at lipid content <10% of sample weight. Our results emphasize the importance of reporting lipid contents and progesterone concentrations in both units (ng g<sup>-1 </sup>of tissue and ng g<sup>-1</sup> of lipid) when sample mass is low. By highlighting ways to circumvent potential biases in field sampling associated with capturability of different segments of a population, this study also enhances the usefulness of the technique for estimating pregnancy rate of free-ranging population. </span></p>
Dataset for: Estimating pregnancy rate from blubber progesterone levels of a blindly biopsied beluga population poses methodological, analytical and statistical challenges
Open the record for dataset details and reuse information.
Data for: Intracranial entrainment reveals statistical learning across levels of abstraction
Open the record for dataset details and reuse information.
Settlement-Era Temperature Statistical Model, Upper Midwest: Level 2
Scientific records of temperature and precipitation have been kept for several hundred years, but for many areas, only a shorter record exists. To understand climate change, there is a need for rigorous statistical reconstructions of paleoclimate using proxy data. Paleoclimate proxy data are often sparse, noisy, indirect measurements of the climate process of interest, making each proxy uniquely challenging to model statistically. We reconstruct spatially-explicit temperature surfaces from sparse and noisy measurements recorded at historical United States military forts and other observer stations from 1820-1894. One common method for reconstructing paleoclimate from proxy data is principal component regression (PCR). With PCR, one learns a statistical relationship between the paleoclimate proxy data and a set of climate observations that are used as patterns for potential reconstruction scenarios. We explore PCR in a Bayesian hierarchical framework, extending classical PCR in a variety of ways. First, we model the latent principal components probabilistically, accounting for measurement error in the observational data. Next, we extend our method to better accommodate outliers that occur in the proxy data. Finally, we explore alternatives to the truncation of lower order principal components using different regularization techniques. One fundamental challenge in paleoclimate reconstruction efforts is the lack of out-of-sample data for predictive validation. Cross-validation is of potential value, but is computationally expensive and potentially sensitive to outliers in sparse data scenarios. To overcome the limitations that a lack of out-of-sample records presents, we test our methods using a simulation study, applying proper scoring rules including a computationally efficient approximation to leave-one-out cross-validation using the log score to validate model performance. The result of our analysis is a spatially explicit reconstruction of spatio-temporal temperat
Steady states of Λ-type three-level systems excited by quantum light with various photon statistics in lossy cavities
<p>Dataset of the publication "Steady states of Λ-type three-level systems excited by quantum light with various photon statistics in lossy cavities" H. Rose, O. V. Tikhonova, T. Meier, and P. R. Sharapova, New J. Phys.<strong> 24</strong>, 063020 (2022). ( https://doi.org/10.1088/1367-2630/ac74d8 ). The zip file includes the data on which the plots shown in figures 2,4,5,6,7, B1, and B2 are based.</p>
Precipitation, water level and descriptive statistics
<p>Extreme weather events and the presence of mega-hydroelectric dams, when combined, present an emerging threat to natural habitats in the Amazon region. To understand the magnitude of these impacts, we used remote sensing data to assess forest loss in areas affected by the extreme 2014 flood in the entire Madeira River basin, the location of two mega-dams. In addition, forest plots (26 ha) were monitored between 2011 and 2015 (14,328 trees) in order to evaluate changes in tree mortality, aboveground biomass (AGB), species composition and community structure around the Jirau reservoir (distance between plots varies from 1 to 80 km). We showed that the mega-dams were the main driver of tree mortality in Madeira basin forests after the 2014 extreme flood. Forest loss in the areas surrounding the reservoirs was 56 km² in Santo Antônio, 190 km² in Jirau (7.4-9.2% of the forest cover before flooding), and 79.9% above that predicted in environmental impact assessments. We also show that climatic anomalies, albeit with much smaller impact than that created by the mega-dams, resulted in forest loss along different Madeira sub-basins not affected by dams (34-173 km²; 0.5-1.7%). The impact of flooding was greater in <i>várzea</i> and transitional forests, resulting in high rates of tree mortality (88-100%), AGB decrease (89-100%), and reduction of species richness (78-100%). Conversely, <i>campinarana</i> forests were more flood-tolerant with a slight decrease in species richness (6%) and similar AGB after flooding. Taking together satellite and field measurements, we estimate that the 2014 flood event in the Madeira basin resulted in 8.81-12.47 ∙ 10<sup>6</sup> tons of dead biomass. Environmental impact studies required for environmental licensing of mega-dams by governmental agencies should consider the increasing trend of climatic anomalies and the high vulnerability of different habitats to minimize the serious impacts of dams on Amazonian biodiversity and carbon stocks.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.