Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,243
datasets available to search
ShareScore release 0.7.1
Dataset results
1,243 results for “Statistics”
Statistical Process Control Benchmark Dataset
<p>Datasets to the planned publication "Generalized Statistical Process Control via 1D-ResNet Pretraining" by Tobias Schulze, Louis Huebser, Sebastian Beckschulte and Robert H. Schmitt (Chair for Intelligence in Quality Sensing, Laboratory for Machine Tools and Production Engineering, WZL of RWTH Aachen University)</p> <p>Data for benchmarking SPC against other process monitoring methods. The data consist of a one-dimensional timeseries of floats (x.csv). Addititionally information whether the data are within the specifications are provided as another time series (y.csv). The data are generated by solving an optimization problem for each time to generate a mixture distribution of different probability distributions. Then for each timestep one record is sampled. Inputs for the optimization problem are the given probability distributions, the lower and upper limit of the tolerance interval as well as the desired median of the data. Additionally weights of the different probability distributions can be given as boundary condions for the different time steps. Metadata generated from the solving are stored in k_matrix.csv (wheights at each time step) and distribs (probability distribution objects according to https://doi.org/10.5281/zenodo.8249487). The data consists of phases with data from a stable mixture distribution and phases with data from a mixture distribution that do not fulfill the stability criteria.</p> <p>The train data were used to train the G-SPC model. The test data were used for benchmarking purposes</p> <p>Funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy – EXC-2023 Internet of Production – 390621612.</p>
Electron Donor-Functionalized Pyrenes with Amplified Spontaneous Emission for Violet-Blue Electroluminescent Devices Beyond the Spin Statistical Limit
<p>Quantum Chemical TD-DFT Data on the <span>M062X-GD3/def2-TZVP level of theory. Ground state geometries, first excited state geometries, and single point calculations for donor functionalized pyrenes. </span> </p>
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Quasi-stationarity of scalar turbulent mixing statistics
<p>Dataset underlying the conclusions of the paper "Quasi-stationarity of scalar turbulent mixing statistics" to appear in Physics of Fluids.</p> <p>A copy of the article for exclusive personal use has been added. Any other use requires prior permission of author and AIP Publishing. The article appeared in <em>Physics of Fluids 33(5):055109 (2021) </em>and may be found at https://doi.org/10.1063/5.0044400</p>
Code and data to "Statistical learning and topkriging improve spatio-temporal low-flow estimation"
<p>This data and software supports the manuscript "Statistical learning and topkriging improve spatio-temporal low-flow estimation" (https:://doi.org/<span>10.1029/2024WR038329</span>).</p> <p>The dataset consists of:</p> <ul> <li>all produced predictions of the models (data/predictions.RDS and data/predictions_csv/*)</li> <li>observational data (data/observations.csv)</li> <li>additional catchment data (data/catchment_data.csv) used for presenting the figures</li> <li>state boundaries of Austria as a shape file (data/boundaries.*)</li> <li>partial predictions of a model-based boosting approach (data/partial_predictions.csv)</li> <li>Example output of number of EOF, due to long computational time (data/number_eofs.RDS)</li> <li>IDs of near natural catchments (data/ids_low_flow.csv)</li> </ul> <p>Additionally, the code is provided to:</p> <ul> <li>Compute the number of EOFs (functions/number_eofs.R)</li> <li>Produce all the figures and tables in the paper (scripts/plotting_results.R)</li> </ul>
Data for publication "Statistical characteristics of extreme daily precipitation during 1501 BCE - 1849 CE in the Community Earth System Model".
<p>Here, the data used in Kim, W. M., Blender, R., Sigl, M., Messmer, M., & Raible, C. C. (2021). "Statistical characteristics of extreme daily precipitation during 1501 BCE–1849 CE in the Community Earth System Model" in <em>Climate of the Past </em>(<a href="https://doi.org/10.5194/cp-2021-61">https://doi.org/10.5194/cp-2021-6</a>) are provided.</p> <p>Two simulations covering the period 1501 BCE - 2008 CE are performed with CESM 1.2.2: the orbital-only and the full-forcing simulations. The full-forcing transient simulation includes the new long record of volcanic eruptions (<a href="https://doi.org/10.1594/PANGAEA.928646">https://doi.org/10.1594/PANGAEA.928646</a>) that covers the last 3500 years. The output from the simulations is used to examine the long-term variability and characteristics of daily extreme precipitation during 1501BCE-1849 CE.</p> <p>The following files are provided:</p> <ul> <li> <strong>CESM122.transient.PRECT.anom.above99th.1501BCE-1849CE_I and II</strong>: Daily precipitation anomalies above the 99th percentiles relative to 1501BCE-1849CE from the full forcing simulation. The file is split into two parts, with the first file containing the first 50% of extremes (I) and the second file containing the rest 50% (II).</li> <li><strong>CESM122.orbital.PRECT.anom.above99th.1501BCE-1849CE I and II:</strong> Daily precipitation anomalies above the 99th percentiles relative to 1501BCE-1849CE from the orbital-only simulation.</li> <li> <strong>CESM122.transient.variables.mon.1979-2008CE:</strong> monthly precipitation, temperature, and geopotential height at 500 hPa for 1979-2008CE from the full-forcing simulation.</li> <li><strong>CESM122.trans.variable_names.years:</strong> Monthly variables from the full-forcing simulation. The simulation starts from the model year 1, which corresponds to the actual year 1501BCE. The variables are solar insolation (SOLIN), clear-sky net surface shortwave radiation (FSNSC), geopotential height at 500hPa (Z500), and surface temperature (TS).</li> <li><strong>CESM122.orbital.variable_names.years:</strong> Monthly variables from the orbital-only simulation. The simulation starts from the model year 1, which corresponds to the actual year 1501BCE.</li> <li><strong>CESM122.*.log-likelihood-GPDmodel-ExtForcing</strong>: Negative log-likelihood for the stationary and non-stationary Generalized Pareto Distribution models for external forcings.</li> <li> <strong>CESM122.*.log-likelihood-GPDmodel-ModesVar</strong>: Negative log-likelihood for the stationary and non-stationary Generalized Pareto Distribution models for modes of variability.</li> <li><strong>Evolk_EVA_distribution_1501BCE-2015CE</strong>: Distribution of volcanic aerosol for CAM5, produced based on Kim et al. (2021).</li> </ul> <p>If you use this dataset, please cite:</p> <p><em>Kim, W. M., Blender, R., Sigl, M., Messmer, M., & Raible, C. C. (2021). Statistical characteristics of extreme daily precipitation during 1501 BCE–1849 CE in the Community Earth System Model. Climate of the Past Discussions, 1-38. <a href="https://doi.org/10.5194/cp-2021-61">https://doi.org/10.5194/cp-2021-61</a></em></p>
Lagrangian statistics in turbulent channel flow
<p>A set of Lagrangian statistics of passive tracers in a turbulent channel flow. The particle trajectories are obtained by means of integration of a simulated flow field, computed via Direct Numerical Simulation, at four different Reynolds numbers. The Reynolds numbers here employed are <span class="math-tex">\(\mathrm{Re}_{\tau} = \frac{u_{\tau}\delta}{\nu} = 180,\,395,\,590,\,950\)</span>, where <span class="math-tex">\(u_{\tau}\)</span>is the frictional velocity, <span class="math-tex">\(\delta\)</span> is the channel half height and <span class="math-tex">\(\nu\)</span> is the kinematic viscosity.</p> <p>Additional information about the database is provided in the included documentation.</p> <p>v1.0 -> Added statistics at ReT = 950</p> <p>v1.1 -> Added statistics at ReT = [180, 395, 590]</p> <p>v1.2 -> Added statistics at ReT = 265</p>
First Street Foundation Property Level Flood Risk Statistics V1.3
<p>The property level flood risk statistics generated by the First Street Foundation Flood Model Version 1.3 come in CSV format. The data that is included in the CSV includes:</p> <ul> <li> <p>An FSID; a First Street ID (FSID) is a unique identifier assigned to each location.</p> </li> <li> <p>The latitude and longitude of a parcel as well as the zip code, census block group, census tract, county, congressional district, and state of a given parcel.</p> </li> <li> <p>The property’s Flood Factor as well as data on economic loss.</p> </li> <li> <p>The flood depth in centimeters at the low, medium, and high CMIP 4.5 climate scenarios for the 2, 5, 20, 100, and 500 year storms in 2021, 2036, and 2051.</p> </li> <li> <p>Data on the cumulative probability of a flood event exceeding the 0cm, 15cm, and 30cm threshold depth is provided at the low, medium, and high climate scenarios for years 2021, 2036, and 2051.</p> </li> <li> <p>Information on historical events and flood adaptation, such as ID and name.</p> </li> </ul> <p>You can download a sample of the property level flood risk statistics generated by First Street's Flood Model on this page. You can purchase the property level data for areas within the contiguous United States on the First Street website <a href="https://firststreet.org/data-access/paid-access/?utm_source=Property_Statistics&utm_medium=Purchase_Data&utm_campaign=Zenodo#pricing-component">here</a>. You can find the data dictionary which breaks down the data that is available with each property-level data purchase <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/data-dictionary/?utm_source=Property_Statistics&utm_medium=Data_Dictionary&utm_campaign=Zenodo">here</a>. If you are also interested in the hazard layers, you can find more information <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-hazard-dictionary/?utm_source=Property_Statistics&utm_medium=Hazard_Dictionary&utm_campaign=Zenodo">here</a>.</p>
Ensemble statistics for modelled Eddy Kinetic Energy in the Southern Ocean
<p>This dataset contains surface eddy kinetic energy over the Southern Ocean region, sourced from a 50-member ensemble of 0.25° ocean model simulations. It is used in the paper "Circumpolar variations in the chaotic nature of Southern Ocean eddy dynamics" published in Journal of Geophysical Research - Oceans.</p> <p>This dataset has been computed from the OceaniC Chaos – ImPacts, strUcture, predicTability (OCCIPUT) global ocean/sea-ice ensemble simulation. It is composed of 50 members with a horizontal resolution of 1/4° and 75 geopotential levels (<a href="http://doi.org/10.5194/gmd-10-1091-2017">Bessières et al., 2017</a>, Penduff et al., 2014). The numerical configuration is based on the version 3.5 of the NEMO model (<a href="https://www.nemo-ocean.eu/doc">Madec, 2008</a>). The 50 members were started on January 1st 1960 from a common 21-year spinup. A small stochastic perturbation is applied to the equation of state of sea water (as in <a href="https://doi.org/10.1016/j.ocemod.2013.02.004">Brankart, 2013</a>) within each member during 1960, then switched off during the rest of the simulation. This 1-year perturbation generates an ensemble spread which grows and saturates after a few months up to a few years depending on the region. The 50 members are driven through bulk formulae during the whole 1960-2015 simulation by the same realistic 6-hourly atmospheric forcing (Drakkar Forcing Set DFS5.2, Dussin et al., 2016) derived from ERA interim atmospheric reanalysis. Data is for the period 1979-2015.</p> <p>The sea level anomaly is found according to <a href="http://doi.org/10.1016/j.pocean.2020.102314">Close et al (2020)</a> and converted into surface geostrophic velocity anomaly using the geostrophic relation. This velocity field is then used to calculate the eddy kinetic energy (EKE). Data is averaged over calendar month, and restricted to the latitude range 40°-60°S. A full description of this process is included in the companion paper.</p> <p>The dataset includes EKE files (eke_0??.nc), with monthy EKE saved for the period 1979-2015 for each ensemble member, and a single file (tau.nc) for the monthly-averaged wind stress over the same period.</p>
Data and statistical analysis scripts for manuscript on wheat root response to nitrate using X-ray CT and OpenSimRoot
<p>Data and statistical analysis scripts for manuscript on wheat root response to nitrate using X-ray CT and OpenSimRoot</p> <blockquote> <p><strong>X-ray CT reveals 4D root system development and lateral root responses to nitrate in soil </strong>- [<a href="https://doi.org/10.1002/ppj2.20036">https://doi.org/10.1002/ppj2.20036</a>]</p> </blockquote> <p>The ZIP file contains:</p> <ul> <li><code>MCT1_Rcode.R</code> - Statistics script for candidate single-timepoint experiment. Requires all CSV data files in the directory. User needs to set working directory to location of this script and the CSV data files before running.</li> <li><code>MCT1... .csv</code> - 3 CSV data files required by the R script.</li> <li><code>MCT2_Rcode.R</code> - Statistics script for time-series experiment. Requires all CSV data files in the directory. User needs to set working directory to location of this script and the CSV data files before running.</li> <li><code>MCT2... .csv</code> - 3 CSV data files required by the R script.</li> <li><code>R_RooThProcessing.R</code> - R code for aggregating root traits from RooTh software.</li> <li><code>Modelling folder</code> - OpenSimRoot with model parameters and root data used in manuscript.</li> </ul>
AMS and FTIR measurements and the corresponding codes for their statistical combination
<p>This dataset includes the post-processed FTIR and AMS data for the particulate phase obtained by Yazdani et al., https://doi.org/10.5194/amt-2021-186 form wood and coal burning experiments in the PSI environmental simulation chamber. It also contains the codes for the statistical combination of AMS and FTIR measurements to estimate the high-time-resolution functional group composition of organic aerosols. </p>
KMA Mapping and alignment statistics : livestock fecal metagenomes against ResFinder and genomes
<p>Three zip archives are included used in the analysis of the European livestock resistome.</p> <p>Two of them contain 'mapstat' files produced by the KMA software using the 'extended features' flag.<br> Each mapstat file thus summarize the mapping and alignment statistics when using KMA on a metagenome against a database.</p> <p>The last archive contains the 'refdata' file used to annotate the genomic mapstat hits. It encodes the taxonomic affilication of sequences hit by one or more samples.<br> </p>
Statistical characterization of Andalusian wave climate for several combinations of Global Climate Models and Regional Climate Models and periods 2026 - 2045 and 2081 - 2100.
<p>The following text is an extract of the extended abstract entitled "<strong>Parametric Characterization of Wave Climate along the Andalusian Coast for Non-Stationary Stochastic Simulation</strong>" whose authors are Manuel Cobos, Pedro Magaña, Pedro Otiñar and Asunción Baquerizo, and that was included in proceedings of <em>39th IAHR World Congress</em> where this dataset is included.</p> <p><em>Processed data comes from PIMA Adapta Costas project (Ramírez et al., 2019), in particular, from projections of maritime climate for 2026-2045 and 2081-2100. Sea climate contains, among other information, time series of the significant wave height (H<sub>s</sub>) obtained for several combinations of GCM-RCM projections of EUR-11 for the RCP 8.5. GCM-RCM combinations ACCE, CMCC, CNRM, GFDL, HADG, IPSL, MIRO with a 0.1 degrees grid were used for the Atlantic facade while CNRM, HADG, IPSL, MIRO, MEDC, MPIE, ESM2, EART models with 1/11 degrees were used for the Mediterranean one. A total of 210 locations were analyzed, 54 at the Atlantic facade and 156 at the Mediterranean one (Figure 1). The data was bias adjusted using the Empirical Quantile Mapping (Déqué et al., 2007; Michelangeli et al., 2009). Information of the significant wave height and the dependence between the values at a given time with previous values with a VAR(q) model is already available. </em></p> <p><em>At each location, the methodology of Lira-Loarca et al. (2021) was applied, using the software described in Cobos et al. (2022a). More precisely, for every GCM-RCM (hereinafter, model n for n = 1, .., N where N = 7 for Atlantic data and N = 8 for the Mediterranean data), a non-stationary marginal distribution of H<sub>s</sub>, , assuming that the year was the largest periodicity of the climate, was fitted to data using a lognormal model for the central part and two generalized Pareto distribution for the lower and upper tails, as in Solari and Losada (2011). The non- stationarity is considered by assuming a decomposition of the parameters of the distribution and of the percentiles of the common end points of the interval into a trigonometric truncated expansion.</em></p> <p><em>In addition, the coefficients of the matrix, C<sub>n</sub>, of a VAR(q) model with q up to 92 hours were estimated. The ensemble multi-model characteristics of the data were obtained from the compound distributions and the weighted averaged matrix coefficients. </em></p> <p><em>Soon, the results of the peak period (T<sub>p</sub>) and mean incoming wave direction (ϑ<sub>m</sub>) and the coefficients of the multivariate VAR model will also be included.</em></p> <p> </p> <p> </p>
First Street Foundation Property Level Flood Risk Statistics V2.0
<p>The property level flood risk statistics generated by the First Street Foundation Flood Model Version 2.0 come in CSV format. </p> <p>The data that is included in the CSV includes:</p> <ul> <li> <p>An FSID; a First Street ID (FSID) is a unique identifier assigned to each location.</p> </li> <li> <p>The latitude and longitude of a parcel as well as the zip code, census block group, census tract, county, congressional district, and state of a given parcel.</p> </li> <li> <p>The property’s Flood Factor as well as data on economic loss.</p> </li> <li> <p>The flood depth in centimeters at the low, medium, and high CMIP 4.5 climate scenarios for the 2, 5, 20, 100, and 500 year storms this year and in 30 years.</p> </li> <li> <p>Data on the cumulative probability of a flood event exceeding the 0cm, 15cm, and 30cm threshold depth is provided at the low, medium, and high climate scenarios for this year and in 30 years.</p> </li> <li> <p>Information on historical events and flood adaptation, such as ID and name.</p> </li> </ul> <p> </p> <p>This dataset includes <a href="https://firststreet.org/">First Street</a>'s aggregated flood risk summary statistics. The data is available in CSV format and is aggregated at the congressional district, county, and zip code level. The data allows you to compare FSF data with FEMA data. You can also view aggregated flood risk statistics for various modeled return periods (5-, 100-, and 500-year) and see how risk changes due to climate change (compare FSF 2020 and 2050 data). There are various <a href="https://floodfactor.com/">Flood Factor</a> risk score aggregations available including the average risk score for all properties (flood factor risk scores 1-10) and the average risk score for properties with risk (i.e. flood factor risk scores of 2 or greater). This is version 2.0 of the data and it covers the 50 United States and Puerto Rico. There will be updated versions to follow.</p> <p>If you are interested in acquiring First Street flood data, you can request to access the data <a href="https://firststreet.org/data-access/paid-access/?utm_source=Summary_Statistics_v1.3&utm_medium=Purchase_Data&utm_campaign=Zenodo#pricing-component">here</a>. More information on First Street's flood risk statistics can be found <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-data-dictionaryv2/">here</a> and information on First Street's hazards can be found <a href="https://firststreet.org/data-access/getting-started-with-first-street-data/documentation-hazard-dictionary/?utm_source=Summary_Statistics_v1.3&utm_medium=Hazard_Dictionary&utm_campaign=Zenodo">here</a>.</p> <p>The data dictionary for the parcel-level data is below.</p> <table> <tbody> <tr> <td> <p><strong>Field Name</strong></p> </td> <td> <p><strong>Type</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>fsid</p> </td> <td> <p>int</p> </td> <td> <p>First Street ID (FSID) is a unique identifier assigned to each location</p> </td> </tr> <tr> <td> <p>long</p> </td> <td> <p>float</p> </td> <td> <p>Longitude</p> </td> </tr> <tr> <td> <p>lat</p> </td> <td> <p>float</p> </td> <td> <p>Latitude</p> </td> </tr> <tr> <td> <p>zcta</p> </td> <td> <p>int</p> </td> <td> <p>ZIP code tabulation area as provided by the US Census Bureau</p> </td> </tr> <tr> <td> <p>blkgrp_fips</p> </td> <td> <p>int</p> </td> <td> <p>US Census Block Group FIPS Code</p> </td> </tr> <tr> <td> <p>tract_fips</p> </td> <td> <p>int</p> </td> <td> <p>US Census Tract FIPS Code</p> </td> </tr> <tr> <td> <p>county_fips</p> </td> <td> <p>int</p> </td> <td> <p>County FIPS Code</p> </td> </tr> <tr> <td> <p>cd_fips</p> </td> <td> <p>int</p> </td> <td> <p>Congressional District FIPS Code for the 116th Congress</p> </td> </tr> <tr> <td> <p>state_fips</p> </td> <td> <p>int</p> </td> <td> <p>State FIPS Code</p> </td> </tr> <tr> <td> <p>floodfactor</p> </td> <td> <p>int</p> </td> <td> <p>The property's Flood Factor, a numeric integer from 1-10 (where 1 = minimal and 10 = extreme) based on flooding risk to the building footprint. Flood risk is defined as a combination of cumulative risk over 30 years and flood depth. Flood depth is calculated at the lowest elevation of the building footprint (largest if more than 1 exists, or property centroid where footprint does not exist)</p> </td> </tr> <tr> <td> <p>CS_depth_RP_YY</p> </td> <td> <p>int</p> </td> <td> <p>Climate Scenario (low, medium or high) by Flood depth (in cm) for the Return Period (2, 5, 20, 100 or 500) and Year (today or 30 years in the future). Today as year00 and 30 years as year30. ex: low_depth_002_year00</p> </td> </tr> <tr> <td> <p>CS_chance_flood_YY</p> </td> <td> <p>float</p> </td> <td> <p>Climate Scenario (low, medium or high) by Cumulative probability (percent) of at least one flooding event that exceeds the threshold at a threshold flooding depth in cm (0, 15, 30) for the year (today or 30 years in the future). Today as year00 and 30 years as year30. ex: low_chance_00_year00</p> </td> </tr> <tr> <td> <p>aal_YY_CS</p> </td> <td> <p>int</p> </td> <td> <p>The annualized economic damage estimate to the building structure from flooding by Year (today or 30 years in the future) by Climate Scenario (low, medium, high). Today as year00 and 30 years as year30. ex: aal_year00_low</p> </td> </tr> <tr> <td> <p>hist1_id</p> </td> <td> <p>int</p> </td> <td> <p>A unique First Street identifier assigned to a historic storm event modeled by First Street</p> </td> </tr> <tr> <td> <p>hist1_event</p> </td> <td> <p>string</p> </td> <td> <p>Short name of the modeled historic event</p> </td> </tr> <tr> <td> <p>hist1_year</p> </td> <td> <p>int</p> </td> <td> <p>Year the modeled historic event occurred</p> </td> </tr> <tr> <td> <p>hist1_depth</p> </td> <td> <p>int</p> </td> <td> <p>Depth (in cm) of flooding to the building from this historic event</p> </td> </tr> <tr> <td> <p>hist2_id</p> </td> <td> <p>int</p> </td> <td> <p>A unique First Street identifier assigned to a historic storm event modeled by First Street</p> </td> </tr> <tr> <td> <p>hist2_event</p> </td> <td> <p>string</p> </td> <td> <p>Short name of the modeled historic event</p> </td> </tr> <tr> <td> <p>hist2_year</p> </td> <td> <p>int</p> </td> <td> <p>Year the modeled historic event occurred</p> </td> </tr> <tr> <td> <p>hist2_depth</p> </td> <td> <p>int</p> </td> <td> <p>Depth (in cm) of flooding to the building from this historic event</p> </td> </tr> <tr> <td> <p>adapt_id</p> </td> <td> <p>int</p> </td> <td> <p>A unique First Street identifier assigned to each adaptation project</p> </td> </tr> <tr> <td> <p>adapt_name</p> </td> <td> <p>string</p> </td> <td> <p>Name of adaptation project</p> </td> </tr> <tr> <td> <p>adapt_rp</p> </td> <td> <p>int</p> </td> <td> <p>Return period of flood event structure provides protection for when applicable</p> </td> </tr> <tr> <td> <p>adapt_type</p> </td> <td> <p>string</p> </td> <td> <p>Specific flood adaptation structure type (can be one of many structures associated with a project)</p> </td> </tr> <tr> <td> <p>fema_zone</p> </td> <td> <p>string</p> </td> <td> <p>Specific FEMA zone categorization of the property ex: A, AE, V. Zones beginning with "A" or "V" are inside the Special Flood Hazard Area which indicates high risk and flood insurance is required for structures with mortgages from federally regulated or insured lenders</p> </td> </tr> <tr> <td> <p>footprint_flag</p> </td> <td> <p>int</p> </td> <td> <p>Statistics for the property are calculated at the centroid of the building footprint (1) or at the centroid of the parcel (0)</p> </td> </tr> </tbody> </table> <p> </p>
Input data set for the statistical analsysis of rockfall reach probabilities
<p>These files contain reach probability values extracted from 3D rockfall simulations for field-mapped block deposits as well as a series of attributes characterising the deposits. They served for the statistical analysis of the reach probability values as a function of site, forest and rockfall characteristics. The results of the analysis are published in Dorren et al. 2022: Delimiting rockfall runout zones using reach probability values simulated with a Monte-Carlo based 3D trajectory model. Natural Hazards and Earth System Scienses.</p>
Photon-emission statistics induced by electron tunnelling in plasmonic nanojunctions
<p>OPEN DATA related to the research publication:</p> <p>R. Avriller, Q. Schaeverbeke, T. Frederiksen, and F. Pistolesi<br> <em>Photon-emission statistics induced by electron tunnelling in plasmonic nanojunctions</em><br> Phys. Rev. B <strong>104</strong>, L241403 (2021) [arXiv:2107.07860]</p>
Coherent and non-coherent Eddy Kinetic Energy and gridded coherent eddy statistics
<p>This dataset includes the processed data used for the paper titled "Climatology, seasonality and trends of oceanic coherent eddies". The original data was obtained from AVISO+ SSH altimetry, Martínez-Moreno, J. <em>et.al (</em>2019)<em> </em>and Chelton, D. B., & Schlax, M. G. (2013).</p> <p> </p> <p>Further information and scripts to reproduce the result of the manuscript can be found at: https://github.com/josuemtzmo/CEKE_climatology</p>
Regional summary statistics for 1107 protein targets based on the Olink technology
<p>This data set contains regional summary statistics (±500kb around the protein coding gene) for a total of 1107 protein - gene combinations as measured by the Olink Proximity Extension Assay in the Fenland study (https://www.mrc-epid.cam.ac.uk/research/studies/fenland/) among 485 individuals. A detailed description of the genetic analysis can be found here https://www.nature.com/articles/s41467-021-27164-0. </p>
Summary Statistics from "Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2"
<p>GWAMA summary statistics of PCSK9 levels using fixed-effect model. Genome-wide data is given for Europeans with statin adjustment and Europeans without statin treatment only (subset of the population). In addition, locus-wide data of the PCSK9 gene locus for African-Americans without statin treatment is listed.</p> <p>When using this data, please cite: Pott J, Gadin J, Theusch E, et al.. Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2. Hum Mol Genet. 2021 Sep 30:ddab279. doi: 10.1093/hmg/ddab279. PMID: 34590679</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>
Sentinel-3 NDVI ARD and Long Term Statistics (1999-2019) from the Copernicus Global Land Service over Lombardia
<p>Sentinel-3 NDVI Analysis Ready Data (ARD) (C_GLS_NDVI_20220101_20220701_Lombardia_S3_2.nc) product provided by the Copernicus Global Land Service [3]. The file C_GLS_NDVI_20220101_20220701_Lombardia_S3_2_masked.nc is derived from C_GLS_NDVI_20220101_20220701_Lombardia_S3_2.nc but values have been scaled (raw_value * ( 1/250) - 0.08) and values lower then -0.08 and greater than 0.92 have been removed (set to missing values).</p> <p>The original dataset can also be discovered through the OpenEO API[5] from the CGLS distributor VITO [4]. Access is free of charge but an <a href="https://aai.egi.eu/">EGI registration</a> is needed.</p> <p>The file called Italy.geojson has been created using the Global Administrative Unit Layers <a href="https://data.apps.fao.org/map/catalog/srv/eng/catalog.search#/metadata/9c35ba10-5649-41c8-bdfc-eb78e9e65654">GAUL G2015_2014</a> provided by FAO-UN (see <a href="https://data.apps.fao.org/map/catalog/srv/api/records/9c35ba10-5649-41c8-bdfc-eb78e9e65654/attachments/GAUL2015_Documentation.zip">Documentation</a>). It only contains information related to Italy.</p> <p> </p> <p>Further info about drought indexes can be found in the Integrated Drought Management Programme [5]</p> <p>[1] <a href="https://www.sciencedirect.com/science/article/abs/pii/027311779500079T">Application of vegetation index and brightness temperature for drought detection</a> [2] <a href="https://en.wikipedia.org/wiki/Normalized_difference_vegetation_index">NDVI</a> [3] <a href="https://land.copernicus.eu/global/index.html">Copernicus Global Land Service</a> [4] <a href="https://vito.be/en">Vito</a> [5] <a href="https://openeo.org/">OpenEO</a> [5] <a href="https://www.droughtmanagement.info/indices">Integrated Drought Management</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.