Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

144

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

144 results for “statistical model”

Learn how ShareScore rates datasets ↗
zenodo40/100

Monthly CDIP:MOP-alongshore modeled wave statistics for California, January 2000 - July 2022

<p><strong>Overview</strong></p> <ul> <li>Monthly wave averages for all 11,594 CDIPS-MOPS alongshore sites in California.</li> <li>Jan 2000 to July 2022 inclusive</li> <li>Sites are defined in the files &quot;CDIP_Transects.csv&quot; and &quot;CDIP_Transects.geojson&quot;. The bounds of each site are listed in the file &quot;CA_region_bounds.csv&quot;</li> <li>Data are described here: https://cdip.ucsd.edu/documents/index/product_docs/mops/mop_intro.html</li> <li>Data are obtained from here: https://thredds.cdip.ucsd.edu/thredds/catalog.html</li> </ul> <p><strong>Methods</strong></p> <ul> <li>Data are computed from hourly inshore wave hindcasts and nowcasts. Data download script is the file &quot;CDIP_MassDownloader.ipynb&quot;</li> <li>wave summary statistics have been computed using the file &quot;Create_stats.ipynb&quot;. All monthly data are simple averages (i.e. mean values) of the hourly data</li> </ul> <p><strong>Data files</strong></p> <ul> <li>Data have been split into 25 regions, defined in &quot;CA_regions.json&quot;</li> <li>Data are provided in geoJSON format, in the form of one file per region, and one file for all regions</li> </ul> <p><strong>Data fields</strong></p> <ul> <li>Hs: significant wave height [meters]</li> <li>Tp: peak wave period [seconds]</li> <li>Ta: average wave period [seconds]</li> <li>Dp: peak wave direction [degrees]</li> <li>Da: average wave direction [degrees]</li> <li>Ea: wave energy density, averaged over wave frequencies</li> <li>Es: wave energy density, summed over wave frequencies</li> <li>QC: quality flag</li> <li>waveTime: UTC time string</li> <li>metaWaterDepth: water depth of modeled wave data (range is 10-15m)</li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo40/100

FIGURE 4. P and S in BoneProfileR: The next step to quantify, model, and statistically compare bone section compactness profiles

FIGURE 4. P and S values ffor radial compactness analysis of (A, B) Erinaceus europaeus and (C, D) Eryops megacephalus femur using the flexit model.

opencc-by-4.0Dec 2022View details →
zenodo40/100

FIGURE 3 in BoneProfileR: The next step to quantify, model, and statistically compare bone section compactness profiles

FIGURE 3. Posterior distribution of K1 and K2 parameters of the flexit model applied on Erinaceus europaeus femur (A, B) and on Eryops megacephalus femur (C, D). A -1000 to +1000 uniform prior distribution was used. K1 = K2 = 1 shown in interrupted line is the logistic equation.

opencc-by-4.0Dec 2022View details →
zenodo40/100

FIGURE 2 in BoneProfileR: The next step to quantify, model, and statistically compare bone section compactness profiles

FIGURE 2. Comparison of logistic and flexit fits (A, B) of the Erinaceus europaeus femur shown in Figure 1A and (C, D) of the Eryops megacephalus femur shown in Figure 1B. Table 1 shows the AIC and Akaike weight values for these models.

opencc-by-4.0Dec 2022View details →
zenodo40/100

FIGURE 1 in BoneProfileR: The next step to quantify, model, and statistically compare bone section compactness profiles

FIGURE 1. Example of background and foreground automatic detection in (A) Erinaceus europaeus femur, and (B) Eryops megacephalus femur. Centers were automatically detected in (A) and manually positioned in (B). Here, sections are segmented in 100 concentric circles and 60 slices for measures of global and radial compactness. Details on bone section preparation can be found in Laurin et al. (2004) and Quémeneur et al. (2013).

opencc-by-4.0Dec 2022View details →
zenodo40/100

Stable Modeling on Resource Usage Parameters of MapReduce Application-Figure 4. Statistical Metrics distribution of model on RIO as response of Terasort application

<p>Figure 4 shows the statistical metrics distribution of regression model on read rate as the response of Terasort application. The filled triangle point-up indicates the minimum stable sampling time for statistical metrics. The top-half of figure 4 shows the residual standard error (RSE) distribution as training data size increase. The remaining half is for the distribution of R2.</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Stable Modeling on Resource Usage Parameters of MapReduce Application-Figure 8. Minimum sample time of statistical metrics of MapReduce applications

<p>Figure 8 presents the minimum sampling time distribution of statistic metrics which ensures the stable modeling. Overall, the minimum sampling time of statistic metrics is smaller than sampling time of estimated coefficients. For different applications, a time-consuming application like Terasort needs the largest sampling time to tend to be stable. The Pi application shows the smallest minimum sampling time to reach stability.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Meandering evolution and width variations: a physics-statistics based modeling approach

<p>Coordinates (x,y) of the central bank&nbsp; (field 1 and 2), the distance of the central axis (field 3), coordinates (x,y) of the left bank (fields 4 and 5),&nbsp;coordinates (x,y) of the right bank (fields 6 and 7)</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Annex B to the technical report on the raw primary commodity (RPC) model - Summary statistics of the output data

<p><strong>The raw primary commodity&nbsp;model</strong>:</p> <p>Dietary exposure is typically calculated by combining food consumption data with occurrence data. EFSA&rsquo;s food consumption data are stored in the Comprehensive European Food Consumption Database (Comprehensive Database). Some of these data, however, cannot be used in exposure assessments when the occurrence data are reported for the raw primary commodities (RPCs). The RPC model aims to bridge this gap by transforming the Comprehensive Database into RPC consumption data. Using the RPC model, EFSA successfully developed a new RPC Consumption Database, which contains 51 dietary surveys from 23 different countries. These surveys cover a total of 94,532 subjects and 26,573,088 RPC consumption records. The consumption data generated by the RPC model were manually checked and validated by means of case studies. These case studies demonstrated that the RPC consumption data are suitable for assessing dietary exposure to chemicals where the occurrence data are predominantly available for RPCs.</p> <p><strong>Annex B to the technical report on the&nbsp;raw primary commodity model:</strong></p> <p>Annex B is an excel file which presents summary statistics of the output data generated by the RPC model. The following tables are included in Annex B:</p> <p>Table B.1 :Summary statistics of chronic RPC consumption expressed in g/kg bw per day (total population)</p> <p>Table B.2 :Summary statistics of chronic RPC consumption expressed in g/day (total population)</p> <p>Table B.3 :Summary statistics of acute RPC consumption expressed in g/kg bw (consumers only)</p> <p>Table B.4 :Summary statistics of acute RPC consumption expressed in g (consumers only)</p> <p>Table B.5 :Comparison of the RPC consumption data with RPC consumption data used in EFSA&#39;s Pesticides Residues Intake Model (PRIMo)</p> <p>Table B.6 :Contribution of processed products to the average chronic RPC consumption</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

Hyperparameter tuning and performance assessment of statistical and machine-learning models using spatial data.

<p>This is a research compendium (RC) for the publication &quot;Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data&quot;.</p> <p>The code (including figures, appendices and the manuscript) is packed in <strong>pathogen-modeling-3.zip&nbsp;</strong>or can be found directly in the <a href="https://github.com/pat-s/pathogen-modeling">Github repository</a>.</p> <ul> <li><strong>Publication figures</strong>:&nbsp;analysis/paper/submission/3/latex-source-files/</li> <li><strong>Appendices</strong>: analysis/paper/submission/3/</li> </ul> <p>This RC represents a static snapshot at the time of submission. The Github repository will receive changes after the publication was published.</p> <p><strong>Data sources</strong></p> <ul> <li>Atlas Climatico:&nbsp;<a href="http://opengis.uab.es/wms/iberia/index.htm">http://opengis.uab.es/wms/iberia/index.htm</a></li> <li>DEM:&nbsp;ftp://ftp.geo.euskadi.eus/lidar/MDE_LIDAR_2016_ETRS89/</li> <li>Lithology:&nbsp;<a href="http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home">http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home</a></li> <li>pH:&nbsp;<a href="https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0">https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0</a></li> <li>soil:&nbsp;<a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a></li> </ul> <p><strong>Licenses</strong></p> <p>All files are shared via the given license with the exception of &quot;soil.tif&quot; which is shared via the&nbsp;<strong>ODbL </strong>license<strong>.</strong></p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Great Lakes monthly water balance components from the Large Lakes Statistical Water Balance Model (L2SWBM)

<p>**Note that an updated version of the data (v3.0) was uploaded on October 2, 2024, which supersedes earlier versions**</p> <p>These data sets are the results of leveraging bi-national data and the Large Lakes Statistical Water Balance Model (L2SWBM) specifically tailored for the Laurentian Great Lakes to produce value-added time series of water supply components, including expressions of uncertainty, that ultimately close the water balance across the interconnected Great Lakes system.&nbsp;</p> <p>The model serves as a new cornerstone for bi-national coordination of hydrologic data throughout this international transboundary basin, providing an improved means of capturing data patterns, revealing seasonal variabilities, as well as short-term and long-term trends.</p> <p>This repository includes monthly output from the L2SWBM. Output datasets include over-lake precipitation, over-lake evaporation, lateral tributary inflow (runoff), connecting channel flow (cms and also included in mm normalized to lake area), diversion flow (cms and also included in mm normalized to lake area), and component net basin supply. Data is included for lakes Superior, Michigan-Huron, Erie, and Ontario.</p> <p>This version contains data from 1950 to 2022.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Dataset for publication: Statistically Equivalent Virtual Microstructures for Modeling of Complex Polycrystalline Alloys Using a Generative Adversarial Network (GAN)-Enabled Computational Platform

<p>This dataset provides the necessary data to get the images and results shown in the paper "Statistically Equivalent Virtual Microstructures for Modeling of Complex Polycrystalline Alloys Using a Generative Adversarial Network (GAN)-Enabled Computational Platform".&nbsp;</p> <p>Source Data Raw.zip has the entire data set used to generate the images.</p> <p>Source Data.zip contains the processed data&nbsp; from "Source Data Raw.zip".&nbsp; &nbsp;</p> <p>Files with extension .dream3d are accompained by a file with extension .xdmf. This files can be opened with Paraview. And their data can be accesible using python or matlab.</p> <p>For more information contact Proffesor Somnath Ghosh at Johns Hopkins University, Civil and Systems Engineering Department.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Output from Linear Inverse Models (LIMs) emulating the observed spatiotemporal statistics of Australian precipitation and global sea surface temperatures

<p><strong>Data repository for <em>How unusual was Australia's 2017&ndash;2019 Tinderbox Drought?</em></strong></p> <p>This repository contains LIM data underpinning the paper&nbsp;<em>How unusual was Australia's 2017&ndash;2019 Tinderbox Drought?</em> [doi: 10.1016/j.wace.2024.100734 <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.wace.2024.100734" target="_blank" rel="noopener">available online in&nbsp;<em>Weather and Climate Extremes</em> 17 October 2024</a>]. All other datasets used in the paper are freely available online (see Data Availability statement in the paper for details).&nbsp;</p> <p>The repository contains 12 netcdf files, which together comprise the Linear Inverse Model (LIM) outputs described in the paper. <strong>In all cases, please see the paper for important details on the data and how they were produced.</strong>&nbsp;</p> <p><em>Global LIMs</em></p> <ul> <li>`LIM5000_COBE-globalSST_prec-anoms-aus_monthly.nc` <ul> <li>contains 5000 years of emulated Australian precipitation variability, modelled using Australian rainfall data from the Australian Gridded Climate Dataset v2 (AGCD) and global SST data from 'Centennial in situ Observation-Based Estimates of the Variability of SST and Marine Meteorological Variables version 2' (COBE)</li> </ul> </li> <li>`LIM5000_ERSST-globalSST_prec-anoms-aus_monthly.nc` <ul> <li>contains 5000 years of emulated Australian precipitation variability, modelled using Australian rainfall data from the AGCD and global SST data from US National Oceanic and Atmospheric Administration 'Extended Reconstruction SST version 5&rsquo; (ERSST)</li> </ul> </li> <li>`LIM5000_COBE-globalSST_SST-anoms-global_monthly.nc` <ul> <li>contains 5000 years of emulated global SST variability, modelled using global SST data from COBE</li> </ul> </li> <li>`LIM5000_ERSST-globalSST_SST-anoms-global_monthly.nc` <ul> <li>contains 5000 years of emulated global SST variability, modelled using global SST data from ERSST</li> </ul> </li> </ul> <p><em>Tropical Pacific Ocean LIMs</em></p> <ul> <li>`LIM5000_COBE-TropicalPacificSST_prec-anoms-aus_monthly.nc` <ul> <li>contains 5000 years of emulated Australian precipitation variability, modelled using Australian rainfall data from the AGCD and tropical Pacific Ocean SST data from COBE</li> </ul> </li> <li>`LIM5000_ERSST-TropicalPacificSST_prec-anoms-aus_monthly.nc` <ul> <li>contains 5000 years of emulated Australian precipitation variability, modelled using Australian rainfall data from the AGCD and tropical Pacific Ocean SST data from ERSST</li> </ul> </li> <li>`LIM5000_COBE-TropicalPacificSST_SST-anoms-TropicalPacific_monthly.nc` <ul> <li>contains 5000 years of emulated global SST variability, modelled using tropical Pacific Ocean SST data from COBE</li> </ul> </li> <li>`LIM5000_ERSST-TropicalPacificSST_SST-anoms-TropicalPacific_monthly.nc` <ul> <li>contains 5000 years of emulated global SST variability, modelled using tropical Pacific Ocean SST data from ERSST</li> </ul> </li> </ul> <p><em>Indian Ocean LIMs</em></p> <ul> <li>`LIM5000_COBE-IndianOceanSST_prec-anoms-aus_monthly.nc` <ul> <li>contains 5000 years of emulated Australian precipitation variability, modelled using Australian rainfall data from the AGCD and Indian Ocean SST data from COBE</li> </ul> </li> <li>`LIM5000_ERSST-IndianOceanSST_prec-anoms-aus_monthly.nc` <ul> <li>contains 5000 years of emulated Australian precipitation variability, modelled using Australian rainfall data from the AGCD and Indian Ocean SST data from ERSST</li> </ul> </li> <li>`LIM5000_COBE-IndianOceanSST_SST-anoms-TropicalPacific_monthly.nc` <ul> <li>contains 5000 years of emulated global SST variability, modelled using Indian Ocean SST data from COBE</li> </ul> </li> <li>`LIM5000_ERSST-IndianOceanSST_SST-anoms-TropicalPacific_monthly.nc` <ul> <li>contains 5000 years of emulated global SST variability, modelled using Indian Ocean SST data from ERSST</li> </ul> </li> </ul> <p><strong>How to cite this</strong> <strong>repository</strong></p> <p>If using this data, please cite the original publication, available from <a href="https://www.sciencedirect.com/science/article/pii/S2212094724000951" target="_blank" rel="noopener">https://www.sciencedirect.com/science/article/pii/S2212094724000951.</a>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
dryad40/100

Statistical analysis code for output from a model used to simulate foot-and-mouth disease dynamics in the United Kingdom

<p>Epidemics can sometimes be managed through reductions of host density, such as social distancing for human diseases, reducing plant density through cultural and genetic means, and host culling for epizootics. These approaches allow for a certain density of hosts to remain within a targeted area. By contrast, total ring depopulation is often used as a management strategy for emerging infectious diseases in livestock. In this study, we explore the trade-offs of a density-based culling strategy to determine if fewer livestock farms can be culled within rings while maintaining a decrease in disease transmission. To do so, we evaluated a farm-density-based ring culling strategy to control foot-and-mouth disease (FMD) in the United Kingdom. This strategy may allow for some farms within rings around infected premises (IPs) to escape depopulation, with the aim to prevent over-culling during outbreaks. Using a spatially-explicit, stochastic, state-transition simulation algorithm originally developed by Keeling et al. 2001 to model FMD spread in the United Kingdom, we simulated this reduced-farm-density, or "target density" strategy. We modeled FMD disease spread in four counties in the UK (Aberdeenshire, Cumbria, Devon, and North Yorkshire) that have different farm demographies. We ran 740,000 simulations in a full-factorial analysis of epidemic impact measurements (i.e. culled animals, culled farms, epidemic length) and cull strategy parameters (i.e. target farm density, daily farm cull capacity, cull radius). We found that all of the cull strategy parameters were drivers of epidemic impact. We found that outbreaks in Cumbria had higher epidemic impacts and were more likely to take off compared with other counties with more outbreaks being likely to take off in Cumbria. Most importantly, in all counties, our proposed target density strategy was more effective at combatting FMD compared with traditional 'total ring depopulation' when considering average culled animals and culled farms. The differences in epidemic impact between the counties are likely driven by farm demography, especially differences in cattle and farm density. This target density strategy can be applied to many different systems, including other livestock and agricultural systems, to reduce host density as opposed to over-culling hosts.</p>

opencc-zeroAug 2021View details →
dryad40/100

Data for: Nitrogen deposition in forests: Statistical modeling of total deposition from throughfall loads

<p><strong>Introduction:</strong> Nitrogen (N) gradient studies in some cases use N deposition in throughfall as measure of N deposition to forests. For evaluating critical loads of N, however, information on total N deposition is required, i.e., the sum of estimates of dry, wet and occult deposition.</p> <p><strong>Methods: </strong>The present paper collects a number of studies in Europe where throughfall and total N deposition were compared in different forest types. From this dataset a function was derived which allows to estimate total N deposition from throughfall N deposition.</p> <p><strong>Results: </strong>At low throughfall N deposition values, the proportion of canopy uptake is high and thus the underestimation of total deposition by throughfall N needs to be corrected. At throughfall N deposition values &gt;20 kg N ha<sup>-1</sup> yr<sup>-1</sup> canopy uptake is getting less important.</p> <p><strong>Conclusions: </strong>This work shows that throughfall clearly underestimates total deposition of nitrogen. With the present data set covering large parts of Europe it is possible to derive a critical load estimate from gradient studies using throughfall data.</p>

opencc-zeroDec 2022View details →
zenodo40/100

Dataset: Modelling surface color discrimination under different lighting environments using image chromatic statistics and convolutional neural networks

<p><strong>Associated publication</strong></p> <p>[1] Samuel Ponting*, <strong>Takuma Morimoto</strong>*, Hannah E. Smithson, &ldquo;Modelling surface color discrimination under different lighting environments using image chromatic statistics and convolutional neural networks&rdquo;, *equal contribution, bioRxiv, <a href="https://www.google.com/url?q=https%3A%2F%2Fdoi.org%2F10.1101%2F2022.11.02.514864&amp;sa=D&amp;sntz=1&amp;usg=AOvVaw3KwSo7KmqPzBR1UMc1MHmk">https://doi.org/10.1101/2022.11.02.514864</a></p> <p>[2] Takuma Morimoto, and Hannah E. Smithson, &ldquo;Discrimination of spectral reflectance under complex environmental illumination,&rdquo; Journal of the Optical Society of America A, 35, 4, B244-B255 (2018) https://doi.org/10.1364/JOSAA.35.00B244</p> <p>&nbsp;</p> <p>Datasets contain 2 folders and 1 mat file.</p> <p>&nbsp;</p> <p><strong>(Folder 1) Stimuli</strong></p> <p><strong>(Folder 2) Psychophysics_data</strong></p> <p><strong>(Mat file) stimulusMagnitudeToMacLeodBoynton.mat</strong></p> <p>&nbsp;</p> <p>Details are described below.</p> <p>&nbsp;</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>(Folder 1) Stimuli</strong></p> <p>&nbsp;</p> <p><strong>Overview of datasets</strong></p> <p>This Image dataset includes 57,600 images (2 gloss levels * 3 environments * 100 stimulus magnitudes * 8 hue directions * 12 camera angles from 0 to 330 degree in 30 degree step) in .mat format.</p> <p>&nbsp;</p> <p>The half of images were used in psychophysical experiment (camera angles: 0, 60, 120, 180, 240, 300 degrees).</p> <p>Other half images were used for testing chromatic statistics models and CNN-based models [1] (camera angles: 30, 90, 150, 210, 270, 330 degrees).</p> <p>&nbsp;</p> <p><strong>Each image file</strong></p> <p>Filename denotes a condition name and the camera angle as formatted in a following way.</p> <p>&nbsp;</p> <p>stim_&rdquo;environment&rdquo; _&rdquo;glossiness&rdquo;_&rdquo;hueAngle&rdquo;_&rdquo;magnitude&rdquo;_&rdquo;cameraAngle&rdquo;.mat</p> <p>e.g. &ldquo;stim_en1_glossy_hue45_n45_cameraAngle90.mat&rdquo;</p> <p>&nbsp;</p> <p>Stimulus magnitude 100 is a maximum saturation, and 1 corresponds to equal energy white (which was used as a distractor object).</p> <p>&nbsp;</p> <p>Each image file contains two valuables : MacLeodBoynton, XYZ</p> <p>&nbsp;</p> <p>Each variable contains an image of 128*128*3 pixels (height*width*channel).</p> <p>&nbsp;</p> <p>MacLeod-Boynton: MacLeod-Boynton chromaticity image (1st channel: L/(L+M), 2nd channel: S/(L+M), and 3rd channel L+M)</p> <p>XYZ: XYZ coordinates calculated based on 2-degree CIE 1931 xyz color matching function (1st channel: X, 2nd channel: Y, and 3rd channel Z)</p> <p>&nbsp;</p> <p>Luminance and L+M are both relative (normalised by the maximum luminance across all 57,600 images).</p> <p>&nbsp;</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>(Folder 2) Psychophysics_data</strong></p> <p>Filename denotes the condition and observers formatted in a following way.</p> <p>&nbsp;</p> <p>data_&rdquo;environment&rdquo; _&rdquo;specularities&rdquo;_&rdquo;sessionNumber&rdquo;_&rdquo;obsever&rdquo;.mat</p> <p>e.g. data_en2_matte_session4_JH.mat or .csv</p> <p>&nbsp;</p> <p>Each file includes following variables:</p> <p>&nbsp;</p> <p>(Variable 1) threshold</p> <p>Thresholds are stored in MacLeod-Boynton (MB) chromaticity coordinates for all 8 hue directions (from 0 to 315 degree in 45 degree step).</p> <p>&nbsp;</p> <p>MacLeod-Boynton chromaticity coordinates were calculated in a following way.   </p> <p>These scalings are in accordance with description in CVRL main site (Chromaticity coordinates tab ).</p> <p>&nbsp;</p> <p>First of all, L, M, and S cone signals were calculated based on Stockman &amp; Sharpe cone fundamentals (energy in linear scale available at at http://www.cvrl.org).</p> <p>Each sensitivity curve was normalised to have 1.0 at the peak.</p> <p>&nbsp;</p> <p>Then, MB coordinates were calculated using equation (1-3).</p> <p>&nbsp;</p> <p>L/(L+M) = Lw*L/(Lw*L+Mw*M) - (1)</p> <p>S/(L+M) = Sw*S/(Lw*L+Mw*M) - (2)</p> <p>L+M = Lw*L+Mw*M - (3)</p> <p>&nbsp;</p> <p>where Lw = 0.689903; Mw = 0.348322;Sw = 1.93540.</p> <p>&nbsp;</p> <p>L, M and S denote L-cone, M-cone, S-cone excitations, respectively.</p> <p>&nbsp;</p> <p>Under this calculation, equal energy white becomes L/(L+M) = 0.7078 and S/(L+M) = 1.</p> <p>&nbsp;</p> <p>(Variable 2) staircase</p> <p>&nbsp;</p> <p>Since we ran 8 interleaved staircase (for 8 hue angles), information about 8 staircases are stored in this single variable.</p> <p>(staircase(1) corresponds to 0 degree, and staircase(8) corresponds to 315 degree)</p> <p>&nbsp;</p> <p>There are 5 fields:</p> <p>(i) groundtruth,    (ii) response,    (iii) correct, (iv) magnitude, (v) cameraAngle</p> <p>&nbsp;</p> <p>For each trial, the location of objects was defined in a following way.</p> <p>| 1 3 |</p> <p>| 2 4 |</p> <p>&nbsp;</p> <p>And each field stores following information for all trials in the staircase.</p> <p>&nbsp;</p> <p>(i) groundtruth</p> <p>Location of the target object</p> <p>&nbsp;</p> <p>(ii) response</p> <p>Location that the participant chose</p> <p>&nbsp;</p> <p>(iii) correct</p> <p>If the response was correct (1) or incorrect    (0)</p> <p>&nbsp;</p> <p>(iv) magnitude</p> <p>Stimulus magnitude of target object in each trial from 1 to 100 (1 for equal energy white and 100 for maximum saturation).</p> <p>&nbsp;</p> <p>(v) Camera angle</p> <p>Camera angles assigned for four objects in each trial.</p> <p>This data and (i) groundtruth allow reconstruct of the exact image for each trial.</p> <p>&nbsp;</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>(Mat file) stimulusMagnitudeToMacLeodBoynton.mat</strong></p> <p>This file stores a variable &lsquo;stimulusMagnitudeToMacLeodBoynton&rsquo; (8*100*2) which describes correspondence map between stimulus magnitude and MacLeod-Boynton chromaticity.</p> <p>&nbsp;</p> <p>1st channel: hue direction from 0 degree to 315 degree, 45 degree step</p> <p>2nd channel: magnitude from 1 to 100</p> <p>3rd channel: MacLeod-Boynton coordinate, 1 being L/(L+M) and 2 being S/(L+M)</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Codes in R for spatial statistics analysis, ecological response models and spatial distribution models

<p>In the last decade, a plethora of algorithms have been developed for spatial ecology studies. In our case, we use some of these codes for underwater research work in applied ecology analysis of threatened endemic fishes and their natural habitat. For this, we developed codes in Rstudio&reg; script environment to run spatial and statistical analyses for ecological response and spatial distribution models (e.g., Hijmans &amp; Elith, 2017; Den Burg <em>et al.</em>, 2020). The employed R packages are as follows: caret (Kuhn et al., 2020), corrplot (Wei &amp; Simko, 2017), devtools (Wickham, 2015), dismo (Hijmans &amp; Elith, 2017), gbm (Freund &amp; Schapire, 1997; Friedman, 2002), ggplot2 (Wickham et al., 2019), lattice (Sarkar, 2008), lattice (Musa &amp; Mansor, 2021), maptools (Hijmans &amp; Elith, 2017), modelmetrics (Hvitfeldt &amp; Silge, 2021), pander (Wickham, 2015), plyr (Wickham &amp; Wickham, 2015), pROC (Robin et al., 2011), raster (Hijmans &amp; Elith, 2017), RColorBrewer (Neuwirth, 2014), Rcpp (Eddelbeuttel &amp; Balamura, 2018), rgdal (Verzani, 2011), sdm (Naimi &amp; Araujo, 2016), sf (e.g., Zainuddin, 2023), sp (Pebesma, 2020) and usethis (Gladstone, 2022).</p> <p>It is important to follow all the codes in order to obtain results from the ecological response and spatial distribution models. In particular, for the ecological scenario, we selected the Generalized Linear Model (GLM) and for the geographic scenario we selected DOMAIN, also known as Gower&#39;s metric (Carpenter <em>et al.</em>, 1993). We selected this regression method and this distance similarity metric because of its adequacy and robustness for studies with endemic or threatened species (<em>e.g.</em>, Naoki <em>et al.</em>, 2006). Next, we explain the statistical parameterization for the codes immersed in the GLM and DOMAIN running:</p> <p>In the first instance, we generated the background points and extracted the values of the variables (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code2_Extract_values_DWp_SC.R?versionId=c1ea0c61-53fe-4f95-ab88-0c1cb28399cb">Code2_Extract_values_DWp_SC.R</a>). Barbet-Massin <em>et al. </em>(2012) recommend the use of 10,000 background points when using regression methods (<em>e.g.</em>, Generalized Linear Model) or distance-based models (<em>e.g.</em>, DOMAIN). However, we considered important some factors such as the extent of the area and the type of study species for the correct selection of the number of points (Pers. Obs.).&nbsp; Then, we extracted the values of predictor variables (<em>e.g.</em>, bioclimatic, topographic, demographic, habitat) in function of presence and background points (<em>e.g.</em>, Hijmans and Elith, 2017).</p> <p>Subsequently, we subdivide both the presence and background point groups into 75% training data and 25% test data, each group, following the method of Sober&oacute;n &amp; Nakamura (2009) and Hijmans &amp; Elith (2017). For a training control, the 10-fold (cross-validation) method is selected, where the response variable presence is assigned as a factor. In case that some other variable would be important for the study species, it should also be assigned as a factor (Kim, 2009).</p> <p>After that, we ran the code for the GBM method (Gradient Boost Machine; <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code3_GBM_Relative_contribution.R?versionId=1656bbae-66aa-409e-bb91-d8007dee8f95">Code3_GBM_Relative_contribution.R</a> and <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code4_Relative_contribution.R?versionId=0e1d9352-e6b2-43da-984b-d6853a914258">Code4_Relative_contribution.R</a>), where we obtained the relative contribution of the variables used in the model. We parameterized the code with a Gaussian distribution and cross iteration of 5,000 repetitions (<em>e.g.</em>, Friedman, 2002; kim, 2009; Hijmans and Elith, 2017). In addition, we considered selecting a validation interval of 4 random training points (Personal test). The obtained plots were the partial dependence blocks, in function of each predictor variable.</p> <p>Subsequently, the correlation of the variables is run by Pearson&#39;s method (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code5_Pearson_Correlation.R?versionId=275f8dd4-b056-44d2-bfe5-f6264bc3298b">Code5_Pearson_Correlation.R</a>) to evaluate multicollinearity between variables (Guisan &amp; Hofer, 2003). It is recommended to consider a bivariate correlation &plusmn; 0.70 to discard highly correlated variables (<em>e.g.</em>, Awan <em>et al.</em>, 2021).</p> <p>Once the above codes were run, we uploaded the same subgroups (<em>i.e.</em>, presence and background groups with 75% training and 25% testing) (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code6_Presence&amp;backgrounds.R?versionId=d797b528-782f-4a19-bd61-cfb197f38513">Code6_Presence&amp;backgrounds.R</a>) for the GLM method code (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code7_GLM_model.R?versionId=e4aca276-d601-49ec-a62c-a9223b05a7ed">Code7_GLM_model.R</a>). Here, we first ran the GLM models per variable to obtain the <em>p</em>-significance value of each variable (alpha &le; 0.05); we selected the value one (<em>i.e.</em>, presence) as the likelihood factor. The generated models are of polynomial degree to obtain linear and quadratic response (<em>e.g.</em>, Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006). From these results, we ran ecological response curve models, where the resulting plots included the probability of occurrence and values for continuous variables or categories for discrete variables. The points of the presence and background training group are also included.</p> <p>On the other hand, a global GLM was also run, from which the generalized model is evaluated by means of a 2 x 2 contingency matrix, including both observed and predicted records. A representation of this is shown in Table 1 (adapted from Allouche et al., 2006). In this process we select an arbitrary boundary of 0.5 to obtain better modeling performance and avoid high percentage of bias in type I (omission) or II (commission) errors (e.g., Carpenter et al., 1993; Fielding and Bell, 1997; Allouche et al., 2006; Kim, 2009; Hijmans and Elith, 2017).</p> <p>Table 1. Example of 2 x 2 contingency matrix for calculating performance metrics for GLM models. A represents true presence records (true positives), B represents false presence records (false positives - error of commission), C represents true background points (true negatives) and D represents false backgrounds (false negatives - errors of omission).</p> <table align="center"> <tbody> <tr> <td> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> </td> <td> <p>Validation set</p> </td> </tr> <tr> <td> <p>Model</p> </td> <td> <p>True</p> </td> <td> <p>False</p> </td> </tr> <tr> <td> <p>Presence</p> </td> <td> <p>A</p> </td> <td> <p>B</p> </td> </tr> <tr> <td> <p>Background</p> </td> <td> <p>C</p> </td> <td> <p>D</p> </td> </tr> </tbody> </table> <p>We then calculated the Overall and True Skill Statistics (TSS) metrics. The first is used to assess the proportion of correctly predicted cases, while the second metric assesses the prevalence of correctly predicted cases (Olden and Jackson, 2002). This metric also gives equal importance to the prevalence of presence prediction as to the random performance correction (Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006).</p> <p>The last code (<em>i.e.</em>, <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code8_DOMAIN_SuitHab_model.R?versionId=d951a8f2-d3a4-4804-b862-1b2762061876">Code8_DOMAIN_SuitHab_model.R</a>) is for species distribution modelling using the DOMAIN algorithm (Carpenter <em>et al.</em>, 1993). Here, we loaded the variable stack and the presence and background group subdivided into 75% training and 25% test, each. We only included the presence training subset and the predictor variables stack in the calculation of the DOMAIN metric, as well as in the evaluation and validation of the model.</p> <p>Regarding the model evaluation and estimation, we selected the following estimators:</p> <p>1) partial ROC, which evaluates the approach between the curves of positive (<em>i.e.</em>, correctly predicted presence) and negative (i.e., correctly predicted absence) cases. As farther apart these curves are, the model has a better prediction performance for the correct spatial distribution of the species (Manzanilla-Qui&ntilde;ones, 2020).</p> <p>2) ROC/AUC curve for model validation, where an optimal performance threshold is estimated to have an expected confidence of 75% to 99% probability (De Long <em>et al.</em>, 1988).</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

Discretized U.S. drought data to support statistical modeling

<p>Drought is a costly and disruptive natural disaster, with widespread implications for agriculture, wildfire, and urban planning.  We present a novel data set on US drought built to enable computationally efficient spatio-temporal statistical and probabilistic models of drought. We converted drought data obtained from the widely-used US Drought Monitor (USDM) from continuous shape files to a 0.5-degree regular lattice. These data cover the Continental US from 2000 to mid-2022. Known environmental drivers of drought include those obtained from the North American Land Data Assimilation System (NLDAS-2), US Geological Survey (USGS) streamflow data, and National Oceanic and Atmospheric Administration (NOAA) teleconnections data. The USGS streamflow data is itself a new gridded data product, aggregating point-referenced stream discharges from across the US to a common lattice using watersheds to combine nearby stream data. The resulting data set permits statistical and probabilistic modeling of drought with explicit spatial and/or temporal dependence.  Such models could be used to forecast short-range and even season-to-season future droughts with uncertainty, extending the reach and value of the current US Drought Outlook produced by the National Weather Service Climate Prediction Center. </p>

opencc-zeroMay 2023View details →
zenodo40/100

Data points for "Modelling sorption of hydrocarbons in polyethylene with the SAFT-γ Mie approach combined with a statistical-mechanical model to describe semi-crystalline polymers"

<p>A variety of thermodynamic calculations (VLE, sorption isotherms, etc.) performed&nbsp;with a combination of the SAFT-&gamma; equation of state and a novel model to account for the constraints affecting the amorphous domains in semi-crystalline polyethylene (PE). Please refer to the original article (published in Macromolecules) for the bibliography and more details.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Data set: Statistically parameterizing and evaluating a positive degree-day model to estimate surface melt in Antarctica from 1979 to 2022

<p><strong>Version 2:</strong></p> <p><strong>Updates from version 1: Monthly, daily, and hourly dist-PDD and uni-PDD outputs have been added.</strong></p> <p><strong>https://doi.org/10.5194/tc-17-3667-2023</strong></p> <p>&nbsp;</p> <p>Version 1:</p> <p>This dataset accompanies Zheng et al. (2023):&nbsp;Statistically parameterizing and evaluating a positive degree-day<br> model to estimate surface melt in Antarctica from 1979 to 2022, The Cryosphere.</p> <p>This dataset contains annual PDD model output.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record