Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
173
datasets available to search
ShareScore release 0.9.0
Dataset results
173 results for “Statistical analysis”
Data and statistical analysis for: Bacterial nanotubes are a manifestation of cell death
<p>Contains all data and code to reproduce the statistical analysis in the Supplementary file 2 for the paper "Bacterial nanotubes are a manifestation of cell death" to be published in Nature Communications.</p> <p><strong>Contents:</strong></p> <ul> <li>statistical_analysis.Rmd is the main document written in R Markdown</li> <li>statistical_analysis.html is a compiled version of statistical_analysis.Rmd showing all the computed results.</li> <li>The Source data.xls file contains raw data used for the analysis - the sheet names indiciate the figure they refer to. See the main file for code that can read the data.</li> </ul> <p>The code can also be accessed at <a href="https://github.com/cas-bioinf/nanotubes-death">https://github.com/cas-bioinf/nanotubes-death</a></p>
Genome-wide association summary statistics for sex- and age-specific analysis of chronic back pain
<p>The dataset comprises summary-level statistics for age- and sex-specific genome-wide association study of chronic back pain (cBP) in individuals of European descent from UK Biobank (<a href="https://www.ukbiobank.ac.uk/">https://www.ukbiobank.ac.uk/</a>). The study was carried out under UK Biobank approved project #18219. </p> <p><strong>The dataset accompanies the paper (please cite if using the dataset):</strong></p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/33021770/">Freidin, Maxim B.; Tsepilov, Yakov A.; Stanaway, Ian B.; Meng, Weihua; Hayward, Caroline; Smith, Blair H.; Khoury, Samar; Parisien, Marc; Bortsov, Andrey; Diatchenko, Luda; Børte, Sigrid; Winsvold, Bendik S.; Brumpton, Ben M.; Zwart, John-Anker; HUNT All-In Pain; Aulchenko, Yurii S.; Suri, Pradeep; Williams, Frances M.K. Sex- and age-specific genetic analysis of chronic back pain. Pain. 2020. doi:10.1097/j.pain.0000000000002100.</a></p> <p>The phenotype of cBP was defined as back pain for 3+ months. Linear mixed-effects additive model was fitted adjusting for age, genotyping array type, and 10 genetic PCs provided by UK Biobank. The following filters were applied: minor allele frequency >0.001, genotyping and individual call rates >0.98%, imputation quality score (INFO) >0.7. GWAS were carried out in males and females separately in the whole sample (<strong>allages</strong>) as well as in groups of younger than 65 years (<strong>under65</strong>) and 65+ years old (<strong>65plus</strong>) as detailed in the paper. Accordingly, 6 files are deposited here, corresponding to each group. </p> <p><strong>Column headers:</strong></p> <p>SNP, SNP rsID </p> <p>CHR, chromosome</p> <p>BP, genomic position (GRCh37 build)</p> <p>EA, effect allele (coded as "1")</p> <p>OTHER, other allele (coded as "0")</p> <p>A1FREQ, frequency of effect allele</p> <p>INFO, imputation quality</p> <p>BETA, effect size (for effect allele)</p> <p>SE, standard error of effect size</p> <p>PVAL, p-value for association</p>
Training dataset: Statistical analysis of a HEK/Ecoli Spike-in DIA dataset using MSstats
<p>The uploaded files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 °C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in two different ratios and four replicates of each Spike/in ratio were measured and analysed using OpenSwathWorkflow in Galaxy. Results were exported using PyProphet and can be used for the statistical analysis and detection of the two different Spike-in Ratios. The Spike-in ratios were the following:</p> <p>Sample HEK E.coli <br> Spike_in_1 2.5 0.15<br> Spike_in_2 2.5 0.80 </p> <p>Besides the two PyProphet export files, we uploaded a sample annotation file as well as a comparison matrix file.<br> Additionally, we uploaded the Galaxy MSstats training result files: MSstats_ComparisonResult_export_tabular and MSstats_ComparisonResult_msstats_input.</p>
Data for analysis in "Towards optimal cosmological parameter recovery from compressed bispectrum statistics"
<p>Measures of the three point function extracted from a suite of simulations using 4 different estimators: namely, the bispectrum, modal estimator, integrated bispectrum, line correlation function. Also supplied are power spectrum measures across the same simulations. <br> <br> The measures are done across 200 fiducial and 60 non-fiducial cosmology simulations, at 3 redshifts. Further details on what was done can be attained by reading the document, Overview.md/Overview.pdf, attached to the bundle. Even more details can be acquired by reading the paper this data was prepared for at https://arxiv.org/abs/1705.04392! </p>
Dataset, statistical analysis code, and supplementary material of juvenile ravens' responses towards acoustic cues of different social categories
<p>Social competence i.e., defined as the ability to adjust the expression of social behaviour to the available social information, is known to be influenced by early-life conditions. Brood size might be one of the factors determining such early conditions, particularly in species with extended parental care. We here tested in ravens, whether growing up in families of different sizes affects the chicks' responsiveness to social information. We experimentally manipulated the brood size of 20 captive raven families, creating either small or large families. Simulating dispersal, juveniles were separated from their parents and temporarily housed in one of two captive non-breeder groups. After five weeks of socialization, each raven was individually tested in a playback setting with food-associated calls from three social categories: sibling, familiar unrelated raven they were housed with, and unfamiliar unrelated raven from the other non-breeder aviary. We found that individuals reared in small families were more attentive than birds from large families, in particular towards the familiar unrelated peer. These results indicate that variation in family size during upbringing can affect how juvenile ravens value social information. Whether the observed attention patterns translate into behavioural preferences under daily life conditions remains to be tested in future studies.</p>
The Relationship between LRP5 (rs556442 and rs638051) Polymorphisms and Mutation with Bone Metabolism in Xinjiang women with Type 2 Diabetes after Menopause(Table 1 and Table 2 Statistical Values of Analysis Process)
<p>The Relationship between LRP5 (rs556442 and rs638051) Polymorphisms and Mutation with Bone Metabolism in Xinjiang women with Type 2 Diabetes after Menopause(Table 1 and Table 2 Statistical Values of Analysis Process)</p>
Optimized summary-statistic-based single-cell meta-analysis. Input files
<p>This dataset contains information about the input files used in the Optimized summary-statistic-based single-cell meta-analysis research project. </p> <p> </p>
The New Acropolis Museum: Short Statistical Analysis for a Sustainable Operation with Active Visitors Based on a Sample of Students of the University of Athens
<p>The purpose of this study is to examine the new Acropolis Museum and its potential visits, with a focus on the number of students at the University of Athens visiting it. The dimensions studied are the new museum’s functionality, accessibility, and the intention to and motives for visiting it. The new museum has been open for 14 years and is viewed as a symbol of an exceptional cultural experience by both Greek and foreign visitors. As the focus of our field study, the students replied to mainly quantitative questions via computer, and we then performed a statistical analysis of their replies.</p>
Supplemental Materials to Accompany: Abowd and Schmutte "An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices"
<p>Materials to supplement "<a href="https://www.aeaweb.org/articles?id=10.1257/aer.20170627">An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices</a>" by John Abowd and Ian Schmutte.</p>
Statistical analysis of chlorate occurrence data in food
<p>In accordance with Article 29 (1) (a) of Regulation (EC) No 178/2002, the European Commission asked the European Food Safety Authority (EFSA) in 2014 for a scientific opinion on the risks for human health related to the presence of chlorate in food from all sources, taking also into account its presence in drinking water. The opinion found that “Chronic exposure of adolescent and adult age classes did not exceed the TDI. However, at the 95th percentile the TDI was exceeded in all surveys in ‘Infants’ and ‘Toddlers’ and in some surveys in ‘Other children’. Chronic exposures are of concern in particular in younger age groups with mild or moderate iodine deficiency.” Food manufacturers have started to optimise their manufacturing processes to lower chlorate residue level in foods and the European Commission in 2017 provided a revised Guidance document on good hygiene practices regarding the use of chlorinated disinfectant. It can therefore be expected that chlorate levels in foods are now lower compared with the levels found in the samples from 2011 to 2014. The European Commission (EC) requested EFSA in 2018 to provide an updated statistical analysis on chlorate occurrence levels in foods. A set of 14 Excel tables containing the statistical analysis of reported results for the analysis of chlorates from pesticides monitoring and contaminants monitoring programmes have been prepared. The data is presented at three levels of aggregation using FoodEx product categories. Analysis was performed for two time points, the 2011-2017 dataset contained 15,741 valid results and the 2015-2017 dataset contained 28,033 valid results. Mean, median and percentile (75th, 90th, 95th) for lower bound, middle bound and upper bound concentration values were calculated. Caution should be applied to percentile values calculated from a limited number of results.</p>
Supporting data for: "Diaphysator: an online application for the exhaustive cartography and user-friendly statistical analysis of long bone diaphyses"
<p>Example of dataset to be used with the R-shiny application “Diaphysator”, composed of right tibiae and femora.</p> <p>These data file have been published in: Lacoste Jeanson, A., Santos, F., Villa, C., Banner, J., & Bruzek, J. (2018). Architecture of the femoral and tibial diaphyses in relation to body mass and composition: Research from whole-body CT. <em>American Journal of Physical Anthropology</em>, 167, 813– 826. doi: <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/ajpa.23713">10.1002/ajpa.23713</a></p> <p>This zip file contains:</p> <ul> <li>an “Information file” in CSV format</li> <li>various data files for human femora and tibiae in CSV format</li> </ul> <p>For all CSV files, the field separator is the comma “,” and the character used for decimal points is the dot “.”</p>
Replication package for "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Replication package for "Evolution of statistical analysis in ESE research"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
FIGURE 2 in RNames, a stratigraphical database designed for the statistical analysis of fossil occurrences - the Ordovician diversification as a case study
FIGURE 2. Structure of algorithm for time binning of stratigraphical units of the RNames Database (available under https://github.com/bjoekroe/RNames). Time bins are selected via three correlation routes (colour codes) and six rules resulting in six tables with referenced bins from which only those are selected which are most precise (i.e., range through lowest number of bins). Abbreviations: bio.unit, biostratigraphic unit; non-bio. unit, non-biostratigraphic unit. Colour code: red, correlation exclusively based on biostratigraphy; orange; correlation indirectly based on biostratigraphy; yellow, correlation based on direct or indirect assignments to time bins. -> arrow refers to referenced relations in RNames.
FIGURE 1 in RNames, a stratigraphical database designed for the statistical analysis of fossil occurrences - the Ordovician diversification as a case study
FIGURE 1. Simplified structure of the RNames Database (rnames.luomus.fi/). The database contains eight related tables (blue and red objects) of which the object "Relations" is central. In "Relations" correlated stratigraphic units are listed by reference. Three output tables (yellow objects) list time binned stratigraphic units based on a search algorithm that uses "Relations" via R-Package RMySQL (the scripts are available under https://github.com/bjoekroe/ RNames). Global Stages after Cooper et al. (2012). Abbreviations: ID, identifier; StS, Stage Slice (Bergström et al., 2009); TS, Time Slice (Webby et al., 2004)
FIGURE 5 in RNames, a stratigraphical database designed for the statistical analysis of fossil occurrences - the Ordovician diversification as a case study
FIGURE 5. Quality of PaleobioDB data used for diversity calculations. 1. Number of collections available per time bin. 2. Mean stratigraphic range of collections through time bins. Diamonds, two-time-bin resolution; triangles, one-time bin resolution; squares, all collections. Red, Global Stages after Cooper et al. (2012), green; Stage Slices, Bergström et al. (2009); blue, Time Slices, Webby et al. (2004).
FIGURE 4 in RNames, a stratigraphical database designed for the statistical analysis of fossil occurrences - the Ordovician diversification as a case study
FIGURE 4. Ordovician genus-level diversity trends of PaleobioDB occurrence data, based on three different time binning approaches. 1. Total mean standing diversity (after Cooper, 2004). 2. Rarefied diversity with time bins of <100 collections culled, with quota 600. Diamonds, two-time-bin resolution; triangles, one-time bin resolution; stars, all collections. Red, Global Stages after Cooper et al. (2012), green; Stage Slices, Bergström et al. (2009); blue, Time Slices, Webby et al. (2004). Error bars reflect 95% confidence interval.
Statistical analysis code for output from a model used to simulate foot-and-mouth disease dynamics in the United Kingdom
<p>Epidemics can sometimes be managed through reductions of host density, such as social distancing for human diseases, reducing plant density through cultural and genetic means, and host culling for epizootics. These approaches allow for a certain density of hosts to remain within a targeted area. By contrast, total ring depopulation is often used as a management strategy for emerging infectious diseases in livestock. In this study, we explore the trade-offs of a density-based culling strategy to determine if fewer livestock farms can be culled within rings while maintaining a decrease in disease transmission. To do so, we evaluated a farm-density-based ring culling strategy to control foot-and-mouth disease (FMD) in the United Kingdom. This strategy may allow for some farms within rings around infected premises (IPs) to escape depopulation, with the aim to prevent over-culling during outbreaks. Using a spatially-explicit, stochastic, state-transition simulation algorithm originally developed by Keeling et al. 2001 to model FMD spread in the United Kingdom, we simulated this reduced-farm-density, or "target density" strategy. We modeled FMD disease spread in four counties in the UK (Aberdeenshire, Cumbria, Devon, and North Yorkshire) that have different farm demographies. We ran 740,000 simulations in a full-factorial analysis of epidemic impact measurements (i.e. culled animals, culled farms, epidemic length) and cull strategy parameters (i.e. target farm density, daily farm cull capacity, cull radius). We found that all of the cull strategy parameters were drivers of epidemic impact. We found that outbreaks in Cumbria had higher epidemic impacts and were more likely to take off compared with other counties with more outbreaks being likely to take off in Cumbria. Most importantly, in all counties, our proposed target density strategy was more effective at combatting FMD compared with traditional 'total ring depopulation' when considering average culled animals and culled farms. The differences in epidemic impact between the counties are likely driven by farm demography, especially differences in cattle and farm density. This target density strategy can be applied to many different systems, including other livestock and agricultural systems, to reduce host density as opposed to over-culling hosts.</p>
Immersive haptic simulation for training nurses in emergency medical procedures - Data collected and statistical analysis
<p>Data collected during the evaluation presented in "Haptic simulation for emergency procedures in nursing training" paper.</p> <table> <caption>HR ALL</caption> <thead> <tr> <th>Measure 1</th> <th> </th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>2.857</td> <td>29</td> <td>0.008</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-8.089</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>7.567</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-2.962</td> <td>29</td> <td>0.006</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Paired samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>HR FIRST MANN</caption> <thead> <tr> <th>Measure 1</th> <th> </th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>1.665</td> <td>14</td> <td>0.118</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-7.104</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>6.498</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-1.461</td> <td>14</td> <td>0.166</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Paired samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>HR FIRST VR</caption> <thead> <tr> <th>Measure 1</th> <th> </th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>2.341</td> <td>14</td> <td>0.035</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-4.612</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>4.482</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-2.688</td> <td>14</td> <td>0.018</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Paired samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>HR BETWEEN GROUPS</caption> <thead> <tr> <th> </th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-1.958</td> <td>28</td> <td>0.060</td> </tr> <tr> <td>Mann post HR</td> <td>-1.902</td> <td>28</td> <td>0.068</td> </tr> <tr> <td>VR pre HR</td> <td>-4.013</td> <td>28</td> <td>< .001</td> </tr> <tr> <td>VR post HR</td> <td>-2.344</td> <td>28</td> <td>0.026</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Independent samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>Physiological T-Test results for the participants who started the experiment performing the procedure in the mannequin.</caption> <thead> <tr> <th>First variable</th> <th>μ</th> <th>σ</th> <th>Second variable</th> <th>μ</th> <th>σ</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>SBP pre-mannequin</td> <td>128.333</td> <td>10.715</td> <td>SBP pre-simulator</td> <td>134.533</td> <td>11.819</td> <td>-1.870</td> <td>14</td> <td>0.083</td> </tr> <tr> <td>SBP post-mannequin</td> <td>125.600</td> <td>11.648</td> <td>SBP post-simulator</td> <td>131.467</td> <td>14.643</td> <td>-2.094</td> <td>14</td> <td>0.055</td> </tr> <tr> <td>DBP pre-mannequin</td> <td>80.133</td> <td>5.527</td> <td>DBP pre-simulator</td> <td>81.533</td> <td>9.039</td> <td>-0.623</td> <td>14</td> <td>0.544</td> </tr> <tr> <td>DBP post-mannequin</td> <td>78.667</td> <td>6.956</td> <td>DBP post-simulator</td> <td>81.400</td> <td>8.475</td> <td>-2.073</td> <td>14</td> <td>0.057</td> </tr> <tr> <td>HR pre-mannequin</td> <td>92.133</td> <td>14.837</td> <td>HR pre-simulator</td> <td>75.733</td> <td>9.9625</td> <td>6.498</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>HR post-mannequin</td> <td>87.400</td> <td>9.132</td> <td>HR post-simulator</td> <td>91.400</td> <td>14.217</td> <td>-1.461</td> <td>29</td> <td>0.166</td> </tr> </tbody> </table> <p>SBP = Systolic blood pressure. DBP = Diastolic blood pressure. HR = Heart Rate.</p> <table> <caption>Physiological T-Test results for the participants who started the experiment performing the procedure in the ParaVR simulator.</caption> <thead> <tr> <th>First variable</th> <th>μ</th> <th>σ</th> <th>Second variable</th> <th>μ</th> <th>σ</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>SBP pre-mannequin</td> <td>119.067</td> <td>12.898</td> <td>SBP pre-simulator</td> <td>130.600</td> <td>12.188</td> <td>-3.799</td> <td>14</td> <td>0.002</td> </tr> <tr> <td>SBP post-mannequin</td> <td>117.533</td> <td>13.410</td> <td>SBP post-simulator</td> <td>128.200</td> <td>13.385</td> <td>-4.022</td> <td>14</td> <td>0.001</td> </tr> <tr> <td>DBP pre-mannequin</td> <td>76.533</td> <td>8.943</td> <td>DBP pre-simulator</td> <td>80.200</td> <td>6.899</td> <td>-1.815</td> <td>14</td> <td>0.091</td> </tr> <tr> <td>DBP post-mannequin</td> <td>74.333</td> <td>8.541</td> <td>DBP post-simulator</td> <td>79.133</td> <td>7.864</td> <td>-2.003</td> <td>14</td> <td>0.065</td> </tr> <tr> <td>HR pre-mannequin</td> <td>102.067</td> <td>12.876</td> <td>HR pre-simulator</td> <td>91.533</td> <td>11.825</td> <td>4.482</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>HR post-mannequin</td> <td>95.867</td> <td>14.623</td> <td>HR post-simulator</td> <td>103.667</td> <td>14.450</td> <td>-2.688</td> <td>29</td> <td>0.018</td> </tr> </tbody> </table>
Codes in R for spatial statistics analysis, ecological response models and spatial distribution models
<p>In the last decade, a plethora of algorithms have been developed for spatial ecology studies. In our case, we use some of these codes for underwater research work in applied ecology analysis of threatened endemic fishes and their natural habitat. For this, we developed codes in Rstudio® script environment to run spatial and statistical analyses for ecological response and spatial distribution models (e.g., Hijmans & Elith, 2017; Den Burg <em>et al.</em>, 2020). The employed R packages are as follows: caret (Kuhn et al., 2020), corrplot (Wei & Simko, 2017), devtools (Wickham, 2015), dismo (Hijmans & Elith, 2017), gbm (Freund & Schapire, 1997; Friedman, 2002), ggplot2 (Wickham et al., 2019), lattice (Sarkar, 2008), lattice (Musa & Mansor, 2021), maptools (Hijmans & Elith, 2017), modelmetrics (Hvitfeldt & Silge, 2021), pander (Wickham, 2015), plyr (Wickham & Wickham, 2015), pROC (Robin et al., 2011), raster (Hijmans & Elith, 2017), RColorBrewer (Neuwirth, 2014), Rcpp (Eddelbeuttel & Balamura, 2018), rgdal (Verzani, 2011), sdm (Naimi & Araujo, 2016), sf (e.g., Zainuddin, 2023), sp (Pebesma, 2020) and usethis (Gladstone, 2022).</p> <p>It is important to follow all the codes in order to obtain results from the ecological response and spatial distribution models. In particular, for the ecological scenario, we selected the Generalized Linear Model (GLM) and for the geographic scenario we selected DOMAIN, also known as Gower's metric (Carpenter <em>et al.</em>, 1993). We selected this regression method and this distance similarity metric because of its adequacy and robustness for studies with endemic or threatened species (<em>e.g.</em>, Naoki <em>et al.</em>, 2006). Next, we explain the statistical parameterization for the codes immersed in the GLM and DOMAIN running:</p> <p>In the first instance, we generated the background points and extracted the values of the variables (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code2_Extract_values_DWp_SC.R?versionId=c1ea0c61-53fe-4f95-ab88-0c1cb28399cb">Code2_Extract_values_DWp_SC.R</a>). Barbet-Massin <em>et al. </em>(2012) recommend the use of 10,000 background points when using regression methods (<em>e.g.</em>, Generalized Linear Model) or distance-based models (<em>e.g.</em>, DOMAIN). However, we considered important some factors such as the extent of the area and the type of study species for the correct selection of the number of points (Pers. Obs.). Then, we extracted the values of predictor variables (<em>e.g.</em>, bioclimatic, topographic, demographic, habitat) in function of presence and background points (<em>e.g.</em>, Hijmans and Elith, 2017).</p> <p>Subsequently, we subdivide both the presence and background point groups into 75% training data and 25% test data, each group, following the method of Soberón & Nakamura (2009) and Hijmans & Elith (2017). For a training control, the 10-fold (cross-validation) method is selected, where the response variable presence is assigned as a factor. In case that some other variable would be important for the study species, it should also be assigned as a factor (Kim, 2009).</p> <p>After that, we ran the code for the GBM method (Gradient Boost Machine; <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code3_GBM_Relative_contribution.R?versionId=1656bbae-66aa-409e-bb91-d8007dee8f95">Code3_GBM_Relative_contribution.R</a> and <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code4_Relative_contribution.R?versionId=0e1d9352-e6b2-43da-984b-d6853a914258">Code4_Relative_contribution.R</a>), where we obtained the relative contribution of the variables used in the model. We parameterized the code with a Gaussian distribution and cross iteration of 5,000 repetitions (<em>e.g.</em>, Friedman, 2002; kim, 2009; Hijmans and Elith, 2017). In addition, we considered selecting a validation interval of 4 random training points (Personal test). The obtained plots were the partial dependence blocks, in function of each predictor variable.</p> <p>Subsequently, the correlation of the variables is run by Pearson's method (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code5_Pearson_Correlation.R?versionId=275f8dd4-b056-44d2-bfe5-f6264bc3298b">Code5_Pearson_Correlation.R</a>) to evaluate multicollinearity between variables (Guisan & Hofer, 2003). It is recommended to consider a bivariate correlation ± 0.70 to discard highly correlated variables (<em>e.g.</em>, Awan <em>et al.</em>, 2021).</p> <p>Once the above codes were run, we uploaded the same subgroups (<em>i.e.</em>, presence and background groups with 75% training and 25% testing) (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code6_Presence&backgrounds.R?versionId=d797b528-782f-4a19-bd61-cfb197f38513">Code6_Presence&backgrounds.R</a>) for the GLM method code (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code7_GLM_model.R?versionId=e4aca276-d601-49ec-a62c-a9223b05a7ed">Code7_GLM_model.R</a>). Here, we first ran the GLM models per variable to obtain the <em>p</em>-significance value of each variable (alpha ≤ 0.05); we selected the value one (<em>i.e.</em>, presence) as the likelihood factor. The generated models are of polynomial degree to obtain linear and quadratic response (<em>e.g.</em>, Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006). From these results, we ran ecological response curve models, where the resulting plots included the probability of occurrence and values for continuous variables or categories for discrete variables. The points of the presence and background training group are also included.</p> <p>On the other hand, a global GLM was also run, from which the generalized model is evaluated by means of a 2 x 2 contingency matrix, including both observed and predicted records. A representation of this is shown in Table 1 (adapted from Allouche et al., 2006). In this process we select an arbitrary boundary of 0.5 to obtain better modeling performance and avoid high percentage of bias in type I (omission) or II (commission) errors (e.g., Carpenter et al., 1993; Fielding and Bell, 1997; Allouche et al., 2006; Kim, 2009; Hijmans and Elith, 2017).</p> <p>Table 1. Example of 2 x 2 contingency matrix for calculating performance metrics for GLM models. A represents true presence records (true positives), B represents false presence records (false positives - error of commission), C represents true background points (true negatives) and D represents false backgrounds (false negatives - errors of omission).</p> <table align="center"> <tbody> <tr> <td> <p> </p> </td> <td> <p>Validation set</p> </td> </tr> <tr> <td> <p>Model</p> </td> <td> <p>True</p> </td> <td> <p>False</p> </td> </tr> <tr> <td> <p>Presence</p> </td> <td> <p>A</p> </td> <td> <p>B</p> </td> </tr> <tr> <td> <p>Background</p> </td> <td> <p>C</p> </td> <td> <p>D</p> </td> </tr> </tbody> </table> <p>We then calculated the Overall and True Skill Statistics (TSS) metrics. The first is used to assess the proportion of correctly predicted cases, while the second metric assesses the prevalence of correctly predicted cases (Olden and Jackson, 2002). This metric also gives equal importance to the prevalence of presence prediction as to the random performance correction (Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006).</p> <p>The last code (<em>i.e.</em>, <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code8_DOMAIN_SuitHab_model.R?versionId=d951a8f2-d3a4-4804-b862-1b2762061876">Code8_DOMAIN_SuitHab_model.R</a>) is for species distribution modelling using the DOMAIN algorithm (Carpenter <em>et al.</em>, 1993). Here, we loaded the variable stack and the presence and background group subdivided into 75% training and 25% test, each. We only included the presence training subset and the predictor variables stack in the calculation of the DOMAIN metric, as well as in the evaluation and validation of the model.</p> <p>Regarding the model evaluation and estimation, we selected the following estimators:</p> <p>1) partial ROC, which evaluates the approach between the curves of positive (<em>i.e.</em>, correctly predicted presence) and negative (i.e., correctly predicted absence) cases. As farther apart these curves are, the model has a better prediction performance for the correct spatial distribution of the species (Manzanilla-Quiñones, 2020).</p> <p>2) ROC/AUC curve for model validation, where an optimal performance threshold is estimated to have an expected confidence of 75% to 99% probability (De Long <em>et al.</em>, 1988).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.