Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
466
datasets available to search
ShareScore release 0.9.0
Dataset results
466 results for “code analysis”
Genetic analysis of mycobacteria isolated from suspect bovine tuberculosis lesions in Wolaita, Ethiopia - code and datasets.
<p>Bovine tuberculosis (bTB), caused by Mycobacterium bovis and other members of the Mycobacterium tuberculosis complex (MTBC), is a significant concern for livestock and public health in Ethiopia. This study aimed to assess the prevalence and causative agents of bTB in cattle from four abattoirs in the Wolaita region of Ethiopia. </p>
MITgcm simulations of sea level response to freshwater injected at the surface and at depth in southern high latitudes: Model output and analysis code
<p>Model output (netcdf) and python code (included in both py and ipynb formats) to create the figures in Eisenman et al. (2024).</p> <div> <p>See https://eisenman-group.github.io for further details.</p> </div>
Data and analysis code for Repo et al., "Contrasting forest management strategies: impacts on biodiversity and ecosystem services under changing climate and disturbance regimes"
<p>This repository contains analysis code and pre-processed data for the study "Contrasting forest management strategies: impacts on biodiversity and ecosystem services under changing climate and disturbance regimes" by Repo et al.<br>Data processing and analysis mainly done by Aapo Jantunen, Katharina Albrich<br>Due to respository space limitations, the original model outputs are archived in the Finnish "Allas" data storage service. For access, contact katharina.albrich@luke.fi<br>The code used to process the raw data is included here for reproducibility.</p> <p>If you are interested in using iLand, visit https://iland-model.org/ and https://iland-model.org/iland-book/ for information on using the model and a guide to setting up a landscape.</p> <p><span>This work was supported by the Ministry of Agriculture and Forestry by funding project Future multifunctional forests and their disturbance risk in the changing climate (Foster) through the “Catch the Carbon” initiative (<span>project number VN/28654/2020)</span>. A.R. has been supported by the grant [TRACY Trade-offs and synergies in land-based climate change mitigation and biodiversity conservation decision 322066 by the Academy of Finland.], J. H by the grant [CASCADE - Changing Disturbance Regimes and Forest Landscapes of Fennoscandia 342569 by the Academy of Finland]. </span></p> <p> </p>
Data and analysis code for Zhou et al. 2024, Global Change Biology
<p>This dataset accompanies the paper:</p> <p>Zhou, J., Zhu, P., Kluger, D.M., Lobell, D.B., and Jin, Z. 2024. Changes in the yield effect of the preceding crop in the US Corn Belt under a warming climate. Global Change Biology</p>
Datasets and R Code for native-invasive shrub analysis
<p>Datasets and R code to reproduce statistical analysis comparing functional traits of native and invasive shrub and liana species of North America.</p>
Statistical analysis code for output from a model used to simulate foot-and-mouth disease dynamics in the United Kingdom
<p>Epidemics can sometimes be managed through reductions of host density, such as social distancing for human diseases, reducing plant density through cultural and genetic means, and host culling for epizootics. These approaches allow for a certain density of hosts to remain within a targeted area. By contrast, total ring depopulation is often used as a management strategy for emerging infectious diseases in livestock. In this study, we explore the trade-offs of a density-based culling strategy to determine if fewer livestock farms can be culled within rings while maintaining a decrease in disease transmission. To do so, we evaluated a farm-density-based ring culling strategy to control foot-and-mouth disease (FMD) in the United Kingdom. This strategy may allow for some farms within rings around infected premises (IPs) to escape depopulation, with the aim to prevent over-culling during outbreaks. Using a spatially-explicit, stochastic, state-transition simulation algorithm originally developed by Keeling et al. 2001 to model FMD spread in the United Kingdom, we simulated this reduced-farm-density, or "target density" strategy. We modeled FMD disease spread in four counties in the UK (Aberdeenshire, Cumbria, Devon, and North Yorkshire) that have different farm demographies. We ran 740,000 simulations in a full-factorial analysis of epidemic impact measurements (i.e. culled animals, culled farms, epidemic length) and cull strategy parameters (i.e. target farm density, daily farm cull capacity, cull radius). We found that all of the cull strategy parameters were drivers of epidemic impact. We found that outbreaks in Cumbria had higher epidemic impacts and were more likely to take off compared with other counties with more outbreaks being likely to take off in Cumbria. Most importantly, in all counties, our proposed target density strategy was more effective at combatting FMD compared with traditional 'total ring depopulation' when considering average culled animals and culled farms. The differences in epidemic impact between the counties are likely driven by farm demography, especially differences in cattle and farm density. This target density strategy can be applied to many different systems, including other livestock and agricultural systems, to reduce host density as opposed to over-culling hosts.</p>
Figure 6 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 6. Plots of length against PC1 for all eight species of Lavigeria studied with regressed lines and R2 values. Length values are log transformed.
Figure 3 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 3. Character state APPL. An adult specimen of L. n. sp. X (left) and a juvenile (right). Notice the perimetric, wrinkle-like, lines on the front surface of the apertural lip of the adult. Scale bar = 0.2 cm.
Figure 7 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 7. Fifty per cent majority-rule consensus trees of five, nine, ten and 31 trees (from top to bottom, respectively) from four matrices. Matrices are coding the data of Table 1. See Analysis for explanation of the matrices. Tree and character statistics are given in Table 3. Optimality criterion: maximum parsimony, exhaustive search. All characters binary, of equal weight and unordered. Numbers indicate percentage of topologies that include the respective branches.
Figure 5 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 5. Character states DFST, AXRB and UEPW. Adult specimens of (A) L. n. sp. W, (B) L. n. sp. F, (C) L. n. sp. K, (D) L. n. sp. J showing the aperture in side view and (E) an apertural view of an adult L. n. sp. W. In A-D the trajectory of the suture tends to deviate downwards in comparison to the trajectory of the spiral cord of the previous whorl immediately above the suture (DFST). A and D also show the loss of, or irregularities in the appearance of axial sculpture (AXRB). In E, arrowheads show the undulations formed at the edge of the parietal side of the aperture (UEPW). Scale bar = 0.2 cm.
Figure 2 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 2. Character states WGPW and APLT. An adult specimen of L. n. sp. A (left) and a juvenile (right). The adult shows a thickened (APLT) and opaque (WGPW) inner surface of the apertural lip in comparison to the juvenile. Scale bar = 0.2 cm.
Figure 1 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 1. Eight species used in this study, apertural and side views of adult specimens. (A) Lavigeria new species N. (B) L. n. sp. F. (C) L. n. sp. J. (D) L. n. sp. X. (E) L. n. sp. K. (F) L. n. sp. W. (G) L. n. sp. A. (H) L. n. sp. U. C-H all belong to the same clade. A and B belong to different clades within the genus. Scale bar = 0.2 cm.
Figure 4 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 4. Character state APDT. An adult specimen of L. n. sp. J (left) and a juvenile on the right. Notice in the adult how the parietal side of the apertural lip is completely detached from the previous whorl and a false umbilicus has developed. Scale bar = 0.2 cm.
Data and code for analysis in "Fighting over defence chemicals disrupts mating behaviour"
<p>Data and annotated code for analysis in "Fighting over defence chemicals disrupts mating behaviour". The point at which each data sheet is used in the analysis is specified in the code and code for each respective figure in paper is also given. A renv lockfile is also included for version control, but all package versions are also included in paper's methods section.</p>
QTL mapping and transcriptome analysis of Sclerotinia-resistance in the wild cabbage species Brassica oleracea var. villosa [Main code]
<p>This is the main code supplement for my computational analysis for the manuscript: "QTL mapping and transcriptome analysis of Sclerotinia-resistance in the wild cabbage species <em>Brassica oleracea </em>var<em>. villosa".</em> The main code is availabe in separate html-files. DOI will be added if available.</p>
Supplemental Data and Code for "An exact version of Life Table Response Experiment analysis, and the R package exactLTRE"
<p>This dataset enables the user to repeat the analyses presented in the manuscript "An exact version of Life Table Response Experiment analysis, and the R package exactLTRE." It is comprised of two compressed archives: one which contains code, and one which contains data.</p>
Data and code for reproducing analysis in 'Producing indicative allocations for Community Led Local Development funding in Scotland (2022-23)'
<p>The data and code in this folder can be used to reproduce work used to generate indicative allocations of Community Led Local Development funding (2022-23) to 21 Local Action Group (LAG) areas in Scotland. It accompanies a note ('Producing indicative allocations for Community Led Local Development funding in Scotland (2022-23)', <a href="https://doi.org/10.5281/zenodo.7418862">https://doi.org/10.5281/zenodo.7418862</a>) providing an overview of the analysis and its key outputs.</p>
Codes in R for spatial statistics analysis, ecological response models and spatial distribution models
<p>In the last decade, a plethora of algorithms have been developed for spatial ecology studies. In our case, we use some of these codes for underwater research work in applied ecology analysis of threatened endemic fishes and their natural habitat. For this, we developed codes in Rstudio® script environment to run spatial and statistical analyses for ecological response and spatial distribution models (e.g., Hijmans & Elith, 2017; Den Burg <em>et al.</em>, 2020). The employed R packages are as follows: caret (Kuhn et al., 2020), corrplot (Wei & Simko, 2017), devtools (Wickham, 2015), dismo (Hijmans & Elith, 2017), gbm (Freund & Schapire, 1997; Friedman, 2002), ggplot2 (Wickham et al., 2019), lattice (Sarkar, 2008), lattice (Musa & Mansor, 2021), maptools (Hijmans & Elith, 2017), modelmetrics (Hvitfeldt & Silge, 2021), pander (Wickham, 2015), plyr (Wickham & Wickham, 2015), pROC (Robin et al., 2011), raster (Hijmans & Elith, 2017), RColorBrewer (Neuwirth, 2014), Rcpp (Eddelbeuttel & Balamura, 2018), rgdal (Verzani, 2011), sdm (Naimi & Araujo, 2016), sf (e.g., Zainuddin, 2023), sp (Pebesma, 2020) and usethis (Gladstone, 2022).</p> <p>It is important to follow all the codes in order to obtain results from the ecological response and spatial distribution models. In particular, for the ecological scenario, we selected the Generalized Linear Model (GLM) and for the geographic scenario we selected DOMAIN, also known as Gower's metric (Carpenter <em>et al.</em>, 1993). We selected this regression method and this distance similarity metric because of its adequacy and robustness for studies with endemic or threatened species (<em>e.g.</em>, Naoki <em>et al.</em>, 2006). Next, we explain the statistical parameterization for the codes immersed in the GLM and DOMAIN running:</p> <p>In the first instance, we generated the background points and extracted the values of the variables (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code2_Extract_values_DWp_SC.R?versionId=c1ea0c61-53fe-4f95-ab88-0c1cb28399cb">Code2_Extract_values_DWp_SC.R</a>). Barbet-Massin <em>et al. </em>(2012) recommend the use of 10,000 background points when using regression methods (<em>e.g.</em>, Generalized Linear Model) or distance-based models (<em>e.g.</em>, DOMAIN). However, we considered important some factors such as the extent of the area and the type of study species for the correct selection of the number of points (Pers. Obs.). Then, we extracted the values of predictor variables (<em>e.g.</em>, bioclimatic, topographic, demographic, habitat) in function of presence and background points (<em>e.g.</em>, Hijmans and Elith, 2017).</p> <p>Subsequently, we subdivide both the presence and background point groups into 75% training data and 25% test data, each group, following the method of Soberón & Nakamura (2009) and Hijmans & Elith (2017). For a training control, the 10-fold (cross-validation) method is selected, where the response variable presence is assigned as a factor. In case that some other variable would be important for the study species, it should also be assigned as a factor (Kim, 2009).</p> <p>After that, we ran the code for the GBM method (Gradient Boost Machine; <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code3_GBM_Relative_contribution.R?versionId=1656bbae-66aa-409e-bb91-d8007dee8f95">Code3_GBM_Relative_contribution.R</a> and <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code4_Relative_contribution.R?versionId=0e1d9352-e6b2-43da-984b-d6853a914258">Code4_Relative_contribution.R</a>), where we obtained the relative contribution of the variables used in the model. We parameterized the code with a Gaussian distribution and cross iteration of 5,000 repetitions (<em>e.g.</em>, Friedman, 2002; kim, 2009; Hijmans and Elith, 2017). In addition, we considered selecting a validation interval of 4 random training points (Personal test). The obtained plots were the partial dependence blocks, in function of each predictor variable.</p> <p>Subsequently, the correlation of the variables is run by Pearson's method (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code5_Pearson_Correlation.R?versionId=275f8dd4-b056-44d2-bfe5-f6264bc3298b">Code5_Pearson_Correlation.R</a>) to evaluate multicollinearity between variables (Guisan & Hofer, 2003). It is recommended to consider a bivariate correlation ± 0.70 to discard highly correlated variables (<em>e.g.</em>, Awan <em>et al.</em>, 2021).</p> <p>Once the above codes were run, we uploaded the same subgroups (<em>i.e.</em>, presence and background groups with 75% training and 25% testing) (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code6_Presence&backgrounds.R?versionId=d797b528-782f-4a19-bd61-cfb197f38513">Code6_Presence&backgrounds.R</a>) for the GLM method code (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code7_GLM_model.R?versionId=e4aca276-d601-49ec-a62c-a9223b05a7ed">Code7_GLM_model.R</a>). Here, we first ran the GLM models per variable to obtain the <em>p</em>-significance value of each variable (alpha ≤ 0.05); we selected the value one (<em>i.e.</em>, presence) as the likelihood factor. The generated models are of polynomial degree to obtain linear and quadratic response (<em>e.g.</em>, Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006). From these results, we ran ecological response curve models, where the resulting plots included the probability of occurrence and values for continuous variables or categories for discrete variables. The points of the presence and background training group are also included.</p> <p>On the other hand, a global GLM was also run, from which the generalized model is evaluated by means of a 2 x 2 contingency matrix, including both observed and predicted records. A representation of this is shown in Table 1 (adapted from Allouche et al., 2006). In this process we select an arbitrary boundary of 0.5 to obtain better modeling performance and avoid high percentage of bias in type I (omission) or II (commission) errors (e.g., Carpenter et al., 1993; Fielding and Bell, 1997; Allouche et al., 2006; Kim, 2009; Hijmans and Elith, 2017).</p> <p>Table 1. Example of 2 x 2 contingency matrix for calculating performance metrics for GLM models. A represents true presence records (true positives), B represents false presence records (false positives - error of commission), C represents true background points (true negatives) and D represents false backgrounds (false negatives - errors of omission).</p> <table align="center"> <tbody> <tr> <td> <p> </p> </td> <td> <p>Validation set</p> </td> </tr> <tr> <td> <p>Model</p> </td> <td> <p>True</p> </td> <td> <p>False</p> </td> </tr> <tr> <td> <p>Presence</p> </td> <td> <p>A</p> </td> <td> <p>B</p> </td> </tr> <tr> <td> <p>Background</p> </td> <td> <p>C</p> </td> <td> <p>D</p> </td> </tr> </tbody> </table> <p>We then calculated the Overall and True Skill Statistics (TSS) metrics. The first is used to assess the proportion of correctly predicted cases, while the second metric assesses the prevalence of correctly predicted cases (Olden and Jackson, 2002). This metric also gives equal importance to the prevalence of presence prediction as to the random performance correction (Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006).</p> <p>The last code (<em>i.e.</em>, <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code8_DOMAIN_SuitHab_model.R?versionId=d951a8f2-d3a4-4804-b862-1b2762061876">Code8_DOMAIN_SuitHab_model.R</a>) is for species distribution modelling using the DOMAIN algorithm (Carpenter <em>et al.</em>, 1993). Here, we loaded the variable stack and the presence and background group subdivided into 75% training and 25% test, each. We only included the presence training subset and the predictor variables stack in the calculation of the DOMAIN metric, as well as in the evaluation and validation of the model.</p> <p>Regarding the model evaluation and estimation, we selected the following estimators:</p> <p>1) partial ROC, which evaluates the approach between the curves of positive (<em>i.e.</em>, correctly predicted presence) and negative (i.e., correctly predicted absence) cases. As farther apart these curves are, the model has a better prediction performance for the correct spatial distribution of the species (Manzanilla-Quiñones, 2020).</p> <p>2) ROC/AUC curve for model validation, where an optimal performance threshold is estimated to have an expected confidence of 75% to 99% probability (De Long <em>et al.</em>, 1988).</p>
Codes and data related to the article: Renard et al. Floods and Heavy Precipitation at the Global Scale: 100-year Analysis and 180-year Reconstruction. Journal of Geophysical Research - Atmospheres.
<p>This package contains R codes and data related to the article:</p> <p>B. Renard, D. McInerney, S. Westra, M. Leonard, D. Kavetski, M. Thyer and J.-P. Vidal. Floods and Heavy Precipitation at the Global Scale: 100-year Analysis and 180-year Reconstruction. <em>Journal of Geophysical Research - Atmospheres</em>. DOI: <a href="https://doi.org/10.1029/2022JD037908">10.1029/2022JD037908</a></p> <p><strong>Analyses</strong></p> <p>This folder contains the R scripts used to set up models, analyse results and prepare figures. See README file for details.</p> <p><strong>ShinyApp</strong></p> <p>This folder contains an interactive Shiny App to explore the data and the results from the article.</p> <p>An online version can be found at <a href="https://hydroapps.recover.inrae.fr/HEGS-paper">https://hydroapps.recover.inrae.fr/HEGS-paper</a></p> <p> </p>
Red deer growth data and R code for analysis
<p>Dataset and R code (Rmd-file) for analysis of seasonal growth of body weight in red deer. </p> <p>Supplementary material for the paper "Shifting seasonality of annual growth through ontogeny for red deer at northern latitudes".</p> <p>This study was part of the AgriDeer project (318575), funded by the Research Council of Norway.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.