Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
149
datasets available to search
ShareScore release 0.9.0
Dataset results
149 results for “R package”
specleanr: An R package for automated flagging of environmental outliers in ecological data for modeling workflows
Open the record for dataset details and reuse information.
Code and example images from: recolorize: An R package for flexible color segmentation of biological images
Open the record for dataset details and reuse information.
spectre: An R package to estimate spatially-explicit community composition using sparse data
Open the record for dataset details and reuse information.
Data from: hespdiv: an R package for spatially constrained, hierarchical and contiguous regionalization in palaeobiogeography
Open the record for dataset details and reuse information.
Data from: The article Euclimatch: An R package for climate matching with Euclidean distance metrics
Open the record for dataset details and reuse information.
Data from: Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories
Open the record for dataset details and reuse information.
EcoPhyloMapper: an R package for integrating geographic ranges, phylogeny, and morphology
Open the record for dataset details and reuse information.
Data from: aniMotum, an R package for animal movement data: rapid quality control, behavioural estimation and simulation
Open the record for dataset details and reuse information.
GapAnalysis: An R package to calculate conservation indicators using spatial information
Open the record for dataset details and reuse information.
MERRA-2 subset for evaluation of renewables with merra2ools R-package: 1980-2020 hourly, 0.5° lat x 0.625° lon global grid
Open the record for dataset details and reuse information.
The R package enerscape: A general energy landscape framework for terrestrial movement ecology
Open the record for dataset details and reuse information.
Data from: imageseg: An R package for deep learning-based image segmentation
Open the record for dataset details and reuse information.
Data from: nlstimedist: an R package for the biologically meaningful quantification of unimodal phenology distributions
Phenological investigation can provide valuable insights into the ecological effects of climate change. Appropriate modelling of the time distribution of phenological events is key to determining the nature of any changes, as well as the driving mechanisms behind those changes. Here we present the nlstimedist R package, a distribution function and modelling framework that describes the temporal dynamics of unimodal phenological events. The distribution function is derived from first principles and generates three biologically interpretable parameters. Using seed germination at different temperatures as an example, we show how the influence of environmental factors on a phenological process can be determined from the quantitative model parameters. The value of this model is its ability to represent various unimodal temporal processes statistically. The three intuitively meaningful parameters of the model can make useful comparisons between different time periods, geographical locations or species' populations, in turn allowing exploration of possible causes.
Data from: An R package and online resource for macroevolutionary studies using the ray-finned fish tree of life
1. Comprehensive, time-scaled phylogenies provide a critical resource for many questions in ecology, evolution, and biodiversity. Methodological advances have increased the breadth of taxonomic coverage in phylogenetic data; however, accessing and reusing these data remain challenging. 2. We introduce the Fish Tree of Life website and associated R package fishtree to provide convenient access to sequences, phylogenies, fossil calibrations, and diversification rate estimates for the most diverse group of vertebrate organisms, the ray-finned fishes. The Fish Tree of Life website presents subsets and visual summaries of phylogenetic and comparative data, and is complemented by the R package, which provides flexible programmatic access to the same underlying data source for advanced users wishing to extend or reanalyze the data. 3. We demonstrate functionality with an overview of the website, and show three examples of advanced usage through the R package. First, we test for the presence of long branch attraction artifacts across the fish tree of life. The second example examines the effects of habitat on diversification rate in the pufferfishes. The final example demonstrates how a community phylogenetic analysis could be conducted with the package. 4. This resource makes a large comparative vertebrate dataset easily accessible via the website, while the R package enables the rapid reuse and reproducibility of research results via its ability to easily integrate with other R packages and software for molecular biology and comparative methods.
Wrap your model in an R package!
<p>The groundwater drawdown model WTAQ-2, provided by the United States Geological Survey for free, has been "wrapped" into an R package, which contains functions for writing input files, executing the model engine and reading output files. By calling the functions from the R package a sensitivity analysis, calibration or validation requiring multiple model runs can be performed in an automated way. Automation by means of programming improves and simplifies the modelling process by ensuring that the WTAQ-2 wrapper generates consistent model input files, runs the model engine and reads the output files without requiring the user to cope with the technical details of the communication with the model engine. In addition the WTAQ-2 wrapper automatically adapts cross-dependent input parameters correctly in case one is changed by the user. This assures the formal correctness of the input file and minimises the effort for the user, who normally has to consider all cross-dependencies for each input file modification manually by consulting the model documentation. Consequently the focus can be shifted on retrieving and preparing the data needed by the model. Modelling is described in the form of version controlled R scripts so that its methodology becomes transparent and modifications (e.g. error fixing) trackable. The code can be run repeatedly and will always produce the same results given the same inputs. The implementation in the form of program code further yields the advantage of inherently documenting the methodology. This leads to reproducible results which should be the basis for smart decision making.</p>
Replication package for How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists
<p>This dataset was used in the paper: "How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists", Journal of Empirical Software Engineering, to appear.</p>
Scientific Mapping Example with R package `bibliometrix`
<p>Example of the use of the <strong>R</strong> package <i>bibliometrix </i>for bibliometric analysis</p>
NextClone and CloneDetective: An Integrated Nextflow Pipeline and R Package for Clonal Barcode Extraction and Quantification
<p>Raw FASTQ files for the DNA-seq data required to replicate the analyses presented at: https://phipsonlab.github.io/NextClone-analysis/.</p>
Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie
<p>Diet analysis integrates a wide variety of visual, chemical and biological identification of prey. Samples are often treated as compositional data, where each prey is analyzed as a continuous percentage of the total. However, analyzing compositional data results in analytical challenges, e.g., highly parameterized models or prior transformation of data. Here, we present a novel approximation involving a Tweedie generalized linear model (GLM). We first review how this approximation emerges from considering predator foraging as a thinned and marked point process (with marks representing prey species and individual prey size). This derivation can motivate future theoretical and applied developments. We then provide a practical tutorial for the Tweedie GLM using new package <i>mvtweedie</i> that extends capabilities of widely used packages in R (<i>mgcv</i> and <i>ggplot2</i>) by transforming output to calculate prey compositions. We demonstrate this approach and software using two examples. Tufted puffins (<i>Fratercula cirrhata</i>) provisioning their chicks on a colony in the northern Gulf of Alaska show decadal prey switching among sand lance and prowfish (1980-2000) and then Pacific herring and capelin (2000-2020), while wolves (<i>Canis lupus ligoni</i>) in Southeast Alaska forage on mountain goats and marmots in northern uplands and marine mammals in seaward island coastlines. </p>
comspat: an R package to analyze within-community spatial organization using species combinations
<p>The diversity of species combinations observable in sampling units reflects a species' uneven distribution and preference for specific abiotic and biotic conditions – a phenomenon most commonly expressed in terms of ecological assembly rules of plant communities and other sessile organisms (e.g., subtidal algae, invertebrates and coral reefs). We present comspat, a new R package that uses grid or transect data sets to measure the number of realized (observed) species combinations (NRC) and the Shannon diversity of realized species combinations (compositional diversity; CD) as a function of spatial scale. NRC and CD represent two measures from a model family developed by Pál Juhász-Nagy based on Information Theory. Classical Shannon diversity measures biodiversity based on the number and relative abundance of species, whereas the specific version of Shannon diversity presented here characterizes biodiversity and provides information on species coexistence relationships; both measures operate at fine-scale within the sampling unit or within the community. comspat offers two commonly applied null models, complete spatial randomness and random shift, to disentangle the textural, intraspecific, and interspecific effects on the observed spatial patterns. Combined, these models assist users in detecting and interpreting spatial associations and inferring assembly mechanisms. Our open-sourced package provides a vignette that describes the method and reproduces the figures from this paper to help users contextualize and apply functions to their data.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.