Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
149
datasets available to search
ShareScore release 0.9.0
Dataset results
149 results for “R package”
Data from: PrimerMiner: an R package for development and in silico validation of DNA metabarcoding primers
1. DNA metabarcoding is a powerful tool to assess biodiversity by amplifying and sequencing a standardized gene marker region. Its success is often limited due to variable binding sites that introduce amplification biases. Thus the development of optimized primers for communities or taxa under study in a certain geographic region and/or ecosystems is of critical importance. However, no tool for obtaining and processing of reference sequence data in bulk that can serve as a backbone for primer design is currently available. 2. We developed the R package PrimerMiner, which batch downloads DNA barcode gene sequences from BOLD and NCBI databases for specified target taxonomic groups and then applies sequence clustering into operational taxonomic units (OTUs) to reduce biases introduced by the different number of available sequences per species. Additionally, PrimerMiner offers functionalities to evaluate primers in silico, which are in our opinion more realistic then the strategy employed in another available software for that purpose, ecoPCR. 3. We used PrimerMiner to download cytochrome c oxidase subunit I (COI) sequences for 15 important freshwater invertebrate groups, relevant for ecosystem assessment. By processing COI markers from both databases, we were able to increase the amount of reference data 249-fold on average, compared to using complete mitochondrial genomes alone. Furthermore, we visualized the generated OTU sequence alignments and describe how to evaluate primers in silico using PrimerMiner. 4. With PrimerMiner we provide a useful tool to obtain relevant sequence data for targeted primer development and evaluation. The OTU based reference alignments generated with PrimerMiner can be used for manual primer design, or processed with bioinformatic tools for primer development.
Data from: pcadapt: an R package to perform genome scans for selection based on principal component analysis
The R package pcadapt performs genome scans to detect genes under selection based on population genomic data. It assumes that candidate markers are outliers with respect to how they are related to population structure. Because population structure is ascertained with principal component analysis, the package is fast and works with large-scale data. It can handle missing data and pooled sequencing data. By contrast to population-based approaches, the package handle admixed individuals and does not require grouping individuals into populations. Since its first release, pcadapt has evolved in terms of both statistical approach and software implementation. We present results obtained with robust Mahalanobis distance, which is a new statistic for genome scans available in the 2.0 and later versions of the package. When hierarchical population structure occurs, Mahalanobis distance is more powerful than the communality statistic that was implemented in the first version of the package. Using simulated data, we compare pcadapt to other computer programs for genome scans (BayeScan, hapflk, OutFLANK, sNMF). We find that the proportion of false discoveries is around a nominal false discovery rate set at 10% with the exception of BayeScan that generates 40% of false discoveries. We also find that the power of BayeScan is severely impacted by the presence of admixed individuals whereas pcadapt is not impacted. Last, we find that pcadapt and hapflk are the most powerful in scenarios of population divergence and range expansion. Because pcadapt handles next-generation sequencing data, it is a valuable tool for data analysis in molecular ecology.
Data from: TipDatingBeast: an R package to assist the implementation of phylogenetic tip-dating tests using BEAST
Molecular tip-dating of phylogenetic trees is a growing discipline that uses DNA sequences sampled at different points in time to co-estimate the timing of evolutionary events with rates of molecular evolution. In this context, BEAST, a program for Bayesian analysis of molecular sequences, is the most widely used phylogenetic tool. Here, we introduce TipDatingBeast, an R package built to assist the implementation of various phylogenetic tip-dating tests using BEAST. TipDatingBeast currently contains two main functions. The first one allows preparing date-randomization analyses, which assess the temporal signal of a dataset. The second function allows performing leave-one-out analyses, which test for the consistency between independent calibration sequences and allow pinpointing those leading to potential bias. We apply those functions to an empirical dataset and supply practical guidance for results interpretation.
bipl5: An R Package for Reactive Calibrated Axes PCA Biplots
Open the record for dataset details and reuse information.
Supplementary material 1 from: Niehues A, de Visser C, Hagenbeek FA, Karu N, Kindt ASD, Kulkarni P, Pool R, Boomsma DI, van Dongen J, van Gool AJ, `t Hoen PAC (2022) A Multi-omics Data Analysis Workflow Packaged as a FAIR Digital Object. Research Ideas and Outcomes 8: e94042. https://doi.org/10.3897/rio.8.e94042
Members of the ACTION Consortium
Data from the rfPred R package
<p>Large dataset to be used with the rfPred R package</p>
climetrics: an R package to quantify multiple dimensions of climate change
Open the record for dataset details and reuse information.
Figure 8b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 8b "Point Pattern Edition" features. - Information that is displayed (marks of the point pattern, if available, as defined by the user) when an event is clicked
Figure 5a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 5a Example of use of the SimplifyLinearNetwork function. - A road network introduced as input in which there is an excess of road segments and vertex
Figure 1 from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 1 Workflow that describes all the steps that could be carried out in order to perform a spatial analysis on a point pattern that lies on a linear network. Some of these steps which lead to the final statistical analysis may be skipped but, at least, all of them should be considered. The blocks pointing the steps of the process include some of the R packages that would allow to successfully achieve each of them.
Figure 3b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 3b "Network Edition" example of use (I). - Network resulting from clicking on "Rebuild linear network" in the situation of a
Figure 2a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 2a "Network Edition" features. - Overview of the "Network Edition" section of the SpNetPrep application
Figure 6b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 6b "Network Direction" features. - Manual addition of traffic flow to the network by using the options "Add flow" and "Add long flow"
Figure 3a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 3a "Network Edition" example of use (I). - Use of the "Join vertex" (in green), "Remove edge" (in red) and "Add point" options (in green) in the SpNetPrep application
Figure 8a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 8a "Point Pattern Edition" features. - An example of a point pattern that lies on a road network as it can be visualized in SpNetPrep
Figure 4b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 4b "Network Edition" example of use (II). - Network resulting from clicking on "Rebuild linear network" in the situation of a
Figure 7 from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 7 Example of a linear road network following usual notation for the edges (\documentclass[12pt]{standalone} \usepackage{varwidth} \usepackage[utf8x]{inputenc} \usepackage[T1]{fontenc} \usepackage{lmodern} \usepackage{amsmath, amssymb, graphics, setspace} \newcommand{\mathsym}[1]{{}} \newcommand{\unicode}[1]{{}} \newcounter{mathematicapage} \begin{document} \begin{varwidth}{50in} \begin{equation*} e_{i} \end{equation*} \end{varwidth} \end{document} ) and vertex (\documentclass[12pt]{standalone} \usepackage{varwidth} \usepackage[utf8x]{inputenc} \usepackage[T1]{fontenc} \usepackage{lmodern} \usepackage{amsmath, amssymb, graphics, setspace} \newcommand{\mathsym}[1]{{}} \newcommand{\unicode}[1]{{}} \newcounter{mathematicapage} \begin{document} \begin{varwidth}{50in} \begin{equation*} v_{i} \end{equation*} \end{varwidth} \end{document} ). Arrows represent the direction of traffic flow.
Figure 6a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 6a "Network Direction" features. - A zone of a road network introduced as an input in the "Network Direction" section of the SpNetPrep application
Figure 4a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 4a "Network Edition" example of use (II). - Another use of the "Join vertex" (in green) option of the "Network Edition" section
Figure 5b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521
Figure 5b Example of use of the SimplifyLinearNetwork function. - Simplified version of the network in a after the application of the SimplifyLinearNetwork function with parameters Angle = 25 and Length = 65
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.