Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,249
datasets available to search
ShareScore release 0.9.0
Dataset results
1,249 results for “R data”
Data from: PrimerMiner: an R package for development and in silico validation of DNA metabarcoding primers
1. DNA metabarcoding is a powerful tool to assess biodiversity by amplifying and sequencing a standardized gene marker region. Its success is often limited due to variable binding sites that introduce amplification biases. Thus the development of optimized primers for communities or taxa under study in a certain geographic region and/or ecosystems is of critical importance. However, no tool for obtaining and processing of reference sequence data in bulk that can serve as a backbone for primer design is currently available. 2. We developed the R package PrimerMiner, which batch downloads DNA barcode gene sequences from BOLD and NCBI databases for specified target taxonomic groups and then applies sequence clustering into operational taxonomic units (OTUs) to reduce biases introduced by the different number of available sequences per species. Additionally, PrimerMiner offers functionalities to evaluate primers in silico, which are in our opinion more realistic then the strategy employed in another available software for that purpose, ecoPCR. 3. We used PrimerMiner to download cytochrome c oxidase subunit I (COI) sequences for 15 important freshwater invertebrate groups, relevant for ecosystem assessment. By processing COI markers from both databases, we were able to increase the amount of reference data 249-fold on average, compared to using complete mitochondrial genomes alone. Furthermore, we visualized the generated OTU sequence alignments and describe how to evaluate primers in silico using PrimerMiner. 4. With PrimerMiner we provide a useful tool to obtain relevant sequence data for targeted primer development and evaluation. The OTU based reference alignments generated with PrimerMiner can be used for manual primer design, or processed with bioinformatic tools for primer development.
Data from: pcadapt: an R package to perform genome scans for selection based on principal component analysis
The R package pcadapt performs genome scans to detect genes under selection based on population genomic data. It assumes that candidate markers are outliers with respect to how they are related to population structure. Because population structure is ascertained with principal component analysis, the package is fast and works with large-scale data. It can handle missing data and pooled sequencing data. By contrast to population-based approaches, the package handle admixed individuals and does not require grouping individuals into populations. Since its first release, pcadapt has evolved in terms of both statistical approach and software implementation. We present results obtained with robust Mahalanobis distance, which is a new statistic for genome scans available in the 2.0 and later versions of the package. When hierarchical population structure occurs, Mahalanobis distance is more powerful than the communality statistic that was implemented in the first version of the package. Using simulated data, we compare pcadapt to other computer programs for genome scans (BayeScan, hapflk, OutFLANK, sNMF). We find that the proportion of false discoveries is around a nominal false discovery rate set at 10% with the exception of BayeScan that generates 40% of false discoveries. We also find that the power of BayeScan is severely impacted by the presence of admixed individuals whereas pcadapt is not impacted. Last, we find that pcadapt and hapflk are the most powerful in scenarios of population divergence and range expansion. Because pcadapt handles next-generation sequencing data, it is a valuable tool for data analysis in molecular ecology.
Data from: paco: implementing Procrustean Approach to Cophylogeny in R
1. The concordance of evolutionary histories and extant species interactions provides a useful metric for addressing questions of how the structure of ecological communities is influenced by macro-evolutionary processes. 2. We introduce paco (v.0.3.1), an R package to perform Procrustean Approach to Cophylogeny. This method assesses the phylogenetic congruence, or evolutionary dependence, of two groups of interacting species using both ecological interaction networks and their phylogenetic history. 3. We demonstrate the functionality of paco through its application to empirical host-parasite and plant-pollinator communities 4. Although the package is intended to assess phylogenetic congruence between groups of interacting species, the method is also directly applicable to other scenarios that may show phylogenetic congruence including historical biogeography, molecular systematics, and cultural evolution.
Data from: Population genomics of pearl millet (Pennisetum glaucum (L.) R. Br.): comparative analysis of global accessions and Senegalese landraces
Background: Pearl millet is a staple food for people in arid and semi-arid regions of Africa and South Asia due to its high drought tolerance and nutritional qualities. A better understanding of the genomic diversity and population structure of pearl millet germplasm is needed to support germplasm conservation and genetic improvement of this crop. Here we characterized two pearl millet diversity panels, (i) a set of global accessions from Africa, Asia, and the America, and (ii) a collection of landraces from multiple agro-ecological zones in Senegal. Results: We identified 83,875 single nucleotide polymorphisms (SNPs) in 500 pearl millet accessions, comprised of 252 global accessions and 248 Senegalese landraces, using genotyping by sequencing (GBS) of PstI-MspI reduced representation libraries. We used these SNPs to characterize genomic diversity and population structure among the accessions. The Senegalese landraces had the highest levels of genetic diversity (π), while accessions from southern Africa and Asia showed lower diversity levels. Principal component analyses and ancestry estimation indicated clear population structure between the Senegalese landraces and the global accessions, and among countries in the global accessions. In contrast, little population structure was observed across in the Senegalese landraces collections. We ordered SNPs on the pearl millet genetic map and observed much faster linkage disequilibrium (LD) decay in Senegalese landraces compared to global accessions. A comparison of pearl millet GBS linkage map with the foxtail millet (Setaria italica) and sorghum (Sorghum bicolor) genomes indicated extensive regions of synteny, as well as some large-scale rearrangements in the pearl millet lineage. Conclusions: We identified 83,875 SNPs as a genomic resource for pearl millet improvement. The high genetic diversity in Senegal relative to other regions of Africa and Asia supports a West African origin of this crop, followed by wide diffusion. The rapid LD decay and lack of confounding population structure along agro-ecological zones in Senegalese pearl millet will facilitate future association mapping studies. Comparative population genomics will provide insights into panicoid crop evolution and support improvement of these climate-resilient crops.
Data from: TipDatingBeast: an R package to assist the implementation of phylogenetic tip-dating tests using BEAST
Molecular tip-dating of phylogenetic trees is a growing discipline that uses DNA sequences sampled at different points in time to co-estimate the timing of evolutionary events with rates of molecular evolution. In this context, BEAST, a program for Bayesian analysis of molecular sequences, is the most widely used phylogenetic tool. Here, we introduce TipDatingBeast, an R package built to assist the implementation of various phylogenetic tip-dating tests using BEAST. TipDatingBeast currently contains two main functions. The first one allows preparing date-randomization analyses, which assess the temporal signal of a dataset. The second function allows performing leave-one-out analyses, which test for the consistency between independent calibration sequences and allow pinpointing those leading to potential bias. We apply those functions to an empirical dataset and supply practical guidance for results interpretation.
R scripts for data simulation and model assessment in ShuqingNTeng MEE 2017
<p>R scripts for reproducing results in ShuqingNTeng MEE 2017</p>
Data and R code for cluster analysis and machine learning modelling of favourite places for outdoor recreation
Open the record for dataset details and reuse information.
Sequence data processing R script
<p>R script used for processing the 16S Illumina paired end read data using the DADA2 pipeline.</p><p> </p>
Dataset in R format, containing labor force data for a synthetic population
Open the record for dataset details and reuse information.
Supplementary material 3 from: Luo T, Mao M-L, Lan C-T, Song L-X, Zhao X-R, Yu J, Wang X-L, Xiao N, Zhou J-J, Zhou J (2023) Four new hypogean species of the genus Triplophysa (Osteichthyes, Cypriniformes, Nemacheilidae) from Guizhou Province, Southwest China, based on molecular and morphological data. ZooKeys 1185: 43-81. https://doi.org/10.3897/zookeys.1185.105499
Morphological characters and measurement data
Supplementary material 2 from: Luo T, Mao M-L, Lan C-T, Song L-X, Zhao X-R, Yu J, Wang X-L, Xiao N, Zhou J-J, Zhou J (2023) Four new hypogean species of the genus Triplophysa (Osteichthyes, Cypriniformes, Nemacheilidae) from Guizhou Province, Southwest China, based on molecular and morphological data. ZooKeys 1185: 43-81. https://doi.org/10.3897/zookeys.1185.105499
GPS information on the geographical distribution of 39 hypogean species of the genus Triplophysa
Data and R code linked to the paper "The human metatarsal from Sedia del Diavolo"
<p>Data and R code to reproduce the results reported in "The human metatarsal from Sedia del Diavolo"</p>
Supplementary material 1 from: Borisenko A, Young R, Hanner R (2024) A lab-centric, workflow-based data management system for environmental DNA research. Research Ideas and Outcomes 10: e120483. https://doi.org/10.3897/rio.10.e120483
eDNA Laboratory Database Schema Outline
C. Ott, R. Torres et al. cilium annotation data
Open the record for dataset details and reuse information.
Supplementary material 1 from: Inäbnit T, Jochum A, Slapnik R, Neubert E (2021) New genetic data reveals a new species of Zospeum in Bosnia (Gastropoda, Ellobioidea, Carychiinae). ZooKeys 1071: 175-193. https://doi.org/10.3897/zookeys.1071.66417
Figure S1
Figure 3 from: Inäbnit T, Jochum A, Slapnik R, Neubert E (2021) New genetic data reveals a new species of Zospeum in Bosnia (Gastropoda, Ellobioidea, Carychiinae). ZooKeys 1071: 175-193. https://doi.org/10.3897/zookeys.1071.66417
Figure 3 Specimens sequenced in this study. Zospeum troglobalcanicum: NMBE 568052 & 568053 (both from Špilja Jezero); Zospeum simplex sp. nov.: NMBE 568054 (Špilja Dahna), NMBE 568055–568057 (Jama u kamenolomu), NMBE 568059 (Vranjača), NMBE 568060 (Holotype, Jama Dobravljevac), NMBE 568061–568063 (Paratypes, Jama Dobravljevac)
Figure 1 from: Inäbnit T, Jochum A, Slapnik R, Neubert E (2021) New genetic data reveals a new species of Zospeum in Bosnia (Gastropoda, Ellobioidea, Carychiinae). ZooKeys 1071: 175-193. https://doi.org/10.3897/zookeys.1071.66417
Figure 1 Map showing the distribution of the Zospeum pretneri group and the Zospeum alpestre group (except Z. isselianum). Austrian specimens from Kruckenhauser et al. (2019) are labelled as "Z. cf. amoenum".
Figure 2 from: Inäbnit T, Jochum A, Slapnik R, Neubert E (2021) New genetic data reveals a new species of Zospeum in Bosnia (Gastropoda, Ellobioidea, Carychiinae). ZooKeys 1071: 175-193. https://doi.org/10.3897/zookeys.1071.66417
Figure 2 Bayesian tree of the genus Zospeum. Node support values of both the Bayesian Inference (front) and the Maximum Likelihood analysis (back) are given. Branches are coloured to denote the informal species groups within the eastern radiation of Zospeum following Inäbnit et al. (2019). Coloured sample names indicate specimens not included in the tree in Inäbnit et al. (2019): blue: Austrian specimens from Kruckenhauser et al. (2019); dark green: Zospeum troglobalcanicum; light green: Zospeum simplex sp. nov.
R code: Large-scale assessment of bird data quality in a citizen science platform
<p>R code archived to Zenodo for publication</p>
Supplementary material 1 from: Hatami R, Inglis G, Lane SE, Growcott A, Kluza D, Lubarsky C, Jones-Todd C, Seaward K, Robinson AP (2022) Modelling the likelihood of entry of marine non-indigenous species from internationally arriving vessels to maritime ports: a case study using New Zealand data. NeoBiota 72: 183-203. https://doi.org/10.3897/neobiota.72.77266
Supplementary materials
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.