Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
149
datasets available to search
ShareScore release 0.9.0
Dataset results
149 results for “R package”
Data for testing the SIAMCAT R package
<p>Datasets needed for the vignettes of the <a href="https://siamcat.embl.de/">SIAMCAT</a> R package</p>
Dataset used for the validation of stemv: An R package for calculating tree stem volume in Japan
<p>This repository contains the dataset for validating an R package "stemv", which provides the functions for calculating tree stem volume in Japan.<br>The source of the package can be found on <a href="https://github.com/dulvrq/stemv" target="_blank" rel="noopener">GitHub </a>(https://github.com/dulvrq/stemv).</p>
Example Data for the R package lpjmlkit (LPJmLData Vignette)
<div>This repository contains example data for the R package lpjmlkit to provide usable large datasets for LPJmLData Vignette.</div> <div> </div> <div><em>Breier, J., Ostberg, S., Wirth, S. B., Minoli, S., Stenzel, F., & Müller, C. (2024). lpjmlkit: Toolkit for Basic LPJmL Handling (Version 1.7.0) [Computer software]. https://doi.org/10.5281/zenodo.7773134</em></div> <div> </div> <div>Usage information: We split zip files to bypass Zenodo's upload restrictions. Please download all and open one of the zip files to unpack all automatically.</div>
RZooRoH: an R package to characterize individual genomic autozygosity and identify homozygous-by-descent segments
<p>these are two files used to test the RZooRoH package in the manuscript:<br> RZooRoH: an R package to characterize genomic autozygosity and identify homozygous-by-descent segments<br> A.R. Bertrand, N.K. Kadri, L. Flori, M. Gautier and T. Druet</p> <p>Each row represents a SNP. The files contain first 5 information columns followed by the genotyes (1 row per individual):<br> 1 Chromosome number<br> 2 Marker name<br> 3 Position in bp<br> 4 Allele 1<br> 5 Allele 2</p> <p>Genotypes are coded as 0, 1 and 2 for AA, AB and BB and 9 for missing.</p> <p>The file geno_sim15_5col.txt contains the simulated genotypes: 25000 markers and 500 individuals.<br> The file filtered_soay_5col.txt contains the sheep genotype data: 47365 markers and 110 individuals.</p> <p>The origin of the files is described in the paper.<br> The simulated data set is from Druet and Gautier (2017) whereas the sheep data comes from Kijas et al. (2012).</p>
Review data for: SnowQM 1.0: A fast R Package for bias-correcting spatial fields of snow water equivalent using quantile mapping
<p>Climatology of snow water equivalent of Switzerland between winters 1962 and 2021. Obtained using quartile mapping between a model using data assimilation since 1998 and a model without data assimilation. This version of the dataset corresponds to the publication revision time. The publication has been submitted to GMD Copernicus journal as: <em>SnowQM 1.0: A fast R Package for bias-correcting spatial fields of snow water equivalent using quantile mapping</em></p>
Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie
Open the record for dataset details and reuse information.
Data from: An R package and online resource for macroevolutionary studies using the ray-finned fish tree of life
Open the record for dataset details and reuse information.
comspat: an R package to analyze within-community spatial organization using species combinations
Open the record for dataset details and reuse information.
Data from: Chronospaces: an R package for the statistical exploration of divergence times promotes the assessment of methodological sensitivity
Open the record for dataset details and reuse information.
Data from: nlstimedist: an R package for the biologically meaningful quantification of unimodal phenology distributions
Open the record for dataset details and reuse information.
Replicate analysis from: tinyVAST: R package with an expressive interface to specify lagged and simultaneous effects in multivariate spatio-temporal models
Open the record for dataset details and reuse information.
Data for: Multi-trait/environment sparse genomic prediction using the SFSI R-package
Open the record for dataset details and reuse information.
Data from: Non-random association of MHC-I alleles in favor of high diversity haplotypes in wild songbirds revealed by computer-assisted MHC haplotype inference using the R package MHCtools
<p><strong>Data set from:</strong></p> <p>Roved J., Hansson B., Stervander, M., Hasselquist D., & Westerdahl H. (2020). Non-random association of MHC-I alleles in favor of high diversity haplotypes in wild songbirds revealed by computer-assisted MHC haplotype inference using the R package MHCtools.</p>
Data from: PolyPatEx: an R package for paternity exclusion in autopolyploids
Microsatellite markers have demonstrated their value for performing paternity exclusion and hence exploring mating patterns in plants and animals. Methodology is well established for diploid species and several software packages exist for elucidating paternity in diploids, however these issues are not so readily addressed in polyploids due to the increased complexity of the exclusion problem and a lack of available software. We introduce PolyPatEx, an R package for paternity exclusion analysis using microsatellite data in autopolyploid, monoecious or dioecious/bisexual species with a ploidy of 4n, 6n or 8n. Given marker data for a set of offspring, their mothers, and a set of candidate fathers, PolyPatEx uses allele matching to exclude candidates whose marker alleles are incompatible with the alleles in each offspring-mother pair. PolyPatEx can analyse marker data sets in which allele copy numbers are known (genotype data) or unknown (allelic phenotype data) – for data sets in which allele copy numbers are unknown, comparisons are made taking into account all possible genotypes that could arise from the compared allele sets. PolyPatEx is a software tool that provides population geneticists with the ability to investigate the mating patterns of autopolyploids using paternity exclusion analysis on data from codominant markers having multiple alleles per locus.
Data from: Multi-DICE: R package for comparative population genomic inference under hierarchical co-demographic models of independent single-population size changes
Population genetic data from multiple taxa can address comparative phylogeographic questions about community-scale response to environmental shifts, and a useful strategy to this end is to employ hierarchical co-demographic models that directly test multi-taxa hypotheses within a single, unified analysis while benefiting in statistical power from aggregating datasets. This approach has been applied to classical phylogeographic datasets such as mitochondrial barcodes as well as reduced-genome polymorphism datasets that can yield 10,000s of SNPs, produced by emergent technologies such as RAD-seq and GBS. A strategy for the latter had been accomplished by adapting the site frequency spectrum to a novel summarization of population genomic data across multiple taxa called the aggregate site frequency spectrum (aSFS), which potentially can be deployed under various inferential frameworks including approximate Bayesian computation, random forest, and composite likelihood optimization. Here, we introduce the R package Multi-DICE, a wrapper program that exploits existing simulation software for straight-forward and flexible execution of hierarchical model-based inference using the aSFS, which is derived from genomic-scale data, as well as mitochondrial data. We validate several novel software features such as applying alternative inferential frameworks, enforcing a minimal threshold of time surrounding event pulses, and specifying flexible hyperprior distributions. In sum, Multi-DICE provides comparative analysis within the familiar R environment while allowing a high degree of user customization, and will thus serve as a valuable tool for comparative phylogeography and population genomics.
Data from: An R package for analyzing survival using continuous-time open capture-recapture models
Capture–recapture software packages have proven to be very powerful tools for analysing factors affecting survival in wild populations. However, all such packages are limited to discrete-time protocols. Appropriate survival analysis tools are still lacking for data acquired from continuous-time protocols. We have developed a statistical method and propose an r package for analysing such data based on an extension of classical survival analysis models incorporating an inhomogeneous Poisson process for modelling capture histories. First, data were simulated from a continuous-time protocol. These data were used to (i) compare survival estimation biases of discrete- and continuous-time approaches and (ii) investigate the performance and accuracy of our r package for four types of covariates: factors varying between individuals (like sex), in time (like climatic factors), both in time and between individuals (like physical condition) and age (as a categorical factor). Secondly, the r package has been applied to a real data set for survival analysis of cats in the Kerguelen archipelago (regrouping 682 cats over 20 years) as an illustrative example. Results of the simulated data analysis show that the method performs better than its discrete-time counterpart for analysing data acquired from continuous-time protocols. It provides unbiased parameter estimates for all parameters except those that vary both in time and between individuals – which is not surprising, since in our case, these factors were not updated in continuous time (i.e. only upon capture). When applied to the Kerguelen cat data set, the results suggest that survival is lower in juveniles than in adults and subadults, varies between study sites and increases with physical condition, and this latter effect being more important in females than in males. Sex, season, temporal linear trend in survival and the NDVI vegetation index were also tested but were not found to be significant. However, confidence intervals were too large (due to a low recapture rate) for excluding such effects. Further analyses are still needed for rigorous covariate testing in this context. In conclusion, continuous-time approaches – such as that presented in this paper – should be preferred when data acquired from continuous-time protocols is analysed.
Data from: SIDER: an R package for predicting trophic discrimination factors of consumers based on their ecology and phylogenetic relatedness
Stable isotope mixing models (SIMMs) are an important tool used to study species' trophic ecology. These models are dependent on, and sensitive to, the choice of trophic discrimination factors (TDF) representing the offset in stable isotope delta values between a consumer and their food source when they are at equilibrium. Ideally, controlled feeding trials should be conducted to determine the appropriate TDF for each consumer, tissue type, food source, and isotope combination used in a study. In reality however, this is often not feasible nor practical. In the absence of species-specific information, many researchers either default to an average TDF value for the major taxonomic group of their consumer, or they choose the nearest phylogenetic neighbour for which a TDF is available. Here, we present the SIDER package for R, which uses a phylogenetic regression model based on a compiled dataset to impute (estimate) a TDF of a consumer. We apply information on the tissue type and feeding ecology of the consumer, all of which are known to affect TDFs, using Bayesian inference. Presently, our approach can estimate TDFs for two commonly used isotopes (nitrogen and carbon), for species of mammals and birds with or without previous TDF information. The estimated posterior probability provides both a mean and variance, reflecting the uncertainty of the estimate, and can be subsequently used in the current suite of SIMM software. SIDER allows users to place a greater degree of confidence on their choice of TDF and its associated uncertainty, thereby leading to more robust predictions about trophic relationships in cases where study-specific data from feeding trials is unavailable. The underlying database can be updated readily to incorporate more stable isotope tracers, replicates and taxonomic groups to further increase the confidence in dietary estimates from stable isotope mixing models, as this information becomes available.
NextClone and CloneDetective: An Integrated Nextflow Pipeline and R Package for Clonal Barcode Extraction and Quantification
<p>BAM file (chunk 26-50 of 50) for the scRNAseq data required to replicate the analyses presented at: https://phipsonlab.github.io/NextClone-analysis/.</p><p>Chunk 1-25 can be downloaded from https://zenodo.org/records/10129134.</p><p>The original BAM file was too big to fit in an entry. Thus it was split into 50 using picard:</p><blockquote><p>SplitSamByNumberOfReads -I possorted_genome_bam.bam -O split_sam/ -N_FILES 50 <i>--CREATE_MD5_FILE</i></p></blockquote><p>Before re-running all the analyses in https://phipsonlab.github.io/NextClone-analysis/, make sure you merge all 50 chunks first using picard (change xxx to point to the directory storing all the 50 chunks you have downloaded):</p><blockquote><p>outdir="xxx"</p><p>args=""</p><p><i># Loop through each BAM file</i></p><p>for file in ${outdir}/*.bam; do</p><p> args+="-I $file "</p><p>done</p><p># Do the actual merging</p><p>MergeSamFiles $args -O $outdir/merged_v2/merged_bam_v2.bam --USE_THREADING --CREATE_MD5_FILE</p></blockquote><p>Picard can be downloaded from: https://github.com/broadinstitute/picard</p>
NextClone and CloneDetective: An Integrated Nextflow Pipeline and R Package for Clonal Barcode Extraction and Quantification
<p>BAM file (chunk 1-25 of 50) for the scRNAseq data required to replicate the analyses presented at: https://phipsonlab.github.io/NextClone-analysis/.</p><p>Chunk 26-50 can be downloaded from https://zenodo.org/uploads/10129625</p><p>The original BAM file was too big to fit in an entry. Thus it was split into 50 using picard:</p><blockquote><p>SplitSamByNumberOfReads -I possorted_genome_bam.bam -O split_sam/ -N_FILES 50 <i>--CREATE_MD5_FILE</i></p></blockquote><p>Before re-running all the analyses in https://phipsonlab.github.io/NextClone-analysis/, make sure you merge all 50 chunks first using picard (change xxx to point to the directory storing all the 50 chunks you have downloaded):</p><blockquote><p>outdir="xxx"</p><p>args=""</p><p><i># Loop through each BAM file</i></p><p>for file in ${outdir}/*.bam; do</p><p> args+="-I $file "</p><p>done</p><p># Do the actual merging</p><p>MergeSamFiles $args -O $outdir/merged_v2/merged_bam_v2.bam --USE_THREADING --CREATE_MD5_FILE</p></blockquote><p>Picard can be downloaded from: https://github.com/broadinstitute/picard</p>
Lemna minor annotation R package (org.Lminor.eg.db) and corresponding files behind the custom built
<p>This public repository containing the following files:</p> <ol> <li>Custom built annotation R package for the <em>Lemna minor </em>reference genome [<a href="https://zenodo.org/api/files/33f1633e-9232-4a1c-af07-5bcf19db9304/org.Lminor.eg.db.7z?versionId=c0a06176-0e46-4732-87ae-c6d9f9c68d0c">org.Lminor.eg.db.7z</a>]. The package was built via AnnotationForge using sequence homology of protein coding genes for functional characterisation (Description, PFAMs, GO terms). A combined approach using <a href="https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Web&PAGE_TYPE=BlastDocs&DOC_TYPE=Download">blastx </a>and <a href="http://eggnog5.embl.de/#/app/home">EMBL's eggNOG mapper</a> was used for this task. This package is compatible with <a href="https://bioconductor.org/packages/release/bioc/html/clusterProfiler.html">clusterProfiler</a> for downstream functional enrichment analysis (ORA / GSEA) of <em>L. minor</em> transcriptomic / proteomic data.<br> <br> For <strong>how to install and use this package</strong> in your R session,<strong> check the R code example below</strong>.<br> </li> <li>Reference genome, genome annotation (gtf), gene coding sequences (cds) and cds translated peptide sequences (cds.pep) of the duckweed <em>Lemna minor </em>[<a href="https://zenodo.org/api/files/33f1633e-9232-4a1c-af07-5bcf19db9304/Lminor_refGenome_GTF_CDS.7z?versionId=c25b53a8-1fc3-4bfb-b559-718ab8b230a9">Lminor_refGenome_GTF_CDS.7z</a>].<br> The reference genome assembly fasta was downloaded from <a href="http://www.lemna.org">www.lemna.org</a>. Matching GTF annotation file was generated via '<em>gffread</em>', from the GFF<em> </em>annotation file available <a href="https://genomevolution.org/coge/LoadGenome.pl?wid=47218">here</a>.<br> </li> <li>With the cds translated peptide file, a blastp search was performed against a custom plant protein sequence database [<a href="https://zenodo.org/api/files/33f1633e-9232-4a1c-af07-5bcf19db9304/Lminor_ref.org.Db4blastp.7z?versionId=d1e69820-edfc-494d-9d44-f6624b190841">Lminor_ref.org.Db4blastp.7z</a>]. The custom database was built from the proteomes of well annotated reference plant species. (For details refer to the readme file within the compressed folder)</li> </ol> <p>For more details please refer to our publication in <a href="https://doi.org/10.1021/acs.est.2c01777">Environmental Science & Technology</a>:<br> Loll, Alexandra, Hannes Reinwald, Steve U. Ayobahan, Bernd Göckener, Gabriela Salinas, Christoph Schäfers, Karsten Schlich, Gerd Hamscher, and Sebastian Eilebrecht. <em><strong>“Short-Term Test for Toxicogenomic Analysis of Ecotoxic Modes of Action in Lemna Minor.”</strong></em> Environmental Science & Technology 56, no. 16 (August 16, 2022): 11504–15.<br> DOI: <a href="https://doi.org/10.1021/acs.est.2c01777">https://doi.org/10.1021/acs.est.2c01777</a></p> <pre><code class="language-bash"># 1. Download and unzip (7zip format) the org.Lminor.eg.db package. # 2. Install the package via: orgDb = "path/to/org.Lminor.eg.db/" install.packages(orgDb, type="source", repos=NULL) # 3. Restart R session then load package require(org.Lminor.eg.db) require(AnnotationDbi) # to check for columns and keytypes: columns(org.Lminor.eg.db) keytypes(org.Lminor.eg.db) # query the org.Lminor.eg.db for particular Lemna gene IDs (GID) gid = keys(org.Lminor.eg.db, keytype="GID") col = columns(org.Lminor.eg.db)[c(5,17,9,15,1,8,14)] df = select(org.Lminor.eg.db, keys=gid[1000:1100], columns=col, keytype="GID") View(df) ### Running overrepresenation analysis in clusterProfiler using the Lminor annotation package ### # ORA for multiple gene sets via compareCluster() require(clusterProfiler) genLs = list(setA = gid[1:40], setB = gid[100:140], setC = gid[1000:1040]) res = compareCluster(genLs, fun = "enrichGO", OrgDb = "org.Lminor.eg.db", keyType = "GID", ont = "BP", universe = gid) # Compute semantic similiarities among GO terms: d = GOSemSim::godata('org.Lminor.eg.db', ont="BP", computeIC=FALSE, keytype = "GID") res = enrichplot::pairwise_termsim(res, method = "Wang", semData = d) # Rmv GO terms with redudant biological information resS = simplify(res, .8) # resort results after pvalues resS@compareClusterResult = resS@compareClusterResult[order(resS@compareClusterResult$pvalue),] View(res@compareClusterResult) # Network plot emapplot(resS, showCategory = 30)</code></pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.