Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data from: Comparison of target-capture and restriction-site associated DNA sequencing for phylogenomics: a test in cardinalid tanagers (Aves, genus: Piranga)
Restriction-site associated DNA sequencing (RAD-seq) and target capture of specific genomic regions, such as ultraconserved elements (UCEs), are emerging as two of the most popular methods for phylogenomics using reduced-representation genomic datasets. These two methods were designed to target different evolutionary timescales: RAD-seq was designed for population-genomic level questions and UCEs for deeper phylogenetics. The utility of both datasets to infer phylogenies across a variety of taxonomic levels has not been adequately compared within the same taxonomic system. Additionally, the effects of uninformative gene trees on species tree analyses (for target capture data) have not been explored. Here, we utilize RAD-seq and UCE data to infer a phylogeny of the bird genus Piranga. The group has a range of divergence dates (0.5 my – 6 my), contains eleven recognized species, and lacks a resolved phylogeny. We compared two species tree methods for the RAD-seq data and six species tree methods for the UCE data. Additionally, in the UCE data, we analyzed a complete matrix as well as datasets with only highly informative loci. A complete matrix of 189 UCE loci with ten or more parsimony informative (PI) sites, and an ~80% complete matrix of 1128 PI SNPs (from RAD-seq) yield the same fully resolved phylogeny of Piranga. We inferred non-monophyletic relationships of P. lutea individuals, with all other a priori species identified as monophyletic. Finally, we found that species tree analyses that included predominantly uninformative gene trees provided strong support for different topologies, with consistent phylogenetic results when limiting species tree analyses to highly informative loci or only using less informative loci with concatenation or methods meant for SNPs alone.
Data from: Variation in chick-a-dee call sequences, not in the fine structure of chick-a-dee calls, influences mobbing behaviour in mixed-species flocks
When animals vocalize under the threat of predation, variation in the structure of calls can play a vital role in survival. The chick-a-dee calls of chickadees and titmice provide a model system for studying communication in such contexts. In previous studies, birds' responses to chick-a-dee calls covaried with call structure, but also with unmeasured and correlated parameters of the calling sequence, including duty cycle (the proportion of the calling sequence when a signal was present). In this study, we exposed flocks of Black-capped Chickadees (Poecile atricapillus) and heterospecific birds to playback of chick-a-dee calls and taxidermic models of predators. We quantified birds' responses to variation in number of D-notes and duty cycle of the signalling sequence. Chickadees and heterospecific birds responded more intensely to high-duty-cycle treatments, and equally to treatments where duty cycle was held constant and the number of D-notes varied. Although our study does not disentangle the effects of call rate and duty cycle, it is the first to investigate independently the behavioural responses of birds to variation in structural and sequence-level parameters of the chick-a-dee call during a predator confrontation. Critically, our results confirm that the pattern previously observed in a feeding context holds true in a mobbing context: variation in calling sequences, not in call structure, is the salient acoustic feature of chick-a-dee calls. These results call into question the idea that chick-a-dee call structure carries allometric information about predator size, suggesting instead that sequence-level parameters play a central role in communication in a mobbing context.
Data from: Using Illumina Next Generation Sequencing Technologies to sequence multigene families in de novo species
The advent of Next Generation Sequencing Technology (NGST) has revolutionized molecular biology research, allowing for rapid gene/genome sequencing from a multitude of diverse species. As high throughput sequencing becomes more accessible, more efficient workflows must be developed to deal with the amounts of data produced and better assemble the genomes of de novo lineages. We combine traditional laboratory methods with Illumina NGST to amplify and sequence the largest mammalian multigene family, the Olfactory Receptor gene family, for species with and without a reference genome. We develop novel assembly methods to annotate and filter these data, which can be utilized for any gene family or any species. We find no significant difference between the ratio of genes within their respective gene families of our data compared with available genomic data. Using simulated data we explore the limitations of short-read sequence data and our assembly in recovering this gene family. We highlight the benefits and shortcomings of these methods. Compared with data generated from traditional polymerase chain reaction, cloning and Sanger sequencing methodologies, sequence data generated using our pipeline increases yield and sequencing efficiency without reducing the number of unique genes amplified. A cloning step is not required, therefore shortening data generation time. The novel downstream methodologies and workflows described provide a tool to be utilized by many fields of biology, to access and analyze the vast quantities of data generated. By combining laboratory and in silico methods, we provide a means of extracting genomic information for multigene families without complete genome sequencing.
Data from: Distribution of MICB diversity in the Zhejiang Han population: PCR sequence-based typing for exons 2–6 and identification of five novel MICB alleles
The polymorphism of major histocompatibility complex class I chain-related gene B (MICB) and variations in MICB alleles in a variety of populations have been characterized using several genotyping approaches. In the present study, a novel polymerase chain reaction sequence-based typing (PCR-SBT) method was established for the genotyping of MICB exons 2–6, and the allelic frequency of MICB in the Zhejiang Han population was investigated. Among 400 unrelated healthy Han individuals from Zhejiang Province, China, a total of 20 MICB alleles were identified, of which MICB*005:02:01, MICB*002:01:01, and MICB*004:01:01 were the most predominant alleles, with frequencies of 0.57375, 0.1225, and 0.08375, respectively. Nine MICB alleles were detected on only one occasion, giving a frequency of 0.00125. Of the 118 distinct MICB ∼ HLA-B haplotypes identified, 42 showed significant linkage disequilibrium (P < 0.05). Haplotypes MICB*005:02:01 ∼ B*46:01, MICB*005:02:01 ∼ B*40:01, and MICB*008 ∼ B*58:01 were the most common haplotypes, with frequencies of 0.0978, 0.0761, and 0.0616, respectively. Five novel alleles, MICB*005:07, MICB*005:08, MICB*027, MICB*028, and MICB*029 were identified. Compared with the MICB*005:02:01 sequence, a G > A substitution was observed at nucleotide position 210 in MICB*005:07, and a 1,134 T > C substitution in MICB*005:08 and an 862 G > A substitution in MICB*027 were detected. In addition, it appears that MICB*028 probably arose from MICB*004:01:01 with an A to G substitution at position 1,147 in exon 6. MICB*029 had a G > T transversion at nucleotide position 730 in exon 4, compared with that of MICB*002:01:01. On the basis of the new PCR-SBT assay, these observed results demonstrated MICB allelic variations in the Zhejiang Han population.
Data from: Characterization of 42 polymorphic microsatellite loci in Mimulus ringens (Phrymaceae) using Illumina sequencing
Premise of the study: Microsatellite markers were isolated and characterized in Mimulus ringens (Phrymaceae), a herbaceous wetland perennial, to facilitate studies of mating patterns and population genetic structure. Methods and Results: A total of 42 polymorphic loci were identified from a sample of 24 individuals from a single popula- tion in Ohio, USA. The number of alleles per locus ranged from two to nine, and median observed heterozygosity was 0.435. Conclusions: This large number of polymorphic loci will enable researchers to quantify male fitness, patterns of multiple pa- ternity, selfing, and biparental inbreeding in large natural populations of this species. These markers will also permit detailed study of fine-scale patterns of genetic structure.
Data from: Targeted sequencing of venom genes from cone snail genomes improves understanding of conotoxin molecular evolution
To expand our capacity to discover venom sequences from the genomes of venomous organisms, we applied targeted sequencing techniques to selectively recover venom gene superfamilies and non-toxin loci from the genomes of 32 cone snail species (family, Conidae), a diverse group of marine gastropods that capture their prey using a cocktail of neurotoxic peptides (conotoxins). We were able to successfully recover conotoxin gene superfamilies across all species with high confidence (> 100X coverage) and used these data to provide new insights into conotoxin evolution. First, we found that conotoxin gene superfamilies are composed of 1-6 exons and are typically short in length (mean = ~85bp). Second, we expanded our understanding of the following genetic features of conotoxin evolution: (a) positive selection, where exons coding the mature toxin region were often three times more divergent than their adjacent noncoding regions, (b) expression regulation, with comparisons to transcriptome data showing that cone snails only express a fraction of the genes available in their genome (24%-63%), and (c) extensive gene turnover, where Conidae species varied from 120-859 conotoxin gene copies. Finally, using comparative phylogenetic methods, we found that while diet specificity did not predict patterns of conotoxin evolution, dietary breadth was positively correlated with total conotoxin gene diversity. Overall, the targeted sequencing technique demonstrated here has the potential to radically increase the pace at which venom gene families are sequenced and studied, reshaping our ability to understand the impact of genetic changes on ecologically relevant phenotypes and subsequent diversification.
Data from: Phylogenetic relationships of Agaric fungi based on nuclear large subunit ribosomal DNA sequences
Phylogenetic relationships of mushrooms and their relatives within the order Agaricales were addressed using nuclear large subunit ribosomal DNA sequences. Approximately 900 bases of the 5' end of the nucleus-encoded large subunit RNA gene (nLSU-rDNA) were sequenced for 154 selected taxa representing most families within the Agaricales. Several phylogenetic methods were used, including weighted and equally weighted parsimony (MP), maximum likelihood (ML), and distance methods (NJ). The starting tree for branch swapping in the ML analyses was the tree with the highest ML score among previously produced MP and NJ trees. A high degree of consensus was observed between phylogenetic estimates obtained through MP and ML. NJ trees differed according to the distance model that was used, however, all NJ trees still supported most of the same terminal groupings as MP and ML trees. NJ trees were always significantly suboptimal when evaluated against the best MP and ML trees, using both parsimony and likelihood tests. Our analyses suggest that weighted parsimony and ML provide the best estimates of Agaricales phylogeny. Similar support was observed between bootstrapping and jackknifing methods for evaluation of tree robustness. Phylogenetic analyses revealed many groups of agaricoid fungi that are supported by moderate to high bootstrap or jackknife levels or are consistent with morphology-based classification schemes. Analyzes also support separate placement of the boletes and russules, which are basal to the main core group of gilled mushrooms (the Agaricineae of Singer). Examples of monophyletic groups include the families Amanitaceae, Coprinaceae (excluding Coprinus comatus and subfamily Panaeolideae), Agaricaceae (excluding the Cystodermateae), and Strophariaceae pro parte (Stropharia, Pholiota, and Hypholoma); the mycorrhizal species of Tricholoma (including Leucopaxillus, also mycorrhizal); Mycena and Resinomycena; Termitomyces, Podabrella, and Lyophyllum; and Pleurotus with Hohenbuehelia. Several nonmonophyletic groups revealed by these data include the families Tricholomataceae, Cortinariaceae, and Hygrophoraceae and the genera Clitocybe, Omphalina, and Marasmius. This study provides a framework for future systematics studies in the Agaricales and suggestions for analyzing large molecular data sets.
Data from: A branch-heterogeneous model of protein evolution for efficient inference of ancestral sequences
Most models of nucleotide or amino acid substitution used in phylogenetic studies assume that the evolutionary process has been homogeneous across lineages and that composition of nucleotides or amino acids has remained the same throughout the tree. These oversimplified assumptions are refuted by the observation that compositional variability characterizes extant biological sequences. Branch-heterogeneous models of protein evolution that account for compositional variability have been developed, but are not yet in common use because of the large number of parameters required, leading to high computational costs and potential overparameterization. Here, we present a new branch-nonhomogeneous and nonstationary model of protein evolution that captures more accurately the high complexity of sequence evolution. This model, henceforth called Correspondence and likelihood analysis (COaLA), makes use of a correspondence analysis to reduce the number of parameters to be optimized through maximum likelihood, focusing on most of the compositional variation observed in the data. The model was thoroughly tested on both simulated and biological data sets to show its high performance in terms of data fitting and CPU time. COaLA efficiently estimates ancestral amino acid frequencies and sequences, making it relevant for studies aiming at reconstructing and resurrecting ancestral amino acid sequences. Finally, we applied COaLA on a concatenate of universal amino acid sequences to confirm previous results obtained with a nonhomogeneous Bayesian model regarding the early pattern of adaptation to optimal growth temperature, supporting the mesophilic nature of the Last Universal Common Ancestor.
Data from: Sequence data for Clostridium autoethanogenum using three generations of sequencing technologies
During the past decade, DNA sequencing output has been mostly dominated by the second generation sequencing platforms which are characterized by low cost, high throughput and shorter read lengths for example, Illumina. The emergence and development of so called third generation sequencing platforms such as PacBio has permitted exceptionally long reads (over 20 kb) to be generated. Due to read length increases, algorithm improvements and hybrid assembly approaches, the concept of one chromosome, one contig and automated finishing of microbial genomes is now a realistic and achievable task for many microbial laboratories. In this paper, we describe high quality sequence datasets which span three generations of sequencing technologies, containing six types of data from four NGS platforms and originating from a single microorganism, Clostridium autoethanogenum. The dataset reported here will be useful for the scientific community to evaluate upcoming NGS platforms, enabling comparison of existing and novel bioinformatics approaches and will encourage interest in the development of innovative experimental and computational methods for NGS data.
Data from: "You are not what you eat: massive parallel sequencing reveals that gut microbiome is not diet-related in larval Dilophus febrilis (Diptera: Bibionidae)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015
This article documents the public availability of metagenome sequence data from 454 amplicon sequencing of larval dipteran gut (Dilophus febrilis) and their potential food sources dwarf shrub litter (Vaccinium gaultheroides), grass litter (Dactylis glomerata), and cow dung (Bos primigenius taurus).
Data from: Chromosomal inversions and ecotypic differentiation in Anopheles gambiae: the perspective from whole-genome sequencing
The molecular mechanisms and genetic architecture that facilitate adaptive radiation of lineages remain elusive. Polymorphic chromosomal inversions, due to their recombination-reducing effect, are proposed instruments of ecotypic differentiation. Here we study an ecologically diversifying lineage of An. gambiae, known as the Bamako chromosomal form based on its unique complement of three chromosomal inversions, to explore the impact of these inversions on ecotypic differentiation. We used pooled and individual genome sequencing of Bamako, typical (non-Bamako) An. gambiae, and the sister species An. coluzzii to investigate evolutionary relationships and genome-wide patterns of nucleotide diversity and differentiation among lineages. Despite extensive shared polymorphism and limited differentiation from the other taxa, Bamako clusters apart from the other taxa, and forms a maximally supported clade in neighbor-joining trees based on whole genome data (including inversions) or solely on collinear regions. Nevertheless, FST outlier analysis reveals that the majority of differentiated regions between Bamako and typical An. gambiae are located inside chromosomal inversions, consistent with their role in the ecological isolation of Bamako. Exceptionally differentiated genomic regions were enriched for genes implicated in nervous system development and signaling. Candidate genes associated with a selective sweep unique to Bamako contain substitutions not observed in sympatric samples of the other taxa, and several insecticide resistance gene alleles shared between Bamako and other taxa segregate at sharply different frequencies in these samples. Bamako represents a useful window into the initial stages of ecological and genomic differentiation from sympatric populations in this important group of malaria vectors.
Data from: An examination of the accuracy of a sequential PCR and sequencing test used to detect the incursion of an invasive species: the case of the red fox in Tasmania
1. Polymerase Chain Reaction (PCR) diagnostic tests are increasingly applied to the identification of wildlife. Yet rigorous verification is rare and the estimation of test accuracy (the probability that true positive and true negative samples are correctly identified – test sensitivity and specificity, respectively), particularly in combination with sequencing, is uncommon. This is important because PCR-based tests are prone to contamination in sampling and the laboratory. 2. Here, we use an experimental case–control approach to estimate the sensitivity and specificity of a sequential PCR-based wildlife detection test used to identify incursions of red foxes into Tasmania from predator faeces (scats). 3. Our results show that the sensitivity of the fox test is high (~94%) for the PCR-based test on its own, but this decreases to ~84% when combined with the DNA sequencing step. In contrast, the specificity increases from ~96% in the PCR only test to ~99.6% after inclusion of the DNA sequencing step. 4. The intense public scrutiny of the fox eradication program in Tasmania, has undoubtedly influenced the application of a sequential PCR test that maximises specificity at the expense of sensitivity and so increases the risk that scats containing fox DNA would not be detected. This could lead to the establishment of foxes in Tasmania as a consequence. 5. Synthesis and applications. Importantly, the estimation of the sensitivity and specificity of sequential tests enables decisions about the risk associated with mistaken identification (i.e. false negatives vs false positives) to be quantified for decision makers. The cost of false negative errors should be balanced against the costs of false positive errors, which could include the expenditure incurred in the application of unnecessary management actions were foxes not in fact present. Understanding the risks and costs associated with both false negative and false positive errors is therefore a key component to the decision making process for the management of the Tasmanian fox incursion.
Data from: Kakusan4 and Aminosan: two programs for comparing nonpartitioned, proportional, and separate models for combined molecular phylogenetic analyses of multilocus sequence data
Proportional and separate models able to apply different combination of substitution rate matrix and among-site rate variation model to each locus are frequently used in phylogenetic studies of multilocus data. However, the selection from among nonpartitioned (i.e., a common combination of models is applied to all-loci concatenated sequences), proportional, and separate models is usually based on the researcher's preference rather than on any information criteria. The present study describes two programs, "Kakusan4" (for DNA sequences) and "Aminosan" (for amino-acid sequences), that allow the selection of evolutionary models based on several types of information criteria. The programs can handle both multilocus and single-locus data, in addition to providing an easy-to-use wizard interface and a non-interactive command line interface. In the case of multilocus data, substitution rate matrices and among-site rate variation models are compared at each locus and at all-loci concatenated sequences, after which nonpartitioned, proportional, and separate models are compared based on information criteria. The programs also provide model configuration files for MrBayes, PAUP*, PHYML, RAxML, and Treefinder to support further phylogenetic analysis using a selected model. The best-fit models were found to differ depending on the data set. Furthermore, differences in the information criteria among nonpartitioned, proportional, and separate models were much larger than those among the nonpartitioned models. These findings suggest that selecting from nonpartitioned, proportional, and separate models results in a better phylogenetic tree. Kakusan4 and Aminosan are available at http://www.fifthdimension.jp/. They are licensed under GNU GPL Ver.2, and are able to run on Windows, MacOS X, and Linux.
Data from: Resolving ambiguity of concatenation in multi-locus sequence data for the construction of phylogenetic supermatrices
The construction of supermatrices from mining of DNA metadata is problematic due to incomplete species identification and incongruence of gene trees that hamper sequence concatenation based on Linnaean binomials. We applied methods from graph theory to minimize ambiguity of concatenation globally over a large data set. An initial step establishes sequence clusters for each locus that broadly correspond to Linnaean species. These clusters frequently are not consistent with binomials and specimen identifiers, which greatly complicates the concatenation of clusters across multiple loci. A multipartite heuristic algorithm is used to match clusters across loci and to generate a global set of concatenates that minimizes conflict of taxonomic names. The procedure was applied to all available data on GenBank for the Coleoptera (beetles) including >10500 taxon labels for >23500 sequences of four loci. The BlastClust algorithm was used in the initial clustering step, resulting in 11241 clusters or divergent singletons. Clusters were first used for name assignment of unidentified sequences resulting in 510 new identifications (13.9% of total unidentified sequences) of which nearly half were by clustering of a specimen at a secondary locus. Concatenation was straightforward only for 12.8% of all binomials represented by a singleton sequence at each locus with an available entry, while the majority of binomials were associated to multi-sequence clusters in at least one locus. Concatenation of clusters is particularly problematic where limits of DNA-based clusters are inconsistent with the Linnaean binomials, either containing more than one binomial or splitting a binomial among multiple clusters. The current data set contained 1518 such clusters (13.5% of total). By applying a scoring scheme for full and partial name matches in pairs of clusters, the maximum weight set of concatenates produced a matrix of minimally 7366 terminals. Varying the match weights for partial matches had little effect on the number of terminals, although if partial matches were disallowed, the number of terminals increased greatly. Trees from the resulting supermatrices generally produced tree topologies in good agreement with the Linnaean taxonomy, with fewer terminals compared to trees generated according to standard species labels. The study illustrates a strategy for assembling the Tree-of-Life from an ever more complex primary database.
Data from: Seabird and Louse Coevolution: Complex Histories Revealed by 12S rRNA Sequences and Reconciliation Analyses
We investigated the coevolutionary history of seabirds (orders Procellariiformes and Sphenisciformes) and their lice (order Phthiraptera). Independent trees were produced for the seabirds (tree derived from 12S ribosomal RNA (rRNA), isozyme, and behavioral data) and their lice (trees derived from 12S rRNA data). Brookâ s parsimony analysis (BPA) supported a general history of cospeciation (consistency index = 0.84, retention index = 0.81). We inferred that the homoplasy in the BPA was caused by one intrahost speciation, one potential host switching and eight or nine sorting events. Using reconciliation analysis we quantified the cost of fitting the louse tree onto the seabird tree. The reconciled TreeMap tree postulated one host switching, nine cospeciation, three or four intrahost speciation and 11 to 14 sorting events. The number of cospeciation events was significantly more than would be expected due to chance. The sequence data were used to test for rate heterogeneity for both seabirds and lice. The seabird tree showed no significant rate heterogeneity over all of its branches whereas part of the louse tree did show rate heterogeneity. An examination of the codivergent nodes revealed that seabirds and lice have cospeciated synchronously, and that lice have evolved at about 5.5 times the rate of seabirds. Sequence data supported some of the postulated intrahost speciation events (Halipeurus pre-dated the evolution of their present hosts). Sequence data also supported some of the postulated host-switching events. These results demonstrate the value of sequence data and reconciliation analyses in unraveling complex histories between hosts and their parasites.
Data from: Whole-genome sequencing approaches for conservation biology: advantages, limitations, and practical recommendations
Whole-genome resequencing (WGR) is a powerful method for addressing fundamental evolutionary biology questions that have not been fully resolved using traditional methods. WGR includes four approaches: the sequencing of individuals to a high depth of coverage with either unresolved (huWGR) or resolved haplotypes (hrWGR), the sequencing of population genomes to a high depth by mixing equimolar amounts of unlabelled-individual DNA (Pool-seq), and the sequencing of multiple individuals from a population to a low depth (lcWGR). These techniques require the availability of a reference genome. This, along with the still high cost of shotgun sequencing and the large demand for computing resources and storage, has limited their implementation in non-model species with scarce genomic resources and in fields such as conservation biology. Our goal here is to describe the various WGR methods, their pros and cons, and potential applications in conservation biology. WGR offers an unprecedented marker density and surveys a wide diversity of genetic variations not limited to single nucleotide polymorphisms (e.g. structural variants and mutations in regulatory elements), increasing their power for the detection of signatures of selection and local adaptation as well as for the identification of the genetic basis of phenotypic traits and diseases. Currently though, no single WGR approach fulfills all requirements of conservation genetics, and each method has its own limitations and sources of potential bias. We discuss proposed ways to minimize such biases. We envision a not distant future where the analysis of whole genomes becomes a routine task in many non-model species and fields including conservation biology.
Data from: Sequencing of seven haloarchaeal genomes reveals patterns of genomic flux
We report the sequencing of seven genomes from two haloarchaeal genera, Haloferax and Haloarcula. Ease of cultivation and the existence of well-developed genetic and biochemical tools for several diverse haloarchaeal species make haloarchaea a model group for the study of archaeal biology. The unique physiological properties of these organisms also make them good candidates for novel enzyme discovery for biotechnological applications. Seven genomes were sequenced to ~20×coverage and assembled to an average of 50 contigs (range 5 scaffolds - 168 contigs). Comparisons of protein-coding gene compliments revealed large-scale differences in COG functional group enrichment between these genera. Analysis of genes encoding machinery for DNA metabolism reveals genera-specific expansions of the general transcription factor TATA binding protein as well as a history of extensive duplication and horizontal transfer of the proliferating cell nuclear antigen. Insights gained from this study emphasize the importance of haloarchaea for investigation of archaeal biology.
Data from: Target capture and massively parallel sequencing of ultraconserved elements for comparative studies at shallow evolutionary time scales
Comparative genetic studies of non-model organisms are transforming rapidly due to major advances in sequencing technology. A limiting factor in these studies has been the identification and screening of orthologous loci across an evolutionarily distant set of taxa. Here, we evaluate the efficacy of genomic markers targeting ultraconserved DNA elements (UCEs) for analyses at shallow evolutionary timescales. Using sequence capture and massively parallel sequencing to generate UCE data for five co-distributed Neotropical rainforest bird species, we recovered 776–1516 UCE loci across the five species. Across species, 53–77% of the loci were polymorphic, containing between 2.0 and 3.2 variable sites per polymorphic locus, on average. We performed species tree construction, coalescent modeling, and species delimitation, and we found that the five co-distributed species exhibited discordant phylogeographic histories. We also found that species trees and divergence times estimated from UCEs were similar to the parameters obtained from mtDNA. The species that inhabit the understory had older divergence times across barriers, contained a higher number of cryptic species, and exhibited larger effective population sizes relative to the species inhabiting the canopy. Because orthologous UCEs can be obtained from a wide array of taxa, are polymorphic at shallow evolutionary timescales, and can be generated rapidly at low cost, they are an effective genetic marker for studies investigating evolutionary patterns and processes at shallow timescales.
Data from: Mining microsatellite markers from public expressed sequence tags databases for the study of threatened plants
Background: Simple Sequence Repeats (SSRs) are widely used in population genetic studies but their classical development is costly and time-consuming. The ever-increasing available DNA datasets generated by high-throughput techniques offer an inexpensive alternative for SSRs discovery. Expressed Sequence Tags (ESTs) have been widely used as SSR source for plants of economic relevance but their application to non-model species is still modest. Methods: Here, we explored the use of publicly available ESTs (GenBank at the National Center for Biotechnology Information-NCBI) for SSRs development in non-model plants, focusing on genera listed by the International Union for the Conservation of Nature (IUCN). We also search two model genera with fully annotated genomes for EST-SSRs, Arabidopsis and Oryza, and used them as controls for genome distribution analyses. Overall, we downloaded 16 031 555 sequences for 258 plant genera which were mined for SSRsand their primers with the help of QDD1. Genome distribution analyses in Oryza and Arabidopsis were done by blasting the sequences with SSR against the Oryza sativa and Arabidopsis thaliana reference genomes implemented in the Basal Local Alignment Tool (BLAST) of the NCBI website. Finally, we performed an empirical test to determine the performance of our EST-SSRs in a few individuals from four species of two eudicot genera, Trifolium and Centaurea. Results: We explored a total of 14 498 726 EST sequences from the dbEST database (NCBI) in 257 plant genera from the IUCN Red List. We identify a very large number (17 102) of ready-to-test EST-SSRs in most plant genera (193) at no cost. Overall, dinucleotide and trinucleotide repeats were the prevalent types but the abundance of the various types of repeat differed between taxonomic groups. Control genomes revealed that trinucleotide repeats were mostly located in coding regions while dinucleotide repeats were largely associated with untranslated regions. Our results from the empirical test revealed considerable amplification success and transferability between congenerics. Conclusions: The present work represents the first large-scale study developing SSRs by utilizing publicly accessible EST databases in threatened plants. Here we provide a very large number of ready-to-test EST-SSR (17 102) for 193 genera. The cross-species transferability suggests that the number of possible target species would be large. Since trinucleotide repeats are abundant and mainly linked to exons they might be useful in evolutionary and conservation studies. Altogether, our study highly supports the use of EST databases as an extremely affordable and fast alternative for SSR developing in threatened plants.
Data from: Integrating sequence evolution into probabilistic orthology analysis
Orthology analysis, that is, finding out whether a pair of homologous genes are orthologs — stemming from a speciation — or paralogs — stemming from a gene duplication - is of central importance in computational biology, genome annotation, and phylogenetic inference. In particular, an orthologous relationship makes functional equivalence of the two genes highly likely. A major approach to orthology analysis is to reconcile a gene tree to the corresponding species tree, (most commonly performed using the most parsimonious reconciliation, MPR). However, most such phylogenetic orthology methods infer the gene tree without considering the constraints implied by the species tree and, perhaps even more importantly, only allow the gene sequences to influence the orthology analysis through the a priori reconstructed gene tree. We propose a sound, comprehensive Bayesian Markov chain Monte Carlo-based method, DLRSOrthology, to compute orthology probabilities. It efficiently sums over the possible gene trees and jointly takes into account the current gene tree, all possible reconciliations to the species tree, and the, typically strong, signal conveyed by the sequences. We compare our method with PrIME-GEM, a probabilistic orthology approach built on a probabilistic duplication-loss model, and MRBAYESMPR, a probabilistic orthology approach that is based on conventional Bayesian inference coupled with MPR. We find that DLRSOrthology outperforms these competing approaches on synthetic data as well as on biological data sets and is robust to incomplete taxon sampling artifacts.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.