Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data from: Congruent species delimitation of two controversial gold-thread nanmu tree species based on morphological and restriction site-associated DNA sequencing data
Species delimitation is fundamental to conservation and sustainable use of economically important forest tree species. However, the delimitation of two highly valued gold-thread nanmu species (Phoebe bournei and P. zhennan) has been confusing and debated. To address this problem, we integrated morphology and restriction site-associated DNA sequencing (RADseq) to define their species boundaries. We obtained highly consistent results from both data sets, supporting two distinct lineages corresponding to P. bournei and P. zhennan. In Phoebe bournei, higher order leaf venation is more prominent, petioles are thicker and leaf apex angle is narrower, compared to P. zhennan. Both data sets also showed that putative P. bournei localities from north-eastern Guizhou were P. zhennan. The two species have different distributions and only overlap in the Wuling Mountains. Phoebe bournei occurs mainly in Central Fujian, southern Jiangxi, the Nanling Mountains and the Wuling Mountains, whereas P. zhennan is found in the adjoining eastern regions of the Qionglai Mountains, the Southern Sichuan Hills and the Wuling Mountains. The improved delimitation of P. bournei and P. zhennan and clarification of their ranges provide a better guidance for conservation and sustainable utilization of these tree species.
Data from: Species tree estimation of North American chorus frogs (Hylidae: Pseudacris) with parallel tagged amplicon sequencing
The field of phylogenetics is changing rapidly with the application of high-throughput sequencing to non-model organisms. Cost-effective use of this technology for phylogenetic studies, which often include a relatively small portion of the genome but several taxa, requires strategies for genome partitioning and sequencing multiple individuals in parallel. In this study we estimated a multilocus phylogeny for the North American chorus frog genus Pseudacris using anonymous nuclear loci that were recently developed using a reduced representation library approach. We sequenced 27 nuclear loci and three mitochondrial loci for 44 individuals on 1/3 of an Illumina MiSeq run, obtaining 96.5% of the targeted amplicons at less than 20% of the cost of traditional Sanger sequencing. We found heterogeneity among gene trees, although four major clades (Trilling Frog, Fat Frog, crucifer, and West Coast) were consistently supported, and we resolved the relationships among these clades for the first time with strong support. We also found discordance between the mitochondrial and nuclear datasets that we attribute to mitochondrial introgression and a possible selective sweep. Bayesian concordance analysis in BUCKy and species tree analysis in *BEAST produced largely similar topologies, although we identify taxa that require additional investigation in order to clarify taxonomic and geographic range boundaries. Overall, we demonstrate the utility of a reduced representation library approach for marker development and parallel tagged sequencing on an Illumina MiSeq for phylogenetic studies of non-model organisms.
Data from: Discrimination of grasshopper (Orthoptera: Acrididae) diet and niche overlap using next-generation sequencing of gut contents
Species of grasshopper have been divided into three diet classifications based on mandible morphology: forbivorous (specialist on forbs), graminivorous (specialist on grasses), and mixed feeding (broad-scale generalists). For example, Melanoplus bivittatus and Dissosteira carolina are presumed to be broad-scale generalists, Chortophaga viridifasciata is a specialist on grasses, and Melanoplus femurrubrum is a specialist on forbs. These classifications, however, have not been verified in the wild. Multiple specimens of these four species were collected, and diet analysis was performed using DNA metabarcoding of the gut contents. The rbcLa gene region was amplified and sequenced using Illumina MiSeq sequencing. Levins' measure and the Shannon–Wiener measure of niche breadth were calculated using family-level identifications and Morisita's measure of niche overlap was calculated using operational taxonomic units (OTUs). Gut contents confirm both D. carolina and M. bivittatus as generalists and C. viridifasciata as a specialist on grasses. For M. femurrubrum, a high niche breadth was observed and species of grasses were identified in the gut as well as forbs. Niche overlap values did not follow predicted patterns, however, the low values suggest low competition between these species.
Data from: Megaphylogenetic specimen-level approaches to the Carex (Cyperaceae) phylogeny using ITS, ETS, and matK sequences: implications for classification
We present the first large-scale phylogenetic hypothesis for the genus Carex based on 996 of the 1983 accepted species (50.23%). We used a supermatrix approach using three DNA regions: ETS, ITS and matK. Every concatenated sequence was derived from a single specimen. The topology of our phylogenetic reconstruction largely agreed with previous studies. We also gained new insights into the early divergence structure of the two largest clades, core Carex and Vignea clades, challenging some previous evolutionary hypotheses about inflorescence structure. Most sections were recovered as non-monophyletic. Homoplasy of characters traditionally selected as relevant for classification, historical misunderstanding of how morphology varies across Carex, and regional rather than global views of Carex diversity seem to be the main reasons for the high levels of polyphyly and paraphyly in the current infrageneric classification.
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Despite the recent advances in generating molecular data, reconstructing species-level phylogenies for non-models groups remains a challenge. The use of a number of independent genes is required to resolve phylogenetic relationships, especially for groups displaying low polymorphism. In such cases, low-copy nuclear exons and non-coding regions, such as 3′ untranslated regions (3′-UTRs) or introns, constitute a potentially interesting source of nuclear DNA variation. Here, we present a methodology meant to identify new nuclear orthologous markers using both public-nucleotide databases and transcriptomic data generated for the group of interest by using next generation sequencing technology. To identify PCR primers for a non-model group, the genus Leucadendron (Proteaceae), we adopted a framework aimed at minimizing the probability of paralogy and maximizing polymorphism. We anchored when possible the right-hand primer into the 3′-UTR and the left-hand primer into the coding region. Seven new nuclear markers emerged from this search strategy, three of those included 3′-UTRs. We further compared the phylogenetic potential between our new markers and the ribosomal internal transcribed spacer region (ITS). The sequenced 3′-UTRs yielded higher polymorphism rates than the ITS region did. We did not find strong incongruences with the phylogenetic signal contained in the ITS region and the seven new designed markers but they strongly improved the phylogeny of the genus Leucadendron. Overall, this methodology is efficient in isolating orthologous loci and is valid for any non-model group given the availability of transcriptomic data.
Data from: Assemblage Accumulation Curves: A framework for resolving species accumulation in biological communities using chloroplast genome sequences
The timing and tempo of the processes involved in community assembly are of substantial concern to community ecologists and conservation managers. The fossil record is a valuable source of data for studying past changes in community composition, but it is not always detailed enough to allow the process of community assembly to be resolved at regional or site scales while tracing the trajectories of known species with associated known traits. We present a three‐step framework for studying present‐day species accumulation through time: DNA sampling from multiple individuals from multiple species within a community; estimates of coalescence times for each species using molecular dating methods; and plotting the accumulation of present‐day species through time using the inferred population ages. Our approach is illustrated using whole chloroplast genomes from plants from three rainforest communities in eastern Australia. Expected times to coalescence for multiple species in each community were inferred from pooled high‐throughput sequence libraries. Local assemblage accumulation curves for each community were constructed. We also explored the variation in assemblage accumulation curves of species with different functional traits. Models of equilibrium species richness informed our null hypothesis and largely explained the shape of the assemblage accumulation curves and indicated that the complexities of the accumulation process should be explored with additional parameters, for example allowing species classes with different extinction rates. The assemblage accumulation curves for the study sites showed evidence of recent population expansions within each of the communities. This signal of recent accumulation is consistent with the increase in suitable rainforest habitat that followed the Last Glacial Maximum. Our method of constructing assemblage accumulation curves provides a simple approach for visualizing species‐accumulation data. It can be used to test hypotheses such as the relative survival potential of species‐specific ecological attributes. Although our example used single‐nucleotide polymorphisms derived from whole‐chloroplast sequencing, this framework can be applied to mitochondrial genomes and to communities of other organisms.
Data from: Heterochromatin-enriched assemblies reveal the sequence and organization of the Drosophila melanogaster Y chromosome
Heterochromatic regions of the genome are repeat-rich and poor in protein coding genes, and are therefore underrepresented in even the best genome assemblies. One of the most difficult regions of the genome to assemble are sex-limited chromosomes. The Drosophila melanogaster Y chromosome is entirely heterochromatic, yet has wide-ranging effects on male fertility, fitness, and genome-wide gene expression. The genetic basis of this phenotypic variation is difficult to study, in part because we do not know the detailed organization of the Y chromosome. To study Y chromosome organization in D. melanogaster, we develop an assembly strategy involving the in silico enrichment of heterochromatic long single-molecule reads and use these reads to create targeted de novo assemblies of heterochromatic sequences. We assigned contigs to the Y chromosome using Illumina reads to identify male-specific sequences. Our pipeline extends the D. melanogaster reference genome by 11.9 Mb, closes 43.8% of the gaps, and improves overall contiguity. The addition of 10.6 MB of Y-linked sequence permitted us to study the organization of repeats and genes along the Y chromosome. We detected a high rate of duplication to the pericentric regions of the Y chromosome from other regions in the genome. Most of these duplicated genes exist in multiple copies. We detail the evolutionary history of one sex-linked gene family—crystal-Stellate. While the Y chromosome does not undergo crossing over, we observed high gene conversion rates within and between members of the crystal-Stellate gene family, Su(Ste), and PCKR, compared to genome-wide estimates. Our results suggest that gene conversion and gene duplication play an important role in the evolution of Y-linked genes.
Data from: Two new phragmotic ant species from Africa: morphology and next-generation sequencing solve a caste association problem in the genus Carebara Westwood
Phragmotic or "door head" ants have evolved independently in several ant genera across the world, but in Africa only one case has been documented until now. Carebara elmenteitae (Patrizi) is known from only a single phragmotic major worker collected from sifted leaf-litter near Lake Elmenteita in Kenya, but here the worker castes of two species collected from Kakamega Forest, a small rainforest in Western Kenya, are studied. Phragmotic major workers were previously identified as Carebara elmenteitae and non-phragmotic major and minor workers were assigned to C. thoracica (Weber). Using evidence of both morphological and next-generation sequencing analysis, it is shown that phragmotic and non-phragmotic workers of the two different species are actually the same and that neither name – C. elmenteitae or C. thoracica – correctly applies to them. Instead, this and another closely related species from Ivory Coast are both morphologically different from C. elmenteitae, and thus they are described as the new species Carebara phragmotica sp. n. and Carebara lilith sp. n.
Data from: Transcriptome sequencing reveals both neutral and adaptive genome dynamics in a marine invader
Species invasions cause significant ecological and economic damage, and genetic information is important to understanding and managing invasive species. In the ocean, many invasive species have high dispersal and gene flow, lowering the discriminatory power of traditional genetic approaches. High-throughput sequencing holds tremendous promise for increasing resolution and illuminating the relative contributions of selection and drift in marine invasion, but has not yet been used to compare the diversity and dynamics of a high-dispersal invader in its native and invaded ranges. We test a transcriptome-based approach in the European green crab (Carcinus maenas), a widespread invasive species with high gene flow and a well-known invasion history, in two native and five invasive populations. A panel of 10 809 transcriptome-derived nuclear SNPs identified significant population structure among highly bottlenecked invasive populations that were previously undifferentiated with traditional markers. Comparing the full data set and a subset of 9246 putatively neutral SNPs strongly suggested that non-neutral processes are the primary driver of population structure within the species' native range, while neutral processes appear to dominate in the invaded range. Non-neutral native range structure coincides with significant differences in intraspecific thermal tolerance, suggesting temperature as a potential selective agent. These results underline the importance of adaptation in shaping intraspecific differences even in high geneflow marine invasive species. They also demonstrate that high-throughput approaches have broad utility in determining neutral structure in recent invasions of such species. Together, neutral and non-neutral data derived from high-throughput approaches may increase the understanding of invasion dynamics in high-dispersal species.
Data from: Design of a 9K SNP chip for polar bears (Ursus maritimus) from RAD and transcriptome sequencing
Single-nucleotide polymorphisms (SNPs) offer numerous advantages over anonymous markers such as microsatellites, including improved estimation of population parameters, finer-scale resolution of population structure and more precise genomic dissection of quantitative traits. However, many SNPs are needed to equal the resolution of a single microsatellite, and reliable large-scale genotyping of SNPs remains a challenge in nonmodel species. Here, we document the creation of a 9K Illumina Infinium BeadChip for polar bears (Ursus maritimus), which will be used to investigate: (i) the fine-scale population structure among Canadian polar bears and (ii) the genomic architecture of phenotypic traits in the Western Hudson Bay subpopulation. To this end, we used restriction-site associated DNA (RAD) sequencing from 38 bears across their circumpolar range, as well as blood/fat transcriptome sequencing of 10 individuals from Western Hudson Bay. Six-thousand RAD SNPs and 3000 transcriptomic SNPs were selected for the chip, based primarily on genomic spacing and gene function respectively. Of the 9000 SNPs ordered from Illumina, 8042 were successfully printed, and – after genotyping 1450 polar bears – 5441 of these SNPs were found to be well clustered and polymorphic. Using this array, we show rapid linkage disequilibrium decay among polar bears, we demonstrate that in a subsample of 78 individuals, our SNPs detect known genetic structure more clearly than 24 microsatellites genotyped for the same individuals and that these results are not driven by the SNP ascertainment scheme. Here, we present one of the first large-scale genotyping resources designed for a threatened species.
Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages
Molecular phylogenetic studies of homologous sequences of nucleotides often assume that the underlying evolutionary process was globally stationary, reversible and homogeneous (SRH), and that a model of evolution with one or more site-specific and time-reversible rate matrices (e.g., the GTR rate matrix) is enough to accurately model the evolution of data over the whole tree. However, an increasing body of data suggests that evolution under these conditions is an exception, rather than the norm. To address this issue, several non-SRH models of molecular evolution have been proposed, but they either ignore heterogeneity in the substitution process across sites (HAS) or assume it can be modelled accurately using the Γ distribution. As an alternative to these models of evolution, we introduce a family of mixture models that approximate HAS without the assumption of an underlying predefined statistical distribution. This family of mixture models is combined with non-SRH models of evolution that account for heterogeneity in the substitution process across lineages (HAL). We also present two algorithms for searching model space and identifying an optimal model of evolution that is less likely to over- or under-parameterize the data. The performance of the two new algorithms was evaluated using alignments of nucleotides with 10,000 sites simulated under complex non-SRH conditions on a 25-tipped tree. The algorithms were found to be very successful, identifying the correct HAL model with a 75% success rate (the average success rate for assigning rate matrices to the tree's 48 edges was 99.25%) and, for the correct HAL model, identifying the correct HAS model with a 98% success rate. Finally, parameter estimates obtained under the correct HAL-HAS model were found to be accurate and precise. The merits of our new algorithms were illustrated with an analysis of 42,337 second codon sites extracted from a concatenation of 106 alignments of orthologous genes encoded by the nuclear genomes of Saccharomyces cerevisiae, S. paradoxus, S. mikatae, S. kudriavzevii, S. castellii, S. kluyveri, S. bayanus, and Candida albicans. Our results show that second codon sites in the ancestral genome of these species contained 49.1% invariable sites, 39.6% variable sites belonging to one rate category (V1), and 11.3% variable sites belonging to a second rate category (V2). The ancestral nucleotide content was found to differ markedly across these 3 sets of sites, and the evolutionary processes operating at the variable sites were found to be non-SRH and best modelled by a combination of 8 edge-specific rate matrices (4 for V1 and 4 for V2). The number of substitutions per site at the variable sites also differed markedly, with sites belonging to V1 evolving slower than those belonging to V2 along the lineages separating the 7 species of Saccharomyces. Finally, sites belonging to V1 appeared to have ceased evolving along the lineages separating S. cerevisiae, S. paradoxus, S. mikatae, S. kudriavzevii, and S. bayanus, implying that they might have become so selectively constrained that they could be considered invariable sites in these species.
Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates
Recent studies have advocated biomonitoring using DNA techniques. In this study, two high-throughput sequencing (HTS)-based methods were evaluated: amplicon metabarcoding of the cytochrome C oxidase subunit I (COI) mitochondrial gene and gene enrichment using MYbaits (targeting nine different genes including COI). The gene-enrichment method does not require PCR amplification and thus avoids biases associated with universal primers. Macroinvertebrate samples were collected from 12 New Zealand rivers. Macroinvertebrates were morphologically identified and enumerated, and their biomass determined. DNA was extracted from all macroinvertebrate samples and HTS undertaken using the illumina miseq platform. Macroinvertebrate communities were characterized from sequence data using either six genes (three of the original nine were not used) or just the COI gene in isolation. The gene-enrichment method (all genes) detected the highest number of taxa and obtained the strongest Spearman rank correlations between the number of sequence reads, abundance and biomass in 67% of the samples. Median detection rates across rare (<1% of the total abundance or biomass), moderately abundant (1–5%) and highly abundant (>5%) taxa were highest using the gene-enrichment method (all genes). Our data indicated primer biases occurred during amplicon metabarcoding with greater than 80% of sequence reads originating from one taxon in several samples. The accuracy and sensitivity of both HTS methods would be improved with more comprehensive reference sequence databases. The data from this study illustrate the challenges of using PCR amplification-based methods for biomonitoring and highlight the potential benefits of using approaches, such as gene enrichment, which circumvent the need for an initial PCR step.
Data from: The transcriptomics of sympatric dwarf and normal lake whitefish (Coregonus clupeaformis spp., Salmonidae) divergence as revealed by next-generation sequencing
Gene expression divergence is one of the mechanisms thought to be involved in the emergence of incipient species. Next-generation sequencing has become an extremely valuable tool for the study of this process by allowing whole transcriptome sequencing, or RNA-Seq. We have conducted a 454 GS-FLX pyrosequencing experiment in order to refine our understanding of adaptive divergence between dwarf and normal lake whitefish species (Coregonus clupeaformis spp.). The objectives were to: (1) investigate transcriptomic divergence as measured by liver RNA-Seq; (2) test the correlation between divergence in expression and sequence polymorphism and (3) investigate the extent of allelic imbalance. We also compared the results of RNA-seq with those of a previous microarray study performed on the same fish. Following de novo assembly, results showed that normal whitefish over-expressed more contigs associated with protein synthesis while dwarf fish over-expressed more contigs related to energy metabolism, immunity and DNA replication and repair. Moreover, 63 SNPs showed significant allelic imbalance, and this phenomenon prevailed in the recently diverged dwarf whitefish. Results also showed an absence of correlation between gene expression divergence as measured by RNA-Seq and either polymorphism rate or sequence divergence between normal and dwarf whitefish. This study reiterates an important role for gene expression divergence, and provides evidence for allele-specific expression divergence as well as evolutionary decoupling of regulatory and coding sequences in the adaptive divergence of normal and dwarf whitefish. It also demonstrates how next-generation sequencing can lead to a more comprehensive understanding of transcriptomic divergence in a young species pair.
Data from: Phylogenetic relationships of Iranian Allium sect. Allium (Amaryllidaceae, Allioideae) as inferred from nrDNA ITS, cpDNA rps16 and trnL–F sequences
Allium is a particularly species rich (more than 800 species) and economically important genus, with numerous taxonomic problems at all levels of classification. In this study, we try to uncover the phylogenetic relationships in the common leek (A. ampeloprasum) based on selected samples of this species and its putative relatives in sect. Allium from Iran. The silica-dried leaf samples of 56 accessions representing 23 species of Allium were sequenced for this study, 53 sequences of nrDNA ITS, 35 sequences of plastid rps16 and 52 sequences of trnL-F were generated and several accessions were extracted from GenBank in order to cover all recognized main lineages in the genus. Maximum Parsimony and Bayesian Inference generated similar trees, but the placement of A. ampeloprasum and its relatives differs slightly in the nuclear versus plastid datasets. In the nrITS tree A. ampeloprasum is retrieved in a highly supported clade with A. iranicum, while in the combined plastid tree A. ampeloprasum formed a highly supported clade with A. vineale. This supports the hypothesis of a possible hybrid origin of A. ampeloprasum. Allium iranicum formed a clade in the plastid tree, but was resolved as paraphyletic in the nrITS tree, probably due to presence of multiple non-concerted copies of nrITS. Close relationships are suggested between following species: A. aznavense and A. wendelboi with A. talyschense, A. erubescens and A. rotundum with A. scorodoprasum, and A. abbasii with A. phanerantherum.
Data from: Predicting function from sequence in a large multifunctional toxin family
Venoms contain active substances with highly specific physiological effects and are increasingly being used as sources of novel diagnostic, research and treatment tools for human disease. Experimental characterisation of individual toxin activities is a severe rate-limiting step in the discovery process, and in-silico tools which allow function to be predicted from sequence information are essential. Toxins are typically members of large multifunctional families of structurally similar proteins that can have different biological activities, and minor sequence divergence can have significant consequences. Thus, existing predictive tools tend to have low accuracy. We investigated a classification model based on physico-chemical attributes that can easily be calculated from amino-acid sequences, using over 250 (mostly novel) viperid phospholipase A2 toxins. We also clustered proteins by sequence profiles, and carried out in-vitro tests for four major activities on a selection of isolated novel toxins, or crude venoms known to contain them. The majority of detected activities were consistent with predictions, in contrast to poor performance of a number of tested existing predictive methods. Our results provide a framework for comparison of active sites among different functional sub-groups of toxins that will allow a more targeted approach for identification of potential drug leads in the future.
Data from: Phylogenomic analyses of Sabal (Arecaceae) species relationships using targeted sequence capture
With the increasing availability of high-throughput sequencing, phylogenetic analyses are no longer constrained by the limited availability of a few loci. Here, we describe a sequence capture methodology, which we used to collect data for analyses of diversification within Sabal (Arecaceae), a palm genus native to the south-eastern USA, Caribbean, Bermuda and Central America. RNA probes were developed and used to enrich DNA samples for putatively low copy nuclear genes and the plastomes for all Sabal species and two outgroup species. Sequence data were generated on an Illumina MiSeq sequencer and target sequences were assembled using custom workflows. Both coalescence and supermatrix analyses of 133 nuclear genes were used to estimate species trees relationships. Plastid genomes were also analysed, yielding generally poor resolution with regard to species relationships. Species relationships described in both nuclear gene and plastome sequences largely reflect the biogeography of the group and, to a lesser extent, previous morphology-based hypotheses. Beyond the biological implications, this research validates a high-throughput methodology for generating a large number of genes for coalescence-based phylogenetic analyses in plant lineages.
Data from: RAD sequencing resolves fine-scale population structure in a benthic invertebrate: implications for understanding phenotypic plasticity
The field of molecular ecology is transitioning from the use of small panels of classical genetic markers such as microsatellites to much larger panels of single nucleotide polymorphisms (SNPs) generated by approaches like RAD sequencing. However, few empirical studies have directly compared the ability of these methods to resolve population structure. This could have implications for understanding phenotypic plasticity, as many previous studies of natural populations may have lacked the power to detect genetic differences, especially over micro-geographic scales. We therefore compared the ability of microsatellites and RAD sequencing to resolve fine-scale population structure in a commercially important benthic invertebrate by genotyping great scallops (Pecten maximus) from nine populations around Northern Ireland at 13 microsatellites and 10 539 SNPs. The shells were then subjected to morphometric and colour analysis in order to compare patterns of phenotypic and genetic variation. We found that RAD sequencing was superior at resolving population structure, yielding higher Fst values and support for two distinct genetic clusters, whereas only one cluster could be detected in a Bayesian analysis of the microsatellite dataset. Furthermore, appreciable phenotypic variation was observed in size-independent shell shape and coloration, including among localities that could not be distinguished from one another genetically, providing support for the notion that these traits are phenotypically plastic. Taken together, our results suggest that RAD sequencing is a powerful approach for studying population structure and phenotypic plasticity in natural populations.
Data from: Exploring evolution and diversity of Chinese Dipterocarpaceae using next-generation sequencing
Tropical forests, a key-category of land ecosystems, are faced with the world's highest levels of habitat conversion and associated biodiversity loss. In tropical Asia, Dipterocarpaceae are one of the economically and ecologically most important tree families, but their genomic diversity and evolution remain understudied, hampered by a lack of available genetic resources. Southern China represents the northern limit for Dipterocarpaceae, and thus changes in habitat ecology, community composition and adaptability to climatic conditions are of particular interest in this group. Phylogenomics is a tool for exploring both biodiversity and evolutionary relationships through space and time using plastome, nuclear and mitochondrial genome. We generated full plastome and Nuclear Ribosomal Cistron (NRC) data for Chinese Dipterocarpaceae species as a first step to improve our understanding of their ecology and evolutionary relationships. We generated the plastome of Dipterocarpus turbinatus, the species with the widest distribution using it as a baseline for comparisons with other taxa. Results showed low level of genomic diversity among analysed range-edge species, and different evolutionary history of the incongruent NRC and plastome data. Genomic resources provided in this study will serve as a starting point for future studies on conservation and sustainable use of these dominant forest taxa, phylogenomics and evolutionary studies.
Data from: Restriction-site-associated DNA sequencing reveals a cryptic viburnum species on the North American coastal plain
Species are the starting point for most studies of ecology and evolution, but the proper circumscription of species can be extremely difficult in morphologically variable lineages, and there are still few convincing examples of molecularly-informed species delimitation in plants. We focus here on the Viburnum nudum complex, a highly variable clade that is widely distributed in eastern North America. Taxonomic treatments have mostly divided this complex into northern (V. nudum var. cassinoides) and southern (V. nudum var. nudum) entities, but additional names have been proposed. We used multiple lines of evidence, including RADseq, morphological, and geographic data, to test how many independently evolving lineages exist within the V. nudum complex. Genetic clustering and phylogenetic methods revealed three distinct groups—one lineage that is highly divergent, and two others that are recently diverged and morphologically similar. A combination of evidence that includes reciprocal monophyly, lack of introgression, and discrete rather than continuous patterns of variation supports the recognition of all three lineages as separate species. These results identify a surprising case of cryptic diversity in which two broadly sympatric species have consistently been lumped in taxonomic treatments. The clarity of our findings is directly related to the dense sampling and high quality genetic data in this study. We argue that there is a critical need for carefully sampled and integrative species delimitation studies to clarify species boundaries even in well-known plant lineages. Studies following the model that we have developed here are likely to identify many more cryptic lineages and will fundamentally improve our understanding of plant speciation and patterns of species richness.
Data from: A high-density linkage map for Astyanax mexicanus using genotyping-by-sequencing technology
The Mexican tetra, Astyanax mexicanus, is a unique model system consisting of cave-adapted and surface-dwelling morphotypes which diverged >1My ago. This remarkable natural experiment has enabled powerful genetic analyses of cave adaptation. Here, we describe the application of next-generation sequencing technology to the creation of a high-density linkage map. Our map comprises over 2200 markers populating 25 linkage groups constructed from genotypic data generated from a single genotyping-by-sequencing project. We leveraged emergent genomic and transcriptomic resources to anchor hundreds of anonymous Astyanax markers to the genome of the zebrafish (Danio rerio), the most closely related model organism to our study species. This facilitated the identification of 784 distinct connections between our linkage map and the Danio rerio genome, highlighting several regions of conserved genomic architecture between the two species despite ~150My of divergence. Using a Mendelian cave-associated trait as a proof-of-principle, we successfully recovered the genomic position of the albinism locus near the gene Oca2. Further, our map successfully informed the positions of unplaced Astyanax genomic scaffolds within particular linkage groups. This ability to identify the relative location, orientation and linear order of unaligned genomic scaffolds will facilitate ongoing efforts to improve upon the current early draft and assemble future versions of the Astyanax physical genome. Moreover, this improved linkage map will enable higher resolution genetic analyses and catalyze the discovery of the genetic basis for cave-associated phenotypes.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.