Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: 16S rRNA amplicon sequencing for epidemiological surveys of bacteria in wildlife

The human impact on natural habitats is increasing the complexity of human-wildlife interactions and leading to the emergence of infectious diseases worldwide. Highly successful synanthropic wildlife species, such as rodents, will undoubtedly play an increasingly important role in transmitting zoonotic diseases. We investigated the potential for recent developments in 16S rRNA amplicon sequencing to facilitate the multiplexing of the large numbers of samples needed to improve our understanding of the risk of zoonotic disease transmission posed by urban rodents in West Africa. In addition to listing pathogenic bacteria in wild populations, as in other high-throughput sequencing (HTS) studies, our approach can estimate essential parameters for studies of zoonotic risk, such as prevalence and patterns of coinfection within individual hosts. However, the estimation of these parameters requires cleaning of the raw data to mitigate the biases generated by HTS methods. We present here an extensive review of these biases and of their consequences, and we propose a comprehensive trimming strategy for managing these biases. We demonstrated the application of this strategy using 711 commensal rodents, including 208 Mus musculus domesticus, 189 Rattus rattus, 93 Mastomys natalensis, and 221 Mastomys erythroleucus, collected from 24 villages in Senegal. Seven major genera of pathogenic bacteria were detected in their spleens: Borrelia, Bartonella, Mycoplasma, Ehrlichia, Rickettsia, Streptobacillus, and Orientia. Mycoplasma, Ehrlichia, Rickettsia, Streptobacillus, and Orientia have never before been detected in West African rodents. Bacterial prevalence ranged from 0% to 90% of individuals per site, depending on the bacterial taxon, rodent species, and site considered, and 26% of rodents displayed coinfection. The 16S rRNA amplicon sequencing strategy presented here has the advantage over other molecular surveillance tools of dealing with a large spectrum of bacterial pathogens without requiring assumptions about their presence in the samples. This approach is therefore particularly suitable to continuous pathogen surveillance in the context of disease-monitoring programs.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genotyping by sequencing reveals contrasting patterns of population structure, ecologically mediated divergence and long-distance dispersal in North American palms

Comparative studies can provide powerful insights into processes that affect population divergence and thereby help to elucidate the mechanisms by which contemporary populations may respond to environmental change. Furthermore, approaches such as genotyping by sequencing (GBS) provide unprecedented power for resolving genetic differences among species and populations. We therefore used GBS to provide a genome-wide perspective on the comparative population structure of two palm genera, Washingtonia and Brahea, on the Baja California peninsula, a region of high landscape and ecological complexity. First, we used phylogenetic analysis to address taxonomic uncertainties among five currently recognised species. We resolved three main clades, the first corresponding to W. robusta and W. filifera, the second to B. brandegeei and B. armata, and the third to B. edulis from Guadalupe Island. Focusing on the first two clades, we then delved deeper by investigating the underlying population structure. Striking differences were found, with GBS uncovering four distinct Washingtonia populations and identifying a suite of loci associated with temperature, consistent with ecologically mediated divergence. By contrast, individual mountain ranges could be resolved in Brahea and few loci were associated with environmental variables, implying a more prominent role of neutral divergence. Finally, evidence was found for long-distance dispersal events in Washingtonia but not Brahea, in line with knowledge of the dispersal mechanisms of these palms including the possibility of human-mediated dispersal. Overall, our study demonstrates the power of GBS together with a comparative approach to elucidate markedly different patterns of genome-wide divergence mediated by multiple effectors.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Extent and variability of interstitial telomeric sequences and their effects on estimates of telomere length

Telomeres often shorten with time, although this varies between tissues, individuals and species, and their length and/or rate of change may reflect fitness and rate of senescence. Measurement of telomeres is increasingly important to ecologists, yet the relative merits of different methods for estimating telomere length are not clear. In particular the extent to which interstitial telomere sequences (ITSs), telomere repeats located away from chromosomes ends, confound estimates of telomere length is unknown. Here we present a method to estimate the extent of ITS within a species and variation among individuals. We estimated the extent of ITS by comparing the amount of label hybridized to in-gel telomere restriction fragments (TRF) before and after the TRFs were denatured. This protocol produced robust and repeatable estimates of the extent of ITS in birds. In five species, the amount of ITS was substantial, ranging from 15% to 40% of total telomeric sequence DNA. In addition, the amount of ITS can vary significantly among individuals within a species. Including ITSs in telomere length calculations always underestimated telomere length because most ITSs are shorter than most telomeres. The magnitude of that error varies with telomere length and is larger for longer telomeres. Estimating telomere length using methods that incorporate ITSs, such as Southern blot TRF and quantitative PCR analyses reduces an investigator's power to detect difference in telomere dynamics between individuals or over time within an individual.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Contrasting patterns of divergence at the regulatory and sequence level in European Daphnia galeata natural populations

Understanding the genetic basis of local adaptation has long been the focus of evolutionary biology. Recently there has been increased interest in deciphering the evolutionary role of Daphnia's plasticity and the molecular mechanisms of local adaptation. Using transcriptome data, we assessed the differences in gene expression profiles and sequences within and between four European Daphnia galeata populations. To distinguish neutral from adaptive differentiation, we corrected for phylogenetic differentiation of Daphnia populations. We also applied a "transcriptome scan" approach to investigate the role of natural selection in shaping divergent expression profiles among populations. Furthermore, a SNP analysis allowed inferring population structure and the distribution of genetic variation. Using sequence information, the transcripts were annotated using a comparative genomics approach. In total, ~33% of 32903 transcripts were differentially expressed between populations. Among 10280 differentially expressed transcripts, 5209 transcripts deviated from neutral expectations and were likely involved in local adaptation. The population divergence at the sequence level was higher than at the gene expression level by several orders of magnitude and revealed very distinct clusters according to population origin. Our analysis revealed the respective roles of genetic drift and selection in the four Daphnia populations. This study is a first attempt to understand the genetic background of adaptation to environmental changes in a key species of aquatic ecosystems in absence of any laboratory induced stressor.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Dissecting molecular evolution in the highly diverse plant clade Caryophyllales using transcriptome sequencing

Many phylogenomic studies based on transcriptomes have been limited to "single-copy" genes due to methodological challenges in homology and orthology inferences. Only a relatively small number of studies have explored analyses beyond reconstructing species relationships. We sampled 69 transcriptomes in the hyperdiverse plant clade Caryophyllales and 27 outgroups from annotated genomes across eudicots. Using a combined similarity- and phylogenetic tree-based approach, we recovered 10,960 homolog groups, where each was represented by at least eight ingroup taxa. By decomposing these homolog trees, and taking gene duplications into account, we obtained 17,273 ortholog groups, where each was represented by at least ten ingroup taxa. We reconstructed the species phylogeny using a 1,122-gene data set with a gene occupancy of 92.1%. From the homolog trees, we found that both synonymous and nonsynonymous substitution rates in herbaceous lineages are up to three times as fast as in their woody relatives. This is the first time such a pattern has been shown across thousands of nuclear genes with dense taxon sampling. We also pinpointed regions of the Caryophyllales tree that were characterized by relatively high frequencies of gene duplication, including three previously unrecognized whole-genome duplications. By further combining information from homolog tree topology and synonymous distance between paralog pairs, phylogenetic locations for 13 putative genome duplication events were identified. Genes that experienced the greatest gene family expansion were concentrated among those involved in signal transduction and oxidoreduction, including a cytochrome P450 gene that encodes a key enzyme in the betalain synthesis pathway. Our approach demonstrates a new approach for functional phylogenomic analysis in nonmodel species that is based on homolog groups in addition to inferred ortholog groups.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Insights into the genetic relationships and breeding patterns of the African tea germplasm based on nSSR markers and cpDNA sequences

Africa is one of the key centers of global tea production. Understanding the genetic diversity and relationships of cultivars of African tea is important for future targeted breeding efforts for new crop cultivars, specialty tea processing, and to guide germplasm conservation efforts. Despite the economic importance of tea in Africa, no research work has been done so far on its genetic diversity at a continental scale. Twenty-three nSSRs and three plastid DNA regions were used to investigate the genetic diversity, relationships, and breeding patterns of tea accessions collected from eight countries of Africa. A total of 280 African tea accessions generated 297 alleles with a mean of 12.91 alleles per locus and a genetic diversity (HS) estimate of 0.652. A STRUCTURE analysis suggested two main genetic groups of African tea accessions which corresponded well with the two tea types Camellia sinensis var. sinensis and C. sinensis var. assamica, respectively, as well as an admixed "mosaic" group whose individuals were defined as hybrids of F2 and BC generation with a high proportion of C. sinensis var. assamica being maternal parents. Accessions known to be C. sinensis var. assamica further separated into two groups representing the two major tea breeding centers corresponding to southern Africa (Tea Research Foundation of Central Africa, TRFCA), and East Africa (Tea Research Foundation of Kenya, TRFK). Tea accessions were shared among countries. African tea has relatively lower genetic diversity. C. sinensis var. assamica is the main tea type under cultivation and contributes more in tea breeding improvements in Africa. International germplasm exchange and movement among countries within Africa was confirmed. The clustering into two main breeding centers, TRFCA, and TRFK, suggested that some traits of C. sinensis var. assamica and their associated genes possibly underwent selection during geographic differentiation or local breeding preferences. This study represents the first step toward effective utilization of differently inherited molecular markers for exploring the breeding status of African tea. The findings here will be important for planning the exploration, utilization, and conservation of tea germplasm for future breeding efforts in Africa.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genotyping-by-sequencing of genome-wide microsatellite loci reveals fine-scale harvest composition in a coastal Atlantic salmon fishery

Individual assignment and genetic mixture analysis are commonly utilized in contemporary wildlife and fisheries management. Although microsatellite loci provide unparalleled numbers of alleles per locus, their use in assignment applications is increasingly limited. However, next-generation sequencing, in conjunction with novel bioinformatic tools allows large numbers of microsatellite loci to be simultaneously genotyped, presenting new opportunities for individual assignment and genetic mixture analysis. Here we scanned the published Atlantic salmon genome to identify 706 microsatellite loci, from which we developed a final panel of 101 microsatellites distributed across the genome (average 3.4 loci per chromosome). Using samples from 35 Atlantic salmon populations (n=1485 individuals) from coastal Labrador, Canada, a region characterized by low levels of differentiation in this species, this panel identified 844 alleles (average of 8.4 alleles per locus). Simulation-based evaluations of assignment and mixture identification accuracy revealed unprecedented resolution, clearly identifying 26 rivers or groups of rivers spanning 500 km of coastline. This baseline was used to examine the stock composition of 696 individuals harvested in the Labrador Atlantic salmon fishery and revealed that coastal fisheries largely targeted regional groups (<300km). This work suggests that the development and application of large sequenced microsatellite panels presents great potential for stock resolution in Atlantic salmon and more broadly in other exploited anadromous and marine species.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genotyping HapSTR loci: phase determination from direct sequencing of PCR products

HapSTRs combine information from a microsatellite (or simple tandem repeat, STR) with one or more single nucleotide polymorphisms (SNPs) in the DNA sequence immediately flanking the STR. These loci may offer increased power for the estimation of demographic parameters, but also present some challenges for data collection and analysis. We describe a process for inferring HapSTR alleles, including the flanking haplotypes, STR alleles, and their phase relative to each other, directly from DNA sequence electropherograms of PCR products from heterozygous individuals. Our approach eliminates the need for more costly and time-consuming processes such as cloning or acrylamide gel electrophoresis to separate alleles prior to sequencing.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Genome-wide SNPs resolve a key conflict between sequence and allozyme data to confirm another threatened candidate species of river blackfishes (Teleostei: Percichthyidae: Gadopsis)

Conflicting results from different molecular datasets have long confounded our ability to characterise species boundaries. Here we use genome-wide SNP data and an expanded allozyme dataset to resolve conflicting systematic hypotheses on an enigmatic group of fishes (Gadopsis, river blackfishes, Percichthyidae) restricted to southeastern Australia. Previous work based on three sets of molecular markers: mtDNA, nuclear intron DNA and 51 allozyme loci was unable to clearly resolve the status of a putative fifth candidate species (SWV) within Gadopsis marmoratus. Resolving the taxonomic status of candidate species SWV is particularly critical as based on IUCN criteria this taxon would be considered Critically Endangered. After all filtering steps we retained a subset of 10,862 putatively unlinked SNP loci for population genetic and phylogenomic analyses. Analyses of SNP loci based on maximum likelihood, fastSTRUCTURE and DAPC were all consistent with the previous and updated allozyme results supporting the validity of the candidate Gadopsis species SWV. Immediate conservation actions should focus on preventing take by anglers, protection of water resources to sustain perennial reaches and drought refuge pools, and aquatic and riparian habitat protection and improvement. In addition, a formal morphological taxonomic review of the genus Gadopsis is urgently required.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Development of a genotype-by-sequencing immunogenetic assay as exemplified by screening for variation in red fox with and without endemic rabies exposure

Pathogens are recognized as major drivers of local adaptation in wildlife systems. By determining which gene variants are favored in local interactions among populations with and without disease, spatially explicit adaptive responses to pathogens can be elucidated. Much of our current understanding of host responses to disease comes from a small number of genes associated with an immune response. High-throughput sequencing (HTS) technologies, such as genotype-by-sequencing (GBS), facilitate expanded explorations of genomic variation among populations. Hybridization-based GBS techniques can be leveraged in systems not well characterized for specific variants associated with disease outcome to "capture" specific genes and regulatory regions known to influence expression and disease outcome. We developed a multiplexed, sequence capture assay for red foxes to simultaneously assess ~300-kbp of genomic sequence from 116 adaptive, intrinsic, and innate immunity genes of predicted adaptive significance and their putative upstream regulatory regions along with 23 neutral microsatellite regions to control for demographic effects. The assay was applied to 45 fox DNA samples from Alaska, where three arctic rabies strains are geographically restricted and endemic to coastal tundra regions, yet absent from the boreal interior. The assay provided 61.5% on-target enrichment with relatively even sequence coverage across all targeted loci and samples (mean = 50×), which allowed us to elucidate genetic variation across introns, exons, and potential regulatory regions (4,819 SNPs). Challenges remained in accurately describing microsatellite variation using this technique; however, longer-read HTS technologies should overcome these issues. We used these data to conduct preliminary analyses and detected genetic structure in a subset of red fox immune-related genes between regions with and without endemic arctic rabies. This assay provides a template to assess immunogenetic variation in wildlife disease systems.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Seascape genetics along environmental gradients in the Arabian Peninsula: insights from ddRAD sequencing of anemonefishes

Understanding the processes that shape patterns of genetic structure across space is a central aim of landscape genetics. However, it remains unclear how geographic features and environmental variables shape gene flow, particularly for marine species in large complex seascapes. Here, we evaluated the genomic composition of the two-band anemonefish Amphiprion bicinctus across its entire geographic range in the Red Sea and Gulf of Aden, as well as its close relative, Amphiprion omanensis endemic to the southern coast of Oman. Both the Red Sea and the Arabian Sea are complex and environmentally heterogeneous marine systems that provide an ideal scenario to address these questions. Our findings confirm the presence of two genetic clusters previously reported for A. bicinctus in the Red Sea. Genetic structure analyses suggest a complex seascape configuration, with evidence of both Isolation by Distance (IBD) and Isolation by Environment (IBE). In addition to IBD and IBE, genetic structure among sites was best explained when two barriers to gene flow were also accounted for. One of these coincides with a strong oligotrophic-eutrophic gradient at around 16-20˚N in the Red Sea. The other agrees with an historical bathymetric barrier at the straight of Bab al Mandab. Finally, these data support the presence of inter-specific hybrids at an intermediate suture zone at Socotra and indicate complex patterns of genomic admixture in the Gulf of Aden with evidence of introgression between species. Our findings highlight the power of recent genomic approaches to resolve subtle patterns of gene flow in marine seascapes.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Do cryptic species matter in macroecology? Sequencing European groundwater crustaceans yields smaller ranges but does not challenge biodiversity determinants

Ecologists increasingly rely on molecular delimitation methods (MMs) to identify species boundaries, thereby potentially increasing the number of putative species because of the presence of morphologically cryptic species. It has been argued that cryptic species could challenge our understanding of what determine large-scale biodiversity patterns which have traditionally been documented from morphology alone. Here, we used morphology and three MMs to derive four different sets of putative species among the European groundwater crustaceans. Then, we used regression models to compare the relative importance of spatial heterogeneity, productivity and historical climates, in shaping species richness and range size patterns across sets of putative species. We tested three predictions. First, MMs would yield many more putative species than morphology because groundwater is a constraining environment allowing little morphological changes. Second, for species richness, MMs would increase the importance of spatial heterogeneity because cryptic species are more likely along physical barriers separating ecologically similar regions than along resource gradients promoting ecologically-based divergent selection. Third, for range size, MMs would increase the importance of historical climates because of reduced and asymmetrical fragmentation of large morphological species ranges at northern latitudes. MMs yielded twice more putative species than morphology and decreased by 10-fold the average species range size. Yet, MMs strengthened the mid-latitude ridge of high species richness and the Rapoport effect of increasing range size at higher latitudes. Species richness predictors did not vary between morphology and MMs but the latter increased the proportion of variance in range size explained by historical climates. These findings demonstrate that our knowledge of groundwater biodiversity determinants is robust to overlooked cryptic species because the latter are homogeneously distributed along environmental gradients. Yet, our findings call for incorporating multiple species delimitation methods into the analysis of large-scale biodiversity patterns across a range of taxa and ecosystems.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Development of diagnostic microsatellite markers from whole-genome sequences of Ammodramus sparrows for assessing admixture in a hybrid zone

Studies of hybridization and introgression and, in particular, the identification of admixed individuals in natural populations benefit from the use of diagnostic genetic markers that reliably differentiate pure species from each other and their hybrid forms. Such diagnostic markers are often infrequent in the genomes of closely related species, and genomewide data facilitate their discovery. We used whole-genome data from Illumina HiSeqS2000 sequencing of two recently diverged (600,000 years) and hybridizing, avian, sister species, the Saltmarsh (Ammodramus caudacutus) and Nelson's (A. nelsoni) Sparrow, to develop a suite of diagnostic markers for high-resolution identification of pure and admixed individuals. We compared the microsatellite repeat regions identified in the genomes of the two species and selected a subset of 37 loci that differed between the species in repeat number. We screened these loci on 12 pure individuals of each species and report on the 34 that successfully amplified. From these, we developed a panel of the 12 most diagnostic loci, which we evaluated on 96 individuals, including individuals from both allopatric populations and sympatric individuals from the hybrid zone. Using simulations, we evaluated the power of the marker panel for accurate assignments of individuals to their appropriate pure species and hybrid genotypic classes (F1, F2, and backcrosses). The markers proved highly informative for species discrimination and had high accuracy for classifying admixed individuals into their genotypic classes. These markers will aid future investigations of introgressive hybridization in this system and aid conservation efforts aimed at monitoring and preserving pure species. Our approach is transferable to other study systems consisting of closely related and incipient species.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Parallel tagged next-generation sequencing on pooled samples – a new approach for population genetics in ecology and conservation

Next-generation sequencing (NGS) on pooled samples has already been broadly applied in human medical diagnostics and plant and animal breeding. However, thus far it has been only sparingly employed in ecology and conservation, where it may serve as a useful diagnostic tool for rapid assessment of species genetic diversity and structure at the population level. Here we undertake a comprehensive evaluation of the accuracy, practicality and limitations of parallel tagged amplicon NGS on pooled population samples for estimating species population diversity and structure. We obtained 16S and Cyt b data from 20 populations of Leiopelma hochstetteri, a frog species of conservation concern in New Zealand, using two approaches – parallel tagged NGS on pooled population samples and individual Sanger sequenced samples. Data from each approach were then used to estimate two standard population genetic parameters, nucleotide diversity (π) and population differentiation (FST), that enable population genetic inference in a species conservation context. We found a positive correlation between our two approaches for population genetic estimates, showing that the pooled population NGS approach is a reliable, rapid and appropriate method for population genetic inference in an ecological and conservation context. Our experimental design also allowed us to identify both the strengths and weaknesses of the pooled population NGS approach and outline some guidelines and suggestions that might be considered when planning future projects.

opencc-zeroDec 2012View details →
dryad32/100

Data from: A public database of memory and naive B-cell receptor sequences

The vast diversity of B-cell receptors (BCR) and secreted antibodies enables the recognition of, and response to, a wide range of epitopes, but this diversity has also limited our understanding of humoral immunity. We present a public database of more than 37 million unique BCR sequences from three healthy adult donors that is many fold deeper than any existing resource, together with a set of online tools designed to facilitate the visualization and analysis of the annotated data. We estimate the clonal diversity of the naive and memory B-cell repertoires of healthy individuals, and provide a set of examples that illustrate the utility of the database, including several views of the basic properties of immunoglobulin heavy chain sequences, such as rearrangement length, subunit usage, and somatic hypermutation positions and dynamics.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Biodiversity assessment using next-generation sequencing: comparison of phylogenetic and functional diversity between Nebraska grasslands

Global biodiversity is declining rapidly as a consequence of anthropogenic changes to the environment. Traditional diversity indices such as species richness have been used to assess biodiversity, but recent arguments call for a more comprehensive assessment that includes both phylogenetic and functional diversity (PD and FD, respectively). Many PD metrics have been developed, but few empirical studies have compared metrics across sites with the goal of understanding their application to characterizing biodiversity. In this study, 17 PD metrics, four traditional diversity indices, and one measure of FD were calculated and compared between two Nebraska grasslands. PD metrics were calculated from robust phylogenies estimated from next-generation sequencing data of 45 species. Traditional indices were calculated using species abundance data, and FD was quantified by measuring the phylogenetic signal, K, of specific leaf area (SLA). Results showed that PD metrics and traditional indices were not always correlated, and various PD metrics characterized biodiversity differently. In addition, phylogenies estimated from >80 genes were more robust than single- or dual-gene phylogenies resulting in more reliable PD metrics. K of SLA indicated random trait assembly in all sites. Results suggested that metrics that identify phylogenetic structure and relatedness can provide information to conservation planners about the ability of a community to persist in an unpredictable future. A combination of these results with those of future investigations applying PD and FD metrics to varying communities will support concrete recommendations to conservation planners about how to incorporate these metrics into the selection of priority regions.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Utility of pooled sequencing for association mapping in non-model organisms

High density genome-wide sequencing increases the likelihood of discovering genes of major effect and genomic structural variation in organisms. While there is an increasing availability of reference genomes across broad taxa, the greatest limitation to whole-genome sequencing of multiple individuals continues to be the costs associated with sequencing. To alleviate excessive costs, pooling multiple individuals with similar phenotypes and sequencing the homogenized DNA (Pool-Seq) can achieve high genome coverage, but at the loss of individual genotypes. Although Pool-Seq has been an effective method for association mapping in model organisms, it has not been frequently utilized in natural populations. To extend bioinformatic tools for rapid implementation of Pool-Seq data in non-model organisms, we developed a pipeline called PoolParty and illustrate its effectiveness in genetic association mapping. Alignment expectations based on five pooled Chinook salmon (Oncorhynchus tshawytscha) libraries showed that approximately 48% genome coverage per library could be achieved with reasonable sequencing effort. We additionally examined male and female O. tshawytscha libraries to illustrate how Pool-Seq techniques can successfully map known genes associated with functional differences among sexes such as growth hormone 2. Finally, we compared pools of individuals of different spawning ages for each sex to discover novel genes involved with age at maturity in O. tshawytscha such as opsin4 and transmembrane protein19. While not appropriate for every system, Pool-Seq data processed by the PoolParty pipeline is a practical method for identifying genes of major effect in non-model organisms when high genome coverage is necessary and cost is a limiting factor.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Pleistocene climate change and phylogeographic structure of the Gymnocarpos przewalskii (Caryophyllaceae) in the northwest China: Evidence from plastid DNA, ITS sequences, and Microsatellite

Northwestern China has a wealth of endemic species, which has been hypothesized to be affected by the complex paleoclimatic and paleogeographic history during Quaternary. In this paper, we used Gymnocarpos przewalskii as a model to address the evolutionary history and current population genetic structure of species in northwestern China. We employed two chloroplast DNA fragments (rps16 and psbB‐psbI), one nuclear DNA fragment (ITS), and simple sequence repeat (SSRs) to investigate the spatial genetic pattern of G. przewalskii. High genetic diversity (cpDNA: hS = 0.330, hT = 0.866; ITS: hS = 0.458, hT = 0.872) was identified in almost all populations, and most of the population have private haplotypes. Moreover, multimodal mismatch distributions were observed and estimates of Tajima's D and Fu's FS tests did not identify significantly departures from neutrality, indicating that recent expansion of G. przewalskii was rejected. Thus, we inferred that G. przewalskii survived generally in northwestern China during the Pleistocene. All data together support the genotypes of G. przewalskii into three groups, consistent with their respective geographical distributions in the western regions—Tarim Basin, the central regions—Hami Basin and Hexi Corridor, and the eastern regions—Alxa Desert and Wulate Prairie. Divergence among most lineages of G. przewalskii occurred in the Pleistocene, and the range of potential distributions is associated with glacial cycles. We concluded that climate oscillation during Pleistocene significantly affected the distribution of the species.

opencc-zeroDec 2018View details →
dryad32/100

Data from: DNA sequence variation among conspecific accessions of the legume Coursetia caribaea reveals geographically localized clades here ranked as species

Coursetia caribaea is geographically and morphologically the most variable species in the genus Coursetia and in the tribe Robinieae (Leguminosae, Papilionoideae). Because of potentially undetected species, we assessed the phylogenetic relationships among the eight taxonomic varieties of C. caribaea. Sampling included nuclear ribosomal internal transcribed spacer sequences from 489 Robinieae accessions representing all varieties of C. caribaea and 38 of the 40 species of Coursetia, in addition to chloroplast trnD-trnT sequences from 186 accessions. Separate and combined phylogenetic analyses resolved a clade of conspecific accessions of the Bolivian C. caribaea var. astragalina as sister to the central Andean Coursetia grandiflora clade. Also distantly related to Coursetia caribaea var. caribaea accessions were those of the coastal Oaxacan C. caribaea var. pacifica, which formed the sister clade to accessions of the central Andean C. caribaea var. ochroleuca. The estimated mean ages of the stem clades for these three lineages, 11, 7.7, and 7.7 Ma, respectively, contrasted to the estimated mean ages of the corresponding crown clades of 0, 0, and 1.5 Ma. The contrasting stem and crown ages suggest that these taxa, appropriately ranked as species, Coursetia astragalina, Coursetia diversifolia, and Coursetia ochroleuca, each have persisted over evolutionary time frames as distinct geographically localized populations in seasonally dry tropical forests and woodlands.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Repetitive flanking sequences challenge SSR marker development: a case study in the lepidopteran Melanargia galathea

Microsatellite DNA families (MDF) are stretches of DNA that share similar or identical sequences beside nuclear simple-sequence repeat (nSSR) motifs, potentially causing problems during nSSR marker development. Primers positioned within MDFs can bind several times within the genome and might result in multiple banding patterns. It is therefore common practice to exclude MDF loci in the course of marker development. Here, we propose an approach to deal with multiple primer binding sites by purposefully positioning primers within the detected repetitive element. We developed a new protocol to determine the family type and the primer position in relation to MDFs using the software packages RepARK and RepeatMasker together with an in-house R script. We re-evaluated newly developed nSSR markers for the lepidopteran Marbled White (Melanargia galathea) and explored the implications of our results with regard to published data sets of the butterfly Ephydryas aurinia, the grasshopper Stethophyma grossum, the conifer Pinus cembra, and the crucifer Arabis alpina. For M. galathea, we show that it is not only possible to develop reliable nSSR markers for MDF loci, but even to benefit from their presence in some cases: We used one unlabeled primer, successfully binding within an MDF, for two different loci in a multiplex PCR, combining this family primer with uniquely binding and fluorescently labeled primers outside of MDFs, respectively. As MDFs are abundant in many taxa, we propose to consider these during nSSR marker development in taxa concerned. Our new approach might help in reducing the number of tested primers during nSSR marker development.

opencc-zeroDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record