Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data from: DNA barcoding meets molecular scatology: short mtDNA sequences for standardized species assignment of carnivore noninvasive samples
Although species assignment of scats is important to study carnivoran biology, there is still no standardized assay for the identification of carnivores worldwide, which would allow large-scale routine assessments and reliable cross-comparison of results. Here we evaluate the potential of two short mtDNA fragments (ATP6 [126 bp] and COI [187 bp]) to serve as standard markers for the Carnivora. Samples of 66 species were sequenced for one or both of these segments. Alignments were complemented with archival sequences, and analyzed with three approaches (tree-based, distance-based and character-based). Intraspecific genetic distances were generally lower than between-species distances, resulting in diagnosable clusters for 86% (ATP6) and and 85% (COI) of the species. Notable exceptions were recently diverged species, most of which could still be identified using diagnostic characters, uniqueness of haplotypes, or by reducing the geographic scope of the comparison. In silico comparative analyses were also performed with a 110-bp cytochrome b (cytb) segment, whose identification success was lower (70%), possibly due to the smaller number of informative sites and/or the influence of misidentified sequences obtained from GenBank. Finally, we performed case-studies with faecal samples, which supported the suitability of our two focal markers for poor-quality DNA, and allowed an assessment of prey-DNA co-amplification. No evidence of prey DNA contamination was found for ATP6, while some cases were observed for COI and subsequently eliminated by the design of more specific primers. Overall, our results indicate that these segments hold good potential as standard markers for accurate species-level identification in the Carnivora.
Data from: Testing models of speciation from genome sequences: divergence and asymmetric admixture in Island Southeast Asian Sus species during the Plio-Pleistocene climatic fluctuations
In many temperate regions, ice ages promoted range contractions into refugia resulting in divergence (and potentially speciation), while warmer periods led to range expansions and hybridization. However, the impact these climatic oscillations had in many parts of the tropics remains elusive. Here, we investigate this issue using genome sequences of three pig (Sus) species, two of which are found on islands of the Sunda-shelf shallow seas in Island Southeast Asia (ISEA). A previous study revealed signatures of inter-specific admixture between these Sus species (Frantz et al. (2013) Genome sequencing reveals fine scale diversification and reticulation history during speciation in Sus. Genome biology, 14, R107). However, the timing, directionality and extent of this admixture remain unknown. Here we use a likelihood based model comparison to more finely resolve this admixture history and test whether it was mediated by humans or occurred naturally. Our analyses suggest that inter-specific admixture between Sunda-shelf species was most likely asymmetric and occurred long before the arrival of humans in the region. More precisely, we show that these species diverged during the late Pliocene but around 23% of their genomes have been affected by admixture during the later Pleistocene climatic transition. In addition, we show that our method provides a significant improvement over D-statistics which are uninformative about the direction of admixture.
Data from: Genotyping-by-sequencing provides the discriminating power to investigate the subspecies of Daucus carota (Apiaceae)
Background: The majority of the subspecies of Daucus carota have not yet been discriminated clearly by various molecular or morphological methods and hence their phylogeny and classification remains unresolved. Recent studies using 94 nuclear orthologs and morphological characters, and studies employing other molecular approaches were unable to distinguish clearly many of the subspecies. Fertile intercrosses among traditionally recognized subspecies are well documented. We here explore the utility of single nucleotide polymorphisms (SNPs) generated by genotyping-by-sequencing (GBS) to serve as an effective molecular method to discriminate the subspecies of the D. carota complex. Results: We used GBS to obtain SNPs covering all nine Daucus carota chromosomes from 162 accessions of Daucus and two related genera. To study Daucus phylogeny, we scored a total of 10,814 or 38,920 SNPs with a maximum of 10 or 30 % missing data, respectively. To investigate the subspecies of D. carota, we employed two data sets including 150 accessions: (i) rate of missing data 10 % with a total of 18,565 SNPs, and (ii) rate of missing data 30 %, totaling 43,713 SNPs. Consistent with prior results, the topology of both data sets separated species with 2n = 18 chromosome from all other species. Our results place all cultivated carrots (D. carota subsp. sativus) in a single clade. The wild members of D. carota from central Asia were on a clade with eastern members of subsp. sativus. The other subspecies of D. carota were in four clades associated with geographic groups: (1) the Balkan Peninsula and the Middle East, (2) North America and Europe, (3) North Africa exclusive of Morocco, and (4) the Iberian Peninsula and Morocco. Daucus carota subsp. maximus was discriminated, but neither it, nor subsp. gummifer (defined in a broad sense) are monophyletic. Conclusions: Our study suggests that (1) the morphotypes identified as D. carota subspecies gummifer (as currently broadly circumscribed), all confined to areas near the Atlantic Ocean and the western Mediterranean Sea, have separate origins from sympatric members of other subspecies of D. carota, (2) D. carota subsp. maximus, on two clades with some accessions of subsp. carota, can be distinguished from each other but only with poor morphological support, (3) D. carota subsp. capillifolius, well distinguished morphologically, is an apospecies relative to North African populations of D. carota subsp. carota, (4) the eastern cultivated carrots have origins closer to wild carrots from central Asia than to western cultivated carrots, and (5) large SNP data sets are suitable for species-level phylogenetic studies in Daucus.
Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning
Reconstructing the phylogenetic relationships between species is one of the most formidable tasks in evolutionary biology. Multiple methods exist to reconstruct phylogenetic trees, each with their own strengths and weaknesses. Both simulation and empirical studies have identified several "zones" of parameter space where accuracy of some methods can plummet, even for four-taxon trees. Further, some methods can have undesirable statistical properties such as statistical inconsistency and/or the tendency to be positively misleading (i.e. assert strong support for the incorrect tree topology). Recently, deep learning techniques have made inroads on a number of both new and longstanding problems in biological research. Here we designed a deep convolutional neural network (CNN) to infer quartet topologies from multiple sequence alignments. This CNN can readily be trained to make inferences using both gapped and ungapped data. We show that our approach is highly accurate on simulated data, often outperforming traditional methods, and is remarkably robust to bias-inducing regions of parameter space such as the Felsenstein zone and the Farris zone. We also demonstrate that the confidence scores produced by our CNN can more accurately assess support for the chosen topology than bootstrap and posterior probability scores from traditional methods. While numerous practical challenges remain, these findings suggest that deep learning approaches such as ours have the potential to produce more accurate phylogenetic inferences.
Data from: Degenerate adaptor sequences for detecting PCR duplicates in reduced representation sequencing data improve genotype calling accuracy
RAD-tag is a powerful tool for high-throughput genotyping. It relies on PCR amplification of the starting material, following enzymatic digestion and sequencing adaptor ligation. Amplification introduces duplicate reads into the data, which arise from the same template molecule and are statistically nonindependent, potentially introducing errors into genotype calling. In shotgun sequencing, data duplicates are removed by filtering reads starting at the same position in the alignment. However, restriction enzymes target specific locations within the genome, causing reads to start in the same place, and making it difficult to estimate the extent of PCR duplication. Here, we introduce a slight change to the Illumina sequencing adaptor chemistry, appending a unique four-base tag to the first index read, which allows duplicate discrimination in aligned data. This approach was validated on the Illumina MiSeq platform, using double-digest libraries of ants (Wasmannia auropunctata) and yeast (Saccharomyces cerevisiae) with known genotypes, producing modest though statistically significant gains in the odds of calling a genotype accurately. More importantly, removing duplicates also corrected for strong sample-to-sample variability of genotype calling accuracy seen in the ant samples. For libraries prepared from low-input degraded museum bird samples (Mixornis gularis), which had low complexity, having been generated from relatively few starting molecules, adaptor tags show that virtually all of the genotypes were called with inflated confidence as a result of PCR duplicates. Quantification of library complexity by adaptor tagging does not significantly increase the difficulty of the overall workflow or its cost, but corrects for differences in quality between samples and permits analysis of low-input material.
Data from: Chlamydomonas genome resource for laboratory strains reveals a mosaic of sequence variation, identifies true strain histories, and enables strain-specific studies
Chlamydomonas reinhardtii is a widely used reference organism in studies of photosynthesis, cilia, and biofuels. Most research in this field uses a few dozen standard laboratory strains that are reported to share a common ancestry, but exhibit substantial phenotypic differences. In order to facilitate ongoing Chlamydomonas research and explain the phenotypic variation, we mapped the genetic diversity within these strains using whole-genome resequencing. We identified 524,640 single nucleotide variants and 4812 structural variants among 39 commonly used laboratory strains. Nearly all (98.2%) of the total observed genetic diversity was attributable to the presence of two, previously unrecognized, alternate haplotypes that are distributed in a mosaic pattern among the extant laboratory strains. We propose that these two haplotypes are the remnants of an ancestral cross between two strains with ∼2% relative divergence. These haplotype patterns create a fingerprint for each strain that facilitates the positive identification of that strain and reveals its relatedness to other strains. The presence of these alternate haplotype regions affects phenotype scoring and gene expression measurements. Here, we present a rich set of genetic differences as a community resource to allow researchers to more accurately conduct and interpret their experiments with Chlamydomonas.
Data from: Defining the alloreactive T cell repertoire using high-throughput sequencing of mixed lymphocyte reaction culture
The cellular immune response is the most important mediator of allograft rejection and is a major barrier to transplant tolerance. Delineation of the depth and breadth of the alloreactive T cell repertoire and subsequent application of the technology to the clinic may improve patient outcomes. As a first step toward this, we have used MLR and high-throughput sequencing to characterize the alloreactive T cell repertoire in healthy adults at baseline and 3 months later. Our results demonstrate that thousands of T cell clones proliferate in MLR, and that the alloreactive repertoire is dominated by relatively high-abundance T cell clones. This clonal make up is consistently reproducible across replicates and across a span of three months. These results indicate that our technology is sensitive and that the alloreactive TCR repertoire is broad and stable over time. We anticipate that application of this approach to track donor-reactive clones may positively impact clinical management of transplant patients.
Data from: DNA barcodes from century-old type specimens using next generation sequencing
Type specimens have high scientific importance because they provide the only certain connection between the application of a Linnean name and a physical specimen. Many other individuals may have been identified as a particular species, but their linkage to the taxon concept is inferential. Because type specimens are often more than a century old and have experienced conditions unfavorable for DNA preservation, success in sequence recovery has been uncertain. The present study addresses this challenge by employing next generation sequencing (NGS) to recover sequences for the barcode region of the cytochrome c oxidase 1 gene from small amounts of template DNA. DNA quality was first screened in more than 1800 century-old type specimens of Lepidoptera by attempting to recover 164bp and 94bp reads via Sanger sequencing. This analysis permitted the assignment of each specimen to one of three DNA quality categories – high (164bp sequence), medium (94bp sequence), or low (no sequence). Ten specimens from each category were subsequently analyzed via a PCR-based NGS protocol requiring very little template DNA. It recovered sequence information from all specimens with average read lengths ranging from 458bp to 610bp for the three DNA categories. By sequencing ten specimens in each NGS run, costs were similar to Sanger analysis. Future increases in the number of specimens processed in each run promise substantial reductions in cost, making it possible to anticipate a future where barcode sequences are available from most type specimens.
Data from: De novo sequencing and variant calling with nanopores using PoreSeq
The accuracy of sequencing single DNA molecules with nanopores is continually improving, but de novo genome sequencing and assembly using only nanopore data remain challenging. Here we describe PoreSeq, an algorithm that identifies and corrects errors in nanopore sequencing data and improves the accuracy of de novo genome assembly with increasing coverage depth. The approach relies on modeling the possible sources of uncertainty that occur as DNA transits through the nanopore and finds the sequence that best explains multiple reads of the same region. PoreSeq increases nanopore sequencing read accuracy of M13 bacteriophage DNA from 85% to 99% at 100× coverage. We also use the algorithm to assemble Escherichia coli with 30× coverage and the λ genome at a range of coverages from 3× to 50×. Additionally, we classify sequence variants at an order of magnitude lower coverage than is possible with existing methods.
Data from: Cross-species transferability of SSR loci developed from transcriptome sequencing in lodgepole pine
With the advent of next generation sequencing technologies, transcriptome level sequence collections are arising as prominent resources for the discovery of gene-based molecular markers. In a previous study more than 15 000 simple sequence repeats (SSRs) in expressed sequence tag (EST) sequences resulting from 454 pyrosequencing of Pinus contorta cDNA were identified. From these we developed PCR primers for approximately 4000 candidate SSRs. Here, we tested 184 of these SSRs for successful amplification across P. contorta and eight other pine species and examined patterns of polymorphism and allelic variability for a subset of these SSRs. Cross-species transferability was high, with high percentages of loci producing PCR products in all species tested. In addition, 50% of the loci we screened across panels of individuals from three of these species were polymorphic and allelically diverse. We examined levels of diversity in a subset of these SSRs by collecting genotypic data across several populations of Pinus ponderosa in northern Wyoming. Our results indicate the utility of mining pyrosequenced EST collections for gene-based SSRs and provide a source of molecular markers that should bolster evolutionary genetic investigations across the genus Pinus.
Data from: Congruent deep relationships in the grape family (Vitaceae) based on sequences of chloroplast genomes and mitochondrial genes via genome skimming
Vitaceae is well-known for having one of the most economically important fruits, i.e., the grape (Vitis vinifera). The deep phylogeny of the grape family was not resolved until a recent phylogenomic analysis of 417 nuclear genes from transcriptome data. However, it has been reported extensively that topologies based on nuclear and organellar genes may be incongruent due to differences in their evolutionary histories. Therefore, it is important to reconstruct a backbone phylogeny of the grape family using plastomes and mitochondrial genes. In this study, next-generation sequencing data sets of 27 species were obtained using genome skimming with total DNAs from silica-gel preserved tissue samples on an Illumina HiSeq 2500 instrument. Plastomes were assembled using the combination of de novo and reference genome (of V. vinifera) methods. Sixteen mitochondrial genes were also obtained via genome skimming using the reference genome of V. vinifera. Extensive phylogenetic analyses were performed using maximum likelihood and Bayesian methods. The topology based on either plastome data or mitochondrial genes is congruent with the one using hundreds of nuclear genes, indicating that the grape family did not exhibit significant reticulation at the deep level. The results showcase the power of genome skimming in capturing extensive phylogenetic data: especially from chloroplast and mitochondrial DNAs.
Data from: In silico site-directed mutagenesis informs species-specific predictions of chemical susceptibility derived from the Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool
Chemical hazard assessment requires extrapolation of information from model organisms to all species of concern. The Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool was developed as a rapid, cost effective method to aid cross-species extrapolation of susceptibility to chemicals acting on specific protein targets through evaluation of protein structural similarities and differences. The greatest resolution for extrapolation of chemical susceptibility across species involves comparisons of individual amino acid residues at key positions involved in protein-chemical interactions. However, a lack of understanding of whether specific amino acid substitutions among species at key positions in proteins affect interaction with chemicals made manual interpretation of alignments time consuming and potentially inconsistent. Therefore, this study used in silico site-directed mutagenesis coupled with docking simulations of computational models for acetylcholinesterase (AChE) and ecdysone receptor (EcR) to investigate how specific amino acid substitutions impact protein-chemical interaction. This study found that computationally derived substitutions in identities of key amino acids caused no change in protein-chemical interaction if residues share the same side chain functional properties and have comparable molecular dimensions, while differences in these characteristics can change protein-chemical interaction. These findings were considered in the development of capabilities for automatically generated species-specific predictions of chemical susceptibility in SeqAPASS. These predictions for AChE and EcR were shown to agree with SeqAPASS predictions comparing the primary sequence and functional domain sequence of proteins for more than 90 % of the investigated species, but also identified dramatic species-specific differences in chemical susceptibility that align with results from standard toxicity tests. These results provide a compelling line-of-evidence for use of SeqAPASS in deriving screening level, species-specific, susceptibility predictions across broad taxonomic groups for application to human and ecological hazard assessment.
Data from: Relationship between the sequencing and timing of vocal motor elements in birdsong
Accurate coordination of the sequencing and timing of motor gestures is important for the performance of complex and evolutionarily relevant behaviors. However, the degree to which motor sequencing and timing are related remains largely unknown. Birdsong is a communicative behavior that consists of discrete vocal motor elements ('syllables') that are sequenced and timed in a precise manner. To reveal the relationship between syllable sequencing and timing, we analyzed how variation in the probability of syllable transitions at branch points, nodes in song with variable sequencing across renditions, correlated with variation in the duration of silent gaps between syllable transitions ('gap durations') for adult Bengalese finch song. We observed a significant negative relationship between transition probability and gap duration: more prevalent transitions were produced with shorter gap durations. We then assessed the degree to which long-term age-dependent changes and acute context-dependent changes to syllable sequencing and timing followed this inverse relationship. Age- but not context-dependent changes to syllable sequencing and timing were inversely related. On average, gap durations at branch points decreased with age, and the magnitude of this decrease was greater for transitions that increased in prevalence than for transitions that decreased in prevalence. In contrast, there was no systematic relationship between acute context-dependent changes to syllable sequencing and timing. Gap durations at branch points decreased when birds produced female-directed courtship song compared to when they produced undirected song, and the magnitude of this decrease was not related to the direction and magnitude of changes to transition probabilities. These analyses suggest that neural mechanisms that regulate syllable sequencing could similarly control syllable timing but also highlight mechanisms that can independently regulate syllable sequencing and timing.
Genotyping by sequencing data of five legume tree species widespread in the rainforests of West and Central Africa
<p>Although today the forest cover is continuous in Central Africa this may have not always been the case, as the scarce fossil record in this region suggests that arid conditions might have significantly reduced tree density during the Ice Ages. Our aim was to investigate whether the dry ice-age periods left a genetic signature on tree species that can be used to infer the date of the past fragmentation of the rainforest. We sequenced reduced representation libraries of 182 samples representing five widespread Legume trees and seven outgroups. Phylogenetic analyses identified an early divergent lineage for all species in West Africa (Upper Guinea), and two clades in Central Africa: Lower Guinea-North and Lower Guinea-South. As the structure separating the Northern and Southern clades -congruent across species- cannot be explained by geographic barriers, we tested other hypotheses with demographic model testing using ∂a∂I. The best estimates indicate that the two clades split between the Upper Pliocene and the Pleistocene, a date compatible with forest fragmentation driven by ice-age climatic oscillations. Furthermore, we found remarkably older split dates for the shade-tolerant tree species with non-assisted seed dispersal than for light-demanding species with long-distance wind dispersal. Different recolonisation abilities after recurrent cycles of forest fragmentation seem to explain why species with long-distance dispersal show more recent genetic admixture between the two clades than species with limited seed dispersal. Despite their old history, our results depict the African rainforests as a dynamic biome where tree species have expanded relatively recently after the last glaciation.</p>
Data from: Transcriptome sequencing and microarray development for the woodrat (Neotoma spp.): custom genetic tools for exploring herbivore ecology
Massively parallel sequencing has enabled the creation of novel, in-depth genetic tools for nonmodel, ecologically important organisms. We present the de novo transcriptome sequencing, analysis and microarray development for a vertebrate herbivore, the woodrat (Neotoma spp.). This genus is of ecological and evolutionary interest, especially with respect to ingestion and hepatic metabolism of potentially toxic plant secondary compounds. We generated a liver transcriptome of the desert woodrat (Neotoma lepida) using the Roche 454 platform. The assembled contigs were well annotated using rodent references (99.7% annotation), and biotransformation function was reflected in the gene ontology. The transcriptome was used to develop a custom microarray (eArray, Agilent). We tested the microarray with three experiments: one across species with similar habitat (thus, dietary) niches, one across species with different habitat niches and one across populations within a species. The resulting one-colour arrays had high technical and biological quality. Probes designed from the woodrat transcriptome performed significantly better than functionally similar probes from the Norway rat (Rattus norvegicus). There were a multitude of expression differences across the woodrat treatments, many of which related to biotransformation processes and activities. The pattern and function of the differences indicate shared ecological pressures, and not merely phylogenetic distance, play an important role in shaping gene expression profiles of woodrat species and populations. The quality and functionality of the woodrat transcriptome and custom microarray suggest these tools will be valuable for expanding the scope of herbivore biology, as well as the exploration of conceptual topics in ecology.
Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.
Next-generation sequencing (NGS) technologies are increasingly applied in many organisms, including non-model organisms that are important for ecological and conservation purposes. Illumina and 454 sequencing are among the most used NGS technologies and have been shown to produce optimal results at reasonable costs when used together. Here, we describe the combined application of these two NGS technologies to characterize the transcriptome of a plant species of ecological and conservation relevance for which no genomic resource is available, Scabiosa columbaria. We obtained 528,557 reads from a 454 GS-FLX run and a total of 28,993,627 reads from two lanes of an Illumina GAII single run. After reads trimming, the de novo assembly of both types of reads produced 109,630 contigs. Both the contigs and the >75 bp remaining singletons were blasted against Uniprot/Swissprot database, resulting in 29,676 and 10,515 significant hits, respectively. Based on sequence similarity with known gene products, these sequences represent at least 12,516 unique genes, most of which are well covered by contig sequences. In addition, we identified 4,320 microsatellite loci, of which 856 had flanking sequences suitable for PCR primer design. We also identified 75,054 putative SNPs. This annotated sequence collection and the relative molecular markers represent a main genomic resource for S. columbaria which should contribute to future research in conservation and population biology studies. Our results demonstrate the utility of NGS technologies as starting point for the development of genomic tools in nonmodel but ecologically important species.
Data from: Genome-level homology and phylogeny of Vibrionaceae (Gammaproteobacteria: Vibrionales) with three new complete genome sequences
Background: Phylogenetic hypotheses based on complete genome data are presented for the Gammaproteobacteria family Vibrionaceae. Two taxon samplings are presented: one including all those taxa for which the genome sequences are complete in terms of arrangement (chromosomal location of fragments; 19 taxa) and one for which the genome sequences contain multiple contigs (44 taxa). Analyses are presented under the Maximum Parsimony and Maximum Likelihood optimality criteria for total evidence datasets, the two chromosomes separately, and individual analyses of locally collinear blocks. Three of the genomes included in the 44 taxon dataset, those of Vibrio gazogenes, Salinivibrio costicola, and Aliivibrio logei have been newly sequenced and their genome sequences are documented here. Results: Phylogenetic results for the 19-taxon datasets show similar levels of collinear subset of dataset incongruence as a previous study of 22 taxa from the sister family Shewanellaceae, while also echoing the strong phylogenetic performance of random subsets of data also shown in this study. Phylogenetic results for both the 19-taxon and 44-taxon datasets corroborate previous hypotheses about the placement of Photobacterium and Aliivibrio within Vibrionaceae and also highlight problems with how Photobacterium is delimited and indicate that it likely should be dissolved into Vibrio to produce a phylogenetic taxonomy. The 19-taxon and 44-taxon trees based on the large chromosome are congruent for the majority of taxa that are present in both datasets. Analyses of the 44-taxon sampling based on the second, small chromosome are quite different from those based on the large chromosome, which is not surprising given the dramatically divergent nature of the small chromosome and the difficulty in postulating primary homologies. Conclusions: The phylogenetic analyses presented here represent the most comprehensive genome-level phylogenetic analyses in terms of taxa and data. Based on the availability of genome data for many bacterial species on GenBank, many other bacterial groups would also be amenable to similar genome-scale phylogenetic analyses even when present in multiple contigs. The result that collinear subsets of data are incongruent with the concatenated dataset and with each other while random data subsets show very little incongruence echoes the result of previous work on Shewanellaceae. The 44-taxon phylogenetic analysis presented here thus represents the future of phylogenomic analyses in scope and complexity.
Data from: Regulation of transposable elements: interplay between TE-encoded regulatory sequences and host-specific trans-acting factors in Drosophila melanogaster
Transposable elements (TEs) are mobile genetic elements that can move around the genome, and their expression is one precondition for this mobility. Because the insertion of TEs in new genomic positions is largely deleterious, the molecular mechanisms for transcriptional suppression have been extensively studied. In contrast, very little is known about their primary transcriptional regulation. Here, we characterize the expression dynamics of TE families in Drosophila melanogaster across a broad temperature range (13–29°C). In 71% of the expressed TE families, the expression is modulated by temperature. We show that this temperature-dependent regulation is specific for TE families and strongly affected by the genetic background. We deduce that TEs carry family-specific regulatory sequences, which are targeted by host-specific trans-acting factors, such as transcription factors. Consistent with the widespread dominant inheritance of gene expression, we also find the prevailing dominance of TE family expression. We conclude that TE family expression across a range of temperatures is regulated by an interaction between TE family-specific regulatory elements and trans-acting factors of the host.
Data from: ASSET: analysis of sequences of synchronous events in massively parallel spike trains
With the ability to observe the activity from large numbers of neurons simultaneously using modern recording technologies, the chance to identify sub-networks involved in coordinated processing increases. Sequences of synchronous spike events (SSEs) constitute one type of such coordinated spiking that propagates activity in a temporally precise manner. The synfire chain was proposed as one potential model for such network processing. Previous work introduced a method for visualization of SSEs in massively parallel spike trains, based on an intersection matrix that contains in each entry the degree of overlap of active neurons in two corresponding time bins. Repeated SSEs are reflected in the matrix as diagonal structures of high overlap values. The method as such, however, leaves the task of identifying these diagonal structures to visual inspection rather than to a quantitative analysis. Here we present ASSET (Analysis of Sequences of Synchronous EvenTs), an improved, fully automated method which determines diagonal structures in the intersection matrix by a robust mathematical procedure. The method consists of a sequence of steps that i) assess which entries in the matrix potentially belong to a diagonal structure, ii) cluster these entries into individual diagonal structures and iii) determine the neurons composing the associated SSEs. We employ parallel point processes generated by stochastic simulations as test data to demonstrate the performance of the method under a wide range of realistic scenarios, including different types of non-stationarity of the spiking activity and different correlation structures. Finally, the ability of the method to discover SSEs is demonstrated on complex data from large network simulations with embedded synfire chains. Thus, ASSET represents an effective and efficient tool to analyze massively parallel spike data for temporal sequences of synchronous activity.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Multiple sequence aligners typically work by progressively aligning the most closely related sequences or group of sequences according to guide trees. In PNAS, Boyce et al. report that alignments reconstructed using simple chained trees (i.e., comb-like topologies) with random leaf assignment performed better in protein structure-based benchmarks than those reconstructed using phylogenies estimated from the data as guide trees. The authors state that this result could turn decades of research in the field on its head. In light of this statement, it is important to check immediately whether their result holds under evolutionary criteria: recovery of homologous sequence residues and inference of phylogenetic trees from the alignments. We have done this and the results are entirely opposed to Boyce et al.'s findings.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.