Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Detection of SNPs based on transcriptome sequencing in Norway spruce (Picea abies (L.) Karst)

A novel set of SNPs was derived from transcriptome data of ten Norway spruce (Picea abies) trees from the Bavarian Forest National Park in Germany (BaFoNP). SNPs were identified by mapping against a de-novo transcriptome assembly and against pre-mRNAs of predicted genes of the reference genome assembly. This resulted in 111,849 and 366,577 SNPs, respectively. Out of these, 311 were either randomly selected or chosen because of their pronounced divergence between sampling sites and genotyped in 218 trees with an Illumina Infinium HD iSelect BeadChip.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Random sequences are an abundant source of bioactive RNAs or peptides

It is generally assumed that new genes arise through duplication and/or recombination of existing genes. The probability that a new functional gene could arise out of random non-coding DNA is so far considered to be negligible, as it seems unlikely that such an RNA or protein sequence could have an initial function that influences the fitness of an organism. Here, we have tested this question systematically, by expressing clones with random sequences in Escherichia coli and subjecting them to competitive growth. Contrary to expectations, we find that random sequences with bioactivity are not rare. In our experiments we find that up to 25% of the evaluated clones enhance the growth rate of their cells and up to 52% inhibit growth. Testing of individual clones in competition assays confirms their activity and provides an indication that their activity could be exerted by either the transcribed RNA or the translated peptide. This suggests that transcribed and translated random parts of the genome could indeed have a high potential to become functional. The results also suggest that random sequences may become an effective new source of molecules for studying cellular functions, as well as for pharmacological activity screening.

opencc-zeroDec 2016View details →
dryad28/100

Data from: PCR-Free enrichment of mitochondrial DNA from human blood and cell lines for high quality next-generation DNA sequencing

Recent advances in sequencing technology allow for accurate detection of mitochondrial sequence variants, even those in low abundance at heteroplasmic sites. Considerable sequencing cost savings can be achieved by enriching samples for mitochondrial (relative to nuclear) DNA. Reduction in nuclear DNA (nDNA) content can also help to avoid false positive variants resulting from nuclear mitochondrial sequences (numts). We isolate intact mitochondrial organelles from both human cell lines and blood components using two separate methods: a magnetic bead binding protocol and differential centrifugation. DNA is extracted and further enriched for mitochondrial DNA (mtDNA) by an enzyme digest. Only 1 ng of the purified DNA is necessary for library preparation and next generation sequence (NGS) analysis. Enrichment methods are assessed and compared using mtDNA (versus nDNA) content as a metric, measured by using real-time quantitative PCR and NGS read analysis. Among the various strategies examined, the optimal is differential centrifugation isolation followed by exonuclease digest. This strategy yields >35% mtDNA reads in blood and cell lines, which corresponds to hundreds-fold enrichment over baseline. The strategy also avoids false variant calls that, as we show, can be induced by the long-range PCR approaches that are the current standard in enrichment procedures. This optimization procedure allows mtDNA enrichment for efficient and accurate massively parallel sequencing, enabling NGS from samples with small amounts of starting material. This will decrease costs by increasing the number of samples that may be multiplexed, ultimately facilitating efforts to better understand mitochondria-related diseases.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Sequence specificity despite intrinsic disorder: how a disease-associated Val/Met polymorphism rearranges tertiary interactions in a long disordered protein

The role of electrostatic interactions and mutations that change charge states in intrinsically disordered proteins (IDPs) is well-established, but many disease-associated mutations in IDPs are charge-neutral. The Val66Met single nucleotide polymorphism (SNP) in precursor brain-derived neurotrophic factor (BDNF) is one of the earliest SNPs to be associated with neuropsychiatric disorders, and the underlying molecular mechanism is unknown. Here we report on over 250 μ s of fully-atomistic, explicit solvent, temperature replica exchange molecular dynamics (MD) simulations of the 91 residue BDNF prodomain, for both the V66 and M66 sequence. The simulations were able to correctly reproduce the location of both local and non-local secondary changes due to the Val66Met mutation when compared with NMR spectroscopy. We find that the change in local structure is mediated via entropic and sequence specific effects. We developed a hierarchical sequence-based framework for analysis and conceptualization, which first identifies "blobs" of 5-15 residues representing local globular regions or linkers. We use this framework within a novel test for enrichment of higher-order (tertiary) structure in disordered proteins; the size and shape of each blob is extracted from MD simulation of the real protein (RP), and used to parameterize a self-avoiding heterogenous polymer (SAHP). The SAHP version of the BDNF prodomain suggested a protein segmented into three regions, with a central long, highly disordered polyampholyte linker separating two globular regions. This effective segmentation was also observed in full simulations of the RP, but the Val66Met substitution significantly increased interactions across the linker, as well as the number of participating residues. The Val66Met substitution replaces β -bridging between Val66 and Val94 (on either side of the linker) with specific side-chain interactions between Met66 and Met95.The protein backbone in the vicinity of Met95 is then free to form β -bridges with residues 31-41 near the N-terminus, which condenses the protein. A significant role for Met/Met interactions is consistent with previously-observed non-local effects of the Val66Met SNP, as well as established interactions between the Met66 sequence and a Met-rich receptor that initiates neuronal growth cone retraction.

opencc-zeroOct 2019View details →
dryad28/100

Data from: Complex models of sequence evolution require accurate estimators as exemplified with the invariable site plus Gamma model

The invariable site plus Γ model is widely used to model rate heterogeneity among alignment sites in maximum likelihood and Bayesian phylogenetic analyses. The proof that the invariable site plus continuous Γ model is identifiable (model parameters can be inferred correctly given enough data) has increased the creditability of its application to phylogeny reconstruction. However, most phylogenetic software implement the invariable site plus discrete Γ model, whose identifiability is likely but unproven. How well the parameters of the invariable site plus discrete Γ model are estimated is still disputed. Especially the correlation of the fraction of invariable sites with the fractions of sites with a slow evolutionary rate is discussed as being problematic. We show that optimization heuristics as implemented in frequently used phylogenetic software cannot always reliably estimate the shape parameter, the proportion of invariable sites and the tree length. Here, we propose an improved optimization heuristic that accurately estimates the three parameters. While research efforts mainly focus on tree search methods, our results signify the equal importance of verifying and developing effective estimation methods for complex models of sequence evolution.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Development of an Arabis alpina genomic contig sequence dataset and application to single nucleotide polymorphisms discovery

The alpine plant Arabis alpina is an emerging model in the ecological genomic field which is well-suited to identifying the genes involved in local adaptation in contrasted environmental conditions, a subject which remains poorly understood at molecular level. This paper presents the assembly of a pool of A. alpina genomic fragments using Next Generation Sequencing technologies. These contigs cover 172 Mb of the A. alpina genome (i.e. 50% of the genome) and were shown to contain sequences giving positive hits against 96% of the 458 CEGMA core genes (Core Eukaryotic Genes Mapping Approach), a set of highly conserved eukaryotic genes. Regions presenting high nucleic sequence identity with 77% of the close relative Arabidopsis thaliana's genes were found, with an unbiased distribution across the different functional categories of A. thaliana genes. This new resource was tested using a resequencing assay to identify polymorphic sites. Sixteen samples were successfully analyzed and 127,041 Single Nucleotide Polymorphisms identified. This contig dataset will contribute to improving understanding of the ecology of Arabis alpina, thus constituting an important resource for future ecological genomic studies.

opencc-zeroDec 2012View details →
dryad28/100

Data from: aTRAM - automated Target Restricted Assembly Method: a fast method for assembling loci across divergent taxa from next-generation sequencing data

Background: Assembling genes from next-generation sequencing data is not only time consuming but computationally difficult, particularly for taxa without a closely related reference genome. Assembling even a draft genome using de novo approaches can take days, even on a powerful computer, and these assemblies typically require data from a variety of genomic libraries. Here we describe software that will alleviate these issues by rapidly assembling genes from distantly related taxa using a single library of paired-end reads: aTRAM, automated Target Restricted Assembly Method. The aTRAM pipeline uses a reference sequence, BLAST, and an iterative approach to target and locally assemble the genes of interest. Results: Our results demonstrate that aTRAM rapidly assembles genes across distantly related taxa. In comparative tests with a closely related taxon, aTRAM assembled the same sequence as reference-based and de novo approaches taking on average < 1 min per gene. As a test case with divergent sequences, we assembled >1,000 genes from six taxa ranging from 25 – 110 million years divergent from the reference taxon. The gene recovery was between 97 – 99% from each taxon. Conclusions: aTRAM can quickly assemble genes across distantly-related taxa, obviating the need for draft genome assembly of all taxa of interest. Because aTRAM uses a targeted approach, loci can be assembled in minutes depending on the size of the target. Our results suggest that this software will be useful in rapidly assembling genes for phylogenomic projects covering a wide taxonomic range, as well as other applications. The software is freely available: http://www.github.com/juliema/aTRAM

opencc-zeroDec 2014View details →
dryad28/100

Data from: Survey sequencing reveals elevated DNA transposon activity, novel elements, and variation in repetitive landscapes among vesper bats

The repetitive landscapes of mammalian genomes typically display high Class I (retrotransposon) transposable element (TE) content, usually around half of the genome. In contrast, the Class II (DNA transposon) contribution is typically small (<3% in model mammals). Most mammalian genomes also exhibit a precipitous decline in Class II activity beginning roughly 40 million years ago (Ma). The first signs of more recently active mammalian Class II TEs were obtained from the little brown bat, Myotis lucifugus and are reflected by higher genome content (~5%). To aid in determining taxonomic limits and potential impacts of this elevated Class II activity, we performed 454 survey sequencing of a second Myotis species as well as four additional taxa within the family Vespertilionidae and an outgroup species from Phyllostomidae. Graph-based clustering methods were used to reconstruct the major repeat families present in each species and novel elements were identified in several taxa. Retrotransposons remained the dominant group with regard to overall genome mass. Elevated Class II TE composition (3-4%) was observed in all five vesper bats while less than 0.5% of the phyllostomid reads were identified as Class II derived. Differences in satellite DNA and Class I TE content are also described among vespertilionid taxa. These analyses present the first cohesive description of TE evolution across closely related mammals, revealing genome-scale differences in TE content within a single family.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Not all sequence tags are created equal: designing and validating sequence identification tags robust to indels

Ligating adapters with unique synthetic oligonucleotide sequences (sequence tags) onto individual DNA samples before massively parallel sequencing is a popular and efficient way to obtain sequence data from many individual samples. Tag sequences should be numerous and sufficiently different to ensure sequencing, replication, and oligonucleotide synthesis errors do not cause tags to be unrecoverable or confused. However, many design approaches only protect against substitution errors during sequencing and extant tag sets contain too few tag sequences. We developed an open-source software package to validate sequence tags for conformance to two distance metrics and design sequence tags robust to indel and substitution errors. We use this software package to evaluate several commercial and non-commercial sequence tag sets, design several large sets (maxcount=7,198) of edit metric sequence tags having different lengths and degrees of error correction, and integrate a subset of these edit metric tags to polymerase chain reaction (PCR) primers and sequencing adapters. We validate a subset of these edit metric tagged PCR primers and sequencing adapters by sequencing on several platforms and subsequent comparison to commercially available alternatives. We find that several commonly used sets of sequence tags or design methodologies used to produce sequence tags do not meet the minimum expectations of their underlying distance metric, and we find that PCR primers and sequencing adapters incorporating edit metric sequence tags designed by our software package perform as well as their commercial counterparts. We suggest that researchers evaluate sequence tags prior to use or evaluate tags that they have been using. The sequence tag sets we design improve on extant sets because they are large, valid across the set, and robust to the suite of substitution, insertion, and deletion errors affecting massively parallel sequencing workflows on all currently used platforms.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Large-scale genotyping of highly polymorphic loci by next generation sequencing: how to overcome the challenges to reliably genotype individuals?

Studying the different roles of adaptive genes is still a challenge in evolutionary ecology and requires reliable genotyping of large numbers of individuals. Next-generation sequencing (NGS) techniques enable such large-scale sequencing, but stringent data processing is required. Here, we develop an easy to use methodology to process amplicon-based NGS data and we apply this methodology to reliably genotype four major histocompatibility complex (MHC) loci belonging to MHC class I and II of Alpine marmots (Marmota marmota). Our post-processing methodology allowed us to increase the number of retained reads. The quality of genotype assignment was further assessed using three independent validation procedures. A total of 3069 high-quality MHC genotypes were obtained at four MHC loci for 863 Alpine marmots with a genotype assignment error rate estimated as 0.21%. The proposed methodology could be applied to any genetic system and any organism, except when extensive copy-number variation occurs (that is, genes with a variable number of copies in the genotype of an individual). Our results highlight the potential of amplicon-based NGS techniques combined with adequate post-processing to obtain the large-scale highly reliable genotypes needed to understand the evolution of highly polymorphic functional genes.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Assembly and comparative analysis of transposable elements from low coverage genomic sequence data in Asparagales

The research field of comparative genomics is moving from a focus on genes to a more holistic view including the repetitive complement. This study aimed to characterize relative proportions of the repetitive fraction of large, complex genomes in a non-model system. The monocotyledonous plant order Asparagales (onion, asparagus, agave) comprises some of the largest angiosperm genomes and represents variation in both genome size and structure (karyotype). Anonymous, low coverage, single-end Illumina data from eleven exemplar Asparagales taxa were assembled using a de novo method. Resulting contigs were annotated using a reference library of available monocot repetitive sequences. Mapping reads to contigs provided rough estimates of relative proportions of each type of transposon in the nuclear genome. The results were parsed into general repeat types and synthesized with genome size estimates and a phylogenetic context to describe the pattern of transposable element evolution among these lineages. The major finding is that while some lineages in Asparagales exhibit conservation in repeat proportions, there is generally wide variation in types and frequency of repeats. This approach is an appropriate first step in characterizing repeats in evolutionary lineages with a paucity of genomic resources.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Development of highly reliable in silico SNP resource and genotyping assay from exome capture and sequencing: an example from black spruce (Picea mariana)

Picea mariana is a widely distributed boreal conifer across Canada and the subject of advanced breeding programs for which population genomics and genomic selection approaches are being developed. Targeted sequencing was achieved after capturing P. mariana exome with probes designed from the sequenced transcriptome of Picea glauca, a distant relative. A high capture efficiency of 75.9% was reached although spruce has a complex and large genome including gene sequences interspersed by some long introns. The results confirmed the relevance of using probes from congeneric species to perform successfully interspecific exome capture in the genus Picea. A bioinformatics pipeline was developed including stringent criteria that helped detect a set of 97 075 highly reliable in silico SNPs. These SNPs were distributed across 14 909 genes. Part of an Infinium iSelect array was used to estimate the rate of true positives by validating 4267 of the predicted in silico SNPs by genotyping trees from P. mariana populations. The true positive rate was 96.2%, for in silico SNPs compared to a genotyping success rate of 96.7% for a set 1115 P. mariana control SNPs recycled from previous genotyping arrays. These results indicate the high success rate of the genotyping array and the relevance of the selection criteria used to delineate the new P. mariana in silico SNP resource. Furthermore, in silico SNPs were generally of medium to high frequency in natural populations, thus providing high informative value for future population genomics applications.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Identification of multiple QTL hotspots in sockeye salmon (Oncorhynchus nerka) using genotyping-by-sequencing and a dense linkage map

Understanding the genetic architecture of phenotypic traits can provide important information about the mechanisms and genomic regions involved in local adaptation and speciation. Here, we used genotyping-by-sequencing and a combination of previously published and newly generated data to construct sex-specific linkage maps for sockeye salmon (Oncorhynchus nerka). We then used the denser female linkage map to conduct quantitative trait locus (QTL) analysis for 4 phenotypic traits in 3 families. The female linkage map consisted of 6322 loci distributed across 29 linkage groups and was 4082 cM long, and the male map contained 2179 loci found on 28 linkage groups and was 2291 cM long. We found 26 QTL: 6 for thermotolerance, 5 for length, 9 for weight, and 6 for condition factor. QTL were distributed nonrandomly across the genome and were often found in hotspots containing multiple QTL for a variety of phenotypic traits. These hotspots may represent adaptively important regions and are excellent candidates for future research. Comparing our results with studies in other salmonids revealed several regions with overlapping QTL for the same phenotypic trait, indicating these regions may be adaptively important across multiple species. Altogether, our study demonstrates the utility of genomic data for investigating the genetic basis of important phenotypic traits. Additionally, the linkage map created here will enable future research on the genetic basis of phenotypic traits in salmon.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Genome assembly improvement and mapping convergently evolved skeletal traits in sticklebacks with genotyping-by-sequencing

Marine populations of the threespine stickleback (Gasterosteus aculeatus) have repeatedly colonized and rapidly adapted to freshwater habitats, providing a powerful system to map the genetic architecture of evolved traits. Here, we developed and applied a binned genotyping-by-sequencing (GBS) method to build dense genome-wide linkage maps of sticklebacks using two large marine by freshwater F2 crosses of more than 350 fish each. The resulting linkage maps significantly improve the genome assembly by anchoring 78 new scaffolds to chromosomes, reorienting 40 scaffolds, and rearranging scaffolds in 4 locations. In the revised genome assembly, 94.6% of the assembly was anchored to a chromosome. To assess linkage map quality, we mapped quantitative trait loci (QTL) controlling lateral plate number, which mapped as expected to a 200-kb genomic region containing Ectodysplasin, as well as a chromosome 7 QTL overlapping a previously identified modifier QTL. Finally, we mapped eight QTL controlling convergently evolved reductions in gill raker length in the two crosses, which revealed that this classic adaptive trait has a surprisingly modular and nonparallel genetic basis.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Transcriptome sequencing and marker development for four underutilized legumes

Premise of the study: Combating threats to food and nutrition security in the context of climate change and global population increase is one of the highest priorities of major international organizations. Hundreds of species are grown on a small scale in some of the most drought/flood-prone regions of the world and as such may harbor some of the most environmentally tolerant crops (and alleles). Methods and Results: In this study, transcriptomes were sequenced, assembled, and annotated for four underutilized legume crops. Microsatellite markers were identified in each species, as well as a conserved orthologous set of markers for cross-family phylogenetics and comparative mapping, which were ground-truthed on a panel of diverse legume germplasm. Conclusions: An understanding of these underutilized legumes will inform crop selection and breeding by allowing the investigation of genetic variation and the genetic basis of adaptive traits to be established.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Phylogenetic Relationships of Fig Wasps Pollinating Functionally Dioecious Ficus Based on Mitochondrial DNA Sequences and Morphology

The obligate mutualism between pollinating fig wasps in the family Agaonidae (Hymenoptera: Chalcidoidea) and Ficus species (Moraceae) is often regarded as an example of coevolution but little is known about the history of the interaction and understanding the origin of functionally dioecious fig pollination has been especially difficult. The phylogenetic relationships of fig wasps pollinating functionally dioecious Ficus were inferred from mitochondrial cytochrome oxidase gene sequences (mtDNA) and morphology. Separate and combined analyses indicated that the pollinators of functionally dioecious figs are not monophyletic. However, pollinator relationships were generally congruent with host phylogeny and support a revised classification of Ficus. Ancestral changes in pollinator ovipositor length were also correlated with changes in fig breeding system. In particular, the relative elongation of the ovipositor was associated with the repeated loss of functionally dioecious pollination. The concerted evolution of interacting morphologies may bias estimates of phylogeny based on female head characters but homoplasy is not so concerted in other morphological traits. The lesser phylogenetic utility of morphology compared to mtDNA is not due to rampant convergence in morphology but rather to the greater number of potentially informative characters in DNA sequence data and patterns of nucleotide substitution also limit the utility of mtDNA. None the less, inferring the ancestral associations of fig pollinators from the best-supported phylogeny provided strong evidence of host conservatism in this highly specialized mutualism.

opencc-zeroDec 2008View details →
dryad28/100

Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences

There is a lack of consensus on how next-generation sequence data should be considered for phylogenetic and phylogeographic estimates, with some studies excluding loci with missing data, while others include them, even when sequences are missing from a large number of individuals. Here we use simulations, focusing specifically on RAD sequences, to highlight some of the unforeseen consequence of excluding missing data from next-generation sequencing. Specifically, we show that in addition to the obvious effects associated with reducing the amount of data used to make historical inferences, the decisions we make about missing data (such as the minimum number of individuals with a sequence for a locus to be included in the study) also impact the types of loci sampled for a study. In particular, as the tolerance for missing data becomes more stringent, the mutational spectrum represented in the sampled loci becomes truncated such that loci with the highest mutation rates are disproportionately excluded. This effect is exacerbated further by factors involved in the preparation of the genomic library (i.e., the use of reduced representation libraries, as well as the coverage) and the taxonomic diversity represented in the library (i.e., the level of divergence among the individuals). We demonstrate that the intuitive appeals about being conservative by removing loci may be misguided.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Delimitation of the Thoracosphaeraceae (Dinophyceae), including the calcareous dinoflagellates, based on large amounts of ribosomal RNA sequence data

The phylogenetic relationships of the Dinophyceae (Alveolata) are not sufficiently resolved at present. The Thoracosphaeraceae (Peridiniales) are the only group of the Alveolata that include members with calcareous coccoid stages; this trait is considered apomorphic. Although the coccoid stage apparently is not calcareous, Bysmatrum has been assigned to the Thoracosphaeraceae based on thecal morphology. We tested the monophyly of the Thoracosphaeraceae using large sets of ribosomal RNA sequence data of the Alveolata including the Dinophyceae. Phylogenetic analyses were performed using Maximum Likelihood and Bayesian approaches. The Thoracosphaeraceae were monophyletic, but included also a number of non-calcareous dinophytes (such as Ensiculifera and Pfiesteria) and even parasites (such as Duboscquodinium and Tintinnophagus). Bysmatrum had an isolated and uncertain phylogenetic position outside the Thoracosphaeraceae. The phylogenetic relationships among calcareous dinophytes appear complex, and the assumption of the single origin of the potential to produce calcareous structures is challenged. The application of concatenated ribosomal RNA sequence data may prove promising for phylogenetic reconstructions of the Dinophyceae in future.

opencc-zeroDec 2010View details →
dryad28/100

Data from: The Strepsiptera Problem: Phylogeny of the Holometabolous Insect Orders Inferred from 18S and 28S Ribosomal DNA Sequences and Morphology

Phylogenetic relationships among the holometabolous insect orders were inferred from cladistic analysis of nucleotide sequences of 18S ribosomal DNA (rDNA) (85 exemplars) and 28S rDNA (52 exemplars) and morphological characters. Exemplar outgroup taxa were Collembola (1 sequence), Archaeognatha (1), Ephemerida (1), Odonata (2), Plecoptera (2), Blattodea (1), Mantodea (1), Dermaptera (1), Orthoptera (1), Phasmatodea (1), Embioptera (1), Psocoptera (1), Phthiraptera (1), Hemiptera (4), and Thysanoptera (1). Exemplar ingroup taxa were Coleoptera: Archostemata (1), Adephaga (2), and Polyphaga (7); Megaloptera (1); Raphidioptera (1); Neuroptera (sensu stricto ;eq Planipennia): Mantispoidea (2), Hemerobioidea (2), and Myrmeleontoidea (2); Hymenoptera: Symphyta (4) and Apocrita (19); Trichoptera: Hydropsychoidea (1) and Limnephiloidea (2); Lepidoptera: Ditrysia (3); Siphonaptera: Pulicoidea (1) and Ceratophylloidea (2); Mecoptera: Meropeidae (1), Boreidae (1), Panorpidae (1), and Bittacidae (2); Diptera: Nematocera (1), Brachycera (2), and Cyclorrhapha (1); and Strepsiptera: Corioxenidae (1), Myrmecolacidae (1), Elenchidae (1), and Stylopidae (3). We analyzed ~1 kilobase of 18S rDNA, starting 398 nucleotides downstream of the 5' end, and ~400 bp of 28S rDNA in expansion segment D3. Multiple alignment of the 18S and 28S sequences resulted in 1,116 nucleotide positions with 24 insert regions and 398 positions with 14 insert regions, respectively. All Strepsiptera and Neuroptera have large insert regions in 18S and 28S. The secondary structure of 18S insert 23 is composed of long stems that are GC rich in the basal Strepsiptera and AT rich in the more derived Strepsiptera. A matrix of 176 morphological characters was analyzed for holometabolous orders. Incongruence length difference tests indicate that the 28S + morphological data sets are incongruent but that 28S + 18S, 18S + morphology, and 28S + 18S + morphology fail to reject the hypothesis of congruence. Phylogenetic trees were generated by parsimony analysis, and clade robustness was evaluated by branch length, Bremer support, percentage of extra steps required to force paraphyly, and sensitivity analysis using the following parameters: gap weights, morphological character weights, methods of data set combination, removal of key taxa, and alignment region. The following are monophyletic under most or all combinations of parameter values: Holometabola, Polyphaga, Megaloptera + Raphidioptera, Neuroptera, Hymenoptera, Trichoptera, Lepidoptera, Amphiesmenoptera (Trichoptera + Lepidoptera), Siphonaptera, Siphonaptera + Mecoptera, Strepsiptera, Diptera, and Strepsiptera + Diptera (Halteria). Antliophora (Mecoptera + Diptera + Siphonaptera + Strepsiptera), Mecopterida (Antliophora + Amphiesmenoptera), and Hymenoptera + Mecopterida are supported in the majority of total evidence analyses. Mecoptera may be paraphyletic because Boreus is often placed as sister group to the fleas; hence, Siphonaptera may be subordinate within Mecoptera. The 18S sequences for Priacma (Coleoptera: Archostemata), Colpocaccus (Coleoptera: Adephaga), Agulla (Raphidioptera), and Corydalus (Megaloptera) are nearly identical, and Neuropterida are monophyletic only when those two beetle sequences are removed from the analysis. Coleoptera are therefore paraphyletic under almost all combinations of parameter values. Halteria and Amphiesmenoptera have high Bremer support values and long branch lengths. The data do not support placement of Strepsiptera outside of Holometabola nor as sister group to Coleoptera. We reject the notion that the monophyly of Halteria is due to long branch attraction because Strepsiptera and Diptera do not have the longest branches and there is phylogenetic congruence between molecules, across the entire parameter space, and between morphological and molecular data.

opencc-zeroDec 2008View details →
dryad28/100

Data from: Tree imbalance causes a bias in phylogenetic estimation of evolutionary timescales using heterochronous sequences

Phylogenetic estimation of evolutionary timescales has become routine in biology, forming the basis of a wide range of evolutionary and ecological studies. However, there are various sources of bias that can affect these estimates. We investigated whether tree imbalance, a property that is commonly observed in phylogenetic trees, can lead to reduced accuracy or precision of phylogenetic timescale estimates. We analysed simulated data sets with calibrations at internal nodes and at the tips, taking into consideration different calibration schemes and levels of tree imbalance. We also investigated the effect of tree imbalance on two empirical data sets: mitogenomes from primates and serial samples of the African swine fever virus. In analyses calibrated using dated, heterochronous tips, we found that tree imbalance had a detrimental impact on precision and produced a bias in which the overall timescale was underestimated. A pronounced effect was observed in analyses with shallow calibrations. The greatest decreases in accuracy usually occurred in the age estimates for medium and deep nodes of the tree. In contrast, analyses calibrated at internal nodes did not display a reduction in estimation accuracy or precision due to tree imbalance. Our results suggest that molecular-clock analyses can be improved by increasing taxon sampling, with the specific aims of including deeper calibrations, breaking up long branches and reducing tree imbalance.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record