Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data from: "Identification of SNP markers for the endangered Ugandan red colobus (Procolobus rufomitratus tephrosceles) using RAD sequencing" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015
Despite dramatic growth in the field of primate genomics over the past decade, studies of primate population and conservation genomics in the wild have been hampered due to the difficulties inherent in studying non-model organisms and endangered species, such as lack of a reference genome and current challenges in de novo primate genome assembly. Here, we used Restriction-site Associated DNA (RAD) sequencing to develop a population-based SNP panel for the Ugandan red colobus (P. rufomitratus tephrosceles), which is a highly threatened monkey due to habitat loss. We analyzed blood samples from 24 individuals from Kibale National Park (Uganda) using single-end RAD sequencing. We obtained 70,773,857 reads, of which 58,814,906 passed the filtering steps. Using the program STACKS v. 1.11 we identified 113,376 loci, of which 50,558 were polymorphic and had a mean observed heterozygosity of 0.25. These data will be used to study the effects of habitat fragmentation on genomic diversity, dispersal, and disease transmission in this species. Our approach provides a good example of the potential of RAD sequencing in studies of wild primate populations.
Data from: Deep COI sequencing of standardized benthic samples unveils overlooked diversity of Jordanian coral reefs in the Northern Red Sea
High-Throughput Sequencing (HTS) of DNA barcodes (metabarcoding), particularly when combined with standardized sampling protocols, is one of the most promising approaches for censusing overlooked cryptic invertebrate communities. We present biodiversity estimates based on sequencing of the cytochrome c oxidase subunit 1 (COI) gene for coral reefs of the Gulf of Aqaba, a semi-enclosed system in the Northern Red Sea. Samples were obtained from standardized sampling devices [Autonomous Reef Monitoring Structures (ARMS)] deployed for 18 months. DNA barcoding of non-sessile specimens >2mm revealed 83 OTUs in six phyla, of which only 25% matched a reference sequence in public databases. Metabarcoding of the 2mm-500μm and sessile bulk fractions revealed 1197 OTUs in 15 animal phyla, of which only 4.9% matched reference barcodes. These results highlight the scarcity of COI data for cryptobenthic organisms of the Red Sea. Compared with data obtained using similar methods, our results suggest that Gulf of Aqaba reefs are less diverse than two Pacific coral reefs but much more diverse than an Atlantic oyster reef at a similar latitude. The standardized approaches used here show promise for establishing baseline data on biodiversity, monitoring the impacts of environmental change, and quantifying patterns of diversity at regional and global scales.
Data from: Phylogeography of the arid shrub Atraphaxis frutescens (Polygonaceae) in northwestern China: evidence from cpDNA sequences
Climatic fluctuations during the Pleistocene are usually considered as a significant factor in shaping intraspecific genetic variation and influencing demographic histories. To well-understand these processes in desert northwest China, we selected arid adapted Atraphaxis frutescens as the study species. Two cpDNA regions (psbK-psbI, psbB-psbH) were sequenced in 272 individuals from 33 natural populations across the range of this shrub, and 10 haplotypes were identified. It was found to contain high levels of total gene diversity (H T = 0.858), and low levels of within-population diversity (H S = 0.092). Analysis of molecular variance (AMOVA) indicates that genetic differentiation primarily occurs among groups of populations. Based on BEAST (Bayesian Evolutionary Analysis Sampling Trees) analysis, we suggest that intraspecific differentiation of the species, resulting from isolated populations, accompanied enhanced desertification during the middle and late Pleistocene. The expansion of the Gurbantunggut and Kumtag deserts in this area appears to have triggered divergence among populations of the western, central, and eastern portions of the region and shaped genetic differentiation among them. Two possible independent glacial refugia were predicted, the Ili Valley and the northern Junggar Basin. Extensive development of arid habitats (desert margin and arid piedmont grassland) coupled with a more equable climate because the early Holocene are factors likely to have generated recent expansion of A. frutescens.
Data from: Who's for dinner? High-throughput sequencing reveals bat diet differentiation in a biodiversity hotspot where prey taxonomy is largely undescribed
Effective management and conservation of biodiversity requires understanding of predator–prey relationships to ensure the continued existence of both predator and prey populations. Gathering dietary data from predatory species, such as insectivorous bats, often presents logistical challenges, further exacerbated in biodiversity hot spots because prey items are highly speciose, yet their taxonomy is largely undescribed. We used high-throughput sequencing (HTS) and bioinformatic analyses to phylogenetically group DNA sequences into molecular operational taxonomic units (MOTUs) to examine predator–prey dynamics of three sympatric insectivorous bat species in the biodiversity hotspot of south-western Australia. We could only assign between 4% and 20% of MOTUs to known genera or species, depending on the method used, underscoring the importance of examining dietary diversity irrespective of taxonomic knowledge in areas lacking a comprehensive genetic reference database. MOTU analysis confirmed that resource partitioning occurred, with dietary divergence positively related to the ecomorphological divergence of the three bat species. We predicted that bat species' diets would converge during times of high energetic requirements, that is, the maternity season for females and the mating season for males. There was an interactive effect of season on female, but not male, bat species' diets, although small sample sizes may have limited our findings. Contrary to our predictions, females of two ecomorphologically similar species showed dietary convergence during the mating season rather than the maternity season. HTS-based approaches can help elucidate complex predator–prey relationships in highly speciose regions, which should facilitate the conservation of biodiversity in genetically uncharacterized areas, such as biodiversity hotspots.
Data from: Skin swabbing of amphibian larvae yields sufficient DNA for efficient sequencing and reliable microsatellite genotyping
Skin swabbing, a minimally invasive DNA sampling method recently developed on adult amphibians, was tested on larvae of fire salamanders (Salamandra salamandra). The quality and quantity of the sampled DNA was evaluated by (i) measuring DNA concentration in DNA extracts, (ii) sequencing part of the mtDNA cytochrome b gene (692 bp) and (iii) genotyping eight polymorphic nuclear microsatellite loci. The multiple-tubes approach was used for calculating allelic dropout (ADO) and false allele (FA) rates to evaluate the reliability of the genotypes. DNA extracts from tissue samples of road-killed individuals were included in the study as positive controls. Our results showed that skin swabs of fire salamander larvae can provide DNA in sufficient quantity and quality, as sequencing was successful and no allelic dropouts or false alleles were detected. This method, tested for the first time on amphibian larvae, has proven to be an efficient and reliable alternative to the controversial tail fin clipping procedure.
Data from: Genotyping-in-Thousands by sequencing (GT-seq) panel development and application to minimally-invasive DNA samples to support studies in molecular ecology
Minimally-invasive sampling (MIS) is widespread in wildlife studies; however, its utility for massively parallel DNA sequencing (MPS) is limited. Poor sample quality and contamination by exogenous DNA can make MIS challenging to use with modern genotyping-by-sequencing approaches, which have been traditionally developed for high-quality DNA sources. Given that MIS is often more appropriate in many contexts, there is a need to make such samples practical for harnessing MPS. Here, we test the ability for Genotyping-in-Thousands by sequencing (GT-seq), a multiplex amplicon sequencing approach, to effectively genotype minimally-invasive cloacal DNA samples collected from the Western Rattlesnake (Crotalus oreganus), a threatened species in British Columbia, Canada. As there was no previous genetic information for this species, an optimized panel of 362 SNPs was selected for use with GT-seq from a de novo restriction-site associated DNA sequencing (RADseq) assembly. Comparisons of genotypes generated within and among RADseq and GT-seq for the same individuals found low rates of genotyping error (GT-seq: 0.50%; RADseq: 0.80%) and discordance (2.57%), the latter likely due to the different genotype calling models employed. GT-seq mean genotype discordance between blood and cloacal swab samples collected from the same individuals was also minimal (1.37%). Estimates of population diversity parameters were similar across GT-seq and RADseq datasets, as were inferred patterns of population structure. Overall, GT-seq can be effectively applied to low quality DNA samples, minimizing the inefficiencies presented by exogenous DNA typically found in minimally-invasive samples and continuing the expansion of molecular ecology and conservation genetics in the genomics era.
Data from: A phylogeny of birds based on over 1,500 loci collected by target enrichment and high-throughput sequencing
Evolutionary relationships among birds in Neoaves, the clade comprising the vast majority of avian diversity, have vexed systematists due to the ancient, rapid radiation of numerous lineages. We applied a new phylogenomic approach to resolve relationships in Neoaves using target enrichment (sequence capture) and high-throughput sequencing of ultraconserved elements (UCEs) in avian genomes. We collected sequence data from UCE loci for 32 members of Neoaves and one outgroup (chicken) and analyzed data sets that differed in their amount of missing data. An alignment of 1,541 loci that allowed missing data was 87% complete and resulted in a highly resolved phylogeny with broad agreement between the Bayesian and maximum-likelihood (ML) trees. Although results from the 100% complete matrix of 416 UCE loci was similar, the Bayesian and ML trees differed to a greater extent in this analysis, suggesting that increasing from 416 to 1,541 loci led to increased stability and resolution of the tree. Novel results of our study include surprisingly close relationships between phenotypically divergent bird families, such as tropicbirds (Phaethontidae) and the sunbittern (Eurypygidae) as well as between bustards (Otididae) and turacos (Musophagidae). This phylogeny bolsters support for monophyletic waterbird and landbird clades and also strongly supports controversial results from previous studies, including the sister relationship between passerines and parrots and the non-monophyly of raptorial birds in the hawk and falcon families. Although significant challenges remain to fully resolving some of the deep relationships in Neoaves, especially among lineages outside the waterbirds and landbirds, this study suggests that increased data will yield an increasingly resolved avian phylogeny.
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy has better precision and recall (with respect to the true alignments) than the other alignment methods on the simulated data sets, but has consistently lower recall on the biological benchmarks (with respect to the reference alignments) than many of the other methods. In other words, we find that BAli-Phy systematically under-aligns when operating on biological sequence data, but shows no sign of this on simulated data. There are several potential causes for this change in performance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments, and future research is needed to determine the most likely explanation. We conclude with a discussion of the potential ramifications for each of these possibilities.
Data from: Evolutionary history of endemic Sulawesi squirrels constructed from UCEs and mitogenomes sequenced from museum specimens
Background: The Indonesian island of Sulawesi has a complex geological history. It is composed of several landmasses that have arrived at a near modern configuration only in the past few million years. It is the largest island in the biodiversity hotspot of Wallacea—an area demarcated by the biogeographic breaks between Wallace's and Lydekker's lines. The mammal fauna of Sulawesi is transitional between Asian and Australian faunas. Sulawesi's three genera of squirrels, all endemic (subfamily Nannosciurinae: Hyosciurus, Rubrisciurus and Prosciurillus), are of Asian origin and have evolved a variety of phenotypes that allow a range of ecological niche specializations. Here we present a molecular phylogeny of this radiation using data from museum specimens. High throughput sequencing technology was used to generate whole mitochondrial genomes and a panel of nuclear ultraconserved elements providing a large genome-wide dataset for inferring phylogenetic relationships. Results: Our analysis confirmed monophyly of the Sulawesi taxa with deep divergences between the three endemic genera, which predate the amalgamation of the current island of Sulawesi. This suggests lineages may have evolved in allopatry after crossing Wallace's line. Nuclear and mitochondrial analyses were largely congruent and well supported, except for the placement of Prosciurillus murinus. Mitochondrial analysis revealed paraphyly for Prosciurillus, with P. murinus between or outside of Hyosciurus and Rubrisciurus, separate from other species of Prosciurillus. A deep but monophyletic history for the four included species of Prosciurillus was recovered with the nuclear data. Conclusions: The divergence of the Sulawesi squirrels from their closest relatives dated to ~9.7–12.5 million years ago (MYA), pushing back the age estimate of this ancient adaptive radiation prior to the formation of the current conformation of Sulawesi. Generic level diversification took place around 9.7 MYA, opening the possibility that the genera represent allopatric lineages that evolved in isolation in an ancient proto-Sulawesian archipelago. We propose that incongruence between phylogenies based on nuclear and mitochondrial sequences may have resulted from biogeographic discordance, when two allopatric lineages come into secondary contact, with complete replacement of the mitochondria in one species.
Data from: Spatial dynamics and mixing of bluefin tuna in the Atlantic Ocean and Mediterranean Sea revealed using next generation sequencing
The Atlantic bluefin tuna is a highly migratory species emblematic of the challenges associated with shared fisheries management. In an effort to resolve the species' stock dynamics, a genome-wide search for spatially informative single nucleotide polymorphisms (SNPs) was undertaken, by way of sequencing reduced representation libraries. An allele frequency approach to SNP discovery was used, combining the data of 555 larvae and young-of-the-year (LYOY) into pools representing major geographical areas and mapping against a newly assembled genomic reference. From a set of 184,895 candidate loci, 384 were selected for validation using 167 LYOY. A highly discriminatory genotyping panel of 95 SNPs was ultimately developed by selecting loci with the most pronounced differences between western Atlantic and Mediterranean Sea LYOY. The panel was evaluated by genotyping a different set of LYOY (n= 326) and from these 77.8% and 82.1% were correctly assigned to western Atlantic and Mediterranean Sea origins, respectively. The panel revealed temporally persistent differentiation among LYOY from the western Atlantic and Mediterranean Sea (FST = 0.008, p=0.034). The composition of six mixed feeding aggregations in the Atlantic Ocean and Mediterranean Sea was characterized using genotypes from medium (n = 184) and large (n = 48) adults, applying population assignment and mixture analyses. The results provide evidence of persistent population structuring across broad geographic areas and extensive mixing in the Atlantic Ocean, particularly in the mid-Atlantic Bight and Gulf of St. Lawrence. The genomic reference and genotyping tools presented here constitute novel resources useful for future research and conservation efforts.
Data from: Genome-wide RAD sequence data provide unprecedented resolution of species boundaries and relationships in the Lake Victoria cichlid adaptive radiation
Although population genomic studies using next generation sequencing (NGS) data are becoming increasingly common, studies focusing on phylogenetic inference using these data are in their infancy. Here, we use NGS data generated from reduced representation genomic libraries of restriction-site-associated DNA (RAD) markers to infer phylogenetic relationships among 16 species of cichlid fishes from a single rocky island community within Lake Victoria's cichlid adaptive radiation. Previous attempts at sequence-based phylogenetic analyses in Victoria cichlids have shown extensive sharing of genetic variation among species and no resolution of species or higher-level relationships. These patterns have generally been attributed to the very recent origin (<15 000 years) of the radiation, and ongoing hybridization between species. We show that as we increase the amount of sequence data used in phylogenetic analyses, we produce phylogenetic trees with unprecedented resolution for this group. In trees derived from our largest data supermatrices (3 to >5.8 million base pairs in width), species are reciprocally monophyletic with high bootstrap support, and the majority of internal branches on the tree have high support. Given the difficulty of the phylogenetic problem that the Lake Victoria cichlid adaptive radiation represents, these results are striking. The strict interpretation of the topologies we present here warrants caution because many questions remain about phylogenetic inference with very large genomic data set and because we can with the current analysis not distinguish between effects of shared ancestry and post-speciation gene flow. However, these results provide the first conclusive evidence for the monophyly of species in the Lake Victoria cichlid radiation and demonstrate the power that NGS data sets hold to resolve even the most difficult of phylogenetic challenges.
Data from: De novo sequencing and assembly of Azadirachta indica fruit transcriptome
Azadirachta indica (neem) is a unique, versatile and important tree species. Many parts of the plant are traditionally used as pesticide, insecticide, fungicide and for other medicinal purposes. Azadirachta fruits and seeds, a good source of oil, are widely used for agriculturally important pest management. Neem oil and its derivatives also support multiple cottage industries in India. Past efforts have been mostly concentrated towards identifying, characterizing and synthesizing one of its principal components, i.e. azadirachtin from seed kernels. Despite diverse use of the neem plant, a modern drug-development programme which systematically exploits the therapeutic ability of Azadirachta fruits remains to be fully established. Next generation sequencing technology that helps decode genomes and transcriptomes has transformational impact on medicine, agriculture, bio-fuel and biodiversity studies. Here, we report sequencing, assembly and analysis of Azadirachta fruit transcriptome using next-generation sequencing technology. We believe that our study shall offer valuable insights towards realizing the larger vision of understanding the key medicinally active compounds and their pathways.
Data from: Use of RAD sequencing for delimiting species
RAD-tag sequencing is a promising method for conducting genome-wide evolutionary studies. However, to date, only a handful of studies empirically tested its applicability above the species level. In this communication, we use RAD tags to contribute to the delimitation of species within a diverse genus of deep-sea octocorals, Chrysogorgia, for which few classical genetic markers have proved informative. Previous studies have hypothesized that single mitochondrial haplotypes can be used to delimit Chrysogorgia species. On the basis of two lanes of Illumina sequencing, we inferred phylogenetic relationships among 12 putative species that were delimited using mitochondrial data, comparing two RAD analysis pipelines (Stacks and PyRAD). The number of homologous RAD loci decreased dramatically with increasing divergence, as >70% of loci are lost when comparing specimens separated by two mutations on the 700-nt long mitochondrial phylogeny. Species delimitation hypotheses based on the mitochondrial mtMutS gene are largely supported, as six out of nine putative species represented by more than one colony were recovered as discrete, well-supported clades. Significant genetic structure (correlating with geography) was detected within one putative species, suggesting that individuals characterized by the same mtMutS haplotype may belong to distinct species. Conversely, three mtMutS haplotypes formed one well-supported clade within which no population structure was detected, also suggesting that intraspecific variation exists at mtMutS in Chrysogorgia. Despite an impressive decrease in the number of homologous loci across clades, RAD data helped us to fine-tune our interpretations of classical mitochondrial markers used in octocoral species delimitation, and discover previously undetected diversity.
Figure 6 in The species of the varius group of Coccophagus (Hymenoptera: Aphelinidae) from China, with description of a new species, DNA sequence data, and a new country record
Figure 6. Maximum likelihood (ML) tree inferred using IQ-TREE, version 1.6. Bootstrap support values indicated on branches; scale bar represents the number of nucleotide substitutions per site.
Figure 5 in The species of the varius group of Coccophagus (Hymenoptera: Aphelinidae) from China, with description of a new species, DNA sequence data, and a new country record
Figure 5. Coccophagus yunnana sp.nov. female. (a) antenna; (b) fore wing; (c) F3 and clubs; (d) stigma vein of fore wing.
Figure 4 in The species of the varius group of Coccophagus (Hymenoptera: Aphelinidae) from China, with description of a new species, DNA sequence data, and a new country record
Figure 4. Coccophagus yunnana sp. nov. female. (a) head and mesosoma; (b) head in face view; (c) mesosoma; (d) ovipositor.
Figure 1 in The species of the varius group of Coccophagus (Hymenoptera: Aphelinidae) from China, with description of a new species, DNA sequence data, and a new country record
Figure 1. Coccophagus anchoroides (Huang), female. (a) antenna; (b) fore wing; (c) mesosoma and metasoma; (d) ovipositor, mid-tibia and tarsus. Scale bars = μm (from Huang 1994).
Figure 2 in The species of the varius group of Coccophagus (Hymenoptera: Aphelinidae) from China, with description of a new species, DNA sequence data, and a new country record
Figure 2. Coccophagus fumadus Hayat, female. (a) meso- and metasoma (gaster); (b) antenna; (c) fore wing; (d) dorsal mesosoma; (e) head.
Figure 3 in The species of the varius group of Coccophagus (Hymenoptera: Aphelinidae) from China, with description of a new species, DNA sequence data, and a new country record
Figure 3. Coccophagus yunnana sp. nov. female. (a) coccid scale host with C. yunnana pupa and meconia visible; (b) pupa in dorsal view; (c) pupa in ventral view; (d) body in dorsal view; (e) body in ventral view.
Data from: Digital fragment analysis of short tandem repeats by high-throughput amplicon sequencing
High-throughput sequencing has been proposed as a method to genotype microsatellites and overcome the four main technical drawbacks of capillary electrophoresis: amplification artifacts, imprecise sizing, length homoplasy, and limited multiplex capability. The objective of this project was to test a high-throughput amplicon sequencing approach to fragment analysis of short tandem repeats and characterize its advantages and disadvantages against traditional capillary electrophoresis. We amplified and sequenced 12 muskrat microsatellite loci from 180 muskrat specimens and analyzed the sequencing data for precision of allele calling, propensity for amplification or sequencing artifacts, and for evidence of length homoplasy. Of the 294 total alleles, we detected by sequencing, only 164 alleles would have been detected by capillary electrophoresis as the remaining 130 alleles (44%) would have been hidden by length homoplasy. The ability to detect a greater number of unique alleles resulted in the ability to resolve greater population genetic structure. The primary advantages of fragment analysis by sequencing are the ability to precisely size fragments, resolve length homoplasy, multiplex many individuals and many loci into a single high-throughput run, and compare data across projects and across laboratories (present and future) with minimal technical calibration. A significant disadvantage of fragment analysis by sequencing is that the method is only practical and cost-effective when performed on batches of several hundred samples with multiple loci. Future work is needed to optimize throughput while minimizing costs and to update existing microsatellite allele calling and analysis programs to accommodate sequence-aware microsatellite data.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.