Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,109
datasets available to search
ShareScore release 0.9.0
Dataset results
3,109 results for “sequence analysis”
Data from: Whole genome-sequencing and phylogenetic analysis of a historical collection of Bacillus anthracis strains from Danish cattle
Bacillus anthracis, the causative agent of anthrax, is known as one of the most genetically monomorphic species. Canonical single-nucleotide polymorphism (SNP) typing and whole-genome sequencing were used to investigate the molecular diversity of eleven B. anthracis strains isolated from cattle in Denmark between 1935 and 1988. Danish strains were assigned into five canSNP groups or lineages, i.e. A.Br.001/002 (n = 4), A.Br.Ames (n = 2), A.Br.008/011 (n = 2), A.Br.005/006 (n = 2) and A.Br.Aust94 (n = 1). The match with the A.Br.Ames lineage is of particular interest as the occurrence of such lineage in Europe is demonstrated for the first time, filling an historical gap within the phylogeography of the lineage. Comparative genome analyses of these strains with 41 isolates from other parts of the world revealed that the two Danish A.Br.008/011 strains were related to the heroin-associated strains responsible for outbreaks of injection anthrax in drug users in Europe. Eight novel diagnostic SNPs that specifically discriminate the different sub-groups of Danish strains were identified and developed into PCR-based genotyping assays.
Data from: Carnivore diet analysis based on next-generation sequencing: application to the leopard cat (Prionailurus bengalensis) in Pakistan
Diet analysis is a prerequisite to fully understand the biology of a species and the functioning of ecosystems. For carnivores, traditional diet analyses mostly rely upon the morphological identification of undigested remains in the feces. Here, we developed a methodology for carnivore diet analyses based on next generation sequencing. We applied this approach to the analysis of the vertebrate component of leopard cat diet in two ecologically distinct regions in northern Pakistan. Despite being a relatively common species with a wide distribution in Asia, little is known about this elusive predator. We analyzed a total of 38 leopard cat feces. After a classical DNA extraction, the DNA extracts were amplified using primers for vertebrates targeting about 100 bp of the mitochondrial 12S rRNA gene, with and without a blocking oligonucleotide specific to the predator sequence. The amplification products were then sequenced on a next generation sequencer. We identified a total of 18 prey taxa, including eight mammals, eight birds, one amphibian, and one fish. In general, our results confirmed that the leopard cat has a very eclectic diet, and feeds mainly on rodents, and particularly on the Muridae family. The DNA-based approach we propose here represents a valuable complement to current conventional methods. It can be applied to other carnivore species with only a slight adjustment relating to the design of the blocking oligonucleotide. It is robust, simple to implement, and allows the possibility of very large-scale analyses.
Data from: Comparative population genetic analysis of bocaccio rockfish Sebastes paucispinis using anonymous and gene-associated simple sequence repeat loci
Comparative population genetic analyses of traditional and emergent molecular markers aid in determining appropriate use of new technologies. The bocaccio rockfish Sebastes paucispinis is a high-gene-flow marine species off the west coast of North America that experienced strong population decline over the past three decades. We used 18 anonymous and 13 gene associated simple sequence repeat loci (EST-SSRs) to characterize range-wide population structure with temporal replicates. No FST-outliers were detected using the LOSITAN program, suggesting that neither balancing nor divergent selection affected the loci surveyed. Consistent hierarchical structuring of populations by geography or year class was not detected regardless of marker class. The EST-SSRs were less variable than the anonymous SSRs, but no correlation between FST and variation or marker class was observed. General Linear Model analysis showed that low EST-SSR variation was attributable to low mean repeat number. Comparative genomic analysis with Gasterosteus aculeatus, Takifugu rubripes, and Oryzias latipes showed consistently lower repeat number in EST-SSRs than SSR loci that were not in ESTs. Purifying selection likely imposed functional constraints on EST-SSRs resulting in low repeat numbers that affected diversity estimates, but did not affect the observed pattern of population structure.
Data from: Extending RAD tag analysis to microbial ecology: a comparison between multi locus sequence typing (MLST) and 2b-RAD to investigate Listeria monocytogenes genetic structure
The advent of next-generation sequencing (NGS) has dramatically changed bacterial typing technologies, increasing our ability to differentiate bacterial isolates. Despite it is now possible to sequence a bacterial genome in a few days and at reasonable costs, most genetic analyses do not require whole-genome sequencing, which also remains impractical for large population samples due to the cost of individual library preparation and bioinformatics. More traditional sequencing approaches, however, such as MultiLocus Sequence Typing (mlst) are quite laborious and time-consuming, especially for large-scale analyses. In this study, a genotyping approach based on restriction site-associated (RAD) tag sequencing, 2b-RAD, was applied to characterize Listeria monocytogenes strains. To verify the feasibility of the method, an in silico analysis was performed on 30 available complete genomes. For the same set of strains, in silico mlst analysis was conducted as well. Subsequently, 2b-RAD and mlst analyses were experimentally carried out on 58 isolates collected from food samples or food-processing sites. The obtained results demonstrate that 2b-RAD predicts mlst types and often provides more detailed information on population structure than mlst. Moreover, the majority of variants differentiating identical sequence type isolates mapped against accessory fragments, thus providing additional information to characterize strains. Although mlst still represents a reliable typing method, large-scale studies on molecular epidemiology and public health, as well as bacterial phylogenetics, population genetics and biosafety could benefit of a low cost and fast turnaround time approach such as the 2b-RAD analysis proposed here.
Data from: Phylogeny of Mycoplasma bovis isolates from Hungary based on multi locus sequence typing and multiple-locus variable-number tandem repeat analysis
Background: Mycoplasma bovis is an important pathogen causing pneumonia, mastitis and arthritis in cattle worldwide. As this agent is primarily transmitted by direct contact and spread through animal movements, efficient genotyping systems are essential for the monitoring of the disease and for epidemiological investigations. The aim of this study was to compare and evaluate the multi locus sequence typing (MLST) and the multiple-locus variable-number tandem-repeat (VNTR) analysis (MLVA) through the genetic characterization of M. bovis isolates from Hungary. Results: Thirty one Hungarian M. bovis isolates grouped into two clades by MLST. Two strains had the same sequence type (ST) as reference strain PG45, while the other twenty nine Hungarian isolates formed a novel clade comprising five subclades. Isolates originating from the same herds had the same STs except for one case. The same isolates formed two main clades and several subclades and branches by MLVA. One clade contained the reference strain PG45 and three isolates, while the other main clade comprised the rest of the strains. Within-herd strain divergence was also detected by MLVA. Little congruence was found between the results of the two typing systems. Conclusions: MLST is generally considered an intermediate scale typing method and it was found to be discriminatory among the Hungarian M. bovis isolates. MLVA proved to be an appropriate fine scale typing tool for M. bovis as this method was able to distinguish closely related strains isolated from the same farm. We recommend the combined use of the two methods for the genotyping of M. bovis isolates. Strains have to be characterized first by MLST followed by the fine scale typing of identical STs with MLVA.
Data from: Phylogenomic analysis of the Chilean clade of Liolaemus lizards (Squamata: Liolaemidae) based on sequence capture data
The genus Liolaemus is one of the most ecologically diverse and species-rich genera of lizards worldwide. It currently includes more than 250 recognized species, which have been subject to many ecological and evolutionary studies. Nevertheless, Liolaemus lizards have a complex taxonomic history, mainly due to the incongruence between morphological and genetic data, incomplete taxon sampling, incomplete lineage sorting and hybridization. In addition, as many species have restricted and remote distributions, this has hampered their examination and inclusion in molecular systematic studies. The aims of this study are to infer a robust phylogeny for a subsample of lizards representing the Chilean clade (subgenus Liolaemus sensu stricto), and to test the monophyly of several of the major species groups. We use a phylogenomic approach, targeting 541 ultra-conserved elements (UCEs) and 44 protein-coding genes for 16 taxa. We conduct a comparison of phylogenetic analyses using maximum-likelihood and several species tree inference methods. The UCEs provide stronger support for phylogenetic relationships compared to the protein-coding genes; however, the UCEs outnumber the protein-coding genes by 10-fold. On average, the protein-coding genes contain over twice the number of informative sites. Based on our phylogenomic analyses, all the groups sampled are polyphyletic. Liolaemus tenuis tenuis is difficult to place in the phylogeny, because only a few loci (nine) were recovered for this species. Topologies or support values did not change dramatically upon exclusion of L. t. tenuis from analyses, suggesting that missing data did not had a significant impact on phylogenetic inference in this data set. The phylogenomic analyses provide strong support for sister group relationships between L. fuscus, L. monticola, L. nigroviridis and L. nitidus, and L. platei and L. velosoi. Despite our limited taxon sampling, we have provided a reliable starting hypothesis for the relationships among many major groups of the Chilean clade of Liolaemus that will help future work aimed at resolving the Liolaemus phylogeny.
Data from: Analysis of transposable elements in the genome of Asparagus officinalis from high coverage sequence data
Asparagus officinalis is an economically and nutritionally important vegetable crop that is widely cultivated and is used as a model dioecious species to study plant sex determination and sex chromosome evolution. To improve our understanding of its genome composition, especially with respect to transposable elements (TEs), which make up the majority of the genome, we performed Illumina HiSeq2000 sequencing of both male and female asparagus genomes followed by bioinformatics analysis. We generated 17 Gb of sequence (12×coverage) and assembled them into 163,406 scaffolds with a total cumulated length of 400 Mbp, which represent about 30% of asparagus genome. Overall, TEs masked about 53% of the A. officinalis assembly. Majority of the identified TEs belonged to LTR retrotransposons, which constitute about 28% of genomic DNA, with Ty1/copia elements being more diverse and accumulated to higher copy numbers than Ty3/gypsy. Compared with LTR retrotransposons, non-LTR retrotransposons and DNA transposons were relatively rare. In addition, comparison of the abundance of the TE groups between male and female genomes showed that the overall TE composition was highly similar, with only slight differences in the abundance of several TE groups, which is consistent with the relatively recent origin of asparagus sex chromosomes. This study greatly improves our knowledge of the repetitive sequence construction of asparagus, which facilitates the identification of TEs responsible for the early evolution of plant sex chromosomes and is helpful for further studies on this dioecious plant.
FIGURES 510. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 510. Lamyctes hellyeri n. sp. 5, 10, QVMAG 23:23044, holotype female. 5, dorsal habitus, scale 1 mm; 10, ventral view of posterior segments and gonopods, scale 100 m. 69, QVMAG 23:23045, female, scale 0.5 mm. 6, leg 12; 7, leg 13; 8, leg 14; 9, leg 15.
FIGURES 2633. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 2633. Lamyctes hellyeri n. sp. 2631, QVMAG 23:23046, female. 2627, gnathal edge of mandible and detail of ventral part, scales 10 m; 28, aciculae, scale 10 m; 29, 30, fringe of branching bristles, on successively more dorsal part of mandible, scales 10 m; 31, sternite of segment 15 and posterior margin of sternite 14, scale 100 m. 3233, QVMAG 23:23047, female, gonopod and detail of spurs and claw, scales 50 m, 10 m.
FIGURE 38 in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURE 38. Cladograms based on molecular sequence data. Cladograms at left are shortest based on parameter set (121) that minimises incongruence between genes; cladograms at right are strict consensus of all 15 explored parameter sets. Numbers at nodes are parsimony jackknife frequencies. From left to right, top to bottom: cladograms based on combined molecular data (2198 steps); cladograms based on 18S rRNA (562 steps); cladograms based on 28S rRNA (98 steps); cladograms based on 16S rRNA (616 steps); cladograms based on COI (901 steps).
FIGURES 3437. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 3437. Lamyctes hellyeri n. sp. QVMAG 23:23048, female, pretarsus of leg 14, scales 10 m. 3436, anterior, posterior, and ventral views; 37, detail of lateral pore and ornament on scutes of main claw.
FIGURES 1825. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 1825. Lamyctes hellyeri n. sp. QVMAG 23:23046, female. 18, ventral view of maxillipede, scale 100 m; 1920, dental margin of maxillipede coxosternite, scales 50 m, 10 m; 21, tarsus and claw of second maxilla, scale 50 m; 22, distal part of tarsus and claw of second maxilla, scale 10 m; 23, coxal projections and telopods of first maxillae, scale 50 m; 24, first maxillae, scale 100 m; 25, plumose setae on inner margins of telopods of first maxillae, scale 10 m.
FIGURES 1117. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 1117. Lamyctes hellyeri n. sp. 11, 1417, QVMAG 23:23046, female. 11, anterior part of head shield and basal part of antennae, scale 100 m; 14, sensilla on dorsal side of antenna, scale 10 m; 1516, antennal articles, dorsal side, scales 50 m; 17, cephalic pleurite with Tömösváry organ, scale 50 m. 1213, QVMAG 23:23047, female. 12, ventral view of clypeus and labrum, scale 100 m; 13, labral midpiece and inner parts of sidepieces, scale 30 m.
FIGURES 14 in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 14. Lamyctes coeculus (Brölemann). 1, 3, AM KS57961, female, Mellong Range, NSW, Australia. 2, 4, MCZ DNA100472, female, Cerro San Javier, Tucumán, Argentina. 12, ventral view of head, scales 100 m; 34, dental margin of maxillipede coxosternite, scales 50 m.
FIGURE 2 in The morphology and SSU rRNA gene sequence analysis of a poorly-known brackish water ciliate, Pinacocoleps tesselatus (Kahl, 1930) (Ciliophora, Colepidae) from Hangzhou Bay, China
FIGURE 2. Photomicrographs of Pinacocoleps tesselatus (Kahl, 1930) from live cells (A–G), after silver carbonate impregnation (H, J), and after protargol impregnation (I). (A) Lateral view of a typical individual. (B) Anterior secondary plate. (C) Anterior main plate. (D) Posterior main plate. (E) Posterior secondary plate. (F) A squashed specimen showing the arrangement of plates. (G) Posterior view, arrows mark the posterior spines. (H) Lateral view, arrows denote the oral basket, arrowheads mark the extrusomes. (I) Anterior view, showing the oral structure, arrows indicate the adoral organelles. (J) Lateral view, showing the ciliary pattern. Scale bars (in A, J) = 30 μm; in (B–H) = 10 μm.
FIGURE 1 in The morphology and SSU rRNA gene sequence analysis of a poorly-known brackish water ciliate, Pinacocoleps tesselatus (Kahl, 1930) (Ciliophora, Colepidae) from Hangzhou Bay, China
FIGURE 1. Morphology and infraciliature of Pinacocoleps tesselatus (Kahl, 1930) Foissner et al. 2008 (A–E), P. similis (Kahl, 1933) Chen et al. 2010 (F, G), P. heteracanthus (Noland, 1937) Chen et al. 2010 (H), P. arenarius (Bock, 1952) Chen et al. 2010 (I), P. spiralis (Noland, 1937) Chen et al. 2010 (J), P. i n c u r v u s (Ehrenberg, 1833) Foissner et al. 2008 (K), and P. pulcher (Spiegel, 1926) Foissner et al. 2008 (L). (A) Lateral view of typical individual. (B) One row of plates. The circumoral plate is omitted. (C) Ciliary pattern at apical end of body. (D) Ciliary pattern of P. t e s s e l a t u s. (E) P. tesselatus (Kahl, 1930) (from Kahl 1930). (F) P. similis (Kahl, 1933) (from Chen et al. 2010). (G) P. similis (Kahl, 1933) (from Borror 1972). (H) P. heteracanthus (Noland, 1937) (from Noland 1937). (I) P. arenarius (Bock, 1952) (from Bock 1952). (J) P. spiralis (Noland, 1937) (from Noland 1937). (K) P. i n c u r v u s (Ehrenberg, 1933) (from Kahl 1930). (L) P. pulcher (Spiegel, 1926) (from Kahl 1930). AO = adoral organelle; AS = anterior spine; CC = caudal cilium; CK = circumoral kinety; Ma = macronucleus; Mi = micronucleus; PC = perioral ciliature; PS = posterior spine; SK = somatic kinety. Scale bars = 30 μm.
FIGURE 3 in The morphology and SSU rRNA gene sequence analysis of a poorly-known brackish water ciliate, Pinacocoleps tesselatus (Kahl, 1930) (Ciliophora, Colepidae) from Hangzhou Bay, China
FIGURE 3. Maximum likelihood (ML) phylogenetic tree based on the small subunit (SSU) rDNA of Pinacocoleps tesselatus and other colepids. Numbers at branching points show bootstrap values of 1,000 replicates for ML tree and posterior probability for Bayesian (BI) tree, respectively. Fully supported (100%/1.00) branches are marked with solid circles. The scale bar corresponds to 10 substitutions per 100 nucleotide positions. Taxonomic classification mainly follows Lynn (2008). GenBank numbers follow species names.
FIGURE 2 in Molecular Phylogenetic Analysis of the Orthoptera (Arthropoda, Insecta) based on Hexamerin Sequences
FIGURE 2. Bayesian phylogenetic tree resulting from analysis of thirty-four the hexamerins sequences in insects. Next to nodes are bootstrap values. The outgroup species of proteins as follows: AmeHex70c: Apis mellifera, XM-392869; CfeHex2: Camponotus festinatus, AJ251271; BheHex: Bracon hebetor, I25974; AmeHex70b: Apis mellifera, AY601637; CfrHx1: Campodea fragilis, JX867269; CfrHx2: Campodea fragilis, JX867270; CspHex1: Campodea sp., CAX63173.
FIGURE 1 in Molecular Phylogenetic Analysis of the Orthoptera (Arthropoda, Insecta) based on Hexamerin Sequences
FIGURE 1. Multiple alignment of Orthoptera hexamerin sequences. Putative hexamerins from L. migratoria (LmiHx2), R. microptera (RmiHx2), A. cinerea (AciHx2), C. italicus (CitHx2), M. wardi (MwaHx2), O. tibetanus (OtiHx2), C. versicolor (CveHx2), A. sinensis (AsiHx2), C. brunneus (CbrHx2), H. brunneriana (HbrHx1 and HbrHx2), X. japonicus (XjaHx1,2,4,5), P. soochowensis (PsoHx2), P. teretrirsostris (PteHx2), T. subulata (TsuHx1,2,5), Gryllotalpa sp. (GspHx1 and GspHx2), T. commodus (TcoHx1 and TcoHx2) and Ceuthophilus sp. (CespHx2 and CespHx3) were compared. The copper-binding histidines are shaded in gray; other strictly conserved residues are shaded in blue, Putative signal peptides are underlined.
FIGURE 3 in Molecular Phylogenetic Analysis of the Orthoptera (Arthropoda, Insecta) based on Hexamerin Sequences
FIGURE 3. Neighbor-joining phylogenetic tree resulting from analysis of thirty-four the hexamerins sequences in insects. Next to nodes are bootstrap values.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.