Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Data from: Sequencing of seven haloarchaeal genomes reveals patterns of genomic flux
We report the sequencing of seven genomes from two haloarchaeal genera, Haloferax and Haloarcula. Ease of cultivation and the existence of well-developed genetic and biochemical tools for several diverse haloarchaeal species make haloarchaea a model group for the study of archaeal biology. The unique physiological properties of these organisms also make them good candidates for novel enzyme discovery for biotechnological applications. Seven genomes were sequenced to ~20×coverage and assembled to an average of 50 contigs (range 5 scaffolds - 168 contigs). Comparisons of protein-coding gene compliments revealed large-scale differences in COG functional group enrichment between these genera. Analysis of genes encoding machinery for DNA metabolism reveals genera-specific expansions of the general transcription factor TATA binding protein as well as a history of extensive duplication and horizontal transfer of the proliferating cell nuclear antigen. Insights gained from this study emphasize the importance of haloarchaea for investigation of archaeal biology.
Data from: Genome sequence of M6, a diploid inbred clone of the high glycoalkaloid-producing tuber-bearing potato species Solanum chacoense, reveals residual heterozygosity
Cultivated potato (Solanum tuberosum L.) is a highly heterozygous autotetraploid that presents challenges in genome analyses and breeding. Wild potato species serve as a resource for the introgression of important agronomic traits into cultivated potato. One key species is Solanum chacoense and the diploid, inbred clone M6, which is self-compatible and has desirable tuber market quality and disease resistance traits. Sequencing and assembly of the genome of the M6 clone of S. chacoense generated an assembly of 825,767,562 bp in 8,260 scaffolds with an N50 scaffold size of 713,602 bp. Pseudomolecule construction anchored 508 Mb of the genome assembly into 12 chromosomes. Genome annotation yielded 49,124 high confidence gene models representing 37,740 genes. Comparative analyses of the M6 genome with six other Solanaceae species revealed a core set of 158,367 Solanaceae genes and 1,897 genes unique to three potato species. Analysis of single nucleotide polymorphisms across the M6 genome revealed enhanced residual heterozygosity on chromosomes 4, 8 and 9 relative to the other chromosomes. Access to the M6 genome provides a resource for identification of key genes for important agronomic traits and aids in genome-enabled development of inbred diploid potatoes with the potential to accelerate potato breeding.
Data from: GIbPSs: a toolkit for fast and accurate analyses of genotyping-by-sequencing data without a reference genome
Genotyping-by-sequencing (GBS) and related methods are increasingly used for studies of non-model organisms from population genetic to phylogenetic scales. We present GIbPSs, a new genotyping toolkit for the analysis of data from various protocols such as RAD, double-digest RAD, GBS, and two-enzyme GBS without a reference genome. GIbPSs can handle paired-end GBS data and is able to assign reads from both strands of a restriction fragment to the same locus. GIbPSs is most suitable for population genetic and phylogeographic analyses. It avoids genotyping errors due to indel variation by identifying and discarding affected loci. GIbPSs creates a genotype database that offers rich functionality for data filtering and export in numerous formats. We performed comparative analyses of simulated and real GBS data with GIbPSs and another program, pyRAD. This program accounts for indel variation by aligning homologous sequences. GIbPSs performed better than pyRAD in several aspects. It required much less computation time and displayed higher genotyping accuracy. GIbPSs retained smaller numbers of loci overall in analyses of real GBS data. It nevertheless delivered more complete genotype matrices with greater locus overlap between individuals and greater numbers of loci sampled in all individuals.
Data from: Genomic variation underlying complex life history traits revealed by genome sequencing in Chinook salmon
A broad portfolio of phenotypic diversity in natural organisms can buffer against exploitation and increase species persistence in disturbed ecosystems. The study of genomic variation that accounts for ecological and evolutionary adaptation can represent a powerful approach to extend understanding of phenotypic variation in nature. Here we present a chromosome-level reference genome assembly for Chinook salmon (Oncorhynchus tshawytscha; 2.36 Gb) that enabled association mapping of life history variation and phenotypic traits for this species. Whole genome resequencing of populations with distinct life history traits provided evidence that divergent selection was extensive throughout the genome within and among phylogenetic lineages, indicating a broad portfolio of phenotypic diversity exists in this species that is related to local adaptation and life history variation. Association mapping with millions of genome-wide SNPs revealed that a genomic region of major effect on chromosome 28 was associated with phenotypes for premature and mature arrival to spawning grounds and was consistent across three distinct phylogenetic lineages. Our results demonstrate how genomic resources can enlighten the genetic basis of known phenotypes in exploited species and assist in clarifying phenotypic variation that may be difficult to observe in naturally occurring organisms.
Data from: Genomic-scale capture and sequencing of endogenous DNA from feces
Genomic-level analyses of DNA from non-invasive sources would facilitate powerful conservation and evolutionary studies in natural populations of endangered and otherwise elusive species. However, the typical low quantity and poor quality of DNA that is extracted from non-invasive samples have generally precluded such work. Here we apply a modified DNA capture protocol that, when used in combination with massively-parallel sequencing technology, facilitates efficient and highly-accurate resequencing of megabases of specified nuclear genomic regions from fecal DNA samples. We validated our approach by comparing genetic variants identified from corresponding fecal and blood DNA samples of six western chimpanzees (Pan troglodytes verus) across more than 1.5 megabases of chromosome 21, chromosome X, and the complete mitochondrial genome. Our results suggest that it is now feasible to conduct genomic studies in natural populations for which constraints on invasive sampling have otherwise long been a barrier. The data we collected also provided an opportunity to examine western chimpanzee genetic diversity at unprecedented scale. Despite high mitochondrial genome diversity (pi = 0.585%), western chimpanzees have a low ratio (0.42) of X chromosomal (pi = 0.034%) to autosomal (chromosome 21 pi = 0.081%) sequence diversity, a pattern that may reflect an unusual demographic history of this subspecies.
Data from: Development of an Arabis alpina genomic contig sequence dataset and application to single nucleotide polymorphisms discovery
The alpine plant Arabis alpina is an emerging model in the ecological genomic field which is well-suited to identifying the genes involved in local adaptation in contrasted environmental conditions, a subject which remains poorly understood at molecular level. This paper presents the assembly of a pool of A. alpina genomic fragments using Next Generation Sequencing technologies. These contigs cover 172 Mb of the A. alpina genome (i.e. 50% of the genome) and were shown to contain sequences giving positive hits against 96% of the 458 CEGMA core genes (Core Eukaryotic Genes Mapping Approach), a set of highly conserved eukaryotic genes. Regions presenting high nucleic sequence identity with 77% of the close relative Arabidopsis thaliana's genes were found, with an unbiased distribution across the different functional categories of A. thaliana genes. This new resource was tested using a resequencing assay to identify polymorphic sites. Sixteen samples were successfully analyzed and 127,041 Single Nucleotide Polymorphisms identified. This contig dataset will contribute to improving understanding of the ecology of Arabis alpina, thus constituting an important resource for future ecological genomic studies.
Data from: Assembly and comparative analysis of transposable elements from low coverage genomic sequence data in Asparagales
The research field of comparative genomics is moving from a focus on genes to a more holistic view including the repetitive complement. This study aimed to characterize relative proportions of the repetitive fraction of large, complex genomes in a non-model system. The monocotyledonous plant order Asparagales (onion, asparagus, agave) comprises some of the largest angiosperm genomes and represents variation in both genome size and structure (karyotype). Anonymous, low coverage, single-end Illumina data from eleven exemplar Asparagales taxa were assembled using a de novo method. Resulting contigs were annotated using a reference library of available monocot repetitive sequences. Mapping reads to contigs provided rough estimates of relative proportions of each type of transposon in the nuclear genome. The results were parsed into general repeat types and synthesized with genome size estimates and a phylogenetic context to describe the pattern of transposable element evolution among these lineages. The major finding is that while some lineages in Asparagales exhibit conservation in repeat proportions, there is generally wide variation in types and frequency of repeats. This approach is an appropriate first step in characterizing repeats in evolutionary lineages with a paucity of genomic resources.
Data from: Genome assembly improvement and mapping convergently evolved skeletal traits in sticklebacks with genotyping-by-sequencing
Marine populations of the threespine stickleback (Gasterosteus aculeatus) have repeatedly colonized and rapidly adapted to freshwater habitats, providing a powerful system to map the genetic architecture of evolved traits. Here, we developed and applied a binned genotyping-by-sequencing (GBS) method to build dense genome-wide linkage maps of sticklebacks using two large marine by freshwater F2 crosses of more than 350 fish each. The resulting linkage maps significantly improve the genome assembly by anchoring 78 new scaffolds to chromosomes, reorienting 40 scaffolds, and rearranging scaffolds in 4 locations. In the revised genome assembly, 94.6% of the assembly was anchored to a chromosome. To assess linkage map quality, we mapped quantitative trait loci (QTL) controlling lateral plate number, which mapped as expected to a 200-kb genomic region containing Ectodysplasin, as well as a chromosome 7 QTL overlapping a previously identified modifier QTL. Finally, we mapped eight QTL controlling convergently evolved reductions in gill raker length in the two crosses, which revealed that this classic adaptive trait has a surprisingly modular and nonparallel genetic basis.
Data from: The Solanum commersonii genome sequence provides insights into adaptation to stress conditions and genome evolution of wild potato relatives
Here, we report the draft genome sequence of Solanum commersonii, which consists of ∼830 megabases with an N50 of 44,303 bp anchored to 12 chromosomes, using the potato (Solanum tuberosum) genome sequence as a reference. Compared with potato, S. commersonii shows a striking reduction in heterozygosity (1.5% versus 53 to 59%), and differences in genome sizes were mainly due to variations in intergenic sequence length. Gene annotation by ab initio prediction supported by RNA-seq data produced a catalog of 1703 predicted microRNAs, 18,882 long noncoding RNAs of which 20% are shown to target cold-responsive genes, and 39,290 protein-coding genes with a significant repertoire of nonredundant nucleotide binding site-encoding genes and 126 cold-related genes that are lacking in S. tuberosum. Phylogenetic analyses indicate that domesticated potato and S. commersonii lineages diverged ∼2.3 million years ago. Three duplication periods corresponding to genome enrichment for particular gene families related to response to salt stress, water transport, growth, and defense response were discovered. The draft genome sequence of S. commersonii substantially increases our understanding of the domesticated germplasm, facilitating translation of acquired knowledge into advances in crop stability in light of global climate and environmental changes.
Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C
The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.
Data from: The evolutionary history of Xiphophorus fish and their sexually selected sword: a genome-wide approach using restriction site-associated DNA sequencing
Next-generation sequencing (NGS) techniques are now key tools in the detection of population genomic and gene expression differences in a large array of organisms. However, so far few studies have utilized such data for phylogenetic estimations. Here, we use NGS data obtained from genome-wide restriction site-associated DNA (RAD) (∼66000 SNPs) to estimate the phylogenetic relationships among all 26 species of swordtail and platyfish (genus Xiphophorus) from Central America. Past studies, both sequence and morphology-based, have differed in their inferences of the evolutionary relationships within this genus, particularly at the species-level and among monophyletic groupings. We show that using a large number of markers throughout the genome, we are able to infer the phylogenetic relationships with unparalleled resolution for this genus. The relationships among all three major clades and species within each of them are highly resolved and consistent under maximum likelihood, Bayesian inference and maximum parsimony. However, we also highlight the current cautions with this data type and analyses. This genus exhibits a particularly interesting evolutionary history where at least two species may have arisen through hybridization events. Here, we are able to infer the paternal lineages of these putative hybrid species. Using the RAD-marker-based tree we reconstruct the evolutionary history of the sexually selected sword trait and show that it may have been present in the common ancestor of the genus. Together our results highlight the outstanding capacity that RAD sequencing data has for resolving previously problematic phylogenetic relationships, particularly among relatively closely related species.
Data from: Genome-wide identification of microsatellites and transposable elements in the dromedary camel genome using whole genome sequencing data
Transposable elements (TEs) along with simple sequence repeats (SSRs) are prevalent in eukaryotic genome, especially in mammals. Repetitive sequences form approximately one-third of the camelid genomes, so study on this part of genome can be helpful in providing deeper information from the genome and its evolutionary path. Here, in order to improve our understanding regarding the camel genome architecture, the whole genome of the two dromedaries (Yazdi and Trodi camels) was sequenced. Totally, 92- and 84.3-Gb sequence data were obtained and assembled to 137,772 and 149,997 contigs with a N50 length of 54,626 and 54,031 bp in Yazdi and Trodi camels, respectively. Results showed that 30.58% of Yazdi camel genome and 30.50% of Trodi camel genome were covered by TEs. Contrary to the observed results in the genomes of cattle, sheep, horse, and pig, no endogenous retrovirus-K (ERVK) elements were found in the camel genome. Distribution pattern of DNA transposons in the genomes of dromedary, Bactrian, and cattle was similar in contrast with LINE, SINE, and long terminal repeat (LTR) families. Elements like RTE-BovB belonging to LINEs family in cattle and sheep genomes are dramatically higher than genome of dromedary. However, LINE1 (L1) and LINE2 (L2) elements cover higher percentage of LINE family in dromedary genome compared to genome of cattle. Also, 540,133 and 539,409 microsatellites were identified from the assembled contigs of Yazdi and Trodi dromedary camels, respectively. In both samples, di-(393,196) and tri-(65,313) nucleotide repeats contributed to about 42.5% of the microsatellites. The findings of the present study revealed that non-repetitive content of mammalian genomes is approximately similar. Results showed that 9.1 Mb (0.47% of whole assembled genome) of Iranian dromedary's genome length is made up of SSRs. Annotation of repetitive content of Iranian dromedary camel genome revealed that 9,068 and 11,544 genes contain different types of TEs and SSRs, respectively. SSR markers identified in the present study can be used as a valuable resource for genetic diversity investigations and marker-assisted selection (MAS) in camel-breeding programs.
Data from: "NGS based generation of expressed sequence tags for Lymantria dispar and Lymantria monacha, two closely related lepidopteran species with different responses to parasitism by Glyptapanteles liparidis" in Genomic Resources Notes accepted 1 December 2013 to 31 January 2014
Introduction: The gypsy moth, Lymantria dispar, and the nun moth, Lymantria monacha, are closely related species (Lepidoptera, Lymantriidae), co-seasonal and economically important forest pests on broadleaf and coniferous trees. In Central Europe, gypsy moth larvae are frequently parasitized by the gregarious, endoparasitic wasp Glyptapanteles liparidis (Hymenoptera, Braconidae). At oviposition, the female wasp injects between 10 and up to 100 eggs into the hemocoel of a single host larva, together with venom and calyx fluid containing polydnavirus (PDV) particles that subsequently play a critical role in suppressing the host immune response so that successful development of the parasitoid can proceed (Schopf 2007). These viruses, which are integrated in the genomic DNA of the wasp and undergo replication only in the female's ovary, rapidly enter host hemocytes, fat body, and nervous system following parasitization, and viral genes are expressed. In L. dispar larvae parasitized by G. liparidis, the host's hemocytes alter their behavior, fail to spread properly (thereby inhibiting the encapsulation response) and partly undergo programmed cell death (apoptosis), resulting in a dramatic drop in the host's total hemocyte number (Schafellner and Schläger 2009).
Data from: Genome sequencing and comparative analysis of three Chlamydia pecorum strains associated with different pathogenic outcomes
Background: Chlamydia pecorum is the causative agent of a number of acute diseases, but most often causes persistent, subclinical infection in ruminants, swine and birds. In this study, the genome sequences of three C. pecorum strains isolated from the faeces of a sheep with inapparent enteric infection (strain W73), from the synovial fluid of a sheep with polyarthritis (strain P787) and from a cervical swab taken from a cow with metritis (strain PV3056/3) were determined using Illumina/Solexa and Roche 454 genome sequencing. Results: Gene order and synteny was almost identical between C. pecorum strains and C. psittaci. Differences between C. pecorum and other chlamydiae occurred at a number of loci, including the plasticity zone, which contained a MAC/perforin domain protein, two copies of a >3400 amino acid putative cytotoxin gene and four (PV3056/3) or five (P787 and W73) genes encoding phospholipase D. Chlamydia pecorum contains an almost intact tryptophan biosynthesis operon encoding trpABCDFR and has the ability to sequester kynurenine from its host, however it lacks the genes folA, folKP and folB required for folate metabolism found in other chlamydiae. A total of 15 polymorphic membrane proteins were identified, belonging to six pmp families. Strains possess an intact type III secretion system composed of 18 structural genes and accessory proteins, however a number of putative inc effector proteins widely distributed in chlamydiae are absent from C. pecorum. Two genes encoding the hypothetical protein ORF663 and IncA contain variable numbers of repeat sequences that could be associated with persistence of infection. Conclusions: Genome sequencing of three C. pecorum strains, originating from animals with different disease manifestations, has identified differences in ORF663 and pseudogene content between strains and has identified genes and metabolic traits that may influence intracellular survival, pathogenicity and evasion of the host immune system.
Data from: Deconstruction of archaeal genome depict strategic consensus in core pathways coding sequence assembly
A comprehensive in silico analysis of 71 species representing the different taxonomic classes and physiological genre of the domain Archaea was performed. These organisms differed in their physiological attributes, particularly oxygen tolerance and energy metabolism. We explored the diversity and similarity in the codon usage pattern in the genes and genomes of these organisms, emphasizing on their core cellular pathways. Our thrust was to figure out whether there is any underlying similarity in the design of core pathways within these organisms. Analyses of codon utilization pattern, construction of hierarchical linear models of codon usage, expression pattern and codon pair preference pointed to the fact that, in the archaea there is a trend towards biased use of synonymous codons in the core cellular pathways and the Nc-plots appeared to display the physiological variations present within the different species. Our analyses revealed that aerobic species of archaea possessed a larger degree of freedom in regulating expression levels than could be accounted for by codon usage bias alone. This feature might be a consequence of their enhanced metabolic activities as a result of their adaptation to the relatively O2-rich environment. Species of archaea, which are related from the taxonomical viewpoint, were found to have striking similarities in their ORF structuring pattern. In the anaerobic species of archaea, codon bias was found to be a major determinant of gene expression. We have also detected a significant difference in the codon pair usage pattern between the whole genome and the genes related to vital cellular pathways, and it was not only species-specific but pathway specific too. This hints towards the structuring of ORFs with better decoding accuracy during translation. Finally, a codon-pathway interaction in shaping the codon design of pathways was observed where the transcription pathway exhibited a significantly different coding frequency signature.
Supplementary dataset for "Plasticity of repetitive sequences demonstrated by the complete mitochondrial genome of Eucalyptus camaldulensis"
Open the record for dataset details and reuse information.
Whole genome sequencing data for Saccharomyces cerevisiae CBS 493.94
<p>SNP distance matrix.</p>
FIG. 3 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations
FIG. 3.—Populationstructure inference based on STRUCTURE analysis of 199,921 sites for individual bats from four Hawaiian Islands. (A) Ad hoc statistic delta Kanalysis indicates a peak at the Κ = 5; (B) STRUCTURE population inference with Κ = 3, 4, 5. Sample information included in supplementary table S4, Supplementary Material online.
Draft genome sequence of Xylaria bambusicola isolate GMP-LS, the root and basal stem rot pathogen of sugarcane in Indonesia
<p>Supplementary Figure 1</p><p>Supplementary Table 1</p><p>Supplementary Table 2</p>
Whole genome sequencing of two Canine Herpesvirus 1 (CHV-1) isolates and clinicopathological outcomes of infection in French Bulldog puppies
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.