Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Data from: Evolutionary and phylogenetic insights from a nuclear genome sequence of the extinct, giant subfossil koala lemur Megaladapis edwardsi
<p><span>No endemic Madagascar animal with body mass >10 kg survived a relatively recent wave of extinction on the island. From morphological and isotopic analyses of skeletal 'subfossil' remains we can reconstruct some of the biology and behavioral ecology of giant lemurs (primates; up to ~160 kg), elephant birds (up to ~860 kg), and other extraordinary Malagasy megafauna that survived well into the past millennium. Yet much about the evolutionary biology of these now extinct species remains unknown, along with persistent phylogenetic uncertainty in some cases. Thankfully, despite the challenges of DNA preservation in tropical and sub-tropical environments, technical advances have enabled the recovery of ancient DNA from some Malagasy subfossil specimens. Here we present a nuclear genome sequence (~2X coverage) for one of the largest extinct lemurs, the koala lemur <i>Megaladapis edwardsi </i>(~85kg). To support the testing of key phylogenetic and evolutionary hypotheses we also generated new high-coverage complete nuclear genomes for two extant lemur species, <i>Eulemur rufifrons</i> and <i>Lepilemur mustelinus</i>, and we aligned these sequences with previously published genomes for three other extant lemur species and 47 non-lemur vertebrates. Our phylogenetic results confirm that <i>Megaladapis</i> is most closely related to the extant Lemuridae (typified in our analysis by <i>E. rufifrons</i>) to the exclusion of <i>L. mustelinus</i>, which contradicts morphology-based phylogenies. Our evolutionary analyses identified significant convergent evolution between <i>M. edwardsi</i> and extant folivorous primates (colobine monkeys) and ungulate herbivores (horses) in genes encoding protein products that function in the biodegradation of plant toxins and nutrient absorption. These results suggest that koala lemurs were highly adapted to a leaf-based diet, which may also explain their convergent craniodental morphology with the small-bodied folivore <i>Lepilemur</i>.</span></p>
Raw sequencing metrics from two different prep methods for obtaining SARS-CoV2 genomes
<p>Sequencing_metrics_IlluminaDNAprep.xlsx contains metrics used for the determination of SARS-CoV2 genome assembly success using Illumina's DNA prep adapter-tagmentation method, coupled with the ARTIC protocol for gene-specific amplification of SARS-CoV2</p> <p>Sequencing_metrics_Seqwellprep.xlsx contains metrics used for the determination of SARS-CoV2 genome assembly success using Seqwell's Plexwell384 adapter-tagmentation method, coupled with the ARTIC protocol for gene-specific amplification of SARS-CoV2</p>
The complete genome sequence and comparative genomic analyses of four phages (NJ-P3, NB-P21, NC-P34 and NN-P42)
<p>We downloaded and reanalyzed the raw data of four phages genomes (NJ-P3, NB-P21, NC-P34, NN-P42). This is the reassembled whole genomes and comparative genomic analyses of four phages.</p>
Two leaves that cannot die: the genome sequence of Welwitschia mirabilis reveals its unique biology and evolutionary history
<p>Welwitschia mirabilis (hereafter Welwitschia), the sole species in Welwitschiales, belongs to gnetophytes, an ancient, enigmatic gymnosperm lineage. It is a strikingly bizarre plant with distinctive morphology of just two large ever-elongating leaves and is remarkable in being able to survive extreme environmental stresses of the Namibian and Angolan deserts. Here, we provide a chromosome-level assembly of its genome (6.8 Gb/1C) and extensive methylome and transcriptome data to reveal the genetics underpinning its intriguing biology. The Welwitschia genome has been shaped by a lineage-specific ancient whole genome duplication ~ 86 million years ago, and more recently (within 10 million years) by bursts of retrotransposon activity. In addition, high levels of cytosine methylation, extremely so for CHH motifs, are associated with retrotransposons, whilst their long-term deamination has resulted in an exceptionally GC-poor genome. High levels of methylation are likely to be responses to maintain genomic integrity in the face of stress-induced retroelement mobility while reduced GC content will confer a genomic advantage under nutrient limitation. Changes in the copy number and/or expression of key gene families and specific transcription factors (e.g. R2R3MYB, SAUR) controlling cell growth, differentiation and metabolism underpin the plant's extreme longevity under increasing temperature, nutrient and water stress. The Welwitschia chromosome level assembly here, along with a new high-quality assembly for Gnetum montanum, enhances our understanding of genome evolution in gnetophytes. It also provides critical new insights into the extraordinary development of Welwitschia's ever-growing leaves, enabling its survival in such hostile conditions.</p>
Data from: The genome sequence and insights into the immunogenetics of the bananaquit (Passeriformes: Coereba flaveola)
Avian genomics, especially of non-model species, is in its infancy relative to mammalian genomics. Here, we describe the sequencing, assembly, and annotation of a new avian genome, that of the bananaquit Coereba flaveola (Passeriformes: Thraupidae). We produced ∼30-fold coverage of the genome with an assembly size of ca. 1.2 Gb, including approximately 16,500 annotated genes. Passerine birds, such as the bananaquit, are commonly infected by avian malarial parasites (Haemosporida), which presumably drive adaptive evolution of immunogenetic loci within the host genome. In the context of our research on the distribution of avian Haemosporida, we specifically characterized immune loci, including toll-like receptor (TLR) and major histocompatibility complex (MHC) genes. Additionally, we identified novel molecular markers in the form of single nucleotide polymorphisms (SNPs), both genome-wide and within identified immune loci. We discovered nine TLR genes and four MHC genes and identified five other TLR- or MHC- associated genes. Genome-wide, over 6 million high-quality SNPs were annotated, including 568 within TLR genes and 102 in MHC genes. This newly described genome and immune characterization expands the knowledge base for avian genomics and phylogenetics and allows for immune genotyping in the bananaquit, providing tools for the investigation of host-parasite coevolution.
Data from: Genotyping by sequencing and genome–environment associations in wild common bean predict widespread divergent adaptation to drought
Drought will reduce global crop production by >10% in 2050 substantially worsening global malnutrition. Breeding for resistance to drought will require accessing crop genetic diversity found in the wild accessions from the driest high stress ecosystems. Genome–environment associations in crop wild relatives reveal natural adaptation, and therefore can be used to identify adaptive variation. We explored this approach in the food crop Phaseolus vulgaris L., characterizing 86 geo-referenced wild accessions using Genotyping by Sequencing (GBS) to discover single-nucleotide-polymorphisms (SNPs). The wild beans represented Mesoamerica, Guatemala, Colombia, Ecuador/Northern Peru and Andean groupings. We found high polymorphism with a total of 22,845 SNPs across the 86 accessions loci that confirmed genetic relationships for the groups. As a second objective, we quantified allelic associations with a bioclimatic-based drought index using 10 different statistical models that accounted for population structure. Based on the optimum model, 115 SNPs in 90 regions, widespread in all 11 common bean chromosomes, were associated with the bioclimatic-based drought index. A gene coding for an Ankyrin repeat-containing protein and a phototropic-responsive NPH3 gene were identified as potential candidates. Genomic windows of 1Mb containing associated SNPs had more positive Tajima's D scores than windows without associated markers. This indicates that adaptation to drought, as estimated by bioclimatic variables, has been under natural divergent selection, suggesting that drought tolerance may be favorable under dry conditions but harmful in humid conditions. Our work exemplifies that genomic signatures of adaptation are useful for germplasm characterization, potentially enhancing future marker-assisted selection and crop improvement.
Ultra-deep sequencing of HIV-1 near full-length and partial proviral genomes reveals high genetic diversity among Brazilian blood donors
<p>Here, we aimed to gain a comprehensive picture of the HIV-1 diversity in the northeast and southeast part of Brazil. To this end, a high-throughput sequencing-by-synthesis protocol and instrument were used to characterize the near full length (NFLG) and partial HIV-1 proviral genome in 259 HIV-1 infected blood donors at four major blood centers in Brazil: Pro-Sangue foundation (São Paulo state (SP), n 51), Hemominas foundation (Minas Gerais state (MG), n 41), Hemope foundation (Recife state (PE), n 96) and Hemorio blood bank (Rio de Janeiro (RJ), n 70).</p>
ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data. Additional Datasets.
<p>Data supporting the main figures of the "ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data" manuscript and a copy of the ParaMask software and scripts for analysis, and SV calls from longreads. README files are included.</p>
Supplementary data 'Mitochondrial genome sequence of the protist Ancyromonas sigmoides Kent, 1881 (Ancyromonadida) from the Sugluk Inlet, Hudson Strait, Nunavik, Québec'
<p>Fasta file with the transcripts obtained from RNAseq sequencing of Ancryomonas sigmoides (kmer 35) and file with datamining results.</p><p>Fasta file of the transcript matching the cox1 gene.</p><p>Scaffolds of the kmer 85 assembly of the genomic data of Ancryomonas sigmoides.</p><p>Databse used for datamining (fasta file)</p>
Next-generation Sequencing Data Associated with "Genome Editing Outcomes Reveal Mycobacterial NucS Participates in a Short-Patch Repair of DNA Mismatches"
Open the record for dataset details and reuse information.
MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects
<p>This dataset represents all results files described in the paper 'MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects'.</p>
A root-specific NLR network confers resistance to plant parasitic nematodes - genomic sequences and annotations
<p>Sequence and annotation data associated with "A root-specific NLR network confers resistance to plant parasitic nematodes"</p>
A pan-cetacean MHC amplicon sequencing panel developed and evaluated in combination with genome assemblies
<p>The major histocompatibility complex (MHC) is a highly polymorphic gene family that is crucial in immunity, and its diversity can be effectively used as a fitness marker for populations. Despite this, MHC remains poorly characterised in non-model species (e.g., cetaceans: whales, dolphins and porpoises) as high gene copy number variation, especially in the fast-evolving class I region, makes analyses of genomic sequences difficult. To date, only small sections of class I and IIa genes have been used to assess functional diversity in cetacean populations. Here, we undertook a systematic characterisation of the MHC class I and IIa regions in available cetacean genomes. We extracted full-length gene sequences to design pan-cetacean primers that amplified the complete exon2 from MHC class I and IIa genes in one combined sequencing panel. We validated this panel in 19 cetacean species and described 354 alleles for both classes. Furthermore, we identified likely assembly artefacts for many MHC class I assemblies based on the presence of class I genes in the amplicon data compared to missing genes from genomes. Finally, we investigated MHC diversity using the panel in 25 humpback and 30 southern right whales, including four paternity trios for humpback whales. This revealed copy-number variable class I haplotypes in humpback whales, which is likely a common phenomenon across cetaceans. These MHC alleles will form the basis for a cetacean branch of the Immuno-Polymorphism Database (IPD-MHC), a curated resource intended to aid in the systematic compilation of MHC alleles across several species, to support conservation initiatives.</p>
The alignments of chloroplast genome sequences and nuclear ribosomal DNA fragments of six oak species sampled in the hot-dry valley of the Jinsha River, southwestern China
<p>Both chloroplast (cp) genome sequences and nuclear ribosomal (nr) DNA were assembled using GetOrganelle v.1.7.6.1 for 18 oak trees sampled in the Panzhihua Cycad National Nature Reserve, Sichuan Province, China. These trees belong to six oak species, including Quercus cocciferoides, Q. dolicholepis, Q. franchetii, Q. griffithii, Q. longispica, and Q. variabilis. We used PhyloSuite v.1.1.152 to extract coding sequences (CDSs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) of the 18 oak cp genomes. These sequences were aligned separately using MAFFT v.7.3.13 and manually adjusted with BioEdit v.7.2.5. Length variations in mononucleotide repeats were excluded and inversions were replaced with their reverse complements because of their tendency for homoplasy. Other indels were coded as binary characters according to the simple gap coding method using GapCoder. Separate assignments were concatenated according to their respective positions in the cp genome to obtain the alignments of LSC, SSC, IRb, and the whole cp genome.</p>
Data from: Genomic footprint of cladogenesis revealed through RADseq and Sanger sequencing demonstrates congruent patterns in the velvet worm Peripatopsis sedgwicki species complex (Onychophora: Peripatopsidae)
<p>In the present study, first generation DNA sequencing (mitochondrial cytochrome c oxidase subunit one, <em>COI</em>) and reduced-representative genomic RADseq data were used to understand the patterns and processes of diversification of the velvet worm, <em>Peripatopsis sedgwicki</em> species complex across its distribution range in South Africa. For the RADseq data, three datasets (two primary and one supplementary) were generated corresponding to 1259 - 11,468 SNPs, in order to assess the species diversity and phylogeographic of the species complex. Tree topologies for the two primary datasets were inferred using maximum likelihood and Bayesian inferences methods. Phylogenetic analyses using the <em>COI </em>datasets retrieved four distinct, statistically well-supported clades within the species complex. Five species delimitation methods applied to the <em>COI </em>data (ASAP, bPTP, bGMYC, STACEY, and iBPP) all showed support for the distinction of the Fort Fordyce Nature Reserve specimens. In the main <em>P. sedgwicki </em>species complex, the species delimitation methods revealed a variable number of operational taxonomic units and overestimated the number of putative taxa. Divergence time estimates coupled with the geographic exclusivity of species and phylogeographic results suggest recent cladogenesis during the Plio/Pleistocene. The RADseq were subjected to a principal components analysis and a discriminant analysis of principal components, under a maximum-likelihood framework. The latter results corroborate the four main clades observed using the <em>COI</em> data, however, applying additional filtering revealed additional diversity. The high overall congruence observed between the RADseq and <em>COI </em>data suggests that first generation sequence data remain a cheap and effective method for evolutionary studies, although RADseq does provide a far greater resolution of contemporary temporo-spatial patterns. </p>
Genomics polymorphisms of Staphylococcus aureus strain NCTC 8325 in the lab stock maintained at TUM (WT), after 30 passes in BHI media (D) and after 30 passes detecting 4 -fold MIC increase to isocyanide -code I16- 3 biological replicates (A,B,C), and 3 independent colonies sequenced per replicate at the end of the experiment.
<p>Genomics polymorphisms of Staphylococcus aureus strain NCTC 8325 in the lab stock maintained at TUM (WT), after 30 passes in BHI media (D) and after 30 passes detecting 4 -fold MIC increase to isocyanide -code I16- 3 biological replicates (A,B,C), and 3 independent colonies sequenced per replicate at the end of the experiment. Determined from Illumina shotgun genomic sequencing datasets, mapping and analyses vs the reference genome of the strain https://www.ncbi.nlm.nih.gov/nuccore/NC_007795.1/</p>
Data required for "Low mutation rate of spontaneous mutants enables detection of causative genes by comparing whole genome sequences"
<p>In the early 1900s,mutation breeding to select varieties with desirable traits using spontaneous mutation was actively conducted around the world, including Japan. In rice, the number of fixed mutations per generation was estimated to be 1.38-2.25. Although this low mutation rate was a major problem for breeding in those days, in the modern era with the development of NGS technology, it was conversely considered to be an advantage for efficient gene identification. In this paper, we proposed an in silico approach using next-generation sequencing (NGS) to compare the whole genome sequence of a spontaneous mutant with that of a closely related strain with a nearly identical genome, to find polymorphisms that differ between them, and to identify the causal gene by predicting the functional variation of the gene caused by the polymorphism. Using this approach, we found four causal genes for the dwarf mutation, the round shape grain mutation and the awnless mutation. Three of these genes were the same as those previously reported, but one was a novel gene involved in awn formation. The novel gene was isolated from Bozu-Aikoku, a mutant of Aikoku with the awnless trait, in which nine polymorphisms were predicted to alter gene function by their whole-genome comparison. Based on the information on gene function and tissue-specific expression patterns of these candidate genes, Os03g0115700/LOC_Os03g02460, annotated as a shortchain dehydrogenase/reductase SDR family protein, is most likely to be involved in the awnless mutation. Indeed, complementation tests by transformation showed that it is involved in awn formation. Thus, this method is an effective way to accelerate genome breeding of various crop species by enabling the identification of useful genes that can be used for crop breeding with minimal effort for NGS analysis.</p>
Whole genome sequencing (WGS) data from invasive pine sawfly Diprion similis
<p>Biological introductions are unintended "natural experiments" that provide unique insights into evolutionary processes. Invasive phytophagous insects are of particular interest to evolutionary biologists studying adaptation, as introductions often require rapid adaptation to novel host plants. However, adaptive potential of invasive populations may be limited by reduced genetic diversity—a problem known as the "genetic paradox of invasions". One potential solution to this paradox is if there are multiple invasive waves that bolster genetic variation in invasive populations. Evaluating this hypothesis requires characterizing genetic variation and population structure in the invaded range. To this end, we assemble a reference genome and describe patterns of genetic variation in the introduced white pine sawfly, <em>Diprion</em> <em>similis</em>. This species was introduced to North America in 1914, where it has rapidly colonized the thin-needled eastern white pine (<em>Pinus</em> <em>strobus</em>), making it an ideal invasion system for studying adaptation to novel environments. To evaluate evidence of multiple introductions, we generated whole-genome resequencing data for 64 <em>D</em>. <em>similis</em> females sampled across the North American range. Both model-based and model-free clustering analyses supported a single population for North American <em>D</em>. <em>similis</em>. Within this population, we found evidence of isolation-by-distance and a pattern of declining heterozygosity with distance from the hypothesized introduction site. Together, these results support a single-introduction event. We consider implications of these findings for the genetic paradox of invasion and discuss priorities for future research in <em>D</em>. <em>similis</em>, a promising model system for invasion biology.</p>
Complete mitochondrial genome sequence of the Atlantic Mudskipper (Periophthalmus barbarus) (Linnaeus, 1766) (Perciformes: Gobiidae)
<p>The complete mitochondrial genome of the Atlantic mudskipper (<em>Periophthalmus barbarus</em>) was determined in this study. The specimen was collected from Abonnema, Nigeria (4.73075, 6.77565). The complete mitogenome sequence of <em>P. barbarus</em> would be useful for further studies on molecular phylogenetic relationship and population genetics of the subfamily Oxudercinae.</p>
HiFi Metagenomic Sequencing Enables Assembly of Accurate and Complete Genomes from Human Gut Microbiota.
<p>We reported 102 complete metagenome assembled genomes (cMAGs) from five human fecal HiFi sequencing samples.</p> <p>102_cMAGs_fna.tar.gz: Fasta sequence files of 102 cMAGs.</p> <p>gc_skew_figures.tar.gz: GC-skew pattern figures of 102 cMAGs. (SVG format)</p> <p>coverage_plots.tar.gz: Genome coverage plot of 102 cMAGs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.