Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
A chromosome-scale high-contiguity genome assembly of the threatened cheetah (Acinonyx jubatus)
<p><span>The cheetah (<em>Acinonyx</em> <em>jubatus</em>, SCHREBER 1775) is a large felid and is considered the fastest land animal. Historically, it inhabited open grassland across Africa, the Arabian Peninsula, and southwestern Asia; however, only small and fragmented populations remain today. Here, we present a de novo genome assembly of the cheetah based on PacBio continuous long reads and Hi-C proximity ligation data. The final assembly (VMU_Ajub_asm_v1.0) has a total length of 2.38 Gb, of which 99.7% are anchored into the expected 19 chromosome-scale scaffolds. The contig and scaffold N50 values of 96.8 Mb and 144.4 Mb, respectively, a BUSCO completeness of 95.4% and a k-mer completeness of 98.4%, emphasize the high quality of the assembly. Furthermore, annotation of the assembly identified 23,622 genes and a repeat content of 40.4%. This new highly contiguous and chromosome-scale assembly will greatly benefit conservation and evolutionary genomic analyses and will be a valuable resource, e.g., to gain a detailed understanding of the function and diversity of immune response genes in felids.</span></p>
Dryococelus australis genome assembly supplementary data
<p>We present a chromosome-scale genome assembly for the critically endangered Lord Howe Island stick insect <em>Dryococelus australis. </em>Contained in this repository are the original unfiltered annotation .gff file, the .fasta file of repeat families identified by RepeatScout, and the .gff file of repetitive elements throughout the genome assembly.</p>
Chloroplast genome assemblies and comparative analyses of commercially important Vaccinium berry crops
<p><em>Vaccinium</em> is a large genus of shrubs that includes a handful of economically important berry crops. Given the numerous hybridizations and polyploidization events, the taxonomy of this genus has remained the subject of long debate. In addition, berries and berry-based products are liable to adulteration, either fraudulent or unintentional due to misidentification of species. The availability of more genomic information could help achieve higher phylogenetic resolution for the genus, provide molecular markers for berry crop identification, and a framework for efficient genetic engineering of chloroplasts. Therefore, in this study, we assembled five <em>Vaccinium</em> chloroplast sequences representing the economically relevant berry types: northern highbush blueberry (<em>V. corymbosum</em>), southern highbush blueberry (<em>V. corymbosum</em> hybrids), rabbiteye blueberry (<em>V. virgatum</em>), lowbush blueberry (<em>V. angustifolium</em>), and bilberry (<em>V. myrtillus</em>). Comparative analyses showed that the <em>Vaccinium</em> chloroplast genomes exhibited an overall highly conserved synteny and sequence identity among them. Polymorphic regions included the expansion/contraction of inverted repeats, gene copy number variation, simple sequence repeats, indels, and single nucleotide polymorphisms. Based on their in silico discrimination power, we suggested variants that could be developed into molecular markers for berry crop identification. Phylogenetic analysis revealed multiple origins of highbush blueberry plastomes, likely due to the hybridization events that occurred during northern and southern highbush blueberry domestication.</p>
Oceanic Prokaryotes Metagenome-Assembled Genomes reconstructed using metagenomic distances
<p>Sets of reconstructed Metagenome-Assembled Genomes (MAGs) from Tara Oceans dataset. The reconstructed MAGs belong to Magneto paper: https://doi.org/10.1128/msystems.00432-22</p> <p>The dataset is composed of 93 oceanic metagenomes sampled from non-polar oceanic regions.</p> <p>The Metagenomic Distance MAGs were reconstructed following a co-assembly protocol driven by nucleotidic composition similarity, as detailed in the publication.</p> <p>The Oceanic Region MAGs were reconstructed by co-assembly of samples belonging to the same Oceanic Regions.</p> <p>The file clusters.tsv sum up the metagenomic distance cluster and the oceanic region each sample belongs to.</p>
Chromosome-scale genome assembly of the African spiny mouse (Acomys cahirinus)
<p>Genomic DNA was extracted from blood from a single male A. cahirinus animal using a Monarch HMW DNA Extraction Kit for Cells & Blood (T3050, New England Biolabs, Ipswich MA) following the manufacturer’s recommended protocol. DNA was quantified prior to library construction using the Qubit DNA HS Assay (ThermoFischer, Waltham MA) and DNA fragment lengths were assessed using the Agilent Femto Pulse System (Santa Clara, CA). Libraries were prepared for sequencing using the Oxford Nanopore ligation kit (SQK-LSK110) following the manufacturers’ instructions, except that DNA repair and A-tailing was performed for 30 min and the ligation was allowed to continue for 1 hr. Prepared libraries were quantified using a Qubit fluorometer and 30 fmol of the library was loaded onto a Nanopore version R.9.4.1 flow cell and loaded on a PromethION running MinKNOW version (21.05.20). To increase output, the flow cell was washed after approximately 24 hr of sequencing then an additional 12 fmol of library was added to the flow cell and run for an additional 48 hr. Basecalling was performed using Guppy 5.0.12 (Oxford Nanopore) using the superior model (dna_r9.4.1_450bps_sup_prom.cfg). FASTQ files for assembly were extracted from unaligned bam files using samtools (Li et al. 2009) then Flye version 2.9 for assembly using the --nano-hq flag (Kolmogorov et al. 2019). Haplotigs and overlaps in the assembly were purged using purge_dups (https://github.com/dfguan/purge_dups). The assembly was then polished using Medaka version 1.4.2 (https://github.com/nanoporetech/medaka) followed by a second polishing step with pilon version 1.24 (Walker et al. 2014). Assembly statistics at each step were generated using Quast (Gurevich et al. 2013) and BUSCO (Simão et al. 2015) (Table S2). The primary contigs assembled from the Nanopore data were anchored to chromosomes using 505,210,505 read pairs of a Hi-C library isolated from another A. cahirinus individual of unknown sex downloaded from the NCBI Short Read Archive (SRX13258644) (Wang et al. 2022). After aligning the Hi-C reads with the ArimaHi-C Mapping Pipeline (https://github.com/ArimaGenomics/mapping_pipeline), YaHS v1.0 (Zhou et al. 2023) was used with default error correction for scaffolding, and Juicebox v1.11.08 (Dudchenko et al. 2018) was used to generate a Hi-C contact map. Progressive Cactus was used (Armstrong et al. 2020) to perform a whole-genome alignment of the A. cahirinus draft assembly to the Mus musculus GRCm39 reference genome (RefSeq GCF_000001635.27_GRCm39). Comparative annotation of the draft genomes was then performed using the Comparative Annotation Toolkit (CAT) (Fiddes et al. 2018). Briefly, the M. musculus RefSeq annotation GFF was parsed and validated with the “parse_ncbi_gff3” and “validate_gff3” programs (respectively) from CAT. The M. musculus reference transcript cDNA sequences were downloaded and mapped to the M. musculus draft genome with minimap2 (Li 2018) and provided to CAT as long-read RNA-seq reads in the “[ISO_SEQ_BAM]” field of the configuration file. For A. cahirinus, bulk RNA-seq data obtained from multiple pooled organs were downloaded from NCBI SRA BioProject PRJNA342864 (Bellofiore et al. 2017) and mapped to the draft assembly with STAR (Dobin et al. 2013) then provided to CAT in the “[BAMS]” field. CpG islands were identified using the cpg_lh utility from the UCSC suite of tools (Kent et al. 2002).</p>
Genome assemblies: Horizontal transfer of pOXA-48 from a hypervirulent Klebsiella pneumoniae ST23/KL57 to Serratia marcescens
<p>Hypervirulent Klebsiella pneumonia isolates express a range of virulence factors, often encoded on virulence plasmids. Occasionally these hypervirulent lineages acquire antimicrobial resistance genes rendering them multidrug-resistant. We describe the nosocomial transmission of a hypervirulent K. pneumoniae ST23/KL57 isolate carrying an NDM-1 and OXA-48 beta-lactamase among COVID-19 patients in a Danish university hospital. Furthermore, we characterize the plasmid structure and describe a within-patient horizontal transfer of a plasmid carrying a blaOXA-48 gene from K. pneumoniae ST23/KL57 to a Serratia marcescens isolate.</p>
A high-quality, long-read genome assembly of the whitelined sphinx moth (Lepidoptera: Sphingidae: Hyles lineata)
<p><span>The sphinx moth genus <em>Hyles</em> comprises 29 described species inhabiting all continents except Antarctica. The genus diverged relatively recently (40 – 25 mya), arising in the Americas and rapidly establishing a cosmopolitan distribution. The whitelined sphinx moth, <em>Hyles lineata</em>, represents the oldest extant lineage of this group and is one of the most widespread and abundant sphinx moths in North America. <em>Hyles lineata </em>exhibits the large body size and adept flight control characteristic of the sphinx moth family (Sphingidae), but is unique in displaying extreme larval color variation and broad host plant use. These traits, in combination with its broad distribution and high relative abundance within its range, have made <em>H. lineata</em> a strong model organism for studying phenotypic plasticity, plant-herbivore interactions, physiological ecology, and flight control. Despite being one of the most well-studied sphinx moths, little data exists on genetic variation or regulation of gene expression. Here we report a high-quality draft genome showing high contiguity (N50 of 14.2 Mb) and completeness (98.2% of Lepidoptera BUSCO genes), an important first characterization to facilitate such studies. We also annotate the core melanin synthesis pathway genes and confirm that they have high sequence conservation with other moths and are most similar to those of another, well-characterized sphinx moth, the tobacco hornworm (<em>Manduca sexta</em>).</span></p>
FUN-LDA scores for human genome assembly GRCh37
<p>FUN-LDA is based on a Latent Dirichlet Allocation (LDA) model for predicting functional effects of non-coding genetic variants in a cell type and tissue-specific way by integrating diverse epigenetic annotations for specific cell types and tissues from large-scale genomics projects such as ENCODE and Roadmap Epigenomics. Using this unsupervised approach, we predict tissue-specific functional effects for every position in the human genome for 127 tissues and cell types in ENCODE and Roadmap Epigenomics. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Format</strong></p> <p>The FUN-LDA scores are stored in the UCSC Genome Browser bigWig Track Format. </p> <p>To extract FUN-LDA scores, the bigWigAverageOverBed utility is required. It can be downloaded from the Genome Browser website at <a href="https://hgdownload.cse.ucsc.edu/admin/exe">https://hgdownload.cse.ucsc.edu/admin/exe</a>.</p> <p>User should prepare a UCSC Genome Browser bed file, a tab-separated four-column file. The first column is the chromosome; the second is the zero-based coordinate of the position of interest; the third is that zero-based coordinate plus one; and the fourth is a unique identifier for the position. The command below is an example of extracting FUN-LDA scores in tissue E007 for the positions defined in a bed file, "input.bed".</p> <pre><code class="language-bash">bigWigAverageOverBed E007.valley9.c89.bigwig input.bed output.tab</code></pre> <p>It produces a six-column file, output.tab. The first column includes the unique position id defined in the input bed file. The last column, average over the covered bases, is the score for this position.</p> <p><strong>Reference</strong></p> <p>Daniel Backenroth, Zihuai He, Krzysztof Kiryluk, Valentina Boeva, Lynn Pethukova, Ekta Khurana, Angela Christiano, Joseph Buxbaum, Iuliana Ionita-Laza. FUN-LDA: A latent Dirichlet allocation model for predicting tissue-specific functional effects of noncoding variation: Methods and applications. American Journal of Human Genetics, 2018.<br> </p> <p> </p>
MAC genome assembly and gene prediction of Tetrahymena thermophila SB210
<p>Corrected genome assembly and gene prediction of the MAC genome of T. thermophila SB210. These data were generated and analysed in the manuscript "Single-nucleotide polymorphism landscape of the macronuclear genome of <em>Tetrahymena thermophila".</em></p> <p>Please see the Material & methods and Supplementary data files of this manuscript for more details about these files.</p>
Genome annotation file containing predicted genome features of Phytophthora agathidicida (Strain: 3770, Assembly:ASM2572299v1)
<p>This is the genome annotation file (gff3) containing predicted genome features of the <em>Phytophthora agathidicida </em>(Strain: 3770) genome published in Cox et al (2022). This annotation file is associated with the following entries at Genbank:</p> <p>Assembly: ASM2572299v1<br> Biosample: SAMN19597867<br> BioProject: PRJNA734652</p> <p>Included in the file are predicted functional annotations from Blastp search of all predicted proteins sequences against the Swiss-Prot sequence database (Release 23/02).</p>
179 high quality metagenome-assembled genomes sequences and annotations
<p>We analyzed seven sediment samples collected adjacent to ferromanganese nodules from the Clarion–Clipperton Fracture Zone (CCFZ) in the eastern Pacific Ocean. Through deep metagenomic sequencing, assembly, and binning, we reconstructed 179 high quality metagenome-assembled genomes (MAGs). This archive contains these genomes sequences and annotations. </p>
Chromosome-level assemblies of the Pieris mannii butterfly genome suggest Z-origin and rapid evolution of the W chromosome
<p><span>The insect order Lepidoptera (butterflies and moths) represents the largest group of organisms with ZW/ZZ sex determination. While the origin of the Z chromosome predates the evolution of the Lepidoptera, the W chromosomes are considered younger, but their origin is debated. To shed light on the origin of the lepidopteran W, we here produce chromosome-level genome assemblies for the butterfly <em>Pieris</em> <em>mannii</em>, and compare the sex chromosomes within and between <em>P. mannii </em>and its sister species <em>P. rapae</em>. Our analyses clearly indicate a common origin of the W chromosomes of the two <em>Pieris</em> species, and reveal similarity between the Z and W in chromosome sequence and structure. This supports the view that the W in these species originates from Z-autosome fusion rather than from a redundant B chromosome. We further demonstrate the extremely rapid evolution of the W relative to the other chromosomes and argue that this may preclude reliable conclusions about the origins of W chromosomes based on comparisons among distantly related Lepidoptera. Finally, we find that sequence similarity between the Z and W chromosomes is greatest toward the chromosome ends, perhaps reflecting selection for the maintenance of recognition sites essential to chromosome segregation. Our study highlights the utility of long-read sequencing technology for illuminating chromosome evolution.</span></p>
Human genome assemblies enhanced by LOCLA
<p>This is a data repository for the genome assemblies of three human samples enhanced by LOCLA (DOI: 10.5281/zenodo.8280853 ). LOCLA is a novel genome assembly optimization tool, LOCLA, that iteratively improves the quality of an assembly by locating sequencing reads on partially assembled scaffolds and thus enable gap filling and further scaffolding. </p> <p>The three human genome assemblies and the assembly statistics are compressed into one single zip file. File names are explained as follows:</p> <ol> <li>LLD0021C_locla.fasta : Whole genome assembly of a Taiwanese male individual generated by LOCLA</li> <li>LLD0021C_locla_quality.txt : Assembly statistics of LLD0021C_locla.fasta</li> <li>chm13_locla.fasta : Whole genome assembly of the CHM13 cell line generated by LOCLA</li> <li>chm13_locla_quality.txt : Assembly statistics of chm13_locla.fasta</li> <li>hg002_gma_locla.fasta : Whole genome assembly of the HG002 sample generated by LOCLA </li> <li>hg002_gma_locla_quality.txt : Assembly statistics of hg002_gma_locla.fasta</li> </ol> <p> </p>
Poecilia picta female genome assembly files
<p>Sex chromosome dosage compensation is a model to understand the coordinated evolution of transcription, however, the advanced age of the sex chromosomes in model systems makes it difficult to study how the complex regulatory mechanisms underlying chromosome-wide dosage compensation can evolve. The sex chromosomes of <em>Poecilia picta</em> have undergone recent and rapid divergence, resulting in widespread gene loss on the male Y, coupled with complete X Chromosome dosage compensation, the first case reported in a fish. The recent de novo origin of dosage compensation presents a unique opportunity to understand the genetic and evolutionary basis of coordinated chromosomal gene regulation. By combining a new chromosome-level assembly of <em>P. picta</em> with whole-genome bisulfite sequencing and RNA-seq data, we determine that the Yin Yang 1 (YY1) DNA-binding motif is associated with male-specific hypomethylated regions on the X, but not the autosomes. These YY1 motifs are the result of a recent and rapid repetitive element expansion on the <em>P. picta</em> X Chromosome, which is absent in closely related species that lack dosage compensation. Taken together, our results present compelling support that a disruptive wave of repetitive element insertions carrying YY1 motifs resulted in the remodeling of the X Chromosome epigenomic landscape and the rapid de novo origin of a dosage compensation system.</p>
Cosmopolites sordidus genome assemblies
<p><span>PacBio HiFi sequencing was employed in combination with metagenomic binning to produce a high-quality reference genome of <em>Cosmopolites</em> <em>sordidus</em>. We compared k-mer and alignment reference-based pre-binning and post-binning approaches to remove contamination. We were also interested to know if the post-binning approach had interspersed Bacterial contamination within intragenic regions of Arthropoda-binned contigs. Our analyses identified 3,433 genes that were composed with reads identified as of putative bacterial origins. The pre-binning approach yielded a <em>C. sordidus</em> genome of 1.07Gb genome composed of 3,089 contigs with 98.6% and 97.1% complete and single copy genome and protein <em>BUSCO</em> scores respectively. In this paper, we demonstrate that in this case, the pre-binning approach does not sacrifice assembly quality for more stringent metagenomic filtering. We also determine post-binning allows for increased intragenic contamination increased with increasing coverage, but the frequency of gene contamination increased with lower coverage. Finally, NCBI's new FCS-GX program was used as a final post-assembly classification approach and identified contamination in both pre- and post-binning assemblies. This indicates that both pre- and post-binning approaches are required to fully remove contamination. Future work should focus on developing reference-free pre-binning approaches for HiFi reads produced from eukaryotic-based metagenomic samples. </span></p>
Detailed information for the 2032 Saccharomyces cerevisiae genome assemblies studied
<p>Detailed information for the 2032 Saccharomyces cerevisiae genome assemblies studied</p>
Protein families for 2032 Saccharomyces cerevisiae genome assemblies
<p>The protein families obtained with different cluster cutoffs for different sets of genome assemblies as well as the marker genes are presented in text files. In each file, a row indicates a family and a column (separated by TAB) indicates a genome assembly. Protein IDs for multiple homologues from the same assembly are separated by '|'. The family ID is shown in the first column and the tag for each assembly is shown in the first row. Cluster cutoffs used are 50%, 60%, 70%, 80%, and 90%. Genome sets shown are all genomes (all), non-redundant (nr) genomes, medium-high-quality genomes (mhq), and high-quality genomes (hq).</p>
Genome assemblies for halophilic bacteria with potential contamination, sampled from Northern California in 2022
Open the record for dataset details and reuse information.
Chloroplast genome assemblies and comparative analyses of commercially important Vaccinium berry crops
Open the record for dataset details and reuse information.
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.