Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Metagenomics assemblies and high-quality MAGs for "Long-read metagenomics to retrieve high-quality metagenome-assembled genomes from canine feces"
<p>This dataset includes the different metagenomics assemblies analyzed and its summary (_info.txt file):</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/100_assembly.fasta">100_assembly.fasta</a> is the Flye 2.7 metagenomics assembly merging HMW and non-HMW datasets</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/75_assembly.fasta">75_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 75% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/50_assembly.fasta">50_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 50% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/HMW_assembly.fasta?versionId=749ff6fd-2642-4ad1-971a-7f3404baa595">HMW_assembly.fasta</a> is the Flye 2.7 metagenomics assembly for HMW dataset.</p> <p>Moreover, it also includes the eight frameshift-corrected high-quality MAGs analyzed in the manuscript. </p>
Data from: Assembly ASM291031v2 (Genbank: GCA_002910315.2) identified as assembly of the Northern Dolly Varden (Salvelinus malma malma) genome, and not the Arctic char (S. alpinus) genome
<p>Here is the data that is a supplementary to the preprint: Shedko S.V. 2019. Assembly ASM291031v2 (Genbank: GCA_002910315.2) identified as assembly of the Northern Dolly Varden (Salvelinus malma malma) genome, and not the Arctic char (S. alpinus) genome // arXiv:1912.02474 <a href="https://arxiv.org/abs/1912.02474">https://arxiv.org/abs/1912.02474</a></p>
Worldwide Fraxinus Genome Assemblies, Annotations, and Gene Families v0.2
<p>The worldwide <em>Fraxinus</em> genome project was conducted to assess the pathogenic resistance of 34 Ash tree species to Ash Dieback and Emerald Ash Borer. The project is led by Dr. Richard Buggs. v0.1 genomes are available at ashgenome.org and on ENA. As part of Josiah Seaman's PhD thesis, he improved the assembly of 13 genomes included here as v0.2. New de novo annotations, gene families, and all associated files are included for future studies and reproducibility. </p> <p>Annotations are done with GeMoMa using F. excelsior as a reference (Keilwagen et al. 2016). Gene families are defined as genes originating from a single copy at the last common ancestor with Solanum. Orthofinder outputs reconciled gene trees, aligned CDS, and gene families (Emms and Kelly 2015; Tekaia 2016). Species tree was calibrated based on fossil evidence using r8s, RAxML across 25,182,399 sites (SpeciesTreeAlignment.fa). More methods details can be found in the full Chapter two of Josiah Seaman's PhD thesis (2021).</p> <p>I'd be happy to talk with you if you'd like any additional information or help visualizing your genomic data. You can find the tools used to browse this data at https://fluentdna.com/ and http://graphgenome.org/ Contact me at josiah@newline.us</p>
Metagenome assemblies and metagenome-assembled genomes from the Daphnia magna microbiota
<p>Metagenome assemblies generated from raw reads not mapping to the Daphnia magna genome for six samples assembled individually (G4, G14, S1-S4) and a coassembly of all six samples (a_assembly) using metaSPAdes in SPAdes v3.14. Assemblies can be found in metagenome_assemblies.zip.</p> <p>Metagenome-assembled genomes generated using VAMB v3.0.2 (vamb_bins.zip) and ProxiMeta (proximeta_bins.zip). These MAGs were taxonomically identified using GTDB-Tk v1.3 and quality checked using CheckM v1.1. Outputs from GTDB-Tk and CheckM can be found in the .tsv and .tab files, respectively.</p>
Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi
<p>The Puma lineage within the family Felidae consists of three species that last shared a common ancestor around 4.9 million years ago. Whole-genome sequences of two species from the lineage were previously reported: the cheetah (<em>Acinonyx jubatus</em>) and the mountain lion (<em>Puma concolor</em>). The present report describes a whole-genome assembly of the remaining species, the jaguarundi (<em>Puma yagouaroundi</em>). We sequenced the genome of a male jaguarundi with 10X Genomics linked reads and assembled the whole-genome sequence. The assembled genome contains a series of scaffolds that reach the length of chromosome arms and is similar in scaffold contiguity to the genome assemblies of cheetah and puma, with a contig N50 = 100.2 kbp and a scaffold N50 = 49.27 Mbp. We assessed the assembled sequence of the jaguarundi genome using BUSCO, aligned reads of the sequenced individual and another published female jaguarundi to the assembled genome, annotated protein-coding genes, repeats, genomic variants and their effects with respect to the protein-coding genes, and analyzed differences of the two jaguarundis from the reference mitochondrial genome. The jaguarundi genome assembly and its annotation were compared in quality, variants and features to the previously reported genome assemblies of puma and cheetah. Computational analyzes used in the study were implemented in transparent and reproducible way to allow their further reuse and modification.</p>
Assembly and comparison of two closely related Brassica napus genomes
<p>Here we present the <em>de novo</em> assembly of the <em>B. napus</em> cultivar Tapidor and comparison with an improved assembly of the <em>B. napus</em> cultivar Darmor<em>-bzh</em>. Both cultivars were annotated using the same method to allow comparison of gene content. We identified genes unique to each cultivar and differentiate these from artefacts due to variation in the assembly and annotation. We demonstrate that using a common annotation pipeline can result in different gene predictions, even for closely related cultivars, and repeat regions which collapse during assembly impact whole genome comparison. After accounting for differences in assembly and annotation, we demonstrate that the genome of Darmor<em>-bzh</em> contains a greater number of genes than the genome of Tapidor.</p>
Bacterial training dataset for Galaxy training network tutorials on Genome assembly
<p>This training dataset is from an imaginary <em>Staphylococcus aureus</em> bacterium with a miniature genome. There is a reference genome in various formats as well as some fastq reads of a closely related but also imaginary mutant strain.</p> <p>It is a useful dataset for demonstrating:</p> <ul> <li>de novo genome assembly</li> <li>read mapping and variant calling</li> <li>genome annotation</li> </ul> <p>The files included are:</p> <ul> <li><strong>wildtype.fna</strong>: the reference genome sequence of the wildtype strain in fasta format (a header line, then the nucleotide sequence of the genome.)</li> <li><strong>wildtype.gff</strong>: the reference genome sequence of the wildtype strain in general feature format (a list of features - one feature per line, then the nucleotide sequence of the genome.)</li> <li><strong>wildtype.gbk</strong>: the reference genome sequence in genbank format.</li> <li><strong>mutant_R1.fastq</strong> and <strong>mutant_R2.fastq</strong>: Fastq sequence reads of a closely related mutant strain. <ul> <li>The reads are paired-end.</li> <li>Each read is 150 bases long.</li> <li>The number of bases sequenced is equivalent to 19x the genome sequence of the wildtype strain. (Read coverage 19x - rather low!).</li> </ul> </li> </ul>
Supplemental dataset for Northern Spotted Owl (<i>Strix occidentalis caurina</i>) genome assembly version 1.0
<p><strong>StrOccCau_1.0_nuc.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the file that we deposited at DDBJ/ENA/GenBank as a Whole Genome Shotgun (WGS) project under accession NIFN00000000. It is is the file that you will most likely want to download if you would like to perform an alignment to this genome assembly. This file is the assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012) without any contigs and scaffolds less than 1,000 nt and also without the contigs and scaffolds that we identified either as the mitochondrial genome sequence or as contaminant sequences.</p> <p><strong>StrOccCau_1.0_nuc_masked.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the repeat-masked (hard-masked) assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012) without any contigs and scaffolds less than 1,000 nt and also without the contigs and scaffolds that we identified either as the mitochondrial genome sequence or as contaminant sequences.</p> <p><strong>StrOccCau_1.0_mito.fa</strong> : This FASTA format file is the mitochondrial-genome-derived scaffold from the assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012).</p> <p><strong>StrOccCau_1.0.gff.bz2</strong> : This gff format file compressed this file with bzip2 contains the gene annotations of StrOccCau_1.0_nuc.fa.</p> <p><strong>StrOccCau_1.0_transcripts.fa.bz2</strong> : This FASTA format file compressed this file with bzip2 contains the sequences of the gene transcript sequences of the genes annotated in StrOccCau_1.0.gff.</p> <p><strong>StrOccCau_1.0_proteins.fa.bz2</strong> : This FASTA format file compressed this file with bzip2 contains the protein sequences of the genes annotated in StrOccCau_1.0.gff.</p> <p><strong>StrOccCau_1.0_RM_homology_includes_LowComplexity.out.bz2</strong> : This file provides the repeat annotations produced by the homology-based masking of StrOccCau_1.0_nuc.fa that included masking of low complexity regions and simple repeats. We compressed this file with bzip2.</p> <p><strong>StrOccCau_1.0_RM_DeNovo_includes_LowComplexity.out</strong> : This file provides the repeat annotations produced by the de novo masking (which followed after first performing homology-based masking) of StrOccCau_1.0_nuc.fa that included masking of low complexity regions and simple repeats.</p> <p><strong>StrOccCau_1.0_RM_homology_no_LowComplexity.out.bz2</strong> : This file provides the repeat annotations produced by the homology-based masking of StrOccCau_1.0_nuc.fa that did not include masking of low complexity regions and simple repeats. We compressed this file with bzip2.</p> <p><strong>StrOccCau_1.0_RM_DeNovo_no_LowComplexity.out</strong> : This file provides the repeat annotations produced by the de novo masking (which followed after first performing homology-based masking) of StrOccCau_1.0_nuc.fa that did not include masking of low complexity regions and simple repeats.</p> <p><strong>StrOccCau_1.0_alignments_of_light_associated_genes.txt</strong> : This file provides alignments of light-associated gene orthologs as well as assemblies of transcriptome sequences in NEXUS format.</p> <p><strong>StrOccCau_1.0_nuc_masked_SpottedBarredOwl_variant_file.vcf.bz2</strong> : This is a raw, unfiltered variant call format file compressed with bzip2 that was generated after aligning both spotted owl and barred owl short read data aligned to StrOccCau_1.0_nuc_masked.fa.</p> <p><strong>StrOccCau_0.1.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012).</p> <p><strong>StrOccCau_0.1_masked.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the repeat-masked assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012).</p> <p><strong>StrOccCau_0.2.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012) without any contigs and scaffolds less than 1,000 nt.</p> <p><strong>StrOccCau_0.2_masked.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the repeat-masked assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012) without any contigs and scaffolds less than 1,000 nt.</p> <p><strong>StrOccCau_GapCloser_output_NoContamNoMito.fa.bz2</strong> : This FASTA format file compressed with bzip2 is the assembly output from SOAPdenovo2 toolkit GapCloser version 1.12-r6 (Luo et al. 2012) without the contigs and scaffolds that we later identified either as the mitochondrial genome sequence or as contaminant sequences.</p> <p><strong>Citations</strong> - if you utilize these data, please include these citations:</p> <p>Hanna ZR., Henderson JB., Wall JD., Emerling CA., Fuchs J., Runckel C., Mindell DP., Bowie RCK., DeRisi JL., Dumbacher JP. 2017a. Supplemental dataset for Northern Spotted Owl (<em>Strix occidentalis caurina</em>) genome assembly version 1.0. <em>Zenodo</em>. DOI: 10.5281/zenodo.822859.</p> <p>Hanna ZR., Henderson JB., Wall JD., Emerling CA., Fuchs J., Runckel C., Mindell DP., Bowie RCK., DeRisi JL., Dumbacher JP. 2017b. Northern Spotted Owl (Strix occidentalis caurina) Genome: Divergence with the Barred Owl (<em>Strix varia</em>) and Characterization of Light-Associated Genes. <em>Genome Biology and Evolution</em> 9:2522–2545. DOI: 10.1093/gbe/evx158.</p>
Liftover of DGRP D.melanogaster genotypes to reference genome assembly v6.0, with QC graph
<p>Output are vcf and plink-format genotype files. Also provided are the run code in bash and R, the logs, summary statistics, and a graph showing how the positions of SNPs have changed.</p>
Tandem repeat catalog of the human genome generated from long-read assemblies
<p>Allele sequences of polymorphic loci (VCF) and README for all version 2 (2.0 + 2.1) files</p>
Kellet's whelk genome and transcriptome assembly
<p>Understanding genomic characteristics of non-model organisms can help bridge gaps in ecology and evolutionary sciences, but lack of a reference genome and transcriptome for these species challenges their study. We advance this goal by conducting the first full genome and transcriptome sequence assembly and analysis of the non-model organism Kellet's whelk, <em>Kelletia kelletii</em>, a marine gastropod and fisheries species exhibiting a northern range expansion along the US west coast that is potentially driven by climate change. We used a combination of Oxford Nanopore Technologies, PacBio, and Illumina platforms for sequencing, and integrated a set of bioinformatic pipelines to create a comprehensive and contiguous de novo genome assembly. Our results represent the most complete and continuous documented genome among the <em>Buccinoidea</em> superfamily to date. Genome validation revealed its relatively high completeness with low missing metazoan BUSCOs, and an average coverage of ~70x for all contigs, indicating a robust assembly. Characteristics of the <em>K. kelletii</em> genome showed that short-read data contributed significantly to genome coverage and accuracy; however, long-read data was imperative to the completeness and continuity of the genome assembly. Genome annotation identified a large number of protein-coding genes compared to other closely related species, suggesting the presence of a complex genome structure. We conducted the transcriptome assembly and analysis of individuals during their period of peak embryonic development, and revealed highly expressed genes associated with specific GO terms and metabolic pathways, most notably lipid, carbohydrate, glycan, and phospholipid metabolism. We also identified numerous heat shock proteins (HSPs) in the transcriptome and genome with a potential association between the transcriptional expansion of HSP families and the marine environment experienced by the sessile life history stage of the developing embryo. This study offers a valuable reference genome and transcriptome for conducting comprehensive bioinformatic analyses of the non-model organism <em>K. kelletii</em>. Such resources will enhance our understanding of its ecology and evolution, as well as that of other coastal marine species facing environmental changes.</p>
PacBio HiFi de-novo assembled genome and mitochondrial genome for Orbicella faveolata
<p>Final assembly using Funannotate of <i>Orbicella faveolata</i> from PacBio HiFi reads. For full methods please see the publication. </p>
gymnome-assembly: OpenAI Environments for Genome Assembly
<p>OpenAI Environments to simulate genome assemblies. These environments are described in the paper "<em>A step toward a reinforcement learning de novo genome assembler</em>". More information at http://github.com/%20kriowloo/gymnome-assembly.</p>
Genome assemblies of 382 carbapenem-resistant Pseudomonas aeruginosa isolates collected from Japanese hospitals in 2019−2020
<p>This dataset provides genome assemblies used in the study of "Nationwide genome surveillance of carbapenem-resistant Pseudomonas aeruginosa in Japan".</p> <p> </p>
Annotation files related to the Telomere-to-Telomere genome assembly of the clubroot pathogen Plasmodiophora brassicae (GCA_036867785.1)
<p>This repository contains annotation files related to the T2T genome aseembly of <em>Plasmodiophora brassicae</em>. Link to the NCBI genome submission- https://www.ncbi.nlm.nih.gov/bioproject/1071157</p> <p><strong>Description of the files :</strong></p> <p><strong>GCA_036867785.1_ULAVAL_Pb3A_genomic.fna</strong> - Soft-masked genome sequence FASTA file representing 20 chromosomes.</p> <p><strong>sequence_report.jsonl</strong> - Detailed information about individual chromosome seqeunce.</p> <p><strong>PBTT_annotation.gtf</strong> - GTF file corresponding to the genomic FASTA file.The GTF file was generated by BRAKER3 and contains information about all possible transcripts.</p> <p><strong>PBTT_CDS_longest_isoform.fasta</strong> - Contains 10521 FASTA sequences representing the CDS of only the longest isoform of the gene models.</p> <p><strong>PBTT_protein_longest_isoform.fasta</strong> - Contains 10521 FASTA sequences representing the amino acid sequences of only the longest isoform of the gene models.</p>
Genome annotation file for a draft genome assembly for Nucella lapillus
<p><span>A male specimen of wild <em>Nucella lapillus</em>, measuring approximately 1.5–3 cm in length, was collected <span>from a rocky shore (mid-upper shore) at low tide from near the quay at <span>Portnahaven, Isle of Islay, Argyll and Bute, Scotland (National Grid Reference NR 16614 51966)</span> on <span>26th June 2023. Genomic DNA was extracted from the non-shell tissue of the specimen, and sequenced using PacBio HIFI and Oxford Nanopore technolgies (ONT) long read sequencing platforms. </span></span>The genome assembly was derived from 40.6 Gb of PacBio HiFi reads (read <span>N50, 11291; N90, 9246</span>), and 61.1 Gb of ONT data (read <span>N50, </span>3643<span>; N90, </span>1546). </span>Annotation of protein-coding genes in the cleaned and masked genome assembly of <em>Nucella lapillus</em> was performed using GALBA v1.0.11, an automated pipeline that uses proteins from a closely related species to assist in the training of gene prediction using AUGUSTUS. Proteins from <em>Rapana venosa</em> were provided for this purpose, and the miniprot option was used to perform the protein-to-genome alignments. Functional annotation of predicted protein-coding genes was performed using eggNOG-mapper v2.1.12, and additionally annotated with best hit BLAST results (v2.16.0) against the proteomes of the following marine gastropod species: <em>Rapana venosa</em>, <em>Littorina. saxatilis</em>, <em>Pomocea canaliculata, Stramonita haemastoma</em> and <em>Haliotis rufescens</em>.</p>
Syzygium jambos genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Syzygium jambos</em>.</p> <p>The following files are available:</p> <ul> <li>sjam.fa.gz: reference genome sequence in fasta format</li> <li>sjam.gff3.gz: gene annotation in GFF3 format</li> <li>sjam.gtf.gz: gene annotation in GTF format</li> <li>sjam.tx.fa.gz: transcript sequences in fasta format</li> <li>sjam.cds.fa.gz: coding sequences in fasta format</li> <li>sjam.prot.fa.gz: protein sequences in fasta format</li> <li>sjam.tsv.gz: gene functional annotation in TSV format</li> <li>sjam.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Nicotiana tomentosiformis genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tomentosiformis</em>.</p> <p>The following files are available:</p> <ul> <li>ntom.fa.gz: reference genome sequence in fasta format</li> <li>ntom.gff3.gz: gene annotation in GFF3 format</li> <li>ntom.gtf.gz: gene annotation in GTF format</li> <li>ntom.tx.fa.gz: transcript sequences in fasta format</li> <li>ntom.cds.fa.gz: coding sequences in fasta format</li> <li>ntom.prot.fa.gz: protein sequences in fasta format</li> <li>ntom.tsv.gz: gene functional annotation in TSV format</li> <li>ntom.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntom.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntom.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntom.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Syzygium malaccense genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Syzygium malaccense</em>.</p> <p>The following files are available:</p> <ul> <li>smal.fa.gz: reference genome sequence in fasta format</li> <li>smal.gff3.gz: gene annotation in GFF3 format</li> <li>smal.gtf.gz: gene annotation in GTF format</li> <li>smal.tx.fa.gz: transcript sequences in fasta format</li> <li>smal.cds.fa.gz: coding sequences in fasta format</li> <li>smal.prot.fa.gz: protein sequences in fasta format</li> <li>smal.tsv.gz: gene functional annotation in TSV format</li> <li>smal.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Syzygium aqueum genome assembly and annotation
<p><em>De novo </em>genome assembly and annotation of <em>Syzygium aqueum</em>.</p> <p>The following files are available:</p> <ul> <li>saqu.fa.gz: reference genome sequence in fasta format</li> <li>saqu.gff3.gz: gene annotation in GFF3 format</li> <li>saqu.gtf.gz: gene annotation in GTF format</li> <li>saqu.tx.fa.gz: transcript sequences in fasta format</li> <li>saqu.cds.fa.gz: coding sequences in fasta format</li> <li>saqu.prot.fa.gz: protein sequences in fasta format</li> <li>saqu.tsv.gz: gene functional annotation in TSV format</li> <li>saqu.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.