Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Adapter sequences used for trimming of genomic sequences in the assembly of the Northern Spotted Owl (<i>Strix occidentalis caurina</i>) genome assembly version 1.0
<p>These files provide the sequences of the adapters used in the construction of the genomic libraries Hanna et al. (2017a) sequenced and used to assemble the Northern Spotted Owl (<em>Strix occidentalis caurina</em>) genome assembly version 1.0 (Hanna et al. 2017b). These files also contain relevant supplemental adapter sequences from the adapter files included with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011595_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011595. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011596_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011596. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011597_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011597. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011614_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011614. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011615_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011615. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" file distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011616_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011616. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" file distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).<br> <br> <strong>SRR4011617_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011617. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" file distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p>
Processed Hi-C contact matrices for "Three invariant Hi-C interaction patterns: applications to genome assembly"
<p>Processed Hi-C interaction matrices, saved in numpy npz format.</p> <p>Matrices were processed using Dekker lab cMapping pipeline.</p> <p>Raw sequence data was taken from:</p> <p>Hap1: Haarhuis et al 10.1016/j.cell.2017.04.013</p> <p>IMR90, H1ESC, MESC, MCORTEX: Dixon et al 10.1038/nature11082</p> <p>Worm: Crane et al 10.1038/nature14450</p> <p>Caulobacter: Le et al 10.1126/science.1242059</p> <p> </p>
Genome assembly of E. coli C
<p>This record contains the following files:</p> <ol> <li><code>Ecoli_C_assembly.fna</code> - finished assembly in FASTA format. Contains two sequences: nuclear genome of <em>E. coli </em>C and genomes of bacteriophage phiX174, which was used as a spike-in</li> <li> <p><code>Ecoli_C_ONT.pass.fast5.tar.gz</code> - fast5 files from Oxford Nanopore Run</p> </li> <li> <p><code>ont.fastqsanger.gz</code> - oxford nanopore reads in fastqsanger format</p> </li> <li> <p><code>forward.fastqsange</code> - forward Illumina reads in fastqsanger format</p> </li> <li> <p><code>reverse.fastqsanger</code> - reverse Illumina reads in fastqsanger format</p> </li> </ol>
An ultra-dense haploid genetic map for evaluating the highly fragmented genome assembly of Norway spruce (Picea abies)
<p>Data files for construction of the haploid genetic map for Norway spruce (<em>Picea abies</em>). Available at <a href="https://doi.org/10.1101/292151">https://doi.org/10.1101/292151</a></p>
Chardonnay genome assembly and annotations
<p>Annotations relating to the Chardonnay genome assembly (the study is published here: https://doi.org/10.1371/journal.pgen.1007807). The assembly is also available at NCBI: BioProject: PRJNA399599.</p> <p><strong>chardonnay_p-ctg.fasta / chardonnay_h-ctg.fasta:</strong></p> <p>Contig assembly for Chardonnay. chardonnay_p-ctg.fasta = primary contigs (haploid representation), chardonnay_h-ctg.fasta = haplotigs (alt contigs for phased regions in haploid assembly).</p> <p><strong>p-h.maker.gff / p-h.maker.proteins.faa / p-h.maker.transcripts.fna:</strong></p> <p>Maker-predicted gene annotations.</p> <p><strong>p-h.maker.draft-names.tsv / p-h.orthomcl.orthoGroups.tsv / p-h.KEGG.tsv:</strong></p> <p>Draft names (based on UniprotKB blastP hits), OrthoMCL annotations, and KEGG annotations for maker-predicted genes.</p> <p><strong>p-h.repeats.gff:</strong></p> <p>RepeatMasker-based repeat annotations.</p> <p><strong>chardonnay_primary_contigs_chromosome-order.fa / chardonnay_haplotigs_chromosome-order.fa:</strong></p> <p>Contigs placed in chromosome-order (using PN40024 as reference).</p> <p><strong>chardonnay_primary_contig_mappings.tsv / chardonnay_haplotig_mappings.tsv:</strong></p> <p>Mapping coordinates for chromosome-ordered contigs.</p> <p><strong>kmer-based_parentage.primary_contigs.bed / kmer-based_parentage.haplotigs.bed:</strong></p> <p>Parentage assignments using the kmer-based method described in the study.</p> <p><strong>SNP-based_parentage.primary_contigs.bed / SNP-based_parentage.haplotigs.bed:</strong></p> <p>Parentage assignments using SNP-based method (view haplotig assignments against primary contigs) described in the study.</p> <p><strong>p-ctg.gene-expansion-candidates.tsv / h-ctg.gene-expansion-candidates.tsv / p-ctg.gene-expansion-candidates.bed / h-ctg.gene-expansion-candidates.bed:</strong></p> <p>Gene expansion candidates (TSV = 1 row per predicted orthogroup, BED = annotations for viewing).</p> <p><strong>PN_and_CH.FAR2.msa.png: </strong></p> <p>Multi Sequence Alignment for expansion of FAR2-like genes in described in Chardonnay genome assembly publication described in the study.</p>
Draft genome assemblies of killifish from the Fundulus genus with ONT and Illumina sequencing platforms
<p>Four species from the genus Fundulus were selected for genome sequencing to study the physiological and genetic mechanisms that diverge between euryhaline and stenohaline freshwater species within this cyprinodontiform order of ray-finned fishes.</p>
Dataset for "Progessive improvement of the Australian blacklip abalone (Haliotis rubra) genome assembly with Nanopore long reads, hybrid meta assembly and haplotig purging
<p>This Zenodo archive contains the blacklip abalone genome assemblies and and their BUSCO completeness calculations. Genome annotation (gff3 format), CDS, protein sequences and Orthofinder2 output were also included.</p>
The First Highly Contiguous Genome Assembly of Pikeperch (Sander lucioperca), an Emerging Aquaculture Species in Europe
<p><strong>Supporting data for "The First Highly Contiguous Genome Assembly of Pikeperch (<em>Sander lucioperca</em>), an Emerging Aquaculture Species in Europe"</strong></p> <p>===========================================================================================</p> <p><strong>Abstract:</strong></p> <p>--------</p> <p>The pikeperch (<em>Sander lucioperca</em>) is a fresh and brackish water Percid fish natively inhabiting the northern hemisphere. This species is emerging as a promising candidate for intensive aquaculture production in Europe. Specific traits like cannibalism, growth rate and meat quality require genomics based understanding, for an optimal husbandry and domestication process. Still, the aquaculture community is lacking an annotated genome sequence to facilitate genome-wide studies on pikeperch. Here, we report the first highly contiguous draft genome assembly <em>S. lucioperca</em>. In total, 413 and 66 giga base pairs of DNA sequencing raw data were generated with Illumina platform and PacBio Sequel System, respectively. The PacBio data were assembled into a final assembly size of ~900 Mb covering 89% of the 1,014 Mb estimated genome size. The draft genome consisted of 1,966 contigs ordered into 1,313 scaffolds. The contig and scaffold N50 lengths are 3.0 Mb and 4.9 Mb, respectively. The identified repetitive structures accounted for 39% of the genome. We utilized homologies to other ray-finned fishes, and ab initio gene prediction methods to predict 21,249 protein-coding genes in the <em>S. lucioperca </em>genome, of which 88% were functionally annotated by either sequence homology or protein domains and signatures search. The assembled genome spans 97.6% and 96.3% of Vertebrate respectively Actinopterygii single-copy orthologs. The outstanding mapping rate (99.9%) of genomic PE-reads on the assembly suggests an accurate and nearly complete genome reconstruction. This draft genome sequence is the first genomic resource for this promising aquaculture species. It will provide an impetus for genomic-based breeding studies targeting phenotypic and performance traits of captive pikeperch.</p> <p> </p> <p><strong>Files:</strong></p> <p>------</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu.cds.renamed.fa">sanlu.cds.renamed.fa </a> - Coding sequences of predicted protein-coding genes </p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu.genes.filt.gff3">sanlu.genes.filt.gff3 </a> - gff3 file of predicted protein coding genes</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu.genes.pep.fa">sanlu.genes.pep.fa </a> - predicted peptide sequences </p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu.genome.ctg.fasta">sanlu.genome.ctg.fasta </a> - <em>Sander lucioperca</em> genome assembly at contig-level</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu.genome.scf.fa">sanlu.genome.scf.fa </a> - <em>Sander lucioperca</em> genome assembly at scaffold-level</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/Sanlu.genome.masked.fasta">Sanlu.genome.masked.fasta </a>- Repeats-masked <em>Sander lucioperca</em> genome assembly at scaffold-level</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/Sanlu.genome.repeats.gff">Sanlu.genome.repeats.gff </a> - Gff3 file of predicted repeats in <em>Sander lucioperca</em> genome</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/Additional_File_2.xlsx">Additional_File_2.xlsx </a> - Functional annotations of <em>Sander lucioperca </em>genes by SwissProt, NR RefSeq, TrEMBL and InterPro databases</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu.repeats.lib.fasta">sanlu.repeats.lib.fasta </a> - Predicted repeats library in <em>Sander lucioperca </em>in FASTA format</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu_miRNA.csv">sanlu_miRNA.csv </a> Predicted micro RNA families in CSV tab file </p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu_miRNA.bed">sanlu_miRNA.bed </a> - Predicted micro RNA families in BED file format</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu_miRNA.html">sanlu_miRNA.html </a><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu_miRNA.bed"> </a> - Predicted micro RNA families in HTML</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu_rRNA.fasta">sanlu_rRNA.fasta </a> - Predicted ribosomal RNA (rRNA) sequences in FASTA file format</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/sanlu_rRNA.gff">sanlu_rRNA.gff </a> - Predicted ribosomal RNA (rRNA) sequences in GFF file format</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/trna.genes.csv">trna.genes.csv </a> - Predicted transfer RNA (tRNA) genes in CSV tab file</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/SpeciesTree_rooted_node_labels.txt">SpeciesTree_rooted_node_labels.txt </a> - Predicted phylogenetic tree in NEWICK format</p> <p><a href="https://zenodo.org/api/files/808d4d80-6012-4046-bc3c-73b9792b5d8c/SpeciesTreeAlignment.fa">SpeciesTreeAlignment.fa </a> - Species tree alignment in FASTA, based on 1.1 single copy orthologs</p> <p> </p>
TOPC_bin_586 metagenome assembled genome (MAG)
<p><strong>Contig, gene sequences and functional annotation of the <em>TOPC_bin_586</em> metagenome assembled genome (MAG)</strong></p> <p>Data available:</p> <ol> <li>Nucleotide sequences of the contigs composing the MAG [<em>topc.bin.586.fna</em>]</li> <li>Amino acid sequences of the genes (open reading frames, ORFs) [<em>topc.bin.586_ORFs.faa</em>]</li> <li>Functional annotation table (tab-delimited) for the ORFs [<em>topc.bin.586_ORFs_annotation.tsv</em>]</li> </ol>
ST131_4071_genome_assemblies
<p>The genome assemblies of 4,071 E. coli ST131 genomes (see Decano & Downing 2019).</p>
Improved genome assembly and annotation of the soybean aphid (Aphis glycines Matsumura)
<p>Updated genome assembly and annotation of <em>Aphis glycines</em> biotype 4.</p> <p><strong>Overview of files included in this release:</strong></p> <p><strong>Frozen release:</strong></p> <p>Updated <em>A. glycines </em>biotype 4 genome assembly: Aphis_glycines_4.v2.1.scaffolds.fa.gz </p> <p>BRAKER2 gene models for updated <em>A. glycines </em>biotype 4 genome assembly: Aphis_glycines_4.v2.1.scaffolds.fa.gff</p> <p>BRAKER2 protein sequences: Aphis_glycines_4.v2.1.scaffolds.fa.gff.aa.fa</p> <p>BRAKER2 nucleotide coding sequences: Aphis_glycines_4.v2.1.scaffolds.fa.gff.CDS.fa</p> <p><strong>Unfiltered raw intermediate genome assemblies:</strong></p> <p>Canu assembly of biotype 4 PacBio data from Wenger et. al. (2017): canu.fa.gz</p> <p>DBG2OLC hybrid assembly of selected biotype 4 MiSeq data and biotype 4 PacBio data from Wenger et. al. (2017): DBG2OLC.fa.gz</p> <p>Merged Canu and DBG2OLC assembly created with quickmerge: quickmerge.fa.gz</p> <p>Pilon polished (2 rounds) quickmerge assembly: quickmerge.pilon_r2.fa.gz</p> <p><strong>Mitochondrial and endosymbiont contigs extracted from the pilon polished quickmerge assembly: </strong></p> <p><em>A. glycines </em>biotype 4 mitochondrial genome: Aphis_glycines_4_Buchnera_v1.fa</p> <p><em>A. glycines </em>biotype 4 <em>Buchnera aphidicola</em> contigs: Aphis_glycines_4_Buchnera_v1.fa</p> <p><em>A. glycines </em>biotype 4 <em>Wolbachia</em> contigs: Aphis_glycines_4_Buchnera_v1.fa</p> <p><strong>Other files:</strong></p> <p>MUSCLE alignment of <em>A. glycines </em>v1, <em>A. glycines </em>biotype 4 v2.1 and <em>Drosophila melanogaster</em> R6.22 Osiris proteins in fasta format: D_mel_v1_v2_osiris.prots.muscle.fasta</p> <p>FastTree Maximum Likelihood phylogeny based on the MUSCLE alignment of Osiris genes in newick format: D_mel_v1_v2_osiris.prots.muscle.FastTree.nwk</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
De novo whole genome assembly of the giant tiger prawn (Penaeus monodon) from Vietnam
<p>Basecalled Nanopore FastQ files for the Vietnamese giant tiger prawn and its genome assemblies.</p> <p>FastQ files (LSK109 sample prep sequenced on a MinION device for 48 hours). Read stats are in *_stat.txt:</p> <p>TP_A.fastq : gDNA was extracted using Zymo quick DNA minikit from ethanol-preserved muscle tissue</p> <p>TP_B.fastq : Same as TP_A.fastq</p> <p>TP_C.fastq : gDNA was extracted using conventional salting out method (longer read length but reduced yield)</p> <p>Assemblies:</p> <p>v1_MaSuRCA.fasta: Assembly using poly-G trimmed Illumina reads</p> <p>v2_NanoporeScaf.fasta : Scaffolding with Nanopore long reads</p> <p>v3_RNA_NanoporeScaf.fasta: Scaffolding of v2 with RNA reads</p> <p>v4_NCBI_Filt.fasta: post NCBI contaminant and carry-over adapter removal (final version)</p> <p>Annotation:</p> <p>Braker2_annotation.gff3.gz: Inititial Braker2 gff3 output</p> <p>Braker2_CDS.fna.gz: Initial Braker2 predicted genes</p> <p>Braker2_prot.faa.gz: Protein translation of Braker2_CDS.fna</p> <p>CAZy.tar.gz: CAZy annotation for four crustacean species</p> <p>Filtered_Gene.tar.gz: List of genes with functional annotation and/or orthologs</p> <p>Interproscan_result.tsv: Raw InterProScan output</p> <p>OrthoFinder2.tar.gz: OrthoFinder2 output. Proteins used to infer orthologs are included as ".faa".</p>
Chromosome-level genome assembly of a living fossil, the Atlantic Horseshoe Crab Limulus polyphemus
<p>Associated data for male Atlantic horseshoe crab <em>Limulus polyphemus </em>chromosome-scale genome and annotations, including genome (FASTA), structural gene annotations (GFF3), functional annotations (TSV), coding sequences (CDS), protein sequences (PEP), RepeatModeler library (FA.CLASSIFIED), and repeat annotations (OUT).</p> <p>qaLimPoly3.1 - Publication analyses were completed with this genome. </p> <p>qaLimPoly3.3 - This is the current reference assembly. Assembly updated with Sanger sequencing based edits of Chr11 and removal of adapter contamination. Gene annotations updated with curation of canonical proclotting genes. Repeat annotations updated with curation of repeat elements ltr-1_family-1, ltr-1_family-4, and ltr-1_family-26. </p> <p>HSC_Genomic_FacC_860F_PREMIX_CNNJ42_1.ab1 and HSC_Genomic_FacC_1544R_PREMIX_CNNJ43_2.ab1 are Sanger sequenced PCR products for Lp_g42129 (Factor C) from primers FacC_860F.fasta and FacC_1544R.fasta.</p>
Genome assembly and annotation of an apple variety 'RubyMac'
<p>In this dataset, we provided the contig-level genome assembly and annotation of an apple tree called 'RubyMac', which is growing in Michigan, USA (43°04'53.1"N 85°43'13.5"W). In this tree, the upper branches carried a sport mutation as compared with lower branches.</p>
Polished Assemblies for "GoldPolish-Target: Targeted long-read genome assembly polishing"
<p>GoldPolish-Target is a targeted genome assembly polishing tool that uses long reads. We tested GoldPolish-Target on Oxford Nanopore Technologies datasets with a human cell line (NA24385) and Drosophila melanogaster (fruit fly). Here, we provide the data for the GoldRush baseline (unpolished) assembly and the GoldPolish-Target and medaka polished assemblies of these long-read datasets.</p>
Data from: A de novo chromosome-level genome assembly of Coregonus sp. "Balchen": one representative of the Swiss Alpine whitefish radiation
<p>Salmonids are of particular interest to evolutionary biologists due to their incredible diversity of life-history strategies and the speed at which many salmonid species have diversified. In Switzerland alone, over 30 species of Alpine whitefish from the subfamily Coregoninae have evolved since the last glacial maximum, with species exhibiting a diverse range of morphological and behavioural phenotypes. This, combined with the whole genome duplication which occurred in the ancestor of all salmonids, makes the Alpine whitefish radiation a particularly interesting system in which to study the genetic basis of adaptation and speciation and the impacts of ploidy changes and subsequent rediploidization on genome evolution. Although well curated genome assemblies exist for many species within Salmonidae, genomic resources for the subfamily Coregoninae are lacking. To assemble a whitefish reference genome, we carried out PacBio sequencing from one wild-caught <i>Coregonus sp. "Balchen" </i>from Lake Thun to ~90x coverage. PacBio reads were assembled independently using three different assemblers, Falcon, Canu and wtdbg2 and subsequently scaffolded with additional Hi-C data. All three assemblies were highly contiguous, had strong synteny to a previously published <i>Coregonus</i>linkage map, and when mapping additional short-read data to each of the assemblies, coverage was fairly even across most chromosome-scale scaffolds. Here, we present the first <i>de novo</i>genome assembly for the Salmonid subfamily Coregoninae. The final 2.2 Gb wtdbg2 assembly included 40 scaffolds, an N50 of 51.9 Mb, and was 93.3% complete for BUSCOs. The assembly consisted of ~52% TEs and contained 44,525 genes.</p>
A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation
<p>Supplementary data to "A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation"</p>
Genome sequencing and assembly of Lathyrus sativus
<p>The dataset contains the whole-genome assembly and protein sequences of <em>Lathyrus sativus</em> cultivar Pusa-24.</p>
Catalog of metagenome-assembled bacterial genomes from Antarctic endolithic communities
<p>The dataset consists of 2 rar archives and 2 files ( tab-separated values ). Here is a brief summary of their contents:</p> <ul> <li><strong>MAGs_taxonomy: </strong>GTDB classification for each MAG.</li> <li><strong>MAGs_genome_info: </strong>genome size, completeness, contamination, length, N50.</li> <li><strong>MAGs - candidate species: </strong>high quality (HQ) and medium quality (MQ) bacterial metagenome assembled genomes.</li> <li><strong>MAGs_Annotation: </strong>EggNOG annotation files. For each MAG, the following files are included: <ul> <li>eggnog.emapper.annotations: the final EggNOG annotation;</li> <li>eggnog.emapper.hmm_hits: list of significant hits to eggNOG Orthologous Groups</li> <li>eggnog.emapper.seed_orthologs: best match of each query within the best Orthologous Group (OG) reported in the eggnog.emapper.hmm_hits file<strong>.</strong></li> </ul> </li> </ul>
Lathyrus sativus LS007 genome assembly and annotation Rbp1.0
<p>Genome assembly of grass pea (<em>Lathyrus sativus</em> L.) genotype LS007, assembled from PromethION nanopore data and polished using Illumina HiSeq PE data. The assembly was annotated using the mikado-minos pipeline developed by the Earlham Institute. Also included is a separate annotation track for repeat sequences produced using the DANTE pipleline. </p> <p> </p> <p>For any questions regarding this dataset, contact peter.emmrich@jic.ac.uk</p> <p> </p> <p>Note: ctg14433 has been manually corrected based on sequenced amplicon data. Files have been updated accordingly.</p> <p> </p> <p><strong>Assembly files:</strong></p> <p>Lsativus_LS007_Rbp1.0.7z - compressed complete assembly without scaffolding. The annotation refers to this assembly</p> <p>Rbp_9 largest HiC scaffolds.7z - compressed fasta file of the largest 9 scaffolds following HiC scaffolding</p> <p>Lsat_LS007_Rbp_chloroplast.fasta - fasta file of the complete LS007 chloroplast genome</p> <p>Lsat_LS007_Rbp_mitochondrion.fasta - fasta file of the complete LS007 mitochondrial genome</p> <p> </p> <p><strong>Annotation tracks:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3</p> <p>DANTE_transposable_element_protein_domains.gff3</p> <p>Full_length_LTR_retrotransposons.gff3</p> <p>Repeat_annotation_classI_classII_satellites.gff3</p> <p> </p> <p><strong>Annotation FASTA files:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3.cds.fasta</p> <p>LATSA3860_EIv1.0.annotation.gff3.cdna.fasta</p> <p>LATSA3860_EIv1.0.annotation.gff3.pep.fasta</p> <p> </p> <p><strong>Summaries and statistics:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3.final_table.tsv</p> <p>LATSA3860_EIv1.0.annotation.gff3.mikado_stats.txt</p> <p>LATSA3860_EIv1.0.annotation.gff3.biotype_conf.summary</p> <p>LATSA3860_EIv1.0.annotation.gff3.final_table.tsv</p> <p>LATSA3860_EIv1.0.annotation.gff3.pep.fasta.functional_annotation.tsv</p> <p>NOT_UPDATED_LATSA3860_EIv1.0.annotation.gff3.metrics.tsv *</p> <p>Blobtools_passed_contigs.txt - list of all contigs of the assembly that pass the BlobTools filter (Streptophyta, 20-100x coverage, >50 kbp) </p> <p> </p> <p>*this file has not been updated to reflect the correction to ctg14433</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.