Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Draft genome assembly version 1 of the meadow spittlebug Philaenus spumarius (Linnaeus, 1758) (Hemiptera, Aphrophoridae)
<p>We sequenced the genome of the meadow spittlebug, <em>Philaenus spumarius </em>(Linnaeus, 1758), the main insect vector of <em>Xylella fastidiosa </em>Wells et al. 1987 in Europe (Saponari et al., 2014), using 10x Chromium linked-reads. A single <em>P. spumarius</em> adult female from Portugal (Fontanelas, Sintra; GPS location: 38°50'15.75"N; 9°25'20.77"W), collected in September of 2018, was selected for genome sequencing. This population was initially surveyed for colour polymorphism in 1988 (Quartau & Borges, 1997) and was later included in phylogeographic and population genomic studies of this species (Rodrigues et al., 2014; Seabra et al., unpublished). It is also geographically close to the population from which the individual used for the first partial genome assembly was collected (Rodrigues et al., 2016). The availability of this previous genetic information contributed to the choice of this population as the source of genomic material for whole genome sequencing. A subset of males from the same collection date were analysed for genitalia morphology to confirm species identification, as the best diagnostic characters are the appendages of the aedeagus (Drosopoulos & Quartau, 2002).</p> <p>The genomic DNA of the <em>P. spumarius</em> adult from Sintra was extracted using Illustra Nucleon Phytopure kit according to the manufacturer’s instructions (GE Healthcare). We assessed the quality and concentration of the DNA using Femto fragment analyser (Agilent). 10x Chromium library preparation and Illumina genome sequencing (HiSeq X, 150bp paired-end) were performed by Novogene Bioinformatics Technology Co, Beijing, China, in accordance with standard protocols.</p> <p>To create the <em>de novo</em> 10x Chromium assembly we ran Supernova 2.1.1 (Weisenfeld et al., 2017) on the 10x Chromium linked-read data with default parameters, using 1.0 billion reads corresponding to 56X coverage. To improve the initial supernova assembly, we performed iterative scaffolding using all of the 10x raw data (2.3 billion of reads). We ran two rounds of Scaff10x (https://github.com/wtsi-hpag/Scaff10X), followed by mis-assembly detection and correction with Tigmint (Jackman et al., 2018). This was followed by a final round of scaffolding with ARCS (Yeo et al., 2018). The assembly was checked for contamination using the BlobTools pipeline (version 0.9.19; Laetsch and Blaxter 2017; Kumar et al., 2013) and k-mer content was analysed with the KAT comp tool (Mapleson et al., 2017). In order to perform these analyses, it was necessary to remove the 10x linked barcodes from the reads with the script process_10xReads.py (https://github.com/ucdavis-bioinformatics/proc10xG). We assessed the quality of our draft genome assembly by searching for conserved, single copy, arthropod genes (n=1,066) with Benchmarking Universal Single-Copy Orthologs (BUSCO) v3.0 (Waterhouse et al., 2018).</p> <p>With the above assembly procedure, we obtained a final assembly of 2.7 Gb, having a scaffold N50 length of 116 Kb (contig N50 = 18 Kb) and the longest scaffold was 3.7 Mb. The length of the assembly was consistent with the genome size estimated by flow cytometry (Rodrigues et al., 2016). The k-mer distribution indicated high heterozygosity, estimated at 2.3%. BlobTools analyses revealed the presence of contigs assigned to <em>Sodalis </em>spp. (Enterobacteriaceae), a symbiont in members of tribe Philaenini (Koga et al., 2013). These contigs were filtered from the final assembly. Gene completeness assessment shows that 956 (89.6%) among 1,066 BUSCOs were found as complete copies, with only 26 (2.4%) missing. Of the BUSCOs that were detected, 878 (82.4%) were complete and single-copy, 78 (7.3%) were complete and duplicated and 84 (7.9%) were fragmented.</p> <p>In conclusion, due in part to high (2.3%) heterozygosity levels, the <em>P. spumarius</em> version 1 genome assembly is highly fragmented. Nonetheless, the assembly is considered complete and is likely to contain the majority of the gene content of <em>P. spumarius.</em></p>
Gene Annotations of 49 Bacillariophyta Genome Assemblies
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div> </div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div> </div> <div>To extract the dataset, execute the following command:</div> <div> </div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div> </div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div> </div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div> </div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>
Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly
<p>Data and conda software environment file for the chapter '<em>de novo</em> Genome Assembly' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>
Genome, repeat, and functional annotation associated with the naked mole-rat genome assembly, mHetGlaV3 (GCA_964261345.1)
<p>The naked mole-rat (NMR; Heterocephalus glaber) is a eusocial subterranean rodent with a highly unusual set of physiological traits, such as extreme longevity, that has attracted great interest amongst the scientific community. However, the genetic basis of most of these traits has not been elucidated. To facilitate our understanding of the molecular mechanisms underlying NMR physiology and behaviour, we generated a long-read chromosomal-level genome assembly of the NMR. This genome, mHetGlaV2, was subsequently annotated and incorporated into a “91 eutherian mammals” multiple whole genome alignment in Ensembl. </p> <p>We identified intra-chromosomal misassemblies within mHetGlaV2. We fixed these misassemblies by comparing syntenic blocks between this assembly and the Canadian Porcupine (EreDor) genome assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_028451465.1/) and a FISH-Karyotype of the naked mole-rat completed by Romanenko et al., 2023 (PMID: 380307020) to address any misassemblies and place centromeres. Chromosome numbering was identified from a composite karyogram of karyotypes from over 350 cells. This scaffold-corrected assembly is labelled mHetGlaV3 (https://www.ebi.ac.uk/ena/browser/view/GCA_964261345.1).</p> <p>This repository stores the repeat, genome, and epigenome annotations for HetGlaV3.</p> <p>mHetGlaV3.primary.gtf.gz. Gene structures and gene symbols are transferred from ENSEMBL annotations of mHetGlaV2 using liftOff with default parameters. Additional gene symbols were identified using TOGA and manual curation.</p> <p>mHetGlaV3.primary.gtf.gz. Simple repetitive regions and transposable elements were annotated using EarlGrey (https://github.com/TobyBaril/EarlGrey) using "Rodentia" annotations for RepeatMasker.</p> <p>mHetGlaV3.primary.genesymbol_table.txt.txt.gz. A tab-delimited file where rows are gene IDs and columns are gene symbols generated with each method. "Consensus" shows the best matching gene symbol for each gene ID.</p> <p>mHetGlaV3.primary_annotated_blacklist.bed.gz. Provides an assembly "blacklist" for mHetGlaV3. This blacklist is a bed file annotating assembly breakpoints between HetGlaV2 and HetGlaV3. This blacklist contains additional columns (e.g., closest gene, overlapping TE etc.) and should therefore be filtered to the first column before being incorporated into traditional genomic pipelines.</p> <p>mHetGlaV3.primary_hypothalamus_ABC_enhancer.bedpe.gz. Activity-By-Contact enhancers (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction) generated in the female subordinate naked mole-rat hypothalamus using Hi-C-seq, ChIP-seq of H3K27Ac data, ATAC-seq, and RNA-seq information.</p> <p>mHetGlaV3.primary_hypothalamus_chromHMM.bed.gz. Chromatin states (using Chromhmm) annotating the female subordinate naked mole-rat hypothalamus using H3K4me3 (promoter), H4K4me2 (promoter-enhancer), H3K27Ac (active enhancer), H3K36me3 (elongated), H3K27me3 (polycomb repressed), H3K9me3 (heterochromatin), and CTCF (whole brain) ChIP-seq data, as well as ATAC-seq and RNA-seq data.</p> <p>mHetGlaV3.primary.fa.gz. Genome assembly fasta file for the naked mole-rat (V3, primary assembly). This assembly matches the primary assembly stored on ENA, however the chromosome names match these files, rather than have chromosome names processed by ENA (e.g. chr 1 instead of "OZ179169.1 Heterocephalus glaber genome assembly, chromosome: 1").</p> <p> </p> <p>UPDATES:</p> <p>* The 1.2 update fixed unscaffolded contig names from those used in-lab to those compatible with ENA.</p> <p>* The 1.3 update added small (50~100kbp) contigs onto mHetGlaV3.primary.fa.gz that were filtered before the ENA submission.</p> <p>* The 1.4 update fixed a small chromosome naming inconsistency spotted in the 1.3 update.</p>
Bin-assembled Escherichia coli genomes from a study in Punjab, Pakistan
<h2>Bin-assembled <em>Escherichia coli</em> genomes from Punjab, Pakistan</h2> <p>These assemblies are a part of a cross-sectional study conducted in Punjab, Pakistan aimed at investigating <em>E. coli</em> colonisation diversity in healthy carriage with the use of CLED enrichment plates.</p> <h3><strong>About</strong></h3> <h4><strong>Version history</strong></h4> <p><strong>v0.1.1 (current version)</strong></p> <ul> <li>Added reference to the study.</li> </ul> <p><strong>v0.1.0</strong></p> <ul> <li>Added brief description with a few missing parts.</li> </ul> <h4><strong>Distribution</strong></h4> <p>If you use these assemblies in your study please cite the source as appropriate. These assemblies are made available under a CC-BY 4.0 license.</p> <h4><strong>Citation</strong></h4> <p>Khawaja, T., Mäklin, T., Kallonen, T. et al. Deep sequencing of <em>Escherichia coli</em> exposes colonisation diversity and impact of antibiotics in Punjab, Pakistan. Nature Communications 15, 5196 (2024). <a href="https://doi.org/10.1038/s41467-024-49591-5">https://doi.org/10.1038/s41467-024-49591-5</a></p> <h3><strong>Methods briefly</strong></h3> <h4><strong>Species identification</strong></h4> <p>Sequencing data from the ENA project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB36642">PRJEB36642</a> was error-corrected with <a href="https://github.com/opengene/fastp">fastp</a> and pseudoaligned with <a href="https://github.com/algbio/themisto">Themisto</a> against a species-level index (available from <a href="https://doi.org/10.5281/zenodo.6656881">https://doi.org/10.5281/zenodo.6656881</a>). Reads were assigned to species using the <a href="https://doi.org/10.1099%2Fmgen.0.000691">mSWEEP/mGEMS pipeline</a> as described in <a href="https://www.nature.com/articles/s41467-022-35178-5">https://www.nature.com/articles/s41467-022-35178-5</a>.</p> <h4><strong>Lineage identification</strong></h4> <p>Read from the species-level bins were again pseudoaligned with Themisto against an <em>E. coli</em> index (will be made available in a later version). Lineage-level assignment was performed using mSWEEP and mGEMS at the level of <a href="https://genome.cshlp.org/content/29/2/304">PopPUNK</a> sequence clusters. The created bins were screened with <a href="https://github.com/tmaklin/coreutils_demix_check">demix_check</a> and bins that received a score of 1 or 2 were kept. Data in the kept bins were assembled with <a href="https://github.com/tseemann/shovill">shovill</a> and the bin-assembled genomes (BAGs) were quality controlled with <a href="https://genome.cshlp.org/content/25/7/1043">checkm</a> for >= 90% completeness and <= 10% contamination. Finally, BAGs shorter than 4 Mb or longer than 6 Mb were removed.</p> <h3><strong>Contact</strong></h3> <p>Tommi Mäklin <tommi'at'maklin.fi>.</p>
Kiwifruit Genome Assembly Red5 Version PS1.68.5
<p>Draft assembly pseudomolecules plus repeat and gene model annotation files for kiwifruit <em>Actinidia chinensis</em> var. <em>chinensis</em> 'Red5'. </p>
Actinidia eriantha Accession EA01_01 low coverage genome assembly
Draft assembly scaffolds of kiwifruit <i>Actinidia eriantha</i> 'EA01_01'. This is a female vine derived from seed collected on Qi-Yuan Mt., Min-Qing, Fukien on 15/11/75 by Li Lai-Yung, Professor of Subtropical Pomology, University of Fukien, Peoples Republic of China and provided to the New Zealand DSIR in 1975
Metagenome-assembled genomes from Stordalen Mire, Sweden (MAGs v2)
<p><strong>This release (MAGs v2) is a major new version of this metagenome-assembled genome (MAG) set.</strong> All previous releases on this page (which only differ in the metadata) are designated "MAGs v1." The current release (MAGs v2) uses<strong> </strong>CheckM2 v1.0.2 filtering (≥70% completeness, ≤10% contamination) to expand this dataset to include <strong>36,419 MAGs</strong>, with the following subcategories:</p> <ul> <li>Cronin_v1: Manually-curated subset of the "Field" category from MAGs v1.</li> <li>Cronin_v2: MAGs from raw bin filtering on the same assemblies used to generate Cronin_v1.</li> <li>Woodcroft_v2: MAGs from raw bin filtering on the same assemblies used to generate the MAGs reported in <a href="https://doi.org/10.1038/s41586-018-0338-1">Woodcroft & Singleton et al. (2018)</a>.</li> <li>SIPS: Updated genomes from samples originating from a stable isotope probing (SIP) incubation experiment by Moira Hough et al. ("SIP" in MAGs v1), re-analyzed due to read truncation and sample linkage issues in MAGs v1.</li> <li>JGI: Expanded set of genomes from the Joint Genome Institute's metagenome annotation pipeline.</li> </ul> <p> </p> <p>FILES:</p> <ul> <li><strong>Emerge_MAGs_v2.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_v2_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB taxonomy information, CheckM2 quality reports, NCBI GenomeBatch- and MIMAG(6.0)-formatted sample attributes and other metadata for the MAGs. </li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data collected at the Joint Genome Institute was generated under the following awards:</p> <ul> <li>The majority of sequencing at JGI was supported by BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</li> <li>Sequencing of SIP samples was performed under the Facilities Integrating Collaborations for User Science (FICUS) initiative (proposal 503547; award DOI: <a href="https://doi.org/10.46936/fics.proj.2017.49950/60006215">10.46936/fics.proj.2017.49950/60006215</a>) and used resources at the DOE Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://ror.org/04rc0xn13">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities. Both facilities are sponsored by the Office of Biological and Environmental Research and operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</li> </ul>
Metagenome-Assembled Genomes Abundance & Activity Tables. Environmental Parameters Associated with the dataset.
<p>Lake Mendota, WI, USA, is a temperate lake subject to annual temperature and oxygen fluctuations. Each summer, the water column becomes anoxic (no-oxygen). In 2020, we sampled the lake at weekly intervals, at different depths (5, 10, 15, 20 and 23.5m). For each time+depth sample, we collected metagenomes, viromes and metatranscriptomes. Environmental data profiles were collected on-site for each sampling day. </p> <p>Following standard metagenomic binning best practices, we obtained 431 metagenomes-assembled-genomes (MAGs).</p> <p>This record comprises the microbial abundance and expression table for these MAGs, and the environmental profiles collected each day.</p>
BBS phase 1 & phase 2 high quality E. coli bin assembled genomes
<p>1,402 <em>Escherichia coli</em> bin assembled genomes derived from the metagenome data collected as part of the <a href="https://www.ucl.ac.uk/global-health/research/a-z/baby-biome-study">BabyBiome study (BBS)</a> phase 1 & phase 2.</p> <p>The data in this upload was first published as part of "<em>Group 2 and 3 ABC-transporter dependant K-antigen loci contribute significantly to variation in the invasive potential of Escherichia coli"</em> (Gladstone et al. 2024, to be released).</p> <h2>Files</h2> <p>Assembly data:</p> <ul> <li>BBS_E_coli_BAGs.tar: Archive containing sequences of the 1,402 bin assembled genomes.</li> <li>BBS_E_coli_metadata.tsv: Table linking the sequence assemblies to the subject data.</li> </ul> <p>Capsule predictions:</p> <ul> <li>BBS_E_coli_Kaptive_output.csv: Capsule predictions for all sequence data.</li> <li>BBS_E_coli_deduplicated_sequences_IDs.txt: Filenames for assemblies that constitute the 873 deduplicated sequences analysed in Gladstone et al. 2024.</li> </ul> <p>Quality control data:</p> <ul> <li>BBS_E_coli_demix_check_scores.tsv: Output from demix_check for the sequence assemblies.</li> <li>BBS_E_coli_checkm_results.tsv: Output from checkm.</li> <li>BBS_E_coli_gunc_results.tsv: Output from gunc.</li> </ul> <h2>Methods</h2> <h3>Bin assembled genomes</h3> <p>Source data:</p> <ul> <li>BBS phase 1: <a href="https://doi.org/10.1038/s41586-019-1560-1">Shao et al. 2019</a></li> <li>BBS phase 2: <a href="https://doi.org/10.1038/s41564-024-01804-9">Shao et al. 2024</a></li> </ul> <p>The data was produced using the mSWEEP and mGEMS pipeline (<a href="https://doi.org/10.12688/wellcomeopenres.15639.2">Mäklin et al. 2020</a> & <a href="https://doi.org/10.1099/mgen.0.000691">Mäklin et al. 2021</a>) following the steps described in <a href="https://doi.org/10.1038/s41467-024-49591-5">Khawaja, Mäklin, Kallonen, et al. 2024</a>.</p> <h3>Quality control</h3> <p>The BAGs in this upload were filtered with demix_check (<a href="https://github.com/harry-thorpe/demix_check">https://github.com/harry-thorpe/demix_check</a>) and only those with a quality score 1 or 2 are included. For the capsule type annotations, contigs shorter than 5,000bp were removed but the short contigs are still present in the uploaded files). Further QC data is available from checkm (<a href="https://genome.cshlp.org/content/25/7/1043.short">Parks et al. 2015</a>) and gunc (<a href="https://link.springer.com/article/10.1186/s13059-021-02393-0">Orakov et al. 2022</a>) results.</p> <h3>Multilocus sequence typing</h3> <p>Sequence type (ST) was determined using fastmlst (<a href="https://journals.sagepub.com/doi/10.1177/11779322211059238">Guerrero-Araya et al. 2021</a>) with the `ecoli#1` database.</p> <h3>PopPUNK clustering</h3> <p>Sequence clusters (SC) correspond to the database available from <a href="https://zenodo.org/records/12528310">https://zenodo.org/records/12528310</a> and were created using PopPUNK (<a href="https://genome.cshlp.org/content/29/2/304.short">Lees et al. 2019</a>). Construction is described in <a href="https://doi.org/10.1038/s41467-024-49591-5">Khawaja, Mäklin, Kallonen, et al. 2024</a>.</p> <h3>Capsule type annotations</h3> <p>The capsule type annotations were created using Kaptive (<a href="https://doi.org/10.1099/mgen.0.000800">Lam et al. 2022</a>) with an <em>E. coli</em> specific database available from <a href="https://github.com/rgladstone/EC-K-typing">https://github.com/rgladstone/EC-K-typing</a> and described in Gladstone et al. 2024.</p>
Eigen scores for human genome assembly GRCh38 Part 4 (Chr1 - Chr2)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Eigen scores for human genome assembly GRCh38 Part 3 (Chr3 - Chr5)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Supplementary dataset to "Draft genome assembly of the biofuel grass crop Miscanthus sacchariflorus"
<p><em>Miscanthus sacchariflorus</em> (Maxim.) Hack. is a C4 perennial rhizomatous biofuel grass crop. <em>M. sacchariflorus</em> is among the most widely distributed species within the genus, particularly at cold northern latitudes, and one of the progenitor species of the main biomass commercial crop <em>M. × giganteus</em>. We generated a 2.54 Gbps whole-genome assembly of the diploid <em>M. sacchariflorus</em> “Robustus 297” genotype, which represented ~59% of the expected genome size. We later anchored this assembly in the chromosomal-scale <em>M. sinensis</em> genome to improve its contiguity. We annotated 86,767 and 69,049 protein-coding genes in the unanchored and anchored, respectively. We estimated our assemblies include ~85% of the <em>M. sacchariflorus</em> genes based on homology, core markers and RNA-seq alignments stats. Raw data and further metadata are available under Bioproject PRJNA435476.</p> <ul> <li>Msac_v2.fasta: Unanchored whole-genome assembly (WGA) of M. sacchariflorus in FASTA format.</li> <li>Msac_v3.fasta: The previous WGA re-scaffolded with the M. sinensis public reference.</li> <li>Msac_v3.agp: Chromosomal position in the M. sinensis reference of the previous scaffolds in Msac_v3.fasta</li> <li>Msac_v2.gff3: Gene annotation of the unanchored WGA in GFF3 format, which contains 86,767 coding genes</li> <li>Msac_v3.gff3: Gene annotation of the anchored WGA in GFF3 format, which contains 69,049 coding genes</li> <li>Msac_v2.func_annot.tsv: Text table containing the functional annotation of the 86,767 coding genes in Msac_v2.gff3</li> <li>Msac_v2.repeats_annotation.gff3: Repeats annotation (Repeatmasker) of the unanchored reference.</li> <li>Msac_v2.masked.fasta.gz: Repeats-masked version (Repeatmasker) of Msac_v2.fasta</li> <li>all.satsuma.blocks_Msac_v2-vs-Msin.gz: Every alignment from scaffolds in Msac_v3.fasta into M. sinensis reference</li> <li>Msac_v2.orthology_Msin.tsv: Ortologous between Msac_v2 and M. sinensis</li> <li>Msac_v3-vs-Msin.tsv: Ortologous between Msac_v3 and M. sinensis</li> </ul>
Predicted genes from the Amblyomma americanum draft genome assembly
<p>Data for pub "Predicted genes from the <em>Amblyomma americanum </em>draft genome assembly."</p> <ul> <li>Amblyomma_americanum_filtered_assembly.fasta: Decontaminated A. americanum genome with bacterial contigs removed</li> <li>Amblyomma_americanum_bacterial_contigs_info.tsv: Information about contigs classified as bacteria that were removed</li> <li>Amblyomma_americanum_annotation_data.tar.gz: Directory of annotation data produced by EvidenceModeler as part of the nf-core/genomeannotator workflow. Includes files in FASTA format (predicted genes and proteins), set of proteins clustered at 99% identity in FASTA format, and annotations in both GFF3 and GTF formats. GTF file produced from the GFF3 file with AGAT.</li> <li>Amblyomma_americanum_transcriptome_assembly_data.tar.gz: Directory of data generated for the transcriptome assembly that was used for gene prediction</li> </ul>
Penicillium fuscoglaucum Pf_T2 Genome Assembly and Annotation
<p>During routine culturing on selective media in the lab, we obtained an isolate of P. fuscoglaucum Pf_T2 and sequenced its genome. The Pf_T2 genome is far superior to available genomic resources for the species. Our assembly exhibits a length of 35.1 Mb, a BUSCO score of 97.9% complete, and consists of 5 scaffolds/contigs representing the four expected chromosomes. It was determined that the Pf_T2 genome was colinear with a type specimen P. fuscoglaucum, and contained a lineage specific, intact cylcopaizonic acid (CPA) gene cluster.</p>
Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)
<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 & 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by <a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads. </p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a> (v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1) using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18) and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0) and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2) and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p> </p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>
Catalog of stool metagenome-assembled genomes from patients with different cancer types
<p><strong>A non-redundant catalog of 3,816 genomes with at least 75% completeness and no more than 15% contamination assembled from metagenomes. Samples of 976 metagenomes were obtained from patients receiving immunotherapy for the treatment of different types of cancers.</strong></p>
Data from: Genome assembly of Danaus chrysippus and comparison with the Monarch Danaus plexippus
<p><strong>GENOME ASSEMBLY DATA</strong></p> <p><strong>Dchry2_unaltered.fa.gz</strong><br> Unaltered version of Dchry2 before manual edits</p> <p><strong>Dchry2.haplotigs.fasta.gz</strong><br> Haplotypic contigs removed in the Purge_haplotigs step</p> <p><strong>Dchry2_to_Dchry2.2_transfers.txt</strong><br> Edits of the Dchry2 assembly to produce Dchry2.2 (contig breaks, reverse complements and name changes). Columns are: new contig name, new contig start, new contig end, original contig name, original contig start, original contig end, orientation. New contig numbers indicate the chromosome they correspond to.</p> <p><strong>mxv1.200520.ragoo.rnm.fa.gz</strong><br> Fasta file from of MEX_DaPlex assembly with mxdp_ fasta headers</p> <p><br> <strong>GENE AND REPEAT ANNOTATION</strong></p> <p><strong>danaus_plex_mex_braker_a002.sequences.tidy.gff3.gz</strong><br> Sorted and tidied gff3 from MEX_DaPlex re-annotation </p> <p><strong>danaus_plexv4_braker_a006.sequences.tidy.gff3.gz</strong><br> Sorted and tidied gff3 from Dplex_v4 re-annotation </p> <p><strong>mxv1.200520.ragoo.rnm.wDpv3.gff3.gz</strong><br> gff3 file from Cei of MEX_DaPlex annotation</p> <p><strong>dplex2_uniprot-proteome_UP000596680_and_dplex_mex.fasta.gz</strong><br> Protein set used for annotation of all three assemblies</p> <p><strong>functional_annotation_output.tar.gz</strong><br> all output from pannzer2 regarding functinal annotation of the Dchry2.2 genome</p> <p><strong>Lepidoptera_and_danaus_chrysippus2.2.repeatmasker.gz</strong><br> Custom repeat library (combination of lepidoptera and specific dchry2.2 library)</p>
Actinidia chinensis Red5 genome assembly (version 2) and annotation files
<p>We present version 2 of the genome assembly for <em>Actinidia chinensis</em> var. <em>chinensis</em> genotype Red5. The Red5 genome was originally assembled using short read Illumina data (Pilkington et al, 2018; <a href="https://doi.org/10.1186/s12864-018-4656-3">https://doi.org/10.1186/s12864-018-4656-3</a>). In version 2 we employed Pacific BioSciences Sequel Single Molecule Real Time (SMRT) sequencing technology in place of Illumina paired end read sequencing for the main assembly but leveraged that short read data (Pilkington et al, 2018) for post assembly base correction of long read assembly contigs. Additionally the Illumina long insert libraries from Pilkington et al (2018) were used for post assembly scaffolding of contigs. Scaffold assignment to linkage groups leveraged the genetic map described in Pilkington et al (2018) as well as consensus evidence from DNA synteny comparisons to existing whole genome sequences from <em>Actinidia</em>.</p> <p>To meet the file size restrictions some dataset components have been split into multiple parts.</p> <p><strong>Assembly</strong></p> <p>The assembly work flow used the FALCON/FALCON-unzip assembly suite is described in Red5_version_2_genome_assembly.md. The assembly yielded both primary and haplotig contig data sets, the metrics for which are documented in this file. The CDS and predicted peptide fasta and GFF3 gene annotation for the primary and haplotig sets are provided in separate files.</p> <p><strong>File Descriptions</strong></p> <ul> <li>Files named chr1.fasta to chr29.fasta represent the primary assembly linkage group level assembly units</li> <li>Files named haplotig_part_1.fasta to haplotig_part_10.fasta represent the haplotig contig sets split into 10 parts to meet upload file size restrictions</li> <li>Files named primary_assembly.primary.gff3 and haplotig.gff3 contain the gene model annotations for the primary and haplotig assembly datasets respectively</li> <li>primary_assembly.cds.fasta and primary_assembly.pep.fasta contain the CDS and peptide sequences for the annotations on the primary contigs</li> <li>haplotig.cds.fasta and haplotig.pep.fasta contain the CDS and peptide sequences for the annotations on the haplotig contigs</li> <li>haplotigs.placements.tsv and haplotigs.reassignments.tsv describe the placement of haplotigs relative to the primary contigs as derived from purge_haplotigs</li> <li>The file Red5_version_2_genome_assembly.md describes the assembly work flow and code steps used as well as assembly metrics</li> <li>Files HYV3_1.v.R5V2_1.png to HYV3_29.v.R5V2_29.png depict Circos plots of DNA:DNA synteny based on 1coords alignment filter of nucmer alignments using dnadiff</li> </ul> <p>See Red5_version_2_genome_assembly.md for description of assembly methods and assembly metrics.</p> <p><strong>Funding</strong></p> <p>This work was funded by Kiwifruit Royalty Investment Program by The New Zealand Institute for Plant & Food Research Ltd. with support from Zespri, and the CORE grant Endeavour Smart Idea Fund (UOOX1801) from the New Zealand Ministry of Business, Innovation and Employment (MBIE). The funding bodies had no role in the design of the study, the collection, analysis, or interpretation of data or writing this manuscript.</p>
Annotation of the the assembled genome of Fusarium oxysporum f. sp. albedinis strain 133, the causal agent of date palm dieback.
<p>Annotation of the the assembled genome of <em>Fusarium oxysporum f. sp. albedinis</em> strain 133 (Khayi et al., 2020). Gene prediction and annotation were carried out using funnotate pipeline v1.8.1 (Stajich, 2020), which includes masking, ab initio gene-prediction training, using Augustus and Genmark, with the EST dataset reported to the Ganoderma mycocosm repository, gene prediction, and the assignment of functional annotation to protein-coding gene models.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.