Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
Fraxinus pennsylvanica genome assembly and annotation
<p>We report the first chromosome-level assembly for green ash (<em>Fraxinus pennsylvanica</em>) to assist in ash breeding efforts to propagate resistance to the emerald ash borer. The final haploid assembly consists of 23 chromosomes and 87 unplaced scaffolds of 10 kb or more. Over 99% of the bases anchored to the chromosomes. The assembly spans 757 Mb and consists of 49.43% repetitive DNA. Gene annotation yielded 35,470 high-confidence gene models, all located on the chromosomes and assigned to 22,976 Asterid Orthogroups.</p>
Mycobacteroides abscessus subp. bolletii strain associated with a persistent infection (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a <strong><em>Mycobacteroides abscessus subp. bolletti </em></strong>strain associated with a persistente infection. (genome anotation was performed using Bakta v1.2.2 https://github.com/oschwengers/bakta)</p> <p>The raw sequence reads were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB57933; Run Accession: ERR10554471).</p>
The OHEJP BeONE Project – Escherichia coli genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 308 <em>Escherichia coli</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120057">https://zenodo.org/record/7120057</a>), comprising genome assemblies of 1,999 <em>E. coli</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Ec_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Ec_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>E. coli</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57098">PRJEB57098</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 308 isolates passed the dataset curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
The OHEJP BeONE Project – Campylobacter jejuni genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 610 <em>Campylobacter jejuni </em>samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120166">https://zenodo.org/record/7120166</a>), comprising genome assemblies of 3,076 <em>C. jejuni</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Cj_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and <em>in-silico</em> Multi Locus Sequence Type, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Cj_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>C. jejuni</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57119">PRJEB57119</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 610 isolates passed the dataset curation step and were included in the final dataset.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
The OHEJP BeONE Project – Listeria monocytogenes genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 1,426 <em>Listeria monocytogenes</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7116878">https://zenodo.org/record/7116878</a>), comprising genome assemblies of 1,874 <em>L. monocytogenes</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Lm_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and <em>in-silico</em> Multi Locus Sequence Type, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Lm_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>L. monocytogenes </em>genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57166">PRJEB57166</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,426 isolates passed the dataset curation step and were included in the final dataset.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p> <p> </p>
The OHEJP BeONE Project – Salmonella enterica genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 1,540 <em>Salmonella enterica</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7119735">https://zenodo.org/record/7119735</a>), comprising genome assemblies of 1,434 <em>S. enterica</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Se_metadata.xls</strong>x” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Se_assemblies.zi</strong>p” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>S. enterica</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="http://www.ebi.ac.uk/ena/browser/view/PRJEB57179">PRJEB57179</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,540 isolates passed the dataset curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31540993/">Zhang et al. 2019</a>).</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Genome assembly and annotation files for Corylus americana accessions 'Rush' and 'Winkler'
<p>The native shrub American hazelnut (<em>Corylus americana</em>) is currently used in breeding programs that are aiming to develop commercially viable hazelnut varieties for the U.S. Upper Midwestern U.S. This species provides significant ecological benefits as it is a perennial crop and well-adapted to this region. Breeding cycles for perennial species are long, and may benefit from the use of predictive methods such as genomic selection to reduce cycle time and increase the efficiency of field trials.</p> <p>High-quality reference genome assemblies are very useful for the implementation marker-assisted selection and genomic prediction, and we therefore developed the first chromosome-scale reference assemblies for <em>C. americana</em>, using the accessions 'Rush' and 'Winkler'. Initial draft assemblies were created using HiFi PacBio reads and Arima Hi-C sequencing to assemble genomes into 11 pseudomolecules. We then utilized Oxford Nanopore reads and a high-density genetic map in order to perform error correction. N50 scores were calculated to be 31.9 Mb and 35.3 Mb for 'Rush' and 'Winkler', respectively, while 97.1% (for 'Winkler') and 90.2% (for 'Rush') of the total genome was assembled into the 11 pseudomolecules. Gene prediction was performed using both RNAseq libraries as well as protein homology data. 'Rush' had a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while 'Winkler' had corresponding scores of 96.9 and 96.5, indicating extremely high-quality assemblies.</p> <p>These two independent, de novo assemblies enable unbiased assessment of structural variation across the genome, as well as patterns of syntenic relationships within C. americana and the <em>Corylus</em> genus. These assemblies are also an important first step in providing a resource for using next-generation sequencing data in the improvement of <em>C. americana</em>. We demonstrate this utility through the generation of high-density SNP marker sets from genotyping-by-sequencing data for 1,343 <em>C. americana</em>, <em>C. avellana</em>, and <em>C. americana</em> x <em>C. avellana</em> hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published <em>Corylus</em> genomes, were utilized to perform phylogenetic analysis of sporophytic self-incompatibility (SSI) in hazelnut, providing further evidence of unique molecular pathways governing self-incompatibility in Corylus not exhibited in other well-studied SSI systems. We hope these assemblies will aide in the application of modern breeding methods to the development of commercially viable hazelnut varieties for the U.S. Upper Midwest.</p>
Genome Assembly of Fusarium avenaceum
<p>Here, we present a complete genome assembly for <em>Fusarium avenaceum</em> and associated annotation using long-read sequencing generated from the Oxford Nanopore Technologies (ONT;London, UK) platform for both DNA and RNA obtained from fruit-sampled cultures. </p>
ScRAPv20230731: Telomere-to-telomere assemblies of 142 strains characterize the genome structural landscape in Saccharomyces cerevisiae
<p><strong><em>Saccharomyces cerevisiae </em>Reference Assembly Panel (ScRAP) v20230731 </strong>></p> <p>The haplotype-resolved and/or collapsed T2T genome assemblies for 142 <em>S. cerevisiae</em> strains isolated from diverse geographical and ecological niches.</p>
New Soil Metagenome-Assembled Genomes Catalogue Boosts Genetic Resources
<p><strong>Soil harbors a vast expanse of unidentified microbes, termed as microbial dark matter, presenting an untapped reservoir of microbial biodiversity and genetic resources, but has yet to be fully explored. In this study, we conducted the first large-scale excavation of soil microbial dark matter by reconstructing 40,039 metagenome-assembled genome bins (the SMAG catalog) from 3,304 soil metagenomes. We identified 16,530 of 21,077 species-level genome bins (SGBs) as unknown SGBs (uSGBs), which greatly expand archaeal and bacterial diversity across the tree of life. We also illustrate the pivotal role of uSGBs in augmenting soil microbiome's functional landscape and intra-species genome diversity, providing large proportions of the 43,169 biosynthetic gene clusters and 8,545 CRISPR-Cas genes. Additionally, we determined that uSGBs contributed 84.6% of novel viral-host associations identified from the SMAG catalog. Our results propose the SMAG catalog, a novel and expansive genomic resource that brings the soil microbial biodiversity and novel genetic resources to light.</strong></p>
Draft de novo genome assemblies of a male and female Amphibolurus muricatus (jacky dragon)
<p>Four de novo nuclear genome assemblies of <em>Amphibolurus muricatus</em></p> <p><strong>Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> • AmpMurF_1.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_1.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 1.1: Further scaffolding of assembly 1.0 using RNA-seq data</strong><br> • AmpMurF_1.1.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_1.1.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 2.0: Further scaffolding of assembly 1.0 using SLR-superscaffolder</strong><br> • AmpMurF_2.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_2.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing assembly</strong><br> • AmpMurF_3.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_3.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Methods<br> Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> Male and female <em>A. muricatus</em> genome sequencing libraries were constructed on the Chromium system (10x Genomics, Pleasanton, CA, USA) by the Ramaciotti Centre for Genomics (Sydney, Australia). The Chromium instrument enables unique barcoding of long stretches of DNA on gel beads. The barcodes allow later reconstruction of long DNA fragments from a series of short DNA fragments with the same barcode (i.e., linked-reads). After barcoding, DNA was sheared into smaller fragments and sequenced on the NovaSeq 6000 platform (Illumina, CA, USA) to generate 151 bp paired-end (PE) reads. A total of 904.9 M raw 10x Genomics Chromium linked-reads were generated. Raw 10x data were assembled with Supernova v2.1.1 (Weisenfeld et al., 2017) and a FASTA file was generated using the ‘pseudohap style’ option in Supernova mkoutput. All female (~450 M) and male (~550 M) read pairs were utilised (female sequencing depth ca 50.3×; male, ca 47.8×). The resulting assemblies was further scaffolded with ARKS v1.0.3 (Coombe et al., 2018), reusing the 10x reads, and the companion LINKS program (v1.8.7) (Warren et al., 2015). ARKS employs a <em>k</em>-mer approach to map linked barcodes to the contigs in the initial Supernova assembly to generate a scaffold graph with estimated distances for LINKS input. These assemblies were denoted AmpMurF_1.0 (female) and AmpMurM_1.0 (male). We used GapCloser v1.12 (part of SOAPdenovo2) (Luo et al., 2012) to fill gaps in the assembly. GapCloser was run using the parameter -l 150) and clean 10x Genomics reads PE reads. </p> <p><strong>Assembly 1.1: Further scaffolding using RNA-seq data</strong><br> We attempted to improve the v1.0 genome assemblies’ contiguity using RNA-sequencing reads. RNA-seq reads (from brain, ovary, and testis; see below) were filtered (i.e., cleaned) to remove adapters and low-quality reads using Flexbar v3.4.0 and used to further re-scaffold the v1.0 assemblies (FASTA files before gapclosing) with P_RNA_scaffolder (Zhu et al., 2018). The default Flexbar settings discards all reads with any uncalled bases. A final round of scaffolding was performed on the resulting assemblies using L_RNA_scaffolder (Xue et al., 2013). These assemblies were denoted AmpMurF_1.1 (female) and AmpMurM_1.1 (male). As before, GapCloser and clean 10x Genomics reads were used to fill gaps. </p> <p><strong>Assembly 2.0: Further scaffolding using SLR-superscaffolder</strong><br> As an alternative approach, we attempted to improve the v1.0 genome assemblies’ contiguity using SLR-superscaffolder (Guo et al., 2021). Briefly, SLR-superscaffolder employs single tube long fragment read (stLFR) sequencing (Wang et al., 2019) reads (see section below) to generate hybrid genome assemblies. The software was run with default parameters except for PE_SEED_MIN=300 (minimum contig size to fill; default 1000). These assemblies were denoted AmpMurF_2.0 (female) and AmpMurM_2.0 (male). GapCloser and clean stLFR reads (with the barcode removed using https://github.com/BGI-Qingdao/stLFR_barcode_split) were used to fill gaps. </p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing and supernova assembly</strong><br> We also generated independent assemblies for the individuals sequenced on the 10x Genomics Chromium system using single tube long fragment read (stLFR) sequencing (Wang et al., 2019). BGI (Brisbane, Australia) generated ~100×-coverage 100-bp paired-end reads (plus a 42-bp stLFR barcode on the right/_2 read). Low-quality reads, PCR duplicates, and adaptors were removed using SOAPnuke v1.5 (Chen et al. 2018). The stLFRdenovo pipeline (<a href="https://github.com/BGI-biotools/stLFRdenovo">https://github.com/BGI-biotools/stLFRdenovo</a>), which is based on Supernova and customized for stLFR data, was used to generate a <em>de novo</em> genome assembly. The stLFRdenovo tool ‘FillGaps’ was used to fill gaps.</p> <p><strong>References</strong><br> Chen, Y., Chen, Y., Shi, C., Huang, Z., Zhang, Y., Li, S., Li, Y., Ye, J., Yu, C., Li, Z., et al. (2018). SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 7, 1-6.<br> Coombe, L., Zhang, J., Vandervalk, B.P., Chu, J., Jackman, S.D., Birol, I., and Warren, R.L. (2018). ARKS: chromosome-scale scaffolding of human genome drafts with linked read kmers. BMC Bioinformatics 19, 234.<br> Guo, L., Xu, M., Wang, W., Gu, S., Zhao, X., Chen, F., Wang, O., Xu, X., Seim, I., Fan, G., et al. (2021). SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme. BMC Bioinformatics 22, 158.<br> Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18.<br> Wang, O., Chin, R., Cheng, X., Wu, M.K.Y., Mao, Q., Tang, J., Sun, Y., Anderson, E., Lam, H.K., Chen, D., et al. (2019). Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly. Genome Res 29, 798-808.<br> Warren, R.L., Yang, C., Vandervalk, B.P., Behsaz, B., Lagman, A., Jones, S.J., and Birol, I. (2015). LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads. Gigascience 4, 35.<br> Weisenfeld, N.I., Kumar, V., Shah, P., Church, D.M., and Jaffe, D.B. (2017). Direct determination of diploid genome sequences. Genome Res 27, 757-767.<br> Xue, W., Li, J.T., Zhu, Y.P., Hou, G.Y., Kong, X.F., Kuang, Y.Y., and Sun, X.W. (2013). L_RNA_scaffolder: scaffolding genomes with transcripts. BMC Genomics 14, 604.<br> Zhu, B.H., Xiao, J., Xue, W., Xu, G.C., Sun, M.Y., and Li, J.T. (2018). P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads. BMC Genomics 19, 175.</p>
Genome assemblies of four MDR B. fragilis isolates using PacBio data - supporting the PhD Thesis
<p>Unicycler and Canu assemblies using PacBio data of four MDR B. fragilis isolates.</p> <p>Data supporting the PhD Thesis <em>Epidemiology and Genomics of antimicrobial resistance in the Bacteroides fragilis group </em>by Thomas Vognbjerg Sydenham, The research unit of Clinical Microbiology, Department of Clinical Research, Faculty of Health Sciences, University of Southern Denmark September 2019.</p> <p> </p>
MetaBAT 2.12.1 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly
Genome binning of the gold standard pooled assembly <br><strong>Software: </strong>MetaBAT<br><strong>SoftwareVersion: </strong>2.12.1<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://bitbucket.org/berkeleylab/metabat<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> bowtie2-build anonymous_gsa_pooled.fasta anonymous_gsa_pooled.fasta<br>for i in {0..63}; do bowtie2 -q --threads 30 --fr -x anonymous_gsa_pooled.fasta --interleaved sample_${i}/anonymous_reads.fq -S anonymous_reads_sample_${i}.sam ; done<br>for i in {0..63}; do samtools view -b sample_${i}.sam -o anonymous_reads_sample_${i}.bam & done<br>for i in {0..63}; do samtools sort anonymous_reads_sample_${i}.bam -o anonymous_reads_sample_${i}.sorted.bam ; done<br>for i in {0..63}; do samtools index anonymous_reads_sample_${i}.sorted.bam ; done<br>runMetaBat.sh -l anonymous_gsa_pooled.fasta anonymous_reads_sample_*.sorted.bam
GC-MS data set for Generation of a chromosome-scale genome assembly of the insect-repellant terpenoid-producing Lamiaceae species, Callicarpa americana
<p>RAW GC/MS data set for characterization of class II terpene synthases from <em>Callicarpa americana </em></p>
Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""
<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. rufipogon</em> DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. glaberrima</em> IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. barthii</em> W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. nivara</em> W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p>
Prediction of nucleosome dyads for the K562 cell line in the hg19 genome assembly
<p><strong><a href="https://andre-rendeiro.com/2015/05/12/predicting_dyads_from_mnase">Predicting dyads from MNase-seq data</a></strong></p> <p>I needed the location of nucleosomal dyads in the K562 cell line (ENCODE tier 1 line). Surprisingly, although plenty of MNase-seq data for that cell line is available, no nucleosome and dyad prediction exists.</p> <p><strong>Running NuMap</strong></p> <p>I found the <a href="http://www-hsc.usc.edu/~valouev/NuMap/NuMap.html">NuMap</a> software by Anton Valouev to do exactly what I intended.</p> <p>Since it is in a somewhat obscure page and this seemed to be the only place where this software was, I have <a href="https://github.com/afrendeiro/NuMap">uploaded it into a Github repository</a> for the sake of preservation (<a href="https://github.com/orphancode/NuMap">https://github.com/orphancode/NuMap</a>).</p> <p>Predicting dyads from MNase-seq data with NuMap seemed trivial: I downloaded the <a href="http://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeSydhNsome/">K562 MNase-seq data set</a> (11 replicates ~85Gb!!), combined all replicates and ran NuMap on the data(instructions on the Github README).</p> <p>From NuMap output there are <a href="https://www.dropbox.com/s/asmp7bi40lrvtjb/K562_dyads.bed?dl=0">dyad positions in bed format</a> and you can also produce several metrics to evaluate how good the prediction was.</p> <p><strong>Distograms & phasograms</strong></p> <p>Valouev describes two measurements of the frequencies of distances between MNase-seq reads. The frequency of distances between reads mapping to opposite strands can be used to build a “distogram”, which ilustrates the expected nucleosome length (147 bp) - this is consistent across most eukaryotic cells. The frequency of distances between reads mapping to the same strand gives a measurement of the distance between nucleosomes, as they’re separated by some linker DNA - (Valouev calls this plot a “phasogram”). This measurement, on the other hand tends to be species and cell-type specific.</p> <p><strong>K562 predictions:</strong></p> <p>The expected 147 bp nucleosome length in K562 cells.</p> <p>The average distance between dyads in K562 cells seems to be 185 bp.</p> <p><strong>References:</strong></p> <p>Valouev, A., Johnson, S. M., Boyd, S. D., Smith, C. L., Fire, A. Z., Sidow, A. (2011). Determinants of nucleosome organization in primary human cells. Nature, 474(7352), 516–520. <a href="http://doi.org/10.1038/nature10002">http://doi.org/10.1038/nature10002</a></p>
A chromosome-level genome assembly of the woolly apple aphid, Eriosoma lanigerum (Hausman) (Hemiptera: Aphididae)
<p><strong><em>Eriosoma lanigerum</em> v1.0 frozen release</strong></p> <p>Genome assembly: Eriosoma_lanigerum.v1.0.scaffolds.fa.gz</p> <p>BRAKER2 gene models: Eriosoma_lanigerum.v1.0.scaffolds.gff</p> <p>BRAKER2 protein sequences: Eriosoma_lanigerum.v1.0.scaffolds.gff.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only): Eriosoma_lanigerum.v1.0.scaffolds.gff.aa.LTPG.fa</p> <p>BRAKER2 coding sequences: Eriosoma_lanigerum.v1.0.scaffolds.gff.cds.fa</p> <p><em>Buchnera aphidicola</em> scaffolds: Buchnera_aphidicola.scaffolds.fa</p> <p><strong>Aphid orthogroups</strong></p> <p>OrthoFinder run files (see for details <a href="https://github.com/davidemms/OrthoFinder/blob/master/OrthoFinder-manual.pdf">https://github.com/davidemms/OrthoFinder/blob/master/OrthoFinder-manual.pdf</a>): OrthoFinder_run.tar.gz</p>
Chromosomal-level genome assembly of the scimitar‐horned oryx: insights into diversity and demography of a species extinct in the wild
<p>Captive populations provide a valuable insurance against extinctions in the wild. However, they are also vulnerable to the negative impacts of inbreeding, selection and drift. Genetic information is therefore considered a critical aspect of conservation management. Recent developments in sequencing technologies have the potential to improve the outcomes of management programmes; however, the transfer of these approaches to applied conservation has been slow. The scimitar‐horned oryx (<i>Oryx dammah)</i> is a North African antelope that has been extinct in the wild since the early 1980s and is the focus of a large‐scale and long‐term reintroduction project. To enable the selection of suitable founder individuals, facilitate post‐release monitoring and improve captive breeding management, comprehensive genomic resources are required. Here, we used 10X Chromium sequencing together with Hi‐C contact mapping to develop a chromosomal‐level genome assembly for the species. The resulting assembly contained 29 chromosomes with a scaffold N50 of 100.4 Mb, and displayed strong chromosomal synteny with the cattle genome. Using resequencing data from six additional individuals, we demonstrated relatively high genetic diversity in the scimitar‐horned oryx compared to other mammals, despite it having experienced a strong founding event in captivity. Additionally, the level of diversity across populations varied according to management strategy. Finally, we uncovered a dynamic demographic history that coincided with periods of climate variation during the Pleistocene. Overall, our study provides a clear example of how genomic data can uncover valuable insights into captive populations and contributes important resources to guide future management decisions of an endangered species.</p>
Draft genome assembly of a Japanese Oikopleura dioica male individual (O3), using Nanopore long reads.
<p>This draft assembly was used to validate the chrY scaffolds of the OSKA2016 reference genome in the publication “A genome database for a Japanese population of the larvacean Oikopleura dioica”, Development Growth and Differentiation, Wang and coll., 2020 (in press). It is provided as supplemental data for the reproducibility of this work; please note that no further polishing has been done to correct sequencing errors.</p> <p>Genome sequence reads were produced on a MinION sequencer (Oxford Nanopore Technologies) using high-molecular weight DNA from a male individual of the Oikopleura dioica species of zooplankton. The individual was related to the laboratory strain established from a western Japanese population that was used to produce the OSKA2016 reference genome. The raw reads were basecalled with the Guppy software version 3.3.0 using its dna_r9.4.1_450bps algorithm, and deposited in the European Nucleotide Archive (Study ID: PRJEB38559). The draft assembly was made with the Flye software version 2.7 with the options --genome-size 65m and --min-overlap 3000.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.