Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Training Data for "Making sense of a newly assembled genome"
<p>The data provided here is part of the Galaxy Training Network tutorial for Making sense of a newly assembled genome. This data was sourced from NCBI on 2019-08-29</p>
Virulence and antibiotic resistance plasticity of Arcobacter butzleri: insights on the genomic diversity of an emerging human pathogen (genome assembly, annotation dataset, core- and pan-genome loci)
<p>This dataset refers to the analysis of 49 <em>Arcobacter butzleri</em> genomes and includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the predicted transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files), the respective amino acid sequences of the translated CDS sequences (.faa files), the nucleotide alignments of all the 1165 core-genome loci, the nucleotide alignments of the genes <em>hecA</em>, <em>tetR </em>and <em>porA</em>, the categorized amino acid sequences of the six hypervariable regions of PorA, and the nucleotide sequences of the first allele of each of the 7474 pan-genome loci with the respective complete allelic profile matrix.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB34441).</p>
Genome assemblies of Staphylococcus pseudintermedius
<p>Supplementary Dataset to the manuscript "Global phylogenomic analysis of Staphylococcus pseudintermedius reveals genomic and prophage diversity in multi-drug resistant lineages", by Lucy F. Grist, Alice Brown, Noel Fitzpatrick, Giuseppina Mariano, Roberto M. La Ragione, Arnoud H. M. van Vliet and Jai W. Mehat.</p> <p>The dataset contains 2,276 genome assemblies of Staphylococcus pseudintermedius, combining three different types of data, made available here for open research purposes:</p> <ul> <li>Genome assemblies of 110 new UK isolates not previously available. These have also been uploaded into NCBI genome, and have a SCpseud_UoSXXX filename and the SRR code for NCBI SRA.</li> <li>Genome assemblies (1,144) obtained from the NCBI Genome repository, recognisable by the StaphpseudXXXX_GCA filenames</li> <li>Genome assemblies (1,022) generated from sequencing reads obtained from NCBI SRA and the ENA repositories, recognisable by the StaphpseudXXXX_SRR and StaphpseudXXXX_ERR filenames.</li> </ul> <p>Genomes were assembled using Shovill version 1.1.0 (https://github.com/tseemann/shovill) using the default settings and the Spades assembler. Only genomes with N50>50 kb, L50<20, and a number of contigs <200 were included in this study. </p>
Metagenome Assembled Genomes (MAGs) from faecal microbiomes of great tits and blue tits
<h2><span>Overview:</span></h2> <p><span>The vertebrate gut microbiome plays crucial roles in host health and disease. However, there is limited data on the microbiomes of wild birds, most of which is restricted to barcode sequences. We therefore explored the use of shotgun metagenomics on the faecal microbiomes of two wild bird species widely used as model organisms in ecological studies: the great tit (<em>Parus major</em>) and the Eurasian blue tit (<em>Cyanistes caeruleus</em>). High and Medium quality Metagenome Assembled Genomes (MAGs) were assembled from these metagenomes and are made available as a catalogue in this archive.</span></p> <h2><span>Methods:</span></h2> <p><span><span>Metagenomic reads were trimmed, and quality controlled using FastP configured to a minimum phred score of 20 and minimum length of 50 bp</span><span>. In order to avoid contamination of the bins by eukaryotic sequences, Tiara v1.0.3 was used to classify contigs longer than 3.000 kb into their high-level kingdoms, allowing to exclude sequences of a eukaryotic or of an organelle origin, and only retaining all unclassified contigs and prokaryotic contigs for the binning step. <br>Contigs were binned using MaxBin2 v2.2.7 , SemiBin2 v2.1.0 and Metabat2 v 2.15 independently. The bins were refined using DasTool v 1.1.7 using a min score threshold of 0.3. The quality of the refined bins was obtained using CheckM2 v 1.0.2, and any bin with a contamination above 10% were excluded. The final MAGs were classified as Low-quality (<50% completeness, <10% contamination), medium-quality (>50% completeness, <10% contamination) and high-quality (>90% completeness, <5% contamination), as recommended by the MIMAG specification . Finally, the MAGs were dereplicated using an dRep v 3.4.3 with an ANI of 95% and classified using gtdb-tk v2.4.0 using the gtdb database release220.<br></span></span></p> <p><span><span>Files:</span></span></p> <ul> <li><span><span>The <strong>MAGs_sequences_v1.0.0 </strong>contains the fasta sequence for the individual MAGs assembled in this project</span></span></li> <li><span><span>The <strong>MAGs_catalogue_v1.0.0.xlsx</strong> contains a description of the quality, taxonomic annotation and characteristics of each MAGs in the dataset</span></span></li> </ul>
PkA1HT reference genome assembly (version 1.0)
<p> </p> <p><strong><em><span>PkA1HT reference genome assembly</span></em></strong></p> <p><span><span>Long-read PacBio HiFi sequencing was performed on two distinct <em>P. knowlesi piggyBac </em>clones (PkBc38 and PkBc44) to fully investigate potential structural variation (including gene duplication or deletion) amongst mutants that may have been overlooked by short-read sequencing. The <em>piggyBac</em> insertion was cut out of the PkBc38 clone assembly to generate a reference genome for our parental PkA1-H.1 line which we’ve named PkA1HT. The final PkA1HT reference genome was annotated with Companion (v2.2.8) using the <em>P. knowlesi</em> H strain as reference, specifying the assembly option with no contiguation</span><span>. The new PkA1HT reference genome has 18 sequences (comprised of 14 chromosomes, the mitochondrial genome, the apicoplast genome, and two unordered contigs). It has no sequencing gaps and is 25.29MB long, compared to 142 gaps and a smaller 24.32Mb for PkA1-H.1. Sequences for ten chromosomes reach into the telomeric heptamer repeats on both ends, and sequences for the four remaining chromosomes reach into the telomeric repeats for just one end (the two unordered contigs contain the remaining two unassembled chromosome ends).</span></span></p> <p> </p> <p><em><span>Additional methods</span></em></p> <p><span>Long-read PacBio HiFi sequencing was performed on two distinct <em>P. knowlesi piggyBac </em>clones (PkBc38 and PkBc44) to fully investigate potential structural variation (including gene duplication or deletion) amongst mutants that may have been overlooked by short-read sequencing. High molecular weight (HMW) DNA was extracted from parasite-infected human RBCs using the MagAttract HMW DNA Kit (Qiagen #67563), following the manufacturer’s protocol for the manual purification of genomic DNA from whole blood. DNA quality and quantity were assessed using a Qubit fluorometer and agarose gel electrophoresis, with samples meeting the criteria of a Qubit concentration >50 ng/µL and intact bands on the gel. Two µg of HMW genomic DNA was then sheared to an average size of ~15-20 kb, and SMRTbell libraries were prepared using the PacBio SMRTbell Prep Kit 3.0 (PacBio #102-141-700), which included DNA end-repair, adapter ligation, and nuclease treatment to remove incomplete molecules. An additional gel-based size selection step on the PippinHT was performed to remove fragments smaller than 10 kb. The final library was purified, quantified, and sequenced on the PacBio Revio platform (1x Revio Cell) to generate long-read data. Samples were generated at the University of South Florida and sequenced at the Wellcome Sanger Institute.</span></p> <p><span>Sequencing data were processed for quality control using standard PacBio workflows and a genome assembly was generated for each clone. We performed our long-read assemblies using the Canu assembler (v2.2) with default parameters, followed by polishing with ILRA (v1.5.1)</span><span>. The ILRA workflow included running ABACAS against the current PkA1-H.1 reference genome (v. 55) followed by two iterations of short-read correction using Pilon with Illumina reads of the parental line</span><span>. We performed manual finishing with the Artemis Comparison Tool (ACT)</span><span>, informed by long reads mapped back against the draft assemblies. We obtained >500x coverage from our PacBio sequencing with a median read length of 15kbp, and we had telomere-to-telomere completion on several chromosomes.</span></p> <p><span>Assemblies were then compared against each other using ACT. We found no significant structural variation between the two clonal lines, with near complete co-linearity save for each transposon insertion, indicating as expected that the <em>piggyBac </em>transposon insertion introduces no wider genomic changes. To compare differences between our parental line and the PkA1-H.1 reference, we used ACT to identify possible regions of recombination, followed by a mapping approach and manual analysis in Artemis for validation</span><span>. We found a duplication of 14 genes on chromosome 7 in both clones. We further found several synteny breaks between our <em>de novo</em> assemblies and the current PkA1-H.1 reference. Analyzing those “breakpoints” more closely in ACT, we found that they are actually misassemblies in the current reference This finding of misassemblies motivated us to generate a more complete reference genome (see next section). We otherwise found no evidence of recombination or large indels vs. the reference for either clone. <span>It should be noted that our PkA1HT assembly also supersedes both the PkH1 and PkA1 assemblies in terms of stats (details to be reported elsewhere).</span></span></p> <p> </p> <p> </p> <p> </p>
BIL20 and BIL24 Assembled, Annotated, and Mapped Genomes
<p>The dataset is featured in the article <em>Beyond Buro</em>, which investigates the probiotic potential of <em>Limosilactobacillus fermentum</em> BIL20 and BIL24 using genomic analysis and probiotic assays. These strains were isolated from <em>burong isda</em>, a traditional fermented fish product from Arayat, Pampanga. The study adds to the limited research from the Philippines that applies genomic analysis and probiotic testing on LAB isolates from fermented foods, revealing their promising probiotic attributes and potential health advantages. The dataset includes genome assemblies, RAST and Prokka annotations, and mapped assemblies locating the probiotic-related genes and their functional categories.</p>
Supplementary Tables on Gene Annotations of 49 Bacillariophyta Genome Assemblies
<p>Supplementary Tables attaining to the manuscript entitled <strong>Annotation of protein-coding genes in 49 diatom genomes from the Bacillariophyta clade. </strong>These Supplementary Tables describe in part the data foundation and results of the actual annotation dataset that is available at <a href="https://doi.org/10.5281/zenodo.13767023" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13767023</a>.</p> <p><strong>Supplementary Table S1</strong>: Genome assemblies available at NCBI Datasets in June 2024.</p> <p><strong>Supplementary Table S2</strong>: Genome assemblies excluded from annotation.</p> <p><strong>Supplementary Table S3</strong>: Accession numbers of genome assemblies and RNAseq libraries used for annotating 49 diatom genomes. Table also lists repeat content of genome assemblies after masking with RepeatModeler2/RepeatMasker.</p> <p><strong>Supplementary Table S4</strong>: Software and container versions used for annotating 49 diatom genomes.</p> <p><strong>Supplementary Table S5</strong>: Summary of EnTAP functional annotation results. The actual functional annotations are included in the annotation dataset that is available at <a href="https://doi.org/10.5281/zenodo.13767023" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13767023</a>.</p>
Wheat phyllosphere metagenome assembled genomes collected in Ringsted, Denmark
<p><span>We present a completely novel </span><span>and comprehensive wheat phyllosphere metagenomic dataset of 211 samples and </span><span>1261 MAGs. This dataset represents a significant contribution to the field, as it </span><span>provides insights into the poorly studied microbial communities associated with </span><span>wheat leaf surfaces, an ecosystem of considerable agricultural importance.</span></p>
Metagenome-Assembled Genomes and Annotations for McGivern et al
<h3>Files:</h3> <ul> <li><code>reactorEMERGE_annotations.txt</code>: DRAM annotations for MAGs</li> <li><code>gene_lengths.txt</code>: gene length file used to calculate geTMM</li> <li><code>genes.gff.tar.gz</code>: gff file needed for metaT processing</li> <li><code>genes.faa.tar.gz</code>: amino acid sequences for MAG genes</li> <li><code>genes.fna.tar.gz</code>: nucleotide sequences for MAG genes, used as database for metaT mapping</li> </ul>
AMF genome assemblies
<p>Genome assemblies of <em>Innospora majewskii</em> (Paraglomeraceae), <em>Claroideoglomus lamellosum </em>(Cloroideoglomeraceae) and <em>Septoglomus turnauae </em>(Glomeraceae). Please cite: </p> <p>Analyses of transposable elements in arbuscular mycorrhizal fungi support evolutionary parallels with filamentous plant pathogens<br>Jordana IN Oliveira, Catrina Lane, Ken Mugambi, Gokalp Yildirir, Ariane M Nicol, Vasilis Kokkoris, Claudia Banchini, Kasia Dadej, Jeremy Dettman, Franck Stefani, Nicolas Corradi. bioRxiv 2024.11.04.621924; doi: https://doi.org/10.1101/2024.11.04.621924</p>
Code used for the assembly of a wisent genome
<p>copy of the GitHub repository https://github.com/cbortoluzzi/WisentGenomeAssembly</p>
Viral metagenome assembled genomes (vMAGs) from Columbia River hyporheic sediments
<p>Fasta file containing 111 viral metagenome assembled genomes (vMAGs) from publication to be submitted titled "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Metagenome assembled genome (MAG) annotations for Columbia River sediment bacteria and archaea
<p>Excel spreadsheet containing all annotations for metagenome assembled genomes (MAGs) that form part of a publication to be submitted titled: "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Novel canine high-quality metagenome-assembled genomes by long-read metagenomics together with Hi-C proximity ligation
<p>We characterized a canine fecal sample of a healthy dog by combining a long-read metagenomics assembly (Nanopore sequencing) with Hi-C cross-linking data, and further correction of the frameshift errors. We retrieved and characterized 27 HQ MAGs and seven MQ MAGs considering MIMAG criteria.</p> <p>Find in this repository the final Hi-C genomics bins (CanMAG_XX-HiCbin.fa), including both the genome and the extra-chromosomal elements within the bin. </p> <p> </p>
Metagenome-assembled genomes(MAGs) generated from CRC human gut (PRJEB27928).
<p>MAGs generated from CRC human gut (PRJEB27928) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes(MAGs) generated from dog gut (PRJEB20308).
<p>MAGs generated from dog gut (PRJEB20308) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes(MAGs) generated from ocean (PRJEB1787).
<p>MAGs generated from ocean (PRJEB1787) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Nontuberculous mycobacteria persistence in a cell model mimicking alveolar macrophages (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following Nontuberculous mycobacteria (NTM) strains: <em>Mycobacterium smegmatis </em>mc<sup>2</sup>155 (reference strain), <em>Mycobacterium avium</em> ATCC25921 (reference strain), <em>M. avium </em>60/08 (clinical strain), <em>Mycobacterium fortuitum</em> ATCC6841 (reference strain) and <em>M. fortuitum</em> 747/08 (clinical strain).</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB30455).</p> <p>The associated article can be found here: <a href="https://www.ncbi.nlm.nih.gov/pubmed/31035520">https://www.ncbi.nlm.nih.gov/pubmed/31035520</a></p> <p> </p>
Chromosome-level genome assembly of Pterygoplichthys pardalis reveals its genetic basis of extensive invasion
<p>The catfish, <em>Pterygoplichthys</em> <em>pardalis</em>, which belongs to the Loricariidae family, an invasive species which has caused huge damage to the ecological environment. However, the high-quality reference genome for the catfish has not yet been reported. In this study, we successfully assembled the first chromosome-level high-quality genome of <em>P</em>. <em>pardalis</em> using the data we produced from multiple sequencing platforms, which contains 26 chromosomes and with a scaffold N50 of 49.47 Mb. Different evaluation methods all indicate the high connectivity and accuracy of the <em>P</em>. <em>pardalis</em> genome we got. We predicated 23,859 protein-coding genes in the <em>P</em>. <em>pardalis</em> genome, and 22,169 (~92.92%) coding genes could be functionally annotated in public databases. Phylogenetic relationship analysis found <em>P</em>. <em>pardalis</em> was clustered with all the catfishes we used and diverged with them 132.5 million years ago. Besides, whole-genome collinearity analysis found that chromosome 6 of <em>P</em>. <em>pardalis</em> was aligned to two distinct chromosomes both for <em>Ameiurus</em> <em>melas</em>, <em>Pangasianodon</em> <em>hypophthalmus</em> and <em>Ictalurus</em> <em>punctatus</em>, indicating that there may have been a chromosomal fusion/fission event occurred. Furthermore, many immune-system-related genes were large-scale expanded in <em>P</em>. <em>pardalis</em> genome, which may make great contributions to their adaptive traits, even for the highly polluted environmental conditions, and successful invasion. Taken together, this study not only provides insights into the genetic basis of the successful invasion of <em>P</em>. <em>pardalis</em>, but also provides important data resources for comparative genomic analysis of <em>P</em>. <em>pardalis</em> in Siluriformes in the future.</p>
Assembly and annotation of eleven Salix (shrub willow) genomes
<p>The shrub willows (<em>Salix</em> section <em>Vetrix</em>) are an emerging bioenergy crop in North America and Eurasia. However, genomics resources in this section are still quite limited, with only a few reference genomes available, despite many species in use in breeding programs. Here we present de novo assemblies and annotations of eleven shrub willow genomes from six species. Copy number variation of candidate sex determination genes within each genome was characterized and revealed remarkable differences in putative master regulator gene duplication and deletion. We also analyzed copy number and expression of candidate genes involved in floral secondary metabolism and identified substantial variation across genotypes, which can be used for parental selection in breeding programs. Lastly, we report on a genotype that produces only female descendants and identified gene presence/absence variation in the mitochondrial genome that may be responsible for this unusual inheritance.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.