Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Gila monster (Heloderma suspectum) genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of a male Gila monster (<em>Heloderma suspectum</em>). We annotated the genome using the Comparative Annotation Toolkit (CAT), and we have also included GFF3 files of the consensus gene set, output for each taxon included in this process.</p>
Myxococcus xanthus DZ2 Genome Assembly
<p><strong>We report a high-quality assembly and annotation of the <em>Myxococcus xanthus </em>strain DZ2 (CP080538) genome, using a combination of short Reads generated by the DNBSEQ™ (BGI Genomics), and Long High-Fidelity (HiFi) Reads generated by Pacific Biosciences (PacBio) Technologies.</strong></p>
GenoNet scores for human genome assembly GRCh38
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type-specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
The first annotated genome assembly of Macrophomina tecta associated with charcoal rot of sorghum
<p>Raw reads of Macrophomina tecta were obtained from Nanopore, Illumina, and NextSeq (RNA). Files with information about the genome annotation, functional prediction, repeats, effectors and orthologous genes are included. </p>
Genome assemblies of Xanthomonas oryzae pv. oryzae (PXO35, FXO38, Huang604) and Xanthomonas oryzae pv. oryzicola (BAI35, MAI23)
<p>Genome assemblies of <em>Xanthomonas oryzae</em> pv. <em>oryzae</em> (<em>Xoo</em>) and <em>Xanthomonas oryzae</em> pv. <em>oryzicola</em> (<em>Xoc</em>). Genome assemblies of the <em>Xoo</em> strains PXO35, FXO38, Huang604 and the <em>Xoc</em> strains BAI35, MAI23, have been generated with Flye based on ONT reads. For each of these strains, we corrected the sequences encoding for transcription activator-like effectors (TALEs) with our TALE-correction pipeline (https://github.com/Jstacs/Jstacs/tree/master/projects/talecorrect). For Xoo PXO35, we additionally provide assemblies based on reads obtained from different sequencing methods (Illumina, PacBio, ONT) generated by a collection of (hybrid) assembly strategies and different polishing approaches applied to combinations of these.</p>
Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction
<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the <em>P. teres </em>f.<em> maculata </em>isolate FGOB10Ptm-1. </p>
Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction
<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the <em>P. teres </em>f.<em> maculata </em>isolate P-A14. </p>
Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5
<p>Gene annotation files for <em>Fraxinus excelsior</em> genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB: The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 3,076 Campylobacter jejuni isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 2,794-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/4/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 3,076 <em>Campylobacter jejuni </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of Sequence Type [ST]). In total, 476 different STs are represented in this dataset, with ST21, ST50, ST48, ST45 and ST257 being the most represented ones and, together, corresponding to 29.1% of the dataset.</p> <p>File “Cj_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Cj_profiles_wgMLST.tsv” corresponds to a tab separated file with the 2,794-loci wgMLST profiles of each solate presented in the metadata file. The files “profiles/Cj_profiles_cgMLST_95.tsv”, “profiles/Cj_profiles_cgMLST_98.tsv” and “profiles/Cj_profiles_cgMLST_100.tsv” correspond to a 1,012-loci, 987-loci and 29-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>C. jejuni</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://pubmlst.org/organisms/campylobacter-jejunicoli">PubMLST</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 3,539 samples. The majority of them are associated with the INNUENDO project (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>). The remaining ones are associated with five BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB31119">PRJEB31119</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB38253">PRJEB38253</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB40238">PRJEB40238</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB4165">PRJEB4165</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA350537">PRJNA350537</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 3,076 isolates passed this curation step and were included in the final dataset. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 2,794-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/4">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 2,794-loci wgMLST profiles of the 3,076 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 1,012-loci, 987-loci and 29-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,999 Escherichia coli isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 7,601-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/5/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,999 <em>Escherichia coli </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 411 different serotypes are represented in this dataset, with O157:H7 being the most represented one, corresponding to 37.1% of the dataset.</p> <p>File “Ec_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Ec_profiles_wgMLST.tsv” corresponds to a tab separated file with the 7,601-loci wgMLST profiles of each isolate presented in the metadata file. The files “profiles/Ec_profiles_cgMLST_95.tsv”, “profiles/Ec_profiles_cgMLST_98.tsv” and “profiles/Ec_profiles_cgMLST_100.tsv” correspond to a 2,826-loci, 2,704-loci and 465-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>E. coli </em>genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 2,688 samples associated with three BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA230969">PRJNA230969</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB27020">PRJEB27020</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA248042">PRJNA248042</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,999 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 7,601-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/5">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 7,601-loci wgMLST profiles of the 1,999 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 2,826-loci, 2,704-loci and 465-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,434 Salmonella enterica isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 8,558-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/8/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,434 <em>Salmonella enterica </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 125 different serotypes are represented in this dataset, with Typhimurium (including monophasic), Enteritidis and Infantis being the most represented ones and, together, corresponding to 56.2% of the dataset.</p> <p>File “Se_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Se_profiles_wgMLST.tsv” corresponds to a tab separated file with the 8,558-loci wgMLST profiles of each isolate presented in the metadata file. The files “profiles/Se_profiles_cgMLST_95.tsv”, “profiles/Se_profiles_cgMLST_98.tsv” and “profiles/Se_profiles_cgMLST_100.tsv” correspond to a 3,261-loci, 3,179-loci and 874-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>S. enterica</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/senterica">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 1,779 samples associated with four BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB16326">PRJEB16326</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB20997">PRJEB20997</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB30335">PRJEB30335</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB39988">PRJEB39988</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,434 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31540993/">Zhang et al. 2019</a>). wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 8,558-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/8">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31<sup>st</sup>, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 8,558-loci wgMLST profiles of the 1,434 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 3,261-loci, 3,179-loci and 874-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
A Simulated Heterozygous Diploid Genome for Third-gen Sequencing, Assembly, and Curation
<p>A simulated heterozygous diploid genome based on <em>Saccharomyces</em> <em>cerevisiae</em>, and <em>S. paradoxus</em> homologous chromosomes.</p> <p>Simulated PacBio subreads were generated from both parent haplomes and mixed together. A phased assembly was produced using FALCON assembler and FALCON Unzip (doi:10.1038/nmeth.4035). This dataset and assembly were then used to validate the Purge Haplotigs pipeline (https://bitbucket.org/mroachawri/purge_haplotigs). See workflow.sh for commands, comments and file descriptions.</p>
Genome assemblies for "Versatile genome assembly evaluation with QUAST-LG"
<p>Supplementary data for A. Mikheenko, A. Prjibelski, V. Saveliev, D. Antipov, A. Gurevich. Versatile genome assembly evaluation with QUAST-LG. ISMB 2018 PROCEEDINGS (<em>Bioinformatics </em>journal)</p> <p>Reference genomes of</p> <ol> <li><em>Saccharomyces cerevisiae </em>(yeast) version R64-1-1 </li> <li><em>Caenorhabditis elegans </em>(worm) version WBcel235</li> <li><em>Drosophila melanogaster </em>(fruit fly) version BDGP6</li> </ol> <p>And<em> de novo</em> genome assemblies of</p> <ol> <li>Yeast_PB (<em>S. cerevisiae</em>, genome size: 12.1 Mb): Canu, FALCON, Flye, MaSuRCA, Miniasm (from Illumina pair-ends and PacBio SMRT)</li> <li>Yeast_NP (<em>S. cerevisiae</em>, genome size: 12.1 Mb): Canu, Flye, MaSuRCA, Miniasm (from Illumina pair-ends and Oxford Nanopore)</li> <li>Worm_PB (<em>C. elegans</em>, genome size: 100.3 Mb): Canu, FALCON, Flye, MaSuRCA, Miniasm (from Illumina pair-ends and PacBio SMRT)</li> <li>Fly_MP (<em>D. melanogaster</em>, genome size: 137.6 Mb): ABySS2, MaSuRCA, Meraculous, Platanus, SOAPdenovo2, SPAdes (from Illumina pair-ends and mate-pairs)</li> <li>Human_MP (<em>H. sapiens</em>, genome size: 3.1 Gb): UpperBound assembly only (from Illumina pair-ends and mate-pairs)</li> <li>Human_NP (<em>H. sapiens</em>, genome size: 3.1 Gb): UpperBound assembly only (from Illumina pair-ends and Oxford Nanopore)</li> </ol> <p>Each pack (items 1-4) is accompanied with the upper bound assembly created with QUAST-LG (for computing theoretical limits on assembly correctness and completeness for a particular genome and set of reads). For more information, interactive QUAST-LG reports, and links to <em>de novo</em> assemblies of the human datasets please visit http://cab.spbu.ru/software/quast-lg/ or write to <em>quast.support@cab.spbu.ru.</em></p>
Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes
<p> </p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p> </p> <p> RELEASE MAG-2018/01<br> --------------------------------------</p> <p> </p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p> </p> <p>2. LOCATION</p> <p> </p> <p> MAGs sequences are available under ENA project PRJEB25176.</p> <p> </p> <p> ENA_accession Isolate<br> -------------------- --------------<br> ERZ494218 AM_0118<br> ERZ494219 AM_0219<br> ERZ494220 AM_0226<br> ERZ494221 AM_0228<br> ERZ494222 AM_0233<br> ERZ494223 AM_0240<br> ERZ494224 AM_0244<br> ERZ494225 AM_0256<br> ERZ494226 AM_0268<br> ERZ494227 AM_0275<br> ERZ494228 AM_0466<br> ERZ494229 AM_0507<br> ERZ494230 AM_0510<br> ERZ494231 AM_0519<br> ERZ494232 AM_0528<br> ERZ494233 AM_0546<br> ERZ494234 AM_0608<br> ERZ494235 AM_0615<br> ERZ494236 AM_0616<br> ERZ494237 AM_0619<br> ERZ494238 AM_0621<br> ERZ494239 AM_0630<br> ERZ494240 AM_0643<br> ERZ494241 AM_0729<br> ERZ494242 AM_0764<br> ERZ494243 AM_0832<br> ERZ494244 AM_0849<br> ERZ494245 AM_0854<br> ERZ494246 AM_0876<br> ERZ494247 AM_0902<br> ERZ494248 AM_0936<br> ERZ494249 AM_1003<br> ERZ494250 AM_1104<br> ERZ494251 AM_1111<br> ERZ494252 AM_1205<br> ERZ494253 AM_1312<br> ERZ494254 AM_1409<br> ERZ494255 AM_1503<br> ERZ494256 AM_1603<br> ERZ494257 AM_1606<br> ERZ494258 AM_1801<br> ERZ494259 AM_1811<br> ERZ494260 AM_2104<br> ERZ494261 AM_2116<br> ERZ494262 AM_2124<br> ERZ494263 AM_2202<br> ERZ494264 AM_2207<br> ERZ494265 AM_2208<br> ERZ494266 AM_2324<br> ERZ494267 AM_2502<br> ERZ494268 AM_2804<br> </p> <p>3. ACKNOWLEDGEMENTS<br> </p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of São Carlos, São Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq), as well as, the spanish funding organ Consejo Superior de Investigaciones Científicas (CSIC).</p> <p>This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) - Finance Code 001.</p> <p> </p> <p>4. CONTACT INFORMATION</p> <p> Current curators:</p> <p> - Célio Dias Santos Júnior (celio.diasjunior@gmail.com)<br> - Flavio Henrique-Silva (dfhs@ufscar.br)<br> - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> </p> <p>5. COPYRIGHT NOTICE</p> <p> Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> Copyright (C) 2018 The AMnrGC consortium.</p> <p> This database is provided “as is” and without any warranty of any kind,<br> of openly available. You can redistribute and/or modify it<br> as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p> https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>
GenoNet scores for human genome assembly GRCh37
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
De novo genome assembly of the meadow brown butterfly, Maniola jurtina
<p>1. Whole-genome GFF file (raw and filtered for min. gene length) [<em>Maniola.jurtina.gff3</em>, <em>Maniola_jurtina_filtered.gff3</em>]</p> <p>2. List of <em>M. jurtina</em> proteins [<em>Mjurtina_proteins.fa</em>].</p> <p>3. Results of spot pattern genes BLAST against <em>M. jurtina</em> proteome [<em>Lepidoptera_MJ_protein_matches.xlsx</em>]. </p> <p>4. Annotations [blast2go_export.txt]</p>
ST131_4071_genome_assembly_annotation_files
<p>The genome assembly annotation files of 4,071 E. coli ST131 genomes (see Decano & Downing 2019).</p>
Sardinops sagax genome assemblies and annotations
<p>Files included are the genome assemblies for each haplotype (hap 1 and hap 2) of Sardinops sagax and the corresponding annotation files for each haplotype.</p> <p> </p> <p> </p>
LyBar v2.0 Genome Assembly and Annotation for Lycium barbarum
<p>LyBar v2.0 genome assembly and annotation files for <em>Lycium barbarum</em>.</p>
The Metagenome-Assembled Genome Inventory for Children (MAGIC)
<div> <div>Existing microbiota databases are biased towards adult samples, hampering accurate profiling of the infant gut microbiome. Here, we generated a **M**etagenome-**A**ssembled **G**enome **I**nventory for **C**hildren (**MAGIC**) from a large collection of bulk and viral-like particle-enriched metagenomes from 0-7 years of age, encompassing `3,299` prokaryotic and `139,624` viral species-level genomes, `8.5%` and `63.9%` of which are unique to MAGIC. MAGIC improves early-life microbiome profiling, with the greatest improvement in read mapping observed in Africans. We then identified `54` candidate keystone species, including several *Bifidobacterium spp.* and four phages, forming guilds that fluctuated in abundance with time. Their abundances were reduced in preterm infants and were associated with childhood allergies. By analyzing the *B. longum* pangenome, we found evidence of phage-mediated evolution and quorum sensing-related ecological adaptation. Together, the MAGIC database recovers genomes that enable characterization of dynamics of early-life microbiomes, identification of candidate keystone species, and strain-level study of target species.</div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.