Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

660

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

660 results for “genome assembly”

Learn how ShareScore rates datasets ↗
zenodo44/100

Gila monster (Heloderma suspectum) genome assembly and annotation

<p><em>De novo</em>&nbsp;genome assembly and annotation of a male Gila monster (<em>Heloderma suspectum</em>). We annotated the genome&nbsp;using the Comparative Annotation Toolkit (CAT), and we have also included GFF3 files of the consensus gene set,&nbsp;output for each taxon included in this process.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Myxococcus xanthus DZ2 Genome Assembly

<p><strong>We report a high-quality assembly and annotation of the <em>Myxococcus xanthus </em>strain DZ2 (CP080538) genome, using a combination of short Reads generated by the DNBSEQ&trade; (BGI Genomics), and Long High-Fidelity (HiFi) Reads generated by Pacific Biosciences (PacBio) Technologies.</strong></p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

GenoNet scores for human genome assembly GRCh38

<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type-specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files.&nbsp;</p> <p>Each row represents a genomic region with 131 columns. Please find the header line in &quot;genonet.header.txt&quot;.&nbsp;</p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 &nbsp; &nbsp;10000 &nbsp; &nbsp;10025 &nbsp; &nbsp;chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&amp;usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37&nbsp;<a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover&nbsp;<a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

The first annotated genome assembly of Macrophomina tecta associated with charcoal rot of sorghum

<p>Raw reads of Macrophomina tecta were obtained from Nanopore, Illumina,&nbsp;and NextSeq (RNA). Files with information about the genome annotation, functional prediction, repeats, effectors and orthologous genes are included.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Genome assemblies of Xanthomonas oryzae pv. oryzae (PXO35, FXO38, Huang604) and Xanthomonas oryzae pv. oryzicola (BAI35, MAI23)

<p>Genome assemblies of <em>Xanthomonas oryzae</em> pv. <em>oryzae</em> (<em>Xoo</em>) and <em>Xanthomonas oryzae</em> pv. <em>oryzicola</em> (<em>Xoc</em>). Genome assemblies of the <em>Xoo</em> strains PXO35, FXO38, Huang604 and the <em>Xoc</em> strains BAI35, MAI23, have been generated with Flye based on ONT reads. For each of these strains, we corrected the sequences encoding for transcription activator-like effectors (TALEs) with our TALE-correction pipeline (https://github.com/Jstacs/Jstacs/tree/master/projects/talecorrect). For Xoo PXO35, we additionally provide assemblies based on reads obtained from different sequencing methods (Illumina, PacBio, ONT) generated by a collection of (hybrid) assembly strategies and different polishing approaches applied to combinations of these.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction

<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the&nbsp;<em>P. teres&nbsp;</em>f.<em>&nbsp;maculata&nbsp;</em>isolate FGOB10Ptm-1.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction

<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the&nbsp;<em>P. teres&nbsp;</em>f.<em>&nbsp;maculata&nbsp;</em>isolate P-A14.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5

<p>Gene annotation files for&nbsp;<em>Fraxinus excelsior</em>&nbsp;genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These&nbsp;files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB:&nbsp;The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>

opencc-by-4.0Dec 2016View details →
zenodo44/100

Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 3,076 Campylobacter jejuni isolates

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 2,794-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/4/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 3,076 <em>Campylobacter jejuni </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of Sequence Type [ST]). In total, 476 different STs are represented in this dataset, with ST21, ST50, ST48, ST45 and ST257 being the most represented ones and, together, corresponding to 29.1% of the dataset.</p> <p>File &ldquo;Cj_metadata.xlsx&rdquo; contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST.</p> <p>The directory &ldquo;assemblies/&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.&nbsp;</p> <p>The file &ldquo;profiles/Cj_profiles_wgMLST.tsv&rdquo; corresponds to a tab separated file with the 2,794-loci wgMLST profiles of each solate presented in the metadata file. The files &ldquo;profiles/Cj_profiles_cgMLST_95.tsv&rdquo;, &ldquo;profiles/Cj_profiles_cgMLST_98.tsv&rdquo; and &ldquo;profiles/Cj_profiles_cgMLST_100.tsv&rdquo; correspond to a 1,012-loci, 987-loci and 29-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>C. jejuni</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://pubmlst.org/organisms/campylobacter-jejunicoli">PubMLST</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 3,539 samples. The majority of them are associated with the INNUENDO project (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>). The remaining ones are associated with five BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB31119">PRJEB31119</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB38253">PRJEB38253</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB40238">PRJEB40238</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB4165">PRJEB4165</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA350537">PRJNA350537</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 3,076 isolates passed this curation step and were included in the final dataset. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 2,794-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/4">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mix&atilde;o et al. 2022</a>) using the 2,794-loci wgMLST profiles of the 3,076 isolates as input and setting distinct &ldquo;--site-inclusion&rdquo; thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 1,012-loci, 987-loci and 29-loci allelic matrices, respectively).</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,999 Escherichia coli isolates

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 7,601-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/5/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,999 <em>Escherichia coli </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 411 different serotypes are represented in this dataset, with O157:H7 being the most represented one, corresponding to 37.1% of the dataset.</p> <p>File &ldquo;Ec_metadata.xlsx&rdquo; contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory &ldquo;assemblies/&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.&nbsp;</p> <p>The file &ldquo;profiles/Ec_profiles_wgMLST.tsv&rdquo; corresponds to a tab separated file with the 7,601-loci wgMLST profiles of each isolate presented in the metadata file. The files &ldquo;profiles/Ec_profiles_cgMLST_95.tsv&rdquo;, &ldquo;profiles/Ec_profiles_cgMLST_98.tsv&rdquo; and &ldquo;profiles/Ec_profiles_cgMLST_100.tsv&rdquo; correspond to a 2,826-loci, 2,704-loci and 465-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>E. coli </em>genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 2,688 samples associated with three BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA230969">PRJNA230969</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB27020">PRJEB27020</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA248042">PRJNA248042</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,999 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 7,601-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/5">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mix&atilde;o et al. 2022</a>) using the 7,601-loci wgMLST profiles of the 1,999 isolates as input and setting distinct &ldquo;--site-inclusion&rdquo; thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 2,826-loci, 2,704-loci and 465-loci allelic matrices, respectively).</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,434 Salmonella enterica isolates

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 8,558-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/8/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,434 <em>Salmonella enterica </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 125 different serotypes are represented in this dataset, with Typhimurium (including monophasic), Enteritidis and Infantis being the most represented ones and, together, corresponding to 56.2% of the dataset.</p> <p>File &ldquo;Se_metadata.xlsx&rdquo; contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory &ldquo;assemblies/&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.&nbsp;</p> <p>The file &ldquo;profiles/Se_profiles_wgMLST.tsv&rdquo; corresponds to a tab separated file with the 8,558-loci wgMLST profiles of each isolate presented in the metadata file. The files &ldquo;profiles/Se_profiles_cgMLST_95.tsv&rdquo;, &ldquo;profiles/Se_profiles_cgMLST_98.tsv&rdquo; and &ldquo;profiles/Se_profiles_cgMLST_100.tsv&rdquo; correspond to a 3,261-loci, 3,179-loci and 874-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>S. enterica</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/senterica">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 1,779 samples associated with four BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB16326">PRJEB16326</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB20997">PRJEB20997</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB30335">PRJEB30335</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB39988">PRJEB39988</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,434 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31540993/">Zhang et al. 2019</a>). wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 8,558-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/8">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31<sup>st</sup>, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mix&atilde;o et al. 2022</a>) using the 8,558-loci wgMLST profiles of the 1,434 isolates as input and setting distinct &ldquo;--site-inclusion&rdquo; thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 3,261-loci, 3,179-loci and 874-loci allelic matrices, respectively).</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

A Simulated Heterozygous Diploid Genome for Third-gen Sequencing, Assembly, and Curation

<p>A simulated heterozygous diploid genome based on <em>Saccharomyces</em> <em>cerevisiae</em>, and <em>S. paradoxus</em> homologous chromosomes.</p> <p>Simulated PacBio subreads were generated from both parent haplomes and mixed together. A phased assembly was produced using FALCON assembler and FALCON Unzip (doi:10.1038/nmeth.4035). This dataset and assembly were then used to validate the Purge Haplotigs pipeline (https://bitbucket.org/mroachawri/purge_haplotigs). See workflow.sh for commands, comments and file descriptions.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Genome assemblies for "Versatile genome assembly evaluation with QUAST-LG"

<p>Supplementary data for&nbsp;A. Mikheenko, A. Prjibelski, V. Saveliev, D. Antipov, A. Gurevich.&nbsp;Versatile genome assembly evaluation with QUAST-LG.&nbsp;ISMB 2018 PROCEEDINGS (<em>Bioinformatics </em>journal)</p> <p>Reference genomes of</p> <ol> <li><em>Saccharomyces cerevisiae </em>(yeast)&nbsp;version&nbsp;R64-1-1&nbsp;</li> <li><em>Caenorhabditis&nbsp;elegans&nbsp;</em>(worm)&nbsp;version&nbsp;WBcel235</li> <li><em>Drosophila&nbsp;melanogaster&nbsp;</em>(fruit fly) version&nbsp;BDGP6</li> </ol> <p>And<em> de novo</em> genome assemblies of</p> <ol> <li>Yeast_PB (<em>S. cerevisiae</em>, genome size: 12.1 Mb): Canu, FALCON, Flye, MaSuRCA,&nbsp;Miniasm (from Illumina pair-ends&nbsp;and&nbsp;PacBio SMRT)</li> <li>Yeast_NP (<em>S. cerevisiae</em>, genome size: 12.1 Mb):&nbsp;Canu, Flye, MaSuRCA,&nbsp;Miniasm (from Illumina pair-ends&nbsp;and Oxford Nanopore)</li> <li>Worm_PB (<em>C. elegans</em>, genome size:&nbsp;100.3 Mb):&nbsp;Canu, FALCON, Flye, MaSuRCA,&nbsp;Miniasm (from Illumina pair-ends&nbsp;and&nbsp;PacBio SMRT)</li> <li>Fly_MP (<em>D. melanogaster</em>,&nbsp;genome size:&nbsp;137.6 Mb): ABySS2, MaSuRCA,&nbsp;Meraculous, Platanus, SOAPdenovo2, SPAdes (from Illumina pair-ends&nbsp;and&nbsp;mate-pairs)</li> <li>Human_MP (<em>H. sapiens</em>, genome size: 3.1 Gb): UpperBound assembly only (from Illumina pair-ends&nbsp;and&nbsp;mate-pairs)</li> <li>Human_NP (<em>H. sapiens</em>, genome size: 3.1 Gb): UpperBound assembly only (from Illumina pair-ends&nbsp;and&nbsp;Oxford Nanopore)</li> </ol> <p>Each pack (items 1-4) is accompanied with the upper bound&nbsp;assembly created with QUAST-LG (for computing theoretical limits on assembly correctness and completeness for a particular genome and set of reads). For more information, interactive QUAST-LG reports, and links to <em>de novo</em> assemblies of the human datasets please visit&nbsp;http://cab.spbu.ru/software/quast-lg/ or write to&nbsp;<em>quast.support@cab.spbu.ru.</em></p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes

<p>&nbsp;</p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p>&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; RELEASE MAG-2018/01<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --------------------------------------</p> <p>&nbsp;</p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p>&nbsp;</p> <p>2. LOCATION</p> <p>&nbsp;</p> <p>&nbsp;&nbsp; MAGs sequences are available under ENA project PRJEB25176.</p> <p>&nbsp;&nbsp;</p> <p>&nbsp;&nbsp;&nbsp; ENA_accession&nbsp;&nbsp; Isolate<br> &nbsp;&nbsp; &nbsp;--------------------&nbsp;&nbsp; &nbsp;--------------<br> &nbsp;&nbsp; &nbsp;ERZ494218&nbsp;&nbsp; &nbsp;AM_0118<br> &nbsp;&nbsp; &nbsp;ERZ494219&nbsp;&nbsp; &nbsp;AM_0219<br> &nbsp;&nbsp; &nbsp;ERZ494220&nbsp;&nbsp; &nbsp;AM_0226<br> &nbsp;&nbsp; &nbsp;ERZ494221&nbsp;&nbsp; &nbsp;AM_0228<br> &nbsp;&nbsp; &nbsp;ERZ494222&nbsp;&nbsp; &nbsp;AM_0233<br> &nbsp;&nbsp; &nbsp;ERZ494223&nbsp;&nbsp; &nbsp;AM_0240<br> &nbsp;&nbsp; &nbsp;ERZ494224&nbsp;&nbsp; &nbsp;AM_0244<br> &nbsp;&nbsp; &nbsp;ERZ494225&nbsp;&nbsp; &nbsp;AM_0256<br> &nbsp;&nbsp; &nbsp;ERZ494226&nbsp;&nbsp; &nbsp;AM_0268<br> &nbsp;&nbsp; &nbsp;ERZ494227&nbsp;&nbsp; &nbsp;AM_0275<br> &nbsp;&nbsp; &nbsp;ERZ494228&nbsp;&nbsp; &nbsp;AM_0466<br> &nbsp;&nbsp; &nbsp;ERZ494229&nbsp;&nbsp; &nbsp;AM_0507<br> &nbsp;&nbsp; &nbsp;ERZ494230&nbsp;&nbsp; &nbsp;AM_0510<br> &nbsp;&nbsp; &nbsp;ERZ494231&nbsp;&nbsp; &nbsp;AM_0519<br> &nbsp;&nbsp; &nbsp;ERZ494232&nbsp;&nbsp; &nbsp;AM_0528<br> &nbsp;&nbsp; &nbsp;ERZ494233&nbsp;&nbsp; &nbsp;AM_0546<br> &nbsp;&nbsp; &nbsp;ERZ494234&nbsp;&nbsp; &nbsp;AM_0608<br> &nbsp;&nbsp; &nbsp;ERZ494235&nbsp;&nbsp; &nbsp;AM_0615<br> &nbsp;&nbsp; &nbsp;ERZ494236&nbsp;&nbsp; &nbsp;AM_0616<br> &nbsp;&nbsp; &nbsp;ERZ494237&nbsp;&nbsp; &nbsp;AM_0619<br> &nbsp;&nbsp; &nbsp;ERZ494238&nbsp;&nbsp; &nbsp;AM_0621<br> &nbsp;&nbsp; &nbsp;ERZ494239&nbsp;&nbsp; &nbsp;AM_0630<br> &nbsp;&nbsp; &nbsp;ERZ494240&nbsp;&nbsp; &nbsp;AM_0643<br> &nbsp;&nbsp; &nbsp;ERZ494241&nbsp;&nbsp; &nbsp;AM_0729<br> &nbsp;&nbsp; &nbsp;ERZ494242&nbsp;&nbsp; &nbsp;AM_0764<br> &nbsp;&nbsp; &nbsp;ERZ494243&nbsp;&nbsp; &nbsp;AM_0832<br> &nbsp;&nbsp; &nbsp;ERZ494244&nbsp;&nbsp; &nbsp;AM_0849<br> &nbsp;&nbsp; &nbsp;ERZ494245&nbsp;&nbsp; &nbsp;AM_0854<br> &nbsp;&nbsp; &nbsp;ERZ494246&nbsp;&nbsp; &nbsp;AM_0876<br> &nbsp;&nbsp; &nbsp;ERZ494247&nbsp;&nbsp; &nbsp;AM_0902<br> &nbsp;&nbsp; &nbsp;ERZ494248&nbsp;&nbsp; &nbsp;AM_0936<br> &nbsp;&nbsp; &nbsp;ERZ494249&nbsp;&nbsp; &nbsp;AM_1003<br> &nbsp;&nbsp; &nbsp;ERZ494250&nbsp;&nbsp; &nbsp;AM_1104<br> &nbsp;&nbsp; &nbsp;ERZ494251&nbsp;&nbsp; &nbsp;AM_1111<br> &nbsp;&nbsp; &nbsp;ERZ494252&nbsp;&nbsp; &nbsp;AM_1205<br> &nbsp;&nbsp; &nbsp;ERZ494253&nbsp;&nbsp; &nbsp;AM_1312<br> &nbsp;&nbsp; &nbsp;ERZ494254&nbsp;&nbsp; &nbsp;AM_1409<br> &nbsp;&nbsp; &nbsp;ERZ494255&nbsp;&nbsp; &nbsp;AM_1503<br> &nbsp;&nbsp; &nbsp;ERZ494256&nbsp;&nbsp; &nbsp;AM_1603<br> &nbsp;&nbsp; &nbsp;ERZ494257&nbsp;&nbsp; &nbsp;AM_1606<br> &nbsp;&nbsp; &nbsp;ERZ494258&nbsp;&nbsp; &nbsp;AM_1801<br> &nbsp;&nbsp; &nbsp;ERZ494259&nbsp;&nbsp; &nbsp;AM_1811<br> &nbsp;&nbsp; &nbsp;ERZ494260&nbsp;&nbsp; &nbsp;AM_2104<br> &nbsp;&nbsp; &nbsp;ERZ494261&nbsp;&nbsp; &nbsp;AM_2116<br> &nbsp;&nbsp; &nbsp;ERZ494262&nbsp;&nbsp; &nbsp;AM_2124<br> &nbsp;&nbsp; &nbsp;ERZ494263&nbsp;&nbsp; &nbsp;AM_2202<br> &nbsp;&nbsp; &nbsp;ERZ494264&nbsp;&nbsp; &nbsp;AM_2207<br> &nbsp;&nbsp; &nbsp;ERZ494265&nbsp;&nbsp; &nbsp;AM_2208<br> &nbsp;&nbsp; &nbsp;ERZ494266&nbsp;&nbsp; &nbsp;AM_2324<br> &nbsp;&nbsp; &nbsp;ERZ494267&nbsp;&nbsp; &nbsp;AM_2502<br> &nbsp;&nbsp; &nbsp;ERZ494268&nbsp;&nbsp; &nbsp;AM_2804<br> &nbsp; &nbsp;</p> <p>3. ACKNOWLEDGEMENTS<br> &nbsp; &nbsp;</p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of S&atilde;o Carlos, S&atilde;o Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Cient&iacute;fico e Tecnol&oacute;gico (CNPq), as well as, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; the spanish funding organ Consejo Superior de Investigaciones Cient&iacute;ficas (CSIC).</p> <p>This study was financed in part by the Coordena&ccedil;&atilde;o de Aperfei&ccedil;oamento de Pessoal de N&iacute;vel Superior - Brasil (CAPES) - Finance Code 001.</p> <p>&nbsp;</p> <p>4. CONTACT INFORMATION</p> <p>&nbsp;&nbsp; Current curators:</p> <p>&nbsp;&nbsp; - C&eacute;lio Dias Santos J&uacute;nior (celio.diasjunior@gmail.com)<br> &nbsp;&nbsp; - Flavio Henrique-Silva (dfhs@ufscar.br)<br> &nbsp;&nbsp; - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> &nbsp;</p> <p>5. COPYRIGHT NOTICE</p> <p>&nbsp;&nbsp; Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> &nbsp;&nbsp; Copyright (C) 2018 The AMnrGC consortium.</p> <p>&nbsp;&nbsp; This database is provided &ldquo;as is&rdquo; and without any warranty of any kind,<br> &nbsp;&nbsp; of openly available. You can redistribute and/or modify it<br> &nbsp;&nbsp; as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p>&nbsp;&nbsp; &nbsp;https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

GenoNet scores for human genome assembly GRCh37

<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files.&nbsp;</p> <p>Each row represents a genomic region with 131 columns. Please find the header line in &quot;genonet.header.txt&quot;.&nbsp;</p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 &nbsp; &nbsp;10000 &nbsp; &nbsp;10025 &nbsp; &nbsp;chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&amp;usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37&nbsp;<a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover&nbsp;<a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

De novo genome assembly of the meadow brown butterfly, Maniola jurtina

<p>1. Whole-genome GFF file (raw and filtered for min. gene length) [<em>Maniola.jurtina.gff3</em>, <em>Maniola_jurtina_filtered.gff3</em>]</p> <p>2. List of <em>M. jurtina</em> proteins [<em>Mjurtina_proteins.fa</em>].</p> <p>3. Results of spot pattern genes BLAST&nbsp;against <em>M. jurtina</em> proteome [<em>Lepidoptera_MJ_protein_matches.xlsx</em>].&nbsp;</p> <p>4. Annotations [blast2go_export.txt]</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

ST131_4071_genome_assembly_annotation_files

<p>The genome assembly annotation files of 4,071 E. coli ST131 genomes (see Decano &amp; Downing 2019).</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Sardinops sagax genome assemblies and annotations

<p>Files included are the genome assemblies for each haplotype (hap 1 and hap 2) of Sardinops sagax and the corresponding annotation files for each haplotype.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

LyBar v2.0 Genome Assembly and Annotation for Lycium barbarum

<p>LyBar v2.0 genome assembly and annotation files for <em>Lycium barbarum</em>.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

The Metagenome-Assembled Genome Inventory for Children (MAGIC)

<div> <div>Existing microbiota databases are biased towards adult samples, hampering accurate profiling of the infant gut microbiome. Here, we generated a **M**etagenome-**A**ssembled **G**enome **I**nventory for **C**hildren (**MAGIC**) from a large collection of bulk and viral-like particle-enriched metagenomes from 0-7 years of age, encompassing `3,299` prokaryotic and `139,624` viral species-level genomes, `8.5%` and `63.9%` of which are unique to MAGIC. MAGIC improves early-life microbiome profiling, with the greatest improvement in read mapping observed in Africans. We then identified `54` candidate keystone species, including several *Bifidobacterium spp.* and four phages, forming guilds that fluctuated in abundance with time. Their abundances were reduced in preterm infants and were associated with childhood allergies. By analyzing the *B. longum* pangenome, we found evidence of phage-mediated evolution and quorum sensing-related ecological adaptation. Together, the MAGIC database recovers genomes that enable characterization of dynamics of early-life microbiomes, identification of candidate keystone species, and strain-level study of target species.</div> </div>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record