Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
660
datasets available to search
ShareScore release 0.7.1
Dataset results
660 results for “genome assembly”
Metagenome assembled genome database of a human cohort and fecal reactors
<p><strong>HumanCohort_annotations.tsv.zip:</strong> This is the custom MAG database (n=2447 MAGs) and corresponding annotations that were used in Borton 2022: "Targeted curation of the gut microbial gene content modulating human cardiovascular disease". The citation will be updated upon publication of the manuscript. Metagenome assembled genomes were generated from fecal metagenomes derived from a 54 person cohort and anoxic methylated amine enrichments. </p> <p><strong>HumanCohortmetabolism_summary.xlsx.zip: </strong> This is the annotation summary for 2447 MAGs in the cohort database. </p> <p><strong>Quality_Abundance_CohortMAGs.xlsx: </strong>This is a genome inventory of the 2447 MAGs in the cohort database including genome statistics and relative abundance. </p> <p><strong>orig_1D_NMR_fids.zip: </strong>NMR data derived from anoxic methylated amine enrichments. </p>
Baseline assemblies for "ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads" protocol
<p>ntLink is a flexible <em>de novo</em> genome scaffolding toolkit which can be run in various modes depending on the desired user output, with multiple new functionalities recently introduced. Here, we provide the baseline assembly datasets used in the ntLink protocol paper "ntLink: a toolkit for <em>de novo </em>genome assembly scaffolding and mapping using long reads". The provided assemblies are ABySS (short-read) and Flye (long-read) assemblies of <em>Caenorhabditis elegans </em>genome sequencing data. The ABySS (v2.1.4) assembly utilized paired-end short reads (accession DRR008444), and was run with the following parameters: k=64 l=40 s=1000 q=15 B=10G j=8 kc=3 H=4 S=1000-10000 N=9.The <em>C. elegans</em> Flye (v2.5) assembly was run using Oxford Nanopore long reads (accession SRR10028109) and the following parameters: --nano-raw SRR10028109.fastq -g100m -t48.</p>
Chromosome-scale genome assembly and insights into the metabolome and gene regulation of leaf color transition in an important oak species, Quercus dentata
<p><em>Quercus dentata</em> Thunb., a dominant forest tree species in northern China, has significant ecological and ornamental value due to its adaptability and beautiful autumn coloration, with color changes from green to yellow into red resulting from the autumnal shifts in leaf pigmentation. However, the key genes and molecular regulatory mechanisms for leaf color transition remain to be investigated. First, we presented a high-quality chromosome-scale assembly for <em>Q. dentata</em>. This 893.54 Mb sized genome (contig N50=4.21 Mb, scaffold N50=75.55 Mb; 2n=24) harbors 31,584 protein-coding genes. Second, our metabolome analyses uncovered pelargonidin-3-O-glucoside, cyanidin-3-O-arabinoside, and cyanidin-3-O-glucoside as the main pigments involved in leaf color transition. Third, gene co-expression further identified the MYB-bHLH-WD40 (MBW) transcription activation complex as central to anthocyanin biosynthesis regulation. Notably, transcription factor (TF) <em>QdNAC </em>(<em>QD08G038820</em>) was highly co-expressed with this MBW complex and may regulate anthocyanin accumulation and chlorophyll degradation during leaf senescence through direct interaction with another TF, <em>QdMYB </em>(<em>QD01G020890</em>), as revealed by our further protein-protein and DNA-protein interaction assays. Our high-quality genome assembly, metabolome and transcriptome resources further enrich <em>Quercus </em>genomics, and will facilitate upcoming exploration of ornamental values and environmental adaptability in this important genus.</p>
Metagenome-Assembled Genome DRAM Annotations (EMERGE 97% dereplicated MAGs)
<p>This is the combined DRAM annotation outputs for the 1,864 97% dereplicated metagenome-assembled genomes from Stordalen Mire, Sweden. </p> <ul> <li>1864_97percentmags_annotations_combined.tsv.gz</li> <li>1864_97percentmags_metabolism_summary.xlsx</li> <li>product_0.html</li> <li>product_1.html</li> </ul> <p>METHODS:</p> <p>MAGs were annotated and distilled using DRAM (v1.4.0).</p> <p>FUNDING:<br> This research is a contribution of the EMERGE Biology Integration Institute ((https://emerge-bii.github.io/), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.<br> We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.<br> This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.<br> A portion of this research was performed under the Facilities Integrating Collaborations for User Science (FICUS) program (proposal: 10.46936/fics.proj.2017.49950/60006215 and 10.46936/10.25585/60001148) and used resources at the DOE Joint Genome Institute (<a href="https://www.google.com/url?q=https://ror.org/04xm1d337&sa=D&source=docs&ust=1674859614742521&usg=AOvVaw2XgXYw9eI4JIXRMKn3S9Se">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://www.google.com/url?q=https://ror.org/04rc0xn13&sa=D&source=docs&ust=1674859614742655&usg=AOvVaw3UXdoHIFmVjc-mXUhDXYQt">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</p>
Virtual NGS data used in the paper "Optimal k-mer values for good-quality triploid genome assemblies" and its assembly results
<p>Despite technological advancements, whole-genome sequencing remains technically challenging for organisms with higher-ploidy genomes. Therefore, the use of short-read sequencing platforms for this purpose has been attempted, but the conditions that result in poor-quality genomes have not been elucidated. Therefore, in the present study, simulated sequences mimicking the accumulation of insertion/deletion mutations were created to clarify the permissible differences between homologous chromosomes and the k-mer sizes for good-quality genome assembly from short-read sequencing data for triploid species. The results illustrated that a narrow range of k-mers permits the generation of high-quality assemblies for any level of difference between homologous chromosomes. <br> This dataset consists of the virtual haploid genome (O.fasta), triploid genome data (reference_sequense), NGS read data (NGS_reads), and assembly results (all_contigs) created in this study.<br> </p>
Supplemental Files for "A highly contiguous genome assembly reveals sources of genomic novelty in the symbiotic fungus Rhizophagus irregularis"
<p>Supplemental files for "A highly contiguous genome assembly reveals sources of genomic novelty in the symbiotic fungus Rhizophagus irregularis". This data is linked to the bioRxiv pre-print doi: https://doi.org/10.1101/2022.10.19.511543, an updated version of which is in press at G3: Genes|Genomes|Genetics, and corresponds to the NCBI BioProject PRJNA885267 and NCBI BioSample SAMN31081226.</p> <p> </p> <p><strong>Nuclear genome assembly</strong></p> <p>Rhizophagus_irregularis_DAOM197198_assembly.fasta</p> <p> </p> <p><strong>Illumina and Illumina+Nanopore gene annotations</strong></p> <p>Rhizophagus_irregularis_DAOM197198_Illumina+ONT_curated.gff3</p> <p>Rhizophagus_irregularis_DAOM197198_Illumina_curated.gff3</p> <p> </p> <p><strong>Illumina and Illumina+Nanopore functional gene annotations</strong></p> <p>Rhizophagus_irregularis_DAOM197198_annotations_Illumina+ONT.txt</p> <p>Rhizophagus_irregularis_DAOM197198_annotations_Illumina.txt</p> <p> </p> <p><strong>Illumina and Illumina+Nanopore CDS sequences</strong></p> <p><span>Rhizophagus_irregularis_DAOM197198_cds-transcripts_Illumina+ONT_curated.fa</span></p> <p>Rhizophagus_irregularis_DAOM197198_cds-transcripts_Illumina_curated.fa</p> <p> </p> <p><strong>Illumina and Illumina+Nanopore mRNA sequences</strong></p> <p>Rhizophagus_irregularis_DAOM197198_mrna-transcripts_Illumina+ONT_curated.fa</p> <p>Rhizophagus_irregularis_DAOM197198_mrna-transcripts_Illumina_curated.fa</p> <p> </p> <p><strong>Illumina and Illumina+Nanopore protein sequences</strong></p> <p><span>Rhizophagus_irregularis_DAOM197198_proteins_Illumina+ONT_curated.fa</span></p> <p>Rhizophagus_irregularis_DAOM197198_proteins_Illumina_curated.fa</p> <p> </p> <p><strong>GO terms for g:Profiler</strong><br> Rhizophagus_irregularis_DAOM197198_Illumina+ONT_GOterms.gmt<br> *Or use token gp__xfGY_dQeI_yx4</p> <p> </p> <p><strong>Repetitive and transposable element library and annotation</strong></p> <p>Rhizophagus_irregularis_DAOM197198_curatedrepeatlibrary.fasta</p> <p>Rhizophagus_irregularis_DAOM197198_repeatmasker.out</p> <p>Rhizophagus_irregularis_DAOM197198_repeats.gff3</p> <p> </p> <p><strong>DNA methylome (sequenced from spores)</strong></p> <p>Rhizophagus_irregularis_DAOM197198_mCG_mods_frequency.tsv</p> <p> </p> <p><strong>Poly(A) signal and tail sequences</strong></p> <p>Rhizophagus_irregularis_DAOM197198_pasa_polyAsite_analysis.out</p> <p>Rhizophagus_irregularis_DAOM197198_pasa_polyAsites.fasta</p> <p> </p> <p><strong>Small RNA annotation</strong></p> <p>Rhizophagus_irregularis_DAOM197198_sRNA.gff3</p> <p>Rhizophagus_irregularis_DAOM197198_sRNA.tsv</p> <p> </p> <p><strong>Mitochondrial genome assembly and annotation</strong></p> <p>Rhizophagus_irregularis_DAOM197198_mtDNA.fasta</p> <p>Rhizophagus_irregularis_DAOM197198_mtDNA.gff</p> <p> </p> <p><strong><em>R. irregularis</em> phylostratigraphy</strong></p> <p>Rhizophagus_irregularis_DAOM197198_1432141_phyloranks.tsv</p> <p>Rhizophagus_irregularis_DAOM197198_1432141_high-confidence_phyloranks.tsv</p> <p> </p> <p><strong>Mucoromycota fungi phylostratigraphy</strong></p> <p>Disdec1_101101_phyloranks.tsv</p> <p>Geopyr1_50956_phyloranks.tsv</p> <p>Gigmar1_4874_phyloranks.tsv</p> <p>Morel2_1314771_phyloranks.tsv</p> <p>Phybl2_4837_phyloranks.tsv</p> <p>Radspe1_64574_phyloranks.tsv</p> <p> </p> <p><strong>Fatty acid synthase phylogeny</strong></p> <p>FAS_genes_muscle5_msa.fa (alignments)</p> <p>FAS_genes.raxml.support (ML tree)</p>
De novo assembly of a long-read Amblyomma americanum genome
<p>Genome assemblies of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the unphased diploid assembly generated by the Flye assembler (Arcadia_Amblyomma_americanum_asm001.fasta). In addition, there are two associated fasta files containing sequences generated by submitting the unphased diploid assembly to separation by the Purge_Dups pipeline (purged pseudo-haploid assembly and haplotig assembly).</p> <p>Flye assembler: https://github.com/fenderglass/Flye</p> <p>Purge_Dups pipeline: https://github.com/dfguan/purge_dups</p> <p>NCBI Bioproject: PRJNA932813</p>
De novo assembly of a long-read Amblyomma americanum genome (NCBI/Genbank deposited genome)
<p>Genome assembly of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the phased pseudo-haploid tick genome generated after assembly using Flye, phasing using Purge_Dups, and clean-up using custom python scripts generated in-house. </p> <p>NCBI Bioproject: PRJNA932813</p>
Additional annotation, alignment, and results from Ka/Ks analysis for Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus)
<p><strong>Annotation files, alignments, and results summaries from Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus).</strong></p> <p>Pairwise genome alignments contain the .maf suffix</p> <p>FASTA alignments from stitched gene blocks contain the .fasta suffix</p> <p>CSV file containing the Ka/Ks results</p> <p>RepeatMasker .out file</p>
Data for: Genome assemblies of the simultaneously hermaphroditic flatworms Macrostomum cliftonense and Macrostomum hystrix
<p>The free-living, simultaneously hermaphroditic flatworms of the genus <em>Macrostomum, </em>are increasingly used as model systems in various contexts. In particular, <em>M. lignano</em>, the only species of this group with a published genome assembly, has emerged as a model for the study of regeneration, reproduction, and stem-cell function. However, challenges have emerged due to <em>M. lignano</em> being a hidden polyploid, having recently undergone whole-genome duplication and chromosome fusion events. This complex genome architecture presents a significant roadblock to the application of many modern genetic tools. Hence, additional genomic resources for this genus are needed. Here we present such resources for <em>M. cliftonense</em> and <em>M. hystrix</em>, which represent<em> </em>the contrasting mating behaviors of reciprocal copulation and hypodermic insemination found in the genus. We use a combination of PacBio long-read sequencing and Illumina shot-gun sequencing, along with several RNA-Seq datasets, to assemble and annotate highly contiguous genomes for both species. The assemblies span ~227Mb and ~220Mb and are represented by 399 and 42 contigs for<em> M. cliftonense</em> and<em> M. hystrix,</em> respectively. Furthermore, high BUSCO completeness (~84-85%), low BUSCO duplication rates (8.3-6.2%), and low k-mer multiplicity indicate that these assemblies do not suffer from the same assembly ambiguities of the<em> M. lignano</em> genome assembly, that can be attributed to the complex karyology of this species. We also show that these resources, in combination with the prior resources from <em>M. lignano, </em>offer excellent foundations for comparative genomic research in this group of organisms.</p>
Supplementary data for "WebQUAST: online evaluation of genome assemblies"
<p>Supplementary data for A. Mikheenko, V. Saveliev, P. Hirsch, A. Gurevich. WebQUAST: online evaluation of genome assemblies.</p> <p><br> Reference genomes of two <em>Escherichia coli</em> K-12 substrains (MG1655 and W3110) and genome annotation of MG1655.</p> <p><em>De novo</em> assemblies of the <em>Escherichia coli</em> K-12 MG1655 short-read Illumina dataset (<a href="https://trace.ncbi.nlm.nih.gov/Traces/index.html?view=run_browser&acc=ERR008613&display=metadata">ERR008613</a>) with ABySS, MEGAHIT, SPAdes, and Velvet.</p> <p>WebQUAST reports in three evaluation modes: <br> * Use Case 1: reference-free evaluation (<em>sample_data_no_ref</em>)<br> * Use Case 2: reference-based evaluation (<em>sample_data_true_ref</em>) <br> * Use Case 3: evaluation based on a close reference (<em>sample_data_close_ref</em>)</p> <p>The snapshot (<a href="https://github.com/ablab/quast/commit/2bd50600bcf63ee826965c10f6ca4cdc2aa046e5">commit 2bd5060</a>) of the QUAST command-line version used by WebQUAST for generating the reports.</p>
Assemblies for "Linear time complexity de novo long read genome assembly with GoldRush"
<p>GoldRush is a <em>de novo</em> genome assembly algorithm with linear time complexity in the number of input long sequencing reads. We tested GoldRush on Oxford Nanopore Technologies datasets with different base error profiles describing the genomes of three human cell lines (NA24385, HG01243 and HG02055), Oryza sativa (rice), and Solanum lycopersicum (tomato). Here, we provide the assemblies for the GoldRush, Flye, Redbean and Shasta assemblies of these long read datasets.</p>
Novel metagenome assembled genomes (MAGs) that best represent novel species level taxa within the phylum Chloroflexota
<p>1280 Chloroflexita MAGs from the study "Taxonomic re-classification and expansion of the phylum Chloroflexota based on over 5000 genomes and metagenome-assembled genomes". Only MAGs that improved the representation of a species-level genome cluster within the phylum <em>Chloroflexota</em> were included in this deposition.</p> <p>Most of these MAGs were assembled from publicly availabe metagenome sequence data obtained from the NCBI sra database.</p> <p>An overview of the here deposited MAGs can be found in <a href="https://zenodo.org/api/files/2a6a7fa1-489c-426d-8e05-ada23038dfdf/Zenodo_deposited_MAGS_overview.xlsx?versionId=9cde1488-8388-4fdb-a45e-2d48dc066f9a"> Zenodo_deposited_MAGS_overview.xlsx</a>, for more details please refer to the abive mentioned publication.</p> <p>MAG assemblies are deposited as gzip compressed tar.archive. Three tar.gz archives have been deposited, containing the same MAG assemblies but sorted by different criteria:</p> <ol> <li>All MAGs sorted by category of the source environment</li> <li>All MAGs sorted by class designation</li> <li>All MAGs sorted by MIMAG quality (high or moderate)</li> </ol>
Picea mariana isolate 40-10-1 mitochondrial genome assembly
<p><em>Picea mariana</em> isolate 40-10-1 mitochondrial genome assembly generated with Illumina HiSeq and 10x Genomics Chromium reads using ABySS v2.1.0, Tigmint v1.1.2, and ARCS v1.0.6.</p>
Picea glauca isolate WS77111 mitochondrial genome assembly
<p><em>Picea glauca</em> isolate WS77111 mitochondrial genome assembly generated with Illumina HiSeq reads using ABySS v2.1.4.</p>
Picea engelmannii isolate Se404-851 mitochondrial genome assembly
<p><em>Picea engelmannii</em> isolate Se404-851 mitochondrial genome assembly generated with Illumina HiSeq reads using ABySS v2.1.4.</p>
The genome assemble of Bemisia tabaci MED
<p><span><strong>Aim</strong>: </span><span>The sweet potato whitefly, <em>Bemisia tabaci </em>MED is a globally invasive species that causes serious economic damage to agroecosystems. Despite the significant threat it poses to agricultural and economic crops worldwide, the global perspective of the invasion patterns and genetic mechanism contributing to the success of this notorious pest is still poorly understood. The objective of this research was to enhance genome and population genetic analyses to better understand the intricate invasion patterns of <em>B. tabaci</em> MED. </span></p> <p><span><strong>Location</strong>: </span><span>Samples were collected in native (Spain, Croatia, Bosnia and Herzegovina, Cyprus, and Israel) and invaded regions (China, South Korea and North America).</span></p> <p><span><strong>Methods</strong>: </span><span>We first assembled a chromosome-scale reference genome of <em>B. tabaci</em> MED, and then employed the restriction site‐associated 2b‐RAD method to genotype over 20, 000 high‐quality single nucleotide polymorphisms from 29 geographical populations.</span></p> <p><span><strong>Results</strong>: </span><span>A reference genome of <em>B. tabaci </em>MED, with a size of 637.47 Mb, was available. The majority of the assembled sequences (99%) were anchored onto ten linkage groups, with an N50 size of 58.76 Mb, representing a significant improvement over previous whitefly genome assemblies. We identified rapidly expanded gene families and positively selected genes, probably contributing to successful invasion and rapid adaptation to the new environment. Population genomics analysis showed that </span><span>three </span><span>highly differentiated genetic groups</span><span> were formed</span><span>, and </span><span>c</span><span>omplex and e</span><span>xtensive gene flow occurred across the Mediterranean populations</span><span>. The genetic admixture patterns in East Asia populations were distinct from those in North America, indicating that they had different source populations.</span> </p> <p><span><strong>Conclusions</strong>: </span><span>The high-quality, chromosome-scale genome of <em>B. tabaci</em> MED </span><span>offered opportunities for more comprehensive genome-wide studies, and provided a solid foundation for the complex introduction events and the differential invasiveness of <em>B. tabaci</em> MED worldwide.</span></p>
Supplementary data for: Chromosome-level genome assembly and circadian gene repertoire of the Patagonia blennie Eleginops maclovinus
<p>This dataset contains the genome assembly and associated annotation of the Patagonian Blennie (<em>Eleginops maclovinus</em>), the closest extant taxon to the Antarctic notothenioid radiation. In addition to the characterization of the <em>E. maclovinus </em>genome, the dataset includes a description of circadian rhythm orthologs for <em>E. maclovinus</em>, other notothenenioid taxa, and teleost outgroups, as well as a copy of the bioinformatic scripts used for the assembly, annotation, and other downstream analysis.</p>
Metagenome-assembled genomes(MAGs) generated by MetaCC binning
<p>MAGs generated by MetaCC binning from the human gut short-read, the wastewater (WW) short-read, the cow rumen long-read, and the sheep gut long-read metaHi-C datasets</p>
The LakePulse Metagenome-Assembled Genome catalogue
<p>Lakes are heterogenous ecosystems inhabited by a rich microbiome whose genomic diversity is poorly defined. We present a continental-scale study of metagenomes representing 6.5-million km<sup>2</sup> of the most lake-rich landscape on Earth. Analysis of 308 Canadian lakes resulted in a metagenome-assembled genome (MAG) catalogue of 1,008 mostly novel bacterial genomospecies. Lake trophic state was a leading driver of taxonomic and functional diversity among MAG assemblages, reflecting the responses of communities profiled by 16S rRNA amplicons and gene-centric metagenomics. Coupling the MAG catalogue with watershed geomatics revealed terrestrial influences of soils and land use on assemblages. Agriculture and human population density were drivers of turnover, indicating detectable anthropogenic imprints on lake bacteria at the continental scale. The sensitivity of bacterial assemblages to human impact reinforces lakes as sentinels of environmental change. Overall, the LakePulse MAG catalogue greatly expands the freshwater genomic landscape, advancing an integrative view of diversity across Earth's microbiomes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.