Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

915

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

915 results for “metagenomics”

Learn how ShareScore rates datasets ↗
zenodo44/100

iMGMC - integrated Mouse Gut Metagenomic Catalog

<p><em>Creation of an new mouse gut gene catalog with special features:</em></p> <ul> <li>more diverse samples from different studies (12 Vendors incl. wild mice and various gut locations)</li> <li>clustering-free approach: all-in-one assembly, keeping track of each ORF to contigs to bins</li> <li>higher taxonomic resolution and more accuracy by using contigs for annotation</li> <li>16S rRNA gene integration via linkage to bins</li> <li>expansion by 20,927 MAGs from sample-wise assembly of 871 mouse gut metagenomic samples, representing 1,296 species</li> </ul> <p>Code used:&nbsp;<a href="https://github.com/tillrobin/iMGMC">https://github.com/tillrobin/iMGMC</a></p> <p>The vast complexity of host-associated microbial ecosystems requires host-specific reference catalogs to survey the functions and diversity of these communities. We generated a comprehensive resource, the integrated mouse gut metagenome catalog (iMGMC), comprising 4.6 million unique genes and 660 metagenome-assembled genomes (MAGs) with many of them (485 MAGs, 73%) linked to reconstructed full-length 16S rRNA gene sequences. iMGMC enables unprecedented coverage and taxonomic resolution of the mouse gut microbiota, i.e. more than 92% of MAGs lack species-level representatives in public repositories (&lt;95% ANI match). The integration of MAGs and 16S rRNA gene data allows a more accurate prediction of functional profiles of communities than based on 16S rRNA amplicons alone. Accompanying iMGMC we provide a set of MAGs representing 1,296 gut bacteria obtained through complementary assembly strategies. We envision that integrated resources such as iMGMC together with MAG collections will enhance the resolution of numerous existing and future sequencing-based studies.</p> <p>Genecatalog:</p> <p>Description&nbsp;&nbsp; &nbsp;Size&nbsp;&nbsp; Filename<br> Catalog ORF sequences&nbsp;&nbsp; &nbsp;1 GB&nbsp;&nbsp; &nbsp;iMGMC-GeneID.fasta.gz<br> Full assembly contigs&nbsp;&nbsp; &nbsp;1.3 GB&nbsp;&nbsp; &nbsp;iMGMC-ConitgID.fasta.gz<br> Mapping File (GeneID-&gt;ContigID-&gt;BinID)&nbsp;&nbsp; &nbsp;30 MB&nbsp;&nbsp; &nbsp;iMGMC-map-Gene-Contig-Bin.tab.gz<br> Taxonomic annotations&nbsp;&nbsp; &nbsp;40 MB&nbsp;&nbsp; &nbsp;iMGMC_map_taxonomy.tar.gz<br> Functional annotations&nbsp;&nbsp; &nbsp;36 MB&nbsp;&nbsp; &nbsp;iMGMC_map_functionality.tar.gz<br> 16S rRNA sequences&nbsp;&nbsp; &nbsp;2 MB&nbsp;&nbsp; &nbsp;iMGMC-16SrRNAgenes.fasta</p> <p>Metagenome-assembled genomes (MAGs) :</p> <p>Description&nbsp;&nbsp; &nbsp;Size&nbsp;&nbsp; &nbsp;Filename<br> integrated MAGs&nbsp;&nbsp; &nbsp;0.5 GB&nbsp;&nbsp; &nbsp;iMGMC_MAGs.tar.gz<br> representave mMAGs (n=1296)&nbsp;&nbsp; &nbsp;1 GB&nbsp;&nbsp; &nbsp;iMGMC-mMAGs-dereplicated_genomes.tar.gz<br> representave hqMAGs (n=830)&nbsp;&nbsp; &nbsp;0.7 GB&nbsp;&nbsp; &nbsp;iMGMC-hqMAGs-dereplicated_genomes.tar.gz<br> all mMAGs (n=20,927)&nbsp;&nbsp; &nbsp;15 GB&nbsp;&nbsp; &nbsp;iMGMC-mMAGs.tar.gz<br> Annotations by CheckM, dRep-Clustering, GTDB-Tk&nbsp;&nbsp; &nbsp;2 MB&nbsp;&nbsp; &nbsp;MAG-annotation_CheckM_dRep_GTDB-Tk.tar.gz<br> Functional annotations (hqMAGs by eggNOG mapper v2)&nbsp;&nbsp; &nbsp;187 MB&nbsp;&nbsp; &nbsp;hqMAGs.emapper.annotations.gz</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

An evaluation of the accuracy and speed of metagenome analysis tools

<p>Metagenome studies are becoming increasingly widespread, yielding important insights into microbial communities covering diverse environments from terrestrial and aquatic ecosystems to human skin and gut. With the advent of high-throughput sequencing platforms, the use of large scale shotgun sequencing approaches is now commonplace. However, a thorough independent benchmark comparing state-of-the-art metagenome analysis tools is lacking. Here, we present a benchmark where the most widely used tools are tested on complex, realistic data sets. Our results clearly show that the most widely used tools are not necessarily the most accurate, that the most accurate tool is not necessarily the most time consuming and that there is a high degree of variability between available tools. These findings are important as the conclusions of any metagenomics study are affected by errors in the predicted community composition and functional capacity.</p>

opencc-by-4.0Jan 2016View details →
zenodo44/100

Viral Metagenomes from water systems in the Mediterranean Sea

<p>The Water Framework Directive (WFD; 2000/60/EC) is the European umbrella for the assessment and regulation of ecological quality of water systems (lakes, rivers, transitional waters, coastal waters). The scope of WFD is to improve the ecological status of the aquatic ecosystems. To do so, a long time-series of monitoring campaigns within WFD exists with the ultimate goal to protect coastal ecosystems from degradation. Further, WFD project employs the calculation and improvement of ecological quality indices for the definition and assessment of eutrophication. In the submitted sub-project within WFD, the viral metagenome of 15 samples collected in 2014 and 2015 is sequenced and analyzed for several genes for the study of taxonomy and potential function of double-stranded DNA viruses.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

Metagenomics of a pustular microbial mat from Shark Bay, Australia: Raw sequences and assembled MAGs

<p>This data accompanies the paper, &quot;<a href="https://www.nature.com/articles/s43705-022-00128-1">Metagenomic,&nbsp;(bio)chemical, and microscopic analyses reveal the potential for the cycling of sulfated EPS in Shark Bay pustular mats</a>&quot;&nbsp;which looks at the cycling of sulfated polysaccharides in peritidal pustular mats from Shark Bay, Australia. The microbial community was sequenced, assembled, and binned. The&nbsp;raw sequencing reads used in this analysis are the following:</p> <ul> <li>SB_forward_paired_copy.fastq.gz&nbsp;</li> <li>SB_reverse_paired_copy.fastq.gz&nbsp;</li> </ul> <p>The resulting metagenome-assembled genomes (MAGs) are presented in the following folder:</p> <ul> <li>MAGs.zip</li> </ul> <p>&nbsp;</p> <ul> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Noccaea root and soil metagenomes

<p>Datasets comprise metagenomic data of hyperaccumulating <em>Noccaea praecox</em> and <em>N. caerulescens</em> from bulk soil, rhizosphere, and root compartments.</p>

opencc-by-4.0Apr 2025View details →
zenodo44/100

Aqueous geochemical measurements and speciation calculations with concurrent copper resistance gene counts from sediment metagenomes over a seasonal cycle from 2015 to 2016 on Silver Bow Creek and Blacktail Creek near Butte, MT

<p>This dataset contains information from concurrently gathered geochemical and metagenomic samples collected from Silver Bow Creek and Blacktail Creek near Butte, MT (SBC/BC) during 2015 and 2016. SBC/BC is recovering from metal contamination related to extensive mining in the area. Full geochemical measurements, geochemical speciation calculations, and gene counts of sequences mapping to copper resistance genes using MG-RAST are included.&nbsp;&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)

<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 &amp; 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by&nbsp;<a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads.&nbsp;</p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a>&nbsp;(v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with&nbsp;<a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a>&nbsp;(v2.9-b1768) using&nbsp;<a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using&nbsp;<a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1)&nbsp;using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18)&nbsp;and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0)&nbsp;and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2)&nbsp;and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p>&nbsp;</p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p>&nbsp;</p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Catalog of stool metagenome-assembled genomes from patients with different cancer types

<p><strong>A non-redundant catalog of 3,816 genomes with at least 75% completeness and no more than 15% contamination assembled from metagenomes. Samples of 976 metagenomes were obtained from patients receiving immunotherapy for the treatment of different types of cancers.</strong></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Metagenomes: gene function and family annotations

<p>Functional annotations of genes for all contigs in 1,782 metagenomes.</p> <p>Genes were annotated to three sources: (1) COGs, (2) Pfams, and (3) <em>de novo</em> families from reference sequences. These gene annotations are used to train and run PlasX.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

KMA Mapping and alignment statistics : livestock fecal metagenomes against ResFinder and genomes

<p>Three zip archives are included used in the analysis of the European livestock resistome.</p> <p>Two of them contain &#39;mapstat&#39; files produced by the KMA software using the &#39;extended features&#39; flag.<br> Each mapstat file thus summarize the mapping and alignment statistics when using KMA on a metagenome against a database.</p> <p>The last archive contains the &#39;refdata&#39; file used to annotate the genomic mapstat hits. It encodes the taxonomic affilication of sequences hit by one or more samples.<br> &nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Supplementary data for publication Global distribution of mcr gene variants in 214K metagenomic samples

<p># Supplementary data for the manuscript &quot;Global distribution of mcr gene variants in 214,095 metagenomic samples&quot;</p> <p>SD1_mapped_runids.csv : tab-separated file with columns of run_accessions downloaded from ENA and whether the metagenome were positive for at least one of the mcr genes.</p> <p>SD2_mcr_df.csv : compositional table of mcr-positive metagenomes with associated metadata (collection_year, country, and host) for each run_accession, as well as mapping results.</p> <p>SD3_mcr_contigs.fa : FASTA file with contigs carrying mcr genes. The header contains the run_accession ID.</p> <p>SD4_aldex2_results.csv: CSV file containing ALDEx2 results. The columns are as follows:<br> * group: metadata category (year, country or host). If the column contains more than one label, e.g., &quot;Denmark - 2020 - Pigs&quot;, significance is tested within Danish pig samples from 2020.<br> * rab.all:&nbsp; median clr value for all samples in the feature<br> * rab.win.conditionA:&nbsp; median clr value for the condition A of samples<br> * rab.win.conditionB: median clr value for the condition B of samples<br> * diff.btw: median difference in clr values between A and B conditions<br> * diff.win: median of the largest difference in clr values within A and B conditions<br> * effect : median effect size: diff.btw / max(diff.win) for all instances<br> * overlap : proportion of effect size that overlaps 0 (i.e. no effect)<br> * we.ep: Expected P value of Welch&rsquo;s t test<br> * we.eBH: Expected Benjamini-Hochberg corrected P value of Welch&rsquo;s t test<br> * wi.ep: Expected P value of Wilcoxon rank test<br> * wi.eBH: Expected Benjamini-Hochberg corrected P value of Wilcoxon test<br> * parts: gene name<br> * conditionA: label of condition A that is compared against condition B<br> * conditionB: label of condition B that is compared against condition A<br> * conditions.A.vs.B: label to explain condition A compared against condition B<br> NOTE: see for more explanation of the output of ALDEx2 https://www.bioconductor.org/packages/release/bioc/vignettes/ALDEx2/inst/doc/ALDEx2_vignette.html#5_ALDEx2_outputs</p> <p>SD5: Multi-VCF file containing SNP information on mcr alleles. Can be used to construct consensus sequences.</p> <p>SD6: FASTA file containing all unique consensus sequences reported in the manuscript.</p> <p>SD7: CSV file with an overview of which metagenome contains which unique consensus sequence.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Supplementary material for "Towards a metagenomics machine learning interpretable model for understanding the transition from adenoma to colorectal cancer"

<p>Supplementary files for&nbsp;&quot;Towards a metagenomics machine learning interpretable model for understanding the transition from adenoma to colorectal cancer&quot;.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Plasmids Identified in Air Metagenomes

<p>&nbsp;Metagenomic data were selected in Web of Science (Clarivate) on October 2022 using keywords: txid655179[Organism:noexp] AND metagenome [Filter]; AIR Metagenome; Air microbiome; Troposphere; Aerosol; Atmosphere. Data were manually curated to remove sequencing originated from metabarcoding data (i.e., 16S). The assembled data supplied by MetaSUB consortium (Danko et al., 2021) when available was used for air metagenome in the built environments.&nbsp;</p> <div> <p>Plasmid contents were predicted using the assembled data. Metagenomes sequencing by Illumina (paired-illumina reads) were assembled by using megahit 1.2.9 with metalarge option (Li et al., 2015) after cleaning data with bbduk2 (qtrim=rl trimq=28 minlen=25 maq=20 ktrim=r k=25 mink=11 and a list of adaptators to remove) from bbtools suite (<a href="https://jgi.doe.gov/data-and-tools/software-tools/bbtools/" target="_blank" rel="noreferrer noopener">https://jgi.doe.gov/data-and-tools/software-tools/bbtools/</a>)&nbsp;&nbsp;</p> <div> <p><span><span>Plasmids were predicted for each assembling by using scripts describing in-depth in Hilpert et al. (Hilpert </span></span><span><span>et al.</span></span><span><span>, 2021; </span><span>Hennequin</span> </span><span><span>et al.</span></span><span><span>, 2022) and available in </span><span>github</span><span> website (</span></span><span><span><span>https://github.com/meb-team/PlasSuite/</span></span></span><span><span>). Briefly, contigs were analyzed using both reference-based and reference-free approaches.</span></span><span><span>&nbsp;The databases employed included those for chromosomes (archaea and bacteria) and plasmids from NCBI, as well as the MOB-suite tool (Robertson and Nash, 2018</span><span>) ,</span><span> SILVA (Quast </span></span><span><span>et al.</span></span><span><span>, 2013) and phylogenetic markers harbored by chromosomes (Wu </span></span><span><span>et al.</span></span><span><span>, 2013). Two reference-free methods were applied to contigs that were not affiliated with chromosomes (discarded) or plasmids (</span><span>retained</span><span> in the first step): </span><span>PlasFlow</span><span> (Krawczyk et al., 2018) and </span><span>PlasClass</span><span> (Pellow </span></span><span><span>et al.</span></span><span><span>, 2020). Viruses were removed by using </span><span>viralVerify</span><span> (</span></span><span><span><span>https://github.com/ablab/viralVerify</span></span></span><span><span>) (Antipov </span></span><span><span>et al.</span></span><span><span>, 2020) that provides in parallel provide plasmid/non-plasmid classification</span><span>. </span><span>&nbsp;</span><span>The database built for this purpose is available at this address </span></span><span><span><span>https://github.com/meb-team/PlasSuite/?tab=readme-ov-file#1-prepare-or-download-your-databases</span></span></span> <span><span>&nbsp;Eukaryotes contaminants were removed by aligning the sequences against NT databases and human chromosomes (GRCh38) with minimap2 with -x asm5 </span><span>option (Li, 2018)</span><span>. Contigs mapping with an identity of 95% and a coverage of 80% were removed.</span></span><span> &nbsp;the final plasmidome set was clustered by mmseqs (Mirdita, Steinegger and S&ouml;ding, 2019) with 80% of coverage and 90% of identity (--min-seq-id 0.90 -c 0.8 --cov-mode 1 --cluster-mode 2 --alignment-mode 3 --kmer-per-seq-scale 0.2).&nbsp;</span></p> </div> </div>

opencc-by-4.0May 2024View details →
zenodo44/100

Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes

<p>&nbsp;</p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p>&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; RELEASE MAG-2018/01<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --------------------------------------</p> <p>&nbsp;</p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p>&nbsp;</p> <p>2. LOCATION</p> <p>&nbsp;</p> <p>&nbsp;&nbsp; MAGs sequences are available under ENA project PRJEB25176.</p> <p>&nbsp;&nbsp;</p> <p>&nbsp;&nbsp;&nbsp; ENA_accession&nbsp;&nbsp; Isolate<br> &nbsp;&nbsp; &nbsp;--------------------&nbsp;&nbsp; &nbsp;--------------<br> &nbsp;&nbsp; &nbsp;ERZ494218&nbsp;&nbsp; &nbsp;AM_0118<br> &nbsp;&nbsp; &nbsp;ERZ494219&nbsp;&nbsp; &nbsp;AM_0219<br> &nbsp;&nbsp; &nbsp;ERZ494220&nbsp;&nbsp; &nbsp;AM_0226<br> &nbsp;&nbsp; &nbsp;ERZ494221&nbsp;&nbsp; &nbsp;AM_0228<br> &nbsp;&nbsp; &nbsp;ERZ494222&nbsp;&nbsp; &nbsp;AM_0233<br> &nbsp;&nbsp; &nbsp;ERZ494223&nbsp;&nbsp; &nbsp;AM_0240<br> &nbsp;&nbsp; &nbsp;ERZ494224&nbsp;&nbsp; &nbsp;AM_0244<br> &nbsp;&nbsp; &nbsp;ERZ494225&nbsp;&nbsp; &nbsp;AM_0256<br> &nbsp;&nbsp; &nbsp;ERZ494226&nbsp;&nbsp; &nbsp;AM_0268<br> &nbsp;&nbsp; &nbsp;ERZ494227&nbsp;&nbsp; &nbsp;AM_0275<br> &nbsp;&nbsp; &nbsp;ERZ494228&nbsp;&nbsp; &nbsp;AM_0466<br> &nbsp;&nbsp; &nbsp;ERZ494229&nbsp;&nbsp; &nbsp;AM_0507<br> &nbsp;&nbsp; &nbsp;ERZ494230&nbsp;&nbsp; &nbsp;AM_0510<br> &nbsp;&nbsp; &nbsp;ERZ494231&nbsp;&nbsp; &nbsp;AM_0519<br> &nbsp;&nbsp; &nbsp;ERZ494232&nbsp;&nbsp; &nbsp;AM_0528<br> &nbsp;&nbsp; &nbsp;ERZ494233&nbsp;&nbsp; &nbsp;AM_0546<br> &nbsp;&nbsp; &nbsp;ERZ494234&nbsp;&nbsp; &nbsp;AM_0608<br> &nbsp;&nbsp; &nbsp;ERZ494235&nbsp;&nbsp; &nbsp;AM_0615<br> &nbsp;&nbsp; &nbsp;ERZ494236&nbsp;&nbsp; &nbsp;AM_0616<br> &nbsp;&nbsp; &nbsp;ERZ494237&nbsp;&nbsp; &nbsp;AM_0619<br> &nbsp;&nbsp; &nbsp;ERZ494238&nbsp;&nbsp; &nbsp;AM_0621<br> &nbsp;&nbsp; &nbsp;ERZ494239&nbsp;&nbsp; &nbsp;AM_0630<br> &nbsp;&nbsp; &nbsp;ERZ494240&nbsp;&nbsp; &nbsp;AM_0643<br> &nbsp;&nbsp; &nbsp;ERZ494241&nbsp;&nbsp; &nbsp;AM_0729<br> &nbsp;&nbsp; &nbsp;ERZ494242&nbsp;&nbsp; &nbsp;AM_0764<br> &nbsp;&nbsp; &nbsp;ERZ494243&nbsp;&nbsp; &nbsp;AM_0832<br> &nbsp;&nbsp; &nbsp;ERZ494244&nbsp;&nbsp; &nbsp;AM_0849<br> &nbsp;&nbsp; &nbsp;ERZ494245&nbsp;&nbsp; &nbsp;AM_0854<br> &nbsp;&nbsp; &nbsp;ERZ494246&nbsp;&nbsp; &nbsp;AM_0876<br> &nbsp;&nbsp; &nbsp;ERZ494247&nbsp;&nbsp; &nbsp;AM_0902<br> &nbsp;&nbsp; &nbsp;ERZ494248&nbsp;&nbsp; &nbsp;AM_0936<br> &nbsp;&nbsp; &nbsp;ERZ494249&nbsp;&nbsp; &nbsp;AM_1003<br> &nbsp;&nbsp; &nbsp;ERZ494250&nbsp;&nbsp; &nbsp;AM_1104<br> &nbsp;&nbsp; &nbsp;ERZ494251&nbsp;&nbsp; &nbsp;AM_1111<br> &nbsp;&nbsp; &nbsp;ERZ494252&nbsp;&nbsp; &nbsp;AM_1205<br> &nbsp;&nbsp; &nbsp;ERZ494253&nbsp;&nbsp; &nbsp;AM_1312<br> &nbsp;&nbsp; &nbsp;ERZ494254&nbsp;&nbsp; &nbsp;AM_1409<br> &nbsp;&nbsp; &nbsp;ERZ494255&nbsp;&nbsp; &nbsp;AM_1503<br> &nbsp;&nbsp; &nbsp;ERZ494256&nbsp;&nbsp; &nbsp;AM_1603<br> &nbsp;&nbsp; &nbsp;ERZ494257&nbsp;&nbsp; &nbsp;AM_1606<br> &nbsp;&nbsp; &nbsp;ERZ494258&nbsp;&nbsp; &nbsp;AM_1801<br> &nbsp;&nbsp; &nbsp;ERZ494259&nbsp;&nbsp; &nbsp;AM_1811<br> &nbsp;&nbsp; &nbsp;ERZ494260&nbsp;&nbsp; &nbsp;AM_2104<br> &nbsp;&nbsp; &nbsp;ERZ494261&nbsp;&nbsp; &nbsp;AM_2116<br> &nbsp;&nbsp; &nbsp;ERZ494262&nbsp;&nbsp; &nbsp;AM_2124<br> &nbsp;&nbsp; &nbsp;ERZ494263&nbsp;&nbsp; &nbsp;AM_2202<br> &nbsp;&nbsp; &nbsp;ERZ494264&nbsp;&nbsp; &nbsp;AM_2207<br> &nbsp;&nbsp; &nbsp;ERZ494265&nbsp;&nbsp; &nbsp;AM_2208<br> &nbsp;&nbsp; &nbsp;ERZ494266&nbsp;&nbsp; &nbsp;AM_2324<br> &nbsp;&nbsp; &nbsp;ERZ494267&nbsp;&nbsp; &nbsp;AM_2502<br> &nbsp;&nbsp; &nbsp;ERZ494268&nbsp;&nbsp; &nbsp;AM_2804<br> &nbsp; &nbsp;</p> <p>3. ACKNOWLEDGEMENTS<br> &nbsp; &nbsp;</p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of S&atilde;o Carlos, S&atilde;o Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Cient&iacute;fico e Tecnol&oacute;gico (CNPq), as well as, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; the spanish funding organ Consejo Superior de Investigaciones Cient&iacute;ficas (CSIC).</p> <p>This study was financed in part by the Coordena&ccedil;&atilde;o de Aperfei&ccedil;oamento de Pessoal de N&iacute;vel Superior - Brasil (CAPES) - Finance Code 001.</p> <p>&nbsp;</p> <p>4. CONTACT INFORMATION</p> <p>&nbsp;&nbsp; Current curators:</p> <p>&nbsp;&nbsp; - C&eacute;lio Dias Santos J&uacute;nior (celio.diasjunior@gmail.com)<br> &nbsp;&nbsp; - Flavio Henrique-Silva (dfhs@ufscar.br)<br> &nbsp;&nbsp; - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> &nbsp;</p> <p>5. COPYRIGHT NOTICE</p> <p>&nbsp;&nbsp; Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> &nbsp;&nbsp; Copyright (C) 2018 The AMnrGC consortium.</p> <p>&nbsp;&nbsp; This database is provided &ldquo;as is&rdquo; and without any warranty of any kind,<br> &nbsp;&nbsp; of openly available. You can redistribute and/or modify it<br> &nbsp;&nbsp; as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p>&nbsp;&nbsp; &nbsp;https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Supplementary data to accompany "phyloFlash: A pipeline for rapid SSU rRNA-targeted profiling of metagenomes"

<p>Usage examples of the phyloFlash pipeline applied to shotgun metagenomic data sets.</p> <p>The phyloFlash software is available from https://github.com/HRGV/phyloFlash. Examples were generated with phyloFlash v3.3b.</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Data for the publication "Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer"

<p>This dataset encompasses all data needed to reproduce the analyses presented in&nbsp;<a href="https://www.nature.com/articles/s41591-019-0406-6">Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer</a></p> <p>You can also check the&nbsp;<a href="https://github.com/zellerlab/crc_meta">GitHub repository</a></p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Raw metagenomic data from sweep net samples collected in 2016 as part of the Slikok Creek Watershed Biotic Inventory

<p>We set out to inventory vascular plants, bryophytes, lichens, birds, arthropods, and earthworms on a grid of sites in the portion of Slikok Creek watershed that is on the Kenai National Wildlife Refuge, Kenai Peninsula, Alaska. Occurrence data, images, and field data sheets from this project are available via an <a href="https://arctosdb.org/">Arctos</a> project page at <a href="http://arctos.database.museum/project/10002227">http://arctos.database.museum/project/10002227</a>.</p> <p>This dataset includes the raw FASTQ files from metagenomic processing and associated collection data. Of the 160 sweep net samples collected, 125 were selected for High Throughput Sequencing and shipped to RTL Genomics (<a href="http://rtlgenomics.com">http://rtlgenomics.com</a>) for extraction and sequencing steps. Sequencing was performed on an Illumina MiSeq platform and reads were processed using RTL Genomics&rsquo; standard methods with the mlCOIlintF/HCO2198 primer set of Leray et al. (2013), yielding a 313 bp region of the COI gene.</p> <p>Collection data are included in the file <code>ArctosData_43C6167EB1.csv</code> downloaded from Arctos. Extraction methods and sequencing methods provided by RTL Genomics are included in the files <code>Bowser 4869.pdf</code> and <code>Illumina MiSeq Two-Step Method 454 profile only.docx</code>. Primers used are provided in the file <code>Bowser_4869M.txt</code>. The archive <code>FASTQ.zip</code> contains all of the resulting FASTQ files.</p> <p>These sequence data have also been been published to GenBank&#39;s Sequence Read Archive in accessions&nbsp;<a href="http://trace.ncbi.nlm.nih.gov/Traces/sra/?run=SRR10454582">SRR10454582</a>&ndash;<a href="http://trace.ncbi.nlm.nih.gov/Traces/sra/?run=SRR10454706">SRR10454706</a>&nbsp;under BioProject&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA427721">PRJNA427721</a>.</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

The Metagenome-Assembled Genome Inventory for Children (MAGIC)

<div> <div>Existing microbiota databases are biased towards adult samples, hampering accurate profiling of the infant gut microbiome. Here, we generated a **M**etagenome-**A**ssembled **G**enome **I**nventory for **C**hildren (**MAGIC**) from a large collection of bulk and viral-like particle-enriched metagenomes from 0-7 years of age, encompassing `3,299` prokaryotic and `139,624` viral species-level genomes, `8.5%` and `63.9%` of which are unique to MAGIC. MAGIC improves early-life microbiome profiling, with the greatest improvement in read mapping observed in Africans. We then identified `54` candidate keystone species, including several *Bifidobacterium spp.* and four phages, forming guilds that fluctuated in abundance with time. Their abundances were reduced in preterm infants and were associated with childhood allergies. By analyzing the *B. longum* pangenome, we found evidence of phage-mediated evolution and quorum sensing-related ecological adaptation. Together, the MAGIC database recovers genomes that enable characterization of dynamics of early-life microbiomes, identification of candidate keystone species, and strain-level study of target species.</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

A curated data resource of 214K metagenomes for characterization of the global resistome

<p><strong>Data files of the curated resource of 214K metagenomes </strong> <strong>for characterization of the global resistome.</strong></p> <p>We have retrieved 214K metagenomic samples and now share the results here on Zenodo of our large-scale read mapping effort.</p> <p>There are five tables uploaded in three formats (TSV, HDF and MySQL dump):</p> <ul> <li>metadata.* : contains metadata for all sequencing runs.</li> <li>ARG.* : contain read alignment counts of antimicrobial resistance genes (ARGs).</li> <li>rRNA.* : contain read alignment counts of 16S/18S rRNA genes.<sup>1</sup></li> <li>diversity.* : contain diversity measures for ARGs and two taxonomic groups of rRNA genes (phylum, genus).</li> <li>ResFinder_anno.* : contain sequence information on the different ARGs, such as gene_lengths, resistance class, etc.</li> </ul> <p>Note that the HDF file rRNA.h5 is split into batches of 10,000 rows. To load it, the keys are in the format of &quot;table_{i}&quot;, where i=0,1,2,..,4736</p> <p>Details on the different tables are available at https://hmmartiny.github.io/mARG/</p> <p>Additionaly, we have shared the data used to create the figures in the manuscript in the ZIP file named &quot;figure_data.zip&quot;.</p> <p>Any further questions or issues, please contact H.-M. Martiny at hanmar@food.dtu.dk</p> <p>&nbsp;</p> <p><strong>Update log</strong>:</p> <p>* 2023-01-20: Update Diversity tables due to wrong total_fragments entered for ~250 run_accessions.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Reconstruction of prokaryotic genomes from ten termite gut metagenomes using two distinct workflows: SnakeMAGs and ATLAS.

<p><strong><em>SnakeMAGs</em></strong> (Nachida Tadrent, Franck Dedeine, Vincent Herv&eacute; (Submitted).&nbsp;<em>SnakeMAGs</em>: a simple, efficient, flexible and scalable workflow to reconstruct prokaryotic genomes from metagenomes<em>.</em> <a href="https://doi.org/10.5281/zenodo.7303463">https://doi.org/10.5281/zenodo.7303463</a>; https://github.com/Nachida08/SnakeMAGs) is a workflow for building MAGs (Metagenome Assembled Genomes) from raw Illumina metagenomic reads. During the test phase of the development of this tool, a comparative analysis with another workflow called ATLAS v2.9.1 (<em>Kieser </em>et al, 2020) was performed. To compare these two workflows, we analyzed ten publicly available termite gut metagenomes (accession numbers: SRR10402454; SRR14739927; SRR8296321; SRR8296327; SRR8296329; SRR8296337; SRR8296343; DRR097505; SRR7466794; SRR7466795) from five different studies :&nbsp;Waidele et al, 2019; Tokuda et al, 2018; Romero Victorica et al, 2020; Moreira et al, 2021; and Calusinska et al, 2020.</p> <p>In this repository, we provide the configuration files that were used to launch each of the workflows (SnakeMAGs_config.yaml and ATLAS_config.yaml), &nbsp;as well as the obtained results, <em>i.e. </em>the MAGs reconstructed from each metagenome and their taxonomic classification.</p>

opencc-by-4.0Nov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record