Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

12

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

12 results for “genome binning”

Learn how ShareScore rates datasets ↗
zenodo48/100

Bin-assembled Escherichia coli genomes from a study in Punjab, Pakistan

<h2>Bin-assembled <em>Escherichia coli</em> genomes from Punjab, Pakistan</h2> <p>These assemblies are a part of a cross-sectional study conducted in Punjab, Pakistan aimed at investigating <em>E. coli</em> colonisation diversity in healthy carriage with the use of CLED enrichment plates.</p> <h3><strong>About</strong></h3> <h4><strong>Version history</strong></h4> <p><strong>v0.1.1 (current version)</strong></p> <ul> <li>Added reference to the study.</li> </ul> <p><strong>v0.1.0</strong></p> <ul> <li>Added brief description with a few missing parts.</li> </ul> <h4><strong>Distribution</strong></h4> <p>If you use these assemblies in your study please cite the source as appropriate.&nbsp;These assemblies are made available under a CC-BY 4.0 license.</p> <h4><strong>Citation</strong></h4> <p>Khawaja, T., M&auml;klin, T., Kallonen, T. et al. Deep sequencing of <em>Escherichia coli</em> exposes colonisation diversity and impact of antibiotics in Punjab, Pakistan. Nature Communications 15, 5196 (2024).&nbsp;<a href="https://doi.org/10.1038/s41467-024-49591-5">https://doi.org/10.1038/s41467-024-49591-5</a></p> <h3><strong>Methods briefly</strong></h3> <h4><strong>Species identification</strong></h4> <p>Sequencing data from the ENA project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB36642">PRJEB36642</a> was error-corrected with <a href="https://github.com/opengene/fastp">fastp</a> and pseudoaligned with <a href="https://github.com/algbio/themisto">Themisto</a> against a species-level index (available from <a href="https://doi.org/10.5281/zenodo.6656881">https://doi.org/10.5281/zenodo.6656881</a>). Reads were assigned to species using the <a href="https://doi.org/10.1099%2Fmgen.0.000691">mSWEEP/mGEMS pipeline</a> as described in <a href="https://www.nature.com/articles/s41467-022-35178-5">https://www.nature.com/articles/s41467-022-35178-5</a>.</p> <h4><strong>Lineage identification</strong></h4> <p>Read from the species-level bins were again pseudoaligned with Themisto against an <em>E. coli</em> index (will be made available in a later version). Lineage-level assignment was performed using mSWEEP and mGEMS at the level of <a href="https://genome.cshlp.org/content/29/2/304">PopPUNK</a> sequence clusters. The created bins were screened with <a href="https://github.com/tmaklin/coreutils_demix_check">demix_check</a> and bins that received a score of 1 or 2 were kept. Data in the kept bins were assembled with <a href="https://github.com/tseemann/shovill">shovill</a> and the bin-assembled genomes (BAGs) were quality controlled with <a href="https://genome.cshlp.org/content/25/7/1043">checkm</a> for &gt;= 90% completeness and &lt;= 10% contamination. Finally, BAGs shorter than 4 Mb or longer than 6 Mb were removed.</p> <h3><strong>Contact</strong></h3> <p>Tommi M&auml;klin &lt;tommi'at'maklin.fi&gt;.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

BBS phase 1 & phase 2 high quality E. coli bin assembled genomes

<p>1,402&nbsp;<em>Escherichia coli</em> bin assembled genomes derived from the metagenome data collected as part of the <a href="https://www.ucl.ac.uk/global-health/research/a-z/baby-biome-study">BabyBiome study (BBS)</a> phase 1 &amp; phase 2.</p> <p>The data in this upload was first published as part of "<em>Group 2 and 3 ABC-transporter dependant K-antigen loci contribute significantly to variation in the invasive potential of Escherichia coli"</em>&nbsp; (Gladstone et al. 2024, to be released).</p> <h2>Files</h2> <p>Assembly data:</p> <ul> <li>BBS_E_coli_BAGs.tar: Archive containing sequences of the 1,402 bin assembled genomes.</li> <li>BBS_E_coli_metadata.tsv: Table linking the sequence assemblies to the subject data.</li> </ul> <p>Capsule predictions:</p> <ul> <li>BBS_E_coli_Kaptive_output.csv: Capsule predictions for all sequence data.</li> <li>BBS_E_coli_deduplicated_sequences_IDs.txt: Filenames for assemblies that constitute the 873 deduplicated sequences analysed in Gladstone et al. 2024.</li> </ul> <p>Quality control data:</p> <ul> <li>BBS_E_coli_demix_check_scores.tsv: Output from demix_check for the sequence assemblies.</li> <li>BBS_E_coli_checkm_results.tsv: Output from checkm.</li> <li>BBS_E_coli_gunc_results.tsv: Output from gunc.</li> </ul> <h2>Methods</h2> <h3>Bin assembled genomes</h3> <p>Source data:</p> <ul> <li>BBS phase 1: <a href="https://doi.org/10.1038/s41586-019-1560-1">Shao et al. 2019</a></li> <li>BBS phase 2: <a href="https://doi.org/10.1038/s41564-024-01804-9">Shao et al. 2024</a></li> </ul> <p>The data was produced using the mSWEEP and mGEMS pipeline (<a href="https://doi.org/10.12688/wellcomeopenres.15639.2">M&auml;klin et al. 2020</a> &amp; <a href="https://doi.org/10.1099/mgen.0.000691">M&auml;klin et al. 2021</a>) following the steps described in <a href="https://doi.org/10.1038/s41467-024-49591-5">Khawaja, M&auml;klin, Kallonen, et al. 2024</a>.</p> <h3>Quality control</h3> <p>The BAGs in this upload were filtered with demix_check (<a href="https://github.com/harry-thorpe/demix_check">https://github.com/harry-thorpe/demix_check</a>) and only those with a quality score 1 or 2 are included. For the capsule type annotations, contigs shorter than 5,000bp were removed but the short contigs are still present in the uploaded files). Further QC data is available from checkm (<a href="https://genome.cshlp.org/content/25/7/1043.short">Parks et al. 2015</a>) and gunc (<a href="https://link.springer.com/article/10.1186/s13059-021-02393-0">Orakov et al. 2022</a>) results.</p> <h3>Multilocus sequence typing</h3> <p>Sequence type (ST) was determined using fastmlst (<a href="https://journals.sagepub.com/doi/10.1177/11779322211059238">Guerrero-Araya et al. 2021</a>) with the `ecoli#1` database.</p> <h3>PopPUNK&nbsp; clustering</h3> <p>Sequence clusters (SC) correspond to the database available from <a href="https://zenodo.org/records/12528310">https://zenodo.org/records/12528310</a> and were created using PopPUNK (<a href="https://genome.cshlp.org/content/29/2/304.short">Lees et al. 2019</a>). Construction is described in <a href="https://doi.org/10.1038/s41467-024-49591-5">Khawaja, M&auml;klin, Kallonen, et al. 2024</a>.</p> <h3>Capsule type annotations</h3> <p>The capsule type annotations were created using Kaptive (<a href="https://doi.org/10.1099/mgen.0.000800">Lam et al. 2022</a>) with an&nbsp;<em>E. coli</em> specific database available from <a href="https://github.com/rgladstone/EC-K-typing">https://github.com/rgladstone/EC-K-typing</a> and described in Gladstone et al. 2024.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

MetaBAT 2.12.1 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly

Genome binning of the gold standard pooled assembly <br><strong>Software: </strong>MetaBAT<br><strong>SoftwareVersion: </strong>2.12.1<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://bitbucket.org/berkeleylab/metabat<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> bowtie2-build anonymous_gsa_pooled.fasta anonymous_gsa_pooled.fasta<br>for i in {0..63}; do bowtie2 -q --threads 30 --fr -x anonymous_gsa_pooled.fasta --interleaved sample_${i}/anonymous_reads.fq -S anonymous_reads_sample_${i}.sam ; done<br>for i in {0..63}; do samtools view -b sample_${i}.sam -o anonymous_reads_sample_${i}.bam &amp; done<br>for i in {0..63}; do samtools sort anonymous_reads_sample_${i}.bam -o anonymous_reads_sample_${i}.sorted.bam ; done<br>for i in {0..63}; do samtools index anonymous_reads_sample_${i}.sorted.bam ; done<br>runMetaBat.sh -l anonymous_gsa_pooled.fasta anonymous_reads_sample_*.sorted.bam

opencc-by-4.0Jan 2020View details →
zenodo40/100

"Genome binning of viral entities from bulk metagenomics data" - CAMISIM simulated datasets and genomes

<p><strong>Genome binning of viral entities from bulk metagenomics data</strong></p> <p>&nbsp;</p> <p><strong>Authors</strong></p> <p><strong>Joachim Johansen1,2, Damian R. Plichta2, Jakob Nybo Nissen1,3, Marie Louise Jespersen1,4, Shiraz A. Shah5, Ling Deng6, Jakob Stokholm5,6, Hans Bisgaard5, Dennis Sandris Nielsen6, S&oslash;ren S&oslash;rensen7, Simon Rasmussen1</strong></p> <p>&nbsp;</p> <p><strong>Affiliations</strong></p> <p>1 Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen N, Denmark</p> <p>2 Infectious Disease and Microbiome Program, Broad Institute of MIT and Harvard, Cambridge, MA, USA</p> <p>3 Statens Serum Institut, Viral &amp; Microbial Special diagnostics, Copenhagen, Denmark</p> <p>4 National Food Institute, Technical University of Denmark, Kongens Lyngby, Denmark</p> <p>5 Copenhagen Prospective Studies on Asthma in Childhood (COPSAC), Herlev and Gentofte Hospital, University of Copenhagen, Copenhagen, Denmark</p> <p>6 Section of Food Microbiology and Fermentation, Department of Food Science, Faculty of Science, University of Copenhagen, Copenhagen, Denmark</p> <p>7 Section of Microbiology, Department of Biology, University of Copenhagen, Copenhagen, Denmark</p> <p><strong>Methods description</strong></p> <p>We compared the viral binning performance of VAMB and MetaBAT2 using the official CAMI consortium method to create assemblies and metagenome profiles. To this end we generated 3 different metagenome compositions with up to 308 reference genomes; one mixed with bacteria, plasmids and viruses to test binning in complex samples i.e. high diversity (1), one with only crass-like viruses to test binning with highly similar viruses i.e. high relatedness (2) and a set of small-viruses (&lt;6,000 bp) including members of the Microviridae family to address the bias of size (3). Bacterial genomes were gathered from NCBIs refseq genome repository 2021, plasmids from the PLSDB database (v. 2021_06_23)&nbsp;and viral genomes from the recent MGV database.&nbsp;&nbsp;</p> <p>Dataset A contained a mixture of bacteria (N=8), plasmids (N=20) and viruses (N=280) to test binning in complex samples, i.e. high diversity. Dataset B contained only crass-like viruses (N=80) to test binning with highly similar viruses i.e. high relatedness. Dataset C contained small-viruses (N=50, &lt;6,000 bp) of the Microviridae family to address the bias of size. Bacterial genomes were sampled from the Refseq genome repository 2021, plasmids from the PLSDB database&nbsp; and viral genomes from the recent MGV database (Nayfach, et al. Nature Microbiology 2021).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

TOPC_bin_586 metagenome assembled genome (MAG)

<p><strong>Contig,&nbsp;gene sequences and functional annotation of the&nbsp;<em>TOPC_bin_586</em> metagenome assembled genome (MAG)</strong></p> <p>Data available:</p> <ol> <li>Nucleotide sequences of the contigs composing the MAG [<em>topc.bin.586.fna</em>]</li> <li>Amino acid sequences of the genes (open reading frames, ORFs) [<em>topc.bin.586_ORFs.faa</em>]</li> <li>Functional annotation table (tab-delimited) for the ORFs [<em>topc.bin.586_ORFs_annotation.tsv</em>]</li> </ol>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Metagenome-assembled genomes(MAGs) generated by MetaCC binning

<p>MAGs&nbsp;generated by MetaCC binning from the human gut short-read, the wastewater (WW) short-read, the cow rumen long-read, and the sheep gut long-read metaHi-C datasets</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

DAS Tool 1.1.2 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly

Genome binning of the gold standard pooled assembly. Refinement of the binning output of MaxBin 2.2.7, MetaBAT 2.12.1, CONCOCT 1.0.0, and DAS Tool 1.1.2.<br><strong>Software: </strong>DAS Tool<br><strong>SoftwareVersion: </strong>1.1.2<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://github.com/cmks/DAS_Tool<br><strong>DockerImage:</strong> cami/das_tool:1.1.2<br><strong>IsBiobox:</strong> No<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> DAS_Tool -i binning_concoct1.0.0,binning_maxbin2.2.7,binning_metabat2.12.1 -c anonymous_gsa_pooled.fasta -o output --search_engine diamond

opencc-by-4.0Jan 2020View details →
zenodo32/100

CONCOCT 1.0.0 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly

Genome binning of the gold standard pooled assembly <br><strong>Software: </strong>CONCOCT<br><strong>SoftwareVersion: </strong>1.0.0<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://github.com/BinPro/CONCOCT<br><strong>DockerImage:</strong> quay.io/biocontainers/concoct:1.0.0--py37h88e4a8a_5<br><strong>IsBiobox:</strong> No<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> for i in {0..63}; do bowtie2 -q --threads 30 --fr -x anonymous_gsa_pooled.fasta --interleaved sample_${i}/anonymous_reads.fq -S anonymous_reads_sample_${i}.sam ; done<br>for i in {0..63}; do samtools view -b sample_${i}.sam -o anonymous_reads_sample_${i}.bam &amp; done<br>for i in {0..63}; do samtools sort anonymous_reads_sample_${i}.bam -o anonymous_reads_sample_${i}.sorted.bam ; done<br>for i in {0..63}; do samtools index anonymous_reads_sample_${i}.sorted.bam ; done<br>cut_up_fasta.py anonymous_gsa_pooled.fasta -c 10000 -o 0 --merge_last -b contigs_10K.bed &gt; contigs_10K.fa<br>concoct_coverage_table.py contigs_10K.bed /host/benchmarking/fmeyer/output/bowtie2/mouse_gut/sorted_bam/anonymous_reads_sample_*.sorted.bam &gt; coverage_table.tsv<br>concoct --composition_file contigs_10K.fa --coverage_file coverage_table.tsv -b<br>merge_cutup_clustering.py clustering_gt1000.csv &gt; clustering_merged.csv

opencc-by-4.0Jan 2020View details →
zenodo32/100

MaxBin 2.2.7 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly

Genome binning of the gold standard pooled assembly <br><strong>Software: </strong>MaxBin<br><strong>SoftwareVersion: </strong>2.2.7<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://sourceforge.net/projects/maxbin/<br><strong>DockerImage:</strong> cami/maxbin:2.2.7<br><strong>IsBiobox:</strong> No<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> run_MaxBin.pl -thread 16 -contig anonymous_gsa_pooled.fasta -out output -reads sample_0/reads/anonymous_reads.fq -reads2 sample_1/reads/anonymous_reads.fq -reads3 sample_2/reads/anonymous_reads.fq -reads4 sample_3/reads/anonymous_reads.fq -reads5 sample_4/reads/anonymous_reads.fq -reads6 sample_5/reads/anonymous_reads.fq -reads7 sample_6/reads/anonymous_reads.fq -reads8 sample_7/reads/anonymous_reads.fq -reads9 sample_8/reads/anonymous_reads.fq -reads10 sample_9/reads/anonymous_reads.fq -reads11 sample_10/reads/anonymous_reads.fq -reads12 sample_11/reads/anonymous_reads.fq -reads13 sample_12/reads/anonymous_reads.fq -reads14 sample_13/reads/anonymous_reads.fq -reads15 sample_14/reads/anonymous_reads.fq -reads16 sample_15/reads/anonymous_reads.fq -reads17 sample_16/reads/anonymous_reads.fq -reads18 sample_17/reads/anonymous_reads.fq -reads19 sample_18/reads/anonymous_reads.fq -reads20 sample_19/reads/anonymous_reads.fq -reads21 sample_20/reads/anonymous_reads.fq -reads22 sample_21/reads/anonymous_reads.fq -reads23 sample_22/reads/anonymous_reads.fq -reads24 sample_23/reads/anonymous_reads.fq -reads25 sample_24/reads/anonymous_reads.fq -reads26 sample_25/reads/anonymous_reads.fq -reads27 sample_26/reads/anonymous_reads.fq -reads28 sample_27/reads/anonymous_reads.fq -reads29 sample_28/reads/anonymous_reads.fq -reads30 sample_29/reads/anonymous_reads.fq -reads31 sample_30/reads/anonymous_reads.fq -reads32 sample_31/reads/anonymous_reads.fq -reads33 sample_32/reads/anonymous_reads.fq -reads34 sample_33/reads/anonymous_reads.fq -reads35 sample_34/reads/anonymous_reads.fq -reads36 sample_35/reads/anonymous_reads.fq -reads37 sample_36/reads/anonymous_reads.fq -reads38 sample_37/reads/anonymous_reads.fq -reads39 sample_38/reads/anonymous_reads.fq -reads40 sample_39/reads/anonymous_reads.fq -reads41 sample_40/reads/anonymous_reads.fq -reads42 sample_41/reads/anonymous_reads.fq -reads43 sample_42/reads/anonymous_reads.fq -reads44 sample_43/reads/anonymous_reads.fq -reads45 sample_44/reads/anonymous_reads.fq -reads46 sample_45/reads/anonymous_reads.fq -reads47 sample_46/reads/anonymous_reads.fq -reads48 sample_47/reads/anonymous_reads.fq -reads49 sample_48/reads/anonymous_reads.fq -reads50 sample_49/reads/anonymous_reads.fq -reads51 sample_50/reads/anonymous_reads.fq -reads52 sample_51/reads/anonymous_reads.fq -reads53 sample_52/reads/anonymous_reads.fq -reads54 sample_53/reads/anonymous_reads.fq -reads55 sample_54/reads/anonymous_reads.fq -reads56 sample_55/reads/anonymous_reads.fq -reads57 sample_56/reads/anonymous_reads.fq -reads58 sample_57/reads/anonymous_reads.fq -reads59 sample_58/reads/anonymous_reads.fq -reads60 sample_59/reads/anonymous_reads.fq -reads61 sample_60/reads/anonymous_reads.fq -reads62 sample_61/reads/anonymous_reads.fq -reads63 sample_62/reads/anonymous_reads.fq -reads64 sample_63/reads/anonymous_reads.fq

opencc-by-4.0Jan 2020View details →
zenodo28/100

Gold standard genome and taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly

<p>Gold standard genome and taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate</p>

opencc-by-4.0Jan 2020View details →
zenodo28/100

QDNAseq.hg38: QDNAseq bin annotation for the human genome build hg38

<p><strong>QDNAseq</strong> bin annotations of size 1, 5, 10, 15, 30, 50, 100, 500, and 1000 kbp for the human genome build hg38.</p>

openother-openNov 2020View details →
geo20/100

High-density deletion bin map of wheat D genome

GEO Series GSE71190. Aegilops tauschii; Triticum aestivum. 96 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenJul 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record