Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

99

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

99 results for “bacterial genome”

Learn how ShareScore rates datasets ↗
zenodo36/100

Parallel dynamics of bacterial genome reduction across independent transitions to endosymbiosis.

<p>The establishment of symbiosis dramatically alters the evolution of the associated species, making symbiotic systems ideal models for studying the impact of lifestyle changes on genomes. Here, we focused on Enterobacterales, a large and ancient bacterial lineage that includes endosymbionts with diverse host associations, ranging from gut inhabitants to intracellular environments, and from horizontal to vertical transmission. Leveraging over two hundred genomes, along with cutting-edge single-copy gene concatenation and multi-copy gene family approaches, we inferred a robust phylogenetic framework that supports eleven independent transitions to endosymbiosis. Inferences on patterns of genome evolution confirm previous hypotheses about the processes underlying genome reduction: a substantial spike in gene loss always occurs simultaneously with the establishment of endosymbiosis, while a reduction in gene acquisition mechanisms is associated with the subsequent genome erosion. Furthermore, gene family loss frequencies were correlated across independent endosymbiotic clades; genes with more conserved functions and stronger constraints on sequence evolution are lost less frequently, suggesting that differences in gene essentiality and dispensability drive the observed parallelism. Our analyses contribute to the coming of age of the theory of genome evolution in symbiotic associations and provide novel insights into the importance of recombination as an opposing force against genome erosion.</p>

opencc-by-4.0Jun 2024View details →
ClinicalTrials.gov36/100

Bacterial Genomic Sequencing in Overactive Bladder

ClinicalTrials.gov study NCT01642277. IPD Sharing: Not stated. Countries: 1. Publications: 17.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad36/100

Integrated genome-wide investigations of the housefly, a global vector of diseases reveal unique dispersal patterns and bacterial communities across farms

Open the record for dataset details and reuse information.

publicFeb 2020View details →
dryad36/100

Data from: Entangled fates of holobiont genomes during invasion: nested bacterial and host diversities in Caulerpa taxifolia

Open the record for dataset details and reuse information.

publicJan 2017View details →
dryad36/100

Data from: Feature sequence-based genome mining uncovers the hidden diversity of bacterial siderophore pathways

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad36/100

Dynamics of bacterial operons during genome-wide stresses are influenced by premature terminations and internal promoters

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad32/100

Genome reduction is associated with bacterial pathogenicity across different scales of temporal and ecological divergence - between species core gene alignments

<p><span>Emerging bacterial pathogens threaten global health and food security, and so it is important to ask whether these transitions to pathogenicity have any common features. We present a systematic study of the claim that pathogenicity is associated with genome reduction and gene loss. We compare broad-scale patterns across all bacteria, with detailed analyses of <i>Streptococcus suis</i>, an emerging zoonotic pathogen of pigs, which has undergone multiple transitions between disease and carriage forms. We find that pathogenicity is consistently associated with reduced genome size across three scales of divergence (between species within genera, and between and within genetic clusters of <i>S. suis</i>). While genome reduction is also found in mutualist and commensal bacterial endosymbionts, genome reduction in pathogens cannot be solely attributed to the features of their ecology that they share with these species, i.e. host restriction or intracellularity. Moreover, other typical correlates of genome reduction in endosymbionts (reduced metabolic capacity, reduced GC content, and the transient expansion of non-functional elements) are not consistently observed in pathogens. Together, our results indicate that genome reduction is a predictive marker of pathogenicity in bacteria.</span></p>

opencc-zeroNov 2020View details →
dryad32/100

Data from: A genomic perspective on a new bacterial genus and species from the Alcaligenaceae family, Basilea psittacipulmonis

Background: A novel Gram-negative, non-haemolytic, non-motile, rod-shaped bacterium was discovered in the lungs of a dead parakeet (Melopsittacus undulatus) that was kept in captivity in a petshop in Basel, Switzerland. The organism is described with a chemotaxonomic profile and the nearly complete genome sequence obtained through the assembly of short sequence reads. Results: Genome sequence analysis and characterization of respiratory quinones, fatty acids, polar lipids, and biochemical phenotype is presented here. Comparison of gene sequences revealed that the most similar species is Pelistega europaea, with BLAST identities of only 93% to the 16S rDNA gene, 76% identity to the rpoB gene, and a similar GC content (~43%) as the organism isolated from the parakeet, DSM 24701 (40%). The closest full genome sequences are those of Bordetella spp. and Taylorella spp. High-throughput sequencing reads from the Illumina-Solexa platform were assembled with the Edena de novo assembler to form 195 contigs comprising the ~2 Mb genome. Genome annotation with RAST, construction of phylogenetic trees with the 16S rDNA (rrs) gene sequence and the rpoB gene, and phylogenetic placement using other highly conserved marker genes with ML Tree all suggest that the bacterial species belongs to the Alcaligenaceae family. Analysis of samples from cages with healthy parakeets suggested that the newly discovered bacterial species is not widespread in parakeet living quarters. Conclusions: Classification of this organism in the current taxonomy system requires the formation of a new genus and species. We designate the new genus Basilea and the new species psittacipulmonis. The type strain of Basilea psittacipulmonis is DSM 24701 (= CIP 110308 T, 16S rDNA gene sequence Genbank accession number JX412111 and GI 406042063).

opencc-zeroDec 2013View details →
dryad32/100

Data from: Host‐derived population genomics data provides insights into bacterial and diatom composition of the killer whale skin

Recent exploration into the interactions and relationship between hosts and their microbiota has revealed a connection between many aspects of the host's biology, health and associated micro‐organisms. Whereas amplicon sequencing has traditionally been used to characterize the microbiome, the increasing number of published population genomics data sets offers an underexploited opportunity to study microbial profiles from the host shotgun sequencing data. Here, we use sequence data originally generated from killer whale Orcinus orca skin biopsies for population genomics, to characterize the skin microbiome and investigate how host social and geographical factors influence the microbial community composition. Having identified 845 microbial taxa from 2.4 million reads that did not map to the killer whale reference genome, we found that both ecotypic and geographical factors influence community composition of killer whale skin microbiomes. Furthermore, we uncovered key taxa that drive the microbiome community composition and showed that they are embedded in unique networks, one of which is tentatively linked to diatom presence and poor skin condition. Community composition differed between Antarctic killer whales with and without diatom coverage, suggesting that the previously reported episodic migrations of Antarctic killer whales to warmer waters associated with skin turnover may control the effects of potentially pathogenic bacteria such as Tenacibaculum dicentrarchi. Our work demonstrates the feasibility of microbiome studies from host shotgun sequencing data and highlights the importance of metagenomics in understanding the relationship between host and microbial ecology.

opencc-zeroDec 2017View details →
zenodo32/100

C-Sibelia bacterial genome comparison tool: dataset and scripts

<p>Input data and scripts used for comparing the&nbsp;Staphylococcus aureus RN4220 assembly genome and the NCTC 8325 reference genome using C-Sibelia.</p>

opengpl-2.0Nov 2013View details →
zenodo32/100

Supplementary Information: CHAPTER 2 - Unveiling genomic features linked to traits of plant-growth-promoting bacterial communities from sugarcane

<p>Appendix A. Summary of counts of subreads and circular consensus sequencing (CCS) sequences obtained for PacBio sequencing of SMRT libraries. (EMS_1.xlsx)</p> <p>Appendix B. Taxonomy assignment of MAGs at the higher taxonomic rank obtained from GTDB-tk and Kraken tools. (EMS_2.xlsx)</p> <p>Appendix C. Report of the classification workflow using GTDB-tk. (EMS_3.xlsx)</p> <p>Appendix D.&nbsp; Matrix of the KEGG Orthology (KOs) frequencies annotated by the EnrichM tool. (EMS_4.xlsx)</p> <p>Appendix E. Reconstruction and completeness of KEGG modules&nbsp; annotated by EnrichM. The asterisks (*) in the header represent additional values obtained by the script &lsquo;classKEGGModules.pl&rsquo; (https://github.com/dgpinheiro/bioinfoutilities) to estimate PGPTs in KEGG modules. (EMS_5.xlsx)</p> <p>Appendix F. The secondary metabolite biosynthesis gene clusters (BGCs) identified with AntiSMASH. (EMS_6.xlsx)</p> <p>Appendix G. The raw count of&nbsp; plant growth-promoting traits (PGPTs) annotations, according to KEGG Orthology (KO) predictions for MAGs. (EMS_7.xlsx)</p> <p>Appendix H.&nbsp; The raw count of plant growth-promoting traits (PGPTs) that comprises the 39 classes (level 5 hierarchy) identified as enriched according to the results of&nbsp; Pearson's Chi-square test (qvalue &le; 0.1). (EMS_8.xlsx)</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Dataset for "Bacterial genome assembly" workflow

<p>This dataset is associated with the workflow "Bacterial genome assembly for paired end data".</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Differences in the genomic potential of soil bacterial and phage communities between urban greenspaces and natural arid soils.

<p>This repository holds the final data products from metagenomics processing of bacteria and viruses from the article : "Differences in the genomic potential of soil bacterial and phage communities between urban greenspaces and natural arid soils"</p> <p>Contents:&nbsp;</p> <ul> <li>LU_metadata.csv: information on the samples</li> <li>soil_chemistry.txt: physicochemical information on samples</li> <li>*_len.csv: tables containing the length information for annotated genes, divided by database. These are used to calculate RPKM abundances from count tables.&nbsp;</li> <li>BACTERIA</li> <li>ko_table, ko_unknown, ko2level, ko_description: count table of KEGG annotations, total counts for unnanotated genes, match of ko number to level and description</li> <li>all_bracken.csv: count table of taxonomic bacterial annotations using kraken2 and bracken</li> <li>mags_tax.csv: taxonomy assignments to MAGs (metagenome assembled genomes)</li> <li>mags_count_table.csv: abundante table of MAGs in counts</li> <li>ags_result, gc_mean, gc_variance: functional traits results, average genome size, and gc content</li> <li>lu_c_count, lu_n_count, card_d0, metals_count_table: abundance tables of genes annotated with Cazy (carbon), Ncydb (nitrogen), CARD (antibiotic resistance genes), and Bacmet (heavy metal resistance genes)</li> <li>VIRUS</li> <li>amg_summary.csv: results from AMG annotation with DRAM-V, filtered to keep genes of interest</li> <li>genomad_virus_summary.tsv: viral taxonomy annotations with geNomad</li> <li>virus_len.txt: length of inferred viruses (used for calculation of RPKM from count tables)</li> <li>all_host_prediction_to_genus.csv: virus host annotation with IPhop</li> <li>final_checkv.tsv: table of final viral inferences with quality estimates</li> <li>viral_species_count_table.txt: abundance table of infered viral contigs in counts</li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Dataset for "Metagenome-Assembled Bacterial Genomes from Long Accurate Reads Associated with Capilliphycus salinus ALCB114379"

<p>We present the raw genomic dataset of the marine cyanobacterium Capilliphycus salinus ALCB114379, sequenced using PacBio HiFi long-read technology.</p> <p><strong>ABSTRACT</strong></p> <p><span>We report the complete genome sequences of five bacteria linked to the marine cyanobacterium <em>Capilliphycus salinus</em> ALCB114379 from the phylum Pseudomonadota. This genetic diversity offers new insights into the genetic landscape and potential symbiotic relationships of cyanobacteria-associated microbiota.</span></p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Mash Sketch of RefSeq Bacterial Representative Genomes v215

<p>Prepared as described in&nbsp;https://github.com/UPHL-BioNGS/DesolationCanyon.</p> <p>&nbsp;</p> <p>This is the representative genomes of refseq 215 with their taxon name.</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Annotated sequences extracted from bacterial genomes

<p>Three files containing sequences extracted from 1,049,210&nbsp;bacterial genomes available from GenBank (release 252). Protein coding sequences were annotated with IDTAXA (PMID:&nbsp;34541527) using taxon-specific KEGG groups (Bacteria_Protein_subset.fas.gz). These annotations were transferred to their corresponding (nucleotide) coding sequences (Bacteria_Nucleotide_subset.fas.gz). Intergenic regions were extracted from each genome and annotated by FindNonCoding (PMID:&nbsp;34636849) for their overlap with any&nbsp;of 25 common bacterial non-coding RNAs in Rfam (v14). Intergenic regions were required to be at least 100 nucleotides long and contain no ambiguities (Bacteria_Intergenic_subset.fas.gz). Each subset contains only distinct sequences&nbsp;randomly ordered.</p> <p><strong><em>Headers</em></strong></p> <p>Sequence headers contain the assembly accession followed by the annotation and separated by a &quot;|&quot; character. For example:</p> <p><strong>Bacteria_Intergenic_subset.fas.gz</strong></p> <p>&gt;GCA_022121725.1|RF00000<br> ATGTTACCTTCTTGAGTGATACGGGATGAA[...]</p> <p><strong>Bacteria_Protein_subset.fas.gz</strong></p> <p>&gt;GCA_014764685.1|K02049<br> MPRDLIRISGLEKTYADGSVHALSNIDLSIKD[...]</p> <p><strong>Bacteria_Nucleotide_subset.fas.gz</strong></p> <p>&gt;GCA_015948525.1|K02197<br> GTGAACCTGCGACGTAAAAACCGGCTAYG[...]</p> <p><strong><em>Annotations</em></strong></p> <p>Protein and protein coding sequences are labeled with their KEGG group, starting with &quot;K&quot;. Intergenic sequences are named by any overlapping Rfam families, starting with a &quot;RF&quot;, and separated by commas when multiple are predicted. &quot;RF00000&quot; is a placeholder for the absence of any predicted RF families.</p>

opencc-by-4.0May 2023View details →
dryad32/100

Data from: Analysis of bacterial genomes from an evolution experiment with horizontal gene transfer shows that recombination can sometimes overwhelm selection

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad32/100

Data from: A genomic perspective on a new bacterial genus and species from the Alcaligenaceae family, Basilea psittacipulmonis

Open the record for dataset details and reuse information.

publicFeb 2015View details →
dryad32/100

Data from: Host‐derived population genomics data provides insights into bacterial and diatom composition of the killer whale skin

Open the record for dataset details and reuse information.

publicNov 2018View details →
dryad32/100

Genome reduction is associated with bacterial pathogenicity across different scales of temporal and ecological divergence - between species core gene alignments

Open the record for dataset details and reuse information.

publicNov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record