Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,438
datasets available to search
ShareScore release 0.7.1
Dataset results
1,438 results for “bacteria”
Soil Bacteria and Archaea in Macrosystems Biodiversity Project at Harvard Forest 2012
Patterns of biodiversity, such as the increase toward the tropics and the peaked curve during ecological succession, are fundamental phenomena for ecology. Such patterns have multiple, interacting causes, but temperature emerges as a dominant factor across organisms from microbes to trees and mammals, and across terrestrial, marine, and freshwater environments. However, there is little consensus on the underlying mechanisms, even as global temperatures increase and the need to predict their effects becomes more pressing. The purpose of this project is to generate and test theory for how temperature impacts biodiversity through its effect on biochemical processes and metabolic rate. A combination of standardized surveys in the field and controlled experiments in the field and laboratory measure diversity of three taxa -- trees, invertebrates, and microbes -- and key biogeochemical processes of decomposition in seven forests distributed along a geographic gradient of increasing temperature from cold temperate to warm tropical. This field experiment focused on soil microbes. DNA was extracted and purified from soil cores from an array of 21 1m2 subplots. The V4 region of the 16S rRNA genes for bacteria and archaea were amplified and sequenced using Illumina MiSeq by the University of Oklahoma Institute for Environmental Genomics as part of a macrosystems biodiversity and latitude project supported by the National Science Foundation under Cooperative Agreement DEB#1065836.
Number of bacteria in the water column of lakes sampled near Toolik Lake LTER Alaska, throughout summer season, 1992-2000.
Number of bacteria (using of DAPI for identifying and counting) in Toolik Lake water column and other lakes sample throught the summer from 1992-2000. There is no data for 2006.
Abundance, biovolume, and biomass of Synechococcus, eukaryote pico- and nano- phytoplankton, and heterotrophic bacteria from flow cytometry for water column bottle samples on NES-LTER Transect cruises, ongoing since 2018
These data represent the abundance, biovolume, and biomass of prokaryotic phytoplankton, eukaryotic pico- and nano- phytoplankton, and heterotrophic bacteria from discrete flow cytometry samples collected during the Northeast U.S. Shelf Long-Term Ecological Research (NES-LTER) Transect cruises, ongoing since 2018. Samples were collected and preserved from the water column at multiple depths using Niskin bottles on a CTD rosette system along the NES-LTER transect, and analyzed post cruise. Cells were identified and enumerated from the flow cytometry data files based on their scattering, SYBR (525 nm), phycoerythrin (575 nm) and chlorophyll (680 nm) fluorescence signals. Gating was completed manually in the Attune NXT software interface.
Human intestinal Bacteria Collection (HiBC): Isolates and genomes metadata
<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the taxonomy of the isolates, as well as metadata regarding their cultivation and isolation. We also provide metadata regarding the sequencing, genome assembly process and the biological sequences.</p> <p><strong>UPDATE v7</strong>: INSDC accession for <em>Segatella sinensis</em> CLA-AA-H117 was a missing value and is now the correct value of GCA_040324585.2.</p> <p><strong>UPDATE v6: </strong>The growth atmosphere is now indicated by anaerobic or aerobic instead of "Anaerobe/Aerobe" that was a misleading term. The risk group of these two isolates went from 1 to 2:</p> <ul> <li>CLA-AA-H205: <em>Anaerostipes caccae </em></li> <li>CLA-AA-H83: <em>Bacteroides fragilis</em></li> </ul> <p>The risk group of the following isolates has been updated (usually from unknown to 1, or from 2 to 1):</p> <ul> <li>CLA-SR-H026: <em>Aedoeadaptatus acetigenes</em></li> <li>CLA-KB-H139:<em> Bacteroides xylanisolvens</em></li> <li>CLA-SR-H015: <em>Bacteroides xylanisolvens</em></li> <li>CLA-AA-H187: <em>Blautia fusiformis</em></li> <li>CLA-AA-H274: <em>Brotaphodocola catenula</em></li> <li>CLA-AA-H286: <em>Butyricimonas faecihominis</em></li> <li>CLA-AA-H278:<em> Clostridium fessum</em></li> <li>CLA-AA-H147: <em>Dorea ammoniilytica</em></li> <li>CLA-SR-H027: D<em>orea formicigenerans</em></li> <li>CLA-KB-H89: <em>Dorea longicatena</em></li> <li>CLA-KB-H94: <em>Dorea longicatena</em></li> <li>CLA-SR-H022: <em>Enterococcus lactis</em></li> <li>CLA-AA-H250: <em>Hominenteromicrobium mulieris</em></li> <li>CLA-AA-H232: H<em>ominilimicola fabiformis</em></li> <li>CLA-AA-H246: <em>Hominisplanchenecus faecis</em></li> <li>CLA-AA-H276:<em> Hominiventricola filiformis</em></li> <li>CLA-AA-H213:<em> Oliverpabstia intestinalis</em></li> <li>CLA-AA-H241: <em>Oliverpabstia intestinalis</em></li> <li>CLA-AA-H58: <em>Pilosibacter fragilis</em></li> <li>CLA-KB-H110: <em>Ruthenibacterium lactatiformans</em></li> <li>CLA-AA-H174: <em>Segatella sinensis</em></li> <li>CLA-AA-H2: <em>Veillonella parvula</em></li> <li>CLA-AA-H273: <em>Waltera acetigignens</em></li> </ul> <p>Typos in media list have been fixed. </p> <p><strong>UPDATE v5</strong>: The accessions number for the genomes on INSDC databases are added under the column Accession. Plus two typos in the risk group column have been corrected as follow:</p> <ul> <li>CLA-AA-H173: from Risk Group 4 (!) to 2 like the other strain of <em>Sutterella wadsworthensis</em></li> <li>CLA-AA-H198: from Risk Group 4 (!) to 1 like the other <em>Bifidobacterium </em>species.</li> </ul> <p><strong>UPDATE v4</strong>: Only the taxonomy of a couple of isolates has been changed, as follow:</p> <ul> <li>CLA-ER-H4: <em>Collinsella sp900547855</em> instead of <em>Collinsella sp900544645</em></li> <li>CLA-AA-H142: <em>Pilosibacter fragilis</em> (<em>f__Clostridiaceae</em>) instead of <em>Sakamotonia hominis gen. nov.</em> (<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H58: <em>Pilosibacter fragilis </em>(<em>f__Clostridiaceae</em>) instead of <em>Sakamotonia hominis gen. nov. </em>(<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H89B: <em>Lachnospira intestinalis sp. nov.</em> instead of <em>Lachnospira hominis sp. nov.</em></li> <li>CLA-JM-H10: <em>Lachnospira hominis sp. nov.</em> instead of <em>Lachnospira intestinalis sp. nov.</em></li> <li>CLA-JM-H7B: <em>Faecalibacterium taiwanense</em> instead of <em>Faecalibacterium faecis sp. nov.</em></li> <li>CLA-JM-H45: <em>Merdimmobilis hominis</em> instead of <em>Hominicola intestinalis gen. nov.</em></li> </ul> <p><strong>UPDATE v3</strong>: The genome of one of our isolate had been unfortunately swapped. This mistake has been now corrected on Zenodo and Coscine. The genome of <em>Segatella sinensis</em> CLA-AA-H117 should be considered correct with 103 contigs and 3 671 232 nt. Please note that the genome available at the NCBI is the correct one (GCA_040324585.2). Two typos regarding taxonomy have been corrected as well: <em>Maccoya intestinihominis</em> has been corrected to <em>Maccoyia intestinihominis</em> and <em>Faecousia faecis</em> to <em>Faecousia intestinalis</em>.</p>
Data for Decomposing Fomes fomentarius fruiting bodies, unlike fresh ones, represent a bacteria-rich habitat primarily driven by Arthropoda
<p>These files represent the data files necessary to run the R script analysis for the paper "Decomposing<i> Fomes fomentarius</i> fruiting bodies, unlike fresh ones, represent a bacteria-rich habitat primarily driven by Arthropoda."</p><p>A link to the paper and code will be provided once the paper has been published.</p>
We are the bacteria
<p>Video in English (see also the Swahili version), warns of the dangers of the development of Antimicrobial Resistance (AMR) as a result of taking unprescribed antibiotics in subtherapeutic courses. This was originally developed as an audio piece, but still images and cartoons have been added in post-production to create a video. </p><p>Audio and images produced as part of a Participatory Action Research (PAR) Workshop with young professionals in Mwanza, Tanzania to create public health messages on Antimicrobial Resistance in a post-COVID East Africa in June 2022. The project built upon data gathered in two international, interdisciplinary research projects (HATUA – 'Holistic Approaches to Understanding Antimicrobial Resistance in East Africa' and CARE – 'COVID-19 and Antimicrobial Resistance in East Africa – Impact and Response'), seeking to understand the wider medical and societal drivers of AMR in East Africa and identify possible interventions to curb the spread of AMR. The workshop ran for 9 days over a 3 week period and consisted of focus group style discussions with participants to explore issues surrounding AMR, antibiotic use, and public health messaging awareness in local communities (days 1-2), participant-led design of poster, radio, and video messages with feedback from the research team and introduction to filming/recording equipment (days 3-4), filming, shooting and recording materials within local settings in Mwanza with participants serving as actors, directors, and crew with guidance from research team (days 4-8) and a final in-person review and hands-on feedback of preliminary mock-ups of posters and videos (day 9). Participants have continued to collaborate via email and WhatsApp as materials were finalised. Final video production and editing was undertaken by the researchers. </p><p>Correspondence: mgk@st-andrews.ac.uk; kjf4@st-andrews.ac.uk</p><p> </p>
Genomic Epidemiology Dataset for Important Nosocomial Pathogenic Bacteria Acinetobacter baumannii
<p>The<strong> </strong>infections caused by various bacterial pathogens both in clinical and community settings represent a significant threat to public healthcare worldwide. The growing resistance to antimicrobial drugs acquired by bacterial species causing healthcare-associated infections has already become a life-threatening danger noticed by the World Health Organization. Several groups or lineages of bacterial isolates usually called 'the clones of high risk' often drive the spread of resistance within particular species. </p><p>Thus, it is vitally important to reveal and track the spread of such clones and the mechanisms by which they acquire antibiotic resistance and enhance their survival skills. Currently, the analysis of whole genome sequences for bacterial isolates of interest is increasingly used for these purposes, including epidemiological surveillance and developing of spread prevention measures. However, the availability and uniformity of the data derived from the genomic sequences often represents a bottleneck for such investigations. </p><p>In this dataset, we present the results of a comprehensive genomic epidemiology analysis of 17,546 genomes of a dangerous bacterial pathogen <i>Acinetobacter baumannii</i>. Important typing information including multilocus sequence typing (MLST)-based sequence types (STs), intrinsic<i> blaOXA-51-like</i> gene variants, capsular (KL) and oligosaccharide (OCL) types, CRISPR-Cas systems, and cgMLST profiles are presented, as well as the assignment of particular isolates to nine known international clones of high risk. The presence of antimicrobial resistance genes within the genomes is also reported. </p><p>These data will be useful for researchers in the field of <i>A. baumannii</i> genomic epidemiology, resistance analysis and prevention measure development.</p>
DADA2 formatted 16S rRNA gene sequences for both bacteria & archaea
<p><strong><em>This version is to stay up to date with the improvements and increase in 16S rRNA gene sequences (SSU) added to the GTDB release 220. Please read this post for the stats on the updates. </em></strong><strong><em>https://gtdb.ecogenomic.org/stats/r220 </em></strong><strong><em>.</em></strong><strong><em> </em></strong></p> <p><strong><em>There has been no change to the RDP-RefSeq reference database please use previous versions.</em></strong></p> <p><strong><em>If anyone has concerns with MAG extracted 16S rRNA gene contamination concerns, then I suggest that they contact the curators of GTDB themselves because it is outside of my role with these resources designed for DADA2 usage only. </em></strong></p> <p><strong><em>Another concern that was raised was the orientation of the DB sequences, to get past this problem please use the tryRC = TRUE argument in the assignTaxonomy command within DADA2, this will search your ASVs in the reverse complement as well. </em></strong></p> <p>The bacterial and archaeal 16S rRNA gene sequence databases were collated from various sources and formatted to use the "assignTaxonomy" command within the DADA2 pipeline. The data was converted to suite DADA2 format by Alishum Ali.</p> <ol> <li>Genome Taxonomy Database (GTDB): The new version of our dada2 formatted GTDB reference sequences now contains 58102 bacteria and 3672 archaea full 16S rRNA gene sequences. If you wonder why there are fewer species with 16S rRNA, that is because some metagenomics-assembled genomes (MAGs) lack the 16S gene and thus cannot be extracted. The database was downloaded from <a href="https://data.ace.uq.edu.au/public/gtdb/data/releases/release95/">https://data.ace.uq.edu.au/public/gtdb/data/releases/</a> on 24/10/2024. Please read the release notes and file descriptions. </li> </ol> <p>The formatting to DADA2 was done using simple awk bash scripts. The script takes as input a fasta file and a tab-delimited taxonomy file (slightly edited to remove special characters) and then it outputs a fasta file with all 7 taxonomy ranks separated by ";" as required for DADA2 compatibility. Additionally, we have concatenated the unique sequence GTDB ID to the species entry (but replaced the "." with an " _". We see this as an important QC step to highlight the issues/confidence associated with short-read taxonomy assignment at the finer rank levels.</p> <p>Also, this update includes two other files that you can use with the assignTaxonomy and addSpecies commands in DADA2.</p>
SBC LTER: Cross-shelf Study 2008-2009: Profiles of CTD, biogeochemistry, primary production, abundance of phytoplankton groups, and abundance and production of bacteria
These data were used by the following papers: Goodman, J., M. A. Brzezinski, E. R. Halewood and C. A. Carlson. 2012. Sources of phytoplankton to the inner continental shelf in the Santa Barbara Channel inferred from cross-shelf gradients in biological, physical and chemical parameters. Continental Shelf Research, 48: 27-39. (DOI: 10.1016/j.csr.2012.08.011) Halewood, E. R., C. A. Carlson, M. A. Brzezinski, D. C. Reed and J. Goodman. 2012. Annual cycle of organic matter partitioning and its availability to bacteria across the Santa Barbara Channel continental shelf. Aquatic Microbial Ecology, 67:189-209. (DOI:10.3354/ame01586) These data were collected on monthly day cruises from January 2008 to April 2009 on the RV Kelp Fish at five stations across the shelf starting at Mohawk Reef in the nearshore area of the Santa Barbara Channel, California, USA. Data were collected with a SBE19-Plus and rosette sampler. Measurements include standard CTD parameters in 1 m bins (e.g. salinity, temperature, density). At selected depths (1, 5, 10, 20 m), rosette bottle samples were collected for nutrients, pigments, particulate and dissolved organic carbon and nitrogen, and bacterial abundance, community structure and productivity. Phytoplankton abundances (to genus) were obtained from the 5 m sample only.
Dataset - Characterization of Kazachstania humilis and Lactic Acid Bacteria interactions in French sourdoughs
<p>Here you can find the dataset and Rmarkdonw script associated to the scientific paper : Dataset - Characterization of Kazachstania humilis and Lactic Acid Bacteria interactions in French sourdoughs</p>
Coevolving plasmids drive gene flow and genome plasticity in host-associated intracellular bacteria
<p>Comparative genomics and modeling of plasmids of the obligate host-associated intracellular phylum chlamydiae. </p>
Source data for publication "Bacteria use exogenous peptidoglycan as a danger signal to trigger biofilm formation"
<p><strong>This dataset contains the source data for the figures in the following publication: </strong></p> <p><strong>Title: </strong>Bacteria use exogenous peptidoglycan as a danger signal to trigger biofilm formation</p> <p><strong>Authors: </strong>Sanika Vaidya, Dibya Saha, Daniel K.H. Rode, Gabriel Torrens, Mads F. Hansen, Praveen K. Singh, Eric Jelli, Kazuki Nosho, Hannah Jeckel, Stephan Göttig, Felipe Cava, Knut Drescher</p> <p><strong>Journal: </strong>Nature Microbiology, 2025</p> <p> </p> <p><strong>Description of the dataset: </strong></p> <p>The data is organized by figures in the publication receferenced above. For each figure, there is a XLSX-file that contains the processed data and a ZIP-archive that contains the raw image data. The ZIP-archive also contains readme documents with more detailed descriptions for every set of raw image data. </p> <p>Example: For Figure 1 in the main text of the publication, there following data are available:</p> <ul> <li>raw data: Figure_01.zip</li> <li>processed data: Processed_Source_Data_Figure1.xlsx</li> </ul> <p>Similarly, XLSX-files and ZIP-archives are available for the main text Figures 1-6. For the Extended Data (ED) Figures 1-10, there are XLSX-files available that present the data shown in each figure. Only some of the Extended Data Figures present results based on image data - therefore raw image data ZIP-archives are only available for ED Figures 3, 6, 7, 8, 9, 10.</p>
Genomic Typing, Antimicrobial Resistance Gene, Virulence Factor and Plasmid Replicon Dataset for the Important Pathogenic Bacteria Klebsiella pneumoniae
<p>The infections caused by various bacterial pathogens both in clinical and community settings represent a significant threat to public healthcare worldwide. The growing resistance to antimicrobial drugs acquired by bacterial species causing healthcare-associated infections has already become a life-threatening danger noticed by the World Health Organization. Several groups or lineages of bacterial isolates usually called 'the clones of high risk' often drive the spread of resistance within particular species. </p> <p>Thus, it is vitally important to reveal and track the spread of such clones and the mechanisms by which they acquire antibiotic resistance and enhance their survival skills. Currently, the analysis of whole genome sequences for bacterial isolates of interest is increasingly used for these purposes, including epidemiological surveillance and developing of spread prevention measures. However, the availability and uniformity of the data derived from the genomic sequences often represents a bottleneck for such investigations. </p> <p>In this dataset, we present the results of a genomic epidemiology analysis of 61,857 genomes of a dangerous bacterial pathogen <em>Klebsiella pneumoniae</em> obtained from NCBI Genbank database. Important typing information including multilocus sequence typing (MLST)-based sequence types (STs), capsular (KL) and oligosaccharide (OL) types, CRISPR-Cas systems, and cgMLST profiles are presented, as well as the assignment of particular isolates to clonal groups (CG). The presence of antimicrobial resistance and virulence genes, as well as plasmid replicons, within the genomes is also reported. </p> <p>These data will be useful for researchers in the field of <em>K. pneumoniae</em> genomic epidemiology, resistance analysis and prevention measure development.</p>
Rafael et al., 2018: Deep Learning-Based Culture-Free Bacteria Detection in Urine Using Large-Volume Microscopy (Dataset)
<p>Dataset containing 1um polystyrene beads, urine samples, urine samples mixed with ecoli and homogenous ecoli. Dataset is the post-processing version of the images to remove static background. Original model was trained on the post-processed images exclusiviely.</p>
Soil Bacteria Community-Weighted rrn Operon Copy Number Estimation
<p>Datasets and R-Scripts for estimating community-weighted rrn operon copy number for soil bacteria communities collected from the Yukon-Kuskokwim River Delta, AK, USA, and from La Selva Biological Station, Costa Rica. File descriptions follow:</p> <p>"rrnDB_copy_number_database.csv": The Ribosomal RNA Database downloaded from <a href="rrndb.umms.med.umich.edu.">rrndb.umms.med.umich.edu.</a> Citation: </p> <ul> <li>Stoddard S.F, Smith B.J., Hein R., Roller B.R.K. and Schmidt T.M. (2015) <em>rrn</em>DB: improved tools for interpreting rRNA gene abundance in bacteria and archaea and a new foundation for future development. <em>Nucleic Acids Research</em> 2014; doi: 10.1093/nar/gku1201 [<a href="http://www.ncbi.nlm.nih.gov/pubmed/25414355">PMID:25414355</a></li> </ul> <p>"AK_16S_Genus_Abundance.csv": Count of ASVs by taxon (assigned to genus level) present in each soil sample collected in the Yukon_Kuskokwim River Delta, AK, USA.</p> <p>"Costa_Rica_16S_OTU_Abundance": Count of OTUs by taxon present in each soil sample collected in La Selva Biological Station, Costa Rica.</p> <p>"Alaska_rrn_copy_number_estimation_script.R": an R script for processing Alaska ASV count table and estimating community-weighted rrn operon copy numbers for each soil sample.</p> <p>"CostaRica_rrn_copy_number_estimation_script.R": an R script for processing Costa Rica OTU count table and estimating community-weighted rrn operon copy numbers for each soil sample.</p>
Human intestinal Bacteria Collection (HiBC): 16S rRNA gene sequences
<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the sequences of the 16S rRNA gene sequences of the isolates in the FASTA nucleotide format. Sequences ending in Sanger were obtained using the Sanger dideoxy sequencing technology. Sequences ending in Genome were obtained from the genome sequence using barrnap.</p>
Data and code from: "Building multidimensional tolerance landscapes to predict the population dynamics of bacteria exposed to antibiotics in urban sewers"
<p>City sewers harbor diverse bacterial communities exposed to various antibiotic residues resulting from human consumption and excretion. Although these residues typically occur at sub-inhibitory concentrations, they can still impact the growth rate and yield of susceptible wastewater bacteria. Many bacteria exhibit antibiotic tolerance through transient phenotypic changes. Antibiotic residues, combined with complex environmental factors like temperature and salinity, especially in coastal cities, contribute to non-additive interactions that modulate antibiotic tolerance and affect population dynamics.</p> <p>To better understand these interactions, we developed continuous multivariate tolerance landscapes for three bacterial species: <strong><em><span>Escherichia coli</span></em></strong>, the emerging pathogen <strong><em><span>Streptococcus suis</span></em></strong>, and the sewer-inhabiting <strong><em><span>Arcobacter cryaerophilus</span></em></strong>. We modeled their intrinsic growth rates and carrying capacities across complex environments, incorporating temperature, salinity, and concentrations of two antibiotics (ciprofloxacin and azithromycin).<span> Using</span> these multivariate tolerance curves, we predicted microbial population dynamics in two sewers of Barcelona, highlighting the importance of environmental complexity in shaping microbial responses to antibiotic stressors.</p> <p> </p> <p><strong>Usage</strong></p> <p>Users can perform the analysis by running the R script (TC3D.R) after the installation of all</p> <p>package mentioned in the preamble,<span> </span></p> <p>This folder contains:</p> <p>- 3 datasets with OD measures for the 3 species:</p> <p><span> </span>* data_acrya.xlsx</p> <p><span> </span>* data_ecoli.xlsx</p> <p><span> </span>* data_ssuis.xlsx</p> <p>- 1 excel files with metadata (plate, well, species, environmental conditions)</p> <p><span> </span>* map_plate_all.xlsx</p> <p>- 4 datasets giving time series of the flow and several measures including <span> </span>conductivity and <span> </span>temperaturefor 2 sewers of Barcelona obtained from sample cabines <span> </span>set during the implementation of SCOREWATER (ID:820751)</p> <p><span> </span>* carmel_flow.csv</p> <p><span> </span>* carmel_quality.csv</p> <p><span> </span>* poblenou_flow.csv</p> <p><span> </span>* poblenou_quality.csv</p> <p>- 1 C++ script compiled and run with the R TMB package:</p> <p><span> </span>* fit_growth_r_K_SS_treatment.cpp : computes the negative loglikelihood for r and K, and state DOs, given the observed DO, for the populations under one same environmental treatment (salinity * temperature * antibiotic), and computes the density-dependence parameter alpha from r and K using the Delta Method.</p> <p><br><br></p>
gapseq reference sequence databases for Bacteria and Archaea
<p>The repository contains the protein sequences used by <a href="https://github.com/jotech/gapseq">gapseq</a> to predict the presence of metabolic reactions and to construct metabolic models.</p> <p>The workflow using gapseq to generate this set of reference protein sequences:</p> <p> </p> <p>```sh</p> <p># delete all "old" data<br>rm dat/seq/Bacteria/rev/*.fasta<br>rm dat/seq/Bacteria/unrev/*.fasta<br>rm dat/seq/Bacteria/rxn/*.fasta<br>rm dat/seq/Archaea/rev/*.fasta<br>rm dat/seq/Archaea/unrev/*.fasta<br>rm dat/seq/Archaea/rxn/*.fasta</p> <p># run gapseq find to re-download everything#<br># the genome is irrelevant as no blasting is performed ('-x')<br>gapseq find -p all -t Bacteria -n -x -U toy/ecoli.faa.gz > bac_update.log 2>&1<br>gapseq find -p all -t Archaea -n -x -U toy/ecoli.faa.gz > ar_update.log 2>&1</p> <p># create all sequence .tar.gz archives (rev/unrev/rxn)<br>cd dat/seq/Bacteria/rev/ && tar -czvf sequences.tar.gz ./*.fasta && cd ../../../../<br>cd dat/seq/Bacteria/unrev/ && tar -czvf sequences.tar.gz ./*.fasta && cd ../../../../<br>cd dat/seq/Bacteria/rxn/ && tar -czvf sequences.tar.gz ./*.fasta && cd ../../../../<br>cd dat/seq/Archaea/rev/ && tar -czvf sequences.tar.gz ./*.fasta && cd ../../../../<br>cd dat/seq/Archaea/unrev/ && tar -czvf sequences.tar.gz ./*.fasta && cd ../../../../<br>cd dat/seq/Archaea/rxn/ && tar -czvf sequences.tar.gz ./*.fasta && cd ../../../../</p> <p># create md5sum table for all tar.gz archives<br>cd dat/seq/<br>find -mindepth 2 -type f -name "*.tar.gz" -exec md5sum {} \; > md5sums.txt</p> <p># create taxon-specific final archive for Zenodo upload<br>tar -czvf Bacteria.tar.gz Bacteria/*/*.tar.gz<br>tar -czvf Archaea.tar.gz Archaea/*/*.tar.gz</p> <p># Upload Bacteria.tar.gz, Archaea.tar.gz, and md5sums.txt to Zenodo via the web-interface</p> <p>```</p>
Effects of nutrients and organic carbon on the relative proportion of primary producers (microalgae) and heterotrophic decomposers (bacteria and fungi) during aquatic biofilm development in boreal peatland located near Fairbanks Alaska - 2018
1. Producer-decomposer interactions within aquatic biofilms can range from mutualistic associations to competition depending on available resources. The outcomes of such interactions have implications for biogeochemical cycling, and as such, may be especially important in northern peatlands, which are a global carbon sink and are expected to experience changes in resource availability with climate change. The purpose of this study was to evaluate the effects of nutrients and organic carbon on the relative proportion of primary producers (microalgae) and heterotrophic decomposers (bacteria and fungi) during aquatic biofilm development in a boreal peatland. Given that decomposers are often better competitors for nutrients than primary producers in aquatic ecosystems, we predicted that labile carbon subsidies would shift the biofilm composition towards heterotrophy owing to the ability of decomposers to outcompete primary producers for available nutrients in the absence of carbon limitation. 2. We manipulated nutrients (nitrate and phosphate) and organic carbon (glucose) in a full factorial design using nutrient-diffusing substrates in an Alaskan fen. 3. Heterotrophic bacteria were limited by organic carbon and algae were limited by inorganic nutrients. However, the outcomes of competitive interactions depended on background nutrient levels. Heterotrophic bacteria were able to outcompete algae for available nutrients when organic carbon was elevated and nutrient levels remained low, but not when organic carbon and nutrients were both elevated through enrichment. 4. Fungal biomass was significantly lower in the presence of glucose alone, possibly owing to antagonistic interactions with heterotrophic bacteria. In contrast to bacteria, fungi were stimulated along with algae following nutrient enrichment. 5. The decoupling of algae and heterotrophic bacteria in the presence of glucose alone shifted the biofilm trophic status towards heterotrophy. This effect was overturned
Picophytoplankton and bacteria abundances analyzed with flow cytometry (FCM) from CCE-CalCOFI Augmented cruises in the California Current System, 2004 - 2023 (ongoing).
Picophytoplankton populations and non-pigmented prokaryotes are sampled within the California Current Ecosystem (CCE) for abundances from 3 to 8 depths at CalCOFI stations. Seawater is collected from Niskin bottles and cells are fixed in the field aboard the survey cruises (since 2004, ongoing) with paraformaldehyde, and stained with a DNA-specific dye back in the laboratory. The cells are enumerated by an Altra flow cytometer (with a syringe pump for volumetric sample delivery) simultaniously with argon ion lasers, to distinguish three major populations of photoautotrophs (Prochlorococcus, Synechococcus, and pico-eukaryotes) and the assemblage of heterotrophic prokaryotes collectively referred to as H-Bact.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.