Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
52
datasets available to search
ShareScore release 0.9.0
Dataset results
52 results for “Microbial control”
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
Deep reinforcement learning for the control of microbial co-cultures in bioreactors
<p>Data for the figures in the paper:<br> <a href="https://www.biorxiv.org/content/10.1101/457366v2">https://www.biorxiv.org/content/10.1101/457366v2</a><br> (in press PLoS Comp Biol.)</p> <p>Abstract:<br> Multi-species microbial communities are widespread in natural ecosystems. When employed for biomanufacturing, engineered synthetic communities have shown increased productivity in comparison with monocultures and allow for the reduction of metabolic load by compartmentalising bioprocesses between multiple sub-populations. Despite these benefits, co-cultures are rarely used in practice because control over the constituent species of an assembled community has proven challenging. Here we demonstrate, in silico, the efficacy of an approach from artificial intelligence – reinforcement learning – for the control of co-cultures within continuous bioreactors. We confirm that feedback via reinforcement learning can be used to maintain populations at target levels, and that model-free performance with bang-bang control can outperform a traditional proportional integral controller with continuous control, when faced with infrequent sampling. Further, we demonstrate that a satisfactory control policy can be learned in one twenty-four hour experiment by running five bioreactors in parallel. Finally, we show that reinforcement learning can directly optimise the output of a co-culture bioprocess. Overall, reinforcement learning is a promising technique for the control of microbial communities.</p>
Fig. 1 in Long-term exposure of Aedes aegypti to Bacillus thuringiensis svar. israelensis did not involve altered susceptibility to this microbial larvicide or to other control agents
Fig. 1 Resistance ratios (RR) betseen the lethal concentrations of Bti and its toxins (Cru11Aa, Cru4Ba), temephos (Tem) and diflubenzuron (Dif) for third-instar Ae. aegypti larvae from the RecBti strain compared to that of the reference strain. a RR at LC50. b RR at LC90
Figure 5 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 5. Bacterioplankton average annual (2010 - 2011) abundance (N), cell volume (V), biomass (B), intracellular nucleic acids (FL1) and integral metabolic activity (FL1 × N) (± 95% CI) at St. 1 (grey) and St. 2 (black). Significant differences are marked (* p <0.05, ** p <0.01).
Figure 7 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 7. Fraction of live organisms (FLO) as a function of the decomposition-to-mortality ratio (d/m) in the model under steady-state conditions (mortality and specific growth rates are balanced, µ = m) and projections of natural zooplankton communities (St. 1 and 2) onto the model curve.
Figure 1 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 1. Fluorescein diacetate- (FDA) and neutral red (NR) -based estimates of the average annual FLO in the open coastal waters (St. 1 in this study) and the polluted bay (St. 2 in this study) in 2010 – 2011. Calculated from the data presented in Litvinyuk et al. (2011).
Figure 6 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 6. Fraction of live organisms (FLO) in zooplankton versus bacterioplankton abundance (N). Data on FLO (2010-2011) are from Litvinyuk et al. (2011).
Figure 4 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 4. Initial bacterial abundances in the experiment (No, left plot) and frequency distribution of the copepod decomposition stages on the fourth day of exposition at St. 1 and 2 (right plot). Means and standard deviations are presented.
Distribution System Environmental and Sequencing Datasets for Assessing the Impacts of Lead Corrosion Control on the Microbial Ecology and Abundance of Drinking Water Associated Pathogens in a Full-Scale Drinking Water Distribution System
<p>The dataset of environmental parameters and sequence fastqs used to create figures and do analysis in the paper <strong>Assessing the Impacts of Lead Corrosion Control on the Microbial Ecology and Abundance of Drinking Water Associated Pathogens in a Full-Scale Drinking Water Distribution System </strong>submitted to Environmental Science & Technology</p>
Urban Stream Environmental and Sequencing Datasets for Exploring the Impacts of Full-Scale Distribution System Orthophosphate Corrosion Control Implementation on the Microbial Ecology of Hydrologically Connected Urban Streams
<p>The dataset of environmental parameters and sequence fastqs used to create figures and do analysis in the paper <strong>Exploring the Impacts of Full-Scale Distribution System Orthophosphate Corrosion Control Implementation on the Microbial Ecology of Hydrologically Connected Urban Streams </strong>submitted to Applied and Environmental Microbiology. </p>
Figure 3 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 3. Stages of decomposition of copepod carcasses. Explanations are in the text.
Figure 2 in Microbial control of live/dead zooplankton ratio in Sevastopol Bay
Figure 2. Sampling sites in Sevastopol Bay and adjacent coastal waters.
Imaging Data for: Enabling oxygen-controlled microfluidic cultures for spatiotemporal microbial single-cell analysis
<p>This dataset contains the microfluidic microscopy time-lapse data for the publication <a href="https://www.frontiersin.org/articles/10.3389/fmicb.2023.1198170/">"Enabling oxygen-controlled microfluidic cultures for spatiotemporal microbial single-cell analysis"</a>.</p> <p>The imaging data is recorded as raw 16-bit tif-stacks. The sequences <span>17406 - 17410 and 17411 - </span><span>17415 contain the aerobic and anaerobic conditions, respectively. For details about cultivation conditions and image processing please have a look into our <a href="http://www.frontiersin.org/articles/10.3389/fmicb.2023.1198170/">paper</a> or the public code repository <a href="https://github.com/JuBiotech/Supplement-to-Kasahara-et-al.-2023a">Supplement-to-Kasahara-et-al.-2023a</a>.</span></p>
Spatial controls on eco-evolutionary processes in microbial communities
<div> <p><span>Microorganisms are the most biodiverse life forms on our planet, yet we know little about the spatial processes underlying their ecology and evolution. Here, we highlight the importance of two spatial processes that act on individual cells – spatial intermixing of different populations and mechanical cell shoving during growth – to improve our understanding of microbial eco-evolutionary dynamics. Using an individual-based model, we show that the coexistence between slow- and fast-growing populations becomes highly constrained under two conditions: when the slow- and fast-growing populations are highly spatially intermixed and when the ability to shove other cells (both conspecific and heterospecific) is weak. The potential for evolution through plasmid-mediated horizontal gene transfer between slow- and fast- growing populations also becomes restricted in the same scenario. Our modeling highlights that ecological constraints can dampen evolutionary opportunities within microbial communities due to variation in spatial intermixing and mechanical shoving at the cellular scale.</span></p> </div>
Data from: Community composition and physiological plasticity control microbial carbon storage across natural and experimental soil fertility gradients
<p>Data associated with 'Community composition and physiological plasticity control microbial carbon storage across natural and experimental soil fertility gradients' by Butler, Manzoni, and Warren, published in The ISME Journal (accepted 28/9/2023).</p>
Prospective Evaluation of Immunological, Molecular-genetic, Image-based and Microbial Analyzes to Characterize Tumor Response and Control in Patients With Inoperable Stage III NSCLC Treated With Chemo
ClinicalTrials.gov study NCT05027165. IPD Sharing: NO. Countries: 1. Publications: 1.
Spatial controls on eco-evolutionary processes in microbial communities
Open the record for dataset details and reuse information.
Lithological controls on soil geochemistry regulate microbial carbon use efficiency and carbon storage
<p><span>The data supporting the findings of the study titled "Lithological controls on soil geochemistry regulate microbial carbon use efficiency and carbon storage". lithology mediates the effects of soil aggregates and minerals on microbial carbon use efficiency and microbial necromass stability. Furthermore, despite high mineral abundance reduced microbial carbon use efficiency, it enhanced microbial necromass stabilization through organo-mineral associations.</span></p>
Images of the work entitled "The spatial distribution of rhizosphere microbial activities under drought: water availability is more important than root-hair controlled exudation"
<p>These images are the images of zymography, <sup>14</sup>C imaging and neutron radiography of the work entitled "The spatial distribution of rhizosphere microbial activities under drought: water availability is more important than root-hair controlled exudation". Raw data on optimal water conditions were partially overlapping with the data of Bilyera et al., 2021, Soil Biology and Biochemistry, 162, 108426.</p>
dataset of the work entitled "The spatial distribution of rhizosphere microbial activities under drought: water availability is more important than root-hair controlled exudation"
<p>This is the dataset of the enzyme kinetics, and other biochemical properties obtained from zymography, <sup>14</sup>C images and water images of the work entitled "The spatial distribution of rhizosphere microbial activities under drought: water availability is more important than root-hair controlled exudation". </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.