Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
871
datasets available to search
ShareScore release 0.7.1
Dataset results
871 results for “escherichia coli”
Slimfield: Escherichia coli DNA repair proteins (RecA-mGFP and RecB-sfGFP)
<p>Imaging modality / instrument: <em>Brightfield</em> + <em>Slimfield</em></p> <p>Image format:<em> OME TIFF (16 bit) + MicroManager metadata files</em></p> <p>Microscope settings:</p> <p><em>488 nm triggered excitation; split detection, cropped to GFP or RFP/GFP (left/right) channels; 3 ms/frame laser exposure; Photometrics Prime95b CMOS</em></p> <p>Samples and acquisitions:</p> <p>Fluorescent fusions in live E.coli cells. MMC = mitomycin C</p> <table> <tbody> <tr> <td> <p>No. fields of view</p> </td> <td> <p>MMC-</p> </td> <td> <p>MMC+ (0.5 ug/ml 3h)</p> </td> </tr> <tr> <td> <p>RecA-mGFP</p> </td> <td> <p>7</p> </td> <td> <p>15</p> </td> </tr> <tr> <td> <p>RecB-sfGFP</p> </td> <td> <p>17</p> </td> <td> <p>21</p> </td> </tr> <tr> <td> <p>MG1655 control</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> </tr> </tbody> </table> <p>Approx. size before/after compression: 21 GB / 7 GB</p>
Infection Inspection: Classifications and images of ciprofloxacin-treated Escherichia coli clinical isolates
<p>This dataset includes a .csv file with the image metadata and a folder of RGB images of <i>E. coli</i> grown from clinical isolates with varying concentrations of the antibiotic ciprofloxacin and varying minimum inhibitory concentrations. The <i>E. coli</i> cell membranes are stained with Nile Red and the DNA is stained with DAPI. The details of the image data collection are included in: https://doi.org/10.1038/s42003-023-05524-4. The classification data come from a Zooniverse citizen science project, Infection Inspection. (https://www.zooniverse.org/projects/conor-feehily/infection-inspection) Volunteers learned how to interpret ciprofloxacin response phenotypes as antibiotic-sensitive or antibiotic-resistant, and their classifications are included in the Metadata.csv file.</p><p>This dataset could be used for further analysis into the volunteer classifications, or the image data could be used for further image feature analysis of the ciprofloxacin response phenotypes.</p>
Bin-assembled Escherichia coli genomes from a study in Punjab, Pakistan
<h2>Bin-assembled <em>Escherichia coli</em> genomes from Punjab, Pakistan</h2> <p>These assemblies are a part of a cross-sectional study conducted in Punjab, Pakistan aimed at investigating <em>E. coli</em> colonisation diversity in healthy carriage with the use of CLED enrichment plates.</p> <h3><strong>About</strong></h3> <h4><strong>Version history</strong></h4> <p><strong>v0.1.1 (current version)</strong></p> <ul> <li>Added reference to the study.</li> </ul> <p><strong>v0.1.0</strong></p> <ul> <li>Added brief description with a few missing parts.</li> </ul> <h4><strong>Distribution</strong></h4> <p>If you use these assemblies in your study please cite the source as appropriate. These assemblies are made available under a CC-BY 4.0 license.</p> <h4><strong>Citation</strong></h4> <p>Khawaja, T., Mäklin, T., Kallonen, T. et al. Deep sequencing of <em>Escherichia coli</em> exposes colonisation diversity and impact of antibiotics in Punjab, Pakistan. Nature Communications 15, 5196 (2024). <a href="https://doi.org/10.1038/s41467-024-49591-5">https://doi.org/10.1038/s41467-024-49591-5</a></p> <h3><strong>Methods briefly</strong></h3> <h4><strong>Species identification</strong></h4> <p>Sequencing data from the ENA project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB36642">PRJEB36642</a> was error-corrected with <a href="https://github.com/opengene/fastp">fastp</a> and pseudoaligned with <a href="https://github.com/algbio/themisto">Themisto</a> against a species-level index (available from <a href="https://doi.org/10.5281/zenodo.6656881">https://doi.org/10.5281/zenodo.6656881</a>). Reads were assigned to species using the <a href="https://doi.org/10.1099%2Fmgen.0.000691">mSWEEP/mGEMS pipeline</a> as described in <a href="https://www.nature.com/articles/s41467-022-35178-5">https://www.nature.com/articles/s41467-022-35178-5</a>.</p> <h4><strong>Lineage identification</strong></h4> <p>Read from the species-level bins were again pseudoaligned with Themisto against an <em>E. coli</em> index (will be made available in a later version). Lineage-level assignment was performed using mSWEEP and mGEMS at the level of <a href="https://genome.cshlp.org/content/29/2/304">PopPUNK</a> sequence clusters. The created bins were screened with <a href="https://github.com/tmaklin/coreutils_demix_check">demix_check</a> and bins that received a score of 1 or 2 were kept. Data in the kept bins were assembled with <a href="https://github.com/tseemann/shovill">shovill</a> and the bin-assembled genomes (BAGs) were quality controlled with <a href="https://genome.cshlp.org/content/25/7/1043">checkm</a> for >= 90% completeness and <= 10% contamination. Finally, BAGs shorter than 4 Mb or longer than 6 Mb were removed.</p> <h3><strong>Contact</strong></h3> <p>Tommi Mäklin <tommi'at'maklin.fi>.</p>
Data from an investigation into colibactin-producing Escherichia coli endemicity globally
<p>This upload is a part of the study "Geographical variation in the incidence of colorectal cancer and urinary tract cancer is associated with population exposure to colibactin-producing <em>Escherichia coli</em>" published in <em>Lancet Microbe</em> on 5 December 2024, doi: <a href="https://doi.org/10.1016/j.lanmic.2024.101015">10.1016/j.lanmic.2024.101015</a>.</p> <p><strong>Contents</strong></p> <p>Data and scripts from the "<em>Geographical variation in colorectal and urinary tract linked cancer incidence is associated with population exposure to colibactin-producing Escherichia coli</em>" study.</p> <p>For the <em>E. coli</em> assembly collection presented in the study, see the separate upload at <a href="../records/13374348">https://zenodo.org/records/13374348</a>.</p>
Deciphering polymorphism in 61,157 Escherichia coli genomes via epistatic sequence landscapes
<p>We use computational models based on Direct Coupling Analysis - DCA - trained on PFAM domains of distant distant homologues to accurately predict the polymorphisms segregating in a panel of 61,157 <em>Escherichia coli </em>genomes.</p> <p>We show that the genetic context (<em>i.e. </em>the rest of the protein sequence) strongly constrains the tolerable amino acids in 30% to 50% of amino-acid sites. Our study also suggests the gradual build-up of genetic context over long evolutionary timescales by the accumulation of small epistatic contributions.</p> <p>Please refer to the README file for additional information on the structure of this dataset.</p> <p>Code to analyse this dataset is available at https://github.com/GiancarloCroce/DCA_polymorphism_Ecoli.</p> <p> </p>
The self-cleaning properties of biomimetic surfaces to repel Escherichia coli and Listeria monocytogenes attachment, adhesion, and retention
<p>Surface hydrophobicity and roughness were determined for unmodified wax surfaces (control), biomimetic wax surfaces, and Gladioli leaves. The self-cleaning properties of the biomimetic and control surfaces were compared by measuring their propensity to repel <em>Escherichia coli</em> and <em>Listeria monocytogenes</em> attachment, adhesion, and retention in mono- and co-culture conditions.</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,999 Escherichia coli isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 7,601-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/5/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,999 <em>Escherichia coli </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 411 different serotypes are represented in this dataset, with O157:H7 being the most represented one, corresponding to 37.1% of the dataset.</p> <p>File “Ec_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Ec_profiles_wgMLST.tsv” corresponds to a tab separated file with the 7,601-loci wgMLST profiles of each isolate presented in the metadata file. The files “profiles/Ec_profiles_cgMLST_95.tsv”, “profiles/Ec_profiles_cgMLST_98.tsv” and “profiles/Ec_profiles_cgMLST_100.tsv” correspond to a 2,826-loci, 2,704-loci and 465-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>E. coli </em>genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 2,688 samples associated with three BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA230969">PRJNA230969</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB27020">PRJEB27020</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA248042">PRJNA248042</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,999 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 7,601-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/5">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 7,601-loci wgMLST profiles of the 1,999 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 2,826-loci, 2,704-loci and 465-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Assembly of enterohemorrhagic Escherichia coli type IV pilin PpdD and its variants
<p>Assembly of the EHEC major type IV pilin PpdD was analyzed in a reconstituted TP assembly system described in LunaRico et al. Mol Microbiol. 2019 Mar;111(3):732-749. doi: 10.1111/mmi.14188. Bacteria of strain BW25113 F'tet harboring plasmids pMS41 and pCHAP8565 (or its variants) were grown for 2 days at 30°C on M9 plates containing 0.5% glycerol, amplicillin (100 ug/ml) chloramphenicol (25 ug/ml) and 1 mM IPTG.</p> <p>Bacteria were collected and fractionated as described in Luna Rico et al Methods Mol Biol. 2018;1764:291-305. doi: 10.1007/978-1-4939-7759-8_18. Cell and sheared fractions were analysed by electrophoresis on 10 % Tris-Tricin gels, transferred on nitrocellulose and probed with anti-MalE-PpdD polyclonal antibodies. The fluorescence signal was developed with ECL2 (Thermo) and recorded with Typhoon FLA9000 imager (GE).</p> <p>The signal was quantified using ImageJ. The fractions of PpdD assembled into pili were quantified and analysed using Prism9.</p> <p>The images uploaded here are the raw data used to produce the Fig. 4B of the article Karami et al., Structure, 2021.</p>
Riboswitch-inspired toehold riboregulators for gene regulation in Escherichia coli
<p>This dataset comprises flow cytometry data accompanying a publication on the development of synthetic riboregulators.</p> <p>These riboregulators were inspired by the architecture of naturally occurring riboswitches and toehold-mediated strand displacement. Specifically, we adopt the toehold switch hairpin and inserted regulatory sequences within the loop region of which accessibility can be controlled by toehold-mediated strand displacement. We utilized this design principle to develop toehold translation repressor and toehold transcriptional repressor, which regulate mCherry expression in <em>E. coli </em>in translational and transcriptional levels with certain ON/OFF ratios. Furthermore, we combined these two riboregulators and developed them into a NOR gate switch that can regulate downstream GFP expression in <em>E. coli </em>with different input conditions of trigger RNA. We used flow cytometry to quantify the expression level of the NOR gate switch under different inputs.</p>
Synthetic Escherichia coli mixture samples with variable coverage
<p>This dataset contains the synthetic mixture samples and reference sequences - as well as the appropriate metadata - that were originally used in the 2021 revision of the mSWEEP manuscript.<br> <br> There are 87 samples in total, each containing 100bp paired-end Illumina sequencing reads from 10 different <em>Escherichia coli </em>strains from 10 different lineages. The number of reads is set so that the sequencing coverage of the individual strains varies between 50x and 0.10x and sums up to 100x.</p>
The OHEJP BeONE Project – Escherichia coli genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 308 <em>Escherichia coli</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120057">https://zenodo.org/record/7120057</a>), comprising genome assemblies of 1,999 <em>E. coli</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Ec_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Ec_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>E. coli</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57098">PRJEB57098</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 308 isolates passed the dataset curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Confocal microscopy images of Escherichia coli cells treated with various antibiotics (LB, exponential phase)
<p>This dataset and CARE model is part of the publication "<strong>Transertion and cell geometry organize the </strong><i><strong>Escherichia coli</strong></i><strong> nucleoid during rapid growth</strong>".</p><p>It contains all CLSM images that were used for the publication, as well as the single-cell regions of interest for analyses.</p><p>Cells were grown to exponential phase in LB Lennox and antibiotics were added for 0-60 min. Cultures were then chemically fixed, immobilized and stained for DNA (DAPI) and membrane (Nile Red). The strain (NO34) expresses a MreBsw-sfGFP fusion protein from the native chromosomal locus. It was a kind gift from Zemer Gitai (<a href="http://doi:10.1016/j.bpj.2016.07.017">Ouzounov et al., 2016</a>).</p><p>More information can be found in the publication.</p>
Multi-colour SMLM images of untreated and drug-treated Escherichia coli (LB, exponential phase)
<p>This dataset and CARE model is part of the publication "<strong>Transertion and cell geometry organize the <em>Escherichia coli</em> nucleoid during rapid growth</strong>".</p> <p>It contains all SMLM images that were used for the publication, as well as the single-cell regions of interest for analyses.</p> <p>Cells were grown to exponential phase in LB Lennox and antibiotics were added for 0-60 min. Cultures were then chemically fixed, permeabilised and imaged for the nucleoid (JF<sub>646</sub>-Hoechst) and membranes (Nile Red) using PAINT. The strain (NO34) expresses a MreB<sup>sw</sup>-sfGFP fusion protein from the native chromosomal locus. It was a kind gift from Zemer Gitai (<a href="http://doi:10.1016/j.bpj.2016.07.017">Ouzounov et al., 2016</a>).</p> <p>More information can be found in the publication.</p>
Cell growth dataset for "Suppression of bacterial cell death underlies the antagonistic interaction between ciprofloxacin and tetracycline in Escherichia coli"
<div>This dataset contains growth data of <em>E. coli</em> cells measured by optical density at a wavelenght of 600 nm (OD600) to determine the conditions where the combination of ciprofloxacin (CIP) and tetracycline (TET) is antagonistic (or suppressive), using a medium that supports fast, intermediate and slow growth (M9 based medium supplemented with: glucose and amino acid, glucose, and glycerol, respectively).</div> <div>Two types of assays were performed: a checkerboard assay with a 2-dimension gradient of antibiotic concentrations shown in "Fig-S1-Growth_rates-Checkerboard_assay.zip", and bulk growth rates assay shown in "Fig-S2-Bulk_doubling_rates-Bioreactor.zip". </div> <div>These experiments are presented in the supplementary figures of the following manuscript published as preprint in bioRxiv: https://doi.org/10.1101/2024.04.18.590101.</div> <div> </div> <p><strong>Fig-S1-Growth_rates-Checkerboard_assay.zip:</strong></p> <div>- growth-OD600-Blank_correct.xlsx: contains the blank corrected (subtracted by initial OD of sterile medium) OD600 data for each well measured in the checkerboard assay. The negative control is named as "Blank B" in the spreasheet and the antibiotic concentrations for each well is defined in column C (Content).</div> <div>- growth_data.xlsx: contains the processed data from "Growth-OD600-Blank_correct.xlsx":growth curves were smoothed with a 5-window moving median and outliers corrected using the filloutliers function in MATLAB. Outliers that were missed were manually corrected.</div> <div>- time_h.xlsx: contains the time in hours used by the matlab function to build the dose response figure S1B.</div> <div>- analysis_code_checkerboard_assay.m: matlab function used to build the dose response figure S1B.</div> <p> </p> <p><strong>Fig-S2-Bulk_doubling_rates-Bioreactor.zip:</strong></p> <p>- "bulk_doubling rates_gly/glu/gluaa.csv": contains the calculated doubling rates for each of three experimental replicates (rows) and for each antibiotic treatment (columns: Ctrl, CIP, TET, CIP-TET) as shown in Figure S2. These values were used to calculated the Bliss independence (for further details see https://gitlab.com/MEKlab/single-cell-suppression-2024/figure plotting.ipynb<br> - Folders containing OD600 measurements and calculated growth rates for each growth medium, which are further separated into folders for each experimental replicate. Each replicate folder is labelled as the date the experiment was performed (denoted as "*" hereon). Each contains the following: raw optical density data (*.txt files), summary of experiment and results (*.docx file), the function compute_growth_rates.m, and the script Growth_curves_*.m.</p> <div>For more information on methods, strains used and table header descriptions, please refer to the README.txt. </div>
Supplementary material Comparative Genomic Analysis of Antimicrobial-Resistant Escherichia coli from South American Camelids in Central Germany
<p>Supplementary material for publication González-Santamarina, B.; Weber, M.; Menge, C.; Berens, C. Comparative Genomic Analysis of Antimicrobial-Resistant <i>Escherichia coli</i> from South American Camelids in Central Germany. <i>Microorganisms</i> <strong>2022</strong>, <i>10</i>, 1697. https://doi.org/10.3390/microorganisms10091697 </p>
Escherichia coli MYb137
This is one of the Wormbiome database archive files.<br>This entry includes all the genome annotation files related to Escherichia coli MYb137, a\(n\) Gammaproteobacteria.<br>The Wormbiome collection is an online database dedicated to centralizing all the information related to bacteria associated with C. elegans. More information on <a href="https://bitbucket.org/the-samuel-lab/wbm_scripts/src/master/DOCS/Annotations_output.md" target="_blank" rel="noopener noreferrer">the documentation page</a>.<br><br>
Escherichia coli OP50
This is one of the Wormbiome database archive files.<br>This entry includes all the genome annotation files related to Escherichia coli OP50, a\(n\) Gammaproteobacteria.<br>The Wormbiome collection is an online database dedicated to centralizing all the information related to bacteria associated with C. elegans. More information on <a href="https://bitbucket.org/the-samuel-lab/wbm_scripts/src/master/DOCS/Annotations_output.md" target="_blank" rel="noopener noreferrer">the documentation page</a>.<br><br>
Escherichia coli MYb5
This is one of the Wormbiome database archive files.<br>This entry includes all the genome annotation files related to Escherichia coli MYb5, a\(n\) Gammaproteobacteria.<br>The Wormbiome collection is an online database dedicated to centralizing all the information related to bacteria associated with C. elegans. More information on <a href="https://bitbucket.org/the-samuel-lab/wbm_scripts/src/master/DOCS/Annotations_output.md" target="_blank" rel="noopener noreferrer">the documentation page</a>.<br><br>
DeepBacs – Escherichia coli nucleoid denoising dataset and CARE model
<p>Training and test images of H-NS-mScarlet-I expressing <em>E. coli </em>cells for image denoising, as well as a trained CARE model.</p> <p>Additional information can be found on our <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>The example images show confocal images of labelled <em>E. coli</em> nucleoids at low and high SNR.</p> <p> </p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired microscopy images (fluorescence)</p> <p><strong>Microscopy data type</strong>: Confocal fluorescence images</p> <p><strong>Microscope</strong>: Leica SP8 confocal microscope with a 1.40 NA 63x oil immersion objective </p> <p><strong>Cell type</strong>: <em>E. coli</em> strain CS01 expressing H-NS-mScarlet-I fusion protein (H-NS-mScarlet-I) in NO34 parental strain (MreB-sfGFPsw, kindly provided by Zemer Gitai) </p> <p><strong>File format</strong>: .tif (16-bit)</p> <p><strong>Image size</strong>: 512 x 512 px<sup>2</sup> (Pixel size: 45 nm)</p> <p> </p> <p><strong>CARE model</strong>:</p> <p>The CARE 2D model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained from scratch for 100 epochs (600 steps/epoch) on 1400 paired image patches (image dimensions: (512 x 512 px²), patch size: (64 x 64 px²), 50 patches/image) with a batch size of 8 and a laplace loss function, using the CARE 2D ZeroCostDL4Mic notebook (v 1). Key python packages used include tensorflow (v 0.1.12), Keras (v2.3.1), csbdeep (v 0.6.2), numpy (v 1.19.5), cuda (v 11.0.221). The training was accelerated using a Tesla T4 GPU and data was augmented by a factor of 4 using rotation and flipping.</p> <p>The model weights can be used with the ZeroCostDL4Mic CARE 2D notebook and the CSBDeep Fiji plugin.</p> <p> </p> <p><strong>Author(s)</strong>: Christoph Spahn<sup>1,2</sup>, Mike Heilemann<sup>1,3</sup></p> <p><strong>Contact email</strong>: christoph.spahn@mpi-marburg.mpg.de</p> <p> </p> <p><strong>Affiliation(s)</strong>: </p> <p>1) Institute of Physical and Theoretical Chemistry, Max-von-Laue Str. 7, Goethe-University Frankfurt, 60439 Frankfurt, Germany</p> <p>2) ORCID: 0000-0001-9886-2263 </p> <p>3) ORCID: 0000-0002-9821-3578</p>
DeepBacs – Escherichia coli antibiotic phenotyping object detection dataset and YOLOv2 model
<p>Training and test images of <em>E. coli</em> cells treated with different antibiotics for antibiotic phenotyping using YOLOv2 object detection.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>Example images show predictions of drug-treated <em>E. coli</em> cells.</p> <p> </p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired microscopy images (confocal fluorescence) and manual annotations</p> <p><strong>Microscopy data type</strong>: Confocal fluorescence images of fixed <em>E. coli</em> cells stained for membrane (Nile Red) and DNA (DAPI) paired with annotations in PASCAL VOC format</p> <p><strong>Microscope</strong>: Zeiss LSM710 confocal microscope with a Plan-Apo 63x oil objective (1.4 NA)</p> <p><strong>Cell type</strong>: Chemically fixed <em>E. coli</em> NO34 cells (MreB-sfGFPsw, kindly provided by Zemer Gitai) (untreated or drug-treated);</p> <p><strong>File format</strong>: .png (RGB)</p> <p><strong>Image size</strong>: 400 x 400 px² (Pixel size: 84 nm)</p> <p> </p> <p><strong>YOLOv2 model</strong></p> <p>The YOLOv2 model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained from scratch for 97 epochs on 153 manually annotated images (image dimensions: (400, 400, 3)) with a batch size of 16 and a custom loss function combining MSE and crossentropy losses, using the YOLOv2 ZeroCostDL4Mic notebook (v 1.12) (von Chamier & Laine et al., 2020). Key python packages used include tensorflow (v0.1.12), Keras (v 2.3.1), numpy (v 1.19.5), cuda (v 10.1.243). The training was accelerated using a Tesla P100GPU and data was augmented by a factor of 8 using rotation and flipping.</p> <p>The model weights can be used with the ZeroCostDL4Mic YOLOv2 notebook.</p> <p> </p> <p><strong>Author(s)</strong>: Christoph Spahn<sup>1,2</sup>, Mike Heilemann<sup>1,3</sup></p> <p><strong>Contact email</strong>: christoph.spahn@mpi-marburg.mpg.de</p> <p> </p> <p><strong>Affiliation(s)</strong>: </p> <p>1) Institute of Physical and Theoretical Chemistry, Max-von-Laue Str. 7, Goethe-University Frankfurt, 60439 Frankfurt, Germany</p> <p>2) ORCID: 0000-0001-9886-2263 </p> <p>3) ORCID: 0000-0002-9821-3578</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.