Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
340
datasets available to search
ShareScore release 0.9.0
Dataset results
340 results for “plasmid”
Genomes plasmids MDR B. fragilis ONT sequence read files in fastq format
<p>Supporting data for the manuscript <em>Complete genome assembly of clinical multidrug resistant Bacteroides fragilis isolates enables comprehensive identification of antimicrobial resistance genes and plasmids.</em></p> <p>Oxford Nanopore reads demultiplexed with <a href="https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=1&cad=rja&uact=8&ved=2ahUKEwjNjqL5tY7iAhUawMQBHZHfDasQFjAAegQIAhAB&url=https%3A%2F%2Fgithub.com%2Frrwick%2FDeepbinner&usg=AOvVaw0wikvIUagLuFV38CwKZtia">Deepbinner</a> v0.2.0 and base-called (with demultiplexing) using Albacore v2.3.3. Barcodes and adapters were removed with <a href="https://github.com/rrwick/Porechop">Porechop</a> v0.2.4 with the --discard_middle option.</p> <p>Data from each isolate was produced from two runs per isolate. Data for the individual runs are included here. They can easily be concatenated eg with cat. Runs are named TVS_01,. TVS_02, TVS_03 and TVS_04.</p> <p>Fast5 (only demultiplexed with deepbinner and basecalled with albacore) as well as illumina reads and genome assemblies can be found via the NCBI bioproject accessions:</p> <p>Isolates, NCBI bioproject accession no:</p> <p>CCUG4856T, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA525024">PRJNA525024</a></p> <p>BFO17, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244943">PRJNA244943</a></p> <p>BFO18, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244944">PRJNA244944</a></p> <p>S01, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244942">PRJNA244942</a></p> <p>BFO42, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA253771">PRJNA253771</a></p> <p>BFO67, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254401">PRJNA254401</a></p> <p>BFO85, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254455">PRJNA254455</a></p> <p> </p> <p><strong>md5sum's (also found in the file md5.md5):</strong></p> <p>135d0570a1e49e25c8fde59f321cca68 BFO17_TVS_03_99377.barcode02_trimmed.fastq.gz<br> 3c6ca800a1f735937c0cffccd263fb82 BFO18_TVS_01_97673.barcode03_trimmed.fastq.gz<br> dc37823950f529d3a859b7a7c514af8c BFO18_TVS_03_99377.barcode03_trimmed.fastq.gz<br> abd707404f9ebbc38652e45ed378ca21 BFO42_TVS_02.barcode10_trimmed.fastq.gz<br> 9761e9ab082e276624c5778dcbeaffd2 BFO42_TVS_04.barcode10_trimmed.fastq.gz<br> ffc27009c0f7fead1af84ce045dbe5f3 BFO67_TVS_02.barcode09_trimmed.fastq.gz<br> 06c363a7feeeb88f9d195b3769b37f2b BFO67_TVS_04.barcode09_trimmed.fastq.gz<br> 5553c95cc98f4b9d4cfb38c4f8f8f037 BFO85_TVS_02.barcode08_trimmed.fastq.gz<br> b1a8013cba7a079cee6d3bdd6cd97ff2 BFO85_TVS_04.barcode08_trimmed.fastq.gz<br> 8169225219a5fb20935d5f0304aa80c5 CCUG4856T_TVS_01_97673.barcode01_trimmed.fastq.gz<br> a509b0ae912a798a91e677795198c1c6 CCUG5846T_TVS_03_99377.barcode01_trimmed.fastq.gz<br> 8a9d2eb8b626a6e87ed267d31aa220e3 S01_TVS_01_97673.barcode04_trimmed.fastq.gz<br> 69e239a17becc25a4f103dfc3bf5886a S01_TVS_03_99377.barcode04_trimmed.fastq.gz</p>
Coevolving plasmids drive gene flow and genome plasticity in host-associated intracellular bacteria
<p>Comparative genomics and modeling of plasmids of the obligate host-associated intracellular phylum chlamydiae. </p>
Genomic Typing, Antimicrobial Resistance Gene, Virulence Factor and Plasmid Replicon Dataset for the Important Pathogenic Bacteria Klebsiella pneumoniae
<p>The infections caused by various bacterial pathogens both in clinical and community settings represent a significant threat to public healthcare worldwide. The growing resistance to antimicrobial drugs acquired by bacterial species causing healthcare-associated infections has already become a life-threatening danger noticed by the World Health Organization. Several groups or lineages of bacterial isolates usually called 'the clones of high risk' often drive the spread of resistance within particular species. </p> <p>Thus, it is vitally important to reveal and track the spread of such clones and the mechanisms by which they acquire antibiotic resistance and enhance their survival skills. Currently, the analysis of whole genome sequences for bacterial isolates of interest is increasingly used for these purposes, including epidemiological surveillance and developing of spread prevention measures. However, the availability and uniformity of the data derived from the genomic sequences often represents a bottleneck for such investigations. </p> <p>In this dataset, we present the results of a genomic epidemiology analysis of 61,857 genomes of a dangerous bacterial pathogen <em>Klebsiella pneumoniae</em> obtained from NCBI Genbank database. Important typing information including multilocus sequence typing (MLST)-based sequence types (STs), capsular (KL) and oligosaccharide (OL) types, CRISPR-Cas systems, and cgMLST profiles are presented, as well as the assignment of particular isolates to clonal groups (CG). The presence of antimicrobial resistance and virulence genes, as well as plasmid replicons, within the genomes is also reported. </p> <p>These data will be useful for researchers in the field of <em>K. pneumoniae</em> genomic epidemiology, resistance analysis and prevention measure development.</p>
Plasmids Identified in Air Metagenomes
<p> Metagenomic data were selected in Web of Science (Clarivate) on October 2022 using keywords: txid655179[Organism:noexp] AND metagenome [Filter]; AIR Metagenome; Air microbiome; Troposphere; Aerosol; Atmosphere. Data were manually curated to remove sequencing originated from metabarcoding data (i.e., 16S). The assembled data supplied by MetaSUB consortium (Danko et al., 2021) when available was used for air metagenome in the built environments. </p> <div> <p>Plasmid contents were predicted using the assembled data. Metagenomes sequencing by Illumina (paired-illumina reads) were assembled by using megahit 1.2.9 with metalarge option (Li et al., 2015) after cleaning data with bbduk2 (qtrim=rl trimq=28 minlen=25 maq=20 ktrim=r k=25 mink=11 and a list of adaptators to remove) from bbtools suite (<a href="https://jgi.doe.gov/data-and-tools/software-tools/bbtools/" target="_blank" rel="noreferrer noopener">https://jgi.doe.gov/data-and-tools/software-tools/bbtools/</a>) </p> <div> <p><span><span>Plasmids were predicted for each assembling by using scripts describing in-depth in Hilpert et al. (Hilpert </span></span><span><span>et al.</span></span><span><span>, 2021; </span><span>Hennequin</span> </span><span><span>et al.</span></span><span><span>, 2022) and available in </span><span>github</span><span> website (</span></span><span><span><span>https://github.com/meb-team/PlasSuite/</span></span></span><span><span>). Briefly, contigs were analyzed using both reference-based and reference-free approaches.</span></span><span><span> The databases employed included those for chromosomes (archaea and bacteria) and plasmids from NCBI, as well as the MOB-suite tool (Robertson and Nash, 2018</span><span>) ,</span><span> SILVA (Quast </span></span><span><span>et al.</span></span><span><span>, 2013) and phylogenetic markers harbored by chromosomes (Wu </span></span><span><span>et al.</span></span><span><span>, 2013). Two reference-free methods were applied to contigs that were not affiliated with chromosomes (discarded) or plasmids (</span><span>retained</span><span> in the first step): </span><span>PlasFlow</span><span> (Krawczyk et al., 2018) and </span><span>PlasClass</span><span> (Pellow </span></span><span><span>et al.</span></span><span><span>, 2020). Viruses were removed by using </span><span>viralVerify</span><span> (</span></span><span><span><span>https://github.com/ablab/viralVerify</span></span></span><span><span>) (Antipov </span></span><span><span>et al.</span></span><span><span>, 2020) that provides in parallel provide plasmid/non-plasmid classification</span><span>. </span><span> </span><span>The database built for this purpose is available at this address </span></span><span><span><span>https://github.com/meb-team/PlasSuite/?tab=readme-ov-file#1-prepare-or-download-your-databases</span></span></span> <span><span> Eukaryotes contaminants were removed by aligning the sequences against NT databases and human chromosomes (GRCh38) with minimap2 with -x asm5 </span><span>option (Li, 2018)</span><span>. Contigs mapping with an identity of 95% and a coverage of 80% were removed.</span></span><span> the final plasmidome set was clustered by mmseqs (Mirdita, Steinegger and Söding, 2019) with 80% of coverage and 90% of identity (--min-seq-id 0.90 -c 0.8 --cov-mode 1 --cluster-mode 2 --alignment-mode 3 --kmer-per-seq-scale 0.2). </span></p> </div> </div>
Plasmid Maps for a Nuclear Transformation Vector in Chlamydomonas reinhardtii for the Expression and Secretion of the Plastic-Degrading Enzyme (PHL7)
<p><strong>pJP32PHL7 Vector:</strong></p> <ul> <li> <p><strong>Size:</strong> 5692 bp</p> </li> <li> <p><strong>Key Features:</strong></p> <ul> <li><strong>HSP70 Promoter:</strong> A heat shock protein promoter fused with the <em>rbcS2</em> promoter to drive expression of downstream genes.</li> <li><strong>Ble Resistance Gene:</strong> Confers resistance to bleomycin, useful for selection in <em>Chlamydomonas reinhardtii</em>.</li> <li><strong>PHL7 Gene:</strong> Encodes the plastic-degrading enzyme PHL7, inserted downstream of the <em>F2A</em> site for expression in the host.</li> <li><strong>Intron Sequences:</strong> Contains multiple <em>rbcS2</em> introns for enhancing expression in <em>Chlamydomonas</em>.</li> <li><strong>Selectable Marker (AmpR):</strong> Confers ampicillin resistance for selection in <em>E. coli</em>.</li> <li><strong>Replication Origin:</strong> Includes <em>ori</em> and <em>F1 ori</em> for replication in <em>E. coli</em>.</li> </ul> <p> </p> </li> <li> <p><strong>Applications:</strong> This vector is designed for nuclear transformation in <em>Chlamydomonas reinhardtii</em>, enabling the expression and secretion of the plastic-degrading enzyme (PHL7) under the control of a hybrid <em>HSP70</em>rbcS2 promoter.</p> </li> </ul> <p><strong>pJP32PHL7dg Vector:</strong></p> <ul> <li> <p><strong>Size:</strong> 5692 bp</p> </li> <li> <p><strong>Key Features:</strong></p> <ul> <li><strong>HSP70 Promoter:</strong> Retains the HSP70 and <em>rbcS2</em> fusion promoter for gene expression.</li> <li><strong>LacZ Alpha Fragment:</strong> Includes a LacZ alpha fragment for blue/white screening.</li> <li><strong>PHL7 Gene:</strong> Encodes the plastic-degrading enzyme PHL7, linked downstream of the <em>F2A</em> site, allowing for expression in the host.</li> <li><strong>Ble Resistance Gene:</strong> Also confers bleomycin resistance for selection in <em>Chlamydomonas</em>.</li> <li><strong>Selectable Marker (AmpR):</strong> Confers ampicillin resistance for selection in <em>E. coli</em>.</li> <li><strong>Intron Sequences:</strong> Contains <em>rbcS2</em> introns for optimizing gene expression in the host organism.</li> </ul> <p> </p> </li> <li> <p><strong>Applications:</strong> The pJP32PHL7dg vector is similarly designed for nuclear transformation in <em>Chlamydomonas reinhardtii.</em> It also facilitates the expression and secretion of the plastic-degrading enzyme PHL7, driven by the hybrid <em>HSP70</em>rbcS2 promoter, but without glycosilation sites.</p> </li> </ul>
All plasmid sequences
<p>Here are the full sequences of 33 plasmids used in the following BioRxiv manuscript: https://doi.org/10.1101/2021.09.15.460454</p> <p> </p>
Assembly-based analysis of the infant gut microbiome reveals novel ubiquitous plasmids
<p>Assembly-based plasmids found in the gut microbiome of 12 infants born in Norway (BabyBiome project).</p>
Data from: Genes for cooperation are not more likely to be carried by plasmids
<p>Cooperation is prevalent across bacteria, but risks being exploited by non-cooperative cheats. Horizontal gene transfer, particularly via plasmids, has been suggested as a mechanism to stabilize cooperation. A key prediction of this hypothesis is that genes that are more likely to be transferred, such as those on plasmids, should be more likely to code for cooperative traits. Testing this prediction requires identifying all genes for cooperation in bacterial genomes. However, previous studies used a method that likely misses some of these genes for cooperation. To solve this, we used a new genomics tool, SOCfinder, which uses three distinct modules to identify all kinds of genes for cooperation. We compared where these genes were located across 4648 genomes from 146 bacterial species. In contrast to the prediction of the hypothesis, we found no evidence that plasmid genes are more likely to code for cooperative traits. Instead, we found the opposite - that genes for cooperation were more likely to be carried on chromosomes. Overall, the vast majority of genes for cooperation are not located on plasmids, suggesting that the more general mechanism of kin selection is sufficient to explain the prevalence of cooperation across bacteria.</p>
PlasX model, Predicted plasmids, and Known plasmids
<p>PlasX model and all analyses of known and predicted plasmids</p>
Galaxy Training Data for "Designing plasmids encoding predicted pathways by using the BASIC assembly method"
<p>This dataset provides the data needed for the Galaxy BASIC assembly workflow training tutorial (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). This workflow provides a pathway to design plasmids encoding predicted metabolic pathways using the BASIC assembly method (<a href="https://doi.org/10.1021/sb500356d">https://doi.org/10.1021/sb500356d</a>). It generates scripts allowing the automatic construction of these plasmids using an Opentrons liquid handling robot. After downloading these scripts on a computer connected to an Opentrons (<a href="https://opentrons.com">https://opentrons.com</a>), the user can perform the automatic construction of the plasmids on the bench.</p> <p>The content of the dataset is as follows:</p> <ul> <li> <p>an SBML file modeling a heterologous pathway producing lycopene such as those produced by the Pathway Analysis Workflow (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>).</p> </li> <li> <p>a CSV file listing the parts to be used (linkers, backbone and promoters) in the constructions.</p> </li> <li> <p>two YAML files providing two examples of settings, i.e. providing the identifiers of the laboratory equipment and the parameters of the DNA robot.</p> </li> </ul>
Comprehensive discovery of CRISPR-targeted terminally redundant sequences in the human gut metagenome: viruses, plasmids, and more
<p>S1 Data</p> <p>Dataset including the discovered CRISPR spacers, direct repeats, protospacers, co-occurrence-based spacer clustering results, predicted protein sequences, built HMMs, database comparison results, phylogenetic analysis results, predicted targeting hosts, and CRISPR-targeted TR sequences.</p>
Global Salmonella plasmids - Supplemental Data
<p>Sequence data and analysis tool results from a large-scale plasmid characterization from global Salmonella data</p>
METADATA for results of irradiation-induced complex DNA damage measurements using plasmid pBR322 along a typical Proton Treatment Plan at the MedAustron proton and carbon beam therapy facility (energy 137–198 MeV and Linear Energy Transfer (LET) range 1–9 keV/μm), by means of Agarose Gel Electrophoresis and DNA fragmentation using Atomic Force Microscopy (AFM)
Open the record for dataset details and reuse information.
plasmid_masking:v21.1.1
<div> <p>kraken2 DB for plasmids built with masking option in Jan 2021. Contains 22283 accession numbers corresponding to 3357 taxons.</p> <p> </p> </div>
Nanopore sequencing of plasmid cleavage fragments produced with type III CRISPR-associated nucleases NucC, Can1 and Can2
<p>Included datasets were generated in the study "<strong>Sequence-specific capture and concentration of viral RNA </strong><strong>by type III CRISPR system enhances diagnostic"</strong> by Nemudraia et al., 2022</p> <p> </p> <p>For questions contact: Artem Nemudryi (artem.nemudryi@gmail.com) or Blake Wiedenheft (bwiedenheft.com)</p>
PlasBin-flow: A flow-based MILP algorithm for plasmid contigs binning
<p>PlasBin-flow is a method for detecting plasmid contigs bins from the assembly graph for a given bacterial sample. The method is based on a Mixed-Integer Linear Programming (MILP) formulation.</p> <p>The data shared consists of output files of PlasBin-flow as well as those of 5 other plasmid binning methods, namely, PlasBin, HyAsP, MOB-recon, plasmidSPAdes and gplas for 66 test samples. Details about each sample have been provided in the file <strong>samples.csv</strong>. </p> <p>The output folder contains one folder per sample. Each sample folder contains the following files: </p> <ul> <li>PlasBin-flow was executed using 7 different weight combinations for the objective function of the MILP.<br> For every weight combination <span class="math-tex">\((a,b,c)\)</span> , we have a file name <em>plasbin_flow_a_b_c_bins.out.</em></li> <li>PlasBin: <em>plasbin_contig_chains.csv</em>,</li> <li>HyAsP: <em>hyasp_plasmid_contigs.fasta</em>,</li> <li>MOB-recon: <em>mob_recon_contig_report.txt</em>,</li> <li>plasmidSPAdes: <em>plasmidspades_contigs.fasta</em>,</li> <li>gplas: <em>gplas_bins.tab</em>.</li> </ul> <p> </p> <p> </p>
Development and validation of a novel plasmid chassis system for screening of metabolite-responsive transcription factors
<p>This dataset contains the raw data that lie at the basis of the results discussed in <strong>Chapter 3: Development and validation of a novel plasmid chassis system for screening of metabolite-responsive transcription factors </strong>of the PhD thesis of Amber Bernauw. The README.txt file provides more information on the different data files.</p>
Data from: Genes for cooperation are not more likely to be carried by plasmids
Open the record for dataset details and reuse information.
Investigating huntingtin DNA binding – plasmid EMSA with full-length HTT Q23/Q46
<p>Huntingtin structure-function open lab notebook.</p> <p> </p> <p>NB: plasmid concentration should read 0.5 mg/mL not 0.5 mg/uL.</p>
Figures for: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean
<p>Figures created for the short communication: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.