Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,696
datasets available to search
ShareScore release 0.9.0
Dataset results
1,696 results for “DNA sequences”
Phylogeny of "Philoceanus complex" seabird lice (Phthiraptera: Ischnocera) inferred from mitochondrial DNA sequences
<p>Data from "Phylogeny of “<em>Philoceanus </em>complex” seabird lice (Phthiraptera: Ischnocera) inferred from mitochondrial DNA sequences". See the file index.html for details. Data includes NEXUS files for sequences, tree files output by MrBayes and PAUP, and host-parasite association files for TreeMap.</p>
Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing
<p><strong>OV2295 Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz: Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz: Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number </li> <li>minor_cn: HMM predicted minor copy number </li> <li>major_cn: HMM predicted major copy number </li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz: Table of SNVs per clone for OV2295 samples. Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het: is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format. Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: ‘SA922’: ‘OV2295(R2)’, ‘SA921’: ‘TOV2295(R)’, ‘SA1090’: ‘OV2295’,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>
Graphic Illustration of Neal Platt's Talk: Targeted sequencing of pathogen DNA from museum specimens
<p><a href="https://lib.ku.edu/people/courtney-foat" target="_blank" rel="noopener">Courtney Foat</a>, Advisor for Strategic Initiatives & Organizational Engagement at the University of Kansas, graphically recorded this invited talk by Neal Platt at an NSF-supported Workshop: Digital Collections Data and Tracking Disease.</p>
The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.
<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Université Claude Bernard Lyon 1, VetAgro Sup, Research Team “Bacterial Opportunistic Pathogens and Environment” (BPOE), 69280 Marcy L’Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bonté et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer’s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O'LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project “Iouqmer” EST 2016/1/120, l'Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rhône, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: Céline COLINON, Romain MARTI, Emilie BOURGEOIS, Sébastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257–4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161–168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bonté, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153–164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>
Catalog of GenBank sequence read archive (SRA) entries of metagenomic DNA sequence analyses of bacterial and archaeal water column communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2012
In contrast to temperate systems, Arctic lagoons that span the Alaska Beaufort Sea coast face extreme seasonality. Nine months of ice cover up to ∼1.7 m thick is followed by a spring thaw that introduces an enormous pulse of freshwater, nutrients, and organic matter into these lagoons over a relatively brief 2–3 week period. Prokaryotic communities link these subsidies to lagoon food webs through nutrient uptake, heterotrophic production, and other biogeochemical processes, but little is known about how the genomic capabilities of these communities respond to seasonal variability. This study characterizes the metabolic capabilities of microbial communities across three seasons in two lagoons and one open coastal site along the eastern Alaska Beaufort Sea coast. We used metagenomic DNA sequence data of bacterial and archaeal water column communities to identify genes of relevant biogeochemical pathways. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA642637 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA642637. This data package is associated with the following publication: Baker, Kristina D., Colleen T. E. Kellogg, James W. McClelland, Kenneth H. Dunton, and Byron C. Crump. “The Genomic Capabilities of Microbial Communities Track Seasonal Variation in Environmental Conditions of Arctic Lagoons.” Frontiers in Microbiology 12 (2021). https://doi.org/10.3389/fmicb.2021.601901. Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provi
Multiple alignment of DNA-B sequences from CMMGV, EACMCV, EACMV, EACMKV, EACMMV, EACMZV, SACMV (7 "species")
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Multiple alignment of ACMV and ACMBFV DNA-B sequences
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Transcriptome analysis of the effect of over-expressing H2A.J mutants in proliferating WI38 fibroblasts for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences
<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>H2A.J differs from canonical H2A only by a valine at position 11 instead of alanine, and the 7 C-terminal amino acids containing a potential minimal phosphorylation site SQ for DNA-damage response kinases. To test the functional importance of these H2A.J-specific sequences, we mutated Val-11 to Ala as is found in all canonical H2A sequences, and we mutated Ser-123 to either Glu to mimic a phospho-serine residue or to Ala to prevent phosphorylation. We also substituted the C-terminus of H2A.J with the C-terminus of H2A. These mutants, WT-H2A.J and canonical H2A-type1 were ectopically expressed in proliferating fibroblasts, and their microarray transcriptomes were compared to that of proliferating and senescent fibroblasts without ectopic histone expression. Genome-wide transcriptome analysis indicated that senescent fibroblasts clustered distinctly from proliferating fibroblasts, and proliferating fibroblasts expressing the H2A.J-V11A and H2A.J-S123E mutants clustered distinctly from fibroblasts expressing the other H2A.J mutants, WT-H2A.J, and H2A. Hallmark gene set enrichment analysis of the transcriptomes of fibroblasts expressing H2A.J-V11A or H2A.J-S123E versus control proliferating fibroblasts indicated that they showed the same highly significant enrichment for the Epithelial-Mesenchyme Transition, TNF-Alpha Signaling Via NF-kB, and Inflammatory Response gene sets. Notable inflammatory genes including IL1A, IL1B, IL6, CXCL8, and CCL2 are contained in these gene sets and are often induced in senescence as part of the senescence-associated secretory phenotype. Heat maps showed that the H2A.J-V11A and H2A.J-S123E mutants were particularly apt at activating the expression of these inflammatory genes in proliferating fibroblasts</p>
Tracking Down Chimeric Assemblies In The TrackIt DNA Ladder Using Nanopore Sequencing
<h2>Dataset Description</h2><p>These files represent two different LSK114 sequencing runs on a TrackIt 1kb Plus DNA Ladder sample, and associated data analysis.</p><ul><li>July 20 2023 Flongle Run (191 Mb; 465k reads)<ul><li>pod5_files_2023-Jul-20_DAE_DNA_Ladder.tar.gz<br>- raw POD5 format files</li><li>called_2023-Jul-20_DAE_DNA_Ladder_duplex.bam<br>- duplex called reads, called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Jul-20_DAE_DNA_Ladder.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Jul-20_DAE_DNA_Ladder_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Jul-20_DAE_DNA_Ladder.txt<br>- Length / QC summary statistics</li></ul></li><li>October 12 2023 P2 Solo Run (1.95 Gb, 1.11M reads)<ul><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_fail.tar.gz<br>- raw POD5 format files (all failed reads)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_000-059.tar.gz<br>- raw POD5 format files (passed reads, bundle #000-059)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_060-119.tar.gz<br>- raw POD5 format files (passed reads, bundle #060-119)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_120-179.tar.gz<br>- raw POD5 format files (passed reads, bundle #120-179)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_180-222.tar.gz<br>- raw POD5 format files (passed reads, bundle #180-222)</li><li>called_2023-Oct-12_DNA-Ladder-1kbplus_duplex.bam<br>- duplex called reads [October 12, 2023], called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Oct-12_DNA-Ladder-1kbplus.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Oct-12_DNA-Ladder-1kbplus_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Oct-12_DNA-Ladder-1kbplus.txt<br>- Length / QC summary statistics</li><li>ladder_seqs.fa<br>- assembled DNA ladder sequences, based on simplex reads</li></ul></li></ul><h3>Methods</h3><h3>Sample preparation</h3><p>Preparation of DNA for sequencing was carried out following the ONT Ligation Sequencing DNA V14 (SQK-LSK114) protocol, with modifications to exclude DNA repair, and keeping the sample in the same 1.5ml tube to reduce sample loss.</p><h4>Tris-buffered Saline (TBS) buffer preparation</h4><ol><li>1M stock of NaCl was made by adding 2.922g of NaCl into a 50 ml Falcon tube, then made up to 50 ml with MilliPore water</li><li>A 50 mM TBS stock was created by adding 750 μl 1M NaCl solution to a 15 ml Falcon tube, then made up to 15 ml using Qiagen Elution Buffer (EB, i.e. 10 mM Tris-HCl at pH 8.0)</li><li>The pH was confirmed to be 7.9-8.1 using a pH indicator strip (e.g. MColorpHast 6.5 - 10.0; MER1095430001)</li></ol><h4>End prep</h4><ol><li>1 μg DNA ladder (i.e. 10 μl of 0.1 μg / μl DNA ladder) was transferred into a 1.5ml Eppendorf DNA LoBind tube</li><li>The volume was topped up to 43.5 μl with TBS (i.e. 33.5 μl TBS)</li><li>3.5 μl Ultra II End-prep Reaction Buffer and 3 μl Ultra II End-prep Enzyme Mix was added</li><li>After mixing by gentle pipetting, the mixture was incubated at RT for 5 minutes, then 65 \degrees for 5 minutes</li></ol><h4>Bead cleanup</h4><ol><li>The mixture was combined with 60 μl Ampure XP beads, and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 150 μl of an 80% ethanol solution</li><li>The sample was dried briefly for 30s, then eluted in 60 μl TBS</li></ol><h4>Adapter ligation and final bead cleanup</h4><ol><li>To the sample tube was added 25 μl ONT Ligation buffer (LNB), 5μl NEBNext Quick T4 DNA Ligase (reduced from the protocol-suggested 10μl because that was all that was left in the tube), and 5μl ONT Ligation Adapter (LA)</li><li>The tube was mixed by gentle pipetting, spun down for 1-3s on a mini centrifuge, then incubated for 10 minutes at RT</li><li>The mixture was combined with 40 μl Ampure XP beads (100μl Ampure XP beads were used for the Flongle sample), and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 250 μl of ONT Long Fragment Buffer (LFB) for the P2 Solo run, and 250μl ONT Short Fragment Buffer <br>(SFB) for the Flongle run</li><li>The sample was dried briefly for 30s, then eluted for 10 minutes at 37 \degrees in 15 μl ONT Elution buffer (EB)</li></ol><h4>Addition of sequencing library buffers</h4><ol><li>A flow cell was prepared by flushing with ONT Flow Cell Flush (FCF) mixed with ONT Flow Cell Tether (FCT). For the P2 Solo, I used 500 μl of a 1170μl FCF solution that had 30μl FCT added to it; for the Flongle I used 60μl of a 117μl FCF solution that had 3 μl FCT added to it</li><li>1 μl of the eluted library was quantified on a Quantus Fluorometer, and approximately 50 fmol (assuming 1kb average length) was transferred to a new 1.5μl tube</li><li>For the P2 Solo run, the volume was topped up to 32 μl TBS; for the Flongle run, the volume was topped up to 12 μl TBS</li><li>To the sample tube was added ONT Sequencing Buffer (SB; P2 Solo - 100μl; Flongle - 30μl) and ONT Library Beads (LIB; P2 Solo - 68μl; Flongle - 20μl)</li><li>The flow cell was re-flushed with additional FCF/FCT mixture (500 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The sequencing library was then added to the flow cell (200 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The prepared flow cell was left for 10 minutes to allow the library to settle before starting sequencing</li></ol><h4>DNA Sequencing and basecalling</h4><ol><li>Sequencing was carried out using MinKNOW v23.04.6, sequencing in fast mode at 400 bases per second with a 20bp minimum sequence length and <br>5 kHz sampling rate, with reads output as POD5 files</li><li>The Flongle flow cell was run for a full standard run length (24h), whereas the PromethION flow cell was run for 1.5 hours (after which the <br>counts of 15kb reads exceeded 200)</li><li>Sequenced reads were recalled in standard (simplex) mode using Dorado v0.4.0 and the 2023-09-22 bacterial methylation model [res_dna_r10.4.1_e8.2_400bps_sup@2023-09-22_bacterial-methylation]</li></ol><h3>Bioinformatics Analysis of Ladder Sequences </h3><h4>Sequence assembly</h4><p>Assembly process for bands that are 3k in length and greater (done on LFB-depleted P2 Solo sequences): </p><ol><li>Filter >q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not sufficient for the 15kb band; all reads were needed]</li><li>Chop the reads up with a 1000bp overlap (e.g. 3000bp for the 5k band). This works around a Canu expectation that any read overlaps should be less than X% of the read.</li><li>Assemble the reads with Canu v2.2 [#REF], treating them as "pacbio" reads (for correction and homopolymer compression), with the GenomeSize parameter set to the expected band length (e.g. GenomeSize=5000).</li><li>Extract the first reported assembled contig.</li><li>Map the contig to the nanopore adapter sequences, and trim to exclude any matching sequence.</li></ol><p>[Canu has a default genome size and read length cutoff of 1kb, and performs poorly on sequences shorter than this] <br><br>Assembly process for bands under 3k in length (done on LFB-depleted P2 Solo sequences):</p><ol><li>Filter >q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not in sufficient abundance for the 100bp band; all reads were needed]</li><li>Assemble using a<a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kassembler.pl"> kmer-based de-bruijn assembler</a>, trimming off low-count kmers</li><li>Extract the first reported trimmed assembled chain</li><li>Map the assembled chain to the nanopore adapter sequences, and trim to exclude any matching sequence</li><li>Use web BLASTn [#REF] to help trim any additional trailing non-matching sequence</li></ol><h4>Mapping</h4><ol><li>Use a <a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kmapper.pl">kmer-based lightweight mapper</a> to map reads to assembled bands</li><li>Created LAST mismatch matrix using `last-train` on the 5k reads together, using the full assembled ladder sequences as a reference:<br>last-train -Q 1 ladder_seqs.fa 5k_reads.fq.gz</li><li>Mapped all reads to the assembled ladder sequences (only the reference corresponding to the most likely band source), retaining (for each read) the mapping that had the longest combined proportion of read and reference sequence mapped:<br>lastal -p bacterial.mat -P 10 ladder_seqs.fa reads_2023-Oct-12_DNA-Ladder-1kbplus_called_all.fq.gz | \ <br> ~/scripts/maf2csv.pl | \ <br> awk -F ',' '{print $0","($8/100 * $13/100)}' | \ <br> sort -t ',' -k 16rg,16 | sort -t ',' -k 1,1 -u | sort -t ',' -k 1r,1 | \ <br> perl -pe 's/,[^,]*$/\n/' > LAST_reads_vs_ladder_longestMatch.csv.gz</li></ol>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Three new species of Torrent Treefrogs (Anura: Hylidae) of the Hyloscirtus bogotensis group from the eastern Andean slopes and the biogeographic history of the genus"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "Three new species of Torrent Treefrogs (Anura: Hylidae) of the Hyloscirtus bogotensis group from the Amazon foothills and the biogeographic history of the genus"</p> <p>The matrix is in NEXUS format and has 3259 bp and 25 terminals.</p> <p>Partitions are as follows:</p> <div>charset 12S = 1-955;</div> <div>charset ND1_nonCoding1 = 956-1279;</div> <div>charset ND1_Pos1 = 1280-2240\3;</div> <div>charset ND1_Pos2 = 1281-2241\3;</div> <div>charset ND1_Pos3 = 1282-2242\3;</div> <div>charset ND1_nonCoding2 = 2243-2361;</div> <div>charset cmyc_Pos1 = 2362-2779\3;</div> <div>charset cmyc_Pos2 = 2363-2780\3;</div> <div>charset cmyc_Pos3 = 2364-2781\3;</div> <div>charset Rag1_Pos1 = 2782-3415\3;</div> <div>charset Rag1_Pos2 = 2783-3416\3;</div> <div>charset Rag1_Pos3 = 2784-3417\3;</div>
DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region
<p>This dataset has been prepared as part of the Interreg Alpine Space project Eco-AlpsWater (ASP569) - <em>Innovative Ecological Assessment and Water Management Strategy for the Protection of Ecosystem Services in Alpine Lakes and Rivers</em>, <a href="https://www.alpine-space.eu/projects/eco-alpswater/en/home">https://www.alpine-space.eu/projects/eco-alpswater/en/home</a></p> <p>Individual archives include 16S rRNA (cyanobacteria) and 18S rRNA (microalgae) FASTA sequences and associated blastn results obtained from the high throughput sequencing of plankton and biofilm bulk/eDNA samples collected in 2019 in 37 lakes and 22 rivers across the Alpine region. These are supporting files for the paper by Salmaso et al., 2022. DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region. Science of the Total Environment, in press.</p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "A new glassfrog of the genus Centrolene (Amphibia: Centrolenidae) from the Subandean Kutukú Cordillera, eastern Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "A new glassfrog of the genus Centrolene (Amphibia: Centrolenidae) from the Subandean Kutukú Cordillera, eastern Ecuador"</p> <p>The matrix is in NEXUS format and has 6626 bp and 239 terminals.</p> <p>Partitions are as follows:</p> <div> <div>charset 12S = 1-967;</div> <div>charset 16S = 968-2130;</div> <div> </div> <div>charset BNDFcodonPos1 = 2133-2829\3;</div> <div>charset BNDFcodonPos2 = 2131-2830\3;</div> <div>charset BNDFcodonPos3 = 2132-2828\3;</div> <div> </div> <div> </div> <div>charset ND1codonPos1 = 2832-3786\3;</div> <div>charset ND1codonPos2 = 2833-3787\3;</div> <div>charset ND1codonPos3 = 2831-3788\3;</div> <div> </div> <div> </div> <div>charset CXCR4codonPos1 = 3790-4144\3;</div> <div>charset CXCR4codonPos2 = 3791-4142\3;</div> <div>charset CXCR4codonPos3 = 3789-4143\3;</div> <div> </div> <div> </div> <div>charset cmyccodonPos1 = 4145-4547\3;</div> <div>charset cmyccodonPos2 = 4146-4548\3;</div> <div>charset cmyccodonPos3 = 4147-4549\3;</div> <div> </div> <div> </div> <div>charset POMCcodonPos1 = 4551-5160\3;</div> <div>charset POMCcodonPos2 = 4552-5161\3;</div> <div>charset POMCcodonPos3 = 4550-5162\3;</div> <div> </div> <div> </div> <div>charset RAG1codonPos1 = 5163-5616\3;</div> <div>charset RAG1codonPos2 = 5164-5617\3;</div> <div>charset RAG1codonPos3 = 5165-5618\3;</div> <div> </div> <div> </div> <div>charset SLC8A1codonPos1 = 5620-6160\3;</div> <div>charset SLC8A1codonPos2 = 5621-6158\3;</div> <div>charset SLC8A1codonPos3 = 5619-6159\3;</div> <div> </div> <div> </div> <div> </div> <div>charset SLC8A3codonPos1 = 6162-6627\3;</div> <div>charset SLC8A3codonPos2 = 6163-6625\3;</div> <div>charset SLC8A3codonPos3 = 6161-6626\3;</div> </div> <p> </p>
Supplemental_Data_S1 for "Kmer Manifold Approximation and Projection for visualizing DNA sequences"
<p>This dataset includes the results generated by KMAP software applied to the htselexdata dataset. Each folder within the dataset contains outputs from multiple dimensionality reduction techniques, including KMAP, UMAP, t-SNE, and MDS. Additionally, motifs and logos have been derived using both KMAP and MEME methods. This data provides insights into motif patterns and structures, which can be beneficial for further bioinformatics and computational biology analyses.</p>
Quantifying the Tissue-Specific Regulatory Information within Enhancer DNA Sequences
<p>Tables of mouse embryonic differential enhancers and promoters.</p> <p>The tables list the genomic coordinates of differential regulatory elements, the tissues in which they are active (labels column), a differential score, and the DNA sequence of the element. Tissues are numbered from 0 to 7, which correspond to heart, kidney, liver, limb, lung, forebrain, midbrain, and hindbrain, respectively. The tool <a href="https://github.com/pbenner/gonetics/tree/master/tools/segmentationDifferential">segmentationDifferential</a> from the <em>gonetics library</em> was used to compute differential elements. The length of all regions was set to 1000 base pairs around the center. Specifics of the experimental data can be found in the referenced publication.</p>
Subset of nucleosomal DNA sequences from mouse brain nucleus accumbens tissue (GEO dataset GSE54263)
<p>This dataset contains a subset of nucleosomal DNA sequences of +1 nucleosomes from mouse brain nucleus accumbens cells (NAC) used to analyze nucleosome positioning sequence (NPS) patterns in <a href="https://doi.org/10.1371/journal.pcbi.1007365">Pranckeviciene, Erinija and Hosid, Sergey and Liang, Nathan and Ioshikhes, Ilya (2020). Nucleosome positioning sequence patterns as packing or regulatory. In PLoS computational biology, 16 (1), pp. e1007365.</a></p> <ul> <li>controlm.fa.gz contains sequences of <strong>control</strong> mice (GSE54263 subset Con_H3 GSM1311267)</li> <li> resilientm.fa.gz contains sequences of mice <strong>resilient to social stress</strong> (GSE54263 subset Res_H3 GSM1311268)</li> <li> susceptiblem.fa.gz contains sequences of<strong> </strong>mice <strong>susceptible to social stress</strong> (GSE54263 subset Sus_H3 GSM1311269)</li> </ul> <p>This dataset originates from the GEO accession GSE54263 data from <a href="https://www.nature.com/articles/nm.3939">Sun H, Damez-Werno DM, Scobie KN, Shao NY et al. ACF chromatin-remodeling complex mediates stress-induced depressive-like behavior. <em>Nat Med</em> 2015 Oct;21(10):1146-53.</a></p>
Fig. 8 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 8. Thyropygus sutchariti sp. nov., from Kaeng Krachan, holotype (CUMZ-D00090), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Left telopodite, posterior-mesal view. D. Left telopodite, anterior-lateral view.
Fig. 11. A in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 11. A. Thyropygus navychula sp. nov., specimen from Surin Islands, living ♂ (paratype, CUMZ-D00089-1). B. Thyropygus forceps sp. nov., specimen from Namwang Srithammasokrach, living ♂ (paratype, CUMZ-D00073-1).
Fig. 5 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 5. Thyropygus mesocristatus sp. nov., from Srikasorn, holotype (CUMZ-D00094), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Lateral view. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view.
Fig. 2 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 2. Thyropygus cimi sp. nov., from Namwang Srithammasokrach, holotype (CUMZ-D00086), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Lateral view. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view.
Fig. 1 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 1. Phylogenetic relationships of Thyropygus species based on maximum likelihood analysis (ML) and Bayesian Inference (BI) of 1147 bp of concatenated gene fragments of COI (660 bp) and 16S rRNA (487 bp). Numbers at nodes indicate branch support based on bootstrapping (ML) / posterior probability (BI). Scale bar = 0.06 substitutions/site. # indicates branches which received <50% ML bootstrap support, - indicates non-supported branches by posterior probability. Clade memberships and designations are shown as vertical bars; 1A1 = T. allevatus, 1A2 = cuisinieri subgroup, 1A3 = opinatus subgroup and 1A4 = induratus subgroup. The coloured area marks the T. opinatus subgroup. Abbreviations after species names refer to locality names as shown in Table 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.