Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

19

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

19 results for “chimeric sequences”

Learn how ShareScore rates datasets ↗
zenodo44/100

Tracking Down Chimeric Assemblies In The TrackIt DNA Ladder Using Nanopore Sequencing

<h2>Dataset Description</h2><p>These files represent two different LSK114 sequencing runs on a TrackIt 1kb Plus DNA Ladder sample, and associated data analysis.</p><ul><li>July 20 2023 Flongle Run (191 Mb; 465k reads)<ul><li>pod5_files_2023-Jul-20_DAE_DNA_Ladder.tar.gz<br>- raw POD5 format files</li><li>called_2023-Jul-20_DAE_DNA_Ladder_duplex.bam<br>- duplex called reads, called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Jul-20_DAE_DNA_Ladder.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Jul-20_DAE_DNA_Ladder_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Jul-20_DAE_DNA_Ladder.txt<br>- Length / QC summary statistics</li></ul></li><li>October 12 2023 P2 Solo Run (1.95 Gb, 1.11M reads)<ul><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_fail.tar.gz<br>- raw POD5 format files (all failed reads)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_000-059.tar.gz<br>- raw POD5 format files (passed reads, bundle #000-059)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_060-119.tar.gz<br>- raw POD5 format files (passed reads, bundle #060-119)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_120-179.tar.gz<br>- raw POD5 format files (passed reads, bundle #120-179)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_180-222.tar.gz<br>- raw POD5 format files (passed reads, bundle #180-222)</li><li>called_2023-Oct-12_DNA-Ladder-1kbplus_duplex.bam<br>- duplex called reads [October 12, 2023], called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Oct-12_DNA-Ladder-1kbplus.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Oct-12_DNA-Ladder-1kbplus_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Oct-12_DNA-Ladder-1kbplus.txt<br>- Length / QC summary statistics</li><li>ladder_seqs.fa<br>- assembled DNA ladder sequences, based on simplex reads</li></ul></li></ul><h3>Methods</h3><h3>Sample preparation</h3><p>Preparation of DNA for sequencing was carried out following the ONT Ligation Sequencing DNA V14 (SQK-LSK114) protocol, with modifications to exclude DNA repair, and keeping the sample in the same 1.5ml tube to reduce sample loss.</p><h4>Tris-buffered Saline (TBS) buffer preparation</h4><ol><li>1M stock of NaCl was made by adding 2.922g of NaCl into a 50 ml Falcon tube, then made up to 50 ml with MilliPore water</li><li>A 50 mM TBS stock was created by adding 750 μl 1M NaCl solution to a 15 ml Falcon tube, then made up to 15 ml using Qiagen Elution Buffer (EB, i.e. 10 mM Tris-HCl at pH 8.0)</li><li>The pH was confirmed to be 7.9-8.1 using a pH indicator strip (e.g. MColorpHast 6.5 - 10.0; MER1095430001)</li></ol><h4>End prep</h4><ol><li>1 μg DNA ladder (i.e. 10 μl of 0.1 μg / μl DNA ladder) was transferred into a 1.5ml Eppendorf DNA LoBind tube</li><li>The volume was topped up to 43.5 μl with TBS (i.e. 33.5 μl TBS)</li><li>3.5 μl Ultra II End-prep Reaction Buffer and 3 μl Ultra II End-prep Enzyme Mix was added</li><li>After mixing by gentle pipetting, the mixture was incubated at RT for 5 minutes, then 65 \degrees for 5 minutes</li></ol><h4>Bead cleanup</h4><ol><li>The mixture was combined with 60 μl Ampure XP beads, and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 150 μl of an 80% ethanol solution</li><li>The sample was dried briefly for 30s, then eluted in 60 μl TBS</li></ol><h4>Adapter ligation and final bead cleanup</h4><ol><li>To the sample tube was added 25 μl ONT Ligation buffer (LNB), 5μl NEBNext Quick T4 DNA Ligase (reduced from the protocol-suggested 10μl because that was all that was left in the tube), and 5μl ONT Ligation Adapter (LA)</li><li>The tube was mixed by gentle pipetting, spun down for 1-3s on a mini centrifuge, then incubated for 10 minutes at RT</li><li>The mixture was combined with 40 μl Ampure XP beads (100μl Ampure XP beads were used for the Flongle sample), and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 250 μl of ONT Long Fragment Buffer (LFB) for the P2 Solo run, and 250μl ONT Short Fragment Buffer&nbsp;<br>(SFB) for the Flongle run</li><li>The sample was dried briefly for 30s, then eluted for 10 minutes at 37 \degrees in 15 μl ONT Elution buffer (EB)</li></ol><h4>Addition of sequencing library buffers</h4><ol><li>A flow cell was prepared by flushing with ONT Flow Cell Flush (FCF) mixed with ONT Flow Cell Tether (FCT). For the P2 Solo, I used 500 μl of a 1170μl FCF solution that had 30μl FCT added to it; for the Flongle I used 60μl of a 117μl FCF solution that had 3 μl FCT added to it</li><li>1 μl of the eluted library was quantified on a Quantus Fluorometer, and approximately 50 fmol (assuming 1kb average length) was transferred to a new 1.5μl tube</li><li>For the P2 Solo run, the volume was topped up to 32 μl TBS; for the Flongle run, the volume was topped up to 12 μl TBS</li><li>To the sample tube was added ONT Sequencing Buffer (SB; P2 Solo - 100μl; Flongle - 30μl) and ONT Library Beads (LIB; P2 Solo - 68μl; Flongle - 20μl)</li><li>The flow cell was re-flushed with additional FCF/FCT mixture (500 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The sequencing library was then added to the flow cell (200 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The prepared flow cell was left for 10 minutes to allow the library to settle before starting sequencing</li></ol><h4>DNA Sequencing and basecalling</h4><ol><li>Sequencing was carried out using MinKNOW v23.04.6, sequencing in fast mode at 400 bases per second with a 20bp minimum sequence length and &nbsp;<br>5 kHz sampling rate, with reads output as POD5 files</li><li>The Flongle flow cell was run for a full standard run length (24h), whereas the PromethION flow cell was run for 1.5 hours (after which the&nbsp;<br>counts of 15kb reads exceeded 200)</li><li>Sequenced reads were recalled in standard (simplex) mode using Dorado v0.4.0 and the 2023-09-22 bacterial methylation model [res_dna_r10.4.1_e8.2_400bps_sup@2023-09-22_bacterial-methylation]</li></ol><h3>Bioinformatics Analysis of Ladder Sequences&nbsp;</h3><h4>Sequence assembly</h4><p>Assembly process for bands that are 3k in length and greater (done on LFB-depleted P2 Solo sequences):&nbsp;</p><ol><li>Filter &gt;q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not sufficient for the 15kb band; all reads were needed]</li><li>Chop the reads up with a 1000bp overlap (e.g. 3000bp for the 5k band). This works around a Canu expectation that any read overlaps should be less than X% of the read.</li><li>Assemble the reads with Canu v2.2 [#REF], treating them as "pacbio" reads (for correction and homopolymer compression), with the GenomeSize parameter set to the expected band length (e.g. GenomeSize=5000).</li><li>Extract the first reported assembled contig.</li><li>Map the contig to the nanopore adapter sequences, and trim to exclude any matching sequence.</li></ol><p>[Canu has a default genome size and read length cutoff of 1kb, and performs poorly on sequences shorter than this]&nbsp;<br><br>Assembly process for bands under 3k in length (done on LFB-depleted P2 Solo sequences):</p><ol><li>Filter &gt;q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not in sufficient abundance for the 100bp band; all reads were needed]</li><li>Assemble using a<a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kassembler.pl"> kmer-based de-bruijn assembler</a>, trimming off low-count kmers</li><li>Extract the first reported trimmed assembled chain</li><li>Map the assembled chain to the nanopore adapter sequences, and trim to exclude any matching sequence</li><li>Use web BLASTn [#REF] to help trim any additional trailing non-matching sequence</li></ol><h4>Mapping</h4><ol><li>Use a <a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kmapper.pl">kmer-based lightweight mapper</a> to map reads to assembled bands</li><li>Created LAST mismatch matrix using `last-train` on the 5k reads together, using the full assembled ladder sequences as a reference:<br>last-train -Q 1 ladder_seqs.fa 5k_reads.fq.gz</li><li>Mapped all reads to the assembled ladder sequences (only the reference corresponding to the most likely band source), retaining (for each read) the mapping that had the longest combined proportion of read and reference sequence mapped:<br>lastal -p bacterial.mat -P 10 ladder_seqs.fa reads_2023-Oct-12_DNA-Ladder-1kbplus_called_all.fq.gz | \&nbsp;<br>&nbsp;&nbsp;&nbsp;~/scripts/maf2csv.pl | \&nbsp;<br>&nbsp;&nbsp;&nbsp;awk -F ',' '{print $0","($8/100 * $13/100)}' | \&nbsp;<br>&nbsp;&nbsp;&nbsp;sort -t ',' -k 16rg,16 | sort -t ',' -k 1,1 -u | sort -t ',' -k 1r,1 | \&nbsp;<br>&nbsp;&nbsp;&nbsp;perl -pe 's/,[^,]*$/\n/' &gt; LAST_reads_vs_ladder_longestMatch.csv.gz</li></ol>

opencc-by-4.0Oct 2023View details →
dryad36/100

Strains and constructs for: A chimeric nuclease substitutes a phage CRISPR-Cas system to provide sequence specific immunity against subviral parasites

Open the record for dataset details and reuse information.

publicJul 2021View details →
zenodo32/100

TA B L E 2 Estimates of pairwise sequence divergence (cyt-b gene) in pale-bellied Micronycteris, where M. minuta is divided in three clades. Below the diagonal: pairwise distance using the Kimura 2-parameter model (percentage). On the diagonal: within-clade distance using the Kimura 2-parameter model (percentage). Above the diagonal: pairwise p-distance values. Number of specimens sequenced in parenthesis. *Chimeric sequence obtained from two paratypes (Siles et al., 2013). in Revision of the pale-bellied Micronycteris Gray, 1866 (Chiroptera, Phyllostomidae) with descriptions of two new species

TA B L E 2 Estimates of pairwise sequence divergence (cyt-b gene) in pale-bellied Micronycteris, where M. minuta is divided in three clades. Below the diagonal: pairwise distance using the Kimura 2-parameter model (percentage). On the diagonal: within-clade distance using the Kimura 2-parameter model (percentage). Above the diagonal: pairwise p-distance values. Number of specimens sequenced in parenthesis. *Chimeric sequence obtained from two paratypes (Siles et al., 2013).

opennotspecifiedJun 2020View details →
dryad32/100

Data from: ITS all right mama: Investigating the formation of chimeric sequences in the ITS2 region by DNA metabarcoding analyses of fungal mock communities of different complexities

The formation of chimeric sequences can create significant methodological bias in PCR-based DNA metabarcoding analyses. During mixed-template amplification of barcoding regions, chimera formation is frequent and well documented. However, profiling of fungal communities typically uses the more variable rDNA region ITS. Due to a larger research community, tools for chimera detection have been developed mainly for the 16S/18S markers. However, these tools are widely applied to the ITS region without verification of their performance. We examined the rate of chimera formation during amplification and 454 sequencing of the ITS2 region from fungal mock communities of different complexities. We evaluated the chimera detecting ability of two common chimera-checking algorithms: Perseus and UCHIME. Large proportions of the chimeras reported were false positives. No false negatives were found in the dataset. Verified chimeras accounted for only 0.2% of the total ITS2 reads, which is considerably less than what is typically reported in 16S and 18S metabarcoding analyses. Verified chimeric "parent sequences" had significantly higher percent identity to one another than to random members of the mock communities. Community complexity increased the rate of chimera formation. GC content was higher around the verified chimeric break points, potentially facilitating chimera formation through base pair mismatching in the neighboring regions of high similarity in the chimeric region. We conclude that the hypervariable nature of the ITS region seem to buffer the rate of chimera formation in comparison to other, less variable barcoding regions, due to shorter regions of high sequence similarity.

opencc-zeroDec 2015View details →
zenodo32/100

Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of Ms4a3Ai14, BM chimeric mice (CD45.2 Csf2rb-/-: CD45.1 Csf2rb+/+ and CD45.2 Ifngr1-/-: CD45.1 Ifngr1+/+) using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.

<p><strong>Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE,&nbsp; BM chimeric mice (CD45.2 <em>Csf2rb</em><sup>-/-</sup>: CD45.1 <em>Csf2rb</em><sup>+/+</sup> and CD45.2 <em>Ifngr1<sup>-/-</sup></em>: CD45.1 <em>Ifngr1<sup>+/+</sup></em>) using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer&#39;s protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger&#39;s in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

APPENDIX. GenBank accession numbers of all DNA sequences of Cophyla used in this study. NA, not applicable. Asterisks mark cases where sequences from different samples were combined to chimeric terminals for analysis. in Description of the lucky Cophyla (Microhylidae, Cophylinae), a new arboreal frog from Marojejy National Park in north-eastern Madagascar

APPENDIX. GenBank accession numbers of all DNA sequences of Cophyla used in this study. NA, not applicable. Asterisks mark cases where sequences from different samples were combined to chimeric terminals for analysis.

opennotspecifiedAug 2019View details →
dryad32/100

Data from: ITS all right mama: Investigating the formation of chimeric sequences in the ITS2 region by DNA metabarcoding analyses of fungal mock communities of different complexities

Open the record for dataset details and reuse information.

publicOct 2016View details →
zenodo28/100

Figure 1 from: Hughes KW, Morris SD, Reboredo-Segovia A (2015) Cloning of ribosomal ITS PCR products creates frequent, non-random chimeric sequences – a test involving heterozygotes between Gymnopus dichrous taxa I and II. MycoKeys 10: 45-56. https://doi.org/10.3897/mycokeys.10.5126

Figure 1 - The ITS2 region of two haplotypes of Gymnopus dichrous: Haplotypes DI (represented by TENN68084) and DII (represented by TENN68078). The TC pair at position 15 and the indel at position 25 were used to determine which the haplotype was represented by the 5' end of a cloned sequence. Bases in red are points where DI and DII haplotypes differ in sequence and were used to determine if template switching had occurred in a cloned PCR product. Eight base pairs at which template switching can be detected are indicated by numbers 1-8. The possible area in which template switching (ts) could have occurred is indicated by vertical arrows and the number of observed template switching events is given above the vertical arrow. Bases that may be involved in intra-strand base pairing as determined by MFOLD are outlined with black boxes. Ambiguity codes indicate intraspecific variation.

opencc-by-4.0Jun 2015View details →
geo24/100

RNA sequencing of primary acute lymphoblastic leukemia xenograft brains with and without treatment with chimeric antigen receptor T (CAR-T) cell therapy

GEO Series GSE121591. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2020View details →
geo24/100

RNA sequencing of CD19-directed chimeric antigen receptor T (CART19) cells with and without TP-0903 treatment

GEO Series GSE199257. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2023View details →
geo24/100

Genome-wide and Cell-type Selective Profiling of In Vivo Small Noncoding RNA:Target RNA Interactions by Chimeric RNA Sequencing

GEO Series GSE263988. Mus musculus. 12 samples. Type: Other.

openGEO-OpenJul 2024View details →
ClinicalTrials.gov24/100

A Study to Evaluate Next-Generation Sequencing (NGS) Testing and Monitoring of B-cell Recovery to Guide Management Following Chimeric Antigen Receptor T-cell (CART) Induced Remission in Children and Y

ClinicalTrials.gov study NCT05621291. IPD Sharing: YES. Countries: 1. Publications: 0.

controlledIPD-YESFeb 2026View details →
geo24/100

Next Generation Sequencing of Chimeric antigen receptor T (CAR-T) cells polarized under Th9 condition

GEO Series GSE156075. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2020View details →
geo24/100

Transcriptome sequencing identify a recurrent CRYL1-IFT88 chimeric transcript in hepatocellular carcinoma

GEO Series GSE97214. Homo sapiens. 18 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2017View details →
geo20/100

Single-cell RNA sequencing analysis of mouse totipotent potential stem cells, extended pluripotent stem cells, embryonic stem cells, E17.5 TPS-derived chimeric placental cells, TPS-derived teratomas a

GEO Series GSE201747. Mus musculus. 11 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2022View details →
geo20/100

ChimeraScan: A tool for identifying chimeric transcription in sequencing data

GEO Series GSE29098. Homo sapiens. 3 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2011View details →
geo20/100

CapID sequencing of 40 chimeric antigen receptor T-cell infusion products were reported

GEO Series GSE150992. Homo sapiens. 40 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2020View details →
geo16/100

mRNA transcriptome sequencing of siglecf low macrophages from tumor lungs of chimeric mice that were intranasally injected with Ad-Cre

GEO Series GSE301096. Mus musculus. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2025View details →
geo16/100

Single-Cell Transcriptome Sequencing for investigating tumor microenvironment remodeling mediated by a chimeric oncolytic adenovirus combined with PD-1 antibody

GEO Series GSE185575. Mus musculus. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record