Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22
datasets available to search
ShareScore release 0.9.0
Dataset results
22 results for “genomic barcoding”
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Fig. 3 in Chloroplast genome of the conserved Aster altaicus var. uchiyamae B2015-0044 as genetic barcode
Fig. 3. The variable sites in the chloroplast genomes of Aster altaicus var. uchiyamae. Variable sequences are marked in red. GG: Yeoju, Gyeonggi Province, CB: Cheongju, Chungcheongbuk Province.
Fig. 2 in Chloroplast genome of the conserved Aster altaicus var. uchiyamae B2015-0044 as genetic barcode
Fig. 2. The sequence alignment of variable sites in the chloroplast genomes of Aster altaicus var. uchiyamae. Variable sequences are marked in red. GG: Yeoju, Gyeonggi Province, CB: Cheongju, Chungcheongbuk Province.
Fig. 2 in Chloroplast genome of white wild chrysanthemum, Dendranthema sp. K247003, as genetic barcode
Fig. 2. Comparison of chloroplast genomes of Dendranthema sp. K247003 and D. boreale IT121002 using mVISTA program. Grey arrows and thick black lines above the alignment indicate genes with their orientation and the position of the IRs, respectively. The Yscale represents the percent identity between 50-100%. Genome regions are color-coded: Coding regions in blue; noncoding sequences (CNS) in red.
Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns (repository for Genome Research paper, 2022)
<p>Simulated ONT and PacBio RNA-Seq data for "Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns" paper (Mikheenko et al., Genome Research, 2022). All details can be found in the Methods section of the paper.</p> <p><strong>PacBio.simulated_uniform_coverage.fasta.gz</strong> and <strong>ONT.simulated_uniform_coverage.fasta.gz files</strong> were used in Supplemental Note “Benchmarking of the read-to-isoform assignment algorithm”.</p> <p><strong>ONT.simulated_real_expression.fasta.gz</strong> file and all GTF files were used in the Section "Splice site correction improves transcript discovery precision". <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> was used as the annotation file for all tools. <strong>mouse.gencode.M26.spatial.15percent.expressed.gtf </strong>contains the set of all expressed isoforms. <strong>mouse.gencode.M26.spatial.15percent.expressed_kept.gtf</strong> contains those of the isoforms that are in presented in the annotation file ("known" transcripts), <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> contains expressed isoforms that were removed from the annotation ("novel" transcripts).</p>
Data and processing scripts for PRISM barcode sequencing data used in "Massively parallel pooled screening reveals genomic determinants of nanoparticle-cell interactions"
<p>Sequencing data for the PRISM barcodes generated after nano-particle treatment is presented in this repository alongside the code to process the sequencing counts to generate the binning probabilities and weighted scores. <br> <br> For the details please see the original publication or the bioarxiv preprint: https://doi.org/10.1101/2021.04.05.438521<br> <br> The raw data is provided in PILOT_DATA_COUNTS.csv and EXPERIMENT_DATA_COUNTS.csv files, for the pilot and the actual experiment. <br> <br> For each of these files an R script is provided to process them, along with the output of the scripts (PILOT_DATA_PROBABILITIES.csv and EXPERIMENT_DATA_PROBABILITIES.csv)</p>
Fig. 1 in Chloroplast genome of the conserved Aster altaicus var. uchiyamae B2015-0044 as genetic barcode
Fig. 1. Plastid genomic map of Aster altaicus var. uchiyamae.
Fig. 1 in Chloroplast genome of white wild chrysanthemum, Dendranthema sp. K247003, as genetic barcode
Fig. 1. Plastid genomic map of Dendranthema sp. K247003.
Data for paper: Genomic data reveals new species and the limits of mtDNA barcode diagnostics to contain a global pest species complex (Diptera: Tephritidae: Dacinae)
<p>Files in this repository:</p><p>"COI_alignment.fas.zip" Zipped file of the FASTA alignment of the COI sequences.</p><p>"COI_IQtree.treefile" Newick treefile resulting from the IQ-tree analysis of the COI alignment.</p><p>"RAD-loci_alignment.nex.zip" Zipped file of the NEXUS alignment of RAD loci of 2295 samples.</p><p>"RAD-loci_IQtree.tre" Newick treefile resulting from the IQ-tree analysis of the RAD-loci alignment.</p><p>"RAD-SNP_alignment.usnps.nex" NEXUS alignment of the RAD-SNP data of 50 samples.</p><p>"RAD-SNP_SNAPP.trees" Set of Newick trees resulting from the BEAST SNAPP analysis.</p>
Rapid and Inexpensive Whole-Genome Sequencing of SARS-CoV2 using 1200 bp Tiled Amplicons and Oxford Nanopore Rapid Barcoding
<p>Description of 1200bp amplicon primer sets and .bed and .tsv files for SARS-CoV-2 assembly using the ARTIC bioinformatics pipeline.</p>
Supplementary material 1 from: Zúñiga JD, Gostel MR, Mulcahy DG, Barker K, Hill A, Sedaghatpour M, Vo SQ, Funk VA, Coddington JA (2017) Data Release: DNA barcodes of plant species collected for the Global Genome Initiative for Gardens Program, National Museum of Natural History, Smithsonian Institution. PhytoKeys 88: 119-122. https://doi.org/10.3897/phytokeys.88.14607
List of samples collected for the Global Genome Initiative for Gardens project selected for DNA barcoding, with GenBank accession numbers and genetic sample identification numbers. All the sequences are included in the GGI-Gardens BioProject. : Explanation note: List of samples collected for the Global Genome Initiative for Gardens project selected for DNA barcoding, with GenBank accession numbers and genetic sample identification numbers.
Data from: Deep sequencing of mixed total DNA without barcodes allows efficient assembly of highly plastic ascidian mitochondrial genomes
Open the record for dataset details and reuse information.
Data from: Practical low-coverage genomewide sequencing of hundreds of individually barcoded samples for population and evolutionary genomics in nonmodel species
Today most population genomic studies of nonmodel organisms either sequence a subset of the genome deeply in each individual or sequence pools of unlabelled individuals. With a step-by-step workflow, we illustrate how low-coverage whole-genome sequencing of hundreds of individually barcoded samples is now a practical alternative strategy for obtaining genomewide data on a population scale. We used a highly efficient protocol to generate high-quality libraries for ~6.5 USD from each of 876 Atlantic silversides (a teleost fish with a genome size ~730 Mb) that we sequenced to 1–4× genome coverage. In the absence of a reference genome, we developed a bioinformatic pipeline for mapping the genomic reads to a de novo assembled reference transcriptome. This provides an 'in silico' method for exome capture that avoids the complexities and expenses of using wet chemistry for target isolation. Using novel tools for analysis of low-coverage data, we extracted population allele frequencies, individual genotype likelihoods and polymorphism data for 2 504 335 SNPs across the exome for the 876 fish. To illustrate the use of the resulting data, we present a preliminary analysis of geographical patterns in the exome data and a comparison of complete mitochondrial genome sequences for each individual (constructed from the low-coverage data) that show population colonization patterns along the US east coast. With a total cost per sample of less than 50 USD (including sequencing) and ability to prepare 96 libraries in only 5 h, our approach adds a viable new option to the population genomics toolbox.
Data from: Practical low-coverage genomewide sequencing of hundreds of individually barcoded samples for population and evolutionary genomics in nonmodel species
Open the record for dataset details and reuse information.
CellTag Indexing: genetic barcode-based sample multiplexing for single-cell genomics
GEO Series GSE130065. Mus musculus; Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.
Human lineage tracing enabled by mitochondrial mutations and single cell genomics [TF1_barcoding_scRNA]
GEO Series GSE118203. Homo sapiens. 384 samples. Type: Expression profiling by high throughput sequencing.
A barcoded genome-scale library of inducible alleles reveals principles of synthetic gene control
GEO Series GSE158319. Saccharomyces cerevisiae. 357 samples. Type: Expression profiling by array; Expression profiling by high throughput sequencing; Other.
Human lineage tracing enabled by mitochondrial mutations and single cell genomics [TF1_barcoding_ATAC]
GEO Series GSE118202. Homo sapiens. 1 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Slide-tags: scalable, single-nucleus barcoding for multi-modal spatial genomics
GEO Series GSE244355. Mus musculus. 2 samples. Type: Expression profiling by high throughput sequencing; Other.
Investigation of spatial vulnerabilities of Bacterium Escherichia coli genome to spontaneous mutations by molecularly barcoded deep Sequencing
GEO Series GSE116453. Escherichia coli ATCC 8739. 1 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.