Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
86
datasets available to search
ShareScore release 0.9.0
Dataset results
86 results for “sequence tagging”
Data from: Mining microsatellite markers from public expressed sequence tags databases for the study of threatened plants
Background: Simple Sequence Repeats (SSRs) are widely used in population genetic studies but their classical development is costly and time-consuming. The ever-increasing available DNA datasets generated by high-throughput techniques offer an inexpensive alternative for SSRs discovery. Expressed Sequence Tags (ESTs) have been widely used as SSR source for plants of economic relevance but their application to non-model species is still modest. Methods: Here, we explored the use of publicly available ESTs (GenBank at the National Center for Biotechnology Information-NCBI) for SSRs development in non-model plants, focusing on genera listed by the International Union for the Conservation of Nature (IUCN). We also search two model genera with fully annotated genomes for EST-SSRs, Arabidopsis and Oryza, and used them as controls for genome distribution analyses. Overall, we downloaded 16 031 555 sequences for 258 plant genera which were mined for SSRsand their primers with the help of QDD1. Genome distribution analyses in Oryza and Arabidopsis were done by blasting the sequences with SSR against the Oryza sativa and Arabidopsis thaliana reference genomes implemented in the Basal Local Alignment Tool (BLAST) of the NCBI website. Finally, we performed an empirical test to determine the performance of our EST-SSRs in a few individuals from four species of two eudicot genera, Trifolium and Centaurea. Results: We explored a total of 14 498 726 EST sequences from the dbEST database (NCBI) in 257 plant genera from the IUCN Red List. We identify a very large number (17 102) of ready-to-test EST-SSRs in most plant genera (193) at no cost. Overall, dinucleotide and trinucleotide repeats were the prevalent types but the abundance of the various types of repeat differed between taxonomic groups. Control genomes revealed that trinucleotide repeats were mostly located in coding regions while dinucleotide repeats were largely associated with untranslated regions. Our results from the empirical test revealed considerable amplification success and transferability between congenerics. Conclusions: The present work represents the first large-scale study developing SSRs by utilizing publicly accessible EST databases in threatened plants. Here we provide a very large number of ready-to-test EST-SSR (17 102) for 193 genera. The cross-species transferability suggests that the number of possible target species would be large. Since trinucleotide repeats are abundant and mainly linked to exons they might be useful in evolutionary and conservation studies. Altogether, our study highly supports the use of EST databases as an extremely affordable and fast alternative for SSR developing in threatened plants.
Data from: Mining for single nucleotide polymorphisms and insertions / deletions in expressed sequence tag libraries of oil palm
The oil palm is a tropical oil bearing tree. Recently EST-derived SNPs and SSRs are a free by-product of the currently expanding EST (Expressed Sequence Tag) data bases. The development of high-throughput methods for the detection of SNPs (Single Nucleotide Polymorphism) and small indels (insertion / deletion) has led to a revolution in their use as molecular markers. Available (5452) Oil palm EST sequences were mined from dbEST of NCBI. CAP3 program was used to assemble EST sequences into contigs. Candidate SNPs and Indel polymorphisms were detected using the perl script auto_snip version 1.0 which has used 576 ESTs for detecting SNPs and Indel sites. We found 1180 SNP sites and 137 indel polymorphisms with frequency 1.36 SNPs / 100 bp. Among the six tissues from which the EST libraries had been generated, mesocarp had high frequency of 2.91 SNPs and indels per 100 bp whereas the zygotic embryos had lowest frequency of 0.15 per 100 bp. We also used the Shannon index to analyze the proportion of ten possible types of SNP/indels. ESTs from tissues of normal apex showed highest values of Shannon index (0.60) whereas abnormal apex had least value (0.02). The present report deals the use of Shannon index for comparing SNP/ indel frequencies mined from ESTlibraries and also confirm that the frequency of SNP occurrence in oil palm to use them as markers for genetic studies.
Data from: Not all sequence tags are created equal: designing and validating sequence identification tags robust to indels
Ligating adapters with unique synthetic oligonucleotide sequences (sequence tags) onto individual DNA samples before massively parallel sequencing is a popular and efficient way to obtain sequence data from many individual samples. Tag sequences should be numerous and sufficiently different to ensure sequencing, replication, and oligonucleotide synthesis errors do not cause tags to be unrecoverable or confused. However, many design approaches only protect against substitution errors during sequencing and extant tag sets contain too few tag sequences. We developed an open-source software package to validate sequence tags for conformance to two distance metrics and design sequence tags robust to indel and substitution errors. We use this software package to evaluate several commercial and non-commercial sequence tag sets, design several large sets (maxcount=7,198) of edit metric sequence tags having different lengths and degrees of error correction, and integrate a subset of these edit metric tags to polymerase chain reaction (PCR) primers and sequencing adapters. We validate a subset of these edit metric tagged PCR primers and sequencing adapters by sequencing on several platforms and subsequent comparison to commercially available alternatives. We find that several commonly used sets of sequence tags or design methodologies used to produce sequence tags do not meet the minimum expectations of their underlying distance metric, and we find that PCR primers and sequencing adapters incorporating edit metric sequence tags designed by our software package perform as well as their commercial counterparts. We suggest that researchers evaluate sequence tags prior to use or evaluate tags that they have been using. The sequence tag sets we design improve on extant sets because they are large, valid across the set, and robust to the suite of substitution, insertion, and deletion errors affecting massively parallel sequencing workflows on all currently used platforms.
Data from: Mining of expressed sequence tag libraries of cacao for microsatellite markes using five computational tools
Expressed Sequence Tags (ESTs) provide researchers with a quick and inexpensive route for discovering new genes, and data on gene expression and regulation and provide genic markers that help in constructing genome maps. Cacao is an important perennial crop of humid tropics. Cacao EST sequences as available in public domain were downloaded and made into contigs. A total of 769 contigs were made using contigs assembly program pharp. Puative information of contigs were identified using NCBI and ExPASy tools such as BlastX, tblastn.
Data from: Analysis of expressed sequence tags from the placenta of the live-bearing fish Poeciliopsis (Poeciliidae)
Matrotrophic fish in the genus Poeciliopsis (Poeciliidae) have a placenta-like structure used in post-fertilization maternal provisioning of the developing embryo. To understand better the structure and function of the Poeciliopsis placenta we derived cDNA libraries from the maternal follicular placenta of two matrotrophic Poeciliopsis sister species, P. turneri and P. presidionis. These species inherited their placenta from a common ancestor and represent one of three independent origins of placentas in Poeciliopsis. Expressed sequence tags were generated and putative function was determined using BLASTX homology searches and Gene Ontology annotation. Reverse transcriptase-PCR was used to verify placenta tissue expression of a putative candidate gene, alpha-2 macroglobulin. 1956 (71.5% of the total submitted ESTs) and 924 (71.0% of the total submitted ESTs) unique transcripts were identified for the P. turneri and P. presidionis placenta, respectively. Homology search and Gene Ontology annotation revealed putative genes whose products may be involved in specific transport functions of the maternal follicle. These putative genes are excellent candidates for future research on the evolution of the placenta. We discuss our results in light of the parent-offspring conflict theory of placental evolution and in terms of the Poeciliid placenta structure and function.
Data from: "NGS based generation of expressed sequence tags for Lymantria dispar and Lymantria monacha, two closely related lepidopteran species with different responses to parasitism by Glyptapanteles liparidis" in Genomic Resources Notes accepted 1 December 2013 to 31 January 2014
Introduction: The gypsy moth, Lymantria dispar, and the nun moth, Lymantria monacha, are closely related species (Lepidoptera, Lymantriidae), co-seasonal and economically important forest pests on broadleaf and coniferous trees. In Central Europe, gypsy moth larvae are frequently parasitized by the gregarious, endoparasitic wasp Glyptapanteles liparidis (Hymenoptera, Braconidae). At oviposition, the female wasp injects between 10 and up to 100 eggs into the hemocoel of a single host larva, together with venom and calyx fluid containing polydnavirus (PDV) particles that subsequently play a critical role in suppressing the host immune response so that successful development of the parasitoid can proceed (Schopf 2007). These viruses, which are integrated in the genomic DNA of the wasp and undergo replication only in the female's ovary, rapidly enter host hemocytes, fat body, and nervous system following parasitization, and viral genes are expressed. In L. dispar larvae parasitized by G. liparidis, the host's hemocytes alter their behavior, fail to spread properly (thereby inhibiting the encapsulation response) and partly undergo programmed cell death (apoptosis), resulting in a dramatic drop in the host's total hemocyte number (Schafellner and Schläger 2009).
Data from: Mining for single nucleotide polymorphisms and insertions / deletions in expressed sequence tag libraries of oil palm
Open the record for dataset details and reuse information.
Data from: Mining of expressed sequence tag libraries of cacao for microsatellite markes using five computational tools
Open the record for dataset details and reuse information.
Data from: "NGS based generation of expressed sequence tags for Lymantria dispar and Lymantria monacha, two closely related lepidopteran species with different responses to parasitism by Glyptapanteles liparidis" in Genomic Resources Notes accepted 1 December 2013 to 31 January 2014
Open the record for dataset details and reuse information.
Data from: Not all sequence tags are created equal: designing and validating sequence identification tags robust to indels
Open the record for dataset details and reuse information.
Data from: Mining microsatellite markers from public expressed sequence tags databases for the study of threatened plants
Open the record for dataset details and reuse information.
Data from: Tag jumps illuminated – reducing sequence-to-sample misidentifications in metabarcoding studies
Open the record for dataset details and reuse information.
Data from: Investigating the genetics of Bti resistance using mRNA tag sequencing: application on laboratory strains and natural populations of the dengue vector Aedes aegypti
Open the record for dataset details and reuse information.
Data from: Analysis of expressed sequence tags from the placenta of the live-bearing fish Poeciliopsis (Poeciliidae)
Open the record for dataset details and reuse information.
Data from: Viral tagging reveals discrete populations in Synechococcus viral genome sequence space
Open the record for dataset details and reuse information.
Data from: Sequencing degraded DNA from non-destructively sampled museum specimens for RAD-tagging and low-coverage shotgun phylogenetics
Open the record for dataset details and reuse information.
Data from: Parallel tagged amplicon sequencing of relatively long PCR products using the Illumina HiSeq platform and transcriptome assembly
Open the record for dataset details and reuse information.
Integrated detection of both 5-mC and 5-hmC by high-throughput tag sequencing technology highlights methylation reprogramming of bivalent genes during cellular differentiation [DGE]
GEO Series GSE40953. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.
RNA-sequencing in mESCs with overexpression of mCherry-tagged DUX and DUX domain derivatives
GEO Series GSE224296. Mus musculus. 22 samples. Type: Expression profiling by high throughput sequencing.
EZH2 inhibition remodels the inflammatory senescence-associated secretory phenotype to potentiate pancreatic cancer immune surveillance [CUT&Tag sequencing for H3K27me3]
GEO Series GSE221417. Mus musculus. 13 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.