Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

86

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

86 results for “sequence tagging”

Learn how ShareScore rates datasets ↗
zenodo40/100

Fig. 4 in Expressed sequence tags in venomous tissue of Scorpaena plumieri (Scorpaeniformes: Scorpaenidae)

Fig. 4. Sequence alignment of putative lectin from Scorpaena plumieri. Alignment of a lectin-like EST in silico translated sequence from S. plumieri (ClustalW2 EBI) with fish-egg lectin from Oplegnathus fasciatus (BAL618145), Dicentrarchus labrax (CBK52298), Maylandia zebra (XP_004574029), and Oreochromis niloticus (XP003443389). The recombinant clone was isolated with antibody fraction derived from S. plumieri venom. * identifies and identical residue;: identifies a conserved residue. Underlined residues represent invariable sites, underlined IRLS = N-acetylation site.

opencc-by-4.0Oct 2014View details →
zenodo40/100

Fig. 3 in Expressed sequence tags in venomous tissue of Scorpaena plumieri (Scorpaeniformes: Scorpaenidae)

Fig. 3. The classification of EST from Scorpaena plumieri based on their putative fractions. Three-hundred fifty-six EST edited sequences were initially analyzed with Blast and Swiss protein databanks. The consensus sequence was attributed a function based on the strongest match.

opencc-by-4.0Oct 2014View details →
zenodo40/100

Fig. 2 in Expressed sequence tags in venomous tissue of Scorpaena plumieri (Scorpaeniformes: Scorpaenidae)

Fig. 2. Agarose gel electrophoresis of DNA isolated from clones. White colonies containing insert were grown and the plasmidial DNA isolated and digested with EcoRI enzyme. An aliquot from each clone (1-27) was electrophoresed on 1% Agarose gel and stained with ethidium bromide.

opencc-by-4.0Oct 2014View details →
zenodo40/100

Fig. 1 in Expressed sequence tags in venomous tissue of Scorpaena plumieri (Scorpaeniformes: Scorpaenidae)

Fig. 1. Agarose- formaldehyde electrophoresis of RNA from Scorpaena plumieri. A) 1) 2 µg of E. coli tRNA; 2) 2 µg de rRNA de Rattus norvegicus; 3) and 4) 2 µg total RNA from S. plumieri spine gland. B) 1) 2 µg de total RNA from S. plumieri; 2) the same sample incubated 2 h a 37ºC before electrophoresis.

opencc-by-4.0Oct 2014View details →
zenodo36/100

A genomic data set of single‐nucleotide polymorphisms (SNPs) generated by ddRAD tag sequencing in Q. petraea (Matt.) Liebl. populations from Central-Eastern Europe and Balkan Peninsula

<p>This genomic dataset provides highly variable single-nucleotide polymorphism&nbsp;(SNP) markers from georeferenced natural <em>Quercus petraea</em> (Matt.) Liebl. populations collected in Bulgaria, Hungary, Romania, Serbia, Bosnia and Herzegovina, Kosovo and Albania. These SNP loci can be used to assess genetic diversity, differentiation, population structure, and can also be used to detect signatures of selection and local adaptation.</p>

opencc-by-4.0Jun 2020View details →
dryad36/100

The topological nature of tag jumping in environmental DNA metabarcoding studies (sequencing raw data)

<p>Metabarcoding of environmental DNA constitutes a state-of-the-art tool for environmental studies. One fundamental principle implicit in most metabarcoding studies is that individual sample amplicons can still be identified after being pooled with others – based on their unique combinations of tags – during the so-called demultiplexing step that follows sequencing. Nevertheless, it has been recognized that tags can sometimes be changed (i.e. tag jumping), which ultimately leads to sample crosstalk. Here, using four DNA metabarcoding datasets derived from the analysis of soils and sediments, we show that tag jumping follows very specific and systematic patterns. Specifically, we find a strong correlation between the number of reads in blank samples and their topological position in the tag matrix (described by vertical and horizontal vectors). This observed spatial pattern of artefactual sequences could be explained by polymerase activity, which leads to the exchange of the 3' tag of single stranded tagged sequences through the formation of heteroduplexes with mixed barcodes. Importantly, tag jumping substantially distorted our datasets – despite our use of methods suggested to minimize this error. We developed a topologic model to estimate the noise based on the counts in our blanks, which suggested that 40-80% of the taxa in our soil and sedimentary samples were likely false positives introduced through tag jumping. We highlight that the amount of false positive detections caused by tag jumping strongly biased our community analyses. </p>

opencc-zeroNov 2022View details →
dryad36/100

The topological nature of tag jumping in environmental DNA metabarcoding studies (sequencing raw data)

Open the record for dataset details and reuse information.

publicNov 2022View details →
zenodo32/100

Sequence Tagging of FA-KES Dataset

<p>We used the BIOE Sequence Tagging&nbsp;strategy which was utilized in OpenTag [1] in the aim of getting every word in the dataset<br> associated with a label called &lsquo;tag&rsquo;. A tag consists of one of these letters B, I, O, or E, that stand respectively for beginning, inside, outside, or end of an attribute, followed by a &lsquo;-&rsquo; sign, followed by three letters that represent the type of information that was initially&nbsp;extracted. Tokens&nbsp;were labeled with one of the following tags: &lsquo;B-LOC&rsquo;, &lsquo;I-LOC&rsquo;, &lsquo;E-LOC&rsquo;, &lsquo;B-CIV&rsquo;, &lsquo;I-CIV&rsquo;, &lsquo;E-CIV&rsquo;, &lsquo;B-NCV&rsquo;, &lsquo;I-NCV&rsquo;, &lsquo;E-NCV&rsquo;, &lsquo;B-WMN&rsquo;, &lsquo;I-WMN&rsquo;, &lsquo;E-WMN&rsquo;, &lsquo;B-CHD&rsquo;, &lsquo;I-CHD&rsquo;, &lsquo;E-CHD&rsquo;, &lsquo;B-ACT&rsquo;, &lsquo;I-ACT&rsquo;,&lsquo;E-ACT&rsquo;, &lsquo;B-COD&rsquo;, &lsquo;I-COD&rsquo;, &lsquo;E-COD&rsquo;, &lsquo;B-DAT&rsquo;, &lsquo;I-DAT&rsquo;, &lsquo;E-DAT&rsquo;, or &lsquo;O&rsquo; (where O stands for words outside the scope, LOC for&nbsp;the incident location, CIV for the number of civilians dead, NCV for the number of non-civilians dead, WMN for the number of women targeted, CHD for the number of children killed, ACT for actor/authority responsible for the incident, COD for the cause of death, and DAT for date of incident). This was done by creating a parser that would automatically tag each word with the appropriate tag.</p> <p>We created three subsets of the FA-KES dataset. The first one consists of the articles&#39; titles, the second one&nbsp;of the articles&#39; titles concatenated with the articles&#39; first paragraphs, and the third one consists of the articles&#39; titles along with their contents. The first column in each of the three CSV files&nbsp;represents the article number in the dataset, the second column contains the sequence of&nbsp;words for each article and the third one holds&nbsp;the tags linked to&nbsp;the tokens&nbsp;of&nbsp;the previous column.</p> <p>[1]:&nbsp;G. Zheng, S. Mukherjee, X. L. Dong, and F. Li, &ldquo;Opentag: Open attribute value extraction from product profiles,&rdquo; CoRR, vol. abs/1806.01264, 2018.</p>

opencc-by-4.0Nov 2019View details →
dryad32/100

Data from: Exploitation of a turbot (Scophthalmus maximus L.) immune-related expressed sequence tag (EST) database for microsatellite screening and validation

In this study, we identified and characterized 160 microsatellite loci from an expressed sequence tag (EST) database generated from immune-related organs of turbot (Scophthalmus maximus). A final set of 83 new polymorphic microsatellites were validated after the analysis of 40 individuals from Atlantic origin including both wild and farmed individuals. The allele number and the expected heterozygosity ranged from 2 to 18 and from 0.021 to 0.951, respectively. Evidences of null alleles at moderate-high frequencies were detected at six loci using population data. None of the analyzed loci showed deviations from Mendelian segregation after analysis of five full-sib families including ~92 individuals/family. The markers are used to consolidate the turbot genetic map and, since they are mostly EST-derived, they will be very useful for comparative genomic studies within flatfishes and with model fish species. Using an in silico approach, we detected significant homologies of microsatellite sequences with the EST databases of the flatfish species with highest genomic resources (Senegalese sole, Atlantic halibut, bastard halibut) at 31% of these turbot markers. The conservation of these microsatellites within Pleuronectiformes will pave the way for anchoring genetic maps of different species and identifying genomic regions related to productive traits.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Parallel tagged next-generation sequencing on pooled samples – a new approach for population genetics in ecology and conservation

Next-generation sequencing (NGS) on pooled samples has already been broadly applied in human medical diagnostics and plant and animal breeding. However, thus far it has been only sparingly employed in ecology and conservation, where it may serve as a useful diagnostic tool for rapid assessment of species genetic diversity and structure at the population level. Here we undertake a comprehensive evaluation of the accuracy, practicality and limitations of parallel tagged amplicon NGS on pooled population samples for estimating species population diversity and structure. We obtained 16S and Cyt b data from 20 populations of Leiopelma hochstetteri, a frog species of conservation concern in New Zealand, using two approaches – parallel tagged NGS on pooled population samples and individual Sanger sequenced samples. Data from each approach were then used to estimate two standard population genetic parameters, nucleotide diversity (π) and population differentiation (FST), that enable population genetic inference in a species conservation context. We found a positive correlation between our two approaches for population genetic estimates, showing that the pooled population NGS approach is a reliable, rapid and appropriate method for population genetic inference in an ecological and conservation context. Our experimental design also allowed us to identify both the strengths and weaknesses of the pooled population NGS approach and outline some guidelines and suggestions that might be considered when planning future projects.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Species tree estimation of North American chorus frogs (Hylidae: Pseudacris) with parallel tagged amplicon sequencing

The field of phylogenetics is changing rapidly with the application of high-throughput sequencing to non-model organisms. Cost-effective use of this technology for phylogenetic studies, which often include a relatively small portion of the genome but several taxa, requires strategies for genome partitioning and sequencing multiple individuals in parallel. In this study we estimated a multilocus phylogeny for the North American chorus frog genus Pseudacris using anonymous nuclear loci that were recently developed using a reduced representation library approach. We sequenced 27 nuclear loci and three mitochondrial loci for 44 individuals on 1/3 of an Illumina MiSeq run, obtaining 96.5% of the targeted amplicons at less than 20% of the cost of traditional Sanger sequencing. We found heterogeneity among gene trees, although four major clades (Trilling Frog, Fat Frog, crucifer, and West Coast) were consistently supported, and we resolved the relationships among these clades for the first time with strong support. We also found discordance between the mitochondrial and nuclear datasets that we attribute to mitochondrial introgression and a possible selective sweep. Bayesian concordance analysis in BUCKy and species tree analysis in *BEAST produced largely similar topologies, although we identify taxa that require additional investigation in order to clarify taxonomic and geographic range boundaries. Overall, we demonstrate the utility of a reduced representation library approach for marker development and parallel tagged sequencing on an Illumina MiSeq for phylogenetic studies of non-model organisms.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Extending RAD tag analysis to microbial ecology: a comparison between multi locus sequence typing (MLST) and 2b-RAD to investigate Listeria monocytogenes genetic structure

The advent of next-generation sequencing (NGS) has dramatically changed bacterial typing technologies, increasing our ability to differentiate bacterial isolates. Despite it is now possible to sequence a bacterial genome in a few days and at reasonable costs, most genetic analyses do not require whole-genome sequencing, which also remains impractical for large population samples due to the cost of individual library preparation and bioinformatics. More traditional sequencing approaches, however, such as MultiLocus Sequence Typing (mlst) are quite laborious and time-consuming, especially for large-scale analyses. In this study, a genotyping approach based on restriction site-associated (RAD) tag sequencing, 2b-RAD, was applied to characterize Listeria monocytogenes strains. To verify the feasibility of the method, an in silico analysis was performed on 30 available complete genomes. For the same set of strains, in silico mlst analysis was conducted as well. Subsequently, 2b-RAD and mlst analyses were experimentally carried out on 58 isolates collected from food samples or food-processing sites. The obtained results demonstrate that 2b-RAD predicts mlst types and often provides more detailed information on population structure than mlst. Moreover, the majority of variants differentiating identical sequence type isolates mapped against accessory fragments, thus providing additional information to characterize strains. Although mlst still represents a reliable typing method, large-scale studies on molecular epidemiology and public health, as well as bacterial phylogenetics, population genetics and biosafety could benefit of a low cost and fast turnaround time approach such as the 2b-RAD analysis proposed here.

opencc-zeroDec 2014View details →
dryad32/100

Obovaria olivaria maf filtered vcf file from: RAD-tag and mitochondrial DNA sequencing reveal the genetic structure of a widespread and regionally imperiled freshwater mussel, Obovaria olivaria (Bivalvia: Unionidae)

<p><em>Obovaria olivaria</em> is a species of freshwater mussel native to the Mississippi River and Laurentian Great Lakes-St. Lawrence River drainages of North America. This mussel has experienced population declines across large parts of its distribution and is imperiled in many jurisdictions. <em>Obovaria olivaria </em>uses the similarly imperiled <em>Acipenser fulvescens</em> (Lake Sturgeon) as a host for its glochidia. We employed mitochondrial DNA sequencing and Restriction-site Associated DNA sequencing (RAD-seq) to assess patterns of genetic diversity and population structure of <em>O. olivaria</em> from 19 collection locations including the St. Lawrence River drainage, the Great Lakes drainage, the Upper Mississippi River drainage, the Ohioan River drainage and the Mississippi Embayment. Heterozygosity was highest in Upper Mississippi and Great Lakes populations, followed by a reduction in diversity and relative effective population size in the St. Lawrence populations. Pairwise <em>F</em><sub>ST</sub> ranged from 0.00 to 0.20, and analyses of genetic structure revealed two major ancestral populations, one including all St. Lawrence River/Ottawa River sites and the other including remaining sites; however, significant admixture and isolation by river distance across the range were evident. The genetic diversity and structure of <em>O. olivaria</em> is consistent with the existing literature on <em>Acipenser fulvescens</em> and suggest that, although northern and southern <em>O. olivaria</em> populations are genetically distinct, genetic structure in <em>O. olivaria</em> is largely clinal rather than discrete across its range. Conservation and restoration efforts of <em>O. olivaria</em> should prioritize the maintenance and restoration of locations where <em>O. olivaria </em>remain, especially in northern rivers, and to ensure connectivity that will facilitate dispersal of <em>Acipenser fulvescens</em> and movement of encysted glochidia.</p>

opencc-zeroFeb 2024View details →
dryad32/100

Data from: Parallel tagged amplicon sequencing reveals major lineages and phylogenetic structure in the North American tiger salamander (Ambystoma tigrinum) species complex

Open the record for dataset details and reuse information.

publicAug 2012View details →
dryad32/100

Data from: Species tree estimation of North American chorus frogs (Hylidae: Pseudacris) with parallel tagged amplicon sequencing

Open the record for dataset details and reuse information.

publicFeb 2015View details →
dryad32/100

Data from: Parallel tagged amplicon sequencing of transcriptome-based genetic markers for Triturus newts with the Ion Torrent next-generation sequencing platform

Open the record for dataset details and reuse information.

publicFeb 2014View details →
dryad32/100

Data from: Parallel tagged next-generation sequencing on pooled samples – a new approach for population genetics in ecology and conservation

Open the record for dataset details and reuse information.

publicMay 2013View details →
dryad32/100

Data from: Extending RAD tag analysis to microbial ecology: a comparison between multi locus sequence typing (MLST) and 2b-RAD to investigate Listeria monocytogenes genetic structure

Open the record for dataset details and reuse information.

publicNov 2015View details →
dryad32/100

Data from: Exploitation of a turbot (Scophthalmus maximus L.) immune-related expressed sequence tag (EST) database for microsatellite screening and validation

Open the record for dataset details and reuse information.

publicJan 2012View details →
dryad32/100

Obovaria olivaria maf filtered vcf file from: RAD-tag and mitochondrial DNA sequencing reveal the genetic structure of a widespread and regionally imperiled freshwater mussel, Obovaria olivaria (Bivalvia: Unionidae)

Open the record for dataset details and reuse information.

publicFeb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record