Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

96

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

96 results for “nucleotide sequences”

Learn how ShareScore rates datasets ↗
dryad36/100

Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

Genotyping-by-sequencing Single-nucleotide Polymorphism Dataset for Corynorhinus rafinesquii (CORA) and Myotis austroriparius (MYAU)

Open the record for dataset details and reuse information.

publicAug 2024View details →
zenodo32/100

III Average nucleotide distances (%) based on the Kimura 2-parameter (K2P) model between Aselliscus spp., and associated outgroups based on complete mitochondrial Cytb (1,140 bp, below the diagonal) and COI (657 bp, above the diagonal) gene sequences in Description of a new species of the genus Aselliscus (Chiroptera, Hipposideridae) from Vietnam

III Average nucleotide distances (%) based on the Kimura 2-parameter (K2P) model between Aselliscus spp., and associated outgroups based on complete mitochondrial Cytb (1,140 bp, below the diagonal) and COI (657 bp, above the diagonal) gene sequences

opennotspecifiedNov 2015View details →
dryad32/100

Data from: Characterization of the transcriptome, nucleotide sequence polymorphism, and natural selection in the desert adapted mouse Peromyscus eremicus

As a direct result of intense heat and aridity, deserts are thought to be among the most harsh of environments, particularly for their mammalian inhabitants. Given that osmoregulation can be challenging for these animals, with failure resulting in death, strong selection should be observed on genes related to the maintenance of water and solute balance. One such animal, Peromyscus eremicus, is native to the desert regions of the southwest United States and may live its entire life without oral fluid intake. As a first step toward understanding the genetics that underlie this phenotype, we present a characterization of the P. eremicus transcriptome. We assay four tissues (kidney, liver, brain, testes) from a single individual and supplement this with population level renal transcriptome sequencing from 15 additional animals. We identified a set of transcripts undergoing both purifying and balancing selection based on estimates of Tajima's D. In addition, we used the branch-site test to identify a transcript—Slc2a9, likely related to desert osmoregulation—undergoing enhanced selection in P. eremicus relative to a set of related non-desert rodents.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Development of nuclear microsatellite loci and mitochondrial single nucleotide polymorphisms for the natterjack toad, Bufo (Epidalea) calamita (Bufonidae), using next generation sequencing and Competitive Allele Specific PCR (KASPar)

Amphibians are undergoing a major decline worldwide and the steady increase in the number of threatened species in this particular taxa highlights the need for conservation genetics studies using high-quality molecular markers. The natterjack toad, Bufo (Epidalea) calamita, is a vulnerable pioneering species confined to specialized habitats in Western Europe. To provide efficient and cost-effective genetic resources for conservation biologists, we developed and characterized 22 new nuclear microsatellite markers using next-generation sequencing. We also used sequence data acquired from Sanger sequencing to develop the first mitochondrial markers for KASPar assay genotyping. Genetic polymorphism was then analyzed for 95 toads sampled from 5 populations in France. For polymorphic microsatellite loci, number of alleles and expected heterozygosity ranged from 2 to 14 and from 0.035 to 0.720, respectively. No significant departures from panmixia were observed (mean multilocus F IS = −0.015) and population differentiation was substantial (mean multilocus F ST = 0.222, P < 0.001). From a set of 18 mitochondrial SNPs located in the 16S and D-loop region, we further developed a fast and cost-effective SNP genotyping method based on competitive allele-specific PCR amplification (KASPar). The combination of allelic states for these mitochondrial DNA SNP markers yielded 10 different haplotypes, ranging from 2 to 5 within populations. Populations were highly differentiated (G ST = 0.407, P < 0.001). These new genetic resources will facilitate future parentage, population genetics and phylogeographical studies and will be useful for both evolutionary and conservation concerns, especially for the set-up of management strategies and the definition of distinct evolutionary significant units.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages

Molecular phylogenetic studies of homologous sequences of nucleotides often assume that the underlying evolutionary process was globally stationary, reversible and homogeneous (SRH), and that a model of evolution with one or more site-specific and time-reversible rate matrices (e.g., the GTR rate matrix) is enough to accurately model the evolution of data over the whole tree. However, an increasing body of data suggests that evolution under these conditions is an exception, rather than the norm. To address this issue, several non-SRH models of molecular evolution have been proposed, but they either ignore heterogeneity in the substitution process across sites (HAS) or assume it can be modelled accurately using the Γ distribution. As an alternative to these models of evolution, we introduce a family of mixture models that approximate HAS without the assumption of an underlying predefined statistical distribution. This family of mixture models is combined with non-SRH models of evolution that account for heterogeneity in the substitution process across lineages (HAL). We also present two algorithms for searching model space and identifying an optimal model of evolution that is less likely to over- or under-parameterize the data. The performance of the two new algorithms was evaluated using alignments of nucleotides with 10,000 sites simulated under complex non-SRH conditions on a 25-tipped tree. The algorithms were found to be very successful, identifying the correct HAL model with a 75% success rate (the average success rate for assigning rate matrices to the tree's 48 edges was 99.25%) and, for the correct HAL model, identifying the correct HAS model with a 98% success rate. Finally, parameter estimates obtained under the correct HAL-HAS model were found to be accurate and precise. The merits of our new algorithms were illustrated with an analysis of 42,337 second codon sites extracted from a concatenation of 106 alignments of orthologous genes encoded by the nuclear genomes of Saccharomyces cerevisiae, S. paradoxus, S. mikatae, S. kudriavzevii, S. castellii, S. kluyveri, S. bayanus, and Candida albicans. Our results show that second codon sites in the ancestral genome of these species contained 49.1% invariable sites, 39.6% variable sites belonging to one rate category (V1), and 11.3% variable sites belonging to a second rate category (V2). The ancestral nucleotide content was found to differ markedly across these 3 sets of sites, and the evolutionary processes operating at the variable sites were found to be non-SRH and best modelled by a combination of 8 edge-specific rate matrices (4 for V1 and 4 for V2). The number of substitutions per site at the variable sites also differed markedly, with sites belonging to V1 evolving slower than those belonging to V2 along the lineages separating the 7 species of Saccharomyces. Finally, sites belonging to V1 appeared to have ceased evolving along the lineages separating S. cerevisiae, S. paradoxus, S. mikatae, S. kudriavzevii, and S. bayanus, implying that they might have become so selectively constrained that they could be considered invariable sites in these species.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing

This study aimed to develop a set of SNP markers with high resolution and accuracy within the African buffalo. Such a set can be used, among others, to depict subtle population genetic structure for a better understanding of buffalo population dynamics. In total, 18.5 million DNA sequences of 76 bp were generated by next generation sequencing on an Illumina Genome Analyzer II from a reduced representation library using DNA from a panel of 13 African buffalo representative of the four subspecies. We identified 2534 SNPs with high confidence within the panel by aligning the short sequences to the cattle genome (Bos taurus). The average sequencing depth of the complete aligned set of reads was estimated at 5x, and at 13x when only considering the final set of putative SNPs that passed the filtering criterion. Our set of SNPs was validated by PCR amplification and Sanger sequencing of 15 SNPs. Of these 15 SNPs, 14 amplified successfully and 13 were shown to be polymorphic (success rate: 87%). The fidelity of the identified set of SNPs and potential future applications are finally discussed.

opencc-zeroDec 2015View details →
zenodo32/100

SUPPLEMENTARY FIGURE 2. Tree generated from the nucleotide sequence for the mitochondrial gene region, igr1–cox1 in A taxonomic revision of Anthothela (Octocorallia: Scleraxonia: Anthothelidae) and related genera, with the addition of new taxa, using morphological and molecular data

SUPPLEMENTARY FIGURE 2. Tree generated from the nucleotide sequence for the mitochondrial gene region, igr1–cox1 of Anthothela-like specimens. Bayesian posterior probabilities shown above branch, ML bootstrap values below branch; HKY+G (Bayesian results split freq = 0.0019, 10000000 gen, burnin=25000). (* indicates nodes present only in Bayesian analysis).

opennotspecifiedDec 2017View details →
dryad32/100

Nucleotide sequences in Procambarus clarkii

<p>Crayfish is a model for studying the effect of light on locomotor activity and neuroendocrine functions. In this study, we have described 62 transcripts from the pleonal nerve cord of the crayfish, using bioinformatics tools that identify phylogenetic families of genes related to the light interaction in the extraretinal photoreceptors. We deposited all sequencing data in the GenBank database. Here shows supplemental data from the freshwater crayfish <i>Procambarus clarkii</i>. The results suggest that the genes related to ocular and extraocular light perception in the crayfish <i>P. clarkii</i> use common biosynthesis pathways and phototransduction cascades.</p>

opencc-zeroNov 2021View details →
dryad32/100

Analysis of RNA-seq, DNA target enrichment, and Sanger nucleotide sequence data resolves deep splits in the phylogeny of cuckoo wasps (Hymenoptera: Chrysididae)

<p>The wasp family Chrysididae (cuckoo wasps, gold wasps) comprises exclusively parasitoid and kleptoparasitic species, many of which feature a stunning iridescent coloration and phenotypic adaptations to their parasitic life style. Previous attempts to infer phylogenetic relationships among the family's major lineages (subfamilies, tribes, genera) based on Sanger sequence data were insufficient to statistically resolve the monophyly and the phylogenetic position of the subfamily Amiseginae and the phylogenetic relationships among the tribes Allocoeliini, Chrysidini, Elampini, and Parnopini (Chrysidinae). Here, we present a phylogeny inferred from nucleotide sequence data of 492 nuclear single-copy genes (230,915 aligned amino acid sites) from 94 species of Chrysidoidea (representing Bethylidae, Chrysididae, Dryinidae, Plumariidae) and 45 outgroup species by combining RNA-seq and DNA target enrichment data. We find support for Amiseginae being more closely related to Cleptinae than to Chrysidinae. Furthermore, we find strong support for Allocoeliini being the sister lineage of all remaining Chrysidinae, while Elampini represent the sister lineage of Chrysidini and Parnopini. Our study corroborates results from a recent phylogenomic investigation which revealed Chrysidoidea as likely paraphyletic</p>

opencc-zeroOct 2021View details →
zenodo32/100

FIGURE 6 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 6. Phylogenetic tree of four sequenced assassin bugs. Bayesian inference and Maximum likelihood analysis inferred from all genes recovered the same topological structure. Bootstrap values and Bayesian posterior probabilities are indicated at each node.

opennotspecifiedJun 2013View details →
zenodo32/100

FIGURE 3 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 3. Predicted secondary structure of the rrnL in S. flavipes. Regions in red indicate the high variability in the four assassin bugs. Roman numerals denote the conserved domain structure. Dashed (-) indicate Watson-Crick base pairing and dot () indicate G-U base pairing.

opennotspecifiedJun 2013View details →
zenodo32/100

FIGURE 2 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 2. Inferred secondary structures of 22 tRNAs of S. flavipes. The tRNAs are labeled with the abbreviations of their corresponding amino acids. Dashed (-) indicate Watson-Crick base pairing and dot () indicate G-U base pairing.

opennotspecifiedJun 2013View details →
zenodo32/100

FIGURE 4 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 4. Predicted secondary structure of the rrnS in S. flavipes. Regions in red indicate the high variability in the four assassin bugs. Roman numerals denote the conserved domain structure. Dashed (-) indicate Watson-Crick base pairing and dot () indicate G-U base pairing.

opennotspecifiedJun 2013View details →
zenodo32/100

FIGURE 5 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 5. (A) The conserved region of the mitochondrial control region of S. flavipes, A. dohrni, T. dimidiata and V. hoffmanni. (B) The structural organization of the mitochondrial control region of S. flavipes. The control region flanking genes rrnS, trnI (I), trnQ (Q), and trnM (M) are represented in purple and green boxes. The light blue boxes with roman numerals indicate the tandem repeat region. "G+C" indicates high G+C content region. "A+T" indicates high A+T content region. The black box indicates G element.

opennotspecifiedJun 2013View details →
zenodo32/100

FIGURE 1 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 1. Map of the mtochondrial genome of S. flavipes. Direction of gene transcription is indicated by the arrows. PCGs are shown as blue arrows, rRNA genes as purple arrows, tRNA genes as red arrows and large non-coding regions (&gt;100 bp) as cyan rectangles. tRNA genes are labeled according to single-letter IUPAC-IUB abbreviations (L1: UUR; L2:CUN; S1:AGN; S2:UCN). The GC content is plotted using a black sliding window, as the deviation from the average GC content of the entire sequence. GC Skew is plotted as the deviation from the average GC skew of the entire sequence. Ticks in the inner cycle indicate the sequence length.

opennotspecifiedJun 2013View details →
zenodo32/100

Figure 3 in Nucleotide sequence data confirm diagnosis and local endemism of variable morphospecies of Andean astroblepid catfishes (Siluriformes: Astroblepidae)

Figure 3. Results of the phylogenetic analysis of astroblepid morphospecies obtained from maximum likelihood analysis of the combined DNA sequence data set. Numerals at nodes represent bootstrap proportions (values less than 50% not shown); stars represent nodes supported by bootstrap values of 80% or greater. Sample numbers correspond with materials listed in Table 1. Letters designate morphospecies; shaded boxes denote monophyletic assemblages of population samples.

opennotspecifiedFeb 2011View details →
zenodo32/100

Figure 1 in Nucleotide sequence data confirm diagnosis and local endemism of variable morphospecies of Andean astroblepid catfishes (Siluriformes: Astroblepidae)

Figure 1. Variation in pigmentation in Astroblepus morphospecies A–I. A, morphospecies A, ANSP (Academy of Natural Sciences of Philadelphia) 180586 (4793), 51.6 mm standard length (SL), Araza River. B, morphospecies B, ANSP 180587 (4779), 75 mm SL, Araza River. C, morphospecies B, ANSP 180582 (4801), 80.4 mm SL, Araza drainage (Dr.) D, morphospecies B, ANSP 180582 (4800), 54.5 mm SL, Araza Dr. E, morphospecies C, ANSP 180581 (4805), 27.2 mm SL, Araza Dr. F, morphospecies C, ANSP 180586 (4794), 58 mm SL, Araza River. G, morphospecies D, ANSP 180599 (4822), 51.7 mm SL, Urubamba Dr. H, morphospecies D, ANSP 180602 (4499), 85 mm SL, Urubamba Dr. I, morphospecies H, ANSP 180618 (4423), 46.3 mm SL, Apurimac Dr. J, morphospecies H, ANSP 180616 (4436), 79.2 mm SL, Apurimac Dr. K, morphospecies E, ANSP 180595 (4785), 61.3 mm SL, Urubamba Dr. L, morphospecies E, ANSP 180605 (4490), 110.5 mm SL, Apurimac Dr. M, morphospecies F, ANSP 180606 (4487), 75.7 mm SL, Apurimac Dr. N, morphospecies F, ANSP 180601 (4759), 52.6 mm SL, Urubamba Dr. O, morphospecies G, ANSP 180588 (4787), 59.5 mm SL, Urubamba Dr. P, morphospecies I, ANSP 180607 (4477), 39.4 mm SL, Apurimac Dr. Photo in (A) by S. A. S.; photos in (B–P) by M. H. S. P.

opennotspecifiedFeb 2011View details →
zenodo32/100

Figure 2 in Nucleotide sequence data confirm diagnosis and local endemism of variable morphospecies of Andean astroblepid catfishes (Siluriformes: Astroblepidae)

Figure 2. Distribution of astroblepid morphospecies and study region. Circled letters correspond with the morphospecies designations (Table 1) and may represent more than one lot or collection locality.

opennotspecifiedFeb 2011View details →
zenodo32/100

FIGURE. RAxML tree based on a combined dataset of partial LSU and ITS sequence analyses. Bootstrap support values for ML equal to or greater than 60 %, Bayesian posterior probabilities (BYPP) equal to or greater than 0.95 are shown as ML/ BYPP above the nodes. New isolates are in red bold. The tree is rooted to Conioscypha lignicola and Conioschypha minutispora (FMR11245) and Conioscyphascus varius. The scale bar represents the expected number of nucleotide substitutions per site. in Yunnan-Guizhou Plateau: a mycological hotspot

FIGURE. RAxML tree based on a combined dataset of partial LSU and ITS sequence analyses. Bootstrap support values for ML equal to or greater than 60 %, Bayesian posterior probabilities (BYPP) equal to or greater than 0.95 are shown as ML/ BYPP above the nodes. New isolates are in red bold. The tree is rooted to Conioscypha lignicola and Conioschypha minutispora (FMR11245) and Conioscyphascus varius. The scale bar represents the expected number of nucleotide substitutions per site.

opennotspecifiedOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record