Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Breakdown of phylogenetic signal: a survey of microsatellite densities in 454 shotgun sequences from 154 non model eukaryote species

Microsatellites are ubiquitous in Eukaryotic genomes. A more complete understanding of their origin and spread can be gained from a comparison of their distribution within a phylogenetic context. Although information for model species is accumulating rapidly, it is insufficient due to a lack of species depth, thus intragroup variation is necessarily ignored. As such, apparent differences between groups may be overinflated and generalizations cannot be inferred until an analysis of the variation that exists within groups has been conducted. In this study, we examined microsatellite coverage and motif patterns from 454 shotgun sequences of 154 Eukaryote species from eight distantly related phyla (Cnidaria, Arthropoda, Onychophora, Bryozoa, Mollusca, Echinodermata, Chordata and Streptophyta) to test if a consistent phylogenetic pattern emerges from the microsatellite composition of these species. It is clear from our results that data from model species provide incomplete information regarding the existing microsatellite variability within the Eukaryotes. A very strong heterogeneity of microsatellite composition was found within most phyla, classes and even orders. Autocorrelation analyses indicated that while microsatellite contents of species within clades more recent than 200 Mya tend to be similar, the autocorrelation breaks down and becomes negative or non-significant with increasing divergence time. Therefore, the age of the taxon seems to be a primary factor in degrading the phylogenetic pattern present among related groups. The most recent classes or orders of Chordates still retain the pattern of their common ancestor. However, within older groups, such as classes of Arthropods, the phylogenetic pattern has been scrambled by the long independent evolution of the lineages.

opencc-zeroDec 2011View details →
dryad28/100

Data from: From benchtop to desktop: important considerations when designing amplicon sequencing workflows

Amplicon sequencing has been the method of choice in many high-throughput DNA sequencing (HTS) applications. To date there has been a heavy focus on the means by which to analyse the burgeoning amount of data afforded by HTS. In contrast, there has been a distinct lack of attention paid to considerations surrounding the importance of sample preparation and the fidelity of library generation. No amount of high-end bioinformatics can compensate for poorly prepared samples and it is therefore imperative that careful attention is given to sample preparation and library generation within workflows, especially those involving multiple PCR steps. This paper redresses this imbalance by focusing on aspects pertaining to the benchtop within typical amplicon workflows: sample screening, the target region, and library generation. Empirical data is provided to illustrate the scope of the problem. Lastly, the impact of various data analysis parameters is also investigated in the context of how the data was initially generated. It is hoped this paper may serve to highlight the importance of pre-analysis workflows in achieving meaningful, future-proof data that can be analysed appropriately. As amplicon sequencing gains traction in a variety of diagnostic applications from forensics to environmental DNA (eDNA) it is paramount workflows and analytics are both fit for purpose.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Mining of expressed sequence tag libraries of cacao for microsatellite markes using five computational tools

Expressed Sequence Tags (ESTs) provide researchers with a quick and inexpensive route for discovering new genes, and data on gene expression and regulation and provide genic markers that help in constructing genome maps. Cacao is an important perennial crop of humid tropics. Cacao EST sequences as available in public domain were downloaded and made into contigs. A total of 769 contigs were made using contigs assembly program pharp. Puative information of contigs were identified using NCBI and ExPASy tools such as BlastX, tblastn.

opencc-zeroDec 2007View details →
dryad28/100

Data from: The Solanum commersonii genome sequence provides insights into adaptation to stress conditions and genome evolution of wild potato relatives

Here, we report the draft genome sequence of Solanum commersonii, which consists of ∼830 megabases with an N50 of 44,303 bp anchored to 12 chromosomes, using the potato (Solanum tuberosum) genome sequence as a reference. Compared with potato, S. commersonii shows a striking reduction in heterozygosity (1.5% versus 53 to 59%), and differences in genome sizes were mainly due to variations in intergenic sequence length. Gene annotation by ab initio prediction supported by RNA-seq data produced a catalog of 1703 predicted microRNAs, 18,882 long noncoding RNAs of which 20% are shown to target cold-responsive genes, and 39,290 protein-coding genes with a significant repertoire of nonredundant nucleotide binding site-encoding genes and 126 cold-related genes that are lacking in S. tuberosum. Phylogenetic analyses indicate that domesticated potato and S. commersonii lineages diverged ∼2.3 million years ago. Three duplication periods corresponding to genome enrichment for particular gene families related to response to salt stress, water transport, growth, and defense response were discovered. The draft genome sequence of S. commersonii substantially increases our understanding of the domesticated germplasm, facilitating translation of acquired knowledge into advances in crop stability in light of global climate and environmental changes.

opencc-zeroDec 2018View details →
dryad28/100

Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C

The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.

opencc-zeroJun 2019View details →
dryad28/100

Data from: Consistency of VDJ rearrangement and substitution parameters enables accurate B cell receptor sequence annotation

VDJ rearrangement and somatic hypermutation work together to produce antibody-coding B cell receptor (BCR) sequences for a remarkable diversity of antigens. It is now possible to sequence these BCRs in high throughput; analysis of these sequences is bringing new insight into how antibodies develop, in particular for broadly-neutralizing antibodies against HIV and influenza. A fundamental step in such sequence analysis is to annotate each base as coming from a specific one of the V, D, or J genes, or from an N-addition (a.k.a. non-templated insertion). Previous work has used simple parametric distributions to model transitions from state to state in a hidden Markov model (HMM) of VDJ recombination, and assumed that mutations occur via the same process across sites. However, codon frame and other effects have been observed to violate these parametric assumptions for such coding sequences, suggesting that a non-parametric approach to modeling the recombination process could be useful. In our paper, we find that indeed large modern data sets suggest a model using parameter-rich per-allele categorical distributions for HMM transition probabilities and per-allele-per-position mutation probabilities, and that using such a model for inference leads to significantly improved results. We present an accurate and efficient BCR sequence annotation software package using a novel HMM "factorization" strategy. This package, called partis (https://github.com/psathyrella/partis/), is built on a new general-purpose HMM compiler that can perform efficient inference given a simple text description of an HMM.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Analysis of expressed sequence tags from the placenta of the live-bearing fish Poeciliopsis (Poeciliidae)

Matrotrophic fish in the genus Poeciliopsis (Poeciliidae) have a placenta-like structure used in post-fertilization maternal provisioning of the developing embryo. To understand better the structure and function of the Poeciliopsis placenta we derived cDNA libraries from the maternal follicular placenta of two matrotrophic Poeciliopsis sister species, P. turneri and P. presidionis. These species inherited their placenta from a common ancestor and represent one of three independent origins of placentas in Poeciliopsis. Expressed sequence tags were generated and putative function was determined using BLASTX homology searches and Gene Ontology annotation. Reverse transcriptase-PCR was used to verify placenta tissue expression of a putative candidate gene, alpha-2 macroglobulin. 1956 (71.5% of the total submitted ESTs) and 924 (71.0% of the total submitted ESTs) unique transcripts were identified for the P. turneri and P. presidionis placenta, respectively. Homology search and Gene Ontology annotation revealed putative genes whose products may be involved in specific transport functions of the maternal follicle. These putative genes are excellent candidates for future research on the evolution of the placenta. We discuss our results in light of the parent-offspring conflict theory of placental evolution and in terms of the Poeciliid placenta structure and function.

opencc-zeroDec 2010View details →
dryad28/100

Data from: The evolutionary history of Xiphophorus fish and their sexually selected sword: a genome-wide approach using restriction site-associated DNA sequencing

Next-generation sequencing (NGS) techniques are now key tools in the detection of population genomic and gene expression differences in a large array of organisms. However, so far few studies have utilized such data for phylogenetic estimations. Here, we use NGS data obtained from genome-wide restriction site-associated DNA (RAD) (∼66000 SNPs) to estimate the phylogenetic relationships among all 26 species of swordtail and platyfish (genus Xiphophorus) from Central America. Past studies, both sequence and morphology-based, have differed in their inferences of the evolutionary relationships within this genus, particularly at the species-level and among monophyletic groupings. We show that using a large number of markers throughout the genome, we are able to infer the phylogenetic relationships with unparalleled resolution for this genus. The relationships among all three major clades and species within each of them are highly resolved and consistent under maximum likelihood, Bayesian inference and maximum parsimony. However, we also highlight the current cautions with this data type and analyses. This genus exhibits a particularly interesting evolutionary history where at least two species may have arisen through hybridization events. Here, we are able to infer the paternal lineages of these putative hybrid species. Using the RAD-marker-based tree we reconstruct the evolutionary history of the sexually selected sword trait and show that it may have been present in the common ancestor of the genus. Together our results highlight the outstanding capacity that RAD sequencing data has for resolving previously problematic phylogenetic relationships, particularly among relatively closely related species.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Bacterial characterization of Beijing drinking water by flow cytometry and MiSeq sequencing of the 16S rRNA gene

Flow cytometry (FCM) and 16S rRNA gene sequencing data are commonly used to monitor and characterize microbial differences in drinking water distribution systems. In this study, to assess microbial differences in drinking water distribution systems, 12 water samples from different sources water (groundwater, GW; surface water, SW) were analyzed by FCM, heterotrophic plate count (HPC), and 16S rRNA gene sequencing. FCM intact cell concentrations varied from 2.2 × 103 cells/mL to 1.6 × 104 cells/mL in the network. Characteristics of each water sample were also observed by FCM fluorescence fingerprint analysis. 16S rRNA gene sequencing showed that Proteobacteria (76.9–42.3%) or Cyanobacteria (42.0–3.1%) was most abundant among samples. Proteobacteria were abundant in samples containing chlorine, indicating resistance to disinfection. Interestingly, Mycobacterium, Corynebacterium, and Pseudomonas, were detected in drinking water distribution systems. There was no evidence that these microorganisms represented a health concern through water consumption by the general population. However, they provided a health risk for special crowd, such as the elderly or infants, patients with burns and immune-compromised people exposed by drinking. The combined use of FCM to detect total bacteria concentrations and sequencing to determine the relative abundance of pathogenic bacteria resulted in the quantitative evaluation of drinking water distribution systems. Knowledge regarding the concentration of opportunistic pathogenic bacteria will be particularly useful for epidemiological studies.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Plastome sequencing of ten nonmodel crop species uncovers a large insertion of mitochondrial DNA in cashew

In plant evolution, intracellular gene transfer (IGT) is a prevalent, ongoing process. While nuclear and mitochondrial genomes are known to integrate foreign DNA via IGT and horizontal gene transfer (HGT), plastid genomes (plastomes) have resisted foreign DNA incorporation and only recently has IGT been uncovered in the plastomes of a few land plants. In this study, we completed plastome sequences for l0 crop species and describe a number of structural features including variation in gene and intron content, inversions, and expansion and contraction of the inverted repeat (IR). We identified a putative rpl22 in cinnamon (Cinnamomum verum J. Presl) and other sequenced Lauraceae and an apparent functional transfer of rpl23 to the nucleus of quinoa (Chenopodium quinoa Willd.). In the orchard tree cashew (Anacardium occidentale L.), we report the insertion of an ∼6.7-kb fragment of mitochondrial DNA into the plastome IR. BLASTn analyses returned high identity hits to mitogenome sequences including an intact ccmB open reading frame. Using three plastome markers for five species of Anacardium, we generated a phylogeny to investigate the distribution and timing of the insertion. Four species share the insertion, suggesting that this event occurred <20 million yr ago in a single clade in the genus. Our study extends the observation of mitochondrial to plastome IGT to include long-lived tree species. While previous studies have suggested possible mechanisms facilitating IGT to the plastome, more examples of this phenomenon, along with more complete mitogenome sequences, will be required before a common, or variable, mechanism can be elucidated.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Sequence Capture using PCR-generated Probes (SCPP): a cost-effective method of targeted high-throughput sequencing for non-model organisms

Recent advances in high-throughput sequencing library preparation and subgenomic enrichment methods have opened new avenues for population genetics and phylogenetics of non-model organisms. To multiplex large numbers of indexed samples while sequencing predominantly orthologous, targeted regions of the genome, we propose modifications to an existing, in-solution capture that utilizes PCR products as target probes to enrich library pools for the genomic subset of interest. The sequence capture using PCR-generated probes (SCPP) protocol requires no specialized equipment, is highly flexible, and significantly reduces experimental costs for projects where a modest scale of genetic data is optimal (25-100 genomic loci). Our alterations enable application of this method across a wider phylogenetic range of taxa and result in higher capture efficiencies and coverage at each locus. Efficient and consistent capture over multiple SCPP experiments and at various phylogenetic distances is demonstrated, extending the utility of this method to both phylogeographic and phylogenomic studies.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Amplicon pyrosequencing late Pleistocene permafrost: the removal of putative contaminant sequences and small-scale reproducibility

DNA sequencing of ancient permafrost samples can be used to reconstruct past plant, animal and bacterial communities. In this study, we assess the small-scale reproducibility of taxonomic composition obtained from sequencing four molecular markers (mitochondrial 12S ribosomal DNA (rDNA), prokaryote 16S rDNA, mitochondrial cox1 and chloroplast trnL intron) from two soil cores sampled 10 cm apart. In addition, sequenced control reactions were used to produce a contaminant library that was used to filter similar sequences from sample libraries. Contaminant filtering resulted in the removal of 1% of reads or 0.3% of operational taxonomic units. We found similar richness, overlap, abundance and taxonomic diversity from the 12S, 16S and trnL markers from each soil core. Jaccard dissimilarity across the two soil cores was highest for metazoan taxa detected by the 12S and cox1 markers. Taxonomic community distances were similar for each marker across the two soil cores when the chi-squared metric was used; however, the 12S and cox1 markers did not cluster well when the Goodall similarity metric was used. A comparison of plant macrofossil vs. read abundance corroborates previous work that suggests eastern Beringia was dominated by grasses and forbs during cold stages of the Pleistocene, a habitat that is restricted to isolated sites in the present-day Yukon.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Phylogeographical patterns of an alpine plant, Rhodiola dumulosa (Crassulaceae), inferred from chloroplast DNA sequences

The phylogeographical patterns of Rhodiola dumulosa, an alpine plant species restrictedly growing in the crevices of rock piles, were investigated based on 4 fragments of the chloroplast genome. To cover the full distribution of R. dumulosa in China, 19 populations from 3 major disjunct distribution areas (northern, central, and northwestern China) were sampled. A total of 5881bp (after alignment) of chloroplast DNA (cpDNA) from 100 individuals were sequenced. The combined cpDNA data set yielded 36 haplotypes. The total genetic diversity of R. dumulosa was remarkably high (H T = 0.981). The interpopulation genetic differentiation was significantly large (F ST = 0.8537, P < 0.001), possibly due to the long-term isolation of the natural populations. N ST was significantly larger than G ST (P < 0.001), indicating the presence of phylogeographical structure among the R. dumulosa populations. We propose 2 migration steps to explain the current distribution of R. dumulosa in China. First, this species migrated from refugia in the Qinghai-Tibetan Plateau to northern areas via the intervening highlands when temperatures increased; second, the highland populations migrated toward the mountaintops when temperatures increased further because R. dumulosa is adapted to cold environments. During the second migration step, the common ancestral haplotypes may have been gradually lost.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Genome-wide identification of microsatellites and transposable elements in the dromedary camel genome using whole genome sequencing data

Transposable elements (TEs) along with simple sequence repeats (SSRs) are prevalent in eukaryotic genome, especially in mammals. Repetitive sequences form approximately one-third of the camelid genomes, so study on this part of genome can be helpful in providing deeper information from the genome and its evolutionary path. Here, in order to improve our understanding regarding the camel genome architecture, the whole genome of the two dromedaries (Yazdi and Trodi camels) was sequenced. Totally, 92- and 84.3-Gb sequence data were obtained and assembled to 137,772 and 149,997 contigs with a N50 length of 54,626 and 54,031 bp in Yazdi and Trodi camels, respectively. Results showed that 30.58% of Yazdi camel genome and 30.50% of Trodi camel genome were covered by TEs. Contrary to the observed results in the genomes of cattle, sheep, horse, and pig, no endogenous retrovirus-K (ERVK) elements were found in the camel genome. Distribution pattern of DNA transposons in the genomes of dromedary, Bactrian, and cattle was similar in contrast with LINE, SINE, and long terminal repeat (LTR) families. Elements like RTE-BovB belonging to LINEs family in cattle and sheep genomes are dramatically higher than genome of dromedary. However, LINE1 (L1) and LINE2 (L2) elements cover higher percentage of LINE family in dromedary genome compared to genome of cattle. Also, 540,133 and 539,409 microsatellites were identified from the assembled contigs of Yazdi and Trodi dromedary camels, respectively. In both samples, di-(393,196) and tri-(65,313) nucleotide repeats contributed to about 42.5% of the microsatellites. The findings of the present study revealed that non-repetitive content of mammalian genomes is approximately similar. Results showed that 9.1 Mb (0.47% of whole assembled genome) of Iranian dromedary's genome length is made up of SSRs. Annotation of repetitive content of Iranian dromedary camel genome revealed that 9,068 and 11,544 genes contain different types of TEs and SSRs, respectively. SSR markers identified in the present study can be used as a valuable resource for genetic diversity investigations and marker-assisted selection (MAS) in camel-breeding programs.

opencc-zeroJul 2019View details →
dryad28/100

Data from: Species level phylogeny and polyploid relationships in Hordeum (Poaceae) inferred by next-generation sequencing and in-silico cloning of multiple nuclear loci

Polyploidization is an important speciation mechanism in the barley genus Hordeum. To analyze evolutionary changes after allopolyploidization, knowledge of parental relationships is essential. One chloroplast and 12 nuclear single-copy loci were amplified by polymerase chain reaction (PCR) in all Hordeum plus six out-group species. Amplicons from each of 96 individuals were pooled, sheared, labeled with individual-specific barcodes and sequenced in a single run on a 454 platform. Reference sequences were obtained by cloning and Sanger sequencing of all loci for nine supplementary individuals. The 454 reads were assembled into contigs representing the 13 loci and, for polyploids, also homoeologues. Phylogenetic analyses were conducted for all loci separately and for a concatenated data matrix of all loci. For diploid taxa, a Bayesian concordance analysis and a coalescent-based dated species tree was inferred from all gene trees. Chloroplast matK was used to determine the maternal parent in allopolyploid taxa. The relative performance of different multilocus analyses in the presence of incomplete lineage sorting and hybridization was also assessed. The resulting multilocus phylogeny reveals for the first time species phylogeny and progenitor-derivative relationships of all di- and polyploid Hordeum taxa within a single analysis. Our study proves that it is possible to obtain a multilocus species-level phylogeny for di- and polyploid taxa by combining PCR with next-generation sequencing, without cloning and without creating a heavy load of sequence data.

opencc-zeroDec 2014View details →
dryad28/100

Data from: A long PCR based approach for DNA enrichment prior to next-generation sequencing for systematic studies

Premise of the study: We present an alternative approach for molecular systematic studies that combines long PCR and next-generation sequencing (NGS). Our approach can be used to generate templates from any DNA source for NGS. Here we test our approach by amplifying complete chloroplast genomes and we present a set of 58 potentially universal primers for angiosperms to do so. Additionally, this approach is likely to be particularly useful for nuclear regions. Methods and Results: Chloroplast genomes of 30 species across angiosperms were amplified to test our approach. Amplification success varied depending on whether PCR conditions were optimized for a given taxon. To further test our approach, some amplicons were sequenced on an Illumina HiSeq 2000. Conclusions: Although here we tested this approach by sequencing plastomes, long PCR amplicons could be generated using DNA from any genome, expanding the possibilities of this approach for molecular systematic studies.

opencc-zeroDec 2013View details →
dryad28/100

Data from: "NGS based generation of expressed sequence tags for Lymantria dispar and Lymantria monacha, two closely related lepidopteran species with different responses to parasitism by Glyptapanteles liparidis" in Genomic Resources Notes accepted 1 December 2013 to 31 January 2014

Introduction: The gypsy moth, Lymantria dispar, and the nun moth, Lymantria monacha, are closely related species (Lepidoptera, Lymantriidae), co-seasonal and economically important forest pests on broadleaf and coniferous trees. In Central Europe, gypsy moth larvae are frequently parasitized by the gregarious, endoparasitic wasp Glyptapanteles liparidis (Hymenoptera, Braconidae). At oviposition, the female wasp injects between 10 and up to 100 eggs into the hemocoel of a single host larva, together with venom and calyx fluid containing polydnavirus (PDV) particles that subsequently play a critical role in suppressing the host immune response so that successful development of the parasitoid can proceed (Schopf 2007). These viruses, which are integrated in the genomic DNA of the wasp and undergo replication only in the female's ovary, rapidly enter host hemocytes, fat body, and nervous system following parasitization, and viral genes are expressed. In L. dispar larvae parasitized by G. liparidis, the host's hemocytes alter their behavior, fail to spread properly (thereby inhibiting the encapsulation response) and partly undergo programmed cell death (apoptosis), resulting in a dramatic drop in the host's total hemocyte number (Schafellner and Schläger 2009).

opencc-zeroDec 2013View details →
dryad28/100

Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture

The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Secondary structure models of 18S and 28S rRNAs of the true bugs based on complete rDNA sequences of Eurydema maracandica Oshanin, 1871 (Heteroptera: Pentatomidae)

The sequences of 18S and 28S rDNAs have been used as molecular markers to resolve phylogenetic relationships of Heteroptera for two decades. The complete sequences of 18S rDNAs have been used in many studies, while in most studies only partial sequences of 28S rDNAs have been used due to technical difficulties of amplifying the complete lengths. In this study, we amplified the complete 18S and 28S rDNA sequences of Eurydema maracandica Oshanin, 1871, and reconstructed the secondary structure models of the corresponding rRNAs. In addition, and more importantly, all of the length variable regions of 18S rRNA were compared among 37 families of Heteroptera based on 140 sequences, and the D3 region of 28S rRNA was compared among 51 families based on 84 sequences. It was found that 8 length variable regions could potentially serve as molecular synapomorphies for some monophyletic groups. Therefore discoveries of more molecular synapomorphies for specific clades can be anticipated from amplification of complete 18S and 28S rDNAs of more representatives of Heteroptera.

opencc-zeroDec 2012View details →
dryad28/100

Data from: NREM2 and sleep spindles are instrumental to the consolidation of motor sequence memories

Although numerous studies have convincingly demonstrated that sleep plays a critical role in motor sequence learning (MSL) consolidation, the specific contribution of the different sleep stages in this type of memory consolidation is still contentious. To probe the role of stage 2 non-REM sleep (NREM2) in this process, we used a conditioning protocol in three different groups of participants who either received an odor during initial training on a motor sequence learning task and were re-exposed to this odor during different sleep stages of the post-training night (i.e., NREM2 sleep [Cond-NREM2], REM sleep [Cond-REM], or were not conditioned during learning but exposed to the odor during NREM2 [NoCond]). Results show that the Cond-NREM2 group had significantly higher gains in performance at retest than both the Cond-REM and NoCond groups. Also, only the Cond-NREM2 group yielded significant changes in sleep spindle characteristics during cueing. Finally, we found that a change in frequency of sleep spindles during cued-memory reactivation mediated the relationship between the experimental groups and gains in performance the next day. These findings strongly suggest that cued-memory reactivation during NREM2 sleep triggers an increase in sleep spindle activity that is then related to the consolidation of motor sequence memories.

opencc-zeroDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record