Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Genome sequence of M6, a diploid inbred clone of the high glycoalkaloid-producing tuber-bearing potato species Solanum chacoense, reveals residual heterozygosity

Cultivated potato (Solanum tuberosum L.) is a highly heterozygous autotetraploid that presents challenges in genome analyses and breeding. Wild potato species serve as a resource for the introgression of important agronomic traits into cultivated potato. One key species is Solanum chacoense and the diploid, inbred clone M6, which is self-compatible and has desirable tuber market quality and disease resistance traits. Sequencing and assembly of the genome of the M6 clone of S. chacoense generated an assembly of 825,767,562 bp in 8,260 scaffolds with an N50 scaffold size of 713,602 bp. Pseudomolecule construction anchored 508 Mb of the genome assembly into 12 chromosomes. Genome annotation yielded 49,124 high confidence gene models representing 37,740 genes. Comparative analyses of the M6 genome with six other Solanaceae species revealed a core set of 158,367 Solanaceae genes and 1,897 genes unique to three potato species. Analysis of single nucleotide polymorphisms across the M6 genome revealed enhanced residual heterozygosity on chromosomes 4, 8 and 9 relative to the other chromosomes. Access to the M6 genome provides a resource for identification of key genes for important agronomic traits and aids in genome-enabled development of inbred diploid potatoes with the potential to accelerate potato breeding.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Plasticity of promoter-core sequences allows bacteria to compensate for the loss of a key global regulatory gene

Transcription regulatory networks (TRNs) are of central importance for both short-term phenotypic adaptation in response to environmental fluctuations and long-term evolutionary adaptation, with global regulatory genes often being targets of natural selection in laboratory experiments. Here, we combined evolution experiments, whole-genome resequencing, and molecular genetics to investigate the driving forces, genetic constraints, and molecular mechanisms that dictate how bacteria can cope with a drastic perturbation of their TRNs. The crp gene, encoding a major global regulator in Escherichia coli, was deleted in four different genetic backgrounds, all derived from the Long-Term Evolution Experiment (LTEE) but with different TRN architectures. We confirmed that crp deletion had a more deleterious effect on growth rate in the LTEE-adapted genotypes; and we showed that the ptsG gene, which encodes the major glucose-PTS transporter, gained CRP dependence over time in the LTEE. We then further evolved the four crp-deleted genotypes in glucose minimal medium, and we found that they all quickly recovered from their growth defects by increasing glucose uptake. We showed that this recovery was specific to the selective environment and consistently relied on mutations in the cis regulatory region of ptsG, regardless of the initial genotype. These mutations affected the interplay of transcription factors acting at the promoters, changed the intrinsic properties of the existing promoters, or produced new transcription initiation sites. Therefore, the plasticity of even a single promoter region can compensate by three different mechanisms for the loss of a key regulatory hub in the E. coli TRN.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Chloroplast phylogenomic data from the green algal order Sphaeropleales (Chlorophyceae, Chlorophyta) reveal complex patterns of sequence evolution

Chloroplast sequence data are widely used to infer phylogenies of plants and algae. With the increasing availability of complete chloroplast genome sequences, the opportunity arises to resolve ancient divergences that were heretofore problematic. On the flip side, properly analyzing large multi-gene data sets can be a major challenge, as these data may be riddled with systematic biases and conflicting signals. Our study contributes new data from nine complete and four fragmentary chloroplast genome sequences across the green algal order Sphaeropleales. Our phylogenetic analyses of a 56-gene data set show that analyzing these data on a nucleotide level yields a well-supported phylogeny – yet one that is quite different from a corresponding amino acid analysis. We offer some possible explanations for this conflict through a range of analyses of modified data sets. In addition, we characterize the newly sequenced genomes in terms of their structure and content, thereby further contributing to the knowledge of chloroplast genome evolution.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Identification of SNP markers for inferring phylogeny in temperate bamboos (Poaceae: Bambusoideae) using RAD sequencing

Phylogenetic relationships among temperate species of bamboo are difficult to resolve, owing to both the challenge of detecting sufficiently variable markers and their polyploid history. Here, we use restriction site–associated DNA sequencing to identify candidate loci with fixed allelic differences segregating between and within two temperate species of bamboos: Arundinaria faberi and Yushania brevipaniculata. Approximately 27 million paired-end sequencing reads were generated across four samples. From pooled data, we assembled 67 685 and 70 668 de novo contigs from partial overlap among paired-end reads, with an average length of 240 and 241 bp for the two species, respectively, which were used to investigate functional classification of RAD tags in a blastx search. Analysed separately by population, we recovered 29 443 putatively orthologous RAD tags shared across the four sampled populations, containing 28 023 sequence variants, of which c. 13 000 are segregating between species, and c. 3000 segregating between populations within each species. Analyses based on these RAD tags yielded robust phylogenetic inferences, even with data set constructed from surprisingly few loci. This study illustrates the potential for reduced-representation genome data to resolve difficult phylogenetic relationships in temperate bamboos.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Sequence analysis of European maize inbred line F2 provides new insights into molecular and chromosomal characteristics of presence/absence variants

Maize is well known for its exceptional structural diversity, including copy number variants (CNVs) and presence/absence variants (PAVs), and there is growing evidence for the role of structural variation in maize adaptation. While PAVs have been described in this important crop species, the extent of presence/absence variation and the relative position of inbred-specific regions remain to be elucidated. De novo genome sequencing of the F2 maize inbred line which played a key role in European breeding programs over the past 50 years revealed thousands of novel genomic regions, making up 88Mb of DNA, that are present in the F2 but not in B73. Comparison of B73 and F2 PAV localization revealed contrasted chromosomal distributions between the two inbreds and specific evolutionary dynamics of PAVs as compared to SNPs. Detailed sequence and functional annotation of F2 PAV sequences revealed hundreds of new genes with transcriptional support, but also a large fraction of repetitive sequences. Detailed analysis of sequence breakpoint highlights the role of double strand break repair, but also transposon insertion in PAV generation. Typing of the B73 and F2 PAVs in maize temperate inbreds revealed that some PAVs are found only in European Flint material, thus pinpointing structural features that may be at the origin of adaptive traits involved in the success of this material. Linkage disequilibrium (LD) analysis revealed that LD is strong within PAVs, as expected by the absence of recombination in crosses where PAV is missing in one of the parents.

opencc-zeroDec 2017View details →
dryad28/100

Data from: A phylogenomic perspective on the radiation of ray-finned fishes based upon targeted sequencing of ultraconserved elements (UCEs)

Ray-finned fishes constitute the dominant radiation of vertebrates with over 32,000 species. Although molecular phylogenetics has begun to disentangle major evolutionary relationships within this vast section of the Tree of Life, there is no widely available approach for efficiently collecting phylogenomic data within fishes, leaving much of the enormous potential of massively parallel sequencing technologies for resolving major radiations in ray-finned fishes unrealized. Here, we provide a genomic perspective on longstanding questions regarding the diversification of major groups of ray-finned fishes through targeted enrichment of ultraconserved nuclear DNA elements (UCEs) and their flanking sequence. Our workflow efficiently and economically generates data sets that are orders of magnitude larger than those produced by traditional approaches and is well-suited to working with museum specimens. Analysis of the UCE data set recovers a well-supported phylogeny at both shallow and deep time-scales that supports a monophyletic relationship between Amia and Lepisosteus (Holostei) and reveals elopomorphs and then osteoglossomorphs to be the earliest diverging teleost lineages. Our approach additionally reveals that sequence capture of UCE regions and their flanking sequence offers enormous potential for resolving phylogenetic relationships within ray-finned fishes.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Mode and tempo of sequence and floral evolution within the Anserineae

The Chenopodiaceae Tribe Anserineae Dumort was proposed to include the genus Spinacia and the genus Blitum. In addition to the recent domestication of Spinacia, the tribe demonstrates extensive evolution within its floral development. We test whether the development of dioecy, monoecy, and protogyny is reflected differentially among floral developmental versus non-floral developmental genes, and whether recent domestication leaves traces in the phylogenetic relationship within the genus Spinacia. The phylogenetic predictions consistently support the sister relationship of Spinacia sp. to a Blitum clade consisting of Blitum bonus-henricus, Blitum virgatum, and Blitum nuttallianum. Relative rates tests indicate a generally faster rate of nucleotide substitutions within Spinacia. Tests of selection indicate that there is generally purifying selection acting on the sequences. In addition, insertion/deletion (indel) events occur more prominently within the Spinacia clade and occur in both coding and intron regions. The phylogenetic relationships within this tribe calls into question the hypothesis that dioecy in Spinacia evolved from a monoecious grade. The evidence for purifying selection in Spinacia suggests that the increased nucleotide substitution rates are not driving protein evolution, in contrast to evidence of protein sequence and structure evolution driven by indels. There is no footprint of domestication on sequence evolution, and we cannot detect phylogenetic signals that would support separation of the Spinacia accessions into three distinct taxa.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Restriction site-associated DNA sequencing (RAD-seq) reveals an extraordinary number of transitions among gecko sex-determining systems

Sex chromosomes have evolved many times in animals and studying these replicate evolutionary "experiments" can help broaden our understanding of the general forces driving the origin and evolution of sex chromosomes. However this plan of study has been hindered by the inability to identify the sex chromosome systems in the large number of species with cryptic, homomorphic sex chromosomes. Restriction site-associated DNA sequencing (RAD-seq) is a critical enabling technology that can identify the sex chromosome systems in many species where traditional cytogenetic methods have failed. Using newly generated RAD-seq data from twelve gecko species, along with data from the literature, we reinterpret the evolution of sex-determining systems in lizards and snakes and test the hypothesis that sex chromosomes can routinely act as evolutionary traps. We uncovered between 17 and 25 transitions among gecko sex-determining systems. This is approximately ½ to ⅔ of the total number of transitions observed among all lizards and snakes. We find support for the hypothesis that sex chromosome systems can readily become trap-like and show that adding even a small number of species from understudied clades can greatly enhance hypothesis testing in a model-based phylogenetic framework. RAD-seq will undoubtedly prove useful in evaluating other species for male or female heterogamety, particularly the majority of fish, amphibian, and reptile species that lack visibly heteromorphic sex chromosomes, and will significantly accelerate the pace of biological discovery.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Counting with DNA in metabarcoding studies: how should we convert sequence reads to dietary data?

Advances in DNA sequencing technology have revolutionised the field of molecular analysis of trophic interactions and it is now possible to recover counts of food DNA sequences from a wide range of dietary samples. But what do these counts mean? To obtain an accurate estimate of a consumer's diet should we work strictly with datasets summarising frequency of occurrence of different food taxa, or is it possible to use relative number of sequences? Both approaches are applied to obtain semi-quantitative diet summaries, but occurrence data is often promoted as a more conservative and reliable option due to taxa-specific biases in recovery of sequences. We explore representative dietary metabarcoding datasets and point out that diet summaries based on occurrence data often overestimate the importance of food consumed in small quantities (potentially including low-level contaminants) and are sensitive to the count threshold used to define an occurrence. Our simulations indicate that using relative read abundance (RRA) information often provide a more accurate view of population-level diet even with moderate recovery biases incorporated; however, RRA summaries are sensitive to recovery biases impacting common diet taxa. Both approaches are more accurate when the mean number of food taxa in samples is small. The ideas presented here highlight the need to consider all sources of bias and to justify the methods used to interpret count data in dietary metabarcoding studies. We encourage researchers to continue addressing methodological challenges, and acknowledge unanswered questions to help spur future investigations in this rapidly developing area of research.

opencc-zeroDec 2017View details →
dryad28/100

Data from: BrAD-seq: Breath Adapter Directional sequencing: a streamlined, ultra-simple and fast library preparation protocol for strand specific mRNA library construction

Next Generation Sequencing (NGS) is driving rapid advancement in biological understanding and RNA-sequencing (RNA-seq) has become an indispensable tool for biology and medicine. There is a growing need for access to these technologies although preparation of NGS libraries remains a bottleneck to wider adoption. Here we report a novel method for the production of strand specific RNA-seq libraries utilizing inherent properties of double-stranded cDNA to capture and incorporate a sequencing adapter. Breath Adapter Directional sequencing (BrAD-seq) reduces sample handling and requires far fewer enzymatic steps than most available methods to produce high quality strand-specific RNA-seq libraries. The method we present is optimized for 3-prime Digital Gene Expression (DGE) libraries and can easily extend to full transcript coverage shotgun (SHO) type strand-specific libraries and is modularized to accommodate a diversity of RNA and DNA input materials. BrAD-seq offers a highly streamlined and inexpensive option for RNA-seq libraries.

opencc-zeroDec 2014View details →
dryad28/100

Data from: GIbPSs: a toolkit for fast and accurate analyses of genotyping-by-sequencing data without a reference genome

Genotyping-by-sequencing (GBS) and related methods are increasingly used for studies of non-model organisms from population genetic to phylogenetic scales. We present GIbPSs, a new genotyping toolkit for the analysis of data from various protocols such as RAD, double-digest RAD, GBS, and two-enzyme GBS without a reference genome. GIbPSs can handle paired-end GBS data and is able to assign reads from both strands of a restriction fragment to the same locus. GIbPSs is most suitable for population genetic and phylogeographic analyses. It avoids genotyping errors due to indel variation by identifying and discarding affected loci. GIbPSs creates a genotype database that offers rich functionality for data filtering and export in numerous formats. We performed comparative analyses of simulated and real GBS data with GIbPSs and another program, pyRAD. This program accounts for indel variation by aligning homologous sequences. GIbPSs performed better than pyRAD in several aspects. It required much less computation time and displayed higher genotyping accuracy. GIbPSs retained smaller numbers of loci overall in analyses of real GBS data. It nevertheless delivered more complete genotype matrices with greater locus overlap between individuals and greater numbers of loci sampled in all individuals.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference

Phylogenetic inference is generally performed on the basis of multiple sequence alignments (MSA). Because errors in an alignment can lead to errors in tree estimation, there is a strong interest in identifying and removing unreliable parts of the alignment. In recent years several automated filtering approaches have been proposed, but despite their popularity, a systematic and comprehensive comparison of different alignment filtering methods on real data has been lacking. Here, we extend and apply recently introduced phylogenetic tests of alignment accuracy on a large number of gene families and contrast the performance of unfiltered versus filtered alignments in the context of single-gene phylogeny reconstruction. Based on multiple genome-wide empirical and simulated data sets, we show that the trees obtained from filtered MSAs are on average worse than those obtained from unfiltered MSAs. Furthermore, alignment filtering often leads to an increase in the proportion of well-supported branches that are actually wrong. We confirm that our findings hold for a wide range of parameters and methods. Although our results suggest that light filtering (up to 20% of alignment positions) has little impact on tree accuracy and may save some computation time, contrary to widespread practice, we do not generally recommend the use of current alignment filtering methods for phylogenetic inference. By providing a way to rigorously and systematically measure the impact of filtering on alignments, the methodology set forth here will guide the development of better filtering algorithms.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Elucidating the functional evolution of heat sensors among Xenopus species adapted to different thermal niches by ancestral sequence reconstruction

Ambient temperature fluctuations are detected via the thermosensory system which allows animals to seek preferable thermal conditions or escape from harmful temperatures. Evolutionary changes in thermal perception have thus potentially played crucial roles in niche selection. The genus Xenopus (clawed frog) is suitable for investigating the relationship between thermal perception and niche selection due to their diverse latitudinal and altitudinal distributions. Here we performed comparative analyses of the neuronal heat sensors TRPV1 and TRPA1 among closely related Xenopus species (X. borealis, X. muelleri, X. laevis, and X. tropicalis) to elucidate their functional evolution and to assess whether their functional differences correlate with thermal niche selection among the species. Comparison of TRPV1 among four extant Xenopus species and reconstruction of the ancestral TRPV1 revealed that TRPV1 responses to repeated heat stimulation were specifically altered in the lineage leading to X. tropicalis which inhabits warmer niches. Moreover, the thermal sensitivity of TRPA1 was lower in X. tropicalis than the other species, although the thermal sensitivity of TRPV1 and TRPA1 was not always lower in species that inhabit warmer niches than the species inhabit cooler niches. However, a clear correlation was found in species differences in TRPA1 activity. Heat-evoked activity of TRPA1 in X. borealis and X. laevis, which are adapted to cooler niches, was significantly higher than in X. tropicalis and X. muelleri which are adapted to warmer niches. These findings suggest that the functional properties of heat sensors changed during Xenopus evolution, potentially altering the preferred temperature ranges among species.

opencc-zeroJun 2019View details →
dryad28/100

Data from: The phylogenetic utility and functional constraint of microRNA flanking sequences

MicroRNAs (miRNAs) have recently risen to prominence as novel factors responsible for post-transcriptional regulation of gene expression. miRNA genes have been posited as highly conserved in the clades in which they exist. Consequently, miRNAs have been used as rare genome change characters to estimate phylogeny by tracking their gain and loss. However, their short length (21–23 bp) has limited their perceived utility in sequenced-based phylogenetic inference. Here, using reference taxa with established phylogenetic relationships, we demonstrate that miRNA sequences are of high utility in quantitative, rather than in qualitative, phylogenetic analysis. The clear orthology among miRNA genes from different species makes it straightforward to identify and align these sequences from even fragmentary datasets. We also identify significant sequence conservation in the regions directly flanking miRNA genes, and show that this too is of utility in phylogenetic analysis, as well as highlighting conserved regions that will be of interest to other fields. Employing miRNA sequences from 12 sequenced drosophilid genomes, together with a Tribolium castaneum outgroup, we demonstrate that this approach is robust using Bayesian and maximum-likelihood methods. The utility of these characters is further demonstrated in the rhabditid nematodes and primates. As next-generation sequencing makes it more cost-effective to sequence genomes and small RNA libraries, this methodology provides an alternative data source for phylogenetic analysis. The approach allows rapid resolution of relationships between both closely related and rapidly evolving species, and provides an additional tool for investigation of relationships within the tree of life.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Inferring the mode of origin of polyploid species from next-generation sequence data

Many eucaryote organisms are polyploid. However, despite their importance, evolutionary inference of polyploid origins and modes of inheritance has been limited by a need for analyses of allele segregation at multiple loci using crosses. The increasing availability of sequence data for non-model species now allows the application of established approaches for the analysis of genomic data in polyploids. Here, we ask whether approximate Bayesian computation (ABC), applied to realistic traditional and next-generation sequence data, allows correct inference of the evolutionary and demographic history of polyploids. Using simulations, we evaluate the robustness of evolutionary inference by ABC for tetraploid species as a function of the number of individuals and loci sampled, and the presence or absence of an outgroup. We find that ABC adequately retrieves the recent evolutionary history of polyploid species on the basis of both old and new sequencing technologies. Application of ABC to sequence data from diploid and polyploid species of the plant genus Capsella confirms its utility. Our analysis strongly supports an allopolyploid origin of C. bursa-pastoris about 80,000 years ago. This conclusion runs contrary to previous findings based on the same dataset but using an alternative approach and is in agreement with recent findings based on whole-genome sequencing. Our results indicate that ABC is a promising and powerful method for revealing the evolution of polyploid species, without the need to attribute alleles to a homeologous chromosome pair. The approach can readily be extended to more complex scenarios involving higher ploidy levels.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Coestimating reticulate phylogenies and gene trees from multilocus sequence data

The multispecies network coalescent (MSNC) is a stochastic process that captures how gene trees grow within the branches of a phylogenetic network. Coupling the MSNC with a stochastic mutational process that operates along the branches of the gene trees gives rise to a generative model of how multiple loci from within and across species evolve in the presence of both incomplete lineage sorting (ILS) and reticulation (e.g., hybridization). We report on a Bayesian method for sampling the parameters of this generative model, including the species phylogeny, gene trees, divergence times, and population sizes, from DNA sequences of multiple independent loci. We demonstrate the utility of our method by analyzing simulated data and reanalyzing an empirical data set. Our results demonstrate the significance of not only co-estimating species phylogenies and gene trees, but also accounting for reticulation and ILS simultaneously. In particular, we show that when gene flow occurs, our method accurately estimates the evolutionary histories, coalescence times, and divergence times. Tree inference methods, on the other hand, underestimate divergence times and overestimate coalescence times when the evolutionary history is reticulate. While the MSNC corresponds to an abstract model of ``intermixture," we study the performance of the model and method on simulated data generated under a gene flow model. We show that the method accurately infers the most recent time at which gene flow occurs. Finally, we demonstrate the application of the new method to a 106-locus yeast data set.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Mining for single nucleotide polymorphisms and insertions / deletions in expressed sequence tag libraries of oil palm

The oil palm is a tropical oil bearing tree. Recently EST-derived SNPs and SSRs are a free by-product of the currently expanding EST (Expressed Sequence Tag) data bases. The development of high-throughput methods for the detection of SNPs (Single Nucleotide Polymorphism) and small indels (insertion / deletion) has led to a revolution in their use as molecular markers. Available (5452) Oil palm EST sequences were mined from dbEST of NCBI. CAP3 program was used to assemble EST sequences into contigs. Candidate SNPs and Indel polymorphisms were detected using the perl script auto_snip version 1.0 which has used 576 ESTs for detecting SNPs and Indel sites. We found 1180 SNP sites and 137 indel polymorphisms with frequency 1.36 SNPs / 100 bp. Among the six tissues from which the EST libraries had been generated, mesocarp had high frequency of 2.91 SNPs and indels per 100 bp whereas the zygotic embryos had lowest frequency of 0.15 per 100 bp. We also used the Shannon index to analyze the proportion of ten possible types of SNP/indels. ESTs from tissues of normal apex showed highest values of Shannon index (0.60) whereas abnormal apex had least value (0.02). The present report deals the use of Shannon index for comparing SNP/ indel frequencies mined from ESTlibraries and also confirm that the frequency of SNP occurrence in oil palm to use them as markers for genetic studies.

opencc-zeroDec 2007View details →
dryad28/100

Data from: Targeted sequence capture and resequencing implies a predominant role of regulatory regions in the divergence of a sympatric lake whitefish species pair (Coregonus clupeaformis)

Latest technological developments in evolutionary biology bring new challenges in documenting the intricate genetic architecture of species in the process of divergence. Sympatric populations of lake whitefish represent one of the key systems to investigate this issue. Despite the value of random genotype-by-sequencing methods and decreasing cost of sequencing technologies, it remains challenging to investigate variation in coding regions, especially in the case of recently duplicated genomes as in salmonids, as this greatly complicates whole genome resequencing. We thus designed a sequence capture array targeting 2773 annotated genes to document the nature and the extent of genomic divergence between sympatric dwarf and normal whitefish. Among the 2728 genes successfully captured, a total of 2182 coding and 10 415 noncoding putative single-nucleotide polymorphisms (SNPs) were identified after applying a first set of basic filters. A genome scan with a quality-refined selection of 2203 SNPs identified 267 outlier SNPs in 210 candidate genes located in genomic regions potentially involved in whitefish divergence and reproductive isolation. We found highly heterogeneous FST estimates among SNP loci. There was an overall low level of coding polymorphism, with a predominance of noncoding mutations among outliers. The heterogeneous patterns of divergence among loci confirm the porous nature of genomes during speciation with gene flow. Considering that few protein-coding mutations were identified as highly divergent, our results, along with previous transcriptomic studies, imply that changes in regulatory regions most likely had a greater role in the process of whitefish population divergence than protein-coding mutations. This study is the first to demonstrate the efficiency of large-scale targeted resequencing for a nonmodel species with such a large and unsequenced genome.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Genomic variation underlying complex life history traits revealed by genome sequencing in Chinook salmon

A broad portfolio of phenotypic diversity in natural organisms can buffer against exploitation and increase species persistence in disturbed ecosystems. The study of genomic variation that accounts for ecological and evolutionary adaptation can represent a powerful approach to extend understanding of phenotypic variation in nature. Here we present a chromosome-level reference genome assembly for Chinook salmon (Oncorhynchus tshawytscha; 2.36 Gb) that enabled association mapping of life history variation and phenotypic traits for this species. Whole genome resequencing of populations with distinct life history traits provided evidence that divergent selection was extensive throughout the genome within and among phylogenetic lineages, indicating a broad portfolio of phenotypic diversity exists in this species that is related to local adaptation and life history variation. Association mapping with millions of genome-wide SNPs revealed that a genomic region of major effect on chromosome 28 was associated with phenotypes for premature and mature arrival to spawning grounds and was consistent across three distinct phylogenetic lineages. Our results demonstrate how genomic resources can enlighten the genetic basis of known phenotypes in exploited species and assist in clarifying phenotypic variation that may be difficult to observe in naturally occurring organisms.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Genomic-scale capture and sequencing of endogenous DNA from feces

Genomic-level analyses of DNA from non-invasive sources would facilitate powerful conservation and evolutionary studies in natural populations of endangered and otherwise elusive species. However, the typical low quantity and poor quality of DNA that is extracted from non-invasive samples have generally precluded such work. Here we apply a modified DNA capture protocol that, when used in combination with massively-parallel sequencing technology, facilitates efficient and highly-accurate resequencing of megabases of specified nuclear genomic regions from fecal DNA samples. We validated our approach by comparing genetic variants identified from corresponding fecal and blood DNA samples of six western chimpanzees (Pan troglodytes verus) across more than 1.5 megabases of chromosome 21, chromosome X, and the complete mitochondrial genome. Our results suggest that it is now feasible to conduct genomic studies in natural populations for which constraints on invasive sampling have otherwise long been a barrier. The data we collected also provided an opportunity to examine western chimpanzee genetic diversity at unprecedented scale. Despite high mitochondrial genome diversity (pi = 0.585%), western chimpanzees have a low ratio (0.42) of X chromosomal (pi = 0.034%) to autosomal (chromosome 21 pi = 0.081%) sequence diversity, a pattern that may reflect an unusual demographic history of this subspecies.

opencc-zeroDec 2009View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record