Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
268
datasets available to search
ShareScore release 0.7.1
Dataset results
268 results for “polyploid”
Data from: Genetic structure and diversity of populations of polyploid Tibouchina pulchra Cogn. (Melastomataceae) under different environmental conditions in extremes of an elevational gradient
Open the record for dataset details and reuse information.
Effects of glaciation and whole genome duplication on the distribution of the Campanula rotundifolia polyploid complex
Open the record for dataset details and reuse information.
Data from: Insights into the genetic basis of blueberry fruit-related traits using diploid and polyploid models in a GWAS context
Open the record for dataset details and reuse information.
Data from: Evolution and systematics of polyploid Nigritella (Orchidaceae)
Open the record for dataset details and reuse information.
Data from: Evolution and phylogeography analysis of diploid and polyploid Misgurnus anguillicaudatus populations across China
Open the record for dataset details and reuse information.
Tracking the ancestry of known and ‘ghost’ homeologous subgenomes in model grass Brachypodium polyploids
Open the record for dataset details and reuse information.
Data from: Genetic diversity and distribution patterns of diploid and polyploid hybrid water frog populations (Pelophylax esculentus complex) across Europe
Open the record for dataset details and reuse information.
Data from: Invasion success in polyploids: the role of inbreeding in the contrasting colonization abilities of diploid versus tetraploid populations of Centaurea stoebe s.l
Open the record for dataset details and reuse information.
Genetic variation in the brooding brittle-star: A global hybrid polyploid complex?
Open the record for dataset details and reuse information.
Data from: Genetic and ecotypic differentiation in a Californian plant polyploid complex (Grindelia, Asteraceae)
Open the record for dataset details and reuse information.
Data from: Centromere satellite repeats have undergone rapid changes in polyploid wheat subgenomes
Centromeres mediate the pairing of homologous chromosomes during meiosis; this pairing is particularly challenging for polyploid plants such as hexaploid wheat (Triticum aestivum), as their meiotic machinery must differentiate homologs from similar homoeologs. However, the sequence compositions (especially functional centromeric satellites) and evolutionary history of wheat centromeres are largely unknown. Here, we mapped T. aestivum centromeres by chromatin immunoprecipitation sequencing using antibodies to the centromeric-specific histone CENH3; this identified two types of functional centromeric satellites that are abundant in two of the three subgenomes. These centromeric satellites had unit sizes over 500 bp and contained specific sites with highly phased binding to CENH3 nucleosomes. Phylogenetic analysis revealed that the satellites have diverged in the three T. aestivum subgenomes, and the more homogeneous satellite arrays are associated with CENH3. Satellite signals decreased and the degree of satellites variation increased from diploid to hexaploid wheat. Moreover, several T. aestivum centromeres lack satellite repeats. Rearrangements, including local expansion and satellite variations, inversions, and changes in gene expression, occurred during the evolution from diploid to tetraploid and hexaploid wheat. These results reveal the asymmetry in centromere organization among the wheat subgenomes, which may play a role in proper homolog pairing during meiosis.
Diploid VCF files used for publication "Haplotype Threading: Accurate Polyploid Phasing from Long Reads"
<p>Diploid VCF files for three human samples, from which polyploid datasets were generated as evaluation data for the publication given in the title. Files contain gold standard haplotypes, which have been conducted by trio phasing from the samples themselves and their respective parents.</p>
Data from: Maximum parsimony inference of phylogenetic networks in the presence of polyploid complexes
<p>Phylogenetic networks provide a powerful framework for modeling and analyzing reticulate evolutionary histories. While polyploidy has been shown to be prevalent not only in plants but also in other groups of eukaryotic species, most work done thus far on phylogenetic network inference assumes diploid hybridization. These inference methods have been applied, with varying degrees of success, to data sets with polyploid species, even though polyploidy violates the mathematical assumptions underlying these methods. Statistical methods were developed recently for handling specific types of polyploids and so were parsimony methods that could handle polyploidy more generally yet while excluding processes such as incomplete lineage sorting.</p> <p>In this paper, we introduce a new method for inferring most parsimonious phylogenetic networks on data that include polyploid species. Taking gene trees as input, the method seeks a phylogenetic network that minimizes deep coalescences while accounting for polyploidy. The method could also infer trees, thus potentially distinguishing between auto- and allo-polyploidy. We demonstrate the performance of the method on both simulated and biological data. The inference method as well as a method for evaluating given phylogenetic networks are implemented and publicly available in the PhyloNet software package.</p>
Phylogenomics of the Andean tetraploid clade of the American Amaryllidaceae (subfamily Amaryllidoideae): unlocking a polyploid generic radiation abetted by continental geodynamics
<p>The second large clade of the endemic American Amaryllidaceae subfam. Amaryllidoideae constitutes the tetraploid-derived (<i>n</i> = 23) Andean-centered tribes, most of which have 46 chromosomes. Despite progress in resolving phylogenetic relationships of the group with nrDNA, certain subclades were poorly resolved or weakly supported in those previous studies. Sequence capture using anchored hybrid enrichment was employed across 95 species of the clade along with five outgroups and generated sequences of 524 nuclear genes and a partial plastome. Maximum likelihood phylogenetic analyses were conducted on concatenated supermatrices, and coalescent species tree analyses were run on the gene trees, followed by hybridization network, age diversification and biogeographic analyses. The four tribes Clinantheae, Eucharideae, Eustephieae (the first branch), and Hymenocallideae (sister to <i>Clinanthus</i>) are resolved in all analyses with 100% support. Nuclear gene supermatrix and species tree results were largely in concordance; however cytonuclear discordance was evident. Hybridization network analysis identified significant reticulation in <i>Clinanthus</i>, <i>Hymenocallis</i>, <i>Stenomesson</i> and the subclade of Eucharideae comprising <i>Eucharis</i>, <i>Caliphruria</i>, and <i>Urceolina</i>. Our data support a previous treatment of the latter as a single genus, <i>Urceolina</i>, with the addition of <i>Eucrosia</i> <i>dodsonii</i>. Biogeographic analysis and penalized likelihood age estimation suggests an origin in the central Andean region (north and central Peru) for the complex in the mid-Oligocene, with more dispersals than vicariances in its history, but no extinctions. The Eucharideae experienced a sudden lineage radiation ca. 10 Mya. We tie much of the divergences in the Andean-centered lineages to the rise of the Andes, directly and indirectly, and suggest that the Amotape-Huancabamba Zone functioned as both a corrider (dispersal) and a barrier to migration (vicariance). Several taxonomic changes are made. This is the largest DNA sequence data set to be applied within Amaryllidaceae to date.The second large clade of the endemic American Amaryllidaceae subfam. Amaryllidoideae constitutes the tetraploid-derived (<i>n</i> = 23) Andean-centered tribes, most of which have 46 chromosomes. Despite progress in resolving phylogenetic relationships of the group with nrDNA, certain subclades were poorly resolved or weakly supported in those previous studies. Sequence capture using anchored hybrid enrichment was employed across 95 species of the clade along with five outgroups and generated sequences of 524 nuclear genes and a partial plastome. Maximum likelihood phylogenetic analyses were conducted on concatenated supermatrices, and coalescent species tree analyses were run on the gene trees, followed by hybridization network, age diversification and biogeographic analyses. The four tribes Clinantheae, Eucharideae, Eustephieae (the first branch), and Hymenocallideae (sister to <i>Clinanthus</i>) are resolved in all analyses with 100% support. Nuclear gene supermatrix and species tree results were largely in concordance; however cytonuclear discordance was evident. Hybridization network analysis identified significant reticulation in <i>Clinanthus</i>, <i>Hymenocallis</i>, <i>Stenomesson</i> and the subclade of Eucharideae comprising <i>Eucharis</i>, <i>Caliphruria</i>, and <i>Urceolina</i>. Our data support a previous treatment of the latter as a single genus, <i>Urceolina</i>, with the addition of <i>Eucrosia</i> <i>dodsonii</i>. Biogeographic analysis and penalized likelihood age estimation suggests an origin in the central Andean region (north and central Peru) for the complex in the mid-Oligocene, with more dispersals than vicariances in its history, but no extinctions. The Eucharideae experienced a sudden lineage radiation ca. 10 Mya. We tie much of the divergences in the Andean-centered lineages to the rise of the Andes, directly and indirectly, and suggest that the Amotape-Huancabamba Zone functioned as both a corrider (dispersal) and a barrier to migration (vicariance). Several taxonomic changes are made. This is the largest DNA sequence data set to be applied within Amaryllidaceae to date.</p>
Data from: Nitrate reductase phylogeny of potato (Solanum sect. Petota) genomes with emphasis on the origins of the polyploid species
Solanum section Petota is taxonomically difficult, partly because of interspecific hybridization at both the diploid and polyploid levels. There is much disagreement regarding species boundaries and affiliation of species to series. Elucidating the phylogenetic relationships within the polyploids is crucial for an effective taxonomic treatment of the section and for the utilization of wild potato germplasm in breeding programs. We here infer relationships among the potato diploids and polyploids using nitrate reductase (NIA) sequence data in comparison to prior plastid phylogenies and: 1) examine genome types within section Petota, 2) show species in the polyploid series Conicibaccata, Longipedicellata, and in the Iopetalum group to be derived from allopolyploidization, 3) support an earlier hypothesis by confirming S. verrucosum as the maternal genome donor for the polyploid species S. demissum as well as species in the Iopetalum Group, 4) demonstrate that S. verrucosum is the closest relative to the maternal genome donor for species in ser. Longipedicellata, 5) support the close relationship between S. acaule and diploid species from series Megistacroloba and Tuberosa, and 6) show the North and Central American B genome species to be well distinguished from the A genome species of South America.
Data from: Performance of gene expression analyses using de novo assembled transcripts in polyploid species
Motivation: Quality of gene expression analyses using de novo assembled transcripts in species that experienced recent polyploidization remains unexplored. Results: Differential gene expression (DGE) analyses using putative genes inferred by Trinity, Corset and Grouper performed slightly differently across five plant species that experienced various poly-ploidy histories. In species that lack recent polyploidy events that occurred in the past several millions of years, DGE analyses using de novo assembled transcriptomes identified 54–82% of the differen-tially expressed genes recovered by mapping reads to the reference genes. However, in species that experienced more recent polyploidy events, the percentage decreased to 21–65%. Gene co-expression network analyses using de novo assemblies vs. mapping to the reference genes recov-ered the same module that significantly correlated with treatment in one species that lacks recent polyploidization.
Data from: The case of the missing ancient fungal polyploids
Polyploidy—the increase in the number of whole chromosome sets—is an important evolutionary force in eukaryotes. Polyploidy is well recognized throughout the evolutionary history of plants and animals, where several ancient events have been hypothesized to be drivers of major evolutionary radiations. However, fungi provide a striking contrast: while numerous recent polyploids have been documented, ancient fungal polyploidy is virtually unknown. We present a survey of known fungal polyploids that confirms the absence of ancient fungal polyploidy events. Three hypotheses may explain this finding. First, ancient fungal polyploids are indeed rare, with unique aspects of fungal biology providing similar benefits without genome duplication. Second, fungal polyploids are not successful in the long term, leading to few extant species derived from ancient polyploidy events. Third, ancient fungal polyploids are difficult to detect, causing the real contribution of polyploidy to fungal evolution to be underappreciated. We consider each of these hypotheses in turn and propose that failure to detect ancient events is the most likely reason for the lack of observed ancient fungal polyploids. We examine whether existing data can provide evidence for previously unrecognized ancient fungal polyploidy events but discover that current resources are too limited. We contend that establishing whether unrecognized ancient fungal polyploidy events exist is important to ascertain whether polyploidy has played a key role in the evolution of the extensive complexity and diversity observed in fungi today and, thus, whether polyploidy is a driver of evolutionary diversifications across eukaryotes. Therefore, we conclude by suggesting ways to test the hypothesis that there are unrecognized polyploidy events in the deep evolutionary history of the fungi.
Data from: Inferring the mode of origin of polyploid species from next-generation sequence data
Many eucaryote organisms are polyploid. However, despite their importance, evolutionary inference of polyploid origins and modes of inheritance has been limited by a need for analyses of allele segregation at multiple loci using crosses. The increasing availability of sequence data for non-model species now allows the application of established approaches for the analysis of genomic data in polyploids. Here, we ask whether approximate Bayesian computation (ABC), applied to realistic traditional and next-generation sequence data, allows correct inference of the evolutionary and demographic history of polyploids. Using simulations, we evaluate the robustness of evolutionary inference by ABC for tetraploid species as a function of the number of individuals and loci sampled, and the presence or absence of an outgroup. We find that ABC adequately retrieves the recent evolutionary history of polyploid species on the basis of both old and new sequencing technologies. Application of ABC to sequence data from diploid and polyploid species of the plant genus Capsella confirms its utility. Our analysis strongly supports an allopolyploid origin of C. bursa-pastoris about 80,000 years ago. This conclusion runs contrary to previous findings based on the same dataset but using an alternative approach and is in agreement with recent findings based on whole-genome sequencing. Our results indicate that ABC is a promising and powerful method for revealing the evolution of polyploid species, without the need to attribute alleles to a homeologous chromosome pair. The approach can readily be extended to more complex scenarios involving higher ploidy levels.
Data from: Species level phylogeny and polyploid relationships in Hordeum (Poaceae) inferred by next-generation sequencing and in-silico cloning of multiple nuclear loci
Polyploidization is an important speciation mechanism in the barley genus Hordeum. To analyze evolutionary changes after allopolyploidization, knowledge of parental relationships is essential. One chloroplast and 12 nuclear single-copy loci were amplified by polymerase chain reaction (PCR) in all Hordeum plus six out-group species. Amplicons from each of 96 individuals were pooled, sheared, labeled with individual-specific barcodes and sequenced in a single run on a 454 platform. Reference sequences were obtained by cloning and Sanger sequencing of all loci for nine supplementary individuals. The 454 reads were assembled into contigs representing the 13 loci and, for polyploids, also homoeologues. Phylogenetic analyses were conducted for all loci separately and for a concatenated data matrix of all loci. For diploid taxa, a Bayesian concordance analysis and a coalescent-based dated species tree was inferred from all gene trees. Chloroplast matK was used to determine the maternal parent in allopolyploid taxa. The relative performance of different multilocus analyses in the presence of incomplete lineage sorting and hybridization was also assessed. The resulting multilocus phylogeny reveals for the first time species phylogeny and progenitor-derivative relationships of all di- and polyploid Hordeum taxa within a single analysis. Our study proves that it is possible to obtain a multilocus species-level phylogeny for di- and polyploid taxa by combining PCR with next-generation sequencing, without cloning and without creating a heavy load of sequence data.
FIGURE 3 in A new polyploid species of Limonium (Plumbaginaceae) from the Western Mediterranean basin
FIGURE 3. Distribution map of Limonium irtaensis; natural plant in the Irta mountains.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.