Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
549
datasets available to search
ShareScore release 0.9.0
Dataset results
549 results for “SNP data”
Pakistani historical wheat panel 37K SNP data
Open the record for dataset details and reuse information.
Obuasi case study data: Performance of neutral SNP barcodes to determine genetic diversity and structure of Plasmodium falciparum in Africa
Open the record for dataset details and reuse information.
Data from: Genome-wide SNP identification and association mapping for seed mineral concentration in Mung bean (Vigna radiata L.)
<p><span><span>Mung bean (<i>Vigna radiata</i> L.) quality is dependent on seed chemical composition, which in turn determines the benefits of mung bean consumption for human health. While rich in a range of nutritional components, such as protein, macro- and micro- nutrients, carbohydrates and vitamins, mung bean remains less well studied than other legume crops. Mung bean genomics and genetic resources are relatively sparse. To further improve nutritional levels of mung bean grain requires genome-wide marker system tools. The objectives of this research were to develop these tools and conduct nutrient analysis in order to 1) identify single nucleotide polymorphisms (SNPs) using genotyping by sequencing (GBS) and to 2) perform genome-wide association studies (GWAS) for levels of calcium, iron, potassium, manganese, phosphorous, sulfur, and zinc in mung bean grain produced over two years of field experiment. A total of 112 GWAS models were explored using 6,486 high quality SNPs discovered in 92 cultivated mung bean accessions chosen from USDA core collection that represented 13 countries. The data obtained allowed for the identification of 43 associated SNPs and 20 main genomic regions that explained on average 22 % of the overall variation in seed macro- and micro- nutrients concentration on the basis of a multiple-year analysis. Most of the regions discovered in this study provide valuable candidate gene to use in future breeding of new varieties of mung bean with novel nutritional properties. Identification of the <a>underlying genes</a> will help to reveal the genetic control of mung bean seed nutritional property. Other SNPs identified in this study will serve as important resources to enable marker-assisted selection (MAS) in the species <i>V</i>. <i>radiata</i>, including wide and narrow crosses with / between cultivated and wild mung bean.</span></span></p>
Data from: Development of a Chinook salmon sex identification SNP assay based on the growth hormone pseudogene
Genotypic sex identification assays can provide valuable information about fish populations when phenotypic sex determination is difficult. Here we describe the development of a TaqMan® assay (Ots_SexID) designed to identify the genotypic sex of Winter-Run Chinook salmon collected from the Sacramento River and spawned at the Livingston Stone National Fish Hatchery. The TaqMan® assay targets a region previously examined in the growth hormone pseudogene. Accuracy of the marker was assessed by comparing genotypic sex assignments for Chinook salmon spawned at Livingston Stone National Fish hatchery in 2012 (n = 84) to phenotypic sex recorded during spawning. Genotypic sex was observed to be concordant with phenotypic sex identified using Ots_SexID in 83/84 individuals, suggesting that the assay could be used to predict phenotypic sex with ~99% accuracy. To evaluate the utility of the TaqMan® assay in other parts of the species' range, we examined collections from 29 other populations ranging from Alaska to California. Sex assignments based on the assay were generally concordant with observed phenotypes, but there were some strong exceptions. These results suggest that the new assay will be very useful in Sacramento River Winter-Run Chinook salmon, but also highlight the importance of thoroughly testing any sex identification assay prior to application in a population of interest.
Data from: A study of applicability of SNP chips developed for bovine and ovine species to whole-genome analysis of reindeer Rangifer tarandus
Two sets of commercially available single nucleotide polymorphisms (SNPs) developed for cattle (BovineSNP50 BeadChip) and sheep (OvineSNP50 BeadChip) have been trialed for whole-genome analysis of 4 female samples of Rangifer tarandus inhabiting Russia. We found out that 43.0% of bovine and 47.0% of Ovine SNPs could be genotyped, while only 5.3% and 2.03% of them were respectively polymorphic. The scored and the polymorphic SNPs were identified on each bovine and each ovine chromosome, but their distribution was not unique. The maximal value of runs of homozygosity (ROH) was 30.93Mb (for SNPs corresponding to bovine chromosome 8) and 80.32Mb (for SNPs corresponding to ovine chromosome 7). Thus, the SNP chips developed for bovine and ovine species can be used as a powerful tool for genome analysis in reindeer R. tarandus.
Data from: RAD sequencing yields a high success rate for westslope cutthroat and rainbow trout species-diagnostic SNP assays
Hybridization with introduced rainbow trout threatens most native westslope cutthroat trout populations. Understanding the genetic effects of hybridization and introgression requires a large set of high-throughput, diagnostic genetic markers to inform conservation and management. Recently, we identified several thousand candidate single nucleotide polymorphism (SNP) markers based on RAD sequencing of 11 westslope cutthroat trout and 13 rainbow trout individuals. Here we used flanking sequence for 56 of these candidate SNP markers to design high-throughput genotyping assays. We validated the assays on a total of 92 individuals from 22 populations and seven hatchery strains. Forty-six assays (82%) amplified consistently and allowed easy identification of westslope cutthroat and rainbow trout alleles as well as heterozygote controls. The 46 SNPs will provide high power for early detection of population admixture and improved identification of hybrid and non-hybridized individuals. This technique shows promise as a very low-cost, reliable, and relatively rapid method for developing and testing SNP markers for non-model organisms with limited genomic resources.
Data from: Insights into the genetic history of French cattle from dense SNP data on 47 worldwide breeds
BACKGROUND: Modern cattle originate from populations of the wild extinct aurochs through a few domestication events which occurred about 8,000 years ago. Newly domesticated populations subsequently spread worldwide following breeder migration routes. The resulting complex historical origins associated with both natural and artificial selection have led to the differentiation of numerous different cattle breeds displaying a broad phenotypic variety over a short period of time. METHODOLOGY/PRINCIPAL FINDINGS: This study gives a detailed assessment of cattle genetic diversity based on 1,121 individuals sampled in 47 populations from different parts of the world (with a special focus on French cattle) genotyped for 44,706 autosomal SNPs. The analyzed data set consisted of new genotypes for 296 individuals representing 14 French cattle breeds which were combined to those available from three previously published studies. After characterizing SNP polymorphism in the different populations, we performed a detailed analysis of genetic structure at both the individual and population levels. We further searched for spatial patterns of genetic diversity among 23 European populations, most of them being of French origin, under the recently developed spatial Principal Component analysis framework. CONCLUSIONS/SIGNIFICANCE: Overall, such high throughput genotyping data confirmed a clear partitioning of the cattle genetic diversity into distinct breeds. In addition, patterns of differentiation among the three main groups of populations—the African taurine, the European taurine and zebus—may provide some additional support for three distinct domestication centres. Finally, among the European cattle breeds investigated, spatial patterns of genetic diversity were found in good agreement with the two main migration routes towards France, initially postulated based on archeological evidence.
Data from: Genome-wide SNP data reveal cryptic phylogeographic structure and microallopatric divergence in a rapids-adapted clade of cichlids from the Congo River
The lower Congo River (LCR) is a freshwater biodiversity hotspot in Africa characterized by some of the world's largest rapids. However, little is known about the evolutionary forces shaping this diversity, which include numerous endemic fishes. We investigated phylogeographic relationships in Teleogramma, a small clade of rheophilic cichlids, in the context of regional geography and hydrology. Previous studies have been unable to resolve phylogenetic relationships within Teleogramma due to lack of variation in nuclear genes and discrete morphological characters among putative species. To sample more broadly across the genome we analyzed double-digest restriction-associated sequencing (ddRAD) data from 53 individuals across all described species in the genus. We also assessed body shape and mitochondrial variation within and between taxa. Phylogenetic analyses reveal previously unrecognized lineages and instances of microallopatric divergence across as little as ~1.5 km. Species ranges appear to correspond to geographic regions broadly separated by major hydrological and topographic barriers, indicating these features are likely important drivers of diversification. Mitonuclear discordance indicates one or more introgressive hybridization events, but no clear evidence of admixture is present in nuclear genomes, suggesting these events were likely ancient. A survey of female fin patterns hints that previously undetected lineage-specific patterning may be acting to reinforce species cohesion. These analyses highlight the importance of hydrological complexity in generating diversity in certain freshwater systems, as well as the utility of ddRAD-Seq data in understanding diversification processes operating both below and above the species level.
Data from: "Polar bear (Ursus maritimus) transcriptome assembly and SNP discovery" in Genomic Resources Notes accepted 1 August 2013-30 September 2013
Polar bears (Ursus maritimus) in the Western Hudson Bay subpopulation have been declining in size and body condition for decades, as climate change causes earlier sea ice breakup, reduced hunting time on the ice, and an increasingly long fasting season. As Western Hudson Bay females have decreased in size, rates of litter production and average litter size have also decreased, while cub mortality and average time to independence have increased. Although these changes have potential evolutionary consequences, little is yet known about the adaptive genetic variation in body size or fat accumulation that would have to underlie any such change. In this study, we used high-throughput Illumina sequencing to develop SNPs from pooled blood and fat transcriptomes, using samples from five adult female polar bears and five (unrelated) dependent cubs. In total, we generated 371,258 transcripts of which 36,755 were deemed to be "full length" (i.e., covered more than 90% of their best BLAST hit), and we identified 63,020 SNPs. Since this study was conducted, we have used a subset of these SNPs to develop an Illumina BeadArray for quantitative genetics research in Western Hudson Bay.
Data from: Identifying litchi (Litchi chinensis Sonn.) cultivars and their genetic relationships using single nucleotide polymorphism (SNP) markers
Litchi is an important fruit tree in tropical and subtropical areas of the world. However, there is widespread confusion regarding litchi cultivar nomenclature and detailed information of genetic relationships among litchi germplasm is unclear. In the present study, the potential of single nucleotide polymorphism (SNP) for the identification of 96 representative litchi accessions and their genetic relationships in China was evaluated using 155 SNPs that were evenly spaced across litchi genome. Ninety SNPs with minor allele frequencies above 0.05 and a good genotyping success rate were used for further analysis. A relatively high level of genetic variation was observed among litchi accessions, as quantified by the expected heterozygosity (He = 0.305). The SNP based multilocus matching identified two synonymous groups, 'Heiye' and 'Wuye', and 'Chengtuo' and 'Baitangli 1'. A subset of 14 SNPs was sufficient to distinguish all the non-redundant litchi genotypes, and these SNPs were proven to be highly stable by repeated analyses of a selected group of cultivars. Unweighted pair-group method of arithmetic averages (UPGMA) cluster analysis divided the litchi accessions analyzed into four main groups, which corresponded to the traits of extremely early-maturing, early-maturing, middle-maturing, and late-maturing, indicating that the fruit maturation period should be considered as the primary criterion for litchi taxonomy. Two subpopulations were detected among litchi accessions by STRUCTURE analysis, and accessions with extremely early- and late-maturing traits showed membership coefficients above 0.99 for Cluster 1 and Cluster 2, respectively. Accessions with early- and middle-maturing traits were identified as admixture forms with varying levels of membership shared between the two clusters, indicating their hybrid origin during litchi domestication. The results of this study will benefit litchi germplasm conservation programs and facilitate maximum genetic gains in litchi breeding programs.
Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates
Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage >30, but remained >0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from >5 to >30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.
Data from: "White-tailed deer (Odocoileus virginianus) transcriptome assembly and SNP discovery" in Genomic Resources Notes accepted 1 June 2013-31 July 2013
White-tailed deer (Odocoileus virginianus) are among the most abundant and widespread large mammals in the Americas, comprising up to 38 subspecies ranging from Northern Canada to Peru. Although believed to have high genetic diversity, surprisingly few genomic resources are currently available, despite the species' ecological and economic importance. White-tailed deer and other cervids throughout central North America are currently being afflicted by chronic wasting disease (CWD), one of the degenerative prion diseases collectively known as transmissible spongiform encephalopathies. Although CWD is of major importance to white-tailed deer management, little is currently known about innate resistance or susceptibility to CWD outside of polymorphisms in the prion protein gene, Prnp, though a recent study using microsatellites suggests that the disease may have additional underlying genetic components. Further association analysis is hindered by low marker density. In this study, we used high-throughput SOLiD sequencing to create novel sequence data for white-tailed deer and identify single-nucleotide polymorphisms, using the pooled blood transcriptomes of six individuals. In total, we generated 14,010 contigs of length ≥ 200 nt, representing 4,104,760 nt of unique sequence data, and we identified 66,596 SNPs. This data represents one of the largest genetic resources currently available for any cervid. We hope it will facilitate future research for population genomics and assist with the identification of genetic factors that underlie disease resistance and other traits relevant for conservation and management.
Data from: Transatlantic secondary contact in Atlantic salmon, comparing microsatellites, a SNP array, and Restriction Associated DNA sequencing for the resolution of complex spatial structure
Identification of discrete and unique assemblages of individuals or populations is central to the management of exploited species. Advances in population genomics provide new opportunities for re-evaluating existing conservation units but comparisons among approaches remain rare. We compare the utility of RAD-seq, a single nucleotide polymorphism (SNP) array and a microsatellite panel to resolve spatial structuring under a scenario of possible trans-Atlantic secondary contact in a threatened Atlantic Salmon, Salmo salar, population in southern Newfoundland. Bayesian clustering indentified two large groups subdividing the existing conservation unit and multivariate analyses indicated significant similarity in spatial structuring among the three data sets. mtDNA alleles diagnostic for European ancestry displayed increased frequency in southeastern Newfoundland and were correlated with spatial structure in all marker types. Evidence consistent with introgression among these two groups was present in both SNP data sets but not the microsatellite data. Asymmetry in the degree of introgression was also apparent in SNP data sets with evidence of gene flow towards the east or European type. This work highlights the utility of RAD-seq based approaches for the resolution of complex spatial patterns, resolves a region of trans-Atlantic secondary contact in Atlantic Salmon in Newfoundland and demonstrates the utility of multiple marker comparisons in identifying dynamics of introgression.
Data from: Microevolution in time and space: SNP analysis of historical DNA reveals dynamic signatures of selection in Atlantic cod
Little is known about how quickly natural populations adapt to changes in their environment and how temporal and spatial variation in selection pressures interact to shape patterns of genetic diversity. We here address these issues with a series of genome scans in four overfished populations of Atlantic cod (Gadus morhua) studied over an 80-year period. Screening of >1000 gene-associated single-nucleotide polymorphisms (SNPs) identified 77 loci that showed highly elevated levels of differentiation, likely as an effect of directional selection, in either time, space or both. Exploratory analysis suggested that temporal allele frequency shifts at certain loci may correlate with local temperature variation and with life history changes suggested to be fisheries induced. Interestingly, however, largely nonoverlapping sets of loci were temporal outliers in the different populations and outliers from the 1928 to 1960 period showed almost complete stability during later decades. The contrasting microevolutionary trajectories among populations resulted in sequential shifts in spatial outliers, with no locus maintaining elevated spatial differentiation throughout the study period. Simulations of migration coupled with observations of temporally stable spatial structure at neutral loci suggest that population replacement or gene flow alone could not explain all the observed allele frequency variation. Thus, the genetic changes are likely to at least partly be driven by highly dynamic temporally and spatially varying selection. These findings have important implications for our understanding of local adaptation and evolutionary potential in high gene flow organisms and underscore the need to carefully consider all dimensions of biocomplexity for evolutionarily sustainable management.
Data from: SNP-array reveals genome wide patterns of geographical and potential adaptive divergence across the natural range of Atlantic salmon (Salmo salar)
Atlantic salmon (Salmo salar) is one of the most extensively studied fish species in the world due to its significance in aquaculture, fisheries and ongoing conservation efforts to protect declining populations. Yet, limited genomic resources have hampered our understanding of genetic architecture in the species and the genetic basis of adaptation to the wide range of natural and artificial environments it occupies. In this paper, we describe the development of a medium density Atlantic salmon SNP-array based on Expressed Sequence Tags (ESTs) and genomic sequencing. The array was used in the most extensive assessment of population genetic structure performed to date in this species. A total of 6176 informative SNPs were successfully genotyped in 38 anadromous and freshwater wild populations distributed across the species natural range. Principal component analysis clearly differentiated European and North American populations, and within Europe, three major regional genetic groups were identified for the first time in a single analysis. We assessed the potential for the array to disentangle neutral and putative adaptive divergence of SNP allele frequencies across populations and among regional groups. In Europe, secondary contact zones were identified between major clusters where endogenous and exogenous barriers could be associated, rendering the interpretation of environmental influence on potentially adaptive divergence equivocal. A small number of markers highly divergent in allele frequencies (outliers) were observed between (multiple) freshwater and anadromous populations, between northern and southern latitudes, and when comparing Baltic populations to all others. We also discuss the potential future applications of the SNP-array for conservation, management and aquaculture.
Data from: SNP-skimming: a fast approach to map loci generating quantitative variation in natural populations
Genome-wide association mapping (GWAS) is a method to estimate the contribution of segregating genetic loci to trait variation. A major challenge for applying GWAS to non-model species has been generating dense genome-wide markers that satisfy the key requirement that marker data is error-free. Here we present an approach to map loci within natural populations using inexpensive shallow genome sequencing. This 'SNP skimming' approach involves two steps: an initial genome-wide scan to identify putative targets followed by deep sequencing for confirmation of targeted loci. We apply our method to a test dataset of floral dimension variation in the plant Penstemon virgatus, a member of a genus that has experienced dynamic floral adaptation that reflects repeated transitions in primary pollinator. The ability to detect SNPs that generate phenotypic variation depends on population genetic factors such as population allele frequency, effect size, and epistasis as well as sampling effects contingent on missing data and genotype uncertainty. However, both simulations and the Penstemon data suggest that the most significant tests from the initial SNP skim are likely to be true positives – loci with subtle but significant quantitative effects on phenotype. We discuss the promise and limitations of this method and consider optimal experimental design for a given sequencing effort. Simulations demonstrate that sampling a larger number of individual at the expense of average read depth per individual maximizes the power to detect loci.
Data from: Phylogeography and adaptation genetics of stickleback from the Haida Gwaii archipelago revealed using genome-wide SNP genotyping
Threespine stickleback populations are model systems for studying adaptive evolution and the underlying genetics. In lakes on the Haida Gwaii archipelago (off western Canada), stickleback have undergone a remarkable local radiation and show phenotypic diversity matching that seen throughout the species distribution. To provide a historical context for this radiation, we surveyed genetic variation at >1000 single nucleotide polymorphism (SNP) loci in stickleback from over 100 populations. SNPs included markers evenly distributed throughout genome and candidate SNPs tagging adaptive genomic regions. Based on evenly distributed SNPs, the phylogeographic pattern differs substantially from the disjunct pattern previously observed between two highly divergent mtDNA lineages. The SNP tree instead shows extensive within watershed population clustering and different watersheds separated by short branches deep in the tree. These data are consistent with separate colonizations of most watersheds, despite underlying genetic connections between some independent drainages. This supports previous suppositions that morphological diversity observed between watersheds has been shaped independently, with populations exhibiting complete loss of lateral plates and giant size each occurring in several distinct clades. Throughout the archipelago, we see repeated selection of SNPs tagging candidate freshwater adaptive variants at several genomic regions differentiated between marine–freshwater populations on a global scale (e.g. EDA, Na/K ATPase). In estuarine sites, both marine and freshwater allelic variants were commonly detected. We also found typically marine alleles present in a few freshwater lakes, especially those with completely plated morphology. These results provide a general model for postglacial colonization of freshwater habitat by sticklebacks and illustrate the tremendous potential of genome-wide SNP data sets hold for resolving patterns and processes underlying recent adaptive divergences.
Data from: Genome-wide SNP analysis unveils genetic structure and phylogeographic history of snow sheep (Ovis nivicola) populations inhabiting the Verkhoyansk Mountains and Momsky Ridge (northeastern Siberia)
Insights into the genetic characteristics of a species provide important information for wildlife conservation programs. Here, we used the OvineSNP50 BeadChip developed for domestic sheep to examine population structure and evaluate genetic diversity of snow sheep (Ovis nivicola) inhabiting Verkhoyansk Range and Momsky Ridge. A total of 1121 polymorphic SNPs were used to test 80 specimens representing five populations, including four populations of the Verkhoyansk Mountain chain: Kharaulakh Ridge–Tiksi Bay (TIK, n = 22), Orulgan Ridge (ORU, n = 22), the central part of Verkhoyansk Range (VER, n = 15), Suntar-Khayata Ridge (SKH, n = 13), and Momsky Ridge (MOM, n = 8). We showed that the studied populations were genetically structured according to a geographical pattern. Pairwise FST values ranged from 0.044 to 0.205. Admixture analysis identified K = 2 as the most likely number of ancestral populations. A Neighbor-Net tree showed that TIK was an isolated group related to the main network through ORU. TreeMix analysis revealed that TIK and MOM originated from two different ancestral populations and detected gene flow from MOM to ORU. This was supported by the f3 statistic, which showed that ORU is an admixed population with TIK and MOM/SKH heritage. Genetic diversity in the studied groups was increasing southward. Minimum values of observed (Ho) and expected (He) heterozygosity and allelic richness (Ar) were observed in the most northern population–TIK, and maximum values were observed in the most southern population–SKH. Thus, our results revealed clear genetic structure in the studied populations of snow sheep and showed that TIK has a different origin from MOM, SKH and VER even though they are conventionally considered a single subspecies known as Yakut snow sheep (Ovis nivicola lydekkeri). Most likely, TIK was an isolated group during the late Pleistocene glaciations of Verkhoyansk Range.
Data from: Genome-wide SNP data suggests complex ancestry of sympatric North Pacific killer whale ecotypes
Three ecotypes of killer whale occur in partial sympatry in the North Pacific. Individuals assortatively mate within the same ecotype, resulting in correlated ecological and genetic differentiation. A key question is whether this pattern of evolutionary divergence is an example of incipient sympatric speciation from a single panmictic ancestral population, or whether sympatry could have resulted from multiple colonisations of the North Pacific and secondary contact between ecotypes. Here, we infer multilocus coalescent trees from >1000 nuclear single-nucleotide polymorphisms (SNPs) and find evidence of incomplete lineage sorting so that the genealogies of SNPs do not all conform to a single topology. To disentangle whether uncertainty in the phylogenetic inference of the relationships among ecotypes could also result from ancestral admixture events we reconstructed the relationship among the ecotypes as an admixture graph and estimated f4-statistics using TreeMix. The results were consistent with episodes of admixture between two of the North Pacific ecotypes and the two outgroups (populations from the Southern Ocean and the North Atlantic). Gene flow may have occurred via unsampled 'ghost' populations rather than directly between the populations sampled here. Our results indicate that because of ancestral admixture events and incomplete lineage sorting, a single bifurcating tree does not fully describe the relationship among these populations. The data are therefore most consistent with the genomic variation among North Pacific killer whale ecotypes resulting from multiple colonisation events, and secondary contact may have facilitated evolutionary divergence. Thus, the present-day populations of North Pacific killer whale ecotypes have a complex ancestry, confounding the tree-based inference of ancestral geography.
Data from: Development of genomic tools in a widespread tropical tree, Symphonia globulifera L.f.: a new low-coverage draft genome, SNP and SSR markers
Population genetic studies in tropical plants are often challenging because of limited information on taxonomy, phylogenetic relationships and distribution ranges, scarce genomic information and logistic challenges in sampling. We describe a strategy to develop robust and widely applicable genetic markers based on a modest development of genomic resources in the ancient tropical tree species Symphonia globulifera L.f. (Clusiaceae), a keystone species in African and Neotropical rainforests. We provide the first low-coverage (11X) fragmented draft genome sequenced on an individual from Cameroon, covering 1.027 Gbp or 67.5% of the estimated genome size. Annotation of 565 scaffolds (7.57 Mbp) resulted in the prediction of 1046 putative genes (231 of them containing a complete open reading frame) and 1523 exact simple sequence repeats (SSRs, microsatellites). Aligning a published transcriptome of a French Guiana population against this draft genome produced 923 high-quality single nucleotide polymorphisms. We also preselected genic SSRs in silico that were conserved and polymorphic across a wide geographical range, thus reducing marker development tests on rare DNA samples. Of 23 SSRs tested, 19 amplified and 18 were successfully genotyped in four S. globulifera populations from South America (Brazil and French Guiana) and Africa (Cameroon and São Tomé island, FST = 0.34). Most loci showed only population-specific deviations from Hardy–Weinberg proportions, pointing to local population effects (e.g. null alleles). The described genomic resources are valuable for evolutionary studies in Symphonia and for comparative studies in plants. The methods are especially interesting for widespread tropical or endangered taxa with limited DNA availability.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.