Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
84
datasets available to search
ShareScore release 0.7.1
Dataset results
84 results for “genetic resources”
Dissecting the genetic basis of variation in Drosophila sleep using a multiparental QTL mapping resource
There is considerable variation in sleep duration, timing and quality in human populations, and sleep dysregulation has been implicated as a risk factor for a range of health problems. Human sleep traits are known to be regulated by genetic factors, but also by an array of environmental and social factors. These uncontrolled, non-genetic effects complicate powerful identification of the loci contributing to sleep directly in humans. The model system, Drosophila melanogaster, exhibits a behavior that shows the hallmarks of mammalian sleep, and here we use a multitiered approach, encompassing high-resolution QTL mapping, expression QTL data, and functional validation with RNAi to investigate the genetic basis of sleep under highly controlled environmental conditions. We measured a battery of sleep phenotypes in >750 genotypes derived from a multiparental mapping panel and identified several, modest-effect QTL contributing to natural variation for sleep. Merging sleep QTL data with a large head transcriptome eQTL mapping dataset from the same population allowed us to refine the list of plausible candidate causative sleep loci. This set includes genes with previously characterized effects on sleep and circadian rhythms, in addition to novel candidates. Finally, we employed adult, nervous system-specific RNAi on the Dopa decarboxylase, dyschronic, and timeless genes, finding significant effects on sleep phenotypes for all three. The genes we resolve are strong candidates to harbor causative, regulatory variation contributing to sleep.
Microsatellite exploration in the climbing hydrangea (Hydrangea petiolaris Siebold & Zucc.) transcriptome: A resource for population genetics and functional genomics
<p><strong>Background</strong></p> <p><em>Hydrangea petiolaris</em> Siebold & Zucc., also known as climbing hydrangea, is a vine native to the woodlands of Korea, Japan, and Sakhalin Island. It is an economically important ornamental plant with fertile and sterile flowers. Despite the recent increase in <em>Hydrangea</em> breeding and interest in germplasm conservation, relatively little is known about the relationships between <em>Hydrangea</em> species.</p> <p><strong>Results</strong></p> <p>We employed Illumina NovaSeq 6000 sequencing technology to generate a total of 39,945,480 reads, which were assembled into 137,715 contigs. A total of 109,092 filtered transcripts were used to identify microsatellites, and 54,587 microsatellite repeat motifs were revealed within 33,556 contigs. Among these, 4,510 transcripts harboring microsatellites had Gene Ontology annotations, and numerous microsatellite-containing transcripts exhibited associations with genes, including those encoding PPR proteins, aldehyde dehydrogenases, and bHLH transcription factors, related to the <em>restorer of fertility</em> (<em>Rf</em>) genes, which play a critical role in restoring fertility in plants with cytoplasmic male sterility. Validation of transcriptomic SSR markers demonstrated high levels of polymorphism, revealing significant genetic diversity within populations. However, null alleles and deviations from Hardy‒Weinberg equilibrium at specific loci suggested caution in genotyping accuracy. Population-level analysis disclosed high genetic differentiation and distinct clustering of populations.</p> <p><strong>Conclusions</strong></p> <p>The <em>H</em>. <em>petiolaris</em> transcriptomic SSR markers offer valuable insight for gaining insights into the population genetics, evolutionary background, and practical strategies for conserving this species. Moreover, the microsatellite loci we have identified and their associations with annotated genes hold promise for creating functional markers specifically tailored for <em>H</em>. <em>petiolaris</em>. These markers include valuable resources of transcriptomic SSR markers suitable for population genetic investigations and have a reasonable degree of applicability across different taxa.</p>
Genetic admixture and evolutionary history of Han Chinese in the Shandong Peninsula inferred from integrative modern and ancient genomic resources
<p>The allele frequency data of 264 individuals from Shandong Province and supplementary table.</p>
Fraxinus excelsior genotype data for "Genetic resources of common ash (Fraxinus excelsior L.) in Poland"
<p>The data set contains microsatellite genotypes (3 chloroplast + 10 nuclear loci) of Fraxinus excelsior trees, together with the information about sampling sites.<br> </p>
Male diversity matters: Genetic structuring of insular male date palm (Phoenix dactylifera L.) revealed valuable breeding and conservation resources
Most male date palms available for pollinating different female cultivars have mainly originated from seed propagation resulting in many different local males that represent a source of genetic diversity. Favorable fruit production is related to quality of pollen and its compatibility with a certain female variety. Therefore, the genetic characterization of the male pollen as a superior one for each female cultivar should be the first step to establish an intensive program to produce superior males through different procedures. In this study, the genetic diversity and population structure of 72 male date palm accessions were investigated using 15 microsatellite loci. A total number of 125 alleles was detected with an average of 8.33 alleles per locus. Bayesian model-based clustering analysis indicated the presence of two differentiated endemic male date palm genetic clusters, continental and insular, with the presence of introduced accessions originating from the Middle East. The diversity analysis in the insular region of Tunisia, which had never been performed before, revealed that this germplasm enclosed valuable endemic resources supporting the hypothesis of the presence of wild material. These findings are crucial for identifying interesting genotypes that can be integrated into international coordinated actions of Phoenix dactylifera L. breeding programs as well as for the protection and conservation of valuable resources.
Data from: genetic resources of macroalgae: development of an efficient method using microsatellite markers in non-model organisms
<p><span>Red and brown seaweeds are species with high ecological and economic importance. Here we report the feasibility of cost-effective molecular marker development in 6 species from different clades. Microsatellites markers of two brown seaweed species <em>Alaria esculenta</em>, <em>Pylaiella littoralis</em>, and of four red seaweed species <em>Calliblepharis jubata</em>, <em>Gracilaria gracilis</em>, <em>Gracilaria dura </em>and <em>Palmaria palmata</em> were identified and characterized using genomic sequences of Double-Digest Restriction site Associated DNA (ddRAD). A total of 64,623,186 reads were generated from two runs of multiplexed Illumina Miseq sequencing for which 30,636 reads containing microsatellites and 15,443 microsatellite loci with primers pairs were found. Five hundred seventy-six primers pairs were selected for amplification trials and levels of polymorphism. From the 338 that gave a positive amplification, 142 primers pairs were polymorphic. For genetic analyses two or three populations per species from 13 different geographic locations were used. A total of 28 usable polymorphic markers for <em>A. esculenta</em>, 18 for <em>P. littoralis</em>, 11 for <em>C. jubata</em>, 14 for <em>G. gracilis</em>, 21 for <em>G. dura </em>and 13 for <em>P. palmata </em>were developed. The overall number of alleles per locus ranged from 2 to 22. These 105 new microsatellite markers will be useful for further studies of population genetics, breeding programs and conservation genetics of these species. Compared with traditional approaches, our study yielded thousands of microsatellite loci in a short tim</span><span>e with affordable costs in six different species. This study based on ddRAD-sequencing for the development of microsatellite markers provides preliminary data u</span><span>sing a few individuals from two distinct populations on the genetic structure and reproduction mode of a non-model species as shown </span>with the detection of clonality for the two red algae, <em>C. jubata </em>and <em>G. dura</em> and the detection of highly genetically divergent populations corresponding probably to different cryptic species under the name of<em> P. littoralis</em>.</p>
Reference genome resources associated with the project: Functional genetic diversity is correlated with intensity of genetic drift in populations of an endangered rattlesnake
<p class="MsoNormal">Theory predicts that genetic erosion in small, isolated populations of endangered species can be assessed using estimates of neutral genetic variation reflecting long-term impacts of genetic drift, yet this widely used approach has been questioned in the genomics era. Here we leverage a chromosome-level assembly and whole genome resequencing data (N=110 individuals) from an endangered rattlesnake (<em>Sistrurus catenatus</em>) to evaluate the relationship between genome-wide neutral and functional diversity over long- and short-term timescales. As predicted for populations at long-term equilibrium, we found a positive correlation between population-level estimates of neutral genetic diversity (π) and the mean number of highly detrimental loss-of-function mutations, and a negative relationship between neutral genetic diversity and an estimate of genetic load. In contrast, we found only a weak, non-significant positive correlation between levels of neutral and adaptive variation. Additional analyses using estimates of drift at more recent time scales (> 100 generations) show expected correlations between both measures of genetic load, but a lack of a significant correlation with levels of adaptive variation. Individual-based demographic metrics that capture drift impacts over recent time scales confirm these results. Broadly, our results confirm that estimates of diversity and demography based on neutral genetic variation provide an accurate measure of a key component of genetic erosion – genetic load – in populations of a threatened vertebrate. Our findings also provide nuance to the neutral-functional diversity controversy by demonstrating that neutral genetic diversity is useful in predicting some, but not all, components of functional genetic diversity.</p>
Resource for teaching human genetics at the college level to combat harmful misconceptions
<p><span>How we teach human genetics matters for social equity. The biology curriculum appears to be a crucial locus of intervention for either reinforcing or undermining students' racial essentialist views. The Mendelian genetic models dominating textbooks, particularly in combination with racially inflected language sometimes used when teaching about monogenic disorders, can increase middle and high school students' racial essentialism and opposition to policies to increase equity. These findings are of particular concern given the increasing spread of racist misinformation online and misappropriation of human genomics research by white supremacists, who take advantage of low levels of genetics literacy in the general public. Encouragingly, however, teaching updated information about the geographic distribution of human genetic variation and the complex, multifactorial basis of most human traits, reduces students' endorsement of racial essentialism. The genetics curriculum is therefore a key tool in combating misinformation and scientific racism. Here, we describe a framework and example teaching materials for teaching students key concepts in genetics, human evolutionary history, and human phenotypic variation at the undergraduate level. This framework can be flexibly applied in biology and anthropology classes and adjusted based on time availability. Our goal is to provide college instructors with varying levels of expertise with a set of evidence-informed tools for teaching human genetics to combat scientific racism, including an evolving set of instructional resources, as well as learning goals and pedagogical approaches instructors can apply when teaching genetics. Additionally, we hope to generate conversation about integrating modern genetics into the undergraduate curriculum, in light of recent findings about the risks and opportunities associated with teaching genetics.</span></p>
Resource for teaching human genetics at the college level to combat harmful misconceptions
Open the record for dataset details and reuse information.
Data from: genetic resources of macroalgae: development of an efficient method using microsatellite markers in non-model organisms
Open the record for dataset details and reuse information.
Dissecting the genetic basis of variation in Drosophila sleep using a multiparental QTL mapping resource
Open the record for dataset details and reuse information.
Microsatellite exploration in the climbing hydrangea (Hydrangea petiolaris Siebold & Zucc.) transcriptome: A resource for population genetics and functional genomics
Open the record for dataset details and reuse information.
Reference genome resources associated with the project: Functional genetic diversity is correlated with intensity of genetic drift in populations of an endangered rattlesnake
Open the record for dataset details and reuse information.
Datasets of "Whole genome sequencing of European autochthonous and commercial pig breeds provides selection signatures of adaptation of genetic resources to different breeding and production systems"
<p>Results of the F<sub>ST</sub> and H<sub>P</sub> analyses.</p>
Datasets of the study: "Describing variability in pig genes involved in coronavirus infections: towards a One Health perspective in conservation of animal genetic resources"
<p><strong>Dataset description</strong></p> <p>Sequencing data (*.bam files) of four pig genes (<em>ACE2</em>, <em>ANPEP</em>, <em>DPP4</em> and <em>TMPRSS2</em>)<em> </em>that can serve as receptors or protease for priming the infection of coronaviruses.</p> <p>The datasets are related to 22 European pig breeds and wild boars (Alentejana, AL; Apulo-Calabrese, AC; Basque, BA; Bísara, BI; Black Slavonian, BS; Casertana, CA; Cinta Senese, CS; Gascon, GA; Krškopolje, KR; Lithuanian Indigenous Wattle, LIW; Lithuanian White Old Type, LWOT; Majorcan Black, MB; Mora Romagnola, MR; Moravka, MO; Nero Siciliano, NS; Sarda, SA; Schwäbisch-Hällisches Schwein, SHS; Swallow-Bellied Mangalitsa, SBMA; Turopolje, TU; Italian Duroc, IDU; Italian Large White, ILW; Italian Landrace, ILA; Wild Boar, WB). This work took advantage of a study design developed within the Horizon 2020 TREASURE project.</p> <p>Each folder contains *.bam files and the related indexes *.bai. The name of the investigated breed and gene is part of the file name (e.g. ILW.ACE2.bam identifies the sequencing data related to the ACE2 gene in the Italian Large White pig breed). Details of sequencing and the bioinformatic pipeline are below reported.</p> <p><strong>Sequencing data</strong></p> <p>A total of 22 DNA pools were constructed from the European pig breeds and one DNA pool was constructed from European wild boars, including in each pool 30 or 35 individual DNA samples pooled at equimolar concentration. For the 22 DNA pools, libraries were prepared and fed into an Illumina HiSeq X Ten sequencer for paired-end sequencing, obtaining 150 bp length reads. The wild boar DNA pool was sequenced from 250 bp fragment libraries, with 100 bp long paired-end reads, on the BGISeq 500 platform, following the provider’s procedures.</p> <p><strong>Data processing</strong></p> <p>Reads that were obtained from the sequenced libraries were cleaned by removing adapter sequences and filtering out sequences presenting more than 10% unknown bases (N) and/or containing low-quality bases (Q ≤ 5) over 50% of the total sequenced bases. Then, filtered high-quality reads were mapped on the latest version of the <em>Sus scrofa</em> reference genome (Sscrofa11.1; https://ftp.ncbi.nlm.nih.gov/genomes/all/GCF/000/003/025/GCF_000003025.6_Sscrofa11.1/GCF_000003025.6_Sscrofa11.1_genomic.fna.gz) using the BWA-MEM algorithm v.0.7.17 and the parameters for paired-end data. Picard v.2.1.1 (https://broadinstitute.github.io/picard/) was used to remove duplicated reads. Whole sequence data are available in the EMBL-EBI European Nucleotide Archive (ENA) repository (http://www.ebi.ac.uk/ena), under the study accession PRJEB36830. </p> <p>Reads covering the four genes (ACE2: NC_010461.5:12094853-12156275; ANPEP: NC_010449.5:55346083-55378881; DPP4: NC_010457.5:68655849-68748818; TMPRSS2: NC_010455.5:204871561-204907561) were extracted with samtools v.1.7 and exported as aligned, sorted and indexed *.bam files. Gene length includes UTRs and flanking regions of 5 kbp upstream [flanking (5’-UTR)] and downstream [flanking (3’-UTR)].</p>
Data from: Multi-objective optimization for plant germplasm collection conservation of genetic resources based on molecular variability
Germplasm collections play a significant role among strategies for conservation of diversity. It is common to select a core collection to represent the genetic diversity of a germplasm collection, in order to minimize the cost of conservation, while ensuring the maximization of genetic variation. We aimed to solve two main problems: (1) to select a set of individuals, from an in situ data set, that is genetically complementary to an existing germplasm collection, and (2) to define a core collection for a germplasm collection. We proposed a new multi-objective optimization (MOO) approach based on principles of systematic conservation planning (SCP) incorporating heterozygosity information; therefore, optimization takes genotypic diversity and variability patterns into account as well. As a case study, we used Dipteryx alata microsatellite loci information from two sources, an ex situ germplasm collection located at the Agronomy School of the Federal University of Goiás (UFG-AS), and an in situ data set composed of 642 sampled individual trees. We were able to identify within a population of several individuals, the exact accessions/samples that should be chosen in order to preserve the species diversity. We found that material from nine in situ individual trees are enough to complement the UFG-AS germplasm collection as it is, and that it is possible to define a core collection of 20 individual trees representing all studied genetic diversity. Moreover, we defined a method (a protocol) to deal with large amounts of accessions in the context of MOO. The proposed approach can be used to help constructing collections with maximal allelic richness and can also be extended to the in situ conservation. As far as we know, this is the first time that principles of SCP and the MOO approach are applied to the problem of complementing a germplasm collection and of finding a core collection for a germplasm collection.
Data from: Genetic resources of teak (Tectona grandis Linn. f.) – strong genetic structure among natural populations
Twenty-nine provenances of teak (Tectona grandis Linn. f.) representing the full natural distribution range of the species were genotyped with microsatellite DNA markers to analyse genetic diversity and population genetic structure. Provenances originating from the semi-moist east coast of India had the highest genetic diversity while provenances from Laos showed the lowest. In the eastern part of the natural distribution area, comprising Myanmar, Thailand and Laos, there was a strong clinal decrease in genetic diversity the further east the provenance was located. Overall, the pattern of genetic diversity supports the hypothesis that teak has its centre of origin in India, from where it spread eastwards. The analysis of molecular variance (AMOVA) gave an overall highly significant F st value of 0.227—population pairwise F st values were in the range 0.01–0.48. Applying the G″st differentiation parameter, the estimated overall differentiation was 0.632, implying a strong genetic structure among populations. A neighbour-joining (NJ) tree, using the pairwise population matrix of G″st values as input, contained three distinct groups: (1) the eight provenances from Thailand and Laos, (2) the Indian provenances from the dry interior and the moist west coast and (3) the provenances from northern Myanmar. The provenances from southern Myanmar were placed close to the root of the tree together with the three provenances from the semi-moist east coast of India. A Bayesian cluster analysis using the STRUCTURE software gave very similar results, with three main clusters, each containing two sub-clusters, while Bayesian cluster analysis in the Geneland software, exploiting the spatial coordinates of the provenances, resulted in five clusters in accordance with the former results. The implications of the findings for conservation and use of genetic resources of the species are discussed.
Data from: Complex genetic effects on early vegetative development shape resource allocation differences between Arabidopsis lyrata populations
Costs of reproduction due to resource allocation trade-offs have long been recognized as key forces in life history evolution, but little is known about their functional or genetic basis. Arabidopsis lyrata, a perennial relative of the annual model plant A. thaliana with a wide climatic distribution, has populations that are strongly diverged in resource allocation. In this study, we evaluated the genetic and functional basis for variation in resource allocation in a reciprocal transplant experiment, using four A. lyrata populations and F2 progeny from a cross between North Carolina (USA) and Norway parents, which had the most divergent resource allocation patterns. Local alleles at quantitative trait loci (QTL) at a North Carolina field site increased reproductive output while reducing vegetative growth. These QTL had little overlap with flowering date QTL. Structural equation models incorporating QTL genotypes and traits indicated that resource allocation differences result primarily from QTL effects on early vegetative growth patterns, with cascading effects on later vegetative and reproductive development. At a Norway field site, North Carolina alleles at some of the same QTL regions reduced survival and reproductive output components, but these effects were not associated with resource allocation trade-offs in the Norway environment. Our results indicate that resource allocation in perennial plants may involve important adaptive mechanisms largely independent of flowering time. Moreover, the contributions of resource allocation QTL to local adaptation appear to result from their effects on developmental timing and its interaction with environmental constraints, and not from simple models of reproductive costs.
Data from: Resource allocation during ontogeny is influenced by genetic, developmental, and ecological factors in the horned beetle, Onthophagus taurus
Resource allocation trade-offs arise when developing organs are in competition for a limited pool of resources to sustain growth and differentiation. Such competition may constrain the maximal size to which structures can grow and may force a situation in which the evolutionary elaboration of one structure may only be possible at the expense of another. However, recent studies have called into question both the consistency and evolutionary importance of resource allocation trade-offs. This study focuses on a well-described trade-off between the horns and eyes of Onthophagus beetles and assesses the degree to which it is influenced by genetic, developmental and ecological conditions. Contrary to expectations, we observed that trade-off signatures (i) were mostly absent within natural populations, (ii) mostly failed to match naturally evolved divergences in horn investment among populations, (iii) were subject to differential changes in F1 populations derived from divergent field populations and (iv) remained largely unaffected by developmental genetic manipulations of horn investment. Collectively, our results demonstrate that populations subject to different ecological conditions exhibit different patterns of, and differential plasticity in, resource allocation. Further, variation in ecological conditions, rather than canalized developmental mechanisms, may determine whether and to what degree morphological structures engage in resource allocation trade-offs.
Data from: Does local adaptation to resources explain genetic differentiation among Daphnia populations?
Substantial genetic differentiation is frequently observed among populations of cyclically parthenogenetic zooplankton despite their high dispersal capabilities and potential for gene flow. Local adaptation has been invoked to explain population genetic differentiation despite high dispersal, but several neutral models that account for basic life history features also predict high genetic differentiation. Here, we study genetic differentiation among four populations of Daphnia pulex in east central Illinois. As with other studies of Daphnia, we demonstrate substantial population genetic differentiation despite close geographic proximity (< 50 km; mean θ = 0.22). However, we explicitly tested, and failed to find evidence for, the hypothesis that local adaptation to food resources in these populations. Recognizing that local adaptation can occur in traits unrelated to resources, we estimated contemporary migration rates (m) and tested for admixture to evaluate the hypothesis that observed genetic differentiation is consistent with local adaptation to other untested ecological factors. Using Bayesian assignment methods, we detected migrants in three of the four study populations including substantial evidence for successful reproduction by immigrants in one pond, allowing us to reject the hypothesis that local adaptation limits gene flow for at least this population. Thus, we suggest that local adaptation does not explain genetic differentiation among these Daphnia populations, and that other factors related to extinction / colonization dynamics, a long approach to equilibrium FST, or substantial genetic drift due to a low number of individuals hatching from the egg bank each season may explain genetic differentiation.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.