Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,153
datasets available to search
ShareScore release 0.9.0
Dataset results
5,153 results for “Genetic data”
Data from: Do genetic loci that cause reproductive isolation in the lab inhibit gene flow in nature?
<p>The genetic dissection of reproductive barriers between diverging lineages provides enticing clues into the origin of species. One strategy uses linkage analysis in experimental crosses to identify genomic locations involved in phenotypes that mediate reproductive isolation. A second framework searches for genomic regions that show reduced rates of exchange across natural hybrid zones. It is often assumed that these approaches will point to the same loci, but this assumption is rarely tested. In this perspective, we discuss the factors that determine whether loci connected to postzygotic reproductive barriers in the laboratory are inferred to reduce gene flow in nature. We synthesize data on the genetics of postzygotic isolation in house mice, one of the most intensively studied systems in speciation genetics. In a rare empirical comparison, we measure the correspondence of loci tied to postzygotic barriers via genetic mapping in the laboratory and loci at which gene flow is inhibited across a natural hybrid zone. We find no evidence that the two sets of loci overlap beyond what is expected by chance. In light of these results, we recommend avenues for empirical and theoretical research to resolve the potential incongruence between the two predominant strategies for understanding the genetics of speciation.</p>
Data for: Harvest and decimation affect genetic drift and the effective population size in wild reindeer
<p>Harvesting and culling are methods used to monitor and manage wildlife diseases. An important consequence of these practices is a change in the genetic dynamics of affected populations that may threaten their long-term viability. The effective population size (N<sub>e</sub>) is a fundamental parameter for describing such changes as it determines the amount of genetic drift in a population. Here, we estimate N<sub>e</sub> of a harvested wild reindeer population in Norway. Then we use simulations to investigate the genetic consequences of management efforts for handling a recent spread of chronic wasting disease, including increased adult male harvest and population decimation. The N<sub>e</sub>/N ratio in this population was found to be 0.124 at the end of the study period, compared to 0.239 in the preceding 14-year period. The difference was caused by increased harvest rates with a high proportion of adult males (older than 2.5 years) being shot (15.2 % in 2005-2018 and 44.8 % in 2021). Increased harvest rates decreased N<sub>e</sub> in the simulations, but less sex-biased harvest strategies had a lower negative impact. For harvest strategies that yield stable population dynamics, shifting the harvest from calves to adult males and females increased N<sub>e</sub>. Population decimation always resulted in decreased genetic variation in the population, with higher loss of heterozygosity and rare alleles with more severe decimation or longer periods of low population size. A very high proportion of males in the harvest had the most severe consequences for the loss of genetic variation. This study clearly shows how the effects of harvest strategies and changes in population size interact to determine the genetic drift of a managed population. The long-term genetic viability of wildlife populations subject to disease will also depend on the population impacts of the disease and how these interact with management actions.</p>
Figure 6 in Description of a new species of the genus Rana (Anura: Ranidae) from western Guizhou, China, integrating morphological and molecular genetic data
Figure 6. Variation of the live adult male paratype GZNU20220705001 of Rana zhijinensis Luo, Xiao & Zhou, sp. nov. A. Dorsolateral view; B. Dorsal view; C. Ventral view.
Figure 1 in Description of a new species of the genus Rana (Anura: Ranidae) from western Guizhou, China, integrating morphological and molecular genetic data
Figure 1. Sampling localities of Rana zhijinensis Luo, Xiao & Zhou, sp. nov., R. culaiensis, R. hanluica, and R. omeimontis in Guizhou Province, China. A. Guiguo Town, Zhijin County; B. Supu Town, Qianxi County; C. Zhujianshan Nature Reserve, Huangping County; D. Leigongshan National Nature Reserve, Leishan County.
Figure 2 in Description of a new species of the genus Rana (Anura: Ranidae) from western Guizhou, China, integrating morphological and molecular genetic data
Figure 2. Phylogenetic tree based on three mitochondrial genes and six nuclear genes. A. Maternal tree; B. Nuclear gene tree. In both phylogenetic tree, ultrafast bootstrap support (UFB) values from ML analyses/Bayesian posterior probabilities (BPP) from BI analyses are given beside nodes. Scale bars denote nucleotide substitutions per sites for mitochondrial and nuclear genes.
Figure 5 in Description of a new species of the genus Rana (Anura: Ranidae) from western Guizhou, China, integrating morphological and molecular genetic data
Figure 5. Morphological features of the live adult male holotype GZNU2018081606 of Rana zhijinensis Luo, Xiao & Zhou, sp. nov. A. Dorsolateral view; B. Dorsal view; C. Ventral view; D. Egg cluster; E. Ventral view of hand and dark gray-blackish nuptial pad; F. Ventral view of foot.
Figure 4 in Description of a new species of the genus Rana (Anura: Ranidae) from western Guizhou, China, integrating morphological and molecular genetic data
Figure 4. Haplotype networks of Rana zhijinensis Luo, Xiao & Zhou, sp. nov. and its related species constructed based on the nuclear gene sequences. Different species of the R. japonica group are shown as different colors.
Figure 3 in Description of a new species of the genus Rana (Anura: Ranidae) from western Guizhou, China, integrating morphological and molecular genetic data
Figure 3. Phylogenetic tree based on four mitochondrial genes and six nuclear genes. In this phylogenetic tree, UFB from ML analyses/ BPP from BI analyses are given beside nodes. The scale bar represents 0.03 nucleotide substitutions per site. Red lines represent species delimitation results of bPTP and BPP.
Figure 1 in Why we should develop guidelines and quantitative standards for using genetic data to delimit subspecies for data-poor organisms like cetaceans
Figure 1. Depiction of the divergence of lineages with four times (T1–T4) chosen to illustrate different levels of biological organization. At T1 the yellow lineage is found across the distribution and although there are likely Demographically Independent Populations (DIPs) that differ in frequencies of the blue, yellow, and red lineages, there are no discontinuities. At T2 some lineages may be diagnosable but likely do not yet appear to be separate lineages. At T3 three groups (the blue/green, yellow, and orange/red lineages) meet the subspecies definition (they are diagnosable and appear to be diverging separately). The divergence level is not sufficient that reconvergence can be ruled out. Between T3 and T4, barriers to gene flow change such that the yellow lineage comes into contact with the blue/green and red-dominated lineages. Blue has diverged in a manner by which gene flow does not resume and the green/yellow lineage dies out. The yellow lineage reconverges and persists alongside the red lineage with a small level of gene flow (orange). At T4 the blue lineage is a species evolving separately from the yellow/red species. The yellow/red species has two subspecies that are both diagnosable and partially diverged.
Figure 3 in Guidelines and quantitative standards to improve consistency in cetacean subspecies and species delimitation relying on molecular genetic data
Figure 3. Flow diagram for subspecies delineation using combined quantitative and qualitative standards. The threshold values assume the user is evaluating a case relying on mtDNA control region data. Percent Diagnosable (PD) is the smallest strata-specific correct classification score in a given comparison (e.g., PD50 in two-strata comparisons in Archer et al. 2017). The second box in the second row (other evidence to meet subspecies definition) allows for subspecies delineation when both conditions are not met using mtDNA. This box could be used either for the case when one condition is met and one unmet or when both just barely miss meeting the standards. For example, consider the case with PD <95% and dA> 0.004. Diagnosability could be achieved with morphological data or nuclear data that are sufficient for subspecies but not for full species.
Figure 2. A in Guidelines and quantitative standards to improve consistency in cetacean subspecies and species delimitation relying on molecular genetic data
Figure 2. A comparison of the pairs of populations (red triangles), subspecies (green squares) and species (blue circles) estimated by Rosel et al. (2017a). Net nucleotide divergence (dA) is shown on a natural log scale to better illustrate differences between the pairwise comparisons at low levels of divergence. Bars show the central 95th-pecentile of the estimate distributions. The solid vertical line at dA = 0.020 delimits all but one species and correctly excludes all subspecies pairs. The vertical dashed line at dA = 0.004 delimits all populations from the higher taxonomic levels and correctly delimits seven of eleven subspecies. The horizontal dashed lines are two potential thresholds for percent diagnosable (80% and 95%) that are discussed in the text.
Data from: Context matters: the landscape matrix determines the population genetic structure of temperate forest herbs across Europe
<p>Context. Plant populations in agricultural landscapes are mostly fragmented and their functional connectivity often depends on seed and pollen dispersal by animals. However, little is known about how the interactions of seed and pollen dispersers with the agricultural matrix translate into gene flow among plant populations.</p> <p>Objectives. We aimed to identify effects of the landscape structure on the genetic diversity within, and the genetic differentiation among, spatially isolated populations of three temperate forest herbs. We asked, whether different arable crops have different effects, and whether the orientation of linear landscape elements relative to the gene dispersal direction matters.</p> <p>Methods. We analysed the species' population genetic structures in seven agricultural landscapes across temperate Europe using microsatellite markers. These were modelled as a function of landscape composition and configuration, which we quantified in buffer zones around, and in rectangular landscape strips between, plant populations.</p> <p>Results. Landscape effects were diverse and often contrasting between species, reflecting their association with different pollen- or seed dispersal vectors. Differentiating crop types rather than lumping them together yielded higher proportions of explained variation. Some linear landscape elements had both a channelling and hampering effect on gene flow, depending on their orientation.</p> <p>Conclusions. Landscape structure is a more important determinant of the species' population genetic structure than habitat loss and fragmentation <i>per se</i>. Landscape planning with the aim to enhance the functional connectivity among spatially isolated plant populations should consider that even species of the same ecological guild might show distinct responses to the landscape structure.</p>
Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: discovery cohort meta data and parsed TCR repertoire data
<p>Meta data corresponding the the discovery cohort for the paper, "Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities" by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include: </p> <p>(1) a file mapping the SNP data subject IDs to the TCR repertoire data subject IDs (gwas_id_mapping.tsv)<br> (2) a file including the PCAir PCs, self-reported ancestry, and genomic ancestry for each subject (all_pc_air.txt)<br> (3) a file including the PCAir variance explained by each PC (all_pc_air_variance.txt)<br> (3) a file including the SNP ID, chromosome, hg19 position, allele, rsid, and quality control metrics for each SNP in the SNP array (emerson_snp_rs_data.tsv)<br> (4) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (5) a file including predicted TRBD2 allele genotypes for each subject (emerson_trbd2_alleles.tsv)<br> (6) Parsed TCRB repertoire data. These raw data were first published in Emerson et. al, <em>Nature Genetics </em>2017. (emerson_parsed_tcrb.tgz)</p> <p><strong>Corresponding discovery cohort raw TCR repertoire data is available here: </strong>https: //doi.org/10.21417/B7001Z (ImmuneACCESS database)<br> <strong>Corresponding discovery cohort SNP data is available here:</strong> https: //www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs001918.v1.p1 (The database of Genotypes and Phenotypes, accession number: phs001918)<br> <br> <strong>Software tools designed to work with these data are available here:</strong> https://github.com/phbradley/tcr-gwas</p>
Genetic diversity of wild and cultivated Coffea canephora in northeastern DR Congo and the implications for conservation - Additional Data
<p>List of wild and cultivated <em>Coffea canephora </em>accessions from northeastern Democratic Republic of the Congo included in Vanden Abeele et al. 2021 - American Journal of Botany, and the corresponding alleles for each of the 18 microsatellite markers (0 indicates missing alleles).</p>
[Data from:] Genetic Analysis Reveals Three Novel QTLs Underpinning a Butterfly Egg-Induced Hypersensitive Response-Like Cell Death in Brassica Rapa
<p><strong>Background</strong></p> <p>Cabbage white butterflies (<em>Pieris</em> spp.) can be severe pests of <em>Brassica</em> crops such as Chinese cabbage, Pak choi (<em>Brassica rapa</em>) or cabbages (<em>B. oleracea</em>). Eggs of <em>Pieris</em> spp. can induce a hypersensitive response-like (HR-like) cell death which reduces egg survival in the wild black mustard (<em>B. nigra</em>). Unravelling the genetic basis of this egg-killing trait in <em>Brassica</em> crops could improve crop resistance to herbivory, reducing major crop losses and pesticides use. Here we investigated the genetic architecture of a HR-like cell death induced by <em>P. brassicae</em> eggs in <em>B. rapa.</em></p> <p><strong>Results</strong></p> <p>A germplasm screening of <em>B. rapa</em> 56 accessions, representing the genetic and geographical diversity of a <em>B. rapa</em> core collection, showed phenotypic variation for cell death. An image-based phenotyping protocol was developed to accurately measure size of HR-like cell death and was then used to identify two accessions that consistently showed weak (R-o-18) or strong cell death response (L58). Screening of 160 RILs derived from these two accessions resulted in three novel QTLs for P<em>ieris</em> b<em>rassicae-</em>induced cell death on chromosomes A02 (<em>Pbc1</em>), A03 (<em>Pbc2</em>), and A06 (<em>Pbc3</em>). The three QTLs <em>Pbc1-3</em> contain cell surface receptors, intracellular receptors and other genes involved in plant immunity processes, such as ROS accumulation and cell death formation. Synteny analysis with <em>A. thaliana</em> suggested that <em>Pbc1</em> and <em>Pbc2</em> are novel QTLs associated with this trait, while <em>Pbc3</em> contains also LecRK-I.1, a gene of <em>A. thaliana</em> previously associated with cell death induced by a <em>P. brassicae</em> egg extract.</p> <p><strong>Conclusions</strong></p> <p>This study provides the first genomic regions associated with the <em>Pieris</em> egg-induced HR-like cell death in a <em>Brassica</em> crop species. It is a step closer towards unravelling the genetic basis of an egg-killing crop resistance trait, paving the way for breeders to further fine-map and validate candidate genes.</p>
Data from: Chrysolaena obovata, A SPECIES NATIVE OF BRAZILIAN CERRADO: GENETIC DIVERSITY AND STRUCTURE OF NATURAL POPULATIONS AND POTENTIAL FOR INULIN PRODUCTION
<p><em>Chrysolaena obovata</em> (Less.) M. Dematteis, an herbaceous Asteraceae species widely distributed across different Brazilian Cerrado physiognomies, has underground organs, named rhizophores, that accumulate high concentrations of inulin-type fructans. These carbohydrates are recognized as beneficial soluble fibers for human health and are currently used in the food and pharmaceutical industries. Considering that fructans, in addition to their economic potential, provide plants with greater tolerance to drought, heat and cold, it is important to understand whether their metabolism is conserved in natural populations. In this work, we aimed to investigate if the levels of genetic diversity in the populations studied allow the selection of localities with a high genetic base and higher fructan content for future programs of <em>in</em> <em>situ</em> conservation and genetic improvement for inulin production. Therefore, we characterized the diversity, structure, and gene flow of seven natural populations from Brazilian Cerrado, using nine microsatellite loci (SSR). In addition, we compared whether the fructan composition varied between populations of different Cerrado phytophysiognomies. Overall, we found that <em>C. obovata</em> populations exhibited moderate levels of genetic diversity, low genetic differentiation, and high gene flow. This study identified two populations with less genetic diversity and therefore, greater attention should be given to conservation programs including these populations. Fructan metabolism is conserved in all populations, indicating that <em>C. obovata</em> is an important genetic resource with high potential for inulin production.</p> <p><strong>File descriptions</strong></p> <p>Population_code.txt - Contains a matrix that indicates the population_code, Population_name, Brazilian-state, Phytophysiognomy, Collection coordinates and Altitudes (m).</p> <p>Date_ Diaz et al.xlsx – Contains Genotypes crude of the individuals analyzed. Primer used for nine microsatellite loci (Camacho <em>et al</em> 2017). </p> <p>Carbohydrates_Diaz et al - Contains data for carbohydrates in <em>C. obovata</em> plant rhizophores in each population (BRA, UB, SD, SP).</p> <p><strong>Location: Brazilian Cerrado</strong></p>
Data and scripts for: Genetic dissection of seasonal vegetation index dynamics in maize through aerial based high-throughput phenotyping
<p>Plant phenotyping under field conditions plays an important role in agricultural research. Efficient and accurate high-throughput phenotyping strategies enable a better connection between genotype and phenotype. Unmanned aerial vehicle-based high-throughput phenotyping platforms (UAV-HTPPs) provide novel opportunities for large-scale proximal measurement of plant traits with high efficiency, high resolution, and low cost. The objective of this study was to use time series normalized difference vegetation index (NDVI) extracted from UAV-based multispectral imagery to characterize its pattern across development and conduct genetic dissection of NDVI in a large maize population. The time series NDVI data from the multispectral sensor were obtained at 5 time points across the growing season for 1,752 diverse maize accessions with a UAV-HTPP. Cluster analysis of the acquired measurements classified 1,752 maize accessions into 2 groups with distinct NDVI developmental trends. To capture the dynamics underlying these static observations, penalized-splines (P-splines) model was used to obtain genotype-specific curve parameters. Genome-wide association study (GWAS) using static NDVI values and curve parameters as phenotypic traits detected signals significantly associated with the traits. Additionally, GWAS using the projected NDVI values from the P-splines models revealed the dynamic change of genetic effects, indicating the role of gene-environment interplay in controlling NDVI across the growing season. Our results demonstrated the utility of ultra-high spatial resolution multispectral imagery, as that acquired using a UAV-based remote sensing, for genetic dissection of NDVI.</p>
Data from: Seeing the Forest for the trees: Assessing genetic offset predictions from Gradient Forest
<p><span>Gradient Forest (GF) is a machine learning algorithm designed to analyze spatial patterns of biodiversity as a function of environmental gradients. An offset measure between the GF predicted environmental association of adapted alleles and a new environment (GF Offset), is increasingly being used to predict the loss of environmentally adapted alleles under rapid environmental change, but remains mostly untested for this purpose. Here we explore the robustness of GF Offset to assumption violations, and its relationship to measures of fitness, using SLiM simulations with explicit genome architecture and a spatial metapopulation. We evaluate measures of GF Offset in: (1) a neutral model with no environmental adaptation; (2) a monogenic "population genetic" model with a single environmentally adapted locus; and (3) a polygenic "quantitative genetic" model with two adaptive traits, each adapting to a different environment. We found GF Offset to be broadly correlated with fitness offsets under both single locus and polygenic architectures. However, neutral demography, genomic architecture, and the nature of the adaptive environment can all confound relationships between GF Offset and fitness. GF Offset is a promising tool, but it is important to understand its limitations and underlying assumptions, especially when used in the context of predicting maladaptation.</span></p>
Data accompanying Polyphenisms and polymorphisms: genetic variation in plasticity and color variation within and among bluefin killifish populations
<p>The presence of stable color polymorphisms within populations begs the question of how genetic variation is maintained. Consistent variation among populations in coloration, especially when correlated with environmental variation, raises questions about whether environmental conditions affect either the fulcrum of those balanced polymorphisms, the plastic expression of coloration, or both. Color patterns in male bluefin killifish provoke both types of questions. Red and yellow morphs are common in all populations. Blue males are more common in tannin-stained swamps relative to clear springs. Here we combined crosses with a manipulation of light to explore how genetic variation and phenotypic plasticity shape these patterns. We found that the variation in coloration is attributable mainly to two axes of variation: (1) a red-yellow axis with yellow being dominant to red, and (2) a blue axis that can override red-yellow and is controlled by genetics, phenotypic plasticity, and genetic variation for phenotypic plasticity. The variation among populations in plasticity suggests it is adaptive in some populations but not others. The variation among sires in plasticity within the swamp population suggests balancing selection may be acting not only on the red-yellow polymorphism but also on plasticity for blue coloration.</p>
Data and Software for "Complex Dynamics in a Synchronized Cell-Free Genetic Clock"
<p>Contains raw data, analysis scripts, simulation scripts, device operation software for the publication: "Complex Dynamics in a Synchronized Cell-Free Genetic Clock".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.