Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
899
datasets available to search
ShareScore release 0.9.0
Dataset results
899 results for “allele”
Evolution of allele frequencies in two chicken lines divergently selected for meat ultimate pH
<p>Two lines of chicken were divergently selected during 5 generations for high or low meat ultimate pH. Genotypes at about 50K SNPs were obtained for a sample of individuals in each line and at each generation (including the founder population G0). The present dataset provides the allele frequencies for all SNP and generations in the two lines, at plink frq.strat format. The position of the SNP on the chicken genome are provided in another file at plink .map format.</p> <p>These data were first ued in the following publication:</p> <p>Le Bihan-Duval, E., Hennequet-Antier, C., Berri, C., Beauclercq, S. A., Bourin, M. C., Boulay, M., ... & Boitard, S. (2018). Identification of genomic regions and candidate genes for chicken meat ultimate pH by combined detection of selection signatures and QTL. <em>BMC genomics</em>, <em>19</em>(1), 294.</p>
Per-gene per-strain data: expression divergence between strains and alleles in F1s in wild C. elegans
<p>This dataset comprises p<span>er-gene per-strain data (used to perform all analyses and generate all figures), including regulatory pattern and inheritance mode classifications and underlying statistical differential expression results</span>.</p> <p>This is supplemental data for the linked preprint/publication describing insights derived from comparing gene expression (RNA-seq) between seven wild <em>C. elegans</em> strains and the laboratory reference strain N2, as well as the allelic expression of the wild and N2 alleles in F1s of crosses between all these wild strains and the reference strain.</p> <p>The PDF file <code>column_names_descriptions_worm_ase_data_pergene_perstrain.pdf</code> and excel spreadsheet <code>column_names_descriptions_worm_ase_data_pergene_perstrain.xlsx</code> serve as READMEs for the data file by providing details of the data held in each column of the data file <code>worm_ase_data_pergene_perstrain.txt.gz</code></p> <p>If you use this dataset (we hope someone does!), please cite the latest version of the accompanying preprint/publication.</p> <p>To query each gene in a user-friendly, visual format, see our shiny app <a href="https://wildworm.biosci.gatech.edu/ase/" target="_blank" rel="noopener">https://wildworm.biosci.gatech.edu/ase/</a></p>
Data from: "From cultivar mixtures to allelic mixtures: opposite effects of allelic richness between genotypes and genotype richness in wheat"
<p><em><strong>Data and code used for the study : "From cultivar mixtures to allelic mixtures: opposite effects of allelic richness between genotypes and genotype richness in wheat".</strong></em></p> <p>The script "Manuscript_Analyses.R" contains all code for the statistical analysis presented in the manuscript (main text & supplementary information). This script uses files produced in the folder "Locus-by-locus analysis" as inputs, and "manhattan_custom.R" as a source function ("manhattan_custom.R" is used to highlight SNPs in a given interval and to write specified SNPs name on manhattan plots). The file "Traits_monocultures.csv" contains the 20 functional traits measured on the 179 monoculture plots (see Supplementary Methods for more information on trait measurement). This file is used as an input in the script "Manuscript_Analyses.R".</p> <p>The "Locus-by-locus analysis" folder contains all analyses conducted to test the effect of allelic richness on the four variables of interest: Grain Yield (GY, g/m²), Spike Number per m² (SNb, nb spikes/m²), Thousand Kernel Weight (TKW, g), and Septoria tritici blotch (STB) severity. The locus-by-locus analysis is performed with the script "Allelic_richness_locus_by_locus_analysis.R". This analysis generates a list of .csv files with one file per chromosome. Each file contains the pvalues and estimated effect sizes of the tested SNPs for the given chromosome. These output files are stored in folders named after the variables for which the effect of allelic richness was tested ("RAW_GY", "RAW_SNb", "RAW_TKW", and "RAW_severity"). The script "Allelic_richness_locus_by_locus_output_processing.R" combines all .csv files into a single dataframe and produces three diagnostic plots: Manahattan plots, histograms of p-value distributions, and p-value q-q plots. p-value thresholds were computed based on a Family-Wise Error Rate of 5% using the Galwey correction. This is done in the "pvalue_thresholds" folder with the "Meff_computation.R" script. "Meff_computation.R" uses the "Meff_function.R" as a source function and generates "GY_thresholds.csv" and "STB_thresholds.csv" as outputs (these files contains different thresholds computed according to different methods but we only retained the Galwey method (most recent) for the analyses. Since GY, SNb, and TKW were analyzed with the same number of SNPs (~19K), we used the same significance threshold for the three variables ("GY_thresholds.csv"), whereas we computed a different thresholds for STB ("STB_thresholds.csv") for which we could only include ~6K SNPs in the analysis. The "geno_pos.csv" file contains the physical positions of the SNPs.</p> <p>Upstream the locus-by-locus analysis, phenotypic and genotypic files are prepared in the "Phenoytpic file preparation" and "Genotypic file preparation" folders, respecively.</p> <p>The phenotypic file preparation includes the correction of yield-related variables (GY, SNb, and TKW) for spatial auto-correlation in the "Spatial_analyses_YLD_variables" folder, and the computation of plot-level variables from individual-level variables with the "Allelic_richness_phenotypic_file_prep.R" script. In this script, we compute both absolute plot values (termed "RAW_...) and relative plot values (termed "RYT_..., only for mixture plots). All phenotypic files have the same structure with the same first 6 columns: "focal" = identity of the focal genotype (the one for which the variable is measured, only relevant for variables measured at the individual-level), "neighbor" = identity of the neighbor genotype (the neighbor of the genotype for which the variable is measured, only relevant for variables measured at the individual-level), "pair" = identity of the genotypic pair (combines the identity of the focal and the neighbor genotypes), "assoc" = type of plot ("M" = monoculture or pure stand plot, "P" = mixture plot), "row" = position of the plot along the smallest dimension of the grid (see Figure 1), "column" = position of the plot along the largest dimension of the grid (see Figure 1).</p> <p>The genotypic file preparation is done with the "Allelic_richness_genotypic_file_prep.R" script and includes SNP filtering, computation of matrices of allelic richness, and computation of matrices of genetic similarity between genotypic pairs. The analysis is done separatly for yield-related variables and for STB severity since the two types of variable were not measured on the same set of plots.</p>
OlSiFaComp : A database for Olea: Olive Self-Incompatibility Flower Allelic Composition
<p><strong>Introduction</strong> : Data on olive cross and selfing studies were numerous and dispersed in literature. Varieties were list in lines, and each line correspond to one data – bag, pollen test, paternity tests, ...To compare data in bags standardization was achieved based on 100 hermaphroditic flowers.</p> <p><strong>Materials and methods</strong> : Data from publications were recorded keeping the name of origin for each variety. Varieties in pair wise combinations in crosses enabled to decipher the S-allele pair (Breton and Bervillé 2012, Farinelli et al. 2014) leading to attribute the PASI pair. The G group (Saumitou-Laprade et al. 2017) was recorded from Mariotti et al. 2021.</p> <p><strong>Results</strong> : The DSI pair and the PASI pair were introduced and sorting data by screening G1xG2 (1 for fruit) and G2xG1 (1) whereas by screening G1xG1 (0) and G2XG2 (0) show that all crosses display fruit. The DSSM reconciles DSI and PASI to explain Self-incompatibility. Selfing was shown appearing when crosses were 1-0 or 0-1, but not when 1-1, whereas some expected 0-0 combinations lead to fruit. Moreover, paternity tests (column embryo) revealed most of the time a father compatible in DSI , but incompatible in PASI, this is due to DS-D, that shows compatible pollen is insufficient.</p> <p>The database is useful to check whether the variety has already been studied for SI. It enables to check the homogeneity of cross data in literature.</p>
Alliance of Genome Resources Alleles
<p>Tab separated formatted spreadsheets of allele annotations from the Alliance of Genome Resources.</p> <p>Files include annotations for</p> <ul> <li>Caenorhabditis elegans (nematode; NCBITaxon 6239)</li> <li>Danio rerio (zebrafish;NCBITaxon 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBITaxon 7227)</li> <li>Mus musculus (mouse; NCBITaxon 10090)</li> <li>Rattus norvegicus (rat; NCBITaxon 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBITaxon 559292 )</li> </ul>
Code repository for: Base editing mutagenesis maps functional alleles to tune human T cell activity
<p>Jupyter notebook and supplemental datasets required to created critical figures for the publication.</p>
Data of: Imputation-free reconstructions of three-dimensional chromosome architectures in human diploid single-cells using allele-specified contacts
Open the record for dataset details and reuse information.
Data from: Only rare classical MHC-I alleles are highly expressed in the European house sparrow
Open the record for dataset details and reuse information.
Benefits and limits of phasing alleles for network inference of allopolyploid complexes
Open the record for dataset details and reuse information.
Distinct signals of clinal and seasonal allele frequency change at eQTLs in Drosophila melanogaster
Open the record for dataset details and reuse information.
Reconstructing NOD-like receptor alleles with high internal conservation in Podospora anserina using long-read sequencing
Open the record for dataset details and reuse information.
Multi-allele species reconstruction using ASTRAL
Open the record for dataset details and reuse information.
Data from: Spatial patterns and rarity of the white-phased 'Spirit Bear' allele reveals gaps in habitat protection
<p>Preserving genetic and phenotypic diversity can help safeguard not only biodiversity but also cultural and economic values.</p> <p>Here, we present data that emerged from Indigenous-led research at the intersection of evolution and ecology to support conservation planning of a culturally salient, economically valuable, and rare phenotypic variant. We addressed three conservation objectives for the white-phased 'Spirit bear' polymorphism, a rare and endemic white-coated phenotype of black bear (Ursus americanus) in Kitasoo/Xai'xais and Gitga'at Territories and beyond in coastal British Columbia, Canada. First, we used non-invasively collected hair samples (n = 385 bears over ~18,000 km<sup>2</sup>) to assess the spatial variation in the frequency of the allele that controls the white-coloured morph (mc1r). Second, we compared our observed allele frequencies at mc1r with those expected under Hardy-Weinberg equilibrium. Finally, we examined how well current protected areas in the region aligned with spatial hotspots of Spirit bear alleles.</p> <p>We found that landscape-level allele frequency was lower than previously reported. For example, our systematic sampling estimated a frequency of 0.25 (95% CI 0.13-0.41) on Gribbell Island compared with the previously reported estimate of 0.56. Also, in contrast with previous reports, we failed to detect a statistically significant departure from Hardy-Weinberg equilibrium at mc1r, which calls into question the previously-posited role of homozygote gene flow, heterozygote disadvantage, and positive assortative mating in the maintenance of this polymorphism. Finally, we found a discrepancy between the placement of protected areas and the 90th percentile hotspots (upper 10% of all estimated values) of Spirit bear alleles, with ~50% of hotspots falling outside of protected areas.</p> <p>These results provide new insight into hypotheses related to the maintenance of this rare polymorphism, and directly relevant information to support evidence-based opportunities for Indigenous Nations of the area to attend to gaps in conservation planning.</p>
Severe inbreeding depression is predicted by the "rare allele load" in Mimulus guttatus
<p>Most flowering plants are hermaphroditic and experience strong pressures to evolve self‐pollination (automatic selection, reproductive assurance). Inbreeding depression (ID) can oppose selection for selfing, but it remains unclear if ID is typically strong enough to maintain outcrossing. To measure the full cost of sustained inbreeding on fitness, and its genomic basis, we planted highly homozygous, fully genome‐sequenced inbred lines of yellow monkeyflower (<i>Mimulus guttatus)</i> in the field next to outbred plants from crosses between the same lines. The cost of full homozygosity is severe: 65% for survival, 86% for lifetime seed production. Accounting for the unmeasured effect of lethal and sterile mutations, we estimate that the average fitness of fully inbred genotypes is only 3–4% that of outbred competitors. The genome sequence data provides no indication of simple overdominance, but the number of rare alleles carried by a line, especially within rare allele clusters nonrandomly distributed across the genome, is a significant negative predictor of fitness measurements. These findings are consistent with a deleterious allele model for ID. High variance in rare allele load among lines and the genomic distribution of rare alleles both suggest that migration might be an important source of deleterious alleles to local populations.</p>
Data from: Risk alleles for tuberculosis infection associate with reduced immune reactivity in a wild mammalian host
Integrating biological processes across scales remains a central challenge in disease ecology. Genetic variation drives differences in host immune responses, which, along with environmental factors, generates temporal and spatial infection patterns in natural populations that epidemiologists seek to predict and control. However, genetics and immunology are typically studied in model systems, whereas population-level patterns of infection status and susceptibility are uniquely observable in nature. Despite obvious causal connections, organizational scales from genes to host outcomes to population patterns are rarely linked explicitly. Here we identify two loci near genes involved in macrophage (phagocyte) activation and pathogen degradation that additively increase risk of bovine tuberculosis infection by up to 9-fold in wild African buffalo. Furthermore, we observe genotype-specific variation in IL-12 production indicative of variation in macrophage activation. Here we provide measurable differences in infection resistance at multiple scales by characterizing the genetic and inflammatory variation driving patterns of infection in a wild mammal.
Early Onset TAAD cohort logR Ratio and B allele frequency data
<p>Recurrent Rare Genomic Copy Number Variants and Bicuspid Aortic Valve Are Enriched in Early Onset Thoracic Aortic Aneurysms and Dissections</p> <p>Abstract:</p> <p>Thoracic Aortic Aneurysms and Dissections (TAAD) are a major cause of death in the United States. The spectrum of TAAD ranges from genetic disorders, such as Marfan syndrome, to sporadic isolated disease of unknown cause. We hypothesized that genomic copy number variants (CNVs) contribute causally to early onset TAAD (ETAAD). We conducted a genome-wide SNP array analysis of ETAAD patients of European descent who were enrolled in the National Registry of Genetically Triggered Thoracic Aortic Aneurysms and Cardiovascular Conditions (GenTAC). Genotyping was performed on the Illumina Omni-Express platform, using PennCNV, Nexus and CNVPartition for CNV detection. ETAAD patients (n = 108, 100% European American, 28% female, average age 20 years, 55% with bicuspid aortic valves) were compared to 7013 dbGAP controls without a history of vascular disease using downsampled Omni 2.5 data. For comparison, 805 sporadic TAAD patients with late onset aortic disease (STAAD cohort) and 192 affected probands from families with at least two affected relatives (FTAAD cohort) from our institution were screened for additional CNVs at these loci with SNP arrays. We identified 47 recurrent CNV regions in the ETAAD, FTAAD and STAAD groups that were absent or extremely rare in controls. Nine rare CNVs that were either very large (>1 Mb) or shared by ETAAD and STAAD or FTAAD patients were also identified. Four rare CNVs involved genes that cause arterial aneurysms when mutated. The largest and most prevalent of the recurrent CNVs were at Xq28 (two duplications and two deletions) and 17q25.1 (three duplications). The percentage of individuals harboring rare CNVs was significantly greater in the ETAAD cohort (32%) than in the FTAAD (23%) or STAAD (17%) cohorts. We identified multiple loci affected by rare CNVs in one-third of ETAAD patients, confirming the genetic heterogeneity of TAAD. Alterations of candidate genes at these loci may contribute to the pathogenesis of TAAD.</p> <p> </p>
HLA and KIR allele genotyping for HPRC-frz2 haplotype assemblies
<p>HLA/KIR annotation for available long reads assemblies, constructed by the pipeline : https://github.com/YingZhou001/Immuannot</p> <p>IPD-KIR version: V2.13.0, IPD-IMGT/HLA version: V3.59.0</p> <p>472 haploid assemblies included</p>
Allele dataset of western grasswren for use in VORTEX (PVA analysis)
<p>Conservation translocations have become an increasingly popular method to restore or secure vulnerable populations. However, translocations greatly vary in success. The use of population viability analysis (PVA) may increase the likelihood of meeting translocation goals. However, the quality of PVAs to inform translocations is dependent on the availability of ecological data and clear translocation objectives to guide them. Here, we used PVAs to inform the planned conservation translocation of the Western Grasswren (<em>Amytornis textilis textilis</em>) from mainland Shark Bay onto Dirk Hartog Island, Western Australia. A range of translocation scenarios was modelled and scored against success criteria as determined by translocation objectives. Simulations of 20-year outcomes found that a minimum founder population of 112 individuals meets all success criteria. PVA supported sourcing individuals from two subpopulations to maximise genetic diversity. No impact to source populations was detected for the proposed harvest quantities despite conservative estimates of initial source population sizes. Here we demonstrate that creating clear, measurable objectives alongside a PVA lessens ambiguity about which translocation scenarios could be viable. In doing so, we have identified the minimum translocation sizes needed to maintain genetic diversity and population growth, thus increasing the likelihood of translocation success without impacting the source population.</p>
Diversity and host specificity of Borrelia burgdorferi's outer surface protein C (ospC) alleles in synanthropic mammals, with a notable ospC allele U absence from mixed infections
<p>Interactions among pathogen genotypes that vary in host specificity may affect overall transmission dynamics in multi-host systems. <em>Borrelia burgdorferi</em>, a bacterium that causes Lyme disease, is typically transmitted among wildlife by <em>Ixodes</em> ticks. Despite the existence of many alleles of <em>B. burgdorferi</em>'s <em>sensu stricto</em> outer surface protein C (<em>ospC</em>) gene, most human infections are caused by a small number of <em>ospC</em> alleles ["human infectious alleles" (HIAs)], suggesting variation in host specificity associated with <em>ospC</em>. To characterize the wildlife host association of <em>B. burgdorferi</em>'s <em>ospC</em> alleles, we used metagenomics to sequence <em>ospC</em> alleles from 68 infected individuals belonging to eight mammalian species trapped at three sites in suburban New Brunswick, New Jersey (USA). We found that multiple allele ("mixed") infections were common. HIAs were most common in mice (<em>Peromyscus</em> spp.) and only one HIA was detected at a site where mice were rarely captured. <em>OspC </em>allele U was exclusively found in chipmunks (<em>Tamias striatus</em>), and although a significant number of different alleles were observed in chipmunks, including HIAs, allele U never co-occurred with other alleles in mixed infections. Our results suggest that allele U may be excluding other alleles, thereby reducing the capacity of chipmunks to act as reservoirs for HIAs.</p>
Microsatellite allele length of Phytophthora ramorum in San Mateo County California
<p>We implement a population genetics approach to clarify the role that temporal and environmental variability, spatially distinct locations, and different hosts may have in the epidemiology of plant disease and the microevolution of its causative pathogen. In California and Southern Oregon (USA), the introduction of the invasive pathogen <em>Phytophthora ramorum</em>, causal agent of the widespread disease Sudden Oak Death (SOD), has resulted in extensive mortality of various oaks (<em>Quercus sp</em>.) and of tanoak (<em>Notholithocarpus densiflorus</em>). Although the disease can infect over a hundred hosts, California bay laurel (<em>Umbellularia californica</em>) is the most competent transmissive host but is not lethally affected by the disease. Using population genetics data, we identify the relationship among<em> P. ramorum </em>populations in bay laurels, oaks and tanoaks to clarify the contribution of each host on the epidemiology of SOD and the microevolution of its causal agent and to explore differences in population structure across sites and years. We conclude that bay laurel is the primary source for infections of both tanoak and oak, and that tanoak contributes minimally to oak infection but can infect bay laurel, creating a secondary pathogen amplification process. Overall, pathogen diversity is associated with rainfall and presence of bay laurels, which sustain the largest populations of the pathogen. Additionally, we clarify that while bay laurels are a common source of inoculum, oaks and tanoaks act as sinks that maintain host-specific pathogen genotypes not observed in bay laurel populations. Finally, we conclude that different sites support a dominance of different pathogen genotypes. Some genotypes were widespread, while others were limited to a subset of the plots. Sites with higher bay laurel densities sustained a higher genotypic diversity of the pathogen. This work provides novel insight into the ecology and evolutionary trajectories of SOD epidemics in natural ecosystems.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.