Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “paralogy”
Data from: Gene conversion yields novel gene combinations in paralogs of GOT1 in the copepod Tigriopus californicus
Background: Gene conversion of duplicated genes can slow the divergence of paralogous copies over time but can also result in other interesting evolutionary patterns. Islands of genetic divergence that persist in the face of gene conversion can point to gene regions undergoing selection for new functions. Novel combinations of genetic variation that differ greatly from the original sequence can result from the transfer of genetic variation between paralogous genes by rare gene conversion events. Genetically divergent populations of the copepod Tigriopus californicus provide an excellent model to look at the patterns of divergence among paralogs across multiple independent evolutionary lineages. Results: In this study the evolution of a set of paralogous genes encoding putative aspartate transaminase proteins (called GOT1 here) are examined in populations of the copepod T. californicus. One pair of duplicated genes, GOT1p1 and GOT1p2, has regions of high divergence between the copies in the face of apparent on-going gene conversion. The GOT1p2 gene also has unique haplotypes in two populations that appear to have resulted from a transfer of genetic variation via inter-paralog gene conversion. A second pair of duplicated genes GOT1Sr and GOT1Sd also shows evidence of gene conversion, but this gene conversion does not appear to have maintained each as a functional copy in all populations. Conclusions: The patterns of conservation and sequence divergence across this set of paralogous genes among populations of T. californicus suggest that some interesting evolutionary patterns are occurring at these loci. The results for the GOT1p1/GOT1p2 paralogs illustrate how gene conversion can factor in the creation of a mosaic pattern of regions of high divergence and low divergence. When coupled with rare gene conversion events of divergent regions, this pattern can result in the formation of novel proteins differing substantially from either original protein. The evolutionary patterns across these paralogs show how gene conversion can both constrain and facilitate diversification of genetic sequences.
Supplementary files from paper "Revisiting the evolution and function of NIP2 paralogs in the Rhynchosporium spp. complex"
<p>This dataset contains additional files related with the "Revisiting the evolution and function of NIP2 parologs in the <em>Rhynchosporium</em> spp. complex" paper. The second version contains four additional files: EarlGrey_annotation.bed, The_RIPper_annotation.gff3, Figure_S7_alignment.fasta, and WAI_strains_attributes.xlsx</p> <p>Folder 1: R script for differential gene expression analysis and RPKM calculation (EdgeR_WAI453.html) including the input files (gene_count_matrix.csv and for_RPKM_featureCounts.csv).</p> <p>Folder 2: Genome assembly (R.communeWAI453_20contigs.fasta), gene annotation of <em>R. commune</em> WAI453 (R.communeWAI453.gff3), <strong>transposable elements annotation of <em>R. commune</em> WAI453 (EarlGrey_annotation.bed)</strong>, and<strong> RIP-affected region annotation of <em>R. commune</em> WAI453 (The_RIPper_annotation.gff3)</strong>.</p> <p>Folder 3: Sequence of <em>NIP2</em> and <em>NLP</em> paralogs from <em>R. commune</em> global populations and <em>R. commune</em> sister species used in this study (nip2.1.all.fasta-nlp4.all.fasta) the summary of presence-absence polymorphism of all <em>NIP2</em> and <em>NLP</em> genes in <em>R. commune</em> and three <em>R. commune</em> sister species (gene_summary.csv), <strong>alignment of <em>NIP2 </em>and <em>NLP</em> genes (Figure_S7_alignment.fasta), </strong>and <strong>attributes of all <em>R. commune</em> WAI strains used in this study (WAI_strains_attributes.xlsx).</strong></p> <p>Folder 4: Pdb files of structural prediction of NIP2 proteins generated by AlphaFold 2 (nip2.1.pdb, nip2.3.pdb, and nip2.6.pdb).</p> <p>Folder 5: A python script to parse local blast result into fasta file (BLASTtoGFF_multiple.py).</p>
Haploid, diploid, and pooled exome capture recapitulate features of biology and paralogy in two non-model tree species
<p>Despite their suitability for studying evolution, many conifer species have large and repetitive giga-genomes (16-31Gbp) that create hurdles to producing high coverage SNP datasets that capture diversity from across the entirety of the genome. Due in part to multiple ancient whole genome duplication events, gene family expansion and subsequent evolution within <i>Pinaceae</i>, false diversity from the misalignment of paralog copies creates further challenges in accurately and reproducibly inferring evolutionary history from sequence data. Here, we leverage the cost-saving benefits of pool-seq and exome-capture to discover SNPs in two conifer species, Douglas-fir (<i>Pseudotsuga menziesii</i> var. <i>menziesii </i>(Mirb.) Franco, <i>Pinaceae</i>) and jack pine (<i>Pinus banksiana</i> Lamb., <i>Pinaceae</i>). We show, using minimal baseline filtering, that allele frequencies estimated from pooled individuals show a strong positive correlation with those estimated by sequencing the same population as individuals (r > 0.948), on par with such comparisons made in model organisms. Further, we highlight the utility of haploid megagametophyte tissue for identifying sites that are likely due to misaligned paralogs. Together with additional minor filtering, we show that it is possible to remove many of the loci with large frequency estimate discrepancies between individual and pooled sequencing approaches, improving the correlation further (r > 0.973). Our work addresses bioinformatic challenges in non-model organisms with large and complex genomes, highlights the use of megagametophyte tissue for the identification of paralog sites, and suggests the combination of pool-seq and exome capture to be robust for further evolutionary hypothesis testing in these systems.</p>
CYCLOIDEA paralogs function redundantly to specify dorsal flower development in Mimulus lewisii (Phrymaceae)
<p><strong>Premise</strong>: Duplicated genes (paralogs) are abundant in plant genomes and their retention may influence the function of genetic programs and contribute to evolutionary novelty. How gene duplication affects genetic modules, and the forces that contribute to paralog retention are outstanding questions. The CYCLOIDEA(CYC)-dependent flower symmetry program is a model for understanding the evolution of gene duplication, providing multiple examples of paralog partitioning and novelty. However, a novel CYC gene lineage duplication event near the origin of Higher Core Lamiales (HCL) has received little attention.</p> <p><strong>Methods</strong>: To understand the evolutionary fate of duplicated HCL CYC2 genes, we determined the effects on flower symmetry of suppressing MlCYC2A and MlCYC2B expression using RNA interference (RNAi). We determined flower symmetry phenotypic effects in single and double silenced backgrounds and coupled this with expression surveys of MlCYC2A, MlCYC2B, and a putative downstream RADIALIS (MlRAD5) ortholog.</p> <p><strong>Key</strong> <strong>results</strong>: MlCYC2A and MlCYC2B jointly contribute to bilateral flower symmetry. MlCYC2B exhibits a clear dorsal flower identity function and may additionally function in carpel development. MlCYC2A functions in establishing dorsal petal shape. Further, our results suggest an MlCYC2A–MlCYC2B regulatory interaction which may affect pathway homeostasis.</p> <p><strong>Conclusions</strong>: Our results suggest that Higher Core Lamiales-specific CYC paralogs may be selectively retained for their joint contribution to bilateral flower symmetry, similar to the independently-derived CYC paralogs in the Lamiales model for bilateral flower symmetry research, <em>Antirrhinum</em> <em>majus</em> (snapdragon).</p>
Paralogs and off-target sequences improve phylogenetic resolution in a densely-sampled study of the breadfruit genus (Artocarpus, Moraceae)
Open the record for dataset details and reuse information.
Data from: Differential requirements for the RAD51 paralogs in genome repair and maintenance in human cells
Open the record for dataset details and reuse information.
Data from: Gene conversion yields novel gene combinations in paralogs of GOT1 in the copepod Tigriopus californicus
Open the record for dataset details and reuse information.
Analysis of paralogs in target enrichment data pinpoints multiple ancient polyploidy events in Alchemilla s.l. (Rosaceae)
Open the record for dataset details and reuse information.
Data from: Mouse fitness measures reveal incomplete functional redundancy of Hox paralogous group 1 proteins
Open the record for dataset details and reuse information.
Data from: Extensive local gene duplication and functional divergence among paralogs in Atlantic salmon
Open the record for dataset details and reuse information.
Data from: Identifying pathogenicity of human variants via paralog-based yeast complementation
Open the record for dataset details and reuse information.
Haploid, diploid, and pooled exome capture recapitulate features of biology and paralogy in two non-model tree species
Open the record for dataset details and reuse information.
Data from: Congruent population structure across paralogous and non-paralogous loci in Salish Sea chum salmon (Oncorhynchus keta)
Open the record for dataset details and reuse information.
CYCLOIDEA paralogs function redundantly to specify dorsal flower development in Mimulus lewisii (Phrymaceae)
Open the record for dataset details and reuse information.
Systemic paralogy and function of retinal determination network homologs in arachnids
Open the record for dataset details and reuse information.
Evaluating the use of paralogous protein domains to increase data availability for missense variant classification [dataset]
Open the record for dataset details and reuse information.
The invasive land flatworm Arthurdendyus triangulatus: repeated sequences in the mitogenome, extra-long cox2 gene and paralogous rRNA clusters
<p>Fasta and tbl files for the mitogenomes of various Rhynchodeminae</p>
ProfileView on the paralogs of the Calvin-Benson-Bassham cycle
Open the record for dataset details and reuse information.
Data from: Paralogs are revealed by proportion of heterozygotes and deviations in read ratios in genotyping by sequencing data from natural populations
Whole genome duplications have occurred in the recent ancestors of many plants, fish, and amphibians, resulting in a pervasiveness of paralogous loci and the potential for both disomic and tetrasomic inheritance in the same genome. Paralogs can be difficult to reliably genotype and are often excluded from genotyping-by-sequencing (GBS) analyses; however, removal requires paralogs to be identified which is difficult without a reference genome. We present a method for identifying paralogs in natural populations by combining two properties of duplicated loci: 1) the expected frequency of heterozygotes exceeds that for singleton loci, and 2) within heterozygotes, observed read ratios for each allele in GBS data will deviate from the 1:1 expected for singleton (diploid) loci. These deviations are often not apparent within individuals, particularly when sequence coverage is low; but, we postulated that summing allele reads for each locus over all heterozygous individuals in a population would provide sufficient power to detect deviations at those loci. We identified paralogous loci in three species: Chinook salmon (Oncorhynchus tshawytscha) which retains regions with ongoing residual tetrasomy on eight chromosome arms following a recent whole genome duplication, mountain barberry (Berberis alpina) which has a large proportion of paralogs that arose through an unknown mechanism, and dusky parrotfish (Scarus niger) which has largely re-diploidized following an ancient whole genome duplication. Importantly, this approach only requires the genotype and allele-specific read counts for each individual, information which is readily obtained from most GBS analysis pipelines.
SimPhy configuration scripts for simulations reported in the study titled: Species tree inference methods intended to deal with incomplete lineage sorting are robust to the presence of paralogs
<p>Many recent phylogenetic methods have focused on accurately inferring species trees when there is gene tree discordance due to incomplete lineage sorting (ILS). For almost all of these methods, and for phylogenetic methods in general, the data for each locus is assumed to consist of orthologous, single-copy sequences. Loci that are present in more than a single copy in any of the studied genomes are excluded from the data. These steps greatly reduce the number of loci available for analysis. The question we seek to answer in this study is: What happens if one runs such species tree inference methods on data where paralogy is present, in addition to or without ILS being present? Through simulation studies and analyses of two large biological data sets, we show that running such methods on data with paralogs can still provide accurate results. We use multiple different methods, some of which are based directly on the multispecies coalescent (MSC) model, and some of which have been proven to be statistically consistent under it. We also treat the paralogous loci in multiple ways: from explicitly denoting them as paralogs, to randomly selecting one copy per species. In all cases the inferred species trees are as accurate as equivalent analyses using single-copy orthologs. Our results have significant implications for the use of ILS-aware phylogenomic analyses, demonstrating that they do not have to be restricted to single-copy loci. This will greatly increase the amount of data that can be used for phylogenetic inference.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.