Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
479
datasets available to search
ShareScore release 0.7.1
Dataset results
479 results for “genome evolution”
Genome evolution and introgression in the New Zealand mud snails Potamopyrgus estuarinus and Potamopyrgus kaitunuparaoa
<p>We have sequenced, assembled, and analyzed the nuclear and mitochondrial genomes and transcriptomes of <i>Potamopyrgus estuarinus</i> and <i>Potamopyrgus kaitunuparaoa</i>, two prosobranch snail species native to New Zealand that together span the continuum from estuary to freshwater.<i> </i>These two species are the closest known relatives of the freshwater species <i>P. antipodarum—</i>a model for studying the evolution of sex, host-parasite coevolution, and biological invasiveness—and thus provide key evolutionary context for understanding its unusual biology. The <i>P. estuarinus</i> and <i>P. kaitunuparaoa </i>genomes are very similar in size and overall gene content. Comparative analyses of genome content indicate that these two species harbor a near-identical set of genes involved in meiosis and sperm functions, including seven genes with meiosis-specific functions. These results are consistent with obligate sexual reproduction in these two species and provide a framework for future analyses of <i>P. antipodarum—</i>a species comprising both obligately sexual and obligately asexual lineages, each separately derived from a sexual ancestor. Genome-wide multigene phylogenetic analyses indicate that <i>P. kaitunuparaoa</i> is likely the closest relative to <i>P. antipodarum. </i>We nevertheless show that there has been considerable introgression between <i>P. estuarinus</i> and <i>P. kaitunuparaoa.</i> That introgression does not extend to the mitochondrial genome, which appears to serve as a barrier to hybridization between <i>P. estuarinus </i>and <i>P. kaitunuparaoa.</i> Nuclear-encoded genes whose products function in joint mitochondrial-nuclear enzyme complexes exhibit similar patterns of non-introgression, indicating that incompatibilities between the mitochondrial and the nuclear genome may have prevented more extensive gene flow between these two species.<i> </i> </p>
The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma
<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup> (IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>
Genomics of extreme ecological specialists: multiple convergent evolution but no genetic divergence between ecotypes of Maculinea alcon butterflies
<p>Biotic interactions are often acknowledged as catalysers of genetic divergence and eventual explanation of processes driving species richness. We address the question, whether extreme ecological specialization is always associated with lineage sorting, by analysing polymorphisms in morphologically similar ecotypes of the myrmecophilous butterfly <em>Maculinea alcon</em>. The ecotypes occur in either hygric or xeric habitats, use different larval host plants and ant species, but no significant distinctive molecular traits have been revealed so far. We apply genome-wide RAD-sequencing to specimens originating from both habitats across Europe in order to get a view of the potential evolutionary processes at work. Our results confirm that genetic variation is mainly structured geographically but not ecologically — specimens from close localities are more related to each other than populations of each ecotype from distant localities. However, we found two loci for which the association with xeric versus hygric habitats is supported by segregating alleles, suggesting convergent evolution of habitat preference. Thus, ecological divergence between the forms probably does not represent an early stage of speciation, but may result from independent recurring adaptations involving few genes. We discuss the implications of these results for conservation and suggest preserving biotic interactions and main genetic clusters.</p>
A chromosome-level genome resource for studying virulence mechanisms and evolution of the coffee rust pathogen Hemileia vastatrix
<p>Recurrent epidemics of coffee leaf rust, caused by the fungal pathogen <em>Hemileia vastatrix,</em> have constrained the sustainable production of Arabica coffee for over 150 years. The ability of <em>H. vastatrix </em>to overcome resistance in coffee cultivars and evolve new races is inexplicable for a pathogen that supposedly only utilizes clonal reproduction. Understanding the evolutionary complexity between <em>H. vastatrix</em> and its only known host, including determining how the pathogen evolves virulence so rapidly is crucial for disease management. Achieving such goals relies on the availability of a comprehensive and high-quality genome reference assembly. To date, two reference genomes have been assembled and published for <em>H. vastatrix</em> that, while useful, remain fragmented and do not represent chromosomal scaffolds. Here, we present a complete scaffolded pseudochromosome-level genome resource for <em>H. vastatrix </em>strain 178a (Hv178a). Our initial assembly revealed an unusually high degree of gene duplication (over 50% BUSCO basidiomycota_odb10 genes). Upon inspection, this was predominantly due to a single scaffold that itself showed 91.9% BUSCO Completeness. Taxonomic analysis of predicted BUSCO genes placed this scaffold in Exobasidiomycetes and suggests it is a distinct genome, which we have named Hv178a associated fungal genome (Hv178a AFG). The high depth of coverage and close association with Hv178a raises the prospect of symbiosis, although we cannot completely rule out contamination at this time. The main Ca. 546 Mbp Hv178a genome was primarily (97.7%) localised to 11 pseudochromosomes (51.5 Mb N50), building the foundation for future advanced studies of genome structure and organization. Citation: https://doi.org/10.1101/2022.07.29.502101</p>
Genomic incongruence accompanies the evolution of flower symmetry in Eudicots: a case study in the poppy family (Papaveraceae, Ranunculales)
<p>Nuclear and plastid datasets and phylogenomic workflow associated to "Genomic Incongruence Accompanies the Evolution of Flower Symmetry in Eudicots: a case study in the poppy family (Papaveraceae, Ranunculales)", published in <em>Frontiers in Plant Science </em>15:1340056.<br>This compressed file (poppy_repo.zip) contains a markdown readme file (poppy_readme.md) describing the phylogenomic workflow followed, as well as two dataset folders (poppy_nuc and poppy_pl) divided into four (aln_nuc, gtr_nuc, sptr_nuc, and chrono_nuc) and three (aln_pl, sptr_pl, and chrono_pl) subfolders, respectively.<br>The nuclear folder (poppy_nuc) comprises shrunk and trimmed alignments (aln_nuc), ML gene trees (gtr_nuc), coalescent species trees (sptr_nuc), and a time tree (chrono_nuc).<br>The plastid folder (poppy_pl) comprises shrunk and trimmed alignments (aln_pl), a concatenated ML species tree (sptr_pl), and a time tree (chrono_pl).<br>The research article is available at https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2024.1340056 (doi: 10.3389/fpls.2024.1340056).</p>
Supplementary data and scripts for Willemsen and Bravo 2019 "Origin and evolution of papillomavirus (onco)genes and genomes"
<p>Supplementary data for Willemsen and Bravo 2019 "Origin and evolution of papillomavirus (onco)genes and genomes". The data set consists of two folders: “Bali-Phy” and “RandomPermutationTests”. The “Bali-Phy” folder contains the final results and convergence diagnostics of the Common Ancestry tests obtained by using the Bali-Phy software. The “RandomPermutationTests” folder contains all the data and scripts to repeat the random permutation tests described in the manuscript. Please see the corresponding README files for more information.</p>
Alignments from "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates"
<p>Compressed file containing the alignments at both nucleotide and amino acid level for the manuscript "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates" </p>
Analysis of the P. lividus sea urchin genome highlights contrasting trends of genomic and regulatory evolution in deuterostomes
<p><br> Supplementary datasets accompanying paper: </p> <p>stage_peaks_anc_sel.xlsx : ATAC peaks with classification, conservation and binding sites<br> Pliv.mfuzz.enrichGO.txt : GO enrichment in MFuzz cluster<br> Pliv_genes_master_filt.xlsx : Gene models with corresponding information<br> bindetect_results_anf.txt : results of TOBIAS<br> hits_pprx_cl0_ord3vrr+Et_red.fa : alignment of homeobox sequences<br> Pliv_aH2p.gn.gtf.gz : annotation in GTF format<br> Pliv_PqN3S_sm.fa.gz : genome of P. livius <br> ansr_*_network.tsv.gz : Stage specific networks from ANANSE analysis<br> ATAC_pks_normcov.tsv : Coverage of unified peaks for ATAC-seq<br> Cttg_pks_normcov.tsv : Coverage of unified peaks for Cut-and-tag H3K27Ac data<br> lncRNA_stgSpe_fpkm.tsv : Expression levels (FPKM) for predicted lncRNAs for available RNA-seq samples <br> Split_Urchin_FPKMs.clean.txt.gz : Expression levels for unified ATAc-seq peaks following direct and reverse orientation</p> <p> </p> <p> </p> <p> </p> <p> </p>
Data from: The genomics and evolution of inter-sexual mimicry and female-limited polymorphisms in damselflies
<p>The dataset contains intermediate output files required to reproduce the figures in the main text and Supporting Material of Willink <em>et al</em>. 2023. The genomics and evolution of inter-sexual mimicry and female-limited polymorphisms in damselflies.</p> <p>FILE OVERVIEW:</p> <p>1. Morph-specific assemblies<br> A. File names: Afem_1354_ragtag.fasta.gz, Ifem_1049_ragtag.fa.gz, Ofem_0081_ragtag.fa.gz, O054_Shasta_run2.PMDV.HAP1.purged.fasta.gz, A059_Shasta_run1.PMDV.HAP1.purged.fa.gz<br> B. Description: genome assemblies for different morphs of <em>Ischnura elegans</em> (Afem_1354, Ifem_1049, and Ofem_0081) and <em>Ischnura senegalensis</em> (A059 and O054), generated in this study from long-read Nanopore data using Shasta v 0.7.0 (https://github.com/paoloshasta/shasta).</p> <p>2. Assembly statistics<br> A. File names: Assembly_statistics.csv, Assembly_statistics_sen.csv<br> B. Description: Completeness and quality metrics for <em>de novo</em> genome assemblies of <em>I. elegans</em> and <em>I. senegalensis</em> female morphs. See Fig. S1-S2.</p> <p>3. Repetitive content annotation<br> A. File names: A1354_ragtag_RED.bed.repeats.bed.gz, Afem_Shasta1_polished_ragtag_UPPER.fa.out.gz, Ifem_Shasta2_polished_ragtag_UPPER.fa.out.gz, ioIscEleg1.1.primary_UPPER.fa.out.gz, ToL_RED.repeats.bed.gz<br> B. Description: Annotation of repetitive sequences in morph-specific assemblies. All morph assemblies (A, I and Darwin Tree of Life assemblies) were annotated using RepeatModeler v 2.0.1 and RepeatMasker v 1.0.93 (http://www.repeatmasker.org). The A morph and DToL assemblies were additionally annotated using Red v 0.0.1 (https://github.com/BioinformaticsToolsmith/Red). RepeatMasker annotations were then used to estimate TE coverage. See Extended Data Fig. 4 and Fig. S7.</p> <p>4. GWAS output<br> A. File names: A1354_ragtag_AvI.assoc_filtered.txt.gz, A1354_ragtag_AvO.assoc_filtered.txt.gz, A1354_ragtag_IvO.assoc_filtered.txt.gz, ToL_AvI.assoc_filtered.txt.gz, ToL_AvO.assoc_filtered.txt.gz, ToL_IvO.assoc_filtered.txt.gz<br> B. Description: filtered SNPs in pairwise association tests between morphs (n = 19 resequencing samples per morph) of<em> I. elegans</em>. Analyses were conducted in PLINK v 1.9 (http://pngu.mgh.harvard.edu/purcell/plink/), using either the A morph assembly (Fig. 2a-b), or the Darwin Tree of Life (DToL) reference assembly (Extended Data Figure 8a-b) as mapping reference.</p> <p>5. Population statistics<br> A. File names: Afem_pixy_30K_fst.txt.gz, A1354_30kb.Tajima.D.gz, Afem_pi_30K_pi.txt.gz, ToL_30K_fst.txt.gz, ToL_30kb.Tajima.D.gz, ToL_30K_onepop_pi.txt.gz<br> B. Description: Genetic differentiation (fst) between morphs, Tajima's D statistics, and nucleotide diversity across 30 kb windows of the<em> I. elegans</em> genome. Population statistics were computed using either the A morph assembly (Fig. 2c-e), or the DToL reference assembly (Extended Data Figure 8c-e) as mapping reference.</p> <p>6. k-mer based GWAS<br> A. File names: AvI_kmers.fa.gz, AvO_kmers.fa.gz, OvAI_kmers.fa.gz, AvI_kmers.fa_v_A1354_Shasta_run1_table.tsv.gz, AvO_kmers.fa_v_A1354_Shasta_run1_table.tsv.gz, OvAI_kmers.fa_v_A1354_Shasta_run1_table.tsv.gz, OvAI_kmers.fa_v_Ifem_1049_ragtag_table.tsv.gz<br> B. Description: List of significant k-mers (in fasta format) in three k-mer based association analyses (n = 19 resequencing samples per morph) between morphs of<em> I. elegans</em>. Significant k-mers were then mapped to morph-specific assemblies using Blast v 2.22.28 (https://blast.ncbi.nlm.nih.gov/Blast.cgi) for short sequences. We include mapping results shown in Fig. 3a-b.</p> <p>7. Read-depth coverage<br> A. File names: reseq_coverage_norepeat_500_window.bed.gz, nano_coverage_norepeat_500_window.bed.gz, Ifem_nano_coverage_norepeat_500_window.bed.gz, Ifem_reseq_coverage_norepeat_500_window_15Mb.bed.gz, poolseq_coverage_norepeat_500_window.bed.gz, morph_coverage_norepeat_diff_500.tsv.gz, SwD_popmap<br> B. Description: Read depth coverage of the morph locus and a 15 mb region used to estimate baseline read depths. 19 Illumina resequencing samples, and one long-read Nanopore sample of each morph of <em>I. elegans</em> were mapped to both the A and I assemblies to estimate read depth. Two poolseq samples (each pool consisting of 30 females of each morph) of<em> I. senegalensis</em> were mapped to the A assembly of<em> I. elegans</em> to estimate read depth. Read depth was estimated in mosdepth v 0.2.8 (https://github.com/brentp/mosdepth) across 500 bp windows after filtering windows with more than 10% repetitive content. For poolseq samples, the difference in coverage values between the A and O pools was computed across the entire genome. Sample information for resequencing samples is recorded in the file SwD_popmap. See Fig. 3c-d, 5b, and S8.</p> <p>8. Assembly alignment<br> A. File names: nucmer_aln_Ifem_1049_ragtag_Afem_1354_ragtag.qr1_filter.reformat.coords.gz, nucmer_aln_Ofem_0081_ragtag_Afem_1354_ragtag.qr1_filter.reformat.coords.gz, nucmer_aln_Afem_Isen_Afem_Iele.qr1_filter.reformat.coords.gz, nucmer_aln_Ofem_Isen_Afem_Iele.qr1_filter.reformat.coords.gz, karyotype_AI_RagTag.csv, karyotype_AO_RagTag.csv, karyotype_AIsen_AIele.cs, karyotype_OIsen_AIele.csv<br> B. Description: Assembly alignments using nucmer v 4.0.0 (https://github.com/mummer4/mummer) and contig synteny for plotting using RIdeogram v 0.2.2 (https://cran.r-project.org/web/packages/RIdeogram/vignettes/RIdeogram.html) in R v 4.2.2 (https://www.r-project.org/). The A morph assembly of <em>I. elegans</em> was aligned to the I and O morph assemblies of<em> I. elegans</em> and to the A and O-like assemblies of <em>I. senegalensis</em>. See Fig. 4a, 5c.</p> <p>9. Genotyping the Darwin Tree of Life assembly<br> A. File names: nucmer_aln_Afem_ragtag_ToL-haplotigs.qr1_filter.reformat.coords.gz, nucmer_aln_Afem_ragtag_ToL-primary.qr1_filter.reformat.coords.gz, ToL_500_norepeat.regions.bed.gz, karyotype_AToL_13_unloc_RagTag.csv, karyotype_AToL_RagTag_haplotigs.csv<br> B. Description: To genotype the DToL reference assembly of<em> I. elegans</em>, we estimated read-depth coverage of the DToL long-read Pacbio data mapped to the A morph assembly of <em>I. elegans</em> generated in this study, and aligned the A morph assembly to both the primary DToL assembly and to the purged haplotigs. Read depth was estimated in mosdepth v 0.2.8 (https://github.com/brentp/mosdepth) and assembly alignments were conducted using nucmer v 4.0.0 (https://github.com/mummer4/mummer). See Fig. S3.</p> <p>10. SV calling<br> A. File names: A_to_A.bam, A_to_A.bam.bai, A_to_I.bam, A_to_I.bam.bai, A_to_O.bam, A_to_O.bam.bai, A_to_ToL_2mb.bam, A_to_ToL_2mb.bam.bai, I_to_A.bam, I_to_A.bam.bai, I_to_I.bam, I_to_I.bam.bai, I_to_O.bam, I_to_O.bam.bai, I_to_ToL_2mb.bam, I_to_ToL_2mb.bam.bai, O_to_A.bam, O_to_A.bam.bai, O_to_I.bam, O_to_I.bam.bai, O_to_O.bam, O_to_O.bam.bai, O_to_ToL_2mb.bam, O_to_ToL_2mb.bam.bai<br> B. Description: mergede alignements of resequencing samples (n = 19 per morph) to alternative reference assemblies (A, I, O, and DToL) for<em> I. elegans</em>. The alignments have been filtered by quality and to contain only the unlocalized scaffold 2 of chromosome 13, which includes the morph locus. These files were used to call morph-specific structural variants using samplot v 1.3.0 (https://github.com/ryanlayer/samplot). See Extended Data Figs 2, 7, and Fig. S5-S6.</p> <p>11. Mapping of inversion breakpoint reads<br> A. File names: AvO_3K.tsv.gz, AvO_22K.tsv.gz, AvO_sen_3K.tsv.gz, AvO_sen_22K.tsv.gz, IvO_3K.tsv.gz<br> B. Description: Signatures of an inversion with breakpoints at ~ 3 kb and ~ 22 kb of the unlocalized scaffold 2 of chromosome 13 on the O assembly were found in A and I resequencing samples of <em>I. elegans</em> and in poolseq samples of A females of <em>I. senegalensis</em>. We queried the reads mapping to the inversion breakpoints and then tabulated their mapping locations of the A morph assembly of<em> I. elegans</em> (Fig. 6 and Extended Data Fig. 3, 7b-c). For the first inversion breakpoint, we also mapped reads on the I morph assembly of <em>I.</em> elegans (Fig. S12).</p> <p>12. Evidence of translocation in I<br> A. File names: Ifem_nano_SUPER_13_unloc_2.bam, Ifem_nano_SUPER_13_unloc_2.bam.bai<br> B. Description: Long-read Nanopore data of a I morph female of <em>I. elegans</em> mapped to the A morph of <em>I. elegans</em> and filtered to contain the entire unlocalized scaffold 2 of chromosome 13. Read mapping was conducted in minimap2 v 2.22-r1110 (https://github.com/lh3/minimap2) and used to identify a translocation signature in the I morph, relative to the A morph of <em>I. elegans</em>. See Extended Data Fig. 6.</p> <p>13. PCA output<br> A. File names: A1354_all.eigenval, A1354_all.eigenvec, I1049_all.eigenval, I1049_all.eigenvec<br> B. Description: Eigenvectors and eigenvalues of PCA analyses of population structure between morphs of <em>I. elegans</em>. PCA analysis were conducted on morph locus, using either the A morph or the I morph assembly as mapping reference in PLINK v 1.9 (http://pngu.mgh.harvard.edu/purcell/plink/). See Fig. S4.</p> <p>14. Linkage disequilibrium<br> A. File names: A1354_SUPER_1_allr.ld.gz, A1354_SUPER_2_allr.ld.gz, A1354_SUPER_3_allr.ld.gz, A1354_SUPER_4_allr.ld.gz, A1354_SUPER_5_allr.ld.gz, A1354_SUPER_6_allr.ld.gz, A1354_SUPER_7_allr.ld.gz, A1354_SUPER_8_allr.ld.gz, A1354_SUPER_9_allr.ld.gz, A1354_SUPER_10_allr.ld.gz, A1354_SUPER_11_allr.ld.gz, A1354_SUPER_12_allr.ld.gz, A1354_SUPER_13_allr.ld.gz, A1354_SUPER_13_unloc_1_allr.ld.gz, A1354_SUPER_13_unloc_2_allr.ld.gz, A1354_SUPER_13_unloc_3_allr.ld.gz, A1354_SUPER_13_unloc_4_allr.ld.gz, A1354_SUPER_X_allr.ld.gz<br> B. Description: Estimates of recombination rate (R2) between SNPs across the first 15 mb of each chromosome and unlocalized segments of chromosome 13 of <em>I. elegans</em>. Recombination rates were estimated based on 57 resequencing samples and using the A morph assembly as mapping reference in PLINK v 1.9 (http://pngu.mgh.harvard.edu/purcell/plink/). See Extended Data Fig. 5.</p> <p>15. Gene annotations<br> A. File names: Afem_all_ragtag.gtf.gz, Afem_all_transcripts.transdecoder.genome.gff3.gz, Isen.gtf.gz<br> B. Description: Annotation of the A morph assembly of <em>I. elegans</em> using RNAseq data to assemble transcripts <em>de novo </em>for <em>I. elengans</em> and <em>I. senegalensis</em> in Stringtie v 2.1.4 (https://ccb.jhu.edu/software/stringtie/). Peptide sequences for the <em>I. elegans</em> transcripts were then predicted using Transdecoder v 5.5.0 (https://github.com/TransDecoder/TransDecoder).</p> <p>16. Gene annotations in the morph locus<br> A. File names: gene_models_shared_trancripts_simple.csv, gene_models_shared_trancripts_simple_I.csv<br> B. Description: locations of exon features for genes in the morphs locus and expressed in at least one adult sample of both<em> I. elegans</em> and <em>I. senegalensis</em>. Locations are given for the A and I assemblies. See Fig. 6 and S12.</p> <p>17. Gene expression<br> A. File names: DToL_gene_count_matrix.csv.gz, DToL_transcript_count_matrix.csv.gz, gene_count_matrix.csv.gz, transcript_count_matrix.csv.gz, Isen_gene_count_matrix.csv.gz, Isen_transcript_count_matrix.csv.gz, Iele_phenodata.csv, Isen_phenodata.csv<br> B. Description: Sample information (phenodata), gene and transcript count matrices for gene expression analysis. For <em>I. elegans</em>, gene expression was quantified on thoracic tissue of six adult females of each morph and six adult males (three sexually mature and three sexually immature in each group). Reads were mapped to both the A morph assembly and the DToL reference assembly. For <em>I. senegalensis</em>, we used previously published data (NCBI BioProject PRJDB11387) from different tissues of adult females of each morph and males (one upon emergence and one two days after emergence for each group) mapped to the A morph assembly. Gene and transcript counts were generated using Stringtie v 2.1.4 (https://ccb.jhu.edu/software/stringtie/). See Fig. 6, S9-S11, S13.</p> <p>18. SNPs in the morph locus<br> A. File names: A1354-ragtag-allsites-candidate_gene_cds.vcf.gz, A1354-ragtag-allsites-candidate_gene_cds.vcf.gz.tbi, vcf_popmap<br> B. Description: SNPs in 57 resequencing samples across coding sequences of the morph locus of <em>I. elegans</em>. The A morph assembly was used as mapping reference. Sample information for resequencing samples is recorded in the file vcf_popmap. See Fig. S14a.</p> <p>19. Domains and orthologues of Gastrula zinc-finger transcription factor in the morph locus<br> A. File names: GZnf_domain_annot.csv, GZnF_orthologue.tre, GZnf_orthologue_annot.txt<br> B. Description: Functional domains were annotated using InterProScan (https://www.ebi.ac.uk/interpro/). The gene orthologue tree was inferred using OrthoFinder v 2.5.2 (https://github.com/davidemms/OrthoFinder). See Fig. S14.</p>
Data from: Repetitive DNA profiles reveal evidence of rapid genome evolution and reflect species boundaries in ground beetles
Genome architecture is a complex, multidimensional property of an organism defined by the content and spatial organization of the genome's component parts. Comparative study of entire genome architecture in model organisms is shedding light on mechanisms underlying genome regulation, evolution, and diversification; but such studies require costly analytical approaches which make extensive comparative study impractical for most groups. However, lower-cost methods that measure a single architectural component (e.g., distribution of one class of repeats) have potential as a new data source for evolutionary studies insofar as that measure correlates with more complex biological phenomena, and for which it could serve as part of an explanatory framework. We investigated copy number variation (CNV) profiles in ribosomal DNA (rDNA) as a simple measure reflecting the distribution of rDNA subcomponents across the genome. We find that signatures present in rDNA CNV profiles strongly correlate with species boundaries in the <i>breve</i> species group of <i>Bembidion</i>, and vary across broader taxonomic sampling in <i>Bembidion</i> subgenus <i>Plataphus</i>. Profiles of several species show evidence of re-patterning of rDNA-like sequences throughout the genome, revealing evidence of rapid genome evolution (including among sister pairs) not evident from analysis of traditional data sources such as multi-gene data sets. Major re-patterning of rDNA-like sequences has occurred frequently within the evolutionary history of <i>Plataphus</i>. We confirm that CNV profiles represent an aspect of genomic architecture (i.e., the linear distribution of rDNA components across the genome) via fluorescence in-situ hybridization. In at least one species, novel rDNA-like elements are spread throughout all chromosomes. We discuss the potential of copy number profiles of rDNA, or other repeats, as a low-cost tool for incorporating signal of genomic architecture variation in studies of species delimitation and genome evolution.
The hagfish genome and the evolution of vertebrates
<p><strong>Genomic datasets for the hagfish </strong><i><strong>Eptatretus atami</strong></i></p><p><strong>Pata_MYAQb.rep.bed.gz:</strong> bed file with position of repeats masked by repeat masker</p><p><strong>Pata_ah2p.pep.fa</strong>: proteins from genome annotation </p><p><strong>Pata_MYAQb_rnm_ah2p.gtf</strong>: genome annotation performed using mikado, augustus and refined with Pasa (see methods). </p><p><strong>Pata_MYAQb_rnm.fa.gz: </strong>genome assembly in fasta format. </p><p> </p><p> </p>
Implications of the three-dimensional chromatin organization for genome evolution in a fungal plant pathogen
<p><span>The spatial organization of eukaryotic genomes is linked to their biological functions, although it is not clear how this impacts the overall evolution of a genome. Here, we uncover the three-dimensional (3D) genome organization of the phytopathogen <em>Verticillium dahliae</em>,<em> </em>known to possess distinct genomic regions, designated adaptive genomic regions (AGRs), enriched in transposable elements and genes that mediate host infection. Short-range DNA interactions form clear topologically associating domains (TADs) with gene-rich boundaries that show reduced levels of gene expression and reduced genomic variation. Intriguingly, TADs are less clearly insulated in AGRs than in the core genome. At a global scale, the genome contains bipartite long-range interactions, particularly enriched for AGRs and more generally containing segmental duplications. Notably, the patterns observed for <em>V. dahliae </em>are also present in other <em>Verticillium</em> species. Thus, our analysis links 3D genome organization to evolutionary features conserved throughout the <em>Verticillium</em> genus.</span></p>
apomixis_parallel_evolution, and Fortunella hindsii (Citrus hindsii )Genome sequencing and assembly
<p>##The assemble(genome) files</p> <p>Citrus hindsii (Mini citrus)Genome assembly</p> <p>Citrus hindsii (Mini citrus)Genome assembly gene model gff3 file</p> <p>Citrus hindsii (Mini citrus)Genome assembly function annotation</p> <p>Citrus hindsii (Mini citrus)Genome assembly TE gff3 file</p> <p>##The population dataset</p> <p>LD_pur.vcf.gz //The LD purning SNP vcfs (1.4 M sites) used in analysis</p> <p><br> log10_auxin.txt //The expression (log10) related to auxin pathway</p> <p><br> sjg.temergedref.fasta.gz //The TE insertion modify genome in popTE2 analysis</p> <p><br> SVs.vcf.gz //The SV vcfs used in the paper</p> <p><br> te-hierarchy.txt //The TE classfication in popTE2 analysis</p> <p><br> TPM_count.txt //All samples expression in TMP count</p> <p><br> unfiltered.vcf.gz //The unfiltered vcfs file (7.3 M sites, within 0.4 M indels) </p>
Pollinator loss causes rapid adaptive evolution of selfing and dramatically reduces genome-wide genetic variability
<p>While selfing populations harbor little genetic variation limiting evolutionary potential, the causes are unclear. We experimentally evolved large, replicate populations of <em>Mimulus guttatus </em>for nine generations in greenhouses with or without pollinating bees and studied DNA polymorphism in descendants. Populations without bees adapted to produce more selfed seed yet exhibited striking reductions in DNA polymorphism despite large population sizes. Importantly, the genome-wide pattern of variation cannot be explained by a simple reduction in effective population size, but instead reflects the complicated interaction between selection, linkage, and inbreeding. Simulations demonstrate that the spread of favored alleles at few loci depresses neutral variation genome-wide in large populations containing fully selfing lineages. It also generates greater heterogeneity among chromosomes than expected with neutral evolution in small populations. Genome-wide deviations from neutrality were documented in populations with bees, suggesting widespread influences of background selection. After applying outlier tests to detect loci under selection, two genome regions were found in populations with bees, yet no adaptive loci were otherwise mapped. Large amounts of stochastic change in selfing populations compromise evolutionary potential and undermine outlier tests for selection. This occurs because genetic draft in highly selfing populations makes even the largest changes in allele frequency unremarkable.</p>
Figure 4 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains
Figure 4. Maximum likelihood-based phylogeny, with 100 bootstrap replicates and using all the 26 mitochondrial genomes generated in this study. The dataset was supplemented with Thyropygus sp. and Abacion magnum as outgroups, with sequences derived from GenBank. GenBank accession numbers are given in parentheses. Colours represent Tropostreptus sample origins. The upper right inset shows the topology of the Tropostreptus hamatus lineage, enlarged to clarify the branching order. Only support values <100 are shown. *Thyropygus sp. (red font) is very likely to be a species misidentification; for more information, see Discussion text.
Figure 5. Bayesian phylogeny, with species divergence age estimates reconstructed with BEAST using all the 26 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains
Figure 5. Bayesian phylogeny, with species divergence age estimates reconstructed with BEAST using all the 26 mitochondrial genomes generated in this study. The dataset was supplemented with Thyropygus sp. and Abacion magnum as outgroups, derived from GenBank. GenBank accession numbers are provided in parentheses. Blue bars indicate the 95% highest probability density intervals for node ages. Age estimation for lineage divergence was based on a general arthropod mitochondrial DNA substitution rate and should be considered with caution. *Thyropygus sp. (red font) is very likely to be a misidentification; for more information, see the Discussion.
Figure 1 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains
Figure 1. Typical Tropostreptus appearance exemplified by a Tropostreptus hamatus individual from Udzungwa Mountains, Tanzania (photograph credit: Nikolaj Scharff).
Figure 2 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains
Figure 2. Map showing the origin of the millipede specimens used in the study, with the accuracy of location restricted to mountain blocks. Coloured circles all represent Tropostreptus species, whereas grey symbols represent species from other millipede genera. Base map published by permission of the Eastern Arc Mountains Conservation Endowment Fund.
Figure 3 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains
Figure 3. The gene order of mitochondrial coding sequences shared among all analysed millipede species in this study, which include all known species of Tropostreptus (T. droides, T. hamatus, T. kipunji, T. microcephalus, T. severus and T. sigmatospinus), in addition to Archispirostreptus gigas, Chaleponcus netus, Macrolenostreptus orestes, Prionopetalum kraepelini and Pseudotibiozus cerasopus. Colour key: red, ribosomal RNA (rRNA); pink, transfer RNA (tRNA); yellow, protein-coding sequences (CDS). Arrows indicate gene transcription orientation.
Evolution of the correlated genomic variation landscape across a divergence continuum in the genus Castanopsis
<p>The heterogeneous landscape of genomic variation has been well documented in population genomic studies. However, disentangling the intricate interplay of evolutionary forces influencing the genetic variation landscape over time remains challenging. In this study, we assembled a chromosome-level genome for <em>Castanopsis eyrei</em> and sequenced the whole genomes of 276 individuals from 12 <em>Castanopsis</em> species, spanning a broad divergence continuum. We found highly correlated genomic variation landscapes across these species. Furthermore, variations in genetic diversity and differentiation along the genome were strongly associated with recombination rates and gene density. These results suggest that long-term linked selection and conserved genomic features have contributed to the formation of a common genomic variation landscape. By examining how correlations between population summary statistics change throughout the species divergence continuum, we determined that background selection alone does not fully explain the observed patterns of genomic variation; the effects of recurrent selective sweeps must be considered. We further revealed that extensive gene flow has significantly influenced patterns of genomic variation in <em>Castanopsis</em> species. The estimated admixture proportion correlated positively with recombination rate and negatively with gene density, supporting a scenario of selection against gene flow. Additionally, putative introgression regions exhibited strong signals of positive selection, an enrichment of functional genes, and reduced genetic burdens, indicating that adaptive introgression has played a role in shaping the genomes of hybridizing species. This study provides insights into how different evolutionary forces have interacted in driving the evolution of the genomic variation landscape.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.