Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
380
datasets available to search
ShareScore release 0.7.1
Dataset results
380 results for “Transposable elements”
A pangenome-guided manually curated library of transposable elements for Zymoseptoria tritici
<p>A manually-curated TE consensus library generated using a panel of 19 reference genomes for <em>Zymoseptoria tritici</em><sup>1-3</sup> along with reference genome assemblies for the sister species <em>Z. ardabiliae</em>, <em>Z. brevis</em>, <em>Z. pseudotritici</em>, and <em>Z. passerinii<sup>4</sup></em>. </p> <p> </p> <p><strong>Methods</strong></p> <p>Putative TE consensus sequences were first obtained by annotating all 23 genome assemblies<sup>1–4</sup> with Earl Grey with default settings (v3.0; <a href="https://github.com/TobyBaril/EarlGrey">https://github.com/TobyBaril/EarlGrey</a>)<sup>5,6</sup>. Consensus sequences generated from each reference genome were clustered using CD-Hit-Est (v4.8.1)<sup>7,8</sup> to group sequences with 90% similarity across 80% of the longer sequence length (<em>-n 8 -d 0 -aL 0.8 -c 0.90 -G 0 -g 1 -b 500 -r 1</em>) to reduce redundancy whilst preventing the collapsing of chimeric sequences. Consensus sequences <100bp were removed, as these are unlikely to represent true TE sequences. Each consensus sequence was then subject to manual curation as described by Goubert et al. (2022)<sup>9</sup>. Briefly, genomic copies of each TE were obtained using a “BLAST, Extract, Extend” process to recover genomic copies from each of the 23 reference genome assemblies with 1,000 flanking bases at either end<sup>9,10</sup>. For families with >100 BLASTN hits, the 25 longest hits were selected, along with 75 random hits. Multiple alignments were generated for each putative TE family using MAFFT (v7.505) with the --auto flag<sup>11</sup>. Columns composed of >=80% gaps were removed with T-COFFEE (v13.45.0.4846264)<sup>12</sup>. Subsequently, all sequence alignments were manually curated to define TE boundaries and remove regions of low conservation and rare insertions. Following manual curation, new majority-rule consensus sequences were generated with EMBOSS (v6.6.0.0) cons<sup>13</sup>. TE-Aid (<a href="https://github.com/clemgoub/TE-Aid/">https://github.com/clemgoub/TE-Aid/</a>) was used to aid visual inspection and to identify diagnostic features for classification of extended consensus sequences. Following this, TIRs were recorded if present, and nhmmscan (HMMER v3.3.2)<sup>14</sup> was used to identify homology to known curated elements in Dfam (v3.7). Combining this information, each TE consensus sequence was manually classified using available information following the naming convention ‘>ZymTri_2023_family_[n]#[Classification]/[Family]’. Consensus sequences classified with low confidence have a ‘?’ added to the name, as well as the string ‘_LowConf’. To reduce redundancy in the final TE library, sequences were clustered to the family-level using the 80-80-80 rule implemented in CD-hit-est<sup>9,15 </sup>(<em>-d 0 -aS 0.8 -c 0.8 -G 0 -g 1 -b 500 -r 1</em>). The representative sequence for each cluster was manually selected to select the sequence with the highest classification confidence, also defined as the ‘most intact consensus’. Chimeric sequences erroneously clustered were manually separated to retain sequences for the chimeric TE and the individual elements that generated the chimer.</p> <p> </p> <p><strong>References</strong></p> <p>1. Badet, T., Oggenfuss, U., Abraham, L., McDonald, B. A. & Croll, D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen Zymoseptoria tritici. <em>BMC Biol.</em> <strong>18</strong>, 12 (2020).</p> <p>2. Goodwin, S. B. <em>et al.</em> Finished genome of the fungal wheat pathogen Mycosphaerella graminicola reveals dispensome structure, chromosome plasticity, and stealth pathogenesis. <em>PLoS Genet.</em> <strong>7</strong>, e1002070 (2011).</p> <p>3. Plissonneau, C., Hartmann, F. E. & Croll, D. Pangenome analyses of the wheat pathogen Zymoseptoria tritici reveal the structural basis of a highly plastic eukaryotic genome. <em>BMC Biol.</em> <strong>16</strong>, 5 (2018).</p> <p>4. Feurtey, A. <em>et al.</em> Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria. <em>BMC Genomics</em> <strong>21</strong>, 588 (2020).</p> <p>5. Baril, T., Imrie, R. M. & Hayward, A. Earl Grey: a fully automated user-friendly transposable element annotation and analysis pipeline. (2022) doi:10.21203/rs.3.rs-1812599/v1.</p> <p>6. Baril, T., Galbraith, J. & Hayward, A. <em>Earl Grey</em>. (Zenodo, 2023). doi:10.5281/ZENODO.8116025.</p> <p>7. Li, W. & Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. <em>Bioinformatics</em> <strong>22</strong>, 1658–1659 (2006).</p> <p>8. Fu, L., Niu, B., Zhu, Z., Wu, S. & Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data. <em>Bioinformatics</em> <strong>28</strong>, 3150–3152 (2012).</p> <p>9. Goubert, C. <em>et al.</em> A beginner’s guide to manual curation of transposable elements. <em>Mob. DNA</em> <strong>13</strong>, 7 (2022).</p> <p>10. Camacho, C. <em>et al.</em> BLAST+: Architecture and applications. <em>BMC Bioinformatics</em> <strong>10</strong>, 1–9 (2009).</p> <p>11. Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. <em>Mol. Biol. Evol.</em> <strong>30</strong>, 772–780 (2013).</p> <p>12. Notredame, C., Higgins, D. G. & Heringa, J. T-coffee: a novel method for fast and accurate multiple sequence alignment. <em>J. Mol. Biol.</em> <strong>302</strong>, 205–217 (2000).</p> <p>13. Rice, P., Longden, L. & Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite. <em>Trends Genet.</em> <strong>16</strong>, 276–277 (2000).</p> <p>14. Wheeler, T. J. & Eddy, S. R. nhmmer: DNA homology search with profile HMMs. <em>Bioinformatics</em> <strong>29</strong>, 2487–2489 (2013).</p> <p>15. Wicker, T. <em>et al.</em> A unified classification system for eukaryotic transposable elements. <em>Nat. Rev. Genet.</em> <strong>8</strong>, 973–982 (2007).</p>
Datasets and Pipeline V1.0 from An Atlas of Plant Transposable Elements
<p>In this repository, we deposited support data for the article "An Atlas of Plant Transposable Elements", available at <a href="http://apte.cp.utfpr.edu.br/">http://apte.cp.utfpr.edu.br/</a>.</p> <p>Here, we included:</p> <p><strong>1.) Supplementary material data:</strong><br> A) SuppMat_1.xlsx: The genome assembly reference access from Ensembl Plants species used.<br> B) SuppMat_2.docx: A brief transposable elements annotation steps are used in this work.</p> <p><strong>2.) Code and software: </strong>all script code create, third-party software, how we are used it, are detailed using Arabidopsis thaliana genome as an example in the GitHub: <a href="https://github.com/alerpaschoal/apte_pipeline">https://github.com/alerpaschoal/apte_pipeline</a> under the MIT license (please see details in licence.txt file). For the third part-software, consult their terms.</p> <p>To report bugs, to ask for help, and to give any feedback, please contact Alexandre R. Paschoal (paschoal@utfpr.edu.br) or Douglas S. Domingues (douglas.domingues@unesp.br).</p>
Paramecium Polycomb Repressive Complex 2 physically interacts with the small RNA binding PIWI protein to repress transposable elements
<p>Polycomb Repressive Complex 2 (PRC2) maintains transcriptionally silent genes in a repressed state via deposition of histone H3 K27 trimethyl (me3) marks. PRC2 has also been implicated in silencing transposable elements (TEs), yet how PRC2 is targeted to TEs remains unclear. To address this question, we identified proteins that physically interact with the <em>Paramecium</em> Enhancer-of-zeste Ezl1 enzyme, which catalyzes H3K9me3 and H3K27me3 deposition at TEs. We show that the <em>Paramecium</em> PRC2 core complex comprises four subunits, each required <em>in vivo</em> for catalytic activity. We also identify PRC2 cofactors, including the RNA interference (RNAi) effector Ptiwi09, which are necessary to target H3K9me3 and H3K27me3 to TEs. We find that the physical interaction between PRC2 and the RNAi pathway is mediated by a RING finger protein and that small RNA recruitment of PRC2 to TEs is analogous to the small RNA recruitment of H3K9 methylation SU(VAR)3-9 enzymes.</p>
Transposable elements in rice detected by TEF
<p>Next generation sequence data of '<a href="https://www.gene.affrc.go.jp/databases-core_collections_wr_en.php">World Rice Core Collection</a>' and '<a href="https://www.gene.affrc.go.jp/databases-core_collections_jr_en.php">Rice Core Collection of Japanese Landraces</a>' distributing NARO genebank have been analyzed by the software '<a href="https://pubmed.ncbi.nlm.nih.gov/36418944/">Transposable Elements Finder</a>'. This data is an additional supplementary data of the TEF paper.</p> <p>Transposition evidences were detected by direct comparison of NGS short reads between Japonica rice Nipponbare (wrc01) and other cultivars. Nearby 21,000 kinds of head and tail sequence pairs of TE have been identified by TEF. Head and tail sequences of TE, chromosome number, position, and name of detected rice cultivar were listed. </p> <ul> <li>This version is TE list of <em>Oryza sativa</em> detected by TEFv1.4.</li> <li>Positions of TE transpositions were mapped on <a href="https://rapdb.dna.affrc.go.jp/download/archive/irgsp1/IRGSP-1.0_genome.fasta">Os-Nipponbare-Reference-IRGSP-1.0</a> distributed from <a href="https://rapdb.dna.affrc.go.jp/">The Rice Annotation Project Database</a>.</li> <li>Accession Numbers of NGS data are listed in Japanese page '<a href="https://www.gene.affrc.go.jp/databases-core_collections_wr.php">World Rice Core Collection</a>' and in NCBI SRA page '<a href="https://www.ncbi.nlm.nih.gov/Traces/study/?acc=DRP006572&o=acc_s%3Aa">Rice Core Collection of Japanese Landraces</a>'.</li> <li>The TEF software is available at: <a href="https://github.com/akiomiyao/tef">https://github.com/akiomiyao/tef</a></li> <li>Miyao, A., Yamanouchi, U. Transposable element finder (TEF): finding active transposable elements from next generation sequencing data. <em>BMC Bioinformatics</em> <strong>23</strong>, 500 (2022). <a href="https://doi.org/10.1186/s12859-022-05011-3">https://doi.org/10.1186/s12859-022-05011-3</a></li> </ul>
Data and Code for "Cell Type-specific Genome Scans of DNA Methylation Diversity Indicate an Important Role for Transposable Elements"
<p>This is a release of the gitlab repository "meta-methylome" (https://gitlab.com/okartal/meta-methylome.git) that, in addition to the code, also contains the resulting genomic data.</p> <p>Extract the directory on the command line using</p> <pre><code class="language-bash">$ tar -xhzvf meta-methylome.tar.gz</code></pre> <p>to preserve the symbolic links.</p>
Generation of transcriptional novelty by transposable element insertions in Arabidopsis, Genome Sequencing and eccDNA Data
<p><strong>Raw Illumina sequencing data from the Manuscript entitled "Generation of transcriptional novelty by transposable element insertions in Arabidopsis"</strong></p> <p><strong>A. Illumina genome sequencing reads of Arabidopsis control and hcLines that contain novel transposable element insertions.</strong></p> <p>To identify the genomic position of the new <em>ONSEN</em> insertions, the extracted DNA of the 11 selected lines (nine lines with new insertions and two control lines) was sent to BGI, Hong-Kong for Illumina paired-end 150 bp sequencing, aiming for a minimum of 20X sequencing coverage. Quality control of the raw reads was done using FastQC (Andrews S. (2010). FastQC: a quality control tool for high throughput sequence data. Available online at: <a href="http://www.bioinformatics.babraham.ac.uk/projects/fastqc">http://www.bioinformatics.babraham.ac.uk/projects/fastqc</a>) and trimming/clipping was done using Trimmomatic with parameters ILLUMINACLIP: TruSeq3:2:30:10 LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 and MINLEN:36. Quality of the reads was deemed excellent and no further actions were taken.</p> <p>Samples identifications: genome_hcLineX with "_1" indicating the forward and "_2" the reverse reads.</p> <p><strong>B. Illumina eccDNA sequencing of Arabidopsis control and hcLines following stress treatments</strong></p> <p>Extrachromosomal circular DNA was prepared and sequenced as follows: twenty plants from each petri dish were pooled separately and DNA was extracted using the CTAB method (<a href="https://dx.doi.org/10.17504/protocols.io.quidwue">dx.doi.org/10.17504/protocols.io.quidwue</a>). Following the mobilome-seq method described in (Lanciano et al., 2017), for all samples, we digested linear DNA from 2 µg of total DNA for 17 hours at 37<sup>o</sup>C using 10 U of PlasmidSafe (<em>LubioScience cat# E3101K</em>), followed by enzyme denaturation (30 mins at 70<sup>o</sup>C). Digested DNA was precipitated with isopropanol supplemented with 1 µg of GlycoBlue coprecipitant (<em>Fisher Scientific cat# 10391565</em>). Circular DNA was then amplified through rolling circle amplification (RCA) with the Illustra TempliPhi kit (<em>GE Healthcare cat# 25-6400-10</em>), following the manufacturer recommendation and leaving the reaction for 16h at 30<sup>o</sup>C. DNA was once again precipitated with isopropanol and sent for Illumina paired end 150 bp sequencing at BGI, Hong Kong. </p> <p>Samples identification: </p> <p>eccDNA_A.thaliana_ctrl: control reads</p> <p>eccDNA_A.thaliana_HS: heat stressed plants reads</p> <p>eccDNA_A.thaliana_AZ_HS: reads of alpha-amanitin, zebularine and heat-stressed plants</p> <p>"R1" indicates forward and "R2" reverse reads.</p> <p> </p>
1000 Genomes Project Transposable Element database
<p>Multi-sample VCF with transposable elements across individuals in the 1KGP dataset. Transposable elements were called using RetroSeq </p>
Transposable element annotation Rhynchosporium commune isolate UK7
<p>To obtain a consensus sequence for each TE family, RepeatModeler v. open-4.0.7 (http://www.repeatmasker.org/RepeatModeler/) was run on the <em>R. commune</em> UK7 reference genome. The classification was based on the GIRI Repbase (v. 2018) using RepeatMasker v. open-4.0.7. (Smit, Hubley, and P. 2015; Bao, Kojima, and Kohany 2015). We used WICKERsoft to finalize the classification of TE consensus sequences (Breen et al. 2010). Specifically, we used WICKERsoft to screen for copies of known consensus sequences from other fungal species with blastn filtering for sequence identity > 80% and sequence length > 80%. (Altschul et al. 1997). Then, using WICKERsoft, flanks of 10000 bp were added and visually inspected for sequence similarity and terminal repeats with dot plots. Subsequent multiple sequence alignments were performed with 10-15 sequences using ClustalW (Thompson, Higgins, and Gibson 1994). Alignment boundaries were visually inspected in WICKERsoft and trimmed if necessary. Using WICKERsoft, consensus sequences were classified according to the presence and type of terminal repeats, as well as homology of the encoded proteins based on blastx against the NCBI protein database. Consensus sequences were named according to the three-letter classification system (Wicker et al. 2007). The reference genome was annotated with the curated consensus sequences using RepeatMasker v. open-4.0.7 with a cut-off value of 250 (Smit, Hubley, and P. 2015). Simple repeats, low complexity regions and annotated elements shorter than 100 bp were filtered out and adjacent identical TEs overlapping by more than 100 bp were merged as belonging to the same TE family. Different TE families overlapping by more than 100 bp were considered as nested insertions and were renamed accordingly. Identical elements separated by less than 200 bp are indicative of interrupted elements and were grouped into a single element. TEs overlapping genes were recovered using the bedtools v. 2.27.1 suite and the “overlap” function (Quinlan and Hall 2010).</p>
Data from: Hybrid incompatibility between D. virilis and D. lumei is stronger in the presence of transposable elements
<p>Mismatches between parental genomes in selfish elements are frequently hypothesized to underlie hybrid dysfunction and drive speciation. However, because the genetic basis of most hybrid incompatibilities is unknown, testing the contribution of selfish elements to reproductive isolation is difficult. Here we evaluated the role of transposable elements (TEs) in hybrid incompatibilities between Drosophila virilis and D. lummei by experimentally comparing hybrid incompatibility in a cross where active TEs are present in D. virilis (TE+) and absent in D. lummei, to a cross where these TEs are absent from both D. virilis (TE-) and D. lummei genotypes. Using genomic data, we confirmed copy number differences in TEs between the D. virilis (TE+) strain and the D. virilis (TE-) strain and D. lummei. We observed F1 postzygotic reproductive isolation exclusively in the interspecific cross involving TE+ D. virilis but not in the cross involving TE- D. virilis. This precisely mirrors the intraspecies dysgenic phenotype where teste atrophy only occurs when TE+ D. virilis is the paternal parent. A series of backcross experiments, designed to account for alternative models of hybrid incompatibility, showed that both F1 hybrid incompatibility and intrastrain dysgenesis is consistent with the action of TEs rather than other, genic, interactions. A further Y-autosome interaction contributes to additional, sex-specific, inviability in one direction of this cross combination. These experiments demonstrate that TEs that cause intraspecies dysgenesis can increase reproductive isolation between closely related lineages, thereby adding to the processes that consolidate speciation.</p>
The pan-genome unearths gene content and transposable element variations in modern pigs
<p>Genes, gene annotations, proteins, and sequences identified in the non-reference genome of the pig pan-genome. Transposable insertion polymorphisms (TIP) indentified in the pig mobolome.</p>
Generation of transcriptional novelty by transposable element insertions in Arabidopsis, RNAseq Control Condition Sequencing Data
<p><strong>Arabidopsis stranded 150 bp paired end RNA sequencing data (Illumina) of plants that were grown under control conditions for the manuscript "Generation of transcriptional novelty by transposable element insertions in Arabidopsis"</strong></p> <p><strong><strong>Plant growth conditions</strong></strong></p> <p>Sequenced F4 seeds were sterilized for 10 minutes in 10% bleach, rinsed, and stratified at 4°C for four days in the dark before being sown on 0.5x Murashige & Skoog media (Du<em>schefa cat# M0222</em>) and transferred to growth chambers under long day conditions (16h of light at 24°C followed by 8h of darkness at 21°C; 20 seeds per plate, 6 replicate plates). Ten days after sowing, plants were subjected to 6°C for 24 hours and control plants were returned to normal long day growing conditions for 24 hours before harvesting (3 replicate plates per condition).</p> <p><strong><strong>RNA extraction and sequencing</strong></strong></p> <p>Seedlings were harvested and RNA extractions were done on pools of 5 plants. RNA extractions were performed for 3 biological replicate samples for each line in each condition (n=96) using the Macherey-Nagel NucleoSpin RNA kit (cat# 740955.50). Samples were sent to Novogene for Illumina 150bp paired-end sequencing using a stranded poly-A library.</p> <p><strong>RNAseq sample descriptions of the plants grown under control conditions</strong></p> <p>wt_control: wild-type plants.</p> <p>wtHS_control: wild-type plants that have been submitted to heat stress in a previous generation.</p> <p>wtAZ_control: wild-type plants that have been submitted to epigenetic drug treatments (alpha-amanitin and zebularine) in a previous generation.</p> <p>htLine#: plants carrying additional <em>ONSEN</em> transposable element insertions.</p> <p>Files description: Forward and reverse strand RNA seq data are combined in one file. The numbering at the end ("_1") denominates the biological replicate number.</p>
Datasets from An Atlas of Plant Transposable Elements
<p>In this repository, we deposited support data for the article "An Atlas of Plant Transposable Elements", available at <a href="http://apte.cp.utfpr.edu.br/">http://apte.cp.utfpr.edu.br/</a>.</p> <p>Here, we included:</p> <p><strong>1.) Supplementary material data:</strong><br> A) SuppMat_1.xlsx: The genome assembly reference access from Ensembl Plants species used.<br> B) SuppMat_2.docx: A brief transposable elements annotation steps used in this work.</p> <p><strong>2.) Code and software: </strong>all script code create, third part software, how we used it, are detailed using Arabidopsis thaliana genome as an example in the GitHub: <a href="https://github.com/daniellonghi/te_pipeline">https://github.com/daniellonghi/te_pipeline</a> under the MIT license (please see details in licence.txt file). For third part-software, consult their terms.</p> <p>To report bugs, to ask for help, and to give any feedback, please contact Alexandre R. Paschoal (paschoal@utfpr.edu.br) or Douglas S. Domingues (douglas.domingues@unesp.br).</p>
Datasets and Scripts associated with "Transposable elements are associated with the variable response to influenza infection"
<p>Scripts and datasets included here were used for the main analyses in the article "Transposable elements are associated with the variable response to influenza infection" (BioRxiv doi: <a href="https://doi.org/10.1101/2022.05.10.491101">https://doi.org/10.1101/2022.05.10.491101</a>).</p> <p>Study Summary:</p> <p>Influenza A virus (IAV) infections are frequent every year and result in a range of disease severity. Given that the regulation of transposable elements (TEs) contributes to the activation of innate immunity, we wanted to explore their potential role in this variability. Transcriptome profiling in monocyte-derived macrophages from 39 individuals following IAV infection revealed significant inter-individual variation in viral load post-infection. Using ATAC-seq we identified a set of TE families with either enhanced or reduced accessibility upon infection. Of the enhanced families, 15 showed high variability between individuals and had distinct epigenetic profiles. Motif analysis showed an association with known immune regulators (e.g., BATFs, FOSs/JUNs, IRFs, STATs, NFkBs, NFYs, and RELs) in stably enriched TE families and with other factors in variable families, including KRAB-ZNFs. We also observed a strong association between basal TE transcripts and viral load post infection and showed that TEs, and host factors regulating TEs, were predictive of the response. Our findings shed light on the variable transcriptional and epigenetic response to infection and the role TEs and KRAB-ZNFs may play in inter-individual variation in immunity.</p> <p> </p> <p> </p> <p> </p> <p> </p>
Transposable element libraries from 101 fish
<p><span>Repetitive DNA make up a considerable fraction of most eukaryotic genomes. In fish, transposable element (TE) activity has coincided with rapid species diversification. Here, we annotated the repetitive content in 100 genome assemblies, covering the major branches of the diverse lineage of teleost fish. We investigated if TE content correlates with family level net diversification rates and found support for a weak negative correlation. Further, we found that </span><span>TE proportion correlate to genome size, but not to the proportion of short tandem repeats (STRs), </span><span>which implies independent evolutionary paths. </span><span>Marine</span><span> and freshwater fish have large differences in STR content. The most extreme propagation was found in the genomes of codfish species and Atlantic herring. Such a high density of STRs is likely to increase the mutational load, which we propose could be counterbalanced by high fecundity as seen in codfishes and herring.</span></p>
Transposable element libraries from 101 fish
Open the record for dataset details and reuse information.
Combined analysis of transposable elements and structural variation in maize genomes reveals genome contraction outpaces expansion
Open the record for dataset details and reuse information.
Data from: Transposable element diversity and activity patterns in neotropical salamanders
Open the record for dataset details and reuse information.
Data from: Hybrid incompatibility between D. virilis and D. lumei is stronger in the presence of transposable elements
Open the record for dataset details and reuse information.
Comparative analyses of the Hymenoscyphus fraxineus and Hymenoscyphus albidus genomes reveals potentially adaptive differences in secondary metabolite and transposable element repertoires
<p><strong>Background </strong>The dieback epidemic decimating common ash (<em>Fraxinus excelsior</em>) in Europe is caused by the invasive fungus <em>Hymenoscyphus fraxineus</em>. In this study we analyzed the genomes of <em>H. fraxineus</em> and <em>H. albidus</em>, its native but, now essentially displaced, non-pathogenic sister species, and compared them with several other members of <em>Helotiales</em>. The focus of the analyses was to identify signals in the genome that may explain the rapid establishment of <em>H. fraxineus</em> and displacement of <em>H. albidus</em>.</p> <p><strong>Results</strong> The genomes of <em>H. fraxineus</em> and <em>H. albidus </em>showed a high level of synteny and identity. The assembly of <em>H. fraxineus </em>is 13 Mb longer than that of <em>H. albidus’, </em>most of this difference can be attributed to higher dispersed repeat content ((i.e transposable elements [TEs]) in <em>H. fraxineus</em>. In general, TE families in <em>H. fraxineus</em>showed more signals of repeat-induced point mutations (RIP) than in <em>H. albidus</em>, especially in Long-terminal repeat (LTR)/Copia and LTR/Gypsy elements. Comparing gene family expansions and 1:1 orthologs, relatively few genes show signs of positive selection between species. However, several of those that did appeared to be associated with secondary metabolite genes families, including gene families containing two of the genes in the <em>H. fraxineus-</em>specific, <em>hymenosetin </em>biosynthetic gene cluster (BGC).</p> <p><strong>C</strong><strong>onclusion </strong>The genomes of <em>H. fraxineus</em> and <em>H. albidus</em> show a high degree of synteny, and are rich in both TEs and BGCs, but the genomic signatures also indicated that <em>H. albidus</em> may be less well equipped to adapt and maintain its ecological niche in a rapidly changing environment. </p> <p><strong>Data included</strong></p> <p>This post contains the alternate structural and functional annotations of the genomes of Helotealean fungi used in the study.</p>
Data from: Diversity and evolution of the transposable element repertoire in arthropods with particular reference to insects
Background: Transposable elements (TEs) are a major component of metazoan genomes and are associated with a variety of mechanisms that shape genome architecture and evolution. Despite the ever-growing number of insect genomes sequenced to date, our understanding of the diversity and evolution of insect TEs remains poor. Results: Here, we present a standardized characterization and an order-level comparison of arthropod TE repertoires, encompassing 62 insect and 11 outgroup species. The insect TE repertoire contains TEs of almost every class previously described, and in some cases even TEs previously reported only from vertebrates and plants. Additionally, we identified a large fraction of unclassifiable TEs. We found high variation in TE content, ranging from less than 6 % in the antarctic midge (Diptera), the honey bee and the turnip sawfly (Hymenoptera) to more than 58 % in the malaria mosquito (Diptera) and the migratory locust (Orthoptera), and a possible relationship between the content and diversity of TEs and the genome size. Conclusion: While most insect orders exhibit a characteristic TE composition, we also observed intraordinal differences, e.g., in Diptera, Hymenoptera, and Hemiptera. Our findings shed light on common patterns and reveal lineage-specific differences in content and evolution of TEs in insects. We anticipate our study to provide the basis for future comparative research on the insect TE repertoire.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.