Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
380
datasets available to search
ShareScore release 0.7.1
Dataset results
380 results for “Transposable elements”
Data from: Analysis of transposable elements in the genome of Asparagus officinalis from high coverage sequence data
Asparagus officinalis is an economically and nutritionally important vegetable crop that is widely cultivated and is used as a model dioecious species to study plant sex determination and sex chromosome evolution. To improve our understanding of its genome composition, especially with respect to transposable elements (TEs), which make up the majority of the genome, we performed Illumina HiSeq2000 sequencing of both male and female asparagus genomes followed by bioinformatics analysis. We generated 17 Gb of sequence (12×coverage) and assembled them into 163,406 scaffolds with a total cumulated length of 400 Mbp, which represent about 30% of asparagus genome. Overall, TEs masked about 53% of the A. officinalis assembly. Majority of the identified TEs belonged to LTR retrotransposons, which constitute about 28% of genomic DNA, with Ty1/copia elements being more diverse and accumulated to higher copy numbers than Ty3/gypsy. Compared with LTR retrotransposons, non-LTR retrotransposons and DNA transposons were relatively rare. In addition, comparison of the abundance of the TE groups between male and female genomes showed that the overall TE composition was highly similar, with only slight differences in the abundance of several TE groups, which is consistent with the relatively recent origin of asparagus sex chromosomes. This study greatly improves our knowledge of the repetitive sequence construction of asparagus, which facilitates the identification of TEs responsible for the early evolution of plant sex chromosomes and is helpful for further studies on this dioecious plant.
Manually Curated Library of Transposable Elements (TEs) and TE Annotations for Drosophila amaguana
<p>This data collection provides a manually curated library of transposable elements (TEs) for <em>Drosophila amaguana</em>, including consensus sequences, genome-wide TE annotations, and individual TE copy sequences. The library was built using <em>de novo</em> generated by EDTA (Extensive <em>de novo</em> TE Annotator) and RepeatModeler, curated with MCHelper, and further processed for genome annotation using RepeatMasker and OneCodeToFindThemAll. </p> <p>Below is a description of the included files:</p> <ul> <li><strong><code>Dama_curated_TE_library.fasta</code>:</strong> Contains 737 consensus TE sequences manually curated for <em>D. amaguana</em>. Sequence identifiers include classification and origin (e.g., new families or similarity to known elements). The identifier for each sequence in the FASTA file includes information about its classification:<br><br> <ul> <li><strong>For sequences corresponding to potentially new families</strong>: The identifier consists of a three-letter abbreviation for <em>D. amaguana</em> (Dama), followed by the new family identifier and the superfamily name, all separated by underscores. <em>Example: </em>Dama_NF_BELPAO_1.</li> <li><strong>For consensus sequences that show similarity to TE sequences previously reported in other species</strong>: The identifier includes the abbreviation Dama, followed by the superfamily name and an abbreviation for the species in which the TE was previously reported, all separated by underscores. Example: Dama_Helitron-1_DVir.<br><br></li> </ul> </li> <li><strong><code>Dama_TE_annotations.out</code>:</strong> Genome-wide annotation file of TE insertions produced with RepeatMasker and post-processed using OneCodeToFindThemAll to merge fragmented elements.</li> <li><strong><code>Dama_TE_copies.fasta</code>: </strong>FASTA file containing the extracted sequences of all annotated TE copies from the <em>D. amaguana</em> genome.</li> <li><strong><code>TEcopies_sequences.sh</code>:</strong> Shell script used to extract TE copy sequences from the genome using the annotation coordinates.</li> <li><strong><code>Dynamics_Dama.ipynb</code>: </strong>Jupyter Notebook for the analysis of transposable element (TE) dynamics in<strong> </strong><em>D. amaguana.</em></li> </ul>
Transposable element products, functions, and regulatory networks in Arabidopsis thaliana
<h1>README</h1> <p>This dataset includes the main outputs from the work titled <strong>Transposable element products, functions, and regulatory networks in <em>Arabidopsis thaliana</em>.</strong></p> <h2><strong>Summary</strong></h2> <p>Transposable elements (TEs) are DNA sequences with the ability to propagate themselves within and across genomes. Their mobilization is catalyzed by self-encoded factors, yet these factors have been poorly investigated due to difficulties in defining TE genes in genomes. Here, we leveraged extensive long- and short-read transcriptome data, together with structural predictions, transcription factor binding site identification, and transcriptional network analyses, to construct a comprehensive atlas of TE transcripts and TE-encoded products in the model organism <em>Arabidopsis thaliana</em>. We uncovered hundreds of transcriptionally competent TEs, each potentially encoding multiple proteins either through distinct genes, alternative splicing, or post-translational processing. Structural-based protein analyses revealed dozens of hitherto unidentified domains of unknown function, enabling us to predict proteins with multimerization and DNA binding domains forming macromolecular complexes involved in transposition. Furthermore, we demonstrate that TE expression is highly intertwined with the transcriptional network of cellular genes, and identified transcription factors and cis-regulatory elements associated with their coordinated expression during development or in response to environmental cues. This comprehensive atlas of TE-genes and TE-proteins provides a valuable resource for studying the mechanisms involved in transposition and their consequences for genome and organismal function.</p> <h2><strong>File description</strong></h2> <p>It includes the following data:</p> <ol> <li><code>annots/TE_Functional_Annotation.Borreda2024.gtf</code> - Annotation file including Arabidopsis TEs and TE-genes. TE-genes defined in our work are indicated in the 'Source' column of the gtf. TAIR10-defined TEs for whom we did not annotate new transcripts are also included.</li> <li><code>seqs</code> - This folders includes all the transcript sequences (cDNAs.tsv) and the first and longest ORFs found in each of them (prot.csv), which were used for further analyses. The specific copy, gene, isoform and, in the case of proteins, ORF, is indicated for each sequence.</li> <li><code>structures</code> - The zipped folder <code>full_length_prots_pdbs.zip</code> includes all the 3D structures from full-length TE proteins. Note that identical proteins, which would result in identical structures, have been collapsed to reduce the total dataset size; equivalences can be found in <code>identical_proteins</code>.</li> <li><code>structures/SD_Cluster_Functions.tsv</code> - We clustered all Structural Domains (SD) based on 3D similarity and assigned a function to each cluster based on the database hits. This table indicated, for each of these SDs, to which cluster it belongs, the superfamily, family and element containing it, the number of Conserved Domains included within it and the number of hits with resolved (retrieved from the RCSB-PDB database) or predicted (AlphaFold2) protein structures. The last column includes the putative function assigned to each cluster.</li> <li><code>coexpression</code> - Coexpressed genes were classiffied into modules using WGCNA. In the table <code>Gene_Modules.tsv</code> we include, for each gene and TE-gene (provided it has expression in at least one sample, see methods on the publication for details), the TE family and superfamily when applicable and the module to which it belongs. The modules were named based on the results of the GO enrichment analysis of the genes contained. The results of this GO enrichment are included in <code>GO_Enrichment.tsv</code>, where we include the main funciton of the associated GOs, the number of entries and TE-genes within the module, a list of GO terms enriched in that specific module and finally a list of TE families enriched in each module.</li> <li><code>dapseq</code> - We reanalized the DAP-seq dataset from O'Malley 2016, selecting only TFBS with a binding site within a DAP-seq peak. The list of filtered peaks we found is reported in <code>DAPseq_TFBS_Motifs.tsv</code>. The columns include the coordinates of the TFBS (which have been filtered to fall within a DAP-seq peak and include the TFBS motif), the strand of the motif, the score of the motif reported by FIMO, the motif sequence, the Sequence Read identified for the original DAP-seq data, and the family, name and gene of the TF associated with that specific peak.</li> </ol>
Investigating the cis-regulatory functions of transposable elements
<p>This repository contains the data and code associated with the article ‘<strong><span>Multi</span></strong><strong>-</strong><strong><span>omics analysis reveals critical cis-regulatory roles of transposable elements in livestock genomes</span></strong>'</p>
Data From - TE Density: a tool to investigate the biology of transposable elements
<p><strong>Background</strong>: Transposable elements (TEs) are powerful creators of genotypic and phenotypic diversity due to theirinherent mutagenic capabilities and in this way they serve as a deep reservoir of sequences for genomic variation. As agents of genetic disruption, a TE's potential to impact phenotype is partially a factor of its location in the genome. Previous research has shown TEs' ability to impact the expression of neighboring genes, however our understanding of this trend is hampered by the exceptional amount of diversity in the TE world, and a lack of publicly availablecomputational methods that quantify the presence of TEs relative to genes.</p> <p><strong>Results: </strong>Here, we have developed a tool to more easily quantify TE presence relative to genes through the useof only a gene and TE annotation, yielding a new metric we call TE density. Briefly defined as the proportion of TE-occupied base-pairs relative to a window-size of the genome. This new pipeline reports TE density for each gene in the genome, for each type descriptor of TE (order and superfamily), and for multiple positions and distances relative to the gene (upstream, intragenic, and downstream) over sliding, user-defined windows. In this way, we overcome previous limitations to the study of TE-gene relationships by focusing on all TE types present in the genome, utilizing flexible genomic distances for measurement, and reporting a TE presence metric for every gene in the genome.</p> <p><strong>Conclusions: </strong>Together, this new tool opens up new avenues for studying TE-gene relationships, genome architecture,comparative genomics, along with the tremendous diversity present in the TE world.</p> <p><strong>Data Availability: </strong>TE Density is open-source and freely available at: https://github.com/sjteresi/TE_Density.</p>
A Unified Framework to Analyze Transposable Element Insertion Polymorphisms using Graph Genomes
<p>GraffiTE_analyses_scripts_and_data_052924 <br>**********************************************</p> <p>Contact/info: goubert.clement@gmail.com </p> <p>Figures: <br>^^^^^^<br>This directory contains all the files and R script necessary to reproduce the data figures of the manuscript. </p> <p>GraffiTE pMEs datasets:<br>^^^^^^^^^^^^^^^^^^<br>This directory contains new pMEs datasets produced by GraffiTE for the different models presented in the manuscript. All datasets have been filtered according to the descriptions given in the Methods and Supplementary Methods sections of the manuscript. These data, along with more intermediate file, metadata and analysis script used to produce the figure are also available in the "Figures" directory.</p> <p> </p>
De novo assemblies for the manuscrip "Candida albicans isolates contain frequent heterozygous structural variants and transposable elements within genes and centromeres"
Open the record for dataset details and reuse information.
Data from: Chironomus riparius (Diptera) genome sequencing reveals the impact of minisatellite transposable elements on population divergence
Active transposable elements (TEs) may result in divergent genomic insertion and abundance patterns among conspecific populations. Upon secondary contact, such divergent genetic backgrounds can theoretically give rise to classical Dobzhansky-Muller incompatibilities (DMI), thus contributing to the evolution of endogenous genetic barriers and eventually cause population divergence. We investigated differential TE abundance among conspecific populations of the non-biting midge Chironomus riparius and evaluated their potential role in causing endogenous genetic incompatibilities between these populations. We focussed on a Chironomus-specific TE, the minisatellite-like Cla-element, whose activity is associated with speciation in the genus. Using a newly generated and annotated draft genome for a genomic study with five natural C. riparius populations, we found highly population-specific TE insertion patterns with many private insertions. A significant correlation of the pairwise FST estimated from genome-wide single nucleotide polymorphisms (SNPs) and the FST estimated from TEs, is consistent with drift as the major force driving TE population differentiation. However, the significantly higher Cla-element FST level due to a high proportion of differentially fixed Cla-element insertions also indicates selection against segregating (i.e. heterozygous) insertions. With reciprocal crossing experiments and fluorescent in-situ hybridisation of Cla-elements to polytene chromosomes, we documented phenotypic effects on female fertility and chromosomal mispairings. We propose that the inferred negative selection on heterozygous Cla-element insertions may cause endogenous genetic barriers and therefore acts as DMI among C. riparius populations. The intrinsic genomic turnover exerted by TEs may thus have a direct impact on population divergence that is operationally different from drift and local adaptation.
Transposable element annotation of Sitophilus oryzae
<p>Sitophilus oryzae TE annotation (GFF), and consensus sequences (fasta).</p>
VCF files for D. serrata transposable elements
<p><span><span>Transposable elements are an important element of the complex genomic ecosystem. Transposable element insertion also appears to be bursty – either due to invasion of new transposable elements that are not yet repressed, de-repression due to instability of organismal defense systems, stress, or genetic variation in hosts. Here, we characterize the transposable element landscape in an important model <i>Drosophila</i>, <i>D. serrata</i>, and investigate variation in transposable element copy number between genotypes and in the population at large. We find that a subset of transposable elements are clearly related to elements annotated in <i>D. melanogaster</i> and <i>D. simulans</i>, suggesting they spread between species more recently than other transposable elements. We also find that some transposable elements proliferate in particular genotypes compared to population levels. In natural populations an active transposable element and a potentially permissive background would not be held in association as in inbred lines, thus this could be a product of inbreeding. Yet many of the inbred lines have actively proliferating transposable elements suggesting that active transposable elements are not uncommon and that genotypes vary in their permissiveness in populations. </span></span></p>
Dataset for "Population-level transposable element expression dynamics influence trait evolution in a fungal crop pathogen"
<p><strong>Supplementary Tables</strong></p> <p><strong>Supplementary Table S1: </strong>SRA accession list of RNAseq reads.</p> <p><strong>Supplementary Table S2:</strong> Genomic localization of TEs in gene elements and 10 kb windows upstream and downstream of the transcription start site (TSS) in the reference genome IPO323.</p> <p><strong>Supplementary Table S3:</strong> Genome-wide TE insertion polymorphism (TIPs) in the pathogen population. 0 represents TE absence and 1 represents TE presence.</p> <p><strong>Supplementary Table S4:</strong> Gene expression (log-transformed RPKM) values across the population.</p> <p><strong>Supplementary Table S5:</strong> Locus-specific transcript abundance at individual TE loci (FPKM) across individuals.</p> <p><strong>Supplementary Table S6:</strong> Percent of expressed copies within each TE family in the reference genome IPO323 and percent expressed TE copies in each TE family across the population.</p> <p><strong>Supplementary Table S7:</strong> Linkage disequilibrium of TIP in the genome and neighboring SNPs within the 600bp distance from the TE loci.</p> <p><strong>Supplementary Table S8:</strong> Genome-wide association mapping of the virulence-associated trait (PLACP: percent leaf area covered by pycnidia) and TE insertion polymorphisms in the genome.</p> <p><strong>Supplementary Table S9:</strong> TIPs in the genome significantly associated with metabolite peak intensity variation in the pathogen population (filtered by Bonferroni threshold).</p>
Selective control of transposable element expression during T cell exhaustion and anti-PD1 treatment
<p>Supplementary data from (Bonté et al.) which contains Human and Mouse RNA-seq data including : </p> <ul> <li><strong>Raw counts and TPM/FPKM expression matrices of data analyzed </strong>(genes, individual TEs, TE subfamilies) <ul> <li>raw.counts / fpkm.count / tpm.counts rds files</li> </ul> </li> <li><strong>SQuiRE raw output files </strong> <ul> <li>SQuiRE.results rds files</li> </ul> </li> <li><strong>Seurat processed single cell object</strong> (with genes, TEs, Subfamilies) <ul> <li>SC.object rds files</li> </ul> </li> </ul> <p>In details, Mouse and Human data are labelled as follows :</p> <ul> <li><strong>Mouse.SC :</strong> <ul> <li><strong>Type of data</strong> : Single cell RNA-seq data of sorted CD8+ TILs (model B16 melanoma tumors )</li> <li><strong>Label of samples </strong>: Naïve like, Early activated, EffectorMemory, Tpex, Tex</li> <li><strong>Origin of the data</strong> : S. J. Carmona, I. Siddiqui, M. Bilous, W. Held, D. Gfeller, Deciphering the transcriptomic landscape of tumor-infiltrating CD8 lymphocytes in B16 melanoma tumors with single-cell RNA-Seq. Oncoimmunology 9, 1737369 (2020).</li> </ul> </li> <li><strong>Mouse.tumor.TILs :</strong> <ul> <li><strong>Type of data</strong> : bulk RNA-seq data of sorted CD8+ TILs (model B16 melanoma tumors)</li> <li><strong>Label of samples </strong>: Tumor Slamf6+ or Tumor Tim3+</li> <li><strong>Origin of the data</strong> : B. C. Miller<em> et al.</em>, Subsets of exhausted CD8+ T cells differentially mediate tumor control and respond to checkpoint blockade. <em>Nature Immunology</em> <strong>20</strong>, 326-336 (2019).</li> </ul> </li> <li><strong>Mouse.LCMV.TILs :</strong> <ul> <li><strong>Type of data</strong> : bulk RNA-seq data of sorted CD8+ TILs (model LCMV clone 13)</li> <li><strong>Label of samples </strong>: LCMV Slamf6+ or LCMV Tim3+</li> <li><strong>Origin of the data</strong> : B. C. Miller<em> et al.</em>, Subsets of exhausted CD8+ T cells differentially mediate tumor control and respond to checkpoint blockade. <em>Nature Immunology</em> <strong>20</strong>, 326-336 (2019).</li> </ul> </li> <li><strong>Mouse.LCMV_Fli1.TILs :</strong> <ul> <li><strong>Type of data</strong> : bulk RNA-seq data of sorted CD8+ TILs (model LCMV clone 13)</li> <li><strong>Label of samples </strong>: WT or Fli1KO</li> <li><strong>Origin of the data</strong> : Z. Chen<em> et al.</em>, In vivo CD8(+) T cell CRISPR screening reveals control by Fli1 in infection and cancer. <em>Cell</em> <strong>184</strong>, 1262-1280 e1222 (2021).</li> </ul> </li> <li><strong>Human.TILs </strong>: <ul> <li><strong>Type of data</strong> : bulk RNA-seq data of sorted CD8+ TILs (NSCLC tumor tissue )</li> <li><strong>Label of samples </strong>: Tex or Tpex</li> <li><strong>Origin of the data</strong> : Bonté et al. (ongoing submission)</li> </ul> </li> </ul> <p> </p> <p>Type of features is labeled as follows:</p> <ul> <li><strong>Gene expression</strong> : genes</li> <li><strong>Individual TE expression</strong> : TEs</li> <li><strong>TE subfamiliy expression</strong> : SF</li> </ul> <p>Metadata corresponding to each expression matrix can be found <a href="https://github.com/pebonte/TE_TILs">here</a>.</p> <p> </p>
Imaging files for: ACD15, ACD21, and SLN regulate accumulation and mobility of MBD6 to silence genes and transposable elements
<p>It's well known that DNA methylation is linked to gene silencing but the mechanisms of how proteins that bind the DNA methylation cause gene silencing remains unclear. We demonstrated that the novel MBD5/6 complex contains three chaperone proteins, called ACD15, ACD21, and SLN, which specifically mediate the gene silencing function. ACD15 and ACD21 bridge the interaction of SLN to MBD5 and/or MBD6 while also functioning to drive the accumulation of the MBD5/6 complex at CG methylation sites. We further discovered that SLN also regulates the accumulation of the MBD5/6 complex and regulates the turnover of all protein members once accumulated at meCG sites. </p> <p>To demonstrate these results we primarily used fluorescence, confocal microscopy using RFP tagged MBD6, YFP tagged ACD15, and CFP tagg ACD21 and SLN imaging the roots of <em>Arabidopsis thaliana</em>. We expressed these constructs either alone or together in multiple mutant lines including <em>mbd5 mbd6, acd15, acd21, acd15 acd21,</em> <em>sln, </em>and <em>acd15 acd21 </em><em>sln </em>mutant plants. We also truncated MBD6 to remove the C terminus, called the C term deletion, as well as a version of MBD6 lacking the C terminus but with the domain necessary for interaction with ACD15 added back (MBD6 plus StkyC). With these microscopy experiments, we were able to demonstrate the specificity of ACD15 for the StkyC domain, the function of the chaperone proteins, and linke protein accumulation with gene silencing.</p>
Supplementary Tables for "Population-level transposable element expression dynamics influence trait evolution in a fungal crop pathogen"
<p><strong>Supplementary Table S1:</strong> Genomic localization of TEs in gene elements and 10 kb windows upstream and downstream of the transcription start site (TSS) in the reference genome IPO323.</p><p><strong>Supplementary Table S2:</strong> Genome-wide TE insertion polymorphism (TIPs) in the pathogen population. 0 represents TE absence and 1 represents TE presence.</p><p><strong>Supplementary Table S3:</strong> Gene expression (log transformed RPKM) values across the population.</p><p><strong>Supplementary Table S4:</strong> Locus-specific transcript abundance at individual TE loci (FPKM) across individuals.</p><p><strong>Supplementary Table S5:</strong> Percent of expressed copies within each TE family in the reference genome IPO323 and percent expressed TE copies in each TE family across the population.</p><p><strong>Supplementary Table S6:</strong> Linkage disequilibrium of TIP in the genome and neighboring SNPs within the 600bp distance from the TE loci.</p><p><strong>Supplementary Table S7:</strong> Genome-wide association mapping of the virulence-associated trait (PLACP: percent leaf area covered by pycnidia) and TE insertion polymorphisms in the genome.</p><p><strong>Supplementary Table S8:</strong> metabolite peak variation in the pathogen population for individual isolates</p><p><strong>Supplementary Table S9:</strong> TIPs in the genome significantly associated with metabolite peak intensity variation in the pathogen population (filtered by Bonferroni threshold).</p><p><strong>Supplementary Table S10: </strong>SRA accession list of RNAseq reads.</p>
Imaging files for: ACD15, ACD21, and SLN regulate accumulation and mobility of MBD6 to silence genes and transposable elements
Open the record for dataset details and reuse information.
Data from: TERAD: Extraction of transposable element composition from RADseq data
Open the record for dataset details and reuse information.
Data from: High-throughput sequencing of transposable element insertions suggests adaptive evolution of the invasive Asian Tiger Mosquito towards temperate environments
Open the record for dataset details and reuse information.
The sunflower (Helianthus annuusL.) genome reflects a recent history of biased accumulation of transposable elements
Open the record for dataset details and reuse information.
Data from: Chironomus riparius (Diptera) genome sequencing reveals the impact of minisatellite transposable elements on population divergence
Open the record for dataset details and reuse information.
Data from: Recent and dynamic transposable elements contribute to genomic divergence under asexuality
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.