Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

380

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

380 results for “Transposable elements”

Learn how ShareScore rates datasets ↗
dryad36/100

Wide spectrum and high frequency of genomic structural variation, including transposable elements, in large double stranded DNA viruses

Our knowledge of the diversity and frequency of genomic structural variation segregating in populations of large double stranded (ds) DNA viruses is limited. Here we sequenced the genome of a baculovirus (AcMNPV) purified from beet armyworm (Spodoptera exigua) larvae at depths >195,000X using both short-read (Illumina) and long-read (PacBio) technologies. Using a pipeline relying on hierarchical clustering of structural variants (SVs) detected in individual short- and long-reads by six variant callers, we identified a total of 1,141 SVs in AcMNPV, including 464 deletions, 443 inversions, 160 duplications and 74 insertions. These variants are considered robust and unlikely to result from technical artifacts because they were independently detected in at least three long reads as well as at least three short reads. SVs are distributed along the entire AcMNPV genome and may involve large genomic regions (30,496 bp on average). We show that no less than 39.9% of genomes carry at least one SV in AcMNPV populations, that the vast majority of SVs (75%) segregate at very low frequency (<0.01%) and that very few SVs persist after 10 replication cycles, consistent with a negative impact of most SVs on AcMNPV fitness. Using short-read sequencing datasets, we then show that populations of two iridoviruses and one herpesvirus are also full of SVs, as they contain between 426 and 1102 SVs carried by 52.4 to 80.1% of genomes. Finally, AcMNPV long reads allowed us to identify 1,757 transposable elements (TEs) insertions, 895 of which are truncated and occur at one extremity of the reads. This further supports the role of baculoviruses as possible vectors of horizontal transfer of TEs. Altogether, we found that SVs, which evolve mostly under rapid dynamics of gain and loss in viral populations, represent an important feature in the biology of large dsDNA viruses.

opencc-zeroDec 2019View details →
dryad36/100

Data from: The fire ant social supergene is characterized by extensive gene and transposable element copy number variation

In the fire ant Solenopsis invicta, a supergene composed of ~600 genes and having two variants, SB and Sb, regulates colony social form. In single queen colonies all individuals carry only the SB allele, while in multiple queen colonies, some individuals carry the Sb allele. In this study we characterized genes with copy number variation between SB and Sb-carrying individuals. We showed extensive acquisition of gene duplicates in Sb genome, with some likely involved in polygyne-related phenotypes. We found 260 genes with differences in copy number between SB and Sb, of which 239 are in greater copy number in Sb. We observed TE accumulation on Sb, likely due to the accumulation of repetitive elements on the non-recombining chromosome. We found a weak correlation between TE copy number and differential expression, suggesting some TEs may still be proliferating in Sb, however many of the duplicated TEs were already silenced. Among the 115 non-TE genes with higher copy in Sb, enzymes responsible for cuticular hydrocarbon synthesis were highly represented. These include a desaturase and an elongase; both potentially responsible for differential queen odor and likely beneficial for polygyne ants. These genes seem to have translocated into the supergene from other chromosomes and proliferated by multiple duplication events. While the presence of transposable elements (TEs) in supergenes is well documented, little is known about duplication of non-TE genes and their possible adaptive role. Overall, our results suggest that gene duplications may be an important factor leading to monogyne and polygyne ant societies.

opencc-zeroJan 2020View details →
dryad36/100

Data from: high repeat content in the genomes of sparrows: the importance of genome assembly completeness for transposable element discovery

<p>Transposable elements (TE) play critical roles in shaping genome evolution. However, the highly repetitive sequence content of TEs is a major source of assembly gaps. This makes it difficult to decipher the impact of these elements on the dynamics of genome evolution. The increased capacity of long-read sequencing technologies to span highly repetitive regions of the genome should provide novel insights into patterns of TE diversity. Here we report the generation of highly contiguous reference genomes using PacBio long read and Omni-C technologies for three species of sparrows in the family Passerellidae. To assess the influence of sequencing technology on TE annotation, we compared these assemblies to three chromosome-level sparrow assemblies recently generated by the Vertebrate Genomes Project and nine other sparrow species generated using a variety of short- and long-read technologies. All long-read based assemblies were longer in length (range: 1.12-1.41 Gb) than short-read assemblies (0.91-1.08 Gb). Assembly length was strongly correlated with the amount of repeat content, with longer genomes showing much higher levels of repeat content than typically reported for the avian order Passeriformes. Repeat content for the Bell's sparrow (31.2% of genome) was the highest level reported to date for a songbird genome assembly and was more in line with woodpecker (order Piciformes) genomes. CR1 LINE elements retained from an expansion that occurred 25-30 million years ago were the most abundant TEs in the song sparrow genome. Although the other five sparrow species also exhibit evidence for a spike in CR1 LINE activity at 25-30 million years ago, LTR elements stemming from more recent expansions were the most abundant elements in these species. LTRs were uniquely abundant in the Bell's sparrow genome deriving from two recent peaks of activity. Higher levels of repeat content (79.2-93.7%) were found on the W chromosome relative to the Z (20.7-26.5) or autosomes (16.1-30.9%). These patterns support a dynamic model of transposable element expansion and contraction underpinning the seemingly constrained and small sized genomes of birds. Our work highlights how the resolution of difficult-to-assemble regions of the genome with new sequencing technologies promises to transform our understanding of avian genome evolution.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Transposable elements mark a repeat-rich region associated with migratory phenotypes of willow warblers (Phylloscopus trochilus).

<p>This upload contains the relevant data (alignments, trees, html TE landscape files, and datasheets) used in the manuscript: &quot;The importance of repeat-rich regions in the migratory phenotypes of the willow warbler (Phylloscopus trochilus): transposable elements as a marker.&quot;</p>

opencc-by-4.0Apr 2021View details →
zenodo36/100

Simulated data set of chimeric transposable elements.

<p>This dataset is composed of 9,000 sequences of transposable elements (TEs) of 20,000 bp. Three cases of transposable elements with artifacts at the ends, with another chimeric TE or with simple repeats were considered for the generation of the sequences. The sequences of the first case consist of a DNA fragment + first TE + DNA fragment + second TE + DNA fragment. The sequences of the second case consist of a first TE + second TE + repeat of the first TE. The third case sequences consist of a microsatellite that is repeated in tandem by placing the extracted TE at position 10,000 (in the middle), occupying both sides to the ends. All TEs used in this dataset were taken from Dfam, for the species Drosophila melanogaster.</p> <p>The identifier of each sequence has the information about the case, TE Dfam identifier, TE initial position inside the sequence, and the TE length, all separated by "_". For example:</p> <p>Caso1_DF000001548.2_6926_5126</p> <p>File description:</p> <p>dataset.zip: The fasta file containing all the sequences</p> <p>features_data.npy.zip: A numpy file containing the numerical representation of the four TE+Aid plots generated for the sequences presented in the dataset.zip file. This data is actually a numpy array with dimensions 9000x256x256x3x4</p> <p>labels_data.numpy.zip: A numpy file containing the starting and ending position (normalized between 0 and 1) of each TE presented in the dataset.zip file.</p> <p>The last two files were generated to train a neural network for trimming out automatically artifacts in DNA sequences.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

TransposonDB - database of transposable elements

<p>This record contains some of the supplementary files for the TransposonUltimate package by Riehl et al. Please see https://github.com/DerKevinRiehl/TransposonUltimate for code.</p> <pre>Folder &quot;Assemblies&quot; contains the 40 polished scaffolded assemblies of 20 C. elegans wild isolates, created by Cristian Riccio. 20 assemblies were scaffolded using the VC2010 reference. 20 assemblies were scaffolded using the CB4856 reference. We release the data under the usual Fort Lauderdale/Bermuda agreements. We will publish a paper about these genomes ourselves but we allow you to mine the dataset for a gene-specific analysis. Please contact us if you would like to publish anything using these data before our genome publication. Folder &quot;Classification&quot; contains the TransposonDB in FASTA file format, that was introduced in the context of the transposon classification module &quot;RFSB&quot;. Folder &quot;Annotation&quot; contains the transposon annotations of the C. elegans and O. Sativa ssp. Japonica genomes, that were generated using the &quot;reasonaTE&quot; transposon annotation module. Folder &quot;Detection&quot; contains the detected transposition events in the context of the transposon detection module &quot;deTEct&quot;. All 20 wild isolates assemblies of C. elegans were used as probe genomes, and compared to the two reference genomes &quot;VC2010&quot; and &quot;CB4856&quot;. All combinations were investigated using PBMM2 alignments + PBSV structural variants, and NGMLR alignments + Sniffles structural variants.</pre>

opencc-by-4.0Apr 2021View details →
zenodo36/100

Magnaporthe oryzae transposable elements manuscript additional datasets

<p>Additional datasets for manuscript on transposable elements in <em>Magnaporthe oryzae</em>.&nbsp;</p> <ul> <li>analysis_files.tar.gz - contains selected analysis files generated by scripts in this GitHub repository:&nbsp;<a href="https://github.com/annenakamoto/moryzae_tes">https://github.com/annenakamoto/moryzae_tes</a>. File names correspond to those in the scripts.</li> <li>FungGAP_out.tar.gz - contains gene prediction outputs for all genomes used</li> <li>OrthoFinder_out.tar.gz - contains output from OrthoFinder, run on proteomes of all genomes used</li> <li>tabular_data_for_figures.tar.gz - contains tabular data used to generate figures</li> <li>visualization_files.tar.gz - contains bed and seg files used to visualize all genes, SCOs, effectors, TEs, solo LTRs, and GC content in each representative genome</li> </ul>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Resolution of structural variation in diverse mouse genomes reveals chromatin remodeling due to transposable elements

<p>Structural variant calls, RepeatMasker annotations, and genome assemblies of diverse mouse genomes.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Transposable elements potentiate radiotherapy-induced cellular immune reactions via RIG-I-mediated virus-sensing pathways

<p>Mass spectrometry-based proteomics source data.</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Data from: Diversity and evolution of the transposable element repertoire in arthropods with particular reference to insects

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad36/100

Data from: The fire ant social supergene is characterized by extensive gene and transposable element copy number variation

Open the record for dataset details and reuse information.

publicJan 2020View details →
dryad36/100

Wide spectrum and high frequency of genomic structural variation, including transposable elements, in large double stranded DNA viruses

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad36/100

Data from: high repeat content in the genomes of sparrows: the importance of genome assembly completeness for transposable element discovery

Open the record for dataset details and reuse information.

publicDec 2023View details →
dryad36/100

Transposable element landscape in Drosophila populations selected for longevity

Open the record for dataset details and reuse information.

publicFeb 2021View details →
dryad36/100

DNA gains and losses in gigantic genomes do not track differences in transposable element-host silencing interactions

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad36/100

Transposable element accumulation drives size differences among polymorphic Y chromosomes in Drosophila

Open the record for dataset details and reuse information.

publicApr 2022View details →
zenodo32/100

Austropuccinia psidii, causing myrtle rust, has a gigabase-sized genome shaped by transposable elements

<p><em>Austropuccinia psidii</em>, originating in South America, is a globally invasive fungal plant pathogen that causes rust disease on Myrtaceae. Several biotypes are recognized, with the most widely distributed pandemic biotype spreading throughout the Asia-Pacific and Oceania regions over the last decade. <em>Austropuccinia</em><em> psidii</em> has a broad host range with more than 480 myrtaceous species. Since first detected in Australia in 2010, the pathogen has caused the near extinction of at least three species and negatively affected commercial production of several Myrtaceae. To enable molecular and evolutionary studies into <em>A. psidii</em> pathogenicity, we assembled a highly contiguous genome for the pandemic biotype. With an estimated haploid genome size of just over 1 Gb (gigabases), it is the largest assembled fungal genome to date. The genome has undergone massive expansion via distinct transposable element (TE) bursts. Over 90% of the genome is covered by TEs predominantly belonging to the Gypsy superfamily. These TE bursts have likely been followed by deamination events of methylated cytosines to silence the repetitive elements. This in turn led to the depletion of CpG sites in transposable elements and a very low overall GC content of 33.8%. The overall gene content is highly conserved, when compared to other closely related Pucciniales, yet the intergenic distances are increased by an order of magnitude indicating a general insertion of TEs between genes.&nbsp; Overall, we show how transposable elements shaped the genome evolution of <em>A. psidii</em> and provide a greatly needed resource for strategic approaches to combat disease spread. Please cite the authors if using this data:&nbsp;https://academic.oup.com/g3journal/article/11/3/jkaa015/6007476</p>

opencc-by-4.0Feb 2020View details →
dryad32/100

The sunflower (Helianthus annuusL.) genome reflects a recent history of biased accumulation of transposable elements

<p>Aside from polyploidy, transposable elements are the major drivers of genome size increases in plants. Thus, understanding the diversity and evolutionary dynamics of transposable elements in sunflower (<i>Helianthus annuus</i> L.), especially given its large genome size (∼3.5 Gb) and the well‐documented cases of amplification of certain transposons within the genus, is of considerable importance for understanding the evolutionary history of this emerging model species. By analyzing approximately 25% of the sunflower genome from random sequence reads and assembled bacterial artificial chromosome (BAC) clones, we show that it is composed of over 81% transposable elements, 77% of which are long terminal repeat (LTR) retrotransposons. Moreover, the LTR retrotransposon fraction in BAC clones harboring genes is disproportionately composed of chromodomain‐containing <i>Gypsy</i> LTR retrotransposons ('chromoviruses'), and the majority of the intact chromoviruses contain tandem chromodomain duplications. We show that there is a bias in the efficacy of homologous recombination in removing LTR retrotransposon DNA, thereby providing insight into the mechanisms associated with transposable element (TE) composition in the sunflower genome. We also show that the vast majority of observed LTR retrotransposon insertions have likely occurred since the origin of this species, providing further evidence that biased LTR retrotransposon activity has played a major role in shaping the chromatin and DNA landscape of the sunflower genome. Although our findings on LTR retrotransposon age and structure could be influenced by the selection of the BAC clones analyzed, a global analysis of random sequence reads indicates that the evolutionary patterns described herein apply to the sunflower genome as a whole.</p>

opencc-zeroDec 2019View details →
zenodo32/100

Dynamics of transposable elements in recently diverged fungal pathogens: lineage-specific transposable element content and efficiency of genome defenses

<p>Transposable elements (TEs) impact genome plasticity, architecture and evolution in fungal plant pathogens. The wide range of TE content observed in fungal genomes reflects diverse efficacy of host-genome defence mechanisms that can counter-balance TE expansion and spread. Closely related species can harbour drastically different TE repertoires. The evolution of fungal effectors, which are crucial determinants of pathogenicity, has been linked to the activity of TEs in pathogen genomes. Here we describe how TEs have shaped genome evolution of the fungal wheat pathogen <em>Zymoseptoria tritici</em> and four closely related species. We compared <em>de novo</em> TE annotations and Repeat-Induced Point mutation signatures in twenty-six genomes from the <em>Zymoseptoria</em> species-complex. Then, we assessed the relative insertion ages of TEs using a comparative genomics approach. Finally, we explored the impact of TE insertions on genome architecture and plasticity. The twenty-six genomes of <em>Zymoseptoria</em> species reflect different TE dynamics with a majority of recent insertions. TEs associate with accessory genome compartments, with chromosomal rearrangements, with gene presence/absence variation and with effectors in all <em>Zymoseptoria </em>species. We find that the extent of RIP-like signatures varies among <em>Z. tritici</em> genomes compared to genomes of the sister species. The detection of a reduction of RIP-like signatures and TE recent insertions in <em>Z. tritici</em> reflects ongoing but still moderate TE mobility.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
dryad32/100

Data from: Small RNAs from a big genome: the piRNA pathway and transposable elements in the salamander species Desmognathus fuscus

Most of the largest vertebrate genomes are found in salamanders, a clade of amphibians that includes 686 species. Salamander genomes range in size from 14 to 120 Gb, reflecting the accumulation of large numbers of transposable element (TE) sequences from all three TE classes. Although DNA loss rates are slow in salamanders relative to other vertebrates, high levels of TE insertion are also likely required to explain such high TE loads. Across the Tree of Life, novel TE insertions are suppressed by several pathways involving small RNA molecules. In most known animals, TE activity in the germline is primarily regulated by the Piwi-interacting RNA (piRNA) pathway. In this study, we test the hypothesis that salamanders' unusually high TE loads reflect the loss of the ancestral piRNA-mediated TE-silencing machinery. We characterized the small RNA pool in the female and male adult gonads, testing for the presence of small RNA molecules that bear the characteristics of TE-targeting piRNAs. We also analyzed the amino acid sequences of piRNA pathway proteins from salamanders and other vertebrates, testing whether the overall patterns of sequence divergence are consistent with conserved pathway function across the vertebrate clade. Our results do not support the hypothesis of piRNA pathway loss; instead, they suggest that the piRNA pathway is expressed in salamanders. Given these results, we propose hypotheses to explain how the extraordinary TE loads in salamander genomes could have accumulated, despite the expression of TE-silencing machinery.

opencc-zeroDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record