Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

386

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

386 results for “transposons”

Learn how ShareScore rates datasets ↗
zenodo48/100

Enhanced Biosafety of the Sleeping Beauty Transposon System by Using mRNA as Source of Transposase to Efficiently and Stably Transfect Retinal Pigment Epithelial Cells

<p>Raw data of the publication &quot;Enhanced Biosafety of the Sleeping Beauty Transposon System by Using mRNA as Source of Transposase to Efficiently and Stably Transfect Retinal Pigment Epithelial Cells&quot;.</p> <p>Abstract:&nbsp; Neovascular age-related macular degeneration (nvAMD) is characterized by choroidal<br> neovascularization (CNV), which leads to retinal pigment epithelial (RPE) cell and photoreceptor<br> degeneration and blindness if untreated. Since blood vessel growth is mediated by endothelial cell<br> growth factors, including vascular endothelial growth factor (VEGF), treatment consists of repeated,<br> often monthly, intravitreal injections of anti-angiogenic biopharmaceuticals. Frequent injections are<br> costly and present logistic difficulties; therefore, our laboratories are developing a cell-based gene<br> therapy based on autologous RPE cells transfected ex vivo with the pigment epithelium derived factor<br> (PEDF), which is the most potent natural antagonist of VEGF. Gene delivery and long-term expression<br> of the transgene are enabled by the use of the non-viral Sleeping Beauty (SB100X) transposon system<br> that is introduced into the cells by electroporation. The transposase may have a cytotoxic effect and a<br> low risk of remobilization of the transposon if supplied in the form of DNA. Here, we investigated<br> the use of the SB100X transposase delivered as mRNA and showed that ARPE-19 cells as well as<br> primary human RPE cells were successfully transfected with the Venus or the PEDF gene, followed<br> by stable transgene expression. In human RPE cells, secretion of recombinant PEDF could be detected<br> in cell culture up to one year. Non-viral ex vivo transfection using SB100X-mRNA in combination<br> with electroporation increases the biosafety of our gene therapeutic approach to treat nvAMD while<br> ensuring high transfection efficiency and long-term transgene expression in RPE cells.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Developmental timing of programmed DNA elimination in Paramecium tetraurelia recapitulates germline transposon evolutionary dynamics

<p>With its nuclear dualism, the ciliate <em>Paramecium</em> constitutes an original model to study how host genomes cope with transposable elements (TEs). <em>P. tetraurelia</em> harbors two germline micronuclei (MIC) and a polyploid somatic macronucleus (MAC) that develops from the MIC at each sexual cycle. Throughout evolution, the MIC genome has been continuously colonized by TEs and related sequences that are removed from the somatic genome during MAC development. Whereas TE elimination is generally imprecise, excision of ~45000 TE-derived Internal Eliminated Sequences (IESs) is precise, allowing for functional gene assembly. Programmed DNA elimination is concomitant with genome amplification. It is guided by non-coding RNAs and repressive chromatin marks. A subset of IESs are excised independently of this epigenetic control, raising the question of how they are targeted for elimination. To gain insight into the determinants of IES excision, we determined the developmental timing of DNA elimination genome-wide by combining fluorescence-assisted nuclear sorting with next-generation sequencing. Essentially all IESs are excised within one endoduplication round only (32C to 64C), while TEs are eliminated at a later stage. We show that time, rather than replication, controls the progression of DNA elimination. Further analyses defined four IES classes according to excision timing and revealed that the earliest excised IESs tend to be independent of epigenetic factors, display strong sequence signals at their ends and originate from the most ancient integration events. We conclude that old IESs have been optimized during evolution for early and accurate excision, by acquiring stronger sequence determinants and escaping epigenetic control.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Dataset for the genome of medicinal plant Sophora flavescens has undergone significant expansion of both transposons and genes

<p><em>Sophora flavescens</em> is a medicinal plant in the genus Sophora of the Fabaceae family. The root of <em>S. flavescens</em> is known in China as Kushen and has a long history of wide use in multiple formulations of Traditional Chinese Medicine (TCM). However, there is little genomic information available for <em>S. flavescens</em>, which has greatly hindered the breeding of <em>S. flavescens</em> and characterisation of bioactive compounds. Therefore, in this study, we used third-generation Nanopore long-read sequencing technology combined with Hi-C scaffolding technology to <em>de novo</em> assemble the <em>S. flavescens</em> genome. We obtained a chromosomal level high-quality <em>S. flavescens</em> draft genome. The draft genome size is approximately 2.08 Gb, with more than 80% annotated as Transposable Elements (TEs). We also annotated 60,485 genes and examined their expression profiles in leaf, stem and root tissues. We also characterised the genes and pathways involved in the biosynthesis of major bioactive compounds, including alkaloids, flavonoids and isoflavonoids. The assembled genome provides valuable resources for conservation, genetic research and breeding of <em>S. flavescens</em>.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Sample datasets for Transposon insertion sequencing analysis tutorial

<p>The dataset contains five files:</p> <ol> <li>Tnseq-Tutorial-reads.fastqsanger.gz&nbsp; - A subset of TnSeq reads published in&nbsp; `Santiago, M., Matano, L. M., Moussa, S. H., Gilmore, M. S., Walker, S., &amp; Meredith, T. C. (2015). A new platform for ultra-high density Staphylococcus aureus transposon libraries. <em>BMC Genomics</em>, <em>16</em>(1), 1&ndash;18. http://doi.org/10.1186/s12864-015-1361-3`</li> <li>condition_barcodes.fasta&nbsp;-&nbsp; Set of barcodes to separate reads from different experimental conditions</li> <li>construct_barcodes.fasta&nbsp; - Set of barcodes to separate reads from different transposon constructs</li> <li>staph_aur.fasta&nbsp;: Genome&nbsp;file for <em>Staphylococcus aureus&nbsp;</em></li> <li>staph_aur.fasta&nbsp;: Annotation file for <em>Staphylococcus aureus&nbsp;</em></li> </ol>

opencc-by-4.0Feb 2019View details →
dryad40/100

Supplementary data for: Transposon mutagenesis identifies cooperating genetic drivers during keratinocyte transformation and cutaneous squamous cell carcinoma progression

<p><strong>Supplementary Note 1:</strong></p> <ul> <li>S1 Text: Oncogenomic comparisons between SB candidate Trunk driver genes and their direct orthologs in human Cancer Gene Census; Pyrosequencing analysis of SB-driven keratinocyte cancer models; References.</li> </ul> <p><strong>Supplementary Figures 1-11:</strong></p> <ul> <li>S1 Fig: Overview of genetic crosses to generate SB|Trp53|Onc3 mouse model.</li> <li>S2 Fig: SB insertion patterns in activated and inactivated drivers.</li> <li>S3 Fig: Evaluating the reproducibility of SBCapSeq results from bulk cuSCC and normal skin specimens.</li> <li>S4 Fig. Hierarchical two-dimensional clustering of recurrent events in cuKA and cuSCC.</li> <li>S5 Fig. Curated biological pathways and processes enriched within SB-induced cuSCC.</li> <li>S7 Fig: ZMIZ1 metagene within the TCGA Head &amp; Neck Squamous Cell Carcinoma (hnSCC) RNA-seq dataset.</li> <li>S8 Fig: Clonally selected SB insertions affect trunk driver proto-oncogene expression in SB-cuSCC genomes.</li> <li>S9 Fig: Clonally selected SB insertions affect trunk driver genes by inactivating expression in SB-cuSCC genomes.</li> <li>S10 Fig: CREBBP knockdown does not alter proliferation rate in cuSCC cell lines.</li> <li>S11 Fig: Gross photographs of cuSCC xenograft masses collected at necropsy showing robust TurboGFP expression.</li> <li>S12 Fig: SB T2/Onc3 TG.12740 allele donor position mapping and exclusion for SB Driver Analysis.</li> </ul> <p><strong>Supplementary Tables 1-20:</strong></p> <ul> <li>S1 Table: Tumor incidence and subgroup classifications by cohort.</li> <li>S2 Table: Specimen metafile data for projects sequenced using SBCapSeq protocol with Ion Torrent Proton sequencer.</li> <li>S3 Table: Discovery and progression SB Driver Analysis for cuSCC60_SBC.</li> <li>S4 Table: Trunk SB Driver Analysis for cuSCC60_SBC.</li> <li>S5 Table: Discovery and progression SB Driver Analysis for cuKA11_SBC.</li> <li>S6 Table: Trunk SB Driver Analysis for cuKA11_SBC.</li> <li>S7 Table: Discovery and progression SB Driver Analysis for cuSK32_SBC.</li> <li>S8 Table: SBCapSeq read depth and analysis for 4 cuSCC genomes selected for multi-region resequencing because they had intermixing of cuSCC and cuKA histologies.</li> <li>S9 Table: Enrichr gene set pathway enrichment analysis of cuSCC drivers.</li> <li>S10 Table: Summary of 7 cuSCC transcriptomes selected for whole transcriptome RNAseq analysis.</li> <li>S11 Table: BED file of SBfusion insertions in 7 cuSCC genomes by whole transcriptome RNAseq analysis.</li> <li>S12 Table: Venn diagram for overlap of genes with SBfusion reads detected by whole transcriptome RNAseq analysis and cuSCC60_SBC discovery driver.</li> <li>S13 Table: Venn diagram for overlap of genes with SBfusion reads detected by whole transcriptome RNAseq analysis and all cuSCC drivers.</li> <li>S14 Table: Transcripts per million (TPM) normalized whole transcriptome RNAseq values per gene from RNA isolated from cuSCC genomes with and without Zmiz1 insertions.</li> <li>S15 Table: Fragments Per Kilobase of Transcripts per Million (FPKM) normalized whole transcriptome RNAseq values per gene transcript from RNA isolated from cuSCC genomes with and without Zmiz1 insertions.</li> <li>S16 Table: Normalized microarray values per gene from RNA isolated from cuSCC genomes with and without <em>Zmiz1</em> insertions.</li> <li>S17 Table: Normalized microarray values per probe from RNA isolated from cuSCC genomes with and without <em>Zmiz1</em> insertions.</li> <li>S18 Table: All 289 genes with differential expression analysis from microarray data from RNA isolated from cuSCC genomes with and without Zmiz1 insertions with P&lt;0.0001 and q&lt;0.05.</li> <li>S19 Table: Lentiviral vectors containing shRNAs used in this study.</li> <li>S20 Table: TaqMan probes used in this study.</li> </ul> <p><strong>Supplementary Datasets 1-5:</strong></p> <ul> <li>S1 Data: BED file of SB insertions for cuSCC60_SBC.</li> <li>S2 Data: BED file of SB insertions for cuKA11_SBC.</li> <li>S3 Data: BED file of SB insertions for cuSK32_SBC</li> <li>S4 Data: BED file of SB insertions for 4 cuSCC genomes selected for multi-region resequencing because they had intermixing of cuSCC and cuKA histologies.</li> <li>S5 Data: Numerical data for graphs pertaining to Figure Panels Fig1A; Fig5A–E; Fig6A–B,D; Fig7C–G; Fig8A–B,D–F; Fig9A–I in the paper on the publicly availble <em>PLOS Genetics</em> Web site.</li> </ul>

opencc-zeroAug 2021View details →
dryad40/100

Supplementary data for: Transposon mutagenesis identifies cooperating genetic drivers during keratinocyte transformation and cutaneous squamous cell carcinoma progression

Open the record for dataset details and reuse information.

publicAug 2021View details →
dryad36/100

The Enterprise, a massive transposon carrying Spok meiotic drive genes

<p>The genomes of eukaryotes are full of parasitic sequences known as transposable elements (TEs). Most TEs studied to date are relatively small (50 – 12000 bp), but can contribute to very large proportions of genomes. Here we report the discovery of a putative giant tyrosine-recombinase-mobilized DNA transposon, <em>Enterprise</em>, from the model fungus <em>Podospora anserina</em>. Previously, we described a large genomic feature called the <em>Spok</em> block which is notable due to the presence of meiotic drive genes of the <em>Spok</em> gene family. The <em>Spok</em> block ranges from 110 kb to 247 kb and can be present in at least four different genomic locations within <em>P. anserina</em>, despite what is an otherwise highly conserved genome structure. We propose that the reason for its varying positions is that the <em>Spok</em> block is not only capable of meiotic drive, but is also capable of transposition. More precisely, the <em>Spok</em> block represents a unique case where the <em>Enterprise</em> has captured the <em>Spoks</em>, thereby parasitizing a resident genomic parasite to become a genomic hyperparasite. Furthermore, we demonstrate that <em>Enterprise</em> (without the <em>Spoks</em>) is found in other fungal lineages, where it can be as large as 70 kb. Lastly, we provide experimental evidence that the Spok block is deleterious, with detrimental effects on spore production in strains which carry it. This union of meiotic drivers and a transposon has created a selfish element of impressive size in <em>Podospora</em>, challenging our perception of how TEs influence genome evolution and broadening the horizons in terms of what the upper limit of transposition may be.</p>

opencc-zeroDec 2020View details →
zenodo36/100

ProTInSeq: transposon insertion tracking by ultra-deep DNA sequencing applied to identify small and large translated ORFs

<p>ProTInSeq is a novel -omics technique designed to characterize proteomes by using DNA ultra-deep sequencing. The technique is based on transposons engineered to have a positive or negative protein selection marker expressed when the transposon is inserted in-frame into a protein-coding gene. In the genome-reduced bacterium Mycoplasma pneumoniae, ProTInSeq identifies 80% of known expressed proteins, as well as 5 new open reading frames (ORFs; &gt;100 amino acids); and 153 novel small ORF-encoded proteins (SEPs; &le;100 aa) that represent up to 18% of this bacterium&rsquo;s proteome. ProTInSeq can be used to detect translational noise, for protein quantification and to provide insight into functional protein aspects such as relative half-life, stability, and membrane topology. Herein, we describe a methodology that can be easily implemented in any living system and allows the deep understanding of proteomes and more importantly the identification of small proteins by DNA ultra-sequencing.</p> <p>We include the following files:</p> <p>- processed_inscalling.zip: output obtain after running FASTQINS transposon calling tool over the datasets found at&nbsp;<a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee</a>. This include every genome position in <em>M. pneumoniae</em>, the number of times an insertion has been mapped to that position and the total read count value.</p> <p>- separated_library_metrics.zip: insertion and read count processed from&nbsp;processed_inscalling files associated to every ORF&nbsp; and intergenic region in&nbsp;<em>M.&nbsp;penumoniae. </em>Columns include frame measured (0 - whole gene, 1 - in-frame, 2 and 3 for following positions) and metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant).</p> <p>- allmetrics.xlsx: merged table with the combination of results from separated_library_metrics.zip tab-delimited files.</p> <p>- selective_metrics_allannotations.xlsx:&nbsp; this table includes all the available information about the 30,112 sequences that could encode for a coding sequence in <em>M. pneumoniae</em>. For each identifier (column B), we include coordinates information and nucleotide and amino acid length information (columns C-H). Column I includes the gene name when the entry is found annotated in <em>M. pneumoniae</em>. Localization and function are described in columns J and K. Column L includes the operon number in which the annotation would be expressed. We also included transcription-related information average expression (column M; as log2(gene read count/gene length) and estimated average RNA copies per cell (column N) considering 4 RNA sequencing samples covering different growth times (6, 24 and 48 hours, ArrayExpress identifier E‐MTAB‐6203). Column O accounts for the number of mass spectrometry experiments detecting that entry (to a maximum of 116) and column P accounts for the total number of unique tryptic peptides detected. This comes <a href="https://paperpile.com/c/BImj5N/eljz">[5]</a>, available for 12,426 sequences that present an amino acid length &ge;19 (from 116 mass spectrometry experiments, ID PRIDE: PXD008243). Columns Q to T recapitulate protein copies per cell under different conditions (overall, extracting with urea, extracting with SDS and mean, respectively). Column U includes half-lives of the proteins. Columns V and W describe the reference density of insertion and essentiality assigned in previous studies. Column X-AA includes the predicted RanSEPs score, ribosome binding site presence, homology into seven groups: 0&mdash;no hits passed the thresholds defined; 1&mdash;conserved with an annotated function; 2&mdash;conserved as an annotated SEP in NCBI but no associated function; 3&mdash;conserved in a different species but target and homologous sequence not found in NCBI; 4&mdash;sequence is completely or partially (&gt; 75%) repeated &ge; 3 times in the reference genome; 5&mdash;potential pseudogene; and 6&mdash;to depict those annotations that are found in the reference NCBI annotation file. , and function expected by homology, respectively. Columns AB to AD cover the output provided by Phobius, including the number of transmembrane segments, presence of signal peptide and transmembrane topology predicted by TM-HMM. Column AE includes the complex information where 1 implies that entry is functional as a monomer, 2 as dimer, and so on. Finally, columns AF-AH will be 1 if the protein is a Lon protease target, a lipoprotein, and/or a truncated gene or pseudogene, respectively, 0 otherwise.&nbsp;Following columns include for every sample presenting selective insertion rates in-frame using the following identifiers separated by underscores: marker (BarnB, Cm or Ery), type (control-AC or selection-BD, antibiotic concentration, sample replicate, frame measured, metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant). Last columns combine the number of samples each annotation has been identified. Notice for barnase library the results need to be interpreted considering it is a negative selection marker inverting the 0 and 1 meaning.</p>

openOct 2023View details →
dryad36/100

Data from: An intronic transposon insertion associates with a trans-species color polymorphism in Midas cichlid fishes

<p><span><span><span><span><span><span><span><span><span><span><span>Polymorphisms have fascinated biologists for a long time, but their genetic underpinnings often remained elusive. Here, we aimed to uncover the genetic basis of the gold/dark polymorphism that is eponymous of Midas cichlid fish (<i>Amphilophus </i>spp.) adaptive radiations in Nicaraguan crater lakes. While most Midas cichlids are of the melanic "dark morph", about 10% of individuals lose their melanic pigmentation during their ontogeny and transition into a conspicuous "gold morph". Using a new haplotype-resolved long-read assembly we discovered an 8.2kb, transposon-derived inverted repeat in an intron of an undescribed gene, which we term <i>goldentouch</i> in reference to the Greek myth of King Midas. The gene <i>goldentouch</i> is differentially expressed between morphs, likely due to structural implications of inverted repeats in both DNA and RNA (cruciform and hairpin formation). The near-perfect association with the phenotype across several independent populations suggests that this insertion likely underlies this trans-specific, stable polymorphism.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroNov 2021View details →
zenodo36/100

Horizontal transfer and subsequent explosive expansion of a DNA transposon in sea kraits (Laticauda)

<p><strong>Abstract</strong></p> <p>Transposable elements (TEs) are self replicating genetic sequences and are often described as important &ldquo;drivers of evolution&rdquo;. This driving force is because TEs promote genomic novelty by enabling rearrangement, and through exaptation as coding and regulatory elements. However, most TE insertions will be neutral or harmful, therefore host genomes have evolved machinery to supress TE expansion. Through horizontal transposon transfer (HTT) TEs can colonise new genomes, and since new hosts may not be able to shut them down, these TEs may proliferate rapidly. Here we describe HTT of the&nbsp;<em>Harbinger-Snek</em>&nbsp;DNA transposon into sea kraits (<em>Laticauda</em>), and its subsequent explosive expansion within&nbsp;<em>Laticauda</em>&nbsp;genomes. This HTT occurred following the divergence of&nbsp;<em>Laticauda</em>&nbsp;from terrestrial Australian elapids ~15-25 Mya. This has resulted in numerous insertions into introns and regulatory regions, with some insertions into exons which appear to have altered UTRs or added sequence to coding exons.&nbsp;<em>Harbinger-Snek</em>&nbsp;has rapidly expanded to make up 8-12% of&nbsp;<em>Laticauda</em>&nbsp;spp. genomes; this is the fastest known expansion of TEs in amniotes following HTT. Genomic changes caused by this rapid expansion may have contributed to adaptation to the amphibious-marine habitat.</p> <p><strong>Dataset</strong></p> <p>The deposited dataset contains scripts used in analysis, GFFs of the <em>Laticauda&nbsp;</em>genome gene annotations produced using Liftoff, repeat sequences of all <em>Harbinger-Snek variants and&nbsp;Harbinger-Snek</em>-like TEs, repeat library used in RepeatMasker repeat annotation, repeat annotation of&nbsp;<em>Laticauda, Notechis</em> and <em>Pseudonaja</em>&nbsp;genomes, screenshots of IGV showing RNASeq reads mapped to gene&nbsp;exons and UTRs containing <em>Harbinger-Snek</em>&nbsp;insertions, and all phylogenetic trees and the sequence data used in generating them.</p>

opencc-by-4.0Jun 2021View details →
dryad36/100

Data from: Critical role of insertion preference for invasion trajectory of transposons

<p>It is unclear how mobile DNA sequences (transposable elements, hereafter TEs) invade eukaryotic genomes and reach stable copy numbers, as transposition can decrease host fitness. This challenge is particularly stark early in the invasion of a TE family at which point hosts may lack the specialized machinery to repress the spread of these TEs. One possibility (in addition to the evolution of host regulation of TEs), is that TE families may evolve to preferentially insert into chromosomal regions that are less likely to impact host fitness. This may allow the mean TE copy number to grow while minimizing the risk for host population extinction. To test this, we constructed simulations to explore how the transposition probability and insertion preference of a TE family influence the evolution of mean TE copy number and host population size, allowing for extinction. We find that the effect of a TE family's insertion preference depends on a host's ability to regulate this TE family. Without host repression, a neutral insertion preference increases the frequency of and decreases the time to population extinction. With host repression, a preference for neutral insertions minimizes the cumulative deleterious load, increases population fitness, and, ultimately, avoids triggering an extinction vortex.</p>

opencc-zeroJul 2023View details →
dryad36/100

The Enterprise, a massive transposon carrying Spok meiotic drive genes

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad36/100

Dual randomly barcoded transposon sequencing (Dual Tn-seq) data for <em>Streptococcus pneumoniae</em> D39

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad36/100

Data from: A platform supporting generation and isolation of random transposon mutants in Chlamydia trachomatis

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Data from: Critical role of insertion preference for invasion trajectory of transposons

Open the record for dataset details and reuse information.

publicJul 2023View details →
dryad36/100

Data from: An intronic transposon insertion associates with a trans-species color polymorphism in Midas cichlid fishes

Open the record for dataset details and reuse information.

publicNov 2021View details →
zenodo32/100

Accurate SNV detection in single cells by transposon-based whole-genome amplification of complementary strands

<p>Common SNPs from gnomAD</p>

opencc-by-4.0Feb 2021View details →
dryad32/100

Data from: Longevity and transposon defense, the case of termite reproductives

Social insects are promising new models in aging research. Within single colonies, longevity differences of several magnitudes exist that can be found elsewhere only between different species. Reproducing queens (and, in termites, also kings) can live for several decades, whereas sterile workers often have a lifespan of a few weeks only. We studied aging in the wild in a highly social insect, the termite Macrotermes bellicosus, which has one of the most pronounced longevity differences between reproductives and workers. We show that gene-expression patterns differed little between young and old reproductives, implying negligible aging. By contrast, old major workers had many genes up-regulated that are related to transposable elements (TEs), which can cause aging. Strikingly, genes from the PIWI-interacting RNA (piRNA) pathway, which are generally known to silence TEs in the germline of multicellular animals, were down-regulated only in old major workers but not in reproductives. Continued up-regulation of the piRNA defense commonly found in the germline of animals can explain the long life of termite reproductives, implying somatic cooption of germline defense during social evolution. This presents a striking germline/soma analogy as envisioned by the superorganism concept: the reproductives and workers of a colony reflect the germline and soma of multicellular animals, respectively. Our results provide support for the disposable soma theory of aging.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genome-wide patterns of transposon proliferation in an evolutionary young hybrid fish

Hybridization can induce transposons to jump into new genomic positions, which may result in their accumulation across the genome. Alternatively, transposon copy numbers may increase through non-allelic (ectopic) homologous recombination in highly repetitive regions of the genome. The relative contribution of transposition bursts versus recombination-based mechanisms to evolutionary processes remains unclear because studies on transposon dynamics in natural systems are rare. We assessed the genome-wide distribution of transposon insertions in a young hybrid lineage ("invasive Cottus", n=11) and its parental species Cottus rhenanus (n=17) and Cottus perifretum (n=9) using a reference genome assembled from long single molecule PacBio reads. An inventory of transposable elements was reconstructed from the same data and annotated. Transposon copy numbers in the hybrid lineage increased in 120 (15.9%) out of 757 transposons studied here. The copy number increased on average by 69% (range: 10 – 197 %). Given the age of the hybrid lineage, this suggests that they have proliferated within a few hundred generations since admixture began. However, frequency spectra of transposon insertions revealed no increase of novel and rare insertions across assembled parts of the genome. This implies that transposons were added to repetitive regions of the genome that remain difficult to assemble. Future studies will need to evaluate whether recombination-based mechanisms rather than genome-wide transposition may explain the majority of the recent transposon proliferation in the hybrid lineage. Irrespectively of the underlying mechanism, the observed over-abundance in repetitive parts of the genome suggests that gene-rich regions are unlikely to be directly affected.

opencc-zeroDec 2017View details →
dryad32/100

Duck pan-genome reveals two transposon-derived structural variations caused bodyweight enlarging and white plumage phenotype formation during evolution

<p><span>Structural variations (SVs) are a major source of domestication and improvement traits. We present the first duck pan-genome constructed using five genome assemblies capturing ~40.98 Mb new sequences. This pan-genome together with high-depth sequencing data (&gt;46.5X) identified 101,041 SVs, of which substantial proportions were derived from transposable element (TE) activity. Many TE-derived SVs anchored in a gene body or regulatory region are linked to domestication and improvement. By combining quantitative genetics with molecular experiments, we dissect how TE-derived SVs change gene expression of <em>IGF2BP1</em> and generate novel transcripts of <em>MITF</em>, shaping body weight and plumage color. In the <em>IGF2BP1</em> locus, the TE-derived SV explains the largest effect on body weight among avian species (27.61% of phenotypic variation). Our findings highlight the </span><span>importance of using a pan-genome as a reference in genomics studies</span><span> and explore the roles of TE-derived SVs in trait formation and in livestock breeding.</span></p>

opencc-zeroNov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record