Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

49

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

49 results for “pseudogenization”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data and code for 'Pseudogenes act as a neutral reference for detecting selection in prokaryotic pangenomes'

<p>This repository contains the code and files for reproducing the analyses and results reported in 'Pseudogenes act as a neutral reference for detecting selection in prokaryotic pangenomes' by Gavin M. Douglas and B. Jesse Shapiro&nbsp;(<a href="https://doi.org/10.1038/s41559-023-02268-6">https://doi.org/10.1038/s41559-023-02268-6</a>).</p> <p>File organization and descriptions:</p> <ul> <li><strong>code/</strong> - Contains GitHub repository releases of code used in manuscript (the other folders contain datafiles only). This code is provided here as well as on GitHub to ensure long-term access. <ul> <li><strong>handy_pop_gen-1.1.0/</strong> - release v1.1.0 of the convenience repository (used for specific data processing and analysis steps referred to in the manuscript).</li> <li><strong>pangenome_pseudogene_null-1.1.0/&nbsp;</strong>- Main code repository for manuscript.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>broad_pangenome_analysis/</strong> <ul> <li><strong>element_info/element_counts.tsv.gz</strong> - Counts of (filtered) pseudogenes and intact genes called per genome accession.</li> <li><strong>element_info/gene_sizes.tsv.gz</strong> - Gene sizes in base-pairs.</li> <li><strong>element_info/pseudogene_sizes.tsv.gz</strong> - Filtered pseudogene sizes in base-pairs.</li> <li><strong>element_info/element_percent_coverage/*tsv.gz</strong> - Tables containing the percent genome coverage of genes and pseudogenes, by accession and averaged over accessions per species separately.</li> <li><strong>example_Mycoplasmopsis_bovis_panaroo_output.csv.gz</strong> - Panaroo output table for <em>Mycoplasmopsis bovis</em>, which was used for an example. Corresponds to the&nbsp;<em>gene_presence_absence.csv</em>&nbsp;file in the raw Panaroo output.</li> <li><strong>focal_and_non.focal_full_to_short.tsv.gz</strong> - Mapfile of full to short (and unique) species ids used in analysis. Primarily to include species ids in cluster names without making them unnecessarily long.</li> <li><strong>genome_info/accessions.tsv.gz</strong> - Genome accessions used for broad pangenome analysis (note that not all genome accessions could be downloaded [and were ignored], which is indicated in the "could_download" column).</li> <li><strong>genome_info/genome_sizes.tsv.gz</strong> - Sizes of all genomes used for the broad pangenome analysis.</li> <li><strong>metrics_additional_subsamples.tsv.gz</strong> - Contains columns also found in the <em>pangenome_and_related_metrics.tsv.gz</em>&nbsp;file below, but based on genome subsamplings of 3 and 20, rather than 9.</li> <li><strong>model_output/pangenome_linear_models.rds</strong> - R Data Serialization&nbsp;files containing the&nbsp;output of R linear model objects (generated by lm and provided as an R list object). There are separate elements in the list for the mean number of genes, genomic fluidity, percentage&nbsp;singletons (si), and si/sp.</li> <li><strong>model_output/linear_model_coef.tsv.gz</strong> - Coefficient summary table for all linear models.</li> <li><strong>pangenome_and_related_metrics.tsv.gz</strong> - Metrics used for broad pangenome analysis across 670 prokaryotic species. Note that this table was filtered down to 668 species after excluding those with &lt; 9 genomes.</li> <li><strong>pangenome_and_related_metrics_filt.tsv.gz</strong> - Filtered table, as described above.</li> <li><strong>taxonomy.tsv.gz</strong> - Taxonomy for all species used for this analysis, taken from GTDB. Row names are species names.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>indepth_10_species_analysis/</strong> <ul> <li><strong>cluster_breakdown_tables/</strong> - Folder containing tables providing breakdown of how clusters are distributed by element type, pangenome partition, and species. Provided for easy plotting.</li> <li><strong>cluster_COG_annot.tsv.gz</strong> - Mapping of cluster IDs to COG annotations.</li> <li><strong>cluster_filt_lengths_and_additional.tsv.gz</strong> - Metadata on clusters, most pertinently the length of the representative sequence in the cluster (which was used to filter out some clusters, below the cut-off which pseudogenes could not be called).</li> <li><strong>cluster_member_breakdown.tsv.gz </strong>- Table providing information on each element (called pseudogenes and intact genes) and provides information such as what cluster they are part of, what species and genome accession they are found in, etc.</li> <li><strong>cluster_types.rds</strong> - R Data Serialization file containing R list providing breakdown of all clusters into categories (intact/pseudogene/mixed, where mixed means containing both pseudogene and intact elements).</li> <li><strong>COG_enrichment_results/ultra.cloud-COG-gene-enrichments.tsv.gz</strong> - Output file with enrichment test summaries for COG IDs in significant COG categories, which was run for the ultra-cloud pangenome partition model only.</li> <li><strong>element_glmm_input.tsv.gz </strong>- Table containing all information used for fitting generalized linear mixed models.</li> <li><strong>focal_species.txt</strong> - Names of species used for the in-depth analysis.</li> <li><strong>genome_info/ </strong>- Folder containing the genome accessions (and the corresponding genome sizes) for all ten analyzed species.</li> <li><strong>glmm_output/</strong> - Folder containing R Data Serialization files containing output R objects after fitting generalized linear mixed models (only ultra-rare files are present, due to file size constraints).</li> <li><strong>per_genome_element.type_percent_coverages.rds</strong> - R Data Serialization&nbsp;file containing R list providing the percent coverage by intact genes vs pseudogenes per accession (nested by species)</li> </ul> </li> </ul>

opencc-by-4.0May 2023View details →
dryad36/100

Do pseudogenes pose a problem for metabarcoding marine animal communities?

<p>Because DNA metabarcoding typically employs sequence diversity among mitochondrial amplicons to estimate species composition, nuclear mitochondrial pseudogenes (NUMTs) can inflate diversity. This study quantifies the incidence and attributes of NUMTs derived from the 658 bp barcode region of cytochrome c oxidase I (COI) in 156 marine animal genomes. NUMTs were examined to ascertain if they could be recognized by their possession of indels or stop codons. In total, 309 NUMTs  150 bp were detected, with an average of 1.98 per species (range = 0–33) and a mean length of 391 bp  200 bp. Among this total, 75 (23.4%) lacked indels or stop codons. NUMTs appear to pose the greatest interpretational risk when short (&lt; 313 bp) amplicons are used, such as in eDNA studies, dietary analyses, or processed fish identification. Employing the standard amplicon length (313 bp) for marine metabarcoding, NUMTs could potentially inflate the OTU count by 21% above the true species count while also raising intraspecific variation at COI by 15%. However, when both amplicon length and position are considered, inflation in OTU counts and in barcode variation were just 9% and 10%, respectively, suggesting NUMTs will not seriously distort biodiversity assessments. There was a weak positive correlation between genome size and NUMT count but no variation among phyla or trophic groups. Until bioinformatic advances improve NUMT detection, the best defense involves targeting long amplicons and developing reference databases that include both mitochondrial sequences and their NUMT derivatives. </p>

opencc-zeroJun 2022View details →
zenodo36/100

Electrophysiology Data for "Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion"

<p>Here we provide&nbsp;the electrophysiology data for&nbsp;the manuscript &quot;Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion&quot;, published in Molecular Biology and Evolution (2021):&nbsp;msab271 (doi: 10.1093/molbev/msab271).&nbsp;Data are reported as current values in Excel format, sorted according to the appearance in Figures and supplemented by explanatory text on the procedures/data presentation.</p>

opencc-by-4.0Sep 2021View details →
dryad36/100

Supplementary information for: NUMT PARSER: Automated identification and removal of nuclear mitochondrial pseudogenes (numts) for accurate mitochondrial genome reconstruction in Panthera

<p>Nuclear mitochondrial pseudogenes (numts) may hinder the reconstruction of mtDNA genomes and affect the reliability of mtDNA datasets for phylogenetic and population genetic comparisons. Here, we present the program Numt Parser, which allows for the identification of DNA sequences that likely originate from numt pseudogene DNA. Sequencing reads are classified as originating from either numt or true cytoplasmic mitochondrial (cymt) DNA by direct comparison against cymt and numt reference sequences. Classified reads can then be parsed into cymt or numt datasets. We tested this program using whole genome shotgun-sequenced data from two ancient Cape lions (<em>Panthera</em> <em>leo</em>) because mtDNA is often the marker of choice for ancient DNA studies, and the genus <em>Panthera</em> is known to have numt pseudogenes. Numt Parser decreased sequence disagreements that were likely due to numt pseudogene contamination and equalized read coverage across the mitogenome by removing reads that likely originated from numts. We compared the efficacy of Numt Parser to two other bioinformatic approaches that can be used to account for numt contamination. We found that Numt Parser outperformed approaches that rely only on read alignment or Basic Local Alignment Search Tool (BLAST) properties, and was effective at identifying sequences that likely originated from numts while having minimal impacts on the recovery of cymt reads. Numt Parser therefore improves the reconstruction of true mitogenomes, allowing for more accurate and robust biological inferences.</p>

opencc-zeroDec 2022View details →
dryad36/100

Do pseudogenes pose a problem for metabarcoding marine animal communities?

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad36/100

Supplementary information for: NUMT PARSER: Automated identification and removal of nuclear mitochondrial pseudogenes (numts) for accurate mitochondrial genome reconstruction in Panthera

Open the record for dataset details and reuse information.

publicDec 2022View details →
dryad32/100

Datasets from: Validated removal of nuclear pseudogenes and sequencing artefacts from mitochondrial metabarcode

<p>Metabarcoding of Metazoa using mitochondrial genes may be confounded by both the accumulation of PCR and sequencing artefacts and the co-amplification of nuclear mitochondrial pseudogenes (NUMTs). The application of read abundance thresholds and denoising methods is efficient in reducing noise accompanying authentic mitochondrial amplicon sequence variants (ASVs). However, these procedures do not fully account for the complex nature of concomitant sequences and the highly variable DNA contribution of specimens in a metabarcoding sample. We propose, as a complement to denoising, the metabarcoding Multidimensional Abundance Threshold Evaluation (<i>metaMATE</i>) framework, a novel approach that allows comprehensive examination of multiple dimensions of abundance filtering and the evaluation of the prevalence of unwanted concomitant sequences in denoised metabarcoding datasets. <i>metaMATE</i> requires a denoised set of ASVs as input, and designates a subset of ASVs as being either authentic (mtDNA haplotypes) or non-authentic ASVs (NUMTs and erroneous sequences) by comparison to external reference data and by analysing nucleotide substitution patterns. <i>metaMATE</i> (i) facilitates the application of read abundance filtering strategies, which are structured with regard to sequence library and phylogeny and applied for a range of increasing abundance threshold values, and (ii) evaluates their performance by quantifying the prevalence of non-authentic ASVs and the collateral effects on the removal of authentic ASVs. The output from <i>metaMATE</i> facilitates decision-making about required filtering stringency and can be used to improve the reliability of intraspecific genetic information derived from metabarcode data. The framework is implemented in the <i>metaMATE</i> software, available at https://github.com/tjcreedy/metamate).</p>

opencc-zeroJan 2021View details →
dryad32/100

Data from: Development of a Chinook salmon sex identification SNP assay based on the growth hormone pseudogene

Genotypic sex identification assays can provide valuable information about fish populations when phenotypic sex determination is difficult. Here we describe the development of a TaqMan® assay (Ots_SexID) designed to identify the genotypic sex of Winter-Run Chinook salmon collected from the Sacramento River and spawned at the Livingston Stone National Fish Hatchery. The TaqMan® assay targets a region previously examined in the growth hormone pseudogene. Accuracy of the marker was assessed by comparing genotypic sex assignments for Chinook salmon spawned at Livingston Stone National Fish hatchery in 2012 (n = 84) to phenotypic sex recorded during spawning. Genotypic sex was observed to be concordant with phenotypic sex identified using Ots_SexID in 83/84 individuals, suggesting that the assay could be used to predict phenotypic sex with ~99% accuracy. To evaluate the utility of the TaqMan® assay in other parts of the species' range, we examined collections from 29 other populations ranging from Alaska to California. Sex assignments based on the assay were generally concordant with observed phenotypes, but there were some strong exceptions. These results suggest that the new assay will be very useful in Sacramento River Winter-Run Chinook salmon, but also highlight the importance of thoroughly testing any sex identification assay prior to application in a population of interest.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Trpc2 Pseudogenization dynamics in bats reveal ancestral vomeronasal signaling, then pervasive loss

Comparative methods are often used to infer loss or gain of complex phenotypes, but few studies take advantage of genes tightly linked with complex traits to test for shifts in the strength of selection. In mammals vomerolfaction detects chemical cues mediating many social and reproductive behaviors and is highly conserved, but all bats exhibit degraded vomeronasal structures with the exception of two families (Phyllostomidae and Miniopteridae). These families either regained vomerolfaction after ancestral loss, or there were many independent losses after diversification from an ancestor with functional vomerolfaction. In this study, we use the Transient receptor potential cation channel 2 (Trpc2) as a molecular marker for testing the evolutionary mechanisms of loss and gain of the mammalian vomeronasal system. We sequenced Trpc2 exon 2 in over 100 bat species across 17 of 20 chiropteran families. Most families showed independent pseudogenizing mutations in Trpc2, but the reading frame was highly conserved in phyllostomids and miniopterids. Phylogeny-based simulations suggest loss of function occurred after bat families diverged, and purifying selection in two families has persisted since bats shared a common ancestor. As most bats still display pheromone-mediated behavior, they might detect pheromones through the main olfactory system without using the Trpc2 signaling mechanism.

opencc-zeroDec 2016View details →
zenodo32/100

Datasets for phylogeny-guided analysis of pseudogenes in Pseudomonas aeruginosa

<p>In the study entitled &quot;A large-scale phylogeny-guided analysis of pseudogenes in <em>Pseudomonas aeruginosa</em>&nbsp;bacterium&quot;&nbsp;we analyzed the genomic data of 4699 strains of the bacterium <em>Pseudomonas aeruginosa P. aeruginosa)</em>&nbsp;as they exhibit high variability in the number of annotated pseudogenes. &nbsp;In particular, we&nbsp;looked for correlations between the number of pseudogenes and other genomic- and meta- features of the strains. We identified orthologous genes and pseudogenes and compared cluster size distributions and length homogeneity within clusters. We mapped and examined orthology relationships between genes and pseudogenes. We generated a phylogenetic tree of the strains and found that phylogenetically related strains are more homogeneous in the number of pseudogenes and share a significant amount of pseudogenes. Finally, we dived into clusters of orthologues genes and pseudogenes and quantified their phylogenetic neighborhood, classifying pseudogenes into evolutionary preserved pseudogenes, misannotated pseudogenes, or pseudogenes formed by failed horizontal transfer events. This in-depth study provides important insights that can be incorporated into pseudogene annotation pipelines in the future.<br> <br> We provide the following files:</p> <p>protein_seq.zip:&nbsp;A fasta file containing&nbsp;28,948,105 protein sequences of coding genes across all strains.</p> <p>pseudo_seq.zip:&nbsp;A fasta file containing&nbsp;86,273&nbsp;DNA sequences of pseudogenes across all strains.&nbsp;</p> <p>protein_clustering_output.zip: Clustering of protein sequences with CD-HIT clustering algorithm and the following parameters: similarity threshold of 70%, word size of 5, and the slow mode version.</p> <p>pseudo_clustering_output.zip:&nbsp;Clustering of pseudogene sequences with CD-HIT-EST clustering algorithm and the following parameters:&nbsp;similarity threshold of 80\%, word size of 5, and the slow mode version.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
dryad32/100

Data from: Development of a Chinook salmon sex identification SNP assay based on the growth hormone pseudogene

Open the record for dataset details and reuse information.

publicFeb 2015View details →
dryad32/100

Data from: Trpc2 Pseudogenization dynamics in bats reveal ancestral vomeronasal signaling, then pervasive loss

Open the record for dataset details and reuse information.

publicJan 2017View details →
dryad32/100

Datasets from: Validated removal of nuclear pseudogenes and sequencing artefacts from mitochondrial metabarcode

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad28/100

Data from: Alternative translation initiation codons for the plastid maturase MatK: unraveling the pseudogene misconception in the Orchidaceae

Background: The plastid maturase MatK has been implicated as a possible model for the evolutionary "missing link" between prokaryotic and eukaryotic splicing machinery. This evolutionary implication has sparked investigations concerning the function of this unusual maturase. Intron targets of MatK activity suggest that this is an essential enzyme for plastid function. The matK gene, however, is described as a pseudogene in many photosynthetic orchid species due to presence of premature stop codons in translations, and its high rate of nucleotide and amino acid substitution. Results: Sequence analysis of the matK gene from orchids identified an out-of-frame alternative AUG initiation codon upstream from the consensus initiation codon used for translation in other angiosperms. We demonstrate translation from the alternative initiation codon generates a conserved MatK reading frame. We confirm that MatK protein is expressed and functions in sample orchids currently described as having a matK pseudogene using immunodetection and reverse-transcription methods. We demonstrate using phylogenetic analysis that this alternative initiation codon emerged de novo within the Orchidaceae, with several reversal events at the basal lineage and deep in orchid history. Conclusion: These findings suggest a novel evolutionary shift for expression of matK in the Orchidaceae and support the function of MatK as a group II intron maturase in the plastid genome of land plants including the orchids.

opencc-zeroDec 2014View details →
zenodo28/100

Raw data for pseudogene identification

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Figure 1 from: Arthofer W, Avtzis D, Riegler M, Stauffer C (2010) Mitochondrial phylogenies in the light of pseudogenes and Wolbachia: re-assessment of a bark beetle dataset. ZooKeys 56: 269-280. https://doi.org/10.3897/zookeys.56.531

Figure 1 - in situ hybridization with wsp specific probe and staining with NBT/BCIP solution on uninfected A and Wolbachia infected Drosophila simulans B. An accumulation of dark color is observed only in ovarioles of Wolbachia infected Drospohila simulans. C Results of in situ hybridization of ovarial tissue excised from one Pityogenes chalcographus individual with accumulation of dark color (arrows). Three specimens were analysed. All ictures taken with 40-fold magnification.

opencc-by-4.0Sep 2010View details →
dryad28/100

Data from: Alternative translation initiation codons for the plastid maturase MatK: unraveling the pseudogene misconception in the Orchidaceae

Open the record for dataset details and reuse information.

publicSep 2015View details →
geo24/100

Exhaustive profiling in Arabidopsis reveals abundant polysome-associated 24-nt small RNAs including hitherto undescribed AGO5-associated pseudogene-derived siRNAs (psiRNAs)

GEO Series GSE99828. Arabidopsis thaliana. 10 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenDec 2017View details →
geo24/100

Gene expression profile of breast cancer cells transfected with antisense oligonucleotides (ASO) against BRCA1 pseudogene (BRCA1P1)

GEO Series GSE112572. Homo sapiens. 12 samples. Type: Expression profiling by array.

openGEO-OpenMar 2021View details →
geo24/100

Pseudogene INTS6P1 regulates its cognate gene INTS6 through competitive binding of miR-17-5p in hepatocellular carcinoma [mRNA, lncRNA]

GEO Series GSE64631. Homo sapiens. 6 samples. Type: Expression profiling by array; Non-coding RNA profiling by array.

openGEO-OpenJan 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record