Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

132

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

132 results for “paralogy”

Learn how ShareScore rates datasets ↗
zenodo48/100

Protein structure files for the paper "Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth" in Nature Cell Biology by Tang et al.

<p>This archive contains models of HRAS, KRAS, and NRAS homo- and heterodimers with various mutations discussed in the paper,&nbsp; &quot;Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth&quot; in Nature Cell Biology by Tang et al.<br> as well as crystallographic dimers of these proteins as identified by the ProtCAD database, http://dunbrack2.fccc.edu/ProtCAD/Results/PfamArchClusterInfo.aspx?GroupId=8 (cluster 5). Several of the models are shown in Supp. Figure 11b and the crystallographic dimers of RAS that provide evidence for the possible biological relevance of these models are shown in Supp. Figure 11a.</p> <p>The crystallographic dimers were identified by clustering all possible interfaces generated by symmetry operators in crystals of HRAS, KRAS, and NRAS as described in the paper: Xu, Q., Dunbrack, R.L. ProtCID: a data resource for structural information on protein interactions. <em>Nat Commun</em> <strong>11</strong>, 711 (2020). https://doi.org/10.1038/s41467-020-14301-4.</p> <p>The models were created by superposing monomers of HRAS, KRAS, or NRAS onto the alpha4-alpha5 dimer present in the crystal of PDB entry 3k8y. Mutations were made in PyMOL. The structures were relaxed with the FastRelax protocol and the Ref2015 scoring function in the program Rosetta, which uses the backbone-dependent rotamer library of Shapovalov and Dunbrack to repack side chains.</p> <p>The crystallographic dimers are contained in a zipped PyMOL session. The mmCIF format for all the structures is present in a zip file, Tang_et_al_crystallographic_and_modeled_RAS_dimer_ciffiles.zip. The PyMOL session and zip file contains 87 HRAS dimers, 14 KRAS dimers, and 1 NRAS dimer, all having the interface consisting of the alpha4 and alpha5 helices. The PyMOL session also contains the modeled structures. Only Mg ions and GTP/GNP/GDP ligands are shown. Others are present but hidden and may be displayed by PyMOL (&quot;show sticks, het&quot;).</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Paralog variant classification and scoring

<p><em>Para_zscore </em>data</p> <p>Input data, annotation of all hg19 missense variants, score for every gene having a paralog in the human genes. This dataset is a supplement for the publication Lal. et al.</p> <p>Information on the files, scripts to generate and use the <em>para_zscore</em> are available under</p> <p>https://git-r3lab.uni.lu/genomeanalysis/paralogs.</p> <p>Version 3582386 updates:</p> <p>- Annovar annotation file for hg38 added</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

common function paralog pairs and sequence similarity features

<p>This repository contains datasets of selected paralog pairs (Ensembl 111), labeled with various "common function" annotations, including PPI, SL, and GO datasets for both human (<em>Homo sapiens</em>) and budding yeast (<em>Saccharomyces cerevisiae</em>). These paralog pairs are characterized using different sequence similarity features, such as AlphaFold-predicted structures, Protein Language Model embeddings, and similarity searches from various databases.</p> <p>&nbsp;</p> <p>These datasets are used in the following manuscript:&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2024.10.11.617835v1">Evaluating Sequence and Structural Similarity Metrics for Predicting Shared Paralog Functions</a></p> <p>&nbsp;</p> <p>For the corresponding analysis notebooks, see:&nbsp;<a href="https://github.com/cancergenetics/paralog_seq_similarity/tree/main">github.com/cancergenetics/paralog_seq_similarity/</a></p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

FIG. 5 in Temporal paralogy, cladograms, and the quality of the fossil record

FIG. 5. — Pectinate cladogram illustrating Pongidae relationships as currently understood; A, all nodes are temporally informative (orthologous). The age of diversification of Pan is more recent than the age of origin of (Homo, Pan); B, the addition of a fossil taxon (Homo neanderthalensis) defines a paralogous node and two new arrows of time (black arrows). Consequently the fit of ages of diversification of both Homo and Pan to stratigraphy is meaningless.

opencc-zeroDec 2004View details →
zenodo40/100

FIG. 4 in Temporal paralogy, cladograms, and the quality of the fossil record

FIG. 4. — Temporal paralogy and the origin of tetrapods. The possibility of osteolepiforms being ancestors of tetrapods has been reinterpreted based on a parsimony analysis (Ahlberg &amp; Johanson 1998). Among osteolepiforms, the Tristichopteridae appear as the closest relative to the Tetrapoda. The common ancestor of both groups is represented by a paralogous node, and according to temporal hierarchy the relative inclusiveness of each of the sister-groups cannot be decided. As a consequence, the Tetrapoda can be supposed to have occurred before the first appearance of osteolepiforms (Eifelian). There are no arguments to see osteolepiforms as possible ancestors of tetrapods, if this question has any meaning when argued from a parsimony analysis.

opencc-zeroDec 2004View details →
zenodo40/100

FIG. 3 in Temporal paralogy, cladograms, and the quality of the fossil record

FIG. 3. — Temporal information and cladograms; A, maximally informative cladogram of taxa (A-F) and their ages (6-1). All nodes are orthologous, and the ages of the fossil specimens can be either consistent or not with the temporal hierarchy; B, the effect of a better knowledge of the fossil record by addition of the age of a supplementary taxon (N, 4) is the decrease in temporal resolution. The number of informative (orthologous) nodes decreases (white circles) as temporally ambiguous nodes appear (shown as grey circles) node leading to two terminals or a terminal and a paralogous node, the black circle corresponds a new paralogous node, indicating a temporal paralogy. The ages of sister-taxa (C, N) and (D, E, F) are temporal paralogs. Each arrow represents a semi-independent temporal hierarchy.

opencc-zeroDec 2004View details →
zenodo36/100

AF2 models for "Co-translational assembly promotes functional diversification of paralogous proteins" by Mallik, et al.

<p>AF2 models used for analyses presented in "Co-translational assembly promotes functional diversification of paralogous proteins" by Saurav Mallik, Angel F. Cisneros, Christian R. Landry, and Emmanuel D. Levy.</p> <p>These models were used to analyze the structural divergence of 3703 Obligatory Homomer, 697 Mixed, and 181 Obligatory Heteromer pairs.</p> <p>Folders are separated into different categories:<br>- Models of homomers (from Schweke et al., 2024. Cell):<br>&nbsp; &nbsp; . AF2_HM_full_models: &nbsp;Full structures of homodimeric models.<br>&nbsp; &nbsp; . AF2_HM_nodiso3: Core structures of homomeric models, trimmed using scripts from Schweke et al. (2023).</p> <p>- Models of heteromers (generated in this work):<br>&nbsp; &nbsp; . AF2_HET_full_models: Full structures of heterodimeric models.<br>&nbsp; &nbsp; . AF2_HET_nodiso3: Core structures of heteromeric models, trimmed using a modified version of the code from Schweke et al. (2023) to work with heterodimers.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Supplemental Dataset Excel files and Source Data Excel file for "START domains generate paralog-specific regulons from a single network architecture"

<p>Supplemental Dataset Excel files and Source Data Excel file for "START domains generate paralog-specific regulons from a single network architecture" in Nat Comms</p>

opengpl-3.0-or-laterOct 2024View details →
dryad36/100

A new pipeline for removing paralogs in target enrichment data

<p><span><span>Target enrichment (such as Hyb-Seq) is a well-established high throughput sequencing method that has been increasingly used for phylogenomic studies. Unfortunately, current widely used pipelines for analysis of target enrichment data do not have a vigorous procedure to remove paralogs in target enrichment data. In this study, we develop a pipeline we call Putative Paralogs Detection (PPD) to better address putative paralogs from enrichment data. The new pipeline is an add-on to the existing HybPiper pipeline, and the entire pipeline applies criteria in both sequence similarity and heterozygous sites at each locus in the identification of paralogs. Users may adjust the thresholds of sequence identity and heterozygous sites to identify and remove paralogs according to the level of phylogenetic divergence of their group of interest. The new pipeline also removes highly polymorphic sites attributed to errors in sequence assembly and gappy regions in the alignment. We demonstrated the value of the new pipeline using empirical data generated from Hyb-Seq and the Angiosperm 353 kit for two woody genera <i>Castanea</i> (Fagaceae, Fagales) and <i>Hamamelis</i> (Hamamelidaceae, Saxifragales). Comparisons of datasets showed that the PPD identified many more putative paralogs than the popular method HybPiper. Comparisons of tree topologies and divergence times showed evident differences between data from HybPiper and data from our new PPD pipeline. We further evaluated the accuracy and error rates of PPD by BLAST mapping of putative paralogous and orthologous sequences to a reference genome sequence of<i> Castanea mollissima</i>. Compared to HybPiper alone, PPD identified substantially more paralogous gene sequences that mapped to multiple regions of the reference genome (31 genes for PPD compared with 4 genes for HybPiper alone). In conjunction with HybPiper, paralogous genes identified by both pipelines can be removed resulting in the construction of more robust orthologous gene datasets for phylogenomic and divergence time analyses. Our study demonstrates the value of Hyb-Seq with data derived from the Angiosperm 353 probe set for elucidating species relationships within a genus, and argues for the importance of additional steps to filter paralogous genes and poorly aligned regions (e.g., as occur through assembly errors), such as our new PPD pipeline described in this study.</span></span></p>

opencc-zeroJun 2021View details →
dryad36/100

Data from: Dissection of the role of a SH3 domain in the evolution of binding preference of paralogous proteins

<p><span>Protein-protein interactions drive many cellular processes. Some protein interactions are directed by Src homology 3 (SH3) domains that bind proline-rich motifs on other proteins. The evolution of the binding specificity of SH3 domains is not completely understood, particularly following gene duplication. Paralogous genes accumulate mutations that can modify protein functions and, for SH3 domains, their binding preferences. Here, we examined how the binding of the SH3 domains of two paralogous yeast type I myosins, Myo3 and Myo5, evolved following duplication. We found that the paralogs have subtly different SH3-dependent interaction profiles. However, by swapping SH3 domains between the paralogs and characterizing the SH3 domains freed from their protein context, we find that few of the differences in interactions, if any, depend on the SH3 domains themselves. We used ancestral sequence reconstruction to resurrect the pre-duplication SH3 domains and examined, moving back in time, how the binding preference changed. Although the closest ancestor of the two domains had a very similar binding preference as the extant ones, older ancestral domains displayed a gradual loss of interaction with the modern interaction partners when inserted in the extant paralogs. Molecular docking and experimental characterization of the free ancestral domains showed that their affinity with the proline motifs is likely not the cause for this loss of binding. Taken together, our results suggest that the SH3 and its host protein could create intramolecular or allosteric interactions essential for the SH3-dependent PPIs, making domains not functionally equivalent even when they have the same binding specificity. </span></p>

opencc-zeroSep 2023View details →
dryad36/100

Data from: Addressing incomplete lineage sorting and paralogy in the inference of uncertain salmonid phylogenetic relationships

Open the record for dataset details and reuse information.

publicMay 2020View details →
dryad36/100

Data from: Evidence of functional divergence in MSP7 paralogous proteins: a molecular-evolutionary and phylogenetic analysis

Open the record for dataset details and reuse information.

publicDec 2016View details →
dryad36/100

Data from: Advancing Pyrus phylogeny: Deep genome skimming-based inference coupled with paralogy analysis yields a robust phylogenetic backbone and an updated infrageneric classification of the pear genus (Maleae, Rosaceae)

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad36/100

Data from: Dissection of the role of a SH3 domain in the evolution of binding preference of paralogous proteins

Open the record for dataset details and reuse information.

publicSep 2023View details →
dryad36/100

A new pipeline for removing paralogs in target enrichment data

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad36/100

ASTRAL-Pro: Quartet-based species-tree inference despite paralogy

Open the record for dataset details and reuse information.

publicJul 2023View details →
dryad32/100

Systemic paralogy and function of retinal determination network homologs in arachnids

<p>Arachnids are important components of cave ecosystems and display many examples of troglomorphisms, such as blindness, depigmentation, and elongate appendages. Little is known about how the eyes of arachnids are specified genetically, let alone the mechanisms for eye reduction and loss in troglomorphic arachnids. Additionally, paralogy of Retinal Determination Gene Network (RDGN) homologs in spiders has convoluted functional inferences extrapolated from single-copy homologs in pancrustacean models. Here, we investigated a sister species pair of Israeli cave whip spiders (Arachnopulmonata, Amblypygi, <i>Charinus</i>) of which one species has reduced eyes. We generated the first embryonic transcriptomes for Amblypygi, and discovered that several RDGN homologs exhibit duplications. We show that paralogy of RDGN homologs is systemic across arachnopulmonates (arachnid orders that bear book lungs), rather than being a spider-specific phenomenon. A differential gene expression (DGE) analysis comparing the expression of RDGN genes in field-collected embryos of both species identified candidate RDGN genes involved in the formation and reduction of eyes in whip spiders. To ground bioinformatic inference of expression patterns with functional experiments, we interrogated the function of three candidate RDGN genes identified from DGE in a spider, using RNAi in the spider <i>Parasteatoda tepidariorum</i>. We provide functional evidence that one of these paralogs, <i>sine oculis/Six1 A </i>(<i>soA</i>), is necessary for the development of all arachnid eye types. Our results support the conservation of at least one RDGN component across Arthropoda and establish a framework for investigating the role of gene duplications in arachnid eye diversity.</p>

opencc-zeroNov 2020View details →
dryad32/100

Data from: Differential requirements for the RAD51 paralogs in genome repair and maintenance in human cells

Deficiency in several of the classical human RAD51 paralogs [RAD51B, RAD51C, RAD51D, XRCC2 and XRCC3] is associated with cancer predisposition and Fanconi anemia. To investigate their functions, isogenic disruption mutants for each were generated in non-transformed MCF10A mammary epithelial cells and in transformed U2OS and HEK293 cells. In U2OS and HEK293 cells, viable ablated clones were readily isolated for each RAD51 paralog; in contrast, with the exception of RAD51B, RAD51 paralogs are cell-essential in MCF10A cells. Underlining their importance for genomic stability, mutant cell lines display variable growth defects, impaired sister chromatid recombination, reduced levels of stable RAD51 nuclear foci, and hyper-sensitivity to mitomycin C and olaparib. Altogether these observations underscore the contributions of RAD51 paralogs in diverse DNA repair processes, and demonstrate essential differences in different cell types. Finally, this study will provide useful reagents to analyze patient-derived mutations and to investigate mechanisms of chemotherapeutic resistance deployed by cancers.

opencc-zeroSep 2019View details →
dryad32/100

Data from: Mouse fitness measures reveal incomplete functional redundancy of Hox paralogous group 1 proteins

Here we assess the fitness consequences of the replacement of the Hoxa1 coding region with its paralog Hoxb1 in mice (Mus musculus) residing in semi-natural enclosures. Previously, this Hoxa1B1 swap was reported as resulting in no discernible embryonic or physiological phenotype (i.e., functionally redundant), despite the 51% amino acid sequence differences between these two Hox proteins. Within heterozygous breeding cages no differences in litter size nor deviations from Mendelian genotypic expectations were observed in the outbred progeny; however, within semi-natural population enclosures mice homozygous for the Hoxa1B1 swap were out-reproduced by controls resulting in the mutant allele being only 87.5% as frequent as the control in offspring born within enclosures. Specifically, Hoxa1B1 founders produced only 77.9% as many offspring relative to controls, as measured by homozygous pups, and a 22.1% deficiency of heterozygous offspring was also observed. These data suggest that Hoxa1 and Hoxb1 have diverged in function through either sub- or neo-functionalization and that the HoxA1 and HoxB1 proteins are not mutually interchangeable when expressed from the Hoxa1 locus. The fitness assays conducted under naturalistic conditions in this study have provided an ultimate-level assessment of the postulated equivalence of competing alleles. Characterization of these differences has provided greater understanding of the forces shaping the maintenance and diversifications of Hox genes as well as other paralogous genes. This fitness assay approach can be applied to any genetic manipulation and often provides the most sensitive way to detect functional differences.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Congruent population structure across paralogous and non-paralogous loci in Salish Sea chum salmon (Oncorhynchus keta)

Whole genome duplications are major evolutionary events with a lasting impact on genome structure. Duplication events complicate genetic analyses as paralogous sequences are difficult to distinguish; consequently paralogs are often excluded from studies. The effects of an ancient whole genome duplication (approximately 88MYA) are still evident in salmonids through the persistence of numerous paralogous gene sequences and partial tetrasomic inheritance. We use restriction site-associated DNA sequencing (RADseq) on ten collections of chum salmon from the Salish Sea in the USA and Canada to investigate genetic diversity and population structure in both tetrasomic and re-diploidized regions of the genome. We use a pedigree and high-density linkage map to identify paralogous loci and to investigate genetic variation across the genome. By applying multivariate statistical methods, we show that it is possible to characterize paralogous genetic loci and that they display similar patterns of population structure as the diploidized portion of the genome. We find genetic associations with the adaptively important trait of run timing in both sets of loci. By including paralogous loci in genome scans, we can observe evolutionary signals in genomic regions that have routinely been excluded from population genetic studies in other polyploid-derived species.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record