Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
846
datasets available to search
ShareScore release 0.9.0
Dataset results
846 results for “homologs”
Molecular and electrophysiological features of GABAergic neurons in the dentate gyrus reveal limited homology with cortical interneurons
Open the record for dataset details and reuse information.
Ancestral origin and structural characteristics of non-syntenic homologous chromosomes in abalones (<em>Haliotis</em>)
Open the record for dataset details and reuse information.
Data from: CAnDI: a new tool to investigate conflict in homologous gene trees and explain convergent trait evolution
Open the record for dataset details and reuse information.
Data from: Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC
Open the record for dataset details and reuse information.
Scripts and data sets associated with: On testing homogeneity of the evolutionary process using alignments of homologous sequences
Open the record for dataset details and reuse information.
Phylogenetic analysis of Harmonin Homology Domains - Datasets
<p>Datasets associated to the article Phylogenetic analysis of Harmonin Homology Domains.</p> <p>HHD_starting-profile.hmm -> Profile HMM used to screen the UniprotKB</p> <p>HHD_all-hits_aligned.fa -> All hits aligned</p> <p>Other *.fa correspond to sequences identified for each cluster described in th article</p>
Homology analysis of HCV sequences obtained from a dialysis unit of northeast India
<p>This is a homology analysis (by DNASTAR MegaAlign Version 5.00) of 29 HCV Sequences (5'UTR-Core region sequencing) obtained from a dialysis unit of a tertiary care teaching hospital in northeast India. It is a part of a manuscript titled- "Circulation of an atypical hepatitis C virus (HCV) strain in a dialysis unit in northeast India". </p>
Supplementary table with homology analysis of twenty nine HCV sequences ((5'UTR-Core region) )
<p>It is a supplementary data table with homology analysis of 29 HCV sequences ((5’UTR-Core region) obtained from a dialysis unit of a tertiary care teaching hospital of Northeast India (Assam). It is a part (supplementary data) of the manuscript titled - "Circulation of an atypical hepatitis C virus (HCV) strain in a dialysis unit in northeast India" </p>
X-ray structure ensemble refinement of the second bromodomain of Pleckstrin homology domain interacting protein (PHIP) (space group P21212)
<p>X-ray structure ensemble refinement of the second bromodomain of Pleckstrin homology domain interacting protein (PHIP) (space group P21212). Raw diffraction images are available on Zenodo: 10.5281/zenodo.4086066. A single conformer model was deposited in the Protein Data Bank under accession code <a href="https://www.ebi.ac.uk/pdbe/entry/pdb/7AV8">7AV8</a>. Refinement was carried out with <a href="https://www.phenix-online.org/documentation/reference/ensemble_refinement.html">phenix.ensemble_refinement</a> and the repository contains all input and output files:</p> <p>Input:</p> <ul> <li>mx8421v63_xPHIPAx1521_free.mtz</li> <li>refine9.pdb</li> </ul> <p>Output:</p> <ul> <li>PHIPA-P21212_ensemble_refinement_ensemble.geo</li> <li>PHIPA-P21212_ensemble_refinement_ensemble.pdb</li> <li>PHIPA-P21212_ensemble_refinement_ensemble.mtz</li> <li>PHIPA-P21212_ensemble_refinement_ensemble.log</li> </ul>
Data from: Complex evolution of insect insulin receptors and homologous decoy receptors, and functional significance of their multiplicity
<p></p><p>Evidence accumulates that the functional plasticity of insulin and insulin-like growth factor signaling in insects could spring, among others, from the multiplicity of insulin receptors (InRs). Their multiple variants may be implemented in the control of insect polyphenism, such as wing or caste polyphenism. Here, we present a comprehensive phylogenetic analysis of insect InR sequences in 118 species from 23 orders and investigate the role of three InRs identified in the linden bug, Pyrrhocoris apterus, in wing polymorphism control. We identified two gene clusters (Clusters I and II) resulting from an ancestral duplication in a late ancestor of winged insects, which remained conserved in most lineages, only in some of them being subject to further duplications or losses. One remarkable yet neglected feature of InR evolution is the loss of the tyrosine kinase catalytic domain, giving rise to decoys of InR in both clusters. Within the Cluster I, we confirmed the presence of the secreted decoy of insulin receptor in all studied Muscomorpha. More importantly, we described a new tyrosine kinase-less gene (DR2) in the Cluster II, conserved in apical Holometabola for ∼300 My. We differentially silenced the three P. apterus InRs and confirmed their participation in wing polymorphism control. We observed a pattern of Cluster I and Cluster II InRs impact on wing development, which differed from that postulated in planthoppers, suggesting an independent establishment of insulin/insulin-like growth factor signaling control over wing development, leading to idiosyncrasies in the co-option of multiple InRs in polyphenism control in different taxa.</p><p></p>
Performance of virtual screening against GPCR homology models: Impact of template selection and treatment of binding site plasticity
<p>Rational drug design for G protein-coupled receptors (GPCRs) is limited by the small number of available atomic resolution structures. We assessed the use of homology modeling to predict the structures of two therapeutically relevant GPCRs and strategies to improve the performance of virtual screening against modeled binding sites. Homology models of the D<sub>2</sub> dopamine (D<sub>2</sub>R) and serotonin 5-HT<sub>2A</sub> receptors (5-HT<sub>2A</sub>R) were generated based on crystal structures of 16 different GPCRs. Comparison of the homology models to D<sub>2</sub>R and 5-HT<sub>2A</sub>R crystal structures showed that accurate predictions could be obtained, but not necessarily using the most closely related template. Assessment of virtual screening performance was based on molecular docking of ligands and decoys. The results demonstrated that several templates and multiple models based on each of these must be evaluated to identify the optimal binding site structure. Models based on aminergic GPCRs displayed ligand enrichment and there was a trend toward improved virtual screening performance with increasing binding site accuracy. The best models even displayed ligand enrichment better than that of the D<sub>2</sub>R and 5-HT<sub>2A</sub>R crystal structures. Methods to consider binding site plasticity were explored to further improve predictions. Molecular docking to ensembles of structures did not outperform the best individual binding site models, but could increase the diversity of hits from virtual screens and be advantageous for GPCR targets with few known ligands. Molecular dynamics refinement resulted in moderate improvements of structural accuracy and the virtual screening performance of snapshots was either comparable to or worse than that of the raw homology models. These results provide guidelines for successful application of structure-based ligand discovery using GPCR homology models.</p>
Supplemental Movie Files for "Agent-Based Modeling of a Nuclear Chromosome Ensemble Identifies Determinants of Homolog Pairing During Meiosis" by Chriss et al.
<p>This set of Supplemental Information contains two movies made from the simulations from the model developed in the manuscript "<strong>Agent-Based Modeling of a Nuclear Chromosome Ensemble Identifies Determinants of Homolog Pairing During Meiosis</strong>" by A. Chriss, G. V. Börner, and S. D. Ryan. </p> <p> </p> <p>Supplemental Movie S1: <strong>WT Chromosome Trajectories during Prophase I. </strong>The first file "movie_WT..." contains the file for the results of simulations for the wild-type chromosomes and the exact parameter values can be found in Table 1 of the manuscript. The movie shows one realization of the agent-based model. The simulation movie covers the homology search process from <em>t = 3h </em>to <em>t = 9h</em>. Matching colors correspond to homologous pairs. True chromosome lengths are incorporated and scale the relevant interaction radii. The radius represents the attractive and non-homologous repulsive region.</p> <p> </p> <p>Supplemental Movie S2: <strong>WT Chromosome Trajectories during Prophase I with active dumbbell model. </strong> The second file "movie<em>WT</em>_activedumbbell..." contains the file for the results of the simulations for the modeling of chromosomes as active dumbbells (from polymers) to allow for the study of the effects of elongation, orientation, and flexibility. The movie shows one realization of the agent-based active dumbbell model which is closer to modeling a chromosome as a polymer. The simulation movie covers the homology search process from <em>t = 3h </em>to <em>t = 9h</em>. Matching colors correspond to homologous pairs. True chromosome lengths are incorporated and scale the relevant interaction radii, but are allowed to change in time as the two beads expand and contract. The radius represents the attractive and non-homologous repulsive region.</p> <p> </p> <p>Supplemental Movie S3: <strong><em>spo11</em> hypomorph (30% WT DSB levels) Chromosome Trajectories during Prophase I (parameters from Fig 7B)</strong>. The third file "movie_spo11..." contains the file for the results of the simulations for the spo-11 hypomorph and the associated parameter values can be found in Table 1 of the manuscript. The movie depicts one realization of the agent-based model for the {\it spo11} hypomorphic mutant. The simulation movie covers the homology search process from <em>t = 3h</em> to <em>t = 9h</em> where mutant <em>spo11</em> is associated with a weaker attractive and repulsive force (e.g., reduction to 77% of WT values). True chromosome lengths are incorporated and scale the relevant interaction radii. Matching colors correspond to homologous pairs. The radii represent the homologous attractive and the non-homologous repulsive region. Note that the reduction in interaction strength delays homologous pairing consistent with experimental observations in [13]. </p> <p> </p> <p> </p> <p>The codes that generated these movies were written in Matlab and freely available via GitHub: <a href="https://github.com/sdryan/ChromosomeDynamicsProphase1">https://github.com/sdryan/ChromosomeDynamicsProphase1</a></p> <p> </p> <p>For questions please contact the corresponding authors: G. Valentin Börner <a href="mailto:g.boerner@csuohio.edu">g.boerner@csuohio.edu</a> (Biology) or Shawn D. Ryan <a href="mailto:s.d.ryan@csuohio.edu">s.d.ryan@csuohio.edu</a> (Math).</p>
Enhancing comparative T-cell receptor repertoire analysis in small biological samples through pooling homologous cell samples from multiple mice
<p>All data files used to generate the figures in the paper are shared in this project.</p> <p>Scripts are available on <a href="https://github.com/i3-unit/CRM_24" target="_blank" rel="noopener">GitHub</a>.</p>
Data and Weights for Reverse Homology
<p>Training data, weights, and classification datasets for "Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning". </p> <p><strong>Training data:</strong></p> <ul> <li>scer_idr_homologues and human_idr_homologues contain a zip file of the fasta files of IDR homologues used to train the yeast and human model, respectively. Note that these fasta files are aligned, but we strip away the alignment symbol "-" before input into our model. disprot_idr_homologues contain a zip file of IDR homologues corresponding to the DisProt database.</li> <li>human_protein_alignments contains fasta files for the full proteins containing these IDRs</li> </ul> <p><strong>Weights:</strong></p> <ul> <li>scer_idr_model and human_idr_model contain a zip file of the weights for the yeast and human model respectively, which can be loaded into the model files at github.com/alexxijielu/reverse_homology. Likewise, disprot_idr_model contains a zip file of our model trained on DisProt IDRs exclusively.</li> </ul> <p><strong>Logo Websites:</strong></p> <ul> <li>scer_idr_logo_website, human_idr_logo_website, and disprot_idr_logo_website contain a zip file of an HTML file for the yeast, human, and DisProt model respectively, showing sequence logos of the features learned by each model and their enrichments.</li> </ul> <p><strong>Features:</strong></p> <ul> <li>human_idr_features contains the raw feature activations for all human IDRs in our human model. (We didn't include this file in the supplementary for the paper due to size.) </li> </ul> <p><strong>Classification datasets:</strong></p> <ul> <li>IDR_classification_datasets contains datasets used in our benchmarks. These datasets are encoded as binary csv matrixes. cdc28_classification contains IDRs labeled as Cdc28 phosphorylation sites, mitochondrial_targeting_classification contains IDRs labeled as mitochondrial targeting signals, evosig_cluster_classification contains IDRs labeled by clusters assigned in previous computational work by Zarin <em>et al</em>. eLife 2019, and go_SLIM_classification contains proteins labeled by GO Slim annotations. </li> </ul>
Acute pseudo-landmarking and Constellation homologies: A generalized workflow to identify and track segmented structures in plant time series images
<p>Assessing plant phenotypes throughout the lifecycle is integral to exploring the development, genetics, and evolution of morphology, and can be critical for agronomic and basic research studies. Although various automated or semi-automated phenomic approaches have been developed, it has been challenging to analyze differential growth because of difficulties in segmenting and annotating specific structures or positions in the plant body and maintaining their identities throughout time-series data. To address this gap, we have developed a generalized workflow linking our previously published function, <i>Acute</i>, with a companion homology workflow, <i>Constellation</i>, in the PlantCV environment. <i>Acute</i> identifies acute shapes (pseudo-landmarks) in the plant body, most often corresponding to leaf tips and ligular regions. <i>Constellation</i> uses a strategy of dimensionality reduction via <i>starscape</i> followed by hierarchical clustering through <i>constella </i>to identify 'constellations' of segments in eigenspace that represent the same landmark in consecutive images of a time-series. We devised a quality control function, <i>constellaQC</i>, to test the accuracy of the clustering approach, and use it to show that the approach appropriately clusters the pseudo-landmarks derived from <i>Acute</i>, with 80-90% accuracy. We discuss the reasons for and consequences of this lack of 100% accuracy in automated workflows and suggest how to develop these functions for other phenomics datasets that may vary in dimensional complexity.</p>
Specimen alignment with limited point-based homology: 3D morphometrics of disparate bivalve shells (Mollusca: Bivalvia)
<p>Supplemental data and code for Edie, Collins, and Jablonski, Specimen alignment with limited point-based homology: 3D morphometrics of disparate bivalve shells (Mollusca: Bivalvia).</p>
Synapsed homologs of meiotic mouse chromosomes visualized by TIRFM
<p>Mouse testes were removed from euthanized animals, followed by decapsulation and maceration in high-glucose MEM medium. Suspension was mixed thoroughly and left to settle; the supernatant was then collected and centrifuged at 7200 rpm for 1 min. The pellet was then resuspended in a 0.5 M sucrose solution and added to PFA-treated (1% in 0.015% Triton X-100) coverslides, which were incubated for at room temperature for 2 hours in a humidified environment. After incubation, slides were rinsed twice with a wetting agent solution (Kodak, 1464510) in water and allowed to air dry. SYCP3 was labeled with primary SCP-3 (D-1) antibody (Santa Cruz Biotechnology, SC-74569) and a secondary anti-mouse antibody fused with Alexa-568 (Thermo, A11004). Surface chromosome spreads were imaged in an Elyra 7 microscope (Zeiss), with a 60x 1.4 NA oil immersion objective and an a 1.4x magnification lens. Image reconstruction was performed in ZEN Black with parameters set to default.</p> <p>Experimental procedures were approved by the “Ministero della Salute" of Italy, authorization n.701/2018-PR.</p>
GPI-anchored protein homolog IcFBR1 functions directly in morphogenesis of Isaria cicadae
<p><em>Isaria cicadae</em> is a famous edible and medicinal fungus in China and Asia. The molecular basis of morphogenesis and synnemal formation needs to be understood in more detail, because this is the main source of biomass production in <em>I. cicadae</em>. In the present study, a fruiting body formation-related gene with a glycosylphosphatidylinositol (GPI) anchoring protein (GPI-Ap) gene homolog, <em>IcFBR1</em>, was identified by screening random insertion mutants. Targeted deletion of <em>IcFBR1</em> resulted in abnormal, non-formation of synnemata, impairing aerial hyphae growth and sporulation. The <em>IcFBR1</em> mutants were defective in utilization of carbon sources with reduced polysaccharide contents and in the regulation of amylase and protease activities. Transcriptome analysis of <em>ΔIcfbr1</em> showed that <em>IcFBR1</em> deletion influenced 49 gene ontology terms, including 23 biological processes, nine molecular functions, and 14 cellular components. IcFBR1 is therefore necessary for regulating synnemal development, secondary metabolism, and nutrient utilization in this important edible and medicinal fungus. This is the first report illustrating that the function of IcFBR1 is associated with the synnemata in <em>I. cicadae</em>.</p>
Raw NGS Data for "Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins"
<p>This directory contains relevant fastq files used for deep sequencing analysis in the publication “Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins”. </p> <p>Fastq files are provided for presorted, uninduced and induced populations from DMS experiments of four homologs (TtgR, TetR, RolR, and MphR). Three replicates were performed for each sample.</p> <p>Data analysis of this deep sequencing data was performed using custom scripts, which are described in the methods section of the publication.</p>
Bacterial homologs of innate eukaryotic antiviral defenses provide phage protection.
<p>Supplementary data regarding the 'Bacterial homologs of innate eukaryotic antiviral defenses provide phage protection' manuscript.</p><p> </p><p><strong>Abstract:</strong><br>Prokaryotes have evolved a multitude of defense systems to protect themselves from bacteriophage predation. Here, we discovered new phage defense systems related to innate antiviral genes from vertebrates and plants. Our search uncovered over 400 candidates from which eleven were selected and six novel phage defense systems validated. We identified a DNA replication helicase/nuclease 2 (Prometheus) which may act on transcription R-loops, an inositol-monophosphatase-like phage defense protein (Pan), and two ATPases from the NACHT family coupled to novel effectors NucS and SfsA (Nyx and Hypnos). In addition, a fused ubiquitin-like E1-E2-JAB protein combined with a putative MBL nuclease (6A-MBL) was found, and a novel member of the Thoeris family that contains four essential TIR domains with a putative effector SLOG domain (Thoeris type III). Collectively, these defense systems support the concept of deep evolutionary links and shared antiviral mechanisms between prokaryotes and eukaryotes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.