Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

121

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

121 results for “orthology”

Learn how ShareScore rates datasets ↗
zenodo48/100

Bedrock radioactivity influences the rate and spectrum of mutation - Orthologous genes

<p>Alignments of the 2490 orthologous genes used in the article &quot;Natural Bedrock radioactivity influences the rate and spectrum of mutation&quot; to estimate the mutational spectrum and synonymous substitution rate.</p> <p>To compute accurate synonymous substitution rate, we removed genes with short sequences (&lt;half of the alignment) and genes strongly supporting another phylogeny using ProfileNJ <a href="https://paperpile.com/c/Klqlpb/W5sS">(Noutahi et al. 2016)</a> with a bootstrap threshold of 90%, resulting in a subset of 769 genes listed in the file &quot;List_769_1-to-1_orthologs_EvolutionRate.txt&quot;.</p> <p>Transcriptome paired-end reads used to define these orthologous genes have been deposited to the European Nucleotide Archive and are available under the study ID PRJEB14193.</p> <p>Sequences were aligned with Prank<a href="https://paperpile.com/c/Klqlpb/pilh"> (L&ouml;ytynoja &amp; Goldman 2008)</a> using a codon model and sites ambiguously aligned were removed with Gblocks <a href="https://paperpile.com/c/Klqlpb/c5kb">(Castresana 2000)</a>.</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Collation and orthology-based identification of hormone-related genes in bread wheat

<p>Plant hormones coordinate a plethora of developmental processes in plants, including responses to abiotic and biotic stressors. Here, we collate the findings of previous studies identifying bread wheat (<em>Triticum aestivum</em>) genes related to hormonal processes (<strong>biosynthesis</strong>, <strong>transport</strong>, <strong>signalling</strong>, and <strong>catabolism</strong>) and collect wheat orthologues from hundreds of additional hormone-related genes utilising the Ensembl Plants Compara database. We have initially conducted this procedure for <strong>abscisic acid</strong>, <strong>auxins </strong>(IAA and IBA), <strong>brassinosteroids</strong>, <strong>cytokinins</strong>, <strong>ethylene</strong>, <strong>gibberellins</strong>, and <strong>strigolactone</strong>, yielding a total of over 1,700 putative wheat orthologues. We aim to provide a community resource to aid gene annotation and subsequent analyses. We warmly welcome feedback from the community.</p> <p>Please refer to the file <strong>README.pdf</strong> for further details, including methods and references.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Allliance of Genome Resources Orthology

<p>Tab separated formatted spreadsheet of orthology annotations from the Alliance of Genome Resources.</p> <p>The Alliance provides the results of all methods that have been benchmarked by the <a href="https://questfororthologs.org/">Quest for Orthologs Consortium (QfO)</a>, as well as curated ortholog inferences from HGNC (for human and mouse genes), Xenbase (for frog genes), and ZFIN (relating zebrafish genes to orthologs in human, mouse, and fly).</p> <p>The ortholog inferences from the different methods have been integrated using the DRSC Integrative Ortholog Prediction Tool (DIOPT). DIOPT integrates a number of existing methods including those used by the Alliance: Ensembl Compara, HGNC, Hieranoid, InParanoid, OMA, OrthoFinder, OrthoInspector, PANTHER, PhylomeDB, SonicParanoid, Xenbase, and ZFIN. See the <a href="https://fgr.hms.harvard.edu/diopt-documentation">DIOPT documentation</a> for additional information and references related to the included methods. DIOPT assigns a score/count based on the number of methods that call a specific ortholog. For noncoding RNA genes, currently only HGNC and ZFIN curated orthologs are included.</p> <p>File includes orthology relationships among genes from the following organisms:</p> <ul> <li>Homo sapiens (human; NCBI:txid 9606)</li> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> <li>Xenopus tropicalis (Western clawed frog; NCBI:txid 8364)</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Orthologs for Homo Sapiens and Drosophila Melanogaster from the Roundup Orthology Database Version 4

<p>This dataset contains orthologs for Homo sapiens and Drosophila melanogaster, computed with&nbsp;the Reciprocal Smallest Distance algorithm using a divergence threshold&nbsp;of 0.8 and an e-value threshold of 1e-5. The orthologs were downloaded from version 4 of the Roundup database, which is no longer available.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Data sets for orthologous target pair analysis

<p>The set of all 803 originally identified orthologous target pairs (OTPs) and the subset of 222 OTPs with at least 10 shared compounds are provided herein. For each OTP both organisms, the target, the number of shared compounds,the OTP category, and the number of reference articles is reported. In addtion, the list of all 1149 candidate compounds and their human target assignments is provided.&nbsp;</p>

opencc-zeroMay 2015View details →
zenodo40/100

Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue for single nucleotide polymorphisms discovery (data set)

<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of <em>Festuca pratensis</em> and ten samples (corresponding to six genotypes) of <em>Lolium multiflorum</em> and fourteen samples of<em> Lolium perenne</em> (corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant <em>F. pratensis</em> transcripts were assembled by combining the contigs of all six genotypes based on orthology with the <em>Brachypodium distachyon </em>proteome. Similarly, <em>19,036</em> non-redundant<em> L. multiflorum</em> transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em> L. perenne</em> transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>

opencc-zeroFeb 2016View details →
zenodo40/100

Heatmaps of orthology and protein domain preservation in RNA Processing complexes throughout the fungal kingdom

<p>An analysis of the presence/absence of orthologues for Fungal RNA Processing protein complexes, and the presence/absence of the known PFAM protein domains within each protein within these complexes in the organism's proteome.  </p> <p>Each image represents one RNA Processing protein complex.</p> <p>Orthology (far left panel in each image) is relative to Yeast, and taken from a query against the EnsEMBL orthology database (black = no orthologue; red = orthologue).  <br> <br> Each orthologue was then queried for its PFAM domains, and the non-redundant set of PFAM domains representing each set of orthologous proteins, spanning all species, was then scanned against the complete proteome of each species.  The resulting heatmap indicates the presence or absence of that PFAM domain anywhere in the proteome of that species.  (black = absent; red = 1 copy; grey-&gt;blue = more than one copy)</p>

opencc-by-4.0Mar 2016View details →
zenodo40/100

Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue (update data set)

<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of&nbsp;<em>Festuca pratensis</em>&nbsp;and ten samples (corresponding to six genotypes) of&nbsp;<em>Lolium multiflorum</em>&nbsp;and fourteen samples of<em>&nbsp;Lolium perenne</em>&nbsp;(corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant&nbsp;<em>F. pratensis</em>&nbsp;transcripts were assembled by combining the contigs of all six genotypes based on orthology with the&nbsp;<em>Brachypodium distachyon&nbsp;</em>proteome. Similarly,&nbsp;<em>19,036</em>&nbsp;non-redundant<em>&nbsp;L. multiflorum</em>&nbsp;transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em>&nbsp;L. perenne</em>&nbsp;transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>

opencc-zeroJun 2016View details →
zenodo40/100

Datasets for "Deep learning-based design of synthetic orthologs of SH3 signaling domains"

<p>Description of data for "Deep learning-based design of synthetic orthologs of SH3 signaling domains":<br><br>biochemistry_data.zip --&gt; contains the binding assay, melting temperature, and enthalpy measurements that reproduce table 1 in the main text.<br>sequence_data.zip --&gt; contains the sequences with relative enrichment measurements and other meta data information (e.g. latent embeddings, paralog labels, etc.).<br><br>sequence_data.zip &gt; SH3_Library_Natural.xlsx --&gt; contains the natural alleles with normalized relative enrichment scores, paralog labels, mmd latent coordinates, and among other meta data.<br>sequence_data.zip &gt; SH3_Library_Design.xlsx --&gt; contains the design alleles with normalized relative enrichment scores, mmd latent coordinates, and among other meta data.<br>sequence_data.zip &gt; paralog_mapping.xlsx --&gt; contains the mapping between paralog name and COG labels.<br><br>note: to access these spreadsheets for analysis, we recommend using pandas library in Python. To reproduce figures within the manuscript, please follow this github repo link: https://github.com/chemgeeklian/SH3_orthology_paper_analysis</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Ortholog data from the tuatara genome project

<p>This record contains orthology predictions based on the Maker gene annotation<br> of the tuatara (Sphenodon punctatus) and a set of 25 other species, using the<br> Ensembl methodology.</p> <p>See http://doi.org/10.5281/zenodo.1489354 for the Maker annotation in GFF<br> format.</p> <p>The files all_trees.emf.gz and all_homologies.tsv.gz contain the phylogenetic<br> trees (all_trees.emf.gz, in the EMF alignment format) and the derived pairwise<br> orthologies and paralogies (all_homologies.tsv.gz, in tabular format).</p> <p>From the phylogenetic trees, sets of 1-to-1 orthologues across all 26 species<br> were extracted (pure_one2one_orthologies.txt).&nbsp; Sets of orthologues that span<br> all the species but include paralogues were reduced to 1 copy per species<br> using gene order conservation and sequence similarity. This extra dataset is<br> available in promoted_one2one_orthologies.txt</p> <p>The main Ensembl entry point for tuatara is:<br> &nbsp; http://www.ensembl.org/Sphenodon_punctatus/</p> <p>This work is supported by Ngatiwai iwi, Allan Wilson Centre, University of<br> Otago, New Zealand Genomics Limited, Illumina, National eScience<br> Infrastructure (NeSI NZ).</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Data for manuscript "Synteny identifies reliable orthologs for phylogenomics and comparative genomics of the Brassicaceae"

<p>Data and code for manuscript &quot;Synteny identifies reliable orthologs for phylogenomics and comparative genomics of the Brassicaceae&quot;. Preprint available at bioRxiv: https://doi.org/10.1101/2022.09.07.506897.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

KEGG Orthology annotation of Avena sativa (cv. Sang) proteins

<p>KofamScan v1.3.0 was used to assign KEGG Orthologs (KOs) to the <em>Avena sativa</em> (cv. Sang) proteins.</p> <p>exec_annotation Asativa_sang.v1.1.aa.fa -o Asativa_sang.v1.1.aa.ko.txt --profile profiles/eukaryote.hal --ko-list ko_list</p> <p>The Oat proteins (Asativa_sang.v1.1.aa.fa) were downloaded from https://doi.org/10.5447/ipk/2022/2<br> The KOfam database was downloaded from https://www.genome.jp/ftp/db/kofam/archives/2021-12-01/</p>

opencc-by-4.0Jul 2023View details →
dryad40/100

Finding orthologs for aminoacyl tRNA synthetases in parasitic plants

<p>Eukaryotic nuclear genomes often encode distinct sets of protein translation machinery for function in the cytosol vs. organelles (mitochondria and plastids). This phenomenon raises questions about why multiple translation systems are maintained even though they are capable of comparable functions, and whether they evolve differently depending on the compartment where they operate. These questions are particularly interesting in land plants because translation machinery, including aminoacyl-tRNA synthetases (aaRS), is often dual-targeted to both the plastids and mitochondria. These two organelles have quite different metabolisms, with much higher rates of translation in plastids to supply the abundant, rapid-turnover proteins required for photosynthesis. Previous studies have indicated that plant organellar aaRS evolve more slowly compared to mitochondrial aaRS in other eukaryotes that lack plastids. Thus, we investigated the evolution of nuclear-encoded organellar and cytosolic translation machinery across a broad sampling of angiosperms, including non-photosynthetic (heterotrophic) plant species with reduced rates of plastid gene expression to test the hypothesis that translational demands associated with photosynthesis constrain the evolution of bacterial-like enzymes involved in organellar tRNA metabolism. Remarkably, heterotrophic plants exhibited wholesale loss of many organelle-targeted aaRS and other enzymes, even though translation still occurs in their mitochondria and plastids. These losses were often accompanied by apparent retargeting of cytosolic enzymes and tRNAs to the organelles, sometimes preserving aaRS-tRNA charging relationships but other times creating surprising mismatches between cytosolic aaRS and mitochondrial tRNA substrates. Our findings indicate that the presence of a photosynthetic plastid drives the retention of specialized systems for organellar tRNA metabolism.</p>

opencc-zeroAug 2023View details →
dryad40/100

Finding orthologs for aminoacyl tRNA synthetases in parasitic plants

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad40/100

Immune-associated orthologous genes across primates

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad40/100

Different orthology inference algorithms generate similar predicted orthogroups among Brassicaceae species

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad36/100

Total Ortholog Median Matrix (TOMM): an alternative unsupervised approach for phylogenomics based on evolutionary distance between protein coding genes

<p>The increasing number of available genomic data allowed the development of phylogenomic analytical tools. Current methods compile information from single gene phylogenies, whether based on topologies or multiple sequence alignments. Generally, phylogenomic analyses elect gene families or genomic regions to construct phylogenomic trees. Here, we presented an alternative approach for Phylogenomics, named TOMM (Total Ortholog Median Matrix), to construct a representative phylogram composed by amino acid distance measures of all pairwise ortholog protein sequence pairs from desired species inside a group of organisms. The procedure is divided two main steps, (1) ortholog detection and (2) creation of a matrix with the median amino acid distance measures of all pairwise orthologous sequences. We tested this approach within three different group of organisms: Kinetoplastida protozoa, hematophagous Diptera vectors and Primates. Our approach was robust and efficacious to reconstruct the phylogenetic relationships for the three groups. Moreover, novel branch topologies could be achieved, providing insights about some phylogenetic relationships between some taxa.</p>

opencc-zeroDec 2020View details →
dryad36/100

Artifactual orthologs and the need for diligent data exploration in complex phylogenomic datasets: A museomic case study from the Andean flora

<p>The Andes mountains of western South America are a globally important biodiversity hotspot, yet there is a paucity of resolved phylogenies for plant clades from this region. Filling an important gap to our understanding of the World's richest flora, we present the first phylogeny of <em>Freziera</em> (Pentaphylacaceae), an Andean-centered, cloud forest radiation. Our dataset was obtained via hybrid-enriched target sequence capture of Angiosperms353 universal loci for 50 of the ca. 75 spp., obtained almost entirely from herbarium specimens. We identify high phylogenomic complexity in <em>Freziera</em>, including a significant proportion of paralogous loci and a high degree of gene tree discordance. Via gene tree filtering, by-eye observation of gene trees, and detailed examination of warnings from recently improved assembly pipelines, we identified that cryptic paralogs (i.e., the presence of only one copy of a multi-copy gene due to assembly errors) were a major source of gene tree heterogeneity that had a negative impact on phylogenetic inference and support. These cryptic paralogs likely result from limitations in data collection that are common in museomics, combined with a history of genome duplication; they may be common in plant phylogenomic datasets. After accounting for cryptic paralogs as source of gene tree error, we identified a significant, but non-specific signal of introgression using Patterson's D and f4 statistics. Despite phylogenomic complexity, we were able to resolve <em>Freziera</em> into nine well-supported subclades whose histories have been shaped by myriad evolutionary processes, including incomplete lineage sorting, historical gene flow, and gene duplication. Our results highlight the complexities of plant phylogenomics, and point to the need to test for multiple sources of gene tree discordance via careful examination of empirical datasets.</p>

opencc-zeroJan 2024View details →
zenodo36/100

Dataset associated to "Annotation matters: the effect of structural gene annotation on orthology"

<p>Dataset including the input and output files in "Annotation matters: the effect of structural gene annotation on orthology".</p> <ul> <li>Input proteomes for OMA and their corresponding splice files are in the OMAproteomes zipped folder. The OMA results for each annotation method are in the zipped folders with the method name (e.g. Augustus.zip).</li> <li>The fasta files (proteomes) are the same for OrthoFinder input in the cases of UniProt and Augustus (as they only have one isoform per gene). In these cases, the OrthoFinder folders (e.g. OFUniProt.zip), include the OrthoFinder output for that proteomes set. For Ensembl and NCBI, given the different approach each orthology method follows, the specific orthofinder proteomes are also included in the OrthoFinder (OF) zipped folder (e.g. OFtopNCBI.zip), in their corresponding primary_transcripts subfolder.&nbsp;</li> <li>The folder GSTDBenchmarOutput.zip contains the results from the Generalized Species Tree Discordance Benchmark.</li> <li>topNCBI/topEnsembl correspond to the original proteomes downloaded from the databases.</li> <li>priNCBI/primEnsembl correspond to the proteomes sets which include only the genes found on the primary assembly (reference sequences).</li> <li>For the species code to species name correspondance, please check the Code-Species.csv file.</li> </ul>

opencc-by-4.0Mar 2024View details →
dryad36/100

Concatenated amino acid (AA) phylogenetic dataset of nuclear gene orthologs for Ephydroidea (Diptera)

<p>The schizophoran superfamily Ephydroidea (Diptera: Cyclorrhapha) includes eight families, ranging from the well-known vinegar flies (Drosophilidae) and shore flies (Ephydridae), to several small, relatively unusual groups, the phylogenetic placement of which has been particularly challenging for systematists. Extraordinary diversity in life histories, feeding habits, and morphology are hallmarks of fly biology, and the Ephydroidea are no exception. Extreme specialization can lead to "orphaned" taxa with no clear evidence for their phylogenetic position. To resolve relationships among a diverse sample of Ephydroidea, including the highly modified flies in the families Braulidae and Mormotomyiidae, we conducted phylogenomic sampling. Using exon capture from Anchored Hybrid Enrichment and transcriptomics to obtain 320 orthologous nuclear genes sampled for 32 species of Ephydroidea and 11 outgroups, we evaluate a new phylogenetic hypothesis for representatives of the superfamily. These data strongly support monophyly of Ephydroidea with Ephydridae as an early branching radiation and the placement of Mormotomyiidae as a family-level lineage sister to all remaining families. We confirm the placement of Cryptochetidae as a sister taxon to a large clade containing both Drosophilidae and Braulidae – the latter a family of honeybee ectoparasites. Our results reaffirm that sampling of both taxa and characters is critical in hyperdiverse clades and that these factors have a major influence on phylogenomic reconstruction of the history of the schizophoran fly radiation.</p>

opencc-zeroSep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record