Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

121

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

121 results for “orthology”

Learn how ShareScore rates datasets ↗
zenodo36/100

SonicParanoid2: fast, accurate, and comprehensive orthology inference with machine learning and language models

<p>This repository contains the documentation, test datasets and scripts used in the following study:</p> <p>"SonicParanoid2: fast, accurate, and comprehensive orthology inference with machine learning and language models"<br><br>- `sonic-manuscript-master.zip` contains all the scripts to reproduce the study, including those for generating the figures and tables included in the manuscript.</p> <p>- <a href="../api/records/11361985/draft/files/sonicparanoid2.wiki.tar.xz/content" target="_blank" rel="noopener noreferrer">sonicparanoid2.wiki.tar.xz</a> contains a snapshot fo the wiki for SonicParanoid2 as of May 30, 2024</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Theropithecus gelada and Microcebus murinus TBC1D3 Orthologous Sequence

<p>Genomic Sequence of TBC1D3 orthologous regions in Theropithecus gelada hap1 and hap2 used for Guitart et al.2024 journal article: "<strong>Independent expansion, selection and hypervariability of the TBC1D3 gene family in humans</strong>"&nbsp;</p> <p>https://www.biorxiv.org/content/10.1101/2024.03.12.584650v1</p>

opencc-by-4.0Jul 2024View details →
dryad36/100

Genetic dissection of triplicated chromosome 21 orthologs yields varying skeletal traits in Down syndrome model mice

<p><span class="normaltextrun">Down syndrome (DS) phenotypes result from triplicated genes, but effects of three copy genes are not well known. A mouse mapping panel genetically dissecting human chromosome 21 (Hsa21) syntenic regions was used to investigate the contributions and interactions of triplicated Hsa21 orthologous genes on mouse chromosome 16 (Mmu16) on skeletal phenotypes. Skeletal structure and mechanical properties were assessed in femurs of male and female Dp9Tyb, Dp2Tyb, Dp3Tyb, Dp4Tyb, Dp5Tyb, Dp6Tyb, Ts1Rhr, and Dp1Tyb;<em>Dyrk1a</em><sup>+/+/-</sup> mice. Dp1Tyb mice, with the entire Hsa21 homologous region of Mmu16 triplicated, display bone deficits similar to those of humans with DS and served as a baseline for other strains in the panel. Bone phenotypes varied based on triplicated gene content, sex, and bone compartment. Three copies of <em>Dyrk1a</em> played a sex-specific, essential role in trabecular deficits and may interact with other genes to influence cortical deficits related to DS. Triplicated genes in Dp9Tyb and Dp2Tyb mice improved some skeletal parameters. As triplicated genes can both improve and worsen bone deficits, it is important to understand the interaction between and molecular mechanisms of skeletal alterations affected by these genes.</span></p>

opencc-zeroApr 2023View details →
zenodo36/100

Single-copy orthologous genes used for Ricefish phylogeny

<p>Ortholog set</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;We generated a reference set consisting of 8390 single-copy protein-coding genes derived from OrthoDB v.9.1&nbsp;(Waterhouse et al., 2013)&nbsp;available for the following species:&nbsp;<em>Austrofundulus limnaeus, Centrocoris variegatus, Fundulus heteroclitus, Kryptolebias marmoratus, Nothobranchius furzeri, Oryzias latipes, O. melastigma, Poecilia formosa, P. latipinna ,P. mexicana, P. reticulata</em>&nbsp;and&nbsp;<em>Xiphophorus maculatus&nbsp;</em>(NCBI Accession numbers in Table S7). The hierarchical split was set to Actinopterygii (ID 7898). We used the script &ldquo;make-ogs-corresponding.pl&rdquo; to check for inconsistencies between the amino acid sequences and the corresponding nucleotide sequences and removed 96 problematic genes (Tab. S7).&nbsp;</p> <p>Identification of orthologs for transcripts and genome and&nbsp;alignment of single-copy genes</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Ortholog identification among 16 ricefish species and four outgroups (DS1, supplementary tables Tab. S1a) was carried out with Orthograph v0.7.1&nbsp;(Petersen et al., 2017). Forward search for candidate transcript was left at default. Best reciprocal hit: Ortholog candidate genes needed at least one hit in either&nbsp;<em>O. latipes</em>&nbsp;or&nbsp;<em>O. melastigma&nbsp;</em>and we allowed concatenation of hits if they met the criteria and did not overlap. Max-blast-searches were set to 50, blast-max-hits were also set to 50. &ldquo;U&rdquo; in the amino acid sequences was changed to &ldquo;X&rdquo; to avoid issues in downstream analysis. The results of the orthology prediction were summarized for all species using a custom perl script coming with the orthograph package. Sequences of only those orthologs with all species present were aligned using&nbsp;MAFFT v7.221 with the L-INS-I algorithm&nbsp;on amino acid level&nbsp;(Katoh &amp; Standley, 2013). 915 orthologs with outliers were identified according to Misof et al. 2014 and were subsequently removed from further analysis. We used the amino-acid alignments as blue print to generate corresponding nucleotide alignments with&nbsp;a modified version of Pal2Nal v14&nbsp;(Misof et al., 2014; Suyama et al., 2006). To check each amino acid alignment for ambiguously aligned regions, we ran ALISCORE v2.0 with the maximal number of possible sequence selected pairs to analyze (-r)&nbsp;(K&uuml;ck et al., 2010; Misof et al., 2014; Misof &amp; Misof, 2009). Sites which needed masking were cut out using ALICUT v2.3&nbsp;(K&uuml;ck, 2009)&nbsp;from the amino acid alignments and correspondingly also from the nucleotide alignments. For further analyses we only proceeded with the data set on nucleotide level.</p>

opencc-by-4.0Jul 2023View details →
dryad36/100

Total Ortholog Median Matrix (TOMM): an alternative unsupervised approach for phylogenomics based on evolutionary distance between protein coding genes

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad36/100

Concatenated amino acid (AA) phylogenetic dataset of nuclear gene orthologs for Ephydroidea (Diptera)

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad36/100

Supplementary data from: Lacewing-specific universal single-copy orthologs designed towards resolution of backbone phylogeny of Neuropterida

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad36/100

Genetic dissection of triplicated chromosome 21 orthologs yields varying skeletal traits in Down syndrome model mice

Open the record for dataset details and reuse information.

publicApr 2023View details →
dryad36/100

Artifactual orthologs and the need for diligent data exploration in complex phylogenomic datasets: A museomic case study from the Andean flora

Open the record for dataset details and reuse information.

publicJan 2024View details →
zenodo32/100

Phylogenetic comparative methods are problematic when applied to gene trees with speciation and duplication nodes: correcting for biases in testing the ortholog conjecture

<p>This repository contains &ldquo;manuscript_dunn.RData&rdquo; file, which is reproduced by using the files and scripts of Dunn et al. (Dunn CW, Zapata F, Munro C, Siebert S, Hejnol A (2018) Pairwise comparisons across species are problematic when analyzing functional genomic data. Proc Natl Acad Sci U S A 115: E409&ndash;E417. <a href="http://dx.doi.org/10.1073/pnas.1707515115">doi:10.1073/pnas.1707515115</a>).</p> <p>In this repository, we also supplied &ldquo;Data_TMRR_latest.rda&rdquo; file, containing the results generated by using our own scripts. Our scripts are available on GitHub: <a href="https://github.com/tbegum/Testing_the_ortholog_conjecture">https://github.com/tbegum/Testing_the_ortholog_conjecture</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2019View details →
dryad32/100

Data from: Mining from transcriptomes: 315 single-copy orthologous genes concatenated for the phylogenetic analyses of Orchidaceae

Phylogenetic relationships are hotspots for orchid studies with controversial standpoints. Traditionally, the phylogenies of orchids are based on morphology and subjective factors. Although more reliable than classic phylogenic analyses, the current methods are based on a few gene markers and PCR amplification, which are labor intensive and cannot identify the placement of some species with degenerated plastid genomes. Therefore, a more efficient, labor-saving and reliable method is needed for phylogenic analysis. Here, we present a method of orchid phylogeny construction using transcriptomes. Ten representative species covering five subfamilies of Orchidaceae were selected, and 315 single-copy orthologous genes extracted from the transcriptomes of these organisms were applied to reconstruct a more robust phylogeny of orchids. This approach provided a rapid and reliable method of phylogeny construction for Orchidaceae, one of the most diversified family of angiosperms. We also showed the rigorous systematic position of holomycotrophic species, which has previously been difficult to determine because of the degenerated plastid genome. We concluded that the method presented in this study is more efficient and reliable than methods based on a few gene markers for phylogenic analyses, especially for the holomycotrophic species or those whose DNA sequences have been difficult to amplify. Meanwhile, a total of 315 single-copy orthologous genes of orchids are offered and more informative loci could be used in the future orchid phylogenetic studies.

opencc-zeroDec 2014View details →
zenodo32/100

Genome alignments from "Split-alignment of genomes finds orthologies more accurately"

<p>Here are the genome alignments resulting from &quot;Split-alignment of genomes finds orthologies more accurately&quot;, Genome Biology 2015, 16:106.</p> <ul> <li>In the terminology of that paper, these are 2-split, post-masked alignments.</li> <li>They are in MAF format (http://genome.ucsc.edu/FAQ/FAQformat.html), with &quot;p&quot; lines (http://last.cbrc.jp/doc/last-split.html).</li> <li>Alignments with high ambiguity have not been removed, so if you want unambiguously 1-to-1 alignments, remove those annotated with mismap &gt; 0.00001 or so.</li> <li>The chimp and orangutan alignments were made using the 500/30 gap costs suggested at the end of the paper.</li> </ul> <p>&nbsp;</p>

opencc-zeroMay 2015View details →
zenodo32/100

Curated protostome sequences to validate deuterostome specific orthologous groups

<p>This dataset contains protostome peptide sequences that have been used to test the validity of deuterostome specific orthologous groups. BLAST was used to find potential homologous protostome sequences that could invalidate the specificity of the deuterostome orthogroups in question.</p>

opencc-by-4.0Dec 2018View details →
dryad32/100

Data from: DISCOMARK: nuclear marker discovery from orthologous sequences using draft genome data

Open the record for dataset details and reuse information.

publicJul 2016View details →
dryad32/100

Tumour suppressor genes fasta files for orthologous groups and hierarchical orthologous groups

Open the record for dataset details and reuse information.

publicFeb 2021View details →
dryad32/100

Data from: Hemiptera phylogenomic resources: tree-based orthology prediction and conserved exon identification

Open the record for dataset details and reuse information.

publicMay 2020View details →
dryad32/100

Data from: Mining from transcriptomes: 315 single-copy orthologous genes concatenated for the phylogenetic analyses of Orchidaceae

Open the record for dataset details and reuse information.

publicJul 2016View details →
dryad32/100

One-to-one orthologs between Danaus plexippus and other insect species

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad28/100

Divergence time estimation of genus Tribolium by extensive sampling of highly conserved orthologs

<p><i>Tribolium castaneum</i>, the red flour beetle, is among the most well-studied eukaryotic genetic model organisms. <i>Tribolium</i> often serves as a comparative bridge from highly derived <i>Drosophila</i> traits to other organisms. Simultaneously, as a member of the most diverse order of metazoans, Coleoptera,  <i>Tribolium</i> informs us about innovations that accompany hyper diversity. However, understanding the tempo and mode of evolutionary innovation requires well-resolved, time-calibrated phylogenies, which are not available for <i>Tribolium</i>. The most recent effort to understand <i>Tribolium</i>phylogenetics used two mitochondrial and three nuclear markers. The study concluded that the genus may be paraphyletic and reported a broad range for divergence time estimates. Here we employ recent advances in Bayesian methods to estimate the relationships and divergence times among <i>Tribolium castaneum</i>, <i>T. brevicornis</i>, <i>T. confusum</i>, <i>T. freemani</i>, and <i>Gnatocerus cornutus </i>using 1368 orthologs conserved across all five species and an independent substitution rate estimate. We find that the most basal split within <i>Tribolium</i> occurred ~86 Mya [95% HPD 85.90–87.04 Mya] and that the most recent split was between <i>T. freemani</i> and <i>T. castaneum</i> at ~14 Mya [95% HPD 13.55-14.00]. Our results are consistent with broader phylogenetic analyses of insects and suggest that Cenozoic climate changes played a role in the <i>Tribolium </i>diversification.</p>

opencc-zeroFeb 2021View details →
dryad28/100

Data from: Integrating sequence evolution into probabilistic orthology analysis

Orthology analysis, that is, finding out whether a pair of homologous genes are orthologs — stemming from a speciation — or paralogs — stemming from a gene duplication - is of central importance in computational biology, genome annotation, and phylogenetic inference. In particular, an orthologous relationship makes functional equivalence of the two genes highly likely. A major approach to orthology analysis is to reconcile a gene tree to the corresponding species tree, (most commonly performed using the most parsimonious reconciliation, MPR). However, most such phylogenetic orthology methods infer the gene tree without considering the constraints implied by the species tree and, perhaps even more importantly, only allow the gene sequences to influence the orthology analysis through the a priori reconstructed gene tree. We propose a sound, comprehensive Bayesian Markov chain Monte Carlo-based method, DLRSOrthology, to compute orthology probabilities. It efficiently sums over the possible gene trees and jointly takes into account the current gene tree, all possible reconciliations to the species tree, and the, typically strong, signal conveyed by the sequences. We compare our method with PrIME-GEM, a probabilistic orthology approach built on a probabilistic duplication-loss model, and MRBAYESMPR, a probabilistic orthology approach that is based on conventional Bayesian inference coupled with MPR. We find that DLRSOrthology outperforms these competing approaches on synthetic data as well as on biological data sets and is robust to incomplete taxon sampling artifacts.

opencc-zeroDec 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record