Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

121

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

121 results for “orthology”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Identification of putative orthologous genes for the phylogenetic reconstruction of temperate woody bamboos (Poaceae: Bambusoideae)

The temperate woody bamboos (Arundinarieae) are highly diverse in morphology but lack a substantial amount of genetic variation. The taxonomy of this lineage is intractable, and the relationships within the tribe have not been well resolved. Recent studies indicated that this tribe could have a complex evolutionary history. Although phylogenetic studies of the tribe have been carried out, most of these phylogenetic reconstructions were based on plastid data, which provide lower phylogenetic resolution compared to nuclear data. In this study, we intended to identify a set of desirable nuclear genes for resolving the phylogeny of the temperate woody bamboos. Using two different methodologies, we identified 209 and 916 genes, respectively, as putative single copy orthologous genes. A total of 112 genes was successfully amplified and sequenced by next-generation sequencing technologies in five species sampled from the tribe. As most of the genes exhibited intra-individual allele heterozygotes (IIAHs), we investigated phylogenetic utility by reconstructing the phylogeny based on individual genes. Discordance among gene trees was observed and, to resolve the conflict, we performed a range of analyses using BUCKy and HybTree. While caution should be taken when inferring a phylogeny from multiple conflicting genes, our analysis indicated that 74 of the 112 investigated genes are potential markers for resolving the phylogeny of the temperate woody bamboos.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture

The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.

opencc-zeroDec 2015View details →
zenodo28/100

An Orthology Graph Database

<p>The graph database used in the study....</p> <p>More info will be added following the publication...</p>

opencc-by-4.0Jun 2024View details →
zenodo28/100

Patterns of selection on tuatara orthologs

<p>PAML analyses of tuatara genome.</p> <p>File 1 = PAML results</p> <p>File 2. KEGG and GO analyses on PAML output</p>

opencc-by-4.0Mar 2019View details →
zenodo28/100

C.difficile - WhatsGNU ortholog database

<p>WhatsGNU ortholog database for <em>Clostridioides difficile.</em></p>

opencc-by-4.0Aug 2024View details →
zenodo28/100

S.aureus - WhatsGNU ortholog database

<p>WhatsGNU ortholog database for <em>Staphylococcus aureus</em>.</p>

opencc-by-4.0Aug 2024View details →
zenodo28/100

Biosynthetic Gene Cluster Synteny - Orthologous Polyketide Synthases in Hypogymnia physodes, Hypogymnia tubulosa and Parmelia sulcata

<p>Supplementary Material of Publication</p>

opencc-by-4.0Aug 2023View details →
dryad28/100

Data from: Adaptive functional divergence of the warm temperature acclimation-related protein (WAP65) in fishes and the ortholog Hemopexin (HPX) in mammals

Open the record for dataset details and reuse information.

publicOct 2013View details →
dryad28/100

Data from: Integrating sequence evolution into probabilistic orthology analysis

Open the record for dataset details and reuse information.

publicJul 2015View details →
dryad28/100

Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture

Open the record for dataset details and reuse information.

publicMay 2016View details →
dryad28/100

Data from: Identification of putative orthologous genes for the phylogenetic reconstruction of temperate woody bamboos (Poaceae: Bambusoideae)

Open the record for dataset details and reuse information.

publicMar 2014View details →
dryad28/100

Data from: Polymorphisms in a desaturase 2 ortholog associate with cuticular hydrocarbon and male mating success variation in a natural population of Drosophila serrata

Open the record for dataset details and reuse information.

publicNov 2015View details →
dryad28/100

Divergence time estimation of genus Tribolium by extensive sampling of highly conserved orthologs

Open the record for dataset details and reuse information.

publicFeb 2021View details →
geo24/100

A conserved transcriptional regulator governs fungal morphology in widely diverged species [ChIP-chip, Transcriptional regulation by Mit1 and orthologs]

GEO Series GSE32557. Saccharomyces cerevisiae. 10 samples. Type: Genome binding/occupancy profiling by genome tiling array.

openGEO-OpenOct 2011View details →
geo24/100

The C. elegans ortholog of human selenium binding protein 1 is a pro-aging factor protecting against selenite toxicity

GEO Series GSE134196. Caenorhabditis elegans. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2019View details →
geo24/100

Transcriptome sequencing of WT and chloroquine resistance transporter (TgCRT) ortholog-deficient Toxoplasma parasites

GEO Series GSE116539. Toxoplasma gondii. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2019View details →
geo24/100

Effect of global Caspase-1 (Casp1) knockout (KO) on kidney gene expression in an orthologous mouse model of polycystic kidney disease (PKD)

GEO Series GSE207957. Mus musculus. 12 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2022View details →
geo24/100

Genome wide co-localization of Polycomb orthologs and their effects on gene expression in human fibroblasts

GEO Series GSE40740. Homo sapiens. 27 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenDec 2013View details →
geo24/100

The ortholog of DDM1 is mainly required for CHG and CG methylation of heterochromatin and is involved in DRM2-mediated CHH methylation that targets mostly genic regions of the rice genome

GEO Series GSE81436. Oryza sativa Japonica Group. 17 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenJun 2016View details →
geo24/100

The Peptide Genomic Therapy Increases Antibacterial Immunity and Survival in Sepsis by Reprograming the Gene Orthologs of Human Immunodeficiencies in the Spleen and Lungs

GEO Series GSE308045. Mus musculus. 11 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record