Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
121
datasets available to search
ShareScore release 0.7.1
Dataset results
121 results for “orthology”
Data from: Identification of putative orthologous genes for the phylogenetic reconstruction of temperate woody bamboos (Poaceae: Bambusoideae)
The temperate woody bamboos (Arundinarieae) are highly diverse in morphology but lack a substantial amount of genetic variation. The taxonomy of this lineage is intractable, and the relationships within the tribe have not been well resolved. Recent studies indicated that this tribe could have a complex evolutionary history. Although phylogenetic studies of the tribe have been carried out, most of these phylogenetic reconstructions were based on plastid data, which provide lower phylogenetic resolution compared to nuclear data. In this study, we intended to identify a set of desirable nuclear genes for resolving the phylogeny of the temperate woody bamboos. Using two different methodologies, we identified 209 and 916 genes, respectively, as putative single copy orthologous genes. A total of 112 genes was successfully amplified and sequenced by next-generation sequencing technologies in five species sampled from the tribe. As most of the genes exhibited intra-individual allele heterozygotes (IIAHs), we investigated phylogenetic utility by reconstructing the phylogeny based on individual genes. Discordance among gene trees was observed and, to resolve the conflict, we performed a range of analyses using BUCKy and HybTree. While caution should be taken when inferring a phylogeny from multiple conflicting genes, our analysis indicated that 74 of the 112 investigated genes are potential markers for resolving the phylogeny of the temperate woody bamboos.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.
An Orthology Graph Database
<p>The graph database used in the study....</p> <p>More info will be added following the publication...</p>
Patterns of selection on tuatara orthologs
<p>PAML analyses of tuatara genome.</p> <p>File 1 = PAML results</p> <p>File 2. KEGG and GO analyses on PAML output</p>
C.difficile - WhatsGNU ortholog database
<p>WhatsGNU ortholog database for <em>Clostridioides difficile.</em></p>
S.aureus - WhatsGNU ortholog database
<p>WhatsGNU ortholog database for <em>Staphylococcus aureus</em>.</p>
Biosynthetic Gene Cluster Synteny - Orthologous Polyketide Synthases in Hypogymnia physodes, Hypogymnia tubulosa and Parmelia sulcata
<p>Supplementary Material of Publication</p>
Data from: Adaptive functional divergence of the warm temperature acclimation-related protein (WAP65) in fishes and the ortholog Hemopexin (HPX) in mammals
Open the record for dataset details and reuse information.
Data from: Integrating sequence evolution into probabilistic orthology analysis
Open the record for dataset details and reuse information.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
Open the record for dataset details and reuse information.
Data from: Identification of putative orthologous genes for the phylogenetic reconstruction of temperate woody bamboos (Poaceae: Bambusoideae)
Open the record for dataset details and reuse information.
Data from: Polymorphisms in a desaturase 2 ortholog associate with cuticular hydrocarbon and male mating success variation in a natural population of Drosophila serrata
Open the record for dataset details and reuse information.
Divergence time estimation of genus Tribolium by extensive sampling of highly conserved orthologs
Open the record for dataset details and reuse information.
A conserved transcriptional regulator governs fungal morphology in widely diverged species [ChIP-chip, Transcriptional regulation by Mit1 and orthologs]
GEO Series GSE32557. Saccharomyces cerevisiae. 10 samples. Type: Genome binding/occupancy profiling by genome tiling array.
The C. elegans ortholog of human selenium binding protein 1 is a pro-aging factor protecting against selenite toxicity
GEO Series GSE134196. Caenorhabditis elegans. 6 samples. Type: Expression profiling by high throughput sequencing.
Transcriptome sequencing of WT and chloroquine resistance transporter (TgCRT) ortholog-deficient Toxoplasma parasites
GEO Series GSE116539. Toxoplasma gondii. 4 samples. Type: Expression profiling by high throughput sequencing.
Effect of global Caspase-1 (Casp1) knockout (KO) on kidney gene expression in an orthologous mouse model of polycystic kidney disease (PKD)
GEO Series GSE207957. Mus musculus. 12 samples. Type: Expression profiling by high throughput sequencing.
Genome wide co-localization of Polycomb orthologs and their effects on gene expression in human fibroblasts
GEO Series GSE40740. Homo sapiens. 27 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
The ortholog of DDM1 is mainly required for CHG and CG methylation of heterochromatin and is involved in DRM2-mediated CHH methylation that targets mostly genic regions of the rice genome
GEO Series GSE81436. Oryza sativa Japonica Group. 17 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.
The Peptide Genomic Therapy Increases Antibacterial Immunity and Survival in Sepsis by Reprograming the Gene Orthologs of Human Immunodeficiencies in the Spleen and Lungs
GEO Series GSE308045. Mus musculus. 11 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.