Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

24

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

24 results for “Phylogenetic profile”

Learn how ShareScore rates datasets ↗
edi48/100

Inventory of High-resolution phylogenetic profiles of the planktonic microbial communities (via 16S and 18S rRNA gene amplicons) from Shark River Slough and Taylor Slough, Everglades National Park (FCE LTER), Florida, USA, 2017 - ongoing

Planktonic microbial communities mediate many vital biogeochemical processes in wetland ecosystems, yet compared to other aquatic ecosystems, like oceans, lakes, rivers, or estuaries, they remain relatively underexplored. Our study site, the Florida Everglades (USA)—a vast iconic wetland consisting of a slow-moving system of shallow rivers connecting freshwater marshes with coastal mangrove forests and seagrass meadows—is a highly threatened model ecosystem for studying salinity and nutrient gradients, as well as the effects of sea level rise and saltwater intrusion. This dataset provides the first high-resolution phylogenetic profiles of planktonic bacterial and eukaryotic microbial communities (using 16S and 18S rRNA gene amplicons) from these environments. The dataset contains 16S and 18S rRNA data from 2017, and contains 16S rRNA data for monthly (2019) and quarterly water samples (2020-ongoing). The 2017 data are published in Laas et al. 2022. A detailed list of sequence data and their accession numbers in GenBank is provided and will be updated as more data are published. This data package is an inventory of sequence read archive (SRA) entries available through GenBank BioProject PRJNA525456 (at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA525456) and BioProject PRJNA1018945 (at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1018945). This data package is associated with the following publication: Laas, P., Ugarelli, K., Travieso, R., Stumpf, S., Gaiser, E. E., Kominoski, J. S., & Stingl, U. (2022). Water column microbial communities vary along salinity gradients in the Florida Coastal Everglades wetlands. Microorganisms, 10(2), 215. https://doi.org/10.3390/microorganisms10020215 Instead of citing this package, which is an inventory, please cite the original GenBank data or journal article, as appropriate. Citation guidance for the journal article is available on the respective publisher's website.

openCC (other)Feb 2024View details →
zenodo40/100

Phylogenetic analysis, morphological studies, element profiling, and muscarine detection reveal a new toxic Inosperma (Inocybaceae, Agaricales) species from tropical China

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo40/100

Phylogenetic profile of 100 annotated low complexity proteins against the Uniprot Reference Proteome dataset

<p>Phylogenetic profile of&nbsp;100 human proteins with characteristic compositional bias,&nbsp;previously recorded by&nbsp;Mier et al (2020)&nbsp;against the Uniprot&nbsp;Reference Proteome,&nbsp;containing a total of 11297 proteomes, excluding viruses. The counts for each protein correspond to homologs found in each proteome.&nbsp;</p> <p>Detailed description of included columns:</p> <p><strong>ref_proteome_identifier</strong>: The<strong>&nbsp;</strong>Uniprot&nbsp;Reference Proteome&nbsp;identifier</p> <p><strong>ncbi_taxid</strong>: The NCBI taxonomy ID</p> <p><strong>species_name</strong>: NCBI common name corresponding to taxonomy ID</p> <p><strong>species_code</strong>: internal species code composed of 9 characters</p> <p><strong>taxonomic domain</strong>: E/B/A for Eukaryota/Bacteria/Archaea classification of proteome</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

The Community Coevolution Model with application to the study of evolutionary relationships between genes based on phylogenetic profiles

<p>Organismal traits can evolve in a coordinated way, with correlated patterns of gains and losses reflecting important evolutionary associations. Discovering these associations can reveal important information about the functional and ecological linkages among traits. Phylogenetic profiles treat individual genes as traits distributed across sets of genomes and can provide a fine-grained view of the genetic underpinnings of evolutionary processes in a set of genomes. Phylogenetic profiling has been used to identify genes that are functionally linked, and to identify common patterns of lateral gene transfer in microorganisms. However, comparative analysis of phylogenetic profiles and other trait distributions should take into account the phylogenetic relationships among the organisms under consideration.</p> <p>Here we propose the Community Coevolution Model (CCM), a new coevolutionary model to analyze the evolutionary associations among traits, with a focus on phylogenetic profiles. In the CCM, traits are considered to evolve as a community with interactions, and the transition rate for each trait depends on the current states of other traits. Surpassing other comparative methods for pairwise trait analysis, CCM has the additional advantage of being able to examine multiple traits as a community to reveal more dependency relationships. We also develop a simulation procedure to generate phylogenetic profiles with correlated evolutionary patterns that can be used as benchmark data for evaluation purposes.</p> <p>A simulation study demonstrates that CCM is more accurate than other methods including the Jaccard Index and three tree-aware methods. The parameterization of CCM makes the interpretation of the relations between genes more direct, which leads to Darwin's scenario being identified easily based on the estimated parameters. We show that CCM is more efficient and fits real data better than other methods resulting in higher likelihood scores with fewer parameters. An examination of 3786 phylogenetic profiles across a set of 659 bacterial genomes highlights linkages between genes with common functions, including many patterns that would not have been identified under a non-phylogenetic model of common distribution. We also applied the CCM to 44 proteins in the well-studied Mitochondrial Respiratory Complex I and recovered associations that mapped well onto the structural associations that exist in the complex.</p>

opencc-zeroAug 2022View details →
zenodo36/100

Data supplementing the article "Boosting DNA metabarcoding for biomonitoring with phylogenetic estimation of OTUs' ecological profiles" F. Keck, V. Vasselon, F. Rimet, A. Bouchez, and M. Kahlert submitted to Molecular Ecology Resources journal

<p>These data supplement the article &quot;Enhancing DNA metabarcoding for biomonitoring with phylogenetic estimation of OTUs&#39; ecological profiles&quot; F. Keck, V. Vasselon, F. Rimet, A. Bouchez, and M. Kahlert&nbsp; submitted to Molecular Ecology Resources journal</p> <p>The directory contains the following files:</p> <p><strong>278 (139 x 2 replicates) samples fastq files.rar&nbsp;</strong>- contains the 278 fastq files provided by the sequencing platform with demultiplexed and contig DNA reads&nbsp;corresponding to the 139 samples with 2 sequencing replicates (A and B).</p> <p><strong>Counts_diatoms.xlsx&nbsp;</strong>- contains the morphological inventories with species list (Omnidia code) and valve abundances&nbsp;for the 139 samples.</p> <p><strong>Sites_list.xlsx&nbsp;</strong>- contains information regarding the 139 samples, including: River name, GPS coordinates, code used for molecular analysis and corresponding to sequencing&nbsp;fastq names.</p>

opencc-by-4.0Oct 2017View details →
zenodo36/100

Supplementary materials for the manuscript entitled "Comprehensive Identification, Phylogenetic Analysis and Expression Profiling of Multicopper Oxidase Genes in Maize"

<p><strong>Table S1. </strong>Gene IDs and names of maize and Arabidopsis MCOs</p> <p><strong>Figure S1.</strong> Phylogenetic tree of protein sequences of multicopper oxidase (MCO) of maize and Arabidopsis with rectangular layout. Branches are not cladogram-transformed. Shown in nodes are bootstrap support. The tree is polar layout, and branches are cladogram-transformed. Maize IDs are labeled in brown. Clades representative of<em> SKS</em>, <em>LAC</em> and <em>AAO</em> are labeled in blue, red and green, respectively. The tree is reconstructed by PhyML with 1000 bootstrap replicates.</p> <p><strong>Figure S2. </strong>Chromosome map depicting location of maize <em>MCO</em>s, with gene IDs shown in the map. Figure legend is the same as Figure 2.</p> <p><strong>Figure S3. </strong>Phylogenetic tree of maize and Arabidopsis <em>AAO</em>s. The trees are rectangular layout. The tree is reconstructed by PhyML with 1000 bootstrap replicates. Shown in nodes are bootstrap support.</p>

opencc-by-4.0May 2019View details →
dryad36/100

The Community Coevolution Model with application to the study of evolutionary relationships between genes based on phylogenetic profiles

Open the record for dataset details and reuse information.

publicAug 2022View details →
zenodo32/100

Fig. 31. Petiole, left profile view. A in A phylogenetic analyis of ant morphology (Hymenoptera: Formicidae) with special reference to the poneromoprh subfamilies

Fig. 31. Petiole, left profile view. A. Adetomyrma sp1. B. Ponera pennsylvanica. Abbreviations: IIS, sternum of petiole; IIIS, sternum of third abdominal segment; IT, tergum of propodeum; IIT, tergum of petiole; IIIT, tergum of third abdominal segment.

opennotspecifiedDec 2011View details →
zenodo32/100

Figure 6 in A novel epidermal gland type in lizards (α-gland): structural organization, histochemistry, protein profile and phylogenetic origins

Figure 6. Transmission electron microscopy of the epithelium from the femoral area of Tropidurus catalanensis. A, general view of the glandular epithelium showing major skin layers and abundance of vesicles inside glandular cells. B–G, magnified view of glandular tissue showing: B, clusters of melanin granules and numerous vesicles; C, aggregations of four different types of vesicles (V1, V2, V3 and V4); D, E, Golgi apparatuses amid secretory vesicles; F, autophagocytic events (arrowheads indicate membrane projections); G, iridophores present in the apical portion of the dermis. Legend: Go, Golgi complex; Ir, iridophore; Me, melanin granule; Nu, nucleus; V1, vesicles type 1; V2, vesicles type 2; V3, vesicles type 3; V4, vesicles type 4.

opennotspecifiedJul 2021View details →
zenodo32/100

Figure 3 in A novel epidermal gland type in lizards (α-gland): structural organization, histochemistry, protein profile and phylogenetic origins

Figure 3. Maderson &amp; Chiu's (1970) model of epidermal gland evolution. Steps of the model are briefly described in items A–G.

opennotspecifiedJul 2021View details →
zenodo32/100

Figure 2 in A novel epidermal gland type in lizards (α-gland): structural organization, histochemistry, protein profile and phylogenetic origins

Figure 2. Histological structure of unspecialized skin and epidermal glands. Diagrammatic representation of the unspecialized squamate skin and major epidermal gland types, illustrating their respective secretion mechanisms as hypothesized by Maderson (1972).

opennotspecifiedJul 2021View details →
zenodo32/100

Figure 5. A in A novel epidermal gland type in lizards (α-gland): structural organization, histochemistry, protein profile and phylogenetic origins

Figure 5. A, unspecialized scales from the pre-cloacal flap of a male Stenocercus caducus (MZUSP-R 82815). B, unspecialized scales from the femoral area of a female Tropidurus chromatops (MZUSP-R 106266). C, scales with α-glands from the femoral area of male T. chromatops (MZUSP-R 106263). Note that β-keratin layers are not present in A and C because they were lost during sample preparation. D, unspecialized scale (stage I) from the humeral region of a female T. chromatops (MZUSP-R 106266). E, detail of a scale with α-gland (stage IV) from the femoral area of a male T. xanthochilus (MZUSP-R 106342). F, detail of a scale with α-gland (stage V) from the pre-cloacal flap of a male T. xanthochilus (MZUSP-R 106342). G, detail of a scale with α-gland (stage VI) from the femoral areal of a male T. xanthochilus (MZUSP-R 106336). H, part of the inner generation of an α-gland from the femoral area of a male Plica plica (MTR 18918) showing the glandular stratum with a large number of secretory vesicles (indicated with an asterisk). I, part of the inner generation of the α-gland from the femoral area of a male T. catalanensis (MZUSP-R 106470) with melanophores transferring melanin granules into glandular cells and numerous melanin granules accumulated in their cytoplasm. J, glandular tissue of the outer generation

opennotspecifiedJul 2021View details →
zenodo32/100

Figure 9 in A novel epidermal gland type in lizards (α-gland): structural organization, histochemistry, protein profile and phylogenetic origins

Figure 9. Ancestral state reconstructions of α-gland/flash mark related characters of tropidurid lizards.

opennotspecifiedJul 2021View details →
zenodo32/100

Figure 4 in A novel epidermal gland type in lizards (α-gland): structural organization, histochemistry, protein profile and phylogenetic origins

Figure 4. Ventral view of (A) a male Tropidurus chromatops Harvey &amp; Gutberlet, 1998 (MHNC-R 3018) from ~30 km W Florida, Santa Cruz, Bolivia, (B) a male T. melanopleurus Boulenger, 1902 (IBIGEO-R 5331) from Aguas Blancas, Salta, Argentina, (C) a female T. xanthochilus Harvey &amp; Gutberlet, 1998 (MZUSP-R 106321) from Santo Antônio do Leverger, Mato Grosso, Brazil, and (D) a male T. etheridgei Cei, 1982 (AMNH-R 176273) from Filadelfia, Boquerón, Paraguay, illustrating the location and coloration of flash marks observed (or not) on the ventral body of tropidurines. Black squares in (A) indicate body areas from which we collected skin samples for histological examination. In (D) the yellow coloration covering the background of the black flash-marks of T. etheridgei might either represent a transient ontogenetic state or an instance in which a yellow background persists throughout life.

opennotspecifiedJul 2021View details →
zenodo32/100

Figure 2. Lithological profile. A in Osteology and phylogenetic relationships of Ligabuesaurus leanzai (Dinosauria: Sauropoda) from the Early Cretaceous of the Neuquén Basin, Patagonia, Argentina

Figure 2. Lithological profile. A schematic log of the the lower section of the Cullin Grande Member (Bajada del Agrio Group, Lohan Cura Formation, Lower Cretaceous, Albian) that outcrops at the Cerro de los Leones locality (modified from Martinelli et al., 2007). Abbreviations: CS, crevasse channel; FF, floodplain fines; FL, fossiliferous level; LA, lateral accretion; LS, laminated sand sheets; LV, levee; SB, sandy bedforms. Architectural element codes follow Miall (1996).

opennotspecifiedNov 2022View details →
dryad28/100

Data from: PhyloBayes MPI: phylogenetic reconstruction with infinite mixtures of profiles in a parallel environment

Modeling across site variation of the substitution process is increasingly recognized as important for obtaining more accurate phylogenetic reconstructions. Both finite and infinite mixture models have been proposed, and have been shown to significantly improve on classical single-matrix models. Compared to their finite counterparts, infinite mixtures have a greater expressivity. However, they are computationally more challenging. This has resulted in practical compromises in the design of infinite mixture models. In particular, a fast but simplified version of a Dirichlet process model over equilibirum frequency profiles implemented in PhyloBayes (Lartillot et al, 2007) has often been used in recent phylogenomics studies, while more refined model structures, more realistic and empirically more fit, have been practically out of reach. We introduce an Message Passing Interface (MPI) version of PhyloBayes, implementing the Dirichlet process mixture models as well as more classical empirical matrices and finite mixtures. The parallelization is made efficient thanks to the combination of two algorithmic strategies: a partial Gibbs sampling update of the tree topology, and the use of a truncated stick-breaking representation for the Dirichlet process prior. The implementation shows close to linear gains in computational speed for up to 64 cores, thus allowing faster phylogenetic reconstruction under complex mixture models. PhyloBayes MPI is freely available from our website www.phylobayes.org.

opencc-zeroDec 2012View details →
dryad28/100

Data from: An improved hypergeometric probability method for identification of functionally linked proteins using phylogenetic profiles

Predicting functions of proteins and alternatively spliced isoforms encoded in a genome is one of the important applications of bioinformatics in the post-genome era. Due to the practical limitation of experimental characterization of all proteins encoded in a genome using biochemical studies, bioinformatics methods provide powerful tools for function annotation and prediction. These methods also help minimize the growing sequence-to-function gap. Phylogenetic profiling is a bioinformatics approach to identify the influence of a trait across species and can be employed to infer the evolutionary history of proteins encoded in genomes. Here we propose an improved phylogenetic profile-based method which considers the co-evolution of the reference genome to derive the basic similarity measure, the background phylogeny of target genomes for profile generation and assigning weights to target genomes. The ordering of genomes and the runs of consecutive matches between the proteins were used to define phylogenetic relationships in the approach. We used Escherichia coli K12 genome as the reference genome and its 4195 proteins were used in the current analysis. We compared our approach with two existing methods and our initial results show that the predictions have outperformed two of the existing approaches. In addition, we have validated our method using a targeted protein-protein interaction network derived from protein-protein interaction database STRING. Our preliminary results indicates that improvement in function prediction can be attained by using coevolution-based similarity measures and the runs on to the same scale instead of computing them in different scales. Our method can be applied at the whole-genome level for annotating hypothetical proteins from prokaryotic genomes.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Dissecting signal and noise in diatom chloroplast protein encoding genes with phylogenetic information profiling

Open the record for dataset details and reuse information.

publicJul 2016View details →
dryad28/100

Data from: PhyloBayes MPI: phylogenetic reconstruction with infinite mixtures of profiles in a parallel environment

Open the record for dataset details and reuse information.

publicApr 2013View details →
dryad28/100

Data from: An improved hypergeometric probability method for identification of functionally linked proteins using phylogenetic profiles

Open the record for dataset details and reuse information.

publicJun 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record