Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

15 results for “Compositional bias”

Learn how ShareScore rates datasets ↗
zenodo44/100

Supplementary Data Files for the paper "Intrinsically disordered compositional bias in proteins: Sequence traits, region clustering, and generation of hypothetical functional associations"

<div> <div> <div> <div> <p><strong>Supplementary data files relating to <a href="https://doi.org/10.1177/11779322241287485">https://doi.org/10.1177/11779322241287485.&nbsp;</a></strong></p> <p><strong><span>Suppl. File 1: Protein Family Clusters.</span></strong></p> <p><strong><span>Suppl. File 2: Cluster GO enrichments/depletions. </span></strong></p> <p><strong><span>Suppl. File 3: The raw ID-CBR data with annotations. </span></strong></p> <p><strong><span>Suppl. File 4: &shy;ID-CBR Cluster membership.</span></strong></p> <p><strong><span>Each file has an explanatory header.&nbsp;</span></strong></p> <p>&nbsp;</p> </div> </div> </div> </div>

opencc-by-4.0Oct 2024View details →
dryad36/100

Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches

<p><strong>Background</strong></p> <p>The application of reduced metagenomic sequencing approaches holds promise as a middle ground between targeted amplicon sequencing and whole metagenome sequencing approaches but has not been widely adopted as a technique. A major barrier to adoption is the lack of read simulation software built to handle characteristic features of these novel approaches. Reduced metagenomic sequencing (RMS) produces unique patterns of fragmentation per genome that are sensitive to restriction enzyme choice, and the non-uniform size selection of these fragments may introduce novel challenges to taxonomic assignment as well as relative abundance estimates.</p> <p><strong>Results</strong></p> <p>Through the development and application of simulation software, readsynth, we compare simulated metagenomic sequencing libraries with existing RMS data to assess the influence of multiple library preparation and sequencing steps on downstream analytical results. Based on read depth per position, readsynth achieved 0.79 Pearson's correlation and 0.94 Spearman's correlation to these benchmarks. Application of a novel estimation approach, fixed length taxonomic ratios, improved quantification accuracy of simulated human gut microbial communities when compared to estimates of mean or median coverage.</p> <p><strong>Conclusions</strong></p> <p>We investigate the possible strengths and weaknesses of applying the RMS technique to profiling microbial communities via simulations with readsynth. The choice of restriction enzymes and size selection steps in library prep are non-trivial decisions that bias downstream profiling and quantification. The simulations investigated in this study illustrate the possible limits of preparing metagenomic libraries with a reduced representation sequencing approach, but also allow for the development of strategies for producing and handling the sequence data produced by this promising application.</p>

opencc-zeroApr 2024View details →
zenodo36/100

DATA for Exploration of O-GlcNAc-transferase (OGT) glycosylation sites reveals a target sequence compositional bias

<p>Mass spectrometry data for identification of glycosylation sites in CBP ID3 and EWS LCRn</p> <p>Perl based implementation of glycosylation Site Predictor OGTcomPred</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad32/100

Psocodea Phylogenomic dataset from: Phylogenomics of parasitic and non-parasitic lice (Insecta: Psocodea): combining sequence data and Exploring compositional bias solutions in Next Generation Datasets

<p>This dataset includes all alignments used for the phylogenomic analysis of Psocodea. In this dataset, includes all result files of phylogenomic analyses completed. This includes maximum likelihood, astral, MCMCtree, quartet sampling, and all gene trees. Any relevant input files are included, and any materials are available upon request.</p> <p>The insect order Psocodea is a diverse lineage comprising both parasitic (Phthiraptera) and non-parasitic members (Psocoptera). The extreme age and ecological diversity of the group may be associated with major genomic changes, such as base compositional biases expected to affect phylogenetic inference. Divergent morphology between parasitic and non-parasitic members has also obscured the origins of parasitism within the order. We conducted a phylogenomic analysis on the order Psocodea utilizing both transcriptome and genome sequencing to obtain a data set of 2,370 orthologous genes. All phylogenomic analyses, including both concatenated and coalescent methods suggest a single origin of parasitism within the order Psocodea, resolving conflicting results from previous studies. This phylogeny allows us to propose a stable ordinal level classification scheme that retains significant taxonomic names present in historical scientific literature and reflects the evolution of the group as a whole. A dating analysis, with internal nodes calibrated by fossil evidence, suggests an origin of parasitism that predates the K-Pg boundary. Nucleotide compositional biases are detected in third and first codon positions and result in the anomalous placement of the Amphientometae as sister to Psocomorpha when all nucleotide sites are analyzed. Likelihood-mapping and quartet sampling methods demonstrate that base compositional biases can also have an effect on quartet-based methods.</p>

opencc-zeroSep 2020View details →
dryad32/100

Data from: Support for a clade of Placozoa and Cnidaria in genes with minimal compositional bias

The phylogenetic placement of the morphologically simple placozoans is crucial to understanding the evolution of complex animal traits. Here, we examine the influence of adding new genomes from placozoans to a large dataset designed to study the deepest splits in the animal phylogeny. Using site-heterogeneous substitution models, we show that it is possible to obtain strong support, in both amino acid and reduced-alphabet matrices, for either a sister-group relationship between Cnidaria and Placozoa, or for Cnidaria and Bilateria as seen in most published work to date, depending on the orthologues selected to construct the matrix. We demonstrate that a majority of genes show evidence of compositional heterogeneity, and that support for the Cnidaria+Bilateria clade can be assigned to this source of systematic error. In interpreting these results, we caution against a peremptory reading of placozoans as secondarily reduced forms of little relevance to broader discussions of early animal evolution.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Support for a clade of Placozoa and Cnidaria in genes with minimal compositional bias

Open the record for dataset details and reuse information.

publicNov 2018View details →
dryad32/100

Psocodea Phylogenomic dataset from: Phylogenomics of parasitic and non-parasitic lice (Insecta: Psocodea): combining sequence data and Exploring compositional bias solutions in Next Generation Datasets

Open the record for dataset details and reuse information.

publicSep 2020View details →
dryad32/100

Data from: Female-female competition leads to female-biased sex allocation and dimorphism in brood sex composition in a gall-forming aphid

Open the record for dataset details and reuse information.

publicNov 2018View details →
dryad28/100

Data from: Conflicting phylogenies for early land plants are caused by composition biases among synonymous substitutions

Plants are the primary producers of the terrestrial ecosystems that dominate much of the natural environment. Occurring approximately 480 MYA (Sanderson 2003; Kenrick et. al. 2012), the evolutionary transition of plants from an aquatic to a terrestrial environment was accompanied by several major developmental innovations. The freshwater charophyte ancestors of land plants have a haplobiontic life cycle with a single haploid multicellular stage, whereas land plants, which include the bryophytes (liverworts, hornworts, and mosses) and tracheophytes (also called vascular plants, namely, lycopods, ferns, and seed plants), exhibit a marked alternation of generations with a diplobiontic life-cycle with both haploid and diploid multicellular stages and where the embryo remains attached to, and is nourished by, the gametophyte (Haig 2008). The interjection of a multicellular diploid phase into the land plant life cycle was an important adaptation that enabled long-distance dispersal via mitotic spores where water-borne male gametes have restricted motility in dry terrestrial environments. Despite the similarity among land-plant life-cycles, they differ in one significant aspect: in the three bryophyte groups, the haploid gametophytic stage is the dominant vegetative stage, whereas in vascular plants the diploid sporophyte dominates. A common assumption, and one implied by the tradition of referring to bryophytes as "lower plants" - in contrast to the "higher plants", the tracheophytes - is that the bryophytes and their life-cycle are primitive (Kato and Akiyama 2005). However, without a strong phylogenetic hypothesis of land-plant relationships, it is not clear which (if either) of the gametophyte or sporophyte was the dominant ancestral vegetative state present in the earliest land plants (Renzaglia et al. 2007; Qiu et al. 2012). Early land plants have a relatively poor fossil record with few intermediate forms (Kenrick and Crane 1997; Wellman et al. 2003; Clarke et al. 2011), so most of the evidence for early land plant evolution has been based upon the patterns of morphological change that are implied by phylogenetic trees of relationships among extant land plant and algal groups. In this context, several recent studies based on large molecular data sets have converged upon a phylogenetic solution to land plant origins wherein tracheophytes are derived from bryophyte ancestors (Karol et al. 2001; Qiu et al. 2006; Gao et al. 2010; Karol et al. 2010; Chang and Graham 2011). In this hypothesis, the three bryophyte groups, namely liverworts, mosses, and hornworts, diverged sequentially and form a paraphyletic group with the hornworts sister to the tracheophytes. This phylogeny supports an intuitively elegant evolutionary trajectory whereby plants increased in morphological complexity from single-celled algae to seed plants via bryophyte intermediates (Karol et al. 2001; McCourt et al. 2004). Specifically, it implies that the gametophyte-dominant bryophyte life-cycle was ancestral among land plants and that the complex modular growth form of the vascular plant sporophyte evolved from the simplistic bryophyte sporophyte that consists only of a single growth module (Kato and Akiyama 2005; Barthélémy and Caraglio 2007).

opencc-zeroDec 2013View details →
dryad28/100

Data from: Mitochondrial phylogenomics of early land plants: mitigating the effects of saturation, compositional heterogeneity, and codon-usage bias

Phylogenetic analyses using concatenation of genomic-scale data have been seen as the panacea to resolving the incongruences among inferences from few or single genes. However, phylogenomics may also suffer from systematic errors, due to the, perhaps cumulative, effects of saturation, among-taxa compositional (GC content) heterogeneity, or codon-usage bias plaguing the individual nucleotide loci that are concatenated. Here we provide an example of how these factors affect the inferences of the phylogeny of early land plants based on mitochondrial genomic data. Mitochondrial sequences evolve slowly in plants and hence are thought to be suitable for resolving deep relationships. We newly assembled mitochondrial genomes from 20 bryophytes, complemented these with 40 other streptophytes (land plants plus algal outgroups), compiling a data matrix of 60 taxa and 41 mitochondrial genes. Homogeneous analyses of the concatenated nucleotide data resolve mosses as sister-group to the remaining land plants. However, the corresponding translated amino acid data support the liverwort lineage in this position. Both results receive weak to moderate support in maximum likelihood analyses, but strong support in Bayesian inferences. Tests of alternative hypotheses using either nucleotide or amino-acid data provide implicit support for the respective optimal topologies. By analyzing the nucleotide data, we found that the 3rd codon positions are more saturated than the 1st and 2nd codon positions, and excluding these from the analyses leads to a topology congruent with that obtained using amino-acid data. Further, we determined that land plant lineages differ in their nucleotide composition, and in their usage of synonymous codon variants. Composition heterogeneous Bayesian analyses employing a non-stationary model that accounts for variation in among-lineage composition, and inferences from degenerated nucleotide data that avoids the effects of synonymous mutations that underlie codon-usage bias, again recovered liverworts being sister to the remaining land plants. These analyses indicate that the discrepancy between the nucleotide-based and the amino acid-based trees is caused by the lineage specific, parallel compositional bias, or synonymous mutations driving codon-usage bias, as well as saturation in the 3rd codon positions. While genomic data may generate highly supported phylogenetic trees, these inferences may be artifacts. We suggest that phylogenomic analyses should assess the possible impact of potential biases through comparisons of protein coding gene data and their amino-acids translations, by analyzing data modeling compositional bias, and by excluding nucleotide noisy signals due to saturation or codon-usage bias. We caution against relying on any one presentation of the data (nucleotide or amino acid) or any one type of analysis even when analyzing large-scale data sets, no matter how well-supported, without fully exploring the effects of substitution models.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Conflicting phylogenies for early land plants are caused by composition biases among synonymous substitutions

Open the record for dataset details and reuse information.

publicJan 2014View details →
dryad28/100

Data from: Mitochondrial phylogenomics of early land plants: mitigating the effects of saturation, compositional heterogeneity, and codon-usage bias

Open the record for dataset details and reuse information.

publicJul 2014View details →
geo16/100

Systematic Bias Introduced by Ficoll-Based Isolation in AML Sample Composition and Downstream Analyses

GEO Series GSE313042. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2025View details →
geo12/100

Splicing factors and chromatin organization enhance exon recognition by alleviating constraints generated by gene nucleotide composition bias

GEO Series GSE138397. Homo sapiens. 12 samples. Type: Other.

openGEO-OpenOct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record