Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
53
datasets available to search
ShareScore release 0.9.0
Dataset results
53 results for “principal components analysis”
Data from: Principal component analysis as an alternative treatment for morphometric characters: phylogeny of caseids as a case study
In a recent study, the phylogeny of Caseidae (a herbivorous family of Palaeozoic synapsids belonging to the paraphyletic grade known as pelycosaurs) was analysed with a dataset employing more than three hundred continuous morphological characters in an effort to follow the principles of total evidence. Continuous characters are a source of great debate, with disagreements surrounding their suitability for and treatment in phylogenetic analysis. A number of shortcomings were identified in the handling of continuous characters in this study of caseids, including the use of gap weighting to discretize the characters and potential issues with redundancy and character non-independence. Therefore, an alternative treatment for these characters is suggested here. First, rather than using gap weighting, the continuous characters were analysed in the program TNT, in which the raw values can be treated as continuous rather than discrete. Second, prior to the phylogenetic analysis, the continuous characters were subjected to a log-ratio principal component analysis, and then the principal components were included in the character matrix rather than the raw ratios. Analysing the original data in TNT produced little difference in the results, but using the principal components as continuous characters resulted in alternative positions for Caseopsis agilis, Ennatosaurus tecton and Caseoides sanangeloensis. The differences are judged to be due to the reduced redundancy of the characters, the smaller number of principal components not overwhelming the discrete characters and the use of a scaling method which allows principal components with a higher variance to have a greater influence on the analysis. The positions of highly fragmentary fossils depended heavily on the method used to treat the missing characters in the principal component analysis, and so the method proposed here is not recommended for analysing very incomplete taxa.
Fig. 2 in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis
Fig. 2. Discriminant function graph of seven hornbill species based on the enter independents together model: A — genus Rhyticeros; B — genus Buceros; C — genus Anthracoceros.
A Potential Method for Identifying Milk Adulteration and Pb(II) Contamination Scenarios Using Principal Component Analysis from Smartphone Photographs
<p>Early Research Data</p>
Data from: Comparative analysis of principal components can be misleading
Most existing methods for modeling trait evolution are univariate, although researchers are often interested in investigating evolutionary patterns and processes across multiple traits. Principal components analysis (PCA) is commonly used to reduce the dimensionality of multivariate data so that univariate trait models can be fit to individual principal components. The problem with using standard PCA on phylogenetically structured data has been previously pointed out yet it continues to be widely used in the literature. Here we demonstrate precisely how using standard PCA can mislead inferences: The first few principal components of traits evolved under constant-rate multivariate Brownian motion will appear to have evolved via an "early burst" process. A phylogenetic PCA (pPCA) has been proprosed to alleviate these issues. However, when the true model of trait evolution deviates from the model assumed in the calculation of the pPCA axes, we find that the use of pPCA suffers from similar artifacts as standard PCA. We show that data sets with high effective dimensionality are particularly likely to lead to erroneous inferences. Ultimately, all of the problems we report stem from the same underlying issue—by considering only the first few principal components as univariate traits, we are effectively examining a biased sample of a multivariate pattern. These results highlight the need for truly multivariate phylogenetic comparative methods. As these methods are still being developed, we discuss potential alternative strategies for using and interpreting models fit to univariate axes of multivariate data.
Data from: Comparative analysis of principal components can be misleading
Open the record for dataset details and reuse information.
Data from: pcadapt: an R package to perform genome scans for selection based on principal component analysis
Open the record for dataset details and reuse information.
PERMANOVA results from Principal component analysis of avian hind limb and foot morphometrics and the relationship between ecology and phylogeny
Open the record for dataset details and reuse information.
Data from: Principal component analysis as an alternative treatment for morphometric characters: phylogeny of caseids as a case study
Open the record for dataset details and reuse information.
Risk prediction models for dementia constructed by supervised principal component analysis using miRNA expression data
GEO Series GSE120584. Homo sapiens. 1601 samples. Type: Non-coding RNA profiling by array.
Figure 1 in Interpopulation differences in shell forms of the pearl oyster, Pinctada imbricata radiata (Bivalvia: Pterioida), in the northern Persian Gulf inferred from principal component analysis and elliptic Fourier analysis
Figure 1. Map of the study area showing the fishing grounds.
JPSS-1 CrIS Level 1B Principal Component Analysis / Rapid Event Detection V3.0 (SNDRJ1CrISL1BPCARED) at GES DISC
This sample data collection contains L1B radiance values that are compressed and denoised via Principal Component Analysis (PCA). Additionally it contains a new Rapid Event Detection (RED) product, which for multiple spectral regions identify where rare signals are observed, providing an efficient and useful way to locate and study interesting phenomena. Some of the potential uses for these RED components are the detection of atmospheric gases where fires and volcanoes occur. PCA/RED might be a desirable alternative to the existing L1B product because (1) it is many times smaller in data volume, (2) about 80% of the random noise is removed, and (3) it comes with the new RED component. The Cross-track Infrared Sounder (CrIS) Level 1B Full Spectral Resolution (FSR) data files contain radiance measurements along with ancillary spacecraft, instrument, and geolocation data of the CrIS instrument on the Joint Polar Satellite System-1 (JPSS-1) platform. This platform is also know as NOAA-20 (National Oceanic and Atmospheric Administration).
pcaGoPromoter - An R package for functional interpretation of principal component analysis of genome-wide gene expression data
GEO Series GSE27071. Homo sapiens. 13 samples. Type: Expression profiling by array.
PRINCIPAL COMPONENT ANALYSIS FOR EXPERIMENTS
GEO Series GSE31375. Mus musculus; Rattus norvegicus. 0 samples. Type: Third-party reanalysis; Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.