Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
102
datasets available to search
ShareScore release 0.7.1
Dataset results
102 results for “outlier”
Data for "PTP Over Wide Area Networks With Offset Measurement Outlier Filtering"
<p>Dataset used in the manuscript "PTP Over Wide Area Networks With Offset Measurement Outlier Filtering". This dataset contains synchronization accuracy measurements over long distance links using both NTP and PTP, as well as synthetically generated PTP replays used for offline testing.</p> <p>A detailed description of the contents is found in the <code>README.md</code> file at the root of the dataset.</p>
DWCox: A Density-Weighted Cox Model for Outlier-Robust Prediction of Prostate Cancer Survival
<p>This package, <strong>DWCox</strong>, implements a <strong>d</strong>ensity-<strong>w</strong>eighted <strong>Cox</strong> regression model that is more robust against outliers in the training data. DWCox gives more accurate predictions than the standard Cox regression on prostate cancer survival, especially in cases where the training data are expected to contain a lot of outliers. More details can be found in our paper (coming soon) and the README file inside this package.</p>
Lipidomics LC-MS analysis support tools for outlier detection
<p>Identification of features with high levels of confidence in liquid chromatography-mass spectrometry (LC MS) lipidomics research is an essential part of biomarker discovery, but existing software platforms can give inconsistent results, even from identical spectral data. This poses a clear challenge for reproducibility in bioinformatics work, and highlights the importance of data-driven outlier detection in assessing spectral outputs – here demonstrated using a machine learning approach based on support vector machine regression combined with leave-one-out cross validation – as well as manual curation, in order to identify software-driven errors driven by closely related lipids and by co-elution issues.</p> <p>The lipidomics case study dataset used in this work analysed a lipid extraction of a human pancreatic adenocarcinoma cell line (PANC-1, Merck, UK, cat no. 87092802) analysed using an Acquity M-Class UPLC system (Waters, UK) coupled to a ZenoToF 7600 mass spectrometer (Sciex, UK). Raw output files are included alongside processed data using MS DIAL (v4.9.221218) and Lipostar (v2.1.4) and a Jupyter notebook with Python code to analyse the outputs for outlier detection.</p>
Semantic-Discrepant Outliers on CIFAR-10 Dataset
<p>We provide synthetic Out-of-distibution (OOD) dataset, which is called Semantic-Discrepant (SD) outliers, on CIFAR-10 dataset. SD outliers can be utilized for boosting OOD detection model performance. For the details, SD outliers are realistic OOD samples that contains incoherent semantic shift while preserving nuisances with in-distribution (ID). SD-outliers are generated from ID training samples using semantic-discrepant sampling in the diffusion model. so SD-outliers on CIFAR-10 contains 50000 32X32 images which is same as CIFAR-10 training dataset size. The dataset has a capacity of 768MB.</p>
Supplementary Data to *Robust adaptive distance functions for approximate Bayesian inference on outlier-corrupted data*
<p>Supplementary code and data to <strong>Robust adaptive distance functions for approximate Bayesian inference on outlier-corrupted data</strong> by <strong>Y. Schaelte et al., 2021</strong>.</p> <p>The archive contains a <strong>README.rst </strong>for information on what is where and how to execute the study and generate the figures. The underlying code without the data can be found at the repository https://github.com/yannikschaelte/study_abc_rad, of which this archive is a snapshot.</p> <p> </p>
Multi-Domain Outlier Detection Dataset
<p>The Multi-Domain Outlier Detection Dataset contains datasets for conducting outlier detection experiments for four different application domains:</p> <ol> <li>Astrophysics - detecting anomalous observations in the Dark Energy Survey (DES) catalog (data type: feature vectors)</li> <li>Planetary science - selecting novel geologic targets for follow-up observation onboard the Mars Science Laboratory (MSL) rover (data type: grayscale images)</li> <li>Earth science: detecting anomalous samples in satellite time series corresponding to ground-truth observations of maize crops (data type: time series/feature vectors)</li> <li>Fashion-MNIST/MNIST: benchmark task to detect anomalous MNIST images among Fashion-MNIST images (data type: grayscale images)</li> </ol> <p>Each dataset contains a "fit" dataset (used for fitting or training outlier detection models), a "score" dataset (used for scoring samples used to evaluate model performance, analogous to test set), and a label dataset (indicates whether samples in the score dataset are considered outliers or not in the domain of each dataset). </p> <p>To read more about the datasets and how they are used for outlier detection, or to cite this dataset in your own work, please see the following citation:</p> <p>Kerner, H. R., Rebbapragada, U., Wagstaff, K. L., Lu, S., Dubayah, B., Huff, E., Lee, J., Raman, V., and Kulshrestha, S. (2022). Domain-agnostic Outlier Ranking Algorithms (DORA)-A Configurable Pipeline for Facilitating Outlier Detection in Scientific Datasets. Under review for <em>Frontiers in Astronomy and Space Sciences</em>. </p>
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Abstract Phylogenetic reconstruction using concatenated loci ("phylogenomics" or "supermatrix phylogeny") is a powerful tool for solving evolutionary splits that are poorly resolved in single gene/protein trees (SGTs). However, recent phylogenomic attempts to resolve the eukaryote root have yielded conflicting results, along with claims of various artefacts hidden in the data. We have investigated these conflicts using two new methods for assessing phylogenetic conflict. ConJak uses whole marker (gene or protein) jackknifing to assess deviation from a central mean for each individual sequence, while ConWin uses a sliding window to screen for incongruent protein fragments (mosaics). Both methods allow selective masking of individual sequences or sequence fragments in order to minimize missing data, an important consideration for resolving deep splits with limited data. Analyses focused on a set of 76 eukaryotic proteins of bacterial-ancestry previously used in various combinations to assess the branching order among the three major divisions of eukaryotes: Amorphea (mainly animals, fungi and Amoebozoa), Diaphoretickes (most other well-known eukaryotes and nearly all algae) and Excavata, represented here by Discoba (Jakobida, Heterolobosea, and Euglenozoa). ConJak analyses found strong outliers to be concentrated in under-sampled lineages, while ConWin analyses of Discoba, the most under-sampled of the major lineages, detected potentially incongruent fragments scattered throughout. Phylogenetic analyses of the full data using an LG-gamma model support a Discoba sister scenario (neozoan-excavate root), which rises to 99-100% bootstrap support with data masked according to either protocol. However, analyses with two site-specific (CAT) mixture models yielded widely inconsistent results and a striking sensitivity to missing data. The neozoan-excavate root places Amorphea and Diaphoretickes as more closely related to each other than either is to Discoba, a fundamental relationship that should remain unaffected by additional taxa.
Data to accompany the outlier-waveform-detection Github repository (internal globus pallidus, GPi)
<p>This repository contains data based on neuronal recordings from two monkeys (G and I, in the pre- and post-MPTP states) that serve as input to the code provided at <a href="https://github.com/turner-lab-pitt/outlier-waveform-detection">https://github.com/turner-lab-pitt/outlier-waveform-detection</a>. Text files located within that Github repository provide detailed instructions on how these data may be used with that code. As described in those text files, extra data are provided for Monkey G, in the pre-MPTP state.</p> <p>The data-description.txt file provides detailed information regarding the contents of each zipped tar archive. Briefly, the most important components of the files are the "snips" (individual spike waveforms) from the two monkeys and MPTP states, as extracted for each of a series of single sorted units from the internal globus pallidus (GPi). The additional G-Pre data provides examples of the high-pass filtered voltage signals from which these snips were extracted. All data are stored in the Matlab .mat format.</p> <p>All zipped files can be decompressed with 7-zip: <a href="https://www.7-zip.org/" target="_blank" rel="noopener">https://www.7-zip.org/</a></p> <p>These data and the associated Github code were used for analyses reported in an in-preparation manuscript (Kase et al., "Movement-related activity in the internal globus pallidus of the parkinsonian macaque"), and also with a preprint that is currently under review:</p> <div> <div>Detecting rhythmic spiking through the power spectra of point process model residuals</div> </div> <div>Karin M. Cox, Daisuke Kase, Taieb Znati, Robert S. Turner</div> <div>bioRxiv 2023.09.08.556120; doi: <a href="https://doi.org/10.1101/2023.09.08.556120" target="_blank" rel="noopener">https://doi.org/10.1101/2023.09.08.556120</a></div> <div> </div> <p>This research was funded in part by Aligning Science Across Parkinson's [ASAP-020519] through the Michael J. Fox Foundation for Parkinson's Research (MJFF). For the purpose of open access, the authors have applied a Creative Commons Attribution 4.0 International (CC BY) public copyright license to this dataset. </p>
Synthetic Dataset for Outlier Detection
<p>This synthetically generated dataset can be used to evaluate outlier detection algorithms. It has 10 attributes and 1000 observations, of which 100 are labeled as outliers. Two-dimensional combinations of attributes form differently shaped clusters.</p> <ul> <li>Attribute 0 & Attribute 1: Two circular clusters</li> <li>Attribute 2 & Attribute 3: Two banana shaped clusters</li> <li>Attribute 4 & Attribute 5: Three point clouds</li> <li>Attribute 6 & Attribute 7: Two point clouds with variances</li> <li>Attribute 8 & Attribute 9: Three anisotropic shaped clusters. </li> </ul> <p>The "outlier" column states whether an observation is an outlier or not. Additionally, the .zip file contains 10 stratified randomized train test splits (70% train, 30% test).</p>
Wiki-based Communities of Interest: Demographics and Outliers
<p>These datasets contains statements about demographics and outliers of Wiki-based Communities of Interest. </p> <p><strong>Group-centric dataset (sample):</strong></p> <pre><code class="language-json">{ "title": "winners of Priestley Medal", "recorded_members": 83, "topics": ["STEM.Chemistry"], "demographics": [ "occupation-chemist", "gender-male", "citizen-U.S." ], "outliers": [ { "reason": "NOT(chemist) unlike 82 recorded members", "members": [ "Francis Garvan (lawyer, art collector)" ] }, { "reason": "NOT(male) unlike 80 recorded members", "members": [ "Mary L. Good (female)", "Darleane Hoffman (female)", "Jacqueline Barton (female)" ] } ] }</code></pre> <p><strong>Subject-centric dataset (sample):</strong></p> <pre><code class="language-json">{ "subject": "Serena Williams", "statements": [ { "statement": "NOT(sport-basketball) but (tennis) unlike 4 recorded winners of Best Female Athlete ESPY Award.", "score": 0.36 }, { "statement": "NOT(occupation-politician) but (tennis player, businessperson, autobiographer) unlike 20 recorded winners of Michigan Women's Hall of Fame.", "score": 0.17 } ] }</code></pre> <p><strong>This data can be also browsed at: <a href="https://wikiknowledge.onrender.com/demographics/">https://wikiknowledge.onrender.com/demographics/</a></strong></p>
Key triggers of adaptive genetic variability of sessile oak [Q. petraea (Matt.) Liebl.] from the Balkan refugia: outlier detection and association of SNP loci from ddRAD-seq data
<p>Knowledge on the genetic composition of <em>Quercus petraea</em> in south-eastern Europe is limited despite the species' significant role in the re-colonisation of Europe during the Holocene, and the diverse climate and physical geography of the region. Therefore, it is imperative to conduct research on adaptation in sessile oak to better understand its ecological significance in the region. While large sets of SNPs have been developed for the species, there is a continued need for smaller sets of SNPs that are highly informative about the possible adaptation to this varied landscape. By using double digest restriction site associated DNA sequencing data from our previous study, we mapped RAD-tag sequences to the <em>Quercus robur</em> reference genome and identified a set of SNPs putatively related to drought stress-response. A total of 179 individuals from eighteen natural populations at sites covering heterogeneous climatic conditions in the southeastern natural distribution range of <em>Q. petraea</em> were genotyped. The detected highly polymorphic variant sites revealed three genetic clusters with a generally low level of genetic differentiation and balanced diversity among them but showed a north–southeast gradient. Selection tests showed nine outlier SNPs positioned in different functional regions. Genotype-environment association analysis of these markers yielded a total of 53 significant associations, explaining 2.4–16.6% of the total genetic variation. Our work exemplifies that adaptation to drought may be under natural selection in the examined <em>Q. petraea</em> populations.</p>
Рис. 8–13. ΔанΑшафты Южного УраΛа (8–11) и Русской равнины (12–13). 8 – разнотравная степь у поΑножия горы ВербΛюжка, местообитание Cionus rossicus; 9 – ксерофитные Λуга в пойме реки УраΛ вбΛизи горы ВербΛюжка, местообитание Cionus rossicus; 10 – южные степи в районе КзыΛаΑырского карстового поΛя, местообитание Cionus gebleri; 11 – степи низкогорий Южного УраΛа бΛиз с. КиΑрясово, местообитание Smicronyx albopictus; 12 – КаменноброΑские меΛовые горы на юго-запаΑе ПривоΛжской возвышенности, местообитание Mecinus janthiniformis, Smicronyx robustus и S. albopictus; 13 – меΛовой останец КобыΛья ГоΛова в прироΑном парке «Àонской», местообитание Mecinus janthiniformis. Figs 8–13. Landscapes of the Southern Urals (8–11) and the Russian Plain (12–13). 8 – forb steppe at the down of Verblyuzhka Mt., habitat of Cionus rossicus; 9 – xerophytic meadows in the floodplain of the Ural River near Verblyuzhka Mt., habitat of Cionus rossicus; 10 – southern steppes in the Kzyladyr karst area, habitat of Cionus gebleri; 11 – steppes of the low mountains of the Southern Urals near Kidryasovo village, habitat of Smicronyx albopictus; 12 – Kamennobrodsky chalk mountains in the southwest of the Volga Upland, habitat of Mecinus janthiniformis, Smicronyx robustus, and S. albopictus; 13 – Cretaceous outlier Kobyl'ya Golova in the Donskoy Nature Park, habitat of Mecinus janthiniformis. in Interesting records of weevils (Coleoptera: Curculionidae: Curculioninae) in the steppe zone of the European part of Russia and the Urals
Рис. 8–13. ΔанΑшафты Южного УраΛа (8–11) и Русской равнины (12–13). 8 – разнотравная степь у поΑножия горы ВербΛюжка, местообитание Cionus rossicus; 9 – ксерофитные Λуга в пойме реки УраΛ вбΛизи горы ВербΛюжка, местообитание Cionus rossicus; 10 – южные степи в районе КзыΛаΑырского карстового поΛя, местообитание Cionus gebleri; 11 – степи низкогорий Южного УраΛа бΛиз с. КиΑрясово, местообитание Smicronyx albopictus; 12 – КаменноброΑские меΛовые горы на юго-запаΑе ПривоΛжской возвышенности, местообитание Mecinus janthiniformis, Smicronyx robustus и S. albopictus; 13 – меΛовой останец КобыΛья ГоΛова в прироΑном парке «Àонской», местообитание Mecinus janthiniformis. Figs 8–13. Landscapes of the Southern Urals (8–11) and the Russian Plain (12–13). 8 – forb steppe at the down of Verblyuzhka Mt., habitat of Cionus rossicus; 9 – xerophytic meadows in the floodplain of the Ural River near Verblyuzhka Mt., habitat of Cionus rossicus; 10 – southern steppes in the Kzyladyr karst area, habitat of Cionus gebleri; 11 – steppes of the low mountains of the Southern Urals near Kidryasovo village, habitat of Smicronyx albopictus; 12 – Kamennobrodsky chalk mountains in the southwest of the Volga Upland, habitat of Mecinus janthiniformis, Smicronyx robustus, and S. albopictus; 13 – Cretaceous outlier Kobyl'ya Golova in the Donskoy Nature Park, habitat of Mecinus janthiniformis.
specleanr: An R package for automated flagging of environmental outliers in ecological data for modeling workflows
Open the record for dataset details and reuse information.
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Open the record for dataset details and reuse information.
Outliers and Missing Gait_Human, IDS, JavaScript vulnerability Datasets
<p>Human Gait Dataset (CASIA-A) [1] is available at: http://www.cbsr.ia.ac.cn/english/Gait\%20Databases.asp, JavaScript vulnerability<br>dataset is publicly available at [2] and KDD CUP 99 dataset is publicly available at [3].</p> <p> </p> <p> </p> <p>[1] Wang L, Tan T, Ning H, Hu W. Silhouette analysis-based gait recognition for<br>human identification. IEEE transactions on pattern analysis and machine<br>intelligence. 2003;25(12):1505–1518</p> <p>[2] Ferenc R, Heged ̋us P, Gyimesi P, Antal G, B ́an D, Gyim ́othy T. Challenging<br>machine learning algorithms in predicting vulnerable javascript functions. In:<br>2019 IEEE/ACM 7th International Workshop on Realizing Artificial Intelligence<br>Synergies in Software Engineering (RAISE). IEEE; 2019. p. 8–14.</p> <p>[3] Tavallaee M, Bagheri E, Lu W, Ghorbani AA. A detailed analysis of the KDD<br>CUP 99 data set. In: 2009 IEEE symposium on computational intelligence for<br>security and defense applications. Ieee; 2009. p. 1–6.</p> <p> </p>
TreeShrink: fast and accurate detection of outlier long branches in collections of phylogenetic trees
Open the record for dataset details and reuse information.
ClinePlotR: Visualizing genomic clines and detecting outliers in R
<p class="Normal1">Patterns of multi-locus differentiation (i.e., genomic clines) often extend broadly across hybrid zones and their quantification can help diagnose how species boundaries are shaped by adaptive processes, both intrinsic and extrinsic. In this sense, the transitioning of loci across admixed individuals can be contrasted as a function of the genome-wide trend, in turn allowing an expansion of clinal theory across a much wider array of biodiversity. However, computational tools that serve to interpret and consequently visualize 'genomic clines' are limited.</p> <p>Here, we introduce the <span class="MsoSubtleReference">ClinePlotR R</span>-package for visualizing genomic clines and detecting outlier loci using output generated by two popular software packages, <span class="MsoSubtleReference">bgc </span>and <span class="MsoSubtleReference">Introgress.</span></p> <p><span class="MsoSubtleReference">ClinePlotR </span>bundles both input generation (i.e, filtering datasets and creating specialized file formats) and output processing (e.g., MCMC thinning and burn-in) with functions that directly facilitate interpretation and hypothesis testing. Tools are also provided for post-hoc analyses that interface with external packages such as <span class="MsoSubtleReference">ENMeval </span>and <span class="MsoSubtleReference">RIdeogram</span></p> <p>Our package increases the reproducibility and accessibility of genomic cline methods, thus allowing an expanded user base and promoting these methods as mechanisms to address diverse evolutionary questions in both model and non-model organisms.</p>
Data from: Pacman profiling: a simple procedure to identify stratigraphic outliers in high-density deep-sea microfossil data
The deep-sea microfossil record is characterized by an extraordinarily high density and abundance of fossil specimens, and by a very high degree of spatial and temporal continuity of sedimentation. This record provides a unique opportunity to study evolution at the species level for entire clades of organisms. Compilations of deep-sea microfossil species occurrences are, however, affected by reworking of material, age model errors, and taxonomic uncertainties, all of which combine to displace a small fraction of the recorded occurrence data both forward and backwards in time, extending total stratigraphic ranges for taxa. These data outliers introduce substantial errors into both biostratigraphic and evolutionary analyses of species occurrences over time. We propose a simple method—Pacman—to identify and remove outliers from such data, and to identify problematic samples or sections from which the outlier data have derived. The method consists of, for a large group of species, compiling species occurrences by time and marking as outliers calibrated fractions of the youngest and oldest occurrence data for each species. A subset of biostratigraphic marker species whose ranges have been previously documented is used to calibrate the fraction of occurrences to mark as outliers. These outlier occurrences are compiled for samples, and profiles of outlier frequency are made from the sections used to compile the data; the profiles can then identify samples and sections with problematic data caused, for example, by taxonomic errors, incorrect age models, or reworking of sediment. These samples/sections can then be targeted for re-study.
Data from: Integration of Random Forest with population-based outlier analyses provides insight on the genomic basis and evolution of run timing in Chinook salmon (Oncorhynchus tshawytscha)
Anadromous Chinook salmon populations vary in the period of river entry at the initiation of adult freshwater migration, facilitating optimal arrival at natal spawning. Run timing is a polygenic trait that shows evidence of rapid parallel evolution in some lineages, signifying a key role for this phenotype in the ecological divergence between populations. Studying the genetic basis of local adaptation in quantitative traits is often impractical in wild populations. Therefore, we used a novel approach, Random Forest, to detect markers linked to run timing across 14 populations from contrasting environments in the Columbia River and Puget Sound, USA. The approach permits detection of loci of small effect on the phenotype. Divergence between populations at these loci was then examined using both principle component analysis and FST outlier analyses, to determine whether shared genetic changes resulted in similar phenotypes across different lineages. Sequencing of 9107 RAD markers in 414 individuals identified 33 predictor loci explaining 79.2% of trait variance. Discriminant analysis of principal components of the predictors revealed both shared and unique evolutionary pathways in the trait across different lineages, characterized by minor allele frequency changes. However, genome mapping of predictor loci also identified positional overlap with two genomic outlier regions, consistent with selection on loci of large effect. Therefore, the results suggest selective sweeps on few loci and minor changes in loci that were detected by this study. Use of a polygenic framework has provided initial insight into how divergence in a trait has occurred in the wild.
Data from: Outlier SNPs detect weak regional structure against a background of genetic homogeneity in the Eastern Rock Lobster, Sagmariasus verreauxi
Genetic differentiation is characteristically weak in marine species making assessments of population connectivity and structure difficult. However the advent of genomic methods have increased genetic resolution, enabling studies to detect weak, but significant population differentiation within marine species. With an increasing number of studies employing high resolution genome-wide techniques, we are realising the connectivity of marine populations is often complex and quantifying this complexity can provide an understanding of the processes shaping marine species genetic structure and to inform long-term, sustainable management strategies. This study aims to assess the genetic structure, connectivity and local adaptation of the Eastern Rock Lobster (Sagmariasus verreauxi), which has a maximum pelagic larval duration of 12 months and inhabits both subtropical and temperate environments. We used 645 neutral and 15 outlier SNPs to genotype lobsters collected from the only two known breeding populations and a third episodic population — encompassing S. verreauxi's known range. Through examination of the neutral SNP panel, we detected genetic homogeneity across the three regions, which extended across the Tasman Sea encompassing both Australian and New Zealand populations. We discuss differences in neutral genetic signature of S. verreauxi and a closely-related, co-distributed rock lobster, Jasus edwardsii, determining a regional pattern of genetic disparity between the species, which have largely similar life histories. Examination of the outlier SNP panel detected weak genetic differentiation between the three regions. Outlier SNPs showed promise in assigning individuals to their sampling origin and may prove useful as a management tool for species exhibiting genetic homogeneity.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.