Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.9.0
Dataset results
1,782 results for “Algorithm”
Comparison of RNA-seq and Microarray Platforms for Splice Event Detection using a Cross-Platform Algorithm
GEO Series GSE104974. Homo sapiens. 60 samples. Type: Expression profiling by array; Expression profiling by high throughput sequencing.
A comprehensive assessment of array-based platforms and calling algorithms for detection of copy number variants (Mapping250K_Nsp)
GEO Series GSE28105. Homo sapiens. 14 samples. Type: Genome variation profiling by SNP array.
CoGAPS matrix factorization algorithm identifies AP-2alpha as a feedback mechanism from therapeutic inhibition of the EGFR network
GEO Series GSE80667. Homo sapiens. 96 samples. Type: Expression profiling by array.
A COMPARATIVE STUDY OF ALGORITHMS FOR LAND COVER CHANGE
A COMPARATIVE STUDY OF ALGORITHMS FOR LAND COVER CHANGE SHYAM BORIAH*, VARUN MITHAL*, ASHISH GARG*, VIPIN KUMAR*, MICHAEL STEINBACH*, CHRIS POTTER**, AND STEVE KLOOSTER*** Abstract. Ecosystem-related observations from remote sensors on satellites offer huge potential for understanding the location and extent of global land cover change. This paper presents a comparative study of three time series based algorithms for detecting changes in land cover. The techniques are evaluated quantitatively using forest fire ground truth from the state of California for 2000–2009. On relatively high quality data sets, all three schemes perform reasonably well, but their ability to handle noise and natural variability in the vegetation data differs dramatically. In particular, one of the algorithms significantly outperforms the other two since it accounts for variability in the time series.
Empirical Evaluation of Diagnostic Algorithm Performance Using a Generic Framework
A variety of rule-based, model-based and datadriven techniques have been proposed for detection and isolation of faults in physical systems. However, there have been few efforts to comparatively analyze the performance of these approaches on the same system under identical conditions. One reason for this was the lack of a standard framework to perform this comparison. In this paper we introduce a framework, called DXF, that provides a common language to represent the system description, sensor data and the fault diagnosis results; a run-time architecture to execute the diagnosis algorithms under identical conditions and collect the diagnosis results; and an evaluation component that can compute performance metrics from the diagnosis results to compare the algorithms. We have used DXF to perform an empirical evaluation of 13 diagnostic algorithms on a hardware testbed (ADAPT) at NASA Ames Research Center and on a set of synthetic circuits typically used as benchmarks in the model-based diagnosis community. Based on these empirical data we analyze the performance of each algorithm and suggest directions for future development.
Anomaly Detection and Diagnosis Algorithms for Discrete Symbols
We present a set of novel algorithms which we call sequenceMiner that detect and characterize anomalies in large sets of high-dimensional symbol sequences that arise from recordings of switch sensors in the cockpits of commercial airliners. While the algorithms we present are general and domain-independent, we focus on a specific problem that is critical to determining the system-wide health of a fleet of aircraft. The approach taken uses unsupervised clustering of sequences using the normalized length of the longest common subsequence (nLCS) as a similarity measure, followed by detailed outlier analysis to detect anomalies. In this method, an outlier sequence is defined as a sequence that is far away from the cluster centre. We present new algorithms for outlier analysis that provide comprehensible indicators as to why a particular sequence is deemed to be an outlier. The algorithms provide a coherent description to an analyst of the anomalies in the sequence when compared to more normal sequences. In the final section of the paper we demonstrate the effectiveness of sequenceMiner for anomaly detection on a real set of discrete sequence data from a fleet of commercial airliners. We show that sequenceMiner discovers actionable and operationally significant safety events. We also compare our innovations with standard HiddenMarkov Models, and show that our methods are superior.
Towards a Framework for Evaluating and Comparing Diagnosis Algorithms
Diagnostic inference involves the detection of anomalous system behavior and the identification of its cause, possibly down to a failed unit or to a parameter of a failed unit. Traditional approaches to solving this problem include expert/rule-based, model-based, and data-driven methods. Each approach (and various techniques within each approach) use different representations of the knowledge required to perform the diagnosis. The sensor data is expected to be combined with these internal representations to produce the diagnosis result. In spite of the availability of various diagnosis technologies, there have been only minimal efforts to develop a standardized software framework to run, evaluate, and compare different diagnosis technologies on the same system. This paper presents a framework that defines a standardized representation of the system knowledge, the sensor data, and the form of the diagnosis results – and provides a run-time architecture that can execute diagnosis algorithms, send sensor data to the algorithms at appropriate time steps from a variety of sources (including the actual physical system), and collect resulting diagnoses. We also define a set of metrics that can be used to evaluate and compare the performance of the algorithms, and provide software to calculate the metrics.
Comparison of Algorithms for Anomaly Detection in Flight Recorder Data of Airline Operations
Published at 12th AIAA Aviation Technology, Integration, and Operations (ATIO) Conference and 14th AIAA/ISSM 17 - 19 September 2012, Indianapolis, Indiana
ICESat-2 Derived Sea Ice Melt Pond Characteristics from the Density-Dimension Algorithm V003
This data set provides locations and depths of melt ponds in the Multi-Year Arctic Sea Ice Region, calculated from ATLAS/ICESat-2 L2A Global Geolocated Photon Data, Version 5 (ATL03) using an autoadaptive algorithm.
Simulating Degradation Data for Prognostic Algorithm Development
**PHM08 Challenge Dataset is now publicly available at the NASA Prognostics Respository + [Download](http://ti.arc.nasa.gov/tech/dash/pcoe/prognostic-data-repository/)** **INTRODUCTION - WHY SIMULATE DEGRADATION DATA?** Of various challenges encountered in prognostics algorithm development, the non-availability of suitable validation data is most often the bottleneck in the technology certification process. Prognostics imposes several requirements on the training data in addition to what is commonly available from various applications. It not only requires data containing fault signatures but also that contains fault evolution trends with corresponding time indexes (in number of hours or number of operational cycles). In general there are three sources from which data is usually available, namely: Fielded applications, experimental test-beds, and computer simulations (see [Figure 1](https://c3.ndc.nasa.gov/dashlink/static/media/other/3_modes.bmp)). From prognostics point of view, data collection paradoxically suffers from the situation that the systems that do run to failure often did not have warning instrumentation installed, hence no or little record of what went wrong. In the other situation, those that are continuously monitored are prevented from running to failure or are subject to maintenance that eliminates the signatures of fault evolution. Conducting experiments that replicate real world situations is extremely expensive in terms of time required for a healthy system to run to failure and is often dangerous. Accelerated ageing may be useful to some extent but may not emulate normal wear patterns. Furthermore, to manage uncertainty multiple datasets must be collected to quantify variations resulting from multiple sources, which makes it all the way more unattainable. Simulations can be fast, inexpensive, and provide a number of options to design experiments, but their usefulness is contingent on the availability of high fidelity models that represent the real systems fairly well. However, once such a model is available, simulations offer the flexibility to rerun various experiments with added knowledge from the system as it becomes available. Where, availability of real fault evolution data from the fielded systems would be more desirable, generating data using a high fidelity model and integrating it with the knowledge gathered from the partial data obtained from the real systems is by far the most practical approach for prognostics algorithm development, validation, and verification. In this presentation we discuss some key elements that must be kept in mind while generating datasets suitable for prognostics. Furthermore, with the help of an example it has been shown how a dynamical system model can be supported with suitable degradation models available from respective domain knowledge to create suitable data. The example is discussed next. **APPLICATION DOMAIN** Tracking and Predicting the progressionof damage in aircraft turbo machinery has been an active area of study within the Condition Based Maintenance (CBM) community. A general approach has been to correlate flow and effciency losses to degradation signtures in various components of the engine. Once such mapping is available, the next task is to estimate this loss of flow and eficiency inferring information from measurable sensor outputs, which ultimtely is used to assess the level of degradation in the system. **SYSTEM MODEL: C-MAPSS** The C-MAPSS (Commercial Modular Aero Propulsion System Simulation) is a tool, recently released, for simulating a realistic large commercial turbofan engine. C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) that simulates a realistic large (~90,000lb) commercial turbofan engine. It allows the user to choose and design operational profiles, controllers, environmental conditions, thrust levels, etc. to simualte a scenario of interest. An extensive list of output va
Entropy-based probabilistic fatigue damage prognosis and algorithmic performance comparison
In this paper, a maximum entropy-based general framework for probabilistic fatigue damage prognosis is investigated. The proposed methodology is based on an underlying physics-based crack growth model. V arious uncertainties from measurements, modeling, and parameter estimations are considered to describe the stochastic process of fatigue damage accumulation. A probabilistic prognosis updating procedure based on the maximum relative entropy concept is proposed to incorporate measurement data. Markov Chain Monte Carlo (MCMC) technique is used to provide the posterior samples for model updating in the maximum entropy approach. Experimental data are used to demonstrate the operation of the proposed probabilistic prognosis methodology. A set of prognostics-based metrics are employed to quantitatively evaluate the prognosis performance and compare the proposed method with the classical Bayesian updating algorithm. In particular, model accuracy, precision and convergence are rigorously evaluated in* addition to the qualitative visual comparison.
Uncertainty Representation and Interpretation in Model-based Prognostics Algorithms based on Kalman Filter Estimation
This article discusses several aspects of uncertainty represen- tation and management for model-based prognostics method- ologies based on our experience with Kalman Filters when applied to prognostics for electronics components. In par- ticular, it explores the implications of modeling remaining useful life prediction as a stochastic process and how it re- lates to uncertainty representation, management, and the role of prognostics in decision-making. A distinction between the interpretations of estimated remaining useful life probability density function and the true remaining useful life probabil- ity density function is explained and a cautionary argument is provided against mixing interpretations for the two while considering prognostics in making critical decisions.
Algorithms for Spectral Decomposition with Applications
The analysis of spectral signals for features that represent physical phenomenon is ubiquitous in the science and engineering communities. There are two main approaches that can be taken to extract relevant features from these high-dimensional data streams. The first set of approaches relies on extracting features using a physics-based paradigm where the underlying physical mechanism that generates the spectra is used to infer the most important features in the data stream. We focus on a complementary methodology that uses a data-driven technique that is informed by the underlying physics but also has the ability to adapt to unmodeled system attributes and dynamics. We discuss the following four algorithms: Spectral Decomposition Algorithm (SDA), Non-Negative Matrix Factorization (NMF), Independent Component Analysis (ICA) and Principal Components Analysis (PCA) and compare their performance on a spectral emulator which we use to generate artificial data with known statistical properties. This spectral emulator mimics the real-world phenomena arising from the plume of the space shuttle main engine and can be used to validate the results that arise from various spectral decomposition algorithms and is very useful for situations where real-world systems have very low probabilities of fault or failure. Our results indicate that methods like SDA and NMF provide a straightforward way of incorporating prior physical knowledge while NMF with a tuning mechanism can give superior performance on some tests. We demonstrate these algorithms to detect potential system-health issues on data from a spectral emulator with tunable health parameters.
A Local Scalable Distributed Expectation Maximization Algorithm for Large Peer-to-Peer Networks
This paper describes a local and distributed expectation maximization algorithm for learning parameters of Gaussian mixture models (GMM) in large peer-to-peer (P2P) environments. The algorithm can be used for a variety of well-known data mining tasks in distributed environments such as clustering, anomaly detection, target tracking, and density estimation to name a few, necessary for many emerging P2P applications in bioinformatics, webmining and sensor networks. Centralizing all or some of the data to build global models is impractical in such P2P environments because of the large number of data sources, the asynchronous nature of the P2P networks, and dynamic nature of the data/network. The proposed algorithm takes a two-step approach. In the monitoring phase, the algorithm checks if the model ‘quality’ is acceptable by using an efficient local algorithm. This is then used as a feedback loop to sample data from the network and rebuild the GMM when it is outdated. We present thorough experimental results to verify our theoretical claims.
Evaluating Prognostics Performance for Algorithms Incorporating Uncertainty Estimates
Uncertainty Representation and Management (URM) are an integral part of the prognostic system development.1As capabilities of prediction algorithms evolve, research in developing newer and more competent methods for URM is gaining momentum.2Beyond initial concepts, more sophisticated prediction distributions are obtained that are not limited to assumptions of Normality and unimodal characteristics. Most prediction algorithms yield non-parametric distributions that are then approximated as known ones for analytical simplicity, especially for performance assessment methods. Although applying the prognostic metrics introduced earlier with their simple definitions has proven useful, a lot of information about the distributions gets thrown away. In this paper, several techniques have been suggested for incorporating information available from Remaining Useful Life (RUL) distributions, while applying the prognostic performance metrics. These approaches offer a convenient and intuitive visualization of algorithm performance with respect to metrics like prediction horizon and α-λ performance, and also quantify the corresponding performance while incorporating the uncertainty information. A variety of options have been shortlisted that could be employed depending on whether the distributions can be approximated to some known form or cannot be parameterized. This paper presents a qualitative analysis on how and when these techniques should be used along with a quantitative comparison on a real application scenario. A particle filter based prognostic framework has been chosen as the candidate algorithm on which to evaluate the performance metrics due to its unique advantages in uncertainty management and flexibility in accommodating non-linear models and non-Gaussian noise. We investigate how performance estimates get affected by choosing different options of integrating the uncertainty estimates. This allows us to identify the advantages and limitations of these techniques and their applicability towards a standardized performance evaluation method.
Algorithms and their Impact on Integrated Vehicle Health Management - Chapter 7
This chapter discussed some of the algorithmic choices one encounters when designing an IVHM system. While it would be generally desirable to be able to pick a particular set of algorithms for a particular problem, the reality is a bit more complex. Depending on the budget, the performance requirements, the computational constraints, sensor availability, access to historical data, operational and environmental conditions, robustness to changing system configurations, algorithm maintenance needs, etc., no one algorithm will perform best in all situations. Indeed, it is necessary to evaluate these constraints during the algorithm design process and determine the best choice on a case-by-case analysis. The trade-offs between different choices are very real, and sometimes no solution can be found, which means that some of the constraints have to be relaxed. The simplest solution is generally preferred over a more complex one, but it is also important to consider that there is no free lunch. Finally, any health management solution also has to undergo verification and validation (V&V) and, in some cases, certification. Some of these issues are topics of other chapters in this book.
Comparison of Prognostic Algorithms for Estimating Remaining Useful Life of Batteries
We were interested here in particular in conditions where un-modeled effects are present as manifested by the different degradation curve at 45°C. Although all algorithms were given the same amount of information to the degree practical, there were considerable differences in performance. Specifically, the combined Bayesian regression-estimation approach implemented as a RVM-PF framework has significant advantages over conventional methods of RUL estimation like ARIMA and EKF. ARIMA, being a purely data-driven method, does not incorporate any physics of the process into the computation, and hence ends up with wide uncertainty margins that make it unsuitable for long-term predictions. Additionally, it may not be possible to eliminate all non-stationarity from a dataset even after repeated differencing, thus adding to prediction inaccuracy. EKF, though robust against non-stationarity, suffers from the inability to accommodate un-modeled effects and can diverge quickly as shown. We did not explore other variations of the Kalman Filter that might provide better performance such as the unscented Kalman Filter. The Bayesian statistical approach, on the other hand, appears to be well suited to handle various sources of uncertainties since it defines probability distributions over both parameters and variables and integrates out the nuisance terms. Also, it does not simply provide a mean estimate of the time-to-failure; rather it generates a probability distribution over time that best encapsulates the uncertainties inherent in the system model and measurements and in the core concept of failure prediction.
A Local Scalable Distributed EM Algorithm for Large P2P Networks
his paper describes a local and distributed expectation maximization algorithm for learning parameters of Gaussian mixture models (GMM) in large peer-to-peer (P2P) environments. The algorithm can be used for a variety of well-known data mining tasks in distributed environments such as clustering, anomaly detection, target tracking, and density estimation to name a few, necessary for many emerging P2P applications in bioinformatics, webmining and sensor networks. Centralizing all or some of the data to build global models is impractical in such P2P environments because of the large number of data sources, the asynchronous nature of the P2P networks, and dynamic nature of the data/network. The proposed algorithm takes a two-step approach. In the monitoring phase, the algorithm checks if the model ‘quality’ is acceptable by using an efficient local algorithm. This is then used as a feedback loop to sample data from the network and rebuild the GMM when it is outdated. We present thorough experimental results to verify our theoretical claims.
ARC Code TI: Multiple Kernel Anomaly Detection (MKAD) Algorithm
The Multiple Kernel Anomaly Detection (MKAD) algorithm is designed for anomaly detection over a set of files.
ADAPTIVE FAULT DETECTION ON LIQUID PROPULSION SYSTEMS WITH VIRTUAL SENSORS: ALGORITHMS AND ARCHITECTURES
Prior to the launch of STS-119 NASA had completed a study of an issue in the flow control valve (FCV) in the Main Propulsion System of the Space Shuttle using an adaptive learning method known as Virtual Sensors. Virtual Sensors are a class of algorithms that estimate the value of a time series given other potentially nonlinearly correlated sensor readings. In the case presented here, the Virtual Sensors algorithm is based on an ensemble learning approach and takes sensor readings and control signals as input to estimate the pressure in a subsystem of the Main Propulsion System. Our results indicate that this method can detect faults in the FCV at the time when they occur. We use the standard deviation of the predictions of the ensemble as a measure of uncertainty in the estimate. This uncertainty estimate was crucial to understanding the nature and magnitude of transient characteristics during startup of the engine. This paper overviews the Virtual Sensors algorithm and discusses results on a comprehensive set of Shuttle missions and also discusses the architecture necessary for deploying such algorithms in a real-time, closed-loop system or a human-in-the-loop monitoring system. These results were presented at a Flight Readiness Review of the Space Shuttle in early 2009.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.