Indian Agricultural Statistics Research Institute, New Delhi-110 012, (India) e-mail: arrao@iasri.res.in
An empirical investigation is carried out to compare the performance of four multivariate outlier detection methods. Multivariate breeding data of maize genotypes obtained from Directorate of Maize Research, New Delhi is used to estimate the mean vectors and dispersion matrices for clean and outlier data. These estimates are used to simulate the samples from clean and outlier populations under different situations, viz. shift outliers, scale outliers and both shift and scale outliers. All the four methods are compared using total probability of wrong identification [(1-p1) + p2], based on highest probability of correctly identified outliers (p1) and lowest probability of wrongly identified outliers (p2). The method, which identifies lowest total probability of wrong identification of outliers is judged as the best one. It is found that the method M4 (Filzmoser et al., 2008) can be used with advantage for detection of multivariate shift outliers and M3 (Filzmoser et al., 2005) for multivariate scale and both scale & shift outliers.
Multivariate outliers, Breeding data, Simulation