scieee AI-readable full text Open interactive document viewer

Improved motion segmentation based on shadow detection

Kampel, Martin; Hanbury, Allan; Blauensteiner, Philipp; Wildenauer, Horst

Abstract

In this paper, we discuss common colour models for background subtraction and problems related to their utilisation are discussed. A novel approach to represent chrominance information more suitable for robust background modelling and shadow suppression is proposed. Our method relies on the ability to represent colours in terms of a 3D-polar coordinate system having saturation independent of the brightness function; specifically, we build upon an Improved Hue, Luminance, and Saturation space (IHLS). The additional peculiarity of the approach is that we deal with the problem of unstable hue values at low saturation by modelling the hue-saturation relationship using saturation-weighted hue statistics. The effectiveness of the proposed method is shown in an experimental comparison with approaches based on RGB, Normalised RGB and HSV.

Full text

Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 Improved motion segmentation based on shadow detection M. Kampel∗and H. Wildenauer+and P. Blauensteiner∗and A. Hanbury∗ ∗Pattern Recognition and Image Processing Group, Vienna University of Technology, Favoritenstr.9, A-1040 Vienna, Austria +Automation and Control Institute, Vienna University of Technology, Gusshausstr.27, A-1040 Vienna, Austria Abstract In this paper, we discuss common colour models for background subtraction and problems related to their utilisation are discussed. A novel approach to represent chrominance information more suitable for robust background modelling and shadow suppression is proposed. Our method relies on the ability to represent colours in terms of a 3D-polar coordinate system having saturation independent of the brightness function; specifically, we build upon an Improved Hue, Luminance, and Saturation space (IHLS). The additional peculiarity of the approach is that we deal with the problem of unstable hue values at low saturation by modelling the hue-saturation relationship using saturation-weighted hue statistics. The effectiveness of the proposed method is shown in an experimental comparison with approaches based on RGB, Normalised RGB and HSV. Key Words: Motion detection, shadow detection, background subtraction, colour spaces. 1 Introduction The underlying step of visual surveillance applications like target tracking and scene understanding is the detection of moving objects. Background subtraction algorithms are commonly applied to detect these objects of interest by the use of statistical colour background models. Many present systems exploit the properties of the Normalised RGB to achieve a certain degree of insensitivity with respect to changes in scene illumination. Hong and Woo [1] apply the Normalised RGB space in their background segmentation system. McKenna et al. [2] use this colour space in addition to gradient information for their adaptive background subtraction. The AVITRACK project [3] utilises Normalised RGB for change detection and adopts the shadow detection proposed by Horprasert et al. [4]. Beside Normalised RGB, representations of the RGB colour space in terms of 3D-polar coordinates (hue, saturation, and brightness) are used for change detection and shadow suppresion in surveillance applications. Franc¸ois and Medioni [5] suggest the application of HSV for background modelling for real-time video segmentation. In their work, a complex set of rules is introduced to reflect the relevance of observed and background colour information during change detection and model update. Cucchiara et al. [6] propose a RGB-based background model which they transform to the HSV representation in order to utilise the properties of HSV chrominance information for shadow suppression. Correspondence to: <[email protected]> Recommended for acceptance by U. Pal and P. Nagabhushan ELCVIA ISSN:1577-5097 Published by Computer Vision Center / Universitat Aut`onoma de Barcelona, Barcelona, Spain 2M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 Our approach differs from the aforementioned in the way that we build upon the IHLS colour space, which is more suitable for background subtraction. Additionally, we propose the application of saturation-weighted hue statistics [7] to deal with unstable hue values at weakly saturated colours. Also, a technique to efficiently classify changes in scene illumination (e.g. shadows), modelling the relationship between saturation and hue has been devised. The remainder of this paper is organised as follows: Section 2 reviews the Normalised RGB and the Improved Hue, Luminance and Saturation (IHLS) colour space. Furthermore it gives a short overview over circular colour statistics, which have to be applied on the hue as angular value. Section 3 presents how these statistics can be applied in order to model the background in image sequences. In Section 4 we describe metrics for the performance evaluation of our motion segmentation. The conducted experiments and their results are presented in Section 5. Section 6 concludes this paper and gives an outlook. 2 Colour Spaces In this section, the Normalised RGB and IHLS colour spaces used in this paper are described. It also gives a short overview over circular colour statistics and a review of saturation weighted hue statistics. 2.1 Normalised RGB The Normalised RGB space aims to separate the chromatic components from the brightness component. The red, green and blue channel can be transformed to their normalised counterpart by using the formulae l=R+G+B, r =R/l, g =G/l, b =B/l (1) if l6= 0 and r=g=b= 0 otherwise [8]. One of these normalised channels is redundant, since by definition r,g, and bsum up to 1. Therefore, the Normalised RGB space is sufficiently represented by two chromatic components (e.g. rand g) and a brightness component l. From Kender [9] it is known that the practical application of Normalised RGB suffers from a problem inherent to the normalisation; namely, that noise (such as, e.g. sensor or compression noise) at low intensities results in unstable chromatic components. For an example see Figure 1. Note the artefacts in dark regions such as the bushes (top left) and the shadowed areas of the cars (bottom right). Figure 1: Examples of chromatic components. Lexicographically ordered - Image from the PETS2001 dataset, it’s normalised blue component b, normalised saturation (cylindrical HSV), IHLS saturation. M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 3 2.2 IHLS Space The Improved Hue, Luminance and Saturation (IHLS) colour space was introduced in [10]. It is obtained by placing an achromatic axis through all the grey (R=G=B) points in the RGB colour cube, and then specifying the coordinates of each point in terms of position on the achromatic axis (brightness), distance from the axis (saturation s) and angle with respect to pure red (hue θH). The IHLS model is improved with respect to the similar colour spaces (HLS, HSI, HSV, etc.) by removing the normalisation of the saturation by the brightness. This has the following advantages: (a) the saturation of achromatic pixels is always low and (b) the saturation is independent of the brightness function used. One may therefore choose any function of R,Gand Bto calculate the brightness. It is interesting that this normalisation of the saturation by the brightness, which results in the colour space having the shape of a cylinder instead of a cone or double-cone, is usually implicitly part of the transformation equations from RGB to a 3D-polar coordinate space. This is mentioned in one of the first papers on this type of transformation [11], but often in the literature the equations for a cylindrically-shaped space (i.e. with normalised saturation) are shown along with a diagram of a cone or double-cone (for example in [12, 13]). Figure 1 shows a comparison of the different formulations of saturation. The undesirable effects created by saturation normalisation are easily perceivable, as some dark, colourless regions (eg., the bushes and the side window of the driving car) reach higher saturation values than their more colourfull surroundings. Also, note the artefacts resulting from the singularity of the saturation at the black vertex of the RGB-cube (again, the bushes and the two bottom right cars). The following formulae are used for the conversion from RGB to hue θH, luminance yand saturation sof the IHLS space: s= max(R, G, B)−min(R, G, B) y= 0.2125R+ 0.7154G+ 0.0721B crx=R−G+B 2, cry=√3 2(B−G) cr =qcr2 x+cr2 y(2) θH=     undefined if cr = 0 arccos crx cr elseif cry≤0 360◦−arccos crx cr else where crxand crydenote the chrominance coordinates and cr ∈[0,1] the chroma. The saturation assumes values in the range [0,1] independent of the hue angle (the maximum saturation values are shown by the circle on the chromatic plane in Figure 2). The chroma has the maximum values shown by the dotted hexagon in Figure 2. When using this representation, it is important to remember that the hue is undefined if s= 0, and that it does not contain much useable information when sis low (i.e. near to the achromatic axis). 2.3 Hue statistics In a 3D-polar coordinate space, standard (linear) statistical formulae can be utilised to calculate statistical descriptors for brightness and saturation coordinates. The hue, however, is an angular value, and consequently the appropriate methods from circular statistics are to be used. Now, let θH i,i= 1,...,n be nobservations sampled from a population of angular hue values. Then, the vector hipointing from O= (0,0)Tto the point on the circumference of the unit circle, corresponding to θH i, is given by the Cartesian coordinates (cos θH i,sin θH i)T.∗ ∗Note that, when using the IHLS space (Eq. 3), no costly trigonometric functions are involved in the calculation of hi, since cos(θH i) = crx/cr and sin(θH i) = −cry/cr. 4M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 Figure 2: The chromatic plane of the IHLS color space. The mean direction θHis defined to be the direction of the resultant of the unit vectors h1,...,hnhaving directions θH i. That is, we have θH= arctan2 (S,C),(3) where C= n X i=1 cos θH i,S= n X i=1 sin θH i(4) and arctan2(y, x)is the four-quadrant inverse tangent function. The mean length of the resultant vector R=√C2+S2 n.(5) is an indicator of the dispersion of the observed data. If the nobserved directions θH icluster tightly about the mean direction θHthen Rwill approach 1. Conversely, if the angular values are widely dispersed Rwill be close to 0. The circular variance is defined as V= 1 −R (6) While the circular variance differs from the linear statistical variance in being limited to the range [0,1], it is similar in the way that lower values represent less dispersed data. Further measures of circular data distribution are given in [14]. 2.4 Saturation-weighted hue statistics The use of statistics solely based on the hue has the disadvantage of ignoring the tight relationship between the chrominance components hue and saturation. For weakly saturated colours the hue channel is unimportant and behaves unpredictably in the presence of colour changes induced by image noise. In fact, for colours with zero saturation the hue is undefined. As one can see in Figure 2, the chromatic components may be represented by means of Cartesian coordinate vectors ciwith direction and length given by hue and saturation respectively. Using this natural approach, we introduce the aforementioned relationship into the hue statistics by weighting the unit hue vectors hiby their corresponding saturations si. Now, let (θH i, si),i= 1,...,nbe npairs of observations sampled from a population of hue values and associated saturation values. We proceed as described in Section 2.3, with the difference that instead of calculating the resultant of unit vectors, the vectors ci, which we will dub chrominance vectors throughout this paper, have length si. M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 5 That is, we weight the vector components in Eq. 4 by their saturations si Cs= n X i=1 sicos θH i,Ss= n X i=1 sisin θH i,(7) and choose the mean resultant length of the chrominance vectors (for other possible formulations see, e.g. [7]) to be Rn=pC2 s+S2 s n.(8) Consequently, for the mean resultant chrominance vector we get cn= (Cs/n, Ss/n)T.(9) Here, the length of the resultant is compared to the length obtained if all vectors had the same direction and maximum saturation. Hence, Rngives an indication of the saturations of the vectors which gave rise to the mean of the chrominance vector, as well as an indication of the angular dispersion of the vectors. To test if a mean chrominance vector cnis similar to a newly observed chrominance vector, we use the Euclidean distance in the chromatic plane: D=q(cn−co)T(cn−co),(10) with co=soho. Here, hoand sodenote the observed hue vector and saturation respectively. 3 The IHLS Background Model With the foundations laid out in Section 2.4 we proceed with devising a simple background subtraction algorithm based on the IHLS colour model and saturation-weighted hue statistics. Specifically, each background pixel is modelled by its mean luminance µyand associated standard deviation σy, together with the mean chrominance vector cnand the mean Euclidean distance σDbetween cnand the observed chrominance vectors (see Eq. 10). On observing the luminance yo, saturation so, and a Cartesian hue vector hofor each pixel in a newly acquired image, the pixel is classified as foreground if: |(yo−µy)|> ασy∨ kcn−sohok> ασD(11) where αis the foreground threshold, usually set between 2and 3.5. In order to decide whether a foreground detection was caused by a moving object or by its shadow cast on the static background, we exploit the chrominance information of the IHLS space. A foreground pixel is considered as shaded background if the following three conditions hold: yo< µy∧ |yo−µy|< βµy,(12) so−Rn< τds (13) khoRn−cnk< τh,(14) where Rn=kcnk(see Eq. 8,). These equations are designed to reflect the empirical observations that cast shadows cause a darkening of the background and usually lower the saturation of a pixel, while having only limited influence on its hue. The first condition (Eq. 12) works on the luminance component, using a threshold βto take into account the strength of the predominant light source. Eq. 13 performs a test for a lowering in saturation, as proposed by Cucchiara et al. [6]. Finally, the lowering in saturation is compensated by scaling the observed hue vector hoto the same length as the mean chrominance vector cnand the hue deviation is tested using the Euclidean distance (Eq. 14). 6M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 This, in comparison to a check of angular deviation (see Eq. 31 or [6]), also takes into account the model’s confidence in the learned chrominance vector. That is, using a fixed threshold τhon the Euclidean distance relaxes the angular error-bound in favour of stronger hue deviations at lower model saturation value Rn, while penalising hue deviations for high saturations (where the hue is usually more stable). 4 Metrics for Motion Segmentation The quality of motion segmentation can in principle be described by two characteristics. Namely, the spatial deviation from the reference segmentation, and the fluctuation of spatial deviation over time. In this work, however, we concentrate on the evaluation of spatial segmentation characteristics. That is, we will investigate the capability of the error metrics listed below, to describe the spatial accuracy of motion segmentations. •Detection rate (DR) and false alarm rate (FR) DR =T P FN +TP (15) FR =FP N−(FN +TP)(16) where TP denotes the number of true positives, FN the number of false negatives, FP the number of false positives, and Nthe total number of pixels in the image. •Misclassification penalty (MP) The obtained segmentation is compared to the reference mask on an object-by-object basis; misclassified pixels are penalized by their distances from the reference objects border [15]. MP =MPfn +MPfp (17) with MPfn =PNf n j=1 dj fn D(18) MPfp =PNf p k=1 dk fp D(19) Here, dj fn and dk fp stand for the distances of the jth false negative and kth false positive pixel from the contour of the reference segmentation. The normalised factor Dis the sum of all pixel-to-contour distances in a frame. •Rate of misclassifications (RM) The average normalised distance of detection errors from the contour of a reference object is calculated using [16]: RM =RMfn +RMfp (20) with RMfn =1 Nfn Nfn X j=1 dj fn Ddiag (21) RMfp =1 Nfp Nfp X k=1 dk fp Ddiag (22) M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 7 Nfn and Nfp denote the number of false negative and false positive pixels respectively. q Ddiag is the diagonal distance within the frame. •Weighted quality measure (QMS) This measure quantifies the spatial discrepancy between estimated and reference segmentation as the sum of weighted effects of false positive and false negative pixels [17]. QMS =QMSfn +QMSfp (23) with QMSfn =1 N Nfn X j=1 wfn(dj fn)dj fn (24) QMSfp =1 N Nfp X k=1 wfp(dk fp)dk fp (25) Nis the area of the reference object in pixels. Following the argument that the visual importance of false positives and false negatives is not the same, and thus they should be treated differently, the weighting functions wfp and wfn were introduced: wfp(dfp) = B1+B2 dfp +B3 (26) wfn(dfn) = C·dfn (27) In our work for a fair comparison of the change detection algorithms with regard to their various decision parameters, receiver operating characteristics (ROC) based on detection rate (DR) and false alarm rate (FR) were utilised. 5 Experiments and Results We compared the proposed IHLS method with three different approaches from literature. Namely, a RGB background model using either NRGB- (RGB+NRGB), or HSV-based (RGB+HSV) shadow detection, and a method relying on NRGB for both background modelling and shadow detection (NRGB+NRGB). All methods were implemented using the Colour Mean and Variance approach tomodel the background [18]. A pixel is considered foreground if |co−µc|> ασcfor any channel c, where c∈ {r, g, l}for the Normalised RGB and c∈ {R, G, B}for the RGB space respectively. ocdenotes the observed value, µcits mean, σcthe standard deviation, and αthe foreground threshold. The tested background models are maintained by means of exponentially weighted averaging [18] using different learning rates for background and foreground pixels. During the experiments the same learning and update parameters were used for all background models, as well as the same number of training frames. For Normalised RGB (RGB+NRGB,NRGB+NRGB), shadow suppression was implemented based on Horprasert’s approach [3, 4]. Each foreground pixel is classified as shadow if: lo< µl∧lo> βµl |ro−µr|< τc∧ |go−µg|< τc(28) where βand τcdenote thresholds for the maximum allowable change in the intensity and colour channels, so that a pixel is considered as shaded background. 8M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 In the HSV-based approach (RGB+HSV) the RGB background model is converted into HSV (specifically, the reference luminance µv, saturation µs, and hue µθ) before the following shadow tests are applied. A foreground pixel is classified as shadow if: β1≤vo µv≤β2(29) so−µs≤τs(30) |θH o−µθ| ≤ τθ(31) The first condition tests the observed luminance vofor a significant darkening in the range defined by β1and β2. On the saturation soa threshold on the difference is performed. Shadow lowers the saturation of points and the difference between images and the reference is usually negative for shadow points. The last condition takes into account the assumption that shading causes only small deviation of the hue θH o[6]. For the evaluation of the algorithms, three video sequences were used. As an example for a typical indoor scene Test Sequence 1, recorded by an AXIS-211 network camera, shows a moving person in a stairway. For this sequence, ground truth was generated manually for 35 frames. Test Sequence 2 was recorded with the same equipment and shows a person waving books in front of a coloured background. For this sequence 20 ground truth frames were provided. Furthermore in Test Sequence 3 the approaches were tested on 25 ground truth frames from the PETS2001 dataset 1 (camera 2, testing sequence). Example pictures of the dataset can be found in Figure 3. (a) (b) (c) Figure 3: Evaluation dataset: Test Sequence 1 (a), Test Sequence 2 (b), Test Sequence 3 (c) For a dense evaluation, we experimentaly determined suitable ranges for all parameters and sub-sampled them in ten steps. Figure 7 shows the convex hulls of the points in ROC space obtained for all parameter combinations. We also want to point out that RGB+HSV was tested with unnormalised and normalised saturation; however, since the normalised saturation consistently performed worse, we omit the results in the ROC for clarity of presentation. As one can see, our approach outperforms its competitors on Test Sequence 1. One reason for this is the insensitivity of the RGB+NRGB and NRGB+NRGB w.r.t. small colour differences at light, weakly saturated colours. RGB+HSV, however, suffered from the angular hue test reacting strongly to unstable hue values close to the achromatic axis. For conservative thresholds (i.e. small values for τcor τθ) all three approaches either detected shadows on the wall as foreground, or, for larger thresholds failed to classify the beige t-shirt of the person as forground. Figure 4 shows output images from Test Sequence 1. We present the source image (a), the ground truth image (b), the resulting image from our approach (c), and the resulting images from the algorithms we compared with. I.a. it is shown that the shirt of the person in image (c) is detected with higher precision as in the images (d), (e), and (f), where it is mostly marked as shadow. For Test Sequence 2 the advantageous behaviour of our approach is even more evident. Although the scene is composed of highly saturated, stable colours, RGB+NRGB and NRGB+NRGB show rather poor results, again stemming from their insufficient sensitivity for bright colours. RGB+HSV gave better results, but could not take full advantage of the colour information. Similar hue values for the books and the background resulted M. Kampel et al. / Electronic Letters on Computer Vision and Image Analysis 6(3):1-12, 2007 9 (a) (b) (c) (d) (e) (f) Figure 4: Output images Test Sequence 1:Source Image (a), Ground Truth (b), Our Approach (c), RGB+NRGB (d), NRGB+NRGB (e), RGB+HSV (f) in incorrectly classified shadow regions. Figure 5 shows output images from Test Sequence 2. Especially the lower left part of the images (c), (d), (e), and (f) visualizes a better performance of the IHLS approach. (a) (b) (c) (d) (e) (f) Figure 5: Output images Test Sequence 2:Source Image (a), Ground Truth (b), Our Approach (c), RGB+NRGB (d), NRGB+NRGB (e), RGB+HSV (f) The Test Sequence 3 sequence shows the problems of background modelling using NRGB already mentioned in Section 2. Due to the low brightness and the presence of noise in this scene, the chromatic components are unstable and therefore the motion detection resulted in an significantly increased number of false positives. RGB+NRGB and our approach exhibit similar performance (our approach having the slight edge), mostly relying on brightness checks, since there was not much useable information in shadow regions. RGB+HSV performed less well, having problems to cope with the unstable hue information in dark areas. Figure 6 shows output images Test Sequence 3.