scieee Science in your language
[en] (orig)

Discovery of motifs to forecast outlier occurrence in time series

Abstract

The forecasting process of real-world time series has to deal with especially unexpected values, commonly known as outliers. Outliers in time series can lead to unreliable modeling and poor forecasts. Therefore, the identification of future outlier occurrence is an essential task in time series analysis to reduce the average forecasting error. The main goal of this work is to predict the occurrence of outliers in time series, based on the discovery of motifs. In this sense, motifs will be those pattern sequences preceding certain data marked as anomalous by the proposed metaheuristic in a training set. Once the motifs are discovered, if data to be predicted are preceded by any of them, such data are identified as outliers, and treated separately from the rest of regular data. The forecasting of outlier occurrence has been added as an additional step in an existing time series forecasting algorithm (PSF), which was based on pattern sequence similarities. Robust statistical methods have been used to evaluate the accuracy of the proposed approach regarding the forecasting of both occurrence of outliers and their corresponding values. Finally, the methodology has been tested on six electricity-related time series, in which most of the outliers were properly found and forecasted.

Read accessible full text

Discovery of motifs to forecast outlier occurrence in time series

Author: Martínez Álvarez, Francisco; Troncoso Lora, Alicia; Riquelme Santos, José Cristóbal; Aguilar Ruiz, Jesús Salvador
Publisher: Elsevier
Year: 2011
DOI: 10.1016/j.patrec.2011.05.002
Source: https://idus.us.es/bitstreams/5e251d59-5070-411a-a144-65f3566cee4b/download
Disco e y o mo i s o o ecas ou lie occu ence in ime se ies
F. Ma ínez–Ál a ez , A. T oncoso, J.C. Riquelme , J.S. Aguila –Ruiz
Keywo ds:
Time se ies o ecas ing
Pa e n ecogni ion
Mo i s
Ou lie s
abs ac
The o ecas ing p ocess o eal-wo ld ime se ies has o deal wi h especially unexpec ed alues, com-
monly known as ou lie s. Ou lie s in ime se ies can lead o un eliable modeling and poo o ecas s.
The e o e, he iden ifica ion o u u e ou lie occu ence is an essen ial ask in ime se ies analysis o
educe he a e age o ecas ing e o . The main goal o his wo k is o p edic he occu ence o ou lie s
in ime se ies, based on he disco e y o mo i s. In his sense, mo i s will be hose pa e n sequences p e-
ceding ce ain da a ma ked as anomalous by he p oposed me aheu is ic in a aining se . Once he mo i s
a e disco e ed, i da a o be p edic ed a e p eceded by any o hem, such da a a e iden ified as ou lie s,
and ea ed sepa a ely om he es o egula da a. The o ecas ing o ou lie occu ence has been added
as an addi ional s ep in an exis ing ime se ies o ecas ing algo i hm (PSF), which was based on pa e n
sequence simila i ies. Robus s a is ical me hods ha e been used o e alua e he accu acy o he p oposed
app oach ega ding he o ecas ing o bo h occu ence o ou lie s and hei co esponding alues. Finally,
he me hodology has been es ed on six elec ici y- ela ed ime se ies, in which mos o he ou lie s we e
p ope ly ound and o ecas ed.
1. In oduc ion
This wo k p oposes a new s a egy o p edic he occu ence o
ou lying da a in ime se ies, as well as p o iding accu a e o ecas s
o hem. I is wo h highligh ing ha he goal o his me hodology
is o o ecas hei appea ance, ins ead o de ec ing hem in an
al eady known se o alues, which is a common goal in obus s a-
is ics (Ma onna e al., 2007). The majo i y o obus s a is ical
echniques pe o m a pos e io i de ec ion, ha is, hey de e mine
whe he a da um is an ou lie o no , bu once i has al eady
occu ed. Howe e , a compa ison wi h hese echniques will be
o he u mos impo ance in o de o e alua e he accu acy o he
p oposed me aheu is ic.
A gene al-pu pose o ecas ing algo i hm, called PSF, was p e-
sen ed in Ma ínez-Ál a ez e al. (in p ess). I s main ea u e lied
in pe o ming a disc e iza ion o he ime se ies by means o ce -
ain clus e ing echnique. Then, i only used he gene a ed labels
o make p edic ions. Based on ha p e ious disc e iza ion, his
wo k a emp s o disco e pa e n sequences (hence o h called
mo i s) in he his o ical da a o o ecas he occu ence o ou lie s
and hei associa ed alues. This no el me hodology is inse ed in
he gene al scheme o PSF.
The p edic ion o ou lie s plays an impo an ole as w ong mod-
els and poo o ecas s a e ob ained when igno ing ou lie s. This
wo k p esen s a me aheu is ic o disco e mo i s in ime se ies
and, hen,i da a obep edic eda ep ecededbyanyo hesedisco -
e ed mo i s, conside hese da a as ou lie . The mo i s a e de e -
mined du ing he aining phase as hose pa e n sequences ha
p ecede da a wi h ema kable o ecas ing e o . Thus, he exis ing
PSF algo i hm is modified by adding a new mo i ex ac ion s ep.
The enhanced e sion is capable o p edic ing he appea ance o
such ou lie s wi h g ea eliabili y when he mo i s ex ac ion s ep
is added. In ac , he app oach has been success ully es ed on six
eal-wo ld elec ici y- ela ed ime se ies, in pa icula , on ene gy
p ices and demand o h ee di e en ma ke s, eaching sensi i i y
alues g ea e han 82%, and specifici y alues g ea e han 95%.
Fu he mo e, esul s abou he e ec o ou lie s on he a e age
o ecas ing e o s a e epo ed o all he six ime se ies, exhibi ing
ema kable o ecas ing e o educ ion.
Despi e he as a ie y o wo ks ela ed o ou lie s de ec ion
and mo i s disco e y in ime se ies, he e is no app oach in ime
se ies in o de o o ecas he occu ence o ou lie s, o he au ho s’
knowledge.
The emaining o he pape is o ganized as ollows. A e iew o
he mos ecen ly published wo ks ega ding ene gy ime se ies
o ecas ing, mo i s disco e y and ou lie s de ec ion can be ound
in Sec ion 2. Sec ion 3p o ides o mal desc ip ion o sensi i e
e ms, and p esen s a b ie explana ion o he o iginal algo i hm.
As o Sec ion 4, i in oduces he p oposed me hodology, showing
how o inse he ou lie occu ence o ecas ing in he o iginal algo-
i hm’s gene al scheme. The esul s ob ained o he six elec ici y
p ices and demand ime se ies a e epo ed and discussedin Sec ion
5. Finally, Sec ion 6summa izes he main conclusions achie ed.
2. Rela ed wo k
This sec ion p o ides use ul and ecen e e ences abou he
h ee main opics in ol ed in his pape : Ene gy ime se ies o e-
cas ing, mo i s disco e y in empo al da a and obus s a is ical
me hods o de ec ou lie s. Fo he sake o cla i y, hese opics ha e
been sepa a ed in h ee di e en sec ions.
2.1. Ene gy ime se ies o ecas ing
The in e es o analyzing elec ici y p ice ime se ies esides in
he p og essi e de egula ion o elec ic powe ma ke s. Fu he -
mo e, elec ici y p ice ime se ies possess ce ain ea u es ha u n
he p edic ion in o a di ficul ask: non-cons an mean/ a iance
and equen ly ou lie occu ences. Fo his eason, elec ici y-p o-
duce companies wan op imized bidding s a egies as well as
needing assessmen abou he isk o us ing o ecas s (Plazas
e al., 2005).
On he o he hand, he p ocess o o ecas ing he quan i y o
elec ici y equi ed o a specific geog aphical a ea du ing a ime
pe iod is called load o ecas ing o demand o ecas ing. This p o-
cess is key since cu en echnology allows o s o e a limi ed
amoun o elec ici y in ba e ies. The e o e, he demand o ecas -
ing plays an impo an ole o elec ici y powe supplie s because
bo h excess and insu ficien ene gy p oduc ion may lead o in-
c eased cos s and a significan educ ion o p ofi s.
The pu sui o accu a e o ecas ing in elec ici y p ice ime se -
ies has mo i a ed esea ch wo ks by many au ho s (Agga wal e al.,
2009). Thus, he use o mixed models was p oposed in Ga cía-
Ma os e al. (2007) o o ecas p ices o di e en ho izons o
p edic ion. Also ema kable was he wo k in oduced in T oncoso
e al. (2007) ha , by means o weigh ed nea es neighbo s me hod-
ology, o ecas ed nex -day elec ici y p ices. Also, he use o an
a ificial neu al ne wo k o ulfil he same goal can be ound in Pino
e al. (2008). E en he use o classical au o eg essi e models has
been ecen ly used o o ecas p ices in se e al ma ke s (We on
and Misio ek, 2008).
On he con a y, he au ho s in T oncoso e al. (2004) p oposed
a weigh ed nea es neighbo s-based me hodology o o ecas elec-
ici y demand. The p oposed app oach was es ed o e he nex -
day Spanish load o ecas ing. Also, he au ho s in El-Telbany and
El-Ka mi (2008) o ecas ed he Jo danian elec ici y demand wi h
an a ificial neu al ne wo k, which was ained by means o pa i-
cle swa m op imiza ion echniques. This ma ke was also s udied
in Bad an e al. (2008) bu , his ime, he au ho s p e e ed o con-
cen a e on sho and medium- e m load o ecas ing by using
eg ession models. Finally, Wang and Wang (2008) p oposed a
new p edic ion app oach based on suppo ec o machines
(SVM) echniques wi h a p e ious selec ion o ea u es om da a
se s by using an e olu iona y me hod.
The disco e y o ou lie s in elec ici y p ice ime se ies has also
been widely discussed in li e a u e. Thus, he au ho s in Lu e al.
(2005) p oposed a model based on he analysis o se e al a iables
(among which he elec ici y demand is highligh ed) by means o
Bayesian classifica ion (BC) and simila i y sea ching echniques.
A hyb id me hodology ha combined SVMs and BC was de eloped
in Wu e al. (2006) o classi y bo h spikes and no mal elec ici y
p ices. Al e na i ely, a da a mining amewo k based on SVM and
p obabili y classifie s was desc ibed in Zhao e al. (2007) wi h
he aim o o ecas ing spikes in p ices accu a ely.
By con as , he p edic ion o peaks in elec ici y demand was
add essed in Saini (2008). This wo k o ecas ed demand peaks up
o se en days ahead using eed o wa d neu al ne wo k and adap-
i e backp opaga ion lea ning me hods. In he wo k in oduced in
Ismail e al. (2009), he au ho s de eloped a ule-based me hod
ha combined eg ession models and uzzy sys ems o analyze
daily elec ici y peak load demands in Malaysia. Also, he au ho s
in Hyndman and Fan (2010) desc ibed a semi-pa ame ic addi i e
model o disco e ela ionships be ween he demand and exogen
a iables. The app oach was applied o long- e m peaks o he
Sou h Aus alian ma ke .
2.2. Mo i s disco e y
The disco e y o mo i s in con inuous da a, also known as unc-
ional da a in many wo ks (Valde ama, 2008), was o iginally o -
malized in Lin e al. (2002), in which he au ho s in oduced
se e al algo i hms o mine mo i s in ime se ies, among which
he k-mo i algo i hm highligh s. Howe e , he main d awback o
his algo i hm is i s dependence on a p e-fixed pa e n leng h.
La e , he au ho s in Tang and Liao (2008) p oposed a modified
e sion ha imp o ed, p ecisely, his ea u e. Fu he mo e, hey
gene a ed o iginal pa e ns by conside ing he disco e ed mo i s.
The de ec ion o on-line mo i s in con inuous da a has also been
add essed. Pa icula ly, a new me hodology o de ec on-line
mo i s in ime se ies by combining p obabilis ic models and poly-
nomial leas -squa es app oxima ions was p oposed in Fuchs e al.
(2009). This opic was also s udied in Mueen and Keogh (2010) in
which he au ho s ound and main ained ime se ies mo i s om
obo ics, online comp ession and wildli e managemen domains.
Besides hese goals, di e en objec i es ha e been ulfilled by
disco e ing mo i s ecen ly. Hence, he wo k p esen ed in Tanaka
e al. (2005) p oposed an algo i hm de o ed o disco e mo i s
based on he minimum desc ip ion leng h p inciple. The app oach
also allowed o ob ain mo i s om mul i-dimensional ime se ies
da a by using p incipal componen analysis. Mueen e al. (2009)
p oposed in 2009 an exac algo i hm o find ime se ies mo i s
much as e han b u e- o ce sea ching s a egies do. Also, an ap-
p oach based on ee-cons uc ion sea ch o disco e mo i s in
mul i a ia e ime se ies was p oposed in Wang e al. (2010) and
applied o senso y da ase s.
Since S o mo (2000) fi s e iewed s a egies o find DNA mo i s
(meaning ul base sequence pa e ns ha iden i y binding si es
esponsible o ansc ip ion ac o s) in 2000, a la ge amoun o
algo i hms ha e been de eloped. Thus, an ensemble algo i hm
a emp ing o disco e egula o y mo i s in DNA sequences was
p oposed in Hu and Kiha a (2006). Ano he algo i hm was p o-
posed in Wijaya e al. (2007) which, gi en a se o sequences, exe-
cu es mdi e en mo i finde s, each o hem epo ing nmo i s.
Finally, Sha o and Mino u (2009) p esen ed CisFinde , a so wa e
ha gene a es a comp ehensi e lis o mo i s en iched in a se o
DNA sequences and desc ibes hem wi h posi ion equency
ma ices.
2.3. Robus s a is ical me hods o de ec ou lie s
The p oblem o a pos e io i ou lie s de ec ion in ime se ies has
been widely s udied in he li e a u e, and aced om many di e -
en poin s o iew. In ac , he exis ence o e en ew ou lie s
usually leads o inaccu a e models and no sa is ac o y o ecas s
(Galeano e al., 2006), since hey may deeply influence he
es ima es ha classical me hods p opose (Ca ne o e al., 2007).
Fo his eason, he e is a la ge amily o obus s a is ical
me hods (Rousseeuw and Hube , 2011) ha deal wi h ou lie s
and, pa icula ly, p opose app oaches o de ec hei exis ence in
he da ase s subjec ed o analysis. Gelpe e al. p oposed an
adap ed e sion o he classical exponen ial and Hol -Win e s
smoo hing me hodologies, p o iding hem wi h obus ness
(Gelpe e al., 2010). Ano he e sion o a obus mul i a ia e
exponen ial smoo hing applied o ime se ies can be ound in
C oux e al. (2010). Following wi h classical me hods, a wo k ha
enhanced ARMA by adding obus ness can be ound in Mule e al.
(2009), in which he au ho s succeeded in limi ing he e ec o
ou lying da a o he ime s amp in which hey happen.
Suppo ec o machines (SVM) ha e also been adap ed o deal
wi h ou lie s. Ac ually an app oach o model ime se ies using o-
bus SVM was p oposed in Camps-Valls e al. (2004). In pa icula ,
he au ho s claimed ha hei p oposal p o ides s able models
and allows he analysis o models’ memo y dep h. Recen ly, a wo-
s eps me hodology was p oposed in Chuang and Lee (2011) ha
combined heuseo obus SVM o emo eanomalousobse a ions,
and non- obus SVM o ob ain es ima es om ha educed da ase .
Many o hese p oposals ha e been implemen ed and eely dis-
ibu ed in so wa e packages. Bu om all o hem he e a e wo
ha highligh . LIBRA (Ve bo en and Hube , 2005) is a Ma lab li-
b a y o obus analysis, ha con ains (among o he s) obus
co a iance es ima ion, eg ession, p incipal componen analysis,
p incipal componen eg ession o pa ial leas squa es, as well
as me hodologies o de ec ou lying obse a ions in da ase s. The
TOMCAT oolbox Daszykowski e al. (2007), also de eloped in he
Ma lab en i onmen , includes almos he same me hods ha
LIBRA does, bu also includes a g aphical in e ace.
3. Fundamen als
This Sec ion fi s defines some e ms in o de o p e en possi-
ble misin e p e a ions in sensi i e e ms. Since he p oposed
me hodology is based on an exis ing algo i hm, his Sec ion also
p o ides a b ie summa y o he ma hema ical undamen als
unde lying he PSF algo i hm. No e ha a mo e de ailed explana-
ion can be ound in Ma ínez-Ál a ez e al. (in p ess).
3.1. Defini ions
This wo k uses ce ain concep s –such as ou lie o mo i – ha
can be in e p e ed in many di e en senses, depending on he
applica ion o e en he au ho . Fo his eason, his Sec ion p o-
ides a o mal defini ion o hese sensi i e e ms.
Defini ion 1 (Hou ly ime se ies). An hou ly ime se ies Tis a se o
eal- alued da a in successi e o de , occu ing e e y hou . In his
wo k, T=[
1
,...,
p
], whe e pis he leng h o he ime se ies and
usually a mul iple o 24.
Defini ion 2 (Daily ime se ies). F om an hou ly ime se ies, a daily
ime se ies Dis o med by uples in R
24
,D=[d
1
,...,d
p/24
], whe e
d
i
=[
24(i1)+1
,...,
24i
].
Defini ion 3 (Label). In his wo k, he e m label is used o iden i y
a se o possible ca ego ical alues. Thus, L¼ l
1
;...;l
K
g, whe e Kis
a p e-fixed numbe .
Defini ion 4 (Sequence). A sequence Sis a se o labels occu ing in
successi e o de . In his wo k, S=[s
1
,...,s
q
], whe e qis he leng h
o he sequence and s
i
2L.
Defini ion 5 (Ou lie ). Gi en a es se , an ou lie is an obse a ion
which appea s o be inconsis en wi h he es o he da a, ela i e
o an assumed model (E e i , 2006). Ou lie s a e usually
ep esen ed by a bina y andom a iable
i
, o i=1,...,p ha
models hei occu ence (
i
= 1 i i occu s, and
i
= 0 o he wise),
and by ano he eal andom a iable z
i
, o i=1,...,p ha models
hei magni ude (Ma onna e al., 2007). Al hough di e en ou lie
ypes can be ound in he li e a u e, only he addi i e ou lie model
is conside ed in his wo k, due o he na u e o he s udied da a:
i
¼x
i
þ
i
z
i
;ð1Þ
whe e
i
he obse ed alue, and x
i
he i  h cleaned da a modeled
by any app oach.
Defini ion 6 (Mo i ). A mo i M
W
is a sequence o Wconsecu i e
labels conside ed o occu jus be o e an ou lie , whe e Wis he
p e-fixed leng h o he sequence. In addi ion, M
W
is a subsequence
ound in Sand, consequen ly: M
W
¼½s
0
1
;...;s
0
W
, whe e s
0
i
2L.
3.2. Time se ies o ecas ing: he PSF algo i hm
The PSF algo i hm is a gene al-pu pose ime se ies o ecas ing
algo i hm whose main ea u e is ha i only makes use o ce ain
labels –ob ained by means o a clus e ing p ocess– o o ecas a bi-
a y ho izons o p edic ion. Howe e , he ou pu is no composed
by labels bu by eal alues.
PSF can deal wi h an a bi a y numbe o samples pe day. How-
e e , specifically in his pape , he ime se ies conside ed consis s
o wen y- ou samples pe day. Tha is, gi en he hou ly alues up
o day i o a ime se ies, he PSF algo i hm p o ides he 24 hou ly
alues co esponding o day i+ 1. Fo mally, le d
i
2R
24
be a ec o
ha comp ises he 24 hou ly alues o a ce ain day, i.
Fi s , PSF applies clus e ing echniques o such da a in o de o
assign a label o each day. Fo mally, i uses a unc ion F
K
ha as-
signs a label l
i
2L o he alues d
i
2Do each day by means o a
clus e ing p ocess, F
K
:D!L, ha is, e e y 24 h a e iden ified by
a label. Once Kis fixed, his p ocess ans o ms he daily ime se ies
Din o a sequence o labels S, hus disc e izing he o iginal da a. Le
l
i
be he label assigned o he day iob ained by means o he appli-
ca ion o a clus e ing echnique. Le S
i
W
be he labels’ subsequence
o Wconsecu i e days, om day ibackwa d:
S
i
W
¼½l
iðW1Þ
;l
iðW2Þ
;...;l
i1
;l
i
ð2Þ
whe e he leng h o he window, W, is a pa ame e o be
de e mined.
Le W
⁄
be he leng h o he window de e mined by PSF. Fo a day i
and leng h o window W
⁄
, he PSF algo i hm sea ches o he subse-
quences o labels which a e exac ly equals o S
i
W

in he da ase , p o-
iding he equal subsequences se , ES, defined by he equa ion,
ESði;W

Þ¼ days j2Dsuch ha S
j
W

¼S
i
W

no
ð3Þ
I is wo h ema king ha i no subsequence equal o S
i
W

was ound
in he da ase , ha is, ES(i,W
⁄
)=;, he leng h o he window would
dec ease by one uni , W
0
=W
⁄
1, and he PSF would sea ch o
subsequences equal o S
i
W
0
. This p ocess may be epea ed un il
any subsequence is ound, ha is, ES(i,W)–;.
The e o e, he W
⁄
consecu i e labels ha p ecede he day o be
p edic ed a e ex ac ed and sea ched o in he his o ical da a.
Once all occu ences o S
i
W

a e ound, he 24 hou ly alues o he
day i+ 1 a e p edic ed by a e aging he eal alues ound immedi-
a ely a e each S
i
W

ma ch. Ma hema ically,
^
d
iþ1
ðW

Þ¼ 1
#ESði;W

ÞX
j2ESði;W

Þ
d
jþ1
ð4Þ
Finally, he daily e o o any day iis defined by:
e
day
ði;W

Þ¼j
^
d
i
ðW

Þd
i
jð5Þ
4. Ou lie o ecas ing in ime se ies
This sec ion explains he me hodology p oposed o imp o e he
o ecas ing p ocess p o ided by he PSF algo i hm. The disco e y o
mo i s is included in he a o emen ioned algo i hm as a c ucial
s ep o o ecas ing he occu ence o ou lie s and, hen, p o iding
accu a e es ima es o such anomalous obse a ions.
The alue o wo pa ame e s had o be de e mined in he PSF
p ocess: The numbe o clus e s Kand he leng h o he window
W. Wi h ega d o K, he new app oach ac s exac ly he same as
wha was p oposed in he o iginal PSF, ha is, i applies h ee
well-known alidi y indices –Silhoue e, Dunn and Da ies–Boul-
din– and de e mines he op imal numbe o clus e s by means o
a majo i y o e sys em.
On he o he hand, he n old c oss- alida ion is used o ob-
ain he op imal alue o W. Twel e olds ha e been c ea ed in his
wo k (n= 12) o all he da ase s, whe e each old ep esen s a
mon h. The e o e he aining se consis s o one yea . The
12  old c oss– alida ion is hen e alua ed. The o ecas ing e o s
a e calcula ed in e e y old by a ying he leng h o W. Fo each
window size W, he mon hly e o s a e deno ed by e
mon h
(W) and
a e calcula ed as ollows:
e
mon h
ðWÞ¼ 1
#mon h X
i2mon h
e
day
ði;WÞð6Þ
o W=1,...,W
max
and W
max
= 10, since no longe sequences we e
ound in daily ime se ies. Then, he a e age e o s a e calcula ed
o each window size as ollows,

eðWÞ¼1
nX
mon h
e
mon h
ðWÞð7Þ
whe e n= 12 and mon h ={Jan,...,Dec}.
The W
⁄
selec ed is he one ha minimizes he a e age e o co -
esponding o he 12 olds (mon hs) e alua ed.
W

¼a gmin 
eðWÞg wi h W¼1;...;W
max
ð8Þ
I is now –jus a e he aining s ep and be o e he p edic ion p o-
cess– ha he disco e y o mo i s plays a c ucial ole, as i a emp s
a o ecas ing he occu ence o anomalous days in ime se ies.
Specifically, he mo i s o be ound a e hose which gene a e a
p edic ion e o g ea e han he a e age e o in he c oss– alida-
ion p ocess. The e o e, a se o days ibelonging o he aining se
(TS) ha sa isfies e
day
ði;W

Þ>
eðW

Þis cons uc ed. This se , CS o
candida es se , ga he s all he candida e days o be p eceded by a
sequence ha will e en ually be a mo i . Fo mally,
CS ¼ i2TS such ha e
day
ði;W

Þ>
eðW

Þg ð9Þ
Ne e heless, no all he sequences ha p ecede hese candida es
ha e he same p obabili y o be e en ually conside ed as ou lie
p ecu so s, since he associa ed e o s ange om alues close o
he mean e o ( hese candida e sequences should be e en ually
disca ded by he app oach) o significan ly high alues. Fo his ea-
son, each candida e is co-labeled by using clus e ing echniques,
mo e specifically, he K-means algo i hm. The decision on how
many clus e s ha e o be c ea ed is always an open ques ion and
many indices could be used. Howe e , i is wo hless o ha e a la ge
numbe o clus e s and he e o e, only h ee concep ual classes will
be c ea ed: A class ha ga he s days wi h low e o s (C
l
) o nea es
e o s o he 
eðW

Þ, a class con aining medium e o s (C
m
) and, fi-
nally, a class de o ed o iden i y high e o s (C
h
) o a hes e o s
o he 
eðW

Þ. Thus, o all days i2C
l
,j2C
m
and k2C
h
he ollowing
inequali ies a e ulfilled:
e
day
ðk;W

Þ>e
day
ðj;W

Þ>e
day
ði;W

Þ>
eðW

Þð10Þ
Fig. 1 illus a es an imagina y e o dis ibu ion in he TS, de e min-
ing he candida es ha will e en ually o m CS as he union o can-
dida es in C
l
,C
m
and C
h
. In o he wo ds, CS =C
l
SC
m
SC
h
.
A p io i easoning e eals ha hose candida es belonging o C
h
mus be mo e p obable o be p eceded by mo i s p eceding ou lie s
han he candida es in C
l
o C
m
. Resul s co esponding o each clus-
e o da a will be sepa a ely analyzed in Sec ion 5.
The nex s ep consis s o compu ing he sequences o labels
occu ing be o e he candida es in o de o de e mine which se-
quences will be conside ed ou lie p ecu so s, since no all hese
sequences will be mo i s. Hence, he app oach has o decide
whe he he sequences p eceding he candida es a e mo i s
esponsible o ou lie s o no . In pa icula , he sequences ha
only appea be o e he candida es will be conside ed mo i s p e-
ceding an ou lie occu ence. Tha is, i a sequence o labels p eced-
ing a candida e is ound p io o any o he day ha does no belong
o he CS, he sequence is disca ded and no conside ed o be a mo-
i . Thus, a se o mo i s MS is defined by:
MS ¼S
i
W

such ha i2CS and ESði;WÞ#CS
no
ð11Þ
No e ha MS can be also exp essed as ollows: MS =M
l
SM
m
SM
h
,
whe e M
l
,M
m
and M
h
a e he disco e ed mo i s associa ed o he se-
quences ound in he classes C
l
,C
m
, and C
h
espec i ely.
Fig. 4 depic s he p ocess o disco e ing he sequences o labels
ha ep esen he mo i s p eceding ou lie s. This figu e shows an
illus a i e example in which six clus e s we e c ea ed, K=6
(labels a e digi s 1 o 6). The e o e, he ime se ies appea s disc e -
ized, making only use o six di e en labels o be assigned one o
each day. The labels in bold e e o hose days ini ially included
in CS, ha is, hose ha ob ained o ecas ing e o g ea e han
he a e age. By con as , he labels ollowed by a bulle a e hose
days ha do no belong o CS bu a e p eceded by a sequence equal
o ano he ha p ecedes a candida e. Then, he h ee labels p eced-
ing each day in CS (W
⁄
= 3 in his example) a e ex ac ed. Finally,
he cases in which he sequences be o e he candida es also
appea ed be o e any day no belonging o CS we e disca ded.
O he wise, hese sequences a e conside ed mo i s o pa e n se-
quences p eceding an especially unexpec ed alue. Fo his pa ic-
ula example, no e ha only he sequence {1,5,4} would ha e been
conside ed as mo i .
The goals a e now: (i) o o ecas he occu ence o an ou lie
and (ii) p o ide accu a e es ima ion o i . Thus, once MS is
50 100 150 200 250 300 350
0
5
10
15
Days
E o (%)
E o dis ibu ion in TS
Cl
Ch
Cm
MEAN
Fig. 1. Illus a i e dis ibu ion o candida es in C
l
,C
m
and C
h
.
cons uc ed, he gene al scheme o o ecas ing is as ollows. O igi-
nal PSF jus ex ac ed S
W

and sea ched o i in he his o ical da a.
Bu now, be o e i s sea ch, i has o be de e mined i his pa e n
sequence ma ches any o he mo i s o ming MS. Gi en his si ua-
ion, wo cases may a ise:
(1) S
W

does no ma ch any o he mo i s in MS. The app oach
de e mines ha he day o be o ecas ed is no an ou lie
and i would con inue wi h no mal PSF p ocedu e. Tha is,
^
d
iþ1
ðW

Þis calcula ed as PSF does om Eq. (4).
(2) S
W

ma ches any o he mo i s in MS. The app oach de e -
mines ha he day o be o ecas ed is an ou lie . In his case
he challenge is o p o ide accu a e es ima ions o hese
obse a ions, i.e. o p o ide accu a e alues o z
i
. To ulfill
his goal, a simple s a egy is p oposed: To a e age he alue
o he ou lie s ound by he p oposed app oach in he his o -
ical da a. Fo mally, he se o ou lying days om he aining
se is defined by:
OS ¼ i2TS such ha S
i
W

2MSgð12Þ
Then, he o ecas o he ou lie is based on he a pos e io i
de ec ed ou lie s in he aining se :
^
d
iþ1
ðW

Þ¼ 1
#OS X
i2OS
d
iþ1
ð13Þ
whe e #OS is he numbe o ou lie s de ec ed by he app oach
in he his o ical da a ( he numbe o elemen s in OS), and ^
d
i
he alues o he ime se ies o hese ou lying days ha o m
OS.
Fig. 2 illus a es he en i e p ocess o p edic ion when he dis-
co e y o mo i s is included in he PSF algo i hm. No e ha his
s ep has o be pe o med immedia ely a e he clus e ing (c ea ion
o he sequence o labels) and be o e he o ecas ing. In addi ion,
he s eps co esponding o disco e y o mo i s and p edic ion a e
u he de ailed in Fig. 3.
Finally, he pseudocode o he p oposed me hodology is p e-
sen ed in Fig. 5, and ha o he disco e y o mo i s p ocess in Fig. 6.
5. Resul s
This sec ion p esen s he esul s ob ained by he applica ion o
he p oposed me hodology o six ime se ies. The mo i s disco e y
p ocess o six eal-wo ld ime se ies is desc ibed in Sec ion 5.1.
Then, a s a is ical analysis has been ca ied ou o de e mine he
alidi y o he assump ions made when o ecas ing he occu ence
o ou lie s in he ime se ies. This analysis can be ound in Sec ion
5.2. Finally, o compa e he esul s ob ained, Sec ion 5.3 epo s
a e age o ecas ing e o s o he new me hodology and o he
echniques.
5.1. Mo i s disco e y in eal-wo ld ime se ies
The disco e y o mo i s on eal-wo ld ime se ies is now de-
sc ibed. In pa icula , six public elec ici y- ela ed ( h ee o p ices
and h ee o demand) ime se ies ha e been conside ed o show
ha he p oposed me hodology p ope ly wo ks on di e en da a-
se s. Thus, he new app oach has been applied o he Spanish
(OMEL), New Yo ke (NYISO) and Aus alian (ANEM) ma ke s,
whose da a a e a ailable on-line in Spanish Elec ici y P ice Ma ke
Ope a o (h p://www.omel.es), The New Yo k Independen Sys-
em Ope a o (h p://www.nyiso.com) and Aus alia’s Na ional
Elec ici y Ma ke (h p://www.nemmco.com.au), espec i ely.
The o ecas ing p ocess is applied o he yea 2006 o he h ee
ma ke s, wi h a his o ical da a o one yea and wi h a ho izon o
p edic ion o one mon h. As wel e mon hs a e going o be e alu-
a ed o each ma ke , he me hodology is going o be es ed on 72
da ase s. Gi en his si ua ion, e e y ime a mon h is o ecas ed he
aining se changes. Fo ins ance, when Janua y 2006 is o ecas ed,
he aining se comp ises he whole yea o 2005. Howe e , when
Feb ua y 2006 is o ecas ed he his o ical da a anges om Feb u-
a y 2005 o Janua y 2006, and so on.
These changes in he aining se in ol e changes in he config-
u a ion o PSF. Fi s o all, bo h Kand Wha e o be de e mined
acco ding o he me hodology p esen ed in Sec ion 4.Table 1 sum-
ma izes he alues o hese pa ame e s o he six ma ke s in he
yea 2006.
The esul s a e he mo i s ex ac ion s ep o he h ee ma ke s
a e summa ized inTables2–4.Tha is, a summa yo all encoun e ed
classes, sequences and mo i s can be ound in heses Tables o he
Spanish,NewYo ke andAus alianma ke s, espec i ely.Howe e ,
onlyelec ici y p ices esul so Janua y 2006 o heAus alianma -
ke a e now desc ibed as he explana ion o he emaining ele en
mon hs o each yea and ma ke is simila . The e o e, all he com-
men s abou he esul s p o ided below e e o p ices shown in
Table 4. Fi s , he pa ame e s o be se in he PSF a e equal o:
(K,W) (3,6), acco ding o Table 1. The CS can be now cons uc ed.
Fo hispu pose, he 
eðWÞ(seeEq.(7))has o beconside ed since he
candida es a e hose days belonging o he aining se (Janua y o
Decembe 2005) ha ob ained an e o g ea e han 
eðWÞ.The alue
o he mean e o , calcula ed acco ding o he me hodology in
Sec ion 4is 
eð6Þ¼5:81%. The e o e, CS would be o med by
all days in 2005 wi h o ecas ing e o g ea e han 5.81%.
Now he h ee classes a e cons uc ed by applying K-means,
wi h K= 3 as men ioned in Sec ion 4. The h ee clus e s a e defined
as: C
l
is he class ha con ains he candida es days wi h e o om
5.81% o 7.13%, C
m
he one ha ga he s he candida es wi h e o s
anging om 7.13% o 9.47% and C
h
he class ha con ains he can-
dida es wi h e o s g ea e han 9.47%. The e o dis ibu ion,
acco ding o hese h ee clus e s, is shown in Fig. 7.
F om he365 dayso 2005 ha comp ise he ainingse ,137(see
Table 4, ow 1: #C
l
+#C
m
+#C
h
= 101 + 32 + 4 = 137) had an e o
g ea e han 5.81% so he cons uc ed CS con ains 137 candida es
days. Once he candida es a e selec ed, he numbe o di e en se-
quences ha gene a ed hem a e conside ed. F om he candida es
in C
l
, 5 di e en sequences we e ound (S
l
= 5); om he candida es
in C
m
,3(S
m
= 3) and om he candida es in C
h
,2(S
h
= 2). This ac in-
ol es ha om all he K
W
= 729 possible sequences, only 10 caused
e o s g ea e han he a e age.
No e ha he e we e si ua ions in which a pa icula sequence
appea ed be o e di e en candida es ha belong o di e en
classes. Fo hese cases, only he sequence ha appea ed in he
class wi h highe associa ed e o (C
l
was de o ed o include days
wi h lowe e o s, C
m
o medium e o s and C
h
o highe ones) was
coun ed.
Finally, he numbe o mo i s ha iden i y ou lie s a e de e -
mined. F om he sequences C
l
, only one appea ed exclusi ely as
Da a Clus e ing MOTIFS
DISCOVERY P edic ion
Labeled da a
Mo i s
Inse
p edic ed sample
Mo e
days?
Yes
End
No
Fig. 2. Illus a ion o he p oposed me hodology.

an ou lie p ecu so , M
l
= 1. Wi h e e ence o he sequences in C
m
,
one ou o h ee, M
m
= 1. Las , bo h sequences in C
h
we e exclusi e,
M
h
=2.
Fig. 8 is p o ided o de e mine he use ulness o di iding he CS
in o h ee g oups. These his og ams show he mo i s dis ibu ion
along wi h C
l
,C
m
and C
h
and, o be p ecise, he ela ion be ween di -
e en sequences and mo i s o each class and ma ke , exp essed
as %. Thus, each ba is calcula ed by di iding he numbe o di e -
en sequences ha p ecede he candida es and he numbe o mo-
i s ha a e finally selec ed. As i is possible o obse e, he
p obabili y ha a sequence becomes a mo i is di ec ly ela ed
wi h he e o associa ed o he candida e o which i p ecedes.
Fo ins ance, he pe cen age o mo i s o he ANEM’s elec ici y
p ice ime se ies a e 20.26% o he sequences in C
l
, 29.82% o
he sequences in C
m
and 52.94% o he sequences in C
h
.
The mo i s ound in Janua y 2006, ep esen ed as a nume ical
sequence o labels, a e shown in Tables 5 and 6 co esponding o
p ices and demand, espec i ely. Fo ins ance, no e ha he fi s
mo i ound in ANEM’s p ices is M
1
l
¼ 1;3;2;2;3;1g. Each label
is ep esen ed by one o he Kclus e s gene a ed du ing he ain-
ing o he PSF (whe e K= 3 in his case) and iden ifies 24 h. As he
leng h o he window was se o W= 6, hese six labels in ac ep-
esen 144 h.
Figs. 9 and 10 illus a e he mos ep esen a i e mo i s ound in
all he ma ke s when o ecas ing Janua y 2006. Ac ually, hese mo-
i s ep esen he ime se ies alues ha will p ecede an ou lie . As
each mo i was ep esen ed by Wconsecu i e labels (see Defini ion
6), hese figu es depic he alues associa ed o e e y label, which
ha e been ob ained by means o clus e ing echniques.
Cons uc
CS
Clus e ing
CS
Find
mo i s
Candida es
DISCOVERY OF MOTIFS
{Cl,C
m,C
h}
Sea ch o S
in his o ical da a
PREDICTION
Mo i s se , MS
Window's leng h, W*
any mo i ?
NO
YES
Fo ecas s
S ma ches
w* w*
Sea ch o mo i s Fo ecas da a
om Eq. (13)
in his o ical da a
Fo ecas da a
om Eq. (4)
Fo ecas s
(Ou lie occu ence)
Mo i s se , MS
Window's leng h,
W
*
T aining se , TS
Numbe o clus e s, K Se W W*
Fig. 3. De ail o disco e y o mo i s and p edic ion s eps.
Fig. 4. Illus a i e example o mo i s disco e y.
Fig. 5. A gene al scheme o he p oposed me hodology.
Rega ding he elec ici y p ices, no e ha o he Spanish ma -
ke , six mo i s we e ound; six o New Yo k, and ou o he Aus-
alian ma ke . As o he elec ici y demand, he numbe o mo i s
ound we e se en, fi e and ou o he Spanish, New Yo k and
Aus alian ma ke s, espec i ely. In ac ual ac , he cu es in hese
figu es ep esen he a e age e olu ion o he fi e (OMEL), ou
(NYISO) and six (ANEM) days p io o an ou lie o ecas in p ices
ime se ies and he a e age e olu ion o he h ee (OMEL and NYI-
SO) and ou (ANEM) days p io o an ou lie o ecas in demand
ime se ies.
Fig. 6. Pseudocode o he mo i s disco e y.
Table 1
Se ing he PSF. OMEL e e s o he Spanish ma ke , NYISO o he New Yo ke ma ke and ANEM co esponds o he Aus alian ma ke .
Mon h P ices Demand
OMEL NYISO ANEM OMEL NYISO ANEM
KWKWKWKWKWKW
Janua y 4 5 5 4 3 6 7 3 4 3 6 4
Feb ua y 4 5 5 3 3 6 7 4 3 5 6 5
Ma ch 4 5 5 4 3 6 7 3 4 4 4 3
Ap il 4 5 5 4 4 6 6 4 4 5 4 4
May 45 63 46 54 44 64
June 6 4 5 3 3 6 6 3 4 5 5 5
July 5 5 6 4 3 5 6 4 3 5 6 4
Augus 6 4 6 3 4 6 5 4 4 3 5 4
Sep embe 6 4 5 3 3 6 7 3 5 4 6 5
Oc obe 6 4 5 4 3 6 5 4 5 4 5 4
No embe 6 4 5 4 3 6 5 3 4 4 6 4
Decembe 5 5 5 3 3 5 6 3 5 3 4 6
Table 2
Mo i s dis ibu ion o Spanish ma ke s.
Mon h P ices Demand
#C
l
(S
l
)[M
l
]#C
m
(S
m
)[M
m
]#C
h
(S
h
)[M
h
]#C
l
(S
l
)[M
l
]#C
m
(S
m
)[M
m
]#C
h
(S
h
)[M
h
]
Janua y 98(7)[1] 25(4)[3] 8(2)[2] 103(6)[2] 35(6)[2] 6(3)[3]
Feb ua y 87(6)[0] 31(7)[2] 5(2)[2] 121(8)[1] 22(5)[3] 3(2)[1]
Ma ch 73(5)[2] 16(3)[1] 8(1)[1] 99(7)[1] 30(7)[4] 6(1)[0]
Ap il 103(9)[1] 30(6)[1] 6(3)[2] 83(4)[0] 17(5)[2] 5(3)[3]
May 65(4)[0] 51(6)[0] 10(4)[2] 101(6)[2] 26(8)[3] 5(3)[3]
June 97(6)[0] 38(5)[2] 4(0)[0] 124(7)[3] 27(5)[1] 7(1)[1]
July 180(8)[2] 27(3)[1] 12(5)[3] 97(8)[2] 13(3)[0] 8(4)[3]
Augus 101(8)[3] 26(5)[0] 9(4)[2] 113(11)[4] 29(6)[3] 11(3)[2]
Sep embe 110(8)[1] 25(5)[3] 5(0)[0] 105(9)[2] 19(5)[4] 3(2)[2]
Oc obe 108(7)[1] 23(4)[2] 6(1)[0] 98(8)[3] 32(3)[2] 4(1)[1]
No embe 120(9)[1] 40(6)[0] 6(3)[1] 99(9)[2] 20(5)[3] 9(3)[3]
Decembe 169(10)[2] 38(9)[3] 10(3)[1] 114(8)[2] 24(3)[2] 12(2)[0]
Fig. 11 shows ANEM’s elec ici y p ices o July 2006. This
mon h is illus a ed since i p esen s la ge ou lie s. As o his
mon h W= 5 (see Table 1), he numbe o labels conside ed o ou -
lie occu ence o ecas ing is fi e. Also, he numbe o mo i s ound
in TS we e fi e (see Table 4): M
1
l
¼ 1;2;3;3;1g;M
2
l
¼ 2;2;
3;1;3g;M
3
l
¼ 3;1;1;1;1g;M
4
h
¼ 2;1;3;3;1gand M
5
h
¼ 2;1;1;
1;3gThus, g ey ba s ep esen ou lie s a pos e io i de ec ed by
means o he obus s a is ical me hod p esen ed in Gelpe e al.
(2010), ha is, days 3, 13, 18 and 22 July we e he ou lie s iden i-
fied. I can be app ecia ed ha sequences M
3
l
;M
4
h
and M
1
l
a e h ee
o he mo i s ound in TS ha e en ually p eceded days 13, 18 and
22 July, espec i ely. Fu he mo e, 3 d July was p eceded by he
mo i M
5
h
. Only he las wo labels o his mo i (1,3) a e depic ed
in Fig. 11 because he fi s h ee labels (2,1,1) co espond o days
in June. Finally, no e ha one o he mo i s ound in TS, M
2
l
, did no
occu when o ecas ing July 2006.
5.2. E alua ion o ou lie occu ence o ecas ing
Once all he MS ha e been cons uc ed, he p oposed me hod
p edic a p io i i he day o be o ecas ed will be an ou lie . This
Sec ion is de o ed o s a is ically quan i y he ou lie occu ences
Table 3
Mo i s dis ibu ion o New Yo k ma ke s.
Mon h P ices Demand
#C
l
(S
l
)[M
l
]#C
m
(S
m
)[M
m
]#C
h
(S
h
)[M
h
]#C
l
(S
l
)[M
l
]#C
m
(S
m
)[M
m
]#C
h
(S
h
)[M
h
]
Janua y 101(8)[3] 34(3)[1] 12(2)[2] 96(9)[1] 28(3)[1] 14(4)[3]
Feb ua y 92(11)[2] 36(4)[1] 14(3)[2] 88(9)[3] 14(2)[1] 7(2)[2]
Ma ch 89(7)[2] 45(5)[1] 11(2)[1] 101(9)[2] 33(2)[0] 9(3)[2]
Ap il 110(13)[3] 21(4)[1] 9(5)[4] 93(8)[2] 29(3)[2] 11(4)[2]
May 121(7)[2] 31(2)[1] 6(3)[1] 114(9)[3] 30(4)[2] 13(6)[3]
June 142(5)[1] 32(0)[0] 7(0)[0] 103(7)[0] 26(3)[1] 10(4)[2]
July 92(10)[3] 41(5)[1] 18(4)[4] 86(5)[1] 29(4)[2] 8(3)[2]
Augus 84(7)[2] 39(6)[2] 9(6)[5] 76(7)[4] 14(2)[2] 4(1)[1]
Sep embe 107(7)[0] 40(4)[0] 10(1)[1] 84(8)[1] 21(4)[1] 8(3)[3]
Oc obe 141(9)[2] 28(3)[0] 4(0)[0] 95(6)[0] 32(4)[2] 9(3)[1]
No embe 99(12)[3] 32(8)[2] 15(4)[3] 115(8)[3] 38(6)[2] 15(6)[2]
Decembe 87(8)[1] 44(3)[1] 8(0)[0] 109(8)[3] 29(4)[2] 12(4)[3]
Table 4
Mo i s dis ibu ion o Aus alian ma ke s.
Mon h P ices Demand
#C
l
(S
l
)[M
l
]#C
m
(S
m
)[M
m
]#C
h
(S
h
)[M
h
]#C
l
(S
l
)[M
l
]#C
m
(S
m
)[M
m
]#C
h
(S
h
)[M
h
]
Janua y 101(5)[1] 32(3)[1] 4(2)[2] 131(7)[0] 46(7)[2] 5(3)[2]
Feb ua y 165(5)[0] 25(6)[1] 9(1)[1] 115(6)[1] 56(7)[3] 6(3)[3]
Ma ch 133(8)[2] 13(2)[0] 3(0)[0] 92(6)[0] 72(9)[1] 7(3)[3]
Ap il 190(13)[3] 8(2)[0] 11(4)[2] 102(7)[2] 64(5)[0] 5(4)[3]
May 187(17)[5] 13(4)[2] 6(3)[1] 123(7)[2] 41(3)[1] 11(4)[0]
June 169(12)[5] 22(5)[1] 3(3)[1] 183(16)[3] 33(3)[0] 4(4)[2]
July 172(22)[3] 9(0)[0] 2(2)[2] 117(9)[1] 40(5)[0] 8(5)[3]
Augus 142(20)[2] 34(3)[2] 6(4)[2] 139(11)[2] 41(4)[3] 7(3)[1]
Sep embe 102(18)[4] 20(8)[3] 12(2)[1] 110(8)[2] 27(3)[2] 9(3)[3]
Oc obe 81(13)[2] 53(13)[4] 9(2)[0] 142(12)[3] 42(8)[3] 5(1)[1]
No embe 112(9)[1] 43(7)[1] 14(5)[4] 137(12)[2] 38(7)[3] 7(3)[2]
Decembe 121(11)[3] 39(4)[2] 13(6)[2] 134(10)[4] 29(5)[2] 8(4)[2]
0 50 100 150 200 250 300 350
0
2
4
6
8
10
12
14
16
18
20
Days
E o (%)
9.47%
7.13%
5.81%
Ch
Cm
Cl
Fig. 7. Fo ecas ing e o s in TS dis ibu ed in C
l
,C
m
and C
h
ob ained by applying he K-means o he candida e se CS wi h K=3.
F. Ma ínez–Ál a ez e al. /Pa e n Recogni ion Le e s 32 (2011) 1652–1665 1659
p ope ly o ecas ed by he p oposed me hodology. Thus, he qual-
i y pa ame e s used o e alua e he accu acy o he app oach a e
fi s in oduced in Sec ion 5.2.1 and, hen, he conduc ed s a is ical
analysis is epo ed in Sec ion 5.2.2.
5.2.1. Pa ame e s o quali y
The pa ame e s used o assess he accu acy o he app oach a e
now in oduced. A pos e io i analysis has been ca ied ou o
de e mine he exis ence o ou lie s in all examined se ies. In pa -
icula , he obus me hod p oposed in Gelpe e al. (2010) (he ea -
e called RHW o simplici y) o de ec ou lie s in ime se ies has
been conside ed. Hence, a o ecas o an ou lie occu ence is said
o be p ope ly made by he p oposed app oach, i RHW also poin s
he obse a ion as anomalous. Thus, a p io i o ecas ing ( he p o-
posed app oach in his wo k) is compa ed o a pos e io i de ec ion
( he me hod p oposed in Gelpe e al. (2010)).
No e ha he au ho s in Gelpe e al. (2010) de e mined ha
ou lie s a e hose da a ha do no ulfil any o he wo bounds hey
define:
UB
¼^
y
þ2^
ð14Þ
LB
¼^
y
2^
ð15Þ
whe e UB
and LB
a e he uppe and lowe bounds espec i ely, ^
y
is
he fi ed alue, and ^
is he s anda d de ia ion o he eg ession
esiduals hey ob ain.
Hence, in subsequen equa ions, ue posi i es o TP is he num-
be ou lie occu ences p ope ly o ecas ed, ha is, he numbe o
days p eceded by a mo i in MS ha a e ou lie s acco ding o RHW;
ue nega i es o TN is he numbe o days ha we e no p eceded
by a mo i in MS and we e no conside ed ou lie by RHW ei he ;
alse posi i es o FP is he numbe o days p eceded by a mo i in
MS ha we e no conside ed ou lie by RHW; and alse nega i es
o FN is he numbe days no conside ed ou lie s (no p eceded
by any mo i in MS) and e en ually conside ed ou lie s by RHW.
Acco ding o hese defini ions, he sensi i i y is he p obabili y
ha a mo i disco e ed p ecedes a eal ou lie s. I s o mula is de-
fined as ollows:
Sensi i
i y ¼TP
TP þFN ð16Þ
Ano he ele an pa ame e is he specifici y, o he a io o se-
quences p eceding he day o be o ecas ed p ope ly disca ded by
he app oach. The ma hema ical exp ession is:
Speci ici y ¼TN
TN þFP ð17Þ
The posi i e p edic i e alue (PPV) is he p obabili y ha a o e-
cas ed ou lie is indeed a eal one. I s o mula is:
PPV ¼TP
TP þFP ð18Þ
0.00%
10.00%
20.00%
30.00%
40.00%
50.00%
60.00%
70.00%
80.00%
90.00%
OMEL
ANEM
NYISO
CmCh
Cl
0.00%
10.00%
20.00%
30.00%
40.00%
50.00%
60.00%
70.00%
80.00%
90.00%
OMEL
ANEM
NYISO
ClCmCh
Fig. 8. Mo i s dis ibu ion in C
l
,C
m
and C
h
.
Table 5
Mo i s ound in TS o he elec ici y p ice when
o ecas ing Janua y 2006.
Ma ke Mo i
OMEL M
1
l
¼ 1;1;3;4;4g
M
2
m
¼ 3;4;4;1;3g
M
3
m
¼ 3;4;4;2;1g
M
4
m
¼ 2;3;4;4;1g
M
5
h
¼ 1;1;1;2;3g
M
6
h
¼ 3;2;3;3;4g
NYISO M
1
l
¼ 1;3;5;4g
M
2
l
¼ 2;2;3;1g
M
3
l
¼ 5;1;1;3g
M
4
m
¼ 3;4;2;5g
M
5
h
¼ 3;2;4;1g
M
6
h
¼ 3;1;5;1g
ANEM M
1
l
¼ 1;3;2;2;3;1g
M
2
m
¼ 2;1;3;3;3;1g
M
3
h
¼ 3;2;1;1;2;3g
M
4
h
¼ 3;3;3;1;1;2g
Table 6
Mo i s ound in TS o he elec ici y demand when
o ecas ing Janua y 2006.
Ma ke Mo i
OMEL M
1
l
¼ 6;4;4g
M
2
l
¼ 7;1;7g
M
3
m
¼ 5;5;4g
M
4
m
¼ 4;7;5g
M
5
h
¼ 4;4;7g
M
6
h
¼ 6;5;6g
M
7
h
¼ 7;7;2g
NYISO M
1
l
¼ 2;3;3g
M
2
m
¼ 4;3;4g
M
3
h
¼ 1;3;2g
M
4
h
¼ 3;4;1g
M
5
h
¼ 2;2;3g
ANEM M
1
m
¼ 6;5;3;4g
M
2
m
¼ 2;6;3;4g
M
3
h
¼ 2;3;1;4g
M
4
h
¼ 4;6;6;5g