A Theo e ical Demons a ion o Rein o cemen
Lea ning o PI Con ol Dynamics o Op imal
Speed Con ol o DC Mo o s by using Twin Delay
Deep De e minis ic Policy G adien Algo i hm
TUFENKCI, S.; ALAGOZ, B. B.; KAVURAN, G.; YEROGLU, C.; HERENCSÁR, N.; MAHATA, S.
Applied Ma e ials Today
Volume 213, Pa C, Ma ch 2023, 119192, Pages 1-16
ISSN: 0957-4174
DOI: h ps://doi.o g/10.1016/j.eswa.2022.119192
Accep ed manusc ip
© 2022. This manusc ip e sion is made a ailable unde he CC-BY-NC-ND 4.0 license
h p://c ea i ecommons.o g/licenses/by-nc-nd/4.0/
dspace. u b .cz
1
A Theo e ical Demons a ion o Rein o cemen Lea ning o PI
Con ol Dynamics o Op imal Speed Con ol o DC Mo o s by using
Twin Delay Deep De e minis ic Policy G adien Algo i hm
Se ilay TUFENKCI1*, Ba is Baykan ALAGOZ2, Gu kan KAVURAN3, Celaleddin YEROGLU4,
No be HERENCSAR5, Shibendu MAHATA6
1Mala ya Tu gu Ozal Uni e si y, Depa men o Compu e Technology, 44100, Mala ya, Tu key, se il[email p o ec ed]
2Inonu Uni e si y, Depa men o Compu e Enginee ing, 44100, Mala ya, Tu key, b[email p o ec ed]
3Mala ya Tu gu Ozal Uni e si y, Depa men o Elec ical-Elec onics Enginee ing, 44100, Mala ya, Tu key,
gu kan.ka u
[email protected].
4Inonu Uni e si y, Depa men o Compu e Enginee ing, 44100, Mala ya, Tu key, c.ye
[email protected].
5B no Uni e si y o Technology, Facul y o Elec ical Enginee ing and Communica ions, Dep . o Telecommunica ions, Technicka
12, 616 00 B no, Czechia, he enc[email p o ec ed]
6D . B. C. Roy Enginee ing College, Depa men o Elec ical Enginee ing, Du gapu , Wes Bengal 713206, India,
[email p o ec ed]
*Co esponding au ho : se ilay. u
[email protected].
2
Abs ac : To bene i om he ad an ages o Rein o cemen Lea ning (RL) in indus ial con ol
applica ions, RL me hods can be used o op imal uning o he classical con olle s based on he
simula ion scena ios o ope a ing condi ions. In his s udy, he Twin Delay Deep De e minis ic
(TD3) policy g adien me hod, which is an e ec i e ac o -c i ic RL s a egy, is implemen ed o
lea n op imal P opo ional In eg al (PI) con olle dynamics om a Di ec Cu en (DC) mo o
speed con ol simula ion en i onmen . Fo his pu pose, he PI con olle dynamics a e in oduced
o he ac o -ne wo k by using he PI-based obse e s a es om he con ol simula ion en i onmen .
A sui able Simulink simula ion en i onmen is adap ed o pe o m he aining p ocess o he TD3
algo i hm. The ac o -ne wo k lea ns he op imal PI con olle dynamics by using he ewa d
mechanism ha implemen s he minimiza ion o he op imal con ol objec i e unc ion. A se poin
il e is used o desc ibe he desi ed se poin esponse, and s ep dis u bance signals wi h andom
ampli ude a e inco po a ed in he simula ion en i onmen o imp o e dis u bance ejec ion con ol
skills wi h he help o expe ience based lea ning in he designed con ol simula ion en i onmen .
When he aining ask is comple ed, he op imal PI con olle coe icien s a e ob ained om he
weigh coe icien s o he ac o -ne wo k. The pe o mance o he op imal PI dynamics, which we e
lea ned by using he TD3 algo i hm and Deep De e minis ic Policy G adien algo i hm, a e
compa ed. Mo eo e , con ol pe o mance imp o emen o his RL based PI con olle uning
me hod (RL-PI) is demons a ed ela i e o pe o mances o bo h in ege and ac ional o de PI
con olle s ha we e uned by using se e al popula me aheu is ic op imiza ion algo i hms such as
Gene ic Algo i hm, Pa icle Swa m Op imiza ion, G ey Wol Op imiza ion and Di e en ial
E olu ion.
Keywo ds: Deep ein o cemen lea ning, DC mo o , PI con olle , Twin-delayed deep de e minis ic
policy g adien , me aheu is ic op imiza ion
1. In oduc ion
Rein o cemen Lea ning (RL) is an e ec i e machine lea ning me hod ha is designed o
lea ning om expe ience (Kaelbling e al., 1996; Mnih e al., 2013). In ecen yea s, i has been
u ilized o in elligen con ol o eal sys ems ( Mnih e al., 2015; Lillic ap e al., 2016; Rabaul e
al., 2019). Today, many sys ems in daily use include DC mo o s o con e elec ical ene gy o
mechanical ene gy. The e o e, pe o mance imp o emen in DC mo o con ol con ibu es o many
3
inno a i e applica ion a eas such as elec ic ehicles (Wu e al., 2004), Unmanned Ae ial Vehicles
(UAV) (Solomon e al., 2006; Solomon, 2007), e c.
DC mo o s ha e been equen ly u ilized in many applica ion a eas (Cui e al., 2012; Be ahim,
2014). Thei lowe p ice, ease o use, lexibili y, and obus ness a e he main easons o p e e ing
DC mo o s in applica ions. Today, hey a e used in nume ous a eas, such as obo s, indus ial
machine y, home equipmen , and elec ic ehicles . In daily li e applica ions,
op imal speed con ol o DC mo o s has impo ance in e ms o e iciency and com o . The main
objec i e o DC mo o speed con ol is o p oduce he desi ed speed wi hin a ce ain e e ence alue
ange in he sho es ime and o ejec en i onmen al dis u bances such as load al e a ions and
changes in ope a ing condi ions.
The mos commonly used DC mo o speed con ol echniques is based he P opo ional
In eg al De i a i e (PID) con olle amily (Sabi & Khan, 2014), which can conside he e o
signal, he change in he e o signal, and he sum o he e o signal in o de o p oduce a con ol
signal. The PID con olle amily is a widely p e e ed indus ial con ol s anda d because o hei
e ec i eness and hei simplici y in he con olle ealiza ion (Ekinci & Hekimoglu, 2019). The e is
a la ge amoun o esea ch collec ion o heo y and p ac ice o PID con ol sys ems in he li e a u e
(Ås öm & Hägglund, 1995; Visioli, 2006). Howe e , analy ical op imal uning me hods a e no
commonly conside ed en i onmen al unce ain ies and dis u bance impac s. The e o e, he
con olle uned analy ically may no exhibi op imal design pe o mance when hey a e applied o a
eal sys em. The e a e many app oaches o add ess he e iciency and pe o mance o he PID
con olle s in design asks (Ås öm & Hägglund, 1995; Visioli, 2006; Young e al., 1999; Alagoz e
al., 2015). Especially, he mul i-loop model- e e ence con ol s uc u es can imp o e he
dis u bance ejec ion pe o mance o he classical PID con olle loops (Bu le e al., 1989; Alagoz
e al., 2020); howe e , hese me hods do no gua an ee an empi ically alida ed dis u bance
ejec ion con ol pe o mance. In eal con ol applica ions, unp edic able sys em pe u ba ions and
en i onmen al dis u bances may occu . The e o e, he con ol sys em mus be designed o be obus
enough agains pa ame ic unce ain y, measu emen noise and unp edic able dis u bance in o de o
main ain he desi ed con ol pe o mance in p ac ice. A solu ion o his design issue may be he
expe ience-based op imal uning o con olle s in a ealis ic con ol simula ion en i onmen ha can
simula e he s ochas ic na u e o eal-wo ld sys ems by inco po a ing andom sys em pe u ba ion
and dis u bances. This simula ion en i onmen p o ides a use ul aining en i onmen o con olle
uning algo i hms o each he bes p ac ical pe o mance agains unce ain y, noise and
en i onmen al dis u bances. RL me hods a e inhe en ly sui able o he expe ience-based lea ning
and pe o mance imp o emen acco ding o simula ion en i onmen esul s, and his poin was a
cen al mo i a ion o he au ho s in he cu en s udy. We pa icula ly ocused on designing a
4
sui able simula ion en i onmen o e ec i e RL based con olle uning. Then, we illus a ed
implemen a ion o he TD3 algo i hm in o de o manage expe ience-based lea ning o obus
pe o mance PI con olle dynamics o he mo o con ol applica ion in he designed simula ion
en i onmen . Recen ly, T aue e al. poin ed ou he impo ance o RL en i onmen design o
in elligen mo o con ol (T aue e al., 2022). speci ically ocus on
con ol pe o mance obus ness and dis u bance ejec ion imp o emen p oblems in he design s age
o he RL aining en i onmen . To con ibu e o hei pe spec i es, we add essed he p oblem o
applica ion speci ic simula ion en i onmen design o imp o emen o obus con ol pe o mance
o PI con olle s. Simila o T aue e al., Book e al. implemen ed he Deep De e minis ic Policy
G adien (DDPG) algo i hm o RL con ol o elec ic mo o s and p esen ed expe imen al esul s
(Book e al., 2021). In ano he ecen wo k, applica ion o he deep Q-ne wo ks (DQN) algo i hm
was demons a ed o he speed con ol o pe manen magne synch onous mo o s and compa ed
wi h pe o mance o classical PI con olle s (Song e al., 2021). Howe e , hese wo ks employed
he ained RL ne wo k as a di ec RL con olle ins ead o uning a s anda d con olle . Di ec RL
con ol may lead o p ac ical ealiza ion di icul ies such as he compu a ional complexi y o RL
aining and he s abili y conce ns.
Table 1 lis s some inhe en p ope ies o he me aheu is ic op imiza ion and RL algo i hms
ha make hem ad an ageous so compu a ion ools o con ol applica ions. The me aheu is ic
me hods can be ad an ageous o sol e well-de ined op imal con ol p oblems because o hei
e ec i e global sea ch capabili ies. Howe e , he expe ience based lea ning acco ding o he agen -
en i onmen in e ac ion is a c i ical p ope y o p e e he RL me hod o con olle uning
acco ding o well-designed simula ion en i onmen s ha can mee applica ion speci ic con ol
equi emen s.
Table 1. Some use ul p ope ies o me aheu is ic op imiza ion and RL algo i hms o con ol applica ions
P ope ies Me aheu is ic
op imiza ion based
uning me hods
RL based uning
me hods
Suppo ing popula ion based global sea ching High No
Suppo ing local sea ching High High
Suppo ing o -line con olle uning High High
Suppo ing online con olle uning Low High
Sui abili y o algo i hms o he expe ience
based lea ning om agen -en i onmen
in e ac ion
Low High
Sui abili y o algo i hms o wo k as con olle Low High
The RL has also been used o implemen a neu al con olle (Koch e al., 2019) and o uning
classical con olle s (Chen e al., 2022; Esmaeili e al., 2017) in con ol applica ions. The uzzy Q-
lea ning me hod was sugges ed o uning he PID con olle based on RL and applied o
empe a u e con ol (Chen e al., 2022). Esmaeili e al. p oposed an imp o ed RL-based uzzy-PID
5
con olle o load equency con ol o an island mic og id (Esmaeili e al., 2017). The main
conce n may be he ques ion whe he he RL algo i hms can lea n om whole sys em dynamics and
con ol hem. A sui able implemen a ion o he RL algo i hm can manage con ol o sys em
dynamics based on he aining expe iences. In DC mo o speed con ol, he PI con olle om he
PID con ol amily has been widely p e e ed (Naga ajan e al., 2016; Kanojiya e al., 2012;
Sunda eswa an & Vasu, 2000). A PI con olle can be sui able o DC mo o speed con ol because
de i a i e elemen s o PID may cause e y as changes and cha e in he con ol signal and may
ampli y noise signals in he con ol loop o DC mo o s. The e o e, o a smoo he and mo e
s aigh o wa d con ol solu ion, PI con ol can be p e e ed o speed con ol o DC mo o s.
In his s udy, a RL app oach is implemen ed o lea n he PI con ol dynamics ha can p o ide
a desi ed DC mo o con ol pe o mance wi hin he con ol simula ion en i onmen . The RL
me hod needs a ealis ic simula ion en i onmen o lea n om agen expe iences. Majo p ac ical
con ol complica ions such as unp edic able ex e nal dis u bance and measu emen noise can be
in oduced in his simula ion en i onmen and con ibu e o lea ning an e ec i e and op imal PI
con ol dynamic om expe ience in he simula ion en i onmen . Acco ding o RL policies, he
agen ac ing in his en i onmen ies o ind he mos app op ia e ac ion ha maximizes he ewa d
(Kaelbling e al., 1996). The e o e, he op imal PI con ol dynamics ha deal wi h he complica ion
can be lea ned om he con ol simula ion en i onmen in o de o pe o m he bes PI ac ion by
using a sui ably designed ewa d and obse e mechanism.
His o ically, he RL s a egy, which was ounded by Richa d Bellman in he 1950s, i s
gained success in playing a backgammon game. As a esul , combining i wi h an a i icial neu al
ne wo k o e ime achie ed supe io success and pe o med be e han humans in sol ing complex
p oblems (Mnih e al., 2015; G aepel, 2016). La e , Q-Lea ning was i s o come ou in 1989,
which is a special ype o RL app oach. Wa kins used he le e Q o he alue unc ion, which is
based on he heo y o Ma ko decision p ocesses (Wa kins, 1989; Wa kins & Dayan, 1992).
Howe e , Q-Lea ning has no a ac ed in e es in i s domains un il DQN algo i hms we e
de eloped (Mnih e al., 2013). A e wa ds, he De e minis ic Policy G adien (DPG) algo i hm,
which uses he ac o -c i ic s uc u e, was de eloped (Lillic ap e al., 2016; Sil e e al., 2014).
Finally, he TD3 Policy G adien algo i hm was p oposed, and he TD3 algo i hm can wo k mo e
e ec i ely han many RL me hods in applica ions (Fujimo o e al., 2018).
RL-based me hods a e widely used in he ield o con ol. I can enhance in elligence in
con ol sys ems (Lillic ap e al., 2016; Rabaul e al., 2019; Chen e al., 2018; B andi e al., 2020;
Sa heeshbabu e al., 2019). In he cu en s udy, he TD3 algo i hm is implemen ed o lea ning
op imal PI con ol dynamics om a closed-loop DC mo o simula ion en i onmen . A e e ence
model in oduced by a se poin il e is used in his simula ion en i onmen o de ine desi ed s ep
6
esponses wi hin a ange o s ep inpu ampli udes o he TD3 algo i hm. The addi i e inpu
dis u bance model wi h he andom ampli ude s ep dis u bance was implemen ed in he simula ion
en i onmen o ain he RL-based PI con ol dynamics o imp o ed dis u bance ejec ion. The
ac o -ne wo k was designed o yield he PI coe icien s o he lea ned op imal PI dynamics in his
en i onmen . Thus, an RL uned PI con olle was ob ained o he DC mo o speed con ol, and i s
pe o mance was es ed and compa ed wi h he esul s o he me aheu is ic op imiza ion based
op imal PI and ac ional o de PI (FOPI) uning algo i hm. The main con ibu ions o his s udy
can be summa ized as ollows:
(i) This s udy demons a es he design o con ol simula ion en i onmen s in o de o lea n
op imal and dis u bance ejec PI dynamics om a se -poin e e ence model by using he
TD3 algo i hm. Thus, he e e ence model esponse based op imal pe o mance uning o he
PI con olle acco ding o he ealis ic en i onmen al condi ions (e.g., andom dis u bance
inse ion, andom change o inpu signal in a wide ange) was pe o med by using he
ein o ced lea ning s a egy. In his lea ning pe spec i e, he e e ence model desc ibes a
desi ed con ol sys em esponse (o e shoo , se ling ime e c.) and imp o es RL based
con olle uning p ocess acco ding o con ol applica ion equi emen s and speci ica ions.
(ii) To in es iga e pe o mance imp o emen o he RL based con olle uning me hod, con olle
uning pe o mance o he TD3 algo i hm was compa ed wi h pe o mances o s a e-o -a
con olle uning me hods in he cu en s udy. We illus a e he pe o mance imp o emen o
RL-based PI uning in o e shoo - ee, smoo h con ol o DC mo o speed compa ed o se e al
me aheu is ic PI and FOPI con olle uning me hods in he e e ence model based con olle
uning. Also, pe o mance imp o emen o RL-PI con olle wi h he TD3 algo i hm is
demons a ed o e ha o he DDPG algo i hm in his DC mo o con ol p oblem.
In he es o he pape , Sec ion 2 p o ides p elimina y knowledge o RL and b ie ly
in oduces TD3 policy g adien me hod algo i hm. A e p esen ing he speed con ol model o DC
mo o s, which is used in he simula ion en i onmen , a Simulink simula ion en i onmen o
aining o he TD3 algo i hm o lea n an op imal and obus PI con olle dynamics o DC mo o
speed con ol is explained in Sec ion 3. Sec ion 4 p esen s a simula ion s udy ha compa es esul s
o RL based PI con olle design and esul s o PI and FOPI con olle designs o o he popula
me hods. Some conclusions a e summa ized in Sec ion 5.
7
2. P elimina ies and P oblem S a emen
2.1. P elimina ies o Rein o cemen Lea ning
The RL is a machine lea ning echnique ha lea ns om he di ec in e ac ion o asse s wi h
he en i onmen (Su on & Ba o, 1998). The agen lea ns he en i onmen and ules h ough
expe ience o achie e goals in he en i onmen . The ewa d and punishmen mechanisms a e used
in he aining p ocess, and hese mechanisms guide he agen 's lea ning (Hoshino & Kamei, 2003;
Russell & No ig, 2003).
The main di e ence om o he lea ning me hods is ha he sys em only lea ns by ial and
e o in i s own en i onmen wi hou using any p ede ined ins uc ion se s (Su on & Ba o, 1998;
Na end a & Tha hacha , 2012). The ewa d mo i a es he agen o choose he bes ac ion in he
en i onmen . In o de o ind he bes way o sol e he p oblem, he sys em's lea ning p ocess
depends on a su icien numbe o ials (Na end a & Tha hacha , 2012). The agen aims o ob ain
an op imal policy by maximizing he o al ewa d. This lea ning me hod can be assumed as a
Ma ko ian p ocess (Bellman, 1957), and i s heo e ical ounda ion assumes RL o be a s ochas ic
lea ning p ocess ha can p og essi ely enhance se and ial esul s by sui ably ewa ding agen
ac ions in he en i onmen .
A each disc e e ime inc emen , he agen selec s an ac ion acco ding o i s policy
depending on he s a e . Then, he agen ecei es a ewa d , and a new s a e
is ecei ed om he en i onmen o he ac ion (Fujimo o e al., 2018). In he Q-Lea ning,
he quali y o a s a e-ac ion couple is e alua ed acco ding o he unc ion , and i exp esses
expec ed ewa ds om an ac ion a he s a e . The is calcula ed by using he Bellman
equa ion. The o m o he Bellman equa ion is adap ed as
. (1)
This equa ion implies ha he maximum u u e ewa d can be possible wi h he ewa d o he
cu en ac ion ( ) plus he maximum o u u e ewa d es ima ions . To sol e he
equa ion i e a i ely, he Bellman e o l
B
ea i e a ion is de ined as
. (2)
Then, an i e a i e scheme o each op imal s a e-ac ion alue was gi en by using
empo al di e ence lea ning (Su on, 1988) as
, (3)
8
whe e, is a discoun ac o be ween ze o and one, de e mining he weigh s o sho - e m
ewa ds (Luu, 2015). The pa ame e is he lea ning a e. The unc ion is he ewa d
o ac ion
a
and obse e s a e . This o mula ion de ines he RL p ocess as a Ma ko p ocess. In
o he wo ds, u u e si ua ions depend on only p esen si ua ions in his p ocess. The e o e,
condi ional p obabili y is widely used o model RL p ocesses.
Ma ko p ocess o a s a e can be desc ibed as
. (4)
I he cu en s a e o a sys em is conside ed, p e ious ac ion om a he s a e om
leads o a new s a e om and a ewa d om . The p obabili y densi y unc ion o he andom
a iables and can be exp essed o he s ochas ic modeling o he RL as
, (5)
whe e , ' ,s s S R and . I can be seen ha Equa ion (1) depends on he s a e o he ewa d
a he ime and he maximum o s a e-ac ion quali y. The e o e, he possibili ies o ansi ion o
is w i en by conside ing he ma ginal dis ibu ion o he s a e (Mo ales & Za agoza, 2011).
(6)
Acco dingly, he expec ed ewa d can be w i en by using a ma ginal dis ibu ion o s a es as
(Mo ales & Za agoza, 2011)
. (7)
To implemen his scheme, app oxima ion unc ions such as neu al ne wo ks a e used o lea n
ac ions ha maximize ewa ds om he en i onmen .
P og ess in RL me hods con inues and applica ion ields o RL ha e expanded in many ields
such as a i ude con ol o hype sonic een y ehicles (Y. Liu e al., 2022), zone scheduling
op imiza ion o pumps in wa e dis ibu ion ne wo ks (Xu e al., 2021), lexible con ol o disc e e
e en sys ems using en i onmen simula ion (Zielinski e al., 2021) and au onomous na iga ion o
UAV in mul i-obs acle en i onmen s (Zhang e al., 2022). Ka u an con ibu ed o his discussion
by in es iga ing he e ec o deep RL on he ac ional-o de oscilla o (Ka u an, 2022).
2.2. A B ie In oduc ion o Twin Delayed Deep De e minis ic Policy G adien Algo i hm
The TD3 policy g adien algo i hm, which is an imp o emen o he DDPG algo i hm, akes
in o accoun he unc ion app oach e o (Lillic ap e al., 2016; Fujimo o e al., 2018). The objec i e
o he RL is o ind he op imal policy ha maximizes he expec ed ewa d by uning he
15
wi h one neu on o he ou pu . The o he ne wo ks a e he ac o neu al ne wo ks (ac o -
ne wo k) ha yield he ac ion ( ) by using obse e s a es (
s
). Based on alues, he TD3
algo i hm ains c i ic-ne wo ks and ac o -ne wo ks o maximize he ewa d as an op imal con ol
objec i e gi en by Equa ion (21) in his s udy.
Ac ion Inpu S a e Inpu
Fully Connec ed Laye Fully Connec ed Laye
Conca ena ion Laye
ReLU Laye
Fully Connec ed Laye
ReLU Laye
Fully Connec ed Laye
S a e Pa hAc ion Pa h
Figu e 5. The a chi ec u e o he c i ic neu al ne wo k ha is implemen ed in his applica ion
The s a e obse e moni o s impo an s a es o he con ol sys em om he simula ion
en i onmen . To ma ch he weigh coe icien o he ac o -ne wo k wi h he PI con olle
coe icien s in he Ma lab RL con ol example, he in eg al o he e o signal is obse ed o
implemen he in eg al s a e ( ), and he e o signal i sel is obse ed o implemen p opo ional
s a e ( ) o PI dynamics. The s a e obse e was implemen ed by moni o ing hese wo s a es o PI
dynamics as
. (22)
Ac o ne wo k is designed by using a wo-inpu ully connec ed laye wi h a linea ac i a ion
unc ion. Two obse e s a es eed he ac o ne wo k. Acco dingly, he ac ion o he ac o can be
easily w i en by using he obse e s a es as ollows
. (23)
A e he aining o he ac o ne wo k is comple ed, he weigh s ( , ) o he neu al
ne wo k exp esses he lea ned PI con olle coe icien s. He e, he weigh co esponds o he
16
coe icien , and he weigh co esponds o he coe icien o he disc e e PI con olle
implemen a ion. To demon ollow:
(24)
whe e exp esses he disc e e- ime PI con olle unc ion, whe e he
p opo ional coe icien o he PI con olle is , and he in eg al coe icien o he PI con olle
is .
4. Simula ion S udy
In o de o ob ain he desi ed speed con ol pe o mance o he DC mo o by using he TD3
policy g adien algo i hm, he aining o he RL sys em was ca ied ou by using he Simulink
simula ion en i onmen shown in Figu e 3. Con ol simula ions we e pe o med o 20 seconds o
each episode, and an e o signal om each simula ion was used o ob ain obse e s a es (Equa ion
22). E o and con ol signals om he simula ion en i onmen a e used o calcula e he ewa d
(Equa ion 19) in he aining o he TD3 algo i hm. Table 4 lis s some hype pa ame e s used o he
con igu a ion o he aining p ocess o TD3 agen s. The ial and e o me hod was used o se
sui able alues o hese pa ame e s.
Table 4. Va iables o aining p ocess o he TD3 algo i hm
Maximum Numbe o Episodes (T) 100
Mini Ba ch Size 128
Sampling ime 0.1 s
Lea ning Ra e 0.001
Du ing he aining o he RL sys em, e e ence s ep inpu signals wi h andom ampli udes
(wi hin he ange o 0-155) we e applied o he inpu o he sys em in o de o imp o e se poin
con ol pe o mance, and he s ep dis u bance signals wi h andom ampli ude (wi hin he ange o [-
5,5]) we e used o ain he sys em o ob ain an imp o ed dis u bance ejec ion pe o mance. To
gi e a ewa d acco ding o he op imal con ol o mula ion (Equa ion (19)), pa ame e s and
a e se o 0.9 and 0.01, espec i ely.
The RL aining in he simula ion en i onmen (Figu e 3) allows he ac o neu al ne wo k o
espond simila o a PI con olle by adjus ing he ac o -ne wo k weigh s ( , ). The ac o -
ne wo k is he ully connec ed single neu on wi h wo inpu s om obse ed s a es. Equa ions (23)
and (24) p o e ha a combina ion o ac o -ne wo k and obse e block implemen s a PI con olle .
The ou pu o he ac o -ne wo k is he con ol signal o he DC mo o . Figu e 6 shows he change
17
o ins an alues and a e age alues o ewa ds ob ained du ing he aining p ocess in he
simula ion en i onmen . The a e age ewa d is he a e age o he ins an ewa ds and exp esses
o e all pe o mance du ing he aining p ocess.
Figu e 6. Change o ins an alues and a e age alues o ewa ds in he simula ion en i onmen
When he aining was comple ed, he p
w was ob ained 1.6626, which is he equi alen
alue o a disc e e- ime PI con olle , and he was ob ained 2.0830, which is he equi alen o
a disc e e- ime PI con olle . Acco ding o hese aining esul s, he lea ned disc e e PI con olle
dynamics a e equal o
1
0830.26626.1)( z
T
zC s. (25)
In his s udy, he RL-PI con ol was ewa ded o lea n he s ep esponse o he se poin il e
1
1
)( s
sFse , and i enables smoo hly se ling o elec ic mo o speed o he e e ence speed 100
pm wi hou any o e shoo and ipples. This is an impo an p ope y o elec ical mo o con ol
because he o e shoo s and ipples can cause edundan accele a ion and luc ua ions in he DC
mo o 's speed esponses. These e ec s esul in discom o and unnecessa y ene gy consump ion in
con ol applica ions, whe e DC mo o s a e used, o example applica ions in elec ical ehicles,
UAVs, e c.
To analyze pe o mances o he designed PI con olle s, a con ol sys em es simula ion o
100 seconds du a ion was ca ied ou , and an addi i e inpu dis u bance wi h a s ep wa e o m ( he
ampli ude o s ep dis u bance is 20) is applied a he simula ion ime o 50 seconds. The
pe o mance o he RL-PI con ol is compa ed wi h he esponses o he PI and FOPI con olle s
ha we e uned by using se e al me aheu is ics such as Gene ic Algo i hm (GA) (Holland, 1973),
A e age Rewa d
*
Ins an Rewa d
o
18
Pa icle Swa m Op imiza ion (PSO) (Kennedy & Ebe ha ,1995), G ey Wol Op imiza ion (GWO)
(Mi jalili e al., 2014) and Di e en ial E olu ion (DE) (S o n & P ice,1997). Simila o he RL
algo i hms, me aheu is ic op imiza ion me hod can be e ec i ely used o seeking op imal solu ions
in decision making and scheduling p oblems such as op imal cha ging scheduling p oblem o
elec ic ehicles (Liu e al., 2020), ba ch-p ocessing machine scheduling p oblem (Zhou e al., 2021
), ene gy-e icien dis ibu ed no-idle low-shop scheduling p oblem (Zhao e al., 2020; Zhao e al.,
2021a; Zhao e al., 2021b). Also, success ul use o se e al me aheu is ic me hods was shown in
con olle uning p oblems (Tu enkci e al., 2020; Zheng & Pi, 2016; Koma hi & Umamaheswa i,
2020; Liu & Hsu, 2010). Besides, hype heu is ic app oaches ha in ol e high-le el heu is ics and
low-le el heu is ics, can p o ide imp o ed pe o mance in a wide applica ion domain owing o
high-le el ac ics (Zhao e al., 2021c). The me aheu is ic me hods op imized he PI con olle
coe icien s o yield a s ep esponse ha esembles he esponse o he se poin il e . This
objec i e is he same as in he RL-PI con ol sys em, whe e he il e is a supe iso (a
e e ence model) ha was used o shape he s ep esponse o he con ol sys em acco ding o
esponses o he e e ence model , which p o ides smoo h and o e shoo - ee con ol o he
DC mo o . Me aheu is ic me hods ha e been widely used o op imal uning o he con olle
unc ions. The e o e, we demons a ed he obus con ol pe o mance imp o emen s o he RL-PI
con olle wi h TD3 algo i hm compa ed o esul s o GA, PSO, GWO and DE algo i hms. The
sol ed op imiza ion p oblem o op imal uning o PI con olle unc ions ( ) and
FOPI con olle ( ) ia me aheu is ic me hods we e shown in he Appendix. Table
5 shows PI and FOPI con olle uning esul s o GA, PSO, GWO and DE algo i hms. Since he
objec i e o e e ence model based con olle uning aims o app oxima e o he esponse o he
e e ence model 1
1
)( s
sFse , he op imal con olle coe icien s we e ob ained o close each
o he . Figu es 7 and 8 show s ep esponses o hose con olle s in compa ison wi h he esponse o
he RL-PI con olle wi h he TD3 algo i hm. Figu es 9 and 10 show e o signals om hese
con olle simula ions. Figu es 11 and 12 show he change o con ol signals du ing hese con ol
simula ions. Main ad an ages o he RL-PI con olle uning comes om he design o a sui able
simula ion en i onmen ha in ol es andom ampli ude dis u bance inse ion and andom se poin s.
The expe ience based lea ning o he PI con olle dynamics can explo e mo e obus esponses o
he inpu dis u bances in he RL based uning. The me aheu is ic uning me hods only sol ed he
op imiza ion p oblem ha is de ined in he Appendix. The e o e, me aheu is ic uning me hods,
which a e based on analy ical solu ions o an op imal con ol p oblem, canno di ec ly conside
e ec s o andom ampli ude dis u bance signal o he andom ampli ude e e ence signal on con ol
19
pe o mance (Ozbey e al., 2020). This sho coming can limi he uning pe o mance o
me aheu is ic op imiza ion me hods in applica ion speci ic design.
Table 5. Op imal PI and FOPI con olle unc ions ( )(sC ) ha we e uned by me aheu is ic algo i hms
Algo i hms Con olle s Algo i hms Con olle s
GA-PI GWO-PI
GA-FOPI GWO-FOPI
PSO-PI DE-PI
PSO-FOPI DE-FOPI
Figu e 7. S ep esponses ob ained by GA-PI, PSO-PI, GWO-PI, DE-PI con olle s and RL-PI
con olle wi h TD3 algo i hm
Figu e 8. S ep esponses ob ained by GA-FOPI, PSO-FOPI, GWO-FOPI, DE-FOPI con olle s and
RL-PI con olle wi h TD3 algo i hm
20
Figu e 9. E o signals o GA-PI, PSO-PI, GWO-PI, DE-PI and RL-PI con olle s
Figu e 10. E o signals o GA-FOPI, PSO-FOPI, GWO-FOPI, DE-FOPI and RL-PI con olle s
21
Figu e 11. Con ol signals o GA-PI, PSO-PI, GWO-PI, DE-PI and RL-PI con olle s
Figu e 12. Con ol signals o GA-FOPI, PSO-FOPI, GWO-FOPI, DE-FOPI and RL-PI con olle s
We also compa ed pe o mances o wo di e en RL s a egies in he PI con olle uning
p oblem. When he aining o DDPG algo i hm was comple ed, he disc e e- ime con olle
coe icien s we e calcula ed as and o a disc e e- ime PI con olle .
Acco ding o hese aining esul s, he DDPG lea ned disc e e PI con olle dynamics a e equal o
22
. (26)
Figu e 13 shows he change o ins an alues and a e age alues o ewa ds du ing he
aining wi h he DDPG algo i hm in he same simula ion en i onmen . Figu e 14 shows s ep
esponses o he RL-PI con olle s ha we e lea ned by he TD3 algo i hm and he DDPG
algo i hm. The DDPG algo i hm has been used as RL s a egy in he mo o con ol applica ion
(T aue e al., 2022; Book e al., 2021). The same simula ion en i onmen and he same neu al
ne wo k a chi ec u es we e used o hese algo i hms. Also, pa ame e se ings o he DDPG
algo i hm we e con igu ed o be he same wi h he TD3 algo i hm in he aining. Figu e 15 and 16
show compa isons o e o signals and con ol signals. Simula ion esul s showed ha he RL-PI
con olle uned wi h he TD3 algo i hm p o ided be e PI con olle pe o mance compa ed o he
RL-PI con olle uned wi h he DDPG algo i hm. In gene al, he TD3 algo i hm can be mo e
consis en in con e gence o maximum ewa d alues han he DDPG algo i hm, mainly because o
using wo c i ic ne wo k and ac o ne wo ks wi h a delayed upda e and ac ion noise egula iza ion
(clipped noise addi ion as in he equa ion (10) o a clipped double Q-lea ning e ec (Fujimo o e
al., 2018)), which make TD3 mo e ad an ageous o escape om he o e i ing o na ow ewa d
peaks and imp o es he sea ch skill o he TD3 algo i hm. The di e ence in ewa ds cha ac e is ics
in Figu e 6 ( o he TD3 algo i hm) and Figu e 13 ( o he DDPG algo i hm) con i ms
imp o emen s o he TD3 algo i hm o e he DDPG algo i hm in his con ol p oblem.
Figu e 13. Change o ins an alues and a e age alues o ewa ds in he simula ion en i onmen
A e age Rewa d
*
Ins an Rewa d
o
23
Figu e 14. S ep esponses ob ained by he RL-PI con olle wi h TD3 algo i hm and he RL-PI
con olle wi h DDPG algo i hm
Figu e 15. E o signals o he RL-PI con olle wi h TD3 algo i hm and he RL-PI con olle wi h
DDPG algo i hm
24
Figu e 16. Con ol signals o he RL-PI con olle wi h TD3 algo i hm and he RL-PI con olle wi h
DDPG algo i hm
Table 6 lis s he Mean Squa e E o ( ,whe e he is numbe o sampled
da a om simula ion en i onmen ) and A e age o Absolu e Con ol signal ( )
pe o mances o he conside ed algo i hms o he PI con olle uning e o s. We pe o med 5
imes independen ial uns o each algo i hm. In hese es s, o ob ain di e en esul s, he RL-PI
algo i hms we e ini ia ed a 5 di e en ini ial poin s o he PI con olle coe icien s. He e, he MSE
indica es he dissimila i y be ween he esponse o he DC mo o con ol sys em and esponses o
he e e ence model . The AAC exp esses he a e age le el o con ol signal ampli ude ha
is an impo an indica o ela ed o he ene gy e iciency o he con ol ac ions.
Table 6. MSE and AAC pe o mances o he compa ed PI con olle uning me hods
Con olle s
MSE AAC
Min A e age Max S anda d
De ia ion
A e age
GA-PI 26.28310 26.28360 26.28390 3.16876 10-4 111.8900
PSO-PI 26.25929 26.28388 26.31272 1.91939 10-2 111.8901
GWO-PI 26.27081 26.70522 28.38485 9.39030 10-1 111.8901
DE-PI 26.28387 26.28389 26.28390 1.17834 10-5 111.8900
RL-PI (DDPG) 6.011537 111597.0 557907.5 2.49495 105 315.9496
RL-PI (TD3) 9.053762 14.91569 20.75544 5.66412 111.8495
31
h ps://doi.o g/10.1007/S00521-020-05352-1/FIGURES/9
S o n, R., & P ice, K. (1997). Di e en ial e olu ion a simple and e icien heu is ic o global op imiza ion o e
con inuous spaces. Jou nal o global op imiza ion, 11(4), 341-359. h ps://doi.o g/10.1023/A:1008202821328
Sunda eswa an, K., & Vasu, M. (2000). Gene ic uning o PI con olle o speed con ol o DC mo o d i e.
P oceedings o he IEEE In e na ional Con e ence on Indus ial Technology, 1, 521 525.
h ps://doi.o g/10.1109/ici .2000.854212
Su on, R. S. (1988). Lea ning o p edic by he me hods o empo al di e ences. Machine Lea ning, 3(1), 9 44.
h ps://doi.o g/10.1007/b 00115009
Su on, R. S., & Ba o, A. G. (1998). Rein o cemen Lea ning: An In oduc ion. MIT p ess, Camb ige, MA.
T aue, A., Book, G., Ki chgassne , W., & Wallscheid, O. (2022). Towa d a Rein o cemen Lea ning En i onmen
Toolbox o In elligen Elec ic Mo o Con ol. IEEE T ansac ions on Neu al Ne wo ks and Lea ning Sys ems,
33(3), 919 928. h ps://doi.o g/10.1109/TNNLS.2020.3029573
-
domain. In Jou nal o Ad anced Resea ch, 25, 171 180. h ps://doi.o g/10.1016/j.ja e.2020.03.002
Uni e si y o Michigan. (2017). Con ol Tu o ials o MATLAB and Simulink - Mo o Speed: Sys em Modeling.
h ps://c ms.engin.umich.edu/CTMS/index.php?example=Mo o Speed§ion=Sys emModeling
Visioli, A. (2006). P ac ical PID Con ol. In P ac ical PID Con ol. h ps://doi.o g/10.1007/1-84628-586-0
Wa e Tank Rein o cemen Lea ning En i onmen Model - MATLAB & Simulink - Ma hWo ks Swi ze land. (n.d.).
Re ie ed Ma ch 24, 2022, om h ps://ch.ma hwo ks.com/help/ ein o cemen -lea ning/ug/wa e - ank-
ein o cemen -lea ning-en i onmen -model.h ml
Wa kins, C. J. C. H. (1989). Lea ning om delayed ewa ds. In Robo ics and Au onomous Sys ems, 15(4), 233 235.
Wa kins, C. J. C. H., & Dayan, P. (1992). Q-lea ning. Machine Lea ning, 8(3 4), 279 292.
h ps://doi.o g/10.1007/b 00992698
Wu, H. X., Cheng, S. K., & Cui, S. M. (2004). A con olle o b ushless DC Mo o o elec ic ehicle. 2004 12 h
Symposium on Elec omagne ic Launch Technology, 528 533.
Xu, J., Wang, H., Rao, J., & Wang, J. (2021). Zone scheduling op imiza ion o pumps in wa e dis ibu ion ne wo ks
wi h deep ein o cemen lea ning and knowledge-assis ed lea ning. So Compu ing, 25(23), 14757 14767.
IEEE
T ansac ions on Con ol Sys ems Technology, 7(3), 328 342. h ps://doi.o g/10.1109/87.761053
Zhang, S., Li, Y., & Dong, Q. (2022). Au onomous na iga ion o UAV in mul i-obs acle en i onmen s based on a Deep
Rein o cemen Lea ning app oach. Applied So Compu ing, 115, 108194.
Zhao, F., He, X., & Wang, L. (2020). A wo-s age coope a i e e olu iona y algo i hm wi h p oblem-speci ic knowledge
o ene gy-e icien scheduling o no-wai low-shop p oblem. IEEE T ansac ions on Cybe ne ics, 51(11), 5291
5303. h ps://doi.o g/10.1109/TCYB.2020.3025662
Zhou, S., Xing, L., Zheng, X., Du, N., Wang, L., & Zhang, Q. (2021). A Sel -Adap i e Di e en ial E olu ion
Algo i hm o Scheduling a Single Ba ch-P ocessing Machine wi h A bi a y Job Sizes and Release Times. IEEE
T ansac ions on Cybe ne ics, 51(3), 1430 1442. h ps://doi.o g/10.1109/TCYB.2019.2939219
Zhao, F., Ma, R., & Wang, L. (2021a). A Sel -Lea ning Disc e e Jaya Algo i hm o Mul iobjec i e Ene gy-E icien
Dis ibu ed No-Idle Flow-Shop Scheduling P oblem in He e ogeneous Fac o y Sys em. IEEE T ansac ions on
Cybe ne ics. h ps://doi.o g/10.1109/TCYB.2021.3086181
Zhao, F., Zhang, L., Cao, J., & Tang, J. (2021b). A coope a i e wa e wa e op imiza ion algo i hm wi h ein o cemen
lea ning o he dis ibu ed assembly no-idle lowshop scheduling p oblem. Compu e s and Indus ial
Enginee ing, 153. h ps://doi.o g/10.1016/j.cie.2020.107082
32
Zhao, F., Di, S., Cao, J., Tang, J., & Jon inaldi. (2021c). A No el Coope a i e Mul i-S age Hype -Heu is ic o
Combina ion Op imiza ion P oblems. Complex Sys em Modeling and Simula ion, 1(2), 91 108.
Zheng, W., & Pi, Y. (2016). S udy o he ac ional o de p opo ional in eg al con olle o he pe manen magne
synch onous mo o based on he di e en ial e olu ion algo i hm. ISA T ansac ions, 63, 387 393.
h ps://doi.o g/10.1016/j.isa a.2015.11.029
Zielinski, K. M. C., Hendges, L. V., Flo indo, J. B., Lopes, Y. K., Ribei o, R., Teixei a, M., & Casano a, D. (2021).
Flexible con ol o Disc e e E en Sys ems using en i onmen simula ion and Rein o cemen Lea ning. Applied
So Compu ing, 111, 107714. h ps://doi.o g/10.1016/J.ASOC.2021.107714