scieee Open visual document viewer

A theoretical demonstration for reinforcement learning of PI control dynamics for optimal speed control of DC motors by using Twin Delay Deep Deterministic Policy Gradient Algorithm

Tufenkci, Sevilay; Alagoz, Baris Baykant; Kavuran, Gürkan; Yeroglu, Celaleddin; Herencsár, Norbert; Mahata, Shibendu

Abstract

To benefit from the advantages of Reinforcement Learning (RL) in industrial control applications, RL methods can be used for optimal tuning of the classical controllers based on the simulation scenarios of operating con-ditions. In this study, the Twin Delay Deep Deterministic (TD3) policy gradient method, which is an effective actor-critic RL strategy, is implemented to learn optimal Proportional Integral (PI) controller dynamics from a Direct Current (DC) motor speed control simulation environment. For this purpose, the PI controller dynamics are introduced to the actor-network by using the PI-based observer states from the control simulation envi-ronment. A suitable Simulink simulation environment is adapted to perform the training process of the TD3 algorithm. The actor-network learns the optimal PI controller dynamics by using the reward mechanism that implements the minimization of the optimal control objective function. A setpoint filter is used to describe the desired setpoint response, and step disturbance signals with random amplitude are incorporated in the simu-lation environment to improve disturbance rejection control skills with the help of experience based learning in the designed control simulation environment. When the training task is completed, the optimal PI controller coefficients are obtained from the weight coefficients of the actor-network. The performance of the optimal PI dynamics, which were learned by using the TD3 algorithm and Deep Deterministic Policy Gradient algorithm, are compared. Moreover, control performance improvement of this RL based PI controller tuning method (RL-PI) is demonstrated relative to performances of both integer and fractional order PI controllers that were tuned by using several popular metaheuristic optimization algorithms such as Genetic Algorithm, Particle Swarm Opti-mization, Grey Wolf Optimization and Differential Evolution.

Full text

A Theo e ical Demons a ion o Rein o cemen Lea ning o PI Con ol Dynamics o Op imal Speed Con ol o DC Mo o s by using Twin Delay Deep De e minis ic Policy G adien Algo i hm TUFENKCI, S.; ALAGOZ, B. B.; KAVURAN, G.; YEROGLU, C.; HERENCSÁR, N.; MAHATA, S. Applied Ma e ials Today Volume 213, Pa C, Ma ch 2023, 119192, Pages 1-16 ISSN: 0957-4174 DOI: h ps://doi.o g/10.1016/j.eswa.2022.119192 Accep ed manusc ip © 2022. This manusc ip e sion is made a ailable unde he CC-BY-NC-ND 4.0 license h p://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ dspace. u b .cz 1 A Theo e ical Demons a ion o Rein o cemen Lea ning o PI Con ol Dynamics o Op imal Speed Con ol o DC Mo o s by using Twin Delay Deep De e minis ic Policy G adien Algo i hm Se ilay TUFENKCI1*, Ba is Baykan ALAGOZ2, Gu kan KAVURAN3, Celaleddin YEROGLU4, No be HERENCSAR5, Shibendu MAHATA6 1Mala ya Tu gu Ozal Uni e si y, Depa men o Compu e Technology, 44100, Mala ya, Tu key, se il[email p o ec ed] 2Inonu Uni e si y, Depa men o Compu e Enginee ing, 44100, Mala ya, Tu key, b[email p o ec ed] 3Mala ya Tu gu Ozal Uni e si y, Depa men o Elec ical-Elec onics Enginee ing, 44100, Mala ya, Tu key, gu kan.ka u [email protected]. 4Inonu Uni e si y, Depa men o Compu e Enginee ing, 44100, Mala ya, Tu key, c.ye [email protected]. 5B no Uni e si y o Technology, Facul y o Elec ical Enginee ing and Communica ions, Dep . o Telecommunica ions, Technicka 12, 616 00 B no, Czechia, he enc[email p o ec ed] 6D . B. C. Roy Enginee ing College, Depa men o Elec ical Enginee ing, Du gapu , Wes Bengal 713206, India, [email p o ec ed] *Co esponding au ho : se ilay. u [email protected]. 2 Abs ac : To bene i om he ad an ages o Rein o cemen Lea ning (RL) in indus ial con ol applica ions, RL me hods can be used o op imal uning o he classical con olle s based on he simula ion scena ios o ope a ing condi ions. In his s udy, he Twin Delay Deep De e minis ic (TD3) policy g adien me hod, which is an e ec i e ac o -c i ic RL s a egy, is implemen ed o lea n op imal P opo ional In eg al (PI) con olle dynamics om a Di ec Cu en (DC) mo o speed con ol simula ion en i onmen . Fo his pu pose, he PI con olle dynamics a e in oduced o he ac o -ne wo k by using he PI-based obse e s a es om he con ol simula ion en i onmen . A sui able Simulink simula ion en i onmen is adap ed o pe o m he aining p ocess o he TD3 algo i hm. The ac o -ne wo k lea ns he op imal PI con olle dynamics by using he ewa d mechanism ha implemen s he minimiza ion o he op imal con ol objec i e unc ion. A se poin il e is used o desc ibe he desi ed se poin esponse, and s ep dis u bance signals wi h andom ampli ude a e inco po a ed in he simula ion en i onmen o imp o e dis u bance ejec ion con ol skills wi h he help o expe ience based lea ning in he designed con ol simula ion en i onmen . When he aining ask is comple ed, he op imal PI con olle coe icien s a e ob ained om he weigh coe icien s o he ac o -ne wo k. The pe o mance o he op imal PI dynamics, which we e lea ned by using he TD3 algo i hm and Deep De e minis ic Policy G adien algo i hm, a e compa ed. Mo eo e , con ol pe o mance imp o emen o his RL based PI con olle uning me hod (RL-PI) is demons a ed ela i e o pe o mances o bo h in ege and ac ional o de PI con olle s ha we e uned by using se e al popula me aheu is ic op imiza ion algo i hms such as Gene ic Algo i hm, Pa icle Swa m Op imiza ion, G ey Wol Op imiza ion and Di e en ial E olu ion. Keywo ds: Deep ein o cemen lea ning, DC mo o , PI con olle , Twin-delayed deep de e minis ic policy g adien , me aheu is ic op imiza ion 1. In oduc ion Rein o cemen Lea ning (RL) is an e ec i e machine lea ning me hod ha is designed o lea ning om expe ience (Kaelbling e al., 1996; Mnih e al., 2013). In ecen yea s, i has been u ilized o in elligen con ol o eal sys ems ( Mnih e al., 2015; Lillic ap e al., 2016; Rabaul e al., 2019). Today, many sys ems in daily use include DC mo o s o con e elec ical ene gy o mechanical ene gy. The e o e, pe o mance imp o emen in DC mo o con ol con ibu es o many 3 inno a i e applica ion a eas such as elec ic ehicles (Wu e al., 2004), Unmanned Ae ial Vehicles (UAV) (Solomon e al., 2006; Solomon, 2007), e c. DC mo o s ha e been equen ly u ilized in many applica ion a eas (Cui e al., 2012; Be ahim, 2014). Thei lowe p ice, ease o use, lexibili y, and obus ness a e he main easons o p e e ing DC mo o s in applica ions. Today, hey a e used in nume ous a eas, such as obo s, indus ial machine y, home equipmen , and elec ic ehicles . In daily li e applica ions, op imal speed con ol o DC mo o s has impo ance in e ms o e iciency and com o . The main objec i e o DC mo o speed con ol is o p oduce he desi ed speed wi hin a ce ain e e ence alue ange in he sho es ime and o ejec en i onmen al dis u bances such as load al e a ions and changes in ope a ing condi ions. The mos commonly used DC mo o speed con ol echniques is based he P opo ional In eg al De i a i e (PID) con olle amily (Sabi & Khan, 2014), which can conside he e o signal, he change in he e o signal, and he sum o he e o signal in o de o p oduce a con ol signal. The PID con olle amily is a widely p e e ed indus ial con ol s anda d because o hei e ec i eness and hei simplici y in he con olle ealiza ion (Ekinci & Hekimoglu, 2019). The e is a la ge amoun o esea ch collec ion o heo y and p ac ice o PID con ol sys ems in he li e a u e (Ås öm & Hägglund, 1995; Visioli, 2006). Howe e , analy ical op imal uning me hods a e no commonly conside ed en i onmen al unce ain ies and dis u bance impac s. The e o e, he con olle uned analy ically may no exhibi op imal design pe o mance when hey a e applied o a eal sys em. The e a e many app oaches o add ess he e iciency and pe o mance o he PID con olle s in design asks (Ås öm & Hägglund, 1995; Visioli, 2006; Young e al., 1999; Alagoz e al., 2015). Especially, he mul i-loop model- e e ence con ol s uc u es can imp o e he dis u bance ejec ion pe o mance o he classical PID con olle loops (Bu le e al., 1989; Alagoz e al., 2020); howe e , hese me hods do no gua an ee an empi ically alida ed dis u bance ejec ion con ol pe o mance. In eal con ol applica ions, unp edic able sys em pe u ba ions and en i onmen al dis u bances may occu . The e o e, he con ol sys em mus be designed o be obus enough agains pa ame ic unce ain y, measu emen noise and unp edic able dis u bance in o de o main ain he desi ed con ol pe o mance in p ac ice. A solu ion o his design issue may be he expe ience-based op imal uning o con olle s in a ealis ic con ol simula ion en i onmen ha can simula e he s ochas ic na u e o eal-wo ld sys ems by inco po a ing andom sys em pe u ba ion and dis u bances. This simula ion en i onmen p o ides a use ul aining en i onmen o con olle uning algo i hms o each he bes p ac ical pe o mance agains unce ain y, noise and en i onmen al dis u bances. RL me hods a e inhe en ly sui able o he expe ience-based lea ning and pe o mance imp o emen acco ding o simula ion en i onmen esul s, and his poin was a cen al mo i a ion o he au ho s in he cu en s udy. We pa icula ly ocused on designing a 4 sui able simula ion en i onmen o e ec i e RL based con olle uning. Then, we illus a ed implemen a ion o he TD3 algo i hm in o de o manage expe ience-based lea ning o obus pe o mance PI con olle dynamics o he mo o con ol applica ion in he designed simula ion en i onmen . Recen ly, T aue e al. poin ed ou he impo ance o RL en i onmen design o in elligen mo o con ol (T aue e al., 2022). speci ically ocus on con ol pe o mance obus ness and dis u bance ejec ion imp o emen p oblems in he design s age o he RL aining en i onmen . To con ibu e o hei pe spec i es, we add essed he p oblem o applica ion speci ic simula ion en i onmen design o imp o emen o obus con ol pe o mance o PI con olle s. Simila o T aue e al., Book e al. implemen ed he Deep De e minis ic Policy G adien (DDPG) algo i hm o RL con ol o elec ic mo o s and p esen ed expe imen al esul s (Book e al., 2021). In ano he ecen wo k, applica ion o he deep Q-ne wo ks (DQN) algo i hm was demons a ed o he speed con ol o pe manen magne synch onous mo o s and compa ed wi h pe o mance o classical PI con olle s (Song e al., 2021). Howe e , hese wo ks employed he ained RL ne wo k as a di ec RL con olle ins ead o uning a s anda d con olle . Di ec RL con ol may lead o p ac ical ealiza ion di icul ies such as he compu a ional complexi y o RL aining and he s abili y conce ns. Table 1 lis s some inhe en p ope ies o he me aheu is ic op imiza ion and RL algo i hms ha make hem ad an ageous so compu a ion ools o con ol applica ions. The me aheu is ic me hods can be ad an ageous o sol e well-de ined op imal con ol p oblems because o hei e ec i e global sea ch capabili ies. Howe e , he expe ience based lea ning acco ding o he agen - en i onmen in e ac ion is a c i ical p ope y o p e e he RL me hod o con olle uning acco ding o well-designed simula ion en i onmen s ha can mee applica ion speci ic con ol equi emen s. Table 1. Some use ul p ope ies o me aheu is ic op imiza ion and RL algo i hms o con ol applica ions P ope ies Me aheu is ic op imiza ion based uning me hods RL based uning me hods Suppo ing popula ion based global sea ching High No Suppo ing local sea ching High High Suppo ing o -line con olle uning High High Suppo ing online con olle uning Low High Sui abili y o algo i hms o he expe ience based lea ning om agen -en i onmen in e ac ion Low High Sui abili y o algo i hms o wo k as con olle Low High The RL has also been used o implemen a neu al con olle (Koch e al., 2019) and o uning classical con olle s (Chen e al., 2022; Esmaeili e al., 2017) in con ol applica ions. The uzzy Q- lea ning me hod was sugges ed o uning he PID con olle based on RL and applied o empe a u e con ol (Chen e al., 2022). Esmaeili e al. p oposed an imp o ed RL-based uzzy-PID 5 con olle o load equency con ol o an island mic og id (Esmaeili e al., 2017). The main conce n may be he ques ion whe he he RL algo i hms can lea n om whole sys em dynamics and con ol hem. A sui able implemen a ion o he RL algo i hm can manage con ol o sys em dynamics based on he aining expe iences. In DC mo o speed con ol, he PI con olle om he PID con ol amily has been widely p e e ed (Naga ajan e al., 2016; Kanojiya e al., 2012; Sunda eswa an & Vasu, 2000). A PI con olle can be sui able o DC mo o speed con ol because de i a i e elemen s o PID may cause e y as changes and cha e in he con ol signal and may ampli y noise signals in he con ol loop o DC mo o s. The e o e, o a smoo he and mo e s aigh o wa d con ol solu ion, PI con ol can be p e e ed o speed con ol o DC mo o s. In his s udy, a RL app oach is implemen ed o lea n he PI con ol dynamics ha can p o ide a desi ed DC mo o con ol pe o mance wi hin he con ol simula ion en i onmen . The RL me hod needs a ealis ic simula ion en i onmen o lea n om agen expe iences. Majo p ac ical con ol complica ions such as unp edic able ex e nal dis u bance and measu emen noise can be in oduced in his simula ion en i onmen and con ibu e o lea ning an e ec i e and op imal PI con ol dynamic om expe ience in he simula ion en i onmen . Acco ding o RL policies, he agen ac ing in his en i onmen ies o ind he mos app op ia e ac ion ha maximizes he ewa d (Kaelbling e al., 1996). The e o e, he op imal PI con ol dynamics ha deal wi h he complica ion can be lea ned om he con ol simula ion en i onmen in o de o pe o m he bes PI ac ion by using a sui ably designed ewa d and obse e mechanism. His o ically, he RL s a egy, which was ounded by Richa d Bellman in he 1950s, i s gained success in playing a backgammon game. As a esul , combining i wi h an a i icial neu al ne wo k o e ime achie ed supe io success and pe o med be e han humans in sol ing complex p oblems (Mnih e al., 2015; G aepel, 2016). La e , Q-Lea ning was i s o come ou in 1989, which is a special ype o RL app oach. Wa kins used he le e Q o he alue unc ion, which is based on he heo y o Ma ko decision p ocesses (Wa kins, 1989; Wa kins & Dayan, 1992). Howe e , Q-Lea ning has no a ac ed in e es in i s domains un il DQN algo i hms we e de eloped (Mnih e al., 2013). A e wa ds, he De e minis ic Policy G adien (DPG) algo i hm, which uses he ac o -c i ic s uc u e, was de eloped (Lillic ap e al., 2016; Sil e e al., 2014). Finally, he TD3 Policy G adien algo i hm was p oposed, and he TD3 algo i hm can wo k mo e e ec i ely han many RL me hods in applica ions (Fujimo o e al., 2018). RL-based me hods a e widely used in he ield o con ol. I can enhance in elligence in con ol sys ems (Lillic ap e al., 2016; Rabaul e al., 2019; Chen e al., 2018; B andi e al., 2020; Sa heeshbabu e al., 2019). In he cu en s udy, he TD3 algo i hm is implemen ed o lea ning op imal PI con ol dynamics om a closed-loop DC mo o simula ion en i onmen . A e e ence model in oduced by a se poin il e is used in his simula ion en i onmen o de ine desi ed s ep 6 esponses wi hin a ange o s ep inpu ampli udes o he TD3 algo i hm. The addi i e inpu dis u bance model wi h he andom ampli ude s ep dis u bance was implemen ed in he simula ion en i onmen o ain he RL-based PI con ol dynamics o imp o ed dis u bance ejec ion. The ac o -ne wo k was designed o yield he PI coe icien s o he lea ned op imal PI dynamics in his en i onmen . Thus, an RL uned PI con olle was ob ained o he DC mo o speed con ol, and i s pe o mance was es ed and compa ed wi h he esul s o he me aheu is ic op imiza ion based op imal PI and ac ional o de PI (FOPI) uning algo i hm. The main con ibu ions o his s udy can be summa ized as ollows: (i) This s udy demons a es he design o con ol simula ion en i onmen s in o de o lea n op imal and dis u bance ejec PI dynamics om a se -poin e e ence model by using he TD3 algo i hm. Thus, he e e ence model esponse based op imal pe o mance uning o he PI con olle acco ding o he ealis ic en i onmen al condi ions (e.g., andom dis u bance inse ion, andom change o inpu signal in a wide ange) was pe o med by using he ein o ced lea ning s a egy. In his lea ning pe spec i e, he e e ence model desc ibes a desi ed con ol sys em esponse (o e shoo , se ling ime e c.) and imp o es RL based con olle uning p ocess acco ding o con ol applica ion equi emen s and speci ica ions. (ii) To in es iga e pe o mance imp o emen o he RL based con olle uning me hod, con olle uning pe o mance o he TD3 algo i hm was compa ed wi h pe o mances o s a e-o -a con olle uning me hods in he cu en s udy. We illus a e he pe o mance imp o emen o RL-based PI uning in o e shoo - ee, smoo h con ol o DC mo o speed compa ed o se e al me aheu is ic PI and FOPI con olle uning me hods in he e e ence model based con olle uning. Also, pe o mance imp o emen o RL-PI con olle wi h he TD3 algo i hm is demons a ed o e ha o he DDPG algo i hm in his DC mo o con ol p oblem. In he es o he pape , Sec ion 2 p o ides p elimina y knowledge o RL and b ie ly in oduces TD3 policy g adien me hod algo i hm. A e p esen ing he speed con ol model o DC mo o s, which is used in he simula ion en i onmen , a Simulink simula ion en i onmen o aining o he TD3 algo i hm o lea n an op imal and obus PI con olle dynamics o DC mo o speed con ol is explained in Sec ion 3. Sec ion 4 p esen s a simula ion s udy ha compa es esul s o RL based PI con olle design and esul s o PI and FOPI con olle designs o o he popula me hods. Some conclusions a e summa ized in Sec ion 5. 7 2. P elimina ies and P oblem S a emen 2.1. P elimina ies o Rein o cemen Lea ning The RL is a machine lea ning echnique ha lea ns om he di ec in e ac ion o asse s wi h he en i onmen (Su on & Ba o, 1998). The agen lea ns he en i onmen and ules h ough expe ience o achie e goals in he en i onmen . The ewa d and punishmen mechanisms a e used in he aining p ocess, and hese mechanisms guide he agen 's lea ning (Hoshino & Kamei, 2003; Russell & No ig, 2003). The main di e ence om o he lea ning me hods is ha he sys em only lea ns by ial and e o in i s own en i onmen wi hou using any p ede ined ins uc ion se s (Su on & Ba o, 1998; Na end a & Tha hacha , 2012). The ewa d mo i a es he agen o choose he bes ac ion in he en i onmen . In o de o ind he bes way o sol e he p oblem, he sys em's lea ning p ocess depends on a su icien numbe o ials (Na end a & Tha hacha , 2012). The agen aims o ob ain an op imal policy by maximizing he o al ewa d. This lea ning me hod can be assumed as a Ma ko ian p ocess (Bellman, 1957), and i s heo e ical ounda ion assumes RL o be a s ochas ic lea ning p ocess ha can p og essi ely enhance se and ial esul s by sui ably ewa ding agen ac ions in he en i onmen . A each disc e e ime inc emen , he agen selec s an ac ion acco ding o i s policy depending on he s a e . Then, he agen ecei es a ewa d , and a new s a e is ecei ed om he en i onmen o he ac ion (Fujimo o e al., 2018). In he Q-Lea ning, he quali y o a s a e-ac ion couple is e alua ed acco ding o he unc ion , and i exp esses expec ed ewa ds om an ac ion a he s a e . The is calcula ed by using he Bellman equa ion. The o m o he Bellman equa ion is adap ed as . (1) This equa ion implies ha he maximum u u e ewa d can be possible wi h he ewa d o he cu en ac ion ( ) plus he maximum o u u e ewa d es ima ions . To sol e he equa ion i e a i ely, he Bellman e o l B ea i e a ion is de ined as . (2) Then, an i e a i e scheme o each op imal s a e-ac ion alue was gi en by using empo al di e ence lea ning (Su on, 1988) as , (3) 8 whe e, is a discoun ac o be ween ze o and one, de e mining he weigh s o sho - e m ewa ds (Luu, 2015). The pa ame e is he lea ning a e. The unc ion is he ewa d o ac ion a and obse e s a e . This o mula ion de ines he RL p ocess as a Ma ko p ocess. In o he wo ds, u u e si ua ions depend on only p esen si ua ions in his p ocess. The e o e, condi ional p obabili y is widely used o model RL p ocesses. Ma ko p ocess o a s a e can be desc ibed as . (4) I he cu en s a e o a sys em is conside ed, p e ious ac ion om a he s a e om leads o a new s a e om and a ewa d om . The p obabili y densi y unc ion o he andom a iables and can be exp essed o he s ochas ic modeling o he RL as , (5) whe e , ' ,s s S R and . I can be seen ha Equa ion (1) depends on he s a e o he ewa d a he ime and he maximum o s a e-ac ion quali y. The e o e, he possibili ies o ansi ion o is w i en by conside ing he ma ginal dis ibu ion o he s a e (Mo ales & Za agoza, 2011). (6) Acco dingly, he expec ed ewa d can be w i en by using a ma ginal dis ibu ion o s a es as (Mo ales & Za agoza, 2011) . (7) To implemen his scheme, app oxima ion unc ions such as neu al ne wo ks a e used o lea n ac ions ha maximize ewa ds om he en i onmen . P og ess in RL me hods con inues and applica ion ields o RL ha e expanded in many ields such as a i ude con ol o hype sonic een y ehicles (Y. Liu e al., 2022), zone scheduling op imiza ion o pumps in wa e dis ibu ion ne wo ks (Xu e al., 2021), lexible con ol o disc e e e en sys ems using en i onmen simula ion (Zielinski e al., 2021) and au onomous na iga ion o UAV in mul i-obs acle en i onmen s (Zhang e al., 2022). Ka u an con ibu ed o his discussion by in es iga ing he e ec o deep RL on he ac ional-o de oscilla o (Ka u an, 2022). 2.2. A B ie In oduc ion o Twin Delayed Deep De e minis ic Policy G adien Algo i hm The TD3 policy g adien algo i hm, which is an imp o emen o he DDPG algo i hm, akes in o accoun he unc ion app oach e o (Lillic ap e al., 2016; Fujimo o e al., 2018). The objec i e o he RL is o ind he op imal policy ha maximizes he expec ed ewa d by uning he 15 wi h one neu on o he ou pu . The o he ne wo ks a e he ac o neu al ne wo ks (ac o - ne wo k) ha yield he ac ion ( ) by using obse e s a es ( s ). Based on alues, he TD3 algo i hm ains c i ic-ne wo ks and ac o -ne wo ks o maximize he ewa d as an op imal con ol objec i e gi en by Equa ion (21) in his s udy. Ac ion Inpu S a e Inpu Fully Connec ed Laye Fully Connec ed Laye Conca ena ion Laye ReLU Laye Fully Connec ed Laye ReLU Laye Fully Connec ed Laye S a e Pa hAc ion Pa h Figu e 5. The a chi ec u e o he c i ic neu al ne wo k ha is implemen ed in his applica ion The s a e obse e moni o s impo an s a es o he con ol sys em om he simula ion en i onmen . To ma ch he weigh coe icien o he ac o -ne wo k wi h he PI con olle coe icien s in he Ma lab RL con ol example, he in eg al o he e o signal is obse ed o implemen he in eg al s a e ( ), and he e o signal i sel is obse ed o implemen p opo ional s a e ( ) o PI dynamics. The s a e obse e was implemen ed by moni o ing hese wo s a es o PI dynamics as . (22) Ac o ne wo k is designed by using a wo-inpu ully connec ed laye wi h a linea ac i a ion unc ion. Two obse e s a es eed he ac o ne wo k. Acco dingly, he ac ion o he ac o can be easily w i en by using he obse e s a es as ollows . (23) A e he aining o he ac o ne wo k is comple ed, he weigh s ( , ) o he neu al ne wo k exp esses he lea ned PI con olle coe icien s. He e, he weigh co esponds o he 16 coe icien , and he weigh co esponds o he coe icien o he disc e e PI con olle implemen a ion. To demon ollow: (24) whe e exp esses he disc e e- ime PI con olle unc ion, whe e he p opo ional coe icien o he PI con olle is , and he in eg al coe icien o he PI con olle is . 4. Simula ion S udy In o de o ob ain he desi ed speed con ol pe o mance o he DC mo o by using he TD3 policy g adien algo i hm, he aining o he RL sys em was ca ied ou by using he Simulink simula ion en i onmen shown in Figu e 3. Con ol simula ions we e pe o med o 20 seconds o each episode, and an e o signal om each simula ion was used o ob ain obse e s a es (Equa ion 22). E o and con ol signals om he simula ion en i onmen a e used o calcula e he ewa d (Equa ion 19) in he aining o he TD3 algo i hm. Table 4 lis s some hype pa ame e s used o he con igu a ion o he aining p ocess o TD3 agen s. The ial and e o me hod was used o se sui able alues o hese pa ame e s. Table 4. Va iables o aining p ocess o he TD3 algo i hm Maximum Numbe o Episodes (T) 100 Mini Ba ch Size 128 Sampling ime 0.1 s Lea ning Ra e 0.001 Du ing he aining o he RL sys em, e e ence s ep inpu signals wi h andom ampli udes (wi hin he ange o 0-155) we e applied o he inpu o he sys em in o de o imp o e se poin con ol pe o mance, and he s ep dis u bance signals wi h andom ampli ude (wi hin he ange o [- 5,5]) we e used o ain he sys em o ob ain an imp o ed dis u bance ejec ion pe o mance. To gi e a ewa d acco ding o he op imal con ol o mula ion (Equa ion (19)), pa ame e s and a e se o 0.9 and 0.01, espec i ely. The RL aining in he simula ion en i onmen (Figu e 3) allows he ac o neu al ne wo k o espond simila o a PI con olle by adjus ing he ac o -ne wo k weigh s ( , ). The ac o - ne wo k is he ully connec ed single neu on wi h wo inpu s om obse ed s a es. Equa ions (23) and (24) p o e ha a combina ion o ac o -ne wo k and obse e block implemen s a PI con olle . The ou pu o he ac o -ne wo k is he con ol signal o he DC mo o . Figu e 6 shows he change 17 o ins an alues and a e age alues o ewa ds ob ained du ing he aining p ocess in he simula ion en i onmen . The a e age ewa d is he a e age o he ins an ewa ds and exp esses o e all pe o mance du ing he aining p ocess. Figu e 6. Change o ins an alues and a e age alues o ewa ds in he simula ion en i onmen When he aining was comple ed, he p w was ob ained 1.6626, which is he equi alen alue o a disc e e- ime PI con olle , and he was ob ained 2.0830, which is he equi alen o a disc e e- ime PI con olle . Acco ding o hese aining esul s, he lea ned disc e e PI con olle dynamics a e equal o 1 0830.26626.1)( z T zC s. (25) In his s udy, he RL-PI con ol was ewa ded o lea n he s ep esponse o he se poin il e 1 1 )( s sFse , and i enables smoo hly se ling o elec ic mo o speed o he e e ence speed 100 pm wi hou any o e shoo and ipples. This is an impo an p ope y o elec ical mo o con ol because he o e shoo s and ipples can cause edundan accele a ion and luc ua ions in he DC mo o 's speed esponses. These e ec s esul in discom o and unnecessa y ene gy consump ion in con ol applica ions, whe e DC mo o s a e used, o example applica ions in elec ical ehicles, UAVs, e c. To analyze pe o mances o he designed PI con olle s, a con ol sys em es simula ion o 100 seconds du a ion was ca ied ou , and an addi i e inpu dis u bance wi h a s ep wa e o m ( he ampli ude o s ep dis u bance is 20) is applied a he simula ion ime o 50 seconds. The pe o mance o he RL-PI con ol is compa ed wi h he esponses o he PI and FOPI con olle s ha we e uned by using se e al me aheu is ics such as Gene ic Algo i hm (GA) (Holland, 1973), A e age Rewa d * Ins an Rewa d o 18 Pa icle Swa m Op imiza ion (PSO) (Kennedy & Ebe ha ,1995), G ey Wol Op imiza ion (GWO) (Mi jalili e al., 2014) and Di e en ial E olu ion (DE) (S o n & P ice,1997). Simila o he RL algo i hms, me aheu is ic op imiza ion me hod can be e ec i ely used o seeking op imal solu ions in decision making and scheduling p oblems such as op imal cha ging scheduling p oblem o elec ic ehicles (Liu e al., 2020), ba ch-p ocessing machine scheduling p oblem (Zhou e al., 2021 ), ene gy-e icien dis ibu ed no-idle low-shop scheduling p oblem (Zhao e al., 2020; Zhao e al., 2021a; Zhao e al., 2021b). Also, success ul use o se e al me aheu is ic me hods was shown in con olle uning p oblems (Tu enkci e al., 2020; Zheng & Pi, 2016; Koma hi & Umamaheswa i, 2020; Liu & Hsu, 2010). Besides, hype heu is ic app oaches ha in ol e high-le el heu is ics and low-le el heu is ics, can p o ide imp o ed pe o mance in a wide applica ion domain owing o high-le el ac ics (Zhao e al., 2021c). The me aheu is ic me hods op imized he PI con olle coe icien s o yield a s ep esponse ha esembles he esponse o he se poin il e . This objec i e is he same as in he RL-PI con ol sys em, whe e he il e is a supe iso (a e e ence model) ha was used o shape he s ep esponse o he con ol sys em acco ding o esponses o he e e ence model , which p o ides smoo h and o e shoo - ee con ol o he DC mo o . Me aheu is ic me hods ha e been widely used o op imal uning o he con olle unc ions. The e o e, we demons a ed he obus con ol pe o mance imp o emen s o he RL-PI con olle wi h TD3 algo i hm compa ed o esul s o GA, PSO, GWO and DE algo i hms. The sol ed op imiza ion p oblem o op imal uning o PI con olle unc ions ( ) and FOPI con olle ( ) ia me aheu is ic me hods we e shown in he Appendix. Table 5 shows PI and FOPI con olle uning esul s o GA, PSO, GWO and DE algo i hms. Since he objec i e o e e ence model based con olle uning aims o app oxima e o he esponse o he e e ence model 1 1 )( s sFse , he op imal con olle coe icien s we e ob ained o close each o he . Figu es 7 and 8 show s ep esponses o hose con olle s in compa ison wi h he esponse o he RL-PI con olle wi h he TD3 algo i hm. Figu es 9 and 10 show e o signals om hese con olle simula ions. Figu es 11 and 12 show he change o con ol signals du ing hese con ol simula ions. Main ad an ages o he RL-PI con olle uning comes om he design o a sui able simula ion en i onmen ha in ol es andom ampli ude dis u bance inse ion and andom se poin s. The expe ience based lea ning o he PI con olle dynamics can explo e mo e obus esponses o he inpu dis u bances in he RL based uning. The me aheu is ic uning me hods only sol ed he op imiza ion p oblem ha is de ined in he Appendix. The e o e, me aheu is ic uning me hods, which a e based on analy ical solu ions o an op imal con ol p oblem, canno di ec ly conside e ec s o andom ampli ude dis u bance signal o he andom ampli ude e e ence signal on con ol 19 pe o mance (Ozbey e al., 2020). This sho coming can limi he uning pe o mance o me aheu is ic op imiza ion me hods in applica ion speci ic design. Table 5. Op imal PI and FOPI con olle unc ions ( )(sC ) ha we e uned by me aheu is ic algo i hms Algo i hms Con olle s Algo i hms Con olle s GA-PI GWO-PI GA-FOPI GWO-FOPI PSO-PI DE-PI PSO-FOPI DE-FOPI Figu e 7. S ep esponses ob ained by GA-PI, PSO-PI, GWO-PI, DE-PI con olle s and RL-PI con olle wi h TD3 algo i hm Figu e 8. S ep esponses ob ained by GA-FOPI, PSO-FOPI, GWO-FOPI, DE-FOPI con olle s and RL-PI con olle wi h TD3 algo i hm 20 Figu e 9. E o signals o GA-PI, PSO-PI, GWO-PI, DE-PI and RL-PI con olle s Figu e 10. E o signals o GA-FOPI, PSO-FOPI, GWO-FOPI, DE-FOPI and RL-PI con olle s 21 Figu e 11. Con ol signals o GA-PI, PSO-PI, GWO-PI, DE-PI and RL-PI con olle s Figu e 12. Con ol signals o GA-FOPI, PSO-FOPI, GWO-FOPI, DE-FOPI and RL-PI con olle s We also compa ed pe o mances o wo di e en RL s a egies in he PI con olle uning p oblem. When he aining o DDPG algo i hm was comple ed, he disc e e- ime con olle coe icien s we e calcula ed as and o a disc e e- ime PI con olle . Acco ding o hese aining esul s, he DDPG lea ned disc e e PI con olle dynamics a e equal o 22 . (26) Figu e 13 shows he change o ins an alues and a e age alues o ewa ds du ing he aining wi h he DDPG algo i hm in he same simula ion en i onmen . Figu e 14 shows s ep esponses o he RL-PI con olle s ha we e lea ned by he TD3 algo i hm and he DDPG algo i hm. The DDPG algo i hm has been used as RL s a egy in he mo o con ol applica ion (T aue e al., 2022; Book e al., 2021). The same simula ion en i onmen and he same neu al ne wo k a chi ec u es we e used o hese algo i hms. Also, pa ame e se ings o he DDPG algo i hm we e con igu ed o be he same wi h he TD3 algo i hm in he aining. Figu e 15 and 16 show compa isons o e o signals and con ol signals. Simula ion esul s showed ha he RL-PI con olle uned wi h he TD3 algo i hm p o ided be e PI con olle pe o mance compa ed o he RL-PI con olle uned wi h he DDPG algo i hm. In gene al, he TD3 algo i hm can be mo e consis en in con e gence o maximum ewa d alues han he DDPG algo i hm, mainly because o using wo c i ic ne wo k and ac o ne wo ks wi h a delayed upda e and ac ion noise egula iza ion (clipped noise addi ion as in he equa ion (10) o a clipped double Q-lea ning e ec (Fujimo o e al., 2018)), which make TD3 mo e ad an ageous o escape om he o e i ing o na ow ewa d peaks and imp o es he sea ch skill o he TD3 algo i hm. The di e ence in ewa ds cha ac e is ics in Figu e 6 ( o he TD3 algo i hm) and Figu e 13 ( o he DDPG algo i hm) con i ms imp o emen s o he TD3 algo i hm o e he DDPG algo i hm in his con ol p oblem. Figu e 13. Change o ins an alues and a e age alues o ewa ds in he simula ion en i onmen A e age Rewa d * Ins an Rewa d o 23 Figu e 14. S ep esponses ob ained by he RL-PI con olle wi h TD3 algo i hm and he RL-PI con olle wi h DDPG algo i hm Figu e 15. E o signals o he RL-PI con olle wi h TD3 algo i hm and he RL-PI con olle wi h DDPG algo i hm 24 Figu e 16. Con ol signals o he RL-PI con olle wi h TD3 algo i hm and he RL-PI con olle wi h DDPG algo i hm Table 6 lis s he Mean Squa e E o ( ,whe e he is numbe o sampled da a om simula ion en i onmen ) and A e age o Absolu e Con ol signal ( ) pe o mances o he conside ed algo i hms o he PI con olle uning e o s. We pe o med 5 imes independen ial uns o each algo i hm. In hese es s, o ob ain di e en esul s, he RL-PI algo i hms we e ini ia ed a 5 di e en ini ial poin s o he PI con olle coe icien s. He e, he MSE indica es he dissimila i y be ween he esponse o he DC mo o con ol sys em and esponses o he e e ence model . The AAC exp esses he a e age le el o con ol signal ampli ude ha is an impo an indica o ela ed o he ene gy e iciency o he con ol ac ions. Table 6. MSE and AAC pe o mances o he compa ed PI con olle uning me hods Con olle s MSE AAC Min A e age Max S anda d De ia ion A e age GA-PI 26.28310 26.28360 26.28390 3.16876 10-4 111.8900 PSO-PI 26.25929 26.28388 26.31272 1.91939 10-2 111.8901 GWO-PI 26.27081 26.70522 28.38485 9.39030 10-1 111.8901 DE-PI 26.28387 26.28389 26.28390 1.17834 10-5 111.8900 RL-PI (DDPG) 6.011537 111597.0 557907.5 2.49495 105 315.9496 RL-PI (TD3) 9.053762 14.91569 20.75544 5.66412 111.8495 31 h ps://doi.o g/10.1007/S00521-020-05352-1/FIGURES/9 S o n, R., & P ice, K. (1997). Di e en ial e olu ion a simple and e icien heu is ic o global op imiza ion o e con inuous spaces. Jou nal o global op imiza ion, 11(4), 341-359. h ps://doi.o g/10.1023/A:1008202821328 Sunda eswa an, K., & Vasu, M. (2000). Gene ic uning o PI con olle o speed con ol o DC mo o d i e. P oceedings o he IEEE In e na ional Con e ence on Indus ial Technology, 1, 521 525. h ps://doi.o g/10.1109/ici .2000.854212 Su on, R. S. (1988). Lea ning o p edic by he me hods o empo al di e ences. Machine Lea ning, 3(1), 9 44. h ps://doi.o g/10.1007/b 00115009 Su on, R. S., & Ba o, A. G. (1998). Rein o cemen Lea ning: An In oduc ion. MIT p ess, Camb ige, MA. T aue, A., Book, G., Ki chgassne , W., & Wallscheid, O. (2022). Towa d a Rein o cemen Lea ning En i onmen Toolbox o In elligen Elec ic Mo o Con ol. IEEE T ansac ions on Neu al Ne wo ks and Lea ning Sys ems, 33(3), 919 928. h ps://doi.o g/10.1109/TNNLS.2020.3029573 - domain. In Jou nal o Ad anced Resea ch, 25, 171 180. h ps://doi.o g/10.1016/j.ja e.2020.03.002 Uni e si y o Michigan. (2017). Con ol Tu o ials o MATLAB and Simulink - Mo o Speed: Sys em Modeling. h ps://c ms.engin.umich.edu/CTMS/index.php?example=Mo o Speed§ion=Sys emModeling Visioli, A. (2006). P ac ical PID Con ol. In P ac ical PID Con ol. h ps://doi.o g/10.1007/1-84628-586-0 Wa e Tank Rein o cemen Lea ning En i onmen Model - MATLAB & Simulink - Ma hWo ks Swi ze land. (n.d.). Re ie ed Ma ch 24, 2022, om h ps://ch.ma hwo ks.com/help/ ein o cemen -lea ning/ug/wa e - ank- ein o cemen -lea ning-en i onmen -model.h ml Wa kins, C. J. C. H. (1989). Lea ning om delayed ewa ds. In Robo ics and Au onomous Sys ems, 15(4), 233 235. Wa kins, C. J. C. H., & Dayan, P. (1992). Q-lea ning. Machine Lea ning, 8(3 4), 279 292. h ps://doi.o g/10.1007/b 00992698 Wu, H. X., Cheng, S. K., & Cui, S. M. (2004). A con olle o b ushless DC Mo o o elec ic ehicle. 2004 12 h Symposium on Elec omagne ic Launch Technology, 528 533. Xu, J., Wang, H., Rao, J., & Wang, J. (2021). Zone scheduling op imiza ion o pumps in wa e dis ibu ion ne wo ks wi h deep ein o cemen lea ning and knowledge-assis ed lea ning. So Compu ing, 25(23), 14757 14767. IEEE T ansac ions on Con ol Sys ems Technology, 7(3), 328 342. h ps://doi.o g/10.1109/87.761053 Zhang, S., Li, Y., & Dong, Q. (2022). Au onomous na iga ion o UAV in mul i-obs acle en i onmen s based on a Deep Rein o cemen Lea ning app oach. Applied So Compu ing, 115, 108194. Zhao, F., He, X., & Wang, L. (2020). A wo-s age coope a i e e olu iona y algo i hm wi h p oblem-speci ic knowledge o ene gy-e icien scheduling o no-wai low-shop p oblem. IEEE T ansac ions on Cybe ne ics, 51(11), 5291 5303. h ps://doi.o g/10.1109/TCYB.2020.3025662 Zhou, S., Xing, L., Zheng, X., Du, N., Wang, L., & Zhang, Q. (2021). A Sel -Adap i e Di e en ial E olu ion Algo i hm o Scheduling a Single Ba ch-P ocessing Machine wi h A bi a y Job Sizes and Release Times. IEEE T ansac ions on Cybe ne ics, 51(3), 1430 1442. h ps://doi.o g/10.1109/TCYB.2019.2939219 Zhao, F., Ma, R., & Wang, L. (2021a). A Sel -Lea ning Disc e e Jaya Algo i hm o Mul iobjec i e Ene gy-E icien Dis ibu ed No-Idle Flow-Shop Scheduling P oblem in He e ogeneous Fac o y Sys em. IEEE T ansac ions on Cybe ne ics. h ps://doi.o g/10.1109/TCYB.2021.3086181 Zhao, F., Zhang, L., Cao, J., & Tang, J. (2021b). A coope a i e wa e wa e op imiza ion algo i hm wi h ein o cemen lea ning o he dis ibu ed assembly no-idle lowshop scheduling p oblem. Compu e s and Indus ial Enginee ing, 153. h ps://doi.o g/10.1016/j.cie.2020.107082 32 Zhao, F., Di, S., Cao, J., Tang, J., & Jon inaldi. (2021c). A No el Coope a i e Mul i-S age Hype -Heu is ic o Combina ion Op imiza ion P oblems. Complex Sys em Modeling and Simula ion, 1(2), 91 108. Zheng, W., & Pi, Y. (2016). S udy o he ac ional o de p opo ional in eg al con olle o he pe manen magne synch onous mo o based on he di e en ial e olu ion algo i hm. ISA T ansac ions, 63, 387 393. h ps://doi.o g/10.1016/j.isa a.2015.11.029 Zielinski, K. M. C., Hendges, L. V., Flo indo, J. B., Lopes, Y. K., Ribei o, R., Teixei a, M., & Casano a, D. (2021). Flexible con ol o Disc e e E en Sys ems using en i onmen simula ion and Rein o cemen Lea ning. Applied So Compu ing, 111, 107714. h ps://doi.o g/10.1016/J.ASOC.2021.107714