Full text
Dis ibu ed Ne wo k Con ol o QoS Assu ance in Mul i-
Domain Ne wo ks
H. Shakespea -Miles, S. Ba zega , M. Ruiz, and L. Velasco
Op ical Communica ions G oup (GCO), Uni e si a Poli ècnica de Ca alunya (UPC), Ba celona, Spain
e-mail: [email p o ec ed]
ABSTRACT
The deploymen o beyond 5G and 6G ne wo ks in oduces many new se ices wi h s ingen Quali y o Se ice
(QoS) equi emen s. Recen ly machine lea ning has been shown o be a iable solu ion in p oposing adap able
solu ions. Howe e , cen alized machine lea ning based solu ions s ill encoun e hu dles in achie ing eal- ime
esponsi eness due o hei need o a global ne wo k iew. In his pape , we explo e a dis ibu ed app oach aimed
a op imizing ne wo k pe o mance in eal- ime scena ios. By using Mul i-Agen Sys ems (MAS), ou me hod
a ge s nea - eal- ime end- o-end delay assu ance ac oss di e se ne wo k domains, wi hou he need o p io a ic
p o ile knowledge. E alua ed esul s highligh he e ec i eness o ou app oach in educing ou ing cos s and
ensu ing desi ed end- o-end delay le els.
Keywo ds: Dis ibu ed Con ol, Au onomous sys ems, Deep Rein o cemen Lea ning, Mul i-Domain
1. INTRODUCTION
In o de o suppo he demands o beyond 5G and 6G se ices, anspo ne wo ks mus e ol e o handle inc eased
a ic dynamici y and s ic e pe o mance equi emen s. In ac , such suppo equi es inc eased le els o
lexibili y and au oma ion, oge he wi h highe p io i y gi en o ne wo k op imiza ion, secu i y, ene gy
consump ion, and cos e iciency. As a esul , such in as uc u es need o apply A i icial In elligence (AI) /
Machine Lea ning (ML) echniques [1] o c ea e an au oma ed managemen o implemen da a-d i en closed
con ol loops. To achie e au onomic ne wo king, So wa e de ined Ne wo king (SDN) con ol is being augmen ed
wi h ins an aneous da a-d i en decision-making [2]. This app oach is bene icial o many applica ions ha do no
equi e making decisions nea eal ime, like ailu e managemen .
In he case o dynamic a ic condi ions, cen alized decision-making leads o poo esou ce u iliza ion and high
ene gy consump ion. Howe e , p ecisely because o i s cen alized loca ion, (nea ) eal- ime decision making does
no i well wi h SDN con olle s. In pa icula , in he case ha au oma ion needs o deal wi h highly dynamic
a ic condi ions, cen alized decision-making leads o poo esou ce u iliza ion because o long esponse imes.
In his wo k, we ocus on low ou ing, whe e decisions need o be made nea eal- ime o op imize esou ce
u iliza ion while ensu ing he Quali y o Se ice (QoS) o he lows. I 's wo h no ing ha a ic a ia ions may
in oduce bo lenecks impac ing end- o-end (e2e) delay, de ined as he ime equi ed o ansmi ing low a ic
be ween wo bo de packe nodes.
In addi ion, ML algo i hms migh be execu ed as close as possible o he da a sou ces (con a ily o he cen alized
a chi ec u e o SDN) looking a minimizing he amoun o da a o be con eyed, as well as minimizing he esponse
ime. Examples, include he use o Deep Rein o cemen Lea ning (DRL) o he managemen o he capaci y a
packe [3] o op ical connec ions (ligh pa h) [4] in eal ime, whe e he capaci y o he connec ions is managed
nea eal- ime o adap o inpu a ic. No e ha packe and op ical laye s a e closely ela ed in mul i-laye
ne wo ks, whe e links connec ing packe nodes a e suppo ed by ligh pa hs. A possible amewo k o dis ibu ed
solu ions is Mul i-Agen Sys ems (MAS) [5].
This wo k consolida es indings om wo p io s udies [6][7] o unde sco e he in ica e na u e o ou ing packe
lows ac oss di e se ou pu in e aces o mul i-domain ne wo ks, emphasizing he dual objec i es o main aining
Quali y o Se ice (QoS) s anda ds and op imizing esou ce u iliza ion. The subsequen sec ions o his pape
delinea e ou app oach and indings. Sec ion II elabo a es on he dis ibu ed in elligence a chi ec u e acili a ing
nea - eal- ime decision-making in he mul i-domain scena io. Sec ion III delinea es he cons uc ion o he DRL
model, including he au onomous low ou ing ewa d unc ion and special conside a ions ha mus be made in
he mul i-domain case. Sec ion IV p esen s he simula ion esul s, and inally, Sec ion V ou lines he conclusions
de i ed om ou s udies.
2. DISTRIBUTED AUTONOMOUS FLOW ROUTING
Figu e 1 ske ches he dis ibu ed in elligence a chi ec u e, whe e a single node and he cen alized SDN con olle
a e ep esen ed. No e ha in such a chi ec u e, we a e mo ing he in elligence om he cen alized con ol plane
o he nodes hus esul ing in a hyb id cen alized SDN con ol wi h dis ibu ed ne wo k in elligence. Agen nodes
a e able o communica e wi h each o he o exchange da a and models o he sake o coo dina ion. Decision-
making is pe o med by e e y indi idual agen nea eal- ime (sub-second o ew second g anula i y) based on i s
own obse ed da a, as well as on he da a and models ecei ed om o he agen s. The con ol plane o e sees he
o e all ne wo k coo dina ion, gene a ing necessa y guidelines o agen s o ope a e au onomously wi h he desi ed
deg ee o eedom. To illus a e his concep , conside he packe laye ; packe s in a low ollow he ou e ha has
been decided om he SDN con olle . Rou e compu a ion is based on he ne wo k opology and ypically s able
© 2024 IEEE. Pe sonal use o his ma e ial is pe mi ed. Pe mission om IEEE mus be ob ained o all o he uses, in any cu en o u u e
media, including ep in ing/ epublishing his ma e ial o ad e ising o p omo ional pu poses,c ea ing new collec i e wo ks, o esale o
edis ibu ion o se e s o lis s, o euse o any copy igh ed componen o his wo k in o he wo ks.
h ps://dx.doi.o g/0.1109/ICTON62926.2024.10647121
unless ne wo k condi ions al e i . In ou app oach, he SDN con olle gi es deg ees o eedom o he packe laye
by compu ing a se o ou es (including a single one) o e e y low ha a e gi en o he node agen s as guidelines
( oge he wi h some o he pa ame e s), so he ac ual ou e is decided by he node agen s based on AI/ML models
and obse ed da a (e.g., end- o-end delay) and i migh be changed nea eal ime O he in ica e asks, like
mul ilaye issues and ailu e managemen , emain cen alized, elie ing he con ol plane om immedia e
ope a ions o p io i ize long- e m ac i i ies. Figu e 2 illus a es low ou ing wi hin a mul ilaye scena io, whe e he
packe node agen ecei es h ee po en ial ou es om he SDN con olle o a gi en a ic low. I mus use DRL
o de e mine he op imal ou e o combina ion o achie e desi ed QoS while minimizing cos . The igu e depic s
wo in e connec ed blocks: he o me is in cha ge o lea ning he bes ac ions o be aken based on he cu en
s a e and he ecei ed ewa d, whe eas he la e is in cha ge o compu ing he s a e based on he obse ed a ic
and cu en combina ion o ou es (pe cen ages), as well he ob ained ewa d based on he measu ed end- o-end
delay o he selec ed ou es. Mul iple sub- lows ollow dis inc ou es, wi h end- o-end delay measu ed a he
des ina ion and s a is ics elayed o pa icipa ing node agen s.
A speci ically challenging scena io is mul i-domain ne wo ks, whe e a packe low a e ses wo di e en
adminis a i e domains. Examples, include access ne wo ks ( ixed o mobile) and me o/co e ne wo ks. Al hough
he e2e a ic low consis s o wo segmen s, one in each domain, dmaxe2e needs o be ensu ed. E en when each
domain wo ks unde low o mode a e load egime, delay luc ua ions a e p oduced as a esul o a ic a ia ions,
which makes load also a iable in ime. In his case, he SDN con olle s o each domain ha e ecei ed he equi ed
dmaxe2e o he a ic low a p o isioning ime. I is wo h no ing ha i bo h domains ope a e wi hou any
coo dina ion among hem, la ge capaci y o e p o isioning is equi ed o abso b delay a ia ions in oduced no
only by he own domain, bu also by he o he domains a e sed by he low. In iew o ha , we assume some
so o coo dina ion among domains. An example is ep esen ed in Figu e 3. The SDN con olle o domain 1
dynamically ge s he delay ha can be ensu ed o he segmen o he low (dmaxD1) and sha e ha alue wi h he
SDN con olle o domain 2. In esponse, he SDN con olle unes he equi emen o delay o he local segmen
(dmaxD2) so dmaxe2e is ensu ed. We expec ha o e p o isioning can be g ea ly educed and e2e delay gua an eed
by adjus ing domain delay budge s dynamically.
3. DRL OPERATION IN MULTIDOMAIN SCENARIOS
In his sec ion, we de ail he ewa d unc ion o Twin Delayed Deep De e minis ic Policy G adien s (TD3) [7].
The s a e is de ined as he a io a ic o e he capaci y o he in e aces. In addi ion, each ac ion is de ined as
being ela ed o one low and ou pu in e aces and ep esen s he pe cen age o low o be sen h ough hose
in e aces. As an example, o one single low ha can be ou ed h ough 3 di e en in e aces, ac ion [50, 20, 30]
en ails 50% o a ic low h ough he i s in e ace, 20% h ough he second one, and 30% h ough he hi d one.
In ou app oach, he compu a ion o s a es and ac ions is pe o med pe iodically (e.g., e e y second).
A gene ic ewa d unc ion has been de ined in Equa ion 1 wi h he objec i e o penalizing he ac ions causing
ha some a ge delay is exceeded and/o inc easing he cos o ne wo k. In consequence, wo ewa d componen s
Node Agen
Guidelines
and Policies
In elligence
Obse ed da a
and/o models
Algo i hms o con ol deg ee o
eedom and gene a e guidelines
Global iew-
Coo dina ing Algo i hms
SDN Con ol
To o he agen s
Se ice Agen
Fo wa ding
Plane
Obse ed da a
and/o models
Figu e 1: Dis ibu ed in elligence a chi ec u e
T a ic, delay,
Pe cen ages
Lea ning Agen
En i onmen
Rewa d
S a e Ac ion
Packe Node Agen
Requi ed
pe cen ages
DRL-based Flow Rou ing
I 1
I 2
I 3
Flow
des ina ion
T a ic
end- o-end
delay
0%
20%
80%
Packe Node
Agen
delay
Flow
sou ce o
in e media e Packe Laye
Op ical Laye
Figu e 2: Example o dis ibu ed low ou ing based on DRL
Domain 1 Domain 2
SDN Con ol
Domain 1
SDN Con ol
Domain 2
dmaxD1
dmaxe2e
dmaxD2
dmaxD1
R1.B
R1.A
R1.C
R2.B
R2.C
R2.Z
Figu e 3: Example o ope a ion unde a ying
dmax in mul i-domain scena ios.
Sandbox domain
DRL engine
Analyze
model
Flow ou ing manage
model
3
upda e(dmax)
5
4
SDN Con ol D1 SDN Con ol D2
dmax
e2e
dmax
D1
dmax
D2
2
1
Figu e 4: Flow ope a ion in mul i-domain scena ios.
ha e been conside ed, o accoun o he ob ained delay ( delay) and o he cos ( cos ), whe e he inal ewa d is
de ined as ollows; pa ame e s αdelay and αcos ep esen he p opo ion o each o hem. Assuming a gi en
maximum delay o be ensu ed o he low (deno ed Dmax), he ewa d ela ed o he ob ained delay can be de ined
in Equa ion 2, whe e β is a ixed penal y o iola ing he maximum delay. Finally, he ewa d ela ed o he cos
o using he ou pu in e aces is ela ed o he pe cen age o a ic sen h ough each o hem, as well as o he a io
cos capaci y o he in e ace shown in Equa ion 3. We assume ha he capaci y o he in e aces, as well as he
cos o each in e ace (which is ela ed o he ou e o he des ina ion o he low), Dmax and β o each low ha e
been ecei ed om he SDN con olle .
=
∙
+
∙
(1)
=
{
−
−
/
>
0
ℎ
!"#!
(2)
=
−
$%
∙
&
'
(!%!)$*!
'
∙
%#
%$($%+
(3)
In o de o p ope ly adap he au onomous ope a ion o mul i-domain scena ios, coo dina ion be ween domains
is implemen ed o sa is y he e2e delay equi emen . Figu e 4 shows he wo k low ha implemen s such
coo dina ion. Fo he sake o simplici y, we assume he scena io in Figu e 3, whe e he low unde s udy c osses
Domain 1 (D1) be o e en e ing e e ence domain D2. Recall ha dmaxe2e need o be ensu ed and hus, au onomous
ope a ion in D2 needs o gua an ee such equi emen , conside ing he delay in he p e ious domain D1.
Then, once ope a ion s a s, he con olle o D1 asynch onously no i ies i s maximum delay dmaxD1 o he
ne wo k con olle o Domain 2 (labeled 1 in Figu e 4), which compu es he equi emen o i s domain segmen as
dmaxD2 = dmaxe2e – dmaxD1 (2). This alue is pushed o he low ou ing manage , ha will wo k o gua an ee such
upda ed dmaxD2 equi emen . In pa icula , he analyze asks o he sandbox domain o upda ing o he new dmax
(3). The sandbox e alua es whe he he cu en model can p ope ly wo k wi h he new QoS equi emen ; o he wise,
e u n a new model (4) o be loaded in o he DRL engine (5). Since he sandbox s o es moni o ed d( ) o a gi en
his o y, i e alua es whe he pas d( ) measu emen s a e below new dmax. I so, no model upda e is necessa y;
o he wise, a new model speci ically ained o such new dmax is loaded. No e ha , in bo h cases, con inuous
online lea ning will imp o e he model by lea ning om i s ou ing decisions.
4. RESULTS
A Py hon-based simula o was implemen ed and ealis ic a ic low beha io was accu a ely emula ed
ollowing a simila con igu a ion as in [3]. As in Figu e 2, we assume ha i 1, i 2, and i 3 ollow di e en ou es in
he mul ilaye ne wo k, so ha h ee di e en delay beha io s a e emula ed. The cos o each in e ace was
con igu ed in e sely p opo ional o he expec ed end- o-end delay h ough ha in e ace. β=6, Dmax = 0.5 ms,
and compa ed h ee di e en op imiza ion a ge s by de ining di e en con igu a ions o he uple (αdelay, αcos ),
namely: (0,1) (cos minimiza ion), (1,0) (delay assu ance), and (1,1) (mul i-objec i e). Figu e 5 illus a es he
o e all pe o mance unde all h ee cases in e ms o maximum delay and cos . I is wo h no ing ha he DRL
lea ned a model p oducing s able and good- ewa ded ac ions in all he scena ios a e 5000 episodes. As seen in
Figu e 5, he i s con igu a ion (0,1) achie es he expec ed bes solu ion in e ms o cos , a he expense o ha ing
a la ge maximum delay. Figu e 6 de ails a ic ou ing and delay one day pos DRL con e gence. In Figu e 6a,
a ic solely u ilizes he cheapes in e ace, esul ing in delays consis en ly exceeding Dmax, comp omising QoS.
Con e sely, con igu a ion (1,0) in Figu e 6b educes maximum delay below Dmax, albei a a no able inc ease in
ne wo k cos due o a o ing he delay-e icien , expensi e in e ace i 1. Howe e , penaliza ion o delay iola ion
main ains delay well below he limi . No ably, con igu a ion (1,1) in Figu e 6 shows he bes pe o mance, wi h
maximum delay below Dmax and a 58% educ ion in ne wo k cos compa ed o (1,0). De ailed analysis in Figu e
6c e eals DRL's abili y o balance ou ing be ween i 2 and i 3, yielding delays consis en ly below Dmax while
minimizing cos s by a oiding expensi e i 1. This demons a es DRL's capaci y o con e ge o di e se solu ions
o he e ogeneous op imiza ion c i e ia. Addi ionally, a second s udy explo es a ious Dmax scena ios anging
om 0.3 o 2 ms unde ixed con igu a ion (1,1).
Fo e alua ing mul i-domain pe o mance ano he simula ion scena io was un ollowing [7], whe e backg ound
a ic is no cons an in ime. a ime 0 and a p e- ained ini ial model assuming cons an backg ound a ic and
a gi en dmaxD2 was used o ope a ion. This model is con inuously imp o ed h ough online lea ning; no e ha i
now needs o lea n he ac ual cha ac e is ics o he inpu a ic and hose o he ime- a ying backg ound a ic.
A e some ime in ope a ion, a ime 1 he model eaches a s able pe o mance ha canno be signi ican ly u he
imp o ed. Then, a ime 2, he D2 SDN con olle ecei es an asynch onous no i ica ion om D1 SDN con olle
upda ing dmaxD1, which in u n igge s upda ing dmaxD2 and consequen ly, he p oposed analysis and model
upda e p ocedu e is ca ied ou . Then, ope a ion con inues wi h he new dmaxD2. Online lea ning migh imp o e
he model in ope a ion, which will each pe o mance s abili y a ime 3. Figu e 7 summa izes he main esul s o
he simula ions in e ms o ou ing cos and delay measu ed a e e y o he abo emen ioned ime ins an s.
Speci ically, Figu e 7a and Figu e 7b show wo cases, whe e dmaxD2 is elaxed ( om 0.5 o 0.75ms and om 0.25
o 0.75ms, espec i ely), whe eas in Figu e 7c and Figu e 7d dmaxD2 becomes mo e s ingen ( om 0.75 o 0.5
and om 0.75 o 0.25, espec i ely).
We obse e ha online lea ning imp o es ini ial p e- ained models e en in he p esence o ime- a ying
backg ound a ic, since ou ing cos is educed in all he cases om 0 o 1, while dmaxD2 is gua an eed in he
whole pe iod [ 0, 1]. In case o Figu e 7a and Figu e 7b, he e was no change in he model in ime 2 because o
dmaxD2 elaxa ion and hence, pe o mance in 2 equals ha o 1. Howe e , a new model was loaded when dmaxD2
was educed in ime 2, and which educed maximum delay o gua an ee he desi ed QoS pe o mance om 2 on,
as obse ed in Figu e 7c and Figu e 7d. Finally, no e ha ega dless he case, he model was imp o ed a e dmaxD2
upda e, by inc easing maximum delay and/o educing ou ing cos .
5. CONCLUSION
This pape summa izes a dis ibu ed app oach le e aging Mul i-Agen Sys ems (MAS) ailo ed speci ically o
op imizing ne wo k pe o mance in eal- ime scena ios, wi h a pa icula emphasis on ensu ing nea - eal- ime end-
o-end delay assu ance ac oss mul iple ne wo k domains. The esul s demons a e he e ec i eness o his app oach
in educing ou ing cos s and main aining desi ed end- o-end delay le els wi hin mul i-domain scena ios. No ably,
he adap abili y o Deep Rein o cemen Lea ning (DRL) models in achie ing he e ogeneous op imiza ion c i e ia.
The inco po a ion o online lea ning mechanisms enhances model pe o mance, e en in he ace o ime- a ying
backg ound a ic. O e all, hese esul s unde sco e he po en ial o dis ibu ed in elligence app oaches in
add essing he e ol ing challenges o nex -gene a ion ne wo ks wi h s ingen QoS equi emen s in mul i-domain
en i onmen s.
ACKNOWLEDGEMENTS
The esea ch leading o hese esul s has ecei ed unding om he Eu opean Union's Ho izon Eu ope esea ch
and inno a ion p og amme SEASON (G.A. 101096120), he MICINN IBON (PID2020-114135RB-I00) and om
he ICREA Ins i u ion.
REFERENCES
[1] D. Ra ique and L. Velasco, “Machine Lea ning o Op ical Ne wo k Au oma ion: O e iew, A chi ec u e and
Applica ions,” (In i ed Tu o ial) IEEE/OSA Jou nal o Op ical Communica ions and Ne wo king (JOCN), 2018.
[2] L. Velasco e al., “Moni o ing and Da a Analy ics o Op ical Ne wo king: Bene i s, A chi ec u es, and Use Cases,” IEEE
Ne wo k Magazine, 2019.
[3] S. Ba zega , M. Ruiz, and L. Velasco, “Packe Flow Capaci y Au onomous Ope a ion based on Rein o cemen Lea ning,”
MDPI Senso s, ol. 21, pp. 8306, 2021.
[4] L. Velasco e al., “Au onomous and Ene gy E icien Ligh pa h Ope a ion based on Digi al Subca ie Mul iplexing,”
IEEE Jou nal on Selec ed A eas in Communica ions, ol. 39, pp. 2864-2877, 2021.
[5] M. Woold idge, An in oduc ion o mul iagen sys ems, John Wiley & Sons, 2009.
[6] S. Ba zega , M. Ruiz and L. Velasco, "Dis ibu ed and Au onomous Flow Rou ing Based on Deep Rein o cemen
Lea ning," OECC/PSC, Japan, 2022
[7] S. Ba zega , M. Ruiz and L. Velasco, "Au onomous Flow Rou ing o Nea Real-Time Quali y o Se ice Assu ance,"
IEEE T ansac ions on Ne wo k and Se ice Managemen (TNSM) 2023
0
2
4
6
8
10
12
14
16
0.0
0.5
1.0
1.5
2.0
(0,1) (1,0) (1,1)
Thousands
cos max delay
D-./=0.5 ms
pa ams (α
delay
, α
cos
)
Maximum delay (ms)
Cos (c.u.)
58%
Figu e 5: O e all DRL pe o mance
0
20
40
60
80
I 1
I 2
I 3
T a ic (Gb/s)
(a) (0, 1) (b) (1, 0) (c) (1, 1)
0.00
0.25
0.50
0.75
1.00
0 4 8 12 16 20 24
Thousands
Delay (ms)
>1 ms
0 4 8 12 16 20 24
Daily hou 0 4 8 12 16 20 24
D-./=0.5 ms
Figu e 6: De ailed pe o mance o con igu a ion (0,1)(a), (1,0)(b), and (1,1)(c)
0
1
2
3
4
5
0
0.25
0.5
0.75
0 1 2 3
0
1
2
3
4
5
0
0.25
0.5
0.75
0 1 2 3
0
1
2
3
4
5
0
0.25
0.5
0.75
0 1 2 3
0
1
2
3
4
5
0
0.25
0.5
0.75
0 1 2 3
cos
maximum delay
Rou ing cos [c.u.]
Delay [
ms
]
0: T a ic low se up 1: Model w/ s able pe .
2: D1 delay upda e 3: Model imp o ed
Time ins an
(a) dmaxD2 0.5->0.75 ms (b) dmaxD2 0.25->0.75 ms (c) dmaxD2 0.75->0.5 ms (d) dmaxD2 0.75->0.25 ms
Figu e 7: Pe o mance e alua ion in mul i-domain scena ios wi h ime- a ying backg ound a ic