scieee Science in your language
[en] (orig)

Energy and precision evaluation of a systolic array accelerator using a quantization approach for edge computing

Abstract

This paper focuses on the implementation of a neural network accelerator optimized for speed and energy efficiency, for use in embedded machine learning. Specifically, we explore power reduction at the hardware level through systolic array and low-precision data systems, including quantized approaches. We present a comprehensive analysis comparing a full precision (FP16) accelerator with a quantized (INT16) version on an FPGA. We upgraded the FP16 modules to handle INT16 values, employing data shifts to enhance value density while maintaining accuracy. Through single convolution experiments, we assess the energy consumption and error minimization. The paper’s structure includes a detailed description of the FP16 accelerator, the transition to quantization, mathematical and implementation insights, instrumentation for power measurement, and a comparative analysis of power consumption and convolution error. Our results attempt to identify a pattern in 16-bit quantization to achieve significant power savings with minimal loss of accuracy.

Read accessible full text

Energy and precision evaluation of a systolic array accelerator using a quantization approach for edge computing

Author: Sanchez Flores, Alejandra,Fornt Mas, Jordi,Álvarez Martí, Lluc,Alorda Ladaria, Bartomeu
Year: 2024
DOI: 10.3390/electronics13142822
Source: https://upcommons.upc.edu/bitstream/2117/418503/1/electronics-13-02822.pdf
Ci a ion: Sanchez-Flo es, A.; Fo n , J.;
Al a ez, L.; Alo da-Lada ia, B.
Ene gy and P ecision E alua ion o a
Sys olic A ay Accele a o Using a
Quan iza ion App oach o Edge
Compu ing. Elec onics 2024,13, 2822.
h ps://doi.o g/10.3390/
elec onics13142822
Academic Edi o : Ja id Tahe i
Recei ed: 11 June 2024
Re ised: 9 July 2024
Accep ed: 15 July 2024
Published: 18 July 2024
Copy igh : © 2024 by he au ho s.
Licensee MDPI, Basel, Swi ze land.
This a icle is an open access a icle
dis ibu ed unde he e ms and
condi ions o he C ea i e Commons
A ibu ion (CC BY) license (h ps://
c ea i ecommons.o g/licenses/by/
4.0/).
elec onics
A icle
Ene gy and P ecision E alua ion o a Sys olic A ay Accele a o
Using a Quan iza ion App oach o Edge Compu ing
Alejand a Sanchez-Flo es 1,*, Jo di Fo n 2, Lluc Al a ez 2,* and Ba omeu Alo da-Lada ia 1,3,4,*
1Depa men o Indus ial Enginee ing and Cons uc ion, Uni e si a de les Illes Balea s Palma,
07122 Palma, Spain
2Ba celona Supe compu ing Cen e , Uni e si a Poli ècnica de Ca alunya Ba celona, 08034 Ba celona, Spain;
[email p o ec ed]
3Balea ic Islands Heal h Resea ch Ins i u e (IdISBa), 07120 Palma, Spain
4Ins i u e o En i onmen al Ag o-En i onmen al Resea ch and Wa e Economics (INAGEA),
07120 Palma, Spain
*
Co espondence: [email p o ec ed] (A.S.-F.); lluc.al a [email p o ec ed] (L.A.); [email p o ec ed] (B.A.-L.)
Abs ac : This pape ocuses on he implemen a ion o a neu al ne wo k accele a o op imized o
speed and ene gy e iciency, o use in embedded machine lea ning. Speci ically, we explo e powe
educ ion a he ha dwa e le el h ough sys olic a ay and low-p ecision da a sys ems, including
quan ized app oaches. We p esen a comp ehensi e analysis compa ing a ull p ecision (FP16) ac-
cele a o wi h a quan ized (INT16) e sion on an FPGA. We upg aded he FP16 modules o handle
INT16 alues, employing da a shi s o enhance alue densi y while main aining accu acy. Th ough
single con olu ion expe imen s, we assess he ene gy consump ion and e o minimiza ion. The
pape ’s s uc u e includes a de ailed desc ip ion o he FP16 accele a o , he ansi ion o quan iza-
ion, ma hema ical and implemen a ion insigh s, ins umen a ion o powe measu emen , and a
compa a i e analysis o powe consump ion and con olu ion e o . Ou esul s a emp o iden i y a
pa e n in 16-bi quan iza ion o achie e signi ican powe sa ings wi h minimal loss o accu acy.
Keywo ds: a ay sys olic accele a o ; ene gy consump ion; embedded sys ems
1. In oduc ion
Au oma ed lea ning s a egies can be used in many a eas; one o hem is he p ocessing
o la ge amoun s o in o ma ion om senso ne wo ks o p o ide accu a e and eliable
da a, while minimizing he ene gy equi ed. The design o neu al ne wo k accele a o s o
pe o m con olu ional ope a ions, op imizing speed and ene gy consump ion, is an ac i e
esea ch ield.
The issue o powe educ ion in embedded machine lea ning is impo an a bo h he
so wa e and ha dwa e le els. A he ha dwa e le el, powe educ ion has been app oached
om di e en angles. In he con ex o accele a o s, he use o sys olic a ays is a pa icula ly
e ec i e app oach, as i allows o he pa allel execu ion o he same ope a ion wi h di e en
inpu da a, he eby gene a ing an e icien ou pu ec o , as exempli ied by he ec o [
1
].
Fu he mo e, among he mos equen ly u ilized and e icacious s a egies a e he use
o di ec memo y access (DMA) memo ies o educe powe , due o hei e icien da a
access [
2
]. Da a access h ough p o ocols such as he Ad anced eX ensible In e ace (AXI)
imp o es da a a ailabili y be ween he cen al p ocessing uni (CPU) and he accele a o .
Quan ized sys ems ha use low-p ecision in ege da a consume less powe han hose ha
use loa ing poin da a [
3
]. Bina ized sys ems, which use a single bi o s o e ML pa ame e s
o pe o m bi -wise ope a ions, consume e en less powe , and some au ho s [
4
–
6
] p opose
a uni o bi -wise ope a ions in he accele a o . In ligh o he indings o ou p e ious
s udies [
7
,
8
] we ha e concluded ha quan iza ion ep esen s one o he mos e ec i e
app oaches a 8 o 16 bi s.
Elec onics 2024,13, 2822. h ps://doi.o g/10.3390/elec onics13142822 h ps://www.mdpi.com/jou nal/elec onics
Elec onics 2024,13, 2822 2 o 13
In his pape , we will ocus on he implemen a ion o a ull-p ecision e sion o he
accele a o and he design and implemen a ion o a quan ized e sion, on an FPGA. The
di e ences in he beha io o some o he basic con olu ions commonly used in machine
lea ning p o ide oppo uni ies o educe powe consump ion by compa ing he esul s o
FP16 and INT16.
Fi s , we use an FP16 accele a o [
9
], which exhibi s o e all op imized and ene gy-
e icien cha ac e is ics. The p oposal is o upg ade he accele a o modules o be used
wi h INT16 alues and o use a da a shi o adjus he esul owa ds he highe densi y
alues. Fu he mo e, we emain cognizan o he impo ance o accu acy. To explain he
op imali y condi ions o he con olu ion ope a ions and he a iables in ol ed, expe i-
men s a e pe o med a he single-con olu ion le el, obse ing he beha io o low ene gy
consump ion and minimum e o alue. The s uc u e o he pape is as ollows: Sec ion 2
desc ibes he FP16 accele a o and he modules in ol ed in he ansi ion o he quan i-
za ion. Sec ion 3is de o ed o he ma hema ical and implemen a ion explana ions o he
quan iza ion.
Sec ion 4
is dedica ed o he ins umen a ion equi ed o compu e powe and
ene gy. Sec ion 5compa es he powe consump ion o he wo e sions o he accele a o
and he e ec o low-p ecision da a on he con olu ion e o . Finally, in Sec ion 6, we
analyze whe he his s a egy has a posi i e impac on ene gy sa ings and discuss whe he
he unques ionable loss o accu acy in he esul s is wo h he ene gy sa ings.
2. Accele a o Desc ip ion
The op imiza ion app oach and e alua ion esul s a e ob ained on an FP16 accele a o
s uc u e epo ed in a p e ious wo k [
9
]—see Figu e 1a. The accele a o sys em consis s
o a sys olic a ay (SA) ha in e changes da a wi h a DMA. The DMA con olle applies
da a o he SA and ecei es he co esponding ou pu . I manages he ac i a ion enso
( enso A), he weigh enso ( enso B), and he p e-load alue enso ( enso C). All alues
o enso s a e FP16.
Elec onics 2024, 13, x FOR PEER REVIEW 2 o 13
indings o ou p e ious s udies [7,8] we ha e concluded ha quan iza ion ep esen s one
o he mos effec i e app oaches a 8 o 16 bi s.
In his pape , we will ocus on he implemen a ion o a ull-p ecision e sion o he
accele a o and he design and implemen a ion o a quan ized e sion, on an FPGA. The
diffe ences in he beha io o some o he basic con olu ions commonly used in machine
lea ning p o ide oppo uni ies o educe powe consump ion by compa ing he esul s o
FP16 and INT16.
Fi s , we use an FP16 accele a o [9], which exhibi s o e all op imized and ene gy-
efficien cha ac e is ics. The p oposal is o upg ade he accele a o modules o be used
wi h INT16 alues and o use a da a shi o adjus he esul owa ds he highe densi y
alues. Fu he mo e, we emain cognizan o he impo ance o accu acy. To explain he
op imali y condi ions o he con olu ion ope a ions and he a iables in ol ed, expe i-
men s a e pe o med a he single-con olu ion le el, obse ing he beha io o low ene gy
consump ion and minimum e o alue. The s uc u e o he pape is as ollows: Sec ion
2 desc ibes he FP16 accele a o and he modules in ol ed in he ansi ion o he quan i-
za ion. Sec ion 3 is de o ed o he ma hema ical and implemen a ion explana ions o he
quan iza ion. Sec ion 4 is dedica ed o he ins umen a ion equi ed o compu e powe
and ene gy. Sec ion 5 compa es he powe consump ion o he wo e sions o he accel-
e a o and he effec o low-p ecision da a on he con olu ion e o . Finally, in Sec ion 6,
we analyze whe he his s a egy has a posi i e impac on ene gy sa ings and discuss
whe he he unques ionable loss o accu acy in he esul s is wo h he ene gy sa ings.
2. Accele a o Desc ip ion
The op imiza ion app oach and e alua ion esul s a e ob ained on an FP16 accele a-
o s uc u e epo ed in a p e ious wo k [9]—see Figu e 1a. The accele a o sys em con-
sis s o a sys olic a ay (SA) ha in e changes da a wi h a DMA. The DMA con olle ap-
plies da a o he SA and ecei es he co esponding ou pu . I manages he ac i a ion en-
so ( enso A), he weigh enso ( enso B), and he p e-load alue enso ( enso C). All
alues o enso s a e FP16.
Figu e 1. High le el desc ip ion o (a) FP16 accele a o , (b) PSM, and (c) sys olic a ay 4 × 8.
The SA is he p ocessing module based on a pa allel scheme—see Figu e 1c. I con-
sis s o an XY a ay o p ocessing elemen s (PEs), o ganized as a ma ix ecei ing inpu
enso s A, B, and C. Each PE indi idually ecei es inpu da a 𝑎, 𝑏, and 𝑐, which is con-
ained in he co esponding enso s. The PE has a loa ing poin (FP) p ocessing uni ha
Figu e 1. High le el desc ip ion o (a) FP16 accele a o , (b) PSM, and (c) sys olic a ay 4 ×8.
The SA is he p ocessing module based on a pa allel scheme—see Figu e 1c. I consis s
o an XY a ay o p ocessing elemen s (PEs), o ganized as a ma ix ecei ing inpu enso s
A,B, and C. Each PE indi idually ecei es inpu da a
ai
,
bi
, and
ci
, which is con ained in
he co esponding enso s. The PE has a loa ing poin (FP) p ocessing uni ha a emp s
o ob ain he mos accu a e alue possible o he mul iplica ion–accumula ion (MAC)
ope a ion o he con olu ion compu a ion.
Elec onics 2024,13, 2822 3 o 13
The inal and mos pe inen module o ou pu poses is he Pa ial Sum Module
(PSM). I s ope a ion is summa ized in Figu e 1b. Inpu enso C, deno ed as
Ci
, co esponds
o he p eload alues passed om one con olu ion o he nex . This enso is s o ed in he
inpu bu e C be o e p ocessing s a s. When he s a o p ocessing is commanded by he
CPU, he MUX1 selec s he con en s o bu e C, which is ans e ed o he shi egis e
submodule SR. The SR empo a ily s o es he C- enso s and hen ans e s hem o he SA
ia MUX2. The MUX2 swi ches om ze o o he SR con en when he e is a alid alue o
enso s. The C bu e eads he con en s o he SR and ac s as he ou pu bu e when all
ope a ions a e comple ed. The da a ype o all egis e s is main ained a a leng h o 16 bi s
h oughou he module.
3. Quan iza ion and Implemen a ion
In he p e ious sec ion, he ope a ions we e pe o med wi h FP16, and all he egis e s
in he PEs and he SMP we e de ined as hey a e. To con e da a o he INT16 o ma , he
quan iza ion p ocess is conduc ed p io o he da a being ans e ed o he accele a o .
Fu he mo e, i is impo an o conside ha he in e ence ask o ou accele a o mus
be able o wo k independen ly om he aining p ocess. This means ha he da a accu acy
emains a ull accu acy du ing aining, and any pa ame e adjus men s a e scheduled o
he in e ence s age. Some au ho s p opose he adjus men o such pa ame e s in a pos -
aining phase, con e ing hem o low-p ecision da a, wi h good esul s o accu acy [
10
–
13
].
Based on his, we p opose he quan iza ion pos - aining using FP16 da a o be pe o med
in he CPU o ob ain quan ized da a o enso Aand enso B, and enso C(i applicable).
All hese a e ed o he accele a o .
3.1. Desc ip ion o Quan iza ion
The p oposed quan iza ion s a egy wo ks wi h uni o mly quan ized and symme -
ic alues, as p oposed in [
13
,
14
]. Thus, i should be assumed a p io i ha he model
da a in FP16 o ma ha e a Gaussian dis ibu ion wi h mean
µ
= 0 and a iance
σ
2=S
(s anda d de ia ion).
In addi ion, i is es ablished ha , ini ially, any da a in he in e al
[−dmax,dmax]
a e o
he loa ing poin ype and hei alue is ounded o a
B10
decimal alue, as ep esen ed by
Equa ion (1), using powe s o 2 [15,16]:
B10 =(β0. . . βn. . .)2=∑
n∈N
βn2n(1)
Since ncan be any posi i e o nega i e in ege alue, Equa ion (1) esul s in a mixed
in ege / ac ions o ma , as shown in Equa ion (2), whe e each βn akes a alue o 1 o 0.
(βn2n)+βn−12n−1+. . . +β121+β020+β−12−1+β−22−2+. . . (2)
I is wo h men ioning ha he numbe s o ac ions gi en by he elemen s wi h a
nega i e powe a e known as dyadic numbe s. Ha dwa e design has used he dyadic
sys em o decades o op imize p ocesso s [
17
–
19
]. In accele a o s, p e ious wo k [
20
]
has implemen ed his s a egy, which ensu es ha he esul s o a i hme ic ope a ions a e
always kep in powe s o 2, speeding up p ocessing.
To explain quan iza ion, le n be a gi en posi i e in ege . The maximum ep esen able
alue is
dmax =(2n−1
and he in e al o numbe s is gi en by
[−(2n−1),(2n−1)]
. This
se con ains only in ege s, since he smalles numbe ha can be ep esen ed is
dmin =
2
0
= 1,
as shown in Figu e 2a.
I we educe he powe n by wo uni s, which co esponds o a shi o he igh , hen
dmax =(2n−2−1
. The in e al is ede ined as
−2n−2−1,2n−2−1
, as is illus a ed
in Figu e 2b, and
dmin =
2
−2=
0.25. Now, he se o da a con ains loa ing poin numbe s.
Using a bina y no a ion, he con e sion can be pe o med by applying a shi o he igh .
Elec onics 2024,13, 2822 4 o 13
Elec onics 2024, 13, x FOR PEER REVIEW 4 o 13
Figu e 2. The g aph ep esen s a ypical dis ibu ion o he da a o diffe en shi ing. (a) shi = 0,
(b) shi = 2 and, (c) shi = 4.
I we educe he powe n by wo uni s, which co esponds o a shi o he igh , hen
𝑑=(2−1) . The in e al is ede ined as 󰇟−(2−1),(2−1)󰇠, as is illus a ed
in Figu e 2b, and 𝑑= 2=0.25. Now, he se o da a con ains loa ing poin num-
be s. Using a bina y no a ion, he con e sion can be pe o med by applying a shi o he
igh .
Roughly speaking, using n bi s, he in e al o numbe s is de ined as 󰇟−(2−
1),(2−1)󰇠, wi h an accu acy o 2. Then, he la ge he bi shi on he igh , he
smalle he minimum alue, allowing o g ea e p ecision and da a densi y. This allows
he da a in e al o be adjus ed o include mo e alues a ound he alue o σ2. Figu e 2c
illus a es he example o m = 4.
Once he in e al and p ecision a e es ablished, he nex s ep is o de e mine he
leng h o egis e s he PEs will ope a e on. To do his, he de ini ion o he MAC ope a ion
is ecalled. I consis s o one mul iplica ion ope a ion and one addi ion ope a ion. The
p oduc o wo loa ing numbe s, 𝑓 and 𝑓, as de ined in Equa ion (1), is gi en by Equa-
ion (3), as ollows:
𝑓
=𝛼⋅2 and
𝑓
=𝛼⋅2 (3)
which is desc ibed in [15]. Then, he p oduc o wo loa ing numbe s is
(
𝑓
⋅
𝑓
)= 𝛼𝛼⋅2 (4)
Fo ou pu poses 𝛼= 𝛼=1. I θ and η a e 16, hen he mul iplica ion esul will
be a maximum o 32 bi s in iew o he maximum powe o he inpu da a. In conclusion,
he ou pu egis e o he MAC is 34 bi s when he ca y de i ed om he addi ion pa
and he sign bi a e added.
In he emainde o his sec ion, he le -shi ope a ion applied o he da a will be
e e ed o as he quan iza ion, and i is illus a ed in Figu e 3a. The in e se ope a ion o
igh -shi ing o he da a will be e e ed o as dequan iza ion, as shown in Figu e 3b.
Figu e 3. (a) Quan iza ion ope a ion applied o he ou pu egis e o each PE. (b) Dequan iza ion
ope a ion applied o he inpu egis e s o each PE.
Figu e 2. The g aph ep esen s a ypical dis ibu ion o he da a o di e en shi ing. (a) shi = 0,
(b) shi = 2 and, (c) shi = 4.
Roughly speaking, using n bi s, he in e al o numbe s is de ined as
[−(2n−m−1),
(2n−m−1)]
, wi h an accu acy o 2
−m
. Then, he la ge he bi shi on he igh , he smalle
he minimum alue, allowing o g ea e p ecision and da a densi y. This allows he da a
in e al o be adjus ed o include mo e alues a ound he alue o
σ
2. Figu e 2c illus a es
he example o m = 4.
Once he in e al and p ecision a e es ablished, he nex s ep is o de e mine he leng h
o egis e s he PEs will ope a e on. To do his, he de ini ion o he MAC ope a ion is
ecalled. I consis s o one mul iplica ion ope a ion and one addi ion ope a ion. The p oduc
o wo loa ing numbe s,
1
and
2
, as de ined in Equa ion (1), is gi en by Equa ion (3),
as ollows:
1=αθ·2θand 2=αη·2η(3)
which is desc ibed in [15]. Then, he p oduc o wo loa ing numbe s is
( 1· 2)=αθαη·2θ+η(4)
Fo ou pu poses
αθ=αη=
1. I
θ
and
η
a e 16, hen he mul iplica ion esul will be
a maximum o 32 bi s in iew o he maximum powe o he inpu da a. In conclusion, he
ou pu egis e o he MAC is 34 bi s when he ca y de i ed om he addi ion pa and
he sign bi a e added.
In he emainde o his sec ion, he le -shi ope a ion applied o he da a will be
e e ed o as he quan iza ion, and i is illus a ed in Figu e 3a. The in e se ope a ion o
igh -shi ing o he da a will be e e ed o as dequan iza ion, as shown in Figu e 3b.
Elec onics 2024, 13, x FOR PEER REVIEW 4 o 13
Figu e 2. The g aph ep esen s a ypical dis ibu ion o he da a o diffe en shi ing. (a) shi = 0,
(b) shi = 2 and, (c) shi = 4.
I we educe he powe n by wo uni s, which co esponds o a shi o he igh , hen
𝑑=(2−1) . The in e al is ede ined as 󰇟−(2−1),(2−1)󰇠, as is illus a ed
in Figu e 2b, and 𝑑= 2=0.25. Now, he se o da a con ains loa ing poin num-
be s. Using a bina y no a ion, he con e sion can be pe o med by applying a shi o he
igh .
Roughly speaking, using n bi s, he in e al o numbe s is de ined as 󰇟−(2−
1),(2−1)󰇠, wi h an accu acy o 2. Then, he la ge he bi shi on he igh , he
smalle he minimum alue, allowing o g ea e p ecision and da a densi y. This allows
he da a in e al o be adjus ed o include mo e alues a ound he alue o σ2. Figu e 2c
illus a es he example o m = 4.
Once he in e al and p ecision a e es ablished, he nex s ep is o de e mine he
leng h o egis e s he PEs will ope a e on. To do his, he de ini ion o he MAC ope a ion
is ecalled. I consis s o one mul iplica ion ope a ion and one addi ion ope a ion. The
p oduc o wo loa ing numbe s, 𝑓 and 𝑓, as de ined in Equa ion (1), is gi en by Equa-
ion (3), as ollows:
𝑓
=𝛼⋅2 and
𝑓
=𝛼⋅2 (3)
which is desc ibed in [15]. Then, he p oduc o wo loa ing numbe s is
(
𝑓
⋅
𝑓
)= 𝛼𝛼⋅2 (4)
Fo ou pu poses 𝛼= 𝛼=1. I θ and η a e 16, hen he mul iplica ion esul will
be a maximum o 32 bi s in iew o he maximum powe o he inpu da a. In conclusion,
he ou pu egis e o he MAC is 34 bi s when he ca y de i ed om he addi ion pa
and he sign bi a e added.
In he emainde o his sec ion, he le -shi ope a ion applied o he da a will be
e e ed o as he quan iza ion, and i is illus a ed in Figu e 3a. The in e se ope a ion o
igh -shi ing o he da a will be e e ed o as dequan iza ion, as shown in Figu e 3b.
Figu e 3. (a) Quan iza ion ope a ion applied o he ou pu egis e o each PE. (b) Dequan iza ion
ope a ion applied o he inpu egis e s o each PE.
Figu e 3. (a) Quan iza ion ope a ion applied o he ou pu egis e o each PE. (b) Dequan iza ion
ope a ion applied o he inpu egis e s o each PE.
I he dequan iza ion lea es MSB wi hou alues, hey a e illed wi h 1’s i he numbe
o dequan ize is nega i e, and illed wi h 0’s i he alue is posi i e. By shi ing he 16 bi s
o a highe alue posi ion, he leas signi ican bi s a e illed wi h 0’s.
To conclude his sec ion, i should be emphasized ha he selec ion o he bes o se -
shi o each laye is applied du ing he con olu ion p ocess. The HW o he accele a o
mus be designed o handle his.
Elec onics 2024,13, 2822 5 o 13
3.2. Implemen a ion on FPGA
The ollowing explains he changes equi ed o upda e he accele a o and wo k
wi h quan ized alues. The changes a e numbe ed acco ding o he o de in which hey
we e implemen ed.
1.
The da a
IAw
,
IBw
, and
ICw
a e he leng h o he inpu da a,
ai
(ac i a ion da a),
bi
(weigh s), and ci(p eload alues), and a e all upda ed o INT16.
2. Each PE is modi ied o use only in ege a i hme ic, as shown in Figu e 4a.
3.
The PSM, p e iously shown in Figu e 1b, is upda ed, as shown in Figu e 4b. The main
change is he modi ica ion o he ou pu o he PSM, he enso
Co
. As explained in
he las sec ion, he new leng h o OCwis 34 bi s.
4. Co
is ob ained a e he MAC ope a ion by using he da a
ai
,
bi
, and
ci
, which a e
g ouped in o he enso s A,B, and C, espec i ely. Fo each Y ow o he SA ma ix,
a enso
Co
is ou pu , hen a se o
Y∗(OCw−1)
da a (o Ynumbe o Co enso s)
a e ou pu . Since he leng h o
OCw
does no ma ch he
ICw
leng h (16 bi s), a
quan iza ion is implemen ed, as is shown in he le side o Figu e 4b. A new con ol
signal SEL_SHIFT is u ilized o selec he o se .
5.
When he da a lea es he SR block, he da a leng h is 16 bi s, and is subjec ed o
a dequan iza ion p ocess o, again, main ain consis ency wi h he leng h o he SA
egis e s. See he igh side o Figu e 4b.
Elec onics 2024, 13, x FOR PEER REVIEW 5 o 13
I he dequan iza ion lea es MSB wi hou alues, hey a e illed wi h 1’s i he numbe
o dequan ize is nega i e, and illed wi h 0’s i he alue is posi i e. By shi ing he 16 bi s
o a highe alue posi ion, he leas signi ican bi s a e illed wi h 0’s.
To conclude his sec ion, i should be emphasized ha he selec ion o he bes offse -
shi o each laye is applied du ing he con olu ion p ocess. The HW o he accele a o
mus be designed o handle his.
3.2. Implemen a ion on FPGA
The ollowing explains he changes equi ed o upda e he accele a o and wo k wi h
quan ized alues. The changes a e numbe ed acco ding o he o de in which hey we e
implemen ed.
1. The da a 𝐼𝐴, 𝐼𝐵, and 𝐼𝐶 a e he leng h o he inpu da a, 𝑎 (ac i a ion da a), 𝑏
(weigh s), and 𝑐 (p eload alues), and a e all upda ed o INT16.
2. Each PE is modi ied o use only in ege a i hme ic, as shown in Figu e 4a.
3. The PSM, p e iously shown in Figu e 1b, is upda ed, as shown in Figu e 4b. The
main change is he modi ica ion o he ou pu o he PSM, he enso 𝐶. As explained
in he las sec ion, he new leng h o 𝑂𝐶 is 34 bi s.
4. 𝐶 is ob ained a e he MAC ope a ion by using he da a 𝑎, 𝑏, and 𝑐, which a e
g ouped in o he enso s A, B, and C, espec i ely. Fo each Y ow o he SA ma ix,
a enso 𝐶 is ou pu , hen a se o 𝑌∗(𝑂𝐶−1) da a (o Y numbe o Co enso s)
a e ou pu . Since he leng h o 𝑂𝐶 does no ma ch he 𝐼𝐶 leng h (16 bi s), a quan-
iza ion is implemen ed, as is shown in he le side o Figu e 4b. A new con ol signal
SEL_SHIFT is u ilized o selec he offse .
5. When he da a lea es he SR block, he da a leng h is 16 bi s, and is subjec ed o a
dequan iza ion p ocess o, again, main ain consis ency wi h he leng h o he SA eg-
is e s. See he igh side o Figu e 4b.
Figu e 4. (a) PE e sion INT16. (b) PSM modi ied o include educing and expanding sub-mod-
ules.
The quan iza ion submodule is a se o Y numbe s o mul iplexe s applied o each
da a Co. Figu e 5 illus a es he p ocess indi idually o a single da a elemen . The da a
wi h leng h 𝑂𝐶 en e s he module whe e he selec ion offse is applied o he MUX.
Once educed, a leng h enso 𝐼𝐶 is passed o he shi egis e . The alue o SEL_SHIFT
is sha ed by all da a. See le side o Figu e 5.
Figu e 4. (a) PE e sion INT16. (b) PSM modi ied o include educing and expanding sub-modules.
The quan iza ion submodule is a se o Y numbe s o mul iplexe s applied o each
da a Co. Figu e 5illus a es he p ocess indi idually o a single da a elemen . The da a
wi h leng h
OCw
en e s he module whe e he selec ion o se is applied o he MUX. Once
educed, a leng h enso
ICw
is passed o he shi egis e . The alue o SEL_SHIFT is
sha ed by all da a. See le side o Figu e 5.
Elec onics 2024, 13, x FOR PEER REVIEW 6 o 13
Figu e 5. The quan iza ion and dequan iza ion submodules a e implemen ed as a MUX o he o -
me and a DEMUX o he la e . No e ha he SR egis e s a e shi ed om posi ion 0 o X.
The dequan iza ion submodule comp ises a se o Y demul iplexe s, each dedica ed
o a single da a elemen . A he ou pu o he shi egis e , a da a elemen o leng h 𝐼𝐶
en e s he DEMUX, which has a SEL_SHIFT inpu o he de e mina ion o he offse o
he dequan iza ion o he da a o 𝑂𝐶. This offse is sha ed by all da a elemen s lea ing
he SR. See igh side o Figu e 5.
4. Ins umen a ion and Measu emen Me hodology
A e a e iew o he quan iza ion li e a u e, we ound ha he powe o ene gy con-
sump ion alues we e no s anda dized. Some au ho s p esen powe in uni s o mW
[21,22] o W [23], while o he s p esen ene gy in uni s o µJ [24,25]. E en he ene gy effi-
ciency alues a e p esen ed in diffe en uni s, such as pJ/op [26] o MAC/W [27]. The issue
a hand is no he uni in and o i sel . Ra he , i is he lack o speci ica ion as o he me h-
odology and ins umen a ion used in he measu emen p ocess ha gi es ise o ques ions
as o he c i e ia used o e alua e he powe and ene gy consump ion, as well as he ene gy
efficiency.
This sec ion is o ganized as ollows: he i s subsec ion desc ibes he ea u es o he
implemen a ion boa d, he second subsec ion gi es an o e iew o he sys em imple-
men ed in Vi ado, and he las sec ion is de o ed o explaining he es cases.
4.1. Implemen a ion Boa d P ope ies
The quan ized accele a o is in ended o use in embedded applica ions. We ha e
es ablished ha an FPGA is necessa y o his esea ch. Among he FPGAs used o em-
bedded applica ions, we ind he Pynq se ies, which offe s design suppo in Vi ado Xil-
inx, and a Debian Linux ope a ing sys em ha allows unning p og ams using Jupy e
No ebook. Speci ically, he Ul a96- 2 boa d was selec ed as he op imal choice. In addi-
ion, use s can ob ain elec ical in o ma ion om he Ul a96- 2 boa d ia PMBus com-
munica ion using In ineon’s USB005, a USB dongle.
Measu emen s can be made on he FPGA since he Pynq boa d is elec ically di ided
in o wo sec ions—a p ocesso sys em called PS and an FPGA sec ion called PL, as shown
in Figu e 6a. The PL ol age exhibi s a mean alue o 0.85 V wi h a noise le el o 0.016
Vpp. The p ocessing consump ion is mo e clea ly e lec ed in he cu en signal, which
has a s andby alue o 85 mA. Bo h signals a e eco ded and used o each powe calcu-
la ion. The ma k signal is con igu ed by he use o p o ide ime e e ences. See Figu e
6b.
Figu e 5. The quan iza ion and dequan iza ion submodules a e implemen ed as a MUX o he
o me and a DEMUX o he la e . No e ha he SR egis e s a e shi ed om posi ion 0 o X.

Elec onics 2024,13, 2822 6 o 13
The dequan iza ion submodule comp ises a se o Y demul iplexe s, each dedica ed
o a single da a elemen . A he ou pu o he shi egis e , a da a elemen o leng h
ICw
en e s he DEMUX, which has a SEL_SHIFT inpu o he de e mina ion o he o se o
he dequan iza ion o he da a o
OCw
. This o se is sha ed by all da a elemen s lea ing
he SR. See igh side o Figu e 5.
4. Ins umen a ion and Measu emen Me hodology
A e a e iew o he quan iza ion li e a u e, we ound ha he powe o ene gy
consump ion alues we e no s anda dized. Some au ho s p esen powe in uni s o
mW [
21
,
22
]o W[
23
], while o he s p esen ene gy in uni s o
µ
J [
24
,
25
]. E en he ene gy
e iciency alues a e p esen ed in di e en uni s, such as pJ/op [
26
] o MAC/W [
27
]. The
issue a hand is no he uni in and o i sel . Ra he , i is he lack o speci ica ion as o
he me hodology and ins umen a ion used in he measu emen p ocess ha gi es ise o
ques ions as o he c i e ia used o e alua e he powe and ene gy consump ion, as well as
he ene gy e iciency.
This sec ion is o ganized as ollows: he i s subsec ion desc ibes he ea u es o he
implemen a ion boa d, he second subsec ion gi es an o e iew o he sys em implemen ed
in Vi ado, and he las sec ion is de o ed o explaining he es cases.
4.1. Implemen a ion Boa d P ope ies
The quan ized accele a o is in ended o use in embedded applica ions. We ha e
es ablished ha an FPGA is necessa y o his esea ch. Among he FPGAs used o embed-
ded applica ions, we ind he Pynq se ies, which o e s design suppo in Vi ado Xilinx, and
a Debian Linux ope a ing sys em ha allows unning p og ams using Jupy e No ebook.
Speci ically, he Ul a96- 2 boa d was selec ed as he op imal choice. In addi ion, use s can
ob ain elec ical in o ma ion om he Ul a96- 2 boa d ia PMBus communica ion using
In ineon’s USB005, a USB dongle.
Measu emen s can be made on he FPGA since he Pynq boa d is elec ically di ided
in o wo sec ions—a p ocesso sys em called PS and an FPGA sec ion called PL, as shown
in Figu e 6a. The PL ol age exhibi s a mean alue o 0.85 V wi h a noise le el o 0.016 Vpp.
The p ocessing consump ion is mo e clea ly e lec ed in he cu en signal, which has a
s andby alue o 85 mA. Bo h signals a e eco ded and used o each powe calcula ion.
The ma k signal is con igu ed by he use o p o ide ime e e ences. See Figu e 6b.
Elec onics 2024, 13, x FOR PEER REVIEW 7 o 13
Figu e 6. (a) A simpli ied diag am o he Ul a96 2 boa d, di ided in o PS p ocessing sec ions and
PL logic pa s. (b) An example o plo ing made in Jupy e no ebook.
4.2. High-Le el Sys em
The accele a ion sys em desc ibed in Figu e 7 consis s o a CPU p o ided by he Zynq
Ul ascale+ mic op ocesso and he Sau ia subsys em accele a o .
Figu e 7. A comp ehensi e desc ip ion o he sys em implemen ed in Vi ado.
The CPU and he accele a o u ilize he AXI o communica ion be ween hem. The
AXI Li e is employed o ansmi 32-bi con ol and con igu a ion commands, wi h he
CPU ac ing as he mas e and he accele a o as he sla e. The ull 128-bi AXI is employed
as a channel o da a ansmission, wi h he accele a o ac ing as he mas e o he channel
and he CPU ac ing as he sla e. The AXI in e connec blocks a e p o ided by he Pynq
amewo k, and hey a e esponsible o managing da a la ency, synch oniza ion, and a -
bi a ion. Simila ly, he ese and clock signal connec ions a e handled by he amewo k.
The clock equency o he en i e sys em has been se o 25 MHz.
4.3. Ins umen a ion Desc ip ion
To pe o m he equi ed measu emen s o he elec ical signals, he ollowing con-
nec ions mus be es ablished (see Figu e 8 o an illus a ion o he necessa y wi ing):
Figu e 8. Connec ions be ween measu ing elemen s.
PMBus [28] is a 400 KHz I2C ha sends in o ma ion om he Ul a96- 2 boa d’s ol -
age egula o . The pynq.pmbus.Da aReco de class o Py hon p o ides an in e ace o ob-
ain he ol age, cu en , iming, and o he signals om he PL pa , which a e sen o he
hos .
Figu e 6. (a) A simpli ied diag am o he Ul a96 2 boa d, di ided in o PS p ocessing sec ions and
PL logic pa s. (b) An example o plo ing made in Jupy e no ebook.
4.2. High-Le el Sys em
The accele a ion sys em desc ibed in Figu e 7consis s o a CPU p o ided by he Zynq
Ul ascale+ mic op ocesso and he Sau ia subsys em accele a o .
The CPU and he accele a o u ilize he AXI o communica ion be ween hem. The
AXI Li e is employed o ansmi 32-bi con ol and con igu a ion commands, wi h he CPU
ac ing as he mas e and he accele a o as he sla e. The ull 128-bi AXI is employed as
a channel o da a ansmission, wi h he accele a o ac ing as he mas e o he channel
and he CPU ac ing as he sla e. The AXI in e connec blocks a e p o ided by he Pynq
amewo k, and hey a e esponsible o managing da a la ency, synch oniza ion, and
a bi a ion. Simila ly, he ese and clock signal connec ions a e handled by he amewo k.
The clock equency o he en i e sys em has been se o 25 MHz.
Elec onics 2024,13, 2822 7 o 13
Elec onics 2024, 13, x FOR PEER REVIEW 7 o 13
Figu e 6. (a) A simpli ied diag am o he Ul a96 2 boa d, di ided in o PS p ocessing sec ions and
PL logic pa s. (b) An example o plo ing made in Jupy e no ebook.
4.2. High-Le el Sys em
The accele a ion sys em desc ibed in Figu e 7 consis s o a CPU p o ided by he Zynq
Ul ascale+ mic op ocesso and he Sau ia subsys em accele a o .
Figu e 7. A comp ehensi e desc ip ion o he sys em implemen ed in Vi ado.
The CPU and he accele a o u ilize he AXI o communica ion be ween hem. The
AXI Li e is employed o ansmi 32-bi con ol and con igu a ion commands, wi h he
CPU ac ing as he mas e and he accele a o as he sla e. The ull 128-bi AXI is employed
as a channel o da a ansmission, wi h he accele a o ac ing as he mas e o he channel
and he CPU ac ing as he sla e. The AXI in e connec blocks a e p o ided by he Pynq
amewo k, and hey a e esponsible o managing da a la ency, synch oniza ion, and a -
bi a ion. Simila ly, he ese and clock signal connec ions a e handled by he amewo k.
The clock equency o he en i e sys em has been se o 25 MHz.
4.3. Ins umen a ion Desc ip ion
To pe o m he equi ed measu emen s o he elec ical signals, he ollowing con-
nec ions mus be es ablished (see Figu e 8 o an illus a ion o he necessa y wi ing):
Figu e 8. Connec ions be ween measu ing elemen s.
PMBus [28] is a 400 KHz I2C ha sends in o ma ion om he Ul a96- 2 boa d’s ol -
age egula o . The pynq.pmbus.Da aReco de class o Py hon p o ides an in e ace o ob-
ain he ol age, cu en , iming, and o he signals om he PL pa , which a e sen o he
hos .
Figu e 7. A comp ehensi e desc ip ion o he sys em implemen ed in Vi ado.
4.3. Ins umen a ion Desc ip ion
To pe o m he equi ed measu emen s o he elec ical signals, he ollowing connec-
ions mus be es ablished (see Figu e 8 o an illus a ion o he necessa y wi ing):
Elec onics 2024, 13, x FOR PEER REVIEW 7 o 13
Figu e 6. (a) A simpli ied diag am o he Ul a96 2 boa d, di ided in o PS p ocessing sec ions and
PL logic pa s. (b) An example o plo ing made in Jupy e no ebook.
4.2. High-Le el Sys em
The accele a ion sys em desc ibed in Figu e 7 consis s o a CPU p o ided by he Zynq
Ul ascale+ mic op ocesso and he Sau ia subsys em accele a o .
Figu e 7. A comp ehensi e desc ip ion o he sys em implemen ed in Vi ado.
The CPU and he accele a o u ilize he AXI o communica ion be ween hem. The
AXI Li e is employed o ansmi 32-bi con ol and con igu a ion commands, wi h he
CPU ac ing as he mas e and he accele a o as he sla e. The ull 128-bi AXI is employed
as a channel o da a ansmission, wi h he accele a o ac ing as he mas e o he channel
and he CPU ac ing as he sla e. The AXI in e connec blocks a e p o ided by he Pynq
amewo k, and hey a e esponsible o managing da a la ency, synch oniza ion, and a -
bi a ion. Simila ly, he ese and clock signal connec ions a e handled by he amewo k.
The clock equency o he en i e sys em has been se o 25 MHz.
4.3. Ins umen a ion Desc ip ion
To pe o m he equi ed measu emen s o he elec ical signals, he ollowing con-
nec ions mus be es ablished (see Figu e 8 o an illus a ion o he necessa y wi ing):
Figu e 8. Connec ions be ween measu ing elemen s.
PMBus [28] is a 400 KHz I2C ha sends in o ma ion om he Ul a96- 2 boa d’s ol -
age egula o . The pynq.pmbus.Da aReco de class o Py hon p o ides an in e ace o ob-
ain he ol age, cu en , iming, and o he signals om he PL pa , which a e sen o he
hos .
Figu e 8. Connec ions be ween measu ing elemen s.
PMBus [
28
] is a 400 KHz I2C ha sends in o ma ion om he Ul a96- 2 boa d’s
ol age egula o . The pynq.pmbus.Da aReco de class o Py hon p o ides an in e ace o
ob ain he ol age, cu en , iming, and o he signals om he PL pa , which a e sen o
he hos .
4.4. Tes s Desc ip ion
The pu pose o he con olu ion es s is o examine he esponse o he accele a o
unde di e en scena ios. The block diag am in Figu e 9shows he gene al es sequence.
Elec onics 2024, 13, x FOR PEER REVIEW 8 o 13
4.4. Tes s Desc ip ion
The pu pose o he con olu ion es s is o examine he esponse o he accele a o
unde diffe en scena ios. The block diag am in Figu e 9 shows he gene al es sequence.
Figu e 9. P ocess desc ip ion o one con olu ion calcula ion.
The con igu a ion phase consis s o se ing he HW cha ac e is ics ha mus ma ch
he RTL implemen a ion o he accele a o , de ined as hype pa ame e s (see Table 1).
These pa ame e s include he dimensions o he sys olic a ay o he accele a o (X and Y),
he inpu da a ype, and he BRAM wid h.
Table 1. Hype pa ame e s o he accele a o .
Hype pa ame e Value
A ay Shape (X, Y) 8 × 4
A i hme ic in 16
Ze o Ga ing Mul + Add
BRAMA wid h 128
BRAMB wid h 128
BRAMC wid h 128
BRAM dep h (all) 2048
In he ini ializa ion phase, he AXI communica ion channels a e ini ialized; he pa-
ame e s ela ing o he ype o con olu ion a e sen o he accele a o ; and he ac i a ion
enso s, weigh s, and p eload a e w i en o he DMA.
The s a -p ocessing block ep esen s he ime ha he accele a o pe o ms a con o-
lu ion, om he ime he CPU sends he s a command un il he accele a o sends he
comple ion esponse.
Pos -p ocessing asks include eading he enso esul ing om he con olu ion and
swi ching he double ou pu buffe . The epo ing s age indica es he end o he con olu-
ion e alua ion by gi ing in e nal accele a o in o ma ion.
Tes s ha e epe i ion cycles—see Figu e 9. These cycles a e o de ed by execu ion le -
els. The i s one, in he blue line, indica es he loop in which he N es s (N = 11) a e
execu ed o e alua e he accele a o esponse unde diffe en con olu ion condi ions.
The pu ple line ep esen s he con olu ion epe i ions loop (R es s = 10,000), which
se es o s abilize he measu ed elec ical signals. Only one con olu ion is p ocessed in
nanoseconds, which is no enough o keep he cu en and ol age a a measu able alue.
In pa icula , one con olu ion execu ion implies all he MAC ope a ions, pa ial sums,
and da a ans e s o memo y. Du ing he CNN in e ence, p ocessing he con olu ion is
execu ed on nume ous occasions, po en ially a hund ed o mo e imes.
The g een line indica es he wai loop while he con olu ion is being comple ed. This
is ca ied ou ins ead o implemen ing an in e up signal which, when es ed, inc eased
he ime be ween es s.
4.5. Con olu ion Fea u es o E e y Tes
We conduc ed he e alua ion using 11 es cases, as shown in Table 2. To es he
sys em unde diffe en condi ions, each es is designed wi h diffe en con olu ion pa-
ame e s. The in en ion is o explo e diffe en con olu ion ea u es.
Figu e 9. P ocess desc ip ion o one con olu ion calcula ion.
The con igu a ion phase consis s o se ing he HW cha ac e is ics ha mus ma ch
he RTL implemen a ion o he accele a o , de ined as hype pa ame e s (see Table 1). These
pa ame e s include he dimensions o he sys olic a ay o he accele a o (X and Y), he
inpu da a ype, and he BRAM wid h.
Table 1. Hype pa ame e s o he accele a o .
Hype pa ame e Value
A ay Shape (X, Y) 8 ×4
A i hme ic in 16
Ze o Ga ing Mul + Add
BRAMA wid h 128
BRAMB wid h 128
BRAMC wid h 128
BRAM dep h (all) 2048
Elec onics 2024,13, 2822 8 o 13
In he ini ializa ion phase, he AXI communica ion channels a e ini ialized; he pa-
ame e s ela ing o he ype o con olu ion a e sen o he accele a o ; and he ac i a ion
enso s, weigh s, and p eload a e w i en o he DMA.
The s a -p ocessing block ep esen s he ime ha he accele a o pe o ms a con o-
lu ion, om he ime he CPU sends he s a command un il he accele a o sends he
comple ion esponse.
Pos -p ocessing asks include eading he enso esul ing om he con olu ion and
swi ching he double ou pu bu e . The epo ing s age indica es he end o he con olu ion
e alua ion by gi ing in e nal accele a o in o ma ion.
Tes s ha e epe i ion cycles—see Figu e 9. These cycles a e o de ed by execu ion le els.
The i s one, in he blue line, indica es he loop in which he N es s (N = 11) a e execu ed
o e alua e he accele a o esponse unde di e en con olu ion condi ions.
The pu ple line ep esen s he con olu ion epe i ions loop (R
es s
= 10,000), which
se es o s abilize he measu ed elec ical signals. Only one con olu ion is p ocessed in
nanoseconds, which is no enough o keep he cu en and ol age a a measu able alue.
In pa icula , one con olu ion execu ion implies all he MAC ope a ions, pa ial sums,
and da a ans e s o memo y. Du ing he CNN in e ence, p ocessing he con olu ion is
execu ed on nume ous occasions, po en ially a hund ed o mo e imes.
The g een line indica es he wai loop while he con olu ion is being comple ed. This
is ca ied ou ins ead o implemen ing an in e up signal which, when es ed, inc eased
he ime be ween es s.
4.5. Con olu ion Fea u es o E e y Tes
We conduc ed he e alua ion using 11 es cases, as shown in Table 2. To es he sys em
unde di e en condi ions, each es is designed wi h di e en con olu ion pa ame e s.
The in en ion is o explo e di e en con olu ion ea u es.
Table 2. Fea u es o he 11 es s used in his wo k o measu e powe and ene gy.
Tes Numbe Memo y Requi emen s
Ac i a ions, Weigh s, P eload/Ou pu
Con olu ion Fea u es. Dimensions
o Tenso s: Ac i a ion (A), Weigh s (B),
P eload/Ou pu (C)
0 448, 1728, 288 A (16, 4, 14), B (24, 16, 3, 3), C (24, 2, 12)
1 283, 1176, 128 A (3, 9, 21), B (16, 3, 7, 7), C (16, 2, 8)
2 1012, 2400, 128 A (3, 15, 45), B (16, 3, 10, 10), C (16, 2, 8)
3 55, 1332, 12 A (111, 1, 1), B (24, 111, 1, 1), C (24, 1, 1)
4 108, 216, 128 A (3, 6, 12), B (16, 3, 3, 3), C (16, 2, 8)
5 5920, 288, 16 A (8, 37, 40), B (8, 8, 3, 3), C (8, 1, 4)
6 420, 600, 128 A (3, 14, 20), B (16, 3, 5, 5), C (16, 2, 8)
7 36, 36, 288 A (3, 2, 12), B (24, 3, 1, 1), C (24, 2, 12)
8 1664, 2496, 192 A (208, 2, 8), B (24, 208, 1, 1), C (24, 2, 8)
9 84, 324, 288 A (3, 4, 14), B (24, 3, 3, 3), C (24, 2, 12)
10 36, 40, 12 A (3, 4, 6), B (3, 3, 3, 3), C (3, 2, 4)
4.6. Ene gy Calcula ion
Once we ha e es ablished all he in o ma ion abou he sys em, he measu emen
signals, and he es s, we p oceed o ob ain he powe and ene gy esul s.
Knowing ha a single con olu ion is ca ied ou in a ime ha is no su icien o mea-
su emen s (app ox. 68 ms a e age sampling ime), we epea he con olu ion calcula ion
block o Figu e 9up o 10,000 imes o ob ain a measu emen o ol age and cu en o
calcula ing he powe alue in one ins an ime.
Elec onics 2024,13, 2822 9 o 13
On he o he hand, in o de o accu a ely iden i y he cu en co esponding o he
ac i i y o he accele a o , we no ice di e en cu en le els in he measu emen s (Figu e 10).
In he i s one, when he boa d is u ned on and be o e he s a o he es , we ha e a
s eady s a e wi h some noise. Du ing his pe iod o inac i i y,
Tnop oc
, he mean PL cu en ,
is calcula ed and s o ed as a cons an
Imean
. In he second, du ing he execu ion o he
con olu ion (
Tp oc
),
Ip oc
ep esen s he cu en consumed by he p ocess. Thus, o ob ain
Ip oc
, we ake he di e ence be ween each ins an aneous measu ed cu en I
INT
and he
p ep ocessing cu en Imean.
Ip oc(i)=IINT(i)−Imean(i)(5)
The powe calcula ion is pe o med o each sample i:
P(i)=VINT(i)∗Ip oc(i)(6)
The ene gy consumed by he accele a o du ing he p ocessing ime is ob ained by:
EINT =
∑
Tp oc
P(i)∗∆ (i)
(7)
whe e
∆ (i)
is he ime inc emen du ing
Tp oc
, co esponding o each sample, and has he
median alue o 60 ms
±
15 ms. The i egula sample iming allows some alea o y peaks o
be de ec ed.
EINT
is he ene gy calcula ed o he INT e sion.
EFP
is he ene gy calcula ed
o he FP e sion.
Elec onics 2024, 13, x FOR PEER REVIEW 9 o 13
Table 2. Fea u es o he 11 es s used in his wo k o measu e powe and ene gy.
Tes Numbe Memo y Requi emen s
Ac i a ions, Weigh s, P eload/Ou pu
Con olu ion Fea u es. Dimensions
o Tenso s: Ac i a ion (A), Weigh s (B),
P eload/Ou pu (C)
0 448, 1728, 288 A (16, 4, 14), B (24, 16, 3, 3), C (24, 2, 12)
1 283, 1176, 128 A (3, 9, 21), B (16, 3, 7, 7), C (16, 2, 8)
2 1012, 2400, 128 A (3, 15, 45), B (16, 3, 10, 10), C (16, 2, 8)
3 55, 1332, 12 A (111, 1, 1), B (24, 111, 1, 1), C (24, 1, 1)
4 108, 216, 128 A (3, 6, 12), B (16, 3, 3, 3), C (16, 2, 8)
5 5920, 288, 16 A (8, 37, 40), B (8, 8, 3, 3), C (8, 1, 4)
6 420, 600, 128 A (3, 14, 20), B (16, 3, 5, 5), C (16, 2, 8)
7 36, 36, 288 A (3, 2, 12), B (24, 3, 1, 1), C (24, 2, 12)
8 1664, 2496, 192 A (208, 2, 8), B (24, 208, 1, 1), C (24, 2, 8)
9 84, 324, 288 A (3, 4, 14), B (24, 3, 3, 3), C (24, 2, 12)
10 36, 40, 12 A (3, 4, 6), B (3, 3, 3, 3), C (3, 2, 4)
4.6. Ene gy Calcula ion
Once we ha e es ablished all he in o ma ion abou he sys em, he measu emen
signals, and he es s, we p oceed o ob ain he powe and ene gy esul s.
Knowing ha a single con olu ion is ca ied ou in a ime ha is no sufficien o
measu emen s (app ox. 68 ms a e age sampling ime), we epea he con olu ion calcu-
la ion block o Figu e 9 up o 10,000 imes o ob ain a measu emen o ol age and cu en
o calcula ing he powe alue in one ins an ime.
On he o he hand, in o de o accu a ely iden i y he cu en co esponding o he
ac i i y o he accele a o , we no ice diffe en cu en le els in he measu emen s (Figu e
10). In he i s one, when he boa d is u ned on and be o e he s a o he es , we ha e
a s eady s a e wi h some noise. Du ing his pe iod o inac i i y, 𝑇, he mean PL cu -
en , is calcula ed and s o ed as a cons an 𝐼. In he second, du ing he execu ion o
he con olu ion (𝑇), 𝐼 ep esen s he cu en consumed by he p ocess. Thus, o
ob ain 𝐼, we ake he diffe ence be ween each ins an aneous measu ed cu en IINT and
he p ep ocessing cu en 𝐼.
Figu e 10. Diffe en cu en le els measu ed in he PL sec ion.
𝐼(𝑖)=𝐼(𝑖)−𝐼(𝑖) (5)
The powe calcula ion is pe o med o each sample i:
𝑃(𝑖)=𝑉(𝑖)∗𝐼(𝑖) (6)
The ene gy consumed by he accele a o du ing he p ocessing ime is ob ained by:
Figu e 10. Di e en cu en le els measu ed in he PL sec ion.
To compa e he ene gy consump ion o INT and FP, we will use a pe cen age o ene gy.
This a io shows he pe cen age o ene gy consumed by he INT e sion compa ed o he
FP e sion, using Equa ion (7), o ob ain EINT and EFP:
PEconsum(k) = 100% ∗EINT
EFP (8)
5. Compa a i e Analysis
In o de o de ine he ype o analysis ha we a e going o do, we ha e o speci y he
objec i e o his wo k. In his ega d, we s a e ha ou objec i e is o obse e he beha io
o ene gy a he con olu ion un le el. This will allow us o iden i y a ela ionship o ene gy
wi h con olu ion cha ac e is ics.
The expe imen pe o med consis s o unning he con olu ion cases men ioned in he
p e ious sec ion. The enso da a A,B, and Ca e ini ialized wi h FP16 and a e andomized
wi h a no mal o Gaussian dis ibu ion. In addi ion, di e en s anda d densi ies a e used
o keep he con olu ion esul s wi h solu ions wi hin he ange o possible esul s.