Improving ductal carcinoma in situ classification by convolutional neural network with exponential linear unit and rank-based weighted pooling
Abstract
British Heart Foundation Accelerator Award, UK
Full text
Complex & Intelligent Systems https://doi.org/10.1007/s40747-020-00218-4 ORIGINAL ARTICLE Improving ductal carcinoma in situ classification by convolutional neural network with exponential linear unit and rank-based weighted pooling Yu-Dong Zhang1,2 ·Suresh Chandra Satapathy3·Di Wu5·David S. Guttery4·Juan Manuel Górriz6· Shui-Hua Wang2,7 Received: 8 August 2020 / Accepted: 7 October 2020 © The Author(s) 2020 Abstract Ductal carcinoma in situ (DCIS) is a pre-cancerous lesion in the ducts of the breast, and early diagnosis is crucial for optimal therapeutic intervention. Thermography imaging is a non-invasive imaging tool that can be utilized for detection of DCIS and although it has high accuracy (~88%), it is sensitivity can still be improved. Hence, we aimed to develop an automated artificial intelligence-based system for improved detection of DCIS in thermographs. This study proposed a novel artificial intelligence based system based on convolutional neural network (CNN) termed CNN-BDER on a multisource dataset containing 240 DCIS images and 240 healthy breast images. Based on CNN, batch normalization, dropout, exponential linear unit and rank-based weighted pooling were integrated, along with L-way data augmentation. Ten runs of tenfold cross validation were chosen to report the unbiased performances. Our proposed method achieved a sensitivity of 94.08±1.22%, a specificity of 93.58±1.49 and an accuracy of 93.83±0.96. The proposed method gives superior performance than eight state-of-theart approaches and manual diagnosis. The trained model could serve as a visual question answering system and improve diagnostic accuracy. Keywords Ductal carcinoma in situ ·Thermal images ·Deep learning ·Convolutional neural network ·Breast thermography · Exponential linear unit ·Rank-based weighted pooling ·Data augmentation ·Color jittering ·Visual question answering Yu-Dong Zhang, Suresh Chandra Satapathy and Di Wu contributed equally to this paper. Yu-Dong Zhang, Suresh Chandra Satapathy and Di Wu are co-first authors. BDavid S. Guttery [email protected] BJuan Manuel Górriz [email protected] BShui-Hua Wang shuihuaw[email protected]g Yu-Dong Zhang [email protected] Suresh Chandra Satapathy [email protected]g Di Wu [email protected] 1School of Informatics, University of Leicester, Informatics Building, University Road, Leicester LE1 7RH, UK Introduction Ductal carcinoma in situ (DCIS), also named intra-ductal carcinoma is a pre-cancerous lesion of cells that line the breast milk ducts, but have not spread into the surrounding breast tissue. DCIS is considered the earliest stage of breast cancer (Stage 0) [1], and although cure rates are 2Department of Information Systems, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia 3School of Computer Engg, KIIT Deemed to University, Bhubaneswar, India 4Leicester Cancer Research Center, University of Leicester, Leicester LE1 7RH, UK 5University of Melbourne, Melbourne, VIC 3010, Australia 6Department of Signal Theory, Networking and Communications, University of Granada, Granada, Spain 7School of Architecture Building and Civil Engineering, Loughborough University, Loughborough LE11 3TU, UK 123
Complex & Intelligent Systems high the patients still need to be treated, since DCIS can become invasive. Note that there are four other stages: Stage 1 describes invasive breast cancer, the cancer cells of which are invading normal surrounding breast tissues. Stages 2 and 3 describe breast cancers that have invaded regional lymph nodes and Stage 4 represents metastatic cancer which spreads beyond the breast and regional lymph nodes to other distant organs [2]. Upon diagnosis of DCIS, treatment options include breast-conserving surgery (BCS), usually in combination with radiation therapy [3] or mastectomy. Breast thermography (BT) is an alternative imaging tool to mammography, which is the traditional diagnostic tool for DCIS. Unlike mammography (which uses ionizing radiation to generate an image of the breast), BT utilizes infra-red (IR) images of skin temperature to assist in the diagnosis of numerous medical conditions, and has been suggested to detect breast cancer up to 10 years earlier than mammography [4]. Furthermore, due to its use of ionizing radiation, mammography can increase the risk of breast cancer by 2% with each scan [5]. Automatic interpretation of DCIS [6] by BT images consists of three phases: (1) segmentation of the region of interest, separating the breast from the image; (2) feature extraction, choosing distinguishing features that can help recognize the suspicious lesion; (3) classification, identifying the image as DCIS or healthy. Previous studies have developed a number of effective artificial intelligence (AI) methods for DCIS detection using BT. Milosevic et al. [7] utilized 50 IR breast images to develop a co-occurrence matrix (COM) and run length matrix (RLM) as IR image descriptors. In the classification stage, a support vector machine (SVM) and naive Bayesian classifier (NBC) were used. Their methods are abbreviated as CRSVM and CRNBC. In addition, Nicandro et al. [8] employed NBC, whereas Chen [9] utilized wavelet energy entropy (WEE) as features to classify breast cancers with promising results. Zadeh et al. [10] combined self-organizing map and multilayer perceptron abbreviated as SMMP and Nguyen [11] introduced Hu moment invariant (HMI) to detect abnormal breasts. Finally, Muhammad [12] combined statistical measure and fractal dimension (SMFD), and Guo [13] proposed a wavelet energy support vector machine (WESVM) to detect breast cancer. Nevertheless, the above methods require laborious feature engineering (FE), i.e., using domain knowledge to extract features from raw data. To help create an improved, automated AI model quickly and effectively, we proposed to use recent deep learning (DL) technologies, viz, convolutional neural networks (CNNs), which are a broad AI technique combining artificial intelligence and representation learning (RL). Our contributions lie in four parts: (1) we proposed a novel 5-layer CNN; (2) we introduced exponential linear unit to replace traditional rectified linear unit; (3) we introduced rank-based weighted pooling to replace traditional pooling methods and (4) we used data augmentation to enhance the training set, so as to improve the test performance. Background Table 14 in “Appendix A” gives the abbreviations and their explanations for ease of reading. Physical fundamentals BT is a sub-science field within IR imaging sciences. IR cameras detect radiation in the long IR range (9–14 µm), with the thermal images generated being dubbed thermograms. Physically, Planck’s law stated the spectral of a body for frequency ωat absolute temperature Tis given as B(ω, T)2oω3 l2 s×1 θ(ω, T)(1a) θ(ω, T)expoω kBT−1,(1b) where Bstands for the spectral radiance, othe Planck constant, kBthe Boltzmann constant, and lsthe light speed,. If replacing frequency ωby wavelength λusing lsλω, above equation can be written as: B(λ, T)2ol2 s λ5×1 θ(λ, T)(2a) θ(λ, T)expols λkBT−1.(2b) Both charge-coupled device (CCD) and complementary metal-oxide-semiconductor (CMOS) sensors in optical cameras detect visible light, and even near-infra-red (NIR) by utilizing parts of the IR spectrum. Basically, they could produce true thermograms with temperatures beyond 280 °C. In our breast thermogram cases, the thermal imaging cameras have a range of 15–45 °C, and a sensitivity around 0.05 °C. Furthermore, three emitted components (ECs) help generate the following breast thermogram images: (1) EC of the breast, (2) EC of the surrounding medium, and (3) EC in the neighboring tissue. Physiological fundamentals In healthy tissue, the major regulation and control of dermal circulation is neurovascular, i.e. through the sympathetic nervous system. Its sympathetic response includes both adrenergic and cholinergic. The former causes vasoconstriction (VC, narrowing of blood vessels); conversely, the latter 123
Complex & Intelligent Systems Adrenergic Cholinergic Normal Blood VesselNarrowing Widening Vasoconstriction Vasodilation Fig. 1 Difference between VC and VD leads to vasodilation (VD, widening of blood vessels). The difference between VC and VD is presented in Fig. 1. In the early stages of cancer growth, cancer cells produce nitric oxide (NO), resulting in VD. Tumor cells then initiate angiogenesis, which is necessary to sustain breast tumor growth. Both VD and angiogenesis lead to increased blood flow; therefore, the increased heat released as a result of increased blood flow to the tumor results in hotter areas than healthy skin. The thermogram of a healthy person is symmetrical across the midline. Asymmetry in the thermogram might signify an abnormality, or even a tumor. Therefore, the thermogram illustrates the status of the breast and presence of breast diseases by identifying asymmetric temperature distribution. Despite this, previous studies [7,8,14] have not measured asymmetry directly. As an alternative, those papers employed texture or statistical measures. As a result, this study did not use asymmetry information, and treated each side image (left breast or right breast) as individual images. Dataset and preprocessing 240 DCIS breast images and 240 healthy breast (HB) images were obtained from 5 sources: (1) our previous study [12] and further collections after its publication. (2) Ann Arbor thermography [15]; (3) The Breast Thermography Image dataset [16]; (4) The Database for Mastology Research with Infrared Image [17] and (5) online resources using search engines including Google, Yahoo, etc. Since our dataset is multi-source, we normalized all the collected images using preprocessing techniques. These included: (I) crop: remove background contents and only preserve the breast tissue and (II) resize: all images were re-sampled to the size of [128 ×128 ×3]. Suppose original image is x1(t),t∈[1,480]. After Step I, we have x2(t)Crop[x1(t),(lt,rt,tt,bt)](3) where (lt,rt,tt,bt)are four parameters denotes left, right, top, and bottom margins of t-th image to be cropped. Finally, after Step II, we have all the images x3(t)∈X3 x3(t)resize[x2(t),(128,128,3)].(4) Fig. 2 Sample of our dataset Note some BT images used different pseudo colormaps (PCMs). For example, some used yellow to denote high temperature while some used red; conversely, some used blue to denote low temperature while some used green. We did not apply the same PCM to all BT images within our dataset for four reasons: (1) we expected our AI model would learn to determine a diagnosis based on color difference, not the color itself; (2) humans can make a diagnosis regardless of the PCM configuration, so we believed AI can do the same; (3) we expected our AI model can be universal, i.e., PCMindependent and (4) mixing of PCM color schemes in the training set can help make our AI model more robust when analyzing the test set, i.e., it does not require a particular PCM scheme. Figure 2shows a DCIS case, where we can clearly see the temperature difference of the lesion and the surrounding healthy tissues. All the images included in our dataset were checked by agreement of two professional radiologists (R1,R2)with more than 10 years of experience. If their decisions [H(R1),H(R2)]agreed, then the images were labelled correspondingly, otherwise, a senior radiologist (R3)was consulted to achieve a consensus: H[x3(t)]H[x3(t),R1]H[x3(t),R1]H[x3(t),R2] M{H[x3(t),(R1,R2,R3)]}otherwise . (5) Here His the labelling result, Mdenotes the majority voting, H[x3(t),(R1,R2,R3)]denotes the labelling results by all three radiologists. Methodology Improvement 1: exponential linear unit The activation function mimics the influence of an extracellular field on a brain axon/neuron. The real activation 123
Complex & Intelligent Systems function for an axon is quite complicated, and can be written as fn1 cVe n−1−Ve n Rn−1 2+Rn 2 +Ve n+1 −Ve n Rn+1 2+Rn 2 +···,(6) where nmeans the index of axon’s compartment model, cthe membrane capacity, Rnthe axonal resistance of compartment n,Ve nthe extra-cellular voltage outside compartment n relative to the ground [18]. This is difficult to determine in an “artificial neural network”, and thus AI scientists designed some simplistic and ideal activation functions (AFs), which have no direct connection with the axon’s activating function, but those AFs work well for ANNs [19]. An important property of AF is nonlinearity. The reason is stacks of linear function will also be linear, and those kinds of linear AFs can only solve trivial problems and cannot make decisions. Only nonlinear AF can allow neural networks to solve non-trivial problems, such as decision-making. Similar ideas were mentioned as “even our mind is governed by the nonlinear dynamics of complex systems” by Mainzer [20]. Suppose the input is t, traditional rectified linear unit (ReLU) [21]fReLU is defined as fReLU(t)max(0,t),(7) with its derivative as f ReLU(t)0t≤0 0t>0.(8) When t<0, the activation of fReLU values are set to zero, so ReLU cannot train the networks via gradient-based learning. Clevert et al. [22] proposed the exponential linear unit (ELU) fELU(γ,t)γet−1t≤0 tt>0.(9) ELU’s derivative is f ELU(γ,t)fELU(γ,t)+γt≤0 tt>0(10) The default value of γ1. Figure 3represents the shapes of five different but common AFs. Each subplot has the same range on the x-axis and y-axis for easy comparison. Information regarding the three AFs (Sigmoid, HT, and LReLU) can be found in “Appendix B”. Improvement 2: rank-based weighted pooling The activation maps (AMs) after conv layer are usually too large, i.e., the size of their width, length, and channels are too large to handle, which will cause (1) overfitting of the training set and (2) large computational costs. Instead pooling layer (PL) is a form of nonlinear downsampling (NLDS) used to solve the above issue. Further, PL can provide invariance-totranslation properties to the AMs. For a 2 ×2 region, suppose the pixels within the region ϕij,(i1,2,j1,2)are ϕ1,1ϕ1,2 ϕ2,1ϕ2,2.(11) Strided convolution (SC) can be regarded as a convolution followed by a special pooling. If the stride is set to 2, the output of SC is: ySC ϕ1,1.(12) The shortcoming of SC is that it will miss stronger activations if ϕ1,1is not the strongest activation. The advantage of SC is the convolution layer only needs to calculate 1/4 of all outputs in this case, so it can save computation. L2P calculates the l2norm [23] of a given region . Assume the output value after NLDS is y, L2P output yL2P is defined as yL2P sqrt2 i,j1φ2 ij. In this study, we add a constant 1/||, where ||means the number of elements of region .Here||4ifweusea2×2 NLDS pooling. This added new constant 1/4 does not influence training and inference. yL2P 2 i,j1φ2 ij ||.(13) The average pooling (AP) calculates the mean value in the region as yAP average() ϕ1,1+ϕ1,2+ϕ2,1+ϕ2,2 ||.(14) The max pooling (MP) operates on the region and selects the max value. Note that L2P, AP and MP work on every slice separately. yMP max() 2 max i,j1ϕi,j.(15) Rank-based weighted pooling (RWP) was introduced to overcome the down-weight (DW), overfitting, and lack of generation (LG) caused by the above pooling methods (L2P, AP, and MP). Instead of computing the l2norm, average, or the max, the output of the RWP yRWP is calculated based on the rank matrix. 123
Complex & Intelligent Systems (a) Sigmoid (b) HT (c) ReLU (d) LReLU (e) ELU Fig. 3 Shape of five different activation functions. HT hyperbolic tangent, ReLU rectified linear unit, LReLU leaky rectified linear unit, ELU exponential linear unit) First, rank matrix (RM) R{rm}is calculated based on the values of each element ϕm∈, usually lower ranks r∈Rare assigned to higher values (ϕ)as ϕm1ϕm2⇒rm1rm2.(16) In case of tied values (ϕm1ϕm2), a constraint is added as (ϕm1ϕm2)∧(m1>m2)⇒rm1>rm2.(17) Second, (ER) map E{em}is defined as emα×(1−α)rm−1,(18) where αis a hyper-parameter. α0.5 for all RWP layers, so we do not need to tune αin this study. Equation (18) can be updated as emα×αrm−1αrm.(19) Third, RWP [24] is defined as the summation of ϕij and eij as below yRWP 2 i,j1 ϕij ×eij.(20) Figure 7in “Appendix C” gives a schematic comparison of L2P, AP, MP, and RWP. Table 1 Pseudocode of RWP Step 1 For an activation map AM with size of [,] Step 2 for =1: % is row index For =1: % is column index Select the 2×2 region Φ: Φ= (2−1:2 ,2 −1:2 ), Generate rank matrix : ={}, See Eqs. (16)(17), Generate exponential rank : ={}, See Eq. (19), Generate RWP (,), See Eq. (20). end end Step 3 Output RWP pooling result : =(,)|=1: ,=1: . For better understanding, a pseudocode of RWP is presented in Table 1. We suppose there is an activation map XAM with size of [R,C], where Rmeans the number of rows, and Cmeans the number of columns. Note row index is set to rand column index c. The RWP output of XAM is symbolized as yRWP with size of R 2,C 2. Table 2itemizes the equations of every pooling methods. 123
Complex & Intelligent Systems Table 2 Comparison of different pooling methods Approach Output Raw ϕ1,1ϕ1,2 ϕ2,1ϕ2,2 SC ySC ϕ1,1 L2P yL2P 2 i,j1φ2 ij || AP yAP ϕ1,1+ϕ1,2+ϕ2,1+ϕ2,2 || MP yMP max2 i,j1ϕi,j RWP yRWP 2 i,j1 ϕij ×eij Improvement 3: L-way data augmentation Traditional data augmentation is a strategy that enables AI practitioners to radically increase the diversity of training data, without collecting new data actually. In this study, we proposed a L-way data augmentation (LDA) technology to further increase the diversity of the training data. The whole preprocessed image set X3,fromEq.(4), will separate into Kfolds: X3 split −→{X3(k1),...,X3(kK)},(21) where krepresents the fold index. At k-th trial, fold kwill be used as the test set Dk, and other folds will be used as the training set Ck: CkX3−Ck(22a) DkX3(k),(22b) If we do not consider the index k, and just simplify the situations as X3→{C,D}, for each training image c (k)∈C,k1,...,|C|, we will do the following eight DA techniques. Here we suppose each DA technique will generate Wnew images. (1) Gamma correction (GC). The equations are defined as: −−→ c1(k)GC[c(k)] cGC 1k,η GC 1,...cGC Wk,η GC W,(23) where ηGC j(j1,...,W)are GC factors. (2) Rotation. Rotation operation rotates the original image to produce Wnew images [25]: −−→ c2(k)RO[c(k)] cRO 1k,η RO 1,...cRO Wk,η RO W (24) where ηRO j(j1,...,W)are rotation factors. (3) Scaling. All training images c(k)were scaled [25]as −−→ c3(k)SC[c(k)] cSC 1k,η SC 1,...cSC Wk,η SC W,(25) where ηSC j(j1,...,W)are scaling factors. (4) Horizontal shear (HS) transform. Wnew images were generated by HS transform −−→ c4(k)HS[c(k)] cHS 1k,η HS 1,...cHS Wk,η HS W,(26) where ηHS j(j1,...,W)are HS factors. (5) Vertical shear (VS) transform. VS transform was generated similarly to HS transform −−→ c5(k)VS[c(k)] cVS 1k,η VS 1,...cVS Wk,η VS W,(27a) ηVS mηHS m,∀m∈1,2,...,W.(27b) (6) Random translation (RT). All training images c(k)were translated Wtimes with random horizontal shift εxand random vertical shift εy, both values of which are in the range of [−, ], and obey uniform distribution U: −−→ c6(k)RT[c(k)] cRT 1k,εx 1,εy 1,...cRT Wk,εx W,εy W,(28) where εx m∼U[−, ],∀m∈[1,W](29a) εy m∼U[−, ],∀m∈[1,W],(29b) where is the maximum shift factor. (7) Color jittering (CJ). CJ shifts the color values in original images [26] by adding or subtracting a random value. The advantage of CJ is it can help bring in randomness change to the color channels, so it can aid production of fake color images: −−→ c7(k)CJ[c(k)] cCJ 1k,ξr 1,ξg 1,ξb 1,...cCJ Wk,ξr W,ξg W,ξb W. (30) 123
Complex & Intelligent Systems Table 3 Proposed five models Index Inheritance Name Description Model-0 Base CNN model BCNN Base model with NCL conv layers and NFCL fully-connected layers Model-1 Model-0 + BN + DO CNN-BD Add BN and DO to Model-0 Model-2 Model-1 + ELU CNN-BDE Use ELU to replace ReLU in Model-1 Model-3 Model-1 + RWP CNN-BDR Use RWP to replace MP in Model 1 Model-4 Model-1 + ELU + RWP CNN-BDER Use ELU and RWP to replace ReLU and MP in Model-1, respectively The shifted color random values are within the range of [−,+],as ξCC m∼U[−,](31a) ∀m∈[1,W]∧∀CC ∈{r,g,b},(31b) where CC means color channel. means maximum color shift value. (8) Noise injection. The 0-mean 0.01-variance Gaussian noises [27] were added to all training images to produce Wnew noised images: −−−−→ cL/2(k)NO[a(k)] cNO 1(k),...cNO W(k),(32) where NO denotes the noise injection operation. (9) Mirror and concatenation. All the above L/2 results are mirrored, we have −−−−−→ cL/2+1(k)M−−→ c1(k)(33a) −−−−−→ cL/2+2(k)M−−→ c2(k)(33b) ··· −−−→ cL(k)M−−−−→ cL/2(k).(33c) where Mrepresents the mirror function. All the results are finally concatenated as −−−−→ cLDA(k) L×W+1 concat⎧ ⎨ ⎩ c(k) 1 ,−−→ c1(k) W ,...,−−−→ cL(k) W ⎫ ⎬ ⎭ .(34) Thesizeof−−−−→ cLDA(k)is L×W+1 images. Thus, the LDA can be regarded as a function c(k)→ −−−−→ cLDA(k). Proposed models and algorithm We proposed five models in total in this study. Table 3 presents their relationships. Model-0 was the base CNN model with NCL conv layers and NFCL fully connected layers. In Model-0, we used max pooling (MP) and ReLU activation function. Model-1 combined Model-0 with batch normalization (BN) and dropout (DO). Model-2 used ELU to replace ReLU in Model-1, while Model-3 used RWP to replace MP in Model-1. Finally, Model-4 introduced both ELU and RWP to enhance the performance based on Model-1. The top row of Fig. 4a shows the activation maps of the proposed Model-0. Here the size of input was S0128 × 128 ×3, the first conv block is composed of one conv layer, one activation function layer, and one pooling layer. After conv layer, S1128 ×128 ×32. Then after the activation function layer, the output is the same as S1. After the pooling layer, the size is S264 ×64 ×32. The conv block then repeats three times, we have S364 ×64 ×64 and S4 32×32×64 for the second conv block, S532 ×32 ×128, and S616 ×16 ×128 for the third conv block, S7 16 ×16 ×256 and S88×8×256 for the four conv block. Then S8was flattened and passed through the first fully connected layer with output as S91×1×50. The output of the second fully connected layer was S10 1×1×2. Measures The randomness effect of each run reduced performance reliability, so we used K-fold cross validation to analyze unbiased performances. The size of each fold is |X3|/K. Due to there being two balanced classes (DCIS and HB), each class will have |X3|/(2×K)images. The split setting of one trial is shown in Table 6. Within each trial, (K−1) folds were used as training, and the rest fold were used as test. After combining all Ktrials, the test image grew to |X3|. If above K-fold cross validation repeats Zruns, the performance will be reported on |X3|×Zimages. Suppose the ideal confusion matrix Eideal over the test set at k-th trial and z-th run is Eideal(k,z)!|X3| 2×K0 0|X3| 2×K",(35) 123
Complex & Intelligent Systems Fig. 4 Block chart of five proposed models. Ssize, Cconv, BN batch normalization, RReLU,EELU, Ddropout, Ffully connected where the constant 2 is because our dataset is a balanced, i.e., DCIS class has the same size of HB. After combining K trials, the ideal confusion matrix is at z-th run is Eideal(z) K k1 Eideal(k,z) !|X3| 20 0|X3| 2"(36) In realistic inference, we cannot get the perfect diagonal matrix as shown in Eq. (36), suppose the z-th run real confusion matrix is Ereal(z) K k1 Ereal(k,z) a(z)b(z) c(z)d(z)(37) 123
Complex & Intelligent Systems where 0 ≤a,b,c,d≤|X3|/2. The four variables (a,b,c,d)represent TP, FN, FP, and TN, respectively. Here Pmeans DCIS and Nmeans healthy breast (HB). Four simple measures ν1,ν2,ν3,ν4can be defined as ν1(z)a(z) a(z)+b(z)(38a) ν2(z)d(z) c(z)+d(z)(38b) ν3(z)a(z) a(z)+c(z)(38c) ν4(z)a(z)+d(z) a(z)+b(z)+c(z)+d(z).(38d) where ν1(z),ν2(z),ν3(z),ν4(z)means sensitivity, specificity, precision, and accuracy at z-th run, respectively. Besides, F1 score ν5(z), Matthews correlation coefficient (MCC) ν6(z), and Fowlkes–Mallows index (FMI) ν7(z)can be defined as: ν5(z)2×ν3(z)×ν1(z) ν3(z)+ν1(z) 2×a(z) 2×a(z)+b(z)+c(z)(39a) ν6(z)d(z)×a(z)−c(z)×b(z) √γ(z)(39b) γ(z)[c(z)+a(z)]×[a(z)+b(z)] ×[d(z)+c(z)]×[d(z)+b(z)](39c) ν7(z)a(z) a(z)+c(z)×a(z) a(z)+b(z).(39d) After averaging Zruns, we can calculate the mean (M) and standard deviation (SD) of all k-th (∀k∈[1,7])measures as Mνk1 Z× Z z1 νk(z)(40a) SDνk# $ $ %1 Z−1× Z z1νk(z)−Mνk2.(40b) The result is reported in the format of M±SD. For ease of typing, we write it in short as MSD. Experiments and results Parameter setting Table 4shows the parameter setting of variables in this study. The values were obtained using trial-and-error. The total size Table 4 Parameter setting of variables Parameter Meaning Value |X3|Size of preprocessed image set 480 |Ck|Size of training set at k-th trial 432 |Dk|Size of test set at k-th trial 48 KTotal number of k-folds 10 WNumber of new images for each DA 30 LNumber of DA techniques 16 Maximum color shift value 50 ZTotal number of runs of K-fold cross validation 10 NCL Number of conv layers/blocks 4 NFCL Number of fully connected layers/blocks 2 Table 5 LDA parameter setting LDA parameter Values GC factors ηGC 10.4,η GC 20.44,...,η GC 15 0.96, ηGC 16 1.04,η GC 17 1.08,...η GC W1.6 Rotation factors ηRO 1−W◦,η RO 2 −W+2 ◦,...,η RO 15 −2◦,ηRO 16 +2◦,η RO 17 +4◦,...,η RO W+W◦ Scaling factors ηSC 10.7,η SC 20.72,...,η SC 15 0.98, ηSC 16 1.02,η SC 17 1.04,...,η SC W1.3 HS factors ηHS 1−0.15,η HS 2 −0.14,...,η HS 15 −0.01, ηHS 16 +0.01,η HS 17 +0.02,...,η HS W+0.15 Maximum shift factor 15 Maximum color shift value 50 of our dataset was 480, and thus the size of the preprocessed image set is |X3|480. The number of folds and runs were all set to 10, i.e., K10,Z10. Then, each fold contained 48 images, that is 24 DCIS and 24 HB images. The training set contained |C|432 images, and the test set contained |D|48 images. The number of DA ways was L16, the number of new images for each DA technique was W30. Thus, we created L×W480 new images for every training image. The number of conv layers/blocks was NCL 4, and the number of fully connected layers/blocks was NFCL 2. Table 5itemizes the LDA parameter settings. The GC factors ηGC varied from 0.4 to 1.6 with an increase of 0.04, skipping the value of 1. The rotation vector ηRO was in the value from −Wto Wan increase of 2°, skipping ηRO 0. Scaling factor ηSC varied from 0.7 to 1.3 with an increase of 0.02, skipping ηSC 1. HS factors ηHS varied from − 0.15 to 0.15 with an increase of 0.01, skipping the value of ηHS 0. The maximum shift factor 15. The maximum color shift value was 50. 123
Complex & Intelligent Systems 12. Muhammad K (2017) Ductal carcinoma in situ detection in breast thermography by extreme learning machine and combination of statistical measure and fractal dimension. J Ambient Intell Humaniz Comput. https://doi.org/10.1007/s12652-017-0639-5 13. Guo Z-W (2018) Breast cancer detection via wavelet energy and support vector machine. In: 27th IEEE international conference on robot and human interactive communication (ROMAN), Nanjing, China, 2018, pp 758–763 14. Milosevic M, Jankovic D, Peulic A (2014) Thermography based breast cancer detection using texture features and minimum variance quantization. EXCLI J 13:1204–1215 15. Ann Arbor Thermography. https://aathermography.com/breast/ breasthtml/breasthtml.html 16. Breast Thermography Image dataset (2020). https://www.dropbox. com/s/c7gfp2bo1ae466m/database.zip?dl=0 17. Silva LF, Saade DCM, Sequeiros GO, Silva AC, Paiva AC, Bravo RS et al (2014) A new database for breast research with infrared image. J Med Imaging Health Inform 4:92–100 18. Rattay F (1998) Analysis of the electrical excitation of CNS neurons. IEEE Trans Biomed Eng 45:766–772 19. Górriz JM (2020) Artificial intelligence within the interplay between natural and artificial computation: advances in data science, trends and applications. Neurocomputing 410:237–270 20. Mainze K (1997) Introduction: from linear to nonlinear thinking. In Thinking in complexity. Springer, Berlin, pp 1–2 21. Nair V, Hinton GE (2010) Rectified linear units improve restricted Boltzmann machines. In: 27th International conference on machine learning (ICML), Haifa, Israel, 2010, pp 807–814 22. Clevert D-A, Unterthiner T, Hochreiter S (2016) Fast and accurate deep network learning by exponential linear units (ELUs). arXiv. arXiv:1511.07289v5 23. Rezaei M, Yang H, Meinel C (2017) Deep neural network with l2-norm unit for brain lesions detection. In: International conference on neural information processing (ICNIP), Cham, 2017, pp 798–807 24. Jiang YY (2017) Cerebral micro-bleed detection based on the convolution neural network with rank based average pooling. IEEE Access 5:16576–16583 25. Blok PM, van Evert FK, Tielen APM, van Henten EJ, Kootstra G (2020) The effect of data augmentation and network simplification on the image-based detection of broccoli heads with Mask R-CNN. J Field Robot. https://doi.org/10.1002/rob.21975 26. Puttaruksa C, Taeprasartsit P (2018) Color data augmentation through learning color-mapping parameters between cameras. In: 15th International joint conference on computer science and software engineering, Mahidol University, Facility ICT, Thailand, 2018, pp 6–11 27. Pandian JA, Geetharamani G, Annette B, Ieee (2019) Data augmentation on plant leaf disease image dataset using image manipulation and deep learning techniques. In: 9th International conference on advanced computing, MAM College of Engineering and Technology, Tiruchirapalli, India, 2019, pp 199–204 28. Marcot BG, Hanea AM (2020) What is an optimal value of k in k-fold cross-validation in discrete Bayesian network analysis? Comput Stat. https://doi.org/10.1007/s00180-020-00999-9 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. 123