Full text
Proyecto Fin de Carrera Ingeniería de Telecomunicación Formato de Publicación de la Escuela Técnica Superior de Ingeniería Autor: F. Javier Payán Somet Tutor: Juan José Murillo Fuentes Dep. Teoría de la Señal y Comunicaciones Escuela Técnica Superior de Ingeniería Universidad de Sevilla Sevilla, 2013 Doctoral Thesis Advanced Signal Processing Techniques for the Modeling and Linearization of Wireless Communication Systems Author: Juan Antonio Becerra González Advisor: Luis Javier Reina Tosina Co-advisor: Carlos Crespo Cadenas Teoría de la Señal y Comunicaciones Escuela Técnica Superior de Ingeniería Universidad de Sevilla Sevilla, Spain 2018
Doctoral Thesis Advanced Signal Processing Techniques for the Modeling and Linearization of Wireless Communication Systems Author: Juan Antonio Becerra González Advisor: Luis Javier Reina Tosina Profesor Titular Co-advisor: Carlos Crespo Cadenas Profesor Titular Teoría de la Señal y Comunicaciones Escuela Técnica Superior de Ingeniería Universidad de Sevilla Sevilla, Spain 2018
Doctoral Thesis: Advanced Signal Processing Techniques for the Modeling and Linearization of Wireless Communication Systems Author: Juan Antonio Becerra González Advisor: Luis Javier Reina Tosina Co-advisor: Carlos Crespo Cadenas El tribunal nombrado para juzgar la Tesis arriba indicada, compuesto por los siguientes doctores: Presidente: Vocales: Secretario: acuerdan otorgarle la calificación de: El Secretario del Tribunal Fecha:
A mis padres, Juan Antonio y Lucía. To my parents, Juan Antonio and Lucía.
Acknowledgements I would like to express my sincerest gratitude to my research group, María José, Javier, and Carlos. Thank you very much for everything you taught me and the patience you had with me. I appreciate all the guidance during these years and the unconditional support you have shown to every kind of idea. To the young grasshopper of the family, Abraham, for bringing strength, energy, and motivation in this final stage. To my colleagues in the Department of Signal Theory and Communications, for making me feel at home and for everything I learned from them. Special thanks to Irene Fondón, Uxi, Iván, Eva, Irene Santos and Javier Olías. To Michel, for his patience while teaching the basic concepts on which this thesis is based. To Deepa, for proving that connection with other people is above language barriers. To my colleagues at the Université de Limoges and the University of Delaware, for the new dimensions they gave to this project. To Javier Payán and Juanjo, for the design of this template that has been so simple to use. To Esperanza, for the tremendous logistical support at the time of doing paperwork with the University. To María and José, for all the cheering while preparing the gasoline that accelerates the research. To Aurora, a friend more than a sister. Finally, I would like to close the acknowledgments with the same sentence that I started my final undergraduate project nine years ago: "To my parents, for all that I am, and for all that I do not". III
Short Contents Abstract VII Resumen IX Short Contents XI List of Acronyms XVII Notation XIX 1 Overview 1 1.1 Introduction 1 1.2 Aims and Scope 2 1.3 Thesis Outline 3 2 Materials and Methods 5 2.1 Behavioral Modeling Background 5 2.2 Performance of Behavioral Models 7 2.3 Model Parameters and DPD Identification 10 2.4 Experimental Testbenches 16 3 Volterra-Parafac Digital Predistorters 21 3.1 Introduction 21 3.2 Volterra-Parafac Models and Frequency-Domain Techniques 22 3.3 Complexity Assessment 27 4 Formal Deduction of a Volterra Series Model for Complex-valued Systems 31 4.1 Introduction 31 4.2 Volterra Models for Real-valued Systems 32 4.3 Specific Volterra Models for Complex-valued Systems 32 4.4 A General Volterra Model for Complex Systems 33 XI
XII Short Contents 5 Structured Compressed-Sensing Volterra Series 37 5.1 Introduction 37 5.2 Structured Compressed-Sensing for Volterra Series Models 38 6 Transmitter Linearization Adaptable to Power-Varying Operation 41 6.1 Introduction 41 6.2 DPD Structure and Identification Procedure 42 7 Results 49 7.1 Volterra-Parafac Digital Predistorters 49 7.2 Complex-valued Volterra Series Model 55 7.3 Structured Volterra Series Model 57 7.4 Transmitter Linearization Adaptable to Power-Varying Operation 59 8 Conclusions and future work 69 8.1 Conclusions 69 8.2 Future Work 71 8.3 Contributions 71 List of Figures 75 List of Tables 79 Bibliography 81
Contents Abstract VII Resumen IX Short Contents XI List of Acronyms XVII Notation XIX 1 Overview 1 1.1 Introduction 1 1.2 Aims and Scope 2 1.3 Thesis Outline 3 2 Materials and Methods 5 2.1 Behavioral Modeling Background 5 2.1.1 Volterra Series Models 6 2.2 Performance of Behavioral Models 7 2.2.1 Metrics of Performance 7 2.3 Model Parameters and DPD Identification 10 2.3.1 Digital Predistortion 10 2.3.2 Least Squares (LS) 11 2.3.3 Ridge Regression 12 2.3.4 The Least Absolute Shrinkage and Selection Operator (LASSO) 13 2.3.5 Selection of the Tuning Parameter 14 2.3.6 Adaptive Optimization 14 2.3.7 Steepest Descent 14 2.3.8 Least Mean Squares (LMS) Algorithm 16 2.4 Experimental Testbenches 16 2.4.1 University of Seville Testbench 16 2.4.2 Chalmers University of Technology Weblab 17 XIII
XIV Contents 3 Volterra-Parafac Digital Predistorters 21 3.1 Introduction 21 3.2 Volterra-Parafac Models and Frequency-Domain Techniques 22 3.2.1 The Volterra-Parafac Baseband Model 22 3.2.2 Block Processing and FFT Algorithm Applied to the VP Structure (B-VP DPD) 24 3.2.3 Particularization to OFDM Systems (FD-VP DPD) 26 3.3 Complexity Assessment 27 3.3.1 B-VP DPD Approach 28 3.3.2 FD-VP DPD Approach 28 4 Formal Deduction of a Volterra Series Model for Complex-valued Systems 31 4.1 Introduction 31 4.2 Volterra Models for Real-valued Systems 32 4.3 Specific Volterra Models for Complex-valued Systems 32 4.4 A General Volterra Model for Complex Systems 33 5 Structured Compressed-Sensing Volterra Series 37 5.1 Introduction 37 5.2 Structured Compressed-Sensing for Volterra Series Models 38 6 Transmitter Linearization Adaptable to Power-Varying Operation 41 6.1 Introduction 41 6.2 DPD Structure and Identification Procedure 42 6.2.1 Proposed DPD Model for Joint Mitigation in Transmitters 43 6.2.2 Identification Procedure 44 6.2.3 Preceding DPD Models for Joint Mitigation 47 7 Results 49 7.1 Volterra-Parafac Digital Predistorters 49 7.1.1 Measured Performance for the B-VP DPD 49 7.1.2 Measured Performance for the FD-VP DPD 52 7.2 Complex-valued Volterra Series Model 55 7.3 Structured Volterra Series Model 57 7.4 Transmitter Linearization Adaptable to Power-Varying Operation 59 7.4.1 Linearization of a Basic Transmitter 59 7.4.2 Linearization of a Realistic Transmitter 63 7.4.3 Power Adaptability of ACC and CVS Linearizers 65 8 Conclusions and future work 69 8.1 Conclusions 69 8.2 Future Work 71 8.3 Contributions 71
Contents XV List of Figures 75 List of Tables 79 Bibliography 81
List of Acronyms List of Acronyms ACPR Adjacent Channel Power Ratio ACEPR Adjacent Channel Error Power Ratio BIC Bayesian Information Criterion DDR Dynamic Deviation Reduction DPD Digital Predistortion DUT Device Under Test EVM Error Vector Magnitude FPGA Field-Programmable Gate Arrays FV Full Volterra GMP Generalized Memory Polynomial LASSO Least Absolute Shrinkage and Selection Operator LMS Least Mean Squares LS Least Squares LTE Long Term Evolution MP Memory Polynomial NMSE Normalized Mean Square Error OFDM Orthogonal Frequency Division Multiplexing PA Power Amplifier PAPR Peak-to-Average Power Ratio PRSS Penalized Residual Sum of Squares XVII
XVIII Chapter 0. Contents QAM Quadrature Amplitude Modulation RSS Residual Sum of Squares VSA Vector Signal Analyzer VSG Vector Signal Generator
Notation tContinuous-time variable kDiscrete-time variable nIteration index xInput signal in vector notation XMeasurement matrix yOutput signal in vector notation hVolterra kernel vector ˆ yEstimated output signal in vector notation x(t)Input signal in continuous time t x[k]Input signal in discrete time k hiVolterra kernel of i-th order fContinuous-frequency variable X(f)Signal in frequency domain εError signal AHHermitian of matrix A A+Moore-Penrose pseudoinverse of matrix A XIX
6Chapter 2. Materials and Methods where h0 is a constant and hk(τ1,τ2,...,τk) is the k -th order Volterra kernel. The causal discrete-time Volterra filter is equivalently described by y[k] = h0+ ∞ ∑ k1=0 h1(k1)x[k−k1]+ ∞ ∑ k1=0 ∞ ∑ k2=0 h2(k1,k2)x[k−k1]x[k−k2]+ ∞ ∑ k1=0 ∞ ∑ k2=0 ··· ∞ ∑ kp=0 hp(k1,k2,...,kp)x[k−k1]x[k−k2]...x[k−kp]+... (2.2) Volterra series expansions have the following properties [6]: Linearity with respect to the kernel coefficients: Although Volterra series represent a nonlinear behaviour between the input and the output of the system, the kernels are linear with respect to the output of the system. Symmetry of the kernels: The permutation of the indices in a Volterra series kernels results in the same combination of input samples. Multidimensional convolution property: A p -th order Volterra kernel may be written as a p-dimensional convolution. Stability property: A Volterra kernel of order p is bounded-input bounded-output (BIBO) stable if ∞ ∑ k1=0 ··· ∞ ∑ kp=0 |hp(k1,...,kp)|<∞.(2.3) High kernel complexity: The number of parameters in Volterra series increases rapidly with the polynomial order and memory. This is generally known as the curse of dimensionality. 2.1.1 Volterra Series Models PAs are passband devices. By assuming that the bandwidth of the signal is small with respect to the center frequency, instead of using passband samples (at high sample rates), baseband samples (at lower sample rates) can be used. The relation between the passband signal ˜x(t)and its baseband formulation x(t)follows ˜x(t) = ej2πf0tx(t)+e−j2πf0tx∗(t) 2,(2.4) where x∗(t) is the complex conjugate of x(t) . This complex baseband representation is used in the rest of this Thesis. The discrete baseband representation of the general Volterra
2.2 Performance of Behavioral Models 7 series has the structure of y[k] = P ∑ p1=0 Q ∑ q1=0 Q ∑ q2=q1 ··· Q ∑ qp=qp−1 Q ∑ qp+1=0 Q ∑ qp+2=qp+1 ··· Q ∑ q2p−1=q2p−2 h2p−1(q1,q2,...,q2p−1) p ∏ j1=1 x[k−qj1] 2p−1 ∏ j2=p+1 x∗[k−qj2]. (2.5) In the literature, this model is referred to as the Full Volterra ( FV ) model. In a practical application, the number of coefficients of this most-general model increases rapidly with the order 2P−1 and the memory depth Q [ 7 ]. Because of this, pruned versions have been extensively researched and developed. 2.2 Performance of Behavioral Models The aim we pursue when a system is characterized by a behavioral model is to obtain an transfer function that is able to represent the output precisely. For that purpose, several performance metrics are defined and widely used [ 8 ]. In this section, these key indicators are highlighted. 2.2.1 Metrics of Performance The most widely used performance indicator is the Normalized Mean Square Error ( NMSE ). The NMSE represents the ratio between the error power and the reference signal power, and it measures the modeling accuracy capability of the model. It is defined as NMSE =kymodel −yrefk2 2 kyrefk2 2 ,(2.6) where ymodel is the output of the model (stacking the samples in vector form) for an input signal, yref is the reference signal vector and kVk`= Nv ∑ i=0 (vi)`!1 ` (2.7) is the `norm of the vector V= [v1,v2,...,vNv]. The Adjacent Channel Power Ratio ( ACPR ) is also extensively applied to get the model performance. The ACPR takes into account the ratio between the power in the main operation band to the power in the adjacent bands in the spectrum. Depending on which band we are taking as a reference, the ACPR is usually tagged with a plus or a minus symbol and a number. The + symbol indicates the upper channel and the − symbol stands
8Chapter 2. Materials and Methods Figure 2.1 Example of an AM/AM characteristic. for the lower channel. The number indicates the channel (1 for the first, 2 for the second and so on). The ACPR is defined as ACPR(+/−)K=Rchannel0|Y(f)|2d f Rchannel(+/−)K|Y(f)|2d f ,(2.8) where Y(f) represents the spectrum of the signal y . The ACPR is calculated over the error signal when it is taken as a performance indicator of a modeling and it is calculated over the linearized signal when the evaluation of the linearization capabilities is performed. The AM/AM and the AM/PM characteristics are the graphical representations of the instantaneous power of the output and the instantaneous phase of the output versus the instantaneous power of the input. An example of these representations are shown in figures 2.1 and 2.3. The AM/AM characteristic is sometimes represented using the instantaneous gain instead of the instantaneous input power (see figure 2.2). The AM/AM plot gives relevant information about the behaviour of the system, as Linear gain The linear gain of the system is the slope of the AM/AM. Compressed gain The compressed gain is the gain at the highest input power. Memory effects Memory effects are evidenced by the width of the cloud at medium-low powers. Nonlinear effects Nonlinear effects are shown in this representation as the average evolution of the output power with the input power. The Error Vector Magnitude ( EVM ) is a performance indicator of in-band distortion for digitally modulated signals. It is calculated over the received constellation with respect
2.2 Performance of Behavioral Models 9 Figure 2.2 Example of an AM/Gain characteristic. Figure 2.3 Example of an AM/PM characteristic. to the ideal constellation at the transmitter. It measures how far are the points from the ideal locations, giving a measure of the error in the constellation space. EVM =Perror Preference ,(2.9) where Perror is the amplitude of the error vector in the constellation space and Preference is the coordinate associated to the transmitted symbol.
10 Chapter 2. Materials and Methods 2.3 Model Parameters and DPD Identification 2.3.1 Digital Predistortion The basic principle of digital predistortion is to obtain a block that will be placed in front of the nonlinear PA so that the combined action of these two systems produces a scaled version of the input [9]. The input-output characteristic of a typical PA is shown in figure 2.4. In a PA, when the signal level is low, the response is close to linear. At high signal levels, the PA produces compression and generates distortion. When the PA is not able to deliver more power, it is in saturation. The input-output response of a DPD is also presented. While the PA shows a compressive characteristic, the DPD has an expansive characteristic that is the symmetrical complementary of the PA with respect to a linear gain. The combination of these two characteristics results in a linear relationship. The inverse characteristic of the DPD will only work up to the saturation power of the PA, therefore we need to limit the input signal level to prevent the PA becoming saturated. Input signal level Ouput signal level PA DPD Linear gain Figure 2.4 Input-output characteristic of a saturated PA and a DPD. There is a need of upsampling in the input of the DPD. Since it generates harmonics in anti-phase to compensate the nonlinear behavior of the PA, the Nyquist sample rate is no longer valid. The sampling frequency needs to accommodate the new generated spectral bands. The oversample rates will depend on the bandwidth of the input signal and the DPD capability we are expecting. The typical oversampling rate is over five times the signal bandwidth. Digital predistortion is commonly considered as a traditional closed-loop digital adaptive problem, in which the controller (the DPD) is tuned by some algorithm.
2.3 Model Parameters and DPD Identification 11 In an indirect learning scheme [ 10 ], we aim at finding a postdistorter, that is, a block that creates the linear input to the system at its output when it has in its input a scaled version of the PA output. The scale is usually the average power gain of the PA, i.e., the DPD has an expected gain of about the unit. This is required for stability of the DPD loop. The estimation of the DPD coefficients may be performed at once through Least Squares ( LS ) or with an adaptive technique. These algorithms are introduced in the next sections. DPD PA x z y Figure 2.5 Schematic of a predistorter and a power amplifier. 2.3.2 Least Squares (LS) The LS regression, which appeared formally formulated at the beginning of the 19th century [ 11 ], is one of the most common techniques for regressing a linear set of equations, since the error is penalized quadratically and it has only one solution. In the case of Volterra series, although the relation between the input x= [x(n),x(n−1),...,x(n−M−1)] , where M is the maximum number of samples taken into account and the output y= [y(n),y(n−1),...,y(n−M−1)] is nonlinear by its definition, there exists a linear regression between the vector of coefficients h and the output y , where the columns of the measurement matrix X make the transformation from one space to another. The measurement matrix structure is model-dependent. Its columns, X{i} , are the regressors of the model and usually take the form of some kind of function of the input complex envelope x(k) . For example, the Memory Polynomial ( MP ) model [ 12 ] generates p th-order regressors lagged q samples with the form x(k−q)|x(k−q)|p−1 . Its P th-order and maximum memory Q measurement matrix has the following structure XMP = x(k)··· x(k)|x(k)|P−1··· x(k−Q)|x(k−Q)|P−1 x(k−1)··· x(k−1)|x(k−1)|P−1··· x(k−Q−1)|x(k−Q−1)|P−1 . . .. . .. . .. . .. . . x(k−(M−1)) ··· x(k−(M−1))|x(k−(M−1))|P−1··· x(k−Q−(M−1))|x(k−Q−(M−1))|P−1 . (2.10) X x + −y hˆ ye Figure 2.6 Block diagram of Volterra series represented as a measurement process where his the vector of Volterra coefficients to be estimated.
12 Chapter 2. Materials and Methods Taking into account this linear relation and following the block system shown in figure 2.6, the measurement equation ˆ y=Xh, the error vector can be defined as ε=y−ˆ y=y−Xh.(2.11) Minimizing the norm of the error is performed by finding the minimum of kε(h)k2=εHε= (yH−hHXH)(y−Xh) =yyH−yHXh −hHXHy+hHXHXh,(2.12) where AH is the hermitian transpose of matrix A . Taking the derivative of this equation with respect to the kernel vector and setting it to zero, ∂kε(h)k2 ∂h=−2XHy+2XHXh =0,(2.13) we get the so-called deterministic normal equation, XHy=XHXˆ h,(2.14) where h has become ˆ h since its value is the estimated optimum value after setting the derivative equal to zero. Please note that X is not generally square, therefore it is not invertible. Nevertheless, the product XHX is an invertible square matrix, allowing us to rearrange ˆ h= (XHX)−1XHy.(2.15) The term X+= (XHX)−1XH is commonly known as the Moore–Penrose pseudoinverse, which it is a generalization of the inverse matrix for non-square matrices. In a LS regression, the estimate error εis orthogonal to the columns of the measurement matrix. 2.3.3 Ridge Regression Ridge regression is similar to LS , with the exception that the Penalized Residual Sum of Squares (PRSS) is minimized instead of the Residual Sum of Squares (RSS). PRSS =ky−Xhk2 2+λkhk2 2=RSS+λkhk2 2,(2.16) where λ is the tuning parameter. Ridge regression trades off two different criteria. The first term seeks coefficients that fit the data well, producing a small RSS . The shrinkage penalty λkhk2 2 has the effect of shrinking the coefficients towards zero. The tuning parameter λ controls the impact of both terms on the regression. When λ=0 , the shrinkage penalty has no effect and the equation becomes the LS regression. When λ→∞ , the penalty grows and the estimates approach zero. Ridge regression produces a set of estimates instead of a single solution as the LS regression does, therefore a selection procedure for λ is critical. As the solution depends on the scaling of the predictors, a good practice is to standardize
2.3 Model Parameters and DPD Identification 13 the predictors by dividing them by its `2norm: ¯ Xi=Xi kXik2 .(2.17) The advantage of Ridge regression over LS is the existence of the tuning parameter and that it allows to control the bias-variance trade-off. As λ increases, the flexibility of the Ridge regression decreases, leading to decreased variance but increased bias. In cases where the number of samples is higher than the number of regressors, as it is with Volterra series, Ridge regression can perform well by trading off a small increase in bias for a large decrease in variance. An alternative formulation for Ridge regression is solving the problem minimize hky−Xhk2 2subject to khk2 2≤s.(2.18) The estimator of a Ridge regression has a closed form that follows ˆ hRidge = (XHX+λI)−1XHy,(2.19) where Iis the identity matrix. 2.3.4 The LASSO The main disadvantage of Ridge regression is that all the p predictors are included in the final model. The penalty λkhk1 will shrink all the coefficients towards zero, but it will not set any of them exactly to zero, except for λ=∞. The LASSO is an alternative to the Ridge regression that overcomes this disadvantage. ky−Xhk2 2+λkhk1=RSS+λkhk1.(2.20) The formulation is very similar to the Ridge regression but the penalty includes the `1 norm instead of the `2. The LASSO shrinks the coefficients towards zero. However, in comparison with the Ridge regression, the `1 penalty forces the coefficients to be exactly equal to zero when the tuning parameter λ is large. The LASSO yields sparse models —which involve a subset of variables—. In the case of the LASSO regression the selection of λis critical as well. The equivalent to the formulation in 2.19 to the LASSO is minimize hky−Xhk2 2+λkhk1subject to khk1≤s.(2.21) In a two-dimensional plane with coordinates β1 and β2 , Ridge regression seeks the solution with the constraint β2 1+β2 2≤s , while the LASSO constraint has the form |β1|+ |β2| ≤ s . Figure 2.7 shows the constraints for both regressions. The LS solutions are marked as ˆ h and all the points in the same ellipse have a common value in RSS . The RSS increases as the ellipses are further from the LS solution. The Ridge and LASSO solutions are the first point at which an ellipse contacts the constraint region. The constraint of the LASSO has a diamond shape and the Ridge regression’s is sphere-shaped. Since Ridge
14 Chapter 2. Materials and Methods has no sharp points, this intersection will not generally occur in the axis -equivalent to sparse solutions-. However, the LASSO constraint has corners in the axis which will often intercept the ellipse at an axis. When this occurs, those coefficients will be equal to zero. In higher dimensions, many of the coefficients may be equal to zero simultaneously. ˆ β ˆ β 1 β 1 β 2 β 2 β Figure 2.7 Contours of the error and constraint functions for the LASSO (left) and Ridge regression (right) in a space with coordinates β1 and β2 . The solid areas are the constraint regions, |β1|+|β2| ≤ s and β2 1+β2 2≤s , while the ellipses are the contours of the RSS. Figure adapted from [13]. 2.3.5 Selection of the Tuning Parameter The general technique for selecting a value for the tuning parameter is known as crossvalidation. Cross-validation is a simple procedure that consists on choosing a grid of λ values and computing the cross-validation error for each value of λ . The λ that gives the lowest cross-validation error is chosen and the model is re-fit with all the available observations. 2.3.6 Adaptive Optimization Adaptive optimization is a technique where the estimation of the kernel vector is updated dynamically in the time. It has the advantage of being more robust to the change of conditions of the system. This family of algorithms is able to follow the effects that may appear with a change of any variable -signal type, power level, temperature, etc.- with the counterpart of constantly update the model of the system. The scheme is very similar to that of the LS , but taking into account only the last N samples of each of the signals. In this topic, the index nis introduced to express the dependency with the time instant. 2.3.7 Steepest Descent The block diagram followed in this topic is shown in figure 2.8. Please note that it has the same structure than figure 2.6 with the addition of dependence on the time or iteration variable n. For the derivation of the algorithm, we depart from the measurement process equation that takes the form
2.3 Model Parameters and DPD Identification 15 X(n) x(n) + −y(n) h(n)ˆ y(n)e(n) Figure 2.8 Block diagram of Volterra series represented as a measurement process where his the vector of Volterra coefficients to be estimated. ˆ y(n) = hH(n)X(n),(2.22) where ˆ y(n)∈CM is the estimated output in the iteration n , h(n)∈CN is the Volterra vector with the N coefficients of the model, X(n)∈CM×N is the measurement matrix, whose N columns corresponds to the regressors of the model formed with a section of the last N samples of the input signal x(n)=[x(n)x(n−1)...x(n−M+1)] . The optimization error is defined as ε(n) = y(n)−ˆ y(n) = y(n)−hH(n)X(n).(2.23) Considering the `2-norm as the cost function to minimize, J(n) = Ekε(n)k2=Ey(n)−hH(n)X(n)yH(n)−XH(n)h(n),(2.24) and handling the terms considering the random character of X(n)and y(n), J(n) = Ey(n)yH(n) −Ey(n)XH(n)h(n) −hH(n)EX(n)yH(n) +hH(n)EX(n)XH(n)h(n). (2.25) The final cost function remains as J(n) = σ2 y+pHh(n)−hH(n)p+hH(n)Rh(n),(2.26) where pis the cross-correlation between X(n)and yH(n)and Ris the correlation matrix of X(n). If we set the direction of the kernel vector update to −∇J(n) = 2p−2Rh(n), h(n+1) = h(n)+ 1 2µ[−∇J(n)] =h(n)+ µ[p−Rh(n)], (2.27) where µ is the step size, which can be demonstrated that for the algorithm to converge, it should have a value in the range of 0<µ<2 λmax , being λmax is the maximum eigenvalue of the correlation matrix R.
22 Chapter 3. Volterra-Parafac Digital Predistorters Volterra kernels viewed as tensors has been demonstrated in [ 21 ]. This Volterra-Parafac (VP) structure can be represented as an arrangement of parallel branches formed with cascaded linear filters and static homogeneous nonlinearities. Recently, a new baseband VP model with a reduced parametric complexity has been derived by using a doubly symmetric Parafac decomposition [22]–[23]. Based on the encouraging results of PA linearization in [ 24 ], this chapter is dedicated to demonstrate the design of two Volterra-Parafac DPDs under a frequency domain perspective and show their improvement capacity, mainly in terms of computational complexity. In the first approach, the DPD operates on the discrete-time samples of the input waveform to produce the predistorter signal. Block processing and frequency-domain techniques are applied to improve the computational efficiency. We refer to this technique as blockprocessed Volterra-Parafac (B-VP) DPD and it has been favourably compared to a DPD designed with a high-standard model, the GMP. In the second case, the predistorter acts directly on the data symbols, exploiting an attractive feature of the VP model in the particular case of OFDM systems, i.e., its ability to integrate the M -QAM symbols within the signal processing technique. In this approach, designated as frequency-domain Volterra-Parafac (FD-VP) DPD, the DPD input data is the block of M -QAM symbols corresponding to the active subcarriers in the spectral domain, and the output is the discrete-time predistorter signal. The discussion of the Volterra-Parafac approach, block adaptive filtering and frequencydomain techniques applied to VP models is developed in section 3.2. In section 3.3, the computational complexity of the model is analyzed. 3.2 Volterra-Parafac Models and Frequency-Domain Techniques Making use of the fact that the kernels of the full Volterra behavioral model can be seen as symmetric multidimensional arrays, the tensorial Parafac decomposition has been applied to introduce the VP representation [ 21 ]. Since in this work we use the corresponding discrete-time baseband structure published in [ 23 ] to model the DPD, we dedicate the following subsection to expose a brief overview of this VP representation and the parameter identification procedure. 3.2.1 The Volterra-Parafac Baseband Model The VP baseband model is composed of several branches, each one containing a lowpass parallel Wiener structure with an n th-order baseband nonlinearity, as it is shown in figure.3.1. The input to the DPD is given by the complex-envelope samples of the wireless signal x[k] , and the output is the complex-envelope of the predistorted signal u[k] . The number of branches for each nonlinear order is equal to the symmetrical rank of the n th-order kernel, Rn , and the filter for the particular branch (n,r) is defined by its weights w(r) n0[k],··· ,w(r) nQn[k] . Therefore, the output of the filter at discrete time instant k is given by the convolution z(r) n[k] = w(r)T n[k]xn[k],(3.1)
3.2 Volterra-Parafac Models and Frequency-Domain Techniques 23 NL3 NLn ¦ # # NL3 > @ 1 wk > @ 1 3 wk > @ w r n k > @ 3 R 3 wk > @ r n z k > @ 1 z k > @ 1 3 z k > @ 3 R 3 z k # # > @ x k > @ u k Figure 3.1 Structure of the Volterra-Parafac baseband model. where xn[k] = [x[k],x[k−1],··· ,x[k−Qn]]T is the input vector associated to the n th-order Volterra regression term, and the weight vector w(r) n[k]is defined as w(r) n[k] = hw(r) n0[k]···w(r) nQn[k]iT.(3.2) If we write the output of the nonlinearity in the branch (n,r)as u(r) n[k] = z(r) n[k] n−1 z(r) n[k],(3.3) and collect the contribution of all the branches, the output of the VP model can be enunciated as u[k] = N ∑ n=1 0Rn ∑ r=1 u(r) n[k] = N ∑ n=1 0Rn ∑ r=1z(r) n[k] n−1 z(r) n[k].(3.4) The primma in the first sum indicates that only the contribution of the odd-order terms is included. Since the nonlinearity in each branch is already provided with the VP model, it is only necessary to identify the parameters of the filters to evaluate the predistorter output. The goal is that the output y[k] produced by the PA when it is driven by u[k] , is equal to the linearized signal ydes[k] . The complex least mean square (CLMS) algorithm can be implemented to estimate w(r) n[k] by defining the error with respect to the desired output, e[k] = ydes[k]−y[k] , and minimizing the real-valued cost function J(w[k]) = 1 2|e[k]|2 .
24 Chapter 3. Volterra-Parafac Digital Predistorters Observe that the cost function depends on the adjustable weight vector w[k], defined as w[k] = hw(1)T 1[k]w(1)T 3[k]···w(R3)T 3[k]···w(1)T N[k]···w(RN)T N[k]iT.(3.5) A necessary and sufficient condition for J(w)to have a stationary point is ∂J(w) ∂w∗=−1 2e∗[k]∂y(z) ∂w∗−1 2e[k]∂y∗(z) ∂w∗=0,(3.6) where the dependence of the output signal with the aggregated vector z[k] = hz(1) 1[k]···z(N) N[k]iT is explicited. Applying the steepest descent algorithm, we obtain the update equation for the estimated parameters: ˆ wT[k+1] = ˆ wT[k]−µ µ µ∂J(w) ∂w∗ˆ w[k] .(3.7) The diagonal matrix µ µ µ contains the step sizes µ(r) n that control the convergence speed and the steady-state properties of the CLMS algorithm for each branch (n,r) . The derivative can be computed using the chain rule by regarding y[k] as a bivariate function y(z,z∗) and treating z and z∗ as independent variables. The update equation for the branch (n,r) adopts the form w(r) n[k+1] = w(r) n[k]+ µ(r) nε(r) n[k]x∗ n[k].(3.8) The defined variable ε(r) n[k] = n+1 4e[k]|z(r) n[k]|n−1+n−1 4e∗[k]|z(r) n[k]|n−3(z(r) n[k])2(3.9) can be intuitively interpreted as the mapping of the output error e[k] onto the (n,r) -branch filter output. Once the update procedure (3.8) has converged, the system output is easily evaluated by means of the convolutions (3.1) and the nonlinear memoryless equation (3.4) . The direct application of this procedure in a sample-by-sample way is expensive from a computational cost perspective, hence the usefulness of the more efficient block processing techniques in the frequency domain becomes evident. 3.2.2 Block Processing and FFT Algorithm Applied to the VP Structure (B-VP DPD) An immediate strategy to reduce the computational complexity of the filter convolutions is to incorporate, together with the block processing technique, the fast Fourier transform (FFT) algorithms that efficiently perform the convolutions. Assuming this perspective in the estimation procedure, we can segment the input signal in blocks of L samples each, and keep the filter parameters fixed until a new block is received. For the sake of clarity, from now on the input sequence will be written as xk[l] = x[kL +l] , where l=0,1,...L−1 is the time index inside the k th block, and the filter parameters of the current block will be denoted as w(r) n,k . This approach allows to execute the update equation at a lower sampling
3.2 Volterra-Parafac Models and Frequency-Domain Techniques 25 FFT L-points IFFT L-points IFFT L-points IFFT L-points NL3 NL3 1 3 Wm 1 Wm IFFT L-points NLn r n Wm Xm r n Zm 1 3 Zm 1 Zm xl 1 zl 1 3 zl r n zl ul 3 3 R zl 3 3 R Zm 3 3 R Wm Figure 3.2 Volterra-Parafac structure and implementation of block processing with efficient frequency-domain techniques (B-VP DPD). The block index kis omitted. rate, following the recursion for the block k+1 . If the definition ε(r) n,k[l] = ε(r) n[kL +l] is used, this update equation can be written as w(r) n,k+1=w(r) n,k+µ(r) n L−1 ∑ l=0 ε(r) n,k[l]x∗ k[l].(3.10) It is well-known that block processing can be implemented in a computationally more efficient manner by performing the parameter estimation in the frequency domain using the FFT algorithm. We have selected the filter dimension equal to the block length L and followed the procedure given by the unconstrained method [25], [26]. The resulting structure of this approach is depicted in figure 3.2. The input signal x[k] is accumulated in a memory buffer to form L -point data blocks and transformed by the Fourier transform matrix, denoted as F . Let us define the input signal diagonal matrix for the current block kas Xk=diag{F[xk[0],··· ,xk[L−1]]T}(3.11) and the frequency-domain weight vector for each filter as W(r) n,k= [W(r) n,k[0],··· ,W(r) n,k[L−1]]T=Fw(r) n,k,(3.12) then, the frequency-domain output vector of the filter (n,r) can be written as Z(r) n,k=XkW(r) n,k.(3.13)
26 Chapter 3. Volterra-Parafac Digital Predistorters IFFT L-points IFFT L-points IFFT L-points M-QAM Modulator Data CP Shaping Filter CP ↑ov NL3 CP ↑ov NL3 IFFT L-points CP ↑ov NLn 1 3 Wm 1 Wm r n Wm Xm ul 3 3 R Wm Figure 3.3 A DPD with Volterra-Parafac structure for OFDM systems (FD-VP DPD). The block index kis omitted. Once the parameters W(r) n,k have been identified, the computation is implemented by using efficient frequency-domain techniques and inverse-transforming the filter output to the time domain. After applying the corresponding homogeneous n th-order baseband nonlinearity, the predistorted output uk[l] is straightforwardly computed. If this B-VP approach is used in the implementation of the DPD, the model operates over the time-domain input signal to produce the time-domain predistorted signal uk[l]. 3.2.3 Particularization to OFDM Systems (FD-VP DPD) In OFDM systems, the transmitter generates the signal by assembling blocks of the input M -QAM symbols, and mapping them into subcarriers in the frequency domain. Therefore, the use of a DPD with the output in the time domain, but operating directly on the input symbols of the k th block, denoted as Xk(m) , seems very attractive. With reference to figure 3.3, the M -QAM modulator generates the symbol sequence, and the linear branch processes these symbols by filtering, transforming them to the time domain with an inverse FFT (IFFT), adding the cyclic prefix (CP) and filtering to conform a band-limited spectrum. The modules that are represented with a dashed contour in the figure have been integrated in the linear branch of the DPD, but these elements are inherent of a typical OFDM transmitter. In this second approach, the estimation procedure can be performed following the general form of the frequency-domain update algorithm, given by W(r) n,k+1=W(r) n,k+µ µ µ(r) nX∗ kε ε ε(r) n,k.(3.14) The vector ε ε ε(r) n,k represents the frequency-domain transformation of the error vector in the (n,r)-branch output, µ µ µ(r) nis a diagonal matrix with the step sizes, and Xkrefers to (3.11).
3.3 Complexity Assessment 27 The computation of the error term used in (3.14) depends on an output error like in (3.9) , but here this error shall compensate for the frequency-domain error between the OFDM symbols passed through the PA model ( Yk ) and the linearized OFDM symbols (Xk), expressed as: Ek(m) = Yk(m)/Gc−Xk(m),(3.15) where m stands for the frequency bins ( m=0,...,L ), k is the block index, and Gc is the target linearized gain. Since the input symbols Xk(m) and the parameters W(r) n,k(m) are specified in the frequency domain, it is not necessary to compute the FFTs (3.11) and (3.12) , and the filter operation is performed by a simple product. In each nonlinear branch, an inverse transform to the time domain and insertion of the CP modules are included. In the same form, an oversampling operation is applied before the homogeneous n th-order memoryless nonlinear block. This oversampling is necessary to match the sampling rate produced by the spectrum shaping filter of the linear branch. Although in this FD-VP DPD structure, the input band is approximately equal to the bandwidth of the OFDM signal, the approach is sufficient to guarantee in-band equalization as well as minimization of the spectral regrowth in the adjacent channels. 3.3 Complexity Assessment A common measure to evaluate the complexity of a behavioral model for DPD is the number of floating point operations (FLOPs) [ 7 ]. While both the identification complexity, running complexity, and adaptation complexity may be distinguished, the running complexity is the most demanding term. Its consideration entails the complexity associated with the construction of the basis functions and filtering the basis with the model kernels. To set the notation, the complexity associated with the Fourier transform or inverse Fourier transform of a block of L points will be denoted as CFFT−L . Other elementary operations such as the complex-to-complex product, the complex-to-real product, the real-to-real product, the complex summation, and the square-modulus will be expressed in terms of their complexity as Cprodc−c , Cprodc−r , Cprodr−r , Csumc , and C|·|2 , respectively. The number of FLOPs corresponding to these elementary operations is detailed in Table 1 of [ 7 ]. As general remarks, considering that both of the proposed approaches yield a block of L samples each time, in order to compare their complexity with a conventional DPD computed in the time-domain, the per-sample complexity will be calculated by dividing the complexity associated with a block of L samples over the block size. On the other hand, let us recall that filtering a block of L points in the frequency domain with a filter of L taps involves L complex products only, instead of L complex products and L−1 complex summations that would be required, per sample, to compute the convolution in the time domain. Regarding the per-sample complexity of the static nonlinearities, for order n=2p+1 ( p=1,...,P ), p−1 real-to-real products are required, which refer to products among the square-modulus terms to yield the (n−1) -th power of the absolute value, in addition to one complex-to-real product. Let us also recall that the branch-summation operator block for both approaches involves a number of complex summations equal to the number of
28 Chapter 3. Volterra-Parafac Digital Predistorters branches minus one. In order to derive compact expressions for the complexity, the same rank Rwill be assumed for all the nonlinear orders. 3.3.1 B-VP DPD Approach For the construction of the basis functions, the following operations are required: • A Fourier transform, corresponding to the time-domain to frequency-domain conversion of the input signal. •L complex-to-complex products per order and rank, resulting from the computation of Z(r) n,kin (3.13). •An L-point IFFT per order and rank. •Lsquare-modulus per order and rank (with the exception of the linear branch). In summary, the complexity associated with the construction of the basis functions for a block of Lsamples can be expressed as: Cbasis,B−VP =CFFT−L+R(P+1)LCprodc−c+CFFT−L+LRPC|·|2.(3.16) Filtering itself corresponds to passing the block samples through the static nonlinearities plus branch summation. This filtering corresponds to a per-sample complexity given by: Cfilt,B−VP =RP 2(P−1)Cprodr−r+PRCprodc−r+ [R(P+1)−1]Csumc−c.(3.17) The per-sample FLOPs of a B-VP DPD structure of order N=2P+1 and rank R can be obtained by dividing (3.16) over the block size Land combining with (3.17): CB−VP =1 LCFFT−L+R(P+1)Cprodc−c+1 LCFFT−L+RPC|·|2+ +RP 2(P−1)Cprodr−r+PRCprodc−r+[R(P+1)−1]Csumc−c(3.18) 3.3.2 FD-VP DPD Approach Here, it is assumed that the linear processes associated with the linear branch, except the filtering by the equalizer W(1) 1,k , do not contribute to the system complexity, as they belong to the OFDM modulator, i.e. M -QAM symbol mapping, serializer-deserializer, L -point IFFT, CP insertion and spectrum-shape filtering. On the other hand, while CP insertion does not contribute to complexity itself, because it only involves reading a shifted segment of memory, it indirectly affects complexity as it changes the block size, becoming L0=L+LCP , where block-size L is augmented with the CP length ( LCP ). Similarly, oversampling does not increase complexity either, as it is performed through zero-padding, but the interpolation filter represents an additional cost per branch depending on the filter length ( M ) and oversampling rate ( ovs ), equivalent to L0·ovs ·Cfilter FLOPs, where Cfilter =M·Cprodc−c+ (M−1)Csumc . Finally, including the complexity associated
3.3 Complexity Assessment 29 with the complex multiplications required by the equalizer filters W(r) n,k , IFFT blocks of the nonlinear branches, the static nonlinearities and branch summation, the per-sample complexity associated with a FV-VP structure of order N=2P+1 and rank R can be expressed as: CFD−V P =L·R(P+1) L0·ovs Cprodc−c+R·P L0·ovsCFFT−L+ +R·PM·Cprodc−c+(M−1)Csumc−c+ +1 2R·P(P−1)Cprodr−r+P·R·Cprodc−r+ +R·P ovs C|·|2+[R(P+1)−1]Csumc .(3.19)
4 Formal Deduction of a Volterra Series Model for Complex-valued Systems 4.1 Introduction Volterra series (VS) is one of the most popular representations to model modern wireless communication systems [ 27 ]. The baseband signal is used in the modulator to generate the RF signal, possibly with linear and nonlinear impairments, and the power amplifier (PA) delivers the signal at the desired level, adding further nonlinear distortion. Without nonlinear effects, the complex envelope of the RF output signal can be viewed as a linear transformation of the complex-valued baseband signal x[k] = xI[k]+ jxQ[k] . However, to exploit the complete statistical characterization of data in complex-valued signal processing, access to the information contained in the complementary correlation is required through the introduction of widely linear (WL) transformations [ 28 ]–[ 29 ]. When nonlinear effects need to be modeled for real-valued systems, the VS approach is a natural choice, however, its extension to the complex-valued case is not a direct mapping, and particular models are usually deduced to match specific problems. Examples of these are the proposals in array processing and beamforming [ 30 ]–[ 31 ], or in the case of impairments compensation in I/Q modulators [ 32 ], but the use of these particular models is not justified to describe the output of other systems. The availability of a general complex-valued Volterra nonlinear model is a relevant topic for the design of nonlinear communication systems. In this chapter, the VS approach to a general nonlinear system with complex-valued signals by using Wirtinger calculus [33] is extended. 31
38 Chapter 5. Structured Compressed-Sensing Volterra Series Algorithm 1 Summary of the Structured Compressed-Sensing Algorithm for Volterra Series Models Input: nmax >0,α∈[0,1],f(·),X∈Cm×n,y∈Cm Output: S(t),nopt Initialization : 1: r(0)←y,S(0)←∅ 2: for t=1to nmax do 3: θt←(1−α)·maxi/∈S(t)XH {i}·r(t−1) kX{i}k2 4: i(t) pre ←*i XH {i}·r(t−1) kX{i}k2>θt+ 5: i(t)←argmini∈i(t) pre fX{i} 6: S(t)←S(t)∪i(t) 7: ˆ h←XH S(t)XS(t)−1 XH S(t)y 8: ˆ y(t)←XS(t)ˆ h 9: r(t)←y−ˆ y(t) 10: end for 11: ˆ σ2 e=ky−ˆ y(nmax)k2 2 12: nopt ←argminnc2mln ˆ σ2 e+2ncln(2m) 5.2 Structured Compressed-Sensing for Volterra Series Models The structural compressed-sensing algorithm presented in this chapter can be considered a particularization of the stagewise orthogonal matching pursuit (StOMP) algorithm, which selects a fixed number of regressors in each iteration based on a threshold. The improvement consists on the inclusion of a priority function that assigns the significance of the coefficient within the model in the subset defined after the thresholding. The new greedy algorithm for pruning Volterra model matrices taking into account the structural information is summarized in Algorithm 1. The initialization, which corresponds to line 1 of Algorithm 1, consists on the definition of the residual r(0)=y , that will be used for keeping the remaining part of the output still to be modeled. The support set S(0) is empty in the first iteration, as no regressor is still selected. In each iteration t , the algorithm calculates the correlation between the residual r(t) and each of the columns of the measurement matrix X normalized by its `2 -norm (line 3). A first preselection is performed where all the regressors with absolute value of the correlation greater than a fraction (1−α)∈[0,1] of the maximum are included in the subset i(t) pre , shown in line 4. When the span α is equal to 0, the selection becomes that of the classic OMP, which chooses only the maximum correlation within all the regressors, and if it is equal to 1, no correlation-based sorting is made and only the structural information of the model is evaluated for this arrangement. Then, the regressor with the lowest score given by the priority function f(·) is included onto the support set. Then, the estimation
5.2 Structured Compressed-Sensing for Volterra Series Models 39 of the Volterra vector ˆ h is obtained by a LS regression and the estimated output of the model ˆ y(t) and the residual ˆ r(t) are updated (lines 7-9). An example of this function is given in the experimental design section. Finally, when all the regressors are sorted or a fixed maximum of regressors to sort nmax has been reached, a Bayesian Information Criterion ( BIC ) is applied to obtain the optimum number of Volterra kernels nopt . The model with the lowest BIC is selected according to line 12, where ˆ σ2 e is the estimation of the error variance and ncis the number of components.
6 Transmitter Linearization Adaptable to Power-Varying Operation 6.1 Introduction The evident interest of the current published work about the linearization of wireless communications transmitters reveals the convenience of adequate behavioral models for the design of DPD s. The conventional baseband Volterra models, such as the FV [ 27 ], the MP [ 12 ], or the GMP [ 18 ], are specific for PA , and have insufficient accuracy to represent more general nonlinear systems with complex-valued input signals. To comply with the objective of linearizing a transmitter, designers address several challenging issues. In particular, the joint compensation of the I/Q-modulator and the PA impairments [ 41 ]–[ 42 ], and the problem of concurrent dual-band PAs [ 43 ] have been subject of concern in the last years. Another important goal is the design of a DPD that is capable of following the changes in the PA operating conditions. This has been performed by updating the DPD parameters in order to follow the power level changes [44]–[45]. Joint compensation of the I/Q modulator impairments and the PA nonlinearities are addressed in [ 46 ] by using a structure composed of a MP model dependent on the signal x(k) and a filter, augmented with a similar second branch dependent on the image signal x∗(k) . A similar approach is presented in [ 47 ] for the case of a multiple-input multipleoutput (MIMO) transmitter. The generalized twin-box model [ 42 ] incorporates parallelHammerstein (PH)-based branches that ignore the out-of-diagonal kernels and may contain non-significant diagonal terms. In these approaches, some terms belonging to the GMP or the FV models, which can be important, are not included. In [ 44 ], a power-scalable DPD based on a two-box architecture with a set of precalculated memoryless look-up table and a MP function to compensate the memory effects is proposed. The need of a power-adaptive DPD to compensate distortion in level-variable conditions was solved in [ 45 ] by adjusting the model parameters in accordance with the changes in the 41
42 Chapter 6. Transmitter Linearization Adaptable to Power-Varying Operation input power level. The procedure introduces extended scaling factors to track the varying signal power with an interpolation approach. The model extraction in both proposals requires PA measurements at multiple power levels and the number of coefficients is fixed independently of the varying PA nonlinear operation. In this chapter, we focus on a DPD design to linearize the transmitter in a range of output power levels incorporating a model with a complete set of parameters to jointly compensate the I/Q modulator impairments and the PA nonlinearities. The mathematical deduction of the general complex-valued Volterra series (CVS) representation [ 48 ] is based on Wirtinger calculus by considering the nonlinear system dependent on the input x(k) and its complex-conjugate x∗(k) , operating these variables as real-valued. Then, the transmitter can be viewed as a two-input system and analyzed using a double Volterra series approach. Under the assumption of a Volterra representation, the CVS model is able to describe the nonlinear behavior of a wireless communications transmitter. The present proposal requires signal acquisition at only a single power operating point and the identified model is directly extended to a wide range of power levels. The rise of the number of coefficients in the CVS model makes necessary a suitable procedure to identify the model parameters in an efficient and robust manner, as in [ 49 ]–[ 37 ], where a technique for model reduction using the sparse structure of Volterra kernels was introduced. A thresholding procedure contributes to the model effectiveness by reducing the number of coefficients as the system enters into weakly nonlinear modes. In [ 50 ], the robustness of the procedure is confirmed by verifying that the set of parameters identified at a given power level are applicable to the accurate estimation of the system output over a wide dynamic range. Here, the proposed approach is applied to the design of a DPD to accomplish the joint linearization of transmitter I/Q modulator and PA impairments, under power-varying conditions. The next sections of this chapter are organized as follows: After this introduction, section 6.2 presents the rationale for the proposed identification method in regression models. First, subsection 6.2.1 describes the framework of the CVS model in the context of general discrete-time complex-valued nonlinear systems. Next, subsection 6.2.2 gives a detailed theoretical justification of the proposed identification procedure and establishes the power-scalable law for the model parameters. Subsection 6.2.3 reviews related works for the joint compensation of I/Q modulator and PA impairments. 6.2 DPD Structure and Identification Procedure Compensation of transmitter RF impairments, originated predominantly by I/Q imbalance and nonlinearities, is one of the most important challenges for DPD designers. This problem can be approached by following two different points of view. The first approach is based on the knowledge of how the different blocks are assembled inside the transmitter (see figure 6.1a) and an approximate input-output representation is deduced for the whole system. An alternative perspective, where the predistorter is designed considering the transmitter as a black box, is presented below. For comparison, the first approach is discussed at the end of this section.
6.2 DPD Structure and Identification Procedure 43 s1(k) ↓ ĥ(1) s2(k) ↓ ĥ(2) s3(k) ↓ ĥ(3) … sN(k) ↓ ĥ(N) k y(k) I/Q Modulator PA Transmitter under test (TUT) xy e y0 (a) (b) Figure 6.1 Acquisition model with additive noise (a) and signal segmentation (b). 6.2.1 Proposed DPD Model for Joint Mitigation in Transmitters In the context of Volterra series representation, any given system with complex-valued input can be modeled by the CVS model [ 48 ]. Viewed as a black box, the relationship for the input and output complex envelopes, x(k) and y0(k) , of the DPD in a wireless communications system can be expressed as y0(k) = h0,0+ ∞ ∑ n=1(n ∑ m=0 Q ∑ qn=0 Q ∑ pm=0 hn−m,m(qn−m,pm)× × n−m ∏ r=1 x(k−qr) m ∏ s=1 x∗(k−ps)).(6.1) The vector of delays are qn= [q1,q2,··· ,qn]T and pm , defined in a similar way, with a maximum delay Q for all indices. The product of the input signal samples is denoted as ∏n r=1x(k−qr) = x(k−q1)x(k−q2)···x(k−qn) and the same notation is used for the product of the image samples x∗(k) . For m=0 and m=n , hn,0(qn) and h0,n(pn) are standard Volterra kernels and the products do not contain x∗(k−ps) and x(k−qr) , respectively. The factor hn−m,m(qn−m,pm) is a Volterra kernel of order n dependent on the two kinds of indices. The CVS model (6.1) with h0,0=0 can be viewed also as a linear combination of the Volterra regressors ∏n−m r=1x(k−qr)∏m s=1x∗(k−ps) . The M samples of the input signal can be disposed to form the column vector x= [x(0),x(1),··· ,x(M−1)]T and, if the Volterra regressors are likewise organized to construct the regressor vectors ξ ξ ξi , the truncated version of (6.1) can be rewritten as y0= NR ∑ i=1 hiξ ξ ξi,(6.2)
44 Chapter 6. Transmitter Linearization Adaptable to Power-Varying Operation where the i th regression coefficient hi comes from the Volterra kernels arranged in an ordered-fashion, y0 is a vector with the output samples disposed in the same way as x , and NR is the number of regressor vectors of the model. Examples of Volterra regressor vectors are ξ ξ ξ1=x , corresponding to the memoryless linear regressor, and ξ ξ ξi0= [x(−q)|x(−q)|2,x(1−q)|x(1−q)|2,··· ,x(M−1−q)|x(M−1−q)|2]T , corresponding to the third-order memory polynomial regressor with a delay q . The regressor vectors constitute a complete set in the case of the CVS model. Although ξ ξ ξi are non-orthogonal, a feasible recursive algorithm to estimate the NR coefficients of (6.2) is the Orthogonal Matching Pursuit (OMP) [51]. The so-called observation matrix X and the coefficients vector h are constructed by stacking all column vectors ξ ξ ξi and the model coefficients hi , respectively. Therefore, (6.2) can be expressed in a compact matrix form [18],[37]. 6.2.2 Identification Procedure In an experimental setup, the vector of the acquired samples, y , is contaminated by equipment distortion and noise. Assuming an approximately distortionless setup, for instance in laboratory conditions with high-performance equipment, the schematic of the acquisition experiment can be modeled as in figure 6.1. Therefore, the input-output relationship is y=Xh +e,(6.3) with a zero-mean additive white Gaussian (AWGN) noise e , stationary and complex-valued. The model parameters can be identified by using a direct LS solution, given by ˆ h=XHX−1XHy+w,(6.4) where H represents the Hermitian transpose operation. The difference between the estimated vector ˆ h and the true vector h is originated by the noise process w . Observe that the selection of a model with lacking regressors is another important source of identification error. For example, the FV and the GMP representations, as particular cases of (6.1) , are insufficient to model a transmitter with impairments. The general character of the CVS model provides a complete set of regressors with a high number of coefficients and, at the same time, an associated suitability for a pruning procedure without a presumption on the significant regressors. The application of the Bayesian information criterion (BIC) was proposed in [ 37 , 50 , 52 ] to select the significant parameters. Whereas [ 52 ] is based on a simulated annealing algorithm with all the possible variants of the model, the approach in [ 37 ] is based on the application of the OMP algorithm to represent the PA output as its projection onto the span of Volterra regressors, and on the BIC rule to discard the irrelevant coefficients, maintaining only the active regressors. It is worth noticing that while LS identification is affected by regressor correlation, the OMP algorithm guarantees the recovery of the exact value of the coefficients in a noiseless environment in a given number of iterations [53]. The alternative procedure applied in [ 50 ] is mainly a parameters detection based on statistical hypotheses testing. Referred to figure 6.1b), the entire set of acquired samples is divided in N segments. Focusing on the i th coefficient of the estimated vector ˆ h , each
6.2 DPD Structure and Identification Procedure 45 segment returns an independent measurement of this particular coefficient. We can define the vector ˜ hi constructed with the N different realizations of this random variable to decide whether or not the i th coefficient has to be incorporated as an active parameter of the model. We use the Neyman-Pearson (NP) approach to make the decision, based on two hypotheses [54]: •The measurement is produced by noise, H0: if ˜ hi=wi, •The measurement is produced by the presence of a model coefficient plus noise, H1: if ˜ hi=hi+wi. The probability density function under H1is p(˜ hi;H1) = 1 πNσ2Ne−1 σ2(˜ hi−hi)H(˜ hi−hi),(6.5) where σ2 is the variance of the complex-valued AWGN wi . Likewise, the probability density function under H0is expressed as p(˜ hi;H0) = 1 πNσ2Ne−1 σ2˜ hH i˜ hi.(6.6) For a particular probability of erroneous measurement induced by noise, the NP theorem states that the probability of true detection is maximized if the detector decides the hypothesis H1when the likelihood ratio L(˜ hi)exceeds a given threshold γ, i.e., L(˜ hi) = p(˜ hi;H1) p(˜ hi;H0)>γ.(6.7) Substituting (6.5) and (6.6) in the likelihood ratio and taking lnL(˜ hi), we obtain lnL(˜ hi) = 2 σ2RehH i˜ hi−1 σ2|hi|2(6.8) and the decision Re(hH i˜ hi) = Re N ∑ r=1 h∗ i(r)˜ hi(r)!>γ0(6.9) is equivalent to the likelihood ratio test (6.7) . Under steady conditions, any coefficient is an unknown constant hi(r) = hi, with its estimator being the average value hi≈¯ ˜ hi=1 N N ∑ r=1 ˜ hi(r).(6.10)
46 Chapter 6. Transmitter Linearization Adaptable to Power-Varying Operation Table 6.1 Summary of the Proposed Identification Procedure 1: Divide the input x(k) and measured output y(k) signal into N segments, each with M samples. 2: Estimate the parameter vectors ˆ h(r) , r=1,...,N using the OMP algorithm with the same pre-defined NRnumber of components for all models. 3: For all the CVS components, compute the average coefficients ¯ ˜ hi. 4: Sweep the thresholding level to apply the hypothesis test to all coefficients. 5: Take the thresholding level that optimizes the BIC rule. Operating with (6.9), we can decide H1if Re ¯ ˜ h∗ i 1 N N ∑ r=1 ˜ hi(r)!=Re¯ ˜ h∗ i¯ ˜ hi=|¯ ˜ hi|2>γ0 N.(6.11) There is a trade off between the deficit of model coefficients (if the threshold is high), and the inclusion of undesired noise (if the threshold is chosen too low). The proposed procedure is as follows. Once the set of coefficients has been estimated by the OMP algorithm for each segment of the acquired signal (see figure 6.1b), the average ¯ ˜ hi and the hypothesis test (6.11) are computed for all coefficients. The lower the threshold level, the more coefficients (regressors) will be incorporated to the model and the BIC criterion is used to decide the optimum number of regressors nc [ 37 ]. Due to the sparse character of the Volterra kernels, many coefficients can be discarded without a significant loss in accuracy. Here the BIC rule with its explicit dependence on the NMSE . If the NMSE is expressed in dB, the variance is given by ˆ σ2 e= 1 M M−1 ∑ m=0 |y(m)|2!×10NMSE/10 (6.12) and substituting in (23) of [37], the BIC rule becomes nc0=argmin ncnNMSE +nc M10log(2M)o.(6.13) This procedure, summarized in Table 6.1, and that published in [ 37 ], start with the OMP, a method that estimates the set of coefficients by iteratively adding new components to the coefficients vector. The difference is that in this work the coefficients computed with several segments of the signal are averaged and a further thresholding step is implemented to get the sparse model structure. According to the results in [ 50 ], the normalized coefficients originated in the nonlinearities of the I/Q branches are not dependent on level variations at the modulator output and the coefficients associated to the FV regressors are normalized according to the PA input level [ 37 ]. For example, if h(n)Pi0 is a FV n th-order normalized coefficient ( n odd) for an
6.2 DPD Structure and Identification Procedure 47 input signal with a power level Pi0 and this level changes to P=A2Pi0 ( A>0 , real-valued), the corresponding normalized coefficients follow the relation h(n)P h(n)Pi0 =1 An =Pi0 Pn/2 .(6.14) Once the normalized parameters of the DPD have been computed at a given input level, Pi0 , they can be straightforwardly scaled to adapt the coefficients to other level P . Odd-order normalized parameters shared with the FV model follow an exponential scaling with the average input power, so that the magnitudes at Pi0and Pare related as: h(n)P=h(n)Pi0 ·10−n·∆P(dB) 20 ,(6.15) where ∆P(dB) = P(dBm)−Pi0(dBm) and n is the coefficient order. Those normalized coefficients not included in the FV model (e.g., the image or the even-order parameters associated with impairments of the I/Q modulator), are not dependent on the power level. In that case, the way the DPD is adapted to a decrement in power level is by excluding parameters with values below a given threshold, and not recalculating the coefficients [44, 45]. 6.2.3 Preceding DPD Models for Joint Mitigation In the previous procedure, the transmitter is viewed as a black box and the CVS model is adopted for the DPD. Another perspective is to assume the internal architecture of the transmitter (figure 6.1a) and deduce the DPD structure to join the mitigation of PA and I/Q-modulator impairments under reasonable approximations [ 41 ]-[ 47 ]. In [ 41 ], the rationale is to compensate the impairments in the reverse order that they appear. Based on the knowledge of the ensemble arranged by the modulator and the PA, a first MP structure is proposed for the PA DPD and then the modulator compensation is implemented with a widely linear (WL) model. This point of view is further extended to a general FV model [ 46 ] and to the augmented complex conjugate (ACC) model for joint mitigation of distortion in MIMO transmitters [ 47 ]. The output for the ACC model in a single-input single-output (SISO) transmitter is written as y(k) = N ∑ n=1 0¯ Hn{x(k)}+ N ∑ n=1 0¯ Hcn{x∗(k)}+h0,(6.16) where the prima symbols in the sums indicate that only odd orders are considered, ¯ Hn{x}= Q ∑ q=0 hn(q)|x(k−q)|n−1x(k−q)(6.17) and ¯ Hcn{x∗(k)}is defined in the same form.
54 Chapter 7. Results 3570 3580 3590 3600 3610 3620 3630 −70 −60 −50 −40 −30 −20 −10 0 Frequency (MHz) Normalized Power Spectral Density (dB/Hz) W/o DPD GMP DPD FD−VP DPD Figure 7.4 Measured power spectral density of the test PA with different DPD linearizers. The unpredistorted PA was driven at an input level to yield the same output power than the linearized PA (+26.2 dBm). a similar performance to the DPD based on the GMP model, while significantly reducing the computational cost per sample associated to the generation of the predistorted signal. The results of the ACPR confirm that both the FD-VP DPD and the GMP DPD provide an improvement over 17 dB with respect to the nonlinear signal at the same average output power level. Despite the fact that the spectral regrowth of the un-predistorted PA signal failed to comply with the requirements of the LTE-downlink standard, the ACPR requirements were fulfilled after applying both DPDs. Figure 7.4 shows the power spectral density of the LTE signal with and without DPD, demonstrating the reduction of the spectral regrowth and the similarity between the results obtained by both DPDs. It is worth mentioning that these results were achieved by deriving the DPD settings from the aforementioned adaptive algorithm, but using a single estimate of PA model. Although both approaches led to similar values, there is room for improvement in the reduction of the ACPR by using more accurate estimates of the PA model in an iterative fashion. In-band distortion mitigation of the DPDs can be assessed in terms of the EVM. Table 7.3 reveals that the proposed DPD reduces the EVM from 3.0% to 0.6% , while the GMP DPD produces a higher dispersion of the received symbols around the reference constellation points, with an associated EVM of 1.4% . The more notable mitigation of the in-band distortion for the proposed DPD can be observed in figure 7.5, where one symbol of the 16-QAM constellation of the received LTE signal has been zoomed in. Finally, the NMSE between the sampled signal in the receiver and the theoretical reference is calculated to evaluate the overall linearization achieved. As it can be observed in Table 7.3, the residual error of the LTE signal significantly decreases with the proposed DPD, reducing the NMSE
7.2 Complex-valued Volterra Series Model 55 Figure 7.5 16-QAM constellation of a received 15-MHz LTE-downlink signal without DPD (red dots), with a GMP DPD (blue dots), and with the proposed FD-VP DPD (black dots). The reference constellation point is plotted with a green ‘+’ mark. One symbol has been zoomed in. Table 7.3 DPD performance comparison for a 15-MHz LTE-downlink signal ( Pout = +26.2 dBm) Cases under NMSE ACPR (dBc) EVM Computational study (dB) Lower Upper (%) cost (FLOPs/sa) W/o DPD −30.2−39.5−39.6 3.0− GMP DPD −41.7−57.2−58.0 1.4 939 FD-VP DPD −44.5−56.6−56.9 0.6 255 of the nonlinear signal by more than 14 dB. Regarding the complexity of the DPD data generation, the cost per sample associated with (3.19) is 255 FLOPs under the aforementioned settings. This figure can be fairly compared with the particular implementation of this GMP DPD, as shown in Table 7.3: it would require 939 FLOPs per sample. 7.2 Complex-valued Volterra Series Model An experimental study based on the testbench introduced in section 2.4.1 has been conducted. The VSG served as an I/Q modulator, providing a flexible interface to introduce
56 Chapter 7. Results Table 7.4 Modeling performance in terms of the NMSE and ACEPR for the test PA with impairments in the I/Q modulator Model ACEPR (dBc) NMSE # −30 MHz −15 MHz In-Band +15 MHz +30 MHz (dB) Coeff. FV-CS −62.8−54.7−31.7−54.6−66.4−31.616 WNL-CS −68.8−65.2−54.5−65.1−68.2−52.7129 I/Q impairments. The VSG output was fed into a commercial PA constructed with the CGH40010 GaN HEMT transistor, from Cree Inc. The test signal was a 15-MHz orthogonal OFDM signal generated at 3.6 GHz according to the Long Term Evolution ( LTE ) downlink standard. In the receiver side, the RF signal was measured at the VSA with a sampling rate of 92.16 MS/s providing over 360000 samples. The test signal was composed of 56 OFDM symbols and model identification was performed with only one symbol. Following the usual methodology for behavioral modeling of transmitter architectures, model parameters were identified by applying a conventional LS procedure to the inputoutput measurement datasets, providing the NMSE between the modeled and measured signal as a quality metric, in addition to the Adjacent Channel Error Power Ratio ( ACEPR ), the latter evaluated for both the upper and lower adjacent channels. Modeling performance of the proposed WNL approach was compared to the FV, given its demonstrated accuracy in the context of PAs. Considering the PA operation point, both models were configured with thirteenth-order, a memory length Q=3 for orders one to five, and memoryless kernels for higher orders. Due to the relatively large number of coefficients in the general proposal (248 for the FV model and 1370 for the WNL model), the compressed-sensing algorithm [ 37 ] was applied to select the appropriate kernels of both models. According to Table 7.4, under I/Q modulator impairments, the FV compressed-sensing (FV-CS) model reduces the number of coefficients to 16 and degrades the NMSE to −31.6 dB, while the pruned WNL-CS model consists of 129 coefficients and presents an NMSE of −52.7 dB that outperforms the former in about 20 dB. While the number of the WNL-CS coefficients can be high, an increase of the FV-CS coefficients leads to overfitting because the appropriate regressors are missing. This behavior is confirmed by the predicted spectra plotted in figure 7.6, showing that the model tightly matches the measured signal, while the FV-CS model fails in the description of spectral regrowth in the adjacent channels. These outcomes are even more evident in figure 7.7, where the spectra of the error are depicted. As a reference we include in this figure the results for the test PA without I/Q modulator impairments (dash-dot line), showing a very low spectrum error inside the signal band for the FV-CS model. Let us clarify that impairments were removed from the I/Q modulator output by using a precompensator based on [ 32 ], and therefore this case resembles a nonlinear system for which the baseband PA model holds. However, when I/Q impairments are present, the FV model cannot describe accurately the output and the in-band error degrades by almost 30 dB. On the contrary, the WNL preserves the performance in that case. The spectrum of the error predicted by the WNL-CS under compensated I/Q impairments, not shown, is almost the same as the FV-CS model.
7.3 Structured Volterra Series Model 57 3570 3580 3590 3600 3610 3620 3630 −80 −70 −60 −50 −40 −30 −20 −10 0 10 Frequency (MHz) PSD (dB/Hz) Meas. Data FV−CS WNL−CS Figure 7.6 Normalized power spectral density of the measured (dotted line) and modeled (solid lines) output of the PA driven at 3.6 GHz with an impaired I/Q modulator. 7.3 Structured Volterra Series Model A set of measurements were acquired in order to validate the predistortion capabilities of the algorithm. The weblab introduced in section 2.4.2 was used. The measurements were taken at an output power of +35.2 dBm. The test signal is compound of 56 OFDM symbols of a 15-MHz OFDM signal generated from 16-QAM symbols modulated onto 900 subcarriers and filtered by a raised-cosine with roll-off factor of 0.1, according to the LTE-downlink standard. This signal exhibits a PAPR of about 11 dB and a hard clipping to the 7 samples with the highest absolute value was applied to reduce the PAPR to 10 dB. An oversampling of 1 to 6 applied to the original signal results in a sampling frequency of 92.16 MHz which was also used to record the measurements. The test signal contains over 360000 samples. A GMP model was selected to test the algorithm. This model has the structure shown in (7.1) , where a configuration of seventeenth order and a maximum distance from the diagonal of 10 was set. This corresponds to Ka=Kb=Kc=16 , La=10 and Lb=Mb=
58 Chapter 7. Results Frequency (MHz) 3570 3580 3590 3600 3610 3620 3630 Error PSD (dB/Hz) -80 -70 -60 -50 -40 -30 -20 -10 0 10 FV-CS (w/o IQ impair.) FV-CS (with IQ impair.) WNL-CS (with IQ impair.) Figure 7.7 Spectrum of the error signal between the measured and modeled output of the PA driven at 3.6 GHz with an impaired I/Q modulator (solid lines) and an I/Q modulator with precompensated impairments (dot-dashed line). Lc=Mc=5.The resulting Volterra model contains 1087 components. y(k) = Ka ∑ p=0 La ∑ l=0 aplx(k−l)|x(k−l)|p+ + Kb ∑ p=1 Lb ∑ l=0 Mb ∑ m=1 bplmx(k−l)|x(k−l−m)|p+ + Kc ∑ p=1 Lc ∑ l=0 Mc ∑ m=1 cplmx(k−l)|x(k−l+m)|p. (7.1) The thresholding function designed for this experiment was f(l,m,p) = |l∓m|+(p+1) according to the form of the GMP regressors of x(k−l)|x(k−l±m)|p . This function assigns a higher score —which corresponds to less priority— to high orders and lags, considering a memory fading behavior [ 57 ]. The number of maximum components to be considered nmax was set to 200. The predistorter was calculated through the indirect learning scheme [ 10 ]. The 30% of the signal with the highest maximum value at the output of the PA was used for modeling and the span value from 0 to 1 with an increment of 0.2 was swept. Once the predistorter was identified, the validation was carried out with the complete signal and the predistorted signal was obtained placing the previous PA input signal at the input of the DPD. The
7.4 Transmitter Linearization Adaptable to Power-Varying Operation 59 output of the DPD was then sent to the web platform and the performance parameters of NMSE, ACPR and EVM were measured in the returned signal. Table 7.5 Performance results of DPD in a sweep of span αvalues. ACPR-2 ACPR-1 ACPR+1 ACPR+2 EVM NMSE # case (dBc) (dBc) (dBc) (dBc) (%) (dB) Coeff. w/o DPD -38.04 -32.82 -32.93 -38.00 5.00 -25.52 - α 0.00 -39.76 -36.83 -36.82 -39.80 2.82 -30.46 200 0.20 -42.57 -40.46 -40.51 -42.75 1.89 -33.90 198 0.40 -44.15 -41.35 -41.29 -44.17 1.78 -34.49 190 0.60 -44.65 -42.92 -42.91 -44.74 1.46 -36.18 148 0.80 -50.04 -48.58 -48.71 -50.24 0.80 -41.52 17 1.00 -55.09 -53.88 -54.13 -55.37 0.52 -45.62 194 The performance parameters of the DPD are presented in Table 7.5. The ACPR, EVM and NMSE values experience a decrease with the span, which indicates that the introduction of the span enhances the predistortion capabilities of the model. The signal without predistortion is characterized by a first lower and upper ACPRs of −32.82 dBc and −32.93 dBc. The predistortion enhances this values from about 4 dB for α=0 to 21 dB in the case of α=1 , where the minimum values of −53.88 dBc and −54.13 dBc are reached. Similar improvements are achieved for the second lower and upper ACPRs, where the minimum is also obtained in the same case and take the values of −55.09 dBc and −55.37 dBc, corresponding to an improvement of about 17 dB with respect to the case without DPD. These values are in accordance with figure 7.8, where the power spectral density (PSD) of the predistorted signal for the best case is compared to the case without DPD. The reduction of the spectral regrowth achieved by the DPD agrees with the PSD of the error signal, which shows a low in-band and out-of-band error densities. The EVM is decreased from a 5 % to a 2.8 % for α=0 and a 0.52 % in the case of α=1 . This enhancement is clear in the constellation shown in figure 7.9, where the predistorted case shows a reduction in dispersion compared to the signal without predistorsion. The NMSE is reduced from −25.52 dBm without predistortion to −30.45 dB for the lowest value of the span and to −45.62 dB for α=1 . Finally, the number of components in the pruned model results in roughly the same value for all the range of the span, which indicates that the pruning capabilities of the algorithm remains at the same performance level. 7.4 Transmitter Linearization Adaptable to Power-Varying Operation To illustrate the proposed method, the joint compensation and linearization of two transmitters, with adaptable capability in a wide dynamic range, is demonstrated. 7.4.1 Linearization of a Basic Transmitter The first case of study is the basic transmitter modeled in [ 50 ], referred to as TUT-1. The TUT-1 was arranged with the I/Q modulator integrated in the commercial generator
60 Chapter 7. Results Frequency (MHz) 3560 3580 3600 3620 3640 PSD (dB/Hz) -60 -50 -40 -30 -20 -10 0w/o DPD DPD α = 1 Frequency (MHz) 3560 3580 3600 3620 3640 Error PSD (dB/Hz) -65 -60 -55 -50 -45 -40 -35 -30 -25 w/o DPD DPD α = 1 Figure 7.8 Spectral densities for one realization of the experiment for the cases without DPD and with DPD in the case of α=1. SMU200A of Rohde & Schwarz, and a PA based on Cree’s board for the evaluation of the power GaN HEMT CGH40010, operated at a carrier frequency of 3.6 GHz. Later on a DPD linearizer for a more realistic transmitter (TUT-2) will be tested. The probing signal was designed with an OFDM format and 15-MHz bandwidth, according to the LTE-downlink standard. The input vector x containing over 300000 samples of the complex envelope was uploaded into the generator and transmitted with
7.4 Transmitter Linearization Adaptable to Power-Varying Operation 61 Figure 7.9 15-MHz bandwidth signal constellation for the original signal without DPD and with DPD in the case of α=1. a peak power of about +30 dBm, corresponding to a measured output average level of +19 dBm. Since the peak-to-average power ratio (PAPR) level is 11 dB, the upper limit to the transmitter power is fixed by the maximum modulator peak level, about +15 dBm (+6 dBm of average power). The samples of the output signal were gathered by averaging 300 acquisitions in the vector signal analyzer. Driven with this signal, the TUT-1 serves as a first proof of concept of the proposed linearization approach. To demonstrate the DPD capability to compensate the nonlinerities as well as the I/Q impairments, the settings of the modulator were configured with a quadrature error of 1 deg. Whereas commercial I/Q modulators show quadrature errors in the range of 1–3 deg [ 58 ]–[ 59 ], the choice of 1 deg was intended to show to which extent a reduced error can have a significant impact on the linearization capability of a DPD. The acquired baseband samples were subsequently off-line postprocessed with Matlab and used in a conventional indirect learning architecture to obtain the DPD coefficients at a maximum average output power of +20.8dBm (+31.8dBm of peak output power). The model parameters were chosen to reduce the error between y(k)/G , the measured output scaled with the target gain of the linearized TUT-1, and x(k) . The underlying CVS structure was configured with thirteenth order, a maximum delay Q=3 for orders 1 to 5, and memoryless (ML) for the higher orders, comprising 1370 coefficients. The OMP algorithm was executed to provide the ordering of 200 coefficients per segment. The synthesized DPD was composed of 22 normalized coefficients above a selected −68 dB
62 Chapter 7. Results Ouput level (dBm) 8 10 12 14 16 18 20 22 ACPR (dB) -75 -70 -65 -60 -55 -50 -45 ACPR1 and ACPR2 w/o DPD ACPR1 and ACPR2 with DPD Figure 7.10 Adjacent channel power ratios of the output spectrum with (solid lines) and without (dashed lines) DPD. First adjacent channel ( + marks) and second adjacent channel (×marks). Basic transmitter TUT-1. threshold. The results of the linearization produced by the DPD at the output power of +20.8 dBm are plotted with filled circles in figure 7.10, where an ACPR of about −66 dB in the first adjacent channel is satisfactorily compared to the value of −46 dB obtained with the non-linearized transmitter. A reduction of about 10 dB in ACPR is also observed in the second adjacent channel. Likewise, about 6 dB of NMSE reduction was achieved with the TUT-1 driven at the maximum level, as it is shown with filled marks in figure 7.11. The normalized coefficients of the DPD were extended at other power levels following the explained procedure, and those below the threshold were discarded. A second series of measurements was then performed with the new predistorted signals calculated according to the model adjusted at the different output levels, yielding the ACPR displayed with solid lines in figure 7.10. To have a reference of the DPD performance, the ACPRs of the transmitter without DPD are also shown (dashed lines), demonstrating a reduction over 21 dB in the first adjacent channel at Po= +20.8 dBm. The DPD performs in all the dynamic range, demonstrating values of ACPR 1 and ACPR 2 better than −65 dB and −70 dB, respectively, without any further adjustment of the coefficients. This capability is also observed in the results of the NMSE displayed in figure 7.11 for the transmitter without DPD, in dashed line, and with DPD, in solid line. NMSE levels of about −41 dB in a wide range of operating powers demonstrate a flexible functioning of the linearizer. The results of the DPD performance for the TUT-1 were encouraging and led to apply the procedure to a new transmitter where the power capability of the output device was exploited thoroughly.
7.4 Transmitter Linearization Adaptable to Power-Varying Operation 63 Ouput level (dBm) 8 10 12 14 16 18 20 22 NMSE (dB) -42 -41 -40 -39 -38 -37 -36 -35 -34 -33 -32 w/o DPD with DPD Figure 7.11 NMSE of the output signal without DPD (dashed line) and with DPD (solid line). Basic transmitter TUT-1. 7.4.2 Linearization of a Realistic Transmitter The objective of this subsection is the joint compensation and linearization of a transmitter operating near its maximum level of +40 dBm. In this transmitter, referred to as TUT-2, a ZHL42W preamplifier of MiniCircuits is connected at the input of the Cree’s evaluation board to drive it to a higher nonlinear operating point and delivering a maximum average output power of +26 dBm ( +37 dBm of peak power), with a small-signal chain gain of about 46.7 dB. Again, 1 deg of quadrature error was considered for the I/Q modulator. The AM/AM and AM/PM characteristics are plotted in figure 7.12, revealing a gain compression of about 1.6 dB. The output spectrum is shown in figure 7.13, where we observe an important spectral regrowth that fails to comply with the standard ACPR of −45 dB. Let us also remark that not only the PA is contributing to the nonlinear distortion, but the I/Q modulator is also driving the preamplifier beyond its linear operation level. Again, a thirteenth-order CVS structure with 1370 coefficients was pruned. After repeating the procedure described in section 7.4.1, the corresponding DPD was designed. The linearization capability is also shown in the same figure, demonstrating a spectral regrowth reduction of about 20 dB in the first adjacent channel. Although a second iteration will usually provide better results, this initial design is adequate enough for the objectives of the present experiment. The normalized magnitudes of the 23 most relevant coefficients at Po= +26 dBm ( Pi=−20 dBm) are plotted with large marks in figure 7.14 and, in agreement with our previous discussion, the values adapted to other levels are displayed with straight lines. Notice that all the values are referred to the ML linear coefficient and that the number of parameters can be reduced as the PA is operated at lower levels. Since the precision requirements allow neglecting the normalized coefficients below a −50 dB threshold, the model at a PA input of Pi=−20 dBm is composed of only 14 coefficients, a value that
70 Chapter 8. Conclusions and future work Next, in chapter 4, a formal deduction of a Volterra model for complex-valued nonlinear systems has been analyzed. It has also been demonstrated that the input-output relationships derived ad-hoc for several communication systems –a power amplifier, a beamformer, an I/Q modulator– are particular cases of the CVS model. To manage the increment of the parameters, the kernel sparsity in this approach has been exploited and the structure complexity with the proposed WNL-CS model has been reduced. Through experimental validation, it has been also shown that this model is able to accurately represent a diverse range of distortion sources in communication subsystems, including impairments at the I/Q modulator combined with the nonlinear distortion associated to PAs. Although conventional behavioral models are accurate for the description of nonlinearities in PAs, they show a degraded performance when the input signal is impaired by a quadrature modulator. On the contrary, the richness of its regressors makes the proposed model a robust representation of a general transmitter-receiver architecture. In chapter 5, an improved method for the sparse recovery of Volterra series models has been presented. It adds structural information of the PA model in the selection process to achieve a reduction in the optimum number of coefficients while maintaining the fidelity. The benefits of the introduction of the span parameter has been experimentally demonstrated. This method is susceptible to be applied to any kind of Volterra series behavioral model. Predistortion with models given by the algorithm have been performed showing a high reduction of the complexity of the model with the same level of performance that the ones given by the complete model before pruning. Finally, in chapter 6, a reliable procedure to identify the model coefficients of a powerscalable DPD for the joint compensation of wireless communications transmitters has been reported. It has been applied to a DPD based on the CVS model and has been approvingly compared with other recent DPD proposals that have been advanced to overcome the limitations of conventional models under the presence of I/Q modulator impairments. The proposed DPD design uses a regressor search algorithm and a threshold to limit the number of parameters, allowing a significant reduction guaranteed by discarding negligible coefficients. Several causes are behind the model-order reduction: irrelevant kernel types, truncated nonlinear order, truncated memory depth, sparse memory delays, etc. The coefficients were identified at a power level where the PA is near saturation, and the invariance of the denormalized kernels in the measured dynamic range was exploited to estimate the values at other operating conditions without the need of interpolation or a new series of measurements. Tailoring of the number of parameters at other levels is performed with a simple rejection of regressors with normalized coefficients below the predetermined threshold. In consequence, the number of coefficients is further reduced as the transmitter enters into the less nonlinear operating range. In terms of model-order reduction and precision, a better performance with respect to other alternative approaches has been demonstrated in a wide range of measured power levels below the DPD upper operating point.
8.2 Future Work 71 8.2 Future Work The work presented in this Thesis provides solutions to some of the practical problems in PA and transmitter modeling and linearization under sparsity hypothesis and can be extended in any of the following ways. The use of iterative signal processing techniques for the sparse regression of Volterra models has a natural continuity with the use of machine learning and deep learning techniques. Although it has been proved that compressive-sensing techniques allow to recover a sparse regression that works with the same level of performance than those attained by the LS regression, the use of artificial intelligence techniques as those afore-mentioned are promising in this field. The adaptation of multi-layered recognition algorithms to the unsupervised recovering of DPDs is an interesting field still to be discovered. 8.3 Contributions This Thesis is based on the following list of contributions: [2017a] C. Crespo-Cadenas, M. J. Madero-Ayora, J. Reina-Tosina, and J. A. Becerra , “Formal deduction of a Volterra series model for complex-valued systems,” Signal Processing, vol. 131, no. Supplement C, pp. 245 – 248, 2017. [2017b] Crespo-Cadenas, C., M. J. Madero-Ayora, Reina-Tosina, J., and J. A. Becerra , “Transmitter linearization adaptable to power-varying operation,” IEEE Transactions on Microwave Theory and Techniques, vol. 65, no. 10, pp. 3624–3632, Oct 2017. [2017c] J. A. Becerra , M. J. Madero-Ayora, J. Reina-Tosina, and C. Crespo-Cadenas, “Digital predistortion of power amplifiers using structured compressed-sensing Volterra series,” Electronics Letters, vol. 53, no. 2, pp. 89–91, 2017. [2017d] M. J. Madero-Ayora, C. Crespo-Cadenas, J. Reina-Tosina, and J. A. Becerra , “Effect of an impaired I/Q modulator in the structure of a behavioral model for power amplifiers,” in 2017 IEEE 18th Wireless and Microwave Technology Conference (WAMICON), April 2017, pp. 1–4. [2017e] C. Crespo-Cadenas, J. Reina-Tosina, M. J. Madero-Ayora, and J. A. Becerra , “Identification of volterra model parameters in wireless systems,” in 2017 IEEE Topical Conference on RF/Microwave Power Amplifiers for Radio and Wireless Applications (PAWR), Jan 2017, pp. 96–99. [2017f] M. J. Madero-Ayora, C. Crespo-Cadenas, J. Reina-Tosina, and J. A. Becerra , “Relationship Between a Pruned Volterra Model Structure and Impairments in an I/Q Modulator,” in 47th European Microwave Confecence, Oct 2017. [2016a] C. Crespo-Cadenas, M. J. Madero-Ayora, J. Reina-Tosina, and J. A. Becerra , “Joint compensation of modulator and power amplifier nonlinearities based on a complexvalued Volterra model,” in 2016 46th European Microwave Conference (EuMC), Oct 2016, pp. 1067–1070.
72 Chapter 8. Conclusions and future work [2016b] J. A. Becerra , M. J. Madero-Ayora, J. Reina-Tosina, and C. Crespo-Cadenas, “Podado de modelos de Volterra con técnicas de muestro compresivo.” in XXXI Simposium Nacional de la Unión Científica Internacional de Radio (XXXI Spanish National Symposium of the International Scientific Union of Radio), September 2016. [2015a] J. A. Becerra , Madero-Ayora, M. J., Reina-Tosina, J., and Crespo-Cadenas, C, “Modelo ampliamente no lineal para la compensación de moduladores I/Q,” in XXX Simposium Nacional de la Unión Científica Internacional de Radio (XXX Spanish National Symposium of the International Union of Radio Science), September 2015. [2015b] C. Crespo-Cadenas, M. J. Madero-Ayora, J. Reina-Tosina, and J. A. Becerra , “A widely nonlinear approach to compensate impairments in I/Q modulators,” in 2015 European Microwave Conference (EuMC), Sept 2015, pp. 506–509. [2015c] M. J. Madero-Ayora, J. Reina-Tosina, C. Crespo-Cadenas, and J. A. Becerra , “On the robustness of a widely nonlinear approach to model impairments in I/Q modulators,” in 2015 Integrated Nonlinear Microwave and Millimetre-wave Circuits Workshop (INMMiC), Oct 2015, pp. 1–3. [2014a] C. Crespo-Cadenas, J. Reina-Tosina, M. J. Madero-Ayora, and J. A. Becerra , “Volterra-parafac digital predistorter for OFDM wireless systems,” in 2014 44th European Microwave Conference, Oct 2014, pp. 829–832. [2014b] C. Crespo-Cadenas, P. Aguilera-Bonet, J. A. Becerra, and S. Cruces, “On nonlinear amplifier modeling and identification using baseband Volterra-Parafac models,” Signal Processing, vol. 96, no. Part B, pp. 401 – 405, 2014. [2014c] J. A. Becerra , M. J. Madero-Ayora, Reina-Tosina, J., and Crespo-Cadenas, C, “Predistorsionador adaptativo por bloques basado en un modelo Volterra-Parafac,” in XXIX Simposium Nacional de la Unión Científica Internacional de Radio (XXIX Spanish National Symposium of the International Union of Radio Science), September 2014.
8.3 Contributions 73 The author has been involved in the following publications that are not part of the thesis. [2018a] J. A. Becerra , M. J. Madero-Ayora, J. Reina-Tosina, C. Crespo-Cadenas, J. García-Frías, and G. Arce, “A doubly orthogonal matching pursuit algorithm for sparse predistortion of power amplifiers,” IEEE Microwave and Wireless Components Letters, DOI: 10.1109/LMWC.2018.2845947. 1–3, 2018. [2018b] J. A. Becerra , D. Herrera, M. J. Madero-Ayora, and C. Crespo-Cadenas, “Sparse model selection of digital predistorters using Subspace Pursuit,” in 13th European Microwave Integrated Circuits Conference (EuMIC) , September 2018. [2018c] A. Pérez-Hernández, J. A. Becerra , M. J. Madero-Ayora, and C. Crespo-Cadenas, “A Comprehensive Performance Benchmark of Component Selection Techniques for Volterra Behavioral Models,” in XXXIII Simposium Nacional de la Unión Científica Internacional de Radio (XXXIII Spanish National Symposium of the International Union of Radio Science), September 2018. [2018d] J. A. Becerra , M. J. Madero-Ayora, A. Pérez-Hernández, and C. Crespo-Cadenas, “ Algoritmo Subspace Pursuit para la Selección de Parámetros de Modelos de Volterra,” in XXXIII Simposium Nacional de la Unión Científica Internacional de Radio (XXXIII Spanish National Symposium of the International Union of Radio Science), September 2018. [2018e] J. A. Becerra , A. Pérez-Hernández, M. J. Madero-Ayora, and C. Crespo-Cadenas, “Efficient linearization of a RF Transmitter under 5G waveforms through Iterated Ridge Regression,” in 91st ARFTG Microwave Measurement Conference, June 2018. [2017c] A. Pérez-Hernández, D. Herrera, M. J. Madero-Ayora, and J. A. Becerra , “Modelado Sin Memoria de Amplificadores de Potencia con Curvas de Bézier Cúbicas,” in XXXII Simposium Nacional de la Unión Científica Internacional de Radio (XXXII Spanish National Symposium of the International Union of Radio Science), September 2017. [2017c] M. J. Madero-Ayora, J. Reina-Tosina, C. Crespo-Cadenas, and J. A. Becerra , “Revisión de Modelos de Volterra para Sistemas no Lineales con Señales Complejas,” in XXXII Simposium Nacional de la Unión Científica Internacional de Radio (XXXII Spanish National Symposium of the International Union of Radio Science), September 2017. [2016] J. A. Becerra , M. J. Madero-Ayora, J. Reina-Tosina, and C. Crespo-Cadenas, “Structured compressed-sensing for Volterra series models,” in 2016 IEEE 17th Annual Wireless and Microwave Technology Conference (WAMICON), April 2016, pp. 1–4. [2012] J. A. Becerra , J. Reina-Tosina, E. A. Ojeda-García, and C. Crespo-Cadenas, “Modelo simple con características electrotérmicas para dispositivos HEMT de GaN,” in XXVII Simposium Nacional de la Unión Científica Internacional de Radio (XXVII Spanish National Symposium of the International Union of Radio Science), September 2012.
List of Figures 2.1 Example of an AM/AM characteristic. 8 2.2 Example of an AM/Gain characteristic. 9 2.3 Example of an AM/PM characteristic. 9 2.4 Input-output characteristic of a saturated PA and a DPD. 10 2.5 Schematic of a predistorter and a power amplifier. 11 2.6 Block diagram of Volterra series represented as a measurement process where his the vector of Volterra coefficients to be estimated. 11 2.7 Contours of the error and constraint functions for the LASSO (left) and Ridge regression (right) in a space with coordinates β1and β2. The solid areas are the constraint regions, |β1|+|β2| ≤ sand β2 1+β2 2≤s, while the ellipses are the contours of the RSS. Figure adapted from [13]. 14 2.8 Block diagram of Volterra series represented as a measurement process where his the vector of Volterra coefficients to be estimated. 15 2.9 Experimental test bench. From left to right. Front: ZHL42W Minicircuits’ driver amplifier, CGH40010 Cree’s evaluation board, coupler and power load. Back: SMU200A vector signal generator, PXA N9030A vector signal analyzer. 17 2.10 Evaluation board of the CGH40010 GaN HEMT transistor. 18 2.11 Picture of the Chalmers University of Technology weblab. 18 2.12 Diagram of the Chalmers University of Technology weblab. 19 3.1 Structure of the Volterra-Parafac baseband model. 23 3.2 Volterra-Parafac structure and implementation of block processing with efficient frequency-domain techniques (B-VP DPD). The block index kis omitted. 25 3.3 A DPD with Volterra-Parafac structure for OFDM systems (FD-VP DPD). The block index kis omitted. 26 6.1 Acquisition model with additive noise (a) and signal segmentation (b). 43 7.1 AM/AM (a) and AM/PM (b) characteristics of the unpredistorted and linearized GaN PA driven by a 15-MHz LTE signal. 51 75
76 List of Figures 7.2 Normalized power spectral density of the GaN PA at an average output level of +30 dBm without DPD (blue), employing the GMP-based DPD (green) and the proposed B-VP DPD (red). Dashed lines correspond to the error spectrum. 52 7.3 Transmitted constellation of the GaN PA associated with the 15-MHz LTE signal at an average output level of +30 dBm without DPD (blue), employing the GMP-based DPD (green) and the proposed B-VP DPD (red). The +marks correspond to the reference constellation for the transmitted symbols. 53 7.4 Measured power spectral density of the test PA with different DPD linearizers. The unpredistorted PA was driven at an input level to yield the same output power than the linearized PA (+26.2 dBm). 54 7.5 16-QAM constellation of a received 15-MHz LTE-downlink signal without DPD (red dots), with a GMP DPD (blue dots), and with the proposed FD-VP DPD (black dots). The reference constellation point is plotted with a green ‘+’ mark. One symbol has been zoomed in. 55 7.6 Normalized power spectral density of the measured (dotted line) and modeled (solid lines) output of the PA driven at 3.6 GHz with an impaired I/Q modulator. 57 7.7 Spectrum of the error signal between the measured and modeled output of the PA driven at 3.6 GHz with an impaired I/Q modulator (solid lines) and an I/Q modulator with precompensated impairments (dot-dashed line). 58 7.8 Spectral densities for one realization of the experiment for the cases without DPD and with DPD in the case of α=1. 60 7.9 15-MHz bandwidth signal constellation for the original signal without DPD and with DPD in the case of α=1. 61 7.10 Adjacent channel power ratios of the output spectrum with (solid lines) and without (dashed lines) DPD. First adjacent channel (+marks) and second adjacent channel (×marks). Basic transmitter TUT-1. 62 7.11 NMSE of the output signal without DPD (dashed line) and with DPD (solid line). Basic transmitter TUT-1. 63 7.12 Compression curves of the TUT-2 transmitter for average output levels of +26 dBm (blue), +23.4dBm (red), 20.6dBm (orange), and +17.8dBm (magenta). Gain compression (a) and AM/PM (b) characteristics. 64 7.13 Normalized output spectrum and spectrum of the error for an average output level of Po= +26 dBm (+37 dBm, peak), with and without DPD. Realistic transmitter TUT-2. 64 7.14 Measured normalized coefficients at Po= +26 dBm (large marks) and computed by using (6.15) (straight lines). (a) Odd-order and (b) even-order coefficients. The PA input power is represented in the abscissas axis. Realistic transmitter TUT-2. 65 7.15 Adjacent channel power ratios of the output spectrum with and without DPD. First adjacent channel (solid lines) and second adjacent channel (dotted lines). Realistic transmitter TUT-2. 66 7.16 NMSE of the output signal without DPD (dashed line) and with DPD (solid line). Realistic transmitter TUT-2. 66
List of Figures 77 7.17 Normalized spectral density of the linearized output for power levels ranging from +17.8dBm to +26.6dBm. As a reference, the trace of the output spectrum without DPD is also shown. Realistic transmitter TUT-2. 67 7.18 Comparison of a) lower and upper ACPR1and b) NMSE for the SISO ACC model [47] (circles) and the present proposal with the CVS model (triangles). The filled marks correspond to performance at the identification level (Po= +26 dBm). Realistic transmitter TUT-2. 68
List of Tables 6.1 Summary of the Proposed Identification Procedure 46 7.1 DPDs performance comparison in terms of out-of-band emissions in the first and second adjacent channels with the GaN PA at Pout = +30 dBm 52 7.2 DPDs performance comparison for LTE-downlink signals with the GaN PA at Pout = +30 dBm 53 7.3 DPD performance comparison for a 15-MHz LTE-downlink signal (Pout = +26.2 dBm) 55 7.4 Modeling performance in terms of the NMSE and ACEPR for the test PA with impairments in the I/Q modulator 56 7.5 Performance results of DPD in a sweep of span αvalues. 59 79
86 Bibliography [62] H. Zhou, G. Wan, and L. Chen, “A nonlinear memory power amplifier behavior modeling and identification based on memory polynomial model in soft-defined shortwave transmitter,” in Proc. 6th Int. Conf. Wireless Communications Networking and Mobile Computing (WiCOM), Sep. 2010, pp. 1–4. [63] G. T. Zhou and G. B. Giannakis, “Nonlinear channel identification and performance analysis with PSK inputs,” in Proc. First IEEE Signal Processing Workshop Signal Processing Advances in Wireless Communications, Apr. 1997, pp. 337–340. [64] M. F. Duarte and Y. C. Eldar, “Structured compressed sensing: From theory to applications,” IEEE Trans. Signal Process., vol. 59, no. 9, pp. 4053–4085, 2011. [65] D. L. Donoho, “Compressed sensing,” IEEE Trans. Inf. Theory, vol. 52, no. 4, pp. 1289–1306, 2006. [66] A. Montanari, Graphical models concepts in compressed sensing. Cambridge University Press, 2012, p. 394–438. [67] M. Lustig, D. Donoho, and J. M. Pauly, “Sparse MRI: The application of compressed sensing for rapid MR imaging,” Magn. Reson. Med., vol. 58, no. 6, pp. 1182–1195, 2007. [68] G. Shi, D. Liu, D. Gao, Z. Liu, J. Lin, and L. Wang, “Advances in theory and application of compressed sensing,” Acta Electronica Sinica, vol. 37, no. 5, pp. 1070–1081, 2009. [69] M. Davies, “Compressed sensing: Theory and applications.” [70] M. A. Figueiredo, R. D. Nowak, and S. J. Wright, “Gradient projection for sparse reconstruction: Application to compressed sensing and other inverse problems,” IEEE Journal of Selected Topics in Signal Processing, vol. 1, no. 4, pp. 586–597, 2007. [71] S. Foucart and H. Rauhut, A Mathematical Introduction to Compressive Sensing. Birkhäuser Basel, 2013, vol. 1, no. 3. [72] S. Qaisar, R. M. Bilal, W. Iqbal, M. Naureen, and S. Lee, “Compressive sensing: From theory to applications, a survey,” J. Commun. Networks, vol. 15, no. 5, pp. 443–456, 2013. [73] J. Haupt, W. U. Bajwa, G. Raz, and R. Nowak, “Toeplitz compressed sensing matrices with applications to sparse channel estimation,” IEEE Trans. Inf. Theory, vol. 56, no. 11, pp. 5862–5875, 2010. [74] W. Yin, S. Osher, D. Goldfarb, and J. Darbon, “Bregman iterative algorithms for `1 -minimization with applications to compressed sensing,” SIAM J. Imag. Sci., vol. 1, no. 1, pp. 143–168, 2008. [75] L. Gan, “Block compressed sensing of natural images,” in 15th International Conference on Digital Signal Processing. IEEE, 2007, pp. 403–406.
Bibliography 87 [76] M. Lustig, D. L. Donoho, J. M. Santos, and J. M. Pauly, “Compressed sensing MRI,” IEEE Signal Process. Mag., vol. 25, no. 2, pp. 72–82, 2008. [77] G. Kutyniok, “Theory and applications of compressed sensing,” GAMMMitteilungen, vol. 36, no. 1, pp. 79–101, 2013. [78] E. J. Candès and M. B. Wakin, “An introduction to compressive sampling,” IEEE Signal Process. Mag., vol. 25, no. 2, pp. 21–30, 2008. [79] M. C. Vuran, Ö. B. Akan, and I. F. Akyildiz, “Spatio-temporal correlation: theory and applications for wireless sensor networks,” Comput. Networks, vol. 45, no. 3, pp. 245–259, 2004. [80] E. T. Hale, W. Yin, and Y. Zhang, “A fixed-point continuation method for `1 - regularized minimization with applications to compressed sensing,” Computational and Applied Mathematics (CAAM) Technical Report, Rice University, vol. 43, p. 44, 2007. [81] J. K. Cavers, “The effect of quadrature modulator and demodulator errors on adaptive digital predistorters for amplifier linearization,” IEEE Trans. Veh. Technol., vol. 46, no. 2, pp. 456–466, May 1997. [82] T. Blumensath and M. E. Davies, “Stagewise weak gradient pursuits,” IEEE Trans. Signal Process., vol. 57, no. 11, pp. 4333–4346, Nov. 2009. [83] A. Maleki, L. Anitori, Z. Yang, and R. G. Baraniuk, “Asymptotic analysis of complex LASSO via complex approximate message passing (CAMP),” CoRR, vol. abs/1108.0477, 2011. [Online]. Available: http://arxiv.org/abs/1108.0477 [84] J. Tan, Y. Ma, H. Rueda, D. Baron, and G. R. Arce, “Approximate message passing in coded aperture snapshot spectral imaging,” in 2015 IEEE Global Conference on Signal and Information Processing (GlobalSIP), Dec. 2015, pp. 609–613. [85] D. Donoho, A. Maleki, and A. Montanari, “Message passing algorithms for compressed sensing: I. motivation and construction,” in 2010 IEEE Information Theory Workshop on Information Theory (ITW 2010, Cairo), Jan. 2010, pp. 1–5. [86] D. Donoho, A. Maleki, and A. Montanari, “Message passing algorithms for compressed sensing: II. analysis and validation,” in 2010 IEEE Information Theory Workshop on Information Theory (ITW 2010, Cairo), Jan. 2010, pp. 1–5. [87] M. Younes, O. Hammi, A. Kwan, and F. Ghannouchi, “An accurate complexityreduced plume model for behavioral modeling and digital predistortion of RF power amplifiers,” IEEE Trans. Ind. Electron., vol. 58, no. 4, pp. 1397–1405, Apr. 2011. [88] M. Schetzen, The Volterra and Wiener Theories of Nonlinear Systems. New York: Wiley, 1980.
88 Bibliography [89] S. Rangan, P. Schniter, and A. Fletcher, “On the convergence of approximate message passing with arbitrary matrices,” in IEEE International Symposium on Information Theory (ISIT), June 2014, pp. 236–240. [90] A. Montanari, “Graphical models concepts in compressed sensing,” CoRR, vol. abs/1011.4328, 2010. [Online]. Available: http://arxiv.org/abs/1011.4328 [91] V. Kekatos and G. Giannakis, “Sparse Volterra and polynomial regression models: Recoverability and estimation,” IEEE Trans. Signal Process., vol. 59, no. 12, pp. 5907–5920, Dec. 2011. [92] V. Kekatos and G. Giannakis, “Compressed sensing for Volterra and polynomial regression models,” in Signal Processing Conference, 2011 19th European, Aug. 2011, pp. 373–377. [93] M. Khalifa, A. Abdelhafiz, A. Kwan, F. Ghannouchi, and O. Hammi, “Behavioral modeling of envelope tracking power amplifier using Volterra series model and compressed sampling,” in EEE InternationalWireless Symposium (IWS),, Mar. 2015, pp. 1–4. [94] A. Bolstad and B. Miller, “Sparse Volterra systems: Theory and practice,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May. 2013, pp. 5740–5744. [95] A. Zhu, J. Pedro, and T. Brazil, “Dynamic deviation reduction-based Volterra behavioral modeling of rf power amplifiers,” IEEE Trans. Microw. Theory Techn., vol. 54, no. 12, pp. 4323–4332, Dec. 2006. [96] Z. Yu and E. Zhu, “A comparative study of learning architecture for digital predistortion,” in Proc. Asia-Pacific Microwave Conf. (APMC), vol. 1, Dec. 2015, pp. 1–3. [97] T. Magesacher and P. Singerl, “Benchmarking of learning architectures for digital predistortion,” in Proc. Systems and Computers 2016 50th Asilomar Conf. Signals, Nov. 2016, pp. 648–651. [98] H. Paaso and A. Mammela, “Comparison of direct learning and indirect learning predistortion architectures,” in Proc. IEEE Int. Symp. Wireless Communication Systems, Oct. 2008, pp. 309–313. [99] M. Schetzen, “Theory of pth-order inverses of nonlinear systems,” IEEE Trans. Circuits Syst., vol. 23, no. 5, pp. 285–291, May 1976. [100] M. Ergen, Mobile Broadband - Including WiMAX and LTE, 1st ed. Springer Publishing Company, Incorporated, 2009. [101] D. H. Holma and D. A. Toskala, LTE for UMTS - OFDMA and SC-FDMA Based Radio Access. Wiley Publishing, 2009.
Bibliography 89 [102] M. Allegue-Martínez, “Modeling and compensation of non-linear effects in wireless communications systems,” Ph.D. dissertation, Universidad de Sevilla, 2012. [103] A. S. Tehrani, “Behavioral modeling of wireless transmitters for distortion mitigation,” Ph.D. dissertation, Chalmers University of Technology, 2012. [104] C. Crespo-Cadenas, M. Madero-Ayora, J. Reina-Tosina, and J. Becerra-Gonzalez, “A widely nonlinear approach to compensate impairments in I/Q modulators,” in 2015 European Microwave Conf. (EuMC), Sept 2015, pp. 506–509.