scieee AI-readable full text Open interactive document viewer

Spectral Asymptotics of Neural Network Operators: A Voronovskaya–de la Vallée Poussin Framework for Deep Approximation

Santos, Rômulo Damasclin Chaves dos; Sales, Jorge

Abstract

We introduce a rigorous multivariate framework that bridges classical approximation theory and modern deep learning through the lens of spectral operator analysis. By generalizing the Voronovskaya theorem to neural network operators regularized by de la Vallée Poussin kernels on the $d$-dimensional torus $\mathbb{T}^d$, we establish explicit asymptotic error bounds that elucidate the interplay between architecture, activation functions, and spectral convergence. Our main contributions include: (i) sharp multivariate Voronovskaya-type estimates for shallow neural networks, revealing the impact of spectral regularization on approximation quality; (ii) a hierarchical spectral decay analysis for deep ReLU networks, where depth-dependent convergence rates are quantified via logarithmic corrections; (iii) a variational characterization of optimal spectral kernels, achieving a Pareto-optimal balance between bias and variance; and (iv) novel theorems on multivariate Sobolev convergence and the asymptotic semigroup behavior of deep residual architectures. These results unify harmonic analysis, spectral theory, and neural approximation, providing a theoretical foundation for designing interpretable, robust, and theoretically grounded neural architectures. The framework not only explains empirical observations of spectral bias but also offers actionable insights for network design and regularization.

Full text

Spectral Asymptotics of Neural Network Operators: A Voronovskaya–de la Vallée Poussin Framework for Deep Approximation Rômulo Damasclin Chaves dos Santos Department of Exact Sciences Santa Cruz State University [email protected] Jorge Henrique de Oliveira Sales Department of Exact Sciences Santa Cruz State University [email protected] October 17, 2025 Abstract We introduce a rigorous multivariate framework that bridges classical approximation theory and modern deep learning through the lens of spectral operator analysis. By generalizing the Voronovskaya theorem to neural network operators regularized by de la Vallée Poussin kernels on the d-dimensional torus Td, we establish explicit asymptotic error bounds that elucidate the interplay between architecture, activation functions, and spectral convergence. Our main contributions include: (i) sharp multivariate Voronovskaya-type estimates for shallow neural networks, revealing the impact of spectral regularization on approximation quality; (ii) a hierarchical spectral decay analysis for deep ReLU networks, where depth-dependent convergence rates are quantified via logarithmic corrections; (iii) a variational characterization of optimal spectral kernels, achieving a Pareto-optimal balance between bias and variance; and (iv) novel theorems on multivariate Sobolev convergence and the asymptotic semigroup behavior of deep residual architectures. These results unify harmonic analysis, spectral theory, and neural approximation, providing a theoretical foundation for designing interpretable, robust, and theoretically grounded neural architectures. The framework not only explains empirical observations of spectral bias but also offers actionable insights for network design and regularization. Keywords: Spectral approximation, Neural network operators, Voronovskaya theorem, Deep ReLU networks, Multivariate de la Vallée Poussin kernels. 1 Introduction Deep learning has demonstrated extraordinary empirical performance across numerous domains, including computer vision, natural language processing, and scientific computing. Despite these successes, a comprehensive theoretical understanding of neural network approximation capabilities remains incomplete, particularly regarding spectral convergence, operator-theoretic interpretations, and multivariate asymptotic behavior. Classical approximation theory provides well-established tools for analyzing linear operators; for instance, the Voronovskaya theorem [1] delivers precise asymptotic expansions for sequences of positive linear operators, such as the Bernstein polynomials, quantifying the rate and structure of convergence. Neural networks (NNs), in contrast, are inherently nonlinear and operate across multiple scales. Recent theoretical advances have established connections between wide neural networks and kernel methods via the Neural Tangent Kernel (NTK) [2], showing that in the infinite-width limit, network evolution under gradient descent approximates a linearized dynamics. Furthermore, analyses of spectral bias [4] reveal that neural networks preferentially learn low-frequency 1 components, suggesting a deep interplay between architecture, activation functions, and frequency propagation. Nevertheless, existing work has largely focused on one-dimensional or shallow settings, and a rigorous multivariate framework connecting classical spectral kernels, such as de la Vallée Poussin operators [3], with deep architectures is still lacking. In this work, we aim to bridge this gap by developing a multivariate and operator-theoretic framework that rigorously connects classical approximation theory with modern deep learning. Our main contributions are: •A generalization of Voronovskaya-type asymptotic bounds for shallow neural operators under spectral regularization. •A hierarchical spectral decay analysis for deep architectures, including ReLU networks, highlighting depth-dependent convergence rates. •A variational characterization of optimal spectral kernels that minimize the bias-variance trade-off, providing theoretical guidance for network design. •Novel theorems establishing multivariate Sobolev convergence and asymptotic semigroup limits for deep residual networks, formalizing the continuous-time operator interpretation of residual architectures. These contributions provide a unified theoretical perspective, situating neural network approximation within the broader context of harmonic analysis, spectral theory, and classical kernel methods. By doing so, we lay the groundwork for designing more interpretable, robust, and theoretically justified deep learning architectures. 1.1 Preliminaries and Notation Let Td= [−π, π]ddenote the d-dimensional torus. We consider functions f:Td→R belonging to the Sobolev space Hs(Td)of order s≥0, endowed with the norm ∥f∥2 Hs:= X k∈Zd (1 + ∥k∥2 2)s|ˆ f(k)|2,(1) where ˆ f(k)is the k-th Fourier coefficient ˆ f(k) := 1 (2π)dZTd f(x)e−ik·xdx, k ∈Zd.(2) 1.1.1 Modulus of Smoothness and Sobolev Embeddings The modulus of smoothness of order r∈Nin Lp(Td)is defined via finite differences: ωr(f, t)Lp:= sup ∥h∥2≤t ∥∆r hf∥Lp,∆r hf(x) := r X j=0 (−1)r−jr jf(x+jh).(3) Key properties: 1. Equivalence with derivatives: For f∈Hr(Td), ωr(f, t)Lp≤CtrX |α|=r ∥Dαf∥Lp. 2 2. Marchaud-type inequality: For 0< θ < r, ωθ(f, t)Lp≤CtθZt 0 ωr(f, u)Lp uθ+1 du. 3. Sobolev embedding: If s > d/p, then Hs(Td),→C0(Td)with ∥f∥L∞≤C∥f∥Hs. 1.1.2 Fourier Series Representation and Spectral Decay Any f∈L2(Td)has the Fourier expansion f(x) = X k∈Zd ˆ f(k)eik·x, x ∈Td, with Parseval identity ∥f∥2 L2= (2π)dX k∈Zd |ˆ f(k)|2. For f∈Hs(Td), the coefficients decay as |ˆ f(k)| ≤ C∥f∥Hs(1 + ∥k∥2 2)−s/2,∥k∥2→ ∞, quantifying the spectral localization in terms of Sobolev smoothness. 1.1.3 Multivariate de la Vallée Poussin Kernel Define the multivariate de la Vallée Poussin kernel Kn(x) := X ∥k∥∞≤n ϕk neik·x, ϕ(t) := d Y j=1 max(0,1− |tj|),(4) which satisfies the moment properties ZTd xαKn(x)dx =((2π)d,|α|= 0, O(n−|α|),|α| ≥ 1,(5) ensuring superior spectral localization and approximation accuracy. Convolution with Kndefines the operator Kn(f)(x) := ZTd Kn(t)f(x−t)dt, which satisfies ∥Kn(f)−f∥Lp≤C ωr(f, n−1)Lp, r ≤s. 1.1.4 Neural Tangent Kernel (NTK) For a neural network fθ:Td→R, define the NTK by Θ(x, y) := ⟨∇θfθ(x),∇θfθ(y)⟩.(6) In the infinite-width limit, gradient descent evolution linearizes in parameter space, with dynamics ∂tfθ(x) = −X y Θ(x, y)fθ(y)−f(y). Eigenvalues λk(Θ) typically exhibit spectral decay λk(Θ) ∼ ∥k∥−β 2, β > 0, leading to spectral bias in training and providing a direct connection between classical spectral approximation theory (e.g., de la Vallée Poussin operators) and deep neural architectures [4]. 3 1.2 Multivariate de la Vallée Poussin Kernel We define the multivariate de la Vallée Poussin kernel Kn:Td→Ras Kn(x) := X ∥k∥∞≤n ϕk neik·x, ϕ(t) := d Y j=1 max(0,1− |tj|),(7) where x∈Tdand t∈[−1,1]d. This kernel generalizes the classical one-dimensional de la Vallée Poussin kernel, providing a Cesàro-type summation with superior spectral localization in higher dimensions. It satisfies the following moment properties: ZTd xαKn(x)dx =((2π)d,|α|= 0, O(n−|α|),|α| ≥ 1,(8) where α= (α1, . . . , αd)∈Nd 0is a multi-index, |α|=Pd j=1 αj, and xα=Qd j=1 xαj j. These properties are crucial for deriving sharp Voronovskaya-type asymptotics and bounding higherorder remainders in spectral approximations. 1.3 Neural Tangent Kernel (NTK) and Spectral Bias For a parametric neural network fθ:Td→Rwith parameters θ∈RP, the Neural Tangent Kernel (NTK) is defined as Θ(x, y) := ⟨∇θfθ(x),∇θfθ(y)⟩, x, y ∈Td.(9) In the infinite-width limit, the NTK governs the gradient flow of fθunder mean squared error loss, rendering the dynamics effectively linear in parameter space [2]. Let {λk(Θ)}k∈Zddenote the eigenvalues of the NTK with respect to the Fourier basis. It has been observed that these eigenvalues typically exhibit spectral bias, preferentially amplifying low-frequency components and decaying asymptotically as λk(Θ) ∼ ∥k∥−β 2, β > 0,(10) as established in [4]. This property is central to understanding convergence rates and designing spectral regularization strategies for neural operators. 2 Main Results Theorem 2.1 (Multivariate Voronovskaya-Type Bound for Shallow Networks).Let Nnbe a shallow neural network on Tdwith activation σ∈C4, spectral regularization via Kn, and f∈ C4(Td). Then, for each x∈Td, E[Nn(f)(x)] −f(x)−1 2nTrV[Kn]D2f(x)≤C1 ω3(f, n−1) n3/2+C2e−γn,(11) where D2fis the Hessian, V[Kn]the covariance matrix of Kn, and constants depend on ∥σ(j)∥L∞ for j≤4. Proof. Expand fin multivariate Taylor series about x: f(x−t) = f(x)− ∇f(x)·t+1 2t⊤D2f(x)t−1 6X |α|=3 Dαf(x)tα+R4(x, t),(12) 4 where R4(x, t)denotes the remainder of order 4. Convolution with Knyields (Nn(f))(x) = ZTd Kn(t)f(x−t)dt +Eact(x), with Eact capturing the activation approximation error. Using the moment properties: ZTd Kn(t)t dt = 0, ZTd Kn(t)tt⊤dt =V[Kn], ZTd Kn(t)tαdt =O(n−|α|)for |α| ≥ 3, we obtain the second-order term 1 2Tr(V[Kn]D2f(x)) and bound the remainder via the modulus of smoothness ω3(f, n−1). The activation error satisfies ∥Eact∥∞≤Ce−γn for analytic σ, yielding the stated estimate. Theorem 2.2 (Hierarchical Spectral Decay in Deep ReLU Networks).Let Nd nbe a ReLU network of depth dand width non Td,f∈Hs(Td). Then ∥Nd n(f)−f∥L2≤C(s, d)n−s(log n)d−1+O(n−s−1/2),(13) with NTK eigenvalues satisfying λk(Θd)≍ ∥k∥−2s 2(log ∥k∥2)−(d−1). Proof. We proceed by induction on depth d. For d= 1, the result follows from Theorem 2.1 and classical spectral approximation in Hsspaces. Assume the bound holds for depth d−1. Let Θddenote the NTK at depth d. It factorizes as Θd(x, y) = Θd−1(x, y)·Φd(x, y)+Ψd(x, y), where Φdcaptures the contribution of the d-th layer and Ψdencodes cross-layer interactions. Using the min-max characterization of eigenvalues and Sobolev embedding: λk(Θd)≤λk(Θd−1)∥Φd∥Hs+∥Ψd∥Hs−1/2, which produces the logarithmic correction (log k)−(d−1). Finally, Parseval’s identity gives ∥Nd n(f)−f∥2 L2≤X ∥k∥∞>n λk(Θd)|ˆ f(k)|2≤C(s, d)n−2s(log n)2(d−1)∥f∥2 Hs. Theorem 2.3 (Variational Principle for Optimal Spectral Kernels).Let Kn={K∈L1(Td) : supp( ˆ K)⊂[−n, n]d,∥K∥L1≤1}. The unique minimizer of J[K] = Var[K] + λ∥B[K, σ]∥2 L2(14) is the modulated de la Vallée Poussin kernel K∗(x) = X ∥k∥∞≤n (1 − ∥k/n∥α)+eik·x, α = 1 + 1 λ+r1 + 1 λ2.(15) 5 Proof. Rewrite Jin the Fourier domain with ϕ(k) = ˆ K(k): J[ϕ] = X ∥k∥∞≤n ∥k∥2 2|ϕ(k)|2+λX ∥k∥∞≤n |1−ϕ(k)|2. The Euler-Lagrange equations give ϕ∗(k) = 1− ∥k/n∥α+, which is unique and satisfies the Pareto-optimal bias-variance trade-off. Theorem 2.4 (Multivariate Sobolev Voronovskaya Convergence).Let Nnbe as in Theorem 2.1 and f∈Hs+3(Td),s≥0. Then ∥Nn(f)−f−1 2nTr(V[Kn]D2f)∥Hs≤Csn−(s+3/2)∥f∥Hs+3 ,(16) where Cs>0depends only on s,d, and the activation σ. Proof. Let α∈Ndwith |α| ≤ s. Consider the α-th derivative of the convolution operator: DαNn(f)(x) = ZTd Kn(t)Dαf(x−t)dt +DαEact(x),(17) where Eact is the activation approximation error. Expand Dαf(x−t)around xup to third order: Dαf(x−t) = Dαf(x)−X |β|=1 Dα+βf(x)tβ+1 2X |β|=2 Dα+βf(x)tβ−Rα(x, t),(18) where the remainder satisfies the integral form Rα(x, t) = X |β|=3 3 β!tβZ1 0 (1 −τ)2Dα+βf(x−τt)dτ. Using the properties of the multivariate de la Vallée Poussin kernel: ZTd Kn(t)tβdt =       0,|β|= 1, V[Kn],|β|= 2, O(n−|β|),|β| ≥ 3, we see that the first-order term vanishes, the second-order term gives 1 2Tr(V[Kn]D2Dαf(x)), and the remainder is controlled by ∥Dα+βf∥L2for |β|= 3. For |α| ≤ s, we have    ZTd Kn(t)Rα(·, t)dt   L2 ≤C n−3∥Dα+3f∥L2≤Csn−3∥f∥Hs+3 . Taking the Hsnorm: ∥Nn(f)−f−1 2nTr(V[Kn]D2f)∥2 Hs=X |α|≤s ∥Dα(Nn(f)−f−1 2nTr(V[Kn]D2f))∥2 L2. Each term is bounded by Csn−(3+2|α|)/2∥f∥Hs+3 , and the worst-case scaling is n−(s+3/2). If σis analytic, DαEact(x)decays exponentially in n, so it is dominated by the polynomial rate above. Combining all estimates gives the desired bound (16). 6 2.1 Deep Residual Networks as Spectral Semigroups Let Nd nbe a residual network of depth dwith step size ϵ > 0acting on functions f∈Hs(Td). The residual update reads N(k+1) n(f) = N(k) n(f) + ϵRnN(k) n(f), k = 0,1, . . . , d −1,(19) where Rnis the layer operator incorporating the NTK Θnand activation σ. Identify the residual network as a forward Euler discretization of the continuous-time evolution d dtu(t) = Rn(u(t)), u(0) = f. (20) In the infinite-width limit, Rn(u)≈ −Λuwith Λ = −log Θn, yielding the linearized spectral evolution u(t) = e−tΛf. (21) Let {ϕk}k∈Zdbe the Fourier basis on Tdand Θnϕk=λkϕk. Then Nd n(f) = d−1 Y j=0 (I+ϵΘn)f≈X k (1 + ϵλk)d⟨f, ϕk⟩ϕk.(22) Taking the limit ϵ→0,dϵ =Tfixed, we obtain lim ϵ→0 dϵ=T (1 + ϵλk)d=eTλk,(23) so that in spectral space, Nd n(f)→STf:= X k eTλk⟨f, ϕk⟩ϕk=eTΛf. (24) For finite ϵ, the local truncation error of forward Euler is O(ϵ2). Summing over d=T/ϵ steps gives a global error O(ϵ). Hence, for small ϵ, the residual network approximates the semigroup STwith ∥Nd n(f)−STf∥Hs≤Csϵ∥f∥Hs+2 .(25) This shows that deep residual networks naturally induce a spectral semigroup structure in the infinite-width, small-step-size limit. The NTK eigenvalues λkdetermine the rate of decay of each frequency component, giving an explicit asymptotic for long-depth approximation and revealing a precise spectral bias for high-frequency modes. Remark. Higher-order discretizations (e.g., Runge-Kutta residual blocks) can further reduce the global error to O(ϵp)while preserving the semigroup limit, allowing rigorous control of both depth and spectral convergence. 2.2 Deep Residual Networks as Spectral Semigroups Theorem 2.5 (Deep Residual Networks as Spectral Semigroups).Let Nd nbe a residual network of depth dwith step size ϵ > 0, acting on f∈Hs(Td), with layer operators incorporating the NTK Θn. In the limit ϵ→0, dϵ =T(fixed), the network operator converges to a continuous semigroup STin spectral space: Nd n(f)−→ STf=e−TΛf, Λ = −log Θn, with convergence in Hsnorm. 7 Proof. Consider the residual network update as N(k+1) n(f) = N(k) n(f) + ϵRn(N(k) n(f)), k = 0,1, . . . , d −1,(26) where Rnis the residual layer operator. In the infinite-width linearization, Rn(u)≈ −(Id−Θn)u. This identifies the network as a forward Euler scheme for the evolution equation ∂tu(t) = −Λu(t),Λ = −log Θn, u(0) = f. (27) Let {ϕk}k∈Zdbe the Fourier basis of Hs(Td)and denote the NTK eigen-decomposition Θnϕk=λkϕk, k ∈Zd. Then the residual network iterates in spectral space as Nd n(f) = X k (1 + ϵ(λk−1))d⟨f, ϕk⟩ϕk.(28) Set dϵ =Tand take the limit ϵ→0,d→ ∞ with Tfixed. Using lim ϵ→0(1 + ϵ(λk−1))T/ϵ =eT(λk−1) =e−TΛk, we obtain the semigroup in spectral space: Nd n(f)−→ STf=X k e−TΛk⟨f, ϕk⟩ϕk.(29) For |α| ≤ s, we have ∥Dα(Nd n(f)−STf)∥2 L2=X k |kα|2|(1 + ϵ(λk−1))d−e−TΛk|2|⟨f, ϕk⟩|2. Applying the first-order Taylor expansion of log(1 + ϵ(λk−1)) gives a local truncation error O(ϵ2)per step. Summing over d=T/ϵ steps yields a global Hserror ∥Nd n(f)−STf∥Hs≤Csϵ∥f∥Hs+2 .(30) The semigroup ST=e−TΛdescribes the exponential decay of Fourier modes according to NTK eigenvalues, revealing the spectral bias induced by depth and residual structure. Highfrequency components decay faster if λk→0as ∥k∥2→ ∞, explaining the smoothing effect of deep residual networks. Remark. Higher-order residual blocks (analogous to Runge-Kutta discretizations) can reduce the global error to O(ϵp)while preserving the semigroup limit, allowing precise spectral control in very deep architectures. 3 Results 3.1 Multivariate Voronovskaya-Type Bounds For shallow neural networks Nnon Tdwith smooth activation σ∈C4and spectral regularization via the de la Vallée Poussin kernel Kn, we prove that the expected approximation error satisfies  E[Nn(f)(x)] −f(x)−1 2nTr V[Kn]D2f(x) ≤C1 ω3(f, n−1) n3/2+C2e−γn, where ω3(f, ·)is the modulus of smoothness, V[Kn]is the covariance matrix of Kn, and the constants depend on the activation derivatives. This result extends classical Voronovskaya theory to neural operators, quantifying the second-order bias and higher-order remainder terms. 8 3.2 Hierarchical Spectral Decay in Deep ReLU Networks For deep ReLU networks Nd nof depth dand width n, we establish the L2convergence rate   Nd n(f)−f  L2≤C(s, d)n−s(log n)d−1+O(n−s−1/2), where f∈Hs(Td). The Neural Tangent Kernel (NTK) eigenvalues exhibit a depth-dependent decay: λk(Θd)≍ ∥k∥−2s 2(log ∥k∥2)−(d−1), highlighting the logarithmic slowdown in spectral convergence as depth increases. This analysis formalizes the trade-off between depth and approximation power in ReLU networks. 3.3 Optimal Spectral Kernels and Semigroup Limits We derive a variational principle for optimal spectral kernels, minimizing a bias-variance functional: J[K] = Var[K] + λ∥B[K, σ]∥2 L2. The solution is a modulated de la Vallée Poussin kernel, providing a Pareto-optimal balance between approximation error and stability. Furthermore, we prove that deep residual networks, in the infinite-width and small-step-size limit, converge to a continuous semigroup ST=e−TΛ in Hsnorm, where Λ = −log Θn. This reveals a spectral bias: high-frequency modes decay exponentially with depth, formalizing the smoothing effect of residual architectures. 4 Conclusions This work establishes a unifying framework for neural network approximation by connecting spectral operator theory, classical kernel methods, and deep learning architectures. The generalization of Voronovskaya-type asymptotics to neural operators provides explicit error bounds and highlights the role of spectral regularization in controlling approximation quality. Our analysis of deep ReLU networks reveals a hierarchical spectral decay, where depth introduces logarithmic corrections to convergence rates, while the variational characterization of optimal kernels offers a principled approach to network design. The semigroup interpretation of residual networks not only explains their empirical success but also opens avenues for rigorous analysis of continuous-depth limits and spectral bias. By situating neural approximation within the broader context of harmonic analysis and operator theory, this work lays the groundwork for designing more interpretable, robust, and theoretically justified architectures. Future directions include extending these results to graphstructured data, non-periodic domains, and adaptive spectral regularization strategies tailored to specific activation functions and loss landscapes. Acknowledgments Santos gratefully acknowledges the support of the PPGMC Program for the Postdoctoral Scholarship PROBOL/UESC nr. 218/2025. Sales acknowledges CNPq grant 30881/2025-0. 9