Figures
Abstract
Affine frequency division multiplexing (AFDM) demonstrates exceptional resilience to Doppler effects in doubly dispersive channels, making it a promising waveform for 6G high-mobility communications, but its high peak-to-average power ratio (PAPR) problem severely limits system energy efficiency. When migrated to AFDM, existing mask-based deep learning PAPR reduction schemes face the challenges of contextual loss caused by local sampling and the lack of physical interpretability in black-box models. Therefore, this paper proposes a physics-aware mask-based PAPR reduction scheme tailored for AFDM. First, a full-frame input strategy is introduced to exploit the global time-domain correlation of AFDM signals for precisely recovering the impaired symbols. Second, to address the nonlinear distortion induced by masking, a PolyNet-Volterra network is proposed by integrating Volterra series theory. Departing from the traditional design of blindly stacking layers, this model explicitly constructs first-order linear and third-order power feature layers, which substantially enhances the nonlinear reconstruction accuracy while avoiding overfitting. Simulation results demonstrate that, under a threshold of 0.9, the proposed scheme achieves a significant PAPR reduction of approximately 5.9 dB. Furthermore, requiring only about 165,000 parameters, the PolyNet-Volterra model comprehensively outperforms baseline models such as DNN and ResNet in terms of both bit error rate (BER) and mean squared error (MSE). Compared with the conventional PolyNet, the proposed model reduces the parameter and computational overhead, suggesting its potential as a low-complexity receiver-side reconstruction module for resource-constrained future wireless systems. Further hardware-oriented and over-the-air validation is still needed before drawing deployment-level conclusions.
Citation: Liu X, Zhang Y, Guo J, Zhu R, Wang F, Wang L (2026) A mask-based peak-to-average power ratio reduction scheme for affine frequency division multiplexing systems using Volterra Series-Driven PolyNet. PLoS One 21(7): e0354675. https://doi.org/10.1371/journal.pone.0354675
Editor: Xuebo Zhang, Northwest Normal University, CHINA
Received: April 21, 2026; Accepted: July 9, 2026; Published: July 28, 2026
Copyright: © 2026 Liu et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data and code generated or analyzed during this study are included in the manuscript and supporting files. The simulated data can be reproduced using the code provided in the supplementary materials.
Funding: This work was supported by the National Social Science Foundation of China (Grant No. 23AGL039), the Shaanxi Provincial Key Research and Development Program (Grant No. 2021GY341), the Guangdong Provincial Key Laboratory of Prevention and Control for Severe Clinical Animal Diseases (Grant No. 20190102), and the Shaanxi Provincial Department of Education General Special Scientific Research Project (Grant No. 24JK0693), all awarded to Yushuai Zhang. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
With the rapid evolution of mobile communication technologies, the sixth-generation (6G) wireless network has become a global focal point for breakthrough technologies [1–3]. It aims to provide highly reliable and ubiquitous connectivity services for high-mobility scenarios, such as high-speed railways, low Earth orbit (LEO) satellite internet, and unmanned aerial vehicle (UAV) communications [4]. In such demanding scenarios, the high relative velocity between terminals and base stations induces significant Doppler shifts, causing the wireless channel to exhibit severe doubly selective fading characteristics. This leads to inter-symbol interference and frequency-selective fading, which makes channel estimation and equalization more difficult. To tackle these technical challenges, affine frequency division multiplexing (AFDM) has emerged as a novel multi-carrier modulation technique [5–7]. By incorporating the discrete affine Fourier transform (DAFT), AFDM can achieve full diversity gain in doubly dispersive channels. Furthermore, it not only delivers superior resilience to Doppler shifts compared to orthogonal frequency division multiplexing (OFDM), but also entails lower overhead for channel estimation and multi-user multiplexing than orthogonal time frequency space (OTFS) modulation [8–10].
AFDM is promising for high-mobility scenarios, but its time-domain waveform is still formed by the superposition of multiple chirp-modulated subcarriers. Therefore, the high peak-to-average power ratio (PAPR) problem still needs to be addressed [11,12]. When these signals with extremely high instantaneous peaks pass through a high power amplifier (HPA), they are driven into the saturation region, exceeding the linear amplification limit and suffering from clipping. This phenomenon inevitably causes in-band distortion, which elevates the bit error rate, and out-of-band radiation, which interferes with adjacent channels. Therefore, energy-constrained satellite, vehicular, and low-altitude communication scenarios provide a strong motivation for studying low-complexity AFDM PAPR reduction methods while maintaining communication reliability [13].
Currently, research on PAPR reduction is quite mature, encompassing traditional methods such as clipping, partial transmit sequence (PTS), and selective mapping (SLM) [14–16]. However, while clipping is simple and straightforward, it introduces irreversible nonlinear distortion, severely degrading the BER at the receiver [17,18]. Conversely, probabilistic schemes like PTS and SLM require the transmission of additional side information, and their computational complexity grows exponentially with the number of candidate sequences. In recent years, with the widespread application of deep learning in communication physical layer technologies [19], existing schemes [20–23] have successfully achieved PAPR reduction by leveraging deep neural networks (DNNs). Most notably, He et al. [24] proposed a mask-based PAPR reduction architecture for OFDM systems. This scheme masks a small fraction of high-peak symbols at the transmitter and uses a DNN at the receiver for signal reconstruction. In this way, PAPR reduction and signal recovery are handled in two separate steps, which improves the overall performance.
However, several critical issues in existing research remain unresolved. First, the input to their mask reconstruction neural network comprises only the symbols at the masked positions within the received signal, along with the indices of these positions. Specifically, rather than utilizing the entire received signal, the model’s input is strictly confined to the masked symbols. Although this input strategy reduces computational complexity, it discards much of the time-domain context around the masked samples. As a result, the network has limited information for estimating the local waveform trajectory, which may lead to similar reconstruction limits across different model architectures [25]. Furthermore, these existing black-box models ignore the physical mechanisms of communication signal distortion. They typically rely on blindly stacking massive numbers of neurons and network layers to approximate the nonlinear mapping, resulting in an enormous number of parameters and sluggish training convergence. This not only exacerbates the computational burden on the receiver but also fails to meet the stringent requirements of future 6G networks for low-latency and low-power communications [26–28].
To address the aforementioned dual challenges, this paper proposes a physics-aware mask-based PAPR reduction scheme tailored for AFDM systems. Initially, the scheme employs a dynamic masking mechanism to precisely locate high-peak regions within the signal, strategically replacing these high instantaneous power peaks to significantly reduce the PAPR at the source. Subsequently, a lightweight polynomial neural network designed based on Volterra series theory (termed PolyNet-Volterra) is utilized to perform nonlinear reconstruction of the impaired symbols within these peak regions. This innovative approach substantially lowers the PAPR while maintaining a low bit error rate (BER).
The main contributions of this paper are summarized as follows:
- We introduce a mask-based PAPR reduction scheme into the AFDM system, validating its effectiveness under chirp modulation. We also extend the masked-symbol input strategy to a full-frame input strategy. This allows the network to use frame-level signal context when reconstructing the masked samples.
- To handle the nonlinear amplitude distortion introduced by masking, we propose PolyNet-Volterra, a neural network model inspired by the Volterra series. The model uses power-related feature layers to improve interpretability. A mask-specific loss function is also used to improve reconstruction accuracy at high-energy masked samples.
- Simulation results demonstrate that, under identical PAPR reduction levels, the proposed PolyNet-Volterra model achieves faster convergence and superior BER performance compared to conventional DNN and ResNet models, while drastically reducing the parameter count. These results indicate that PolyNet-Volterra is a promising low-complexity candidate for AFDM PAPR reduction, while its practical deployment still requires further validation under hardware and over-the-air conditions.
The remainder of this paper is organized as follows. The Related Work section reviews existing approaches. The AFDM System Model section introduces the system model. The Simulation Results Analysis section presents and analyzes the experimental results. Finally, the Conclusion section summarizes the study.
Related work
In multi-carrier communication systems, the high peak-to-average power ratio (PAPR) has been a persistent technical challenge. For a continuous time-domain signal x(t),its PAPR is defined as the ratio of the peak power to the average power, which is mathematically expressed as:
where denotes the maximum instantaneous power (i.e., the peak power) attainable by the signal within a continuous symbol period T, and
represents the expected (average) power of the signal over the same period. An excessively high PAPR inevitably drives the signal into the nonlinear saturation region of the high-power amplifier (HPA), causing in-band signal distortion and out-of-band spectral regrowth. To address the severe PAPR issue in wireless multi-carrier systems, academia has proposed various solutions, including clipping, companding transforms, and precoding-based PAPR reduction schemes [29,30,31]. Existing research can generally be categorized into conventional signal processing methods and deep learning (DL)-based optimization approaches.
Conventional PAPR reduction techniques
Conventional PAPR reduction techniques are quite mature and mainly include signal distortion techniques and probabilistic techniques.
Signal distortion techniques include methods such as clipping and companding. This approach directly compresses signal peaks through nonlinear operations, featuring simple implementation and high efficiency. However, such nonlinear processing destroys the orthogonality of the signal, introducing in-band distortion and out-of-band radiation, which leads to significant degradation of the bit error rate (BER) at the receiver. Although filtering can be applied to mitigate out-of-band radiation, it typically causes the peak regrowth phenomenon. In addition, techniques such as companding are also commonly used to smooth power peaks [32].
Probabilistic techniques include methods such as partial transmit sequence (PTS) and selective mapping (SLM). These methods generate multiple candidate signals by applying phase rotations to the signal in the frequency domain and select the one with the minimum PAPR for transmission. Although this category of techniques does not cause signal distortion, it requires executing multiple IFFT operations, and the computational complexity escalates exponentially with the number of candidate sequences. In AFDM systems, given the introduction of the discrete affine Fourier transform (DAFT) and additional chirp modulation parameters, the overhead of such multi-branch parallel computation would further increase, making it even more prohibitive for the system to tolerate.
Existing deep learning-based PAPR reduction schemes
In recent years, deep learning has been widely applied to PAPR reduction tasks due to its powerful nonlinear fitting capabilities [33–35]. Some studies utilize autoencoders to completely replace traditional modulation and demodulation modules, jointly optimizing the PAPR and BER performance. For example, by designing a weighted loss function, the network is guided to generate low-PAPR waveforms. However, such tightly coupled architectures typically require collaborative training between the transmitter and the receiver, resulting in an enormous number of model parameters. Moreover, models trained for specific channels are difficult to generalize to the complex doubly selective fading scenarios of AFDM. Another category of methods employs deep learning to optimize the key parameters of traditional algorithms, such as learning the optimal phase factors in SLM or the extension vectors in ACE. Although these methods yield performance improvements, they remain confined by the high-complexity frameworks of traditional algorithms.
Mask-based schemes and their limitations
In existing mask recovery schemes, the network typically takes only the impaired symbols at the masked positions as input. This design effectively reduces the input dimension of the model and performs excellently in OFDM systems. The underlying reason is that the subcarriers in OFDM exhibit relative independence, making local recovery sufficient to meet efficiency requirements. However, in AFDM systems, after the signal undergoes chirp modulation and the discrete affine Fourier transform (DAFT), a tighter coupling relationship emerges among the time-domain samples. This coupling originates from the non-orthogonality of the modulation scheme and the global nature of the transformation, causing the signal components to interweave with each other. If the recovery relies solely on the local information at the masked positions, it becomes difficult to fully exploit the rich contextual correlations embedded in the uncorrupted neighboring symbols, thereby degrading the accuracy and robustness of the recovery. This issue becomes particularly pronounced under high-noise or complex channel conditions.
Meanwhile, most existing reconstruction schemes employ generic fully connected neural networks (DNNs) for end-to-end recovery. Such generic architectures possess strong generalization and nonlinear fitting capabilities, enabling them to implicitly learn the nonlinear inverse transformations by deepening the network layers and adapt to various signal environments and impairment patterns. However, overly generic network structures are usually accompanied by massive parameter counts and high computational overhead. Such substantial parameter volumes and computational demands can lead to sluggish training processes and significantly increased memory footprints, posing feasibility challenges for deployment in real-time applications or on resource-constrained devices. In addition, generic DNNs may lack the effective encoding of prior knowledge in the signal processing domain and cannot be optimized for specific modulation or transformation characteristics. This limitation restricts the room for further improvements in model efficiency and performance. A comparison of existing schemes is presented in Table 1.
AFDM system model
This section details the proposed AFDM communication system architecture, which encompasses the complete signal transmission, reception, and processing link. At the transmitter, the bit stream sequentially undergoes quadrature amplitude modulation (QAM) mapping, pilot insertion, and inverse discrete affine Fourier transform (IDAFT) to generate the time-domain AFDM waveform. Subsequently, a time-domain masking module strategically intercepts and replaces the high instantaneous power symbols prior to transmission. This reduces the system PAPR at the source, effectively preventing the signal from being driven into the nonlinear saturation region of the power amplifier. The signal is then transmitted to the receiver over a doubly dispersive fading channel affected by multipath delays, Doppler shifts, and additive white Gaussian noise (AWGN). At the receiver, the signal captured by the antenna first undergoes preliminary linear channel equalization via a minimum mean square error (MMSE) equalizer before entering the PolyNet-Volterra neural network module cascaded thereafter. The PolyNet-Volterra module uses first-order and third-order power-related features inspired by the Volterra series to reconstruct the nonlinear distortion caused by masking before demodulation. Finally, the signal is transformed back to the symbol domain via the discrete affine Fourier transform (DAFT) and undergoes QAM demapping to recover the original data. The overall architecture of the proposed scheme is illustrated in Fig 1.
Mask-based PAPR reduction mechanism
The core of the AFDM system lies in utilizing the discrete affine Fourier transform (DAFT) to process the signal, enabling it to maintain orthogonality in doubly dispersive channels. The time-domain transmitted signal of AFDM, denoted as , is generated by the inverse discrete affine Fourier transform (IDAFT):
where denotes the standard inverse discrete Fourier transform (IDFT) matrix, and
represents the transmitted symbol vector in the discrete affine Fourier transform (DAFT) domain. The matrices
and
are two diagonal chirp modulation matrices with their diagonal elements given by
and
, respectively. The parameters c1 and c2 are selected based on the optimal parameter design criteria of AFDM as
and
, where M is an integer, to ensure that the signal achieves full diversity gain in the delay-Doppler domain. By expanding the aforementioned formula, the time-domain sample x[n] can be expressed as:
This equation reveals that, structurally, the signal is essentially an inverse discrete Fourier transform (IDFT) encapsulated by a dual-layer chirp term, namely the transmitter pre-processing and the time-domain post-processing
. From a physical perspective, this time-domain waveform is intrinsically a coherent superposition of N quadratically phase-modulated subcarriers at time index n. When the instantaneous phases of these subcarriers align in the complex plane at a specific moment due to constructive interference, their amplitudes accumulate linearly in phase. This causes the instantaneous power at that specific point to significantly exceed the average power of the entire frame. Although chirp modulation alters the phase distribution, it does not eliminate the fundamental mathematical core of multi-carrier linear summation. Consequently, AFDM inherits a high PAPR drawback similar to that of OFDM, where the peak power can theoretically reach N times the average power under conditions of extreme phase alignment.
A high PAPR not only degrades the operational efficiency of the high-power amplifier (HPA) but may also drive the signal into the nonlinear saturation region of the HPA, thereby inducing signal distortion. For AFDM signals employing chirp modulation, the nonlinear distortion from the HPA destroys the orthogonality of the chirps, resulting in severe inter-carrier interference (ICI) during DAFT demodulation at the receiver.
To reduce the PAPR without incurring high computational complexity, this paper conducts a statistical analysis of the amplitude characteristics of AFDM signals. The probability density function of the AFDM signal amplitude exhibits a prominent long-tail distribution characteristic, as illustrated in Fig 2. A long-tail distribution is a probability distribution in statistics characterized by a higher frequency in its tail; that is, within a dataset, although the majority of data points are concentrated within a certain range, a significant number of rare events or outliers exist in the tail. This implies that only a minute fraction of symbols carry extremely high instantaneous amplitudes. This contrasts with conventional companding transform methods, which alter the entire signal, as shown in Fig 3. Based on this observation, this paper proposes employing a nonlinear processing mechanism based on time-domain masking. The core concept of this mechanism is to strategically replace the high-amplitude peaks in the time-domain signal directly at the transmitter, thereby limiting the peak power of the signal at the source. However, this nonlinear operation inevitably introduces in-band distortion, which is precisely the target that the deep learning network must recover in subsequent sections.
Neural network-based signal recovery at the receiver
Before detailing the neural network architecture, this paper first formally defines the signal reconstruction task at the receiver from a mathematical perspective. Assuming the original time-domain AFDM signal is , the masking operation at the transmitter can be essentially abstracted as a nonlinear function
. Then, the transmitted signal after mask processing, denoted as
, can be expressed as:
After this signal propagates through the doubly selective fading channelHand is corrupted by additive white Gaussian noisez, the observed signalyat the receiver can be expressed as:
At the receiver, the signal first passes through a conventional channel equalizer (such as an MMSE equalizer) to eliminate the channel fading effects caused by multipath and Doppler shifts. After ideal equalization, the signal can be approximately expressed as:
where is the residual noise after the equalization operation. Thus, it is evident that for a conventional receiver, even if the channel equalization is extremely perfect, the signal still retains the nonlinear distortion
introduced by the masking operation. If symbol decision is performed directly at this stage, this nonlinear distortion will inevitably lead to a severe error floor.
Consequently, this paper designs a neural network parameterized by
to approximate the inverse function of the masking operation,
, as closely as possible under complex noise interference. The core task of the network is to remove the effect of
from the impaired signal and reconstruct the original signal:
In this receiver structure, neural reconstruction is performed after MMSE equalization. The MMSE equalizer first suppresses the dominant linear doubly selective channel effect, so that the neural network can focus on recovering the nonlinear distortion introduced by time-domain masking. This processing order makes the learning task more targeted and reduces the burden on the reconstruction network.
To approximate this inverse mapping with a compact model, we use the physical structure of the masking distortion to construct prior features for the network. Considering that the nonlinear amplitude distortion of the signal is strongly correlated with its instantaneous power, we design the following power-aware feature expansion vector and employ it as the physical feature input layer for the neural network . For any complex sample
within the received signal vector, where r is the real part, i is the imaginary part, and the instantaneous power is
, its feature expansion vector is defined as:
The retained feature terms are chosen from the dominant structure of amplitude dependent baseband distortion. The first order components r and i preserve the main linear signal content, while the third order power components r|s|2 and i|s|2 provide the lowest order nonlinear correction related to the instantaneous power. This treatment is consistent with Volterra inverse modeling and memory polynomial predistortion, where low order complex baseband terms such as x and x|x|2 are commonly used to model amplitude dependent RF distortion and inverse compensation. The second order terms are not included because they introduce even order cross products that are less consistent with the odd symmetric baseband distortion considered here, while also increasing the feature dimension.
The polynomial feature expansion is applied only inside the receiver-side reconstruction network after channel equalization; it does not modify the transmitter-side AFDM modulation matrix, chirp basis, or DAFT/IDAFT definition. Therefore, the expansion does not directly preserve or destroy AFDM chirp orthogonality. Its role is to reconstruct the masked time-domain samples before DAFT-domain demodulation.
Based on this physical prior, the complete forward propagation process of the neural network can be reformulated as a composite mapping of feature expansion and a multilayer perceptron (MLP):
Through this explicit construction, the network is no longer required to guess the patterns of nonlinear envelope distortion via deep neuron stacking. Instead, it performs weighted fusion directly based on the key physical features provided by , which comprise the first-order fundamental components (r,i) containing the basic information of the original signal, and the third-order power distortion components
. This architecture, combining feature engineering with a lightweight network, tremendously boosts the efficiency of approximating
. In terms of mathematical formulation, this feature construction method draws inspiration from the nonlinear system modeling theory of the Volterra series [36]. For a memoryless nonlinear system, its output y(n) can be approximately expressed as an odd-order polynomial expansion of its input x(n):
Based on the above theory, this paper designs the PolyNet-Volterra network, as shown in Fig 4. Unlike conventional DNNs, the proposed PolyNet-Volterra neural network adopts an architecture that combines physics-prior-based feature expansion with a deep multilayer perceptron (MLP). The network is designed to compensate for the nonlinear distortion induced by the masking operation. The entire forward propagation process is divided into three main stages:
- 1. The network receives an estimated complex signal frame of dimension N as input. Based on Volterra series theory, the network first decouples the signal and extracts four-dimensional physical features, including the first-order real part (r) and imaginary part (i), as well as the third-order nonlinear terms
and
constructed from the instantaneous power
. The four N-dimensional feature vectors are concatenated into a real-valued tensor of dimension 4N, which serves as the input to the subsequent network.
The full-frame input strategy is adopted because AFDM samples are globally coupled through chirp modulation and DAFT/IDAFT operations. Using only the masked positions discards neighboring and frame-level waveform context, which is important for reconstructing the trajectory around high-power peaks. The receiver therefore feeds the entire equalized AFDM frame into the reconstruction network and updates only the masked positions at the output.
For AFDM frames with sizes different from N = 64, the same mask-and-reconstruct principle can still be applied after adjusting the input/output dimension and retraining the network for the new frame size. The current fully connected implementation should therefore not be regarded as a size-free model, but as an architecture that can be resized and retrained for other AFDM frame lengths.
- 2. The concatenated high-dimensional features are fed into the MLP backbone network consisting of two fully connected layers. The first layer maps the 4N-dimensional input to a hidden space with H neurons. Each linear mapping is followed by a cascade of layer normalization and the Tanh activation function. This structure not only effectively mitigates the internal covariate shift problem, but also endows the network with the expressive capacity required to fit the complex nonlinear amplitude-phase distortion.
Tanh is used in the PolyNet variants because the polynomial expansion can create large feature magnitudes near masked peaks, and a bounded activation helps keep normalized complex-baseband reconstruction stable. GELU and SiLU/Swish are well-established smooth activation functions [37,38] and are reasonable alternatives, but they were not selected as the default here because their unbounded positive side may amplify large polynomial features unless additional normalization or tuning is added. A dedicated activation-function ablation, including GELU and SiLU, is a meaningful future extension.
The hidden dimension was selected as 256 because the AFDM simulations use N = 64 subcarriers and the PolyNet-Volterra expansion forms four real-valued features per subcarrier, giving an expanded feature dimension of 4N = 256. This width lets the first hidden layer preserve the full expanded physical-feature representation instead of imposing an immediate bottleneck, while keeping the model compact enough for fair comparison with the main baselines. The same 256-dimensional setting was then used for the principal neural baselines and PolyNet variants so that performance differences mainly reflect the feature design and architecture rather than a narrower proposed model.
- 3. The final linear layer compresses the hidden layer features to 2N dimensions. Finally, these 2N real numbers are reconstructed into an N-dimensional complex tensor, where the first half serves as the real part and the second half as the imaginary part, thereby outputting the restored AFDM signal frame with both low PAPR and low distortion characteristics.
Simulation results analysis
Simulation setup
To evaluate the performance of the proposed PolyNet-Volterra scheme in the AFDM system, a deep learning simulation platform based on PyTorch is built in this paper. The simulation architecture strictly follows the standard settings in the fundamental theoretical literature of the AFDM system [5–7]. The detailed simulation parameter configurations are listed in Table 2.
To ensure the optimal performance of the PolyNet-Volterra network in achieving nonlinear signal reconstruction within the AFDM system, this paper designs a comprehensive training framework encompassing a specific-region loss function, a dynamic SNR adaptation mechanism, and a numerical stability control strategy. Given the spatial sparsity of the masking operation, the traditional global mean square error (Global MSE) is prone to being dominated by minute noise errors at numerous unmasked positions, which consequently dilutes the network’s focus on recovering high-amplitude distortions. To address this, we propose a mask-specific region loss function, which calculates the reconstruction error exclusively at the masked symbol positions. This loss function is defined as follows:
where K denotes the index set of the masked symbols, and |K| represents its cardinality. and
denote the reconstructed output of the network and the original transmitted symbol at position k, respectively, while
represents the set of network parameters. Through this focusing mechanism, the backpropagated gradients can bypass the limitations of the activation function’s gradients–much like the cross-entropy loss in network training–and guide the network to optimize its nonlinear reconstruction capability based directly on the discrepancy between the predicted and actual values. To endow the model with robust generalization capabilities under time-varying channel conditions, we abandon the traditional paradigm of fixed-SNR training in favor of a dynamic SNR adaptation mechanism. During each training batch, the signal-to-noise ratio (SNR) is randomly sampled from a uniform distribution
, where
and
. This data augmentation strategy forces the network to learn the intrinsic structural features of the signal rather than memorizing specific noise patterns, thereby significantly enhancing the model’s adaptability across diverse channel environments.
All numerical results are obtained from synthetic AFDM simulations generated from the system model described above. The global random seed is fixed to 42 for the main experiments. The implementation uses Python, PyTorch, NumPy, and Matplotlib. No early stopping is used; all enabled neural models are trained for the scheduled 150 epochs with Adam optimization and the StepLR learning-rate scheduler. The fixed training workload, random seeds, model configurations, and numerical data behind the added tables are provided to improve reproducibility.
Training convergence results
The convergence process under a fixed signal-to-noise ratio (SNR) environment of 20 dB is illustrated in Fig 5. Since the background noise power remains constant across all training epochs, the random interference induced by noise is eliminated, resulting in curves that exhibit a clear and smooth downward trend. Consequently, the fitting limits and convergence speeds of the various models are distinctly observable. It is evident that the traditional DNN and ResNet converge slowly and suffer from higher error lower bounds, whereas the proposed PolyNet-Volterra network, after rapidly plateauing, consistently maintains the lowest error level among all evaluated models. This demonstrates that incorporating the Volterra physical mechanism effectively enhances both the learning efficiency and reconstruction accuracy of the network.
The training state under a dynamic SNR environment, which randomly fluctuates between 5 dB and 25 dB, is depicted in Fig 6. Due to the severe fluctuations in the background noise intensity of the input data for each batch, the mean square error (MSE) calculated by the network also exhibits substantial variations, manifesting as dense sawtooth oscillations. Although the curves fluctuate strongly, PolyNet-Volterra remains lower than the other tested models over most training epochs. This result suggests that the proposed model can maintain stable reconstruction performance under varying SNR conditions.
Empirical verification confirms that this dynamic SNR training method does not significantly increase the computational burden. The total number of training samples remains constant; the samples are simply generated online to be uniformly distributed across various SNRs and channel conditions. Since the number of training samples is unchanged and the samples are generated online, dynamic SNR training does not add an obvious training-time burden. It improves generalization without increasing the dataset size.
Model selection and evaluation
To determine the optimal signal reconstruction architecture, this paper is not limited to a single model. Instead, we conduct extensive reproduction and comparative experiments (training from scratch) on various representative network architectures in the current deep learning domain, as shown in Table 3:
The benchmark models were selected to represent complementary lightweight reconstruction families under a common fully connected input-output interface. The DNN baseline reflects a generic black-box multilayer perceptron; ResNet tests residual learning; GatedMLP examines gated feature modulation; SE-ResNet adds channel-wise recalibration; and PolyNet provides a blind polynomial-expansion comparison. This set allows the effect of the physically selected Volterra-inspired first-plus-third-order features to be isolated under the same data, optimizer, batch size, epoch number, and evaluation metrics. More recent attention-based, sequence, graph, or generative reconstruction models are relevant future baselines, but they introduce different inductive biases and usually larger hardware-dependent costs that should be evaluated in a separate benchmarking study.
The computational complexity and performance scores of the various models are also quantified, with the experimental results presented in Table 4. This demonstrates that the proposed PolyNet-Volterra network possesses an overwhelming advantage in parameter efficiency. Due to its fully connected nature, the traditional DNN typically involves a massive number of parameters, leading to high computational load and memory consumption during training. In contrast to ResNet or DNN, the PolyNet-Volterra network requires significantly fewer parameters to achieve superior MSE and BER, obtaining the highest comprehensive performance score among all tested models. This indicates that the proposed scheme provides a favorable accuracy-complexity tradeoff among the tested simulation baselines. Its practical deployment performance should still be evaluated with hardware platforms and measured RF signals.
BatchNorm is used in the generic DNN, ResNet, GatedMLP, and SE-ResNet baselines because these models follow common batch-statistics-based MLP/residual designs, where mini-batch normalization helps stabilize deep real-valued hidden layers [39]. LayerNorm is used in the PolyNet variants because the polynomial feature distribution is frame dependent and can vary strongly with the masked peaks; normalizing within each sample is therefore more suitable for the frame-level polynomial representation [40]. All enabled models use the same optimizer, batch size, epoch number, SNR setting, and evaluation protocol.
BER performance analysis
The error rates and performance comparison of the various schemes tested in a random SNR environment ranging from 0 to 25 dB are illustrated in Fig 7 and Table 5. Among all comparative models, the BER curve of the PolyNet-Volterra network is the closest to the ideal curve. Particularly in the high-SNR region (e.g., 20 dB), compared with traditional ResNet and DNN models, its BER is further reduced while the number of parameters drops substantially. This improvement is attributed to its unique physics-aware feature layer, which can accurately reversely simulate and eliminate the nonlinear distortion introduced by the masking operation.
Volterra order ablation study and minimalist architecture justification
In the previous comparisons, the PolyNet-Volterra network demonstrated outstanding signal recovery capabilities with an extremely low number of parameters. This section delves into its underlying mechanisms to verify the scientific validity of relying exclusively on the 1st-order (r,i) and 3rd-order physics-prior features. This paper comprehensively analyzes the effectiveness of this minimalist architecture from multiple perspectives, including multi-dimensional feature ablation, feature-distortion correlation, power-dependent error distribution, and neural network interpretability.
Ablation of physics-prior features and parameter efficiency analysis
To explicitly determine the independent contribution of the features at each order, we construct various feature-subset models to compare their mean square errors and parameter efficiencies. The constructed feature-subset models are detailed in Table 6:
Relying exclusively on the third-order features results in the highest MSE, indicating that the first-order linear features serve as the cornerstone of signal reconstruction. Despite possessing a massive number of parameters, the DNNBaseline lacks physical priors, rendering its performance even inferior to that of the simplest LinearOnly model. Simply introducing an independent power term (PowerAug) fails to effectively enhance performance, which demonstrates that the network struggles to autonomously couple phase and power. Furthermore, although the StandardPolyNet, which incorporates all first- to third-order cross-terms, exhibits improvements in MSE, its performance in the high-SNR region remains constrained by feature redundancy, accompanied by a substantially larger parameter count. These results support the compact first and third order design used in PolyNet-Volterra. The broader first-to-third-order feature set increases the model size, but it does not provide a better high-SNR accuracy and complexity balance than the selected Volterra subset. The specific comparisons are illustrated in Fig 8.
In contrast, the proposed PolyNet-Volterra exclusively retains four groups of core physical features , as depicted in Fig 9. While drastically compressing the parameter count to 165K under a 15 dB environment, it simultaneously achieves the lowest MSE across all SNR levels. This succinctly and powerfully verifies the correctness of utilizing Volterra physical priors to eliminate redundant polynomials, thereby constructing a minimalist yet highly efficient compensation architecture.
Correlation analysis between features and nonlinear distortion
We calculate the Pearson correlation coefficients and partial correlation coefficients between the model input features and the actual distortion errors induced by the masking operation. As the SNR increases from 10 dB to 20 dB, the Pearson correlation between the features and the distortion errors gradually strengthens, as illustrated in Fig 10, Fig 11, and Fig 12. According to the evaluation criteria for Pearson correlation strength, an absolute value between 0.4 and 0.6 indicates a moderate correlation. Notably, the absolute value of the correlation coefficient between the third-order features and the errors reaches 0.571 at 20 dB, falling precisely into the moderate correlation range and exhibiting a considerably significant correlation. The partial correlation analysis on the right further confirms that even after controlling for the influence of the first-order features, the third-order features still maintain an independent and significant correlation. This intuitively explains the necessity of introducing the third-order physical features to accurately compensate for nonlinear distortion.
Physical validation of power-dependent reconstruction errors
The essence of clipping or amplifier distortion lies in the fact that the degree of signal distortion is highly dependent on instantaneous power; namely, high-power peaks suffer severe impairment, whereas low-power signals essentially maintain linearity. We extract the scatter relationship between the prediction error and the original signal power |s|2 at an SNR of 15 dB, as shown in Fig 13. In Fig 13(a), the original error escalates exponentially with the increase in signal power. Fig 13(b) reveals that although the LinearOnly model lowers the overall error baseline, its error reduction in the left region of the coordinate axis is negligible. In Fig 13(c), constrained by low learning efficiency due to an excess of non-independent features, the StandardPolyNet still exhibits a pronounced divergent scatter distribution across the entire power range. In contrast, Fig 13(d) demonstrates that the PolyNet-Volterra outperforms the other models, validating that its nonlinear terms can better manage the dynamic error envelope induced by clipping peak suppression and nonlinear saturation.
Data-driven perspective: gradient response and weight allocation
To verify the self-adaptive mechanism of this minimalist architecture in a data-driven environment, this paper unveils the black box of the neural network and analyzes the gradient importance of the features along with the first-layer weight distribution. The results are illustrated in Fig 14.
During the backpropagation process, the network not only allocates the primary gradient flow–accounting for approximately 74.3%–to the first-order linear features to ensure the fundamental signal reconstruction performance, but it also reserves a stable and significant optimization gradient–totaling 25.6%–for the third-order features. The weight distribution histogram in Fig 15 further indicates that upon convergence, the weights of the third-order features exhibit a healthy Gaussian-like distribution. Their L2 norm remains above 4.0, showing absolutely no signs of redundant failure such as collapsing toward zero under regularization. From the perspective of low-level network optimization, this thoroughly validates the effectiveness and indispensability of every single physical feature within the minimalist architecture.
Fig 15 shows that the first-layer weights are distributed over both the first-order features and the third-order power-coupled features. The linear real and imaginary components carry the largest weight energy, while the nonlinear rP and iP terms also keep clear nonzero contributions. This result suggests that the network does not simply rely on a black-box mapping, but uses the manually constructed low-order polynomial features during masked-sample reconstruction. Since these features are close in form to the basis terms used in truncated Volterra-type nonlinear inverse modeling [41,42], the proposed architecture can be viewed as a lightweight Volterra-inspired learned inverse approximation. It should not be regarded as a fixed polynomial kernel regression model, because the expanded features are still processed by learned nonlinear layers and optimized end to end for the AFDM reconstruction task.
Reconstruction error analysis in the masked region
To delve deeper into the specific performance of different network models in signal recovery, the mean square error (MSE) comparison of various schemes at the masked positions is illustrated in Fig. 16.
Unlike the traditional global MSE metric, this paper calculates the reconstruction error exclusively for the signal samples that were replaced by the mask at the transmitter. This metric effectively eliminates the interference from unmasked normal signal samples, precisely measuring the accuracy of the neural network in recovering the lost peak information.
A pronounced performance stratification phenomenon can be observed from the figure, where the gray dashed line represents the error of the received signal without any processing. Since the peaks are replaced by non-peak symbols, an inherent amplitude and phase deviation exists at these positions. Consequently, its MSE consistently remains at a high level and exhibits negligible variation with the SNR, indicating that the primary source of error is the nonlinear distortion at the transmitter rather than channel noise. Compared to the baseline, traditional deep learning models significantly reduce the MSE, demonstrating their capacity for signal repair to some extent. However, in the high-SNR region, their MSE curves tend to plateau, suggesting that these models struggle to further eliminate residual nonlinear errors. The MSE curve of the proposed PolyNet-Volterra model ultimately reaches the lowest error level. This implies that the signal peaks predicted by this model are the closest to the originally transmitted signals.
This MSE result is highly consistent with the aforementioned BER performance, explaining the source of the PolyNet-Volterra’s advantage from a signal processing perspective: by utilizing the Volterra series features to more accurately fit the nonlinear amplitude compression process, the model achieves higher-fidelity waveform reconstruction at the physical layer, which subsequently translates into a lower bit error rate at the demodulation end.
Additional validation and practical notes
To enhance the quantitative comparison, this paper conducted multiple independent statistical analyses. Table 7 reports 10 independent Monte Carlo simulation runs, each using 10,000 AFDM frames and an independent random seed, so the repeated-run variability can be seen directly. Table 8 reports aggregate uncertainty quantification using Wilson 95% confidence intervals for BER and run-level 95% confidence intervals for masked-region MSE. Table 9 reports formal significance testing against PolyNet-Volterra using an aggregate two-proportion z test for BER and a paired sign test across independent runs. Fig. 17 summarizes the repeated-run distribution, uncertainty intervals, and significance levels.
The masking threshold is a sensitivity parameter that controls how aggressively the transmitter compresses large time-domain peaks. Table 10 and Fig 18 give the sensitivity trend. When the threshold is reduced, peak suppression becomes stronger, but the reconstruction problem becomes harder. When the threshold is increased, the reconstruction burden is relaxed, but the peak-compression effect becomes weaker.
To examine high-order modulation more directly, PolyNet-Volterra was trained and evaluated under QPSK, 16-QAM, and 64-QAM, respectively. Table 11 and Fig 19 shows the demodulation accuracy and masked-region reconstruction error for different modulation schemes after modulation-specific training. As the constellation becomes denser, the decision regions become narrower, leading to a higher bit error rate (BER). Meanwhile, the masked-region mean square error (MSE) also increases. Since channel coding, interleaving, and BER-oriented detector optimization are not included in this study, the BER and MSE values remain relatively high under dense constellations. Even so, the results suggest that high-order modulation can still be supported through retraining.
The Doppler robustness test in Table 12 and Fig 20 increases the maximum integer Doppler index from 1 to 4 while keeping the same receiver structure. The BER and MSE increase moderately, indicating that the full-frame reconstruction remains usable under the tested higher-Doppler simulated channels. Regarding unseen channel models, this Doppler stress test provides an initial robustness check, but alternative channel families, model mismatch, and measured channels should be evaluated before making broad channel-generalization claims.
Table 13 reports parameter counts, linear-layer FLOPs per AFDM frame, and RTX 3090 GPU inference latency. The latency benchmark measures only the neural reconstruction module and excludes channel equalization and data generation. Single-frame latency and batch-256 latency are reported separately because they correspond to different inference modes.
Training wall-clock time was organized as a separate hardware-dependent metric. Table 14 reports the measured 150-epoch training time on the NVIDIA GeForce RTX 3090 GPU used for the main training runs.
A full-frame input ablation was added by retraining the enabled neural models with masked-only inputs, where the unmasked samples were removed from the network input. Table 15 and Fig 21 show that masked-only input consistently increases masked-region MSE and BER for all tested neural models. This confirms that, without full-frame contextual information, the models lose important AFDM waveform features and the advantage of the proposed architecture is weakened.
The receiver assumes conventional MMSE equalization with available channel information; synchronization errors and channel-estimation mismatch are important practical impairments that remain to be tested. For nonlinear RF impairments, Rapp/Saleh HPA models, memory-polynomial HPA behavior, spectral regrowth, and adjacent-channel leakage remain future validation tasks.
The present masking rule is based on time-domain amplitude and does not explicitly exploit DAFT-domain sparsity. Combining peak suppression with DAFT-domain sparse structure is a promising direction for adaptive masking design. The framework cannot be directly extended to OTFS without redesign. Only the general mask-and-reconstruct idea may be transferable; the input representation and physical feature basis would need to be rebuilt for the OTFS delay-Doppler transform structure.
A formal relation between masking sparsity and reconstruction-error bounds is not derived here. Such a bound would require assumptions on the AFDM signal distribution, channel model, masking pattern, and network approximation error. A learned adaptive masking policy may outperform the fixed threshold under changing SNR, channel, or modulation conditions. The fixed threshold is retained here because it provides a simple and reproducible setting for evaluating the reconstruction network.
Conclusion
Addressing the severe PAPR problem encountered by AFDM systems in 6G high-mobility scenarios, this paper proposes a physics-aware mask-based PAPR reduction scheme that achieves both high energy efficiency and high reliability. To overcome the two major limitations of existing deep learning-based mask reconstruction models during signal recovery–namely, the lack of contextual information and unclear physical mechanisms–this paper first introduces a full-frame input strategy to thoroughly exploit the time-domain correlation of AFDM signals; secondly, it innovatively constructs the PolyNet-Volterra network based on the Volterra series theory. Through theoretical analysis and rigorous ablation studies, this paper breaks the conventional paradigm in the deep learning field of blindly stacking layers to enhance nonlinear fitting capabilities, establishing a minimalist network architecture that extracts exclusively the first-order linear terms and third-order power terms of the signal.
Simulation results demonstrate that the proposed masking mechanism effectively compresses signal peaks. Regarding signal reconstruction at the receiver, compared to traditional DNN and ResNet baseline models, the PolyNet-Volterra network accurately captures the core physical features of amplitude distortion. With an extremely low parameter count of approximately 160K, it not only effectively circumvents the overfitting risk introduced by higher-order features but also achieves a comprehensive outperformance in both reconstruction mean square error (MSE) and bit error rate (BER). This research provides a promising nonlinear inverse reconstruction approach for relieving the PAPR bottleneck in multi-carrier communication systems. The compact structure of PolyNet-Volterra also suggests possible value for future resource-constrained wireless receivers, such as vehicular, low Earth orbit satellite, and UAV-related scenarios. However, these applications still require hardware implementation, measured-channel evaluation, and over-the-air validation before practical deployment conclusions can be drawn.
References
- 1. Chen S, Liang Y-C, Sun S, Kang S, Cheng W, Peng M. Vision, Requirements, and Technology Trend of 6G: How to Tackle the Challenges of System Coverage, Capacity, User Data-Rate and Movement Speed. IEEE Wireless Commun. 2020;27(2):218–28.
- 2. Saad W, Bennis M, Chen M. A vision of 6G wireless systems: Applications, trends, technologies, and open research problems. IEEE Network. 2020;34(3):134–42.
- 3. Alwis CD, Kalla A, Pham QV. Survey on 6G frontiers: trends, applications, requirements, technologies and future research. IEEE Open Journal of Communications Society. 2021;2:836–86.
- 4. Zhang L, Liang Y-C, Niyato D. 6G Visions: Mobile ultra-broadband, super internet-of-things, and artificial intelligence. China Commun. 2019;16(8):1–14.
- 5.
Bemani A, Cuozzo G, Ksairi N, Kountouris M. Affine Frequency Division Multiplexing for Next-Generation Wireless Networks. In: 2021 17th International Symposium on Wireless Communication Systems (ISWCS), 2021. 1–6. https://doi.org/10.1109/iswcs49558.2021.9562168
- 6.
Bemani A, Ksairi N, Kountouris M. AFDM: A Full Diversity Next Generation Waveform for High Mobility Communications. In: 2021 IEEE International Conference on Communications Workshops (ICC Workshops), 2021. 1–6. https://doi.org/10.1109/iccworkshops50388.2021.9473655
- 7. Bemani A, Ksairi N, Kountouris M. Affine Frequency Division Multiplexing for Next Generation Wireless Communications. IEEE Trans Wireless Commun. 2023;22(11):8214–29.
- 8.
Das SS. Orthogonal Time Frequency Space Modulation. 1st ed. Aalborg: River Publishers; 2022.
- 9.
Chockalingam A. OTFS modulation: theory and applications. 2022.
- 10. Wei Z, Yuan W, Li S. Orthogonal Time-Frequency Space Modulation: A Promising Next-Generation Waveform. IEEE Wireless Commun. 2021;28(4):136–44.
- 11. Rahmatallah Y, Mohan S. Peak-To-Average Power Ratio Reduction in OFDM Systems: A Survey And Taxonomy. IEEE Commun Surv Tutorials. 2013;15(4):1567–92.
- 12.
Wang X, Burkert S, ten Brink S. On peak to average power ratio of universal filtered OFDM signals. In: 2017 IEEE 28th Annual International Symposium on Personal, Indoor, and Mobile Radio Communications (PIMRC), 2017. 1–7. https://doi.org/10.1109/pimrc.2017.8292301
- 13.
Paredes MCP, Grijalva F, Carvajal-Rodriguez J, Sarzosa F. Performance analysis of the effects caused by HPA models on an OFDM signal with high PAPR. In: 2017 IEEE Second Ecuador Technical Chapters Meeting (ETCM), 2017. 1–5.
- 14.
Ku S-J. An improved low-complexity PTS scheme for PAPR reduction in OFDM systems. In: 2016 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), 2016. 1–5. https://doi.org/10.1109/icspcc.2016.7753598
- 15.
More AP, Somani SB. The reduction of PAPR in OFDM systems using clipping and SLM method. In: 2013 International Conference on Information Communication and Embedded Systems (ICICES), 2013. 593–7. https://doi.org/10.1109/icices.2013.6508385
- 16.
Alharbi FS, Chambers JA. Peak-to-average power ratio mitigation in quasi-orthogonal space time block coded MIMO-OFDM systems using selective mapping. In: 2008 Loughborough Antennas and Propagation Conference, 2008. 157–60. https://doi.org/10.1109/lapc.2008.4516890
- 17.
Rao KD, Murthy TSN. Analysis of Effects of Clipping and Filtering on the Performance of MB-OFDM UWB Signals. In: 2007 15th International Conference on Digital Signal Processing, 2007. 559–62. https://doi.org/10.1109/icdsp.2007.4288643
- 18. He C, Armstrong J. Clipping Noise Mitigation in Optical OFDM Systems. IEEE Commun Lett. 2017;21(3):548–51.
- 19. O’Shea T, Hoydis J. An Introduction to Deep Learning for the Physical Layer. IEEE Trans Cogn Commun Netw. 2017;3(4):563–75.
- 20. Wang X, Jin N, Wei J. A Model-Driven DL Algorithm for PAPR Reduction in OFDM System. IEEE Commun Lett. 2021;25(7):2270–4.
- 21. Kim M, Lee W, Cho D-H. A Novel PAPR Reduction Scheme for OFDM System Based on Deep Learning. IEEE Commun Lett. 2018;22(3):510–3.
- 22.
Wang H, Du L, Yang M, Chen G, Chen Y. A Novel PAPR Reduction of OFDM Based on Deep Learning. In: 2024 33rd Wireless and Optical Communications Conference (WOCC), 2024. 45–9.
- 23.
Tang X, Tan Z, Xiao L, Zhao M, Li Y. A New Approach to Dealing with the PAPR Problem for OFDM Systems via Deep Learning. In: 2022 14th International Conference on Wireless Communications and Signal Processing (WCSP), 2022. 252–6. https://doi.org/10.1109/wcsp55476.2022.10039468
- 24. He R, Zhang X, Liu Z, Cui Q, Tao X. Mask-Based PAPR Reduction Scheme With Deep Learning for OFDM Systems. IEEE Trans Wireless Commun. 2026;25:1813–26.
- 25. LeCun Y, Bengio Y, Hinton G. Deep learning. Nature. 2015;521(7553):436–44. pmid:26017442
- 26. Zhang Z, Xiao Y, Ma Z. 6G Wireless Networks: Vision, Requirements, Architecture, and Key Technologies. IEEE Veh Technol Mag. 2019;14(3):28–41.
- 27. Letaief KB, Chen W, Shi Y, Zhang J, Zhang Y-JA. The Roadmap to 6G: AI Empowered Wireless Networks. IEEE Commun Mag. 2019;57(8):84–90.
- 28.
Mahmood NH, Alves H, Lopez OA, Shehab M, Osorio DPM, Latva-Aho M. Six Key Features of Machine Type Communication in 6G. In: 2020 2nd 6G Wireless Summit (6G SUMMIT), 2020. 1–5. https://doi.org/10.1109/6gsummit49458.2020.9083794
- 29.
Kaur P, Singh M. Performance analysis of GA-PTS for PAPR reduction in OFDM system. In: 2016 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET), 2016. 2076–9. https://doi.org/10.1109/wispnet.2016.7566507
- 30.
Kang BM, Ryu H-G, Ryu SB. A PAPR Reduction Method using New ACE (Active Constellation Extension) with Higher Level Constellation. In: 2007 IEEE International Conference on Signal Processing and Communications, 2007. 724–7. https://doi.org/10.1109/icspc.2007.4728421
- 31.
Hsu C-Y, Liao H-C. PAPR reduction using the combination of precoding and Mu-Law companding techniques for OFDM systems. In: 2012 IEEE 11th International Conference on Signal Processing, 2012. 1–4. https://doi.org/10.1109/icosp.2012.6491517
- 32. Reddy VM, Bitra H. PAPR in AFDM: Upper Bound and Reduction With Normalized μ-Law Companding. IEEE Access. 2025;13:86553–61.
- 33. Ye H, Li GY, Juang B-H. Power of Deep Learning for Channel Estimation and Signal Detection in OFDM Systems. IEEE Wireless Commun Lett. 2018;7(1):114–7.
- 34. Zappone A, Di Renzo M, Debbah M. Wireless networks design in the era of deep learning: model-based, ai-based, or both?. IEEE Transactions on Communications. 2019;67(10):7331–76.
- 35. Farsad N, Goldsmith A. Neural Network Detection of Data Sequences in Communication Systems. IEEE Trans Signal Process. 2018;66(21):5663–78.
- 36.
Valet P, Schwingshackl D, Gaier U, Giotta D. Digital Pre-Distortion for CS-DACs based on Inverse Modelling of Volterra Series Expansion. In: 2022 Austrochip Workshop on Microelectronics (Austrochip). Villach, Austria: IEEE; 2022. p. 49–52.
- 37. Dubey SR, Singh SK, Chaudhuri BB. Activation functions in deep learning: A comprehensive survey and benchmark. Neurocomputing. 2022;503:92–108.
- 38. Elfwing S, Uchibe E, Doya K. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural Netw. 2018;107:3–11. pmid:29395652
- 39.
Ioffe S, Szegedy C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Proceedings of the 32nd International Conference on Machine Learning, 2015. 448–56.
- 40. Luo P, Zhang R, Ren J, Peng Z, Li J. Switchable Normalization for Learning-to-Normalize Deep Representation. IEEE Trans Pattern Anal Mach Intell. 2021;43(2):712–28. pmid:31380746
- 41. Ding L, Zhou GT, Morgan DR, Ma Z, Kenney JS, Kim J, et al. A Robust Digital Baseband Predistorter Constructed Using Memory Polynomials. IEEE Trans Commun. 2004;52(1):159–65.
- 42. Morgan DR, Ma Z, Kim J, Zierdt MG, Pastalan J. A Generalized Memory Polynomial Model for Digital Predistortion of RF Power Amplifiers. IEEE Trans Signal Process. 2006;54(10):3852–60.