Figures
Abstract
Deepfake anomaly detection has become increasingly critical in visual security applications. However, many existing methods rely heavily on large-scale supervised forgery annotations and often struggle to capture subtle manipulation traces distributed across both spatial and frequency domains. To address these limitations, we propose a Self-Supervised Dual-Domain Alignment (SDDA) framework that jointly learns spatial and spectral representations without requiring labeled forged samples. Specifically, SDDA employs two parameter-sharing encoders to extract complementary features from RGB images and their frequency-domain counterparts, while a cross-domain alignment module enforces consistency between spatial textures and spectral signatures. To provide explicit and reproducible self-supervised supervision, four complementary descriptors are further derived from FAN feature maps: Local Structural Deviation (LSD) for local gradient inconsistency, Global Pattern Discrepancy (GPD) for holistic channel-correlation differences, Local Residual Difference (LRD) for residual-level inconsistency, and Total Consistency Deviation (TCD) for multi-layer feature deviation. These descriptors are predicted from the fused spatial–frequency embedding, encouraging the model to learn manipulation-sensitive representations in the absence of fake samples. In addition, a dual-domain contrastive objective enhances the discrimination of subtle forgery-related anomalies while improving robustness against real-world degradations such as compression and illumination variations. Experimental results demonstrate that SDDA consistently outperforms state-of-the-art baselines, particularly under challenging unseen-manipulation scenarios. Overall, the proposed framework provides an interpretable and label-efficient solution for dual-domain deepfake anomaly detection.
Citation: He Y (2026) Dual-domain self-supervised feature alignment via spectral–spatial representation learning for deepfake anomaly detection. PLoS One 21(9): e0358049. https://doi.org/10.1371/journal.pone.0358049
Editor: Suo Gao, Dalian Polytechnic University, CHINA
Received: November 26, 2025; Accepted: August 26, 2026; Published: September 11, 2026
Copyright: © 2026 Yi He. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All data used in this study are from publicly available datasets and the corresponding references have been provided in the article. The URLs corresponding to the dataset is as follows: FaceForensics++: https://github.com/ondyari/FaceForensics. Celeb-DF: https://github.com/yuezunli/celeb-deepfakeforensics. Flickr-Faces-HQ: https://github.com/NVlabs/ffhq-dataset. WildDeepfake: https://github.com/OpenTAI/wild-deepfake.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
1 Introduction
Deepfake anomaly detection has become increasingly crucial in visual security applications, including identity authentication, surveillance, and multimedia forensics [1,2]. The main objective is to identify facial manipulations that deviate from genuine human appearance. However, reliable detection remains challenging due to complex illumination conditions, compression artifacts, and subtle low-level inconsistencies introduced by advanced generative models [3,4].
Early studies primarily relied on supervised deep networks such as Xception-based classifiers [5], capsule networks [6], and attention-guided CNNs [7]. Although these supervised models achieved strong results on known manipulation types, they tend to overfit dataset-specific visual cues, resulting in poor generalization to unseen forgeries [2]. To address this issue, unsupervised and one-class anomaly detection approaches—such as autoencoders [8], DeepSVDD [9], and GANomaly [10]—have been explored to model normal facial appearance and detect deviations as anomalies. Despite their advantages, spatial reconstruction-based frameworks often struggle to capture subtle, high-frequency inconsistencies introduced during synthesis and compression.
To enhance robustness, frequency-domain cues have recently been incorporated into deepfake and anomaly detection. Frank et al. [11] demonstrated that spectral magnitude discrepancies offer discriminative power across different forgery generators. Similarly, dual-domain architectures such as DCT-enhanced discriminators [12] and multi-branch frequency CNNs [13] attempted to fuse spatial and spectral features. Nevertheless, these approaches still depend heavily on supervised labels or handcrafted fusion strategies, limiting adaptability and interpretability [14]. As illustrated in Fig 1, existing deepfake anomaly detection frameworks often rely solely on pixel-level reconstruction or frequency-only artifacts. Such single-domain modeling restricts their ability to learn joint representations that capture both global structural integrity and local spectral inconsistencies. Moreover, reconstruction-based learning tends to produce overly smooth outputs, losing the fine-grained discriminative cues vital for anomaly localization.
Traditional methods rely solely on spatial reconstruction or supervised classification, failing to capture frequency-domain inconsistencies. Our method introduces a dual-domain self-supervised learning paradigm that jointly models spatial and spectral representations through shared-weight encoders, enhancing robustness and discriminability.
To overcome these limitations, we propose a dual-domain self-supervised deepfake anomaly detection framework. Our method employs shared-weight encoders to process both the spatial image and its Fourier spectrum, enabling consistent and aligned cross-domain representation learning. To further enhance discriminability, we introduce a multi-feature prediction task comprising four complementary objectives: Local Structural Deviation (LSD), Global Pattern Discrepancy (GPD), Local Residual Difference (LRD) and Total Consistency Deviation (TCD), as summarized in Table 1. These self-supervised objectives encourage the encoder to capture fine-grained variations at both the global and local levels and to model intrinsic correlations between spatial and frequency domains. Finally, we design a unified anomaly scoring mechanism that combines global reconstruction discrepancies with local frequency-prediction errors, improving robustness and interpretability for open-world detection scenarios.
We evaluate our framework on five public datasets—FaceForensics++ [5], Celeb-DF [2], Flickr-Faces-HQ [15], WildDeepfake [16], and ForgeryNet [17]. Extensive experiments demonstrate that our method consistently surpasses state-of-the-art baselines, yielding substantial improvements in AUC and average precision (AP). Unlike prior frequency-aware or self-supervised methods operating on single-domain representations, our approach explicitly enforces spatial–frequency alignment via multi-feature prediction, providing both enhanced generalization and greater interpretability in deepfake anomaly detection.
The major contributions of this work are as follows:
- We propose a dual-domain self-supervised framework that jointly models spatial and frequency representations via shared-weight encoders.
- We introduce a multi-feature prediction strategy to estimate four domain-consistent features (
,
,
,
), strengthening local–global alignment across domains.
- We design a unified anomaly scoring mechanism that integrates global reconstruction errors with local frequency prediction discrepancies for robust and interpretable detection.
- Extensive experiments on four benchmarks confirm the superiority of our method over existing state-of-the-art approaches in both AUC and AP.
2 Related works
2.1 Traditional methods
Early deepfake detection and general anomaly identification [18] approaches primarily relied on handcrafted features and shallow classifiers to distinguish authentic facial images from manipulated ones. Zhang et al. [19] provided a comprehensive review of deepfake audio detection, summarizing forgery mechanisms, detection pipelines, datasets, and evaluation protocols while also discussing emerging challenges related to privacy, robustness, and fairness. Traditional visual forgery detection typically focused on spatial inconsistencies, illumination artifacts, or codec-induced traces. For instance, Li et al. [20] exploited physiological cues such as eye blinking and head pose abnormalities, while Matern et al. [21] examined texture descriptors and facial color statistics. In the frequency domain, Durall et al. [22] demonstrated that upsampling operations in generative models introduce distinctive spectral artifacts that can be used for detection.
To overcome the brittleness of handcrafted features under compression, perturbations, or unseen manipulation types, reconstruction-based anomaly detection frameworks gained popularity. Autoencoders [8], DeepSVDD [9], and GANomaly [10] learn compact representations of normal samples and identify anomalies through reconstruction deviation or embedding distance. Further developments such as AnoGAN [23], DRAEM [24], and PatchCore [25] enhanced robustness using adversarial priors or patch-level reasoning. Nevertheless, purely reconstruction-driven models often fail to capture subtle, high-frequency inconsistencies characteristic of realistic deepfake manipulations, motivating the exploration of frequency-aware and self-supervised strategies.
2.2 Deep learning methods
Recent advances in deep learning have significantly improved the ability to capture spatial, temporal, and spectral irregularities in manipulated content. Early CNN-based models such as MesoNet [26] demonstrated strong performance on low-resolution datasets, while RNN-based architectures [27] incorporated temporal dynamics across video frames. Frequency-guided architectures, including F3-Net [12] and FreqFake [22], explicitly modeled Fourier-domain anomalies caused by generative upsampling, improving robustness under cross-manipulation and compression settings. To further enhance generalization ability, Yang et al. [28] introduced CSTAN, which integrates multi-dimensional attention and dynamic convolution to better resist previously unseen manipulation types.
More recently, several studies have explored fine-grained spatial–frequency cues and noise-sensitive representations to improve robustness against post-processing and unseen manipulations. Chen et al. [29] proposed a signal–noise separation strategy to suppress content-related interference and emphasize manipulation artifacts, while Liao et al. [30] leveraged subtle facial muscle motion patterns to enhance deepfake detection under heavy compression. Related investigations in multimedia forensics further indicate that manipulations introduce non-local structural and frequency inconsistencies that are difficult to capture using single-domain representations alone [31,32]. Beyond supervised paradigms, self-supervised and contrastive learning methods have emerged as powerful alternatives for representation learning without relying on labeled fake samples. CLFD [33] leveraged contrastive objectives to learn manipulation-agnostic features, whereas FATNet [34] incorporated frequency-sensitive transformers to emphasize cross-scale spectral inconsistencies. Generative and diffusion-based methods, including DiffusionAD [35] and Reverse Distillation [36], have also shown promise in modeling normality with stronger generative priors. Meanwhile, flow-based approaches such as FastFlow [37], CutPaste [38], and CFlow [39] demonstrated the effectiveness of density estimation and self-distillation for visual anomaly detection. Additional works, such as Walczyna et al. [40], examined the destructive effects of deepfake face swapping on digital watermarks, revealing how manipulation propagates non-locally across image regions—further motivating techniques capable of capturing global–local inconsistency. In additional, very recent studies published in 2025 have further explored hybrid spatial–frequency modeling and transformer-based architectures for deepfake detection. Liu et al. [41] proposed a frequency-aware vision transformer that explicitly captures long-range spectral dependencies to improve robustness against high-quality forgeries. Similarly, Chen et al. [42] introduced a hybrid CNN–Transformer framework that jointly encodes spatial textures and frequency residuals under a supervised learning paradigm. In addition, Stamnas et al. [43] proposed DiffFake, which transforms the detection task into “differential anomaly detection.” First, they trained the EfficientNet-b4 backbone with self-mixed pseudo Deepfake (simultaneously implanting global and local artifacts) to make the model sensitive to forgery traces. Then, they extracted 1792-d face embeddings and performed (A–B)² fusion using only real image pairs with the same identity, and learned the “natural variation” distribution using a three-component Gaussian mixture model.
Despite these advances, most existing deepfake detection approaches require large-scale labeled fake datasets or lack mechanisms to jointly model spatial and spectral consistencies. In contrast, our proposed dual-domain self-supervised framework addresses these limitations by integrating shared-weight spatial–frequency encoding, multi-feature prediction, and unified global–local anomaly scoring. This design enables the model to effectively capture both structural context and high-frequency artifacts, providing strong detection performance and superior cross-dataset generalization without explicit fake supervision.
2.3 Anomaly detection-based deepfake detection methods
Anomaly detection has become an important and rapidly developing research direction in the field of deepfakes in recent years, especially in addressing the problem of insufficient generalization ability of traditional supervised classification methods in open set scenarios, which has shown significant advantages.
Reconstruction-based methods train reconstruction models (such as autoencoders) using only real faces, allowing them to learn the distribution of real data. During testing, real images are typically reconstructed accurately, while forged images, due to containing aberrant tampering traces, often produce significant reconstruction errors. Deep forgery detection can be achieved by analyzing the differences between the original and reconstructed images. These methods do not require a large number of forged samples and have good generalization ability against unknown forgery methods. Shi et al. [44] proposed FRG2D, an end-to-end face reconstruction-based framework for generalized deepfake detection. The method learns the distribution of authentic faces through reconstruction and identifies manipulated images by analyzing reconstruction discrepancies. To enhance feature extraction, FRG2D incorporates CBAM, CAB, and a Residual Outlook Attention (ROA) module. Experimental results demonstrate that FRG2D achieves strong generalization performance across unseen deepfake domains.
One-class classification methods are trained exclusively on authentic facial images to learn the feature distribution or decision boundary of real data. During inference, samples that deviate from this learned distribution are regarded as anomalies and are therefore classified as deepfakes. Since these methods do not rely on manipulated samples for supervised training, they generally exhibit strong generalization ability to unseen forgery methods. Khalid et al. [45] proposed OC-FakeDect, a one-class deepfake detection method that trains a Variational Autoencoder (VAE) solely on authentic facial images and identifies manipulated images as anomalies, enabling effective generalization to unseen forgery methods. Cao et al. [46] proposed OC-SAN, a one-class style autoencoder network for deepfake detection. The model learns identity and personalized style representations from authentic facial images and reconstructs only real faces effectively. Manipulated images are identified based on abnormal reconstruction patterns, enabling strong generalization to unseen forgery methods without requiring fake samples during training.
Feature distribution modeling methods learn the statistical distribution of authentic facial features and classify samples that deviate significantly from this distribution as deepfakes. By modeling the global structure of real facial representations, these methods generally exhibit strong generalization to unseen forgery methods. Khodabakhsh et al. [47] proposed a two-stage anomaly detection method that learns the distribution of pristine facial images and detects manipulated images based on deviations from the estimated pixel-wise probabilities, achieving strong generalization to unseen synthesis methods. Maiano et al. [48] proposed two out-of-distribution (OOD) detection methods for generalized deepfake detection. The first method detects manipulated images through reconstruction errors, while the second incorporates an attention mechanism to enhance anomaly localization. Both methods are designed to identify samples that deviate from the distribution of authentic images, demonstrating strong robustness to unseen deepfake generation methods.
3 Methodology
In this section, we introduce the proposed SDDA (Dual-domain Feature Alignment for Deepfake Anomaly Detection) framework. The SDDA model is designed as a self-supervised anomaly detection system that jointly learns spatial-frequency consistency and multi-feature prediction, without requiring fake data during training. The overall framework is depicted in Fig 2. SDDA consists of three key components: (1) a dual-domain encoder with shared weights for consistent representation learning, (2) a multi-feature prediction head for structural and frequency-based feature regression, and (3) a unified loss function that integrates feature reconstruction, domain consistency, and hypersphere compactness. The architecture contains three core components: (1) a shared-weight spatial–frequency encoder that learns domain-aligned representations through joint feature extraction; (2) a multi-objective feature-prediction head that estimates the four consistency descriptors (,
,
,
), enforcing fine-grained spatial–frequency correspondence and enhancing cross-domain discriminability; (3) a global–local anomaly scoring module that integrates multi-scale reconstruction with attention-guided deviation estimation to capture both structural and spectral abnormalities. By explicitly aligning dual-domain embeddings and constraining their consistency through self-supervision, the framework establishes a theoretically grounded representation space that generalizes effectively to unseen manipulation types.
3.1 Dual-domain encoder with shared weights
In this section, we present the proposed Self-Supervised Dual-Domain Alignment (SDDA) framework for deepfake anomaly detection. Unlike conventional supervised methods that rely on large quantities of forged samples, SDDA is trained exclusively on real facial images and aims to learn domain-invariant, frequency-aware representations through a combination of spatial–spectral feature coupling, multi-objective self-supervised prediction, and hypersphere regularization. The complete pipeline is shown in Fig 2.
SDDA consists of four tightly coupled modules:
- A data preprocessing module that extracts semantic features from a pretrained FAN network and computes the corresponding frequency-domain representation.
- A dual-domain shared-weight encoder that jointly encodes spatial and spectral representations.
- A multi-feature prediction head that enforces four complementary self-supervised constraints.
- A multi-objective optimization mechanism integrating feature prediction, domain alignment, and hypersphere compactness for anomaly scoring.
3.2 Data preprocessing: Semantic feature extraction and spectral transformation
Throughout this paper, we denote the input face image in the spatial (RGB) domain as . Its corresponding frequency-domain representation, obtained via a two-dimensional discrete Fourier transform (DFT), is denoted as
. Unless otherwise specified,
and
always refer to raw spatial images and their frequency spectra, respectively. Feature maps extracted by the pretrained FAN network are explicitly denoted as
and
to avoid ambiguity.
We adopt the fast Fourier transform (FFT) as the frequency-domain transformation for two primary reasons. First, FFT explicitly reveals periodic patterns and high-frequency artifacts introduced by common deepfake generation operations, such as upsampling, blending, and post-processing compression, which are difficult to capture purely in the spatial domain. Second, FFT provides a deterministic and parameter-free spectral decomposition, avoiding additional learnable parameters that may bias the normality modeling process in self-supervised anomaly detection. Compared to alternative frequency representations such as wavelet transforms or learnable spectral filters, FFT offers superior computational efficiency and stable frequency responses, making it well-suited for large-scale training and cross-dataset generalization. This design choice allows SDDA to effectively leverage spatial–spectral consistency while maintaining robustness to unseen manipulation patterns.
3.3 Dual-domain shared-weight encoder
The dual-domain encoder aims to learn a domain-invariant representation that aligns spatial and frequency semantics. A single encoder with shared parameters is applied to both domains:
where denote spatial and frequency embeddings.
Weight sharing forces the encoder to capture cross-domain semantic invariances and suppress domain-specific artifacts. This improves the robustness and generalization of the learned embedding space, especially under unseen manipulation types.
3.4 Self-supervised descriptor prediction module
The four self-supervised descriptors are originally defined on the paired spatial and frequency representations and
, as summarized in Table 1. To keep the formulation consistent with the original image-level definition, we do not redefine these descriptors at the feature level. Instead, we instantiate
and
with the response maps extracted from different layers of the two encoders during implementation.
Given the feature maps and
from the
-th layer of the spatial and frequency encoders, we first obtain their channel-averaged response maps:
Here, and
are regarded as layer-wise instantiations of the original spatial and frequency representations
and
. Therefore, the following descriptors preserve the same mathematical meaning as the image-level definitions, while allowing them to be computed from the selected encoder layers.
Local Structural Deviation. Local Structural Deviation (LSD) measures local structural inconsistency between the spatial and frequency representations. Following the original definition, LSD is computed as the average gradient discrepancy:
where and
denote the response maps extracted from the second encoder layer, and N2 is the number of spatial positions in the response map. The gradient operator ∇ is implemented by local horizontal and vertical gradient filters. This descriptor encourages the model to preserve fine-grained local structural differences between the two domains.
Global Pattern Discrepancy. Global Pattern Discrepancy (GPD) captures the global distribution difference between the spatial and frequency representations. Consistent with the original image-level formulation, GPD is computed as the discrepancy between the mean responses of the two domains:
where and
are obtained from the third encoder layer. Compared with LSD, GPD focuses more on holistic response distribution and global cross-domain consistency.
Local Residual Difference. Local Residual Difference (LRD) measures residual-level inconsistency between the spatial and frequency representations. Following the original definition, we first obtain the locally smoothed responses:
where denotes average pooling with an
kernel. The LRD descriptor is then computed as:
This descriptor emphasizes local residual differences and is useful for capturing subtle manipulation traces that may be weakened in global representations.
Total Consistency Deviation. Total Consistency Deviation (TCD) measures the consistency deviation between deeper spatial and frequency representations. In accordance with the original definition, TCD is computed over multiple encoder layers. In our implementation, we use the second, third and fourth layers, denoted as . For each layer, the feature representation is obtained by global average pooling:
The TCD descriptor is formulated as:
This descriptor provides a multi-level consistency constraint by aggregating deviations from different semantic layers.
The four descriptors are then concatenated to form the self-supervised target vector:
A prediction head takes the concatenated spatial–frequency embedding as input:
and predicts the descriptor vector:
The descriptor prediction loss is formulated as:
where , and
denotes the loss weight of each descriptor. In our implementation, all
are set to 1 unless otherwise specified. To avoid scale imbalance among different descriptors, each descriptor is normalized before loss computation.
3.5 Domain alignment loss
To explicitly enforce cross-domain consistency, we introduce an alignment loss defined as the squared Euclidean distance between spatial and frequency embeddings:
This objective encourages the shared-weight encoder to learn domain-invariant representations across spatial and spectral inputs.
3.6 Hypersphere compactness constraint
Inspired by DeepSVDD [9], SDDA regularizes normal samples into a compact hypersphere. The hypersphere center is defined as:
The compactness constraint is:
This hypersphere formulation increases separation between normal and anomalous embeddings in the joint latent space.
3.7 Total optimization objective
SDDA combines the three complementary objectives into a unified loss:
where controls the trade-off between self-supervised prediction accuracy and structural embedding consistency.
3.8 Anomaly scoring
The final anomaly score integrates both global hypersphere deviation and local self-supervised feature prediction error:
where denotes the latent embedding,
is the hypersphere center, and
and
are the predicted and target self-supervised features, respectively.
The SDDA training procedure is summarized in Algorithm 2.
3.9 Design rationale and key innovations scoring
To clarify the novelty of our framework, SDDA is not a simple integration of existing components. Instead, it establishes a unified dual-domain self-supervised learning paradigm, where spatial and frequency representations are explicitly aligned through a consistency constraint rather than directly fused. In addition, a dual-domain contrastive objective is designed to enforce compactness of real samples while implicitly separating manipulated ones as out-of-distribution deviations. Finally, a hypersphere constraint is applied on the aligned representation space to further enhance intra-class compactness and inter-class separability. These designs jointly enable SDDA to effectively capture subtle forgery traces across both domains.
4 Experiments
In this section, we describe the datasets, evaluation metrics, baseline methods, and implementation details used in our study. We then comprehensively evaluate the performance of the proposed SDDA framework across multiple public Deepfake anomaly detection benchmarks. Furthermore, ablation studies are conducted to analyze the contribution of each module, including the dual-domain encoder, feature prediction head, and unified loss components. Finally, qualitative visualizations and statistical analyses are provided to demonstrate the interpretability and practical effectiveness of SDDA in detecting subtle and unseen Deepfake manipulations. The general training process of our proposed SDDA model is shown in Table 2.
4.1 Datasets
To rigorously assess the efficacy and generalization of SDDA, we performed experiments on five widely used public datasets: FaceForensics++ [5], Celeb-DF [2], Flickr-Faces-HQ (FFHQ) [15], WildDeepfake [16], and ForgeryNet [17]. Collectively, these datasets cover a wide range of Deepfake generation techniques, compression levels, and acquisition conditions, providing a diverse benchmark for evaluating both spatial and frequency-domain anomaly detection capabilities. Brief information about the datasets used in this study is shown in Table 3.
FaceForensics++ (FF++): FF++ contains 1,000 original video sequences sourced from YouTube and 977 processed video clips. Each clip typically features a single face with minimal occlusion or motion blur, facilitating stable face tracking and manipulation. Four manipulation methods (DeepFakes, Face2Face, FaceSwap, NeuralTextures) are applied to generate corresponding forged videos. For our experiments, we extracted one frame every 10 frames to construct the image dataset.
Celeb-DF: Celeb-DF is a challenging large-scale dataset containing diverse facial identities and high-quality manipulations. Compared with previous datasets, Celeb-DF exhibits fewer visual artifacts and more realistic facial dynamics, making it ideal for evaluating model generalization on subtle and unseen forgeries.
Flickr-Faces-HQ (FFHQ): FFHQ consists of 70,000 high-resolution () PNG images of real human faces collected from Flickr. The dataset contains substantial diversity in age, ethnicity, facial expression, and illumination conditions, and is therefore well suited for modeling normal facial appearance.
In our anomaly detection setting, FFHQ is used exclusively as a source of real samples for training. Since FFHQ does not provide native forged images, fake samples used in the validation and test phases are constructed by incorporating manipulated face images from external Deepfake datasets (e.g., FaceForensics++ and Celeb-DF). These fake samples are never used during training, and are only introduced at evaluation time to assess the model’s ability to distinguish real faces from unseen forgeries.
WildDeepfake: WildDeepfake comprises 7,314 face sequences from 707 Deepfake videos collected from the internet. It features uncontrolled lighting, compression artifacts, and motion blur, representing real-world challenges for forgery detection. This dataset is primarily used to evaluate SDDA’s robustness in-the-wild.
ForgeryNet: ForgeryNet is a large-scale face forgery benchmark containing 1.44 million real images, 1.46 million fake images, and over 220,000 videos from more than 5,400 identities. It includes 15 manipulation methods across face swap, face transfer, face reenactment, and face editing, with 36 perturbation types to simulate real-world distortions. The dataset provides annotations for image and video forgery classification, as well as spatial and temporal forgery localization. In our experiments, we use the image subset and extract one frame every 10 frames to construct the dataset.
Data Preprocessing and Partitioning: For video datasets (FF++, Celeb-DF, WildDeepfake, ForgeryNet), frames were uniformly sampled at a rate of 1/10. For image-based datasets (FFHQ), all images were directly used. Data were split into training, validation, and test sets at a 6:2:2 ratio. The training set contains only authentic samples, while validation and test sets include a balanced mix of real and fake samples.
4.2 Experimental design and evaluation metrics
4.2.1 Experimental setup.
Data Preprocessing. All input face images are first detected and aligned using a standard facial alignment pipeline. The aligned RGB images are resized to a fixed spatial resolution and normalized to [0,1] before being fed into the model. For frequency-domain modeling, we apply a two-dimensional discrete Fourier transform (DFT) to each color channel independently and retain the amplitude spectrum as the spectral representation. To reduce numerical instability, the amplitude spectra are logarithmically scaled and normalized across samples.
Self-supervised Feature Targets. High-level semantic feature maps are extracted from a frozen pretrained facial alignment network (FAN). Based on these feature maps, four self-supervised regression targets are constructed, corresponding to local structural deviation, global pattern discrepancy, local residual difference, and total consistency deviation. These targets serve as supervision signals during training and are not used at inference time.
Network Architecture and Training Details. The dual-domain encoder adopts a lightweight four-layer convolutional architecture with shared weights across spatial and frequency branches. The specific configuration parameters of the encoder backbone and prediction head are shown in Table 4. Each convolutional layer is followed by LeakyReLU activation and batch normalization. The latent embedding dimension is fixed to Z = 8 for all experiments. A shallow multi-layer perceptron (MLP) prediction head is appended to regress the four self-supervised feature targets. All models are trained using the Adam optimizer with a learning rate of and a batch size of 128. Training is conducted for up to 1000 epochs with early stopping based on validation performance and a patience of 200 epochs. To ensure statistical robustness, all experiments are repeated using six different random seeds, and the mean results are reported.
The encoder employs a lightweight convolutional architecture optimized for compact representation learning.
It consists of four sequential 2D convolutional layers with progressively increasing channels, interleaved with non-linear activation and normalization layers. Both spatial-domain RGB images and frequency-domain amplitude spectra are processed in their native two-dimensional form, without reshaping into one-dimensional sequences.
Spatial and frequency branches share identical encoder weights to enforce domain-consistent embeddings. A shallow MLP-based prediction head is appended for multi-feature regression, composed of fully connected layers with LeakyReLU activation.
4.2.2 Evaluation metrics.
The anomaly detection performance was quantified using two widely adopted metrics: AUC (Area Under the ROC Curve) and AP (Average Precision), in accordance with prior works [9,10]. Each experiment was repeated six times with different random seeds, and both mean and standard deviation values are reported to ensure robustness. A higher AUC indicates better separability between normal and anomalous samples, while a higher AP reflects improved ranking consistency of anomaly scores. Together, these metrics provide a comprehensive evaluation of both model discriminability and generalization across datasets and manipulation types.
4.3 Baselines
To comprehensively evaluate the performance of the proposed SDDA framework, we selected twelve state-of-the-art baseline models spanning reconstruction-, adversarial-, flow-, and diffusion-based anomaly detection approaches, which are widely adopted in visual anomaly detection and Deepfake forensics. The details are summarized below:
- AE-CNN [8] is a classical unsupervised reconstruction-based method that learns to reconstruct normal facial images. Anomalies are detected by measuring pixel-wise reconstruction errors.
- DeepSVDD [9] is an unsupervised one-class classification approach that maps normal samples to a compact hypersphere in latent space. Samples with large deviations from the hypersphere center are considered anomalous.
- GANomaly [10] employs an encoder–decoder–encoder architecture to enforce bidirectional feature consistency, using reconstruction and latent-space discrepancies to detect anomalies.
- AnoGAN [23] reconstructs input images using a pretrained GAN generator via iterative latent-space optimization. The anomaly score combines reconstruction and discriminator losses.
- DRAEM [24] introduces self-supervised reconstruction using synthetically corrupted images and trains a dual-branch network to restore and discriminate local anomalies.
- PatchCore [25] leverages memory banks of patch-level embeddings extracted from normal samples. During inference, nearest-neighbor distances are used to detect localized anomalies.
- FastFlow [37] applies normalizing flows to map high-dimensional visual features into a Gaussian latent space, enabling efficient detection and localization based on feature density.
- CutPaste [38] is a self-supervised augmentation technique that generates synthetic anomalies via patch cutting and pasting, training models to distinguish real from altered images.
- Reverse Distillation [36] employs a teacher–student self-distillation mechanism where reconstruction errors in intermediate feature layers indicate anomaly likelihood.
- CFlow [39] utilizes conditional normalizing flows to estimate feature-space likelihoods, achieving high efficiency and real-time performance for facial anomaly detection.
- FATNet [34] is a frequency-aware transformer tailored for Deepfake detection, leveraging dual-domain joint attention to capture subtle cross-domain manipulation traces.
- DiffusionAD [35] is a diffusion-based anomaly detection model that iteratively denoises input images to reconstruct clean samples, enabling robust detection of subtle manipulations.
- RepDFD [49] repurposes a pre-trained Vision-Language Model (e.g., CLIP) for deepfake detection without altering the model’s internal parameters. It utilizes learnable visual perturbations and adaptive text prompts to guide the optimization process, significantly improving generalization and efficiency in deepfake detection.
All baseline models were trained under identical data splits and preprocessing pipelines to ensure fair comparison. Only authentic (non-manipulated) facial data were used during training, while both real and fake samples were employed for validation and testing.
4.4 Experimental results
4.4.1 Performance analysis.
In this section, we quantitatively evaluate the anomaly detection performance of our proposed SDDA framework. The results of all the compared methods on the five benchmark datasets are summarized in Tables 5 and 6. The proposed SDDA consistently achieves the highest AUC and AP values across all datasets, demonstrating superior generalization to diverse forgery types and compression levels. On the FaceForensics dataset, SDDA attains an AUC of 82.36% and an AP of 80.41%, outperforming the strongest baseline DiffusionAD by 7.71% and 13.01%, respectively. On the more challenging Celeb-DF dataset, which contains complex post-processing and realistic facial reenactments, SDDA achieves 81.56% AUC and 78.47% AP, surpassing DiffusionAD by 13.84% in AUC and 11.58% in AP, indicating its robustness under strong compression and illumination variations. In the Flickr-Faces-HQ dataset, SDDA reaches 73.78% AUC and 71.38% AP, improving over the best-performing baseline by 18.44% in AUC and 9.26% in AP, highlighting its capability to capture fine-grained local frequency inconsistencies. Similarly, on the WildDeepfake dataset, which includes uncontrolled in-the-wild conditions, SDDA achieves 76.45% AUC and 72.31% AP, exceeding DiffusionAD by 9.33% and 8.09%, respectively.
These results clearly demonstrate the effectiveness of the proposed dual-domain self-supervised learning paradigm. By jointly modeling spatial and frequency representations through shared-weight encoders, SDDA captures both global structural patterns and local spectral inconsistencies, enabling stronger discrimination of subtle manipulations. Unlike conventional autoencoder- or GAN-based frameworks that focus solely on spatial reconstruction, SDDA introduces a self-supervised feature prediction strategy that estimates four domain-consistent descriptors (,
,
,
), thereby enhancing the semantic alignment between the spatial and frequency domains. This design not only improves detection accuracy but also enhances interpretability by linking feature deviations to specific spectral or spatial distortions.
Furthermore, SDDA achieves superior performance without relying on labeled fake data or handcrafted frequency fusion. This property underscores its potential for real-world deployment in open-world deepfake forensics, where unseen manipulation types frequently emerge. Overall, SDDA achieves a robust balance between discriminability, generalization, and interpretability, outperforming existing state-of-the-art deepfake anomaly detection methods across all evaluation benchmarks.
4.4.2 Ablation study.
To evaluate the contribution of each key component in the proposed SDDA framework, we conduct comprehensive ablation experiments on all five benchmark datasets: FaceForensics++, Celeb-DF, Flickr-Faces-HQ, WildDeepfake, and ForgeryNet. Five ablation variants are designed by selectively removing individual modules or objectives, as described below:
- wo-F: The frequency-domain branch is removed, retaining only the spatial encoder, thereby eliminating the spatial–spectral consistency constraint.
- wo-SD: The Local Structural Deviation (LSD) prediction task is disabled, preventing the model from enforcing local structural alignment between spatial and frequency representations.
- wo-PD: The Global Pattern Discrepancy (GPD) objective is removed, weakening the learning of global structural consistency.
- wo-RD: The Local Residual Difference (LRD) predictor is excluded, reducing the model’s sensitivity to fine-grained local anomalies.
- wo-TD: The Total Consistency Deviation (TCD) constraint is removed, disabling unified cross-domain consistency calibration.
As presented in Tables 7 and 8, the removal of any individual module results in a noticeable degradation in both AUC and AP across all datasets, confirming the contribution of each component to the overall performance of SDDA. The most substantial decline occurs when the frequency-domain branch is removed, with average decreases of 12.83% (AUC) and 12.75% (AP) across datasets, highlighting the critical role of frequency-domain cues in capturing subtle spectral artifacts introduced by manipulations. Among the four self-supervised predictive features, eliminating TCD leads to the second-largest performance drop (average decreases of 8.84% AUC and 8.65% AP), underscoring its importance in integrating cross-domain information. LSD and GPD primarily contribute to learning local and global structural patterns, whereas LRD and TCD reinforce consistency-based discrimination. Overall, the complete SDDA framework achieves the highest AUC and AP across all datasets, demonstrating that the collaborative effect of all components enhances dual-domain feature alignment, anomaly sensitivity, and the compactness and discriminability of learned representations.
4.4.3 Generalization experiments.
In this section, we evaluate the generalization performance of the SDDA model. We used data from the ForgeryNet dataset as the training dataset and the FaceForensics and Celeb-DF datasets as the test datasets. Similarly, we used data from the Flickr-Faces-HQ dataset as the training dataset and the WildDeepfake dataset as the test dataset. The results of the generalization experiments are shown in the Table 9 below.
The results show that the SDDA model exhibits superior generalization performance compared to the models listed. Overall, when trained on the ForgeryNet dataset and tested on the Celeb-DF dataset, the SDDA model maintains good performance, outperforming other comparative models. However, when trained on the Flickr-Faces-HQ dataset and tested on the WildDeepfake dataset, the performance of the SDDA model decreases compared to training and testing directly on the WildDeepfake dataset. Nevertheless, this model still outperforms other comparative models. We speculate that the performance degradation is due to significant data differences between the Flickr-Faces-HQ dataset, which primarily uses images, and the WildDeepfake dataset, which processes videos into images.
For visualization, we selected four representative baseline methods: DeepSVDD, PatchCore, DiffusionAD, and SDDA. These methods represent three major paradigms in anomaly detection: hypersphere-based (DeepSVDD), memory-anchored patch reasoning (PatchCore), diffusion-driven reconstruction (DiffusionAD), and our proposed dual-domain self-supervised feature alignment framework (SDDA). To ensure a comprehensive evaluation, we conducted visual analyses across five benchmark datasets: FaceForensics++, Celeb-DF, Flickr-Faces-HQ (FFHQ), WildDeepfake, and ForgeryNet. This setup offers an extensive view of how each model generalizes to various real-world manipulation scenarios. While the quantitative superiority of SDDA has been confirmed in previous experiments, the visualization results further demonstrate that SDDA produces more compact, structured, and separable representations compared to the other representative baselines.
- Feature embedding visualization and analysis: To assess the discriminability of the learned representations, we visualized the latent embeddings of the test samples using the t-SNE algorithm. As illustrated in Fig 3, green points correspond to real (authentic) samples, while red points denote fake (manipulated) samples. The visualization is conducted separately on five datasets—FaceForensics++, Celeb-DF, FFHQ, WildDeepfake, and ForgeryNet—and the results are compared across four methods: DeepSVDD, PatchCore, DiffusionAD, and SDDA (from left to right). The embeddings of DeepSVDD show weak separation between normal and fake regions, indicating that hypersphere mapping struggles to capture complex facial manipulations. PatchCore demonstrates moderate clustering improvement due to local patch memory matching but still suffers from overlapping regions between normal and abnormal points. DiffusionAD generates smoother embeddings yet exhibits partial entanglement between fake and real clusters, reflecting its diffusion-driven stochastic reconstruction behavior. In contrast, SDDA forms distinct, compact clusters with clearly delineated boundaries between authentic and manipulated samples. This result confirms that the proposed dual-domain feature alignment and self-supervised prediction module effectively enhance both representation compactness and discrimination capability across diverse datasets.
- Anomaly distribution visualization and analysis: To further analyze the separability of normal and abnormal samples, we visualize the anomaly score distributions obtained by the four methods on the same datasets, as shown in Fig 4. Each model computes anomaly scores using its respective mechanism—distance from the hypersphere center (DeepSVDD), feature deviation from memory patches (PatchCore), reconstruction residuals (DiffusionAD), and the unified global–local deviation score (SDDA). In these histograms, green bars indicate normal samples and red bars indicate fake samples. The results show that SDDA exhibits two clearly separated peaks with minimal overlap, whereas DeepSVDD and PatchCore yield wide overlapping areas, and DiffusionAD presents a partially mixed distribution due to its generative noise. SDDA’s sharp boundary between normal and fake score distributions demonstrates its superior capability to learn consistent, domain-aligned representations and accurate anomaly boundaries, making it robust against unseen manipulations across all five benchmark datasets.
Green points represent real (normal) samples, and red points represent fake (abnormal) samples. SDDA produces the most compact and separable clusters across all datasets, highlighting its superior dual-domain representation alignment and anomaly discriminability.
Green bars represent real (normal) samples, and red bars represent fake (abnormal) samples. SDDA produces the most compact and separable clusters across all datasets, highlighting its superior dual-domain representation alignment and anomaly discriminability.
4.4.4 Parameter sensitivity analysis.
We conducted a parameter sensitivity analysis to evaluate the influence of the weighting coefficient in the SDDA loss function (Eq. 16) on anomaly detection performance across the five benchmark datasets. The coefficient
balances the contribution between the feature prediction loss and the combined domain alignment and hypersphere compactness losses.
As shown in Fig 5, SDDA achieves optimal performance on the FaceForensics++, Celeb-DF, and FFHQ datasets when , indicating a balanced trade-off between feature prediction and domain alignment objectives yields the best discriminability and generalization. For the WildDeepfake dataset, which contains highly diverse in-the-wild conditions, the best performance is observed at
, suggesting a slightly higher emphasis on the feature prediction loss is beneficial under more challenging scenarios. Overall, the results demonstrate that SDDA exhibits stable performance across a reasonable range of
values, confirming the robustness of the proposed dual-domain self-supervised learning framework to hyperparameter variations.
4.5 Computational analysis
We further performed a computational complexity analysis to assess the efficiency and deployability of SDDA in practical deepfake anomaly detection scenarios. Table 10 summarizes the number of parameters, FLOPs, and average training/testing time per epoch and per batch for three representative baseline models and our SDDA framework. The results indicate that SDDA maintains a moderate parameter size (0.96 MB) and computational cost (4.12 G FLOPs), which is significantly lower than that of DiffusionAD (7.85 G) while providing faster inference (0.78 s). Although PatchCore achieves slightly faster testing speed (0.65 s), it suffers from lower feature expressiveness due to its non-learnable memory bank design. Overall, SDDA achieves a balanced trade-off between accuracy and computational efficiency, making it suitable for real-time deployment in resource-constrained or embedded deepfake forensics applications.
4.6 Failure cases and limitations
Despite the strong empirical performance of SDDA across multiple benchmarks, several limitations merit further discussion. One limitation arises when the input faces suffer from severe quality degradation, such as heavy compression, strong motion blur, or extreme downsampling. Under these conditions, both spatial textures and frequency amplitudes may be significantly distorted, thereby weakening the effectiveness of spatial–frequency consistency modeling. In addition, identity-preserving manipulations that introduce only subtle or highly localized artifacts remain challenging. Advanced facial reenactment techniques that closely mimic authentic facial statistics may exhibit limited spectral deviation, reducing the anomaly margin learned during self-supervised training. Moreover, the adopted FFT-based frequency representation primarily captures global spectral characteristics. While this design is effective for identifying common generative artifacts, it may be less sensitive to localized or non-periodic distortions. Incorporating multi-scale or adaptive frequency decomposition strategies could further enhance sensitivity to fine-grained anomalies. Finally, the hypersphere-based anomaly modeling assumes a relatively compact distribution of normal samples in the latent space. Although this assumption works well in practice, it may restrict expressiveness when normal data exhibit complex or multi-modal distributions. Exploring more flexible density modeling approaches constitutes an important direction for future research.
Conclusion
In this work, we introduced SDDA, a novel dual-domain feature alignment framework for deepfake anomaly detection. Unlike conventional supervised or reconstruction-based approaches, SDDA leverages a self-supervised learning paradigm that jointly enforces spatial-frequency consistency while predicting multiple discriminative features, without requiring any forged samples during training. Extensive experiments on five benchmark datasets demonstrate that SDDA consistently outperforms state-of-the-art anomaly detection methods in both AUC and AP metrics. These results highlight its superior robustness, generalization to unseen manipulations, and sensitivity to subtle spatial and spectral discrepancies. Our findings validate the effectiveness of explicitly aligning spatial and spectral representations for anomaly detection and underscore the potential of dual-domain self-supervised learning as a foundation for future research. This framework opens avenues for developing more efficient, interpretable, and scalable deepfake detection systems, particularly in real-world and open-world scenarios.
References
- 1. Zhang H, Hu C, Min S, Sui H, Zhou G. TSFF-Net: A two-stream feature domain fusion network for deepfake video detection. PLOS ONE. 2024;19(12):e0311366.
- 2.
Li Y, Yang X, Sun P, Qi H, Lyu S. Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2020. p. 3207–16.
- 3. Tolosana R, Vera-Rodriguez R, Fierrez J, Morales A, Ortega-Garcia J. Deepfakes and beyond: A Survey of face manipulation and fake detection. Inf Fusion. 2020;64:131–48.
- 4. Yadav S, Tiwari N, Singh P. Combining spatial and temporal cues with CNN–BiLSTM–Transformer for robust video deepfake detection. PLOS ONE. 2025;20(5):e0334980.
- 5.
Rössler A, Cozzolino D, Verdoliva L, Riess C, Thies J, Nießner M. FaceForensics++: Learning to Detect Manipulated Facial Images. In: Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV). 2019. p. 1–11.
- 6.
Nguyen HH, Yamagishi J, Echizen I. Capsule-Forensics: Using Capsule Networks to Detect Forged Images and Videos. In: Proc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP). 2019. p. 2307–11.
- 7. Zhao Y, Wang H, Chen Y, Yu Y. Multi-Attentional Deepfake Detection. IEEE Trans Circuits Syst Video Technol. 2021;32:600–14.
- 8.
Bergmann P, Fauser M, Sattlegger D, Steger C. MVTec AD—A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2019. p. 9592–600.
- 9.
Ruff L, Vandermeulen R, Görnitz N, Deecke L, Siddiqui S, Binder A, et al. Deep One-Class Classification. In: Proc. Int. Conf. Mach. Learn. (ICML). 2018. p. 4393–402.
- 10.
Akçay S, Atapour-Abarghouei A, Breckon TP. GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training. In: Asian Conf. Comput. Vis. (ACCV). 2018. p. 622–37.
- 11.
Frank J, Eisenhofer T, Schönherr L, Fischer A, Kolossa D, Holz T. Leveraging Frequency Analysis for Deep Fake Image Recognition. In: Proc. ICML Workshops. 2020.
- 12.
Qian Y, Dong J, Wang W, Tan T, Yang X, Li S. Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues. In: Eur. Conf. Comput. Vis. (ECCV). 2020. p. 86–103.
- 13. Liu Y, Liu J, Huang Y, Liu T. Spatial–Frequency Domain Correlation Learning for Face Forgery Detection. Pattern Recognit. 2021;120:108157.
- 14. Zhang H, Hu C, Min S, Sui H, Zhou G. TSFF-Net: A Two-Stream Feature Fusion Approach for Deepfake Video Forgery Detection. PLOS ONE. 2024;19(12):e0311366.
- 15.
Karras T, Laine S, Aila T. A Style-Based Generator Architecture for Generative Adversarial Networks. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2019. p. 4401–10.
- 16.
Zi B, Chang M, Chen J, Ma X, Jiang Y. WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection. In: Proc. ACM Int. Conf. Multimedia (MM). 2020. p. 2382–90.
- 17.
He Y, Gan B, Chen S, Zhou Y, Yin G, Song L, et al. Forgerynet: A versatile benchmark for comprehensive forgery analysis. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. p. 4360–9.
- 18. Amerini I, Barni M, Battiato S, Bestagini P, Boato G, Guarnera L, et al. Deepfake media forensics: State of the art and open challenges. PLOS ONE. 2024;19(5):e0285890.
- 19. Zhang B, Cui H, Nguyen V, Whitty M. Audio deepfake detection: what has been achieved and what lies ahead. Sensors. 2025;25(7).
- 20.
Li Y, Lyu S. Exposing DeepFake Videos by Detecting Face Warping Artifacts. In: Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshops. 2018. p. 46–52.
- 21.
Matern F, Riess C, Stamminger M. Exploiting Visual Artifacts to Expose DeepFakes and Face Manipulations. In: Proc. IEEE Winter Conf. Appl. Comput. Vis. Workshops (WACVW). 2019. p. 83–92.
- 22.
Durall R, Keuper M, Keuper J. Watch Your Up-Convolution: CNN-Based Generative Deep Neural Networks Are Failing to Reproduce Spectral Distributions. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2020. p. 7890–9.
- 23.
Schlegl T, Seeböck P, Waldstein SM, Schmidt-Erfurth U, Langs G. Unsupervised Anomaly Detection with Generative Adversarial Networks to Guide Marker Discovery. In: Inf. Process. Med. Imaging (IPMI). 2017. p. 146–57.
- 24.
Zavrťanik V, Kristan M, Skočaj D. DRAEM: A Discriminatively Trained Reconstruction Embedding for Surface Anomaly Detection. In: Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV). 2021. p. 8330–9.
- 25.
Roth K, Pemula L, Zepeda J, Schiele B, Hospedales T, Ommer B. Towards Total Recall in Industrial Anomaly Detection. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2022. p. 14318–28.
- 26.
Afchar D, Nozick V, Yamagishi J, Echizen I. MesoNet: A Compact Facial Video Forgery Detection Network. In: Proc. IEEE Int. Workshop Inf. Forensics Secur. (WIFS). 2018. p. 1–7.
- 27.
Guera D, Delp EJ. Deepfake Video Detection Using Recurrent Neural Networks. In: Proc. IEEE Int. Conf. Adv. Video Signal Based Surveillance (AVSS). 2018. p. 1–6.
- 28. Yang R, You K, Pang C, Luo X, Lan R. CSTAN: A Deepfake Detection Network with CST Attention for Superior Generalization. Sensors (Basel). 2024;24(22):7101. pmid:39598878
- 29. Chen J, Liao X, Wang W, Qian Z, Qin Z, Wang Y. SNIS: A Signal Noise Separation-Based Network for Post-Processed Image Forgery Detection. IEEE Trans Circuits Syst Video Technol. 2023;33(2):935–51.
- 30. Liao X, Wang Y, Wang T, Hu J, Wu X. FAMM: Facial Muscle Motions for Detecting Compressed Deepfake Videos Over Social Networks. IEEE Trans Circuits Syst Video Technol. 2023;33(12):7236–51.
- 31. Fu L, Liao X, Guo J, Dong L, Qin Z. WaveRecovery: Screen-Shooting Watermarking Based on Wavelet and Recovery. IEEE Trans Circuits Syst Video Technol. 2024;35(4):3603–18.
- 32. Li Y, Liao X, Wu X. Screen-Shooting Resistant Watermarking With Grayscale Deviation Simulation. IEEE Trans Multimedia. 2024;26:10908–23.
- 33.
Sun Z, Yang L, Wang X, Yu J. CLFD: Contrastive Learning for Generalizable Deepfake Detection. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2023. p. 14012–22.
- 34. Wang Y, Zhang Z, Liu C, Zheng H, Wang J. FATNet: Frequency-Aware Transformer for Generalizable Deepfake Detection. IEEE Trans Inf Forensics Secur. 2024;19:1234–47.
- 35.
Li W, Chen K, Zheng J, Zhang H, Liu Y. DiffusionAD: Diffusion Model-Based Visual Anomaly Detection. arXiv preprint arXiv:240104511. 2024.
- 36.
Deng H, Li X, Xu C, Tian Y. Anomaly Detection via Reverse Distillation from One-Class Embedding. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR); 2022. p. 9737–46.
- 37.
Yu J, Zheng H, Xu D, Zhao P, Li H, Yang Y. FastFlow: Unsupervised Anomaly Detection and Localization via 2D Normalizing Flows. arXiv preprint arXiv:211107677. 2021.
- 38.
Li Y, Sohn K, Yoon J, Pfister T. CutPaste: Self-Supervised Learning for Anomaly Detection and Localization. In: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). 2021. p. 9664–74.
- 39.
Gudovskiy D, Ishizaka S, Kozuka K. CFlow-AD: Real-Time Unsupervised Anomaly Detection with Localization via Conditional Normalizing Flows. In: Proc. IEEE Winter Conf. Appl. Comput. Vis. (WACV). 2022. p. 98–107.
- 40. Walczyna T, Piotrowski Z. Mutual Effects of Face-Swap Deepfakes and Digital Watermarking-A Region-Aware Study. Sensors (Basel). 2025;25(19):6015. pmid:41094837
- 41. Liu Y, Zhang H, Wang J. Frequency-Aware Vision Transformer for Generalized Deepfake Detection. IEEE Trans Circuits Syst Video Technol. 2025;35(3):2145–57.
- 42. Chen R, Li M, Zhao Y. Hybrid spatial–frequency representation learning for robust deepfake detection. IEEE Trans Multimedia. 2025;27:3890–903.
- 43.
Stamnas S, Sanchez V. Difffake: Exposing deepfakes using differential anomaly detection. In: Proceedings of the Winter Conference on Applications of Computer Vision. 2025. p. 695–705.
- 44. Shi Z, Liu W, Chen H. Face Reconstruction-Based Generalized Deepfake Detection Model with Residual Outlook Attention. ACM Trans Multimedia Comput Commun Appl. 2025;21(4):1–19.
- 45.
Khalid H, Woo SS. Oc-fakedect: Classifying deepfakes using one-class variational autoencoder. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops. 2020. p. 656–7.
- 46.
Cao Y, Tong Y, Bao H, Liao X, Zhu M. OC-SAN: Unsupervised Deepfake Detection for Specific Individual Protection Based on Deep One-Class Classification. In: Chinese Conference on Pattern Recognition and Computer Vision (PRCV). Springer Nature Singapore; 2024. p. 327–41.
- 47.
Khodabakhsh A, Busch C. A generalizable deepfake detector based on neural conditional distribution modelling. In: 2020 International Conference of the Biometrics Special Interest Group (BIOSIG). IEEE; 2020. p. 1–5.
- 48. Maiano L, Casadei F, Amerini I. Enhancing abnormality identification: Robust out-of-distribution strategies for deepfake detection. Forensic Sci Int: Digit Investig. 2026;56:302062.
- 49. Lin K, Lin Y, Li W, Yao T, Li B. Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection. AAAI. 2025;39(5):5262–70.