Overview
This research introduces TwinViT-DeepJSCC, a system designed for adversarially robust semantic image communication. The system functions as a preventive-corrective semantic image transceiver, operating within a fixed channel-use budget. Its primary objective is to address vulnerabilities in learning-based semantic communication arising from adversarial perturbations introduced either before semantic encoding or across wireless channels.
Research Context
Learning-based semantic communication systems face susceptibility to adversarial perturbations. These perturbations can originate at the source domain, prior to semantic encoding, or within the wireless channel itself. Protecting the integrity and accuracy of transmitted semantic information under such conditions is a key challenge.
Approach
TwinViT-DeepJSCC incorporates two Vision Transformer (ViT)-based deep joint source-channel coding (DeepJSCC) branches. These branches are engineered to learn complementary latent representations. Protection for these representations is achieved through sensitivity-aware masking.
At the receiver side, the system employs a multi-faceted mitigation strategy. This includes confidence-aware fusion, blind corruption-severity estimation, and signal-to-noise ratio (SNR)-severity-conditioned denoising diffusion implicit model (DDIM) purification. These mechanisms collectively aim to mitigate residual corruption without requiring specific metadata about the attack.
Findings
Experiments were conducted on the Canadian Institute for Advanced Research 100-class (CIFAR-100) dataset. The evaluation considered various source-domain attacks, including fast gradient sign method (FGSM), projected gradient descent (PGD), natural evolution strategies (NES), and Carlini-Wagner (CW). Channel conditions included additive white Gaussian noise (AWGN) and block-flat Rayleigh fading, with additional consideration for random jamming and channel-aware adversarial waveforms.
Under matched channel-use and attack budgets, TwinViT-DeepJSCC demonstrated substantial performance improvements:
- It achieved maximum peak signal-to-noise ratio (PSNR) gains of approximately 9.5 dB when subjected to 20-step PGD attacks.
- Under channel-aware waveform attacks over block-flat Rayleigh fading, the system recorded maximum PSNR gains of approximately 10.8 dB.
- In the presence of PGD attacks, TwinViT-DeepJSCC improved Top-1 accuracy by up to approximately 38 percentage points when compared to an undefended baseline.
- Against the strongest competing defense, TwinViT-DeepJSCC exhibited an improvement in Top-1 accuracy of up to approximately 13 percentage points under PGD attacks.
Ablation results provided further evidence, confirming the complementary contributions of both the transmitter-side and receiver-side mechanisms proposed within the TwinViT-DeepJSCC framework.
Why This Matters
The system addresses the vulnerability of learning-based semantic communication to adversarial perturbations in both source and channel domains, enhancing robustness.
Potential Applications
While the source does not explicitly discuss potential applications, the context of semantic image communication and its robustness to adversarial attacks implies utility in scenarios where secure and reliable image transmission is critical, particularly in the presence of malicious interference or challenging channel conditions.