Masked Image Modeling for Mutual Information-based Adversarial Robustness

[Submitted on 8 Dec 2023 (v1), last revised 15 Dec 2025 (this version, v5)]

View a PDF of the paper titled MIMIR: Masked Image Modeling for Mutual Information-based Adversarial Robustness, by Xiaoyun Xu and 3 other authors

View PDF
HTML (experimental)

Abstract:Vision Transformers (ViTs) have emerged as a fundamental architecture and serve as the backbone of modern vision-language models. Despite their impressive performance, ViTs exhibit notable vulnerability to evasion attacks, necessitating the development of specialized Adversarial Training (AT) strategies tailored to their unique architecture. While a direct solution might involve applying existing AT methods to ViTs, our analysis reveals significant incompatibilities, particularly with state-of-the-art (SOTA) approaches such as Generalist (CVPR 2023) and DBAT (USENIX Security 2024). This paper presents a systematic investigation of adversarial robustness in ViTs and provides a novel theoretical Mutual Information (MI) analysis in its autoencoder-based self-supervised pre-training. Specifically, we show that MI between the adversarial example and its latent representation in ViT-based autoencoders should be constrained via derived MI bounds. Building on this insight, we propose a self-supervised AT method, MIMIR, that employs an MI penalty to facilitate adversarial pre-training by masked image modeling with autoencoders. Extensive experiments on CIFAR-10, Tiny-ImageNet, and ImageNet-1K show that MIMIR can consistently provide improved natural and robust accuracy, where MIMIR outperforms SOTA AT results on ImageNet-1K. Notably, MIMIR demonstrates superior robustness against unforeseen attacks and common corruption data and can also withstand adaptive attacks where the adversary possesses full knowledge of the defense mechanism. Our code and trained models are publicly available at: this https URL.

Submission history

From: Xiaoyun Xu [view email]
[v1]
Fri, 8 Dec 2023 10:50:02 UTC (3,893 KB)
[v2]
Wed, 17 Jan 2024 13:47:32 UTC (3,894 KB)
[v3]
Fri, 16 Aug 2024 12:31:38 UTC (2,579 KB)
[v4]
Tue, 15 Apr 2025 10:50:18 UTC (3,839 KB)
[v5]
Mon, 15 Dec 2025 22:24:07 UTC (4,841 KB)

What's Hot

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Masked Image Modeling for Mutual Information-based Adversarial Robustness

Bridging Modality Gap with Temporal Evolution Semantic Space

How to Effectively Review Claude Code Output

Everything You Need to Know About Recursive Language Models

[2601.15871] Why Inference in Large Models Becomes Decomposable After Training

Self-Hosting Your First LLM | Towards Data Science

To See is Not to Master: Teaching LLMs to Use Private Libraries for Code Generation

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Bridging Modality Gap with Temporal Evolution Semantic Space

How to Effectively Review Claude Code Output

Google adds video visibility to Performance Max reporting

Everything You Need to Know About Recursive Language Models

[2601.15871] Why Inference in Large Models Becomes Decomposable After Training

Top Blog Platforms for SEO: Which Sites to Conside

Most Popular

13 Trending Songs on TikTok in Nov 2025 (+ How to Use Them)

How to watch the 2026 GRAMMY Awards online from anywhere

Corporate Reputation Management Strategies | Sprout Social

Our Picks

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Subscribe to Updates

What's Hot

Masked Image Modeling for Mutual Information-based Adversarial Robustness

Submission history

Related Posts

Subscribe to Updates