Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

[Submitted on 7 Jan 2025 (v1), last revised 18 Feb 2026 (this version, v4)]

View a PDF of the paper titled PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models, by Lingzhi Yuan and 9 other authors

View PDF
HTML (experimental)

Abstract:Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, these models are vulnerable to misuse, particularly generating not-safe-for-work (NSFW) content, such as sexually explicit, violent, political, and disturbing images, raising serious ethical concerns. In this work, we present PromptGuard, a novel content moderation technique that draws inspiration from the system prompt mechanism in large language models (LLMs) for safety alignment. Unlike LLMs, T2I models lack a direct interface for enforcing behavioral guidelines. Our key idea is to optimize a safety soft prompt that functions as an implicit system prompt within the T2I model’s textual embedding space. This universal soft prompt (P*) directly moderates NSFW inputs, enabling safe yet realistic image generation without altering the inference efficiency or requiring proxy models. We further enhance its reliability and helpfulness through a divide-and-conquer strategy, which optimizes category-specific soft prompts and combines them into holistic safety guidance. Extensive experiments across five datasets demonstrate that PromptGuard effectively mitigates NSFW content generation while preserving high-quality benign outputs. PromptGuard achieves 3.8 times faster than prior content moderation methods, surpassing eight state-of-the-art defenses with an optimal unsafe ratio down to 5.84%.

Submission history

From: Xinfeng Li [view email]
[v1]
Tue, 7 Jan 2025 05:39:21 UTC (21,071 KB)
[v2]
Fri, 4 Apr 2025 05:56:04 UTC (21,064 KB)
[v3]
Fri, 5 Sep 2025 04:44:48 UTC (2,376 KB)
[v4]
Wed, 18 Feb 2026 05:55:40 UTC (2,376 KB)

What's Hot

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

Manifold-Matching Autoencoders

One Model to Rule Them All? SAP-RPT-1 and the Future of Tabular Foundation Models

Bridging Facts for Cross-Document Reasoning at Index Time

SpecMoE: Spectral Mixture-of-Experts Foundation Model for Cross-Species EEG Decoding

How a Neural Network Learned Its Own Fraud Rules: A Neuro-Symbolic AI Experiment

Bridging Modality Gap with Temporal Evolution Semantic Space

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

70+ AI art styles to use in your AI prompts

Manifold-Matching Autoencoders

One Model to Rule Them All? SAP-RPT-1 and the Future of Tabular Foundation Models

Why customer personas help you win earlier in AI search

Google expands Personal Intelligence to AI Mode, Gemini, Chrome

Google AI Overviews Cut Germany’s Top Organic CTR By 59%

Most Popular

13 Trending Songs on TikTok in Nov 2025 (+ How to Use Them)

How to watch the 2026 GRAMMY Awards online from anywhere

Corporate Reputation Management Strategies | Sprout Social

Our Picks

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Subscribe to Updates

What's Hot

Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

Submission history

Related Posts

Subscribe to Updates