← The frontier
Technology & AIAug 2, 2026

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR).

We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.