← The frontier
Technology & AIJul 22, 2026

Test-Time Training for Modality Order Consistency in Vision-Language Models

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented.

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.