Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
arXiv:2609.00067v2 Announce Type: replace-cross Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic thatโฆ