Present object (dot size = P(next)) Absent object — a phantom risk Decoding score field w_v·e + w_l·p

VLM Object Hallucination Lab — Evidence/Prior Phase Plane (2D)

Vision-language models like GPT-4V, Gemini Vision, LLaVA and Florence-2 caption images by blending two signals at every generated word: real visual evidence from the image encoder, and a language prior learned from co-occurrence statistics in training captions. Rather than rendering the scene the model is looking at, this 2D build plots the decoder's own decision space directly — an evidence-vs-language-prior phase plane, with a live decoding-score field in the background and every candidate object drawn as a dot sized by its real softmax probability of being the next word. Push the language-prior weight or the occlusion/noise slider up and watch phantom objects (dots on absent items) swell and get named anyway — a hallucinated mention, tracked live by the same CHAIRi metric researchers use to benchmark real VLMs.