HomeArticlesPerception

Gestalt Grouping: How the Eye Assembles a Scene From Fragments

The visual system doesn't perceive a scene as a pile of independent dots and edges — it actively organizes them into structured wholes before you're even aware it happened.

mysimulator teamUpdated June 2026≈ 7 min read▶ Open the simulation

Wertheimer and the whole that isn't the sum of its parts

In 1912, psychologist Max Wertheimer, building on his earlier work on the phi phenomenon — the illusion of motion produced by two lights flashing in quick succession — founded what became known as the Berlin school of Gestalt psychology together with Wolfgang Köhler and Kurt Koffka. Their central claim, usually rendered in English as "the whole is different from the sum of its parts," was a direct challenge to the then-dominant idea that perception is assembled piece by piece from independent sensory elements. Wertheimer argued instead that the visual system actively organizes raw input into structured wholes before we ever become consciously aware of the individual pieces: grouping isn't a later interpretive step layered on top of seeing, it is built into the initial act of seeing itself. Over the following decade the school formalized a small set of grouping laws that predict, with striking reliability, which elements in a scene will be read as belonging together.

Proximity and similarity

The two simplest laws work on spatial arrangement and surface properties respectively. Proximity states that elements positioned close together are perceived as one group — a grid of evenly spaced dots stretched slightly wider in one direction reads as columns rather than rows, purely because of the smaller gap between neighbours. Similarity groups elements that share a visual property such as colour, shape, size or orientation, even when they are scattered and not adjacent at all — a handful of red dots mixed into a field of blue ones pop out as a set no matter where they land on the page. When the two cues are put in direct conflict, whichever one is visually stronger tends to win: strong colour contrast can override weak spacing differences, and tight spacing can override weak similarity.

Proximity      elements close together are read as one group
Similarity     elements sharing colour, shape, size or orientation group together
Closure        the visual system completes incomplete contours into whole shapes
Continuity     smooth continuous paths are preferred over abrupt turns
Common fate    elements moving together are grouped as one unit
Figure-ground  the scene splits into a foreground shape and a background

Closure and continuity

Closure describes the visual system's tendency to complete an incomplete figure into a whole, familiar shape rather than perceiving disconnected fragments. The canonical demonstration is the Kanizsa triangle (Gaetano Kanizsa, 1955): three pac-man-like disks with wedges cut out, and three right-angle brackets, are arranged so their gaps line up, and observers reliably perceive a solid white triangle sitting on top of them — complete with edges, or illusory contours, that are not physically present in the image at all. This is not merely a cognitive guess dressed up as perception: neurons in visual area V2 have been found that respond to these illusory contours as though they were real edges, which means the brightness and edge percept is constructed surprisingly early in the visual pathway, before the signal ever reaches conscious awareness. Continuity, or "good continuation," is a related law: when two lines cross, the visual system prefers to interpret the image as two smooth continuous lines passing through each other rather than as four separate segments that happen to meet at a single point, because a smooth path is a simpler and more probable explanation of the image than four disconnected fragments coincidentally touching.

live demo · dots organizing into groups by proximity and similarity● LIVE

Common fate and figure-ground

Common fate groups elements by shared motion: dots moving in the same direction at the same speed are perceived as a unit even when they share no proximity or similarity with each other. This is why an animal breaking camouflage the instant it moves becomes so conspicuous, and why biological motion — a handful of light points attached to a walking person's joints — is instantly read as a walking figure the moment it starts moving, even though the same points look like a random scatter when frozen. Figure-ground organization works at the level of the whole scene rather than clusters of elements: the visual system assigns one side of a contour a shape as the "figure," and treats the other side as a shapeless "background" that appears to continue behind it. Edgar Rubin's 1915 vase/faces figure is the classic demonstration — the identical contour can be read either as a white vase against two black profiles, or as two black faces against a white vase, and perception flips between the two readings but never holds both at the same time. None of this is an academic curiosity: proximity, similarity and closure are the working vocabulary behind UI layout, dashboard design, iconography and camouflage, precisely because a designer who understands them can predict what a viewer will group together without writing a single word of instruction.

Frequently asked questions

What is the Kanizsa triangle and why do we see a triangle that isn't there?

The Kanizsa triangle is an arrangement of three pac-man-shaped disks and three angle brackets whose gaps line up so the visual system completes them into a solid white triangle, complete with edges that are not physically drawn. This is the Gestalt law of closure at work, and neurons in visual area V2 have been found that respond to these illusory contours as though they were real edges.

Can proximity and similarity conflict, and which one wins?

Yes — a grid of dots can be spaced to suggest columns while being coloured to suggest rows, and the two groupings compete. Whichever cue is visually stronger tends to dominate: strong colour or shape contrast will override weak spacing differences, and tight spacing will override weak similarity, so the outcome depends on the relative strength of each cue rather than a fixed rule.

Why does Rubin's vase flip between a vase and two faces?

The image contains one ambiguous contour that can be assigned to either side as the shaped "figure," with the other side read as shapeless background. Because the visual system commits to only one figure-ground assignment at a time, perception flips between the vase reading and the two-faces reading but never holds both interpretations simultaneously.

Try it live

Everything above runs in your browser — open Gestalt Grouping and rearrange the dots yourself to see proximity and similarity compete in real time. Nothing is installed, nothing is uploaded, the whole model lives in one tab.

▶ Open Gestalt Grouping simulation

What did you find?

Add reproduction steps (optional)