Drag to rotate the volume. Click a patch in the frame view to inspect attention links.
Clean spacetime signal Noise-dominated patch Selected / attended
↻ drag to rotate · zoom slider to scale
Frame t — full 2D token grid
Attention cost vs. patch size ps

Spacetime Patch Tokenizer 2D: How Sora Turns Video Into Transformer Tokens

Before a video Diffusion Transformer can denoise a clip, the clip has to become a sequence of tokens — and Sora's key architectural idea is to patchify a compressed video latent in space and time at once, producing "spacetime patches" instead of per-frame image patches. This 2D companion renders the same latent volume with an independently-computed isometric projection you rotate by dragging, a full top-down 2D token grid for any single time-group, and a live bar chart of the N² self-attention cost as patch size shrinks — change the latent resolution, clip length, and spatial/temporal patch size to watch every panel react together, blend patches toward Gaussian noise to see the tensor a DiT actually denoises, and click any patch to compare full spacetime self-attention against a cheaper per-frame windowed variant.