Fertility ≈ English (efficient) Moderate fertility gap Severe fertility gap
drag to pan · scroll to rescale

Tokenizer Fertility Gap (2D): The Multilingual Vocabulary Tax

Every multilingual language model tokenizes text with one shared subword vocabulary, trained once on a corpus dominated by a handful of high-resource languages. This 2D build renders that vocabulary as eight animated bar-chart columns of stacked token blocks with a live context-window gauge: pick a focus language, drag its share of the training corpus up or down, and watch how many token-blocks the same sentence needs — the "fertility" of the tokenizer for that language. Low-resource and morphologically rich languages routinely need several times as many tokens as English for identical meaning, which quietly taxes them with higher API cost per sentence and less usable context inside a fixed token window.