drag to pan the token stream

Toolformer Self-Annotation Loss Filter (2D)

Toolformer-style training builds its tool-use dataset without a single human label: the model samples candidate API calls at points in ordinary text, executes them, and keeps only the ones that measurably reduce its own loss on the tokens that follow. This 2D simulator renders that filtering step as three linked panels — a pannable token stream with paired loss bars at each candidate site, a live histogram of the loss-reduction distribution against the threshold τ, and a running strip chart of the acceptance rate across batches. Tune the threshold, how many candidate calls are sampled per site, how spread out candidate usefulness is, and the fixed cost charged for actually calling a tool, and watch every panel respond live.