Each task has a real target weight vector and a Fisher-information importance per axis. Training does actual gradient descent on the current task's quadratic loss, plus — when EWC is on — a real elastic penalty λ·F·(w − w*) pulling weights back toward every earlier task's optimum. The ellipsoids show how tightly each finished task is protected.
Drag to orbit · Scroll to zoom · Train tasks in order and compare EWC on vs off