Tumor: 5 · Toxicity: 0 Action: —
⚠ Couldn't load the 3D engineThree.js failed to load from the CDN. Check your connection and reload.

Reinforcement-Learning Chemotherapy Dosing Agent

This simulator trains a tabular Q-learning agent to choose chemotherapy doses for a simulated patient, purely from trial-and-error reward signal — no hand-coded treatment rules. A 3D landscape shows every (tumor burden, toxicity) state as a column whose height is the agent's learned value estimate and whose colour is its currently preferred dose; a marble marks the agent's live position and hops between states as training plays out in real time. Tune the learning rate, exploration decay and toxicity aversion to watch the learned policy shift between aggressive and conservative dosing, with live episode count, exploration rate, moving-average reward and cure rate.