Mode: Undetermined
drag to pan · scroll to zoom
Undetermined activation
k-means convergence & distance-to-centroid histogram

Activation Clustering: Detecting Neural Network Backdoors (2D)

This simulator plots every sample a poisoned classifier assigns to one output class as a point on a 2D activation-space grid. A hidden trigger pushes a poisoned minority of that class into a tight, separated cluster, exactly as it does in a real backdoored network. Adjust the poison rate, trigger separation and clean-activation spread, then run the real k-means Activation Clustering defense (Chen et al., 2018) to flag the suspected cluster, watch its centroids converge iteration by iteration in the panel below, or reveal the planted ground truth to check how the defense performed — with live precision, recall and a silhouette-style separation score. Drag to pan and scroll to zoom the scatter plot.