Edge AI does not always mean an all-or-nothing choice between running inference on-device or shipping the whole job to a server — a neural network can be sliced layer by layer, running its early stages locally and its later stages on a nearby edge server, with a single activation tensor crossing the network at the cut. This simulator models a real 8-layer network (choose a convolutional or transformer-style layer profile, each with its own per-layer compute cost and activation size) and lets you drag the split point across every possible layer boundary, watching device compute time, the network transfer time for that boundary's activation, and edge-server compute time recombine into a live total latency. Tune the device's compute power, the edge server's compute power and the link bandwidth, or hit auto-optimize to sweep every split point and jump straight to the fastest one — the same layer-partitioning trade-off used by real split-computing and collaborative-inference systems on mobile and IoT hardware.