A probe near Jupiter or beyond can be 40+ light-minutes from Earth β by the time a fault report arrives and a ground command is sent back, the spacecraft could already be dead. So flight software must handle faults itself, in three steps:
- Detect β a watchdog notices a sensor reading leave its expected envelope or a subsystem stop responding to heartbeat polls.
- Isolate β the faulty unit is switched off the power/data bus immediately, before its bad behaviour (a short, a stuck actuator, corrupted data) can damage or confuse neighboring subsystems.
- Recover β a redundant backup unit is switched in, or the spacecraft drops into a safe, degraded operating mode, and mission operations continue without waiting for a ground round-trip.
Without FDIR, an undetected fault can cascade: a power-bus short can starve the attitude-control system of voltage, which can send the spacecraft tumbling, which can point solar panels away from the sun and antennas away from Earth β a single component failure snowballing into total mission loss.