Most firmware tutorials start with “blink an LED.” Production systems should start with “what happens when power is ugly, a sensor is missing, or the bus goes quiet?”
Fail-safe design is not a single feature. It is a habit that shows up in reset vectors, peripheral init order, watchdog strategy, and how you represent invalid state.
Start from the safe state
Before you write a control loop, define the de-energized configuration of every actuator. Outputs should default passive in hardware where possible (pull-downs, disabled gate drivers) and again in software before any application task runs.
void board_init_safe(void) {
actuators_all_off();
fault_latches_clear();
watchdog_init();
}Separate detection from reaction
Sensor glitches are normal. Debounce and classify faults in one module; transition to safe state in another. That split keeps control code readable and makes unit tests feasible.
Heartbeats are contracts
If a PLC stops talking, your module should not hold the last command forever unless that is an explicit, documented requirement. A missing heartbeat is information — treat it like a fault class with a clear timeout and recovery path.
Measure what you claim
Safe defaults are only as good as the paths you test: brownout, flash corruption, stuck bus, and “user updates config mid-run.” Automated hardware-in-the-loop tests for these cases are boring — and invaluable.
Fail-safe thinking will slow the first prototype. It will save the product.