chapter eight
8 Root cause thinking for developers
This chapter covers
- Breaking the symptom-fix cycle that creates endless rework
- Applying 5 Whys, Fishbone, Pareto, and Is/Is Not to expose root causes
- Tracing from stack traces to systemic failures, not code-level guesses
- Catching AI-generated patches before they become technical debt
- Writing A3 reports that communicate problems without drama
Your payment service throws 100+ errors every few hours. The pattern looks random. You restart the service, the errors stop, and you move on. Two days later, the same alert fires at 3 a.m. You restart again. A week later, it happens a third time. On the third restart, you stop and actually trace the failure. The service crashes on large transactions because the amount_cents column is a 32-bit INTEGER. Any transaction over $21,474,836.47 overflows the column and throws a SQLException. Three restarts fixed nothing because the column type was the problem the entire time.