11 Diagnostics
This chapter covers:
- Surveying the diagnostic signals
- Detecting memory leaks through a systematic flow
- Distinguishing garbage collection pressure from memory leaks
- Diagnosing CPU bottlenecks and event loop delays
- Introducing event loop utilization as a scaling
- Reasoning about performance problems when the diagnostic tools are unavailable
In one of my old consulting jobs, I was asked to come out and help the customer figure out why their application was performing so poorly. They were being plagued by terrible throughput, high latency, and memory issues that were driving up their infrastructure costs. In my initial consultation with them their lead architect walked me over to a giant monitor they had set up that displayed the real time metrics they were collecting. It was a screen full of graphs and meters and logs. After a very brief orientation he said, “If you can look at this board and tell me what problem we’re having, I’ll hire your team to help us out.”
It was a bit of a ridiculous request given that I had not yet looked at a single line of their code, but I humored him and stared at the board for a couple minutes. One graph stood out immediately. It was the memory heap usage. It would steadily climb, hit a max, hold for 30 seconds, then drop suddenly. Climb. Hold for 30 seconds. Drop. The pattern was clear.