The 3 AM Phone Rings Again: 7 Engineering Decisions for On-Call Rotation Design and Fatigue Management

Overview 3 AM. Your phone buzzes. P1 alert — order service 5xx error rate spiking. You drag yourself out of bed, fire up the laptop, pull logs, check Grafana, trace it to a saturated upstream database connection pool. Restart, scale up, recover. It’s 4:30 AM. The next morning, nursing dark circles in the standup, someone mentions that last night’s alert was actually a false positive. If you’ve worked in operations, you know this feeling....

August 25, 2026 · 24 mins · 4917 words · Xu Baojin

MTTR from 40 Minutes to 8: Cutting Fault Localization Time by 5x with a Four-Layer Defense

Overview 3 AM, phone vibrating. The alert channel exploded with 200 messages, and user complaints are already flooding customer service. You scramble to your laptop, log into the bastion host, and discover a core API endpoint’s P99 latency spiked to 8 seconds. For the next 30 minutes, you dig through logs, check Grafana dashboards, and ping upstream and downstream teams—only to find that a config center pushed an incorrect parameter....

July 28, 2026 · 22 mins · 4552 words · Xu Baojin

Incident Management and On-Call Mechanism Design

Overview There’s a saying in SRE: “Systems will fail — the question is whether you’ll be woken up by them or actively managing them.” Incident management is not “dealing with things after they break” — it’s a complete engineering system spanning prevention, detection, response, and learning. This article systematically covers how to build a practical on-call system across five dimensions: incident grading, on-call rotation, incident response processes, postmortem culture, and alert governance....

August 23, 2024 · 13 mins · 2559 words · XuBaojin