Cutting 80% of Annual Downtime: 6 Engineering Decisions and Hard-Learned Lessons Going from 99.5% to 99.9% Availability

Overview 99.5% availability sounds pretty good — until you do the math. A year has 8,760 hours. 99.5% means you’re allowed 43.8 hours of downtime. Per month, that’s roughly 3.65 hours of service interruption. If this is a core transaction system, those 3.65 hours could mean tens of thousands in lost orders. During a major promotion, the loss multiplies by orders of magnitude. 99.9%? Annual downtime drops to 8.77 hours, less than 44 minutes per month....

August 20, 2026 · 22 mins · 4628 words · Xu Baojin

3 Days of Approval, Still Crashed: Replacing Manual CAB with Error Budget Gates Cut Change-Related Incidents by 60%

Overview 1:47 AM. Phone buzzes. The alert group explodes with 200+ messages. Order success rate drops from 99.8% to 92%. The on-call engineer restarts the service, rolls back the version, and fights for 40 minutes to restore service. Next day’s postmortem: root cause was a one-line database connection pool config change pushed during the afternoon release. This change went through the full approval process — developer submission, test verification, manager sign-off, SRE review....

August 6, 2026 · 23 mins · 4818 words · Xu Baojin

Reliability Measurement Is Not About Piling Metrics: Turning 'Is the System Stable?' into Decision-Driven Numbers with a Four-Layer Model

Overview 3 AM. You get woken up by an alert. You check: CPU normal, memory normal, pods running. But users are complaining the system is unusable. You open Grafana, stare at a wall of charts, and cannot answer one simple question: is this an incident or not? I have encountered this scenario repeatedly across multiple enterprise SRE consulting engagements. The core problem is not insufficient monitoring — it is that the measurement system is designed wrong: we keep monitoring internal system metrics instead of user-perceived service quality....

July 30, 2026 · 24 mins · 5062 words · Xu Baojin

Error Budget Consumption Strategies and Action Guidelines

Overview The Error Budget is the most ingeniously designed mechanism in the SRE framework. It transforms the long-standing “stability vs. iteration speed” debate — previously settled by opinion and politics — into a quantifiable engineering decision framework: your system has an “unavailability allowance,” and when it’s spent, you stop and fix things. In practice, however, many teams define SLOs and error budgets but stop at displaying a percentage number on a dashboard....

April 26, 2024 · 17 mins · 3556 words · XuBaojin