Cutting 80% of Annual Downtime: 6 Engineering Decisions and Hard-Learned Lessons Going from 99.5% to 99.9% Availability

Overview 99.5% availability sounds pretty good — until you do the math. A year has 8,760 hours. 99.5% means you’re allowed 43.8 hours of downtime. Per month, that’s roughly 3.65 hours of service interruption. If this is a core transaction system, those 3.65 hours could mean tens of thousands in lost orders. During a major promotion, the loss multiplies by orders of magnitude. 99.9%? Annual downtime drops to 8.77 hours, less than 44 minutes per month....

August 20, 2026 · 22 mins · 4628 words · Xu Baojin

Reliability Measurement Is Not About Piling Metrics: Turning 'Is the System Stable?' into Decision-Driven Numbers with a Four-Layer Model

Overview 3 AM. You get woken up by an alert. You check: CPU normal, memory normal, pods running. But users are complaining the system is unusable. You open Grafana, stare at a wall of charts, and cannot answer one simple question: is this an incident or not? I have encountered this scenario repeatedly across multiple enterprise SRE consulting engagements. The core problem is not insufficient monitoring — it is that the measurement system is designed wrong: we keep monitoring internal system metrics instead of user-perceived service quality....

July 30, 2026 · 24 mins · 5062 words · Xu Baojin

Alerting Strategy Design: From Noise to Signal

Overview Alerting is the “last mile” of a monitoring system — and the hardest to get right. A common predicament: servers run dozens of alerting rules generating hundreds of alert notifications daily. On-call engineers, bombarded by WeChat/DingTalk/email, gradually become desensitized — truly urgent alerts are drowned in noise until customer complaints reveal the system has been broken for hours. The golden rule of SRE: every alert must have a clear action....

August 19, 2024 · 14 mins · 2928 words · XuBaojin

Error Budget Consumption Strategies and Action Guidelines

Overview The Error Budget is the most ingeniously designed mechanism in the SRE framework. It transforms the long-standing “stability vs. iteration speed” debate — previously settled by opinion and politics — into a quantifiable engineering decision framework: your system has an “unavailability allowance,” and when it’s spent, you stop and fix things. In practice, however, many teams define SLOs and error budgets but stop at displaying a percentage number on a dashboard....

April 26, 2024 · 17 mins · 3556 words · XuBaojin