The 3 AM Phone Rings Again: 7 Engineering Decisions for On-Call Rotation Design and Fatigue Management

Overview 3 AM. Your phone buzzes. P1 alert — order service 5xx error rate spiking. You drag yourself out of bed, fire up the laptop, pull logs, check Grafana, trace it to a saturated upstream database connection pool. Restart, scale up, recover. It’s 4:30 AM. The next morning, nursing dark circles in the standup, someone mentions that last night’s alert was actually a false positive. If you’ve worked in operations, you know this feeling....

August 25, 2026 · 24 mins · 4917 words · Xu Baojin

Cutting 80% of Annual Downtime: 6 Engineering Decisions and Hard-Learned Lessons Going from 99.5% to 99.9% Availability

Overview 99.5% availability sounds pretty good — until you do the math. A year has 8,760 hours. 99.5% means you’re allowed 43.8 hours of downtime. Per month, that’s roughly 3.65 hours of service interruption. If this is a core transaction system, those 3.65 hours could mean tens of thousands in lost orders. During a major promotion, the loss multiplies by orders of magnitude. 99.9%? Annual downtime drops to 8.77 hours, less than 44 minutes per month....

August 20, 2026 · 22 mins · 4628 words · Xu Baojin

3 Days of Approval, Still Crashed: Replacing Manual CAB with Error Budget Gates Cut Change-Related Incidents by 60%

Overview 1:47 AM. Phone buzzes. The alert group explodes with 200+ messages. Order success rate drops from 99.8% to 92%. The on-call engineer restarts the service, rolls back the version, and fights for 40 minutes to restore service. Next day’s postmortem: root cause was a one-line database connection pool config change pushed during the afternoon release. This change went through the full approval process — developer submission, test verification, manager sign-off, SRE review....

August 6, 2026 · 23 mins · 4818 words · Xu Baojin

Reliability Measurement Is Not About Piling Metrics: Turning 'Is the System Stable?' into Decision-Driven Numbers with a Four-Layer Model

Overview 3 AM. You get woken up by an alert. You check: CPU normal, memory normal, pods running. But users are complaining the system is unusable. You open Grafana, stare at a wall of charts, and cannot answer one simple question: is this an incident or not? I have encountered this scenario repeatedly across multiple enterprise SRE consulting engagements. The core problem is not insufficient monitoring — it is that the measurement system is designed wrong: we keep monitoring internal system metrics instead of user-perceived service quality....

July 30, 2026 · 24 mins · 5062 words · Xu Baojin

Microservice Rate Limiting, Circuit Breaking, and Degradation: A Practical Guide from Avalanche Prevention to Fine-Grained Traffic Governance

Overview 3 AM, phone buzzing wildly. Open it — the alert group is already on fire. Payment service P99 latency jumped from 50ms to 12 seconds, error rate past 80%. Worse, the failure cascaded like dominoes: user service timeout → auth failures → inventory service retry storm → message queue backed up → payment callbacks all timed out. The monitoring dashboard went entirely red in thirty seconds. This isn’t a unique disaster....

July 24, 2026 · 26 mins · 5536 words · Xu Baojin

Postmortem Action Items Tracking: Engineering a Closed-Loop Process from Checklist to Completion

Overview Anyone who has done production operations has probably lived through this scenario: you get woken up by an alert at 3 AM, fight the fire until service is restored, hold a postmortem the next day, list a dozen improvement actions in the meeting, send the minutes to the group chat, and everyone says “got it.” Then what? A month later you dig it up and find maybe two or three items actually got done....

July 21, 2026 · 24 mins · 5022 words · Xu Baojin

Building an SRE Team: From Hiring to Organizational Capability Model

Overview 3 AM. The core trading system is down. The on-call engineer frantically flips through the Runbook. The DBA says it’s not a database issue. The network team says the links are fine. The developers say nothing changed. Three teams point fingers at each other. Incident recovery drags on for 47 minutes. This is the reality of operations in many companies. The problem isn’t that people aren’t trying hard enough. The problem is there’s no engineering-driven reliability team to decompose the problem....

July 13, 2026 · 14 mins · 2883 words · Xu Baojin

Data-Driven SRE Decisions: From Metrics to Actions

Overview Data-Driven SRE Decisions: From Metrics to Actions is an essential skill in SRE operations. In production environments, mastering these techniques can significantly improve system stability and operational efficiency. Why Data-Driven SRE Decisions Matters As systems grow in scale and complexity, traditional operations approaches struggle to meet the demands of modern distributed systems. Data-Driven SRE Decisions helps operations teams: Rapid Problem Resolution: Systematic tools and methods reduce troubleshooting time Improved System Visibility: Establish comprehensive monitoring and observability Proactive Fault Prevention: Identify and fix potential risks before they cause outages Resource Optimization: Allocate and schedule resources efficiently Core Concepts and Principles Basic Concepts The core of Data-Driven SRE Decisions lies in establishing standardized processes and automated toolchains....

January 2, 2026 · 3 mins · 569 words · XuBaojin

SRE Perspective on FinOps: Cloud Cost Visibility and Optimization Strategies

Overview SRE Perspective on FinOps: Cloud Cost Visibility and Optimization Strategies is an essential skill in SRE operations. In production environments, mastering these techniques can significantly improve system stability and operational efficiency. Why SRE Perspective on FinOps Matters As systems grow in scale and complexity, traditional operations approaches struggle to meet the demands of modern distributed systems. SRE Perspective on FinOps helps operations teams: Rapid Problem Resolution: Systematic tools and methods reduce troubleshooting time Improved System Visibility: Establish comprehensive monitoring and observability Proactive Fault Prevention: Identify and fix potential risks before they cause outages Resource Optimization: Allocate and schedule resources efficiently Core Concepts and Principles Basic Concepts The core of SRE Perspective on FinOps lies in establishing standardized processes and automated toolchains....

December 27, 2025 · 3 mins · 577 words · XuBaojin

Platform Engineering Introduction: Designing and Implementing Internal Developer Platforms

Overview Platform Engineering Introduction: Designing and Implementing Internal Developer Platforms is an essential skill in SRE operations. In production environments, mastering these techniques can significantly improve system stability and operational efficiency. Why Platform Engineering Introduction Matters As systems grow in scale and complexity, traditional operations approaches struggle to meet the demands of modern distributed systems. Platform Engineering Introduction helps operations teams: Rapid Problem Resolution: Systematic tools and methods reduce troubleshooting time Improved System Visibility: Establish comprehensive monitoring and observability Proactive Fault Prevention: Identify and fix potential risks before they cause outages Resource Optimization: Allocate and schedule resources efficiently Core Concepts and Principles Basic Concepts The core of Platform Engineering Introduction lies in establishing standardized processes and automated toolchains....

December 22, 2025 · 3 mins · 571 words · XuBaojin