Microservice Rate Limiting, Circuit Breaking, and Degradation: A Practical Guide from Avalanche Prevention to Fine-Grained Traffic Governance

Overview 3 AM, phone buzzing wildly. Open it — the alert group is already on fire. Payment service P99 latency jumped from 50ms to 12 seconds, error rate past 80%. Worse, the failure cascaded like dominoes: user service timeout → auth failures → inventory service retry storm → message queue backed up → payment callbacks all timed out. The monitoring dashboard went entirely red in thirty seconds. This isn’t a unique disaster....

July 24, 2026 · 26 mins · 5536 words · Xu Baojin

System Resilience Engineering: From Reactive Firefighting to Proactive Defense

Overview 2 AM. Your phone buzzes. The payment service is timing out, thread pools are exhausted, upstream order services are queuing up, and ten minutes later the entire transaction pipeline collapses. You check the logs: a downstream cache cluster hiccuped for 3 seconds. During those 3 seconds, upstream services retried frantically, maxing out connection limits, and everything sharing that connection pool went down together. This story isn’t new. Almost anyone who’s done operations for a few years has lived through a similar cascading failure....

July 23, 2026 · 20 mins · 4219 words · Xu Baojin

Service Dependency Mapping and Failure Domain Analysis: From Topology Discovery to Blast Radius Control

Overview In modern microservice architectures, a seemingly simple user request may traverse dozens of service nodes. When an incident occurs, the first question an SRE engineer faces is often not “how to fix it” but “what is the scope of impact.” Without a fast answer to this question, incident recovery gets bogged down in endless investigation. Service Dependency Maps and Failure Domain Analysis are the engineering methodologies that address this problem....

December 16, 2024 · 25 mins · 5210 words · XuBaojin