Don't Fall for the No-Ops Myth: 5 Production-Level Governance Decisions for Serverless Cold Starts, Cost Explosions, and Observability Blind Spots

Overview At 2 AM, I got a phone call. Users of a ride-hailing project reported that opening the App to view trip history during late-night hours caused a 3-5 second blank screen. After investigation, I found that this API group used Serverless function compute. During the day, traffic was normal, but after traffic dropped off at night, function instances were reclaimed. The first request triggered a cold start, pushing P99 latency to 4....

August 22, 2026 · 22 mins · 4662 words · Xu Baojin

Redis Operations: Persistence, High Availability, and Performance Monitoring

Overview Redis Operations: Persistence, High Availability, and Performance Monitoring is an essential skill in SRE operations. In production environments, mastering these techniques can significantly improve system stability and operational efficiency. Why Redis Operations Matters As systems grow in scale and complexity, traditional operations approaches struggle to meet the demands of modern distributed systems. Redis Operations helps operations teams: Rapid Problem Resolution: Systematic tools and methods reduce troubleshooting time Improved System Visibility: Establish comprehensive monitoring and observability Proactive Fault Prevention: Identify and fix potential risks before they cause outages Resource Optimization: Allocate and schedule resources efficiently Core Concepts and Principles Basic Concepts The core of Redis Operations lies in establishing standardized processes and automated toolchains....

September 15, 2025 · 3 mins · 565 words · XuBaojin