Starting from Docker Machine Deprecation: 7 Production Decisions for GitLab CI Runner Elastic Architecture and Cache Governance

Overview 1 AM. You get an alert: GitLab CI pipeline queue has 47 pending jobs. The dev chat explodes—“code pushed 40 minutes ago, still not running"“is the Runner down?““just add more machines!” You check the console: all 3 Runners are at capacity, each running 10 concurrent jobs, 30 slots fully occupied. Add machines? Docker Machine executor needs 3 minutes to spin up a new EC2, then another 2 minutes to become Ready....

August 13, 2026 · 17 mins · 3581 words · Xu Baojin

Redis Operations: Persistence, High Availability, and Performance Monitoring

Overview Redis Operations: Persistence, High Availability, and Performance Monitoring is an essential skill in SRE operations. In production environments, mastering these techniques can significantly improve system stability and operational efficiency. Why Redis Operations Matters As systems grow in scale and complexity, traditional operations approaches struggle to meet the demands of modern distributed systems. Redis Operations helps operations teams: Rapid Problem Resolution: Systematic tools and methods reduce troubleshooting time Improved System Visibility: Establish comprehensive monitoring and observability Proactive Fault Prevention: Identify and fix potential risks before they cause outages Resource Optimization: Allocate and schedule resources efficiently Core Concepts and Principles Basic Concepts The core of Redis Operations lies in establishing standardized processes and automated toolchains....

September 15, 2025 · 3 mins · 565 words · XuBaojin