Don't Wait for the Crash to Check Monitoring: A Practical Methodology for Linux Performance Baselines and Anomaly Detection

Overview You’ve probably been here: woken up at 3 AM by a pager alert, scrambled to check Grafana, and found CPU usage spiked to 85%. You panic, investigate for half an hour, and finally realize — this database server runs a scheduled backup job every day from 2:50 to 3:10 AM. CPU is supposed to be at this level. What went wrong? You didn’t know what “normal” looks like for this machine....

August 30, 2026 · 23 mins · 4689 words · Xu Baojin

AIOps Anomaly Detection: From Static Thresholds to Intelligent Alerting

Overview AIOps Anomaly Detection: From Static Thresholds to Intelligent Alerting is an essential skill in SRE operations. In production environments, mastering these techniques can significantly improve system stability and operational efficiency. Why AIOps Anomaly Detection Matters As systems grow in scale and complexity, traditional operations approaches struggle to meet the demands of modern distributed systems. AIOps Anomaly Detection helps operations teams: Rapid Problem Resolution: Systematic tools and methods reduce troubleshooting time Improved System Visibility: Establish comprehensive monitoring and observability Proactive Fault Prevention: Identify and fix potential risks before they cause outages Resource Optimization: Allocate and schedule resources efficiently Core Concepts and Principles Basic Concepts The core of AIOps Anomaly Detection lies in establishing standardized processes and automated toolchains....

June 5, 2026 · 3 mins · 571 words · XuBaojin