API Gateway Monitoring in Practice: Metrics Collection and Alert Design from Kong to APISIX
Overview At 2 AM, the payment system alerts exploded. After digging through logs for ages, we discovered the downstream services weren’t down — the API gateway’s connection pool was exhausted. Requests piled up at the gateway layer and never made it through. But our monitoring dashboard only showed CPU and memory curves, completely blind to gateway-level latency, connection counts, and error rates. This isn’t an isolated case. I’ve seen too many teams treat API gateways as “fancy Nginx” and only monitor nginx_active_connections, only digging into logs when things break....
Don't Wait for the Crash to Check Monitoring: A Practical Methodology for Linux Performance Baselines and Anomaly Detection
Overview You’ve probably been here: woken up at 3 AM by a pager alert, scrambled to check Grafana, and found CPU usage spiked to 85%. You panic, investigate for half an hour, and finally realize — this database server runs a scheduled backup job every day from 2:50 to 3:10 AM. CPU is supposed to be at this level. What went wrong? You didn’t know what “normal” looks like for this machine....
From "@ Ops in Chat" to Self-Service: 6-Layer Architecture and 4 Landing Phases for Platformizing 20+ Operations
Overview Friday, 4:30 PM. The business chat starts flooding: “@ops help me restart the order service” “@ops this endpoint is erroring, grab me the logs” “@ops is the disk full again? clear it for me” Three requests, three @s. The ops engineer puts down the inspection script they were writing, SSHs in, runs commands, pastes screenshots back to chat, then gets asked “is it done yet?” An afternoon, shattered. This isn’t an edge case—it’s the default mode of ops work in most teams: high-frequency, low-difficulty, heavily manual....
The 3 AM Phone Rings Again: 7 Engineering Decisions for On-Call Rotation Design and Fatigue Management
Overview 3 AM. Your phone buzzes. P1 alert — order service 5xx error rate spiking. You drag yourself out of bed, fire up the laptop, pull logs, check Grafana, trace it to a saturated upstream database connection pool. Restart, scale up, recover. It’s 4:30 AM. The next morning, nursing dark circles in the standup, someone mentions that last night’s alert was actually a false positive. If you’ve worked in operations, you know this feeling....
Don't Fall for the No-Ops Myth: 5 Production-Level Governance Decisions for Serverless Cold Starts, Cost Explosions, and Observability Blind Spots
Overview At 2 AM, I got a phone call. Users of a ride-hailing project reported that opening the App to view trip history during late-night hours caused a 3-5 second blank screen. After investigation, I found that this API group used Serverless function compute. During the day, traffic was normal, but after traffic dropped off at night, function instances were reclaimed. The first request triggered a cold start, pushing P99 latency to 4....
Cutting 80% of Annual Downtime: 6 Engineering Decisions and Hard-Learned Lessons Going from 99.5% to 99.9% Availability
Overview 99.5% availability sounds pretty good — until you do the math. A year has 8,760 hours. 99.5% means you’re allowed 43.8 hours of downtime. Per month, that’s roughly 3.65 hours of service interruption. If this is a core transaction system, those 3.65 hours could mean tens of thousands in lost orders. During a major promotion, the loss multiplies by orders of magnitude. 99.9%? Annual downtime drops to 8.77 hours, less than 44 minutes per month....
False Positive Rate from 70% to 5%: A Four-Layer Security Gate Design and False Positive Management in CI/CD
Overview Plug a SAST tool into your Jenkins pipeline, let the scan finish with zero alerts, then tell your boss “we’ve adopted DevSecOps” — I’ve seen this playbook too many times. During a Level 2 cybersecurity protection audit for a ride-hailing project, the client’s security team required us to integrate security scanning into our CI/CD pipeline. The first version was textbook standard: SonarQube for code scanning + Trivy for image scanning, with the gate set to “block Critical....
Don't Turn Terraform Modules Into Black Boxes: IaC Layered Design Decisions and 6 Production Anti-Patterns Broken Down
Overview Last year I was setting up an IaC system for an e-commerce client. When I took over, their Terraform codebase looked roughly like this: three environments (dev/staging/prod), each with its own copy of the configuration, totaling over 3,000 lines—90% of which was copy-paste. Changing a VPC CIDR meant editing three files in sync. Miss one, and you get environment drift. The worst part: a security group rule was missing in prod for three months without anyone noticing....
Starting from Docker Machine Deprecation: 7 Production Decisions for GitLab CI Runner Elastic Architecture and Cache Governance
Overview 1 AM. You get an alert: GitLab CI pipeline queue has 47 pending jobs. The dev chat explodes—“code pushed 40 minutes ago, still not running"“is the Runner down?““just add more machines!” You check the console: all 3 Runners are at capacity, each running 10 concurrent jobs, 30 slots fully occupied. Add machines? Docker Machine executor needs 3 minutes to spin up a new EC2, then another 2 minutes to become Ready....
Disaster Recovery Is Not Backup: Cross-Datacenter RPO<5min, RTO<30min Architecture Decisions and Lessons Learned
Overview At 2:17 AM, my phone buzzed with an alert: core datacenter A network equipment failure, database primary-standby sync interrupted. The on-call SRE switched to the disaster recovery datacenter B, only to find that after applications connected to the new primary, some order data was missing — the async replication window had lost 47 seconds of data. Recovery took 38 minutes, exceeding the RTO target by 8 minutes. This was a real failover incident....