Pods Stuck for 6 Minutes After Network Partition Injection: Blast Radius Control and 5 Production-Grade Decisions in Fault Injection Testing
Overview Netflix says chaos engineering reduced major production incidents by 70%. Many teams have seen this number, but those who have actually run fault injection in production are rare. The reason is simple: fear. Inject a network partition and Pods freeze—what then? Simulate disk full and the database crashes—then what? Manually injecting faults at 3 AM and scrambling to roll back—that’s not a drill, that’s manufacturing an incident. During a K8s migration for a new-energy logistics platform with 120+ microservices, I needed to validate disaster recovery capabilities before the full cutover....