Incidents
You are paged into a cluster that is already failing. The fault is always a control-loop decision and the fix is always a field. Where the fault is in the application instead, that is debug.liter8.sh, and this links there rather than teaching it twice.
Checkout is returning 500s during every deploy
Probes · realSupport has had a handful of complaints during each of the last four deploys. Errors spike for about twenty seconds and then stop on their own. Every pod is Running, every pod is Ready, nothing has restarted, and the rollout reports success. The database is fine.
The node upgrade has been running since Tuesday
Disruption · realSomebody started draining a node three days ago and it has not finished. The node is cordoned. Nothing has errored, nothing has paged, every pod on the cluster is Running and Ready, and the drain command is still sitting there. There is nothing in the logs.