Skip to content
Triage a namespace

Triage a namespace

Something is wrong in a namespace and you don’t yet know what. The kubectl version of this is four commands and a lot of scrolling: get pods, then describe the ones that look off, then logs, then get events.

kx diag

sweeps the current namespace — Deployments, StatefulSets, DaemonSets, Jobs, CronJobs, Services, PersistentVolumeClaims and Ingresses, plus pods nothing owns — and prints what’s unhealthy, worst first.

Healthy resources are left out of the terminal table. --full puts them back.

Each row shows one finding, so which of several equally severe findings sorts first decides what the whole sweep reads like. Rows are ordered by severity, and within a severity by how specific the finding is: a concrete cause — a container state, a scheduling refusal, an exceeded limit — outranks a rollup like “Only 0/3 replicas ready”, which outranks a raw warning event. Without that, five differently-broken Deployments all headlined the same replica count and the actual cause sat one screen down.

The rows are indexed

That’s the part that makes it a starting point rather than a report:

kx diag       # 1  api-badimage   ImagePullBackOff
              # 2  cache-oom      CrashLoopBackOff
kx diag 1     # the full diagnosis of row 1
kx logs 2     # straight into the logs of row 2

One resource

kx diag <index> diagnoses a single resource and prints, on one screen:

  • a verdict banner
  • a SUMMARY of findings
  • a per-pod status table
  • recent log tails from broken containers
  • warning events

The findings it looks for include CrashLoopBackOff, image pull failures, OOMKills, unschedulable pods, stalled rollouts, Services with no endpoints, Pending PVCs, failed CronJob runs, and Ingresses pointing at Services that don’t exist.

Usage as a signal, not just state

Findings also draw on live resource usage, the same data kx top reports. A pod running hot against its memory limit is flagged as an OOMKill risk before it gets killed, which is the one finding you cannot get from describe.

Nodes

A Node is diagnosed the same way, by index rather than by sweep — it is cluster-scoped, so it appears in neither a namespace sweep nor -A.

kx get nodes
kx diag 1

See taking a node out of service.

Wider than one namespace

kx diag -n prod    # a namespace you aren't in
kx diag -A         # every namespace

-A indexes the sweep too, and adds a NAMESPACE column beside the numbers, so kx logs 7 reaches whichever namespace row 7 came from.

As a check

kx diag -A --json
kx diag -A --fail-on critical

--json prints the same sweep as a document, and --fail-on exits 2 when any resource reaches that verdict. See using kx in CI.

In a browser

kx diag --html

renders the same sweep as a filterable, sortable page and opens it — with a group-by for larger sweeps, and each row expanding into that resource’s full report. See browser reports.

kx diag --html dashboard in the github-dark theme kx diag --html dashboard in the dracula theme kx diag --html dashboard in the nord theme kx diag --html dashboard in the gruvbox theme kx diag --html dashboard in the solarized-dark theme kx diag --html dashboard in the catppuccin-mocha theme kx diag --html dashboard in the tokyo-night theme kx diag --html dashboard in the rose-pine theme kx diag --html dashboard in the mono theme kx diag --html dashboard in the light theme