Operations · 8 min read

Kubernetes network policy: a safe default-deny rollout

Default-deny sounds simple until a liveness probe fails at 3 a.m. This is the rollout order we use across validator clusters, the canaries that catch mistakes early, and the rollback plan for when one slips through.

Every managed Kubernetes guide eventually lands on the same advice: deny all traffic by default, then allow only what each workload needs. On a stateless web app this is an afternoon of work. On a validator cluster where a missed peer connection means missed attestations, the order of operations matters more than the policy itself.

Start with audit mode, not enforcement

The first pass deploys every intended policy in audit mode. Nothing is blocked; the CNI only logs what would have been dropped. Two full weeks of logs across a cluster tells you what your workloads actually talk to, which is never quite what the architecture diagram claims.

Canary namespaces first

Enforcement starts in namespaces running non-critical workloads: monitoring, internal tooling, block explorers on testnets. Validators come last, one cluster at a time, with a scheduled window and a rollback manifest applied in the same change request. A policy that cannot be reverted in ninety seconds does not ship.

The probes will find your mistakes

The most common failure is not peer traffic at all. It is the kubelet liveness probe, the metrics scrape, the node-local DNS cache. These live outside the namespace boundary everyone draws on the whiteboard, and they are the first things a default-deny policy silently kills. Allow-list them explicitly, per namespace, before any workload rule lands.

The result is a cluster where every allowed flow is written down, reviewed, and owned by a named service. That is the actual point of the exercise.