PodDisruptionBudgets: Survive a Node Drain
A cluster upgrade drains a node and evicts every replica of your service at once. How a PodDisruptionBudget caps voluntary disruptions and keeps you online.
The first time this bit me, no deploy was involved. A platform engineer was upgrading the kubelet across the cluster, node by node, the way you are supposed to. They ran kubectl drain node-7, and it did what drain does. Two of the three replicas of an API service happened to be on that node, and for about eight seconds both went down at once. The third was mid-restart from an unrelated config reload.
For those eight seconds the service had zero Ready pods. Upstream health checks flipped red, and a queue backed up behind it. Nobody deployed anything, and nobody wrote a bad line of code. Kubernetes did exactly what it was told, and what it was told left no floor under the service.
A PodDisruptionBudget (PDB) gives you that floor. This post is for engineers who run their own workloads on Kubernetes and have felt, or want to avoid, a maintenance operation taking down a service that no deploy touched. I ran the platform side of CMS workflow operations at CERN. There, the WMCore services and their operator console sit on a shared cluster that gets patched, drained, and rescheduled on someone else’s timetable. A PDB is how you tell that timetable where your limits are.
Voluntary versus involuntary disruptions: only one is yours to shape
Kubernetes splits pod deaths into two buckets, and PDBs exist because of that split.
An involuntary disruption is one nobody chose: a node’s kernel panics, the hardware fails, the kubelet gets OOM-killed, or a cloud provider reclaims a spot instance. You cannot budget for these, because there is no request to approve or deny; the pod is simply gone. Your only defence is running enough replicas across enough nodes that losing one does not matter.
A voluntary disruption is one initiated through the API. Examples include draining a node for an upgrade, the cluster autoscaler removing a node, or an operator deliberately deleting a pod. These go through the Eviction API, and that is the hook a PDB uses. A PDB does not stop nodes from failing. It sets a rule that every voluntary eviction has to respect. That way the routine, planned operations that happen constantly in a healthy cluster do not stack up into an outage.
How a drain respects the budget
kubectl drain is the case you hit most. It does not delete pods directly; it calls the Eviction API for each one. Here is what happens for each pod:
- The Eviction API checks the pod against any PDB that selects it.
- If evicting the pod would drop the number of healthy pods below the budget, the API refuses with HTTP
429 Too Many Requests. - Drain does not give up on a 429. It backs off and retries.
So the drain proceeds as fast as the budget allows, and no faster. One replica goes, a replacement schedules elsewhere and becomes Ready, and then the next one is allowed to go. What used to be a simultaneous cull becomes a controlled, rolling handoff.
Writing the budget
A PDB is a small object: a label selector, plus exactly one of two limits. This one protects the wmagent-console pods:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: wmagent-console-pdb
spec:
minAvailable: 2
selector:
matchLabels:
app: wmagent-consoleIt says: across all pods labelled app: wmagent-console, keep at least two Ready at all times during voluntary disruptions. If an eviction would leave fewer than two Ready, for example because a replacement pod is not Ready yet, the eviction waits.
The selector works like any other label selector, and it should match the same pods your Deployment manages. But the PDB is a separate object, so nothing enforces that the labels line up. Point it at the wrong label, and you have a budget that guards nothing.
minAvailable or maxUnavailable
You express the limit as either minAvailable (how many pods must stay up) or maxUnavailable (how many may go down), never both. Both accept an integer or a percentage.
On a five-replica service, minAvailable: 4 and maxUnavailable: 20% describe the same state today: at most one pod down. They stop agreeing the moment the replica count changes.
That difference decides which one to use:
- A service behind the Horizontal Pod Autoscaler, whose replica count moves with load, is usually better served by
maxUnavailableas a percentage. The budget then grows and shrinks with the fleet instead of pinning a stale absolute number. - A fixed-size service where a specific quorum matters (say, three nodes of a consensus system) wants
minAvailableas an integer. “Keep two up” should mean two, not two-thirds of whatever the count happens to be.
The single-replica trap
The most common way a PDB backfires also looks the most logical. You have one replica of something, you do not want it disrupted, so you write minAvailable: 1.
Now kubectl drain on that pod’s node hangs forever. Evicting the only pod would leave zero Ready, which violates the budget, and there is no second pod whose readiness the eviction could wait for. The drain cannot make progress. The platform engineer trying to patch that node is now blocked on your service, with no way forward except overriding you.
A PDB does not create availability; it protects availability you already have. With one replica, no PDB configuration gives you a zero-downtime drain. The single pod has to move, and nothing covers for it while it does. The honest fix is more than one replica across more than one node.
If a workload genuinely cannot run two copies, it cannot be drained without downtime, and the useful thing a PDB does there is make that explicit. Better to force a conversation about a maintenance window than to discover the limitation during an unplanned node failure, where no PDB helps anyway.
Unhealthy pods that block a drain, and the setting that fixes it
For years there was a nastier version of the stuck-drain problem. Suppose your budget is minAvailable: 2 on three replicas, and one of the three is wedged. It is crash-looping, or stuck in a bad Running-but-not-Ready state after a config push. You are already at the floor of two Ready pods.
Now someone drains the node holding the broken pod. Under the old behaviour, the eviction was refused. The budget had only two healthy pods, and evicting anything risked dropping below that, even though the pod being evicted was the broken one doing you no good. The drain stalled on a pod you would have been happy to lose.
Kubernetes 1.27 made the unhealthyPodEvictionPolicy field generally available to break this deadlock. This budget sets it to AlwaysAllow:
spec:
minAvailable: 2
unhealthyPodEvictionPolicy: AlwaysAllow
selector:
matchLabels:
app: wmagent-consoleThe field has two values:
IfHealthyBudget(the default) only lets you evict a not-Ready pod when the application currently has more healthy pods than its minimum. That is safe, but it blocks exactly the case above.AlwaysAllowsays a not-Ready pod can always be evicted, whatever the budget. The reasoning is that a pod not serving traffic is not part of the availability you are protecting.
I now default to AlwaysAllow for stateless HTTP services. Keep the conservative default for anything where an unhealthy pod might still be doing useful work you cannot see from its readiness probe, such as draining connections or finishing a write.
Where the mental model breaks
A few things about PDBs consistently trip people up. All of them come from expecting the budget to govern more than it does.
A PDB does not throttle your rolling deploy. This is the big one. When you kubectl apply a new Deployment spec, the Deployment controller replaces pods according to its own maxUnavailable and maxSurge settings under strategy.rollingUpdate. It deletes pods directly, without going through the Eviction API, so the PDB has no say. A PDB with minAvailable: 3 will not stop a Deployment with maxUnavailable: 50% from taking half your pods down during a rollout.
The two settings guard two different operations: the Deployment strategy governs deploys, and the PDB governs evictions. Configure both, and configure them consistently, because a service is only as available as the weaker of the two.
Direct deletion bypasses the budget. kubectl delete pod removes a pod immediately without consulting any PDB, because it is not an eviction. The budget is a contract with the Eviction API specifically, not a lock on the pod. That is a feature, since you need an escape hatch. But it means a PDB is a guardrail for well-behaved automation, not a hard guarantee against everything that can delete a pod.
Two PDBs matching one pod is an error, not a merge. If more than one PDB selects a pod, the Eviction API cannot decide which budget applies, so it refuses the eviction outright. Overlapping selectors, such as a broad namespace-wide PDB plus a per-app one, will quietly wedge your drains. Keep selectors disjoint, and be careful with empty selectors, which match every pod in the namespace.
A percentage minAvailable rounds up. minAvailable: 50% on three replicas requires two Ready pods, not one and a half rounded down to one. That seems minor until it decides whether a drain flows or stalls. Do the arithmetic against your real replica count instead of assuming the friendly number.
What I actually put on a service
For a normal stateless service on our cluster, the pattern is boring, and that is the point:
- Run at least three replicas, with a pod anti-affinity rule so they do not all land on one node.
- Set a
maxUnavailable: 1PDB (or a percentage, if the service autoscales). - Set
unhealthyPodEvictionPolicy: AlwaysAllow. - Make sure the pods actually shut down cleanly when the SIGTERM from an eviction arrives.
That last piece is where a lot of the value leaks away. A PDB gets you a graceful, one-at-a-time handoff, and then the application drops in-flight requests on shutdown anyway because it never handled the signal. I wrote about closing that gap in graceful shutdown for FastAPI on Kubernetes. It pairs directly with a PDB, and each piece does a different job:
- the budget decides when a pod is allowed to go;
- readiness probes decide whether the replacement is ready to take its place;
- clean shutdown decides whether the departing pod takes any requests down with it.
None of the three does much alone. Together, they separate a cluster upgrade that nobody notices from the eight quiet seconds that started this post.
A PDB is cheap to write and easy to get subtly wrong. The only way to know yours works is to drain a node in staging and watch what the Eviction API actually allows. That test costs five minutes. Finding out during a real upgrade costs a lot more.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.