Kubernetes Native Sidecars: Fix Startup Order
Kubernetes native sidecars are init containers that keep running. They start before your app, restart on their own, and get shut down last. Here is how.
For years, the sidecar pattern in Kubernetes rested on a lie of omission. You put two containers in a pod, called one of them a “sidecar”, and pretended the kubelet (the agent that runs pods on each node) knew which was which. It did not. Both were just entries in the containers list. They started at roughly the same time and were torn down at roughly the same time. Everything built on that fiction was a workaround.
I ran into this on the deployment side of CMS workflow operations at CERN, where services sit behind proxy and log-shipping containers on Kubernetes. Two failures in particular cost real time: the app crashing on boot, and logs lost on shutdown. Native sidecars fix both at the spec level instead of with shell tricks.
This post is for anyone who ships a pod with more than one container and has watched it crash on boot or lose data on shutdown. It covers the two races, what a native sidecar is, the full pod lifecycle, probes, and the edges that still catch people.
The two races: one at startup, one at shutdown
Startup: the app dials a proxy that is not listening yet
Say your application container talks to a database through a proxy sidecar: a connection pooler, a CERN-SSO auth proxy, whatever sits in front. In the old model both containers start together. So the app often dials the proxy a few hundred milliseconds before the proxy is listening.
The connection is refused, and the app exits non-zero. Kubernetes restarts it, and you get a crash loop that clears itself once the proxy happens to win the race. It looks flaky because it is flaky.
Shutdown: the log shipper leaves before the app
Shutdown is the same race running backwards. When a pod is deleted, the kubelet sends SIGTERM to every container at once. Your app wants a few seconds to finish in-flight requests and flush its last log lines. But the log-shipping sidecar next to it got the same signal and is already gone. The final logs, the ones you actually want when something died, never leave the node.
The workarounds
People papered over both races with the same kinds of hack:
- an init container that blocks until the proxy answers;
- a
preStophook that sleeps; - an app that retries its first connection for thirty seconds.
They mostly work, until the day a timeout is a touch too short. The regular init container already had the ordering guarantee you wanted, but it had the wrong shape. It must run to completion before the app starts, so it cannot host a process that needs to stay up.
What a native sidecar is: an init container that keeps running
The fix is almost anticlimactic. It reached beta and became on by default in Kubernetes 1.29, and went stable in 1.33. A native sidecar is an init container with restartPolicy: Always. That one field changes how the kubelet treats it.
A normal init container runs and exits, and only then does the next one start. An init container marked restartPolicy: Always behaves differently. It starts, and the kubelet moves on to the next init container (or to the app containers) as soon as this one has started, not finished. So it stays running. It gets the init sequence’s ordering for free, and it keeps living into the pod’s main phase like an ordinary sidecar. Always is the only value the field accepts here; any other value is rejected.
In this Deployment, the database proxy is a native sidecar with a startup probe, and the app container is unchanged:
apiVersion: apps/v1
kind: Deployment
metadata:
name: wmcore-console
spec:
template:
spec:
initContainers:
- name: db-proxy
image: registry.cern.ch/db-proxy:1.8
restartPolicy: Always # this line makes it a sidecar
startupProbe:
tcpSocket:
port: 5432
periodSeconds: 2
failureThreshold: 30
containers:
- name: app
image: registry.cern.ch/wmcore-console:2026.9
# by the time this starts, db-proxy has passed its startup probeThe sidecar moved out of containers and into initContainers, and it gained one line. What changed is the contract. The kubelet now guarantees that the proxy is up before the app starts, and it keeps the proxy alive until after the app is gone.
The full pod lifecycle, in order
The ordering guarantees are the whole point, so it helps to see the pod’s life as one timeline.
Reading the timeline left to right, startup goes like this:
- Ordinary init containers still run first, each to completion. A schema migration or a config fetch happens before anything else.
- Sidecars start next, in the order they appear. Each one waits until the previous sidecar is up.
- Only after every sidecar has started do the app containers start, all together as before.
On the way down, the order reverses:
- App containers get
SIGTERMfirst and get their fullterminationGracePeriodSecondsto drain. - Sidecars stop after that, in the reverse of the order they started.
Your log shipper is the last thing standing. That is exactly what you want, because the interesting logs are the ones from the app’s final seconds.
Probes: something init containers never had
A regular init container cannot have a liveness or readiness probe. The concept makes no sense for something that runs once and exits. A native sidecar can have all three probe types, and each does something specific:
startupProbecloses the boot race. The kubelet does not consider the sidecar “started”, and so does not let the app container start, until the startup probe passes. That is what the TCP check in the snippet above does: the app cannot start until port 5432 answers.readinessProbefeeds the pod’s overall readiness. A proxy that has lost its upstream connection can pull the whole pod out of a Service’s endpoints.livenessProbelets a wedged sidecar restart on its own, without taking the app down with it.
That independent restart deserves a closer look. During the pod’s running phase, the kubelet restarts a crashed native sidecar on its own, following the pod’s restartPolicy, without disturbing the app container. Under the old two-container model, a sidecar crash and an app crash were tangled together through the pod’s restart behavior. Now the proxy can die and come back while the app keeps serving through the blip.
Where it still bites
Native sidecars do not remove the need to think, and a few edges catch people.
The startup probe is load-bearing, not decorative. If you declare a sidecar without a startup probe, “started” means only that the container process launched, not that it is ready to serve. The app can still start before your proxy is actually listening. The ordering guarantee is about container start; the probe is what upgrades “the process exists” into “the port answers”. Leave it off, and you have quietly rebuilt the original race inside the new mechanism.
The grace period is shared, and it starts at the app’s SIGTERM. The pod has one terminationGracePeriodSeconds, and the clock starts when the app containers are signaled. The sidecar’s own shutdown happens in whatever is left of that window after the app drains. If your app uses the whole grace period and the log shipper needs three seconds to flush, budget for both when you set the number. I have written separately about getting FastAPI to shut down cleanly under Kubernetes; the sidecar’s flush time is now part of that same budget.
A sidecar that fails to start can wedge the pod. This works the same way as a failing regular init container. With the pod’s restartPolicy set to Never, a sidecar that cannot start means the pod does not start, full stop. That is usually what you want (no proxy, no app). But it does mean a broken sidecar image causes a pod-level outage, not a degraded-but-running pod. Watch the init phase in kubectl describe pod: a sidecar stuck starting shows up there, not in the main container status.
You need a new enough Kubernetes version. The feature is stable from 1.33 and on by default from 1.29. But if any cluster in your fleet is older than 1.28, the restartPolicy field on an init container is either ignored or rejected, depending on how old the cluster is. If it is ignored, your “sidecar” silently reverts to a blocking init container that never exits, and that hangs the pod. When you manage deployments across clusters with something like ArgoCD ApplicationSets, confirm every target is new enough before you rely on this. On an old node, the failure is a pod that never becomes ready, not a clear error.
What I would change first
If you have pods using the old two-container sidecar pattern, the migration is mechanical and low-risk:
- Move each sidecar from
containersintoinitContainers. - Add
restartPolicy: Always. - Give every sidecar the app depends on a
startupProbethat tests real readiness, not just that the process is alive.
The payoff is that a whole category of boot-order flakiness and lost shutdown logs stops being your problem and becomes the kubelet’s.
This matters to me for the same reason I care about readiness and liveness probes and about sizing memory limits so a pod fails predictably. The operational failures that wake someone up are almost never the exciting ones. They are ordering, timing, and cleanup. Native sidecars take one of those three and turn a pile of hooks and retry loops into a guarantee the platform makes for you. For the workflow tooling I ran for CMS, fewer moving parts in the boot path was worth more than any feature I could add on top.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.