Kubernetes multi-zone reliability is not achieved by adding three replicas and hoping the scheduler separates them. Define which nodes are eligible, spread matching Pods across zones and hosts, decide whether degraded topology should block placement, reserve capacity for a lost zone, and test the failure. Pod topology spread constraints are the right primitive when you want a bounded imbalance such as 2-2-1; anti-affinity is better for a strict one-per-domain rule. Neither replaces application redundancy, data replication, traffic health checks or a disruption budget.
The official Kubernetes documentation referenced here was verified on 6 October 2026. The field behavior is documented fact; the rollout thresholds, capacity model and drills below are engineering recommendations.
Start with the failure you must survive
A topology domain is a set of nodes sharing a label value, commonly `topology.kubernetes.io/zone` or `kubernetes.io/hostname`. A scheduler cannot infer your business failure domains from replica count. If all three replicas are eligible for the same large node, a healthy cluster can still place them together and one node failure can remove the service.
Write the availability contract first: one node may fail without losing all serving replicas; one zone may fail while the remaining zones can carry the required traffic; a rolling update must not collapse healthy replicas into one zone; and a scale-up must not wait forever for a topology domain that cannot exist. These are different guarantees and require different constraints plus capacity.
The Kubernetes topology spread documentation describes `spec.topologySpreadConstraints` as a way to control Pod distribution across failure domains. It also states an important limit: the scheduler places incoming Pods, but scaling down or later node changes can leave the existing distribution imbalanced. Scheduling policy is not a continuous rebalancer.
Build the eligible-node set before calculating skew
The scheduler first needs a valid candidate set. Use node affinity for requirements such as architecture, compliance tier or dedicated node pool. Use taints and tolerations to repel general workloads from specialized nodes. A toleration only permits placement on a tainted node; it does not require that placement.
After eligibility, topology spread counts matching Pods in each eligible domain. This order matters. If a workload is allowed only in zones A and B, counting zone C would create a misleading minimum. `nodeAffinityPolicy: Honor` makes the spread calculation follow node affinity or selectors. `nodeTaintsPolicy: Honor` counts only untainted nodes or tainted nodes the incoming Pod tolerates. Kubernetes documents both inclusion policies as GA from v1.33.
Keep topology labels consistent. Missing `topology.kubernetes.io/zone` labels make nodes invisible to that constraint, and private label keys can mean different things across environments. Validate labels during node provisioning, not after a placement incident.
Understand maxSkew and minDomains precisely
`maxSkew` bounds the difference between matching Pod counts and the global minimum. With counts 2, 2 and 1 across three zones, `maxSkew: 1` is satisfied. With 3, 1 and 1, the skew is 2 and a strict constraint rejects a new placement that keeps the imbalance.
`minDomains` says how many eligible domains must exist before the ordinary minimum applies. When fewer eligible domains exist, Kubernetes treats the global minimum as zero. With `minDomains: 3`, losing one of three zones can therefore make a strict `DoNotSchedule` constraint intentionally block new Pods that would exceed the allowed skew. That can protect the declared topology contract, but it can also stop recovery during an outage.
Choose the behavior from the service objective. Use `DoNotSchedule` for quorum-sensitive systems where concentrating replicas would create a worse failure mode. Use `ScheduleAnyway` for stateless frontends that should keep serving in a degraded topology; the scheduler prefers the less-populated domain but can place elsewhere when capacity is constrained. A strict policy without spare capacity often turns a zone incident into Pending Pods rather than resilience.
A production Deployment pattern
The following fragment spreads a Deployment across zones and then across hosts. The zone rule is strict, while the host rule is preferred so a small cluster does not deadlock. The selector must match the Pod's own labels; otherwise Kubernetes warns about "ghost Pods" that do not count themselves.
apiVersion: apps/v1
kind: Deployment
metadata:
name: checkout
spec:
replicas: 6
selector:
matchLabels:
app: checkout
template:
metadata:
labels:
app: checkout
spec:
topologySpreadConstraints:
- maxSkew: 1
minDomains: 3
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: checkout
matchLabelKeys:
- pod-template-hash
nodeAffinityPolicy: Honor
nodeTaintsPolicy: Honor
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: checkout
containers:
- name: app
image: registry.example.com/checkout@sha256:<reviewed-digest>
resources:
requests:
cpu: 500m
memory: 512Mi`matchLabelKeys: [pod-template-hash]` separates Deployment revisions in the spread calculation so a rolling update does not mix old and new ReplicaSets as one population. The Kubernetes documentation notes that since v1.34 the matching key-value labels are explicitly merged into the selector. Confirm the behavior on the exact Kubernetes version you run.
Do not copy `minDomains: 3` into a two-zone cluster. Do not make both zone and hostname constraints strict until you have modelled every replica count, rollout surge and node-pool limit. Multiple topology constraints are combined with logical AND; the intersection can be empty even when each constraint looks reasonable alone.
Capacity is part of the scheduling contract
A 2-2-2 placement across three zones survives one zone only if the remaining zones can run the displaced replicas or if four replicas are enough for the service objective. Requests drive scheduler feasibility, so unrealistic CPU or memory requests make the topology contract impossible; undersized requests admit placements that later throttle or OOM. The related Kubernetes requests and limits guide covers that capacity layer.
Model N-1 capacity explicitly. For each critical workload, record the minimum serving replicas after a zone loss, required headroom in surviving zones, maximum acceptable Pending duration and autoscaler reaction time. If a node group can scale to zero in a topology domain, Kubernetes documents that the scheduler may not know that domain exists until at least one node is present. Use an autoscaler aware of topology spread, or retain a minimal footprint for required domains.
Pod Scheduling Readiness adds scheduling gates that keep a Pod out of active scheduling until an external controller removes the gate. The official documentation positions it as a way to avoid unnecessary scheduling attempts. It is useful for coordinated placement, but a stuck controller becomes a Pending-Pod failure mode, so expose gate age and ownership as metrics.
Best practices and anti-patterns
Prefer topology spread over required Pod anti-affinity when the real requirement is a bounded distribution rather than exactly one Pod per domain. Required anti-affinity across hostname with more replicas than nodes guarantees Pending Pods. Spread constraints can express a tolerable skew and combine zone and host objectives.
Keep the selector narrow enough to represent one failure unit. Do not count unrelated applications together because they share a generic label. Keep the same constraints across every Pod in the group, and use admission policy or a workload library to prevent drift.
Avoid hidden contradictions: a node selector for one zone plus `minDomains: 3`; a toleration that changes which nodes count; a strict zone rule with a single-zone volume; or a deployment surge that temporarily needs more eligible slots than the cluster can offer. Inspect scheduler events and the exact eligible domains before loosening policy.
Do not use `ScheduleAnyway` and call the result guaranteed high availability. It is a preference. Do not use `DoNotSchedule` without an incident plan for Pending replicas. And do not assume current balance survives scale-down; if rebalancing is important, use an approved operational process and disruption-aware tooling.
Test the guarantee, not only the manifest
In staging, record the initial Pod counts by zone and node. Cordon and drain one node, then verify replacement placement, service health and rollout behavior. Next simulate an unavailable zone or remove all eligible capacity from it. Confirm whether strict constraints produce the expected Pending state or whether a degraded-placement policy maintains service.
Test replica counts below the number of zones, equal to it and above it. Test a rolling update with `maxSurge`, an HPA scale-up, a cluster autoscaler delay and a tainted node pool. Verify the selector matches the Pods and that node labels remain correct. Alert on unschedulable Pods by reason, topology imbalance, eligible-domain count, gate age and capacity headroom.
Pair scheduling tests with readiness, graceful termination and disruption budgets. The safe Kubernetes rollout guide explains how those controls determine whether a correctly placed Pod actually serves and exits safely. Placement protects against correlated infrastructure loss; readiness and draining protect traffic during change.
When not to use topology spread
A single-replica development service gains no zone availability from spread constraints. A workload tied to one zonal disk cannot become multi-zone through scheduler configuration alone. A database with its own consensus operator may require operator-defined placement rules rather than a generic Deployment policy. Small clusters can also be safer with a preferred rule and clear degraded-mode behavior than with a strict configuration that routinely blocks deployments.
If the requirement is co-location for latency, use Pod affinity. If the requirement is hardware or compliance eligibility, use node affinity and taints. If the requirement is exactly one replica per node, consider required anti-affinity or a DaemonSet. Topology spread is best when the requirement is a measurable distribution across failure domains.
Production decision
Treat scheduling as an availability contract: define eligible nodes, choose the topology domains, set a justified skew, decide how degraded topology behaves, provision N-1 capacity and verify failure modes. A balanced dashboard is not the goal; continued correct service after a node or zone loss is.
Start with a preferred hostname spread and a carefully tested zone rule. Make the zone rule strict only when the workload, storage and capacity can honor it during rollouts and failures. Connect this policy to cloud and delivery engineering so node provisioning, autoscaling, observability and incident response enforce the same contract.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




