A Kubernetes control plane can be healthy at the node and pod layers while its API server is becoming unusable. A buggy operator may relist thousands of objects, a deployment controller may retry aggressively during an outage, or a fleet of AI agents may create short-lived jobs and watch many resources at once. CPU limits on those pods do not protect `kube-apiserver`; the scarce resource is API-server concurrency.
API Priority and Fairness (APF) is Kubernetes' built-in overload-control system. It has been stable since Kubernetes v1.29 and is enabled by default. APF classifies requests with `FlowSchema` objects, assigns them to `PriorityLevelConfiguration` objects, and then applies isolated concurrency limits plus fair queues. The goal is not to make every request fast. It is to preserve useful progress when demand exceeds what the API server can execute.
This matters more as clusters gain custom controllers and autonomous clients. A controller that is merely inefficient in quiet conditions can become a control-plane incident when many tenants reconcile simultaneously. APF gives platform teams a way to contain that blast radius without granting every critical client an unlimited exemption.
APF is not pod priority or CPU QoS
Pod priority affects scheduling and preemption on worker nodes. Resource requests and limits govern container CPU and memory. APF governs HTTP requests entering `kube-apiserver`. A pod can use little CPU yet issue expensive `LIST` calls that return thousands of objects; Kubernetes notes that large list operations can occupy more than one execution seat.
The request path is: match a `FlowSchema`, choose a priority level, distinguish the request's flow, then dispatch, queue or reject it. Lower numeric `matchingPrecedence` values are evaluated first, and the first match wins. That ordering is easy to get wrong: a broad rule placed earlier can capture traffic intended for a more specific rule.
A `FlowSchema` can match users, groups or service accounts and combine that identity with verbs, API groups, resources and namespaces. Its `distinguisherMethod` can separate flows by user or namespace. Separation matters because fairness works between flows inside a priority level; without a useful distinguisher, many tenants can collapse into one flow and share the same fate.
Design priority levels around failure domains
Start from traffic classes, not individual applications. Keep Kubernetes' mandatory and suggested defaults visible, especially node health, system controllers and leader election. Add a limited priority level for custom automation that can tolerate delay, such as report generators, deployment bots, policy scanners and AI agents. Use a separate level only when its failure behavior, latency objective or ownership is meaningfully different.
A limited priority level defines nominal concurrency shares, whether unused capacity is lendable, and how overload behaves. `Reject` immediately returns HTTP 429 when the limit is reached. `Queue` absorbs short bursts and dispatches work fairly. Queueing is usually appropriate for well-behaved controllers because Kubernetes clients are expected to retry with backoff, but a queue is not free capacity: a long queue converts overload into latency and memory use.
apiVersion: flowcontrol.apiserver.k8s.io/v1
kind: PriorityLevelConfiguration
metadata:
name: automation-low
spec:
type: Limited
limited:
nominalConcurrencyShares: 10
lendablePercent: 50
limitResponse:
type: Queue
queuing:
queues: 64
handSize: 6
queueLengthLimit: 50Treat those values as a starting experiment, not universal tuning. More queues reduce collisions between unrelated flows but consume more memory. A larger `queueLengthLimit` absorbs a bigger burst but permits more waiting. A larger `handSize` reduces the chance that two specific flows collide, yet allows a small number of flows to influence more queues. Kubernetes uses shuffle sharding to choose a small hand of queues per flow; the isolation comes from that probabilistic separation, not from a dedicated queue for every client.
Match the client narrowly and keep an escape path
Bind the priority level to a narrow `FlowSchema`. For example, isolate a service account used by an automation controller and restrict the match to the verbs and resources it truly needs. Avoid a wildcard rule for all service accounts unless the intent is explicitly global.
apiVersion: flowcontrol.apiserver.k8s.io/v1
kind: FlowSchema
metadata:
name: automation-jobs
spec:
matchingPrecedence: 800
distinguisherMethod:
type: ByNamespace
priorityLevelConfiguration:
name: automation-low
rules:
- subjects:
- kind: ServiceAccount
serviceAccount:
name: job-orchestrator
namespace: automation
resourceRules:
- verbs: [get, list, watch, create, patch]
apiGroups: ["", batch]
resources: [pods, jobs]
namespaces: ["*"]Before applying a rule, inspect the existing objects and determine which schema currently matches representative requests. Preserve a break-glass administration path, but do not casually put application clients in the `exempt` level. Exempt traffic bypasses flow control; a compromised or looping exempt client can consume the very capacity APF is meant to protect.
Recursive request paths need extra care. Admission webhooks and aggregated API servers may call back into `kube-apiserver` while the original request is still holding capacity. Kubernetes warns that applying priority control to both layers can create priority inversion or deadlock. Model these call graphs explicitly and ensure the dependent request can progress at a sufficiently high priority.
Reduce API work before buying more seats
APF is a safety system, not permission to keep inefficient clients. Prefer shared informers and watches over repeated full lists. Paginate large lists, scope caches by namespace when possible, add jitter to fleet-wide resynchronization, and use exponential backoff for 429 and transient failures. Do not launch one identical polling loop per tenant if a single watch can feed tenant-specific work.
For AI agents, keep cluster observation behind a bounded tool service. The service can cache discovery, enforce namespace and verb scopes, collapse duplicate reads and queue mutations. An agent should not receive broad credentials and freely explore the API server in a reasoning loop. APF contains demand after authentication; it does not replace RBAC, admission control or tool-level budgets.
If APF rejects requests across many priority levels while workload volume is stable, investigate control-plane latency rather than only increasing shares. Slow storage, admission latency or an overloaded API server causes requests to hold seats longer. Redistributing a fixed pool cannot repair a saturated dependency.
Operate from four signals
The Kubernetes metrics reference exposes the signals needed to distinguish a burst from persistent starvation. Alert on the rate of `apiserver_flowcontrol_rejected_requests_total` by `flow_schema`, `priority_level` and `reason`. Track `apiserver_flowcontrol_current_inqueue_requests` and the p95/p99 of `apiserver_flowcontrol_request_wait_duration_seconds`. Compare them with `apiserver_flowcontrol_nominal_limit_seats` and occupied-seat metrics.
Interpret them together. Rejections with `queue-full` and rising wait time indicate sustained overload. `concurrency-limit` means the level rejects rather than queues. High wait time without rejection may still violate controller deadlines. Broad rejection across levels, combined with rising execution latency, points to a control-plane bottleneck rather than a single noisy flow.
Roll out one traffic class at a time. Capture a baseline, apply the FlowSchema and priority level in a noncritical cluster, replay a controlled burst, and verify that leader election and system controllers keep progressing. Then test 429 handling, queue delay, rule precedence and rollback. Store APF objects as versioned infrastructure code, but compare against server-managed defaults during Kubernetes upgrades because suggested objects have auto-update behavior.
The practical invariant is simple: nonessential automation may become slower or receive 429 responses, but it must not prevent node health, leader election and essential reconciliation from advancing. When that invariant is measurable, APF turns control-plane overload from an undifferentiated outage into an isolated, diagnosable failure mode.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




