A normal load balancer sees requests. A large-language-model server sees work whose cost can differ by orders of magnitude: one prompt may reuse a warm prefix, another may fill the context window, and a third may target a LoRA adapter loaded on only one replica. Sending one request to each Pod in turn can therefore create an apparently even request count and a badly uneven queue.
The current Gateway API Inference Extension addresses that mismatch for self-hosted generative models on Kubernetes. An HTTPRoute can target an InferencePool rather than a plain Service. The Gateway asks an Endpoint Picker, or EPP, which model-server endpoint best satisfies the request, then forwards to that endpoint. The official documentation says the decision can use model-serving metrics and capabilities such as prefix-cache state and LoRA availability; the InferencePool guide specifically names KV-cache utilization, pending queue length and active adapters.
This is a current-documentation explainer, not a claim that the project launched today. The documentation was verified on 29 September 2026. It now labels InferencePool v1 as stable, while also documenting an active migration of project pieces into Gateway API and llm-d repositories. Treat the API contract, your chosen Gateway implementation and your production EPP implementation as three related but separately versioned components.
Start with the scheduling failure, not the CRD
Round robin answers a network question: which healthy endpoint has received fewer turns? It does not answer an inference question: which replica can start this request soonest and finish it with the least avoidable work? Request count hides prompt length, generated-token budget, batch state, cache reuse and adapter placement.
Least-connections is only a little better. A long decode may hold one connection while consuming accelerator time; many short requests may be batched efficiently. CPU and GPU utilization are lagging aggregate signals, and a Pod that looks busy may still be the best destination because it already holds the request's prefix in cache. The scheduler needs a small set of signals that predict user-visible latency, not a dashboard's entire telemetry catalog.
Define the objective before choosing weights. Interactive chat usually cares about time to first token and tail latency. Offline summarization may care more about throughput and cost per generated token. The project documentation supports serving priority so latency-sensitive online traffic can outrank tolerant work, but priority is not free capacity. If every caller marks itself critical, the queue has merely acquired a new label.
Separate pool selection from endpoint selection
The Gateway still performs the familiar first step: HTTPRoute selects an InferencePool or Service. Use that layer for stable policy boundaries such as model family, version, tenant tier, region or safety profile. Once a request targets a pool, the Gateway sends request information to that pool's EPP. The EPP evaluates the pool's endpoints and returns the selected destination; the Gateway remains responsible for forwarding.
This split keeps routing ownership understandable. Platform teams define pools and Gateway policy. Model-serving teams expose the metrics and capabilities required by the model-server protocol. The EPP implementation owns endpoint-scoring logic. Application teams request a documented model name and service class rather than learning Pod identities.
The InferencePool documentation says Pods in a pool share compute configuration, accelerator type, base model and model server. That is an important boundary: do not hide radically different hardware or incompatible serving behavior behind one score unless the implementation explicitly normalizes their capacity. HTTPRoute may reference multiple pools when higher-level traffic splitting or rollout is required.
Score work with a few defensible signals
A practical scorer begins by filtering endpoints that cannot serve the requested model or adapter. It then estimates saturation from queue depth and KV-cache pressure, adds a bonus for a useful prefix match, and applies the service priority. The exact formula is implementation-specific; the official project defines the extension model, not a universal production scoring algorithm.
Keep units explicit. Queue length is not queue time. Cache utilization is not cache hit probability. A prefix-match score is valuable only if the saved prefill work exceeds the cost of waiting behind existing work. Normalize metrics per model-server and accelerator class, cap every bonus, and record the inputs used for the decision. When two candidates are effectively equal, add bounded randomization so all gateways do not herd toward the same replica.
Guard against stale data. The documented request flow allows metrics probing to happen asynchronously. Attach an observation timestamp to every endpoint snapshot, reject or penalize data beyond a short freshness budget, and fall back to a simpler healthy-endpoint policy when signals disappear. A sophisticated score built from stale telemetry is often worse than a modest deterministic fallback.
Design failure behavior before production traffic
InferencePool exposes an Endpoint Picker reference and a failure mode. In the official example, FailOpen lets the Gateway forward to an endpoint it chooses when the EPP is unavailable or cannot select one. That preserves availability, but it also removes the inference-aware guarantees that may have protected latency and accelerator utilization. FailClose protects a strict routing or isolation requirement at the cost of rejecting the request.
Choose by consequence. Fail open may be appropriate for a general assistant where degraded latency is better than an outage. Fail close may be safer when the pool boundary encodes a mandatory model, adapter or locality constraint that an arbitrary endpoint cannot satisfy. Test an EPP timeout, malformed response, zero eligible endpoints, stale metrics and a partial pool outage. The fallback must be observable as its own state, not hidden inside ordinary success metrics.
The project's current FAQ is another operational constraint: the project now provides a lightweight reference extension for conformance, not a default production controller. Gateway implementations may provide their own extension or integrate an existing router. Conformance to the API does not prove that a particular scoring policy, metric pipeline or overload behavior fits your workload.
Trace the complete request path
Instrument five moments: arrival at the Gateway, pool selection, EPP decision start, selected endpoint and first token. Add a stable request ID across the Gateway, EPP and model server. Record the requested model, pool, candidate count, selection reason, signal ages, failure mode used and whether the endpoint accepted or rejected the request. Avoid logging prompt text by default.
The diagram below is the control loop to test. The Gateway chooses a pool through HTTPRoute. The EPP receives request context and reads recent endpoint signals. It returns one endpoint. The Gateway forwards, while the model servers continue exporting queue, cache and capability state. No single metric should silently turn into routing authority.
Measure time to first token, inter-token latency, end-to-end latency and errors by model and service class. Pair them with queue wait, cache reuse, tokens processed, EPP decision latency, stale-signal share, fail-open rate and endpoint-selection distribution. Average latency can improve while the slowest users get worse, so make p95 and p99 visible and separate admission delay from generation time.
Roll out as a routing experiment
Begin with shadow decisions. Let the existing balancer serve traffic while the EPP computes and logs the endpoint it would have chosen. Compare predicted queue time, actual first-token latency and cache reuse without changing requests. A difference between shadow and serving endpoints is useful evidence; it is not automatically proof that the EPP was right.
Next, route a small cohort and keep a control group on the old policy. Use the same model versions, generation settings and traffic class. Define rollback thresholds before the test: error rate, EPP timeout rate, time-to-first-token regression, tail-latency regression and accelerator saturation. Traffic splitting by model name can support incremental model rollouts, but do not change the model and the routing algorithm in the same experiment if you want an interpretable result.
Capacity still matters. Inference-aware routing can place work better; it cannot create GPU memory or decode throughput. Pair routing with admission control, bounded queues and autoscaling signals that reflect pending work. Protect online traffic from batch starvation, but reserve a measurable share for background work so priority does not become permanent exclusion.
Know when a Service is enough
Use a plain Service when requests are similarly sized, replicas are genuinely interchangeable, cache locality is irrelevant and operational simplicity matters more than marginal utilization. Use InferencePool and an EPP when model identity, adapters, queue state or reusable prefixes materially affect latency or cost, and when the team can operate the extra decision plane.
The engineering rule is not “smart routing everywhere.” It is “route with the smallest model of reality that changes the outcome.” Start with capability filtering and fresh queue signals. Add prefix awareness only after measuring meaningful reuse. Add priorities only with explicit budgets. The value of inference-aware routing is not that it sounds intelligent; it is that each decision can be explained, observed and rolled back.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
- Kubernetes Gateway API Inference Extension — Introduction, verified 29 September 2026
- Kubernetes Gateway API Inference Extension — InferencePool, verified 29 September 2026
- Kubernetes Gateway API Inference Extension — API overview, verified 29 September 2026
- Kubernetes Gateway API Inference Extension — v1 API reference, verified 29 September 2026
- Kubernetes Gateway API Inference Extension — FAQ and implementation boundaries, verified 29 September 2026
- Kubernetes Blog — Introducing Gateway API Inference Extension
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




