Skip to main content

Command Palette

Search for a command to run...

Gateway API Inference Extensions: An AI Aware Load Balancer

Updated
•10 min read•View as Markdown
Gateway API Inference Extensions: An AI Aware Load Balancer
K
Hands-on deep dives into AWS Networking, Kubernetes, and AI security. Covering EKS, Cilium, service mesh, MCP, and more. Practical guides for platform engineers and cloud architects.

Introduction

If you’ve been anywhere near the Kubernetes ecosystem lately, you’ve probably heard the buzz around Gateway API — and the ongoing “Ingress is going away” narrative. I’ve covered Gateway API before, and at first glance, it felt like the natural evolution of Kubernetes traffic management.

With a normal web application, requests are usually short-lived and predictable. AI inference is different and unfortunately the techniques we relied on with classic load balancing (Round Robin, weighted routing, least connections) will not suffice for AI Inferencing.

Figure 1 below aims to present the problem clearly. From the outside, both requests are HTTP requests hitting the same API endpoint. The first request is small in terms of input/output tokens. It uses minimal GPU resources, finishes quickly, and can be handled by almost any inference pod. The second request usually happens in RAG use cases involves 1000s of more input tokens and possibly output tokens, cause a higher load on the GPU in terms of KV cache and overall processing and the output will take longer to generate.

The problem in a few words is that traditional load balancing is not aware that this is an AI workload & is not aware of each of the vLLM pods situation. Imagine if the 1st VLLM pod is very low on KV cache, has queued requests that haven't been served and the AI Gateway by virtue of round-robin decides to send the request regarding the 200-pages of financial data.

Gateway API Inference Extensions aims to provide the vLLM Server/GPU context to the Gateways so that we can make smarter routing decisions that take into consideration the state of those workloads.

Figure1. Explaining the Challenge with current LB implementations

Solution: At a Glance

The solution at a high level works as follows:

  • Request from client hits the AI Gateway which has a set of possible pods to forward to called Inference Pool. You can think of it like a Kubernetes Service on steroids in the sense that it will choose an a pod using the selector mechanism but what makes it special is that it will need to get the advice of Extension Pod Placement (EPP).

  • The EPP checks the request and looks for the model and/or LoRA adapters required to narrow down the choice of possible pods, and from those remaining looks into a few metrics that the vLLM server exposes such as the following:

    • KV Cache Locality: I already have all previous context stored in the GPU and thus it will be very easy to process this request

    • Queue Depth: this measures the amount of requests waiting in the queue.

    • Running Requests; How many requests are currently being served by this vLLM Server

    • GPU Memory: Measure how much of my KV cache is actually utliized

    • Time To First Token: This is a good measure of latency

Based on these metrics, the EPP generates an overall score for each of the vLLM Pods and select the Pod with the highest score.

Figure2. AI Inference Gateway Solution Architecture

The End-to-End Request Flow

Figure3. End-to-End Request Flow

In this section we will go through the end-to-end request flow depicted in Figure 3.

  1. Client Request: The client sends a request to the AI Gateway which specifies the model in addition to the message. The model can specify either the base model or the name of a specific LoRA adapter. A LoRA adapter is a lightweight add-on that fine-tunes a base AI model for a specific task.

  2. HTTP Route: In a similar fashion to any Application Load Balancer, the AI Gateway compares the incoming request path or host header to find the best match and sends it to an InferencePool.

  3. InferencePool: You can think of an InferencePool as a Kubernetes Service on steroids whereby it uses selectors to match the set of pods with a particular label, but on top of that it adds AI awareness by referencing an Endpoint Picker Extension (EPP). This is inline with Figure 2 and shows how AI Gateway will consult EPP to choose the best Pod . (Ref: Fig 4)

  4. ModelRewrite: This is an optional configuration whereby the objective might be to hide the exact models used in the backend and give them aliases. This provides the flexibility of swapping one model to another in blue/green deployment fashion.

  5. Endpoint Picker Extension (EPP): This is the component that provides the inference extensions to Gateway API. This component goes through each of the available pods within the Inferencepool to infer the running model, LoRA adapter(s) loaded, number of Queued requests, GPU memory..etc Thus, the gateway forwards enough context about the request (model, number of tokens, message hash..etc) to the EPP and the EPP will return back you can use this pod.

  6. vLLM Pod: This is where the request gets to its destination pod to be processed.

# ── InferencePool ─────────────────────────────────────────────────────────
apiVersion: inference.networking.k8s.io/v1
kind: InferencePool
metadata:
  name: vllm-inference-pool
  namespace: ai-workloads
spec:
  targetPorts:
  - number: 8000               
  appProtocol: http
  selector:
    matchLabels:
      app: vllm-inference-pool  
  endpointPickerRef:
    name: vllm-inference-pool-epp   # EPP service
    port:
      number: 9002              # EPP ext-proc gRPC port
    failureMode: FailOpen  
💡
The Endpoint Picker Extension (EPP) is mapped 1:1 with an inferencePool. Thus you can have two inference pools each with two different EPPs. Each EPP can value different signals or metrics based on the workload

EPP Decision Process

In this section, we get a bit deeper into the EPP to understand more about its decision process.

It starts with filtering the pods that actually serve the model or the LoRA adapter that is being requested and that are healthy. Any pods that don't meet any of the previous criteria will be eliminated.

Apparently vLLM Servers can dynamically load LoRA Adapters. More on this here

Scoring: On scoring this is where we evaluate each endpoint (Pod) metrics such as:

  • Queue Depth: This represents the number of pending requests in the queue with a preference for a lower number.

  • Prefix Cache: This is probably the most important criteria which is pertaining to any of the pods processing previous context. The preference is for a higher caching value. For example, if I send an initial prompt like “Explain Kubernetes Networking”, the model performs significant processing and stores Key/Value attention matrices in the GPU memory as part of the KV cache.

    Later, if I send “What about kube-proxy?”, the model still sees the full conversation context. Thus, if we send this request to the same pod that already did the previous work the whole process will be more efficient and we can expect a faster response.

Explain Kubernetes Networking
What about kube-proxy?
  • KV Cache Utilization: this represents the percentage of used GPU memory. The preference here is for a lower KV Cache utilization

Figure4. EPP Decision Process

This is a very simple example showing two vLLM pods and how the overall score is derived. The winner in our case is vllm-2.

Endpoint Queue (x2) KV Cache (x2) Prefix Cache (x3) Total Score
vllm-1 0.8 1 0 3.6
vllm-2 1 1 1 7.0
💡
The Github link here shows how you can modify the EPP metrics used and their relevant weights for your workload

The below are live logs gathered from EPP where I added the additional context as comments.

# ─────────────────────────────────────────────────────────────
# COMPLETE EPP REQUEST TRACE — Request c05468c5
# Scenario:
# - 2 candidate vLLM pods
# - Both already have ru-lora loaded
# - Non-zero KV cache on both pods
# - Prefix cache hit exists only on pod 5ldmj
# ─────────────────────────────────────────────────────────────
 ─────────────────────────────────────────────────────────────
# STEP 1 — MODEL REWRITE
# Before endpoint selection begins, the incoming alias "qwen2"
# is rewritten into the actual backend model name.
# No pods are evaluated yet at this stage.
# ─────────────────────────────────────────────────────────────

{"msg":"LLM request assembled",
 "incomingModelName":"qwen2",
 "targetModelName":"Qwen/Qwen2-0.5B-Instruct"}


─────────────────────────────────────────────────────────────
# STEP 2 — FILTER PHASE
# The EPP evaluates all candidate inference endpoints.
# Both pods survive filtering.
#
# Key observations:
# - Both pods already have ru-lora loaded
# - Both have empty queues
# - Both have non-zero KV cache utilization
# - Pod 5ldmj is handling significantly more running requests
#
─────────────────────────────────────────────────────────────

{"msg":"Before running filter plugins",
 "pods":[
   {"pod":"x7pqq",
    "kv":0.00735,
    "queue":0,
    "running":5,
    "activeModels":{"ru-lora":0},
    "waitingModels":{"ru-lora":0}},

   {"pod":"5ldmj",
    "kv":0.03238,
    "queue":0,
    "running":74,
    "activeModels":{"ru-lora":0},
    "waitingModels":{"ru-lora":0}}
 ]}

{"msg":"Completed running filter plugins",
 "remainingEndpoints":2}


# ─────────────────────────────────────────────────────────────
# STEP 3A — QUEUE SCORER (weight = 2)
#
# The queue scorer prefers lower waiting queue depth.
# Both pods currently have queue=0.
#
# Result:
# - Both receive the maximum queue score of 1.0
# ─────────────────────────────────────────────────────────────

{"msg":"Calculated score",
 "plugin":"queue-scorer",
 "endpoint":"x7pqq",
 "score":1}

{"msg":"Calculated score",
 "plugin":"queue-scorer",
 "endpoint":"5ldmj",
 "score":1}


# ─────────────────────────────────────────────────────────────
# STEP 3B — KV CACHE UTILIZATION SCORER (weight = 2)
#
# The KV cache scorer evaluates memory/cache pressure.
# Lower KV cache utilization produces a slightly higher score.
#
# x7pqq:
# - KV cache = 0.74%
# - Less GPU memory pressure
#
# 5ldmj:
# - KV cache = 3.24%
# - Slightly more memory pressure
#
# Result:
# - x7pqq receives the better KV score
# ─────────────────────────────────────────────────────────────

{"msg":"Calculated score",
 "plugin":"kv-cache-utilization-scorer",
 "endpoint":"x7pqq",
 "score":0.9927}

{"msg":"Calculated score",
 "plugin":"kv-cache-utilization-scorer",
 "endpoint":"5ldmj",
 "score":0.9676}


# ─────────────────────────────────────────────────────────────
# STEP 3C — PREFIX CACHE SCORER (weight = 3)
#
# This scorer checks whether the pod has previously processed
# similar prompt prefixes.
#
# x7pqq:
# - No matching cached prefixes
#
# 5ldmj:
# - Existing prefix cache hit detected
# - Prompt locality already exists on this pod
#
# Result:
# - 5ldmj receives a full score of 1.0
# - This becomes the dominant scoring factor because
#   the plugin weight is 3
# ─────────────────────────────────────────────────────────────

{"msg":"Calculated score",
 "plugin":"prefix-cache-scorer",
 "endpoint":"x7pqq",
 "score":0}

{"msg":"Calculated score",
 "plugin":"prefix-cache-scorer",
 "endpoint":"5ldmj",
 "score":1}


# ─────────────────────────────────────────────────────────────
# STEP 4 — WEIGHTED TOTAL CALCULATION
#
# Final Score Formula:
#
# (queue-score × 2)
# + (kv-cache-score × 2)
# + (prefix-cache-score × 3)
#
# x7pqq:
# (1×2) + (0.9927×2) + (0×3)
# = 3.9853
#
# 5ldmj:
# (1×2) + (0.9676×2) + (1×3)
# = 6.9352
#
# Even though 5ldmj has:
# - higher KV cache usage
# - more running requests
#
# it wins because:
# - both queues are empty
# - both already have ru-lora loaded
# - ONLY 5ldmj has a prefix cache hit
#
# Prefix locality outweighs KV cache pressure in this case.
# ─────────────────────────────────────────────────────────────

{"msg":"Candidate pods for picking",
 "weighted":[
   {"pod":"x7pqq","score":3.9853},
   {"pod":"5ldmj","score":6.9352}
 ]}


# ─────────────────────────────────────────────────────────────
# STEP 5 — FINAL SELECTION
#
# Pod 5ldmj is selected by the picker plugin.
#
# This demonstrates one of the most important principles
# of inference-aware routing:
#
# The least busy pod is NOT always the best pod.
#
# Cache locality and prompt reuse can be more valuable
# than raw utilization alone.
# ─────────────────────────────────────────────────────────────

{"msg":"SELECTED",
 "pod":"vllm-6685564c67-5ldmj"}

Try It Yourself

If you'd like to follow along or test this in your own environment, I’ve published the manifests and configs here.