On September 24, 2026, AWS announced Amazon SageMaker HyperPod Inference Gateway: a Kubernetes-native, GPU-aware routing layer that installs as one EKS managed add-on on existing HyperPod clusters. Instead of blind round-robin across model pods, it places each OpenAI-compatible request using live inference signals — and AWS reports large cuts in time-to-first-token (TTFT) on mixed hardware and bursty traffic.
This guide shows what to draw on a SageMaker HyperPod Inference Gateway architecture diagram: clients → Envoy Endpoint → Body-Based Router → Endpoint Picker → labeled GPU pools (vLLM / SGLang) → observability and failure edges. For generic L4/L7 front doors, keep a separate load balancer architecture; this post is about LLM-aware routing on HyperPod, not classic ALB patterns. Cluster primitives still map to a Kubernetes cluster architecture underneath.
Why naive Kubernetes routing wastes GPUs
Default Service load balancing (round-robin, least connections) has no view inside GPU memory. It cannot see saturated KV caches, deep generation queues, or which pod already holds the LoRA adapter your request needs. Requests stack behind busy pods while peers sit underutilized; first-token latency spikes; teams over-provision H100/A10G capacity to compensate.
HyperPod Inference Gateway replaces that with real-time scoring on every model-serving pod. AWS’s what’s-new post cites up to 82% lower first-token latency and p99 TTFT reductions of 97–98% in mixed-hardware and burst scenarios versus unintelligent round-robin — the exact conditions production fleets usually face.
Architecture layers to draw
| Layer | What it is | Diagram tip |
|---|---|---|
| Clients | Apps, agents, batch jobs | Left: OpenAI /v1/chat/completions |
| Envoy Endpoint | HTTPS terminate; one private URL per cluster | Amber card inside blue Gateway shell |
| Body-Based Router | Reads model from JSON body → GPU pool | Purple card; multi-model, one endpoint |
| Endpoint Picker | Weighted score across 6 signals | Teal card; arrow to chosen pod |
| Model pools | Labeled pods (vLLM, SGLang, …) | Right: per-model pools on HyperPod/EKS |
| Signals | Prometheus metrics into EPP | Footnote band under pools |
| Tier 2 (soon) | Global Inference Router | Dashed “coming soon” box — do not claim GA |
Three core components (Tier 1)
Tier 1 ships inside the amazon-sagemaker-hyperpod-inference EKS add-on (Inference Gateway available from add-on v2.0.0-class builds). All three pieces build on the open-source Gateway API Inference Extension:
- Envoy Endpoint — high-performance L7 proxy that terminates HTTPS and exposes a single private endpoint for the cluster. Clients keep one base URL.
- Body-Based Router (BBR) — inspects the OpenAI-compatible request body, extracts the model name, and routes to the correct model pool. One gateway serves many models with no client-side routing logic.
- Endpoint Picker (EPP) — scores every candidate pod and picks the best backend for that request. Scorers are weighted so you can bias latency-sensitive chat vs throughput batch.
Diagram rule: draw BBR before EPP. Model identity selects the pool; GPU signals select the pod inside that pool.
Endpoint Picker: the six signals
Label these on the diagram (or in a callout) so reviewers see why this is not “smart least-conn”:
- KV cache utilization — avoid pods whose key-value memory is nearly full.
- Queue depth — avoid deep request backlogs.
- LoRA adapter residency — prefer pods that already loaded the requested adapter (cuts swap latency).
- Prefix cache hit rate — prefer pods likely to reuse shared prompt prefixes (multi-turn chat, document Q&A).
- Predicted latency — steer away from pods expected to be slow for this request.
- Running requests — balance active work across the fleet.
On a uniform fleet under steady traffic, AWS notes the gateway can perform on par with round-robin — gains concentrate when hardware mixes, traffic bursts, or prefixes are shared. Annotate that honesty on the figure so design reviews do not over-claim.
Deploy path: add-on + InferenceGatewayConfig
- Enable the gateway in the HyperPod inference add-on (
inferenceGateway.enabled: truealongside the Inference Operator as needed). - Label model server Deployments so the gateway can discover them (for example
app: vllm-llama). - Apply an
InferenceGatewayConfigCRD: TLS options, BBR enabled, and per-model schedulers withmodelName, label selectors, and target port. - Send traffic to the gateway’s OpenAI-compatible URL — existing clients and curl against
/v1/chat/completionswork without SDK or SigV4 changes for inference traffic.
No sidecars, no service mesh, and no application code changes are required for the Tier 1 path. GitOps (kubectl, Helm, Argo CD) continues to own the CRD.
When to use it (and when not to)
| Use HyperPod Inference Gateway | Prefer something else |
|---|---|
| Multi-model GPU fleets on HyperPod/EKS with OpenAI-compatible servers | Single tiny model, one replica, steady load (gains may be small) |
| Mixed GPU generations or bursty chat/agent traffic | Non-OpenAI protocols you have not validated through BBR |
| Shared LoRA bases or heavy prefix reuse | Teams that only need classic L7 LB without LLM signals — use a standard load balancer diagram |
| Want kubectl/CRD lifecycle, not a separate mesh | Need GA cross-region failover today — Tier 2 Global Inference Router is still coming soon |
Pitfalls to label on the diagram
- Tier 2 is not GA — cross-cluster / cross-region routing, global rate limits, and cost-tier shaping are announced as coming soon. Keep a dashed GIR box; do not draw them as live today.
- Add-on version — gateway needs a v2.x HyperPod inference add-on that recognizes
InferenceGatewayConfig; older clusters will not show the CRD. - Pool exhaustion — gateway can return HTTP 429 with Retry-After; pair the diagram with HPA / capacity notes.
- Stale metrics — EPP excludes pods with stale Prometheus data; show a “healthy metrics” dependency edge.
- Not a mesh replacement for all traffic — this is inference-aware routing for model servers, not a general-purpose service mesh for every microservice.
Observability and graceful failure
Draw metrics flowing up from pods (KV, queue, running requests, adapter residency) into pool-level histograms and cluster-level CloudWatch views (average KV, error rate, P99). Failure table worth a footnote: pod failure → exclude and route elsewhere; pool full → 429; cluster/region failure stories belong to Tier 2 GIR once available.
Example prompt for AI diagram generation
Drop this into ByteDiagram for a first draft that matches the hero figure:
Left: Clients/apps sending OpenAI /v1 chat completions
Center blue shell: Inference Gateway — Envoy Endpoint, Body-Based Router, Endpoint Picker
Right green: HyperPod/EKS GPU pools (llama-70b, qwen-32b) with vLLM/SGLang pods
Bottom: 6 EPP signals + TTFT gains vs round-robin + dashed Tier 2 GIR coming soon
Left-to-right flow, amber Envoy, purple BBR, teal EPP.
FAQ
Does it work only with SageMaker-hosted runtimes?
The gateway targets OpenAI-compatible model servers on HyperPod/EKS — AWS explicitly calls out vLLM and SGLang. You still run on HyperPod infrastructure with the official inference add-on.
Is this the same as an Application Load Balancer?
No. ALB/NLB (or kube-proxy) can still sit in front for network ingress, but they do not score KV cache or LoRA residency. Put classic LBs outside or below the LLM-aware path; put Envoy + BBR + EPP as the inference intelligence layer.
Can one endpoint serve many models?
Yes — that is the Body-Based Router’s job. Multiple schedulers in InferenceGatewayConfig map model names to labeled pod pools behind one gateway URL.
Conclusion
SageMaker HyperPod Inference Gateway is AWS’s September 2026 answer for production LLM routing on HyperPod: install one EKS add-on, declare InferenceGatewayConfig, and let Envoy, Body-Based Router, and Endpoint Picker place each request using GPU-level signals. A clear architecture diagram shows clients → Tier 1 gateway → multi-model GPU pools, the six EPP scorers, honest Tier 2 “coming soon,” and failure/observability edges. Draw that once for platform docs and capacity reviews, then reuse it when comparing round-robin baselines to GPU-aware routing. For agent-tool reverse proxies rather than GPU pools, see the MCP gateway architecture guide.
Diagram your HyperPod inference path
Generate an Envoy → BBR → Endpoint Picker → GPU pool layout in ByteDiagram, then animate TTFT-sensitive traffic for platform onboarding.
Open Diagram Editor