← Back to Blog
DevOps

Hybrid Cloud Orchestration Architecture Diagram: EKS Anywhere Step Functions & Redfish

On September 1, 2026, the AWS Architecture Blog published the first part of a hybrid orchestration pattern: AWS serverless services coordinate on-premises servers and Amazon EKS Anywhere clusters. This hybrid cloud orchestration EKS Anywhere architecture diagram is that pattern for one on-prem site. It is not the later implementation post, and it is not a measured rollout. October 3, 2026, the date on this page, is the publish date. September 1, 2026 is the source date. The control plane in AWS is Amazon API Gateway, AWS Lambda, AWS Step Functions, Amazon EventBridge, and Amazon DynamoDB. At the site, Redfish talks to bare metal, and EKS Anywhere runs a management cluster and a workload cluster. A hybrid link joins the Region to that one site. Facts follow only that article.

Hybrid Cloud Orchestration Architecture Diagram: EKS Anywhere Step Functions & Redfish
API Gateway and Lambda start Step Functions. EventBridge and DynamoDB sit in the AWS control plane. A hybrid link reaches one on-prem site: Redfish servers, an EKS Anywhere management cluster, and an EKS Anywhere workload cluster. ADOT to Amazon Managed Prometheus is a dashed observability side path, not a callback.

What this hybrid cloud orchestration EKS Anywhere architecture diagram shows

The headline on the art is Step Functions orchestrating Redfish and EKS Anywhere for one site. In the Region the boxes are API Gateway for operator requests, Lambda to validate the request, DynamoDB for orders and inventory, EventBridge to route order events, and Step Functions for lifecycle workflows. The edge out of that control plane is labeled Hybrid link. The site frame reads one on-prem site: Redfish servers as the bare-metal hardware API, an EKS Anywhere management cluster that runs cluster lifecycle, and an EKS Anywhere workload cluster that hosts applications. ADOT, the AWS Distro for OpenTelemetry, sits on the clusters and sends metrics to Amazon Managed Prometheus. That arrow is dashed and labeled as a metrics side path. It is not a resume signal into Step Functions. Part 1 of the source post is the pattern. Infrastructure-as-code and runbooks are deferred to a later part.

The problem this architecture is solving

The authors describe operations that drift: different hardware, different cluster versions, different add-ons, manual BIOS and firmware work, and no single place to ask which servers are behind. Tools built for one data center, they say, do not stay reliable as the machine count grows. Some environments also cannot put the Kubernetes control plane in AWS, because of data sovereignty, regulation, or a network that is disconnected, disrupted, intermittent, or limited. EKS Anywhere is the choice in the post when the whole cluster, including the control plane, must stay on your hardware. If the site has reliable connectivity to a Region and can use a managed control plane, the same post says Amazon EKS Hybrid Nodes is the recommended approach instead. Hybrid Nodes are a different placement. They are not a box in the EKS Anywhere flow this diagram draws.

Main components and trust boundaries

An order is the unit of work. An API call such as terminating a cluster creates a DynamoDB order and returns an id immediately. EventBridge routes the operation to a Step Functions workflow. The workflow can pause on a task token until the on-premises step, which may take hours, reports that it finished. That pause is state inside the workflow. This diagram does not draw it as an edge. A Distributed Map state is how one server operation fans out across the servers at this site. A new order is denied when another order is already running on the same resource.

Redfish, from the DMTF, is the hardware boundary: firmware, power, health, and BIOS templates without a per-vendor API. EKS Anywhere is the cluster boundary. The management cluster manages the workload cluster, and both are rows in the inventory. A new cluster takes a blueprint plus a hardware CSV, then the EKS Anywhere CLI runs through Systems Manager and AWS Batch. IAM Roles Anywhere issues short-lived credentials from a certificate. Systems Manager hybrid activations register on-premises nodes. Private CA and cert-manager handle certificates. The September 1, 2026 post describes the network between the Region and the hardware. This diagram labels that path Hybrid link. It does not draw Direct Connect or a VPN as a box.

Request or data path, step by step

An operator calls the API. Lambda writes an order, EventBridge starts the state machine, and the id returns immediately. Hardware steps call Redfish for power, firmware, or a BIOS template. Cluster steps select machines, wait while EKS Anywhere comes up, then install add-ons. DynamoDB tracks status. Webhooks can notify an outside system. The post also names DynamoDB Streams for follow-on DNS updates and Portworx as an example of external storage, not a required product.

Metrics are a separate side path. An ADOT collector on the clusters sends server, Kubernetes, and application metrics to Amazon Managed Service for Prometheus. The post names Redfish events or node-exporter, and kube-state-metrics. It points at a related write-up for collector detail. This article does not add that second URL. The dashed edge is only that observability side path. It does not re-enter Step Functions, and it is not how a paused workflow learns that the site finished.

The diagram: labeled boxes and failure or isolation edges

Boxes: API Gateway, Lambda, Step Functions, EventBridge, DynamoDB for orders and inventory, Hybrid link, Redfish servers, an EKS Anywhere management cluster, an EKS Anywhere workload cluster, ADOT, and Amazon Managed Prometheus. The site box says one on-prem site. The order arrows are request, write order, event, workflow, and state. The only dashed edge is ADOT to Amazon Managed Prometheus, the metrics side path. Isolation edges: the Kubernetes control plane stays on the hardware in this pattern; AWS credentials on the cluster are short-lived through IAM Roles Anywhere; a conflicting order stops at the inventory check. Failure edges: the link can be limited or down, which is why the post keeps the cluster local; a firmware or cluster job can run for hours, which is why the workflow waits on its own state instead of holding the API call.

What the source does not claim

This is a September 1, 2026 thought-leadership post, Part 1 of a series. It is not a case study, not a preview feature announcement, and not the infrastructure-as-code installment. It does not publish a latency, an error rate, or a count of servers actually managed. EKS Anywhere cluster lifecycle remains your responsibility. The orchestration is what the pattern automates. Hybrid Nodes stay the source’s other choice when the control plane can live in the Region. They are outside the EKS Anywhere flow on this page.

How this differs from a nearby pattern on ByteDiagram

A single-cluster picture such as the Kubernetes cluster architecture diagram is only a primer for nodes, control plane, and workloads. This page is one on-prem site. The Kubernetes control plane for EKS Anywhere stays on the hardware at that site, and AWS holds orders, workflows, and inventory. The same AWS article says to use EKS Hybrid Nodes instead when the site can rely on a managed control plane in the Region. Those are two placements. This diagram shows only the EKS Anywhere placement, with Redfish on the servers beside the clusters.

FAQ

When does the post choose EKS Anywhere instead of EKS Hybrid Nodes?

EKS Anywhere when the whole Kubernetes cluster, including the control plane, must stay on your hardware. The September 1, 2026 post names data sovereignty, regulation, or a disconnected, disrupted, intermittent, or limited network. It says EKS Hybrid Nodes is the recommended approach when on-premises nodes can use a managed control plane in an AWS Region over reliable connectivity. Hybrid Nodes are not part of the EKS Anywhere flow in this diagram.

What is an order in this design?

An order is a tracked lifecycle operation. The API writes it to DynamoDB and returns an id while Step Functions runs the work. EventBridge routes the API operation to the workflow. The workflow can pause until the site finishes a long hardware or cluster step. A second order on the same resource is denied while one is already running.

Is this post the implementation guide?

No. It is Part 1, a thought-leadership description of the pattern. The authors defer infrastructure-as-code, workflow definitions, and runbooks to a later part. It is not a case study and it does not publish a measured server count, latency, or error rate.

Conclusion

The September 1, 2026 pattern, drawn here for one on-prem site, is API Gateway and Lambda starting Step Functions, with EventBridge and DynamoDB in the AWS control plane. A hybrid link reaches Redfish servers and the EKS Anywhere management and workload clusters. ADOT to Amazon Managed Prometheus is a dashed observability side path, not a callback into the workflow. The source is the AWS Architecture Blog post. More diagrams are on the ByteDiagram blog.

Diagram hybrid orchestration for one site

Map orders in Step Functions, Redfish on the servers, and the EKS Anywhere clusters at one on-prem site in ByteDiagram.

Open Diagram Editor