← Back to Blog
News

Uber ServiceScale Architecture Diagram: Multi-Writer Kubernetes Scaling Intent

In September 2026, Uber Engineering published how its Container Platform evolved from a single horizontal scaling path to a multi-orchestrator model — and InfoQ summarized the same story days later. The architectural move is ServiceScale: a CRD plus Service Scale Controller (SSC) that separates scaling intent from execution so Up / Uber Deployment Controller (UDC) and a failover orchestrator can both influence the same Kubernetes workloads without overloading the deploy hot path. This Uber ServiceScale architecture diagram maps that shape for platform reviews. Treat it as news / engineering-blog coverage of Uber’s internal platform — not a public CRD you can install tomorrow — and do not invent field names beyond what the sources show.

Uber ServiceScale Architecture Diagram: Multi-Writer Kubernetes Scaling Intent
Steady-state orchestrator (Up / UDC) + failover orchestrator → ServiceScale CRs → SSC reconciler → Kubernetes objects / workloads. Callout: scaling intent vs execution.

What an Uber ServiceScale architecture diagram shows

Draw two writers into one intent layer, then one executor into Kubernetes — not a single controller that both decides and mutates replica counts. On the left: a steady-state orchestrator path (Uber’s Up federation layer feeding UDC / UberDeployment lifecycle) expressing the normal desired scale for services. Beside it: a failover orchestrator that needs to influence scale during regional failover — scale down low-tier work, scale up high-tier work — without rewriting UDC. Both write into ServiceScale custom resources. The SSC reconciles combined intent into underlying Kubernetes objects (Deployments, OpenKruise CloneSets, and related workload primitives in Uber’s fleet).

Stamp an explicit callout: intent vs execution. Intent lives in ServiceScale (inspectable with ordinary Kubernetes tooling during an incident). Execution is SSC’s job: materialize the combined desire into workload objects with concurrency controls. For a broader Kubernetes cluster shape primer, see the optional companion Kubernetes cluster architecture diagram guide.

Scaling intent vs execution

For years, Uber’s stateless scaling story was effectively one path: service owners used Up to deploy builds and set scaling expectations; UDC reconciled UberDeployment intent into Kubernetes primitives and reported status back so Up could advance lifecycle workflows (including continuous deployment). That model kept ownership simple — until regional failover needed a second source of scaling desire.

Intent is what each orchestrator wants the scale to be (steady-state vs temporary failover adjustments). Execution is who is allowed to translate that desire into Kubernetes object mutations. Uber’s bet, as Egor Grishechko and Srikar Paruchuru describe on the Uber Engineering blog (Sep 9, 2026), was to keep intent multi-writer and execution centralized in SSC — without standing up an external coordination database that would be harder to debug under incident pressure. Materializing intent in the cluster made “which orchestrator wanted what?” a kubectl-shaped question, and made failback converge by restoring prior intent already on the object rather than replaying forgotten API calls.

ServiceScale CRDs and the SSC reconciler

Conceptually, ServiceScale sits at the level of UberDeployment: a dedicated place to make scale decisions explicit. One orchestrator writes steady-state desired scale; another writes failover-related adjustments; SSC reconciles the combined intent into Kubernetes objects. Uber deliberately kept the model simple — no extra external database, no separate coordination service, no control plane that becomes opaque when paging on-call.

Do not invent CRD field names, status enums, or merge algorithms beyond what the published figures and prose show. Uber’s post includes a ServiceScale CRD specification figure and a controller architecture figure; treat those as the source of truth for labels on your diagram. InfoQ’s Sep 28, 2026 write-up restates the same separation: each orchestrator expresses desire through ServiceScale; SSC reconciles into Kubernetes objects. On the diagram, a single arrow from ServiceScale CRs into the SSC box, then out to “Kubernetes objects / workloads,” is enough — resist drawing speculative webhook chains or invented priority stacks.

Multi-writer orchestrators: steady-state and failover

Why not extend UDC? UDC already sat on the hot path for deploys, scaling changes, and day-to-day lifecycle. Failover is rare by design. Adding failover-specific behavior there would raise complexity on the controller that powers the most common and most critical workflows. As Uber’s authors put it, a regression in failover handling would not stay isolated to failover — it could affect normal deployments across the fleet. They also expected more scaling orchestrators over time (for example future HPA / hybrid autoscaling actors). Dumping every new writer into UDC would grow blast radius; a dedicated ServiceScale path creates a cleaner ownership model.

Production reality was harsher than the CRD sketch. When UDC and SSC both updated related Kubernetes resources, near-simultaneous writes plus stale informer caches produced failure modes optimistic concurrency and server-side apply did not fully prevent. Under specific timing, ReplicaSet metadata drifted from spec — breaking proportional scaling for rolling updates and occasionally sticking workloads until healed. Uber’s layered response: fleet-wide observability for metadata–spec drift, an automated healer in UDC to patch affected ReplicaSets, and a longer-term fix in the scaling path. Separately, Up treated certain status fields as terminal workflow gates; stale cache reads were expensive, so controllers attached generation annotations and verified read-your-own-write consistency before reporting status. Put a small “concurrency / stale-cache guardrails” callout on the SSC / UDC interaction — multi-orchestrator systems are hard because of everything that happens between writes, not because of the APIs.

Regional failover capacity reuse

Uber runs active-active data centers across regions. When a region degrades, traffic may fail over to a surviving region that needs enough capacity for the surge. Historically, the simple safety net was reserved idle capacity everywhere — workable, but expensive underutilization. The new schema: reuse capacity from low-tier workloads to allocate high-tier ones during failover — scale the low tier down, scale the high tier up. That creates a second writer of scaling intent while Up/UDC still own normal desired state.

InfoQ notes Uber’s Container Platform manages over 100 compute clusters across data centers and cloud providers (including Oracle and Google), with roughly 4,000 services on about 3 million cores and 1.5 million daily pod launches — fleet context for why idle 2× regional headroom hurts. An academic paper on Uber’s Unified Failover Architecture (cited via InfoQ; arXiv Jan 2026) reports reducing steady-state provisioning from 2× to 1.3× and eliminating over one million CPU cores.

What not to invent from marketing copy

Is (sourced): ServiceScale CRD + SSC separating scaling intent from execution; multi-writer steady-state (Up/UDC) and failover orchestrators; intent materialized in Kubernetes for inspectability and simpler failback; deliberate choice not to overload UDC; production lessons on stale informer caches, read-your-own-write generation annotations, ReplicaSet metadata–spec drift, healer + root-cause fix; year-long rollout with staging, canaries, and kind-based integration tests supporting native Deployments and OpenKruise CloneSets; fleet-scale context and Unified Failover Architecture provisioning claims as reported by Uber/InfoQ.

Is not: a public open-source CRD install path; invented OpenAPI field lists or merge-priority algorithms; a claim that Kubernetes multi-writer races are “solved” by ServiceScale alone; dollar ROI figures beyond sourced provisioning claims; or conflating ServiceScale with unrelated scaling products. Mark the piece news / engineering blog. Uber’s own takeaway stands: multi-orchestrator systems are not hard because of the APIs — they are hard because of everything that happens between writes.

FAQ: ServiceScale on Kubernetes

What is Uber ServiceScale?

A CRD + Service Scale Controller (SSC) that lets multiple orchestrators each write their own scaling desire, while SSC reconciles the combined intent into Kubernetes objects. It separates scaling intent from execution so Up/UDC can keep owning steady-state lifecycle without absorbing failover logic.

Why not put failover scaling inside Uber Deployment Controller (UDC)?

UDC already sits on the hot path for deploys and day-to-day lifecycle. Failover is rare by design. Folding failover into UDC would raise complexity on the most critical workflows — a regression would not stay isolated to failover. A dedicated ServiceScale path keeps ownership clean and leaves room for future orchestrators.

How does ServiceScale help regional failover capacity reuse?

Instead of carrying large reserved idle capacity in every region, Uber wanted to scale down low-tier workloads and scale up high-tier ones during failover. ServiceScale lets a failover orchestrator express those temporary adjustments alongside steady-state desires from Up/UDC. Failback is simpler because both steady-state and temporary failover information live in the CRD spec.

Conclusion

Keep the diagram honest to the Sep 2026 engineering posture: steady-state orchestrator + failover orchestrator → ServiceScale CRs → SSC reconciler → Kubernetes objects / workloads, with an intent vs execution callout and a small concurrency-guardrails note on the multi-writer path. Cite Uber Engineering — Evolving Uber’s compute platform (Sep 9, 2026) and InfoQ’s Sep 28 coverage; mark CRD details as source-bounded. Browse more architecture diagrams on the ByteDiagram blog.

Diagram multi-orchestrator Kubernetes scaling

Map steady-state and failover writers into ServiceScale CRs, SSC execution into Kubernetes workloads, and an intent-vs-execution callout in ByteDiagram before your next platform review.

Open Diagram Editor