← Back to Blog
News

Meta ZGateway Architecture Diagram: ZippyDB Proxy at 1B+ Ops/s

On September 3, 2026, Meta Engineering published ZGateway — the proxy tier unifying traffic through ZippyDB, Meta’s most widely used distributed key-value store. InfoQ covered the architecture on September 28. ZGateway already handles more than 1 billion operations per second and about 40% of ZippyDB traffic (heading past 60%), while Meta’s model estimates roughly a 19× drop in total persistent connections — at about 6% computational overhead for a typical use case.

This guide shows what to draw on a Meta ZGateway architecture diagram: left, the failed direct-access mesh; right, clients with sticky regional pools → ZGateway (TLS, ACL, Discriminant Load Shedding, shard resolve, cache/batch) → ZServer replicas. For L4/L7 front doors in general, keep the load balancer patterns guide nearby; for edge caches, see the CDN architecture post. This article is about connection economics and shared traffic management in front of a hyperscale KV store — not GPU inference routing.

Meta ZGateway ZippyDB proxy architecture diagram showing direct many-to-many client mesh versus regional ZGateway path with auth, Discriminant Load Shedding, cache batching, and ZServer replicas
Left: unbounded client↔ZServer mesh. Right: sticky regional ZGateway hops that bound fan-in and centralize admission, batching, and failover.

Why ZippyDB needed a proxy

ZippyDB backs product metadata, counters, configuration, and other high-QPS workloads across Meta’s globally distributed fleet. In the direct-access model every client connected to every database host it needed. A single client can touch tens of thousands of shards; those shards sit across hundreds of thousands of database hosts. The result is a dense many-to-many TLS mesh: typical clients hold tens of thousands of outbound connections, and typical database hosts accept tens of thousands of inbound ones.

That mesh is wasteful and fragile. Idle connections still burn memory, CPU, and file descriptors on both ends. Inbound fan-in grows with the client population — Meta cites more than one million client hosts — so every new cohort makes every database host worse. Reconnection storms after deploys or routing bugs have caused file-descriptor exhaustion and OOM reboot loops. Client-side pooling cannot fix this alone: client fleets and the database fleet move on different schedules.

Diagram rule: draw direct access as a dashed red mesh (clients ↔ many ZServers). Draw ZGateway as two bounded hops. The visual contrast teaches the architecture faster than the connection math.

Architecture layers to draw

LayerWhat it isDiagram tip
ClientsApps, services, jobs (>1M hosts)Left of main path: sticky pool to regional gateway
ServiceRouterHyperscale discovery / meshLabel on the client→gateway edge
ZGateway hostStateless regional proxyAmber card: auth, DLS, shard, cache/batch
Tier flavorsPure proxy + read-through cacheFootnote under the gateway card
ZServer replicasSharded ZippyDB storageTeal card: fan-in only from gateway fleet
Direct contrastMany-to-many meshDashed red box on the left

What ZGateway is (and is not)

ZGateway is a stateless proxy tier between ZippyDB clients and the ZServer fleet. It runs as regional tiers discovered through ServiceRouter so clients stay near their gateway. Internally it runs Meta’s thick C++ ZippyDB client as the request engine — “a ZippyDB client run as a managed service.” Two flavors share one pipeline: a pure proxy and a read-through cache.

Responsibilities that stay put: TLS in the Thrift/ServiceRouter stack, key-to-shard mapping in the shard locator, replica selection and hedging in the embedded client. ZGateway owns traffic management — authorization, admission, batching, caching, balancing, cross-region failover — not a reimplementation of the database.

DimensionDirect accessVia ZGateway
Client connectionsTens of thousands per client to shardsSticky pool to regional gateway hosts
ZServer fan-inLinear in client populationBounded by gateway fleet × shard density
Shared workReimplemented in every client binarySolved once in the managed tier
Storm containmentHits every database hostContained at a fleet Meta controls
Overhead (typical)n/a~6% compute (Meta)

Request path to sketch

  1. Client sends over a sticky connection to a regional ZGateway host.
  2. Terminate + authorize — TLS ends; request checked against the use case’s ACLs.
  3. Admission control — per-tenant Discriminant Load Shedding (DLS) buckets plus AIMD concurrency control.
  4. Shard resolve — map key to physical shard / replicas.
  5. Cache (optional) — on cache tiers, check local cache; miss takes a per-key fill lock.
  6. Batch / coalesce — merge same-shard work across unrelated callers; flush on linger, size, or count.
  7. Forward to the correct ZServer replicas; demultiplex responses; emit per-use-case metrics and quota usage.

Connection math worth annotating

Meta models the fleet as balls-into-bins (shards a host touches into hosts). With round figures — 20 regions, 500,000 database hosts, 30,000 proxy hosts, 1,000,000 clients, 50,000 shards per client — their model yields roughly 97–98% fewer connections per host, and about 19× fewer persistent connections end-to-end because each backend connection multiplexes many clients. Label these as model-based estimates, not measured SLOs from your own stack.

The deeper point is scaling behavior: under direct access, database fan-in grows with every new client cohort. With ZGateway, client population drops out of the fan-in formula — what remains is roughly regions × shard density per host, levers the storage team owns.

Capabilities the shared vantage point unlocks

  • Cross-client batching and coalescing — a client library can only merge its own process; the gateway collapses hot-key stampedes into one backend read and amortizes Thrift/syscall overhead.
  • Discriminant Load Shedding — every request maps to a per-tenant priority bucket; noisy neighbors fill their own bucket while others keep draining. In a controlled overload above 90% CPU across ~1,350 buckets, only six shed; the rest executed 99.9% of requests with zero rejections.
  • Read caching with CDC invalidation — in-process cache, per-key fill lock, change-data-capture for invalidate/refill under a bounded-staleness contract, keyspace sliced by consistent hashing.
  • Weighted load balancing — control plane nudges ServiceRouter weights from recent CPU so mixed ~26–126 core hosts do not create hot outliers.
  • Cross-region resilience — global routing tables, mega-regions, and backup rings with percentage knobs; failover keyed off sharp regional signals, not a bland CPU average.
  • Safe migration — client-side config flags by service and shard prefix: percentage ramp, region filter, global kill switch — no client code change for rollback.

When to diagram a ZGateway-style tier

Prefer a shared proxy tierKeep direct / client-side pooling
Huge, diverse client fleets you cannot roll out quicklySmall fleet where mesh cost is still cheap
Connection storms or FD/OOM incidents on storage hostsStorage already sits behind a mature pooler/mesh you trust
Need cross-tenant batching, admission, or progressive cutoverStrict single-hop latency budget with no shared work to centralize
Want one control plane for LB, cache, and failoverRegulatory requirement for client→DB path with no intermediary

Pitfalls to annotate

  • Extra hop is deliberate — Meta argues the hop can improve overall latency by freeing database nodes from connection overhead; still measure p99 for your workload.
  • Batching memory risk — idle eviction + in-flight caps are mandatory safety valves; omit them from a production diagram and you understate ops cost.
  • Not a generic API gateway — compare thoughtfully to Kong/Envoy/MCP gateways; ZGateway’s engine is ZippyDB’s own thick client.
  • Cache staleness — read-through tiers use an explicit bounded-staleness contract via CDC; do not draw “always consistent” boxes.
  • Model vs production — 19× / 97–98% figures are Meta’s model estimates; cite the engineering post when you reuse them.

Example prompt for AI diagram generation

Drop this into ByteDiagram for a first draft that matches the hero figure:

Left dashed red: Direct ZippyDB access — 1M clients mesh to ZServer (FD/OOM risk)
Right purple shell: ZGateway path — Clients sticky pool → Regional ZGateway
  (TLS+ACL, DLS admission, shard resolve, cache/batch) → ZServer replicas
Bottom chips: % migration kill switch, cross-region rings, weighted LB, CDC invalidation
Annotate 1B+ ops/s, ~19x fewer connections, ~6% overhead

FAQ

Is ZGateway only for Meta’s ZippyDB?

The product is Meta-internal, but the pattern — put a managed proxy in front of a shared datastore when clients outnumber and outpace the storage team — applies to Redis/Valkey fleets, sharded SQL proxies, and connection poolers. Draw the same two-hop contrast for your own stack.

How is this different from a service mesh sidecar?

Sidecars usually sit next to each client. ZGateway is a shared regional tier that sees many clients at once, which is what enables cross-client coalescing and discriminant shedding. Meshes and ZGateway-style proxies can coexist; do not draw them as the same box.

Does the proxy replace ZippyDB’s own sharding?

No. ZippyDB still owns managed sharding, replication, failure detection, and capacity. ZGateway adds the shared traffic-management layer in front.

Conclusion

ZGateway is Meta’s September 2026 answer to ZippyDB connection sprawl: keep the thick client’s smarts, run them as a managed regional proxy, and turn an unbounded many-to-many mesh into two bounded hops with admission control, batching, caching, and progressive cutover. A clear Meta ZGateway architecture diagram shows direct-access contrast → clients → regional ZGateway capabilities → ZServer replicas → migration and resilience footnotes. Draw that once for platform reviews, then reuse it when arguing for a shared proxy in front of any hyperscale KV or cache fleet. For disaggregated database compute/storage rather than connection proxies, see the MongoDB Atlas Infinite guide.

Diagram your ZippyDB-style proxy path

Generate a direct-mesh vs regional-gateway contrast in ByteDiagram, then annotate admission, batching, and failover for architecture reviews.

Open Diagram Editor