← Back to Blog
System Design

ProxySQL Architecture Diagram: Query Rules, Hostgroups & Multiplexing

ProxySQL is a high-performance, protocol-aware proxy for MySQL and PostgreSQL — written in C++, asynchronous, multi-threaded, and event-driven. This evergreen ProxySQL architecture diagram maps the shape operators actually review: clients land on worker threads, a query processor decides route / rewrite / cache / block, hostgroups split read and write traffic across backend pools with multiplexing, and a monitor thread watches lag and health so failover can shrink hostgroup membership without restarting the proxy. Facts below follow the official ProxySQL Architecture documentation — no invented metrics or CRD fields.

ProxySQL Architecture Diagram: Query Rules, Hostgroups & Multiplexing
Clients → ProxySQL workers → query processor (route / rewrite / cache / block) → hostgroups (RW / RO) → MySQL backends; monitor thread dashed to backends for lag and health.

What a ProxySQL architecture diagram shows

Draw a left-to-right data plane and a supporting control plane. On the left: application clients (app servers, ORMs, BI tools) speaking the MySQL or PostgreSQL wire protocol. In the middle: a ProxySQL process whose main thread owns specialized thread pools — MySQL/PgSQL worker threads, an admin thread (MySQL admin on port 6032, PostgreSQL admin on 6132), monitor threads, optional cluster and idle threads. Workers feed every statement into the query processor, which matches query rules and asks the HostGroups manager for a backend connection. On the right: backend databases grouped into hostgroups — typically a write (RW) hostgroup for the primary and a read (RO) hostgroup for replicas. Under the data plane, draw a dashed monitor edge to those backends: pings, connection tests, replication lag, and read_only tracking. Core principles from the docs to stamp on the art: maximal uptime (runtime config, no restarts), event-driven scalability via libev, protocol awareness, and resilience through auto-shunning of unhealthy nodes.

Client connections and worker threads

Client TCP sessions terminate on worker threads, not on a single acceptor that also runs queries. MySQL worker threads handle authentication, query parsing, and result-set relay for the bulk of traffic; dedicated PgSQL worker threads cover the PostgreSQL protocol path. The design is explicitly multi-threaded so different thread types avoid contention: workers own the client data path, the admin thread owns the SQLite-backed configuration interface, monitor threads own health checks, and cluster threads (when used) own peer sync. On the diagram, show many client arrows funneling into a small set of labeled worker boxes inside the ProxySQL boundary — that fan-in is the first half of why multiplexing matters later. Configuration itself is multi-layer: an internal SQLite database separates persistent disk state from active runtime state so operators can load changes transactionally without bouncing the process.

Query processor: route, rewrite, cache, block

The query processor is the “brain” of ProxySQL. For every query the worker identifies the user and schema, then the processor matches query rules and decides four orthogonal outcomes:

  • Routing — which hostgroup should handle this statement?
  • Rewriting — does the SQL need modification (for example index hints) before it hits a backend?
  • Caching — is there a valid in-memory cached result to return without touching MySQL?
  • Blocking — should security or policy rules deny the statement entirely?

Under the hood, ProxySQL fingerprints statements into a query digest by normalizing literal values, so statistics and rule matching work on query shapes rather than specific data. Compiled regex patterns (RE2 or PCRE) are cached for high-speed matching across large rule sets. On the diagram, draw the processor as a decision diamond or labeled box between workers and the HostGroups manager, with four callouts — route, rewrite, cache, block — and a short-circuit arrow for cache hits that never leave ProxySQL.

Hostgroups and read/write split

Backend servers are grouped into hostgroups. Each hostgroup maintains its own pool of persistent connections. The classic platform pattern is hostgroup N for writers and hostgroup M for readers: query rules send transactional writes and primary-required statements to the RW group, and send read-only SELECTs to the RO group. The HostGroups manager also tracks latency to prefer faster responders and monitors read_only status so writers and readers stay correctly classified after topology changes. That is the architectural answer to “how do we do MySQL / PostgreSQL proxy routing without teaching every app about replica DSNs?” — the app talks to ProxySQL; ProxySQL’s rules and hostgroups own the split. Draw RW and RO hostgroup boxes distinctly, each with one or more MySQL backends underneath, and label the rule edge that chooses the hostgroup ID.

Multiplexing to backend connections

Multiplexing is why ProxySQL is not just a dumb TCP forwarder. Many frontend sessions can share a smaller number of backend connections when statements are compatible with that sharing. The lifecycle in the docs is: client → worker → query processor (match rules; maybe return cache) → HostGroups manager requests a connection for hostgroup X → backend executes → result relays back to the client. Thread-local statistics keep counters off the global lock path so workers scale with cores; jemalloc reduces fragmentation on long-lived proxies moving large result sets. Draw the fan-in clearly: wide client fan into workers, narrower pool arrows from each hostgroup into backends — and a callout that multiplexing is rule-sensitive, not a blind promise of “1,000 clients → 1,000 MySQL threads.”

Monitor thread: lag and failover

Monitor threads continuously health-check backends: pings, connection tests, and replication lag. Combined with read_only tracking, that feedback lets ProxySQL auto-shun unhealthy nodes and keep hostgroup membership honest during failover. On the architecture diagram this must be a dashed control-plane path under the solid client→worker→processor→hostgroup→backend data plane — otherwise reviewers confuse monitoring with query traffic. Cluster deployments add peer-to-peer checksum sync so nodes converge on the same config without a single control plane SPOF; keep cluster threads as an optional side box unless your review scope includes ProxySQL Cluster.

FAQ: ProxySQL vs app-side pooling

How does ProxySQL differ from app-side connection pooling?

App-side pools sit inside each application process and usually only reuse TCP sessions to one endpoint. ProxySQL is a protocol-aware proxy in front of MySQL (and PostgreSQL) backends: worker threads accept client connections, a query processor applies route / rewrite / cache / block rules, hostgroups implement read/write split across backends, and multiplexing shares a smaller backend pool across many frontend sessions. Failover and lag awareness live in the monitor path, not in every app.

What does multiplexing mean in a ProxySQL architecture diagram?

Multiplexing lets many frontend client sessions share a smaller number of persistent backend connections inside each hostgroup’s connection pool. It is not a blind 1:1 map of clients to MySQL sessions. Statement types that break sharing need explicit query rules so the processor keeps those sessions on dedicated backend connections.

What does the ProxySQL monitor thread do for lag and failover?

Monitor threads continuously health-check backends (pings, connection tests, replication lag) and track read_only status to distinguish writers from readers. Unhealthy or overly lagged nodes can be auto-shunned so the HostGroups manager stops routing traffic to them until they recover — draw that as a dashed path under the client→backend data plane.

Conclusion

Keep the diagram evergreen and honest to the published architecture: clients → workers → query processor (route / rewrite / cache / block) → hostgroups (RW / RO) with multiplexing into MySQL backends, plus a dashed monitor path for lag and failover. Cite the ProxySQL Architecture documentation for thread types, query lifecycle, and HostGroups manager behavior. Browse more architecture diagrams on the ByteDiagram blog.

Diagram ProxySQL query rules and hostgroups

Map clients, worker threads, the query processor, RW/RO hostgroups with multiplexing, and the monitor path for lag and failover in ByteDiagram — then take the picture into your next database platform review.

Open Diagram Editor