← Back to Blog
System Design

Neon Lakebase Architecture Diagram: Compute-Storage Separation & Copy-on-Write Branching

Neon’s architecture overview describes Lakebase Postgres as a serverless database that splits Postgres into an ephemeral compute layer and a durable storage layer joined by write-ahead log (WAL), not by a VM-local filesystem. This evergreen diagram is that split: clients use a compute endpoint, commits land when a safekeeper quorum acknowledges WAL, a pageserver reconstructs pages, and object storage keeps immutable history off the hot path. A copy-on-write branch is a metadata fork of that history. The overview says Neon and Databricks run the same database with different surroundings. Heikki Linnakangas’s March 30, 2023 deep dive names GetPage@LSN and is more specific.

Neon Lakebase Architecture Diagram: Compute-Storage Separation & Copy-on-Write Branching
Clients attach to ephemeral Postgres compute. WAL commits at a safekeeper quorum. The pageserver reconstructs a page at an LSN. A branch is a metadata fork, not a byte copy; the second compute attach follows a copy-on-write branch.

What this Neon lakebase architecture diagram branching shows

The compute box is ordinary Postgres. The overview says the engine still parses SQL, plans, executes, enforces MVCC, and manages locks and indexes. It uses RAM for shared buffers, session state, and hot data, and local NVMe as a page cache. It does not own durable state, so it can start, stop, scale, or fail without being the system of record. Storage is three boxes: safekeepers (WAL quorum), the pageserver (WAL into pages), and object storage (long-term immutable history). The overview says object storage is kept off the critical path and is never placed in front of query execution. Logically, a Project contains branches, a branch has databases and roles, and a compute endpoint attaches to a branch. An Operation (create branch, start compute) is async control plane, not the query path.

The problem this architecture is solving

The 2023 deep dive contrasts one node owning its disk — plus separate backups and replicas — with one storage system that serves the primary, read-only attaches, and time travel. The current overview states the serverless consequence: compute can scale up, down, to idle, and back without moving data. Creating a branch, restoring a snapshot, adding a read replica, or attaching compute become references to existing history rather than file copies. Neither page publishes a latency SLO or a throughput multiplier, so the article makes no unsupported numerical claim.

Main components and trust boundaries

Clients trust compute only as an executor. If it dies, the overview says queries stop and data remains safe because a new compute attaches to the same history. Commit trust is the safekeeper quorum. Compute streams WAL to multiple safekeepers; a transaction is committed once a quorum acknowledges the record via Paxos. The overview contrasts that with a compute-local fsync: commit latency is primarily quorum- and network-bound, with safekeepers batching WAL flushes, and no single machine defines durable state. Compute reads RAM, then NVMe, and only then asks the pageserver for a page at a log sequence number (LSN). It does not read object storage. The pageserver returns a version it already has, or loads a base page from object storage and replays WAL up to that LSN. Persistence of materialized pages is asynchronous, and commits do not wait for it. The 2023 post adds a key-value model (relation id plus block number; 8 KB page or WAL record) and immutable files for object storage such as S3 or Google Cloud Storage. That is the deep dive, not a new API.

Request or data path, step by step

On a write, the overview’s order is: Postgres updates memory and generates WAL; compute streams that WAL to multiple safekeepers instead of a local filesystem; the client hears success at quorum acknowledgement; page reconstruction happens later. The 2023 post adds three safekeepers, a Paxos-like algorithm, and layer files after 1 GB of buffered WAL; those details belong to that dated source. On a read, prefer local cache. The overview says an object-storage read may take hundreds of milliseconds and occurs only inside the pageserver, which is a design warning, not a cache-hit benchmark. GetPage@LSN, in the deep dive, finds the latest image at or before the requested LSN, applies WAL on top, and returns the page. Primaries ask for the latest LSN. A branch or node at an older LSN is how that post describes point-in-time restore.

The diagram: labeled boxes and solid arrows

Solid arrows only: clients to compute (SQL), compute to safekeepers (WAL), compute to the pageserver (get page), safekeepers to the pageserver (records), a second compute to the pageserver (attach), and the pageserver to object storage (history). The overview says a new branch points at existing history and diverges with copy-on-write, so only new or modified data consumes more storage. The 2023 post says the child starts empty, stores its own WAL, and reads unmodified pages from the parent. The article therefore shows one shared history rather than a cloned object store. Scale-to-zero is a property of ephemeral compute in the overview, not a label on the art. The five-minute idle shutdown and Kubernetes compute container in the 2023 post are that article’s operational description, not a figure the overview restates.

What the source does not claim (preview, case study, or limits)

These pages are evergreen technical descriptions. They are not marked preview, beta, or case study. They do not fix a production safekeeper count, an idle timeout, or a measured commit time. “Hundreds of milliseconds” qualifies object-storage reads inside the pageserver. “Five minutes,” three safekeepers, the 1 GB layer threshold, 8 KB pages, and Kubernetes packaging belong to the March 30, 2023 deep dive. The scope stays on the OLTP path described by the sources, without unrelated OLAP or AI engines. No custom resource fields appear in either source.

FAQ

When is a Neon lakebase transaction committed?

The architecture overview says compute streams WAL to multiple safekeepers and the transaction is committed once a quorum acknowledges the record via Paxos. The client is told success then. Page materialization and object-storage uploads happen later and are not on the commit path.

What is a copy-on-write database branch in this diagram?

The overview says a branch does not duplicate files or pages. It points at an existing point in history and diverges with copy-on-write, so only new or modified data consumes additional storage. The 2023 deep dive adds that the branch starts empty, its WAL is stored separately, and unmodified pages are read from the parent.

Does compute read Neon object storage on the query path?

No. The overview says compute checks RAM, then local NVMe, and on a miss asks the pageserver for a page at an LSN. Only the pageserver may read object storage while reconstructing that page. Those reads are described as possibly taking hundreds of milliseconds, and they are not in front of query execution.

Conclusion

The architecture centers on replaceable compute, a safekeeper quorum as the commit line, pageserver reconstruction at an LSN, and object storage off the hot path. A branch is a metadata fork, not a second copy of the bytes. Browse more architecture diagrams on the ByteDiagram blog.

Diagram Neon lakebase compute and branches

Map ephemeral Postgres compute, the safekeeper WAL quorum, pageserver reconstruction, object storage, and a copy-on-write branch fork in ByteDiagram — then take the picture into your next database review.

Open Diagram Editor