← Back to Blog
News

ClickHouse CAS Architecture Diagram: Compute-Storage Separation for MergeTree

On September 21–22, 2026, Altinity introduced experimental Content Addressed Storage (CAS) in Antalya 26.6 — a drop-in extension so MergeTree table data can live as content-hashed blobs on shared object storage while ClickHouse compute nodes scale independently. This ClickHouse CAS compute storage separation architecture diagram maps that shape: query nodes through a metadata_type=cas disk (plus cache) into one shared S3/GCS pool, with immutable manifests, mutable refs, and Keeper still owning part-set consensus. Facts below follow the Altinity Introducing CAS post. Treat it as experimental Antalya — not a vanilla ClickHouse Cloud feature and not a revival of deprecated zero-copy replication.

ClickHouse CAS compute storage separation architecture diagram showing MergeTree query nodes sharing content-hashed parts on object storage via Altinity Antalya CAS
ClickHouse query nodes A/B → optional local/hot volume → CAS disk (metadata_type=cas + cache) → shared S3/GCS pool (content-hashed blobs, immutable manifests, mutable refs). Keeper keeps the replication log / part-set only. Labels: Experimental Antalya 26.6; not ClickHouse Cloud; scale storage ≠ compute.

What a ClickHouse CAS compute storage separation architecture diagram shows

Draw two (or more) ClickHouse query nodes that both attach to the same CAS-backed storage policy. An optional local/hot volume sits above a CAS disk configured with type=object_storage, metadata_type=cas, and usually a cache disk in front. Under the CAS disk, a single shared object storage pool (AWS S3 or GCP GCS) holds three layers Visual should label distinctly: content-hashed blobs, immutable manifests, and mutable refs that map part name → ref → manifest. Off to the side, ClickHouse Keeper still coordinates the ReplicatedMergeTree replication log and part-set — not blob ownership. Key arrows: both nodes read and write the CAS pool; a part-publish sequence (manifest → precommit → blobs → committed ref); a GC reachability chain; and a dashed “not zero-copy bookkeeping” callout so reviewers do not confuse this with the deprecated design.

Shared-nothing MergeTree vs shared object storage (why zero-copy failed)

Classic open-source ClickHouse is a textbook shared-nothing MPP system: every replica owns CPUs and its own local slice or copy of data. That delivers parallelism, but it cannot scale storage and compute independently — growing the query fleet means more disks, more copies, or slow hydrate. ClickHouse added S3 disks years ago, yet replicas still kept their own data copies on object storage. Zero-copy replication tried to share one remote copy across replicas, but the design coordinated three systems transactionally: local data references, remote blobs, and Keeper state per replica. Orphan-file GC was a practical failure mode; upstream deprecated zero-copy in 2025 (tests removed). CAS is Altinity’s ClickHouse-only answer after Project Antalya’s Iceberg/swarm path: keep MergeTree, change how parts are addressed on the bucket, and avoid Keeper-mediated refcounts over shared object paths.

Content-addressed parts: hash → manifest → mutable ref

CAS is a MetadataStorage backend for object-storage disks. Every MergeTree part file is keyed by a hash of its content. If another part contains the same file bytes, it resolves to the same object — no second copy. The part name does not point at blobs directly. The chain is:

part name → mutable ref → immutable manifest → immutable hashed blobs

Small metadata files (for example count.txt / columns.txt) can embed in the manifest; larger files are separate content-addressed blobs. Publication is ordered so a crash cannot leave a visible part pointing at missing data: create the immutable manifest, record a durable precommit ref, upload or adopt blobs (conditional object writes; losers adopt instead of duplicating), then promote to a committed ref. Reading walks the chain the other way: resolve ref, validate manifest, range-read blobs. Coordination data — refs, manifests, mount leases, fencing, GC metadata — lives in the bucket as the source of truth for bytes; there is no mutable per-blob refcount in Keeper.

Architecture diagram: query nodes ↔ CAS disk ↔ shared S3/GCS

Operationally, CAS is a storage-policy change after upgrading to Antalya 26.6+ (use a build at or above the format change in 26.6.4). Define an object-storage disk with metadata_type=cas and a unique cas_server_root_id (typically the {replica} or {server_uuid} macro), wrap it in a cache disk, and attach it via a policy — CAS-only or tiered with a local default volume ahead of CAS. All MergeTree engines work; you can create tables on CAS, move parts between local / regular S3 / CAS, or use CAS as a cold tier with TTL. On the diagram, show query node A and query node B both talking to the same pool through their CAS+cache disks; label “scale storage ≠ compute” on the pool edge. Altinity’s ontime demo notes a new replica initializing against CAS in seconds rather than network-copying part bytes — treat hydrate timing as qualitative, not a benchmark you invent.

Replication, GC, and what Keeper still owns

CAS does not replace ClickHouse replication. Keeper still owns the replication log and part-set consensus — how other replicas learn a new part arrived. What changes is how the receiving replica obtains bytes: zero-copy said “replica 2 should reference the same remote path replica 1 uses” (Keeper bookkeeping); CAS says “the file’s hash identifies the object; any replica can independently publish a reference.” Relink is a three-step confirm-then-publish flow; uncertain results retry, and bytes download only when relink cannot be confirmed. Garbage collection follows reachability: a live ref keeps its manifest; a live manifest keeps its blobs; removed refs make blobs eligible after later GC rounds that recheck and delete only the exact marked object version. That reachability model is the architectural answer to zero-copy orphan files — draw it as a short chain under the pool, not as Keeper refcounts.

Experimental status, Antalya 26.6, and operational pitfalls

Stamp the diagram clearly:

  • Experimental in Antalya 26.6 — open source; Altinity targets OSS GA by end of 2026 (hedge, not a promise).
  • Not ClickHouse Cloud feature parity — this is an Antalya build extension, not a managed Cloud announcement.
  • AWS S3 and GCP GCS functionally complete; Azure not validated.
  • Inserts not fully performant yet — prefer local/hot for high-frequency inserts and merges, then TTL / move to CAS as cold tier; queries are described as working well qualitatively.
  • Watch S3 503 Slow Down under heavy direct write; raise compact-part thresholds (min_bytes_for_wide_part, min_level_for_wide_part); shard pools with {shard} or separate table endpoints when clusters are large.
  • Do not invent Iceberg parity claims: Antalya Iceberg/swarm and CAS solve different trade-offs; CAS’s advantage is drop-in MergeTree storage policy.

FAQ: CAS, Antalya, and compute-storage separation

What is ClickHouse CAS (Content Addressed Storage)?

An experimental metadata_type=cas for object-storage disks in Altinity Antalya 26.6. MergeTree part files are stored once by content hash on a shared S3 or GCS pool so query nodes scale compute without per-replica data copies. It works with all MergeTree engines via a storage policy.

Is CAS the same as ClickHouse Cloud or zero-copy replication?

No. CAS ships in experimental Antalya builds — not as a vanilla ClickHouse Cloud feature. It is also not a revival of deprecated zero-copy replication: CAS shares content hashes and manifests on object storage instead of Keeper bookkeeping over shared remote object paths.

What does ClickHouse Keeper still own with CAS?

Keeper still coordinates the ReplicatedMergeTree replication log and part-set consensus — how replicas learn that a new part exists. CAS changes how replicas obtain part bytes (hash identity, manifests, and refs on the bucket), not the part-set agreement itself.

Conclusion

Draw the compute-storage backbone honestly: query nodes → optional hot volume → CAS disk + cache → shared S3/GCS (hashed blobs, manifests, refs), with Keeper on the side for part-set only, a publish sequence and GC reachability chain, and labels for experimental Antalya 26.6 / not ClickHouse Cloud / not zero-copy. Cite the Altinity Introducing CAS post for scope, cloud-provider status, and operational hints. Browse more architecture diagrams on the ByteDiagram blog.

Diagram MergeTree compute-storage separation

Map ClickHouse query nodes, a CAS disk with content-hashed parts on shared S3/GCS, mutable refs and manifests, and Keeper’s part-set role in ByteDiagram — then stamp experimental Antalya so reviews do not confuse it with ClickHouse Cloud or zero-copy.

Open Diagram Editor