On October 1, 2026 Cloudflare’s launch post said the Cloudflare Data Platform is generally available and renamed Basin. Basin docs, updated the same day, say the old names — Cloudflare Pipelines, R2 Data Catalog, and R2 SQL — are now Basin Pipelines, Basin Catalog, and Basin SQL, and that existing resources keep working. This news diagram is that GA path: ingest and transform, Iceberg tables on R2, then serverless SQL. Pricing sentences and customer quotes in the launch post are vendor claims, not independent measurements.
What this Cloudflare Basin Iceberg architecture diagram shows
Three products, in the order both sources use. Basin Pipelines receives events from HTTP endpoints, Workers bindings, or Cloudflare Logpush, runs a SQL transform, and writes Apache Iceberg tables in Basin Catalog or JSON or Parquet files in R2. Basin Catalog is an Iceberg REST catalog on an R2 bucket. The docs say it maintains tables. Basin SQL queries those tables without a cluster you size: the launch post says it splits work using statistics from Catalog and runs tasks on Workers. The docs’ flow diagram is Pipelines into Catalog, then Catalog queried by Basin SQL and compatible external engines. The launch post names PyIceberg, DuckDB, Snowflake, and Apache Spark as engines that can read and write the same Iceberg data. On this diagram those engines are a dashed Iceberg REST path, not a second warehouse.
The problem this architecture is solving
The launch post points at Iceberg as an open table format and at R2 without egress fees so more than one engine can read the same bytes. “Zero egress,” cost quotes, and the Anomaly and Bobsled statements are vendor-published claims, not measured benchmarks. The GA description is narrower than a warehouse story: Pipelines, Catalog, and SQL are separate, files stay Iceberg on R2, and Basin SQL does not ask you to size a cluster. Usage-based billing with no hourly charge is the post’s pricing description — confirm numbers in the docs.
Main components and trust boundaries
Pipelines is the place SQL can drop or reshape fields before they are stored. The post’s HTTP-log example hashes a client IP and keeps only error statuses. That is a sample transform, not a default policy. Data-quality errors — missing fields, type mismatches, parse failures, nulls — are visible in the dashboard and a GraphQL API, according to the post. Worker bindings are schema-aware: wrangler types generates TypeScript from the stream schema. Terraform can describe catalog, stream, sink, and the SQL between them. Those are control-plane edges. The data plane is R2 holding table files and Catalog holding Iceberg metadata. Query engines, Basin SQL or an external one, trust Catalog’s metadata to find files. The October 1, 2026 launch post says unreferenced data-file cleanup reclaims storage when snapshots expire, without a separate Spark maintenance job. That Spark line covers unreferenced data-file cleanup only, not compaction, snapshot expiration, or manifest optimization, which the same post lists as separate maintenance items.
Request or data path, step by step
Enable a catalog (wrangler basin catalog create in the post, or enable on a bucket in the docs), then a stream, sink, and pipeline SQL. Events arrive by HTTP, a Worker binding, or Logpush. Pipelines writes Iceberg or files. The post’s maintenance list is per-table compaction targets, snapshot expiration that still keeps a minimum of recent snapshots, unreferenced-file cleanup, and manifest grouping. Queries go through Wrangler, the API, or the dashboard editor. The post says SQL at GA includes joins, windows, and more than 190 functions — that count is the post’s claim. External engines use the Iceberg REST catalog instead of Basin SQL.
The diagram: labeled boxes and failure or isolation edges
Left to right: Workers, HTTP, and Logpush into Pipelines; Pipelines into Iceberg tables in Catalog on R2; Basin SQL workers reading Catalog and returning results; external engines on a dashed Iceberg REST path, not a second warehouse. Isolation: transforms that hash or filter happen before the write, so the stored table is not a raw mirror of the event unless your SQL says so. That Iceberg REST path is the only dashed edge. The sources do not name a warehouse cluster or a node count. Failure behavior for a pipeline outage, exactly-once delivery, or a query retry policy is not specified on these two pages.
What the source does not claim (preview, case study, or limits)
GA is the October 1, 2026 claim for Pipelines, Catalog, and SQL. The open-beta period is what the launch post says came before that, including internal Cloudflare use. Ingest rate is inconsistent across the two pages fetched here. The launch post says Pipelines can ingest up to 3 GB/s per stream. The docs’ “What’s new” note the same day says streams can ingest up to 1 GB/s each, up from 5 MB/s. Those two pages do not agree on one ingest rate. Roadmap items are not GA: custom partitioning, schema migration and editable pipeline SQL, Iceberg V3 and the Variant type, stateful streaming aggregations, sort-and-cluster compaction, finer catalog auth, jurisdiction controls, full DDL in Basin SQL, and adaptive scheduling. The post lists those as in progress or planned.
How this differs from a nearby pattern on ByteDiagram
Basin SQL’s distributed query diagram is the query tier alone: how a serverless engine reads Iceberg on R2. This diagram is the platform around it — Pipelines in, Iceberg tables in Catalog on R2, then Basin SQL workers — and external engines on a dashed Iceberg REST path. Use the other picture when the review is only the query workers. That query-only diagram does not include Logpush, transforms, or snapshot expiration.
FAQ
What did Cloudflare rename when Basin became generally available?
The October 1, 2026 launch post and the Basin docs say the Cloudflare Data Platform is now Basin. Cloudflare Pipelines, R2 Data Catalog, and R2 SQL are Basin Pipelines, Basin Catalog, and Basin SQL. Existing configurations continue to work.
Where do Basin tables live, and who can query them?
Pipelines writes Apache Iceberg tables through Basin Catalog, which stores table data in R2, or it writes JSON or Parquet files to R2. Basin SQL queries catalog tables without a sized cluster. The launch post also names PyIceberg, DuckDB, Snowflake, and Apache Spark as Iceberg-compatible engines. Zero egress is Cloudflare’s R2 claim, not a third-party measurement.
What ingest rates do the October 1, 2026 sources publish?
The October 1, 2026 launch post says up to 3 GB/s per stream. The docs’ what’s-new note that day says up to 1 GB/s per stream, raised from 5 MB/s. The two sources disagree. Planned features such as Iceberg V3 and stateful processing are not current GA behavior.
Conclusion
As of the October 1, 2026 GA notes, the path is Workers, HTTP, or Logpush into Pipelines, Iceberg tables in Basin Catalog on R2, then Basin SQL workers, with other Iceberg engines on a dashed Iceberg REST path rather than a second warehouse. The launch post’s up to 3 GB/s per stream and the docs’ up to 1 GB/s per stream stay a disagreement between those pages, not a single rate. Planned items — Iceberg V3, stateful Pipelines, full DDL — are not in that GA description. The query-tier companion is Basin SQL / R2 distributed query. Browse more on the ByteDiagram blog.
Diagram Cloudflare Basin on Iceberg
Map Pipelines, Catalog on R2, and Basin SQL in ByteDiagram.
Open Diagram Editor