Distributed systems · Event streaming

Cloudflare K2 changes the trade-offs of event streaming

K2 builds a durable serverless log on object storage. Here is what its batching, leases and at-least-once delivery mean in production.

TOPIC HUBAPIs, SaaS & System Architecture
Original conceptual illustration of event batches entering durable log segments on an object-storage plane and fanning out to independent consumers.
An editorial interpretation of the topic, followed by a practical execution diagram.

Cloudflare launched K2 in public beta on 1 October 2026 as a serverless durable event-streaming primitive. Producers append batches to an ordered log. Independent subscriptions consume retained records at their own pace. Underneath, Cloudflare says K2 stores partitioned log segments on R2 object storage rather than running a traditional broker cluster with local disks.

That architecture is the important part of the release. K2 is not “Kafka without servers” in every behavioral sense, and it is not a faster replacement for a task queue. Moving log durability into object storage changes the operating model, cost shape, latency floor and failure modes. Teams evaluating K2 should begin with those contracts—not the convenience of creating a stream in one command.

The service is currently a public beta for Workers Paid accounts. Cloudflare documents a 10 GB account storage limit during beta, configurable retention from one hour to 30 days, and no usage billing during the beta. Those facts make it suitable for evaluation, not an automatic production migration.

What Cloudflare actually released

The K2 overview defines a stream as a durable log. A producer writes records, the stream retains them for a configured period, and one or more subscriptions track independent positions in the log. A subscription may distribute batches across a pool of workers, while separate subscriptions let multiple applications each see the same records.

Records contain binary content plus optional string headers. Producers send records in batches, and the produce contract says a batch is atomic: all records in the batch are stored or none are. Production is available through an HTTP endpoint or a Workers binding.

Consumers pull from a subscription. K2 leases a batch to a worker for five minutes. The worker may acknowledge it, negatively acknowledge it for redelivery, or extend the lease. If the lease expires before acknowledgement, K2 can deliver the records again. Cloudflare therefore documents K2 as at-least-once delivery, not exactly once.

This distinction controls the consumer design. A successful HTTP response from the producer proves that K2 accepted a batch. It does not prove that every downstream database, warehouse or API applied each record once.

Why an object-storage log is different

Traditional brokers commonly keep active log segments close to broker compute and replicate them across nodes. K2 instead uses R2 as the durable state layer. Cloudflare describes R2 as providing durable, strongly consistent object operations, allowing K2 to push replication and consensus responsibilities into storage while keeping the application layer simpler.

Object stores do not support appending bytes to an existing object in the way a local log file does. K2 accumulates writes in memory, waits briefly for a batch, and writes a complete segment object. It uses R2 atomic operations to assign strictly increasing offsets without a separate coordination service.

This creates useful separation. Compute can be ephemeral and scale independently from retained data. Historical records can live on storage designed for durability and capacity rather than on a permanently provisioned broker fleet. A consumer outage does not require the producer to remain connected to that consumer.

The cost is latency. Cloudflare reports that the initial release adds roughly one second of produce latency at the 99th percentile because it waits for a segment batch and writes it to object storage. That is a company-reported figure for the beta architecture, not an independent benchmark or a promise for every workload.

The decision is therefore not “managed versus self-hosted.” It is whether a workload benefits more from serverless durability, retention and fan-out than it suffers from the higher write-latency floor and a younger operational surface.

Do not confuse a stream with a task queue

Cloudflare positions Queues around individual units of asynchronous work. A queue offers message-oriented features such as retries, delays and dead-letter handling. K2 is designed for high-volume data movement, retained history and fan-out subscriptions. It processes batches efficiently, but the trade-off is less message-level control.

Use a task queue when one item represents work that must succeed, fail, retry or move to a dead-letter path independently—for example generating an invoice PDF or resizing an uploaded image. Use a retained stream when multiple consumers need an ordered history, replay or independent progress—for example product analytics, audit events, change propagation or model-observation data.

Basin Pipelines is different again. Cloudflare recommends Pipelines when the destination is R2 or an Iceberg table and the main job is ingestion and transformation. K2 is the more general primitive when custom consumers process the records or write to other destinations.

The boundaries matter because adopting a stream does not remove workflow semantics. If an order event requires compensation, timeout ownership and a business state machine, those responsibilities still belong to the application.

A safe K2 pipeline accepts duplicate delivery, commits sink state idempotently and acknowledges a leased batch only after durable processing.
A safe K2 pipeline accepts duplicate delivery, commits sink state idempotently and acknowledges a leased batch only after durable processing. Open for a larger view

At-least-once means duplicate-safe consumers

The K2 getting-started guide explicitly warns that K2 does not deduplicate records. A producer retry can store the same logical event more than once. A consumer can also receive a leased batch again after a lost response, a lease expiry or a negative acknowledgement.

Every business event therefore needs a stable event_id generated before the first produce attempt. The payload should also carry an event type, schema version, aggregate identifier, occurred_at timestamp and producer identity. K2 headers can help routing, but the durable sink must enforce idempotency.

For a relational sink, insert the event_id into an inbox table with a unique constraint in the same transaction that applies the projection. If the insert conflicts, treat the event as already processed. For an external API, use the provider's idempotency key when available and keep an outbox or result ledger locally. A cache-only deduplication key is not enough when its TTL is shorter than stream retention or replay.

Acknowledge the K2 batch only after every accepted effect is durable. If a batch contains one invalid record, blindly retrying the whole batch can trap healthy records behind a poison event. K2 does not advertise message-level dead-letter behavior for this stream contract, so the consumer should validate each record, quarantine permanent failures in its own durable store, and make the batch decision explicit.

Leases define the concurrency contract

Each consume call identifies a worker_id. K2 gives that worker a leased batch, and the consumer documentation states that one worker can hold one lease at a time. A subscription supports up to 128 active leases during the current beta.

When processing may exceed five minutes, extend the lease before it expires. If extension returns a conflict because the lease is no longer owned by that worker, stop applying side effects and request a new batch. Continuing after lease loss can race with a second worker processing the redelivered records.

Autoscaling should follow useful signals: available records, processing duration, lease utilization, redelivery rate, sink latency and error classes. CPU alone cannot tell whether more workers help. A slow shared database can become less reliable when the consumer pool scales aggressively.

Keep worker identifiers unique per live process and stable for the lifetime of its in-flight request. Reusing one identifier across replicas collapses lease ownership. Generating a new identifier on every retry makes response recovery harder because K2 returns an existing leased batch to the same worker.

Ordered log does not mean ordered side effects

K2 describes the stream as ordered and assigns increasing offsets. That does not mean parallel consumers finish work in offset order. The documentation says K2 does not guarantee processing order across workers sharing one subscription.

If order matters for one customer, account or aggregate, do not infer a per-key partition guarantee that the beta does not document. Key-based ordering is listed on Cloudflare's roadmap, which means teams should not design as if it already exists.

Until that contract arrives, options include a single consumer for the ordered subset, downstream serialization by aggregate key, version checks that reject stale transitions, or a domain model whose events commute. Each option sacrifices throughput, simplicity or latency. Make the choice in the consumer architecture rather than hiding it behind the word “stream.”

Independent subscriptions are useful but multiply work. Analytics, fraud detection and search indexing can each track their own position and replay history, yet each needs separate capacity, monitoring, retention expectations and failure handling.

Design the event contract before the transport

The first schema should be an envelope, not an unversioned copy of a database row. Include event_id, event_type, schema_version, aggregate_id, occurred_at and a typed data object. Decide whether personally identifiable or regulated data is allowed before it enters a retained log.

Consumers should support at least the current and previous schema versions during migration. Additive changes are easier than changing meaning. When semantics must change, publish a new event type or version and run old and new consumers side by side until reconciliation proves the new projection.

Record the producer's business timestamp separately from K2's server receive timestamp. The receive timestamp helps measure ingestion delay; it cannot replace when the business event actually occurred. Validate clock assumptions before using either timestamp for ordering.

Do not let retention become accidental governance. K2 permits retention from one hour to 30 days in the published beta limits, with longer periods available by request. Choose the shortest period that covers outage recovery, replay and audit requirements. A durable log is not automatically the system of record.

Treat beta limits as architecture inputs

The published limits currently include 20 streams per account, 100 subscriptions per stream, a 5 MB produce request, an approximately 1 MB record limit, 10 MB per consume response and 10,000 records per consume request. The launch post lists up to 30 MB/s produce throughput per stream and 10 GB of storage across the account during beta.

These limits may change, but today's design must respect them. Split streams by lifecycle, security boundary and retention requirement—not by every event type. Too many streams consume the account limit and complicate replay. One universal stream, however, couples permissions, noisy producers and unrelated retention policies.

Load-test with realistic record sizes and consumer behavior. Measure producer p50, p95 and p99; batch fill efficiency; end-to-end freshness; lease extensions; duplicates; sink throughput; and recovery time after a consumer outage. Cloudflare's roadmap mentions lower-latency tiers, key ordering, push-based Worker consumers and Kafka-client compatibility, but roadmap items are not current guarantees.

A safe evaluation plan

Start with a non-critical fan-out workload such as product analytics or audit-derived projections. Keep the current path as the source of truth. Dual-publish with a stable event_id, consume K2 into a shadow destination and reconcile counts, ordering assumptions and duplicates against the existing pipeline.

Inject failure. Lose the producer response and retry. Kill a consumer after applying the database transaction but before ack. Let a lease expire. Quarantine a malformed record. Stop all consumers long enough to build backlog, then measure recovery without overwhelming the sink.

Define exit criteria before increasing scope: acceptable freshness percentiles, zero unexplained loss, bounded duplicate rate, proven idempotency, tested replay, controlled cost and an operational owner. Also define rollback: stop new production, preserve offsets and continue the old path until reconciliation finishes.

The architecture decision

K2 is interesting because it moves a durable ordered log onto object storage and exposes it as a serverless primitive. That can remove broker provisioning and make retained fan-out practical at the edge. It also introduces a visible latency trade-off and leaves application teams responsible for duplicate safety, poison records, domain ordering and sink transactions.

The right question is not whether K2 can replace Kafka in a diagram. It is whether its documented contracts fit the workload: batch-oriented production, at-least-once delivery, independent pull subscriptions, finite retention and current beta limits.

If those boundaries match, K2 can simplify the infrastructure around event history. If the workload needs sub-second tail latency, mature Kafka compatibility, strict per-key ordering or message-level workflow controls today, the launch documentation itself gives reasons to wait or choose another primitive.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceBackend engineering & API integrationsRelevant projectLogistics at scale