A Transactional Outbox solves one precise problem: your service must change its database and announce that change, but the database and message broker do not share one atomic transaction. Write the business row and an immutable outbox row in the same local database transaction, commit once, and publish the outbox asynchronously through a polling relay or change data capture. Design for at-least-once publication: every event needs a stable ID, consumers must deduplicate, ordering must be scoped to an aggregate, and operators need lag, retry and retention controls.
The AWS, Debezium, PostgreSQL and Apache Kafka documentation referenced here was verified on 6 October 2026. Those sources document the pattern and platform behavior; the schema, rollout gates and operational thresholds below are engineering recommendations.
Why the normal dual write fails
The tempting implementation is `UPDATE orders`, commit, then publish `OrderConfirmed`. If the process crashes after the commit but before the broker acknowledges the event, the order exists and downstream systems never learn about it. Reversing the order is not safer: an event can escape before the database transaction rolls back. Retrying the entire request can duplicate one side.
AWS Prescriptive Guidance describes this as the dual-write problem and recommends storing the event with the database change in one transaction. The guarantee is not "exactly once everywhere." The guarantee is narrower and useful: a committed business change has a committed event record that can be published later; a rolled-back change has neither.
Kafka producer idempotence or broker transactions do not automatically include your PostgreSQL transaction. They protect writes inside Kafka under their documented scope. Unless a single supported transaction coordinator owns both resources, the application still has two commits. The outbox removes that cross-system atomicity requirement.
Commit one durable fact
The command handler validates the request, changes the aggregate and inserts an outbox row before one database commit. Do not call the broker inside this transaction: network latency lengthens locks, broker failure blocks database work, and an uncertain timeout still cannot tell you safely whether the publish happened.
A useful outbox row carries `event_id`, `aggregate_type`, `aggregate_id`, `event_type`, an explicit schema version, payload, occurrence time, trace context and creation time. The event ID is generated once before insert and never changes. The aggregate ID becomes the partition or ordering key. The payload is a published contract, not a dump of every internal column.
BEGIN;
UPDATE orders
SET status = 'confirmed', version = version + 1
WHERE id = :order_id AND version = :expected_version;
INSERT INTO outbox_events (
event_id, aggregate_type, aggregate_id, event_type,
schema_version, payload, occurred_at, created_at
) VALUES (
:event_id, 'order', :order_id, 'order.confirmed',
1, :payload::jsonb, :occurred_at, now()
);
COMMIT;Generate the payload from the same validated state written by the transaction. If a trigger creates events from arbitrary row diffs, business meaning can become ambiguous. Application-owned event creation is usually clearer; triggers are reasonable when database-level enforcement is intentional and well tested.
Treat the outbox schema as a contract
Use an append-only table. Updating the payload after commit destroys the audit trail and can make two publication attempts carry different facts under one ID. Debezium's Outbox Event Router documentation likewise expects inserts and exposes a unique event ID, aggregate key, type and payload. It notes that the aggregate key is important for preserving order inside Kafka partitions.
Version the event envelope independently from the application deployment. Prefer additive schema changes and make consumers tolerate unknown fields. When a breaking change is unavoidable, publish a new event version or topic contract and migrate consumers deliberately. Do not reuse an old event name with incompatible meaning.
Keep tenant and authorization boundaries explicit. A multi-tenant service should carry a tenant identifier needed for routing and policy, but the event must not leak secrets or fields that consumers are not entitled to receive. Encrypt broker transport and storage, restrict topic access, and apply payload retention as a data-governance decision.
Option one: a polling publisher
A polling relay is the simplest design when traffic is moderate and you own PostgreSQL. Multiple workers can claim batches using `FOR UPDATE SKIP LOCKED`; PostgreSQL documents that `SKIP LOCKED` gives an inconsistent general-purpose view but is appropriate for multiple consumers of a queue-like table. Use a deterministic `ORDER BY`, a small batch and short transactions.
WITH claimed AS (
SELECT id
FROM outbox_events
WHERE published_at IS NULL
AND next_attempt_at <= now()
ORDER BY created_at, id
FOR UPDATE SKIP LOCKED
LIMIT 100
)
UPDATE outbox_events o
SET claimed_at = now(), claimed_by = :worker
FROM claimed
WHERE o.id = claimed.id
RETURNING o.*;Do not hold row locks while waiting for the broker. Claim a short lease, commit, publish, then mark success. If the worker dies after the broker accepts the message but before `published_at` is stored, the lease expires and the event is published again. That duplicate is expected. A unique event ID and idempotent consumer make recovery safe.
Use exponential backoff with jitter for temporary failures, a maximum attempt count or quarantine state for poison events, and an operator-visible replay action. Never delete a failing row merely to unblock the queue. Partition or archive published rows so the hot unpublished index remains small.
Option two: CDC with Debezium
CDC reads committed database changes from the transaction log. Debezium can capture the outbox table and transform each insert through the Outbox Event Router. This removes application polling and often reduces publication latency, but adds connector, replication-slot, schema and broker operations.
PostgreSQL logical decoding turns WAL changes into an application-consumable stream. Its documentation warns that after a crash a slot can return to an earlier LSN and resend recent changes, so clients must tolerate repeats. A stalled replication slot can also retain WAL and consume storage. Monitor slot lag in bytes and time, connector heartbeat, last published event time and disk headroom.
Choose CDC when the platform already operates Kafka Connect or Debezium, throughput or latency makes polling unattractive, and the team can own replication-slot recovery. Choose polling when operational simplicity matters more, volumes are bounded and database queries can be tuned safely. Both implement the same contract; neither removes consumer idempotency.
Delivery semantics belong end to end
At-least-once means the system may deliver the same event more than once but should not silently lose a committed event. Exactly-once claims are always scoped. Apache Kafka producer idempotence prevents duplicate writes caused by producer retries under Kafka's configuration rules; it does not deduplicate a business action applied in another database.
Each consumer should insert `event_id` into a processed-events table under the same local transaction as its side effect. A unique constraint turns a duplicate into a no-op. For a balance change or inventory movement, prefer recording an immutable movement keyed by event ID rather than reading a value and applying a blind increment.
BEGIN;
INSERT INTO processed_events (consumer, event_id, processed_at)
VALUES ('billing', :event_id, now())
ON CONFLICT DO NOTHING;
-- Continue only if one row was inserted.
UPDATE invoices SET order_status = :status WHERE order_id = :order_id;
COMMIT;The deduplication record and side effect must share one transaction. A Redis key written separately can expire too early or succeed while the database change fails, recreating a dual write at the consumer. Keep deduplication retention at least as long as the maximum broker replay and recovery window.
Preserve only the ordering you actually need
Global ordering is expensive and rarely required. Most domains need per-aggregate order: changes for one order, shipment or account must arrive sequentially, while unrelated aggregates can run in parallel. Use `aggregate_id` as the partition key and include an aggregate version or sequence in the payload.
A consumer should reject or defer a version gap, ignore an already-applied version and alert when a missing predecessor does not arrive. Timestamps alone are not a reliable sequence across hosts or concurrent transactions. Database commit order, outbox creation order and broker observation order can differ unless the relay and partitioning contract make the boundary explicit.
Do not promise ordering across two aggregates unless the business invariant truly requires it. If one workflow spans services and needs compensating actions, use a saga or explicit process manager on top of reliable events; the outbox transports facts but does not coordinate the whole business transaction.
Operate backlog and recovery as product behavior
Measure the age of the oldest unpublished row, unpublished count, publish latency percentiles, attempts by error class, quarantined events, relay lease expiry, connector or slot lag, consumer lag and deduplication hits. A queue can look healthy by depth while its oldest critical event is stuck. Alert on age and business importance, not only volume.
Run failure drills: terminate the service after commit; kill the relay after broker acknowledgement; pause the broker; duplicate a message; deliver versions out of order; stop a CDC connector long enough to grow WAL; and replay a quarantined event. Confirm there is no lost committed event and no duplicated business side effect.
Define retention separately for unpublished, published and processed-event records. Never purge unpublished rows. Archive published rows only after broker retention and recovery requirements are understood. Protect the cleanup job with a conservative watermark, dry-run metrics and a rollback-free rule: deletion is allowed only for terminal published rows older than policy.
Anti-patterns
Do not publish inside a database transaction and assume a broker timeout means failure. Do not mark the outbox row published before the broker acknowledges it. Do not create a new event ID on every retry. Do not update outbox payloads in place. Do not use a cron job that selects unlocked rows without a lease or row claim, because multiple workers will race.
Avoid routing every event through one partition for "ordering" when only per-order sequencing is needed. Avoid a generic event envelope with dozens of optional fields and no schema ownership. Avoid sending sensitive snapshots because the broker is convenient. And do not call duplicate delivery a rare edge case; design and test it as normal recovery.
When not to use the pattern
A monolith that performs all required work in one database transaction does not need an outbox for internal function calls. A synchronous request that must return the downstream result may need an API workflow, not an asynchronous event. If the source datastore already provides a durable change stream with the required semantics, an extra outbox table may duplicate infrastructure.
Do not use the outbox as a substitute for event sourcing. An outbox publishes integration events; it is normally not the source of truth from which the entire aggregate is rebuilt. Do not add Kafka, CDC and schema registry to a low-volume system merely for fashion—a well-indexed polling relay can be more reliable for the team that operates it.
Production decision
The Transactional Outbox replaces an impossible cross-system promise with two durable local contracts: the service commits state and event together, and the delivery pipeline keeps retrying until downstream processing is safe. Pick polling or CDC from operational capability, not ideology.
Use stable event IDs, aggregate-scoped ordering, immutable versioned payloads, idempotent consumers, explicit leases or replication-slot monitoring, and tested replay procedures. Connect this pattern to backend and API integration engineering and the guide to reliable webhook processing: outbound events and inbound webhooks need the same discipline around duplicates, retries and visible recovery.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




