An AI agent that can call tools for minutes or wait days for approval is a distributed system, not a long HTTP request. Conversation history may help the model remember what was said, but it does not by itself prove which tool already ran, who approved an action, whether an external API committed before a timeout, or which code version must resume the job.
Current OpenAI Agents SDK documentation separates sessions and server-managed conversation state from `RunState`, which can resume paused work. It also points to durable orchestration integrations for runs that span long waits, retries or process restarts. Microsoft's Durable Task guidance describes automatic checkpointing of state transitions and recovery from the last checkpoint. These are useful primitives, but production safety comes from the contract around them.
A reliable design has four layers: conversation memory, execution state, durable orchestration and a side-effect ledger. Treating any one of them as the whole solution creates the failures that matter most: duplicated refunds, repeated emails, stale approvals and runs that resume under incompatible code.
Separate memory from execution truth
Memory answers what context the next model call should see. Execution state answers what the system has committed. Keep them separate even if one SDK can store both. A session may contain user messages and tool outputs, while the execution record should carry a run ID, current step, workflow version, pending tool calls, approval state, attempt counters, deadlines and references to external receipts.
OpenAI documents `RunState` as a serializable pause-and-resume boundary that includes model responses, generated items, interruptions and approval state. Its reference warns that independently restored snapshots must not resume concurrently against the same session. That warning exposes an application responsibility: add an optimistic version or lease so only one worker owns a transition.
A minimal durable record can look like this:
run_id, tenant_id, workflow_version
state_version, status, current_step
pending_action_id, approval_id
checkpoint_uri, session_ref
lease_owner, lease_expires_at
created_at, updated_atStore bulky model context or artifacts separately and reference them by immutable IDs. Encrypt sensitive state, apply tenant isolation and give it a retention policy. Do not put a client-supplied serialized snapshot directly back into the runner: the SDK documentation says serialized state contains approvals and pending tool arguments but does not authenticate the snapshot or the reviewer.
Checkpoint business transitions, not every token
A checkpoint should represent a stable boundary that can be replayed and inspected: plan accepted, tool result recorded, approval requested, approval resolved or final output committed. Persisting every streaming token adds cost without making an external action safer. Persisting only at the end loses expensive model work and leaves ambiguous operations after a crash.
Write the state transition and the job that schedules the next step atomically when possible. An outbox is useful when the database and queue cannot share one transaction: commit `step_completed` plus an outbox row, then let a dispatcher publish the next job. Consumers acknowledge only after the next durable state is committed.
The checkpoint must include the workflow and prompt/tool schema versions that produced it. A run waiting for approval today may resume after a deployment tomorrow. If the new worker interprets old state differently, recovery becomes a migration problem rather than a retry. Either route old runs to compatible workers or migrate snapshots explicitly with tested version adapters.
Design every external effect for at-least-once execution
Durable orchestration prevents lost progress; it does not magically make a payment, email or ERP mutation execute exactly once. Microsoft documents that Durable Task activities have at-least-once execution: if an activity finishes but its result is not recorded, the runtime may run it again. The official guidance therefore recommends idempotent activity logic such as upserts or checking existing results before creation.
Give each consequential tool call an application-owned operation key such as `run_id + action_id + semantic_version`. Store an action ledger row before dispatch, send the same key to providers that support idempotency, and save the provider receipt. A retry first reads the ledger. If the state is `succeeded`, return the stored result; if `in_flight` has expired, reconcile with the provider before deciding to call again.
prepared -> dispatched -> succeeded
\-> outcome_unknown -> reconciledDo not generate a new key on every retry. Do not mark failure merely because the client timed out: the remote side may have committed. For providers without idempotency keys, use a natural business key, a uniqueness constraint, a preflight lookup and reconciliation. Some effects—sending an email to the outside world, for example—cannot be perfectly undone; the goal is a measurable protocol for ambiguity, not a fictional exactly-once guarantee.
Make approval a durable, atomic decision
The OpenAI human-in-the-loop guide lets a run pause on a tool call, serialize state, apply an approval or rejection, and resume. It also gives important security guidance: keep the complete state on the server, authenticate and authorize the reviewer, validate decision IDs against pending server-owned requests, and atomically consume each decision so concurrent or replayed submissions cannot resume the same snapshot twice.
Approval should bind to an exact action, arguments, policy version, tenant and expiry—not to a vague sentence such as approve this agent. Show the reviewer a normalized diff: recipient, amount, resource, expected change and rollback or compensation path. Escape tool names and arguments in the UI because model-produced content is untrusted display data.
After approval, validate mutable preconditions again immediately before execution. A customer balance, deployment target or access role may change while the run is paused. Approval authorizes the reviewed intent; it should not bypass current authorization, limits or input guardrails. Record who decided, when, under which policy, and which action fingerprint the decision covered.
Keep orchestration replay-safe and versioned
Durable engines often rebuild control state by replaying history. Temporal's official workflow documentation requires deterministic workflow code and places API calls, database queries, model invocations and other nondeterministic work in Activities outside the replay path. It also warns that changing command order for a running workflow can cause nondeterminism and provides worker versioning or patching strategies.
The practical boundary is clear: orchestration decides what step should happen next from recorded history; activities perform uncertain I/O. Avoid reading wall-clock time, randomness or mutable configuration directly inside replayed logic unless the runtime records that value. Model calls belong in activities because the same prompt can return a different answer later.
Version prompts, tool schemas and routing policy independently. A compatible code deployment does not imply that a changed prompt is safe for an old approval. Keep the exact model result that selected a tool, but enforce authorization and business rules in deterministic application code rather than asking the model to remember policy on resume.
Test crashes at the commit boundaries
A happy-path agent demo proves almost nothing about durability. Build failure tests around the narrow windows where outcomes become ambiguous: after the provider commits but before the receipt is stored; after the checkpoint commits but before the next job is published; after approval is recorded while two workers race to resume; during a deployment with old runs waiting; and after the session append succeeds but the worker crashes.
Track recovery time, duplicate effects prevented, unknown outcomes awaiting reconciliation, checkpoint age, lease contention, approval latency, expired approvals, retries per activity, token cost per completed run and runs pinned to old workflow versions. Trace IDs help investigate, but traces are observation—not the source of truth for execution.
Start with one consequential workflow and write its invariants before selecting a framework: one owner per transition, one durable action identity per effect, one authorized decision per approval, and one compatible version path per checkpoint. Then choose whether a serialized SDK state plus your queue is enough or a durable orchestrator is justified. The strongest architecture is not the one with the most agent abstractions; it is the one that can explain, after a crash, exactly what happened and what is safe to do next.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
- OpenAI Agents SDK — Human in the loop, verified 27 September 2026
- OpenAI Agents SDK — Running agents and durable integrations, verified 27 September 2026
- OpenAI Agents SDK — RunState reference, verified 27 September 2026
- Microsoft Learn — Durable Task for AI agents, verified 27 September 2026
- Microsoft Learn — Durable Task programming model, verified 27 September 2026
- Temporal Docs — Workflow definition, determinism and versioning, verified 27 September 2026
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




