GPT‑6 adds two runtime capabilities that change how an interactive agent can behave: asynchronous tool calling lets the model continue independent work while an application-run tool is still executing, and mid-turn steering lets a user add requirements over a Responses API WebSocket before the current task finishes. Together they can reduce avoidable waiting and make long-running work easier to redirect. They do not provide a workflow engine, cancel completed actions or remove the need for authorization, durable state and recovery. A production design must therefore treat the model as one participant in an explicit state machine.
OpenAI introduced these GPT‑6 controls on 3 September 2026 and documents them separately. The documentation was verified on 5 October 2026. Statements about supported API behavior below come from OpenAI; the state model, budgets, rollout rules and observability design are engineering recommendations.
What OpenAI actually added
The OpenAI API changelog records async tool calling and mid-turn steering as GPT‑6 Responses API capabilities. The async tool guide says that setting `async: true` on an application-run function or custom tool allows the model to continue after issuing the call. Your application still executes the work and later returns the result using the original `call_id`. This is different from background mode, which makes response generation itself asynchronous.
The steering guide documents `response.steer` over a WebSocket. It queues new user input against a running response and can produce an automatic continuation. Acceptance means the input was queued, not that the model has acted on it. Steering does not rewrite output already emitted, undo side effects or cancel tools already started.
These details matter because an agent can now have several truths at once: a response may still be producing output, one or more tools may be pending, the user may have changed the objective, and a side effect may already have completed. A single `isRunning` boolean cannot represent that safely.
Model the agent as three cooperating state machines
Separate response state, tool state and product state. Response state tracks created, streaming, steered, completed, incomplete and failed outcomes. Tool state tracks proposed, authorized, dispatched, succeeded, failed, timed out and reconciled. Product state records the business facts that actually changed: an order was reserved, a draft was saved or a deployment was requested.
Persist their identifiers and relationships. A practical execution record contains `conversation_id`, `response_id`, `call_id`, an application `task_handle`, tool name, tenant, actor, authorization decision, idempotency key, deadline, attempt, status and a pointer to the result. The response ID is useful for model continuation; it is not your business transaction ID.
{
"execution_id": "exec_42",
"response_id": "resp_1",
"call_id": "call_inventory",
"task_handle": "inventory_7",
"tool": "check_inventory",
"status": "running",
"deadline_at": "2026-10-05T20:05:00Z",
"idempotency_key": "tenant:7:inventory:sku-91:v1"
}This is an application record, not an OpenAI schema. Keep it durable enough to survive a process restart, socket reconnect or worker retry.
Use async tools only for genuinely independent work
Marking every tool asynchronous creates concurrency without value. Strong candidates are slow read-only lookups, independent retrieval requests, file analysis, user questions and calculations whose results are needed later. The model can draft the stable portion, launch multiple lookups or explain progress while those jobs run.
Keep a tool synchronous when the next decision depends immediately on its result, when ordering is a business invariant or when executing it changes authority. A payment authorization must usually complete before capture. A schema migration must not race the validation that decides whether it is allowed. Parallel work is safe only when dependencies are explicit.
The OpenAI guide says async execution applies to functions and custom tools run by your application, not hosted built-in tools. It is supported by GPT‑6 Astra and later models. It also warns that multi-agent mode should not combine async tools with parallel tool calls. Treat compatibility as a deployment check, not an assumption inherited from a different model or tool type.
Build a pending-work registry and a real wait primitive
The provider returns a `call_id`, but the application needs its own registry for jobs, results and retries. OpenAI's guide proposes a unique `task_handle` for each async call and an ordinary synchronous `wait_for_tasks` function. That wait tool is an application pattern, not a built-in OpenAI tool. It lets the model pause only when its next step genuinely needs selected pending results.
Register each job before honoring a dependent wait. Bind the handle to the original call ID and reject handle reuse across the conversation. When jobs complete, return each `function_call_output` on its original `call_id` before the wait tool's status. Store completion before sending it to the model so a lost connection does not force the external action to run again.
Do not use the model context as the queue. A database row or durable workflow engine should own leases, heartbeats, deadlines, cancellation requests and terminal results. The model receives a concise projection of that state.
Treat steering as a new requirement, not cancellation
A steering event can say “limit the plan to one engineer” or “do not contact the vendor.” The API may end the interrupted response as incomplete with reason `steered`, or allow a completed response followed by the steering continuation. Each continuation inherits request settings, while token and tool-call limits apply separately to each response.
The application must classify the update. A presentation preference can affect future output immediately. A reduced scope may invalidate queued read-only work. A new security restriction should block any undispatched effect. But an already completed effect requires reconciliation; text cannot make it disappear. Expose this boundary in the product: “I applied the new constraint to remaining work; the earlier reservation was already created and is being reversed.”
Never translate steering directly into a hard process kill. Record a cancellation intent, ask each worker to stop at a safe checkpoint, and preserve whatever completed. For irreversible or expensive actions, require an approval token tied to the exact action, arguments, tenant, actor and expiry. A later steer can revoke future permission but cannot rewrite a consumed approval.
Reconcile outputs against side effects
An interrupted response can leave useful tool work behind. Before continuing, load the execution registry and divide calls into not started, running, completed, failed and uncertain. For completed calls, provide verified outputs to the continuation. For running calls, keep or cancel them according to dependency and policy. For uncertain calls, query the authoritative system using the idempotency key before retrying.
Side-effect tools should return stable resource identifiers and versions, not only prose. A `create_shipment` result should include shipment ID, status and source version. The final answer can then cite real state. If verification fails, the agent should abstain or request review instead of narrating success.
Keep compensation explicit. A refund is a new operation with its own authorization and failure modes, not a magical rollback of the original charge. Design sagas around business semantics, and let steering choose among allowed transitions rather than inventing reversal behavior.
Engineer WebSocket lanes and recovery
Mid-turn steering requires WebSocket mode. OpenAI documents FIFO execution for requests sharing a `stream_id`, concurrency across different lanes and a maximum connection lifetime of 60 minutes. A socket is a transport optimization, not durable state. Reconnect before or at the limit and resume each lane from stored application state.
If a stored response can be resolved, continue with `previous_response_id`. If the server returns `previous_response_not_found`, start a new response with the full relevant input or a compacted context. Do not replay side effects merely because the model continuation changed. Rehydrate completed tool results from the registry and preserve their original business identifiers.
Use one lane for the primary user task and separate lanes only for work that can truly overlap. Bound the number of lanes per tenant and user. Without admission control, WebSocket concurrency can turn a slow external dependency into a large pending-work backlog.
Enforce deadlines, retry classes and budgets
Every response, tool and user-wait task needs a deadline. Pass the remaining budget to workers and stop starting work that cannot finish in time. OpenAI's error guide distinguishes ramp-rate `429 slow_down` from temporary-capacity `503 server_is_overloaded`; both may provide `Retry-After`. Respect that header, use bounded jittered backoff when absent and keep the original task deadline.
Retry transport failures only when the operation is idempotent or can be reconciled. Do not automatically retry invalid arguments, denied approvals, spend-limit errors or deterministic business rejection. Count SDK retries and application retries together. A hidden retry storm can continue after a user steers away from the work.
Attach budgets for model tokens, tool calls, pending jobs, wall-clock time and external spend. Steering creates a new response and its own limits, so also enforce an execution-level ceiling across the whole user task.
Keep authorization outside the model
The model may propose an async tool call, but trusted code authenticates the actor, derives tenant scope, validates arguments and decides whether approval is required. Persist the authorization result beside the call. Revalidate time-sensitive facts before dispatch, especially after a long user wait.
Never place credentials in tool output or WebSocket events. Workers receive narrow, short-lived access to the resource they need. A steer can reduce future scope but must never expand authority without a fresh trusted policy decision. Apply the same rule to retries and compensating actions.
Treat tool descriptions as untrusted scheduling hints rather than policy. A prompt injection in retrieved content must not change which jobs can run, which tenant they target or whether an irreversible action requires approval.
Observe and evaluate the complete trajectory
Trace response IDs, steering IDs, call IDs and task handles under one execution ID. Record queue delay, tool duration, time to first useful output, number of steers, discarded work, cancellation latency, retries, reconciliations, total cost and verified task outcome. Avoid logging prompts, credentials or sensitive tool payloads.
Evaluation should include late user corrections, two tools completing in different orders, a socket reconnect, duplicate delivery, a timeout, a tool that succeeds after cancellation was requested, a missing previous response and an approval revoked before dispatch. Score whether the final answer reflects authoritative state, whether side effects happened once and whether the user update affected all remaining eligible work.
Compare the async design with a synchronous baseline. Parallelism is valuable only if it reduces completion time or improves the interaction without increasing wrong actions, wasted work or operational complexity beyond its benefit.
When not to use these features
Use ordinary synchronous function calling for a short linear workflow with fast tools and strict ordering. Use a background job or workflow engine when the user does not need an interactive stream. Use a deterministic service or SQL query when the path does not require model reasoning. Mid-turn steering is unnecessary for one-shot extraction whose input is fixed before execution.
Avoid async tools when external systems cannot provide idempotency or status reconciliation, when policy requires serial approvals, or when the team cannot observe pending work. Avoid steering for operations that the interface presents as atomic unless the product can explain partial completion honestly.
Production rollout checklist
Begin with read-only async tools and shadow the scheduler decisions. Add one durable pending-work table, explicit deadlines and idempotency before enabling side effects. Test WebSocket reconnects and `previous_response_not_found`. Then enable steering for output preferences and scope reductions before allowing it near mutating workflows.
Roll out to a small cohort with a synchronous control group. Define rollback thresholds for duplicate effects, uncertain outcomes, abandoned pending jobs, latency regression, retry amplification, cost per verified completion and user-correction failures. Keep a kill switch that turns async tools back into synchronous calls without changing authorization or schemas.
The core design rule is simple: concurrency and steering change the schedule, not the source of truth. Durable application state owns jobs and business effects; trusted policy owns authority; the model plans and explains within those boundaries. Connect this runtime to GPT‑6.1 Sol routing, MCP agent trajectory evaluation and durable AI agent execution.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




