OpenAI introduced Dots on 30 September 2026 as always-on agents powered by GPT-6 Astra. OpenAI says a Dot has its own cloud computer, can work across connected tools, continue while a user's device is offline and learn from feedback over time. Those are product claims and documented capabilities, not independent evidence that every workflow will complete correctly.
The important engineering change is temporal. A request-response assistant owns seconds or minutes. An always-on agent owns work across pauses, events, schedules, changing permissions, expired sessions and human decisions. Prompt quality still matters, but the dominant failure modes move into distributed-systems territory.
This article treats Dots as a concrete product update and extracts architecture lessons for teams building long-running agents. It does not claim that Dots itself is a general developer API. OpenAI documents the Agents API, Agents SDK and Responses API as separate starting points for developers.
The product announcement changes the time boundary
The OpenAI launch post describes Dots as cloud-based, always-on agents that can connect to tools and remain reachable through ChatGPT, Slack or Teams. The Meet Dots guide adds that the agent has a cloud computer and browser and can continue working when the user's computer is off.
That turns “the agent run” into a poor unit of ownership. The product operation may outlive one model call, one browser session, one message channel and one compute environment. A durable task becomes the unit: a named outcome with authorized sources, current state, evidence, approvals, outputs and an explicit completion condition.
Do not infer an API contract from the consumer product. For application builders, OpenAI's agent guide distinguishes the managed Agents API from the Agents SDK and the Responses API. The Agents API manages a Codex harness and saved progress; the SDK gives the application more control over runtime, storage and approvals; Responses exposes lower-level model and tool orchestration. Choose based on who must own state and recovery.
The architectural lesson is portable: if the agent can wake up later, “continue” must mean reload verified task state—not ask a model to reconstruct reality from a long conversation.
Put every responsibility in a durable task ledger
OpenAI's tasks and memory documentation says a Dot can track multiple responsibilities, pause and wake, create background work, respond to time or supported events and continue across conversations. It also warns that a completed run does not by itself prove that the requested result was achieved or delivered.
Model the task as a state machine stored outside the model context. Useful states include `planned`, `ready`, `running`, `waiting_for_tool`, `waiting_for_approval`, `paused`, `verifying`, `succeeded`, `failed`, `cancel_requested` and `cancelled`. Transitions should be conditional writes with version checks, not free-form status text. A worker that resumes an old lease must not overwrite a newer decision.
Keep one immutable intent record: requested outcome, scope, allowed sources, forbidden side effects, notification policy and definition of done. Store mutable execution separately: current step, attempt, environment reference, permission snapshot, last heartbeat, pending approval and output references. This distinction lets the user change priorities without silently rewriting what an earlier action was authorized to do.
task_id: task_01J...
intent_version: 3
state: waiting_for_approval
step: send_customer_message
permission_snapshot: perm_94
input_evidence: [artifact_17, event_203]
idempotency_key: task_01J:send_customer_message:v3
lease_owner: worker_8
lease_expires_at: 2026-10-02T06:07:00+03:00
definition_of_done: message_delivered_and_loggedAppend task events rather than updating one opaque JSON blob. The event stream should explain who requested a change, which evidence the planner used, why execution paused and which postcondition was verified. It is an audit and recovery surface—not hidden chain-of-thought.
Separate conversation context, durable notes and operational truth
The Dots guide distinguishes conversation context, ChatGPT memory and a Dot's own saved notes. It also states that a delegated task receives relevant instructions and context but not automatically every conversation. This is a useful boundary: more memory is not always safer or more correct.
Create three stores with different rules. Conversation context is short-lived input selected for the current interaction. Preference memory contains reviewed facts such as writing style, usual notification threshold or an ongoing project's name. Operational truth contains task state, permissions, external resource IDs, approvals and verified results. Only the last store decides whether a mutation may continue.
Every durable memory entry needs provenance, scope, owner, sensitivity, creation time, review time and expiry. A note inferred from one private conversation must not be disclosed in a team channel simply because both reach the same agent. OpenAI explicitly notes that continuity across channels does not grant permission to disclose private information to another audience.
Use retrieval by task and audience, not a global “relevant memory” query. Before a message is sent to Slack, resolve the target audience and filter context accordingly. Before a worker resumes, load the task's immutable intent, current permission snapshot and verified artifacts; treat narrative summaries as hints.
Memory correction must be first-class. A user changing a preference should create a superseding record and invalidate affected future plans, without rewriting the audit history of completed actions.
Bind permissions to the action, environment and moment
The computers and apps guide separates the Dot's cloud browser, a connected personal computer and connected plugins. Cloud and local sessions are not interchangeable, and app permissions remain distinct from messaging channels. The controls guide describes custom rules for acting automatically, acting only after a request, asking before an action or handing the action back to the user.
In your own system, compile these inputs into an immutable permission snapshot when a risky step is planned. Include principal, tenant, environment, tool, action class, resource scope, data class, approval mode, expiry and policy version. Check again immediately before execution because tokens, plugin access, workspace policy and task intent can change while the agent waits.
Do not confuse authentication with authorization. An active browser session proves that a session can reach an account; it does not prove that this task may send a message, delete a file or approve a payment. Put deterministic policy enforcement between the model's proposed tool call and the tool adapter.
Narrow tools by business effect. `search_mail` and `draft_reply` are safer contracts than a general browser. `send_reply` should require a recipient, content digest, approved task version and idempotency key. Destructive operations should return a previewable plan and require a separate approval token bound to that exact plan.
Credentials must stay outside model-visible context and task logs. Pass credential references to a server-side resolver, redact tool errors and expire environment access after the relevant work. If execution moves from cloud to a connected computer, create a new environment binding; do not pretend the original browser state transferred.
Treat schedules and events as at-least-once inputs
The Dots documentation separates recurring schedules from active work and says connecting a source alone does not create event monitoring. This suggests two different resources: a trigger definition and a task execution. Keep them separate so canceling today's run does not silently delete tomorrow's schedule, and deleting the schedule does not leave an active mutation running.
Assume time and event triggers can be delivered twice, late or after a restart. Deduplicate on trigger identity plus occurrence, then create or attach to one task operation. If “new bug report” fires for an edited message, decide whether it updates the existing task or creates a new one. Store the source event ID and version instead of comparing free-form text.
Use a transactional outbox when a task state change must emit a notification or enqueue work. A database commit followed by a failed queue publish otherwise leaves the ledger and executor inconsistent. Consumers need idempotent handlers because retries can occur after the side effect succeeded but before acknowledgment.
Schedules need time zones, end dates, missed-run policy and concurrency rules. “Run every hour” is incomplete: should a slow run overlap, skip or queue? For an agent that may make changes, default to non-overlapping execution and reconcile the latest source state before acting.
Event filters are security boundaries. A keyword match in a public channel should not grant a task access to private repositories. Resolve source permissions and task permissions independently, then intersect them.
Verify outcomes instead of trusting completion signals
A tool returning success proves only that the call returned. A background agent saying “done” proves only that it stopped. Define postconditions at the business boundary: the document exists at the expected version, the deployment is serving the intended commit, the message has a provider delivery ID or the ticket is in the target state.
The Agents API architecture guide documents streaming and webhooks for progress and state changes, and notes that failures in lifecycle or function-tool handlers can interrupt progress. Treat these events as observations, not the only source of truth. Reconcile the task ledger against the target system before marking success.
Record four times: queue delay, active execution, time waiting for a human and time to verified outcome. A single duration hides whether the agent is slow, the environment is offline or an approval is blocking. Track retry counts, approval age, lease expiry, unknown-outcome writes, event deduplication and verification failures.
Every result should name its evidence. A code task can attach test output and the source revision. A research task can attach source URLs and retrieval dates. A mutation can attach a sanitized provider receipt and a read-after-write check. Do not store secrets or unrestricted personal data as proof.
User-facing status must distinguish working, waiting, blocked, completed and verified. If the agent needs a decision, show the exact action, expected effect, affected resource and what happens if the request expires.
Design stop, revoke and recover as separate operations
Stopping a worker, canceling a task, disabling a schedule, revoking an app and disconnecting a computer are not the same operation. The Dots documentation explicitly separates stopping active work from canceling a recurring schedule. Your architecture should make each lifecycle visible.
A cancel request first prevents new side effects, then signals active workers, expires pending approvals and moves the task toward a safe checkpoint. If a remote write may already have committed, the state is not “cancelled”; it is `reconciliation_required` until the target system is checked. Compensation is a business operation, not a generic rollback.
Revoking credentials should invalidate future resolutions immediately, but it cannot undo work already accepted by an external service. Record the last verified side effect and surface unfinished cleanup. When a connected computer goes offline, pause tasks that require it; do not silently reroute them to a cloud environment with different sessions or data access.
Recovery starts from the ledger and target-system evidence, not from the last generated message. Acquire a new lease, confirm permission and environment health, inspect uncertain steps, then continue from a postcondition boundary. Limit automatic retries and escalate repeated ambiguity to a human.
Test the ugly paths: duplicate triggers, approval after cancellation, worker crash after mutation, expired website session, plugin permission removed mid-run, connected computer offline and schedule deleted while execution is active. These scenarios determine whether the system is genuinely durable.
The practical decision
OpenAI Dots make the always-on-agent category concrete: a cloud agent can hold multiple responsibilities, coordinate background work, react to time or events, use connected environments and ask for human decisions. The launch is significant, but the marketing promise is not a production guarantee.
For teams building similar systems, the durable architecture matters more than the agent loop. Make the task ledger authoritative, separate preference memory from operational truth, bind permissions to each action and environment, deduplicate triggers, verify business outcomes and give cancellation explicit semantics.
A capable model can plan the next step. A trustworthy product must still prove that the right task ran with the right authority, against the right state, and either reached a verified result or stopped safely.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
- OpenAI — Introducing Dots, published 30 September 2026 and verified 2 October 2026
- OpenAI Developers — Meet Dots, verified 2 October 2026
- OpenAI Developers — Dots tasks and memory, verified 2 October 2026
- OpenAI Developers — Dots computers and apps, verified 2 October 2026
- OpenAI Developers — Dots controls, verified 2 October 2026
- OpenAI Developers — Choosing Agents API, Agents SDK or Responses API, verified 2 October 2026
- OpenAI Developers — Agents API architecture, verified 2 October 2026
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




