AI developer tools · Claude Code

Claude Code Mods: event middleware for safer AI coding

Anthropic’s new Mods can rewrite Claude Code events and UI. Here is how to use them safely, test their order and choose between Mods, hooks, Skills and MCP.

TOPIC HUBAI, RAG & Vector Search
Editorial concept of a Claude Code event chain with prompt, tool, permission, policy and interface stages connected as modular middleware.
An editorial interpretation of the topic, followed by a practical execution diagram.

Claude Code Mods are TypeScript or JavaScript functions that intercept Claude Code events and can observe, rewrite, replace or wrap behavior. Released by Anthropic on 1 October 2026 for Claude Code 2.1.287 or later, they can change prompts, tool calls, permission flows and interface components. That power makes Mods useful for organization-specific developer workflows, but also makes them privileged code: Anthropic states that Mods run with the user's permissions and are not sandboxed. Treat them like production dependencies, not prompt snippets.

This guide explains the official launch and documentation reviewed on 5 October 2026. It separates documented behavior from engineering recommendations. Examples are illustrative; verify the generated type declarations and APIs that ship with the exact Claude Code version you deploy because Anthropic says the API can change between releases.

What Anthropic released on 1 October 2026

Anthropic's launch announcement describes a Mod as a small TypeScript function that changes how Claude Code behaves or looks. A Mod can rewrite a prompt, block or retry a tool call, respond to a permission request, redact tool output, replace a built-in feature or draw new terminal and desktop UI. Mods ship inside Claude Code plugins, so installation, distribution and administrative plugin controls remain the packaging layer.

The important change is not merely “custom UI.” Existing settings hooks execute commands or endpoints at lifecycle events; Skills provide reusable instructions; MCP servers provide external tools. Mods register functions inside the Claude Code event pipeline. They can call the next handler, change the event passed to it, or return an answer without continuing. That makes them middleware around an agent runtime.

The feature is available in Claude Code CLI and desktop. The official Mods overview requires Claude Code 2.1.287 or later and says Mods are enabled by default. UI rendering depends on the surface: hooks can still execute in some non-interactive sessions where custom panes or bands are not displayed. A production Mod must therefore keep its policy correct when the UI is absent.

Think in middleware chains, not isolated callbacks

The official getting-started guide shows the core contract as `register(on)` plus handlers that receive the Mods API, event data and `next`. Calling `next(e)` passes the event to the next plugin and eventually to Claude Code. Code before `next` is a pre-stage; code after it is a post-stage; returning a modified event rewrites behavior; returning a result without `next` answers or blocks the event.

This composition model creates order-dependent behavior. Anthropic's announcement says multiple Mods handling the same event run in load order: the first loaded sees the request first and the result last. That is equivalent to nested middleware. A redaction Mod loaded outside an audit Mod may cause the audit to see redacted data on the request path but raw data on a different return path, depending on implementation. Document the intended chain rather than assuming every Mod is independent.

A useful internal contract classifies each handler as one of four types: observer, transformer, policy gate or UI adapter. Observers should not modify events. Transformers should state exactly which fields they may rewrite. Policy gates should fail closed when their evidence is unavailable. UI adapters should never be the only enforcement layer. This classification makes reviews and tests narrower.

Choose Mods, hooks, Skills or MCP deliberately

Use a Mod when the requirement must participate in Claude Code's in-process event chain: add a pane, rewrite a tool call, implement a custom command, maintain session state or wrap another handler. Use a settings hook when a shell script, HTTP endpoint or prompt can block, allow, log or enrich a lifecycle event without custom UI. The hooks reference documents session, prompt, permission, tool, task, compaction and model-switch events and their decision controls.

Use a Skill when the problem is reusable expertise or procedure—how to review a migration, generate a report or follow a team's conventions. A Skill changes what Claude knows and does; it should not impersonate a security boundary. Use MCP when Claude needs a typed interface to an external system such as a deployment service, issue tracker or database. MCP expands the tool surface, while a Mod can observe or govern how Claude Code uses that surface.

These mechanisms can coexist inside one plugin, but avoid building a single opaque package that mixes policy, knowledge, connectivity and UI without boundaries. Keep the Skill readable, the MCP server independently authorized, the policy Mod small and the UI Mod optional. A team should be able to disable the presentation layer without disabling an enforcement control.

A production Mod chain keeps managed policy outside user customizations, makes order explicit and leaves final authorization with the resource-owning service.
A production Mod chain keeps managed policy outside user customizations, makes order explicit and leaves final authorization with the resource-owning service. Open for a larger view

Treat every Mod as privileged supply-chain code

Anthropic explicitly warns that Mods are not sandboxed and run with the same access as Claude Code. The official security section says a loaded Mod can read and write files available to the user, start programs, make network requests, read environment variables and settings, see prompts and tool calls, alter a session, approve some tool calls and spend model usage. Claude Code's Bash sandbox does not automatically contain processes that a Mod starts.

The practical security model is therefore closer to an IDE extension or package-manager dependency than a prompt template. Pin the plugin source and version, review the manifest and hook module, validate what events and Mods API calls it declares, and deploy through an allowlisted marketplace. Do not install arbitrary Mods merely because their output looks useful. Separate developer experimentation from managed production workstations.

Run `claude plugin validate ./plugin-path` before loading a Mod. Anthropic documents that validation can list registered hooks and requested calls without executing the Mod. This is valuable static evidence, not a proof of safety: code review, dependency review and runtime egress controls are still necessary. Capture the validation output as a release artifact and diff it between versions.

Secrets require special treatment. Prefer short-lived credentials outside the interactive environment, scope tokens to the smallest resources, and redact tool results before they reach the model or logs. A Mod that promises to redact secrets still sees them. If a requirement cannot tolerate the Mod process seeing a credential, move the sensitive operation behind a separately authorized service rather than relying on in-process redaction.

Put organizational controls first and make order explicit

Load order is part of the security policy. Anthropic documents a managed `sec-default` Mod for supported organization setups and says it loads before user-installed Mods to restrict risky overrides. If administrators replace the managed first-loaded list, Anthropic recommends retaining `sec-default`. Regardless of product defaults, a team should inventory the effective order on every managed surface.

Design the outermost organizational Mod as a small policy kernel. It should establish immutable constraints, record a correlation ID, derive the environment class and deny actions that violate organization rules. User productivity Mods may run inside that boundary. Audit logic should record both the original intent and the final executed tool call when policy permits, without logging secrets.

Never make a warning pane the only safeguard. Anthropic's example “Blast Radius” can hold commands and ask for confirmation, but its own guide says command-text classification is a safety net rather than a permission system: aliases, scripts and command substitution can bypass string matching. Enforce hard rules through permission policy, environment isolation, least-privilege credentials and server-side authorization.

Design handlers for reliability and bounded latency

An event interceptor sits on the critical path. A slow Mod increases perceived agent latency; a hung policy handler can stop work; an exception may leave the user unsure whether a tool ran. Define a time budget for each handler, use cancellation signals, and keep remote calls out of synchronous paths unless the policy genuinely depends on them. Cache stable configuration, but never cache authorization longer than its validity.

Every transformer should be deterministic for the same event and policy version. Attach a Mod version and decision reason to traces. When retrying a tool call, distinguish “not executed” from “execution result unknown”; otherwise a retry can duplicate an external side effect. For mutating tools, pass an idempotency key to the downstream service or reconcile state before retrying.

Choose fail-open or fail-closed per handler. A decorative token meter can disappear when it fails. A production safeguard should normally fail closed, but it needs an emergency recovery path that does not silently disable all controls. Test safe mode and organization-managed behavior before relying on them during an incident.

State must survive hot reload correctly. Anthropic's tutorial recommends host-managed `$.state` for data that should persist across module reloads and shows typed state declarations. Avoid module globals for durable session decisions; a reload resets them. Keep stored state bounded and versioned so an upgrade can migrate or ignore incompatible values.

Example: a production-change gate with auditable decisions

The following simplified pattern guards Bash tool calls that appear to target production. It does not parse shell semantics and is not a complete security control. Its purpose is to show a bounded policy gate: derive risk, require an explicit approval service for high-risk operations, preserve an audit record and deny on uncertainty.

export function register(on) {
  on('tool.call', { tool: 'Bash' }, async ($, event, next) => {
    const decisionId = crypto.randomUUID();
    const risk = classifyProductionRisk(event.command);
    if (risk === 'none') return next(event);

    const decision = await withDeadline(
      $.http.request('https://policy.example/decide', {
        method: 'POST',
        body: { decisionId, risk, commandDigest: sha256(event.command) },
      }),
      1500,
    );

    if (decision?.allow !== true) {
      return { deny: `Production policy denied \${decisionId}` };
    }

    const result = await next(event);
    await auditBestEffort({ decisionId, result: summarize(result) });
    return result;
  });
}

In a real design, the policy service authenticates the workstation, authorizes the user and environment, and never accepts client-declared roles. The command digest supports correlation without storing raw secrets. The 1.5-second timeout is illustrative, not a recommendation. Side-effecting execution still needs server-side controls because a local Mod can be removed or altered by anyone who controls the machine unless organization management prevents it.

Validate, test and release Mods like software

Anthropic provides `claude plugin validate` and `claude plugin test`. The tutorial demonstrates tests against the real Claude Code runtime, including event stubs and UI mounting. Use unit tests for classifiers and transformations, chain tests for ordering, permission tests for deny and approve paths, cancellation tests, hot-reload tests and surface tests for terminal, desktop and headless sessions.

Add adversarial cases: nested shell commands, encoded arguments, huge tool results, prompt injection in tool output, network failure, stale policy, duplicated events and a Mod that does not call `next`. Test combinations, not only individual plugins. Two correct Mods can produce an unsafe composition when their order changes.

Pin the minimum Claude Code version and test against the version being distributed. Anthropic says generated type declarations inside `.claude-plugin/types/` are authoritative for that build and that the API can change between releases. Promote updates through a canary group, compare decision and latency telemetry, and retain a rollback package. Do not hot-reload unreviewed code on managed workstations.

When not to use a Mod

Do not use a Mod when a declarative permission rule can enforce the requirement more simply. Do not move external authorization into local TypeScript when the resource-owning service can enforce it. Do not use a Mod merely to teach Claude a checklist; that is a Skill. Do not wrap a remote business API directly when an MCP server with scoped credentials and auditable schemas is the clearer boundary.

Avoid Mods for controls that must hold outside Claude Code. A local interceptor cannot protect actions performed through another client, a direct API call or compromised credentials. It can improve the developer workflow and add defense in depth, but the system of record must remain authoritative.

For small personal conveniences, a status line or existing setting may be cheaper to maintain. A Mod is justified when event rewriting, persistent in-session behavior or integrated UI materially improves the workflow enough to carry the compatibility and security cost.

Production rollout checklist

Before rollout, identify the exact events and fields the Mod may observe or change; choose Mods versus hooks, Skills and MCP explicitly; pin Claude Code and plugin versions; review all dependencies; validate declared hooks and API calls; document load order; keep managed security controls first; scope credentials; define timeouts and cancellation; make mutation retries idempotent; bound state; test every supported surface; test Mod combinations and adversarial inputs; measure handler latency and denials; canary the release; keep rollback and safe recovery procedures; and enforce final authorization in the system that owns the resource.

For related architecture, read MCP agent evaluation, durable AI agent execution, OpenAI computer-use trust architecture and secure tenant context propagation.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectAI Action Studio