DevOps

AI agent permissions: Docker bakes them into the image

AI agent permissions now ship inside the OCI image: Docker's Sandbox Kit spec declares network, secrets and volumes. What it changes for agents in production.

AI agent permissions: Docker bakes them into the image

AI agent permissions: Docker bakes them into the image

Ask a platform team what their coding agent is allowed to do, and the honest answer is usually "whatever the token it was started with allows". AI agent permissions have lived in READMEs, local config files and people's heads. On September 24, at the opening keynote of WeAreDevelopers, Docker proposed a fix: the Sandbox Kit Specification, an Apache 2.0 open source format that records, inside an ordinary OCI image, which network hosts, credentials and volumes an agent may reach. The same day, Docker and the CNCF announced a partnership to turn it into a vendor-neutral standard.

The pitch is simple. The answer to "what may this agent do?" should travel with the agent, pinned to the same digest as its code, and it should be something a reviewer can read.

The context

A container isolates a process using the host kernel's namespaces and cgroups. A Dockerfile describes how an image is built. Neither says anything about what the software will need from the outside world once it runs: which domains it calls, which secrets it expects, which directories it mounts.

For a conventional web service, infrastructure fills that gap: firewall rules, secrets injected by the orchestrator, Kubernetes network policies. Agents are harder. A coding agent decides at runtime to call an API, push a branch or read a file. Its behaviour is not fixed at build time, so the boundary of what it can reach becomes the only defence that actually holds.

MCP (Model Context Protocol) standardised how an agent talks to its tools. What was missing was a standard way to declare what the agent is allowed to do with them. That is the gap the Kit targets.

Anatomy of a Kit

According to Docker's technical write-up, a Kit is a plain OCI image. Capability declarations sit in a manifest annotation, vnd.docker.sandbox.kit.descriptor. That design choice matters: you build a Kit with docker buildx build, push it to any registry, and sign and scan it with the tools you already run. No new artifact type, no new registry.

The spec defines two kinds:

  • Workload Kits supply the root filesystem and run the agent.
  • Mixin Kits are overlays that add a CLI, network rules, credentials or context.

A working environment is one workload plus any number of mixins, for instance a coding agent with a GitHub CLI mixin on top.

Typed, versioned permissions

Every declaration carries a type and a version. Docker's GitHub CLI example allows github.com and the GitHub API for a list of HTTP methods, then explicitly denies DELETE on /repos/**:

capabilities:
  - type: com.docker.sandbox/network-policy@2
    config:
      runtime:
        allow:
          - github.com
          - hosts: [api.github.com]
            methods: [GET, HEAD, POST, PATCH, PUT, DELETE]
        deny:
          - hosts: [api.github.com]
            methods: [DELETE]
            paths: [/repos/**]

deny wins over allow. Versions are independent per type, so network-policy@1 and @2 can coexist while runtimes catch up.

Secrets the agent never holds

The credential type tackles the problem every team wiring up agents runs into: giving an agent a token without letting it read, log or leak that token. With proxyManaged: true, the agent only sees a sentinel value. The runtime's proxy injects the real token into the Authorization header, and only for requests to the declared domain. The agent process never touches the secret.

Entries can be marked optional: true. If a required entry cannot be satisfied, the launch is refused. The default is to fail closed.

The host decides, the agent asks

The core principle: a Kit requests permissions, and the host decides whether to grant them. A Kit is not an authorisation. It is a request that can be read, diffed and approved.

Composition resolved at build time

When Kits are combined, coherence rules apply: exactly one workload, no duplicate providers (nothing silently shadows anything), and mixins ordered by their provides / requires dependency graph rather than by the order you typed them. Network rules union, hooks run in dependency order, licences union. A kind: set descriptor freezes an assembly, and an incoherent set fails at build time instead of at launch.

Updates that cannot quietly widen access

This is the part platform teams will care about most. The runtime records the normalised set of permissions it granted. A new Kit version that stays inside that set applies without prompting. Anything that widens access needs fresh approval, and removing a deny rule counts as widening. The approver sees a diff of authority, not a diff of code.

A reference runtime, local and in the cloud

The spec is meant to be portable, but it ships with an implementation. Docker Sandboxes is described as the first conforming runtime: each agent runs in a microVM with its own kernel, which puts the isolation boundary below the agent's reach, unlike a container that shares the host kernel. The CLI is sbx:

brew install docker/tap/sbx
sbx run ./hello --kit ./gh .

Docker also made Cloud Sandboxes generally available the same day. They run the same Kits under the same trust model on Docker-managed capacity, and a session moves either way with sbx move my-project --to cloud. Billing is per second, from $0.07 an hour for a 1 vCPU / 2 GiB microVM to $1.12 for 16 vCPU / 32 GiB. Sessions default to one hour and can run up to 24.

On the ecosystem side, Docker lists Kits built with AWS, Box, Datadog, Dynatrace, JFrog, NanoClaw, OpenClaw, Palo Alto Networks and Snyk. The CNCF partnership is about neutral governance, which is what other runtimes will need before they adopt the format.

What this means for AI engineering teams

Most teams industrialised their models before they industrialised their agents. The pattern is familiar: a coding agent launched with a personal GitHub token that has broad scopes, unfiltered egress, and a setup that differs from one laptop to the next. AI agent permissions exist, but they are implicit.

A Kit changes three things.

One source of truth. "What can this agent do?" is attached to a digest. It goes through code review, lives in the same repository as the image, and can be compared between versions.

A supply chain you already have. Because a Kit is an OCI image, signing, vulnerability scanning, admission policies and registry replication work unchanged. If you sign your images today, you sign your permissions too.

Controlled widening. "Any widening needs approval" turns an agent upgrade into an explicit security decision, which is exactly the guardrail security teams ask for before letting an agent near a production repository.

Keep expectations grounded. The spec is days old, and Docker Sandboxes is so far the only announced conforming runtime. Nothing yet guarantees that a Kubernetes-based platform or another sandbox vendor will enforce these declarations. And a Kit describes a perimeter; it does not replace reviewing what the agent actually did, logging, or evaluating its behaviour.

Our practical advice: even if you never run sbx, start writing the equivalent of a Kit for every agent you operate. List the hosts it really needs, the tokens and their narrowest scopes, the volumes. That inventory is the prerequisite for any AI agent permissions policy, whichever runtime you pick next year.

A starting checklist

If you want to turn that advice into a sprint task, here is the order we would follow with a client team:

  1. Inventory the agents. List every agent that touches code or data: IDE assistants, CI bots, internal copilots, scheduled jobs. Note who launches each one and with which identity.
  2. Map real egress. Run each agent for a representative week behind a logging proxy and record the hosts it actually calls. The gap between "what it can reach" and "what it uses" is your first reduction.
  3. Narrow the tokens. Replace personal tokens with service identities scoped to one repository or one API, and move them behind a proxy so the agent process never holds them.
  4. Write the deny rules first. Destructive methods on production resources, such as deleting repositories or dropping tables, belong in an explicit deny list that survives upgrades.
  5. Make widening a review event. Whatever the tooling, any new host, scope or volume should go through the same pull request process as a code change.

None of these steps depends on Docker. All of them make adopting a Kit, or any future equivalent, a short migration rather than a project.

In short

  • On September 24 Docker published the Sandbox Kit Specification, an Apache 2.0 format that declares an agent's permissions inside an OCI image.
  • Network, credential and volume entries are typed and versioned; deny beats allow, and proxy-managed secrets never reach the agent.
  • The host decides: any permission widening in an update requires new approval.
  • Docker Sandboxes (local microVMs) and Cloud Sandboxes (GA, billed per second) are the first conforming runtimes.
  • For AI teams, agent permissions become a versioned, signed, reviewable artifact.

Industrialising AI agents? SeedVision offers 3-5 day AI audits and 15-30 day production rollouts. See the packages or book a 30-min call.

Cover photo: Photo by Barrett Ward on Unsplash.