Skip to content

Team isolation (isolation: shared / dedicated)

Backlog tasks 90ad7902 and d74d9c67. docs/design.md's template example:

isolation: shared            # or: dedicated workers and NATS account

Two separable guarantees are bundled under "dedicated" (design doc "Transport" → "Rules": "Teams with dedicated workers also get their own NATS account"):

  1. A worker process not shared with other teams -- implemented today, this doc.
  2. That team's own NATS account (credential/subject-permission isolation, not just the agent-scoped MCP token every team already gets) -- partially implemented (backlog task d74d9c67, own scope, not part of this doc). contracts/a2a-nats.md's "Dedicated NATS accounts" section has the full design. Landed so far: kapelle_controller.natsauth/nats_provisioning mint and store a real operator → account → user JWT chain in OpenBao, and TeamService.create()/archive() provision/revoke it automatically for isolation: dedicated teams. Not yet landed: an operator-mode NATS server deployment for these credentials to actually authenticate against (dev compose stays plain-auth; this needs its own profile), and worker-side wiring to fetch its own minted user's .creds and connect with them instead of the shared deployment's plain auth. Until both of those land, a dedicated team's minted NATS credentials sit in OpenBao unused -- see "What this does NOT give you (yet)" below.

What isolation: dedicated actually does today

Nothing provisions or starts anything automatically -- an operator sets this up by hand (see below). What the controller does do, on its own, once a team's template sets isolation: dedicated: kapelle_controller.scale_out.ScaleOutPolicy._check_dedicated_worker_ present checks, once per check_all pass (default every 60s, same loop scale-out itself runs on), whether a pinned durable JetStream consumer exists for each of that team's roles. If one doesn't -- the common case until an operator actually starts a pinned worker for it -- it logs a warning and posts a one-time notice to the team's home room ("Role '' is configured for dedicated isolation, but no pinned worker process is running for it..."), edge-triggered like a budget notice: one notice per disappearance, not one every tick, and it fires again if a previously-present pinned worker later goes away.

Before this, a misconfigured isolation: dedicated team (one with no worker actually pinned to it) was indistinguishable from a genuinely idle one -- work would queue on its A2A_RPC subject forever with no signal anywhere that anything was wrong.

How to actually start a dedicated worker

Worker "pinned mode" already exists and is unchanged by any of this -- kapelle_worker.a2a.config.WorkerConfig.from_env() reads KAPELLE_TEAM/ KAPELLE_ROLE: unset (both) is shared mode (the deploy/compose default, one process serves every team's traffic on a wildcard consumer), set (both) pins the process to that exact (team, role)'s own durable consumer (kapelle_contracts.subjects.rpc_queue_group/rpc_consumer_ config) -- mixed (one set, one not) is a startup error.

KAPELLE_TEAM=<team> KAPELLE_ROLE=<role> uv run --package kapelle-worker python -m kapelle_worker.a2a.main

One process per role that needs dedicating -- KAPELLE_TEAM/ KAPELLE_ROLE pin a single (team, role) pair, not a whole team's every role at once.

The hard constraint: pinned and shared workers can't share one A2A_RPC stream

A2A_RPC is a NATS JetStream work-queue stream. A work-queue stream rejects a second durable consumer whose filter subject overlaps an existing one ("filtered consumer not unique on workqueue stream", err_code=10100 -- see kapelle_worker.a2a.transport.nats_server. NatsA2AServer's own docstring). Shared mode's wildcard consumer filters on team.*.a2a.*.rpc -- every team's subject, including a "dedicated" one's. So a pinned consumer for a dedicated team's role and any shared-mode worker's wildcard consumer cannot both bind to the same A2A_RPC stream -- whichever binds second fails outright, hard error, not silent double-delivery.

This means a dedicated team's worker needs its own NATS server (or at minimum its own separately-configured A2A_RPC stream on a different NATS deployment) from whatever shared-mode workers use -- a deployment- level decision, not something a flag in this codebase can arbitrate for you today. KAPELLE_NATS_URL (or however the worker process is pointed at NATS in your deployment) must differ between a dedicated team's pinned worker(s) and the shared-mode fleet.

What this does NOT give you (yet)

A "dedicated" team's traffic is still only as isolated as putting it on its own NATS deployment makes it -- there is no per-team credential/ subject-permission boundary INSIDE a single NATS server yet, even though the credentials for one now exist in OpenBao the moment the team is created (see above). Two pieces are still missing before that boundary is real:

  1. An operator-mode NATS server for those credentials to authenticate against. contracts/a2a-nats.md's "Operator mode is all-or-nothing on a server" explains why this can't just be a config flip on the existing dev NATS -- it needs its own deployment/profile.
  2. Worker-side credential consumption. Nothing yet reads a dedicated team's minted .creds (kapelle_controller.openbao.layout. team_nats_user_creds_path) and passes it to nats.connect(user_ credentials=...) instead of the shared deployment's plain auth.

Until both land, team-boundary enforcement inside one shared NATS deployment is still the agent-scoped MCP token (contracts/a2a-nats.md), exactly as it is for every shared-mode team.