Dev environment¶
just dev-up starts everything a developer needs except the VM host pieces
(vmd, Firecracker, the credential gateway), which run natively on Linux
with KVM (see docs/design.md §8) — Docker Compose can't sandbox a
microVM host. Everything else (NATS, Postgres, LiteLLM, OpenBao, the OTEL
collector, Jaeger, and the optional RustFS and Mattermost profiles) is plain
containers, defined in deploy/compose/docker-compose.yaml.
Prerequisites¶
- Docker and Docker Compose v2 (
docker compose version). - For
vmd/Firecracker later (not covered by this compose file): a Linux host with/dev/kvm, andsudofor creating network namespaces, taps and nftables rules. just(see the rootjustfile).
First run¶
- Copy
deploy/compose/env.exampletodeploy/compose/.envand fill it in..envis gitignored (root.gitignore,*.env) — never commit it. just dev-up— brings upnats,postgres,openbao(dev mode, in memory; see "Keeping OpenBao's data across restarts" for the durable variant),otel-collector,jaeger,litellm. Verified in this repo to reach all-healthy in well under a minute on a warm image cache.just dev-seed— creates the NATS streams anda2a-cardsKV bucket (contracts/a2a-nats.md), and the OpenBao policies, KV layout and AppRole auth from backlog task 13 (kapelle_controller.openbao.seed) — prints a dev RoleID/SecretID pair for the single "dev" host'scredgwandvmdAppRoles.- Check it worked:
docker compose -f deploy/compose/docker-compose.yaml --env-file deploy/compose/.env exec nats-box nats --server nats://nats:4222 stream ls(or bring upnats-boxand runnats stream lsinside it) — should listA2A_RPC,A2A_CANCEL,A2A_TASKS,A2A_ACTIVITY,KV_a2a-cards.curl http://localhost:4000/health/liveliness— LiteLLM.curl http://localhost:8200/v1/sys/health?standbyok=true— OpenBao ("sealed":false).- open
http://localhost:16686— Jaeger UI.
Keeping OpenBao's data across restarts¶
By default openbao runs as server -dev: in memory. A recreate of the
container (a host reboot, an image bump) empties it, and with it the
secret document of every agent: their LiteLLM and MCP credentials and, for
Firecracker agents, the key that decrypts their data disk (which cannot be
recovered). just doctor names the teams that lost theirs ("agent documents
in the secret store").
To keep the data, set OPENBAO_STATIC_UNSEAL_KEY in deploy/compose/.env
with this command (bash and fish; it does not print the key, and the leading
newline keeps it on its own line when .env does not end with one):
printf '\nOPENBAO_STATIC_UNSEAL_KEY=%s\n' "$(openssl rand -base64 32)" >> deploy/compose/.env
From then on just adds deploy/compose/openbao-persistent.yaml:
raft storage on the openbao-data volume (/openbao/file, the directory the
image ships owned by its non-root user, so the volume needs no setup) and
auto-unseal with that key, config in deploy/compose/openbao/openbao.hcl.
Steps, once:
- Add the key, recreate the container (
just dev-up). The nextjust dev-upafter adding the line recreates OpenBao with an empty store (the in-memory data is not carried over), so add the key at a moment agreed with whoever operates the stack. just dev-upalso runsjust dev-openbao-init(idempotent;kapelle_controller.openbao.dev_init): a new store is initialized (one recovery share, not kept), gets the fixed dev token (OPENBAO_DEV_ROOT_TOKEN, the id every service and test already uses: a root-policy token without expiry, created without a parent), thesecret/KV v2 mount that dev mode used to provide, and the initial root token is revoked.just dev-seedas before. The NATS operator identity and the PKI root it writes now stay the same across restarts.
The server unseals itself after every restart; a store that stays sealed
means the key in .env is not the one it was initialized with. Losing the key
makes the volume unreadable. Back up with bao operator raft snapshot save
before an OpenBao image bump. The isolated test projects
(deploy/compose/tests/) never use the override: their openbao stays the
in-memory dev server.
Pointing LiteLLM at a model¶
Every model comes from the Aiven AI Gateway (https://ai.aiven.io,
OpenAI-compatible; decided 2026-09-24, when the DSV4 dev endpoint and
the LOCAL_LARGE_* placeholder were removed). deploy/compose/litellm/
config.yaml holds one pass-through entry per model the gateway serves
plus the role aliases teams/templates/*.yaml refer to (backlog task
22, design doc §12) -- frontier, frontier-large, frontier-xlarge,
fast, open-coder, and the compatibility names local-coder/local-
large for older docs and fixtures. deploy/compose/litellm/aliases.yaml
is the source of truth for which alias points at which model; just
litellm-sync-aliases regenerates config.yaml's model block from it
and the gateway's live /v1/models (see "Model providers" below).
The gateway needs AIVEN_AI_BASE_URL/AIVEN_AI_API_KEY, and .envrc
is the single, structural source for these -- not deploy/compose/.env
(agents/teammates may not edit that file; it's the user's alone), and
not the calling shell's own environment either. Every compose invocation
in this repo (just dev-up and friends, the eval/e2e and Jaeger
integration test fixtures) runs through direnv exec
/home/mbocevski/dev/factory docker compose ... specifically so it always
picks these up regardless of whether the invoking shell itself has
.envrc loaded -- a bare docker compose up -d run from a shell that
never loaded it sees them as empty and force-recreates every container
whose env depends on them (this is what repeatedly knocked over the
shared dev LiteLLM container when several sessions were running tests
concurrently). Always go through just dev-up/one of the direnv
exec-wrapped fixtures, or prefix your own docker compose command with
direnv exec /home/mbocevski/dev/factory yourself -- never call docker
compose bare against this file. just dev-up also fails fast with a
clear message if .envrc hasn't set AIVEN_AI_BASE_URL/
AIVEN_AI_API_KEY, rather than silently starting LiteLLM with a broken
model config. After any alias change, restart LiteLLM (just dev-up
again, or direnv exec /home/mbocevski/dev/factory docker compose ...
restart litellm).
Every role in a team template must resolve to an alias with a working
endpoint before an agent using that role can actually run a conversation
-- feature-team.yaml uses open-coder for the coder and frontier
for planner/reviewer, all served by the gateway.
Model providers: Aiven AI Gateway¶
config.yaml also has a second, real inference provider wired in: the
Aiven AI Gateway (OpenAI-compatible; GET /v1/models
lists every model it currently serves). Set, in .envrc:
AIVEN_AI_BASE_URL=https://ai.aiven.io
AIVEN_AI_API_KEY=<your key>
deploy/compose/litellm/aliases.yaml is the source of truth for which
semantic alias points at which Aiven model -- not config.yaml itself,
which is generated from it. As of the user's 2026-09-13 decision:
frontier: claude-sonnet-5 # planner/reviewer default
frontier-large: claude-opus-5
frontier-xlarge: claude-fable-5-1
fast: claude-haiku-4-5
open-coder: qwen3-coder-30b
Every model in Aiven's own catalog ALSO gets its own pass-through alias
(its own id, e.g. claude-opus-4-8, gpt-5.6-sol, qwen3-32b, ...), so
a caller can address a model directly instead of through one of the
named aliases above. frontier replacing the old FRONTIER_BASE_URL-
pointed placeholder means planner/reviewer roles have a real, working
default without any extra setup beyond the two env vars above.
To change an alias, edit aliases.yaml, then:
just litellm-sync-aliases
This regenerates config.yaml's Aiven block (between its own # >>>
BEGIN GENERATED >>>/# <<< END GENERATED <<< marker comments -- never
hand-edit between those) from aliases.yaml + Aiven's live /v1/models,
and REFUSES to write anything if any alias's target model isn't in that
live list anymore (the "does it actually resolve on the provider" check,
done before config.yaml is touched). It prints the exact docker
compose restart litellm command afterward -- run that yourself; the
recipe never restarts the shared dev container automatically (a restart
mid-run can kill another teammate's in-flight work). Then just doctor's
"litellm model aliases" check (which also reads aliases.yaml directly,
so it flags an alias you added before you've even regenerated anything)
should go green.
Per-token prices for the Aiven models come from
deploy/compose/litellm/pricing.yaml (keyed by Aiven's own model id, not
alias -- null until a real price is filled in, since LiteLLM can't
compute spend for a model with no price and a null avoids a fabricated
0-cost budget that would silently never trip); just litellm-sync-aliases
picks up new prices there too.
One non-obvious detail if you ever touch generate_aiven_models.py
itself: api_base is baked in as a literal https://ai.aiven.io/v1
string, not an os.environ/AIVEN_AI_BASE_URL reference -- LiteLLM's
OpenAI-compatible client appends /chat/completions onto api_base
verbatim (no automatic /v1), and os.environ/VAR substitution replaces
a field's entire value, so it can't concatenate a literal /v1 suffix
onto an env var either; the bare base URL 404s every real request
("fault filter abort" from Aiven's own edge) even though the model name
and config are otherwise correct. api_key stays a live env-var
reference (it's a secret; api_base isn't).
Verified live (2026-09-13): just doctor's own "litellm model aliases"/
"litellm upstream providers" checks pass, and a real one-token completion
against fast and against local-coder both return PONG.
If you hand-craft a raw /v1/chat/completions call against an Aiven
alias for manual testing, use max_completion_tokens, not max_tokens
-- confirmed live: Aiven's endpoint rejects the older max_tokens field
name OUTRIGHT regardless of its value ("Unrecognized request argument
supplied: max_tokens", tested at both 1 and 16384), not a minimum-
value or range check. This is NOT a regression in the real agent path:
the OpenHands SDK's own LLM client (what executor.py's max_output_
tokens= -- backlog cf1b54c -- actually drives) already sends the
current max_completion_tokens field, confirmed by a real completion
succeeding with it and failing only when a hand-rolled request used the
deprecated name instead.
The gpt-5.6-* models (frontier points at gpt-5.6-sol since
2026-09-28) additionally enforce a minimum output limit: a request with
max_completion_tokens: 10 comes back HTTP 400 integer_below_min_value
(on max_output_tokens), while 200 succeeds. Use a limit of a few
hundred or leave it out for a manual probe; the worker's own 16384 is
unaffected.
Storage (profile: storage)¶
just dev-up-storage starts RustFS (S3-compatible object storage, for the
snapshot/data-disk backups in docs/design.md §8) — not MinIO: MinIO's
last container image release was 2025-10 and the project has been in
maintenance mode since its AGPL relicensing fallout, so it no longer meets
this repo's "latest, actively maintained image" bar (CLAUDE.md). RustFS
(Apache-2.0, Rust, S3-compatible) is pinned to its GA release (1.0.1),
never to latest/rc floating tags. A data volume written by the earlier
1.0.0-rc.6 image is migrated in place on first start; if that ever fails,
remove the volume and run just dev-seed-storage again (dev data is test data).
Console at http://localhost:9001, S3 API at http://localhost:9000
(same ports MinIO used), credentials from RUSTFS_ACCESS_KEY/
RUSTFS_SECRET_KEY in .env. just dev-seed-storage proves the S3 API
actually works: an ephemeral amazon/aws-cli container puts and gets an
object through it and diffs the round trip.
What a test run is allowed to touch¶
Backlog task 55ab4f2d: several integration tests used to start a throwaway
Docker container, or write into the shared dev stack's NATS/OpenBao/
Postgres/LiteLLM, on nothing more than "the port happens to be reachable" --
Docker being installed, or a just dev-up stack happening to be running, is
not consent. Four classes now each need their own explicit opt-in
(contracts/python/src/kapelle_contracts/testing_consent.py), checked
before the test even asks whether the thing it needs is actually there:
| Variable | The test... | Set by |
|---|---|---|
KAPELLE_TEST_CONTAINERS=1 |
starts its own throwaway container (an ephemeral port, torn down by the test itself) | just test-integration, just demo/e2e, CI's python job, nightly |
KAPELLE_TEST_DEV_STACK=1 |
connects to a FIXED port of the shared dev stack (127.0.0.1:4222 NATS, :8200 OpenBao, :5432 Postgres, :4000 LiteLLM) and writes real data into it |
just test-integration, just demo/e2e, nightly (it runs just dev-up first); not CI (no dev stack there) |
KAPELLE_TEST_LIVE_PLATFORMS=1 |
calls a real external platform (Nevia, GitHub, Jira, Linear, Mattermost, Slack) over the network, ON TOP OF that platform's own existing credential check | nobody -- given by hand, for one run, because every live action on an external platform needs its own go |
KAPELLE_TEST_LONG_RUNNING=1 |
not about what it touches -- its own wall-clock sleep runs minutes, not seconds, checked IN ADDITION to whichever other class(es) it also needs (e.g. CONTAINERS) |
nobody -- given by hand; just test-integration never sets it, so a slow test stays out of the ordinary run even when its other consent is already given |
just test-unit and a bare pytest <anything> set none of these, so they
never touch a container or the shared dev stack, whatever happens to be
running or installed. just doctor's own reachability checks (Mattermost,
LiteLLM upstream providers, ...) are a different thing -- they only ever
read, never write, and were never gated by this.
eval/e2e/ is its own case: not in the root pyproject.toml's testpaths,
so just test-unit/a bare pytest never reaches it regardless -- but its
own conftest.py starts a throwaway NATS container (isolated_nats_
container) and brings up (and writes to) the shared dev stack (dev_stack),
gated inside those two fixtures rather than at the conftest module's own
top level (pytest.skip(..., allow_module_level=True) inside a conftest.py
crashes pytest's own config/plugin-loading phase instead of skipping
cleanly -- confirmed with a standalone repro; the fixture form is the safe
one). An explicit pytest eval/e2e/... needs both variables set by hand;
just demo (aliased just e2e) sets both itself, and the recipe fails
outright (parsing its own eval/report/out/e2e-junit.xml) if every single
test skipped -- pytest's own exit code alone would read that as a pass. A
by-hand pytest eval/e2e/... does not have that check: nothing reads the
junit file back for you, so an all-skipped run there still exits 0.
Go has the equivalent gate already, just enforced differently: every real-
sudo, real-Firecracker-boot or real-OpenBao Go test in vmd carries
//go:build integration (three files that write into the shared dev OpenBao
were missing it until this task and had no gate at all -- vmd/cmd/vmdctl/
openbao_integration_test.go, vmd/internal/pkicert/integration_test.go,
vmd/internal/openbao/pki_test.go). go_packages := "./vmd/... ./credgw/..."
and every just/CI invocation of go build/go vet/golangci-lint/
go test uses that bare path with no -tags flag, so every tagged file is
excluded from all of them at COMPILE time. Run one by hand, against a real
dev OpenBao (just dev-up first) or a real KVM host, with the tag:
KAPELLE_TEST_DEV_STACK=1 go test -tags=integration ./vmd/internal/openbao/...
KAPELLE_TEST_DEV_STACK=1 go test -tags=integration ./vmd/internal/pkicert/...
KAPELLE_TEST_DEV_STACK=1 go test -tags=integration ./vmd/cmd/vmdctl/...
KAPELLE_TEST_DEV_STACK=1 go test -tags=integration -run TestEnsure ./vmd/... # real Firecracker boot
credgw's own *_integration_test.go files need no gate beyond the next
paragraph: each starts a real iron-proxy binary bound to loopback
addresses it owns and t.Skips if the binary isn't available
(credgw/scripts/fetch-iron-proxy.sh).
Go tests without the tag that still reach beyond the process use the same
two consent variables as the Python suites, through internal/testconsent
(same variable names and skip text as kapelle_contracts.testing_consent),
checked at the top of the test's own require... helper before anything is
connected to: KAPELLE_TEST_DEV_STACK=1 for the RustFS-backed tests
(vmd/internal/objectstore, the disk backup tests in vmd/internal/manager:
the fixed 127.0.0.1:9000, the buckets kapelle-vmd-test and
kapelle-vmd-manager-test) and KAPELLE_TEST_CONTAINERS=1 for credgw's
OpenTelemetry collector test (one throwaway container). The tagged tests
that write into the shared dev OpenBao, Postgres, NATS or RustFS check
KAPELLE_TEST_DEV_STACK too, on top of the tag. A plain go test and the
CI go job's RustFS tests therefore skip with opt-in: set
KAPELLE_TEST_DEV_STACK=1 (...); the CI job sets KAPELLE_TEST_CONTAINERS=1
(it has docker and no RustFS). On purpose, on a dev box:
KAPELLE_TEST_DEV_STACK=1 KAPELLE_TEST_CONTAINERS=1 go test -race -count=1 ./vmd/internal/objectstore/ ./vmd/internal/manager/ ./credgw/manager/
just test-integration # includes the same Go run with both variables
Surfaces (profile: surfaces)¶
The surfaces profile holds Mattermost, the chat surface of this stack
(just dev-up-mattermost, below).
Mattermost¶
just dev-up-mattermost starts Mattermost
(mattermost/mattermost-team-edition) and its own Postgres. It listens on
127.0.0.1:8065 only; the public way in is the edge profile below.
Variable (deploy/compose/.env) |
Meaning |
|---|---|
MATTERMOST_SITE_URL |
The URL people open, default http://localhost:8065. Mattermost builds every link from it and checks the browser's WebSocket against it. |
MATTERMOST_DB_PASSWORD |
Database password, URL-safe characters only. Set it before the first start. |
MATTERMOST_PORT |
Host port on loopback, default 8065. |
Sign-up is closed (EnableOpenServer=false): people join by invitation
(Mattermost: main menu, "Invite people"). No mail server is configured, so
use the invitation link, not the invitation by email.
just mattermost-seed creates what Kapelle needs, through mmctl --local
inside the container and the REST API on 127.0.0.1:8065:
MATTERMOST_ADMIN_USERNAME=you MATTERMOST_ADMIN_EMAIL=you@example.org \
MATTERMOST_SITE_URL=https://chat.example.org just mattermost-seed
| Created | Notes |
|---|---|
Team kapelle |
Private: joining needs an invitation. |
| Administrator | First password in ~/.config/kapelle/mattermost-admin-password (mode 0600). Change it after the first login. |
Bot kapelle |
Its access token goes into the file below. |
Slash command /kapelle |
Calls http://host.docker.internal:8101/webhooks/mattermost/command, the gateway's webhook listener (MATTERMOST_COMMAND_URL changes it). |
The script prints no secret. It writes ~/.config/kapelle/mattermost.env
(mode 0600) with KAPELLE_MATTERMOST_URL, KAPELLE_MATTERMOST_TEAM,
KAPELLE_MATTERMOST_BOT_TOKEN and KAPELLE_MATTERMOST_COMMAND_TOKEN: copy
the four lines into .envrc. Running it again changes nothing; it needs
the same MATTERMOST_ADMIN_USERNAME every time. just mattermost-role-bots
below reads the same variable: when the administrator is not called admin,
run it with MATTERMOST_ADMIN_USERNAME=you too (it mints and revokes a
short-lived token for that account; unset, it looks for admin).
MATTERMOST_BOT_REGENERATE_TOKEN=1 replaces the bot's token.
deploy/compose/mattermost/seed.py's docstring lists every variable.
Things to know:
mmctl bot createrefuses to run in local mode. The script creates the bot over the REST API with a personal access token of the administrator that it generates and revokes again, and switchesEnableUserAccessTokenson only for that moment.- The container runs without
no-new-privileges: with it,docker execfails on a host with AppArmor, and the image's health check and the seed are both adocker exec. - Verified on 2026-09-28 with Mattermost 11.11.1 and PostgreSQL 18.6.
just mattermost-role-bots gives every team-template role its own bot
account, in place of the single fallback kapelle bot every role posts
through until it has one (design doc §5.3, backlog task 01e62a5f). Run it
after just mattermost-seed and just dev-seed (it needs the team/
administrator the first already made, and reads the dev OpenBao the second
already seeded):
MATTERMOST_ADMIN_USERNAME=you just mattermost-role-bots
(Leave the variable out when the administrator is admin, the default.)
Each bot is added to the Mattermost team by this recipe, but to a room's
channel only when it first posts there: the gateway adds it with kapelle's
own token at that post and retries once, so a new channel needs nothing
done by hand.
Each role bot is shown with a display name and a description: the display
name is display_name from the front matter of roles/<role>.md when the
role has one, otherwise the role's name with a capital letter (Coder), and
the description reads "Posts for the role coder of Kapelle teams. Write to it
in your team's channel." The recipe sets both on every run, so a bot made
earlier (which was created with the display name "Kapelle") is corrected by
running it again. The picture of a role bot is roles/avatars/<role>.png when that file exists
(POST /api/v4/users/{id}/image, uploaded on every run, a failure reported on
the role's line and not fatal); a role without a file keeps the server's
default picture. The repository ships plain placeholders for planner, coder and
reviewer (a colour per role with its initial, written by
scripts/make_role_avatars.py with the standard library only); save a real
picture under the same name to replace one, and do not run the script again
over it.
A role's own name comes from every team template's roles: mapping
(teams/templates/*.yaml); a name already used by a person, or one of
Mattermost's four reserved usernames (all, channel, matterbot,
system), is printed and skipped, and that role keeps using the fallback
bot. Each role's bot token goes straight to the dev OpenBao
(kapelle/surfaces/mattermost/bots/<role>, read by credgw/T3's own
gateway wiring, never Kapelle's own .env files) -- never a command line,
a child process's environment, a file or a log line. Safe to run again:
a role whose token still works is left alone, and a role added to a
template later gets its own bot the next time this runs. If OpenBao is not
reachable, or the path it writes to refuses a write, it creates nothing at
all and exits non-zero -- every role stays on the fallback bot, same as
before the run.
Kapelle's side¶
The gateway and the controller read these from the environment the stack is
started in (.envrc, not deploy/compose/.env); just mattermost-seed
writes the first four.
| Variable | Read by | Meaning |
|---|---|---|
KAPELLE_MATTERMOST_URL |
gateway, controller | Base URL, e.g. https://chat.example.org. |
KAPELLE_MATTERMOST_BOT_TOKEN |
gateway, controller | The bot's access token. With the URL it switches the surface on; one of the two alone leaves it off with a warning. |
KAPELLE_MATTERMOST_TEAM |
controller | The Mattermost team, default kapelle. |
KAPELLE_MATTERMOST_COMMAND_TOKEN |
gateway | Token of the /kapelle slash command. Unset, the command is refused (the mention and direct messages still work). |
KAPELLE_MATTERMOST_CA_FILE, KAPELLE_MATTERMOST_TLS_VERIFY |
gateway, controller, just doctor |
TLS: a CA bundle to trust (a self-signed server), or false to skip verification (dev only). |
KAPELLE_MATTERMOST_CATCH_UP_MAX_AGE_SECONDS |
gateway | How far back a reconnect or a restart replays missed posts, default 86400. Older posts are counted in one log line and skipped. |
KAPELLE_DEFAULT_HOME_SURFACE |
controller | mattermost (default) or slack: the room /team create makes without --home. none is refused here: it is only valid as --home none. |
KAPELLE_DEFAULT_TEMPLATE |
controller | Unset (default): /team create without --template is refused with the usage line and the templates that exist -- a template decides where the agents run and what they may cost, so the platform picks none in silence. Set to a template name, it is what a create without --template uses, and the reply says so; a name that matches no template file is answered with the same list. scripts/stack.py sets feature-team, deploy/nevia/start.sh sets feature-team-nevia. |
KAPELLE_OPENBAO_ADDR, KAPELLE_OPENBAO_TOKEN |
gateway | Where role bots' own tokens are read from (backlog task 2a281d66). Both unset is a valid state: no role gets a bot, every one posts through kapelle as before, logged once. just stack-up sets both itself (below); outside it, export them by hand or leave them unset. |
- Rooms.
/team create demo --home mattermostcreates a public channeldemoin the Mattermost team, makes the bot a member and adds the person who typed it when their Mattermost account is linked (/me link mattermost <user id>). Nothing is created for an existing channel of that name: it is reused. - Work. A top-level post by a linked person in a team's channel starts a
work item; the thread carries the rest (each change of a role's status as a
new post, questions with an
@mention, artifacts, one activity summary post that is edited in place). A reply in a thread that is not a work item is ignored, as are posts of the bot itself and of other bots. - Commands.
/kapelle status,/kapelle cancel,/kapelle team create demo --home mattermost,/kapelle help(inside a thread the answer goes into the thread), or@kapelle statusin a channel, or a direct message to the bot (status). Only the one slash command is registered, so/statusand the like never collide with Mattermost's own commands. - Role bots (backlog task 2a281d66, needs
just mattermost-role-botsabove already run).just stack-up(scripts/stack.py) mints the gateway's ownKAPELLE_OPENBAO_ADDR/KAPELLE_OPENBAO_TOKENfor you: a fresh, short-lived OpenBao token scoped to thegatewaypolicy alone (a child of the dev root token, which never itself reaches the gateway's own environment), requested with a 768-hour lifetime and logged with whatever OpenBao actually granted. Runningservices/gatewaysome other way (barejust gateway-dev, a container, ...) needs the two variables set by hand, or left unset entirely -- see the table above. A role with a bot of its own (coder,planner, ...) posts its own status/questions/artifacts as that account, notkapelle-- the account's own name is the only thing that says who is speaking, so the role's name no longer repeats in the text. A role without a bot, or one the gateway's own OpenBao read can't reach, still posts throughkapellewith the role named in the text, exactly as before role bots existed -- nothing about a work item ever fails for this. The gateway learns which roles have a bot once at its own start, and again for a role it hasn't found yet, at most once a minute; a role bot minted or rotated while the gateway is already running is picked up without a restart. The dev OpenBao keeps everything in memory: after its own container restarts, every role bot's token is gone and roles post throughkapelleagain untiljust mattermost-role-botshas run again (same as OpenBao's own root token, its policies, and everything elsejust dev-seedwrote). Backlog task 34d416c1 (T4): a role's bot that has joined the team (just mattermost- role-botsabove) but not yet a given room's channel is added to it by the adapter itself, withkapelle's own token, the first time it is asked to post there, then retried once -- no controller change, no restart. Refused (the add itself failed, or the retry is still refused), the adapter waits ten minutes before trying that same role and room again, posting throughkapellewith the role named in the text meanwhile, same as a role with no bot at all. - Callback. Mattermost calls
POST /webhooks/mattermost/commandon the gateway over the Docker network (host.docker.internal:8101), so the gateway needs the webhook listener on an address the Mattermost container can reach:KAPELLE_GATEWAY_WEBHOOK_LISTENbound to the Docker bridge address (172.17.0.1:8101on the dev host; the proxy and Mattermost both reach it throughhost.docker.internal).0.0.0.0:8101also works but is wider and then needs the host firewall to keep the port from the LAN. The edge proxy does not forward that path. The command's token, compared in constant time, is the route's only authentication: there is no delivery id to deduplicate on, a wrong token or a body over 64 KiB gets a bare200with{}. - Restarts. The gateway keeps the create time of the newest post it handled
in the
mattermost_checkpointstable (Alembic migration in the gateway's tree); after a dropped socket, a sequence gap or a restart it fetches what it missed. A gateway older than the migration starts without a checkpoint and replays nothing. - Live tests.
services/gateway/tests/test_mattermost_live_integration.pyandservices/controller/tests/test_mattermost_live_integration.pystart their own PostgreSQL and Mattermost on a private Docker network, seed them withseed.pyand remove everything (nameskapelle-gateway-test-mm-*). They are opt-in: withoutKAPELLE_TEST_CONTAINERS=1(see "What a test run is allowed to touch" above) they skip, and Docker is not even asked; they also need both images on the host (mattermost/mattermost-team-editionandpostgres, the tags indeploy/compose/mattermost/live_server.py; nothing is pulled). Never run against the dev Mattermost.
env KAPELLE_TEST_CONTAINERS=1 direnv exec . uv run --package kapelle-gateway \
pytest services/gateway/tests/test_mattermost_live_integration.py -q
env KAPELLE_TEST_CONTAINERS=1 direnv exec . uv run --package kapelle-controller \
pytest services/controller/tests/test_mattermost_live_integration.py -q
The team loop with the team's home in Mattermost is
eval/e2e/test_vertical_slice_mattermost.py: real workers and models plus a
throwaway Mattermost, skipped unless KAPELLE_E2E_MATTERMOST=1:
env KAPELLE_E2E_MATTERMOST=1 direnv exec . uv run --package kapelle-gateway \
python -m pytest eval/e2e/test_vertical_slice_mattermost.py -v -rA
just test-unit never starts any of them (their file names contain
integration or they are outside its paths).
- just doctor checks a configured Mattermost: the ping, and that the bot
token is accepted.
Public names (profile: edge)¶
just dev-up-edge starts Traefik (deploy/compose/edge/dynamic.yaml) for
a host that has one public address and several names pointing at it. A
router forwards its port 443 to this host's 9444 (EDGE_HTTPS_PORT); Traefik
routes by name (nothing is routed on port 80, EDGE_HTTP_PORT, 9080):
| Name | Variable | What happens |
|---|---|---|
| Webhooks | EDGE_HOOKS_HOST, required |
TLS ends at Traefik with a Let's Encrypt certificate (TLS-ALPN-01 on 443, stored in the edge-acme volume). Only /webhooks/ is forwarded, to EDGE_HOOKS_UPSTREAM (default http://host.docker.internal:8101, the gateway's webhook listener, see below). Every other path answers 404, and so does /webhooks/mattermost/: Mattermost calls that route over the compose network, and the slash command's token is its only authentication. |
| Mattermost | EDGE_CHAT_HOST, optional |
TLS ends at Traefik with a Let's Encrypt certificate, like the webhook name. Everything, the WebSocket included, is forwarded to EDGE_CHAT_UPSTREAM (default http://mattermost:8065). Unset, the route does not exist. Set MATTERMOST_SITE_URL to https://<EDGE_CHAT_HOST> as well. |
EDGE_ACME_EMAIL defaults to admin@example.com; Let's Encrypt rejects
example.com addresses, so set it.
Things to know:
- Start it only once the router forwards 443 to it. Traefik asks for the
certificate at start, and a failed validation counts against Let's
Encrypt's limit of 5 failures per name and hour. After changing the
router,
docker restart kapelle-dev-edge-1retries. - Traefik runs on the compose network with published ports, not on the host network: Docker publishes ports past the host firewall (ufw), a host-networked proxy would need firewall rules of its own.
- The gateway's webhook listener is opt-in: export
KAPELLE_GATEWAY_WEBHOOK_LISTEN(host:port, e.g.0.0.0.0:8101) beforejust stack-up(the gateway inherits it) and the gateway starts a second HTTP server that serves ONLY/webhooks/github,/webhooks/linear,/webhooks/jira(each when its surface is configured) andGET /webhooks/healthz(200 without touching anything, for checking the path through the proxy). The main gateway port (8100: the same webhook routes plus the unauthenticated A2A mount) stays on loopback. Signatures are verified by the same handlers and the same idempotency store as on 8100;X-Forwarded-Foris logged for display and nothing else looks at it. The address must be reachable from the Traefik container:host.docker.internalresolves to the host, so bind0.0.0.0(or the Docker bridge address) and keep the port closed to the LAN with the host firewall. The Jira Forge route is not on it: it is an A2A transport, not a webhook receiver. Unset, no second listener exists. - The credential gateway's public listener is not behind it. Its clients are agents that are handed the exact certificate to trust, and it needs no public certificate.
- Verified on 2026-09-28 against a UniFi gateway with the name
hooks.yggdrasill.network: the webhook name got its certificate on the first attempt after the router pointed at the proxy.chat.yggdrasill.networkfollowed the same day: certificate on the first attempt, WebSocket upgrade answered 101 through the proxy.
Observability (profile: observability)¶
just dev-up-observability starts Prometheus and Grafana (backlog task
18, design doc §13) alongside just dev-up's core profile.
- Traces: Jaeger at
http://localhost:16686. The worker, gateway and controller export spans (viacontracts/python/kapelle_contracts/ telemetry.py) to the collector's OTLP/HTTP port whenKAPELLE_OTEL_TRACES_ENDPOINTis set to its base URL, e.g.http://127.0.0.1:4318; unset, tracing is a no-op.just stack-upsets it to that address for all three itself. It is deliberately not the standardOTEL_EXPORTER_OTLP_*name: the OpenHands SDK starts a Laminar exporter at import whenever one of those is present, and that exporter speaks gRPC, which fails against the HTTP port withUNAVAILABLEretries around every task's end. Spans are correlated across NATS messages throughtraceparentheaders. What you see:GET http://127.0.0.1:16686/api/v3/serviceslistskapelle-gatewayandkapelle-worker-<role>(kapelle-worker-sharedfor a worker serving every team), and a trace of one message holds the gateway's spans and the worker'sa2a.nats.<method>spans under one trace id (/api/v3/traces?query.service_name=kapelle-gateway&query.start_time_min=... &query.start_time_max=...; Jaeger 2.21 dropped the old/api/query API). The controller configures tracing the same way but did not show as a service in the first check. Agent VMs export nothing by default. The route exists (backlog task 157f068f): every role policy has anoteldestination (POST/v1/traceson the loopback alias127.0.0.5), reached through the guest's HTTP_PROXY like llm and mcp, soKAPELLE_AGENT_OTEL_ENDPOINT=http://127.0.0.5:4318on the process that provisions agents (controller and worker, both providers) makes the agent-server export spans, withOTEL_EXPORTER_OTLP_TRACES_PROTOCOL= http/protobufset beside the endpoint. Left unset instack.py, and it stays a decision, not a default: the agent-server's spans carry the LLM prompts and completions and the tool inputs and outputs, and no setting in the installed openhands SDK and Laminar turns that off completely (LMNR_TRACE_CONTENT=falsereaches only Laminar's openai, anthropic and groq instrumentations, not the LiteLLM callback or the SDK's ownobservespans), so the collector and Jaeger would hold agent conversations. - Metrics: Grafana at
http://localhost:3000(anonymous admin in dev), one provisioned dashboard ("Kapelle overview"). vmd's own/metrics(default127.0.0.1:9092) and credgwd's (127.0.0.1:9192) are scraped directly by this profile's Prometheus (deploy/compose/ prometheus/prometheus.yml) — both are native-host processes, not compose services, so Prometheus runs withnetwork_mode: hostto reach their loopback-bound ports (see that file's own comment for whyhost.docker.internaldoesn't work here). Everyvmd_*metric documented invmd/internal/metrics(sandbox counts by state, boot/ sleep duration, snapshot size, capacity, per-agent disk usage and configured memory, watchdog kills, jailer setup wait) shows up here as soon asvmd-dev(or a realvmd) is running with the default-metrics-listen, no extra wiring needed. Theservicesprofile's worker and gateway containers export their own/metricstoo (kapelle_worker.metrics/kapelle_gateway.metrics,kapelle_*- prefixed, no per-service namespace) onlocalhost:9292/localhost:9392respectively (KAPELLE_METRICS_PORT, fixed in compose since each container owns its whole port space) — same scrape/host-networking story as vmd/credgw above. - Metrics via OTLP: the collector's own Prometheus exporter
(
localhost:8889,otel-collectorjob inprometheus.yml) is wired and scraped too, but nothing pushes OTLP metrics into it yet — today only traces flow through the collector. If/when a service starts exporting OTLP metrics (vmd doing so, matching how the Python services already push traces, is an open question as of this writing — see vmd's own task-28-era annotation), they'd appear via that job automatically; no compose changes needed to receive them. - Prometheus listens on
http://localhost:9091(--web.listen-addressindeploy/compose/docker-compose.yaml, which also sets the compose healthcheck and Grafana's datasource URL to it), not its default 9090. Every credential-gateway instance used to bind 9090 on its own address, so a Prometheus on*:9090made every agent boot fail withbind: address already in usewhile the profile was up. Local instances now pick a free metrics port of their own (Nevia ones already did), so the profile runs beside the stack; keep 9090 free anyway if you run an older credgwd. - Prometheus/Grafana provisioning lives in
deploy/compose/prometheus/anddeploy/compose/grafana/.
Services (profile: services)¶
just dev-up-services runs the control plane itself -- controller,
gateway, worker -- as containers alongside just dev-up's infrastructure
(backlog task 38), built from the uv workspace by one shared
deploy/compose/services.Dockerfile (parameterized by build args, not
three near-identical files). Needs just dev-up and just dev-seed
first (NATS streams/KV, OpenBao policies). Each service's schema
migration runs as a one-shot <service>-migrate container (alembic
upgrade head) before the real service starts -- up's own dependency
graph (depends_on: ... condition: service_completed_successfully)
handles the ordering, nothing extra to run by hand.
This profile is for proving the containerized control plane works, not
day-to-day development -- most iteration still happens against a native
just controller-dev/gateway-dev/worker-dev process, which starts in
seconds and doesn't need a rebuild after every code change. The
controller doesn't publish a fixed host port here (unlike
postgres/nats/openbao/litellm) -- every real caller reaches it
over the kapelle network by service name (controller:8300), and a
fixed publish would collide with a teammate's own native controller-dev
run on a shared host; use docker compose ... port controller 8300 for
one-off host-side debugging. The gateway's HTTP port (webhooks plus the
/a2a/<team>/<role> A2A JSON-RPC interface, backlog task 36) IS
published, since external callers of that port are exactly the point --
but with no FIXED number: KAPELLE_GATEWAY_HOST_PORT defaults unset,
which Compose resolves to an ephemeral host port rather than the fixed
8100 a teammate's native gateway-dev binds, so the two never collide;
look the assigned port up the same way (docker compose ... port gateway
8100).
The worker container runs worker SHARED mode (backlog task 0deba590):
KAPELLE_TEAM/KAPELLE_ROLE are both unset, so it serves every team's
coder role instead of being pinned to one fixed team -- design doc §5
("workers are stateless and shared"), and the only way just eval
(which creates a uniquely-named team per run) can ever get an answer
from this profile at all. It also runs a REAL model, not the echo stub:
KAPELLE_EXECUTOR=openhands plus FakeProvider's real-agent-server
mode (KAPELLE_FAKE_PROVIDER_REAL_AGENT_SERVER=1, pointed at the
compose litellm -- see "Running a real model without Firecracker"
above, now the compose default instead of only reachable natively). The
image gets the real-agent-server extra and a git binary via
services.Dockerfile's EXTRA build arg (only this service passes it
-- controller/gateway stay lean), plus a placeholder repo with one
commit already baked in at /workspace/project (the coder role's
default repo_checkout_dir) so a team with no repo_url -- every eval
task kind that doesn't need a real git remote -- still has something to
work from.
Known gap, found wiring this up: just eval/kapelle_eval.stack.
NatsStackDriver can't actually complete a real run against this
profile yet -- TeamService.create()'s admin-membership grant (546116d)
needs the creator's users.id to already exist, and nothing creates
that row yet for a synthetic caller (gateway's DirectoryAuthorizer is
gaining auto-provisioning for real surface identities; NatsStackDriver's
own hardcoded admin_person_id="eval-runner" isn't even a valid
users.id and needs its own fix once that lands). Tracked separately;
this profile's worker/controller/NATS wiring itself is verified working
in isolation (a fresh docker compose -p project, so it doesn't share
the dev NATS instance's A2A_RPC stream with any already-running pinned
consumer -- JetStream's own work-queue rule refuses to mix a shared-mode
wildcard consumer with a pinned one on the same stream, so don't run
just dev-up-services and a native pinned worker-dev against the same
NATS at once either).
deploy/compose/tests/test_services_smoke_integration.py (backlog task
38's Docker-gated smoke test, part of just test-integration) brings up
an entirely separate, isolated copy of this same profile (a unique
docker compose -p project, its own network/volumes/NATS -- exactly the
isolation the paragraph above depends on -- and
deploy/compose/tests/ports-override.yaml republishing the infra
services' ports with no fixed number so it can run alongside the shared
kapelle-dev stack without colliding) and drives a real /team create
through the containerized controller, then a real model response
(not the echo stub, now that the worker runs one) through the
containerized GATEWAY's own HTTP A2A surface (a plain httpx POST to
its published port's /a2a/smoke/coder/a2a/jsonrpc JSON-RPC route,
A2A-Version: 1.0 header required) which forwards over NATS to the
containerized worker -- controller -> gateway container HTTP A2A ->
worker, the real external path, not kapelle_gateway.a2a_client.
GatewayA2AClient called directly (that would only prove the NATS wire
protocol, not the gateway container's actual HTTP surface). The "smoke"
team is created and the gateway container started only afterwards, so
the gateway's one-shot startup refresh_once() mounts the route before
the test needs it, rather than waiting on run_a2a_mount_loop's 60s
periodic pass. Proof the composed controller/gateway/worker images and
their real wire protocols actually work together end to end against a
real model, not just that each image imports cleanly.
Dedicated NATS accounts (profile: nats-operator)¶
just dev-up-nats-operator (backlog task d74d9c67; see docs/
team-isolation.md and contracts/a2a-nats.md's "Dedicated NATS
accounts" section for the full design) starts a SECOND nats-server,
separate from the default nats service, running in NATS's decentralized
JWT auth mode (operator -> account -> user) instead of plain/no auth.
This is where an isolation: dedicated team's own NATS account actually
lives and authenticates, once one is provisioned -- the default nats
service stays exactly as it is (plain auth, every isolation: shared
team plus the control plane), unaffected by this profile existing or not.
Needs just dev-up and just dev-seed first: dev-seed (via
kapelle_controller.openbao.seed.seed()) mints the deployment's one NATS
operator identity/signing key and its $SYS system account into OpenBao
the first time it runs (idempotent after that -- a re-seed never re-mints
either, which would orphan every already-minted team account). just
dev-up-nats-operator then renders nats.conf from that OpenBao material
(kapelle_controller.openbao.render_nats_operator_conf, the same
host-side-script pattern dev-seed itself uses -- no separate built
image) before starting the container. Re-running it after the operator's
signing key rotates picks up the change; an unchanged operator leaves the
existing rendered config alone.
Published on 4223/8223 (not 4222/8222) so it runs alongside the
default nats service without colliding. just dev-down-nats-operator
stops it and removes its own two volumes (nats-operator-jwt,
nats-operator-js) without touching the core profile's containers/data
-- just dev-down alone does NOT stop this profile's container (same
"a profile only adds to down's scope, a bare down -v without service
names would remove every named volume in the file, including the shared
dev stack's" reasoning dev-down-services above already documents).
Rendered files under deploy/compose/nats/operator/ (gitignored -- a
freshly-minted operator JWT per OpenBao instance, not source) hold only
public material: the operator's own JWT, the $SYS account's JWT, and
plain resolver settings. No seed, and no .creds, is ever written there
-- the one secret this profile actually needs (the $SYS admin's
.creds, for anyone who needs to publish $SYS.REQ.CLAIMS.UPDATE/
DELETE -- provisioning/revoking a dedicated team's own account) stays
in OpenBao (kapelle_controller.openbao.layout.NATS_SYSTEM_ACCOUNT's own
admin_creds field) and is read from there directly, never from a file
on disk.
Recipes¶
| Recipe | Does |
|---|---|
just dev-up |
Start the core profile (nats, postgres, openbao, otel-collector, jaeger, litellm). |
just dev-up-storage |
Also start RustFS. |
just dev-up-observability |
Also start Prometheus + Grafana. |
just dev-up-services |
Also build and start the containerized controller/gateway/worker. |
just dev-down |
Stop everything (volumes, and so their data, survive). |
just dev-seed |
NATS streams/KV + OpenBao policies/paths/AppRole auth; prints a LiteLLM liveliness check. |
just dev-seed-storage |
Puts and gets an object through RustFS's S3 API to prove it works. |
just dev-logs [service] |
Follow logs for everything, or one service. |
just controller-dev |
Run the team controller natively against the compose services (no container image yet). |
just credgw-dev [role] |
Run credgwd natively with one gateway instance started immediately. |
just vmd-dev |
Run vmd natively against deploy/vmd/dev.yaml. |
just prune-teams [--older-than 2h] [--apply] |
Archive leaked e2e-*/eval-*/demo-*/repro-* teams through the real controller path (see Housekeeping below). |
just prune-follower-consumers [--apply] |
Delete stale gateway-follower-<team> durable consumers on A2A_TASKS (see Housekeeping below). |
Housekeeping¶
A long-running shared dev NATS/Postgres accumulates leftover state from
killed test runs and interrupted just eval/just demo invocations --
neither cleans up after itself if it never reaches its own teardown
step. Three scripts, all dry-run by default:
scripts/nats_prune_consumers.py-- deletes a durableA2A_RPCconsumer whose team is notactivein the controller's ownteamstable. Nojustrecipe; run directly:uv run --package kapelle-controller python scripts/nats_prune_consumers.py [--apply].just prune-teams [--older-than 2h] [--apply](scripts/archive_stale_teams.py) -- archives a leakede2e-*/eval-*/demo-*(just demo's vertical-slice scenario)/repro-*(one-off teammate repros) named team (stillactive, older than--older-than) through the realTeamService.archive()path -- releases its agents/credentials, deletes its ownA2A_RPCconsumers, revokes its LiteLLM team/keys, then marks it archived. Prefer this over the consumer script alone when a team's leak is the actual root cause (a consumer-only prune just lets a new one reappear the next time anything touches that still-active team) -- see its own module docstring for why a shared-mode worker's wildcard consumer in particular can never bind while ANY of these leftover pinned consumers exist on the same stream.just prune-follower-consumers [--apply](scripts/nats_prune_a2a_tasks_consumers.py, backlog task 4bb2516f) -- the sibling of the first script above, for the GATEWAY's owngateway-follower-<team>durable consumers onA2A_TASKSinstead of the worker'sA2A_RPCones (backlog task cd4c8f70 found the gateway following 60+ archived teams forever, since nothing ever deleted these). Reports each stale consumer's OWN reason separately --archived,no row(noteamsrow at all), orno link(active, but nosurface_linksrow on any surface) -- with its age and whether a pull is currently waiting on it. Two further guards a deletion here can't take back need: a consumer with a pull waiting is never touched (a follower alive on it right now, whatever the tables say), and one younger than an hour is never touched either (identical to a team mid-creation, or a short-lived e2e team whose test is still running).
None of these three scripts' prefix/staleness rules cover every
ephemeral naming convention this repo's own test suites might invent in
the future -- a caller using a new throwaway team-naming convention
outside _STALE_TEAM_NAME_PREFIXES (scripts/archive_stale_teams.py)
won't be swept up automatically; check just prune-teams (dry-run) and
a plain read of the teams table's active rows if leaks reappear under
a name shape not already listed there.
Running the controller natively¶
The team controller (backlog task 16, services/controller) has no
container image yet — like vmd/credgw, it runs as a plain process against
the compose services just dev-up brings up, not as a docker-compose.yaml
service. just controller-dev runs it with NATS_URL/DATABASE_URL/
LITELLM_BASE_URL/LITELLM_MASTER_KEY/OPENBAO_ADDR/OPENBAO_TOKEN filled
in from deploy/compose/.env and the compose services' host-published
ports (see services/controller/src/kapelle_controller/config.py for the
full env var list, including the optional GitHub App credentials).
Before the first run:
just dev-upandjust dev-seed.uv run --package kapelle-controller alembic -c services/controller/alembic.ini upgrade headagainstDATABASE_URL— the controller does not create its own tables at runtime.just controller-dev.
Without a Mattermost or Slack client configured, /team create without
--home is refused with a message naming the variable to set; --home none
creates the team with no room. Without KAPELLE_GITHUB_APP_ID set, the
controller logs a warning and GitHub linking/minting is disabled — enough to
exercise /team create --home none, link|archive|show and /budget end to
end over NATS, and the
internal Controller gRPC service (contracts/proto/controller.proto,
default 127.0.0.1:8300 -- not 8200, which is OpenBao's compose port) that
credgw/vmd call for GitHub token minting and
LiteLLM key rotation.
A team that must not get a chat room (probes, e2e runs, tests) is created with --home none (or
TeamService(default_home="none")): no channel is created, no surface link is stored, and its notices are
dropped at INFO; the controller default stays mattermost, and none is refused as
KAPELLE_DEFAULT_HOME_SURFACE, so a room can only be skipped by asking for it. --home none is
deliberately absent from the /team create usage text; a team create without --home on a controller
whose default surface is not configured now fails with a readable error before anything is created, naming
the variable to set and saying to pass --home.
Team memory semantic search¶
Team memory (design doc §14) searches by keyword alone unless
KAPELLE_EMBED_ALIAS is set to the name of a LiteLLM alias backed by
an embedding model. Since 2026-09-24 nothing sets it: the Aiven AI
Gateway (the only model provider) serves no embedding model, so
scripts/stack.py leaves the variable unset and search is keyword-only
until the gateway offers one. With it set, /memory search (and the worker's own
memory_search MCP tool) blends keyword full-text ranking with
pgvector cosine-similarity ranking, so a query can find a related
entry that shares no keywords with it at all -- see
kapelle_controller.memory.store's own module docstring for the exact
blending rule. Unset (or the alias erroring), search silently falls
back to keyword-only; this is never a hard dependency.
Running a real model without Firecracker¶
A worker doesn't need a real Firecracker VM to talk to a real model:
FakeProvider (KAPELLE_SANDBOX_PROVIDER=fake, the default) has a
second mode, use_real_agent_server, that starts a real local
python -m openhands.agent_server subprocess instead of its usual
pure-in-memory stub -- a real, reachable agent, just not inside a VM.
Set KAPELLE_FAKE_PROVIDER_REAL_AGENT_SERVER=1 on a worker started with
KAPELLE_EXECUTOR=openhands to turn this on (kapelle_worker.sandbox.
registry's default "fake" registration reads it); needs the worker
package's real-agent-server extra installed
(uv sync --extra real-agent-server --package kapelle-worker).
Other env vars this mode reads (all optional, sensible dev defaults):
KAPELLE_FAKE_LLM_UPSTREAM/KAPELLE_FAKE_LLM_KEY-- where the real agent-server's LLM calls actually go. Point these at the compose LiteLLM (http://127.0.0.1:4000anddeploy/compose/.env'sLITELLM_MASTER_KEY) to get real completions; leaveKAPELLE_FAKE_LLM_ KEYunset and the agent-server still boots, it just can't complete a real conversation. Same env var nameseval/e2e/conftest.py's own "fake-real" provider registration already uses, for its harness-only worker subprocess -- this is that same construction, promoted onto the default"fake"name so a realkapelle_worker.a2a.mainprocess reaches it too, not just that one test harness.KAPELLE_MCP_HTTP_PORT/KAPELLE_MCP_TOKEN_SECRET-- reused verbatim from the worker's own MCP server config, so the two can't drift; only set these if you're already overriding them for the worker itself (e.g. to avoid colliding with agateway-devon the same default port, 8100).KAPELLE_OPENBAO_ADDR/KAPELLE_OPENBAO_TOKEN/KAPELLE_OPENBAO_HOST-- when set, the agent-server's LLM calls go out under THIS agent's own minted LiteLLM virtual key (read back from the same OpenBao document a realEnsureAgentwrites it to) instead of the flatKAPELLE_FAKE_LLM_ KEYfor every agent. Leave unset for a quick local check where attribution doesn't matter; set them (dev compose:http://127.0.0.1: 8200anddeploy/compose/.env'sOPENBAO_DEV_ROOT_TOKEN) to actually exercise per-agent budgets/rpm/tpm limits.KAPELLE_OPENBAO_HOSTmust match whateverhostthe controller that ran this agent'sEnsureAgentused, not this machine's real hostname.KAPELLE_ROLE_CODER_REPO_CHECKOUT_DIR(or_<ROLE>_for another role) -- the real agent-server clones the team'srepo_url(/team create ... --repo <url>) intorepo_checkout_dir, which defaults to/workspace/project-- root-owned and non-writable outside a real Firecracker guest. Point this at a real, writable directory on this host instead (RoleConfig.from_env's docstring has the full reasoning for why the default is what it is).
Verified for real: a native controller-dev + gateway-dev +
a worker with these four env vars set, given a team pointed at a local
bare git repo, answered a plain chat message with a real model
completion (not the echo: stub) end to end over NATS.
Known gap: a worker started this way (or any way today) is pinned to
one team/role for its whole lifetime (KAPELLE_TEAM/KAPELLE_ROLE),
so it can't serve just eval's per-run, uniquely-named teams --
eval/e2e/run_real_eval.py works around this by spawning its own
worker subprocess per task instead of reusing a long-lived one. A
worker "shared mode" (subscribing every team's queue via a wildcard
subject instead of one fixed team) is planned to close this gap; once
it lands, just dev-up-services's worker and just eval won't need
this workaround.
Running a team on Nevia¶
Backlog task 65a295d0 (design doc §8 "Nevia provider", §12 "Templates"): a team's roles pick their sandbox provider per role, in the template -- no Nevia-specific env var flips a whole deployment over, and firecracker and nevia roles coexist in one running controller/worker pair.
- Build (or reuse) a golden checkpoint.
just nevia-image(seeimages/README.md's "Nevia image" section) creates or wakes a build computer, installs the pinned OpenHands venv, and prints a checkpoint id (chk_...). That id isvm.checkpoint_idbelow -- the build computer itself must stay around (sleeping is fine, deleting is not: a checkpoint 404s on restore once its own source computer is gone,images/README.md's own finding). - Set
vm.provider: neviaon the role(s) that should run there, inteams/templates/<name>.yaml:
roles:
coder:
agent: coder
model: open-coder
vm:
provider: nevia
profile: small # small | medium | large -- see
# kapelle_worker.sandbox.providers.nevia.PROFILES
checkpoint_id: chk_c1503396c4a9
replicas: { max: 3 }
concurrent_runs: 1
A bare vm: small (no provider:/checkpoint_id:) still means
provider: firecracker, unchanged from before this task -- every
existing template keeps working as written. /team create validates
this at template-load time (kapelle_controller.templates.VmProfile);
/team show <name> prints each role's resolved vm=<profile>@
<provider>[:<checkpoint_id>]. An existing team's role can move
between providers without a full archive+recreate: /team set <name>
<role> vm_checkpoint_id chk_... first, then /team set <name> <role>
sandbox_provider nevia (the reverse order is refused with a readable
error -- a role can never end up provider: nevia with no checkpoint
configured).
3. Set the Nevia credentials in the environment just stack-up
(scripts/stack.py) runs from -- forwarded, when set, to BOTH the
controller (it calls SandboxProvider.ensure() directly,
kapelle_controller.agents.ensure_agent) and the worker (wake()/
touch()/... for an already-ensure()-d Nevia agent):
- KAPELLE_NEVIA_TOKEN -- the Nevia API bearer token, used as is. Nevia user
tokens expire after 15 minutes, so this only suits an automation bearer that
something else keeps fresh; on a 401 the provider fails saying the token is
expired or invalid. Nevia has no service account or long-lived token in v1.
- KAPELLE_NEVIA_TOKEN_COMMAND -- instead of the token: a command whose stdout is
the bearer, normally nevia auth access-token after a one-time nevia login
(that command refreshes the CLI's own session; the worker never touches the
credentials file). Run without a shell, cached, re-run 60 s before the token's
exp and once after a 401. Use an ABSOLUTE path (for example
/home/you/.local/bin/nevia auth access-token): the worker's PATH must find
nevia otherwise. Set exactly one of the two. There is no default: just
stack-up forwards it only when you set it. It makes the worker act as the user
logged in to that CLI, in every workspace that user can reach (a development
bridge; a machine identity is the production answer once Nevia has one), and a
session that has expired or was revoked needs nevia login again.
- KAPELLE_NEVIA_WORKSPACE -- the Nevia workspace id (ws_cb2f66c918e9, the
kapelle workspace, for this repo's dev setup).
- KAPELLE_NEVIA_GOLDEN_SOURCE -- optional; the golden checkpoint's
own build computer id (images/nevia/build.sh's output, e.g.
cmp_27b314baee70 in the kapelle workspace). Unset means NeviaProvider resolves it itself
by scanning the workspace's computers the first time any given
checkpoint id is used (one extra request, then cached) -- set this
to skip that scan once you know the id.
Neither is required for a controller/worker running with no nevia
role configured anywhere -- get_provider("nevia") (and so these two
env vars) is only ever reached the first time a role actually resolves
to that provider (kapelle_worker.sandbox.registry._require_env
fails loudly, at that point, if either is missing).
4. just stack-up as usual. A firecracker-provider role in the SAME team
(or a different team) is completely unaffected -- both providers are
registered and cached independently in the same controller/worker
processes (kapelle_worker.sandbox.registry.get_cached_provider).
Credential gateway on Nevia (backlog task ec29a80f): with
KAPELLE_NEVIA_CREDGW_ADDR set, wake() starts the reverse relay
(images/nevia/kapelle-nevia-relay, placed in the golden image by
just nevia-image; an image built before that task has no such script, so
rebuild it) and the agent's HTTP_PROXY/HTTPS_PROXY point at
http://127.0.0.1:18080 inside the computer. The worker dials the agent's own
credgw instance on its private tunnel listener, so the worker must be able to reach
the host credgwd runs on: KAPELLE_NEVIA_CREDGW_RELAY_HOST (default 127.0.0.1,
right for the dev stack where both run on this host). Also read here:
KAPELLE_NEVIA_CREDGW_MODE (relay, the default, or public for the earlier
client-token mode against credgwd's shared public listener) and
KAPELLE_NEVIA_RELAY_PORT (default 18080, the loopback port in the computer).
Measured live on 2026-09-28 (test_nevia_reverse_relay_integration.py, opt-in with
KAPELLE_NEVIA_RELAY_LIVE=1): wake() with bootstrap and relay start about 13 s; a
no-model request adds about 43 ms round trip through the relay against a direct
call to the credgw instance; a streamed completion's first byte added about 45 ms
(median); model time varies far more than that between runs. Cutting the relay stream
abruptly left no orphan relay process in the computer (Nevia ended it with the
connection), and the session had a new relay listening 1.6 s later; only probes in
that first 1.6 s failed. A connection the kernel queued on the old relay's listener
in the milliseconds around a stream renewal's DRAIN is reset (the two relays share
the port), so a client that retries connection errors never notices.
Reading the forward relay's own log (backlog: stage 2 runs 9-11, 2026-09-28):
_handle_relay_connection (kapelle_worker.sandbox.providers.nevia) logs one line
per LOCAL connection through the relay -- the agent key, the request's own first
line (method and path only, never headers or body), reads=[...] (the byte count
of every individual local read() on that connection, in order, so a request that
arrived in one piece looks like reads=[12456] and one split across several looks
like reads=[4096, 4096, 4017]), the total bytes read, the bytes received back, and
how the connection ended. It logs at WARNING when nothing at all came back (grep
"ended without a response" for that case specifically) and at INFO otherwise, so
the line is always there for a normal request too, not just a failure -- useful on
its own for seeing how large a live request actually was, not only when one dies.
When a request into a Nevia computer fails with no other explanation (worker log:
httpx.RemoteProtocolError: Server disconnected without sending a response or
similar, with no corresponding error inside the guest's own agent-server log), read
this line for that connection first: a reads total or single entry at or above a
few tens of KB, together with ended: ok from the relay's own side but nothing
useful from the client's, points at the transport, not the request's content --
that is exactly how the run 9/10 regression (_pump_connect_stream handing h2 a
single un-chunked frame past its negotiated size) first showed up here, before it
was fixed (send_h2_data, commit ff1e62b).
TLS to intercepted HTTPS destinations (github.com for the agent's git clone/push):
a remote instance's own ca_cert_pem is only the public listener's certificate, so
wake() fetches credgwd's interception CA (GetCACertificate, once per worker process)
and hands the computer a bundle of it (plus the listener certificate in public mode)
as KAPELLE_GATEWAY_CA_PEM; the bootstrap installs every certificate of it next to the
system roots. If credgwd cannot supply the CA, wake() fails with a clear message. The
bootstrap applies the bundle when it (re)starts the agent-server, not when it re-attaches
to one that is already running, so an agent that was awake before this change needs a
fresh computer. Requests the relay carries are never logged, and a stream URL is never
kept in a log line.
Rejected requests in credgwd's audit log that are expected (seen in every Nevia run, in
the instance's iron-proxy.log under credgw/run/<computer name>/): a GET of
github.com/<owner>/<repo>/info/refs and a second GET /v1/model/info, both 403 with
"rejected_by":"secrets" and the annotation "rejected":"…/creds/github_installation_token"
or …/litellm_virtual_key. The secrets transform is rendered with require: true: a request
to a destination that carries a swapped secret but arrives without the placeholder in its
Authorization header is rejected instead of being forwarded unauthenticated (credgw/README:
the contract). Those two are credential-less probes (git's first unauthenticated request, an
SDK model-info request without a key); the requests with the placeholder are allowed right
around them, with "swapped":[{"secret":…,"locations":["header:Authorization"]}]. The audit
does not print headers, so "no placeholder" is inferred from the rule and the pattern. A
git ls-remote typed by hand in a computer gets the same 403. Nothing to chase unless a
request that should carry the placeholder is rejected.
Other KAPELLE_NEVIA_* env vars (KAPELLE_NEVIA_GATEWAY_LLM_URL/
_GATEWAY_MCP_URL/_GATEWAY_PROXY_URL/_GATEWAY_CA_PEM/
_GATEWAY_CLIENT_TOKEN/_MCP_TOOLS_TIMEOUT_SECONDS/_CREDGW_ADDR) wire
the credential gateway's remote-mode path into a Nevia agent
(kapelle_worker.sandbox.providers.nevia.NeviaProvider's own
constructor docstring) -- out of scope for a first stack-up run above,
not required to get a Nevia agent running at all.
The control plane running on Nevia¶
Backlog task 56fabe56 (design doc §8, control-plane spec doc of task
5c856e9b): stage 3b proved the platform's own services -- OpenBao, NATS,
credgwd, LiteLLM, controller, gateway, worker -- running natively (no
containers, no just) inside one Nevia computer, against a managed
Postgres, instead of on the workstation. deploy/nevia/start.sh is that
proof landed as a real, reusable script (idempotent, refuses a second
concurrent start); deploy/nevia/nats_seed.py and pg_ensure_dbs.py are
its own companion steps. Bringing the computer up, uploading a git
archive of a real commit into it, and building credgwd/iron-proxy
from that same commit are still done by hand from outside (see stage 3a/3b
in task 5c856e9b's doc for the exact commands) -- start.sh only covers
what runs once that's in place.
Backlog task 5c856e9b (v2 of that script, real credentials): scripts/nevia_cp.py
runs on the operator's machine and drives one computer (default kapelle-cp, workspace
KAPELLE_NEVIA_WORKSPACE): create, sync (a git archive HEAD plus checksummed
binaries), install (deploy/nevia/install.sh: uv, bao and nats-server at pinned versions
verified against the published sha256, then the uv and LiteLLM environments), secrets, up,
health, stage 1 (runs eval/e2e/run_nevia_stage1.py in the computer against its own stack,
the addresses and credentials read from the secret files there; branches are checked by hand
because the computer has no gh), down --confirm <name>; --dry-run prints every
nevia call. secrets reads AIVEN_AI_API_KEY, PG_URL, the GitHub App key file and ids,
the chat bot variables and a copy of a Nevia session
(KAPELLE_NEVIA_CP_CREDENTIALS_FILE) from its own environment and pipes them over
nevia exec -i stdin to deploy/nevia/install_secrets.py, which writes 0600 files under
/data/nevia/kapelle/secrets and prints names and sizes only; a secret is never an
argv. up --init --unseal-key-out FILE initialises OpenBao on the first start (the key
goes to that file on the operator's machine, the root token is revoked after the layout
and the AppRole ids are written); every later up takes the unseal key from
KAPELLE_CP_BAO_UNSEAL_KEY or a prompt and passes it on stdin. credgwd and the
controller log in with AppRole; the gateway and worker log in with their own AppRoles too
(role files, token renewed by the process). A keep-alive loop (start.sh --tick every 30 s) restarts what died; a sealed
OpenBao stays down until the next up. Every checkpoint of this computer carries these
secrets: keep it in workspace kapelle, delete it when done.
LiteLLM on its own computer (backlog task eb4d004f). One computer holds every
component in 2 GB, and LiteLLM alone takes about 750 MB of it, so the control plane can
be split in two: kapelle-llm runs LiteLLM alone (--role llm), kapelle-cp runs
everything else (--role core) and reaches it through the computer's own ingress URL
(Nevia issues https://<cmp>.live.nevia.cloud since 2026-10-01; a published port is 443
or 8000-20000 outside, nevia_cp.py create --port 443:4001 maps it to LiteLLM's port
inside; plain http on a published port is refused by the ingress). The agents get that
URL from credgwd (-llm-tls, an https llm_url) and CONNECT-tunnel to it through
iron-proxy like any other TLS destination, the virtual key injected into the decrypted
request; the services use it directly. The master key is one local 0600 file,
secrets --role llm|core --litellm-key-file FILE (generated there when missing; llm-key
--out FILE copies an all-in-one computer's existing key out first, so the virtual keys
LiteLLM already issued stay valid), installed as /secrets/litellm_master_key on both
computers, plus litellm_url on the core one. In order, on a fresh split:
nevia_cp.py --name kapelle-llm create --port 443:4001
nevia_cp.py --name kapelle-llm sync
nevia_cp.py --name kapelle-llm install --role llm # uv and the LiteLLM venv only
nevia_cp.py --name kapelle-llm secrets --role llm --litellm-key-file FILE
nevia_cp.py --name kapelle-llm up --role llm # no unseal key
nevia_cp.py create && ... sync --bin ... && ... install # the core as before
nevia_cp.py secrets --role core --litellm-key-file FILE # adds the key and the URL
nevia_cp.py up --role core --init --unseal-key-out FILE
up --role records the role in /role, which the keep-alive loop and health read;
health on the core computer checks LiteLLM through the URL. The public URL answers
anyone who knows the computer id: LiteLLM's own key check is the only gate on it.
Checkpoint and restore, in this order, always: checkpoint, then SUSPEND
the source, then restore -- never delete a computer whose checkpoint is
still needed. Found live in stage 3b: a checkpoint lives under its
source computer (the restore endpoint is POST /v1/computers/{cmp}/
checkpoints/{chk}/restore, scoped by that computer's id), and deleting
the source cascade-deletes its own checkpoints with it -- confirmed
through GET /v1/operations, a checkpoint.delete operation appears
automatically right after the computer.delete. Nothing is left orphaned
or silently billing storage, but the checkpoint is simply gone: restoring
a checkpoint whose source was already deleted is not possible on this
API, in any order. Suspending the source instead (nevia suspend) keeps
the checkpoint reachable while guaranteeing the source and the restored
copy are never both running at once (the source is asleep, not deleted --
delete it only after the restore is confirmed healthy). A restore resumes
every process in place with its prior memory state; expect about a
30 second window with no usable network right after, and any service that
held an open database connection reconnects on its own within about a
minute (LiteLLM/asyncpg both did, live, no restart needed) -- restart a
service only if it is still unhealthy after that.
Running vmd/credgwd natively¶
vmd and credgwd can't run inside just dev-up's Docker Compose stack
— Compose can't sandbox a microVM host, and vmd needs /dev/kvm, real
network namespaces/taps/nftables, and the jailer's own privilege model
(see the Prerequisites section above). Both run as plain processes on the
host instead, against the compose services for everything else (OpenBao,
LiteLLM, ...), the same pattern controller-dev already uses.
Neither daemon needs to run as root: every privileged operation each one
performs execs a specific, pinned binary (ip, nft, jailer,
firecracker, ...), and internal/host.runPrivileged (vmd) transparently
prefixes sudo -n to those calls when the process isn't already root.
deploy/host/sudoers.kapelle (generated by vmdctl sudoers, checked in)
is the minimal grant: exact command+argument patterns, not the bare
binaries — a plain NOPASSWD: /usr/bin/tee (or cp/chown) is
effectively unrestricted root for anything running as that user, since
any of those three can overwrite an arbitrary file, so this replaced that
shape of grant entirely. Validate it first, then install it as
/etc/sudoers.d/kapelle-dev:
visudo -cf deploy/host/sudoers.kapelle
sudo install -m 0440 deploy/host/sudoers.kapelle /etc/sudoers.d/kapelle-dev
A fresh host, or one whose username/configured directories differ from
what's checked in, regenerates its own copy first:
vmdctl sudoers -config <your vmd.yaml> -o deploy/host/sudoers.kapelle
(defaults to the current user and config.Default()'s own directories if
you omit -config/-user). Add -dev-test-roots to also grant the
test suites' fixtures under the root-owned base /var/tmp/vmd-it-base
(one entry, /var/tmp/vmd-it-base/*: every fixture works under it since
3188034e — jail, jail-<label>, <prefix>-<hex>, each with its own
images, disks and snap directories) the same jailer/cp/chown/
mkdir/rm patterns; the checked-in file itself is generated with
vmdctl sudoers -config deploy/vmd/dev.yaml -jailer-bin
/usr/local/lib/kapelle/bin/jailer -dev-test-roots -alias KAPELLE_VMD -o
deploy/host/sudoers.kapelle (what scripts/doctor.py compares the
installed grant against) so those suites' real privileged calls aren't
403'd by a grant anchored to only the one directory vmd.yaml itself
configures. The base is root-owned and never a prune target; its children
are removed through vmd-priv fileop (just prune-workdirs). Or skip
sudoers entirely and just run the daemon as root the way
deploy/host/systemd/vmd.service does in production.
-
just credgw-dev [role]— runs credgwd with one gateway instance forrole(defaultcoder) started immediately, against the compose LiteLLM (seecredgw/README.md's Dev section for what this does and doesn't exercise). Needsjust dev-upfirst, andjust credgw-fetch-iron-proxyto have fetched the pinned iron-proxy binary intocredgw/.bin/(done automatically if missing). -
just vmd-dev— builds vmd and runs the compiled binary (exec'd, notgo run, so a signal sent to this recipe's process reaches vmd directly rather than being swallowed bygo run's own wrapper process — that gap orphaned real Firecracker/jailer processes on this shared host when ago run-basedjust/pytest teardown killed only the wrapper) againstdeploy/vmd/dev.yaml, a dev-only config pointing every state directory at/var/tmp/vmd-dev/*(real disk — see below) instead of production's root-owned/srv/vmd, usingjust dev-up's OpenBao directly via its well-known dev root token (no AppRole login). Needs: just dev-upfirst (OpenBao must be reachable, or vmd starts but everySleepcall fails immediately).- The pinned Firecracker/jailer binaries already installed at
/usr/local/lib/kapelle/bin/(deploy/host/fetch-firecracker.sh, a one-time step — seedeploy/host/README.md). - A guest kernel and image:
just vmd-devfetches the kernel automatically if missing (just image-kernel-fetch); publish a real agent image built byjust image-build profile=...into/var/tmp/vmd-dev/imageswithimages/publish-to-vmd.sh <manifest.json> /var/tmp/vmd-dev/imagesto get animage_digestforSandboxSpec. deploy/vmd/dev.yaml'snetwork.cidr(default172.28.0.0/16) must not collide with another vmd instance's pool already running on this host — several teammates runvmd/integration_test.goconcurrently on shared dev hosts, and each picks its own range (see the comment indeploy/vmd/dev.yaml); checkip addrfirst ifEnsureever fails to allocate a network slot.just dev-up-storagetoo, and theobject_storeblock indeploy/vmd/dev.yamluncommented, only if you want to exercise backlog task 23's data-disk backup/restore against RustFS.- Don't run
just vmd-devat the same time asgo test -tags=integration ./vmd/...on the same host: both boot real Firecracker VMs, and this shared workstation genuinely runs out of RAM/tmpfs when several Firecracker-heavy processes overlap (see the note on/tmp/kapelle-firecracker.lockbelow) —vmd-devitself doesn't take that lock (it's a long-lived interactive daemon, not a one-shot test), so this is on you to avoid by hand.
Sleep will still fail for any agent_id with no
kapelle/hosts/dev/agents/<id>/... disk-encryption key already written
in OpenBao (normally done by the real controller when it provisions an
agent) — dev.yaml gets vmd talking to a real OpenBao, it doesn't seed
per-agent keys for you.
Talking to it (backlog task 40): vmd's default — and, unless tls:
is set in the config, only — listener is a Unix socket
(/var/tmp/vmd-dev/vmd.sock here; production defaults to
/run/vmd/vmd.sock), not a TCP port; grpc_listen's 127.0.0.1:9090
in dev.yaml is commented out and inert until tls: is also
configured. The socket is mode 0660, group-owned by whatever
-grpc-socket-group the vmd-dev recipe passes (your own primary
group, $(id -gn) — see the recipe), and vmd's own interceptor
double-checks a caller's real SO_PEERCRED uid/gid against that group on
every call (internal/unixcreds) as defense in depth on top of the
file permission itself — root, or a process whose primary (not just
supplementary) group matches, gets through.
- vmdctl (go run ./vmd/cmd/vmdctl <cmd>) defaults to
-addr unix:///run/vmd/vmd.sock; point it at the dev socket instead:
go run ./vmd/cmd/vmdctl capacity -addr unix:///var/tmp/vmd-dev/vmd.sock.
- Python (FirecrackerProvider, services/worker) accepts the same
unix:// address verbatim (grpc.aio resolves it natively); set
KAPELLE_VMD_ADDR=unix:///var/tmp/vmd-dev/vmd.sock to point a local
worker process at vmd-dev instead of its own default
(unix:///run/vmd/vmd.sock, production's path).
- A remote vmd (a different host than the caller) needs the TCP+mTLS
listener instead: set grpc_listen and all three of tls.cert_file/
key_file/ca_file in its config, and pass vmdctl -cert -key -ca
(or FirecrackerProvider(address, channel_credentials=grpc.
ssl_channel_credentials(...))) from the caller — vmd never accepts
an unauthenticated network connection, by construction, so there's no
partial/insecure TCP mode to fall into by omission.
- Certs from OpenBao PKI, not files (backlog task 36): just
dev-seed (kapelle_controller.openbao.seed) mounts a PKI secrets
engine at pki, generates an internal root CA (kapelle-internal-ca,
10y TTL, only if pki has no root yet -- regenerating one is NOT
idempotent and invalidates every cert already issued under the old
root) and defines one role, vmd, allowing vmd.kapelle.internal (and
subdomains), IP SANs, both server_flag/client_flag, 720h max TTL.
A caller holding a vmd-<host> AppRole login (the same one used for
KV) can POST pki/issue/vmd to fetch a cert.
- vmd's server listener: set tls.pki.mount/role/common_name
(and optionally alt_names/ip_sans/ttl, comma-separated)
instead of tls.cert_file/key_file/ca_file -- mutually
exclusive with the file-based trio, and requires openbao to also
be configured (see config.TLSPKIConfig). vmd fetches an initial
certificate before the TCP listener opens and renews it forever in
the background (vmd/internal/pkicert.Manager, halfway through
each certificate's validity window by default) with no restart --
the listener's *tls.Config.GetConfigForClient always reads the
manager's current certificate and CA pool.
- vmdctl: pass -openbao-addr (+ -openbao-token or
-openbao-role-id-file/-openbao-secret-id-file, plus
-pki-mount/-pki-role/-pki-common-name) instead of
-cert/-key/-ca to fetch a client certificate from the same
PKI role. One fetch per invocation, not a renewal loop --
vmdctl is a short-lived CLI, not a daemon.
- FirecrackerProvider (Python): build channel credentials with
kapelle_worker.sandbox.providers.openbao_pki.
fetch_channel_credentials(openbao_addr=..., openbao_token=... or
role_id=.../secret_id=..., common_name=...) and pass the result
as FirecrackerProvider(address, channel_credentials=...), instead
of grpc.ssl_channel_credentials(...) built from files. Also a
single fetch, not a renewal loop -- a long-lived worker process
that needs to outlive the certificate's TTL should call this again
and rebuild its channel.
- vmd/internal/openbao.Client.IssueCert (Go) and
kapelle_worker.sandbox.providers.openbao_pki.issue_cert (Python)
are the two low-level clients underneath all three, both verified
against a real OpenBao PKI mount
(vmd/internal/openbao/pki_test.go,
services/worker/tests/sandbox/test_openbao_pki.py).
/var/tmp, not /tmp: vmd refuses to start with a jail_base_dir or
snapshot_dir on a tmpfs (RAM-backed) filesystem, which /tmp usually
is on Linux — staging a large data disk or snapshotting a large VM's
guest memory there can exhaust the host's actual RAM (this shared host
hit exactly that: tmpfs /tmp filled to 29 GB across several
teammates' concurrent Firecracker runs). /var/tmp is real disk on a
standard Linux layout (verify with findmnt /var/tmp if unsure).
Every Firecracker-spawning test process on this host must flock
/tmp/kapelle-firecracker.lock before booting a real VM (see
vmd/integration_test.go's acquireFirecrackerLock) — one
Firecracker-using test at a time, host-wide, across every teammate's
test binary. This existed because several teammates' vmd
integration/conformance suites running real Firecracker VMs
concurrently exhausted this host's RAM and left ~17 orphaned
Firecracker processes with dangling taps/nftables rules after a test
process died mid-run. If you write a new Firecracker-booting test
(Go or Python), take this same lock before it boots a VM. The wait
gives up after KAPELLE_FIRECRACKER_LOCK_TIMEOUT (a Go duration
string, e.g. 30m; default 15m) in case the holder is genuinely
stuck rather than just slow — raise it for a legitimately long real
workload (a full e2e coder-flow run held it well past the original 5m
default without being stuck) rather than assuming a timeout means a
wedged lock. Never wrap a test run in your own outer flock on this
path (e.g. flock /tmp/kapelle-firecracker.lock pytest ...) — the
fixture itself already acquires the lock, so an outer flock just
deadlocks the run against itself (oh-spike lost 15 minutes to exactly
this). Just run the test directly; let it take the lock. The opposite
case is real too, for anything that does NOT take the lock itself (a
manual go run ./vmd ... repro, images/test-boot.sh before it grew
its own internal flock -w 900 re-exec) — those need an outer
flock /tmp/kapelle-firecracker.lock <command> wrapped around them, or
they're not actually protected by the convention at all.
Never nft flush table inet kapelle on this shared host — it wipes
every other live session's tap rules too, not just your own orphans
(this table has no per-tenant isolation; vmd/test-boot.sh instances tell
their own rules apart only by the vmd:<agent-id> comment on each one).
Remove only your own: sudo nft -a list table inet kapelle to find the
handles for your iifname/comment, then
sudo nft delete rule inet kapelle input handle <N> for each one.
- The Python privileged test suites need
KAPELLE_RUN_PRIVILEGED=1, checked before anything else.eval/e2e/test_vertical_slice_ microvm.py,test_vertical_slice.py's two privileged tests, andservices/worker/tests/sandbox/test_firecracker_provider_integration.pyall boot real Firecracker VMs, gated by a capability check (eval.e2e.conftest.privileged_env_reason()/test_firecracker_provider_integration.py's own_skip_reason()) — but every one of those capability checks (passwordless sudo,/dev/ kvm, the pinned binaries, a built image) can already be satisfied by simple accident of this host's own setup, with nobody having actually decided "run real Firecracker VMs right now."KAPELLE_RUN_PRIVILEGED=1is required first, on top of the capability checks, precisely so a barepytest/uv run pytestinvocation — nodirenv exec, no conscious intent to touch real Firecracker — skips these instead of silently booting real VMs (found the hard way once already: exactly that invocation booted three real VMs on this shared host).just demo/just e2eandjust test-integration(and so the nightly workflow,.github/workflows/nightly-integration.yml) set it explicitly; set it yourself for a deliberate manual run of one of these files.
Troubleshooting¶
A guest can't reach its credential-gateway instance (curl/git/clone hangs from inside a real microVM). Backlog task 29's real hang took several live debugging sessions to diagnose by hand (tap RX/TX counters, per-rule nftables counters, listener state, guest routing) before it became one command:
vmdctl net-diag -agent-id <agent-id> [-addr unix:///run/vmd/vmd.sock]
Prints a plain-text, one-screen report for that agent: its lifecycle
state, touch/idle status, last error, allocated guest/gateway IP, the
tap's real kernel-level RX/TX packet counters (ip -s link show), every
per-tap nftables rule with its live packet/byte counter and position
(these counters are permanent on every rule vmd installs — see
internal/host/netslot.go's install()/AddGatewayAccept), and the
credential-gateway instance's listener ports with a real, direct
listen-state check (not parsed ss text).
How to read it:
- Every rule's counter is 0, including
ct state established,related(the very first rule in the chain). No guest-originated packet ever reached the tap at all — this is not aninet kapelleproblem (nftables never got a chance to run on THAT table). Check the guest's own routing (ip route/ip neighfrom inside the guest, e.g. viaexecute_bash) and your host firewall — see the next section, since this was the actual root cause found on the dev host task 29 was debugged on. - The
ct state established,relatedand anti-spoof rules show hits, but a later drop rule (MMDS/RFC1918/catch-all) shows the hits instead of the expected accept rule. The packet reached the tap and matched the wrong rule — check rule order in the report (rules are listed in real evaluation order): the gateway-accept rules must appear before every destination-based drop, not just the catch-all (this was a real bug, fixed once already — seenetslot.go'sfirstDestinationDropHandledoc comment). - The accept rule shows hits, but the gateway instance's listener shows
NOT LISTENING. iron-proxy isn't actually bound where vmd thinks it is — check credgwd's own logs/config for that agent, and confirm thehttp_listen/https_listenportsnet-diagreports match what credgwd actually configured (-http-port/-https-porton a dev host withoutCAP_NET_BIND_SERVICE). - Everything above looks correct. The problem is downstream of vmd entirely — inside iron-proxy's own request handling, or the destination it's proxying to.
Your host firewall (ufw or similar) can silently drop guest traffic
that net-diag's report never sees at all. This was the actual root
cause found on the dev host task 29 was debugged on: a restrictive
firewall manages its own nftables table (e.g. ufw's table ip filter),
hooked at the same input priority as vmd's own inet kapelle table but
completely independent of it — an INPUT policy of drop with an allow
rule scoped to a narrower range than the vmd network pool's CIDR silently
ate every guest packet, invisible to inet kapelle's own per-tap
counters since it's a different table entirely. vmdctl net-diag cannot
see this (it only reports vmd's own table); vmd itself checks for it
automatically at every startup and logs a WARN naming the offending
table if it finds one (see warnOnForeignFirewallTables in vmd/main.go
— best-effort, read-only, and it never modifies your firewall for you).
To check by hand: sudo nft list ruleset (add sudo ufw status verbose
if ufw specifically is installed) and look for a table with policy
drop; on its input-hooked chain and no rule accepting your vmd pool
CIDR (network.cidr in your vmd config, 172.20.0.0/16 by default —
see internal/config.Default). Fix it with any of:
sudo ufw allow in on tap+
The preferred fix if you're using ufw: every tap vmd creates is named
tap<12 hex chars> (see internal/host.tapName), and ufw has supported
+ as an interface-name wildcard since 0.30 — this trusts traffic
arriving on any vmd tap at the ufw layer and leaves the actual per-agent
filtering to vmd's own inet kapelle table (anti-spoof, allowed ports,
RFC1918/MMDS drops), which does not depend on ufw at all. Narrower than a
source-CIDR rule, since it only matches vmd's own interfaces regardless
of what else is on the network.
sudo ufw allow from <your vmd pool CIDR>
A source-CIDR rule instead, if you'd rather not special-case interface
names, or your firewall isn't ufw (the nft equivalent is an ip saddr
<cidr> accept rule in whichever chain your firewall's own input path
jumps to).
Or, without touching the firewall at all: reconfigure network.cidr to
fall inside a range the firewall already allows (this is what
eval/e2e/conftest.py's own pool-CIDR picker does — it deliberately
stays inside 172.16.0.0/12 for exactly this reason).
A real agent-server rejects something the worker sends with Unknown
kind '<Name>' for <Base>; Expected one of: [...] (a 422, or an
Event.model_validate() crash on the way back). Anything the worker
sends to the agent-server that resolves to an OpenHands SDK
discriminated union — tools, security analyzers, condensers — must be a
kind the AGENT-SERVER's own process has actually imported, never a
custom subclass that exists only in the worker codebase. A
DiscriminatedUnionMixin subclass's wire "kind" is just its own
__name__, and the receiving side's validator only recognizes whatever
classes happen to be imported there — which is a completely different,
independent set from what the worker process has imported, even
though both sides installed the identical openhands-sdk/
openhands-tools/openhands-agent-server version. Hit twice in one
session (backlog tasks afdeb090/7fe12ac8): Agent(tools=[{"name":
"grep"}, {"name": "glob"}]) sent real, importable GrepTool/
GlobTool instances the agent-server itself built and echoed back —
but the worker's own process never imports openhands.tools.grep/
.glob anywhere, so its SystemPromptEvent.tools: list[ToolDefinition]
validator crashed deserializing them; separately, a
KapelleSecurityAnalyzer(PatternSecurityAnalyzer) subclass crashed the
agent-server outright (POST .../security_analyzer → 422) the instant
it was sent, since the agent-server has never heard of a class that
exists only in kapelle_worker. Fix in both directions: only ever
send a tool/analyzer/condenser kind by NAME through the SDK's own
registry-backed shorthand ({"name": "terminal"}, never construct
then hand over a custom subclass instance), and for anything genuinely
custom, configure the vendor's OWN class via its real data fields
(PatternSecurityAnalyzer(high_patterns=..., medium_patterns=...))
rather than subclassing it. Test by actually round-tripping through
the real event model in the worker's own process (e.g. resolve a tool
name via openhands.sdk.tool.registry.resolve_tool and re-parse it
through SystemPromptEvent, or just assert
type(analyzer) is PatternSecurityAnalyzer) — a same-process check
that a class merely exists and looks right is not proof it will
deserialize on the other side of a real agent-server round trip.
No real secrets in the repo¶
deploy/compose/.env and any *.env file are gitignored. Only
deploy/compose/env.example (placeholder values) is committed. OpenBao
runs in -dev mode here purely for local convenience — it is unsealed with
a single well-known root token and stores everything in memory; never point
it at anything containing real production data.