Skip to content

Upgrade runbook

Backlog task 25; docs/design.md §16 "Upgrades" / §8. Companion to versions.yaml (every pinned component) and deploy/host/check-versions.sh (just versions — compares every pin against its upstream feed).

Rules

  1. One component at a time. Never bump two unrelated pins in the same change — if the evaluation suite (task 34, not built yet) or a smoke test regresses, you need to know which bump caused it.
  2. Gate on the evaluation suite (task 34) once it exists: run it against the new version before rolling out past a canary host. Until task 34 lands, use each component's smoke test below instead — it's a floor, not a substitute.
  3. Check before you pin. just versions first, every time — never assume a "latest" from memory (CLAUDE.md's own rule, made executable). It exits non-zero if anything is behind and prints the real current upstream version to pin to.
  4. Drain before any host-wide change. vmdctl drain -addr <vmd-addr> (backlog task 24) before an OpenHands/Firecracker/kernel upgrade that touches a host's running agents. See "OpenHands upgrade" and "Firecracker / guest kernel upgrade" below for what "drained" means for each.
  5. Update every source: file together. versions.yaml's per- component source list is exactly the set of files one pin bump must touch in the same commit — a component pinned in two places (openhands-sdk is in both services/worker/pyproject.toml and images/agent/Containerfile) drifting apart is its own bug.
  6. Query the endpoints we actually use on the new image before the pin is committed, not after. Real incident: Jaeger 2.21.0 removed its legacy /api/traces//api/services query API outright — a genuine breaking change, the FIRST bullet under "Breaking Changes" in that release's own notes ("remove v1 http endpoints the ui no longer calls") — and this repo's own telemetry test, the e2e trace check, and Grafana's Jaeger datasource all depended on it. The bump was committed, recreated live, broke the telemetry test, and had to be reverted (e8e4c15) before being redone correctly (2b67b6a0). The release notes said so plainly; nobody checked them against what this repo actually calls before committing the pin. For any component this repo queries at runtime (not just imports/links against), read that version's own release notes for API-surface changes AND hit the specific endpoints/calls this repo makes against a throwaway container on the new version, before touching versions.yaml.

Per-component procedures

Firecracker / jailer

just versions | grep -E '^(firecracker|jailer)\b'   # confirm what's behind, and the real latest tag
# Edit deploy/host/fetch-firecracker.sh: VERSION="vX.Y.Z"
./deploy/host/fetch-firecracker.sh /tmp/fc-upgrade-check   # verify checksum + `jailer --version` before rolling to any real host
  • Invalidates every snapshot on the fleet (see "Firecracker / guest kernel upgrade" below) — this is not a routine bump, plan the drain/cold-boot cost in.
  • Rollback: revert VERSION in fetch-firecracker.sh, re-run provision.sh --apply on affected hosts to reinstall the old binaries (they're content-addressed by checksum, not overwritten in place until the script runs).

Guest kernel

cat images/kernel/MANIFEST.json   # current url/sha256/caveat

No automated upstream check exists (versions.yaml's guest-kernel entry is type: manual — see its note: Firecracker's CI artifact bucket isn't a stable, versioned feed). To bump:

  1. Get a fresh kernel build from the same CI bucket (or a real upstream release, once one exists — see images/README.md's tracked follow-up).
  2. Independently verify its checksum yourself (there is no published one to diff against) before trusting it.
  3. Update images/kernel/MANIFEST.json's url/sha256 together, in the same commit as the checksum you just verified.
  4. Invalidates every snapshot on the fleet — same rollout plan as a Firecracker bump.
  5. Rollback: revert MANIFEST.json, re-run images/kernel/fetch.sh.

iron-proxy

just versions | grep iron-proxy
# Edit credgw/scripts/fetch-iron-proxy.sh: VERSION="vX.Y.Z"
./credgw/scripts/fetch-iron-proxy.sh /tmp/iron-proxy-upgrade-check
  • Smoke test: go test -race ./credgw/manager/... (the real-iron-proxy integration test — credgw/manager/integration_test.go) against the new binary (CREDGW_IRON_PROXY_BIN=/tmp/iron-proxy-upgrade-check/iron-proxy).
  • Re-check on every bump (Console K7, mode: allow_all): the catch-all rule is {host: "*"} in the allowlist transform, which matches every host through hostmatch.MatchGlob, and the rule config has no port field, so CONNECT to any port is allowed in allow-all mode. The real-iron-proxy test TestIntegration_EgressModeAllowlistRefusesAndAllowAllReaches fails first if a new version stops matching *.
  • Does not invalidate snapshots (it isn't part of a VM's memory state) — safe to roll out independently of a Firecracker/kernel bump.
  • Rollback: revert VERSION, re-run the fetch script, restart credgwd (systemctl restart credgwd).

OpenHands (openhands-sdk/-tools/-workspace/-agent-server)

Always bumped together, in both places (versions.yaml's source list):

just versions | grep openhands
# services/worker/pyproject.toml AND images/agent/Containerfile:
#   every openhands-{sdk,tools,workspace,agent-server}==X.Y.Z -> the same new version, in both files
uv lock --package kapelle-worker
just image-build profile=python   # and profile=node, rebuild the agent image

This is design doc §16's deep-sleep cycle, not a rolling restart — persisted conversation events aren't guaranteed to load across OpenHands versions:

# 1. Drain every host running the old image (backlog task 24):
vmdctl drain -addr <vmd-addr>   # per host; moves every RUNNING sandbox to ASLEEP

# 2. In-flight work items finish on the OLD image (drain doesn't kill
#    them, it lets Sleep's own idle check decide) -- wait for the fleet
#    to actually reach ASLEEP/DEEP_SLEEP before proceeding:
vmdctl list -addr <vmd-addr>   # repeat until no RUNNING sandboxes remain

# 3. Archive conversations (design doc §16: "conversations are archived"
#    before the new image boots with empty state) -- per-agent conversation
#    data lives in that agent's data disk; back it up before the next
#    step touches the image every NEW boot uses:
#    (task 11/21's sandbox provider owns the actual archival mechanism;
#    not yet implemented as of this runbook -- placeholder step, do not
#    skip silently once it exists)

# 4. Roll out the new image (images/build.sh's output) to hosts.

# 5. New Wakes cold-boot on the new image with empty conversation state
#    (expected, not a bug -- this is the whole point of the deep-sleep
#    cycle for an OpenHands bump specifically).
  • Smoke test: the evaluation suite (task 34) once it exists; until then, manually run one real task end-to-end against the new image (just credgw-dev + a throwaway repo, or FakeProvider's real- agent-server path, services/worker/src/kapelle_worker/sandbox/providers/fake.py).
  • Rollback: revert both files' version pins, uv lock, rebuild the image, roll out — same deep-sleep-cycle procedure in reverse (old conversations are already archived/gone either way, so rollback is not "free": it's a second version bump, not an undo).

LiteLLM

v1.100.1 -> v1.103.0 (backlog task c96850b4): without a Redis-compatible store, a key's token counter could be stuck below its real limit for minutes at a time (docs/upstream-reports/litellm-in-memory-rate-limiter.md has the full trace) -- v1.103.0 doesn't fix the underlying defect but erases its visible effect. Reproduced clean on v1.103.0 across 3/3 throwaway replay runs of the real incident's own traffic, and 2/2 stalled on v1.100.1 under the same traffic.

just versions | grep litellm
# deploy/compose/docker-compose.yaml: image: ghcr.io/berriai/litellm:vX.Y.Z
docker compose --env-file deploy/compose/.env -f deploy/compose/docker-compose.yaml up -d litellm
  • Smoke test (tool-calling, design doc §10's requirement — "Tool calling must be enabled explicitly"): send a tool-call-shaped request through the running proxy and check the response actually contains a tool call, not just prose:
curl -s http://localhost:4000/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"local-coder","messages":[{"role":"user","content":"What is the weather in Boston?"}],"tools":[{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"location":{"type":"string"}}}}}]}' \
  | jq '.choices[0].message.tool_calls'

A null result (no tool call produced for an unambiguous tool-call prompt) is a smoke-test failure — do not roll the new version out further. This is a floor, not the real gate: task 34's evaluation suite is the actual admission bar per design doc §10 ("A model is admitted for a role only after passing the evaluation suite"). - Rollback: revert the image tag, docker compose up -d litellm. - Trying a new image before the pin: deploy/compose/tests/ test_services_smoke_integration.py reads KAPELLE_SMOKE_COMPOSE_OVERRIDE (a compose override file's path) and merges it in as a third -f, so the smoke module runs against a different LiteLLM image without ever modifying docker-compose.yaml -- its own output states which image the container actually ran, read from docker inspect, not from either file.

vLLM

Design doc §10's target self-hosted model-serving version (versions.yaml's vllm entry) — not yet deployed by this repo (no vLLM service exists in deploy/compose/ as of this runbook). Once deployed:

just versions | grep vllm
  • Must be started with --enable-auto-tool-choice and the model's own --tool-call-parser (design doc §10) — a bump that changes the correct parser name for a given model family is a real regression risk, not just a version number; check the new version's release notes for tool-call-parser changes specifically before bumping.
  • Same tool-calling smoke test as LiteLLM above, run directly against the vLLM endpoint before LiteLLM is even pointed at it.
  • Rollback: same as LiteLLM — revert the pin, redeploy.

NATS / Postgres / OpenBao / RustFS / Mattermost / otel-collector / Jaeger / Prometheus / Grafana

All plain image bumps in deploy/compose/docker-compose.yaml:

just versions   # confirm which of these are actually behind
# edit the relevant image: line(s) in deploy/compose/docker-compose.yaml
docker compose --env-file deploy/compose/.env -f deploy/compose/docker-compose.yaml up -d <service>
just test       # full suite -- several integration tests are Docker-gated against these
  • Postgres major-version bumps specifically need a real upgrade path (pg_upgrade or dump/restore), never just swapping the image tag on an existing volume — a minor/patch bump (18.5-alpine -> 18.6-alpine) is safe to roll by image swap alone.
  • Rollback: revert the image tag; for anything with a persistent volume (Postgres, NATS JetStream, OpenBao in non-dev mode), confirm the old image can still read the new image's on-disk format before you need to find out under pressure — check release notes, don't assume.
  • OpenBao (docs/dev-environment.md, "Keeping OpenBao's data across restarts"): with OPENBAO_STATIC_UNSEAL_KEY set the dev store is durable (raft on the openbao-data volume), so a bump is an ordinary image swap: first bao operator raft snapshot save inside the container, then change the tag; rollback is the same swap with that snapshot as the safety net (OpenBao does not promise that an older version reads a newer one's data). Without the key the dev server is in memory and ANY recreate, bump or rollback, empties it, including every agent's secret document (just doctor names the teams that lost theirs).
  • NATS 2.14 -> 2.15: the upgrade guide only supports going back to 2.14.7 or later, so a pin below 2.14.7 is not a rollback target. A scale or move begun on 2.15 (the new desired-state metalayer) does not complete after a downgrade to 2.14 -- send a no-op stream/consumer update with the exact same config first to strip that state before rolling back, if one was in flight; dev never runs one. Ordinary stream/message storage does not change format on a 2.15 write -- confirmed against the server's own source and the upgrade guide -- so 2.14.7+ reads back what 2.15 wrote for the data itself, that one caveat aside. 2.15 also caps a stream at 1000 consumers by default (default_max_consumers in the server's JetStream limits, or max_consumers on the stream, raises it; dev is far below -- A2A_RPC, the one work-queue stream, gets exactly one durable consumer per (team, role) pair). just nats-backup <dir> [<server>] backs up the streams that hold real history (A2A_TASKS, A2A_ACTIVITY) and the a2a-cards KV bucket before a recreate; nats backup restore stream <dir> (against a server where that stream doesn't already exist) is the way back for one of them.
  • Jaeger 2.20 -> 2.21 (real incident, backlog task 2b67b6a0): 2.21.0 removed the legacy /api/traces//api/services query API outright — a breaking change stated plainly in that release's own notes, missed before the pin was first bumped and reverted. /api/v3/* (OTLP-shaped JSON, different response structure entirely) works on both versions; confirm anything that queries Jaeger directly (not just the UI) uses v3 before bumping past 2.20 again. Grafana's own type: jaeger datasource already speaks v3 internally as of Grafana 13.0.0+ — nothing in this repo's Grafana config needs to change for that specifically.

Go / Python toolchains

just versions | grep -E '^(go|python)\b'
# go.mod: go X.Y   (+ .github/workflows/ci.yml's go-version)
# .python-version, pyproject.toml's requires-python, contracts/python/pyproject.toml's requires-python
just lint && just test   # full suite, both languages
  • Rollback: revert the version file(s), re-run just lint && just test.

Python dev tooling (ruff / ty / pytest) and GitHub Actions

just versions | grep -E '^(ruff|ty|pytest|actions-|golangci-lint)\b'
# root pyproject.toml's [dependency-groups] dev list, or the relevant
# `uses:` line in .github/workflows/ci.yml
uv lock
just lint
  • Low risk, no fleet impact — bump freely once just lint/just test pass, no drain or eval-suite gate needed.

Database migrations (three independent Alembic trees)

services/{worker,gateway,controller} each have their own migrations/ tree, all targeting the same kapelle Postgres database (alembic.ini's sqlalchemy.url, overridable via DATABASE_URL) but with no cross-tree foreign keys today (checked: only services/controller's own internal users table is referenced by anything, and only from within its own tree) — so tree order doesn't matter yet. If a future migration adds a cross-tree foreign key, order will start to matter and this section must be updated to say so explicitly, not left silently wrong.

Before upgrading anything that ships a new migration:

# One head per tree, always -- more than one means a branched migration
# history that needs `alembic merge` before anything else:
for svc in controller worker gateway; do
  echo "=== $svc ==="
  uv run --package "kapelle-$svc" alembic -c "services/$svc/alembic.ini" heads
done

# Apply, in this repo's convention order (foundational entities first --
# controller owns teams/hosts, worker/gateway don't reference them today
# but might; keep this order even though nothing enforces it yet):
for svc in controller worker gateway; do
  uv run --package "kapelle-$svc" alembic -c "services/$svc/alembic.ini" upgrade head
done
  • Rollback: alembic downgrade -1 per tree, same order reversed (gateway, worker, controller) — never downgrade a tree whose migration another tree's code already depends on without checking first.

entries.embedding: converting to vector after the fact (backlog task 84811aac)

0c9a7a964933 and ec544ab1edb4 (backlog task 9359758c) both check pg_available_extensions before touching entries.embedding, and leave it as Text when pgvector isn't installed yet — correctly, since CREATE EXTENSION on a Postgres that doesn't have pgvector would fail the whole migration. Alembic marks a revision applied the moment it runs, successfully or as a no-op, and never re-runs it, so a database that migrated before pgvector was available does not pick the extension up automatically once it's installed later (this is exactly the state EmbeddingColumnNotVectorError catches at controller startup if KAPELLE_EMBED_ALIAS is then set).

Once pgvector is available on that database, run the explicit, operator-run command instead of a migration:

# Dry run (default) -- reports the column's current type, how many rows
# have a non-NULL value, and names any that would not convert, without
# changing anything:
uv run --package kapelle-controller python \
  scripts/convert_memory_embeddings_to_vector.py --database-url "$DATABASE_URL"

# Once the dry run reports every row would convert cleanly:
uv run --package kapelle-controller python \
  scripts/convert_memory_embeddings_to_vector.py --database-url "$DATABASE_URL" --apply

Safe to run twice (reports "nothing to do", exit 0, once the column is already vector) and safe in a start script. If any existing row's text isn't a well-formed 1024-dimension vector literal, the command stops before any ALTER and names the bad rows by id — it never deletes or rewrites a value on its own, since what to do about a bad row is a person's decision. If the connected role lacks CREATE EXTENSION rights, it says so and exits 2 without changing anything.

Firecracker / guest kernel upgrade — snapshot invalidation

Both bumps change vmd/internal/fingerprint.Current()'s FirecrackerVersion/KernelDigest fields. Design doc §16: "A Firecracker or kernel upgrade invalidates snapshots, so agents cold-boot afterwards." Concretely:

  1. vmdctl drain -addr <vmd-addr> — every agent reaches at least ASLEEP (its own encrypted snapshot bundle, tagged with the OLD fingerprint).
  2. Roll out the new Firecracker/jailer binaries and/or kernel (deploy/host/provision.sh --apply picks up a new fetch-firecracker.sh pin; images/kernel/fetch.sh for a kernel bump) and restart vmd (systemctl restart vmd).
  3. The next Wake for each agent computes a fresh fingerprint, finds it doesn't match the snapshot's recorded one, and falls back to cold boot automatically (vmd/internal/manager/restore.go — this is the existing, already-implemented fallback path, not new work this runbook is asking for). No manual snapshot deletion needed; expect every agent's next wake to pay full cold-boot latency instead of restore latency, once.
  4. Rollback: revert the binary/kernel pin, restart vmd — snapshots taken on the OLD fingerprint are still valid once you're back on it (nothing deletes them on a fingerprint mismatch, it just skips using them).

Host OS-level changes (KSM / swap / SMT / cgroups / kernel line)

Not a "component" in versions.yaml, but still a one-host-at-a-time, drain-first operation:

vmdctl drain -addr <vmd-addr>
sudo deploy/host/provision.sh --check   # confirm exactly what will change
sudo deploy/host/provision.sh --apply

A host kernel-line switch or cgroup_favordynmods/ systemd.unified_cgroup_hierarchy=1 change needs a reboot (provision.sh warns, never reboots anything itself) — that's also a Firecracker/kernel-upgrade-shaped event for any agent still on that host (see above): drain first, reboot, expect cold boots on the next wake.