Upgrade runbook¶
Backlog task 25; docs/design.md §16 "Upgrades" / §8. Companion to
versions.yaml (every pinned component) and deploy/host/check-versions.sh
(just versions — compares every pin against its upstream feed).
Rules¶
- One component at a time. Never bump two unrelated pins in the same change — if the evaluation suite (task 34, not built yet) or a smoke test regresses, you need to know which bump caused it.
- Gate on the evaluation suite (task 34) once it exists: run it against the new version before rolling out past a canary host. Until task 34 lands, use each component's smoke test below instead — it's a floor, not a substitute.
- Check before you pin.
just versionsfirst, every time — never assume a "latest" from memory (CLAUDE.md's own rule, made executable). It exits non-zero if anything is behind and prints the real current upstream version to pin to. - Drain before any host-wide change.
vmdctl drain -addr <vmd-addr>(backlog task 24) before an OpenHands/Firecracker/kernel upgrade that touches a host's running agents. See "OpenHands upgrade" and "Firecracker / guest kernel upgrade" below for what "drained" means for each. - Update every
source:file together.versions.yaml's per- componentsourcelist is exactly the set of files one pin bump must touch in the same commit — a component pinned in two places (openhands-sdkis in bothservices/worker/pyproject.tomlandimages/agent/Containerfile) drifting apart is its own bug. - Query the endpoints we actually use on the new image before the pin
is committed, not after. Real incident: Jaeger 2.21.0 removed its
legacy
/api/traces//api/servicesquery API outright — a genuine breaking change, the FIRST bullet under "Breaking Changes" in that release's own notes ("remove v1 http endpoints the ui no longer calls") — and this repo's own telemetry test, the e2e trace check, and Grafana's Jaeger datasource all depended on it. The bump was committed, recreated live, broke the telemetry test, and had to be reverted (e8e4c15) before being redone correctly (2b67b6a0). The release notes said so plainly; nobody checked them against what this repo actually calls before committing the pin. For any component this repo queries at runtime (not just imports/links against), read that version's own release notes for API-surface changes AND hit the specific endpoints/calls this repo makes against a throwaway container on the new version, before touchingversions.yaml.
Per-component procedures¶
Firecracker / jailer¶
just versions | grep -E '^(firecracker|jailer)\b' # confirm what's behind, and the real latest tag
# Edit deploy/host/fetch-firecracker.sh: VERSION="vX.Y.Z"
./deploy/host/fetch-firecracker.sh /tmp/fc-upgrade-check # verify checksum + `jailer --version` before rolling to any real host
- Invalidates every snapshot on the fleet (see "Firecracker / guest kernel upgrade" below) — this is not a routine bump, plan the drain/cold-boot cost in.
- Rollback: revert
VERSIONinfetch-firecracker.sh, re-runprovision.sh --applyon affected hosts to reinstall the old binaries (they're content-addressed by checksum, not overwritten in place until the script runs).
Guest kernel¶
cat images/kernel/MANIFEST.json # current url/sha256/caveat
No automated upstream check exists (versions.yaml's guest-kernel
entry is type: manual — see its note: Firecracker's CI artifact
bucket isn't a stable, versioned feed). To bump:
- Get a fresh kernel build from the same CI bucket (or a real upstream
release, once one exists — see
images/README.md's tracked follow-up). - Independently verify its checksum yourself (there is no published one to diff against) before trusting it.
- Update
images/kernel/MANIFEST.json'surl/sha256together, in the same commit as the checksum you just verified. - Invalidates every snapshot on the fleet — same rollout plan as a Firecracker bump.
- Rollback: revert
MANIFEST.json, re-runimages/kernel/fetch.sh.
iron-proxy¶
just versions | grep iron-proxy
# Edit credgw/scripts/fetch-iron-proxy.sh: VERSION="vX.Y.Z"
./credgw/scripts/fetch-iron-proxy.sh /tmp/iron-proxy-upgrade-check
- Smoke test:
go test -race ./credgw/manager/...(the real-iron-proxy integration test —credgw/manager/integration_test.go) against the new binary (CREDGW_IRON_PROXY_BIN=/tmp/iron-proxy-upgrade-check/iron-proxy). - Re-check on every bump (Console K7,
mode: allow_all): the catch-all rule is{host: "*"}in the allowlist transform, which matches every host throughhostmatch.MatchGlob, and the rule config has no port field, soCONNECTto any port is allowed in allow-all mode. The real-iron-proxy testTestIntegration_EgressModeAllowlistRefusesAndAllowAllReachesfails first if a new version stops matching*. - Does not invalidate snapshots (it isn't part of a VM's memory state) — safe to roll out independently of a Firecracker/kernel bump.
- Rollback: revert
VERSION, re-run the fetch script, restartcredgwd(systemctl restart credgwd).
OpenHands (openhands-sdk/-tools/-workspace/-agent-server)¶
Always bumped together, in both places (versions.yaml's source
list):
just versions | grep openhands
# services/worker/pyproject.toml AND images/agent/Containerfile:
# every openhands-{sdk,tools,workspace,agent-server}==X.Y.Z -> the same new version, in both files
uv lock --package kapelle-worker
just image-build profile=python # and profile=node, rebuild the agent image
This is design doc §16's deep-sleep cycle, not a rolling restart — persisted conversation events aren't guaranteed to load across OpenHands versions:
# 1. Drain every host running the old image (backlog task 24):
vmdctl drain -addr <vmd-addr> # per host; moves every RUNNING sandbox to ASLEEP
# 2. In-flight work items finish on the OLD image (drain doesn't kill
# them, it lets Sleep's own idle check decide) -- wait for the fleet
# to actually reach ASLEEP/DEEP_SLEEP before proceeding:
vmdctl list -addr <vmd-addr> # repeat until no RUNNING sandboxes remain
# 3. Archive conversations (design doc §16: "conversations are archived"
# before the new image boots with empty state) -- per-agent conversation
# data lives in that agent's data disk; back it up before the next
# step touches the image every NEW boot uses:
# (task 11/21's sandbox provider owns the actual archival mechanism;
# not yet implemented as of this runbook -- placeholder step, do not
# skip silently once it exists)
# 4. Roll out the new image (images/build.sh's output) to hosts.
# 5. New Wakes cold-boot on the new image with empty conversation state
# (expected, not a bug -- this is the whole point of the deep-sleep
# cycle for an OpenHands bump specifically).
- Smoke test: the evaluation suite (task 34) once it exists; until then,
manually run one real task end-to-end against the new image
(
just credgw-dev+ a throwaway repo, orFakeProvider's real- agent-server path,services/worker/src/kapelle_worker/sandbox/providers/fake.py). - Rollback: revert both files' version pins,
uv lock, rebuild the image, roll out — same deep-sleep-cycle procedure in reverse (old conversations are already archived/gone either way, so rollback is not "free": it's a second version bump, not an undo).
LiteLLM¶
v1.100.1 -> v1.103.0 (backlog task c96850b4): without a Redis-compatible
store, a key's token counter could be stuck below its real limit for
minutes at a time (docs/upstream-reports/litellm-in-memory-rate-limiter.md
has the full trace) -- v1.103.0 doesn't fix the underlying defect but
erases its visible effect. Reproduced clean on v1.103.0 across 3/3
throwaway replay runs of the real incident's own traffic, and 2/2 stalled
on v1.100.1 under the same traffic.
just versions | grep litellm
# deploy/compose/docker-compose.yaml: image: ghcr.io/berriai/litellm:vX.Y.Z
docker compose --env-file deploy/compose/.env -f deploy/compose/docker-compose.yaml up -d litellm
- Smoke test (tool-calling, design doc §10's requirement — "Tool calling must be enabled explicitly"): send a tool-call-shaped request through the running proxy and check the response actually contains a tool call, not just prose:
curl -s http://localhost:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"local-coder","messages":[{"role":"user","content":"What is the weather in Boston?"}],"tools":[{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"location":{"type":"string"}}}}}]}' \
| jq '.choices[0].message.tool_calls'
A null result (no tool call produced for an unambiguous tool-call
prompt) is a smoke-test failure — do not roll the new version out
further. This is a floor, not the real gate: task 34's evaluation
suite is the actual admission bar per design doc §10 ("A model is
admitted for a role only after passing the evaluation suite").
- Rollback: revert the image tag, docker compose up -d litellm.
- Trying a new image before the pin: deploy/compose/tests/
test_services_smoke_integration.py reads KAPELLE_SMOKE_COMPOSE_OVERRIDE
(a compose override file's path) and merges it in as a third -f, so the
smoke module runs against a different LiteLLM image without ever
modifying docker-compose.yaml -- its own output states which image the
container actually ran, read from docker inspect, not from either file.
vLLM¶
Design doc §10's target self-hosted model-serving version
(versions.yaml's vllm entry) — not yet deployed by this repo
(no vLLM service exists in deploy/compose/ as of this runbook). Once
deployed:
just versions | grep vllm
- Must be started with
--enable-auto-tool-choiceand the model's own--tool-call-parser(design doc §10) — a bump that changes the correct parser name for a given model family is a real regression risk, not just a version number; check the new version's release notes for tool-call-parser changes specifically before bumping. - Same tool-calling smoke test as LiteLLM above, run directly against the vLLM endpoint before LiteLLM is even pointed at it.
- Rollback: same as LiteLLM — revert the pin, redeploy.
NATS / Postgres / OpenBao / RustFS / Mattermost / otel-collector / Jaeger / Prometheus / Grafana¶
All plain image bumps in deploy/compose/docker-compose.yaml:
just versions # confirm which of these are actually behind
# edit the relevant image: line(s) in deploy/compose/docker-compose.yaml
docker compose --env-file deploy/compose/.env -f deploy/compose/docker-compose.yaml up -d <service>
just test # full suite -- several integration tests are Docker-gated against these
- Postgres major-version bumps specifically need a real upgrade path
(
pg_upgradeor dump/restore), never just swapping the image tag on an existing volume — a minor/patch bump (18.5-alpine->18.6-alpine) is safe to roll by image swap alone. - Rollback: revert the image tag; for anything with a persistent volume (Postgres, NATS JetStream, OpenBao in non-dev mode), confirm the old image can still read the new image's on-disk format before you need to find out under pressure — check release notes, don't assume.
- OpenBao (
docs/dev-environment.md, "Keeping OpenBao's data across restarts"): withOPENBAO_STATIC_UNSEAL_KEYset the dev store is durable (raft on theopenbao-datavolume), so a bump is an ordinary image swap: firstbao operator raft snapshot saveinside the container, then change the tag; rollback is the same swap with that snapshot as the safety net (OpenBao does not promise that an older version reads a newer one's data). Without the key the dev server is in memory and ANY recreate, bump or rollback, empties it, including every agent's secret document (just doctornames the teams that lost theirs). - NATS 2.14 -> 2.15: the upgrade guide only supports going back to
2.14.7 or later, so a pin below 2.14.7 is not a rollback target. A scale or
move begun on 2.15 (the new desired-state metalayer) does not complete
after a downgrade to 2.14 -- send a no-op stream/consumer update with the
exact same config first to strip that state before rolling back, if one
was in flight; dev never runs one. Ordinary stream/message storage does
not change format on a 2.15 write -- confirmed against the server's own
source and the upgrade guide -- so 2.14.7+ reads back what 2.15 wrote for
the data itself, that one caveat aside. 2.15 also caps a stream at 1000
consumers by default (
default_max_consumersin the server's JetStream limits, ormax_consumerson the stream, raises it; dev is far below -- A2A_RPC, the one work-queue stream, gets exactly one durable consumer per (team, role) pair).just nats-backup <dir> [<server>]backs up the streams that hold real history (A2A_TASKS, A2A_ACTIVITY) and the a2a-cards KV bucket before a recreate;nats backup restore stream <dir>(against a server where that stream doesn't already exist) is the way back for one of them. - Jaeger 2.20 -> 2.21 (real incident, backlog task 2b67b6a0): 2.21.0
removed the legacy
/api/traces//api/servicesquery API outright — a breaking change stated plainly in that release's own notes, missed before the pin was first bumped and reverted./api/v3/*(OTLP-shaped JSON, different response structure entirely) works on both versions; confirm anything that queries Jaeger directly (not just the UI) uses v3 before bumping past 2.20 again. Grafana's owntype: jaegerdatasource already speaks v3 internally as of Grafana 13.0.0+ — nothing in this repo's Grafana config needs to change for that specifically.
Go / Python toolchains¶
just versions | grep -E '^(go|python)\b'
# go.mod: go X.Y (+ .github/workflows/ci.yml's go-version)
# .python-version, pyproject.toml's requires-python, contracts/python/pyproject.toml's requires-python
just lint && just test # full suite, both languages
- Rollback: revert the version file(s), re-run
just lint && just test.
Python dev tooling (ruff / ty / pytest) and GitHub Actions¶
just versions | grep -E '^(ruff|ty|pytest|actions-|golangci-lint)\b'
# root pyproject.toml's [dependency-groups] dev list, or the relevant
# `uses:` line in .github/workflows/ci.yml
uv lock
just lint
- Low risk, no fleet impact — bump freely once
just lint/just testpass, no drain or eval-suite gate needed.
Database migrations (three independent Alembic trees)¶
services/{worker,gateway,controller} each have their own migrations/
tree, all targeting the same kapelle Postgres database
(alembic.ini's sqlalchemy.url, overridable via DATABASE_URL) but
with no cross-tree foreign keys today (checked: only
services/controller's own internal users table is referenced by
anything, and only from within its own tree) — so tree order doesn't
matter yet. If a future migration adds a cross-tree foreign key, order
will start to matter and this section must be updated to say so
explicitly, not left silently wrong.
Before upgrading anything that ships a new migration:
# One head per tree, always -- more than one means a branched migration
# history that needs `alembic merge` before anything else:
for svc in controller worker gateway; do
echo "=== $svc ==="
uv run --package "kapelle-$svc" alembic -c "services/$svc/alembic.ini" heads
done
# Apply, in this repo's convention order (foundational entities first --
# controller owns teams/hosts, worker/gateway don't reference them today
# but might; keep this order even though nothing enforces it yet):
for svc in controller worker gateway; do
uv run --package "kapelle-$svc" alembic -c "services/$svc/alembic.ini" upgrade head
done
- Rollback:
alembic downgrade -1per tree, same order reversed (gateway, worker, controller) — never downgrade a tree whose migration another tree's code already depends on without checking first.
entries.embedding: converting to vector after the fact (backlog task 84811aac)¶
0c9a7a964933 and ec544ab1edb4 (backlog task 9359758c) both check
pg_available_extensions before touching entries.embedding, and leave
it as Text when pgvector isn't installed yet — correctly, since
CREATE EXTENSION on a Postgres that doesn't have pgvector would fail
the whole migration. Alembic marks a revision applied the moment it
runs, successfully or as a no-op, and never re-runs it, so a database
that migrated before pgvector was available does not pick the
extension up automatically once it's installed later (this is exactly
the state EmbeddingColumnNotVectorError catches at controller startup
if KAPELLE_EMBED_ALIAS is then set).
Once pgvector is available on that database, run the explicit, operator-run command instead of a migration:
# Dry run (default) -- reports the column's current type, how many rows
# have a non-NULL value, and names any that would not convert, without
# changing anything:
uv run --package kapelle-controller python \
scripts/convert_memory_embeddings_to_vector.py --database-url "$DATABASE_URL"
# Once the dry run reports every row would convert cleanly:
uv run --package kapelle-controller python \
scripts/convert_memory_embeddings_to_vector.py --database-url "$DATABASE_URL" --apply
Safe to run twice (reports "nothing to do", exit 0, once the column is
already vector) and safe in a start script. If any existing row's text
isn't a well-formed 1024-dimension vector literal, the command stops
before any ALTER and names the bad rows by id — it never deletes or
rewrites a value on its own, since what to do about a bad row is a
person's decision. If the connected role lacks CREATE EXTENSION
rights, it says so and exits 2 without changing anything.
Firecracker / guest kernel upgrade — snapshot invalidation¶
Both bumps change vmd/internal/fingerprint.Current()'s
FirecrackerVersion/KernelDigest fields. Design doc §16: "A Firecracker
or kernel upgrade invalidates snapshots, so agents cold-boot afterwards."
Concretely:
vmdctl drain -addr <vmd-addr>— every agent reaches at least ASLEEP (its own encrypted snapshot bundle, tagged with the OLD fingerprint).- Roll out the new Firecracker/jailer binaries and/or kernel
(
deploy/host/provision.sh --applypicks up a newfetch-firecracker.shpin;images/kernel/fetch.shfor a kernel bump) and restartvmd(systemctl restart vmd). - The next
Wakefor each agent computes a fresh fingerprint, finds it doesn't match the snapshot's recorded one, and falls back to cold boot automatically (vmd/internal/manager/restore.go— this is the existing, already-implemented fallback path, not new work this runbook is asking for). No manual snapshot deletion needed; expect every agent's next wake to pay full cold-boot latency instead of restore latency, once. - Rollback: revert the binary/kernel pin, restart
vmd— snapshots taken on the OLD fingerprint are still valid once you're back on it (nothing deletes them on a fingerprint mismatch, it just skips using them).
Host OS-level changes (KSM / swap / SMT / cgroups / kernel line)¶
Not a "component" in versions.yaml, but still a one-host-at-a-time,
drain-first operation:
vmdctl drain -addr <vmd-addr>
sudo deploy/host/provision.sh --check # confirm exactly what will change
sudo deploy/host/provision.sh --apply
A host kernel-line switch or cgroup_favordynmods/
systemd.unified_cgroup_hierarchy=1 change needs a reboot
(provision.sh warns, never reboots anything itself) — that's also a
Firecracker/kernel-upgrade-shaped event for any agent still on that host
(see above): drain first, reboot, expect cold boots on the next wake.