# Kapelle # Kapelle documentation Kapelle runs teams of coding agents in Firecracker microVMs, talking A2A over NATS. This documentation is written once, in Markdown, and published two ways: rendered for people, and as raw Markdown for the agents that manage Kapelle for them. - Every page is also available as Markdown: replace `.html` with `.md`, or ask for `Accept: text/markdown`. - [`llms.txt`](llms.txt) lists every page with one line each; [`llms-full.txt`](llms-full.txt) holds them all in one file. - New here? Start with [Getting started](getting-started.md), then the [architecture](design.md). - Operating a deployment: [the dev environment](dev-environment.md) and [upgrades](upgrades.md). --- # What is Kapelle Kapelle runs teams of coding and other agents. A team is a lead agent plus roles such as coder and reviewer. Work arrives from a chat room, an issue tracker, a repository or the Console, the lead plans it and hands parts to the other roles, and the results go back to where the work came from. - **A team has a home.** A Mattermost or Slack channel, or none (then you work from the Console). A team can also be linked to a GitHub repository, a Linear team or a Jira project; work arriving there is a task for the lead. - **Each agent is a sandbox.** One Firecracker microVM per agent, running the OpenHands agent-server, with a data disk that holds repositories, worktrees and conversation history. An idle agent is snapshotted and its VM process stops, so it costs disk only; new work wakes it. - **No credential enters a sandbox.** A credential gateway adds GitHub, model and registry credentials to the agents' outgoing requests, and only to the hosts you allowed. See [egress and destinations](egress-and-destinations.md). - **No vendor lock-in.** Models go through a LiteLLM proxy with a capped key per agent; sandboxes sit behind a provider interface; agents talk A2A over NATS. - **People can watch and steer.** Every conversation is readable, and a person can answer a question, cancel a run or start work from chat or from the Console. ## What you do with it | You want to | Page | |---|---| | Create a team and give it a task | [Getting started](getting-started.md), [teams](teams.md), [runs](runs.md) | | Use the web Console | [Console tour](console-tour.md) | | Let your own coding agent manage Kapelle | [Onboard your agent](onboard-your-agent.md), [agents](agents.md) | | Change what an agent is told or may reach | [agents and versions](agents-and-versions.md), [egress and destinations](egress-and-destinations.md) | | Connect Slack, GitHub, Linear or Jira | [integrations](integrations.md) | | Run a deployment | [deployment](deployment.md), [dev environment](dev-environment.md), [upgrades](upgrades.md) | ## Not goals Kapelle does not publish a standard A2A broker transport, inject messages into an agent run already in progress, migrate running agents between hosts or fork many sandboxes from one snapshot. The [architecture](design.md) has the full list of goals, constraints and decisions. --- # Getting started From a fresh clone of this repo on this host to a team answering a real task — every command below was actually run, in this order, against a real Postgres/NATS/LiteLLM/OpenBao (`deploy/compose/`), not a mock. For the deeper reference on any of these pieces, see `docs/dev-environment.md` (the exhaustive version of this doc); for what still needs you personally (Slack/GitHub apps, credentials), see `docs/needs-user.md`. ## Prerequisites - `direnv allow` in the repo root once, so `.envrc` (gitignored, holds the Aiven AI Gateway key and other API keys) loads automatically on `cd`. `just dev-up` fails fast with a clear message if these aren't set yet. - Docker + Docker Compose, `uv`, Go 1.27+ — see `docs/dev-environment.md`'s own Prerequisites section for exact versions and how to check them. - `just doctor` checks all of the above at once (`/dev/kvm`, the sudoers NOPASSWD allowlist and its drift from `vmdctl sudoers`' own rendered grant, the pinned `firecracker`/`jailer` binaries, direnv, Docker Compose v2, `deploy/compose/.env`, `uv`/`go`/`just` versions, free `/var/tmp` space in both absolute GiB and percent-free, and `vmdctl host-check` for a provisioned host's own kernel/KSM/swap/SMT settings) — one `[PASS]`/`[WARN]`/`[FAIL]` line per check, nonzero exit only on a real `FAIL` (a workstation failing `vmdctl host-check` is an expected `WARN`, not a failure — that check only turns into a hard `FAIL` when `KAPELLE_HOST_ROLE=server` says this is meant to be a provisioned host). Pass `--json` for a machine-readable report instead of the plain-text lines (what the nightly CI workflow uploads alongside the human-readable run). Run it first on a new host, before anything below; verified for real, 15/15 passed on this one (2 with warnings: `vmdctl host-check` — expected on a plain dev workstation, whose kernel/swap/SMT settings aren't tuned like a provisioned Kapelle host's — and sudoers drift, here because `vmdctl` itself hasn't been built on this host yet to compute the comparison). - `direnv exec /home/mbocevski/dev/factory ` wraps every `docker compose` call in this doc and in the `justfile`'s recipes — never run `docker compose` bare against this repo's compose file; a bare invocation loses the `.envrc` vars several containers need and will force-recreate them, losing real model credentials. ## 1. Bring up the core stack ``` just dev-up just dev-seed ``` `dev-up` starts NATS, Postgres, OpenBao, the OTel collector, Jaeger, and LiteLLM (`deploy/compose/docker-compose.yaml`, no profile flag needed for these — they're the default set). `dev-seed` creates NATS's JetStream streams/KV buckets and OpenBao's policies/AppRole auth, and checks LiteLLM is actually answering. Both are idempotent — safe to re-run. ## 2. Run the three services Schema migrations first (each service owns its own `alembic` tree; none `create_all()`s at runtime): ``` DATABASE_URL="postgresql+asyncpg://:@127.0.0.1:5432/" \ uv run --package kapelle-controller alembic -c services/controller/alembic.ini upgrade head uv run --package kapelle-gateway alembic -c services/gateway/alembic.ini upgrade head uv run --package kapelle-worker alembic -c services/worker/alembic.ini upgrade head ``` (same `DATABASE_URL` for all three — read the actual values out of `deploy/compose/.env`, gitignored, created from `.env.example` — see `docs/dev-environment.md`'s "Running the controller natively" for the exact `grep`-based one-liners the `justfile` itself uses to build this URL from that file, since teammates/agents aren't allowed to read/edit `.env` directly by hand.) The `a2a` and `litellm` databases these services also need already exist at this point — `deploy/compose/postgres/init-databases.sh` creates them as part of the Postgres container's own first-boot init, before any migration runs. Two ways to run the services themselves; pick one. **Native** (what you want for actual development — starts in seconds, no rebuild after a code change): ``` just controller-dev # backgrounded; logs "controller: gRPC listening on 127.0.0.1:8300" just gateway-dev # backgrounded; mounts /a2a// routes as teams appear ``` Team memory (design doc §14) searches by keyword only: the Aiven AI Gateway serves no embedding model yet, so `KAPELLE_EMBED_ALIAS` stays unset -- see docs/dev-environment.md's "Team memory semantic search" section for what would turn semantic search on. There's no `worker-dev` recipe (a real deployment always names a fixed team/role per worker process, so there's nothing generic to default it to) — start one by hand, matching `controller-dev`/`gateway-dev`'s own env-var pattern: ``` KAPELLE_TEAM=demo KAPELLE_ROLE=coder KAPELLE_EXECUTOR=echo \ NATS_URL="nats://127.0.0.1:4222" \ DATABASE_URL="" \ A2A_DATABASE_URL="" \ uv run --package kapelle-worker python -m kapelle_worker.a2a.main ``` `KAPELLE_EXECUTOR=echo` (task 15's stub, no LLM/sandbox in the loop) is the fast path for proving the wiring end to end, exactly what this doc does below. A real agent needs `KAPELLE_EXECUTOR=openhands` plus a real or fake sandbox provider and real model credentials — see `eval/e2e/run_real_eval.py` for a complete, real (not stubbed) example of that whole setup, or `docs/dev-environment.md`'s "Running vmd/credgwd natively" section for the real-Firecracker path. **Containerized** (`services` profile, backlog task 38 — proves the actual container images work, not day-to-day iteration): ``` just dev-up-services ``` Builds and starts `controller`/`gateway`/`worker` from the uv workspace (one shared `deploy/compose/services.Dockerfile`). The compose worker is fixed to `KAPELLE_TEAM=smoke KAPELLE_ROLE=coder KAPELLE_EXECUTOR=echo`. The gateway's HTTP port publishes to a random host port (so it never collides with a teammate's native `gateway-dev` on the same host) — look it up with `direnv exec /home/mbocevski/dev/factory docker compose -f deploy/compose/docker-compose.yaml port gateway 8100`. Everything below in this doc used the native path; swap in the compose gateway's own port and `smoke` in place of `demo` to do the same thing against this profile instead. **Controller migration `b7d2e4f6a8c1` (child tables keyed by `team_id`).** `team_roles`, `surface_links`, `team_members`, the memory `entries`, `skills` and `mcp_servers` were keyed by the team's name; with reusable names (an archived team's name is free for a new team) they now carry `team_id integer NOT NULL`, a foreign key to `teams.id`, and every uniqueness rule is on that id. The `team` name column stays and is still written. What the upgrade does to existing rows: a row goes to the team of its name that is not archived, else to the archived team of that name with the highest id; a memory entry renamed `.archived-` goes to team ``; a row whose name matches no team at all is deleted, with one log line per table naming the count and the names. The downgrade refuses, with a `RuntimeError` naming the duplicate keys, when two teams now hold rows under one name that the name-based constraints would reject; otherwise it restores the name-based schema. Run the controller's migration before starting a gateway of this version: the gateway checks at start that `team_members` and `surface_links` have `team_id` and exits naming this migration if they do not. Back up the database first if it holds data you care about; the delete of unmatched rows is not undone by the downgrade. ## 3. Create a team from the template No CLI tool sends `/team create` outside a real chat surface yet — the most direct way, and how this doc actually did it, is a raw NATS request matching `kapelle_gateway.controller_client.ControllerCommandClient`'s own wire shape, via the `nats-box` container `deploy/compose/` already has: ``` direnv exec /home/mbocevski/dev/factory docker compose --env-file deploy/compose/.env \ -f deploy/compose/docker-compose.yaml run --rm nats-box \ nats --server nats://nats:4222 req kapelle.controller.command \ '{"command":"team","args":["create","demo","--template","feature-team","--repo","https://github.com/mbocevski/factory-playground.git"],"team":null,"surface":"eval","user":{"person_id":"1","display_name":"you","role":"admin"}}' ``` ``` {"ok": true, "text": "Team 'demo' created from template 'feature-team'."} ``` `feature-team` and `feature-team-nevia` set `requires_repo: true`: every role works in a git checkout of the team's repository, so `/team create` refuses them without `--repo ` and answers `this template needs --repo ; usage: /team create --template feature-team --repo [--home slack|mattermost]`. A team that was created without a repository before this rule (its roles have no `repo_url`) is not repaired: a task on it fails at once with `team has no repository, so has nothing to work on; create the team with /team create --template