Skip to content

Getting started

From a fresh clone of this repo on this host to a team answering a real task — every command below was actually run, in this order, against a real Postgres/NATS/LiteLLM/OpenBao (deploy/compose/), not a mock. For the deeper reference on any of these pieces, see docs/dev-environment.md (the exhaustive version of this doc); for what still needs you personally (Slack/GitHub apps, credentials), see docs/needs-user.md.

Prerequisites

  • direnv allow in the repo root once, so .envrc (gitignored, holds the Aiven AI Gateway key and other API keys) loads automatically on cd. just dev-up fails fast with a clear message if these aren't set yet.
  • Docker + Docker Compose, uv, Go 1.27+ — see docs/dev-environment.md's own Prerequisites section for exact versions and how to check them.
  • just doctor checks all of the above at once (/dev/kvm, the sudoers NOPASSWD allowlist and its drift from vmdctl sudoers' own rendered grant, the pinned firecracker/jailer binaries, direnv, Docker Compose v2, deploy/compose/.env, uv/go/just versions, free /var/tmp space in both absolute GiB and percent-free, and vmdctl host-check for a provisioned host's own kernel/KSM/swap/SMT settings) — one [PASS]/[WARN]/[FAIL] line per check, nonzero exit only on a real FAIL (a workstation failing vmdctl host-check is an expected WARN, not a failure — that check only turns into a hard FAIL when KAPELLE_HOST_ROLE=server says this is meant to be a provisioned host). Pass --json for a machine-readable report instead of the plain-text lines (what the nightly CI workflow uploads alongside the human-readable run). Run it first on a new host, before anything below; verified for real, 15/15 passed on this one (2 with warnings: vmdctl host-check — expected on a plain dev workstation, whose kernel/swap/SMT settings aren't tuned like a provisioned Kapelle host's — and sudoers drift, here because vmdctl itself hasn't been built on this host yet to compute the comparison).
  • direnv exec /home/mbocevski/dev/factory <command> wraps every docker compose call in this doc and in the justfile's recipes — never run docker compose bare against this repo's compose file; a bare invocation loses the .envrc vars several containers need and will force-recreate them, losing real model credentials.

1. Bring up the core stack

just dev-up
just dev-seed

dev-up starts NATS, Postgres, OpenBao, the OTel collector, Jaeger, and LiteLLM (deploy/compose/docker-compose.yaml, no profile flag needed for these — they're the default set). dev-seed creates NATS's JetStream streams/KV buckets and OpenBao's policies/AppRole auth, and checks LiteLLM is actually answering. Both are idempotent — safe to re-run.

2. Run the three services

Schema migrations first (each service owns its own alembic tree; none create_all()s at runtime):

DATABASE_URL="postgresql+asyncpg://<POSTGRES_USER>:<POSTGRES_PASSWORD>@127.0.0.1:5432/<POSTGRES_DB>" \
  uv run --package kapelle-controller alembic -c services/controller/alembic.ini upgrade head
uv run --package kapelle-gateway alembic -c services/gateway/alembic.ini upgrade head
uv run --package kapelle-worker alembic -c services/worker/alembic.ini upgrade head

(same DATABASE_URL for all three — read the actual values out of deploy/compose/.env, gitignored, created from .env.example — see docs/dev-environment.md's "Running the controller natively" for the exact grep-based one-liners the justfile itself uses to build this URL from that file, since teammates/agents aren't allowed to read/edit .env directly by hand.) The a2a and litellm databases these services also need already exist at this point — deploy/compose/postgres/init-databases.sh creates them as part of the Postgres container's own first-boot init, before any migration runs.

Two ways to run the services themselves; pick one.

Native (what you want for actual development — starts in seconds, no rebuild after a code change):

just controller-dev   # backgrounded; logs "controller: gRPC listening on 127.0.0.1:8300"
just gateway-dev       # backgrounded; mounts /a2a/<team>/<role> routes as teams appear

Team memory (design doc §14) searches by keyword only: the Aiven AI Gateway serves no embedding model yet, so KAPELLE_EMBED_ALIAS stays unset -- see docs/dev-environment.md's "Team memory semantic search" section for what would turn semantic search on.

There's no worker-dev recipe (a real deployment always names a fixed team/role per worker process, so there's nothing generic to default it to) — start one by hand, matching controller-dev/gateway-dev's own env-var pattern:

KAPELLE_TEAM=demo KAPELLE_ROLE=coder KAPELLE_EXECUTOR=echo \
NATS_URL="nats://127.0.0.1:4222" \
DATABASE_URL="<same URL as above>" \
A2A_DATABASE_URL="<same URL, database a2a instead of kapelle>" \
  uv run --package kapelle-worker python -m kapelle_worker.a2a.main

KAPELLE_EXECUTOR=echo (task 15's stub, no LLM/sandbox in the loop) is the fast path for proving the wiring end to end, exactly what this doc does below. A real agent needs KAPELLE_EXECUTOR=openhands plus a real or fake sandbox provider and real model credentials — see eval/e2e/run_real_eval.py for a complete, real (not stubbed) example of that whole setup, or docs/dev-environment.md's "Running vmd/credgwd natively" section for the real-Firecracker path.

Containerized (services profile, backlog task 38 — proves the actual container images work, not day-to-day iteration):

just dev-up-services

Builds and starts controller/gateway/worker from the uv workspace (one shared deploy/compose/services.Dockerfile). The compose worker is fixed to KAPELLE_TEAM=smoke KAPELLE_ROLE=coder KAPELLE_EXECUTOR=echo. The gateway's HTTP port publishes to a random host port (so it never collides with a teammate's native gateway-dev on the same host) — look it up with direnv exec /home/mbocevski/dev/factory docker compose -f deploy/compose/docker-compose.yaml port gateway 8100. Everything below in this doc used the native path; swap in the compose gateway's own port and smoke in place of demo to do the same thing against this profile instead.

Controller migration b7d2e4f6a8c1 (child tables keyed by team_id). team_roles, surface_links, team_members, the memory entries, skills and mcp_servers were keyed by the team's name; with reusable names (an archived team's name is free for a new team) they now carry team_id integer NOT NULL, a foreign key to teams.id, and every uniqueness rule is on that id. The team name column stays and is still written. What the upgrade does to existing rows: a row goes to the team of its name that is not archived, else to the archived team of that name with the highest id; a memory entry renamed <name>.archived-<id> goes to team <id>; a row whose name matches no team at all is deleted, with one log line per table naming the count and the names. The downgrade refuses, with a RuntimeError naming the duplicate keys, when two teams now hold rows under one name that the name-based constraints would reject; otherwise it restores the name-based schema. Run the controller's migration before starting a gateway of this version: the gateway checks at start that team_members and surface_links have team_id and exits naming this migration if they do not. Back up the database first if it holds data you care about; the delete of unmatched rows is not undone by the downgrade.

3. Create a team from the template

No CLI tool sends /team create outside a real chat surface yet — the most direct way, and how this doc actually did it, is a raw NATS request matching kapelle_gateway.controller_client.ControllerCommandClient's own wire shape, via the nats-box container deploy/compose/ already has:

direnv exec /home/mbocevski/dev/factory docker compose --env-file deploy/compose/.env \
  -f deploy/compose/docker-compose.yaml run --rm nats-box \
  nats --server nats://nats:4222 req kapelle.controller.command \
  '{"command":"team","args":["create","demo","--template","feature-team","--repo","https://github.com/mbocevski/factory-playground.git"],"team":null,"surface":"eval","user":{"person_id":"1","display_name":"you","role":"admin"}}'
{"ok": true, "text": "Team 'demo' created from template 'feature-team'."}

feature-team and feature-team-nevia set requires_repo: true: every role works in a git checkout of the team's repository, so /team create refuses them without --repo <url> and answers this template needs --repo <url>; usage: /team create <name> --template feature-team --repo <url> [--home slack|mattermost]. A team that was created without a repository before this rule (its roles have no repo_url) is not repaired: a task on it fails at once with team <name> has no repository, so <role> has nothing to work on; create the team with /team create <name> --template <template> --repo <url>. Archive that team and create it again with --repo (the name is free again after archiving).

feature-team.yaml (teams/templates/) is the real template this loads — planner/coder/reviewer roles, model aliases, VM profiles, sleep timeouts, iteration budgets; see the file itself, it's short and commented. The gateway's periodic run_a2a_mount_loop picks up a new team and mounts its /a2a/<team>/<role> routes within ~60s of creation (instant if the gateway process starts after the team already exists).

/team archive//team link//team members add|remove all require the caller to hold admin team membership — TeamService.create() grants created_by that role automatically, in the same transaction that creates the team, so the demo team created above can be archived by that same person_id (verified for real against the dev stack):

direnv exec /home/mbocevski/dev/factory docker compose --env-file deploy/compose/.env \
  -f deploy/compose/docker-compose.yaml run --rm nats-box \
  nats --server nats://nats:4222 req kapelle.controller.command \
  '{"command":"team","args":["archive","demo"],"team":"demo","surface":"eval","user":{"person_id":"1","display_name":"you","role":"admin"}}'
{"ok": true, "text": "Team 'demo' archived."}

/team members <name> (list), /team members add|remove <name> <user_id> [admin|member] (both admin-only) manage who else holds that role on a team.

4. Send a task, watch it complete

A real external caller speaks JSON-RPC 2.0 over HTTP to the gateway's /a2a/<team>/<role>/a2a/jsonrpc route (A2A-Version: 1.0 header required, or a2a-sdk's own version validation rejects the request as protocol 0.3):

curl -sS -X POST http://127.0.0.1:8100/a2a/demo/coder/a2a/jsonrpc \
  -H "Content-Type: application/json" -H "A2A-Version: 1.0" \
  -d '{
    "jsonrpc": "2.0", "id": "1", "method": "SendMessage",
    "params": {"message": {"messageId": "11111111-1111-1111-1111-111111111111",
                            "role": "ROLE_USER",
                            "parts": [{"text": "add a --version flag"}]}}
  }'

With the echo executor this responds immediately, status.state: "TASK_STATE_COMPLETED", one artifact whose text is "echo: add a --version flag". A real agent takes real wall time to work, so "watching" a work item means polling GetTask with the same route and the id SendMessage returned:

curl -sS -X POST http://127.0.0.1:8100/a2a/demo/coder/a2a/jsonrpc \
  -H "Content-Type: application/json" -H "A2A-Version: 1.0" \
  -d '{"jsonrpc": "2.0", "id": "2", "method": "GetTask", "params": {"id": "<task id from above>"}}'

This is the raw wire protocol — the point of this step is proving the plumbing (gateway HTTP → NATS → worker → back), not how a person actually uses this system day to day. A real user watches a work item in a chat room instead: Mattermost is scriptable in this dev environment (just dev-up-mattermost + just mattermost-seed, docs/dev-environment.md), but Slack needs an app only you can create — docs/slack-app.md has the one-time manifest setup and the KAPELLE_SLACK_* env vars to put in .envrc; same story for Linear/Jira (docs/linear-app.md/docs/jira-app.md) and GitHub (PR creation/status, docs/github-app.md — its "Local development" section covers running against a real App without hosting anything publicly: an inactive webhook is enough for token minting/PR creation, gh webhook forward covers live inbound deliveries if you need those too). See docs/needs-user.md's Open section for the exact current list of what's blocked on you specifically.

5. Run the e2e suite and the eval suite

just e2e

Brings up just dev-up itself, then runs every landed step of the demo scenario (docs/demo.md) against the real stack: team creation, a real gateway↔worker HTTP/NATS round trip, worker-kill/reconciler redispatch, a coder asking a clarifying question and pushing a branch, a real planner→coder→reviewer→planner delegation chain (the reviewer leg, test_reviewer_completes_the_loop_and_the_room_sees_every_hop, went green in commit 4755a17 — see docs/demo.md), sleep/restore and credential-hygiene on real Firecracker (needs passwordless sudo and /dev/kvm — see docs/dev-environment.md's Prerequisites; skips itself if you don't have them), and a real Jaeger trace assertion. test_pr_ appears_with_closes_reference stays xfailed — needs a real GitHub App, see docs/needs-user.md. Check the task tracker (backlog task 29) for the current point-in-time pass/fail count before reading too much into any one run — the reviewer leg in particular has an open, filed non-determinism bug (a durable-reply ping-pong that can balloon a run to ~200 LLM turns before settling) tracked separately, not a regression in this doc.

Every run writes eval/report/out/e2e-junit.xml (structured per-test pass/fail) and eval/report/out/e2e-output.log (the full raw output, every worker subprocess's inherited stdout/stderr included, for passing tests too via -rA's PASSES section) regardless of outcome -- backlog task c3e7346b's own real-run finding: this suite starts real subprocesses that stream real, multi-minute output straight to the terminal with no other durable record, so scrolling past the FAILURES section (or piping through your own | tail) loses a failure's root cause for good. Both files are gitignored; check them first before rerunning a failed suite.

The separate test_vertical_slice_microvm.py (a coder actually working inside a real Firecracker guest, not FakeProvider) is confirmed green for real (as of confirm31, 2026-09-13 — see docs/demo.md's own record: hard assertions throughout, no xfail left in the file) — budget several real minutes for it regardless of outcome (a real cold boot + real work through a real credential gateway); a timeout that's too short cuts it off mid-run and leaves a real orphaned firecracker process behind that you'll have to find (ps -eo pid,cmd | grep firecracker) and kill by hand. Task 43's multi-host placement (a registered-and-reachable hosts table row plus a live mTLS Capacity RPC, for a real multi-host deployment) does NOT block a single dev host: ensure_agent's own host-selection (services/controller/src/ kapelle_controller/agents.py's _select_firecracker_host) falls back to the single implicit host (KAPELLE_VMD_ADDR) whenever HostsStore has no rows registered — exactly this test's own setup, and every dev/test/e2e setup that predates task 43.

just eval --offline

Exercises the fixture/report/threshold pipeline against a fake driver — no real stack needed, defaults --model to local-coder so this runs with no flags at all. Correctly exits nonzero: --offline's fixtures are deliberately scripted to fail most of the time, to prove the threshold gate actually gates.

uv run --package kapelle-gateway python eval/e2e/run_real_eval.py [task-name ...]

The real, no-VM path (FakeProvider(use_real_agent_server=True), a real local OpenHands agent-server, real LiteLLM calls, no Firecracker/sudo needed): brings up the shared dev stack, creates a temp team per task, launches a real worker, submits the task, checks the agent's real pushed branch, writes eval/report/out/real_run.json. Defaults to the 5 core tasks with no args, or name one or more of eval/tasks/'s 10 fixtures. See docs/eval-results.md for the full, dated trail of every real run — as of this writing, 9/10 fixtures pass for real (the one miss is a genuine model-capability gap in local-coder's review, not an infra bug).

To try a stronger model on one fixture without editing anything, add --model-alias ROLE=ALIAS (repeatable; env fallback KAPELLE_EVAL_MODEL_ALIAS_<ROLE>), e.g. the review fixture that fails on local-coder:

uv run --package kapelle-gateway python eval/e2e/run_real_eval.py \
    --model-alias coder=frontier review-payments-defects

Note it's coder=, not reviewer= — this harness drives every fixture, review included, through the team's coder role (there's only one worker per run). Refused up front, before any team is created, if LiteLLM doesn't have the alias configured (deploy/compose/litellm/config.yaml's model_list); the alias actually used is recorded per fixture in real_run.json and docs/eval-results.md.

That same real-model, no-VM setup is also reachable from a plain kapelle_worker.a2a.main process, not just this eval harness — see docs/dev-environment.md's "Running a real model without Firecracker" section (KAPELLE_FAKE_PROVIDER_REAL_AGENT_SERVER=1 plus the KAPELLE_FAKE_LLM_*/KAPELLE_MCP_*/repo-checkout-dir env vars) if you want a native controller-dev/gateway-dev/worker trio in section 2 above to answer with a real model completion instead of the echo: stub, without needing /dev/kvm or sudo at all.

just spend

Real LiteLLM cost for a date window (defaults to today), broken down by model, by agent key, and by still-active team — see docs/design.md's "Cost and attribution" (§ Teams) for why per-agent keys attribute correctly even in this no-VM dev setup, and scripts/spend_report.py's own docstring for this report's real limitations (Enterprise-only LiteLLM endpoints it can't use, archived teams' spend only visible in the total).

Run a real team

Sections 1-5 above all run the fake-provider path (FakeProvider, no real Firecracker/credgw) driven by raw NATS requests standing in for a real chat surface. just stack-up (backlog task 6eabe175, scripts/stack.py's own docstring has the full process-by-process wiring) replaces all of that with the real thing: vmd, credgwd (real-guest flags), the controller on the real firecracker provider, the gateway, and a shared-mode worker — so the only steps left, once it's up, are typed in the chat room itself (Mattermost, docs/dev-environment.md).

just stack-up
just stack-status   # confirm all five processes are running

Refuses if anything looks already running (--replace to restart), and checks seven preconditions first (the core compose containers are up, just dev-seed's own JetStream stream exists, the guest kernel and iron-proxy binaries are present, this host's sudoers grant is installed, /dev/kvm is accessible, and a local vmd image manifest exists from a prior just image-build) with a clear [OK]/[FAIL] <name>: <detail> line for each. Once every precondition passes, stack-up itself runs alembic upgrade head for all three services' own migration trees (controller/gateway/worker) — nothing to do by hand — then makes sure that manifest's image is actually published (auto-publishing it if it isn't yet) before starting any process. just stack-up --dry-run prints every process's exact command and environment (secrets redacted) without starting anything, useful for checking the wiring before a real run. Never run just stack-up from a teammate agent session — real image builds and real microVM boots trip the per-session sandbox memory guard (docs/ci.md's own finding); this is a human, or lead-session, call.

Before any of that, stack-up refuses to start at all if a previous run leaked teams or NATS consumers. It dry-runs scripts/archive_stale_teams.py/scripts/nats_prune_consumers.py first; if either finds candidates (an interrupted prior run that never archived its own e2e-*/demo-*/tally-* teams, or left an orphaned per-team NATS consumer behind), it prints their own output plus a one-line refusal ("...would block the shared worker's wildcard consumer -- rerun with --prune to archive them first" for teams, the consumer-equivalent wording for consumers) and exits nonzero — it does NOT clean up on its own. just stack-up --prune runs both tools with --apply first, then continues. This is usually the FIRST thing you hit running stack-up a second time on a shared host: someone's earlier session (including your own, killed mid-run) left something behind.

This pre-flight check does NOT yet dry-run scripts/nats_prune_a2a_ tasks_consumers.py/just prune-follower-consumers (backlog task 4bb2516f) — a leaked gateway-follower-<team> durable consumer on A2A_TASKS costs a stale gateway subscription, not a blocked shared worker consumer the way the two tools above do, so stack-up doesn't refuse to start over it today. Run just prune-follower-consumers by hand (dry-run first) if a growing consumer count on A2A_TASKS suggests one has piled up.

/team create can be typed straight into a channel the bot is in even though that channel isn't linked to any team yet (backlog task f7d19b2a, gap 4: the gateway handles /team ... and /help on any subscribed-but-unlinked channel, specifically to make this chicken-and-egg case work) — as the bot's admin user, today's own real run:

/team create tally-5 --template feature-team --repo https://github.com/mbocevski/factory-playground.git

This creates the real team, provisions its home room, links it, and (per teams/templates/feature-team.yaml) wires up planner/coder/ reviewer roles against real model aliases (frontier/local-coder/ frontier) and real small/medium/small firecracker VM profiles — dispatch goes to the template's own lead: role (planner here), not a hardcoded lead (gap 1). Because the gateway process that handled this /team create itself just mutated the team, it triggers a full refresh immediately afterwards (gap 3): the new team's /a2a/tally-5/<role> routes are mounted and the Mattermost adapter follows the new channel right away, no restart needed (the ~60s periodic-refresh window still applies to a team created some other way — e.g. straight against the controller, see "3. Create a team from the template" above). Today's own run still showed the room's subscription becoming visibly active within about 60s of the create reply — open the team's own channel and confirm it, then send a task message the same way section "4. Send a task, watch it complete" describes.

Real timings from today's own run (local-coder/frontier per the template above, one host, no contention): the planner's own microVM cold-boots in ~10s; a role's microVM that's already been booted once and gone to sleep restores from its own snapshot in well under 1s instead of booting cold again; the coder role (the real implementation work — read, edit, run tests, fix, repeat) took ~5-8 minutes on local-coder; the reviewer role (read the diff, decide, possibly ask a confirm_risky question) took ~3 minutes. A real PR appears on the playground repo (mbocevski/factory-playground; the throwaway repo, mbocevski/throwaway, is reserved for the automated live tests — test_team_create_live_integration.py) once the coder→reviewer→planner pipeline (teams/templates/feature-team.yaml's own pipeline:) settles.

Once the team exists, an admin can inspect and tune it without editing the template file or recreating the team (backlog task 22278ae1):

/team show tally-5
/team refresh tally-5
/team set tally-5 coder max_iterations 60

/team show <name> prints each role's current effective settings (model, vm profile, max_iterations) and the template version (a content hash) the team is on. /team refresh <name> re-reads the template file fresh and re-syncs the team's roles to whatever it currently says, reporting each field it actually changed — the fix for exactly the gap that motivated it: feature-team.yaml's own planner.max_iterations was raised from 15 to 40 after a live run showed a planner running out of budget mid-exploration before it ever delegated, and an already- created team had no way to pick up that change short of archiving and recreating itself. /team set <name> <role> <field> <value> tunes one role's max_iterations/model_alias/vm_profile directly (validated against real LiteLLM aliases / kapelle_worker.sandbox.providers. firecracker.PROFILES) — an explicit override survives a LATER /team refresh, reported back as a "kept override" rather than silently overwritten.

Two more commands worth knowing, both team-stored and both used for real today:

/skills list
/skills accept run-tests
/mcp list
/mcp accept github

/skills//mcp (list|show <name>|propose <name> ...|accept <name>|forget <name>) manage what a team's own agents can reach beyond the platform default: roles/skills/run-tests/SKILL.md (a platform skill — "look for a task runner first: just/Makefile/package.json scripts/CI config before guessing a raw test command") and roles/mcp/github.yaml (GitHub's own hosted MCP server, deliberately default: false — a real regression once meant every coder conversation got its full 41-tool catalog unconditionally, and a small model (local-coder) got distracted exploring with search tools instead of doing the task; the curated tools.allow here is 7 read/comment-only tools: pull_request_read, list_pull_requests, get_file_contents, issue_read, add_issue_comment, add_comment_to_pending_review, pull_request_review_write) are both opt-in this way, or by naming them in a template's own skills:/mcp_servers: list — feature-team.yaml's own coder role already lists mcp_servers: [github].

What a person sees in the room

How the gateway renders a work item in the room (threads, status posts, questions, artifacts, completion) is in docs/design.md's surface table and the Mattermost section of docs/dev-environment.md.

just stack-down stops all five in reverse order (worker, gateway, controller, credgwd, vmd) with SIGTERM then a 15s grace period before SIGKILL, leaving any already-running microVMs to vmd's own shutdown handling. just stack-logs [name] prints one process's log (or every process's, with no name given) from /var/tmp/kapelle-stack/<name>.log. just stack-status also WARNs (not just up/down) if the worker's own log shows repeated "transient JetStream fetch error" — running, but its own NATS consumer likely never bound (see just prune-teams/ scripts/nats_prune_consumers.py, same tool stack-up --prune itself wraps).

6. Host hygiene

A shared dev host accumulates leftovers from killed test runs and other sessions' experiments — three recipes clean specific kinds of mess, all dry-run by default:

just prune-teams

Archives leaked eval-*/e2e-*/demo-*/repro-* teams still marked active (older than 2h by default) through the real controller archive path — a leaked team's own pinned NATS consumer permanently blocks the compose worker's shared-mode wildcard consumer from ever binding, so this isn't just tidiness. Pass --apply to actually archive; needs just dev-up's Postgres/NATS/LiteLLM/OpenBao running.

just prune-workdirs

Removes leaked test-fixture work directories that a killed test process left behind: kapelle-e2e-* under /var/tmp, and the children of /var/tmp/vmd-it-base (jail, jail-<label>, <prefix>-<hex>), where the vmd integration suites work since 3188034e. Directories of the old layout (vmd-it-*, vmd-python-it-*) are still found. Skips anything newer than 6h by default, or still in use by a live process. Pass --apply to actually remove; the base's children go through sudo -n vmd-priv fileop. Never touches /var/tmp/vmd-dev (the persistent vmd-dev recipe's own state) or the base directory itself.

just scan-secrets <path> [path ...]

Audits real files/directories (a log file, a mounted secrets dir, ...) for real credential values leaked into them in plaintext — exits nonzero with redacted hits printed (never a raw value) if anything leaked. See scripts/scan_leaks.py's own docstring for the full flag set (--env, --openbao-addr/--openbao-token/--openbao-agent, ...).

Contributor note

Keep the shared tree buildable at every commit. This is a live, shared working tree — several sessions build and test from it while you edit. New code goes in its own package with its own tests first; add a dependency with uv add/go get and commit the updated uv.lock/ go.mod+go.sum before you write the first import of it; wire an existing call site to the new code only in your final commit. An uncommitted half-edit that breaks uv sync or go build ./vmd/... ./credgw/... blocks everyone else's just lint/just test until it's fixed — see CLAUDE.md's Git workflow section for the full rule and pathspec-commit discipline this shared tree relies on.

CHANGELOG entries: use just changelog <Added|Changed|Deprecated |Removed|Fixed|Security> "<one-sentence bullet>" right after the code commit that lands a user-visible change, rather than hand-editing CHANGELOG.md — it inserts the bullet under [Unreleased] and commits that one-line change immediately (scripts/changelog_add.py), and refuses outright if the file already has someone else's uncommitted bullet, rather than racing to overwrite it. --commit-message "<text>" overrides the default commit message if you want one.