Skip to content

CI

Two workflows, two different jobs:

  • .github/workflows/ci.yml — GitHub-hosted runners, every push/PR. Lint, unit tests, and every Docker-gated integration suite that can run without real KVM/Firecracker or real model credentials (the Firecracker- booting suite and the compose smoke test both self-skip here, same pattern as everywhere else in this repo — see their own module docstrings). No secrets, no self-hosted runner.
  • .github/workflows/nightly-integration.yml — a self-hosted runner, nightly (03:17 UTC) plus workflow_dispatch. Runs just doctor first (continue-on-error, so its own JSON report — just doctor --json, always uploaded as the doctor-report artifact — lands even on a real FAIL; a following step re-raises that failure for the job as a whole right after, before dev-up wastes any time), just dev-up, just test-integration (this time including the real-Firecracker suite, since the runner has KVM), brings up the containerized control plane (just dev-up-services — the compose worker runs shared mode against a real model, see "Why plain just eval now, not eval/e2e/run_real_eval.py" below), and a real OpenHands eval run via plain just eval --database-url ... (non-offline), then uploads eval/report/out as an artifact, scans for leaked credentials (just nightly-scan-secrets, below — same continue-on-error + always-run-upload + re-raise pattern as the doctor step, so the redacted report still gets uploaded even on a real hit), and tears the stack down with just dev-down-services (not dev-down — see that recipe's own comment in the justfile: a bare docker compose down doesn't remove a profile's containers).

Credential-hygiene scan

Right after the real eval run and before teardown (containers still up, so docker compose logs has something to read), just nightly-scan-secrets runs scripts/scan_leaks.py (backlog task 47) against the controller/gateway/worker containers' own collected logs plus any native-process KEEP_WORKDIR=1 work dirs a real microVM run left under /var/tmp (vmd.log, credgwd.log, every iron-proxy.log — a glob matching nothing is silently skipped, so this is harmless on a run that didn't produce any). The values it checks for are discovered, not hardcoded: --openbao-list-agents compose recursively walks every agent actually minted under the compose host (docker-compose.yaml's own KAPELLE_CONTROLLER_HOST/ KAPELLE_OPENBAO_HOST pin) for its real litellm_virtual_key/ mcp_token/disk_encryption_key, since just eval mints a fresh team name every run and the workflow has no way to know it in advance. The scan always writes a report (redacted context only, never a raw value) before failing the job on a real hit, at ${RUNNER_TEMP:-/var/tmp}/kapelle-nightly-scan-secrets/secret-scan-report.txt -- GitHub Actions' own per-job scratch dir when run there, /var/tmp locally, never the repo checkout (it used to write .nightly-logs/ and secret-scan-report.txt straight into the repo root, where nothing ever cleaned them up). nightly-integration.yml's scan-secrets step runs it with continue-on-error: true; an always-run step right after uploads that same report as the secret-scan-report workflow artifact (from ${{ runner.temp }}/kapelle-nightly-scan-secrets/secret-scan-report.txt -- the same path, via GitHub Actions' own expression context rather than the shell's $RUNNER_TEMP); a final step re-raises the failure for the job as a whole if the scan step's outcome was failure -- the same three-step shape doctor-report above uses, so a real hit is reviewable from the artifact without re-running the job.

Registering the self-hosted runner

Nobody has registered one yet — this is on docs/needs-user.md. What it needs, once you do:

  • Labels: self-hosted, linux, kapelle (exactly what nightly-integration.yml's runs-on asks for — a plain self-hosted runner with no labels won't be picked for this job).
  • Host prerequisites: everything just doctor checks — /dev/kvm, the sudoers NOPASSWD allowlist, the pinned firecracker/jailer binaries at /usr/local/lib/kapelle/bin, Docker + Compose v2, uv/go/just, and (if bin/vmdctl is built — just build) vmdctl host-check's KSM/swap/SMT/kernel-version settings, design doc §8's Host bullet. just doctor treats a non-zero host-check exit as only a WARNING unless KAPELLE_HOST_ROLE=server is set in the environment it runs in — this workflow's own just doctor step doesn't set it today, so those settings are informational here, not gating; set it in the workflow's env: if this runner should actually be held to a real provisioned host's tuning. Run just doctor by hand on the machine before registering it; the workflow's own first step just re-checks the same things and fails fast if anything regresses. See docs/dev-environment.md's Prerequisites section and spikes/firecracker/FINDINGS.md for how to get the sudoers entries and static binaries in place the first time.
  • deploy/compose/.env: must already exist in the runner's checkout (copy deploy/compose/env.example, fill it in) — this is a one-time, by-hand step. It's gitignored and, per this repo's own rule, teammates/ agents may not create or edit it; the workflow doesn't try to.
  • Repo secrets — Settings → Secrets and variables → Actions:
  • AIVEN_AI_BASE_URL, AIVEN_AI_API_KEY — the Aiven AI Gateway, the only model provider (2026-09-24): every role alias (open-coder for the coder, frontier for planner/reviewer) resolves to a model there, and just dev-up's own fail-fast check requires both to be non-empty. These are the exact same two variables a developer's own .envrc sets (docs/dev-environment.md's "Pointing LiteLLM at a model") — the workflow passes them through as plain job-level environment variables instead. That works without the runner needing its own .envrc file at all: direnv exec (every compose invocation in this repo goes through it) passes an already-exported environment variable straight through when the target directory has no .envrc of its own — verified empirically, not assumed.

No other secrets: deploy/compose/.env already holds POSTGRES_PASSWORD/LITELLM_MASTER_KEY/OPENBAO_DEV_ROOT_TOKEN, and nothing in this workflow touches GitHub/Slack/Linear/Jira credentials.

Why plain just eval now, not eval/e2e/run_real_eval.py

Previously, just eval --runtime openhands --model <alias> (non-offline) had no way to reach a real sandbox: just dev-up-services's worker ran KAPELLE_EXECUTOR=echo, and the one real-model path anyone had actually run end to end (FakeProvider(use_real_agent_server=True) — no VM, no sudo, a real local OpenHands agent-server subprocess; see docs/eval-results.md's Status section) was only ever constructed directly in Python, not reachable through any kapelle_worker.a2a.main environment variable. eval/e2e/run_real_eval.py was the workaround: it brought up its own copy of the dev stack, stood up an in-process controller and a real local agent-server, and ran every core + safety-suite fixture against it directly.

That gap is closed (backlog task 44/0deba590): just dev-up-services's worker now runs shared mode (KAPELLE_TEAM/KAPELLE_ROLE both unset) with KAPELLE_EXECUTOR=openhands and KAPELLE_FAKE_PROVIDER_REAL_AGENT_SERVER=1 — a real local OpenHands agent-server behind the same containerized control plane, able to answer any team's work, not just a fixed one. So the nightly workflow now brings that stack up with just dev-up-services and runs the eval suite against it with plain just eval --database-url ..., same as a developer would locally (docs/dev-environment.md's "Services" section).

Full-repo health sweep baseline (2026-09-13)

Neither workflow above runs just lint + the complete non-privileged just test (Python workspace + go test -race for vmd/credgw + the compose/scripts suites) as ONE combined pass with the numbers recorded anywhere -- CI splits lint/unit/integration into separate jobs/steps and doesn't total them. After today's ~620 commits, team-lead had one manual sweep run against the live dev stack (just dev-up, no KAPELLE_RUN_PRIVILEGED) to get a real baseline:

  • Lint (just lint, run step by step): ruff check/format clean, golangci-lint clean (0 issues), go build ./vmd/... ./credgw/... clean. ty check failed on 1 real bug (since fixed, b11172b).
  • Test (just test, privileged Firecracker suite excluded per the no-KAPELLE_RUN_PRIVILEGED instruction): 1867 passed / 3 failed / 12 skipped / 1465 deselected across the Python workspace (test-unit + test-integration); go test -race ./vmd/... ./credgw/... all packages passed, 0 failures.
  • Wall time: ~27 minutes end to end (test suites' own summed time ~34 minutes; some integration suites were run concurrently rather than the justfile's documented sequential order, to save wall time).

The 3 failures were routed to their owning teammates rather than fixed in this sweep (see backlog/team chat for the individual repro commands and root causes): a SandboxProvider protocol-conformance gap in services/controller/tests/test_scale_out.py's FakeProvider (fixed, b11172b), a missing enable_approle() self-setup step in services/worker/tests/sandbox/test_openbao_pki.py's own fixture, an uncaught guest-unreachable exception in services/worker/src/kapelle_worker/agent/executor.py, and a deploy/compose/tests/test_services_smoke_integration.py "nats: timeout" failure -- root-caused (not a NATS bug, and not a regression from the NATS-account/operator commits, both ruled out by bisecting to a commit before that work landed) to a real data mistake in this host's own deploy/compose/.env: DSV4_BASE_URL held a literal copy of DSV4_API_KEY's own value, so the isolated compose stack's own litellm container (which, unlike the shared dev stack, never goes through direnv exec and so never gets .envrc's correct value overriding it) sent every real LLM call to a non-URL "endpoint", which surfaced several layers downstream as the worker's task never completing and the controller/gateway timing out waiting for a reply. Fixed by the user; just doctor's new deploy/compose/.env URL/secret hygiene check (792ef12) catches this whole class going forward.

A teammate session can't actually verify this specific test passes. Every attempt from a teammate's own session to run the compose smoke test (docker compose ... --build controller worker, a real image build) gets killed by that session's own per-session sandbox memory guard reacting to buff/cache growth during the build -- confirmed NOT a real host memory shortage (free -h showed comfortable headroom, tens of GB available, at the moment of every kill) and NOT dependent on how quiet the rest of the host is at the time; it reproduced 5 times in a row across multiple confirmed-quiet windows. Verifying this one test needs a session without that guard -- team-lead's own lead session, or the nightly CI runner (.github/workflows/nightly-integration.yml, already outside any teammate sandbox) -- not a teammate session, however quiet the host looks.

Confirmed green as of 2026-09-13 22:54, run from the lead session (no per-session guard to trip): test_composed_control_plane_completes_a_real_work_item, 1 passed in 157.42s, on HEAD with the user's fixed .env. With this, the full sweep above is fully green -- the last of the original 3 failures is now resolved, not just root-caused.

This isn't a CI gate -- there's no automated job enforcing "0 failures" against this baseline -- but it's a useful number to compare the next manual sweep against.