CI¶
Two workflows, two different jobs:
.github/workflows/ci.yml— GitHub-hosted runners, every push/PR. Lint, unit tests, and every Docker-gated integration suite that can run without real KVM/Firecracker or real model credentials (the Firecracker- booting suite and the compose smoke test both self-skip here, same pattern as everywhere else in this repo — see their own module docstrings). No secrets, no self-hosted runner..github/workflows/nightly-integration.yml— a self-hosted runner, nightly (03:17 UTC) plusworkflow_dispatch. Runsjust doctorfirst (continue-on-error, so its own JSON report —just doctor --json, always uploaded as thedoctor-reportartifact — lands even on a real FAIL; a following step re-raises that failure for the job as a whole right after, beforedev-upwastes any time),just dev-up,just test-integration(this time including the real-Firecracker suite, since the runner has KVM), brings up the containerized control plane (just dev-up-services— the compose worker runs shared mode against a real model, see "Why plainjust evalnow, noteval/e2e/run_real_eval.py" below), and a real OpenHands eval run via plainjust eval --database-url ...(non-offline), then uploadseval/report/outas an artifact, scans for leaked credentials (just nightly-scan-secrets, below — samecontinue-on-error+ always-run-upload + re-raise pattern as thedoctorstep, so the redacted report still gets uploaded even on a real hit), and tears the stack down withjust dev-down-services(notdev-down— see that recipe's own comment in thejustfile: a baredocker compose downdoesn't remove a profile's containers).
Credential-hygiene scan¶
Right after the real eval run and before teardown (containers still up,
so docker compose logs has something to read), just
nightly-scan-secrets runs scripts/scan_leaks.py (backlog task 47)
against the controller/gateway/worker containers' own collected logs
plus any native-process KEEP_WORKDIR=1 work dirs a real microVM run
left under /var/tmp (vmd.log, credgwd.log, every iron-proxy.log
— a glob matching nothing is silently skipped, so this is harmless on a
run that didn't produce any). The values it checks for are discovered,
not hardcoded: --openbao-list-agents compose recursively walks every
agent actually minted under the compose host
(docker-compose.yaml's own KAPELLE_CONTROLLER_HOST/
KAPELLE_OPENBAO_HOST pin) for its real litellm_virtual_key/
mcp_token/disk_encryption_key, since just eval mints a fresh team
name every run and the workflow has no way to know it in advance. The
scan always writes a report (redacted context only, never a raw value)
before failing the job on a real hit, at
${RUNNER_TEMP:-/var/tmp}/kapelle-nightly-scan-secrets/secret-scan-report.txt
-- GitHub Actions' own per-job scratch dir when run there, /var/tmp
locally, never the repo checkout (it used to write .nightly-logs/ and
secret-scan-report.txt straight into the repo root, where nothing ever
cleaned them up). nightly-integration.yml's scan-secrets step runs it
with continue-on-error: true; an always-run step right after uploads
that same report as the secret-scan-report workflow artifact (from
${{ runner.temp }}/kapelle-nightly-scan-secrets/secret-scan-report.txt
-- the same path, via GitHub Actions' own expression context rather than
the shell's $RUNNER_TEMP); a final step re-raises the failure for the
job as a whole if the scan step's outcome was failure -- the same
three-step shape doctor-report above uses, so a real hit is reviewable
from the artifact without re-running the job.
Registering the self-hosted runner¶
Nobody has registered one yet — this is on docs/needs-user.md.
What it needs, once you do:
- Labels:
self-hosted,linux,kapelle(exactly whatnightly-integration.yml'sruns-onasks for — a plainself-hostedrunner with no labels won't be picked for this job). - Host prerequisites: everything
just doctorchecks —/dev/kvm, the sudoers NOPASSWD allowlist, the pinned firecracker/jailer binaries at/usr/local/lib/kapelle/bin, Docker + Compose v2,uv/go/just, and (ifbin/vmdctlis built —just build)vmdctl host-check's KSM/swap/SMT/kernel-version settings, design doc §8's Host bullet.just doctortreats a non-zerohost-checkexit as only a WARNING unlessKAPELLE_HOST_ROLE=serveris set in the environment it runs in — this workflow's ownjust doctorstep doesn't set it today, so those settings are informational here, not gating; set it in the workflow'senv:if this runner should actually be held to a real provisioned host's tuning. Runjust doctorby hand on the machine before registering it; the workflow's own first step just re-checks the same things and fails fast if anything regresses. Seedocs/dev-environment.md's Prerequisites section andspikes/firecracker/FINDINGS.mdfor how to get the sudoers entries and static binaries in place the first time. deploy/compose/.env: must already exist in the runner's checkout (copydeploy/compose/env.example, fill it in) — this is a one-time, by-hand step. It's gitignored and, per this repo's own rule, teammates/ agents may not create or edit it; the workflow doesn't try to.- Repo secrets — Settings → Secrets and variables → Actions:
AIVEN_AI_BASE_URL,AIVEN_AI_API_KEY— the Aiven AI Gateway, the only model provider (2026-09-24): every role alias (open-coderfor the coder,frontierfor planner/reviewer) resolves to a model there, andjust dev-up's own fail-fast check requires both to be non-empty. These are the exact same two variables a developer's own.envrcsets (docs/dev-environment.md's "Pointing LiteLLM at a model") — the workflow passes them through as plain job-level environment variables instead. That works without the runner needing its own.envrcfile at all:direnv exec(every compose invocation in this repo goes through it) passes an already-exported environment variable straight through when the target directory has no.envrcof its own — verified empirically, not assumed.
No other secrets: deploy/compose/.env already holds
POSTGRES_PASSWORD/LITELLM_MASTER_KEY/OPENBAO_DEV_ROOT_TOKEN, and
nothing in this workflow touches GitHub/Slack/Linear/Jira credentials.
Why plain just eval now, not eval/e2e/run_real_eval.py¶
Previously, just eval --runtime openhands --model <alias> (non-offline)
had no way to reach a real sandbox: just dev-up-services's worker ran
KAPELLE_EXECUTOR=echo, and the one real-model path anyone had actually
run end to end (FakeProvider(use_real_agent_server=True) — no VM, no
sudo, a real local OpenHands agent-server subprocess; see
docs/eval-results.md's Status section) was only ever constructed
directly in Python, not reachable through any kapelle_worker.a2a.main
environment variable. eval/e2e/run_real_eval.py was the workaround: it
brought up its own copy of the dev stack, stood up an in-process
controller and a real local agent-server, and ran every core +
safety-suite fixture against it directly.
That gap is closed (backlog task 44/0deba590): just dev-up-services's
worker now runs shared mode (KAPELLE_TEAM/KAPELLE_ROLE both
unset) with KAPELLE_EXECUTOR=openhands and
KAPELLE_FAKE_PROVIDER_REAL_AGENT_SERVER=1 — a real local OpenHands
agent-server behind the same containerized control plane, able to answer
any team's work, not just a fixed one. So the nightly workflow now
brings that stack up with just dev-up-services and runs the eval
suite against it with plain just eval --database-url ..., same as a
developer would locally (docs/dev-environment.md's "Services"
section).
Full-repo health sweep baseline (2026-09-13)¶
Neither workflow above runs just lint + the complete non-privileged
just test (Python workspace + go test -race for vmd/credgw + the
compose/scripts suites) as ONE combined pass with the numbers recorded
anywhere -- CI splits lint/unit/integration into separate jobs/steps and
doesn't total them. After today's ~620 commits, team-lead had one
manual sweep run against the live dev stack (just dev-up, no
KAPELLE_RUN_PRIVILEGED) to get a real baseline:
- Lint (
just lint, run step by step): ruff check/format clean, golangci-lint clean (0 issues),go build ./vmd/... ./credgw/...clean.ty checkfailed on 1 real bug (since fixed,b11172b). - Test (
just test, privileged Firecracker suite excluded per the no-KAPELLE_RUN_PRIVILEGEDinstruction): 1867 passed / 3 failed / 12 skipped / 1465 deselected across the Python workspace (test-unit + test-integration);go test -race ./vmd/... ./credgw/...all packages passed, 0 failures. - Wall time: ~27 minutes end to end (test suites' own summed time ~34 minutes; some integration suites were run concurrently rather than the justfile's documented sequential order, to save wall time).
The 3 failures were routed to their owning teammates rather than fixed
in this sweep (see backlog/team chat for the individual repro commands
and root causes): a SandboxProvider protocol-conformance gap in
services/controller/tests/test_scale_out.py's FakeProvider (fixed,
b11172b), a missing enable_approle() self-setup step in
services/worker/tests/sandbox/test_openbao_pki.py's own fixture, an
uncaught guest-unreachable exception in
services/worker/src/kapelle_worker/agent/executor.py, and a
deploy/compose/tests/test_services_smoke_integration.py "nats:
timeout" failure -- root-caused (not a NATS bug, and not a regression
from the NATS-account/operator commits, both ruled out by bisecting to
a commit before that work landed) to a real data mistake in this host's
own deploy/compose/.env: DSV4_BASE_URL held a literal copy of
DSV4_API_KEY's own value, so the isolated compose stack's own litellm
container (which, unlike the shared dev stack, never goes through
direnv exec and so never gets .envrc's correct value overriding it)
sent every real LLM call to a non-URL "endpoint", which surfaced several
layers downstream as the worker's task never completing and the
controller/gateway timing out waiting for a reply. Fixed by the user;
just doctor's new deploy/compose/.env URL/secret hygiene check
(792ef12) catches this whole class going forward.
A teammate session can't actually verify this specific test passes.
Every attempt from a teammate's own session to run the compose smoke
test (docker compose ... --build controller worker, a real image
build) gets killed by that session's own per-session sandbox memory
guard reacting to buff/cache growth during the build -- confirmed NOT a
real host memory shortage (free -h showed comfortable headroom, tens
of GB available, at the moment of every kill) and NOT dependent on how
quiet the rest of the host is at the time; it reproduced 5 times in a
row across multiple confirmed-quiet windows. Verifying this one test
needs a session without that guard -- team-lead's own lead session, or
the nightly CI runner (.github/workflows/nightly-integration.yml,
already outside any teammate sandbox) -- not a teammate session, however
quiet the host looks.
Confirmed green as of 2026-09-13 22:54, run from the lead session
(no per-session guard to trip):
test_composed_control_plane_completes_a_real_work_item, 1 passed in
157.42s, on HEAD with the user's fixed .env. With this, the full
sweep above is fully green -- the last of the original 3 failures is
now resolved, not just root-caused.
This isn't a CI gate -- there's no automated job enforcing "0 failures" against this baseline -- but it's a useful number to compare the next manual sweep against.