Skip to content

Changelog

All notable changes to this project are documented in this file.

The format is based on Keep a Changelog 1.1.0. This project has not made a first release yet — everything below is still [Unreleased]; entries move under a version heading at the first tag.

[Unreleased]

Added

  • Contracts (design §6, §11): the A2A-over-NATS wire contract and the shared factory-contracts Python package (also generates vmd's Python gRPC stubs); the standard A2A JSON-RPC HTTP binding for external callers, alongside the NATS transport.
  • Controller (design §12): team lifecycle (create/archive/link), templates, the OpenBao secrets layout, GitHub App token minting, LiteLLM key lifecycle, per-team budget enforcement and notices, EnsureAgent/ GetTeamRole RPCs for stateless workers, replica scale-out/scale-in policy, and team memory (pgvector-backed, with a /memory command).
  • Gateway (design §13): surface adapters for Zulip, Slack, GitHub Issues, Linear and Jira (Forge remote agent plus OAuth 3LO fallback); a user directory and permissions matrix; a delegation follower so a room sees a whole delegated chain, not just the first task; a /budget chat command; the gateway process mounting /a2a/<team>/<role> JSON-RPC routes on itself, refreshed on an interval as teams/roles appear.
  • Worker (design §7, §11): the A2A-over-NATS server with task leases and a reconciler; an OpenHands-driven AgentExecutor; a sandbox provider interface with a FakeProvider and a real Firecracker-backed provider; an MCP tools server with durable-reply auth; team-memory run-start injection; delegated-branch handoff between roles (push/ fetch/checkout, not a diff paste); confirm_risky INPUT_REQUIRED pauses backed by a real security analyzer, with an approve/deny reply flow and a per-role timeout that cancels a stalled confirmation/ question pause instead of leaving it stuck forever.
  • Worker shared mode (design §5, "workers are stateless and shared"): a worker process can now serve every team's traffic (leaving FACTORY_TEAM/FACTORY_ROLE unset) via one wildcard NATS consumer, instead of being pinned to a single (team, role) — pinned mode remains the default/explicit option.
  • vmd (design §8): the host daemon — a gRPC lifecycle API (Ensure/ Wake/Sleep/Destroy/Capacity/List), jailer integration, scale-to-zero via encrypted snapshot sleep/restore, per-agent data-disk encryption with RustFS backup, snapshot compression, a Unix-socket gRPC listener with SO_PEERCRED auth, and a TCP+mTLS listener sourced from OpenBao PKI for remote hosts; vmdctl (capacity/list/drain/net-diag/destroy/host-check/ sudoers).
  • credgw (design §9): the credential gateway — per-agent iron-proxy instances, OpenBao-backed credential minting scoped by team/role, deterministic git push-ref scoping, OTEL audit export and denial notices.
  • Agent VM image (design §8): the OCI-to-ext4 image pipeline and the stdlib-only Python guest agent (readiness, disk-usage/cleanup, MMDS checks).
  • Dev/prod deploy (design §16): the Docker Compose dev stack (NATS, Postgres, OpenBao, LiteLLM, OTel/Jaeger, optional RustFS/Zulip profiles); production host provisioning (deploy/host/provision.sh); a services compose profile running the containerized controller/ gateway/worker built from the uv workspace (just dev-up-services); Prometheus + Grafana observability.
  • Roles & templates (design §12): planner/coder/reviewer role prompts and the feature-team team template.
  • Eval suite (design §17): the end-to-end vertical-slice harness and demo scenario (just e2e/just demo) driving real team creation, a gateway↔worker round trip, worker-kill/reconciler redispatch, sleep/ restore and credential hygiene on real Firecracker, and a real Jaeger trace assertion; the offline evaluation runner (just eval) with its fixtures/report/threshold pipeline.
  • Scaffold: just doctor, a host-prerequisite checker (KVM, sudoers NOPASSWD allowlist, direnv, Docker Compose v2, tool versions vs. versions.yaml, free disk).
  • Deterministic pipeline hops (design §7/§12): a team template can declare pipeline: {role: next_role} (feature-team.yaml's real example: coder: reviewer, reviewer: planner); when a role's run completes with a pushed branch, the worker delegates to that next role itself, regardless of whether the model called send_task — closing a real gap (gateway's rerun on 0f3f736) where a coder pushed a branch but the model chose not to hand it to the reviewer, silently ending the pipeline early. When the model's own final message is empty (an empty finish call), the hop and the room now see a synthesized summary (pushed branch, HEAD commit, diffstat) instead of a generic "the agent finished the task" placeholder; coder/reviewer's role prompts also now require a substantive closing message, the reviewer's an explicit APPROVED/CHANGES REQUESTED verdict.
  • Per-role model-alias override for run_real_eval.py (--model-alias ROLE=ALIAS, repeatable, env fallback FACTORY_EVAL_MODEL_ALIAS_), validated against LiteLLM's /v1/model/info before any team is created
  • A reusable secret-leak scanner (factory_contracts.leak_scan) and CLI: audits real files/directories for real credential values leaked into them in plaintext, given real minted values from OpenBao and/or a .env file -- promoted out of task 47's Docker-gated dynamic scan so the identical check can run against the native credgwd/iron-proxy/vmd/worker logs that compose profile doesn't cover.
  • just doctor now supports --json output (one object per check: name, status, detail, hint), and the nightly self-hosted workflow runs it first, uploading the report as a workflow artifact.
  • Per-team runtime-configurable daily budget: /budget set (team admins only) updates the team's budget and the corresponding LiteLLM team's max_budget in one call; /budget now also shows the reset time
  • The vertical-slice PR step (test_pr_appears_with_closes_reference) is real now instead of an unconditional stub: a full REST round trip against a throwaway repo -- issue, label, real branch/commit/PR with a Closes #N body, and a merge that proves GitHub itself treats the PR as closing the issue.
  • A person accepting, editing, or forgetting a team-memory entry via /memory now posts a notice in the team's home room, the same way a budget or credgw notice does.
  • scripts/scan_leaks.py gains --openbao-list-agents HOST, recursively discovering every agent minted under a host so a scan does not need every team name known in advance
  • The per-(context_id, role) task cap (the reply-ping-pong safety cap) is now a per-role team-template field (task_cap_per_context_role, default 8) resolved via GetTeamRole, instead of shipping only as a worker-wide FACTORY_TASK_CAP_PER_CONTEXT_ROLE env-var default -- one shared-mode worker process can now enforce a different limit for each team it serves.
  • just github-webhook-forward forwards real GitHub webhook deliveries to a locally running gateway, and just gateway-dev/docs/github-app.md document the full loop -- the one GitHub path the live tests' synthetic payloads couldn't cover.
  • A team configured for dedicated isolation now gets a warning and a one-time notice in its home room when no pinned worker process is actually running for it, instead of silently looking idle -- see docs/team-isolation.md for how to start one.
  • New command: /team model lets a team admin pick a role's model at runtime (validated against LiteLLM's own known aliases), without editing a template and recreating the team. /team roles now shows each role's current alias.
  • Runtime comparison for the eval suite: --runtime pydantic_ai/deep_agents alongside openhands, same sandbox/tool-surface/LiteLLM alias, results recorded per fixture in real_run.json
  • A ReassignAgentHost controller RPC lets an operator move a sleeping agent's recorded host without re-running its one-time placement, backing vmdctl migrate: validates the agent is asleep/deep sleep (never running) and the target host is registered and reachable before updating the record.
  • credgwd mints real GitHub installation tokens for the coder role's github-api destination via the controller's MintGitHubToken RPC (-controller-addr/-github-repo), scoped to pulls/issues only, with the minted token redacted from iron-proxy's own logs
  • vmdctl migrate moves a DEEP_SLEEP agent to another host: backs up and releases its data disk on the source vmd (Destroy), then points the controller's record at the target host (new ReassignAgentHost RPC) so the agent's next Ensure/Wake boots there and restores the backup automatically.
  • Model aliases (frontier, frontier-large, frontier-xlarge, fast, open-coder) are now configurable via deploy/compose/litellm/aliases.yaml, applied with just litellm-sync-aliases; feature-team.yaml's planner role now defaults to frontier instead of a placeholder.
  • gateway now exposes a Prometheus /metrics endpoint (FACTORY_METRICS_HOST/FACTORY_METRICS_PORT), following the worker's own metrics server: inbound events, outbound posts, token mints, webhook signature failures, and A2A task submissions per surface.
  • the coder role opens its own pull request from inside its microVM, through credgw's github-api destination with its own minted installation token, right after a successful push -- Closes #N when the work item is a GitHub issue
  • controller MintGitHubToken now resolves and returns the requesting team's own repo (parsed from Team.repo_url) alongside the minted token -- informational for now, lets credgw stop relying on a static per-host repo flag.
  • Teams with isolation: dedicated now get a real NATS account (operator/account/user JWT auth) provisioned at creation and revoked on archive, instead of relying only on the shared agent-scoped MCP token for isolation -- see contracts/a2a-nats.md's Dedicated NATS accounts section. Requires an operator-mode NATS deployment and controller OpenBao access; creating a dedicated team without OpenBao configured now fails fast instead of silently degrading isolation.
  • just doctor now warns about leaked test databases on the shared dev Postgres (factory_controller_test_, factory_gateway_test_, factory_vmd_it_migrate_*); just prune-test-databases [--apply] drops them.
  • Teams with isolation: dedicated can now provision a real NATS account on a separate operator-mode NATS server (profile nats-operator, just dev-up-nats-operator/dev-down-nats-operator)
  • just doctor now warns when deploy/compose/.env has a URL-shaped variable (*_BASE_URL) that isn't a real http(s) URL, or that matches a secret-shaped variable's value (the real DSV4_BASE_URL/DSV4_API_KEY mixup found 2026-09-13).
  • Teams with isolation: dedicated get real worker-side NATS connectivity: a pinned worker connects to its team's own operator-mode NATS account instead of the shared deployment, and archiving a dedicated team now revokes it server-side too
  • Worker: the Pydantic AI and Deep Agents comparison runtimes now attach the shared Factory MCP server (send_task, ask_requester, memory tools), so they can ask a clarifying question mid-run instead of only ever passing or failing at the end.
  • the agent user inside the guest microVM now has full passwordless sudo (root inside the throwaway VM is the sandbox posture; the agent-server itself still runs unprivileged)
  • coder/reviewer role prompts say the sandbox is theirs -- sudo apt-get install, pip/uv, npm for whatever the task needs, discarded after the run
  • Every role's credgw policy now allows pass-through egress to package registries (apt, pip, npm) with no credential injection, so an agent can install what a task needs
  • vmd attaches each agent's rootfs drive writable, with a private per-agent copy that persists across sleep/wake (fresh again after a deep-sleep restore or migration)
  • vmd-priv: a root-owned privilege helper that replaces the sudoers scheme in production (execPrivileged prefers it whenever its socket exists), for hosts whose sudo is sudo-rs and can't parse the wildcard-shaped sudoers grants this used to depend on
  • just stack-up|stack-down|stack-status|stack-logs: the full native real-team stack (vmd, credgwd, controller, gateway, worker) in one command, with pidfiles/logs and precondition checks, so the remaining steps to run a real team are typing /team create in Zulip and a task message.
  • LiteLLM embed alias (qwen3-embedding-0.6B on the dev endpoint) for team-memory semantic search; enable it in the controller with FACTORY_EMBED_ALIAS=embed.
  • team memory semantic search live via LiteLLM's embed alias (qwen3-embedding-0.6B, 1024 dims) -- keyword+vector hybrid search, a backfill script for pre-existing entries
  • vmd refuses to create/encrypt a new data disk or seed a new agent's rootfs copy when disks_dir/snapshot_dir/jail_base_dir would drop under a configured free-space headroom (default 20 GiB / 10%, vmd.yaml's min_free_disk_bytes/min_free_disk_percent), reporting RESOURCE_EXHAUSTED with the exact path and numbers instead of a bare mkfs ENOSPC; vmdctl prune-agents (dry-run by default, --apply to act) finds and removes agent records whose data disk is missing or whose team is archived; just doctor warns at the same threshold
  • External MCP servers (backlog task d6b93b0c step 1): any role's credgw policy can declare a proxy-mode destination for an external MCP server (Linear, Context7, GitHub's own MCP server, ...), with a per-team OpenBao-backed secret convention (inject_secret: mcp:) for servers whose credential isn't already available another way; a role's team template, repository (.factory/mcp/.yaml on the default branch), or a skill then references one by name (merged across platform/template/repository/skills-bundled sources) -- credentials never enter the VM, and a declaration whose destination isn't in the role's own credgw policy is dropped rather than attached.
  • agents can use skills (AgentSkills-format SKILL.md directories) from the platform and team template sources, merged and resolved per role; repository and team sources land in follow-up steps
  • the OpenHands runtime loads and runs agent skills (design doc §15): resolved skills are staged into the guest workspace so scripts/references/assets are reachable under the credential gateway
  • the Pydantic AI and Deep Agents runtimes render the same skill set as an "Available skills" system-prompt section, with triggered skills injected inline
  • a run's activity feed and outcome memory entry record which skills were actually injected (name@version)
  • Team skills: SKILL.md proposals via the worker's skill_propose MCP tool, curated with /skills list|show|accept|forget on any surface (design doc §15).
  • New platform skill roles/skills/run-tests/: how to find and run a repository's test suite and report failures clearly.
  • External MCP servers are now attached to real OpenHands conversations (backlog task d6b93b0c step 3): a team template's mcp_servers: list resolves through the controller and merges with platform/repository/skills-bundled declarations, injected next to the worker's own built-in delegation/memory server. tools.allow/deny is validated but not yet enforced (logged as a WARN when declared).
  • External MCP servers now attach to the Pydantic AI and Deep Agents comparison runtimes too (backlog task d6b93b0c), not just OpenHands -- HTTP-transport declarations only; stdio is not supported for these two runtimes.
  • The Pydantic AI and Deep Agents runtimes now record which external MCP servers were attached to a completed run as an mcp_servers artifact (backlog task d6b93b0c), matching the existing skills attribution; the OpenHands runtime's own equivalent is next.
  • team-stored external MCP servers: /mcp list|show|accept|forget commands and the mcp_propose MCP tool, with per-run attribution on the activity feed and outcome entries
  • Refuses to call the model when a resumed conversation's history ends with an assistant message (some models reject this outright) -- fails loud and clear instead of a raw provider 400.
  • coder role gains access to GitHub's hosted MCP server (issues/PRs/search) through credgw, reusing the existing GitHub App installation token
  • A run that fails because the model rejected the request (e.g. a Bedrock assistant-message-prefill rejection) logs the last 8 events' shape (role/kind/tool-call presence/content length, never content) at WARN, so the next occurrence is diagnosable from the worker's own log.
  • /team refresh <name> re-syncs an existing team's roles to its current template (an admin no longer has to archive and recreate a team just to pick up a template edit like a raised max_iterations), reporting each changed field; /team set <name> <role> <field> <value> tunes one role's max_iterations/model_alias/vm_profile directly, with validation, without touching the shared template file. An explicit /team set//team model override survives a later /team refresh instead of being silently overwritten. /team show now prints each role's current effective values and the template version the team is on.
  • vmd logs an audit line (RPC name, target agent id, and the real caller identity -- SO_PEERCRED pid/uid for a Unix-socket caller, certificate CN for an mTLS one) on every Ensure/Wake/Sleep/Destroy call, written before the handler runs so a caller is on record even if the call itself never completes
  • vmd gets a new Stop RPC: kills a running guest's Firecracker process and releases its network/gateway resources, but -- unlike Destroy -- keeps the persisted record (marked ERROR) and the data disk/rootfs copy untouched, so a later Wake cold-boots from the exact same disks. For a caller that can't tell a genuinely dead guest apart from one that's merely unreachable or stalled under host load
  • worker,gateway: pushed commits now carry a Co-authored-by trailer naming the person who started the work item, when their surface identity is known
  • per-server MCP tool filtering (tools.allow/deny) is now enforced, not just declared
  • credgw: a "remote" instance mode (StartInstance(remote: true)) so Nevia agents (no tap, open egress) reach their own dedicated iron-proxy instance through one shared public TLS listener, authenticated by a per-agent client token instead of network identity
  • Nevia golden image pipeline (images/nevia/build.sh, factory-nevia-bootstrap, test-boot.sh, just nevia-image/nevia-image-test): builds and boot-tests the checkpoint the Nevia SandboxProvider restores per agent
  • Nevia sandbox provider (factory_worker.sandbox.providers.nevia): wakes/sleeps/destroys agent sandboxes against Aiven's Nevia computers API, reaching the agent-server through an exec-stream relay (HTTP and WebSocket).
  • Nevia agents now get a real, per-agent credgw client token wired end to end: the worker calls credgwd's remote StartInstance on ensure() and feeds the resulting credential into the agent's own bootstrap env (FACTORY_GATEWAY_PROXY_URL/LLM_URL/MCP_URL/CA_PEM/CLIENT_TOKEN), replacing the old shared-across-every-agent static fallback
  • Nevia SandboxProvider destroy() can now export an agent's /data/workspace to S3-compatible object storage and delete the computer for real, instead of sleeping and renaming it forever (factory_worker.sandbox.providers.nevia_backup); images/nevia's golden image now includes zstd
  • Team templates can pick a per-role sandbox provider (firecracker or Nevia); the controller and worker resolve each agent's own provider instead of one deployment-wide default.
  • NeviaProvider.destroy() now really deletes the computer (exporting /data/workspace to S3 first) instead of sleeping and archiving it forever when a backup bucket is configured; ensure() restores that backup automatically
  • Team flow: a template's rework_rounds (default 2, /team set <team> rework_rounds <n>) caps the review-and-fix cycles per work item — past it the lead reports the state to the person instead of re-delegating — and the lead's final message always names the branch and PR (the platform appends them when the lead omits them).
  • A delegate's question is answered by the platform: the asking role's requester turn (a lead, a reviewer) ends with an answer that resumes the paused task, at most three questions per task, a note on failure.
  • The Nevia provider can take its bearer from a command (FACTORY_NEVIA_TOKEN_COMMAND, for example nevia auth access-token) that it runs again before the 15-minute token expires and after a 401, instead of one static token.
  • Nevia agents reach the credential gateway through a reverse relay over the worker's exec stream (computers can neither reach credgwd nor be reached): the agent's proxy is a loopback listener in the computer, renewed make-before-break before the stream URL expires; FACTORY_NEVIA_CREDGW_MODE=public keeps the earlier public-listener mode. The golden image needs a rebuild to carry the relay script.
  • A reverse proxy profile (just dev-up-edge) serves Zulip and the webhook endpoint under their own public names behind one address, with Let's Encrypt certificates
  • Gateway: an optional webhook-only listener (FACTORY_GATEWAY_WEBHOOK_LISTEN, e.g. 0.0.0.0:8101) serving just the webhook routes and /webhooks/healthz, so a public reverse proxy can reach webhooks without reaching the A2A mount on the loopback port.
  • Mattermost as a deployable chat server: just dev-up-mattermost (Mattermost 11.11.1, own PostgreSQL, sign-up closed), an optional public name for it on the edge proxy (EDGE_CHAT_HOST, Let's Encrypt), and just mattermost-seed for the team, administrator, bot and /factory slash command.
  • Mattermost as a chat surface next to Zulip and Slack: a team's room is a channel, every top-level post is a work item and its thread carries the agents' status, questions, artifacts and an in-place activity summary; commands run as /factory <command> (also inside a thread) or as @factory <command> and direct messages to the bot; /team create --home mattermost makes the channel; missed posts are caught up after a reconnect or a gateway restart. Set FACTORY_MATTERMOST_URL and FACTORY_MATTERMOST_BOT_TOKEN to turn it on; FACTORY_DEFAULT_HOME_SURFACE picks the room /team create makes without --home (default zulip). just doctor checks a configured Mattermost.
  • just doctor fails and names the teams whose agents have no document in the secret store (after the dev OpenBao lost its data, for example); it compares the agents of active teams with the store by name only.
  • The dev OpenBao can keep its data across restarts: set OPENBAO_STATIC_UNSEAL_KEY in deploy/compose/.env and just runs it with raft storage on the openbao-data volume and auto-unseal (docs/dev-environment.md); just dev-up initializes a new store (just dev-openbao-init). Until the key is set, the in-memory dev server runs as before.
  • A wake that fails because an agent's secret document is missing from the store now fails the task with one clear sentence naming the team, role and cause, instead of a raw gRPC deadline error; it is never retried or re-provisioned automatically.
  • deploy/nevia/start.sh: a native (no containers) start script that brings up the whole Factory control plane -- OpenBao, NATS, credgwd, LiteLLM, controller, gateway, worker -- inside one Nevia computer against a managed Postgres, proven live in stage 3b; deploy/nevia/nats_seed.py seeds the A2A JetStream streams and the a2a-cards KV bucket without the nats CLI.
  • Database connection pool size and max overflow are now configurable per service (FACTORY_DB_POOL_SIZE/FACTORY_DB_MAX_OVERFLOW on the gateway, controller, and worker), defaulting to SQLAlchemy's own defaults so nothing changes unless set.
  • An INFO log line naming the skill and the directory it landed in confirms when skill staging into a guest workspace succeeds.
  • Operator-run command (scripts/convert_memory_embeddings_to_vector.py) to convert entries.embedding from text to vector(1024) once pgvector becomes available on a database that migrated without it (backlog task 84811aac).
  • just prune-follower-consumers to delete stale gateway-follower- durable consumers on A2A_TASKS (dry-run by default; --apply deletes)
  • /budget is now a registered Slack slash command, alongside /team and /ask (/status and /cancel stay thread replies on Slack, a platform limit, not a gap)
  • just mattermost-role-bots mints a Mattermost bot account per team-template role, in place of the single fallback factory bot every role posts through until it has one; tokens go straight to OpenBao, never a file, a command line, a child process's environment or a log line.
  • Mattermost: a team-template role with its own bot (just mattermost-role-bots) now posts its own status, questions and artifacts as that account instead of factory; a role without one, or on any OpenBao failure, still posts through factory exactly as before.
  • Gateway: active reconciliation asks the worker directly for a stale open task's real state (backlog task 4f70cee7), closing a work item whose terminal event was lost to a gateway restart instead of leaving it stuck open forever.
  • Mattermost: a role's bot that has joined the team (just mattermost-role-bots) but not yet a given room is added to that room by the gateway itself, with factory's own token, the first time it is asked to post there -- no controller change, no restart; refused, it tries again at most once every 10 minutes, posting through factory with the role named in the text meanwhile.
  • an archived team's own name becomes free for a new team to take (/team create no longer refuses it)
  • just install-vmd-priv: builds the root-owned privilege-helper binary and installs it at /usr/local/lib/factory/bin/vmd-priv, printing the sudo install command it runs.
  • scripts/nevia_cp.py and deploy/nevia/start.sh v2 bring up the control plane in a Nevia computer with real credentials (secrets over stdin into 0600 files, OpenBao unsealed from stdin, AppRole for credgwd and the controller).
  • A crashed worker no longer leaves its agent-server calling the model: the sandbox pauses a running conversation when no worker connection was seen for five minutes (needs the new agent image and Nevia checkpoint).
  • Agent guests can export OpenTelemetry traces through credgw: every role policy has an otel destination (POST /v1/traces to 127.0.0.5), and FACTORY_AGENT_OTEL_ENDPOINT now reaches Firecracker and Nevia agents with the http/protobuf protocol set. Off by default; the spans carry LLM prompts, completions and tool inputs and outputs (see docs/dev-environment.md).
  • just mattermost-role-bots gives each role bot a display name (display_name in roles/.md, else the role name with a capital) and a description; bots created before keep the display name Factory until the recipe runs again.
  • Role bots get a picture: roles/avatars/.png is uploaded by just mattermost-role-bots (placeholders for planner, coder and reviewer are shipped; replace them with real art).
  • Nevia control plane: LiteLLM can run on its own computer (nevia_cp.py --role llm|core, create --port 443:4001, llm-key; start.sh roles) behind the computer's ingress URL, and credgwd's -llm-tls hands the agents an https LLM URL for it; frees about 750 MB on the control-plane computer
  • Console API (backlog 1a42615c, K3): the controller serves an HTTP+JSON API on 127.0.0.1:8780 (KAPELLE_API_LISTEN) with session login (argon2id passwords, CSRF header, throttled logins), just console-admin to create the first account, and read routes for teams, runs, users, agents, team templates, environments, images, tools, destinations, secret references (never values), MCP servers and integrations; its OpenAPI document is served and checked in at console/api/openapi.yaml.
  • Console write API and importer (backlog cf3caec1, K5): agents, team templates and teams can be created, changed (with versions, If-Match and an audit log) and archived over the Console API; python -m kapelle_controller.console import loads roles/, teams/templates/ and credgw policies into the store (dry run by default, idempotent, conflicts reported), just console-export writes the store back as repository files, and /kapelle team create reads templates from the store first and the files second.
  • The Console web app is served by the edge profile's console service on EDGE_CONSOLE_HOST, with /api/v1 routed to the controller on the same name
  • Console: a team's and a role's memory settings (inject, max_entries, write_outcomes) can be set with PATCH /teams/{name}; the worker follows them
  • Console API: integrations (Mattermost, Slack, GitHub, Linear, Jira) have settings and secrets stored in the controller database and OpenBao, a test-connection action (POST /api/v1/integrations/{kind}/test, probed by the gateway over NATS, GitHub by the controller) and per-team links (PUT /api/v1/teams/{name}/links, GitHub verified against the App); the gateway reads the store at start with the environment as fallback
  • Console API: each integration reports status.applied (current, pending or unknown): the controller announces a change to the gateway over NATS and asks the gateway what it loaded, so a change shows as "pending restart" until the gateway restarts
  • Console API: GET /api/v1/integrations/mattermost/role-bots lists the Mattermost role bots per role of an active team (ok, missing, token unusable, unused), with the display name and picture each shows and where they differ from the agent definition
  • Console API: the platform skill catalogue (GET/PUT/DELETE /skills, text-only, admin writes) and a read-only list of a team's own skills
  • Console API: a team's own MCP servers can be added, replaced, accepted (from an agent's proposal) and removed (POST/PUT/DELETE /api/v1/mcp-servers), validated with the worker's declaration model, and GET /api/v1/mcp-servers reports the credential state of each server's destination; attaching a platform MCP server to an agent version now reaches every team running that agent
  • Console API: GET /api/v1/onboarding and POST /api/v1/onboarding/token give a person the prompt and the AGENTS.md snippet that let their own coding agent manage Kapelle (generated from the routes served, with ten playbooks, token embedded once); the contract is also served as /api/v1/openapi.yaml, GET /api/v1/me says who the credentials are, and python -m kapelle_controller.console onboarding prints the same text
  • KAPELLE_DOCS_BASE_URL (a repository's docs/ URL) makes the onboarding prompt link console-api.md and design.md as URLs
  • The Console is a surface of the gateway: the gateway records every post of every surface, and the person's own words, as console_posts rows, and takes a task, a reply, a cancel or a status from the controller over kapelle.gateway.console.inbound (a team with no chat integration can be given work); the Console API routes follow
  • Console: give a team a task, answer a run's open question and cancel a run from the Console API (POST /teams/{name}/tasks, /runs/{id}/replies, /runs/{id}/cancel), read a run's conversation (GET /runs/{id}/posts); runs show their home and open question.
  • The published documentation: every page as html and as raw Markdown beside it (Accept: text/markdown on a page URL answers with the .md), llms.txt, llms-full.txt, search and a version selector, served under the Console's /docs and, with EDGE_DOCS_HOST, on a host of its own. just docs-build, docs-check, docs-image.
  • The Console can give a Console-homed team a task, answer a run's open question, send a follow-up to a finished run and cancel an open run; a team or run that lives in a chat room says so and points to the room.
  • A changed integration can apply to its own surface without restarting the gateway: with KAPELLE_GATEWAY_LIVE_RESTART on (off by default for one release), the gateway restarts only that surface about three seconds after the change, keeps the old adapter when the new one cannot be built, and the Console's "pending restart" clears; the loaded-state reply gains restarted_at.

Changed

  • The agents' OpenHands runtime moves from 1.49.1 to 1.53.0, with fastmcp 3.4.8 and mcp 1.30 (OpenHands 1.52 and newer require fastmcp below 4), so the guest image has to be rebuilt; the exception recorded in versions.yaml for the older pin is gone.
  • vmd's network model dropped per-VM network namespaces: since VMs are never cloned, each gets a tap in the host namespace with a unique /30 and nftables rules keyed by tap interface instead (design §8).
  • Replica provisioning moved from the worker to the controller's single-flight EnsureAgent RPC — secrets are never minted worker-side (design §9).
  • Worker role config (model alias, iteration budget, tools, timeouts) now resolves from the controller's GetTeamRole RPC per (team, role) instead of FACTORY_ROLE_* env vars, which remain only a dev fallback.
  • vmd's default gRPC listener is now a Unix socket (SO_PEERCRED- authenticated), not TCP; a remote host still needs the TCP+mTLS listener, now sourced from OpenBao PKI rather than file-based certs.
  • The executor commits a dirty worktree (wip(agent): ...) before the safety-net push, instead of silently pushing an unchanged branch.
  • A delegated task now carries factory.branch/factory.base_commit metadata and the receiving role fetches and checks out that branch deterministically, instead of starting review at the checkout's own HEAD.
  • /memory export now uses one stable factory/team-memory branch instead of a fresh branch and PR every time -- docs/team-memory.md is updated in place, an unchanged export is a no-op, and a merged/closed PR just means the next real change opens a fresh one from the same branch.
  • vmdctl sudoers gains repeatable -extra-base: the sudoers render was anchored to exactly one chroot base directory, but this repo own vmd test suites each jail their throwaway VMs under their own /var/tmp/* directory -- a grant anchored to only the configured jail_base_dir 403d every real privileged call those suites made. The checked-in deploy/host/sudoers.factory and just doctor own drift check are now both generated the same way (dev.yaml plus the three test-suite base globs), closing a pre-existing mismatch between them.
  • the worker's run-stall guard now polls the agent-server's own /server_info endpoint for streaming activity during a single long LLM call, closing a gap where a 16-minute single completion with no intermediate events could be misread as a stalled run
  • Per-agent LiteLLM keys allow 400k tokens per minute (was 100k): one coder request with MCP tool schemas exceeded the old cap and looped on 429 forever.
  • confirm_risky no longer pauses for a person on rm -rf commands confined to the agent's own worktree -- only deletes reaching outside it, or any other risky action, still ask
  • The Aiven AI Gateway is now the only model provider: the DSV4 dev endpoint and the LOCAL_LARGE_* placeholder are gone from LiteLLM, compose, the justfile precheck, doctor and CI; open-coder/frontier are the role defaults (local-coder/local-large remain as compatibility aliases for the same models), and team memory runs keyword-only because the gateway serves no embedding model.
  • A worker that loses a task's lease while the run is going (the task was re-dispatched, or its lease row is gone) now stops that run and leaves the task alone: no terminal state, reply, push or hand-off from it; the worktree and conversation stay for the re-dispatched run. Counted in factory_lease_lost_total{reason}.
  • just stack-status now probes each process's own port/socket (not just checking that its pid is alive) and reports a process that is running but not accepting connections as a distinct WARN, exiting non-zero.
  • Jaeger bumped to 2.21.0; the telemetry test, the e2e trace check, and the upgrade runbook now use its v3 query API instead of the removed legacy /api/ endpoints.
  • Integration tests, container-starting tests and live-external-platform tests now require explicit opt-in via one of three environment variables (FACTORY_TEST_CONTAINERS, FACTORY_TEST_DEV_STACK, FACTORY_TEST_LIVE_PLATFORMS) instead of running automatically whenever the resource happens to be reachable; see docs/dev-environment.md's "What a test run is allowed to touch".
  • The real agent-server suite, four container-starting integration tests, and the eval/e2e end-to-end suite now require the same explicit opt-in as every other integration test (FACTORY_TEST_CONTAINERS / FACTORY_TEST_DEV_STACK) instead of running whenever Docker or the dev LiteLLM happened to be reachable.
  • just demo/e2e now fails outright if every single test in the eval/e2e suite skipped, instead of reading as a pass -- e.g. when the two consent variables above weren't set.
  • Role prompts tell an agent not to end a turn with a closing offer or question unless a real answer is needed from the requester.
  • credgw denials rejected_by secrets (a startup probe with no credential yet) are logged but no longer posted to the team's room; allowlist denials are posted in plain words naming the host
  • team create now says a name belongs to an archived team, instead of the generic "already exists", when the collision is with an archived team's name
  • team create now requires --template explicitly and lists the available templates, sorted, when it is missing or names one that doesn't exist
  • A pull request's title now comes from the branch's own first commit (the oldest one, not the latest) when there is one, before the task text, before the branch name — a rework round's later commits no longer change what an already-open pull request would have been titled, and titles are no longer taken from an instruction or review comment's own first line when a perfectly good commit subject already existed.
  • a role's own required skill (named by its front matter or the team's template) now fails the task if it cannot be resolved or fully staged into the guest, instead of silently continuing without it; skills from the repository or the team's curated set stay best-effort
  • a role backed by the lean (PydanticAI/Deep Agents) runtime now gets the same required-skill guarantee a role on OpenHands already has -- a required skill that cannot be resolved or fully staged now fails the task, instead of silently continuing without it
  • The worker now tells a lease expiry the worker itself caused by its own shutdown from one nobody explains: a task survives up to five shutdown-caused redispatches (against one retry for an unexplained lease expiry) before the reconciler gives up on it, and the worker refuses to start against a bookkeeping database missing the columns this needs.
  • LiteLLM counts spend for qwen3-next-80b (0.50 / 10.00 USD per million input / output tokens) and counts gpt-5.6-sol at 1.00 / 10.00 USD instead of its built-in list price; both are the operator's working figures in deploy/compose/litellm/pricing.yaml, so a team's daily budget now binds for the models every role runs on.
  • The dev OpenBao runs 2.7.0 and keeps its data across restarts once OPENBAO_STATIC_UNSEAL_KEY is set in deploy/compose/.env (raft storage on the openbao-data volume, unsealed by itself at start).
  • A successful /team create now says which chat surface and room it used (e.g. "home: mattermost, room #demo"), so a create with no --home no longer leaves that silent
  • The agents' OpenHands runtime moves from 1.47.0 to 1.49.1: a bash command with a working directory outside the conversation's workspace is refused, and a run whose status update stalls is polled for up to 30 seconds instead of hanging. 1.49.2 and newer stay out until OpenHands lifts its cap on fastmcp (recorded in versions.yaml).
  • The A2A library moves to a2a-sdk 1.2.0 and protobuf to 7.36.2: an operation on a task that has already ended is refused as unsupported instead of as an invalid parameter, and a message whose context does not match its task is rejected.
  • credgwd refuses to start when -llm-port or -mcp-port is unset while its host is set, instead of handing agents an llm_url/mcp_url with port 80.
  • The Mattermost thread no longer gets a separate Result post; the lead role's closing post carries the result. The internal conversation branch (openhands/) is hidden on every surface.
  • A role waiting on a long model answer is no longer failed as stalled; a hang with no tool running is caught by the run timeout instead.
  • A reviewer's confirm_risky pause on a worktree-confined delete chained with && to LOW commands (rm -rf pycache && git status) is now auto-approved; any other chain still escalates.
  • A risky-action confirmation now shows what the action runs (the terminal command, the file path or URL, otherwise its arguments) after the agent's own summary.
  • prune-teams (archive_stale_teams.py) takes --vmd-addr / FACTORY_VMD_ADDR and refuses --apply without it; stack-up's prune passes the dev vmd socket; a PARTIAL archive names each agent that was not destroyed
  • A skill a role declares that cannot be staged into the guest (its directory is missing or staging fails) now fails the task before the first model call, naming the skill and the reason; it was only a warning.
  • Mattermost threads: a role's status (working, done, a verdict with its branch and pull request) is a new reply each time it changes, batched to one per role per flush, and a post is never edited; before, each role had one post edited in place, which hid the sequence of events
  • Every role runs without confirmations now (confirmation_policy: never_confirm for the reviewer too, in the role file and both feature-team templates): no agent action pauses for a person's approval; the sandbox is the boundary and a pull request is still merged by a person
  • Mattermost thread cards read as a sequence: a "working" card is bare, a done card carries only that round's branch, commit, pull request and text, the lead posts no "working" cards between delegate replies and never repeats a text, the worker's work-item-record tail is dropped from bodies, and the closing post never doubles the lead's last card
  • The reviewer no longer sends work back for whitespace or formatting alone (it notes it, unless the project's own lint/CI enforces it), and the coder runs the project's formatter over the files it changed before committing.
  • A failed delegate's report to the lead now carries the failed task's brief and the work item's branch, and the lead re-delegates a limit failure once on that branch instead of asking the person for the requirements.
  • A role setting changed with /team set (max iterations, confirmation policy, model, ...) now applies to the next task at once; the worker no longer caches the role config for 30 seconds.
  • The project is renamed from Factory to Kapelle: packages, environment variables (KAPELLE_ prefix), wire subjects and keys, metrics, secret paths, host paths and the chat command, which is now /kapelle. Old names are not read; scripts/rename_env_prefix.py converts a user's own env files.
  • Agent images rebuilt on the refreshed package set: a2a-sdk 1.2.2, mcp 2.3.0, fastapi 0.142.2 and 22 more; new Nevia golden checkpoint chk_255c31d2574e and Firecracker image agent-python-a32a1a036bf2.
  • The OpenTelemetry collector in the dev stack is 0.162.0.
  • The dev stack's object storage (RustFS) is the 1.0.1 GA release; an existing volume is migrated on first start.

Removed

  • Zulip is no longer part of the dev stack: the zulip services, volumes, secrets and edge route, just dev-up-surfaces, just zulip-realm-link and the bridge-bot script are gone; just stack-up no longer requires a chat container
  • Zulip is no longer a chat surface: the gateway has no Zulip adapter, /team create --home accepts slack, mattermost and (for probes and tests, not advertised) none, the default home is Mattermost, and /help no longer lists Zulip. The FACTORY_ZULIP_* settings are ignored, except a FACTORY_DEFAULT_HOME_SURFACE that still says zulip, which stops the controller at startup with a message. A team whose home or link is Zulip gets no replies or notices (a notice for a surface without a poster is dropped with one INFO line); /team link zulip is refused.

Fixed

  • vmd's own OpenBao disk-encryption-key path disagreed with credgw's read path (a bare replica number vs. the full agent id) — found running a real controller+credgw+vmd pipeline end to end for the first time.
  • Several real Firecracker/jailer process-orphaning gaps on test teardown (vmd itself, the jailed Firecracker process, and a go run wrapper process all now guaranteed cleanup on every exit path).
  • wake() minted its own session key instead of using vmd's real, persisted session_api_key.
  • Role prompts (roles/*.md) were written in task 22 but never actually loaded into the executor's system prompt until a later fix — the first eval runs had measured an unprompted agent.
  • A flaky Postgres-container readiness race (pg_isready over the default Unix socket, not TCP) in the contracts integration suite.
  • Compose's ports: merge semantics (concatenates sequences across -f files rather than replacing them) initially let the isolated smoke-test project collide with the shared dev stack's fixed ports; fixed with the !override YAML tag.
  • credgw's per-agent iron-proxy instances relied on iron-proxy's own built-in 30s upstream-response-header timeout, too short for a real, possibly non-streaming LLM completion under load or a large git push — every instance now sets an explicit, generous timeout.
  • /team archive only removed Agent Cards and deleted the team's LiteLLM team record — it never destroyed the team's agents, released their credgw instances, or deleted their per-agent OpenBao secrets/LiteLLM keys, leaking all of them on every archive. The gateway's A2A routes and delegation follower for an archived team now also actively stop instead of running forever. Replica scale-in had the identical gap (never persisted the destroyed state or released the replica's credentials) — both now share one release helper.
  • A worker's durable-reply notification (a task completing/pausing sent a SendMessage back to the requester) was itself recorded as an ordinary, repliable task — its own later completion could trigger another reply back, and so on indefinitely, burning real LLM calls on nothing. A reply notification is now marked and never itself replied to. As defense in depth against the next unknown loop of this shape, a worker now also refuses to create more than a configured number of tasks (default 8) for the same work item and role, failing loudly with a notice instead of looping unboundedly.
  • Even after the reply-notification fix above closed the unbounded ping-pong, every hop of a real delegation chain still left one bounded "echo" behind: a delegator that had already handed off and finished still got a brand-new task and a wasted LLM turn from a plain "completed" notice about work it no longer cared about. A durable COMPLETED reply is now only published when the requester's own task is still actively waiting on it (mid-run or INPUT_REQUIRED); replies reporting FAILED/CANCELED/REJECTED/INPUT_REQUIRED always go through regardless, since those are never "just FYI".
  • Nothing ever deleted a team's per-role NATS durable consumers on archive or on a test process's teardown, leaking hundreds of stale consumers on the dev NATS instance over time and permanently blocking worker shared mode's wildcard consumer if any were left over on the same stream. A worker can now delete its own durable consumer on clean shutdown when explicitly marked as owning that (team, role) for its own short lifetime (test harnesses only, never a real deployment); /team archive now also deletes the archived team's per-role consumers as part of its own teardown, alongside its agents/credentials/cards.
  • A new task could allocate a brand-new sandbox replica while an existing one for the same (team, role) was already awake and genuinely idle — the sandbox's own reported RUNNING/IDLE state wasn't a reliable enough signal to tell "awake and free" apart from "awake and mid-run". Placement now cross-checks whether an awake replica has any in-flight task before treating it as free, reusing one over provisioning a new one whenever possible.
  • The executor's HTTP client used httpx's own 5s default timeout for every phase (connect/read/write/pool alike), too short for a real worktree conversation create (real git-worktree setup inside the guest) under host load — it now gets an explicit, generous read/write budget while keeping a short connect/pool timeout.
  • The executor's keep-awake loop only started once the run itself began, leaving the conversation setup beforehand (resolving role config, building the agent, cloning the repo) covered by nothing but two one-off touches — under host load a sandbox could go idle and sleep mid-setup. It now starts right after the sandbox wakes and covers everything through the post-run push.
  • vmd's per-tap nftables gateway-accept rule was inserted before the catch-all drop but after the three RFC1918 destination-drop rules — since an agent's gateway address is itself RFC1918 space (the whole network pool's CIDR is), one of those drops silently killed every gateway-destined packet before the accept rule was ever reached, so a guest could not reach its own credential-gateway instance on any port. The accept rule now inserts before every destination-based drop rule, not merely the catch-all.
  • vmd derived a credgw instance's real listener ports by parsing its llm_url/mcp_url/proxy_url response fields — those are guest- facing URLs built for a different purpose (llm_url/mcp_url in particular carry the raw upstream address, e.g. LiteLLM's own host, never anything credgwd itself binds), so the gateway-accept rule above ended up allowing unrelated upstream ports and never discovered the https listener at all. credgw's StartInstance response now reports its own bound http_listen/https_listen ports directly.
  • vmd's Destroy deleted an agent's data disk with no backup at all — the nightly ticker and the deep-sleep transition both back up to object storage, but archiving a team (which calls Destroy directly, with no prior sleep step) permanently lost a RUNNING or merely-ASLEEP agent's data disk. Destroy now backs it up first, matching the deep-sleep path, and refuses (rather than deleting anything) if that backup fails while object storage is configured. A new skip_backup request field (vmdctl destroy --skip-backup, and an optional skip_backup kwarg on SandboxProvider.destroy()) opts out explicitly, for operator use and test teardown — the controller's team-archive path never sets it, so archiving always keeps the safety net.
  • vmd's Capacity RPC counted only the brief mid-transition SLEEPING state as "sleeping", missing the settled ASLEEP/DEEP_SLEEP population entirely — the reported count (surfaced via vmdctl capacity and the controller's hosts table) was ~always 0 regardless of how many agents were genuinely asleep. Now counts all three.
  • Firecracker's kvm-pit kernel thread (guest PIT timer emulation) landed in the host's root cgroup by default on every VM, not the VM's own — Firecracker's own docs say only an external agent can fix this, since the kernel creates the thread post-boot, after Firecracker/jailer itself is deprivileged. vmd now moves it into the VM's own cgroup right after every boot and restore.
  • Slack replies now stream (chat.startStream/appendStream/ stopStream, live-updating as an agent works) instead of always falling back to plain chat.postMessage/chat.update — the adapter code was fully built and unit-tested already, but nothing wired it into a real gateway process. On by default whenever Slack is configured; opt out with FACTORY_SLACK_STREAMING=false on a workspace plan that doesn't support streaming.
  • credgwd's -llm-host/-mcp-host are documented host-only, but a host:port value was silently accepted and then 403'd every LLM/MCP request (the allow rule's Host still carried the port, the agent's own request Host didn't) — both flags now fail fast at startup on a port or a URL scheme, naming the flag and the corrected value.
  • The guest-facing llm_url/mcp_url credgwd hands an agent (via vmd's AgentEndpoint) never carried a port, so a real microVM guest's LLM/ MCP client dialed the default port 80 instead of the real one — new -llm-port/-mcp-port flags supply it without affecting the allowlist rule, which must stay port-free.
  • Those same llm_url/mcp_url also carried a trailing slash, and both real consumers (the OpenHands SDK, the worker's MCP client) append their own leading one — the doubled // 404'd against a real LiteLLM, which doesn't normalize it.
  • mcp_url was built with the same bare-host shape as llm_url, but the worker's MCP server mounts at /mcp/<conversation_id>, not the host root — mcp_url now always ends in /mcp (contract decision, matching the fake sandbox provider's own already-correct shape).
  • Every role's mcp destination allowed only Content-Type/User-Agent/ Accept through to the real MCP server, stripping the MCP streamable- HTTP transport's own required Mcp-Session-Id/Mcp-Protocol-Version/ Mcp-Method headers — every call after the initial initialize 400'd as a result.
  • Jira's Forge remote-agent transport (JWKS-verified rovo:agentConnector, design doc's preferred Jira path) was fully implemented and tested but never actually reachable in a running gateway process — nothing mounted it. Setting FACTORY_JIRA_FORGE_APP_ARI/_TEAM/_ROLE now mounts it at /jira/forge/a2a; unset, the gateway behaves exactly as before.
  • The idle detector could sleep a freshly woken microVM before the worker's own first Touch() ever arrived — a real rerun with a short idle_after caught the gap between Wake returning and the executor's first touch (role config resolution, a repo clone), which isIdle's conversation-count check has no visibility into. Wake now sets the same keep-awake deadline Touch does on both the cold-boot and restore paths, so a Sleep call in that window is refused exactly like an explicit Touch would.
  • The compose worker container's Prometheus metrics server bound 127.0.0.1 inside its own network namespace, unreachable from the published 9292:9292 port -- Prometheus's own scrape target for it has been silently failing since the services profile existed; now binds 0.0.0.0, and the smoke test asserts the endpoint actually returns Prometheus text.
  • /memory export only ever rendered curated entries as chat text, the same as /memory search/list — design doc §14 promises a real docs/team-memory.md pull request on the team's linked repo, and docs/memory-commands.md's own wire contract already said that was the gateway's job, not the controller's, but nothing did it on either side. factory_gateway.memory_export.MemoryExportService now opens that PR when GitHub is configured on the gateway and the team has a repo linked; falls back to the previous plain-text rendering otherwise (no GitHub configured, no repo linked, or the PR fails to open for any reason).
  • vmd's jailer invocation dropped --daemonize so the guest's serial console and Firecracker's own operational log are captured to a per-run console.log instead of /dev/null, and vmd now notices and logs a jailed process's real exit status the instant it dies, marking the sandbox FAILED immediately if nothing vmd did caused it, instead of only the watchdog silently noticing later.
  • just eval --model previously had no effect on the target role's LiteLLM alias (NatsStackDriver accepted the parameter but never wrote it to team_roles); it now patches the role's config directly before the run starts
  • vmd's watchdog killed a real VM after a single non-2xx response from Firecracker's own API under real load (found: the actual cause of task 29's reported silent VM deaths) and logged nothing about it. A non-2xx status means Firecracker's API is up but not confirmed healthy, distinct from a connection-level failure meaning the process is actually gone -- the watchdog now requires WatchdogFailuresBeforeKill (default 3) consecutive failures of either kind before killing, logs every individual failure at WARN with status/body, and exposes the running failure count live via Status.
  • A memory proposal an agent made mid-run (the memory_propose MCP tool) was recorded but never surfaced to the activity feed at all -- invisible outside the agent's own conversation transcript. Now publishes to the same per-task activity subject every other run event uses.
  • justfile recipes that take free-form arguments (just changelog, eval, dev-logs, spend, scan-secrets, prune-teams, prune-workdirs) no longer re-parse those arguments through a shell, so a shell-special character like a backtick or an apostrophe reaches the underlying script literally instead of silently corrupting the command or breaking it outright.
  • OpenHandsAgentExecutor bounds a run against a guest that goes unreachable mid-run (default 60s, FACTORY_GUEST_UNREACHABLE_SECONDS) instead of waiting out the OpenHands SDK's own much longer, unbounded retry against a dead connection; the sandbox is marked unhealthy so the next placement gets a fresh replica.
  • vmd's own watchdog was the real cause of task 29's microVM deaths, not the guest: Firecracker's embedded API server caps at 10 simultaneous connections and vmd built a fresh, never-reused HTTP client on every watchdog tick, leaking one idle connection per check until the cap filled in under a minute and every later call -- including the watchdog's own -- got rejected and killed the VM. Each sandbox now reuses one long-lived API client for its whole lifetime instead of constructing a new one per call.
  • The GitHub adapter built a question's @mention from the requester's numeric GitHub user id instead of their login -- GitHub mentions only resolve logins, so the mention silently never notified anyone, found running the live integration test against a real GitHub App for the first time.
  • OpenHandsAgentExecutor also bounds a run that stalls while the guest stays reachable (no new agent-server event for FACTORY_RUN_STALL_SECONDS, default 600s) -- a real run hung for 16 minutes with Touch/reachability checks succeeding throughout; the sandbox is marked unhealthy the same way as the unreachable case.
  • Team-memory run-start injection and the automatic outcome entry each publish an activity-feed event now, the same gap already closed for memory_propose -- injected reads and writes were invisible outside the agent's own conversation transcript (design doc §14 "Audited").
  • Every role mcp destination header_allowlist was missing Mcp-Name, the 4th MCP streamable-HTTP routing header the mcp SDK checks for tool-call-shaped requests (Mcp-Session-Id/Mcp-Protocol-Version/Mcp-Method were fixed already) -- iron-proxy stripped it, failing every real tool call a coder/planner/reviewer made. A new test derives the required header set from the installed mcp package itself so a future SDK addition fails a test, not a real microVM run.
  • The gateway now warns in the room when an agent-opened PR targets a branch other than the repo's default -- GitHub's Closes #N auto-link only fires against the default branch, so a wrong-base PR previously failed silently.
  • The gateway's standalone token minting (RealGitHubTokenProvider, used when no controller is reachable) requested a token missing contents/metadata permissions, causing a 403 on any commit/branch lookup even though the GitHub App installation itself was granted correctly.
  • vmd never released a dead VM's network slot (tap + nftables) after an unexpected process exit (only Sleep/Destroy's own explicit paths did), so a retried Wake for the same agent could fail to boot with Firecracker's own "Open tap device failed: Resource busy" -- found from a real run. onJailExit and the watchdog now release the network slot whenever they mark a sandbox FAILED for an exit vmd didn't cause on purpose, using the process's real PID (not exec.Cmd's own Process.Pid, which is sudo's PID whenever vmd runs non-root) to confirm a delayed report is still about the currently-running process before acting. Each boot/restore attempt's own console.log also now survives a retry instead of being silently truncated.
  • vmd's own Wake RPC tied a newly-booted VM's process lifetime to that RPC call's own gRPC context, which grpc-go cancels the instant the handler returns -- in production (vmd running as root) this would SIGKILL the real Firecracker process within moments of every single successful Wake; in dev (sudo-wrapped) the kill landed on sudo's own PID instead, leaving the real VM running but orphaned from vmd's own bookkeeping, which is what a real run's "exited unexpectedly while RUNNING" mystery turned out to be. The VM's process now outlives the RPC call that started it, as intended.
  • Scale-out could never trigger for the default shared-mode worker deployment -- its backlog check only ever looked at a per-team consumer that only a pinned worker process creates. Now falls back to the real per-team message count on the A2A_RPC stream itself when no such consumer exists.
  • Zulip bridge bot could post and read messages but never received any: a freshly created bot is not auto-subscribed to any stream, so it silently never saw a channel it was linked to. The adapter now subscribes itself on startup and bootstrap-bot.sh gains --streams.
  • Worker sandbox keep-awake touches now carry an explicit TTL (twice the configurable touch interval, FACTORY_TOUCH_INTERVAL_SECONDS) instead of relying on vmd's own default touch window, which a short idle_after could outlast.
  • The guest-unreachable bound no longer misfires on a real agent-server that is simply slow to answer a health check while busy (e.g. mid real-LLM-call) -- only a failed connection attempt now counts toward it, not a read timeout on an already-established connection.
  • vmdctl sudoers now takes a -alias flag (default FACTORY_VMD, was the hardcoded FACTORY_VM) so its generated sudoers grant does not collide with a Cmnd_Alias of the same name in another sudoers file already installed on the host
  • coder no longer tries to rename or push its own git branch: the OpenHands SDK names a conversation's worktree branch itself (openhands/, not configurable), and coder's old instruction to push agent// directly collided with that, wasting the run's iteration budget and getting rejected by the credential gateway's refcheck. The platform's own post-run push already lands the agent's commits on the correct ref; coder now only commits.
  • A deterministic pipeline hop to the next role no longer delays or blanks out a completed tasks own status report, and only fires when the run actually produced new commits to hand on
  • A model that ends its turn by calling finish with its full answer and no separate chat message no longer has that answer silently replaced by a generic completion message
  • A guest killed while its trigger request was in flight no longer crashes the executor uncaught -- the run now fails within guest_unreachable_seconds like any other dead guest
  • CONNECT-based HTTPS destinations (GitHub, npm, any TLS-intercepted destination reached via HTTP_PROXY/HTTPS_PROXY) now actually work end to end -- credgw renders iron-proxy's real tunnel listener and scopes its allowlist to allow the CONNECT itself, which previously rejected or silently mishandled every such request
  • send_task (and the deterministic pipeline hop) now reject an unknown target role immediately instead of waiting the full ~60s delegation acknowledgement timeout
  • Worker git clone/push/fetch now send the placeholder Authorization header credgw requires for github-api/github-git destinations, matching the contract _github_curl already used
  • GitHub git-over-HTTPS pushes/clones/fetches now send Basic auth (matching what GitHub's git smart-HTTP endpoints require) instead of Bearer, which credgw's secrets transform swapped correctly but GitHub itself rejected
  • A guest killed anywhere during a run's setup phase (not just after the trigger request) now fails the task within guest_unreachable_seconds instead of crashing the executor uncaught
  • Worker: fixed a bug where a completed coding run could push the wrong commit to GitHub (or push nothing at all on a resumed conversation) while still reporting success; pushes now verify the remote landed the expected commit and the completion artifacts include the exact pushed branch and commit.
  • credgwd no longer hard-fails StartInstance for a team with no GitHub repo configured -- it drops the github-api/github-git destinations and continues; a real controller/mint failure still fails loudly
  • the guest image now overrides openhands-sdk's hardcoded 30s MCP tools/list timeout locally, via a sitecustomize.py hook gated on a new FACTORY_MCP_TOOLS_TIMEOUT_SECONDS guest env var
  • push_branch retries its own post-push remote verification for a bounded window instead of reporting a landed push as failed on the first read, and always logs the expected vs. observed commit sha
  • pip install --user now works inside the guest microVM (a new home-agent-.local.mount bind mount, the read-only rootfs was silently swallowing the write)
  • Controller: a dedicated team's own pinned worker can now actually use team memory (memory_search/get/propose) -- its own NATS account never had permission to reach the controller's command subject, and nothing ever registered the account with the running NATS server at all, so every call silently timed out.
  • a reply answering an agent's INPUT_REQUIRED question now continues that same task (and its real git branch) instead of silently starting a new one with its own untracked branch
  • credgw's per-team GitHub policy now allows GET on the bare repo path (github-repo-metadata destination), fixing the coder's own PR-opening default-branch lookup that iron-proxy previously 403'd
  • Worker: the pydantic_ai comparison runtime no longer crashes when send_task/ask_requester fails deterministically (e.g. delegating to a role nobody's listening for) -- it now reports the failure to the model instead of exhausting a retry budget and hard-failing the whole run.
  • Worker: the pydantic_ai comparison runtime no longer crashes when list_agents is called with the wrong arguments repeatedly (a zero-arg tool a small model sometimes still guesses kwargs for) -- retries still get their normal chance to self-correct, but exhausting the budget now reports an ordinary failed tool result instead of hard-failing the whole run
  • A team whose template's lead role isn't literally lead (e.g. planner) never received the first message of a new conversation — the gateway always dispatched to a hardcoded lead role instead of the team's own teams.lead_role. The gateway now resolves each team's lead role from its own directory mirror and uses it consistently for dispatch, status/cancel commands, and event tagging, falling back to the old hardcoded default (with a warning) only when the team is unknown.
  • Creating a team with a home Zulip/Slack channel created and subscribed the channel but never recorded it as linked — the channel was unusable until a manual /team link. TeamService.create now writes the surface_links row for the home channel in the same transaction.
  • /team create (and /help) typed in a channel the bot is subscribed to but not yet linked to any team was silently dropped — a chicken-and-egg gap, since /team create is exactly what would have linked that channel in the first place. The gateway now handles /team ... and /help on any subscribed-but-unlinked channel before team lookup; every other message there is still ignored.
  • A channel linked to a team after the gateway process started (including via a freshly typed /team create) was never routed to until a full restart — team_by_channel was only built once at startup. The gateway now refreshes it on the same periodic pass that already keeps A2A routes and delegation followers live, plus immediately after this gateway process itself handles a successful /team create/link/archive.
  • The controller's own Zulip client had no CA/TLS-verify knob (the gateway's already did) — it now honors FACTORY_ZULIP_CA_FILE/FACTORY_ZULIP_TLS_VERIFY the same way, sharing the gateway's httpx_verify helper (moved to factory_contracts).
  • Docs and .envrc named the Zulip realm URL variable FACTORY_ZULIP_URL, but both the gateway and controller only read FACTORY_ZULIP_SITE, silently ignoring it. Both services now accept FACTORY_ZULIP_URL as a fallback, logging which name was used; FACTORY_ZULIP_SITE remains the canonical name.
  • A shared-mode worker no longer redelivers or acts on A2A_RPC messages for an archived (or unknown) team -- the controller's EnsureAgent also refuses to provision an agent for a non-active team, and archiving a team now purges its own stale A2A_RPC messages outright, not just its durable consumers. Real repro: cleaning up leaked consumers for already-archived eval/e2e teams released hundreds of stale messages that a shared worker then acted on, re-provisioning agents (real VMs, real disks) for teams that no longer exist.
  • The controller's real Zulip client never actually worked against a real realm — /team create with a Zulip home channel died with internal error handling command (POST /api/v1/channels/create rejecting the request for a missing subscribers field, then a second real discrepancy in the response shape). RealZulipClient.create_channel now sends subscribers as the bot's own numeric id plus the creating person's own linked Zulip id when they have one (so they see the room they just made), and any Zulip/Slack failure during /team create now surfaces as a clean ok: false message instead of that opaque internal error.
  • credgwd rejects clashing listener ports at startup (a real run's -http-port matching the tunnel listener's own default 8080 used to fail every instance with a bare timeout) and includes the instance's own iron-proxy log tail in a failed StartInstance error, instead of a bare "did not come up" timeout with no clue why.
  • vmd cleanly stops (sleeps, or force-stops as a fallback) every running VM on SIGTERM/SIGINT before exiting instead of leaving jailer/Firecracker children orphaned; Reconcile only adopts a still-running process into memory once it confirms the guest's own agent-server is reachable and its gateway instance can be (re)started, destroying and letting the next Ensure recreate it otherwise; a boot that fails after the jailer process has started now kills that process and cleans up instead of leaving it running with the record marked ERROR
  • A worker whose Wake fails because vmd lost the sandbox (e.g. after a vmd restart) now self-heals: it asks the controller to force a real re-Ensure through the provider and retries Wake once, instead of failing the run outright. The controller's own EnsureAgent does the same re-verification when asked, rather than trusting a possibly-stale placement row. The resulting failure message (when the sandbox genuinely can't be recovered) also now names the real cause instead of just the bare agent id.
  • planner and reviewer agents can now clone the team repository through credgw (both gained a read-only github-git destination; the executor clones for every role, but neither policy had one, so cloning failed with a bare CONNECT-tunnel 403); planner also gained the read-only github-repo-metadata destination it was missing.
  • vmd cleans up a boot failure's tap, nft rules and cgroup under a bounded, non-cancellable context instead of leaking them; the IP allocator reclaims a /30 whose gateway address is already held by a stale host interface and Reconcile sweeps orphaned taps/cgroups on startup; KillJail verifies the process is actually gone and logs the pid it signals; just doctor warns on leaked factory taps
  • A planner (or any role that never writes code) delegating to another role no longer attaches a nonexistent git branch to the delegation -- send_task only attaches factory.branch for a role whose own definition declares it actually pushes one (produces_branch: true, set on coder). Before this fix, a planner delegating to a coder always attached its own never-to-be-pushed branch, and the coder's pre-guard failed the whole delegated task waiting for it.
  • A resumed conversation that once hit its iteration cap no longer crashes every later turn instantly -- the SDK's per-run budget always resets correctly, but the executor never caught the run failure it raises; now it does. Planner's max_iterations raised 15->40 in the feature-team template.
  • Team skills accepted via /skills accept now actually reach agents (the resolver's team source was previously a stub that always returned nothing).
  • vmd's per-agent data disk no longer costs its full nominal size twice on disk: the at-rest encrypted copy and the running VM's decrypted working copy are now compressed/sparse, dropping an 8 GiB small-profile disk's real footprint from ~15 GiB to well under 100 MiB when lightly used; vmdctl prune-agents also finds and removes jail directories with no state record at all; just doctor reports each agent's real per-agent disk footprint
  • A wake/replica-ensure failure before any OpenHands conversation exists (e.g. the controller's own RESOURCE_EXHAUSTED when no host has room for a new replica) used to crash the executor uncaught instead of failing the task cleanly -- now reported as a normal TASK_STATE_FAILED outcome.
  • A durable reply's text can no longer be empty -- falls back to the terminal-state name (e.g. "TASK_STATE_FAILED") if the summary is blank, since an empty user-role message appended to a resumed conversation is one shape a Bedrock-backed model rejects outright.
  • Semantic memory search on the shared dev DB no longer fails: entries.embedding is repaired to a real pgvector column, the controller verifies the column type at startup and falls back to keyword-only search with a WARN on mismatch instead of erroring every /memory search, and a broken embedding parameter no longer dumps its ~16 kB vector into logs.
  • worker: a still-WORKING requester now correctly waits for a delayed reply instead of losing it or getting a duplicate task
  • reconcile no longer destroys an adopted-but-unreachable agent's state record outright -- it kills the stale process and marks the record ERROR, keeping it, so an unreachable-at-restart guest is never indistinguishable from a real operator-requested Destroy; only vmdctl destroy or prune-agents' own criteria remove a record for good
  • worker,controller: pushed commits are authored by the real agent identity (Factory ()) instead of the OpenHands SDK's default 'openhands' identity, configurable per role via a new git_identity_domain template field
  • The run guard (reachability/stall watchdog) no longer destroys a sandbox on a false positive -- a transient host-load spike used to get misjudged as a dead guest and permanently erase its worktree/conversation. The guard now requires several consecutive connect failures with a timeout that scales under load (not a single fixed-timeout check), and its recovery action stops the guest via vmd's new Stop RPC (keeping disks and the sandbox record) instead of destroying it.
  • controller: a coder replica whose sandbox went unhealthy no longer causes an unnecessary extra replica -- ensure_replica now tries to self-heal it first, and every scale-out/self-heal decision is now logged; a role's front matter with a field this controller doesn't yet recognize now warns instead of crashing GetTeamRole
  • Staging a skill's files into the guest workspace no longer fails with a permission error on a fresh sandbox -- falls back to the agent user's own passwordless sudo when the plain workspace path isn't writable yet.
  • platform MCP server declarations no longer attach to every conversation by default -- only when a template, repository, team, or skill explicitly references them
  • A role's own declared tool list is now enforced across every runtime (OpenHands, Pydantic AI, Deep Agents) -- a role that never lists a shell tool can no longer run one, matching what its roles/.md front matter actually says.
  • A completed run's outcome text no longer appears twice in the room (once as an artifact bullet, once as the final message) or in its team-memory outcome entry.
  • Enforcing a role's own declared tool list no longer risks sending an unregistered tool kind to a real agent-server -- grep/glob collapse onto file_editor instead of being sent as their own SDK tool names, which crashed a live run.
  • A confirm_risky risky-action approval question can no longer be delivered to the delegating agent that triggered it -- it only ever reaches a person through the home room now. The gateway used to only track the LEAD role's own pending task, so even a person's reply to a delegated role's own pause (e.g. a reviewer's confirm_risky question) either misfired onto the lead's unrelated pending task or silently started a brand-new one; every role's own pending task is now tracked, and a reply addressed to a delegated role (@role) resumes that role's own task. Also: a message from ANY bot/integration account (not just this gateway's own bridge bot) is now refused before it can be treated as a person's reply or approval.
  • Co-authored-by trailer only amends a genuinely new commit this run's own agent made, never a pre-existing one; reviewer/planner never attempt a post-run push, matching their produces_branch: false
  • confirm_risky's own classifier no longer crashes a real run with a 500-shaped error when configuring the security analyzer -- it now uses a real, vendor-recognized analyzer kind instead of a custom one the agent-server never registers.
  • The controller's proactive self-heal now recognizes a genuinely stuck-in-provisioning replica instead of always leaving it alone, and a real worker's provider now keeps agents.state in sync with the sandbox's actual wake/sleep state
  • Agent-to-agent replies (an approval, an answer to ask_requester) now resume the original delegator's own paused task instead of always starting a new one
  • Aiven-backed model aliases accept streaming requests again: LiteLLM drops the stream_options field the worker now sends.
  • external MCP server tool filtering no longer drops the coder's own shell/file-editing tools
  • the real Firecracker provider integration test suite (services/worker/tests/sandbox/test_firecracker_provider_integration.py) had never actually run: its own prerequisite check probed sudo against firecracker directly (never granted; only jailer is), the kernel path it gave vmd fell outside every sudoers cp grant, and vmd's real touch-window default (300s) meant Sleep could never succeed within the test's own margin. Fixed all three plus a real resource leak found running it for the first time: go run ./vmd decoupled the test's own process-wait from vmd's real shutdown timeline, leaving live Firecracker processes behind on test failure -- now builds and execs the vmd binary directly, matching just vmd-dev, with a wait budget that actually covers vmd's own graceful-shutdown timeout
  • Agent images: SSH-form GitHub remotes are now rewritten to HTTPS correctly; the guest agent's insteadOf key carried literal quote characters, so every git@github.com: remote failed instead of routing through the gateway.
  • A delegate's completion reply is no longer silently dropped when its requester has already completed -- it lands as a fresh task for the requester's role instead of being discarded, so a delegation chain can no longer break because two hops happened to finish in the wrong order. A role finishing a configured pipeline hop (e.g. a reviewer for pipeline: {coder: reviewer}) now routes that landed reply to the team's lead instead of back to whichever role technically delegated it, and any landed reply now carries the requester's own conversation history/work item so it continues with real context instead of starting blank -- an earlier version of the fix above landed context-free on the wrong role and failed outright in a real run.
  • Nevia agent bootstrap no longer excludes the LLM/MCP gateway hosts from NO_PROXY, which bypassed credgw's remote listener entirely for those requests
  • The lead no longer polls get_task after delegating (which failed its first task as stuck): the worker's built-in MCP tools are offered only to roles whose tools: list names them, and the planner prompt says to end the turn after delegating.
  • A coder that created and committed on its own git branch no longer loses its work: the platform now pushes the commit the reviewer hop is based on (the worktree's HEAD after the run), not a fixed branch name, so the reviewer sees the change.
  • A question or approval request a role raises a second time within two minutes is delivered instead of being silently dropped, which used to stall the work item.
  • An empty answer from the model (no output, no tool call, zero tokens; seen from gpt-5.6-sol) no longer reaches an agent as a success: the LiteLLM proxy turns it into a retryable 503, so the agent retries instead of its stuck detector failing the task.
  • A reviewer no longer reviews a coder's rework push twice: the pipeline hop is the notification, the duplicate completion reply is not sent.
  • The LiteLLM empty-answer guard now also catches the empty answer actually seen from gpt-5.6-sol: a completed response with zero tokens and one message item holding only blank text.
  • The Nevia backup export no longer hangs when Nevia's exec stream stalls on a freshly restored computer: a stalled stream is detected within 20 s, its upload aborted, and the export retried on a fresh stream (up to 3 attempts) before failing with a clear error; nothing is deleted on failure.
  • A Nevia computer provides the same /workspace layout as a Firecracker one (a link to /data/workspace, made at image build and re-checked on every wake), so an agent on Nevia can clone its repository into /workspace/project. The golden image needs a rebuild.
  • A Nevia agent can reach HTTPS destinations such as github.com through the credential gateway: the worker fetches credgwd's interception CA and installs it in the computer's trust bundle (a remote instance's own certificate was only the public listener's), and wake() fails clearly when credgwd cannot supply it.
  • A Nevia relay stream that goes silent without closing (Nevia can stall an exec stream) is detected within 30 s and replaced: the in-computer relay answers the worker's 10-second pings, and a stream with no inbound byte for 30 s is aborted and restarted. Needs a golden image rebuild for the relay script; older images keep working without the detection.
  • The Nevia forward relay (worker to agent-server) cancels its in-flight connections when the relay closes (sleep, destroy, shutdown) instead of leaving them to be reaped with a 'Task was destroyed but it is pending' warning, and closes a connection whose stream stalls after the request was sent (120 s without a byte) instead of hanging forever.
  • The Nevia forward relay detects an exec stream whose relay process never started (Nevia can stall a fresh stream) within 20 s and retries the connection on a fresh stream, up to three attempts, instead of hanging; the request the client already sent is replayed.
  • The Nevia relays (the forward relay to the agent-server and the reverse relay to credgw, with its pings) run on a dedicated thread with its own event loop, so a worker whose main loop is blocked by a synchronous SDK call no longer deadlocks its own relay: the task used to hang for 10 minutes with no model call, and the in-computer relay exited on its idle rule.
  • The worker no longer exits on a transient NATS or Postgres error: the reconciler (leader election and sweeps), the lease heartbeat and the NATS pull loop log the failure, back off and retry, and the reconciler stops acting as leader while it cannot renew its lock. The stack script rotates a process log above 64 MiB at start.
  • Releasing a Nevia agent (archive, scale-in) now wipes its OpenBao secret document: the worker used to overwrite the recorded OpenBao path with a nevia marker on wake, so the document (disk key, MCP token) stayed behind.
  • credgwd can run several remote (Nevia) instances at once: each remote instance now gets its own metrics port instead of the shared 9090, which made every instance after the first fail to start ('metrics: ... address already in use').
  • A pull request opened by the worker is titled after the task the agent was given (first line, at most 72 characters), not after the agent's one-line final message; the body keeps the final message and adds a task footer. The safety commit for changes the agent did not commit is now chore(agent): <task title> with the task id in the body instead of wip(agent).
  • Releasing an agent (archiving a team, scale-in) now wipes its OpenBao document and deletes its LiteLLM key for Fake and Firecracker agents too; the provider's own row update on wake used to overwrite the document's path, so both were left behind.
  • The coder role now explicitly commits and runs tests before finishing a task; measured offline against real replayed requests, the old prompt left the platform's uncommitted-changes safety net to catch the work far more often than the new one does.
  • The controller's database migrations now reach head on a Postgres without the pgvector extension (a managed Postgres like Nevia's): team memory's embedding column stays a plain text placeholder there instead of blocking every later migration, and starting the controller with FACTORY_EMBED_ALIAS set over a non-vector column now fails clearly at boot instead of silently falling back.
  • The gateway no longer gets stuck routing a dedicated team's traffic through the shared NATS connection forever when the controller wasn't up yet at gateway startup: it now keeps retrying in the background instead of trying once. The controller now names the missing NATS KV bucket and points at just dev-seed instead of crashing with a bare NATS error.
  • credgwd's StartInstance and vmd's Ensure/Wake/Sleep now answer with a distinct, non-retryable status when an agent's OpenBao secret document is gone (the store lost its data, or something else deleted it), instead of a generic retryable error that looked like a transient failure; the worker's Nevia provider now recognizes the same condition from credgwd directly.
  • Fixed the gateway hanging indefinitely, process alive but not accepting connections, after receiving a shutdown signal (SIGTERM/SIGINT) while a periodic background loop with no stop flag was still running.
  • Pooled database connections in the gateway, controller, and worker are now validated before reuse (pool_pre_ping), so a Postgres restart no longer surfaces as a query error on the first stale connection. The controller's budget and scale-out background loops now survive a single failed pass instead of dying silently and permanently.
  • The gateway's shutdown now gives every background task a bounded grace period to stop cooperatively, then force-cancels whatever is still pending -- a loop added later, or an adapter whose stop does not actually work, can no longer bring back the earlier signal-hang.
  • The worker's shutdown now gives every background task a bounded grace period to stop cooperatively, then force-cancels whatever is still pending, matching the gateway's own fix.
  • The Nevia forward relay no longer resets a connection whose request grows past roughly 12 KB in one local read -- send_h2_data now splits any outbound Connect frame across as many HTTP/2 DATA frames as its negotiated max frame size and flow-control window allow, instead of a single unchunked call that h2 rejected outright.
  • NeviaProvider.destroy() now converges an agent's placement row to DESTROYED even when its idempotent no-op path finds the computer already gone (a harness/eval script's own placement-less cleanup previously left the row stuck at whatever it said before, so a later TeamService.archive found nothing to release).
  • A dedicated team's delegation follower no longer stays stuck on the shared NATS account when its own connection opens, closes, or reconnects after the follower already started -- role updates from a dedicated team no longer go silently missing from the thread.
  • a rework round now pushes to and reuses the same branch and pull request as the round it's reworking, instead of opening a duplicate
  • /budget no longer silently drops an agent whose LiteLLM key is blocked and no longer logs a stack trace on every call for it -- the agent is shown with its last known spend/budget and a blocked marker. A deleted key's agent is still left out, but with one info-level log line instead of a full traceback.
  • Skill staging under the fake sandbox provider's real local agent-server no longer attempts sudo mkdir against the host filesystem; the target itself now declares where to stage and that sudo is unavailable there.
  • a role could be given two delegations for the same work item at once -- most commonly the lead re-delegating to coder while reviewer's own rework hop to coder was still running -- each minting its own branch and pull request; a role now refuses (or, for the deterministic pipeline hop, waits for) a second delegation while its current task for that work item is still open
  • a worker's shutdown ignored a stop signal for as long as a run in flight took (a real incident: over two minutes, its agent-server still calling the model the whole time) -- it now interrupts every run in flight first (no terminal state written, left for the reconciler to redispatch), waits out the rest of its grace period, and force-exits if anything is still stuck instead of risking an indefinite hang; stack down also now reports when it had to escalate a service to SIGKILL instead of looking identical to a clean stop
  • LiteLLM upgraded to v1.103.0 -- a key's rate limit could stay stuck below its real budget for minutes without a Redis-compatible store (backlog task c96850b4).
  • The worker and the gateway exit on SIGTERM/SIGINT within a fraction of a second instead of hanging until they are killed: subscriptions and follower tasks are ended before the NATS connection is drained (works around a nats-py 2.15.0 defect that makes such a drain wait forever), the signal handler acts once, every shutdown cleanup step is named, timed and bounded, and a watchdog thread ends the process if the orderly shutdown does not finish.
  • Worker: a clean stop now releases the reconciler's own leader lock immediately, instead of the next worker having to wait out however much of the lock's TTL was left on the old key.
  • The gateway no longer follows an archived team's delegation-follower NATS consumer forever -- it stops at the next refresh pass after archive, and a person still writing in that team's room is told it's archived instead of being met with silence.
  • A delegation follower that hits an unexpected error now logs it, counts it, and retries with a growing backoff instead of dying silently and never being replaced -- the gateway's own periodic refresh pass now checks each core's real state directly instead of a separate copy that could drift from it.
  • vmd refuses to start, and a wake fails with a clear one-sentence reason, when one of its own nftables chains is declared differently than expected on the host instead of failing unpredictably later
  • A lead re-delegating to the same role within one work item (not a rework round) minted a second branch and opened a second pull request instead of reusing the first — the real cause a person saw on the first live team run. send_task now hands the work item's already-established branch on a later delegation, a delegate's own reply to its caller names the branch/PR the platform already has (not left to the model's prose), and a producing role whose handed branch is confirmed gone from the remote (merged and deleted) starts fresh at once instead of waiting out a 120s retry meant for an unrelated race.
  • the worker's own LLM setup no longer makes a synchronous HTTP call on the event loop, which could stall every other task on the worker (leases, heartbeats, other agents' runs) for the duration of that call
  • On Mattermost the platform's own usage lines name the command as it is typed there.
  • a lean-runtime role (PydanticAI/Deep Agents) is now asked to stop, the same way an OpenHands run already was, when the worker shuts down -- its own run in flight is interrupted and writes no terminal state, instead of running unattended through the shutdown grace period
  • a worker reload (or any graceful shutdown) while a task is running no longer loses that task in silence -- the task's lease is expired in place, not deleted, so the next reconciler poll redispatches it within one poll interval instead of never (the lease used to be released outright, making it invisible to the reconciler's own expired-lease sweep for good)
  • Worker events on A2A_TASKS now carry factory.hop_from_task/delegated_by_task/reply_to_task (whichever the task's current run carries), so a hop or delegation announcement in the room actually closes instead of staying "still waiting" forever.
  • prune-teams now points at the dev vmd's own socket, its summary reports an agent whose destroy silently failed instead of claiming a clean archive, and a new repair-stale-agents command reconciles those rows afterward
  • A reply to the room is no longer lost when two event paths of the gateway see the first event of the same task at the same moment (the second one failed on a duplicate row and stopped handling the task's events).
  • Two vmd instances in one process (multihost/migrate tests, and any future multi-host deployment) each resolve their own configured directories correctly again -- a shared package-level cache could leak one instance's configuration into another's
  • vmd refuses a -kernel outside every configured directory (images_dir/disks_dir/jail_base_dir/snapshot_dir) at startup, with a clear error, instead of failing 11+ seconds into a boot attempt
  • A boot failure at any point (not only after the jailer itself started) now releases the agent's network slot and its rules -- a failure early in Wake used to leave a tap device up until the next vmd restart
  • A /team create that fails partway through no longer leaves a team_roles/team_members/surface_links row behind, and a LiteLLM team it already created is cleaned up again -- a home channel it already created on Zulip/Slack/Mattermost stays, named in the error, since removing it automatically was judged riskier than leaving it
  • /team create now gives a readable, actionable error naming the real cause (an archived team's leftover role rows) when it hits the exact collision the archived-name-reuse migration exists to prevent, instead of "internal error handling command"
  • A team can be created again under the name of a team that was archived before archived names became reusable: a one-time cleanup removes the old team's leftover role, room-link and member rows and moves its memory entries out of the name's reach.
  • Gateway: an archived team's delegation-follower NATS consumer no longer gets recreated by a failed-then-retrying follower shortly after /team archive deletes it.
  • Controller RPCs (ListTeamRoles, GetTeamRole, EnsureAgent, MintGitHubToken, RotateAgentLiteLLMKey, GetDedicatedNatsCredentials, GetTeamControlNatsCredentials) no longer crash with MultipleResultsFound when an archived team and a newly created active team share one name; each now resolves to the active team, matching the team-directory rule.
  • Gateway: reused team names (an archived team's own name taken by a new active one) no longer crash dispatch or mis-follow the wrong team's delegation consumer.
  • team create --home zulip now adopts an existing, same-named Zulip channel (created by a previously archived team whose room outlived it) instead of failing with "Channel already exists".
  • A team created without --repo no longer accepts work it cannot do: /team create refuses templates that need a repository without --repo , and a task on an existing repo-less team fails with one sentence naming the create form.
  • The observability profile's Prometheus listens on 9091, not 9090, so it no longer blocks agent boots by taking the local credential gateway's metrics port.
  • A team that reuses an archived team's name no longer inherits its repository, model or sandbox provider from the worker's cached role config.
  • A task redispatched after a worker restart that interrupted its setup no longer fails with a 500 from a half-built conversation; the conversation is recreated.
  • /cancel in a work item now also cancels the tasks its lead delegated (coder, reviewer); before, they kept running, pushed a branch and woke the canceled lead.
  • A task settled after its worker died (canceled or failed) no longer leaves its agent-server calling the model until the run timeout: the run is interrupted when the replica is awake.
  • The Nevia provider finds an agent's computer by the id it recorded at ensure (new agents.sandbox_id column, migration 9e4b7a1c5d20), so a wake right after creation no longer fails on Nevia's lagging name search.
  • The Nevia control plane's credgwd hands agents named LLM/MCP upstream hosts (llm.factory.internal, mcp.factory.internal) instead of 127.0.0.1, which the agent computer's NO_PROXY excluded from the relay; stage 1 of the Nevia proof passes there
  • The role that delegated a task is now told when that task is canceled explicitly; a work item canceled at its lead still sends its delegates' cancellations no notice.
  • /team create without --template uses the template FACTORY_DEFAULT_TEMPLATE names (feature-team unless set) instead of failing on a template called default; a name that matches no template answers with the list of the ones that exist.
  • A coder task that fails after pushing a safety commit no longer leaves the work item with two branches when the lead re-delegates, and the re-delegated task that adds no commit of its own still hands the work to the reviewer.
  • A task redispatched after a Nevia worker crash no longer fails with 401 from its computer's agent-server: the relaunched worker adopts the keys the computer was bootstrapped with.
  • After a gateway restart the room now gets the events of a running task that were published while the gateway was down, once each (per-task event cursor, migration e5a1c9d47b30 needed).
  • Starting the real agent-server from the fake provider no longer risks blocking the worker's event loop: the spawn uses a launcher that ties the child to the worker instead of preexec_fn.
  • The worker reads the LLM gateway's /v1/model/info with an explicit timeout (FACTORY_LLM_MODEL_INFO_TIMEOUT_SECONDS, 10 s); a gateway that does not answer, is unreachable or answers a server error now fails the task with that answer in the text instead of running on without model info.
  • The gateway's delegation follower no longer loses an update it was handling when it stopped, retries a failing update three times with a growing delay, and then skips it with one line in the room naming the role.
  • A run whose guard tripped (unreachable or stalled guest) no longer leaves the SDK's status poll thread polling the dead guest: the worker closes its HTTP client.
  • Zulip's uploaded files (avatars, attachments) now live in the named volume zulip-data, so recreating the zulip container no longer loses them (files already in an old anonymous volume are not carried over).
  • A task re-dispatched after the guest's orphan watchdog paused its conversation now resumes the run instead of pushing the work so far and reporting it done.
  • /team archive and just prune-teams name the cleanup steps that failed (sandbox destroy, key and secret release, consumers, team config) instead of reporting a clean archive; prune-teams exits non-zero
  • A question raised by a delegated role (a reviewer's approval request, for example) or after a gateway restart is addressed to the person by their name on the surface; it read @<person id> (@2) in Mattermost before
  • A reply in a thread that names no role now answers the delegated role whose question is open (a reviewer's approval request) when the lead has nothing pending; before, a bare yes went to the lead and the reviewer stayed paused until @reviewer yes
  • A reply addressed to a role (@reviewer approve) reaches the role without its own mention, so a risky-action approval is recognised at once instead of being asked for again
  • A long tool output in an agent's history no longer fails the run with Permission denied: /data/workspace on the control plane: the worker's history check decides the last message's role without rendering observations
  • A coder that ends its turn with a chat message instead of an action or a finish call is told to continue, twice at most, before its work is pushed and reported; a clone that fails inside the agent computer is retried three times; a failure in a Mattermost thread is one card, written at once, not two posts
  • A delegate's failure reply no longer joins a lead task that is finishing: a reply resumes a task only while it is paused waiting for it, otherwise it lands as a fresh task in the same context, so the room gets the lead's answer.
  • A role's confirmation policy now follows the team's own setting (template or /team set) instead of only the role file's; a role-file edit reaches a running worker within 30 seconds, no restart
  • The gateway no longer posts 'Still waiting on a reply' for a reply that landed before the end it answers (and no longer announces a work item's closing a second time for a late event).
  • Agents can clone a repository with many refs again: the credential gateway lets git's gzipped smart-HTTP request bodies (Content-Encoding) through to GitHub instead of stripping the header and getting a 400 on git-upload-pack
  • The worker no longer holds its shutdown grace period when the agent-server is unreachable: the SDK's status poll thread is closed after a failed interrupt.
  • Agent trace export no longer loses its compressed batches: the credential gateway lets the OTLP exporter's Content-Encoding header through to the collector (half the exports got a 400 before)
  • A second review round no longer fails with "Could not bring the delegated branch into the worktree" when the reviewer's own test run touched files that the rework changes or deletes (committed bytecode, for example).
  • A work item no longer stays open when the reviewer's last act is its own report to the lead: the gateway no longer waits for a reply the worker never sends.
  • A reply that lands on a planner or lead conversation no longer fails the task with "refusing to call the model -- the history ... ends with an assistant message" when the conversation's local event mirror had not yet received the message just sent.
  • The closing post of a work item is now the last post in a Mattermost thread, and the lead's final text appears once instead of twice (a card under the closing).
  • A failed, canceled or rejected rework task's report now reaches the lead (not only the reviewer that delegated it), so a work item no longer stalls when a rework coder task fails.
  • A failed delegate's report now says when nothing has been pushed for the work item, and the lead names a branch to continue on only if the note names one.
  • A task that committed nothing no longer pushes a branch equal to the default branch (no stray branch, no pull request attempt for it, no "pushed branch" claim).
  • A coder that still ends in an unfinished chat message after its reminders is no longer reported as done as if it had finished: the report says so (and the task fails when it left no commit).
  • The lead can no longer delegate to the coder again while the reviewer is still reviewing the coder's change; a person's change request that reaches the lead during a review is refused until the review reports.
  • A coder's completion that was handed to the reviewer now reaches the lead as a status notice, and the lead acknowledges it in one line instead of delegating again or asking the person.
  • The guard that refuses to call the model on a history ending with an assistant message no longer fails a healthy reply task when its event mirror received other events before the reply's echo (the earlier fix waited on the mirror growing, which was the wrong signal); it now waits for the sent message itself and reads the server's history at the ceiling.
  • A task stopped by an exhausted team budget now says so in a plain sentence (an admin can raise it with /budget set), and the lead reports it once instead of re-delegating.
  • A raised /team set ... max_iterations (or a changed /team model) now applies to the next task of a running work item: its conversation is started afresh when the stored limit or model differs, and the task says so.
  • A late completion notice from a delegate no longer reopens a work item that already closed, so the thread gets one closing post instead of two.
  • /team create with a room surface the controller has no credentials for no longer leaves the team's database rows behind (a retry answered "name taken"): it now refuses before creating anything and names the setting that is missing.
  • A team linked to two chat or code surfaces is followed by one delegation follower, and each delegated role's update reaches the room of its own work item, not whichever surface's follower pulled it first.
  • The onboarding prompt named http:// behind the edge, which terminates TLS: without KAPELLE_PUBLIC_BASE_URL it now builds its links from X-Forwarded-Proto and X-Forwarded-Host, trusted only from loopback and private addresses
  • A run's conversation no longer shows a role's status and completion posts twice after a reply to a question.
  • Console: a run's cost shows up when its team is archived right after the run (the cost pass now includes an archived team with a run from the last 24 hours).

Security

  • No credential ever enters an agent VM: the credential gateway injects GitHub/model/registry credentials on egress; MCP tokens are minted and injected gateway-side, never stored in the guest.
  • Per-agent data-disk and snapshot encryption keys come from OpenBao, keyed identically by vmd (Go), credgw (Go) and the controller (Python) — a cross-language golden fixture (contracts/openbao/agent_identity_paths.json) pins the shared path convention so the three can't silently diverge again.
  • vmd's gRPC socket authenticates callers via SO_PEERCRED (uid/gid), not file permissions alone.
  • credgwd and every Python service (controller, gateway, worker) redact known secret shapes (LiteLLM keys, GitHub tokens, OpenBao/Vault tokens) and the actual values they mint or read (management API keys, AppRole role/secret ids) out of every log line, at the logging boundary — not per call site.
  • The dev host's sudoers grant allowed bare /usr/bin/tee, /usr/bin/cp and /usr/bin/chown with no argument restriction at all — effectively unrestricted root for anything running as vmd's own user, since any of the three can overwrite an arbitrary file. vmdctl sudoers renders the minimal replacement (deploy/host/sudoers.factory, checked in) from internal/host.PrivilegedCommands, the exhaustive list of every command+argument shape vmd's own code actually issues.
  • GitHub's App private key/webhook secret/minted installation tokens and Zulip's bot/test-user API keys are now registered with the redaction layer at the point each is minted or configured, closing a gap where only shape-based patterns protected them
  • credgw now renders each team's github-api destination scoped to that team's own resolved repo (not a generic any-repo glob), and drops the destination entirely for a team with no resolvable repo, enforced by iron-proxy itself
  • Dedicated-team NATS isolation now covers the gateway too: GatewayA2AClient and ControllerCommandClient route a dedicated team's A2A traffic and /team, /budget commands through that team's own NATS account instead of the gateway's shared connection, reached via the controller's GetTeamControlNatsCredentials RPC
  • vmd's jailer starts now route through vmd-priv (backlog task 9cb46684) instead of a separately inlined sudo -n, closing the last privileged call the helper hadn't replaced -- a production host no longer needs any sudo installed at all once vmd-priv is deployed.
  • Dedicated-team isolation now covers Agent Card registration and lookup too: a dedicated team's own account gets its own a2a-cards KV bucket, registered by the controller alongside the shared one and read by the gateway's per-team routing
  • Team skills: a per-run injection cap on triggered skill bodies, and a guard so a skill's allowed_tools can never widen past its role's own tool list (design doc §15).
  • a sudoers placeholder value substituted into an ArgsRegex-shaped grant is now rendered safely -- a glob wildcard becomes a bounded one-path-segment class, every other character is escaped as a literal, and a value that can't be rendered safely refuses the render instead of widening the grant
  • vmd's sudoers grant for cp/chown's own jailer-chroot entries is bounded to the real argument shape (an ArgsRegex, not a bare glob) instead of admitting anything under the configured directory
  • the sudoers filename class for a jailer-chroot file could match a segment of only dots, admitting the parent directory; three cp/chown grant entries with no real caller anywhere in the repo are removed instead of narrowed, and a test now fails for any future entry with no matching call site
  • the sudoers grant for rm/mkdir/jailer now matches vmd's real argument shape instead of a bare wildcard
  • vmd's PrepareWritableInChroot (the first of jailer.go's own privileged file operations to move) is confined to its own configured directory tree by the kernel, closing the symlink-following window every argument-string-based cp/chown call left open
  • FinishWritableInChroot (jailer.go) is confined to its own configured directory tree by the kernel, closing the symlink-following window the old direct chown exec left open
  • MakeReadableInChroot (jailer.go) is confined to its own configured directory tree by the kernel, closing the symlink-following window the old direct chown exec left open
  • StartJail's own stale-chroot removal (jailer.go) is confined to its own configured directory tree by the kernel, closing the symlink-following window the old direct rm exec left open
  • KillJail's own chroot removal (jailer.go) is confined to its own configured directory tree by the kernel, closing the symlink-following window the old direct rm exec left open
  • CopyIntoChroot, LinkOrCopyIntoChroot, and CopyOutOfChroot (jailer.go) are confined to their own configured directory trees by the kernel on both sides of the copy, closing the symlink-following window the old direct cp/chown execs left open
  • vmd (backlog task 7ec3413f-603d-414a-aa17-5c08967f2923): jail_base_dir must now be provisioned root:root mode 0755, checked at startup with no skip flag; console-logs moved out to their own vmd-owned directory so nothing needs to write under it unprivileged anymore; a just doctor check reports the same. The check also walks jail_base_dir's own ancestor chain up to "/", refusing any ancestor that is neither root-owned-and-not-group/other-writable nor sticky -- a root-owned leaf alone doesn't stop its non-root-owned parent's own owner from replacing an entry with a symlink before jailer resolves the path. The dev stack's own jail_base_dir moves out from under the user-owned /var/tmp/vmd-dev (which now fails that walk) to the sibling /var/tmp/vmd-dev-jail, whose parent is the root-owned, sticky /var/tmp.