homelab-forge

← All tasks

TASK-011

Factory control-plane hub, CI watch, and agent observability

review

Goal

Make the forge-site control plane the coordinator between Slack and Cursor SDK agents, with inspectable runs, one stable PR per task, sandbox-aware agents, local lint plus CI-watch until green, and concurrent tasks. Failures flow LLM → control plane → Slack. Slack intake only records intent; it does not run the LLM itself.

Acceptance criteria

  • Cursor SDK plan and implement runs persist a redacted conversation transcript (assistant + tool calls/results, no secrets) in Postgres, keyed to the task and to a durable run/job id; GET API returns that transcript
  • forge-site task detail (/dashboard/[id]) lists runs/jobs for the task and lets the operator inspect the SDK transcript for a selected run
  • Slack plan iteration reuses the first plan PR and pinned tasks.branch; gh pr create must not open a second PR when slack_threads.pr_url or a pr artifact already exists (regression of TASK-009 #16 then #19 and TASK-010 #18 then #20)
  • Planner YAML must not rewrite branch after the first worktree/PR; worker and orchestrator always push to the DB-pinned branch
  • After approve, Slack thread replies are stored on the task via the control plane and are not treated as a fresh plan-PR cycle; implementation requests do not spawn a host planner that tries to sudo/apt/implement
  • Orchestrator and worker prompts include an explicit runtime card covering role (plan-only vs implement), sandbox_profile, no TTY, no host sudo/apt, and ADR-002 restrictions; failed shell/SDK/sudo attempts are reported to the control plane and posted to the Slack thread by the control plane
  • systemd journal for orchestrator and worker logs each Slack/API action, task id, pinned branch, PR URL, SDK run id, subprocess command+exit (redacted), and failures; runbook documents journalctl plus artifact log paths
  • Before every git push that updates a factory PR, the worker (and planner when it commits markdown/YAML) runs the same in-repo linters CI runs, at least markdownlint-cli2 per .markdownlint-cli2.yaml plus any other local checks documented in ci.yml that can run without GitHub-hosted secrets; lint failures are fixed in the same run, not left for the operator to notice on GitHub
  • After a plan or implement push, the control plane (or a claimed watch job it dispatches) polls GitHub checks on that PR until they are green or budget/timeout expires; on failure it enqueues a fix run against the same branch/PR, then re-watches; the operator is not required to paste "CI is failing" into Slack; Slack is notified only when green, when retries are exhausted, or when human merge/review is the next step
  • Slack Socket Mode is a thin intake client: slash and thread events POST to the control plane API; the control plane creates/updates the task, enqueues plan/implement/watch/notify jobs, and is the only path that posts agent progress and failures back to Slack (ADR-010 notify queue); host slack_intake.py must not call Cursor SDK, git, or gh directly
  • Multiple tasks run at the same time: control plane allows concurrent in_progress tasks; host workers claim and execute more than one job (separate worktrees, documented FORGE_WORKER_CONCURRENCY); worker PLAYBOOK "one task at a time" is updated; two Slack /forge plan requests must not serialize behind a single SDK process
  • ADR-010 is amended (or a short follow-on ADR) so forge-site is the communication bridge; Slack bot token for outbound posts comes from Vault via ExternalSecret, never git
  • CI green on this implementation PR; factory/review/CHECKLIST.md completed; merge remains human-gated (ADR-008)
Assignee
cursor-diestrin-interactive
Branch
factory/task-011-factory-agent-observability-pr-identity
Sandbox
agent-cell
Risk
high

Artifacts

Agent runs

No agent runs recorded yet.

Message history