@windyroad/itil 2.5.1-preview.1273 → 3.0.0-preview.1296

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -486,5 +486,5 @@
486
486
  }
487
487
  },
488
488
  "name": "wr-itil",
489
- "version": "2.5.1"
489
+ "version": "3.0.0"
490
490
  }
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "wr-itil",
3
- "version": "2.5.1",
3
+ "version": "3.0.0",
4
4
  "description": "ITIL problem-management workflows for AI coding agents",
5
5
  "author": {
6
6
  "name": "Windy Road Technology",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@windyroad/itil",
3
- "version": "2.5.1-preview.1273",
3
+ "version": "3.0.0-preview.1296",
4
4
  "description": "ITIL-aligned IT service management for Claude Code and Codex",
5
5
  "bin": {
6
6
  "windyroad-itil": "./bin/install.mjs"
@@ -170,18 +170,7 @@ What "work" means depends on the problem's status:
170
170
 
171
171
  **Known Error (root cause identified AND workaround documented; ready for fix proposal):**
172
172
 
173
- **Substance-confirm-before-build guard (the ": Confirm a decision's substance before building dependent work on it" architecture rule — propose-fix surface, the "Problem-RFC-Story framework with mandatory problem-trace and unified problem ontology" architecture rule I13).** BEFORE implementing any fix step below, check whether the fix builds on a genuine decision whose **substance is unconfirmed**. This closes the "Agent implements dependent work on genuine new decisions before human-confirming their SUBSTANCE — surfaces only meta-questions" problem (dependent work built on a born-`proposed` decision the user later rejects):
174
-
175
- 1. Collect the decisions this fix builds on: the `ADR-NNN` references in the problem's `## Fix Strategy` section PLUS the `adrs:` frontmatter array (and body `ADR-NNN` mentions) of each referenced RFC.
176
- 2. For each, run the predicate (PATH shim per the "Plugin-bundled scripts invoked from SKILL.md resolve via `bin/` on `$PATH`" architecture rule — adopter-safe, never source repo-relative lib files):
177
- ```bash
178
- wr-architect-is-decision-unconfirmed ADR-<NNN> docs/decisions
179
- ```
180
- Exit 0 = unconfirmed (born without `human-oversight: confirmed`, not superseded) → the guard fires for that ADR. Exit 1 = confirmed or superseded → OK to build. Exit 2 = not found → treat as not-a-blocker (surface in the report).
181
- 3. **If any referenced decision is unconfirmed**: do NOT implement yet. Surface its **substance** (the chosen option the ADR records — not a grain/meta question) for human confirmation:
182
- - **Interactive**: `AskUserQuestion` presenting the ADR's Decision Outcome + Considered Options so the user confirms / amends / rejects the substantive choice. On confirm, the recording skill writes the `human-oversight: confirmed` marker (the ": Human-oversight marker + `/wr-architect:review-decisions` drain for recorded decisions" architecture rule) before the build proceeds.
183
- - **AFK** (`/wr-itil:work-problems` orchestrator): NEVER ask mid-loop — queue the substance to the iteration's `outstanding_questions` (the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule AFK carve-out) and skip the build; do not guess. This is the **queue-and-continue** universal default per the "Structured User Interaction for Governance-Skill Decisions" architecture rule Rule 6 (the "AFK iter default when a skill needs to ask a question and AskUserQuestion is unavailable — should queue the question and move to the next iteration (not halt, not silently skip)" problem, 2026-06-06 amendment): the iter queues the substance + advances; the orchestrator main turn surfaces the queued question at loop end via the Step 2.5 batched AskUserQuestion.
184
- 4. This ask is **the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule category-1 direction-setting** and is EXCLUDED from the lazy-AskUserQuestion regression metric (it is legitimate, not lazy). The trigger is narrow — detection is mechanical (the predicate); only genuine unconfirmed decisions about to be built on fire it. Do NOT over-fire on confirmed/superseded/obvious decisions (inverse-the "Problem 078: Assistant does not offer to capture a problem ticket when the user delivers strong-signal correction" problem / the "Agents over-ask in interactive sessions — conflating mechanical-stages with user-interactive-stages of multi-stage skill contracts (inverse-)" problem guard). A born-`proposed` marker is fine for *recording*; it is not a licence to *build* (the ": Human-oversight marker + `/wr-architect:review-decisions` drain for recorded decisions" architecture rule carve-out).
173
+ **ADR proposal implementation and acceptance.** A fix may build against a documented proposed ADR before human ratification. Review its substance and preserve the existing story-map, job, risk, and release gates. If substantive direction is genuinely ambiguous, ask or queue that choice; do not halt solely because `human-oversight: unconfirmed`. Gather evidence of successful actual production use before requesting final ratification. The ADR remains proposed until both real production evidence and explicit human ratification of the final substance exist. Local tests and CI alone do not satisfy acceptance.
185
174
 
186
175
  **I13 propose-fix trace gate (the ": RFC-first trace invariant not enforced at fix-time" release design B3/B4).** BEFORE the traversal below, enforce the fix-time trace invariant: a fix proposed on a Known Error requires a fix vehicle that traces the problem (the "Every fix goes through an RFC" architecture rule unconditional; the "RFC required at the propose-fix step on a Known Error" architecture rule places the gate here, conforming to the "Problem lifecycle — add a Verification Pending status between Known Error and Closed" architecture rule Known Error semantics — the fix is proposed *after* Known Error).
187
176
 
@@ -728,7 +728,7 @@ rm -f "$ITER_JSON" "$ITERATION_PROMPT_FILE"
728
728
 
729
729
  1. **Context (the unattended declaration — load-bearing carrier per the "Per-ticket goal anchors each AFK iteration" architecture rule)**: this is one iteration of the AFK work-problems loop. <!-- UNATTENDED-DECLARATION-SOURCE --> **The user is AFK and this run is unattended.** The orchestrator selected `P<NNN> (<title>)` as the highest-WSJF actionable ticket. This sentence is the **discriminator** the singular skill keys its pinned short-circuit on (`/wr-itil:work-problem` Step 1): prose cannot observe a TTY and a `claude -p` subprocess offers nothing to sniff, so the unattended state is **declared by the dispatcher**, never detected by the skill. Do not drop or reword the declaration — a rule keyed on something the skill cannot observe silently collapses to always or never.
730
730
  2. **Task**: run `/wr-itil:work-problem P<NNN>` — the singular skill, pinned to the ticket the orchestrator already selected. It is the loop's per-iteration execution unit (both SKILLs have long documented this; the "Per-ticket goal anchors each AFK iteration" architecture rule makes the dispatch match). It delegates the actual work to `/wr-itil:manage-problem <NNN>`, so the manage-problem workflow still runs verbatim — architect / jtbd / style-guide / voice-tone gate reviews and the commit gate (manage-problem Step 11) all apply. Because this subprocess has the Agent tool in its own surface, the normal review-via-subagent paths work — no inline-verdict fallback needed. Because the dispatch is pinned AND declares the run unattended, the singular skill skips its ranking-freshness check (it must not delegate a README refresh, must not prompt, must not commit a ranking rewrite inside this per-ticket unit of work).
731
- 3. **Constraints**: commit the completed work per the "Governance Skills Commit Their Own Completed Work" architecture rule. Do NOT push, do NOT run `push:watch`, do NOT run `release:watch` — the orchestrator's Step 6.5 owns release cadence. Do NOT invoke `capture-*` background skills mid-iter (AFK carve-out — the "Governance skill invocation patterns — foreground + background with deferred-question resumption" architecture rule), **EXCEPT** (a) **retro-surfaced observations of recurring class-of-behaviour** — those route to `/wr-itil:capture-problem` per the **the "Iter retros queue their own observations as `outstanding-questions.jsonl` entries for user-direction triage instead of auto-ticketing — same trust-boundary as `/wr-retrospective:run-retro` Step 4a" problem mechanical-stage carve-out** (see retro-on-exit constraint #4 below; same trust-boundary as `/wr-retrospective:run-retro` Step 4a verification close-on-evidence — the "Iter retros queue their own observations as `outstanding-questions.jsonl` entries for user-direction triage instead of auto-ticketing — same trust-boundary as `/wr-retrospective:run-retro` Step 4a" problem); and (b) **the I13 fix-time row draw** — when the propose-fix gate inside the delegated `/wr-itil:manage-problem` traversal detects a Known Error nothing yet proposes a fix for (`wr-itil-check-fix-rfc-trace` emits a `no-rfc-trace:` directive), the iter **draws a release row on a story map that already covers the journey**, gives it at least one story card, and makes that card's story name the problem in its own `problems:` list — then proceeds. **A fix proposal is a release row; it is never a new document under `docs/rfcs/`.** Take the identity from the directive, which comes from `wr-itil-next-rfc-id` — the single rule that sees rows, documents and git history at once, and the only one that will not re-issue an identity a row already holds. **UNLESS** an existing vehicle cited in the ticket is already this ticket's fix and merely lacks the trace edge, in which case the iter **wires** that edge — a card on the existing row, or the `problems:` array of a legacy document — rather than drawing a duplicate that fragments the fix across two vehicles (the "manage-problem I13 propose-fix gate auto-creates a new RFC instead of wiring an existing fix-vehicle's trace edge" problem; existing-vehicle-untraced sub-case; vehicle-vs-merely-related is a judgement read of citation context, structured-logged as `I13: wired P<NNN> trace edge into existing fix vehicle <ID>`; the load-bearing branch prose lives in the delegated `/wr-itil:manage-problem` I13 gate). This is NOT an aside-capture distraction: the row is the **mandatory vehicle for THIS iter’s own fix** (the "Every fix goes through an RFC" architecture rule), not a tangential observation — it is in-scope working of the current ticket, framework-mediated (NOT cat-1 direction-setting → NO `AskUserQuestion`, the "Agents over-ask in interactive sessions — conflating mechanical-stages with user-interactive-stages of multi-stage skill contracts (inverse-)" problem), and drawing a row onto a map a person has already approved inherits that approval rather than needing a fresh one. **Two things the iter must NOT do silently**: change an existing map in a way that would alter what its approval covers — a new activity column or a new job on the map’s traces, judged from `oversight_map_substance_keys()` in `lib/story-oversight.sh`, the one place those keys are enumerated — or pick a fix approach no existing decision record covers. Either of those queues ONE entry at `outstanding_questions` and the iter moves to the next problem; the loop is never stopped for it. The predicate can also refuse outright (exit 3), and the two refusals are handled differently: a map edited without being re-rendered is **mechanical** — re-render it with `wr-itil-render-story-map` and ask again, asking nobody — while a repository with no story maps at all triggers complete unconfirmed proposal authoring through `/wr-itil:capture-story-map` and `/wr-itil:capture-rfc`, without `AskUserQuestion`. The iter creates the initial activities, identified release row, cards, and story files; queues exactly one `outstanding_questions` item to ratify the completed map; blocks source and story implementation until ratification; and continues independent work. Structured-log the draw event to the iter summary (`notes`) per the ": Progress the Backlog While I'm Away" user outcome audit-trail. Do NOT use `ScheduleWakeup` under any circumstance (the "Problem 083: work-problems Step 5 iteration-worker prompt does not forbid ScheduleWakeup / time-deferring primitives — subagent can abandon synchronous-completion contract" problem — iteration workers must not self-reschedule). **NEVER call `AskUserQuestion` mid-loop in AFK** (the "Decision-delegation contract — agents over-apply Rule 1's interactive default to framework-resolved decisions; codify the framework-resolution boundary + AFK loop's batched-questions-as-deliverable + lazy-AskUserQuestion measurement" problem / the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule): direction / deviation-approval / one-time-override / silent-framework observations queue at `ITERATION_SUMMARY.outstanding_questions` for loop-end batched presentation. **This includes the manage-problem substance-confirm-before-build guard (the ": Confirm a decision's substance before building dependent work on it" architecture rule (Confirm a decision's substance before building dependent work)):** when the propose-fix step detects that the fix builds on a born-`proposed` decision whose substance is unconfirmed (via `wr-architect-is-decision-unconfirmed`), the iter does NOT implement on it and does NOT ask mid-loop — it queues a `category: "direction"` entry naming the unconfirmed ADR + its Decision Outcome for loop-end confirmation, and routes the ticket to `action: skipped`, `skip_reason_category: user-answerable`. Building on the unconfirmed substance instead (or guessing the choice) is the "Agent implements dependent work on genuine new decisions before human-confirming their SUBSTANCE — surfaces only meta-questions" problem failure this guard exists to prevent. The queued substance-confirm is a legitimate cat-1 direction ask — it is NOT counted as lazy in the Step 2d Ask Hygiene Pass (the ": Confirm a decision's substance before building dependent work on it" architecture rule lazy-count exclusion). Per-iter `AskUserQuestion` calls are sub-contracting framework-resolved decisions back to the user (lazy deferral per Step 2d Ask Hygiene Pass classification). Non-interactive defaults apply per the "Structured User Interaction for Governance-Skill Decisions" architecture rule Rule 6 + the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule's framework-resolution boundary. **Treat the user as transient** (the "`/wr-itil:work-problems` orchestrator defaults to subprocess dispatch even when the user is observably interactive — loses real-time presence advantage" problem): even when observably present at orchestrator dispatch time, the user may answer one question and disappear for hours; presence is not a reliable signal and is not the goal. The iter's job is to progress the ticket and accumulate questions for batched surfacing — not to ask "is it OK to proceed?" at a mechanical-stage boundary. **Do NOT poll `bats` output with a bats-console-summary regex against TAP-format output** (the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem — bash until-loop-deadlock antipattern). The bats-console-summary line `<N> tests, <M> failures` is emitted ONLY by bats's *default* (non-TAP) formatter; `bats --tap` does not emit a console summary, so a polling loop of shape `until [ -f $OUT ] && grep -qE '^[0-9]+ tests?,' $OUT; do sleep 5; done` spins forever after bats completes (silent deadlock — no error, no exit; recovery requires manual SIGTERM with metadata loss per the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem/the "SIGTERM-clean-flush guarantee is conditional on subprocess having emitted ITERATION_SUMMARY before going idle — needs SKILL.md caveat + behavioural-test second-source for stuck-before-emit subclass" problem stuck-before-emit subclass). When you need to wait on a backgrounded bats run, prefer `wait $bg_pid` (Unix idiom — completion signaled by process exit, no regex required) or, for the Bash tool, `run_in_background=true` + `BashOutput` polling on the tool's exit-state field rather than regex-poll on stdout. If you genuinely must regex-poll TAP output, anchor on the TAP plan line `^[0-9]+\.\.[0-9]+` (e.g. `1..1455`) — TAP's plan line is emitted on completion and is format-stable across bats versions; the bats-console-summary line is not. The console-summary vs TAP-format divergence is the load-bearing detail: `bats` and `bats --tap` produce structurally different stdout, and the antipattern assumes the former when iter dispatch typically uses the latter. **Do NOT poll subprocess completion with `pgrep -f '<pattern>'` inside an `until` / `while` loop** (the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem — self-referential pgrep deadlock; sibling variant of the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem). `pgrep -f` matches against the FULL command line of every running process, so the polling loop's own `zsh -c` argument (which contains the literal `pgrep -f '<pattern>'` text) matches itself; with multiple concurrent polling loops, each loop matches the others and spins forever. Worked example of the antipattern: `until ! pgrep -f 'bats --recursive' > /dev/null 2>&1; do sleep 5; done` — the 2026-05-16 the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem deadlock witness; 4 concurrent polling loops each matched the others' command lines while no actual bats process ran; 45 min wall-clock + $20-30 wasted before manual SIGTERM. The same self-reference shape applies to `while pgrep -f ...; do sleep; done` and to `until ! pkill -0 -f '<pattern>'` / `while pkill -0 -f '<pattern>'` (signal-0 polling). The structural fix is the same as the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem: prefer `wait $bg_pid` (Unix idiom — shell-native completion signal, no regex / no pgrep) or Bash-tool `run_in_background=true` + `BashOutput` polling (harness-tracked completion state). The hook `packages/itil/hooks/itil-bash-polling-antipattern-detect.sh` denies these shapes at PreToolUse:Bash, but the prompt rule belongs here too — structural enforcement + prompt discipline together close the class. **Do NOT leave a backgrounded task unreaped at turn-end** (`run_in_background: true` on an Agent or Bash tool call, or a `&`-detached shell job, whose completion you intend to observe in a *later* turn) inside iter dispatch contexts (the "Iter subprocess ends its turn waiting on a backgrounded task and never resumes — `claude -p` has no auto-resume; commit-bearing work is lost" problem — turn-end-mid-background work-loss; sibling-class to the "Problem 083: work-problems Step 5 iteration-worker prompt does not forbid ScheduleWakeup / time-deferring primitives — subagent can abandon synchronous-completion contract" problem / the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem / the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem). The iter subprocess is dispatched via `claude -p`, a single-shot CLI invocation with NO auto-resume affordance: its turn boundary IS its process boundary. A background task that outlives the turn never resumes — the iter exits at turn-end with the task incomplete and its own work staged but uncommitted (witnessed: iter 11 of a prior loop — $8.02 / 17 min / 8 staged files / 11 GREEN bats / ZERO commits; recovery required orchestrator main-turn salvage). **The prohibition is on the cross-turn / turn-end-survivor shape, NOT on backgrounding per se:** the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem/the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem-sanctioned idiom of launching `run_in_background=true` + `BashOutput`-poll-then-`wait $bg_pid` (or plain `wait $bg_pid` on a `&` job) **within the same turn** is fine — it reaps the task before turn-end. Use foreground-synchronous invocation instead: the Agent tool WITHOUT `run_in_background: true` (the result returns in-turn, so the commit step is reached), or intra-turn background that you `wait` on before the turn closes. The distinction from the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem/the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem polling antipatterns: those forbid *how* you wait (regex / pgrep poll loops); this forbids *deferring a task's completion past the turn boundary*, where `claude -p` has no notification re-entry to bring you back. The interactive Claude Code session masks this hazard (notification-driven re-entry); the AFK iter subprocess does not. **Release-preparation boundary (the ": Commit-time changeset enforcement forces premature release metadata" problem):** ordinary implementation commits MUST NOT include changesets and remain `STAGED`. After the exact implementation pipeline passes and release preparation is intentional, the orchestrator inspects the cumulative package scope and creates one complete changeset-only commit. Never create speculative, placeholder, dormant-foundation, or “for later” changesets. **`@jtbd the ": Progress the Backlog While I'm Away" user outcome`** (load-bearing) **`@jtbd the ": Keep Plugins Current Across Projects" user outcome`** (closure-dependent).
731
+ 3. **Constraints**: commit the completed work per the "Governance Skills Commit Their Own Completed Work" architecture rule. Do NOT push, do NOT run `push:watch`, do NOT run `release:watch` — the orchestrator's Step 6.5 owns release cadence. Do NOT invoke `capture-*` background skills mid-iter (AFK carve-out — the "Governance skill invocation patterns — foreground + background with deferred-question resumption" architecture rule), **EXCEPT** (a) **retro-surfaced observations of recurring class-of-behaviour** — those route to `/wr-itil:capture-problem` per the **the "Iter retros queue their own observations as `outstanding-questions.jsonl` entries for user-direction triage instead of auto-ticketing — same trust-boundary as `/wr-retrospective:run-retro` Step 4a" problem mechanical-stage carve-out** (see retro-on-exit constraint #4 below; same trust-boundary as `/wr-retrospective:run-retro` Step 4a verification close-on-evidence — the "Iter retros queue their own observations as `outstanding-questions.jsonl` entries for user-direction triage instead of auto-ticketing — same trust-boundary as `/wr-retrospective:run-retro` Step 4a" problem); and (b) **the I13 fix-time row draw** — when the propose-fix gate inside the delegated `/wr-itil:manage-problem` traversal detects a Known Error nothing yet proposes a fix for (`wr-itil-check-fix-rfc-trace` emits a `no-rfc-trace:` directive), the iter **draws a release row on a story map that already covers the journey**, gives it at least one story card, and makes that card's story name the problem in its own `problems:` list — then proceeds. **A fix proposal is a release row; it is never a new document under `docs/rfcs/`.** Take the identity from the directive, which comes from `wr-itil-next-rfc-id` — the single rule that sees rows, documents and git history at once, and the only one that will not re-issue an identity a row already holds. **UNLESS** an existing vehicle cited in the ticket is already this ticket's fix and merely lacks the trace edge, in which case the iter **wires** that edge — a card on the existing row, or the `problems:` array of a legacy document — rather than drawing a duplicate that fragments the fix across two vehicles (the "manage-problem I13 propose-fix gate auto-creates a new RFC instead of wiring an existing fix-vehicle's trace edge" problem; existing-vehicle-untraced sub-case; vehicle-vs-merely-related is a judgement read of citation context, structured-logged as `I13: wired P<NNN> trace edge into existing fix vehicle <ID>`; the load-bearing branch prose lives in the delegated `/wr-itil:manage-problem` I13 gate). This is NOT an aside-capture distraction: the row is the **mandatory vehicle for THIS iter’s own fix** (the "Every fix goes through an RFC" architecture rule), not a tangential observation — it is in-scope working of the current ticket, framework-mediated (NOT cat-1 direction-setting → NO `AskUserQuestion`, the "Agents over-ask in interactive sessions — conflating mechanical-stages with user-interactive-stages of multi-stage skill contracts (inverse-)" problem), and drawing a row onto a map a person has already approved inherits that approval rather than needing a fresh one. **Two things the iter must NOT do silently**: change an existing map in a way that would alter what its approval covers — a new activity column or a new job on the map’s traces, judged from `oversight_map_substance_keys()` in `lib/story-oversight.sh`, the one place those keys are enumerated — or pick a fix approach no existing decision record covers. Either of those queues ONE entry at `outstanding_questions` and the iter moves to the next problem; the loop is never stopped for it. The predicate can also refuse outright (exit 3), and the two refusals are handled differently: a map edited without being re-rendered is **mechanical** — re-render it with `wr-itil-render-story-map` and ask again, asking nobody — while a repository with no story maps at all triggers complete unconfirmed proposal authoring through `/wr-itil:capture-story-map` and `/wr-itil:capture-rfc`, without `AskUserQuestion`. The iter creates the initial activities, identified release row, cards, and story files; queues exactly one `outstanding_questions` item to ratify the completed map; blocks source and story implementation until ratification; and continues independent work. Structured-log the draw event to the iter summary (`notes`) per the ": Progress the Backlog While I'm Away" user outcome audit-trail. Do NOT use `ScheduleWakeup` under any circumstance (the "Problem 083: work-problems Step 5 iteration-worker prompt does not forbid ScheduleWakeup / time-deferring primitives — subagent can abandon synchronous-completion contract" problem — iteration workers must not self-reschedule). **NEVER call `AskUserQuestion` mid-loop in AFK** (the "Decision-delegation contract — agents over-apply Rule 1's interactive default to framework-resolved decisions; codify the framework-resolution boundary + AFK loop's batched-questions-as-deliverable + lazy-AskUserQuestion measurement" problem / the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule): direction / deviation-approval / one-time-override / silent-framework observations queue at `ITERATION_SUMMARY.outstanding_questions` for loop-end batched presentation. A documented ADR proposal may guide implementation before ratification. Keep the proposal unconfirmed; gather successful actual production-use evidence before final human ratification and acceptance. Preserve the other governance gates. Per-iter `AskUserQuestion` calls are sub-contracting framework-resolved decisions back to the user (lazy deferral per Step 2d Ask Hygiene Pass classification). Non-interactive defaults apply per the "Structured User Interaction for Governance-Skill Decisions" architecture rule Rule 6 + the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule's framework-resolution boundary. **Treat the user as transient** (the "`/wr-itil:work-problems` orchestrator defaults to subprocess dispatch even when the user is observably interactive — loses real-time presence advantage" problem): even when observably present at orchestrator dispatch time, the user may answer one question and disappear for hours; presence is not a reliable signal and is not the goal. The iter's job is to progress the ticket and accumulate questions for batched surfacing — not to ask "is it OK to proceed?" at a mechanical-stage boundary. **Do NOT poll `bats` output with a bats-console-summary regex against TAP-format output** (the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem — bash until-loop-deadlock antipattern). The bats-console-summary line `<N> tests, <M> failures` is emitted ONLY by bats's *default* (non-TAP) formatter; `bats --tap` does not emit a console summary, so a polling loop of shape `until [ -f $OUT ] && grep -qE '^[0-9]+ tests?,' $OUT; do sleep 5; done` spins forever after bats completes (silent deadlock — no error, no exit; recovery requires manual SIGTERM with metadata loss per the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem/the "SIGTERM-clean-flush guarantee is conditional on subprocess having emitted ITERATION_SUMMARY before going idle — needs SKILL.md caveat + behavioural-test second-source for stuck-before-emit subclass" problem stuck-before-emit subclass). When you need to wait on a backgrounded bats run, prefer `wait $bg_pid` (Unix idiom — completion signaled by process exit, no regex required) or, for the Bash tool, `run_in_background=true` + `BashOutput` polling on the tool's exit-state field rather than regex-poll on stdout. If you genuinely must regex-poll TAP output, anchor on the TAP plan line `^[0-9]+\.\.[0-9]+` (e.g. `1..1455`) — TAP's plan line is emitted on completion and is format-stable across bats versions; the bats-console-summary line is not. The console-summary vs TAP-format divergence is the load-bearing detail: `bats` and `bats --tap` produce structurally different stdout, and the antipattern assumes the former when iter dispatch typically uses the latter. **Do NOT poll subprocess completion with `pgrep -f '<pattern>'` inside an `until` / `while` loop** (the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem — self-referential pgrep deadlock; sibling variant of the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem). `pgrep -f` matches against the FULL command line of every running process, so the polling loop's own `zsh -c` argument (which contains the literal `pgrep -f '<pattern>'` text) matches itself; with multiple concurrent polling loops, each loop matches the others and spins forever. Worked example of the antipattern: `until ! pgrep -f 'bats --recursive' > /dev/null 2>&1; do sleep 5; done` — the 2026-05-16 the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem deadlock witness; 4 concurrent polling loops each matched the others' command lines while no actual bats process ran; 45 min wall-clock + $20-30 wasted before manual SIGTERM. The same self-reference shape applies to `while pgrep -f ...; do sleep; done` and to `until ! pkill -0 -f '<pattern>'` / `while pkill -0 -f '<pattern>'` (signal-0 polling). The structural fix is the same as the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem: prefer `wait $bg_pid` (Unix idiom — shell-native completion signal, no regex / no pgrep) or Bash-tool `run_in_background=true` + `BashOutput` polling (harness-tracked completion state). The hook `packages/itil/hooks/itil-bash-polling-antipattern-detect.sh` denies these shapes at PreToolUse:Bash, but the prompt rule belongs here too — structural enforcement + prompt discipline together close the class. **Do NOT leave a backgrounded task unreaped at turn-end** (`run_in_background: true` on an Agent or Bash tool call, or a `&`-detached shell job, whose completion you intend to observe in a *later* turn) inside iter dispatch contexts (the "Iter subprocess ends its turn waiting on a backgrounded task and never resumes — `claude -p` has no auto-resume; commit-bearing work is lost" problem — turn-end-mid-background work-loss; sibling-class to the "Problem 083: work-problems Step 5 iteration-worker prompt does not forbid ScheduleWakeup / time-deferring primitives — subagent can abandon synchronous-completion contract" problem / the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem / the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem). The iter subprocess is dispatched via `claude -p`, a single-shot CLI invocation with NO auto-resume affordance: its turn boundary IS its process boundary. A background task that outlives the turn never resumes — the iter exits at turn-end with the task incomplete and its own work staged but uncommitted (witnessed: iter 11 of a prior loop — $8.02 / 17 min / 8 staged files / 11 GREEN bats / ZERO commits; recovery required orchestrator main-turn salvage). **The prohibition is on the cross-turn / turn-end-survivor shape, NOT on backgrounding per se:** the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem/the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem-sanctioned idiom of launching `run_in_background=true` + `BashOutput`-poll-then-`wait $bg_pid` (or plain `wait $bg_pid` on a `&` job) **within the same turn** is fine — it reaps the task before turn-end. Use foreground-synchronous invocation instead: the Agent tool WITHOUT `run_in_background: true` (the result returns in-turn, so the commit step is reached), or intra-turn background that you `wait` on before the turn closes. The distinction from the "AFK iteration subprocess `bash until`-loop polls bats-output file with bats-console regex against TAP-format output — deadlocks indefinitely, manual SIGTERM required, JSON metadata lost" problem/the "bash until-loop with `pgrep -f 'bats --recursive'` self-references the polling loop's own command line — new variant of stuck-before-emit deadlock; SKILL.md prompt warning insufficient" problem polling antipatterns: those forbid *how* you wait (regex / pgrep poll loops); this forbids *deferring a task's completion past the turn boundary*, where `claude -p` has no notification re-entry to bring you back. The interactive Claude Code session masks this hazard (notification-driven re-entry); the AFK iter subprocess does not. **Release-preparation boundary (the ": Commit-time changeset enforcement forces premature release metadata" problem):** ordinary implementation commits MUST NOT include changesets and remain `STAGED`. After the exact implementation pipeline passes and release preparation is intentional, the orchestrator inspects the cumulative package scope and creates one complete changeset-only commit. Never create speculative, placeholder, dormant-foundation, or “for later” changesets. **`@jtbd the ": Progress the Backlog While I'm Away" user outcome`** (load-bearing) **`@jtbd the ": Keep Plugins Current Across Projects" user outcome`** (closure-dependent).
732
732
 
733
733
  4. **Retro-on-exit (the "Problem 086: AFK iteration subprocess does not run retro before returning — per-iteration lessons learnt are lost when the subprocess exits" problem) + retro-surfaced observation classification (the "Iter retros queue their own observations as `outstanding-questions.jsonl` entries for user-direction triage instead of auto-ticketing — same trust-boundary as `/wr-retrospective:run-retro` Step 4a" problem) + iter-owned BRIEFING commit (the "work-problems iteration boundary leaves run-retro BRIEFING.md edits uncommitted" problem)**: before emitting `ITERATION_SUMMARY`, invoke `/wr-retrospective:run-retro`. Retro runs INSIDE this subprocess so its Step 2b pipeline-instability scan has access to the iteration's rich tool-call history (hook misbehaviour, repeat-workaround patterns, subagent-delegation friction, release-path instability). Tickets retro creates ride a separate path: they delegate through `/wr-itil:manage-problem` which IS the "Governance Skills Commit Their Own Completed Work" architecture rule in-scope and self-commits each ticket per its own Step 11. Those commits land independently and the orchestrator picks them up on the next Step 1 scan.
734
734
 
@@ -181,18 +181,7 @@ What "work" means depends on the problem's status:
181
181
 
182
182
  **Known Error (root cause identified AND workaround documented; ready for fix proposal):**
183
183
 
184
- **Substance-confirm-before-build guard (the ": Confirm a decision's substance before building dependent work on it" architecture rule — propose-fix surface, the "Problem-RFC-Story framework with mandatory problem-trace and unified problem ontology" architecture rule I13).** BEFORE implementing any fix step below, check whether the fix builds on a genuine decision whose **substance is unconfirmed**. This closes the "Agent implements dependent work on genuine new decisions before human-confirming their SUBSTANCE — surfaces only meta-questions" problem (dependent work built on a born-`proposed` decision the user later rejects):
185
-
186
- 1. Collect the decisions this fix builds on: the `ADR-NNN` references in the problem's `## Fix Strategy` section PLUS the `adrs:` frontmatter array (and body `ADR-NNN` mentions) of each referenced RFC.
187
- 2. For each, run the predicate (PATH shim per the "Plugin-bundled scripts invoked from SKILL.md resolve via `bin/` on `$PATH`" architecture rule — adopter-safe, never source repo-relative lib files):
188
- ```bash
189
- wr-architect-is-decision-unconfirmed ADR-<NNN> docs/decisions
190
- ```
191
- Exit 0 = unconfirmed (born without `human-oversight: confirmed`, not superseded) → the guard fires for that ADR. Exit 1 = confirmed or superseded → OK to build. Exit 2 = not found → treat as not-a-blocker (surface in the report).
192
- 3. **If any referenced decision is unconfirmed**: do NOT implement yet. Surface its **substance** (the chosen option the ADR records — not a grain/meta question) for human confirmation:
193
- - **Interactive**: `request_user_input` presenting the ADR's Decision Outcome + Considered Options so the user confirms / amends / rejects the substantive choice. On confirm, the recording skill writes the `human-oversight: confirmed` marker (the ": Human-oversight marker + `/wr-architect:review-decisions` drain for recorded decisions" architecture rule) before the build proceeds.
194
- - **AFK** (`/wr-itil:work-problems` orchestrator): NEVER ask mid-loop — queue the substance to the iteration's `outstanding_questions` (the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule AFK carve-out) and skip the build; do not guess. This is the **queue-and-continue** universal default per the "Structured User Interaction for Governance-Skill Decisions" architecture rule Rule 6 (the "AFK iter default when a skill needs to ask a question and request_user_input is unavailable — should queue the question and move to the next iteration (not halt, not silently skip)" problem, 2026-06-06 amendment): the iter queues the substance + advances; the orchestrator main turn surfaces the queued question at loop end via the Step 2.5 batched request_user_input.
195
- 4. This ask is **the "— Decision-Delegation Contract: when agents act on the framework vs ask the user" architecture rule category-1 direction-setting** and is EXCLUDED from the lazy-request_user_input regression metric (it is legitimate, not lazy). The trigger is narrow — detection is mechanical (the predicate); only genuine unconfirmed decisions about to be built on fire it. Do NOT over-fire on confirmed/superseded/obvious decisions (inverse-the "Problem 078: Assistant does not offer to capture a problem ticket when the user delivers strong-signal correction" problem / the "Agents over-ask in interactive sessions — conflating mechanical-stages with user-interactive-stages of multi-stage skill contracts (inverse-)" problem guard). A born-`proposed` marker is fine for *recording*; it is not a licence to *build* (the ": Human-oversight marker + `/wr-architect:review-decisions` drain for recorded decisions" architecture rule carve-out).
184
+ **ADR proposal implementation and acceptance.** A fix may build against a documented proposed ADR before human ratification. Review its substance and preserve the existing story-map, job, risk, and release gates. If substantive direction is genuinely ambiguous, ask or queue that choice; do not halt solely because `human-oversight: unconfirmed`. Gather evidence of successful actual production use before requesting final ratification. The ADR remains proposed until both real production evidence and explicit human ratification of the final substance exist. Local tests and CI alone do not satisfy acceptance.
196
185
 
197
186
  **I13 propose-fix trace gate (the ": RFC-first trace invariant not enforced at fix-time" release design B3/B4).** BEFORE the traversal below, enforce the fix-time trace invariant: a fix proposed on a Known Error requires a fix vehicle that traces the problem (the "Every fix goes through an RFC" architecture rule unconditional; the "RFC required at the propose-fix step on a Known Error" architecture rule places the gate here, conforming to the "Problem lifecycle — add a Verification Pending status between Known Error and Closed" architecture rule Known Error semantics — the fix is proposed *after* Known Error).
198
187