@mrciphersmith/keryx 0.2.71 → 0.2.72

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/dist/cli.js +4296 -2223
  2. package/package.json +1 -1
  3. package/src/gdskills/bundled/rules/core/code-review-learned-profile.mdc +81 -0
  4. package/src/gdskills/bundled/rules/core/jobs-documentation.mdc +1 -1
  5. package/src/gdskills/bundled/rules/core/review-strict-profile.mdc +8 -4
  6. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.codex.md +2 -2
  7. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.cursor.md +2 -2
  8. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.md +2 -2
  9. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.opencode.md +2 -2
  10. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.zed.md +2 -2
  11. package/src/gdskills/bundled/skills/orchestration/context-collector/orchestrator-prompt.md +2 -2
  12. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.codex.md +3 -3
  13. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.cursor.md +3 -3
  14. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.md +3 -3
  15. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.opencode.md +3 -3
  16. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.zed.md +3 -3
  17. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.codex.md +1 -1
  18. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.cursor.md +1 -1
  19. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.md +2 -2
  20. package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +2 -2
  21. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.codex.md +1 -1
  22. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.cursor.md +1 -1
  23. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.md +1 -1
  24. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.opencode.md +1 -1
  25. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.zed.md +1 -1
  26. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.codex.md +972 -509
  27. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.cursor.md +972 -509
  28. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +943 -513
  29. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.opencode.md +972 -509
  30. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.zed.md +972 -509
  31. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/input-contract.schema.json +38 -28
  32. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/orchestrator-prompt.md +98 -66
  33. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/output-contract.schema.json +27 -5
  34. package/src/gdskills/bundled/skills/orchestration/task-implementer/input-contract.schema.json +1 -1
  35. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.md +2 -2
  36. package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.codex.md +252 -0
  37. package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.cursor.md +252 -0
  38. package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.md +243 -0
  39. package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.opencode.md +252 -0
  40. package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.zed.md +252 -0
  41. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.md +1 -1
  42. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
  43. package/src/gdskills/bundled/skills/review/review-architecture/SKILL.md +1 -1
  44. package/src/gdskills/bundled/skills/review/review-backend/SKILL.md +1 -1
  45. package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +1 -1
  46. package/src/gdskills/bundled/skills/review/review-core-boundaries/SKILL.md +1 -1
  47. package/src/gdskills/bundled/skills/review/review-flow-graph/SKILL.md +1 -1
  48. package/src/gdskills/bundled/skills/review/review-frontend/SKILL.md +1 -1
  49. package/src/gdskills/bundled/skills/review/review-frontend-conventions/SKILL.md +1 -1
  50. package/src/gdskills/bundled/skills/review/review-highload/SKILL.md +1 -1
  51. package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +2 -2
  52. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +10 -10
  53. package/src/gdskills/bundled/skills/review/review-performance/SKILL.md +1 -1
  54. package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +36 -17
  55. package/src/gdskills/bundled/skills/review/review-security-code/SKILL.md +1 -1
  56. package/src/gdskills/bundled/skills/review/review-style/SKILL.md +1 -1
  57. package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +1 -1
  58. package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +1 -1
  59. package/src/gdskills/bundled/skills/shared/git-merge-base.md +1 -1
  60. package/src/gdskills/bundled/rules/core/code-review-b091-profile.mdc +0 -48
  61. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.codex.md +0 -209
  62. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.cursor.md +0 -209
  63. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.md +0 -208
  64. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.opencode.md +0 -209
  65. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.zed.md +0 -209
@@ -37,7 +37,11 @@ Proceed directly with your assigned task.
37
37
 
38
38
  ## Purpose
39
39
 
40
- Dynamic orchestrator that builds execution plans based on user intent. Unlike a fixed pipeline, the orchestrator adapts its workflow to what the user actually needs — from "just analyze this issue" to "implement, review, and create a PR". It dispatches sub-agents (`issue-analyzer`, `context-collector`, `task-implementer`, review skills) and persists all work via `job-documenter`.
40
+ Dynamic orchestrator that builds execution plans based on user intent. Unlike a fixed pipeline, the orchestrator adapts its workflow to what the user actually needs — from "just analyze this issue" to "implement, review, and create a PR". It dispatches sub-agents (`issue-analyzer`, `context-collector`, `tests-creator`, `task-implementer`, `code-verifier`, `review-orchestrator`) and persists every step, document and retry through `keryx job`, which writes `.metaproject/jobs/<job-name>/`.
41
+
42
+ **The package is the state.** `keryx job` is the only writer of `state.json`; it validates every write against the registered contract `job-orchestrator-state` and refuses one that does not conform. Never hand-write `state.json`, and never hold a step's outcome only in this session — a step recorded nowhere is a step that did not happen as far as the next session is concerned.
43
+
44
+ **Execution metrics (opt-in):** when a USER runs this orchestrator directly (not as a dispatched subagent), at the start ask "Collect execution statistics for this run? (yes/no)" per `.metaproject/rules/core/execution-metrics.md`. If yes, append the `## Execution Metrics` section at the end and save it under the job dir (`jobs/<job>/metrics/`). Never ask or emit it when dispatched as a subagent.
41
45
 
42
46
  **Key design principle** (from Anthropic's "Building Effective Agents"):
43
47
  > "The key difference from parallelization is its flexibility — subtasks aren't pre-defined, but determined by the orchestrator based on the specific input."
@@ -70,11 +74,32 @@ Phase 3: COMPLETION → Final report, optional PR, tell user where doc
70
74
 
71
75
  ### 0.0 State Resumption Check
72
76
 
73
- Before asking any questions, check if an interrupted job exists:
74
- 1. Look in `$JOBS_ROOT` for any directory containing an incomplete `state.json`.
75
- 2. If found, ASK the user: "Found paused job '<job-name>'. Do you want to resume it or start a new orchestrated job?"
76
- 3. If resume → Parse `state.json`, restore `JOB_STATE`, and jump directly to the first uncompleted step in Phase 2.
77
- 4. If new → Proceed to 0.1.
77
+ Before asking any questions, list existing job packages:
78
+
79
+ ```bash
80
+ keryx job list --json
81
+ ```
82
+
83
+ Every entry carries `phase`, `stepsDone`/`stepsTotal` and `nextStep`. A job whose
84
+ `phase` is not `COMPLETION` is unfinished.
85
+
86
+ 1. If an unfinished job exists, ASK the user:
87
+ "Found unfinished job '<job-name>' (<stepsDone>/<stepsTotal> steps, next: <nextStep>).
88
+ Resume it or start a new orchestrated job?"
89
+ 2. If resume → read the package and jump directly to the step it names:
90
+
91
+ ```bash
92
+ keryx job status <job-name> --json
93
+ ```
94
+
95
+ `next_step` is the first step that is neither `completed` nor `skipped` — computed
96
+ from the file, not recalled. `retries` gives the recorded attempt count per step, so
97
+ a resumed session continues from the real number instead of restarting at zero, and
98
+ `documents` lists what has already been produced.
99
+ 3. If new → proceed to 0.1.
100
+
101
+ There is no `paused` status and nothing writes one. A job is unfinished exactly when a
102
+ step is still open, and `keryx job status` is what reports that.
78
103
 
79
104
  ### 0.1 Determine User Intent
80
105
 
@@ -82,7 +107,7 @@ Parse the user's request to identify the intent:
82
107
 
83
108
  | User Says | Intent | Plan Type |
84
109
  |-----------|--------|-----------|
85
- | "Implement issue #N" / "Issue to PR" | `implement` | Full: analyze → branch → implement → review → fix → checks → PR |
110
+ | "Implement issue #N" / "Issue to PR" | `implement` | Full: analyze → branch → implement → verify → review → fix → PR |
86
111
  | "Analyze issue #N" / "Study issue" | `analyze` | Analysis only: analyze → report. Then ask if user wants to implement. |
87
112
  | "Review my code" / "Review branch" | `review` | Review only: review → report |
88
113
  | "Analyze and implement" | `implement` | Same as implement |
@@ -121,7 +146,7 @@ For `custom` intent OR any ambiguous request, invoke the `interviewer` skill **b
121
146
 
122
147
  **Invoke:**
123
148
  ```
124
- Load skill: skills/interviewer/SKILL.md
149
+ Load skill: skills/gdskills/planning/interviewer/SKILL.md
125
150
 
126
151
  INPUT:
127
152
  topic: <user's original request>
@@ -163,7 +188,9 @@ The orchestrator MUST collect all required context before proceeding:
163
188
  git -C <project_dir> symbolic-ref refs/remotes/origin/HEAD 2>/dev/null | sed 's@^refs/remotes/origin/@@'
164
189
  # Fallback: check for main, master, develop
165
190
  ```
166
- Present detected branch and ask to confirm. No hardcoded default.
191
+ Present detected branch and ask to confirm. No hardcoded default — and
192
+ `input-contract.schema.json` declares none either, so the contract cannot
193
+ reintroduce one behind the question.
167
194
 
168
195
  **Intent-specific questions:**
169
196
 
@@ -236,13 +263,24 @@ Ready to proceed:
236
263
  Intent: implement
237
264
  Issue: #4141 — Pipeline validation improvements
238
265
  Project: /Users/.../<PROJECT>
239
- Base: develop-2
266
+ Base: <detected base branch>
240
267
  Create PR: yes
241
268
  Job name: issue-4141--pipeline-validation
242
269
 
243
270
  Proceed? (yes / adjust)
244
271
  ```
245
272
 
273
+ This is the **operator** gate and it is not governed by `skip_confirmation`. That
274
+ setting is `{"const": true}` in `input-contract.schema.json` and means exactly one
275
+ thing: dispatched sub-agents run without asking the operator to approve each
276
+ dispatch. It has never covered this question, and the two are named apart here so
277
+ the contract and the prose stop reading as a contradiction. The gate that *can* be
278
+ turned off is `plan_approval` in 1.3.
279
+
280
+ `job_name` must match `^[a-z0-9-]+$` — the pattern `state.schema.json` declares and
281
+ `keryx job init` enforces before it builds a path from the value. `issue-4141--pipeline-validation`
282
+ conforms; anything with a slash, a space or an uppercase letter is refused.
283
+
246
284
  ---
247
285
 
248
286
  ## Phase 1: PLAN BUILDING
@@ -271,22 +309,40 @@ PLAN:
271
309
  15. { id: "deploy", type: "deploy", agent: "deploy", depends: ["pr"], conditional: true }
272
310
  ```
273
311
 
274
- **Conditional step triggers:**
275
- - `sanity-check`: always runs — verifies ≥1 commit was made
276
- - `tests-creator`: always runs — mandatory TDD step before every task-implementer wave
277
- - `verify`: always runs — code-verifier is the mandatory quality gate after implementation
278
- - `security`: diff touches auth/, api/, migrations, schema files, or `.env`
279
- - `fix`: review or verify found CRITICAL/HIGH findings
280
- - `verify-post-fix`: always runs after fix (confirms fix resolved the findings)
281
- - `perf-check`: diff contains *.tsx, *.jsx, *.css, dist/, build/ files
282
- - `security`: diff touches auth/, api/, migrations, schema files, or `.env`
283
- - `fix`: review found CRITICAL/WARNING findings
284
- - `perf-check`: diff contains *.tsx, *.jsx, *.css, dist/, build/ files
285
- - `pr`: `create_pr: true`
286
- - `deploy`: user answers "yes" to post-PR staging deploy prompt
312
+ This is the plan `keryx job init --intent implement` writes, step for step. The `agent`
313
+ field is the **label recorded in the plan**, not a dispatch target: the `review` step is
314
+ executed by `review-orchestrator` (2.6), and `orchestrator` means this skill does the
315
+ step itself. Read the recorded plan back at any time with `keryx job status <job-name>`.
316
+
317
+ **Conditional step triggers** — one row per step, no step listed twice:
318
+
319
+ | Step | Runs when |
320
+ |------|-----------|
321
+ | `sanity-check` | always — verifies ≥1 commit was made |
322
+ | `tests-creator` | always — mandatory TDD step before every task-implementer wave |
323
+ | `verify` | always — `code-verifier` is the mandatory quality gate after implementation |
324
+ | `security` | diff touches `auth/`, `api/`, migrations, schema files, or `.env` |
325
+ | `fix` | review or verify produced a `blocker` or `major` finding |
326
+ | `verify-post-fix` | after `fix` ran — confirms the fix resolved the findings |
327
+ | `perf-check` | diff contains `*.tsx`, `*.jsx`, `*.css`, `dist/` or `build/` files |
328
+ | `pr` | `create_pr: true` |
329
+ | `deploy` | user answers "yes" to the post-PR staging deploy prompt |
330
+
331
+ Severities are the canonical four — `blocker`, `major`, `minor`, `info` — from
332
+ `review-finding.schema.json`. They are the only vocabulary this skill uses, so the
333
+ `fix` trigger and the counts in the report are read off the same field.
287
334
 
288
335
  Note: `security` runs in parallel with `review` (both depend on `implement` results, no overlap).
289
336
 
337
+ **A conditional step is not exempt from the record.** Every step in the plan is
338
+ written into the package by `keryx job init`, and `keryx job complete` refuses while
339
+ any step is neither `completed` nor `skipped`. A condition that did not fire is
340
+ closed explicitly:
341
+
342
+ ```bash
343
+ keryx job step <job-name> perf-check --status skipped --reason "no frontend files in diff"
344
+ ```
345
+
290
346
  **For `analyze` intent:**
291
347
  ```
292
348
  PLAN:
@@ -308,72 +364,90 @@ PLAN:
308
364
  **For `custom` intent:**
309
365
  Build plan dynamically. Each step must have: id, type, agent, dependencies.
310
366
 
311
- ### 1.2 Initialize Job Documentation
367
+ ### 1.2 Create the Job Package
312
368
 
313
- Dispatch `job-documenter` with `init` action:
369
+ Create the package with the CLI. This is one command, run by the orchestrator — not
370
+ a sub-agent dispatch:
314
371
 
372
+ ```bash
373
+ keryx job init --name <job-name> --intent implement|analyze|review|custom --project <project_dir>
315
374
  ```
316
- Task({
317
- description: "Init job docs: <job-name>",
318
- subagent_type: "general",
319
- prompt: |
320
- You are the job-documenter agent.
321
- Load skill: skills/job-documenter/SKILL.md
322
- Follow rules: rules/core/jobs-documentation.mdc
323
375
 
324
- ACTION: init
325
- JOB_NAME: <job-name>
326
- JOBS_ROOT: <JOBS_ROOT>
376
+ It creates `.metaproject/jobs/<job-name>/` containing:
327
377
 
328
- DATA:
329
- TITLE: <job title>
330
- DESCRIPTION: <description>
331
- INTENT: <intent>
332
- SOURCE: <issue URL or description>
333
- PROJECT: <project path>
334
- BRANCH: TBD
335
- BASE_BRANCH: <base branch>
336
- PLAN: <plan steps>
337
-
338
- Execute and return DOCUMENTER_RESULT.
339
- })
378
+ - `state.json` — validated against the registered contract `job-orchestrator-state`
379
+ on **every** write. A state that does not conform is refused, not written.
380
+ - `journal.md` — append-only, one line per recorded event, written by `keryx job`.
381
+ - the plan for the chosen intent, every step `pending`, with `plan.current_step`
382
+ already pointing at the first one.
383
+
384
+ `--intent` defaults to `implement`. `--project` defaults to the current directory;
385
+ pass the path collected in 0.2 explicitly rather than relying on the default.
386
+
387
+ **Refusals to expect, and what each means:**
388
+
389
+ | Message | Cause |
390
+ |---------|-------|
391
+ | `Job package already exists: .metaproject/jobs/<name>` | The package is there. Run `keryx job status <name>` and resume it (0.0) instead of re-initialising. |
392
+ | `Invalid --name "<name>"` | The name is not `^[a-z0-9-]+$`. |
393
+ | `Invalid --intent "<value>"` | Not one of `implement`, `analyze`, `review`, `custom`. |
394
+
395
+ Confirm the result before proceeding:
396
+
397
+ ```bash
398
+ keryx job status <job-name>
340
399
  ```
341
400
 
342
- **Validate response:** status must be `success`. If `error` → report to user, ask how to proceed.
401
+ It prints the phase, the step list with statuses, and `next:` — the step execution
402
+ starts from.
343
403
 
344
404
  ### 1.3 Display Plan + Agent Approval
345
405
 
346
- Show each step with its agent and status, then ask the user to approve or adjust:
406
+ Display the plan the package actually holds — do not retype it from memory:
407
+
408
+ ```bash
409
+ keryx job status <job-name>
410
+ ```
411
+
412
+ There is **one** plan. Every step listed in 1.1 is in it, including the conditional
413
+ ones; a conditional step is one whose trigger may not fire, not one that is absent
414
+ until somebody adds it. For the `implement` intent that is fifteen steps:
347
415
 
348
416
  ```
349
- Execution plan — <N> steps:
417
+ Execution plan — 15 steps (◦ = conditional):
350
418
 
351
419
  Step 1 analyze issue-analyzer → issue #<N>
352
420
  Step 2 context context-collector → project context + test framework
353
- Step 3 prepare orchestrator → feature branch
421
+ Step 3 prepare orchestrator → feature branch worktree
354
422
  Step 4 tests-creator tests-creator × <tasks> → RED test stubs per task (MANDATORY)
355
423
  Step 5 implement task-implementer × <tasks> → <N> tasks make tests GREEN (wave-parallel)
356
424
  Step 6 sanity-check orchestrator → verify commits exist
357
425
  Step 7 verify code-verifier → lint + type-check + tests + imports (MANDATORY)
358
- Step 8 review code-review × 4 → parallel agents
359
- Step 9 fix task-implementer → [conditional: CRITICAL/HIGH findings]
360
- Step 10 verify-post-fix code-verifier → [conditional: after fix]
361
- Step 11 report orchestrator → final summary
362
- Step 12 pr orchestrator + gh CLI → [conditional: create_pr=true]
426
+ Step 8 review review-orchestrator → managed review round
427
+ Step 9 ◦ security security-audit → auth/API/DB/env files touched
428
+ Step 10◦ fix task-implementer → blocker or major findings
429
+ Step 11◦ verify-post-fix code-verifier → after fix
430
+ Step 12◦ perf-check perf-check → frontend/bundle files changed
431
+ Step 13 report orchestrator → final summary
432
+ Step 14◦ pr orchestrator + gh CLI → create_pr=true
433
+ Step 15◦ deploy deploy → user asked for a staging deploy
363
434
 
364
- Optional (not in plan — add if needed):
365
- + security-audit auto-detect: auth/API/DB changes
366
- + perf-check auto-detect: if frontend/bundle files changed
367
- + deploy ask after PR: "Deploy to staging?"
435
+ Proceed? (yes / adjust: "skip fix", "remove pr", etc.)
436
+ ```
368
437
 
369
- Proceed? (yes / adjust: "skip fix", "add security-audit", "remove pr", etc.)
438
+ **If user adjusts:** record the decision in the package rather than holding it in
439
+ this session:
440
+
441
+ ```bash
442
+ # "skip fix" — close it now, with the reason on the record
443
+ keryx job step <job-name> fix --status skipped --reason "operator asked to skip at plan approval"
444
+ # "remove pr"
445
+ keryx job step <job-name> pr --status skipped --reason "create_pr: false"
370
446
  ```
371
447
 
372
- **If user adjusts:**
373
- - Parse natural language: "skip fix" → mark `fix` step as disabled
374
- - "add security-audit" → insert `{ id: "security-audit", agent: "security-audit", depends: ["review"] }` after review
375
- - "remove pr" → set `create_pr: false`
376
- - Re-display updated plan and ask again
448
+ Then re-display with `keryx job status <job-name>` and ask again. A step the operator
449
+ removed is `skipped` with a reason, never silently dropped — that is the difference
450
+ between a plan somebody changed and a plan that quietly shrank.
377
451
 
378
452
  **If `plan_approval: false`** (automation setting) → skip this display and proceed directly.
379
453
 
@@ -385,41 +459,75 @@ Execute each step in plan order, documenting results after each step.
385
459
 
386
460
  ### 2.1 General Execution Loop
387
461
 
462
+ Every step in the loop is bracketed by two `keryx job` calls. The package, not this
463
+ session, is what says a step ran.
464
+
388
465
  ```
389
466
  FOR step in PLAN:
390
467
  IF step.conditional AND condition_not_met:
391
- SKIP step, mark as "skipped"
468
+ keryx job step <job-name> <step-id> --status skipped --reason "<why the trigger did not fire>"
392
469
  CONTINUE
393
470
 
394
- 2.1.1 Mark step as in-progress (update display)
471
+ 2.1.1 Open the step:
472
+ keryx job step <job-name> <step-id> --status in-progress
473
+ Re-entering a step that was already opened increments `metrics.steps[].retries`
474
+ — that counter is the attempt budget, and it survives a session restart.
475
+
395
476
  2.1.2 Execute step (see step-specific instructions below)
396
477
  **CRITICAL RESILIENCE**: If the sub-agent returns a malformed result or fails to follow formatting rules, run an explicit retry:
397
478
  "The previous output was malformed. Fix these errors: [errors] and try again." (Max 2 retries before counting as critical failure).
479
+ Re-open the step before each retry so the retry is counted.
480
+
398
481
  2.1.3 Collect result
399
- 2.1.4 Document result via job-documenter (add-document)
400
- (Also update job state `state.json`)
401
- 2.1.5 Update job README via job-documenter (update-readme)
402
- 2.1.6 Mark step as completed
403
-
482
+
483
+ 2.1.4 Write the document to disk, then record it in the package:
484
+ keryx job document <job-name> --type analysis|implementation-report|review|verification-report --file <path>
485
+ The file must already exist — `job document` refuses a `--file` it cannot
486
+ find with "Write the document first, then record it." It copies the file
487
+ into the package and adds it to `documentation.documents_created`.
488
+ Re-recording the same type replaces the file and leaves one entry.
489
+
490
+ 2.1.5 Confirm what the package now holds:
491
+ keryx job status <job-name>
492
+ The step list, the retry counts and the recorded documents come from
493
+ `state.json`. This is the job index; there is no README to update.
494
+
495
+ 2.1.6 Close the step:
496
+ keryx job step <job-name> <step-id> --status completed
497
+
404
498
  IF step failed critically:
499
+ keryx job step <job-name> <step-id> --status failed --reason "<what failed>"
405
500
  Ask user: "Step '<name>' failed. Continue with remaining steps or abort?"
406
- IF abort: skip to Phase 3 (COMPLETION) with status "aborted"
501
+ IF abort: skip to Phase 3 (COMPLETION)
407
502
  ```
408
503
 
504
+ **`failed` is not terminal.** `keryx job complete` refuses while any step is `failed`
505
+ or still open, and names them. A job that genuinely ends with a step unfinished is
506
+ closed by deciding what happened to that step — `--status skipped --reason "<why>"` —
507
+ which leaves the decision on the record instead of leaving the package half-written.
508
+
509
+ Only four document types exist: `analysis`, `implementation-report`, `review`,
510
+ `verification-report`. Anything else is refused with the valid list.
511
+
409
512
  ### 2.2 Step: ANALYZE
410
513
 
411
514
  Dispatch `issue-analyzer` as a sub-agent.
412
515
 
413
- **Prepare prompt:** Read `skills/issue-analyzer/orchestrator-prompt.md` (if it exists) and fill in:
516
+ **Prepare prompt:** Read `skills/gdskills/orchestration/issue-analyzer/orchestrator-prompt.md`
517
+ and fill in:
414
518
  - Issue URL or repo+number
415
519
  - Codebase paths with roles
416
520
  - Automation settings (skip_confirmation: true, search_depth: focused)
417
521
 
522
+ That file ships with the skill. If the read fails, the path is wrong or the skill is
523
+ not installed — stop and say so. Do not proceed on an improvised prompt: a missed
524
+ template is exactly the failure that hid behind the old "(if it exists)" hedge.
525
+
418
526
  **Launch:**
419
527
  ```
420
528
  Task({
421
529
  description: "Issue analysis: #<N>",
422
- subagent_type: "general",
530
+ subagent_type: "general-purpose",
423
531
  prompt: <constructed prompt>
424
532
  })
425
533
  ```
@@ -437,18 +545,16 @@ ANALYSIS_RESULT:
437
545
 
438
546
  **Validate:** At least 1 task, no circular dependencies, all dependency references valid. Dependency_order array must contain all task_ids exactly once.
439
547
 
440
- **Document:** Send to job-documenter:
441
- ```
442
- ACTION: add-document
443
- DATA:
444
- DOC_TYPE: analysis
445
- TARGET: both
446
- TITLE: Issue Analysis — #<N>
447
- CONTENT: <human-readable summary for man/, raw JSON for ai/>
448
- AGENT: issue-analyzer
449
- TASK: Analyze issue #<N>
548
+ **Document:** write the analysis, then record it:
549
+
550
+ ```bash
551
+ keryx job document <job-name> --type analysis --file <path/to/analysis.md>
450
552
  ```
451
553
 
554
+ It lands in the package as `analysis.md` (the source extension is preserved, so a
555
+ `.json` analysis lands as `analysis.json`) and appears in `documents` on the next
556
+ `keryx job status`.
557
+
452
558
  **For `analyze` intent:** After documenting, present analysis to user. Ask:
453
559
  ```
454
560
  Analysis complete. Found <N> tasks.
@@ -456,24 +562,24 @@ Want me to implement this? I'll create a feature branch and run the full pipelin
456
562
  ○ Yes, implement
457
563
  ○ No, analysis is enough
458
564
  ```
459
- If "Yes" → extend PLAN with context → prepare → implement → review → fix → checks → pr steps. Continue execution.
565
+ If "Yes" → follow Plan Extension below: create an `implement` package and continue there. Do not rewrite this package's plan.
460
566
  If "No" → skip to Phase 3 (COMPLETION).
461
567
 
462
568
  ### 2.3 Step: CONTEXT
463
569
 
464
570
  Dispatch `context-collector` to build the unified context document.
465
571
 
466
- **Prepare prompt:** Use the template from `skills/context-collector/SKILL.md`:
572
+ **Prepare prompt:** Use the template from `skills/gdskills/orchestration/context-collector/SKILL.md`:
467
573
 
468
574
  ```
469
575
  Task({
470
576
  description: "Collect context: <job-name>",
471
- subagent_type: "general",
577
+ subagent_type: "general-purpose",
472
578
  prompt: |
473
579
  You are the context-collector agent. Your task is to research and build
474
580
  a context document for the current job.
475
581
 
476
- Load the skill from: skills/context-collector/SKILL.md
582
+ Load the skill from: skills/gdskills/orchestration/context-collector/SKILL.md
477
583
 
478
584
  ACTION: collect
479
585
  JOB_NAME: <job-name>
@@ -500,16 +606,27 @@ CONTEXT_RESULT:
500
606
 
501
607
  **Validate:** status must be `success`. If `error` → log warning, continue (context is helpful but not blocking).
502
608
 
503
- **After context is collected:** All subsequent sub-agents receive the **versioned** context path from state.json:
609
+ **After context is collected:** the orchestrator holds the context path and puts it
610
+ into every subsequent dispatch prompt:
611
+
504
612
  ```
505
- CONTEXT_LOCATION: <JOBS_ROOT>/<job-name>/ai/context_v<N>.md
613
+ CONTEXT_LOCATION: <JOBS_ROOT>/<job-name>/context_v<N>.md
506
614
  ```
507
615
 
508
- **Context versioning:** Never overwrite `context.md` — save snapshots as `context_v1.md`, `context_v2.md`, etc.
509
- - Version 1 is created during Step 2.3 (first collect)
510
- - Subsequent versions increment on each update
511
- - `state.json → context_doc.version` always points to the latest version
512
- - Sub-agents always read the path from `state.json`, not a hardcoded filename
616
+ **Context versioning:** never overwrite an existing context file — write snapshots as
617
+ `context_v1.md`, `context_v2.md`, and so on. Version 1 comes from the first collect in
618
+ 2.3; each update writes the next number.
619
+
620
+ The current version is the highest-numbered file in the package, which is a fact on
621
+ disk that any session can read:
622
+
623
+ ```bash
624
+ ls .metaproject/jobs/<job-name>/context_v*.md
625
+ ```
626
+
627
+ `state.json` does not carry a context pointer and nothing writes one — do not tell a
628
+ sub-agent to look for one. The orchestrator passes the path (Constructing Subagent
629
+ Context, below); subagents receive, they do not retrieve.
513
630
 
514
631
  **Triggering context updates during execution:**
515
632
 
@@ -518,11 +635,11 @@ If during later steps (implement, review) a sub-agent reports missing context or
518
635
  ```
519
636
  Task({
520
637
  description: "Update context: <job-name>",
521
- subagent_type: "general",
638
+ subagent_type: "general-purpose",
522
639
  prompt: |
523
640
  You are the context-collector agent. Update the existing context.
524
641
 
525
- Load the skill from: skills/context-collector/SKILL.md
642
+ Load the skill from: skills/gdskills/orchestration/context-collector/SKILL.md
526
643
 
527
644
  ACTION: update
528
645
  JOB_NAME: <job-name>
@@ -598,122 +715,145 @@ BRANCH_STATE:
598
715
  run_command: <RUNNER>
599
716
  ```
600
717
 
601
- > **Store `package_manager` and `run_command` in JOB_STATE** — all subsequent steps use these instead of hardcoded `npm`.
718
+ > **Carry `package_manager` and `run_command` into every subsequent dispatch prompt** — all subsequent steps use these instead of hardcoded `npm`. They are not persisted; the orchestrator holds them for the run and states them explicitly in each dispatch.
719
+
720
+ **Record:** close the step and put the branch on the record:
721
+
722
+ ```bash
723
+ keryx job step <job-name> prepare --status completed --reason "feature/<branch-slug> at <worktree_path>"
724
+ ```
602
725
 
603
- **Document:** Update README via job-documenter (update-readme) with branch info.
726
+ `--reason` is appended to the package's `journal.md` with a timestamp, which is where
727
+ "what branch did this job use" is answerable after the session ends.
604
728
 
605
- ### 2.4.1 Step: TESTS-CREATOR + IMPLEMENT — Wave Isolation
729
+ ### 2.5 Step: TESTS-CREATOR + IMPLEMENT
606
730
 
607
731
  **IRON LAW: tests-creator MUST run before task-implementer for every task. No exceptions.**
608
732
 
609
- **CONTEXT BUDGET RULE: Each wave runs as a single isolated sub-agent. The orchestrator never dispatches task-implementers or tests-creator directly. This keeps the orchestrator context bounded to compact wave summaries regardless of job size.**
733
+ There is no `wave-executor` agent. Each wave is two dispatches the orchestrator makes
734
+ itself — `tests-creator`, then `task-implementer` — and both are real, installed
735
+ skills. Nothing is delegated to an intermediary that does not exist.
736
+
737
+ **CONTEXT BUDGET RULE: instruct every dispatched agent to write its full result to a
738
+ file and return only a compact summary line.** The orchestrator's context grows with
739
+ what agents *return*, not with what they do; a returned result file path costs a line,
740
+ an inlined verification log costs thousands. After 3–4 waves of inlined results the
741
+ session freezes on context reload, which is the failure this rule exists to avoid.
610
742
 
611
743
  ---
612
744
 
613
- #### Why wave isolation
745
+ #### Wave ordering
614
746
 
615
- When the orchestrator dispatches task-implementers directly, each sub-agent result (STATUS text + verification output) accumulates in the orchestrator's context. After 3–4 waves this context can reach 100k+ tokens, causing the session to freeze during context reload. Wave isolation prevents this: each wave sub-agent runs in its own context and returns only a compact summary.
616
-
617
- ---
747
+ Waves come from `dependency_order` in `ANALYSIS_RESULT`, which `issue-analyzer` already
748
+ returned topologically sorted and which 2.2 validated. Wave 1 is every task with no
749
+ unsatisfied dependency; wave N+1 is every task whose dependencies are all in waves 1..N.
750
+ Do not re-derive an ordering the analysis already produced.
618
751
 
619
752
  #### Execution pattern
620
753
 
621
754
  ```
622
- WAVES = topological_sort_into_waves(dependency_order, task_dependencies)
623
-
624
755
  FOR wave_index, wave_tasks in enumerate(WAVES):
625
- Dispatch SINGLE Agent("wave-executor") with all tasks in this wave.
626
-
627
- Receive compact WAVE_RESULT:
628
- STATUS: WAVE_DONE | WAVE_PARTIAL | WAVE_FAILED
629
- Wave: <index>
630
- Commits: [hash msg, hash msg, ...]
631
- Tests: <N passed, M failed>
632
- Tasks: task-1 ✅, task-2 ✅
633
- Result files: <JOBS_ROOT>/<job-name>/results/task-*.json
634
-
635
- Decision:
636
- WAVE_DONE → continue to next wave
637
- WAVE_PARTIAL → log warnings, continue (read result files for details)
638
- WAVE_FAILED → STOP, read result files for failed tasks, ask user
756
+
757
+ keryx job step <job-name> tests-creator --status in-progress # wave 1 only
758
+ # Step A — tests-creator (MANDATORY, run first)
759
+ Dispatch one tests-creator per task in this wave, in a SINGLE turn (parallel).
760
+ Wait for ALL of them. Collect TEST_SPECS[task_id] from each response.
761
+ keryx job step <job-name> tests-creator --status completed # last wave only
762
+
763
+ # Parallel safety check, before Step B:
764
+ # if two tasks in this wave share a target_file, dispatch them sequentially.
765
+
766
+ keryx job step <job-name> implement --status in-progress # wave 1 only
767
+ # Step B — task-implementer (after all test stubs are committed)
768
+ Dispatch one task-implementer per task in this wave, in a SINGLE turn (parallel),
769
+ each carrying test_case_specs: TEST_SPECS[task_id].
770
+ Wait for ALL of them.
771
+
772
+ Read each result's STATUS line:
773
+ all DONE → continue to next wave
774
+ any DONE_WITH_CONCERNS → record the concerns, continue
775
+ any BLOCKED → STOP, read the result file, resolve or ask the user
639
776
  ```
640
777
 
641
- #### Wave executor prompt template
778
+ #### tests-creator dispatch (Step A)
642
779
 
643
780
  ```
644
781
  Task({
645
- description: "Wave <N>: implement tasks <task_ids>",
646
- subagent_type: "general",
782
+ description: "Wave <N> tests: <task_id>",
783
+ subagent_type: "general-purpose",
647
784
  prompt: |
648
- You are a wave executor. Implement all tasks in this wave, then return a compact summary.
649
-
650
- ## Wave
651
- Wave <N> of <total>
652
-
653
- ## Tasks
654
- <JSON array of task objects for this wave>
655
-
785
+ Load skill: skills/gdskills/quality/tests-creator/SKILL.md
786
+
787
+ ## Task
788
+ <the single task object>
789
+
656
790
  ## Workspace
657
- - worktree_path: <absolute path>
658
- - branch: <branch name>
659
- - package_manager: <pm>
660
- - run_command: <runner>
661
- - issue_number: <N>
662
- - job_name: <job-name>
663
- - context_path: <path to context_vN.md>
664
-
665
- ## Instructions
666
-
667
- **Step A — tests-creator (MANDATORY, run first):**
668
- For each task in this wave, dispatch tests-creator in parallel:
669
- Load skill: skills/tests-creator/SKILL.md
670
- Pass: task object, workspace, context_path
671
- Collect: TEST_SPECS[task_id] from each response
672
- Wait for ALL tests-creator agents to finish before Step B.
673
-
674
- **Step B — task-implementer (after all test stubs committed):**
675
- For each task in this wave, dispatch task-implementer in parallel (if no file overlap; sequential otherwise):
676
- Load skill: skills/task-implementer/SKILL.md
677
- Pass: task object WITH test_case_specs: TEST_SPECS[task_id], workspace, job_name, context_path
678
- Wait for ALL task-implementer agents to finish.
679
-
680
- **Parallel safety check:** Before Step B, verify no two tasks share target_files.
681
- If overlap → run sequentially within this wave.
682
-
791
+ - worktree_path: <absolute path>
792
+ - branch: <branch name>
793
+ - package_manager: <pm>
794
+ - run_command: <runner>
795
+ - context_path: <JOBS_ROOT>/<job-name>/context_v<N>.md
796
+
797
+ ## Required response
798
+ Begin with STATUS: <STATUS>. Return the test_case_specs for this task and
799
+ nothing else inline; write anything longer to
800
+ <JOBS_ROOT>/<job-name>/results/<task_id>-tests.json and return the path.
801
+ })
802
+ ```
803
+
804
+ #### task-implementer dispatch (Step B)
805
+
806
+ ```
807
+ Task({
808
+ description: "Wave <N> implement: <task_id>",
809
+ subagent_type: "general-purpose",
810
+ prompt: |
811
+ Load skill: skills/gdskills/orchestration/task-implementer/SKILL.md
812
+
813
+ ## Task
814
+ <the single task object, WITH test_case_specs: TEST_SPECS[task_id]>
815
+
816
+ ## Workspace
817
+ - worktree_path: <absolute path>
818
+ - branch: <branch name>
819
+ - package_manager: <pm>
820
+ - run_command: <runner>
821
+ - issue_number: <N>
822
+ - job_name: <job-name>
823
+ - context_path: <JOBS_ROOT>/<job-name>/context_v<N>.md
824
+
683
825
  ## Required response format (compact — no inline JSON)
684
-
685
- STATUS: WAVE_DONE
686
- Wave: <N>
687
- Commits: [abc1234 feat(x): ..., def5678 feat(y): ...]
826
+ STATUS: DONE
827
+ Task: <task_id>
828
+ Commits: [abc1234 feat(x): ...]
688
829
  Tests: <N passed, M failed>
689
- Tasks: task-1 ✅, task-2 ✅
690
- Result files: <JOBS_ROOT>/<job-name>/results/task-1.json, task-2.json
691
-
692
- Use WAVE_PARTIAL if any task is DONE_WITH_CONCERNS.
693
- Use WAVE_FAILED if any task is BLOCKED or failed.
694
- Do NOT include full task output inline — write details to result files.
830
+ Result file: <JOBS_ROOT>/<job-name>/results/<task_id>.json
831
+
832
+ Write full detail to the result file. Do NOT inline it.
695
833
  })
696
834
  ```
697
835
 
698
- **After all waves, document:**
699
- ```
700
- ACTION: add-document
701
- DATA:
702
- DOC_TYPE: implementation-report
703
- TARGET: both
704
- TITLE: Implementation Report
705
- CONTENT: <summary of all waves, commits, test totals>
706
- AGENT: wave-executor
707
- TASK: Implementation phase
836
+ **Each wave runs in ONE worktree.** The worktree created in 2.4 is the whole job's
837
+ workspace — waves are ordered, not isolated from each other, and a later wave sees
838
+ what an earlier one committed. That is what makes the dependency order mean anything.
839
+
840
+ **After all waves, document:** write the implementation report, then record it:
841
+
842
+ ```bash
843
+ keryx job document <job-name> --type implementation-report --file <path/to/implementation-report.md>
844
+ keryx job step <job-name> implement --status completed
708
845
  ```
709
846
 
847
+ The report summarises every wave: commits, files, test totals, and each task's final
848
+ STATUS.
849
+
710
850
  ### 2.5.1 Post-Implementation Checkpoint
711
851
 
712
852
  After all waves complete, check if tests were created. If not, offer `test-gen`:
713
853
 
714
854
  ```
715
- # Derive all modified files from wave summaries and result files
716
- ALL_FILES = collect from WAVE_RESULTS (read result files for details if needed)
855
+ # Derive all modified files from the per-task result files
856
+ ALL_FILES = collect from <JOBS_ROOT>/<job-name>/results/*.json
717
857
 
718
858
  IF no test files in ALL_FILES:
719
859
  Auto-trigger test-gen for new/modified source files
@@ -736,15 +876,23 @@ What's next?
736
876
  ```
737
877
 
738
878
  **Mapping:**
739
- - A → continue to REVIEW step (default if no response in 60s)
879
+ - A → continue to the REVIEW step (2.6)
740
880
  - B → run `git diff <merge_base>..HEAD --stat` and `git diff <merge_base>..HEAD`, then re-ask
741
- - C → skip REVIEW and FIX steps, go to CHECKS → PR
742
- - D → skip to Phase 3 (COMPLETION) with status "paused"
881
+ - C → skip REVIEW and FIX, go to VERIFY (2.8) → PR. Record both:
882
+ `keryx job step <job-name> review --status skipped --reason "operator chose to skip review"`
883
+ - D → close the open steps with a reason and go to Phase 3:
884
+ `keryx job step <job-name> <step-id> --status skipped --reason "operator stopped here to continue manually"`
743
885
 
744
- ### 2.5.5 Step: IMPLEMENT SANITY CHECK
886
+ **Wait for the answer.** There is no default and no timer: this skill runs as a model
887
+ in a turn-based session, and nothing here can observe wall-clock time passing while a
888
+ user does not reply. A "default after N seconds" could never fire, so it is not
889
+ offered.
890
+
891
+ ### 2.5.2 Step: IMPLEMENT SANITY CHECK
745
892
 
746
893
  Lightweight verification after all waves complete, **before** launching review.
747
- This catches the case where a wave sub-agent claims WAVE_DONE but made no actual git changes.
894
+ This catches the case where a task-implementer reports `STATUS: DONE` but made no
895
+ actual git changes.
748
896
 
749
897
  ```bash
750
898
  # Run in worktree directory
@@ -756,43 +904,78 @@ git log <merge_base>..HEAD --oneline
756
904
 
757
905
  | Check | Pass | Fail action |
758
906
  |-------|------|-------------|
759
- | At least 1 commit exists | ≥1 commit | `retryable` — re-dispatch the failed wave-executor with: "No commits were made. Implement the changes and commit them." |
907
+ | At least 1 commit exists | ≥1 commit | `retryable` — re-dispatch the task-implementers for that wave with: "No commits were made. Implement the changes and commit them." |
760
908
  | At least 1 file modified | ≥1 file changed | Same as above |
761
- | Claimed files actually modified | All files in wave result match diff | Log discrepancy as WARNING, continue |
909
+ | Claimed files actually modified | All files named in the result files appear in the diff | Log discrepancy as a concern, continue |
910
+
911
+ Re-open the step before re-dispatching, so the attempt is counted:
762
912
 
763
- **If retry also produces no commits** → classify as `terminal`, ABORT with:
913
+ ```bash
914
+ keryx job step <job-name> implement --status in-progress
764
915
  ```
765
- "wave-executor returned WAVE_DONE twice but made no git changes.
766
- Please implement manually and re-run from the review step."
916
+
917
+ `metrics.steps[].retries` for `implement` goes up by one. Read it back with
918
+ `keryx job status <job-name> --json` — the count is on disk, so it is still right
919
+ after a session restart.
920
+
921
+ **If the retry also produces no commits** → classify as `terminal` and stop:
922
+
923
+ ```bash
924
+ keryx job step <job-name> sanity-check --status failed --reason "task-implementer reported DONE twice with no git changes"
767
925
  ```
768
926
 
769
- **Record:**
770
927
  ```
771
- SANITY_CHECK:
772
- commits: <count>
773
- files_changed: <count>
774
- lines_added: <N>
775
- lines_removed: <N>
776
- verified: true | false
928
+ "task-implementer returned STATUS: DONE twice but made no git changes.
929
+ Please implement manually and re-run from the review step."
777
930
  ```
778
931
 
932
+ **Record the outcome** in the journal, where it survives the session:
933
+
934
+ ```bash
935
+ keryx job step <job-name> sanity-check --status completed \
936
+ --reason "<N> commits, <M> files changed, +<A>/-<R> lines"
937
+ ```
938
+
939
+ There is no `sanity_check` field in `state.json` and nothing writes one — the
940
+ journal line is the record.
941
+
779
942
  ---
780
943
 
781
944
  ### 2.6 Step: REVIEW
782
945
 
783
- #### 2.6.0 Review Strategy Selection
946
+ `review-orchestrator` is the review path. It is not one strategy among several: it is
947
+ the only entry point that produces a **managed review record**, and every round this
948
+ skill runs is a round that must be citable afterwards. The legacy alternatives —
949
+ launching `code-ai-review` / `code-learned-review` / `code-style-review` by hand, or the
950
+ never-bundled `code-review` 4-agent skill — are gone. They emitted prose into a chat
951
+ transcript and nothing else, which is precisely the failure the managed pipeline
952
+ replaced.
784
953
 
785
- If the user didn't specify a review approach, offer options:
954
+ A pull request driven by this orchestrator has to pass the completion gate shipped in
955
+ 0.2.71. Its five conditions are what 2.6 and 2.7 are built to satisfy:
956
+
957
+ | Gate condition | Satisfied by |
958
+ |---|---|
959
+ | every fix round has a managed record | `keryx review start` before, `keryx review ingest` after (2.6.1, 2.7) |
960
+ | every finding has a terminal disposition | `keryx review complete --finding … --disposition … --evidence …` (2.7) |
961
+ | scope B is recorded when a scope-B reviewer ran | `keryx review blast-radius --json` → `review ingest --blast-radius` (2.6.1) |
962
+ | no inbound PR comment is unanswered | `keryx review comments collect` every round, `… reply --final` once (2.6.2) |
963
+ | verification stats exist | `review-verifier` dispatched, passed as `--verifications` (2.6.1) |
964
+
965
+ #### 2.6.0 Review Scope Selection
966
+
967
+ Ask which reviewer set to use. The flags are `review-orchestrator`'s, and they select
968
+ reviewers — there is no "quick vs thorough" mode:
786
969
 
787
970
  ```
788
- How should I review the implementation?
971
+ Which reviewers should run on this branch?
789
972
 
790
- A) 🚀 Quick (code-review 4-agent parallel) — ~30 sec
791
- B) 📋 Thorough (individual reviewers: ai + boss + style + mobx) — ~2 min
792
- C) 🔒 Security-focused (code-review + security-audit) — ~1 min
793
- D) ⏭ Skip review entirely
973
+ A) Auto-detect from the diff (recommended) — review-orchestrator picks from changed files
974
+ B) Named domains — e.g. --backend --security, --frontend --testing-practices
975
+ C) Everything — --all
976
+ D) Skip review entirely
794
977
 
795
- > pick a letter (default: A)
978
+ > pick a letter
796
979
  ```
797
980
 
798
981
  Then ask which optional convention reviewers to include when local convention docs or matching
@@ -812,107 +995,169 @@ Detected reviewers:
812
995
  - review-flow-graph: shared graph/flow abstraction files
813
996
  ```
814
997
 
815
- Only show detected reviewers. If the user chooses B, ask for the exact skill names to include or
816
- exclude, then persist the choice in job state as `convention_reviewers`.
998
+ Which reviewers are even applicable is **detected, not eyeballed**:
817
999
 
818
- **Auto-select** (skip this question) when:
819
- - `review_mode` is explicitly set in automation settings → use that
820
- - `convention_reviewers` is explicitly set in automation settings → use that for optional convention reviewers
821
- - User already chose at Post-Implementation Checkpoint (2.5.1 option A) → use default (A)
822
- - Time pressure (total_job_timeout close) → use A (fastest)
1000
+ ```bash
1001
+ keryx review stack --json
1002
+ ```
1003
+
1004
+ It reads `package.json` once and every installed review-category skill's declared
1005
+ `metadata.stack_requires`, and reports per reviewer whether the requirement is met.
1006
+ Show only what it includes, and carry its exclusions with their reasons into the
1007
+ report — a reviewer silently absent reads as a reviewer that found nothing.
823
1008
 
824
- #### 2.6.1 Execute Review
1009
+ **Auto-select** (skip these questions) when:
1010
+ - `review_flags` is explicitly set in automation settings → use that
1011
+ - `convention_reviewers` is explicitly set in automation settings → use that for optional convention reviewers
1012
+ - User already chose at Post-Implementation Checkpoint (2.5.1 option A) → use auto-detect (A)
825
1013
 
826
- Dispatch review skills on the whole branch. **Launch all reviewers in parallel** for speed.
1014
+ The selection is held for this run and named in the dispatch. There is no
1015
+ `convention_reviewers` field in `state.json` and nothing writes one; the choice is
1016
+ carried in the dispatch prompt and reported in 2.9.
827
1017
 
828
- **Strategy A — `code-review` (4-agent parallel):**
1018
+ #### 2.6.1 Execute the Round
829
1019
 
830
- Dispatches 4 agents in parallel (correctness, security, performance, style) and produces a unified severity report.
1020
+ **Step 1 — check the budget before dispatching, while stopping is still possible.**
831
1021
 
832
- ```
833
- Launch code-review skill with:
834
- scope: git diff <merge_base>..HEAD
835
- output: unified report with CRITICAL/HIGH/MEDIUM/LOW findings
1022
+ ```bash
1023
+ keryx review budget --spent <usd-so-far> --outstanding <subagents this orchestrator has in flight>
836
1024
  ```
837
1025
 
838
- **Fallback — individual reviewers (if code-review unavailable or user prefers):**
1026
+ `--outstanding` is not optional here. `src/review/caps.ts` names `job-orchestrator`
1027
+ as the outermost of the three nesting levels — `job-orchestrator` →
1028
+ `flow-orchestrator` → `review-orchestrator` — that its cap of 4 in-flight reviewers
1029
+ was chosen to survive. keryx is a CLI invoked once per command; it cannot observe
1030
+ subagents running inside another orchestrator's process. **The cap binds the nested
1031
+ total only when the parent declares its own in-flight count.** Omit `--outstanding`
1032
+ and the cap bounds the reviewer fan-out alone, which the record then states plainly.
839
1033
 
840
- Determine and **dispatch all reviewers simultaneously** (not sequentially):
1034
+ A non-zero exit means the spend ceiling (3 USD by default) is reached: stop and ask
1035
+ the user rather than dispatching another fan-out.
841
1036
 
842
- | Reviewer | Condition | Launch |
843
- |----------|-----------|--------|
844
- | `code-ai-review` | Always | Parallel |
845
- | `code-boss-review` | Always | Parallel |
846
- | `code-style-review` | Always | Parallel |
847
- | `code-mobx-store-review` | Only if `*.store.ts` modified | Parallel |
848
- | `review-frontend-conventions` | If selected and frontend files/local frontend docs match | Parallel |
849
- | `review-testing-practices` | If selected and tests/stories/e2e files match | Parallel |
850
- | `review-core-boundaries` | If selected and shared core files match | Parallel |
851
- | `review-flow-graph` | If selected and shared graph/flow files match | Parallel |
1037
+ **Step 2 — open a managed round.**
852
1038
 
1039
+ ```bash
1040
+ keryx review start --target branch --ref <feature-branch> --head "$(git -C <worktree> rev-parse HEAD)"
1041
+ # reviewing an existing PR instead:
1042
+ keryx review start --target pull-request --ref <pr-number> --head <pr-head-sha>
853
1043
  ```
854
- # Launch ALL applicable reviewers in a SINGLE turn (parallel):
855
- Agent 1: code-ai-review (correctness, security)
856
- Agent 2: code-boss-review (architecture, logic)
857
- Agent 3: code-style-review (naming, patterns)
858
- Agent 4: code-mobx-store-review (if applicable)
859
- Agent 5+: selected convention reviewers (if applicable)
860
1044
 
861
- # Wait for all to complete, then merge results
1045
+ **A fix round is managed, not optional.** A round whose findings were never ingested
1046
+ cannot be cited as a completed round, because nothing durable records what it found.
1047
+
1048
+ **Step 3 — collect inbound PR comments, every round.**
1049
+
1050
+ ```bash
1051
+ keryx review comments collect --repo <owner/repo> --pr <n> --sha <head-sha> \
1052
+ --self <our-login> --round <n> --out <JOBS_ROOT>/<job-name>/comments-r<n>.json
862
1053
  ```
863
1054
 
864
- **Review-orchestrator mode (preferred when available):**
1055
+ `--sha` is required and is the commit collected against; the completion gate compares
1056
+ it to the PR head, so a collection that ran before the comments arrived reads as
1057
+ stale rather than clean. Bot reviewers count as reviewers. Do **not** reply yet —
1058
+ replies happen once, in 2.7, after the last round.
865
1059
 
866
- If `review-orchestrator` exists in the skill catalog, dispatch it with the selected review flags
867
- instead of manually launching individual reviewers. Pass selected convention reviewer flags:
868
- `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`,
869
- and/or `--flow-graph`.
1060
+ **Step 4 — build both scopes.**
870
1061
 
871
- Pass review orchestration controls:
872
- ```
873
- context_mode: <review_context_mode automation setting; default "light", ask "full" for high-risk PRs>
874
- token_budget: <review_token_budget automation setting or computed scope budget>
875
- model_strategy: <review_model_strategy automation setting; default "current">
876
- output: unified report with findings, review_context, token_policy, and model metadata
1062
+ ```bash
1063
+ BASE_SHA="$(git -C <worktree> merge-base HEAD <base_branch>)"
1064
+ keryx review scope --ref "$BASE_SHA" --json > <JOBS_ROOT>/<job-name>/scope.json
1065
+ keryx review blast-radius --ref "$BASE_SHA" --json > <JOBS_ROOT>/<job-name>/blast-radius.json
877
1066
  ```
878
1067
 
879
- **Collect and merge findings:**
880
- ```
881
- REVIEW_FINDINGS: [{
882
- reviewer: "<skill-name>",
883
- findings: [{ file, line, severity: CRITICAL|WARNING|INFO, message }]
884
- }]
1068
+ Scope A (`review scope`) is the bounded diff, with every drop recorded and its reason.
1069
+ Scope B (`review blast-radius`) is what the change can break — the regression set.
1070
+ **Keep both files.** `review ingest --blast-radius <file>` is refused on any round that
1071
+ dispatched `review-regression`, which is every recommended and full round, and an
1072
+ ingest carrying a scope-B finding without the record is refused in code.
1073
+
1074
+ **Step 5 — compute the model per dispatch, never by hand.**
1075
+
1076
+ ```bash
1077
+ keryx review tier --scope <scope> --diff-lines <n> --findings <n> [--security] [--verifier reasoning] --json
885
1078
  ```
886
1079
 
887
- **Strategy C — Security-focused:**
1080
+ Paste the `model` block it prints into that dispatch. `model_strategy: "current"` is
1081
+ gone: it meant "do not switch models", which is exactly the behaviour this command
1082
+ replaced. The command names no model — it ranks what the provider reports at runtime
1083
+ and, when it cannot rank anything, prints `inherit: true`, which means the dispatch
1084
+ runs on the session model. That is a correct answer, not a failure.
1085
+
1086
+ **Step 6 — dispatch `review-orchestrator`.**
888
1087
 
889
- Run `code-review` (4-agent) AND `security-audit` in parallel:
890
1088
  ```
891
- Agent group 1: code-review (correctness, security, performance, style)
892
- Agent group 2: security-audit (dependency vulnerabilities, secrets scan, OWASP patterns)
1089
+ Task({
1090
+ description: "Review round <n>: <job-name>",
1091
+ subagent_type: "general-purpose",
1092
+ prompt: |
1093
+ Load skill: skills/gdskills/review/review-orchestrator/SKILL.md
1094
+
1095
+ flags: <selected flags, e.g. --backend --security --testing-practices>
1096
+ commit_range: <BASE_SHA>..HEAD
1097
+ issue_url: <issue URL, when the job has one — enables the Stage 1 spec gate>
1098
+ context_doc: <JOBS_ROOT>/<job-name>/context_v<N>.md
1099
+ verification_mode: annotate
1100
+ managed_review: { mode: "review-flow", target: "branch", target_ref: "<feature-branch>" }
1101
+ is_fix_round: <true on any round after the first>
1102
+ pr_comments: { enabled: <true when a PR exists> }
1103
+
1104
+ Emit the unified report AND the fenced ```json keryx:findings``` block.
1105
+ Dispatch review-verifier (Wave C) over the consolidated findings and return
1106
+ its verification claims as a file path.
1107
+ })
893
1108
  ```
894
- Merge findings from both into unified `REVIEW_FINDINGS`.
895
1109
 
896
- **Deduplicate:** If multiple reviewers flag the same file:line, merge into a single finding with the highest severity.
1110
+ **Step 7 — verification is part of the round, not an extra.** `review-orchestrator`
1111
+ dispatches `review-verifier` in Wave C over the consolidated findings. The verifier
1112
+ **runs something** and can only delete — it never raises a severity, adds a finding,
1113
+ or rewrites one, and it never verifies a finding raised by the same reviewer. Its
1114
+ claims are merged by the CLI, not by hand.
1115
+
1116
+ **Step 8 — ingest the round.** This is what makes it citable.
1117
+
1118
+ ```bash
1119
+ keryx review ingest --report <path/to/review-report.md> --ref <feature-branch> \
1120
+ --head "$(git -C <worktree> rev-parse HEAD)" \
1121
+ --scope <JOBS_ROOT>/<job-name>/scope.json \
1122
+ --blast-radius <JOBS_ROOT>/<job-name>/blast-radius.json \
1123
+ --verifications <path/to/verifications.json> --verification-mode annotate \
1124
+ --refuted <path/to/refuted.json> \
1125
+ --spent <usd-so-far> --outstanding <subagents in flight>
1126
+ ```
1127
+
1128
+ An unrecognised option is **refused, not ignored** — a silently dropped flag writes
1129
+ nothing and still reports success. `--refuted` carries findings this round raised and
1130
+ then dismissed; without it the package keeps only the survivors of an unlogged triage.
1131
+
1132
+ **Findings are the canonical shape.** One vocabulary, everywhere in this skill:
1133
+
1134
+ - severities are `blocker`, `major`, `minor`, `info` — `review-finding.schema.json`;
1135
+ - the report ends with **exactly one** fenced block whose info string is
1136
+ ` ```json keryx:findings ` — ingest reads that block, not the prose, and a round
1137
+ that emits only prose cannot seed the next one;
1138
+ - `reviewer` is the reviewer that actually produced the finding, never the
1139
+ orchestrator;
1140
+ - identity for dedupe and for the stuck check is `dedupe_key` when the finding has
1141
+ one, otherwise reviewer + file + symbol + problem — never the display id, which is
1142
+ per-report.
897
1143
 
898
1144
  **Classify:**
899
1145
  ```
900
- NEEDS_FIX = count(CRITICAL) > 0 OR count(WARNING) > 0
1146
+ NEEDS_FIX = count(blocker) > 0 OR count(major) > 0
901
1147
  ```
902
1148
 
903
- **Document:**
904
- ```
905
- ACTION: add-document
906
- DATA:
907
- DOC_TYPE: review
908
- TARGET: both
909
- TITLE: Code Review Results
910
- CONTENT: <findings summary for man/, structured findings for ai/>
1149
+ **Document:** record the report in the job package too, so the job and the review
1150
+ record point at each other:
1151
+
1152
+ ```bash
1153
+ keryx job document <job-name> --type review --file <path/to/review-report.md>
911
1154
  ```
912
1155
 
913
1156
  #### 2.6.2 PR Review Report Publication
914
1157
 
915
- If this job is reviewing an existing GitHub PR, or if a PR number/URL was resolved before the review step, ask whether to publish the consolidated review report after review findings are documented and before fix decisions. This gives the user a chance to record the current review state before any automatic fix loop changes it.
1158
+ If this job is reviewing an existing GitHub PR, or a PR number was resolved before the
1159
+ review step, ask whether to publish the consolidated review report — after the round is
1160
+ ingested and before any fix decisions.
916
1161
 
917
1162
  Ask unless automation settings explicitly set `publish_pr_review_report`:
918
1163
 
@@ -928,43 +1173,53 @@ Publish the review report to the PR?
928
1173
 
929
1174
  **Rules:**
930
1175
  - The PR comment and AI artifact must be written in English only, regardless of the chat language or reviewer output language.
931
- - Default is C. Never publish to a PR without explicit user confirmation or `publish_pr_review_report: comment`, `publish_pr_review_report: comment-and-ai-artifact`, or legacy `publish_pr_review_report: true`.
932
- - If the job has review findings but no PR number yet, store `pending_pr_review_report_comment` and `pending_review_ai_artifact` in job state. If the later PR step creates a PR, ask the same question after PR creation.
933
- - If the user chooses A, delegate concise comment formatting to `review-orchestrator`'s PR Review Report Publication contract when available.
934
- - If the user chooses B, delegate concise comment formatting and generate `.metaproject/jobs/<job-name>/ai/review-ai-report.md` using `review-orchestrator`'s Detailed AI Markdown Artifact contract.
935
- - If the user chooses B, the PR comment `Meta` section must include both an `AI artifact` link/path and an `AI artifact description` row explaining in human-readable language that the markdown file contains detailed findings, fix guidance, patch guidance, regression coverage, validation plan, and follow-up agent context.
936
- - If using legacy reviewers, normalize findings into the same concise PR comment and AI artifact structures before posting.
937
- - Record the final decision in job state as `publication_plan.mode`: `comment`, `comment-and-ai-artifact`, or `none`.
1176
+ - Default is C. Never publish to a PR without explicit user confirmation or `publish_pr_review_report: comment`, `publish_pr_review_report: comment-and-ai-artifact`.
1177
+ - If the user chooses A, delegate concise comment formatting to `review-orchestrator`'s PR Review Report Publication contract.
1178
+ - If the user chooses B, also generate `.metaproject/jobs/<job-name>/review-ai-report.md` using `review-orchestrator`'s Detailed AI Markdown Artifact contract, and include in the comment's `Meta` section both an `AI artifact` path and an `AI artifact description` row explaining that the file carries detailed findings, fix guidance, patch guidance, regression coverage, validation plan, and follow-up agent context.
1179
+ - **If no PR exists yet**, do not ask now and do not stash a pending decision — nothing persists one. Ask this question again after the PR step (2.10) creates the PR, when the answer can actually be acted on.
1180
+ - The decision is acted on immediately or not at all. There is no `publication_plan` field in `state.json`; what was published is stated in the 2.9 report.
938
1181
 
939
1182
  **Automation values:**
940
1183
  - `publish_pr_review_report: ask` -> ask the question above.
941
- - `publish_pr_review_report: comment` or legacy `true` -> publish the concise PR comment only.
942
- - `publish_pr_review_report: comment-and-ai-artifact` -> publish the concise PR comment and create/link the detailed AI markdown artifact.
943
- - `publish_pr_review_report: none` or legacy `false` -> do not publish.
1184
+ - `publish_pr_review_report: comment` -> publish the concise PR comment only.
1185
+ - `publish_pr_review_report: comment-and-ai-artifact` -> publish the concise PR comment and create the detailed AI markdown artifact.
1186
+ - `publish_pr_review_report: none` -> do not publish.
944
1187
 
945
1188
  #### 2.6.3 Post-Review Checkpoint
946
1189
 
947
- After review completes, present findings and ask user:
1190
+ After the round is ingested, present findings and ask the user:
948
1191
 
949
1192
  ```
950
- Review complete:
951
- 🔴 <N> CRITICAL 🟠 <M> HIGH 🟡 <K> MEDIUM 🔵 <L> LOW
1193
+ Review round <n> complete:
1194
+ 🔴 <N> blocker 🟠 <M> major 🟡 <K> minor 🔵 <L> info
1195
+ verified: <V> claims recorded, <R> findings refuted
1196
+ inbound PR comments this round: <C>
952
1197
 
953
- A) 🔧 Auto-fix and continue (fix CRITICAL + HIGH, skip LOW)
1198
+ A) 🔧 Auto-fix and continue (fix blocker + major)
954
1199
  B) 📋 Show all findings — I'll decide what to fix
955
1200
  C) ⏭ Skip fixes, proceed to PR as-is
956
1201
  D) ⏹ Stop — I'll fix manually
957
1202
  ```
958
1203
 
1204
+ The counts come from the ingested package, not from re-reading the prose:
1205
+
1206
+ ```bash
1207
+ keryx review status <review-id-or-path>
1208
+ ```
1209
+
959
1210
  **Mapping:**
960
- - A → proceed to FIX step (default if CRITICAL > 0)
1211
+ - A → proceed to the FIX step (2.7)
961
1212
  - B → display all findings grouped by file, then re-ask A/C/D
962
- - C → skip FIX step, go to CHECKS (only if 0 CRITICAL — refuse if CRITICAL > 0)
963
- - D → skip to Phase 3 (COMPLETION) with status "paused"
1213
+ - C → skip FIX, go to VERIFY (2.8) — allowed only when 0 blockers; refuse while a blocker stands
1214
+ - D → close the open steps with a reason (2.1) and go to Phase 3
1215
+
1216
+ Whichever branch is taken, **every finding still needs a disposition** before the
1217
+ review can be completed — see 2.7. "Nobody chose to fix it" is `dismissed-wont-fix`
1218
+ with evidence, not silence.
964
1219
 
965
1220
  **Auto-proceed** (skip this question) when:
966
- - 0 findings → skip directly to CHECKS
967
- - Only INFO findings → skip FIX, go to CHECKS
1221
+ - 0 findings → go straight to VERIFY (2.8)
1222
+ - only `minor`/`info` findings → skip FIX, go to VERIFY (2.8)
968
1223
  - `auto_create_pr: true` → auto-select A
969
1224
 
970
1225
  ### 2.7 Step: FIX (conditional)
@@ -984,43 +1239,95 @@ The bound is a ceiling, not a target. Repetition ends the loop earlier and
984
1239
  slowly" from "stuck", and an agent emitting the identical failing output three
985
1240
  times spends the whole budget before anything notices.
986
1241
 
1242
+ **A finding leaves this loop by being dispositioned, never by being absent.** The
1243
+ previous version of this section recomputed "unresolved" as whatever the next round
1244
+ still reported — so a finding the next reviewer simply did not look at was recorded as
1245
+ fixed. That is absence-as-evidence, and the completion gate refuses it.
1246
+
987
1247
  ```
988
- UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
989
- PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
1248
+ UNRESOLVED_FINDINGS = all blocker + major findings from step 2.6
1249
+ PREVIOUS_REVIEW_OUTPUT = <the ingested report from step 2.6>
990
1250
 
991
1251
  FOR iteration in [1, 2, 3]:
992
1252
  IF NOT NEEDS_FIX: BREAK
993
1253
 
1254
+ keryx job step <job-name> fix --status in-progress # increments metrics.steps[].retries
1255
+
994
1256
  1. Group UNRESOLVED_FINDINGS by file
995
1257
  2. Construct fix prompt — MUST include unresolved findings from previous attempt:
996
1258
 
997
1259
  task_type: "fix"
998
- findings: <UNRESOLVED_FINDINGS>
1260
+ findings: <UNRESOLVED_FINDINGS, in the canonical finding shape>
999
1261
  iteration: <N>
1000
1262
  previously_unresolved: <findings that were in UNRESOLVED_FINDINGS last iteration but still present>
1001
1263
  → Prefix: "These specific findings were NOT fixed in iteration <N-1>: [list]"
1002
1264
 
1003
- 3. Launch task-implementer with fix prompt
1004
- 4. Run sanity-check (step 2.5.5 logic) — verify commits were made
1005
- 5. Re-run reviewers (step 2.6) — parallel dispatch
1006
- 6. Recompute NEEDS_FIX from new findings
1007
- 7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
1008
-
1009
- 8. STUCK CHECK — runs before the next iteration and ignores the budget:
1265
+ 3. Launch task-implementer with the fix prompt (subagent_type: "general-purpose"),
1266
+ on the model `keryx review tier --fix-attempt <N> --findings <n> --json` computes
1267
+ 4. Run the sanity check (step 2.5.2 logic) — verify commits were made
1268
+ 5. Run the next managed round — the FULL 2.6.1 sequence, not a bare re-dispatch:
1269
+ keryx review budget --spent <usd> --outstanding <n>
1270
+ keryx review start --target branch --ref <feature-branch> --head <new-head>
1271
+ keryx review comments collect --repo <r> --pr <n> --sha <new-head> --round <N+1> --out <file>
1272
+ keryx review scope --ref "$BASE_SHA" --json > scope.json
1273
+ keryx review blast-radius --ref "$BASE_SHA" --previous blast-radius.json --json > blast-radius.json
1274
+ <dispatch review-orchestrator with is_fix_round: true>
1275
+ keryx review ingest --report <new-report> --ref <feature-branch> --head <new-head> \
1276
+ --scope scope.json --blast-radius blast-radius.json \
1277
+ --verifications <file> --refuted <file> --outstanding <n>
1278
+ 6. Recompute NEEDS_FIX from the ingested findings
1279
+ 7. Record what became of each finding raised in the PREVIOUS round — every one of
1280
+ them, before the next iteration starts:
1281
+
1282
+ keryx review complete <previous-review-id-or-path> \
1283
+ --finding F-001 --disposition acted-on --evidence "fixed in <commit-sha>" \
1284
+ --finding F-002 --disposition dismissed-incorrect --evidence "<what was run, what it showed>" \
1285
+ --finding F-003 --disposition dismissed-out-of-scope --evidence "<decision, where written>"
1286
+
1287
+ States: unknown, acted-on, dismissed-incorrect, dismissed-wont-fix,
1288
+ dismissed-out-of-scope, dismissed-deprioritised. Everything except `unknown`
1289
+ must cite where the outcome is written down. A recorded state and its citation
1290
+ cannot be overwritten by a later close — record a correction as a new round.
1291
+ Closing with no dispositions leaves every finding reading `unknown`, which means
1292
+ "nobody wrote down what happened".
1293
+ 8. UNRESOLVED_FINDINGS = the blocker + major findings of the NEW round that are
1294
+ still without a terminal disposition
1295
+
1296
+ 9. STUCK CHECK — runs before the next iteration and ignores the budget:
1010
1297
  IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
1011
1298
  OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
1012
1299
  THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
1013
1300
  Identity is the finding's dedupe_key when it has one, otherwise
1014
1301
  reviewer + file + symbol + problem — never the display id, which is
1015
1302
  per-report and would fire on every second iteration whatever happened.
1016
- 9. PREVIOUS_REVIEW_OUTPUT = the new review output
1303
+
1304
+ Detection is also available from the durable record rather than this
1305
+ session's memory, which is the version that survives a restart:
1306
+ keryx review loop --flow <flow-id>
1307
+ It escalates with a non-zero exit on a recurring finding or two identical
1308
+ consecutive rounds, regardless of the remaining budget.
1309
+ 10. PREVIOUS_REVIEW_OUTPUT = the new ingested report
1310
+
1311
+ keryx job step <job-name> fix --status completed
1312
+
1313
+ AFTER THE LAST ROUND ONLY — answer every inbound PR comment, once:
1314
+ keryx review comments reply --repo <owner/repo> --pr <n> --outcomes <file> \
1315
+ --sha <head-sha> --final [--flow-link <url>]
1017
1316
 
1018
1317
  IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1019
1318
  Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
1020
1319
  two ended it — a budget exhausted and a loop detected call for different next
1021
- steps → continue to checks
1320
+ steps. Give every surviving finding a disposition (dismissed-wont-fix or
1321
+ dismissed-deprioritised, with evidence) rather than leaving it `unknown`
1322
+ → continue to VERIFY (2.8)
1022
1323
  ```
1023
1324
 
1325
+ `comments reply` **refuses without `--final`**: replying per round turns one review
1326
+ thread into six, and a reply written mid-loop states an intention rather than an
1327
+ outcome. Each reply is cut in code to 2 sentences and 600 characters, threaded where
1328
+ GitHub gives a thread, capped at 30 with one summary comment for the remainder.
1329
+ `--dry-run` rehearses the whole pass without posting.
1330
+
1024
1331
  **Fix prompt escalation pattern:**
1025
1332
  - Iteration 1: "Fix these findings: [list]"
1026
1333
  - Iteration 2: "These findings were NOT fixed in iteration 1: [subset]. Fix them now."
@@ -1028,14 +1335,15 @@ IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1028
1335
 
1029
1336
  ### 2.8 Step: VERIFY (code-verifier)
1030
1337
 
1031
- Dispatch `code-verifier` as a sub-agent. This replaces the orchestrator-internal "checks" step.
1338
+ Dispatch `code-verifier` as a sub-agent. This is the quality gate; there is no separate
1339
+ `CHECKS` step, and nothing in this document jumps to one.
1032
1340
 
1033
1341
  ```
1034
1342
  Task({
1035
1343
  description: "Quality gate: <job-name>",
1036
- subagent_type: "general",
1344
+ subagent_type: "general-purpose",
1037
1345
  prompt: |
1038
- You are code-verifier. Load skill: skills/code-verifier/SKILL.md
1346
+ You are code-verifier. Load skill: skills/gdskills/orchestration/code-verifier/SKILL.md
1039
1347
 
1040
1348
  codebase_path: <worktree_path>
1041
1349
  base_branch: <base_branch>
@@ -1049,24 +1357,22 @@ Task({
1049
1357
  ```
1050
1358
  IF VERIFICATION_RESULT.gate == "PASS" or "PASS_WITH_WARNINGS":
1051
1359
  → Proceed to review
1052
- → Log findings as informational in job docs
1360
+ → Log findings as informational in the job report
1053
1361
 
1054
1362
  IF VERIFICATION_RESULT.gate == "FAIL":
1055
- → Extract CRITICAL/HIGH findings
1056
- → Check if fix step is already scheduled
1057
- - If not → add fix step to plan (dispatch task-implementer in fix mode)
1058
- - If fix already ran 2× → escalate to user, skip to report
1363
+ → Extract blocker/major findings
1364
+ → Check whether the fix step has already run:
1365
+ keryx job status <job-name> --json # retries["fix"] is the recorded count
1366
+ - If it has not → run the fix step (2.7) with these findings
1367
+ - If `retries["fix"]` has reached 3 → escalate to the user and go to report.
1368
+ Three is the bound, and it is the same three everywhere in this skill.
1059
1369
  ```
1060
1370
 
1061
- **Document result:**
1062
- ```
1063
- ACTION: add-document
1064
- DATA:
1065
- DOC_TYPE: verification-report
1066
- TARGET: both
1067
- TITLE: Verification Report — <gate status>
1068
- CONTENT: <VERIFICATION_RESULT formatted>
1069
- AGENT: code-verifier
1371
+ **Document result:** write the verification report, then record it:
1372
+
1373
+ ```bash
1374
+ keryx job document <job-name> --type verification-report --file <path/to/verification-report.md>
1375
+ keryx job step <job-name> verify --status completed --reason "gate: <PASS|PASS_WITH_WARNINGS|FAIL>"
1070
1376
  ```
1071
1377
 
1072
1378
  ### 2.8.1 Step: VERIFY-POST-FIX (code-verifier, conditional)
@@ -1075,14 +1381,20 @@ After fix iterations, dispatch `code-verifier` again with identical parameters.
1075
1381
 
1076
1382
  ```
1077
1383
  IF fix ran:
1384
+ keryx job step <job-name> verify-post-fix --status in-progress
1078
1385
  Dispatch code-verifier (same params as step 2.8)
1079
1386
  IF gate still FAIL:
1080
- Log "Verification failed after fix" → skip to report with warning
1387
+ Log "Verification failed after fix" → go to report with a warning
1081
1388
  IF gate PASS:
1082
1389
  Proceed to report
1390
+ keryx job document <job-name> --type verification-report --file <path/to/verification-post-fix.md>
1391
+ keryx job step <job-name> verify-post-fix --status completed --reason "gate: <status>"
1392
+
1393
+ IF fix did not run:
1394
+ keryx job step <job-name> verify-post-fix --status skipped --reason "no fix round was needed"
1083
1395
  ```
1084
1396
 
1085
- ### 2.8.1 Step: PERF-CHECK (optional)
1397
+ ### 2.8.2 Step: PERF-CHECK (optional)
1086
1398
 
1087
1399
  Auto-trigger `perf-check` when frontend/bundle files were modified:
1088
1400
 
@@ -1094,7 +1406,42 @@ IF any modified file matches: *.tsx, *.jsx, *.css, *.scss, webpack.*, vite.*, ne
1094
1406
  Add findings to report (informational, not blocking)
1095
1407
  ```
1096
1408
 
1097
- Skip if no frontend files changed or no build output exists. Results are advisory — they don't block the PR.
1409
+ Skip if no frontend files changed or no build output exists. Results are advisory — they don't block the PR. Either way the step is closed on the record:
1410
+
1411
+ ```bash
1412
+ keryx job step <job-name> perf-check --status completed|skipped --reason "<result or why it did not run>"
1413
+ ```
1414
+
1415
+ ### 2.8.3 Step: SKILL LEARNING (conditional)
1416
+
1417
+ Close the self-learning loop (see `rules/core/skill-lifecycle.mdc`). Collect the
1418
+ learning signals produced upstream:
1419
+ - `skill_drift` fields from each task-implementer result (`stale:`/`missing:`).
1420
+ - the `## Skill Learning` block from `review-orchestrator`.
1421
+
1422
+ ```
1423
+ IF no skill_drift and Skill Learning == none:
1424
+ → skip this step (log "no skill drift")
1425
+
1426
+ ELSE for each flagged project-skill:
1427
+ 1. Dispatch a subagent to build the learning proposal:
1428
+ - Model: COMPUTED, not chosen — run
1429
+ keryx review tier --scope narrow --json
1430
+ and paste the `model` block into the dispatch. The command names no model:
1431
+ it ranks what the provider reports at runtime, and when it cannot rank
1432
+ anything it prints `inherit: true`, which means the dispatch runs on the
1433
+ session model. See rules/core/model-selection.mdc for what the tiers mean.
1434
+ - Command: keryx skills learn --from-review <review-report-path> \
1435
+ --skill <module>/<skill>
1436
+ (or --from-test / --from-failure when the signal came from verification)
1437
+ - The subagent returns the proposal path. It does NOT apply.
1438
+ 2. The orchestrator (flagship) reads the proposal and either:
1439
+ - keryx skills learn apply <proposal.json> (accept), or
1440
+ - discards it and notes why in the report.
1441
+ ```
1442
+
1443
+ Never apply a proposal unread, and never run `learn` in a hook. Record applied
1444
+ skill updates in the Job Report under "Skill Updates".
1098
1445
 
1099
1446
  ### 2.9 Step: REPORT
1100
1447
 
@@ -1109,7 +1456,7 @@ Aggregate all information into a human-readable summary.
1109
1456
  - **Source:** <issue URL or description>
1110
1457
  - **Branch:** `<branch_name>`
1111
1458
  - **Tasks:** <completed>/<total> completed
1112
- - **Review Iterations:** <N>
1459
+ - **Review Rounds:** <N> (managed records: <review-id list>)
1113
1460
  - **Final Status:** <READY FOR PR | HAS WARNINGS | HAS ISSUES | ANALYSIS ONLY>
1114
1461
 
1115
1462
  ## Analysis
@@ -1122,19 +1469,27 @@ Aggregate all information into a human-readable summary.
1122
1469
  - **Commits:** <hashes>
1123
1470
 
1124
1471
  ## Review Results
1125
- ### code-ai-review
1126
- - CRITICAL: <N>, WARNING: <N>, INFO: <N>
1127
- ### code-boss-review
1128
- - ...
1472
+ Round <n> — `.metaproject/reviews/<review-id>/`
1473
+ | Reviewer | blocker | major | minor | info |
1474
+ |---|---|---|---|---|
1475
+ | review-logic | <N> | <N> | <N> | <N> |
1476
+ | … | | | | |
1477
+
1478
+ Verification: <V> claims recorded, <R> findings refuted, <U> unverified.
1479
+ Reviewers excluded by `keryx review stack`: <name — reason>.
1480
+ Inbound PR comments: <C> collected, <A> answered in the final reply pass.
1129
1481
 
1130
1482
  ## Unresolved Issues
1131
- - [ ] <file>:<line> — <message> (from <reviewer>)
1483
+ - [ ] <file>:<line> — <message> (from <reviewer>, disposition `<state>`, evidence `<ref>`)
1132
1484
 
1133
1485
  ## Final Checks
1134
1486
  - Lint: PASS
1135
1487
  - Type Check: PASS
1136
1488
  - Tests: 42 passed, 0 failed
1137
1489
 
1490
+ ## Skill Updates
1491
+ - `<module>/<skill>` v1.2.0 → v1.3.0 (from review F-012; applied) | none
1492
+
1138
1493
  ## Changes Summary
1139
1494
  ### Files Modified (<N>)
1140
1495
  - `src/...`
@@ -1159,7 +1514,7 @@ JOB_NAME: <job-name>
1159
1514
  BRANCH: <feature_branch>
1160
1515
  BASE: <base_branch>
1161
1516
  ISSUE_NUMBER: <issue_number if available>
1162
- CONTEXT_PATH: <JOBS_ROOT>/<job-name>/ai/context.md
1517
+ CONTEXT_PATH: <JOBS_ROOT>/<job-name>/context_v<N>.md
1163
1518
  ```
1164
1519
 
1165
1520
  `pr-issue-documenter` will analyze the branch diff and produce a structured PR description (Summary + Changes by area + Key Files table). Use its output as the `body` for the PR.
@@ -1197,42 +1552,66 @@ EOF
1197
1552
  )" --base <base_branch> --head <feature_branch> --draft
1198
1553
  ```
1199
1554
 
1555
+ Then record the step, with the PR on the record:
1556
+
1557
+ ```bash
1558
+ keryx job step <job-name> pr --status completed --reason "<PR URL>"
1559
+ ```
1560
+
1561
+ If the job had review findings but no PR until now, ask the 2.6.2 publication
1562
+ question here — this is the point at which it can be acted on.
1563
+
1200
1564
  ---
1201
1565
 
1202
1566
  ## Phase 3: COMPLETION
1203
1567
 
1204
- ### 3.1 Finalize Job Documentation
1568
+ ### 3.1 Close the Job Package
1205
1569
 
1206
- Dispatch job-documenter with `finalize` action:
1570
+ ```bash
1571
+ keryx job complete <job-name>
1572
+ ```
1207
1573
 
1574
+ This is a **gate, not a formality.** It refuses while any step is still open or
1575
+ `failed`, and the refusal names them:
1576
+
1577
+ ```
1578
+ Cannot complete job <name> — 12/15 steps terminal (not terminal: perf-check, deploy; failed: fix).
1579
+ Close each with: keryx job step <name> <step-id> --status completed|skipped [--reason "<why>"]
1208
1580
  ```
1209
- ACTION: finalize
1210
- DATA:
1211
- FINAL_CONTENT: <full report markdown>
1212
- FINAL_STATUS: completed | aborted
1213
- SUMMARY: <1-3 sentence summary>
1581
+
1582
+ So close every remaining step first, with a reason that says what happened:
1583
+
1584
+ ```bash
1585
+ keryx job step <job-name> deploy --status skipped --reason "user declined the staging deploy"
1214
1586
  ```
1215
1587
 
1216
- **Validate response:** status must be `success`.
1588
+ A job that ended badly is closed the same way — each unfinished step recorded as
1589
+ `skipped` with the reason it stopped. There is no "aborted" status to set: what
1590
+ happened is in the step statuses and in `journal.md`, which is a record, not a label.
1591
+
1592
+ On success the package moves to `phase: COMPLETION` and `plan.current_step` is
1593
+ cleared, so 0.0 will no longer offer it for resumption.
1217
1594
 
1218
1595
  ### 3.2 Present Results
1219
1596
 
1220
1597
  Tell user:
1221
1598
  1. What was accomplished (summary)
1222
- 2. Where documentation is stored: `.metaproject/jobs/<job-name>/`
1599
+ 2. Where the package is: `.metaproject/jobs/<job-name>/`
1223
1600
  3. PR URL (if created)
1224
- 4. Metrics summary (time, tokens)
1225
- 5. Any unresolved issues
1601
+ 4. Step durations and retries, read from the package
1602
+ 5. Any unresolved issues, each with its recorded disposition
1226
1603
 
1227
1604
  ```
1228
1605
  ✅ Job completed successfully.
1229
1606
 
1230
- Documentation: <JOBS_ROOT>/<job-name>/
1231
- Branch: feature/<slug> (worktree: <path>)
1232
- PR: <URL or "not created">
1233
- Metrics: <total time>, <total tokens>
1234
-
1235
- See .metaproject/jobs/<job-name>/README.md for the full job index.
1607
+ Package: <JOBS_ROOT>/<job-name>/
1608
+ Branch: feature/<slug> (worktree: <path>)
1609
+ PR: <URL or "not created">
1610
+ Review: <N> managed rounds, .metaproject/reviews/<review-id>/
1611
+ Steps: <done>/<total>, retries <sum>
1612
+
1613
+ keryx job status <job-name> — the step list, retries and recorded documents
1614
+ <JOBS_ROOT>/<job-name>/journal.md — every recorded event, in order
1236
1615
  ```
1237
1616
 
1238
1617
  ### 3.3 Post-Completion Options
@@ -1259,99 +1638,124 @@ What would you like to do next?
1259
1638
 
1260
1639
  When the orchestrator starts with an `analyze` intent and the user then says "yes, implement":
1261
1640
 
1262
- 1. **Keep existing completed steps** (analyze, context, report are already done)
1263
- 2. **Extend plan** with new steps: prepare → implement → review → fix → checks → report → pr
1264
- 3. **Update job documentation** via job-documenter (update-readme with new plan)
1265
- 4. **Continue execution** from the first new step
1641
+ 1. **Keep the existing package** — its completed steps (analyze, context, report) stay
1642
+ completed and stay on the record.
1643
+ 2. **Create the implementation package** and run it as an `implement` job:
1266
1644
 
1267
- This is the core of dynamic planning — the plan grows based on user decisions.
1645
+ ```bash
1646
+ keryx job init --name <analysis-job-name>-impl --intent implement --project <project_dir>
1647
+ ```
1648
+
1649
+ `keryx job` does not rewrite a package's plan after `init`, and this skill does not
1650
+ ask it to: a plan that could be rewritten in place is a plan whose recorded history
1651
+ cannot be trusted. The two packages are linked by naming and by a journal line:
1652
+
1653
+ ```bash
1654
+ keryx job step <analysis-job-name> proposal --status completed \
1655
+ --reason "user accepted; implementation continues in job <analysis-job-name>-impl"
1656
+ ```
1657
+ 3. **Complete the analysis job** (`keryx job complete <analysis-job-name>`) once its
1658
+ steps are closed, so it stops being offered for resumption in 0.0.
1659
+ 4. **Continue execution** from Phase 1.3 of the new package.
1660
+
1661
+ This is the core of dynamic planning — the work grows based on user decisions, and each
1662
+ stage keeps its own auditable package rather than one package quietly changing shape.
1268
1663
 
1269
1664
  ---
1270
1665
 
1271
1666
  ## State Management
1272
1667
 
1273
- The orchestrator maintains state throughout all phases:
1668
+ There are two kinds of state, and confusing them is how a job loses its record.
1669
+
1670
+ **Persisted — written by `keryx job`, survives the session.** This is exactly what
1671
+ `state.schema.json` declares and exactly what the six commands write. The root carries
1672
+ `additionalProperties: false`, so a field that is not on this list cannot be stored:
1274
1673
 
1275
1674
  ```
1276
- JOB_STATE:
1277
- phase: CONTEXT | PLAN | EXECUTION | COMPLETION
1278
- intent: implement | analyze | review | custom
1279
- create_pr: <bool>
1280
- job_name: <string>
1281
-
1675
+ state.json:
1676
+ phase: CONTEXT | PLAN | EXECUTION | COMPLETION (job init, job step, job complete)
1677
+ intent: implement | analyze | review | custom (job init)
1678
+ job_name: <slug matching ^[a-z0-9-]+$> (job init)
1679
+ create_pr: <bool>
1282
1680
  context:
1283
- issue: { number, title, url, type }
1284
- project_dir: <path>
1285
- base_branch: <string>
1286
-
1287
- branch:
1288
- name: <string>
1289
- worktree_path: <path>
1290
- merge_base: <commit hash>
1291
-
1681
+ project_dir: <path> (job init --project)
1682
+ base_branch: <string>
1683
+ issue: { number, title, url, type }
1292
1684
  plan:
1293
- steps: [{ id, type, agent, depends, status: pending|in_progress|completed|skipped|failed, prompt_chars: <int>, prompt_hash: <sha256 first 8 chars> }]
1294
- current_step: <step_id>
1295
-
1296
- analysis:
1297
- total_tasks: <N>
1298
- tasks: [<task objects>]
1299
- dependency_order: [<task_ids>]
1300
-
1301
- context_doc:
1302
- path: <JOBS_ROOT>/<job-name>/ai/context.md
1303
- version: <current version>
1304
- status: collected | updated | not-collected
1305
-
1306
- implementation:
1307
- task_results: {<task_id>: <result>}
1308
- all_commits: [<hash>]
1309
- all_files: [<path>]
1310
-
1311
- review:
1312
- iteration: <N>
1313
- findings: [<findings>]
1314
- needs_fix: <bool>
1315
- unresolved: [<findings>]
1316
-
1317
- final_checks:
1318
- lint: <result>
1319
- type_check: <result>
1320
- tests: <result>
1321
-
1685
+ steps: [{ id, type, agent, depends, conditional,
1686
+ status: pending|in_progress|completed|skipped|failed }] (job step)
1687
+ current_step: <first step that is not terminal> (maintained by job step)
1322
1688
  documentation:
1323
- job_path: <JOBS_ROOT>/<job-name>
1324
- documents_created: [<paths>]
1689
+ job_path: .metaproject/jobs/<job-name>
1690
+ documents_created: [<file name per recorded document>] (job document)
1691
+ metrics:
1692
+ steps: [{ step_id, status, started_at, completed_at, duration_ms, retries }] (job step)
1693
+ jobs_root: .metaproject/jobs
1694
+ updated_at: <ISO 8601, stamped on every write>
1695
+ ```
1696
+
1697
+ `journal.md` sits beside it: append-only, one timestamped line per event, with the
1698
+ `--reason` text where one was given. Between the two, "what happened to this job" is
1699
+ answerable without this session.
1700
+
1701
+ **In-session — held by the orchestrator for this run, and NOT persisted.** Say it in
1702
+ the dispatch prompt, or it does not reach the sub-agent:
1703
+
1325
1704
  ```
1705
+ branch: { name, worktree_path, merge_base, package_manager, run_command }
1706
+ analysis: { total_tasks, tasks, dependency_order }
1707
+ context_doc: the path to the highest-numbered context_v<N>.md in the package
1708
+ review: the current round's findings — the durable copy is the managed review
1709
+ package, not this
1710
+ ```
1711
+
1712
+ Five fields this skill used to claim it recorded — `sanity_check`,
1713
+ `convention_reviewers`, `publication_plan.mode`, `pending_pr_review_report_comment`,
1714
+ `pending_review_ai_artifact` — are **not** persisted and are not in the schema. Nor is
1715
+ a `paused` or `timeout` status. Nothing writes them, so nothing claims them: what would
1716
+ have gone into them goes into a `--reason` on the journal, or into the 2.9 report.
1326
1717
 
1327
1718
  ---
1328
1719
 
1329
1720
  ## state.json Specification
1330
1721
 
1331
- The orchestrator persists JOB_STATE to `.metaproject/jobs/<job-name>/state.json` for job resumption.
1722
+ **Location:** `.metaproject/jobs/<job-name>/state.json`
1332
1723
 
1333
- **Location:** `.metaproject/jobs/<JOB_NAME>/state.json`
1724
+ **Schema:** `skills/gdskills/orchestration/job-orchestrator/state.schema.json`, registered
1725
+ as the contract `job-orchestrator-state`.
1334
1726
 
1335
- **Schema reference:** `skills/job-orchestrator/state.schema.json`
1727
+ **Who writes it:** `keryx job`, and nothing else. Every write is validated against the
1728
+ registered contract first and a non-conforming state is **refused**, not written:
1336
1729
 
1337
- **When to create:** During Phase 1.2 (Initialize Job Documentation) — write initial state after job docs are initialized.
1730
+ ```
1731
+ Refusing to write .metaproject/jobs/<name>/state.json — it does not validate against
1732
+ job-orchestrator-state:
1733
+ - /plan/steps/0/status: must be one of pending, in_progress, completed, skipped, failed
1734
+ ```
1338
1735
 
1339
- **When to update:** After every step completion in Phase 2 (EXECUTION) — update `plan.steps[i].status`, `plan.steps[i].prompt` (store the prompt used), and `plan.current_step`.
1736
+ **Do not hand-write it.** No `cat > state.json`, no `jq` edit, no sub-agent writing it
1737
+ directly. A hand-written state bypasses the validation and the journal, which is how a
1738
+ package ends up describing a job that did not happen.
1739
+
1740
+ Validate any state file against the contract directly if you need to:
1340
1741
 
1341
- **How to write state.json:**
1342
1742
  ```bash
1343
- # Write state (orchestrator handles this directly, not via job-documenter)
1344
- cat > .metaproject/jobs/<JOB_NAME>/state.json << 'EOF'
1345
- {
1346
- "phase": "EXECUTION",
1347
- "intent": "<intent>",
1348
- "job_name": "<job-name>",
1349
- ...
1350
- }
1351
- EOF
1743
+ keryx skills contracts validate .metaproject/jobs/<job-name>/state.json --schema job-orchestrator-state
1352
1744
  ```
1353
1745
 
1354
- **Job resumption (Phase 0.0):** If `state.json` exists and `phase` is not `COMPLETION`, offer to resume. Parse the file, restore JOB_STATE, jump to the first step with `status: "pending"` or `status: "in_progress"`.
1746
+ **When it is written:**
1747
+
1748
+ | Command | What it changes |
1749
+ |---|---|
1750
+ | `keryx job init` | creates the package, the plan, `phase: PLAN` |
1751
+ | `keryx job step` | a step's status, `plan.current_step`, `metrics.steps[]` (including `retries`), `phase: EXECUTION` |
1752
+ | `keryx job document` | `documentation.documents_created`, and copies the file in |
1753
+ | `keryx job complete` | `phase: COMPLETION`, clears `plan.current_step` — refused unless every step is terminal |
1754
+
1755
+ **Job resumption (Phase 0.0):** `keryx job list --json` finds packages whose `phase` is
1756
+ not `COMPLETION`; `keryx job status <name> --json` names `next_step` — the first step
1757
+ that is neither `completed` nor `skipped`. Both answers are computed from the file, so
1758
+ a resumed session does not depend on remembering where it was.
1355
1759
 
1356
1760
  ---
1357
1761
 
@@ -1372,15 +1776,16 @@ Do not attempt to infer status from prose. Do not trust a response that "looks f
1372
1776
  **`STATUS: DONE`**
1373
1777
  - Accept result.
1374
1778
  - Extract structured payload (JSON result, files changed, commits, verification results).
1375
- - Mark step as completed in JOB_STATE.
1779
+ - Record it: `keryx job step <job-name> <step-id> --status completed`.
1376
1780
  - Continue to next step in the plan.
1377
1781
 
1378
1782
  **`STATUS: DONE_WITH_CONCERNS`**
1379
1783
  - Accept result as complete.
1380
1784
  - Read the `## Concerns for orchestrator` section carefully.
1381
1785
  - Decide: (a) log concern and continue, (b) surface concern to user at next checkpoint, or (c) re-dispatch with adjusted scope if the concern affects correctness.
1382
- - Do NOT silently discard concerns. Record them in JOB_STATE and include in the final report.
1383
- - Mark step as completed.
1786
+ - Do NOT silently discard concerns. Put them on the record and include them in the final report:
1787
+ `keryx job step <job-name> <step-id> --status completed --reason "<the concern>"`
1788
+ — the reason lands in `journal.md`, so the concern outlives the session.
1384
1789
 
1385
1790
  **`STATUS: BLOCKED`**
1386
1791
  - Do NOT proceed to any step that depends on this task.
@@ -1417,7 +1822,7 @@ Use this structure for every subagent dispatch:
1417
1822
  ```
1418
1823
  Task({
1419
1824
  description: "<one-line summary for logs>",
1420
- subagent_type: "general",
1825
+ subagent_type: "general-purpose",
1421
1826
  prompt: |
1422
1827
  ## Task
1423
1828
  <Exactly what to do — no ambiguity>
@@ -1439,6 +1844,10 @@ Task({
1439
1844
  })
1440
1845
  ```
1441
1846
 
1847
+ `subagent_type` is **`general-purpose`**. That is the dispatcher's own name for a
1848
+ general agent; `"general"` is not a value any dispatcher accepts, and a dispatch
1849
+ carrying it does not run.
1850
+
1442
1851
  ### Minimality principle
1443
1852
 
1444
1853
  Pass only what the subagent needs for this specific task. Do not dump job state, full analysis JSON, or conversation history. Extraneous context fills the subagent's context window with noise and increases hallucination risk.
@@ -1463,27 +1872,30 @@ The subagent must not fetch orchestrator state independently. If the subagent ne
1463
1872
 
1464
1873
  | Setting | Default | Options | Description |
1465
1874
  |---------|---------|---------|-------------|
1466
- | `skip_confirmation` | `true` | true/false | Skip confirmation for sub-agents |
1467
- | `base_branch` | auto-detect | any | Base branch (auto-detect from repo default, or ask user) |
1468
- | `max_review_iterations` | `3` | 1-5 | Max review → fix iterations |
1875
+ | `skip_confirmation` | `true` | `true` only | Sub-agents run without per-dispatch confirmation. `{"const": true}` in the input contract. Does **not** cover the 0.4 operator gate — that one is `plan_approval`. |
1876
+ | `base_branch` | auto-detect | any | Base branch (auto-detect from repo default, or ask user). No default in the contract. |
1877
+ | `max_review_iterations` | `3` | 1-3 | Max review → fix iterations. Three everywhere: this table, 2.7, and the input contract's `maximum` and `default`. |
1469
1878
  | `create_pr` | `true` | true/false | Whether to propose PR at the end |
1470
1879
  | `auto_create_pr` | `false` | true/false | Auto-create PR without asking |
1471
- | `review_mode` | `"code-review"` | `"code-review"` / `"individual"` | Use 4-agent parallel or individual reviewers |
1472
- | `reviewers` | `["code-ai-review", "code-boss-review", "code-style-review"]` | skill names | Individual reviewers (when review_mode=individual) |
1473
- | `conditional_reviewers` | `{"code-mobx-store-review": "*.store.ts"}` | skill→pattern | Conditional reviewers |
1880
+ | `review_flags` | auto-detect | `review-orchestrator` flags | Reviewer selection passed to `review-orchestrator` (e.g. `--backend --security`). Unset means auto-detect from the diff. |
1474
1881
  | `convention_reviewers` | `"ask"` | `"ask"` / `"all"` / `"none"` / skill names | Optional convention reviewers to include in review |
1882
+ | `verification_mode` | `annotate` | `off`/`annotate`/`filter` | Passed to `review-orchestrator` and to `review ingest --verification-mode` |
1475
1883
  | `run_final_checks` | `true` | true/false | Run lint/type-check/test |
1476
1884
  | `run_interview` | `true` | true/false | Run interview skill in Phase 0 |
1477
1885
  | `dry_run` | `false` | true/false | Plan-only mode: full Phase 0+1, no agent dispatch or git ops |
1478
- | `log_prompt_sizes` | `true` | true/false | Store prompt char count per step in state.json for observability |
1479
1886
  | `plan_approval` | `true` | true/false | Show agent plan and ask approve/adjust before execution (1.3) |
1480
1887
  | `run_test_gen` | `true` | true/false | Auto-run test-gen if implementer skips tests |
1481
1888
  | `run_security_audit` | `true` | true/false | Auto-run security-audit if auth/API/DB files touched |
1482
1889
  | `run_perf_check` | `true` | true/false | Auto-run perf-check if frontend/bundle files changed |
1483
1890
  | `run_changelog` | `true` | true/false | Auto-generate changelog entry and include in PR description |
1484
- | `publish_pr_review_report` | `ask` | `ask`/`comment`/`comment-and-ai-artifact`/`none`/`true`/`false` | Whether to publish a concise PR review comment and optional detailed AI markdown artifact |
1891
+ | `publish_pr_review_report` | `ask` | `ask`/`comment`/`comment-and-ai-artifact`/`none` | Whether to publish a concise PR review comment and optional detailed AI markdown artifact |
1485
1892
  | `run_deploy` | `ask` | `ask`/`true`/`false` | Post-PR deploy: ask user (ask), always deploy (true), never (false) |
1486
1893
 
1894
+ `review_mode` is gone. It defaulted to `"code-review"` — a skill that is not bundled
1895
+ and not catalogued — and its `"individual"` alternative named the legacy hand-dispatch
1896
+ path 2.6 replaced. Reviewer selection is `review_flags`, and the reviewers are
1897
+ `review-orchestrator`'s.
1898
+
1487
1899
  ## Dry-Run Mode
1488
1900
 
1489
1901
  When `dry_run: true` is set (or `--dry-run` is passed):
@@ -1496,16 +1908,21 @@ When `dry_run: true` is set (or `--dry-run` is passed):
1496
1908
  ```
1497
1909
  Dry-run plan for: issue-4141--pipeline-validation
1498
1910
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
1499
- Step 1: analyze [issue-analyzer] → input: issue #4141
1500
- Step 2: context [context-collector] → input: analysis result, project_dir
1501
- Step 3: prepare [orchestrator] → creates: feature/pipeline-validation
1502
- Step 4: implement [task-implementer × 3] → sequential, 3 tasks
1503
- Step 5: sanity-check [orchestrator] → verifies commits exist
1504
- Step 6: review [code-review × 4] → parallel
1505
- Step 7: fix [task-implementer] → conditional: if NEEDS_FIX
1506
- Step 8: checks [orchestrator] → lint + type-check + test
1507
- Step 9: report [orchestrator] → aggregates all results
1508
- Step 10: pr [orchestrator + gh CLI] → conditional: if create_pr
1911
+ Step 1: analyze [issue-analyzer] → input: issue #4141
1912
+ Step 2: context [context-collector] → input: analysis result, project_dir
1913
+ Step 3: prepare [orchestrator] → creates: feature/pipeline-validation worktree
1914
+ Step 4: tests-creator [tests-creator × 3] → RED stubs, one per task
1915
+ Step 5: implement [task-implementer × 3] → wave-parallel, 3 tasks
1916
+ Step 6: sanity-check [orchestrator] → verifies commits exist
1917
+ Step 7: verify [code-verifier] → lint + type-check + test + imports
1918
+ Step 8: review [review-orchestrator] → managed round, ingested
1919
+ Step 9: security [security-audit] → conditional: auth/API/DB/env files
1920
+ Step 10: fix [task-implementer] → conditional: if NEEDS_FIX
1921
+ Step 11: verify-post-fix [code-verifier] → conditional: after fix
1922
+ Step 12: perf-check [perf-check] → conditional: frontend/bundle files
1923
+ Step 13: report [orchestrator] → aggregates all results
1924
+ Step 14: pr [orchestrator + gh CLI] → conditional: if create_pr
1925
+ Step 15: deploy [deploy] → conditional: if user asks
1509
1926
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
1510
1927
  Estimated sub-agent calls: 11-14 (varies with tasks and review findings)
1511
1928
  No changes will be made. Use without --dry-run to execute.
@@ -1513,21 +1930,27 @@ No changes will be made. Use without --dry-run to execute.
1513
1930
 
1514
1931
  5. Ask user: "Execute this plan? (yes / adjust / abort)"
1515
1932
 
1516
- ## Budget Guards & Timeouts
1933
+ ## Budget Guards
1934
+
1935
+ **There are no timeouts, because this execution model has no clock.** This skill runs
1936
+ as a model inside a turn-based session: it cannot observe wall-clock time passing, it
1937
+ cannot kill a sub-agent mid-flight, and it has no persisted start time to measure
1938
+ against. A `step_timeout_ms` that "kills the agent if exceeded" was a guard nothing
1939
+ could ever enforce, and a job could not end with a `timeout` status because no such
1940
+ status exists in `state.schema.json`.
1517
1941
 
1518
- The orchestrator enforces resource limits to prevent runaway sub-agents:
1942
+ What actually bounds this orchestrator:
1519
1943
 
1520
- | Guard | Default | Description |
1521
- |-------|---------|-------------|
1522
- | `step_timeout_ms` | `300000` (5 min) | Max time per step. Kill agent if exceeded. |
1523
- | `implementation_timeout_ms` | `600000` (10 min) | Max time for full implementation phase |
1524
- | `total_job_timeout_ms` | `1800000` (30 min) | Max time for entire job. Abort to Phase 3 if exceeded. |
1525
- | `max_retries_per_step` | `2` | Max retries for a failed step before asking user |
1944
+ | Guard | Bound | Where it is enforced |
1945
+ |-------|-------|----------------------|
1946
+ | review → fix rounds | 3 | 2.7, and `max_review_iterations` (`maximum: 3`) in the input contract |
1947
+ | repetition, whatever the count says | first repeat | the STUCK CHECK in 2.7, and `keryx review loop` against the durable record |
1948
+ | retries per step | recorded, not guessed | `metrics.steps[].retries`, incremented by `keryx job step --status in-progress` and read back with `keryx job status --json` |
1949
+ | reviewer fan-out | 4 in flight | `keryx review budget --outstanding <n>` before every dispatch (2.6.1) |
1950
+ | spend | 3 USD by default | `keryx review budget --spent <usd>` — a non-zero exit means stop and ask |
1526
1951
 
1527
- **Timeout behavior:**
1528
- - When a step times out → mark as `failed`, record partial results if any
1529
- - Ask user: "Step X timed out after Y minutes. Retry / Skip / Abort?"
1530
- - If total job timeout → force transition to Phase 3 (COMPLETION) with status "timeout"
1952
+ Each of these is a number some command reads or writes. A guard that no command can
1953
+ observe is not a guard, and this section no longer lists any.
1531
1954
 
1532
1955
  **Context passing rules (minimal context principle):**
1533
1956
  - `issue-analyzer`: receives only issue data + codebase paths (NOT previous job state)
@@ -1544,7 +1967,7 @@ Each step failure is classified into one of three classes with different recover
1544
1967
  | Class | Meaning | Action |
1545
1968
  |-------|---------|--------|
1546
1969
  | `terminal` | Unrecoverable — cannot continue | ABORT immediately, surface actionable message |
1547
- | `retryable` | Transient failure (bad output, timeout) | Auto-retry up to 2× with **identical prompt**. After 2 failures → escalate to `recoverable` |
1970
+ | `retryable` | Transient failure — malformed output, an unusable reply, a command that failed on something transient | Auto-retry up to 2× with **identical prompt**, re-opening the step each time so `retries` counts it. After 2 failures → escalate to `recoverable` |
1548
1971
  | `recoverable` | Partial success or skippable failure | Ask user with specific "continue from here / skip step / abort" options |
1549
1972
 
1550
1973
  ### Error Table
@@ -1556,25 +1979,38 @@ Each step failure is classified into one of three classes with different recover
1556
1979
  | Branch/worktree creation fails | `terminal` | ABORT — report git error. NEVER fall back to `git checkout -b` |
1557
1980
  | Interviewer `ready_to_proceed: false` | `terminal` | STOP — tell user which blockers remain |
1558
1981
  | Sub-agent returns malformed JSON | `retryable` | Retry with: "Output was malformed. Fix: [errors]. Try again." (max 2×) |
1559
- | Sub-agent timeout | `retryable` | Retry with identical prompt (max 2×) |
1982
+ | Sub-agent returns nothing usable | `retryable` | Re-open the step (`job step --status in-progress`, which counts the retry) and re-dispatch the identical prompt (max 2×) |
1560
1983
  | Task implementation fails | `recoverable` | Ask: "Step failed. Continue remaining tasks / skip this task / abort?" |
1561
- | Job-documenter returns error | `recoverable` | Log warning, continue (documentation is non-blocking) |
1562
- | All reviewers fail | `recoverable` | Skip review, add warning to report, continue to checks |
1563
- | Fix loop exceeds max_review_iterations | `recoverable` | Log unresolved findings, continue to checks |
1984
+ | `keryx job` refuses a write | `terminal` | The message names the field that failed validation. Fix the input; do NOT hand-write `state.json` to route around it. |
1985
+ | `keryx job complete` refuses | `recoverable` | It names the open and failed steps. Close each with `job step --status completed\|skipped --reason "<why>"`. |
1986
+ | `keryx review ingest` refuses a scope-B finding | `terminal` for that round | Recompute `keryx review blast-radius --json` and re-ingest with `--blast-radius`. The round is not recordable until the set is supplied. |
1987
+ | All reviewers fail | `recoverable` | Record the round as failed with a reason, add a warning to the report, continue to VERIFY (2.8) |
1988
+ | Fix loop exceeds max_review_iterations | `recoverable` | Disposition every surviving finding, log which ended the loop, continue to VERIFY (2.8) |
1564
1989
  | Final checks fail | `recoverable` | Include in report, still propose PR (user decides) |
1565
- | gh CLI not available | `recoverable` | Print PR data, user creates manually |
1990
+ | gh CLI not available | `recoverable` | Print PR data, user creates manually. `keryx review comments` needs it too — say so rather than reporting `0 outstanding`. |
1566
1991
 
1567
1992
  ### Retry Protocol (for `retryable` errors)
1568
1993
 
1569
1994
  ```
1570
- attempt 1: run step normally
1995
+ attempt 1: keryx job step <job-name> <step-id> --status in-progress
1996
+ run step normally
1571
1997
  → failure: classify error
1572
- → if retryable: retry with EXACT same prompt + "Fix these errors: [list]"
1998
+ → if retryable: keryx job step <job-name> <step-id> --status in-progress # retries += 1
1999
+ retry with the EXACT same prompt + "Fix these errors: [list]"
1573
2000
  → if fails again: escalate to recoverable → ask user
1574
- → if success: continue
2001
+ → if success: keryx job step <job-name> <step-id> --status completed
1575
2002
  ```
1576
2003
 
1577
- **Critical:** On retry, use the **same prompt** stored in `state.json → step.prompt`. Never re-derive it — re-derivation causes drift.
2004
+ **Critical:** on retry, re-send the **same prompt** — hold it for the duration of the
2005
+ step and re-send it verbatim. Never re-derive it; re-derivation causes drift.
2006
+
2007
+ The prompt itself is **not** persisted: `keryx job` writes no `step.prompt` and no
2008
+ prompt size, so do not instruct a resuming session to read one. What *is* persisted is
2009
+ that the attempt happened — `metrics.steps[].retries`, incremented every time the step
2010
+ re-enters `in_progress`, and the `--reason` line in `journal.md`. A resumed session
2011
+ therefore knows how many attempts a step has had, which is the fact the retry budget
2012
+ needs, and reconstructs the prompt from the plan and the analysis exactly as the first
2013
+ attempt did.
1578
2014
 
1579
2015
  ---
1580
2016
 
@@ -1599,13 +2035,17 @@ The orchestrator must keep the user informed during long-running execution. This
1599
2035
  │ ├─ ✅ task-1: Add validation schema
1600
2036
  │ ├─ ✅ task-2: Implement validator
1601
2037
  │ └─ 🔄 task-3: Add integration tests...
2038
+ ├─ ⏳ Verify
1602
2039
  ├─ ⏳ Review
1603
2040
  ├─ ⏳ Fix (if needed)
1604
- ├─ ⏳ Final checks
1605
2041
  └─ ⏳ PR
1606
2042
  ```
1607
2043
 
1608
- **Minimum notification interval:** Every 30 seconds during long steps (implementation, review). This prevents the user from thinking the process is stuck.
2044
+ **Notify at every step boundary** — before dispatching and after recording the result.
2045
+ Those are the moments this skill actually regains control, so they are the only moments
2046
+ it can say anything; a "notify every 30 seconds" rule would need a timer nothing here
2047
+ has. `keryx job status <job-name>` renders the same tree from the package, which is
2048
+ what to show a user who asks mid-run.
1609
2049
 
1610
2050
  **If notification tools are unavailable** (no MCP, no Telegram): fall back to inline text output between steps.
1611
2051
 
@@ -1615,66 +2055,76 @@ The orchestrator must keep the user informed during long-running execution. This
1615
2055
 
1616
2056
  1. **DO** ALWAYS collect context in Phase 0 — project directory is MANDATORY, never assume.
1617
2057
  2. **DO** build plans dynamically based on intent — not a fixed 8-phase pipeline.
1618
- 3. **DO** initialize job documentation before executing any step.
1619
- 4. **DO** document every step result via job-documenter.
2058
+ 3. **DO** run `keryx job init` before executing any step.
2059
+ 4. **DO** record every step with `keryx job step` and every document with `keryx job document` — the package, not this session, is the record.
1620
2060
  5. **DO** parallelize independent tasks and reviewers where safe.
1621
2061
  6. **DO** respect dependency order — use wave-based execution for implementation.
1622
- 7. **DO** limit review → fix loop to max_review_iterations.
2062
+ 7. **DO** limit review → fix loop to max_review_iterations, and stop earlier on repetition.
1623
2063
  8. **DO** present PR proposal to user before creating (unless auto_create_pr).
1624
- 9. **DO** tell user where documentation is stored at completion.
2064
+ 9. **DO** tell user where the job package is at completion.
1625
2065
  10. **DO** ALWAYS use `git worktree add` for feature branches — NEVER `git checkout -b`.
1626
2066
  11. **DO** run ALL commands in the **worktree directory**, never in the original project.
1627
2067
  12. **DO** ask user for confirmation before extending plan (e.g., analyze → implement).
1628
- 13. **DO** send progress notifications at phase/step transitions and every 30s during long steps.
2068
+ 13. **DO** send progress notifications at phase and step transitions.
1629
2069
  14. **DO** use auto-detected `package_manager` and `run_command` — never hardcode `npm`.
1630
2070
  15. **DO NOT** ask the user anything during execution (after Phase 0) — except for critical failures and plan extension decisions.
1631
2071
  16. **DO NOT** push the branch until user confirms (or auto_create_pr).
1632
- 17. **DO NOT** skip job documentation — it's a core feature, not optional.
1633
- 18. **DO NOT** create job documentation for sub-agent results directly — orchestrator formats and sends to documenter.
1634
- 19. **DO** store the prompt used for each sub-agent step in `state.json → step.prompt` before dispatching — required for retry and resume.
1635
- 20. **DO** classify every step failure as `terminal`, `retryable`, or `recoverable` — never just abort or ask without classifying first.
1636
- 21. **DO** show agent-explicit plan in 1.3 and ask approve/adjust — unless `plan_approval: false`.
1637
- 22. **DO** run `sanity-check` after every implement step before dispatching review.
1638
- 23. **DO** auto-trigger `test-gen` if implementer produced no test files (unless `run_test_gen: false`).
1639
- 24. **DO** auto-trigger `security-audit` if diff touches auth/API/DB/env files (unless `run_security_audit: false`).
1640
- 25. **DO** include changelog entry in PR body (unless `run_changelog: false`).
1641
- 26. **DO NOT** deploy without user confirmation (unless `run_deploy: true` explicitly set).
2072
+ 17. **DO NOT** hand-write `state.json`, or let a sub-agent write it. `keryx job` is the only writer, and it validates every write.
2073
+ 18. **DO NOT** let a sub-agent record its own result in the package — the orchestrator runs `keryx job document`.
2074
+ 19. **DO** run `keryx review start` before a fix round and `keryx review ingest` after synthesis, so the round is citable.
2075
+ 20. **DO** give every finding a terminal disposition with `keryx review complete --finding … --disposition … --evidence …`. A finding never leaves the loop by being absent from the next round.
2076
+ 21. **DO** classify every step failure as `terminal`, `retryable`, or `recoverable` — never just abort or ask without classifying first.
2077
+ 22. **DO** show agent-explicit plan in 1.3 and ask approve/adjust — unless `plan_approval: false`.
2078
+ 23. **DO** run `sanity-check` after every implement step before dispatching review.
2079
+ 24. **DO** auto-trigger `test-gen` if implementer produced no test files (unless `run_test_gen: false`).
2080
+ 25. **DO** auto-trigger `security-audit` if diff touches auth/API/DB/env files (unless `run_security_audit: false`).
2081
+ 26. **DO** include changelog entry in PR body (unless `run_changelog: false`).
2082
+ 27. **DO** dispatch with `subagent_type: "general-purpose"` — `"general"` is not a value any dispatcher accepts.
2083
+ 28. **DO** compute every dispatch's model with `keryx review tier` — never assign a tier by hand, and never write a model id into a dispatch.
2084
+ 29. **DO** pass `--outstanding <n>` on `keryx review budget` and `keryx review ingest` — this orchestrator is the outermost of the three nesting levels the concurrency cap was sized for, and the cap binds the nested total only when the parent declares its in-flight count.
2085
+ 30. **DO NOT** deploy without user confirmation (unless `run_deploy: true` explicitly set).
1642
2086
 
1643
2087
  ---
1644
2088
 
1645
2089
  ## Configurable Jobs Root
1646
2090
 
1647
- The jobs documentation root is configurable, not hardcoded:
1648
-
1649
- **Resolution order:**
1650
- 1. `JOBS_ROOT` passed explicitly by the orchestrator in the sub-agent dispatch prompt
1651
- 2. `GDMETAPRO_JOBS_ROOT` environment variable (if set)
1652
- 3. Default: `.metaproject/jobs/` ← project-local (PROJECT_DIR is known by Phase 0.2)
2091
+ `JOBS_ROOT` in this document is shorthand for **`.metaproject/jobs`, relative to the
2092
+ project directory** — and that is the only value it takes. `keryx job` resolves it from
2093
+ the working directory and records it in `state.json → jobs_root`; there is no
2094
+ environment variable and no override, so do not tell a sub-agent to look one up.
1653
2095
 
1654
2096
  ```bash
1655
- JOBS_ROOT="${GDMETAPRO_JOBS_ROOT:-.metaproject/jobs}"
2097
+ JOBS_ROOT=".metaproject/jobs"
1656
2098
  ```
1657
2099
 
1658
- All references to job paths in sub-agent prompts must use the resolved `JOBS_ROOT`.
2100
+ The project directory is the one collected in Phase 0.2 and passed as
2101
+ `keryx job init --project <path>`. Run `keryx job` commands from that directory, and
2102
+ expand `<JOBS_ROOT>` to the literal path when writing a sub-agent prompt — a subagent
2103
+ receives paths, it does not resolve them.
1659
2104
 
1660
2105
  ---
1661
2106
 
1662
2107
  ## Post-Mortem (for failed/aborted jobs)
1663
2108
 
1664
- When a job ends with status `aborted`, `timeout`, or has unresolved critical issues:
2109
+ When a job ends with a step recorded `failed`, or with unresolved blocker findings:
2110
+
2111
+ 1. **Auto-generate post-mortem** document. The timeline is not recalled — it is read
2112
+ off `journal.md`, which `keryx job` timestamped as the job ran, and the retry counts
2113
+ come from `keryx job status <job-name> --json`:
1665
2114
 
1666
- 1. **Auto-generate post-mortem** document:
1667
2115
  ```markdown
1668
2116
  # Post-Mortem: <job-name>
1669
2117
 
1670
2118
  ## Timeline
1671
- - Phase 0 completed: <timestamp>
1672
- - Phase 2, step "implement" started: <timestamp>
1673
- - Step "task-3" failed after 2 retries: <timestamp>
1674
- - Job aborted by user: <timestamp>
2119
+ (from .metaproject/jobs/<job-name>/journal.md — every line as recorded)
2120
+ - <ISO timestamp> - created
2121
+ - <ISO timestamp> - step: implement in-progress (retries 0)
2122
+ - <ISO timestamp> - step: implement in-progress (retries 1)
2123
+ - <ISO timestamp> - step: implement failed (retries 1) — <reason>
1675
2124
 
1676
2125
  ## What Went Wrong
1677
2126
  - <Step name> failed with: <error class> — <error message>
2127
+ - Recorded retries: <metrics.steps[].retries>
1678
2128
  - Root cause hypothesis: <analysis>
1679
2129
 
1680
2130
  ## What Worked
@@ -1684,43 +2134,56 @@ When a job ends with status `aborted`, `timeout`, or has unresolved critical iss
1684
2134
  ## Recommendations for Retry
1685
2135
  - Fix <specific issue> before re-running
1686
2136
  - Consider splitting task-3 into smaller subtasks
1687
- - Increase step_timeout_ms if timeout was the issue
1688
2137
  ```
1689
2138
 
1690
2139
  2. Save to `.metaproject/jobs/<job-name>/post-mortem.md`
1691
2140
  3. Include in final user message: "Post-mortem saved to `.metaproject/jobs/<job-name>/post-mortem.md`"
1692
2141
 
2142
+ There is no `aborted` or `timeout` job status to key this on and nothing writes one.
2143
+ The trigger is what the package says: a `failed` step, or findings still without a
2144
+ terminal disposition.
2145
+
1693
2146
  ---
1694
2147
 
1695
2148
  ## Metrics Collection
1696
2149
 
1697
- The orchestrator tracks timing and token usage for each step to enable optimization over time.
2150
+ `keryx job step` writes a metrics row per step. Nothing here is collected by hand.
1698
2151
 
1699
- **Collected per step:**
2152
+ **Written per step, into `state.json → metrics.steps[]`:**
1700
2153
  ```json
1701
2154
  {
1702
2155
  "step_id": "implement",
1703
- "started_at": "2024-03-15T10:30:00Z",
1704
- "completed_at": "2024-03-15T10:35:22Z",
2156
+ "status": "completed",
2157
+ "started_at": "2026-08-30T10:30:00.000Z",
2158
+ "completed_at": "2026-08-30T10:35:22.000Z",
1705
2159
  "duration_ms": 322000,
1706
- "total_tokens": 84500,
1707
- "status": "success",
1708
2160
  "retries": 0
1709
2161
  }
1710
2162
  ```
1711
2163
 
1712
- **Saved to:** `.metaproject/jobs/<job-name>/metrics.json`
2164
+ - `started_at` is stamped every time the step enters `in_progress`.
2165
+ - `retries` counts attempts **beyond the first**: the first `--status in-progress`
2166
+ leaves it at 0 and every re-entry adds one. It is on disk, so a resumed session
2167
+ reads the real count instead of restarting at zero.
2168
+ - `duration_ms` is `completed_at - started_at` for the last attempt.
2169
+ - `total_tokens` is declared in the schema and **nothing writes it**. Do not report a
2170
+ token figure as if it came from the package; if you have one, say where it came from.
2171
+
2172
+ **Read it back:**
2173
+ ```bash
2174
+ keryx job status <job-name> --json # `retries` per step, plus phase and next_step
2175
+ ```
1713
2176
 
1714
- **Aggregated in report:**
2177
+ **Aggregated in the report:**
1715
2178
  ```markdown
1716
2179
  ## Metrics
1717
- | Step | Duration | Tokens | Retries |
1718
- |------|----------|--------|---------|
1719
- | Analyze | 45s | 12K | 0 |
1720
- | Context | 30s | 8K | 0 |
1721
- | Implement | 5m 22s | 84K | 0 |
1722
- | Review | 1m 10s | 25K | 0 |
1723
- | **Total** | **7m 47s** | **129K** | **0** |
2180
+ | Step | Duration | Retries |
2181
+ |------|----------|---------|
2182
+ | Analyze | 45s | 0 |
2183
+ | Context | 30s | 0 |
2184
+ | Implement | 5m 22s | 1 |
2185
+ | Review | 1m 10s | 0 |
2186
+ | **Total** | **7m 47s** | **1** |
1724
2187
  ```
1725
2188
 
1726
- This data helps identify which steps are bottlenecks and whether budget guards need adjustment.
2189
+ This data identifies which steps are bottlenecks and which ones needed a second attempt.