@mrciphersmith/keryx 0.2.70 → 0.2.71

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (177) hide show
  1. package/dist/cli.js +10357 -4676
  2. package/docs/README.md +54 -0
  3. package/docs/requirements/shared-agent-context/README.md +104 -0
  4. package/package.json +3 -2
  5. package/src/gdskills/bundled/rules/core/model-selection.mdc +184 -31
  6. package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +36 -0
  7. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.codex.md +1 -1
  8. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.cursor.md +1 -1
  9. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.md +2 -1
  10. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.opencode.md +1 -1
  11. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.zed.md +1 -1
  12. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.codex.md +1 -1
  13. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.cursor.md +1 -1
  14. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.md +1 -1
  15. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.opencode.md +1 -1
  16. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.zed.md +1 -1
  17. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.codex.md +1 -1
  18. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.cursor.md +1 -1
  19. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.md +1 -1
  20. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.opencode.md +1 -1
  21. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.zed.md +1 -1
  22. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.codex.md +1 -1
  23. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.cursor.md +1 -1
  24. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.md +1 -1
  25. package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +78 -19
  26. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.codex.md +1 -1
  27. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.cursor.md +1 -1
  28. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.md +1 -1
  29. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.opencode.md +1 -1
  30. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.zed.md +1 -1
  31. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.codex.md +1 -1
  32. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.cursor.md +1 -1
  33. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.md +2 -1
  34. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.opencode.md +1 -1
  35. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.zed.md +1 -1
  36. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.codex.md +28 -3
  37. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.cursor.md +28 -3
  38. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +28 -3
  39. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.opencode.md +28 -3
  40. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.zed.md +28 -3
  41. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.codex.md +20 -2
  42. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.cursor.md +20 -2
  43. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.md +21 -2
  44. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.opencode.md +20 -2
  45. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.zed.md +20 -2
  46. package/src/gdskills/bundled/skills/planning/autodoc-analyst/SKILL.md +2 -1
  47. package/src/gdskills/bundled/skills/planning/autodoc-architect/SKILL.md +3 -1
  48. package/src/gdskills/bundled/skills/planning/autodoc-assembler/SKILL.md +2 -1
  49. package/src/gdskills/bundled/skills/planning/autodoc-orchestrator/SKILL.md +2 -1
  50. package/src/gdskills/bundled/skills/planning/autodoc-scanner/SKILL.md +2 -1
  51. package/src/gdskills/bundled/skills/planning/autodoc-writer/SKILL.md +2 -1
  52. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.codex.md +1 -1
  53. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.cursor.md +1 -1
  54. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
  55. package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.codex.md +1 -1
  56. package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.cursor.md +1 -1
  57. package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.md +2 -1
  58. package/src/gdskills/bundled/skills/planning/docpack-orchestrator/SKILL.md +1 -1
  59. package/src/gdskills/bundled/skills/planning/docpack-review/SKILL.md +1 -1
  60. package/src/gdskills/bundled/skills/planning/interview/SKILL.codex.md +1 -1
  61. package/src/gdskills/bundled/skills/planning/interview/SKILL.cursor.md +1 -1
  62. package/src/gdskills/bundled/skills/planning/interview/SKILL.md +1 -1
  63. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.codex.md +1 -1
  64. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.cursor.md +1 -1
  65. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
  66. package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.codex.md +1 -1
  67. package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.cursor.md +1 -1
  68. package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.md +2 -1
  69. package/src/gdskills/bundled/skills/planning/planner/SKILL.codex.md +1 -1
  70. package/src/gdskills/bundled/skills/planning/planner/SKILL.cursor.md +1 -1
  71. package/src/gdskills/bundled/skills/planning/planner/SKILL.md +2 -1
  72. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.codex.md +1 -1
  73. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.cursor.md +1 -1
  74. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.md +1 -1
  75. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.opencode.md +1 -1
  76. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.zed.md +1 -1
  77. package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.codex.md +1 -1
  78. package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.cursor.md +1 -1
  79. package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.md +2 -1
  80. package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.codex.md +1 -1
  81. package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.cursor.md +1 -1
  82. package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.md +2 -1
  83. package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.codex.md +1 -1
  84. package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.cursor.md +1 -1
  85. package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.md +2 -1
  86. package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.codex.md +1 -1
  87. package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.cursor.md +1 -1
  88. package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.md +2 -1
  89. package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.codex.md +1 -1
  90. package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.cursor.md +1 -1
  91. package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.md +1 -1
  92. package/src/gdskills/bundled/skills/platform/hookify/SKILL.codex.md +1 -1
  93. package/src/gdskills/bundled/skills/platform/hookify/SKILL.cursor.md +1 -1
  94. package/src/gdskills/bundled/skills/platform/hookify/SKILL.md +1 -1
  95. package/src/gdskills/bundled/skills/quality/changelog/SKILL.codex.md +1 -1
  96. package/src/gdskills/bundled/skills/quality/changelog/SKILL.cursor.md +1 -1
  97. package/src/gdskills/bundled/skills/quality/changelog/SKILL.md +1 -1
  98. package/src/gdskills/bundled/skills/quality/commit/SKILL.codex.md +1 -1
  99. package/src/gdskills/bundled/skills/quality/commit/SKILL.cursor.md +1 -1
  100. package/src/gdskills/bundled/skills/quality/commit/SKILL.md +1 -1
  101. package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.codex.md +1 -1
  102. package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.cursor.md +1 -1
  103. package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.md +1 -1
  104. package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.codex.md +1 -1
  105. package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.cursor.md +1 -1
  106. package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.md +1 -1
  107. package/src/gdskills/bundled/skills/quality/deploy/SKILL.codex.md +1 -1
  108. package/src/gdskills/bundled/skills/quality/deploy/SKILL.cursor.md +1 -1
  109. package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
  110. package/src/gdskills/bundled/skills/quality/metaproject-security/SKILL.md +1 -1
  111. package/src/gdskills/bundled/skills/quality/perf-check/SKILL.codex.md +1 -1
  112. package/src/gdskills/bundled/skills/quality/perf-check/SKILL.cursor.md +1 -1
  113. package/src/gdskills/bundled/skills/quality/perf-check/SKILL.md +1 -1
  114. package/src/gdskills/bundled/skills/quality/pr/SKILL.codex.md +1 -1
  115. package/src/gdskills/bundled/skills/quality/pr/SKILL.cursor.md +1 -1
  116. package/src/gdskills/bundled/skills/quality/pr/SKILL.md +1 -1
  117. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.codex.md +1 -1
  118. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.cursor.md +1 -1
  119. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.md +1 -1
  120. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.opencode.md +1 -1
  121. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.zed.md +1 -1
  122. package/src/gdskills/bundled/skills/quality/push/SKILL.codex.md +1 -1
  123. package/src/gdskills/bundled/skills/quality/push/SKILL.cursor.md +1 -1
  124. package/src/gdskills/bundled/skills/quality/push/SKILL.md +1 -1
  125. package/src/gdskills/bundled/skills/quality/security-audit/SKILL.codex.md +1 -1
  126. package/src/gdskills/bundled/skills/quality/security-audit/SKILL.cursor.md +1 -1
  127. package/src/gdskills/bundled/skills/quality/security-audit/SKILL.md +1 -1
  128. package/src/gdskills/bundled/skills/quality/test-gen/SKILL.codex.md +1 -1
  129. package/src/gdskills/bundled/skills/quality/test-gen/SKILL.cursor.md +1 -1
  130. package/src/gdskills/bundled/skills/quality/test-gen/SKILL.md +1 -1
  131. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.codex.md +1 -1
  132. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.cursor.md +1 -1
  133. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.md +1 -1
  134. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.opencode.md +1 -1
  135. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.zed.md +1 -1
  136. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.codex.md +1 -1
  137. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.cursor.md +1 -1
  138. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.md +1 -1
  139. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.opencode.md +1 -1
  140. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.zed.md +1 -1
  141. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.codex.md +1 -1
  142. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.cursor.md +1 -1
  143. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.md +1 -1
  144. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.opencode.md +1 -1
  145. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.zed.md +1 -1
  146. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.codex.md +1 -1
  147. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.cursor.md +1 -1
  148. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.md +2 -1
  149. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.opencode.md +1 -1
  150. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.zed.md +1 -1
  151. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.codex.md +1 -1
  152. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.cursor.md +1 -1
  153. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
  154. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.opencode.md +1 -1
  155. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.zed.md +1 -1
  156. package/src/gdskills/bundled/skills/review/review-architecture/SKILL.md +37 -10
  157. package/src/gdskills/bundled/skills/review/review-backend/SKILL.md +48 -14
  158. package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +49 -12
  159. package/src/gdskills/bundled/skills/review/review-core-boundaries/SKILL.md +34 -2
  160. package/src/gdskills/bundled/skills/review/review-flow-graph/SKILL.md +33 -2
  161. package/src/gdskills/bundled/skills/review/review-frontend/SKILL.md +70 -29
  162. package/src/gdskills/bundled/skills/review/review-frontend-conventions/SKILL.md +34 -3
  163. package/src/gdskills/bundled/skills/review/review-highload/SKILL.md +49 -15
  164. package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +39 -11
  165. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +590 -21
  166. package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +7 -0
  167. package/src/gdskills/bundled/skills/review/review-orchestrator/verification-claim.schema.json +78 -0
  168. package/src/gdskills/bundled/skills/review/review-performance/SKILL.md +43 -13
  169. package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +8 -2
  170. package/src/gdskills/bundled/skills/review/review-regression/SKILL.md +185 -0
  171. package/src/gdskills/bundled/skills/review/review-security-code/SKILL.md +44 -13
  172. package/src/gdskills/bundled/skills/review/review-style/SKILL.md +26 -6
  173. package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +35 -3
  174. package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +276 -0
  175. package/src/gdskills/contracts/review-finding.schema.json +119 -1
  176. package/src/gdskills/contracts/subagent-dispatch.schema.json +59 -3
  177. package/src/gdskills/bundled/skills/review/review-strict/SKILL.md +0 -328
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "workflow"
14
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode,claude"
16
16
  ---
17
17
 
18
18
  # Feature Development (7-Phase)
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "2.0.0"
13
13
  category: "workflow"
14
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode,claude"
16
16
  ---
17
17
 
18
18
  <SUBAGENT-STOP>
@@ -14,8 +14,8 @@ metadata:
14
14
  author: "MrCipherSmith"
15
15
  version: "1.3.0"
16
16
  category: "orchestration"
17
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
17
18
  license: "MIT"
18
- compatibility: "cursor,codex,zed,opencode,claude"
19
19
  ---
20
20
 
21
21
  # Flow Orchestrator
@@ -84,8 +84,8 @@ flowchart TD
84
84
  H -- "yes" --> I{"Ask user how to finish"}
85
85
  I -- "create PR" --> J["create PR and run review/fix loop"]
86
86
  J --> K{"review clean, PR mergeable?"}
87
- K -- "no, attempts < 6" --> J
88
- K -- "no, attempts = 6" --> R["enrich context and change fix strategy"]
87
+ K -- "no, attempts < 3" --> J
88
+ K -- "no, attempts = 3 or repetition detected" --> R["enrich context and change fix strategy"]
89
89
  R --> J
90
90
  K -- "yes" --> L["merge PR into recorded base branch"]
91
91
  L --> M["keryx flow implemented --pr"]
@@ -128,10 +128,21 @@ it already tried. The flow package does.
128
128
  ```
129
129
 
130
130
  5. Apply the Phase 4 attempt budget against the **persisted** count. If
131
- `attempts.count` for the task has already reached six, do not re-dispatch
132
- the same approach: go to the re-planning step (Phase 4, PR review/fix
133
- loop, step 4) and record the decision in `journal.md`.
134
- 6. If the flow is `blocked`, read the blocking reason from `journal.md`,
131
+ `attempts.count` for the task has already reached **three**, do not
132
+ re-dispatch the same approach: go to the re-planning step (Phase 4, PR
133
+ review/fix loop, step 4) and record the decision in `journal.md`.
134
+ 6. Run the repetition check before spending an attempt, whatever the count
135
+ says:
136
+
137
+ ```bash
138
+ keryx review loop --flow <id> --task <Tn>
139
+ ```
140
+
141
+ A non-zero exit means the same finding has recurred or two consecutive
142
+ rounds produced identical output. Go straight to the re-planning step.
143
+ Do not spend the remaining attempts on the same approach because the
144
+ budget has some left — that is the failure this check exists to catch.
145
+ 7. If the flow is `blocked`, read the blocking reason from `journal.md`,
135
146
  resolve or escalate it, then `keryx flow unblock <id>`.
136
147
  4. If the user wants a new flow, continue at 0.1.
137
148
 
@@ -231,8 +242,13 @@ did not exist and 24 completed flows shipped with an open task:
231
242
  `flow.json`, which `keryx flow init` writes for every flow it creates. A
232
243
  package created before the gate landed does not carry the flag, and for it
233
244
  the gate reports `skipped` and blocks nothing;
234
- - a task fails the gate when its status is not `done`, when its disposition is
235
- `failed`, or when its disposition is `skipped` with no recorded reason;
245
+ - a task fails the gate when its status is not `done`; when its disposition is
246
+ `failed`; when its disposition is `blocked` (terminal, but the work did not
247
+ happen — and the harness emits this disposition on its own for a run that
248
+ ended blocked); when its disposition is `skipped` with no recorded reason; or
249
+ when its disposition is a value this build does not recognise. An
250
+ unrecognised disposition FAILS rather than falling through: a gate whose
251
+ default for the unknown case is "pass" is not a gate;
236
252
  - to close a task as deliberately not needed, record why:
237
253
 
238
254
  ```bash
@@ -346,7 +362,23 @@ Before accepting implementation:
346
362
  1. Run focused tests for touched scope.
347
363
  2. Run `code-verifier`.
348
364
  3. Run `keryx health run` when Code Health is enabled.
349
- 4. Run `review-orchestrator` with relevant domains.
365
+ 4. Check the bounds, then run `review-orchestrator` with relevant domains.
366
+
367
+ ```bash
368
+ keryx review budget --spent <usd-so-far> --outstanding <subagents you already have in flight>
369
+ ```
370
+
371
+ A non-zero exit means the spend ceiling (3 USD by default) has been reached:
372
+ **stop and ask the user** rather than dispatching another fan-out.
373
+
374
+ `--outstanding` is the part that matters here. `review-orchestrator`
375
+ dispatches reviewers in parallel and runs *nested* under this skill, and
376
+ keryx cannot observe subagents in another process. Passing the count you
377
+ already have in flight is the only thing that makes the concurrency cap mean
378
+ anything across the nesting; omit it and the cap bounds the reviewer fan-out
379
+ alone, which the review record then states plainly rather than implying
380
+ otherwise.
381
+
350
382
  5. If findings require code changes, dispatch fix work through `task-implementer`
351
383
  and record the fix task in the flow.
352
384
  6. Close the skill-learning loop (see `rules/core/skill-lifecycle.mdc`). Collect
@@ -392,18 +424,45 @@ How should this flow end?
392
424
  branch state.
393
425
  2. If findings or required check failures remain, create or update a flow fix
394
426
  task, dispatch `task-implementer`, push the fix, and run review again.
395
- 3. Allow at most six review/fix attempts for the current approach. Count an
396
- attempt when review/check results are available, including a clean result,
427
+ 3. Allow at most **three** review/fix attempts for the current approach. Count
428
+ an attempt when review/check results are available, including a clean result,
397
429
  and record it with `keryx flow task attempt <id> <Tn> --outcome ...` so the
398
430
  count survives a session restart. Read the budget from that task's
399
431
  `attempts.count` in `flow.json`, never from this session's memory.
400
- 4. If attempt six is not clean, do not blindly repeat the same loop. Enrich
401
- context from the findings, affected graph, relevant wiki, and
402
- health/testing artifacts; identify the likely cycle cause; choose a
403
- materially different fix strategy or split the work into narrower tasks;
404
- record the decision in `journal.md`; then continue with the enriched
405
- context.
406
- 5. Never merge while findings or required checks remain unresolved. If the
432
+
433
+ Three, and the same three that `task-implementer` and `job-orchestrator`
434
+ already use. This skill said six, which was an outlier with nothing behind
435
+ it. The evidence converges on three: *"the first three to four repair
436
+ iterations account for most achievable gains"*
437
+ ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197)); correctness falls
438
+ **0.820 -> 0.673** across two forced revisions while cumulative ever-correct
439
+ is **0.847** ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)) — the
440
+ agent finds the fix and then destroys it, throwing away ~15 points by not
441
+ stopping. Aider hardcodes `max_reflections = 3`; OpenHands' critic uses 3.
442
+ Rounds four through six were not buying convergence; they were buying
443
+ regressions.
444
+
445
+ 4. **Before** spending an attempt, and regardless of how much budget is left,
446
+ run the repetition check:
447
+
448
+ ```bash
449
+ keryx review loop --flow <id> --task <Tn>
450
+ ```
451
+
452
+ It escalates (non-zero exit) when the same finding recurs in two rounds, or
453
+ two consecutive rounds produce identical review output. It reads the review
454
+ packages on disk and the persisted `attempts.count`, not this session's
455
+ memory, and it deliberately never reads the remaining budget — an agent
456
+ emitting the identical failing output three times must be caught on the
457
+ second, not after the budget runs out.
458
+
459
+ 5. If the third attempt is not clean, **or the repetition check escalated
460
+ earlier**, do not blindly repeat the same loop. Enrich context from the
461
+ findings, affected graph, relevant wiki, and health/testing artifacts;
462
+ identify the likely cycle cause; choose a materially different fix strategy
463
+ or split the work into narrower tasks; record the decision in `journal.md`;
464
+ then continue with the enriched context.
465
+ 6. Never merge while findings or required checks remain unresolved. If the
407
466
  re-planned approach still cannot produce a mergeable PR, leave the flow
408
467
  `in-progress` and report the blocker instead of forcing completion.
409
468
 
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "analysis"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode"
16
16
  ---
17
17
 
18
18
  # Issue Analyzer
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "analysis"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode"
16
16
  ---
17
17
 
18
18
  # Issue Analyzer
@@ -12,8 +12,8 @@ metadata:
12
12
  version: "1.1.0"
13
13
  category: "analysis"
14
14
  agent_worthy: true
15
+ compatible_harnesses: "cursor,codex,zed,opencode"
15
16
  license: "MIT"
16
- compatibility: "cursor,codex,zed,opencode"
17
17
  ---
18
18
 
19
19
  # Issue Analyzer
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "analysis"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode"
16
16
  ---
17
17
 
18
18
  # Issue Analyzer
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "analysis"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode"
16
16
  ---
17
17
 
18
18
  # Issue Analyzer
@@ -10,8 +10,8 @@ metadata:
10
10
  author: "MrCipherSmith"
11
11
  version: "1.0.0"
12
12
  category: "documentation"
13
+ compatible_harnesses: "cursor,codex,zed,opencode"
13
14
  license: "MIT"
14
- compatibility: "cursor,codex,zed,opencode"
15
15
  ---
16
16
 
17
17
  # Job Documenter
@@ -10,8 +10,8 @@ metadata:
10
10
  author: "MrCipherSmith"
11
11
  version: "1.0.0"
12
12
  category: "documentation"
13
+ compatible_harnesses: "cursor,codex,zed,opencode"
13
14
  license: "MIT"
14
- compatibility: "cursor,codex,zed,opencode"
15
15
  ---
16
16
 
17
17
  # Job Documenter
@@ -1,5 +1,6 @@
1
1
  ---
2
2
  name: job-documenter
3
+ model_tier: light
3
4
  description: "Use when a job folder needs to be initialized, or analysis/report/review documents need to be created or updated in jobs/."
4
5
  triggers:
5
6
  - "Document job"
@@ -10,8 +11,8 @@ metadata:
10
11
  author: "MrCipherSmith"
11
12
  version: "1.0.0"
12
13
  category: "documentation"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
13
15
  license: "MIT"
14
- compatibility: "cursor,codex,zed,opencode"
15
16
  ---
16
17
 
17
18
  <SUBAGENT-STOP>
@@ -10,8 +10,8 @@ metadata:
10
10
  author: "MrCipherSmith"
11
11
  version: "1.0.0"
12
12
  category: "documentation"
13
+ compatible_harnesses: "cursor,codex,zed,opencode"
13
14
  license: "MIT"
14
- compatibility: "cursor,codex,zed,opencode"
15
15
  ---
16
16
 
17
17
  # Job Documenter
@@ -10,8 +10,8 @@ metadata:
10
10
  author: "MrCipherSmith"
11
11
  version: "1.0.0"
12
12
  category: "documentation"
13
+ compatible_harnesses: "cursor,codex,zed,opencode"
13
14
  license: "MIT"
14
- compatibility: "cursor,codex,zed,opencode"
15
15
  ---
16
16
 
17
17
  # Job Documenter
@@ -23,8 +23,8 @@ metadata:
23
23
  author: "MrCipherSmith"
24
24
  version: "3.2.0"
25
25
  category: "orchestration"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
27
  license: "MIT"
27
- compatibility: "cursor,codex,zed,opencode,claude"
28
28
  ---
29
29
 
30
30
  <SUBAGENT-STOP>
@@ -971,8 +971,22 @@ Review complete:
971
971
 
972
972
  Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
973
973
 
974
+ Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
975
+ this skill all use it. *"The first three to four repair iterations account for
976
+ most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
977
+ correctness falls **0.820 -> 0.673** across two forced revisions while
978
+ cumulative ever-correct is **0.847**
979
+ ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
980
+ `max_reflections = 3`; OpenHands' critic uses 3.
981
+
982
+ The bound is a ceiling, not a target. Repetition ends the loop earlier and
983
+ **regardless of remaining iterations** — a counter cannot tell "converging
984
+ slowly" from "stuck", and an agent emitting the identical failing output three
985
+ times spends the whole budget before anything notices.
986
+
974
987
  ```
975
988
  UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
989
+ PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
976
990
 
977
991
  FOR iteration in [1, 2, 3]:
978
992
  IF NOT NEEDS_FIX: BREAK
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
992
1006
  6. Recompute NEEDS_FIX from new findings
993
1007
  7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
994
1008
 
995
- IF still NEEDS_FIX after max iterations:
996
- Log "Unresolved after <N> iterations" with finding list → continue to checks
1009
+ 8. STUCK CHECK — runs before the next iteration and ignores the budget:
1010
+ IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
1011
+ OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
1012
+ THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
1013
+ Identity is the finding's dedupe_key when it has one, otherwise
1014
+ reviewer + file + symbol + problem — never the display id, which is
1015
+ per-report and would fire on every second iteration whatever happened.
1016
+ 9. PREVIOUS_REVIEW_OUTPUT = the new review output
1017
+
1018
+ IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1019
+ Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
1020
+ two ended it — a budget exhausted and a loop detected call for different next
1021
+ steps → continue to checks
997
1022
  ```
998
1023
 
999
1024
  **Fix prompt escalation pattern:**
@@ -23,8 +23,8 @@ metadata:
23
23
  author: "MrCipherSmith"
24
24
  version: "3.2.0"
25
25
  category: "orchestration"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
27
  license: "MIT"
27
- compatibility: "cursor,codex,zed,opencode,claude"
28
28
  ---
29
29
 
30
30
  <SUBAGENT-STOP>
@@ -971,8 +971,22 @@ Review complete:
971
971
 
972
972
  Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
973
973
 
974
+ Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
975
+ this skill all use it. *"The first three to four repair iterations account for
976
+ most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
977
+ correctness falls **0.820 -> 0.673** across two forced revisions while
978
+ cumulative ever-correct is **0.847**
979
+ ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
980
+ `max_reflections = 3`; OpenHands' critic uses 3.
981
+
982
+ The bound is a ceiling, not a target. Repetition ends the loop earlier and
983
+ **regardless of remaining iterations** — a counter cannot tell "converging
984
+ slowly" from "stuck", and an agent emitting the identical failing output three
985
+ times spends the whole budget before anything notices.
986
+
974
987
  ```
975
988
  UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
989
+ PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
976
990
 
977
991
  FOR iteration in [1, 2, 3]:
978
992
  IF NOT NEEDS_FIX: BREAK
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
992
1006
  6. Recompute NEEDS_FIX from new findings
993
1007
  7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
994
1008
 
995
- IF still NEEDS_FIX after max iterations:
996
- Log "Unresolved after <N> iterations" with finding list → continue to checks
1009
+ 8. STUCK CHECK — runs before the next iteration and ignores the budget:
1010
+ IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
1011
+ OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
1012
+ THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
1013
+ Identity is the finding's dedupe_key when it has one, otherwise
1014
+ reviewer + file + symbol + problem — never the display id, which is
1015
+ per-report and would fire on every second iteration whatever happened.
1016
+ 9. PREVIOUS_REVIEW_OUTPUT = the new review output
1017
+
1018
+ IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1019
+ Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
1020
+ two ended it — a budget exhausted and a loop detected call for different next
1021
+ steps → continue to checks
997
1022
  ```
998
1023
 
999
1024
  **Fix prompt escalation pattern:**
@@ -23,8 +23,8 @@ metadata:
23
23
  author: "MrCipherSmith"
24
24
  version: "3.2.0"
25
25
  category: "orchestration"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
27
  license: "MIT"
27
- compatibility: "cursor,codex,zed,opencode,claude"
28
28
  ---
29
29
 
30
30
  <SUBAGENT-STOP>
@@ -973,8 +973,22 @@ Review complete:
973
973
 
974
974
  Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
975
975
 
976
+ Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
977
+ this skill all use it. *"The first three to four repair iterations account for
978
+ most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
979
+ correctness falls **0.820 -> 0.673** across two forced revisions while
980
+ cumulative ever-correct is **0.847**
981
+ ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
982
+ `max_reflections = 3`; OpenHands' critic uses 3.
983
+
984
+ The bound is a ceiling, not a target. Repetition ends the loop earlier and
985
+ **regardless of remaining iterations** — a counter cannot tell "converging
986
+ slowly" from "stuck", and an agent emitting the identical failing output three
987
+ times spends the whole budget before anything notices.
988
+
976
989
  ```
977
990
  UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
991
+ PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
978
992
 
979
993
  FOR iteration in [1, 2, 3]:
980
994
  IF NOT NEEDS_FIX: BREAK
@@ -994,8 +1008,19 @@ FOR iteration in [1, 2, 3]:
994
1008
  6. Recompute NEEDS_FIX from new findings
995
1009
  7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
996
1010
 
997
- IF still NEEDS_FIX after max iterations:
998
- Log "Unresolved after <N> iterations" with finding list → continue to checks
1011
+ 8. STUCK CHECK — runs before the next iteration and ignores the budget:
1012
+ IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
1013
+ OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
1014
+ THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
1015
+ Identity is the finding's dedupe_key when it has one, otherwise
1016
+ reviewer + file + symbol + problem — never the display id, which is
1017
+ per-report and would fire on every second iteration whatever happened.
1018
+ 9. PREVIOUS_REVIEW_OUTPUT = the new review output
1019
+
1020
+ IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1021
+ Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
1022
+ two ended it — a budget exhausted and a loop detected call for different next
1023
+ steps → continue to checks
999
1024
  ```
1000
1025
 
1001
1026
  **Fix prompt escalation pattern:**
@@ -23,8 +23,8 @@ metadata:
23
23
  author: "MrCipherSmith"
24
24
  version: "3.2.0"
25
25
  category: "orchestration"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
27
  license: "MIT"
27
- compatibility: "cursor,codex,zed,opencode,claude"
28
28
  ---
29
29
 
30
30
  <SUBAGENT-STOP>
@@ -971,8 +971,22 @@ Review complete:
971
971
 
972
972
  Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
973
973
 
974
+ Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
975
+ this skill all use it. *"The first three to four repair iterations account for
976
+ most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
977
+ correctness falls **0.820 -> 0.673** across two forced revisions while
978
+ cumulative ever-correct is **0.847**
979
+ ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
980
+ `max_reflections = 3`; OpenHands' critic uses 3.
981
+
982
+ The bound is a ceiling, not a target. Repetition ends the loop earlier and
983
+ **regardless of remaining iterations** — a counter cannot tell "converging
984
+ slowly" from "stuck", and an agent emitting the identical failing output three
985
+ times spends the whole budget before anything notices.
986
+
974
987
  ```
975
988
  UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
989
+ PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
976
990
 
977
991
  FOR iteration in [1, 2, 3]:
978
992
  IF NOT NEEDS_FIX: BREAK
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
992
1006
  6. Recompute NEEDS_FIX from new findings
993
1007
  7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
994
1008
 
995
- IF still NEEDS_FIX after max iterations:
996
- Log "Unresolved after <N> iterations" with finding list → continue to checks
1009
+ 8. STUCK CHECK — runs before the next iteration and ignores the budget:
1010
+ IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
1011
+ OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
1012
+ THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
1013
+ Identity is the finding's dedupe_key when it has one, otherwise
1014
+ reviewer + file + symbol + problem — never the display id, which is
1015
+ per-report and would fire on every second iteration whatever happened.
1016
+ 9. PREVIOUS_REVIEW_OUTPUT = the new review output
1017
+
1018
+ IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1019
+ Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
1020
+ two ended it — a budget exhausted and a loop detected call for different next
1021
+ steps → continue to checks
997
1022
  ```
998
1023
 
999
1024
  **Fix prompt escalation pattern:**
@@ -23,8 +23,8 @@ metadata:
23
23
  author: "MrCipherSmith"
24
24
  version: "3.2.0"
25
25
  category: "orchestration"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
27
  license: "MIT"
27
- compatibility: "cursor,codex,zed,opencode,claude"
28
28
  ---
29
29
 
30
30
  <SUBAGENT-STOP>
@@ -971,8 +971,22 @@ Review complete:
971
971
 
972
972
  Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
973
973
 
974
+ Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
975
+ this skill all use it. *"The first three to four repair iterations account for
976
+ most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
977
+ correctness falls **0.820 -> 0.673** across two forced revisions while
978
+ cumulative ever-correct is **0.847**
979
+ ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
980
+ `max_reflections = 3`; OpenHands' critic uses 3.
981
+
982
+ The bound is a ceiling, not a target. Repetition ends the loop earlier and
983
+ **regardless of remaining iterations** — a counter cannot tell "converging
984
+ slowly" from "stuck", and an agent emitting the identical failing output three
985
+ times spends the whole budget before anything notices.
986
+
974
987
  ```
975
988
  UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
989
+ PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
976
990
 
977
991
  FOR iteration in [1, 2, 3]:
978
992
  IF NOT NEEDS_FIX: BREAK
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
992
1006
  6. Recompute NEEDS_FIX from new findings
993
1007
  7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
994
1008
 
995
- IF still NEEDS_FIX after max iterations:
996
- Log "Unresolved after <N> iterations" with finding list → continue to checks
1009
+ 8. STUCK CHECK — runs before the next iteration and ignores the budget:
1010
+ IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
1011
+ OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
1012
+ THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
1013
+ Identity is the finding's dedupe_key when it has one, otherwise
1014
+ reviewer + file + symbol + problem — never the display id, which is
1015
+ per-report and would fire on every second iteration whatever happened.
1016
+ 9. PREVIOUS_REVIEW_OUTPUT = the new review output
1017
+
1018
+ IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
1019
+ Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
1020
+ two ended it — a budget exhausted and a loop detected call for different next
1021
+ steps → continue to checks
997
1022
  ```
998
1023
 
999
1024
  **Fix prompt escalation pattern:**
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "implementation"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode"
16
16
  ---
17
17
 
18
18
  # Task Implementer
@@ -284,7 +284,25 @@ npm run build-storybook # Verify stories compile
284
284
  | Test failures | Fix failing tests, re-commit |
285
285
  | Story build failure | Fix story code, re-commit |
286
286
 
287
- Maximum 3 self-fix attempts per verification step.
287
+ Maximum 3 self-fix attempts per verification step.
288
+
289
+ Three, and it is the same three `job-orchestrator` and `flow-orchestrator`
290
+ use: one round bound, not four. *"The first three to four repair iterations
291
+ account for most achievable gains"*
292
+ ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197)); correctness falls
293
+ **0.820 -> 0.673** across two forced revisions while cumulative ever-correct is
294
+ **0.847** ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)) — the agent
295
+ finds the fix and then destroys it. Aider hardcodes `max_reflections = 3`;
296
+ OpenHands' critic uses 3.
297
+
298
+ **Stop earlier on repetition, whatever the count says.** If an attempt produces
299
+ the same failure output as the previous attempt — the same failing test with the
300
+ same message, the same type error at the same site — do NOT spend the remaining
301
+ attempts. The counter cannot tell "converging slowly" from "stuck", and three
302
+ identical outputs cost the whole budget to learn what the second one already
303
+ said. Report the block instead, naming what repeated.
304
+
305
+
288
306
  **ROLLBACK POLICY**: If implementation fatally fails (e.g. tests still failing after 3 attempts or unresolvable compilation errors), you MUST run `git reset --hard` to clean the worktree before reporting the failure in Phase 6, unless explicitly instructed to leave it dirty.
289
307
 
290
308
  **5.5 Re-commit fixes if any:**
@@ -11,8 +11,8 @@ metadata:
11
11
  author: "MrCipherSmith"
12
12
  version: "1.0.0"
13
13
  category: "implementation"
14
+ compatible_harnesses: "cursor,codex,zed,opencode"
14
15
  license: "MIT"
15
- compatibility: "cursor,codex,zed,opencode"
16
16
  ---
17
17
 
18
18
  # Task Implementer
@@ -284,7 +284,25 @@ npm run build-storybook # Verify stories compile
284
284
  | Test failures | Fix failing tests, re-commit |
285
285
  | Story build failure | Fix story code, re-commit |
286
286
 
287
- Maximum 3 self-fix attempts per verification step.
287
+ Maximum 3 self-fix attempts per verification step.
288
+
289
+ Three, and it is the same three `job-orchestrator` and `flow-orchestrator`
290
+ use: one round bound, not four. *"The first three to four repair iterations
291
+ account for most achievable gains"*
292
+ ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197)); correctness falls
293
+ **0.820 -> 0.673** across two forced revisions while cumulative ever-correct is
294
+ **0.847** ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)) — the agent
295
+ finds the fix and then destroys it. Aider hardcodes `max_reflections = 3`;
296
+ OpenHands' critic uses 3.
297
+
298
+ **Stop earlier on repetition, whatever the count says.** If an attempt produces
299
+ the same failure output as the previous attempt — the same failing test with the
300
+ same message, the same type error at the same site — do NOT spend the remaining
301
+ attempts. The counter cannot tell "converging slowly" from "stuck", and three
302
+ identical outputs cost the whole budget to learn what the second one already
303
+ said. Report the block instead, naming what repeated.
304
+
305
+
288
306
  **ROLLBACK POLICY**: If implementation fatally fails (e.g. tests still failing after 3 attempts or unresolvable compilation errors), you MUST run `git reset --hard` to clean the worktree before reporting the failure in Phase 6, unless explicitly instructed to leave it dirty.
289
307
 
290
308
  **5.5 Re-commit fixes if any:**