@mrciphersmith/keryx 0.2.70 → 0.2.71
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +10357 -4676
- package/docs/README.md +54 -0
- package/docs/requirements/shared-agent-context/README.md +104 -0
- package/package.json +3 -2
- package/src/gdskills/bundled/rules/core/model-selection.mdc +184 -31
- package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +36 -0
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +78 -19
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.codex.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.cursor.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.opencode.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.zed.md +28 -3
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.codex.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.cursor.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.md +21 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.opencode.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.zed.md +20 -2
- package/src/gdskills/bundled/skills/planning/autodoc-analyst/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-architect/SKILL.md +3 -1
- package/src/gdskills/bundled/skills/planning/autodoc-assembler/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-orchestrator/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-scanner/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-writer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/docpack-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/docpack-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/metaproject-security/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/review-architecture/SKILL.md +37 -10
- package/src/gdskills/bundled/skills/review/review-backend/SKILL.md +48 -14
- package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +49 -12
- package/src/gdskills/bundled/skills/review/review-core-boundaries/SKILL.md +34 -2
- package/src/gdskills/bundled/skills/review/review-flow-graph/SKILL.md +33 -2
- package/src/gdskills/bundled/skills/review/review-frontend/SKILL.md +70 -29
- package/src/gdskills/bundled/skills/review/review-frontend-conventions/SKILL.md +34 -3
- package/src/gdskills/bundled/skills/review/review-highload/SKILL.md +49 -15
- package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +39 -11
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +590 -21
- package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +7 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/verification-claim.schema.json +78 -0
- package/src/gdskills/bundled/skills/review/review-performance/SKILL.md +43 -13
- package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +8 -2
- package/src/gdskills/bundled/skills/review/review-regression/SKILL.md +185 -0
- package/src/gdskills/bundled/skills/review/review-security-code/SKILL.md +44 -13
- package/src/gdskills/bundled/skills/review/review-style/SKILL.md +26 -6
- package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +35 -3
- package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +276 -0
- package/src/gdskills/contracts/review-finding.schema.json +119 -1
- package/src/gdskills/contracts/subagent-dispatch.schema.json +59 -3
- package/src/gdskills/bundled/skills/review/review-strict/SKILL.md +0 -328
|
@@ -14,8 +14,8 @@ metadata:
|
|
|
14
14
|
author: "MrCipherSmith"
|
|
15
15
|
version: "1.3.0"
|
|
16
16
|
category: "orchestration"
|
|
17
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
17
18
|
license: "MIT"
|
|
18
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
19
19
|
---
|
|
20
20
|
|
|
21
21
|
# Flow Orchestrator
|
|
@@ -84,8 +84,8 @@ flowchart TD
|
|
|
84
84
|
H -- "yes" --> I{"Ask user how to finish"}
|
|
85
85
|
I -- "create PR" --> J["create PR and run review/fix loop"]
|
|
86
86
|
J --> K{"review clean, PR mergeable?"}
|
|
87
|
-
K -- "no, attempts <
|
|
88
|
-
K -- "no, attempts =
|
|
87
|
+
K -- "no, attempts < 3" --> J
|
|
88
|
+
K -- "no, attempts = 3 or repetition detected" --> R["enrich context and change fix strategy"]
|
|
89
89
|
R --> J
|
|
90
90
|
K -- "yes" --> L["merge PR into recorded base branch"]
|
|
91
91
|
L --> M["keryx flow implemented --pr"]
|
|
@@ -128,10 +128,21 @@ it already tried. The flow package does.
|
|
|
128
128
|
```
|
|
129
129
|
|
|
130
130
|
5. Apply the Phase 4 attempt budget against the **persisted** count. If
|
|
131
|
-
`attempts.count` for the task has already reached
|
|
132
|
-
the same approach: go to the re-planning step (Phase 4, PR
|
|
133
|
-
loop, step 4) and record the decision in `journal.md`.
|
|
134
|
-
6.
|
|
131
|
+
`attempts.count` for the task has already reached **three**, do not
|
|
132
|
+
re-dispatch the same approach: go to the re-planning step (Phase 4, PR
|
|
133
|
+
review/fix loop, step 4) and record the decision in `journal.md`.
|
|
134
|
+
6. Run the repetition check before spending an attempt, whatever the count
|
|
135
|
+
says:
|
|
136
|
+
|
|
137
|
+
```bash
|
|
138
|
+
keryx review loop --flow <id> --task <Tn>
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
A non-zero exit means the same finding has recurred or two consecutive
|
|
142
|
+
rounds produced identical output. Go straight to the re-planning step.
|
|
143
|
+
Do not spend the remaining attempts on the same approach because the
|
|
144
|
+
budget has some left — that is the failure this check exists to catch.
|
|
145
|
+
7. If the flow is `blocked`, read the blocking reason from `journal.md`,
|
|
135
146
|
resolve or escalate it, then `keryx flow unblock <id>`.
|
|
136
147
|
4. If the user wants a new flow, continue at 0.1.
|
|
137
148
|
|
|
@@ -231,8 +242,13 @@ did not exist and 24 completed flows shipped with an open task:
|
|
|
231
242
|
`flow.json`, which `keryx flow init` writes for every flow it creates. A
|
|
232
243
|
package created before the gate landed does not carry the flag, and for it
|
|
233
244
|
the gate reports `skipped` and blocks nothing;
|
|
234
|
-
- a task fails the gate when its status is not `done
|
|
235
|
-
`failed
|
|
245
|
+
- a task fails the gate when its status is not `done`; when its disposition is
|
|
246
|
+
`failed`; when its disposition is `blocked` (terminal, but the work did not
|
|
247
|
+
happen — and the harness emits this disposition on its own for a run that
|
|
248
|
+
ended blocked); when its disposition is `skipped` with no recorded reason; or
|
|
249
|
+
when its disposition is a value this build does not recognise. An
|
|
250
|
+
unrecognised disposition FAILS rather than falling through: a gate whose
|
|
251
|
+
default for the unknown case is "pass" is not a gate;
|
|
236
252
|
- to close a task as deliberately not needed, record why:
|
|
237
253
|
|
|
238
254
|
```bash
|
|
@@ -346,7 +362,23 @@ Before accepting implementation:
|
|
|
346
362
|
1. Run focused tests for touched scope.
|
|
347
363
|
2. Run `code-verifier`.
|
|
348
364
|
3. Run `keryx health run` when Code Health is enabled.
|
|
349
|
-
4.
|
|
365
|
+
4. Check the bounds, then run `review-orchestrator` with relevant domains.
|
|
366
|
+
|
|
367
|
+
```bash
|
|
368
|
+
keryx review budget --spent <usd-so-far> --outstanding <subagents you already have in flight>
|
|
369
|
+
```
|
|
370
|
+
|
|
371
|
+
A non-zero exit means the spend ceiling (3 USD by default) has been reached:
|
|
372
|
+
**stop and ask the user** rather than dispatching another fan-out.
|
|
373
|
+
|
|
374
|
+
`--outstanding` is the part that matters here. `review-orchestrator`
|
|
375
|
+
dispatches reviewers in parallel and runs *nested* under this skill, and
|
|
376
|
+
keryx cannot observe subagents in another process. Passing the count you
|
|
377
|
+
already have in flight is the only thing that makes the concurrency cap mean
|
|
378
|
+
anything across the nesting; omit it and the cap bounds the reviewer fan-out
|
|
379
|
+
alone, which the review record then states plainly rather than implying
|
|
380
|
+
otherwise.
|
|
381
|
+
|
|
350
382
|
5. If findings require code changes, dispatch fix work through `task-implementer`
|
|
351
383
|
and record the fix task in the flow.
|
|
352
384
|
6. Close the skill-learning loop (see `rules/core/skill-lifecycle.mdc`). Collect
|
|
@@ -392,18 +424,45 @@ How should this flow end?
|
|
|
392
424
|
branch state.
|
|
393
425
|
2. If findings or required check failures remain, create or update a flow fix
|
|
394
426
|
task, dispatch `task-implementer`, push the fix, and run review again.
|
|
395
|
-
3. Allow at most
|
|
396
|
-
attempt when review/check results are available, including a clean result,
|
|
427
|
+
3. Allow at most **three** review/fix attempts for the current approach. Count
|
|
428
|
+
an attempt when review/check results are available, including a clean result,
|
|
397
429
|
and record it with `keryx flow task attempt <id> <Tn> --outcome ...` so the
|
|
398
430
|
count survives a session restart. Read the budget from that task's
|
|
399
431
|
`attempts.count` in `flow.json`, never from this session's memory.
|
|
400
|
-
|
|
401
|
-
|
|
402
|
-
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
432
|
+
|
|
433
|
+
Three, and the same three that `task-implementer` and `job-orchestrator`
|
|
434
|
+
already use. This skill said six, which was an outlier with nothing behind
|
|
435
|
+
it. The evidence converges on three: *"the first three to four repair
|
|
436
|
+
iterations account for most achievable gains"*
|
|
437
|
+
([arXiv:2607.05197](https://arxiv.org/abs/2607.05197)); correctness falls
|
|
438
|
+
**0.820 -> 0.673** across two forced revisions while cumulative ever-correct
|
|
439
|
+
is **0.847** ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)) — the
|
|
440
|
+
agent finds the fix and then destroys it, throwing away ~15 points by not
|
|
441
|
+
stopping. Aider hardcodes `max_reflections = 3`; OpenHands' critic uses 3.
|
|
442
|
+
Rounds four through six were not buying convergence; they were buying
|
|
443
|
+
regressions.
|
|
444
|
+
|
|
445
|
+
4. **Before** spending an attempt, and regardless of how much budget is left,
|
|
446
|
+
run the repetition check:
|
|
447
|
+
|
|
448
|
+
```bash
|
|
449
|
+
keryx review loop --flow <id> --task <Tn>
|
|
450
|
+
```
|
|
451
|
+
|
|
452
|
+
It escalates (non-zero exit) when the same finding recurs in two rounds, or
|
|
453
|
+
two consecutive rounds produce identical review output. It reads the review
|
|
454
|
+
packages on disk and the persisted `attempts.count`, not this session's
|
|
455
|
+
memory, and it deliberately never reads the remaining budget — an agent
|
|
456
|
+
emitting the identical failing output three times must be caught on the
|
|
457
|
+
second, not after the budget runs out.
|
|
458
|
+
|
|
459
|
+
5. If the third attempt is not clean, **or the repetition check escalated
|
|
460
|
+
earlier**, do not blindly repeat the same loop. Enrich context from the
|
|
461
|
+
findings, affected graph, relevant wiki, and health/testing artifacts;
|
|
462
|
+
identify the likely cycle cause; choose a materially different fix strategy
|
|
463
|
+
or split the work into narrower tasks; record the decision in `journal.md`;
|
|
464
|
+
then continue with the enriched context.
|
|
465
|
+
6. Never merge while findings or required checks remain unresolved. If the
|
|
407
466
|
re-planned approach still cannot produce a mergeable PR, leave the flow
|
|
408
467
|
`in-progress` and report the blocker instead of forcing completion.
|
|
409
468
|
|
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: job-documenter
|
|
3
|
+
model_tier: light
|
|
3
4
|
description: "Use when a job folder needs to be initialized, or analysis/report/review documents need to be created or updated in jobs/."
|
|
4
5
|
triggers:
|
|
5
6
|
- "Document job"
|
|
@@ -10,8 +11,8 @@ metadata:
|
|
|
10
11
|
author: "MrCipherSmith"
|
|
11
12
|
version: "1.0.0"
|
|
12
13
|
category: "documentation"
|
|
14
|
+
compatible_harnesses: "cursor,codex,zed,opencode"
|
|
13
15
|
license: "MIT"
|
|
14
|
-
compatibility: "cursor,codex,zed,opencode"
|
|
15
16
|
---
|
|
16
17
|
|
|
17
18
|
<SUBAGENT-STOP>
|
|
@@ -23,8 +23,8 @@ metadata:
|
|
|
23
23
|
author: "MrCipherSmith"
|
|
24
24
|
version: "3.2.0"
|
|
25
25
|
category: "orchestration"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
27
|
license: "MIT"
|
|
27
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
28
28
|
---
|
|
29
29
|
|
|
30
30
|
<SUBAGENT-STOP>
|
|
@@ -971,8 +971,22 @@ Review complete:
|
|
|
971
971
|
|
|
972
972
|
Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
|
|
973
973
|
|
|
974
|
+
Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
|
|
975
|
+
this skill all use it. *"The first three to four repair iterations account for
|
|
976
|
+
most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
|
|
977
|
+
correctness falls **0.820 -> 0.673** across two forced revisions while
|
|
978
|
+
cumulative ever-correct is **0.847**
|
|
979
|
+
([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
|
|
980
|
+
`max_reflections = 3`; OpenHands' critic uses 3.
|
|
981
|
+
|
|
982
|
+
The bound is a ceiling, not a target. Repetition ends the loop earlier and
|
|
983
|
+
**regardless of remaining iterations** — a counter cannot tell "converging
|
|
984
|
+
slowly" from "stuck", and an agent emitting the identical failing output three
|
|
985
|
+
times spends the whole budget before anything notices.
|
|
986
|
+
|
|
974
987
|
```
|
|
975
988
|
UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
|
|
989
|
+
PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
|
|
976
990
|
|
|
977
991
|
FOR iteration in [1, 2, 3]:
|
|
978
992
|
IF NOT NEEDS_FIX: BREAK
|
|
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
|
|
|
992
1006
|
6. Recompute NEEDS_FIX from new findings
|
|
993
1007
|
7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
|
|
994
1008
|
|
|
995
|
-
|
|
996
|
-
|
|
1009
|
+
8. STUCK CHECK — runs before the next iteration and ignores the budget:
|
|
1010
|
+
IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
|
|
1011
|
+
OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
|
|
1012
|
+
THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
|
|
1013
|
+
Identity is the finding's dedupe_key when it has one, otherwise
|
|
1014
|
+
reviewer + file + symbol + problem — never the display id, which is
|
|
1015
|
+
per-report and would fire on every second iteration whatever happened.
|
|
1016
|
+
9. PREVIOUS_REVIEW_OUTPUT = the new review output
|
|
1017
|
+
|
|
1018
|
+
IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
|
|
1019
|
+
Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
|
|
1020
|
+
two ended it — a budget exhausted and a loop detected call for different next
|
|
1021
|
+
steps → continue to checks
|
|
997
1022
|
```
|
|
998
1023
|
|
|
999
1024
|
**Fix prompt escalation pattern:**
|
|
@@ -23,8 +23,8 @@ metadata:
|
|
|
23
23
|
author: "MrCipherSmith"
|
|
24
24
|
version: "3.2.0"
|
|
25
25
|
category: "orchestration"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
27
|
license: "MIT"
|
|
27
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
28
28
|
---
|
|
29
29
|
|
|
30
30
|
<SUBAGENT-STOP>
|
|
@@ -971,8 +971,22 @@ Review complete:
|
|
|
971
971
|
|
|
972
972
|
Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
|
|
973
973
|
|
|
974
|
+
Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
|
|
975
|
+
this skill all use it. *"The first three to four repair iterations account for
|
|
976
|
+
most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
|
|
977
|
+
correctness falls **0.820 -> 0.673** across two forced revisions while
|
|
978
|
+
cumulative ever-correct is **0.847**
|
|
979
|
+
([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
|
|
980
|
+
`max_reflections = 3`; OpenHands' critic uses 3.
|
|
981
|
+
|
|
982
|
+
The bound is a ceiling, not a target. Repetition ends the loop earlier and
|
|
983
|
+
**regardless of remaining iterations** — a counter cannot tell "converging
|
|
984
|
+
slowly" from "stuck", and an agent emitting the identical failing output three
|
|
985
|
+
times spends the whole budget before anything notices.
|
|
986
|
+
|
|
974
987
|
```
|
|
975
988
|
UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
|
|
989
|
+
PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
|
|
976
990
|
|
|
977
991
|
FOR iteration in [1, 2, 3]:
|
|
978
992
|
IF NOT NEEDS_FIX: BREAK
|
|
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
|
|
|
992
1006
|
6. Recompute NEEDS_FIX from new findings
|
|
993
1007
|
7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
|
|
994
1008
|
|
|
995
|
-
|
|
996
|
-
|
|
1009
|
+
8. STUCK CHECK — runs before the next iteration and ignores the budget:
|
|
1010
|
+
IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
|
|
1011
|
+
OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
|
|
1012
|
+
THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
|
|
1013
|
+
Identity is the finding's dedupe_key when it has one, otherwise
|
|
1014
|
+
reviewer + file + symbol + problem — never the display id, which is
|
|
1015
|
+
per-report and would fire on every second iteration whatever happened.
|
|
1016
|
+
9. PREVIOUS_REVIEW_OUTPUT = the new review output
|
|
1017
|
+
|
|
1018
|
+
IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
|
|
1019
|
+
Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
|
|
1020
|
+
two ended it — a budget exhausted and a loop detected call for different next
|
|
1021
|
+
steps → continue to checks
|
|
997
1022
|
```
|
|
998
1023
|
|
|
999
1024
|
**Fix prompt escalation pattern:**
|
|
@@ -23,8 +23,8 @@ metadata:
|
|
|
23
23
|
author: "MrCipherSmith"
|
|
24
24
|
version: "3.2.0"
|
|
25
25
|
category: "orchestration"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
27
|
license: "MIT"
|
|
27
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
28
28
|
---
|
|
29
29
|
|
|
30
30
|
<SUBAGENT-STOP>
|
|
@@ -973,8 +973,22 @@ Review complete:
|
|
|
973
973
|
|
|
974
974
|
Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
|
|
975
975
|
|
|
976
|
+
Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
|
|
977
|
+
this skill all use it. *"The first three to four repair iterations account for
|
|
978
|
+
most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
|
|
979
|
+
correctness falls **0.820 -> 0.673** across two forced revisions while
|
|
980
|
+
cumulative ever-correct is **0.847**
|
|
981
|
+
([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
|
|
982
|
+
`max_reflections = 3`; OpenHands' critic uses 3.
|
|
983
|
+
|
|
984
|
+
The bound is a ceiling, not a target. Repetition ends the loop earlier and
|
|
985
|
+
**regardless of remaining iterations** — a counter cannot tell "converging
|
|
986
|
+
slowly" from "stuck", and an agent emitting the identical failing output three
|
|
987
|
+
times spends the whole budget before anything notices.
|
|
988
|
+
|
|
976
989
|
```
|
|
977
990
|
UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
|
|
991
|
+
PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
|
|
978
992
|
|
|
979
993
|
FOR iteration in [1, 2, 3]:
|
|
980
994
|
IF NOT NEEDS_FIX: BREAK
|
|
@@ -994,8 +1008,19 @@ FOR iteration in [1, 2, 3]:
|
|
|
994
1008
|
6. Recompute NEEDS_FIX from new findings
|
|
995
1009
|
7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
|
|
996
1010
|
|
|
997
|
-
|
|
998
|
-
|
|
1011
|
+
8. STUCK CHECK — runs before the next iteration and ignores the budget:
|
|
1012
|
+
IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
|
|
1013
|
+
OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
|
|
1014
|
+
THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
|
|
1015
|
+
Identity is the finding's dedupe_key when it has one, otherwise
|
|
1016
|
+
reviewer + file + symbol + problem — never the display id, which is
|
|
1017
|
+
per-report and would fire on every second iteration whatever happened.
|
|
1018
|
+
9. PREVIOUS_REVIEW_OUTPUT = the new review output
|
|
1019
|
+
|
|
1020
|
+
IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
|
|
1021
|
+
Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
|
|
1022
|
+
two ended it — a budget exhausted and a loop detected call for different next
|
|
1023
|
+
steps → continue to checks
|
|
999
1024
|
```
|
|
1000
1025
|
|
|
1001
1026
|
**Fix prompt escalation pattern:**
|
|
@@ -23,8 +23,8 @@ metadata:
|
|
|
23
23
|
author: "MrCipherSmith"
|
|
24
24
|
version: "3.2.0"
|
|
25
25
|
category: "orchestration"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
27
|
license: "MIT"
|
|
27
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
28
28
|
---
|
|
29
29
|
|
|
30
30
|
<SUBAGENT-STOP>
|
|
@@ -971,8 +971,22 @@ Review complete:
|
|
|
971
971
|
|
|
972
972
|
Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
|
|
973
973
|
|
|
974
|
+
Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
|
|
975
|
+
this skill all use it. *"The first three to four repair iterations account for
|
|
976
|
+
most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
|
|
977
|
+
correctness falls **0.820 -> 0.673** across two forced revisions while
|
|
978
|
+
cumulative ever-correct is **0.847**
|
|
979
|
+
([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
|
|
980
|
+
`max_reflections = 3`; OpenHands' critic uses 3.
|
|
981
|
+
|
|
982
|
+
The bound is a ceiling, not a target. Repetition ends the loop earlier and
|
|
983
|
+
**regardless of remaining iterations** — a counter cannot tell "converging
|
|
984
|
+
slowly" from "stuck", and an agent emitting the identical failing output three
|
|
985
|
+
times spends the whole budget before anything notices.
|
|
986
|
+
|
|
974
987
|
```
|
|
975
988
|
UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
|
|
989
|
+
PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
|
|
976
990
|
|
|
977
991
|
FOR iteration in [1, 2, 3]:
|
|
978
992
|
IF NOT NEEDS_FIX: BREAK
|
|
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
|
|
|
992
1006
|
6. Recompute NEEDS_FIX from new findings
|
|
993
1007
|
7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
|
|
994
1008
|
|
|
995
|
-
|
|
996
|
-
|
|
1009
|
+
8. STUCK CHECK — runs before the next iteration and ignores the budget:
|
|
1010
|
+
IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
|
|
1011
|
+
OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
|
|
1012
|
+
THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
|
|
1013
|
+
Identity is the finding's dedupe_key when it has one, otherwise
|
|
1014
|
+
reviewer + file + symbol + problem — never the display id, which is
|
|
1015
|
+
per-report and would fire on every second iteration whatever happened.
|
|
1016
|
+
9. PREVIOUS_REVIEW_OUTPUT = the new review output
|
|
1017
|
+
|
|
1018
|
+
IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
|
|
1019
|
+
Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
|
|
1020
|
+
two ended it — a budget exhausted and a loop detected call for different next
|
|
1021
|
+
steps → continue to checks
|
|
997
1022
|
```
|
|
998
1023
|
|
|
999
1024
|
**Fix prompt escalation pattern:**
|
|
@@ -23,8 +23,8 @@ metadata:
|
|
|
23
23
|
author: "MrCipherSmith"
|
|
24
24
|
version: "3.2.0"
|
|
25
25
|
category: "orchestration"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
27
|
license: "MIT"
|
|
27
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
28
28
|
---
|
|
29
29
|
|
|
30
30
|
<SUBAGENT-STOP>
|
|
@@ -971,8 +971,22 @@ Review complete:
|
|
|
971
971
|
|
|
972
972
|
Only runs if NEEDS_FIX is true. Default max: **3 iterations** (`max_review_iterations`).
|
|
973
973
|
|
|
974
|
+
Three is the shared round bound: `task-implementer`, `flow-orchestrator` and
|
|
975
|
+
this skill all use it. *"The first three to four repair iterations account for
|
|
976
|
+
most achievable gains"* ([arXiv:2607.05197](https://arxiv.org/abs/2607.05197));
|
|
977
|
+
correctness falls **0.820 -> 0.673** across two forced revisions while
|
|
978
|
+
cumulative ever-correct is **0.847**
|
|
979
|
+
([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)). Aider hardcodes
|
|
980
|
+
`max_reflections = 3`; OpenHands' critic uses 3.
|
|
981
|
+
|
|
982
|
+
The bound is a ceiling, not a target. Repetition ends the loop earlier and
|
|
983
|
+
**regardless of remaining iterations** — a counter cannot tell "converging
|
|
984
|
+
slowly" from "stuck", and an agent emitting the identical failing output three
|
|
985
|
+
times spends the whole budget before anything notices.
|
|
986
|
+
|
|
974
987
|
```
|
|
975
988
|
UNRESOLVED_FINDINGS = all CRITICAL + WARNING findings from step 2.6
|
|
989
|
+
PREVIOUS_REVIEW_OUTPUT = <the review output from step 2.6>
|
|
976
990
|
|
|
977
991
|
FOR iteration in [1, 2, 3]:
|
|
978
992
|
IF NOT NEEDS_FIX: BREAK
|
|
@@ -992,8 +1006,19 @@ FOR iteration in [1, 2, 3]:
|
|
|
992
1006
|
6. Recompute NEEDS_FIX from new findings
|
|
993
1007
|
7. Update UNRESOLVED_FINDINGS = remaining CRITICAL + WARNING
|
|
994
1008
|
|
|
995
|
-
|
|
996
|
-
|
|
1009
|
+
8. STUCK CHECK — runs before the next iteration and ignores the budget:
|
|
1010
|
+
IF any finding identity is in UNRESOLVED_FINDINGS for the SECOND iteration
|
|
1011
|
+
OR the new review output is identical to PREVIOUS_REVIEW_OUTPUT
|
|
1012
|
+
THEN log "stuck: <what repeated>" and BREAK, even with iterations left.
|
|
1013
|
+
Identity is the finding's dedupe_key when it has one, otherwise
|
|
1014
|
+
reviewer + file + symbol + problem — never the display id, which is
|
|
1015
|
+
per-report and would fire on every second iteration whatever happened.
|
|
1016
|
+
9. PREVIOUS_REVIEW_OUTPUT = the new review output
|
|
1017
|
+
|
|
1018
|
+
IF still NEEDS_FIX after max iterations, or the stuck check broke the loop:
|
|
1019
|
+
Log "Unresolved after <N> iterations" with finding list, and say WHICH of the
|
|
1020
|
+
two ended it — a budget exhausted and a loop detected call for different next
|
|
1021
|
+
steps → continue to checks
|
|
997
1022
|
```
|
|
998
1023
|
|
|
999
1024
|
**Fix prompt escalation pattern:**
|
|
@@ -11,8 +11,8 @@ metadata:
|
|
|
11
11
|
author: "MrCipherSmith"
|
|
12
12
|
version: "1.0.0"
|
|
13
13
|
category: "implementation"
|
|
14
|
+
compatible_harnesses: "cursor,codex,zed,opencode"
|
|
14
15
|
license: "MIT"
|
|
15
|
-
compatibility: "cursor,codex,zed,opencode"
|
|
16
16
|
---
|
|
17
17
|
|
|
18
18
|
# Task Implementer
|
|
@@ -284,7 +284,25 @@ npm run build-storybook # Verify stories compile
|
|
|
284
284
|
| Test failures | Fix failing tests, re-commit |
|
|
285
285
|
| Story build failure | Fix story code, re-commit |
|
|
286
286
|
|
|
287
|
-
Maximum 3 self-fix attempts per verification step.
|
|
287
|
+
Maximum 3 self-fix attempts per verification step.
|
|
288
|
+
|
|
289
|
+
Three, and it is the same three `job-orchestrator` and `flow-orchestrator`
|
|
290
|
+
use: one round bound, not four. *"The first three to four repair iterations
|
|
291
|
+
account for most achievable gains"*
|
|
292
|
+
([arXiv:2607.05197](https://arxiv.org/abs/2607.05197)); correctness falls
|
|
293
|
+
**0.820 -> 0.673** across two forced revisions while cumulative ever-correct is
|
|
294
|
+
**0.847** ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)) — the agent
|
|
295
|
+
finds the fix and then destroys it. Aider hardcodes `max_reflections = 3`;
|
|
296
|
+
OpenHands' critic uses 3.
|
|
297
|
+
|
|
298
|
+
**Stop earlier on repetition, whatever the count says.** If an attempt produces
|
|
299
|
+
the same failure output as the previous attempt — the same failing test with the
|
|
300
|
+
same message, the same type error at the same site — do NOT spend the remaining
|
|
301
|
+
attempts. The counter cannot tell "converging slowly" from "stuck", and three
|
|
302
|
+
identical outputs cost the whole budget to learn what the second one already
|
|
303
|
+
said. Report the block instead, naming what repeated.
|
|
304
|
+
|
|
305
|
+
|
|
288
306
|
**ROLLBACK POLICY**: If implementation fatally fails (e.g. tests still failing after 3 attempts or unresolvable compilation errors), you MUST run `git reset --hard` to clean the worktree before reporting the failure in Phase 6, unless explicitly instructed to leave it dirty.
|
|
289
307
|
|
|
290
308
|
**5.5 Re-commit fixes if any:**
|
|
@@ -11,8 +11,8 @@ metadata:
|
|
|
11
11
|
author: "MrCipherSmith"
|
|
12
12
|
version: "1.0.0"
|
|
13
13
|
category: "implementation"
|
|
14
|
+
compatible_harnesses: "cursor,codex,zed,opencode"
|
|
14
15
|
license: "MIT"
|
|
15
|
-
compatibility: "cursor,codex,zed,opencode"
|
|
16
16
|
---
|
|
17
17
|
|
|
18
18
|
# Task Implementer
|
|
@@ -284,7 +284,25 @@ npm run build-storybook # Verify stories compile
|
|
|
284
284
|
| Test failures | Fix failing tests, re-commit |
|
|
285
285
|
| Story build failure | Fix story code, re-commit |
|
|
286
286
|
|
|
287
|
-
Maximum 3 self-fix attempts per verification step.
|
|
287
|
+
Maximum 3 self-fix attempts per verification step.
|
|
288
|
+
|
|
289
|
+
Three, and it is the same three `job-orchestrator` and `flow-orchestrator`
|
|
290
|
+
use: one round bound, not four. *"The first three to four repair iterations
|
|
291
|
+
account for most achievable gains"*
|
|
292
|
+
([arXiv:2607.05197](https://arxiv.org/abs/2607.05197)); correctness falls
|
|
293
|
+
**0.820 -> 0.673** across two forced revisions while cumulative ever-correct is
|
|
294
|
+
**0.847** ([arXiv:2607.24604](https://arxiv.org/abs/2607.24604)) — the agent
|
|
295
|
+
finds the fix and then destroys it. Aider hardcodes `max_reflections = 3`;
|
|
296
|
+
OpenHands' critic uses 3.
|
|
297
|
+
|
|
298
|
+
**Stop earlier on repetition, whatever the count says.** If an attempt produces
|
|
299
|
+
the same failure output as the previous attempt — the same failing test with the
|
|
300
|
+
same message, the same type error at the same site — do NOT spend the remaining
|
|
301
|
+
attempts. The counter cannot tell "converging slowly" from "stuck", and three
|
|
302
|
+
identical outputs cost the whole budget to learn what the second one already
|
|
303
|
+
said. Report the block instead, naming what repeated.
|
|
304
|
+
|
|
305
|
+
|
|
288
306
|
**ROLLBACK POLICY**: If implementation fatally fails (e.g. tests still failing after 3 attempts or unresolvable compilation errors), you MUST run `git reset --hard` to clean the worktree before reporting the failure in Phase 6, unless explicitly instructed to leave it dirty.
|
|
289
307
|
|
|
290
308
|
**5.5 Re-commit fixes if any:**
|