lazycodex-ai 5.0.0-beta.40 → 5.0.0-beta.42

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (101) hide show
  1. package/dist/cli/index.js +21 -15
  2. package/dist/cli-node/index.js +21 -15
  3. package/package.json +1 -1
  4. package/packages/lsp-daemon/dist/cli.js +11 -6
  5. package/packages/lsp-daemon/dist/client.js +11 -6
  6. package/packages/lsp-daemon/dist/index.js +11 -6
  7. package/packages/lsp-tools-mcp/dist/cli.js +11 -6
  8. package/packages/lsp-tools-mcp/dist/lsp/manager.js +11 -6
  9. package/packages/lsp-tools-mcp/dist/mcp.js +11 -6
  10. package/packages/lsp-tools-mcp/dist/tools.js +11 -6
  11. package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
  12. package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
  13. package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
  14. package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
  15. package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
  16. package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
  17. package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
  18. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
  19. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
  20. package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +3 -3
  21. package/packages/omo-codex/plugin/components/lsp/dist/cli.js +11 -6
  22. package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
  23. package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
  24. package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
  25. package/packages/omo-codex/plugin/components/rules/package.json +1 -1
  26. package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
  27. package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
  28. package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
  29. package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
  30. package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
  31. package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
  32. package/packages/omo-codex/plugin/components/ulw-execute-continuation/hooks/hooks.json +2 -2
  33. package/packages/omo-codex/plugin/components/ulw-execute-continuation/package.json +1 -1
  34. package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +4 -4
  35. package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
  36. package/packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/SKILL.md +1 -1
  37. package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
  38. package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
  39. package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
  40. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
  41. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
  42. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
  43. package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
  44. package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
  45. package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
  46. package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
  47. package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
  48. package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
  49. package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
  50. package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
  51. package/packages/omo-codex/plugin/hooks/stop-checking-ulw-execute-continuation.json +1 -1
  52. package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
  53. package/packages/omo-codex/plugin/hooks/subagent-stop-checking-ulw-execute-continuation.json +1 -1
  54. package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
  55. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
  56. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
  57. package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
  58. package/packages/omo-codex/plugin/package-lock.json +12 -12
  59. package/packages/omo-codex/plugin/package.json +1 -1
  60. package/packages/omo-codex/plugin/scripts/sync-skills.mjs +6 -18
  61. package/packages/omo-codex/plugin/skills/ast-grep/SKILL.md +1 -1
  62. package/packages/omo-codex/plugin/skills/coding-agent-sessions/SKILL.md +1 -1
  63. package/packages/omo-codex/plugin/skills/data-scientist/SKILL.md +1 -1
  64. package/packages/omo-codex/plugin/skills/debugging/SKILL.md +1 -1
  65. package/packages/omo-codex/plugin/skills/frontend/SKILL.md +1 -1
  66. package/packages/omo-codex/plugin/skills/git-master/SKILL.md +1 -1
  67. package/packages/omo-codex/plugin/skills/init-deep/SKILL.md +1 -1
  68. package/packages/omo-codex/plugin/skills/lsp-setup/SKILL.md +1 -1
  69. package/packages/omo-codex/plugin/skills/programming/SKILL.md +1 -1
  70. package/packages/omo-codex/plugin/skills/refactor/SKILL.md +1 -1
  71. package/packages/omo-codex/plugin/skills/remove-ai-slops/SKILL.md +1 -1
  72. package/packages/omo-codex/plugin/skills/review-work/SKILL.md +99 -417
  73. package/packages/omo-codex/plugin/skills/ultimate-browsing/ATTRIBUTION.md +1 -1
  74. package/packages/omo-codex/plugin/skills/ultimate-browsing/SKILL.md +24 -9
  75. package/packages/omo-codex/plugin/skills/ultimate-browsing/references/chrome-stealth.md +13 -10
  76. package/packages/omo-codex/plugin/skills/ulw-execute/SKILL.md +1 -1
  77. package/packages/omo-codex/plugin/skills/ulw-loop/SKILL.md +1 -1
  78. package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +5 -2
  79. package/packages/omo-codex/plugin/skills/visual-qa/SKILL.md +2 -2
  80. package/packages/omo-codex/plugin/test/sync-skills-test-support.mjs +1 -1
  81. package/packages/omo-codex/scripts/install-dist/install-local.mjs +3 -2
  82. package/packages/shared-skills/package.json +7 -1
  83. package/packages/shared-skills/skills/ast-grep/SKILL.md +1 -1
  84. package/packages/shared-skills/skills/coding-agent-sessions/SKILL.md +1 -1
  85. package/packages/shared-skills/skills/data-scientist/SKILL.md +1 -1
  86. package/packages/shared-skills/skills/debugging/SKILL.md +1 -1
  87. package/packages/shared-skills/skills/frontend/SKILL.md +1 -1
  88. package/packages/shared-skills/skills/git-master/SKILL.md +1 -1
  89. package/packages/shared-skills/skills/init-deep/SKILL.md +1 -1
  90. package/packages/shared-skills/skills/lsp-setup/SKILL.md +1 -1
  91. package/packages/shared-skills/skills/programming/SKILL.md +1 -1
  92. package/packages/shared-skills/skills/refactor/SKILL.md +1 -1
  93. package/packages/shared-skills/skills/remove-ai-slops/SKILL.md +1 -1
  94. package/packages/shared-skills/skills/review-work/SKILL.md +99 -417
  95. package/packages/shared-skills/skills/ultimate-browsing/ATTRIBUTION.md +1 -1
  96. package/packages/shared-skills/skills/ultimate-browsing/SKILL.md +24 -9
  97. package/packages/shared-skills/skills/ultimate-browsing/references/chrome-stealth.md +13 -10
  98. package/packages/shared-skills/skills/ulw-execute/SKILL.md +1 -1
  99. package/packages/shared-skills/skills/ulw-plan/SKILL.md +1 -1
  100. package/packages/shared-skills/skills/ulw-research/SKILL.md +5 -2
  101. package/packages/shared-skills/skills/visual-qa/SKILL.md +2 -2
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: review-work
3
- description: "Post-implementation review orchestrator. Launches 5 parallel background sub-agents: Oracle (goal/constraint verification), Oracle (code quality), Oracle (security), unspecified-high (hands-on QA execution), unspecified-high (context mining from GitHub/git/Slack/Notion). All must pass for review to pass. MUST USE before a PR handoff or when the user explicitly asks to review completed work. Triggers: 'review work', 'review my work', 'review changes', 'QA my work', 'verify implementation', 'check my work', 'validate changes', 'post-implementation review'."
3
+ description: "Post-implementation gate review: run manual QA on the real surface yourself, then launch ONE gate reviewer (never a panel) to audit goal, constraints, code quality, security, missed context, and QA evidence. Use before a PR handoff or when the user explicitly asks to review completed work."
4
4
  ---
5
5
  ## Codex Harness Tool Compatibility
6
6
 
@@ -59,9 +59,9 @@ deliverable. Preserve completed lane results immediately. If the retry
59
59
  budget is exhausted, keep the lane `INCONCLUSIVE` and still emit a final
60
60
  aggregate result.
61
61
 
62
- # Review Work - 5-Agent Parallel Review Orchestrator
62
+ # Review Work - Gate Review Orchestrator
63
63
 
64
- Launch 5 specialized sub-agents in parallel to review completed implementation work from every angle. All 5 must pass for the review to pass. If even ONE fails, the review fails.
64
+ Review completed implementation work through exactly two lanes: your own hands-on manual QA on the real surface, and ONE gate reviewer sub-agent that audits the whole change set against the goal, the constraints, and your QA evidence. The review passes only when the QA matrix has no failing row AND the gate reviewer returns APPROVE.
65
65
 
66
66
  When `review-work` is used as a final implementation, PR, or `$ulw-execute`
67
67
  gate, it is blocking. A timeout, missing deliverable, ack-only response,
@@ -93,21 +93,18 @@ include raw tokens, credentials, auth headers, cookies, API keys, env dumps,
93
93
  private logs, or PII; summarize with lengths, hashes, and short non-sensitive
94
94
  prefixes when identity is needed.
95
95
 
96
- The 5 agents cover complementary concerns - together they form a comprehensive review that no single reviewer could match:
96
+ One reviewer, not a panel. A single gate reviewer holding the full context (goal, diff, history, QA evidence) catches what a fan-out of narrow reviewers misses between their seams, and it costs one agent instead of five. Never add review lanes; widen the gate reviewer's checklist instead.
97
97
 
98
- | # | Agent | Type | Role | Focus Level |
99
- |---|-------|------|------|-------------|
100
- | 1 | Goal Verifier | Oracle | Did we build what was asked? | MAIN |
101
- | 2 | QA Executor | unspecified-high | Does it actually work? | MAIN |
102
- | 3 | Code Reviewer | Oracle | Is the code well-written? | MAIN |
103
- | 4 | Security Auditor | Oracle | Is it secure? | SUB |
104
- | 5 | Context Miner | unspecified-high | Did we miss any context? | MAIN |
98
+ | Lane | Who runs it | Question it answers |
99
+ |------|-------------|---------------------|
100
+ | Manual QA | You, the orchestrator, on the real surface | Does it actually work? |
101
+ | Gate review | One gate reviewer sub-agent (`oracle` on OpenCode; the surface's gate-reviewer agent elsewhere) | Did we build what was asked - correctly, safely, well, and without missing context? |
105
102
 
106
103
  ---
107
104
 
108
105
  ## Phase 0: Gather Review Context
109
106
 
110
- Before launching agents, collect these inputs. Extract from conversation history first - the user's original request, constraints discussed, and decisions made are usually already in the thread. Only ask if truly missing.
107
+ Before running anything, collect these inputs. Extract from conversation history first - the user's original request, constraints discussed, and decisions made are usually already in the thread. Only ask if truly missing.
111
108
 
112
109
  <required_inputs>
113
110
 
@@ -116,12 +113,12 @@ Before launching agents, collect these inputs. Extract from conversation history
116
113
  - **BACKGROUND**: Why this work was needed. Business context, user stories, related systems, prior decisions that informed the approach.
117
114
  - **CHANGED_FILES**: Auto-collect via `git diff --name-only HEAD~1` or against the appropriate base (branch point, specific commit).
118
115
  - **DIFF**: Auto-collect via `git diff HEAD~1` or against the appropriate base.
119
- - **FILE_CONTENTS**: Read the full content of each changed file (not just the diff). Oracle agents cannot read files - they need full context in the prompt.
116
+ - **FILE_CONTENTS**: The full content of each changed file plus the neighboring files that show the established patterns. Required verbatim when the reviewer cannot read files (`oracle`); when your surface's gate reviewer can read files and run commands, pass the paths and the diff instead of pasting everything.
120
117
  - **RUN_COMMAND**: How to start/run the application. Check `package.json` scripts, `Makefile`, `docker-compose.yml`, or ask the user.
118
+ - **CONTEXT_MINING**: What the history and the trackers say about this area (collected below).
121
119
 
122
120
  </required_inputs>
123
121
 
124
-
125
122
  Review PRs and branches from a dedicated review worktree only: create or attach one with `git worktree add <path> <branch>` before collecting changed files, diff, file contents, or running checks, then immediately lock it with `git worktree lock <path> --reason "review:<pr-or-branch>"`. The main worktree is read-only context; never checkout, test, or edit the review branch there.
126
123
 
127
124
  **Auto-collection sequence:**
@@ -137,34 +134,61 @@ git diff HEAD~1 # or: git diff main...HEAD
137
134
  # Check package.json -> "scripts.dev" or "scripts.start"
138
135
  # Check Makefile -> default target
139
136
  # Check docker-compose.yml -> services
137
+
138
+ # 4. Mine the context the implementation may have missed (keep the output short)
139
+ git log --oneline -20 -- <each changed file> # recent changes and their reasons
140
+ git log --all --oneline --grep="<keywords from goal>" # related commits, reverts
141
+ gh issue list --search "<keywords>" --state all # related issues (when gh is available)
142
+ gh pr list --search "<keywords>" --state all # related PRs and their review comments
143
+ rg -n "TODO|FIXME|HACK" <changed files> # warnings left by previous authors
144
+ # plus: files that import the changed modules, tests touching the same paths,
145
+ # docs and config that reference the changed behavior
140
146
  ```
141
147
 
148
+ Record CONTEXT_MINING as a short list: source -> finding -> why it matters for this change. Slack, Notion, and Discord searches belong here too when those tools exist.
149
+
142
150
  For GOAL, CONSTRAINTS, BACKGROUND - review the full conversation history. The user's original message almost always contains the goal. Constraints often emerge during discussion. If anything critical is ambiguous, ask ONE focused question - not a checklist.
143
151
 
144
152
  ---
145
153
 
146
- ## Phase 1: Launch 5 Agents
154
+ ## Phase 1: Manual QA (you run it)
155
+
156
+ You are the QA lane. Do not delegate hands-on QA to a sub-agent: the orchestrator owns the real-surface proof, exactly as the ulw-loop final gate records `manualQa` under the main session.
157
+
158
+ 1. **Reuse first.** If this session already captured real-surface evidence for the FINAL tree (an ultrawork or ulw-loop evidence directory, a `visual-qa` verdict on this same build), consume it as QA rows instead of re-running. A fix committed after a capture stales that capture: re-run the rows it covered.
159
+ 2. **Pick the channel that faithfully exercises the surface** and capture the artifact:
160
+ - HTTP: `curl -i` (or an API request context) - status line, headers, body.
161
+ - CLI / TUI: a real pty - drive the command and keep the transcript; for color or layout evidence render through a browser-based terminal, never a `tmux capture-pane` dump.
162
+ - Web: the real page in a real browser - the harness's in-process surface (a `Bun.WebView` / `playwright-core` code cell, or Codex's Browser plugin) or the agent-browser CLI - action log plus screenshot.
163
+ - Desktop / GUI: OS-level automation against the running app - action log plus screenshot.
164
+ - Library / SDK: a script that imports and exercises the public API - transcript.
165
+ - Data-shaped work (migrations, configs, generated files): the resulting artifact itself, diffed or dumped.
166
+ 3. **Cover at least**: the happy path the goal names, the riskiest edge (empty, boundary, malformed, or concurrent input), and one regression on adjacent behavior the change could have broken. Add a row for every stated success criterion.
167
+ 4. **Build the QA matrix** - one row per scenario:
147
168
 
148
- Launch ALL 5 in a single turn. Every agent uses `run_in_background=true`. No sequential launches. No waiting between them.
169
+ | # | Scenario | Exact command / action | Expected | Observed | Verdict | Artifact |
170
+ |---|----------|------------------------|----------|----------|---------|----------|
149
171
 
150
- **Oracle agents receive everything in the prompt** (they cannot read files or run commands). Include DIFF + FILE_CONTENTS + all context directly in the prompt text.
172
+ A row without an artifact path is not PASS. If the application cannot even start, that is an immediate FAIL.
151
173
 
152
- **unspecified-high agents are autonomous** - they can read files, run commands, and use tools. Give them goals and pointers, not raw content dumps.
174
+ Any FAIL ends the review here: report **REVIEW FAILED** with the failing rows and skip the gate reviewer - reviewing code that does not work wastes the reviewer. Fix first, then re-enter at Phase 0 with the delta.
153
175
 
154
176
  ---
155
177
 
156
- ### Agent 1: Goal & Constraint Verification (Oracle) - MAIN
178
+ ## Phase 2: Launch the Gate Reviewer (one agent)
157
179
 
158
- This agent answers: "Did we build exactly what was asked, within the rules we were given?"
180
+ Launch exactly one reviewer, in the background, then keep doing independent root work (teardown prep, report scaffolding) while it runs.
181
+
182
+ `oracle` cannot read files or run commands: it receives everything inline (DIFF + FILE_CONTENTS + CONTEXT_MINING + the QA matrix). If your surface's gate reviewer has read and shell tools, still paste the diff and the QA matrix, and hand it file paths instead of full contents.
159
183
 
160
184
  ```
161
185
  task(
162
186
  subagent_type="oracle",
163
187
  run_in_background=true,
164
188
  load_skills=[],
165
- description="Verify implementation against original goal and constraints",
189
+ description="Gate-review the completed work against goal, constraints, and QA evidence",
166
190
  prompt="""
167
- <review_type>GOAL & CONSTRAINT VERIFICATION</review_type>
191
+ <review_type>GATE REVIEW</review_type>
168
192
 
169
193
  <original_goal>
170
194
  {GOAL - paste the user's original request and any clarifications}
@@ -183,425 +207,86 @@ task(
183
207
  </changed_files>
184
208
 
185
209
  <file_contents>
186
- {FILE_CONTENTS - full content of every changed file, clearly delimited per file}
210
+ {FILE_CONTENTS - full content of every changed file plus neighboring files that show existing patterns; or the paths, when the reviewer can read files}
187
211
  </file_contents>
188
212
 
189
213
  <diff>
190
214
  {DIFF - the actual git diff}
191
215
  </diff>
192
216
 
193
- Review whether this implementation correctly and completely achieves the stated goal within the given constraints. Be obsessively thorough - the point of this review is to catch what the implementer missed.
194
-
195
- REVIEW CHECKLIST:
217
+ <context_mining>
218
+ {CONTEXT_MINING - git history, related issues and PRs, docs and config that reference the changed behavior, warnings from previous authors}
219
+ </context_mining>
196
220
 
197
- 1. **Goal Completeness**: Break the goal into every sub-requirement (explicit AND implied). For each, mark ACHIEVED / MISSED / PARTIAL. Missing even one implied requirement that a reasonable engineer would have addressed = PARTIAL at minimum.
221
+ <manual_qa_matrix>
222
+ {QA MATRIX - every row with its artifact path}
223
+ </manual_qa_matrix>
198
224
 
199
- 2. **Constraint Compliance**: List every constraint. For each, verify compliance with specific code evidence. A constraint violated = automatic FAIL.
225
+ Role: final gate reviewer. You do not implement fixes. Assume every success claim is unverified until you reproduce it from the artifacts: executors can be wrong, tests can be too narrow, success prose can be misleading.
200
226
 
201
- 3. **Requirement Gaps**: Requirements the user clearly wanted but didn't spell out. Things implied by the goal or background that a thoughtful engineer would have included.
227
+ Review from the user's perspective first: what did they originally want, what result did they expect to receive, and does the shipped change actually satisfy that outcome? Then work the checklist. Be obsessively thorough - the point of this review is to catch what the implementer missed.
202
228
 
203
- 4. **Over-Engineering**: Anything added that wasn't requested - unnecessary abstractions, extra features, premature optimizations, speculative generality. Flag these as scope creep.
229
+ REVIEW CHECKLIST:
204
230
 
205
- 5. **Edge Cases**: Given the goal, what inputs or scenarios would break this? Trace through at least 5 edge cases mentally.
231
+ 1. **Goal completeness**: break the goal into every sub-requirement (explicit AND implied). Mark each ACHIEVED / MISSED / PARTIAL with code evidence. An implied requirement a reasonable engineer would have addressed counts.
232
+ 2. **Constraint compliance**: list every constraint and verify each with specific evidence. A violated constraint is a blocker.
233
+ 3. **Behavioral correctness**: trace 3+ representative scenarios and 5+ edge cases (empty, boundary, malformed, concurrent, failure paths) through the code. Logic errors, off-by-one, null handling, races, leaks, unhandled rejections.
234
+ 4. **Code quality**: pattern consistency with the neighboring files, naming, error handling (no swallowed errors), type safety (no `as any` or suppressions), performance on hot paths, abstraction level, tests that are meaningful rather than tautological or implementation-mirroring, API and breaking-change hygiene.
235
+ 5. **Security**: input validation and injection vectors (SQL, XSS, command, path traversal, SSRF), authentication and authorization, secrets in code or logs, data exposure, new dependencies and lockfile consistency, unsafe file and network handling.
236
+ 6. **Missed context**: does the change respect the reasons recorded in history, issues, PR reviews, and docs? Does anything that imports or documents the changed behavior need to move with it?
237
+ 7. **QA evidence audit**: for every matrix row, does the artifact exist and prove what the row claims? Name any success criterion that has no covering row.
238
+ 8. **Scope**: anything added that was not asked for - unnecessary abstraction, speculative generality, unrequested hardening - is a NOTE unless it breaks a constraint.
206
239
 
207
- 6. **Behavioral Correctness**: Walk through the code logic for 3+ representative scenarios. Does the code actually produce the expected behavior in each case?
240
+ APPROVE unless you can cite a specific goal item, constraint, or QA row that the artifact fails, with the evidence that proves it (including an artifact a criterion requires but that is missing). A gap you cannot tie to the stated goal - style preference, alternative design, a scenario the goal never named - is a NOTE, not a blocker. You do not judge approach optimality or hypothetical future requirements.
208
241
 
209
242
  OUTPUT FORMAT:
210
- <verdict>PASS or FAIL</verdict>
243
+ <verdict>APPROVE or REJECT</verdict>
211
244
  <confidence>HIGH / MEDIUM / LOW</confidence>
212
- <summary>1-3 sentence overall assessment</summary>
245
+ <summary>1-3 sentence overall assessment from the user's perspective</summary>
213
246
  <goal_breakdown>
214
- For each sub-requirement:
215
- - [ACHIEVED/MISSED/PARTIAL] Requirement description
216
- - Evidence: specific code reference or gap
247
+ - [ACHIEVED/MISSED/PARTIAL] Requirement - evidence (file:line or artifact)
217
248
  </goal_breakdown>
218
249
  <constraint_compliance>
219
- For each constraint:
220
- - [ACHIEVED/MISSED] Constraint description - evidence
250
+ - [ACHIEVED/MISSED] Constraint - evidence
221
251
  </constraint_compliance>
222
- <findings>
223
- - [PASS/FAIL/WARN] Category: Description
224
- - File: path (line range if applicable)
225
- - Evidence: specific code or logic reference
226
- </findings>
227
- <blocking_issues>Issues that MUST be fixed. Empty if PASS.</blocking_issues>
228
- """)
229
- ```
230
-
231
- ---
232
-
233
- ### Agent 2: QA via App Execution (unspecified-high) - MAIN
234
-
235
- This agent answers: "Does it actually work when you run it?"
236
-
237
- The QA agent follows a structured process: brainstorm scenarios exhaustively first, then self-review and augment, then create a task list, then execute systematically.
238
-
239
- ```
240
- task(
241
- category="unspecified-high",
242
- run_in_background=true,
243
- load_skills=["browser:control-in-app-browser", "playwright", "dev-browser"],
244
- description="QA by actually running and using the application",
245
- prompt="""
246
- <review_type>QA - HANDS-ON APP EXECUTION</review_type>
247
-
248
- <original_goal>
249
- {GOAL}
250
- </original_goal>
251
-
252
- <constraints>
253
- {CONSTRAINTS}
254
- </constraints>
255
-
256
- <changed_files>
257
- {CHANGED_FILES}
258
- </changed_files>
259
-
260
- <run_command>
261
- {RUN_COMMAND - how to start the application, or "unknown" if not determined}
262
- </run_command>
263
-
264
- You are a QA engineer. Your job is to RUN the application and verify it works through hands-on testing. You do not review code - you test behavior.
265
-
266
- If the orchestrator already ran the `visual-qa` dual-oracle gate on this same build, consume that verdict instead of re-running it - your lane covers hands-on behavior the visual gate does not.
267
-
268
- MANDATORY PROCESS (follow in order):
269
-
270
- ### Step 1: Scenario Brainstorm
271
-
272
- Before touching the app, write down EVERY test scenario you can think of. Be exhaustive. Think about:
273
-
274
- - **Happy paths**: The primary use cases this implementation enables. What's the main thing the user wanted to do?
275
- - **Boundary conditions**: Empty inputs, maximum-length inputs, zero values, negative numbers, special characters, unicode, very large datasets.
276
- - **Error paths**: Invalid inputs, network failures, missing files, permission denied, timeout conditions.
277
- - **Regression scenarios**: Existing features that touch the same code paths. Things that worked before and must still work.
278
- - **State transitions**: What happens when you do things out of order? Rapid repeated actions? Concurrent usage?
279
- - **UX scenarios** (if applicable): Layout on different sizes, keyboard navigation, screen reader compatibility, loading states, error messages.
280
- - **Integration points**: Does this feature interact with external services, databases, or other modules? Test those boundaries.
281
-
282
- Write each scenario as a one-liner with expected behavior. Aim for 15-30 scenarios minimum.
283
-
284
- ### Step 2: Scenario Augmentation
285
-
286
- Review your scenario list with fresh eyes. For each scenario, ask:
287
- - "What could go wrong here that I haven't considered?"
288
- - "What would a malicious or careless user do?"
289
- - "What environmental conditions could affect this?" (disk full, slow network, expired tokens)
290
-
291
- Add at least 5 more scenarios from this reflection. Group scenarios by priority: P0 (must pass), P1 (should pass), P2 (nice to pass).
292
-
293
- ### Step 3: Create Task List
294
-
295
- Convert your augmented scenario list into a structured task list (use TaskCreate/TaskUpdate or your todo system). Each task = one test scenario with:
296
- - Test name
297
- - Steps to execute
298
- - Expected result
299
- - Priority (P0/P1/P2)
300
-
301
- ### Step 4: Execute Systematically
302
-
303
- Work through the task list in priority order (P0 first). For each test:
304
-
305
- 1. Execute the test steps
306
- 2. Record actual result
307
- 3. Compare with expected result
308
- 4. Mark PASS or FAIL
309
- 5. If FAIL: capture evidence (screenshot, terminal output, error message)
310
- 6. Mark the task complete
311
-
312
- **Execution guidance by app type:**
313
- - **Web app**: In Codex, use `browser:control-in-app-browser` first for browser work that does not need an authenticated user session. Fall back to playwright/dev-browser when the Browser plugin is unavailable, lacks the needed action, or the test specifically needs a persistent/authenticated browser profile. Navigate, click, fill forms, and verify visual output through the chosen browser surface.
314
- - **CLI tool**: Run commands with various arguments, pipe inputs, check exit codes and output.
315
- - **Library/SDK**: Write and execute a test script that imports and exercises the public API.
316
- - **Backend API**: Use curl/httpie to hit endpoints with various payloads, verify response codes and bodies.
317
- - **Mobile/Desktop**: If not directly runnable, write integration tests and execute them.
318
-
319
- If the app cannot be started (build failure), that's an immediate FAIL - no need to continue.
320
-
321
- ### Step 5: Compile Results
322
-
323
- OUTPUT FORMAT:
324
- <verdict>PASS or FAIL</verdict>
325
- <confidence>HIGH / MEDIUM / LOW</confidence>
326
- <summary>1-3 sentence overall assessment</summary>
327
- <scenario_coverage>
328
- Total scenarios: N
329
- P0: X tested, Y passed
330
- P1: X tested, Y passed
331
- P2: X tested, Y passed
332
- </scenario_coverage>
333
- <test_results>
334
- For each test:
335
- - [PASS/FAIL] Test name (Priority)
336
- - Steps: What you did
337
- - Expected: What should happen
338
- - Actual: What actually happened
339
- - Evidence: Screenshot path or terminal output snippet (if FAIL)
340
- </test_results>
341
- <blocking_issues>P0 or P1 failures only. Empty if PASS.</blocking_issues>
252
+ <qa_audit>
253
+ - [VERIFIED/UNPROVEN] Row # - what the artifact shows, or what is missing
254
+ </qa_audit>
255
+ <blocking_issues>
256
+ - Violated goal item / constraint / QA row -> observation -> file:line or artifact pointer -> exact fix. Empty if APPROVE.
257
+ </blocking_issues>
258
+ <notes>Non-blocking findings, grouped by theme (quality, security, context). Keep them short.</notes>
342
259
  """)
343
260
  ```
344
261
 
345
262
  ---
346
263
 
347
- ### Agent 3: Code Quality Review (Oracle) - MAIN
348
-
349
- This agent answers: "Is the code well-written, maintainable, and consistent with the codebase?"
350
-
351
- ```
352
- task(
353
- subagent_type="oracle",
354
- run_in_background=true,
355
- load_skills=[],
356
- description="Review overall code quality, patterns, and architecture",
357
- prompt="""
358
- <review_type>CODE QUALITY REVIEW</review_type>
359
-
360
- <changed_files>
361
- {CHANGED_FILES}
362
- </changed_files>
363
-
364
- <file_contents>
365
- {FILE_CONTENTS - full content of changed files AND neighboring files that show existing patterns}
366
- </file_contents>
367
-
368
- <diff>
369
- {DIFF}
370
- </diff>
371
-
372
- <background>
373
- {BACKGROUND}
374
- </background>
375
-
376
- You are a senior staff engineer conducting a code review. Your standard: "Would I approve this PR without comments?"
377
-
378
- REVIEW DIMENSIONS (examine each):
264
+ ## Phase 3: Wait & Collect
379
265
 
380
- 1. **Correctness**: Logic errors, off-by-one, null/undefined handling, race conditions, resource leaks, unhandled promise rejections.
266
+ Wait for the reviewer in bounded cycles while doing independent root work. Do not treat a timeout, an ack-only reply, or an empty result as APPROVE.
381
267
 
382
- 2. **Pattern Consistency**: Does new code follow the codebase's established patterns? Compare with the neighboring files provided. Introducing a new pattern where one already exists = finding.
268
+ Store the lane verdicts independently:
383
269
 
384
- 3. **Naming & Readability**: Clear variable/function/type names? Self-documenting code? Would another engineer understand this without explanation?
270
+ | Lane | Verdict | Notes |
271
+ |------|---------|-------|
272
+ | Manual QA (Phase 1) | PASS/FAIL | rows, artifact paths |
273
+ | Gate review | pending/APPROVE/REJECT/INCONCLUSIVE | - |
385
274
 
386
- 4. **Error Handling**: Errors properly caught, logged, and propagated? No empty catch blocks? No swallowed errors? User-facing errors are helpful?
275
+ If the reviewer stays silent after the reliability followup, record the lane INCONCLUSIVE and respawn one smaller reviewer scoped to the same inputs. If that retry also ends without a verdict, close the still-running agent if safe, keep the lane INCONCLUSIVE, and emit the final result naming the incomplete lane. Do not spin in repeated wait/followup cycles, and do not use a queued followup as a cancellation.
387
276
 
388
- 5. **Type Safety**: Any `as any`, `@ts-ignore`, `@ts-expect-error`? Proper generic usage? Correct type narrowing? (If TypeScript/typed language)
389
-
390
- 6. **Performance**: N+1 queries? Unnecessary re-renders? Blocking I/O on hot paths? Memory leaks? Unbounded growth?
391
-
392
- 7. **Abstraction Level**: Right level of abstraction? No copy-paste duplication? But also no premature over-abstraction?
393
-
394
- 8. **Testing**: New behaviors covered by tests? Tests are meaningful, not just coverage padding? Test names describe scenarios?
395
-
396
- 9. **API Design**: Public interfaces clean and consistent with existing APIs? Breaking changes flagged?
397
-
398
- 10. **Tech Debt**: Does this introduce new tech debt? Or create coupling that will be painful to change?
399
-
400
- Categorize each finding by severity:
401
- - **CRITICAL**: Will cause bugs, data loss, or crashes in production
402
- - **MAJOR**: Significant quality issue that should be fixed before merge
403
- - **MINOR**: Improvement worth making but not blocking
404
- - **NITPICK**: Style preference, optional
405
-
406
- OUTPUT FORMAT:
407
- <verdict>PASS or FAIL</verdict>
408
- <confidence>HIGH / MEDIUM / LOW</confidence>
409
- <summary>1-3 sentence overall assessment</summary>
410
- <findings>
411
- - [CRITICAL/MAJOR/MINOR/NITPICK] Category: Description
412
- - File: path (line range)
413
- - Current: what the code does now
414
- - Suggestion: how to improve
415
- </findings>
416
- <blocking_issues>CRITICAL and MAJOR items only. Empty if PASS.</blocking_issues>
417
- """)
418
- ```
419
-
420
- ---
421
-
422
- ### Agent 4: Security Review (Oracle) - SUB
423
-
424
- This agent answers: "Are there security vulnerabilities in these changes?"
425
-
426
- This is supplementary - it focuses exclusively on security. It does NOT comment on code style, architecture, or functionality unless those directly create a security risk.
427
-
428
- ```
429
- task(
430
- subagent_type="oracle",
431
- run_in_background=true,
432
- load_skills=[],
433
- description="Security-focused review of implementation changes",
434
- prompt="""
435
- <review_type>SECURITY REVIEW (supplementary)</review_type>
436
-
437
- <changed_files>
438
- {CHANGED_FILES}
439
- </changed_files>
440
-
441
- <file_contents>
442
- {FILE_CONTENTS - full content of changed files}
443
- </file_contents>
444
-
445
- <diff>
446
- {DIFF}
447
- </diff>
448
-
449
- You are a security engineer. Review this diff exclusively for security vulnerabilities and anti-patterns. Ignore code style, naming, architecture - unless it directly creates a security risk.
450
-
451
- SECURITY CHECKLIST:
452
-
453
- 1. **Input Validation**: User inputs sanitized? SQL injection, XSS, command injection, SSRF vectors?
454
- 2. **Auth & AuthZ**: Authentication checks where needed? Authorization verified for each action? Privilege escalation paths?
455
- 3. **Secrets & Credentials**: Hardcoded secrets, API keys, tokens in code or config? Secrets in logs?
456
- 4. **Data Exposure**: Sensitive data in logs? PII in error messages? Over-exposed API responses?
457
- 5. **Dependencies**: New dependencies added? Known CVEs? Suspicious or unnecessary packages?
458
- 6. **Cryptography**: Proper algorithms? No custom crypto? Secure random? Proper key management?
459
- 7. **File & Path**: Path traversal? Unsafe file operations? Symlink following?
460
- 8. **Network**: CORS configured correctly? Rate limiting? TLS enforced? Certificate validation?
461
- 9. **Error Leakage**: Stack traces exposed to users? Internal details in error responses?
462
- 10. **Supply Chain**: Lockfile updated consistently? Dependency pinning?
463
-
464
- OUTPUT FORMAT:
465
- <verdict>PASS or FAIL</verdict>
466
- <severity>CRITICAL / HIGH / MEDIUM / LOW / NONE</severity>
467
- <summary>1-3 sentence overall assessment</summary>
468
- <findings>
469
- - [CRITICAL/HIGH/MEDIUM/LOW] Category: Description
470
- - File: path (line range)
471
- - Risk: What could an attacker do?
472
- - Remediation: Specific fix
473
- </findings>
474
- <blocking_issues>CRITICAL and HIGH items only. Empty if PASS.</blocking_issues>
475
- """)
476
- ```
477
-
478
- ---
479
-
480
- ### Agent 5: Context Mining (unspecified-high) - MAIN
481
-
482
- This agent answers: "Did we miss any context that should have informed this implementation?"
483
-
484
- ```
485
- task(
486
- category="unspecified-high",
487
- run_in_background=true,
488
- load_skills=["git-master"],
489
- description="Mine all accessible contexts for missed requirements or background knowledge",
490
- prompt="""
491
- <review_type>CONTEXT MINING - MISSED REQUIREMENTS & BACKGROUND</review_type>
492
-
493
- <original_goal>
494
- {GOAL}
495
- </original_goal>
496
-
497
- <constraints>
498
- {CONSTRAINTS}
499
- </constraints>
500
-
501
- <changed_files>
502
- {CHANGED_FILES}
503
- </changed_files>
504
-
505
- <background>
506
- {BACKGROUND}
507
- </background>
508
-
509
- You are an investigator. Your mission: search every accessible information source to find context that should have informed this implementation but might have been missed. The question: "Is there something we should have known but didn't?"
510
-
511
- SOURCES TO SEARCH (use every available tool):
512
-
513
- 1. **Git History** (ALWAYS search):
514
- - `git log --oneline -20 -- {each changed file}` - recent changes and their reasons
515
- - `git blame {critical sections}` - who wrote what and when
516
- - `git log --all --grep="{keywords from goal}"` - related commits
517
- - Look for reverted commits, TODO/FIXME/HACK comments in history
518
-
519
- 2. **GitHub** (if `gh` CLI available):
520
- - `gh issue list --search "{keywords}"` - related open/closed issues
521
- - `gh pr list --search "{keywords}" --state all` - related PRs and their review comments
522
- - Check if any issue is specifically linked to this work
523
- - Look at review comments on past PRs touching these files
524
-
525
- 3. **Communication Channels** (if MCP tools available):
526
- - Slack: search for messages mentioning the feature, file names, or related keywords
527
- - Notion: search for design docs, RFCs, ADRs related to this feature
528
- - Discord: relevant discussions
529
-
530
- 4. **Codebase Cross-References** (ALWAYS search):
531
- - Files that import or reference the changed modules
532
- - Tests that might need updating due to behavior changes
533
- - Documentation (README, docs/, comments) that references changed behavior
534
- - Config files that might need corresponding updates
535
- - Related features in the same domain
536
-
537
- WHAT TO LOOK FOR:
538
-
539
- - Requirements mentioned in issues/PRs that the implementation misses
540
- - Past decisions explaining WHY code was written a certain way - and whether new changes respect those reasons
541
- - Related systems or features affected by these changes
542
- - Warnings from previous developers (PR review comments, inline TODOs, commit messages)
543
- - Migration or deprecation notes that affect the changed code
544
- - Design decisions documented outside the codebase (Notion, Slack, ADRs)
545
-
546
- OUTPUT FORMAT:
547
- <verdict>PASS or FAIL</verdict>
548
- <confidence>HIGH / MEDIUM / LOW</confidence>
549
- <summary>1-3 sentence overall assessment</summary>
550
- <sources_searched>
551
- - [SEARCHED/SKIPPED] Source name - what was searched (or why it wasn't accessible)
552
- </sources_searched>
553
- <discovered_context>
554
- For each discovery:
555
- - Source: Where found (git commit abc123, GitHub issue #42, Slack message, etc.)
556
- - Finding: What was found
557
- - Relevance: How it relates to the current work
558
- - Impact: [BLOCKING / IMPORTANT / FYI]
559
- </discovered_context>
560
- <missed_requirements>Requirements the implementation should address but doesn't. Empty if none.</missed_requirements>
561
- <blocking_issues>BLOCKING items only. Empty if PASS.</blocking_issues>
562
- """)
563
- ```
277
+ After the lane reaches a terminal state and before delivering the verdict, tear down the review worktree: run `git worktree unlock <path>` followed by `git worktree remove <path>`. The reviewer runs inside that worktree, so removing it earlier destroys its working directory; a crashed review leaves the locked tree as a recoverable marker for manual cleanup.
564
278
 
565
- ## Phase 2: Wait & Collect
566
-
567
- After launching all 5 agents in one turn, wait for completions in bounded
568
- cycles. Do not treat a timeout, ack-only reply, or empty child result as
569
- a PASS.
570
-
571
- As each completes, collect via the Codex mapping above (`multi_agent_v1.wait_agent`,
572
- then the child's substantive final result). Preserve completed lane
573
- results immediately; never lose a PASS/FAIL because another lane is
574
- still running. Store each verdict independently:
575
-
576
- | Agent | Verdict | Notes |
577
- |-------|---------|-------|
578
- | 1. Goal Verification | pending/PASS/FAIL/INCONCLUSIVE | - |
579
- | 2. QA Execution | pending/PASS/FAIL/INCONCLUSIVE | - |
580
- | 3. Code Quality | pending/PASS/FAIL/INCONCLUSIVE | - |
581
- | 4. Security | pending/PASS/FAIL/INCONCLUSIVE | - |
582
- | 5. Context Mining | pending/PASS/FAIL/INCONCLUSIVE | - |
583
-
584
- Do NOT deliver the final report until ALL 5 lanes have a terminal state:
585
- PASS, FAIL, or INCONCLUSIVE.
586
- If a lane remains silent after the reliability followup, record it as
587
- inconclusive and respawn a smaller reviewer/worker for that exact lane.
588
- If it still remains unfinished after that retry, close the still-running
589
- agent if safe, keep the lane INCONCLUSIVE, and emit the final aggregate
590
- review result with the incomplete lane named. Do not spin in repeated
591
- wait/followup cycles. Do not use `multi_agent_v1.send_input` as an interrupt; queued
592
- followups are not cancellation.
593
-
594
- After ALL 5 lanes reach a terminal state and before delivering the verdict, tear down the review worktree: run `git worktree unlock <path>` followed by `git worktree remove <path>`. The lanes above run inside that worktree, so removing it earlier destroys their working directory; a crashed review leaves the locked tree as a recoverable marker for manual cleanup.
279
+ A re-review after fixes is a fresh Phase 0 -> 3 pass scoped to the delta plus the current evidence: re-run the affected QA rows and spawn a NEW reviewer with the delta diff and the blockers it must re-check. Never send fixes as a followup to the previous reviewer.
595
280
 
596
281
  ---
597
282
 
598
- ## Phase 3: Deliver Verdict
283
+ ## Phase 4: Deliver Verdict
599
284
 
600
285
  <verdict_logic>
601
286
 
602
- ALL 5 agents returned PASS → **REVIEW PASSED**
603
- ANY agent returned FAIL → **REVIEW FAILED - criteria not met**
604
- ANY lane is INCONCLUSIVE and none failed → **REVIEW INCONCLUSIVE - not approved**
287
+ QA matrix has no FAIL row AND the reviewer returned APPROVE → **REVIEW PASSED**
288
+ Any QA row FAIL OR the reviewer returned REJECT → **REVIEW FAILED - criteria not met**
289
+ Reviewer INCONCLUSIVE and nothing failed → **REVIEW INCONCLUSIVE - not approved**
605
290
 
606
291
  </verdict_logic>
607
292
 
@@ -612,25 +297,22 @@ Compile the final report in this format:
612
297
 
613
298
  ## Overall Verdict: PASSED / FAILED / INCONCLUSIVE
614
299
 
615
- | # | Review Area | Agent Type | Verdict | Confidence |
616
- |---|------------|------------|---------|------------|
617
- | 1 | Goal & Constraint Verification | Oracle | PASS/FAIL/INCONCLUSIVE | HIGH/MED/LOW |
618
- | 2 | QA Execution | unspecified-high | PASS/FAIL/INCONCLUSIVE | HIGH/MED/LOW |
619
- | 3 | Code Quality | Oracle | PASS/FAIL/INCONCLUSIVE | HIGH/MED/LOW |
620
- | 4 | Security (supplementary) | Oracle | PASS/FAIL/INCONCLUSIVE | Severity |
621
- | 5 | Context Mining | unspecified-high | PASS/FAIL/INCONCLUSIVE | HIGH/MED/LOW |
300
+ | Lane | Verdict | Confidence |
301
+ |------|---------|------------|
302
+ | Manual QA (N rows, M artifacts) | PASS/FAIL | - |
303
+ | Gate review | APPROVE/REJECT/INCONCLUSIVE | HIGH/MED/LOW |
622
304
 
623
305
  ## Blocking Issues
624
- [Aggregated from all agents - deduplicated, prioritized]
306
+ [Failing QA rows first, then the reviewer's blockers - deduplicated, in fix order, each with its pointer]
625
307
 
626
308
  ## Key Findings
627
- [Top 5-10 most important findings across all agents, grouped by theme]
309
+ [Top findings across QA and review, grouped by theme]
628
310
 
629
311
  ## Recommendations
630
312
  [If FAILED: exactly what to fix, in priority order]
631
- [If PASSED: non-blocking suggestions worth considering]
313
+ [If PASSED: non-blocking notes worth considering]
632
314
  ```
633
315
 
634
- If FAILED - be specific. The user should know exactly what to fix and in what order. No vague "consider improving X" - state the problem, the file, and the fix.
316
+ If FAILED - be specific. The user should know exactly what to fix and in what order: the problem, the file or artifact, and the fix. No vague "consider improving X".
635
317
 
636
- If PASSED - keep it short. Highlight any non-blocking suggestions, but don't turn a passing review into a lecture.
318
+ If PASSED - keep it short. Highlight the non-blocking notes worth considering, but don't turn a passing review into a lecture.