@tea-agent/loop-agent 0.12.0 → 0.13.0-beta.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (284) hide show
  1. package/AGENTS.md +155 -153
  2. package/CHANGELOG.md +338 -265
  3. package/README.md +345 -298
  4. package/bin/agent-worker.js +22 -22
  5. package/bin/loop-agent.js +21 -21
  6. package/dist/application/dag/generate-task-dag.js +28 -28
  7. package/dist/application/evaluation/candidate-hash.js +75 -0
  8. package/dist/application/evaluation/candidate.js +52 -0
  9. package/dist/application/evaluation/replay.js +289 -0
  10. package/dist/application/evaluation/types.js +130 -0
  11. package/dist/cli/command-definitions.js +27 -7
  12. package/dist/cli/program.js +8 -4
  13. package/dist/commands/cursor-prompt.js +6 -6
  14. package/dist/commands/eval.js +235 -0
  15. package/dist/commands/init.js +544 -506
  16. package/dist/commands/knowledge.js +129 -31
  17. package/dist/commands/loop-benchmark.js +11 -11
  18. package/dist/commands/pi-reuse-benchmark.js +16 -16
  19. package/dist/executors/pi-sdk-executor.js +38 -24
  20. package/dist/executors/shell-executor.js +34 -2
  21. package/dist/executors/shell-presets.js +20 -0
  22. package/dist/executors/shell-verification.js +7 -0
  23. package/dist/governance/manifest-types.js +4 -0
  24. package/dist/infrastructure/evaluation/candidate-store.js +435 -0
  25. package/dist/infrastructure/evaluation/store.js +40 -0
  26. package/dist/sidecars/cursor-prompt/executor.js +1 -1
  27. package/dist/task/config-types.js +28 -1
  28. package/dist/task/runtime.js +27 -27
  29. package/dist/worker/cli.js +96 -1
  30. package/dist/worker/delivery/package.js +3 -3
  31. package/dist/worker/feature/decision-loader.js +37 -6
  32. package/dist/worker/feature/next-action.js +10 -2
  33. package/dist/worker/feature/ready-plan-projection.js +81 -0
  34. package/dist/worker/feature/reducer.js +2 -1
  35. package/dist/worker/feature/review.js +19 -2
  36. package/dist/worker/feature/run.js +27 -2
  37. package/dist/worker/follow-up/approve.js +5 -2
  38. package/dist/worker/follow-up/factory.js +1 -1
  39. package/dist/worker/observability/read-model.js +246 -41
  40. package/dist/worker/observe/routes.js +173 -15
  41. package/dist/worker/observe/spec-evidence.js +281 -0
  42. package/dist/worker/observe/static/api.js +46 -27
  43. package/dist/worker/observe/static/app.js +150 -150
  44. package/dist/worker/observe/static/constants.js +148 -148
  45. package/dist/worker/observe/static/copy.js +67 -67
  46. package/dist/worker/observe/static/dag-helpers.js +172 -172
  47. package/dist/worker/observe/static/dag-layout.d.ts +31 -31
  48. package/dist/worker/observe/static/dag-layout.js +83 -83
  49. package/dist/worker/observe/static/dag-model.js +72 -72
  50. package/dist/worker/observe/static/dom.js +61 -61
  51. package/dist/worker/observe/static/format-pool.js +67 -67
  52. package/dist/worker/observe/static/format.js +292 -292
  53. package/dist/worker/observe/static/index.html +308 -308
  54. package/dist/worker/observe/static/kpi.js +94 -94
  55. package/dist/worker/observe/static/relations.js +133 -128
  56. package/dist/worker/observe/static/router.js +93 -85
  57. package/dist/worker/observe/static/run-processing.js +148 -148
  58. package/dist/worker/observe/static/shell-chrome.js +68 -68
  59. package/dist/worker/observe/static/state.js +253 -253
  60. package/dist/worker/observe/static/styles.css +1902 -1890
  61. package/dist/worker/observe/static/views/batch.js +227 -226
  62. package/dist/worker/observe/static/views/dag-graph.js +172 -172
  63. package/dist/worker/observe/static/views/dag-inspector.js +607 -477
  64. package/dist/worker/observe/static/views/dag.js +362 -362
  65. package/dist/worker/observe/static/views/dashboard.js +445 -442
  66. package/dist/worker/observe/static/views/failures.js +143 -143
  67. package/dist/worker/observe/static/views/feature.js +492 -453
  68. package/dist/worker/observe/static/views/pool.js +350 -347
  69. package/dist/worker/observe/static/views/run.js +453 -453
  70. package/dist/worker/observe/static/views/session-timeline.js +205 -205
  71. package/dist/worker/observe/static/views/shell.js +7 -7
  72. package/dist/worker/observe/static/views/task.js +314 -260
  73. package/dist/worker/observe/static/views/timeline.js +163 -163
  74. package/dist/worker/pool/doctor.js +165 -0
  75. package/dist/worker/pool/migrate-state.js +303 -0
  76. package/dist/worker/pool/run-store.js +205 -17
  77. package/dist/worker/pool/types.js +17 -1
  78. package/dist/worker/pool/validation.js +100 -15
  79. package/dist/worker/report/morning-report.js +12 -2
  80. package/dist/worker/runner/run-ready.js +41 -26
  81. package/dist/worker/task-graph/ready-planner.js +136 -0
  82. package/dist/workflows/dag/backend-test-analysis-contract.js +120 -0
  83. package/dist/workflows/dag/canvas-observer.js +275 -275
  84. package/dist/workflows/dag/convergence/controller.js +16 -8
  85. package/dist/workflows/dag/dynamic-runtime/map.js +90 -2
  86. package/dist/workflows/dag/failure-routing.js +12 -1
  87. package/dist/workflows/dag/init-hybrid.js +2404 -360
  88. package/dist/workflows/dag/node-execution.js +9 -0
  89. package/dist/workflows/dag/prompt.js +9 -0
  90. package/dist/workflows/dag/report.js +35 -1
  91. package/dist/workflows/dag/runner.js +28 -2
  92. package/dist/workflows/dag/task-demand-routing.js +383 -0
  93. package/dist/workflows/dag/types.js +51 -13
  94. package/dist/workflows/dag/upstream-artifacts.js +1 -0
  95. package/dist/workflows/dag/validate.js +59 -1
  96. package/docs/README.md +106 -104
  97. package/docs/agent-dag-recovery-playbook.md +195 -184
  98. package/docs/agent-dag-runner.md +67 -67
  99. package/docs/architecture/README.md +26 -26
  100. package/docs/architecture/dag-execution.md +140 -140
  101. package/docs/architecture/evolution.md +54 -53
  102. package/docs/architecture/facts-and-state.md +71 -58
  103. package/docs/architecture/runtime-boundaries.md +191 -191
  104. package/docs/architecture/system-overview.md +93 -93
  105. package/docs/architecture/worker-and-feature.md +85 -81
  106. package/docs/cursor-prompt-sidecar.md +36 -36
  107. package/docs/decisions/README.md +18 -15
  108. package/docs/design/README.md +167 -77
  109. package/docs/development-principles.md +73 -73
  110. package/docs/exec-plans/README.md +6 -6
  111. package/docs/exec-plans/active/README.md +15 -9
  112. package/docs/exec-plans/completed/README.md +85 -73
  113. package/docs/feature-workflow.md +389 -261
  114. package/docs/harness-methodology-debugging.md +153 -153
  115. package/docs/harness-methodology-tdd.md +130 -130
  116. package/docs/harness-methodology-verification.md +27 -27
  117. package/docs/init-surface.manifest.json +289 -280
  118. package/docs/loop-agent-harness.md +142 -130
  119. package/docs/production-readiness.md +96 -96
  120. package/docs/progress/README.md +64 -54
  121. package/docs/reports/README.md +117 -94
  122. package/docs/skills/README.md +7 -7
  123. package/docs/skills/vetted-skill-registry.md +29 -27
  124. package/docs/templates/adr.md +60 -60
  125. package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
  126. package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
  127. package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
  128. package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
  129. package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
  130. package/docs/templates/agent-dag-report.schema.json +473 -473
  131. package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
  132. package/docs/templates/agent-dag.base.json +190 -190
  133. package/docs/templates/agent-dag.final-verification.json +185 -185
  134. package/docs/templates/agent-dag.schema.json +411 -383
  135. package/docs/templates/agent-dag.supervised-implementation.json +501 -501
  136. package/docs/templates/backend-test-analysis.schema.json +44 -0
  137. package/docs/templates/backend-test-dag.generate-pytest.prompt.md +202 -139
  138. package/docs/templates/backend-test-dag.json +311 -276
  139. package/docs/templates/backend-test-dag.retrospect.prompt.md +125 -125
  140. package/docs/templates/backend-test-dag.review-cases.prompt.md +81 -81
  141. package/docs/templates/exec-plan.md +64 -64
  142. package/docs/templates/feature-spec.md +53 -53
  143. package/docs/templates/frontend-design-contract.md +42 -33
  144. package/docs/templates/frontend-task-constraints.md +35 -25
  145. package/docs/templates/frontend-task-requirement.md +70 -61
  146. package/docs/templates/frontend-test-dag.generate-cases.prompt.md +5 -0
  147. package/docs/templates/frontend-test-dag.json +23 -0
  148. package/docs/templates/frontend-test-dag.retrieve-context.prompt.md +3 -0
  149. package/docs/templates/frontend-test-dag.retrospect.prompt.md +3 -0
  150. package/docs/templates/frontend-test-dag.review-cases.prompt.md +3 -0
  151. package/docs/templates/frontend-test-dag.review-execution.prompt.md +3 -0
  152. package/docs/templates/harness.schema.json +221 -221
  153. package/docs/templates/hybrid-dag.json +188 -188
  154. package/docs/templates/init-evolution-review.md +35 -35
  155. package/docs/templates/interactive-ui-round2-experiment.md +66 -66
  156. package/docs/templates/knowledge-graph-bootstrap-dag.json +118 -0
  157. package/docs/templates/knowledge-sync-dag.json +178 -0
  158. package/docs/templates/knowledge-sync-draft.schema.json +71 -0
  159. package/docs/templates/product-line/AGENTS.md +8 -8
  160. package/docs/templates/product-line/README.md +9 -9
  161. package/docs/templates/product-line/acceptance.yaml +14 -14
  162. package/docs/templates/product-line/closeout.yaml +9 -9
  163. package/docs/templates/product-line/design.md +13 -13
  164. package/docs/templates/product-line/links.md +10 -10
  165. package/docs/templates/product-line/requirement.md +17 -17
  166. package/docs/templates/product-line/task-graph.yaml +15 -15
  167. package/docs/templates/product-line/task.yaml +64 -64
  168. package/docs/templates/product-line/test-plan.md +7 -7
  169. package/docs/templates/production-readiness-checklist.md +57 -57
  170. package/docs/templates/progress-log.md +17 -17
  171. package/docs/templates/project-start-checklist.md +9 -9
  172. package/docs/templates/qa-report.md +48 -48
  173. package/docs/templates/sprint-contract.md +29 -29
  174. package/docs/templates/worker-dogfood-evidence.md +80 -80
  175. package/docs/templates/worker-dogfood-setup.md +68 -68
  176. package/docs/verification-matrix.md +70 -66
  177. package/examples/decision-gate-agent-dag.json +177 -177
  178. package/examples/example-dag.json +46 -46
  179. package/examples/hybrid-loop-agent-dag.json +189 -189
  180. package/harness.json +66 -66
  181. package/package.json +88 -46
  182. package/scripts/check-product-line-docs.sh +29 -29
  183. package/scripts/check-task-pool-root.sh +32 -32
  184. package/scripts/kb-bootstrap-init-skeleton.sh +240 -0
  185. package/scripts/kb-graph-incremental-prepare.mjs +386 -0
  186. package/scripts/kb-graph-incremental-prepare.sh +5 -0
  187. package/scripts/kb-graph-materialize.mjs +105 -0
  188. package/scripts/kb-graph-materialize.sh +4 -0
  189. package/scripts/kb-graph-promote.mjs +164 -0
  190. package/scripts/kb-graph-promote.sh +4 -0
  191. package/scripts/kb-query.mjs +554 -0
  192. package/scripts/kb-query.sh +5 -0
  193. package/skills/agent-worker/SKILL.md +39 -37
  194. package/skills/agent-worker/references/agent-worker-operator.md +60 -43
  195. package/skills/ai-engineering-context/SKILL.md +48 -48
  196. package/skills/analyze-product-dependencies/SKILL.md +67 -0
  197. package/skills/analyze-product-dependencies/agents/openai.yaml +4 -0
  198. package/skills/analyze-product-dependencies/references/api-documentation-schema.md +30 -0
  199. package/skills/analyze-product-dependencies/references/dependency-analysis-schema.md +28 -0
  200. package/skills/analyze-product-dependencies/references/example.md +76 -0
  201. package/skills/analyze-product-dependencies/references/forward-test-cases.md +35 -0
  202. package/skills/analyze-product-dependencies/references/input-contract.md +11 -0
  203. package/skills/analyze-product-dependencies/references/scouting-rules.md +61 -0
  204. package/skills/analyze-product-dependencies/scripts/test-validators.mjs +267 -0
  205. package/skills/analyze-product-dependencies/scripts/validate-api-documentation.mjs +101 -0
  206. package/skills/analyze-product-dependencies/scripts/validate-dependency-analysis.mjs +142 -0
  207. package/skills/analyze-product-dependencies/scripts/validate-product-requirement-input.mjs +76 -0
  208. package/skills/analyze-product-dependencies/scripts/validation-helpers.mjs +146 -0
  209. package/skills/analyze-product-requirements/SKILL.md +90 -0
  210. package/skills/analyze-product-requirements/agents/openai.yaml +4 -0
  211. package/skills/analyze-product-requirements/references/acceptance-criteria.md +91 -0
  212. package/skills/analyze-product-requirements/references/clarification-and-knowledge.md +56 -0
  213. package/skills/analyze-product-requirements/references/example.md +86 -0
  214. package/skills/analyze-product-requirements/references/forward-test-cases.md +66 -0
  215. package/skills/analyze-product-requirements/references/product-analysis-schema.md +32 -0
  216. package/skills/analyze-product-requirements/references/product-requirement-schema.md +33 -0
  217. package/skills/analyze-product-requirements/references/requirement-clarification-schema.md +35 -0
  218. package/skills/analyze-product-requirements/scripts/test-validators.mjs +193 -0
  219. package/skills/analyze-product-requirements/scripts/validate-product-analysis.mjs +69 -0
  220. package/skills/analyze-product-requirements/scripts/validate-product-requirement.mjs +97 -0
  221. package/skills/analyze-product-requirements/scripts/validate-requirement-clarification.mjs +98 -0
  222. package/skills/analyze-product-requirements/scripts/validation-helpers.mjs +156 -0
  223. package/skills/code-review-core/SKILL.md +20 -20
  224. package/skills/codebase-scout/SKILL.md +19 -19
  225. package/skills/frontend-design-review/SKILL.md +66 -59
  226. package/skills/frontend-design-review/references/review-checklist.md +58 -37
  227. package/skills/frontend-implementation/SKILL.md +47 -51
  228. package/skills/frontend-implementation/references/code-standards.md +32 -34
  229. package/skills/frontend-implementation/references/design-spec.md +46 -46
  230. package/skills/frontend-implementation/references/node-contracts.md +76 -32
  231. package/skills/frontend-review/SKILL.md +59 -53
  232. package/skills/frontend-review/references/review-findings.md +47 -42
  233. package/skills/frontend-verification/SKILL.md +53 -40
  234. package/skills/frontend-verification/references/verification-checklist.md +68 -56
  235. package/skills/grill-me/SKILL.md +10 -10
  236. package/skills/grill-with-docs/SKILL.md +88 -88
  237. package/skills/grill-with-docs/adr-format.md +47 -47
  238. package/skills/grill-with-docs/context-format.md +60 -60
  239. package/skills/init-capability-evolution/SKILL.md +70 -70
  240. package/skills/loop-agent/SKILL.md +151 -151
  241. package/skills/loop-agent/references/README.md +67 -67
  242. package/skills/loop-agent/references/command-reference.md +505 -452
  243. package/skills/loop-agent/references/docs-converge.md +126 -126
  244. package/skills/loop-agent/references/harness-policy.md +263 -263
  245. package/skills/loop-agent/references/hybrid-dag.md +238 -233
  246. package/skills/loop-agent/references/learned/README.md +21 -21
  247. package/skills/loop-agent/references/long-running-loop.md +57 -57
  248. package/skills/loop-agent/references/model-routing.md +36 -36
  249. package/skills/loop-agent/references/multi-worktree.md +54 -54
  250. package/skills/loop-agent/references/one-shot-runs.md +85 -85
  251. package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
  252. package/skills/loop-agent/references/pi-prompt.md +23 -23
  253. package/skills/loop-agent/references/pi-subagent-assisted-mode.md +84 -84
  254. package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
  255. package/skills/loop-agent/references/task-workflow.md +89 -89
  256. package/skills/loop-agent/references/verification-and-failure-handling.md +139 -139
  257. package/skills/playwright-cli/SKILL.md +420 -0
  258. package/skills/playwright-cli/references/element-attributes.md +23 -0
  259. package/skills/playwright-cli/references/playwright-tests.md +39 -0
  260. package/skills/playwright-cli/references/request-mocking.md +87 -0
  261. package/skills/playwright-cli/references/running-code.md +241 -0
  262. package/skills/playwright-cli/references/session-management.md +225 -0
  263. package/skills/playwright-cli/references/storage-state.md +275 -0
  264. package/skills/playwright-cli/references/test-generation.md +433 -0
  265. package/skills/playwright-cli/references/tracing.md +139 -0
  266. package/skills/playwright-cli/references/video-recording.md +143 -0
  267. package/skills/playwright-cli-case-generator/SKILL.md +74 -0
  268. package/skills/requesting-code-review/SKILL.md +101 -101
  269. package/skills/requesting-code-review/code-reviewer.md +168 -168
  270. package/skills/systematic-debugging/CREATION-LOG.md +119 -119
  271. package/skills/systematic-debugging/SKILL.md +296 -296
  272. package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
  273. package/skills/systematic-debugging/condition-based-waiting.md +115 -115
  274. package/skills/systematic-debugging/defense-in-depth.md +122 -122
  275. package/skills/systematic-debugging/find-polluter.sh +63 -63
  276. package/skills/systematic-debugging/root-cause-tracing.md +169 -169
  277. package/skills/systematic-debugging/test-academic.md +14 -14
  278. package/skills/systematic-debugging/test-pressure-1.md +58 -58
  279. package/skills/systematic-debugging/test-pressure-2.md +68 -68
  280. package/skills/systematic-debugging/test-pressure-3.md +69 -69
  281. package/skills/test-driven-development/SKILL.md +20 -20
  282. package/skills/using-git-worktrees/SKILL.md +215 -215
  283. package/skills/verification-before-completion/SKILL.md +154 -154
  284. package/skills/webapp-testing/SKILL.md +19 -19
@@ -1,32 +1,76 @@
1
- # Frontend Node Contracts
2
-
3
- ## `frontend-contract-pi`
4
-
5
- - Read task source, constraints, `task.json`, and explicit references; do not edit.
6
- - Define scope, non-goals, routes/components, runtime, user flows, acceptance criteria, states, risks, and verification expectations.
7
- - Preserve requirement IDs and language; never turn an implementation guess into a requirement.
8
- - For every standard UI state, specify behavior or mark it not applicable with a reason.
9
- - Output: `Scope`, `Non-goals`, `Acceptance Criteria`, `UI States`, `Target Runtime Environment`, `Risks`, `Verification Expectations`.
10
-
11
- ## `frontend-scout-pi`
12
-
13
- - Inspect routes, pages, components, styles/tokens, state/data flow, API/mocks, scripts, tests, and reusable assets; do not edit.
14
- - Separate confirmed facts, inferred conventions, and missing information.
15
- - Attempt knowledge-base first. If absent, failed, or unmatched, recursively inspect `<repoRoot>/openSpec/**` before other repo conventions.
16
- - Record source as `knowledge-base`, `openSpec fallback`, `repository fallback`, or `unavailable`, with query/search terms and matched paths.
17
- - Output: `Frontend Stack`, `Routes`, `Components`, `Styling System`, `Existing Design Conventions`, `State / Data Flow`, `Test Entry Points`, `Reuse Opportunities`, `Risks`.
18
-
19
- ## `frontend-plan-pi`
20
-
21
- - Map every acceptance criterion to ordered implementation and verification steps.
22
- - Name target files and reasons; keep them within allowed paths and expected `writeSet`.
23
- - Define states, component/styling reuse, interaction behavior, dependency policy, and exact static/behavior commands.
24
- - Knowledge-base failure is not approval: apply relevant `openSpec/` matches as current-project rules. Request clarification only when neither source resolves required compliance or they conflict materially.
25
- - Output: `Implementation Steps`, `Target Files`, `UI State Handling`, `Styling / Component Strategy`, `Interaction Notes`, `Dependency Policy`, `Verification Plan`, `Residual Risks`.
26
-
27
- ## `frontend-implement-pi`
28
-
29
- - Run only after design gate pass and implement only the approved plan inside `writeSet`.
30
- - Re-read current files; stop instead of crossing forbidden paths or guessing a blocking decision.
31
- - Reuse confirmed project primitives and update tests. Do not claim knowledge-base or downstream verification without evidence.
32
- - Output: `Changed Files`, `Implemented Behavior`, `UI States Covered`, `Styling / Component Notes`, `Verification Attempted`, `Residual Risks`.
1
+ # Frontend Node Contracts
2
+
3
+ All pre-write nodes are read-only. Preserve IDs, source labels, commands, language,
4
+ and required headings.
5
+
6
+ ## `frontend-contract-pi`
7
+
8
+ Read all inputs and define `Scope`, `Non-goals`, `Acceptance Criteria`, `UI States`,
9
+ `Target Runtime Environment`, `Risks`, and `Verification Expectations`. Never turn a
10
+ guess into a requirement.
11
+
12
+ ## `frontend-scout-pi`
13
+
14
+ Inspect routes, components, styles/tokens, data/API/Mock seams, scripts, tests, and
15
+ assets. Separate fact, inference, and gap. Query the knowledge base; otherwise search
16
+ and read `<repoRoot>/openSpec/**` before repository fallback. Output `Frontend Stack`,
17
+ `Routes`, `Components`, `Styling System`, `Existing Design Conventions`, `State / Data
18
+ Flow`, `Test Entry Points`, `Reuse Opportunities`, and `Risks`.
19
+
20
+ ## `frontend-mock-assess-pi` and contract gate
21
+
22
+ Consume contract, scout, generation-time capability seed, and frozen verification
23
+ entrypoints. First non-empty line:
24
+
25
+ `MOCK_STRATEGY: native|browser-intercept|request-adapter|not-needed|blocked`
26
+
27
+ Prefer a proven native service. Use browser interception only with an existing
28
+ browser/e2e harness, or a request adapter/DI seam for reversible local preview.
29
+ `not-needed` requires positive no-remote/stable-real-backend evidence and applicable
30
+ behavior verification; it is invalid when `frontendMock.policy=required` resolved to
31
+ a required decision. Select `blocked` for missing/conflicting contracts, unsafe or
32
+ unauthorized paths/dependencies, unread/conflicting specs, production-default-on
33
+ behavior, or frozen entrypoints that cannot verify the selected mechanism.
34
+
35
+ Output `Mock Decision`, `API Contract Evidence`, `Specification Evidence`, `Service
36
+ Evidence`, `Backend Readiness`, `Selection Evidence`, `Endpoint / Fixture Matrix`,
37
+ `Activation`, `Target Files`, `Production Safety`, `Verification Plan`, `Real
38
+ Integration Gap`, and `Blocking Issues`. Trace methods, paths, fields, statuses, UI
39
+ states, fixtures, and consumers to API/schema evidence. Never invent fields, store
40
+ secrets/real user data, comment real requests, import test mocks from production, or
41
+ claim Mock evidence is real integration.
42
+
43
+ `frontend-mock-contract-gate-shell` accepts only allowed non-blocked strategy lines
44
+ using `first-non-empty`; it never authorizes writes. A generation-time unsafe or
45
+ incomplete explicitly-required contract produces contract/scout/assessment/gate only,
46
+ with no writer.
47
+
48
+ ## `frontend-plan-pi` and design loop
49
+
50
+ Plan directly consumes contract, scout, assessment, and gate. Map criteria to
51
+ ordered steps, in-bound files, UI states, interactions, reuse, dependencies,
52
+ activation/rollback, fixed verification entrypoints, and the real-integration gap.
53
+ Output `Implementation Steps`, `Target Files`, `UI State Handling`, `Styling /
54
+ Component Strategy`, `Interaction Notes`, `Mock / API Strategy`, `Dependency Policy`,
55
+ `Verification Plan`, `Real Integration Gap`, and `Residual Risks`.
56
+
57
+ The first design gate accepts `VERDICT: pass` or `VERDICT: request-revision` for
58
+ read-only revision. On pass, `frontend-plan-revision-pi` outputs
59
+ `PASS_NO_REVISION_NEEDED`; otherwise it returns a complete corrected plan without
60
+ manufacturing evidence. Final design review directly rechecks the original plan,
61
+ first findings, revision, assessment, and all Mock safety boundaries. Only the final
62
+ `VERDICT: pass` gate authorizes writes; failure routes to replan/rerun, not dev-fix.
63
+
64
+ ## `frontend-implement-pi`
65
+
66
+ This is the sole exclusive writer. Consume the approved plan, final review/gate, and
67
+ assessment. Implement only inside `writeSet`; keep the real request default and Mock
68
+ activation reversible, dev/test-only, and production-off. Native handler/fixture,
69
+ browser interception, or request adapter changes stay atomic with their UI consumer
70
+ and tests. Stop on forbidden paths or guesses. Output `Changed Files`, `Implemented
71
+ Behavior`, `UI States Covered`, `Styling / Component Notes`, `Verification Attempted`,
72
+ and `Residual Risks`.
73
+
74
+ When trusted Mock-specific commands were frozen at generation, a read-only
75
+ `frontend-mock-verify-shell` runs after the writer. Static and behavior verification
76
+ always run; behavior must prove page consumption, not merely handler unit tests.
@@ -1,53 +1,59 @@
1
- ---
2
- name: frontend-review
3
- description: Use to review completed frontend code and verification before closeout.
4
- references:
5
- - path: references/review-findings.md
6
- required: true
7
- ---
8
-
9
- # Frontend Review
10
-
11
- Use for `frontend-review-pi`; read the findings guide first. Required inputs are
12
- original task/reference material, derived contract/constraints, approved plan and
13
- design verdict, implementation summary, actual diff, and shell evidence. Missing
14
- actual diff or required evidence forces revision; never infer it from a summary.
15
-
16
- ## Verdict Contract
17
-
18
- The first non-empty line must be exactly `VERDICT: pass` or
19
- `VERDICT: request-revision`. Any Critical/Important finding, failed or missing
20
- required check, forbidden write, or unmet acceptance criterion forces revision.
21
-
22
- ## Review Scope
23
-
24
- - Compare original intent, derived artifacts, approved plan, actual diff, and evidence; report lost or altered requirements.
25
- - Inspect every changed file against allowed, forbidden, and approved write scope.
26
- - Map criteria to behavior, applicable UI states, tests, and shell evidence.
27
- - Review state/data flow, validation, async/error behavior, components/design, responsive behavior, accessibility, dependencies, maintenance, and regression risk when applicable.
28
- - Component/design claims require traceable knowledge-base evidence or, after connection/query failure or no match, relevant `<repoRoot>/openSpec/**` evidence. The connector format is TODO; never claim a query or fallback search without evidence.
29
- - Treat shell exit status as authoritative. Do not edit files.
30
-
31
- ## Evidence And Output
32
-
33
- Findings cite a tight file location, exact command/result, or named DAG artifact.
34
- Separate confirmed defects, missing evidence, and residual risks.
35
-
36
- ```markdown
37
- VERDICT: request-revision
38
-
39
- ## Findings
40
- - [Important] `path:line` issue, impact, and required correction.
41
-
42
- ## Verification Assessment
43
- - ...
44
-
45
- ## UX Assessment
46
- - ...
47
-
48
- ## Residual Risks
49
- - ...
50
- ```
51
-
52
- A pass requires no Critical/Important findings and all required shell checks passed.
53
- Still report knowledge-source status and optional browser/manual gaps.
1
+ ---
2
+ name: frontend-review
3
+ description: Use to review completed frontend code and verification before closeout.
4
+ references:
5
+ - path: references/review-findings.md
6
+ required: true
7
+ maxChars: 2800
8
+ ---
9
+
10
+ # Frontend Review
11
+
12
+ Use for `frontend-review-pi`; read the findings guide first. Required inputs are
13
+ original task/reference material, contract/constraints, Mock assessment, original and
14
+ revised/confirmed plan, final design verdict, implementation summary, actual diff,
15
+ and static/behavior/optional Mock shell evidence. Missing actual diff or required
16
+ evidence forces revision; never infer it from a summary.
17
+
18
+ ## Verdict Contract
19
+
20
+ The first non-empty line must be exactly `VERDICT: pass` or
21
+ `VERDICT: request-revision`. Any Critical/Important finding, failed or missing
22
+ required check, forbidden write, or unmet acceptance criterion forces revision.
23
+
24
+ ## Review Scope
25
+
26
+ - Compare intent, contract, plan, diff, and evidence; report altered requirements.
27
+ - Inspect every changed file against allowed, forbidden, and approved write scope.
28
+ - Map criteria to behavior, applicable UI states, tests, and shell evidence.
29
+ - Review state/data flow, validation, async/error behavior, components/design, responsive behavior, accessibility, dependencies, maintenance, and regression risk when applicable.
30
+ - Inspect static, behavior, and available Mock-specific artifacts directly. For Mock
31
+ strategies, compare the endpoint matrix, handler/fixture/adapter and consumer diff;
32
+ require the real request as default, contract-aligned fixtures, production isolation,
33
+ and no false real-integration claim. `not-needed` needs applicable real/no-remote evidence.
34
+ - Component/design claims require traceable knowledge-base evidence or, after connection/query failure or no match, relevant `<repoRoot>/openSpec/**` evidence. The connector format is TODO; never claim a query or fallback search without evidence. Execute explicit `grep`/`find` to locate spec files and `read` to load them before referencing their rules. Only successful `read` tool calls are observable as "已读取规范文件" in the spec-evidence inspector.
35
+ - Treat shell exit status as authoritative. Do not edit files.
36
+
37
+ ## Evidence And Output
38
+
39
+ Findings cite a tight file location, exact command/result, or named DAG artifact.
40
+ Separate confirmed defects, missing evidence, and residual risks.
41
+
42
+ ```markdown
43
+ VERDICT: request-revision
44
+
45
+ ## Findings
46
+ - [Important] `path:line` — issue, impact, and required correction.
47
+
48
+ ## Verification Assessment
49
+ - ...
50
+
51
+ ## UX Assessment
52
+ - ...
53
+
54
+ ## Residual Risks
55
+ - ...
56
+ ```
57
+
58
+ A pass requires no Critical/Important findings and all required shell checks passed.
59
+ Still report knowledge-source status and optional browser/manual gaps.
@@ -1,42 +1,47 @@
1
- # Frontend Review Findings Guide
2
-
3
- ## Severity
4
-
5
- - **Critical**: blocks primary flow, corrupts data, violates security/privacy, writes forbidden paths, or bypasses required verification.
6
- - **Important**: acceptance/state/validation gap, material convention drift, missing behavior tests, unauthorized dependency, or failed/missing required verification.
7
- - **Minor**: non-blocking maintainability, copy, layout, or cleanup issue.
8
-
9
- ## Evidence
10
-
11
- - Cite tight file locations, exact commands/results, or named DAG artifacts.
12
- - Never invent evidence; name the missing check. An implementation summary is not the actual diff.
13
- - Failed required static/behavior verification is at least Important unless proven unrelated.
14
- - A knowledge-base claim records connector/query, source ID/version, and retrieval time. If absent, failed, or unmatched, review evidence must show `<repoRoot>/openSpec/**` search terms and matched paths/headings; label `openSpec fallback`, `repository fallback`, or `unavailable` accurately.
15
-
16
- ## Review Sequence
17
-
18
- 1. Establish changed-file inventory and write boundaries.
19
- 2. Compare original requirement with derived contract/constraints.
20
- 3. Map each criterion to code, states, tests, and evidence.
21
- 4. Inspect interactions, state/data/API behavior, failure paths, and regression risk.
22
- 5. Check component/design evidence, responsive/accessibility behavior, dependencies, and maintenance fit when applicable.
23
- 6. Classify findings and derive the verdict mechanically.
24
-
25
- Skipping the required `openSpec/` search after knowledge-base failure is Important
26
- when component/design compliance affects acceptance or implementation choices.
27
-
28
- Use one issue per finding:
29
-
30
- ```text
31
- - [Critical|Important|Minor] path:line Problem; impact; required correction; evidence.
32
- ```
33
-
34
- Avoid vague advice. When no source location exists, cite the command or artifact.
35
-
36
- ## Pass Rules
37
-
38
- - No Critical or Important findings remain.
39
- - Required static and behavior nodes ran and passed.
40
- - Changed files are authorized.
41
- - Criteria and applicable states have implementation and evidence.
42
- - Optional unavailable knowledge-base, browser, visual, or manual checks remain explicit risks.
1
+ # Frontend Review Findings Guide
2
+
3
+ ## Severity
4
+
5
+ - **Critical**: blocks primary flow, corrupts data, violates security/privacy, writes forbidden paths, or bypasses required verification.
6
+ - **Important**: acceptance/state/validation gap, material convention drift, missing behavior tests, unauthorized dependency, unsafe mock activation/import, mock-contract drift, misleading real-integration claim, or failed/missing required verification.
7
+ - **Minor**: non-blocking maintainability, copy, layout, or cleanup issue.
8
+
9
+ ## Evidence
10
+
11
+ - Cite tight file locations, exact commands/results, or named DAG artifacts.
12
+ - Never invent evidence; name the missing check. An implementation summary is not the actual diff.
13
+ - Failed required static/behavior verification is at least Important unless proven unrelated.
14
+ - A knowledge-base claim records connector/query, source ID/version, and retrieval time. If absent, failed, or unmatched, review evidence must show `<repoRoot>/openSpec/**` search terms and matched paths/headings; label `openSpec fallback`, `repository fallback`, or `unavailable` accurately.
15
+
16
+ ## Review Sequence
17
+
18
+ 1. Establish changed-file inventory and write boundaries.
19
+ 2. Compare original requirement with derived contract/constraints.
20
+ 3. Map each criterion to code, states, tests, and evidence.
21
+ 4. Inspect interactions, state/data/API behavior, failure paths, and regression risk.
22
+ 5. Check mock selection, contract-to-fixture mapping, activation/default path,
23
+ handler/fixture/adapter and consumer diff, optional Mock-specific verification,
24
+ production imports, evidence scope, and the documented real-integration gap.
25
+ 6. Check component/design evidence, responsive/accessibility behavior, dependencies, and maintenance fit when applicable.
26
+ 7. Classify findings and derive the verdict mechanically.
27
+
28
+ Skipping the required `openSpec/` search after knowledge-base failure is Important
29
+ when component/design compliance affects acceptance or implementation choices.
30
+
31
+ Use one issue per finding:
32
+
33
+ ```text
34
+ - [Critical|Important|Minor] path:line Problem; impact; required correction; evidence.
35
+ ```
36
+
37
+ Avoid vague advice. When no source location exists, cite the command or artifact.
38
+
39
+ ## Pass Rules
40
+
41
+ - No Critical or Important findings remain.
42
+ - Required static and behavior nodes ran and passed.
43
+ - Changed files are authorized.
44
+ - Criteria and applicable states have implementation and evidence.
45
+ - Required Mock-backed behavior passed; any generated Mock-specific verification also
46
+ passed; Mock is not enabled by default in production.
47
+ - Optional unavailable knowledge-base, browser, visual, or manual checks remain explicit risks.
@@ -1,40 +1,53 @@
1
- ---
2
- name: frontend-verification
3
- description: Use to assess frontend evidence and produce closeout.
4
- references:
5
- - path: references/verification-checklist.md
6
- required: true
7
- ---
8
-
9
- # Frontend Verification
10
-
11
- Use for `frontend-closeout-pi`. Shell nodes execute commands; this read-only skill
12
- assesses their evidence. Read the checklist first.
13
-
14
- Required inputs: criteria, implementation summary/change inventory, static and
15
- behavior shell commands/source/status/artifacts, review verdict/findings, and any
16
- required browser, visual, manual, or knowledge-base validation.
17
-
18
- ## Evidence Rules
19
-
20
- - Static evidence covers type/lint/build/schema checks; behavior evidence covers tests or interaction checks that actually exercise the flow.
21
- - Shell exit status is authoritative. Classify checks as `passed`, `failed`, `not-run`, `blocked`, or `unavailable`; only passed satisfies a required check.
22
- - Never use static success as behavior proof, or tests as visual/browser proof they did not exercise.
23
- - Unavailable commands remain gaps, not weaker success claims.
24
- - Resolve component/design evidence as knowledge base first, then `<repoRoot>/openSpec/**` when setup/query fails or has no match. Mark `openSpec fallback` as the active project specification source; do not report it as missing merely because the connector format is TODO.
25
-
26
- ## Method And Output
27
-
28
- Inventory required checks, map each to fresh evidence, classify gaps, confirm review
29
- pass, and state only proven user-visible changes.
30
-
31
- Return Markdown headings:
32
-
33
- - `Changes`: changed behavior and areas.
34
- - `Verification Evidence`: table of check, command/source, status, and artifact/result.
35
- - `Review Result`: exact review verdict and findings.
36
- - `Known Risks`: missing optional checks and environment caveats.
37
- - `Follow-up`: concrete work or `None`.
38
-
39
- Do not edit files. Do not claim complete when review is not pass or a required check
40
- is failed, not-run, blocked, unavailable, stale, or contradicted.
1
+ ---
2
+ name: frontend-verification
3
+ description: Use to assess frontend evidence and produce closeout.
4
+ references:
5
+ - path: references/verification-checklist.md
6
+ required: true
7
+ maxChars: 2800
8
+ ---
9
+
10
+ # Frontend Verification
11
+
12
+ Use for `frontend-closeout-pi`. Shell nodes execute commands; this read-only skill
13
+ assesses their evidence. Read the checklist first.
14
+
15
+ Inputs: criteria, change inventory, shell commands/status/artifacts, review
16
+ verdict/findings, and required browser, visual, manual, or knowledge evidence.
17
+
18
+ ## Evidence Rules
19
+
20
+ - Static evidence covers type/lint/build/schema; behavior evidence must exercise the flow.
21
+ - Shell exit status is authoritative. Classify as `passed`, `failed`, `not-run`,
22
+ `blocked`, or `unavailable`; only passed satisfies a required check.
23
+ - Never use static success as behavior proof, or tests as visual/browser proof they did not exercise.
24
+ - Mock-backed behavior proves frontend rendering and state transitions only. It never
25
+ proves backend readiness, transport compatibility, or real API integration.
26
+ - Unavailable commands remain gaps.
27
+ - Resolve design evidence via knowledge base, then `<repoRoot>/openSpec/**` after
28
+ failure/no match. Its connector format remains TODO; never invent it. An applied
29
+ `openSpec fallback` is available project evidence.
30
+ - Separate Mock service/handler checks from page consumption and record the
31
+ dev/test-only boundary; handler tests alone do not prove page use.
32
+
33
+ ## Method And Output
34
+
35
+ Map required checks to fresh evidence, classify gaps, confirm review pass, and state
36
+ only proven changes.
37
+
38
+ Return Markdown headings:
39
+
40
+ - `Changes`: changed behavior and areas.
41
+ - `Mock Decision`, `Mock Files`, `Mock Verification`, `Production Boundary`: status is `passed`, `failed`, `not-required`, `blocked`, or `unavailable`.
42
+ - `Verification Evidence`: table of check, command/source, status, and artifact/result.
43
+ - `Review Result`: exact review verdict and findings.
44
+ - `Known Risks`: missing optional checks and environment caveats.
45
+ - `Follow-up`: concrete work or `None`.
46
+
47
+ When the backend remains unavailable but required mock-backed checks pass, state
48
+ `Frontend status: mock-validated` and `Real integration: pending`. Use a completed
49
+ `<task-id>-real-api-integration-verify` task before changing the latter to complete;
50
+ the follow-up is explicit, not auto-created or auto-executed.
51
+
52
+ Do not edit files. Do not claim complete when review is not pass or a required check
53
+ is failed, not-run, blocked, unavailable, stale, or contradicted.
@@ -1,56 +1,68 @@
1
- # Frontend Verification Checklist
2
-
3
- ## Static And Behavior Evidence
4
-
5
- - Type/compile, lint/format, build, and schema/client checks ran when required.
6
- - Generated output was authorized.
7
- - Unit/component/integration tests cover changed logic and flows; regressions pass.
8
- - Browser/e2e/manual evidence exists when explicitly required.
9
- - API/mock behavior and applicable loading, empty, error, success, disabled, permission, retry, and boundary states have evidence.
10
-
11
- ## Design And Component Evidence
12
-
13
- - Claims cite traceable knowledge-base retrieval or relevant `<repoRoot>/openSpec/**` fallback evidence.
14
- - Retrieval includes connector/query, collection/document ID, version when available, and time.
15
- - When connector setup/query fails or has no match, evidence shows recursive `openSpec/` search terms, inspected paths, matched headings/lines, and applied rules.
16
- - Relevant `openSpec/` matches become the current-project specification and satisfy source availability; only the knowledge-base connection remains unavailable.
17
- - Missing both sources blocks explicit compliance or an unresolved required design decision.
18
-
19
- ## Status
20
-
21
- - `passed`: fresh successful evidence matches current implementation.
22
- - `failed`: check ran and failed.
23
- - `not-run`: no fresh attempt exists.
24
- - `blocked`: a prerequisite prevented execution.
25
- - `unavailable`: tool, environment, connector, or source was absent.
26
-
27
- Only passed satisfies a required check. Other optional statuses remain disclosed risks.
28
-
29
- ## Closeout Checks
30
-
31
- - List exact commands, source, exit status, and archived output/artifact when available.
32
- - Map every criterion to evidence or a named gap.
33
- - Record review verdict before completion.
34
- - Do not conflate static, behavior, browser/visual/manual, or knowledge-base proof.
35
-
36
- ```markdown
37
- ## Changes
38
- - ...
39
-
40
- ## Verification Evidence
41
- | Check | Command or source | Status | Evidence |
42
- |---|---|---|---|
43
- | ... | ... | passed | ... |
44
-
45
- ## Review Result
46
- - Verdict: `VERDICT: pass`
47
-
48
- ## Known Risks
49
- - ...
50
-
51
- ## Follow-up
52
- - None.
53
- ```
54
-
55
- If review is not pass or a required check is not passed, describe the task as
56
- incomplete and list concrete follow-up.
1
+ # Frontend Verification Checklist
2
+
3
+ ## Static And Behavior Evidence
4
+
5
+ - Type/compile, lint/format, build, and schema/client checks ran when required.
6
+ - Generated output was authorized.
7
+ - Tests cover changed logic/flows and regressions.
8
+ - Browser/e2e/manual evidence exists when explicitly required.
9
+ - Selected-strategy behavior and applicable loading, empty, error, success, disabled,
10
+ permission, retry, and boundary states have evidence from fixed DAG entrypoints.
11
+ - When generated, Mock-specific verification checks service/handler/schema/fixtures;
12
+ behavior evidence separately proves page consumption.
13
+ - For Mock strategies, activation is explicit/non-production and a production/default-
14
+ real-path build with Mock off confirms the real request remains default.
15
+ - `not-needed` has positive readiness/no-remote evidence plus applicable real or
16
+ no-remote behavior evidence. Mock-backed evidence remains frontend-only and never
17
+ satisfies real API integration.
18
+
19
+ ## Design And Component Evidence
20
+
21
+ - Claims cite knowledge-base retrieval or `<repoRoot>/openSpec/**` fallback.
22
+ - Evidence records query/source/version/time or fallback search terms, paths, headings, and applied rules.
23
+ - Relevant `openSpec/` matches become the current-project specification and satisfy source availability; only the knowledge-base connection remains unavailable.
24
+ - Missing both sources blocks explicit compliance or an unresolved required design decision.
25
+
26
+ ## Status
27
+
28
+ - `passed`: fresh successful evidence matches current implementation.
29
+ - `failed`: check ran and failed.
30
+ - `not-run`: no fresh attempt exists.
31
+ - `blocked`: a prerequisite prevented execution.
32
+ - `unavailable`: tool, environment, connector, or source was absent.
33
+
34
+ Only passed satisfies a required check. Other optional statuses remain disclosed risks.
35
+
36
+ ## Closeout Checks
37
+
38
+ - List exact commands, source, exit status, and archived output/artifact when available.
39
+ - Map every criterion to evidence or a named gap.
40
+ - Record review verdict before completion.
41
+ - Do not conflate static, behavior, browser/visual/manual, or knowledge-base proof.
42
+ - If only mock evidence exists, report `Frontend status: mock-validated` and
43
+ `Real integration: pending`, with the actual API verification as follow-up.
44
+
45
+ ```markdown
46
+ ## Changes
47
+ - ...
48
+
49
+ ## Verification Evidence
50
+ | Check | Command or source | Status | Evidence |
51
+ |---|---|---|---|
52
+ | ... | ... | passed | ... |
53
+
54
+ ## Mock Decision / Mock Files / Mock Verification / Production Boundary
55
+ - Status: `passed | failed | not-required | blocked | unavailable`
56
+
57
+ ## Review Result
58
+ - Verdict: `VERDICT: pass`
59
+
60
+ ## Known Risks
61
+ - ...
62
+
63
+ ## Follow-up
64
+ - None.
65
+ ```
66
+
67
+ If review is not pass or a required check is not passed, describe the task as
68
+ incomplete and list concrete follow-up.
@@ -1,10 +1,10 @@
1
- ---
2
- name: grill-me
3
- description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
4
- ---
5
-
6
- Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
7
-
8
- Ask the questions one at a time.
9
-
10
- If a question can be answered by exploring the codebase, explore the codebase instead.
1
+ ---
2
+ name: grill-me
3
+ description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
4
+ ---
5
+
6
+ Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
7
+
8
+ Ask the questions one at a time.
9
+
10
+ If a question can be answered by exploring the codebase, explore the codebase instead.