@tea-agent/loop-agent 0.12.0 → 0.13.0-beta.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (284) hide show
  1. package/AGENTS.md +155 -153
  2. package/CHANGELOG.md +338 -265
  3. package/README.md +345 -298
  4. package/bin/agent-worker.js +22 -22
  5. package/bin/loop-agent.js +21 -21
  6. package/dist/application/dag/generate-task-dag.js +28 -28
  7. package/dist/application/evaluation/candidate-hash.js +75 -0
  8. package/dist/application/evaluation/candidate.js +52 -0
  9. package/dist/application/evaluation/replay.js +289 -0
  10. package/dist/application/evaluation/types.js +130 -0
  11. package/dist/cli/command-definitions.js +27 -7
  12. package/dist/cli/program.js +8 -4
  13. package/dist/commands/cursor-prompt.js +6 -6
  14. package/dist/commands/eval.js +235 -0
  15. package/dist/commands/init.js +544 -506
  16. package/dist/commands/knowledge.js +129 -31
  17. package/dist/commands/loop-benchmark.js +11 -11
  18. package/dist/commands/pi-reuse-benchmark.js +16 -16
  19. package/dist/executors/pi-sdk-executor.js +38 -24
  20. package/dist/executors/shell-executor.js +34 -2
  21. package/dist/executors/shell-presets.js +20 -0
  22. package/dist/executors/shell-verification.js +7 -0
  23. package/dist/governance/manifest-types.js +4 -0
  24. package/dist/infrastructure/evaluation/candidate-store.js +435 -0
  25. package/dist/infrastructure/evaluation/store.js +40 -0
  26. package/dist/sidecars/cursor-prompt/executor.js +1 -1
  27. package/dist/task/config-types.js +28 -1
  28. package/dist/task/runtime.js +27 -27
  29. package/dist/worker/cli.js +96 -1
  30. package/dist/worker/delivery/package.js +3 -3
  31. package/dist/worker/feature/decision-loader.js +37 -6
  32. package/dist/worker/feature/next-action.js +10 -2
  33. package/dist/worker/feature/ready-plan-projection.js +81 -0
  34. package/dist/worker/feature/reducer.js +2 -1
  35. package/dist/worker/feature/review.js +19 -2
  36. package/dist/worker/feature/run.js +27 -2
  37. package/dist/worker/follow-up/approve.js +5 -2
  38. package/dist/worker/follow-up/factory.js +1 -1
  39. package/dist/worker/observability/read-model.js +246 -41
  40. package/dist/worker/observe/routes.js +173 -15
  41. package/dist/worker/observe/spec-evidence.js +281 -0
  42. package/dist/worker/observe/static/api.js +46 -27
  43. package/dist/worker/observe/static/app.js +150 -150
  44. package/dist/worker/observe/static/constants.js +148 -148
  45. package/dist/worker/observe/static/copy.js +67 -67
  46. package/dist/worker/observe/static/dag-helpers.js +172 -172
  47. package/dist/worker/observe/static/dag-layout.d.ts +31 -31
  48. package/dist/worker/observe/static/dag-layout.js +83 -83
  49. package/dist/worker/observe/static/dag-model.js +72 -72
  50. package/dist/worker/observe/static/dom.js +61 -61
  51. package/dist/worker/observe/static/format-pool.js +67 -67
  52. package/dist/worker/observe/static/format.js +292 -292
  53. package/dist/worker/observe/static/index.html +308 -308
  54. package/dist/worker/observe/static/kpi.js +94 -94
  55. package/dist/worker/observe/static/relations.js +133 -128
  56. package/dist/worker/observe/static/router.js +93 -85
  57. package/dist/worker/observe/static/run-processing.js +148 -148
  58. package/dist/worker/observe/static/shell-chrome.js +68 -68
  59. package/dist/worker/observe/static/state.js +253 -253
  60. package/dist/worker/observe/static/styles.css +1902 -1890
  61. package/dist/worker/observe/static/views/batch.js +227 -226
  62. package/dist/worker/observe/static/views/dag-graph.js +172 -172
  63. package/dist/worker/observe/static/views/dag-inspector.js +607 -477
  64. package/dist/worker/observe/static/views/dag.js +362 -362
  65. package/dist/worker/observe/static/views/dashboard.js +445 -442
  66. package/dist/worker/observe/static/views/failures.js +143 -143
  67. package/dist/worker/observe/static/views/feature.js +492 -453
  68. package/dist/worker/observe/static/views/pool.js +350 -347
  69. package/dist/worker/observe/static/views/run.js +453 -453
  70. package/dist/worker/observe/static/views/session-timeline.js +205 -205
  71. package/dist/worker/observe/static/views/shell.js +7 -7
  72. package/dist/worker/observe/static/views/task.js +314 -260
  73. package/dist/worker/observe/static/views/timeline.js +163 -163
  74. package/dist/worker/pool/doctor.js +165 -0
  75. package/dist/worker/pool/migrate-state.js +303 -0
  76. package/dist/worker/pool/run-store.js +205 -17
  77. package/dist/worker/pool/types.js +17 -1
  78. package/dist/worker/pool/validation.js +100 -15
  79. package/dist/worker/report/morning-report.js +12 -2
  80. package/dist/worker/runner/run-ready.js +41 -26
  81. package/dist/worker/task-graph/ready-planner.js +136 -0
  82. package/dist/workflows/dag/backend-test-analysis-contract.js +120 -0
  83. package/dist/workflows/dag/canvas-observer.js +275 -275
  84. package/dist/workflows/dag/convergence/controller.js +16 -8
  85. package/dist/workflows/dag/dynamic-runtime/map.js +90 -2
  86. package/dist/workflows/dag/failure-routing.js +12 -1
  87. package/dist/workflows/dag/init-hybrid.js +2404 -360
  88. package/dist/workflows/dag/node-execution.js +9 -0
  89. package/dist/workflows/dag/prompt.js +9 -0
  90. package/dist/workflows/dag/report.js +35 -1
  91. package/dist/workflows/dag/runner.js +28 -2
  92. package/dist/workflows/dag/task-demand-routing.js +383 -0
  93. package/dist/workflows/dag/types.js +51 -13
  94. package/dist/workflows/dag/upstream-artifacts.js +1 -0
  95. package/dist/workflows/dag/validate.js +59 -1
  96. package/docs/README.md +106 -104
  97. package/docs/agent-dag-recovery-playbook.md +195 -184
  98. package/docs/agent-dag-runner.md +67 -67
  99. package/docs/architecture/README.md +26 -26
  100. package/docs/architecture/dag-execution.md +140 -140
  101. package/docs/architecture/evolution.md +54 -53
  102. package/docs/architecture/facts-and-state.md +71 -58
  103. package/docs/architecture/runtime-boundaries.md +191 -191
  104. package/docs/architecture/system-overview.md +93 -93
  105. package/docs/architecture/worker-and-feature.md +85 -81
  106. package/docs/cursor-prompt-sidecar.md +36 -36
  107. package/docs/decisions/README.md +18 -15
  108. package/docs/design/README.md +167 -77
  109. package/docs/development-principles.md +73 -73
  110. package/docs/exec-plans/README.md +6 -6
  111. package/docs/exec-plans/active/README.md +15 -9
  112. package/docs/exec-plans/completed/README.md +85 -73
  113. package/docs/feature-workflow.md +389 -261
  114. package/docs/harness-methodology-debugging.md +153 -153
  115. package/docs/harness-methodology-tdd.md +130 -130
  116. package/docs/harness-methodology-verification.md +27 -27
  117. package/docs/init-surface.manifest.json +289 -280
  118. package/docs/loop-agent-harness.md +142 -130
  119. package/docs/production-readiness.md +96 -96
  120. package/docs/progress/README.md +64 -54
  121. package/docs/reports/README.md +117 -94
  122. package/docs/skills/README.md +7 -7
  123. package/docs/skills/vetted-skill-registry.md +29 -27
  124. package/docs/templates/adr.md +60 -60
  125. package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
  126. package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
  127. package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
  128. package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
  129. package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
  130. package/docs/templates/agent-dag-report.schema.json +473 -473
  131. package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
  132. package/docs/templates/agent-dag.base.json +190 -190
  133. package/docs/templates/agent-dag.final-verification.json +185 -185
  134. package/docs/templates/agent-dag.schema.json +411 -383
  135. package/docs/templates/agent-dag.supervised-implementation.json +501 -501
  136. package/docs/templates/backend-test-analysis.schema.json +44 -0
  137. package/docs/templates/backend-test-dag.generate-pytest.prompt.md +202 -139
  138. package/docs/templates/backend-test-dag.json +311 -276
  139. package/docs/templates/backend-test-dag.retrospect.prompt.md +125 -125
  140. package/docs/templates/backend-test-dag.review-cases.prompt.md +81 -81
  141. package/docs/templates/exec-plan.md +64 -64
  142. package/docs/templates/feature-spec.md +53 -53
  143. package/docs/templates/frontend-design-contract.md +42 -33
  144. package/docs/templates/frontend-task-constraints.md +35 -25
  145. package/docs/templates/frontend-task-requirement.md +70 -61
  146. package/docs/templates/frontend-test-dag.generate-cases.prompt.md +5 -0
  147. package/docs/templates/frontend-test-dag.json +23 -0
  148. package/docs/templates/frontend-test-dag.retrieve-context.prompt.md +3 -0
  149. package/docs/templates/frontend-test-dag.retrospect.prompt.md +3 -0
  150. package/docs/templates/frontend-test-dag.review-cases.prompt.md +3 -0
  151. package/docs/templates/frontend-test-dag.review-execution.prompt.md +3 -0
  152. package/docs/templates/harness.schema.json +221 -221
  153. package/docs/templates/hybrid-dag.json +188 -188
  154. package/docs/templates/init-evolution-review.md +35 -35
  155. package/docs/templates/interactive-ui-round2-experiment.md +66 -66
  156. package/docs/templates/knowledge-graph-bootstrap-dag.json +118 -0
  157. package/docs/templates/knowledge-sync-dag.json +178 -0
  158. package/docs/templates/knowledge-sync-draft.schema.json +71 -0
  159. package/docs/templates/product-line/AGENTS.md +8 -8
  160. package/docs/templates/product-line/README.md +9 -9
  161. package/docs/templates/product-line/acceptance.yaml +14 -14
  162. package/docs/templates/product-line/closeout.yaml +9 -9
  163. package/docs/templates/product-line/design.md +13 -13
  164. package/docs/templates/product-line/links.md +10 -10
  165. package/docs/templates/product-line/requirement.md +17 -17
  166. package/docs/templates/product-line/task-graph.yaml +15 -15
  167. package/docs/templates/product-line/task.yaml +64 -64
  168. package/docs/templates/product-line/test-plan.md +7 -7
  169. package/docs/templates/production-readiness-checklist.md +57 -57
  170. package/docs/templates/progress-log.md +17 -17
  171. package/docs/templates/project-start-checklist.md +9 -9
  172. package/docs/templates/qa-report.md +48 -48
  173. package/docs/templates/sprint-contract.md +29 -29
  174. package/docs/templates/worker-dogfood-evidence.md +80 -80
  175. package/docs/templates/worker-dogfood-setup.md +68 -68
  176. package/docs/verification-matrix.md +70 -66
  177. package/examples/decision-gate-agent-dag.json +177 -177
  178. package/examples/example-dag.json +46 -46
  179. package/examples/hybrid-loop-agent-dag.json +189 -189
  180. package/harness.json +66 -66
  181. package/package.json +88 -46
  182. package/scripts/check-product-line-docs.sh +29 -29
  183. package/scripts/check-task-pool-root.sh +32 -32
  184. package/scripts/kb-bootstrap-init-skeleton.sh +240 -0
  185. package/scripts/kb-graph-incremental-prepare.mjs +386 -0
  186. package/scripts/kb-graph-incremental-prepare.sh +5 -0
  187. package/scripts/kb-graph-materialize.mjs +105 -0
  188. package/scripts/kb-graph-materialize.sh +4 -0
  189. package/scripts/kb-graph-promote.mjs +164 -0
  190. package/scripts/kb-graph-promote.sh +4 -0
  191. package/scripts/kb-query.mjs +554 -0
  192. package/scripts/kb-query.sh +5 -0
  193. package/skills/agent-worker/SKILL.md +39 -37
  194. package/skills/agent-worker/references/agent-worker-operator.md +60 -43
  195. package/skills/ai-engineering-context/SKILL.md +48 -48
  196. package/skills/analyze-product-dependencies/SKILL.md +67 -0
  197. package/skills/analyze-product-dependencies/agents/openai.yaml +4 -0
  198. package/skills/analyze-product-dependencies/references/api-documentation-schema.md +30 -0
  199. package/skills/analyze-product-dependencies/references/dependency-analysis-schema.md +28 -0
  200. package/skills/analyze-product-dependencies/references/example.md +76 -0
  201. package/skills/analyze-product-dependencies/references/forward-test-cases.md +35 -0
  202. package/skills/analyze-product-dependencies/references/input-contract.md +11 -0
  203. package/skills/analyze-product-dependencies/references/scouting-rules.md +61 -0
  204. package/skills/analyze-product-dependencies/scripts/test-validators.mjs +267 -0
  205. package/skills/analyze-product-dependencies/scripts/validate-api-documentation.mjs +101 -0
  206. package/skills/analyze-product-dependencies/scripts/validate-dependency-analysis.mjs +142 -0
  207. package/skills/analyze-product-dependencies/scripts/validate-product-requirement-input.mjs +76 -0
  208. package/skills/analyze-product-dependencies/scripts/validation-helpers.mjs +146 -0
  209. package/skills/analyze-product-requirements/SKILL.md +90 -0
  210. package/skills/analyze-product-requirements/agents/openai.yaml +4 -0
  211. package/skills/analyze-product-requirements/references/acceptance-criteria.md +91 -0
  212. package/skills/analyze-product-requirements/references/clarification-and-knowledge.md +56 -0
  213. package/skills/analyze-product-requirements/references/example.md +86 -0
  214. package/skills/analyze-product-requirements/references/forward-test-cases.md +66 -0
  215. package/skills/analyze-product-requirements/references/product-analysis-schema.md +32 -0
  216. package/skills/analyze-product-requirements/references/product-requirement-schema.md +33 -0
  217. package/skills/analyze-product-requirements/references/requirement-clarification-schema.md +35 -0
  218. package/skills/analyze-product-requirements/scripts/test-validators.mjs +193 -0
  219. package/skills/analyze-product-requirements/scripts/validate-product-analysis.mjs +69 -0
  220. package/skills/analyze-product-requirements/scripts/validate-product-requirement.mjs +97 -0
  221. package/skills/analyze-product-requirements/scripts/validate-requirement-clarification.mjs +98 -0
  222. package/skills/analyze-product-requirements/scripts/validation-helpers.mjs +156 -0
  223. package/skills/code-review-core/SKILL.md +20 -20
  224. package/skills/codebase-scout/SKILL.md +19 -19
  225. package/skills/frontend-design-review/SKILL.md +66 -59
  226. package/skills/frontend-design-review/references/review-checklist.md +58 -37
  227. package/skills/frontend-implementation/SKILL.md +47 -51
  228. package/skills/frontend-implementation/references/code-standards.md +32 -34
  229. package/skills/frontend-implementation/references/design-spec.md +46 -46
  230. package/skills/frontend-implementation/references/node-contracts.md +76 -32
  231. package/skills/frontend-review/SKILL.md +59 -53
  232. package/skills/frontend-review/references/review-findings.md +47 -42
  233. package/skills/frontend-verification/SKILL.md +53 -40
  234. package/skills/frontend-verification/references/verification-checklist.md +68 -56
  235. package/skills/grill-me/SKILL.md +10 -10
  236. package/skills/grill-with-docs/SKILL.md +88 -88
  237. package/skills/grill-with-docs/adr-format.md +47 -47
  238. package/skills/grill-with-docs/context-format.md +60 -60
  239. package/skills/init-capability-evolution/SKILL.md +70 -70
  240. package/skills/loop-agent/SKILL.md +151 -151
  241. package/skills/loop-agent/references/README.md +67 -67
  242. package/skills/loop-agent/references/command-reference.md +505 -452
  243. package/skills/loop-agent/references/docs-converge.md +126 -126
  244. package/skills/loop-agent/references/harness-policy.md +263 -263
  245. package/skills/loop-agent/references/hybrid-dag.md +238 -233
  246. package/skills/loop-agent/references/learned/README.md +21 -21
  247. package/skills/loop-agent/references/long-running-loop.md +57 -57
  248. package/skills/loop-agent/references/model-routing.md +36 -36
  249. package/skills/loop-agent/references/multi-worktree.md +54 -54
  250. package/skills/loop-agent/references/one-shot-runs.md +85 -85
  251. package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
  252. package/skills/loop-agent/references/pi-prompt.md +23 -23
  253. package/skills/loop-agent/references/pi-subagent-assisted-mode.md +84 -84
  254. package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
  255. package/skills/loop-agent/references/task-workflow.md +89 -89
  256. package/skills/loop-agent/references/verification-and-failure-handling.md +139 -139
  257. package/skills/playwright-cli/SKILL.md +420 -0
  258. package/skills/playwright-cli/references/element-attributes.md +23 -0
  259. package/skills/playwright-cli/references/playwright-tests.md +39 -0
  260. package/skills/playwright-cli/references/request-mocking.md +87 -0
  261. package/skills/playwright-cli/references/running-code.md +241 -0
  262. package/skills/playwright-cli/references/session-management.md +225 -0
  263. package/skills/playwright-cli/references/storage-state.md +275 -0
  264. package/skills/playwright-cli/references/test-generation.md +433 -0
  265. package/skills/playwright-cli/references/tracing.md +139 -0
  266. package/skills/playwright-cli/references/video-recording.md +143 -0
  267. package/skills/playwright-cli-case-generator/SKILL.md +74 -0
  268. package/skills/requesting-code-review/SKILL.md +101 -101
  269. package/skills/requesting-code-review/code-reviewer.md +168 -168
  270. package/skills/systematic-debugging/CREATION-LOG.md +119 -119
  271. package/skills/systematic-debugging/SKILL.md +296 -296
  272. package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
  273. package/skills/systematic-debugging/condition-based-waiting.md +115 -115
  274. package/skills/systematic-debugging/defense-in-depth.md +122 -122
  275. package/skills/systematic-debugging/find-polluter.sh +63 -63
  276. package/skills/systematic-debugging/root-cause-tracing.md +169 -169
  277. package/skills/systematic-debugging/test-academic.md +14 -14
  278. package/skills/systematic-debugging/test-pressure-1.md +58 -58
  279. package/skills/systematic-debugging/test-pressure-2.md +68 -68
  280. package/skills/systematic-debugging/test-pressure-3.md +69 -69
  281. package/skills/test-driven-development/SKILL.md +20 -20
  282. package/skills/using-git-worktrees/SKILL.md +215 -215
  283. package/skills/verification-before-completion/SKILL.md +154 -154
  284. package/skills/webapp-testing/SKILL.md +19 -19
@@ -1,154 +1,154 @@
1
- ---
2
- name: verification-before-completion
3
- description: 在宣称 work complete、fixed 或 passing,或在 commit / 创建 PR 之前使用——须先运行 verification commands 并确认 output,再作任何 success claims;始终 evidence before assertions
4
- ---
5
-
6
- # Verification Before Completion
7
-
8
- ## Overview
9
-
10
- 未经验证就宣称 work complete 是不诚实,不是效率。
11
-
12
- **Core principle:** 始终 evidence before claims。
13
-
14
- **违反本条字面即违反其精神。**
15
-
16
- ## The Iron Law
17
-
18
- ```
19
- NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
20
- ```
21
-
22
- 若本 message 中尚未运行 verification command,不得宣称 passes。
23
-
24
- ## The Gate Function
25
-
26
- ```
27
- BEFORE claiming any status or expressing satisfaction:
28
-
29
- 1. IDENTIFY: What command proves this claim?
30
- 2. RUN: Execute the FULL command (fresh, complete)
31
- 3. READ: Full output, check exit code, count failures
32
- 4. VERIFY: Does output confirm the claim?
33
- - If NO: State actual status with evidence
34
- - If YES: State claim WITH evidence
35
- 5. ONLY THEN: Make the claim
36
-
37
- Skip any step = lying, not verifying
38
- ```
39
-
40
- ## Common Failures
41
-
42
- | Claim | Requires | Not Sufficient |
43
- |-------|----------|----------------|
44
- | Tests pass | Test command output: 0 failures | Previous run, "should pass" |
45
- | Linter clean | Linter output: 0 errors | Partial check, extrapolation |
46
- | Build succeeds | Build command: exit 0 | Linter passing, logs look good |
47
- | Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
48
- | Regression test works | Red-green cycle verified | Test passes once |
49
- | Agent completed | VCS diff shows changes | Agent reports "success" |
50
- | Requirements met | Line-by-line checklist | Tests passing |
51
- | Harness integrity | `scripts/check-repo.sh` exit 0 | Files look correct |
52
- | Harness CI | `scripts/ci.sh` exit 0 | Individual checks pass |
53
- | Contract handoff | Handoff checklist completed + progress/report updated | "Should be fine"
54
-
55
- ## Red Flags - STOP
56
-
57
- - 使用 "should"、"probably"、"seems to"
58
- - 验证前表达满意("Great!"、"Perfect!"、"Done!" 等)
59
- - 未验证就要 commit/push/PR
60
- - 信任 agent success reports
61
- - 依赖 partial verification
62
- - 认为 "just this once"
63
- - 疲惫想结束工作
64
- - **任何未运行 verification 却暗示 success 的措辞**
65
-
66
- ## Rationalization Prevention
67
-
68
- | Excuse | Reality |
69
- |--------|---------|
70
- | "Should work now" | RUN the verification |
71
- | "I'm confident" | Confidence ≠ evidence |
72
- | "Just this once" | No exceptions |
73
- | "Linter passed" | Linter ≠ compiler |
74
- | "Agent said success" | Verify independently |
75
- | "I'm tired" | Exhaustion ≠ excuse |
76
- | "Partial check is enough" | Partial proves nothing |
77
- | "Different words so rule doesn't apply" | Spirit over letter |
78
-
79
- ## Key Patterns
80
-
81
- **Tests:**
82
- ```
83
- ✅ [Run test command] [See: 34/34 pass] "All tests pass"
84
- ❌ "Should pass now" / "Looks correct"
85
- ```
86
-
87
- **Regression tests (TDD Red-Green):**
88
- ```
89
- ✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
90
- ❌ "I've written a regression test" (without red-green verification)
91
- ```
92
-
93
- **Build:**
94
- ```
95
- ✅ [Run build] [See: exit 0] "Build passes"
96
- ❌ "Linter passed" (linter doesn't check compilation)
97
- ```
98
-
99
- **Requirements:**
100
- ```
101
- ✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
102
- ❌ "Tests pass, phase complete"
103
- ```
104
-
105
- **Agent delegation:**
106
- ```
107
- ✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
108
- ❌ Trust agent report
109
- ```
110
-
111
- ## Harness-Specific Verification
112
-
113
- 在 harness-governed repo 中工作(存在 `harness.json`)时:
114
-
115
- - **Docs/structure changes** → `bash scripts/check-repo.sh`
116
- - **Full-repo delivery** → `bash scripts/ci.sh`
117
- - **Cross-platform changes** → 验证 OpenCode 与 Pi-Agent 两条路径
118
- - **Contract changes** → 验证 contract docs 已更新 + tests 对齐
119
- - **Handoff** → 宣称 complete 前运行 `handoff check`
120
-
121
- 完整 command 选择见项目 `docs/verification-matrix.md`。
122
-
123
- ## Why This Matters
124
-
125
- 来自 24 条 failure memories:
126
- - human partner 说 "I don't believe you" — trust 已破裂
127
- - Undefined functions 已 ship — 会 crash
128
- - Missing requirements 已 ship — 功能不完整
129
- - 虚假完成浪费时间 → redirect → rework
130
- - 违反:"Honesty is a core value. If you lie, you'll be replaced."
131
-
132
- ## When To Apply
133
-
134
- **在以下情况之前 ALWAYS:**
135
- - 任何 success/completion claims 的变体
136
- - 任何表达满意
137
- - 任何关于 work state 的正面陈述
138
- - Commit、PR creation、task completion
139
- - 进入 next task
140
- - 委派给 agents
141
-
142
- **规则适用于:**
143
- - 精确短语
144
- - paraphrases 与同义词
145
- - success 的暗示
146
- - 任何暗示 completion/correctness 的沟通
147
-
148
- ## The Bottom Line
149
-
150
- **Verification 无捷径。**
151
-
152
- Run the command. Read the output. THEN claim the result.
153
-
154
- This is non-negotiable.
1
+ ---
2
+ name: verification-before-completion
3
+ description: 在宣称 work complete、fixed 或 passing,或在 commit / 创建 PR 之前使用——须先运行 verification commands 并确认 output,再作任何 success claims;始终 evidence before assertions
4
+ ---
5
+
6
+ # Verification Before Completion
7
+
8
+ ## Overview
9
+
10
+ 未经验证就宣称 work complete 是不诚实,不是效率。
11
+
12
+ **Core principle:** 始终 evidence before claims。
13
+
14
+ **违反本条字面即违反其精神。**
15
+
16
+ ## The Iron Law
17
+
18
+ ```
19
+ NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
20
+ ```
21
+
22
+ 若本 message 中尚未运行 verification command,不得宣称 passes。
23
+
24
+ ## The Gate Function
25
+
26
+ ```
27
+ BEFORE claiming any status or expressing satisfaction:
28
+
29
+ 1. IDENTIFY: What command proves this claim?
30
+ 2. RUN: Execute the FULL command (fresh, complete)
31
+ 3. READ: Full output, check exit code, count failures
32
+ 4. VERIFY: Does output confirm the claim?
33
+ - If NO: State actual status with evidence
34
+ - If YES: State claim WITH evidence
35
+ 5. ONLY THEN: Make the claim
36
+
37
+ Skip any step = lying, not verifying
38
+ ```
39
+
40
+ ## Common Failures
41
+
42
+ | Claim | Requires | Not Sufficient |
43
+ |-------|----------|----------------|
44
+ | Tests pass | Test command output: 0 failures | Previous run, "should pass" |
45
+ | Linter clean | Linter output: 0 errors | Partial check, extrapolation |
46
+ | Build succeeds | Build command: exit 0 | Linter passing, logs look good |
47
+ | Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
48
+ | Regression test works | Red-green cycle verified | Test passes once |
49
+ | Agent completed | VCS diff shows changes | Agent reports "success" |
50
+ | Requirements met | Line-by-line checklist | Tests passing |
51
+ | Harness integrity | `scripts/check-repo.sh` exit 0 | Files look correct |
52
+ | Harness CI | `scripts/ci.sh` exit 0 | Individual checks pass |
53
+ | Contract handoff | Handoff checklist completed + progress/report updated | "Should be fine"
54
+
55
+ ## Red Flags - STOP
56
+
57
+ - 使用 "should"、"probably"、"seems to"
58
+ - 验证前表达满意("Great!"、"Perfect!"、"Done!" 等)
59
+ - 未验证就要 commit/push/PR
60
+ - 信任 agent success reports
61
+ - 依赖 partial verification
62
+ - 认为 "just this once"
63
+ - 疲惫想结束工作
64
+ - **任何未运行 verification 却暗示 success 的措辞**
65
+
66
+ ## Rationalization Prevention
67
+
68
+ | Excuse | Reality |
69
+ |--------|---------|
70
+ | "Should work now" | RUN the verification |
71
+ | "I'm confident" | Confidence ≠ evidence |
72
+ | "Just this once" | No exceptions |
73
+ | "Linter passed" | Linter ≠ compiler |
74
+ | "Agent said success" | Verify independently |
75
+ | "I'm tired" | Exhaustion ≠ excuse |
76
+ | "Partial check is enough" | Partial proves nothing |
77
+ | "Different words so rule doesn't apply" | Spirit over letter |
78
+
79
+ ## Key Patterns
80
+
81
+ **Tests:**
82
+ ```
83
+ ✅ [Run test command] [See: 34/34 pass] "All tests pass"
84
+ ❌ "Should pass now" / "Looks correct"
85
+ ```
86
+
87
+ **Regression tests (TDD Red-Green):**
88
+ ```
89
+ ✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
90
+ ❌ "I've written a regression test" (without red-green verification)
91
+ ```
92
+
93
+ **Build:**
94
+ ```
95
+ ✅ [Run build] [See: exit 0] "Build passes"
96
+ ❌ "Linter passed" (linter doesn't check compilation)
97
+ ```
98
+
99
+ **Requirements:**
100
+ ```
101
+ ✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
102
+ ❌ "Tests pass, phase complete"
103
+ ```
104
+
105
+ **Agent delegation:**
106
+ ```
107
+ ✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
108
+ ❌ Trust agent report
109
+ ```
110
+
111
+ ## Harness-Specific Verification
112
+
113
+ 在 harness-governed repo 中工作(存在 `harness.json`)时:
114
+
115
+ - **Docs/structure changes** → `bash scripts/check-repo.sh`
116
+ - **Full-repo delivery** → `bash scripts/ci.sh`
117
+ - **Cross-platform changes** → 验证 OpenCode 与 Pi-Agent 两条路径
118
+ - **Contract changes** → 验证 contract docs 已更新 + tests 对齐
119
+ - **Handoff** → 宣称 complete 前运行 `handoff check`
120
+
121
+ 完整 command 选择见项目 `docs/verification-matrix.md`。
122
+
123
+ ## Why This Matters
124
+
125
+ 来自 24 条 failure memories:
126
+ - human partner 说 "I don't believe you" — trust 已破裂
127
+ - Undefined functions 已 ship — 会 crash
128
+ - Missing requirements 已 ship — 功能不完整
129
+ - 虚假完成浪费时间 → redirect → rework
130
+ - 违反:"Honesty is a core value. If you lie, you'll be replaced."
131
+
132
+ ## When To Apply
133
+
134
+ **在以下情况之前 ALWAYS:**
135
+ - 任何 success/completion claims 的变体
136
+ - 任何表达满意
137
+ - 任何关于 work state 的正面陈述
138
+ - Commit、PR creation、task completion
139
+ - 进入 next task
140
+ - 委派给 agents
141
+
142
+ **规则适用于:**
143
+ - 精确短语
144
+ - paraphrases 与同义词
145
+ - success 的暗示
146
+ - 任何暗示 completion/correctness 的沟通
147
+
148
+ ## The Bottom Line
149
+
150
+ **Verification 无捷径。**
151
+
152
+ Run the command. Read the output. THEN claim the result.
153
+
154
+ This is non-negotiable.
@@ -1,19 +1,19 @@
1
- ---
2
- name: webapp-testing
3
- description: 任务明确涉及 browser 渲染行为时,用于前端或本地 web UI 验证。
4
- ---
5
-
6
- # Webapp Testing
7
-
8
- 仅当任务包含 browser UI 或本地 web app 时使用本 skill。
9
-
10
- ## 规则
11
-
12
- - 优先使用项目现有的 dev server 与 test tooling。
13
- - 当 visual 或 interaction 行为重要时,用 browser 或文档化的 UI test 命令验证渲染行为。
14
- - UI 有变更时,检查 desktop 与 mobile 布局的 overlap、clipping、blank state。
15
- - 默认不添加 networked services 或第三方 scan。
16
-
17
- ## Output
18
-
19
- 报告确切的 server 命令、URL、browser/test 命令与观察结果。
1
+ ---
2
+ name: webapp-testing
3
+ description: 任务明确涉及 browser 渲染行为时,用于前端或本地 web UI 验证。
4
+ ---
5
+
6
+ # Webapp Testing
7
+
8
+ 仅当任务包含 browser UI 或本地 web app 时使用本 skill。
9
+
10
+ ## 规则
11
+
12
+ - 优先使用项目现有的 dev server 与 test tooling。
13
+ - 当 visual 或 interaction 行为重要时,用 browser 或文档化的 UI test 命令验证渲染行为。
14
+ - UI 有变更时,检查 desktop 与 mobile 布局的 overlap、clipping、blank state。
15
+ - 默认不添加 networked services 或第三方 scan。
16
+
17
+ ## Output
18
+
19
+ 报告确切的 server 命令、URL、browser/test 命令与观察结果。