@tea-agent/loop-agent 0.13.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (272) hide show
  1. package/AGENTS.md +157 -157
  2. package/CHANGELOG.md +116 -305
  3. package/README.md +357 -334
  4. package/bin/agent-worker.js +22 -22
  5. package/bin/loop-agent.js +21 -21
  6. package/dist/commands/cursor-prompt.js +6 -6
  7. package/dist/commands/init.js +505 -505
  8. package/dist/commands/loop-benchmark.js +11 -11
  9. package/dist/commands/pi-reuse-benchmark.js +16 -16
  10. package/dist/executors/pi-event-serializer.js +33 -11
  11. package/dist/sidecars/cursor-prompt/executor.js +1 -1
  12. package/dist/task/runtime.js +27 -27
  13. package/dist/worker/observe/spec-evidence.js +19 -10
  14. package/dist/worker/observe/static/api.js +46 -46
  15. package/dist/worker/observe/static/app.js +151 -150
  16. package/dist/worker/observe/static/constants.js +156 -148
  17. package/dist/worker/observe/static/copy.js +67 -67
  18. package/dist/worker/observe/static/dag-helpers.js +201 -172
  19. package/dist/worker/observe/static/dag-layout.d.ts +31 -31
  20. package/dist/worker/observe/static/dag-layout.js +83 -83
  21. package/dist/worker/observe/static/dag-model.js +72 -72
  22. package/dist/worker/observe/static/dom.js +122 -122
  23. package/dist/worker/observe/static/format-pool.d.ts +71 -0
  24. package/dist/worker/observe/static/format-pool.js +134 -67
  25. package/dist/worker/observe/static/format.js +317 -292
  26. package/dist/worker/observe/static/index.html +350 -308
  27. package/dist/worker/observe/static/kpi.js +100 -94
  28. package/dist/worker/observe/static/markdown-render.js +124 -0
  29. package/dist/worker/observe/static/relations.js +133 -133
  30. package/dist/worker/observe/static/router.js +93 -93
  31. package/dist/worker/observe/static/run-processing.js +148 -148
  32. package/dist/worker/observe/static/shell-chrome.js +74 -68
  33. package/dist/worker/observe/static/state.js +273 -267
  34. package/dist/worker/observe/static/styles.css +2504 -1902
  35. package/dist/worker/observe/static/views/batch.js +227 -227
  36. package/dist/worker/observe/static/views/dag-graph.js +172 -172
  37. package/dist/worker/observe/static/views/dag-inspector.js +530 -627
  38. package/dist/worker/observe/static/views/dag.js +371 -371
  39. package/dist/worker/observe/static/views/dashboard.js +86 -100
  40. package/dist/worker/observe/static/views/failures.js +143 -143
  41. package/dist/worker/observe/static/views/feature.js +492 -492
  42. package/dist/worker/observe/static/views/pool.js +708 -350
  43. package/dist/worker/observe/static/views/run.js +453 -453
  44. package/dist/worker/observe/static/views/session-timeline.js +771 -219
  45. package/dist/worker/observe/static/views/shell.js +7 -7
  46. package/dist/worker/observe/static/views/task.js +314 -314
  47. package/dist/worker/observe/static/views/timeline.js +163 -163
  48. package/dist/workflows/dag/canvas-observer.js +275 -275
  49. package/dist/workflows/dag/init-hybrid.js +27 -11
  50. package/docs/README.md +106 -104
  51. package/docs/architecture/README.md +26 -26
  52. package/docs/architecture/dag-execution.md +140 -140
  53. package/docs/architecture/evolution.md +54 -54
  54. package/docs/architecture/facts-and-state.md +71 -71
  55. package/docs/architecture/runtime-boundaries.md +191 -191
  56. package/docs/architecture/system-overview.md +93 -93
  57. package/docs/architecture/worker-and-feature.md +85 -85
  58. package/docs/harness-methodology-debugging.md +153 -153
  59. package/docs/harness-methodology-tdd.md +130 -130
  60. package/docs/harness-methodology-verification.md +27 -27
  61. package/docs/init-surface.manifest.json +304 -307
  62. package/docs/skills/README.md +7 -7
  63. package/docs/skills/vetted-skill-registry.md +29 -29
  64. package/docs/templates/adr.md +60 -60
  65. package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
  66. package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
  67. package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
  68. package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
  69. package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
  70. package/docs/templates/agent-dag-report.schema.json +473 -473
  71. package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
  72. package/docs/templates/agent-dag.base.json +190 -190
  73. package/docs/templates/agent-dag.final-verification.json +185 -185
  74. package/docs/templates/agent-dag.schema.json +411 -411
  75. package/docs/templates/agent-dag.supervised-implementation.json +620 -620
  76. package/docs/templates/backend-test-analysis.schema.json +44 -44
  77. package/docs/templates/backend-test-case-manifest.schema.json +190 -190
  78. package/docs/templates/backend-test-dag.classify.prompt.md +75 -75
  79. package/docs/templates/backend-test-dag.generate-pytest.prompt.md +204 -204
  80. package/docs/templates/backend-test-dag.json +559 -559
  81. package/docs/templates/backend-test-dag.retrospect.prompt.md +139 -139
  82. package/docs/templates/backend-test-dag.review-cases.prompt.md +83 -83
  83. package/docs/templates/backend-test-execution.schema.json +133 -133
  84. package/docs/templates/backend-test-result.schema.json +99 -99
  85. package/docs/templates/branch-merge-report.md +0 -1
  86. package/docs/templates/exec-plan.md +64 -64
  87. package/docs/templates/feature-spec.md +53 -53
  88. package/docs/templates/frontend-design-contract.md +42 -42
  89. package/docs/templates/frontend-eval/fixtures/failures/01-type-build-error.md +17 -17
  90. package/docs/templates/frontend-eval/fixtures/failures/02-unit-component-test-fail.md +16 -16
  91. package/docs/templates/frontend-eval/fixtures/failures/03-fixture-schema-drift.md +16 -16
  92. package/docs/templates/frontend-eval/fixtures/failures/04-missing-loading-empty-error-state.md +16 -16
  93. package/docs/templates/frontend-eval/fixtures/failures/05-forbidden-write-writeset-expansion.md +16 -16
  94. package/docs/templates/frontend-eval/fixtures/failures/06-unapproved-dependency-add.md +16 -16
  95. package/docs/templates/frontend-eval/fixtures/failures/07-mock-production-on.md +21 -21
  96. package/docs/templates/frontend-eval/fixtures/functional/01-simple-component-style.md +29 -29
  97. package/docs/templates/frontend-eval/fixtures/functional/02-form-validation.md +28 -28
  98. package/docs/templates/frontend-eval/fixtures/functional/03-list-detail-page.md +28 -28
  99. package/docs/templates/frontend-eval/fixtures/functional/04-api-mock.md +29 -29
  100. package/docs/templates/frontend-eval/fixtures/functional/05-permission-auth-gated-ui.md +27 -27
  101. package/docs/templates/frontend-eval/fixtures/functional/06-ssr-server-client-boundary.md +28 -28
  102. package/docs/templates/frontend-eval/fixtures/functional/07-shared-public-component-api.md +28 -28
  103. package/docs/templates/frontend-eval/fixtures/functional/08-pure-local-no-remote.md +27 -27
  104. package/docs/templates/frontend-eval/metrics.md +138 -138
  105. package/docs/templates/frontend-eval/smoke-targets.md +53 -53
  106. package/docs/templates/frontend-implementation-contract.schema.json +27 -27
  107. package/docs/templates/frontend-task-constraints.md +35 -35
  108. package/docs/templates/frontend-task-requirement.md +70 -70
  109. package/docs/templates/frontend-test-dag.generate-cases.prompt.md +5 -5
  110. package/docs/templates/frontend-test-dag.json +23 -23
  111. package/docs/templates/frontend-test-dag.retrieve-context.prompt.md +3 -3
  112. package/docs/templates/frontend-test-dag.retrospect.prompt.md +3 -3
  113. package/docs/templates/frontend-test-dag.review-cases.prompt.md +3 -3
  114. package/docs/templates/frontend-test-dag.review-execution.prompt.md +3 -3
  115. package/docs/templates/harness.schema.json +221 -221
  116. package/docs/templates/hybrid-dag.json +188 -188
  117. package/docs/templates/init-evolution-review.md +35 -35
  118. package/docs/templates/interactive-ui-round2-experiment.md +66 -66
  119. package/docs/templates/knowledge-graph-bootstrap-dag.json +118 -118
  120. package/docs/templates/knowledge-sync-dag.json +178 -178
  121. package/docs/templates/knowledge-sync-draft.schema.json +71 -71
  122. package/docs/templates/product-line/AGENTS.md +8 -8
  123. package/docs/templates/product-line/README.md +9 -9
  124. package/docs/templates/product-line/acceptance.yaml +14 -14
  125. package/docs/templates/product-line/closeout.yaml +9 -9
  126. package/docs/templates/product-line/design.md +13 -13
  127. package/docs/templates/product-line/links.md +10 -10
  128. package/docs/templates/product-line/requirement.md +17 -17
  129. package/docs/templates/product-line/task-graph.yaml +15 -15
  130. package/docs/templates/product-line/task.yaml +64 -64
  131. package/docs/templates/product-line/test-plan.md +7 -7
  132. package/docs/templates/production-readiness-checklist.md +57 -57
  133. package/docs/templates/progress-log.md +17 -17
  134. package/docs/templates/project-start-checklist.md +9 -9
  135. package/docs/templates/qa-report.md +48 -48
  136. package/docs/templates/sprint-contract.md +29 -29
  137. package/docs/templates/worker-dogfood-evidence.md +80 -80
  138. package/docs/templates/worker-dogfood-setup.md +68 -68
  139. package/examples/decision-gate-agent-dag.json +173 -173
  140. package/examples/example-dag.json +46 -46
  141. package/examples/hybrid-loop-agent-dag.json +188 -188
  142. package/harness.json +66 -66
  143. package/package.json +78 -52
  144. package/scripts/kb-bootstrap-init-skeleton.sh +240 -240
  145. package/scripts/kb-graph-incremental-prepare.mjs +386 -386
  146. package/scripts/kb-graph-materialize.mjs +105 -105
  147. package/scripts/kb-graph-promote.mjs +164 -164
  148. package/scripts/kb-query.mjs +554 -554
  149. package/skills/agent-worker/SKILL.md +39 -39
  150. package/skills/agent-worker/references/agent-worker-operator.md +60 -60
  151. package/skills/ai-engineering-context/SKILL.md +48 -48
  152. package/skills/analyze-product-dependencies/SKILL.md +67 -67
  153. package/skills/analyze-product-dependencies/agents/openai.yaml +4 -4
  154. package/skills/analyze-product-dependencies/references/api-documentation-schema.md +30 -30
  155. package/skills/analyze-product-dependencies/references/dependency-analysis-schema.md +28 -28
  156. package/skills/analyze-product-dependencies/references/example.md +76 -76
  157. package/skills/analyze-product-dependencies/references/forward-test-cases.md +35 -35
  158. package/skills/analyze-product-dependencies/references/input-contract.md +11 -11
  159. package/skills/analyze-product-dependencies/references/scouting-rules.md +61 -61
  160. package/skills/analyze-product-dependencies/scripts/test-validators.mjs +267 -267
  161. package/skills/analyze-product-dependencies/scripts/validate-api-documentation.mjs +101 -101
  162. package/skills/analyze-product-dependencies/scripts/validate-dependency-analysis.mjs +142 -142
  163. package/skills/analyze-product-dependencies/scripts/validate-product-requirement-input.mjs +76 -76
  164. package/skills/analyze-product-dependencies/scripts/validation-helpers.mjs +146 -146
  165. package/skills/analyze-product-requirements/SKILL.md +90 -90
  166. package/skills/analyze-product-requirements/agents/openai.yaml +4 -4
  167. package/skills/analyze-product-requirements/references/acceptance-criteria.md +91 -91
  168. package/skills/analyze-product-requirements/references/clarification-and-knowledge.md +56 -56
  169. package/skills/analyze-product-requirements/references/example.md +86 -86
  170. package/skills/analyze-product-requirements/references/forward-test-cases.md +66 -66
  171. package/skills/analyze-product-requirements/references/product-analysis-schema.md +32 -32
  172. package/skills/analyze-product-requirements/references/product-requirement-schema.md +33 -33
  173. package/skills/analyze-product-requirements/references/requirement-clarification-schema.md +35 -35
  174. package/skills/analyze-product-requirements/scripts/test-validators.mjs +193 -193
  175. package/skills/analyze-product-requirements/scripts/validate-product-analysis.mjs +69 -69
  176. package/skills/analyze-product-requirements/scripts/validate-product-requirement.mjs +97 -97
  177. package/skills/analyze-product-requirements/scripts/validate-requirement-clarification.mjs +98 -98
  178. package/skills/analyze-product-requirements/scripts/validation-helpers.mjs +156 -156
  179. package/skills/browser-tools/SKILL.md +196 -196
  180. package/skills/browser-tools/browser-content.js +103 -103
  181. package/skills/browser-tools/browser-cookies.js +35 -35
  182. package/skills/browser-tools/browser-eval.js +53 -53
  183. package/skills/browser-tools/browser-hn-scraper.js +108 -108
  184. package/skills/browser-tools/browser-nav.js +44 -44
  185. package/skills/browser-tools/browser-pick.js +162 -162
  186. package/skills/browser-tools/browser-screenshot.js +34 -34
  187. package/skills/browser-tools/browser-start.js +86 -86
  188. package/skills/browser-tools/package-lock.json +2556 -2556
  189. package/skills/browser-tools/package.json +19 -19
  190. package/skills/code-review-core/SKILL.md +20 -20
  191. package/skills/codebase-scout/SKILL.md +19 -19
  192. package/skills/frontend-design-review/SKILL.md +66 -66
  193. package/skills/frontend-design-review/references/review-checklist.md +40 -58
  194. package/skills/frontend-implementation/SKILL.md +49 -49
  195. package/skills/frontend-implementation/references/code-standards.md +32 -32
  196. package/skills/frontend-implementation/references/design-spec.md +46 -46
  197. package/skills/frontend-implementation/references/node-contracts.md +27 -27
  198. package/skills/frontend-review/SKILL.md +61 -59
  199. package/skills/frontend-review/references/review-findings.md +48 -47
  200. package/skills/frontend-verification/SKILL.md +55 -53
  201. package/skills/frontend-verification/references/verification-checklist.md +59 -68
  202. package/skills/grill-me/SKILL.md +10 -10
  203. package/skills/grill-with-docs/SKILL.md +88 -88
  204. package/skills/grill-with-docs/adr-format.md +47 -47
  205. package/skills/grill-with-docs/context-format.md +60 -60
  206. package/skills/init-capability-evolution/SKILL.md +70 -70
  207. package/skills/loop-agent/SKILL.md +151 -151
  208. package/skills/loop-agent/references/README.md +67 -67
  209. package/skills/loop-agent/references/command-reference.md +527 -527
  210. package/skills/loop-agent/references/docs-converge.md +126 -126
  211. package/skills/loop-agent/references/harness-policy.md +263 -263
  212. package/skills/loop-agent/references/hybrid-dag.md +243 -243
  213. package/skills/loop-agent/references/learned/README.md +21 -21
  214. package/skills/loop-agent/references/long-running-loop.md +57 -57
  215. package/skills/loop-agent/references/model-routing.md +36 -36
  216. package/skills/loop-agent/references/multi-worktree.md +54 -54
  217. package/skills/loop-agent/references/one-shot-runs.md +85 -85
  218. package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
  219. package/skills/loop-agent/references/pi-prompt.md +23 -23
  220. package/skills/loop-agent/references/pi-subagent-assisted-mode.md +84 -84
  221. package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
  222. package/skills/loop-agent/references/task-workflow.md +89 -89
  223. package/skills/loop-agent/references/verification-and-failure-handling.md +141 -141
  224. package/skills/playwright-cli/SKILL.md +420 -420
  225. package/skills/playwright-cli/references/element-attributes.md +23 -23
  226. package/skills/playwright-cli/references/playwright-tests.md +39 -39
  227. package/skills/playwright-cli/references/request-mocking.md +87 -87
  228. package/skills/playwright-cli/references/running-code.md +241 -241
  229. package/skills/playwright-cli/references/session-management.md +225 -225
  230. package/skills/playwright-cli/references/storage-state.md +275 -275
  231. package/skills/playwright-cli/references/test-generation.md +433 -433
  232. package/skills/playwright-cli/references/tracing.md +139 -139
  233. package/skills/playwright-cli/references/video-recording.md +143 -143
  234. package/skills/playwright-cli-case-generator/SKILL.md +74 -74
  235. package/skills/requesting-code-review/SKILL.md +101 -101
  236. package/skills/requesting-code-review/code-reviewer.md +168 -168
  237. package/skills/systematic-debugging/CREATION-LOG.md +119 -119
  238. package/skills/systematic-debugging/SKILL.md +296 -296
  239. package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
  240. package/skills/systematic-debugging/condition-based-waiting.md +115 -115
  241. package/skills/systematic-debugging/defense-in-depth.md +122 -122
  242. package/skills/systematic-debugging/find-polluter.sh +63 -63
  243. package/skills/systematic-debugging/root-cause-tracing.md +169 -169
  244. package/skills/systematic-debugging/test-academic.md +14 -14
  245. package/skills/systematic-debugging/test-pressure-1.md +58 -58
  246. package/skills/systematic-debugging/test-pressure-2.md +68 -68
  247. package/skills/systematic-debugging/test-pressure-3.md +69 -69
  248. package/skills/test-driven-development/SKILL.md +20 -20
  249. package/skills/using-git-worktrees/SKILL.md +215 -215
  250. package/skills/verification-before-completion/SKILL.md +154 -154
  251. package/skills/webapp-testing/SKILL.md +19 -19
  252. package/docs/agent-dag-recovery-playbook.md +0 -195
  253. package/docs/agent-dag-runner.md +0 -67
  254. package/docs/cursor-prompt-sidecar.md +0 -36
  255. package/docs/decisions/README.md +0 -18
  256. package/docs/design/README.md +0 -167
  257. package/docs/development-principles.md +0 -73
  258. package/docs/exec-plans/README.md +0 -6
  259. package/docs/exec-plans/active/README.md +0 -12
  260. package/docs/exec-plans/completed/README.md +0 -107
  261. package/docs/feature-workflow.md +0 -414
  262. package/docs/loop-agent-harness.md +0 -142
  263. package/docs/production-readiness.md +0 -96
  264. package/docs/progress/README.md +0 -80
  265. package/docs/reports/README.md +0 -159
  266. package/docs/verification-matrix.md +0 -70
  267. package/scripts/check-product-line-docs.sh +0 -29
  268. package/scripts/check-task-pool-root.sh +0 -32
  269. package/scripts/kb-graph-incremental-prepare.sh +0 -5
  270. package/scripts/kb-graph-materialize.sh +0 -4
  271. package/scripts/kb-graph-promote.sh +0 -4
  272. package/scripts/kb-query.sh +0 -5
@@ -1,296 +1,296 @@
1
- ---
2
- name: systematic-debugging
3
- description: 遇到任何 bug、test failure 或 unexpected behavior 时使用,且在提出 fixes 之前
4
- ---
5
-
6
- # Systematic Debugging
7
-
8
- ## Overview
9
-
10
- Random fixes 浪费时间并制造新 bug。Quick patches 掩盖 underlying issues。
11
-
12
- **Core principle:** ALWAYS 在尝试 fixes 之前找到 root cause。Symptom fixes 是 failure。
13
-
14
- **违反本流程字面即违反 debugging 精神。**
15
-
16
- ## The Iron Law
17
-
18
- ```
19
- NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST
20
- ```
21
-
22
- 若尚未完成 Phase 1,不得提出 fixes。
23
-
24
- ## When to Use
25
-
26
- 用于 ANY technical issue:
27
- - Test failures
28
- - Production bugs
29
- - Unexpected behavior
30
- - Performance problems
31
- - Build failures
32
- - Integration issues
33
-
34
- **ESPECIALLY 在以下情况使用:**
35
- - 时间压力下(emergencies 使 guessing 诱人)
36
- - "Just one quick fix" 看起来 obvious
37
- - 已尝试 multiple fixes
38
- - Previous fix 无效
39
- - 未完全理解 issue
40
-
41
- **Don't skip when:**
42
- - Issue 看起来 simple(simple bugs 也有 root causes)
43
- - 赶时间(rushing 保证 rework)
44
- - Manager 要求 NOW 修好(systematic 比 thrashing 更快)
45
-
46
- ## The Four Phases
47
-
48
- 进入下一阶段前 MUST 完成每一 phase。
49
-
50
- ### Phase 1: Root Cause Investigation
51
-
52
- **在尝试 ANY fix 之前:**
53
-
54
- 1. **Read Error Messages Carefully**
55
- - 不要跳过 errors 或 warnings
56
- - 它们常含 exact solution
57
- - 完整阅读 stack traces
58
- - 记下 line numbers、file paths、error codes
59
-
60
- 2. **Reproduce Consistently**
61
- - 能否可靠触发?
62
- - Exact steps 是什么?
63
- - 是否每次都发生?
64
- - 若不可 reproduce → 收集更多 data,不要 guess
65
-
66
- 3. **Check Recent Changes**
67
- - 什么变更可能导致此问题?
68
- - Git diff、recent commits
69
- - New dependencies、config changes
70
- - Environmental differences
71
-
72
- 4. **Gather Evidence in Multi-Component Systems**
73
-
74
- **WHEN system 有多个 components(CI → build → signing,API → service → database):**
75
-
76
- **BEFORE proposing fixes,添加 diagnostic instrumentation:**
77
- ```
78
- For EACH component boundary:
79
- - Log what data enters component
80
- - Log what data exits component
81
- - Verify environment/config propagation
82
- - Check state at each layer
83
-
84
- Run once to gather evidence showing WHERE it breaks
85
- THEN analyze evidence to identify failing component
86
- THEN investigate that specific component
87
- ```
88
-
89
- **Example (multi-layer system):**
90
- ```bash
91
- # Layer 1: Workflow
92
- echo "=== Secrets available in workflow: ==="
93
- echo "IDENTITY: ${IDENTITY:+SET}${IDENTITY:-UNSET}"
94
-
95
- # Layer 2: Build script
96
- echo "=== Env vars in build script: ==="
97
- env | grep IDENTITY || echo "IDENTITY not in environment"
98
-
99
- # Layer 3: Signing script
100
- echo "=== Keychain state: ==="
101
- security list-keychains
102
- security find-identity -v
103
-
104
- # Layer 4: Actual signing
105
- codesign --sign "$IDENTITY" --verbose=4 "$APP"
106
- ```
107
-
108
- **This reveals:** Which layer fails (secrets → workflow ✓, workflow → build ✗)
109
-
110
- 5. **Trace Data Flow**
111
-
112
- **WHEN error 在 call stack 深处:**
113
-
114
- 完整 backward tracing 见本目录 `root-cause-tracing.md`。
115
-
116
- **Quick version:**
117
- - Bad value 从哪 originate?
118
- - 谁用 bad value 调用了 this?
119
- - 一直向上 trace 直到 source
120
- - 在 source 修复,而非 symptom
121
-
122
- ### Phase 2: Pattern Analysis
123
-
124
- **Fix 前先找 pattern:**
125
-
126
- 1. **Find Working Examples**
127
- - 在同 codebase 找 similar working code
128
- - 什么能 work、什么 broken?
129
-
130
- 2. **Compare Against References**
131
- - 若实现 pattern,COMPLETE 阅读 reference implementation
132
- - 不要 skim — 读每一行
133
- - 应用前 fully 理解 pattern
134
-
135
- 3. **Identify Differences**
136
- - Working 与 broken 有何不同?
137
- - 列出 every difference,再小也要列
138
- - 不要假设 "that can't matter"
139
-
140
- 4. **Understand Dependencies**
141
- - 还需要哪些 other components?
142
- - 哪些 settings、config、environment?
143
- - 它作哪些 assumptions?
144
-
145
- ### Phase 3: Hypothesis and Testing
146
-
147
- **Scientific method:**
148
-
149
- 1. **Form Single Hypothesis**
150
- - 清楚陈述:"I think X is the root cause because Y"
151
- - 写下来
152
- - 要 specific,不要 vague
153
-
154
- 2. **Test Minimally**
155
- - 做 SMALLEST possible change 以 test hypothesis
156
- - One variable at a time
157
- - 不要一次 fix multiple things
158
-
159
- 3. **Verify Before Continuing**
160
- - 有效?Yes → Phase 4
161
- - 无效?Form NEW hypothesis
162
- - DON'T 在其上叠加更多 fixes
163
-
164
- 4. **When You Don't Know**
165
- - 说 "I don't understand X"
166
- - 不要假装知道
167
- - Ask for help
168
- - Research more
169
-
170
- ### Phase 4: Implementation
171
-
172
- **Fix root cause,不是 symptom:**
173
-
174
- 1. **Create Failing Test Case**
175
- - Simplest possible reproduction
176
- - 可能的话用 automated test
177
- - 无 framework 时用 one-off test script
178
- - MUST 在 fix 之前有
179
- - 遵循 RED-GREEN-REFACTOR:写 failing test,看它 fail,再 fix
180
-
181
- 2. **Implement Single Fix**
182
- - 针对已识别的 root cause
183
- - ONE change at a time
184
- - 无 "while I'm here" improvements
185
- - 无 bundled refactoring
186
-
187
- 3. **Verify Fix**
188
- - Test 现在 pass?
189
- - 无 other tests broken?
190
- - Issue 真的 resolved?
191
-
192
- 4. **If Fix Doesn't Work**
193
- - STOP
194
- - Count:已尝试多少 fixes?
195
- - 若 < 3:Return to Phase 1,用 new information 再分析
196
- - **若 ≥ 3:STOP 并质疑 architecture(见下方 step 5)**
197
- - DON'T 在未做 architectural discussion 前尝试 Fix #4
198
-
199
- 5. **If 3+ Fixes Failed: Question Architecture**
200
-
201
- **表明 architectural problem 的 pattern:**
202
- - 每个 fix 在不同位置 reveal 新的 shared state/coupling/problem
203
- - Fixes 需要 "massive refactoring" 才能实现
204
- - 每个 fix 在其他地方制造新 symptoms
205
-
206
- **STOP 并质疑 fundamentals:**
207
- - 此 pattern fundamentally sound 吗?
208
- - 是否 "sticking with it through sheer inertia"?
209
- - 应 refactor architecture 还是继续 fix symptoms?
210
-
211
- **Discuss with your human partner before attempting more fixes**
212
-
213
- This is NOT a failed hypothesis — this is a wrong architecture.
214
-
215
- ## Red Flags - STOP and Follow Process
216
-
217
- 若发现自己想:
218
- - "Quick fix for now, investigate later"
219
- - "Just try changing X and see if it works"
220
- - "Add multiple changes, run tests"
221
- - "Skip the test, I'll manually verify"
222
- - "It's probably X, let me fix that"
223
- - "I don't fully understand but this might work"
224
- - "Pattern says X but I'll adapt it differently"
225
- - "Here are the main problems: [lists fixes without investigation]"
226
- - Proposing solutions before tracing data flow
227
- - **"One more fix attempt" (when already tried 2+)**
228
- - **Each fix reveals new problem in different place**
229
-
230
- **ALL of these mean: STOP. Return to Phase 1.**
231
-
232
- **If 3+ fixes failed:** Question the architecture (see Phase 4.5)
233
-
234
- ## your human partner's Signals You're Doing It Wrong
235
-
236
- **Watch for these redirections:**
237
- - "Is that not happening?" — You assumed without verifying
238
- - "Will it show us...?" — You should have added evidence gathering
239
- - "Stop guessing" — You're proposing fixes without understanding
240
- - "Ultrathink this" — Question fundamentals, not just symptoms
241
- - "We're stuck?" (frustrated) — Your approach isn't working
242
-
243
- **When you see these:** STOP. Return to Phase 1.
244
-
245
- ## Common Rationalizations
246
-
247
- | Excuse | Reality |
248
- |--------|---------|
249
- | "Issue is simple, don't need process" | Simple issues 也有 root causes。Process 对 simple bugs 很快。 |
250
- | "Emergency, no time for process" | Systematic debugging 比 guess-and-check thrashing 更快。 |
251
- | "Just try this first, then investigate" | First fix 定模式。从一开始就做对。 |
252
- | "I'll write test after confirming fix works" | Untested fixes 不 stick。Test first 证明它。 |
253
- | "Multiple fixes at once saves time" | 无法 isolate what worked。制造新 bugs。 |
254
- | "Reference too long, I'll adapt the pattern" | Partial understanding 保证 bugs。Complete 阅读。 |
255
- | "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause。 |
256
- | "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem。质疑 pattern,不要再 fix。 |
257
-
258
- ## Quick Reference
259
-
260
- | Phase | Key Activities | Success Criteria |
261
- |-------|---------------|------------------|
262
- | **1. Root Cause** | Read errors, reproduce, check changes, gather evidence | Understand WHAT and WHY |
263
- | **2. Pattern** | Find working examples, compare | Identify differences |
264
- | **3. Hypothesis** | Form theory, test minimally | Confirmed or new hypothesis |
265
- | **4. Implementation** | Create test, fix, verify | Bug resolved, tests pass |
266
-
267
- ## When Process Reveals "No Root Cause"
268
-
269
- 若 systematic investigation 表明 issue truly environmental、timing-dependent 或 external:
270
-
271
- 1. You've completed the process
272
- 2. Document what you investigated
273
- 3. Implement appropriate handling (retry, timeout, error message)
274
- 4. Add monitoring/logging for future investigation
275
-
276
- **But:** 95% 的 "no root cause" cases 是 incomplete investigation。
277
-
278
- ## Supporting Techniques
279
-
280
- 本目录中属于 systematic debugging 的技术:
281
-
282
- - **`root-cause-tracing.md`** — Trace bugs backward through call stack 找 original trigger
283
- - **`defense-in-depth.md`** — 找到 root cause 后在 multiple layers 加 validation
284
- - **`condition-based-waiting.md`** — 用 condition polling 替代 arbitrary timeouts
285
-
286
- **Related principles:**
287
- - **RED-GREEN-REFACTOR**(见 `ai_workspace/loop-agent/harness-methodology-tdd.md`)— 用于 creating failing test case(Phase 4, Step 1)
288
- - **Verification discipline** — 宣称 success 前 verify fix worked。Run verification command,读 output,THEN claim result。
289
-
290
- ## Real-World Impact
291
-
292
- 来自 debugging sessions:
293
- - Systematic approach:15-30 分钟 fix
294
- - Random fixes approach:2-3 小时 thrashing
295
- - First-time fix rate:95% vs 40%
296
- - New bugs introduced:Near zero vs common
1
+ ---
2
+ name: systematic-debugging
3
+ description: 遇到任何 bug、test failure 或 unexpected behavior 时使用,且在提出 fixes 之前
4
+ ---
5
+
6
+ # Systematic Debugging
7
+
8
+ ## Overview
9
+
10
+ Random fixes 浪费时间并制造新 bug。Quick patches 掩盖 underlying issues。
11
+
12
+ **Core principle:** ALWAYS 在尝试 fixes 之前找到 root cause。Symptom fixes 是 failure。
13
+
14
+ **违反本流程字面即违反 debugging 精神。**
15
+
16
+ ## The Iron Law
17
+
18
+ ```
19
+ NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST
20
+ ```
21
+
22
+ 若尚未完成 Phase 1,不得提出 fixes。
23
+
24
+ ## When to Use
25
+
26
+ 用于 ANY technical issue:
27
+ - Test failures
28
+ - Production bugs
29
+ - Unexpected behavior
30
+ - Performance problems
31
+ - Build failures
32
+ - Integration issues
33
+
34
+ **ESPECIALLY 在以下情况使用:**
35
+ - 时间压力下(emergencies 使 guessing 诱人)
36
+ - "Just one quick fix" 看起来 obvious
37
+ - 已尝试 multiple fixes
38
+ - Previous fix 无效
39
+ - 未完全理解 issue
40
+
41
+ **Don't skip when:**
42
+ - Issue 看起来 simple(simple bugs 也有 root causes)
43
+ - 赶时间(rushing 保证 rework)
44
+ - Manager 要求 NOW 修好(systematic 比 thrashing 更快)
45
+
46
+ ## The Four Phases
47
+
48
+ 进入下一阶段前 MUST 完成每一 phase。
49
+
50
+ ### Phase 1: Root Cause Investigation
51
+
52
+ **在尝试 ANY fix 之前:**
53
+
54
+ 1. **Read Error Messages Carefully**
55
+ - 不要跳过 errors 或 warnings
56
+ - 它们常含 exact solution
57
+ - 完整阅读 stack traces
58
+ - 记下 line numbers、file paths、error codes
59
+
60
+ 2. **Reproduce Consistently**
61
+ - 能否可靠触发?
62
+ - Exact steps 是什么?
63
+ - 是否每次都发生?
64
+ - 若不可 reproduce → 收集更多 data,不要 guess
65
+
66
+ 3. **Check Recent Changes**
67
+ - 什么变更可能导致此问题?
68
+ - Git diff、recent commits
69
+ - New dependencies、config changes
70
+ - Environmental differences
71
+
72
+ 4. **Gather Evidence in Multi-Component Systems**
73
+
74
+ **WHEN system 有多个 components(CI → build → signing,API → service → database):**
75
+
76
+ **BEFORE proposing fixes,添加 diagnostic instrumentation:**
77
+ ```
78
+ For EACH component boundary:
79
+ - Log what data enters component
80
+ - Log what data exits component
81
+ - Verify environment/config propagation
82
+ - Check state at each layer
83
+
84
+ Run once to gather evidence showing WHERE it breaks
85
+ THEN analyze evidence to identify failing component
86
+ THEN investigate that specific component
87
+ ```
88
+
89
+ **Example (multi-layer system):**
90
+ ```bash
91
+ # Layer 1: Workflow
92
+ echo "=== Secrets available in workflow: ==="
93
+ echo "IDENTITY: ${IDENTITY:+SET}${IDENTITY:-UNSET}"
94
+
95
+ # Layer 2: Build script
96
+ echo "=== Env vars in build script: ==="
97
+ env | grep IDENTITY || echo "IDENTITY not in environment"
98
+
99
+ # Layer 3: Signing script
100
+ echo "=== Keychain state: ==="
101
+ security list-keychains
102
+ security find-identity -v
103
+
104
+ # Layer 4: Actual signing
105
+ codesign --sign "$IDENTITY" --verbose=4 "$APP"
106
+ ```
107
+
108
+ **This reveals:** Which layer fails (secrets → workflow ✓, workflow → build ✗)
109
+
110
+ 5. **Trace Data Flow**
111
+
112
+ **WHEN error 在 call stack 深处:**
113
+
114
+ 完整 backward tracing 见本目录 `root-cause-tracing.md`。
115
+
116
+ **Quick version:**
117
+ - Bad value 从哪 originate?
118
+ - 谁用 bad value 调用了 this?
119
+ - 一直向上 trace 直到 source
120
+ - 在 source 修复,而非 symptom
121
+
122
+ ### Phase 2: Pattern Analysis
123
+
124
+ **Fix 前先找 pattern:**
125
+
126
+ 1. **Find Working Examples**
127
+ - 在同 codebase 找 similar working code
128
+ - 什么能 work、什么 broken?
129
+
130
+ 2. **Compare Against References**
131
+ - 若实现 pattern,COMPLETE 阅读 reference implementation
132
+ - 不要 skim — 读每一行
133
+ - 应用前 fully 理解 pattern
134
+
135
+ 3. **Identify Differences**
136
+ - Working 与 broken 有何不同?
137
+ - 列出 every difference,再小也要列
138
+ - 不要假设 "that can't matter"
139
+
140
+ 4. **Understand Dependencies**
141
+ - 还需要哪些 other components?
142
+ - 哪些 settings、config、environment?
143
+ - 它作哪些 assumptions?
144
+
145
+ ### Phase 3: Hypothesis and Testing
146
+
147
+ **Scientific method:**
148
+
149
+ 1. **Form Single Hypothesis**
150
+ - 清楚陈述:"I think X is the root cause because Y"
151
+ - 写下来
152
+ - 要 specific,不要 vague
153
+
154
+ 2. **Test Minimally**
155
+ - 做 SMALLEST possible change 以 test hypothesis
156
+ - One variable at a time
157
+ - 不要一次 fix multiple things
158
+
159
+ 3. **Verify Before Continuing**
160
+ - 有效?Yes → Phase 4
161
+ - 无效?Form NEW hypothesis
162
+ - DON'T 在其上叠加更多 fixes
163
+
164
+ 4. **When You Don't Know**
165
+ - 说 "I don't understand X"
166
+ - 不要假装知道
167
+ - Ask for help
168
+ - Research more
169
+
170
+ ### Phase 4: Implementation
171
+
172
+ **Fix root cause,不是 symptom:**
173
+
174
+ 1. **Create Failing Test Case**
175
+ - Simplest possible reproduction
176
+ - 可能的话用 automated test
177
+ - 无 framework 时用 one-off test script
178
+ - MUST 在 fix 之前有
179
+ - 遵循 RED-GREEN-REFACTOR:写 failing test,看它 fail,再 fix
180
+
181
+ 2. **Implement Single Fix**
182
+ - 针对已识别的 root cause
183
+ - ONE change at a time
184
+ - 无 "while I'm here" improvements
185
+ - 无 bundled refactoring
186
+
187
+ 3. **Verify Fix**
188
+ - Test 现在 pass?
189
+ - 无 other tests broken?
190
+ - Issue 真的 resolved?
191
+
192
+ 4. **If Fix Doesn't Work**
193
+ - STOP
194
+ - Count:已尝试多少 fixes?
195
+ - 若 < 3:Return to Phase 1,用 new information 再分析
196
+ - **若 ≥ 3:STOP 并质疑 architecture(见下方 step 5)**
197
+ - DON'T 在未做 architectural discussion 前尝试 Fix #4
198
+
199
+ 5. **If 3+ Fixes Failed: Question Architecture**
200
+
201
+ **表明 architectural problem 的 pattern:**
202
+ - 每个 fix 在不同位置 reveal 新的 shared state/coupling/problem
203
+ - Fixes 需要 "massive refactoring" 才能实现
204
+ - 每个 fix 在其他地方制造新 symptoms
205
+
206
+ **STOP 并质疑 fundamentals:**
207
+ - 此 pattern fundamentally sound 吗?
208
+ - 是否 "sticking with it through sheer inertia"?
209
+ - 应 refactor architecture 还是继续 fix symptoms?
210
+
211
+ **Discuss with your human partner before attempting more fixes**
212
+
213
+ This is NOT a failed hypothesis — this is a wrong architecture.
214
+
215
+ ## Red Flags - STOP and Follow Process
216
+
217
+ 若发现自己想:
218
+ - "Quick fix for now, investigate later"
219
+ - "Just try changing X and see if it works"
220
+ - "Add multiple changes, run tests"
221
+ - "Skip the test, I'll manually verify"
222
+ - "It's probably X, let me fix that"
223
+ - "I don't fully understand but this might work"
224
+ - "Pattern says X but I'll adapt it differently"
225
+ - "Here are the main problems: [lists fixes without investigation]"
226
+ - Proposing solutions before tracing data flow
227
+ - **"One more fix attempt" (when already tried 2+)**
228
+ - **Each fix reveals new problem in different place**
229
+
230
+ **ALL of these mean: STOP. Return to Phase 1.**
231
+
232
+ **If 3+ fixes failed:** Question the architecture (see Phase 4.5)
233
+
234
+ ## your human partner's Signals You're Doing It Wrong
235
+
236
+ **Watch for these redirections:**
237
+ - "Is that not happening?" — You assumed without verifying
238
+ - "Will it show us...?" — You should have added evidence gathering
239
+ - "Stop guessing" — You're proposing fixes without understanding
240
+ - "Ultrathink this" — Question fundamentals, not just symptoms
241
+ - "We're stuck?" (frustrated) — Your approach isn't working
242
+
243
+ **When you see these:** STOP. Return to Phase 1.
244
+
245
+ ## Common Rationalizations
246
+
247
+ | Excuse | Reality |
248
+ |--------|---------|
249
+ | "Issue is simple, don't need process" | Simple issues 也有 root causes。Process 对 simple bugs 很快。 |
250
+ | "Emergency, no time for process" | Systematic debugging 比 guess-and-check thrashing 更快。 |
251
+ | "Just try this first, then investigate" | First fix 定模式。从一开始就做对。 |
252
+ | "I'll write test after confirming fix works" | Untested fixes 不 stick。Test first 证明它。 |
253
+ | "Multiple fixes at once saves time" | 无法 isolate what worked。制造新 bugs。 |
254
+ | "Reference too long, I'll adapt the pattern" | Partial understanding 保证 bugs。Complete 阅读。 |
255
+ | "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause。 |
256
+ | "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem。质疑 pattern,不要再 fix。 |
257
+
258
+ ## Quick Reference
259
+
260
+ | Phase | Key Activities | Success Criteria |
261
+ |-------|---------------|------------------|
262
+ | **1. Root Cause** | Read errors, reproduce, check changes, gather evidence | Understand WHAT and WHY |
263
+ | **2. Pattern** | Find working examples, compare | Identify differences |
264
+ | **3. Hypothesis** | Form theory, test minimally | Confirmed or new hypothesis |
265
+ | **4. Implementation** | Create test, fix, verify | Bug resolved, tests pass |
266
+
267
+ ## When Process Reveals "No Root Cause"
268
+
269
+ 若 systematic investigation 表明 issue truly environmental、timing-dependent 或 external:
270
+
271
+ 1. You've completed the process
272
+ 2. Document what you investigated
273
+ 3. Implement appropriate handling (retry, timeout, error message)
274
+ 4. Add monitoring/logging for future investigation
275
+
276
+ **But:** 95% 的 "no root cause" cases 是 incomplete investigation。
277
+
278
+ ## Supporting Techniques
279
+
280
+ 本目录中属于 systematic debugging 的技术:
281
+
282
+ - **`root-cause-tracing.md`** — Trace bugs backward through call stack 找 original trigger
283
+ - **`defense-in-depth.md`** — 找到 root cause 后在 multiple layers 加 validation
284
+ - **`condition-based-waiting.md`** — 用 condition polling 替代 arbitrary timeouts
285
+
286
+ **Related principles:**
287
+ - **RED-GREEN-REFACTOR**(见 `ai_workspace/loop-agent/harness-methodology-tdd.md`)— 用于 creating failing test case(Phase 4, Step 1)
288
+ - **Verification discipline** — 宣称 success 前 verify fix worked。Run verification command,读 output,THEN claim result。
289
+
290
+ ## Real-World Impact
291
+
292
+ 来自 debugging sessions:
293
+ - Systematic approach:15-30 分钟 fix
294
+ - Random fixes approach:2-3 小时 thrashing
295
+ - First-time fix rate:95% vs 40%
296
+ - New bugs introduced:Near zero vs common