@tea-agent/loop-agent 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +135 -133
- package/CHANGELOG.md +88 -63
- package/README.md +171 -168
- package/bin/agent-worker.js +22 -0
- package/bin/loop-agent.js +21 -21
- package/dist/commands/init.js +457 -457
- package/dist/commands/loop-benchmark.js +11 -11
- package/dist/commands/pi-reuse-benchmark.js +16 -16
- package/dist/executors/cursor-executor.js +1 -1
- package/dist/executors/dag-pi-executor.js +8 -1
- package/dist/task/runtime.js +27 -27
- package/dist/worker/cli.js +119 -0
- package/dist/worker/loop-agent/command-result.js +1 -0
- package/dist/worker/loop-agent/loop-agent-client.js +105 -0
- package/dist/worker/loop-agent/parse-json.js +14 -0
- package/dist/worker/materialize/harness-task-materializer.js +157 -0
- package/dist/worker/pool/failure-routing.js +98 -0
- package/dist/worker/pool/run-store.js +117 -0
- package/dist/worker/pool/types.js +1 -0
- package/dist/worker/preflight.js +108 -0
- package/dist/worker/profile-mapping.js +76 -0
- package/dist/worker/progress-reporter.js +81 -0
- package/dist/worker/report/morning-report.js +69 -0
- package/dist/worker/repos/repo-resolver.js +23 -0
- package/dist/worker/run-task/run-task.js +359 -0
- package/dist/worker/runner/run-ready.js +216 -0
- package/dist/worker/task-graph/acceptance-schema.js +25 -0
- package/dist/worker/task-graph/ready-queue.js +23 -0
- package/dist/worker/task-graph/task-graph-schema.js +28 -0
- package/dist/worker/task-graph/types.js +1 -0
- package/dist/worker/task-graph/validate.js +188 -0
- package/dist/worker/task-spec/complexity-mapping.js +8 -0
- package/dist/worker/task-spec/schema.js +116 -0
- package/dist/worker/task-spec/types.js +1 -0
- package/dist/worker/task-spec/validate.js +352 -0
- package/dist/workflows/dag/canvas-observer.js +275 -275
- package/docs/README.md +65 -61
- package/docs/agent-dag-recovery-playbook.md +184 -184
- package/docs/agent-dag-runner.md +42 -42
- package/docs/architecture/runtime-boundaries.md +147 -147
- package/docs/cursor-executor-usage.md +25 -25
- package/docs/decisions/README.md +3 -3
- package/docs/design/README.md +36 -36
- package/docs/development-principles.md +73 -71
- package/docs/dynamic-workflow-dag-engine-roadmap.md +1749 -1749
- package/docs/exec-plans/README.md +6 -6
- package/docs/exec-plans/active/README.md +7 -7
- package/docs/exec-plans/completed/README.md +19 -11
- package/docs/feature-workflow.md +186 -186
- package/docs/harness-methodology-debugging.md +153 -153
- package/docs/harness-methodology-tdd.md +130 -130
- package/docs/harness-methodology-verification.md +27 -27
- package/docs/loop-agent-harness.md +42 -42
- package/docs/production-readiness.md +96 -96
- package/docs/progress/README.md +3 -3
- package/docs/reports/README.md +5 -5
- package/docs/skills/README.md +6 -6
- package/docs/skills/vetted-skill-registry.md +22 -22
- package/docs/templates/adr.md +60 -60
- package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
- package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
- package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
- package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
- package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
- package/docs/templates/agent-dag-report.schema.json +454 -454
- package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
- package/docs/templates/agent-dag.base.json +195 -195
- package/docs/templates/agent-dag.final-verification.json +190 -190
- package/docs/templates/agent-dag.schema.json +316 -316
- package/docs/templates/agent-dag.supervised-implementation.json +500 -500
- package/docs/templates/exec-plan.md +64 -64
- package/docs/templates/feature-spec.md +53 -53
- package/docs/templates/hybrid-dag.json +193 -193
- package/docs/templates/production-readiness-checklist.md +57 -57
- package/docs/templates/progress-log.md +17 -17
- package/docs/templates/project-start-checklist.md +9 -9
- package/docs/templates/qa-report.md +48 -48
- package/docs/templates/sprint-contract.md +29 -29
- package/docs/verification-matrix.md +41 -41
- package/examples/decision-gate-agent-dag.json +123 -123
- package/examples/example-dag.json +51 -51
- package/examples/hybrid-loop-agent-dag.json +194 -194
- package/harness.json +89 -89
- package/package.json +60 -58
- package/skills/ai-engineering-context/SKILL.md +48 -48
- package/skills/code-review-core/SKILL.md +20 -20
- package/skills/codebase-scout/SKILL.md +19 -19
- package/skills/loop-agent/SKILL.md +147 -145
- package/skills/loop-agent/references/README.md +67 -67
- package/skills/loop-agent/references/command-reference.md +368 -340
- package/skills/loop-agent/references/harness-policy.md +259 -258
- package/skills/loop-agent/references/hybrid-dag.md +216 -216
- package/skills/loop-agent/references/learned/README.md +21 -21
- package/skills/loop-agent/references/long-running-loop.md +59 -59
- package/skills/loop-agent/references/model-routing.md +36 -36
- package/skills/loop-agent/references/multi-worktree.md +54 -54
- package/skills/loop-agent/references/one-shot-runs.md +85 -85
- package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
- package/skills/loop-agent/references/pi-prompt.md +23 -23
- package/skills/loop-agent/references/pi-subagent-assisted-mode.md +81 -81
- package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
- package/skills/loop-agent/references/task-workflow.md +84 -84
- package/skills/loop-agent/references/verification-and-failure-handling.md +128 -128
- package/skills/requesting-code-review/SKILL.md +101 -101
- package/skills/requesting-code-review/code-reviewer.md +168 -168
- package/skills/systematic-debugging/CREATION-LOG.md +119 -119
- package/skills/systematic-debugging/SKILL.md +296 -296
- package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
- package/skills/systematic-debugging/condition-based-waiting.md +115 -115
- package/skills/systematic-debugging/defense-in-depth.md +122 -122
- package/skills/systematic-debugging/find-polluter.sh +63 -63
- package/skills/systematic-debugging/root-cause-tracing.md +169 -169
- package/skills/systematic-debugging/test-academic.md +14 -14
- package/skills/systematic-debugging/test-pressure-1.md +58 -58
- package/skills/systematic-debugging/test-pressure-2.md +68 -68
- package/skills/systematic-debugging/test-pressure-3.md +69 -69
- package/skills/test-driven-development/SKILL.md +20 -20
- package/skills/verification-before-completion/SKILL.md +154 -154
- package/skills/webapp-testing/SKILL.md +19 -19
|
@@ -1,296 +1,296 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: systematic-debugging
|
|
3
|
-
description: 遇到任何 bug、test failure 或 unexpected behavior 时使用,且在提出 fixes 之前
|
|
4
|
-
---
|
|
5
|
-
|
|
6
|
-
# Systematic Debugging
|
|
7
|
-
|
|
8
|
-
## Overview
|
|
9
|
-
|
|
10
|
-
Random fixes 浪费时间并制造新 bug。Quick patches 掩盖 underlying issues。
|
|
11
|
-
|
|
12
|
-
**Core principle:** ALWAYS 在尝试 fixes 之前找到 root cause。Symptom fixes 是 failure。
|
|
13
|
-
|
|
14
|
-
**违反本流程字面即违反 debugging 精神。**
|
|
15
|
-
|
|
16
|
-
## The Iron Law
|
|
17
|
-
|
|
18
|
-
```
|
|
19
|
-
NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST
|
|
20
|
-
```
|
|
21
|
-
|
|
22
|
-
若尚未完成 Phase 1,不得提出 fixes。
|
|
23
|
-
|
|
24
|
-
## When to Use
|
|
25
|
-
|
|
26
|
-
用于 ANY technical issue:
|
|
27
|
-
- Test failures
|
|
28
|
-
- Production bugs
|
|
29
|
-
- Unexpected behavior
|
|
30
|
-
- Performance problems
|
|
31
|
-
- Build failures
|
|
32
|
-
- Integration issues
|
|
33
|
-
|
|
34
|
-
**ESPECIALLY 在以下情况使用:**
|
|
35
|
-
- 时间压力下(emergencies 使 guessing 诱人)
|
|
36
|
-
- "Just one quick fix" 看起来 obvious
|
|
37
|
-
- 已尝试 multiple fixes
|
|
38
|
-
- Previous fix 无效
|
|
39
|
-
- 未完全理解 issue
|
|
40
|
-
|
|
41
|
-
**Don't skip when:**
|
|
42
|
-
- Issue 看起来 simple(simple bugs 也有 root causes)
|
|
43
|
-
- 赶时间(rushing 保证 rework)
|
|
44
|
-
- Manager 要求 NOW 修好(systematic 比 thrashing 更快)
|
|
45
|
-
|
|
46
|
-
## The Four Phases
|
|
47
|
-
|
|
48
|
-
进入下一阶段前 MUST 完成每一 phase。
|
|
49
|
-
|
|
50
|
-
### Phase 1: Root Cause Investigation
|
|
51
|
-
|
|
52
|
-
**在尝试 ANY fix 之前:**
|
|
53
|
-
|
|
54
|
-
1. **Read Error Messages Carefully**
|
|
55
|
-
- 不要跳过 errors 或 warnings
|
|
56
|
-
- 它们常含 exact solution
|
|
57
|
-
- 完整阅读 stack traces
|
|
58
|
-
- 记下 line numbers、file paths、error codes
|
|
59
|
-
|
|
60
|
-
2. **Reproduce Consistently**
|
|
61
|
-
- 能否可靠触发?
|
|
62
|
-
- Exact steps 是什么?
|
|
63
|
-
- 是否每次都发生?
|
|
64
|
-
- 若不可 reproduce → 收集更多 data,不要 guess
|
|
65
|
-
|
|
66
|
-
3. **Check Recent Changes**
|
|
67
|
-
- 什么变更可能导致此问题?
|
|
68
|
-
- Git diff、recent commits
|
|
69
|
-
- New dependencies、config changes
|
|
70
|
-
- Environmental differences
|
|
71
|
-
|
|
72
|
-
4. **Gather Evidence in Multi-Component Systems**
|
|
73
|
-
|
|
74
|
-
**WHEN system 有多个 components(CI → build → signing,API → service → database):**
|
|
75
|
-
|
|
76
|
-
**BEFORE proposing fixes,添加 diagnostic instrumentation:**
|
|
77
|
-
```
|
|
78
|
-
For EACH component boundary:
|
|
79
|
-
- Log what data enters component
|
|
80
|
-
- Log what data exits component
|
|
81
|
-
- Verify environment/config propagation
|
|
82
|
-
- Check state at each layer
|
|
83
|
-
|
|
84
|
-
Run once to gather evidence showing WHERE it breaks
|
|
85
|
-
THEN analyze evidence to identify failing component
|
|
86
|
-
THEN investigate that specific component
|
|
87
|
-
```
|
|
88
|
-
|
|
89
|
-
**Example (multi-layer system):**
|
|
90
|
-
```bash
|
|
91
|
-
# Layer 1: Workflow
|
|
92
|
-
echo "=== Secrets available in workflow: ==="
|
|
93
|
-
echo "IDENTITY: ${IDENTITY:+SET}${IDENTITY:-UNSET}"
|
|
94
|
-
|
|
95
|
-
# Layer 2: Build script
|
|
96
|
-
echo "=== Env vars in build script: ==="
|
|
97
|
-
env | grep IDENTITY || echo "IDENTITY not in environment"
|
|
98
|
-
|
|
99
|
-
# Layer 3: Signing script
|
|
100
|
-
echo "=== Keychain state: ==="
|
|
101
|
-
security list-keychains
|
|
102
|
-
security find-identity -v
|
|
103
|
-
|
|
104
|
-
# Layer 4: Actual signing
|
|
105
|
-
codesign --sign "$IDENTITY" --verbose=4 "$APP"
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
**This reveals:** Which layer fails (secrets → workflow ✓, workflow → build ✗)
|
|
109
|
-
|
|
110
|
-
5. **Trace Data Flow**
|
|
111
|
-
|
|
112
|
-
**WHEN error 在 call stack 深处:**
|
|
113
|
-
|
|
114
|
-
完整 backward tracing 见本目录 `root-cause-tracing.md`。
|
|
115
|
-
|
|
116
|
-
**Quick version:**
|
|
117
|
-
- Bad value 从哪 originate?
|
|
118
|
-
- 谁用 bad value 调用了 this?
|
|
119
|
-
- 一直向上 trace 直到 source
|
|
120
|
-
- 在 source 修复,而非 symptom
|
|
121
|
-
|
|
122
|
-
### Phase 2: Pattern Analysis
|
|
123
|
-
|
|
124
|
-
**Fix 前先找 pattern:**
|
|
125
|
-
|
|
126
|
-
1. **Find Working Examples**
|
|
127
|
-
- 在同 codebase 找 similar working code
|
|
128
|
-
- 什么能 work、什么 broken?
|
|
129
|
-
|
|
130
|
-
2. **Compare Against References**
|
|
131
|
-
- 若实现 pattern,COMPLETE 阅读 reference implementation
|
|
132
|
-
- 不要 skim — 读每一行
|
|
133
|
-
- 应用前 fully 理解 pattern
|
|
134
|
-
|
|
135
|
-
3. **Identify Differences**
|
|
136
|
-
- Working 与 broken 有何不同?
|
|
137
|
-
- 列出 every difference,再小也要列
|
|
138
|
-
- 不要假设 "that can't matter"
|
|
139
|
-
|
|
140
|
-
4. **Understand Dependencies**
|
|
141
|
-
- 还需要哪些 other components?
|
|
142
|
-
- 哪些 settings、config、environment?
|
|
143
|
-
- 它作哪些 assumptions?
|
|
144
|
-
|
|
145
|
-
### Phase 3: Hypothesis and Testing
|
|
146
|
-
|
|
147
|
-
**Scientific method:**
|
|
148
|
-
|
|
149
|
-
1. **Form Single Hypothesis**
|
|
150
|
-
- 清楚陈述:"I think X is the root cause because Y"
|
|
151
|
-
- 写下来
|
|
152
|
-
- 要 specific,不要 vague
|
|
153
|
-
|
|
154
|
-
2. **Test Minimally**
|
|
155
|
-
- 做 SMALLEST possible change 以 test hypothesis
|
|
156
|
-
- One variable at a time
|
|
157
|
-
- 不要一次 fix multiple things
|
|
158
|
-
|
|
159
|
-
3. **Verify Before Continuing**
|
|
160
|
-
- 有效?Yes → Phase 4
|
|
161
|
-
- 无效?Form NEW hypothesis
|
|
162
|
-
- DON'T 在其上叠加更多 fixes
|
|
163
|
-
|
|
164
|
-
4. **When You Don't Know**
|
|
165
|
-
- 说 "I don't understand X"
|
|
166
|
-
- 不要假装知道
|
|
167
|
-
- Ask for help
|
|
168
|
-
- Research more
|
|
169
|
-
|
|
170
|
-
### Phase 4: Implementation
|
|
171
|
-
|
|
172
|
-
**Fix root cause,不是 symptom:**
|
|
173
|
-
|
|
174
|
-
1. **Create Failing Test Case**
|
|
175
|
-
- Simplest possible reproduction
|
|
176
|
-
- 可能的话用 automated test
|
|
177
|
-
- 无 framework 时用 one-off test script
|
|
178
|
-
- MUST 在 fix 之前有
|
|
179
|
-
- 遵循 RED-GREEN-REFACTOR:写 failing test,看它 fail,再 fix
|
|
180
|
-
|
|
181
|
-
2. **Implement Single Fix**
|
|
182
|
-
- 针对已识别的 root cause
|
|
183
|
-
- ONE change at a time
|
|
184
|
-
- 无 "while I'm here" improvements
|
|
185
|
-
- 无 bundled refactoring
|
|
186
|
-
|
|
187
|
-
3. **Verify Fix**
|
|
188
|
-
- Test 现在 pass?
|
|
189
|
-
- 无 other tests broken?
|
|
190
|
-
- Issue 真的 resolved?
|
|
191
|
-
|
|
192
|
-
4. **If Fix Doesn't Work**
|
|
193
|
-
- STOP
|
|
194
|
-
- Count:已尝试多少 fixes?
|
|
195
|
-
- 若 < 3:Return to Phase 1,用 new information 再分析
|
|
196
|
-
- **若 ≥ 3:STOP 并质疑 architecture(见下方 step 5)**
|
|
197
|
-
- DON'T 在未做 architectural discussion 前尝试 Fix #4
|
|
198
|
-
|
|
199
|
-
5. **If 3+ Fixes Failed: Question Architecture**
|
|
200
|
-
|
|
201
|
-
**表明 architectural problem 的 pattern:**
|
|
202
|
-
- 每个 fix 在不同位置 reveal 新的 shared state/coupling/problem
|
|
203
|
-
- Fixes 需要 "massive refactoring" 才能实现
|
|
204
|
-
- 每个 fix 在其他地方制造新 symptoms
|
|
205
|
-
|
|
206
|
-
**STOP 并质疑 fundamentals:**
|
|
207
|
-
- 此 pattern fundamentally sound 吗?
|
|
208
|
-
- 是否 "sticking with it through sheer inertia"?
|
|
209
|
-
- 应 refactor architecture 还是继续 fix symptoms?
|
|
210
|
-
|
|
211
|
-
**Discuss with your human partner before attempting more fixes**
|
|
212
|
-
|
|
213
|
-
This is NOT a failed hypothesis — this is a wrong architecture.
|
|
214
|
-
|
|
215
|
-
## Red Flags - STOP and Follow Process
|
|
216
|
-
|
|
217
|
-
若发现自己想:
|
|
218
|
-
- "Quick fix for now, investigate later"
|
|
219
|
-
- "Just try changing X and see if it works"
|
|
220
|
-
- "Add multiple changes, run tests"
|
|
221
|
-
- "Skip the test, I'll manually verify"
|
|
222
|
-
- "It's probably X, let me fix that"
|
|
223
|
-
- "I don't fully understand but this might work"
|
|
224
|
-
- "Pattern says X but I'll adapt it differently"
|
|
225
|
-
- "Here are the main problems: [lists fixes without investigation]"
|
|
226
|
-
- Proposing solutions before tracing data flow
|
|
227
|
-
- **"One more fix attempt" (when already tried 2+)**
|
|
228
|
-
- **Each fix reveals new problem in different place**
|
|
229
|
-
|
|
230
|
-
**ALL of these mean: STOP. Return to Phase 1.**
|
|
231
|
-
|
|
232
|
-
**If 3+ fixes failed:** Question the architecture (see Phase 4.5)
|
|
233
|
-
|
|
234
|
-
## your human partner's Signals You're Doing It Wrong
|
|
235
|
-
|
|
236
|
-
**Watch for these redirections:**
|
|
237
|
-
- "Is that not happening?" — You assumed without verifying
|
|
238
|
-
- "Will it show us...?" — You should have added evidence gathering
|
|
239
|
-
- "Stop guessing" — You're proposing fixes without understanding
|
|
240
|
-
- "Ultrathink this" — Question fundamentals, not just symptoms
|
|
241
|
-
- "We're stuck?" (frustrated) — Your approach isn't working
|
|
242
|
-
|
|
243
|
-
**When you see these:** STOP. Return to Phase 1.
|
|
244
|
-
|
|
245
|
-
## Common Rationalizations
|
|
246
|
-
|
|
247
|
-
| Excuse | Reality |
|
|
248
|
-
|--------|---------|
|
|
249
|
-
| "Issue is simple, don't need process" | Simple issues 也有 root causes。Process 对 simple bugs 很快。 |
|
|
250
|
-
| "Emergency, no time for process" | Systematic debugging 比 guess-and-check thrashing 更快。 |
|
|
251
|
-
| "Just try this first, then investigate" | First fix 定模式。从一开始就做对。 |
|
|
252
|
-
| "I'll write test after confirming fix works" | Untested fixes 不 stick。Test first 证明它。 |
|
|
253
|
-
| "Multiple fixes at once saves time" | 无法 isolate what worked。制造新 bugs。 |
|
|
254
|
-
| "Reference too long, I'll adapt the pattern" | Partial understanding 保证 bugs。Complete 阅读。 |
|
|
255
|
-
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause。 |
|
|
256
|
-
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem。质疑 pattern,不要再 fix。 |
|
|
257
|
-
|
|
258
|
-
## Quick Reference
|
|
259
|
-
|
|
260
|
-
| Phase | Key Activities | Success Criteria |
|
|
261
|
-
|-------|---------------|------------------|
|
|
262
|
-
| **1. Root Cause** | Read errors, reproduce, check changes, gather evidence | Understand WHAT and WHY |
|
|
263
|
-
| **2. Pattern** | Find working examples, compare | Identify differences |
|
|
264
|
-
| **3. Hypothesis** | Form theory, test minimally | Confirmed or new hypothesis |
|
|
265
|
-
| **4. Implementation** | Create test, fix, verify | Bug resolved, tests pass |
|
|
266
|
-
|
|
267
|
-
## When Process Reveals "No Root Cause"
|
|
268
|
-
|
|
269
|
-
若 systematic investigation 表明 issue truly environmental、timing-dependent 或 external:
|
|
270
|
-
|
|
271
|
-
1. You've completed the process
|
|
272
|
-
2. Document what you investigated
|
|
273
|
-
3. Implement appropriate handling (retry, timeout, error message)
|
|
274
|
-
4. Add monitoring/logging for future investigation
|
|
275
|
-
|
|
276
|
-
**But:** 95% 的 "no root cause" cases 是 incomplete investigation。
|
|
277
|
-
|
|
278
|
-
## Supporting Techniques
|
|
279
|
-
|
|
280
|
-
本目录中属于 systematic debugging 的技术:
|
|
281
|
-
|
|
282
|
-
- **`root-cause-tracing.md`** — Trace bugs backward through call stack 找 original trigger
|
|
283
|
-
- **`defense-in-depth.md`** — 找到 root cause 后在 multiple layers 加 validation
|
|
284
|
-
- **`condition-based-waiting.md`** — 用 condition polling 替代 arbitrary timeouts
|
|
285
|
-
|
|
286
|
-
**Related principles:**
|
|
287
|
-
- **RED-GREEN-REFACTOR**(见 `docs/harness-methodology-tdd.md`)— 用于 creating failing test case(Phase 4, Step 1)
|
|
288
|
-
- **Verification discipline** — 宣称 success 前 verify fix worked。Run verification command,读 output,THEN claim result。
|
|
289
|
-
|
|
290
|
-
## Real-World Impact
|
|
291
|
-
|
|
292
|
-
来自 debugging sessions:
|
|
293
|
-
- Systematic approach:15-30 分钟 fix
|
|
294
|
-
- Random fixes approach:2-3 小时 thrashing
|
|
295
|
-
- First-time fix rate:95% vs 40%
|
|
296
|
-
- New bugs introduced:Near zero vs common
|
|
1
|
+
---
|
|
2
|
+
name: systematic-debugging
|
|
3
|
+
description: 遇到任何 bug、test failure 或 unexpected behavior 时使用,且在提出 fixes 之前
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Systematic Debugging
|
|
7
|
+
|
|
8
|
+
## Overview
|
|
9
|
+
|
|
10
|
+
Random fixes 浪费时间并制造新 bug。Quick patches 掩盖 underlying issues。
|
|
11
|
+
|
|
12
|
+
**Core principle:** ALWAYS 在尝试 fixes 之前找到 root cause。Symptom fixes 是 failure。
|
|
13
|
+
|
|
14
|
+
**违反本流程字面即违反 debugging 精神。**
|
|
15
|
+
|
|
16
|
+
## The Iron Law
|
|
17
|
+
|
|
18
|
+
```
|
|
19
|
+
NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
若尚未完成 Phase 1,不得提出 fixes。
|
|
23
|
+
|
|
24
|
+
## When to Use
|
|
25
|
+
|
|
26
|
+
用于 ANY technical issue:
|
|
27
|
+
- Test failures
|
|
28
|
+
- Production bugs
|
|
29
|
+
- Unexpected behavior
|
|
30
|
+
- Performance problems
|
|
31
|
+
- Build failures
|
|
32
|
+
- Integration issues
|
|
33
|
+
|
|
34
|
+
**ESPECIALLY 在以下情况使用:**
|
|
35
|
+
- 时间压力下(emergencies 使 guessing 诱人)
|
|
36
|
+
- "Just one quick fix" 看起来 obvious
|
|
37
|
+
- 已尝试 multiple fixes
|
|
38
|
+
- Previous fix 无效
|
|
39
|
+
- 未完全理解 issue
|
|
40
|
+
|
|
41
|
+
**Don't skip when:**
|
|
42
|
+
- Issue 看起来 simple(simple bugs 也有 root causes)
|
|
43
|
+
- 赶时间(rushing 保证 rework)
|
|
44
|
+
- Manager 要求 NOW 修好(systematic 比 thrashing 更快)
|
|
45
|
+
|
|
46
|
+
## The Four Phases
|
|
47
|
+
|
|
48
|
+
进入下一阶段前 MUST 完成每一 phase。
|
|
49
|
+
|
|
50
|
+
### Phase 1: Root Cause Investigation
|
|
51
|
+
|
|
52
|
+
**在尝试 ANY fix 之前:**
|
|
53
|
+
|
|
54
|
+
1. **Read Error Messages Carefully**
|
|
55
|
+
- 不要跳过 errors 或 warnings
|
|
56
|
+
- 它们常含 exact solution
|
|
57
|
+
- 完整阅读 stack traces
|
|
58
|
+
- 记下 line numbers、file paths、error codes
|
|
59
|
+
|
|
60
|
+
2. **Reproduce Consistently**
|
|
61
|
+
- 能否可靠触发?
|
|
62
|
+
- Exact steps 是什么?
|
|
63
|
+
- 是否每次都发生?
|
|
64
|
+
- 若不可 reproduce → 收集更多 data,不要 guess
|
|
65
|
+
|
|
66
|
+
3. **Check Recent Changes**
|
|
67
|
+
- 什么变更可能导致此问题?
|
|
68
|
+
- Git diff、recent commits
|
|
69
|
+
- New dependencies、config changes
|
|
70
|
+
- Environmental differences
|
|
71
|
+
|
|
72
|
+
4. **Gather Evidence in Multi-Component Systems**
|
|
73
|
+
|
|
74
|
+
**WHEN system 有多个 components(CI → build → signing,API → service → database):**
|
|
75
|
+
|
|
76
|
+
**BEFORE proposing fixes,添加 diagnostic instrumentation:**
|
|
77
|
+
```
|
|
78
|
+
For EACH component boundary:
|
|
79
|
+
- Log what data enters component
|
|
80
|
+
- Log what data exits component
|
|
81
|
+
- Verify environment/config propagation
|
|
82
|
+
- Check state at each layer
|
|
83
|
+
|
|
84
|
+
Run once to gather evidence showing WHERE it breaks
|
|
85
|
+
THEN analyze evidence to identify failing component
|
|
86
|
+
THEN investigate that specific component
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
**Example (multi-layer system):**
|
|
90
|
+
```bash
|
|
91
|
+
# Layer 1: Workflow
|
|
92
|
+
echo "=== Secrets available in workflow: ==="
|
|
93
|
+
echo "IDENTITY: ${IDENTITY:+SET}${IDENTITY:-UNSET}"
|
|
94
|
+
|
|
95
|
+
# Layer 2: Build script
|
|
96
|
+
echo "=== Env vars in build script: ==="
|
|
97
|
+
env | grep IDENTITY || echo "IDENTITY not in environment"
|
|
98
|
+
|
|
99
|
+
# Layer 3: Signing script
|
|
100
|
+
echo "=== Keychain state: ==="
|
|
101
|
+
security list-keychains
|
|
102
|
+
security find-identity -v
|
|
103
|
+
|
|
104
|
+
# Layer 4: Actual signing
|
|
105
|
+
codesign --sign "$IDENTITY" --verbose=4 "$APP"
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
**This reveals:** Which layer fails (secrets → workflow ✓, workflow → build ✗)
|
|
109
|
+
|
|
110
|
+
5. **Trace Data Flow**
|
|
111
|
+
|
|
112
|
+
**WHEN error 在 call stack 深处:**
|
|
113
|
+
|
|
114
|
+
完整 backward tracing 见本目录 `root-cause-tracing.md`。
|
|
115
|
+
|
|
116
|
+
**Quick version:**
|
|
117
|
+
- Bad value 从哪 originate?
|
|
118
|
+
- 谁用 bad value 调用了 this?
|
|
119
|
+
- 一直向上 trace 直到 source
|
|
120
|
+
- 在 source 修复,而非 symptom
|
|
121
|
+
|
|
122
|
+
### Phase 2: Pattern Analysis
|
|
123
|
+
|
|
124
|
+
**Fix 前先找 pattern:**
|
|
125
|
+
|
|
126
|
+
1. **Find Working Examples**
|
|
127
|
+
- 在同 codebase 找 similar working code
|
|
128
|
+
- 什么能 work、什么 broken?
|
|
129
|
+
|
|
130
|
+
2. **Compare Against References**
|
|
131
|
+
- 若实现 pattern,COMPLETE 阅读 reference implementation
|
|
132
|
+
- 不要 skim — 读每一行
|
|
133
|
+
- 应用前 fully 理解 pattern
|
|
134
|
+
|
|
135
|
+
3. **Identify Differences**
|
|
136
|
+
- Working 与 broken 有何不同?
|
|
137
|
+
- 列出 every difference,再小也要列
|
|
138
|
+
- 不要假设 "that can't matter"
|
|
139
|
+
|
|
140
|
+
4. **Understand Dependencies**
|
|
141
|
+
- 还需要哪些 other components?
|
|
142
|
+
- 哪些 settings、config、environment?
|
|
143
|
+
- 它作哪些 assumptions?
|
|
144
|
+
|
|
145
|
+
### Phase 3: Hypothesis and Testing
|
|
146
|
+
|
|
147
|
+
**Scientific method:**
|
|
148
|
+
|
|
149
|
+
1. **Form Single Hypothesis**
|
|
150
|
+
- 清楚陈述:"I think X is the root cause because Y"
|
|
151
|
+
- 写下来
|
|
152
|
+
- 要 specific,不要 vague
|
|
153
|
+
|
|
154
|
+
2. **Test Minimally**
|
|
155
|
+
- 做 SMALLEST possible change 以 test hypothesis
|
|
156
|
+
- One variable at a time
|
|
157
|
+
- 不要一次 fix multiple things
|
|
158
|
+
|
|
159
|
+
3. **Verify Before Continuing**
|
|
160
|
+
- 有效?Yes → Phase 4
|
|
161
|
+
- 无效?Form NEW hypothesis
|
|
162
|
+
- DON'T 在其上叠加更多 fixes
|
|
163
|
+
|
|
164
|
+
4. **When You Don't Know**
|
|
165
|
+
- 说 "I don't understand X"
|
|
166
|
+
- 不要假装知道
|
|
167
|
+
- Ask for help
|
|
168
|
+
- Research more
|
|
169
|
+
|
|
170
|
+
### Phase 4: Implementation
|
|
171
|
+
|
|
172
|
+
**Fix root cause,不是 symptom:**
|
|
173
|
+
|
|
174
|
+
1. **Create Failing Test Case**
|
|
175
|
+
- Simplest possible reproduction
|
|
176
|
+
- 可能的话用 automated test
|
|
177
|
+
- 无 framework 时用 one-off test script
|
|
178
|
+
- MUST 在 fix 之前有
|
|
179
|
+
- 遵循 RED-GREEN-REFACTOR:写 failing test,看它 fail,再 fix
|
|
180
|
+
|
|
181
|
+
2. **Implement Single Fix**
|
|
182
|
+
- 针对已识别的 root cause
|
|
183
|
+
- ONE change at a time
|
|
184
|
+
- 无 "while I'm here" improvements
|
|
185
|
+
- 无 bundled refactoring
|
|
186
|
+
|
|
187
|
+
3. **Verify Fix**
|
|
188
|
+
- Test 现在 pass?
|
|
189
|
+
- 无 other tests broken?
|
|
190
|
+
- Issue 真的 resolved?
|
|
191
|
+
|
|
192
|
+
4. **If Fix Doesn't Work**
|
|
193
|
+
- STOP
|
|
194
|
+
- Count:已尝试多少 fixes?
|
|
195
|
+
- 若 < 3:Return to Phase 1,用 new information 再分析
|
|
196
|
+
- **若 ≥ 3:STOP 并质疑 architecture(见下方 step 5)**
|
|
197
|
+
- DON'T 在未做 architectural discussion 前尝试 Fix #4
|
|
198
|
+
|
|
199
|
+
5. **If 3+ Fixes Failed: Question Architecture**
|
|
200
|
+
|
|
201
|
+
**表明 architectural problem 的 pattern:**
|
|
202
|
+
- 每个 fix 在不同位置 reveal 新的 shared state/coupling/problem
|
|
203
|
+
- Fixes 需要 "massive refactoring" 才能实现
|
|
204
|
+
- 每个 fix 在其他地方制造新 symptoms
|
|
205
|
+
|
|
206
|
+
**STOP 并质疑 fundamentals:**
|
|
207
|
+
- 此 pattern fundamentally sound 吗?
|
|
208
|
+
- 是否 "sticking with it through sheer inertia"?
|
|
209
|
+
- 应 refactor architecture 还是继续 fix symptoms?
|
|
210
|
+
|
|
211
|
+
**Discuss with your human partner before attempting more fixes**
|
|
212
|
+
|
|
213
|
+
This is NOT a failed hypothesis — this is a wrong architecture.
|
|
214
|
+
|
|
215
|
+
## Red Flags - STOP and Follow Process
|
|
216
|
+
|
|
217
|
+
若发现自己想:
|
|
218
|
+
- "Quick fix for now, investigate later"
|
|
219
|
+
- "Just try changing X and see if it works"
|
|
220
|
+
- "Add multiple changes, run tests"
|
|
221
|
+
- "Skip the test, I'll manually verify"
|
|
222
|
+
- "It's probably X, let me fix that"
|
|
223
|
+
- "I don't fully understand but this might work"
|
|
224
|
+
- "Pattern says X but I'll adapt it differently"
|
|
225
|
+
- "Here are the main problems: [lists fixes without investigation]"
|
|
226
|
+
- Proposing solutions before tracing data flow
|
|
227
|
+
- **"One more fix attempt" (when already tried 2+)**
|
|
228
|
+
- **Each fix reveals new problem in different place**
|
|
229
|
+
|
|
230
|
+
**ALL of these mean: STOP. Return to Phase 1.**
|
|
231
|
+
|
|
232
|
+
**If 3+ fixes failed:** Question the architecture (see Phase 4.5)
|
|
233
|
+
|
|
234
|
+
## your human partner's Signals You're Doing It Wrong
|
|
235
|
+
|
|
236
|
+
**Watch for these redirections:**
|
|
237
|
+
- "Is that not happening?" — You assumed without verifying
|
|
238
|
+
- "Will it show us...?" — You should have added evidence gathering
|
|
239
|
+
- "Stop guessing" — You're proposing fixes without understanding
|
|
240
|
+
- "Ultrathink this" — Question fundamentals, not just symptoms
|
|
241
|
+
- "We're stuck?" (frustrated) — Your approach isn't working
|
|
242
|
+
|
|
243
|
+
**When you see these:** STOP. Return to Phase 1.
|
|
244
|
+
|
|
245
|
+
## Common Rationalizations
|
|
246
|
+
|
|
247
|
+
| Excuse | Reality |
|
|
248
|
+
|--------|---------|
|
|
249
|
+
| "Issue is simple, don't need process" | Simple issues 也有 root causes。Process 对 simple bugs 很快。 |
|
|
250
|
+
| "Emergency, no time for process" | Systematic debugging 比 guess-and-check thrashing 更快。 |
|
|
251
|
+
| "Just try this first, then investigate" | First fix 定模式。从一开始就做对。 |
|
|
252
|
+
| "I'll write test after confirming fix works" | Untested fixes 不 stick。Test first 证明它。 |
|
|
253
|
+
| "Multiple fixes at once saves time" | 无法 isolate what worked。制造新 bugs。 |
|
|
254
|
+
| "Reference too long, I'll adapt the pattern" | Partial understanding 保证 bugs。Complete 阅读。 |
|
|
255
|
+
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause。 |
|
|
256
|
+
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem。质疑 pattern,不要再 fix。 |
|
|
257
|
+
|
|
258
|
+
## Quick Reference
|
|
259
|
+
|
|
260
|
+
| Phase | Key Activities | Success Criteria |
|
|
261
|
+
|-------|---------------|------------------|
|
|
262
|
+
| **1. Root Cause** | Read errors, reproduce, check changes, gather evidence | Understand WHAT and WHY |
|
|
263
|
+
| **2. Pattern** | Find working examples, compare | Identify differences |
|
|
264
|
+
| **3. Hypothesis** | Form theory, test minimally | Confirmed or new hypothesis |
|
|
265
|
+
| **4. Implementation** | Create test, fix, verify | Bug resolved, tests pass |
|
|
266
|
+
|
|
267
|
+
## When Process Reveals "No Root Cause"
|
|
268
|
+
|
|
269
|
+
若 systematic investigation 表明 issue truly environmental、timing-dependent 或 external:
|
|
270
|
+
|
|
271
|
+
1. You've completed the process
|
|
272
|
+
2. Document what you investigated
|
|
273
|
+
3. Implement appropriate handling (retry, timeout, error message)
|
|
274
|
+
4. Add monitoring/logging for future investigation
|
|
275
|
+
|
|
276
|
+
**But:** 95% 的 "no root cause" cases 是 incomplete investigation。
|
|
277
|
+
|
|
278
|
+
## Supporting Techniques
|
|
279
|
+
|
|
280
|
+
本目录中属于 systematic debugging 的技术:
|
|
281
|
+
|
|
282
|
+
- **`root-cause-tracing.md`** — Trace bugs backward through call stack 找 original trigger
|
|
283
|
+
- **`defense-in-depth.md`** — 找到 root cause 后在 multiple layers 加 validation
|
|
284
|
+
- **`condition-based-waiting.md`** — 用 condition polling 替代 arbitrary timeouts
|
|
285
|
+
|
|
286
|
+
**Related principles:**
|
|
287
|
+
- **RED-GREEN-REFACTOR**(见 `docs/harness-methodology-tdd.md`)— 用于 creating failing test case(Phase 4, Step 1)
|
|
288
|
+
- **Verification discipline** — 宣称 success 前 verify fix worked。Run verification command,读 output,THEN claim result。
|
|
289
|
+
|
|
290
|
+
## Real-World Impact
|
|
291
|
+
|
|
292
|
+
来自 debugging sessions:
|
|
293
|
+
- Systematic approach:15-30 分钟 fix
|
|
294
|
+
- Random fixes approach:2-3 小时 thrashing
|
|
295
|
+
- First-time fix rate:95% vs 40%
|
|
296
|
+
- New bugs introduced:Near zero vs common
|