@tea-agent/loop-agent 0.1.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +62 -45
- package/CHANGELOG.md +60 -28
- package/README.md +160 -124
- package/bin/loop-agent.js +21 -21
- package/dist/adapters/index.js +3 -2
- package/dist/adapters/loop-agent.js +44 -2
- package/dist/application/dag/args.js +420 -0
- package/dist/application/dag/generate-task-dag.js +280 -0
- package/dist/application/dag/report-dag.js +14 -0
- package/dist/application/dag/run-dag.js +106 -0
- package/dist/application/dag/validate-dag.js +102 -0
- package/dist/application/loop/run-action.js +23 -0
- package/dist/cli/catalog.js +2 -237
- package/dist/cli/command-definitions.js +571 -0
- package/dist/cli/index.js +2 -0
- package/dist/cli/program.js +65 -1
- package/dist/cli/router.js +13 -0
- package/dist/cli-governance/active-residue-check.js +38 -0
- package/dist/commands/dag-report.js +6 -107
- package/dist/commands/dag-run-task.js +8 -466
- package/dist/commands/dag-validate.js +7 -179
- package/dist/commands/examples.js +90 -0
- package/dist/commands/init.js +1518 -0
- package/dist/commands/loop.js +57 -31
- package/dist/commands/pi-prompt.js +2 -9
- package/dist/commands/run-dag.js +7 -180
- package/dist/executors/cursor-executor-artifacts.js +3 -4
- package/dist/executors/cursor-worker-client.js +13 -3
- package/dist/executors/dag-cursor-executor.js +2 -3
- package/dist/executors/dag-pi-executor.js +3 -4
- package/dist/executors/dag-static-executor.js +2 -5
- package/dist/executors/pi-defaults.js +9 -0
- package/dist/executors/shell-executor.js +12 -20
- package/dist/governance/manifest-types.js +1 -0
- package/dist/infrastructure/harness/active-residue-policy.js +73 -0
- package/dist/infrastructure/harness/artifact-store.js +72 -0
- package/dist/infrastructure/harness/atomic-write.js +49 -0
- package/dist/infrastructure/harness/completed-facts-guard.js +40 -0
- package/dist/infrastructure/harness/loop-action-store.js +23 -0
- package/dist/infrastructure/harness/loop-store.js +41 -0
- package/dist/infrastructure/harness/one-shot-run-store.js +94 -0
- package/dist/infrastructure/harness/task-store.js +77 -0
- package/dist/records/one-shot-runs.js +26 -61
- package/dist/records/promotion.js +3 -4
- package/dist/shared/artifacts-core.js +5 -5
- package/dist/shared/logger.js +9 -15
- package/dist/task/delegate.js +4 -4
- package/dist/task/runtime.js +5 -7
- package/dist/task/state.js +6 -20
- package/dist/workflows/dag/convergence/controller.js +277 -0
- package/dist/workflows/dag/dynamic-runtime/condition.js +48 -0
- package/dist/workflows/dag/dynamic-runtime/loop-until.js +156 -0
- package/dist/workflows/dag/dynamic-runtime/map.js +185 -0
- package/dist/workflows/dag/dynamic-runtime/reduction.js +72 -0
- package/dist/workflows/dag/dynamic-runtime/shared.js +133 -0
- package/dist/workflows/dag/failure-routing.js +82 -0
- package/dist/workflows/dag/lifecycle.js +101 -8
- package/dist/workflows/dag/node-execution.js +262 -0
- package/dist/workflows/dag/report.js +73 -1
- package/dist/workflows/dag/run-store.js +36 -0
- package/dist/workflows/dag/runner.js +82 -1341
- package/dist/workflows/dag/scheduler.js +84 -0
- package/dist/workflows/dag/upstream-artifacts.js +20 -18
- package/dist/workflows/loop/actions/cursor-fix.js +191 -0
- package/dist/workflows/loop/actions/dag-action.js +130 -0
- package/dist/workflows/loop/actions/pi-review.js +267 -0
- package/dist/workflows/loop/actions/shared.js +157 -0
- package/dist/workflows/loop/actions/shell-verify.js +82 -0
- package/dist/workflows/loop/actions/types.js +1 -0
- package/dist/workflows/loop/actions/workflow-action.js +255 -0
- package/dist/workflows/loop/actions.js +55 -1212
- package/dist/workflows/loop/closeout.js +5 -4
- package/dist/workflows/loop/context.js +2 -3
- package/dist/workflows/loop/events.js +3 -2
- package/dist/workflows/loop/policy/auto-policy.js +104 -0
- package/dist/workflows/loop/policy/cursor-fix-policy.js +31 -0
- package/dist/workflows/loop/rounds.js +3 -3
- package/dist/workflows/loop/signals.js +4 -7
- package/dist/workflows/loop/state.js +11 -11
- package/docs/README.md +47 -44
- package/docs/agent-dag-recovery-playbook.md +32 -6
- package/docs/agent-dag-runner.md +17 -17
- package/docs/architecture/runtime-boundaries.md +147 -0
- package/docs/cursor-executor-usage.md +5 -5
- package/docs/decisions/README.md +2 -2
- package/docs/design/README.md +24 -24
- package/docs/development-principles.md +50 -50
- package/docs/dynamic-workflow-dag-engine-roadmap.md +6 -6
- package/docs/exec-plans/README.md +4 -4
- package/docs/exec-plans/active/README.md +10 -5
- package/docs/exec-plans/completed/README.md +9 -5
- package/docs/feature-workflow.md +111 -109
- package/docs/harness-methodology-verification.md +18 -18
- package/docs/loop-agent-harness.md +36 -36
- package/docs/production-readiness.md +96 -0
- package/docs/progress/README.md +2 -2
- package/docs/reports/README.md +4 -2
- package/docs/templates/agent-dag-decision-gate-dogfood-report.md +1 -1
- package/docs/templates/agent-dag-process-supervisor.prompt.md +2 -2
- package/docs/templates/agent-dag-report.schema.json +33 -2
- package/docs/templates/agent-dag-review-verdict.prompt.md +1 -1
- package/docs/templates/agent-dag.base.json +195 -195
- package/docs/templates/agent-dag.final-verification.json +190 -190
- package/docs/templates/agent-dag.schema.json +17 -17
- package/docs/templates/agent-dag.supervised-implementation.json +500 -500
- package/docs/templates/hybrid-dag.json +193 -193
- package/docs/templates/production-readiness-checklist.md +57 -0
- package/docs/templates/progress-log.md +7 -7
- package/docs/templates/project-start-checklist.md +8 -8
- package/docs/templates/qa-report.md +17 -11
- package/docs/templates/sprint-contract.md +19 -19
- package/docs/verification-matrix.md +37 -26
- package/examples/example-dag.json +51 -51
- package/examples/hybrid-loop-agent-dag.json +194 -194
- package/harness.json +5 -5
- package/package.json +62 -61
- package/skills/ai-engineering-context/SKILL.md +21 -21
- package/skills/loop-agent/SKILL.md +56 -171
- package/skills/loop-agent/references/README.md +6 -2
- package/skills/loop-agent/references/command-reference.md +107 -65
- package/skills/loop-agent/references/harness-policy.md +115 -115
- package/skills/loop-agent/references/hybrid-dag.md +30 -30
- package/skills/loop-agent/references/learned/README.md +13 -13
- package/skills/loop-agent/references/long-running-loop.md +59 -0
- package/skills/loop-agent/references/model-routing.md +1 -1
- package/skills/loop-agent/references/orchestrator-and-interventions.md +1 -1
- package/skills/loop-agent/references/pi-prompt.md +9 -9
- package/skills/loop-agent/references/pi-subagent-assisted-mode.md +0 -2
- package/skills/loop-agent/references/post-implementation-and-patterns.md +7 -7
- package/skills/loop-agent/references/task-workflow.md +19 -19
- package/skills/loop-agent/references/verification-and-failure-handling.md +54 -0
- package/skills/requesting-code-review/SKILL.md +40 -40
- package/skills/requesting-code-review/code-reviewer.md +4 -4
- package/skills/systematic-debugging/CREATION-LOG.md +43 -43
- package/skills/systematic-debugging/SKILL.md +113 -113
- package/skills/systematic-debugging/condition-based-waiting.md +20 -20
- package/skills/systematic-debugging/defense-in-depth.md +27 -27
- package/skills/systematic-debugging/root-cause-tracing.md +38 -38
- package/skills/systematic-debugging/test-academic.md +6 -6
- package/skills/systematic-debugging/test-pressure-1.md +6 -6
- package/skills/systematic-debugging/test-pressure-2.md +2 -2
- package/skills/systematic-debugging/test-pressure-3.md +6 -6
- package/skills/verification-before-completion/SKILL.md +37 -37
|
@@ -1,76 +1,76 @@
|
|
|
1
1
|
# Creation Log: Systematic Debugging Skill
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
提取、结构化与 bulletproofing 关键 skill 的 reference example。
|
|
4
4
|
|
|
5
5
|
## Source Material
|
|
6
6
|
|
|
7
|
-
|
|
8
|
-
- 4-phase systematic process
|
|
9
|
-
- Core mandate
|
|
10
|
-
-
|
|
7
|
+
从 `~/.claude/CLAUDE.md` 提取 debugging framework:
|
|
8
|
+
- 4-phase systematic process(Investigation → Pattern Analysis → Hypothesis → Implementation)
|
|
9
|
+
- Core mandate:ALWAYS find root cause,NEVER fix symptoms
|
|
10
|
+
- 设计以 resist time pressure 与 rationalization 的规则
|
|
11
11
|
|
|
12
12
|
## Extraction Decisions
|
|
13
13
|
|
|
14
14
|
**What to include:**
|
|
15
|
-
-
|
|
16
|
-
- Anti-shortcuts
|
|
17
|
-
- Pressure-resistant language
|
|
18
|
-
-
|
|
15
|
+
- 完整 4-phase framework 及所有 rules
|
|
16
|
+
- Anti-shortcuts("NEVER fix symptom"、"STOP and re-analyze")
|
|
17
|
+
- Pressure-resistant language("even if faster"、"even if I seem in a hurry")
|
|
18
|
+
- 各 phase 的 concrete steps
|
|
19
19
|
|
|
20
20
|
**What to leave out:**
|
|
21
21
|
- Project-specific context
|
|
22
|
-
-
|
|
23
|
-
- Narrative explanations
|
|
22
|
+
- 同一 rule 的 repetitive variations
|
|
23
|
+
- Narrative explanations(condensed 为 principles)
|
|
24
24
|
|
|
25
25
|
## Structure Following skill-creation/SKILL.md
|
|
26
26
|
|
|
27
|
-
1. **Rich when_to_use**
|
|
28
|
-
2. **Type: technique**
|
|
29
|
-
3. **Keywords**
|
|
30
|
-
4. **Flowchart**
|
|
31
|
-
5. **Phase-by-phase breakdown**
|
|
32
|
-
6. **Anti-patterns section**
|
|
27
|
+
1. **Rich when_to_use** — 含 symptoms 与 anti-patterns
|
|
28
|
+
2. **Type: technique** — 带 steps 的 concrete process
|
|
29
|
+
3. **Keywords** — "root cause"、"symptom"、"workaround"、"debugging"、"investigation"
|
|
30
|
+
4. **Flowchart** — "fix failed" 决策点 → re-analyze vs add more fixes
|
|
31
|
+
5. **Phase-by-phase breakdown** — Scannable checklist format
|
|
32
|
+
6. **Anti-patterns section** — 什么 NOT to do(对本 skill 关键)
|
|
33
33
|
|
|
34
34
|
## Bulletproofing Elements
|
|
35
35
|
|
|
36
|
-
Framework
|
|
36
|
+
Framework 设计以 resist rationalization under pressure:
|
|
37
37
|
|
|
38
38
|
### Language Choices
|
|
39
|
-
- "ALWAYS" / "NEVER"
|
|
39
|
+
- "ALWAYS" / "NEVER"(非 "should" / "try to")
|
|
40
40
|
- "even if faster" / "even if I seem in a hurry"
|
|
41
|
-
- "STOP and re-analyze"
|
|
42
|
-
- "Don't skip past"
|
|
41
|
+
- "STOP and re-analyze"(explicit pause)
|
|
42
|
+
- "Don't skip past"(捕获 actual behavior)
|
|
43
43
|
|
|
44
44
|
### Structural Defenses
|
|
45
|
-
- **Phase 1 required**
|
|
46
|
-
- **Single hypothesis rule**
|
|
47
|
-
- **Explicit failure mode**
|
|
48
|
-
- **Anti-patterns section**
|
|
45
|
+
- **Phase 1 required** — 不能 skip to implementation
|
|
46
|
+
- **Single hypothesis rule** — 强制思考,防止 shotgun fixes
|
|
47
|
+
- **Explicit failure mode** — "IF your first fix doesn't work" 及 mandatory action
|
|
48
|
+
- **Anti-patterns section** — 展示 shortcuts 的确切样子
|
|
49
49
|
|
|
50
50
|
### Redundancy
|
|
51
|
-
- Root cause mandate
|
|
52
|
-
- "NEVER fix symptom"
|
|
53
|
-
-
|
|
51
|
+
- Root cause mandate 在 overview + when_to_use + Phase 1 + implementation rules
|
|
52
|
+
- "NEVER fix symptom" 在不同 contexts 出现 4 次
|
|
53
|
+
- 各 phase 有 explicit "don't skip" guidance
|
|
54
54
|
|
|
55
55
|
## Testing Approach
|
|
56
56
|
|
|
57
|
-
|
|
57
|
+
按 skills/meta/testing-skills-with-subagents 创建 4 个 validation tests:
|
|
58
58
|
|
|
59
59
|
### Test 1: Academic Context (No Pressure)
|
|
60
|
-
- Simple bug
|
|
61
|
-
- **Result:** Perfect compliance
|
|
60
|
+
- Simple bug,无 time pressure
|
|
61
|
+
- **Result:** Perfect compliance,complete investigation
|
|
62
62
|
|
|
63
63
|
### Test 2: Time Pressure + Obvious Quick Fix
|
|
64
|
-
- User "in a hurry"
|
|
65
|
-
- **Result:** Resisted shortcut
|
|
64
|
+
- User "in a hurry",symptom fix 看起来 easy
|
|
65
|
+
- **Result:** Resisted shortcut,followed full process,found real root cause
|
|
66
66
|
|
|
67
67
|
### Test 3: Complex System + Uncertainty
|
|
68
|
-
- Multi-layer failure
|
|
69
|
-
- **Result:** Systematic investigation
|
|
68
|
+
- Multi-layer failure, unclear 能否 find root cause
|
|
69
|
+
- **Result:** Systematic investigation,traced through all layers,found source
|
|
70
70
|
|
|
71
71
|
### Test 4: Failed First Fix
|
|
72
|
-
- Hypothesis
|
|
73
|
-
- **Result:** Stopped
|
|
72
|
+
- Hypothesis 无效,temptation 加 more fixes
|
|
73
|
+
- **Result:** Stopped,re-analyzed,formed new hypothesis(no shotgun)
|
|
74
74
|
|
|
75
75
|
**All tests passed.** No rationalizations found.
|
|
76
76
|
|
|
@@ -99,16 +99,16 @@ Bulletproof skill that:
|
|
|
99
99
|
|
|
100
100
|
## Key Insight
|
|
101
101
|
|
|
102
|
-
**Most important bulletproofing
|
|
102
|
+
**Most important bulletproofing:** Anti-patterns section 展示 moment 里 feel justified 的 exact shortcuts。当 Claude 想 "I'll just add this one quick fix",看到 listed as wrong 的 exact pattern 产生 cognitive friction。
|
|
103
103
|
|
|
104
104
|
## Usage Example
|
|
105
105
|
|
|
106
|
-
|
|
106
|
+
遇到 bug 时:
|
|
107
107
|
1. Load skill: skills/debugging/systematic-debugging
|
|
108
|
-
2. Read overview (10 sec)
|
|
109
|
-
3. Follow Phase 1 checklist
|
|
110
|
-
4. If tempted to skip
|
|
111
|
-
5. Complete all phases
|
|
108
|
+
2. Read overview (10 sec) — reminded of mandate
|
|
109
|
+
3. Follow Phase 1 checklist — forced investigation
|
|
110
|
+
4. If tempted to skip — see anti-pattern,stop
|
|
111
|
+
5. Complete all phases — root cause found
|
|
112
112
|
|
|
113
113
|
**Time investment:** 5-10 minutes
|
|
114
114
|
**Time saved:** Hours of symptom-whack-a-mole
|
|
@@ -1,17 +1,17 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: systematic-debugging
|
|
3
|
-
description:
|
|
3
|
+
description: 遇到任何 bug、test failure 或 unexpected behavior 时使用,且在提出 fixes 之前
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Systematic Debugging
|
|
7
7
|
|
|
8
8
|
## Overview
|
|
9
9
|
|
|
10
|
-
Random fixes
|
|
10
|
+
Random fixes 浪费时间并制造新 bug。Quick patches 掩盖 underlying issues。
|
|
11
11
|
|
|
12
|
-
**Core principle
|
|
12
|
+
**Core principle:** ALWAYS 在尝试 fixes 之前找到 root cause。Symptom fixes 是 failure。
|
|
13
13
|
|
|
14
|
-
|
|
14
|
+
**违反本流程字面即违反 debugging 精神。**
|
|
15
15
|
|
|
16
16
|
## The Iron Law
|
|
17
17
|
|
|
@@ -19,61 +19,61 @@ Random fixes waste time and create new bugs. Quick patches mask underlying issue
|
|
|
19
19
|
NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST
|
|
20
20
|
```
|
|
21
21
|
|
|
22
|
-
|
|
22
|
+
若尚未完成 Phase 1,不得提出 fixes。
|
|
23
23
|
|
|
24
24
|
## When to Use
|
|
25
25
|
|
|
26
|
-
|
|
26
|
+
用于 ANY technical issue:
|
|
27
27
|
- Test failures
|
|
28
|
-
-
|
|
28
|
+
- Production bugs
|
|
29
29
|
- Unexpected behavior
|
|
30
30
|
- Performance problems
|
|
31
31
|
- Build failures
|
|
32
32
|
- Integration issues
|
|
33
33
|
|
|
34
|
-
**
|
|
35
|
-
-
|
|
36
|
-
- "Just one quick fix"
|
|
37
|
-
-
|
|
38
|
-
- Previous fix
|
|
39
|
-
-
|
|
34
|
+
**ESPECIALLY 在以下情况使用:**
|
|
35
|
+
- 时间压力下(emergencies 使 guessing 诱人)
|
|
36
|
+
- "Just one quick fix" 看起来 obvious
|
|
37
|
+
- 已尝试 multiple fixes
|
|
38
|
+
- Previous fix 无效
|
|
39
|
+
- 未完全理解 issue
|
|
40
40
|
|
|
41
|
-
**Don't skip when
|
|
42
|
-
- Issue
|
|
43
|
-
-
|
|
44
|
-
- Manager
|
|
41
|
+
**Don't skip when:**
|
|
42
|
+
- Issue 看起来 simple(simple bugs 也有 root causes)
|
|
43
|
+
- 赶时间(rushing 保证 rework)
|
|
44
|
+
- Manager 要求 NOW 修好(systematic 比 thrashing 更快)
|
|
45
45
|
|
|
46
46
|
## The Four Phases
|
|
47
47
|
|
|
48
|
-
|
|
48
|
+
进入下一阶段前 MUST 完成每一 phase。
|
|
49
49
|
|
|
50
50
|
### Phase 1: Root Cause Investigation
|
|
51
51
|
|
|
52
|
-
|
|
52
|
+
**在尝试 ANY fix 之前:**
|
|
53
53
|
|
|
54
54
|
1. **Read Error Messages Carefully**
|
|
55
|
-
-
|
|
56
|
-
-
|
|
57
|
-
-
|
|
58
|
-
-
|
|
55
|
+
- 不要跳过 errors 或 warnings
|
|
56
|
+
- 它们常含 exact solution
|
|
57
|
+
- 完整阅读 stack traces
|
|
58
|
+
- 记下 line numbers、file paths、error codes
|
|
59
59
|
|
|
60
60
|
2. **Reproduce Consistently**
|
|
61
|
-
-
|
|
62
|
-
-
|
|
63
|
-
-
|
|
64
|
-
-
|
|
61
|
+
- 能否可靠触发?
|
|
62
|
+
- Exact steps 是什么?
|
|
63
|
+
- 是否每次都发生?
|
|
64
|
+
- 若不可 reproduce → 收集更多 data,不要 guess
|
|
65
65
|
|
|
66
66
|
3. **Check Recent Changes**
|
|
67
|
-
-
|
|
68
|
-
- Git diff
|
|
69
|
-
- New dependencies
|
|
67
|
+
- 什么变更可能导致此问题?
|
|
68
|
+
- Git diff、recent commits
|
|
69
|
+
- New dependencies、config changes
|
|
70
70
|
- Environmental differences
|
|
71
71
|
|
|
72
72
|
4. **Gather Evidence in Multi-Component Systems**
|
|
73
73
|
|
|
74
|
-
**WHEN system
|
|
74
|
+
**WHEN system 有多个 components(CI → build → signing,API → service → database):**
|
|
75
75
|
|
|
76
|
-
**BEFORE proposing fixes
|
|
76
|
+
**BEFORE proposing fixes,添加 diagnostic instrumentation:**
|
|
77
77
|
```
|
|
78
78
|
For EACH component boundary:
|
|
79
79
|
- Log what data enters component
|
|
@@ -109,112 +109,112 @@ You MUST complete each phase before proceeding to the next.
|
|
|
109
109
|
|
|
110
110
|
5. **Trace Data Flow**
|
|
111
111
|
|
|
112
|
-
**WHEN error
|
|
112
|
+
**WHEN error 在 call stack 深处:**
|
|
113
113
|
|
|
114
|
-
|
|
114
|
+
完整 backward tracing 见本目录 `root-cause-tracing.md`。
|
|
115
115
|
|
|
116
|
-
**Quick version
|
|
117
|
-
-
|
|
118
|
-
-
|
|
119
|
-
-
|
|
120
|
-
-
|
|
116
|
+
**Quick version:**
|
|
117
|
+
- Bad value 从哪 originate?
|
|
118
|
+
- 谁用 bad value 调用了 this?
|
|
119
|
+
- 一直向上 trace 直到 source
|
|
120
|
+
- 在 source 修复,而非 symptom
|
|
121
121
|
|
|
122
122
|
### Phase 2: Pattern Analysis
|
|
123
123
|
|
|
124
|
-
**
|
|
124
|
+
**Fix 前先找 pattern:**
|
|
125
125
|
|
|
126
126
|
1. **Find Working Examples**
|
|
127
|
-
-
|
|
128
|
-
-
|
|
127
|
+
- 在同 codebase 找 similar working code
|
|
128
|
+
- 什么能 work、什么 broken?
|
|
129
129
|
|
|
130
130
|
2. **Compare Against References**
|
|
131
|
-
-
|
|
132
|
-
-
|
|
133
|
-
-
|
|
131
|
+
- 若实现 pattern,COMPLETE 阅读 reference implementation
|
|
132
|
+
- 不要 skim — 读每一行
|
|
133
|
+
- 应用前 fully 理解 pattern
|
|
134
134
|
|
|
135
135
|
3. **Identify Differences**
|
|
136
|
-
-
|
|
137
|
-
-
|
|
138
|
-
-
|
|
136
|
+
- Working 与 broken 有何不同?
|
|
137
|
+
- 列出 every difference,再小也要列
|
|
138
|
+
- 不要假设 "that can't matter"
|
|
139
139
|
|
|
140
140
|
4. **Understand Dependencies**
|
|
141
|
-
-
|
|
142
|
-
-
|
|
143
|
-
-
|
|
141
|
+
- 还需要哪些 other components?
|
|
142
|
+
- 哪些 settings、config、environment?
|
|
143
|
+
- 它作哪些 assumptions?
|
|
144
144
|
|
|
145
145
|
### Phase 3: Hypothesis and Testing
|
|
146
146
|
|
|
147
|
-
**Scientific method
|
|
147
|
+
**Scientific method:**
|
|
148
148
|
|
|
149
149
|
1. **Form Single Hypothesis**
|
|
150
|
-
-
|
|
151
|
-
-
|
|
152
|
-
-
|
|
150
|
+
- 清楚陈述:"I think X is the root cause because Y"
|
|
151
|
+
- 写下来
|
|
152
|
+
- 要 specific,不要 vague
|
|
153
153
|
|
|
154
154
|
2. **Test Minimally**
|
|
155
|
-
-
|
|
155
|
+
- 做 SMALLEST possible change 以 test hypothesis
|
|
156
156
|
- One variable at a time
|
|
157
|
-
-
|
|
157
|
+
- 不要一次 fix multiple things
|
|
158
158
|
|
|
159
159
|
3. **Verify Before Continuing**
|
|
160
|
-
-
|
|
161
|
-
-
|
|
162
|
-
- DON'T
|
|
160
|
+
- 有效?Yes → Phase 4
|
|
161
|
+
- 无效?Form NEW hypothesis
|
|
162
|
+
- DON'T 在其上叠加更多 fixes
|
|
163
163
|
|
|
164
164
|
4. **When You Don't Know**
|
|
165
|
-
-
|
|
166
|
-
-
|
|
165
|
+
- 说 "I don't understand X"
|
|
166
|
+
- 不要假装知道
|
|
167
167
|
- Ask for help
|
|
168
168
|
- Research more
|
|
169
169
|
|
|
170
170
|
### Phase 4: Implementation
|
|
171
171
|
|
|
172
|
-
**Fix
|
|
172
|
+
**Fix root cause,不是 symptom:**
|
|
173
173
|
|
|
174
174
|
1. **Create Failing Test Case**
|
|
175
175
|
- Simplest possible reproduction
|
|
176
|
-
-
|
|
177
|
-
-
|
|
178
|
-
- MUST
|
|
179
|
-
-
|
|
176
|
+
- 可能的话用 automated test
|
|
177
|
+
- 无 framework 时用 one-off test script
|
|
178
|
+
- MUST 在 fix 之前有
|
|
179
|
+
- 遵循 RED-GREEN-REFACTOR:写 failing test,看它 fail,再 fix
|
|
180
180
|
|
|
181
181
|
2. **Implement Single Fix**
|
|
182
|
-
-
|
|
182
|
+
- 针对已识别的 root cause
|
|
183
183
|
- ONE change at a time
|
|
184
|
-
-
|
|
185
|
-
-
|
|
184
|
+
- 无 "while I'm here" improvements
|
|
185
|
+
- 无 bundled refactoring
|
|
186
186
|
|
|
187
187
|
3. **Verify Fix**
|
|
188
|
-
- Test
|
|
189
|
-
-
|
|
190
|
-
- Issue
|
|
188
|
+
- Test 现在 pass?
|
|
189
|
+
- 无 other tests broken?
|
|
190
|
+
- Issue 真的 resolved?
|
|
191
191
|
|
|
192
192
|
4. **If Fix Doesn't Work**
|
|
193
193
|
- STOP
|
|
194
|
-
- Count
|
|
195
|
-
-
|
|
196
|
-
-
|
|
197
|
-
- DON'T
|
|
194
|
+
- Count:已尝试多少 fixes?
|
|
195
|
+
- 若 < 3:Return to Phase 1,用 new information 再分析
|
|
196
|
+
- **若 ≥ 3:STOP 并质疑 architecture(见下方 step 5)**
|
|
197
|
+
- DON'T 在未做 architectural discussion 前尝试 Fix #4
|
|
198
198
|
|
|
199
199
|
5. **If 3+ Fixes Failed: Question Architecture**
|
|
200
200
|
|
|
201
|
-
|
|
202
|
-
-
|
|
203
|
-
- Fixes
|
|
204
|
-
-
|
|
201
|
+
**表明 architectural problem 的 pattern:**
|
|
202
|
+
- 每个 fix 在不同位置 reveal 新的 shared state/coupling/problem
|
|
203
|
+
- Fixes 需要 "massive refactoring" 才能实现
|
|
204
|
+
- 每个 fix 在其他地方制造新 symptoms
|
|
205
205
|
|
|
206
|
-
**STOP
|
|
207
|
-
-
|
|
208
|
-
-
|
|
209
|
-
-
|
|
206
|
+
**STOP 并质疑 fundamentals:**
|
|
207
|
+
- 此 pattern fundamentally sound 吗?
|
|
208
|
+
- 是否 "sticking with it through sheer inertia"?
|
|
209
|
+
- 应 refactor architecture 还是继续 fix symptoms?
|
|
210
210
|
|
|
211
211
|
**Discuss with your human partner before attempting more fixes**
|
|
212
212
|
|
|
213
|
-
This is NOT a failed hypothesis
|
|
213
|
+
This is NOT a failed hypothesis — this is a wrong architecture.
|
|
214
214
|
|
|
215
215
|
## Red Flags - STOP and Follow Process
|
|
216
216
|
|
|
217
|
-
|
|
217
|
+
若发现自己想:
|
|
218
218
|
- "Quick fix for now, investigate later"
|
|
219
219
|
- "Just try changing X and see if it works"
|
|
220
220
|
- "Add multiple changes, run tests"
|
|
@@ -234,11 +234,11 @@ If you catch yourself thinking:
|
|
|
234
234
|
## your human partner's Signals You're Doing It Wrong
|
|
235
235
|
|
|
236
236
|
**Watch for these redirections:**
|
|
237
|
-
- "Is that not happening?"
|
|
238
|
-
- "Will it show us...?"
|
|
239
|
-
- "Stop guessing"
|
|
240
|
-
- "Ultrathink this"
|
|
241
|
-
- "We're stuck?" (frustrated)
|
|
237
|
+
- "Is that not happening?" — You assumed without verifying
|
|
238
|
+
- "Will it show us...?" — You should have added evidence gathering
|
|
239
|
+
- "Stop guessing" — You're proposing fixes without understanding
|
|
240
|
+
- "Ultrathink this" — Question fundamentals, not just symptoms
|
|
241
|
+
- "We're stuck?" (frustrated) — Your approach isn't working
|
|
242
242
|
|
|
243
243
|
**When you see these:** STOP. Return to Phase 1.
|
|
244
244
|
|
|
@@ -246,14 +246,14 @@ If you catch yourself thinking:
|
|
|
246
246
|
|
|
247
247
|
| Excuse | Reality |
|
|
248
248
|
|--------|---------|
|
|
249
|
-
| "Issue is simple, don't need process" | Simple issues
|
|
250
|
-
| "Emergency, no time for process" | Systematic debugging
|
|
251
|
-
| "Just try this first, then investigate" | First fix
|
|
252
|
-
| "I'll write test after confirming fix works" | Untested fixes
|
|
253
|
-
| "Multiple fixes at once saves time" |
|
|
254
|
-
| "Reference too long, I'll adapt the pattern" | Partial understanding
|
|
255
|
-
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause
|
|
256
|
-
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem
|
|
249
|
+
| "Issue is simple, don't need process" | Simple issues 也有 root causes。Process 对 simple bugs 很快。 |
|
|
250
|
+
| "Emergency, no time for process" | Systematic debugging 比 guess-and-check thrashing 更快。 |
|
|
251
|
+
| "Just try this first, then investigate" | First fix 定模式。从一开始就做对。 |
|
|
252
|
+
| "I'll write test after confirming fix works" | Untested fixes 不 stick。Test first 证明它。 |
|
|
253
|
+
| "Multiple fixes at once saves time" | 无法 isolate what worked。制造新 bugs。 |
|
|
254
|
+
| "Reference too long, I'll adapt the pattern" | Partial understanding 保证 bugs。Complete 阅读。 |
|
|
255
|
+
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause。 |
|
|
256
|
+
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem。质疑 pattern,不要再 fix。 |
|
|
257
257
|
|
|
258
258
|
## Quick Reference
|
|
259
259
|
|
|
@@ -266,31 +266,31 @@ If you catch yourself thinking:
|
|
|
266
266
|
|
|
267
267
|
## When Process Reveals "No Root Cause"
|
|
268
268
|
|
|
269
|
-
|
|
269
|
+
若 systematic investigation 表明 issue truly environmental、timing-dependent 或 external:
|
|
270
270
|
|
|
271
271
|
1. You've completed the process
|
|
272
272
|
2. Document what you investigated
|
|
273
273
|
3. Implement appropriate handling (retry, timeout, error message)
|
|
274
274
|
4. Add monitoring/logging for future investigation
|
|
275
275
|
|
|
276
|
-
**But:** 95%
|
|
276
|
+
**But:** 95% 的 "no root cause" cases 是 incomplete investigation。
|
|
277
277
|
|
|
278
278
|
## Supporting Techniques
|
|
279
279
|
|
|
280
|
-
|
|
280
|
+
本目录中属于 systematic debugging 的技术:
|
|
281
281
|
|
|
282
|
-
- **`root-cause-tracing.md`**
|
|
283
|
-
- **`defense-in-depth.md`**
|
|
284
|
-
- **`condition-based-waiting.md`**
|
|
282
|
+
- **`root-cause-tracing.md`** — Trace bugs backward through call stack 找 original trigger
|
|
283
|
+
- **`defense-in-depth.md`** — 找到 root cause 后在 multiple layers 加 validation
|
|
284
|
+
- **`condition-based-waiting.md`** — 用 condition polling 替代 arbitrary timeouts
|
|
285
285
|
|
|
286
286
|
**Related principles:**
|
|
287
|
-
- **RED-GREEN-REFACTOR
|
|
288
|
-
- **Verification discipline** —
|
|
287
|
+
- **RED-GREEN-REFACTOR**(见 `docs/harness-methodology-tdd.md`)— 用于 creating failing test case(Phase 4, Step 1)
|
|
288
|
+
- **Verification discipline** — 宣称 success 前 verify fix worked。Run verification command,读 output,THEN claim result。
|
|
289
289
|
|
|
290
290
|
## Real-World Impact
|
|
291
291
|
|
|
292
|
-
|
|
293
|
-
- Systematic approach
|
|
294
|
-
- Random fixes approach
|
|
295
|
-
- First-time fix rate
|
|
296
|
-
- New bugs introduced
|
|
292
|
+
来自 debugging sessions:
|
|
293
|
+
- Systematic approach:15-30 分钟 fix
|
|
294
|
+
- Random fixes approach:2-3 小时 thrashing
|
|
295
|
+
- First-time fix rate:95% vs 40%
|
|
296
|
+
- New bugs introduced:Near zero vs common
|
|
@@ -2,9 +2,9 @@
|
|
|
2
2
|
|
|
3
3
|
## Overview
|
|
4
4
|
|
|
5
|
-
Flaky tests
|
|
5
|
+
Flaky tests 常用 arbitrary delays 猜 timing。这制造 race conditions:fast machines 上 pass,load 或 CI 下 fail。
|
|
6
6
|
|
|
7
|
-
**Core principle
|
|
7
|
+
**Core principle:** Wait for 你真正关心的 actual condition,不是猜需要多久。
|
|
8
8
|
|
|
9
9
|
## When to Use
|
|
10
10
|
|
|
@@ -21,15 +21,15 @@ digraph when_to_use {
|
|
|
21
21
|
}
|
|
22
22
|
```
|
|
23
23
|
|
|
24
|
-
**Use when
|
|
25
|
-
- Tests
|
|
26
|
-
- Tests
|
|
27
|
-
-
|
|
28
|
-
-
|
|
24
|
+
**Use when:**
|
|
25
|
+
- Tests 有 arbitrary delays(`setTimeout`、`sleep`、`time.sleep()`)
|
|
26
|
+
- Tests flaky(有时 pass,load 下 fail)
|
|
27
|
+
- Parallel 运行时 timeout
|
|
28
|
+
- 等待 async operations 完成
|
|
29
29
|
|
|
30
|
-
**Don't use when
|
|
31
|
-
-
|
|
32
|
-
-
|
|
30
|
+
**Don't use when:**
|
|
31
|
+
- 测试 actual timing behavior(debounce、throttle intervals)
|
|
32
|
+
- 若用 arbitrary timeout,ALWAYS document WHY
|
|
33
33
|
|
|
34
34
|
## Core Pattern
|
|
35
35
|
|
|
@@ -79,18 +79,18 @@ async function waitFor<T>(
|
|
|
79
79
|
}
|
|
80
80
|
```
|
|
81
81
|
|
|
82
|
-
|
|
82
|
+
完整实现及 domain-specific helpers(`waitForEvent`、`waitForEventCount`、`waitForEventMatch`)见本目录 `condition-based-waiting-example.ts`,来自 actual debugging session。
|
|
83
83
|
|
|
84
84
|
## Common Mistakes
|
|
85
85
|
|
|
86
|
-
**❌ Polling too fast:** `setTimeout(check, 1)`
|
|
86
|
+
**❌ Polling too fast:** `setTimeout(check, 1)` — wastes CPU
|
|
87
87
|
**✅ Fix:** Poll every 10ms
|
|
88
88
|
|
|
89
|
-
**❌ No timeout:**
|
|
89
|
+
**❌ No timeout:** 条件永不满足则 loop forever
|
|
90
90
|
**✅ Fix:** Always include timeout with clear error
|
|
91
91
|
|
|
92
|
-
**❌ Stale data:**
|
|
93
|
-
**✅ Fix:**
|
|
92
|
+
**❌ Stale data:** Loop 前 cache state
|
|
93
|
+
**✅ Fix:** Loop 内 call getter 取 fresh data
|
|
94
94
|
|
|
95
95
|
## When Arbitrary Timeout IS Correct
|
|
96
96
|
|
|
@@ -103,13 +103,13 @@ await new Promise(r => setTimeout(r, 200)); // Then: wait for timed behavior
|
|
|
103
103
|
|
|
104
104
|
**Requirements:**
|
|
105
105
|
1. First wait for triggering condition
|
|
106
|
-
2. Based on known timing
|
|
106
|
+
2. Based on known timing(not guessing)
|
|
107
107
|
3. Comment explaining WHY
|
|
108
108
|
|
|
109
109
|
## Real-World Impact
|
|
110
110
|
|
|
111
|
-
|
|
112
|
-
-
|
|
113
|
-
- Pass rate
|
|
114
|
-
- Execution time
|
|
111
|
+
来自 debugging session (2025-10-03):
|
|
112
|
+
- 修复 3 个文件中 15 个 flaky tests
|
|
113
|
+
- Pass rate:60% → 100%
|
|
114
|
+
- Execution time:40% faster
|
|
115
115
|
- No more race conditions
|