@miphamai/cli 0.24.7 → 0.24.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skills/standard/implement.SKILL.md +155 -0
- package/skills/standard/research.SKILL.md +147 -0
- package/skills/standard/systematic-debugging.SKILL.md +236 -0
- package/src/i18n-core/locales/en-US.json +1 -1
- package/src/i18n-core/locales/zh-CN.json +1 -1
- package/src/index.tsx +1 -0
- package/src/ui/app.tsx +106 -85
- package/src/ui/input.tsx +9 -0
package/package.json
CHANGED
|
@@ -0,0 +1,155 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: implement
|
|
3
|
+
description: Build work from a spec or tickets with systematic discipline — TDD at pre-agreed seams, incremental verification, code review before commit. Use when implementing features, bugfixes, or any planned work.
|
|
4
|
+
version: 1.0.0
|
|
5
|
+
user-invocable: true
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Implement — Structured Build Execution
|
|
9
|
+
|
|
10
|
+
融合 Superpowers executing-plans(计划审阅 + 隔离工作区)+ Matt Pocock implement(TDD 接缝 + 增量验证 + 提交前审查)。
|
|
11
|
+
|
|
12
|
+
## When to Use
|
|
13
|
+
|
|
14
|
+
- Implementing work from a written spec or ticket set
|
|
15
|
+
- Executing a development plan with clear deliverables
|
|
16
|
+
- Building a feature with predefined success criteria
|
|
17
|
+
|
|
18
|
+
## When NOT to Use
|
|
19
|
+
|
|
20
|
+
- Exploratory coding / prototyping → use `prototype` skill
|
|
21
|
+
- Quick one-line fixes → just fix it
|
|
22
|
+
- No spec or tickets exist → use `to-tickets` or `to-spec` first
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Step 1: Load and Review
|
|
27
|
+
|
|
28
|
+
### 1.1 Ensure isolated workspace
|
|
29
|
+
|
|
30
|
+
Use git worktree or a feature branch. Never implement on main/master without explicit consent.
|
|
31
|
+
|
|
32
|
+
### 1.2 Read the plan/spec/tickets
|
|
33
|
+
|
|
34
|
+
Read the full spec or ticket set. Understand:
|
|
35
|
+
|
|
36
|
+
- What is being built?
|
|
37
|
+
- What are the acceptance criteria?
|
|
38
|
+
- What are the pre-agreed seams (where TDD should be applied)?
|
|
39
|
+
|
|
40
|
+
### 1.3 Review critically
|
|
41
|
+
|
|
42
|
+
Before writing any code:
|
|
43
|
+
|
|
44
|
+
- Are there gaps or ambiguities in the spec?
|
|
45
|
+
- Are the success criteria testable?
|
|
46
|
+
- Do you understand every instruction?
|
|
47
|
+
|
|
48
|
+
**If concerns exist, raise them before starting.** Don't guess.
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## Step 2: Execute Tasks
|
|
53
|
+
|
|
54
|
+
For each task in order:
|
|
55
|
+
|
|
56
|
+
### 2.1 At pre-agreed seams: TDD
|
|
57
|
+
|
|
58
|
+
Where the spec specifies (or where interfaces are well-defined):
|
|
59
|
+
|
|
60
|
+
1. Write a **failing test** that asserts the expected behavior
|
|
61
|
+
2. Watch it fail (red)
|
|
62
|
+
3. Write the **minimum code** to make it pass (green)
|
|
63
|
+
4. Refactor if needed, keeping tests green
|
|
64
|
+
|
|
65
|
+
Use the `tdd` skill for full red-green-refactor discipline.
|
|
66
|
+
|
|
67
|
+
### 2.2 Incremental verification
|
|
68
|
+
|
|
69
|
+
During implementation:
|
|
70
|
+
|
|
71
|
+
- **Run typecheck** after each significant change: `pnpm typecheck`
|
|
72
|
+
- **Run relevant test file** after each task: `pnpm test -- <file>`
|
|
73
|
+
- **Don't wait** until everything is done to discover type errors
|
|
74
|
+
|
|
75
|
+
### 2.3 One task at a time
|
|
76
|
+
|
|
77
|
+
- Follow each step exactly — the plan has bite-sized steps for a reason
|
|
78
|
+
- One change at a time. No "while I'm here" improvements.
|
|
79
|
+
- Mark tasks as complete after verification passes
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
## Step 3: Final Verification
|
|
84
|
+
|
|
85
|
+
After all tasks are complete:
|
|
86
|
+
|
|
87
|
+
### 3.1 Full test suite
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
pnpm test
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
All tests must pass. If any fail, fix before proceeding.
|
|
94
|
+
|
|
95
|
+
### 3.2 Lint and format
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
pnpm lint
|
|
99
|
+
pnpm format
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
CI must be green.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## Step 4: Code Review
|
|
107
|
+
|
|
108
|
+
**Before committing**, run code review:
|
|
109
|
+
|
|
110
|
+
Use the `code-review` skill for a two-axis review:
|
|
111
|
+
|
|
112
|
+
- **Standards**: Does the diff follow the repo's coding standards?
|
|
113
|
+
- **Spec**: Does it faithfully implement the originating issue/spec?
|
|
114
|
+
|
|
115
|
+
Fix any findings before committing.
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## Step 5: Commit
|
|
120
|
+
|
|
121
|
+
Commit your work to the current branch.
|
|
122
|
+
|
|
123
|
+
```bash
|
|
124
|
+
git add -A
|
|
125
|
+
git commit -m "<type>: <description>"
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
- Follow Conventional Commits
|
|
129
|
+
- Reference the spec/ticket in the commit message
|
|
130
|
+
- **Do NOT commit unless explicitly asked** (per CLAUDE.md §关键约束)
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
## When to Stop and Ask
|
|
135
|
+
|
|
136
|
+
**STOP immediately when:**
|
|
137
|
+
|
|
138
|
+
- A task is blocked (missing dependency, unclear instruction, verification fails repeatedly)
|
|
139
|
+
- The spec has a critical gap that prevents starting
|
|
140
|
+
- You don't understand an instruction
|
|
141
|
+
- 3+ fix attempts fail — this may be an architectural issue
|
|
142
|
+
|
|
143
|
+
**Ask for clarification rather than guessing.**
|
|
144
|
+
|
|
145
|
+
---
|
|
146
|
+
|
|
147
|
+
## Quick Reference
|
|
148
|
+
|
|
149
|
+
| Step | Key Activities | Done When |
|
|
150
|
+
| -------------- | ------------------------------------------------------------ | -------------------------------- |
|
|
151
|
+
| **1. Review** | Load spec, isolate workspace, review critically | All concerns raised and resolved |
|
|
152
|
+
| **2. Execute** | TDD at seams, incremental typecheck/test, one task at a time | All tasks complete and verified |
|
|
153
|
+
| **3. Verify** | Full test suite, lint, format | CI-ready (all green) |
|
|
154
|
+
| **4. Review** | Two-axis code review (standards + spec) | Findings addressed |
|
|
155
|
+
| **5. Commit** | Conventional Commits, reference spec/ticket | Work committed to branch |
|
|
@@ -0,0 +1,147 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: research
|
|
3
|
+
description: Deep research against primary sources, executed as a background agent. Collects findings into a single cited Markdown file. Use for investigation that requires reading official docs, source code, specs, or first-party APIs — not secondary summaries.
|
|
4
|
+
version: 1.0.0
|
|
5
|
+
user-invocable: true
|
|
6
|
+
allowed-tools:
|
|
7
|
+
- WebSearch
|
|
8
|
+
- WebFetch
|
|
9
|
+
- Agent
|
|
10
|
+
- Bash
|
|
11
|
+
- Write
|
|
12
|
+
- Read
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
# Research — Background Deep Research
|
|
16
|
+
|
|
17
|
+
融合 Mipham web-search v3.0(查询构建+验证)+ Matt Pocock research(后台代理+一手来源+Markdown 报告)。
|
|
18
|
+
|
|
19
|
+
## When to Use
|
|
20
|
+
|
|
21
|
+
- "Research X for me"
|
|
22
|
+
- "Find out everything about Y from primary sources"
|
|
23
|
+
- "Investigate Z and write up findings"
|
|
24
|
+
- Any question where googling + reading multiple sources is the right answer
|
|
25
|
+
|
|
26
|
+
## When NOT to Use
|
|
27
|
+
|
|
28
|
+
- Quick fact lookup → use `/web-search` directly
|
|
29
|
+
- Question answerable from code already in context
|
|
30
|
+
- Pure logic/algorithmic question
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## Phase 0: Route
|
|
35
|
+
|
|
36
|
+
```
|
|
37
|
+
Research task is...
|
|
38
|
+
├── Quick (1-2 sources, immediate answer)?
|
|
39
|
+
│ └── → Use web-search skill directly (Phase 0-4)
|
|
40
|
+
│
|
|
41
|
+
├── Deep (multiple sources, needs synthesis)?
|
|
42
|
+
│ └── → THIS SKILL — background agent
|
|
43
|
+
│
|
|
44
|
+
└── Login-walled / SPA-only sources?
|
|
45
|
+
└── → web-access skill (ComputerUse browser)
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
|
|
50
|
+
## Phase 1: Spin Up Background Agent
|
|
51
|
+
|
|
52
|
+
Launch a **background agent** to do the heavy reading, so you keep working while it researches.
|
|
53
|
+
|
|
54
|
+
The agent's instructions:
|
|
55
|
+
|
|
56
|
+
```
|
|
57
|
+
You are a research agent. Your task:
|
|
58
|
+
|
|
59
|
+
1. Investigate the question against PRIMARY SOURCES ONLY:
|
|
60
|
+
- Official documentation (docs.*.com, *.org)
|
|
61
|
+
- Source code repositories (GitHub, GitLab)
|
|
62
|
+
- Technical specifications (RFCs, standards)
|
|
63
|
+
- First-party API references
|
|
64
|
+
- NOT: blog posts, Medium articles, forum threads, secondary summaries
|
|
65
|
+
|
|
66
|
+
2. For every claim, follow it back to the source that owns it.
|
|
67
|
+
If a secondary source makes a claim, find the primary source and cite that.
|
|
68
|
+
|
|
69
|
+
3. Use WebSearch to find sources.
|
|
70
|
+
Use WebFetch to deep-read promising pages.
|
|
71
|
+
Cross-reference critical claims across 2+ independent primary sources.
|
|
72
|
+
|
|
73
|
+
4. Write findings to a SINGLE Markdown file.
|
|
74
|
+
- Cite every claim with its primary source URL
|
|
75
|
+
- Distinguish between facts (needs citation) and reasoning (your own)
|
|
76
|
+
- Flag outdated content ("article from 2024, may be stale")
|
|
77
|
+
- Note if a source is official docs vs community
|
|
78
|
+
|
|
79
|
+
5. Save the file where the repo already keeps such notes.
|
|
80
|
+
Match existing conventions. If none exist, put it in docs/research/.
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
---
|
|
84
|
+
|
|
85
|
+
## Phase 2: Report Format
|
|
86
|
+
|
|
87
|
+
The agent writes findings in this structure:
|
|
88
|
+
|
|
89
|
+
```markdown
|
|
90
|
+
# [Research Topic]
|
|
91
|
+
|
|
92
|
+
**Date**: YYYY-MM-DD
|
|
93
|
+
**Sources**: N primary, M cross-references
|
|
94
|
+
|
|
95
|
+
## Key Findings
|
|
96
|
+
|
|
97
|
+
- [Finding 1] — [Source](URL)
|
|
98
|
+
- [Finding 2] — [Source](URL)
|
|
99
|
+
|
|
100
|
+
## Detailed Analysis
|
|
101
|
+
|
|
102
|
+
### [Subtopic A]
|
|
103
|
+
|
|
104
|
+
[Claim and citation]
|
|
105
|
+
|
|
106
|
+
### [Subtopic B]
|
|
107
|
+
|
|
108
|
+
[Claim and citation]
|
|
109
|
+
|
|
110
|
+
## Source Evaluation
|
|
111
|
+
|
|
112
|
+
| Source | Type | Authority | Notes |
|
|
113
|
+
| ----------- | ------------- | --------- | --------------------- |
|
|
114
|
+
| [Name](URL) | Official docs | High | Current as of YYYY-MM |
|
|
115
|
+
| [Name](URL) | Source code | High | Tag vX.Y.Z |
|
|
116
|
+
|
|
117
|
+
## Open Questions
|
|
118
|
+
|
|
119
|
+
- [Question 1]
|
|
120
|
+
- [Question 2]
|
|
121
|
+
|
|
122
|
+
Sources:
|
|
123
|
+
|
|
124
|
+
- [Title](URL) — brief note
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
---
|
|
128
|
+
|
|
129
|
+
## Phase 3: Review
|
|
130
|
+
|
|
131
|
+
When the background agent completes:
|
|
132
|
+
|
|
133
|
+
1. Read the output file
|
|
134
|
+
2. Spot-check: did it follow the chain back to primary sources?
|
|
135
|
+
3. Flag any claims that need further verification
|
|
136
|
+
4. Surface uncertainties to the user
|
|
137
|
+
|
|
138
|
+
---
|
|
139
|
+
|
|
140
|
+
## Research Quality Checklist
|
|
141
|
+
|
|
142
|
+
- [ ] Every factual claim has a primary source citation
|
|
143
|
+
- [ ] At least one critical claim is cross-referenced (2+ sources)
|
|
144
|
+
- [ ] Source type is clearly identified (official docs / source code / spec / community)
|
|
145
|
+
- [ ] Outdated content is flagged with publication year
|
|
146
|
+
- [ ] Reasoning vs facts are clearly distinguished
|
|
147
|
+
- [ ] File saved in repo-appropriate location
|
|
@@ -0,0 +1,236 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: debug-loop
|
|
3
|
+
description: Enhanced diagnosis with feedback loop construction (10 methods) — build a tight red/green signal, minimize, falsifiable hypotheses, tagged instrumentation, seam assessment, post-mortem. Complements systematic-debugging. Use together for thorough debugging.
|
|
4
|
+
version: 1.0.0
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Debug Loop — 反馈闭环诊断
|
|
8
|
+
|
|
9
|
+
融合 Matt Pocock diagnosing-bugs(反馈闭环方法论)+ Superpowers systematic-debugging(反猜測紀律)。
|
|
10
|
+
|
|
11
|
+
> **与 `systematic-debugging` 互补**:systematic-debugging 侧重反猜測纪律和根因分析框架;debug-loop 侧重构建可执行的红绿反馈信号。两者配合使用效果最佳。
|
|
12
|
+
|
|
13
|
+
## The Iron Law
|
|
14
|
+
|
|
15
|
+
```
|
|
16
|
+
NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST.
|
|
17
|
+
NO HYPOTHESIS WITHOUT A RED-CAPABLE FEEDBACK LOOP FIRST.
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Phase 0: Decide Whether to Use This Skill
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
Issue is...
|
|
26
|
+
├── Test failure? → USE THIS SKILL
|
|
27
|
+
├── Bug in production? → USE THIS SKILL
|
|
28
|
+
├── Performance regression? → USE THIS SKILL
|
|
29
|
+
├── Build/integration break? → USE THIS SKILL
|
|
30
|
+
├── "It's probably X, quick fix" → USE THIS SKILL (especially now)
|
|
31
|
+
└── Trivial typo/syntax? → fix directly (but still verify)
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
## Phase 1: Build a Feedback Loop 🔴 THE SKILL
|
|
37
|
+
|
|
38
|
+
**This is the centerpiece.** A tight pass/fail signal for the bug — one that goes red on _this_ bug — makes everything else mechanical. No loop = no debugging, only guessing.
|
|
39
|
+
|
|
40
|
+
Spend disproportionate effort here. Be aggressive. Be creative. Refuse to give up.
|
|
41
|
+
|
|
42
|
+
### 1.1 Ways to construct one (try in roughly this order)
|
|
43
|
+
|
|
44
|
+
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
|
|
45
|
+
2. **Curl / HTTP script** against a running dev server.
|
|
46
|
+
3. **CLI invocation** with fixture input, diffing stdout against known-good snapshot.
|
|
47
|
+
4. **Headless browser script** (Playwright/Puppeteer) — drives UI, asserts on DOM/console/network.
|
|
48
|
+
5. **Replay a captured trace.** Save a real network request/payload/event log to disk; replay through the code path in isolation.
|
|
49
|
+
6. **Throwaway harness.** Spin up minimal subset of the system (one service, mocked deps) that exercises the bug path with a single function call.
|
|
50
|
+
7. **Property/fuzz loop.** If "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
|
51
|
+
8. **Bisection harness.** Automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
|
52
|
+
9. **Differential loop.** Run same input through old-version vs new-version and diff outputs.
|
|
53
|
+
10. **Multi-component evidence gathering.** For systems with multiple layers:
|
|
54
|
+
```
|
|
55
|
+
For EACH component boundary:
|
|
56
|
+
- Log what data enters
|
|
57
|
+
- Log what data exits
|
|
58
|
+
- Verify environment/config propagation
|
|
59
|
+
- Check state at each layer
|
|
60
|
+
|
|
61
|
+
Run once to identify WHICH layer fails, THEN investigate that component.
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
### 1.2 Tighten the loop
|
|
65
|
+
|
|
66
|
+
Once you have _a_ loop, make it tighter:
|
|
67
|
+
|
|
68
|
+
- **Faster**: Cache setup, skip unrelated init, narrow test scope.
|
|
69
|
+
- **Sharper signal**: Assert on the specific symptom, not "didn't crash".
|
|
70
|
+
- **More deterministic**: Pin time, seed RNG, isolate filesystem, freeze network.
|
|
71
|
+
|
|
72
|
+
A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is a superpower.
|
|
73
|
+
|
|
74
|
+
### 1.3 Non-deterministic bugs
|
|
75
|
+
|
|
76
|
+
Goal: higher reproduction rate (not clean repro). Loop 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake is debuggable; 1% is not — keep raising it.
|
|
77
|
+
|
|
78
|
+
### 1.4 When you genuinely cannot build a loop
|
|
79
|
+
|
|
80
|
+
Stop explicitly. List everything tried. Ask for: (a) access to the reproducing environment, (b) a redacted captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. **Do not proceed without a loop.**
|
|
81
|
+
|
|
82
|
+
### 1.5 Completion criterion
|
|
83
|
+
|
|
84
|
+
Phase 1 is done when the loop is **tight** and **red-capable**:
|
|
85
|
+
|
|
86
|
+
- [ ] **Red-capable** — drives the actual bug path and asserts the user's exact symptom. Not "runs without erroring".
|
|
87
|
+
- [ ] **Deterministic** — same verdict every run.
|
|
88
|
+
- [ ] **Fast** — seconds, not minutes.
|
|
89
|
+
- [ ] **Agent-runnable** — you can run it unattended.
|
|
90
|
+
|
|
91
|
+
> If you catch yourself reading code to build a theory before this loop exists — **STOP.** No red-capable command, no Phase 2.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## Phase 2: Reproduce + Minimise
|
|
96
|
+
|
|
97
|
+
### 2.1 Reproduce
|
|
98
|
+
|
|
99
|
+
Run the loop. Watch it go red.
|
|
100
|
+
|
|
101
|
+
- [ ] The loop produces the failure mode the **user** described — not a different nearby failure.
|
|
102
|
+
- [ ] The failure is reproducible (or, for flaky bugs, at a high enough rate).
|
|
103
|
+
- [ ] You have captured the exact symptom so later phases can verify the fix.
|
|
104
|
+
|
|
105
|
+
### 2.2 Minimise
|
|
106
|
+
|
|
107
|
+
Shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut.
|
|
108
|
+
|
|
109
|
+
**Why**: A minimal repro shrinks the hypothesis space and becomes the clean regression test in Phase 5.
|
|
110
|
+
|
|
111
|
+
Done when **every remaining element is load-bearing** — removing any one makes the loop go green.
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## Phase 3: Pattern Analysis + Hypothesise
|
|
116
|
+
|
|
117
|
+
### 3.1 Pattern Analysis
|
|
118
|
+
|
|
119
|
+
Before forming hypotheses:
|
|
120
|
+
|
|
121
|
+
- Find similar **working** code in the same codebase.
|
|
122
|
+
- Read the reference implementation completely — don't skim.
|
|
123
|
+
- List every difference between working and broken, however small.
|
|
124
|
+
- Understand dependencies, config, environment, assumptions.
|
|
125
|
+
|
|
126
|
+
### 3.2 Generate 3-5 Ranked Hypotheses
|
|
127
|
+
|
|
128
|
+
Generate multiple hypotheses **before testing any**. Single-hypothesis generation anchors on the first plausible idea.
|
|
129
|
+
|
|
130
|
+
Each hypothesis must be **falsifiable**:
|
|
131
|
+
|
|
132
|
+
> "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
|
133
|
+
|
|
134
|
+
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
|
|
135
|
+
|
|
136
|
+
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly. Don't block on it if they're AFK.
|
|
137
|
+
|
|
138
|
+
---
|
|
139
|
+
|
|
140
|
+
## Phase 4: Instrument
|
|
141
|
+
|
|
142
|
+
Each probe must map to a specific Phase 3 prediction. **Change one variable at a time.**
|
|
143
|
+
|
|
144
|
+
Tool preference:
|
|
145
|
+
|
|
146
|
+
1. **Debugger/REPL** — one breakpoint beats ten logs.
|
|
147
|
+
2. **Targeted logs** at boundaries that distinguish hypotheses.
|
|
148
|
+
3. Never "log everything and grep".
|
|
149
|
+
|
|
150
|
+
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
|
|
151
|
+
|
|
152
|
+
**Perf branch**: For performance regressions, establish a baseline measurement first (timing harness, profiler, query plan), then bisect. Measure first, fix second.
|
|
153
|
+
|
|
154
|
+
---
|
|
155
|
+
|
|
156
|
+
## Phase 5: Fix + Regression Test
|
|
157
|
+
|
|
158
|
+
### 5.1 Seam Assessment
|
|
159
|
+
|
|
160
|
+
Write the regression test **before the fix** — but only if there is a **correct seam**:
|
|
161
|
+
|
|
162
|
+
A correct seam exercises the **real bug pattern** at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers), a regression test there gives false confidence.
|
|
163
|
+
|
|
164
|
+
**If no correct seam exists, that itself is the finding.** Note it. The architecture is preventing the bug from being locked down.
|
|
165
|
+
|
|
166
|
+
### 5.2 If a correct seam exists
|
|
167
|
+
|
|
168
|
+
1. Turn the minimised repro into a failing test at that seam.
|
|
169
|
+
2. Watch it fail.
|
|
170
|
+
3. Apply the fix — **ONE change at a time**.
|
|
171
|
+
4. Watch it pass.
|
|
172
|
+
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
|
|
173
|
+
|
|
174
|
+
### 5.3 If Fix Doesn't Work
|
|
175
|
+
|
|
176
|
+
- Try #1 failed? → Return to Phase 3, form new hypothesis.
|
|
177
|
+
- Try #2 failed? → Return to Phase 1, re-check the loop.
|
|
178
|
+
- **If 3+ fixes failed: STOP.** This is an architectural problem, not a bug:
|
|
179
|
+
- Each fix reveals new problems in different places.
|
|
180
|
+
- Fixes require "massive refactoring" to implement.
|
|
181
|
+
- **Question the architecture, not the symptom.**
|
|
182
|
+
- Discuss with your human partner before attempting more fixes.
|
|
183
|
+
|
|
184
|
+
---
|
|
185
|
+
|
|
186
|
+
## Phase 6: Cleanup + Post-Mortem
|
|
187
|
+
|
|
188
|
+
Required before declaring done:
|
|
189
|
+
|
|
190
|
+
- [ ] Original repro no longer reproduces (re-run Phase 1 loop)
|
|
191
|
+
- [ ] Regression test passes (or absence of seam is documented)
|
|
192
|
+
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
|
193
|
+
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
|
194
|
+
- [ ] The correct hypothesis is stated in the commit/PR message
|
|
195
|
+
|
|
196
|
+
**Then ask: what would have prevented this bug?** If architectural change would have prevented it, note the specifics. You have more information now than when you started.
|
|
197
|
+
|
|
198
|
+
---
|
|
199
|
+
|
|
200
|
+
## Red Flags — STOP Immediately
|
|
201
|
+
|
|
202
|
+
If you catch yourself thinking:
|
|
203
|
+
|
|
204
|
+
| Thought | Reality |
|
|
205
|
+
| ---------------------------------------------- | ---------------------------------------------------------- |
|
|
206
|
+
| "Quick fix for now, investigate later" | First fix sets the pattern. Do it right. |
|
|
207
|
+
| "Just try changing X and see if it works" | Guessing. Build a loop instead (Phase 1). |
|
|
208
|
+
| "Add multiple changes, run tests" | Can't isolate what worked. One variable at a time. |
|
|
209
|
+
| "Skip the test, I'll verify manually" | Untested fixes don't stick. |
|
|
210
|
+
| "It's probably X, let me fix that" | Seeing symptoms ≠ understanding root cause. |
|
|
211
|
+
| "I don't fully understand but this might work" | Return to Phase 1. |
|
|
212
|
+
| "Reference too long, I'll adapt the pattern" | Partial understanding guarantees bugs. Read it completely. |
|
|
213
|
+
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem. Question the pattern. |
|
|
214
|
+
|
|
215
|
+
**ALL of these mean: STOP. Return to the earliest incomplete Phase.**
|
|
216
|
+
|
|
217
|
+
---
|
|
218
|
+
|
|
219
|
+
## Quick Reference
|
|
220
|
+
|
|
221
|
+
| Phase | Key Activities | Success Criteria |
|
|
222
|
+
| -------------------------- | --------------------------------------------------------- | ---------------------------------------------- |
|
|
223
|
+
| **1. Feedback Loop** | Build tight red/green signal for the bug | Deterministic, fast, agent-runnable |
|
|
224
|
+
| **2. Reproduce+Minimise** | Confirm + shrink to smallest load-bearing scenario | Every element is load-bearing |
|
|
225
|
+
| **3. Pattern+Hypothesise** | Compare working examples, rank 3-5 falsifiable hypotheses | Each hypothesis has a testable prediction |
|
|
226
|
+
| **4. Instrument** | One probe per prediction, tagged logs | Identify which hypothesis holds |
|
|
227
|
+
| **5. Fix+Regression** | Assess seam → test → single fix → verify | Bug resolved, test passes, original loop green |
|
|
228
|
+
| **6. Cleanup+Post-Mortem** | Remove instrumentation, document cause | Preventative insight captured |
|
|
229
|
+
|
|
230
|
+
---
|
|
231
|
+
|
|
232
|
+
## Supporting Techniques
|
|
233
|
+
|
|
234
|
+
- **Root Cause Tracing**: Trace bug backward through call stack to find original trigger. Where does the bad value originate? Keep tracing up.
|
|
235
|
+
- **Defense in Depth**: After fixing root cause, add validation at multiple layers so this class of bug can't recur.
|
|
236
|
+
- **Condition-Based Waiting**: Replace arbitrary timeouts (`sleep(5)`) with condition polling (`waitFor(selector)`).
|
package/src/index.tsx
CHANGED
package/src/ui/app.tsx
CHANGED
|
@@ -12,6 +12,8 @@ import { ChatPanel } from './chat'
|
|
|
12
12
|
import { InputBar } from './input'
|
|
13
13
|
import { ModelPicker } from './picker'
|
|
14
14
|
import { AgentFooter, type AgentEntry } from './agent-footer'
|
|
15
|
+
import { AgentViewDashboard } from '../agent-view/dashboard'
|
|
16
|
+
import type { AgentViewManager } from '../agent-view/agent-view-manager'
|
|
15
17
|
import { WorkflowProgress } from './workflow-progress.js'
|
|
16
18
|
import {
|
|
17
19
|
getCommand,
|
|
@@ -33,6 +35,7 @@ interface AppProps {
|
|
|
33
35
|
pluginManager?: PluginManager
|
|
34
36
|
version?: string
|
|
35
37
|
sessionId?: string
|
|
38
|
+
agentViewManager?: AgentViewManager
|
|
36
39
|
}
|
|
37
40
|
|
|
38
41
|
export interface ToolMeta {
|
|
@@ -121,6 +124,7 @@ export function App({
|
|
|
121
124
|
pluginManager,
|
|
122
125
|
version,
|
|
123
126
|
sessionId,
|
|
127
|
+
agentViewManager,
|
|
124
128
|
}: AppProps) {
|
|
125
129
|
const { t } = useI18n()
|
|
126
130
|
const PERMISSION_LABELS = useMemo<Record<PermissionMode, string>>(
|
|
@@ -139,6 +143,7 @@ export function App({
|
|
|
139
143
|
const [providerId, setProviderId] = useState(initialProvider || config.defaultProvider)
|
|
140
144
|
const [modelId, setModelId] = useState(initialModel || config.defaultModel)
|
|
141
145
|
const [pickerOpen, setPickerOpen] = useState(false)
|
|
146
|
+
const [agentViewOpen, setAgentViewOpen] = useState(false)
|
|
142
147
|
const [_sessionTitle, setSessionTitle] = useState('')
|
|
143
148
|
const [_fastMode, setFastMode] = useState(false)
|
|
144
149
|
const [_effort, setEffort] = useState('high')
|
|
@@ -580,99 +585,115 @@ export function App({
|
|
|
580
585
|
{/* Workflow progress — auto-detects active workflows, renders nothing when idle */}
|
|
581
586
|
<WorkflowProgress />
|
|
582
587
|
|
|
583
|
-
{/*
|
|
584
|
-
|
|
585
|
-
|
|
586
|
-
|
|
587
|
-
|
|
588
|
-
|
|
589
|
-
config={config}
|
|
590
|
-
currentProvider={providerId}
|
|
591
|
-
currentModel={modelId}
|
|
592
|
-
onSelect={(newProvider, newModel) => {
|
|
593
|
-
engine.switchProvider(newProvider, newModel)
|
|
594
|
-
setProviderId(newProvider)
|
|
595
|
-
setModelId(newModel)
|
|
596
|
-
setPickerOpen(false)
|
|
597
|
-
setMessages((prev) => [
|
|
598
|
-
...prev,
|
|
599
|
-
{ role: 'system', content: `✓ Switched to ${newProvider}/${newModel}` },
|
|
600
|
-
])
|
|
601
|
-
}}
|
|
602
|
-
onClose={() => setPickerOpen(false)}
|
|
588
|
+
{/* Agent View Dashboard — Ctrl+G overlay (replaces chat + input) */}
|
|
589
|
+
{agentViewOpen && agentViewManager ? (
|
|
590
|
+
<AgentViewDashboard
|
|
591
|
+
manager={agentViewManager}
|
|
592
|
+
onAttach={() => {}}
|
|
593
|
+
onExit={() => setAgentViewOpen(false)}
|
|
603
594
|
/>
|
|
604
595
|
) : (
|
|
605
|
-
|
|
606
|
-
|
|
607
|
-
<
|
|
608
|
-
|
|
609
|
-
|
|
610
|
-
|
|
611
|
-
|
|
612
|
-
|
|
613
|
-
|
|
614
|
-
|
|
615
|
-
|
|
616
|
-
|
|
617
|
-
|
|
618
|
-
|
|
619
|
-
|
|
620
|
-
|
|
621
|
-
|
|
622
|
-
|
|
623
|
-
|
|
624
|
-
|
|
625
|
-
|
|
626
|
-
|
|
627
|
-
|
|
628
|
-
|
|
629
|
-
|
|
630
|
-
|
|
631
|
-
|
|
596
|
+
<>
|
|
597
|
+
{/* Chat panel */}
|
|
598
|
+
<ChatPanel messages={messages} focusMode={focusMode} />
|
|
599
|
+
|
|
600
|
+
{/* Input with separator lines */}
|
|
601
|
+
{pickerOpen ? (
|
|
602
|
+
<ModelPicker
|
|
603
|
+
config={config}
|
|
604
|
+
currentProvider={providerId}
|
|
605
|
+
currentModel={modelId}
|
|
606
|
+
onSelect={(newProvider, newModel) => {
|
|
607
|
+
engine.switchProvider(newProvider, newModel)
|
|
608
|
+
setProviderId(newProvider)
|
|
609
|
+
setModelId(newModel)
|
|
610
|
+
setPickerOpen(false)
|
|
611
|
+
setMessages((prev) => [
|
|
612
|
+
...prev,
|
|
613
|
+
{ role: 'system', content: `✓ Switched to ${newProvider}/${newModel}` },
|
|
614
|
+
])
|
|
615
|
+
}}
|
|
616
|
+
onClose={() => setPickerOpen(false)}
|
|
617
|
+
/>
|
|
618
|
+
) : (
|
|
619
|
+
/* Input bar (hidden when picker is open) */
|
|
620
|
+
<Box flexDirection="column">
|
|
621
|
+
<Text dimColor>──────────────────────────────</Text>
|
|
622
|
+
<InputBar
|
|
623
|
+
onSubmit={handleSubmit}
|
|
624
|
+
isLoading={isLoading}
|
|
625
|
+
onTogglePicker={() => setPickerOpen((prev) => !prev)}
|
|
626
|
+
onToggleFocus={() => setFocusMode((prev) => !prev)}
|
|
627
|
+
onToggleExpand={() => {
|
|
628
|
+
setMessages((prev) => {
|
|
629
|
+
const msgs = [...prev]
|
|
630
|
+
for (let i = msgs.length - 1; i >= 0; i--) {
|
|
631
|
+
if (msgs[i]?.toolMeta) {
|
|
632
|
+
const meta = msgs[i]!.toolMeta!
|
|
633
|
+
if (meta.collapsed) {
|
|
634
|
+
msgs[i] = {
|
|
635
|
+
...msgs[i]!,
|
|
636
|
+
content: `🔧 ${meta.name}: ${meta.input}\n📋 Result: ${meta.output || '(pending)'}`,
|
|
637
|
+
toolMeta: { ...meta, collapsed: false },
|
|
638
|
+
}
|
|
639
|
+
} else {
|
|
640
|
+
const short =
|
|
641
|
+
meta.input.length > 50 ? meta.input.slice(0, 50) + '...' : meta.input
|
|
642
|
+
msgs[i] = {
|
|
643
|
+
...msgs[i]!,
|
|
644
|
+
content: `⏺ ${meta.name} · ${short} (Ctrl+O to expand)`,
|
|
645
|
+
toolMeta: { ...meta, collapsed: true },
|
|
646
|
+
}
|
|
647
|
+
}
|
|
648
|
+
break
|
|
632
649
|
}
|
|
633
650
|
}
|
|
634
|
-
|
|
651
|
+
return msgs
|
|
652
|
+
})
|
|
653
|
+
}}
|
|
654
|
+
onCyclePermission={() => {
|
|
655
|
+
setPermissionMode((prev) => {
|
|
656
|
+
const idx = PERMISSION_MODES.indexOf(prev)
|
|
657
|
+
const next = PERMISSION_MODES[(idx + 1) % PERMISSION_MODES.length]!
|
|
658
|
+
engine.getPermission().setMode(next)
|
|
659
|
+
return next
|
|
660
|
+
})
|
|
661
|
+
}}
|
|
662
|
+
onCancel={() => {
|
|
663
|
+
if (abortRef.current) {
|
|
664
|
+
abortRef.current.abort()
|
|
635
665
|
}
|
|
636
|
-
}
|
|
637
|
-
|
|
638
|
-
|
|
639
|
-
|
|
640
|
-
|
|
641
|
-
|
|
642
|
-
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
|
|
646
|
-
|
|
647
|
-
}
|
|
648
|
-
onCancel={() => {
|
|
649
|
-
if (abortRef.current) {
|
|
650
|
-
abortRef.current.abort()
|
|
651
|
-
}
|
|
652
|
-
}}
|
|
666
|
+
}}
|
|
667
|
+
onToggleAgentView={() => setAgentViewOpen((prev) => !prev)}
|
|
668
|
+
/>
|
|
669
|
+
<Text dimColor>──────────────────────────────</Text>
|
|
670
|
+
</Box>
|
|
671
|
+
)}
|
|
672
|
+
|
|
673
|
+
{/* Agent status footer — shows running background agents */}
|
|
674
|
+
<AgentFooter
|
|
675
|
+
agents={Object.values(runningAgents)}
|
|
676
|
+
gitBranch={gitBranch}
|
|
677
|
+
tick={agentTick}
|
|
653
678
|
/>
|
|
654
|
-
<Text dimColor>──────────────────────────────</Text>
|
|
655
|
-
</Box>
|
|
656
|
-
)}
|
|
657
679
|
|
|
658
|
-
|
|
659
|
-
|
|
660
|
-
|
|
661
|
-
|
|
662
|
-
|
|
663
|
-
|
|
664
|
-
|
|
665
|
-
<
|
|
680
|
+
{/* Status line — Claude Code style */}
|
|
681
|
+
<Box marginTop={1} flexDirection="column">
|
|
682
|
+
{goalText && (
|
|
683
|
+
<Box>
|
|
684
|
+
<Text color="green">🎯 Goal: {goalText}</Text>
|
|
685
|
+
</Box>
|
|
686
|
+
)}
|
|
687
|
+
<Box flexDirection="row">
|
|
688
|
+
<Text color={PERMISSION_COLORS[permissionMode]}>
|
|
689
|
+
⏵⏵ {PERMISSION_LABELS[permissionMode]} ({t('ui.status.shift_tab_cycle')})
|
|
690
|
+
</Text>
|
|
691
|
+
<Text dimColor> · {t('ui.status.esc_to_interrupt')}</Text>
|
|
692
|
+
<Text dimColor> · {t('ui.status.left_for_agents')}</Text>
|
|
693
|
+
</Box>
|
|
666
694
|
</Box>
|
|
667
|
-
|
|
668
|
-
|
|
669
|
-
<Text color={PERMISSION_COLORS[permissionMode]}>
|
|
670
|
-
⏵⏵ {PERMISSION_LABELS[permissionMode]} ({t('ui.status.shift_tab_cycle')})
|
|
671
|
-
</Text>
|
|
672
|
-
<Text dimColor> · {t('ui.status.esc_to_interrupt')}</Text>
|
|
673
|
-
<Text dimColor> · {t('ui.status.left_for_agents')}</Text>
|
|
674
|
-
</Box>
|
|
675
|
-
</Box>
|
|
695
|
+
</>
|
|
696
|
+
)}
|
|
676
697
|
</Box>
|
|
677
698
|
)
|
|
678
699
|
}
|
package/src/ui/input.tsx
CHANGED
|
@@ -15,6 +15,8 @@ interface InputBarProps {
|
|
|
15
15
|
onToggleFocus?: () => void
|
|
16
16
|
/** Ctrl+O → expand last tool call */
|
|
17
17
|
onToggleExpand?: () => void
|
|
18
|
+
/** Ctrl+G → toggle agent view dashboard */
|
|
19
|
+
onToggleAgentView?: () => void
|
|
18
20
|
/** Shift+Tab → cycle permission mode */
|
|
19
21
|
onCyclePermission?: () => void
|
|
20
22
|
/** Escape → cancel loading (when input is empty) */
|
|
@@ -71,6 +73,7 @@ export function InputBar({
|
|
|
71
73
|
onTogglePicker,
|
|
72
74
|
onToggleFocus,
|
|
73
75
|
onToggleExpand,
|
|
76
|
+
onToggleAgentView,
|
|
74
77
|
onCyclePermission,
|
|
75
78
|
onCancel,
|
|
76
79
|
}: InputBarProps) {
|
|
@@ -185,6 +188,12 @@ export function InputBar({
|
|
|
185
188
|
onToggleExpand?.()
|
|
186
189
|
return
|
|
187
190
|
}
|
|
191
|
+
// Ctrl+G → toggle agent view dashboard
|
|
192
|
+
if (key.ctrl && input === 'g') {
|
|
193
|
+
setValue(valueBeforeShortcut.current)
|
|
194
|
+
onToggleAgentView?.()
|
|
195
|
+
return
|
|
196
|
+
}
|
|
188
197
|
|
|
189
198
|
if (vimEngine.current.mode !== 'normal') return
|
|
190
199
|
|