@zhuxixi/pi-agent-board 0.5.2 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/CHANGELOG.md +32 -0
  2. package/README.md +41 -3
  3. package/docs/superpowers/plans/2026-09-03-code-refs-pr-backlink-narrow.md +551 -0
  4. package/docs/superpowers/plans/2026-09-04-evidence-outputpreview.md +209 -0
  5. package/docs/superpowers/plans/2026-09-04-warm-host-reclaim.md +796 -0
  6. package/docs/superpowers/plans/2026-09-05-issue-13-drainnextfollowup-pty-probe.md +114 -0
  7. package/docs/superpowers/plans/2026-09-05-issue-38-windows-wezterm-ime-cursor.md +73 -0
  8. package/docs/superpowers/plans/2026-09-05-issue-39-truncate-codepoint-boundary.md +143 -0
  9. package/docs/superpowers/plans/2026-09-05-issue-61-mention-fallback-guards.md +226 -0
  10. package/docs/superpowers/plans/2026-09-05-issue-63-flaky-manual-completion.md +87 -0
  11. package/docs/superpowers/plans/2026-09-05-issue-64-changelog-release-helper.md +53 -0
  12. package/docs/superpowers/plans/2026-09-05-pty-host-stacking-sock-race.md +731 -0
  13. package/docs/superpowers/specs/2026-08-29-code-refs-badges-design.md +1 -1
  14. package/docs/superpowers/specs/2026-09-03-code-refs-pr-backlink-narrow-design.md +92 -0
  15. package/docs/superpowers/specs/2026-09-04-evidence-outputpreview-design.md +50 -0
  16. package/docs/superpowers/specs/2026-09-04-warm-host-reclaim-design.md +106 -0
  17. package/docs/superpowers/specs/2026-09-05-issue-13-drainnextfollowup-pty-probe-design.md +64 -0
  18. package/docs/superpowers/specs/2026-09-05-issue-38-windows-wezterm-ime-design.md +48 -0
  19. package/docs/superpowers/specs/2026-09-05-issue-39-truncate-codepoint-boundary-design.md +64 -0
  20. package/docs/superpowers/specs/2026-09-05-issue-61-mention-fallback-design.md +71 -0
  21. package/docs/superpowers/specs/2026-09-05-issue-63-flaky-manual-completion-design.md +49 -0
  22. package/docs/superpowers/specs/2026-09-05-issue-64-changelog-helper-design.md +76 -0
  23. package/docs/superpowers/specs/2026-09-05-pty-host-stacking-sock-race-design.md +510 -0
  24. package/package.json +83 -81
  25. package/runner/job-runner.mjs +2 -2
  26. package/runner/pty-runner.mjs +573 -2
  27. package/runner/state-runner.mjs +3 -0
  28. package/runner/title-runner.mjs +1 -1
  29. package/scripts/release_helper.mjs +277 -0
  30. package/src/commands/agent-board.ts +38 -35
  31. package/src/commands/attach-decision.mjs +66 -0
  32. package/src/commands/attach-flow.ts +45 -39
  33. package/src/core/code-refs.mjs +85 -33
  34. package/src/core/evidence.mjs +2 -2
  35. package/src/core/heuristics.mjs +40 -2
  36. package/src/core/host-coordination.mjs +159 -0
  37. package/src/core/host-crash.mjs +43 -3
  38. package/src/core/host-probe.mjs +196 -0
  39. package/src/core/launch.mjs +3 -1
  40. package/src/core/locks.mjs +196 -1
  41. package/src/core/paths.mjs +24 -0
  42. package/src/core/store.mjs +164 -5
  43. package/src/core/types.mjs +17 -1
  44. package/src/core/warm-host-sweeper.mjs +150 -0
  45. package/src/index.ts +29 -1
  46. package/src/runtime/service.mjs +967 -109
  47. package/src/ui/dashboard-decisions.mjs +55 -0
  48. package/src/ui/dashboard.ts +26 -12
@@ -0,0 +1,114 @@
1
+ # Issue #13 Drain-Path PTY Probe Refresh — Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Stop `drainNextFollowUp` (reachable from the 700ms reconcile poll) from forcing a real `pty.spawn` probe on every cycle; switch it to cached-probe semantics and lock that in with a red-green test.
6
+
7
+ **Architecture:** One-line production change (`ptySupport({ refresh: true })` → `ptySupport()` in `drainNextFollowUp`) + one new test in `test/service.test.mjs` driving the public `reconcile()` API with an injected `ptySupport` spy, modeled on the existing "ensureHost probes PTY support with TTL cache" test and the "busy replies queue and drain when idle" row-state pattern.
8
+
9
+ **Tech Stack:** Node built-in test runner; existing `createService` DI seams (`ptySupport`, `launch`).
10
+
11
+ ## Global Constraints
12
+
13
+ - Production delta is exactly the one-line probe switch in `src/runtime/service.mjs` `drainNextFollowUp` (spec A4).
14
+ - `dispatch` (L616) and `reply` (L649) keep `refresh: true` — explicit user actions.
15
+ - No changes to `ptySpawnSupported` / `shouldProbePtySupport`.
16
+ - Test command: `npm test` (= `node --test test/*.test.mjs`).
17
+
18
+ **Execution mode:** inline (executing-plans) — single task, one-line production diff + one test.
19
+
20
+ ---
21
+
22
+ ### Task 1: Red-green test + one-line probe switch
23
+
24
+ **Files:**
25
+ - Modify: `src/runtime/service.mjs:496` (inside `drainNextFollowUp`)
26
+ - Modify: `test/service.test.mjs` (append new test after "ensureHost probes PTY support with TTL cache, not forced refresh", ~L958)
27
+
28
+ **Interfaces:**
29
+ - Consumes: `createService` opts `{ ptySupport?: (opts?) => {ok, reason?}, launch?: typeof launchRun }`; public `svc.reconcile()`, `svc.queueFollowUp(viewId, text)`; helpers `freshRoot()`, `createView(root, {id, name, cwd})`, `readState`/`writeState`, `rmSync`.
30
+ - Produces: nothing consumed by other tasks.
31
+
32
+ - [ ] **Step 1: Write the failing test (red)**
33
+
34
+ Append to `test/service.test.mjs` (after the "ensureHost probes PTY support with TTL cache, not forced refresh" test):
35
+
36
+ ```js
37
+ test("reconcile auto-drain uses the cached PTY probe, never forced refresh", () => {
38
+ const root = freshRoot();
39
+ try {
40
+ createView(root, { id: "v1", name: "a", cwd: "/r" });
41
+ const st = readState(root, "v1");
42
+ st.semanticState = "idle";
43
+ st.processState = "exited";
44
+ writeState(root, st);
45
+ const probeCalls = [];
46
+ const svc = service(root, {
47
+ ptySupport: (opts = {}) => {
48
+ probeCalls.push(opts);
49
+ return { ok: false, reason: "test" };
50
+ },
51
+ launch: () => ({ pid: null, configPath: "/no/config.json" }),
52
+ });
53
+ const queued = svc.queueFollowUp("v1", "next step");
54
+ assert.equal(queued.ok, true);
55
+ const fixed = svc.reconcile();
56
+ assert.ok(fixed >= 1, "reconcile drained the queued follow-up");
57
+ assert.ok(probeCalls.length >= 1, "drain path probed ptySupport");
58
+ for (const opts of probeCalls) {
59
+ assert.notEqual(opts?.refresh, true, "reconcile→drain must not force ptySupport refresh");
60
+ }
61
+ } finally {
62
+ rmSync(root, { recursive: true, force: true });
63
+ }
64
+ });
65
+ ```
66
+
67
+ - [ ] **Step 2: Run it, verify it fails against current production code**
68
+
69
+ Run: `node --test test/service.test.mjs`
70
+ Expected: FAIL — `reconcile→drain must not force ptySupport refresh` (production still passes `{refresh: true}`).
71
+
72
+ - [ ] **Step 3: The one-line fix (green)**
73
+
74
+ In `src/runtime/service.mjs` `drainNextFollowUp` (~L496), change:
75
+
76
+ ```js
77
+ const pty = ptySupport({ refresh: true });
78
+ ```
79
+
80
+ to:
81
+
82
+ ```js
83
+ const pty = ptySupport();
84
+ ```
85
+
86
+ (Only inside `drainNextFollowUp`. The `dispatch` and `reply` occurrences keep `{ refresh: true }`.)
87
+
88
+ - [ ] **Step 4: Run the test file, verify pass**
89
+
90
+ Run: `node --test test/service.test.mjs`
91
+ Expected: PASS including the new test.
92
+
93
+ - [ ] **Step 5: Full suite (A2/A3)**
94
+
95
+ Run: `npm test`
96
+ Expected: all green.
97
+
98
+ - [ ] **Step 6: Verify A4 — scope**
99
+
100
+ Run: `git diff main -- src/ test/`
101
+ Expected: exactly one line changed in `src/runtime/service.mjs`, one test appended in `test/service.test.mjs`.
102
+
103
+ - [ ] **Step 7: Commit**
104
+
105
+ ```bash
106
+ git add src/runtime/service.mjs test/service.test.mjs
107
+ git commit -m "perf: use cached PTY probe on the reconcile drain path (issue #13)"
108
+ ```
109
+
110
+ ## Self-Review
111
+
112
+ - Spec coverage: A1 (Steps 1-2 red, 3-4 green), A2/A3 (Step 5), A4 (Step 6). Covered.
113
+ - Placeholder scan: exact code shown for both edits.
114
+ - Type consistency: spy signature matches `createService` opts type; `launch` fake matches the pattern used by the existing "busy replies queue and drain when idle" test.
@@ -0,0 +1,73 @@
1
+ # Issue #38 Windows WezTerm IME Docs — Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Document that Windows WezTerm attach sessions need a visible hardware cursor (`PI_HARDWARE_CURSOR=1` / `"showHardwareCursor": true`) for IME candidate tracking.
6
+
7
+ **Architecture:** One new Troubleshooting subsection in README.md, placed after "### Attach is slow or keeps reconnecting".
8
+
9
+ **Tech Stack:** Markdown only.
10
+
11
+ ## Global Constraints
12
+
13
+ - Docs-only: README.md is the only non-process file touched (spec A2).
14
+ - Both fixes documented with scope guidance; cosmetic trade-off stated.
15
+ - No upstream comment on earendil-works/pi#5200 (owner's voice).
16
+
17
+ **Execution mode:** inline (executing-plans).
18
+
19
+ ---
20
+
21
+ ### Task 1: Add the Troubleshooting subsection
22
+
23
+ **Files:**
24
+ - Modify: `README.md` (Troubleshooting section, after the "### Attach is slow or keeps reconnecting" block)
25
+
26
+ - [ ] **Step 1: Insert the subsection**
27
+
28
+ Insert after the "### Attach is slow or keeps reconnecting" paragraph, before "### Start & attach falls back to background":
29
+
30
+ ````markdown
31
+ ### IME candidate window is stuck at the window edge (Windows WezTerm)
32
+
33
+ On Windows WezTerm with a WSL2 backend, the IME candidate window may stay
34
+ pinned to the right edge instead of following the text cursor in an attached
35
+ session. Windows WezTerm only tracks the IME candidate position from the
36
+ visible hardware cursor, and Pi hides the hardware cursor by default — the
37
+ block cursor you see in the editor is drawn content, not the hardware cursor.
38
+ Linux terminals are not affected.
39
+
40
+ Make the hardware cursor visible, either way:
41
+
42
+ ```bash
43
+ export PI_HARDWARE_CURSOR=1 # machine-local, e.g. ~/.zshrc.local
44
+ ```
45
+
46
+ or set `"showHardwareCursor": true` in Pi's `settings.json` (syncs across
47
+ machines if the config is version-controlled; harmless on Linux).
48
+
49
+ Trade-off: the real terminal cursor becomes visible inside the TUI. This is
50
+ cosmetic only.
51
+ ````
52
+
53
+ - [ ] **Step 2: Verify A1 (content presence)**
54
+
55
+ Run: `grep -n "PI_HARDWARE_CURSOR=1\|showHardwareCursor\|IME candidate window is stuck" README.md`
56
+ Expected: 3+ hits inside the Troubleshooting section.
57
+
58
+ - [ ] **Step 3: Verify A2 (scope)**
59
+
60
+ Run: `git diff main --stat`
61
+ Expected: README.md + process docs under docs/superpowers/ only.
62
+
63
+ - [ ] **Step 4: Commit**
64
+
65
+ ```bash
66
+ git add README.md
67
+ git commit -m "docs: Windows WezTerm IME needs a visible hardware cursor (issue #38)"
68
+ ```
69
+
70
+ ## Self-Review
71
+
72
+ - Spec coverage: A1 (Step 2), A2 (Step 3). Covered.
73
+ - Placeholder scan: full markdown shown.
@@ -0,0 +1,143 @@
1
+ # Issue #39 Truncate Surrogate-Safe Cut — Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Make `truncate()` in `src/core/heuristics.mjs` never split a surrogate pair (cut-point back-off within the UTF-16 unit budget) and never emit lone surrogates, eliminating the U+FFFD → width-mismatch → wrap → stacked-frames chain at its repo-side root.
6
+
7
+ **Architecture:** Pure-function change inside `truncate` only: check whether the unit before the cut point is a high surrogate whose low surrogate is being dropped; if so, back the cut off by one unit. Strip any lone surrogates from the returned string on both paths via a small private regex helper.
8
+
9
+ **Tech Stack:** Node built-in test runner; no new dependencies.
10
+
11
+ ## Global Constraints
12
+
13
+ - Keep the UTF-16 unit budget semantics of `n` (spec Decision 1) — do NOT switch to code-point counting.
14
+ - Only `src/core/heuristics.mjs` (truncate + private helper) and `test/heuristics.test.mjs` change (spec A5).
15
+ - Existing tests must pass unmodified (spec A3).
16
+ - Test command: `npm test`.
17
+
18
+ **Execution mode:** inline (executing-plans).
19
+
20
+ ---
21
+
22
+ ### Task 1: Red tests at the surrogate boundaries
23
+
24
+ **Files:**
25
+ - Modify: `test/heuristics.test.mjs` (extend the existing "truncate adds ellipsis" test block or add a new test after it)
26
+
27
+ **Interfaces:**
28
+ - Consumes: `truncate(s, n)` exported from `src/core/heuristics.mjs`.
29
+ - Produces: nothing consumed by other tasks.
30
+
31
+ - [ ] **Step 1: Write the failing tests (red)**
32
+
33
+ Replace the existing block:
34
+
35
+ ```js
36
+ test("truncate adds ellipsis", () => {
37
+ assert.equal(truncate("hello", 10), "hello");
38
+ assert.equal(truncate("hello world", 5), "hell…");
39
+ });
40
+ ```
41
+
42
+ with:
43
+
44
+ ```js
45
+ test("truncate adds ellipsis", () => {
46
+ assert.equal(truncate("hello", 10), "hello");
47
+ assert.equal(truncate("hello world", 5), "hell…");
48
+ assert.equal(truncate("一二三四五六", 5), "一二三四…");
49
+ });
50
+
51
+ test("truncate never splits a surrogate pair", () => {
52
+ // Cut point lands between the high and low surrogate of 👍: back off.
53
+ assert.equal(truncate("a👍b", 3), "a…");
54
+ // Long string whose 80-unit cut lands inside the emoji (issue #39 repro shape).
55
+ const s = "x".repeat(59) + "完成 — LGTM " + "👍" + "y".repeat(10);
56
+ const out = truncate(s, 80);
57
+ assert.ok(!/[\uD800-\uDBFF]$/.test(out.replace(/…$/, "")), "no trailing lone high surrogate before ellipsis");
58
+ // The pair survives whole when it fits the budget.
59
+ assert.equal(truncate("👍", 2), "👍");
60
+ assert.equal(truncate("a👍", 3), "a👍");
61
+ });
62
+
63
+ test("truncate strips lone surrogates from input", () => {
64
+ assert.equal(truncate("ab\ud83d", 10), "ab"); // lone high surrogate, short path
65
+ assert.equal(truncate("ab\ud83dcd", 4), "abc…"); // lone high surrogate, truncated path
66
+ assert.equal(truncate("\udc4dab", 10), "ab"); // lone low surrogate, short path
67
+ });
68
+ ```
69
+
70
+ Note on `"ab\ud83dcd"` with n=4: units are a,b,\ud83d,c,d (length 5 > 4) → cut end=3 → slice = "ab\ud83d" → the high surrogate is last-but-0 with no following unit in the *kept* slice; input's `\ud83d` is followed by `c` (not a low surrogate), so it is a lone surrogate in the input and must be stripped → expected "ab…". If the implementation instead computes end=3, sees charCodeAt(2) is a high surrogate and charCodeAt(3)="c" is not a low surrogate, no pair back-off applies (no pair to preserve) — the strip helper removes it → "ab…". Both reasonings converge; assert the observable "ab…".
71
+
72
+ - [ ] **Step 2: Run, verify red**
73
+
74
+ Run: `node --test test/heuristics.test.mjs`
75
+ Expected: the two new tests FAIL (current implementation splits pairs / keeps lone surrogates), "truncate adds ellipsis" still passes.
76
+
77
+ ### Task 2: Implement the surrogate-safe truncate
78
+
79
+ **Files:**
80
+ - Modify: `src/core/heuristics.mjs` (replace `truncate`)
81
+
82
+ - [ ] **Step 3: New implementation**
83
+
84
+ Replace the existing `truncate` function with:
85
+
86
+ ```js
87
+ const LONE_HIGH_SURROGATE_RE = /[\uD800-\uDBFF](?![\uDC00-\uDFFF])/g;
88
+ const LONE_LOW_SURROGATE_RE = /(?<![\uD800-\uDBFF])[\uDC00-\uDFFF]/g;
89
+
90
+ /** @param {string} s */
91
+ function stripLoneSurrogates(s) {
92
+ if (!s) return s;
93
+ return s.replace(LONE_HIGH_SURROGATE_RE, "").replace(LONE_LOW_SURROGATE_RE, "");
94
+ }
95
+
96
+ /**
97
+ * Truncate to `n` chars with an ellipsis (counts characters, not display width).
98
+ * The budget `n` is in UTF-16 units. The cut never splits a surrogate pair
99
+ * (backs off one unit when it would), and lone surrogates in the input are
100
+ * stripped from the result — a lone surrogate renders as U+FFFD, whose
101
+ * terminal width can disagree with the computed width and misalign rows.
102
+ * @param {string} s
103
+ * @param {number} n
104
+ * @returns {string}
105
+ */
106
+ export function truncate(s, n) {
107
+ const str = String(s ?? "");
108
+ if (str.length <= n) return stripLoneSurrogates(str);
109
+ let end = Math.max(0, n - 1);
110
+ const lastUnit = str.charCodeAt(end - 1);
111
+ const nextUnit = str.charCodeAt(end);
112
+ if (lastUnit >= 0xd800 && lastUnit <= 0xdbff && nextUnit >= 0xdc00 && nextUnit <= 0xdfff) end -= 1;
113
+ return `${stripLoneSurrogates(str.slice(0, end))}…`;
114
+ }
115
+ ```
116
+
117
+ - [ ] **Step 4: Run the file, verify green**
118
+
119
+ Run: `node --test test/heuristics.test.mjs`
120
+ Expected: all pass, including the two new tests and the unchanged originals.
121
+
122
+ - [ ] **Step 5: Full suite (A4)**
123
+
124
+ Run: `npm test`
125
+ Expected: all green.
126
+
127
+ - [ ] **Step 6: Verify A5 scope**
128
+
129
+ Run: `git diff main -- src/ test/`
130
+ Expected: only `src/core/heuristics.mjs` and `test/heuristics.test.mjs`.
131
+
132
+ - [ ] **Step 7: Commit**
133
+
134
+ ```bash
135
+ git add src/core/heuristics.mjs test/heuristics.test.mjs
136
+ git commit -m "fix: truncate never splits surrogate pairs or emits lone surrogates (issue #39)"
137
+ ```
138
+
139
+ ## Self-Review
140
+
141
+ - Spec coverage: A1 (Task 1 red + Task 2), A2 (lone-surrogate tests), A3 (unchanged originals + CJK case), A4 (Step 5), A5 (Step 6). Covered.
142
+ - Placeholder scan: full code shown.
143
+ - Type consistency: `stripLoneSurrogates` string→string; `charCodeAt` unit checks match the surrogate ranges used in the regexes.
@@ -0,0 +1,226 @@
1
+ # Issue #61 Mention Fallback Guards — Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Stop the Rule 6 bare `#N` mention fallback from lifting placeholder numbers: left-boundary guard, inline-code-span exclusion, PR-context kind-aware skip, plus a documented providers.json example for internal-platform CLIs.
6
+
7
+ **Architecture:** Three guards inside `mentionFallback`/`MENTION_RE` (pure regex logic), one README example. Tests go through the public `extractCodeRefs` seam like all existing mention tests.
8
+
9
+ **Tech Stack:** Node built-in test runner; no new dependencies.
10
+
11
+ ## Global Constraints
12
+
13
+ - Issue-only fallback stays (no PR mention candidate fabrication) — spec non-goals.
14
+ - Existing mention tests unchanged (spec A5).
15
+ - Only `src/core/code-refs.mjs`, `README.md` (Evidence section), `test/code-refs-extract.test.mjs` change (spec A7).
16
+ - Test command: `npm test`.
17
+
18
+ **Execution mode:** inline (executing-plans).
19
+
20
+ ---
21
+
22
+ ### Task 1: Red tests for the three guards + composite repro
23
+
24
+ **Files:**
25
+ - Modify: `test/code-refs-extract.test.mjs` (after the existing "mention fallback yields no winner" test, ~L283)
26
+
27
+ - [ ] **Step 1: Add failing tests**
28
+
29
+ ```js
30
+ test("mention fallback ignores ##N heading forms and word-prefixed #N (issue #61)", () => {
31
+ const result = extractCodeRefs(
32
+ { commands: [], assistantTexts: ["##1378", "##1378", "##1378", "a#1378", "b#99"], worktreePath: null, branch: null, repoUrl: null, host: null },
33
+ github()
34
+ );
35
+ assert.equal(result.issue, null);
36
+ assert.equal(result.pr, null);
37
+ });
38
+
39
+ test("mention fallback ignores #N inside inline code spans (issue #61)", () => {
40
+ const result = extractCodeRefs(
41
+ { commands: [], assistantTexts: ["`monitor pr #1378`", "`#1378` x3", "see `#1378`"], worktreePath: null, branch: null, repoUrl: null, host: null },
42
+ github()
43
+ );
44
+ assert.equal(result.issue, null);
45
+ });
46
+
47
+ test("mention fallback does not count pr-context #N toward issue candidates (issue #61)", () => {
48
+ const result = extractCodeRefs(
49
+ { commands: [], assistantTexts: ["monitor pr #1378", "monitor pr #1378", "monitor pr #1378", "monitor pr #1378", "monitor pr #1378"], worktreePath: null, branch: null, repoUrl: null, host: null },
50
+ github()
51
+ );
52
+ assert.equal(result.issue, null);
53
+ });
54
+
55
+ test("issue #61 repro shape: guarded placeholder loses, real issue prose wins", () => {
56
+ const result = extractCodeRefs(
57
+ {
58
+ commands: [],
59
+ assistantTexts: [
60
+ "设置 IID=\"#1378\" 会让 \"#$IID\" 展开成 \"##1378\" → 404", // ##1378 heading-guard form
61
+ "`monitor pr #1378` 文档示例", // code-span form
62
+ "monitor pr #1378 用户独立调用", // pr-context form
63
+ "正在处理 issue #18", "issue #18 的修复", "更新 #18 状态", // real issue prose ×3
64
+ ],
65
+ worktreePath: null, branch: null, repoUrl: null, host: null,
66
+ },
67
+ github()
68
+ );
69
+ assert.equal(result.issue.number, 18);
70
+ assert.equal(result.issue.strength, "mention");
71
+ });
72
+ ```
73
+
74
+ Note on the heading test: `##1378` — with the guard, the second `#` is preceded by `#` → no match; the first `#` is followed by `#` not digits → no match. `a#1378`/`b#99`: preceded by word char → no match. Expected result with OLD code: `#1378` counts 3× (from the `##1378` lines' second hash) + more → a winner exists → test fails (red).
75
+
76
+ - [ ] **Step 2: Run, verify red**
77
+
78
+ Run: `node --test test/code-refs-extract.test.mjs`
79
+ Expected: the new tests fail against current code; existing tests green.
80
+
81
+ ### Task 2: Implement the guards
82
+
83
+ **Files:**
84
+ - Modify: `src/core/code-refs.mjs` (MENTION_RE L549; mentionFallback L767)
85
+
86
+ - [ ] **Step 3: Code changes**
87
+
88
+ Replace:
89
+
90
+ ```js
91
+ /** Bare `#N` mentions. */
92
+ const MENTION_RE = /#(\d{1,7})/g;
93
+ ```
94
+
95
+ with:
96
+
97
+ ```js
98
+ /**
99
+ * Bare `#N` mentions. The lookbehind rejects a preceding `#` (markdown
100
+ * headings, `##N` typos) or word character (`a#1`) — those are not issue
101
+ * references (issue #61).
102
+ */
103
+ const MENTION_RE = /(?<![#\w])#(\d{1,7})/g;
104
+ /** Inline code span (single-line) — excluded from mention counting. */
105
+ const INLINE_CODE_SPAN_RE = /`[^`\n]*`/g;
106
+ /** Explicit PR-context tokens immediately before a `#N` match. */
107
+ const PR_CONTEXT_RE = /(?:^|\W)(?:pr|pull|mr|merge)s?(?:\W|$)/i;
108
+ ```
109
+
110
+ In `mentionFallback`, replace the counting loop:
111
+
112
+ ```js
113
+ for (let i = start; i < assistantTexts.length; i++) {
114
+ const text = assistantTexts[i];
115
+ if (typeof text !== "string") continue;
116
+ for (const m of text.matchAll(MENTION_RE)) {
117
+ const n = Number(m[1]);
118
+ counts.set(n, (counts.get(n) ?? 0) + 1);
119
+ lastIndex.set(n, baseIndex + i);
120
+ }
121
+ }
122
+ ```
123
+
124
+ with:
125
+
126
+ ```js
127
+ for (let i = start; i < assistantTexts.length; i++) {
128
+ const text = assistantTexts[i];
129
+ if (typeof text !== "string") continue;
130
+ // Doc examples inside inline code spans are not real references; strip
131
+ // them before counting (issue #61).
132
+ const stripped = text.replace(INLINE_CODE_SPAN_RE, " ");
133
+ for (const m of stripped.matchAll(MENTION_RE)) {
134
+ // `monitor pr #N` is explicitly PR context: do not count it toward
135
+ // the issue fallback (issue #61).
136
+ const before = stripped.slice(Math.max(0, m.index - 16), m.index);
137
+ if (PR_CONTEXT_RE.test(before)) continue;
138
+ const n = Number(m[1]);
139
+ counts.set(n, (counts.get(n) ?? 0) + 1);
140
+ lastIndex.set(n, baseIndex + i);
141
+ }
142
+ }
143
+ ```
144
+
145
+ - [ ] **Step 4: Run test file, verify green**
146
+
147
+ Run: `node --test test/code-refs-extract.test.mjs`
148
+ Expected: all pass (new + existing).
149
+
150
+ - [ ] **Step 5: Full suite (A5)**
151
+
152
+ Run: `npm test`
153
+ Expected: all green.
154
+
155
+ ### Task 3: README providers.json example (G1/A6)
156
+
157
+ **Files:**
158
+ - Modify: `README.md` (Evidence and Code References section)
159
+ - Modify: `test/code-refs-extract.test.mjs` (one JSON-validity test)
160
+
161
+ - [ ] **Step 6: Append after the section's first paragraph**
162
+
163
+ ```markdown
164
+ For internal code platforms (non-github/gitlab hosts) or custom CLIs, add a
165
+ per-store `providers.json` so claim/action rules exist and sessions stop
166
+ depending on the low-confidence mention fallback. Example — an internal CLI
167
+ (`acli`) with claim-strength issue rules:
168
+
169
+ ```json
170
+ {
171
+ "providers": [
172
+ {
173
+ "name": "acode",
174
+ "hosts": ["acode.internal.example.com"],
175
+ "rules": [
176
+ { "pattern": "acli\\s+issue\\s+update\\s+#?(\\d+)(?=[\\s\\S]*--assignee)", "kind": "issue", "strength": "claim" },
177
+ { "pattern": "acli\\s+issue\\s+(?:note|comment|close)\\s+#?(\\d+)", "kind": "issue", "strength": "action" },
178
+ { "pattern": "acli\\s+issue\\s+(?:show|view)\\s+#?(\\d+)", "kind": "issue", "strength": "view" }
179
+ ]
180
+ }
181
+ ]
182
+ }
183
+ ```
184
+
185
+ `hosts` matches the repo's remote host; rules follow the same
186
+ `pattern`/`kind`/`strength` shape as the built-in `gh`/`glab` tables
187
+ (`strength`: `claim` > `action` > `view`), and `#N` in the matched text
188
+ captures the number. Validation errors from a broken file are surfaced in the
189
+ diagnostics panel and never break extraction — the file is simply ignored.
190
+ ```
191
+
192
+ - [ ] **Step 7: JSON-validity test**
193
+
194
+ ```js
195
+ test("README providers.json example parses and validates (issue #61)", async () => {
196
+ const readme = readFileSync(new URL("../README.md", import.meta.url), "utf8");
197
+ const fence = readme.match(/```json\n(\{[\s\S]*?providers[\s\S]*?\})\n```/);
198
+ assert.ok(fence, "README contains a json providers example");
199
+ const parsed = JSON.parse(fence[1]);
200
+ const { provider, errors } = validateProvider(parsed.providers[0]);
201
+ assert.deepEqual(errors, []);
202
+ assert.equal(provider.name, "acode");
203
+ assert.equal(provider.rules.length, 3);
204
+ });
205
+ ```
206
+
207
+ Add `validateProvider` to the import from `../src/core/code-refs.mjs` and
208
+ `readFileSync` to the node:fs import if not already present.
209
+
210
+ - [ ] **Step 8: Full suite + scope (A5/A7)**
211
+
212
+ Run: `npm test` and `git diff main -- src/ README.md test/ --stat`
213
+ Expected: all green; scope = code-refs.mjs, README.md, code-refs-extract.test.mjs.
214
+
215
+ - [ ] **Step 9: Commit**
216
+
217
+ ```bash
218
+ git add src/core/code-refs.mjs README.md test/code-refs-extract.test.mjs
219
+ git commit -m "fix: guard mention fallback against placeholders, code spans, and pr-context (issue #61)"
220
+ ```
221
+
222
+ ## Self-Review
223
+
224
+ - Spec coverage: A1/A2/A3 (Task 1 red + Task 2), A4 (composite test), A5 (Step 5), A6 (Task 3), A7 (Step 8). Covered.
225
+ - Placeholder scan: full code shown.
226
+ - Type consistency: regexes module-level constants; matchAll index usage valid post-strip (indices refer to `stripped`, used only for the window check).
@@ -0,0 +1,87 @@
1
+ # Issue #63 Flaky Manual-Completion Assert — Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Stop the runner.integration "manual completion" test from racing `markCompleted` against the status.json→state.json persist window (test-only fix).
6
+
7
+ **Architecture:** Replace the one-shot `assert.deepEqual(createService({root}).markCompleted("view_1"), {ok:true})` with a `waitFor` poll of `markCompleted` until `{ok:true}`, then assert the exact shape. Durability asserts after runner exit stay unchanged.
8
+
9
+ **Tech Stack:** Node built-in test runner (`node --test`), existing `waitFor(fn, timeoutMs = 15000, intervalMs = 50)` helper in the test file.
10
+
11
+ ## Global Constraints
12
+
13
+ - Test-only: modify `test/runner.integration.test.mjs` and nothing else (spec A4).
14
+ - No production code change under `runner/` or `src/` (spec Non-goals).
15
+ - Test command: `npm test` (= `node --test test/*.test.mjs`).
16
+ - Commit style: conventional commits.
17
+
18
+ **Execution mode:** inline (executing-plans) — single task, 3-line diff; the heavy part is the 10-run verification loop which stays in the controlling session.
19
+
20
+ ---
21
+
22
+ ### Task 1: Poll markCompleted to success in runner.integration test
23
+
24
+ **Files:**
25
+ - Modify: `test/runner.integration.test.mjs` (~L357-364, test "runner does not clobber a manual completion made during post-exit model passes")
26
+
27
+ **Interfaces:**
28
+ - Consumes: `waitFor(fn, timeoutMs = 15000, intervalMs = 50)` (defined at top of same file), `createService({ root }).markCompleted(viewId)` → `{ ok: boolean, error?: string }`.
29
+ - Produces: nothing consumed by other tasks.
30
+
31
+ - [ ] **Step 1: Replace the one-shot assertion with a poll**
32
+
33
+ Replace:
34
+
35
+ ```js
36
+ // User marks the row done manually while the model passes are in flight.
37
+ const { createService } = await import("../src/runtime/service.mjs");
38
+ assert.deepEqual(createService({ root }).markCompleted("view_1"), { ok: true });
39
+ ```
40
+
41
+ with:
42
+
43
+ ```js
44
+ // User marks the row done manually while the model passes are in flight.
45
+ // persist() writes status.json before state.json, so markCompleted can
46
+ // transiently reject ("active run") after endedAt is already visible.
47
+ // Poll it to success, then assert the exact shape; the race window is
48
+ // millisecond-scale and 15s default timeout is ample.
49
+ const { createService } = await import("../src/runtime/service.mjs");
50
+ const svc = createService({ root });
51
+ let manual = null;
52
+ await waitFor(() => {
53
+ manual = svc.markCompleted("view_1");
54
+ return manual.ok ? manual : null;
55
+ });
56
+ assert.deepEqual(manual, { ok: true });
57
+ ```
58
+
59
+ Note: if `waitFor` times out it returns `null` and leaves `manual` falsy-or-`{ok:false}`; the trailing `assert.deepEqual` then fails — timeout is still a test failure, not a silent pass.
60
+
61
+ - [ ] **Step 2: Run the single test file, expect pass**
62
+
63
+ Run: `npm test` (full suite; there is no per-file filter in `node --test` glob mode — use `node --test test/runner.integration.test.mjs` for the quick check)
64
+ Expected: PASS, no other test affected.
65
+
66
+ - [ ] **Step 3: Verify A3 — 10 consecutive full `npm test` runs**
67
+
68
+ Run: `for i in $(seq 1 10); do npm test > /tmp/issue63-run-$i.log 2>&1 || { echo "RUN $i FAILED"; break; }; echo "run $i ok"; done`
69
+ Expected: `run 1 ok` … `run 10 ok`, zero failures (spec A3).
70
+
71
+ - [ ] **Step 4: Verify A4 — only the test file changed**
72
+
73
+ Run: `git diff main --stat`
74
+ Expected: exactly one file: `test/runner.integration.test.mjs`.
75
+
76
+ - [ ] **Step 5: Commit**
77
+
78
+ ```bash
79
+ git add test/runner.integration.test.mjs
80
+ git commit -m "test: poll markCompleted to success instead of racing the persist window (issue #63)"
81
+ ```
82
+
83
+ ## Self-Review
84
+
85
+ - Spec coverage: A1 (Step 1 diff shape), A2 (Step 2 + unchanged durability asserts), A3 (Step 3), A4 (Step 4). Covered.
86
+ - Placeholder scan: no TBD/TODO; exact code shown.
87
+ - Type consistency: `manual` is `{ok:boolean,error?:string}`; `waitFor` returns truthy-on-success, `null` on timeout. Consistent.