@zhuxixi/pi-agent-board 0.5.2 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/CHANGELOG.md +32 -0
  2. package/README.md +41 -3
  3. package/docs/superpowers/plans/2026-09-03-code-refs-pr-backlink-narrow.md +551 -0
  4. package/docs/superpowers/plans/2026-09-04-evidence-outputpreview.md +209 -0
  5. package/docs/superpowers/plans/2026-09-04-warm-host-reclaim.md +796 -0
  6. package/docs/superpowers/plans/2026-09-05-issue-13-drainnextfollowup-pty-probe.md +114 -0
  7. package/docs/superpowers/plans/2026-09-05-issue-38-windows-wezterm-ime-cursor.md +73 -0
  8. package/docs/superpowers/plans/2026-09-05-issue-39-truncate-codepoint-boundary.md +143 -0
  9. package/docs/superpowers/plans/2026-09-05-issue-61-mention-fallback-guards.md +226 -0
  10. package/docs/superpowers/plans/2026-09-05-issue-63-flaky-manual-completion.md +87 -0
  11. package/docs/superpowers/plans/2026-09-05-issue-64-changelog-release-helper.md +53 -0
  12. package/docs/superpowers/plans/2026-09-05-pty-host-stacking-sock-race.md +731 -0
  13. package/docs/superpowers/specs/2026-08-29-code-refs-badges-design.md +1 -1
  14. package/docs/superpowers/specs/2026-09-03-code-refs-pr-backlink-narrow-design.md +92 -0
  15. package/docs/superpowers/specs/2026-09-04-evidence-outputpreview-design.md +50 -0
  16. package/docs/superpowers/specs/2026-09-04-warm-host-reclaim-design.md +106 -0
  17. package/docs/superpowers/specs/2026-09-05-issue-13-drainnextfollowup-pty-probe-design.md +64 -0
  18. package/docs/superpowers/specs/2026-09-05-issue-38-windows-wezterm-ime-design.md +48 -0
  19. package/docs/superpowers/specs/2026-09-05-issue-39-truncate-codepoint-boundary-design.md +64 -0
  20. package/docs/superpowers/specs/2026-09-05-issue-61-mention-fallback-design.md +71 -0
  21. package/docs/superpowers/specs/2026-09-05-issue-63-flaky-manual-completion-design.md +49 -0
  22. package/docs/superpowers/specs/2026-09-05-issue-64-changelog-helper-design.md +76 -0
  23. package/docs/superpowers/specs/2026-09-05-pty-host-stacking-sock-race-design.md +510 -0
  24. package/package.json +83 -81
  25. package/runner/job-runner.mjs +2 -2
  26. package/runner/pty-runner.mjs +573 -2
  27. package/runner/state-runner.mjs +3 -0
  28. package/runner/title-runner.mjs +1 -1
  29. package/scripts/release_helper.mjs +277 -0
  30. package/src/commands/agent-board.ts +38 -35
  31. package/src/commands/attach-decision.mjs +66 -0
  32. package/src/commands/attach-flow.ts +45 -39
  33. package/src/core/code-refs.mjs +85 -33
  34. package/src/core/evidence.mjs +2 -2
  35. package/src/core/heuristics.mjs +40 -2
  36. package/src/core/host-coordination.mjs +159 -0
  37. package/src/core/host-crash.mjs +43 -3
  38. package/src/core/host-probe.mjs +196 -0
  39. package/src/core/launch.mjs +3 -1
  40. package/src/core/locks.mjs +196 -1
  41. package/src/core/paths.mjs +24 -0
  42. package/src/core/store.mjs +164 -5
  43. package/src/core/types.mjs +17 -1
  44. package/src/core/warm-host-sweeper.mjs +150 -0
  45. package/src/index.ts +29 -1
  46. package/src/runtime/service.mjs +967 -109
  47. package/src/ui/dashboard-decisions.mjs +55 -0
  48. package/src/ui/dashboard.ts +26 -12
@@ -63,7 +63,7 @@ repo.mjs remoteHost()(带缓存)──────────────
63
63
  |---|---|---|
64
64
  | claim(最强) | 认领命令:provider 规则里 `strength: "claim"` 的命令模式(GitHub 内置:`gh issue edit N --add-assignee`) | commands 正则 |
65
65
  | claim | worktree 命名:`issue-<N>-<slug>`(issue-driven 工作流强制规范) | worktreePath / branch 结构化解析(非正则配置,引擎内置) |
66
- | claim | PR 回链:`gh pr create` 的 body/后续文本中的 `issue #N` / `Closes #N`(同时定 issue + PR 两个值) | commands + assistantTexts |
66
+ | claim | PR 回链(按证据上下文拆分,issue #65):`gh pr create` body 中的 `Closes #N` / `fixes issue #N` / 兼容裸 `issue #N`;后续 assistant 文本仅认 canonical `Closes/Fixes/Resolves #N`(带单词边界与 7 位编号边界);仅在恰好一个 PR create 命令时扫描 assistant,后续 command 不参与回链 | create 命令自身 + assistantTexts |
67
67
  | action(强) | `issue comment/edit/close N`、`pr checkout/view/merge N`、`pr create`(编号从 outputUrl 或后续 URL 反查,见 D4 限制) | commands 正则 |
68
68
  | view(中) | `issue view N` / `pr view N` | commands 正则,要求频次 ≥2,取最近一次 |
69
69
  | mention(弱,兜底) | 裸 `#N` | assistantTexts,要求频次显著最高(≥3 且 ≥ 第二名的 2 倍) |
@@ -0,0 +1,92 @@
1
+ # Spec: 收窄 code-refs pr-backlink 提取(issue #65)
2
+
3
+ - Issue: https://github.com/zhuxixi/pi-agent-board/issues/65
4
+ - 日期:2026-09-02
5
+ - 状态:已确认(2026-09-03,含 review 修订)
6
+
7
+ ## 1. 背景与根因(已确认)
8
+
9
+ board 行在「PR 已创建、issue 未关闭」窗口期内把 issue 徽标误显示为 `#1`(正确应为 `#19`)。根因分为直接触发和放大因素:
10
+
11
+ 1. **直接触发是后续证据复用了过宽的正则**:`src/core/code-refs.mjs:510` 的 `PR_BACKLINK_RE` 含 `issue\s+#\d+` 分支,无 closing keyword 锚定。CR 报告模板文本「本轮仅验证上轮 issue #1(no-pushback)」中的 `#1` 是 review finding 编号,却被采为 claim 级 pr-backlink 候选。
12
+ 2. **放大因素是 4b 的证据归属过宽**:`resolveBacklinkAfter`(:651)会把 `pr create` 后的后续 command 与 assistant 文本都视为同一 PR 的回链候选;误匹配的 lastIndex≈158 比真实 #19 信号(PR 正文回链≈66、worktree 命名=-1)更靠后,同强度按 lastIndex 决胜 → `#1` 胜出。
13
+
14
+ `buildEngineInput` 将最近 200 条命令与最近 20 条 assistant 文本分别截取,再把两组数组拼成一条“命令在前、assistant 在后”的伪序列;它没有保留两类证据之间的真实 `at` 时序。因此“最后 3 条命令 / 前 3 条 assistant”不能可靠表示“紧随 PR create 的消息”,本 spec 不采用这个索引启发式。
15
+
16
+ 窗口滑动后 winner 回到 #19 的现象,以及旧 #1 通过 `mergeWithExisting` 留在 `allRefs` 的现象均已确认;后者属于历史 artifact 迁移边界,单独列入非目标。
17
+
18
+ ## 2. 修复设计
19
+
20
+ ### F1:按证据上下文拆分 PR 回链正则(主修复)
21
+
22
+ 不再让一个宽正则同时处理 PR create 命令和后续 assistant 文本,改为两个纯数据规则:
23
+
24
+ ```js
25
+ // Explicit PR-create command: preserve the legacy `issue #N` body form.
26
+ const PR_CREATE_BACKLINK_RE =
27
+ /\b(?:(?:close[sd]?|fix(?:e[sd])?|resolve[sd]?)\b\s+(?:issue\s+)?|issue\s+)#(\d{1,7})(?!\w)/i;
28
+
29
+ // Later assistant evidence: accept only canonical closing-keyword syntax.
30
+ const PR_FOLLOWUP_BACKLINK_RE =
31
+ /\b(?:close[sd]?|fix(?:e[sd])?|resolve[sd]?)\b\s+#(\d{1,7})(?!\w)/i;
32
+ ```
33
+
34
+ - 两个规则都增加单词边界,避免 `prefix #1`、`disclose #2`、`unresolved #3` 这类长单词子串误命中;编号后的 `(?!\w)` 防止把超过 7 位的数字或紧随字母/下划线的 token 截断成前 7 位。
35
+ - `applyPrBacklink` 只在 provider 命中的 PR create 命令上使用 `PR_CREATE_BACKLINK_RE`。PR create 命令中的 `issue #N` 是用户显式提供的 body 内容,保留原 PR #43 的兼容契约。
36
+ - `resolveBacklinkAfter` 只对 assistant 文本使用 `PR_FOLLOWUP_BACKLINK_RE`。后续 assistant 仅接受 `close/closed/closes #N`、`fix/fixed/fixes #N`、`resolve/resolved/resolves #N`,不接受裸 `issue #N`,也不接受 `fixes issue #N`。
37
+ - 两类规则都保留 `claim` 强度;不通过降级强度掩盖误报。
38
+
39
+ ### F2:去掉后续 command 的通用回链扫描,不实现伪时序窗口
40
+
41
+ 当前输入模型无法安全实现“PR create 后 N 条证据”这样的时间窗口,因此本 issue 采用可证明的来源边界,而不是增加固定数量常量:
42
+
43
+ - PR create 命令自身仍由 `PR_CREATE_BACKLINK_RE` 处理。
44
+ - `resolveBacklinkAfter` **只遍历 assistantTexts**,不再遍历 marker 后的任意 command。`gh issue close #21`、`gh pr comment 20 --body "fixes #21"`、`echo "closes #40"` 都不能被升级成 `pr-backlink` claim;它们若命中 provider 自己的 command 规则,仍保留其原本的 issue/pr action 语义,但 source 不得是 `pr-backlink`。
45
+ - **只有恰好一个不同 command index 的 PR create marker 时**,Rule 4b 才调用 assistant backlink resolver;同一条 create 命令若因用户规则与内置规则同时命中,仍只算一个 marker。没有 marker 或存在多个不同 PR create 命令时不产生 assistant `pr-backlink` 候选。这样在当前缺少真实时序的输入模型下,宁可漏掉多 PR 场景的 assistant 回链,也不把一条文本错误归属给某个 PR。
46
+ - resolver 接收 `assistantTexts` 与 `commands.length` 这个 `baseIndex`,按 assistantTexts 保留顺序返回第一个 canonical 命中,并以 `baseIndex + assistantIndex` 记录 `lastIndex`。这里的 `baseIndex` 仅用于保持现有排序契约,不表示真实时间。
47
+ - 不按当前扁平索引增加固定 assistant 数量或固定时间窗口。即使只有一个 PR marker,assistant 的真实 `at` 时序目前仍未进入 `buildEngineInput`,晚到的 canonical 句式仍可能被误归属;这是明确记录的残余风险。若要做到“紧邻 assistant 总结”的严格归属,必须先引入保留 `kind/text/at/sequence` 的 timestamped ordered-evidence 输入,另开 issue 设计。
48
+ - 后续 `gh pr edit --body/--body-file` 的回链归属不在本 issue 承诺范围内;只有被纳入 assistant evidence 且符合 canonical closing 语法的文本才可能被识别。
49
+
50
+ ### F3:同步文档、测试和历史边界
51
+
52
+ - 更新 `src/core/code-refs.mjs` 注释,以及 `docs/superpowers/specs/2026-08-29-code-refs-badges-design.md` 中关于 PR 回链语法和 4b 来源的描述:create body 可兼容 `issue #N`;后续 assistant 只认 canonical closing 语法;任意后续 command 不属于 PR 回链。原始 design doc 是已提交历史文档,本次同步更新必须在 worktree 中与代码、测试同一提交完成。
53
+ - 通过 extractor 测试和 store 组合测试验证新 extraction 不会产生 issue #1;不改变 store/渲染运行逻辑。
54
+ - 本 issue 只保证**新一轮 extraction**不再从 CR 文本产生 `#1`;已有 `github.json` 中的旧 `pr-backlink` ref 不回溯清理,`mergeWithExisting` 的 carry-forward 语义不变。若要求升级后立即清除历史误 ref,需要另一个 artifact migration 设计。
55
+
56
+ ### 不采纳的方向
57
+
58
+ - **按扁平索引加“最后 3 条”窗口**:时序信息不存在,边界不可证明,可能同时漏掉真实 assistant 总结并误收更晚文本。不采纳。
59
+ - **把 4b 降级为 action**:改变 spec D1 设计语义(PR 回链 = claim),且误 ref 仍会残留在 `allRefs`/peek。不采纳。
60
+ - **本 issue 内引入 timestamped ordered-evidence**:这是解决长期归属准确性的正确方向,但会扩大 evidence 输入契约和迁移面;作为后续独立设计,不与本次精确误报修复捆绑。
61
+
62
+ ## 3. 验收矩阵
63
+
64
+ 以下验收同时覆盖“候选没有产生”和“用户可见 winner 没被错误候选抢走”两层;表中命令均可在本地纯数据 fixture 中执行。
65
+
66
+ | ID | 功能点 | 验收方式 | 具体验证 | 通过标准 |
67
+ |----|--------|----------|----------|----------|
68
+ | A1 | issue #65 的真实误报路径不再产生 #1 | 自动化验证(unit) | `node --test test/code-refs-extract.test.mjs` | fixture 含 PR body `Closes #19`、worktree `issue-19-*`、assistant 文本「上轮 issue #1」「pushback verdict for issue #1」;`result.issue.number === 19`,且 `result.allRefs` 完全不含 `kind=issue, number=1`(无论 source) |
69
+ | A2 | create body 兼容性与严格 follow-up 语法 | 自动化验证(unit) | 同上 | create 命令中的裸 `issue #40` 仍得到 `source=pr-body`;create body 与 assistant 中的 `Closes/Fixes/Resolves #40` 按各自上下文正确命中;assistant 中裸 `issue #1`、`fixes issue #1`、`prefix #1`、`disclose #2`、`unresolved #3` 及超过 7 位编号均不作为 follow-up backlink,且上述负例不进入 `allRefs` |
70
+ | A3 | 后续 command 不被冒充为 PR 回链;多 PR assistant 不产生歧义归属 | 自动化验证(unit) | 同上 | create `#19` 后的 `gh issue close #21`、`gh pr comment ... fixes #21`、`echo "closes #21"` 不产生 `source=pr-backlink, number=21`;存在两个 PR create marker 时,后续 assistant 的 canonical 回链不产生任何 `pr-backlink` 候选(不猜测归属) |
71
+ | A4 | 持久化组合路径不产生新的误 ref | 自动化验证(integration) | `node --test test/code-refs-store.test.mjs` | 通过 `updateCodeRefsFromEvidence` 写入新 `github.json` 后,winner 为 #19,`allRefs` 不含新产生的 `pr-backlink #1`;不要求清理预先存在的历史 artifact |
72
+ | A5 | 全量回归与发布包完整性 | 自动化验证(unit + static + build) | `npm run verify` | typecheck、全测试、c8 coverage(lines ≥85%、functions ≥80%、branches ≥70%)及 `npm pack --dry-run` 全部通过 |
73
+
74
+ 无用户实测项:本次行为改动限定在 `code-refs.mjs` 纯函数提取器、对应测试和文档,零网络、零新的持久化协议;历史 artifact 清理明确不属于本次验收。
75
+
76
+ ## 4. 可测性拆分设计
77
+
78
+ 改动主体落在 `src/core/code-refs.mjs`(既有纯函数模块,零 I/O、无副作用),并更新 extractor 测试、store 组合测试和设计文档;不修改 store/渲染运行逻辑:
79
+
80
+ - **`PR_CREATE_BACKLINK_RE` + `matchPrCreateBacklink(text)`**:只负责 PR create 命令自身的回链语法(包括兼容的 legacy `issue #N`)。输入/输出为字符串与正整数或 `null`;通过 `extractCodeRefs` 测 body 正例、closing keyword 单词边界和 7 位编号边界。
81
+ - **`PR_FOLLOWUP_BACKLINK_RE` + `matchPrFollowupBacklink(text)`**:只负责后续 assistant 的 canonical closing 语法。测试 `Closes/Fixes/Resolves #N` 正例,以及裸 `issue #N`、`fixes issue #N`、嵌入长单词、超过 7 位编号等负例;同时验证命中后只产生 `source=pr-backlink`,不会把同一文本中的其他裸 `#N` 误升级。
82
+ - **`resolveBacklinkAfter(assistantTexts, baseIndex)`**:保持纯函数,只遍历 assistant 文本并使用 follow-up matcher;不再接收或扫描 command,也不再需要 marker/`stopBefore`。调用方仅在不同 PR create command index 恰好为 1 时调用一次;多 PR 场景直接跳过 resolver,通过 `extractCodeRefs` 黑盒测试这一保守边界,不为内部细节新增公共导出。
83
+ - **store 组合边界**:使用临时 root 和脱敏 evidence 调用 `updateCodeRefsFromEvidence`,确认 extractor 结果经过 artifact 合并后仍不新增 #1;单独断言“历史 artifact 不自动清理”,避免把非目标误写成已修复。
84
+ - **副作用隔离**:不读网络、不改变 `github.json` schema、不改 `mergeWithExisting`;timestamped ordered-evidence 归属模型和 artifact migration 另行设计。
85
+
86
+ ## 5. 非目标
87
+
88
+ - 不按当前扁平 `commands + assistantTexts` 索引实现固定数量/时间窗口;不在本 issue 改造 timestamped ordered-evidence 输入。
89
+ - 不扫描任意后续 command 作为 `pr-backlink`;不承诺多 PR 场景的 assistant 回链归属;不承诺后续 `gh pr edit --body/--body-file` 的独立归属。
90
+ - 不清理已经写入 `github.json` 的旧误 ref;carry-forward 机制本身不变,升级后的历史清理另开 artifact migration issue。
91
+ - 不动 mention 兜底(issue #61 是独立问题)。
92
+ - 不改变 providers.json schema;两个 backlink 正则是引擎内置的上下文规则,不能由 provider 规则配置覆盖。
@@ -0,0 +1,50 @@
1
+ # Design: evidence outputPreview 从 AgentToolResult 正确提取文本(issue #41)
2
+
3
+ 状态:approved(用户确认于 2026-09-04)
4
+ 仓库:zhuxixi/pi-agent-board · 调研留档:`~/.claude/github-issue-driven/zhuxixi/pi-agent-board/issue-41/research/pi-tool-result-structure.md`
5
+
6
+ ## 1. 问题与根因(调研已闭环)
7
+
8
+ - 现象:`evidence.json` 的 `commands[].outputPreview` 对 bash 命令一律为 `"[object Object]"`,真实输出丢失(`~/.pi/agent/agent-board/views/` 抽样 100% 复现)。
9
+ - 根因:`src/core/evidence.mjs` `reduceEvidence` 的 `tool_execution_end` → bash 分支用 `String(event.result ?? "")`;pi 的 `result` 是 `AgentToolResult` 对象(`{ content: (TextContent|ImageContent)[], details, ... }`),文本在 `content[]` 的 `type==="text"` block 里。
10
+ - 漏网原因:现有单测 fixture 把 `result` 写成字符串 `"ok"`,与真实事件结构不符。
11
+ - pi 官方提取模式(`convertToolResultOutput`):`content.filter(c => c.type === "text").map(c => c.text).join("\n")`。
12
+
13
+ ## 2. 设计决策表
14
+
15
+ | ID | 决策 | 理由 |
16
+ |----|------|------|
17
+ | D1 | 新增纯函数 `toolResultText(result)`,放 `src/core/heuristics.mjs`,与既有 `assistantText`(message.content text-block 提取)同文件、同防御模式 | 模式对称,仓库先例;不新建文件 |
18
+ | D2 | 行为:`null/undefined → ""`;`string → 原样`(兼容旧 fixture/历史回放);`对象 + Array.isArray(content) → filter(type==="text" && typeof text==="string") → map → join("\n") → trim`(纯 image → "");**对象无 content 数组 / 无 text block → `""`(宁缺毋滥)**;标量(number/bool 等)→ `String()` | 照抄 `assistantText` 防御式 + pi 官方 join("\n") 语义;兜底遵循 #40「宁可不显示也不错显示」原则——未知形状对象不再产生 `[object Object]`(即本 bug 的兜底复现路径,见 Review F1) |
19
+ | D3 | `reduceEvidence` 仅改一行:`outputPreview: truncate(toolResultText(event.result), 500)` | 单点修复,先提全文再截 500(与现状顺序一致) |
20
+ | D4 | 不动 `EvidenceCommand` 数据结构、不动其他工具分支、不迁移历史 evidence.json | 历史输出物理丢失无法恢复;其他工具本就无 outputPreview 提取 |
21
+ | D5 | 向后兼容:旧字符串 `result` 的既有 fixture `"ok"` 必须继续通过 | 防止修复破坏旧事件回放语义 |
22
+
23
+ ## 3. 可测性拆分设计(硬约束)
24
+
25
+ - `toolResultText(result)`:**纯函数**,零副作用、零依赖 pi 运行时,输入任意 → 输出 string。测试边界:unit 直接构造各形状输入断言输出(单 text block / 多 text block join "\n" / 纯 image → "" / **对象无 content 字段 → ""** / string 原样 / null/undefined → "" / number/bool 标量 → String())。
26
+ - `reduceEvidence` bash 分支:保持现有快照式测试模式(构造 event → 断言 snapshot),仅补真实 `AgentToolResult` 形状 fixture;**不得**把提取逻辑内联进 reduceEvidence(保持函数已拆分,实现阶段不得耦合回去)。
27
+ - 测试层级选择:全部行为 unit 层可证(纯函数 + 快照),无需 integration/E2E mock 整个 pi runtime(成本高于收益)。真实 pi 进程链路(extension 加载 → bash 工具执行 → evidence 落盘)无法在单测内稳定脚本化,划给 U1 用户实测。
28
+
29
+ ## 4. 验收矩阵
30
+
31
+ | ID | 功能点 | 验收方式 | 具体验证 | 通过标准 |
32
+ |----|--------|----------|----------|----------|
33
+ | A1 | `toolResultText` 纯函数各形状行为 | 自动化验证(unit) | `node --test test/heuristics.test.mjs` | 新增用例全过:AgentToolResult 形状提取、多 block join("\n")、纯 image → ""、对象无 content → ""(F1 兜底)、string 原样、null → ""、标量 String() 兜底 |
34
+ | A2 | `reduceEvidence` bash 分支用真实结构 | 自动化验证(unit) | `node --test test/evidence.test.mjs` | 真实 AgentToolResult fixture → outputPreview 为提取文本且 ≠ "[object Object]";既有字符串 fixture `"ok"` 断言不回归(D5) |
35
+ | A3 | 全量回归 | 自动化验证(unit+static) | `npm run verify` | 全部通过(注意勿用 `npm test -- <file>`,glob 会展开全量) |
36
+ | U1 | 真实 board 运行链路 | 用户实测 | 真实 pi session 里跑若干 bash 命令 → 查 `~/.pi/agent/agent-board/views/<view>/evidence.json`(**可执行时机:merge 发布、board 随日常 pi session 重启加载新版 extension 后**) | 新记录的 bash 命令 outputPreview 含真实输出文本(如 `gh pr create` 的 URL),不再出现 `[object Object]` |
37
+
38
+ ## 5. 非目标
39
+
40
+ - 不修历史 evidence.json(数据已丢);不清理历史 `[object Object]` 记录。
41
+ - 不动其他工具的 result 处理与 code-refs 引擎(#41 修复后解锁 #40 的 outputUrl 路径,属后续工作)。
42
+ - 不改 500 字符截断长度与 `upsertEvidenceCommand` 覆盖语义。
43
+ - **不处理 powershell 工具**(Windows 下 pi 的 shell 工具是 powershell,其命令本就不进 commands 列表——既有行为,与本 bug 无关)。
44
+
45
+ ## 6. 改动文件清单
46
+
47
+ - `src/core/heuristics.mjs`:+`toolResultText`(~12 行)
48
+ - `src/core/evidence.mjs`:import + 改 1 行
49
+ - `test/heuristics.test.mjs`:+纯函数用例
50
+ - `test/evidence.test.mjs`:+真实结构 fixture 用例
@@ -0,0 +1,106 @@
1
+ # Spec: warm host 回收 —— 周期 sweep + 退出清理(issue #75)
2
+
3
+ 日期:2026-09-04
4
+ 状态:draft,待用户确认
5
+ 仓库:zhuxixi/pi-agent-board
6
+
7
+ ## 1. 背景与根因
8
+
9
+ detach / 关闭宿主后,空闲 PTY host(runner + child pi)无限期空转(实测 8.5h)。根因:
10
+
11
+ 1. `pruneWarmHosts()` 仅惰性触发(attach/prewarm/dispatch 时),用户 detach 后不再操作 board 则永不触发 → `AGENT_BOARD_WARM_HOST_TTL_MS`(默认 10min)与 `AGENT_BOARD_MAX_WARM_HOSTS`(默认 4)形同虚设
12
+ 2. 宿主 pi 退出/会话切换时无清理钩子(扩展未注册 `session_shutdown`)
13
+ 3. 无周期扫描兜底;Windows 下 runner detached 孤儿,宿主被 kill 后无任何回收
14
+
15
+ ## 2. 设计决策
16
+
17
+ ### M1:提取纯判定函数 `selectIdleHostsToEvict`(可测性核心)
18
+
19
+ 把淘汰判定从 `pruneWarmHosts` 闭包中提取为**纯函数**(无 IO、无 env):
20
+
21
+ ```
22
+ selectIdleHostsToEvict(rows, opts) -> { ttlEvicted: Row[], excessEvicted: Row[] }
23
+
24
+ opts = { now, maxWarm, ttlMs, graceMs, keepViewId, isBusy? }
25
+ ```
26
+
27
+ - idle 定义:`hostAlive && !isBusy(row) && (host.attachedClients ?? 0) === 0 && (row.meta.id !== keepViewId)`
28
+ - **新增 graceMs 豁免**:`host.startedAt` 距今 < graceMs 的 host 不参与淘汰(防 ensureHost→attach 竞态;attach 重连兜底为第二保险)
29
+ - 排序:survivors 按 `lastActivityAt ?? startedAt` 升序,超出 maxWarm 的从最旧淘汰
30
+ - `isBusy` 默认 = 现 `isAgentBusy` 语义(queued/working/hasPendingQuestions),作为参数注入以保持纯性
31
+
32
+ ### M2:新增调度器 `createWarmHostSweeper`(新模块 `src/core/warm-host-sweeper.mjs`)
33
+
34
+ ```
35
+ createWarmHostSweeper({ sweep, intervalMs }) -> { start, stop, sweepNow }
36
+ ```
37
+
38
+ - `setInterval(sweep, intervalMs)`,`unref()`(不阻塞宿主退出);`sweepNow()` 立即执行一次(错误吞掉,best-effort)
39
+ - `stop()` 清定时器;`start()` 幂等
40
+ - 新环境变量 `AGENT_BOARD_SWEEP_INTERVAL_MS`(默认 60_000;0 = 禁用周期 sweep,仅保留启动/session_shutdown 时 sweepNow)
41
+
42
+ ### M3:接线(src/index.ts,扩展模块级)
43
+
44
+ - 扩展加载时:**非 isHostedChild** 才创建 sweeper(`sweep = () => pruneWarmHosts({})`),`start()` + 立即 `sweepNow()`(回收上一个宿主遗留的泄漏 host)
45
+ - 注册 `session_shutdown`:`sweepNow()`(尽力清理)+ `stop()`
46
+ - **isHostedChild 必须跳过**:child pi 内加载扩展时若启动 sweeper,会向自己的 runner 发 terminate(自杀链),绝对禁止
47
+ - 新增环境变量 `AGENT_BOARD_WARM_HOST_GRACE_MS`(默认 30_000;0 = 无豁免,保持旧行为可调)
48
+
49
+ ### M4:dashboard 打开期间顺带 prune(低成本增强)
50
+
51
+ `openDashboard` 的 POLL_MS 轮询回调里加一次 `pruneWarmHosts()`(dashboard 打开 = 用户正用 board,实时回收)。保持既有惰性调用点不变。
52
+
53
+ ### Non-goals(明确不做)
54
+
55
+ - **runner 自退出兜底**:runner 无权威 busy 信号(state.json 是 service 层写的,跨进程读取引入耦合/竞态);误杀风险(PRD stretch "detach while running 转 headless worker" 未实现时,pty child 可能正跑任务)。宿主被杀场景由"下一次任意宿主启动 sweepNow"覆盖,残余缺口(永不再开任何 pi)接受。
56
+ - **detach-while-running 转 headless worker**:PRD stretch,独立议题。
57
+ - 不改变 terminate 消息协议(沿用 `{type:"terminate"}` 既有语义:杀 child,runner 收尾退出)。
58
+
59
+ ## 3. 数据流
60
+
61
+ ```
62
+ 宿主扩展实例(每个 pi session 一个,随 session_shutdown 销毁)
63
+ ├─ module load(非 child): createWarmHostSweeper(...).start() + sweepNow()
64
+ ├─ setInterval(60s, unref): pruneWarmHosts({}) → selectIdleHostsToEvict → sendHostMessage(terminate)
65
+ ├─ dashboard POLL_MS: pruneWarmHosts({})(UI 打开时)
66
+ └─ session_shutdown: sweepNow() + stop()
67
+
68
+ selectIdleHostsToEvict(rows, opts) ← 纯判定(无副作用)
69
+ pruneWarmHosts(opts) ← 薄执行层(listRows + select + sendHostMessage),既有惰性调用点行为等价
70
+ sendHostMessage(row, terminate) ← 既有 fire-and-forget socket 路径(#71 协议)
71
+ ```
72
+
73
+ ## 4. 验收矩阵
74
+
75
+ | ID | 功能点 | 验收方式 | 具体验证 | 通过标准 |
76
+ |----|--------|----------|----------|----------|
77
+ | A1 | selectIdleHostsToEvict:TTL 到期淘汰、未到期保留、keepViewId 豁免 | 自动化(unit) | `node --test test/warm-host-sweeper.test.mjs` | 过期 idle host 进入 ttlEvicted;未过期/keepViewId 不进入 |
78
+ | A2 | busy(queued/working/pendingQuestions)与 attachedClients>0 永不淘汰 | 自动化(unit) | 同上 | 两类 host 均不进入任何淘汰列表 |
79
+ | A3 | graceMs 豁免:startedAt 距今 < grace 的 host 不淘汰 | 自动化(unit) | 同上 | 豁免窗口内不进淘汰列表;窗口外正常淘汰 |
80
+ | A4 | maxWarm 超额时淘汰最旧 survivors(按 lastActivityAt/startedAt) | 自动化(unit) | 同上 | 淘汰对象为排序最旧者,数量 = survivors - maxWarm |
81
+ | A5 | sweeper 调度:start 定时触发、stop 停止、sweepNow 立即执行、unref | 自动化(unit) | 同上(node:test mock timers) | fake timer 前进 intervalMs 后 sweep 被调用;stop 后不再调用;sweepNow 立即调用 |
82
+ | A6 | 接线:非 child 宿主启动即 sweepNow + 周期 sweep;isHostedChild 不接线 | 自动化(integration) | `node --test test/service.test.mjs` 增补:临时 root + node:net 假 server 捕获 terminate | idle host 的 socketPath 收到 terminate;child 模式(flag)下无 sweeper 启动 |
83
+ | A7 | session_shutdown 触发 sweepNow(退出清理) | 自动化(integration) | 同上:调用注册的 shutdown 回调 | terminate 发出;sweeper 已 stop |
84
+ | A8 | 重构无回归:pruneWarmHosts 既有惰性调用点行为等价 | 自动化(全量) | `node --test test/*.test.mjs`(基线 409+ 含既有失败清单) | 与修复前基线一致,无新增失败 |
85
+ | U1 | 真实场景回收 | 用户实测 | 开会话→attach→detach→设 `AGENT_BOARD_WARM_HOST_TTL_MS=10s` + `AGENT_BOARD_SWEEP_INTERVAL_MS=5s`→观察 | TTL 过后 host.json 变 exited,runner/child 进程退出,dashboard 状态正确 |
86
+
87
+ ## 5. 可测性拆分设计(实现硬约束)
88
+
89
+ 1. **`selectIdleHostsToEvict`**:纯函数,输入 rows + opts,输出淘汰分组;不触盘、不连 socket、不读 process.env(所有阈值参数显式传入)。测试直接构造内存 rows。**边界**:判定逻辑全部在此函数内;`pruneWarmHosts` 只做 IO 编排,不得内联任何判定分支。
90
+ 2. **`createWarmHostSweeper`**:调度器封装,`sweep` 回调依赖注入;timer 用全局 setInterval(node:test mock timers 可控)。**边界**:不 import service/store;与业务零耦合。
91
+ 3. **`attachWarmHostSweeperToLifecycle(piLike, deps)`**(index.ts 内接线函数):接收 `pi`(on 事件注册)与 `deps = { createSweeper, isHostedChild, sweep }`,可注入 fake 断言注册行为。**边界**:index.ts 只调这个函数,不内联生命周期逻辑。
92
+ 4. **`pruneWarmHosts`**:重构为薄执行层(listRows → selectIdleHostsToEvict → sendHostMessage),原惰性调用点(dispatch/ensureHost/dashboard)签名不变,行为等价(A8 回归保障)。
93
+
94
+ ## 6. 风险与缓解
95
+
96
+ - 竞态(attach 与 sweep 同时):graceMs + attach 重连兜底(既有 #48 L3)
97
+ - child pi 自杀链:isHostedChild 硬跳过(A6 覆盖)
98
+ - 多宿主并发 sweep 重复 terminate:terminate 幂等(sendHostMessage 失败静默、runner 已死则 socket error 吞掉)
99
+ - sweep 与 archive/stop 竞态:单进程内同 service 数据视图,均 fire-and-forget,最坏重复 terminate(幂等)
100
+
101
+ ## 7. 测试计划摘要
102
+
103
+ - 新增 `test/warm-host-sweeper.test.mjs`:A1-A5(unit,含 mock timers)
104
+ - 增补 `test/service.test.mjs`:A6-A7(integration,net 假 server 捕获 terminate)
105
+ - 全量回归:A8
106
+ - U1 用户实测清单写入 PR 描述与 issue 评论
@@ -0,0 +1,64 @@
1
+ # Spec: issue #13 — drainNextFollowUp forces PTY probe refresh on the 700ms poll path
2
+
3
+ ## Problem
4
+
5
+ `src/runtime/service.mjs` `drainNextFollowUp` (L496) calls
6
+ `ptySupport({ refresh: true })` and is reachable from the 700ms `reconcile()`
7
+ poll (L1007-1011: queued follow-up + `canAutoDrain` → `drainNextFollowUp`).
8
+ When a queued follow-up repeatedly fails to start, every poll cycle forces a
9
+ fresh PTY probe — `ptySpawnSupported` does a **real `pty.spawn`** process.
10
+ This violates the probe-cache discipline (`shouldProbePtySupport`): success
11
+ cached for process lifetime, failures retried on a 2s TTL; only `refresh`
12
+ bypasses it.
13
+
14
+ `dispatch` (L616) and `reply` (L649) also use `refresh: true` but are explicit
15
+ user actions — fine as-is, untouched.
16
+
17
+ ## Decision (issue's fix sketch, verbatim)
18
+
19
+ Replace in `drainNextFollowUp`:
20
+
21
+ ```js
22
+ const pty = ptySupport({ refresh: true });
23
+ ```
24
+
25
+ with default cached-probe semantics:
26
+
27
+ ```js
28
+ const pty = ptySupport();
29
+ ```
30
+
31
+ Plus a red-green test asserting the injected `ptySupport` never receives
32
+ `refresh: true` from the reconcile→drain path (pattern from PR #12's service
33
+ test "ensureHost probes PTY support with TTL cache, not forced refresh",
34
+ test/service.test.mjs:940).
35
+
36
+ ## Behavior contract
37
+
38
+ - Poll-path drains use the cache: after a successful probe, zero further real
39
+ spawns; after a failed probe, at most one real spawn per 2s window.
40
+ - Explicit user actions (`dispatch`, `reply`) keep forced refresh.
41
+ - No other behavior change in drain logic (claim/launch/complete paths
42
+ untouched).
43
+
44
+ ## Non-goals
45
+
46
+ - No change to `dispatch`/`reply` refresh semantics.
47
+ - No change to `ptySpawnSupported`/`shouldProbePtySupport` internals.
48
+
49
+ ## Acceptance matrix
50
+
51
+ | ID | Feature point | Acceptance | Concrete verification | Pass criteria |
52
+ |----|---------------|------------|----------------------|---------------|
53
+ | A1 | reconcile→drain path never forces probe refresh | Automated (unit) | New test in `test/service.test.mjs`: idle row + queued follow-up → `svc.reconcile()` → injected ptySupport spy records opts; assert `probeCalls.length >= 1` and no call has `refresh === true` | New test passes; fails (red) if production line keeps `refresh: true` |
54
+ | A2 | Explicit user actions unaffected | Automated (unit) | Existing service tests (dispatch/reply paths) green in full suite | `npm test` all green |
55
+ | A3 | Full regression | Automated (integration) | Full `npm test` | All pass |
56
+ | A4 | Change scope | Automated (static) | `git diff main` review | Production delta is exactly the one-line probe switch in `src/runtime/service.mjs`; test delta is one new test |
57
+
58
+ ## Testability split design
59
+
60
+ Test-only seam already exists: `createService` opts inject `ptySupport`
61
+ (spy) and `launch` (fake, avoids real spawn). The new test drives the public
62
+ `reconcile()` API so the exact poll path is covered. Row state constructed
63
+ via `readState`/`writeState` (idle + exited) so `canAutoDrain` holds. No new
64
+ production functions.
@@ -0,0 +1,48 @@
1
+ # Spec: issue #38 — document Windows WezTerm IME hardware-cursor requirement
2
+
3
+ ## Problem
4
+
5
+ With #24/#28 shipped (v0.4.3), the IME candidate window is pinned and
6
+ flicker-free on Linux (WezTerm + fcitx5/X11). On **Windows WezTerm (WSL2
7
+ backend)** the candidate window stays stuck at the right edge: per upstream
8
+ earendil-works/pi#5200, Windows WezTerm only updates the IME candidate
9
+ position when the **hardware cursor is visible**, and pi-tui hides it by
10
+ default. The visible block cursor in the editor is a fake (reverse-video
11
+ content), so "I can see the cursor" does not imply the hardware cursor is on.
12
+
13
+ Fix (verified on the real environment, per the issue): `PI_HARDWARE_CURSOR=1`
14
+ env or `"showHardwareCursor": true` in pi's settings.json. Trade-off: the
15
+ real terminal cursor becomes visible in the TUI (cosmetic only; harmless on
16
+ Linux).
17
+
18
+ ## Decision
19
+
20
+ Add one Troubleshooting subsection to README.md:
21
+
22
+ - Title: IME candidate window stuck at the right edge (Windows WezTerm)
23
+ - Content: symptom + platform scope (Windows WezTerm → WSL2, Windows IME/TSF;
24
+ Linux unaffected), root cause in one sentence (hardware cursor hidden by
25
+ default; Windows WezTerm needs it visible to track IME position; the editor
26
+ block cursor is fake content, not the hardware cursor), both fixes with
27
+ sync/scope guidance (machine-local env vs synced settings.json), and the
28
+ cosmetic trade-off.
29
+
30
+ Placement: after "### Attach is slow or keeps reconnecting" (IME is an
31
+ attach-surface concern, keeps related attach topics adjacent).
32
+
33
+ ## Non-goals
34
+
35
+ - No code changes.
36
+ - No upstream comment on earendil-works/pi#5200 (owner's voice, left to the
37
+ repo owner — noted in the issue comment).
38
+
39
+ ## Acceptance matrix
40
+
41
+ | ID | Feature point | Acceptance | Concrete verification | Pass criteria |
42
+ |----|---------------|------------|----------------------|---------------|
43
+ | A1 | Troubleshooting entry exists with both fixes | Automated (static) | grep README.md for the new section + `PI_HARDWARE_CURSOR=1` + `showHardwareCursor` | All three present in the Troubleshooting section |
44
+ | A2 | Scope: docs-only change | Automated (static) | `git diff main --stat` | Only README.md (plus process docs under docs/superpowers/) |
45
+
46
+ ## Testability split design
47
+
48
+ Docs-only: verification is content presence + scope, no behavioral seam.
@@ -0,0 +1,64 @@
1
+ # Spec: issue #39 — truncate() cuts surrogate pairs, producing lone surrogates (U+FFFD) and wrap-misaligned dashboard rows
2
+
3
+ ## Problem
4
+
5
+ `src/core/heuristics.mjs` `truncate(s, n)` uses `str.length` and `str.slice()`
6
+ (UTF-16 code units). When the cut lands inside a surrogate pair (e.g. emoji,
7
+ 2 units), the output keeps a lone high surrogate → renders as U+FFFD →
8
+ pi-tui counts it 1 column while WezTerm renders 2 → line over-wide → terminal
9
+ wrap → TUI diff-render row misalignment → stacked stale frames on the
10
+ dashboard (issue's repro chain).
11
+
12
+ `truncate` feeds `deriveSummary` (max=80), `events.mjs`
13
+ (`latestAssistantPreview`), `evidence.mjs`, `auto-state.mjs` — all display
14
+ text paths.
15
+
16
+ ## Decision (fix direction A, repo-side)
17
+
18
+ Two changes inside `truncate`:
19
+
20
+ 1. **Code-point-safe cut (cut-point back-off, NOT code-point budget).** Keep
21
+ the existing UTF-16 unit budget `n` (all callers pass display budgets; CJK
22
+ chars occupy 2 units — switching to a code-point count would double the
23
+ effective width of CJK summaries from 80 to ~160 columns and *reintroduce*
24
+ over-wide lines). Instead: if the character just before the cut point is a
25
+ high surrogate whose low surrogate is the first dropped unit, back the cut
26
+ off by one unit so the pair stays whole. Output length only shrinks by one
27
+ unit in that rare case; the ellipsis still fits the budget.
28
+
29
+ 2. **Defensive lone-surrogate strip on output.** Input text (from agent
30
+ output stored in status.json) may already contain lone surrogates; strip
31
+ them from the returned string on both the short path and the truncated
32
+ path (`/[\uD800-\uDBFF](?![\uDC00-\uDFFF])/g` and the mirrored
33
+ low-surrogate regex).
34
+
35
+ Out of scope (per goal boundary): pi-tui width-table changes for Ambiguous
36
+ characters (fix direction B) and terminal-side clamping (direction C) —
37
+ upstream/external. Completion here = truncate no longer produces or passes
38
+ through lone surrogates; "visual overlap fully gone" additionally depends on
39
+ those upstream items.
40
+
41
+ ## Non-goals
42
+
43
+ - No `Intl.Segmenter` grapheme segmentation (ZWJ emoji families remain
44
+ multi-codepoint; the defect being fixed is lone surrogates, not grapheme
45
+ splitting).
46
+ - No caller changes; no width-function changes.
47
+
48
+ ## Acceptance matrix
49
+
50
+ | ID | Feature point | Acceptance | Concrete verification | Pass criteria |
51
+ |----|---------------|------------|----------------------|---------------|
52
+ | A1 | Cut never splits a surrogate pair | Automated (unit) | New tests in `test/heuristics.test.mjs`: e.g. `truncate("a👍b", 3) === "a…"` (back-off) and a long-string case whose cut lands on an emoji; assert output matches `/\uD83D$/` never (no trailing lone high surrogate) | Tests pass |
53
+ | A2 | Output is lone-surrogate-free even when input isn't | Automated (unit) | `truncate("ab\ud83d", 10) === "ab"`, `truncate("ab\udc4dzzzzzzzzzz", 5)` has no lone low surrogate | Tests pass |
54
+ | A3 | Existing behavior preserved for ASCII and CJK | Automated (unit) | Existing tests unchanged (`truncate("hello world", 5) === "hell…"`) + new: `truncate("一二三四五", 5) === "一二…"` (2-unit budget semantics kept) | Tests pass |
55
+ | A4 | Full regression | Automated (integration) | Full `npm test` | All green |
56
+ | A5 | Scope | Automated (static) | `git diff main -- src/ test/` | Only `src/core/heuristics.mjs` (truncate) and `test/heuristics.test.mjs` |
57
+
58
+ ## Testability split design
59
+
60
+ `truncate` is already a pure exported function — tests exercise it directly
61
+ at unit level. Two internal helpers (back-off check, lone-surrogate strip)
62
+ stay private; they are covered through truncate's public behavior at the
63
+ exact boundaries (cut on high surrogate, lone surrogate in input on both
64
+ paths).
@@ -0,0 +1,71 @@
1
+ # Spec: issue #61 — mention fallback lifts placeholder numbers (no left-boundary guard, kind-blind, code-span-blind)
2
+
3
+ ## Problem
4
+
5
+ The Rule 6 bare `#N` mention fallback (`src/core/code-refs.mjs`
6
+ `mentionFallback`, L767) misfires on internal-platform sessions: the real
7
+ work object was internal issue `#18` (only `acli` commands — no builtin rules
8
+ for that CLI), while a skill doc's example placeholder `#1378` won the
9
+ fallback and was written into `github.json` as the session's issue
10
+ (`strength: mention, confidence: low`).
11
+
12
+ Four gaps (issue G1-G4):
13
+
14
+ - **G2** `MENTION_RE = /#(\d{1,7})/g` has no left boundary: `##1378`
15
+ (markdown heading / double-hash typo) still matches `#1378` as a substring.
16
+ - **G4** mention counting does not distinguish prose from inline code spans:
17
+ `` `monitor pr #1378` `` (a doc example) counts the same as a real
18
+ reference.
19
+ - **G3** counting is kind-blind: `monitor pr #1378` — explicitly PR context —
20
+ feeds the *issue* candidate.
21
+ - **G1** no documented `providers.json` example for internal-platform CLIs,
22
+ so such sessions systematically slide into the fallback at all.
23
+
24
+ ## Decision (issue's fixes 1-4)
25
+
26
+ In `src/core/code-refs.mjs`:
27
+
28
+ 1. **Left-boundary guard (G2):** `MENTION_RE = /(?<![#\w])#(\d{1,7})/g` —
29
+ rejects a preceding `#` or word character; `#18`, `( #18 )`, `text #18`
30
+ still match.
31
+ 2. **Code-span exclusion (G4):** before counting, strip inline code spans
32
+ from each assistant text (`text.replace(/`[^`\n]*`/g, " ")`).
33
+ 3. **Kind-aware counting (G3):** for each surviving match, inspect a small
34
+ window before the match (16 chars); if it contains a standalone
35
+ `pr|pull|mr|merge` token (case-insensitive, word-bounded), the match does
36
+ not count toward the issue fallback (PR-context numbers must not lift an
37
+ issue). Dropped, not reclassified — the fallback is issue-only by design;
38
+ fabricating a PR mention candidate is out of scope.
39
+ 4. **Docs (G1):** README "Evidence and Code References" gains a compact
40
+ `providers.json` example: an internal-platform provider with a claim
41
+ strength rule (e.g. `acli issue update #N --assignee`) and a hosts entry,
42
+ so users can configure claim/action rules in five minutes and such
43
+ sessions stop depending on the fallback at all.
44
+
45
+ ## Non-goals
46
+
47
+ - No new builtin rules for any specific internal CLI (that's per-user config).
48
+ - No PR-side mention fallback (issue-only fallback stays).
49
+ - No fenced-block stripping (assistantTexts are prose; inline spans are the
50
+ observed vector).
51
+
52
+ ## Acceptance matrix
53
+
54
+ | ID | Feature point | Acceptance | Concrete verification | Pass criteria |
55
+ |----|---------------|------------|----------------------|---------------|
56
+ | A1 | Left boundary guard | Automated (unit) | New tests: `##1378` and `a#1`-style texts produce no mention winner; plain `#18` texts still win | Tests pass |
57
+ | A2 | Code-span exclusion | Automated (unit) | `` `monitor pr #1378` `` inside backticks contributes zero counts | Tests pass |
58
+ | A3 | Kind-aware skip | Automated (unit) | Prose `monitor pr #1378` ×5 does not produce an issue winner | Tests pass |
59
+ | A4 | Issue repro shape regresses fixed | Automated (unit) | Composite fixture: placeholder `#1378` only in guarded forms (heading / code span / pr-context) + real `#18` prose ×3 → winner is 18; without the `#18` prose → no winner | Tests pass |
60
+ | A5 | Existing fallback behavior preserved | Automated (unit + integration) | Existing tests ("picks #40 (5x) over #7 (2x)", "no winner when tied") unchanged and green; full `npm test` green | All green |
61
+ | A6 | providers.json example documented | Automated (static) | README Evidence section contains a valid `providers.json` example with a claim rule for an internal CLI + `hosts` | grep + render check |
62
+ | A7 | Scope | Automated (static) | `git diff main -- src/ README.md test/` | Only `src/core/code-refs.mjs` (MENTION_RE + mentionFallback), README Evidence section, and `test/code-refs-extract.test.mjs` |
63
+
64
+ ## Testability split design
65
+
66
+ `mentionFallback` is exercised through the public `extractCodeRefs` (same
67
+ seam as all existing mention tests) — no new export needed. Guards are pure
68
+ regex/substring logic on the existing pure function; each gap gets one
69
+ focused test plus one composite repro test (A4). The providers.json example
70
+ is docs; its JSON validity is asserted by pasting it through
71
+ `validateProvider` in a unit test (bonus guard, cheap).
@@ -0,0 +1,49 @@
1
+ # Spec: issue #63 — runner.integration "manual completion" flaky
2
+
3
+ ## Problem
4
+
5
+ `test/runner.integration.test.mjs` "runner does not clobber a manual completion
6
+ made during post-exit model passes" asserts `markCompleted()` returns
7
+ `{ok:true}` immediately after `endedAt` becomes visible in status.json. But
8
+ `runner/job-runner.mjs` `persist()` writes status.json **before** state.json
9
+ (no transaction). In the window between the two writes, `markCompleted` →
10
+ `completeView` still sees an active run (pid liveness via `loadRow` →
11
+ `isAgentBusy`) and rejects with
12
+ `{ok:false, error:'Wait for the active run to finish before marking done'}`.
13
+
14
+ CI caught this once on PR #60 (Node 24); rerun + 5 local runs passed. Pure
15
+ timing window, no behavior regression.
16
+
17
+ ## Decision (issue option 1 — poll the assertion target itself)
18
+
19
+ Replace the one-shot assertion with a `waitFor` poll of `markCompleted` until
20
+ it returns `{ok:true}` (then assert the exact shape). The assertion then checks
21
+ **durability after success** rather than racing the exact call. The existing
22
+ `waitFor(fn, timeoutMs = 15000, intervalMs = 50)` helper is reused; default
23
+ 15s timeout is ample (the race window is millisecond-scale: two evidence file
24
+ writes).
25
+
26
+ The durability assertions that are the actual behavior under test stay
27
+ unchanged: after the runner exits, `semanticState === "completed"` and
28
+ `autoState === null`.
29
+
30
+ ## Non-goals
31
+
32
+ - No production code change (`runner/`, `src/` untouched) — test-only fix.
33
+ - No refactor of `persist()` write ordering (would change production behavior;
34
+ the ordering is deliberate: endedAt converges first, see code comment).
35
+
36
+ ## Acceptance matrix
37
+
38
+ | ID | Feature point | Acceptance | Concrete verification | Pass criteria |
39
+ |----|---------------|------------|----------------------|---------------|
40
+ | A1 | Test no longer asserts `markCompleted` return immediately after `endedAt` visibility | Automated (unit/integration) | Inspect diff: one-shot `assert.deepEqual(createService(...).markCompleted(...), {ok:true})` replaced by waitFor-poll + shape assert | Code review of diff; no immediate assert after `endedAt` waitFor remains |
41
+ | A2 | Flaky assertion semantics preserved | Automated (integration) | `npm test` (node --test, the runner.integration test itself) | Test passes; durability asserts (`semanticState === "completed"`, `autoState === null`) unchanged in diff |
42
+ | A3 | Suite stability (the CI flake mode) | Automated (integration) | 10 consecutive full `npm test` runs | All 10 green, zero flakes |
43
+ | A4 | No production behavior change | Automated (static) | `git diff main --stat` in PR | Only `test/runner.integration.test.mjs` modified |
44
+
45
+ ## Testability split design
46
+
47
+ Test-only change; no new production functions. The poll helper (`waitFor`)
48
+ already exists and is exercised by every other use in the file — no new test
49
+ boundary introduced.