@zhuxixi/pi-agent-board 0.5.2 → 0.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +44 -0
- package/README.md +43 -5
- package/docs/superpowers/plans/2026-09-03-code-refs-pr-backlink-narrow.md +551 -0
- package/docs/superpowers/plans/2026-09-04-evidence-outputpreview.md +209 -0
- package/docs/superpowers/plans/2026-09-04-warm-host-reclaim.md +796 -0
- package/docs/superpowers/plans/2026-09-05-issue-13-drainnextfollowup-pty-probe.md +114 -0
- package/docs/superpowers/plans/2026-09-05-issue-38-windows-wezterm-ime-cursor.md +73 -0
- package/docs/superpowers/plans/2026-09-05-issue-39-truncate-codepoint-boundary.md +143 -0
- package/docs/superpowers/plans/2026-09-05-issue-61-mention-fallback-guards.md +226 -0
- package/docs/superpowers/plans/2026-09-05-issue-63-flaky-manual-completion.md +87 -0
- package/docs/superpowers/plans/2026-09-05-issue-64-changelog-release-helper.md +53 -0
- package/docs/superpowers/plans/2026-09-05-pty-host-stacking-sock-race.md +731 -0
- package/docs/superpowers/plans/2026-09-08-attach-ctrl-left-detach.md +30 -0
- package/docs/superpowers/plans/2026-09-08-dashboard-shrink-repaint.md +68 -0
- package/docs/superpowers/plans/2026-09-08-legacy-stale-host-recovery.md +125 -0
- package/docs/superpowers/plans/2026-09-08-spawn-async-error-swallow.md +56 -0
- package/docs/superpowers/plans/2026-09-08-stale-model-attach-guard.md +96 -0
- package/docs/superpowers/specs/2026-08-29-code-refs-badges-design.md +1 -1
- package/docs/superpowers/specs/2026-09-03-code-refs-pr-backlink-narrow-design.md +92 -0
- package/docs/superpowers/specs/2026-09-04-evidence-outputpreview-design.md +50 -0
- package/docs/superpowers/specs/2026-09-04-warm-host-reclaim-design.md +106 -0
- package/docs/superpowers/specs/2026-09-05-issue-13-drainnextfollowup-pty-probe-design.md +64 -0
- package/docs/superpowers/specs/2026-09-05-issue-38-windows-wezterm-ime-design.md +48 -0
- package/docs/superpowers/specs/2026-09-05-issue-39-truncate-codepoint-boundary-design.md +64 -0
- package/docs/superpowers/specs/2026-09-05-issue-61-mention-fallback-design.md +71 -0
- package/docs/superpowers/specs/2026-09-05-issue-63-flaky-manual-completion-design.md +49 -0
- package/docs/superpowers/specs/2026-09-05-issue-64-changelog-helper-design.md +76 -0
- package/docs/superpowers/specs/2026-09-05-pty-host-stacking-sock-race-design.md +510 -0
- package/docs/superpowers/specs/2026-09-08-attach-ctrl-left-detach-design.md +58 -0
- package/docs/superpowers/specs/2026-09-08-dashboard-shrink-repaint-design.md +52 -0
- package/docs/superpowers/specs/2026-09-08-legacy-stale-host-recovery-design.md +87 -0
- package/docs/superpowers/specs/2026-09-08-spawn-async-error-swallow-design.md +56 -0
- package/docs/superpowers/specs/2026-09-08-stale-model-attach-guard-design.md +79 -0
- package/package.json +83 -81
- package/runner/job-runner.mjs +2 -2
- package/runner/pty-runner.mjs +626 -3
- package/runner/state-runner.mjs +3 -0
- package/runner/title-runner.mjs +1 -1
- package/scripts/release_helper.mjs +277 -0
- package/src/commands/agent-board.ts +47 -35
- package/src/commands/attach-decision.mjs +66 -0
- package/src/commands/attach-flow.ts +45 -39
- package/src/commands/bg.ts +9 -0
- package/src/core/code-refs.mjs +85 -33
- package/src/core/evidence.mjs +2 -2
- package/src/core/heuristics.mjs +75 -2
- package/src/core/host-coordination.mjs +182 -0
- package/src/core/host-crash.mjs +43 -3
- package/src/core/host-probe.mjs +196 -0
- package/src/core/launch-options.mjs +17 -0
- package/src/core/launch.mjs +35 -35
- package/src/core/locks.mjs +196 -1
- package/src/core/paths.mjs +24 -0
- package/src/core/store.mjs +164 -5
- package/src/core/types.mjs +17 -1
- package/src/core/warm-host-sweeper.mjs +150 -0
- package/src/index.ts +40 -3
- package/src/runtime/service.mjs +1071 -108
- package/src/ui/dashboard-decisions.mjs +55 -0
- package/src/ui/dashboard.ts +67 -13
- package/src/ui/pty-attach.ts +23 -5
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
# Spec: warm host 回收 —— 周期 sweep + 退出清理(issue #75)
|
|
2
|
+
|
|
3
|
+
日期:2026-09-04
|
|
4
|
+
状态:draft,待用户确认
|
|
5
|
+
仓库:zhuxixi/pi-agent-board
|
|
6
|
+
|
|
7
|
+
## 1. 背景与根因
|
|
8
|
+
|
|
9
|
+
detach / 关闭宿主后,空闲 PTY host(runner + child pi)无限期空转(实测 8.5h)。根因:
|
|
10
|
+
|
|
11
|
+
1. `pruneWarmHosts()` 仅惰性触发(attach/prewarm/dispatch 时),用户 detach 后不再操作 board 则永不触发 → `AGENT_BOARD_WARM_HOST_TTL_MS`(默认 10min)与 `AGENT_BOARD_MAX_WARM_HOSTS`(默认 4)形同虚设
|
|
12
|
+
2. 宿主 pi 退出/会话切换时无清理钩子(扩展未注册 `session_shutdown`)
|
|
13
|
+
3. 无周期扫描兜底;Windows 下 runner detached 孤儿,宿主被 kill 后无任何回收
|
|
14
|
+
|
|
15
|
+
## 2. 设计决策
|
|
16
|
+
|
|
17
|
+
### M1:提取纯判定函数 `selectIdleHostsToEvict`(可测性核心)
|
|
18
|
+
|
|
19
|
+
把淘汰判定从 `pruneWarmHosts` 闭包中提取为**纯函数**(无 IO、无 env):
|
|
20
|
+
|
|
21
|
+
```
|
|
22
|
+
selectIdleHostsToEvict(rows, opts) -> { ttlEvicted: Row[], excessEvicted: Row[] }
|
|
23
|
+
|
|
24
|
+
opts = { now, maxWarm, ttlMs, graceMs, keepViewId, isBusy? }
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
- idle 定义:`hostAlive && !isBusy(row) && (host.attachedClients ?? 0) === 0 && (row.meta.id !== keepViewId)`
|
|
28
|
+
- **新增 graceMs 豁免**:`host.startedAt` 距今 < graceMs 的 host 不参与淘汰(防 ensureHost→attach 竞态;attach 重连兜底为第二保险)
|
|
29
|
+
- 排序:survivors 按 `lastActivityAt ?? startedAt` 升序,超出 maxWarm 的从最旧淘汰
|
|
30
|
+
- `isBusy` 默认 = 现 `isAgentBusy` 语义(queued/working/hasPendingQuestions),作为参数注入以保持纯性
|
|
31
|
+
|
|
32
|
+
### M2:新增调度器 `createWarmHostSweeper`(新模块 `src/core/warm-host-sweeper.mjs`)
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
createWarmHostSweeper({ sweep, intervalMs }) -> { start, stop, sweepNow }
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
- `setInterval(sweep, intervalMs)`,`unref()`(不阻塞宿主退出);`sweepNow()` 立即执行一次(错误吞掉,best-effort)
|
|
39
|
+
- `stop()` 清定时器;`start()` 幂等
|
|
40
|
+
- 新环境变量 `AGENT_BOARD_SWEEP_INTERVAL_MS`(默认 60_000;0 = 禁用周期 sweep,仅保留启动/session_shutdown 时 sweepNow)
|
|
41
|
+
|
|
42
|
+
### M3:接线(src/index.ts,扩展模块级)
|
|
43
|
+
|
|
44
|
+
- 扩展加载时:**非 isHostedChild** 才创建 sweeper(`sweep = () => pruneWarmHosts({})`),`start()` + 立即 `sweepNow()`(回收上一个宿主遗留的泄漏 host)
|
|
45
|
+
- 注册 `session_shutdown`:`sweepNow()`(尽力清理)+ `stop()`
|
|
46
|
+
- **isHostedChild 必须跳过**:child pi 内加载扩展时若启动 sweeper,会向自己的 runner 发 terminate(自杀链),绝对禁止
|
|
47
|
+
- 新增环境变量 `AGENT_BOARD_WARM_HOST_GRACE_MS`(默认 30_000;0 = 无豁免,保持旧行为可调)
|
|
48
|
+
|
|
49
|
+
### M4:dashboard 打开期间顺带 prune(低成本增强)
|
|
50
|
+
|
|
51
|
+
`openDashboard` 的 POLL_MS 轮询回调里加一次 `pruneWarmHosts()`(dashboard 打开 = 用户正用 board,实时回收)。保持既有惰性调用点不变。
|
|
52
|
+
|
|
53
|
+
### Non-goals(明确不做)
|
|
54
|
+
|
|
55
|
+
- **runner 自退出兜底**:runner 无权威 busy 信号(state.json 是 service 层写的,跨进程读取引入耦合/竞态);误杀风险(PRD stretch "detach while running 转 headless worker" 未实现时,pty child 可能正跑任务)。宿主被杀场景由"下一次任意宿主启动 sweepNow"覆盖,残余缺口(永不再开任何 pi)接受。
|
|
56
|
+
- **detach-while-running 转 headless worker**:PRD stretch,独立议题。
|
|
57
|
+
- 不改变 terminate 消息协议(沿用 `{type:"terminate"}` 既有语义:杀 child,runner 收尾退出)。
|
|
58
|
+
|
|
59
|
+
## 3. 数据流
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
宿主扩展实例(每个 pi session 一个,随 session_shutdown 销毁)
|
|
63
|
+
├─ module load(非 child): createWarmHostSweeper(...).start() + sweepNow()
|
|
64
|
+
├─ setInterval(60s, unref): pruneWarmHosts({}) → selectIdleHostsToEvict → sendHostMessage(terminate)
|
|
65
|
+
├─ dashboard POLL_MS: pruneWarmHosts({})(UI 打开时)
|
|
66
|
+
└─ session_shutdown: sweepNow() + stop()
|
|
67
|
+
|
|
68
|
+
selectIdleHostsToEvict(rows, opts) ← 纯判定(无副作用)
|
|
69
|
+
pruneWarmHosts(opts) ← 薄执行层(listRows + select + sendHostMessage),既有惰性调用点行为等价
|
|
70
|
+
sendHostMessage(row, terminate) ← 既有 fire-and-forget socket 路径(#71 协议)
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
## 4. 验收矩阵
|
|
74
|
+
|
|
75
|
+
| ID | 功能点 | 验收方式 | 具体验证 | 通过标准 |
|
|
76
|
+
|----|--------|----------|----------|----------|
|
|
77
|
+
| A1 | selectIdleHostsToEvict:TTL 到期淘汰、未到期保留、keepViewId 豁免 | 自动化(unit) | `node --test test/warm-host-sweeper.test.mjs` | 过期 idle host 进入 ttlEvicted;未过期/keepViewId 不进入 |
|
|
78
|
+
| A2 | busy(queued/working/pendingQuestions)与 attachedClients>0 永不淘汰 | 自动化(unit) | 同上 | 两类 host 均不进入任何淘汰列表 |
|
|
79
|
+
| A3 | graceMs 豁免:startedAt 距今 < grace 的 host 不淘汰 | 自动化(unit) | 同上 | 豁免窗口内不进淘汰列表;窗口外正常淘汰 |
|
|
80
|
+
| A4 | maxWarm 超额时淘汰最旧 survivors(按 lastActivityAt/startedAt) | 自动化(unit) | 同上 | 淘汰对象为排序最旧者,数量 = survivors - maxWarm |
|
|
81
|
+
| A5 | sweeper 调度:start 定时触发、stop 停止、sweepNow 立即执行、unref | 自动化(unit) | 同上(node:test mock timers) | fake timer 前进 intervalMs 后 sweep 被调用;stop 后不再调用;sweepNow 立即调用 |
|
|
82
|
+
| A6 | 接线:非 child 宿主启动即 sweepNow + 周期 sweep;isHostedChild 不接线 | 自动化(integration) | `node --test test/service.test.mjs` 增补:临时 root + node:net 假 server 捕获 terminate | idle host 的 socketPath 收到 terminate;child 模式(flag)下无 sweeper 启动 |
|
|
83
|
+
| A7 | session_shutdown 触发 sweepNow(退出清理) | 自动化(integration) | 同上:调用注册的 shutdown 回调 | terminate 发出;sweeper 已 stop |
|
|
84
|
+
| A8 | 重构无回归:pruneWarmHosts 既有惰性调用点行为等价 | 自动化(全量) | `node --test test/*.test.mjs`(基线 409+ 含既有失败清单) | 与修复前基线一致,无新增失败 |
|
|
85
|
+
| U1 | 真实场景回收 | 用户实测 | 开会话→attach→detach→设 `AGENT_BOARD_WARM_HOST_TTL_MS=10s` + `AGENT_BOARD_SWEEP_INTERVAL_MS=5s`→观察 | TTL 过后 host.json 变 exited,runner/child 进程退出,dashboard 状态正确 |
|
|
86
|
+
|
|
87
|
+
## 5. 可测性拆分设计(实现硬约束)
|
|
88
|
+
|
|
89
|
+
1. **`selectIdleHostsToEvict`**:纯函数,输入 rows + opts,输出淘汰分组;不触盘、不连 socket、不读 process.env(所有阈值参数显式传入)。测试直接构造内存 rows。**边界**:判定逻辑全部在此函数内;`pruneWarmHosts` 只做 IO 编排,不得内联任何判定分支。
|
|
90
|
+
2. **`createWarmHostSweeper`**:调度器封装,`sweep` 回调依赖注入;timer 用全局 setInterval(node:test mock timers 可控)。**边界**:不 import service/store;与业务零耦合。
|
|
91
|
+
3. **`attachWarmHostSweeperToLifecycle(piLike, deps)`**(index.ts 内接线函数):接收 `pi`(on 事件注册)与 `deps = { createSweeper, isHostedChild, sweep }`,可注入 fake 断言注册行为。**边界**:index.ts 只调这个函数,不内联生命周期逻辑。
|
|
92
|
+
4. **`pruneWarmHosts`**:重构为薄执行层(listRows → selectIdleHostsToEvict → sendHostMessage),原惰性调用点(dispatch/ensureHost/dashboard)签名不变,行为等价(A8 回归保障)。
|
|
93
|
+
|
|
94
|
+
## 6. 风险与缓解
|
|
95
|
+
|
|
96
|
+
- 竞态(attach 与 sweep 同时):graceMs + attach 重连兜底(既有 #48 L3)
|
|
97
|
+
- child pi 自杀链:isHostedChild 硬跳过(A6 覆盖)
|
|
98
|
+
- 多宿主并发 sweep 重复 terminate:terminate 幂等(sendHostMessage 失败静默、runner 已死则 socket error 吞掉)
|
|
99
|
+
- sweep 与 archive/stop 竞态:单进程内同 service 数据视图,均 fire-and-forget,最坏重复 terminate(幂等)
|
|
100
|
+
|
|
101
|
+
## 7. 测试计划摘要
|
|
102
|
+
|
|
103
|
+
- 新增 `test/warm-host-sweeper.test.mjs`:A1-A5(unit,含 mock timers)
|
|
104
|
+
- 增补 `test/service.test.mjs`:A6-A7(integration,net 假 server 捕获 terminate)
|
|
105
|
+
- 全量回归:A8
|
|
106
|
+
- U1 用户实测清单写入 PR 描述与 issue 评论
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# Spec: issue #13 — drainNextFollowUp forces PTY probe refresh on the 700ms poll path
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
`src/runtime/service.mjs` `drainNextFollowUp` (L496) calls
|
|
6
|
+
`ptySupport({ refresh: true })` and is reachable from the 700ms `reconcile()`
|
|
7
|
+
poll (L1007-1011: queued follow-up + `canAutoDrain` → `drainNextFollowUp`).
|
|
8
|
+
When a queued follow-up repeatedly fails to start, every poll cycle forces a
|
|
9
|
+
fresh PTY probe — `ptySpawnSupported` does a **real `pty.spawn`** process.
|
|
10
|
+
This violates the probe-cache discipline (`shouldProbePtySupport`): success
|
|
11
|
+
cached for process lifetime, failures retried on a 2s TTL; only `refresh`
|
|
12
|
+
bypasses it.
|
|
13
|
+
|
|
14
|
+
`dispatch` (L616) and `reply` (L649) also use `refresh: true` but are explicit
|
|
15
|
+
user actions — fine as-is, untouched.
|
|
16
|
+
|
|
17
|
+
## Decision (issue's fix sketch, verbatim)
|
|
18
|
+
|
|
19
|
+
Replace in `drainNextFollowUp`:
|
|
20
|
+
|
|
21
|
+
```js
|
|
22
|
+
const pty = ptySupport({ refresh: true });
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
with default cached-probe semantics:
|
|
26
|
+
|
|
27
|
+
```js
|
|
28
|
+
const pty = ptySupport();
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Plus a red-green test asserting the injected `ptySupport` never receives
|
|
32
|
+
`refresh: true` from the reconcile→drain path (pattern from PR #12's service
|
|
33
|
+
test "ensureHost probes PTY support with TTL cache, not forced refresh",
|
|
34
|
+
test/service.test.mjs:940).
|
|
35
|
+
|
|
36
|
+
## Behavior contract
|
|
37
|
+
|
|
38
|
+
- Poll-path drains use the cache: after a successful probe, zero further real
|
|
39
|
+
spawns; after a failed probe, at most one real spawn per 2s window.
|
|
40
|
+
- Explicit user actions (`dispatch`, `reply`) keep forced refresh.
|
|
41
|
+
- No other behavior change in drain logic (claim/launch/complete paths
|
|
42
|
+
untouched).
|
|
43
|
+
|
|
44
|
+
## Non-goals
|
|
45
|
+
|
|
46
|
+
- No change to `dispatch`/`reply` refresh semantics.
|
|
47
|
+
- No change to `ptySpawnSupported`/`shouldProbePtySupport` internals.
|
|
48
|
+
|
|
49
|
+
## Acceptance matrix
|
|
50
|
+
|
|
51
|
+
| ID | Feature point | Acceptance | Concrete verification | Pass criteria |
|
|
52
|
+
|----|---------------|------------|----------------------|---------------|
|
|
53
|
+
| A1 | reconcile→drain path never forces probe refresh | Automated (unit) | New test in `test/service.test.mjs`: idle row + queued follow-up → `svc.reconcile()` → injected ptySupport spy records opts; assert `probeCalls.length >= 1` and no call has `refresh === true` | New test passes; fails (red) if production line keeps `refresh: true` |
|
|
54
|
+
| A2 | Explicit user actions unaffected | Automated (unit) | Existing service tests (dispatch/reply paths) green in full suite | `npm test` all green |
|
|
55
|
+
| A3 | Full regression | Automated (integration) | Full `npm test` | All pass |
|
|
56
|
+
| A4 | Change scope | Automated (static) | `git diff main` review | Production delta is exactly the one-line probe switch in `src/runtime/service.mjs`; test delta is one new test |
|
|
57
|
+
|
|
58
|
+
## Testability split design
|
|
59
|
+
|
|
60
|
+
Test-only seam already exists: `createService` opts inject `ptySupport`
|
|
61
|
+
(spy) and `launch` (fake, avoids real spawn). The new test drives the public
|
|
62
|
+
`reconcile()` API so the exact poll path is covered. Row state constructed
|
|
63
|
+
via `readState`/`writeState` (idle + exited) so `canAutoDrain` holds. No new
|
|
64
|
+
production functions.
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# Spec: issue #38 — document Windows WezTerm IME hardware-cursor requirement
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
With #24/#28 shipped (v0.4.3), the IME candidate window is pinned and
|
|
6
|
+
flicker-free on Linux (WezTerm + fcitx5/X11). On **Windows WezTerm (WSL2
|
|
7
|
+
backend)** the candidate window stays stuck at the right edge: per upstream
|
|
8
|
+
earendil-works/pi#5200, Windows WezTerm only updates the IME candidate
|
|
9
|
+
position when the **hardware cursor is visible**, and pi-tui hides it by
|
|
10
|
+
default. The visible block cursor in the editor is a fake (reverse-video
|
|
11
|
+
content), so "I can see the cursor" does not imply the hardware cursor is on.
|
|
12
|
+
|
|
13
|
+
Fix (verified on the real environment, per the issue): `PI_HARDWARE_CURSOR=1`
|
|
14
|
+
env or `"showHardwareCursor": true` in pi's settings.json. Trade-off: the
|
|
15
|
+
real terminal cursor becomes visible in the TUI (cosmetic only; harmless on
|
|
16
|
+
Linux).
|
|
17
|
+
|
|
18
|
+
## Decision
|
|
19
|
+
|
|
20
|
+
Add one Troubleshooting subsection to README.md:
|
|
21
|
+
|
|
22
|
+
- Title: IME candidate window stuck at the right edge (Windows WezTerm)
|
|
23
|
+
- Content: symptom + platform scope (Windows WezTerm → WSL2, Windows IME/TSF;
|
|
24
|
+
Linux unaffected), root cause in one sentence (hardware cursor hidden by
|
|
25
|
+
default; Windows WezTerm needs it visible to track IME position; the editor
|
|
26
|
+
block cursor is fake content, not the hardware cursor), both fixes with
|
|
27
|
+
sync/scope guidance (machine-local env vs synced settings.json), and the
|
|
28
|
+
cosmetic trade-off.
|
|
29
|
+
|
|
30
|
+
Placement: after "### Attach is slow or keeps reconnecting" (IME is an
|
|
31
|
+
attach-surface concern, keeps related attach topics adjacent).
|
|
32
|
+
|
|
33
|
+
## Non-goals
|
|
34
|
+
|
|
35
|
+
- No code changes.
|
|
36
|
+
- No upstream comment on earendil-works/pi#5200 (owner's voice, left to the
|
|
37
|
+
repo owner — noted in the issue comment).
|
|
38
|
+
|
|
39
|
+
## Acceptance matrix
|
|
40
|
+
|
|
41
|
+
| ID | Feature point | Acceptance | Concrete verification | Pass criteria |
|
|
42
|
+
|----|---------------|------------|----------------------|---------------|
|
|
43
|
+
| A1 | Troubleshooting entry exists with both fixes | Automated (static) | grep README.md for the new section + `PI_HARDWARE_CURSOR=1` + `showHardwareCursor` | All three present in the Troubleshooting section |
|
|
44
|
+
| A2 | Scope: docs-only change | Automated (static) | `git diff main --stat` | Only README.md (plus process docs under docs/superpowers/) |
|
|
45
|
+
|
|
46
|
+
## Testability split design
|
|
47
|
+
|
|
48
|
+
Docs-only: verification is content presence + scope, no behavioral seam.
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# Spec: issue #39 — truncate() cuts surrogate pairs, producing lone surrogates (U+FFFD) and wrap-misaligned dashboard rows
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
`src/core/heuristics.mjs` `truncate(s, n)` uses `str.length` and `str.slice()`
|
|
6
|
+
(UTF-16 code units). When the cut lands inside a surrogate pair (e.g. emoji,
|
|
7
|
+
2 units), the output keeps a lone high surrogate → renders as U+FFFD →
|
|
8
|
+
pi-tui counts it 1 column while WezTerm renders 2 → line over-wide → terminal
|
|
9
|
+
wrap → TUI diff-render row misalignment → stacked stale frames on the
|
|
10
|
+
dashboard (issue's repro chain).
|
|
11
|
+
|
|
12
|
+
`truncate` feeds `deriveSummary` (max=80), `events.mjs`
|
|
13
|
+
(`latestAssistantPreview`), `evidence.mjs`, `auto-state.mjs` — all display
|
|
14
|
+
text paths.
|
|
15
|
+
|
|
16
|
+
## Decision (fix direction A, repo-side)
|
|
17
|
+
|
|
18
|
+
Two changes inside `truncate`:
|
|
19
|
+
|
|
20
|
+
1. **Code-point-safe cut (cut-point back-off, NOT code-point budget).** Keep
|
|
21
|
+
the existing UTF-16 unit budget `n` (all callers pass display budgets; CJK
|
|
22
|
+
chars occupy 2 units — switching to a code-point count would double the
|
|
23
|
+
effective width of CJK summaries from 80 to ~160 columns and *reintroduce*
|
|
24
|
+
over-wide lines). Instead: if the character just before the cut point is a
|
|
25
|
+
high surrogate whose low surrogate is the first dropped unit, back the cut
|
|
26
|
+
off by one unit so the pair stays whole. Output length only shrinks by one
|
|
27
|
+
unit in that rare case; the ellipsis still fits the budget.
|
|
28
|
+
|
|
29
|
+
2. **Defensive lone-surrogate strip on output.** Input text (from agent
|
|
30
|
+
output stored in status.json) may already contain lone surrogates; strip
|
|
31
|
+
them from the returned string on both the short path and the truncated
|
|
32
|
+
path (`/[\uD800-\uDBFF](?![\uDC00-\uDFFF])/g` and the mirrored
|
|
33
|
+
low-surrogate regex).
|
|
34
|
+
|
|
35
|
+
Out of scope (per goal boundary): pi-tui width-table changes for Ambiguous
|
|
36
|
+
characters (fix direction B) and terminal-side clamping (direction C) —
|
|
37
|
+
upstream/external. Completion here = truncate no longer produces or passes
|
|
38
|
+
through lone surrogates; "visual overlap fully gone" additionally depends on
|
|
39
|
+
those upstream items.
|
|
40
|
+
|
|
41
|
+
## Non-goals
|
|
42
|
+
|
|
43
|
+
- No `Intl.Segmenter` grapheme segmentation (ZWJ emoji families remain
|
|
44
|
+
multi-codepoint; the defect being fixed is lone surrogates, not grapheme
|
|
45
|
+
splitting).
|
|
46
|
+
- No caller changes; no width-function changes.
|
|
47
|
+
|
|
48
|
+
## Acceptance matrix
|
|
49
|
+
|
|
50
|
+
| ID | Feature point | Acceptance | Concrete verification | Pass criteria |
|
|
51
|
+
|----|---------------|------------|----------------------|---------------|
|
|
52
|
+
| A1 | Cut never splits a surrogate pair | Automated (unit) | New tests in `test/heuristics.test.mjs`: e.g. `truncate("a👍b", 3) === "a…"` (back-off) and a long-string case whose cut lands on an emoji; assert output matches `/\uD83D$/` never (no trailing lone high surrogate) | Tests pass |
|
|
53
|
+
| A2 | Output is lone-surrogate-free even when input isn't | Automated (unit) | `truncate("ab\ud83d", 10) === "ab"`, `truncate("ab\udc4dzzzzzzzzzz", 5)` has no lone low surrogate | Tests pass |
|
|
54
|
+
| A3 | Existing behavior preserved for ASCII and CJK | Automated (unit) | Existing tests unchanged (`truncate("hello world", 5) === "hell…"`) + new: `truncate("一二三四五", 5) === "一二…"` (2-unit budget semantics kept) | Tests pass |
|
|
55
|
+
| A4 | Full regression | Automated (integration) | Full `npm test` | All green |
|
|
56
|
+
| A5 | Scope | Automated (static) | `git diff main -- src/ test/` | Only `src/core/heuristics.mjs` (truncate) and `test/heuristics.test.mjs` |
|
|
57
|
+
|
|
58
|
+
## Testability split design
|
|
59
|
+
|
|
60
|
+
`truncate` is already a pure exported function — tests exercise it directly
|
|
61
|
+
at unit level. Two internal helpers (back-off check, lone-surrogate strip)
|
|
62
|
+
stay private; they are covered through truncate's public behavior at the
|
|
63
|
+
exact boundaries (cut on high surrogate, lone surrogate in input on both
|
|
64
|
+
paths).
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# Spec: issue #61 — mention fallback lifts placeholder numbers (no left-boundary guard, kind-blind, code-span-blind)
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
The Rule 6 bare `#N` mention fallback (`src/core/code-refs.mjs`
|
|
6
|
+
`mentionFallback`, L767) misfires on internal-platform sessions: the real
|
|
7
|
+
work object was internal issue `#18` (only `acli` commands — no builtin rules
|
|
8
|
+
for that CLI), while a skill doc's example placeholder `#1378` won the
|
|
9
|
+
fallback and was written into `github.json` as the session's issue
|
|
10
|
+
(`strength: mention, confidence: low`).
|
|
11
|
+
|
|
12
|
+
Four gaps (issue G1-G4):
|
|
13
|
+
|
|
14
|
+
- **G2** `MENTION_RE = /#(\d{1,7})/g` has no left boundary: `##1378`
|
|
15
|
+
(markdown heading / double-hash typo) still matches `#1378` as a substring.
|
|
16
|
+
- **G4** mention counting does not distinguish prose from inline code spans:
|
|
17
|
+
`` `monitor pr #1378` `` (a doc example) counts the same as a real
|
|
18
|
+
reference.
|
|
19
|
+
- **G3** counting is kind-blind: `monitor pr #1378` — explicitly PR context —
|
|
20
|
+
feeds the *issue* candidate.
|
|
21
|
+
- **G1** no documented `providers.json` example for internal-platform CLIs,
|
|
22
|
+
so such sessions systematically slide into the fallback at all.
|
|
23
|
+
|
|
24
|
+
## Decision (issue's fixes 1-4)
|
|
25
|
+
|
|
26
|
+
In `src/core/code-refs.mjs`:
|
|
27
|
+
|
|
28
|
+
1. **Left-boundary guard (G2):** `MENTION_RE = /(?<![#\w])#(\d{1,7})/g` —
|
|
29
|
+
rejects a preceding `#` or word character; `#18`, `( #18 )`, `text #18`
|
|
30
|
+
still match.
|
|
31
|
+
2. **Code-span exclusion (G4):** before counting, strip inline code spans
|
|
32
|
+
from each assistant text (`text.replace(/`[^`\n]*`/g, " ")`).
|
|
33
|
+
3. **Kind-aware counting (G3):** for each surviving match, inspect a small
|
|
34
|
+
window before the match (16 chars); if it contains a standalone
|
|
35
|
+
`pr|pull|mr|merge` token (case-insensitive, word-bounded), the match does
|
|
36
|
+
not count toward the issue fallback (PR-context numbers must not lift an
|
|
37
|
+
issue). Dropped, not reclassified — the fallback is issue-only by design;
|
|
38
|
+
fabricating a PR mention candidate is out of scope.
|
|
39
|
+
4. **Docs (G1):** README "Evidence and Code References" gains a compact
|
|
40
|
+
`providers.json` example: an internal-platform provider with a claim
|
|
41
|
+
strength rule (e.g. `acli issue update #N --assignee`) and a hosts entry,
|
|
42
|
+
so users can configure claim/action rules in five minutes and such
|
|
43
|
+
sessions stop depending on the fallback at all.
|
|
44
|
+
|
|
45
|
+
## Non-goals
|
|
46
|
+
|
|
47
|
+
- No new builtin rules for any specific internal CLI (that's per-user config).
|
|
48
|
+
- No PR-side mention fallback (issue-only fallback stays).
|
|
49
|
+
- No fenced-block stripping (assistantTexts are prose; inline spans are the
|
|
50
|
+
observed vector).
|
|
51
|
+
|
|
52
|
+
## Acceptance matrix
|
|
53
|
+
|
|
54
|
+
| ID | Feature point | Acceptance | Concrete verification | Pass criteria |
|
|
55
|
+
|----|---------------|------------|----------------------|---------------|
|
|
56
|
+
| A1 | Left boundary guard | Automated (unit) | New tests: `##1378` and `a#1`-style texts produce no mention winner; plain `#18` texts still win | Tests pass |
|
|
57
|
+
| A2 | Code-span exclusion | Automated (unit) | `` `monitor pr #1378` `` inside backticks contributes zero counts | Tests pass |
|
|
58
|
+
| A3 | Kind-aware skip | Automated (unit) | Prose `monitor pr #1378` ×5 does not produce an issue winner | Tests pass |
|
|
59
|
+
| A4 | Issue repro shape regresses fixed | Automated (unit) | Composite fixture: placeholder `#1378` only in guarded forms (heading / code span / pr-context) + real `#18` prose ×3 → winner is 18; without the `#18` prose → no winner | Tests pass |
|
|
60
|
+
| A5 | Existing fallback behavior preserved | Automated (unit + integration) | Existing tests ("picks #40 (5x) over #7 (2x)", "no winner when tied") unchanged and green; full `npm test` green | All green |
|
|
61
|
+
| A6 | providers.json example documented | Automated (static) | README Evidence section contains a valid `providers.json` example with a claim rule for an internal CLI + `hosts` | grep + render check |
|
|
62
|
+
| A7 | Scope | Automated (static) | `git diff main -- src/ README.md test/` | Only `src/core/code-refs.mjs` (MENTION_RE + mentionFallback), README Evidence section, and `test/code-refs-extract.test.mjs` |
|
|
63
|
+
|
|
64
|
+
## Testability split design
|
|
65
|
+
|
|
66
|
+
`mentionFallback` is exercised through the public `extractCodeRefs` (same
|
|
67
|
+
seam as all existing mention tests) — no new export needed. Guards are pure
|
|
68
|
+
regex/substring logic on the existing pure function; each gap gets one
|
|
69
|
+
focused test plus one composite repro test (A4). The providers.json example
|
|
70
|
+
is docs; its JSON validity is asserted by pasting it through
|
|
71
|
+
`validateProvider` in a unit test (bonus guard, cheap).
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
# Spec: issue #63 — runner.integration "manual completion" flaky
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
`test/runner.integration.test.mjs` "runner does not clobber a manual completion
|
|
6
|
+
made during post-exit model passes" asserts `markCompleted()` returns
|
|
7
|
+
`{ok:true}` immediately after `endedAt` becomes visible in status.json. But
|
|
8
|
+
`runner/job-runner.mjs` `persist()` writes status.json **before** state.json
|
|
9
|
+
(no transaction). In the window between the two writes, `markCompleted` →
|
|
10
|
+
`completeView` still sees an active run (pid liveness via `loadRow` →
|
|
11
|
+
`isAgentBusy`) and rejects with
|
|
12
|
+
`{ok:false, error:'Wait for the active run to finish before marking done'}`.
|
|
13
|
+
|
|
14
|
+
CI caught this once on PR #60 (Node 24); rerun + 5 local runs passed. Pure
|
|
15
|
+
timing window, no behavior regression.
|
|
16
|
+
|
|
17
|
+
## Decision (issue option 1 — poll the assertion target itself)
|
|
18
|
+
|
|
19
|
+
Replace the one-shot assertion with a `waitFor` poll of `markCompleted` until
|
|
20
|
+
it returns `{ok:true}` (then assert the exact shape). The assertion then checks
|
|
21
|
+
**durability after success** rather than racing the exact call. The existing
|
|
22
|
+
`waitFor(fn, timeoutMs = 15000, intervalMs = 50)` helper is reused; default
|
|
23
|
+
15s timeout is ample (the race window is millisecond-scale: two evidence file
|
|
24
|
+
writes).
|
|
25
|
+
|
|
26
|
+
The durability assertions that are the actual behavior under test stay
|
|
27
|
+
unchanged: after the runner exits, `semanticState === "completed"` and
|
|
28
|
+
`autoState === null`.
|
|
29
|
+
|
|
30
|
+
## Non-goals
|
|
31
|
+
|
|
32
|
+
- No production code change (`runner/`, `src/` untouched) — test-only fix.
|
|
33
|
+
- No refactor of `persist()` write ordering (would change production behavior;
|
|
34
|
+
the ordering is deliberate: endedAt converges first, see code comment).
|
|
35
|
+
|
|
36
|
+
## Acceptance matrix
|
|
37
|
+
|
|
38
|
+
| ID | Feature point | Acceptance | Concrete verification | Pass criteria |
|
|
39
|
+
|----|---------------|------------|----------------------|---------------|
|
|
40
|
+
| A1 | Test no longer asserts `markCompleted` return immediately after `endedAt` visibility | Automated (unit/integration) | Inspect diff: one-shot `assert.deepEqual(createService(...).markCompleted(...), {ok:true})` replaced by waitFor-poll + shape assert | Code review of diff; no immediate assert after `endedAt` waitFor remains |
|
|
41
|
+
| A2 | Flaky assertion semantics preserved | Automated (integration) | `npm test` (node --test, the runner.integration test itself) | Test passes; durability asserts (`semanticState === "completed"`, `autoState === null`) unchanged in diff |
|
|
42
|
+
| A3 | Suite stability (the CI flake mode) | Automated (integration) | 10 consecutive full `npm test` runs | All 10 green, zero flakes |
|
|
43
|
+
| A4 | No production behavior change | Automated (static) | `git diff main --stat` in PR | Only `test/runner.integration.test.mjs` modified |
|
|
44
|
+
|
|
45
|
+
## Testability split design
|
|
46
|
+
|
|
47
|
+
Test-only change; no new production functions. The poll helper (`waitFor`)
|
|
48
|
+
already exists and is exercised by every other use in the file — no new test
|
|
49
|
+
boundary introduced.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Spec: issue #64 — CHANGELOG.md driven by conventional commits (port of jfox release helper)
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
No CHANGELOG exists; releases are bare version squashes (`0.5.2 (#74)`).
|
|
6
|
+
npm users and gallery visitors have no "what changed in this version" entry
|
|
7
|
+
point. Commits are already highly conventional (QA baseline #32 onward), so
|
|
8
|
+
generation is mechanical.
|
|
9
|
+
|
|
10
|
+
## Decision (jfox approach, Node port)
|
|
11
|
+
|
|
12
|
+
Port the jfox release helper (`zhuxixi/jfox/.claude/skills/release/release_helper.py`)
|
|
13
|
+
to `scripts/release_helper.mjs` (repo stack: Node ESM). Architecture:
|
|
14
|
+
|
|
15
|
+
- **Pure, exported functions** (unit-testable, no git in unit tests):
|
|
16
|
+
- `parseCommitLines(lines)` — conventional-commit parse: type
|
|
17
|
+
(feat/fix/perf/refactor/docs/chore/test + `!`), scope, message, trailing
|
|
18
|
+
`(#N)` PR number; skip `bump version`-ish and pure merge subjects;
|
|
19
|
+
dedupe.
|
|
20
|
+
- `generateChangelog({ version, date, entries, prevTag })` —
|
|
21
|
+
`## [version] - date` section grouped Features / Fixes / Performance /
|
|
22
|
+
Changes, each entry `- message (#N)`, plus compare link
|
|
23
|
+
`https://github.com/zhuxixi/pi-agent-board/compare/vPREV...vNEXT`.
|
|
24
|
+
- `changelogTopPrs(text)` — PR numbers inside the first `## [...]` section.
|
|
25
|
+
- `verifyFrom({ functionalLines, changelogText })` — functional whitelist
|
|
26
|
+
(feat/fix/refactor/docs/perf, skipping merges / version-bump commits /
|
|
27
|
+
`docs(changelog)` maintenance commits to avoid a verify fix loop), diff
|
|
28
|
+
against top-section PRs → `{ ok, missing, extra, functionalCommits }`,
|
|
29
|
+
fail-closed on bad input.
|
|
30
|
+
- **CLI** (git-backed, thin): `node scripts/release_helper.mjs [--dry-run]
|
|
31
|
+
[patch|minor|major|X.Y.Z]` (default patch; version read from
|
|
32
|
+
package.json; `--dry-run` prints JSON preview without touching files;
|
|
33
|
+
apply inserts the section at the top of CHANGELOG.md) and
|
|
34
|
+
`node scripts/release_helper.mjs verify` (exit 1 + missing list when
|
|
35
|
+
`last v* tag..HEAD` functional commits' PR numbers are absent from the
|
|
36
|
+
CHANGELOG top section — guards against a PR merged after the changelog
|
|
37
|
+
was generated).
|
|
38
|
+
- **Forward-only**: CHANGELOG.md is created with a header stub only; entries
|
|
39
|
+
start from the next release after this lands. No backfill of ≤0.5.2.
|
|
40
|
+
- **npm script**: `"changelog": "node scripts/release_helper.mjs"` so
|
|
41
|
+
`npm run changelog -- --dry-run` works (issue acceptance wording).
|
|
42
|
+
- **README Publishing section** updated to the new flow (changelog BEFORE
|
|
43
|
+
`npm version`, because the bump commit+tag would empty the generation
|
|
44
|
+
range): `npm run verify` → `npm run changelog -- --dry-run` (review) →
|
|
45
|
+
`npm run changelog -- <bump>` (insert section) → `verify` must exit 0 →
|
|
46
|
+
commit CHANGELOG → `npm version <bump>` → `npm publish` → GitHub Release
|
|
47
|
+
with notes from the top section.
|
|
48
|
+
|
|
49
|
+
Tag format: this repo tags `vX.Y.Z` (verified: v0.4.2..v0.5.2), same as
|
|
50
|
+
jfox's `--match v*` logic.
|
|
51
|
+
|
|
52
|
+
## Non-goals
|
|
53
|
+
|
|
54
|
+
- No GitHub Release automation (manual step, per jfox flow).
|
|
55
|
+
- No backfilling historical versions.
|
|
56
|
+
- No changes to the existing `verify` npm script.
|
|
57
|
+
|
|
58
|
+
## Acceptance matrix
|
|
59
|
+
|
|
60
|
+
| ID | Feature point | Acceptance | Concrete verification | Pass criteria |
|
|
61
|
+
|----|---------------|------------|----------------------|---------------|
|
|
62
|
+
| A1 | dry-run preview | Automated (integration) | `npm run changelog -- --dry-run` exits 0, prints JSON with `changelog_preview` containing `## [0.6.0]` or next-patch section and grouped entries from v0.5.2..HEAD; CHANGELOG.md unmodified | Command output assertions |
|
|
63
|
+
| A2 | apply inserts top section | Automated (unit) | Unit test on a temp file via exported `insertSection`/CLI with fixture dir: header-stub CHANGELOG gains the new `## [...]` section above prior content | Test passes |
|
|
64
|
+
| A3 | parse correctness | Automated (unit) | Fixtures: `feat(scope): msg (#12)`, `fix: msg (#13)` with `!`, bare merge, `0.5.2 (#74)` version squash (skipped), non-conventional fallback, dedupe | Test passes |
|
|
65
|
+
| A4 | verify missing detection | Automated (unit + integration) | Unit: constructed lines/text → missing/extra/fail-closed paths. Integration: `node scripts/release_helper.mjs verify` on this branch (CHANGELOG stub, v0.5.2..HEAD has functional PRs #77-#82) exits 1 with missing listing those PRs | Both pass |
|
|
66
|
+
| A5 | Full regression | Automated (integration) | `npm test` | All green |
|
|
67
|
+
| A6 | Publishing docs updated | Automated (static) | README Publishing section mentions the changelog script and flow | grep |
|
|
68
|
+
| A7 | Scope | Automated (static) | `git diff main --stat` | Only scripts/release_helper.mjs, test/release-helper.test.mjs, package.json (one script line), CHANGELOG.md (stub), README.md |
|
|
69
|
+
|
|
70
|
+
## Testability split design
|
|
71
|
+
|
|
72
|
+
Pure functions take plain inputs (string lines, text, options) and return
|
|
73
|
+
plain data — no fs/git in unit tests. The CLI layer (git + fs) is covered by
|
|
74
|
+
two integration assertions (A1 dry-run on the real repo, A4 verify on the
|
|
75
|
+
real repo) executed as plain commands in the workflow, plus a tmp-dir apply
|
|
76
|
+
test via the exported path. This mirrors the repo's existing DI style.
|