tianshu-mcp 0.5.4 → 0.5.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.en.md +46 -2
  2. package/CHANGELOG.md +46 -2
  3. package/README.en.md +502 -483
  4. package/README.md +504 -485
  5. package/dist/agents/builtin.js +63 -0
  6. package/dist/agents/kimicode/adapter.js +67 -0
  7. package/dist/agents/kimicode/cdp.js +636 -0
  8. package/dist/agents/kimicode/dialog.js +295 -0
  9. package/dist/agents/kimicode/discovery.js +159 -0
  10. package/dist/agents/kimicode/dom.js +346 -0
  11. package/dist/agents/kimicode/instance.js +249 -0
  12. package/dist/agents/kimicode/liveness.js +128 -0
  13. package/dist/agents/kimicode/model.js +158 -0
  14. package/dist/agents/kimicode/recovery.js +128 -0
  15. package/dist/agents/kimicode/run.js +985 -0
  16. package/dist/agents/kimicode/selectors.js +371 -0
  17. package/dist/agents/kimicode/session.js +63 -0
  18. package/dist/agents/kimicode/workspace.js +62 -0
  19. package/dist/agents/qoder/adapter.js +67 -0
  20. package/dist/agents/qoder/cdp.js +130 -0
  21. package/dist/agents/qoder/dialog.js +231 -0
  22. package/dist/agents/qoder/discovery.js +181 -0
  23. package/dist/agents/qoder/instance.js +239 -0
  24. package/dist/agents/qoder/liveness.js +10 -0
  25. package/dist/agents/qoder/model.js +82 -0
  26. package/dist/agents/qoder/profile.js +16 -0
  27. package/dist/agents/qoder/questions.js +108 -0
  28. package/dist/agents/qoder/references.js +13 -0
  29. package/dist/agents/qoder/run.js +256 -0
  30. package/dist/agents/qoder/selectors.js +30 -0
  31. package/dist/agents/qoder/workspace.js +48 -0
  32. package/dist/agents/registry.js +53 -2
  33. package/dist/config/schema.js +42 -6
  34. package/dist/loop/fix-loop.js +43 -3
  35. package/dist/mcp/context.js +33 -0
  36. package/dist/mcp/formatter.js +5 -0
  37. package/dist/mcp/handlers.js +42 -2
  38. package/dist/mcp/tools.js +3 -3
  39. package/dist/tasks/task-manager.js +39 -1
  40. package/dist/version.generated.js +1 -1
  41. package/docs/qoder-cdp.en.md +68 -0
  42. package/docs/qoder-cdp.md +68 -0
  43. package/docs/qoder-evidence/custom-model-settings.png +0 -0
  44. package/docs/qoder-evidence/default-model-settings.png +0 -0
  45. package/docs/qoder-evidence/windows-smoke.json +199 -0
  46. package/docs/visual-validation.en.md +11 -1
  47. package/docs/visual-validation.md +11 -1
  48. package/package.json +10 -2
  49. package/scripts/probe-kimicode.mjs +594 -0
  50. package/scripts/probe-qoder.mjs +40 -0
  51. package/skills/tianshu-mcp/SKILL.md +19 -15
  52. package/skills/tianshu-mcp/usage-examples.md +40 -4
package/CHANGELOG.en.md CHANGED
@@ -8,11 +8,45 @@ Chinese version: [CHANGELOG.md](CHANGELOG.md)
8
8
 
9
9
  ---
10
10
 
11
- ## [Unreleased]
11
+ ## [0.5.6] - 2026-09-22
12
+
13
+ ### Added
14
+
15
+ - Add the Qoder CN GUI adapter: installation discovery, full-path workspace binding and import, default/custom model selection, and persisted reasoning settings verified through Model Management.
16
+ - Integrate dispatch, liveness, objective acceptance, and same-session rework. Automatic and manual rework write a plan before sending its filename, full path, and complete text.
17
+ - Report model source and actual settings; retain permission mode and disclose global reasoning preferences. Approvals require the user, questions use dedicated controls, and uncertain submissions are never repeated automatically.
18
+ - Windows default/custom model, new workspace and same-session repair acceptance passed on a real desktop. macOS remains research and dispatch is disabled. See the [Qoder guide](docs/qoder-cdp.en.md).
19
+ - Isolate unit test files to prevent discovery command mocks from contaminating Git baseline tests; exclude temporary probes from lint.
20
+
21
+ ### Tests
22
+
23
+ - Full suite: **826 passed / 12 skipped** (Windows 10 x64, Node 24.18.0; 78 test files passed plus 3 real-browser files skipped by design), roughly 60 cases more than v0.5.5: discovery priority (explicit → D drive → registry/shortcuts → standard directories), wrong installation paths and identity checks, CJK/space/same-name paths, workspace import and read-back failure, cross-group model ambiguity and unsupported tiers, a save that did not persist, uncertain sends and answer submissions never resent, this-turn-bound completion judging (old replies and a static screen do not qualify), same-session repair, unconfirmed cancellation/timeout stops, and macOS branches failing closed.
24
+ - Real-browser gate cases (`TIANSHU_VISUAL_BROWSER_TEST=1`) remain 12/12 on this Windows 10 machine; `npm pack` content validation plus a clean-consumer install and strict stdio check pass locally.
25
+ - Uncovered items are stated plainly: Qoder cancellation, question answering and login/quota/network waiting classification are **covered by hermetic integration tests only**, and macOS has no real GUI validation. See [HANDOFF §9.12](HANDOFF.md).
26
+
27
+ ### Planned
28
+
29
+ - The Qoder CN macOS hardware matrix (stays `research`; dispatch disabled until verified).
30
+ - Hardware verification of Qoder CN cancellation, question answering and login/quota/network waiting classification (currently covered by hermetic integration tests only).
31
+
32
+ ## [0.5.5] - 2026-09-20
33
+
34
+ ### Added
35
+
36
+ - **Added the Kimi Code GUI adapter (`agentId=kimicode`, the fourth GUI agent)**: the Kimi Code desktop app (Moonshot AI, measured 1.0.2) is a **plain Electron install** — injecting `--remote-debugging-port` and driving it over CDP is enough, with **no** MSIX COM activation.
37
+ - **Two-renderer CDP driving**: the model menu, thinking tiers and execution-mode menu are rendered by the app's `browserOverlayOpenMenu()` into a separate `Kimi Browser Overlay` renderer process (measured: after clicking `model-pill` the main window's DOM node count is unchanged and no menu node appears), while the workspace menu and the "switch model" dialog stay in the main window → the client holds both pages and excludes the `Screenshot` target.
38
+ - **Workspace binding and import**: the **normalized full path** is the sole criterion (same-name different-directory always fails closed, never guesses an entry); an unregistered workspace is imported through the native "add workspace" dialog (`#32770`; Win32 coordinate clicks plus `WM_SETTEXT`/`WM_GETTEXT`), and after binding the panel selection and the `ws-chip` text are read back.
39
+ - **Three-stage model selection**: pill read-back → direct pick in the overlay shortcut menu → "more models…" → search and exact row pick in the main window's "switch model" dialog (the only entry for unofficial models). Thinking tiers are validated against **the tier set the UI actually renders** (official `Low/High/Max`, unofficial `On/Off`); requesting a tier the UI does not render fails loudly with `model_mismatch` before sending and is never silently kept.
40
+ - **Execution mode is forced to "fully automatic"** and read back after switching; the `run_task.reasoningLevel` domain grows to `low/medium/high` plus `max/on/off`.
41
+ - **Run detection**: `button.stop` (`aria-label="中断"`) and `button.send.is-starting` are the authoritative run signals; sending uses a marker plus a bounded 60s confirmation (the session id and the landed user message are required anchors) and is **never re-sent**.
42
+ - **Six `needs_user` kinds with `continue_task` recovery**: `close_existing_instance` / `login_required` / `user_confirmation` / `agent_question` / `system_permission` / `setup_recovery`. Answering a question writes back to the recorded session (without re-sending the brief), a user confirmation only re-attaches for observation, and environment kinds re-send the full brief; when the recorded session cannot be located the run fails closed with `session_lost`.
43
+ - **Cancellation and dispatch guard**: following the Codex M14 semantics, `cancel_task` clicks `button.stop` best-effort and bounded-waits (`gui.cancelWaitMs`) for the UI to go idle, stating plainly when the stop is unconfirmed; before dispatching, a still-running managed instance is stopped best-effort, and if it never goes idle the dispatch is rejected with `instance_busy`.
44
+ - **Machine verification (Windows 10 x64 + Kimi Code 1.0.2)**: the success path, unregistered-workspace import with auto-acceptance, and the failure → rework → re-acceptance **same-session loop** all pass. **Cancellation, question answering and same-name workspace ambiguity are covered by hermetic integration tests only** (no hardware stop click, no real question card triggered); macOS is `research` and fail-closed (executable discovery and native-dialog driving are unmeasured on macOS).
12
45
 
13
46
  ### Planned
14
47
 
15
48
  - More external AI-Agent adapters (a new agent = one profile + an optional adapter file).
49
+ - The Kimi Code macOS hardware matrix, plus hardware verification of cancellation / question answering / same-name ambiguity (currently covered by hermetic integration tests only).
16
50
  - TraeWork executable discovery and native-dialog driving on macOS (currently fail-closed).
17
51
  - Optional project-level skill seeding (by default nothing is written into target repos).
18
52
  - Cancel/rework/new-project matrices for the Codex and ZCode GUI drivers on macOS (both remain `research` on darwin).
@@ -22,6 +56,10 @@ Chinese version: [CHANGELOG.md](CHANGELOG.md)
22
56
  - ZCode's **automatic import of an unregistered project** cannot complete on Windows: the native-panel script relies on `SetForegroundWindow` to bring the dialog forward before activating its address bar, but a child process of a background MCP server is refused by Windows, so the address-bar Edit never appears and the script spins until its deadline (measured: 56s, then classified by `budget.check()` as an exhausted setup budget). PowerShell's stdout is also block-buffered through a pipe, so killing the process loses the buffer and not a single `native:` stage reaches the log, misdirecting diagnosis. Workaround: add the target directory to the ZCode project list manually first; the fix direction is to grab the foreground with `AttachThreadInput` inside the script, or to use a supported ZCode registration entry point.
23
57
  - Cross-round verdict-flip circuit breaking for AI content validation (this round covers it with caching plus sampling; reconsider if hardware data still shows churn), cross-task cache sharing, and reference-image/design-diff comparison.
24
58
 
59
+ ### Fixed
60
+
61
+ - **Completed the two contract-mapping assertions issue #13's plan §5 G requires (added after the v0.5.4 release)**: `CONTENT_TIMEOUT` had its classification logic implemented but tests only asserted the transport-level `outcome.timeout`, and the stub's `sleep` mode was never exercised by any case. `visual-content-command` now has an end-to-end assertion (a timeout yields a single `blocked` item with `CONTENT_TIMEOUT`, still warning-only, with temp input files removed), plus a success-path temp-input-removal assertion and a `visual doctor` assertion for the "multi-rule total budget exceeds `roundTimeoutMs` → advisory only" branch. The v0.5.4 tag (`ea797d1`) shipped 644 cases; master now has 647.
62
+
25
63
  ---
26
64
 
27
65
  ## [0.5.4] — 2026-09-16
@@ -776,7 +814,13 @@ project → pick model and reasoning level → send instructions → run detecti
776
814
 
777
815
  ---
778
816
 
779
- [Unreleased]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.4.1...HEAD
817
+ [0.5.6]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.5...v0.5.6
818
+ [0.5.5]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.4...v0.5.5
819
+ [0.5.4]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.3...v0.5.4
820
+ [0.5.3]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.2...v0.5.3
821
+ [0.5.2]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.1...v0.5.2
822
+ [0.5.1]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.0...v0.5.1
823
+ [0.5.0]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.4.1...v0.5.0
780
824
  [0.4.1]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.4.0...v0.4.1
781
825
  [0.4.0]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.3.4...v0.4.0
782
826
  [0.3.4]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.3.3...v0.3.4
package/CHANGELOG.md CHANGED
@@ -7,11 +7,45 @@
7
7
 
8
8
  ---
9
9
 
10
- ## [未发布]
10
+ ## [0.5.6] - 2026-09-22
11
+
12
+ ### 新增
13
+
14
+ - 新增 Qoder CN GUI 适配:安装发现、完整路径工作区绑定与导入、默认/自定义模型选择、模型管理思考等级保存回读。
15
+ - 接入已有 MCP 开发、运行检测、客观验收和原会话返修流程;自动与手动返修均先保存修复计划,再发送文件名、完整路径与全文。
16
+ - 新增模型来源与实际设置报告;保留权限模式,披露全局思考偏好影响;审批不代批,提问使用专用答题控件,提交不明不重发。
17
+ - Windows 默认/自定义模型、新工作区与原会话返修真机验收通过;macOS 保持 research 并禁止派发。详见 [Qoder 操作说明](docs/qoder-cdp.md)。
18
+ - 恢复单元测试文件隔离,修复安装探测命令 mock 泄漏导致 Git 基线测试依赖执行顺序的问题;lint 排除临时探测目录。
19
+
20
+ ### 测试
21
+
22
+ - 全量 **826 passed / 12 skipped**(Windows 10 x64,Node 24.18.0;78 个测试文件通过 + 3 个真实浏览器文件按设计 skip),较 v0.5.5 净增约 60 项:安装发现优先级(显式 → D 盘 → 注册表/快捷方式 → 标准目录)、错误安装路径与身份校验、中文/空格/同名路径、工作区导入与回读失败、模型跨组重名与档位不支持、保存未生效回读、发送/答题提交不明不重发、本轮绑定完成判定(旧回复与静止不触发)、原会话返修、取消与超时停止不确定、macOS 平台分支 fail-closed。
23
+ - 真实浏览器门禁用例(`TIANSHU_VISUAL_BROWSER_TEST=1`)在 Windows 10 本机维持 12/12;`npm pack` 内容校验与干净消费者安装 + 严格 stdio 检查在本机通过。
24
+ - 未覆盖项如实标注:Qoder 的取消真停、提问续答与登录/额度/网络等待分类**仅由 hermetic 集成测试覆盖**;macOS 未做真机 GUI 验证。详见 [HANDOFF §9.12](HANDOFF.md)。
25
+
26
+ ### 计划中
27
+
28
+ - Qoder CN 的 macOS 真机验证矩阵(保持 `research`,未验证前禁止派发)。
29
+ - Qoder CN 取消真停 GUI、提问续答与登录/额度/网络等待分类的真机验证(当前仅 hermetic 集成测试覆盖)。
30
+
31
+ ## [0.5.5] - 2026-09-20
32
+
33
+ ### 新增
34
+
35
+ - **新增 Kimi Code GUI 适配(`agentId=kimicode`,第四个 GUI agent)**:Kimi Code 桌面端(Moonshot AI,实测 1.0.2)是**普通 Electron 安装**,以 `--remote-debugging-port` 注入后经 CDP 驱动,**不需要** MSIX COM 激活。
36
+ - **双渲染进程 CDP 驱动**:模型菜单 / 思考档位 / 执行模式菜单经应用内 `browserOverlayOpenMenu()` 渲染在独立的 `Kimi Browser Overlay` 渲染进程(实测点击 `model-pill` 后主窗口 DOM 节点数不变、不产生任何菜单节点),工作区菜单与「切换模型」对话框仍在主窗口 → 客户端同时持有两个页面,并排除 `Screenshot` target。
37
+ - **工作区绑定与导入**:以**归一化完整路径**为唯一判据(同名不同目录一律 fail-closed,绝不猜一个点),未登记的工作区经原生「添加工作区」对话框(`#32770`;Win32 坐标点击 + `WM_SETTEXT`/`WM_GETTEXT`)导入,绑定后回读面板选中项与 `ws-chip` 文本。
38
+ - **模型三级选择**:pill 回读 → overlay 快捷菜单直选 → 「更多模型…」→ 主窗口「切换模型」对话框搜索精确选行(非官方模型的唯一入口);思考档位按**界面实际渲染的档位集合**校验(官方 `Low/High/Max`,非官方 `On/Off`),请求界面不存在的档位在发送前以 `model_mismatch` 响亮失败,绝不静默沿用。
39
+ - **执行模式强制「完全自动」**并在切换后回读确认;`run_task.reasoningLevel` 取值域扩展为 `low/medium/high` + `max/on/off`。
40
+ - **运行检测**:`button.stop`(`aria-label="中断"`)与 `button.send.is-starting` 为权威运行信号;发送以标记 + 60s 有界确认(会话 id 与用户消息落地为必需锚点),**绝不重发**。
41
+ - **`needs_user` 六类与 `continue_task` 恢复**:`close_existing_instance` / `login_required` / `user_confirmation` / `agent_question` / `system_permission` / `setup_recovery`;提问续答写回原会话(不重发任务书)、用户确认仅重连观察、环境类补发完整任务书;定位不到原会话一律 `session_lost` fail-closed。
42
+ - **取消与重派护栏**:照 Codex M14 语义,`cancel_task` 尽力点 `button.stop` 并在 `gui.cancelWaitMs` 内有界等待界面空闲,未确认停止时终态如实明示;派发前发现未停止的运行先尽力停止,仍不空闲以 `instance_busy` 拒绝。
43
+ - **真机验证(Windows 10 x64 + Kimi Code 1.0.2)**:成功路径、未登记工作区导入 + 自动验收、失败 → 返修 → 再验收**同会话闭环**均已通过。**取消、提问续答、同名工作区歧义仅由 hermetic 集成测试覆盖**(未在真机点停、未触发真实提问卡片);macOS 为 `research` 且 fail-closed(可执行探测与原生对话框驱动未在 macOS 实测)。
11
44
 
12
45
  ### 计划中
13
46
 
14
47
  - 更多外部 AI-Agent 适配(新 agent = 一个 profile +(如需)一个 adapter 文件)。
48
+ - Kimi Code 的 macOS 真机验证矩阵,以及取消 / 提问续答 / 同名歧义的真机验证(当前仅 hermetic 集成测试覆盖)。
15
49
  - TraeWork 在 macOS 下的可执行探测与原生对话框驱动(当前 macOS 分支 fail-closed)。
16
50
  - 可选的项目级技能播种(默认不写入目标项目仓库)。
17
51
  - Codex 与 ZCode GUI 的 macOS 取消/返修/新建项目矩阵(当前两者 darwin 均保持 `research`)。
@@ -20,6 +54,10 @@
20
54
  - ZCode **未登记项目**的自动导入在 Windows 上无法完成:原生面板脚本靠 `SetForegroundWindow` 抢前台来激活地址栏,而后台 MCP server 的子进程会被 Windows 拒绝,地址栏 Edit 永不出现,脚本空转到 deadline(实测 56s 后由 `budget.check()` 归类为 setup 预算耗尽);且 PowerShell 的 stdout 在管道里被缓冲、进程被 kill 后缓冲丢失,日志里连一条 `native:` 阶段都看不到,排障方向被误导。临时对策:先在 ZCode 中手动把目标目录加入项目列表;修复方向是脚本内改用 `AttachThreadInput` 抢前台(或改走 ZCode 受支持的登记入口)。
21
55
  - AI 内容校验的跨轮判定翻转熔断(本轮以缓存 + 采样覆盖;若真机数据显示仍扰动再议)、跨任务缓存共享、参考图/设计稿差异比对。
22
56
 
57
+ ### 修复
58
+
59
+ - **补齐 issue #13 计划 §5 G 要求的两项契约映射断言(v0.5.4 发布后补)**:`CONTENT_TIMEOUT` 此前只实现了分类逻辑、测试仅断言传输层 `outcome.timeout`,判定桩的 `sleep` 模式从未被用例使用;现补 `visual-content-command` 的端到端断言(超时 → 单项 `blocked` + `CONTENT_TIMEOUT` + 仍为仅告警 + 临时输入文件已删除),并补成功路径的临时输入删除断言,以及 `visual doctor` 的「多规则总预算超 `roundTimeoutMs` 只给建议值」分支断言。v0.5.4 tag(`ea797d1`)的用例数为 644,master 现为 647。
60
+
23
61
  ---
24
62
 
25
63
  ## [0.5.4] — 2026-09-16
@@ -671,7 +709,13 @@ Codex 桌面端改为 **GUI 驱动**:新增 `codex-gui` adapter,通过 MSIX
671
709
 
672
710
  ---
673
711
 
674
- [未发布]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.4.1...HEAD
712
+ [0.5.6]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.5...v0.5.6
713
+ [0.5.5]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.4...v0.5.5
714
+ [0.5.4]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.3...v0.5.4
715
+ [0.5.3]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.2...v0.5.3
716
+ [0.5.2]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.1...v0.5.2
717
+ [0.5.1]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.5.0...v0.5.1
718
+ [0.5.0]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.4.1...v0.5.0
675
719
  [0.4.1]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.4.0...v0.4.1
676
720
  [0.4.0]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.3.4...v0.4.0
677
721
  [0.3.4]: https://github.com/lanlan0811/tianshu-mcp/compare/v0.3.3...v0.3.4