tianshu-mcp 0.7.6 → 0.7.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.en.md CHANGED
@@ -8,6 +8,29 @@ Chinese version: [CHANGELOG.md](CHANGELOG.md)
8
8
 
9
9
  ---
10
10
 
11
+ ## [0.7.7] - 2026-10-01
12
+
13
+ ### Added
14
+
15
+ - **Blocking wait primitives `wait_task` / `wait_any` (issue #28)**. `run_task` returning a `taskId` immediately is the right adaptation to the host's constraints, but the intended caller (a Tianshu agent session) is **turn-driven** — it runs only within the turn that received a user message and does nothing between turns, so it cannot poll on its own. The task-completion moment could therefore only be caught by a human sending another message. Two **pure read-only, approval-free** tools now carry the waiting:
16
+ - `wait_task(taskId, timeoutMs?)` — block until a single task reaches a **stop point** (terminal status or `needs_user`) or the timeout elapses;
17
+ - `wait_any(taskIds, timeoutMs?)` — wait for the **first task in array order** among a group (1..20) to reach a stop point, returning its snapshot plus every task's current status.
18
+ - **Stop-point definition** (single decision point `isWaitSettled`): `isTerminal(status) || status === "needs_user"`. The moment a task **stops making progress** is the moment to wake the caller — `needs_user` is not terminal but has stopped awaiting a human (it can be resumed by `continue_task` and may re-enter); without waiting for it the wait would block until the timeout and the caller would know nothing about "the task is waiting for a person".
19
+ - **Timeout policy**: `timeoutMs` defaults to `50000ms` (below the common 60 s client tool timeout, leaving round-trip headroom) and caps at `600000ms`; values above the cap are **clamped and disclosed honestly** (never silently rewritten); a timed-out response steers the caller into a call loop (≈50 s per round; long tasks need several calls).
20
+ - **Lossless guarantee**: the wait is **read-only** — it writes no task state and touches no task body, so a client truncation / connection drop / timeout **never affects the task's continued execution**. On request cancellation / connection close the loop exits immediately via the SDK's `extra.signal`, leaking no background wait.
21
+ - **Tool surface 11 → 13** (`read` family +2); new bilingual [wait primitives](docs/wait-task.en.md) doc.
22
+
23
+ ### Changed
24
+
25
+ - The `registerTool` callback in `src/server.ts` forwards the SDK request `extra` (including `signal`) to the handler — **used only by the wait tools**; every other handler is unchanged.
26
+ - `MetaBlockFields` gains optional `waitSettled` (whether a stop point was reached) and `waitedMs` (actual wait duration) for programmatic checks by the caller.
27
+
28
+ ### Tests
29
+
30
+ - Added `test/unit/wait-task.test.ts` (8 cases: stop-point decision / state transition / timeout / `signal` abort / missing reporting / clamp disclosure) and `test/integration/wait-task.test.ts` (6 cases: real stub long-task end-to-end / short-timeout continuation / not-found error / `cancel_task` effective during a wait / `wait_any` first settled / fail-closed on a missing id).
31
+ - `test/protocol/protocol.test.ts` truth table and tool-name array synced to 13 (the count hard assertion covers it automatically).
32
+ - Full suite **1447 passed / 12 skipped** (1459 tests, 122 files).
33
+
11
34
  ## [0.7.6] - 2026-09-30
12
35
 
13
36
  ### Fixed
package/CHANGELOG.md CHANGED
@@ -7,6 +7,35 @@
7
7
 
8
8
  ---
9
9
 
10
+ ## [0.7.7] - 2026-10-01
11
+
12
+ ### 新增
13
+
14
+ - **阻塞等待原语 `wait_task` / `wait_any`(issue #28)**。`run_task` 秒回 `taskId` 是对宿主约束的正确适配,
15
+ 但目标调用方(天枢 agent 会话)**回合驱动**——只在收到用户消息的回合内运行、回合之间不运行,无法自行轮询,
16
+ 于是「任务完成时刻」只能靠人工再发一条消息触发查询。新增两个**纯只读、免审批**工具承载「等」:
17
+ - `wait_task(taskId, timeoutMs?)`:阻塞等待单任务到达**停点**(终态或 `needs_user`)或超时;
18
+ - `wait_any(taskIds, timeoutMs?)`:等待一组任务(1..20)中**数组顺序首个**到达停点者,返回其快照 + 全部任务当前状态。
19
+ - **停点定义**(单一判定点 `isWaitSettled`):`isTerminal(status) || status === "needs_user"`。任务**停止推进**的时刻即应唤醒调用方——
20
+ `needs_user` 虽非终态但已停等人工(可被 `continue_task` 恢复,之后可能再次进入),不等它会空等到超时,调用方对「任务在等人」一无所知。
21
+ - **超时策略**:`timeoutMs` 缺省 `50000ms`(低于生态常见 60s 客户端超时,留序列化/往返余量)、上限 `600000ms`,
22
+ 显式超上限的值**钳制并如实披露**(不静默改值);超时返回体引导循环调用(每轮 ≈50s,长任务靠多次调用)。
23
+ - **无损保证**:等待是**纯只读**的——不写任务状态、不动任务本体;被客户端截断 / 连接中断 / 超时都**不影响任务继续执行**。
24
+ 请求取消 / 连接关闭时经 SDK 的 `extra.signal` **立即退出**循环,不泄漏后台等待。
25
+ - **工具面 11 → 13**(`read` 族 +2);新增 [等待原语](docs/wait-task.md) 双语文档。
26
+
27
+ ### 变更
28
+
29
+ - `src/server.ts` 的 `registerTool` 回调把 SDK 的请求 `extra`(含 `signal`)透传给 handler——**仅 wait 工具使用**,其余 handler 行为不变。
30
+ - `MetaBlockFields` 新增可选字段 `waitSettled`(是否到停点)与 `waitedMs`(实际等待时长),供调用方可编程判断。
31
+
32
+ ### 测试
33
+
34
+ - 新增 `test/unit/wait-task.test.ts`(8 例:停点判定 / 状态跃迁 / 超时 / `signal` 中止 / 缺失上报 / 钳制披露)
35
+ 与 `test/integration/wait-task.test.ts`(6 例:真实 stub 长任务端到端 / 短超时续等 / 不存在报错 / 等待期间 `cancel_task` 即时生效 / `wait_any` 先停者 / 缺一即报错)。
36
+ - `test/protocol/protocol.test.ts` 真值表与工具面名字数组同步为 13(数量硬断言自动覆盖)。
37
+ - 全量 **1447 passed / 12 skipped**(1459 项,122 文件)。
38
+
10
39
  ## [0.7.6] - 2026-09-30
11
40
 
12
41
  ### 修复
package/README.en.md CHANGED
@@ -68,8 +68,9 @@ Codex · TraeWork · ZCode · Kimi Code · Qoder CN · Open Design
68
68
  target project workspace ← git repo + tests + .tianshu-mcp/
69
69
  ```
70
70
 
71
- - **11 MCP tools** — `run_task / continue_task / query_task / list_tasks / get_task_report / cancel_task / verify_task / rework_task / get_profiles`, plus `prepare_visual_baseline / approve_visual_baseline` for visual acceptance.
72
- - **Async contract — long tasks never block `tools/call`** — `run_task` returns a `taskId` immediately and `query_task` polls; progress is persisted, never pushed, so the caller always sees "the last fact written to disk".
71
+ - **13 MCP tools** — `run_task / continue_task / query_task / list_tasks / get_task_report / cancel_task / verify_task / rework_task / get_profiles / wait_task / wait_any`, plus `prepare_visual_baseline / approve_visual_baseline` for visual acceptance.
72
+ - **Async contract — long tasks never block `tools/call`** — `run_task` returns a `taskId` immediately, then `wait_task` blocks until a stop point (terminal status or `needs_user`), or `query_task` polls; progress is persisted, never pushed, so the caller always sees "the last fact written to disk".
73
+ - **Wait primitives (issue #28)** — `wait_task(taskId)` / `wait_any(taskIds)` return in a single call once a task reaches a **stop point** (terminal status or `needs_user`), designed for turn-driven callers: dispatch with `run_task` and wait for the result within the same turn, no hand-rolled polling. Read-only; timeouts or interruptions never affect the task itself.
73
74
  - **Objective acceptance, fail-closed** — automated command checks plus programmatic code analysis, all relative to the **git baseline captured before work started**, and the server **never auto-commits / stashes / rolls back**. A test check that exits 0 with zero executed tests, or a git project with zero net changes, fails rather than passing green.
74
75
  - **Rework loop** — automatic rework (`autoFixRounds`) plus manual `rework_task`; failure reasons are parsed into **directly executable actions** and fed back to the agent, and exhausted rounds become `needs_attention` awaiting a Tianshu verdict.
75
76
  - **Six GUI execution surfaces (CDP)** — each agent uses an isolated CDP flow to drive its desktop UI and reports fine-grained events at key nodes, so `query_task` can tell "the agent is working" apart from "stuck on a dialog waiting for a human".
@@ -125,7 +126,7 @@ The shape of this project is not a free design: it was forced by a handful of **
125
126
 
126
127
  ## Core features
127
128
 
128
- - **Async dispatch and polling** — `run_task` returns a `taskId` immediately; `query_task` reports status / progress / log tail / recent fine-grained events (`eventLimit`, 1..50, default 10).
129
+ - **Async dispatch and waiting** — `run_task` returns a `taskId` immediately; `wait_task` blocks until a task reaches a stop point (terminal status or `needs_user`) and `wait_any` waits for the first of a group; use `query_task` for progress detail — status / progress / log tail / recent fine-grained events (`eventLimit`, 1..50, default 10). See [wait primitives](docs/wait-task.en.md).
129
130
  - **Objective acceptance engine** — automated command checks (typecheck/lint/test/build, skipped when absent, plus tech-stack derivation) plus programmatic code analysis (changed-file list / diffstat / suspicious signals such as TODO, debugger, secret-like patterns), all relative to the **git baseline**; command checks run **bounded-parallel** by default (`verifyConcurrency`, default 2, range 1–4; `1` makes them fully serial).
130
131
  - **Three fail-closed guards** — a test check fails when its output reports zero executed tests even if the exit code is 0; git projects must produce changes relative to the baseline by default (pure analysis tasks opt out with `"requireChanges": false` in `.tianshu-mcp/acceptance.json`); a round cancelled at any point yields `passed=false`.
131
132
  - **Three-level acceptance config inheritance** (issue #20) — `<data-dir>/acceptance.default.json` (global fallback) → `<project>/.tianshu-mcp/acceptance.json` (project override) → the `acceptanceOverride` argument (transient task override, never written to disk). Inspect the effective configuration with `tianshu-mcp config acceptance <projectPath> [--task <id>]`. See the [acceptance config spec](docs/acceptance-config.en.md).
@@ -183,10 +184,10 @@ In Tianshu, go to **Settings → MCP servers → Add** and fill in the fields be
183
184
  | Command | `npx` | `node` |
184
185
  | Arguments (space-separated) | `-y tianshu-mcp` | `<absolute-repo-path>/dist/index.js` |
185
186
 
186
- > - The server ID is the tool prefix: with `tianshu-mcp` the tools are `mcp__tianshu-mcp__run_task` and the other 10.
187
+ > - The server ID is the tool prefix: with `tianshu-mcp` the tools are `mcp__tianshu-mcp__run_task` and the other 12.
187
188
  > - Arguments are space-separated with **no quotes**; in local development replace `<absolute-repo-path>` with a real absolute path.
188
189
  > - The UI has no environment-variable field; to override the data directory, use the `config.json` route below to set `TIANSHU_MCP_HOME`.
189
- > - Once the connection succeeds you are done; a new session shows all 11 tools.
190
+ > - Once the connection succeeds you are done; a new session shows all 13 tools.
190
191
 
191
192
  ### Or edit config.json (environment variables supported)
192
193
 
@@ -204,23 +205,24 @@ In Tianshu, go to **Settings → MCP servers → Add** and fill in the fields be
204
205
  }
205
206
  ```
206
207
 
207
- After opening a new session the tool surface exposes `mcp__tianshu-mcp__run_task` and the other 10 tools. One typical loop:
208
+ After opening a new session the tool surface exposes `mcp__tianshu-mcp__run_task` and the other 12 tools. One typical loop:
208
209
 
209
210
  ```text
210
211
  run_task(projectPath=D:/xxx/my-app, task="…task brief…", agentId=codex,
211
212
  model="GPT-5.6 Sol", reasoningLevel="high", autoVerify=true, autoFixRounds=5)
212
- → taskId → poll query_task(taskId) → succeeded / failed / needs_attention → read get_task_report
213
+ → taskId → wait_task(taskId) blocks until a stop point → succeeded / failed / needs_attention → read get_task_report
214
+ (turn-driven callers: one wait_task call returns at the stop point; after a timeout call it again to keep waiting, or use query_task for progress detail)
213
215
  ```
214
216
 
215
217
  ### Prompts to give Tianshu (recommended usage)
216
218
 
217
- > "In project `D:\xxx`, use codex to implement 『task』. First run `run_task(autoVerify:true, autoFixRounds:2)`, then check with `query_task`; if the report shows `needs_attention`, pass the failure summary from `get_task_report` as `feedback` to `rework_task` for another round; when everything passes, report `changedFiles` and `diffstat` back to me."
219
+ > "In project `D:\xxx`, use codex to implement 『task』. First run `run_task(autoVerify:true, autoFixRounds:2)`, then `wait_task` until it reaches a stop point; if the report shows `needs_attention`, pass the failure summary from `get_task_report` as `feedback` to `rework_task` for another round; when everything passes, report `changedFiles` and `diffstat` back to me."
218
220
 
219
221
  > "In project `D:\xxx`, use traework with `mode=Code` to implement 『task』; it switches to Code mode, binds the project, sends the task, verifies automatically, and on failure generates a repair plan and reworks."
220
222
 
221
223
  ## Tool surface
222
224
 
223
- 11 tools, split into three capability families: `read` (read/query, no side effects), `write` (side effects, all requiring approval), and `execute` (runs project-side commands without changing source; currently only `verify_task`, still approval-free).
225
+ 13 tools, split into three capability families: `read` (read/query, no side effects), `write` (side effects, all requiring approval), and `execute` (runs project-side commands without changing source; currently only `verify_task`, still approval-free).
224
226
 
225
227
  | Tool | Capability / approval | Purpose |
226
228
  |---|---|---|
@@ -231,6 +233,8 @@ run_task(projectPath=D:/xxx/my-app, task="…task brief…", agentId=codex,
231
233
  | `get_task_report` | read | Full text of one round's acceptance report (`report.md`) |
232
234
  | `cancel_task` | write + approval | Cancel a running task: CLI agents kill the process tree; GUI agents best-effort click stop over CDP and wait boundedly within `gui.cancelWaitMs` (default 15s); for a terminal GUI task it doubles as the manual confirmation entry |
233
235
  | `verify_task` | execute (no source changes, approval-free) | Run acceptance once against a task or a project path. It runs configured commands and may produce build artifacts, so its MCP `readOnlyHint` is `false` — but it **changes no source and still needs no approval**; optional `idempotencyKey` |
236
+ | `wait_task` | read | Block until one task reaches a stop point (terminal status or `needs_user`) or the timeout elapses; `timeoutMs` defaults to 50000, caps at 600000 — call again after a timeout to keep waiting. Read-only, harmless |
237
+ | `wait_any` | read | Block until the first of a group (1..20) reaches a stop point, in array order; returns that task's snapshot plus the current status of every task. Validates all ids exist, failing if any is missing |
234
238
  | `rework_task` | write + approval | Manual rework (feeds the failure report back to the same agent); optional `repairHint` (≤4000 chars) |
235
239
  | `get_profiles` | read | Show agent adapters and executable discovery results |
236
240
  | `prepare_visual_baseline` | write + approval | Capture or import a reference image and produce a candidate and summary for review |
@@ -296,7 +300,7 @@ The built-in `codex` drives the desktop GUI; if you would rather not depend on G
296
300
 
297
301
  | Capability | Meaning | Approval | Tools |
298
302
  |---|---|---|---|
299
- | `read` | read/query only, no side effects | none | `query_task` / `list_tasks` / `get_task_report` / `get_profiles` |
303
+ | `read` | read/query only, no side effects | none | `query_task` / `list_tasks` / `get_task_report` / `get_profiles` / `wait_task` / `wait_any` |
300
304
  | `write` | has side effects | required | `run_task` / `continue_task` / `cancel_task` / `rework_task` / the two visual baseline tools |
301
305
  | `execute` | runs project-side commands, changes no source | none | `verify_task` |
302
306
 
@@ -424,7 +428,8 @@ Decoupling and release boundaries (read before changing anything here):
424
428
 
425
429
  | Document | Contents |
426
430
  |---|---|
427
- | [ARCHITECTURE.en.md](ARCHITECTURE.en.md) | Architecture: layering and module boundaries, state machine, acceptance pipeline, driver-layer contracts, extension points, known gaps |
431
+ | [ARCHITECTURE.en.md](<ARCHITECTURE.en.md>) | Architecture: layering and module boundaries, state machine, acceptance pipeline, driver-layer contracts, extension points, known gaps |
432
+ | [docs/core-principles.en.md](<docs/core-principles.en.md>) | Core principles: how four hard constraints forced the current architecture, the core mechanisms one by one, and why they are self-consistent |
428
433
  | [docs/tianshu-integration.en.md](docs/tianshu-integration.en.md) | The two Tianshu `config.json` integration modes, UI / API steps, smoke procedure, FAQ |
429
434
  | [docs/agent-profiles.en.md](docs/agent-profiles.en.md) | Agent profile field reference plus real-machine samples |
430
435
  | [docs/adapter-matrix.en.md](docs/adapter-matrix.en.md) | Capability research matrix for each agent |
@@ -450,7 +455,8 @@ Decoupling and release boundaries (read before changing anything here):
450
455
  | [docs/repair-directives.en.md](docs/repair-directives.en.md) | Structured repair directives: sources, fallback semantics, known limits |
451
456
  | [docs/dry-run.en.md](docs/dry-run.en.md) | dryRun mode: read-only constraint, zero-change gate, plan document |
452
457
  | [docs/event-stream.en.md](docs/event-stream.en.md) | Fine-grained event stream: vocabulary, persistence, bounded read-side window |
453
- | [docs/notifications.en.md](docs/notifications.en.md) | Task terminal-state notifications: webhook contract, de-duplication, signing |
458
+ | docs/notifications.en.md | Task terminal-state notifications: webhook contract, de-duplication, signing |
459
+ | docs/wait-task.en.md | Wait primitives: `wait_task` / `wait_any` contract, stop-point definition, timeout matrix and loop patterns |
454
460
  | [docs/visual-acceptance.en.md](docs/visual-acceptance.en.md) | Visual acceptance primer and full configuration (including optional AI content validation) |
455
461
  | [docs/visual-validation.en.md](docs/visual-validation.en.md) | Visual acceptance validation progress and platform evidence |
456
462
  | [docs/visual-validation-evidence/](docs/visual-validation-evidence/) | Raw machine-readable records behind that validation |
package/README.md CHANGED
@@ -68,8 +68,9 @@ Codex · TraeWork · ZCode · Kimi Code · Qoder CN · Open Design
68
68
  目标项目工作区 ← git 仓库 + 测试 + .tianshu-mcp/
69
69
  ```
70
70
 
71
- - **11 个 MCP 工具** —— `run_task / continue_task / query_task / list_tasks / get_task_report / cancel_task / verify_task / rework_task / get_profiles`,外加视觉验收的 `prepare_visual_baseline / approve_visual_baseline`。
72
- - **异步契约,长任务不卡 `tools/call`** —— `run_task` 秒回 `taskId`,用 `query_task` 轮询;进度只落盘、不推送,调用方看到的始终是「最后一次落盘的事实」。
71
+ - **13 个 MCP 工具** —— `run_task / continue_task / query_task / list_tasks / get_task_report / cancel_task / verify_task / rework_task / get_profiles / wait_task / wait_any`,外加视觉验收的 `prepare_visual_baseline / approve_visual_baseline`。
72
+ - **异步契约,长任务不卡 `tools/call`** —— `run_task` 秒回 `taskId`,用 `wait_task` 阻塞等到停点(终态或 `needs_user`)、或用 `query_task` 轮询;进度只落盘、不推送,调用方看到的始终是「最后一次落盘的事实」。
73
+ - **等待原语(issue #28)** —— `wait_task(taskId)` / `wait_any(taskIds)` 一次调用即等到任务到达**停点**(终态或 `needs_user`),专为回合驱动调用方设计:`run_task` 后在本回合内直接等结果,无需自行轮询;纯只读、超时/中断对任务本体零影响。
73
74
  - **客观验收,fail-closed** —— 自动命令检查 + 程序化代码分析,全部相对动工前的 **git 基线**,**绝不自动 commit / stash / 回滚**;「测试退出码 0 但零用例」「git 项目零净变更」都判失败,杜绝假绿。
74
75
  - **失败返修闭环** —— 自动返修(`autoFixRounds`)+ 手动 `rework_task`;失败原因被解析为**可直接执行的动作**随计划喂回 agent,轮次耗尽转 `needs_attention` 等天枢裁决。
75
76
  - **六个 GUI 执行面(CDP)** —— 各 agent 使用隔离的 CDP 流程驱动桌面 UI,并在关键节点上报细粒度事件,`query_task` 因此能区分「agent 正在干活」与「卡在弹窗等人工介入」。
@@ -125,7 +126,7 @@ Codex · TraeWork · ZCode · Kimi Code · Qoder CN · Open Design
125
126
 
126
127
  ## 核心特性
127
128
 
128
- - **异步派单与轮询** —— `run_task` 秒回 `taskId`;`query_task` 返回状态 / 进度 / 日志尾 / 最近细粒度事件(`eventLimit`,1..50,默认 10)。
129
+ - **异步派单与等待** —— `run_task` 秒回 `taskId`;`wait_task` 阻塞等到任务到达停点(终态或 `needs_user`),`wait_any` 等一组任务的先到者;需要进度细节时用 `query_task` 看状态 / 进度 / 日志尾 / 最近细粒度事件(`eventLimit`,1..50,默认 10)。详见 [等待原语](docs/wait-task.md)。
129
130
  - **客观验收引擎** —— 自动命令检查(typecheck/lint/test/build,缺则跳过 + 技术栈推导)+ 程序化代码分析(变更清单 / diffstat / TODO·debugger·密钥形态等可疑标记),全部相对 **git 基线**;命令默认**有界并行**(`verifyConcurrency`,默认 2,范围 1–4,`1` 即完全串行)。
130
131
  - **三项 fail-closed 保护** —— 测试退出码为 0 但零用例判失败;git 项目默认要求相对基线产生变更(纯分析任务可在 `.tianshu-mcp/acceptance.json` 设 `"requireChanges": false` 显式关闭);本轮被取消即 `passed=false`。
131
132
  - **验收配置三级继承**(issue #20)—— `<数据目录>/acceptance.default.json`(全局兜底)→ `<项目>/.tianshu-mcp/acceptance.json`(项目覆盖)→ `acceptanceOverride` 参数(任务级临时覆盖,不落盘)。用 `tianshu-mcp config acceptance <projectPath> [--task <id>]` 查看最终生效配置。详见 [验收配置规范](docs/acceptance-config.md)。
@@ -183,10 +184,10 @@ npm install -g tianshu-mcp
183
184
  | 命令 | `npx` | `node` |
184
185
  | 参数(空格分隔) | `-y tianshu-mcp` | `<仓库绝对路径>/dist/index.js` |
185
186
 
186
- > - 服务器 ID 即工具前缀:填 `tianshu-mcp` 后工具名为 `mcp__tianshu-mcp__run_task` 等 11 个。
187
+ > - 服务器 ID 即工具前缀:填 `tianshu-mcp` 后工具名为 `mcp__tianshu-mcp__run_task` 等 13 个。
187
188
  > - 参数按空格分隔填写,**不要加引号**;本地开发模式请把 `<仓库绝对路径>` 换成真实绝对路径。
188
189
  > - 界面未提供环境变量输入框;如需自定义数据目录,改用下面的 `config.json` 方式设置 `TIANSHU_MCP_HOME`。
189
- > - 添加后连接成功即完成;新开会话即可看到 11 个工具。
190
+ > - 添加后连接成功即完成;新开会话即可看到 13 个工具。
190
191
 
191
192
  ### 或改 config.json(可配环境变量)
192
193
 
@@ -204,23 +205,24 @@ npm install -g tianshu-mcp
204
205
  }
205
206
  ```
206
207
 
207
- 新开会话后,工具面出现 `mcp__tianshu-mcp__run_task` 等 11 个工具。一次典型闭环:
208
+ 新开会话后,工具面出现 `mcp__tianshu-mcp__run_task` 等 13 个工具。一次典型闭环:
208
209
 
209
210
  ```text
210
211
  run_task(projectPath=D:/xxx/my-app, task="…任务书…", agentId=codex,
211
212
  model="GPT-5.6 Sol", reasoningLevel="高", autoVerify=true, autoFixRounds=5)
212
- → taskId → query_task(taskId) 轮询 → succeeded / failed / needs_attention → get_task_report 读报告
213
+ → taskId → wait_task(taskId) 阻塞等到停点 → succeeded / failed / needs_attention → get_task_report 读报告
214
+ (回合驱动调用方:wait_task 一次调用即等到停点;超时返回后再次调用本工具继续等待,或用 query_task 看进度细节)
213
215
  ```
214
216
 
215
217
  ### 给天枢的提示语(推荐用法)
216
218
 
217
- > 「在项目 `D:\xxx` 用 codex 实现『任务』。先跑 `run_task(autoVerify:true, autoFixRounds:2)`,完成后用 `query_task` 看结果;若报告显示 `needs_attention`,把 `get_task_report` 的失败项摘要作为 `feedback` 调 `rework_task` 再验一轮;全部通过后向我汇报 `changedFiles` 与 `diffstat`。」
219
+ > 「在项目 `D:\xxx` 用 codex 实现『任务』。先跑 `run_task(autoVerify:true, autoFixRounds:2)`,完成后用 `wait_task` 等到停点再看结果;若报告显示 `needs_attention`,把 `get_task_report` 的失败项摘要作为 `feedback` 调 `rework_task` 再验一轮;全部通过后向我汇报 `changedFiles` 与 `diffstat`。」
218
220
 
219
221
  > 「在项目 `D:\xxx` 用 traework、`mode=Code` 实现『任务』;它会先切到 Code 模式再绑定项目,然后发任务、自动验收,失败自动生成修复计划并返修。」
220
222
 
221
223
  ## 工具面
222
224
 
223
- 11 个工具,按能力分为三族:`read`(读 / 查询,无副作用)、`write`(有副作用,全部需审批)、`execute`(执行项目侧命令但不改源码,当前仅 `verify_task`,仍免审批)。
225
+ 13 个工具,按能力分为三族:`read`(读 / 查询,无副作用)、`write`(有副作用,全部需审批)、`execute`(执行项目侧命令但不改源码,当前仅 `verify_task`,仍免审批)。
224
226
 
225
227
  | 工具 | 能力 / 审批 | 作用 |
226
228
  |---|---|---|
@@ -231,6 +233,8 @@ run_task(projectPath=D:/xxx/my-app, task="…任务书…", agentId=codex,
231
233
  | `get_task_report` | read | 某轮验收报告全文(`report.md`) |
232
234
  | `cancel_task` | write + 审批 | 取消运行中任务:CLI agent kill 进程树;GUI agent 经 CDP 尽力点停止并在 `gui.cancelWaitMs`(默认 15s)内有界等待;对已终态 GUI 任务兼任人工确认入口 |
233
235
  | `verify_task` | execute(不改源码,免审批) | 对任务 / 项目路径做一次验收。会跑项目配置命令、可能产生构建产物,故 MCP `readOnlyHint` 为 `false`,但**不改源码、仍免审批**;可选 `idempotencyKey` |
236
+ | `wait_task` | read | 阻塞等待单任务到达停点(终态或 `needs_user`)或超时;`timeoutMs` 缺省 50000、上限 600000,超时返回后再调一次继续等。纯只读、无害 |
237
+ | `wait_any` | read | 阻塞等待一组任务(1..20)中数组顺序首个到达停点者;返回该任务快照 + 全部任务当前状态。校验全部 id 存在,缺一即报错 |
234
238
  | `rework_task` | write + 审批 | 手动返修(把失败报告喂回同一 agent);可选 `repairHint`(≤4000 字符) |
235
239
  | `get_profiles` | read | 查看 agent 适配与可执行探测结果 |
236
240
  | `prepare_visual_baseline` | write + 审批 | 截图或导入参考图,生成待审阅候选和摘要 |
@@ -296,7 +300,7 @@ run_task(projectPath=D:/xxx/my-app, task="…任务书…", agentId=codex,
296
300
 
297
301
  | 能力 | 含义 | 审批 | 工具 |
298
302
  |---|---|---|---|
299
- | `read` | 只读 / 查询,无副作用 | 免审批 | `query_task` / `list_tasks` / `get_task_report` / `get_profiles` |
303
+ | `read` | 只读 / 查询,无副作用 | 免审批 | `query_task` / `list_tasks` / `get_task_report` / `get_profiles` / `wait_task` / `wait_any` |
300
304
  | `write` | 有副作用 | 需审批 | `run_task` / `continue_task` / `cancel_task` / `rework_task` / 两个视觉基准工具 |
301
305
  | `execute` | 执行项目侧命令,不改源码 | 免审批 | `verify_task` |
302
306
 
@@ -424,7 +428,8 @@ run_task(projectPath=D:/xxx/my-app, task="…任务书…", agentId=codex,
424
428
 
425
429
  | 文档 | 说明 |
426
430
  |---|---|
427
- | [ARCHITECTURE.md](ARCHITECTURE.md) | 架构说明:分层模型与模块边界、状态机、验收流水线、驱动层契约、扩展点与已知缺口 |
431
+ | [ARCHITECTURE.md](<ARCHITECTURE.md>) | 架构说明:分层模型与模块边界、状态机、验收流水线、驱动层契约、扩展点与已知缺口 |
432
+ | [docs/core-principles.md](<docs/core-principles.md>) | 核心原理分析:四条硬约束如何逼出当前架构、核心机制逐条拆解与自洽性总结 |
428
433
  | [docs/tianshu-integration.md](docs/tianshu-integration.md) | 天枢 config.json 两种接入模式、UI / API 操作、冒烟步骤、FAQ |
429
434
  | [docs/agent-profiles.md](docs/agent-profiles.md) | agent profile 字段说明 + 真实机器样例 |
430
435
  | [docs/adapter-matrix.md](docs/adapter-matrix.md) | 各 Agent 能力调研矩阵 |
@@ -450,7 +455,8 @@ run_task(projectPath=D:/xxx/my-app, task="…任务书…", agentId=codex,
450
455
  | [docs/repair-directives.md](docs/repair-directives.md) | 结构化修复指令:来源、回退语义与已知限制 |
451
456
  | [docs/dry-run.md](docs/dry-run.md) | dryRun 干跑模式:只读约束、零改动门禁、方案文档 |
452
457
  | [docs/event-stream.md](docs/event-stream.md) | 细粒度事件流:词表、落盘与读取侧有界窗口 |
453
- | [docs/notifications.md](docs/notifications.md) | 任务终态通知:webhook 契约、去重与签名 |
458
+ | docs/notifications.md | 任务终态通知:webhook 契约、去重与签名 |
459
+ | docs/wait-task.md | 等待原语:`wait_task` / `wait_any` 契约、停点定义、超时矩阵与循环模式 |
454
460
  | [docs/visual-acceptance.md](docs/visual-acceptance.md) | 视觉验收入门与完整配置(含可选 AI 内容校验) |
455
461
  | [docs/visual-validation.md](docs/visual-validation.md) | 视觉验收验证进度与平台证据 |
456
462
  | [docs/visual-validation-evidence/](docs/visual-validation-evidence/) | 上述验证的原始机器可读记录 |
@@ -220,6 +220,46 @@ export const ContinueTaskParamsSchema = z.object({
220
220
  taskId: z.string().min(1),
221
221
  message: z.string().min(1, "message 不能为空"),
222
222
  });
223
+ /* ---------------- 等待原语(issue #28) ---------------- */
224
+ /**
225
+ * `wait_task` 单次等待的**默认**上限(ms)。
226
+ * 刻意低于生态常见的 60s 客户端单次工具超时,留出序列化 / 网络往返余量:
227
+ * 若客户端超时比 50s 更短,截断也只让调用方多调一次(等待无损),不会出错。
228
+ * 真机校准(计划 Wave 5)后如需按客户端调整,只改这一处常量。
229
+ */
230
+ export const WAIT_TASK_TIMEOUT_DEFAULT_MS = 50_000;
231
+ /**
232
+ * 单次 wait 调用可请求的等待**上限**(ms,10 分钟):给「无超时或已知长超时」的调用方。
233
+ * 显式传入超过本值的值会被钳制到本值并**如实披露**(不静默改值);更长场景靠循环调用。
234
+ */
235
+ export const WAIT_TASK_TIMEOUT_MAX_MS = 600_000;
236
+ /** `wait_any` 一次可等待的任务数上限。 */
237
+ export const WAIT_ANY_TASK_IDS_MAX = 20;
238
+ export const WaitTaskParamsSchema = z.object({
239
+ taskId: z.string().min(1),
240
+ /** 本次等待上限(ms);缺省 {@link WAIT_TASK_TIMEOUT_DEFAULT_MS},超 {@link WAIT_TASK_TIMEOUT_MAX_MS} 被钳制。 */
241
+ timeoutMs: z.number().int().positive().optional(),
242
+ });
243
+ export const WaitAnyParamsSchema = z.object({
244
+ /** 一组任务 id(1..{@link WAIT_ANY_TASK_IDS_MAX});开始前校验全部存在,缺一即报错。 */
245
+ taskIds: z.array(z.string().min(1)).min(1).max(WAIT_ANY_TASK_IDS_MAX),
246
+ /** 本次等待上限(ms);语义同 {@link WaitTaskParamsSchema} 的 timeoutMs。 */
247
+ timeoutMs: z.number().int().positive().optional(),
248
+ });
249
+ /**
250
+ * 钳制等待上限:缺省用默认值;显式值超过上限时钳到上限并标记 `clamped`,由 handler
251
+ * 在响应正文里**如实披露**(计划 §2.3 缓解 1:不静默改值)。
252
+ * `timeoutMs` 的正整数约束由 schema 承担,此处只处理缺省与上限。
253
+ */
254
+ export function clampWaitTimeout(timeoutMs) {
255
+ if (timeoutMs === undefined) {
256
+ return { timeoutMs: WAIT_TASK_TIMEOUT_DEFAULT_MS, clamped: false };
257
+ }
258
+ if (timeoutMs > WAIT_TASK_TIMEOUT_MAX_MS) {
259
+ return { timeoutMs: WAIT_TASK_TIMEOUT_MAX_MS, clamped: true };
260
+ }
261
+ return { timeoutMs, clamped: false };
262
+ }
223
263
  /* ---------------- server 配置 config.json ---------------- */
224
264
  /**
225
265
  * webhook 通知可订阅的事件类别(按任务**状态语义**归类,而非原始 status 字符串)。
@@ -1,6 +1,7 @@
1
1
  /**
2
- * 11 个工具的具体 handler。统一返回 ToolResult(文本 + meta 块)。
3
- * run_task / rework / verify 依赖 AppContext 提供的 manager/engine/services。
2
+ * 13 个工具的具体 handler。统一返回 ToolResult(文本 + meta 块)。
3
+ * run_task / rework / verify 依赖 AppContext 提供的 manager/engine/services;
4
+ * wait_task / wait_any(issue #28)额外接收 SDK 的请求 `extra`(用其 `signal` 感知中断)。
4
5
  */
5
6
  import fsp from "node:fs/promises";
6
7
  import { validateQoderReferences } from "../agents/qoder/references.js";
@@ -9,7 +10,7 @@ import { normalizeOpenDesignDirection } from "../agents/opendesign/model.js";
9
10
  import { prepareBaseline, approveBaseline, PrepareBaselineSchema, ApproveBaselineSchema, } from "../visual/baselines.js";
10
11
  import { assertSafeProjectDir, normPath, resolveProjectDir } from "../util/path.js";
11
12
  import { execFileAsync } from "../verify/exec.js";
12
- import { QUERY_TASK_EVENT_LIMIT_DEFAULT } from "../config/schema.js";
13
+ import { QUERY_TASK_EVENT_LIMIT_DEFAULT, WAIT_TASK_TIMEOUT_MAX_MS, clampWaitTimeout, } from "../config/schema.js";
13
14
  import { toAcceptanceDef } from "../config/store.js";
14
15
  import { isDefaultWorkspace, isTerminal, } from "../tasks/task.js";
15
16
  import { canonicalDigest, IdempotencyIndex, keyDigest, } from "../tasks/idempotency.js";
@@ -175,6 +176,8 @@ export function makeHandlers(ctx, defaults) {
175
176
  get_task_report: getReportHandler(ctx),
176
177
  cancel_task: cancelTaskHandler(ctx),
177
178
  verify_task: verifyTaskHandler(ctx, idempotency),
179
+ wait_task: waitTaskHandler(ctx),
180
+ wait_any: waitAnyHandler(ctx),
178
181
  rework_task: reworkTaskHandler(ctx),
179
182
  continue_task: continueTaskHandler(ctx),
180
183
  get_profiles: getProfilesHandler(ctx),
@@ -566,6 +569,91 @@ function queryTaskHandler(ctx) {
566
569
  return formatToolResult(lines.join("\n"), metaFromTask(meta, { recentEvents }));
567
570
  };
568
571
  }
572
+ /* ---------------- 等待原语(issue #28) ---------------- */
573
+ /** 停点后的后续动作指引:终态 → 取报告;needs_user → continue 后再次 wait。 */
574
+ function waitNextStep(status) {
575
+ if (status === "needs_user") {
576
+ return "任务在等待人工处理:请用 continue_task 恢复,恢复后再次调用 wait_task 继续等待。";
577
+ }
578
+ if (status === "succeeded")
579
+ return "可用 get_task_report 查看验收报告。";
580
+ return "可用 get_task_report / query_task 查看详情。";
581
+ }
582
+ /** 钳制披露(仅在显式 timeoutMs 超上限时非空):如实说明已钳制,不静默改值。 */
583
+ function waitClampNote(clamped) {
584
+ return clamped ? `(timeoutMs 超上限,已钳制到 ${WAIT_TASK_TIMEOUT_MAX_MS}ms)` : "";
585
+ }
586
+ function waitTaskHandler(ctx) {
587
+ const { manager } = ctx;
588
+ return async (rawArgs, extra) => {
589
+ const args = rawArgs;
590
+ const meta = await manager.getMeta(args.taskId);
591
+ if (!meta)
592
+ return errorResult(`任务不存在: ${args.taskId}`);
593
+ const { timeoutMs, clamped } = clampWaitTimeout(args.timeoutMs);
594
+ const result = await manager.waitForStops([args.taskId], timeoutMs, extra?.signal);
595
+ const waitedSec = Math.round(result.waitedMs / 1000);
596
+ const clampNote = waitClampNote(clamped);
597
+ if (result.aborted) {
598
+ // 连接已断时本响应自然丢弃;循环已释放(SDK _onclose abort 全部 in-flight handler)。
599
+ return formatToolResult(`等待被取消(调用方中断 / 连接关闭),任务不受影响。${describeStatus(meta)}`, metaFromTask(meta, { waitSettled: false, waitedMs: result.waitedMs }));
600
+ }
601
+ if (result.stopped.length > 0) {
602
+ const hit = result.stopped[0];
603
+ return formatToolResult([
604
+ `任务已到停点(等待 ${waitedSec} 秒):${describeStatus(hit.meta)}`,
605
+ waitNextStep(hit.meta.status),
606
+ ].join("\n"), metaFromTask(hit.meta, { waitSettled: true, waitedMs: result.waitedMs }));
607
+ }
608
+ // 超时:重读一次最新快照,避免回显等待开始前的旧 meta
609
+ const current = (await manager.getMeta(args.taskId)) ?? meta;
610
+ return formatToolResult([
611
+ `等待超时(${waitedSec} 秒):${describeStatus(current)}${clampNote}`,
612
+ "任务本体不受影响;请再次调用 wait_task 继续等待,或用 query_task 查看细节。",
613
+ ].join("\n"), metaFromTask(current, { waitSettled: false, waitedMs: result.waitedMs }));
614
+ };
615
+ }
616
+ function waitAnyHandler(ctx) {
617
+ const { manager } = ctx;
618
+ return async (rawArgs, extra) => {
619
+ const args = rawArgs;
620
+ // 入口预检全部 id(fail-closed):缺一即报错并列出缺失 id,绝不静默跳过。
621
+ const pre = await Promise.all(args.taskIds.map((id) => manager.getMeta(id)));
622
+ const missing = args.taskIds.filter((_, i) => pre[i] === null);
623
+ if (missing.length > 0)
624
+ return errorResult(`任务不存在: ${missing.join(", ")}`);
625
+ const { timeoutMs, clamped } = clampWaitTimeout(args.timeoutMs);
626
+ const result = await manager.waitForStops(args.taskIds, timeoutMs, extra?.signal);
627
+ const waitedSec = Math.round(result.waitedMs / 1000);
628
+ const clampNote = waitClampNote(clamped);
629
+ const current = await Promise.all(args.taskIds.map((id) => manager.getMeta(id)));
630
+ const statusLines = args.taskIds
631
+ .map((id, i) => {
632
+ const m = current[i];
633
+ return `- ${id}: ${m ? describeStatus(m) : "任务不存在"}`;
634
+ })
635
+ .join("\n");
636
+ if (result.aborted) {
637
+ return formatToolResult(["等待被取消(调用方中断 / 连接关闭),任务不受影响。", statusLines].join("\n"), { ok: false, message: "等待被取消", waitSettled: false, waitedMs: result.waitedMs });
638
+ }
639
+ if (result.stopped.length > 0) {
640
+ // stopped 按 taskIds 数组下标升序 → stopped[0] 即「数组顺序首个已停」(确定性优先)。
641
+ const hit = result.stopped[0];
642
+ return formatToolResult([
643
+ `已有任务到达停点(等待 ${waitedSec} 秒):${hit.meta.taskId} —— ${describeStatus(hit.meta)}`,
644
+ waitNextStep(hit.meta.status),
645
+ "全部任务当前状态:",
646
+ statusLines,
647
+ ].join("\n"), metaFromTask(hit.meta, { waitSettled: true, waitedMs: result.waitedMs }));
648
+ }
649
+ return formatToolResult([
650
+ `等待超时(${waitedSec} 秒):暂无任务到达停点。${clampNote}`,
651
+ "任务本体不受影响;请再次调用 wait_any 继续等待,或用 query_task 查看细节。",
652
+ "全部任务当前状态:",
653
+ statusLines,
654
+ ].join("\n"), { ok: false, message: "等待超时", waitSettled: false, waitedMs: result.waitedMs });
655
+ };
656
+ }
569
657
  async function existsFile(p) {
570
658
  try {
571
659
  await fsp.access(p);
package/dist/mcp/tools.js CHANGED
@@ -1,5 +1,5 @@
1
1
  /**
2
- * 工具注册表:11 个工具的 name/description/inputSchema/capability/approval 元数据。
2
+ * 工具注册表:13 个工具的 name/description/inputSchema/capability/approval 元数据。
3
3
  * MCP 层用 inputSchema 声明;capability/requireApproval 供天枢 policy(§5/§11.3)。
4
4
  * 能力标注遵守 R11(三族语义):
5
5
  * - read:读/查询,无副作用;`server.ts` 据此推导 MCP `readOnlyHint: true`。
@@ -9,7 +9,7 @@
9
9
  */
10
10
  import { z } from "zod";
11
11
  import { PrepareBaselineSchema, ApproveBaselineSchema } from "../visual/baselines.js";
12
- import { RunTaskParamsSchema, QueryTaskParamsSchema, ListTasksParamsSchema, GetReportParamsSchema, CancelTaskParamsSchema, VerifyTaskParamsSchema, ReworkTaskParamsSchema, ContinueTaskParamsSchema, } from "../config/schema.js";
12
+ import { RunTaskParamsSchema, QueryTaskParamsSchema, ListTasksParamsSchema, GetReportParamsSchema, CancelTaskParamsSchema, VerifyTaskParamsSchema, ReworkTaskParamsSchema, ContinueTaskParamsSchema, WaitTaskParamsSchema, WaitAnyParamsSchema, } from "../config/schema.js";
13
13
  export const TOOL_DEFS = [
14
14
  {
15
15
  name: "prepare_visual_baseline",
@@ -79,6 +79,20 @@ export const TOOL_DEFS = [
79
79
  capability: "execute",
80
80
  requireApproval: false,
81
81
  },
82
+ {
83
+ name: "wait_task",
84
+ description: "等待任务到达停点(阻塞只读原语):轮询至终态(succeeded/failed/needs_attention/cancelled/interrupted)或 needs_user,或超时(timeoutMs 缺省 50000ms、上限 600000ms)后返回当前状态快照。适合回合驱动的调用方:run_task 后在本回合内等待结果。超时返回时请再次调用本工具继续等待——本调用不影响任务本体,超时/中断均无害。",
85
+ inputSchema: WaitTaskParamsSchema,
86
+ capability: "read",
87
+ requireApproval: false,
88
+ },
89
+ {
90
+ name: "wait_any",
91
+ description: "等待一组任务中首个到达停点(终态或 needs_user)的任务;返回该任务快照与全部任务当前状态。taskIds 1..20 个,开始前校验全部存在,缺一即报错。",
92
+ inputSchema: WaitAnyParamsSchema,
93
+ capability: "read",
94
+ requireApproval: false,
95
+ },
82
96
  {
83
97
  name: "rework_task",
84
98
  description: "手动返修:把终态任务(failed/needs_attention)重新入队续跑,同一 agent/项目与轮次记账。" +
package/dist/server.js CHANGED
@@ -1,6 +1,6 @@
1
1
  /**
2
2
  * server.ts:组装 —— 加载配置、初始化数据目录/日志、TaskManager/AcceptanceEngine/
3
- * Registry、注册 11 个工具到 McpServer、触发技能自检安装。被 index.ts 调用以 stdio 启动。
3
+ * Registry、注册 13 个工具到 McpServer、触发技能自检安装。被 index.ts 调用以 stdio 启动。
4
4
  */
5
5
  import path from "node:path";
6
6
  import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
@@ -53,7 +53,7 @@ export async function buildServer(opts = {}) {
53
53
  });
54
54
  const server = new McpServer({ name: "tianshu-mcp", version: MCP_SERVER_VERSION }, {
55
55
  capabilities: { tools: {} },
56
- instructions: "tianshu-mcp:调度外部 AI-Agent(codex/zcode/traework/kimicode/qoder)完成项目开发、验收、返修闭环。ZCode 提问或等待用户环境处理时进入 needs_user,可用 continue_task 恢复原会话。run_task 异步返回 taskId,再用 query_task 轮询。",
56
+ instructions: "tianshu-mcp:调度外部 AI-Agent(codex/zcode/traework/kimicode/qoder)完成项目开发、验收、返修闭环。ZCode 提问或等待用户环境处理时进入 needs_user,可用 continue_task 恢复原会话。run_task 异步返回 taskId;随后用 wait_task 阻塞等待任务到达停点(终态或 needs_user),超时则再次调用本工具继续等待;需要看进度细节时用 query_task 轮询。",
57
57
  });
58
58
  for (const tool of TOOL_DEFS) {
59
59
  const handler = handlers[tool.name];
@@ -80,7 +80,10 @@ export async function buildServer(opts = {}) {
80
80
  idempotentHint: tool.name === "run_task" || tool.name === "verify_task",
81
81
  title: tool.name,
82
82
  },
83
- }, async (args) => {
83
+ },
84
+ // issue #28:把 SDK 的请求 extra(含 signal)透传给 handler,供 wait_task/wait_any
85
+ // 感知请求取消 / 连接关闭;其余 handler 不读 extra,行为不变。
86
+ async (args, extra) => {
84
87
  try {
85
88
  const parsed = tool.inputSchema.safeParse(args ?? {});
86
89
  if (!parsed.success) {
@@ -92,7 +95,7 @@ export async function buildServer(opts = {}) {
92
95
  isError: true,
93
96
  };
94
97
  }
95
- return await handler(parsed.data);
98
+ return await handler(parsed.data, extra);
96
99
  }
97
100
  catch (e) {
98
101
  const msg = e instanceof Error ? e.message : String(e);
@@ -1,5 +1,6 @@
1
1
  import { ACTIVE_STATUSES, isTerminal, isProjectWorkspace, guiAppNameOf, guiStopDisclosure, } from "./task.js";
2
2
  import { TaskOrchestrator } from "../loop/fix-loop.js";
3
+ import { waitForStops as waitForStopsCore } from "./wait.js";
3
4
  import { genTaskId, nowIso } from "../util/id.js";
4
5
  import { normPath } from "../util/path.js";
5
6
  /**
@@ -120,6 +121,19 @@ export class TaskManager {
120
121
  await this.store.waitForStatusWrite(taskId);
121
122
  return this.tasks.get(taskId) ?? (await this.store.readSnapshot(taskId));
122
123
  }
124
+ /**
125
+ * 阻塞等待一组任务到达停点(issue #28 / 计划 §2.4 C3)。
126
+ *
127
+ * **纯只读**:等待期间不写任务状态、不动任务本体;被客户端截断 / 连接中断 / 超时
128
+ * 都不影响任务继续执行。状态读取复用 `getMeta`(`waitForStatusWrite` 屏障 +
129
+ * 内存优先 + 快照兜底),因此 wait 看到的是与 `query_task` 同一口径的事实。
130
+ * @param taskIds 目标任务 id(`wait_task` 单个 / `wait_any` 一组)
131
+ * @param timeoutMs 本次等待上限(调用方已钳制)
132
+ * @param signal SDK 请求的取消信号(连接关闭 / 请求取消)
133
+ */
134
+ async waitForStops(taskIds, timeoutMs, signal) {
135
+ return waitForStopsCore(taskIds, (id) => this.getMeta(id), { timeoutMs, signal });
136
+ }
123
137
  /** S4:外部对终态任务元数据的更新(如手动 verify 更新报告指针/轮次)——写快照并同步内存 map */
124
138
  async persistMetaUpdate(meta) {
125
139
  this.tasks.set(meta.taskId, meta);
@@ -48,6 +48,19 @@ export const TRANSITIONS = {
48
48
  export function isTerminal(s) {
49
49
  return TERMINAL_STATUSES.includes(s);
50
50
  }
51
+ /**
52
+ * `wait_task` / `wait_any` 的「停点」判定(issue #28 的单一判定点):
53
+ * 任务**停止推进**的时刻 = 调用方应当被唤醒的时刻。
54
+ *
55
+ * 与 cancel 路径的 `settled` 语义刻意分开:`needs_user` 是**非终态**
56
+ * (可被 `continue_task` 恢复到 `queued`,之后可能再次进入),但此刻任务已停止推进、
57
+ * 在等人工处理,必须立即唤醒调用方——否则 wait 会一直空等到 timeout,
58
+ * 调用方对「任务在等人」一无所知。
59
+ * 状态机演进(新增停点)只改这一处。
60
+ */
61
+ export function isWaitSettled(s) {
62
+ return isTerminal(s) || s === "needs_user";
63
+ }
51
64
  /** 归一化工作区模式:缺字段一律按 project(保守,绝不把旧记录或损坏记录当作无项目)。 */
52
65
  export function workspaceModeOf(meta) {
53
66
  return meta.workspaceMode === "default" ? "default" : "project";
@@ -0,0 +1,67 @@
1
+ /**
2
+ * 阻塞等待原语核心(issue #28 / 开发计划 §2.4 C1)。
3
+ *
4
+ * `waitForStops` 轮询一组任务,直到其中任一到达**停点**(终态 ∪ `needs_user`)或超时。
5
+ * 设计要点:
6
+ * - **纯逻辑、依赖注入 `getMeta`**:不依赖文件系统 / TaskManager 构造,可独立单测;
7
+ * - **只读**:不写任何任务状态、不动任务本体——被客户端截断 / 连接中断 / 超时都无副作用;
8
+ * - **`signal` 感知**:连接关闭或请求取消时用 SDK 给的 `extra.signal` 立即退出循环,
9
+ * 不泄漏后台等待(SDK `_onclose` 会 abort 全部 in-flight handler)。
10
+ */
11
+ import { isWaitSettled } from "./task.js";
12
+ /** 可被 `signal` 提前唤醒的 sleep:abort 时立即 resolve(不做悬挂等待)。 */
13
+ function sleepWithSignal(ms, signal) {
14
+ return new Promise((resolve) => {
15
+ if (signal?.aborted)
16
+ return resolve();
17
+ const timer = setTimeout(() => {
18
+ signal?.removeEventListener("abort", onAbort);
19
+ resolve();
20
+ }, ms);
21
+ const onAbort = () => {
22
+ clearTimeout(timer);
23
+ resolve();
24
+ };
25
+ signal?.addEventListener("abort", onAbort, { once: true });
26
+ });
27
+ }
28
+ /**
29
+ * 轮询 `taskIds` 直到有任务到达停点、或超时、或被 signal 中止。
30
+ *
31
+ * 单次轮询内**先读全部任务**再判定:`missing` 优先于 `stopped`(缺一即视为异常,交调用方处理);
32
+ * 否则只要有任一任务 `isWaitSettled` 即返回;再判超时;否则 sleep 到下一轮。
33
+ */
34
+ export async function waitForStops(taskIds, getMeta, opts) {
35
+ const pollIntervalMs = opts.pollIntervalMs ?? 500;
36
+ const { signal } = opts;
37
+ const start = Date.now();
38
+ const deadline = start + opts.timeoutMs;
39
+ const elapsed = () => Date.now() - start;
40
+ for (;;) {
41
+ if (signal?.aborted) {
42
+ return { stopped: [], timedOut: false, waitedMs: elapsed(), missing: [], aborted: true };
43
+ }
44
+ const metas = await Promise.all(taskIds.map((id) => getMeta(id)));
45
+ const missing = [];
46
+ const stopped = [];
47
+ metas.forEach((meta, index) => {
48
+ if (!meta) {
49
+ missing.push(taskIds[index]);
50
+ return;
51
+ }
52
+ if (isWaitSettled(meta.status))
53
+ stopped.push({ index, meta });
54
+ });
55
+ if (missing.length > 0) {
56
+ return { stopped: [], timedOut: false, waitedMs: elapsed(), missing, aborted: false };
57
+ }
58
+ if (stopped.length > 0) {
59
+ return { stopped, timedOut: false, waitedMs: elapsed(), missing: [], aborted: false };
60
+ }
61
+ const now = Date.now();
62
+ if (now >= deadline) {
63
+ return { stopped: [], timedOut: true, waitedMs: elapsed(), missing: [], aborted: false };
64
+ }
65
+ await sleepWithSignal(Math.min(pollIntervalMs, deadline - now), signal);
66
+ }
67
+ }
@@ -4,4 +4,4 @@
4
4
  * 本文件由 scripts/sync-version.mjs 在每次 build 前重新生成。
5
5
  */
6
6
  // generated: 勿手改 —— 运行 `npm run build` 自动同步
7
- export const MCP_SERVER_VERSION = "0.7.6";
7
+ export const MCP_SERVER_VERSION = "0.7.7";
@@ -0,0 +1,142 @@
1
+ # Blocking wait primitives: `wait_task` / `wait_any` (issue #28)
2
+
3
+ Chinese version: wait-task.md
4
+
5
+ `run_task` is an asynchronous contract: it returns a `taskId` immediately and never blocks `tools/call`. But the intended caller — a Tianshu desktop agent session — is **turn-driven**: the agent only runs within the turn that received a user message and does nothing between turns, so it **cannot poll on its own**. In the "turn-driven caller + long task" combination the tool surface therefore missed the completion moment: previously every task completion **required a human to send a message** to trigger a check.
6
+
7
+ This capability fills that gap: `wait_task` / `wait_any` carry the waiting with **one blocking, read-only call**, returning once a task reaches a stop point — no user intervention needed.
8
+
9
+ ## 1. Why it must be a blocking wait
10
+
11
+ Server-side push (`notifications/progress` etc.) cannot be relied on: the host MCP tools **return text only** (`content[].text`), call `tools/call` **synchronously per call**, and do not consume server-side push (see README "Runtime contract" C1/C2). So "waiting" can only be carried by **a single tool call** — a blocking wait is the only viable shape on the tool surface.
12
+
13
+ It **complements** the webhook notification ([issue #22](notifications.en.md)) without overlapping: the webhook receiver is an external HTTP endpoint (a human / a bot) and its result never flows back into the MCP session; `wait_task` delivers the result **back to the session that made the call**.
14
+
15
+ ## 2. Stop-point definition
16
+
17
+ The "returnable point" that `wait_task` / `wait_any` waits for is a **stop point**:
18
+
19
+ ```text
20
+ isWaitSettled(status) = isTerminal(status) || status === "needs_user"
21
+ ```
22
+
23
+ - `isTerminal`: `succeeded` / `failed` / `needs_attention` / `cancelled` / `interrupted`.
24
+ - `needs_user`: **not terminal** — the task has stopped making progress and awaits a human (agent question / login / environment handling); resume it with `continue_task`.
25
+
26
+ **Why `needs_user` is also a stop point**: the moment a task truly stops making progress is the moment to wake the caller. Without waiting for it, once a task enters `needs_user` the wait would block until the timeout, and the caller would **know nothing about "the task is waiting for a person"** — yet that is exactly the status that must be relayed immediately.
27
+
28
+ ## 3. Tool contracts
29
+
30
+ ### `wait_task(taskId, timeoutMs?)`
31
+
32
+ Block until **a single** task reaches a stop point or the timeout elapses.
33
+
34
+ | Argument | Required | Notes |
35
+ |---|---|---|
36
+ | `taskId` | yes | The target task id |
37
+ | `timeoutMs` | no | Wait cap (ms); default `50000`, cap `600000`; values above the cap are **clamped and disclosed in the body** |
38
+
39
+ Returns: text (stop-point line / status line / next-step guidance) + a meta block, whose meta carries `waitSettled` (whether a stop point was reached) and `waitedMs` (actual wait duration).
40
+
41
+ ### `wait_any(taskIds, timeoutMs?)`
42
+
43
+ Block until the **first task in array order** among **a group** reaches a stop point.
44
+
45
+ | Argument | Required | Notes |
46
+ |---|---|---|
47
+ | `taskIds` | yes | 1..20 task ids; all are validated up front and a **single missing id fails closed, listing the missing ids** |
48
+ | `timeoutMs` | no | Same as `wait_task` |
49
+
50
+ Returns: the snapshot of the task that reached a stop point + a meta block, and the body also lists **every task's current status**.
51
+
52
+ > It returns "the first settled in array order", not "the earliest finished": deterministic and predictable, avoiding the sorting ambiguity when `finishedAt` is missing or identical.
53
+
54
+ ## 4. Timeout matrix
55
+
56
+ | Call | `timeoutMs` | Behavior |
57
+ |---|---|---|
58
+ | `wait_task` / `wait_any` | omitted | Uses the default `50000ms` (below the common 60 s client timeout, leaving serialization / round-trip headroom) |
59
+ | same | ≤ 600000 | Waits the given value |
60
+ | same | > 600000 | **Clamped to 600000ms** and the response body **honestly states** "clamped to the cap" |
61
+ | same | expires without a stop point | Returns the **current snapshot** + "please call this tool again to keep waiting" guidance (`waitSettled=false`) |
62
+
63
+ `timeoutMs` must be a positive integer (the schema rejects 0 / negative / non-integer).
64
+
65
+ ## 5. Loop patterns
66
+
67
+ Long tasks (30–50 min) are covered by **repeated calls**: ≈50 s per round until a stop point or the user interrupts.
68
+
69
+ ### 5.1 Single task (success → read report)
70
+
71
+ ```text
72
+ run_task(...) → taskId
73
+ wait_task(taskId, timeoutMs=50000)
74
+ → reached a stop point (waited 37 s): status: [PASS] succeeded
75
+ → get_task_report(taskId)
76
+ ```
77
+
78
+ ### 5.2 Timeout continuation
79
+
80
+ ```text
81
+ wait_task(taskId, timeoutMs=50000)
82
+ → timed out (50 s): status: running. The task itself is unaffected;
83
+ call wait_task again to keep waiting, or use query_task for detail.
84
+ wait_task(taskId) # calling again continues the wait (lossless)
85
+ → … until a stop point
86
+ ```
87
+
88
+ ### 5.3 needs_user loop
89
+
90
+ ```text
91
+ wait_task(taskId)
92
+ → reached a stop point (waited 12 s): status: awaiting user (resumable via continue_task).
93
+ Call continue_task to resume, then wait_task again to keep waiting.
94
+ # The user handles it in the client (answers / logs in / closes the old instance…), then:
95
+ continue_task(taskId, message="done")
96
+ wait_task(taskId) # keep waiting after resuming (needs_user can recur)
97
+ ```
98
+
99
+ ### 5.4 First of several tasks
100
+
101
+ ```text
102
+ wait_any(taskIds=[tsk_a, tsk_b, tsk_c], timeoutMs=50000)
103
+ → a task reached a stop point (waited 8 s): tsk_b — status: [FAIL] failed
104
+ current status of every task: …
105
+ ```
106
+
107
+ ## 6. Lossless guarantee
108
+
109
+ `wait_task` / `wait_any` are **pure read-only** operations (`capability: "read"`, approval-free, MCP `readOnlyHint: true`): they write no task state and touch no task body. A client truncation, a dropped connection, or a timeout — **no path affects the task's continued execution**; the worst case is that the caller calls a few more times, and `query_task` yields the latest fact after reconnecting.
110
+
111
+ Other tool calls proceed normally during the wait (SDK request handling does not block, measured):
112
+
113
+ ```text
114
+ [probe] slow sent, +200ms later; fast returned in 215ms (expected ~200ms, far below 3000ms)
115
+ [probe] conclusion = requests do not block each other (concurrent handling)
116
+ ```
117
+
118
+ So `cancel_task` / `query_task` / `get_profiles` are handled normally during a wait; when the connection / request is cancelled, the wait loop exits immediately via the SDK-injected `extra.signal`, leaking no background wait.
119
+
120
+ ## 7. FAQ
121
+
122
+ **What happens on a timeout?**
123
+ Nothing happens to the task. The wait is read-only; a timeout returns only "current snapshot + call again". The caller can call repeatedly within the same turn (≈50 s per round; a 30–50 min task ≈ 40–60 calls).
124
+
125
+ **What if the task enters `needs_user`?**
126
+ The wait treats `needs_user` as a stop point and returns (it does not block until the timeout). After the user handles it in the client, `continue_task` resumes it and you call `wait_task` once more to keep waiting — `needs_user` can recur, which the loop pattern covers naturally.
127
+
128
+ **What if the client's single `tools/call` timeout is shorter (e.g. 30 s)?**
129
+ Set `timeoutMs` a little below it (for example 20000 ms). Even if truncated it is harmless: the caller just calls again to continue.
130
+
131
+ **Several tasks at once?**
132
+ Use `wait_any`. It returns the first task in `taskIds` array order that reaches a stop point, and lists every task's current status.
133
+
134
+ **What about the original wait call after a server restart?**
135
+ The wait is **in-process**: after a restart the original wait call ends with the connection. The caller reconnects and re-checks with `query_task` — historical leftover tasks are archived as `interrupted` at startup, and a wait returns **immediately** for an already-terminal task.
136
+
137
+ ## 8. Known limits
138
+
139
+ - The wait is **in-process**: after a server restart the original wait call ends with the connection (the caller re-checks with `query_task` after reconnecting).
140
+ - A single call waits at most **600 s** (a constant); longer scenarios rely on repeated calls (lossless).
141
+ - `wait_any` does not know "earliest finished"; it returns the first settled task in array order.
142
+ - The default `50000ms` is a conservative value for "unknown client timeout"; if your client's single-tool timeout is shorter, adjust per §5.
@@ -0,0 +1,142 @@
1
+ # 阻塞等待原语:`wait_task` / `wait_any`(issue #28)
2
+
3
+ 英文版:wait-task.en.md
4
+
5
+ `run_task` 是异步契约:秒回 `taskId`,不阻塞 `tools/call`。但目标调用方(天枢桌面端的 agent 会话)是**回合驱动**的——agent 只在收到用户消息的回合内执行,回合之间不运行,它**无法自行轮询**。于是在「回合驱动调用方 + 长任务」的组合下,工具面漏掉了任务完成时刻:过去每次任务完成都**必须人工发一条消息**触发查询。
6
+
7
+ 本能力补上这个空白:`wait_task` / `wait_any` 用**一次阻塞式只读调用**承载「等」这个动作,任务到达停点时返回,无需用户干预。
8
+
9
+ ## 一、为什么必须是「阻塞等待」
10
+
11
+ 服务端推送(`notifications/progress` 等)不可依赖:宿主 MCP 工具**只回文本**(`content[].text`)、**按次同步**调用 `tools/call`、不消费服务端推送(见 README「运行时契约」C1/C2)。因此「等」只能由**一次工具调用**承载——阻塞等待是工具面唯一可行的形态。
12
+
13
+ 它与 webhook 通知([issue #22](notifications.md))**互补、不重复**:webhook 的接收端是外部 HTTP 端点(人 / 机器人),结果不回流到 MCP 会话;`wait_task` 把结果送**回到发起调用的那个会话**。
14
+
15
+ ## 二、停点定义
16
+
17
+ `wait_task` / `wait_any` 等待的「可返回点」是**停点**:
18
+
19
+ ```text
20
+ isWaitSettled(status) = isTerminal(status) || status === "needs_user"
21
+ ```
22
+
23
+ - `isTerminal`:`succeeded` / `failed` / `needs_attention` / `cancelled` / `interrupted`。
24
+ - `needs_user`:**非终态**,任务已停止推进、在等人工处理(agent 提问 / 登录 / 环境处理),需 `continue_task` 恢复。
25
+
26
+ **为什么 `needs_user` 也是停点**:任务真正停止推进的时刻 = 调用方应当被唤醒的时刻。若不等它,任务进 `needs_user` 后 wait 会一直空等到 timeout,调用方**在超时前对「任务在等人」一无所知**——而这恰是需要立刻转达用户的状态。
27
+
28
+ ## 三、工具契约
29
+
30
+ ### `wait_task(taskId, timeoutMs?)`
31
+
32
+ 阻塞等待**单个**任务到达停点或超时。
33
+
34
+ | 入参 | 必填 | 说明 |
35
+ |---|---|---|
36
+ | `taskId` | 是 | 目标任务 id |
37
+ | `timeoutMs` | 否 | 本次等待上限(ms);缺省 `50000`、上限 `600000`,超上限被**钳制并在正文披露** |
38
+
39
+ 返回:文本(停点行 / 状态行 / 后续动作指引)+ meta 块,meta 含 `waitSettled`(是否到停点)与 `waitedMs`(实际等待时长)。
40
+
41
+ ### `wait_any(taskIds, timeoutMs?)`
42
+
43
+ 阻塞等待**一组**任务中**数组顺序首个**到达停点者。
44
+
45
+ | 入参 | 必填 | 说明 |
46
+ |---|---|---|
47
+ | `taskIds` | 是 | 1..20 个任务 id;开始前校验全部存在,**缺一即 fail-closed 报错并列出缺失 id** |
48
+ | `timeoutMs` | 否 | 同 `wait_task` |
49
+
50
+ 返回:到达停点的那一个任务的快照 + meta 块,正文另列出**全部任务当前状态行**。
51
+
52
+ > 返回「数组顺序首个已停」而非「完成时间最早」:确定性、可预测,避免 `finishedAt` 缺失 / 相同时的排序歧义。
53
+
54
+ ## 四、超时矩阵
55
+
56
+ | 调用 | `timeoutMs` 取值 | 行为 |
57
+ |---|---|---|
58
+ | `wait_task` / `wait_any` | 省略 | 使用默认 `50000ms`(低于生态常见 60s 客户端超时,留序列化 / 往返余量) |
59
+ | 同上 | 传 ≤ 600000 | 按传入值等待 |
60
+ | 同上 | 传 > 600000 | **钳制到 600000ms**,并在响应正文**如实写明**「已钳制到上限」 |
61
+ | 同上 | 等待到期仍未到停点 | 返回**当前快照** + 「请再次调用本工具继续等待」指引(`waitSettled=false`) |
62
+
63
+ `timeoutMs` 必须是正整数(schema 层拒绝 0 / 负数 / 非整数)。
64
+
65
+ ## 五、循环模式
66
+
67
+ 长任务(30–50 分钟)靠**循环调用**覆盖:每轮 ≈50s,直到停点或用户打断。
68
+
69
+ ### 5.1 单任务(成功 → 读报告)
70
+
71
+ ```text
72
+ run_task(...) → taskId
73
+ wait_task(taskId, timeoutMs=50000)
74
+ → 任务已到停点(等待 37 秒):状态: [PASS] 任务成功
75
+ → get_task_report(taskId)
76
+ ```
77
+
78
+ ### 5.2 超时续等
79
+
80
+ ```text
81
+ wait_task(taskId, timeoutMs=50000)
82
+ → 等待超时(50 秒):状态: 运行中(agent 正在开发)。任务本体不受影响;
83
+ 请再次调用 wait_task 继续等待,或用 query_task 查看细节。
84
+ wait_task(taskId) # 再次调用即续等(无损)
85
+ → …直到停点
86
+ ```
87
+
88
+ ### 5.3 needs_user 循环
89
+
90
+ ```text
91
+ wait_task(taskId)
92
+ → 任务已到停点(等待 12 秒):状态: 等待用户处理(可用 continue_task 恢复)。
93
+ 请用 continue_task 恢复,恢复后再次调用 wait_task 继续等待。
94
+ # 用户在客户端处理(回答问题 / 登录 / 关旧实例…),然后:
95
+ continue_task(taskId, message="已处理")
96
+ wait_task(taskId) # 恢复后继续等(needs_user 可多次进入)
97
+ ```
98
+
99
+ ### 5.4 多任务先到者
100
+
101
+ ```text
102
+ wait_any(taskIds=[tsk_a, tsk_b, tsk_c], timeoutMs=50000)
103
+ → 已有任务到达停点(等待 8 秒):tsk_b —— 状态: [FAIL] 任务失败
104
+ 全部任务当前状态:…
105
+ ```
106
+
107
+ ## 六、无损保证
108
+
109
+ `wait_task` / `wait_any` 是**纯只读**操作(`capability: "read"`、免审批、MCP `readOnlyHint: true`):不写任何任务状态、不动任务本体。被客户端截断、连接中断、超时返回——**任何路径都不影响任务继续执行**;最坏结果只是调用方多调几次,重连后 `query_task` 即拿到最新事实。
110
+
111
+ 等待期间**其他工具调用照常**(SDK 请求处理互不阻塞,已实测):
112
+
113
+ ```text
114
+ [probe] slow 发出后 +200ms;fast 返回耗时 = 215ms(期望 ~200ms,远小于 3000ms)
115
+ [probe] 结论 = 请求互不阻塞(并发处理)
116
+ ```
117
+
118
+ 因此等待期间 `cancel_task` / `query_task` / `get_profiles` 正常处理;连接 / 请求被取消时,等待循环经 SDK 注入的 `extra.signal` **立即退出**,不泄漏后台等待。
119
+
120
+ ## 七、FAQ
121
+
122
+ **问:超时会怎样?**
123
+ 任务零影响。wait 是只读的,超时只返回「当前快照 + 请再次调用」;调用方在同一回合内连续多次调用即可(每轮 50s,30–50 分钟任务 ≈ 40–60 次)。
124
+
125
+ **问:任务进入 `needs_user` 怎么办?**
126
+ wait 会把 `needs_user` 当停点返回(不会空等到超时)。让用户在客户端处理完后,`continue_task` 恢复,**再调一次 `wait_task`** 继续等——`needs_user` 可多次进入,循环模式天然覆盖。
127
+
128
+ **问:客户端单次 `tools/call` 超时更短(如 30s)怎么办?**
129
+ 把 `timeoutMs` 调到略低于该超时(例如 20000ms)。即便被截断也无害:调用方再次调用即可续等。
130
+
131
+ **问:多个任务要一起等?**
132
+ 用 `wait_any`。它按 `taskIds` 数组顺序返回首个到停点者,并列出全部任务当前状态。
133
+
134
+ **问:server 重启后原来的 wait 调用呢?**
135
+ 等待是**进程内**的:重启后原 wait 调用随连接终止。调用方重连后用 `query_task` 复核——历史遗留任务在启动时已被归档为 `interrupted`,wait 对已终态任务**立即返回**。
136
+
137
+ ## 八、已知限制
138
+
139
+ - 等待是**进程内**的:server 重启后原 wait 调用随连接终止(调用方重连后 `query_task` 复核)。
140
+ - 单次调用等待上限 **600s**(常量);更长场景靠循环调用(无损)。
141
+ - `wait_any` 不识别「完成时间最早」,只按数组顺序返回首个已停任务。
142
+ - 默认值 `50000ms` 是面向「客户端超时未知」的保守值;如你的客户端单次工具超时更短,按 §五调整。
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "tianshu-mcp",
3
- "version": "0.7.6",
3
+ "version": "0.7.7",
4
4
  "description": "天枢 × AI-Agent 编排 MCP server —— 驱动 Codex、TraeWork、ZCode、Kimi Code、Qoder CN 与 Open Design 完成项目开发、验收、失败返修与再验收闭环。",
5
5
  "type": "module",
6
6
  "license": "Apache-2.0",
@@ -27,6 +27,8 @@
27
27
  "docs/dry-run.en.md",
28
28
  "docs/notifications.md",
29
29
  "docs/notifications.en.md",
30
+ "docs/wait-task.md",
31
+ "docs/wait-task.en.md",
30
32
  "docs/visual-acceptance.md",
31
33
  "docs/visual-acceptance.en.md",
32
34
  "docs/visual-validation.md",
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tianshu-mcp
3
- description: 让外部 AI-Agent(codex/zcode/traework/kimicode/qoder/opendesign)做项目开发并自动验收、失败返修的编排方法。当任务需要“叫一个 AI-Agent 去开发/改代码/补测试并验收,不行就返修”时先加载本技能:按它用 mcp__tianshu-mcp__ 的 11 个工具(run_task/continue_task/query_task/list_tasks/get_task_report/verify_task/rework_task/cancel_task/get_profiles/prepare_visual_baseline/approve_visual_baseline)派活、暂停继续、轮询、查历史、读验收报告、驱动返修、管理视觉基准,并按硬失败错误码快速定位卡点。小改动或纯问答不需要。
3
+ description: 让外部 AI-Agent(codex/zcode/traework/kimicode/qoder/opendesign)做项目开发并自动验收、失败返修的编排方法。当任务需要“叫一个 AI-Agent 去开发/改代码/补测试并验收,不行就返修”时先加载本技能:按它用 mcp__tianshu-mcp__ 的 13 个工具(run_task/continue_task/query_task/list_tasks/get_task_report/verify_task/rework_task/cancel_task/get_profiles/wait_task/wait_any/prepare_visual_baseline/approve_visual_baseline)派活、阻塞等待、暂停继续、查历史、读验收报告、驱动返修、管理视觉基准,并按硬失败错误码快速定位卡点。小改动或纯问答不需要。
4
4
  triggers: '开发|编码|写代码|改代码|实现功能|加功能|修复|重构|补测试|写测试|验收|返修|返工|重做|自动验收|自动返修|任务书|ai.?agent|子代理|外部.?agent|agent|codex|zcode|traework|kimicode|kimi.?code|qoder|opendesign|open.?design|claude|编排|项目开发|派活|派单'
5
5
  ---
6
6
 
@@ -16,7 +16,8 @@ triggers: '开发|编码|写代码|改代码|实现功能|加功能|修复|重
16
16
 
17
17
  ```text
18
18
  run_task(秒回 taskId,异步)
19
- → query_task 轮询(5–10 秒一次)
19
+ → wait_task 阻塞等到停点(终态或 needs_user;超时后再调一次继续等)
20
+ → 需要进度细节时 query_task 轮询(5–10 秒一次)
20
21
  ├─ needs_user → §5:让用户在客户端处理 → continue_task 恢复
21
22
  ├─ 硬失败 → §9:读 agentEndReason,不要当“agent 没做好”重试
22
23
  └─ 终态 → §6:读 get_task_report 的 checks / analysis / visual
@@ -32,12 +33,14 @@ run_task(秒回 taskId,异步)
32
33
 
33
34
  ---
34
35
 
35
- ## 1. 工具面(11 个)
36
+ ## 1. 工具面(13 个)
36
37
 
37
38
  | 工具 | 能力 / 审批 | 作用 | 关键入参 |
38
39
  |---|---|---|---|
39
40
  | `run_task` | write + 审批 | 派活给外部 agent;**异步**返回 `taskId` | 见 §3 |
40
41
  | `query_task` | read | 轮询状态 + agent 日志尾(`tailLines` 缺省 40 行) | `taskId`、`tailLines?` |
42
+ | `wait_task` | read | **阻塞等待**单任务到停点(终态或 `needs_user`)或超时;超时后再次调用继续等。纯只读、无害 | `taskId`、`timeoutMs?`(缺省 50000、上限 600000) |
43
+ | `wait_any` | read | **阻塞等待**一组任务(1..20)中数组顺序首个到停点者;返回其快照 + 全部状态 | `taskIds`、`timeoutMs?` |
41
44
  | `list_tasks` | read | 查历史任务(每行:taskId / status / agent / project / 摘要) | `projectPath?`、`status?`、`limit?`(缺省 50,上限 200) |
42
45
  | `get_task_report` | read | 读某轮验收报告 **Markdown 全文** | `taskId`、`round?`(0-based,缺省最新) |
43
46
  | `verify_task` | execute(不改源码、无需审批) | 对任务或任意项目**独立验收**(会跑项目命令、可产生构建产物,故 `readOnlyHint=false`;不改源码、无需审批) | `taskId` 或 `projectPath` 二选一、`extraChecks?`、`checksMode?`、`baselineRef?` |
@@ -206,7 +209,8 @@ meta 的 `needsUserKind` 给出等待类型,`pendingQuestion` 给出问题原
206
209
 
207
210
  ## 7. 轮询与查询
208
211
 
209
- - `query_task(taskId, tailLines?)` 间隔 **5–10 秒**;返回状态行 + 最近消息 + agent 日志尾(缺省 40 行)。
212
+ - **首选 `wait_task`(issue #28)**:回合驱动调用方无法自行轮询——`run_task` 后在本回合内调 `wait_task(taskId)` **阻塞等到停点**(终态或 `needs_user`),不必等用户再发消息触发查询。超时(`timeoutMs` 缺省 50000、上限 600000)返回后**再次调用本工具继续等待**;等待纯只读、无害,被截断/中断对任务本体零影响。多任务并行等待用 `wait_any(taskIds)`。
213
+ - `query_task(taskId, tailLines?)` 间隔 **5–10 秒**;返回状态行 + 最近消息 + agent 日志尾(缺省 40 行)——要看进度细节时用它。
210
214
  - 查历史用 `list_tasks(projectPath?, status?, limit?)`(缺省 50,上限 200);`projectPath` 与 `run_task` 同样做 realpath 归一。
211
215
  - **同一项目勿重复派单**:每项目串行 + 全局并发(`concurrency.maxRunning`,默认 2);重复派只会排队,反而更慢。
212
216
  - **重试必须带幂等键(issue #15)**:`tools/call` 超时、连接抖动、宿主重启后重发同一意图时,**复用同一条 `idempotencyKey`** 调 `run_task` / `verify_task`——`run_task` 会返回原 `taskId` 与原状态(不会排队第二轮 agent),`verify_task` 执行中返回「进行中」、已完成返回既有报告(不会重跑 `build`/`e2e`/部署类检查)。**换参数就得换 key**:同键异参会直接报冲突。未带 key 时若响应里出现 `projectActiveTask`,说明该工作区已有未结束任务——先 `query_task` 复核,不要盲目再派。
@@ -327,5 +331,5 @@ meta 的 `needsUserKind` 给出等待类型,`pendingQuestion` 给出问题原
327
331
 
328
332
  1. `get_profiles` → 确认目标 agent `[PASS] 可用`(看 profileStatus 与探测来源);不可用就转达用户,别硬试。
329
333
  2. `run_task(projectPath=<绝对路径>, task=<任务书>, agentId=codex, model=<面板模型名>, autoVerify=true, autoFixRounds=5, idempotencyKey=<本次逻辑派单的稳定标识>)` → 拿 `taskId`。**model 以界面实际为准**,示例名不可当真;`idempotencyKey` 建议由宿主按「本次意图」生成一次并在所有重试中复用(见 §7)。
330
- 3. `query_task(taskId)` 每 ~8 秒轮询到终态;`needs_user` 按 §5 处理,硬失败按 §9 定位。
334
+ 3. `wait_task(taskId)` 阻塞等到停点(终态或 `needs_user`,超时后再调一次继续等);要看进度细节用 `query_task(taskId)` 每 ~8 秒轮询。`needs_user` 按 §5 处理,硬失败按 §9 定位。
331
335
  4. 终态按 §6 处理;汇报带 `get_task_report` 的 changedFiles 与 diffstat;启用视觉时一并读 `visual` 段落与离线 HTML。
@@ -224,11 +224,66 @@ run_task(projectPath=/path/to/项目, agentId=codex-cli,
224
224
  ### 2.10 通用约定
225
225
 
226
226
  - `run_task` 是**异步契约**:立即返回 `taskId` + 队列位置,不要当同步调用等结果。
227
- - 轮询间隔 5–10 秒(`query_task` 缺省返回 agent 日志末 40 行);同项目串行 + 全局并发默认 2,重复派单只会排队。
227
+ - **优先用 `wait_task` 等结果(issue #28)**:回合驱动调用方无法自行轮询,`run_task` 后在本回合内直接 `wait_task(taskId)` 阻塞等到停点,无需用户再发消息触发查询;要看进度细节才用 `query_task` 轮询(间隔 5–10 秒,缺省返回 agent 日志末 40 行)。同项目串行 + 全局并发默认 2,重复派单只会排队。
228
228
  - **重试复用同一条 `idempotencyKey`(issue #15)**:`tools/call` 超时、断线、宿主重启后重发同一意图时,`run_task` 会返回**原 `taskId` 与当前状态**(不排队第二轮 agent),`verify_task` 会返回「进行中」或既有报告(不重跑检查)。**参数变了就换 key**——同键异参 fail-closed 报错并回报原记录 id。幂等重放的响应文本以「幂等重放:」开头、meta 带 `idempotencyReplay`,不要汇报成「已重新派单」。
229
229
  - 只有 `needs_user` 能用 `continue_task` 恢复,且当前支持 **codex / zcode / kimicode / qoder / opendesign**(opendesign 会产出 `login_required` / `user_confirmation` / `system_permission` / `setup_recovery` / `close_existing_instance` 五类,均支持 `continue_task` 恢复);traework 与 spawn 类会被明确拒绝。
230
230
  - `autoVerify` 不传时**默认开**;`autoFixRounds` 不传时取 agent 缺省(codex 5 / zcode 2 / kimicode 2 / qoder 3 / traework 落 server 默认 0)。
231
231
 
232
+ ### 2.11 等待任务:`wait_task` / `wait_any`(issue #28)
233
+
234
+ **动机**:`run_task` 秒回 `taskId`,但回合驱动调用方(天枢 agent 会话)只在收到用户消息的回合内运行、无法自行轮询——过去「每次任务完成都必须人工发一条消息触发查询」。`wait_task` 用**一次阻塞只读调用**承载等待:等到任务到达**停点**(终态或 `needs_user`)或超时后返回。
235
+
236
+ **模式 A:单任务等待(成功 → 读报告)**
237
+
238
+ ```text
239
+ run_task(projectPath=D:/repo/app, agentId=codex, model="GPT-5.6 Sol",
240
+ task="…", autoVerify=true, autoFixRounds=2)
241
+ → taskId
242
+ wait_task(taskId=tsk_..., timeoutMs=50000)
243
+ → 任务已到停点(等待 37 秒):状态: [PASS] 任务成功
244
+ 可用 get_task_report 查看验收报告。
245
+ → get_task_report(taskId=tsk_...)
246
+ ```
247
+
248
+ **模式 B:超时循环(任务比单次上限长)**
249
+
250
+ ```text
251
+ wait_task(taskId=tsk_..., timeoutMs=50000)
252
+ → 等待超时(50 秒):状态: 运行中(agent 正在开发)。任务本体不受影响;
253
+ 请再次调用 wait_task 继续等待,或用 query_task 查看细节。
254
+ wait_task(taskId=tsk_...) # 再次调用即续等(无损)
255
+ → …直到停点
256
+ ```
257
+
258
+ **模式 C:needs_user 循环(任务在等人工处理)**
259
+
260
+ ```text
261
+ wait_task(taskId=tsk_...)
262
+ → 任务已到停点(等待 12 秒):状态: 等待用户处理(可用 continue_task 恢复)。
263
+ 任务在等待人工处理:请用 continue_task 恢复,恢复后再次调用 wait_task 继续等待。
264
+ # 让用户在客户端处理(回答问题 / 登录 / 关旧实例…),然后:
265
+ continue_task(taskId=tsk_..., message="已处理")
266
+ wait_task(taskId=tsk_...) # 恢复后继续等(needs_user 可多次进入)
267
+ ```
268
+
269
+ **模式 D:多任务先到者**
270
+
271
+ ```text
272
+ wait_any(taskIds=[tsk_a, tsk_b, tsk_c], timeoutMs=50000)
273
+ → 已有任务到达停点(等待 8 秒):tsk_b —— 状态: [FAIL] 任务失败
274
+ 全部任务当前状态:
275
+ - tsk_a: 运行中(agent 正在开发)
276
+ - tsk_b: [FAIL] 任务失败
277
+ - tsk_c: 排队中(每项目串行,等待前面任务完成)
278
+ ```
279
+
280
+ 要点:
281
+
282
+ - `wait_task` / `wait_any` 是**纯只读**工具(免审批、`readOnlyHint=true`):不写任务状态、不动任务本体;被客户端截断 / 连接中断 / 超时**都无害**,最坏只是多调几次。
283
+ - `timeoutMs` 缺省 **50000ms**(低于生态常见 60s 客户端超时),上限 **600000ms**;显式传超过上限的值会被**钳制并在响应正文写明**(不静默改值)。
284
+ - 停点含 **`needs_user`**(非终态):任务已停止推进、在等人工,必须立即唤醒调用方——这正是需要转达用户的时刻。
285
+ - `wait_any` 按 `taskIds` **数组顺序**返回首个到停点者(确定性优先,不看完成时间);入口校验全部 id 存在,缺一即 fail-closed 报错并列出缺失 id。
286
+
232
287
  ---
233
288
 
234
289
  ## 3. 参数速查(易错项)
@@ -245,6 +300,8 @@ run_task(projectPath=/path/to/项目, agentId=codex-cli,
245
300
  | `extraChecks` / `checksMode` / `baselineRef` | 仅 verify_task | 独立 projectPath 下 `baselineRef` 只能是 git ref,不能是任务 ID |
246
301
  | `idempotencyKey` | run_task / verify_task | trim 后 1..128 字符、不含控制字符;**两工具各自独立命名空间**;同键异参 fail-closed;不传即维持原行为 |
247
302
  | `tailLines` | query_task | 缺省 40 行 |
303
+ | `timeoutMs` | wait_task / wait_any | 缺省 **50000ms**、上限 **600000ms**;超上限被**钳制并在响应正文披露**;超时后再次调用即续等 |
304
+ | `taskIds` | wait_any | 1..20 个;**全部必须存在**,缺一即 fail-closed 报错并列出缺失 id |
248
305
  | `round` | get_task_report | **0-based**;缺省最新;显式 `0` 合法 |
249
306
 
250
307
  ---