@fateforge/xpedition-cli 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. package/.agent/AGENT.md +59 -0
  2. package/.agent/AGENT_zh.md +59 -0
  3. package/.agent/CLI-SPEC.md +1073 -0
  4. package/.agent/CLI-SPEC_zh.md +891 -0
  5. package/.agent/SEC-SPEC.md +158 -0
  6. package/.agent/SEC-SPEC_zh.md +132 -0
  7. package/.agent/SKILL-SPEC.md +266 -0
  8. package/.agent/SKILL-SPEC_zh.md +221 -0
  9. package/.agent/SPEC_VERSION +1 -0
  10. package/AGENTS.md +34 -0
  11. package/AGENTS_zh.md +33 -0
  12. package/CHANGELOG.md +795 -0
  13. package/CODE_OF_CONDUCT.md +35 -0
  14. package/CODE_OF_CONDUCT_zh.md +35 -0
  15. package/CONTRIBUTING.md +50 -0
  16. package/CONTRIBUTING_zh.md +42 -0
  17. package/LICENSE +21 -0
  18. package/NOTICE.md +16 -0
  19. package/NOTICE_zh.md +13 -0
  20. package/README.md +200 -0
  21. package/README_zh.md +178 -0
  22. package/SECURITY.md +108 -0
  23. package/SECURITY_zh.md +83 -0
  24. package/docs/AGENT_HARDENING_EVIDENCE.md +102 -0
  25. package/docs/AGENT_READS.md +74 -0
  26. package/docs/AGENT_READS_METRICS.json +216 -0
  27. package/docs/AGENT_READS_VALIDATION.json +13 -0
  28. package/docs/API_INVENTORY_BINDING_VALIDATION.json +16 -0
  29. package/docs/API_INVENTORY_DESIGN.md +90 -0
  30. package/docs/API_INVENTORY_REVIEW.md +59 -0
  31. package/docs/API_INVENTORY_VALIDATION.json +29 -0
  32. package/docs/API_INVENTORY_WINDOWS_VALIDATION.json +29 -0
  33. package/docs/COMPATIBILITY.md +499 -0
  34. package/docs/CONFIRMATION_CONCURRENCY_VALIDATION.json +33 -0
  35. package/docs/DIAGNOSTIC_BOUNDARIES.md +33 -0
  36. package/docs/DIAGNOSTIC_BOUNDARIES_VALIDATION.json +12 -0
  37. package/docs/E2E.md +445 -0
  38. package/docs/EVALS.md +134 -0
  39. package/docs/MCP.md +20 -0
  40. package/docs/NATIVE_ADAPTER.md +141 -0
  41. package/docs/OPEN_SOURCE_CHECKLIST.md +61 -0
  42. package/docs/OPEN_SOURCE_CHECKLIST_zh.md +61 -0
  43. package/docs/PIN_WORKFLOW_VALIDATION.json +28 -0
  44. package/docs/PLACEMENT_TASKS.md +99 -0
  45. package/docs/PLACEMENT_TASKS_VALIDATION.json +36 -0
  46. package/docs/REFERENCE_ADOPTION.md +67 -0
  47. package/package.json +48 -0
  48. package/scripts/run.js +46 -0
  49. package/skills/xpedition-cli/SKILL.md +300 -0
  50. package/skills/xpedition-cli/reference/agent-hardening.md +58 -0
  51. package/skills/xpedition-cli/reference/api-inventory.md +58 -0
  52. package/skills/xpedition-cli/reference/confirmation-safety.md +55 -0
  53. package/skills/xpedition-cli/test-prompts.json +62 -0
  54. package/skills/xpedition-pcb/SKILL.md +244 -0
  55. package/skills/xpedition-pcb/reference/fabrication.md +26 -0
  56. package/skills/xpedition-pcb/reference/hand-routing.md +33 -0
  57. package/skills/xpedition-pcb/reference/pcb-conventions.md +162 -0
  58. package/skills/xpedition-pcb/reference/placement-tasks.md +28 -0
  59. package/skills/xpedition-pcb/test-prompts.json +62 -0
  60. package/skills/xpedition-schematic/SKILL.md +244 -0
  61. package/skills/xpedition-schematic/reference/pin-assignment.md +61 -0
  62. package/skills/xpedition-schematic/reference/schematic-conventions.md +306 -0
  63. package/skills/xpedition-schematic/reference/schematic-design-format.md +219 -0
  64. package/skills/xpedition-schematic/test-prompts.json +52 -0
@@ -0,0 +1,1073 @@
1
+ # Agent-Facing CLI Design Spec
2
+
3
+
4
+ This document defines the machine contract a CLI must honor when called by an AI agent. The goal: agents can call it reliably, parse it reliably, recover and retry reliably, and never block or mis-write in a non-interactive setting.
5
+
6
+ ## 1. Core rules
7
+
8
+ 1. stdout is the contract: emit a single valid JSON document by default; no logs, progress, prompts, or color codes mixed in.
9
+ 2. stderr is the side channel: progress, warnings, debug, and error explanations all go to stderr.
10
+ 3. Machine-first: default `--format json`; `text` is for humans only; `raw` is for raw bytes, logs, diffs passed through verbatim.
11
+ 4. Non-interactive safe: write operations must not wait on keyboard input; use `--dry-run` + `--confirm <token>`.
12
+ 5. Deterministic: same input produces the same output structure; field names, field order, and schema version stay stable.
13
+ 6. Least surprise: queries don't change state; a write with no valid confirm token must fail rather than proceed.
14
+ 7. Recoverable: error codes, exit codes, and `retryable` must be stable enough for an agent to decide retry, back off, or ask the user.
15
+
16
+ ## 2. Global flags
17
+
18
+ | Flag | Meaning |
19
+ |------|---------|
20
+ | `--format json/text/raw` | Output format, default `json` |
21
+ | `--json` | Compatibility alias for `--format json`; not recommended for new calls |
22
+ | `--fields <a,b,c>` | Return only selected fields, reduces tokens (query commands) |
23
+ | `--compact` | Compact JSON output, strips redundant whitespace (query commands) |
24
+ | `--dry-run` | Simulate a write, return a change preview and `confirm_token` |
25
+ | `--confirm <token>` | Carry the dry-run token to actually execute the write |
26
+ | `--quiet` | Suppress progress/prompts on stderr, never suppress errors |
27
+
28
+ `update` is a **single command, not a write with a confirm gate** (see §14): a
29
+ bare `update` performs the whole self-update in one call. It may add
30
+ tool-specific flags such as `--target-version` or `--channel`, and it keeps
31
+ `--check` and `--dry-run` as **optional read-only** flags, but it does NOT take
32
+ `--confirm <token>` — self-update is exempt from the §7 write gate.
33
+
34
+ Format responsibilities:
35
+
36
+ - `json`: structured machine output, the default, and the only format recommended for agents.
37
+ - `text`: human-readable, may change, must not be parsed programmatically.
38
+ - `raw`: unwrapped bytes / log / diff, passed through verbatim, no JSON envelope.
39
+
40
+ ## 3. Unified output envelope
41
+
42
+ Success and failure share one shape. The agent only needs to check `ok` first.
43
+
44
+ Success:
45
+
46
+ ```json
47
+ {
48
+ "ok": true,
49
+ "schema_version": "1.0",
50
+ "data": {},
51
+ "meta": {
52
+ "duration_ms": 0
53
+ }
54
+ }
55
+ ```
56
+
57
+ Failure:
58
+
59
+ ```json
60
+ {
61
+ "ok": false,
62
+ "schema_version": "1.0",
63
+ "error": {
64
+ "code": "E_NOT_FOUND",
65
+ "message": "human readable message",
66
+ "details": {},
67
+ "retryable": false
68
+ },
69
+ "meta": {
70
+ "duration_ms": 0
71
+ }
72
+ }
73
+ ```
74
+
75
+ Conventions:
76
+
77
+ - Every JSON response must include `ok` and `schema_version`.
78
+ - `data` is always the command's business payload; do not hoist business fields to the envelope top level.
79
+ - `error.code` is a stable semantic enum, prefixed with `E_`.
80
+ - `error.message` is for humans; agents should not parse it.
81
+ - `error.details` holds structured context; must be redacted.
82
+ - `error.retryable` tells the agent whether it may back off and retry automatically.
83
+ - `meta.duration_ms` records command execution time. `meta` is always emitted on
84
+ every response (success and error); do not mark it `omitempty`, since
85
+ `duration_ms: 0` is a valid value the agent should always be able to read.
86
+ - `meta.notices[]` (optional) MAY carry ambient operational notices — currently
87
+ the update-available notice — read **only from the local cache**, never via a
88
+ network call. Each notice has a `severity` (`info` | `warning`).
89
+ Omit the field when there is nothing to report. See §14.
90
+ - A breaking schema change must bump the `schema_version` major version.
91
+
92
+ ### 3.1 Canonical machine contract (`contract.json`) — single source, enforced
93
+
94
+ The prose in §3/§6/§11 has a machine-readable twin: **`contract/contract.json`**,
95
+ maintained **only** in this template repo and the single source of truth for the
96
+ fields every tool shares. It encodes the envelope key sets, the `E_*` ↔ exit ↔
97
+ `retryable` table, the self-description required keys, naming convention,
98
+ pagination/batch shapes, and the `_untrusted` key.
99
+
100
+ - **One source, vendored copies, no drift.** A tool does not hand-author these.
101
+ It vendors `contract.json` (and the `.agent` specs) pinned to a spec tag in
102
+ `.agent/SPEC_VERSION`, regenerates a per-language module (`contract_gen.{go,py}`)
103
+ from it via `scripts/gen-contract.js`, and a fail-closed CI guard
104
+ (`scripts/check-spec.js`) keeps both byte-identical to `template@<pin>` and the
105
+ generated module in sync with `contract.json`. Editing the contract happens in
106
+ the template; tools bump the pin and run `scripts/sync-spec.js`.
107
+ - **Core is frozen; features extend, never redefine.** `error_codes.core` is
108
+ identical in every tool (this is what makes exit-code/retryable behavior
109
+ portable). A tool's unique error codes live in `contract-ext.json` and are
110
+ validated: an ext code is `E_*`, must declare `{exit, retryable}`, must not
111
+ shadow a core code, and exit `9` is reserved for human-action codes
112
+ (`E_HUMAN_REQUIRED`, `E_2FA_REQUIRED`). This is how the contract stays uniform
113
+ while supporting tool-specific fields.
114
+ - **Enforcement calibration.** Envelope keys, the error-code table, `meta` keys,
115
+ and the `schema_version` value are matched **exactly**. Self-description blocks
116
+ (`reference` / `context` / `doctor` / `changelog` / `update`) must contain the
117
+ **required** canonical keys but may add tool-specific ones — so a unique feature
118
+ command is never blocked, only the shared surface is pinned. A runtime
119
+ conformance test (per tool) asserts actual command output against
120
+ `contract.json`.
121
+
122
+ ## 4. stdout / stderr rules
123
+
124
+ - In `json` mode, stdout may contain only one JSON document, or NDJSON for explicitly streaming commands.
125
+ - stderr may carry progress, warnings, and diagnostics.
126
+ - On error in `json` mode, the failure envelope is that single JSON document on stdout — agents always parse stdout and check `ok`, never scrape stderr. stderr may add human-readable explanation.
127
+ - `--quiet` may only suppress non-error info on stderr.
128
+ - No banners, prompts, progress bars, or color codes before/after the JSON on stdout.
129
+ - stdout / stderr are always **UTF-8 encoded, no BOM**, newline `\n`, so agents parse reliably across platforms (especially Windows).
130
+
131
+ ## 5. Streaming output (NDJSON)
132
+
133
+ Large output, log streams, subscription streams, and per-item batch results use NDJSON. Each line must be an independent valid JSON object — easy to consume streaming, low memory, interruptible:
134
+
135
+ ```jsonl
136
+ {"ok":true,"schema_version":"1.0","type":"item","data":{}}
137
+ {"ok":true,"schema_version":"1.0","type":"item","data":{}}
138
+ {"ok":true,"schema_version":"1.0","type":"summary","data":{"count":2}}
139
+ ```
140
+
141
+ Conventions:
142
+
143
+ - Normal queries use a single JSON envelope by default.
144
+ - Use NDJSON only when the command is explicitly a log / stream / subscribe / batched-stream.
145
+ - NDJSON lines must include `ok`, `schema_version`, `type`.
146
+ - The final line should use `type: "summary"`.
147
+ - True binary or plain-text passthrough goes through `--format raw`, not wrapped into one giant JSON.
148
+
149
+ ## 6. Exit code table
150
+
151
+ | Code | Meaning | Agent behavior |
152
+ |------|---------|----------------|
153
+ | 0 | Success | continue |
154
+ | 1 | Generic error | read the error envelope to decide |
155
+ | 2 | Argument/usage error | don't retry, fix args |
156
+ | 3 | Resource not found | don't retry |
157
+ | 4 | Permission/auth/config failure | don't retry, surface credentials or permission |
158
+ | 5 | Confirmation required but token missing | run dry-run for a token, then retry |
159
+ | 6 | Precondition conflict or invalid token | re-read state, then retry |
160
+ | 7 | Retryable transient error (network/rate-limit/server) | back off and retry |
161
+ | 8 | Timeout | back off and retry |
162
+ | 9 | Human action required (see §16.3, optional) | relay to the user, run `resume` once done |
163
+
164
+ Error codes and exit codes must align:
165
+
166
+ - `E_USAGE` / `E_VALIDATION` -> 2
167
+ - `E_NOT_FOUND` -> 3
168
+ - `E_AUTH` / `E_FORBIDDEN` / `E_CONFIG` -> 4
169
+ - `E_CONFIRMATION_REQUIRED` -> 5
170
+ - `E_CONFLICT` -> 6
171
+ - `E_NETWORK` / `E_RATE_LIMITED` / `E_SERVER` -> 7
172
+ - `E_TIMEOUT` -> 8
173
+ - `E_INTEGRITY` -> 1 (release integrity failure: missing/invalid signature or checksum mismatch; **non-retryable**, see §14)
174
+ - `E_IO` -> 1 (local filesystem failure: disk full, file locked, partial write; non-retryable, needs an environment fix; see §14 update replace stage)
175
+ - `E_INTERRUPTED` -> 130 (operation cancelled by signal/user, SIGINT = 128+2; retryable — staged work leaves nothing half-applied, see §14)
176
+ - `E_HUMAN_REQUIRED` -> 9 (optional, only when §16.3 is enabled)
177
+
178
+ When the failure comes from an upstream HTTP call, map the status onto the
179
+ taxonomy so the agent can tell failure modes apart from `error.code` +
180
+ `retryable` — do NOT collapse every 4xx into `E_NETWORK`:
181
+
182
+ - `401` -> `E_AUTH`
183
+ - `403` -> `E_FORBIDDEN`
184
+ - `404` -> `E_NOT_FOUND`
185
+ - `408` -> `E_TIMEOUT` (retryable)
186
+ - `409` -> `E_CONFLICT`
187
+ - `429` -> `E_RATE_LIMITED` (retryable)
188
+ - `5xx` -> `E_SERVER` (retryable)
189
+ - connection refused / DNS / reset -> `E_NETWORK` (retryable)
190
+
191
+ Map by the upstream's own error TYPE/status where available, not by sniffing
192
+ the human-readable message text (substring matching misclassifies messages that
193
+ merely contain words like "not found"). Keep this mapping in ONE function so the
194
+ status->code->exit contract cannot drift between the output layer and the
195
+ command layer. Codes that are declared but never reachable should be annotated
196
+ as reserved so an agent does not plan for a branch that cannot occur.
197
+
198
+ ## 7. Write flow (dry-run -> confirm)
199
+
200
+ A write command must first support `--dry-run`, returning a preview and a token:
201
+
202
+ ```json
203
+ {
204
+ "ok": true,
205
+ "schema_version": "1.0",
206
+ "data": {
207
+ "preview": {
208
+ "changes": [
209
+ {
210
+ "action": "delete",
211
+ "resource": "mail",
212
+ "id": "123",
213
+ "before": {},
214
+ "after": null
215
+ }
216
+ ]
217
+ },
218
+ "confirm_token": "ct_9f2a...",
219
+ "expires_at": "2026-06-05T12:00:00Z"
220
+ },
221
+ "meta": {
222
+ "duration_ms": 0
223
+ }
224
+ }
225
+ ```
226
+
227
+ The second step carries the token to execute:
228
+
229
+ ```bash
230
+ tool resource delete --id 123 --confirm ct_9f2a...
231
+ ```
232
+
233
+ Confirm-token conventions:
234
+
235
+ - The token must bind a hash of the operation content: command path, args, target resource ID, calling account, permission context.
236
+ - The hash must be keyed (HMAC) with a machine-local secret (e.g. `~/.<tool>/confirm.secret`, created on first use, `0600`), so a token cannot be fabricated by recomputing a public hash — it must come from a real `--dry-run` on the same machine.
237
+ - When a resource version is available, also bind it (version, etag, changekey, or updated_at) to prevent state drift.
238
+ - The token must expire; `expires_at` is ISO 8601 UTC.
239
+ - On expiry, changed args, or changed target state, execution returns `E_CONFLICT`, exit code 6.
240
+ - With no token, return `E_CONFIRMATION_REQUIRED`, exit code 5.
241
+ - dry-run must not cause external side effects, but may read state to build the preview.
242
+
243
+ ## 8. Query, pagination, and field selection
244
+
245
+ Query commands support, by default:
246
+
247
+ - `--fields <a,b,c>`: return only selected fields; when dotted paths are supported, declare it in reference.
248
+ - `--compact`: strip JSON whitespace.
249
+ - `--limit`: cap the number of returned items.
250
+ - `--cursor` or `--offset`: pagination cursor or offset.
251
+
252
+ Suggested pagination shape:
253
+
254
+ ```json
255
+ {
256
+ "items": [],
257
+ "count": 0,
258
+ "next_cursor": null,
259
+ "has_more": false
260
+ }
261
+ ```
262
+
263
+ For offset-based upstreams, echo `offset` and return an explicit `next_offset`
264
+ (the value to pass next, present only while `has_more` is true) so the agent
265
+ pages deterministically instead of re-deriving `offset + count`:
266
+
267
+ ```json
268
+ {
269
+ "items": [],
270
+ "count": 0,
271
+ "offset": 0,
272
+ "next_offset": 20,
273
+ "has_more": true
274
+ }
275
+ ```
276
+
277
+ When a list is silently capped (e.g. an auto-paginate ceiling), surface
278
+ `truncated: true` rather than returning a short list that looks complete.
279
+
280
+ Conventions:
281
+
282
+ - All IDs are strings, even if numeric underneath.
283
+ - All times are ISO 8601 UTC.
284
+ - List order must be stable; declare the default sort in reference.
285
+ - Query commands must not fall into an interactive prompt just because an optional filter is missing.
286
+
287
+ ### 8.1 Server-side filters over client-side faking
288
+
289
+ Prefer pushing a filter to the upstream over post-filtering a page client-side. A
290
+ filter applied after pagination silently undercounts: it looks complete but only
291
+ reflects the fetched page. If the upstream gained a filter in a known version, map
292
+ the flag to it and record the minimum version (reference + compatibility doc)
293
+ rather than emulating it; if you must filter locally, page the full set first and
294
+ say so.
295
+
296
+ ### 8.2 Heavy sub-resources are separate, structured, and projectable
297
+
298
+ A sub-resource whose size is unbounded (a diff, a log, an artifact, a full file
299
+ body) is its own command, never inlined into a list. The cheap, bounded summary
300
+ (counts, paths, stats) belongs on the list; the heavy payload is fetched on demand
301
+ for the specific item the agent chose. Return it **structured** (e.g. a diff as
302
+ per-file entries, not one opaque blob) so `--fields` can project it down to an
303
+ inventory (paths + line counts) without shipping the payload. This makes the
304
+ agent's token cost a choice, not a surprise — no bespoke truncation protocol
305
+ required.
306
+
307
+ ### 8.3 Multi-scope queries fan out under the batch contract
308
+
309
+ When a read spans many containers (projects in a group, every project in the
310
+ instance), resolve the scope to a concrete set and fan out as a client-side loop
311
+ (§15, class B): one external command, exactly one scope selector, results
312
+ aggregated with each item annotated by its source container. A container that
313
+ fails to scan is reported in the result (e.g. `projectErrors[]` / `scope` /
314
+ `projectsScanned`), never silently dropped, and a single failure must not abort
315
+ the rest. Aggregating bounded metadata across containers is safe; never aggregate
316
+ an unbounded sub-resource (§8.2) across the whole set. An instance-wide scope that
317
+ only makes sense for one actor (all of a user's commits) must be bound to that
318
+ actor and fail closed otherwise, so a bare unbounded scan is impossible.
319
+
320
+ ## 9. Idempotency and concurrency safety
321
+
322
+ Write commands should support idempotent semantics where possible:
323
+
324
+ - Create-type commands should support `--request-id` or `--idempotency-key`. Where the upstream
325
+ honors an idempotency header (e.g. GitLab's `Idempotency-Key`), forward it; bind the key into the
326
+ confirm scope so the token matches only that key.
327
+ - Retrying the same idempotency key must not create duplicate resources.
328
+ - Update/delete commands should record the target resource version during dry-run.
329
+ - If a version change is detected at confirm time, return `E_CONFLICT`.
330
+ - **Confirm tokens are single-use.** Once a token has been accepted to execute a write, record its
331
+ fingerprint (e.g. under `~/.<tool>/confirm-consumed.json`, pruned by expiry) and reject any replay
332
+ with `E_CONFLICT` ("token already used; re-run `--dry-run`"). This gives agents safe-retry: a
333
+ confirmed write that times out cannot be blindly re-sent — the retry is rejected and re-running
334
+ `--dry-run` reveals the now-current state. This is the universal safe-retry mechanism for upstreams
335
+ that expose no resource version to bind. Mark consumed BEFORE the write executes (a crash mid-write
336
+ conservatively blocks the replay rather than risking a duplicate). A storage failure must degrade
337
+ gracefully and never block the write.
338
+ - Batch writes should return per-item results; don't hide other items' status because one failed.
339
+
340
+ Suggested batch-write result:
341
+
342
+ ```json
343
+ {
344
+ "results": [
345
+ {
346
+ "id": "1",
347
+ "ok": true,
348
+ "action": "deleted"
349
+ },
350
+ {
351
+ "id": "2",
352
+ "ok": false,
353
+ "error": {
354
+ "code": "E_NOT_FOUND"
355
+ }
356
+ }
357
+ ],
358
+ "summary": {
359
+ "ok_count": 1,
360
+ "error_count": 1
361
+ }
362
+ }
363
+ ```
364
+
365
+ ## 10. Sensitive data and auditing
366
+
367
+ - password, token, secret, authorization header, cookie must not appear in stdout, stderr, error.details, or the audit log.
368
+ - dry-run previews must redact sensitive fields.
369
+ - reference/context/doctor must not leak plaintext credentials.
370
+ - context may report whether credentials exist, but only as a boolean or redacted summary.
371
+ - The audit log should record command path, redacted args, calling account, time, exit code, duration.
372
+ - `--quiet` must not disable auditing.
373
+
374
+ ## 11. Self-description commands (reference / context / doctor / changelog)
375
+
376
+ ### reference
377
+
378
+ Declares the tool's capabilities, commands, params, output schema, error codes, and permission levels, so an agent understands the tool first.
379
+
380
+ Each command's `output_schema` MUST be machine-usable, not a stub. Use a string label that
381
+ resolves to an entry in a top-level `schemas` catalog: `{ "shape": "object"|"array", "fields":
382
+ [...], "untrusted_fields": [...] }`, with the field list enumerated from the command's actual
383
+ returned data (the flatten structs / `*ToMap` builders) and `untrusted_fields` listing the
384
+ attacker-controllable keys. Each command SHOULD also carry `examples`: one runnable invocation
385
+ (write commands show the `--dry-run` then `--confirm` pair, dangerous commands include
386
+ `--dangerous`). A guard test SHOULD assert every leaf command resolves to a non-empty schema and has
387
+ at least one example, so `reference` cannot silently regress to a stub.
388
+
389
+ ```json
390
+ {
391
+ "ok": true,
392
+ "schema_version": "1.0",
393
+ "data": {
394
+ "tool": "tool-name",
395
+ "version": "1.0.0",
396
+ "release_readiness": {
397
+ "level": "beta",
398
+ "fcc_required": true,
399
+ "fcc_status": "verified",
400
+ "mock_upstream_required": true,
401
+ "mock_upstream_status": "verified",
402
+ "live_smoke_required_for_stable": true,
403
+ "live_smoke_status": "missing",
404
+ "reason": "Stable requires recorded live smoke/E2E evidence.",
405
+ "required_evidence": [
406
+ "functional_contract_coverage_100",
407
+ "mock_upstream_contract_tests",
408
+ "recorded_live_smoke_for_stable"
409
+ ]
410
+ },
411
+ "commands": [
412
+ {
413
+ "path": "resource delete",
414
+ "type": "write",
415
+ "description": "Delete a resource",
416
+ "params": [
417
+ {
418
+ "name": "id",
419
+ "type": "string",
420
+ "required": true,
421
+ "multiple": false
422
+ }
423
+ ],
424
+ "output_schema": "deleted_resource",
425
+ "examples": [
426
+ "<tool> resource delete <id> --dry-run --compact",
427
+ "<tool> resource delete <id> --confirm <confirm_token> --compact"
428
+ ]
429
+ }
430
+ ],
431
+ "schemas": {
432
+ "deleted_resource": {
433
+ "shape": "object",
434
+ "fields": ["id", "status"],
435
+ "untrusted_fields": []
436
+ }
437
+ },
438
+ "exit_codes": {}
439
+ },
440
+ "meta": {
441
+ "duration_ms": 0
442
+ }
443
+ }
444
+ ```
445
+
446
+ `release_readiness` is the machine-readable release gate. It must appear in
447
+ `reference` for every AI-native CLI:
448
+
449
+ - `level`: `stable`, `beta`, or `unpublishable`.
450
+ - `stable`: FCC is 100%, mock upstream/contract tests cover external behavior,
451
+ and at least one recorded live smoke/E2E run exists for the release candidate.
452
+ - `beta`: FCC is 100% and mock upstream/contract tests exist, but live
453
+ smoke/E2E evidence is missing or explicitly not available yet.
454
+ - `unpublishable`: any public behavior lacks command-level coverage, or mock
455
+ upstream/contract tests cover only happy paths while failure/auth/pagination/
456
+ empty/rate-limit behavior remains untested.
457
+ - `fcc_status`, `mock_upstream_status`, and `live_smoke_status` use
458
+ `verified`, `missing`, `not_applicable`, or `unknown`; `stable` may not use
459
+ `missing` or `unknown` for any required item.
460
+ - `required_evidence[]` names the evidence an agent or release script should
461
+ inspect before trusting the level.
462
+
463
+ ### context
464
+
465
+ Reports the current runtime, config, target, and credential status.
466
+
467
+ ```json
468
+ {
469
+ "ok": true,
470
+ "schema_version": "1.0",
471
+ "data": {
472
+ "env": "prod",
473
+ "account": "user@example.com",
474
+ "config": {},
475
+ "credentials": {
476
+ "configured": true
477
+ }
478
+ },
479
+ "meta": {
480
+ "duration_ms": 0
481
+ }
482
+ }
483
+ ```
484
+
485
+ ### doctor
486
+
487
+ Environment and risk check-up; each item gives an actionable fix.
488
+
489
+ ```json
490
+ {
491
+ "ok": true,
492
+ "schema_version": "1.0",
493
+ "data": {
494
+ "checks": [
495
+ {
496
+ "check": "auth",
497
+ "status": "pass",
498
+ "fix": null
499
+ },
500
+ {
501
+ "check": "network",
502
+ "status": "fail",
503
+ "fix": "set HTTP_PROXY or check VPN"
504
+ },
505
+ {
506
+ "check": "release_readiness",
507
+ "status": "warn",
508
+ "fix": "record live smoke/E2E evidence before declaring stable"
509
+ }
510
+ ]
511
+ },
512
+ "meta": {
513
+ "duration_ms": 0
514
+ }
515
+ }
516
+ ```
517
+
518
+ `doctor` must include `check: "release_readiness"` with the same release level
519
+ reported by `reference`. Use `pass` for `stable`, `warn` for intentional `beta`,
520
+ and `fail` for `unpublishable` or for a declared `stable` state with missing
521
+ evidence. The check should include an actionable `fix` when the status is not
522
+ `pass`.
523
+
524
+ ### changelog
525
+
526
+ Reports **what changed between versions** so an agent that just self-updated can refresh its knowledge instead of reusing stale patterns. This is the time-axis complement to `reference` (which describes current capabilities).
527
+
528
+ ```bash
529
+ tool changelog # all version changes
530
+ tool changelog --since 1.0.3 # only versions newer than 1.0.3
531
+ ```
532
+
533
+ ```json
534
+ {
535
+ "ok": true,
536
+ "schema_version": "1.0",
537
+ "data": {
538
+ "current_version": "1.1.0",
539
+ "since": "1.0.3",
540
+ "entries": [
541
+ {
542
+ "version": "1.1.0",
543
+ "date": "2026-06-07",
544
+ "changes": {
545
+ "added": [
546
+ "..."
547
+ ],
548
+ "changed": [
549
+ "..."
550
+ ],
551
+ "fixed": []
552
+ }
553
+ }
554
+ ]
555
+ },
556
+ "meta": {
557
+ "duration_ms": 0
558
+ }
559
+ }
560
+ ```
561
+
562
+ Conventions:
563
+
564
+ - **Single source of truth**: `changelog` output is derived from `CHANGELOG.md` (embedded into the binary at build time by `## [version]` section); no separate data maintained. Same source as release notes.
565
+ - `--since <version>` returns only entries strictly newer than that version, for an agent that "last saw version X" to pull the delta.
566
+ - Change categories follow Keep a Changelog: `added` / `changed` / `fixed` / `deprecated` / `removed` / `security`.
567
+ - After a successful self-update, the tool should hint the agent to run `changelog --since <old version>` (see §14).
568
+
569
+ ## 12. Command design conventions
570
+
571
+ 1. Use the shortest command that completes a clear task; reduce combinatorial complexity.
572
+ 2. Query commands support `--fields` and `--compact` by default to cut tokens.
573
+ 3. Write commands must support `--dry-run` and `--confirm`.
574
+ 4. Naming uses `<noun> <verb>` or `<verb> <noun>` style, consistent across the tool.
575
+ 5. Don't require agents to parse help text; `--help` is for humans, machine capability is exposed via `reference`.
576
+ 6. All times ISO 8601 UTC; all IDs strings.
577
+ 7. On failure, return a structured error rather than a half-finished success payload.
578
+ 8. Avoid ambiguous params; booleans are flags, enums are bounded choices.
579
+
580
+ ## 13. Functional contract coverage and release gate
581
+
582
+ Functional Contract Coverage (FCC) is the release blocker: every public behavior
583
+ an agent can rely on must have automated command-level test coverage. Numeric
584
+ line or branch coverage is useful engineering telemetry, but it is secondary and
585
+ must not be used as a substitute for missing functional contract tests.
586
+
587
+ A public functional contract is anything declared in:
588
+
589
+ - `README.md` / `README_zh.md`, `SKILL.md`, or Skill reference pages;
590
+ - `tool reference`, `--help`, `context`, `doctor`, `changelog`, or `update`
591
+ output;
592
+ - JSON envelope fields, command output schemas, global flags, error codes, exit
593
+ codes, retryability, and stdout/stderr behavior;
594
+ - documented config/env variables, credential handling, write safety, update
595
+ verification, Skill sync, and `_untrusted` security guarantees.
596
+
597
+ Required coverage for each public command or contract:
598
+
599
+ - success path;
600
+ - missing/invalid arguments;
601
+ - missing config, missing auth, or permission failure when applicable;
602
+ - upstream API failure, network failure, rate limit, or timeout when applicable;
603
+ - JSON envelope shape, output schema, exit code, and stderr/stdout boundary;
604
+ - non-interactive behavior: no prompts, no blocking, and write commands use
605
+ `--dry-run` -> `--confirm <token>`;
606
+ - regression test for every bug fix that changes observable behavior.
607
+
608
+ What `FCC = 100%` means:
609
+
610
+ - every command/flag/output/error behavior listed in the public contract is
611
+ mapped to at least one automated test, or explicitly marked non-applicable;
612
+ - command-level tests validate the CLI boundary, not just internal helpers;
613
+ - broad generated code, version constants, build metadata, or unreachable
614
+ platform guards may be excluded from numeric coverage, but not from FCC if
615
+ they are documented public behavior;
616
+ - a release cannot be tagged while known FCC gaps remain;
617
+ - `fcc_status: "verified"` must be machine-backed by an enumeration guard
618
+ test: enumerate every leaf command from live `reference` output and assert
619
+ each one is invoked by at least one command-level test. The guard skips
620
+ while the status is honestly declared `missing`, and fails if the claim is
621
+ flipped to `verified` without the coverage (the template ships this guard).
622
+
623
+ CI should run the unit and command-level suites for every PR. Numeric coverage
624
+ thresholds may ratchet upward over time per repository, but the release standard
625
+ is absolute: public functional contracts must be covered.
626
+
627
+ ### Release readiness levels
628
+
629
+ Release readiness is deliberately stricter than "tests pass":
630
+
631
+ - **Stable**: FCC is 100%; mock upstream/contract tests cover success,
632
+ validation, config/auth/permission failures, upstream/network/rate-limit/
633
+ timeout failures, empty results, pagination, output schema, exit codes, and
634
+ stdout/stderr boundaries; at least one live smoke/E2E run has been recorded
635
+ for the release candidate.
636
+ - **Beta**: FCC is 100% and mock upstream/contract tests meet the same
637
+ behavioral breadth, but recorded live smoke/E2E evidence is missing or the
638
+ project explicitly declares that live E2E is not available yet.
639
+ - **Unpublishable**: any public command/flag/output/error behavior lacks
640
+ command-level coverage, or mock upstream tests only cover happy paths.
641
+
642
+ `reference.release_readiness` and `doctor.checks[]` are the machine-readable
643
+ surface for this gate. A repository may choose not to publish `beta` artifacts,
644
+ but it must not describe itself as `stable` without the live evidence above.
645
+
646
+ ## 14. Versioning and compatibility
647
+
648
+ - `schema_version` is the output schema version, not the tool version.
649
+ - A breaking schema change bumps the major version, e.g. `1.x` -> `2.0`.
650
+ - Non-breaking added fields may keep the major version.
651
+ - Deprecated fields should keep a compatibility window and be marked deprecated in reference.
652
+ - Compatibility aliases may exist but should not be the recommended usage in new docs.
653
+ - Agents should rely on `reference`, not `--help` or README.
654
+
655
+ ### Version negotiation (tool version ↔ Skill expectation)
656
+
657
+ A Skill is a snapshot of the capabilities the day it was written; once the binary version drifts, things misalign: a Skill written for v1.1 against a v1.0 binary will silently call commands that don't exist.
658
+
659
+ - The tool must report its own version: `tool --version` and `context.data.version`.
660
+ - The Skill declares a minimum compatible version in frontmatter (see SKILL-SPEC `requires.min_version`).
661
+ - `doctor` should include a check "does the current version meet the declared minimum"; if not, give a `fix` (upgrade command), status `fail`.
662
+
663
+ ### Self-update and Skill-sync loop
664
+
665
+ For tools with `self-update`, after a successful update they **must close both
666
+ refresh loops**:
667
+
668
+ 1. the binary/package is current;
669
+ 2. every bundled Agent Skill directory is current, with the same end state as
670
+ running `npx skills add <repo> -y -g`.
671
+
672
+ The user-facing Skill install command stays `npx skills add ...`; the binary
673
+ must not expose a separate `install-skill` command. During update, however, the
674
+ tool owns the full lifecycle and must either sync every Skill directory under
675
+ `skills/` — running `npx skills add <repo> -y -g` does exactly that, and syncing
676
+ only `skills/<tool>/` would leave a tool's other Skills (SKILL-SPEC §7) behind —
677
+ or return an explicit `skill_sync_status` and `skill_sync_command`
678
+ that the agent can execute before using new behavior.
679
+
680
+ Single-command update contract (no leaf commands, no confirm token):
681
+
682
+ - A bare `update` performs the whole update in ONE call: resolve the latest (or
683
+ `--target-version`) release, verify its integrity, replace the binary/package,
684
+ then sync the Skill directories. Self-update is a single-target, non-destructive,
685
+ self-verifying operation, so it is **exempt from the §7 `--dry-run → --confirm
686
+ <token>` write gate** — the safety guarantee is the in-process signature
687
+ verification below, not an agent's review of a preview. There are no `update`
688
+ leaf subcommands.
689
+ - **Install-method dispatch — drive the manager, don't just print the command.**
690
+ "Replace the binary/package" means reaching the upgraded end state in that one
691
+ call for *every* install method, not only standalone binaries:
692
+ - **Standalone binary** (the tool owns the file): download → in-process Sigstore
693
+ signature verify → checksum → atomic in-place swap; `signature_status:
694
+ "verified"`.
695
+ - **Package-manager-managed install** (npm / Go / Homebrew — the manager owns the
696
+ file): the tool MUST NOT mutate the managed file in place (that desyncs the
697
+ manager's metadata) and MUST NOT merely return the command for the user to run.
698
+ It DRIVES the manager — it executes the install command on the user's behalf
699
+ (e.g. `npm install -g <pkg>@<version>`), then syncs the Skill, reaching the same
700
+ end state with `status: "updated"`. Integrity on this path is the package
701
+ manager's own (registry integrity/provenance), so `signature_status` is
702
+ `not_checked`; the new version takes effect on the next invocation. Detection
703
+ must be robust (don't misclassify a standalone binary as managed); a failed
704
+ manager invocation reports the error with `binary_replaced: false` and the
705
+ exact command to run manually.
706
+ - `update --check` is an OPTIONAL read-only probe: report current/target
707
+ versions, install method, whether an update is available, whether Skill sync is
708
+ supported, and signature/checksum availability. It changes nothing.
709
+ - `update --dry-run` is an OPTIONAL read-only preview of the same changes
710
+ (binary/package update, Skill sync, verification plan). It issues NO token and
711
+ is never a required step before `update`.
712
+ - `update` is idempotent: when already on the latest (or requested) version it
713
+ returns `ok` with a no-op result, so an agent may call it freely. This no-op
714
+ check MUST run before any package-manager command; an already-current install
715
+ must not re-run `npm`, `go`, `pip`, `brew`, or an equivalent manager.
716
+ - Successful and no-op `update` results describe the final post-command state,
717
+ not the pre-install comparison. After a successful install, `data` carries
718
+ `previous_version`, `current_version`, `target_version`, `update_available`,
719
+ `signature_verified`, `signature_status`, `skill_sync_status`, and enough
720
+ verification metadata for the agent to audit what happened. `current_version`
721
+ MUST equal `target_version`, and `update_available` MUST be `false`. In a
722
+ no-op result, `current_version` and `target_version` are also equal and
723
+ `update_available` is `false`.
724
+ - If the binary/package updates but Skill sync fails, return partial success
725
+ (`ok: false`, `binary_replaced: true`) with `target_version`,
726
+ `update_available: false`, and `skill_sync_command`; the agent must not use
727
+ newly documented behavior until the Skill sync has completed.
728
+
729
+ Version notification contract:
730
+
731
+ - `update --check` actively checks the latest release and refreshes the local
732
+ update notice cache.
733
+ - `doctor` may actively check with a short timeout; network failure must not
734
+ make `doctor` fail by itself.
735
+ - `context` and `--help` only read the local cache and must not contact remote
736
+ registries or GitHub.
737
+ - The cached notice MAY also be attached to **every command's `meta.notices`**,
738
+ read **only from the local cache** (no network; cost is one local file read).
739
+ Business commands surface the cached notice — they never actively check / phone
740
+ home. Omit `meta.notices` when the cache has nothing to report.
741
+ - After a successful install, partial success after the binary/package commit,
742
+ or an idempotent no-op at the target/latest version, the tool MUST clear or
743
+ suppress any cached `update_available` notice for that installed target before
744
+ later commands can attach `meta.notices`. An update command's own response must
745
+ not carry a stale notice that says the just-installed target is still available.
746
+ - When an update is available, the notice carries `type: "update_available"`,
747
+ `severity`, current/latest versions, install method, `recommended_command`,
748
+ release URL when known, checked-at timestamp, and machine-readable next steps.
749
+ It appears in active-check command `data` (`context` / `doctor` / `update
750
+ --check`) and, read-only from cache, in any command's `meta.notices`. Text/help
751
+ output may append one concise hint.
752
+ - **Severity grading** — computed at check time from the embedded CHANGELOG delta
753
+ between the running version and the latest, and stored in the cache so the
754
+ cached `meta.notices` carries the right level:
755
+ - `info` (default): routine patch/minor with no security entry.
756
+ - `warning`: the changelog delta since the running version contains a
757
+ `security` entry, OR the latest crosses a **major** version (first semver
758
+ component increased) — i.e. likely security-relevant or breaking.
759
+
760
+ Release verification baseline:
761
+
762
+ - **Mandatory signature verification, no skip path**: the binary self-update path
763
+ MUST verify the Sigstore signature on `checksums.txt` in-process, then verify
764
+ the archive SHA256 against it. A missing signature bundle, a signature that does
765
+ not verify, or a checksum mismatch all fail closed — there is no "can't verify,
766
+ proceed anyway" degradation. The whole chain surfaces `E_INTEGRITY` (exit 1,
767
+ non-retryable): a forged or corrupt release is not a transient blip to retry.
768
+ - **Verifier embedded, no user-environment dependency**: verification happens
769
+ inside the tool binary (Go via `sigstore-go`, Python inside the frozen binary
770
+ via `sigstore`) with **no external cosign** and nothing pre-installed on the
771
+ machine. The TUF trust root is bootstrapped from the library's embedded
772
+ `root.json`, not fetched on first-use trust (TOFU).
773
+ - **New bundle format**: the signing side produces a Sigstore protobuf bundle
774
+ (`checksums.txt.sigstore.json`) via `cosign sign-blob --new-bundle-format`, which
775
+ the in-process verifier consumes; the legacy cosign bundle format is not accepted.
776
+ - **Identity binding**: verifiers bind the certificate SAN to this repo's tagged
777
+ release workflow (`…/release.yml@refs/tags/v*`, anchored `^…$`) and validate the
778
+ GitHub OIDC issuer. When the target tag is known, pin the exact identity (stronger
779
+ than a regexp).
780
+ - **Cross-language parity**: Go binaries and Python frozen binaries follow the same
781
+ self-update contract — download archive → in-process signature verify → checksum
782
+ → replace binary. Package managers do not own integrity.
783
+ - Update results carry `signature_status` (`verified` on success; any failure exits
784
+ via the error envelope) and `signature_verified` (true only when in-process
785
+ Sigstore verification actually ran and succeeded). Never imply checksum
786
+ verification is a signature.
787
+
788
+ - After `update` succeeds, return `previous_version` and `current_version` in `data`.
789
+ - Also hint in the result: `run "changelog --since <previous_version>" to see what changed`.
790
+ - Agent convention: after self-update, before continuing, read `changelog --since <old version>` (see the SKILL-SPEC recipe).
791
+
792
+ Failure and interruption contract:
793
+
794
+ Single-command `update` runs as staged work — discover → download → verify
795
+ signature → verify checksum → replace binary → sync Skill — with exactly one
796
+ atomic commit point. The invariant that makes every failure message honest:
797
+
798
+ - Everything BEFORE the binary swap touches only a temp dir; any failure or
799
+ interruption there leaves the installed binary untouched and fully usable
800
+ (`current_version` unchanged, `binary_replaced: false`).
801
+ - The swap itself is atomic. On Windows a running executable can be renamed but
802
+ not overwritten, so the verified `<bin>.new` is moved into place by renaming the
803
+ in-use binary aside to `<bin>.old`, then renaming `<bin>.new` onto the real path
804
+ — the new binary is in place at once, `binary_replaced: true`, and the next
805
+ invocation runs the new version (no restart needed). The displaced `<bin>.old`
806
+ cannot be deleted while the old process is still running, so it is left in place
807
+ and cleaned up by the next `update`, which re-verifies any staged artifact from
808
+ scratch — a leftover is never trusted. On non-Windows the swap is a same-filesystem rename
809
+ of the verified binary into place. A crash mid-swap leaves either the old or the
810
+ new binary, never a hybrid.
811
+ - Skill sync runs AFTER the swap and is independently replayable.
812
+
813
+ So the tool can always determine — and MUST always report — its post-failure
814
+ state. Every update failure envelope carries, in `error.details` (or `data` on
815
+ partial success): `stage`
816
+ (`discover|download|verify_signature|verify_checksum|replace|skill_sync`),
817
+ `current_version`, `binary_replaced`, and `skill_sync_status`.
818
+
819
+ Classify the failure by the agent's next action, not by the raw cause:
820
+
821
+ | Stage | Failure | code / exit | retryable | Post-state | Message must say |
822
+ |-------|---------|-------------|-----------|------------|------------------|
823
+ | discover / download | network / timeout / rate-limit | `E_NETWORK` / `E_TIMEOUT` / `E_RATE_LIMITED` → 7,8 | true | old version, no change | "transient — re-run `update`, it is idempotent" |
824
+ | verify_signature / verify_checksum | missing/invalid signature, identity mismatch, checksum mismatch | `E_INTEGRITY` → 1 | **false** | old version, install refused | "integrity failure — do NOT retry, stop and report" |
825
+ | replace | permission / disk full / file locked | `E_FORBIDDEN` / `E_CONFIG` / `E_IO` → 4,1 | false (needs fix) | old version (atomic not committed) | the concrete fix, then re-run |
826
+ | skill_sync (post-swap) | npx missing / network | partial success (`ok:false`, `binary_replaced:true`) | true | binary NEW, Skill OLD | "binary at vX; run `<skill_sync_command>`, then `changelog --since <prev>`" |
827
+ | any | user/signal interrupt (SIGINT) | `E_INTERRUPTED` → 130 | true | per the stage invariant above | what actually happened + the safe next step |
828
+
829
+ Interruption (Ctrl-C / SIGTERM):
830
+
831
+ - Trap the signal, unwind the current stage to a clean state, and STILL emit the
832
+ terminal JSON envelope on stdout before exiting non-zero — an interrupted agent
833
+ must receive a parseable terminal state, never a bare killed process.
834
+ - Always clean the temp dir on interrupt; a partial download must never be
835
+ trusted by a later run (re-download + re-verify always).
836
+ - The message depends on the interrupted stage: before the swap → "cancelled, no
837
+ change, still on `<current>`"; during the atomic swap → report old-or-new
838
+ truthfully; after the swap during Skill sync → partial success with
839
+ `skill_sync_command`.
840
+
841
+ Three rules the messages must never break:
842
+
843
+ 1. Never misstate the version: every terminal envelope states the version the
844
+ tool is actually running now.
845
+ 2. Never let an agent retry an integrity failure: `E_INTEGRITY` is
846
+ `retryable: false` and verbally distinct from any network failure — a forged
847
+ release is not a transient blip to loop on.
848
+ 3. Never call a partial a success: binary replaced but Skill not synced is
849
+ partial success with `skill_sync_command`, not `ok: true`.
850
+
851
+ ## 15. Batch operations
852
+
853
+ Many write workflows need to act on a batch of objects in one call (close many issues, send to many openids, run one SQL across many instances). A batch command is still **one** agent-facing command with **one** envelope, **one** confirm token, and **one** aggregated result — never a loop the agent has to drive. The contract below is identical whether the batch is served by a native upstream bulk endpoint (class A) or by a client-side loop (class B); the agent must not be able to tell which.
854
+
855
+ ### 15.1 Plural inputs
856
+
857
+ - Batch targets use a plural flag: `--ids`, `--symbols`, `--instances`, `--openids`, etc.
858
+ - Each plural flag accepts both **comma-separated** (`--ids 1,2,3`) and **repeatable** (`--ids 1 --ids 2 --ids 3`) forms; the two are equivalent and may be mixed.
859
+ - A single value degrades gracefully: `--ids 1` is a valid batch of one, same envelope as a batch of many. Where a singular flag (`--symbol`) already exists, keep it as a compatibility alias of the plural and declare it deprecated in reference; do not run two divergent code paths.
860
+ - De-duplicate targets before executing and preserve input order in the result `items[]` so the agent can zip results back to inputs deterministically.
861
+ - An empty target list is a usage error (`E_VALIDATION`, exit 2), not a silent no-op.
862
+
863
+ ### 15.2 Dry-run summary for a batch
864
+
865
+ `--dry-run` on a batch returns **what will happen to N objects** before any write, plus a single `confirm_token` that covers the whole batch:
866
+
867
+ ```json
868
+ {
869
+ "ok": true,
870
+ "schema_version": "1.0",
871
+ "data": {
872
+ "preview": {
873
+ "action": "close",
874
+ "total": 3,
875
+ "targets": ["1042", "1043", "1044"],
876
+ "changes": [
877
+ { "action": "close", "resource": "issue", "id": "1042" },
878
+ { "action": "close", "resource": "issue", "id": "1043" },
879
+ { "action": "close", "resource": "issue", "id": "1044" }
880
+ ]
881
+ },
882
+ "confirm_token": "ct_9f2a...",
883
+ "expires_at": "2026-06-15T12:00:00Z"
884
+ },
885
+ "meta": { "duration_ms": 0 }
886
+ }
887
+ ```
888
+
889
+ - The preview must state the operation and the full target set, so a human or agent can audit the blast radius before confirming.
890
+ - The token binds the **whole resolved target set** (plus command path, args, account, permission context per §7), so adding or removing a target invalidates it with `E_CONFLICT`.
891
+
892
+ ### 15.3 One confirm token covers the whole batch, consumed once
893
+
894
+ - A single `--confirm <token>` from the batch dry-run authorizes the entire batch; the agent does not confirm per item.
895
+ - The token is **single-use** exactly as in §9: it is fingerprinted and marked consumed before the write runs, and any replay is rejected with `E_CONFLICT` ("token already used; re-run `--dry-run`"). This reuses each repo's existing single-consumption confirm infrastructure — batch adds no new token mechanism.
896
+ - On a partial batch failure the token stays consumed; the agent re-runs `--dry-run` (which now resolves to the still-pending targets) rather than replaying the old token.
897
+
898
+ ### 15.4 Dangerous batches: `--dangerous` two-step gate
899
+
900
+ Irreversible or high-blast-radius batches — bulk `delete`, MR `merge`, mass send / broadcast — require an extra gate **on top of** dry-run → confirm:
901
+
902
+ - The command must be invoked with `--dangerous`; without it the command fails with `E_CONFIRMATION_REQUIRED` (exit 5) even when a valid confirm token is present.
903
+ - This is a two-step human-intent gate: `--dangerous` declares intent, the confirm token authorizes the specific resolved batch. Both are required; neither alone executes.
904
+ - Reference marks these commands `dangerous: true`, and their `examples[]` show the `--dangerous` form.
905
+ - Tools may layer stricter local policy on a dangerous batch (e.g. per-item confirm, default `--continue-on-error false`, or night-time dry-run-only); such overrides must be declared in reference, not hidden.
906
+
907
+ ### 15.5 Per-item aggregated result, no whole-batch rollback
908
+
909
+ Batch results aggregate per item. A partial failure does **not** roll back succeeded items:
910
+
911
+ ```json
912
+ {
913
+ "ok": true,
914
+ "schema_version": "1.0",
915
+ "data": {
916
+ "items": [
917
+ { "target": "1042", "ok": true },
918
+ { "target": "1043", "ok": true },
919
+ {
920
+ "target": "1044",
921
+ "ok": false,
922
+ "error": { "code": "E_NOT_FOUND", "retryable": false }
923
+ }
924
+ ],
925
+ "summary": { "total": 3, "succeeded": 2, "failed": 1 }
926
+ },
927
+ "meta": { "duration_ms": 0 }
928
+ }
929
+ ```
930
+
931
+ - Each `items[]` entry carries `target` (the input identifier — `id`, `symbol`, `instance`, …; use the natural key, not an array index), `ok`, and on failure `error` with the same `{ code, retryable }` taxonomy as the top-level envelope (§6).
932
+ - `summary` always reports `{ total, succeeded, failed }`. The counts must equal the item tally.
933
+ - Top-level `ok` is `true` when the batch executed and produced a result, even with per-item failures; per-item status lives in `items[]`. Reserve top-level `ok: false` for a batch that could not run at all (bad args, auth, no targets). Do not hide other items' status because one failed (consistent with §9).
934
+ - `--continue-on-error` controls whether the batch keeps going after the first item failure; **default `true`** (best-effort, finish the batch). Set `--continue-on-error false` to stop at the first failure — already-applied items stay applied (no rollback), and `summary` reflects only attempted items, with the unattempted remainder reported (e.g. `skipped`) so the agent can resume. Dangerous batches may flip the default to `false`; declare it in reference.
935
+
936
+ ### 15.6 Upstream caps: client-side auto-chunking
937
+
938
+ When the native bulk endpoint has a per-call limit, the command splits the batch into chunks and submits them sequentially, presenting a single command to the agent:
939
+
940
+ - Known caps: Jira `/issue/bulk`, agile `backlog/issue` and `sprint/{id}/issue` ≤ 50; WeChat openid batches ≤ 100, mass-send openid lists per upstream limit. The command must chunk to the cap; it must not pass a too-large batch straight through and let the upstream 400.
941
+ - Chunking is invisible in the contract: still one envelope, one `confirm_token`, one aggregated `items[]`/`summary` across all chunks.
942
+ - A chunk-level failure is mapped back onto the affected `items[]`; one failed chunk does not fail the whole command (subject to `--continue-on-error`).
943
+ - Keep the chunk size in one shared helper per tool so the cap cannot drift between commands sharing an endpoint.
944
+
945
+ ### 15.7 A/B classes, one external contract
946
+
947
+ - **Class A** uses a native upstream bulk endpoint (true server-side batch, possibly atomic per the upstream).
948
+ - **Class B** loops client-side over single-target calls because no native bulk exists.
949
+ - The external contract is identical for both: plural inputs, dry-run summary, single confirm token, dangerous gate, aggregated `items[]`/`summary`, and `--continue-on-error`. The agent cannot and need not tell A from B.
950
+ - Atomicity is **not** part of the external contract. A class-B (or capped, chunked class-A) batch must not claim upstream atomicity; where the upstream genuinely is non-atomic or order-unstable, say so in the output schema / reference rather than implying a transaction.
951
+
952
+ ### 15.8 Self-description for batch commands
953
+
954
+ Every batch command carries a real `output_schema` and runnable `examples[]` (per §11):
955
+
956
+ - The schema declares the `items[]` shape (`target`, `ok`, `error{code,retryable}`) and the `summary{total,succeeded,failed}` shape, with attacker-controllable keys listed in `untrusted_fields` (`_untrusted`).
957
+ - `examples[]` show the plural-input dry-run then confirm pair; dangerous batches include `--dangerous`.
958
+ - These must pass the reference guard (non-empty schema + at least one example per leaf command) and count toward FCC (§13) like any other public behavior.
959
+
960
+ ## 16. Optional patterns (enable as needed)
961
+
962
+ These three patterns are **not for everyone**: implement them if your tool needs them, ignore them otherwise — zero overhead. They let the spec scale with tool complexity — a simple tool stays light, a complex tool need not reinvent the wheel. Each is marked "when applicable."
963
+
964
+ ### 16.1 Credential lifecycle (when tokens expire)
965
+
966
+ **When applicable**: credentials are not static but expire / need refresh — OAuth access_token (WeChat Official Account ~2h), cookie / session (Xiaohongshu), temporary STS credentials, etc. Tools with static username/password skip this section.
967
+
968
+ - Beyond "is it configured," `context.data.credentials` should report **validity and expiry** (redacted):
969
+
970
+ ```json
971
+ {
972
+ "credentials": {
973
+ "configured": true,
974
+ "valid": true,
975
+ "expires_at": "2026-06-07T12:00:00Z",
976
+ "refreshable": true
977
+ }
978
+ }
979
+ ```
980
+
981
+ - When a token is expired and cannot auto-refresh, the operation returns `E_AUTH` (exit 4), with `details` indicating re-auth is needed.
982
+ - Tools that can auto-refresh should do so **transparently**, not bothering the agent; degrade to `E_AUTH` only if refresh fails.
983
+ - `doctor` adds a `check: "credentials"` item; for near-expiry give `warn` + a renew `fix`.
984
+ - Refresh tokens and secrets are always redacted — never in stdout / stderr / details.
985
+
986
+ ### 16.2 Async job lifecycle (long jobs: submit -> poll -> fetch result)
987
+
988
+ **When applicable**: the operation can't return a result synchronously — async SQL execution / approval (Archery), bulk send, scrape/crawl jobs, large exports. Commands that return results synchronously skip this section.
989
+
990
+ - The submit command returns a `job_id` and status immediately, without blocking:
991
+
992
+ ```json
993
+ {
994
+ "ok": true,
995
+ "schema_version": "1.0",
996
+ "data": {
997
+ "job_id": "job_abc123",
998
+ "status": "pending",
999
+ "poll": "tool job status --id job_abc123",
1000
+ "result": "tool job result --id job_abc123"
1001
+ },
1002
+ "meta": { "duration_ms": 12 }
1003
+ }
1004
+ ```
1005
+
1006
+ - Status queries return a stable enum: `pending` / `running` / `succeeded` / `failed` / `cancelled`, with progress (e.g. `progress`, `eta_seconds`).
1007
+ - Result and status are fetched separately: after `succeeded`, use the `result` command to pull data (large results via NDJSON / `--format raw`).
1008
+ - A `failed` result uses the standard error envelope; `retryable` indicates whether the whole job can be retried.
1009
+ - Submission of a write-type long job still goes through `dry-run → confirm`; the `job_id` is created only after confirm.
1010
+
1011
+ ### 16.3 Human-in-the-loop checkpoints (when a human must scan / solve captcha / approve)
1012
+
1013
+ **When applicable**: a step mid-flow must be completed by a human — QR login / captcha (Xiaohongshu), approver sign-off (Archery), secondary confirmation. Fully automated tools skip this section.
1014
+
1015
+ - When stuck at a human step, **don't block, don't guess** — return a dedicated signal so the agent hands off to the user:
1016
+
1017
+ ```json
1018
+ {
1019
+ "ok": false,
1020
+ "schema_version": "1.0",
1021
+ "error": {
1022
+ "code": "E_HUMAN_REQUIRED",
1023
+ "message": "Scan the QR code to continue",
1024
+ "details": { "action": "scan_qr", "resume": "tool login resume --id sess_1", "qr_path": "/tmp/qr.png" },
1025
+ "retryable": false
1026
+ },
1027
+ "meta": { "duration_ms": 30 }
1028
+ }
1029
+ ```
1030
+
1031
+ - `E_HUMAN_REQUIRED` uses exit code `9` (added beyond the existing 0–8; not reusing `4`, to distinguish "bad credentials" from "waiting on a human action").
1032
+ - `details.action` is a stable enum describing what the human must do; `details.resume` gives the command to continue after the human is done.
1033
+ - Agent convention: on `E_HUMAN_REQUIRED` → relay `message` and the required action to the user → wait for them → run `resume`; do not auto-retry.
1034
+
1035
+ ## 17. Design checklist
1036
+
1037
+ > Items marked `(optional)` only apply when the corresponding optional pattern is enabled.
1038
+
1039
+ - [ ] Default `--format json`
1040
+ - [ ] stdout contains only valid JSON / NDJSON, no pollution
1041
+ - [ ] Logs and progress all go to stderr
1042
+ - [ ] Success/failure share one envelope, with `ok` and `schema_version`
1043
+ - [ ] `error` has semantic `code`, `details`, `retryable`
1044
+ - [ ] Exit codes tiered and consistent with `retryable`
1045
+ - [ ] Write commands have the dry-run / confirm-token loop
1046
+ - [ ] Confirm token binds operation args, account, permission context, resource version
1047
+ - [ ] Provides `reference` / `context` / `doctor`
1048
+ - [ ] Provides `changelog [--since]`, same source as CHANGELOG/release-notes
1049
+ - [ ] Tool reports its own version (`--version` and `context.version`)
1050
+ - [ ] `reference` reports `release_readiness`, and `doctor` checks it
1051
+ - [ ] (with self-update) `update` is a single command (no leaf subcommands, no confirm token); `--check` / `--dry-run` are optional read-only
1052
+ - [ ] (with self-update) release integrity is verified, and signature status is explicit
1053
+ - [ ] (with self-update) syncing every Skill directory under `skills/` is part of the update result
1054
+ - [ ] (with self-update) post-update returns previous/current version and hints to read changelog
1055
+ - [ ] (with self-update) every update failure/interruption envelope carries `stage` + `current_version` + `binary_replaced` + `skill_sync_status`; `E_INTEGRITY` is non-retryable; binary-replaced-but-Skill-unsynced is partial success, not `ok`
1056
+ - [ ] (with self-update) SIGINT/SIGTERM is trapped, leaves nothing half-applied, and still emits the terminal JSON envelope
1057
+ - [ ] Query commands support `fields` / `compact`
1058
+ - [ ] List commands support pagination or explicitly state none is needed
1059
+ - [ ] Batch commands take plural inputs (`--ids`/`--symbols`/…), comma-separated or repeatable, single value degrades
1060
+ - [ ] Batch dry-run summarizes the full target set; one confirm token covers and is consumed once for the whole batch
1061
+ - [ ] Dangerous batches (delete / merge / mass-send) require the `--dangerous` two-step gate
1062
+ - [ ] Batch results aggregate `items[].{target,ok,error{code,retryable}}` + `summary{total,succeeded,failed}`, no whole-batch rollback, `--continue-on-error` default true
1063
+ - [ ] Capped upstream bulk auto-chunks client-side under one command; A/B classes share one external contract; batch commands ship real `output_schema` + `examples`
1064
+ - [ ] Functional Contract Coverage is 100% for public README / Skill / reference / help / context / doctor / changelog / update behavior
1065
+ - [ ] Stable releases have recorded live smoke/E2E evidence; otherwise the tool declares `beta`
1066
+ - [ ] All times ISO 8601 UTC
1067
+ - [ ] All IDs strings
1068
+ - [ ] Secrets redacted end to end
1069
+ - [ ] Schema changes have a versioning/compat policy
1070
+ - [ ] stdout/stderr are UTF-8 without BOM
1071
+ - [ ] (optional · expiring tokens) `context`/`doctor` report credential validity and expiry; refresh failure degrades to `E_AUTH`
1072
+ - [ ] (optional · long jobs) submit returns `job_id` + status enum, status/result separated
1073
+ - [ ] (optional · human needed) stuck human steps return `E_HUMAN_REQUIRED` (exit 9) + `resume`, no auto-retry