@heihei0299/matt-skills 1.3.1 → 1.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/.agents/skills/ask-matt/PHASE-BOUNDARIES.md +55 -0
  2. package/.agents/skills/ask-matt/SKILL.md +37 -25
  3. package/.agents/skills/ci-guard/SKILL.md +104 -0
  4. package/.agents/skills/ci-guard/agents/openai.yaml +5 -0
  5. package/.agents/skills/code-review/SKILL.md +28 -35
  6. package/.agents/skills/codebase-design/DEEPENING.md +4 -4
  7. package/.agents/skills/codebase-design/DESIGN-IT-TWICE.md +10 -10
  8. package/.agents/skills/codebase-design/SKILL.md +13 -13
  9. package/.agents/skills/diagnosing-bugs/SKILL.md +34 -30
  10. package/.agents/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +3 -0
  11. package/.agents/skills/domain-modeling/ADR-FORMAT.md +11 -11
  12. package/.agents/skills/domain-modeling/CONTEXT-FORMAT.md +3 -3
  13. package/.agents/skills/domain-modeling/SKILL.md +10 -10
  14. package/.agents/skills/grill-me/SKILL.md +1 -1
  15. package/.agents/skills/grill-with-docs/SKILL.md +1 -1
  16. package/.agents/skills/grilling/SKILL.md +20 -4
  17. package/.agents/skills/grilling/agents/openai.yaml +1 -1
  18. package/.agents/skills/handoff/SKILL.md +1 -1
  19. package/.agents/skills/improve-codebase-architecture/HTML-REPORT.md +19 -19
  20. package/.agents/skills/improve-codebase-architecture/SKILL.md +21 -21
  21. package/.agents/skills/instance-test/SKILL.md +36 -27
  22. package/.agents/skills/instance-test/agents/openai.yaml +1 -1
  23. package/.agents/skills/instance-test/references/instances.md +63 -36
  24. package/.agents/skills/prototype/LOGIC.md +30 -42
  25. package/.agents/skills/prototype/SKILL.md +7 -7
  26. package/.agents/skills/prototype/UI.md +23 -23
  27. package/.agents/skills/research/SKILL.md +1 -1
  28. package/.agents/skills/resolving-merge-conflicts/SKILL.md +1 -1
  29. package/.agents/skills/scaffold-functional-test/SKILL.md +77 -0
  30. package/.agents/skills/scaffold-functional-test/agents/openai.yaml +5 -0
  31. package/.agents/skills/setup-matt-pocock-skills/SKILL.md +30 -30
  32. package/.agents/skills/setup-matt-pocock-skills/domain.md +4 -4
  33. package/.agents/skills/setup-matt-pocock-skills/issue-tracker-github.md +5 -5
  34. package/.agents/skills/setup-matt-pocock-skills/issue-tracker-gitlab.md +6 -6
  35. package/.agents/skills/setup-matt-pocock-skills/issue-tracker-local.md +3 -3
  36. package/.agents/skills/tdd/SKILL.md +9 -7
  37. package/.agents/skills/teach/GLOSSARY-FORMAT.md +3 -3
  38. package/.agents/skills/teach/LEARNING-RECORD-FORMAT.md +10 -10
  39. package/.agents/skills/teach/MISSION-FORMAT.md +4 -4
  40. package/.agents/skills/teach/RESOURCES-FORMAT.md +2 -2
  41. package/.agents/skills/teach/SKILL.md +4 -4
  42. package/.agents/skills/to-questionnaire/SKILL.md +54 -0
  43. package/.agents/skills/to-questionnaire/agents/openai.yaml +5 -0
  44. package/.agents/skills/to-spec/SKILL.md +4 -4
  45. package/.agents/skills/to-tickets/SKILL.md +16 -16
  46. package/.agents/skills/triage/AGENT-BRIEF.md +9 -9
  47. package/.agents/skills/triage/OUT-OF-SCOPE.md +15 -15
  48. package/.agents/skills/triage/SKILL.md +29 -29
  49. package/.agents/skills/wait-what/SKILL.md +7 -0
  50. package/.agents/skills/wait-what/agents/openai.yaml +5 -0
  51. package/.agents/skills/wayfinder/SKILL.md +37 -37
  52. package/.agents/skills/wizard/SKILL.md +44 -0
  53. package/.agents/skills/wizard/agents/openai.yaml +3 -0
  54. package/.agents/skills/wizard/template.sh +204 -0
  55. package/.agents/skills/writing-for-agents/SKILL-MECHANICS.md +22 -0
  56. package/.agents/skills/writing-for-agents/SKILL.md +81 -0
  57. package/.agents/skills/writing-for-agents/agents/openai.yaml +3 -0
  58. package/README.md +9 -9
  59. package/bin/cli.js +1 -1
  60. package/config/proprietary.json +8 -1
  61. package/package.json +1 -1
  62. package/scripts/sync-upstream.js +1 -1
  63. package/template/.opencode/CONTEXT.md +2 -2
  64. package/template/.opencode/commands/{writing-great-skills.md → writing-for-agents.md} +1 -1
  65. package/template/.opencode/docs/agents/skill-design.md +3 -3
  66. package/template/.opencode/skills/ci-guard/SKILL.md +104 -0
  67. package/template/.opencode/skills/ci-guard/agents/openai.yaml +5 -0
  68. package/template/.opencode/skills/scaffold-functional-test/SKILL.md +77 -0
  69. package/template/.opencode/skills/scaffold-functional-test/agents/openai.yaml +5 -0
  70. package/template/.pi/CONTEXT.md +55 -0
  71. package/template/.pi/docs/agents/runtime-discipline.md +3 -2
  72. package/template/.pi/docs/agents/skill-design.md +10 -5
  73. package/template/.pi/skills/ci-guard/SKILL.md +104 -0
  74. package/template/.pi/skills/ci-guard/agents/openai.yaml +5 -0
  75. package/template/.pi/skills/scaffold-functional-test/SKILL.md +77 -0
  76. package/template/.pi/skills/scaffold-functional-test/agents/openai.yaml +5 -0
  77. package/template/AGENTS.md +3 -4
  78. package/.agents/skills/writing-great-skills/GLOSSARY.md +0 -201
  79. package/.agents/skills/writing-great-skills/SKILL.md +0 -83
  80. package/.agents/skills/writing-great-skills/agents/openai.yaml +0 -5
  81. package/template/.opencode/skills/instance-test/SKILL.md +0 -61
  82. package/template/.opencode/skills/instance-test/agents/openai.yaml +0 -5
  83. package/template/.opencode/skills/instance-test/references/instances.md +0 -48
  84. package/template/.pi/skills/instance-test/SKILL.md +0 -61
  85. package/template/.pi/skills/instance-test/agents/openai.yaml +0 -5
  86. package/template/.pi/skills/instance-test/references/instances.md +0 -48
@@ -1,61 +1,70 @@
1
1
  ---
2
2
  name: instance-test
3
3
  disable-model-invocation: true
4
- description: "Verify project meets expected goals by running prompt instances in isolated temp dirs"
4
+ description: "matt-skills 专属功能测试示范(由 scaffold-functional-test 从 spec 生成)— 验证 sync 合并 update 后的行为;仅显式调用"
5
5
  ---
6
6
 
7
- # Instance Test
7
+ # Instance Test — matt-skills 专属示范
8
8
 
9
- Run **instance** prompts to verify project meets expected goals via actual functional tests. Each **instance** is a prompt + expected outcome, executed in an isolated temp dir — no mocks, no stubs.
9
+ 本 skill 是 **matt-skills 专属**的功能测试示范,由 `scaffold-functional-test` 从 `.scratch/sync-merge-update/spec.md` 生成(见 `references/instances.md` 头部 `spec hash` + `generatedAt`)。它是生成器产出形态的示例,不随 Template Snapshot 分发,仅保留于 Workspace。旧通用执行器文案已废弃。
10
+
11
+ 兼容别名:`instance-test` 保留原名以兼容历史调用,实际为 `matt-functional-test` 的示范实现。
10
12
 
11
13
  ## Steps
12
14
 
13
15
  ### 1. Gather instances
14
16
 
15
- Collect the **instance** set to run:
17
+ 实例集已由生成器按**受控扩展模型**落盘于 `references/instances.md`(头部含 `spec hash` + `generatedAt`,每实例含**溯源** `spec.md` 章节/行号,`<!-- manual -->` 段受保护)。
16
18
 
17
- - User-provided instances (prompt, command, expected files/stdout/exit code), or
18
- - Derived from `spec.md`/`README` acceptance criteria — extract each verifiable behavior as one instance, then confirm the list with the user before running.
19
+ 执行前校验指纹:若当前 spec 的 `spec hash` 与 `references/instances.md` 头部不一致,提示「spec 已变更,建议重跑 scaffold-functional-test」但不自动覆盖,需用户显式确认才 regenerate(AI 先给 diff 建议)。
19
20
 
20
- Each **instance** must declare: command to run, expected files/content, expected stdout phrases, expected exit code.
21
+ 每实例声明:`prompt/command/expected files/content/expected stdout phrases/expected exit code` 必选,`setup/env/timeout/type/teardown` 可选,默认 `type: cli`。
21
22
 
22
- Completion: instance list is fixed (prompt, expected outcome, verification command) — no instance is added mid-run.
23
+ 完成:实例清单已固定(含溯源与指纹),`<!-- manual -->` 段未被覆盖。
23
24
 
24
25
  ### 2. Run instances
25
26
 
26
27
  For each **instance** in order:
27
28
 
28
- 1. `mktemp -d` isolated dir (or `git worktree` / `--dest` if the project supports it).
29
- 2. Execute the instance's command — capture stdout/stderr and exit code.
30
- 3. Snapshot result files and side effects declared in expected.
29
+ 1. `mktemp -d` 隔离目录(或项目支持的 `git worktree` / `--dest`),单线程串行,不并行。
30
+ 2. 执行实例的 `command` 与可选 `setup`,捕获 stdout/stderr 与 exit code。
31
+ 3. 快照 `expected` 声明的文件与副作用。
31
32
 
32
- Do not run instances in parallel — one **instance** at a time, so failures are isolated and artifacts do not collide.
33
+ 一个 **instance** 一次,失败不阻断后续,产物不碰撞。
33
34
 
34
- Completion: every **instance** has a run dir with captured output and file snapshot.
35
+ 完成:每实例均有独立 run dir 与捕获输出。
35
36
 
36
37
  ### 3. Evaluate
37
38
 
38
- Compare each **instance**'s actual vs expected:
39
+ 对比每实例的 actual vs expected:
39
40
 
40
- - File existence/content (`test -f`, `grep -q`, `diff`).
41
- - Stdout/stderr contains expected phrases.
42
- - Exit code matches expected.
41
+ - 文件存在性/内容(`test -f`/`grep -q`/`diff`)
42
+ - Stdout/stderr 含预期短语
43
+ - Exit code 一致
44
+ - 扩展字段(`env`/`timeout`/`type`)行为符合声明
43
45
 
44
- Mark `PASS`/`FAIL` per **instance** with evidence (file path, stdout line, or diff).
46
+ 标记 `PASS`/`FAIL`,附 `expected vs actual` diff 与 run dir 证据。
45
47
 
46
- Completion: every **instance** has a `PASS` or `FAIL` with evidence — no unevaluated instance.
48
+ 完成:每实例均有 `PASS` 或 `FAIL` 且含证据。
47
49
 
48
50
  ### 4. Report
49
51
 
50
- Summarize in conversation:
52
+ 对话内汇总:
53
+
54
+ - `PASS m/n` + per-instance evidence
55
+ - 失败项列出 gap(expected vs actual)与 run dir 复现路径
56
+ - 成功默认清理临时目录、失败默认保留;`--keep` 保留全部;`--report` 显式开启才落盘报告文件
57
+
58
+ 不以文件刷屏——默认输出在对话,报告文件仅显式开启才写。
51
59
 
52
- - `PASS m/n` with per-instance evidence.
53
- - Failures list the gap (expected vs actual) and the run dir for reproduction.
54
- - Clean up temp dirs unless `--keep` is requested.
60
+ ## 实例来源
55
61
 
56
- Do not write a report file (`report-*.md`) — output stays in conversation. Keep temp dirs only on failure for debugging.
62
+ - 源 spec:`.scratch/sync-merge-update/spec.md`(`spec hash` 见 `references/instances.md` 头部)
63
+ - 推导策略:混合推导(验收标准锚点 + 需求/接口/边界补充),每实例含溯源,无溯源视为幻觉
64
+ - 手工段:`<!-- manual -->` 保护
57
65
 
58
- ## References
66
+ ## 引用
59
67
 
60
- - Instance definitions (if any): `references/instances.md` — example set, auto-loaded only when present, not required.
61
- - Project expected behavior: `spec.md`/`README`/`--help` — the source of truth for what to verify.
68
+ - 生成器:`scaffold-functional-test`(读 spec 产出本 skill)
69
+ - 领域术语:`CONTEXT.md`
70
+ - 技能设计:`docs/agents/skill-design.md`
@@ -1,5 +1,5 @@
1
1
  interface:
2
2
  display_name: "Instance Test"
3
- short_description: "Verify project via prompt instances in isolated temp projects"
3
+ short_description: "matt-skills 专属功能测试示范 — 验证 sync 合并 update 后行为"
4
4
  policy:
5
5
  allow_implicit_invocation: false
@@ -1,48 +1,75 @@
1
- # Instances template
1
+ # Instances for matt-skills — sync 行为功能测试(由 scaffold-functional-test 生成)
2
2
 
3
- Generic template for **instance** functional tests. Each **instance** is a prompt + command + expected outcome. Copy and adapt for your project; the example below is for `matt-skills`.
3
+ > 源 spec:`.scratch/sync-merge-update/spec.md`
4
+ > spec hash: `062a76fc872d` # .scratch/sync-merge-update/spec.md 的 sha256 前 12 位
5
+ > generatedAt: 2026-05-11
6
+ > 推导策略:混合推导(验收标准锚点 + 需求/行为补充),每实例含溯源,无溯源视为幻觉
4
7
 
5
- ## Format
8
+ 本文件由 `scaffold-functional-test` 按**受控扩展模型**生成:必选 `prompt/command/expected files/content/expected stdout phrases/expected exit code`,可选 `setup/env/timeout/type/teardown`,默认 `type: cli`。执行语义:`mktemp -d` 隔离、单线程串行、`PASS m/n` 汇总、证据含 `expected vs actual` diff + `run dir`。
6
9
 
7
- Each instance declares:
10
+ 执行前校验:对比当前 `.scratch/sync-merge-update/spec.md` 的 hash 与本文件头部 `spec hash`,不一致时提示「spec 已变更,建议重跑 scaffold-functional-test」但不自动覆盖。
8
11
 
9
- - Prompt: human intent (what to verify)
10
- - Command: shell command to run in isolated dir
11
- - Expected: files/content, stdout phrases, exit code
12
+ ---
12
13
 
13
- Verification commands are in `SKILL.md` steps.
14
+ ## 1. sync 默认 check(无参不写盘)
14
15
 
15
- ## Example: matt-skills functional behavior
16
+ - Prompt: 验证 `matt-skills sync` 无参等价 check,打印表且不改 AGENTS.md
17
+ - 溯源: spec.md — 需求/行为「`matt-skills sync` 无参:等价 `check`」+ 验收标准「`sync` 无参在已定制的 `pi-switch` 仓库上不改 `AGENTS.md`」
18
+ - type: cli
19
+ - setup: `node bin/cli.js init --dest <tmp>` 后手工改 `AGENTS.md` 加入 `tdd-implement` 定制行
20
+ - Command: `node bin/cli.js sync --dest <tmp>`(无参)
21
+ - Expected:
22
+ - `git diff HEAD -- AGENTS.md` 为空(`AGENTS.md` 未被覆盖)
23
+ - stdout 含 `上游 HEAD` 与 `新增/更新/删除/一致` 表头
24
+ - stdout 含 `--json` 可解析提示或表格行
25
+ - exit 0 或 1(有差异时 exit 1,判 exit code 符合 check 语义)
26
+ - Expected files/content: `AGENTS.md` 保留定制行,无 `AGENTS.md.bak` 新增
27
+ - Expected stdout phrases: `上游 HEAD`, `一致`
28
+ - Expected exit code: 1(有差异时)/ 0(无差异时)— 按实现定义,测试以实际 check 语义为准
16
29
 
17
- ### 1. Fresh init
18
- Prompt: verify fresh project initialization
19
- Command: `node bin/cli.js init --dest <tmp>`
20
- Expected: `AGENTS.md`, `.opencode/skills/tdd-implement/SKILL.md`, `.pi/skills/tdd-implement/SKILL.md`, `.agents/skills/tdd` (22 upstream) exist; stdout `模板:已复制` + `上游技能:已装 22`; no `.bak`; exit 0.
30
+ ## 2. sync --apply 安全增量(AGENTS.md 跳过、上游强制覆盖不删)
21
31
 
22
- ### 2. Init skip on existing
23
- Prompt: verify idempotent init without --force
24
- Command: `init` twice without `--force`, second with local edit to `AGENTS.md`
25
- Expected: second stdout `模板已存在.*跳过`, `上游技能:已装 0、跳过 22`; local edit preserved; exit 0.
32
+ - Prompt: 验证 `sync --apply` 为安全增量,`AGENTS.md` 定制跳过、上游技能被覆盖但 remove 列表不删
33
+ - 溯源: spec.md — 需求/行为「`sync --apply`:安全增量写盘。`AGENTS.md` 若含独有路由则跳过;上游技能 `rm+cp force` 覆盖,跳过 `PROPRIETARY`,不执行 `remove`」+ 验收标准「`sync --apply` 后上游技能被强制更新为上游 `HEAD`,`remove` 列表的技能仍保留」
34
+ - type: cli
35
+ - setup: 在 `<tmp>` 放置旧版上游技能 `test-skill` 过期文件,并手工改 `AGENTS.md`
36
+ - Command: `node bin/cli.js sync --apply --dest <tmp>`
37
+ - Expected:
38
+ - `AGENTS.md` 仍含定制行(未被模板覆盖)
39
+ - 上游技能文件已更新为上游 HEAD 内容(`diff` 无旧版残留)
40
+ - `remove` 列表中的技能目录仍存在(未被删除)
41
+ - Expected files/content: `AGENTS.md` 定制行存在;`test-skill` 被覆盖为新版;无 `AGENTS.md.bak`(安全档不备份)或按实现保留但不覆盖
42
+ - Expected stdout phrases: `已同步` 或 `已更新` 或 `同步`
43
+ - Expected exit code: 0
44
+ - timeout: 30000
26
45
 
27
- ### 3. Init --force with direct overwrite
28
- Prompt: verify forced init directly overwrites
29
- Command: `init --force --dest <tmp>` after local edit
30
- Expected: stdout `已覆盖` + `已装 22`; `AGENTS.md.bak` exists with local edit; `AGENTS.md` restored from template; no `.agents/skills/*.bak`; no `.opencode.bak`/`.pi.bak`; exit 0.
46
+ ## 3. sync --force 硬盖(AGENTS.md 备份后覆盖、全量 add/update/remove)
31
47
 
32
- ### 4. Sync on existing
33
- Prompt: verify sync directly updates existing project
34
- Command: `sync --dest <tmp>` after local edit
35
- Expected: stdout `同步` + `已同步`/`已更新`; `AGENTS.md.bak` exists; `.agents/skills/tdd` updated; exit 0.
36
- Prompt: verify sync backs up existing project
37
- Command: `sync --dest <tmp>` after local edit
38
- Expected: stdout `同步` + `已备份`; `AGENTS.md.bak` exists; exit 0.
48
+ - Prompt: 验证 `sync --force` 硬盖,`AGENTS.md` 备份后被模板覆盖、技能与模板全量同步含删除
49
+ - 溯源: spec.md — 需求/行为「`sync --force`:硬盖。`AGENTS.md` 先 `backupIfExists → .bak` 再 `cp -r force`;技能与模板均 `add/update/remove` 全做」+ 验收标准「`sync --force` 后 `AGENTS.md` 变为模板且 `AGENTS.md.bak` 存在,`remove` 列表的技能被删除」
50
+ - type: cli
51
+ - setup: 在 `<tmp>` 放置 `AGENTS.md` 定制行 + 一个上游已删的本地技能 `obsolete-skill/`
52
+ - Command: `node bin/cli.js sync --force --dest <tmp>`
53
+ - Expected:
54
+ - `AGENTS.md` 已被模板覆盖(定制行消失,与 `template/AGENTS.md` 一致)
55
+ - `AGENTS.md.bak` 存在且含定制行备份
56
+ - `obsolete-skill/` 已被删除
57
+ - Expected files/content: `AGENTS.md` 内容等于 `template/AGENTS.md`;`AGENTS.md.bak` 存在
58
+ - Expected stdout phrases: `已覆盖` 或 `硬盖`
59
+ - Expected exit code: 0
39
60
 
40
- ### 5. Sync --force without backup
41
- Prompt: verify sync --force does not backup
42
- Command: `sync --force --dest <tmp>`
43
- Expected: stdout `已覆盖` without new `.bak`; exit 0.
61
+ ## 4. update 已合并到 sync --apply(删除分支、提示已合并)
44
62
 
45
- ### 6. List
46
- Prompt: verify skill listing
47
- Command: `list` and `list --json`
48
- Expected: 27 skills, includes `tdd` with correct description; `--json` is valid JSON array; exit 0.
63
+ - Prompt: 验证 `matt-skills update` 已删除,执行后报错提示已合并到 `sync --apply`,且 `--help` 不再列 `update`
64
+ - 溯源: spec.md — 需求/行为「`matt-skills update`:删除该分支,`main` 中 `command === 'update'` 改为 `stderr: 'update 已合并到 sync --apply'` 且 `exit 1`,`--help` 不再列 `update`」+ 验收标准「`matt-skills update` 执行后报错提示已合并,`--help` 无 `update`」
65
+ - type: cli
66
+ - Command: `node bin/cli.js update 2>&1; echo "exit:$?"` 与 `node bin/cli.js --help`
67
+ - Expected:
68
+ - `update` 命令 stdout/stderr 含 `已合并到 sync --apply` 且 exit 1
69
+ - `--help` 输出不含独立的 `update` 子命令行(不匹配 `^\s*update`)
70
+ - Expected stdout phrases: `已合并到 sync --apply`
71
+ - Expected exit code: 1(`update` 分支)
72
+ - env: {}
73
+
74
+ <!-- manual -->
75
+ <!-- 以下为人工定制实例保护段:由开发者手写,scaffold-functional-test 重生成时不覆盖此段以上的内容。如需新增手工实例,请在此段后追加。 -->
@@ -1,79 +1,67 @@
1
1
  # Logic Prototype
2
2
 
3
- A tiny interactive terminal app that lets the user drive a state model by hand. Use this when the question is about **business logic, state transitions, or data shape** — the kind of thing that looks reasonable on paper but only feels wrong once you push it through real cases.
3
+ A single, self-contained HTML file (a **shareable demo**) that lets anyone drive a state model by clicking buttons. Use this when the question is about **business logic, state transitions, or data shape**: the kind of thing that looks reasonable on paper but only feels wrong once you push it through real cases.
4
+
5
+ Because it's one file with nothing to install, you can hand it to a non-developer (a designer, a PM, a domain expert) and let them feel the model for themselves. So it speaks their language, not the code's.
4
6
 
5
7
  ## When this is the right shape
6
8
 
7
9
  - "I'm not sure if this state machine handles the edge case where X then Y."
8
10
  - "Does this data model actually let me represent the case where..."
9
11
  - "I want to feel out what the API should look like before writing it."
10
- - Anything where the user wants to **press buttons and watch state change**.
12
+ - Anything where someone wants to **press buttons and watch state change**.
11
13
 
12
- If the question is "what should this look like" — wrong branch. Use [UI.md](UI.md).
14
+ If the question is "what should this look like," this is the wrong branch. Use [UI.md](UI.md).
13
15
 
14
16
  ## Process
15
17
 
16
18
  ### 1. State the question
17
19
 
18
- Before writing code, write down what state model and what question you're prototyping. One paragraph, in the prototype's README or a comment at the top of the file. A logic prototype that answers the wrong question is pure waste — make the question explicit so it can be checked later, whether the user is watching now or returning to it AFK.
19
-
20
- ### 2. Pick the language
21
-
22
- Use whatever the host project uses. If the project has no obvious runtime (e.g. a docs repo), ask.
23
-
24
- Match the project's existing conventions for tooling — don't add a new package manager or runtime just for the prototype.
20
+ Before writing code, write down what state model and what question you're prototyping. One paragraph, at the top of the demo (in a visible intro, not just a comment). A logic prototype that answers the wrong question is pure waste, so make the question explicit so it can be checked later, whether the user is watching now or returning to it AFK.
25
21
 
26
- ### 3. Isolate the logic in a portable module
22
+ ### 2. Isolate the logic in a portable module
27
23
 
28
- Put the actual logic — the bit that's answering the question — behind a small, pure interface that could be lifted out and dropped into the real codebase later. The TUI around it is throwaway; the logic module shouldn't be.
24
+ Put the actual logic (the bit that's answering the question) in a single `<script>` block written as a small, pure module that could be lifted out and dropped into the real codebase later. The page around it is throwaway; this module isn't.
29
25
 
30
26
  The right shape depends on the question:
31
27
 
32
- - **A pure reducer** — `(state, action) => state`. Good when actions are discrete events and state is a single value.
33
- - **A state machine** — explicit states and transitions. Good when "which actions are even legal right now" is part of the question.
34
- - **A small set of pure functions** over a plain data type. Good when there's no implicit current state — just transformations.
28
+ - **A pure reducer**: `(state, action) => state`. Good when actions are discrete events and state is a single value.
29
+ - **A state machine**: explicit states and transitions. Good when "which actions are even legal right now" is part of the question.
30
+ - **A small set of pure functions** over a plain data type. Good when there's no implicit current state, just transformations.
35
31
  - **A class or module with a clear method surface** when the logic genuinely owns ongoing internal state.
36
32
 
37
- Pick whichever shape best fits the question being asked, *not* whichever is easiest to wire to a TUI. Keep it pure: no I/O, no terminal code, no `console.log` for control flow. The TUI imports it and calls into it; nothing flows the other direction.
38
-
39
- This is what makes the prototype useful past its own lifetime: when the question's been answered, the validated reducer / machine / function set can be lifted into the real module on its own.
40
-
41
- ### 4. Build the smallest TUI that exposes the state
42
-
43
- Build it as a **lightweight TUI** — on every tick, clear the screen (`console.clear()` / `print("\033[2J\033[H")` / equivalent) and re-render the whole frame. The user should always see one stable view, not an ever-growing scrollback.
44
-
45
- Each frame has two parts, in this order:
33
+ Pick whichever shape best fits the question being asked, *not* whichever is easiest to wire to a page. Keep it pure: no DOM, no `document`, no button handlers reaching inside it. The page calls into it; nothing flows the other direction. This is what makes the prototype useful past its own lifetime: once the question's answered, the validated reducer / machine / function set lifts into the real module on its own.
46
34
 
47
- 1. **Current state**, pretty-printed and diff-friendly (one field per line, or formatted JSON). Use **bold** for field names or section headers and **dim** for less important context (timestamps, IDs, derived values). Native ANSI escape codes are fine — `\x1b[1m` bold, `\x1b[2m` dim, `\x1b[0m` reset. No need to pull in a styling library unless one is already in the project.
48
- 2. **Keyboard shortcuts**, listed at the bottom: `[a] add user [d] delete user [t] tick clock [q] quit`. Bold the key, dim the description, or vice-versa — whatever reads cleanly.
35
+ ### 3. Build the shareable HTML file
49
36
 
50
- Behaviour:
37
+ One file, plain HTML/CSS/JS: no framework, no bundler, no server, everything inline so it opens by double-click and survives being emailed around. Anyone should be able to run it by opening it.
51
38
 
52
- 1. **Initialise state** — a single in-memory object/struct. Render the first frame on start.
53
- 2. **Read one keystroke (or one line)** at a time, dispatch to a handler that mutates state.
54
- 3. **Re-render** the full frame after every action — don't append, replace.
55
- 4. **Loop until quit.**
39
+ Write it for a non-developer. Every label is in **domain language**, not code: buttons and state read like the business, not the reducer. Explain in plain words what's happening.
56
40
 
57
- The whole frame should fit on one screen.
41
+ Lay it out with a clean hierarchy, top to bottom:
58
42
 
59
- ### 5. Make it runnable in one command
43
+ 1. **Title and one-line explanation** of what this demo lets you explore (the question from step 1).
44
+ 2. **Current state**: the full relevant state, rendered as a readable panel (labelled fields, not a raw JSON dump), re-rendered after every click so the change is visible. Where it helps a non-developer follow, call out what just changed.
45
+ 3. **Free-play buttons**: one button per action, always available, so anyone can poke at the model in any order. Each click dispatches its action and re-renders the state.
46
+ 4. **Guided walkthroughs**: a set of **scenarios**, one per tab. Each tab holds a short plain-language description of the scenario (the situation it sets up and what to watch for) and underneath it, the ordered **buttons to press** for that scenario. Each step is a real button: clicking it performs that action and moves to the next step. Starting a walkthrough resets to a known initial state so the scenario runs the same way every time.
60
47
 
61
- Add a script to the project's existing task runner (`package.json` scripts, `Makefile`, `justfile`, `pyproject.toml`). The user should run `pnpm run <prototype-name>` or equivalent — never need to remember a path.
48
+ Choose scenarios that demonstrate the awkward cases, the ones hard to reason about on paper: the happy path, a tricky edge case, an attempt at something that should be illegal.
62
49
 
63
- If the host project has no task runner, just put the command at the top of the prototype's README.
50
+ Keep it beautiful but restrained: clean typography, generous spacing, one accent colour. No animations, no gimmicks: nothing that competes with the state and the buttons.
64
51
 
65
- ### 6. Hand it over
52
+ ### 4. Hand it over
66
53
 
67
- Give the user the run command. They'll drive it themselves; the interesting moments are when they say "wait, that shouldn't be possible" or "huh, I assumed X would be different" — those are the bugs in the _idea_, which is the whole point. If they want new actions added, add them. Prototypes evolve.
54
+ Send them the file, or open it for them. They'll click through the walkthroughs and free-play whenever they get to it; the interesting moments are when they say "wait, that shouldn't be possible" or "huh, I assumed X would be different"; those are the bugs in the _idea_, which is the whole point. If they want new actions or a new scenario, add them. Prototypes evolve.
68
55
 
69
- ### 7. Capture the answer and the prototype
56
+ ### 5. Capture the answer and the prototype
70
57
 
71
- Once the prototype has answered its question, capture the answer, then capture the prototype the way the [SKILL](SKILL.md) describes. The logic-specific mapping: the validated reducer / machine / function set lifts into the real module (the decision, absorbed); the TUI shell rides along to the throwaway branch that keeps the prototype as a primary source.
58
+ Once the prototype has answered its question, capture the answer, then capture the prototype the way the [SKILL](SKILL.md) describes. The logic-specific mapping: the validated reducer / machine / function set lifts into the real module (the decision, absorbed); the HTML shell rides along to the throwaway branch that keeps the prototype as a primary source, and being one self-contained file, it stays trivially re-runnable there.
72
59
 
73
60
  ## Anti-patterns
74
61
 
75
62
  - **Don't add tests.** A prototype that needs tests is no longer a prototype.
76
- - **Don't wire it to the real database.** Use an in-memory store unless the question is specifically about persistence.
63
+ - **Don't wire it to the real database.** Use in-memory state unless the question is specifically about persistence.
77
64
  - **Don't generalise.** No "what if we wanted to support X later." The prototype answers one question.
78
- - **Don't blur the logic and the TUI together.** If the reducer / state machine references `console.log`, prompts, or terminal escape codes, it's no longer portable. Keep the TUI as a thin shell over a pure module.
79
- - **Don't ship the TUI shell into production.** The shell is optimised for being driven by hand from a terminal. The logic module behind it is the bit worth keeping.
65
+ - **Don't blur the logic and the page together.** If the pure module references the DOM, `document`, or button handlers, it's no longer liftable. Keep the page as a thin shell over a pure module.
66
+ - **Don't reach for a framework, bundler, or server.** One file the recipient double-clicks; a React app or a dev server defeats "shareable".
67
+ - **Don't ship the HTML shell into production.** The page is optimised for being clicked through by hand. The logic module behind it is the bit worth keeping.
@@ -9,18 +9,18 @@ A prototype is **throwaway code that answers a question**. The question decides
9
9
 
10
10
  ## Pick a branch
11
11
 
12
- Identify which question is being answered — from the user's prompt, the surrounding code, or by asking if the user is around:
12
+ Identify which question is being answered, using the user's prompt, the surrounding code, or by asking if the user is around:
13
13
 
14
- - **"Does this logic / state model feel right?"** → [LOGIC.md](LOGIC.md). Build a tiny interactive terminal app that pushes the state machine through cases that are hard to reason about on paper.
14
+ - **"Does this logic / state model feel right?"** → [LOGIC.md](LOGIC.md). Build a single shareable HTML file (free-play buttons plus tabbed guided walkthroughs) that pushes the state machine through cases that are hard to reason about on paper, and that a non-developer can drive.
15
15
  - **"What should this look like?"** → [UI.md](UI.md). Generate several radically different UI variations on a single route, switchable via a URL search param and a floating bottom bar.
16
16
 
17
- The two branches produce very different artifacts — getting this wrong wastes the whole prototype. If the question is genuinely ambiguous and the user isn't reachable, default to whichever branch better matches the surrounding code (a backend module → logic; a page or component → UI) and state the assumption at the top of the prototype.
17
+ The two branches produce very different artifacts, so getting this wrong wastes the whole prototype. If the question is genuinely ambiguous and the user isn't reachable, default to whichever branch better matches the surrounding code (a backend module → logic; a page or component → UI) and state the assumption at the top of the prototype.
18
18
 
19
19
  ## Rules that apply to both
20
20
 
21
- 1. **Throwaway from day one, and clearly marked as such.** Locate the prototype code close to where it will actually be used (next to the module or page it's prototyping for) so context is obvious — but name it so a casual reader can see it's a prototype, not production. For throwaway UI routes, obey whatever routing convention the project already uses; don't invent a new top-level structure.
22
- 2. **One command to run.** Whatever the project's existing task runner supports — `pnpm <name>`, `python <path>`, `bun <path>`, etc. The user must be able to start it without thinking.
23
- 3. **No persistence by default.** State lives in memory. Persistence is the thing the prototype is _checking_, not something it should depend on. If the question explicitly involves a database, hit a scratch DB or a local file with a clear "PROTOTYPE — wipe me" name.
21
+ 1. **Throwaway from day one, and clearly marked as such.** Locate the prototype code close to where it will actually be used (next to the module or page it's prototyping for) so context is obvious, but name it so a casual reader can see it's a prototype, not production. For throwaway UI routes, obey whatever routing convention the project already uses; don't invent a new top-level structure.
22
+ 2. **Trivial to run.** A UI prototype starts from one command in the project's task runner: `pnpm <name>`, `python <path>`, `bun <path>`, etc. A logic demo is a single HTML file the user double-clicks. Either way, no thinking required to start it.
23
+ 3. **No persistence by default.** State lives in memory. Persistence is the thing the prototype is _checking_, not something it should depend on. If the question explicitly involves a database, hit a scratch DB or a local file with a clear "PROTOTYPE, wipe me" name.
24
24
  4. **Skip the polish.** No tests, no error handling beyond what makes the prototype _runnable_, no abstractions. The point is to learn something fast.
25
25
  5. **Surface the state.** After every action (logic) or on every variant switch (UI), print or render the full relevant state so the user can see what changed.
26
- 6. **Capture it when done.** Fold any validated decision into the real code, then capture the prototype itself as a **primary source**: commit it to a throwaway branch, out of main, and leave a context pointer to that branch on the implementation issue. Capture the answer too — the verdict and the question it settled — in the issue or a commit. The main branch keeps only the validated decision.
26
+ 6. **Capture it when done.** Fold any validated decision into the real code, then capture the prototype itself as a **primary source**: commit it to a throwaway branch, out of main, and leave a context pointer to that branch on the implementation issue. Capture the answer too (the verdict and the question it settled) in the issue or a commit. The main branch keeps only the validated decision.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Generate **several radically different UI variations** on a single route, switchable from a floating bottom bar. The user flips between variants in the browser, picks one (or steals bits from each), then throws the rest away.
4
4
 
5
- If the question is about logic/state rather than what something looks like — wrong branch. Use [LOGIC.md](LOGIC.md).
5
+ If the question is about logic/state rather than what something looks like, this is the wrong branch. Use [LOGIC.md](LOGIC.md).
6
6
 
7
7
  ## When this is the right shape
8
8
 
@@ -11,21 +11,21 @@ If the question is about logic/state rather than what something looks like — w
11
11
  - "Try a different layout for the settings screen."
12
12
  - Any time the user would otherwise spend a day picking between three vague mockups in their head.
13
13
 
14
- ## Two sub-shapes — strongly prefer sub-shape A
14
+ ## Two sub-shapes: strongly prefer sub-shape A
15
15
 
16
- A UI prototype is much easier to judge when it's **butting up against the rest of the app** — real header, real sidebar, real data, real density. A throwaway route on its own is a vacuum: every variant looks fine in isolation. Default to sub-shape A whenever there's a plausible existing page to host the variants. Only reach for sub-shape B if the prototype genuinely has no nearby home.
16
+ A UI prototype is much easier to judge when it's **butting up against the rest of the app**: real header, real sidebar, real data, real density. A throwaway route on its own is a vacuum: every variant looks fine in isolation. Default to sub-shape A whenever there's a plausible existing page to host the variants. Only reach for sub-shape B if the prototype genuinely has no nearby home.
17
17
 
18
- ### Sub-shape A — adjustment to an existing page (preferred)
18
+ ### Sub-shape A: adjustment to an existing page (preferred)
19
19
 
20
- The route already exists. Variants are rendered **on the same route**, gated by a `?variant=` URL search param. The existing data fetching, params, and auth all stay — only the rendering swaps. This is the default; pick it unless there's a specific reason not to.
20
+ The route already exists. Variants are rendered **on the same route**, gated by a `?variant=` URL search param. The existing data fetching, params, and auth all stay. Only the rendering swaps. This is the default; pick it unless there's a specific reason not to.
21
21
 
22
- If the prototype is for something that doesn't yet have a page but *would naturally live inside one* (a new section of the dashboard, a new card on the settings screen, a new step in an existing flow) — that's still sub-shape A. Mount the variants inside the host page.
22
+ If the prototype is for something that doesn't yet have a page but *would naturally live inside one* (a new section of the dashboard, a new card on the settings screen, a new step in an existing flow), it's still sub-shape A. Mount the variants inside the host page.
23
23
 
24
- ### Sub-shape B — a new page (last resort)
24
+ ### Sub-shape B: a new page (last resort)
25
25
 
26
- Only use this when the thing being prototyped genuinely has no existing page to live inside — e.g. an entirely new top-level surface, or a flow that can't be embedded anywhere sensible.
26
+ Only use this when the thing being prototyped genuinely has no existing page to live inside (e.g. an entirely new top-level surface, or a flow that can't be embedded anywhere sensible).
27
27
 
28
- Create a **throwaway route** following whatever routing convention the project already uses — don't invent a new top-level structure. Name it so it's obviously a prototype (e.g. include the word `prototype` in the path or filename). Same `?variant=` pattern.
28
+ Create a **throwaway route** following whatever routing convention the project already uses. Don't invent a new top-level structure. Name it so it's obviously a prototype (e.g. include the word `prototype` in the path or filename). Same `?variant=` pattern.
29
29
 
30
30
  Before committing to sub-shape B, sanity-check: is there really no existing page this could be embedded in? An empty route hides design problems that a populated one would expose.
31
31
 
@@ -35,7 +35,7 @@ In both sub-shapes the floating bottom bar is identical.
35
35
 
36
36
  ### 1. State the question and pick N
37
37
 
38
- Default to **3 variants**. More than 5 stops being radically different and starts being noise — cap there.
38
+ Default to **3 variants**. More than 5 stops being radically different and starts being noise, so cap there.
39
39
 
40
40
  Write down the plan in one line, in the prototype's location or a top-of-file comment:
41
41
 
@@ -51,14 +51,14 @@ Draft each variant. Hold each one to:
51
51
  - The project's component library / styling system (TailwindCSS, shadcn, MUI, plain CSS, whatever).
52
52
  - A clear exported component name, e.g. `VariantA`, `VariantB`, `VariantC`.
53
53
 
54
- Variants must be **structurally different** — different layout, different information hierarchy, different primary affordance, not just different colours. Three slightly-tweaked card grids isn't a UI prototype, it's wallpaper. If two drafts come out too similar, redo one with explicit "do not use a card grid" guidance.
54
+ Variants must be **structurally different**: different layout, different information hierarchy, different primary affordance, not just different colours. Three slightly-tweaked card grids isn't a UI prototype, it's wallpaper. If two drafts come out too similar, redo one with explicit "do not use a card grid" guidance.
55
55
 
56
56
  ### 3. Wire them together
57
57
 
58
58
  Create a single switcher component on the route:
59
59
 
60
60
  ```tsx
61
- // pseudo-code — adapt to the project's framework
61
+ // pseudo-code, adapt to the project's framework
62
62
  const variant = searchParams.get('variant') ?? 'A';
63
63
  return (
64
64
  <>
@@ -78,35 +78,35 @@ For sub-shape B (new page): the throwaway route under `/prototype/<name>` mounts
78
78
 
79
79
  A small fixed-position bar at the bottom-centre of the screen with three pieces:
80
80
 
81
- - **Left arrow** — cycles to the previous variant (wraps around).
82
- - **Variant label** — shows the current variant key and, if the variant exports a name, that name too. e.g. `B — Sidebar layout`.
83
- - **Right arrow** — cycles forward (wraps around).
81
+ - **Left arrow**: cycles to the previous variant (wraps around).
82
+ - **Variant label**: shows the current variant key and, if the variant exports a name, that name too. e.g. `B (Sidebar layout)`.
83
+ - **Right arrow**: cycles forward (wraps around).
84
84
 
85
85
  Behaviour:
86
86
 
87
- - Clicking an arrow updates the URL search param (use the framework's router — `router.replace` on Next, `navigate` on React Router, etc) so the variant is shareable and reload-stable.
87
+ - Clicking an arrow updates the URL search param (use the framework's router, e.g. `router.replace` on Next, `navigate` on React Router, etc) so the variant is shareable and reload-stable.
88
88
  - Keyboard: `←` and `→` arrow keys also cycle. Don't intercept arrow keys when an `<input>`, `<textarea>`, or `[contenteditable]` is focused.
89
89
  - Visually distinct from the page (e.g. high-contrast pill, subtle shadow) so it's obviously not part of the design being evaluated.
90
- - Hidden in production builds — gate on `process.env.NODE_ENV !== 'production'` or an equivalent check, so a stray prototype merge can't ship the bar to users.
90
+ - Hidden in production builds: gate on `process.env.NODE_ENV !== 'production'` or an equivalent check, so a stray prototype merge can't ship the bar to users.
91
91
 
92
92
  Put the switcher in a single shared component so both sub-shapes can reuse it. Locate it wherever shared UI lives in the project.
93
93
 
94
94
  ### 5. Hand it over
95
95
 
96
- Surface the URL (and the `?variant=` keys). The user will flip through whenever they get to it. The interesting feedback is usually **"I want the header from B with the sidebar from C"** — that's the actual design they want.
96
+ Surface the URL (and the `?variant=` keys). The user will flip through whenever they get to it. The interesting feedback is usually **"I want the header from B with the sidebar from C"**, which is the actual design they want.
97
97
 
98
98
  ### 6. Capture the answer and clean up
99
99
 
100
- Once a variant has won, capture the answer — which variant and why — then capture the prototype the way the [SKILL](SKILL.md) describes. Fold the winner into the real code and move the rest onto the throwaway branch, not into main:
100
+ Once a variant has won, capture the answer (which variant and why), then capture the prototype the way the [SKILL](SKILL.md) describes. Fold the winner into the real code and move the rest onto the throwaway branch, not into main:
101
101
 
102
- - **Sub-shape A** — fold the winner into the existing page; drop the losing variants and the switcher from main.
103
- - **Sub-shape B** — promote the winning variant to a real route; drop the throwaway route and the switcher from main.
102
+ - **Sub-shape A**: fold the winner into the existing page; drop the losing variants and the switcher from main.
103
+ - **Sub-shape B**: promote the winning variant to a real route; drop the throwaway route and the switcher from main.
104
104
 
105
- The full set of variants is the primary source, so it lands on the throwaway branch, not the bin — variant components and the switcher left in the main branch rot fast and confuse the next reader.
105
+ The full set of variants is the primary source, so it lands on the throwaway branch, not the bin, since variant components and the switcher left in the main branch rot fast and confuse the next reader.
106
106
 
107
107
  ## Anti-patterns
108
108
 
109
109
  - **Variants that differ only in colour or copy.** That's a tweak, not a prototype. Real variants disagree about structure.
110
110
  - **Sharing too much code between variants.** A shared `<Header>` is fine; a shared `<Layout>` defeats the point. Each variant should be free to throw out the layout.
111
- - **Wiring variants to real mutations.** Read-only prototypes are fine. If a variant needs to mutate, point it at a stub — the question is "what should this look like", not "does the backend work".
111
+ - **Wiring variants to real mutations.** Read-only prototypes are fine. If a variant needs to mutate, point it at a stub: the question is "what should this look like", not "does the backend work".
112
112
  - **Promoting the prototype directly to production.** The variant code was written under prototype constraints (no tests, minimal error handling). Rewrite it properly when you fold it in.
@@ -7,6 +7,6 @@ Spin up a **background agent** to do the research, so you keep working while it
7
7
 
8
8
  Its job:
9
9
 
10
- 1. Investigate the question against **primary sources** — official docs, source code, specs, first-party APIs — not a secondary write-up of them. Follow every claim back to the source that owns it.
10
+ 1. Investigate the question against **primary sources** (official docs, source code, specs, first-party APIs), not a secondary write-up of them. Follow every claim back to the source that owns it.
11
11
  2. Write the findings to a single Markdown file, citing each claim's source.
12
12
  3. Save it where the repo already keeps such notes; match the existing convention, and if there is none, put it somewhere sensible and say where.
@@ -9,6 +9,6 @@ description: "Use when you need to resolve an in-progress git merge/rebase confl
9
9
 
10
10
  3. **Resolve each hunk.** Preserve both intents where possible. Where incompatible, pick the one matching the merge's stated goal and note the trade-off. Do **not** invent new behaviour. Always resolve; never `--abort`.
11
11
 
12
- 4. Discover the project's **automated checks** and run them — typically typecheck, then tests, then format. Fix anything the merge broke.
12
+ 4. Discover the project's **automated checks** and run them, typically typecheck, then tests, then format. Fix anything the merge broke.
13
13
 
14
14
  5. **Finish the merge/rebase.** Stage everything and commit. If rebasing, continue the rebase process until all commits are rebased.
@@ -0,0 +1,77 @@
1
+ ---
2
+ name: scaffold-functional-test
3
+ disable-model-invocation: false
4
+ description: "Scaffold a repo-specific functional-test skill from spec — use when the user wants to generate a customized functional-test suite/skill from a spec/README/help; not for running tests (use instance-test) nor for TDD (use tdd-implement)"
5
+ ---
6
+
7
+ # Scaffold Functional Test
8
+
9
+ 从本仓库的 spec 自动脚手架出**仓库专属的功能测试 skill**。本技能为**非 Long-Horizon 轻量 skill**(一次性 scaffold,不做多 seam 红绿循环),一次性完成「读 spec → 推导实例 → 落盘 skill → 自验证」闭环。术语定义见 `CONTEXT.md`。
10
+
11
+ ## 产出物
12
+
13
+ - 定制 skill 目录:`.agents/skills/<repo>-functional-test/`(含 `SKILL.md` + `references/instances.md` + 可选 `scripts/run.sh`)
14
+ - 指纹:`spec hash` + `generatedAt` 写入生成物头部,用于后续执行前校验
15
+ - 保护:`<!-- manual -->` 标记段不被覆盖
16
+
17
+ 生成物纳入 git,可回归复用,不进入 `template/` 再分发(生成器本身才随 Template Snapshot 分发)。
18
+
19
+ ## Steps
20
+
21
+ ### ① 采集 Spec
22
+
23
+ 解析用户传入的 spec 路径,默认 `.scratch/<feature>/spec.md`。
24
+
25
+ - 若 spec 存在:读取 `CONTEXT.md`/`docs/adr/` 相关术语与决策,提取待覆盖行为清单(以验收标准为锚点)。
26
+ - 若 spec 不存在:回退到 `README` + `--help` 输出倒推行为清单,但必须进入 Step ② 的清单确认关卡,不静默臆测。
27
+
28
+ 完成:待覆盖行为清单已固定,无未澄清歧义。
29
+
30
+ ### ② 推导实例
31
+
32
+ 按混合推导策略生成实例草案:
33
+
34
+ - 以验收标准为锚点,需求/接口/边界为补充,可为 spec 未显式写的隐含行为(如 `--help` 文案、错误码、幂等性)补实例,但每条实例必须标注**溯源**(spec 章节/行号或 `README/--help` 来源),无溯源的实例视为幻觉需删除。
35
+ - 每实例声明**受控扩展模型**:必选 `prompt/command/expected files/content/expected stdout phrases/expected exit code`,可选 `setup/env/timeout/type/teardown`,默认 `type: cli`。
36
+ - **强制门禁**:实例清单必须与用户确认后才进入 Step ③;无确认不落盘。
37
+
38
+ 完成:实例清单已获用户确认,每实例含溯源与完整四元组。
39
+
40
+ ### ③ 脚手架落盘
41
+
42
+ 按受控扩展模型写入定制 skill 目录:
43
+
44
+ - `SKILL.md`:执行语义(见下节「执行语义」)
45
+ - `references/instances.md`:实例集(含溯源、必选+可选字段、头部 `spec hash` + `generatedAt`)
46
+ - 不覆盖 `<!-- manual -->` 保护段;覆盖式更新需经用户确认;重生成时先给出 diff 建议,用户确认后才应用。
47
+
48
+ 完成:定制 skill 目录已落盘,指纹正确,人工段受保护。
49
+
50
+ ### ④ 自验证
51
+
52
+ 落盘后立即按实例执行语义串行执行一轮实例集作自验证:
53
+
54
+ - `mktemp -d` 隔离(或项目支持的 `git worktree` / `--dest`),单线程串行,不并行。
55
+ - 每实例捕获 stdout/stderr 与 exit code,按 `test -f`/`grep -q`/`diff` 对比判定 `PASS`/`FAIL`,单 FAIL 不阻断后续。
56
+ - 对话内输出 `PASS m/n` + per-instance evidence(`expected vs actual diff` + `run dir`),失败不回滚生成物但给出 gap 供迭代 `regenerate`。
57
+ - 成功默认清理临时目录、失败默认保留(`--keep` 保留全部);`--report` 显式开启才落盘报告文件。
58
+
59
+ 完成:自验证已执行,对话内汇总完成,证据可复现。
60
+
61
+ ## 执行语义(生成物复用)
62
+
63
+ 生成物本身的执行语义与 `instance-test` 一致:`mktemp -d` 串行、`PASS m/n` 汇总、证据含 `expected vs actual diff` + `run dir`。执行前校验 `spec hash` 指纹:若当前 spec 已变更,提示「spec 已变更,建议重跑 scaffold-functional-test」但不自动覆盖,需用户显式确认才 regenerate。
64
+
65
+ ## 不做什么
66
+
67
+ - 不替代 `tdd`/`tdd-implement` 的红绿循环与 `commit-check` 门禁
68
+ - 不自动织入每次 `tdd-implement` 或 `commit-check`;仅 `tdd-implement --with-functional` 显式 opt-in
69
+ - 不支持并行执行与 `docker` 隔离
70
+ - 不处理超出混合推导锚点范围的源码静态分析隐式行为挖掘
71
+
72
+ ## 引用
73
+
74
+ - 领域术语:`CONTEXT.md`
75
+ - 技能设计规则:`docs/agents/skill-design.md`
76
+ - 示范产物:`.agents/skills/instance-test/`(本仓库专属,见其 SKILL.md)
77
+ - Issue tracker:`docs/agents/issue-tracker.md`
@@ -0,0 +1,5 @@
1
+ interface:
2
+ display_name: "Scaffold Functional Test"
3
+ short_description: "Scaffold a repo-specific functional-test skill from spec — not for running tests nor TDD"
4
+ policy:
5
+ allow_implicit_invocation: true