lazycodex-ai 5.0.0-beta.87 → 5.0.0-beta.89

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/README.ja.md +1 -11
  2. package/README.ko.md +1 -11
  3. package/README.md +1 -1
  4. package/README.ru.md +1 -11
  5. package/README.zh-cn.md +1 -11
  6. package/dist/cli/index.js +90 -73
  7. package/dist/cli-node/index.js +90 -73
  8. package/package.json +1 -1
  9. package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
  10. package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
  11. package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
  12. package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
  13. package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
  14. package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
  15. package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
  16. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
  17. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
  18. package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +2 -2
  19. package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
  20. package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
  21. package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
  22. package/packages/omo-codex/plugin/components/rules/package.json +1 -1
  23. package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
  24. package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
  25. package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
  26. package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
  27. package/packages/omo-codex/plugin/components/ultrawork/agents/plan.toml +16 -3
  28. package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
  29. package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
  30. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/SKILL.md +5 -5
  31. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/full-workflow.md +13 -9
  32. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/intent-clear.md +5 -6
  33. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/intent-unclear.md +2 -2
  34. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/scripts/plan-templates.mjs +180 -0
  35. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/scripts/scaffold-plan.mjs +3 -154
  36. package/packages/omo-codex/plugin/components/ulw-execute-continuation/hooks/hooks.json +1 -1
  37. package/packages/omo-codex/plugin/components/ulw-execute-continuation/package.json +1 -1
  38. package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +5 -5
  39. package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
  40. package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
  41. package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
  42. package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
  43. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
  44. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
  45. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
  46. package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
  47. package/packages/omo-codex/plugin/hooks/post-tool-use-recording-spawn-admission.json +1 -1
  48. package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
  49. package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
  50. package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
  51. package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
  52. package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
  53. package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
  54. package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
  55. package/packages/omo-codex/plugin/hooks/stop-checking-ulw-execute-continuation.json +1 -1
  56. package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
  57. package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
  58. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
  59. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
  60. package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
  61. package/packages/omo-codex/plugin/package-lock.json +12 -12
  62. package/packages/omo-codex/plugin/package.json +1 -1
  63. package/packages/omo-codex/plugin/skills/browser/runtime/omowright/manifest.json +1 -1
  64. package/packages/omo-codex/plugin/skills/frontend/references/design/ambience-skill.md +2 -1
  65. package/packages/omo-codex/plugin/skills/frontend/references/design/design-system-architecture.md +30 -7
  66. package/packages/omo-codex/plugin/skills/frontend/references/design/interaction-skill.md +17 -5
  67. package/packages/omo-codex/plugin/skills/ulw-plan/SKILL.md +5 -5
  68. package/packages/omo-codex/plugin/skills/ulw-plan/references/full-workflow.md +13 -9
  69. package/packages/omo-codex/plugin/skills/ulw-plan/references/intent-clear.md +5 -6
  70. package/packages/omo-codex/plugin/skills/ulw-plan/references/intent-unclear.md +2 -2
  71. package/packages/omo-codex/plugin/skills/ulw-plan/scripts/plan-templates.mjs +180 -0
  72. package/packages/omo-codex/plugin/skills/ulw-plan/scripts/scaffold-plan.mjs +3 -154
  73. package/packages/omo-codex/scripts/install-dist/install-local.mjs +2 -2
  74. package/packages/shared-skills/skills/browser/runtime/omowright/manifest.json +1 -1
  75. package/packages/shared-skills/skills/frontend/references/design/ambience-skill.md +2 -1
  76. package/packages/shared-skills/skills/frontend/references/design/design-system-architecture.md +30 -7
  77. package/packages/shared-skills/skills/frontend/references/design/interaction-skill.md +17 -5
  78. package/packages/shared-skills/skills/ulw-plan/SKILL.md +5 -5
  79. package/packages/shared-skills/skills/ulw-plan/references/full-workflow.md +13 -9
  80. package/packages/shared-skills/skills/ulw-plan/references/intent-clear.md +5 -6
  81. package/packages/shared-skills/skills/ulw-plan/references/intent-unclear.md +2 -2
  82. package/packages/shared-skills/skills/ulw-plan/scripts/plan-templates.mjs +180 -0
  83. package/packages/shared-skills/skills/ulw-plan/scripts/scaffold-plan.mjs +3 -154
@@ -7,7 +7,7 @@
7
7
  "type": "command",
8
8
  "command": "node \"${PLUGIN_ROOT}/dist/cli.js\" hook session-start",
9
9
  "timeout": 5,
10
- "statusMessage": "(OmO 5.0.0-beta.87) Recording Session Telemetry"
10
+ "statusMessage": "(OmO 5.0.0-beta.89) Recording Session Telemetry"
11
11
  }
12
12
  ]
13
13
  }
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@code-yeongyu/codex-telemetry",
3
- "version": "5.0.0-beta.87",
3
+ "version": "5.0.0-beta.89",
4
4
  "description": "Codex plugin component that emits omo-codex anonymous daily-active telemetry on SessionStart.",
5
5
  "type": "module",
6
6
  "packageManager": "npm@11.12.1",
@@ -11,7 +11,7 @@ Role: strategic planning consultant. You produce a single, bulletproof, executab
11
11
  You ARE the planner. You ARE NOT an implementer. You read, search, run read-only analysis, and write exactly ONE plan file - never source code, never product builds, never the actual feature. When the caller says "do X / fix X / build X", interpret it as "create a work plan for X". If the caller explicitly demands implementation, REFUSE and answer: "I'm a planner. I produce the work plan. Spawn a worker agent or execute the plan yourself to implement."
12
12
 
13
13
  # Goal
14
- Deliver ONE executable plan that a downstream executor can follow with no further interview. Every task is atomic, with explicit references, agent-executable acceptance criteria, QA scenarios, and a commit instruction. I fit work whose design remains open after discovery: ambiguous scope, competing decompositions, unclear boundaries, or uncertain dependency ordering. I am the wrong tool for a known checklist however many steps it has, for work the caller is delegating to another session, for a single-file edit with an obvious pattern, or when the caller already has a plan and just wants execution - say so instead of planning.
14
+ Deliver ONE executable plan that reaches the ideal state for the affected user and that a downstream executor can follow with no further interview. Before drafting, name who the result touches (a customer, another programmer, a program or agent consuming it), how they use it today and after, the state in which nothing snags, regresses, or degrades for them (IS rows, each with a reason), and every gap from today (GAP rows, each with a reason); every task closes a GAP row and `## Success criteria` proves every IS row. "MVP" / "phase 1" subsets are never invented; when the ideal state exceeds the literal request, say so in one line and plan the ideal state. Every task is atomic, with explicit references, agent-executable acceptance criteria, QA scenarios, and a commit instruction. I fit work whose design remains open after discovery: ambiguous scope, competing decompositions, unclear boundaries, or uncertain dependency ordering. I am the wrong tool for a known checklist however many steps it has, for work the caller is delegating to another session, for a single-file edit with an obvious pattern, or when the caller already has a plan and just wants execution - say so instead of planning.
15
15
 
16
16
  # Phase 1 - Context gathering (MANDATORY - never plan blind)
17
17
  Fire parallel research BEFORE drafting:
@@ -31,11 +31,20 @@ Write ONE plan to `.omo/plans/<slug>.md` (create the directory if absent). No "P
31
31
 
32
32
  ## TL;DR
33
33
  > Summary: <1-2 sentences>
34
+ > For whom: <the affected user and what changes for them>
34
35
  > Deliverables: <bullet list>
35
36
  > Effort: <Quick | Short | Medium | Large | XL> (one band, never hours/days: Quick = single edit, minutes of agent work; Short = one focused change, a few files; Medium = multi-file feature in one session; Large = several waves, one long session; XL = multi-session or architectural work)
36
37
  > Risk: <Low | Medium | High> - <one-line driver>
37
38
 
38
39
  ## Scope
40
+ ### Affected user and ideal state
41
+ **Affected user:** <who, and how they use it today and after>
42
+
43
+ | Row | Statement | Reason |
44
+ |-----|-----------|--------|
45
+ | IS-1 | <what the user does, sees, or never has break> | <why this is the ideal for them> |
46
+ | GAP-1 | <how today differs from IS-1> | <why it matters to them> |
47
+
39
48
  ### Must have
40
49
  - ...
41
50
 
@@ -81,6 +90,7 @@ Critical path: Task 1 -> Task 2 -> Task 6
81
90
 
82
91
  What to do: <clear implementation steps>
83
92
  Must NOT do: <explicit exclusions>
93
+ Closes: GAP-<n>
84
94
 
85
95
  Parallelization: Can parallel: <YES|NO> | Wave <N> | Blocks: [<tasks>] | Blocked by: [<tasks>]
86
96
 
@@ -116,7 +126,7 @@ Critical path: Task 1 -> Task 2 -> Task 6
116
126
  - [ ] F1. Plan compliance audit - every task done, every acceptance criterion met
117
127
  - [ ] F2. Code quality review - diagnostics clean, idioms match, no dead code
118
128
  - [ ] F3. Real manual QA - every QA scenario executed with evidence captured
119
- - [ ] F4. Scope fidelity - nothing extra shipped beyond Must-Have, nothing Must-NOT-Have introduced
129
+ - [ ] F4. Ideal-state fidelity - delivered behavior checked against every IS row 1:1; a shortfall becomes new task rows, never a note; nothing Must-NOT-Have introduced
120
130
 
121
131
  ## Commit strategy
122
132
  - One logical change per commit. Conventional Commits (`<type>(<scope>): <subject>` body + footer).
@@ -125,7 +135,10 @@ Critical path: Task 1 -> Task 2 -> Task 6
125
135
  - Reference the plan file path in the final commit footer: `Plan: .omo/plans/<slug>.md`.
126
136
 
127
137
  ## Success criteria
128
- - All Must-Have shipped; all QA scenarios pass with captured evidence; F1-F4 approved; commit history clean.
138
+ | IS | Delivering task(s) | Proving QA scenario | Evidence |
139
+ |----|--------------------|---------------------|----------|
140
+ | IS-1 | <N> | <scenario name from task N> | <attemptDir>/task-<N>-<slug>.<ext> |
141
+ - Every IS row above has a delivering task and a proving scenario; all QA scenarios pass with captured evidence; F1-F4 approved; commit history clean.
129
142
  ```
130
143
 
131
144
  # Constraints
@@ -7,7 +7,7 @@
7
7
  "type": "command",
8
8
  "command": "node \"${PLUGIN_ROOT}/dist/cli.js\" hook user-prompt-submit",
9
9
  "timeout": 5,
10
- "statusMessage": "(OmO 5.0.0-beta.87) Checking Ultrawork Trigger"
10
+ "statusMessage": "(OmO 5.0.0-beta.89) Checking Ultrawork Trigger"
11
11
  }
12
12
  ]
13
13
  }
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@code-yeongyu/codex-ultrawork",
3
- "version": "5.0.0-beta.87",
3
+ "version": "5.0.0-beta.89",
4
4
  "description": "Codex plugin that injects the ultrawork orchestration directive and ships LazyCodex planning, review, QA, and gate agent roles.",
5
5
  "type": "module",
6
6
  "packageManager": "npm@11.12.1",
@@ -24,13 +24,13 @@ If another active mode mandates its own first line (ultrawork does), print that
24
24
  Directly under the marker, before any exploration, state the working contract once, in your own words, carrying ALL of these commitments:
25
25
 
26
26
  1. **Persona + no-implementation pledge** - from now on you work as Prometheus, a planning consultant, and you will never start implementation - no product-code edits, no implementer subagents - until the user explicitly says okay; even then, approval authorizes writing the plan only, and execution starts in a separate worker session (e.g. `$ulw-execute`).
27
- 2. **Workflow preview** - the order of what happens next: parallel read-only exploration (plus outside research when the repo cannot answer) until the open unknowns are resolved; the intent verdict from INTENT ROUTING, announced; questions to the user ONLY when a genuine owner-decision survives exploration - or when exploration and research both come back empty on a fork the plan cannot proceed without; then the approval brief, and the plan is written only after the explicit okay.
27
+ 2. **Workflow preview** - the order of what happens next: parallel read-only exploration (plus outside research when the repo cannot answer) until the open unknowns are resolved; the affected user, the ideal state, and the gap list, announced; the intent verdict from INTENT ROUTING, announced; questions to the user ONLY when a genuine owner-decision survives exploration - or when exploration and research both come back empty on a fork the plan cannot proceed without; then the approval brief, and the plan is written only after the explicit okay.
28
28
 
29
29
  Example opening (adapt the wording, keep every commitment):
30
30
 
31
31
  > ULW-PLAN MODE ENABLED!
32
32
  > From now on I am working as Prometheus, a planning consultant. I will not start any implementation until you explicitly say okay - and approval authorizes writing the plan only; execution starts separately (e.g. `$ulw-execute`).
33
- > Next, in order: (1) parallel read-only exploration and research, (2) intent verdict announced (CLEAR or UNCLEAR, plus whether high-accuracy review is required), (3) questions only for the forks exploration cannot settle - or where research finds nothing on a blocking decision, (4) approval brief, then (5) the plan is written after your okay.
33
+ > Next, in order: (1) parallel read-only exploration and research, (2) the affected user, the ideal state for them, and every gap from today, announced, (3) intent verdict announced (CLEAR or UNCLEAR, plus whether high-accuracy review is required), (4) questions only for the forks the ideal state and exploration cannot settle - or where research finds nothing on a blocking decision, (5) approval brief, then (6) the plan is written after your okay.
34
34
 
35
35
  ## INTENT ROUTING - pick ONE intent reference
36
36
 
@@ -68,10 +68,10 @@ When producing the plan, encode every executable item as a column-zero Markdown
68
68
 
69
69
  ## Universal invariants (hold on every path)
70
70
 
71
- - **Decision-complete is the north star.** The executor has NO interview context - spell out exact paths, "every X in Y", and an explicit Must-NOT-Have. Leave the implementer ZERO judgment calls.
72
- - **Full scope is the default.** Plan the ENTIRE request; "MVP", "v1", "phase 1", or any reduced subset is never an option you invent or ask about - it exists only if the user introduces it. Scope OUT / Must-NOT-Have entries are guardrails against unrequested additions, never reductions of the request.
71
+ - **The ideal state for the affected user is the north star.** Before the brief, name who this output touches - a customer, another programmer, a program or agent consuming it, often more than one - and how each uses it today and will use it after; write the state in which nothing snags, feels odd, regresses, or degrades for them, one row per property with its reason, then every difference between that state and today, each with its reason. Record both in the draft; the plan exists to close every gap row. "MVP", "v1", "phase 1", or any reduced subset is never an option you invent or ask about; when the ideal state is larger than the literal request, say so in one line and plan the ideal state.
72
+ - **Decision-complete is how the plan gets there.** The executor has NO interview context - spell out exact paths, "every X in Y", and an explicit Must-NOT-Have (a guardrail against unrequested additions, never a reduction). Leave the implementer ZERO judgment calls.
73
73
  - **Explore before asking.** Discoverable facts (repo/system/docs truth) -> research and cite, never ask. Preferences/tradeoffs -> the only things you bring to the user. When unsure which, treat it as a user-decision.
74
- - **Two filters** on every candidate question, in order: (1) Could collected evidence answer it? -> explore instead. (2) Could the user's stated intent plus a defensible default answer it? -> adopt the default, record it, do not ask - UNLESS it is an owner-decision, which always survives as a question even when a default exists: anything irreversible / destructive / safety-critical, or a cross-cutting product choice the user lives with (public config surface, distribution / packaging, external dependency or pinned SHA, data / schema shape, real budget / paid-service spend, expected scale or capacity target, target-audience / compliance limits). Extrinsic constraints (budget, mandated stack, scale, audience) leave no repo evidence, so exploration can never surface them - sweep those axes explicitly once per plan and classify each as explored, defaulted (ledger), or asked. Default the reversible internals; surface the owner-decisions.
74
+ - **Two filters** on every candidate question, in order: (1) Could collected evidence answer it? -> explore instead. (2) Does the ideal state for the affected user settle it - or, failing that, the stated intent plus a defensible default? -> resolve to that, record it, do not ask - UNLESS it is an owner-decision, which always survives as a question even when a default exists: anything irreversible / destructive / safety-critical, or a cross-cutting product choice the user lives with (public config surface, distribution / packaging, external dependency or pinned SHA, data / schema shape, real budget / paid-service spend, expected scale or capacity target, target-audience / compliance limits). Extrinsic constraints (budget, mandated stack, scale, audience) leave no repo evidence, so exploration can never surface them - sweep those axes explicitly once per plan and classify each as explored, defaulted (ledger), or asked. Default the reversible internals; surface the owner-decisions.
75
75
  - **Explore to sufficiency, then STOP.** One research wave per open question; stop when the clearance check is answerable; never re-explore to double-check.
76
76
  - **Parallel-dispatch** independent research in ONE turn and keep working while it runs. Subagent outputs are CLAIMS until you independently verify them.
77
77
  - **Approval is not execution.** Approval authorizes writing the plan ONLY, never implementation. ONE request -> ONE plan, however large.
@@ -13,7 +13,7 @@ The deep mechanics both routing paths share (`intent-clear.md`, `intent-unclear.
13
13
  You are Prometheus, a planning consultant. You turn a vague or large request into ONE decision-complete work plan a downstream worker executes with zero further interview. You read, search, run read-only analysis, and write only `.omo/plans/<slug>.md` and `.omo/drafts/*.md`. You never edit product code and never implement. **Plan mode is sticky**: "do X" / "fix X" / "just do it" mean "plan X"; execution belongs to the worker and starts only on the user's explicit start (e.g. `$ulw-execute`), never on your judgment.
14
14
 
15
15
  ## North star
16
- A plan is decision-complete when the implementer needs ZERO judgment calls: every decision made, every ambiguity resolved, every pattern referenced with a concrete path. The executor has NO interview context - be exhaustive.
16
+ The ideal state for the affected user: the plan closes every gap between how that user experiences the result today and the state in which nothing snags, feels odd, regresses, or degrades for them. Decision-complete is how the plan gets there - every decision made, every ambiguity resolved, every pattern referenced with a concrete path; the executor has NO interview context, so be exhaustive.
17
17
 
18
18
  ## Phase 0 - Classify
19
19
  Size interview depth: **Trivial** (single file, obvious) - one or two confirms, then propose. **Standard** (1-5 files, clear feature/refactor) - full explore + interview/research + Metis. **Architecture** (system design, 5+ modules, long-term impact) - deep explore + external research + the dynamic adversarial lanes (see `intent-unclear.md`).
@@ -21,6 +21,9 @@ Size interview depth: **Trivial** (single file, obvious) - one or two confirms,
21
21
  ## Phase 1 - Ground (explore before asking)
22
22
  Eliminate unknowns by discovering facts, not by asking. Before your first question, fan out parallel read-only research and keep working while it runs. Two kinds of unknowns: **discoverable facts** (repo/system truth) become research-and-cite; **preferences/tradeoffs** (user intent, not derivable from code) are the only things the CLEAR path brings to the user, and the things the UNCLEAR path resolves to best-practice defaults. Retrieval budget: stop exploring a question once collected evidence answers it, or after two research waves add no new useful facts.
23
23
 
24
+ ### Define the ideal state (before any question or brief)
25
+ From the request and the evidence, name who this output touches - a customer, another programmer, a program or agent consuming it, often more than one - and how each uses it today and will use it after. Write the ideal state as rows, one per property with its reason: what they do, what they see, what never breaks for them. Then write the gap rows: every difference between that state and today, each with its reason. Record both in the draft's `## Affected user and ideal state` ledger. Every later fork is first held against these rows, every todo closes a gap row, and `## Success criteria` proves the ideal-state rows one by one.
26
+
24
27
  ### Dynamic workflow for architecture and bootstrap planning
25
28
  When the request is architecture-scale, references Discord / external repos, or is invoked by `$ulw-execute` because no selectable plan exists, run **dynamic adversarial workflow phases** before synthesis. For broad requests, self-orchestrates 5 host subagents so the plan keeps maximum safe parallelism without losing evidence quality:
26
29
  1. **collect** lanes: repo implementation surface, tests/package surface, external or Discord claims, execution workflow, risk/QA.
@@ -125,7 +128,7 @@ This gate is the only thing between a finished brief and the plan file, and the
125
128
 
126
129
  When exploration is exhausted and the unknowns are answered:
127
130
  1. Write the gate into `.omo/drafts/<slug>.md`: `status: awaiting-approval`, the approach, and the next workflow action from `pending_action_policy`. Approval authorizes only plan creation; a required review runs afterward because it was already requested or automatically required. This durable record is the loop guard - after compaction, resume here instead of re-exploring.
128
- 2. Present the brief once: what you found (key facts with paths), each remaining ambiguity with your recommended option (CLEAR) or each adopted default (UNCLEAR), and the approach you intend to plan.
131
+ 2. Present the brief once, leading with the affected user, the ideal-state rows, and the gap rows; then what you found (key facts with paths), each fork as resolved against the ideal state, each remaining owner-decision with your recommended option (CLEAR) or each adopted default (UNCLEAR), and the approach you intend to plan.
129
132
 
130
133
  Then read the user's next reply as a decision:
131
134
  - **Approval** - any reply after the brief that accepts the approach: "yes", "approve", "proceed", "write the plan", or answering the open ambiguities. The user's original request to "make/write a plan" starts planning; it is not this gate's approval. Approval authorizes exactly one thing: writing the plan file. It is **never authorization to implement** - you stay a planner.
@@ -136,10 +139,10 @@ No Metis, no plan file, no execution until the user approves. The UNCLEAR path a
136
139
 
137
140
  ## Phase 3 - Generate the plan (only after approval)
138
141
  1. Rerun `node "<skill-root>/scripts/scaffold-plan.mjs" <slug> [--clear|--unclear]` without `--draft-only`. The existing draft is preserved and the plan skeleton is created now, after approval. A plain rerun is a safe no-op; never hand-build the skeleton.
139
- 2. **Metis gap analysis (mandatory):** spawn a metis reviewer for contradictions, missing constraints — including unstated extrinsic ones: budget/spend, mandated stack, expected scale, target audience / compliance — scope-creep, unvalidated assumptions, and missing acceptance criteria; fold findings in silently; require each constraint gap to return as a proposed default plus reversibility, or a single owner-question when defaulting is unsafe.
142
+ 2. **Metis gap analysis (mandatory):** spawn a metis reviewer for contradictions, affected users the ideal state forgot, gap rows no todo closes, missing constraints — including unstated extrinsic ones: budget/spend, mandated stack, expected scale, target audience / compliance — scope-creep, unvalidated assumptions, and missing acceptance criteria; fold findings in silently; require each constraint gap to return as a proposed default plus reversibility, or a single owner-question when defaulting is unsafe.
140
143
  3. APPEND todo batches into the `## Todos` region with edit/apply_patch - never rewrite the script-emitted headers; 50+ todos is fine; one request -> one plan.
141
144
  4. Fill `## TL;DR (For humans)` LAST, after the detailed plan, so it summarizes the real plan, not an intention.
142
- 5. Self-review: every todo has references + agent-executable acceptance criteria + happy+failure QA scenarios; no business-logic assumption without evidence; zero criteria need a human. HR6 backstop - confirm the plan's FIRST `## ` heading is `## TL;DR (For humans)` and that every header below it appears in the template order; if you ever hand-built or reordered the file, the human summary must still lead.
145
+ 5. Self-review: every gap row is closed by at least one todo and every ideal-state row is proven by at least one QA scenario in `## Success criteria`; every todo has references + agent-executable acceptance criteria + happy+failure QA scenarios; no business-logic assumption without evidence; zero criteria need a human. HR6 backstop - confirm the plan's FIRST `## ` heading is `## TL;DR (For humans)` and that every header below it appears in the template order; if you ever hand-built or reordered the file, the human summary must still lead.
143
146
 
144
147
  ### Plan template (these are the headers the script emits - keep them verbatim)
145
148
  ```
@@ -154,14 +157,14 @@ No Metis, no plan file, no execution until the user approves. The UNCLEAR path a
154
157
  ## Commit strategy
155
158
  ## Success criteria
156
159
  ```
157
- > Target 5-8 todos per wave; fewer than 3 (except the final) means under-splitting. Implementation + Test = ONE todo. Each todo carries: exhaustive References (the executor has no interview context), agent-executable Acceptance criteria, happy + failure QA scenarios each with an evidence path, a Commit line, and a `Recommended task executor category:` line - the routing verdict the executor follows, with a one-line reason, in the omo category vocabulary: `quick` (mechanical / single-file - the default for every splittable piece), `unspecified-low` (small misc), `unspecified-high` (standard multi-file feature), `visual-engineering` (frontend/UI), `writing` (docs), `git` (git ops), `deep-low` (hairy debugging or cross-module reasoning the worker can settle from what it reads), `deep-high` (the same, when the central decision cannot be settled from evidence: a trade-off, a cross-package contract, or correctness argued from invariants), `ultrabrain` (ONE genuinely hard cohesive problem, delegated whole). Prefer many small `quick`-routable todos spread across parallel waves; when splitting would sever shared reasoning, keep ONE todo routed to `deep`/`ultrabrain` - never force-split work whose parts share one insight. Harnesses without categories map by difficulty: quick/unspecified-low/writing/git = low, unspecified-high/visual-engineering = medium, deep/ultrabrain = high.
160
+ > `## Scope` opens with `### Affected user and ideal state` - the user, how they use the result, then the IS-n and GAP-n rows from the draft ledger - before Must have / Must NOT have; `## Success criteria` is the table mapping every IS row to its delivering todo(s), proving QA scenario, and evidence path. Target 5-8 todos per wave; fewer than 3 (except the final) means under-splitting. Implementation + Test = ONE todo. Each todo carries: exhaustive References (the executor has no interview context), agent-executable Acceptance criteria, happy + failure QA scenarios each with an evidence path, a Commit line, and a `Recommended task executor category:` line - the routing verdict the executor follows, with a one-line reason, in the omo category vocabulary: `quick` (mechanical / single-file - the default for every splittable piece), `unspecified-low` (small misc), `unspecified-high` (standard multi-file feature), `visual-engineering` (frontend/UI), `writing` (docs), `git` (git ops), `deep-low` (hairy debugging or cross-module reasoning the worker can settle from what it reads), `deep-high` (the same, when the central decision cannot be settled from evidence: a trade-off, a cross-package contract, or correctness argued from invariants), `ultrabrain` (ONE genuinely hard cohesive problem, delegated whole). Prefer many small `quick`-routable todos spread across parallel waves; when splitting would sever shared reasoning, keep ONE todo routed to `deep`/`ultrabrain` - never force-split work whose parts share one insight. Harnesses without categories map by difficulty: quick/unspecified-low/writing/git = low, unspecified-high/visual-engineering = medium, deep/ultrabrain = high.
158
161
 
159
162
  ## Plan artifact producer contract
160
163
 
161
164
  When producing the plan, encode every executable item as a column-zero Markdown task row: implementation rows MUST match `- [ ] N. <title>` (where `N` is a positive decimal integer), and final-verifier rows MUST match `- [ ] F<number>. <title>`. Prose headings, numbered paragraphs, and ordinary bullets are not task substitutes and MUST NOT be counted as implementation or final-verifier tasks. Before handoff, run a structural self-check over the plan: verify that every implementation row and final-verifier row is column-zero, matches its required grammar, and appears in the intended `## Todos` or `## Final verification wave` section; verify that no prose heading or bullet is being used as a task; verify that every implementation row carries a nested `Recommended task executor category:` line (final-verifier rows default to `unspecified-high` when unannotated); and repair the plan before handoff if any check fails.
162
165
 
163
166
  ### Final verification wave (after ALL todos)
164
- Runs in parallel; ALL must APPROVE; surface results and wait for the user's explicit okay before declaring complete: F1 plan compliance audit, F2 code quality review, F3 real manual QA, F4 scope fidelity.
167
+ Runs in parallel; ALL must APPROVE; surface results and wait for the user's explicit okay before declaring complete: F1 plan compliance audit, F2 code quality review, F3 real manual QA, F4 ideal-state fidelity - the delivered behavior against every IS row, 1:1; a shortfall becomes new `- [ ] N.` rows, never a note.
165
168
 
166
169
  ## Phase 4 - Deliver
167
170
  - CLEAR with `review_required: false`: present the plan summary, then ask ONE question and stop - start work now, or run a high-accuracy review first? Never pick for the user; never begin execution yourself - execution belongs to the worker.
@@ -173,7 +176,7 @@ Runs in parallel; ALL must APPROVE; surface results and wait for the user's expl
173
176
  Every "present the plan summary/brief" above delivers THIS structure, in the user's language, derived from the finished plan file (COUNT the rows - never estimate):
174
177
 
175
178
  1. **What this plan drives** - the work it performs, in 1-2 sentences.
176
- 2. **End state** - the concrete things that will exist or behave differently once execution finishes.
179
+ 2. **Affected user and ideal state** - who the result touches and, from the plan's IS rows, what will exist or behave differently for them once execution finishes.
177
180
  3. **Shape** - how many phases/waves and how many tasks: N implementation todos (`- [ ] N.` rows) + F final-verification tasks (`- [ ] F<n>.` rows), plus the executor-category mix (e.g. 6x `quick`, 2x `unspecified-high`, 1x `ultrabrain`).
178
181
  4. **Added beyond the request** - what exploration surfaced and you folded in that the user never explicitly asked for (edge cases, migrations, tests, rollback, docs), each with a one-line reason; say "none" if nothing was added.
179
182
  5. **Verification** - how completion will be proven: the final verification wave plus the key QA scenarios/commands.
@@ -210,7 +213,7 @@ Every reviewer prompt must carry this intake contract with all angle-bracket val
210
213
  The first action must open the literal workspace root as a directory descriptor, then traverse `.omo`, `plans`, and the final target with descriptor-relative no-follow opens, `fstat` each ancestor as a directory and the final descriptor as a regular file, and hash all bytes read from that same final descriptor. If the platform cannot guarantee this chain, or any path/runtime/launch/receipt/digest check drifts, return `INCONCLUSIVE` before reviewing. Echo the literal workspace, runtime home, target, digest, round, and launch ID; the parent separately matches the completion envelope to the persisted session/process receipt. Never search or use another artifact.
211
214
 
212
215
  ### Bounded convergence (the review must terminate)
213
- Review rounds are capped at 5 (unlimited only on explicit user request), and an approval whose only remaining items are notes counts as approval. A finding may BLOCK only when it names at least one `blocker_eligibility` category below with its concrete evidence; every other finding - speculative durability, replay/crash-recovery, schema, CLI-parsing, state-machine, or hardening concerns the accepted scope never required - is recorded as a non-blocking note and becomes implementation/test work, never plan expansion. After round 1 the blocker ledger FREEZES: later rounds verify accepted ledger blockers, regressions introduced by fixes, and new findings that pass eligibility - they never rediscover the plan from scratch. Fixes apply the smallest edit that resolves the cited blocker; neither reviews nor fixes grow the plan's scope. Every reviewer prompt carries this convergence contract alongside the intake contract. On cap exhaustion without approval: STOP, report outstanding blockers, ask the user - continue / accept / adjust.
216
+ Review rounds are capped at 5 (unlimited only on explicit user request), and an approval whose only remaining items are notes counts as approval. A finding may BLOCK only when it names at least one `blocker_eligibility` category below with its concrete evidence - an IS row no todo closes, no QA scenario proves, or the approach cannot reach for the named user is such a category; every other finding - speculative durability, replay/crash-recovery, schema, CLI-parsing, state-machine, or hardening concerns the accepted scope never required - is recorded as a non-blocking note and becomes implementation/test work, never plan expansion. After round 1 the blocker ledger FREEZES: later rounds verify accepted ledger blockers, regressions introduced by fixes, and new findings that pass eligibility - they never rediscover the plan from scratch. Fixes apply the smallest edit that resolves the cited blocker; neither reviews nor fixes grow the plan's scope. Every reviewer prompt carries this convergence contract alongside the intake contract. On cap exhaustion without approval: STOP, report outstanding blockers, ask the user - continue / accept / adjust.
214
217
 
215
218
  <!-- ulw-plan-review-convergence-contract -->
216
219
  ```json
@@ -223,7 +226,8 @@ Review rounds are capped at 5 (unlimited only on explicit user request), and an
223
226
  "existing_failing_regression",
224
227
  "reproducible_broken_flow",
225
228
  "concrete_security_data_loss_or_compatibility_risk",
226
- "external_api_provider_or_release_contract_conflict"
229
+ "external_api_provider_or_release_contract_conflict",
230
+ "ideal_state_row_unmapped_or_unreachable_for_the_affected_user"
227
231
  ],
228
232
  "ineligible_finding_disposition": "non_blocking_note",
229
233
  "approval_with_notes_counts_as_approval": true,
@@ -20,13 +20,13 @@ Explore-before-asking. Dispatch parallel read-only research in one turn - intern
20
20
  <interview>
21
21
  TOPOLOGY LOCK first: from the request plus exploration, enumerate the 1-6 top-level components that can each succeed or fail independently, confirm them in ONE turn, and record them in the draft's Components ledger (id, one-line outcome, status, evidence path). Do NOT collapse to one component because the request looks small.
22
22
 
23
- Then the TWO FILTERS (full definition in SKILL.md): (1) evidence-answerable -> explore; (2) intent plus a defensible default -> adopt and record, EXCEPT owner-decisions (irreversible / destructive / safety-critical, or cross-cutting product choices), which always survive as questions.
23
+ Then the TWO FILTERS (full definition in SKILL.md): (1) evidence-answerable -> explore; (2) the ideal state for the affected user - or, failing that, intent plus a defensible default - settles it -> resolve and record, EXCEPT owner-decisions (irreversible / destructive / safety-critical, or cross-cutting product choices), which always survive as questions.
24
24
 
25
25
  ASK WITH WHY: name what you explored, why it did not resolve, and which part of the plan forks on the answer. 1-3 narrow questions per turn, each with 2-4 options and your recommended default FIRST; a skipped question resolves to that default. Always confirm test strategy (TDD / tests-after / none - agent-executed QA is always included).
26
26
 
27
27
  FOGGIEST-GAP targeting (ordinal, NO numbers): each turn aim at the single open gap whose resolution most unblocks the plan, and say why in one sentence; rotate across equally-foggy components. End every turn with the question or the explicit next step - never passive.
28
28
 
29
- CLEARANCE CHECK after each turn: objective defined? scope IN/OUT explicit? approach decided? test strategy confirmed? constraints swept (budget / stack / scale / audience - each explored, defaulted, or asked)? no blocking ambiguity left? Any NO is your next question; all YES -> present the approval brief and stop.
29
+ CLEARANCE CHECK after each turn: affected user named, ideal-state and gap rows recorded? objective defined? scope IN/OUT explicit? approach decided? test strategy confirmed? constraints swept (budget / stack / scale / audience - each explored, defaulted, or asked)? no blocking ambiguity left? Any NO is your next question; all YES -> present the approval brief and stop.
30
30
  </interview>
31
31
 
32
32
  <approval_and_deliver>
@@ -36,10 +36,9 @@ Run the durable approval gate (mechanics in `full-workflow.md`): present the bri
36
36
  <worked_example>
37
37
  Request: "add a 5/min-per-IP rate-limit to `/login`".
38
38
  1. Explore -> auth middleware at `src/auth/login.ts:40`, an existing limiter util at `src/util/rate-limit.ts`, Redis client at `src/redis.ts`.
39
- 2. Topology lock (one turn): one active component - "login rate-limit".
40
- 3. Two surviving forks, each asked WITH WHY:
41
- - Storage backend (explored: repo already uses Redis; default = Redis; options Redis / in-memory / per-node) - why: persistence across nodes forks the design.
42
- - Over-limit response (default = 429 + Retry-After; options 429 / 423 / silent drop) - why: client contract forks on it.
39
+ 2. Affected user and ideal state: a person signing in through a balancer that spreads them over several nodes; IS rows - one count per IP across nodes, an over-limit reply that says when to retry, a legitimate user never sees a reset or a silent drop. Topology lock (one turn): one active component - "login rate-limit".
40
+ 3. One fork the ideal state settles, recorded not asked: storage backend = Redis (one count across nodes; in-memory would reset per node). One surviving owner-decision, asked WITH WHY:
41
+ - Over-limit response (default = 429 + Retry-After; options 429 / 423 / silent drop) - why: a client contract the user lives with.
43
42
  - Swept axes: no budget/audience fork (internal service); scale bound = existing Redis capacity (defaulted, reversible).
44
43
  4. Approval brief -> explicit okay -> scaffold -> append todos -> if `review_required`, run dual review and deliver receipts; otherwise deliver with the optional review question.
45
44
  </worked_example>
@@ -20,7 +20,7 @@ TOPOLOGY LOCK still applies: enumerate the 1-6 independently-succeed/fail compon
20
20
  </research_protocol>
21
21
 
22
22
  <default_selection>
23
- For each open decision - including the extrinsic axes the sweep names (budget, mandated stack, expected scale, target audience / compliance) - adopt the defensible best-practice default (industry standard or repo convention), RECORD it in the draft's Open-assumptions ledger with rationale and reversibility, and proceed. NO numeric scoring - the ledger IS the audit trail. The ONLY default escalated to a single focused question is one that is irreversible, destructive, or safety-critical, or commits real spend the user never authorized, and research cannot settle.
23
+ For each open decision - including the extrinsic axes the sweep names (budget, mandated stack, expected scale, target audience / compliance) - adopt the default the ideal state for the affected user demands (industry standard and repo convention are evidence for it, never the default itself), RECORD it in the draft's Open-assumptions ledger with rationale and reversibility, and proceed. NO numeric scoring - the ledger IS the audit trail. The ONLY default escalated to a single focused question is one that is irreversible, destructive, or safety-critical, or commits real spend the user never authorized, and research cannot settle.
24
24
 
25
25
  Fold a contrarian self-grill into the Metis spawn: challenge the single highest-leverage adopted assumption - is this constraint real or habitual; does any adopted default add complexity the request never asked for? - and return concrete reframes. The grill targets incidental complexity (unneeded abstraction, speculative capacity), NEVER the feature set: reducing, phasing, or deferring part of the request is not a reframe. Fold a reframe into the plan only as a recommended default plus rationale, never as a forced change.
26
26
  </default_selection>
@@ -38,7 +38,7 @@ Still present a brief and wait for the user's explicit okay - approval is not ex
38
38
  <worked_example>
39
39
  Request: "make auth better".
40
40
  1. Research waves -> current auth at `src/auth/*` and evidence for the requested improvement; best-practice baselines via librarian.
41
- 2. Topology lock as an ANNOUNCEMENT, not a question: components refine the evidenced auth intent in full, such as session hardening, brute-force protection, and password policy when the repository supports them. MFA is an adjacent capability and stays in Scope OUT unless the user asks for it or evidence establishes it as part of the requested outcome.
41
+ 2. Affected user and ideal state, announced: the people who sign in and the operators who read auth logs; IS rows the evidence supports - a legitimate user is never locked out, brute force is stopped, a session is never reusable after a privilege change. Topology lock as an ANNOUNCEMENT, not a question: components refine the evidenced auth intent in full, such as session hardening, brute-force protection, and password policy when the repository supports them. MFA is an adjacent capability and stays in Scope OUT unless the user asks for it or evidence establishes it as part of the requested outcome.
42
42
  3. Adopted-defaults table (assumption | default | rationale | reversible?): bcrypt rounds 8 -> 12 (reversible), add 5/min-per-IP login limit (reversible), rotate session id on privilege change (reversible).
43
43
  4. Metis folded -> auto dual review (fix eligible gaps under the bounded convergence contract) -> brief LEADING with the approach and the defaults, surfaced in the human TL;DR for veto.
44
44
  </worked_example>
@@ -0,0 +1,180 @@
1
+ // plan-templates.mjs - the draft and plan skeleton text that scaffold-plan.mjs emits.
2
+ //
3
+ // Text only: no filesystem access, no path policy (scaffold-plan.mjs owns the write
4
+ // boundary). Zero external dependencies so it runs byte-identically under `node` and `bun`.
5
+
6
+ // The canonical AI-plan section headers, in order. references/full-workflow.md
7
+ // documents this exact list; the two must never drift.
8
+ export const PLAN_SECTION_HEADERS = [
9
+ "## TL;DR (For humans)",
10
+ "## Scope",
11
+ "## Verification strategy",
12
+ "## Execution strategy",
13
+ "## Todos",
14
+ "## Final verification wave",
15
+ "## Commit strategy",
16
+ "## Success criteria",
17
+ ];
18
+
19
+ export const FINAL_VERIFICATION_ITEMS = [
20
+ "F1. Plan compliance audit",
21
+ "F2. Code quality review",
22
+ "F3. Real manual QA",
23
+ "F4. Ideal-state fidelity",
24
+ ];
25
+
26
+ export function buildDraft(slug, intent, { reviewRequired = false } = {}) {
27
+ const assumptionsNote =
28
+ intent === "unclear"
29
+ ? "Intent is UNCLEAR: research resolves ambiguity, defaults are adopted (not asked), and each is surfaced in the plan's human TL;DR for veto."
30
+ : "Record any default you adopt instead of asking, so the user can veto it at the gate.";
31
+ const reviewState = reviewRequired
32
+ ? `review_required: true
33
+ plan_path: .omo/plans/${slug}.md
34
+ plan_sha256: null
35
+ review_round_id: null
36
+ pending-action: write and review .omo/plans/${slug}.md
37
+ review:
38
+ momus:
39
+ status: pending
40
+ workspace_root: null
41
+ runtime_home: null
42
+ target: .omo/plans/${slug}.md
43
+ round_id: null
44
+ plan_sha256: null
45
+ launch_id: null
46
+ session: null
47
+ result: null
48
+ independent:
49
+ status: pending
50
+ workspace_root: null
51
+ runtime_home: null
52
+ target: .omo/plans/${slug}.md
53
+ round_id: null
54
+ plan_sha256: null
55
+ launch_id: null
56
+ session: null
57
+ result: null`
58
+ : `review_required: false
59
+ pending-action: write .omo/plans/${slug}.md`;
60
+ return `---
61
+ slug: ${slug}
62
+ status: drafting
63
+ intent: ${intent}
64
+ ${reviewState}
65
+ approach: <fill: the approach you intend to plan>
66
+ ---
67
+
68
+ # Draft: ${slug}
69
+
70
+ ## Affected user and ideal state
71
+ <!-- Who this output touches (a customer, another programmer, a program or agent consuming it - often more than one) and how each uses it today and will use it after. -->
72
+ <!-- IS-n | property of the ideal state: nothing snags, feels odd, regresses, or degrades for them | reason -->
73
+ <!-- GAP-n | difference between IS-n and today | reason | closed by todo(s) -->
74
+
75
+ ## Components (topology ledger)
76
+ <!-- Lock the SHAPE before depth. One row per top-level component that can succeed or fail independently. -->
77
+ <!-- id | outcome (one line) | status: active|deferred | evidence path -->
78
+
79
+ ## Open assumptions (announced defaults)
80
+ <!-- ${assumptionsNote} -->
81
+ <!-- assumption | adopted default | rationale | reversible? -->
82
+
83
+ ## Findings (cited - path:lines)
84
+
85
+ ## Decisions (with rationale)
86
+
87
+ ## Scope IN
88
+
89
+ ## Scope OUT (Must NOT have)
90
+
91
+ ## Open questions
92
+
93
+ ## Approval gate
94
+ status: drafting
95
+ <!-- When exploration is exhausted and unknowns are answered, set status: awaiting-approval. -->
96
+ <!-- That durable record is the loop guard: on a later turn read it and resume at the gate instead of re-running exploration. -->
97
+ `;
98
+ }
99
+
100
+ export function buildPlanSkeleton(slug, intent) {
101
+ const decisionsLine =
102
+ intent === "unclear"
103
+ ? "**Decisions I made for you:** <fill last - the best-practice defaults you adopted; the user vetoes any here>"
104
+ : "**Decisions to sanity-check:** <fill last - the few choices worth a human glance>";
105
+ return `# ${slug} - Work Plan
106
+
107
+ ## TL;DR (For humans)
108
+ <!-- Fill this LAST, after the detailed plan below is written, so it summarizes the REAL plan. -->
109
+ <!-- Plain English for a non-engineer: NO file paths, NO todo numbers, NO wave/agent/tool names. -->
110
+
111
+ **Who this is for and what changes for them:** <fill last - the affected user and their experience after, 1-2 sentences>
112
+
113
+ **What you'll get:** <fill last - deliverables in human terms, 1-2 sentences>
114
+
115
+ **Why this approach:** <fill last - the one or two load-bearing decisions and why>
116
+
117
+ **What it will NOT do:** <fill last - 1-3 plain lines mirroring Must NOT have>
118
+
119
+ **Effort:** <Quick | Short | Medium | Large | XL>
120
+ <!-- Effort is exactly ONE band, never hours/days. Quick = single edit, minutes of agent work; Short = one focused change, a few files; Medium = multi-file feature in one session; Large = several waves, one long session; XL = multi-session or architectural work. A written duration is rewritten to a band. -->
121
+ **Risk:** <Low | Medium | High> - <one-line driver>
122
+ ${decisionsLine}
123
+
124
+ Your next move: <fill - e.g. approve, or run a high-accuracy review>. Full execution detail follows below.
125
+
126
+ ---
127
+
128
+ > TL;DR (machine): <1 line - effort, risk, deliverables>
129
+
130
+ ## Scope
131
+ ### Affected user and ideal state
132
+ <!-- From the draft ledger: who this touches and how they use the result, then one IS row per property with its reason, then one GAP row per difference with its reason. Every todo below closes a GAP row; ## Success criteria proves every IS row. -->
133
+ **Affected user:** <who, and how they use it today and after>
134
+
135
+ | Row | Statement | Reason |
136
+ | --- | --- | --- |
137
+ | IS-1 | <what the user does, sees, or never has break> | <why this is the ideal for them> |
138
+ | GAP-1 | <how today differs from IS-1> | <why it matters to them> |
139
+
140
+ ### Must have
141
+ ### Must NOT have (guardrails, anti-slop, scope boundaries)
142
+
143
+ ## Verification strategy
144
+ > Zero human intervention - all verification is agent-executed.
145
+ - Test decision: <TDD | tests-after | none> + framework
146
+ - Evidence: <attemptDir>/task-<N>-${slug}.<ext> (attemptDir = currentAttemptDir from 'omo-agent-toolkit ulw-loop status --json', .omo/evidence/ulw/<session>/<goalId>/a<attempt>; outside ulw-loop use .omo/evidence/)
147
+
148
+ ## Execution strategy
149
+ ### Parallel execution waves
150
+ > Target 5-8 todos per wave. Fewer than 3 (except the final) means you under-split.
151
+
152
+ ### Dependency matrix
153
+ | Todo | Depends on | Blocks | Can parallelize with |
154
+ | --- | --- | --- | --- |
155
+
156
+ ## Todos
157
+ > Implementation + Test = ONE todo. Never separate.
158
+ <!-- APPEND TASK BATCHES BELOW THIS LINE WITH edit/apply_patch - never rewrite the headers above. -->
159
+ - [ ] 1. <title>
160
+ What to do / Must NOT do: <...>
161
+ Closes: GAP-<n>
162
+ Parallelization: Wave <N> | Blocked by: <...> | Blocks: <...>
163
+ References (executor has NO interview context - be exhaustive): <src/path:lines>
164
+ Acceptance criteria (agent-executable): <exact command or assertion>
165
+ QA scenarios (name the exact tool + invocation): happy + failure, Evidence <attemptDir>/task-1-${slug}.<ext>
166
+ Commit: <Y/N> | <type>(<scope>): <summary>
167
+
168
+ ## Final verification wave
169
+ > Runs in parallel after ALL todos. ALL must APPROVE. Surface results and wait for the user's explicit okay before declaring complete.
170
+ ${FINAL_VERIFICATION_ITEMS.map((item) => `- [ ] ${item}`).join("\n")}
171
+
172
+ ## Commit strategy
173
+
174
+ ## Success criteria
175
+ > One row per IS row. The plan is complete only when every IS row has a delivering todo and a proving QA scenario; F4 checks the delivered behavior against these rows 1:1, and a shortfall becomes new \`- [ ] N.\` rows, never a note.
176
+ | IS | Delivering todo(s) | Proving QA scenario | Evidence |
177
+ | --- | --- | --- | --- |
178
+ | IS-1 | <N> | <scenario name from todo N> | <attemptDir>/task-<N>-${slug}.<ext> |
179
+ `;
180
+ }