oh-my-opencode 5.0.0-beta.79 → 5.0.0-beta.80

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (96) hide show
  1. package/.agents/command/remove-deadcode.md +1 -1
  2. package/.agents/skills/hyperplan/SKILL.md +5 -5
  3. package/.agents/skills/remove-deadcode/SKILL.md +1 -1
  4. package/.agents/skills/security-research/SKILL.md +4 -4
  5. package/dist/cli/index.js +499 -336
  6. package/dist/cli-node/index.js +499 -336
  7. package/dist/config-migration/category-deep-split.d.ts +3 -0
  8. package/dist/config-migration/migration-plans.d.ts +1 -0
  9. package/dist/features/builtin-commands/templates/hyperplan.d.ts +1 -1
  10. package/dist/features/builtin-commands/templates/refactor-sections/team-mode-addendum.d.ts +1 -1
  11. package/dist/features/builtin-commands/templates/remove-ai-slops.d.ts +1 -1
  12. package/dist/index.js +282 -123
  13. package/dist/plugin-config/omo-config-chain.d.ts +1 -0
  14. package/dist/skills/debugging/references/methodology/02-investigate.md +5 -5
  15. package/dist/skills/refactor/SKILL.md +2 -2
  16. package/dist/skills/remove-ai-slops/SKILL.md +4 -4
  17. package/dist/skills/ulw-execute/SKILL.md +3 -2
  18. package/dist/skills/ulw-plan/references/full-workflow.md +1 -1
  19. package/dist/skills/ulw-research/SKILL.md +2 -2
  20. package/dist/tools/delegate-task/openai-categories.d.ts +6 -4
  21. package/dist/tui.js +146 -71
  22. package/package.json +14 -14
  23. package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
  24. package/packages/omo-codex/plugin/components/bootstrap/dist/cli.js +77 -8
  25. package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
  26. package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
  27. package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
  28. package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
  29. package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
  30. package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
  31. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
  32. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
  33. package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +2 -2
  34. package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
  35. package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
  36. package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
  37. package/packages/omo-codex/plugin/components/rules/package.json +1 -1
  38. package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
  39. package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
  40. package/packages/omo-codex/plugin/components/telemetry/dist/cli.js +63 -17
  41. package/packages/omo-codex/plugin/components/telemetry/dist/posthog.js +63 -17
  42. package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
  43. package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
  44. package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
  45. package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
  46. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/full-workflow.md +1 -1
  47. package/packages/omo-codex/plugin/components/ulw-execute-continuation/hooks/hooks.json +1 -1
  48. package/packages/omo-codex/plugin/components/ulw-execute-continuation/package.json +1 -1
  49. package/packages/omo-codex/plugin/components/ulw-loop/dist/checkpoint-template.js +3 -3
  50. package/packages/omo-codex/plugin/components/ulw-loop/dist/cli.js +14 -6
  51. package/packages/omo-codex/plugin/components/ulw-loop/dist/codex-goal-instruction.js +2 -2
  52. package/packages/omo-codex/plugin/components/ulw-loop/dist/surface.js +9 -1
  53. package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +5 -5
  54. package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
  55. package/packages/omo-codex/plugin/components/ulw-loop/src/checkpoint-template.ts +3 -3
  56. package/packages/omo-codex/plugin/components/ulw-loop/src/codex-goal-instruction.ts +2 -2
  57. package/packages/omo-codex/plugin/components/ulw-loop/src/surface.ts +9 -1
  58. package/packages/omo-codex/plugin/components/ulw-loop/test/checkpoint-template.test.ts +1 -1
  59. package/packages/omo-codex/plugin/components/ulw-loop/test/codex-goal-instruction.test.ts +1 -1
  60. package/packages/omo-codex/plugin/components/ulw-loop/test/sdk-contract.test.ts +1 -1
  61. package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
  62. package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
  63. package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
  64. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
  65. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
  66. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
  67. package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
  68. package/packages/omo-codex/plugin/hooks/post-tool-use-recording-spawn-admission.json +1 -1
  69. package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
  70. package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
  71. package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
  72. package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
  73. package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
  74. package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
  75. package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
  76. package/packages/omo-codex/plugin/hooks/stop-checking-ulw-execute-continuation.json +1 -1
  77. package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
  78. package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
  79. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
  80. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
  81. package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
  82. package/packages/omo-codex/plugin/package-lock.json +12 -12
  83. package/packages/omo-codex/plugin/package.json +1 -1
  84. package/packages/omo-codex/plugin/skills/debugging/references/methodology/02-investigate.md +5 -5
  85. package/packages/omo-codex/plugin/skills/refactor/SKILL.md +2 -2
  86. package/packages/omo-codex/plugin/skills/remove-ai-slops/SKILL.md +4 -4
  87. package/packages/omo-codex/plugin/skills/ulw-execute/SKILL.md +3 -2
  88. package/packages/omo-codex/plugin/skills/ulw-plan/references/full-workflow.md +1 -1
  89. package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +2 -2
  90. package/packages/omo-codex/scripts/install-dist/install-local.mjs +167 -51
  91. package/packages/shared-skills/skills/debugging/references/methodology/02-investigate.md +5 -5
  92. package/packages/shared-skills/skills/refactor/SKILL.md +2 -2
  93. package/packages/shared-skills/skills/remove-ai-slops/SKILL.md +4 -4
  94. package/packages/shared-skills/skills/ulw-execute/SKILL.md +3 -2
  95. package/packages/shared-skills/skills/ulw-plan/references/full-workflow.md +1 -1
  96. package/packages/shared-skills/skills/ulw-research/SKILL.md +2 -2
@@ -1,12 +1,12 @@
1
1
  {
2
2
  "name": "@sisyphuslabs/omo-codex-plugin",
3
- "version": "5.0.0-beta.79",
3
+ "version": "5.0.0-beta.80",
4
4
  "lockfileVersion": 3,
5
5
  "requires": true,
6
6
  "packages": {
7
7
  "": {
8
8
  "name": "@sisyphuslabs/omo-codex-plugin",
9
- "version": "5.0.0-beta.79",
9
+ "version": "5.0.0-beta.80",
10
10
  "workspaces": [
11
11
  "components/comment-checker",
12
12
  "components/git-bash",
@@ -90,7 +90,7 @@
90
90
  },
91
91
  "components/comment-checker": {
92
92
  "name": "@code-yeongyu/codex-comment-checker",
93
- "version": "5.0.0-beta.79",
93
+ "version": "5.0.0-beta.80",
94
94
  "license": "MIT",
95
95
  "bin": {
96
96
  "omo-comment-checker": "dist/cli.js"
@@ -111,7 +111,7 @@
111
111
  },
112
112
  "components/git-bash": {
113
113
  "name": "@sisyphuslabs/codex-git-bash-hook",
114
- "version": "5.0.0-beta.79",
114
+ "version": "5.0.0-beta.80",
115
115
  "bin": {
116
116
  "omo-git-bash-hook": "dist/cli.js"
117
117
  },
@@ -125,7 +125,7 @@
125
125
  },
126
126
  "components/lazycodex-executor-verify": {
127
127
  "name": "@code-yeongyu/codex-lazycodex-executor-verify",
128
- "version": "5.0.0-beta.79",
128
+ "version": "5.0.0-beta.80",
129
129
  "license": "MIT",
130
130
  "bin": {
131
131
  "lazycodex-executor-verify": "dist/cli.js"
@@ -142,7 +142,7 @@
142
142
  },
143
143
  "components/lsp": {
144
144
  "name": "@code-yeongyu/codex-lsp",
145
- "version": "5.0.0-beta.79",
145
+ "version": "5.0.0-beta.80",
146
146
  "license": "MIT",
147
147
  "dependencies": {
148
148
  "@code-yeongyu/lsp-daemon": "file:../../../../lsp-daemon",
@@ -163,7 +163,7 @@
163
163
  },
164
164
  "components/rules": {
165
165
  "name": "@code-yeongyu/codex-rules",
166
- "version": "5.0.0-beta.79",
166
+ "version": "5.0.0-beta.80",
167
167
  "license": "MIT",
168
168
  "dependencies": {
169
169
  "picomatch": "^4.0.7"
@@ -185,7 +185,7 @@
185
185
  },
186
186
  "components/teammode": {
187
187
  "name": "@sisyphuslabs/codex-teammode",
188
- "version": "5.0.0-beta.79",
188
+ "version": "5.0.0-beta.80",
189
189
  "devDependencies": {
190
190
  "@types/node": "^26.2.0",
191
191
  "bun-types": "^1.4.2",
@@ -198,7 +198,7 @@
198
198
  },
199
199
  "components/telemetry": {
200
200
  "name": "@code-yeongyu/codex-telemetry",
201
- "version": "5.0.0-beta.79",
201
+ "version": "5.0.0-beta.80",
202
202
  "license": "MIT",
203
203
  "bin": {
204
204
  "omo-telemetry": "dist/cli.js"
@@ -216,7 +216,7 @@
216
216
  },
217
217
  "components/ultrawork": {
218
218
  "name": "@code-yeongyu/codex-ultrawork",
219
- "version": "5.0.0-beta.79",
219
+ "version": "5.0.0-beta.80",
220
220
  "license": "MIT",
221
221
  "bin": {
222
222
  "omo-ultrawork": "dist/cli.js"
@@ -234,7 +234,7 @@
234
234
  },
235
235
  "components/ulw-execute-continuation": {
236
236
  "name": "@code-yeongyu/codex-ulw-execute-continuation",
237
- "version": "5.0.0-beta.79",
237
+ "version": "5.0.0-beta.80",
238
238
  "license": "MIT",
239
239
  "bin": {
240
240
  "omo-ulw-execute-continuation": "dist/cli.js"
@@ -251,7 +251,7 @@
251
251
  },
252
252
  "components/ulw-loop": {
253
253
  "name": "@code-yeongyu/codex-ulw-loop",
254
- "version": "5.0.0-beta.79",
254
+ "version": "5.0.0-beta.80",
255
255
  "license": "MIT",
256
256
  "bin": {
257
257
  "omo-ulw-loop": "dist/cli.js",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@sisyphuslabs/omo-codex-plugin",
3
- "version": "5.0.0-beta.79",
3
+ "version": "5.0.0-beta.80",
4
4
  "description": "Aggregate Codex plugin root for OMO components.",
5
5
  "type": "module",
6
6
  "packageManager": "npm@11.12.1",
@@ -75,22 +75,22 @@ When the `team_*` tools are present, create a **debug-squad** team and split inv
75
75
  "members": [
76
76
  {
77
77
  "kind": "category",
78
- "category": "deep",
78
+ "category": "deep-low",
79
79
  "prompt": "You are the Runtime State Inspector. Your job: attach to the live process, hit breakpoints, read program state (variables, heap, goroutines, stack, registers depending on runtime), and report observed values verbatim. Never guess — if you don't see the value, say so. Report back via team_send_message with file:line / address references and captured values. Never edit source code. Never run git commands. If you need an instrumentation statement added (breakpoint(), debugger;, dbg!, etc.), ask the Lead first."
80
80
  },
81
81
  {
82
82
  "kind": "category",
83
- "category": "deep",
83
+ "category": "deep-low",
84
84
  "prompt": "You are the Log Archaeologist. Your job: grep server logs, stderr streams, SDK-internal debug output (DEBUG env, RUST_LOG, GODEBUG, PYTHONASYNCIODEBUG), and correlate timestamps. Produce a timeline of events with latencies. Flag anything that looks like a silent catch, a swallowed rejection, a panic recovered-and-ignored, a success response that contains failure signals (HTTP 200 with empty body, stopReason=error, exit 0 with error-in-stdout). Never edit source code."
85
85
  },
86
86
  {
87
87
  "kind": "category",
88
- "category": "deep",
88
+ "category": "deep-low",
89
89
  "prompt": "You are the Reproduction Engineer. Your job: build the smallest reliable repro — a curl command, a vitest/pytest/go test, a tmux script, a Playwright script for browser bugs, a pwntools script for binary targets. It must reproduce on first try and be copy-pasteable by the Lead. Document exact input, expected output, observed output. Save repro artifacts under /tmp/ and tell the Lead to journal them. If the bug is browser-based you MUST use Playwright CLI — do not simulate with curl."
90
90
  },
91
91
  {
92
92
  "kind": "category",
93
- "category": "deep",
93
+ "category": "deep-low",
94
94
  "prompt": "You are the Trace Correlator. Your job: take findings from the other members and cross-link them. Build a causal chain from symptom to suspected cause. Identify missing evidence. Propose the next single most-decisive runtime query. Never edit source code; only reason across already-captured evidence. If hypotheses diverge sharply after correlation, tell the Lead immediately — that is the signal for the Oracle Triple."
95
95
  }
96
96
  ]
@@ -117,7 +117,7 @@ task(subagent_type="explore", load_skills=[], run_in_background=true,
117
117
  Runtime state investigation for hypothesis 1: ...")
118
118
  task(subagent_type="explore", load_skills=[], run_in_background=true,
119
119
  prompt="Log/timing investigation for hypothesis 2: ...")
120
- task(category="deep", load_skills=[], run_in_background=true,
120
+ task(category="deep-low", load_skills=[], run_in_background=true,
121
121
  prompt="Reproduction minimizer for hypothesis 3: ...")
122
122
  ```
123
123
 
@@ -708,7 +708,7 @@ Record the chosen path in the TodoWrite list.
708
708
 
709
709
  Rationale for this composition:
710
710
  - **4 workers = team mode's parallel cap.** 5+ just queues.
711
- - **No verifier team member.** Verification needs \`deep\` reasoning (or \`unspecified-high\` fallback). In-team category routing downcasts to the category worker, which is weaker than required — the verifier runs OUTSIDE the team as a \`task(category="deep")\`.
711
+ - **No verifier team member.** Verification needs \`deep-high\` reasoning (or \`unspecified-high\` fallback). In-team category routing downcasts to the category worker, which is weaker than required — the verifier runs OUTSIDE the team as a \`task(category="deep-high")\`.
712
712
  - **quick × 2** for mechanical edits, **unspecified-low × 2** for reasoning edits — mirrors the plan's split.
713
713
 
714
714
  **Team lifecycle** (one team, reused until Phase 6 cleanup):
@@ -740,7 +740,7 @@ While any team task is \`pending | claimed | in_progress\`:
740
740
  - On a worker completion report, immediately dispatch an **external verifier** — verification runs OUTSIDE the team because team-member category routing downcasts to the category worker:
741
741
  \`\`\`
742
742
  task(
743
- category="deep",
743
+ category="deep-high",
744
744
  load_skills=[],
745
745
  run_in_background=true,
746
746
  description="verify step <N>",
@@ -184,9 +184,9 @@ File: src/bar.py
184
184
 
185
185
  Order rule (safest → riskiest): comments → dead code → defensive → duplication → complexity → abstraction/boundary → performance → tests → oversized-modules. This minimizes blast radius of any one change.
186
186
 
187
- ### Phase 4: Parallel slop removal via `deep` agents in batches of 5
187
+ ### Phase 4: Parallel slop removal via `deep-low` agents in batches of 5
188
188
 
189
- Files are processed by `deep` category agents with the `$omo:remove-ai-slops` skill loaded, **batched 5 at a time in parallel**. The executable skill name is `remove-ai-slops`. The `deep` category gives the agent enough thoroughness to correctly evaluate the 9 categories and respect the KEEP rules without slipping into surface fixes; the 5-wide batch is the sweet spot — more than 5 creates result-merging noise and context contention, fewer wastes parallelism.
189
+ Files are processed by `deep-low` category agents with the `$omo:remove-ai-slops` skill loaded, **batched 5 at a time in parallel**. The executable skill name is `remove-ai-slops`. The `deep-low` category gives the agent enough thoroughness to correctly evaluate the 9 categories and respect the KEEP rules without slipping into surface fixes; the 5-wide batch is the sweet spot — more than 5 creates result-merging noise and context contention, fewer wastes parallelism.
190
190
 
191
191
  **Batching protocol** (strict):
192
192
 
@@ -203,7 +203,7 @@ Files are processed by `deep` category agents with the `$omo:remove-ai-slops` sk
203
203
 
204
204
  ```
205
205
  task(
206
- category="deep",
206
+ category="deep-low",
207
207
  load_skills=["remove-ai-slops"],
208
208
  run_in_background=true,
209
209
  description="Slop removal: {filename}",
@@ -229,7 +229,7 @@ For each skipped issue, give reason.
229
229
  )
230
230
  ```
231
231
 
232
- **Batch failure handling**: a `multi_agent_v1.wait_agent` timeout only means no new mailbox update arrived, not that a `deep` agent failed. For long passes, require each child to send `WORKING: <file> - <current phase>` and `BLOCKED: <reason>` only when it cannot progress. Treat a running child as alive. Mark a file for retry only when the child is completed without the deliverable, ack-only after followup, explicitly `BLOCKED:`, or no longer running. Do NOT block the remaining 4 in that batch; collect successful results and retry the failed file once later. If retry also fails, escalate that file under "Issues Found & Fixed" in the final report.
232
+ **Batch failure handling**: a `multi_agent_v1.wait_agent` timeout only means no new mailbox update arrived, not that a `deep-low` agent failed. For long passes, require each child to send `WORKING: <file> - <current phase>` and `BLOCKED: <reason>` only when it cannot progress. Treat a running child as alive. Mark a file for retry only when the child is completed without the deliverable, ack-only after followup, explicitly `BLOCKED:`, or no longer running. Do NOT block the remaining 4 in that batch; collect successful results and retry the failed file once later. If retry also fails, escalate that file under "Issues Found & Fixed" in the final report.
233
233
 
234
234
  ### Phase 5: Verify with quality gates + critical review
235
235
 
@@ -152,13 +152,14 @@ When the plan annotates a todo with `Recommended task executor category:`, follo
152
152
  | `visual-engineering` (medium) | frontend, UI/UX, styling, animation |
153
153
  | `writing` (low) | documentation and prose |
154
154
  | `git` (low) | git operations |
155
- | `deep` (high) | hairy debugging, research-heavy or subtle cross-module work |
155
+ | `deep-low` (medium) | hairy debugging, research-heavy or subtle cross-module work the worker can settle from what it reads |
156
+ | `deep-high` (high) | the same, when the central decision cannot be settled from evidence: a trade-off, a cross-package contract, or correctness argued from invariants |
156
157
  | `ultrabrain` (high) | ONE genuinely hard, logic-heavy problem — hand it the goal, not step-by-step instructions |
157
158
 
158
159
  Sizing is a two-branch decision made per checkbox, before dispatch:
159
160
 
160
161
  - **Splittable work splits.** When the checkbox decomposes into independent pieces, dispatch them as a swarm of `quick`/`unspecified-low` workers in ONE parallel burst — many small cheap workers in parallel beat one large delegation.
161
- - **Cohesive hard work stays whole.** When splitting would sever shared reasoning (one algorithm, one migration, one subtle bug), send the WHOLE problem to `deep` or `ultrabrain` as ONE delegation. Never force-split work whose parts share one insight.
162
+ - **Cohesive hard work stays whole.** When splitting would sever shared reasoning (one algorithm, one migration, one subtle bug), send the WHOLE problem to `deep-low`, `deep-high` or `ultrabrain` as ONE delegation. Never force-split work whose parts share one insight.
162
163
 
163
164
  Each sub-task message must include:
164
165
 
@@ -154,7 +154,7 @@ No Metis, no plan file, no execution until the user approves. The UNCLEAR path a
154
154
  ## Commit strategy
155
155
  ## Success criteria
156
156
  ```
157
- > Target 5-8 todos per wave; fewer than 3 (except the final) means under-splitting. Implementation + Test = ONE todo. Each todo carries: exhaustive References (the executor has no interview context), agent-executable Acceptance criteria, happy + failure QA scenarios each with an evidence path, a Commit line, and a `Recommended task executor category:` line - the routing verdict the executor follows, with a one-line reason, in the omo category vocabulary: `quick` (mechanical / single-file - the default for every splittable piece), `unspecified-low` (small misc), `unspecified-high` (standard multi-file feature), `visual-engineering` (frontend/UI), `writing` (docs), `git` (git ops), `deep` (hairy debugging or cross-module reasoning), `ultrabrain` (ONE genuinely hard cohesive problem, delegated whole). Prefer many small `quick`-routable todos spread across parallel waves; when splitting would sever shared reasoning, keep ONE todo routed to `deep`/`ultrabrain` - never force-split work whose parts share one insight. Harnesses without categories map by difficulty: quick/unspecified-low/writing/git = low, unspecified-high/visual-engineering = medium, deep/ultrabrain = high.
157
+ > Target 5-8 todos per wave; fewer than 3 (except the final) means under-splitting. Implementation + Test = ONE todo. Each todo carries: exhaustive References (the executor has no interview context), agent-executable Acceptance criteria, happy + failure QA scenarios each with an evidence path, a Commit line, and a `Recommended task executor category:` line - the routing verdict the executor follows, with a one-line reason, in the omo category vocabulary: `quick` (mechanical / single-file - the default for every splittable piece), `unspecified-low` (small misc), `unspecified-high` (standard multi-file feature), `visual-engineering` (frontend/UI), `writing` (docs), `git` (git ops), `deep-low` (hairy debugging or cross-module reasoning the worker can settle from what it reads), `deep-high` (the same, when the central decision cannot be settled from evidence: a trade-off, a cross-package contract, or correctness argued from invariants), `ultrabrain` (ONE genuinely hard cohesive problem, delegated whole). Prefer many small `quick`-routable todos spread across parallel waves; when splitting would sever shared reasoning, keep ONE todo routed to `deep`/`ultrabrain` - never force-split work whose parts share one insight. Harnesses without categories map by difficulty: quick/unspecified-low/writing/git = low, unspecified-high/visual-engineering = medium, deep/ultrabrain = high.
158
158
 
159
159
  ## Plan artifact producer contract
160
160
 
@@ -255,7 +255,7 @@ Interest alone is not a trigger. Anything without one stays a queued lead in `ex
255
255
  Settle with executed code, not judgment, whenever sources disagree, a behavior is undocumented, a claim is performance- or compatibility-shaped, or the honest answer is "it should work". Spawn one verification worker per claim:
256
256
 
257
257
  ```
258
- task(category="deep", run_in_background=true, prompt="TASK: verify by execution: <claim>.
258
+ task(category="deep-low", run_in_background=true, prompt="TASK: verify by execution: <claim>.
259
259
  SOURCE: <where it came from>; CONTRADICTION: <opposing source, if any>.
260
260
  Write a minimal self-contained script that tests the claim; run it (uv run --with <deps> python / bun / direct compile); capture full stdout+stderr; pin versions.
261
261
  Reply with: the exact code, the full output, environment (OS, runtime, dependency versions), and a verdict — CONFIRMED / REFUTED / PARTIAL — grounded in the output.")
@@ -335,7 +335,7 @@ Asset workers (background, parallel, each fed `design-spec.md`) — visuals are
335
335
 
336
336
  **Verify the asset manifest before rendering.** List every asset the document references, assert each file exists and is non-empty on disk, and re-render whatever is missing. A document that renders with three broken diagrams is a document you will publish twice.
337
337
 
338
- Assembly worker — `task(category="deep", load_skills=["frontend", "visual-qa", "open-design", "data-scientist", "imagegen", "ulw-loop"], run_in_background=true, ...)`: before writing, read every available design and visualization skill and apply it — the report is a designed artifact, not a text dump; the worker's prompt carries `design-spec.md`. Use the template the user approved; absent a stronger house style the default skeleton is executive summary → key findings by theme → detailed analysis (quotes under 20 words with attribution, charts, Mermaid graphs, generated visuals, SHA-pinned permalinks, verification results) → comparative analysis when options compete → numbered sources with access dates → methodology appendix (workers, waves, searches, verifications, debate rounds) → correction log naming what verification overturned. Write it long and specific: every claim cites `[Source N]`, and the sources section lists every source the run actually used rather than a curated few.
338
+ Assembly worker — `task(category="deep-low", load_skills=["frontend", "visual-qa", "open-design", "data-scientist", "imagegen", "ulw-loop"], run_in_background=true, ...)`: before writing, read every available design and visualization skill and apply it — the report is a designed artifact, not a text dump; the worker's prompt carries `design-spec.md`. Use the template the user approved; absent a stronger house style the default skeleton is executive summary → key findings by theme → detailed analysis (quotes under 20 words with attribution, charts, Mermaid graphs, generated visuals, SHA-pinned permalinks, verification results) → comparative analysis when options compete → numbered sources with access dates → methodology appendix (workers, waves, searches, verifications, debate rounds) → correction log naming what verification overturned. Write it long and specific: every claim cites `[Source N]`, and the sources section lists every source the run actually used rather than a curated few.
339
339
 
340
340
  ### The delivery gate — visual QA must PASS
341
341