oh-my-opencode 5.0.0-beta.79 → 5.0.0-beta.80
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/command/remove-deadcode.md +1 -1
- package/.agents/skills/hyperplan/SKILL.md +5 -5
- package/.agents/skills/remove-deadcode/SKILL.md +1 -1
- package/.agents/skills/security-research/SKILL.md +4 -4
- package/dist/cli/index.js +499 -336
- package/dist/cli-node/index.js +499 -336
- package/dist/config-migration/category-deep-split.d.ts +3 -0
- package/dist/config-migration/migration-plans.d.ts +1 -0
- package/dist/features/builtin-commands/templates/hyperplan.d.ts +1 -1
- package/dist/features/builtin-commands/templates/refactor-sections/team-mode-addendum.d.ts +1 -1
- package/dist/features/builtin-commands/templates/remove-ai-slops.d.ts +1 -1
- package/dist/index.js +282 -123
- package/dist/plugin-config/omo-config-chain.d.ts +1 -0
- package/dist/skills/debugging/references/methodology/02-investigate.md +5 -5
- package/dist/skills/refactor/SKILL.md +2 -2
- package/dist/skills/remove-ai-slops/SKILL.md +4 -4
- package/dist/skills/ulw-execute/SKILL.md +3 -2
- package/dist/skills/ulw-plan/references/full-workflow.md +1 -1
- package/dist/skills/ulw-research/SKILL.md +2 -2
- package/dist/tools/delegate-task/openai-categories.d.ts +6 -4
- package/dist/tui.js +146 -71
- package/package.json +14 -14
- package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/dist/cli.js +77 -8
- package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
- package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
- package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/rules/package.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/dist/cli.js +63 -17
- package/packages/omo-codex/plugin/components/telemetry/dist/posthog.js +63 -17
- package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/full-workflow.md +1 -1
- package/packages/omo-codex/plugin/components/ulw-execute-continuation/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-execute-continuation/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/checkpoint-template.js +3 -3
- package/packages/omo-codex/plugin/components/ulw-loop/dist/cli.js +14 -6
- package/packages/omo-codex/plugin/components/ulw-loop/dist/codex-goal-instruction.js +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/dist/surface.js +9 -1
- package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +5 -5
- package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/checkpoint-template.ts +3 -3
- package/packages/omo-codex/plugin/components/ulw-loop/src/codex-goal-instruction.ts +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/src/surface.ts +9 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/checkpoint-template.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/codex-goal-instruction.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/sdk-contract.test.ts +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-recording-spawn-admission.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-execute-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/package-lock.json +12 -12
- package/packages/omo-codex/plugin/package.json +1 -1
- package/packages/omo-codex/plugin/skills/debugging/references/methodology/02-investigate.md +5 -5
- package/packages/omo-codex/plugin/skills/refactor/SKILL.md +2 -2
- package/packages/omo-codex/plugin/skills/remove-ai-slops/SKILL.md +4 -4
- package/packages/omo-codex/plugin/skills/ulw-execute/SKILL.md +3 -2
- package/packages/omo-codex/plugin/skills/ulw-plan/references/full-workflow.md +1 -1
- package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +2 -2
- package/packages/omo-codex/scripts/install-dist/install-local.mjs +167 -51
- package/packages/shared-skills/skills/debugging/references/methodology/02-investigate.md +5 -5
- package/packages/shared-skills/skills/refactor/SKILL.md +2 -2
- package/packages/shared-skills/skills/remove-ai-slops/SKILL.md +4 -4
- package/packages/shared-skills/skills/ulw-execute/SKILL.md +3 -2
- package/packages/shared-skills/skills/ulw-plan/references/full-workflow.md +1 -1
- package/packages/shared-skills/skills/ulw-research/SKILL.md +2 -2
|
@@ -1,12 +1,12 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@sisyphuslabs/omo-codex-plugin",
|
|
3
|
-
"version": "5.0.0-beta.
|
|
3
|
+
"version": "5.0.0-beta.80",
|
|
4
4
|
"lockfileVersion": 3,
|
|
5
5
|
"requires": true,
|
|
6
6
|
"packages": {
|
|
7
7
|
"": {
|
|
8
8
|
"name": "@sisyphuslabs/omo-codex-plugin",
|
|
9
|
-
"version": "5.0.0-beta.
|
|
9
|
+
"version": "5.0.0-beta.80",
|
|
10
10
|
"workspaces": [
|
|
11
11
|
"components/comment-checker",
|
|
12
12
|
"components/git-bash",
|
|
@@ -90,7 +90,7 @@
|
|
|
90
90
|
},
|
|
91
91
|
"components/comment-checker": {
|
|
92
92
|
"name": "@code-yeongyu/codex-comment-checker",
|
|
93
|
-
"version": "5.0.0-beta.
|
|
93
|
+
"version": "5.0.0-beta.80",
|
|
94
94
|
"license": "MIT",
|
|
95
95
|
"bin": {
|
|
96
96
|
"omo-comment-checker": "dist/cli.js"
|
|
@@ -111,7 +111,7 @@
|
|
|
111
111
|
},
|
|
112
112
|
"components/git-bash": {
|
|
113
113
|
"name": "@sisyphuslabs/codex-git-bash-hook",
|
|
114
|
-
"version": "5.0.0-beta.
|
|
114
|
+
"version": "5.0.0-beta.80",
|
|
115
115
|
"bin": {
|
|
116
116
|
"omo-git-bash-hook": "dist/cli.js"
|
|
117
117
|
},
|
|
@@ -125,7 +125,7 @@
|
|
|
125
125
|
},
|
|
126
126
|
"components/lazycodex-executor-verify": {
|
|
127
127
|
"name": "@code-yeongyu/codex-lazycodex-executor-verify",
|
|
128
|
-
"version": "5.0.0-beta.
|
|
128
|
+
"version": "5.0.0-beta.80",
|
|
129
129
|
"license": "MIT",
|
|
130
130
|
"bin": {
|
|
131
131
|
"lazycodex-executor-verify": "dist/cli.js"
|
|
@@ -142,7 +142,7 @@
|
|
|
142
142
|
},
|
|
143
143
|
"components/lsp": {
|
|
144
144
|
"name": "@code-yeongyu/codex-lsp",
|
|
145
|
-
"version": "5.0.0-beta.
|
|
145
|
+
"version": "5.0.0-beta.80",
|
|
146
146
|
"license": "MIT",
|
|
147
147
|
"dependencies": {
|
|
148
148
|
"@code-yeongyu/lsp-daemon": "file:../../../../lsp-daemon",
|
|
@@ -163,7 +163,7 @@
|
|
|
163
163
|
},
|
|
164
164
|
"components/rules": {
|
|
165
165
|
"name": "@code-yeongyu/codex-rules",
|
|
166
|
-
"version": "5.0.0-beta.
|
|
166
|
+
"version": "5.0.0-beta.80",
|
|
167
167
|
"license": "MIT",
|
|
168
168
|
"dependencies": {
|
|
169
169
|
"picomatch": "^4.0.7"
|
|
@@ -185,7 +185,7 @@
|
|
|
185
185
|
},
|
|
186
186
|
"components/teammode": {
|
|
187
187
|
"name": "@sisyphuslabs/codex-teammode",
|
|
188
|
-
"version": "5.0.0-beta.
|
|
188
|
+
"version": "5.0.0-beta.80",
|
|
189
189
|
"devDependencies": {
|
|
190
190
|
"@types/node": "^26.2.0",
|
|
191
191
|
"bun-types": "^1.4.2",
|
|
@@ -198,7 +198,7 @@
|
|
|
198
198
|
},
|
|
199
199
|
"components/telemetry": {
|
|
200
200
|
"name": "@code-yeongyu/codex-telemetry",
|
|
201
|
-
"version": "5.0.0-beta.
|
|
201
|
+
"version": "5.0.0-beta.80",
|
|
202
202
|
"license": "MIT",
|
|
203
203
|
"bin": {
|
|
204
204
|
"omo-telemetry": "dist/cli.js"
|
|
@@ -216,7 +216,7 @@
|
|
|
216
216
|
},
|
|
217
217
|
"components/ultrawork": {
|
|
218
218
|
"name": "@code-yeongyu/codex-ultrawork",
|
|
219
|
-
"version": "5.0.0-beta.
|
|
219
|
+
"version": "5.0.0-beta.80",
|
|
220
220
|
"license": "MIT",
|
|
221
221
|
"bin": {
|
|
222
222
|
"omo-ultrawork": "dist/cli.js"
|
|
@@ -234,7 +234,7 @@
|
|
|
234
234
|
},
|
|
235
235
|
"components/ulw-execute-continuation": {
|
|
236
236
|
"name": "@code-yeongyu/codex-ulw-execute-continuation",
|
|
237
|
-
"version": "5.0.0-beta.
|
|
237
|
+
"version": "5.0.0-beta.80",
|
|
238
238
|
"license": "MIT",
|
|
239
239
|
"bin": {
|
|
240
240
|
"omo-ulw-execute-continuation": "dist/cli.js"
|
|
@@ -251,7 +251,7 @@
|
|
|
251
251
|
},
|
|
252
252
|
"components/ulw-loop": {
|
|
253
253
|
"name": "@code-yeongyu/codex-ulw-loop",
|
|
254
|
-
"version": "5.0.0-beta.
|
|
254
|
+
"version": "5.0.0-beta.80",
|
|
255
255
|
"license": "MIT",
|
|
256
256
|
"bin": {
|
|
257
257
|
"omo-ulw-loop": "dist/cli.js",
|
|
@@ -75,22 +75,22 @@ When the `team_*` tools are present, create a **debug-squad** team and split inv
|
|
|
75
75
|
"members": [
|
|
76
76
|
{
|
|
77
77
|
"kind": "category",
|
|
78
|
-
"category": "deep",
|
|
78
|
+
"category": "deep-low",
|
|
79
79
|
"prompt": "You are the Runtime State Inspector. Your job: attach to the live process, hit breakpoints, read program state (variables, heap, goroutines, stack, registers depending on runtime), and report observed values verbatim. Never guess — if you don't see the value, say so. Report back via team_send_message with file:line / address references and captured values. Never edit source code. Never run git commands. If you need an instrumentation statement added (breakpoint(), debugger;, dbg!, etc.), ask the Lead first."
|
|
80
80
|
},
|
|
81
81
|
{
|
|
82
82
|
"kind": "category",
|
|
83
|
-
"category": "deep",
|
|
83
|
+
"category": "deep-low",
|
|
84
84
|
"prompt": "You are the Log Archaeologist. Your job: grep server logs, stderr streams, SDK-internal debug output (DEBUG env, RUST_LOG, GODEBUG, PYTHONASYNCIODEBUG), and correlate timestamps. Produce a timeline of events with latencies. Flag anything that looks like a silent catch, a swallowed rejection, a panic recovered-and-ignored, a success response that contains failure signals (HTTP 200 with empty body, stopReason=error, exit 0 with error-in-stdout). Never edit source code."
|
|
85
85
|
},
|
|
86
86
|
{
|
|
87
87
|
"kind": "category",
|
|
88
|
-
"category": "deep",
|
|
88
|
+
"category": "deep-low",
|
|
89
89
|
"prompt": "You are the Reproduction Engineer. Your job: build the smallest reliable repro — a curl command, a vitest/pytest/go test, a tmux script, a Playwright script for browser bugs, a pwntools script for binary targets. It must reproduce on first try and be copy-pasteable by the Lead. Document exact input, expected output, observed output. Save repro artifacts under /tmp/ and tell the Lead to journal them. If the bug is browser-based you MUST use Playwright CLI — do not simulate with curl."
|
|
90
90
|
},
|
|
91
91
|
{
|
|
92
92
|
"kind": "category",
|
|
93
|
-
"category": "deep",
|
|
93
|
+
"category": "deep-low",
|
|
94
94
|
"prompt": "You are the Trace Correlator. Your job: take findings from the other members and cross-link them. Build a causal chain from symptom to suspected cause. Identify missing evidence. Propose the next single most-decisive runtime query. Never edit source code; only reason across already-captured evidence. If hypotheses diverge sharply after correlation, tell the Lead immediately — that is the signal for the Oracle Triple."
|
|
95
95
|
}
|
|
96
96
|
]
|
|
@@ -117,7 +117,7 @@ task(subagent_type="explore", load_skills=[], run_in_background=true,
|
|
|
117
117
|
Runtime state investigation for hypothesis 1: ...")
|
|
118
118
|
task(subagent_type="explore", load_skills=[], run_in_background=true,
|
|
119
119
|
prompt="Log/timing investigation for hypothesis 2: ...")
|
|
120
|
-
task(category="deep", load_skills=[], run_in_background=true,
|
|
120
|
+
task(category="deep-low", load_skills=[], run_in_background=true,
|
|
121
121
|
prompt="Reproduction minimizer for hypothesis 3: ...")
|
|
122
122
|
```
|
|
123
123
|
|
|
@@ -708,7 +708,7 @@ Record the chosen path in the TodoWrite list.
|
|
|
708
708
|
|
|
709
709
|
Rationale for this composition:
|
|
710
710
|
- **4 workers = team mode's parallel cap.** 5+ just queues.
|
|
711
|
-
- **No verifier team member.** Verification needs \`deep\` reasoning (or \`unspecified-high\` fallback). In-team category routing downcasts to the category worker, which is weaker than required — the verifier runs OUTSIDE the team as a \`task(category="deep")\`.
|
|
711
|
+
- **No verifier team member.** Verification needs \`deep-high\` reasoning (or \`unspecified-high\` fallback). In-team category routing downcasts to the category worker, which is weaker than required — the verifier runs OUTSIDE the team as a \`task(category="deep-high")\`.
|
|
712
712
|
- **quick × 2** for mechanical edits, **unspecified-low × 2** for reasoning edits — mirrors the plan's split.
|
|
713
713
|
|
|
714
714
|
**Team lifecycle** (one team, reused until Phase 6 cleanup):
|
|
@@ -740,7 +740,7 @@ While any team task is \`pending | claimed | in_progress\`:
|
|
|
740
740
|
- On a worker completion report, immediately dispatch an **external verifier** — verification runs OUTSIDE the team because team-member category routing downcasts to the category worker:
|
|
741
741
|
\`\`\`
|
|
742
742
|
task(
|
|
743
|
-
category="deep",
|
|
743
|
+
category="deep-high",
|
|
744
744
|
load_skills=[],
|
|
745
745
|
run_in_background=true,
|
|
746
746
|
description="verify step <N>",
|
|
@@ -184,9 +184,9 @@ File: src/bar.py
|
|
|
184
184
|
|
|
185
185
|
Order rule (safest → riskiest): comments → dead code → defensive → duplication → complexity → abstraction/boundary → performance → tests → oversized-modules. This minimizes blast radius of any one change.
|
|
186
186
|
|
|
187
|
-
### Phase 4: Parallel slop removal via `deep` agents in batches of 5
|
|
187
|
+
### Phase 4: Parallel slop removal via `deep-low` agents in batches of 5
|
|
188
188
|
|
|
189
|
-
Files are processed by `deep` category agents with the `$omo:remove-ai-slops` skill loaded, **batched 5 at a time in parallel**. The executable skill name is `remove-ai-slops`. The `deep` category gives the agent enough thoroughness to correctly evaluate the 9 categories and respect the KEEP rules without slipping into surface fixes; the 5-wide batch is the sweet spot — more than 5 creates result-merging noise and context contention, fewer wastes parallelism.
|
|
189
|
+
Files are processed by `deep-low` category agents with the `$omo:remove-ai-slops` skill loaded, **batched 5 at a time in parallel**. The executable skill name is `remove-ai-slops`. The `deep-low` category gives the agent enough thoroughness to correctly evaluate the 9 categories and respect the KEEP rules without slipping into surface fixes; the 5-wide batch is the sweet spot — more than 5 creates result-merging noise and context contention, fewer wastes parallelism.
|
|
190
190
|
|
|
191
191
|
**Batching protocol** (strict):
|
|
192
192
|
|
|
@@ -203,7 +203,7 @@ Files are processed by `deep` category agents with the `$omo:remove-ai-slops` sk
|
|
|
203
203
|
|
|
204
204
|
```
|
|
205
205
|
task(
|
|
206
|
-
category="deep",
|
|
206
|
+
category="deep-low",
|
|
207
207
|
load_skills=["remove-ai-slops"],
|
|
208
208
|
run_in_background=true,
|
|
209
209
|
description="Slop removal: {filename}",
|
|
@@ -229,7 +229,7 @@ For each skipped issue, give reason.
|
|
|
229
229
|
)
|
|
230
230
|
```
|
|
231
231
|
|
|
232
|
-
**Batch failure handling**: a `multi_agent_v1.wait_agent` timeout only means no new mailbox update arrived, not that a `deep` agent failed. For long passes, require each child to send `WORKING: <file> - <current phase>` and `BLOCKED: <reason>` only when it cannot progress. Treat a running child as alive. Mark a file for retry only when the child is completed without the deliverable, ack-only after followup, explicitly `BLOCKED:`, or no longer running. Do NOT block the remaining 4 in that batch; collect successful results and retry the failed file once later. If retry also fails, escalate that file under "Issues Found & Fixed" in the final report.
|
|
232
|
+
**Batch failure handling**: a `multi_agent_v1.wait_agent` timeout only means no new mailbox update arrived, not that a `deep-low` agent failed. For long passes, require each child to send `WORKING: <file> - <current phase>` and `BLOCKED: <reason>` only when it cannot progress. Treat a running child as alive. Mark a file for retry only when the child is completed without the deliverable, ack-only after followup, explicitly `BLOCKED:`, or no longer running. Do NOT block the remaining 4 in that batch; collect successful results and retry the failed file once later. If retry also fails, escalate that file under "Issues Found & Fixed" in the final report.
|
|
233
233
|
|
|
234
234
|
### Phase 5: Verify with quality gates + critical review
|
|
235
235
|
|
|
@@ -152,13 +152,14 @@ When the plan annotates a todo with `Recommended task executor category:`, follo
|
|
|
152
152
|
| `visual-engineering` (medium) | frontend, UI/UX, styling, animation |
|
|
153
153
|
| `writing` (low) | documentation and prose |
|
|
154
154
|
| `git` (low) | git operations |
|
|
155
|
-
| `deep` (
|
|
155
|
+
| `deep-low` (medium) | hairy debugging, research-heavy or subtle cross-module work the worker can settle from what it reads |
|
|
156
|
+
| `deep-high` (high) | the same, when the central decision cannot be settled from evidence: a trade-off, a cross-package contract, or correctness argued from invariants |
|
|
156
157
|
| `ultrabrain` (high) | ONE genuinely hard, logic-heavy problem — hand it the goal, not step-by-step instructions |
|
|
157
158
|
|
|
158
159
|
Sizing is a two-branch decision made per checkbox, before dispatch:
|
|
159
160
|
|
|
160
161
|
- **Splittable work splits.** When the checkbox decomposes into independent pieces, dispatch them as a swarm of `quick`/`unspecified-low` workers in ONE parallel burst — many small cheap workers in parallel beat one large delegation.
|
|
161
|
-
- **Cohesive hard work stays whole.** When splitting would sever shared reasoning (one algorithm, one migration, one subtle bug), send the WHOLE problem to `deep` or `ultrabrain` as ONE delegation. Never force-split work whose parts share one insight.
|
|
162
|
+
- **Cohesive hard work stays whole.** When splitting would sever shared reasoning (one algorithm, one migration, one subtle bug), send the WHOLE problem to `deep-low`, `deep-high` or `ultrabrain` as ONE delegation. Never force-split work whose parts share one insight.
|
|
162
163
|
|
|
163
164
|
Each sub-task message must include:
|
|
164
165
|
|
|
@@ -154,7 +154,7 @@ No Metis, no plan file, no execution until the user approves. The UNCLEAR path a
|
|
|
154
154
|
## Commit strategy
|
|
155
155
|
## Success criteria
|
|
156
156
|
```
|
|
157
|
-
> Target 5-8 todos per wave; fewer than 3 (except the final) means under-splitting. Implementation + Test = ONE todo. Each todo carries: exhaustive References (the executor has no interview context), agent-executable Acceptance criteria, happy + failure QA scenarios each with an evidence path, a Commit line, and a `Recommended task executor category:` line - the routing verdict the executor follows, with a one-line reason, in the omo category vocabulary: `quick` (mechanical / single-file - the default for every splittable piece), `unspecified-low` (small misc), `unspecified-high` (standard multi-file feature), `visual-engineering` (frontend/UI), `writing` (docs), `git` (git ops), `deep` (hairy debugging or cross-module reasoning), `ultrabrain` (ONE genuinely hard cohesive problem, delegated whole). Prefer many small `quick`-routable todos spread across parallel waves; when splitting would sever shared reasoning, keep ONE todo routed to `deep`/`ultrabrain` - never force-split work whose parts share one insight. Harnesses without categories map by difficulty: quick/unspecified-low/writing/git = low, unspecified-high/visual-engineering = medium, deep/ultrabrain = high.
|
|
157
|
+
> Target 5-8 todos per wave; fewer than 3 (except the final) means under-splitting. Implementation + Test = ONE todo. Each todo carries: exhaustive References (the executor has no interview context), agent-executable Acceptance criteria, happy + failure QA scenarios each with an evidence path, a Commit line, and a `Recommended task executor category:` line - the routing verdict the executor follows, with a one-line reason, in the omo category vocabulary: `quick` (mechanical / single-file - the default for every splittable piece), `unspecified-low` (small misc), `unspecified-high` (standard multi-file feature), `visual-engineering` (frontend/UI), `writing` (docs), `git` (git ops), `deep-low` (hairy debugging or cross-module reasoning the worker can settle from what it reads), `deep-high` (the same, when the central decision cannot be settled from evidence: a trade-off, a cross-package contract, or correctness argued from invariants), `ultrabrain` (ONE genuinely hard cohesive problem, delegated whole). Prefer many small `quick`-routable todos spread across parallel waves; when splitting would sever shared reasoning, keep ONE todo routed to `deep`/`ultrabrain` - never force-split work whose parts share one insight. Harnesses without categories map by difficulty: quick/unspecified-low/writing/git = low, unspecified-high/visual-engineering = medium, deep/ultrabrain = high.
|
|
158
158
|
|
|
159
159
|
## Plan artifact producer contract
|
|
160
160
|
|
|
@@ -255,7 +255,7 @@ Interest alone is not a trigger. Anything without one stays a queued lead in `ex
|
|
|
255
255
|
Settle with executed code, not judgment, whenever sources disagree, a behavior is undocumented, a claim is performance- or compatibility-shaped, or the honest answer is "it should work". Spawn one verification worker per claim:
|
|
256
256
|
|
|
257
257
|
```
|
|
258
|
-
task(category="deep", run_in_background=true, prompt="TASK: verify by execution: <claim>.
|
|
258
|
+
task(category="deep-low", run_in_background=true, prompt="TASK: verify by execution: <claim>.
|
|
259
259
|
SOURCE: <where it came from>; CONTRADICTION: <opposing source, if any>.
|
|
260
260
|
Write a minimal self-contained script that tests the claim; run it (uv run --with <deps> python / bun / direct compile); capture full stdout+stderr; pin versions.
|
|
261
261
|
Reply with: the exact code, the full output, environment (OS, runtime, dependency versions), and a verdict — CONFIRMED / REFUTED / PARTIAL — grounded in the output.")
|
|
@@ -335,7 +335,7 @@ Asset workers (background, parallel, each fed `design-spec.md`) — visuals are
|
|
|
335
335
|
|
|
336
336
|
**Verify the asset manifest before rendering.** List every asset the document references, assert each file exists and is non-empty on disk, and re-render whatever is missing. A document that renders with three broken diagrams is a document you will publish twice.
|
|
337
337
|
|
|
338
|
-
Assembly worker — `task(category="deep", load_skills=["frontend", "visual-qa", "open-design", "data-scientist", "imagegen", "ulw-loop"], run_in_background=true, ...)`: before writing, read every available design and visualization skill and apply it — the report is a designed artifact, not a text dump; the worker's prompt carries `design-spec.md`. Use the template the user approved; absent a stronger house style the default skeleton is executive summary → key findings by theme → detailed analysis (quotes under 20 words with attribution, charts, Mermaid graphs, generated visuals, SHA-pinned permalinks, verification results) → comparative analysis when options compete → numbered sources with access dates → methodology appendix (workers, waves, searches, verifications, debate rounds) → correction log naming what verification overturned. Write it long and specific: every claim cites `[Source N]`, and the sources section lists every source the run actually used rather than a curated few.
|
|
338
|
+
Assembly worker — `task(category="deep-low", load_skills=["frontend", "visual-qa", "open-design", "data-scientist", "imagegen", "ulw-loop"], run_in_background=true, ...)`: before writing, read every available design and visualization skill and apply it — the report is a designed artifact, not a text dump; the worker's prompt carries `design-spec.md`. Use the template the user approved; absent a stronger house style the default skeleton is executive summary → key findings by theme → detailed analysis (quotes under 20 words with attribution, charts, Mermaid graphs, generated visuals, SHA-pinned permalinks, verification results) → comparative analysis when options compete → numbered sources with access dates → methodology appendix (workers, waves, searches, verifications, debate rounds) → correction log naming what verification overturned. Write it long and specific: every claim cites `[Source N]`, and the sources section lists every source the run actually used rather than a curated few.
|
|
339
339
|
|
|
340
340
|
### The delivery gate — visual QA must PASS
|
|
341
341
|
|