lazycodex-ai 4.16.0 → 4.16.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (137) hide show
  1. package/dist/cli/codex-ulw-loop.d.ts +8 -0
  2. package/dist/cli/get-local-version/types.d.ts +1 -1
  3. package/dist/cli/index.js +260 -207
  4. package/dist/cli-node/index.js +260 -207
  5. package/docs/reference/web-terminal-visual-qa.md +39 -45
  6. package/package.json +3 -2
  7. package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
  8. package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
  9. package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
  10. package/packages/omo-codex/plugin/components/codegraph/package.json +1 -1
  11. package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
  12. package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
  13. package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
  14. package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
  15. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
  16. package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
  17. package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
  18. package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
  19. package/packages/omo-codex/plugin/components/rules/bundled-rules/hephaestus.md +1 -1
  20. package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
  21. package/packages/omo-codex/plugin/components/rules/package.json +1 -1
  22. package/packages/omo-codex/plugin/components/start-work-continuation/directive.md +3 -3
  23. package/packages/omo-codex/plugin/components/start-work-continuation/hooks/hooks.json +2 -2
  24. package/packages/omo-codex/plugin/components/start-work-continuation/package.json +1 -1
  25. package/packages/omo-codex/plugin/components/start-work-continuation/test/codex-hook.test.ts +1 -1
  26. package/packages/omo-codex/plugin/components/teammode/AGENTS.md +1 -1
  27. package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
  28. package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
  29. package/packages/omo-codex/plugin/components/teammode/skills/teammode/SKILL.md +3 -3
  30. package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
  31. package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
  32. package/packages/omo-codex/plugin/components/ultrawork/directive.md +19 -10
  33. package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
  34. package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
  35. package/packages/omo-codex/plugin/components/ultrawork/skills/ultrawork/SKILL.md +19 -10
  36. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/SKILL.md +13 -0
  37. package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/full-workflow.md +2 -0
  38. package/packages/omo-codex/plugin/components/ultrawork/test/codex-hook.test.ts +22 -1
  39. package/packages/omo-codex/plugin/components/ulw-loop/CHANGELOG.md +4 -0
  40. package/packages/omo-codex/plugin/components/ulw-loop/directive.md +19 -10
  41. package/packages/omo-codex/plugin/components/ulw-loop/dist/cli-commands.js +15 -2
  42. package/packages/omo-codex/plugin/components/ulw-loop/dist/cli-steering.js +2 -1
  43. package/packages/omo-codex/plugin/components/ulw-loop/dist/cli.js +89 -27
  44. package/packages/omo-codex/plugin/components/ulw-loop/dist/plan-io.d.ts +6 -0
  45. package/packages/omo-codex/plugin/components/ulw-loop/dist/plan-io.js +55 -9
  46. package/packages/omo-codex/plugin/components/ulw-loop/dist/steering-snapshot.d.ts +15 -0
  47. package/packages/omo-codex/plugin/components/ulw-loop/dist/steering-snapshot.js +33 -0
  48. package/packages/omo-codex/plugin/components/ulw-loop/dist/steering-types.d.ts +10 -3
  49. package/packages/omo-codex/plugin/components/ulw-loop/dist/steering.js +15 -11
  50. package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +2 -2
  51. package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
  52. package/packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/SKILL.md +3 -1
  53. package/packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/references/full-workflow.md +7 -7
  54. package/packages/omo-codex/plugin/components/ulw-loop/src/cli-commands.ts +17 -2
  55. package/packages/omo-codex/plugin/components/ulw-loop/src/cli-steering.ts +2 -1
  56. package/packages/omo-codex/plugin/components/ulw-loop/src/plan-io.ts +59 -11
  57. package/packages/omo-codex/plugin/components/ulw-loop/src/steering-snapshot.ts +38 -0
  58. package/packages/omo-codex/plugin/components/ulw-loop/src/steering-types.ts +11 -3
  59. package/packages/omo-codex/plugin/components/ulw-loop/src/steering.ts +15 -7
  60. package/packages/omo-codex/plugin/components/ulw-loop/test/cli-create-goals.test.ts +16 -0
  61. package/packages/omo-codex/plugin/components/ulw-loop/test/plan-io.test.ts +260 -2
  62. package/packages/omo-codex/plugin/components/ulw-loop/test/skill-contract.test.ts +1 -1
  63. package/packages/omo-codex/plugin/components/ulw-loop/test/steering-snapshot.test.ts +124 -0
  64. package/packages/omo-codex/plugin/components/ulw-loop/test/steering.test.ts +101 -2
  65. package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
  66. package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
  67. package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
  68. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-codegraph-init-guidance.json +1 -1
  69. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
  70. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
  71. package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
  72. package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
  73. package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
  74. package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
  75. package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
  76. package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
  77. package/packages/omo-codex/plugin/hooks/session-start-checking-codegraph-bootstrap.json +1 -1
  78. package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
  79. package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
  80. package/packages/omo-codex/plugin/hooks/stop-checking-start-work-continuation.json +1 -1
  81. package/packages/omo-codex/plugin/hooks/subagent-stop-checking-start-work-continuation.json +1 -1
  82. package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
  83. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
  84. package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
  85. package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
  86. package/packages/omo-codex/plugin/package-lock.json +13 -13
  87. package/packages/omo-codex/plugin/package.json +1 -1
  88. package/packages/omo-codex/plugin/scripts/auto-update.mjs +64 -17
  89. package/packages/omo-codex/plugin/scripts/hook-status-message.mjs +10 -6
  90. package/packages/omo-codex/plugin/scripts/migrate-codex-config/multi-agent-v2-guard.mjs +186 -20
  91. package/packages/omo-codex/plugin/scripts/migrate-codex-config/subagent-limit-guard.mjs +51 -3
  92. package/packages/omo-codex/plugin/scripts/migrate-codex-config.mjs +33 -5
  93. package/packages/omo-codex/plugin/scripts/sync-skills.mjs +1 -1
  94. package/packages/omo-codex/plugin/skills/frontend/SKILL.md +9 -9
  95. package/packages/omo-codex/plugin/skills/frontend/references/design/README.md +7 -3
  96. package/packages/omo-codex/plugin/skills/frontend/references/design/_INDEX.md +1 -1
  97. package/packages/omo-codex/plugin/skills/frontend/references/design/design-system-architecture.md +24 -2
  98. package/packages/omo-codex/plugin/skills/frontend/references/designpowers/README.md +2 -2
  99. package/packages/omo-codex/plugin/skills/frontend/references/designpowers/lane-b-execution.md +1 -1
  100. package/packages/omo-codex/plugin/skills/init-deep/SKILL.md +1 -1
  101. package/packages/omo-codex/plugin/skills/refactor/SKILL.md +1 -1
  102. package/packages/omo-codex/plugin/skills/remove-ai-slops/SKILL.md +1 -1
  103. package/packages/omo-codex/plugin/skills/review-work/SKILL.md +1 -1
  104. package/packages/omo-codex/plugin/skills/start-work/SKILL.md +3 -3
  105. package/packages/omo-codex/plugin/skills/teammode/SKILL.md +3 -3
  106. package/packages/omo-codex/plugin/skills/ultrawork/SKILL.md +19 -10
  107. package/packages/omo-codex/plugin/skills/ulw-loop/SKILL.md +3 -1
  108. package/packages/omo-codex/plugin/skills/ulw-loop/references/full-workflow.md +7 -7
  109. package/packages/omo-codex/plugin/skills/ulw-plan/SKILL.md +13 -0
  110. package/packages/omo-codex/plugin/skills/ulw-plan/references/full-workflow.md +2 -0
  111. package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +1 -1
  112. package/packages/omo-codex/plugin/skills/visual-qa/SKILL.md +14 -18
  113. package/packages/omo-codex/plugin/test/aggregate-hooks.test.mjs +4 -4
  114. package/packages/omo-codex/plugin/test/aggregate-plugin-fixture.mjs +1 -1
  115. package/packages/omo-codex/plugin/test/auto-update.test.mjs +35 -1
  116. package/packages/omo-codex/plugin/test/bootstrap-hooks.test.mjs +1 -1
  117. package/packages/omo-codex/plugin/test/hook-status-message.test.mjs +22 -9
  118. package/packages/omo-codex/plugin/test/migrate-codex-config.test.mjs +275 -19
  119. package/packages/omo-codex/plugin/test/subagent-limit-migration.test.mjs +33 -0
  120. package/packages/omo-codex/plugin/test/sync-hook-status-messages.test.mjs +6 -6
  121. package/packages/omo-codex/plugin/test/sync-skills-orchestration.test.mjs +4 -2
  122. package/packages/omo-codex/plugin/test/ulw-plan-skill-contract.test.mjs +52 -0
  123. package/packages/omo-codex/scripts/install-dist/install-local.mjs +65 -39
  124. package/packages/shared-skills/skills/frontend/SKILL.md +9 -9
  125. package/packages/shared-skills/skills/frontend/references/design/README.md +7 -3
  126. package/packages/shared-skills/skills/frontend/references/design/_INDEX.md +1 -1
  127. package/packages/shared-skills/skills/frontend/references/design/design-system-architecture.md +24 -2
  128. package/packages/shared-skills/skills/frontend/references/designpowers/README.md +2 -2
  129. package/packages/shared-skills/skills/frontend/references/designpowers/lane-b-execution.md +1 -1
  130. package/packages/shared-skills/skills/review-work/SKILL.md +3 -1
  131. package/packages/shared-skills/skills/start-work/SKILL.md +4 -2
  132. package/packages/shared-skills/skills/ulw-research/SKILL.md +3 -1
  133. package/packages/shared-skills/skills/visual-qa/SKILL.md +13 -17
  134. package/script/qa/strip-ansi.mjs +10 -0
  135. package/script/qa/web-terminal-visual-qa.mjs +112 -195
  136. package/script/qa/xterm-live-terminal.mjs +180 -0
  137. package/script/qa/web-terminal-renderer.mjs +0 -218
@@ -1,9 +1,39 @@
1
+ import { prefersMultiAgentV2, resolveMultiAgentVersionFromConfig } from "./multi-agent-v2-guard.mjs";
2
+
1
3
  const CODEX_AGENTS_HEADER = "[agents]";
2
4
  const CODEX_MULTI_AGENT_V2_HEADER = "[features.multi_agent_v2]";
3
5
  const CODEX_SUBAGENT_THREAD_LIMIT = "1000";
4
6
 
5
- export function ensureSubagentConcurrencyLimit(config) {
6
- return ensureMultiAgentV2ThreadLimit(ensureAgentsMaxThreads(config));
7
+ /**
8
+ * Ensure subagent concurrency limits without writing settings that conflict
9
+ * with MultiAgentV2. When the selected model prefers V2 (catalog `v2`, or a
10
+ * GPT-5.6 family session model with the catalog unavailable) or V2 is already
11
+ * enabled in config, skip `agents.max_threads` because Codex rejects that key
12
+ * while features.multi_agent_v2 is enabled.
13
+ *
14
+ * @param {string} config
15
+ * @param {{ multiAgentVersion?: string | null, sessionModel?: string | null, env?: NodeJS.ProcessEnv, modelsCachePath?: string }} [options]
16
+ */
17
+ export function ensureSubagentConcurrencyLimit(config, options = {}) {
18
+ const multiAgentVersion =
19
+ options.multiAgentVersion !== undefined
20
+ ? options.multiAgentVersion
21
+ : resolveMultiAgentVersionFromConfig(config, options);
22
+ const v2Preferred = prefersMultiAgentV2(multiAgentVersion, options.sessionModel) || isMultiAgentV2Enabled(config);
23
+
24
+ let result = config;
25
+ if (!v2Preferred) {
26
+ result = ensureAgentsMaxThreads(result);
27
+ } else {
28
+ result = removeAgentsMaxThreads(result);
29
+ }
30
+ return ensureMultiAgentV2ThreadLimit(result);
31
+ }
32
+
33
+ function isMultiAgentV2Enabled(config) {
34
+ const section = findSection(config, CODEX_MULTI_AGENT_V2_HEADER);
35
+ if (!section) return false;
36
+ return /^\s*enabled\s*=\s*true[ \t]*(?:#[^\n]*)?$/m.test(section.text);
7
37
  }
8
38
 
9
39
  function ensureAgentsMaxThreads(config) {
@@ -12,6 +42,22 @@ function ensureAgentsMaxThreads(config) {
12
42
  return replaceOrInsertSetting(config, section, "max_threads", CODEX_SUBAGENT_THREAD_LIMIT);
13
43
  }
14
44
 
45
+ function removeAgentsMaxThreads(config) {
46
+ const section = findSection(config, CODEX_AGENTS_HEADER);
47
+ if (!section) return config;
48
+ if (!/^\s*max_threads\s*=/m.test(section.text)) return config;
49
+
50
+ const patched = section.text.replace(/^\s*max_threads\s*=\s*[^\n]*\n?/m, "");
51
+ const bodyLines = patched
52
+ .split("\n")
53
+ .slice(1)
54
+ .filter((line) => line.trim() !== "");
55
+ if (bodyLines.length === 0) {
56
+ return config.slice(0, section.start) + config.slice(section.end).replace(/^\n+/, "");
57
+ }
58
+ return config.slice(0, section.start) + patched + config.slice(section.end);
59
+ }
60
+
15
61
  function ensureMultiAgentV2ThreadLimit(config) {
16
62
  const section = findSection(config, CODEX_MULTI_AGENT_V2_HEADER);
17
63
  if (!section) {
@@ -26,7 +72,9 @@ function ensureMultiAgentV2ThreadLimit(config) {
26
72
  function replaceOrInsertSetting(config, section, key, value) {
27
73
  const pattern = new RegExp(`^(\\s*)${escapeRegExp(key)}\\s*=\\s*[^\\n#]*(#[^\\n]*)?$`, "m");
28
74
  if (pattern.test(section.text)) {
29
- const patched = section.text.replace(pattern, (_match, indent, comment) => comment ? `${indent}${key} = ${value} ${comment}` : `${indent}${key} = ${value}`);
75
+ const patched = section.text.replace(pattern, (_match, indent, comment) =>
76
+ comment ? `${indent}${key} = ${value} ${comment}` : `${indent}${key} = ${value}`,
77
+ );
30
78
  return config.slice(0, section.start) + patched + config.slice(section.end);
31
79
  }
32
80
 
@@ -8,7 +8,10 @@ import { FALLBACK_CATALOG, readModelCatalog } from "./migrate-codex-config/catal
8
8
  import { configPaths } from "./migrate-codex-config/config-paths.mjs";
9
9
  import { removeStaleContext7PlaceholderMcpServer } from "./migrate-codex-config/context7-placeholder-guard.mjs";
10
10
  import { removeUnsupportedRootMultiAgentMode } from "./migrate-codex-config/multi-agent-mode-guard.mjs";
11
- import { forceDisableMultiAgentV2 } from "./migrate-codex-config/multi-agent-v2-guard.mjs";
11
+ import {
12
+ forceDisableMultiAgentV2,
13
+ resolveMultiAgentVersionFromConfig,
14
+ } from "./migrate-codex-config/multi-agent-v2-guard.mjs";
12
15
  import { ensureCodexReasoningConfig as applyReasoningProfile, readRootSettings } from "./migrate-codex-config/root-settings.mjs";
13
16
  import { readState, resolveStatePath, writeState } from "./migrate-codex-config/state.mjs";
14
17
  import { ensureSubagentConcurrencyLimit } from "./migrate-codex-config/subagent-limit-guard.mjs";
@@ -19,7 +22,12 @@ export function ensureCodexReasoningConfig(config, profile = FALLBACK_CATALOG.cu
19
22
  return applyReasoningProfile(config, profile);
20
23
  }
21
24
 
22
- export async function migrateCodexConfig({ env = process.env, cwd = process.cwd() } = {}) {
25
+ export async function migrateCodexConfig({
26
+ env = process.env,
27
+ cwd = process.cwd(),
28
+ sessionModel = null,
29
+ requireSessionModel = false,
30
+ } = {}) {
23
31
  const catalog = await readModelCatalog(env);
24
32
  const statePath = resolveStatePath(env);
25
33
  const state = await readState(statePath);
@@ -31,6 +39,9 @@ export async function migrateCodexConfig({ env = process.env, cwd = process.cwd(
31
39
  const result = await migrateConfigFile(configPath, {
32
40
  catalog,
33
41
  previousState: state.files?.[configPath],
42
+ env,
43
+ sessionModel,
44
+ requireSessionModel,
34
45
  });
35
46
  if (result.changed) changed.push(configPath);
36
47
  if (result.multiAgentModeChanged) modeChanged.push(configPath);
@@ -44,7 +55,16 @@ export async function migrateCodexConfig({ env = process.env, cwd = process.cwd(
44
55
  return { changed, modeChanged };
45
56
  }
46
57
 
47
- export async function migrateConfigFile(configPath, { catalog = FALLBACK_CATALOG, previousState } = {}) {
58
+ export async function migrateConfigFile(
59
+ configPath,
60
+ {
61
+ catalog = FALLBACK_CATALOG,
62
+ previousState,
63
+ env = process.env,
64
+ sessionModel = null,
65
+ requireSessionModel = false,
66
+ } = {},
67
+ ) {
48
68
  const before = await readConfig(configPath);
49
69
  const decision = shouldApplyCatalog(before, catalog, previousState);
50
70
 
@@ -56,7 +76,12 @@ export async function migrateConfigFile(configPath, { catalog = FALLBACK_CATALOG
56
76
  reasoningApplied = config !== before;
57
77
  }
58
78
 
59
- const afterMultiAgentGuard = forceDisableMultiAgentV2(config);
79
+ const multiAgentOptions = { env, sessionModel, requireSessionModel };
80
+ const multiAgentVersion = resolveMultiAgentVersionFromConfig(config, multiAgentOptions);
81
+ const afterMultiAgentGuard = forceDisableMultiAgentV2(config, {
82
+ ...multiAgentOptions,
83
+ multiAgentVersion,
84
+ });
60
85
  const multiAgentChanged = afterMultiAgentGuard !== config;
61
86
  if (multiAgentChanged) config = afterMultiAgentGuard;
62
87
 
@@ -68,7 +93,10 @@ export async function migrateConfigFile(configPath, { catalog = FALLBACK_CATALOG
68
93
  const context7PlaceholderChanged = afterContext7PlaceholderGuard !== config;
69
94
  if (context7PlaceholderChanged) config = afterContext7PlaceholderGuard;
70
95
 
71
- const afterSubagentLimit = ensureSubagentConcurrencyLimit(config);
96
+ const afterSubagentLimit = ensureSubagentConcurrencyLimit(config, {
97
+ ...multiAgentOptions,
98
+ multiAgentVersion,
99
+ });
72
100
  const subagentLimitChanged = afterSubagentLimit !== config;
73
101
  if (subagentLimitChanged) config = afterSubagentLimit;
74
102
 
@@ -52,7 +52,7 @@ This skill may include examples copied from the OpenCode harness. In Codex, do n
52
52
 
53
53
  Role-specific behavior must be described in a self-contained \`message\`. Use \`fork_context: false\` to start the child with only the initial prompt (no parent history); use \`fork_context: true\` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's \`message\`. OMO installs these selectable agent roles into \`~/.codex/agents/\`: \`explorer\`, \`librarian\`, \`plan\`, \`momus\`, \`metis\`, \`lazycodex-code-reviewer\`, \`lazycodex-qa-executor\`, and \`lazycodex-gate-reviewer\` - pass the matching name as \`agent_type\` so the child gets that role's model and instructions. If the spawn tool exposes no \`agent_type\` parameter, omit it and describe the role inside \`message\`. If a code block below conflicts with this section, this section wins.
54
54
 
55
- On \`multi_agent_v2\` sessions the same \`agent_type\` applies (the OMO installer exposes it) with \`fork_turns\` instead of \`fork_context\`. If a code block below conflicts with this section, this section wins.
55
+ Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If \`multi_agent_v1.*\` tools exist, use the table above as written. If instead a flat \`spawn_agent\` with a required \`task_name\` exists (\`multi_agent_v2\`), rewrite every \`multi_agent_v1.*\` example: \`multi_agent_v1.spawn_agent({...,"fork_context":false})\` becomes \`spawn_agent({"task_name":"<lowercase_digits_underscores>","message":...,"agent_type":...,"fork_turns":"none"})\` (\`"all"\` only when full parent history is truly required); \`send_input\` becomes \`send_message\`; do not call \`close_agent\`/\`resume_agent\` (finished agents end on their own; \`followup_task\` re-tasks one, \`interrupt_agent\` stops one); \`wait_agent\` takes only \`timeout_ms\` and returns on any child mailbox activity. \`agent_type\` works the same on both surfaces. If a code block below conflicts with this section, this section wins.
56
56
 
57
57
  When translating \`load_skills=[...]\`, include the requested skill names in the spawned agent's \`message\`. If a code block below conflicts with this section, this section wins.
58
58
 
@@ -16,7 +16,7 @@ This file is a router, not a rulebook. The rules live in four rulesets under `re
16
16
  | ANY UI implementation, styling, redesign, mockup, or visual decision | `references/design/README.md` FIRST. It enforces two mandatory gates — the Design System Gate (a `DESIGN.md` must exist before any component is written) and the React Dev Tooling Gate (react-grab / react-scan / react-doctor installed by default) — then routes to the taste and brand references below. |
17
17
  | Writing or modifying frontend code, OR auditing performance / SEO / accessibility / quality | ALSO `references/perfection/README.md`. Lighthouse 100 in every category, measured on real Playwright Chromium (never the `lighthouse` CLI), achieved through architecture — never by dropping animations or hiding content. |
18
18
  | Looking up a concrete style, color palette, font pairing, chart type, landing-page structure, or UX guideline — or generating a project design system from keywords | `references/ui-ux-db/README.md`. A searchable CSV database with a CLI; a lookup tool, not a posture. Load on demand; `design` stays the source of truth for taste and the `DESIGN.md` contract. |
19
- | Design operating-layer work: personas, cognitive accessibility, design critique, design debt, handoff, synthetic user testing, or designpowers-style guidance | `references/designpowers/README.md`. This is an internal frontend ruleset, not a separate skill. It enriches `/frontend` routing with design brief, role-reference, accessibility, evidence, and debt language while preserving the `design` and `perfection` gates. |
19
+ | ANY implementation or redesign that creates or updates `DESIGN.md` — plus explicit operating-layer asks (personas, critique, debt, handoff, synthetic user testing) | `references/designpowers/README.md` + `references/designpowers/lane-c-review.md`. An internal frontend ruleset, not a separate skill: lane-c is the Phase Final flatness/critique reviewer, and its accessibility-constraints and accepted-debt language fills the required `DESIGN.md` sections. Load other lanes only when their phase applies. |
20
20
 
21
21
  **For implementation work, design + perfection load together.** A page that hits Lighthouse 100 but looks like AI slop has failed; a page that looks beautiful but ships a 2 MB bundle has failed. Both win or neither does.
22
22
 
@@ -28,15 +28,15 @@ Every implementation must choose one of these branches before UI code changes:
28
28
  - **Static visual reference** (screenshot, generated mockup, Stitch/Imagen output, Figma export, overview, or annotated packet): load `references/design/image-to-code-skill.md` plus the relevant design/perfection files, extract the reference's exact tokens, layout geometry, copy, spacing, states, and responsive intent into `DESIGN.md`, then implement reusable primitives against that contract.
29
29
  - **Live site or URL reference** (the user names a site to clone or gives a URL): load `references/design/clone-from-url.md`. Drive a real browser and extract the runtime truth via `getComputedStyle` — tokens, layout geometry, default/hover/focus/active states, transitions and keyframes, and downloaded assets — into `DESIGN.md`, then clone-code reusable primitives against that contract.
30
30
  Final QA for both runs `/visual-qa` in reference-fidelity mode: compare the actual UI against the reference pixel-by-pixel and verify the code is an extensible design-system implementation, not a screenshot-matched one-off.
31
- 2. **Greenfield or fresh setup:** if the user gave no concrete visual reference, design direction is a research deliverable, not a vibe. Fire every available research lane IN PARALLEL before `DESIGN.md` is written; skip a lane only when its tool or network is genuinely unavailable, and name the skip in `DESIGN.md`:
32
- - **Embedded references:** use `references/design/_INDEX.md` to shortlist 2-3 plausible Layer B references, then deeply load exactly one Layer A style skill and one Layer B brand/design-system reference; use `open-design` only when the curated set has no fit; add `ui-ux-db` lookups for palette/type/domain questions.
33
- - **Lazyweb real-product screens:** run the curl-only recipe in `references/design/lazyweb.md` to see how shipped products in the target domain actually look; harvest layout grammar and component patterns, never pixel copies.
34
- - **Imagen concept drafts:** generate 2-3 imagen concept drafts, each seeded with the loaded Layer A + Layer B tokens (palette, type, material); pick the strongest and treat the chosen draft as the reference-fidelity contract.
35
- Synthesize every lane into `DESIGN.md`. Treat sources as source material, not mood labels: extract tokens, layout grammar, component anatomy, interaction states, motion, and taste decisions, then recombine them into project-specific primitives. Never freestyle past the selected references, never copy logos or brand-specific copy. Define Section 5 primitives and their default/hover/active/focus/disabled/loading/empty/error states before code, and pass each through mobile/tablet/desktop visual QA before product screens.
31
+ 2. **Greenfield or fresh setup:** if the user gave no concrete visual reference, design research is a build step with named deliverables — not exploration to be budgeted. Exploration-stop instincts ("enough exploration", two-wave caps) do not apply here. Fire every research lane IN PARALLEL before `DESIGN.md` is written, and open `DESIGN.md` with a `## 0. Research Log` section recording each lane's deliverable — a lane with no Research Log line did not run. Skip a lane only when its tool or network is genuinely unavailable, and name the skip in `DESIGN.md`:
32
+ - **Embedded references:** use `references/design/_INDEX.md` to shortlist 2-3 plausible Layer B references, then read exactly one Layer A style skill and one Layer B reference in full — every line, no partial reads (they are 200-500 lines; a sliced read produces the flattened token set this gate exists to prevent). Log the shortlist, the pick, and why. Use `open-design` only when the curated set has no fit; add `ui-ux-db` lookups for palette/type/domain questions.
33
+ - **Lazyweb real-product screens:** READ `references/design/lazyweb.md` FIRST and run its recipe verbatim — do not improvise curl calls against lazyweb.com; the recipe mints its own anonymous token. Log the queries run, how many screens you actually VIEWED, and the layout grammar harvested never pixel copies.
34
+ - **Imagen concept drafts:** generate 2-3 imagen concept drafts, each seeded with the loaded Layer A + Layer B tokens (palette, type, material); pick the strongest and treat the chosen draft as the reference-fidelity contract. Log the draft paths and the pick.
35
+ Synthesize every lane into `DESIGN.md`. Treat sources as source material, not mood labels: extract tokens, layout grammar, component anatomy, interaction states, motion, and taste decisions, then recombine them into project-specific primitives. Never freestyle past the selected references, never copy logos or brand-specific copy. Then run the Primitive Showcase Gate (`references/design/README.md` Phase 0) before any product screen.
36
36
  3. **Existing project with `DESIGN.md` or a component system:** read it, follow it, and update it before implementation only when the requested work needs a new token, primitive, state, motion rule, accessibility constraint, accepted debt, or reference-fidelity requirement.
37
37
  4. **Existing project with UI but no `DESIGN.md` and no reusable component layer:** STOP and ask the user one focused question: should you preserve the current look with copy-nearby styling, or extract a real `DESIGN.md` plus reusable components before continuing? Do not silently choose.
38
38
 
39
- When `references/designpowers/README.md` is loaded for implementation, redesign, or design-system work, feed its personas, accessibility, critique, debt, handoff, and role-reference guidance into the branch above. The resulting `DESIGN.md` is the implementation contract: tokens, typography, spacing, primitives, motion, responsive behavior, accessibility constraints, and accepted debt must be named there before code uses them. Verify component primitives, states, and final screens with real visual QA evidence; pass design-system decisions, implementation evidence, and unresolved debt into `/review-work` for significant implementation work.
39
+ For implementation, redesign, or design-system work that creates or updates `DESIGN.md`, `references/designpowers/README.md` + `lane-c-review.md` are part of the default load feed their personas, accessibility, critique, debt, handoff, and role-reference guidance into the branch above. The resulting `DESIGN.md` is the implementation contract: tokens, typography, spacing, primitives, motion, responsive behavior, accessibility constraints, and accepted debt must be named there before code uses them. Verify component primitives, states, and final screens with real visual QA evidence; pass design-system decisions, implementation evidence, and unresolved debt into `/review-work` for significant implementation work.
40
40
 
41
41
  ## Ruleset 1 — design (`references/design/`)
42
42
 
@@ -46,7 +46,7 @@ The reference library has one architecture file, 12 taste skills (Layer A — *h
46
46
 
47
47
  | File | Read when |
48
48
  |---|---|
49
- | `design-system-architecture.md` | The project has no `DESIGN.md` (defines the 7-section structure you must create first), or you are extracting a design system from existing UI code. |
49
+ | `design-system-architecture.md` | The project has no `DESIGN.md` (defines the structure you must create first — 8 sections plus a greenfield-only `## 0. Research Log`), or you are extracting a design system from existing UI code. |
50
50
 
51
51
  ### Layer A — taste skills (pick AT MOST ONE style skill; they encode opposing philosophies)
52
52
 
@@ -102,7 +102,7 @@ Domains: `product` `style` `typography` `color` `landing` `chart` `ux` `react` `
102
102
 
103
103
  ## Ruleset 4 — designpowers (`references/designpowers/`)
104
104
 
105
- `README.md` routes design operating-layer guidance from the pinned `Owl-Listener/designpowers` reference corpus into the existing frontend workflow. Load it when a frontend task needs explicit personas, accessibility and cognitive constraints, design critique, design debt, handoff, synthetic user testing, motion guidance, or role-reference prompts. It does not replace this frontend skill, `/visual-qa`, `/ulw-plan`, `/start-work`, or `/review-work`; it supplies richer design context that must first be distilled into the project `DESIGN.md`, then used as the design-system contract for implementation and verification.
105
+ `README.md` routes design operating-layer guidance from the pinned `Owl-Listener/designpowers` reference corpus into the existing frontend workflow. Load it — together with `lane-c-review.md` — for every implementation or redesign that creates or updates `DESIGN.md`, and additionally when a task needs explicit personas, accessibility and cognitive constraints, design critique, design debt, handoff, synthetic user testing, motion guidance, or role-reference prompts. It does not replace this frontend skill, `/visual-qa`, `/ulw-plan`, `/start-work`, or `/review-work`; it supplies richer design context that must first be distilled into the project `DESIGN.md`, then used as the design-system contract for implementation and verification.
106
106
 
107
107
  ## Quick routes — most common requests
108
108
 
@@ -38,12 +38,12 @@ Before touching any UI code, before routing to any reference, before even thinki
38
38
 
39
39
  1. Read `design-system-architecture.md` — it defines the exact structure.
40
40
  2. Identify the branch: greenfield setup, existing UI with implicit patterns/components, or existing UI with no reusable component layer.
41
- 3. **Greenfield setup:** if the user gave no concrete visual reference, use `_INDEX.md` to shortlist 2-3 plausible Layer B references, then deeply load exactly one Layer A style skill and one Layer B brand/design-system reference; use `open-design` only when the curated set has no fit. Treat those references as source material, not mood labels: extract tokens, layout grammar, component anatomy, interaction states, motion, and taste decisions into `DESIGN.md`, then recombine them into project-specific primitives. Customize for the user's product and content, but do not freestyle past the selected references; never copy logos, trademarked assets, or brand-specific copy.
41
+ 3. **Greenfield setup:** if the user gave no concrete visual reference, use `_INDEX.md` to shortlist 2-3 plausible Layer B references, then read exactly one Layer A style skill and one Layer B brand/design-system reference in full — every line, no partial reads; use `open-design` only when the curated set has no fit. Open `DESIGN.md` with a `## 0. Research Log` recording each research lane's deliverable (embedded-reference shortlist + pick, lazyweb screens viewed, imagen drafts — see the SKILL.md workflow); a lane with no line did not run. Treat those references as source material, not mood labels: extract tokens, layout grammar, component anatomy, interaction states, motion, and taste decisions into `DESIGN.md`, then recombine them into project-specific primitives. Customize for the user's product and content, but do not freestyle past the selected references; never copy logos, trademarked assets, or brand-specific copy.
42
42
  - **Commit a distinctive direction BEFORE extracting tokens.** In 1-2 sentences, name the atmosphere, the signature material, the color story, and the one moment a visitor will remember. For an expressive brief, sketch 2-3 genuinely different directions and pick the boldest one you can defend with the loaded reference; do not average them, because the average IS the generic default this skill exists to beat. A locked, never-revisited one-shot decision is how a page ends up flat.
43
43
  - **The reference's distinctive material MUST survive extraction (expressive briefs).** The common failure is loading a rich reference and then distilling it into a generic dark-SaaS token set. Your `DESIGN.md` must carry the *non-default* decisions forward and name which reference each came from: the actual elevation recipe (the specific layers that make a surface read as glass/glossy, not a single blur), a multi-stop perceptual color ramp (not one brand hex reused at varied opacity), the explicit display/body/mono type choices, and one signature interaction. Self-check before writing code: if your `DESIGN.md` could describe any generic dark SaaS, you flattened the reference — go back and put the specific material in.
44
44
  4. **Existing UI with implicit patterns/components:** extract the colors, typography, spacing, primitives, states, and motion already in use. Write `DESIGN.md` to codify what exists before changing UI code.
45
45
  5. **Existing UI with no reusable component layer:** STOP and ask whether to preserve the current style with copy-nearby edits or extract a `DESIGN.md` plus reusable components first. Do not silently choose the cheaper path or the larger refactor.
46
- 6. **Do not proceed to product screens until `DESIGN.md` exists, Section 5 names the reusable primitives and their states, and each primitive plus required state passes mobile/tablet/desktop visual QA in a component showcase or equivalent state harness.**
46
+ 6. Finish the triage at the Primitive Showcase Gate below.
47
47
 
48
48
  #### If YES design system exists → READ IT, FOLLOW IT
49
49
 
@@ -52,7 +52,11 @@ Before touching any UI code, before routing to any reference, before even thinki
52
52
  3. If you need a token that doesn't exist, **add it to `DESIGN.md` first**, then use it.
53
53
  4. Never introduce raw hex codes, arbitrary px values, or ad-hoc component patterns that bypass the system.
54
54
 
55
- **This gate is non-negotiable. No design system = no UI work. Period.**
55
+ **The Design System Gate is non-negotiable. No design system = no UI work. Period.**
56
+
57
+ ### Primitive Showcase Gate (MANDATORY)
58
+
59
+ **Do not proceed to product screens until `DESIGN.md` exists, Section 5 names the reusable primitives and their states, and each primitive plus required state passes mobile/tablet/desktop visual QA in a component showcase or equivalent state harness.** Skipping this gate ships ad-hoc-styled product screens and re-enters the redesign loop.
56
60
 
57
61
 
58
62
  ## Phase 0.5 — React Dev Tooling Gate (MANDATORY for React projects)
@@ -13,7 +13,7 @@ All reference files live flat in this directory. Three layers:
13
13
 
14
14
  | File | Purpose | Load when |
15
15
  |---|---|---|
16
- | `design-system-architecture.md` | Defines the 7-section `DESIGN.md` structure (atmosphere, color tokens, typography scale, spacing system, components, motion, depth). Creation workflow for new and existing projects. Validation rules and memory management. | Phase 0 fires and no `DESIGN.md` exists in the project. Also load when extracting a design system from existing code. |
16
+ | `design-system-architecture.md` | Defines the `DESIGN.md` structure — 8 sections plus a greenfield-only `## 0. Research Log` (atmosphere, color tokens, typography scale, spacing system, components, motion, depth, accessibility constraints & accepted debt). Creation workflow for new and existing projects. Validation rules and memory management. | Phase 0 fires and no `DESIGN.md` exists in the project. Also load when extracting a design system from existing code. |
17
17
 
18
18
  ---
19
19
 
@@ -16,11 +16,19 @@ Every frontend project MUST have a `DESIGN.md` at its root. This file is the sin
16
16
 
17
17
  ## DESIGN.md Structure
18
18
 
19
- The file has 7 sections. Every section is mandatory. Skip nothing.
19
+ The file has 8 sections plus a greenfield-only `## 0. Research Log`. Every section is mandatory. Skip nothing.
20
20
 
21
21
  ```markdown
22
22
  # [Project Name] Design System
23
23
 
24
+ ## 0. Research Log (greenfield only)
25
+
26
+ One line per research lane, written before the sections below — a lane with no line did not run:
27
+ - Embedded refs: shortlisted [2-3 Layer B candidates] → picked [Layer A] + [Layer B] because [reason]
28
+ - Lazyweb: [N] queries, [M] screens viewed → [layout grammar taken]
29
+ - Imagen drafts: [paths] → picked [draft] as the reference-fidelity contract
30
+ - Skipped lanes: [lane] — [tool/network reason]
31
+
24
32
  ## 1. Atmosphere & Identity
25
33
 
26
34
  One paragraph. What this product FEELS like. Not what it does — how it feels to use.
@@ -165,6 +173,19 @@ If borders:
165
173
 
166
174
  If tonal-shift:
167
175
  Surfaces use progressively lighter/darker shades. No borders, no shadows.
176
+
177
+ ## 8. Accessibility Constraints & Accepted Debt
178
+
179
+ ### Constraints
180
+ - WCAG target: [e.g. 2.2 AA] — contrast floor [4.5:1 body / 3:1 large text], visible focus on every
181
+ interactive element, full keyboard reachability, `prefers-reduced-motion` respected (Section 6).
182
+
183
+ ### Accepted Debt
184
+ | Item | Location | Why accepted | Owner / Exit |
185
+ |------|----------|--------------|--------------|
186
+ | [debt] | [file/screen] | [reason + user sign-off] | [when it gets fixed] |
187
+
188
+ New debt is recorded here at the moment it is accepted — never silently.
168
189
  ```
169
190
 
170
191
  ## Creation Workflow
@@ -173,7 +194,7 @@ Surfaces use progressively lighter/darker shades. No borders, no shadows.
173
194
 
174
195
  1. **Select references before taste** — no visual reference means `_INDEX.md` shortlist of 2-3 Layer B candidates, then exactly one Layer A style skill and one Layer B brand/design-system reference. Use `open-design` only when the curated set has no fit.
175
196
  2. **Assemble from references** — extract tokens, layout grammar, component anatomy, states, motion, and taste decisions, then recombine them into project-specific primitives. Customize for the user's product; never copy logos, trademarked assets, or brand-specific copy.
176
- 3. **Define the system** — atmosphere, palette, typography, spacing, and one depth strategy, grounded in the selected references and product semantics.
197
+ 3. **Define the system** — atmosphere, palette, typography, spacing, and one depth strategy, grounded in the selected references and product semantics. Sanity-check the palette and type pairing with one `ui-ux-db` domain search (CLI in `references/ui-ux-db/README.md`).
177
198
  4. **Document initial primitives** — only components you are about to build, including variants and states.
178
199
  5. **Write it to `DESIGN.md`** at project root.
179
200
  6. **Build a primitive showcase first** — exercise each primitive's default, hover, active, focus, disabled, loading, empty, and error states at mobile/tablet/desktop widths before composing product screens.
@@ -199,6 +220,7 @@ After every component implementation, check:
199
220
  - [ ] Component reused 2+ times? Documented in Section 5.
200
221
  - [ ] Motion follows the timing table. No arbitrary durations.
201
222
  - [ ] Component visual QA passed for each primitive and required state before product screens were composed.
223
+ - [ ] Section 8 accessibility constraints hold for the new component; any new debt is recorded in Section 8, not silently accepted.
202
224
 
203
225
  ## Memory Management
204
226
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  This is an internal frontend ruleset, not a standalone skill. `/frontend` remains the only public activation point for web UI, UX, visual design, accessibility, design QA, and frontend implementation routing.
4
4
 
5
- Load this reference from the frontend router when a task needs design operating-layer guidance: personas, cognitive accessibility, critique, design debt, handoff, synthetic user testing, motion guidance, or designpowers-style role references.
5
+ Load this reference from the frontend router for EVERY implementation or redesign that creates or updates `DESIGN.md`, and whenever a task needs design operating-layer guidance: personas, cognitive accessibility, critique, design debt, handoff, synthetic user testing, motion guidance, or designpowers-style role references.
6
6
 
7
7
  The purpose of this ruleset is to enrich the existing frontend workflow while preserving its gates:
8
8
 
@@ -18,7 +18,7 @@ Read these files before applying designpowers guidance:
18
18
  1. `README.md` - this frontend integration contract.
19
19
  2. `routing.md` - how designpowers context feeds existing frontend, planning, execution, visual QA, and review routes.
20
20
  3. `orchestration.md` - shared state, Direct/Auto prompt semantics, safeguards, and role-reference rules.
21
- 4. Phase lane docs, loaded only when relevant:
21
+ 4. Phase lane docs `lane-c-review.md` loads with this README for every implementation or redesign heading into Phase Final review (it is the flatness/critique reviewer); the other lanes load only when their phase applies:
22
22
  - `lane-a-direction.md` for planning, direction, discovery, personas, taste, and accessibility constraints.
23
23
  - `lane-b-execution.md` for execution, UI build prompts, frontend handoff, and implementation evidence.
24
24
  - `lane-c-review.md` for visual QA, design critique, review gates, and objective evidence before judgment.
@@ -44,7 +44,7 @@ Lane B worker DoneClaims must include:
44
44
 
45
45
  - Exact changed files and the `frontend` references loaded.
46
46
  - The real-surface QA invocation required by `start-work` for the UI surface, with captured artifact path.
47
- - Screenshot, browser, HTTP, or tmux artifacts appropriate to the visible surface.
47
+ - Screenshot, browser, HTTP, or xterm.js web-terminal artifacts appropriate to the visible surface.
48
48
  - Accessibility evidence from the existing OpenAgent frontend path, such as Lighthouse, react-doctor, keyboard checks, or other plan-required checks.
49
49
  - A short design trace: which persona, design principle, token, or state requirement each major UI decision satisfies.
50
50
  - Cleanup receipts for any browser session, server, tmux session, temporary artifact, or process used during QA.
@@ -18,7 +18,7 @@ This skill may include examples copied from the OpenCode harness. In Codex, do n
18
18
 
19
19
  Role-specific behavior must be described in a self-contained `message`. Use `fork_context: false` to start the child with only the initial prompt (no parent history); use `fork_context: true` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's `message`. OMO installs these selectable agent roles into `~/.codex/agents/`: `explorer`, `librarian`, `plan`, `momus`, `metis`, `lazycodex-code-reviewer`, `lazycodex-qa-executor`, and `lazycodex-gate-reviewer` - pass the matching name as `agent_type` so the child gets that role's model and instructions. If the spawn tool exposes no `agent_type` parameter, omit it and describe the role inside `message`. If a code block below conflicts with this section, this section wins.
20
20
 
21
- On `multi_agent_v2` sessions the same `agent_type` applies (the OMO installer exposes it) with `fork_turns` instead of `fork_context`. If a code block below conflicts with this section, this section wins.
21
+ Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If `multi_agent_v1.*` tools exist, use the table above as written. If instead a flat `spawn_agent` with a required `task_name` exists (`multi_agent_v2`), rewrite every `multi_agent_v1.*` example: `multi_agent_v1.spawn_agent({...,"fork_context":false})` becomes `spawn_agent({"task_name":"<lowercase_digits_underscores>","message":...,"agent_type":...,"fork_turns":"none"})` (`"all"` only when full parent history is truly required); `send_input` becomes `send_message`; do not call `close_agent`/`resume_agent` (finished agents end on their own; `followup_task` re-tasks one, `interrupt_agent` stops one); `wait_agent` takes only `timeout_ms` and returns on any child mailbox activity. `agent_type` works the same on both surfaces. If a code block below conflicts with this section, this section wins.
22
22
 
23
23
  When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`. If a code block below conflicts with this section, this section wins.
24
24
 
@@ -19,7 +19,7 @@ This skill may include examples copied from the OpenCode harness. In Codex, do n
19
19
 
20
20
  Role-specific behavior must be described in a self-contained `message`. Use `fork_context: false` to start the child with only the initial prompt (no parent history); use `fork_context: true` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's `message`. OMO installs these selectable agent roles into `~/.codex/agents/`: `explorer`, `librarian`, `plan`, `momus`, `metis`, `lazycodex-code-reviewer`, `lazycodex-qa-executor`, and `lazycodex-gate-reviewer` - pass the matching name as `agent_type` so the child gets that role's model and instructions. If the spawn tool exposes no `agent_type` parameter, omit it and describe the role inside `message`. If a code block below conflicts with this section, this section wins.
21
21
 
22
- On `multi_agent_v2` sessions the same `agent_type` applies (the OMO installer exposes it) with `fork_turns` instead of `fork_context`. If a code block below conflicts with this section, this section wins.
22
+ Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If `multi_agent_v1.*` tools exist, use the table above as written. If instead a flat `spawn_agent` with a required `task_name` exists (`multi_agent_v2`), rewrite every `multi_agent_v1.*` example: `multi_agent_v1.spawn_agent({...,"fork_context":false})` becomes `spawn_agent({"task_name":"<lowercase_digits_underscores>","message":...,"agent_type":...,"fork_turns":"none"})` (`"all"` only when full parent history is truly required); `send_input` becomes `send_message`; do not call `close_agent`/`resume_agent` (finished agents end on their own; `followup_task` re-tasks one, `interrupt_agent` stops one); `wait_agent` takes only `timeout_ms` and returns on any child mailbox activity. `agent_type` works the same on both surfaces. If a code block below conflicts with this section, this section wins.
23
23
 
24
24
  When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`. If a code block below conflicts with this section, this section wins.
25
25
 
@@ -19,7 +19,7 @@ This skill may include examples copied from the OpenCode harness. In Codex, do n
19
19
 
20
20
  Role-specific behavior must be described in a self-contained `message`. Use `fork_context: false` to start the child with only the initial prompt (no parent history); use `fork_context: true` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's `message`. OMO installs these selectable agent roles into `~/.codex/agents/`: `explorer`, `librarian`, `plan`, `momus`, `metis`, `lazycodex-code-reviewer`, `lazycodex-qa-executor`, and `lazycodex-gate-reviewer` - pass the matching name as `agent_type` so the child gets that role's model and instructions. If the spawn tool exposes no `agent_type` parameter, omit it and describe the role inside `message`. If a code block below conflicts with this section, this section wins.
21
21
 
22
- On `multi_agent_v2` sessions the same `agent_type` applies (the OMO installer exposes it) with `fork_turns` instead of `fork_context`. If a code block below conflicts with this section, this section wins.
22
+ Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If `multi_agent_v1.*` tools exist, use the table above as written. If instead a flat `spawn_agent` with a required `task_name` exists (`multi_agent_v2`), rewrite every `multi_agent_v1.*` example: `multi_agent_v1.spawn_agent({...,"fork_context":false})` becomes `spawn_agent({"task_name":"<lowercase_digits_underscores>","message":...,"agent_type":...,"fork_turns":"none"})` (`"all"` only when full parent history is truly required); `send_input` becomes `send_message`; do not call `close_agent`/`resume_agent` (finished agents end on their own; `followup_task` re-tasks one, `interrupt_agent` stops one); `wait_agent` takes only `timeout_ms` and returns on any child mailbox activity. `agent_type` works the same on both surfaces. If a code block below conflicts with this section, this section wins.
23
23
 
24
24
  When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`. If a code block below conflicts with this section, this section wins.
25
25
 
@@ -18,7 +18,7 @@ This skill may include examples copied from the OpenCode harness. In Codex, do n
18
18
 
19
19
  Role-specific behavior must be described in a self-contained `message`. Use `fork_context: false` to start the child with only the initial prompt (no parent history); use `fork_context: true` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's `message`. OMO installs these selectable agent roles into `~/.codex/agents/`: `explorer`, `librarian`, `plan`, `momus`, `metis`, `lazycodex-code-reviewer`, `lazycodex-qa-executor`, and `lazycodex-gate-reviewer` - pass the matching name as `agent_type` so the child gets that role's model and instructions. If the spawn tool exposes no `agent_type` parameter, omit it and describe the role inside `message`. If a code block below conflicts with this section, this section wins.
20
20
 
21
- On `multi_agent_v2` sessions the same `agent_type` applies (the OMO installer exposes it) with `fork_turns` instead of `fork_context`. If a code block below conflicts with this section, this section wins.
21
+ Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If `multi_agent_v1.*` tools exist, use the table above as written. If instead a flat `spawn_agent` with a required `task_name` exists (`multi_agent_v2`), rewrite every `multi_agent_v1.*` example: `multi_agent_v1.spawn_agent({...,"fork_context":false})` becomes `spawn_agent({"task_name":"<lowercase_digits_underscores>","message":...,"agent_type":...,"fork_turns":"none"})` (`"all"` only when full parent history is truly required); `send_input` becomes `send_message`; do not call `close_agent`/`resume_agent` (finished agents end on their own; `followup_task` re-tasks one, `interrupt_agent` stops one); `wait_agent` takes only `timeout_ms` and returns on any child mailbox activity. `agent_type` works the same on both surfaces. If a code block below conflicts with this section, this section wins.
22
22
 
23
23
  When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`. If a code block below conflicts with this section, this section wins.
24
24
 
@@ -19,7 +19,7 @@ This skill may include examples copied from the OpenCode harness. In Codex, do n
19
19
 
20
20
  Role-specific behavior must be described in a self-contained `message`. Use `fork_context: false` to start the child with only the initial prompt (no parent history); use `fork_context: true` only when full parent history is truly required. Include any required conversation context, files, diffs, constraints, and requested skill names directly in the spawned agent's `message`. OMO installs these selectable agent roles into `~/.codex/agents/`: `explorer`, `librarian`, `plan`, `momus`, `metis`, `lazycodex-code-reviewer`, `lazycodex-qa-executor`, and `lazycodex-gate-reviewer` - pass the matching name as `agent_type` so the child gets that role's model and instructions. If the spawn tool exposes no `agent_type` parameter, omit it and describe the role inside `message`. If a code block below conflicts with this section, this section wins.
21
21
 
22
- On `multi_agent_v2` sessions the same `agent_type` applies (the OMO installer exposes it) with `fork_turns` instead of `fork_context`. If a code block below conflicts with this section, this section wins.
22
+ Codex exposes ONE of two subagent tool surfaces per session; check your own tool list and route accordingly. If `multi_agent_v1.*` tools exist, use the table above as written. If instead a flat `spawn_agent` with a required `task_name` exists (`multi_agent_v2`), rewrite every `multi_agent_v1.*` example: `multi_agent_v1.spawn_agent({...,"fork_context":false})` becomes `spawn_agent({"task_name":"<lowercase_digits_underscores>","message":...,"agent_type":...,"fork_turns":"none"})` (`"all"` only when full parent history is truly required); `send_input` becomes `send_message`; do not call `close_agent`/`resume_agent` (finished agents end on their own; `followup_task` re-tasks one, `interrupt_agent` stops one); `wait_agent` takes only `timeout_ms` and returns on any child mailbox activity. `agent_type` works the same on both surfaces. If a code block below conflicts with this section, this section wins.
23
23
 
24
24
  When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`. If a code block below conflicts with this section, this section wins.
25
25
 
@@ -108,10 +108,10 @@ Each sub-task message must include:
108
108
  4. Automated verification commands to run.
109
109
  5. One Manual-QA channel, named with the exact tool and exact invocation (the literal `curl`, `send-keys`, `browser:control-in-app-browser` action, `page.click`, payload, selectors, and the binary observable that decides PASS/FAIL), not "verify it works". A LIGHT checkbox needs one real-surface proof of its deliverable, and auxiliary surfaces (CLI stdout, DB state diff, parsed config dump) are first-class when the surface is CLI- or data-shaped:
110
110
  - HTTP call: `curl -i` against the live endpoint.
111
- - tmux: a `tmux` session driven with `send-keys`, dumped via `capture-pane`.
111
+ - Terminal / TUI: drive a real pty; `tmux send-keys` is fine for a boot/behavior smoke, but color/layout/CJK evidence goes through the xterm.js web terminal below, NEVER `tmux capture-pane`.
112
112
  - Browser use: in Codex, use `browser:control-in-app-browser` first when available and the scenario does not need an authenticated or persistent user browser profile; otherwise drive the real page with Chrome, or agent-browser (https://github.com/vercel-labs/agent-browser) when Chrome is unavailable.
113
113
  - Computer use: OS-level GUI automation against the running desktop app when the surface is not a page.
114
- - TUI visual evidence: when a tmux/TUI claim needs visual QA or PR proof, run `node script/qa/web-terminal-visual-qa.mjs --from-file <capture.txt> --evidence-dir <dir>` and attach `terminal.png` plus `metadata.json`.
114
+ - TUI visual evidence: when a TUI claim needs visual QA or PR proof, run `node script/qa/web-terminal-visual-qa.mjs --command "<cmd>" --input "{Enter}" --evidence-dir <dir>` (real pty rendered through xterm.js in Chrome) and attach `terminal.png` plus `metadata.json`.
115
115
  6. The adversarial classes that apply to this sub-task (from the 9 ultraqa classes) and how each is probed.
116
116
  7. Required artifact path and cleanup receipt.
117
117
 
@@ -19,7 +19,7 @@ Use a TEAM when EITHER holds:
19
19
  - one task still needs exploration, yet its GOAL is already clear - parallel investigation under
20
20
  a fixed objective.
21
21
 
22
- Use plain subagents (`$ulw` / `multi_agent_v1.spawn_agent`) - NOT a team - when EITHER holds:
22
+ Use plain subagents (`$ulw` / `multi_agent_v1.spawn_agent` / flat `spawn_agent` on `multi_agent_v2`) - NOT a team - when EITHER holds:
23
23
  - the work IS perfectly isolated, so there is no coordination cost worth paying; or
24
24
  - the GOAL is still ambiguous, where one mind should resolve direction before any fan-out.
25
25
 
@@ -39,7 +39,7 @@ merge), not the keystrokes.
39
39
  ## Compose by part, ownership, or perspective - not by job title
40
40
 
41
41
  A team is ALWAYS two or more members - never a single-member team. One worker on an isolated
42
- job is a subagent (`multi_agent_v1.spawn_agent`), not a team; if you end up with a single member,
42
+ job is a spawned subagent, not a team; if you end up with a single member,
43
43
  either split off a second distinct slice or drop the team and use a subagent.
44
44
 
45
45
  Compose the team from what you actually KNOW about the work. Ground the split in real knowledge
@@ -110,7 +110,7 @@ do not treat the intended mutation as complete; retry after the named command fi
110
110
  lose in the sidebar without this link.
111
111
 
112
112
  Every team member is a real Codex thread created with `codex_app.create_thread` - this is strict,
113
- not a preference. NEVER substitute `multi_agent_v1.spawn_agent`, or any other in-process subagent,
113
+ not a preference. NEVER substitute a spawned in-process subagent (`multi_agent_v1.spawn_agent` or flat `spawn_agent`)
114
114
  for a team member: a spawned agent is an ephemeral helper that does not show up as a team thread,
115
115
  cannot carry the `[team name] <member name>` title, and cannot be inspected, titled, archived, or
116
116
  re-opened with the `codex_app.*` thread tools - which defeats the entire point of a durable team.
@@ -56,9 +56,10 @@ exercises the surface; capture the artifact.
56
56
  1. HTTP call — hit the live endpoint with `curl -i` (or a
57
57
  Playwright APIRequestContext); capture status line + headers +
58
58
  body.
59
- 2. tmux `tmux new-session -d -s ulw-qa-<criterion>`, drive with
60
- `send-keys`, dump via `tmux capture-pane -pS -E -`; transcript
61
- is the artifact.
59
+ 2. Terminal / TUI - drive a real pty and prove it through the
60
+ xterm.js web terminal (see the TUI visual QA note below). tmux
61
+ `send-keys` is fine for a boot smoke; NEVER `tmux capture-pane`
62
+ for color / layout / CJK evidence, which degrades truecolor.
62
63
  3. Browser use — in Codex, use `browser:control-in-app-browser`
63
64
  first when available and no authenticated/persistent user browser
64
65
  profile is required. Otherwise use Chrome to drive the REAL page;
@@ -86,13 +87,13 @@ channel scenario when the behavior is user-facing. `--dry-run`,
86
87
  printing the command, "should respond", and "looks correct" never
87
88
  count.
88
89
 
89
- For TUI visual QA, terminal transcripts alone are not enough when a
90
- visual surface is being evaluated. In this repo, prefer
91
- `node script/qa/web-terminal-visual-qa.mjs --title "<surface>" --from-file <capture.txt> --evidence-dir <dir>`
92
- or the helper's `--command` tmux-backed PTY connector when available.
93
- Outside this repo, capture equivalent browser/computer-use rendered
94
- terminal evidence: screenshot, plain transcript, rendered HTML or action
95
- log, and cleanup receipt.
90
+ For TUI visual QA, render the terminal through the real xterm.js web
91
+ terminal and screenshot it - never a `tmux capture-pane` dump, which
92
+ degrades color and wide-glyph width. In this repo:
93
+ `node script/qa/web-terminal-visual-qa.mjs --title "<surface>" --command "<cmd>" --input "{Enter}" --evidence-dir <dir>`
94
+ (live pty + xterm.js in Chrome; `--from-file <capture>` replays a raw
95
+ stream). Outside this repo, capture equivalent browser-rendered terminal
96
+ evidence: screenshot + plain transcript + cleanup receipt.
96
97
 
97
98
  # Bootstrap (DO ALL FOUR BEFORE ANY OTHER WORK — NO SKIPPING)
98
99
 
@@ -265,6 +266,7 @@ Every `multi_agent_v1.spawn_agent` message is self-contained and starts with
265
266
  handoff. Use `fork_context: false` unless full history is truly
266
267
  required; paste only the context the child needs. Full-history forks can
267
268
  make the child continue old parent context instead of the delegated task.
269
+ If your tool list has a flat `spawn_agent` with a required `task_name` instead of `multi_agent_v1.*` (`multi_agent_v2`), rewrite: `fork_context: false` becomes `fork_turns: "none"`, `send_input` becomes `send_message`, finished agents end on their own (no `close_agent`; `followup_task` re-tasks, `interrupt_agent` stops), and `wait_agent` takes only `timeout_ms`, returning on any child mailbox activity.
268
270
 
269
271
  # TOML-backed subagent routing compatibility
270
272
  Treat TOML-backed role routing as **routing-unverified**. The
@@ -294,6 +296,13 @@ evidence for that step. Do not start dependent implementation until the
294
296
  audit, research, or review result is integrated or explicitly recorded
295
297
  as inconclusive. Do not generate a plan before spawned research lanes
296
298
  that feed the plan have returned or been closed as inconclusive.
299
+ Spawn every independent child for the current wave first. After the wave
300
+ is launched, run `multi_agent_v1.wait_agent` for each spawned child until
301
+ each reaches terminal status (`completed`, `failed`, `blocked`, or
302
+ explicitly recorded inconclusive) before any dependent `update_plan`
303
+ transition, `create_goal` continuation, implementation tool call, plan
304
+ drafting, approval-gate work, PR handoff, or final response. A timeout is
305
+ not terminal status.
297
306
  Do not write the final answer, PR handoff, or completion summary while
298
307
  active child agents remain open. Use short `multi_agent_v1.wait_agent` cycles.
299
308
  After two silent waits send `TASK STILL ACTIVE: return <deliverable> or
@@ -22,7 +22,7 @@ This skill is intentionally compact. The full workflow lives in `references/full
22
22
  - Use the ulw-loop CLI state under `.omo/ulw-loop`; do not hand-edit goal state.
23
23
  - After any compaction or context loss, re-read brief + goals + ledger FIRST plus `omo ulw-loop status --json`, then resume; never re-plan from scratch.
24
24
  - If `omo ulw-loop create-goals` says the existing aggregate is already complete, start unrelated new work with a fresh `--session-id <new-id>` instead of steering or forcing the completed default state. Use `--force` only to intentionally overwrite completed evidence.
25
- - Every success criterion needs observable evidence from a real surface: a channel (tmux, HTTP, browser, computer-use) or, for CLI- or data-shaped criteria, an auxiliary surface (CLI stdout, DB diff, parsed config dump).
25
+ - Every success criterion needs observable evidence from a real surface: a channel (terminal/TUI via the xterm.js web terminal, HTTP, browser, computer-use) or, for CLI- or data-shaped criteria, an auxiliary surface (CLI stdout, DB diff, parsed config dump).
26
26
  - Record evidence through the CLI only after cleanup receipts are available.
27
27
  - Delegate code edits, test writes, fixes, and QA execution to right-sized Codex subagents when the workflow requires it.
28
28
  - Every `multi_agent_v1.spawn_agent` message starts with `TASK:`, then names `DELIVERABLE`, `SCOPE`, and `VERIFY`; put role and specialty instructions inside `message`; use `fork_context: false` unless full history is truly required.
@@ -46,4 +46,6 @@ The full workflow may mention OpenCode-style orchestration examples. In Codex, t
46
46
  | Wait for background result | `multi_agent_v1.wait_agent(...)` |
47
47
  | Clean up finished worker | `multi_agent_v1.close_agent(...)` |
48
48
 
49
+ Flat `spawn_agent` requiring `task_name` instead (`multi_agent_v2`)? Rewrite rows: add `"task_name"`, `"fork_context":false` → `"fork_turns":"none"`, `wait_agent` takes only `timeout_ms`, no `close_agent` — finished agents end on their own.
50
+
49
51
  When translating `load_skills=[...]`, include the requested skill names in the spawned agent's `message`.
@@ -19,16 +19,16 @@ Audit each pass, fail, block, steering change, and checkpoint in `.omo/ulw-loop/
19
19
  Run each criterion's real-surface proof yourself through the channel that faithfully exercises it; capture the artifact before recording PASS.
20
20
 
21
21
  1. **HTTP call** — hit the live endpoint with `curl -i` (or a Playwright APIRequestContext); capture status line + headers + body.
22
- 2. **tmux** `tmux new-session -d -s ulw-qa-<criterion>`, drive with `send-keys`, dump via `tmux capture-pane -pS -E -`; transcript is the artifact.
22
+ 2. **Terminal / TUI** - prove it through the xterm.js web terminal; tmux `send-keys` is fine for a boot smoke, but NEVER `tmux capture-pane` for color/layout/CJK evidence (it degrades truecolor).
23
23
  3. **Browser use** — in Codex, use `browser:control-in-app-browser` first when available and the scenario does not need an authenticated or persistent user browser profile. Otherwise use Chrome to drive the REAL page; if unavailable, use agent-browser. Capture action log + screenshot path. Never downgrade a browser-facing criterion.
24
24
  4. **Computer use** — for desktop/GUI apps, drive the running app via OS automation (computer-use, AppleScript, xdotool, etc.); capture action log + screenshot.
25
25
 
26
- For TUI visual QA, pair the tmux transcript with a browser-rendered terminal
27
- screenshot. In this repo run `node script/qa/web-terminal-visual-qa.mjs
28
- --from-file <capture.txt> --evidence-dir <dir>` and record `terminal.png`,
29
- `terminal.html`, `terminal.txt`, and `metadata.json` as the visual evidence
30
- bundle. This is mandatory when a PR or review needs to inspect the terminal
31
- screen, not just the text.
26
+ For TUI visual QA, render the terminal through the real xterm.js web terminal and
27
+ screenshot it - NEVER a `tmux capture-pane` dump (it degrades color and wide-glyph
28
+ width). In this repo run `node script/qa/web-terminal-visual-qa.mjs --command
29
+ "<cmd>" --input "{Enter}" --evidence-dir <dir>` (live pty + xterm.js in Chrome;
30
+ `--from-file` replays a raw stream) and record `terminal.png`, `terminal.txt`, and
31
+ `metadata.json`. Mandatory when a PR or review must inspect the terminal screen.
32
32
 
33
33
  Auxiliary surfaces (CLI stdout / DB state diff / parsed config dump) are first-class evidence for CLI- or data-shaped criteria; use a channel scenario when the behavior is user-facing. `--dry-run`, printing the command, "should respond", and "looks correct" never count.
34
34
 
@@ -67,6 +67,19 @@ Fan out read-only research before deciding. Every spawn names DELIVERABLE / SCOP
67
67
  multi_agent_v1.spawn_agent({"message":"TASK: act as an explorer. DELIVERABLE: ... SCOPE: ... VERIFY: ...","agent_type":"explorer","fork_context":false})
68
68
  ```
69
69
 
70
+ If your tool list has a flat `spawn_agent` with a required `task_name` instead of `multi_agent_v1.*` (`multi_agent_v2`), rewrite: add `"task_name":"<lowercase_digits_underscores>"`, replace `"fork_context":false` with `"fork_turns":"none"`, and `wait_agent` takes only `timeout_ms`, returning on any child mailbox activity (finished agents end on their own).
71
+
72
+ Spawn every independent child for the current wave first. After the wave
73
+ is launched, use `multi_agent_v1.wait_agent` for each child until each
74
+ reaches terminal status. A timeout is not terminal status. Do not start dependent planning, drafting, approval-gate work, or final handoff until each child result is integrated or recorded as inconclusive.
75
+
76
+ For work likely to exceed one wait cycle, require the child to send
77
+ `WORKING: <task> - <current phase>` before long passes and
78
+ `BLOCKED: <reason>` only when progress stops. A `multi_agent_v1.wait_agent`
79
+ timeout only means no new mailbox update arrived. Treat a running child as
80
+ alive. Fallback only when the child is completed without the deliverable,
81
+ ack-only after followup, explicitly `BLOCKED:`, or no longer running.
82
+
70
83
  Roles: `explorer` (internal patterns/conventions/tests), `librarian` (external docs/contracts), `metis` (gap analysis), `momus` (high-accuracy plan review). Full spawn/wait/fallback discipline is in `references/full-workflow.md`.
71
84
 
72
85
  ## Stop rules