oh-my-opencode 4.19.4 → 5.0.0-beta.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/command/publish.md +44 -16
- package/.agents/skills/publish/SKILL.md +44 -16
- package/.agents/skills/work-with-pr/SKILL.md +37 -23
- package/.opencode/command/publish.md +44 -16
- package/.opencode/skills/work-with-pr/SKILL.md +37 -23
- package/README.ja.md +1 -1
- package/README.ko.md +1 -1
- package/README.md +18 -6
- package/README.ru.md +1 -1
- package/README.zh-cn.md +1 -1
- package/bin/oh-my-opencode.js +14 -1
- package/bin/oh-my-opencode.test.ts +21 -0
- package/dist/agents/atlas/agent.d.ts +0 -1
- package/dist/agents/sisyphus/grok-4.d.ts +20 -0
- package/dist/agents/sisyphus/index.d.ts +2 -0
- package/dist/agents/sisyphus-agent-config.d.ts +6 -0
- package/dist/agents/sisyphus-agent-factory.d.ts +1 -1
- package/dist/agents/sisyphus-runtime-prompt-reconciler.d.ts +15 -4
- package/dist/agents/types.d.ts +2 -2
- package/dist/cli/index.js +1888 -807
- package/dist/cli/run/on-complete-hook.d.ts +2 -0
- package/dist/cli-node/index.js +1888 -807
- package/dist/config/schema/agent-overrides.d.ts +528 -0
- package/dist/config/schema/oh-my-opencode-config.d.ts +495 -0
- package/dist/features/monitor/batcher.d.ts +3 -1
- package/dist/features/monitor/manager-internals.d.ts +1 -0
- package/dist/features/monitor/output-injector-types.d.ts +2 -0
- package/dist/features/monitor/output-injector.d.ts +6 -0
- package/dist/hooks/atlas/final-wave-approval-gate.test-support.d.ts +50 -0
- package/dist/hooks/atlas/system-reminder-templates.d.ts +0 -1
- package/dist/hooks/todo-continuation-enforcer/types.d.ts +1 -0
- package/dist/hooks/todo-continuation-enforcer/unrecoverable-request-error.d.ts +9 -0
- package/dist/hooks/tool-pair-validator/hook.test-support.d.ts +29 -0
- package/dist/hooks/tool-pair-validator/tool-part-ids.d.ts +14 -5
- package/dist/hooks/tool-pair-validator/tool-result-repair.d.ts +4 -3
- package/dist/hooks/tool-pair-validator/types.d.ts +5 -22
- package/dist/index.js +3139 -2232
- package/dist/mcp/lsp.d.ts +1 -0
- package/dist/oh-my-opencode.schema.json +1466 -131
- package/dist/shared/normalize-sdk-response.d.ts +1 -0
- package/dist/shared/shell-env.d.ts +1 -1
- package/dist/shared/tmux/constants.d.ts +1 -1
- package/dist/skills/ast-grep/SOURCE +1 -1
- package/dist/skills/ast-grep/install.ps1 +2 -2
- package/dist/skills/ast-grep/install.sh +1 -1
- package/dist/skills/ast-grep/references/install.md +2 -2
- package/dist/skills/ast-grep/tests/smoke.sh +1 -1
- package/dist/skills/coding-agent-sessions/SKILL.md +3 -2
- package/dist/skills/coding-agent-sessions/references/all-platforms.md +1 -1
- package/dist/skills/coding-agent-sessions/references/senpi.md +4 -4
- package/dist/skills/coding-agent-sessions/scripts/agent_sessions/pi_family.py +1 -1
- package/dist/skills/frontend/SKILL.md +10 -7
- package/dist/skills/frontend/references/design/_INDEX.md +1 -0
- package/dist/skills/frontend/references/design/stylegallery.md +80 -0
- package/dist/skills/start-work/SKILL.md +54 -9
- package/dist/skills/ultimate-browsing/ATTRIBUTION.md +37 -10
- package/dist/skills/ultimate-browsing/SKILL.md +2 -2
- package/dist/skills/ultimate-browsing/engine/AGENTS.md +179 -0
- package/dist/skills/ultimate-browsing/engine/__main__.py +8 -1
- package/dist/skills/ultimate-browsing/engine/bias_check.py +11 -0
- package/dist/skills/ultimate-browsing/engine/fetch_chain.py +90 -52
- package/dist/skills/ultimate-browsing/engine/result_schema.py +10 -1
- package/dist/skills/ultimate-browsing/engine/surrogate.py +214 -0
- package/dist/skills/ultimate-browsing/engine/surrogates.yaml +60 -0
- package/dist/skills/ultimate-browsing/engine/templates/package.json +1 -1
- package/dist/skills/ultimate-browsing/engine/tests/fixtures/amp_redirect_stub.html +7 -0
- package/dist/skills/ultimate-browsing/engine/tests/fixtures/search_interstitial.html +19 -0
- package/dist/skills/ultimate-browsing/engine/tests/fixtures/wayback_available.json +1 -0
- package/dist/skills/ultimate-browsing/engine/tests/fixtures/wayback_snapshot.html +1128 -0
- package/dist/skills/ultimate-browsing/engine/tests/test_surrogate.py +252 -0
- package/dist/skills/ultimate-browsing/engine/tests/test_surrogate_validators.py +78 -0
- package/dist/skills/ultimate-browsing/engine/validators.py +46 -0
- package/dist/skills/ultimate-browsing/engine/waf_detector.py +1 -1
- package/dist/skills/ultimate-browsing/engine/waf_profiles.yaml +10 -5
- package/dist/skills/ultimate-browsing/references/agent-reach/social.md +1 -1
- package/dist/skills/ultimate-browsing/references/chrome-stealth.md +13 -11
- package/dist/skills/ultimate-browsing/references/insane-search/README.md +4 -4
- package/dist/skills/ultimate-browsing/references/insane-search/cache-archive.md +51 -50
- package/dist/skills/ultimate-browsing/references/insane-search/fallback.md +1 -1
- package/dist/skills/ultimate-browsing/references/insane-search/jina.md +8 -2
- package/dist/skills/ultimate-browsing/references/insane-search/naver.md +1 -1
- package/dist/skills/ultimate-browsing/references/insane-search/twitter.md +3 -3
- package/dist/skills/ulw-plan/SKILL.md +3 -3
- package/dist/skills/ulw-plan/references/full-workflow.md +30 -6
- package/dist/skills/ulw-plan/references/intent-clear.md +2 -1
- package/dist/skills/ulw-plan/references/intent-unclear.md +3 -3
- package/dist/skills/ulw-plan/scripts/scaffold-plan.mjs +1 -1
- package/dist/skills/ulw-research/SKILL.md +15 -10
- package/dist/tui.js +477 -36
- package/docs/reference/web-terminal-visual-qa.md +1 -1
- package/package.json +34 -24
- package/packages/lsp-core/src/lsp/client-diagnostics-concurrency.integration.test.ts +44 -0
- package/packages/lsp-core/src/lsp/client-diagnostics-freshness.integration.test.ts +0 -28
- package/packages/lsp-core/src/lsp/client-wrapper.test.ts +60 -7
- package/packages/lsp-core/src/lsp/client-wrapper.ts +69 -16
- package/packages/lsp-core/src/lsp/connection.ts +1 -1
- package/packages/lsp-core/src/lsp/workspace-edit-adversarial.test.ts +20 -1
- package/packages/lsp-core/src/tools/diagnostics.ts +3 -3
- package/packages/lsp-core/src/tools/navigation.ts +4 -2
- package/packages/lsp-core/src/tools/rename.ts +4 -2
- package/packages/lsp-core/src/tools/symbols.ts +1 -1
- package/packages/lsp-daemon/dist/cli.js +108 -40
- package/packages/lsp-daemon/dist/client.js +123 -55
- package/packages/lsp-daemon/dist/ensure-daemon.d.ts +1 -0
- package/packages/lsp-daemon/dist/ensure-daemon.js +18 -5
- package/packages/lsp-daemon/dist/index.js +111 -43
- package/packages/lsp-tools-mcp/dist/cli.js +77 -23
- package/packages/lsp-tools-mcp/dist/lsp/manager.js +1 -1
- package/packages/lsp-tools-mcp/dist/mcp.js +77 -23
- package/packages/lsp-tools-mcp/dist/tools.js +77 -23
- package/packages/omo-codex/plugin/.codex-plugin/plugin.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/dist/cli.js +268 -92
- package/packages/omo-codex/plugin/components/bootstrap/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/package.json +1 -1
- package/packages/omo-codex/plugin/components/bootstrap/src/setup.ts +7 -7
- package/packages/omo-codex/plugin/components/bootstrap/src/worker.ts +3 -0
- package/packages/omo-codex/plugin/components/codegraph/dist/cli.js +220 -10
- package/packages/omo-codex/plugin/components/codegraph/dist/serve.js +220 -10
- package/packages/omo-codex/plugin/components/codegraph/package.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/comment-checker/package.json +1 -1
- package/packages/omo-codex/plugin/components/git-bash/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/git-bash/package.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/package.json +1 -1
- package/packages/omo-codex/plugin/components/lazycodex-executor-verify/test/codex-hook.test.ts +3 -17
- package/packages/omo-codex/plugin/components/lsp/dist/.omo-runtime-manifest.json +3 -3
- package/packages/omo-codex/plugin/components/lsp/dist/cli.js +135 -67
- package/packages/omo-codex/plugin/components/lsp/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/lsp/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/rules/package.json +1 -1
- package/packages/omo-codex/plugin/components/rules/test/bundled-rules-priority.test.ts +11 -16
- package/packages/omo-codex/plugin/components/rules/test/bundled-rules.test.ts +16 -23
- package/packages/omo-codex/plugin/components/rules/test/codex-hook-post-compact-budget.test.ts +9 -7
- package/packages/omo-codex/plugin/components/rules/test/codex-hook-post-compact-context.test.ts +0 -6
- package/packages/omo-codex/plugin/components/rules/test/codex-hook-post-compact-dedup.test.ts +6 -4
- package/packages/omo-codex/plugin/components/rules/test/codex-hook-post-compact-directive.test.ts +12 -9
- package/packages/omo-codex/plugin/components/rules/test/codex-hook.test.ts +28 -37
- package/packages/omo-codex/plugin/components/rules/test/formatter.test.ts +37 -69
- package/packages/omo-codex/plugin/components/rules/test/hook-output.test.ts +2 -3
- package/packages/omo-codex/plugin/components/rules/test/windows-git-bash-bundled-rule.test.ts +1 -15
- package/packages/omo-codex/plugin/components/start-work-continuation/AGENTS.md +4 -2
- package/packages/omo-codex/plugin/components/start-work-continuation/README.md +5 -1
- package/packages/omo-codex/plugin/components/start-work-continuation/directive.md +2 -1
- package/packages/omo-codex/plugin/components/start-work-continuation/dist/cli.js +18 -0
- package/packages/omo-codex/plugin/components/start-work-continuation/hooks/hooks.json +2 -2
- package/packages/omo-codex/plugin/components/start-work-continuation/package.json +1 -1
- package/packages/omo-codex/plugin/components/start-work-continuation/src/codex-hook.ts +21 -0
- package/packages/omo-codex/plugin/components/start-work-continuation/test/cli.test.ts +0 -3
- package/packages/omo-codex/plugin/components/start-work-continuation/test/codex-hook.test.ts +107 -16
- package/packages/omo-codex/plugin/components/teammode/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/package.json +1 -1
- package/packages/omo-codex/plugin/components/teammode/test/thread-title-hook.test.ts +3 -9
- package/packages/omo-codex/plugin/components/telemetry/dist/cli.js +24 -12
- package/packages/omo-codex/plugin/components/telemetry/dist/posthog.js +24 -12
- package/packages/omo-codex/plugin/components/telemetry/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/telemetry/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/CHANGELOG.md +2 -0
- package/packages/omo-codex/plugin/components/ultrawork/agents/lazycodex-code-reviewer.toml +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/agents/lazycodex-gate-reviewer.toml +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/agents/lazycodex-qa-executor.toml +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/agents/lazycodex-worker-high.toml +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/agents/lazycodex-worker-low.toml +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/agents/lazycodex-worker-medium.toml +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/agents/plan.toml +2 -2
- package/packages/omo-codex/plugin/components/ultrawork/directive.md +9 -2
- package/packages/omo-codex/plugin/components/ultrawork/hooks/hooks.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/package.json +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/skills/ultrawork/SKILL.md +9 -2
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/SKILL.md +3 -3
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/full-workflow.md +30 -6
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/intent-clear.md +2 -1
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/references/intent-unclear.md +3 -3
- package/packages/omo-codex/plugin/components/ultrawork/skills/ulw-plan/scripts/scaffold-plan.mjs +1 -1
- package/packages/omo-codex/plugin/components/ultrawork/test/codex-hook.test.ts +0 -136
- package/packages/omo-codex/plugin/components/ultrawork/test/skill-pointer.test.ts +0 -2
- package/packages/omo-codex/plugin/components/ulw-loop/AGENTS.md +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/CHANGELOG.md +2 -0
- package/packages/omo-codex/plugin/components/ulw-loop/README.md +11 -11
- package/packages/omo-codex/plugin/components/ulw-loop/directive.md +9 -2
- package/packages/omo-codex/plugin/components/ulw-loop/dist/checkpoint-reconciliation.js +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/cli-output.d.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/cli-output.js +9 -9
- package/packages/omo-codex/plugin/components/ulw-loop/dist/cli-steering.js +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/cli.js +66 -66
- package/packages/omo-codex/plugin/components/ulw-loop/dist/codex-goal-instruction.js +4 -4
- package/packages/omo-codex/plugin/components/ulw-loop/dist/codex-hook.js +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/plan-crud.js +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/plan-io.js +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/steering.js +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/dist/stop-resume-hook.js +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/hooks/hooks.json +4 -4
- package/packages/omo-codex/plugin/components/ulw-loop/package.json +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/SKILL.md +5 -4
- package/packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/references/define-goal.md +108 -0
- package/packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/references/full-workflow.md +23 -25
- package/packages/omo-codex/plugin/components/ulw-loop/src/checkpoint-reconciliation.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/cli-output.ts +9 -9
- package/packages/omo-codex/plugin/components/ulw-loop/src/cli-steering.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/cli.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/codex-goal-instruction.ts +4 -4
- package/packages/omo-codex/plugin/components/ulw-loop/src/codex-hook.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/plan-crud.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/plan-io.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/steering.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/src/stop-resume-hook.ts +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/test/checkpoint-continuation.test.ts +0 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/cli-commands.test.ts +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/test/cli-entrypoint.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/cli-helpers.test.ts +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/test/cli-steering-kind-guidance.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/codex-goal-instruction.test.ts +2 -2
- package/packages/omo-codex/plugin/components/ulw-loop/test/codex-hook.test.ts +2 -5
- package/packages/omo-codex/plugin/components/ulw-loop/test/fixtures/quality-gate-builder.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/package-smoke.test.ts +7 -40
- package/packages/omo-codex/plugin/components/ulw-loop/test/plan-io.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/quality-gate-roles.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/steering.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/stop-resume-hook.test.ts +1 -1
- package/packages/omo-codex/plugin/components/ulw-loop/test/ultrawork-directive.test.ts +4 -5
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-git-bash-mcp-reminder.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-lsp-diagnostics-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-compact-resetting-project-rule-cache.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-codegraph-init-guidance.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-comments.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-lsp-diagnostics.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-checking-thread-title-hygiene.json +1 -1
- package/packages/omo-codex/plugin/hooks/post-tool-use-matching-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-enforcing-unlimited-goal-budget.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-guarding-ulw-loop-spawns.json +1 -1
- package/packages/omo-codex/plugin/hooks/pre-tool-use-recommending-git-bash-mcp.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-auto-update.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-bootstrap-provisioning.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-checking-codegraph-bootstrap.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/hooks/session-start-recording-session-telemetry.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-start-work-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/stop-checking-ulw-loop-resume.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-checking-start-work-continuation.json +1 -1
- package/packages/omo-codex/plugin/hooks/subagent-stop-verifying-lazycodex-executor-evidence.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ultrawork-trigger.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-checking-ulw-loop-steering.json +1 -1
- package/packages/omo-codex/plugin/hooks/user-prompt-submit-loading-project-rules.json +1 -1
- package/packages/omo-codex/plugin/package-lock.json +20 -20
- package/packages/omo-codex/plugin/package.json +1 -1
- package/packages/omo-codex/plugin/scripts/sync-skills.mjs +4 -4
- package/packages/omo-codex/plugin/skills/ast-grep/SOURCE +1 -1
- package/packages/omo-codex/plugin/skills/ast-grep/install.ps1 +2 -2
- package/packages/omo-codex/plugin/skills/ast-grep/install.sh +1 -1
- package/packages/omo-codex/plugin/skills/ast-grep/references/install.md +2 -2
- package/packages/omo-codex/plugin/skills/ast-grep/tests/smoke.sh +1 -1
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/SKILL.md +3 -2
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/references/all-platforms.md +1 -1
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/references/senpi.md +4 -4
- package/packages/omo-codex/plugin/skills/coding-agent-sessions/scripts/agent_sessions/pi_family.py +1 -1
- package/packages/omo-codex/plugin/skills/frontend/SKILL.md +10 -7
- package/packages/omo-codex/plugin/skills/frontend/references/design/_INDEX.md +1 -0
- package/packages/omo-codex/plugin/skills/frontend/references/design/stylegallery.md +80 -0
- package/packages/omo-codex/plugin/skills/start-work/SKILL.md +54 -9
- package/packages/omo-codex/plugin/skills/ultimate-browsing/ATTRIBUTION.md +37 -10
- package/packages/omo-codex/plugin/skills/ultimate-browsing/SKILL.md +2 -2
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/AGENTS.md +179 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/__main__.py +8 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/bias_check.py +11 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/fetch_chain.py +90 -52
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/result_schema.py +10 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/surrogate.py +214 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/surrogates.yaml +60 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/templates/package.json +1 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/tests/fixtures/amp_redirect_stub.html +7 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/tests/fixtures/search_interstitial.html +19 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/tests/fixtures/wayback_available.json +1 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/tests/fixtures/wayback_snapshot.html +1128 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/tests/test_surrogate.py +252 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/tests/test_surrogate_validators.py +78 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/validators.py +46 -0
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/waf_detector.py +1 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/engine/waf_profiles.yaml +10 -5
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/agent-reach/social.md +1 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/chrome-stealth.md +13 -11
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/README.md +4 -4
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/cache-archive.md +51 -50
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/fallback.md +1 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/jina.md +8 -2
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/naver.md +1 -1
- package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/twitter.md +3 -3
- package/packages/omo-codex/plugin/skills/ultrawork/SKILL.md +9 -2
- package/packages/omo-codex/plugin/skills/ulw-loop/SKILL.md +5 -4
- package/packages/omo-codex/plugin/skills/ulw-loop/references/define-goal.md +108 -0
- package/packages/omo-codex/plugin/skills/ulw-loop/references/full-workflow.md +23 -25
- package/packages/omo-codex/plugin/skills/ulw-plan/SKILL.md +3 -3
- package/packages/omo-codex/plugin/skills/ulw-plan/references/full-workflow.md +30 -6
- package/packages/omo-codex/plugin/skills/ulw-plan/references/intent-clear.md +2 -1
- package/packages/omo-codex/plugin/skills/ulw-plan/references/intent-unclear.md +3 -3
- package/packages/omo-codex/plugin/skills/ulw-plan/scripts/scaffold-plan.mjs +1 -1
- package/packages/omo-codex/plugin/skills/ulw-research/SKILL.md +15 -10
- package/packages/omo-codex/plugin/test/aggregate-agents.test.mjs +19 -173
- package/packages/omo-codex/plugin/test/aggregate-hooks.test.mjs +4 -24
- package/packages/omo-codex/plugin/test/aggregate-plugin-fixture.mjs +175 -13
- package/packages/omo-codex/plugin/test/aggregate.test.mjs +78 -2
- package/packages/omo-codex/plugin/test/auto-update-release-notes.test.mjs +19 -33
- package/packages/omo-codex/plugin/test/bootstrap-binlinks.test.mjs +12 -12
- package/packages/omo-codex/plugin/test/bootstrap-orchestration.test.mjs +36 -4
- package/packages/omo-codex/plugin/test/lcx-contribute-bug-fix-template.test.mjs +21 -27
- package/packages/omo-codex/plugin/test/scaffold-plan.test.mjs +0 -36
- package/packages/omo-codex/plugin/test/sync-skills-codex-compatibility.test.mjs +101 -0
- package/packages/omo-codex/plugin/test/sync-skills-test-support.mjs +4 -4
- package/packages/omo-codex/plugin/test/sync-skills.test.mjs +1 -119
- package/packages/omo-codex/plugin/test/teammode-archive-ambiguity.test.mjs +0 -40
- package/packages/omo-codex/plugin/test/teammode-communication.test.mjs +6 -62
- package/packages/omo-codex/plugin/test/teammode-thread-links.test.mjs +3 -36
- package/packages/omo-codex/plugin/test/teammode-transport.test.mjs +0 -44
- package/packages/omo-codex/plugin/test/teammode-worktree.test.mjs +2 -6
- package/packages/omo-codex/plugin/test/ultrawork-skill-pointer.test.mjs +0 -3
- package/packages/omo-codex/plugin/test/ulw-plan-review-state-contract.test.mjs +0 -3
- package/packages/omo-codex/scripts/install-bin-links.test.mjs +56 -2
- package/packages/omo-codex/scripts/install-delegated-command.test.mjs +6 -6
- package/packages/omo-codex/scripts/install-dist/install-local.mjs +170 -78
- package/packages/omo-codex/scripts/install-lazycodex-version-stamp.test.mjs +7 -2
- package/packages/omo-codex/scripts/install-local-entrypoint.test.mjs +4 -4
- package/packages/omo-codex/scripts/install-local.test.mjs +5 -2
- package/packages/shared-skills/index.mjs +19 -1
- package/packages/shared-skills/skills/ast-grep/SOURCE +1 -1
- package/packages/shared-skills/skills/ast-grep/install.ps1 +2 -2
- package/packages/shared-skills/skills/ast-grep/install.sh +1 -1
- package/packages/shared-skills/skills/ast-grep/references/install.md +2 -2
- package/packages/shared-skills/skills/ast-grep/tests/smoke.sh +1 -1
- package/packages/shared-skills/skills/coding-agent-sessions/SKILL.md +3 -2
- package/packages/shared-skills/skills/coding-agent-sessions/references/all-platforms.md +1 -1
- package/packages/shared-skills/skills/coding-agent-sessions/references/senpi.md +4 -4
- package/packages/shared-skills/skills/coding-agent-sessions/scripts/agent_sessions/pi_family.py +1 -1
- package/packages/shared-skills/skills/frontend/SKILL.md +10 -7
- package/packages/shared-skills/skills/frontend/references/design/_INDEX.md +1 -0
- package/packages/shared-skills/skills/frontend/references/design/stylegallery.md +80 -0
- package/packages/shared-skills/skills/start-work/SKILL.md +54 -9
- package/packages/shared-skills/skills/ultimate-browsing/ATTRIBUTION.md +37 -10
- package/packages/shared-skills/skills/ultimate-browsing/SKILL.md +2 -2
- package/packages/shared-skills/skills/ultimate-browsing/engine/AGENTS.md +179 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/__main__.py +8 -1
- package/packages/shared-skills/skills/ultimate-browsing/engine/bias_check.py +11 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/fetch_chain.py +90 -52
- package/packages/shared-skills/skills/ultimate-browsing/engine/result_schema.py +10 -1
- package/packages/shared-skills/skills/ultimate-browsing/engine/surrogate.py +214 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/surrogates.yaml +60 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/templates/package.json +1 -1
- package/packages/shared-skills/skills/ultimate-browsing/engine/tests/fixtures/amp_redirect_stub.html +7 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/tests/fixtures/search_interstitial.html +19 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/tests/fixtures/wayback_available.json +1 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/tests/fixtures/wayback_snapshot.html +1128 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/tests/test_surrogate.py +252 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/tests/test_surrogate_validators.py +78 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/validators.py +46 -0
- package/packages/shared-skills/skills/ultimate-browsing/engine/waf_detector.py +1 -1
- package/packages/shared-skills/skills/ultimate-browsing/engine/waf_profiles.yaml +10 -5
- package/packages/shared-skills/skills/ultimate-browsing/references/agent-reach/social.md +1 -1
- package/packages/shared-skills/skills/ultimate-browsing/references/chrome-stealth.md +13 -11
- package/packages/shared-skills/skills/ultimate-browsing/references/insane-search/README.md +4 -4
- package/packages/shared-skills/skills/ultimate-browsing/references/insane-search/cache-archive.md +51 -50
- package/packages/shared-skills/skills/ultimate-browsing/references/insane-search/fallback.md +1 -1
- package/packages/shared-skills/skills/ultimate-browsing/references/insane-search/jina.md +8 -2
- package/packages/shared-skills/skills/ultimate-browsing/references/insane-search/naver.md +1 -1
- package/packages/shared-skills/skills/ultimate-browsing/references/insane-search/twitter.md +3 -3
- package/packages/shared-skills/skills/ulw-plan/SKILL.md +3 -3
- package/packages/shared-skills/skills/ulw-plan/references/full-workflow.md +30 -6
- package/packages/shared-skills/skills/ulw-plan/references/intent-clear.md +2 -1
- package/packages/shared-skills/skills/ulw-plan/references/intent-unclear.md +3 -3
- package/packages/shared-skills/skills/ulw-plan/scripts/scaffold-plan.mjs +1 -1
- package/packages/shared-skills/skills/ulw-research/SKILL.md +15 -10
- package/postinstall.mjs +6 -0
- package/dist/tools/call-omo-agent/background-agent-executor.d.ts +0 -5
- package/packages/omo-codex/plugin/test/aggregate-skills.test.mjs +0 -92
- package/packages/omo-codex/plugin/test/sync-skills-orchestration.test.mjs +0 -314
- package/packages/omo-codex/plugin/test/ulw-plan-scope-contract.test.mjs +0 -24
package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/cache-archive.md
CHANGED
|
@@ -1,46 +1,50 @@
|
|
|
1
|
-
# 캐시 & 아카이브
|
|
1
|
+
# 캐시 & 아카이브 (surrogate 경로)
|
|
2
2
|
|
|
3
|
-
> 원본 사이트가 차단되었을 때
|
|
4
|
-
>
|
|
3
|
+
> 원본 사이트가 차단되었을 때 캐시/아카이브된 **사본**으로 접근.
|
|
4
|
+
> 2026-08-09 실측 probe 기준으로 정렬. 각 경로는 생명 주기가 짧다 — 이 파일도
|
|
5
|
+
> 90일마다 재검증 대상. (당일 probe: 기존 기대 경로 6개 중 4개 사망 또는 스텁 반환.)
|
|
5
6
|
|
|
6
7
|
## 의존성
|
|
7
8
|
|
|
8
|
-
없음 (curl만 사용).
|
|
9
|
+
없음 (curl만 사용). 수동 경로이며, 자동화는 엔진 Phase 2.5(`engine/surrogates.yaml`)가 담당한다.
|
|
9
10
|
|
|
10
|
-
##
|
|
11
|
+
## 엔진 자동 폴백 (Phase 2.5)
|
|
11
12
|
|
|
12
|
-
|
|
13
|
+
`waf_profiles.yaml`의 `fallback_when_challenge`가 `surrogate_wayback`을 앞에 두므로,
|
|
14
|
+
그리드 실패 후 브라우저 실행 전에 아카이브 경로를 먼저 시도한다.
|
|
15
|
+
성공 시 `FetchResult.provenance = "snapshot"`, `snapshot_timestamp` = 아카이브의 자체 타임스탬프,
|
|
16
|
+
`trust = "archive"`가 채워진다. **사본이므로 반드시 날짜와 함께 인용할 것.**
|
|
17
|
+
`--allow-proxy` 없이는 `kind: proxy` 엔트리는 절대 실행되지 않으며, 프록시에는
|
|
18
|
+
Cookie/Authorization 헤더를 보내지 않는다 (중계자 = 구조적 MITM).
|
|
19
|
+
|
|
20
|
+
## 1. Wayback Machine (Internet Archive) — 1순위
|
|
21
|
+
|
|
22
|
+
**2026-08-09 probe: 정상 동작.** `available` API가 200 JSON으로 스냅샷 URL과
|
|
23
|
+
타임스탬프를 돌려준다 — 출처(provenance) 확보에 가장 좋은 primitive.
|
|
13
24
|
|
|
14
25
|
```bash
|
|
15
|
-
#
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
url = '{URL}'
|
|
21
|
-
p = urlparse(url)
|
|
22
|
-
domain_sub = p.netloc.replace('.', '-')
|
|
23
|
-
print(f'https://{domain_sub}.cdn.ampproject.org/c/s/{p.netloc}{p.path}')
|
|
24
|
-
"
|
|
25
|
-
|
|
26
|
-
# 변환된 URL로 접근
|
|
27
|
-
curl -sL "https://{domain-with-dashes}.cdn.ampproject.org/c/s/{netloc}{path}"
|
|
26
|
+
# 스냅샷 존재 여부 + 최신 스냅샷 URL/타임스탬프 (진입점으로 이것을 쓸 것)
|
|
27
|
+
curl -sL "https://archive.org/wayback/available?url={URL}"
|
|
28
|
+
|
|
29
|
+
# 반환 JSON의 archived_snapshots.closest.url 로 접근
|
|
30
|
+
curl -sL "https://web.archive.org/web/{timestamp}/{URL}"
|
|
28
31
|
```
|
|
29
32
|
|
|
30
|
-
|
|
31
|
-
|
|
33
|
+
> **CDX API 주의**: 이전 버전이 권장하던 `web.archive.org/cdx/search/cdx`는
|
|
34
|
+
> 2026-08 probe에서 503 반환. 스냅샷 열거가 필요 없으면 `available` API만 사용.
|
|
32
35
|
|
|
33
|
-
|
|
36
|
+
**성공 조건**: 크롤링 대상이었던 공개 URL
|
|
37
|
+
**실패 조건**: robots.txt로 차단된 사이트, 스냅샷이 없는 URL, SPA 스냅샷 (렌더링 안 됨)
|
|
38
|
+
|
|
39
|
+
## 2. archive.today — 2순위
|
|
34
40
|
|
|
35
41
|
사용자 제출 아카이브. 페이월 기사, 삭제된 콘텐츠에 특히 유용.
|
|
36
|
-
|
|
42
|
+
**2026-08 probe: 429 rate-limit이 잦고 도메인이 수시로 회전** (archive.ph → archive.md 관찰).
|
|
43
|
+
하나가 차단되면 다른 도메인을 순회한다 (엔진 `host_rotation`과 동일 패턴).
|
|
37
44
|
|
|
38
45
|
```bash
|
|
39
|
-
# 최신 스냅샷 조회
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
# 도메인 로테이션 (하나가 차단되면 다른 것)
|
|
43
|
-
for domain in archive.ph archive.is archive.md archive.vn archive.li; do
|
|
46
|
+
# 최신 스냅샷 조회 — 도메인 회전은 필수 경로, 예외 처리 아님
|
|
47
|
+
for domain in archive.ph archive.md archive.li archive.is; do
|
|
44
48
|
resp=$(curl -sL -o /dev/null -w "%{http_code}" "https://$domain/newest/{URL}")
|
|
45
49
|
if [ "$resp" = "200" ] || [ "$resp" = "302" ]; then
|
|
46
50
|
echo "성공: https://$domain/newest/{URL}"
|
|
@@ -50,34 +54,31 @@ for domain in archive.ph archive.is archive.md archive.vn archive.li; do
|
|
|
50
54
|
done
|
|
51
55
|
```
|
|
52
56
|
|
|
53
|
-
|
|
54
|
-
**실패 조건**: 아카이브된 적 없는 URL
|
|
57
|
+
**주의**: 429 응답에도 수 KB 본문이 딸려 오므로 상태코드 대신 본문 검증이 필요하다.
|
|
55
58
|
|
|
56
|
-
## 3.
|
|
59
|
+
## 3. AMP 캐시 — 강등 (사실상 무용)
|
|
57
60
|
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
# 최신 스냅샷으로 접근
|
|
63
|
-
curl -sL "https://web.archive.org/web/{URL}"
|
|
61
|
+
과거 1순위였으나 **2026-08 probe에서 사실상 무력화**:
|
|
62
|
+
`{host}.cdn.ampproject.org/c/s/...`가 HTTP 200을 돌려주지만, 실제 본문은
|
|
63
|
+
**322바이트짜리 `<TITLE>Redirecting</TITLE>` meta-refresh** — 대상은 다시 **원본(차단된) 페이지**다.
|
|
64
|
+
이걸 성공으로 착각하면 에이전트가 차단 페이지로 되돌아가는 루프가 생긴다.
|
|
64
65
|
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
**성공 조건**: 크롤링 대상이었던 공개 URL
|
|
70
|
-
**실패 조건**: robots.txt로 차단된 사이트, SPA (렌더링 안 됨), iframe 기반 사이트
|
|
66
|
+
엔진은 `engine/validators.py:is_redirect_stub`으로 이 패턴을 CHALLENGE 판정한다
|
|
67
|
+
(3KB 미만 + meta-refresh/JS redirect + 대상 호스트 재등장 조합).
|
|
68
|
+
수동 사용도 권장하지 않는다.
|
|
71
69
|
|
|
72
|
-
## 4. Google Cache
|
|
70
|
+
## 4. Google Cache — 사망 확정
|
|
73
71
|
|
|
74
|
-
|
|
75
|
-
>
|
|
72
|
+
**2024년 7월 종료** 후로도 `webcache.googleusercontent.com`이 HTTP 200 + 수십 KB의
|
|
73
|
+
본문을 반환하지만, 실제로는 `<title>Google Search</title>` 인터스티셜 + JS 리다이렉트다
|
|
74
|
+
(2026-08 probe 재확인). **캐시가 아니라 검색 홈이다.**
|
|
75
|
+
엔진은 `INTERSTITIAL_TITLE_MARKERS`로 판정해 성공 집계에서 배제한다.
|
|
76
76
|
|
|
77
|
-
## 시도 순서
|
|
77
|
+
## 시도 순서 (probe 근거)
|
|
78
78
|
|
|
79
79
|
```
|
|
80
|
-
1.
|
|
81
|
-
2. archive.today (
|
|
82
|
-
3.
|
|
80
|
+
1. Wayback available API → 스냅샷 URL + 타임스탬프 (provenance까지 확보)
|
|
81
|
+
2. archive.today 도메인 회전 (429 대비, 본문 검증 필수)
|
|
82
|
+
3. AMP 캐시: 시도하지 않음 (redirect stub → 원본으로 회귀)
|
|
83
|
+
4. Google Cache: 시도하지 않음 (사망, 검색 인터스티셜 반환)
|
|
83
84
|
```
|
package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/fallback.md
CHANGED
|
@@ -141,7 +141,7 @@ browser_evaluate → () => document.body.innerText (Light Mode — 먼저)
|
|
|
141
141
|
|
|
142
142
|
| 패턴 | 감지 방법 | 처리 |
|
|
143
143
|
|------|----------|------|
|
|
144
|
-
| X SPA 셸 (247KB) | 200 OK + `Sign in to X` 또는 `hasResults: false` | 실패 —
|
|
144
|
+
| X SPA 셸 (247KB) | 200 OK + `Sign in to X` 또는 `hasResults: false` | 실패 — 웹 검색 도구+oEmbed 폴백 |
|
|
145
145
|
| CAPTCHA 페이지 | 200 OK + `captcha\|recaptcha\|hcaptcha\|cf-turnstile` | 실패 — 다음 Phase |
|
|
146
146
|
| 소프트 페이월 | 200 OK + `member-only\|subscribe to read\|구독하세요` | 부분 성공 — 메타만 채택 |
|
|
147
147
|
| DDG 소프트 리밋 | 202 Accepted + body 15KB 미만 | 실패 — 다른 엔진 폴백 |
|
|
@@ -2,12 +2,18 @@
|
|
|
2
2
|
|
|
3
3
|
> `r.jina.ai/URL` 한 줄로 거의 모든 공개 URL을 마크다운으로 변환.
|
|
4
4
|
> Puppeteer 기반 실제 브라우저 렌더링 — JS SPA까지 처리.
|
|
5
|
-
>
|
|
5
|
+
>
|
|
6
|
+
> **2026-08-09 probe 기준 무료 무키 경로는 종료됨.** 익명 호출은 401이며,
|
|
7
|
+
> 리다이렉트를 따라가면 Cloudflare Turnstile(`Just a moment...`)에 막힌다.
|
|
8
|
+
> **이제 `JINA_API_KEY` 환경 변수가 필요**하다 — `Authorization: Bearer <key>` 헤더.
|
|
9
|
+
> 엔진에서는 `engine/surrogates.yaml`의 `jina_reader` 엔트리가 키가 있을 때만 활성화된다
|
|
10
|
+
> (kind=reader, provenance=live — 서버 측 재렌더링).
|
|
11
|
+
> 예전 "무료 500 RPM" 안내는 모두 폐기되었으므로 따르지 않는다.
|
|
6
12
|
|
|
7
13
|
## 기본 사용
|
|
8
14
|
|
|
9
15
|
```bash
|
|
10
|
-
curl -s "https://r.jina.ai/{URL}"
|
|
16
|
+
curl -s -H "Authorization: Bearer ${JINA_API_KEY}" "https://r.jina.ai/{URL}"
|
|
11
17
|
```
|
|
12
18
|
|
|
13
19
|
## 고급 기능
|
package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/naver.md
CHANGED
|
@@ -90,7 +90,7 @@ r = s.get(f"https://search.naver.com/search.naver?where=news&query={quote('검
|
|
|
90
90
|
|
|
91
91
|
### 한국어 키워드 검색의 핵심 경로
|
|
92
92
|
|
|
93
|
-
|
|
93
|
+
웹 검색 도구는 한국어 신규 콘텐츠 인덱싱이 지연되지만, 네이버 검색은 한국어에 최적화되어 있다.
|
|
94
94
|
**한국 사이트 키워드 검색 → 네이버 검색 직접 접근이 가장 정확하고 빠르다.**
|
|
95
95
|
|
|
96
96
|
## 네이버 카페
|
package/packages/omo-codex/plugin/skills/ultimate-browsing/references/insane-search/twitter.md
CHANGED
|
@@ -5,10 +5,10 @@
|
|
|
5
5
|
## 검색 (트윗 발견)
|
|
6
6
|
|
|
7
7
|
```python
|
|
8
|
-
|
|
8
|
+
<사용 가능한 web search tool>(query="site:x.com {검색어}") # Claude Code: WebSearch / OpenCode 계열: websearch_web_search_exa 등 — 하네스마다 실제 tool 이름이 다르므로 현재 세션의 tool 목록에서 확인할 것
|
|
9
9
|
```
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
웹 검색 도구는 X 포스트를 검색 결과로 반환한다. 제목, snippet, URL을 획득할 수 있지만 트윗 전문이나 engagement 수치는 없다.
|
|
12
12
|
|
|
13
13
|
## 타임라인 조회 — Syndication API
|
|
14
14
|
|
|
@@ -89,7 +89,7 @@ curl -sL "https://publish.twitter.com/oembed?url=https://x.com/{user}/status/{tw
|
|
|
89
89
|
## 조합 패턴 (검색 → 상세)
|
|
90
90
|
|
|
91
91
|
```
|
|
92
|
-
1단계:
|
|
92
|
+
1단계: 웹 검색 도구(query="site:x.com {키워드}") → 트윗 URL 획득
|
|
93
93
|
2단계: curl oEmbed API → 트윗 전문 획득
|
|
94
94
|
```
|
|
95
95
|
|
|
@@ -133,6 +133,12 @@ exactly `objective`; do not include `status`. Only when no goal tool
|
|
|
133
133
|
exists on this surface, open your reply with a `# Goal` block treated
|
|
134
134
|
as binding. Goals are unlimited; never invent a numeric budget or
|
|
135
135
|
limit.
|
|
136
|
+
Check `get_goal` first: continue a matching active goal instead of
|
|
137
|
+
duplicating one; surface a conflicting one. Write the objective
|
|
138
|
+
outcome-first: the concrete thing that will be TRUE when done (an
|
|
139
|
+
outcome, never an activity), the named deliverable surfaces, and
|
|
140
|
+
explicit scope bounds — a vague objective produces vague criteria,
|
|
141
|
+
and vague criteria cannot be proven.
|
|
136
142
|
The criteria MUST list, upfront:
|
|
137
143
|
- The user-visible deliverable in one line, and the tier with its
|
|
138
144
|
justification.
|
|
@@ -242,8 +248,9 @@ library/API/docs/web — delegate to the `librarian` subagent. Spawn them
|
|
|
242
248
|
# Execution loop (PIN → RED → GREEN → SURFACE → CLEAN)
|
|
243
249
|
Until every success criterion PASSES with its evidence captured:
|
|
244
250
|
1. Pick next criterion → mark in_progress → update notepad `## Now`.
|
|
245
|
-
2. PIN + RED: when
|
|
246
|
-
characterization test that passes on
|
|
251
|
+
2. PIN + RED: when refactoring behavior whose regressions the change
|
|
252
|
+
could hide, first pin it with a characterization test that passes on
|
|
253
|
+
the unchanged code. Then
|
|
247
254
|
capture the failing-first proof through the cheapest faithful
|
|
248
255
|
channel — a unit test where a seam exists, an integration/e2e test
|
|
249
256
|
where the behavior lives in wiring, or the criterion's real-surface
|
|
@@ -15,14 +15,15 @@ This skill is intentionally compact. The full workflow lives in `references/full
|
|
|
15
15
|
|
|
16
16
|
1. Open `references/full-workflow.md`.
|
|
17
17
|
2. Read through **Bootstrap** (including its tier triage), **Execution Loop**, the **Manual-QA channels** table, and the **Stop Rules** before running any ULW command or recording evidence.
|
|
18
|
-
3.
|
|
18
|
+
3. Open `references/define-goal.md` and register the run's goal by it. Goal creation is NEVER skipped: shape the objective and every success criterion by that reference before any implementation.
|
|
19
|
+
4. If the task has code edits, tests, QA, or commit work, follow the full workflow's delegation and evidence rules. Tests alone never prove done.
|
|
19
20
|
|
|
20
21
|
## Non-Negotiables
|
|
21
22
|
|
|
22
23
|
- Use the ulw-loop CLI state under `.omo/ulw-loop`; do not hand-edit goal state.
|
|
23
|
-
- Register goals up front (`omo ulw-loop create-goals`, then `create_goal` from the printed handoff) and mirror every atomic step into the live `update_plan` checklist: one ultra-granular step per action, exactly one in_progress, transitions marked the instant they happen.
|
|
24
|
-
- After any compaction or context loss, re-read brief + goals + ledger FIRST plus `omo ulw-loop status --json`, then resume; never re-plan from scratch.
|
|
25
|
-
- If `omo ulw-loop create-goals` says the existing aggregate is already complete, start unrelated new work with a fresh `--session-id <new-id>` instead of steering or forcing the completed default state. Use `--force` only to intentionally overwrite completed evidence.
|
|
24
|
+
- Register goals up front, shaped by `references/define-goal.md` (`omo-agent-toolkit ulw-loop create-goals`, then `create_goal` from the printed handoff), and mirror every atomic step into the live `update_plan` checklist: one ultra-granular step per action, exactly one in_progress, transitions marked the instant they happen.
|
|
25
|
+
- After any compaction or context loss, re-read brief + goals + ledger FIRST plus `omo-agent-toolkit ulw-loop status --json`, then resume; never re-plan from scratch.
|
|
26
|
+
- If `omo-agent-toolkit ulw-loop create-goals` says the existing aggregate is already complete, start unrelated new work with a fresh `--session-id <new-id>` instead of steering or forcing the completed default state. Use `--force` only to intentionally overwrite completed evidence.
|
|
26
27
|
- Every success criterion needs observable evidence from a real surface: a channel (terminal/TUI via the xterm.js web terminal, HTTP, browser, computer-use) or, for CLI- or data-shaped criteria, an auxiliary surface (CLI stdout, DB diff, parsed config dump).
|
|
27
28
|
- Evidence is bound to the tree it was captured at (`git rev-parse --short "HEAD^{tree}"`); it goes stale only when tracked content changes — a rebase or amend that keeps the tree identical keeps it valid. When the tree differs, re-run at the current HEAD and re-record, never relabel or regenerate. Record only after cleanup receipts exist.
|
|
28
29
|
- Delegate code edits, test writes, fixes, and QA execution to right-sized Codex subagents when the workflow requires it.
|
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# Define Goal
|
|
2
|
+
|
|
3
|
+
How to turn a brief into a registered goal the run can be held to. Read this BEFORE calling `create_goal`: the objective you register is the binding contract for the whole run, and the run's quality is capped by the quality of this objective.
|
|
4
|
+
|
|
5
|
+
A goal is a prompt to the agent that executes it, including future-you after compaction. It earns its tokens the way any prompt does: it carries only what the run cannot re-derive later, the outcome, the proof, the bounds, and the stop state. Everything else is noise that steals attention from the parts that decide completion.
|
|
6
|
+
|
|
7
|
+
## The quality bar
|
|
8
|
+
|
|
9
|
+
Before registering, the objective must answer all five:
|
|
10
|
+
|
|
11
|
+
1. What concrete thing will be TRUE when this is done? An outcome, never an activity.
|
|
12
|
+
2. What evidence will prove it? Commands, validators, artifacts someone can open.
|
|
13
|
+
3. What quantitative or binary threshold defines success?
|
|
14
|
+
4. What scope boundaries matter? What is in, and what is explicitly out.
|
|
15
|
+
5. What should make the agent stop and ask instead of grinding?
|
|
16
|
+
|
|
17
|
+
An objective that cannot answer one of these is not ready. Repair it (below) before calling the tool.
|
|
18
|
+
|
|
19
|
+
## Objective anatomy
|
|
20
|
+
|
|
21
|
+
Write the objective outcome-first, in this order:
|
|
22
|
+
|
|
23
|
+
1. **Outcome**: one sentence stating what will be true, naming the artifact, system, repo, or user-facing behavior involved.
|
|
24
|
+
2. **Deliverables**: the named surfaces the work lands on (files, endpoints, packages, environments). Use literal paths and names: the executing agent interprets the objective literally and will not infer surfaces you did not name.
|
|
25
|
+
3. **Success criteria**: sized by tier (below), each one a binary observable with its scenario and evidence named upfront.
|
|
26
|
+
4. **Constraints and scope bounds**: Record the user's stated constraints verbatim, including what is explicitly out of scope wherever ambiguity would let the run expand. Where the user was silent on a bound the work forks on, SET it yourself: derive the clearest defensible bound from repo evidence and best practice (stack already in use, compatibility surfaces, scale the code must serve, audience or compliance the repo implies) and record it inside the objective as `assumed: <constraint> — <rationale>, <reversible?>`, binding until the user vetoes it. Unstated bounds do not exist — which is why you write them.
|
|
27
|
+
5. **WHEN TO STOP**: one line, "I'll stop right away when <the exact observable state that ends this run>". This line is binding: the moment it holds, the run delivers and stops. Work past it is a defect, not diligence.
|
|
28
|
+
|
|
29
|
+
State the motivation when it changes execution ("p95 matters because the checkout SLA is 300ms") and omit it when it does not. Positive statements beat prohibitions: "verify against staging" carries more signal than "do not touch production".
|
|
30
|
+
|
|
31
|
+
## Success criteria construction
|
|
32
|
+
|
|
33
|
+
Count by tier, mirroring the run's tier triage:
|
|
34
|
+
|
|
35
|
+
- LIGHT (known pattern, no open design decisions): 1-2 criteria, happy path plus the riskiest edge.
|
|
36
|
+
- HEAVY (new module or abstraction, auth or security, external integration, schema or migration, concurrency, cross-domain refactor, or the user demanded care): 3+ criteria covering happy path, edge (boundary, empty, malformed, concurrent), adjacent-surface regression named by file and function, and the adversarial risk the change actually creates.
|
|
37
|
+
|
|
38
|
+
Every criterion carries, at definition time, not after the work:
|
|
39
|
+
|
|
40
|
+
- a binary pass condition ("returns 200 and the body matches the schema", never "works correctly");
|
|
41
|
+
- the exact scenario: the literal command, request, page action, or payload that will prove it;
|
|
42
|
+
- the evidence artifact it will capture: transcript, status plus body, screenshot path, diff, parsed dump;
|
|
43
|
+
- the failing-first proof (test id or scenario) that will be captured RED before implementation.
|
|
44
|
+
|
|
45
|
+
A criterion that cannot fail is not a criterion. If no input could make the scenario fail, it measures nothing; rewrite it until failure is possible.
|
|
46
|
+
|
|
47
|
+
## Make it quantitative
|
|
48
|
+
|
|
49
|
+
Prefer numbers that represent real success over decorative precision. A threshold nobody would act on differently is noise.
|
|
50
|
+
|
|
51
|
+
| Domain | Quantify as |
|
|
52
|
+
| --- | --- |
|
|
53
|
+
| Bug fix | reproduction first, fix second: the failing case captured RED, then the same validator green |
|
|
54
|
+
| Tests | the exact command and required pass condition, plus run count for flake-sensitive suites |
|
|
55
|
+
| Performance | metric, target threshold, measurement method, and run count ("p95 under 250ms across 3 consecutive local runs") |
|
|
56
|
+
| Quality work | the observable acceptance bar: lint, typecheck, and test pass; reviewed examples; a user-approved artifact |
|
|
57
|
+
| Research | the decision the research must enable, the sources or systems in scope, and the evidence standard per claim |
|
|
58
|
+
| Operations | healthy state, monitoring window, failure threshold, and the rollback or escalation trigger |
|
|
59
|
+
|
|
60
|
+
## Repair weak goals
|
|
61
|
+
|
|
62
|
+
Reject pure activity objectives: "make progress", "keep investigating", "improve things", "work on X". They cannot fail, so they cannot finish.
|
|
63
|
+
|
|
64
|
+
Rewrite vague goals into measurable ones when local context makes the rewrite safe. Ask ONE narrow question only when the missing detail is an OWNER-DECISION — irreversible, destructive, safety-critical, or a cross-cutting product choice (real budget or spend, public surface, external dependency, data shape, target audience) — that changes the intended outcome or its validation, shaped around the missing validator or bound:
|
|
65
|
+
|
|
66
|
+
- "What metric defines success here: latency, cost, accuracy, or user-visible behavior?"
|
|
67
|
+
- "Which environment do I verify against: local, staging, or production?"
|
|
68
|
+
- "What is the minimum evidence you want before this goal is marked complete?"
|
|
69
|
+
|
|
70
|
+
Every other missing constraint follows Objective anatomy #4: adopt the clearest defensible default, state it in the objective as `assumed:`, and let the user veto.
|
|
71
|
+
|
|
72
|
+
When the user cannot provide a metric, propose the most honest binary validator available and proceed with it stated in the objective.
|
|
73
|
+
|
|
74
|
+
Weak: "Make checkout faster."
|
|
75
|
+
Repaired: "Reduce checkout API p95 below 250ms on the documented slow path with the smallest safe server-side change; prove it with `npm run test:checkout` green plus the local latency benchmark showing p95 under 250ms across 3 consecutive runs; out of scope: client-side changes and new caching layers."
|
|
76
|
+
|
|
77
|
+
Weak: "Keep investigating the PR comments."
|
|
78
|
+
Repaired: "Resolve every open change-requesting review comment on PR 123 touching only the affected auth files and their tests; prove it with the targeted auth test command green plus `gh pr view 123` showing zero unresolved change-request threads."
|
|
79
|
+
|
|
80
|
+
## Registration protocol
|
|
81
|
+
|
|
82
|
+
1. Call `get_goal` first, then act by state:
|
|
83
|
+
|
|
84
|
+
| get_goal shows | Action |
|
|
85
|
+
| --- | --- |
|
|
86
|
+
| no active goal | Register with `create_goal`, passing exactly `objective`. Never include lifecycle fields such as `status`; never register a goal in prose, a notepad, or a plan instead of the tool. |
|
|
87
|
+
| an active goal matching this intent | Continue it. Never register a duplicate. |
|
|
88
|
+
| an active goal conflicting with this intent | Stop and surface the conflict; the user decides whether to finish it, complete it, or branch. |
|
|
89
|
+
|
|
90
|
+
2. Goals are unlimited. Never invent a numeric budget, token limit, or deadline the user did not state — that ban covers run quotas; the `assumed:` work constraints from Objective anatomy #4 are different and required.
|
|
91
|
+
3. In a ulw-loop run, the loop CLI owns per-goal state (`.omo/ulw-loop/goals.json`): `create_goal` registers the aggregate objective from the printed handoff, and this reference shapes both that objective and every goal's `successCriteria` at `create-goals` time.
|
|
92
|
+
|
|
93
|
+
## Completion honesty
|
|
94
|
+
|
|
95
|
+
- Report `update_goal` complete only after auditing every criterion against evidence captured in this run. A green suite is supporting evidence, never completion proof by itself.
|
|
96
|
+
- Waiting is not blocked: while a monitor, background child, or scheduled continuation can wake the run, end the turn and let it fire. Blocked requires a true impasse: no live resumption channel, and the same block recurring across consecutive turns.
|
|
97
|
+
- The moment the WHEN TO STOP line holds with evidence in hand, deliver and stop.
|
|
98
|
+
|
|
99
|
+
## Anti-patterns
|
|
100
|
+
|
|
101
|
+
| Anti-pattern | Why it fails | Instead |
|
|
102
|
+
| --- | --- | --- |
|
|
103
|
+
| Activity objective ("investigate X") | Cannot fail, so cannot finish; the run wanders | Name the outcome the activity must produce and its evidence |
|
|
104
|
+
| Criteria added after implementation | The contract bent to fit the work; nothing was proven | Write criteria and scenarios at registration, before any edit |
|
|
105
|
+
| Decorative precision ("99.97% uptime" nobody measures) | A threshold no validator checks is noise wearing a suit | Only thresholds a named validator will actually check |
|
|
106
|
+
| Padded objective (role prose, restated context, filler) | Every extra token competes with the criteria for attention | Outcome, deliverables, criteria, bounds, stop line; nothing else |
|
|
107
|
+
| Goal registered in prose or a notepad | Nothing binds the run; completion becomes a vibe | `create_goal` with the objective, every time the tool exists |
|
|
108
|
+
| Duplicate goal for the same intent | Two contracts, neither authoritative | Continue the active goal or surface the conflict |
|
|
@@ -66,8 +66,8 @@ Codex subagent reliability:
|
|
|
66
66
|
- `.omo/ulw-loop/goals.json`: goals with embedded `successCriteria` per goal.
|
|
67
67
|
- `.omo/ulw-loop/ledger.jsonl`: append-only audit trail.
|
|
68
68
|
- Read artifacts before resuming, steering, or checkpointing.
|
|
69
|
-
- After compaction or context loss, re-read brief + goals + ledger FIRST, then `omo ulw-loop status --json`. Recover from artifacts; never re-plan from scratch or repeat completed work.
|
|
70
|
-
- Never invent state outside `.omo/ulw-loop` artifacts or `omo ulw-loop status --json`.
|
|
69
|
+
- After compaction or context loss, re-read brief + goals + ledger FIRST, then `omo-agent-toolkit ulw-loop status --json`. Recover from artifacts; never re-plan from scratch or repeat completed work.
|
|
70
|
+
- Never invent state outside `.omo/ulw-loop` artifacts or `omo-agent-toolkit ulw-loop status --json`.
|
|
71
71
|
|
|
72
72
|
## Bootstrap
|
|
73
73
|
Do all three steps before execution. No edits, goal tools, or checkpointing before bootstrap completes.
|
|
@@ -86,10 +86,10 @@ if [ -z "$ULW_LOOP_NODE" ]; then
|
|
|
86
86
|
fi
|
|
87
87
|
|
|
88
88
|
ULW_LOOP_CLI=
|
|
89
|
-
if command -v omo >/dev/null 2>&1 && omo ulw-loop help >/dev/null 2>&1; then
|
|
90
|
-
ULW_LOOP_CLI=omo
|
|
89
|
+
if command -v omo-agent-toolkit >/dev/null 2>&1 && omo-agent-toolkit ulw-loop help >/dev/null 2>&1; then
|
|
90
|
+
ULW_LOOP_CLI=omo-agent-toolkit
|
|
91
91
|
elif [ -n "$ULW_LOOP_NODE" ]; then
|
|
92
|
-
for candidate in "$HOME/.local/bin/omo" "$CODEX_HOME/bin/omo" "$CODEX_HOME"/plugins/cache/sisyphuslabs/omo/*/components/ulw-loop/dist/cli.js; do
|
|
92
|
+
for candidate in "$HOME/.local/bin/omo-agent-toolkit" "$CODEX_HOME/bin/omo-agent-toolkit" "$CODEX_HOME"/plugins/cache/sisyphuslabs/omo/*/components/ulw-loop/dist/cli.js; do
|
|
93
93
|
[ -f "$candidate" ] || [ -x "$candidate" ] || continue
|
|
94
94
|
if "$ULW_LOOP_NODE" "$candidate" ulw-loop help >/dev/null 2>&1; then
|
|
95
95
|
ULW_LOOP_CLI="$candidate"
|
|
@@ -97,15 +97,12 @@ elif [ -n "$ULW_LOOP_NODE" ]; then
|
|
|
97
97
|
fi
|
|
98
98
|
done
|
|
99
99
|
|
|
100
|
-
if [ -n "$ULW_LOOP_CLI" ] && [ -n "$ULW_LOOP_NODE" ]; then
|
|
101
|
-
omo() { "$ULW_LOOP_NODE" "$ULW_LOOP_CLI" "$@"; }
|
|
102
|
-
fi
|
|
103
100
|
fi
|
|
104
101
|
|
|
105
102
|
if [ -z "${ULW_LOOP_CLI:-}" ]; then
|
|
106
103
|
/bin/mkdir -p .omo/ulw-loop 2>/dev/null || mkdir -p .omo/ulw-loop 2>/dev/null || true
|
|
107
104
|
NOTE="${NOTE:-.omo/ulw-loop/bootstrap-notepad.md}"
|
|
108
|
-
printf '%s\n' "No ulw-loop-capable omo executable found; PATH omo may
|
|
105
|
+
printf '%s\n' "No ulw-loop-capable omo-agent-toolkit executable found; PATH omo-agent-toolkit may lack the Codex ulw-loop subcommand, and cached ulw-loop CLI was not found under ${CODEX_HOME:-$HOME/.codex}." >> "$NOTE" 2>/dev/null || true
|
|
109
106
|
printf '%s\n' "Install with npx lazycodex-ai install or set CODEX_LOCAL_BIN_DIR to a PATH directory." >&2
|
|
110
107
|
fi
|
|
111
108
|
```
|
|
@@ -113,17 +110,18 @@ If `ULW_LOOP_CLI` is empty, open the durable notepad first, record the missing C
|
|
|
113
110
|
|
|
114
111
|
Run one form:
|
|
115
112
|
```sh
|
|
116
|
-
omo ulw-loop create-goals --brief "<brief>" [--validation-batch-json <json-or-path>] --json
|
|
117
|
-
omo ulw-loop create-goals --brief-file <path> [--validation-batch-json <json-or-path>] --json
|
|
118
|
-
cat <brief> | omo ulw-loop create-goals --from-stdin [--validation-batch-json <json-or-path>] --json
|
|
113
|
+
omo-agent-toolkit ulw-loop create-goals --brief "<brief>" [--validation-batch-json <json-or-path>] --json
|
|
114
|
+
omo-agent-toolkit ulw-loop create-goals --brief-file <path> [--validation-batch-json <json-or-path>] --json
|
|
115
|
+
cat <brief> | omo-agent-toolkit ulw-loop create-goals --from-stdin [--validation-batch-json <json-or-path>] --json
|
|
119
116
|
```
|
|
120
117
|
If the existing aggregate is already complete, do not steer or force the
|
|
121
118
|
completed default state for unrelated new work. Start a fresh run with
|
|
122
|
-
`omo ulw-loop create-goals --session-id <new-id> ...`; use `--force`
|
|
119
|
+
`omo-agent-toolkit ulw-loop create-goals --session-id <new-id> ...`; use `--force`
|
|
123
120
|
only when deliberately overwriting completed evidence.
|
|
124
121
|
Write state through the CLI path. Do not hand-edit state files.
|
|
125
122
|
|
|
126
123
|
### 2. Refine success criteria + a Prometheus-grade QA and parallelism plan per goal
|
|
124
|
+
Shape every goal's objective and `successCriteria` by `references/define-goal.md`: its quality bar, objective anatomy, and criterion construction govern this step. Where the brief is silent on a constraint the work forks on, derive the default per that reference, record it via `annotate_ledger` (`--evidence` naming the repo fact, `--rationale` the default plus reversibility), and surface the assumed list in the first user-visible report so a wrong default is a one-line veto, not a finished run.
|
|
127
125
|
Gather context BEFORE planning with parallel `explorer` / `librarian` workers plus your own read-only tools.
|
|
128
126
|
First survey available skills: read every loosely-relevant skill's description, deliberately choose which this work uses, and prefer applying genuinely-relevant skills over working raw.
|
|
129
127
|
Then run tier triage per goal — rigor (LIGHT/HEAVY below) and shape (`delivery` default, or `research` when the deliverable is a cited answer, not an artifact) — and record both in an `annotate_ledger` steering entry. Default is LIGHT — a narrow change inside existing layers. Take HEAVY only on a fact you can point to: a new module / abstraction / domain model; auth, security, or session; an external integration; a DB schema or migration; concurrency, transaction boundaries, or cache invalidation; a cross-domain refactor; or the user signaled care or demanded review. When unsure, take HEAVY; upgrade the moment a HEAVY fact surfaces, never downgrade mid-run.
|
|
@@ -138,14 +136,14 @@ Use channel-table evidence verbs — not vibes.
|
|
|
138
136
|
Revise any criterion that lacks observable `expectedEvidence` or a named channel before execution.
|
|
139
137
|
|
|
140
138
|
### 3. Inspect state
|
|
141
|
-
Run `omo ulw-loop status --json`.
|
|
139
|
+
Run `omo-agent-toolkit ulw-loop status --json`.
|
|
142
140
|
Read pending goals, criteria IDs, current ledger head, blockers, and aggregate Codex objective.
|
|
143
141
|
|
|
144
142
|
## Execution Loop
|
|
145
143
|
Loop per goal. Cap at 5 cycles per goal. Cap identical same-criterion failures at 3.
|
|
146
144
|
|
|
147
145
|
### Acquire Next Goal
|
|
148
|
-
1. Run `omo ulw-loop complete-goals --json` and read the handoff, including criteria. After the first goal starts, a successful complete checkpoint normally prints the next goal instruction directly; use `complete-goals` as the manual fallback/resume path.
|
|
146
|
+
1. Run `omo-agent-toolkit ulw-loop complete-goals --json` and read the handoff, including criteria. After the first goal starts, a successful complete checkpoint normally prints the next goal instruction directly; use `complete-goals` as the manual fallback/resume path.
|
|
149
147
|
2. Call `get_goal` and inspect active Codex state.
|
|
150
148
|
3. Apply this table exactly:
|
|
151
149
|
|
|
@@ -154,7 +152,7 @@ Loop per goal. Cap at 5 cycles per goal. Cap identical same-criterion failures a
|
|
|
154
152
|
| no active goal | You MUST call `create_goal` — goal registration goes through the tool, never prose — with objective only from `instruction.json.objective`; do not copy lifecycle fields such as `status`. |
|
|
155
153
|
| same aggregate objective active | Continue the current ulw-loop story. |
|
|
156
154
|
| different goal active | STOP. Checkpoint blocked and surface the conflict. |
|
|
157
|
-
4. If retrying failed work, run `omo ulw-loop complete-goals --retry-failed --json`.
|
|
155
|
+
4. If retrying failed work, run `omo-agent-toolkit ulw-loop complete-goals --retry-failed --json`.
|
|
158
156
|
5. Never create a second Codex goal for the same aggregate objective.
|
|
159
157
|
|
|
160
158
|
### Per-Criterion Cycle
|
|
@@ -166,9 +164,9 @@ Loop per goal. Cap at 5 cycles per goal. Cap identical same-criterion failures a
|
|
|
166
164
|
6. CAPTURE: collect the observable artifact path: transcript, stdout, screenshot, assertion, status+body, diff, or parsed dump. No artifact written at the evidence path — not done; record BLOCKED and respawn QA.
|
|
167
165
|
7. CLEAN (PAIRED, NEVER SKIP): tear down every runtime artifact step 5 spawned BEFORE recording — server PIDs (`kill`, verify `kill -0` fails), `tmux` sessions (`tmux kill-session -t ulw-qa-<criterion>`; confirm `tmux ls`), browser / Playwright contexts (`.close()`), containers (`docker rm -f`), bound ports (`lsof -i :<port>` empty), temp sockets / files / dirs (`rm -rf` the `mktemp` paths), QA-only env vars, AND close every finished worker (v1 `close_agent`; on V2 finished workers end on their own — `interrupt_agent` any still running). Register each teardown as its own todo the moment the QA spawns the resource (scripts, tmux assets, browsers / agent-browser sessions, PIDs, ports) so none is forgotten. Embed a one-line cleanup receipt in the evidence string, e.g. `cleanup: killed 12345; tmux kill-session ulw-qa-foo; rm -rf /tmp/ulw.aB12cD; interrupt_agent w-3`. Missing receipt → record BLOCKED, not PASS.
|
|
168
166
|
8. RECORD one result immediately from the artifact you just wrote — never from memory or a later turn — stamping the capture tree `$(git rev-parse --short "HEAD^{tree}")` into the evidence:
|
|
169
|
-
- PASS: `omo ulw-loop record-evidence --goal-id <id> --criterion-id <id> --status pass --evidence "<observable> @tree:<short-tree> | <cleanup receipt>" --json`
|
|
170
|
-
- FAIL: `omo ulw-loop record-evidence --goal-id <id> --criterion-id <id> --status fail --evidence "<observable> @tree:<short-tree> | <cleanup receipt>" --notes "<diagnosis>" --json`
|
|
171
|
-
- BLOCKED: `omo ulw-loop record-evidence --goal-id <id> --criterion-id <id> --status blocked --evidence "<observable>" --notes "<safety/blocker/leftover-state>" --json`
|
|
167
|
+
- PASS: `omo-agent-toolkit ulw-loop record-evidence --goal-id <id> --criterion-id <id> --status pass --evidence "<observable> @tree:<short-tree> | <cleanup receipt>" --json`
|
|
168
|
+
- FAIL: `omo-agent-toolkit ulw-loop record-evidence --goal-id <id> --criterion-id <id> --status fail --evidence "<observable> @tree:<short-tree> | <cleanup receipt>" --notes "<diagnosis>" --json`
|
|
169
|
+
- BLOCKED: `omo-agent-toolkit ulw-loop record-evidence --goal-id <id> --criterion-id <id> --status blocked --evidence "<observable>" --notes "<safety/blocker/leftover-state>" --json`
|
|
172
170
|
9. If actual does not match expected, diagnose, respawn the right-sized worker with the failure context to fix minimally, and rerun the SAME criterion (including a fresh cleanup).
|
|
173
171
|
10. After 3 same-criterion failures, exit the goal with diagnosis.
|
|
174
172
|
11. After 5 cycles on one goal without required criteria passing, checkpoint failed.
|
|
@@ -177,7 +175,7 @@ Loop per goal. Cap at 5 cycles per goal. Cap identical same-criterion failures a
|
|
|
177
175
|
### Goal Completion
|
|
178
176
|
1. Non-final aggregate goal: confirm every `essential` criterion is `pass`; non-essential criteria may remain pending. Final aggregate goal: confirm every criterion across the whole plan is `pass`.
|
|
179
177
|
2. Call `get_goal` for a fresh snapshot.
|
|
180
|
-
3. Run `omo ulw-loop checkpoint --goal-id <id> --status complete --evidence "<criteria evidence summary>" --codex-goal-json <snapshot> --json`; on success it auto-starts and prints the next eligible goal unless `--no-advance` is passed.
|
|
178
|
+
3. Run `omo-agent-toolkit ulw-loop checkpoint --goal-id <id> --status complete --evidence "<criteria evidence summary>" --codex-goal-json <snapshot> --json`; on success it auto-starts and prints the next eligible goal unless `--no-advance` is passed.
|
|
181
179
|
4. If blocked or failed, checkpoint with `--status blocked` or `--status failed` and include diagnosis evidence.
|
|
182
180
|
5. If this is the final goal, run the final quality gate first and pass `--quality-gate-json`.
|
|
183
181
|
|
|
@@ -189,16 +187,16 @@ Trigger only for the final aggregate goal after every criterion in every goal is
|
|
|
189
187
|
3b. Only then spawn lazycodex-gate-reviewer with those artifact paths.
|
|
190
188
|
3c. The gate's approval binds to the frozen tree and full commit SHA and covers its three lanes — code quality, hands-on QA, and goal verification. Immediately append one durable `.omo/ulw-loop/ledger.jsonl` record per passing lane with the lane name, full SHA, verdict, and report artifact/source. Before reuse after continuation or compaction, re-read the ledger and require the exact lane/SHA pair; memory or an unstamped report is not coverage. A later rebase or amend that keeps the tree identical still has a new SHA and needs fresh lane stamps; changed content needs fresh review of the delta.
|
|
191
189
|
4. Treat timeout, missing deliverable, ack-only, `BLOCKED:`, or inconclusive review as a blocker. Any fix restarts the freeze at the new HEAD: re-run ONLY the proofs it invalidated and stamp the fresh output — never regenerate all evidence or relabel stale output to HEAD — re-review the delta at most TWICE; then record-review-blockers (step 5) and surface to the user.
|
|
192
|
-
5. If review remains blocked, run `omo ulw-loop record-review-blockers --goal-id <id> --title "<...>" --objective "<...>" --evidence "<review findings>" --codex-goal-json <snapshot> --json`.
|
|
190
|
+
5. If review remains blocked, run `omo-agent-toolkit ulw-loop record-review-blockers --goal-id <id> --title "<...>" --objective "<...>" --evidence "<review findings>" --codex-goal-json <snapshot> --json`.
|
|
193
191
|
6. If clean, checkpoint final completion:
|
|
194
192
|
```sh
|
|
195
|
-
omo ulw-loop checkpoint --goal-id <id> --status complete --evidence "<e2e evidence + manual QA notes>" --codex-goal-json <snapshot> --quality-gate-json <json-or-path> --json
|
|
193
|
+
omo-agent-toolkit ulw-loop checkpoint --goal-id <id> --status complete --evidence "<e2e evidence + manual QA notes>" --codex-goal-json <snapshot> --quality-gate-json <json-or-path> --json
|
|
196
194
|
```
|
|
197
195
|
`--quality-gate-json` shape:
|
|
198
196
|
```json
|
|
199
197
|
{
|
|
200
198
|
"codeReview":{"by":"lazycodex-code-reviewer","recommendation":"APPROVE","codeQualityStatus":"CLEAR","reportPath":"test/fixtures/artifacts/code-review.md","evidence":"Diff review passed.","blockers":[]},
|
|
201
|
-
"manualQa":{"by":"lazycodex-qa-executor","status":"passed","evidence":"CLI and data surfaces passed.","surfaceEvidence":[{"id":"surface-cli-pass","criterionRef":"C1","surface":"cli","invocation":"omo ulw-loop checkpoint --quality-gate-json sample-quality-gate.json --json","verdict":"passed","artifactRefs":["artifact-cli-pass"]},{"id":"surface-data-pass","criterionRef":"C2","surface":"data","invocation":"diff -u before-ledger.json after-ledger.json","verdict":"passed","artifactRefs":["artifact-data-diff"]}],"adversarialCases":[{"id":"adv-malformed-input","criterionRef":"C3","scenario":"malformed gate input omits manual QA evidence","expectedBehavior":"validator rejects ULW_LOOP_QUALITY_GATE_INVALID","verdict":"passed","artifactRefs":["artifact-cli-reject"]}],"artifactRefs":[{"id":"artifact-cli-pass","kind":"cli-transcript","description":"CLI pass artifact.","path":"test/fixtures/artifacts/cli-pass.txt"},{"id":"artifact-cli-reject","kind":"log","description":"Reject log artifact.","path":"test/fixtures/artifacts/rejection.txt"},{"id":"artifact-data-diff","kind":"data-diff","description":"Data diff artifact.","path":"test/fixtures/artifacts/data-diff.txt"}]},
|
|
199
|
+
"manualQa":{"by":"lazycodex-qa-executor","status":"passed","evidence":"CLI and data surfaces passed.","surfaceEvidence":[{"id":"surface-cli-pass","criterionRef":"C1","surface":"cli","invocation":"omo-agent-toolkit ulw-loop checkpoint --quality-gate-json sample-quality-gate.json --json","verdict":"passed","artifactRefs":["artifact-cli-pass"]},{"id":"surface-data-pass","criterionRef":"C2","surface":"data","invocation":"diff -u before-ledger.json after-ledger.json","verdict":"passed","artifactRefs":["artifact-data-diff"]}],"adversarialCases":[{"id":"adv-malformed-input","criterionRef":"C3","scenario":"malformed gate input omits manual QA evidence","expectedBehavior":"validator rejects ULW_LOOP_QUALITY_GATE_INVALID","verdict":"passed","artifactRefs":["artifact-cli-reject"]}],"artifactRefs":[{"id":"artifact-cli-pass","kind":"cli-transcript","description":"CLI pass artifact.","path":"test/fixtures/artifacts/cli-pass.txt"},{"id":"artifact-cli-reject","kind":"log","description":"Reject log artifact.","path":"test/fixtures/artifacts/rejection.txt"},{"id":"artifact-data-diff","kind":"data-diff","description":"Data diff artifact.","path":"test/fixtures/artifacts/data-diff.txt"}]},
|
|
202
200
|
"gateReview":{"by":"lazycodex-gate-reviewer","recommendation":"APPROVE","reportPath":"test/fixtures/artifacts/gate-review.md","evidence":"Gate review passed.","blockers":[]},
|
|
203
201
|
"iteration":{"fullRerun":true,"status":"passed","rerunCommands":["bunx vitest run packages/omo-codex/plugin/components/ulw-loop/test/quality-gate-doc.test.ts"],"evidence":"Focused rerun passed."},
|
|
204
202
|
"criteriaCoverage":{"totalCriteria":3,"passCount":3,"originalIntent":"User wanted artifact-backed completion.","desiredOutcome":"Behavior ships with review and QA evidence.","userOutcomeReview":"Result matches brief and goals.","adversarialClassesCovered":["malformed_input","stale_state"]}
|
|
@@ -219,10 +217,10 @@ Use steering only for structured evidence-backed mutation. Reject natural-langua
|
|
|
219
217
|
| annotate_ledger | Audit-only note | `--evidence`, `--rationale` |
|
|
220
218
|
| mark_blocked_superseded | Old story replaced by new evidence | `--goal-id`, `--replacements?`, `--evidence`, `--rationale` |
|
|
221
219
|
|
|
222
|
-
Command form: `omo ulw-loop steer --kind <kind> [<kind-specific-fields>] --evidence "<...>" --rationale "<...>" --json`. For multiple evidence-backed plan-shape changes discovered together, pass `--proposals-json <json-or-path>` with an array of proposals; the batch applies atomically or rejects without partial plan mutation.
|
|
220
|
+
Command form: `omo-agent-toolkit ulw-loop steer --kind <kind> [<kind-specific-fields>] --evidence "<...>" --rationale "<...>" --json`. For multiple evidence-backed plan-shape changes discovered together, pass `--proposals-json <json-or-path>` with an array of proposals; the batch applies atomically or rejects without partial plan mutation.
|
|
223
221
|
|
|
224
222
|
Validation batches are optional aggregate-mode review boundaries declared at create time with `--validation-batch-json`. A batch-final member requires all other members resolved, all member criteria pass, and a member-spanning quality gate; split/supersede steering keeps batch membership updated.
|
|
225
|
-
Structured prompt directives accepted: `OMO_ULW_LOOP_STEER: { ... }`, `omo.ulw-loop.steer: {...}`, `omo ulw-loop steer: {...}`.
|
|
223
|
+
Structured prompt directives accepted: `OMO_ULW_LOOP_STEER: { ... }`, `omo.ulw-loop.steer: {...}`, `omo ulw-loop steer: {...}`, `omo-agent-toolkit ulw-loop steer: {...}`.
|
|
226
224
|
|
|
227
225
|
## Constraints
|
|
228
226
|
1. NEVER call `update_goal` mid-aggregate; only on final story after the quality gate passes.
|
|
@@ -34,7 +34,7 @@ Example opening (adapt the wording, keep every commitment):
|
|
|
34
34
|
|
|
35
35
|
## INTENT ROUTING - pick ONE intent reference
|
|
36
36
|
|
|
37
|
-
**Review modifiers are a gate trigger, not a style cue.** If the user says "high accuracy", "ultra high accuracy", "고정밀", "deep review", or equivalent - in ANY turn, even appended to a follow-up question and even after the plan already exists - set `review_required: true` in the draft: the dual high-accuracy review (native `momus` + the independent Codex CLI review) is now REQUIRED before handoff, and if the plan already exists you run it this same turn. Answering the current question more carefully does NOT satisfy it. This does NOT choose CLEAR/UNCLEAR and does NOT suppress interview.
|
|
37
|
+
**Review modifiers are a gate trigger, not a style cue.** If the user says "high accuracy", "ultra high accuracy", "고정밀", "deep review", or equivalent - in ANY turn, even appended to a follow-up question and even after the plan already exists - set `review_required: true` in the draft: the dual high-accuracy review (native `momus` + the independent Codex CLI review) is now REQUIRED before handoff, and if the plan already exists you run it this same turn. The review runs under the bounded convergence contract in `full-workflow.md`: a 5-round cap (unlimited only on explicit user request), evidence-backed blocker eligibility, and approval-with-notes counting as approval. Answering the current question more carefully does NOT satisfy it. This does NOT choose CLEAR/UNCLEAR and does NOT suppress interview.
|
|
38
38
|
|
|
39
39
|
After grounding, make ONE judgment, record `intent: clear|unclear` plus `review_required`, **ANNOUNCE both to the user in one line**, then load ONE intent reference (you ALSO read `references/full-workflow.md` for the shared mechanics - see below). The test keys on whether the desired **OUTCOME** is clear, NOT on request length. This verdict line and the opening announcement above are the two mandatory user-visible signals of a planning session - it tells the user whether they will be interviewed and whether high-accuracy review is already requested; never skip either.
|
|
40
40
|
|
|
@@ -64,7 +64,7 @@ Both invocations are resume-safe no-ops for artifacts already present. Do NOT ha
|
|
|
64
64
|
|
|
65
65
|
## Plan artifact producer contract
|
|
66
66
|
|
|
67
|
-
When producing the plan, encode every executable item as a column-zero Markdown task row: implementation rows MUST match `- [ ] N. <title>` (where `N` is a positive decimal integer), and final-verifier rows MUST match `- [ ] F<number>. <title>`. Prose headings, numbered paragraphs, and ordinary bullets are not task substitutes and MUST NOT be counted as implementation or final-verifier tasks. Before handoff, run a structural self-check over the plan: verify that every implementation row and final-verifier row is column-zero, matches its required grammar, and appears in the intended `## Todos` or `## Final verification wave` section; verify that no prose heading or bullet is being used as a task; and repair the plan before handoff if any check fails.
|
|
67
|
+
When producing the plan, encode every executable item as a column-zero Markdown task row: implementation rows MUST match `- [ ] N. <title>` (where `N` is a positive decimal integer), and final-verifier rows MUST match `- [ ] F<number>. <title>`. Prose headings, numbered paragraphs, and ordinary bullets are not task substitutes and MUST NOT be counted as implementation or final-verifier tasks. Before handoff, run a structural self-check over the plan: verify that every implementation row and final-verifier row is column-zero, matches its required grammar, and appears in the intended `## Todos` or `## Final verification wave` section; verify that no prose heading or bullet is being used as a task; verify that every implementation row carries a nested `Recommended task executor category:` line (final-verifier rows default to `unspecified-high` when unannotated); and repair the plan before handoff if any check fails.
|
|
68
68
|
|
|
69
69
|
## Universal invariants (hold on every path)
|
|
70
70
|
|
|
@@ -72,7 +72,7 @@ When producing the plan, encode every executable item as a column-zero Markdown
|
|
|
72
72
|
- **Full scope is the default.** Plan the ENTIRE request; "MVP", "v1", "phase 1", or any reduced subset is never an option you invent or ask about - it exists only if the user introduces it. Scope OUT / Must-NOT-Have entries are guardrails against unrequested additions, never reductions of the request.
|
|
73
73
|
- **Explore before asking.** Discoverable facts (repo/system/docs truth) -> research and cite, never ask. Preferences/tradeoffs -> the only things you bring to the user. When unsure which, treat it as a user-decision.
|
|
74
74
|
- **CodeGraph first when present.** Use `codegraph_explore` for repo how/where/what/flow questions before wider reads; if codegraph_* tools are absent, inactive/uninitialized, or cold-start unavailable, continue with Read/Grep/Glob/LSP and the ast-grep skill.
|
|
75
|
-
- **Two filters** on every candidate question, in order: (1) Could collected evidence answer it? -> explore instead. (2) Could the user's stated intent plus a defensible default answer it? -> adopt the default, record it, do not ask - UNLESS it is an owner-decision, which always survives as a question even when a default exists: anything irreversible / destructive / safety-critical, or a cross-cutting product choice the user lives with (public config surface, distribution / packaging, external dependency or pinned SHA, data / schema shape). Default the reversible internals; surface the owner-decisions.
|
|
75
|
+
- **Two filters** on every candidate question, in order: (1) Could collected evidence answer it? -> explore instead. (2) Could the user's stated intent plus a defensible default answer it? -> adopt the default, record it, do not ask - UNLESS it is an owner-decision, which always survives as a question even when a default exists: anything irreversible / destructive / safety-critical, or a cross-cutting product choice the user lives with (public config surface, distribution / packaging, external dependency or pinned SHA, data / schema shape, real budget / paid-service spend, expected scale or capacity target, target-audience / compliance limits). Extrinsic constraints (budget, mandated stack, scale, audience) leave no repo evidence, so exploration can never surface them - sweep those axes explicitly once per plan and classify each as explored, defaulted (ledger), or asked. Default the reversible internals; surface the owner-decisions.
|
|
76
76
|
- **Explore to sufficiency, then STOP.** One research wave per open question; stop when the clearance check is answerable; never re-explore to double-check.
|
|
77
77
|
- **Parallel-dispatch** independent research in ONE turn and keep working while it runs. Subagent outputs are CLAIMS until you independently verify them.
|
|
78
78
|
- **Approval is not execution.** Approval authorizes writing the plan ONLY, never implementation. ONE request -> ONE plan, however large.
|