opencode-agent-skill 10.0.0 → 11.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +67 -0
  2. package/README.md +49 -3
  3. package/bin/ocskill.mjs +330 -5
  4. package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +75 -0
  5. package/docs/V11-PERCEPTION-ADAPTIVE.md +220 -0
  6. package/evals/router-triggers.json +82 -0
  7. package/evals/routing.json +76 -0
  8. package/evals/v11/tasks.json +122 -0
  9. package/global-config/agents/merge-arbiter.md +12 -0
  10. package/global-config/agents/visual-verifier.md +12 -0
  11. package/global-config/plugins/ues-router/index.js +272 -2
  12. package/global-config/plugins/ues-router/router.js +27 -3
  13. package/global-config/skills/browser-qa/SKILL.md +14 -0
  14. package/global-config/skills/browser-qa/references/workflow.md +11 -0
  15. package/global-config/skills/browser-security/SKILL.md +12 -0
  16. package/global-config/skills/component-visual-testing/SKILL.md +10 -0
  17. package/global-config/skills/design-source/SKILL.md +10 -0
  18. package/global-config/skills/design-source/references/workflow.md +12 -0
  19. package/global-config/skills/dynamic-workflow/SKILL.md +18 -0
  20. package/global-config/skills/dynamic-workflow/references/workflow.md +19 -0
  21. package/global-config/skills/responsive-verification/SKILL.md +10 -0
  22. package/global-config/skills/skill-authoring/SKILL.md +12 -0
  23. package/global-config/skills/skill-evaluation/SKILL.md +17 -0
  24. package/global-config/skills/visual-fidelity/SKILL.md +14 -0
  25. package/global-config/skills/visual-fidelity/references/workflow.md +14 -0
  26. package/lib/browser-adapter.mjs +82 -0
  27. package/lib/browser-runtime.mjs +193 -0
  28. package/lib/capability-registry.mjs +109 -0
  29. package/lib/context-engine-v11.mjs +146 -0
  30. package/lib/context-manifest.mjs +16 -3
  31. package/lib/control-center.mjs +12 -2
  32. package/lib/dynamic-workflow.mjs +179 -0
  33. package/lib/eval-ablation.mjs +43 -1
  34. package/lib/eval-report.mjs +72 -0
  35. package/lib/eval-telemetry.mjs +61 -0
  36. package/lib/evidence-budget.mjs +84 -0
  37. package/lib/evidence-store.mjs +178 -0
  38. package/lib/hermes-bridge.mjs +45 -1
  39. package/lib/model-config.mjs +9 -1
  40. package/lib/model-policy.mjs +52 -1
  41. package/lib/orchestrator-policy.mjs +1 -1
  42. package/lib/png-diff.mjs +229 -0
  43. package/lib/prompt-cache.mjs +60 -0
  44. package/lib/skill-quality.mjs +72 -0
  45. package/lib/task-engine.mjs +78 -4
  46. package/lib/ui-inspector.mjs +152 -0
  47. package/lib/v11-metrics.mjs +64 -0
  48. package/lib/visual-spec.mjs +159 -0
  49. package/package.json +10 -5
  50. package/scripts/eval-ablation.mjs +4 -1
  51. package/scripts/validate-v11-suite.mjs +58 -0
  52. package/scripts/validate.mjs +16 -4
package/CHANGELOG.md CHANGED
@@ -6,6 +6,73 @@ The project follows Semantic Versioning.
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [11.0.0] - 2026-09-22
10
+
11
+ ### Released
12
+ - Promoted V11 perception-aware adaptive execution to stable after the full local CI gate passed on Windows with 234 tests total, 232 passed, 0 failed and 2 platform-specific skips.
13
+ - Stable npm installs use the default `latest` dist-tag, so users install with `npm install -g opencode-agent-skill`.
14
+ - Includes content-addressed Evidence Store, adaptive EvidenceBudget/context externalization, stable-prefix prompt telemetry, capability-aware model routing, visual geometry receipts, deterministic PNG diff/crop, responsive/design-token inspection, optional Playwright browser inspection, cost-aware dynamic workflow scheduling, 48 focused skills and 12 subagents.
15
+
16
+ ### Verified
17
+ - Syntax: 178 JavaScript modules.
18
+ - Catalog: 48 skills, 11 commands and 12 subagents.
19
+ - Router: 129 cases, 294/294 required routes and 61/61 negative guards.
20
+ - V11 contracts: 13 tasks across 6 categories and 15 required runtime files.
21
+ - Live/long/polyglot fixture validation: 20 / 5 / 8 tasks.
22
+ - npm pack dry-run, packed-install smoke and plain one-command install/resource auto-sync smoke: PASS.
23
+
24
+ ### Changed
25
+ - Package version is 11.0.0.
26
+ - V11 becomes the stable npm release line; V10 remains in Git history as the previous stable release.
27
+
28
+
29
+ ### V11 dev.2
30
+ - Added project-local, fail-closed Playwright browser inspection that returns bounded semantic elements, bounding boxes, computed visual properties and a screenshot path while treating page content as untrusted evidence.
31
+ - Exposed browser inspection through CLI and the OpenCode V2 router without making Playwright a required package dependency.
32
+ - Fixed duplicate `ui_layout` router tool registration and aligned durable context packs with V11 context schema version 6.
33
+ - Extended V11 validation and regression coverage for the browser runtime.
34
+ - Package development version is 11.0.0-dev.2.
35
+
36
+ ## [11.0.0-dev.1] - 2026-09-22
37
+
38
+ ### Added
39
+ - Adaptive context engine that externalizes oversized source/test/reference excerpts into content-addressed evidence pointers while keeping bounded inline previews.
40
+ - Deterministic UI layout and design-token tools exposed to the OpenCode V2 runtime.
41
+ - Cost- and modality-aware workflow scheduling with inline thresholds, deterministic-first waves, independent LLM/vision concurrency and bounded wave cost.
42
+ - Repeated-stable prompt ratio telemetry and an optional fail-closed ablation gate for repeated input efficiency.
43
+ - Stronger V11 contract coverage for adaptive context, UI inspection and evidence externalization.
44
+
45
+ ### Changed
46
+ - V11 router metadata now reports version 11.
47
+ - Task context packs transport large contextual evidence through `evidence:sha256` pointers and report externalized byte/ref counts.
48
+ - Capability-aware model routing fails closed when an enabled configured model set cannot satisfy required capabilities such as vision/browser.
49
+ - Package development version is 11.0.0-dev.1.
50
+
51
+ ### Fixed
52
+ - V11 eval-report regression tests now account for expanded adaptive telemetry coverage instead of using the old V10-only telemetry shape.
53
+
54
+ ## [11.0.0-dev.0] - 2026-09-22
55
+
56
+ ### Added
57
+ - Content-addressed Evidence Store with bounded retrieval, deduplication and garbage collection.
58
+ - Adaptive Evidence Budget planning and context-manifest integration for evidence-efficient execution.
59
+ - Prompt stable-prefix/cache telemetry for repeated-input measurement.
60
+ - Capability registry and capability-aware model selection for coding, reasoning, tools, vision, browser, filesystem and long-context needs.
61
+ - Visual specification, geometry receipts, responsive viewport matrix, dependency-free PNG decode/diff/crop and bounded visual repair planning.
62
+ - Browser QA adapter with CLI-first verification planning, targeted semantic evidence and explicit untrusted-page security boundaries.
63
+ - Cost-aware dynamic workflow scheduler separating deterministic work from LLM/vision work.
64
+ - Skill-quality linting for entrypoint size, metadata and routing-description collision detection.
65
+ - Nine V11 skills: visual-fidelity, browser-qa, design-source, responsive-verification, component-visual-testing, browser-security, skill-authoring, skill-evaluation and dynamic-workflow.
66
+ - Two V11 subagents: visual-verifier and merge-arbiter.
67
+ - Optional Hermes sidecar workflow contract with evidence-pointer transport.
68
+ - V11 Control Center evidence-store/runtime telemetry and V11-specific tests/eval routing cases.
69
+
70
+ ### Changed
71
+ - Model policy schema supports per-model capability metadata plus cost, latency and quality hints.
72
+ - Adaptive model resolution can select configured models by required task capabilities.
73
+ - OpenCode V2 router recognizes visual/browser/design/skill-workflow intents and keeps FAST routing selective.
74
+ - Package version is 11.0.0-dev.0 while npm latest remains V10 stable until V11 release gates pass.
75
+
9
76
  ## [10.0.0] - 2026-09-22
10
77
 
11
78
  ### Released
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # OpenCode Universal Engineering System (UES)
2
2
 
3
- > **V10 stable: 10.0.0** — bản phát hành chính thức của V10, dùng trực tiếp qua npm `latest`.
3
+ > **V11 stable: 11.0.0** — perception-aware adaptive execution engine với evidence-efficient context, capability routing và visual/browser verification.
4
4
  > UES là bộ công cụ hỗ trợ OpenCode xử lý dự án lớn, tác vụ dài và quy trình kỹ thuật cần kiểm chứng bằng bằng chứng thực tế.
5
5
 
6
6
  [![npm version](https://img.shields.io/npm/v/opencode-agent-skill.svg)](https://www.npmjs.com/package/opencode-agent-skill)
@@ -80,6 +80,49 @@ ocskill dashboard . --serve
80
80
 
81
81
  ---
82
82
 
83
+ ## V11 có gì?
84
+
85
+ V11 chuyển UES từ một reliability harness thành **perception-aware adaptive execution engine**. Mục tiêu là model yếu chỉ nhận đúng bằng chứng cần thiết, dùng đúng capability/model/tool và có thể kiểm chứng UI bằng semantic structure + geometry + pixels thay vì đoán từ screenshot.
86
+
87
+ - content-addressed Evidence Store dưới `.ues-cache/evidence-v1/`; output lớn được lưu theo SHA-256 và context nhận bounded excerpt + `evidence:sha256:...` pointer;
88
+ - adaptive EvidenceBudget chia context cho instructions/task/source/tests/references/history/tools theo rủi ro và loại task thay vì luôn tiêu hết FAST/STANDARD/DEEP budget;
89
+ - prompt-envelope telemetry tách stable prefix và dynamic tail để đo cacheable ratio/repeated stable input;
90
+ - capability-aware model routing bổ sung `vision`, `browser`, `reasoning`, `toolCalling`, `filesystem`, `longContext` nhưng vẫn giữ light/standard/heavy để tương thích;
91
+ - `VISUAL_SPEC.json` + geometry receipt kiểm tra vị trí/kích thước element có tolerance rõ ràng;
92
+ - zero-dependency PNG diff/crop định vị vùng pixel sai rồi chỉ đưa crop cần thiết cho vision model;
93
+ - browser QA dùng targeted semantic/accessibility evidence + bounding boxes + screenshots; webpage content luôn được xem là untrusted;
94
+ - responsive viewport matrix và component visual testing/Storybook routing;
95
+ - dynamic workflow scheduler tách deterministic task khỏi LLM/vision fan-out, serialize overlapping writers và giới hạn concurrency;
96
+ - skill linter đo entrypoint size và description collisions; catalog tăng từ 39 lên **48 skills** nhưng router chỉ load skill có tín hiệu hẹp;
97
+ - hai subagent mới: `ues-visual-verifier` và `ues-merge-arbiter`, tổng **12 subagents**;
98
+ - Hermes được nâng thành optional sidecar: UES vẫn sở hữu durable state/evidence/safety, Hermes chỉ nhận bounded schedule/context khi được dùng;
99
+ - Control Center hiển thị Evidence Store và V11 runtime efficiency state.
100
+
101
+ Các CLI V11 chính:
102
+
103
+ ```cmd
104
+ ocskill store status .
105
+ ocskill store get evidence:sha256:<hash> . --max 12000
106
+ ocskill capabilities "match this screenshot in the browser"
107
+ ocskill visual spec VISUAL_SPEC.json
108
+ ocskill visual geometry VISUAL_SPEC.json actual-boxes.json
109
+ ocskill visual compare expected.png actual.png --threshold 16
110
+ ocskill visual crop actual.png failed-region.png --x 10 --y 20 --width 300 --height 120
111
+ ocskill visual viewports
112
+ ocskill browser capability .
113
+ ocskill browser plan http://localhost:3000 --target Checkout
114
+ ocskill browser inspect http://localhost:3000 . --selector "button" --width 1440 --height 900
115
+ ocskill ui tokens src/styles.css
116
+ ocskill ui layout boxes.json --width 390 --height 844
117
+ ocskill workflow-plan PLAN.json --max-concurrent 4
118
+ ocskill skills lint .
119
+ ocskill models capability provider/model --vision on --browser on --quality 0.9
120
+ ```
121
+
122
+ V11 stable được promote sau full local CI trên Windows PASS: 234 tests, 232 pass, 0 fail, 2 platform-specific skips; router/contract/package/install smoke đều PASS. Stable npm install dùng `npm install -g opencode-agent-skill`.
123
+
124
+ ---
125
+
83
126
  ## V10 có gì?
84
127
 
85
128
  V10 tập trung vào **minimum context necessary for maximum task success**: giảm context luôn nạp nhưng không cắt các lớp correctness, verification hay recovery.
@@ -281,6 +324,8 @@ Child session không được tự động merge, push, publish hoặc deploy.
281
324
  | `ues-critic` | Tìm giả định sai, counterexample và điểm yếu trong phương án |
282
325
  | `ues-verifier` | Xác minh độc lập theo acceptance criteria |
283
326
  | `ues-integration-verifier` | Kiểm tra tích hợp giữa nhiều task và luồng end-to-end |
327
+ | `ues-visual-verifier` | Xác minh độc lập screenshot/geometry/responsive/interaction mà không sửa code |
328
+ | `ues-merge-arbiter` | Hòa giải conflict giữa các task đã verify theo intent/evidence, không push/publish/deploy |
284
329
 
285
330
  `ues-executor` là subagent chính có quyền sửa code theo scope được giao. Các agent còn lại chủ yếu phục vụ phân tích, review và xác minh.
286
331
 
@@ -386,7 +431,7 @@ ocskill work status checkout .
386
431
 
387
432
  ## Skills
388
433
 
389
- UES hiện có **39 skills**. Router chỉ chọn các skill phù hợp thay vì nạp toàn bộ catalog vào mỗi task.
434
+ V11 development hiện có **48 skills**; V10 stable có 39. Router vẫn chỉ chọn tập skill phù hợp thay vì nạp toàn bộ catalog vào mỗi task. Router chỉ chọn các skill phù hợp thay vì nạp toàn bộ catalog vào mỗi task.
390
435
 
391
436
  Một số process skill quan trọng:
392
437
 
@@ -491,7 +536,7 @@ Proposal có `shadowRequired` chỉ được đưa trở lại context sau khi v
491
536
 
492
537
  ## Hermes adapter
493
538
 
494
- UES hỗ trợ Hermes theo dạng adapter tùy chọn, không nhúng Hermes runtime vào core:
539
+ V11 giữ Hermes theo dạng **optional sidecar**, không nhúng Hermes runtime vào core. UES vẫn sở hữu durable state, evidence và safety boundaries:
495
540
 
496
541
  ```cmd
497
542
  ocskill hermes status
@@ -580,6 +625,7 @@ Kiểm tra long-horizon và polyglot suite:
580
625
  ```cmd
581
626
  npm run evals:long:validate
582
627
  npm run evals:polyglot:validate
628
+ npm run evals:v11:validate
583
629
  ```
584
630
 
585
631
  Chạy benchmark với model thật:
package/bin/ocskill.mjs CHANGED
@@ -59,13 +59,22 @@ import {
59
59
  } from "../lib/task-engine.mjs"
60
60
  import { reviewScope } from "../lib/review-scope.mjs"
61
61
  import { buildVerificationPlan } from "../lib/verification-plan.mjs"
62
- import { resolveAdaptiveModel, resolveModel } from "../lib/model-policy.mjs"
62
+ import { resolveAdaptiveModel, resolveCapabilityModel, resolveModel } from "../lib/model-policy.mjs"
63
63
  import { createVerificationReceipt } from "../lib/evidence-receipt.mjs"
64
64
  import { classifyEngineeringTask } from "../lib/orchestrator-policy.mjs"
65
65
  import { createTaskSandbox, integrateTaskSandbox, listTaskSandboxes, removeTaskSandbox } from "../lib/worktree-sandbox.mjs"
66
66
  import { analyzeEvalTraces, saveLearningAnalysis, readLearningState, acceptLearning, promoteLearning } from "../lib/learning-engine.mjs"
67
- import { hermesStatus, buildHermesDelegationPrompt, hermesOneShotArgs } from "../lib/hermes-bridge.mjs"
67
+ import { hermesStatus, buildHermesDelegationPrompt, buildHermesWorkflowPrompt, hermesOneShotArgs, hermesSidecarPlan } from "../lib/hermes-bridge.mjs"
68
68
  import { readModelPolicy, validateModelID, writeModelPolicy } from "../lib/model-config.mjs"
69
+ import { evidenceStoreStatus, gcEvidenceStore, getEvidence, putEvidence } from "../lib/evidence-store.mjs"
70
+ import { inferTaskCapabilities } from "../lib/capability-registry.mjs"
71
+ import { browserCapability, buildBrowserVerificationPlan } from "../lib/browser-adapter.mjs"
72
+ import { inspectBrowserPage, summarizeBrowserInspection } from "../lib/browser-runtime.mjs"
73
+ import { comparePngFiles, cropPngFile } from "../lib/png-diff.mjs"
74
+ import { createGeometryReceipt, responsiveViewportMatrix, validateVisualSpec } from "../lib/visual-spec.mjs"
75
+ import { planDynamicWorkflow } from "../lib/dynamic-workflow.mjs"
76
+ import { lintSkillCatalog } from "../lib/skill-quality.mjs"
77
+ import { designTokenEvidence, extractDesignTokens, inspectResponsiveLayout } from "../lib/ui-inspector.mjs"
69
78
  import {
70
79
  clipOutput,
71
80
  errorMessage,
@@ -117,7 +126,14 @@ Usage:
117
126
  ocskill sandbox <action> ... Create, integrate and clean isolated Git worktree sandboxes
118
127
  Also supports capability/exec for fail-closed container verification
119
128
  ocskill learn <action> ... Analyze eval traces and promote benchmark-validated lessons
120
- ocskill hermes <action> ... Optional Hermes adapter/status
129
+ ocskill hermes <action> ... Optional Hermes sidecar/status/task/workflow planning
130
+ ocskill store <status|put|get|gc> ... Content-addressed evidence storage and bounded retrieval
131
+ ocskill capabilities <text> Infer required execution/model capabilities
132
+ ocskill visual <action> ... Geometry receipts, PNG diff/crop and viewport matrix
133
+ ocskill browser <action> ... Browser capability, plan and bounded Playwright inspection
134
+ ocskill ui <tokens|layout> ... Extract design tokens or verify responsive geometry
135
+ ocskill workflow-plan <plan> Cost-aware deterministic/LLM/vision wave schedule
136
+ ocskill skills lint [dir] Lint skill size, metadata and routing-description collisions
121
137
  ocskill dashboard [dir] [--serve] [--port N]
122
138
  Generate/serve the local UES Control Center
123
139
  ocskill models <status|on|off|set|role> ...
@@ -154,6 +170,7 @@ Usage:
154
170
  ocskill models on|off
155
171
  ocskill models set <light|standard|heavy> <provider/model[#variant]>
156
172
  ocskill models role <role> <light|standard|heavy>
173
+ ocskill models capability <provider/model> [--vision on|off] [--browser on|off] [--reasoning on|off] [--long-context on|off] [--cost low|medium|high] [--latency fast|medium|slow] [--quality 0..1]
157
174
 
158
175
  --force backs up and replaces/removes state owned by another package.
159
176
  `)
@@ -741,7 +758,8 @@ async function modelPolicy() {
741
758
  const normalizedAttempt = Number.isInteger(attempt) && attempt > 0 ? attempt : 1
742
759
  const taskText = optionValue(args, "--text")
743
760
  if (taskText) {
744
- printJson(resolveAdaptiveModel(role, normalizedAttempt, classifyEngineeringTask(taskText), policy))
761
+ const taskPolicy = classifyEngineeringTask(taskText)
762
+ printJson(resolveCapabilityModel(role, normalizedAttempt, taskText, taskPolicy, policy))
745
763
  return
746
764
  }
747
765
  printJson(resolveModel(role, normalizedAttempt, policy))
@@ -794,6 +812,48 @@ async function modelsControl() {
794
812
  return
795
813
  }
796
814
 
815
+ if (action === "capability") {
816
+ const model = args[2]
817
+ if (!validateModelID(model)) {
818
+ console.error("Usage: ocskill models capability <provider/model> [capability flags]")
819
+ process.exitCode = 2
820
+ return
821
+ }
822
+ const current = policy.capabilities?.[model] || {}
823
+ const boolFlag = (name, prior) => {
824
+ const value = optionValue(args, name)
825
+ if (value == null) return prior
826
+ if (!["on", "off", "true", "false"].includes(String(value).toLowerCase())) throw new Error(name + " must be on/off")
827
+ return ["on", "true"].includes(String(value).toLowerCase())
828
+ }
829
+ const qualityRaw = optionValue(args, "--quality")
830
+ const quality = qualityRaw == null ? current.quality : Number(qualityRaw)
831
+ if (qualityRaw != null && (!Number.isFinite(quality) || quality < 0 || quality > 1)) throw new Error("--quality must be from 0 to 1")
832
+ const cost = optionValue(args, "--cost") || current.costClass
833
+ const latency = optionValue(args, "--latency") || current.latencyClass
834
+ if (cost && !["low", "medium", "high"].includes(cost)) throw new Error("--cost must be low|medium|high")
835
+ if (latency && !["fast", "medium", "slow"].includes(latency)) throw new Error("--latency must be fast|medium|slow")
836
+ policy = await writeModelPolicy(getConfigDir(), {
837
+ capabilities: {
838
+ [model]: {
839
+ ...current,
840
+ coding: boolFlag("--coding", current.coding),
841
+ reasoning: boolFlag("--reasoning", current.reasoning),
842
+ toolCalling: boolFlag("--tool-calling", current.toolCalling),
843
+ vision: boolFlag("--vision", current.vision),
844
+ browser: boolFlag("--browser", current.browser),
845
+ filesystem: boolFlag("--filesystem", current.filesystem),
846
+ longContext: boolFlag("--long-context", current.longContext),
847
+ ...(cost ? { costClass: cost } : {}),
848
+ ...(latency ? { latencyClass: latency } : {}),
849
+ ...(quality != null ? { quality } : {}),
850
+ },
851
+ },
852
+ })
853
+ printJson(policy)
854
+ return
855
+ }
856
+
797
857
  console.error("Usage: ocskill models <status|on|off|set|role> ...")
798
858
  process.exitCode = 2
799
859
  }
@@ -956,6 +1016,40 @@ async function hermesControl() {
956
1016
  printJson(hermesStatus())
957
1017
  return
958
1018
  }
1019
+ if (action === "workflow" || action === "exec-workflow") {
1020
+ const slug = args[2]
1021
+ const root = positionalArg(args, 3) || process.cwd()
1022
+ if (!slug) {
1023
+ console.error("Usage: ocskill hermes <workflow|exec-workflow> <slug> [dir] [--max-concurrent N]")
1024
+ process.exitCode = 2
1025
+ return
1026
+ }
1027
+ const planFile = path.join(path.resolve(root), ".ues-work", slug, "PLAN.json")
1028
+ const plan = readJsonFile(planFile)
1029
+ const schedule = planDynamicWorkflow(plan.tasks || [], {
1030
+ maxConcurrent: optionInt(args, "--max-concurrent", 4),
1031
+ })
1032
+ const sidecar = hermesSidecarPlan({ mode: "dynamic-workflow", maxConcurrent: schedule.maxConcurrent })
1033
+ const prompt = buildHermesWorkflowPrompt({ slug, goal: plan.goal || null, plan }, schedule)
1034
+ if (action === "workflow") {
1035
+ printJson({ sidecar, schedule, prompt })
1036
+ return
1037
+ }
1038
+ const status = hermesStatus()
1039
+ if (!status.available) {
1040
+ console.error(status.error || "Hermes CLI is unavailable")
1041
+ process.exitCode = 1
1042
+ return
1043
+ }
1044
+ const result = runCapture("hermes", hermesOneShotArgs(prompt), {
1045
+ cwd: path.resolve(root),
1046
+ maxBuffer: 8 * 1024 * 1024,
1047
+ })
1048
+ if (result.stdout) process.stdout.write(result.stdout)
1049
+ if (result.stderr) process.stderr.write(result.stderr)
1050
+ if ((result.status ?? 1) !== 0) process.exitCode = result.status ?? 1
1051
+ return
1052
+ }
959
1053
  if (action === "prompt" || action === "exec") {
960
1054
  const slug = args[2]
961
1055
  const taskID = args[3]
@@ -986,10 +1080,220 @@ async function hermesControl() {
986
1080
  if ((result.status ?? 1) !== 0) process.exitCode = result.status ?? 1
987
1081
  return
988
1082
  }
989
- console.error("Usage: ocskill hermes <status|prompt|exec> ...")
1083
+ console.error("Usage: ocskill hermes <status|prompt|exec|workflow|exec-workflow> ...")
990
1084
  process.exitCode = 2
991
1085
  }
992
1086
 
1087
+ async function evidenceStoreControl() {
1088
+ const action = args[1] || "status"
1089
+ try {
1090
+ if (action === "status") {
1091
+ printJson(await evidenceStoreStatus(positionalArg(args, 2) || process.cwd()))
1092
+ return
1093
+ }
1094
+ if (action === "put") {
1095
+ const file = args[2]
1096
+ const root = positionalArg(args, 3) || process.cwd()
1097
+ if (!file) throw new Error("Usage: ocskill store put <file> [dir] [--kind <kind>] [--summary <text>]")
1098
+ const content = readTextFile(file)
1099
+ printJson(await putEvidence(root, content, {
1100
+ kind: optionValue(args, "--kind") || "file",
1101
+ source: path.resolve(file),
1102
+ summary: optionValue(args, "--summary"),
1103
+ }))
1104
+ return
1105
+ }
1106
+ if (action === "get") {
1107
+ const ref = args[2]
1108
+ const root = positionalArg(args, 3) || process.cwd()
1109
+ if (!ref) throw new Error("Usage: ocskill store get <evidence-ref> [dir] [--max N] [--start N]")
1110
+ printJson(await getEvidence(root, ref, {
1111
+ maxChars: optionInt(args, "--max", 24_000),
1112
+ start: optionInt(args, "--start", 0),
1113
+ }))
1114
+ return
1115
+ }
1116
+ if (action === "gc") {
1117
+ const root = positionalArg(args, 2) || process.cwd()
1118
+ printJson(await gcEvidenceStore(root, {
1119
+ maxEntries: optionInt(args, "--max-entries", 2000),
1120
+ maxAgeDays: optionInt(args, "--max-age-days", 30),
1121
+ }))
1122
+ return
1123
+ }
1124
+ throw new Error("Usage: ocskill store <status|put|get|gc> ...")
1125
+ } catch (error) {
1126
+ console.error(errorMessage(error))
1127
+ process.exitCode = 1
1128
+ }
1129
+ }
1130
+
1131
+ async function capabilityControl() {
1132
+ const text = args.slice(1).join(" ").trim()
1133
+ if (!text) {
1134
+ console.error("Usage: ocskill capabilities <task text>")
1135
+ process.exitCode = 2
1136
+ return
1137
+ }
1138
+ printJson(inferTaskCapabilities(text))
1139
+ }
1140
+
1141
+ async function visualControl() {
1142
+ const action = args[1]
1143
+ try {
1144
+ if (action === "spec") {
1145
+ const file = args[2]
1146
+ if (!file) throw new Error("Usage: ocskill visual spec <VISUAL_SPEC.json>")
1147
+ printJson(validateVisualSpec(readJsonFile(file)))
1148
+ return
1149
+ }
1150
+ if (action === "geometry") {
1151
+ const specFile = args[2]
1152
+ const actualFile = args[3]
1153
+ if (!specFile || !actualFile) throw new Error("Usage: ocskill visual geometry <VISUAL_SPEC.json> <actual-boxes.json>")
1154
+ printJson(createGeometryReceipt(readJsonFile(specFile), readJsonFile(actualFile)))
1155
+ return
1156
+ }
1157
+ if (action === "compare") {
1158
+ const expected = args[2]
1159
+ const actual = args[3]
1160
+ if (!expected || !actual) throw new Error("Usage: ocskill visual compare <expected.png> <actual.png> [--threshold N] [--max-diff-ratio N]")
1161
+ const threshold = Number(optionValue(args, "--threshold") ?? 16)
1162
+ const maxDiffRatio = Number(optionValue(args, "--max-diff-ratio") ?? 0)
1163
+ printJson(await comparePngFiles(expected, actual, { threshold, maxDiffRatio }))
1164
+ return
1165
+ }
1166
+ if (action === "crop") {
1167
+ const input = args[2]
1168
+ const output = args[3]
1169
+ if (!input || !output) throw new Error("Usage: ocskill visual crop <input.png> <output.png> --x N --y N --width N --height N")
1170
+ printJson(await cropPngFile(input, output, {
1171
+ x: optionInt(args, "--x", 0),
1172
+ y: optionInt(args, "--y", 0),
1173
+ width: optionInt(args, "--width", 1),
1174
+ height: optionInt(args, "--height", 1),
1175
+ }))
1176
+ return
1177
+ }
1178
+ if (action === "viewports") {
1179
+ printJson(responsiveViewportMatrix())
1180
+ return
1181
+ }
1182
+ throw new Error("Usage: ocskill visual <spec|geometry|compare|crop|viewports> ...")
1183
+ } catch (error) {
1184
+ console.error(errorMessage(error))
1185
+ process.exitCode = 1
1186
+ }
1187
+ }
1188
+
1189
+ async function browserControl() {
1190
+ const action = args[1] || "capability"
1191
+ try {
1192
+ if (action === "capability") {
1193
+ printJson(await browserCapability(positionalArg(args, 2) || process.cwd()))
1194
+ return
1195
+ }
1196
+ if (action === "plan") {
1197
+ const url = args[2] || null
1198
+ printJson(buildBrowserVerificationPlan({
1199
+ url,
1200
+ target: optionValue(args, "--target"),
1201
+ }))
1202
+ return
1203
+ }
1204
+ if (action === "inspect") {
1205
+ const url = args[2]
1206
+ if (!url) throw new Error("Usage: ocskill browser inspect <url> [dir] [--selector <css>] [--screenshot <path>] [--width N] [--height N] [--max-elements N]")
1207
+ const root = positionalArg(args, 3) || process.cwd()
1208
+ const report = await inspectBrowserPage(root, url, {
1209
+ selector: optionValue(args, "--selector"),
1210
+ screenshot: optionValue(args, "--screenshot"),
1211
+ width: optionInt(args, "--width", 1440),
1212
+ height: optionInt(args, "--height", 900),
1213
+ maxElements: optionInt(args, "--max-elements", 80),
1214
+ timeoutMs: optionInt(args, "--timeout-ms", 30000),
1215
+ waitMs: optionInt(args, "--wait-ms", 0),
1216
+ fullPage: !args.includes("--viewport-only"),
1217
+ })
1218
+ printJson(args.includes("--full") ? report : summarizeBrowserInspection(report, { limit: optionInt(args, "--limit", 20) }))
1219
+ return
1220
+ }
1221
+ throw new Error("Usage: ocskill browser <capability|plan|inspect> ...")
1222
+ } catch (error) {
1223
+ console.error(errorMessage(error))
1224
+ process.exitCode = 1
1225
+ }
1226
+ }
1227
+
1228
+ async function uiControl() {
1229
+ const action = args[1]
1230
+ try {
1231
+ if (action === "tokens") {
1232
+ const file = args[2]
1233
+ if (!file) throw new Error("Usage: ocskill ui tokens <styles.css>")
1234
+ const tokens = extractDesignTokens(readTextFile(file))
1235
+ printJson({ tokens, evidence: designTokenEvidence(tokens) })
1236
+ return
1237
+ }
1238
+ if (action === "layout") {
1239
+ const file = args[2]
1240
+ if (!file) throw new Error("Usage: ocskill ui layout <boxes.json> --width N --height N [--min-touch N] [--overlap-ratio N]")
1241
+ const payload = readJsonFile(file)
1242
+ const items = Array.isArray(payload) ? payload : payload.elements || payload.boxes || []
1243
+ printJson(inspectResponsiveLayout(items, {
1244
+ width: optionInt(args, "--width", Number(payload.viewport?.width || 0)),
1245
+ height: optionInt(args, "--height", Number(payload.viewport?.height || 0)),
1246
+ }, {
1247
+ minTouchTarget: optionInt(args, "--min-touch", 44),
1248
+ overlapRatio: Number(optionValue(args, "--overlap-ratio") ?? 0.15),
1249
+ }))
1250
+ return
1251
+ }
1252
+ throw new Error("Usage: ocskill ui <tokens|layout> ...")
1253
+ } catch (error) {
1254
+ console.error(errorMessage(error))
1255
+ process.exitCode = 1
1256
+ }
1257
+ }
1258
+
1259
+ async function workflowPlanControl() {
1260
+ const file = args[1]
1261
+ if (!file) {
1262
+ console.error("Usage: ocskill workflow-plan <PLAN.json> [--max-concurrent N]")
1263
+ process.exitCode = 2
1264
+ return
1265
+ }
1266
+ try {
1267
+ const plan = readJsonFile(file)
1268
+ printJson(planDynamicWorkflow(plan.tasks || [], {
1269
+ maxConcurrent: optionInt(args, "--max-concurrent", 4),
1270
+ maxLLMConcurrent: optionInt(args, "--max-llm-concurrent", optionInt(args, "--max-concurrent", 4)),
1271
+ maxVisionConcurrent: optionInt(args, "--max-vision-concurrent", 2),
1272
+ maxWaveCost: optionInt(args, "--max-wave-cost", 24),
1273
+ minAgentCost: optionInt(args, "--min-agent-cost", 5),
1274
+ minVisionAgentCost: optionInt(args, "--min-vision-agent-cost", 4),
1275
+ }))
1276
+ } catch (error) {
1277
+ console.error(errorMessage(error))
1278
+ process.exitCode = 1
1279
+ }
1280
+ }
1281
+
1282
+ async function skillsControl() {
1283
+ const action = args[1] || "lint"
1284
+ if (action !== "lint") {
1285
+ console.error("Usage: ocskill skills lint [dir]")
1286
+ process.exitCode = 2
1287
+ return
1288
+ }
1289
+ try {
1290
+ printJson(await lintSkillCatalog(positionalArg(args, 2) || packageRoot))
1291
+ } catch (error) {
1292
+ console.error(errorMessage(error))
1293
+ process.exitCode = 1
1294
+ }
1295
+ }
1296
+
993
1297
  async function dashboardControl() {
994
1298
  const forwarded = args.slice(1)
995
1299
  const code = run(process.execPath, [path.join(packageRoot, "scripts", "control-center.mjs"), ...forwarded])
@@ -1162,6 +1466,27 @@ switch (command) {
1162
1466
  case "hermes":
1163
1467
  await hermesControl()
1164
1468
  break
1469
+ case "store":
1470
+ await evidenceStoreControl()
1471
+ break
1472
+ case "capabilities":
1473
+ await capabilityControl()
1474
+ break
1475
+ case "visual":
1476
+ await visualControl()
1477
+ break
1478
+ case "browser":
1479
+ await browserControl()
1480
+ break
1481
+ case "workflow-plan":
1482
+ await workflowPlanControl()
1483
+ break
1484
+ case "ui":
1485
+ await uiControl()
1486
+ break
1487
+ case "skills":
1488
+ await skillsControl()
1489
+ break
1165
1490
  case "dashboard":
1166
1491
  await dashboardControl()
1167
1492
  break
@@ -0,0 +1,75 @@
1
+ # V11 Perception & Adaptive Execution
2
+
3
+ ## Goal
4
+
5
+ V11 minimizes context and model cost without deleting evidence needed for correctness. It adds perception-aware UI/browser verification so the system can reason about **what an element is, where it is and how it looks** using different evidence channels.
6
+
7
+ ## Runtime layers
8
+
9
+ ```text
10
+ Task
11
+ -> intent/risk
12
+ -> capability requirements
13
+ -> adaptive evidence budget
14
+ -> semantic/index evidence
15
+ -> content-addressed evidence pointers
16
+ -> focused skills
17
+ -> capability-aware model
18
+ -> fresh executor
19
+ -> deterministic verification
20
+ -> visual/browser verifier when required
21
+ -> diagnosis + evidence expansion only on failure
22
+ ```
23
+
24
+ ## Evidence Store
25
+
26
+ Large raw tool output, durable specs and dependency reports are stored under:
27
+
28
+ ```text
29
+ .ues-cache/evidence-v1/<hash-prefix>/<sha256>.blob
30
+ .ues-cache/evidence-v1/<hash-prefix>/<sha256>.json
31
+ ```
32
+
33
+ A prompt receives a bounded excerpt and a reference such as `evidence:sha256:<hash>`. Use `ocskill store get` to retrieve only the needed slice. `.ues-cache` is runtime state and is excluded from workspace verification fingerprints.
34
+
35
+ ## Adaptive evidence budget
36
+
37
+ FAST/STANDARD/DEEP remain outer safety ceilings. Inside that ceiling V11 allocates characters by evidence role instead of treating all context as equally valuable. Debugging shifts budget toward tests/history; high-risk work shifts toward tests/references; browser/visual work shifts toward deterministic tool evidence.
38
+
39
+ Failure expands evidence through the existing initial -> diagnose -> deep-recovery stages instead of loading maximum context on the first attempt.
40
+
41
+ ## Prompt cache shape
42
+
43
+ Stable material is separated conceptually from dynamic task/evidence. The runtime records stable/dynamic hashes and cacheable ratio. This is telemetry, not a promise that every provider supports prompt caching.
44
+
45
+ ## Capability-aware routing
46
+
47
+ Model profiles may declare coding, reasoning, toolCalling, vision, browser, filesystem, longContext, cost/latency class and a quality hint. UES never invents an unavailable capability. If no configured candidate satisfies a requirement it exposes capability fallback instead of silently claiming that a text-only model can see screenshots.
48
+
49
+ ## Visual fidelity
50
+
51
+ Visual verification uses three complementary layers: semantic DOM/accessibility identity, geometry/bounding boxes, and screenshot pixels. VISUAL_SPEC describes important anchors and tolerances. Geometry receipts prove position/size claims. PNG diff finds changed pixels and their bounding region. A failed region can be cropped so vision only sees the area that needs judgment.
52
+
53
+ Screenshot equality does not prove accessibility or interaction; DOM equality does not prove appearance.
54
+
55
+ ## Browser QA and security
56
+
57
+ Browser workflows prefer bounded deterministic scripts/CLI for ordinary verification. Remote webpage content is untrusted and cannot change UES/tool permissions, request secrets, expand the approved task or authorize external side effects.
58
+
59
+ ## Dynamic workflows
60
+
61
+ The scheduler classifies units as deterministic, LLM judgment or vision judgment. Deterministic work does not spawn agents. Independent tasks may share a wave only when dependencies are ready and file ownership does not conflict. Each wave is integrated and verified before later waves rely on it.
62
+
63
+ ## Skill system
64
+
65
+ V11 keeps progressive disclosure: description metadata for routing, short SKILL.md entrypoint, references only when the selected mode needs them, and deterministic logic in runtime/scripts rather than repeated prompt text.
66
+
67
+ `ocskill skills lint` flags oversized entrypoints and highly overlapping descriptions.
68
+
69
+ ## Hermes sidecar
70
+
71
+ Hermes remains optional. UES owns durable `.ues-work` state, evidence references, task leases/runId fencing, verification receipts and safety/permission boundaries. Hermes may execute a bounded task/workflow when explicitly available but does not become the source of truth.
72
+
73
+ ## Release evidence
74
+
75
+ V11 must pass syntax/resource/router/V11 contract tests, all Node tests, package/install smokes, no-regression live suites, capability-routing tests, visual geometry/pixel fixtures, browser security/targeted-evidence fixtures and token/cache/evidence telemetry benchmarks before stable promotion.