opencode-agent-skill 13.0.0-beta.2 → 14.2.0-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (93) hide show
  1. package/README.md +1556 -607
  2. package/bin/ocskill.mjs +172 -24
  3. package/docs/OPENCODE-COMPAT.md +34 -97
  4. package/docs/PI-COMPAT.md +188 -0
  5. package/docs/V14-CONTEXT-MEMORY-FABRIC.md +70 -0
  6. package/docs/V14.1-QUALITY-PERFORMANCE-FABRIC.md +114 -0
  7. package/docs/V14.2-TURBO-WEAK-MODEL-RUNTIME.md +448 -0
  8. package/evals/v14/tasks.json +46 -0
  9. package/global-config/agents/executor.md +7 -0
  10. package/global-config/agents/visual-verifier.md +22 -3
  11. package/global-config/plugins/ues-router/index.js +13 -8
  12. package/global-config/plugins/ues-router/policy-runtime.js +7 -0
  13. package/global-config/plugins/ues-router/router.js +12 -2
  14. package/global-config/skills/ecommerce-engineering/SKILL.md +1 -1
  15. package/global-config/skills/file-upload-engineering/SKILL.md +1 -1
  16. package/global-config/skills/git-safety/SKILL.md +1 -1
  17. package/global-config/skills/nestjs-engineering/SKILL.md +1 -1
  18. package/global-config/skills/performance-engineering/SKILL.md +1 -1
  19. package/global-config/skills/react-native-engineering/SKILL.md +1 -1
  20. package/global-config/skills/rest-api-design/SKILL.md +1 -1
  21. package/global-config/skills/ui-ux-engineering/SKILL.md +1 -1
  22. package/lib/adaptive-context-budget.mjs +97 -0
  23. package/lib/affected-tests.mjs +260 -0
  24. package/lib/benchmark-confidence.mjs +41 -2
  25. package/lib/browser-mcp-routing.mjs +166 -0
  26. package/lib/capability-fabric.mjs +336 -0
  27. package/lib/capability-registry.mjs +9 -0
  28. package/lib/context-engine-v11.mjs +65 -1
  29. package/lib/context-graph-rank.mjs +118 -0
  30. package/lib/context-manifest.mjs +97 -18
  31. package/lib/control-center.mjs +19 -1
  32. package/lib/dynamic-workflow.mjs +3 -1
  33. package/lib/evidence-store.mjs +82 -1
  34. package/lib/hierarchical-context.mjs +215 -0
  35. package/lib/memory-engine.mjs +465 -0
  36. package/lib/model-performance.mjs +33 -8
  37. package/lib/model-policy.mjs +3 -3
  38. package/lib/orchestrator-policy.mjs +5 -209
  39. package/lib/performance-fabric.mjs +229 -0
  40. package/lib/pi-rpc-pool.mjs +433 -0
  41. package/lib/process-hang-detector.mjs +83 -0
  42. package/lib/process-supervisor.mjs +193 -0
  43. package/lib/prompt-cache.mjs +2 -0
  44. package/lib/repo-graph.mjs +53 -2
  45. package/lib/runtime-config.mjs +31 -0
  46. package/lib/safety.mjs +132 -0
  47. package/lib/semantic-index.mjs +52 -3
  48. package/lib/skill-compiler.mjs +128 -0
  49. package/lib/skill-quality.mjs +48 -2
  50. package/lib/task-engine.mjs +66 -5
  51. package/lib/task-policy.mjs +235 -0
  52. package/lib/verification-broker.mjs +284 -0
  53. package/lib/verification-command.mjs +111 -0
  54. package/lib/windows-shim.mjs +35 -0
  55. package/lib/workspace-fingerprint.mjs +198 -0
  56. package/package.json +52 -42
  57. package/pi/extensions/ues-child-runtime.ts +238 -0
  58. package/pi/extensions/ues.ts +3200 -0
  59. package/pi/prompts/ues-audit.md +9 -0
  60. package/pi/prompts/ues-critique.md +9 -0
  61. package/pi/prompts/ues-debug.md +9 -0
  62. package/pi/prompts/ues-feature.md +9 -0
  63. package/pi/prompts/ues-fix.md +9 -0
  64. package/pi/prompts/ues-plan.md +9 -0
  65. package/pi/prompts/ues-research.md +9 -0
  66. package/pi/prompts/ues-resume.md +9 -0
  67. package/pi/prompts/ues-review.md +7 -0
  68. package/pi/prompts/ues-run.md +17 -0
  69. package/pi/prompts/ues-verify.md +9 -0
  70. package/scripts/check-release-consistency.mjs +119 -185
  71. package/scripts/check-runtime-exports.mjs +66 -0
  72. package/scripts/check-source-integrity.mjs +184 -0
  73. package/scripts/eval-pi.mjs +492 -0
  74. package/scripts/install.mjs +16 -0
  75. package/scripts/smoke-package-closure.mjs +110 -0
  76. package/scripts/smoke-packed-install.mjs +24 -11
  77. package/scripts/smoke-pi-extension.mjs +144 -0
  78. package/scripts/uninstall.mjs +44 -0
  79. package/CHANGELOG.md +0 -415
  80. package/docs/DETERMINISTIC-TOOLS.md +0 -105
  81. package/docs/ENGINEERING-DESIGN.md +0 -194
  82. package/docs/EVALS.md +0 -158
  83. package/docs/GITHUB-RULESET.md +0 -50
  84. package/docs/NPM-PUBLISH.md +0 -116
  85. package/docs/RESEARCH-SOURCES.md +0 -37
  86. package/docs/TRACE-SCHEMA.md +0 -122
  87. package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +0 -75
  88. package/docs/V11-PERCEPTION-ADAPTIVE.md +0 -220
  89. package/docs/V12-WEAK-MODEL-INTELLIGENCE.md +0 -27
  90. package/docs/V13-PARALLEL-WEAK-MODEL-RUNTIME.md +0 -86
  91. package/docs/V7-INTELLIGENCE-RUNTIME.md +0 -166
  92. package/docs/V8-INTELLIGENCE-RELIABILITY.md +0 -206
  93. package/docs/V9-SPEED-INTELLIGENCE.md +0 -102
@@ -1,75 +0,0 @@
1
- # V11 Perception & Adaptive Execution
2
-
3
- ## Goal
4
-
5
- V11 minimizes context and model cost without deleting evidence needed for correctness. It adds perception-aware UI/browser verification so the system can reason about **what an element is, where it is and how it looks** using different evidence channels.
6
-
7
- ## Runtime layers
8
-
9
- ```text
10
- Task
11
- -> intent/risk
12
- -> capability requirements
13
- -> adaptive evidence budget
14
- -> semantic/index evidence
15
- -> content-addressed evidence pointers
16
- -> focused skills
17
- -> capability-aware model
18
- -> fresh executor
19
- -> deterministic verification
20
- -> visual/browser verifier when required
21
- -> diagnosis + evidence expansion only on failure
22
- ```
23
-
24
- ## Evidence Store
25
-
26
- Large raw tool output, durable specs and dependency reports are stored under:
27
-
28
- ```text
29
- .ues-cache/evidence-v1/<hash-prefix>/<sha256>.blob
30
- .ues-cache/evidence-v1/<hash-prefix>/<sha256>.json
31
- ```
32
-
33
- A prompt receives a bounded excerpt and a reference such as `evidence:sha256:<hash>`. Use `ocskill store get` to retrieve only the needed slice. `.ues-cache` is runtime state and is excluded from workspace verification fingerprints.
34
-
35
- ## Adaptive evidence budget
36
-
37
- FAST/STANDARD/DEEP remain outer safety ceilings. Inside that ceiling V11 allocates characters by evidence role instead of treating all context as equally valuable. Debugging shifts budget toward tests/history; high-risk work shifts toward tests/references; browser/visual work shifts toward deterministic tool evidence.
38
-
39
- Failure expands evidence through the existing initial -> diagnose -> deep-recovery stages instead of loading maximum context on the first attempt.
40
-
41
- ## Prompt cache shape
42
-
43
- Stable material is separated conceptually from dynamic task/evidence. The runtime records stable/dynamic hashes and cacheable ratio. This is telemetry, not a promise that every provider supports prompt caching.
44
-
45
- ## Capability-aware routing
46
-
47
- Model profiles may declare coding, reasoning, toolCalling, vision, browser, filesystem, longContext, cost/latency class and a quality hint. UES never invents an unavailable capability. If no configured candidate satisfies a requirement it exposes capability fallback instead of silently claiming that a text-only model can see screenshots.
48
-
49
- ## Visual fidelity
50
-
51
- Visual verification uses three complementary layers: semantic DOM/accessibility identity, geometry/bounding boxes, and screenshot pixels. VISUAL_SPEC describes important anchors and tolerances. Geometry receipts prove position/size claims. PNG diff finds changed pixels and their bounding region. A failed region can be cropped so vision only sees the area that needs judgment.
52
-
53
- Screenshot equality does not prove accessibility or interaction; DOM equality does not prove appearance.
54
-
55
- ## Browser QA and security
56
-
57
- Browser workflows prefer bounded deterministic scripts/CLI for ordinary verification. Remote webpage content is untrusted and cannot change UES/tool permissions, request secrets, expand the approved task or authorize external side effects.
58
-
59
- ## Dynamic workflows
60
-
61
- The scheduler classifies units as deterministic, LLM judgment or vision judgment. Deterministic work does not spawn agents. Independent tasks may share a wave only when dependencies are ready and file ownership does not conflict. Each wave is integrated and verified before later waves rely on it.
62
-
63
- ## Skill system
64
-
65
- V11 keeps progressive disclosure: description metadata for routing, short SKILL.md entrypoint, references only when the selected mode needs them, and deterministic logic in runtime/scripts rather than repeated prompt text.
66
-
67
- `ocskill skills lint` flags oversized entrypoints and highly overlapping descriptions.
68
-
69
- ## Hermes sidecar
70
-
71
- Hermes remains optional. UES owns durable `.ues-work` state, evidence references, task leases/runId fencing, verification receipts and safety/permission boundaries. Hermes may execute a bounded task/workflow when explicitly available but does not become the source of truth.
72
-
73
- ## Release evidence
74
-
75
- V11 must pass syntax/resource/router/V11 contract tests, all Node tests, package/install smokes, no-regression live suites, capability-routing tests, visual geometry/pixel fixtures, browser security/targeted-evidence fixtures and token/cache/evidence telemetry benchmarks before stable promotion.
@@ -1,220 +0,0 @@
1
- # UES V11 — Perception & Adaptive Execution
2
-
3
- Status: stable (`11.0.0`). V11 is npm `latest`.
4
-
5
- ## Goal
6
-
7
- V11 extends UES from a reliability-focused coding harness into a perception-aware adaptive execution system. The core rule remains:
8
-
9
- > minimum context necessary for maximum verified task success
10
-
11
- V11 must improve perception, routing and context efficiency without weakening V10 durability, verification or safety.
12
-
13
- ## Runtime layers
14
-
15
- ```text
16
- user task
17
- -> intent/risk classification
18
- -> capability requirements
19
- -> evidence budget
20
- -> semantic/repository evidence
21
- -> content-addressed evidence pointers
22
- -> selective skill/model routing
23
- -> fresh executor
24
- -> deterministic verification
25
- -> visual/browser verifier when required
26
- -> recovery/escalation only from observed failure
27
- ```
28
-
29
- ## Content-addressed Evidence Store
30
-
31
- Large evidence is stored below:
32
-
33
- ```text
34
- .ues-cache/evidence-v1/
35
- ```
36
-
37
- Records are addressed by SHA-256. Executor context receives bounded slices and evidence references instead of repeatedly embedding large tool output.
38
-
39
- Commands:
40
-
41
- ```cmd
42
- ocskill store status .
43
- ocskill store put report.txt . --kind tool-output
44
- ocskill store get evidence:sha256:<hash> . --max 12000
45
- ocskill store gc . --max-entries 2000 --max-age-days 30
46
- ```
47
-
48
- Evidence pointers are not proof by themselves. Any claim still needs the relevant content or a verification receipt.
49
-
50
- ## Adaptive Evidence Budget
51
-
52
- V11 partitions context among instructions, task, declared code, tests, references, history and tools. Debugging, high-risk, visual and browser work can reallocate budget without automatically consuming the full 48k ceiling.
53
-
54
- Failure expands evidence in stages; it does not justify replaying already-proven exploration.
55
-
56
- ## Prompt cache layout
57
-
58
- The prompt envelope is split into a stable prefix and dynamic tail.
59
-
60
- Stable:
61
- - role and invariants
62
- - loaded skill IDs
63
- - stable project facts
64
- - tool policy
65
-
66
- Dynamic:
67
- - current task
68
- - evidence pointers
69
- - recent failure
70
- - next action
71
- - recent messages
72
-
73
- Telemetry records stable/dynamic hashes, sizes and cacheable ratio. This is diagnostic; provider-side cache behavior is provider-dependent.
74
-
75
- ## Capability-aware model routing
76
-
77
- Models may declare:
78
- - coding
79
- - reasoning
80
- - tool calling
81
- - vision
82
- - browser
83
- - filesystem
84
- - long context
85
- - cost class
86
- - latency class
87
- - quality hint
88
-
89
- Example:
90
-
91
- ```cmd
92
- ocskill models capability provider/model --vision on --browser on --reasoning on --quality 0.9
93
- ```
94
-
95
- Task text is converted to capability requirements. The existing light/standard/heavy tiers remain compatible, but V11 can select an eligible configured model instead of assuming every model has the same modalities.
96
-
97
- ## Visual fidelity
98
-
99
- V11 verifies UI through three independent evidence layers:
100
-
101
- 1. semantic/accessibility identity
102
- 2. geometry/bounding boxes
103
- 3. rendered pixels
104
-
105
- A visual task can use `VISUAL_SPEC.json` to define acceptance elements and tolerances.
106
-
107
- ```cmd
108
- ocskill visual spec VISUAL_SPEC.json
109
- ocskill visual geometry VISUAL_SPEC.json actual-boxes.json
110
- ocskill visual compare expected.png actual.png --threshold 16 --max-diff-ratio 0.01
111
- ocskill visual crop actual.png failed-region.png --x 100 --y 200 --width 300 --height 120
112
- ```
113
-
114
- PNG comparison is deterministic and dependency-free for supported 8-bit non-interlaced PNGs. Vision models are used for appearance judgment or cropped failure regions, not for measurements that DOM/geometry can prove exactly.
115
-
116
- ## Browser QA and trust boundary
117
-
118
- Browser workflows are CLI/script-first. Rich browser tooling is used only when persistent exploration or richer introspection is necessary.
119
-
120
- ```cmd
121
- ocskill browser capability .
122
- ocskill browser plan http://localhost:3000 --target Checkout
123
- ```
124
-
125
- Remote webpage text, DOM, ARIA labels and downloaded content are untrusted evidence. They cannot:
126
- - change UES/system policy
127
- - expand permissions or filesystem scope
128
- - request secrets
129
- - authorize publish/deploy/purchases
130
- - weaken verification requirements
131
-
132
- ## Dynamic workflow
133
-
134
- ```cmd
135
- ocskill workflow-plan .ues-work/<slug>/PLAN.json --max-concurrent 4
136
- ```
137
-
138
- The scheduler separates deterministic work from LLM/vision work, respects dependencies and serializes declared write conflicts. Deterministic commands should not consume agent slots. A schedule is planning evidence, not authorization for external side effects.
139
-
140
- ## Hermes sidecar
141
-
142
- Hermes remains optional. UES owns durable state, safety and verification. Hermes can consume one-task or bounded workflow prompts and evidence pointers but does not own merge/push/publish/deploy.
143
-
144
- ```cmd
145
- ocskill hermes status
146
- ocskill hermes workflow <slug> .
147
- ocskill hermes exec-workflow <slug> .
148
- ```
149
-
150
- If Hermes is absent, core UES remains functional.
151
-
152
- ## New skills and agents
153
-
154
- V11 adds nine focused skills:
155
- - visual-fidelity
156
- - browser-qa
157
- - design-source
158
- - responsive-verification
159
- - component-visual-testing
160
- - browser-security
161
- - skill-authoring
162
- - skill-evaluation
163
- - dynamic-workflow
164
-
165
- V11 adds two subagents:
166
- - `ues-visual-verifier`: read-only independent rendered-evidence verification
167
- - `ues-merge-arbiter`: conflict resolution across already-verified task changes
168
-
169
- More agents are not automatically better. New agents require a distinct capability/verification boundary and benchmark evidence.
170
-
171
- ## Skill quality
172
-
173
- ```cmd
174
- ocskill skills lint .
175
- ```
176
-
177
- The linter checks metadata, entrypoint size and likely routing-description collisions. Skill instructions should use progressive disclosure and deterministic scripts for repeatable mechanics.
178
-
179
- ## Evaluation metrics
180
-
181
- V11 keeps correctness first and additionally measures, when available:
182
- - initial and total tokens
183
- - cacheable prompt ratio
184
- - repeated stable input
185
- - evidence-reference reuse
186
- - visual repair attempts
187
- - context expansion count
188
- - model escalation count
189
- - latency/cost/tool calls
190
-
191
- Missing telemetry is `null`, not zero.
192
-
193
- Reference-vs-candidate gates can optionally require V11 telemetry:
194
-
195
- ```cmd
196
- npm run evals:ablation -- v10-summary.json v11-summary.json --require-gate --min-cacheable-ratio 0.70 --min-evidence-reuse-ratio 0.20 --max-repeated-stable-ratio 0.20
197
- ```
198
-
199
- An explicitly requested metric gate fails closed when its telemetry is unavailable.
200
-
201
- ## Release gates
202
-
203
- Do not promote V11 to stable until all are satisfied:
204
-
205
- 1. `npm run ci` passes on the V11 tree.
206
- 2. No correctness/pass-rate regression versus the accepted V10 reference.
207
- 3. Required initial-input/token efficiency gate passes.
208
- 4. Visual geometry and PNG fixtures pass.
209
- 5. Browser trust-boundary and routing tests pass.
210
- 6. Compaction, timeout, provider recovery, lease recovery and loop guards remain green.
211
- 7. Packed and plain npm-install smoke tests pass.
212
- 8. Real weak-model evaluation shows no suite regression.
213
- 9. Any configured cache/evidence target has sufficient telemetry and passes.
214
- 10. npm `latest` is V11 stable; no release gate prevents promotion.
215
-
216
- ## Compatibility
217
-
218
- OpenCode 1.x continues to receive resource/CLI behavior that does not require V2 runtime hooks.
219
-
220
- OpenCode 2.x receives the managed router plugin and fresh-session runtime, including V11 deterministic helper tools for capability inference, evidence retrieval, visual geometry/diff planning, browser planning and dynamic workflow scheduling.
@@ -1,27 +0,0 @@
1
- # V12 Weak-Model Intelligence Foundation
2
-
3
- Status: beta prerelease (`12.0.0-beta.0`, npm dist-tag `next`). V11 (`11.0.0`) remains the stable `latest` release until V12 earns stable-release evidence.
4
-
5
- ## Goal
6
-
7
- V12 focuses on making weaker coding models more reliable on large repositories by improving measured context quality, empirical model routing, plan identity, bounded autonomous decisions, and repo-scale evaluation. It does not claim that orchestration makes one base model equivalent to a stronger model.
8
-
9
- ## Foundations
10
-
11
- - Empirical model performance: observed pass rate, retries, token use and latency can rerank capability-eligible models by task class.
12
- - Context quality receipts: adaptive context reports required-file recall and irrelevant-context ratio.
13
- - Plan-scoped snapshots: every imported plan gets a SHA-256 keyed snapshot and active-plan fence before execution.
14
- - Decision policy: reversible local engineering choices can be auto-resolvable; publish/deploy/destructive/product decisions remain human-gated.
15
- - Repo-scale contract suite: deterministic generation of a 300-module monorepo fixture for larger-repository validation.
16
-
17
- ## Release policy
18
-
19
- V12 is not stable merely because unit tests pass. Promotion requires healthy GitHub CI/Security gates, repo-scale validation, real weak-model baseline-vs-UES trials, no regression in existing suites, measured context recall, sufficient empirical routing samples, and Windows/Linux package/install smoke evidence.
20
-
21
- ## Beta install
22
-
23
- ```cmd
24
- npm install -g opencode-agent-skill@next
25
- ocskill install
26
- ocskill doctor
27
- ```
@@ -1,86 +0,0 @@
1
- # V13 Parallel Weak-Model Runtime
2
-
3
- Status: beta prerelease (`13.0.0-beta.2`).
4
-
5
- ## Goal
6
-
7
- V13 reduces wall-clock time for large engineering work by letting multiple fresh sessions of the **same configured model** execute independent approved tasks concurrently without sharing mutable context.
8
-
9
- ## Runtime contract
10
-
11
- - `ues.dispatch_parallel` defaults to one shared model for every worker and verifier.
12
- - The root must be a Git repository. A pre-existing dirty working tree is allowed: V13 snapshots that baseline into each sandbox and integrates only the task delta, so valid inherited/user changes are preserved.
13
- - Writers execute in isolated Git worktrees.
14
- - Resource leases serialize overlapping files, unknown scope and shared configuration surfaces.
15
- - The scheduler is event-driven: when one task is independently verified, transactionally integrated and completed, newly unblocked dependencies may start immediately.
16
- - Initial and downstream sandboxes inherit the current dirty root through an internal snapshot commit on the sandbox branch; the user's root branch is never auto-committed.
17
- - A fresh `ues-verifier` session using the same model must return PASS before integration.
18
- - Integration is serialized. If receipt/completion fails after patch application, V13 reverses that task patch before marking the run failed.
19
- - No parallel worker pushes, publishes or deploys.
20
-
21
- ## Same-model execution
22
-
23
- One model is sufficient. Parallelism means multiple isolated sessions, not multiple model families:
24
-
25
- ```text
26
- provider/weak-model
27
- ├─ fresh session A
28
- ├─ fresh session B
29
- ├─ fresh session C
30
- └─ fresh verifier sessions
31
- ```
32
-
33
- If no explicit model is supplied, the configured executor model is shared. If no configured model is selected, all sessions keep the OpenCode default model.
34
-
35
- ## Runtime compatibility
36
-
37
- V13 keeps CLI, durable state, receipts, task graphs and Windows text hardening available on OpenCode 1.x. Native `ues.dispatch_task` and `ues.dispatch_parallel` require the V2 router plugin plus the fresh-session capability surface. The router checks capabilities at runtime and fails closed if create/prompt/wait/interrupt/context/switch-agent are incomplete.
38
-
39
- When npm blocks lifecycle scripts (common with stricter npm 11+ `allowScripts` policy), install still leaves the CLI available; run `ocskill install` to perform the documented resource sync explicitly.
40
-
41
- ### Adaptive /ues-run admission on OpenCode V2
42
-
43
- V2 prompt aliases classify the **actual alias arguments** before choosing the prompt envelope. `/ues-run` no longer forces long-horizon mode merely because the alias was used:
44
-
45
- - FAST requests receive a compact prompt that preserves the exact user outcome and skips durable planning unless repository work is actually required.
46
- - STANDARD requests use targeted repository evidence, bounded edits and affected verification without creating durable state by default.
47
- - DEEP / long-horizon / high-risk requests retain the full durable `.ues-work/<slug>/`, plan-gate, fresh-verifier and integration-gate contract.
48
- - `/ues-resume` remains explicitly durable so an existing long-running workspace is not accidentally downgraded.
49
-
50
- If policy classification is unavailable, the router fails conservatively by preserving the full command contract rather than silently weakening verification.
51
-
52
- ## CLI hardening
53
-
54
- V13 parses `--help` before positional arguments, so commands such as these are safe:
55
-
56
- ```cmd
57
- ocskill work init --help
58
- ocskill work status --help
59
- ocskill work gate-receipt --help
60
- ```
61
-
62
- `ocskill work status .` now lists durable workspaces instead of treating `.` as a slug.
63
-
64
- For machine callers, append `--json` to receive structured errors.
65
-
66
- On Windows, prefer:
67
-
68
- ```cmd
69
- ocskill diff . --out dirty.diff
70
- ```
71
-
72
- This writes UTF-8 directly and avoids PowerShell 5 redirection producing UTF-16 text that generic readers may classify as binary. On OpenCode 1.x, where the V2 `ues.text_read` tool is unavailable, read known text safely without creating a converted copy:
73
-
74
- ```cmd
75
- ocskill text-read .ues-work/<slug>/PLAN.json --json
76
- ```
77
-
78
- Existing UTF-8/UTF-16 text can be normalized with:
79
-
80
- ```cmd
81
- ocskill normalize-text dirty.diff
82
- ```
83
-
84
- ## Release evidence
85
-
86
- V13 remains beta until parallel execution demonstrates measurable wall-clock improvement without reducing correctness, and Windows/Linux CI proves help parsing, encoding, dirty-baseline worktree inheritance, rollback, verifier receipts and package installation.
@@ -1,166 +0,0 @@
1
- # UES 7.7 Intelligence Runtime
2
-
3
- UES 7.7 upgrades the V6 long-horizon harness into a more observable, crash-safe and adaptive engineering runtime. The model is still the model; UES improves how work is decomposed, bounded, verified, recovered and learned from.
4
-
5
- ## V7.0 — Runtime reliability
6
-
7
- - live evals use an async process runner with periodic heartbeats
8
- - hard timeout and idle timeout are separate
9
- - Ctrl+C aborts the active OpenCode process tree instead of leaving an orphan
10
- - durable tasks carry `runId`, owner metadata, heartbeat time and lease expiry
11
- - stale running tasks can be recovered with `ocskill work recover`
12
- - `ocskill work resume` performs stale-lease recovery before reporting ready work
13
- - the V2 plugin probes actual session/runtime capabilities before fresh dispatch
14
-
15
- Example:
16
-
17
- ```bash
18
- ocskill work status checkout .
19
- ocskill work recover checkout .
20
- ocskill work heartbeat checkout T1 . --run-id <run-id>
21
- ```
22
-
23
- ## V7.1 — Structured evidence
24
-
25
- `ocskill work verify-command` executes a concrete verification command and records a receipt in `EVIDENCE.json`.
26
-
27
- A receipt contains:
28
-
29
- - command and arguments
30
- - exit code and pass/fail
31
- - start/end/duration
32
- - SHA-256 of stdout/stderr rather than full output
33
- - workspace fingerprint before and after
34
- - task/runId binding
35
-
36
- This complements narrative evidence. Task state reports whether completed work is `receipt-backed` or `narrative`.
37
-
38
- ```bash
39
- ocskill work verify-command checkout T1 . --run-id <run-id> -- npm test
40
- ```
41
-
42
- ## V7.2 — Context intelligence
43
-
44
- Fresh executor context packs now include a bounded context manifest:
45
-
46
- - declared task files
47
- - local import neighbors and reverse importers
48
- - likely related tests
49
- - top-level repository instruction/manifests
50
- - bounded source excerpts
51
- - repository graph hotspots
52
- - accepted learning items relevant to the task
53
-
54
- The goal is to reduce rediscovery cost without dumping the whole repository into one prompt.
55
-
56
- ## V7.3 — Adaptive orchestration/model policy
57
-
58
- `ocskill task-policy <text>` deterministically classifies work by complexity/risk and recommends:
59
-
60
- - `inline`, `standard` or `long-horizon`
61
- - light/standard/heavy model tier
62
- - context budget
63
- - retry budget
64
- - whether plan/integration gates are required
65
-
66
- `ocskill model-policy <role> --attempt N --text <task>` combines role defaults, task risk and retry escalation.
67
-
68
- ## V7.4 — Isolated parallel execution primitives
69
-
70
- The task graph distinguishes reads from writes. By default linked worktrees are created in a sibling `.REPO.ues-sandboxes/` directory on the same drive, avoiding nested worktrees inside the main checkout:
71
-
72
- - read/read overlap can share a safe wave
73
- - write/read and write/write conflicts serialize
74
-
75
- For parallel writers UES provides Git worktree sandboxes:
76
-
77
- ```bash
78
- ocskill sandbox create <slug> <task-id> .
79
- ocskill sandbox list .
80
- ocskill sandbox remove <worktree-path> . --force
81
- ```
82
-
83
- Sandbox creation is deterministic infrastructure. Integration/merging remains an explicit parent responsibility; UES does not silently merge branches.
84
-
85
- ## V7.5 — Evidence-gated learning loop
86
-
87
- UES can mine its own eval traces for recurring failure patterns:
88
-
89
- ```bash
90
- ocskill learn analyze . --eval-dir .ues-evals
91
- ocskill learn status .
92
- ocskill learn accept <proposal-id> .
93
- ```
94
-
95
- Accepted lessons can appear in later context packs when relevant. UES never auto-edits skills or promotes a lesson from a single run without an explicit accept step.
96
-
97
- Current deterministic proposal classes include:
98
-
99
- - hard timeout
100
- - idle timeout
101
- - agent process failure
102
- - hidden grader failure
103
- - orchestration failure
104
- - telemetry parse drift
105
-
106
- ## V7.6 — Optional Hermes bridge
107
-
108
- Hermes is treated as an optional external executor, not embedded as another runtime layer.
109
-
110
- ```bash
111
- ocskill hermes status
112
- ocskill hermes prompt <slug> <task-id> .
113
- ```
114
-
115
- The bridge emits a bounded UES delegation prompt. Hermes must not independently mutate durable UES state, merge, push, publish or deploy.
116
-
117
- ## V7.7 — Control Center
118
-
119
- Generate a zero-dependency local dashboard:
120
-
121
- ```bash
122
- ocskill dashboard .
123
- ocskill dashboard . --serve --port 4177
124
- ```
125
-
126
- When served with `--serve`, the Control Center refreshes its data every few seconds without requiring a frontend build.
127
-
128
- The Control Center summarizes:
129
-
130
- - long-horizon work items and task status
131
- - attempts, blockers and receipt count
132
- - learning proposals/accepted lessons
133
- - recent baseline/UES evaluation summaries
134
-
135
- Generated state lives under `.ues-dashboard/` and is git-ignored.
136
-
137
- ## Live benchmark observability
138
-
139
- The live harness accepts:
140
-
141
- ```bash
142
- node scripts/eval-live.mjs \
143
- --suite long \
144
- --model provider/model \
145
- --mode both \
146
- --trials 3 \
147
- --heartbeat-ms 30000 \
148
- --idle-timeout-ms 300000 \
149
- --timeout-ms 900000
150
- ```
151
-
152
- Every active trial prints a start line and periodic heartbeat rather than appearing frozen.
153
-
154
- ## Safety boundaries
155
-
156
- V7.7 deliberately keeps these actions explicit:
157
-
158
- - merging sandbox branches
159
- - forceful Git/history actions
160
- - publishing/releases
161
- - deployment
162
- - destructive database/filesystem operations
163
- - automatic promotion of learned rules
164
- - automatic execution through Hermes
165
-
166
- The goal is a stronger runtime without removing human control over irreversible side effects.