opencode-agent-skill 7.7.0 → 9.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (63) hide show
  1. package/CHANGELOG.md +63 -3
  2. package/README.md +675 -581
  3. package/bin/ocskill.mjs +358 -156
  4. package/docs/DETERMINISTIC-TOOLS.md +25 -8
  5. package/docs/ENGINEERING-DESIGN.md +31 -13
  6. package/docs/EVALS.md +34 -12
  7. package/docs/NPM-PUBLISH.md +6 -6
  8. package/docs/OPENCODE-COMPAT.md +11 -8
  9. package/docs/TRACE-SCHEMA.md +15 -2
  10. package/docs/V8-INTELLIGENCE-RELIABILITY.md +206 -0
  11. package/docs/V9-SPEED-INTELLIGENCE.md +102 -0
  12. package/evals/live/tasks.json +6 -6
  13. package/evals/polyglot/fixtures/polyglot-bench/api/generated/client.ts +2 -0
  14. package/evals/polyglot/fixtures/polyglot-bench/api/openapi.json +25 -0
  15. package/evals/polyglot/fixtures/polyglot-bench/db/migrations/20260920_add_order_key.sql +1 -0
  16. package/evals/polyglot/fixtures/polyglot-bench/dotnet/OrderService.cs +8 -0
  17. package/evals/polyglot/fixtures/polyglot-bench/java/PriceService.java +5 -0
  18. package/evals/polyglot/fixtures/polyglot-bench/monorepo/package.json +6 -0
  19. package/evals/polyglot/fixtures/polyglot-bench/monorepo/packages/api/package.json +4 -0
  20. package/evals/polyglot/fixtures/polyglot-bench/monorepo/packages/web/package.json +7 -0
  21. package/evals/polyglot/fixtures/polyglot-bench/monorepo/pnpm-lock.yaml +5 -0
  22. package/evals/polyglot/fixtures/polyglot-bench/next/app/api/products/route.ts +7 -0
  23. package/evals/polyglot/fixtures/polyglot-bench/python/tenant_auth.py +4 -0
  24. package/evals/polyglot/fixtures/polyglot-bench/react-native/keyboard.ts +3 -0
  25. package/evals/polyglot/graders/polyglot-bench.mjs +101 -0
  26. package/evals/polyglot/tasks.json +54 -0
  27. package/global-config/AGENTS.md +10 -7
  28. package/global-config/agents/integration-verifier.md +1 -1
  29. package/global-config/agents/plan-checker.md +1 -1
  30. package/global-config/commands/run.md +9 -5
  31. package/global-config/plugins/ues-router/capabilities.js +4 -0
  32. package/global-config/plugins/ues-router/index.js +414 -21
  33. package/global-config/plugins/ues-router/router.js +140 -23
  34. package/global-config/skills/engineering-orchestrator/references/long-horizon.md +6 -4
  35. package/lib/aci.mjs +128 -0
  36. package/lib/benchmark-confidence.mjs +135 -0
  37. package/lib/cli-utils.mjs +41 -0
  38. package/lib/container-sandbox.mjs +102 -0
  39. package/lib/context-manifest.mjs +300 -22
  40. package/lib/control-center.mjs +36 -3
  41. package/lib/eval-order.mjs +9 -0
  42. package/lib/eval-telemetry.mjs +5 -2
  43. package/lib/gate-receipt.mjs +52 -0
  44. package/lib/installer.mjs +39 -25
  45. package/lib/learning-engine.mjs +236 -38
  46. package/lib/opencode-compat.mjs +25 -10
  47. package/lib/orchestrator-policy.mjs +101 -20
  48. package/lib/process-runner.mjs +30 -9
  49. package/lib/runtime-events.mjs +31 -0
  50. package/lib/semantic-index.mjs +318 -0
  51. package/lib/task-engine.mjs +416 -24
  52. package/lib/trajectory.mjs +89 -0
  53. package/lib/windows-shim.mjs +227 -0
  54. package/lib/worktree-sandbox.mjs +85 -3
  55. package/package.json +10 -4
  56. package/scripts/check-release-tag.mjs +22 -0
  57. package/scripts/control-center.mjs +25 -0
  58. package/scripts/eval-live.mjs +39 -58
  59. package/scripts/eval-matrix.mjs +155 -0
  60. package/scripts/smoke-packed-install.mjs +138 -4
  61. package/scripts/smoke-plain-install.mjs +91 -0
  62. package/scripts/validate-live-suite.mjs +3 -3
  63. package/scripts/validate.mjs +27 -5
@@ -1,6 +1,6 @@
1
1
  # Deterministic evidence and execution tools
2
2
 
3
- UES 6 uses dependency-light Node helpers for work that should not rely on a model guessing or remembering it.
3
+ UES 8 uses dependency-light Node helpers for work that should not rely on a model guessing or remembering it.
4
4
 
5
5
  ## Repository evidence
6
6
 
@@ -54,16 +54,20 @@ Validates plan shape/dependencies/cycles and computes topological and safe waves
54
54
  ```cmd
55
55
  ocskill work init <slug> . --goal "..."
56
56
  ocskill work plan <slug> PLAN.json .
57
- ocskill work approve-plan <slug> . --evidence "..."
57
+ ocskill work gate-receipt <slug> plan . --verifier ues-plan-checker --evidence "PASS" --out .ues-work/<slug>/reports/plan-receipt.json
58
+ ocskill work approve-plan <slug> . --evidence "PASS" --receipt-file .ues-work/<slug>/reports/plan-receipt.json
58
59
  ocskill work start <slug> T1 .
59
- ocskill work complete <slug> T1 . --evidence "..."
60
+ ocskill work verify-command <slug> T1 . --run-id <run-id> -- npm test
61
+ ocskill work complete <slug> T1 . --run-id <run-id> --evidence "verified"
60
62
  ocskill work fail <slug> T1 . --reason "..."
61
- ocskill work verify-integration <slug> . --verdict PASS --evidence "..."
62
- ocskill work finalize <slug> . --evidence "..."
63
+ ocskill work gate-receipt <slug> integration . --verifier ues-integration-verifier --verdict PASS --evidence "PASS" --out .ues-work/<slug>/reports/integration-receipt.json
64
+ ocskill work verify-integration <slug> . --verdict PASS --evidence "PASS" --receipt-file .ues-work/<slug>/reports/integration-receipt.json
65
+ ocskill work finalize <slug> . --evidence "final acceptance verified"
66
+ ocskill work events <slug> . --limit 100
63
67
  ocskill work resume <slug> .
64
68
  ```
65
69
 
66
- State/evidence writes use a per-item lock and atomic replacement.
70
+ State/evidence writes use a per-item lock and atomic replacement. `EVENTS.jsonl` is append-only runtime evidence. Long/high-risk tasks require successful receipts for the active run and current workspace fingerprint.
67
71
 
68
72
  ## Context pack
69
73
 
@@ -71,11 +75,24 @@ State/evidence writes use a per-item lock and atomic replacement.
71
75
  ocskill context-pack <slug> <task> .
72
76
  ```
73
77
 
74
- Returns only the task, bounded spec, dependency reports, decisions, blockers and current task state required for a fresh executor.
78
+ Returns the task, bounded spec, dependency reports, decisions, blockers, current task state and Context Manifest v3: declared files, import neighbors, likely tests, nearby instructions, Git-changed files, task-term relevance, symbol hits, centered excerpts and promoted lessons.
75
79
 
76
80
  ## Runtime dispatch on OpenCode V2
77
81
 
78
- The managed plugin exposes `ues.dispatch_task`, which combines `work start`, context pack, model policy and a fresh OpenCode executor session.
82
+ The managed plugin exposes `ues.dispatch_task`, which combines `work start`, context pack, model policy and a fresh OpenCode executor session with heartbeat, bounded wait and interrupt-on-timeout. It also exposes runtime cancellation/recovery helpers when the OpenCode session API supports them.
83
+
84
+ ## Sandboxes and learning
85
+
86
+ ```cmd
87
+ ocskill sandbox create <slug> <task-id> .
88
+ ocskill sandbox integrate <worktree-path> .
89
+ ocskill sandbox list .
90
+ ocskill learn analyze . --eval-dir .ues-evals
91
+ ocskill learn accept <proposal-id> .
92
+ ocskill learn promote <proposal-id> . --baseline 0.50 --candidate 0.75 --samples 4
93
+ ```
94
+
95
+ Sandbox integration refuses overlap with dirty root files. Shadow-required learning proposals are not retrieved until a measured benchmark improvement is recorded.
79
96
 
80
97
  ## Constraints
81
98
 
@@ -1,6 +1,6 @@
1
1
  # UES engineering design
2
2
 
3
- UES 6 evolves the project from an engineering workflow harness into a **long-horizon execution engine** designed to reduce context pressure on coding models.
3
+ UES 8 evolves the project from an engineering workflow harness into a **long-horizon execution engine** designed to reduce context pressure on coding models.
4
4
 
5
5
  The selected model remains the selected model. UES improves orchestration, evidence, task boundaries, state persistence and verification; it does not claim model equivalence.
6
6
 
@@ -48,11 +48,13 @@ The catalog remains 39 skills. UES prefers a small active skill set and loads de
48
48
 
49
49
  ### Hard gates, not reminders
50
50
 
51
- Three V6 gates are machine-enforced:
51
+ V8 machine-enforces the important boundaries:
52
52
 
53
- 1. imported plans are not executable until plan-checker PASS is recorded;
54
- 2. durable mutations are serialized with a per-work-item lock and atomic replacement;
55
- 3. finalization requires integration PASS and an unchanged workspace fingerprint.
53
+ 1. long/high-risk plans are not executable until a structured plan-verification receipt matches the current plan hash;
54
+ 2. long/high-risk task completion requires a successful verification receipt for the active run and the current workspace fingerprint;
55
+ 3. durable state/evidence mutations are serialized with a per-work-item lock and atomic replacement;
56
+ 4. integration PASS for strict work requires a structured integration receipt bound to the current workspace fingerprint;
57
+ 5. finalization requires PASS and rejects any later workspace change.
56
58
 
57
59
  These checks do not depend on a model remembering an instruction.
58
60
 
@@ -64,6 +66,7 @@ These checks do not depend on a model remembering an instruction.
64
66
  PLAN.json
65
67
  STATE.json
66
68
  EVIDENCE.json
69
+ EVENTS.jsonl
67
70
  tasks/
68
71
  reports/
69
72
  ```
@@ -79,13 +82,15 @@ On OpenCode V2, the managed plugin exposes `ues.dispatch_task`.
79
82
  It:
80
83
 
81
84
  1. calls the state engine to start one ready task;
82
- 2. obtains the bounded context pack;
85
+ 2. obtains the bounded Context Manifest v3 pack;
83
86
  3. resolves the configured model tier for the executor attempt;
84
- 4. creates a fresh OpenCode session;
85
- 5. switches to `ues-executor`;
86
- 6. optionally switches to the configured model;
87
- 7. prompts exactly the approved task;
88
- 8. waits and returns child-session context.
87
+ 4. optionally isolates a concurrent writer in a Git worktree;
88
+ 5. creates a fresh OpenCode session rooted at the execution directory;
89
+ 6. binds the session ID to the durable lease;
90
+ 7. switches to `ues-executor` and optionally to the configured model;
91
+ 8. prompts exactly the approved task;
92
+ 9. heartbeats while waiting under a bounded timeout;
93
+ 10. interrupts the child on timeout/cancel and returns a bounded report.
89
94
 
90
95
  The parent is still responsible for inspecting the child diff and recording completion/failure evidence.
91
96
 
@@ -131,7 +136,9 @@ UES separates:
131
136
  1. **static skill contract** — 34 scenarios covering the 39-skill catalog;
132
137
  2. **V2 router precision matrix** — 120 required-route/negative-guard cases;
133
138
  3. **standard live benchmark** — 20 executable hidden-graded tasks;
134
- 4. **long-horizon benchmark** — 5 tasks, including one 15-source-file integration workload.
139
+ 4. **long-horizon benchmark** — 5 tasks, including one 15-source-file integration workload;
140
+ 5. **polyglot benchmark** — 8 tasks spanning Python, Java, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated contracts;
141
+ 6. **matrix runner** — baseline vs UES across multiple suites/trials with coverage validation.
135
142
 
136
143
  For long-suite UES mode, final behavior alone is insufficient. A PASS also requires a completed durable work item with plan approval, at least two tasks, attempted/completed task records, integration PASS and finalization evidence.
137
144
 
@@ -173,4 +180,15 @@ The V2 plugin probes actual session capabilities before dispatch rather than tre
173
180
 
174
181
  Hermes is deliberately adapter-only. UES can detect Hermes and generate a bounded task handoff, but does not embed Hermes' runtime, memory, scheduler or gateway into core.
175
182
 
176
- The local Control Center is observational. It reads durable artifacts and eval summaries; it does not bypass plan, verification, safety or finalization gates.
183
+ ## V8 intelligence and reliability additions
184
+
185
+ V8 adds six control loops around the V7 runtime:
186
+
187
+ 1. **hard evidence binding** — structured plan/integration receipts plus current-workspace task receipts;
188
+ 2. **bounded executor lifecycle** — session binding, timeout interrupt, cancellation and task-scoped recovery;
189
+ 3. **event sourcing for observability** — append-only `EVENTS.jsonl` alongside snapshot state;
190
+ 4. **context manifest v3** — Git-change awareness, symbol hits, TF-IDF-style ranking and centered excerpts;
191
+ 5. **safe parallel integration** — isolated worktrees with dirty-root conflict refusal and UES-only branch cleanup;
192
+ 6. **benchmark-gated learning** — clustered proposals require explicit acceptance and measured shadow improvement before retrieval.
193
+
194
+ The local Control Center remains a safe local observer/controller. It can inspect receipts/events and request stale recovery, but it does not bypass plan, verification, safety or finalization gates and does not directly own OpenCode sessions.
package/docs/EVALS.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # UES evaluations
2
2
 
3
- UES 6 separates catalog correctness, routing precision, benchmark integrity, final behavior and long-horizon orchestration.
3
+ UES 8 separates catalog correctness, routing precision, benchmark integrity, final behavior, long-horizon orchestration and cross-stack coverage.
4
4
 
5
5
  ## 1. Static skill-routing contract
6
6
 
@@ -44,11 +44,22 @@ npm run evals:long:validate
44
44
 
45
45
  The same broken-fixture rule applies.
46
46
 
47
- ## 5. Live baseline vs UES
47
+ ## 5. Polyglot suite integrity
48
+
49
+ The polyglot suite contains **8 tasks** covering Python, Java, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated contract discipline.
50
+
51
+ ```bash
52
+ npm run evals:polyglot:validate
53
+ ```
54
+
55
+ The hidden graders must reject the intentionally broken fixture before the suite is considered valid.
56
+
57
+ ## 6. Live baseline vs UES
48
58
 
49
59
  ```bash
50
60
  ocskill eval-live --model provider/model --trials 3
51
61
  ocskill eval-live --suite long --model provider/model --trials 3
62
+ ocskill eval-live --suite polyglot --model provider/model --trials 3
52
63
  ```
53
64
 
54
65
  Each task runs as:
@@ -66,15 +77,26 @@ For `--suite long`, a UES-mode result is PASS only if:
66
77
  2. hidden behavior grader passes;
67
78
  3. at least one `.ues-work/<slug>/` item is valid;
68
79
  4. the plan contains at least two tasks;
69
- 5. plan approval status is `passed`;
80
+ 5. plan approval status is `passed` and contains a structured plan-verification receipt for the current plan hash;
70
81
  6. every planned task has an attempt and ends `completed`;
71
- 7. integration verification is `PASS`;
72
- 8. integration and finalization evidence exist;
73
- 9. work item status is `completed`.
82
+ 7. every task is backed by a successful verification receipt;
83
+ 8. integration verification is `PASS` and contains a structured integration-verification receipt for the verified workspace fingerprint;
84
+ 9. integration and finalization evidence exist;
85
+ 10. work item status is `completed`.
74
86
 
75
87
  Therefore a model that directly patches all files in its main context but bypasses the long-task engine is not counted as a successful UES long-horizon run.
76
88
 
77
- ## Authentication isolation
89
+ ## 7. Benchmark matrix
90
+
91
+ Run all three behavioral suites in baseline and UES mode:
92
+
93
+ ```bash
94
+ npm run evals:matrix -- --model provider/model --trials 3
95
+ ```
96
+
97
+ The matrix verifies that the expected number of baseline and UES runs was produced before summarizing pass-rate delta. Use `--long-only`, `--standard-only`, `--polyglot-only`, or `--without-polyglot` to narrow the matrix.
98
+
99
+ ## 8. Authentication isolation
78
100
 
79
101
  Default mode:
80
102
 
@@ -92,7 +114,7 @@ ocskill eval-live --model provider/model --auth current --trials 3
92
114
 
93
115
  `current` copies only the current auth file, not the user's global UES configuration.
94
116
 
95
- ## Telemetry
117
+ ## 9. Telemetry
96
118
 
97
119
  Results may include:
98
120
 
@@ -108,7 +130,7 @@ Results may include:
108
130
 
109
131
  No hidden chain-of-thought is collected.
110
132
 
111
- ## Report aggregation
133
+ ## 10. Report aggregation
112
134
 
113
135
  ```bash
114
136
  ocskill eval-report .ues-evals
@@ -119,7 +141,7 @@ Compare the same model, variant, prompt, fixture, grader and environment. Report
119
141
  A benchmark result is evidence only for the measured workload. UES does not claim to turn one base model into another.
120
142
 
121
143
 
122
- ## V7 live-run observability and evidence gate
144
+ ## V8 live-run observability and evidence gate
123
145
 
124
146
  Live runs accept:
125
147
 
@@ -131,6 +153,6 @@ Live runs accept:
131
153
 
132
154
  The harness prints a start line and heartbeat for an active model run. Hard timeout and idle timeout are recorded separately. Ctrl+C aborts the active OpenCode process tree and sets exit code 130 after the current result is recorded.
133
155
 
134
- V7 long-suite UES mode additionally requires **receipt-backed verification for every planned task**. A narrative completion entry is still preserved for backwards compatibility, but it does not satisfy the V7 long benchmark unless at least one successful structured verification receipt is bound to that task.
156
+ V8 long-suite UES mode requires **receipt-backed verification for every planned task**, a structured plan receipt bound to the current plan hash, and a structured integration receipt bound to the verified workspace fingerprint. Strict task completion additionally rejects a successful command receipt if the workspace changed after that receipt.
135
157
 
136
- This intentionally raises the benchmark bar: final code correctness + durable orchestration + machine-observable verification are all required.
158
+ This intentionally raises the benchmark bar: final code correctness + durable orchestration + current machine-observable verification are all required.
@@ -17,7 +17,7 @@ opencode-agent-skill
17
17
  npm run ci
18
18
  ```
19
19
 
20
- V7.7 CI includes syntax validation, resource validation, static skill routing, the 120-case V2 router matrix, standard and long hidden-grader integrity checks, unit/integration tests, package dry-run, and an isolated packed global-install smoke.
20
+ V9 CI includes syntax validation, resource validation, static skill routing, the 120-case V2 router matrix, standard/long/polyglot hidden-grader integrity checks, unit/integration tests, package dry-run, packed global-install smoke, and a plain global-install compatibility smoke.
21
21
 
22
22
  ## Manual release-like test
23
23
 
@@ -27,7 +27,7 @@ Use:
27
27
 
28
28
  ```cmd
29
29
  npm pack
30
- npm install -g .\opencode-agent-skill-7.7.0.tgz --allow-scripts=opencode-agent-skill
30
+ npm install -g .\opencode-agent-skill-9.0.0.tgz --allow-scripts=opencode-agent-skill
31
31
  ocskill status
32
32
  ocskill doctor
33
33
  ```
@@ -47,19 +47,19 @@ After publication verify:
47
47
 
48
48
  ```cmd
49
49
  npm view opencode-agent-skill versions --json
50
- npm view opencode-agent-skill@7.7.0 version
50
+ npm view opencode-agent-skill@9.0.0 version
51
51
  npm dist-tag ls opencode-agent-skill
52
52
  ```
53
53
 
54
54
  The expected release tag is:
55
55
 
56
56
  ```text
57
- latest: 7.7.0
57
+ latest: 9.0.0
58
58
  ```
59
59
 
60
60
  ## GitHub Actions publishing
61
61
 
62
- The repository's publish workflow is release-ready for token-based publishing and provenance. It runs the same package validation before `npm publish`.
62
+ The repository's publish workflow is OIDC/provenance-ready and runs the same package validation before `npm publish`.
63
63
 
64
64
  For stronger long-term supply-chain security, configure npm Trusted Publishing for:
65
65
 
@@ -71,7 +71,7 @@ Workflow: publish.yml
71
71
 
72
72
  Then the GitHub-hosted workflow can authenticate through OIDC instead of a long-lived npm publish token. npm Trusted Publishing requires the corresponding publisher relationship to be configured on npm; repository code alone cannot create that account-side trust relationship.
73
73
 
74
- Until that npm-side setup is complete, keep a valid publish credential configured as `NPM_TOKEN`.
74
+ Until the npm-side Trusted Publisher relationship is configured, the workflow can fall back to a valid `NPM_TOKEN`. After OIDC publishing is verified, remove long-lived publish-token access where practical.
75
75
 
76
76
  ## Release checklist
77
77
 
@@ -1,6 +1,6 @@
1
1
  # OpenCode compatibility
2
2
 
3
- UES 6 ships one npm package for OpenCode 1.x and 2.x, while only enabling V2-native runtime features when V2 is detected.
3
+ UES 8 ships one npm package for OpenCode 1.x and 2.x, while only enabling V2-native runtime features when V2 is detected.
4
4
 
5
5
  ## Detection
6
6
 
@@ -44,11 +44,12 @@ The plugin provides:
44
44
 
45
45
  - prompt-admission skill routing
46
46
  - long-task context guardrails
47
- - permission safety evaluation
48
- - read-only durable-state/task-graph/context-pack tools
49
- - `ues.dispatch_task` fresh executor runtime
47
+ - permission safety evaluation when that hook exists
48
+ - durable-state/task-graph/context-pack tools
49
+ - `ues.dispatch_task` bounded fresh executor runtime
50
+ - `ues.cancel_task` and `ues.recover_task` when session interruption is supported
50
51
 
51
- `ues.dispatch_task` uses V2 session APIs to create a fresh session, select `ues-executor`, optionally switch to the configured model tier, prompt one approved task and wait for completion.
52
+ `ues.dispatch_task` uses V2 session APIs to create a fresh session rooted at the selected execution directory, bind its session ID to the task lease, select `ues-executor`, optionally switch model tier, prompt one approved task, heartbeat while waiting, and interrupt on timeout. Concurrent writing tasks can be isolated in Git worktrees.
52
53
 
53
54
  ## Router control
54
55
 
@@ -92,18 +93,20 @@ Only UES-managed resources are rewritten/removed. Unrelated user plugins/resourc
92
93
  - https://opencode.ai/v2/docs/skills
93
94
 
94
95
 
95
- ## V7 capability probing
96
+ ## V8 capability probing
96
97
 
97
- Version detection remains useful for install-time compatibility, but V7 runtime dispatch does not assume that a major version proves the availability of every session API.
98
+ Version detection remains useful for install-time compatibility, but V8 runtime dispatch does not assume that a major version proves the availability of every session API.
98
99
 
99
100
  The managed V2 plugin probes for:
100
101
 
101
102
  - session creation
102
103
  - prompting
103
104
  - waiting
105
+ - interruption
104
106
  - context retrieval
105
107
  - agent switching
106
108
  - model switching
107
109
  - session hooks
110
+ - permission hooks
108
111
 
109
- `ues.capabilities` exposes the observed surface. `ues.dispatch_task` fails closed when the minimum fresh-dispatch capability set is unavailable instead of attempting a partially supported execution path.
112
+ `ues.capabilities` exposes the observed surface. Fresh dispatch fails closed when the minimum create/prompt/wait/interrupt/context/switch-agent surface is unavailable. Optional context/prompt/permission hooks degrade safely instead of preventing the plugin from loading.
@@ -75,7 +75,7 @@ Keep constant:
75
75
  Compare observable success, regressions, elapsed time, tool behavior and cost rather than narrative confidence.
76
76
 
77
77
 
78
- ## V7 runtime fields
78
+ ## V8 runtime and evidence fields
79
79
 
80
80
  Each live result may additionally contain:
81
81
 
@@ -86,7 +86,7 @@ Each live result may additionally contain:
86
86
  - configured heartbeat/hard/idle timeout values
87
87
  - long-suite receipt coverage inside orchestration inspection
88
88
 
89
- Long-task `EVIDENCE.json` schema 3 may contain a `receipts` array. Receipt fields include:
89
+ Long-task `EVIDENCE.json` schema 3 may contain `receipts` and `gateReceipts` arrays. Command receipt fields include:
90
90
 
91
91
  ```json
92
92
  {
@@ -107,3 +107,16 @@ Long-task `EVIDENCE.json` schema 3 may contain a `receipts` array. Receipt field
107
107
  ```
108
108
 
109
109
  Full stdout/stderr are not stored in receipts; hashes provide binding without persisting potentially sensitive logs.
110
+
111
+
112
+ ### Structured gate receipts
113
+
114
+ Strict plan/integration gates use receipt schema version 1. A plan receipt includes the exact `planHash`; an integration receipt includes the exact `workspaceFingerprint`. Both include verifier identity, optional session/run IDs, evidence text and an optional report hash.
115
+
116
+ ### Runtime event journal
117
+
118
+ Each long work item may include `EVENTS.jsonl`. Every line is an independent JSON event with schema version, UUID, event type, timestamp and task/work metadata. It records operational events only and never hidden chain-of-thought.
119
+
120
+ ### Benchmark matrix summary
121
+
122
+ `scripts/eval-matrix.mjs` writes a matrix summary containing selected suites/model/trials, expected runs per mode, actual baseline/UES counts, coverage completeness and aggregated pass-rate statistics.
@@ -0,0 +1,206 @@
1
+ # UES 8.0 Intelligence & Reliability
2
+
3
+ UES 8.0 focuses on runtime reliability, stronger evidence gates, context quality, safe parallelism and measurable learning. It does not change the underlying model; it improves how engineering work is selected, executed, verified, recovered and evaluated.
4
+
5
+ ## 1. Hard evidence gates
6
+
7
+ Long-horizon and high-risk work uses strict evidence policy.
8
+
9
+ Plan approval requires a structured `plan-verification` receipt bound to the current `PLAN.json` hash:
10
+
11
+ ```bash
12
+ ocskill work gate-receipt checkout plan . --verifier ues-plan-checker --evidence "plan checker PASS" --out .ues-work/checkout/reports/plan-receipt.json
13
+
14
+ ocskill work approve-plan checkout . --evidence "plan checker PASS" --receipt-file .ues-work/checkout/reports/plan-receipt.json
15
+ ```
16
+
17
+ Task completion requires a successful verification receipt for the active `runId`. In strict mode, the receipt's `workspaceAfter` must also equal the current workspace fingerprint.
18
+
19
+ Integration PASS requires an `integration-verification` receipt bound to the current workspace fingerprint. `finalize` still rejects any later workspace change.
20
+
21
+ ## 2. Durable runtime journal
22
+
23
+ Each long work item now includes:
24
+
25
+ ```text
26
+ .ues-work/<slug>/
27
+ SPEC.md
28
+ PLAN.json
29
+ STATE.json
30
+ EVIDENCE.json
31
+ EVENTS.jsonl
32
+ tasks/
33
+ reports/
34
+ ```
35
+
36
+ `EVENTS.jsonl` is append-only runtime evidence for:
37
+
38
+ - work initialization;
39
+ - plan import and approval;
40
+ - task start/session binding/heartbeat;
41
+ - verification receipts;
42
+ - failure and stale recovery;
43
+ - task completion;
44
+ - integration verification;
45
+ - finalization.
46
+
47
+ Read recent events with:
48
+
49
+ ```bash
50
+ ocskill work events <slug> . --limit 100
51
+ ```
52
+
53
+ ## 3. Bounded executor lifecycle
54
+
55
+ The OpenCode V2 dispatcher probes capabilities instead of assuming them from a version string.
56
+
57
+ A fresh executor has:
58
+
59
+ - a durable `runId`;
60
+ - attached OpenCode session ID;
61
+ - heartbeat and lease expiry;
62
+ - bounded runtime;
63
+ - `session.interrupt` on timeout/cancel;
64
+ - task-scoped stale recovery;
65
+ - process-tree cancellation for external eval processes.
66
+
67
+ On Unix, timed-out external process trees receive SIGTERM followed by SIGKILL after a bounded grace period if necessary. Windows uses `taskkill /T /F`.
68
+
69
+ The V2 plugin exposes task cancellation/recovery tools when the runtime supports session interruption.
70
+
71
+ ## 4. Context manifest v3
72
+
73
+ Context selection now combines:
74
+
75
+ - declared task files;
76
+ - local imports and reverse importers;
77
+ - likely related tests;
78
+ - nearby repository instructions/manifests;
79
+ - current Git-changed files;
80
+ - multilingual task terms;
81
+ - symbol hits;
82
+ - TF-IDF-style content relevance;
83
+ - centered source excerpts around matched terms;
84
+ - accepted benchmark-validated lessons.
85
+
86
+ The context remains bounded by a per-task budget rather than dumping the whole repository.
87
+
88
+ ## 5. Safer parallel writes
89
+
90
+ Safe-wave analysis still serializes declared read/write conflicts.
91
+
92
+ Writer tasks are isolated in Git worktrees by default when the root checkout is clean (unless isolation is explicitly disabled). This keeps the first writer off the canonical root as well as later concurrent writers. Integration:
93
+
94
+ - captures tracked and untracked sandbox changes;
95
+ - rejects overlap with dirty files in the root checkout;
96
+ - applies the patch to the root only after explicit integration;
97
+ - cleans temporary UES worktree branches;
98
+ - refuses to delete non-UES branches.
99
+
100
+ Manual flow:
101
+
102
+ ```bash
103
+ ocskill sandbox create <slug> <task-id> .
104
+ ocskill sandbox integrate <worktree-path> .
105
+ ocskill sandbox list .
106
+ ```
107
+
108
+ ## 6. Learning v2
109
+
110
+ Evaluation failures are clustered into recurring patterns and candidate rules.
111
+
112
+ ```bash
113
+ ocskill learn analyze . --eval-dir .ues-evals
114
+ ocskill learn accept <proposal-id> .
115
+ ```
116
+
117
+ Acceptance alone does not make a shadow-required lesson active. Promotion additionally requires measured benchmark improvement:
118
+
119
+ ```bash
120
+ ocskill learn promote <proposal-id> . --report .ues-evals/matrix/matrix-summary-<timestamp>.json
121
+ ```
122
+
123
+ Promotion reads the benchmark matrix artifact itself, verifies complete/equal baseline-vs-UES coverage, hashes the artifact, and refuses caller-supplied pass-rate claims. Only promoted lessons are eligible for future context retrieval.
124
+
125
+ ## 7. Benchmark matrix
126
+
127
+ Run standard, long-horizon and polyglot baseline-vs-UES evaluations:
128
+
129
+ ```bash
130
+ npm run evals:matrix -- --model provider/model --trials 3
131
+ ```
132
+
133
+ The matrix checks expected baseline/UES coverage before producing a summary. Options include:
134
+
135
+ ```text
136
+ --long-only
137
+ --standard-only
138
+ --polyglot-only
139
+ --without-polyglot
140
+ ```
141
+
142
+ The polyglot suite adds eight tasks covering Python authorization, Java money validation, .NET authorization, Next.js API error handling, React Native platform logic, safe SQL migration, monorepo dependency boundaries and generated-contract discipline.
143
+
144
+ Benchmark results are evidence for the measured tasks only. They are not evidence that UES converts one model into another model.
145
+
146
+ ## 8. Control Center
147
+
148
+ ```bash
149
+ ocskill dashboard . --serve --port 4177
150
+ ```
151
+
152
+ V8 adds:
153
+
154
+ - runtime event visibility;
155
+ - verification-receipt inspection;
156
+ - stale-task recovery action;
157
+ - existing work/learning/eval summaries.
158
+
159
+ Executor cancellation remains a runtime operation because a standalone dashboard server cannot safely interrupt an OpenCode session it does not own.
160
+
161
+ ## 9. Package migration
162
+
163
+ The npm package remains:
164
+
165
+ ```text
166
+ opencode-agent-skill
167
+ ```
168
+
169
+ Users already on 7.7.0 can update normally:
170
+
171
+ ```bash
172
+ ocskill update
173
+ ```
174
+
175
+ The installer now writes:
176
+
177
+ ```text
178
+ <!-- managed-by: opencode-agent-skill -->
179
+ ```
180
+
181
+ It still recognizes the former scoped package owner and marker, then migrates them during re-sync.
182
+
183
+ ## 10. Release and supply-chain checks
184
+
185
+ V8 adds:
186
+
187
+ - CodeQL workflow;
188
+ - dependency-review workflow;
189
+ - Dependabot for GitHub Actions and npm;
190
+ - package/tag version consistency guard;
191
+ - tag-only npm publish workflow;
192
+ - OIDC-only npm Trusted Publishing permissions (no long-lived NODE_AUTH_TOKEN);
193
+ - fail-closed tag/version guard;
194
+ - exact plain global-install compatibility smoke in addition to packed-install smoke.
195
+
196
+ Trusted Publishing still requires the npm account-side trust relationship to be configured for `laivannha0202/opencode-agent-skill-` and `publish.yml`.
197
+
198
+ ## Validation before release
199
+
200
+ Run:
201
+
202
+ ```bash
203
+ npm run ci
204
+ ```
205
+
206
+ Then run a real benchmark matrix with the target model. Merge/publish only after local CI is green and benchmark output has been inspected.