opencode-agent-skill 7.7.0 → 10.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +112 -3
- package/README.md +396 -281
- package/bin/ocskill.mjs +382 -156
- package/docs/DETERMINISTIC-TOOLS.md +25 -8
- package/docs/ENGINEERING-DESIGN.md +31 -13
- package/docs/EVALS.md +34 -12
- package/docs/NPM-PUBLISH.md +6 -6
- package/docs/OPENCODE-COMPAT.md +11 -8
- package/docs/TRACE-SCHEMA.md +15 -2
- package/docs/V8-INTELLIGENCE-RELIABILITY.md +206 -0
- package/docs/V9-SPEED-INTELLIGENCE.md +102 -0
- package/evals/live/tasks.json +6 -6
- package/evals/polyglot/fixtures/polyglot-bench/api/generated/client.ts +2 -0
- package/evals/polyglot/fixtures/polyglot-bench/api/openapi.json +25 -0
- package/evals/polyglot/fixtures/polyglot-bench/db/migrations/20260920_add_order_key.sql +1 -0
- package/evals/polyglot/fixtures/polyglot-bench/dotnet/OrderService.cs +8 -0
- package/evals/polyglot/fixtures/polyglot-bench/java/PriceService.java +5 -0
- package/evals/polyglot/fixtures/polyglot-bench/monorepo/package.json +6 -0
- package/evals/polyglot/fixtures/polyglot-bench/monorepo/packages/api/package.json +4 -0
- package/evals/polyglot/fixtures/polyglot-bench/monorepo/packages/web/package.json +7 -0
- package/evals/polyglot/fixtures/polyglot-bench/monorepo/pnpm-lock.yaml +5 -0
- package/evals/polyglot/fixtures/polyglot-bench/next/app/api/products/route.ts +7 -0
- package/evals/polyglot/fixtures/polyglot-bench/python/tenant_auth.py +4 -0
- package/evals/polyglot/fixtures/polyglot-bench/react-native/keyboard.ts +3 -0
- package/evals/polyglot/graders/polyglot-bench.mjs +101 -0
- package/evals/polyglot/tasks.json +54 -0
- package/global-config/AGENTS.md +78 -160
- package/global-config/agents/integration-verifier.md +1 -1
- package/global-config/agents/plan-checker.md +1 -1
- package/global-config/commands/run.md +9 -5
- package/global-config/plugins/ues-router/capabilities.js +4 -0
- package/global-config/plugins/ues-router/index.js +784 -37
- package/global-config/plugins/ues-router/router.js +175 -23
- package/global-config/plugins/ues-router/runtime-guard.js +265 -0
- package/global-config/skills/engineering-orchestrator/references/long-horizon.md +6 -4
- package/lib/aci.mjs +128 -0
- package/lib/benchmark-confidence.mjs +173 -0
- package/lib/cli-utils.mjs +41 -0
- package/lib/container-sandbox.mjs +102 -0
- package/lib/context-manifest.mjs +300 -22
- package/lib/control-center.mjs +36 -3
- package/lib/eval-ablation.mjs +104 -0
- package/lib/eval-order.mjs +9 -0
- package/lib/eval-report.mjs +11 -0
- package/lib/eval-telemetry.mjs +8 -2
- package/lib/gate-receipt.mjs +52 -0
- package/lib/installer.mjs +39 -25
- package/lib/learning-engine.mjs +236 -38
- package/lib/model-policy.mjs +6 -0
- package/lib/opencode-compat.mjs +25 -10
- package/lib/orchestrator-policy.mjs +195 -21
- package/lib/process-runner.mjs +30 -9
- package/lib/runtime-events.mjs +31 -0
- package/lib/semantic-index.mjs +318 -0
- package/lib/task-engine.mjs +557 -28
- package/lib/trajectory.mjs +89 -0
- package/lib/windows-shim.mjs +227 -0
- package/lib/worktree-sandbox.mjs +85 -3
- package/package.json +11 -4
- package/scripts/check-release-tag.mjs +22 -0
- package/scripts/control-center.mjs +25 -0
- package/scripts/eval-ablation.mjs +44 -0
- package/scripts/eval-live.mjs +39 -58
- package/scripts/eval-matrix.mjs +166 -0
- package/scripts/smoke-packed-install.mjs +138 -4
- package/scripts/smoke-plain-install.mjs +91 -0
- package/scripts/validate-live-suite.mjs +3 -3
- package/scripts/validate.mjs +27 -5
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Deterministic evidence and execution tools
|
|
2
2
|
|
|
3
|
-
UES
|
|
3
|
+
UES 8 uses dependency-light Node helpers for work that should not rely on a model guessing or remembering it.
|
|
4
4
|
|
|
5
5
|
## Repository evidence
|
|
6
6
|
|
|
@@ -54,16 +54,20 @@ Validates plan shape/dependencies/cycles and computes topological and safe waves
|
|
|
54
54
|
```cmd
|
|
55
55
|
ocskill work init <slug> . --goal "..."
|
|
56
56
|
ocskill work plan <slug> PLAN.json .
|
|
57
|
-
ocskill work
|
|
57
|
+
ocskill work gate-receipt <slug> plan . --verifier ues-plan-checker --evidence "PASS" --out .ues-work/<slug>/reports/plan-receipt.json
|
|
58
|
+
ocskill work approve-plan <slug> . --evidence "PASS" --receipt-file .ues-work/<slug>/reports/plan-receipt.json
|
|
58
59
|
ocskill work start <slug> T1 .
|
|
59
|
-
ocskill work
|
|
60
|
+
ocskill work verify-command <slug> T1 . --run-id <run-id> -- npm test
|
|
61
|
+
ocskill work complete <slug> T1 . --run-id <run-id> --evidence "verified"
|
|
60
62
|
ocskill work fail <slug> T1 . --reason "..."
|
|
61
|
-
ocskill work
|
|
62
|
-
ocskill work
|
|
63
|
+
ocskill work gate-receipt <slug> integration . --verifier ues-integration-verifier --verdict PASS --evidence "PASS" --out .ues-work/<slug>/reports/integration-receipt.json
|
|
64
|
+
ocskill work verify-integration <slug> . --verdict PASS --evidence "PASS" --receipt-file .ues-work/<slug>/reports/integration-receipt.json
|
|
65
|
+
ocskill work finalize <slug> . --evidence "final acceptance verified"
|
|
66
|
+
ocskill work events <slug> . --limit 100
|
|
63
67
|
ocskill work resume <slug> .
|
|
64
68
|
```
|
|
65
69
|
|
|
66
|
-
State/evidence writes use a per-item lock and atomic replacement.
|
|
70
|
+
State/evidence writes use a per-item lock and atomic replacement. `EVENTS.jsonl` is append-only runtime evidence. Long/high-risk tasks require successful receipts for the active run and current workspace fingerprint.
|
|
67
71
|
|
|
68
72
|
## Context pack
|
|
69
73
|
|
|
@@ -71,11 +75,24 @@ State/evidence writes use a per-item lock and atomic replacement.
|
|
|
71
75
|
ocskill context-pack <slug> <task> .
|
|
72
76
|
```
|
|
73
77
|
|
|
74
|
-
Returns
|
|
78
|
+
Returns the task, bounded spec, dependency reports, decisions, blockers, current task state and Context Manifest v3: declared files, import neighbors, likely tests, nearby instructions, Git-changed files, task-term relevance, symbol hits, centered excerpts and promoted lessons.
|
|
75
79
|
|
|
76
80
|
## Runtime dispatch on OpenCode V2
|
|
77
81
|
|
|
78
|
-
The managed plugin exposes `ues.dispatch_task`, which combines `work start`, context pack, model policy and a fresh OpenCode executor session.
|
|
82
|
+
The managed plugin exposes `ues.dispatch_task`, which combines `work start`, context pack, model policy and a fresh OpenCode executor session with heartbeat, bounded wait and interrupt-on-timeout. It also exposes runtime cancellation/recovery helpers when the OpenCode session API supports them.
|
|
83
|
+
|
|
84
|
+
## Sandboxes and learning
|
|
85
|
+
|
|
86
|
+
```cmd
|
|
87
|
+
ocskill sandbox create <slug> <task-id> .
|
|
88
|
+
ocskill sandbox integrate <worktree-path> .
|
|
89
|
+
ocskill sandbox list .
|
|
90
|
+
ocskill learn analyze . --eval-dir .ues-evals
|
|
91
|
+
ocskill learn accept <proposal-id> .
|
|
92
|
+
ocskill learn promote <proposal-id> . --baseline 0.50 --candidate 0.75 --samples 4
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
Sandbox integration refuses overlap with dirty root files. Shadow-required learning proposals are not retrieved until a measured benchmark improvement is recorded.
|
|
79
96
|
|
|
80
97
|
## Constraints
|
|
81
98
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# UES engineering design
|
|
2
2
|
|
|
3
|
-
UES
|
|
3
|
+
UES 8 evolves the project from an engineering workflow harness into a **long-horizon execution engine** designed to reduce context pressure on coding models.
|
|
4
4
|
|
|
5
5
|
The selected model remains the selected model. UES improves orchestration, evidence, task boundaries, state persistence and verification; it does not claim model equivalence.
|
|
6
6
|
|
|
@@ -48,11 +48,13 @@ The catalog remains 39 skills. UES prefers a small active skill set and loads de
|
|
|
48
48
|
|
|
49
49
|
### Hard gates, not reminders
|
|
50
50
|
|
|
51
|
-
|
|
51
|
+
V8 machine-enforces the important boundaries:
|
|
52
52
|
|
|
53
|
-
1.
|
|
54
|
-
2.
|
|
55
|
-
3.
|
|
53
|
+
1. long/high-risk plans are not executable until a structured plan-verification receipt matches the current plan hash;
|
|
54
|
+
2. long/high-risk task completion requires a successful verification receipt for the active run and the current workspace fingerprint;
|
|
55
|
+
3. durable state/evidence mutations are serialized with a per-work-item lock and atomic replacement;
|
|
56
|
+
4. integration PASS for strict work requires a structured integration receipt bound to the current workspace fingerprint;
|
|
57
|
+
5. finalization requires PASS and rejects any later workspace change.
|
|
56
58
|
|
|
57
59
|
These checks do not depend on a model remembering an instruction.
|
|
58
60
|
|
|
@@ -64,6 +66,7 @@ These checks do not depend on a model remembering an instruction.
|
|
|
64
66
|
PLAN.json
|
|
65
67
|
STATE.json
|
|
66
68
|
EVIDENCE.json
|
|
69
|
+
EVENTS.jsonl
|
|
67
70
|
tasks/
|
|
68
71
|
reports/
|
|
69
72
|
```
|
|
@@ -79,13 +82,15 @@ On OpenCode V2, the managed plugin exposes `ues.dispatch_task`.
|
|
|
79
82
|
It:
|
|
80
83
|
|
|
81
84
|
1. calls the state engine to start one ready task;
|
|
82
|
-
2. obtains the bounded
|
|
85
|
+
2. obtains the bounded Context Manifest v3 pack;
|
|
83
86
|
3. resolves the configured model tier for the executor attempt;
|
|
84
|
-
4.
|
|
85
|
-
5.
|
|
86
|
-
6.
|
|
87
|
-
7.
|
|
88
|
-
8.
|
|
87
|
+
4. optionally isolates a concurrent writer in a Git worktree;
|
|
88
|
+
5. creates a fresh OpenCode session rooted at the execution directory;
|
|
89
|
+
6. binds the session ID to the durable lease;
|
|
90
|
+
7. switches to `ues-executor` and optionally to the configured model;
|
|
91
|
+
8. prompts exactly the approved task;
|
|
92
|
+
9. heartbeats while waiting under a bounded timeout;
|
|
93
|
+
10. interrupts the child on timeout/cancel and returns a bounded report.
|
|
89
94
|
|
|
90
95
|
The parent is still responsible for inspecting the child diff and recording completion/failure evidence.
|
|
91
96
|
|
|
@@ -131,7 +136,9 @@ UES separates:
|
|
|
131
136
|
1. **static skill contract** — 34 scenarios covering the 39-skill catalog;
|
|
132
137
|
2. **V2 router precision matrix** — 120 required-route/negative-guard cases;
|
|
133
138
|
3. **standard live benchmark** — 20 executable hidden-graded tasks;
|
|
134
|
-
4. **long-horizon benchmark** — 5 tasks, including one 15-source-file integration workload
|
|
139
|
+
4. **long-horizon benchmark** — 5 tasks, including one 15-source-file integration workload;
|
|
140
|
+
5. **polyglot benchmark** — 8 tasks spanning Python, Java, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated contracts;
|
|
141
|
+
6. **matrix runner** — baseline vs UES across multiple suites/trials with coverage validation.
|
|
135
142
|
|
|
136
143
|
For long-suite UES mode, final behavior alone is insufficient. A PASS also requires a completed durable work item with plan approval, at least two tasks, attempted/completed task records, integration PASS and finalization evidence.
|
|
137
144
|
|
|
@@ -173,4 +180,15 @@ The V2 plugin probes actual session capabilities before dispatch rather than tre
|
|
|
173
180
|
|
|
174
181
|
Hermes is deliberately adapter-only. UES can detect Hermes and generate a bounded task handoff, but does not embed Hermes' runtime, memory, scheduler or gateway into core.
|
|
175
182
|
|
|
176
|
-
|
|
183
|
+
## V8 intelligence and reliability additions
|
|
184
|
+
|
|
185
|
+
V8 adds six control loops around the V7 runtime:
|
|
186
|
+
|
|
187
|
+
1. **hard evidence binding** — structured plan/integration receipts plus current-workspace task receipts;
|
|
188
|
+
2. **bounded executor lifecycle** — session binding, timeout interrupt, cancellation and task-scoped recovery;
|
|
189
|
+
3. **event sourcing for observability** — append-only `EVENTS.jsonl` alongside snapshot state;
|
|
190
|
+
4. **context manifest v3** — Git-change awareness, symbol hits, TF-IDF-style ranking and centered excerpts;
|
|
191
|
+
5. **safe parallel integration** — isolated worktrees with dirty-root conflict refusal and UES-only branch cleanup;
|
|
192
|
+
6. **benchmark-gated learning** — clustered proposals require explicit acceptance and measured shadow improvement before retrieval.
|
|
193
|
+
|
|
194
|
+
The local Control Center remains a safe local observer/controller. It can inspect receipts/events and request stale recovery, but it does not bypass plan, verification, safety or finalization gates and does not directly own OpenCode sessions.
|
package/docs/EVALS.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# UES evaluations
|
|
2
2
|
|
|
3
|
-
UES
|
|
3
|
+
UES 8 separates catalog correctness, routing precision, benchmark integrity, final behavior, long-horizon orchestration and cross-stack coverage.
|
|
4
4
|
|
|
5
5
|
## 1. Static skill-routing contract
|
|
6
6
|
|
|
@@ -44,11 +44,22 @@ npm run evals:long:validate
|
|
|
44
44
|
|
|
45
45
|
The same broken-fixture rule applies.
|
|
46
46
|
|
|
47
|
-
## 5.
|
|
47
|
+
## 5. Polyglot suite integrity
|
|
48
|
+
|
|
49
|
+
The polyglot suite contains **8 tasks** covering Python, Java, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated contract discipline.
|
|
50
|
+
|
|
51
|
+
```bash
|
|
52
|
+
npm run evals:polyglot:validate
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
The hidden graders must reject the intentionally broken fixture before the suite is considered valid.
|
|
56
|
+
|
|
57
|
+
## 6. Live baseline vs UES
|
|
48
58
|
|
|
49
59
|
```bash
|
|
50
60
|
ocskill eval-live --model provider/model --trials 3
|
|
51
61
|
ocskill eval-live --suite long --model provider/model --trials 3
|
|
62
|
+
ocskill eval-live --suite polyglot --model provider/model --trials 3
|
|
52
63
|
```
|
|
53
64
|
|
|
54
65
|
Each task runs as:
|
|
@@ -66,15 +77,26 @@ For `--suite long`, a UES-mode result is PASS only if:
|
|
|
66
77
|
2. hidden behavior grader passes;
|
|
67
78
|
3. at least one `.ues-work/<slug>/` item is valid;
|
|
68
79
|
4. the plan contains at least two tasks;
|
|
69
|
-
5. plan approval status is `passed
|
|
80
|
+
5. plan approval status is `passed` and contains a structured plan-verification receipt for the current plan hash;
|
|
70
81
|
6. every planned task has an attempt and ends `completed`;
|
|
71
|
-
7.
|
|
72
|
-
8. integration and
|
|
73
|
-
9.
|
|
82
|
+
7. every task is backed by a successful verification receipt;
|
|
83
|
+
8. integration verification is `PASS` and contains a structured integration-verification receipt for the verified workspace fingerprint;
|
|
84
|
+
9. integration and finalization evidence exist;
|
|
85
|
+
10. work item status is `completed`.
|
|
74
86
|
|
|
75
87
|
Therefore a model that directly patches all files in its main context but bypasses the long-task engine is not counted as a successful UES long-horizon run.
|
|
76
88
|
|
|
77
|
-
##
|
|
89
|
+
## 7. Benchmark matrix
|
|
90
|
+
|
|
91
|
+
Run all three behavioral suites in baseline and UES mode:
|
|
92
|
+
|
|
93
|
+
```bash
|
|
94
|
+
npm run evals:matrix -- --model provider/model --trials 3
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
The matrix verifies that the expected number of baseline and UES runs was produced before summarizing pass-rate delta. Use `--long-only`, `--standard-only`, `--polyglot-only`, or `--without-polyglot` to narrow the matrix.
|
|
98
|
+
|
|
99
|
+
## 8. Authentication isolation
|
|
78
100
|
|
|
79
101
|
Default mode:
|
|
80
102
|
|
|
@@ -92,7 +114,7 @@ ocskill eval-live --model provider/model --auth current --trials 3
|
|
|
92
114
|
|
|
93
115
|
`current` copies only the current auth file, not the user's global UES configuration.
|
|
94
116
|
|
|
95
|
-
## Telemetry
|
|
117
|
+
## 9. Telemetry
|
|
96
118
|
|
|
97
119
|
Results may include:
|
|
98
120
|
|
|
@@ -108,7 +130,7 @@ Results may include:
|
|
|
108
130
|
|
|
109
131
|
No hidden chain-of-thought is collected.
|
|
110
132
|
|
|
111
|
-
## Report aggregation
|
|
133
|
+
## 10. Report aggregation
|
|
112
134
|
|
|
113
135
|
```bash
|
|
114
136
|
ocskill eval-report .ues-evals
|
|
@@ -119,7 +141,7 @@ Compare the same model, variant, prompt, fixture, grader and environment. Report
|
|
|
119
141
|
A benchmark result is evidence only for the measured workload. UES does not claim to turn one base model into another.
|
|
120
142
|
|
|
121
143
|
|
|
122
|
-
##
|
|
144
|
+
## V8 live-run observability and evidence gate
|
|
123
145
|
|
|
124
146
|
Live runs accept:
|
|
125
147
|
|
|
@@ -131,6 +153,6 @@ Live runs accept:
|
|
|
131
153
|
|
|
132
154
|
The harness prints a start line and heartbeat for an active model run. Hard timeout and idle timeout are recorded separately. Ctrl+C aborts the active OpenCode process tree and sets exit code 130 after the current result is recorded.
|
|
133
155
|
|
|
134
|
-
|
|
156
|
+
V8 long-suite UES mode requires **receipt-backed verification for every planned task**, a structured plan receipt bound to the current plan hash, and a structured integration receipt bound to the verified workspace fingerprint. Strict task completion additionally rejects a successful command receipt if the workspace changed after that receipt.
|
|
135
157
|
|
|
136
|
-
This intentionally raises the benchmark bar: final code correctness + durable orchestration + machine-observable verification are all required.
|
|
158
|
+
This intentionally raises the benchmark bar: final code correctness + durable orchestration + current machine-observable verification are all required.
|
package/docs/NPM-PUBLISH.md
CHANGED
|
@@ -17,7 +17,7 @@ opencode-agent-skill
|
|
|
17
17
|
npm run ci
|
|
18
18
|
```
|
|
19
19
|
|
|
20
|
-
|
|
20
|
+
V9 CI includes syntax validation, resource validation, static skill routing, the 120-case V2 router matrix, standard/long/polyglot hidden-grader integrity checks, unit/integration tests, package dry-run, packed global-install smoke, and a plain global-install compatibility smoke.
|
|
21
21
|
|
|
22
22
|
## Manual release-like test
|
|
23
23
|
|
|
@@ -27,7 +27,7 @@ Use:
|
|
|
27
27
|
|
|
28
28
|
```cmd
|
|
29
29
|
npm pack
|
|
30
|
-
npm install -g .\opencode-agent-skill-
|
|
30
|
+
npm install -g .\opencode-agent-skill-9.0.0.tgz --allow-scripts=opencode-agent-skill
|
|
31
31
|
ocskill status
|
|
32
32
|
ocskill doctor
|
|
33
33
|
```
|
|
@@ -47,19 +47,19 @@ After publication verify:
|
|
|
47
47
|
|
|
48
48
|
```cmd
|
|
49
49
|
npm view opencode-agent-skill versions --json
|
|
50
|
-
npm view opencode-agent-skill@
|
|
50
|
+
npm view opencode-agent-skill@9.0.0 version
|
|
51
51
|
npm dist-tag ls opencode-agent-skill
|
|
52
52
|
```
|
|
53
53
|
|
|
54
54
|
The expected release tag is:
|
|
55
55
|
|
|
56
56
|
```text
|
|
57
|
-
latest:
|
|
57
|
+
latest: 9.0.0
|
|
58
58
|
```
|
|
59
59
|
|
|
60
60
|
## GitHub Actions publishing
|
|
61
61
|
|
|
62
|
-
The repository's publish workflow is
|
|
62
|
+
The repository's publish workflow is OIDC/provenance-ready and runs the same package validation before `npm publish`.
|
|
63
63
|
|
|
64
64
|
For stronger long-term supply-chain security, configure npm Trusted Publishing for:
|
|
65
65
|
|
|
@@ -71,7 +71,7 @@ Workflow: publish.yml
|
|
|
71
71
|
|
|
72
72
|
Then the GitHub-hosted workflow can authenticate through OIDC instead of a long-lived npm publish token. npm Trusted Publishing requires the corresponding publisher relationship to be configured on npm; repository code alone cannot create that account-side trust relationship.
|
|
73
73
|
|
|
74
|
-
Until
|
|
74
|
+
Until the npm-side Trusted Publisher relationship is configured, the workflow can fall back to a valid `NPM_TOKEN`. After OIDC publishing is verified, remove long-lived publish-token access where practical.
|
|
75
75
|
|
|
76
76
|
## Release checklist
|
|
77
77
|
|
package/docs/OPENCODE-COMPAT.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# OpenCode compatibility
|
|
2
2
|
|
|
3
|
-
UES
|
|
3
|
+
UES 8 ships one npm package for OpenCode 1.x and 2.x, while only enabling V2-native runtime features when V2 is detected.
|
|
4
4
|
|
|
5
5
|
## Detection
|
|
6
6
|
|
|
@@ -44,11 +44,12 @@ The plugin provides:
|
|
|
44
44
|
|
|
45
45
|
- prompt-admission skill routing
|
|
46
46
|
- long-task context guardrails
|
|
47
|
-
- permission safety evaluation
|
|
48
|
-
-
|
|
49
|
-
- `ues.dispatch_task` fresh executor runtime
|
|
47
|
+
- permission safety evaluation when that hook exists
|
|
48
|
+
- durable-state/task-graph/context-pack tools
|
|
49
|
+
- `ues.dispatch_task` bounded fresh executor runtime
|
|
50
|
+
- `ues.cancel_task` and `ues.recover_task` when session interruption is supported
|
|
50
51
|
|
|
51
|
-
`ues.dispatch_task` uses V2 session APIs to create a fresh session, select `ues-executor`, optionally switch
|
|
52
|
+
`ues.dispatch_task` uses V2 session APIs to create a fresh session rooted at the selected execution directory, bind its session ID to the task lease, select `ues-executor`, optionally switch model tier, prompt one approved task, heartbeat while waiting, and interrupt on timeout. Concurrent writing tasks can be isolated in Git worktrees.
|
|
52
53
|
|
|
53
54
|
## Router control
|
|
54
55
|
|
|
@@ -92,18 +93,20 @@ Only UES-managed resources are rewritten/removed. Unrelated user plugins/resourc
|
|
|
92
93
|
- https://opencode.ai/v2/docs/skills
|
|
93
94
|
|
|
94
95
|
|
|
95
|
-
##
|
|
96
|
+
## V8 capability probing
|
|
96
97
|
|
|
97
|
-
Version detection remains useful for install-time compatibility, but
|
|
98
|
+
Version detection remains useful for install-time compatibility, but V8 runtime dispatch does not assume that a major version proves the availability of every session API.
|
|
98
99
|
|
|
99
100
|
The managed V2 plugin probes for:
|
|
100
101
|
|
|
101
102
|
- session creation
|
|
102
103
|
- prompting
|
|
103
104
|
- waiting
|
|
105
|
+
- interruption
|
|
104
106
|
- context retrieval
|
|
105
107
|
- agent switching
|
|
106
108
|
- model switching
|
|
107
109
|
- session hooks
|
|
110
|
+
- permission hooks
|
|
108
111
|
|
|
109
|
-
`ues.capabilities` exposes the observed surface.
|
|
112
|
+
`ues.capabilities` exposes the observed surface. Fresh dispatch fails closed when the minimum create/prompt/wait/interrupt/context/switch-agent surface is unavailable. Optional context/prompt/permission hooks degrade safely instead of preventing the plugin from loading.
|
package/docs/TRACE-SCHEMA.md
CHANGED
|
@@ -75,7 +75,7 @@ Keep constant:
|
|
|
75
75
|
Compare observable success, regressions, elapsed time, tool behavior and cost rather than narrative confidence.
|
|
76
76
|
|
|
77
77
|
|
|
78
|
-
##
|
|
78
|
+
## V8 runtime and evidence fields
|
|
79
79
|
|
|
80
80
|
Each live result may additionally contain:
|
|
81
81
|
|
|
@@ -86,7 +86,7 @@ Each live result may additionally contain:
|
|
|
86
86
|
- configured heartbeat/hard/idle timeout values
|
|
87
87
|
- long-suite receipt coverage inside orchestration inspection
|
|
88
88
|
|
|
89
|
-
Long-task `EVIDENCE.json` schema 3 may contain
|
|
89
|
+
Long-task `EVIDENCE.json` schema 3 may contain `receipts` and `gateReceipts` arrays. Command receipt fields include:
|
|
90
90
|
|
|
91
91
|
```json
|
|
92
92
|
{
|
|
@@ -107,3 +107,16 @@ Long-task `EVIDENCE.json` schema 3 may contain a `receipts` array. Receipt field
|
|
|
107
107
|
```
|
|
108
108
|
|
|
109
109
|
Full stdout/stderr are not stored in receipts; hashes provide binding without persisting potentially sensitive logs.
|
|
110
|
+
|
|
111
|
+
|
|
112
|
+
### Structured gate receipts
|
|
113
|
+
|
|
114
|
+
Strict plan/integration gates use receipt schema version 1. A plan receipt includes the exact `planHash`; an integration receipt includes the exact `workspaceFingerprint`. Both include verifier identity, optional session/run IDs, evidence text and an optional report hash.
|
|
115
|
+
|
|
116
|
+
### Runtime event journal
|
|
117
|
+
|
|
118
|
+
Each long work item may include `EVENTS.jsonl`. Every line is an independent JSON event with schema version, UUID, event type, timestamp and task/work metadata. It records operational events only and never hidden chain-of-thought.
|
|
119
|
+
|
|
120
|
+
### Benchmark matrix summary
|
|
121
|
+
|
|
122
|
+
`scripts/eval-matrix.mjs` writes a matrix summary containing selected suites/model/trials, expected runs per mode, actual baseline/UES counts, coverage completeness and aggregated pass-rate statistics.
|
|
@@ -0,0 +1,206 @@
|
|
|
1
|
+
# UES 8.0 Intelligence & Reliability
|
|
2
|
+
|
|
3
|
+
UES 8.0 focuses on runtime reliability, stronger evidence gates, context quality, safe parallelism and measurable learning. It does not change the underlying model; it improves how engineering work is selected, executed, verified, recovered and evaluated.
|
|
4
|
+
|
|
5
|
+
## 1. Hard evidence gates
|
|
6
|
+
|
|
7
|
+
Long-horizon and high-risk work uses strict evidence policy.
|
|
8
|
+
|
|
9
|
+
Plan approval requires a structured `plan-verification` receipt bound to the current `PLAN.json` hash:
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
ocskill work gate-receipt checkout plan . --verifier ues-plan-checker --evidence "plan checker PASS" --out .ues-work/checkout/reports/plan-receipt.json
|
|
13
|
+
|
|
14
|
+
ocskill work approve-plan checkout . --evidence "plan checker PASS" --receipt-file .ues-work/checkout/reports/plan-receipt.json
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Task completion requires a successful verification receipt for the active `runId`. In strict mode, the receipt's `workspaceAfter` must also equal the current workspace fingerprint.
|
|
18
|
+
|
|
19
|
+
Integration PASS requires an `integration-verification` receipt bound to the current workspace fingerprint. `finalize` still rejects any later workspace change.
|
|
20
|
+
|
|
21
|
+
## 2. Durable runtime journal
|
|
22
|
+
|
|
23
|
+
Each long work item now includes:
|
|
24
|
+
|
|
25
|
+
```text
|
|
26
|
+
.ues-work/<slug>/
|
|
27
|
+
SPEC.md
|
|
28
|
+
PLAN.json
|
|
29
|
+
STATE.json
|
|
30
|
+
EVIDENCE.json
|
|
31
|
+
EVENTS.jsonl
|
|
32
|
+
tasks/
|
|
33
|
+
reports/
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
`EVENTS.jsonl` is append-only runtime evidence for:
|
|
37
|
+
|
|
38
|
+
- work initialization;
|
|
39
|
+
- plan import and approval;
|
|
40
|
+
- task start/session binding/heartbeat;
|
|
41
|
+
- verification receipts;
|
|
42
|
+
- failure and stale recovery;
|
|
43
|
+
- task completion;
|
|
44
|
+
- integration verification;
|
|
45
|
+
- finalization.
|
|
46
|
+
|
|
47
|
+
Read recent events with:
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
ocskill work events <slug> . --limit 100
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
## 3. Bounded executor lifecycle
|
|
54
|
+
|
|
55
|
+
The OpenCode V2 dispatcher probes capabilities instead of assuming them from a version string.
|
|
56
|
+
|
|
57
|
+
A fresh executor has:
|
|
58
|
+
|
|
59
|
+
- a durable `runId`;
|
|
60
|
+
- attached OpenCode session ID;
|
|
61
|
+
- heartbeat and lease expiry;
|
|
62
|
+
- bounded runtime;
|
|
63
|
+
- `session.interrupt` on timeout/cancel;
|
|
64
|
+
- task-scoped stale recovery;
|
|
65
|
+
- process-tree cancellation for external eval processes.
|
|
66
|
+
|
|
67
|
+
On Unix, timed-out external process trees receive SIGTERM followed by SIGKILL after a bounded grace period if necessary. Windows uses `taskkill /T /F`.
|
|
68
|
+
|
|
69
|
+
The V2 plugin exposes task cancellation/recovery tools when the runtime supports session interruption.
|
|
70
|
+
|
|
71
|
+
## 4. Context manifest v3
|
|
72
|
+
|
|
73
|
+
Context selection now combines:
|
|
74
|
+
|
|
75
|
+
- declared task files;
|
|
76
|
+
- local imports and reverse importers;
|
|
77
|
+
- likely related tests;
|
|
78
|
+
- nearby repository instructions/manifests;
|
|
79
|
+
- current Git-changed files;
|
|
80
|
+
- multilingual task terms;
|
|
81
|
+
- symbol hits;
|
|
82
|
+
- TF-IDF-style content relevance;
|
|
83
|
+
- centered source excerpts around matched terms;
|
|
84
|
+
- accepted benchmark-validated lessons.
|
|
85
|
+
|
|
86
|
+
The context remains bounded by a per-task budget rather than dumping the whole repository.
|
|
87
|
+
|
|
88
|
+
## 5. Safer parallel writes
|
|
89
|
+
|
|
90
|
+
Safe-wave analysis still serializes declared read/write conflicts.
|
|
91
|
+
|
|
92
|
+
Writer tasks are isolated in Git worktrees by default when the root checkout is clean (unless isolation is explicitly disabled). This keeps the first writer off the canonical root as well as later concurrent writers. Integration:
|
|
93
|
+
|
|
94
|
+
- captures tracked and untracked sandbox changes;
|
|
95
|
+
- rejects overlap with dirty files in the root checkout;
|
|
96
|
+
- applies the patch to the root only after explicit integration;
|
|
97
|
+
- cleans temporary UES worktree branches;
|
|
98
|
+
- refuses to delete non-UES branches.
|
|
99
|
+
|
|
100
|
+
Manual flow:
|
|
101
|
+
|
|
102
|
+
```bash
|
|
103
|
+
ocskill sandbox create <slug> <task-id> .
|
|
104
|
+
ocskill sandbox integrate <worktree-path> .
|
|
105
|
+
ocskill sandbox list .
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
## 6. Learning v2
|
|
109
|
+
|
|
110
|
+
Evaluation failures are clustered into recurring patterns and candidate rules.
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
ocskill learn analyze . --eval-dir .ues-evals
|
|
114
|
+
ocskill learn accept <proposal-id> .
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
Acceptance alone does not make a shadow-required lesson active. Promotion additionally requires measured benchmark improvement:
|
|
118
|
+
|
|
119
|
+
```bash
|
|
120
|
+
ocskill learn promote <proposal-id> . --report .ues-evals/matrix/matrix-summary-<timestamp>.json
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
Promotion reads the benchmark matrix artifact itself, verifies complete/equal baseline-vs-UES coverage, hashes the artifact, and refuses caller-supplied pass-rate claims. Only promoted lessons are eligible for future context retrieval.
|
|
124
|
+
|
|
125
|
+
## 7. Benchmark matrix
|
|
126
|
+
|
|
127
|
+
Run standard, long-horizon and polyglot baseline-vs-UES evaluations:
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
npm run evals:matrix -- --model provider/model --trials 3
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
The matrix checks expected baseline/UES coverage before producing a summary. Options include:
|
|
134
|
+
|
|
135
|
+
```text
|
|
136
|
+
--long-only
|
|
137
|
+
--standard-only
|
|
138
|
+
--polyglot-only
|
|
139
|
+
--without-polyglot
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
The polyglot suite adds eight tasks covering Python authorization, Java money validation, .NET authorization, Next.js API error handling, React Native platform logic, safe SQL migration, monorepo dependency boundaries and generated-contract discipline.
|
|
143
|
+
|
|
144
|
+
Benchmark results are evidence for the measured tasks only. They are not evidence that UES converts one model into another model.
|
|
145
|
+
|
|
146
|
+
## 8. Control Center
|
|
147
|
+
|
|
148
|
+
```bash
|
|
149
|
+
ocskill dashboard . --serve --port 4177
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
V8 adds:
|
|
153
|
+
|
|
154
|
+
- runtime event visibility;
|
|
155
|
+
- verification-receipt inspection;
|
|
156
|
+
- stale-task recovery action;
|
|
157
|
+
- existing work/learning/eval summaries.
|
|
158
|
+
|
|
159
|
+
Executor cancellation remains a runtime operation because a standalone dashboard server cannot safely interrupt an OpenCode session it does not own.
|
|
160
|
+
|
|
161
|
+
## 9. Package migration
|
|
162
|
+
|
|
163
|
+
The npm package remains:
|
|
164
|
+
|
|
165
|
+
```text
|
|
166
|
+
opencode-agent-skill
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Users already on 7.7.0 can update normally:
|
|
170
|
+
|
|
171
|
+
```bash
|
|
172
|
+
ocskill update
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
The installer now writes:
|
|
176
|
+
|
|
177
|
+
```text
|
|
178
|
+
<!-- managed-by: opencode-agent-skill -->
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
It still recognizes the former scoped package owner and marker, then migrates them during re-sync.
|
|
182
|
+
|
|
183
|
+
## 10. Release and supply-chain checks
|
|
184
|
+
|
|
185
|
+
V8 adds:
|
|
186
|
+
|
|
187
|
+
- CodeQL workflow;
|
|
188
|
+
- dependency-review workflow;
|
|
189
|
+
- Dependabot for GitHub Actions and npm;
|
|
190
|
+
- package/tag version consistency guard;
|
|
191
|
+
- tag-only npm publish workflow;
|
|
192
|
+
- OIDC-only npm Trusted Publishing permissions (no long-lived NODE_AUTH_TOKEN);
|
|
193
|
+
- fail-closed tag/version guard;
|
|
194
|
+
- exact plain global-install compatibility smoke in addition to packed-install smoke.
|
|
195
|
+
|
|
196
|
+
Trusted Publishing still requires the npm account-side trust relationship to be configured for `laivannha0202/opencode-agent-skill-` and `publish.yml`.
|
|
197
|
+
|
|
198
|
+
## Validation before release
|
|
199
|
+
|
|
200
|
+
Run:
|
|
201
|
+
|
|
202
|
+
```bash
|
|
203
|
+
npm run ci
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
Then run a real benchmark matrix with the target model. Merge/publish only after local CI is green and benchmark output has been inspected.
|