opencode-agent-skill 13.0.0-beta.1 → 14.2.0-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1557 -606
- package/bin/ocskill.mjs +172 -24
- package/docs/OPENCODE-COMPAT.md +34 -95
- package/docs/PI-COMPAT.md +188 -0
- package/docs/V14-CONTEXT-MEMORY-FABRIC.md +70 -0
- package/docs/V14.1-QUALITY-PERFORMANCE-FABRIC.md +114 -0
- package/docs/V14.2-TURBO-WEAK-MODEL-RUNTIME.md +448 -0
- package/evals/v14/tasks.json +46 -0
- package/global-config/agents/executor.md +7 -0
- package/global-config/agents/visual-verifier.md +22 -3
- package/global-config/commands/run.md +27 -20
- package/global-config/plugins/ues-router/command-runtime.js +54 -0
- package/global-config/plugins/ues-router/index.js +30 -24
- package/global-config/plugins/ues-router/policy-runtime.js +7 -0
- package/global-config/plugins/ues-router/router.js +12 -2
- package/global-config/skills/ecommerce-engineering/SKILL.md +1 -1
- package/global-config/skills/file-upload-engineering/SKILL.md +1 -1
- package/global-config/skills/git-safety/SKILL.md +1 -1
- package/global-config/skills/nestjs-engineering/SKILL.md +1 -1
- package/global-config/skills/performance-engineering/SKILL.md +1 -1
- package/global-config/skills/react-native-engineering/SKILL.md +1 -1
- package/global-config/skills/rest-api-design/SKILL.md +1 -1
- package/global-config/skills/ui-ux-engineering/SKILL.md +1 -1
- package/lib/adaptive-context-budget.mjs +97 -0
- package/lib/affected-tests.mjs +260 -0
- package/lib/benchmark-confidence.mjs +41 -2
- package/lib/browser-mcp-routing.mjs +166 -0
- package/lib/capability-fabric.mjs +336 -0
- package/lib/capability-registry.mjs +9 -0
- package/lib/context-engine-v11.mjs +65 -1
- package/lib/context-graph-rank.mjs +118 -0
- package/lib/context-manifest.mjs +97 -18
- package/lib/control-center.mjs +19 -1
- package/lib/dynamic-workflow.mjs +3 -1
- package/lib/evidence-store.mjs +82 -1
- package/lib/hierarchical-context.mjs +215 -0
- package/lib/memory-engine.mjs +465 -0
- package/lib/model-performance.mjs +33 -8
- package/lib/model-policy.mjs +3 -3
- package/lib/orchestrator-policy.mjs +5 -209
- package/lib/performance-fabric.mjs +229 -0
- package/lib/pi-rpc-pool.mjs +433 -0
- package/lib/process-hang-detector.mjs +83 -0
- package/lib/process-supervisor.mjs +193 -0
- package/lib/prompt-cache.mjs +2 -0
- package/lib/repo-graph.mjs +53 -2
- package/lib/runtime-config.mjs +31 -0
- package/lib/safety.mjs +132 -0
- package/lib/semantic-index.mjs +52 -3
- package/lib/skill-compiler.mjs +128 -0
- package/lib/skill-quality.mjs +48 -2
- package/lib/task-engine.mjs +66 -5
- package/lib/task-policy.mjs +235 -0
- package/lib/verification-broker.mjs +284 -0
- package/lib/verification-command.mjs +111 -0
- package/lib/windows-shim.mjs +35 -0
- package/lib/workspace-fingerprint.mjs +198 -0
- package/package.json +52 -42
- package/pi/extensions/ues-child-runtime.ts +238 -0
- package/pi/extensions/ues.ts +3200 -0
- package/pi/prompts/ues-audit.md +9 -0
- package/pi/prompts/ues-critique.md +9 -0
- package/pi/prompts/ues-debug.md +9 -0
- package/pi/prompts/ues-feature.md +9 -0
- package/pi/prompts/ues-fix.md +9 -0
- package/pi/prompts/ues-plan.md +9 -0
- package/pi/prompts/ues-research.md +9 -0
- package/pi/prompts/ues-resume.md +9 -0
- package/pi/prompts/ues-review.md +7 -0
- package/pi/prompts/ues-run.md +17 -0
- package/pi/prompts/ues-verify.md +9 -0
- package/scripts/check-release-consistency.mjs +119 -185
- package/scripts/check-runtime-exports.mjs +66 -0
- package/scripts/check-source-integrity.mjs +184 -0
- package/scripts/eval-pi.mjs +492 -0
- package/scripts/install.mjs +16 -0
- package/scripts/smoke-package-closure.mjs +110 -0
- package/scripts/smoke-packed-install.mjs +24 -11
- package/scripts/smoke-pi-extension.mjs +144 -0
- package/scripts/uninstall.mjs +44 -0
- package/CHANGELOG.md +0 -405
- package/docs/DETERMINISTIC-TOOLS.md +0 -105
- package/docs/ENGINEERING-DESIGN.md +0 -194
- package/docs/EVALS.md +0 -158
- package/docs/GITHUB-RULESET.md +0 -50
- package/docs/NPM-PUBLISH.md +0 -116
- package/docs/RESEARCH-SOURCES.md +0 -37
- package/docs/TRACE-SCHEMA.md +0 -122
- package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +0 -75
- package/docs/V11-PERCEPTION-ADAPTIVE.md +0 -220
- package/docs/V12-WEAK-MODEL-INTELLIGENCE.md +0 -27
- package/docs/V13-PARALLEL-WEAK-MODEL-RUNTIME.md +0 -75
- package/docs/V7-INTELLIGENCE-RUNTIME.md +0 -166
- package/docs/V8-INTELLIGENCE-RELIABILITY.md +0 -206
- package/docs/V9-SPEED-INTELLIGENCE.md +0 -102
|
@@ -1,206 +0,0 @@
|
|
|
1
|
-
# UES 8.0 Intelligence & Reliability
|
|
2
|
-
|
|
3
|
-
UES 8.0 focuses on runtime reliability, stronger evidence gates, context quality, safe parallelism and measurable learning. It does not change the underlying model; it improves how engineering work is selected, executed, verified, recovered and evaluated.
|
|
4
|
-
|
|
5
|
-
## 1. Hard evidence gates
|
|
6
|
-
|
|
7
|
-
Long-horizon and high-risk work uses strict evidence policy.
|
|
8
|
-
|
|
9
|
-
Plan approval requires a structured `plan-verification` receipt bound to the current `PLAN.json` hash:
|
|
10
|
-
|
|
11
|
-
```bash
|
|
12
|
-
ocskill work gate-receipt checkout plan . --verifier ues-plan-checker --evidence "plan checker PASS" --out .ues-work/checkout/reports/plan-receipt.json
|
|
13
|
-
|
|
14
|
-
ocskill work approve-plan checkout . --evidence "plan checker PASS" --receipt-file .ues-work/checkout/reports/plan-receipt.json
|
|
15
|
-
```
|
|
16
|
-
|
|
17
|
-
Task completion requires a successful verification receipt for the active `runId`. In strict mode, the receipt's `workspaceAfter` must also equal the current workspace fingerprint.
|
|
18
|
-
|
|
19
|
-
Integration PASS requires an `integration-verification` receipt bound to the current workspace fingerprint. `finalize` still rejects any later workspace change.
|
|
20
|
-
|
|
21
|
-
## 2. Durable runtime journal
|
|
22
|
-
|
|
23
|
-
Each long work item now includes:
|
|
24
|
-
|
|
25
|
-
```text
|
|
26
|
-
.ues-work/<slug>/
|
|
27
|
-
SPEC.md
|
|
28
|
-
PLAN.json
|
|
29
|
-
STATE.json
|
|
30
|
-
EVIDENCE.json
|
|
31
|
-
EVENTS.jsonl
|
|
32
|
-
tasks/
|
|
33
|
-
reports/
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
`EVENTS.jsonl` is append-only runtime evidence for:
|
|
37
|
-
|
|
38
|
-
- work initialization;
|
|
39
|
-
- plan import and approval;
|
|
40
|
-
- task start/session binding/heartbeat;
|
|
41
|
-
- verification receipts;
|
|
42
|
-
- failure and stale recovery;
|
|
43
|
-
- task completion;
|
|
44
|
-
- integration verification;
|
|
45
|
-
- finalization.
|
|
46
|
-
|
|
47
|
-
Read recent events with:
|
|
48
|
-
|
|
49
|
-
```bash
|
|
50
|
-
ocskill work events <slug> . --limit 100
|
|
51
|
-
```
|
|
52
|
-
|
|
53
|
-
## 3. Bounded executor lifecycle
|
|
54
|
-
|
|
55
|
-
The OpenCode V2 dispatcher probes capabilities instead of assuming them from a version string.
|
|
56
|
-
|
|
57
|
-
A fresh executor has:
|
|
58
|
-
|
|
59
|
-
- a durable `runId`;
|
|
60
|
-
- attached OpenCode session ID;
|
|
61
|
-
- heartbeat and lease expiry;
|
|
62
|
-
- bounded runtime;
|
|
63
|
-
- `session.interrupt` on timeout/cancel;
|
|
64
|
-
- task-scoped stale recovery;
|
|
65
|
-
- process-tree cancellation for external eval processes.
|
|
66
|
-
|
|
67
|
-
On Unix, timed-out external process trees receive SIGTERM followed by SIGKILL after a bounded grace period if necessary. Windows uses `taskkill /T /F`.
|
|
68
|
-
|
|
69
|
-
The V2 plugin exposes task cancellation/recovery tools when the runtime supports session interruption.
|
|
70
|
-
|
|
71
|
-
## 4. Context manifest v3
|
|
72
|
-
|
|
73
|
-
Context selection now combines:
|
|
74
|
-
|
|
75
|
-
- declared task files;
|
|
76
|
-
- local imports and reverse importers;
|
|
77
|
-
- likely related tests;
|
|
78
|
-
- nearby repository instructions/manifests;
|
|
79
|
-
- current Git-changed files;
|
|
80
|
-
- multilingual task terms;
|
|
81
|
-
- symbol hits;
|
|
82
|
-
- TF-IDF-style content relevance;
|
|
83
|
-
- centered source excerpts around matched terms;
|
|
84
|
-
- accepted benchmark-validated lessons.
|
|
85
|
-
|
|
86
|
-
The context remains bounded by a per-task budget rather than dumping the whole repository.
|
|
87
|
-
|
|
88
|
-
## 5. Safer parallel writes
|
|
89
|
-
|
|
90
|
-
Safe-wave analysis still serializes declared read/write conflicts.
|
|
91
|
-
|
|
92
|
-
Writer tasks are isolated in Git worktrees by default when the root checkout is clean (unless isolation is explicitly disabled). This keeps the first writer off the canonical root as well as later concurrent writers. Integration:
|
|
93
|
-
|
|
94
|
-
- captures tracked and untracked sandbox changes;
|
|
95
|
-
- rejects overlap with dirty files in the root checkout;
|
|
96
|
-
- applies the patch to the root only after explicit integration;
|
|
97
|
-
- cleans temporary UES worktree branches;
|
|
98
|
-
- refuses to delete non-UES branches.
|
|
99
|
-
|
|
100
|
-
Manual flow:
|
|
101
|
-
|
|
102
|
-
```bash
|
|
103
|
-
ocskill sandbox create <slug> <task-id> .
|
|
104
|
-
ocskill sandbox integrate <worktree-path> .
|
|
105
|
-
ocskill sandbox list .
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
## 6. Learning v2
|
|
109
|
-
|
|
110
|
-
Evaluation failures are clustered into recurring patterns and candidate rules.
|
|
111
|
-
|
|
112
|
-
```bash
|
|
113
|
-
ocskill learn analyze . --eval-dir .ues-evals
|
|
114
|
-
ocskill learn accept <proposal-id> .
|
|
115
|
-
```
|
|
116
|
-
|
|
117
|
-
Acceptance alone does not make a shadow-required lesson active. Promotion additionally requires measured benchmark improvement:
|
|
118
|
-
|
|
119
|
-
```bash
|
|
120
|
-
ocskill learn promote <proposal-id> . --report .ues-evals/matrix/matrix-summary-<timestamp>.json
|
|
121
|
-
```
|
|
122
|
-
|
|
123
|
-
Promotion reads the benchmark matrix artifact itself, verifies complete/equal baseline-vs-UES coverage, hashes the artifact, and refuses caller-supplied pass-rate claims. Only promoted lessons are eligible for future context retrieval.
|
|
124
|
-
|
|
125
|
-
## 7. Benchmark matrix
|
|
126
|
-
|
|
127
|
-
Run standard, long-horizon and polyglot baseline-vs-UES evaluations:
|
|
128
|
-
|
|
129
|
-
```bash
|
|
130
|
-
npm run evals:matrix -- --model provider/model --trials 3
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
The matrix checks expected baseline/UES coverage before producing a summary. Options include:
|
|
134
|
-
|
|
135
|
-
```text
|
|
136
|
-
--long-only
|
|
137
|
-
--standard-only
|
|
138
|
-
--polyglot-only
|
|
139
|
-
--without-polyglot
|
|
140
|
-
```
|
|
141
|
-
|
|
142
|
-
The polyglot suite adds eight tasks covering Python authorization, Java money validation, .NET authorization, Next.js API error handling, React Native platform logic, safe SQL migration, monorepo dependency boundaries and generated-contract discipline.
|
|
143
|
-
|
|
144
|
-
Benchmark results are evidence for the measured tasks only. They are not evidence that UES converts one model into another model.
|
|
145
|
-
|
|
146
|
-
## 8. Control Center
|
|
147
|
-
|
|
148
|
-
```bash
|
|
149
|
-
ocskill dashboard . --serve --port 4177
|
|
150
|
-
```
|
|
151
|
-
|
|
152
|
-
V8 adds:
|
|
153
|
-
|
|
154
|
-
- runtime event visibility;
|
|
155
|
-
- verification-receipt inspection;
|
|
156
|
-
- stale-task recovery action;
|
|
157
|
-
- existing work/learning/eval summaries.
|
|
158
|
-
|
|
159
|
-
Executor cancellation remains a runtime operation because a standalone dashboard server cannot safely interrupt an OpenCode session it does not own.
|
|
160
|
-
|
|
161
|
-
## 9. Package migration
|
|
162
|
-
|
|
163
|
-
The npm package remains:
|
|
164
|
-
|
|
165
|
-
```text
|
|
166
|
-
opencode-agent-skill
|
|
167
|
-
```
|
|
168
|
-
|
|
169
|
-
Users already on 7.7.0 can update normally:
|
|
170
|
-
|
|
171
|
-
```bash
|
|
172
|
-
ocskill update
|
|
173
|
-
```
|
|
174
|
-
|
|
175
|
-
The installer now writes:
|
|
176
|
-
|
|
177
|
-
```text
|
|
178
|
-
<!-- managed-by: opencode-agent-skill -->
|
|
179
|
-
```
|
|
180
|
-
|
|
181
|
-
It still recognizes the former scoped package owner and marker, then migrates them during re-sync.
|
|
182
|
-
|
|
183
|
-
## 10. Release and supply-chain checks
|
|
184
|
-
|
|
185
|
-
V8 adds:
|
|
186
|
-
|
|
187
|
-
- CodeQL workflow;
|
|
188
|
-
- dependency-review workflow;
|
|
189
|
-
- Dependabot for GitHub Actions and npm;
|
|
190
|
-
- package/tag version consistency guard;
|
|
191
|
-
- tag-only npm publish workflow;
|
|
192
|
-
- OIDC-only npm Trusted Publishing permissions (no long-lived NODE_AUTH_TOKEN);
|
|
193
|
-
- fail-closed tag/version guard;
|
|
194
|
-
- exact plain global-install compatibility smoke in addition to packed-install smoke.
|
|
195
|
-
|
|
196
|
-
Trusted Publishing still requires the npm account-side trust relationship to be configured for `laivannha0202/opencode-agent-skill-` and `publish.yml`.
|
|
197
|
-
|
|
198
|
-
## Validation before release
|
|
199
|
-
|
|
200
|
-
Run:
|
|
201
|
-
|
|
202
|
-
```bash
|
|
203
|
-
npm run ci
|
|
204
|
-
```
|
|
205
|
-
|
|
206
|
-
Then run a real benchmark matrix with the target model. Merge/publish only after local CI is green and benchmark output has been inspected.
|
|
@@ -1,102 +0,0 @@
|
|
|
1
|
-
# UES 9.0 Speed & Intelligence
|
|
2
|
-
|
|
3
|
-
UES 9.0 builds on the V8 evidence and reliability gates. The goal is not to make the underlying model generate tokens faster; it is to reduce time-to-correct-result by finding the right code sooner, shrinking unnecessary context, avoiding unsafe retries and requiring evidence before completion.
|
|
4
|
-
|
|
5
|
-
## Execution profiles
|
|
6
|
-
|
|
7
|
-
UES classifies work into three execution profiles:
|
|
8
|
-
|
|
9
|
-
- **FAST**: small/low-risk work, up to 2 skills, 12k context budget, targeted verification, no durable worktree/critic overhead by default.
|
|
10
|
-
- **STANDARD**: medium work, up to 4 skills, 24k context budget, semantic+Git context and targeted/affected verification.
|
|
11
|
-
- **DEEP**: long-horizon or high-risk work, up to 5 skills, 48k context budget, durable state, worktree isolation, critic/integration verification and full CI.
|
|
12
|
-
|
|
13
|
-
Risk still overrides speed. Authentication, security, payment, schema/migration, production/deploy and public-contract work remains fail-closed.
|
|
14
|
-
|
|
15
|
-
## Incremental evidence index
|
|
16
|
-
|
|
17
|
-
The V9 index caches bounded source metadata under `.ues-cache/semantic-index-v1.json`. Unchanged files reuse cached evidence while changed files are reparsed. It records:
|
|
18
|
-
|
|
19
|
-
- concrete source paths;
|
|
20
|
-
- bounded symbol definitions with line numbers;
|
|
21
|
-
- lexical identifier counts;
|
|
22
|
-
- cache reuse/reparse statistics.
|
|
23
|
-
|
|
24
|
-
This is intentionally labelled **syntax-aware lexical evidence**. It must not be presented as proof of program semantics or as a full AST/LSP call graph.
|
|
25
|
-
|
|
26
|
-
Commands:
|
|
27
|
-
|
|
28
|
-
```bash
|
|
29
|
-
ocskill index status .
|
|
30
|
-
ocskill index build .
|
|
31
|
-
ocskill index rebuild .
|
|
32
|
-
ocskill aci search "checkout total" .
|
|
33
|
-
ocskill aci refs calculateTotal .
|
|
34
|
-
ocskill aci view src/checkout.ts . --line 120 --lines 80
|
|
35
|
-
ocskill aci text "exact marker" .
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
## Weak-model ACI
|
|
39
|
-
|
|
40
|
-
The ACI keeps search/view output bounded and evidence-first. Path traversal outside the repository is rejected. Large/binary files are refused by the bounded viewer. Reference results distinguish lexical references from concrete definition lines so an agent cannot honestly claim deeper semantic certainty than the tool measured.
|
|
41
|
-
|
|
42
|
-
## Runtime traces
|
|
43
|
-
|
|
44
|
-
Operational trace events are appended under `.ues-traces/*.jsonl`. Common credentials/tokens are redacted before persistence, oversized payloads are hashed/truncated, and traces explicitly exclude hidden chain-of-thought.
|
|
45
|
-
|
|
46
|
-
```bash
|
|
47
|
-
ocskill trace show <trace-id> .
|
|
48
|
-
```
|
|
49
|
-
|
|
50
|
-
## Verification sandbox
|
|
51
|
-
|
|
52
|
-
When Docker or Podman is available, deterministic verification commands can run with:
|
|
53
|
-
|
|
54
|
-
- network disabled by default;
|
|
55
|
-
- Linux capabilities dropped;
|
|
56
|
-
- no-new-privileges;
|
|
57
|
-
- read-only container root;
|
|
58
|
-
- bounded PIDs/memory/CPU;
|
|
59
|
-
- only the requested workspace bind-mounted;
|
|
60
|
-
- no host secrets forwarded by default.
|
|
61
|
-
|
|
62
|
-
This protects verification commands; it does **not** claim to sandbox the OpenCode model process itself.
|
|
63
|
-
|
|
64
|
-
## Benchmark confidence gate
|
|
65
|
-
|
|
66
|
-
The matrix now embeds paired baseline/UES evidence and computes:
|
|
67
|
-
|
|
68
|
-
- paired wins/losses/ties;
|
|
69
|
-
- pass-rate delta;
|
|
70
|
-
- exact two-sided sign-test p-value;
|
|
71
|
-
- per-suite regression checks;
|
|
72
|
-
- mean duration ratio;
|
|
73
|
-
- optional cost ratio.
|
|
74
|
-
|
|
75
|
-
Normal reporting:
|
|
76
|
-
|
|
77
|
-
```bash
|
|
78
|
-
npm run evals:matrix -- --model provider/model --trials 3
|
|
79
|
-
```
|
|
80
|
-
|
|
81
|
-
Fail-closed release gate:
|
|
82
|
-
|
|
83
|
-
```bash
|
|
84
|
-
npm run evals:matrix:gate -- --model provider/model --trials 3
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
The gate requires enough paired samples, positive uplift, more UES wins than losses, statistical support, no suite regression and acceptable speed. This is evidence for the measured benchmark only; it is not evidence that UES turns a weaker model into a different model.
|
|
88
|
-
|
|
89
|
-
## Lock and Windows hardening
|
|
90
|
-
|
|
91
|
-
State locking now uses owner tokens and heartbeats. Stale takeover renames the old lock before removal, and release only removes a lock whose token still belongs to the caller. Windows CLI/router invocation resolves Node-backed shims and refuses unrecognized batch shims rather than falling back to shell execution.
|
|
92
|
-
|
|
93
|
-
## Release validation
|
|
94
|
-
|
|
95
|
-
Before publishing 9.0.0:
|
|
96
|
-
|
|
97
|
-
```bash
|
|
98
|
-
npm run ci
|
|
99
|
-
npm run evals:matrix:gate -- --model provider/model --trials 3
|
|
100
|
-
```
|
|
101
|
-
|
|
102
|
-
If GitHub-hosted Actions still fails before the first workflow step, treat that as an infrastructure/repository-action issue rather than test evidence; local CI and the paired benchmark remain required release gates.
|