opencode-agent-skill 13.0.0-beta.1 → 14.2.0-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (95) hide show
  1. package/README.md +1557 -606
  2. package/bin/ocskill.mjs +172 -24
  3. package/docs/OPENCODE-COMPAT.md +34 -95
  4. package/docs/PI-COMPAT.md +188 -0
  5. package/docs/V14-CONTEXT-MEMORY-FABRIC.md +70 -0
  6. package/docs/V14.1-QUALITY-PERFORMANCE-FABRIC.md +114 -0
  7. package/docs/V14.2-TURBO-WEAK-MODEL-RUNTIME.md +448 -0
  8. package/evals/v14/tasks.json +46 -0
  9. package/global-config/agents/executor.md +7 -0
  10. package/global-config/agents/visual-verifier.md +22 -3
  11. package/global-config/commands/run.md +27 -20
  12. package/global-config/plugins/ues-router/command-runtime.js +54 -0
  13. package/global-config/plugins/ues-router/index.js +30 -24
  14. package/global-config/plugins/ues-router/policy-runtime.js +7 -0
  15. package/global-config/plugins/ues-router/router.js +12 -2
  16. package/global-config/skills/ecommerce-engineering/SKILL.md +1 -1
  17. package/global-config/skills/file-upload-engineering/SKILL.md +1 -1
  18. package/global-config/skills/git-safety/SKILL.md +1 -1
  19. package/global-config/skills/nestjs-engineering/SKILL.md +1 -1
  20. package/global-config/skills/performance-engineering/SKILL.md +1 -1
  21. package/global-config/skills/react-native-engineering/SKILL.md +1 -1
  22. package/global-config/skills/rest-api-design/SKILL.md +1 -1
  23. package/global-config/skills/ui-ux-engineering/SKILL.md +1 -1
  24. package/lib/adaptive-context-budget.mjs +97 -0
  25. package/lib/affected-tests.mjs +260 -0
  26. package/lib/benchmark-confidence.mjs +41 -2
  27. package/lib/browser-mcp-routing.mjs +166 -0
  28. package/lib/capability-fabric.mjs +336 -0
  29. package/lib/capability-registry.mjs +9 -0
  30. package/lib/context-engine-v11.mjs +65 -1
  31. package/lib/context-graph-rank.mjs +118 -0
  32. package/lib/context-manifest.mjs +97 -18
  33. package/lib/control-center.mjs +19 -1
  34. package/lib/dynamic-workflow.mjs +3 -1
  35. package/lib/evidence-store.mjs +82 -1
  36. package/lib/hierarchical-context.mjs +215 -0
  37. package/lib/memory-engine.mjs +465 -0
  38. package/lib/model-performance.mjs +33 -8
  39. package/lib/model-policy.mjs +3 -3
  40. package/lib/orchestrator-policy.mjs +5 -209
  41. package/lib/performance-fabric.mjs +229 -0
  42. package/lib/pi-rpc-pool.mjs +433 -0
  43. package/lib/process-hang-detector.mjs +83 -0
  44. package/lib/process-supervisor.mjs +193 -0
  45. package/lib/prompt-cache.mjs +2 -0
  46. package/lib/repo-graph.mjs +53 -2
  47. package/lib/runtime-config.mjs +31 -0
  48. package/lib/safety.mjs +132 -0
  49. package/lib/semantic-index.mjs +52 -3
  50. package/lib/skill-compiler.mjs +128 -0
  51. package/lib/skill-quality.mjs +48 -2
  52. package/lib/task-engine.mjs +66 -5
  53. package/lib/task-policy.mjs +235 -0
  54. package/lib/verification-broker.mjs +284 -0
  55. package/lib/verification-command.mjs +111 -0
  56. package/lib/windows-shim.mjs +35 -0
  57. package/lib/workspace-fingerprint.mjs +198 -0
  58. package/package.json +52 -42
  59. package/pi/extensions/ues-child-runtime.ts +238 -0
  60. package/pi/extensions/ues.ts +3200 -0
  61. package/pi/prompts/ues-audit.md +9 -0
  62. package/pi/prompts/ues-critique.md +9 -0
  63. package/pi/prompts/ues-debug.md +9 -0
  64. package/pi/prompts/ues-feature.md +9 -0
  65. package/pi/prompts/ues-fix.md +9 -0
  66. package/pi/prompts/ues-plan.md +9 -0
  67. package/pi/prompts/ues-research.md +9 -0
  68. package/pi/prompts/ues-resume.md +9 -0
  69. package/pi/prompts/ues-review.md +7 -0
  70. package/pi/prompts/ues-run.md +17 -0
  71. package/pi/prompts/ues-verify.md +9 -0
  72. package/scripts/check-release-consistency.mjs +119 -185
  73. package/scripts/check-runtime-exports.mjs +66 -0
  74. package/scripts/check-source-integrity.mjs +184 -0
  75. package/scripts/eval-pi.mjs +492 -0
  76. package/scripts/install.mjs +16 -0
  77. package/scripts/smoke-package-closure.mjs +110 -0
  78. package/scripts/smoke-packed-install.mjs +24 -11
  79. package/scripts/smoke-pi-extension.mjs +144 -0
  80. package/scripts/uninstall.mjs +44 -0
  81. package/CHANGELOG.md +0 -405
  82. package/docs/DETERMINISTIC-TOOLS.md +0 -105
  83. package/docs/ENGINEERING-DESIGN.md +0 -194
  84. package/docs/EVALS.md +0 -158
  85. package/docs/GITHUB-RULESET.md +0 -50
  86. package/docs/NPM-PUBLISH.md +0 -116
  87. package/docs/RESEARCH-SOURCES.md +0 -37
  88. package/docs/TRACE-SCHEMA.md +0 -122
  89. package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +0 -75
  90. package/docs/V11-PERCEPTION-ADAPTIVE.md +0 -220
  91. package/docs/V12-WEAK-MODEL-INTELLIGENCE.md +0 -27
  92. package/docs/V13-PARALLEL-WEAK-MODEL-RUNTIME.md +0 -75
  93. package/docs/V7-INTELLIGENCE-RUNTIME.md +0 -166
  94. package/docs/V8-INTELLIGENCE-RELIABILITY.md +0 -206
  95. package/docs/V9-SPEED-INTELLIGENCE.md +0 -102
@@ -1,206 +0,0 @@
1
- # UES 8.0 Intelligence & Reliability
2
-
3
- UES 8.0 focuses on runtime reliability, stronger evidence gates, context quality, safe parallelism and measurable learning. It does not change the underlying model; it improves how engineering work is selected, executed, verified, recovered and evaluated.
4
-
5
- ## 1. Hard evidence gates
6
-
7
- Long-horizon and high-risk work uses strict evidence policy.
8
-
9
- Plan approval requires a structured `plan-verification` receipt bound to the current `PLAN.json` hash:
10
-
11
- ```bash
12
- ocskill work gate-receipt checkout plan . --verifier ues-plan-checker --evidence "plan checker PASS" --out .ues-work/checkout/reports/plan-receipt.json
13
-
14
- ocskill work approve-plan checkout . --evidence "plan checker PASS" --receipt-file .ues-work/checkout/reports/plan-receipt.json
15
- ```
16
-
17
- Task completion requires a successful verification receipt for the active `runId`. In strict mode, the receipt's `workspaceAfter` must also equal the current workspace fingerprint.
18
-
19
- Integration PASS requires an `integration-verification` receipt bound to the current workspace fingerprint. `finalize` still rejects any later workspace change.
20
-
21
- ## 2. Durable runtime journal
22
-
23
- Each long work item now includes:
24
-
25
- ```text
26
- .ues-work/<slug>/
27
- SPEC.md
28
- PLAN.json
29
- STATE.json
30
- EVIDENCE.json
31
- EVENTS.jsonl
32
- tasks/
33
- reports/
34
- ```
35
-
36
- `EVENTS.jsonl` is append-only runtime evidence for:
37
-
38
- - work initialization;
39
- - plan import and approval;
40
- - task start/session binding/heartbeat;
41
- - verification receipts;
42
- - failure and stale recovery;
43
- - task completion;
44
- - integration verification;
45
- - finalization.
46
-
47
- Read recent events with:
48
-
49
- ```bash
50
- ocskill work events <slug> . --limit 100
51
- ```
52
-
53
- ## 3. Bounded executor lifecycle
54
-
55
- The OpenCode V2 dispatcher probes capabilities instead of assuming them from a version string.
56
-
57
- A fresh executor has:
58
-
59
- - a durable `runId`;
60
- - attached OpenCode session ID;
61
- - heartbeat and lease expiry;
62
- - bounded runtime;
63
- - `session.interrupt` on timeout/cancel;
64
- - task-scoped stale recovery;
65
- - process-tree cancellation for external eval processes.
66
-
67
- On Unix, timed-out external process trees receive SIGTERM followed by SIGKILL after a bounded grace period if necessary. Windows uses `taskkill /T /F`.
68
-
69
- The V2 plugin exposes task cancellation/recovery tools when the runtime supports session interruption.
70
-
71
- ## 4. Context manifest v3
72
-
73
- Context selection now combines:
74
-
75
- - declared task files;
76
- - local imports and reverse importers;
77
- - likely related tests;
78
- - nearby repository instructions/manifests;
79
- - current Git-changed files;
80
- - multilingual task terms;
81
- - symbol hits;
82
- - TF-IDF-style content relevance;
83
- - centered source excerpts around matched terms;
84
- - accepted benchmark-validated lessons.
85
-
86
- The context remains bounded by a per-task budget rather than dumping the whole repository.
87
-
88
- ## 5. Safer parallel writes
89
-
90
- Safe-wave analysis still serializes declared read/write conflicts.
91
-
92
- Writer tasks are isolated in Git worktrees by default when the root checkout is clean (unless isolation is explicitly disabled). This keeps the first writer off the canonical root as well as later concurrent writers. Integration:
93
-
94
- - captures tracked and untracked sandbox changes;
95
- - rejects overlap with dirty files in the root checkout;
96
- - applies the patch to the root only after explicit integration;
97
- - cleans temporary UES worktree branches;
98
- - refuses to delete non-UES branches.
99
-
100
- Manual flow:
101
-
102
- ```bash
103
- ocskill sandbox create <slug> <task-id> .
104
- ocskill sandbox integrate <worktree-path> .
105
- ocskill sandbox list .
106
- ```
107
-
108
- ## 6. Learning v2
109
-
110
- Evaluation failures are clustered into recurring patterns and candidate rules.
111
-
112
- ```bash
113
- ocskill learn analyze . --eval-dir .ues-evals
114
- ocskill learn accept <proposal-id> .
115
- ```
116
-
117
- Acceptance alone does not make a shadow-required lesson active. Promotion additionally requires measured benchmark improvement:
118
-
119
- ```bash
120
- ocskill learn promote <proposal-id> . --report .ues-evals/matrix/matrix-summary-<timestamp>.json
121
- ```
122
-
123
- Promotion reads the benchmark matrix artifact itself, verifies complete/equal baseline-vs-UES coverage, hashes the artifact, and refuses caller-supplied pass-rate claims. Only promoted lessons are eligible for future context retrieval.
124
-
125
- ## 7. Benchmark matrix
126
-
127
- Run standard, long-horizon and polyglot baseline-vs-UES evaluations:
128
-
129
- ```bash
130
- npm run evals:matrix -- --model provider/model --trials 3
131
- ```
132
-
133
- The matrix checks expected baseline/UES coverage before producing a summary. Options include:
134
-
135
- ```text
136
- --long-only
137
- --standard-only
138
- --polyglot-only
139
- --without-polyglot
140
- ```
141
-
142
- The polyglot suite adds eight tasks covering Python authorization, Java money validation, .NET authorization, Next.js API error handling, React Native platform logic, safe SQL migration, monorepo dependency boundaries and generated-contract discipline.
143
-
144
- Benchmark results are evidence for the measured tasks only. They are not evidence that UES converts one model into another model.
145
-
146
- ## 8. Control Center
147
-
148
- ```bash
149
- ocskill dashboard . --serve --port 4177
150
- ```
151
-
152
- V8 adds:
153
-
154
- - runtime event visibility;
155
- - verification-receipt inspection;
156
- - stale-task recovery action;
157
- - existing work/learning/eval summaries.
158
-
159
- Executor cancellation remains a runtime operation because a standalone dashboard server cannot safely interrupt an OpenCode session it does not own.
160
-
161
- ## 9. Package migration
162
-
163
- The npm package remains:
164
-
165
- ```text
166
- opencode-agent-skill
167
- ```
168
-
169
- Users already on 7.7.0 can update normally:
170
-
171
- ```bash
172
- ocskill update
173
- ```
174
-
175
- The installer now writes:
176
-
177
- ```text
178
- <!-- managed-by: opencode-agent-skill -->
179
- ```
180
-
181
- It still recognizes the former scoped package owner and marker, then migrates them during re-sync.
182
-
183
- ## 10. Release and supply-chain checks
184
-
185
- V8 adds:
186
-
187
- - CodeQL workflow;
188
- - dependency-review workflow;
189
- - Dependabot for GitHub Actions and npm;
190
- - package/tag version consistency guard;
191
- - tag-only npm publish workflow;
192
- - OIDC-only npm Trusted Publishing permissions (no long-lived NODE_AUTH_TOKEN);
193
- - fail-closed tag/version guard;
194
- - exact plain global-install compatibility smoke in addition to packed-install smoke.
195
-
196
- Trusted Publishing still requires the npm account-side trust relationship to be configured for `laivannha0202/opencode-agent-skill-` and `publish.yml`.
197
-
198
- ## Validation before release
199
-
200
- Run:
201
-
202
- ```bash
203
- npm run ci
204
- ```
205
-
206
- Then run a real benchmark matrix with the target model. Merge/publish only after local CI is green and benchmark output has been inspected.
@@ -1,102 +0,0 @@
1
- # UES 9.0 Speed & Intelligence
2
-
3
- UES 9.0 builds on the V8 evidence and reliability gates. The goal is not to make the underlying model generate tokens faster; it is to reduce time-to-correct-result by finding the right code sooner, shrinking unnecessary context, avoiding unsafe retries and requiring evidence before completion.
4
-
5
- ## Execution profiles
6
-
7
- UES classifies work into three execution profiles:
8
-
9
- - **FAST**: small/low-risk work, up to 2 skills, 12k context budget, targeted verification, no durable worktree/critic overhead by default.
10
- - **STANDARD**: medium work, up to 4 skills, 24k context budget, semantic+Git context and targeted/affected verification.
11
- - **DEEP**: long-horizon or high-risk work, up to 5 skills, 48k context budget, durable state, worktree isolation, critic/integration verification and full CI.
12
-
13
- Risk still overrides speed. Authentication, security, payment, schema/migration, production/deploy and public-contract work remains fail-closed.
14
-
15
- ## Incremental evidence index
16
-
17
- The V9 index caches bounded source metadata under `.ues-cache/semantic-index-v1.json`. Unchanged files reuse cached evidence while changed files are reparsed. It records:
18
-
19
- - concrete source paths;
20
- - bounded symbol definitions with line numbers;
21
- - lexical identifier counts;
22
- - cache reuse/reparse statistics.
23
-
24
- This is intentionally labelled **syntax-aware lexical evidence**. It must not be presented as proof of program semantics or as a full AST/LSP call graph.
25
-
26
- Commands:
27
-
28
- ```bash
29
- ocskill index status .
30
- ocskill index build .
31
- ocskill index rebuild .
32
- ocskill aci search "checkout total" .
33
- ocskill aci refs calculateTotal .
34
- ocskill aci view src/checkout.ts . --line 120 --lines 80
35
- ocskill aci text "exact marker" .
36
- ```
37
-
38
- ## Weak-model ACI
39
-
40
- The ACI keeps search/view output bounded and evidence-first. Path traversal outside the repository is rejected. Large/binary files are refused by the bounded viewer. Reference results distinguish lexical references from concrete definition lines so an agent cannot honestly claim deeper semantic certainty than the tool measured.
41
-
42
- ## Runtime traces
43
-
44
- Operational trace events are appended under `.ues-traces/*.jsonl`. Common credentials/tokens are redacted before persistence, oversized payloads are hashed/truncated, and traces explicitly exclude hidden chain-of-thought.
45
-
46
- ```bash
47
- ocskill trace show <trace-id> .
48
- ```
49
-
50
- ## Verification sandbox
51
-
52
- When Docker or Podman is available, deterministic verification commands can run with:
53
-
54
- - network disabled by default;
55
- - Linux capabilities dropped;
56
- - no-new-privileges;
57
- - read-only container root;
58
- - bounded PIDs/memory/CPU;
59
- - only the requested workspace bind-mounted;
60
- - no host secrets forwarded by default.
61
-
62
- This protects verification commands; it does **not** claim to sandbox the OpenCode model process itself.
63
-
64
- ## Benchmark confidence gate
65
-
66
- The matrix now embeds paired baseline/UES evidence and computes:
67
-
68
- - paired wins/losses/ties;
69
- - pass-rate delta;
70
- - exact two-sided sign-test p-value;
71
- - per-suite regression checks;
72
- - mean duration ratio;
73
- - optional cost ratio.
74
-
75
- Normal reporting:
76
-
77
- ```bash
78
- npm run evals:matrix -- --model provider/model --trials 3
79
- ```
80
-
81
- Fail-closed release gate:
82
-
83
- ```bash
84
- npm run evals:matrix:gate -- --model provider/model --trials 3
85
- ```
86
-
87
- The gate requires enough paired samples, positive uplift, more UES wins than losses, statistical support, no suite regression and acceptable speed. This is evidence for the measured benchmark only; it is not evidence that UES turns a weaker model into a different model.
88
-
89
- ## Lock and Windows hardening
90
-
91
- State locking now uses owner tokens and heartbeats. Stale takeover renames the old lock before removal, and release only removes a lock whose token still belongs to the caller. Windows CLI/router invocation resolves Node-backed shims and refuses unrecognized batch shims rather than falling back to shell execution.
92
-
93
- ## Release validation
94
-
95
- Before publishing 9.0.0:
96
-
97
- ```bash
98
- npm run ci
99
- npm run evals:matrix:gate -- --model provider/model --trials 3
100
- ```
101
-
102
- If GitHub-hosted Actions still fails before the first workflow step, treat that as an infrastructure/repository-action issue rather than test evidence; local CI and the paired benchmark remain required release gates.