opencode-agent-skill 13.0.0-beta.2 → 14.2.0-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (93) hide show
  1. package/README.md +1556 -607
  2. package/bin/ocskill.mjs +172 -24
  3. package/docs/OPENCODE-COMPAT.md +34 -97
  4. package/docs/PI-COMPAT.md +188 -0
  5. package/docs/V14-CONTEXT-MEMORY-FABRIC.md +70 -0
  6. package/docs/V14.1-QUALITY-PERFORMANCE-FABRIC.md +114 -0
  7. package/docs/V14.2-TURBO-WEAK-MODEL-RUNTIME.md +448 -0
  8. package/evals/v14/tasks.json +46 -0
  9. package/global-config/agents/executor.md +7 -0
  10. package/global-config/agents/visual-verifier.md +22 -3
  11. package/global-config/plugins/ues-router/index.js +13 -8
  12. package/global-config/plugins/ues-router/policy-runtime.js +7 -0
  13. package/global-config/plugins/ues-router/router.js +12 -2
  14. package/global-config/skills/ecommerce-engineering/SKILL.md +1 -1
  15. package/global-config/skills/file-upload-engineering/SKILL.md +1 -1
  16. package/global-config/skills/git-safety/SKILL.md +1 -1
  17. package/global-config/skills/nestjs-engineering/SKILL.md +1 -1
  18. package/global-config/skills/performance-engineering/SKILL.md +1 -1
  19. package/global-config/skills/react-native-engineering/SKILL.md +1 -1
  20. package/global-config/skills/rest-api-design/SKILL.md +1 -1
  21. package/global-config/skills/ui-ux-engineering/SKILL.md +1 -1
  22. package/lib/adaptive-context-budget.mjs +97 -0
  23. package/lib/affected-tests.mjs +260 -0
  24. package/lib/benchmark-confidence.mjs +41 -2
  25. package/lib/browser-mcp-routing.mjs +166 -0
  26. package/lib/capability-fabric.mjs +336 -0
  27. package/lib/capability-registry.mjs +9 -0
  28. package/lib/context-engine-v11.mjs +65 -1
  29. package/lib/context-graph-rank.mjs +118 -0
  30. package/lib/context-manifest.mjs +97 -18
  31. package/lib/control-center.mjs +19 -1
  32. package/lib/dynamic-workflow.mjs +3 -1
  33. package/lib/evidence-store.mjs +82 -1
  34. package/lib/hierarchical-context.mjs +215 -0
  35. package/lib/memory-engine.mjs +465 -0
  36. package/lib/model-performance.mjs +33 -8
  37. package/lib/model-policy.mjs +3 -3
  38. package/lib/orchestrator-policy.mjs +5 -209
  39. package/lib/performance-fabric.mjs +229 -0
  40. package/lib/pi-rpc-pool.mjs +433 -0
  41. package/lib/process-hang-detector.mjs +83 -0
  42. package/lib/process-supervisor.mjs +193 -0
  43. package/lib/prompt-cache.mjs +2 -0
  44. package/lib/repo-graph.mjs +53 -2
  45. package/lib/runtime-config.mjs +31 -0
  46. package/lib/safety.mjs +132 -0
  47. package/lib/semantic-index.mjs +52 -3
  48. package/lib/skill-compiler.mjs +128 -0
  49. package/lib/skill-quality.mjs +48 -2
  50. package/lib/task-engine.mjs +66 -5
  51. package/lib/task-policy.mjs +235 -0
  52. package/lib/verification-broker.mjs +284 -0
  53. package/lib/verification-command.mjs +111 -0
  54. package/lib/windows-shim.mjs +35 -0
  55. package/lib/workspace-fingerprint.mjs +198 -0
  56. package/package.json +52 -42
  57. package/pi/extensions/ues-child-runtime.ts +238 -0
  58. package/pi/extensions/ues.ts +3200 -0
  59. package/pi/prompts/ues-audit.md +9 -0
  60. package/pi/prompts/ues-critique.md +9 -0
  61. package/pi/prompts/ues-debug.md +9 -0
  62. package/pi/prompts/ues-feature.md +9 -0
  63. package/pi/prompts/ues-fix.md +9 -0
  64. package/pi/prompts/ues-plan.md +9 -0
  65. package/pi/prompts/ues-research.md +9 -0
  66. package/pi/prompts/ues-resume.md +9 -0
  67. package/pi/prompts/ues-review.md +7 -0
  68. package/pi/prompts/ues-run.md +17 -0
  69. package/pi/prompts/ues-verify.md +9 -0
  70. package/scripts/check-release-consistency.mjs +119 -185
  71. package/scripts/check-runtime-exports.mjs +66 -0
  72. package/scripts/check-source-integrity.mjs +184 -0
  73. package/scripts/eval-pi.mjs +492 -0
  74. package/scripts/install.mjs +16 -0
  75. package/scripts/smoke-package-closure.mjs +110 -0
  76. package/scripts/smoke-packed-install.mjs +24 -11
  77. package/scripts/smoke-pi-extension.mjs +144 -0
  78. package/scripts/uninstall.mjs +44 -0
  79. package/CHANGELOG.md +0 -415
  80. package/docs/DETERMINISTIC-TOOLS.md +0 -105
  81. package/docs/ENGINEERING-DESIGN.md +0 -194
  82. package/docs/EVALS.md +0 -158
  83. package/docs/GITHUB-RULESET.md +0 -50
  84. package/docs/NPM-PUBLISH.md +0 -116
  85. package/docs/RESEARCH-SOURCES.md +0 -37
  86. package/docs/TRACE-SCHEMA.md +0 -122
  87. package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +0 -75
  88. package/docs/V11-PERCEPTION-ADAPTIVE.md +0 -220
  89. package/docs/V12-WEAK-MODEL-INTELLIGENCE.md +0 -27
  90. package/docs/V13-PARALLEL-WEAK-MODEL-RUNTIME.md +0 -86
  91. package/docs/V7-INTELLIGENCE-RUNTIME.md +0 -166
  92. package/docs/V8-INTELLIGENCE-RELIABILITY.md +0 -206
  93. package/docs/V9-SPEED-INTELLIGENCE.md +0 -102
package/CHANGELOG.md DELETED
@@ -1,415 +0,0 @@
1
- # Changelog
2
-
3
- All notable changes to this project are documented here.
4
-
5
- The project follows Semantic Versioning.
6
-
7
- ## [Unreleased]
8
-
9
- ## [13.0.0-beta.2] - 2026-09-23
10
-
11
- ### Fixed
12
- - Routed all 11 `/ues-*` commands through the stable V2 `session.prompt` path using router-managed prompt aliases, avoiding `UnsupportedContentType` failures from the native `session.command` transport while preserving compatible command behavior.
13
- - Prevented `/ues-run` from duplicating the full `$ARGUMENTS` payload inside its task-policy example, reducing long-prompt amplification.
14
- - Fixed OpenCode V2 local router loading on clean global configs by removing the unnecessary bare `@opencode/plugin` runtime import; `ues-router` now exports the plain `{ id, setup }` definition accepted by the V2 loader.
15
- - Prevented OpenCode V2 UES router subprocesses (`where`, `ocskill`, Node shim execution and Git probes) from flashing transient CMD windows on Windows by routing them through a hidden-window spawn wrapper.
16
- - Changed `/ues-run` from an unconditional long-horizon declaration into adaptive FAST / STANDARD / DEEP admission driven by the actual user request and task policy.
17
- - Removed the synthetic `Explicit long-horizon engineering request.` policy input for `/ues-run`; only `/ues-resume` retains explicit durable-resume semantics.
18
- - Added compact FAST and STANDARD V2 prompt envelopes so trivial or bounded work does not inherit `.ues-work`, plan-gate, subagent-dispatch and integration-gate overhead.
19
- - Kept a conservative fallback: when policy classification is unavailable, the full command contract is preserved instead of weakening verification.
20
-
21
- ### Regression coverage
22
- - Added V13 tests proving that `/ues-run Chỉ trả lời đúng một từ: OK` stays FAST, bounded fixes stay STANDARD, whole-repository work escalates to DEEP, and `/ues-resume` remains long-horizon.
23
- - Added a contract test ensuring the bundled `/ues-run` template itself is adaptive rather than declaring every invocation long-horizon.
24
-
25
- ## [13.0.0-beta.1] - 2026-09-23
26
-
27
- ### Fixed
28
- - Hardened V2 prompt aliases for ChatGPT-style pasted prompts: BOM/zero-width prefixes and whole-prompt markdown fences are normalized before alias detection.
29
- - Bounded the text sent to `ocskill task-policy` so very large pasted prompts no longer risk Windows command-line length failures; the full user prompt still goes to the model while only a bounded head/tail classification view goes through the CLI.
30
- - Routed skill classification from the actual user request instead of the expanded command template, reducing template-induced over-routing on long prompts.
31
- - Kept installer result shape stable with `promptAliases: []` on foreign-state refusal.
32
-
33
- ### Regression coverage
34
- - Added fenced-paste, BOM, 70k-character policy-input and oversized multiline `/ues-run` regression tests.
35
-
36
- ## [13.0.0-beta.0] - 2026-09-22
37
-
38
- ### Added
39
- - Added native same-model parallel execution through `ues.dispatch_parallel` with bounded adaptive concurrency.
40
- - Added event-driven DAG scheduling so newly unblocked tasks can start without waiting for an entire wave barrier.
41
- - Added read/write resource leases, conservative unknown-scope serialization and shared-config writer serialization.
42
- - Added inherited-root worktree snapshots so downstream tasks see already integrated predecessor changes without committing the user's root branch.
43
- - Added transactional integration with sandbox rollback when receipt/completion fails after patch application.
44
- - Added fresh same-model verifier sessions and agent-verifier receipts tied to the active run and current workspace fingerprint.
45
- - Added global CLI help interception, structured `--json` errors, `work status .` workspace listing, UTF-8/UTF-16 auto-decoding, safe UTF-8 Git diff output and text normalization.
46
-
47
- ### Regression coverage
48
- - Added V13 tests for CLI help parsing, structured errors, UTF-16 PowerShell-style diff decoding, event-driven same-model concurrency, resource conflict serialization and inherited-root rollback behavior.
49
- - Added verifier-runtime regression coverage against forged/user-supplied PASS markers and malformed verdict lines.
50
- - Added structured `verificationCommands` validation coverage for deterministic post-integration receipts.
51
-
52
- ### Fixed
53
- - Fixed the V13 structured-verification regression test wiring so the current HEAD test suite can execute the new plan-schema assertions.
54
- - Fixed npm 11+ lifecycle-script blocking in release smoke expectations: a plain global install may require the documented `ocskill install` resource-sync fallback, and CI now verifies that fallback instead of falsely requiring postinstall execution.
55
- - Clarified OpenCode 1.x versus V2 capability boundaries and fenced canonical long-task state to `.ues-work/<slug>/`; legacy/manual `ues-work/` directories are no longer treated as official UES state.
56
- - Added `repo-graph --compact` and switched initial long-run evidence gathering to compact hotspot/count summaries to reduce large-repository context/tool-output overhead.
57
- - Made `ocskill install/status/doctor` surface the OpenCode 1.x versus V2 native-parallel boundary directly so users do not mistake V13 CLI availability for fresh-session parallel availability.
58
- - Hardened migration from V12/manual Windows artifacts by auto-decoding UTF-8/UTF-16 `PLAN.json`, `STATE.json`, `SPEC.md` and dependency reports across the CLI/task engine and V2 parallel router.
59
- - Added bounded `ocskill text-read` so OpenCode 1.x can recover known UTF text that its generic Read tool classifies as binary, without creating ad-hoc converted copies.
60
- - Extended packed-install smoke and regressions to cover the UTF recovery command, UTF durable state discovery and legacy artifact context packs.
61
- - Removed the stale clean-root-only restriction from native parallel dispatch. Existing dirty repository state is now treated as an inherited baseline, while isolated worktree delta integration, resource leases and rollback continue to protect user changes.
62
- - Fixed inherited untracked-file handling so a parallel worker may safely edit an unchanged pre-existing untracked file, while a user edit that races after sandbox creation is still detected and rejected as a conflict.
63
- - Aligned legacy smoke tests with V13 structured usage exit code 2 and typed verification evidence (`command-receipt-backed` / `independent-agent-receipt-backed`).
64
- - Fixed `work status` receipt coverage accounting for the typed V13 evidence strengths and made inherited-untracked worktree regressions line-ending neutral on Windows.
65
-
66
-
67
- ## [12.0.0-beta.0] - 2026-09-22
68
-
69
- ### Beta
70
- - Added empirical per-task-class model performance history and capability-preserving reranking with a minimum evidence threshold before reranking.
71
- - Added context quality receipts for required-file recall and irrelevant-context ratio.
72
- - Added hash-keyed plan snapshots with active-plan execution fencing.
73
- - Added bounded decision policy for reversible local rulings versus human-gated destructive/external actions.
74
- - Added deterministic repo-scale benchmark fixture generation and V12/repo-scale validation gates.
75
- - Hardened release consistency checks to derive eval counts and validate aggregate workflow structure.
76
- - Fixed missing empirical-history handling so models without benchmark history safely fall back to static capability routing.
77
-
78
- ### Verified locally on Windows
79
- - 255 tests total: 253 passed, 0 failed, 2 platform-specific skips.
80
- - Syntax, catalog validation, docs consistency, routing, V11/V12/repo-scale/live/long/polyglot validation: PASS.
81
- - npm pack, packed-install smoke and plain one-command install/resource auto-sync smoke: PASS.
82
- - Published prerelease intent: npm dist-tag `next`; V11 remains `latest` until V12 stable release gates are satisfied.
83
-
84
-
85
- ## [11.0.0] - 2026-09-22
86
-
87
- ### Released
88
- - Promoted V11 perception-aware adaptive execution to stable after the full local CI gate passed on Windows with 234 tests total, 232 passed, 0 failed and 2 platform-specific skips.
89
- - Stable npm installs use the default `latest` dist-tag, so users install with `npm install -g opencode-agent-skill`.
90
- - Includes content-addressed Evidence Store, adaptive EvidenceBudget/context externalization, stable-prefix prompt telemetry, capability-aware model routing, visual geometry receipts, deterministic PNG diff/crop, responsive/design-token inspection, optional Playwright browser inspection, cost-aware dynamic workflow scheduling, 48 focused skills and 12 subagents.
91
-
92
- ### Verified
93
- - Syntax: 178 JavaScript modules.
94
- - Catalog: 48 skills, 11 commands and 12 subagents.
95
- - Router: 129 cases, 294/294 required routes and 61/61 negative guards.
96
- - V11 contracts: 13 tasks across 6 categories and 15 required runtime files.
97
- - Live/long/polyglot fixture validation: 20 / 5 / 8 tasks.
98
- - npm pack dry-run, packed-install smoke and plain one-command install/resource auto-sync smoke: PASS.
99
-
100
- ### Changed
101
- - Package version is 11.0.0.
102
- - V11 becomes the stable npm release line; V10 remains in Git history as the previous stable release.
103
-
104
-
105
- ### V11 dev.2
106
- - Added project-local, fail-closed Playwright browser inspection that returns bounded semantic elements, bounding boxes, computed visual properties and a screenshot path while treating page content as untrusted evidence.
107
- - Exposed browser inspection through CLI and the OpenCode V2 router without making Playwright a required package dependency.
108
- - Fixed duplicate `ui_layout` router tool registration and aligned durable context packs with V11 context schema version 6.
109
- - Extended V11 validation and regression coverage for the browser runtime.
110
- - Package development version is 11.0.0-dev.2.
111
-
112
- ## [11.0.0-dev.1] - 2026-09-22
113
-
114
- ### Added
115
- - Adaptive context engine that externalizes oversized source/test/reference excerpts into content-addressed evidence pointers while keeping bounded inline previews.
116
- - Deterministic UI layout and design-token tools exposed to the OpenCode V2 runtime.
117
- - Cost- and modality-aware workflow scheduling with inline thresholds, deterministic-first waves, independent LLM/vision concurrency and bounded wave cost.
118
- - Repeated-stable prompt ratio telemetry and an optional fail-closed ablation gate for repeated input efficiency.
119
- - Stronger V11 contract coverage for adaptive context, UI inspection and evidence externalization.
120
-
121
- ### Changed
122
- - V11 router metadata now reports version 11.
123
- - Task context packs transport large contextual evidence through `evidence:sha256` pointers and report externalized byte/ref counts.
124
- - Capability-aware model routing fails closed when an enabled configured model set cannot satisfy required capabilities such as vision/browser.
125
- - Package development version is 11.0.0-dev.1.
126
-
127
- ### Fixed
128
- - V11 eval-report regression tests now account for expanded adaptive telemetry coverage instead of using the old V10-only telemetry shape.
129
-
130
- ## [11.0.0-dev.0] - 2026-09-22
131
-
132
- ### Added
133
- - Content-addressed Evidence Store with bounded retrieval, deduplication and garbage collection.
134
- - Adaptive Evidence Budget planning and context-manifest integration for evidence-efficient execution.
135
- - Prompt stable-prefix/cache telemetry for repeated-input measurement.
136
- - Capability registry and capability-aware model selection for coding, reasoning, tools, vision, browser, filesystem and long-context needs.
137
- - Visual specification, geometry receipts, responsive viewport matrix, dependency-free PNG decode/diff/crop and bounded visual repair planning.
138
- - Browser QA adapter with CLI-first verification planning, targeted semantic evidence and explicit untrusted-page security boundaries.
139
- - Cost-aware dynamic workflow scheduler separating deterministic work from LLM/vision work.
140
- - Skill-quality linting for entrypoint size, metadata and routing-description collision detection.
141
- - Nine V11 skills: visual-fidelity, browser-qa, design-source, responsive-verification, component-visual-testing, browser-security, skill-authoring, skill-evaluation and dynamic-workflow.
142
- - Two V11 subagents: visual-verifier and merge-arbiter.
143
- - Optional Hermes sidecar workflow contract with evidence-pointer transport.
144
- - V11 Control Center evidence-store/runtime telemetry and V11-specific tests/eval routing cases.
145
-
146
- ### Changed
147
- - Model policy schema supports per-model capability metadata plus cost, latency and quality hints.
148
- - Adaptive model resolution can select configured models by required task capabilities.
149
- - OpenCode V2 router recognizes visual/browser/design/skill-workflow intents and keeps FAST routing selective.
150
- - Package version is 11.0.0-dev.0 while npm latest remains V10 stable until V11 release gates pass.
151
-
152
- ## [10.0.0] - 2026-09-22
153
-
154
- ### Released
155
- - Promoted V10 RC.2 to stable after the local full gate passed with 182 tests passing, 0 failing, 2 platform-specific skips, plus package and install smoke validation.
156
- - Stable npm installs use the default `latest` dist-tag, so users install with `npm install -g opencode-agent-skill`.
157
- - Includes adaptive context/routing, weak-model recovery, no-progress watchdog, duplicate/loop guards, bounded exploration output, durable compaction checkpoints, provider recovery and retryable lease recovery.
158
-
159
- ### Changed
160
- - Package version is 10.0.0.
161
- - V10 stable keeps FAST at 8k, STANDARD at 20k and DEEP at 48k while preserving configured executor capability and correctness-first verification gates.
162
-
163
- ## [10.0.0-rc.2] - 2026-09-22
164
-
165
- ### Added
166
- - No-progress watchdog for fresh executor sessions with active-tool grace so legitimate long-running tools are not killed merely for being quiet.
167
- - Duplicate exploration guard and loop detector for repeated read/grep/glob-style calls without new workspace/evidence progress.
168
- - Runtime tool-output budgets for large grep/glob/read/repo-graph style outputs.
169
- - Durable pre-compaction checkpoints containing current task, runId, plan hash, workspace fingerprint, evidence pointers and a structured next action.
170
- - Post-compaction resume enforcement with automatic fresh-session recovery when the next action is not executed.
171
- - Provider failure classification and recovery for no-token, timeout, rate-limit, upstream, quota, auth and context-overflow failures.
172
- - Periodic lease supervisor that recovers expired running executors to an explicit retryable state.
173
-
174
- ### Changed
175
- - Provider stalls retry once in a fresh session with the same model; repeated retryable provider failures use the configured escalation model/provider when available.
176
- - Stale lease recovery now records `retryable` instead of overloading terminal/logical `failed`.
177
- - Explicit `long/high-risk`, `high-risk`, and structured long/high-risk facts hard-override FAST/light routing to DEEP/heavy.
178
- - Package version is 10.0.0-rc.2.
179
- - npm tag releases publish prereleases under `next`; stable versions continue to use `latest`.
180
-
181
- ### Fixed
182
- - A post-compaction session that only emits narrative text without executing the persisted next action is treated as stalled instead of silently completing.
183
-
184
- ## [10.0.0-rc.1] - 2026-09-22
185
-
186
- ### Added
187
- - Initial-input token telemetry and aggregate reporting for baseline-vs-UES evaluation.
188
- - Attempt-aware recovery policy with bounded context/skill escalation from initial execution to diagnosis and deep recovery.
189
- - Policy-aware FAST runtime routing that prioritizes direct domain/debug/review skills over generic orchestration.
190
- - Benchmark efficiency gates for initial input tokens and total tokens.
191
- - Reference-vs-candidate ablation CLI that requires pass-rate preservation and measurable initial-context reduction.
192
-
193
- ### Changed
194
- - FAST context budget is 8k and STANDARD is 20k; DEEP remains 48k to preserve high-risk/long-horizon capability.
195
- - Always-loaded global engineering instructions are compressed while retaining exact-contract, evidence, verification, safety and durable-work invariants.
196
- - Retry context expands only after failure; repeated failures add graph/critic evidence instead of repeating speculative patches.
197
- - Adaptive model policy exposes recovery stage and does not silently downshift the configured/default executor tier for FAST tasks.
198
- - OpenCode V2 router metadata is versioned for the V10 policy-aware path.
199
- - Package version is 10.0.0-rc.1; npm publication/tagging is intentionally deferred until RC verification passes.
200
-
201
- ## [9.0.0] - 2026-09-22
202
-
203
- ### Added
204
- - Persistent incremental source index with bounded syntax-aware symbol/reference evidence and deterministic cache reuse.
205
- - Weak-model-oriented ACI commands for semantic search, exact text search, bounded file viewing and reference lookup.
206
- - FAST / STANDARD / DEEP execution profiles with adaptive context budgets, skill caps and verification depth.
207
- - Redacted operational trajectory JSONL for replay/debugging without recording hidden chain-of-thought.
208
- - Optional Docker/Podman verification sandbox with network-off, dropped capabilities, no-new-privileges and resource bounds.
209
- - Paired baseline-vs-UES confidence analysis with exact sign-test evidence, per-suite no-regression checks and speed limits.
210
- - Fail-closed `npm run evals:matrix:gate -- --model provider/model` release benchmark gate.
211
- - Regression tests for semantic index, ACI, trajectory redaction, sandbox arguments, benchmark confidence and stale lock takeover.
212
-
213
- ### Changed
214
- - Context Manifest v4 consumes the incremental evidence index before broader graph expansion and keeps evidence labels explicit.
215
- - Learning promotion now requires complete paired benchmark evidence, statistically supported uplift and no suite regression.
216
- - OpenCode V2 dispatch records bounded redacted operational traces and adaptive execution-profile metadata.
217
- - Hardened live benchmark fairness by making previously implicit hidden-grader contract details explicit in task prompts without weakening hidden graders.
218
- - Clarified webhook stale/duplicate handling to require returning the exact original state reference, matching the hidden contract.
219
- - Counterbalanced paired live-eval execution order across tasks and trials to reduce provider order/throttling bias.
220
- - Focused small/FAST fixes now follow literal acceptance-contract discipline and avoid repo-wide discovery unless evidence requires it.
221
- - OpenCode live-eval telemetry now reads native `part.tokens` / `tokens` payloads in addition to `usage` payloads.
222
- - Extracted shared CLI parsing/read/truncation/error helpers into `lib/cli-utils.mjs` and added smoke/unit regression coverage.
223
- - Preserved documented CLI edge semantics while removing ad-hoc option parsing and unifying bounded output clipping.
224
- - Resolved README version drift and documented V9 capabilities.
225
- - Package version is 9.0.0.
226
-
227
- ### Fixed
228
- - State locks use unique ownership tokens, heartbeat refresh and rename-based stale takeover so an expired owner cannot delete a replacement lock.
229
- - Generated `.ues-cache/` and `.ues-traces/` state no longer invalidates workspace verification fingerprints.
230
- - Windows CLI/router execution no longer falls back to shell-based `.cmd/.bat` invocation for unrecognized shims.
231
- - Windows doctor/OpenCode compatibility probing safely resolves extensionless Node-backed npm shims and fails closed instead of invoking `cmd.exe`.
232
- - Live benchmark execution now uses the same shell-free Windows resolver as `doctor`, skips unsupported batch shims in favor of safe native executables, and normalizes shim paths cross-platform.
233
- - Windows npm shim resolution can recover from non-standard `.cmd` formatting by resolving only an adjacent package's explicit `package.json` bin mapping, including extensionless Node launchers such as OpenCode.
234
- - Windows shim resolution also accepts validated native PE targets declared by npm package metadata, covering `opencode-ai` installs whose `bin.opencode` points to `bin/opencode.exe`.
235
-
236
-
237
- ## [8.0.0] - 2026-09-20
238
-
239
- ### Added
240
- - Structured plan and integration gate receipts bound to the current plan hash or workspace fingerprint, with gate receipts persisted in `EVIDENCE.json`.
241
- - Strict long/high-risk task completion that requires a successful verification receipt for the active `runId` and the current workspace fingerprint.
242
- - Append-only `EVENTS.jsonl` runtime journal for work initialization, plan import/approval, task start/heartbeat/session binding, verification receipts, failure/recovery, integration verification and finalization.
243
- - Task-scoped stale recovery and OpenCode V2 runtime tools for bounded executor cancellation/recovery.
244
- - Context manifest v3 with multilingual task terms, Git-change awareness, symbol hits, TF-IDF-style relevance scoring, related tests/instructions and adaptive centered excerpts.
245
- - Conflict-aware Git worktree integration plus automatic isolation support for concurrent writing executors.
246
- - Learning v2 with recurring failure clustering, explicit acceptance and shadow-benchmark promotion gates.
247
- - UES benchmark matrix runner for baseline-vs-UES comparison across standard, long-horizon and polyglot suites.
248
- - Eight polyglot benchmark tasks spanning Python, Java/Spring-style code, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated-contract discipline.
249
- - Control Center runtime-event visibility, receipt inspection and stale-task recovery control.
250
- - CodeQL, dependency review, Dependabot maintenance and release-tag/version consistency checks.
251
-
252
- ### Changed
253
- - Fresh OpenCode V2 executor sessions now use bounded waits and `session.interrupt` on timeout, with capability probing and graceful degradation when optional hooks are unavailable.
254
- - Process execution escalates Unix process-tree cancellation from SIGTERM to SIGKILL after a bounded grace period and always reports cancelled runs as nonzero.
255
- - Router/task policy adds multilingual and framework-aware signals while preserving deterministic caps.
256
- - Managed resource ownership now uses `managed-by: opencode-agent-skill` while automatically recognizing and migrating the former scoped marker.
257
- - Package version is 8.0.0. Existing `opencode-agent-skill@7.7.0` users remain on the same package name and can update normally.
258
-
259
- ### Fixed
260
- - Strict verification no longer accepts a receipt after the workspace changed.
261
- - Sandbox cleanup refuses to delete non-UES branches and removes temporary UES branches after successful integration.
262
- - Learning proposals cannot be promoted before explicit acceptance.
263
- - Plain npm-install smoke now reports lifecycle-script auto-sync versus explicit `ocskill install` recovery accurately.
264
-
265
- ## [7.7.0] - 2026-09-20
266
-
267
- ### Added
268
- - Crash-safe long-task leases with per-attempt `runId`, executor owner metadata, heartbeat timestamps, lease expiry, explicit heartbeat command and stale-task recovery.
269
- - Structured verification receipts with command/args, exit code, timing, SHA-256 output digests, task/run fencing and before/after workspace fingerprints.
270
- - Context intelligence manifests that add declared files, import neighbors, likely tests, repository instructions/manifests, bounded excerpts and accepted learnings to fresh executor handoffs.
271
- - Adaptive task classification via `ocskill task-policy` and risk/complexity-aware model selection layered on top of attempt escalation.
272
- - Read/write-aware safe-wave scheduling plus isolated Git worktree sandbox primitives for parallel write tasks.
273
- - Evidence-gated learning loop over `.ues-evals` with deterministic proposals, explicit acceptance and relevant accepted lessons fed back into future context packs.
274
- - Optional Hermes adapter commands that detect Hermes and emit bounded UES delegation prompts without embedding Hermes into the UES runtime.
275
- - Zero-dependency local UES Control Center for durable work state, evidence, learning proposals and recent evaluation summaries.
276
- - V2 runtime capability probing tool and fail-closed fresh-dispatch checks.
277
-
278
- ### Changed
279
- - The public npm distribution name is the unscoped `opencode-agent-skill`, so users install it with `npm install -g opencode-agent-skill`.
280
- - Live evaluation runs are asynchronous and observable: start messages, periodic heartbeats, hard timeout, idle timeout and Ctrl+C process-tree cancellation are supported.
281
- - V2 fresh task dispatch refreshes durable task leases while the executor session runs and reports task policy/runId alongside model policy.
282
- - Workspace fingerprints ignore UES runtime-only learning/dashboard/sandbox directories in addition to `.ues-work`.
283
- - Package version is 7.7.0.
284
-
285
- ### Fixed
286
- - Packed global-install smoke derives the package install path from package metadata instead of assuming the former scoped npm name.
287
- - The installer accepts the former `@laivannha0202/opencode-agent-skill` state owner as legacy UES ownership and re-owns it as `opencode-agent-skill` during the next install.
288
- - Live baseline/UES evaluation detects the OpenCode major version: OpenCode 1.x omits the V2-only `--standalone` flag, while OpenCode 2.x+ keeps it. Eval JSON records the detected OpenCode version/major for reproducibility.
289
- - The V2 automatic router retains domain/impact skills and `ues-engineering-orchestrator` ahead of generic process skills when the configured skill cap is exceeded.
290
-
291
- ## [6.0.0] - 2026-09-19
292
-
293
- ### Added
294
- - Durable long-horizon work engine under `.ues-work/<slug>/` with SPEC, machine-readable PLAN/STATE/EVIDENCE, task briefs and reports.
295
- - Four long-horizon roles: `ues-codebase-mapper`, `ues-plan-checker`, editable `ues-executor`, and `ues-integration-verifier`.
296
- - `/ues-run` and `/ues-resume` commands.
297
- - Deterministic `repo-graph`, `review-scope`, `verification-plan`, `task-graph`, `context-pack`, and `ocskill work` commands.
298
- - Dependency-safe DAG waves with conservative declared-file overlap serialization.
299
- - Machine-enforced plan approval through `ocskill work approve-plan`.
300
- - Machine-enforced integration verdict through `ocskill work verify-integration`.
301
- - Workspace fingerprint gate that invalidates finalization when code changes after integration PASS.
302
- - Per-work-item lock and atomic state/evidence writes for parallel-safe durable state updates.
303
- - OpenCode V2 `ues.dispatch_task` runtime tool that creates a fresh executor session and applies configured attempt-based model escalation.
304
- - Configurable light/standard/heavy model tiers and role mappings.
305
- - V2 permission safety gate for forceful Git, publish, destructive file/database and deployment commands.
306
- - 120-case V2 router trigger evaluation.
307
- - Five long-horizon behavioral tasks, including a combined 15-source-file integration task.
308
-
309
- ### Changed
310
- - Package version is 6.0.0 and release documentation now describes 39 skills, 11 commands and 10 subagents.
311
- - Long-suite UES runs count as PASS only when both the hidden grader and durable orchestration state pass.
312
- - Long-task completion now requires independent plan approval, per-task fresh evidence, integration PASS, and an unchanged post-verification workspace.
313
- - `.ues-work/` is git-ignored because it is runtime execution state.
314
- - V2 runtime context guidance now prefers fresh `ues.dispatch_task` execution for approved tasks.
315
-
316
- ### Fixed
317
- - Prevented concurrent `STATE.json` / `EVIDENCE.json` lost updates during safe-wave completion.
318
- - `ocskill model-policy` now resolves against the user's persisted model-tier configuration instead of the default empty policy.
319
- - Safety detection now recognizes short-form `git push -f`.
320
-
321
-
322
- ## [4.0.0] - 2026-09-19
323
-
324
- ### Added
325
- - OpenCode 1.x/2.x compatibility detection with managed V2-native agent permission frontmatter.
326
- - Optional OpenCode V2 runtime skill router installed as a managed global plugin, with `ocskill router on|off|status --max N` controls.
327
- - Dependency-free deterministic repository helpers: stack detection, test-command detection, repository map, impact search, evidence snapshot, and working-tree inspection.
328
- - `ocskill inspect`, `ocskill impact`, `ocskill evidence`, `ocskill working-tree`, `ocskill detect-stack`, and `ocskill detect-tests`.
329
- - Twenty executable hidden-graded live benchmark tasks across correctness, auth, contracts, data, payments, security, frontend state, dependency compatibility, and multi-file changes.
330
- - Live benchmark authentication modes: isolated environment credentials by default and optional current OpenCode auth-file copy.
331
- - Best-effort OpenCode JSONL telemetry for tool calls, loaded skills, subagents, tokens, cost, and changed workspace files.
332
- - `ocskill eval-report` / `npm run evals:report` for baseline-vs-UES pass-rate and efficiency aggregation.
333
- - Live-suite integrity validation and broad JavaScript syntax validation in `npm run ci`.
334
- - Progressive-disclosure workflow references for 18 previously shallow domain/process skills.
335
- - OpenCode compatibility and deterministic-tool documentation.
336
-
337
- ### Changed
338
- - Static routing evaluation now contains 34 scenarios and requires every installed skill to be represented at least once.
339
- - Packed-install smoke testing now forces the OpenCode V2 compatibility path, checks native permissions, validates the managed router plugin, and exercises the packed `ocskill inspect` command.
340
- - Core workflow guidance now prefers deterministic evidence helpers before broad model-driven repository exploration.
341
- - Installer state schema records the detected OpenCode major and managed plugin resources.
342
- - Package contents now publish the complete `lib/`, `scripts/`, and `global-config/` trees required by V4.
343
-
344
- ### Fixed
345
- - `ocskill update` now resolves the explicit npm `latest` dist-tag with `npm view package@latest version` and falls back to `npm dist-tag ls`, avoiding stale untagged package metadata such as the observed 2.1.0/3.0.0 mismatch.
346
- - Managed V2 router resources are removed safely when uninstalling or re-syncing back to an OpenCode 1.x environment.
347
-
348
-
349
- ## [3.0.0] - 2026-09-19
350
-
351
- ### Added
352
- - Live baseline-vs-UES behavioral evaluation harness with isolated OpenCode configs, executable fixtures, hidden graders, multi-trial support, and JSON traces.
353
- - Independent `ues-critic` subagent and `/ues-critique` command for evidence-grounded falsification before completion.
354
- - Evaluator/repair orchestration reference with bounded critic-repair-reverify cycles.
355
- - Context-ledger guidance that separates confirmed facts, assumptions, rejected hypotheses, decisions, and fresh verification evidence.
356
- - Structured output contracts for architecture, debugging, research, review, critic, and verification subagents.
357
- - Trace schema documentation for live benchmark results.
358
-
359
- ### Changed
360
- - Long-task state now preserves rejected hypotheses and evidence so resumed work does not repeat disproved approaches.
361
- - Completion gates now require substantial/high-risk changes to resolve or explicitly surface evidence-backed blocking critic findings.
362
- - Installer tests now require the expanded command/subagent catalog and progressive-disclosure references.
363
- - Added a packed-install smoke test that installs the tarball into an isolated global npm prefix and verifies the package is a real copy rather than a source link/junction.
364
- - Bumped the package release line to 3.0.0 for the intelligence-loop and behavioral-eval release.
365
-
366
- ### Fixed
367
- - `ocskill update` now checks the published npm version first, refuses accidental downgrades, avoids reinstalling an equal version, performs newer package replacement with lifecycle scripts disabled, then explicitly re-syncs resources from the newly installed CLI.
368
- - `ocskill update` and `ocskill remove` now run npm from the user home directory instead of from inside the package directory being replaced or removed.
369
-
370
- ## [2.1.0] - 2026-09-18
371
-
372
- ### Added
373
- - Evidence-driven `engineering-orchestrator` skill with progressive-disclosure routing, verification, retry, and delegation references.
374
- - `context-engineering`, `research-verification`, `change-impact-analysis`, `long-task-state`, and pragmatic `test-driven-development` process skills.
375
- - Five optional read-only/analysis subagents: architect, debugger, researcher, reviewer, and verifier.
376
- - Four new commands: `/ues-plan`, `/ues-debug`, `/ues-verify`, and `/ues-research`.
377
- - Static routing eval contract and `npm run evals` / `ocskill eval`.
378
- - Engineering design and research-source documentation.
379
-
380
- ### Changed
381
- - Expanded the catalog from 33 to 39 skills and from 4 to 8 commands.
382
- - Installer now copies complete skill directories so references/templates survive installation.
383
- - Installer now manages namespaced OpenCode subagents and tracks them in state.
384
- - Strengthened repository exploration, planning, dependency management, root-cause debugging, code review, and verification.
385
- - `ocskill status` now compares package/resource versions and reports subagent synchronization.
386
- - `ocskill update` explicitly re-syncs resources after npm update.
387
- - `ocskill remove` explicitly cleans managed resources before npm uninstall.
388
- - Removed the Windows `shell: true` execution path that produced Node deprecation warnings.
389
-
390
- ## [2.0.1] - 2026-09-18
391
-
392
- ### Fixed
393
- - Added dedicated npm lifecycle entrypoints and more reliable global-install detection on Windows.
394
-
395
- ## [2.0.0] - 2026-09-18
396
-
397
- ### Changed
398
- - Converted the repository into a standard global npm CLI package.
399
- - Package name became `@laivannha0202/opencode-agent-skill`.
400
- - Standardized development and CI on Node.js + npm.
401
-
402
- ### Added
403
- - Global `ocskill` CLI.
404
- - Managed state under the global OpenCode config.
405
- - Idempotent skill/command synchronization.
406
- - Managed-block integration with an existing global `AGENTS.md`.
407
- - npm packaging validation and Node.js tests.
408
- - GitHub Actions workflow for npm publishing.
409
-
410
- ## [1.0.0] - 2026-09-18
411
-
412
- ### Added
413
- - Initial universal engineering skill collection.
414
- - Automatic engineering workflow rules.
415
- - Commands for fix, feature, review, and audit.
@@ -1,105 +0,0 @@
1
- # Deterministic evidence and execution tools
2
-
3
- V11 uses dependency-light Node helpers for work that should not rely on a model guessing or remembering it.
4
-
5
- ## Repository evidence
6
-
7
- ```cmd
8
- ocskill inspect .
9
- ocskill detect-stack .
10
- ocskill detect-tests .
11
- ocskill impact calculateOrderTotal .
12
- ocskill evidence .
13
- ocskill working-tree .
14
- ```
15
-
16
- These identify stack/package manager, project-native checks, bounded impact hits and Git state.
17
-
18
- ## Repository graph
19
-
20
- ```cmd
21
- ocskill repo-graph .
22
- ```
23
-
24
- Builds a bounded import graph, local edges, external import frequencies and coupling hotspots. It is not a full language server/call graph.
25
-
26
- ## Review scope
27
-
28
- ```cmd
29
- ocskill review-scope main .
30
- ```
31
-
32
- Enumerates changed files and deterministic risk hints for persistence/schema, auth/security, payments, public interfaces, dependencies and delivery/infrastructure.
33
-
34
- `coverageRequired` lets a reviewer account for every changed file instead of relying on memory.
35
-
36
- ## Verification plan
37
-
38
- ```cmd
39
- ocskill verification-plan .
40
- ```
41
-
42
- Combines project-native commands, working-tree evidence and changed-file risk into recommended checks plus risk-specific acceptance prompts.
43
-
44
- ## Task graph
45
-
46
- ```cmd
47
- ocskill task-graph PLAN.json
48
- ```
49
-
50
- Validates plan shape/dependencies/cycles and computes topological and safe waves. Same-wave tasks with overlapping/unknown declared files are serialized.
51
-
52
- ## Durable work state
53
-
54
- ```cmd
55
- ocskill work init <slug> . --goal "..."
56
- ocskill work plan <slug> PLAN.json .
57
- ocskill work gate-receipt <slug> plan . --verifier ues-plan-checker --evidence "PASS" --out .ues-work/<slug>/reports/plan-receipt.json
58
- ocskill work approve-plan <slug> . --evidence "PASS" --receipt-file .ues-work/<slug>/reports/plan-receipt.json
59
- ocskill work start <slug> T1 .
60
- ocskill work verify-command <slug> T1 . --run-id <run-id> -- npm test
61
- ocskill work complete <slug> T1 . --run-id <run-id> --evidence "verified"
62
- ocskill work fail <slug> T1 . --reason "..."
63
- ocskill work gate-receipt <slug> integration . --verifier ues-integration-verifier --verdict PASS --evidence "PASS" --out .ues-work/<slug>/reports/integration-receipt.json
64
- ocskill work verify-integration <slug> . --verdict PASS --evidence "PASS" --receipt-file .ues-work/<slug>/reports/integration-receipt.json
65
- ocskill work finalize <slug> . --evidence "final acceptance verified"
66
- ocskill work events <slug> . --limit 100
67
- ocskill work resume <slug> .
68
- ```
69
-
70
- State/evidence writes use a per-item lock and atomic replacement. `EVENTS.jsonl` is append-only runtime evidence. Long/high-risk tasks require successful receipts for the active run and current workspace fingerprint.
71
-
72
- ## Context pack
73
-
74
- ```cmd
75
- ocskill context-pack <slug> <task> .
76
- ```
77
-
78
- Returns the task, bounded spec, dependency reports, decisions, blockers, current task state and Context Manifest v3: declared files, import neighbors, likely tests, nearby instructions, Git-changed files, task-term relevance, symbol hits, centered excerpts and promoted lessons.
79
-
80
- ## Runtime dispatch on OpenCode V2
81
-
82
- The managed plugin exposes `ues.dispatch_task`, which combines `work start`, context pack, model policy and a fresh OpenCode executor session with heartbeat, bounded wait and interrupt-on-timeout. It also exposes runtime cancellation/recovery helpers when the OpenCode session API supports them.
83
-
84
- ## Sandboxes and learning
85
-
86
- ```cmd
87
- ocskill sandbox create <slug> <task-id> .
88
- ocskill sandbox integrate <worktree-path> .
89
- ocskill sandbox list .
90
- ocskill learn analyze . --eval-dir .ues-evals
91
- ocskill learn accept <proposal-id> .
92
- ocskill learn promote <proposal-id> . --baseline 0.50 --candidate 0.75 --samples 4
93
- ```
94
-
95
- Sandbox integration refuses overlap with dirty root files. Shadow-required learning proposals are not retrieved until a measured benchmark improvement is recorded.
96
-
97
- ## Constraints
98
-
99
- These helpers:
100
-
101
- - do not replace reading exact affected code
102
- - do not pretend text/import scans are complete semantic analysis
103
- - do not auto-merge/push/publish/deploy
104
- - preserve unrelated user work
105
- - use JSON outputs where machine consumption matters