opencode-agent-skill 13.0.0-beta.2 → 14.2.0-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1556 -607
- package/bin/ocskill.mjs +172 -24
- package/docs/OPENCODE-COMPAT.md +34 -97
- package/docs/PI-COMPAT.md +188 -0
- package/docs/V14-CONTEXT-MEMORY-FABRIC.md +70 -0
- package/docs/V14.1-QUALITY-PERFORMANCE-FABRIC.md +114 -0
- package/docs/V14.2-TURBO-WEAK-MODEL-RUNTIME.md +448 -0
- package/evals/v14/tasks.json +46 -0
- package/global-config/agents/executor.md +7 -0
- package/global-config/agents/visual-verifier.md +22 -3
- package/global-config/plugins/ues-router/index.js +13 -8
- package/global-config/plugins/ues-router/policy-runtime.js +7 -0
- package/global-config/plugins/ues-router/router.js +12 -2
- package/global-config/skills/ecommerce-engineering/SKILL.md +1 -1
- package/global-config/skills/file-upload-engineering/SKILL.md +1 -1
- package/global-config/skills/git-safety/SKILL.md +1 -1
- package/global-config/skills/nestjs-engineering/SKILL.md +1 -1
- package/global-config/skills/performance-engineering/SKILL.md +1 -1
- package/global-config/skills/react-native-engineering/SKILL.md +1 -1
- package/global-config/skills/rest-api-design/SKILL.md +1 -1
- package/global-config/skills/ui-ux-engineering/SKILL.md +1 -1
- package/lib/adaptive-context-budget.mjs +97 -0
- package/lib/affected-tests.mjs +260 -0
- package/lib/benchmark-confidence.mjs +41 -2
- package/lib/browser-mcp-routing.mjs +166 -0
- package/lib/capability-fabric.mjs +336 -0
- package/lib/capability-registry.mjs +9 -0
- package/lib/context-engine-v11.mjs +65 -1
- package/lib/context-graph-rank.mjs +118 -0
- package/lib/context-manifest.mjs +97 -18
- package/lib/control-center.mjs +19 -1
- package/lib/dynamic-workflow.mjs +3 -1
- package/lib/evidence-store.mjs +82 -1
- package/lib/hierarchical-context.mjs +215 -0
- package/lib/memory-engine.mjs +465 -0
- package/lib/model-performance.mjs +33 -8
- package/lib/model-policy.mjs +3 -3
- package/lib/orchestrator-policy.mjs +5 -209
- package/lib/performance-fabric.mjs +229 -0
- package/lib/pi-rpc-pool.mjs +433 -0
- package/lib/process-hang-detector.mjs +83 -0
- package/lib/process-supervisor.mjs +193 -0
- package/lib/prompt-cache.mjs +2 -0
- package/lib/repo-graph.mjs +53 -2
- package/lib/runtime-config.mjs +31 -0
- package/lib/safety.mjs +132 -0
- package/lib/semantic-index.mjs +52 -3
- package/lib/skill-compiler.mjs +128 -0
- package/lib/skill-quality.mjs +48 -2
- package/lib/task-engine.mjs +66 -5
- package/lib/task-policy.mjs +235 -0
- package/lib/verification-broker.mjs +284 -0
- package/lib/verification-command.mjs +111 -0
- package/lib/windows-shim.mjs +35 -0
- package/lib/workspace-fingerprint.mjs +198 -0
- package/package.json +52 -42
- package/pi/extensions/ues-child-runtime.ts +238 -0
- package/pi/extensions/ues.ts +3200 -0
- package/pi/prompts/ues-audit.md +9 -0
- package/pi/prompts/ues-critique.md +9 -0
- package/pi/prompts/ues-debug.md +9 -0
- package/pi/prompts/ues-feature.md +9 -0
- package/pi/prompts/ues-fix.md +9 -0
- package/pi/prompts/ues-plan.md +9 -0
- package/pi/prompts/ues-research.md +9 -0
- package/pi/prompts/ues-resume.md +9 -0
- package/pi/prompts/ues-review.md +7 -0
- package/pi/prompts/ues-run.md +17 -0
- package/pi/prompts/ues-verify.md +9 -0
- package/scripts/check-release-consistency.mjs +119 -185
- package/scripts/check-runtime-exports.mjs +66 -0
- package/scripts/check-source-integrity.mjs +184 -0
- package/scripts/eval-pi.mjs +492 -0
- package/scripts/install.mjs +16 -0
- package/scripts/smoke-package-closure.mjs +110 -0
- package/scripts/smoke-packed-install.mjs +24 -11
- package/scripts/smoke-pi-extension.mjs +144 -0
- package/scripts/uninstall.mjs +44 -0
- package/CHANGELOG.md +0 -415
- package/docs/DETERMINISTIC-TOOLS.md +0 -105
- package/docs/ENGINEERING-DESIGN.md +0 -194
- package/docs/EVALS.md +0 -158
- package/docs/GITHUB-RULESET.md +0 -50
- package/docs/NPM-PUBLISH.md +0 -116
- package/docs/RESEARCH-SOURCES.md +0 -37
- package/docs/TRACE-SCHEMA.md +0 -122
- package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +0 -75
- package/docs/V11-PERCEPTION-ADAPTIVE.md +0 -220
- package/docs/V12-WEAK-MODEL-INTELLIGENCE.md +0 -27
- package/docs/V13-PARALLEL-WEAK-MODEL-RUNTIME.md +0 -86
- package/docs/V7-INTELLIGENCE-RUNTIME.md +0 -166
- package/docs/V8-INTELLIGENCE-RELIABILITY.md +0 -206
- package/docs/V9-SPEED-INTELLIGENCE.md +0 -102
|
@@ -1,194 +0,0 @@
|
|
|
1
|
-
# UES engineering design
|
|
2
|
-
|
|
3
|
-
V11 evolves the project from an engineering workflow harness into a **perception-aware adaptive execution engine** designed to reduce context pressure on coding models while keeping evidence needed for correctness.
|
|
4
|
-
|
|
5
|
-
The selected model remains the selected model. UES improves orchestration, evidence, task boundaries, state persistence and verification; it does not claim model equivalence.
|
|
6
|
-
|
|
7
|
-
## Core design
|
|
8
|
-
|
|
9
|
-
### Thin orchestrator, durable artifacts, fresh workers
|
|
10
|
-
|
|
11
|
-
For long work:
|
|
12
|
-
|
|
13
|
-
```text
|
|
14
|
-
main/orchestrator
|
|
15
|
-
↓
|
|
16
|
-
SPEC + PLAN + STATE
|
|
17
|
-
↓
|
|
18
|
-
fresh executor per approved task
|
|
19
|
-
↓
|
|
20
|
-
task report + evidence
|
|
21
|
-
↓
|
|
22
|
-
integration verifier
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
Conversation history is not the source of truth. Durable artifacts are.
|
|
26
|
-
|
|
27
|
-
### Deterministic facts before probabilistic reasoning
|
|
28
|
-
|
|
29
|
-
UES moves cheap/reliable work into code:
|
|
30
|
-
|
|
31
|
-
- stack/test-command discovery
|
|
32
|
-
- repository import graph
|
|
33
|
-
- bounded impact search
|
|
34
|
-
- Git state
|
|
35
|
-
- changed-file review coverage
|
|
36
|
-
- risk hints
|
|
37
|
-
- verification recommendations
|
|
38
|
-
- task dependency validation
|
|
39
|
-
- safe-wave scheduling
|
|
40
|
-
- persistent work state
|
|
41
|
-
- completion gates
|
|
42
|
-
|
|
43
|
-
Models still reason about semantics and read affected code.
|
|
44
|
-
|
|
45
|
-
### Progressive disclosure
|
|
46
|
-
|
|
47
|
-
The catalog spans 48 skills. UES prefers a small active skill set and loads deeper references only when needed.
|
|
48
|
-
|
|
49
|
-
### Hard gates, not reminders
|
|
50
|
-
|
|
51
|
-
V11 machine-enforces the important boundaries:
|
|
52
|
-
|
|
53
|
-
1. long/high-risk plans are not executable until a structured plan-verification receipt matches the current plan hash;
|
|
54
|
-
2. long/high-risk task completion requires a successful verification receipt for the active run and the current workspace fingerprint;
|
|
55
|
-
3. durable state/evidence mutations are serialized with a per-work-item lock and atomic replacement;
|
|
56
|
-
4. integration PASS for strict work requires a structured integration receipt bound to the current workspace fingerprint;
|
|
57
|
-
5. finalization requires PASS and rejects any later workspace change.
|
|
58
|
-
|
|
59
|
-
These checks do not depend on a model remembering an instruction.
|
|
60
|
-
|
|
61
|
-
## Long-task state model
|
|
62
|
-
|
|
63
|
-
```text
|
|
64
|
-
.ues-work/<slug>/
|
|
65
|
-
SPEC.md
|
|
66
|
-
PLAN.json
|
|
67
|
-
STATE.json
|
|
68
|
-
EVIDENCE.json
|
|
69
|
-
EVENTS.jsonl
|
|
70
|
-
tasks/
|
|
71
|
-
reports/
|
|
72
|
-
```
|
|
73
|
-
|
|
74
|
-
The state contains operational facts only: task status, attempts, decisions, blockers, approvals and evidence. It is not chain-of-thought.
|
|
75
|
-
|
|
76
|
-
`PLAN.json` is validated for task IDs, dependencies, cycles, acceptance criteria, verification and risk. Safe waves serialize overlapping or unknown declared file scopes.
|
|
77
|
-
|
|
78
|
-
## Fresh-context execution
|
|
79
|
-
|
|
80
|
-
On OpenCode V2, the managed plugin exposes `ues.dispatch_task`.
|
|
81
|
-
|
|
82
|
-
It:
|
|
83
|
-
|
|
84
|
-
1. calls the state engine to start one ready task;
|
|
85
|
-
2. obtains the bounded Context Manifest v3 pack;
|
|
86
|
-
3. resolves the configured model tier for the executor attempt;
|
|
87
|
-
4. optionally isolates a concurrent writer in a Git worktree;
|
|
88
|
-
5. creates a fresh OpenCode session rooted at the execution directory;
|
|
89
|
-
6. binds the session ID to the durable lease;
|
|
90
|
-
7. switches to `ues-executor` and optionally to the configured model;
|
|
91
|
-
8. prompts exactly the approved task;
|
|
92
|
-
9. heartbeats while waiting under a bounded timeout;
|
|
93
|
-
10. interrupts the child on timeout/cancel and returns a bounded report.
|
|
94
|
-
|
|
95
|
-
The parent is still responsible for inspecting the child diff and recording completion/failure evidence.
|
|
96
|
-
|
|
97
|
-
## Model escalation
|
|
98
|
-
|
|
99
|
-
Roles map to `light`, `standard`, or `heavy`. Attempt number can raise a role one tier up to the configured cap. Model IDs are always user-configured; UES never invents provider/model identifiers.
|
|
100
|
-
|
|
101
|
-
Escalation follows evidence, not panic:
|
|
102
|
-
|
|
103
|
-
```text
|
|
104
|
-
attempt 1 fails
|
|
105
|
-
→ diagnose
|
|
106
|
-
→ fresh retry
|
|
107
|
-
→ stronger tier if configured
|
|
108
|
-
→ repeated causal failure
|
|
109
|
-
→ re-plan / architecture review
|
|
110
|
-
```
|
|
111
|
-
|
|
112
|
-
## OpenCode runtime
|
|
113
|
-
|
|
114
|
-
### V1
|
|
115
|
-
|
|
116
|
-
Uses compatible file-based skills, commands and agents. V2-only plugin behavior is not installed.
|
|
117
|
-
|
|
118
|
-
### V2
|
|
119
|
-
|
|
120
|
-
The managed plugin uses current V2 domains for:
|
|
121
|
-
|
|
122
|
-
- prompt admission skill routing
|
|
123
|
-
- model-context guardrails
|
|
124
|
-
- permission evaluation
|
|
125
|
-
- custom tools
|
|
126
|
-
- fresh session creation
|
|
127
|
-
- agent/model switching
|
|
128
|
-
- session waiting
|
|
129
|
-
|
|
130
|
-
Only UES-managed resources are rewritten/removed.
|
|
131
|
-
|
|
132
|
-
## Evaluation architecture
|
|
133
|
-
|
|
134
|
-
UES separates:
|
|
135
|
-
|
|
136
|
-
1. **static skill contract** — 43 scenarios covering the 48-skill catalog;
|
|
137
|
-
2. **V2 router precision matrix** — 120 required-route/negative-guard cases;
|
|
138
|
-
3. **standard live benchmark** — 20 executable hidden-graded tasks;
|
|
139
|
-
4. **long-horizon benchmark** — 5 tasks, including one 15-source-file integration workload;
|
|
140
|
-
5. **polyglot benchmark** — 8 tasks spanning Python, Java, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated contracts;
|
|
141
|
-
6. **matrix runner** — baseline vs UES across multiple suites/trials with coverage validation.
|
|
142
|
-
|
|
143
|
-
For long-suite UES mode, final behavior alone is insufficient. A PASS also requires a completed durable work item with plan approval, at least two tasks, attempted/completed task records, integration PASS and finalization evidence.
|
|
144
|
-
|
|
145
|
-
This prevents a strong model from bypassing the architecture and still being counted as proof that the long-horizon engine worked.
|
|
146
|
-
|
|
147
|
-
## Package lifecycle and release safety
|
|
148
|
-
|
|
149
|
-
npm owns package installation. UES owns only its marked/namespaced OpenCode resources.
|
|
150
|
-
|
|
151
|
-
CI validates syntax, resource contracts, routing, hidden graders, unit tests, package contents and packed global installation. Update logic resolves npm's explicit `latest` tag and refuses accidental downgrade.
|
|
152
|
-
|
|
153
|
-
## Deliberate limits
|
|
154
|
-
|
|
155
|
-
UES deliberately avoids:
|
|
156
|
-
|
|
157
|
-
- hundreds of agents/tools
|
|
158
|
-
- loading every skill
|
|
159
|
-
- treating keyword routing as truth
|
|
160
|
-
- hidden chain-of-thought storage
|
|
161
|
-
- automatic merge/push/publish/deploy
|
|
162
|
-
- claiming success without fresh evidence
|
|
163
|
-
- claiming that one benchmark proves general model equivalence
|
|
164
|
-
|
|
165
|
-
The target is a small number of strong control loops: correct context, small tasks, durable state, deterministic checks and independent verification.
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
## V7.7 intelligence runtime additions
|
|
169
|
-
|
|
170
|
-
V7 adds four control loops around the V6 state machine:
|
|
171
|
-
|
|
172
|
-
1. **runtime reliability** — task attempts carry run fencing, heartbeats and leases; expired `running` state can be recovered after process/session interruption;
|
|
173
|
-
2. **evidence binding** — verification commands can emit structured receipts containing exit status, output digests and before/after workspace fingerprints;
|
|
174
|
-
3. **context intelligence** — fresh executors receive a bounded manifest of declared files, import neighbors, likely tests, instruction/manifests and accepted learnings;
|
|
175
|
-
4. **adaptive policy + learning** — deterministic task risk/complexity influences workflow/model tier and eval traces can produce explicit learning proposals.
|
|
176
|
-
|
|
177
|
-
Parallelism is now read/write aware. Read/read overlap can share a wave; writers serialize against readers/writers unless the parent intentionally moves them into isolated Git worktree sandboxes.
|
|
178
|
-
|
|
179
|
-
The V2 plugin probes actual session capabilities before dispatch rather than treating a major version number as sufficient proof that every runtime API exists.
|
|
180
|
-
|
|
181
|
-
Hermes is deliberately adapter-only. UES can detect Hermes and generate a bounded task handoff, but does not embed Hermes' runtime, memory, scheduler or gateway into core.
|
|
182
|
-
|
|
183
|
-
## V8 intelligence and reliability additions
|
|
184
|
-
|
|
185
|
-
V8 adds six control loops around the V7 runtime:
|
|
186
|
-
|
|
187
|
-
1. **hard evidence binding** — structured plan/integration receipts plus current-workspace task receipts;
|
|
188
|
-
2. **bounded executor lifecycle** — session binding, timeout interrupt, cancellation and task-scoped recovery;
|
|
189
|
-
3. **event sourcing for observability** — append-only `EVENTS.jsonl` alongside snapshot state;
|
|
190
|
-
4. **context manifest v3** — Git-change awareness, symbol hits, TF-IDF-style ranking and centered excerpts;
|
|
191
|
-
5. **safe parallel integration** — isolated worktrees with dirty-root conflict refusal and UES-only branch cleanup;
|
|
192
|
-
6. **benchmark-gated learning** — clustered proposals require explicit acceptance and measured shadow improvement before retrieval.
|
|
193
|
-
|
|
194
|
-
The local Control Center remains a safe local observer/controller. It can inspect receipts/events and request stale recovery, but it does not bypass plan, verification, safety or finalization gates and does not directly own OpenCode sessions.
|
package/docs/EVALS.md
DELETED
|
@@ -1,158 +0,0 @@
|
|
|
1
|
-
# UES evaluations
|
|
2
|
-
|
|
3
|
-
V11 separates catalog correctness, routing precision, benchmark integrity, final behavior, long-horizon orchestration and cross-stack coverage.
|
|
4
|
-
|
|
5
|
-
## 1. Static skill-routing contract
|
|
6
|
-
|
|
7
|
-
`evals/routing.json` keeps 43 representative scenarios and covers all installed skills.
|
|
8
|
-
|
|
9
|
-
```bash
|
|
10
|
-
npm run evals
|
|
11
|
-
# or
|
|
12
|
-
ocskill eval
|
|
13
|
-
```
|
|
14
|
-
|
|
15
|
-
This checks catalog consistency, not model behavior.
|
|
16
|
-
|
|
17
|
-
## 2. V2 router trigger matrix
|
|
18
|
-
|
|
19
|
-
`evals/router-triggers.json` contains **120 cases** spanning positive routes, negative guards and wording variations.
|
|
20
|
-
|
|
21
|
-
```bash
|
|
22
|
-
npm run evals:router
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
The evaluator reports required-route recall and negative-guard success. It exists to catch deterministic router drift separately from LLM behavior.
|
|
26
|
-
|
|
27
|
-
## 3. Standard live-suite integrity
|
|
28
|
-
|
|
29
|
-
The standard live suite contains **20 executable hidden-graded tasks**.
|
|
30
|
-
|
|
31
|
-
```bash
|
|
32
|
-
npm run evals:live:validate
|
|
33
|
-
```
|
|
34
|
-
|
|
35
|
-
Each grader must reject its intentionally broken fixture with an assertion failure. A grader that already passes or fails for unrelated setup reasons invalidates the suite.
|
|
36
|
-
|
|
37
|
-
## 4. Long-horizon suite integrity
|
|
38
|
-
|
|
39
|
-
The long suite contains **5 tasks**. Four exercise coordinated 3–4 file domains; one combines all four domains into a **15-source-file** integration workload.
|
|
40
|
-
|
|
41
|
-
```bash
|
|
42
|
-
npm run evals:long:validate
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
The same broken-fixture rule applies.
|
|
46
|
-
|
|
47
|
-
## 5. Polyglot suite integrity
|
|
48
|
-
|
|
49
|
-
The polyglot suite contains **8 tasks** covering Python, Java, .NET, Next.js, React Native, SQL migration, monorepo boundaries and generated contract discipline.
|
|
50
|
-
|
|
51
|
-
```bash
|
|
52
|
-
npm run evals:polyglot:validate
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
The hidden graders must reject the intentionally broken fixture before the suite is considered valid.
|
|
56
|
-
|
|
57
|
-
## 6. Live baseline vs UES
|
|
58
|
-
|
|
59
|
-
```bash
|
|
60
|
-
ocskill eval-live --model provider/model --trials 3
|
|
61
|
-
ocskill eval-live --suite long --model provider/model --trials 3
|
|
62
|
-
ocskill eval-live --suite polyglot --model provider/model --trials 3
|
|
63
|
-
```
|
|
64
|
-
|
|
65
|
-
Each task runs as:
|
|
66
|
-
|
|
67
|
-
- **baseline** — isolated empty OpenCode config
|
|
68
|
-
- **ues** — same model/task with repository UES resources installed
|
|
69
|
-
|
|
70
|
-
Use multiple trials because coding-agent behavior is nondeterministic.
|
|
71
|
-
|
|
72
|
-
### Long-suite orchestration gate
|
|
73
|
-
|
|
74
|
-
For `--suite long`, a UES-mode result is PASS only if:
|
|
75
|
-
|
|
76
|
-
1. OpenCode agent process exits successfully;
|
|
77
|
-
2. hidden behavior grader passes;
|
|
78
|
-
3. at least one `.ues-work/<slug>/` item is valid;
|
|
79
|
-
4. the plan contains at least two tasks;
|
|
80
|
-
5. plan approval status is `passed` and contains a structured plan-verification receipt for the current plan hash;
|
|
81
|
-
6. every planned task has an attempt and ends `completed`;
|
|
82
|
-
7. every task is backed by a successful verification receipt;
|
|
83
|
-
8. integration verification is `PASS` and contains a structured integration-verification receipt for the verified workspace fingerprint;
|
|
84
|
-
9. integration and finalization evidence exist;
|
|
85
|
-
10. work item status is `completed`.
|
|
86
|
-
|
|
87
|
-
Therefore a model that directly patches all files in its main context but bypasses the long-task engine is not counted as a successful UES long-horizon run.
|
|
88
|
-
|
|
89
|
-
## 7. Benchmark matrix
|
|
90
|
-
|
|
91
|
-
Run all three behavioral suites in baseline and UES mode:
|
|
92
|
-
|
|
93
|
-
```bash
|
|
94
|
-
npm run evals:matrix -- --model provider/model --trials 3
|
|
95
|
-
```
|
|
96
|
-
|
|
97
|
-
The matrix verifies that the expected number of baseline and UES runs was produced before summarizing pass-rate delta. Use `--long-only`, `--standard-only`, `--polyglot-only`, or `--without-polyglot` to narrow the matrix.
|
|
98
|
-
|
|
99
|
-
## 8. Authentication isolation
|
|
100
|
-
|
|
101
|
-
Default mode:
|
|
102
|
-
|
|
103
|
-
```text
|
|
104
|
-
--auth env-only
|
|
105
|
-
```
|
|
106
|
-
|
|
107
|
-
The harness isolates HOME, USERPROFILE, XDG config/data/cache/state and `OPENCODE_CONFIG_DIR`.
|
|
108
|
-
|
|
109
|
-
If provider auth was established through OpenCode itself:
|
|
110
|
-
|
|
111
|
-
```bash
|
|
112
|
-
ocskill eval-live --model provider/model --auth current --trials 3
|
|
113
|
-
```
|
|
114
|
-
|
|
115
|
-
`current` copies only the current auth file, not the user's global UES configuration.
|
|
116
|
-
|
|
117
|
-
## 9. Telemetry
|
|
118
|
-
|
|
119
|
-
Results may include:
|
|
120
|
-
|
|
121
|
-
- pass/fail
|
|
122
|
-
- process exit status
|
|
123
|
-
- duration
|
|
124
|
-
- changed files
|
|
125
|
-
- bounded stdout/stderr
|
|
126
|
-
- best-effort tool calls
|
|
127
|
-
- loaded skills/subagent targets
|
|
128
|
-
- token/cost data when exposed
|
|
129
|
-
- long-suite orchestration inspection
|
|
130
|
-
|
|
131
|
-
No hidden chain-of-thought is collected.
|
|
132
|
-
|
|
133
|
-
## 10. Report aggregation
|
|
134
|
-
|
|
135
|
-
```bash
|
|
136
|
-
ocskill eval-report .ues-evals
|
|
137
|
-
```
|
|
138
|
-
|
|
139
|
-
Compare the same model, variant, prompt, fixture, grader and environment. Report multiple trials.
|
|
140
|
-
|
|
141
|
-
A benchmark result is evidence only for the measured workload. UES does not claim to turn one base model into another.
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
## V11 live-run observability and evidence gate
|
|
145
|
-
|
|
146
|
-
Live runs accept:
|
|
147
|
-
|
|
148
|
-
```bash
|
|
149
|
-
--heartbeat-ms 30000
|
|
150
|
-
--idle-timeout-ms 300000
|
|
151
|
-
--timeout-ms 900000
|
|
152
|
-
```
|
|
153
|
-
|
|
154
|
-
The harness prints a start line and heartbeat for an active model run. Hard timeout and idle timeout are recorded separately. Ctrl+C aborts the active OpenCode process tree and sets exit code 130 after the current result is recorded.
|
|
155
|
-
|
|
156
|
-
V8 long-suite UES mode requires **receipt-backed verification for every planned task**, a structured plan receipt bound to the current plan hash, and a structured integration receipt bound to the verified workspace fingerprint. Strict task completion additionally rejects a successful command receipt if the workspace changed after that receipt.
|
|
157
|
-
|
|
158
|
-
This intentionally raises the benchmark bar: final code correctness + durable orchestration + current machine-observable verification are all required.
|
package/docs/GITHUB-RULESET.md
DELETED
|
@@ -1,50 +0,0 @@
|
|
|
1
|
-
# GitHub Ruleset Readiness
|
|
2
|
-
|
|
3
|
-
This document records the recommended GitHub repository ruleset configuration for the `main` branch, to be activated only after CI Gate and Security Gate have each achieved at least one successful run.
|
|
4
|
-
|
|
5
|
-
## Recommended ruleset
|
|
6
|
-
|
|
7
|
-
- **Ruleset name:** `main-protection`
|
|
8
|
-
- **Target:** default branch / `main`
|
|
9
|
-
- **Enforcement:** Active
|
|
10
|
-
|
|
11
|
-
## Rules
|
|
12
|
-
|
|
13
|
-
1. **Restrict deletions** — block deleting the default branch or force-pushing over it.
|
|
14
|
-
2. **Block force pushes** — no `git push --force` or `git push -f` to `main`.
|
|
15
|
-
3. **Require pull request before merge** — all changes must go through a PR.
|
|
16
|
-
4. **Require conversation resolution** — PR reviewers must resolve all inline comments before merge.
|
|
17
|
-
5. **Require status checks** — PRs must have passing CI Gate and Security Gate checks before merge.
|
|
18
|
-
6. **Require branch up to date** — PR branches must be up to date with `main` before merge (applies when there are multiple contributors; can be relaxed for single-maintainer setups).
|
|
19
|
-
7. **Required check: CI Gate** — the aggregate CI check from `.github/workflows/ci.yml`.
|
|
20
|
-
8. **Required check: Security Gate** — the aggregate security check from `.github/workflows/security.yml`.
|
|
21
|
-
|
|
22
|
-
## Approval policy for single-maintainer repo
|
|
23
|
-
|
|
24
|
-
This repository currently has one maintainer. Setting `required approval = 1` would prevent the owner from merging their own PRs (they cannot approve their own PR). Therefore:
|
|
25
|
-
|
|
26
|
-
- **Required approval = 0** for now.
|
|
27
|
-
- When a collaborator or external reviewer is added, raise to `required approval = 1` to restore oversight.
|
|
28
|
-
|
|
29
|
-
## Classic branch protection vs ruleset
|
|
30
|
-
|
|
31
|
-
GitHub supports both classic branch protection rules and repository rulesets simultaneously. When both are configured:
|
|
32
|
-
|
|
33
|
-
- They can apply independently to different aspects of branch protection.
|
|
34
|
-
- Be cautious of duplicate or conflicting protections (e.g., two rules both requiring status checks but with different required lists).
|
|
35
|
-
- If migrating from classic protection to ruleset, remove the classic rule after confirming the ruleset is active and working.
|
|
36
|
-
|
|
37
|
-
## Activation prerequisite
|
|
38
|
-
|
|
39
|
-
Do not enable the ruleset until:
|
|
40
|
-
|
|
41
|
-
1. CI Gate has at least one successful run on `main`.
|
|
42
|
-
2. Security Gate has at least one successful run on `main` or a PR.
|
|
43
|
-
|
|
44
|
-
Without these prerequisites, the ruleset would block all merges immediately after activation, effectively locking the repository.
|
|
45
|
-
|
|
46
|
-
## Current status
|
|
47
|
-
|
|
48
|
-
- **CI Gate:** Added in `.github/workflows/ci.yml`. Awaiting first successful run.
|
|
49
|
-
- **Security Gate:** Added in `.github/workflows/security.yml`. Awaiting first successful run.
|
|
50
|
-
- **Ruleset:** NOT YET ACTIVATED. Will be configured via GitHub admin settings after both gates have successful runs.
|
package/docs/NPM-PUBLISH.md
DELETED
|
@@ -1,116 +0,0 @@
|
|
|
1
|
-
# Publishing to npm
|
|
2
|
-
|
|
3
|
-
The package is published as:
|
|
4
|
-
|
|
5
|
-
```text
|
|
6
|
-
opencode-agent-skill
|
|
7
|
-
```
|
|
8
|
-
|
|
9
|
-
## Release prerequisites
|
|
10
|
-
|
|
11
|
-
1. The npm account must have publish rights to the unscoped `opencode-agent-skill` package name.
|
|
12
|
-
2. `package.json` and `package-lock.json` versions must match.
|
|
13
|
-
3. `CHANGELOG.md` must contain the release.
|
|
14
|
-
4. Run the complete local validation:
|
|
15
|
-
|
|
16
|
-
```cmd
|
|
17
|
-
npm run ci
|
|
18
|
-
```
|
|
19
|
-
|
|
20
|
-
CI includes syntax validation, resource validation, static skill routing, the 129-case V2 router matrix, 13 V11 contract tasks, standard/long/polyglot hidden-grader integrity checks, unit/integration tests, package dry-run, packed global-install smoke, and a plain one-command install/resource sync smoke.
|
|
21
|
-
|
|
22
|
-
## Manual release-like test
|
|
23
|
-
|
|
24
|
-
Do not use `npm install -g .` as a release simulation because npm may create a symlink/junction back to the checkout.
|
|
25
|
-
|
|
26
|
-
Use:
|
|
27
|
-
|
|
28
|
-
```cmd
|
|
29
|
-
npm pack
|
|
30
|
-
npm install -g .\opencode-agent-skill-13.0.0-beta.1.tgz --allow-scripts=opencode-agent-skill
|
|
31
|
-
ocskill install
|
|
32
|
-
ocskill status
|
|
33
|
-
ocskill doctor
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
For the current V13 beta, also open a fresh OpenCode V2 session, check `/plugins`, and call `ues.capabilities`. Native parallel should only be exercised when `freshDispatch` is `true`.
|
|
37
|
-
|
|
38
|
-
For routine development, the automated `smoke:pack` test uses an isolated npm prefix/OpenCode config so it does not replace the developer's currently installed UES.
|
|
39
|
-
|
|
40
|
-
## Manual publish
|
|
41
|
-
|
|
42
|
-
Before publishing, verify the exact prerelease version is not already present:
|
|
43
|
-
|
|
44
|
-
```cmd
|
|
45
|
-
npm view opencode-agent-skill@13.0.0-beta.1 version --registry=https://registry.npmjs.org/
|
|
46
|
-
```
|
|
47
|
-
|
|
48
|
-
If it is not present, a manual prerelease publish uses `next`, not `latest`:
|
|
49
|
-
|
|
50
|
-
```cmd
|
|
51
|
-
npm login
|
|
52
|
-
npm whoami
|
|
53
|
-
npm run ci
|
|
54
|
-
npm publish --access public --provenance --tag next
|
|
55
|
-
```
|
|
56
|
-
|
|
57
|
-
After publication verify:
|
|
58
|
-
|
|
59
|
-
```cmd
|
|
60
|
-
npm view opencode-agent-skill versions --json
|
|
61
|
-
npm view opencode-agent-skill@13.0.0-beta.1 version
|
|
62
|
-
npm dist-tag ls opencode-agent-skill
|
|
63
|
-
```
|
|
64
|
-
|
|
65
|
-
For V13 beta the expected dist-tags are:
|
|
66
|
-
|
|
67
|
-
```text
|
|
68
|
-
latest: 11.0.0
|
|
69
|
-
next: 13.0.0-beta.1
|
|
70
|
-
```
|
|
71
|
-
|
|
72
|
-
Do not move `latest` to V13 until the prerelease is intentionally promoted stable.
|
|
73
|
-
|
|
74
|
-
## GitHub Actions publishing
|
|
75
|
-
|
|
76
|
-
The repository's publish workflow is OIDC/provenance-ready and runs the same package validation before `npm publish`.
|
|
77
|
-
|
|
78
|
-
For stronger long-term supply-chain security, configure npm Trusted Publishing for:
|
|
79
|
-
|
|
80
|
-
```text
|
|
81
|
-
GitHub owner: laivannha0202
|
|
82
|
-
Repository: opencode-agent-skill-
|
|
83
|
-
Workflow: publish.yml
|
|
84
|
-
```
|
|
85
|
-
|
|
86
|
-
Then the GitHub-hosted workflow can authenticate through OIDC instead of a long-lived npm publish token. npm Trusted Publishing requires the corresponding publisher relationship to be configured on npm; repository code alone cannot create that account-side trust relationship.
|
|
87
|
-
|
|
88
|
-
The current `publish.yml` is tag-only. A matching prerelease tag such as `v13.0.0-beta.1` runs the full package gate and publishes with npm dist-tag `next`; a stable version publishes to `latest`. The workflow first checks whether that exact version already exists and skips duplicate publication.
|
|
89
|
-
|
|
90
|
-
## Release checklist
|
|
91
|
-
|
|
92
|
-
1. Confirm version/changelog/package-lock consistency.
|
|
93
|
-
2. Run `npm run ci` locally on Windows and at least one Unix-like environment when practical.
|
|
94
|
-
3. Run `npm pack` and inspect the tarball contents.
|
|
95
|
-
4. Verify packed install/state/resource counts.
|
|
96
|
-
5. When installer compatibility changed, exercise both forced V1 and V2 paths through tests.
|
|
97
|
-
6. When updater behavior changed, test explicit latest-tag resolution, equal-version behavior, and downgrade refusal.
|
|
98
|
-
7. Commit and push the release branch.
|
|
99
|
-
8. Merge only after review/local validation is clean.
|
|
100
|
-
9. Create/push the matching `vX.Y.Z` tag or run the publish workflow.
|
|
101
|
-
10. Verify registry version and dist-tags (`next` for prerelease, `latest` for stable).
|
|
102
|
-
11. Install the published package on a clean environment before announcing it.
|
|
103
|
-
|
|
104
|
-
## One-command user install
|
|
105
|
-
|
|
106
|
-
After publication:
|
|
107
|
-
|
|
108
|
-
```cmd
|
|
109
|
-
npm install -g opencode-agent-skill
|
|
110
|
-
```
|
|
111
|
-
|
|
112
|
-
If lifecycle execution is blocked by local npm policy:
|
|
113
|
-
|
|
114
|
-
```cmd
|
|
115
|
-
ocskill install
|
|
116
|
-
```
|
package/docs/RESEARCH-SOURCES.md
DELETED
|
@@ -1,37 +0,0 @@
|
|
|
1
|
-
# Research sources
|
|
2
|
-
|
|
3
|
-
UES is informed by public engineering-agent patterns and current primary platform documentation.
|
|
4
|
-
|
|
5
|
-
## Engineering-agent and skill-system references
|
|
6
|
-
|
|
7
|
-
- Alibaba OpenCodeReview: https://github.com/alibaba/open-code-review
|
|
8
|
-
- Open GSD Core: https://github.com/open-gsd/gsd-core
|
|
9
|
-
- Superpowers: https://github.com/obra/superpowers
|
|
10
|
-
- Agent Skills open specification: https://github.com/agentskills/agentskills
|
|
11
|
-
- Anthropic Skills: https://github.com/anthropics/skills
|
|
12
|
-
- OpenAI Agents SDK: https://github.com/openai/openai-agents-python
|
|
13
|
-
- NVIDIA Skills: https://github.com/NVIDIA/skills
|
|
14
|
-
- Ruflo / Claude Flow: https://github.com/ruvnet/ruflo
|
|
15
|
-
|
|
16
|
-
## OpenCode references
|
|
17
|
-
|
|
18
|
-
V1/current installed-system references:
|
|
19
|
-
- Skills: https://opencode.ai/docs/skills
|
|
20
|
-
- Agents: https://opencode.ai/docs/agents
|
|
21
|
-
- Commands: https://opencode.ai/docs/commands
|
|
22
|
-
|
|
23
|
-
OpenCode V2 compatibility/runtime references:
|
|
24
|
-
- Migration from V1: https://opencode.ai/v2/docs/migrate-v1
|
|
25
|
-
- Agent permissions: https://opencode.ai/v2/docs/permissions
|
|
26
|
-
- Agents: https://opencode.ai/v2/docs/agents
|
|
27
|
-
- Skills: https://opencode.ai/v2/docs/skills
|
|
28
|
-
- Plugin discovery/configuration: https://opencode.ai/v2/docs/plugins
|
|
29
|
-
- Plugin API and prompt hooks: https://opencode.ai/v2/docs/build/plugins
|
|
30
|
-
- V1 plugin migration: https://opencode.ai/v2/docs/build/plugins/migrate-v1
|
|
31
|
-
|
|
32
|
-
## npm release references
|
|
33
|
-
|
|
34
|
-
- Trusted Publishing: https://docs.npmjs.com/trusted-publishers/
|
|
35
|
-
- Provenance: https://docs.npmjs.com/generating-provenance-statements/
|
|
36
|
-
|
|
37
|
-
The project does not vendor or copy these projects. External documentation is used to verify platform behavior and inform UES's own implementation.
|
package/docs/TRACE-SCHEMA.md
DELETED
|
@@ -1,122 +0,0 @@
|
|
|
1
|
-
# UES evaluation trace schema
|
|
2
|
-
|
|
3
|
-
Live evaluations write machine-readable JSON under `.ues-evals/`.
|
|
4
|
-
|
|
5
|
-
Each result item contains:
|
|
6
|
-
|
|
7
|
-
- `task`
|
|
8
|
-
- `mode`: `baseline` or `ues`
|
|
9
|
-
- `model` and optional `variant`
|
|
10
|
-
- `trial`
|
|
11
|
-
- `passed`
|
|
12
|
-
- `agentExit` and `graderExit`
|
|
13
|
-
- `durationMs`
|
|
14
|
-
- `authMode`: `env-only` or `current`
|
|
15
|
-
- `changedFiles`: added/removed/modified workspace paths
|
|
16
|
-
- bounded agent/grader stdout and stderr
|
|
17
|
-
- optional kept workspace path when `--keep` is used
|
|
18
|
-
- timestamp
|
|
19
|
-
|
|
20
|
-
## Telemetry
|
|
21
|
-
|
|
22
|
-
`telemetry` currently has schema version 1:
|
|
23
|
-
|
|
24
|
-
```json
|
|
25
|
-
{
|
|
26
|
-
"schemaVersion": 1,
|
|
27
|
-
"format": "best-effort-opencode-jsonl",
|
|
28
|
-
"jsonLines": 42,
|
|
29
|
-
"parseErrors": 0,
|
|
30
|
-
"toolCalls": 12,
|
|
31
|
-
"tools": {
|
|
32
|
-
"bash": 4,
|
|
33
|
-
"read": 5,
|
|
34
|
-
"skill": 2,
|
|
35
|
-
"subagent": 1
|
|
36
|
-
},
|
|
37
|
-
"skillsLoaded": ["ues-bug-diagnosis"],
|
|
38
|
-
"subagents": ["ues-verifier"],
|
|
39
|
-
"tokens": {
|
|
40
|
-
"input": 12000,
|
|
41
|
-
"output": 2200,
|
|
42
|
-
"total": 14200
|
|
43
|
-
},
|
|
44
|
-
"cost": 0.18
|
|
45
|
-
}
|
|
46
|
-
```
|
|
47
|
-
|
|
48
|
-
OpenCode JSON event shapes can evolve, so telemetry extraction is best-effort. Hidden-grader correctness and process exit status remain the primary benchmark evidence.
|
|
49
|
-
|
|
50
|
-
The trace intentionally does not collect or score hidden chain-of-thought.
|
|
51
|
-
|
|
52
|
-
## Run summary
|
|
53
|
-
|
|
54
|
-
Each result file also contains:
|
|
55
|
-
- suite version
|
|
56
|
-
- auth mode
|
|
57
|
-
- selected model/variant
|
|
58
|
-
- trial count
|
|
59
|
-
- optional task filter
|
|
60
|
-
- modes executed
|
|
61
|
-
- pass counts and pass rates per mode
|
|
62
|
-
|
|
63
|
-
Use `ocskill eval-report` or `npm run evals:report -- <paths>` to aggregate multiple result files.
|
|
64
|
-
|
|
65
|
-
## Fair comparisons
|
|
66
|
-
|
|
67
|
-
Keep constant:
|
|
68
|
-
- model and variant
|
|
69
|
-
- task fixture
|
|
70
|
-
- prompt
|
|
71
|
-
- grader
|
|
72
|
-
- runtime/provider environment
|
|
73
|
-
- trial count when possible
|
|
74
|
-
|
|
75
|
-
Compare observable success, regressions, elapsed time, tool behavior and cost rather than narrative confidence.
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
## V11 runtime and evidence fields
|
|
79
|
-
|
|
80
|
-
Each live result may additionally contain:
|
|
81
|
-
|
|
82
|
-
- `timedOut`
|
|
83
|
-
- `idleTimedOut`
|
|
84
|
-
- `aborted`
|
|
85
|
-
- detected `opencodeVersion` / `opencodeMajor`
|
|
86
|
-
- configured heartbeat/hard/idle timeout values
|
|
87
|
-
- long-suite receipt coverage inside orchestration inspection
|
|
88
|
-
|
|
89
|
-
Long-task `EVIDENCE.json` schema 3 may contain `receipts` and `gateReceipts` arrays. Command receipt fields include:
|
|
90
|
-
|
|
91
|
-
```json
|
|
92
|
-
{
|
|
93
|
-
"schemaVersion": 1,
|
|
94
|
-
"id": "uuid",
|
|
95
|
-
"task": "T1",
|
|
96
|
-
"runId": "attempt-uuid",
|
|
97
|
-
"command": "npm",
|
|
98
|
-
"args": ["test"],
|
|
99
|
-
"exitCode": 0,
|
|
100
|
-
"passed": true,
|
|
101
|
-
"durationMs": 1234,
|
|
102
|
-
"stdoutSha256": "...",
|
|
103
|
-
"stderrSha256": "...",
|
|
104
|
-
"workspaceBefore": "...",
|
|
105
|
-
"workspaceAfter": "..."
|
|
106
|
-
}
|
|
107
|
-
```
|
|
108
|
-
|
|
109
|
-
Full stdout/stderr are not stored in receipts; hashes provide binding without persisting potentially sensitive logs.
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
### Structured gate receipts
|
|
113
|
-
|
|
114
|
-
Strict plan/integration gates use receipt schema version 1. A plan receipt includes the exact `planHash`; an integration receipt includes the exact `workspaceFingerprint`. Both include verifier identity, optional session/run IDs, evidence text and an optional report hash.
|
|
115
|
-
|
|
116
|
-
### Runtime event journal
|
|
117
|
-
|
|
118
|
-
Each long work item may include `EVENTS.jsonl`. Every line is an independent JSON event with schema version, UUID, event type, timestamp and task/work metadata. It records operational events only and never hidden chain-of-thought.
|
|
119
|
-
|
|
120
|
-
### Benchmark matrix summary
|
|
121
|
-
|
|
122
|
-
`scripts/eval-matrix.mjs` writes a matrix summary containing selected suites/model/trials, expected runs per mode, actual baseline/UES counts, coverage completeness and aggregated pass-rate statistics.
|