opencode-agent-skill 7.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +163 -0
- package/LICENSE +9 -0
- package/README.md +581 -0
- package/bin/ocskill.mjs +975 -0
- package/docs/DETERMINISTIC-TOOLS.md +88 -0
- package/docs/ENGINEERING-DESIGN.md +176 -0
- package/docs/EVALS.md +136 -0
- package/docs/NPM-PUBLISH.md +102 -0
- package/docs/OPENCODE-COMPAT.md +109 -0
- package/docs/RESEARCH-SOURCES.md +37 -0
- package/docs/TRACE-SCHEMA.md +109 -0
- package/docs/V7-INTELLIGENCE-RUNTIME.md +166 -0
- package/evals/live/fixtures/engineering-bench/package.json +1 -0
- package/evals/live/fixtures/engineering-bench/src/api-errors.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/authz.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/cache-tags.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/config.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/contract-consumer.mjs +6 -0
- package/evals/live/fixtures/engineering-bench/src/contract-producer.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/dedupe.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/dependency.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/discount.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/inventory.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/migration.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/money.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/pagination.mjs +5 -0
- package/evals/live/fixtures/engineering-bench/src/path-safe.mjs +5 -0
- package/evals/live/fixtures/engineering-bench/src/payment.mjs +5 -0
- package/evals/live/fixtures/engineering-bench/src/query-sort.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/react-state.mjs +8 -0
- package/evals/live/fixtures/engineering-bench/src/retry.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/rn-platform.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/upload.mjs +3 -0
- package/evals/live/fixtures/engineering-bench/src/webhook.mjs +3 -0
- package/evals/live/graders/engineering-bench.mjs +239 -0
- package/evals/live/tasks.json +126 -0
- package/evals/long/fixtures/long-horizon/package.json +1 -0
- package/evals/long/fixtures/long-horizon/src/auth.mjs +5 -0
- package/evals/long/fixtures/long-horizon/src/checkout.mjs +9 -0
- package/evals/long/fixtures/long-horizon/src/inventory.mjs +5 -0
- package/evals/long/fixtures/long-horizon/src/money.mjs +3 -0
- package/evals/long/fixtures/long-horizon/src/payment.mjs +7 -0
- package/evals/long/fixtures/long-horizon/src/product-api.mjs +10 -0
- package/evals/long/fixtures/long-horizon/src/product-cache.mjs +3 -0
- package/evals/long/fixtures/long-horizon/src/product-service.mjs +6 -0
- package/evals/long/fixtures/long-horizon/src/product-state.mjs +8 -0
- package/evals/long/fixtures/long-horizon/src/project-api.mjs +9 -0
- package/evals/long/fixtures/long-horizon/src/project-service.mjs +7 -0
- package/evals/long/fixtures/long-horizon/src/user-consumer.mjs +3 -0
- package/evals/long/fixtures/long-horizon/src/user-migration.mjs +7 -0
- package/evals/long/fixtures/long-horizon/src/user-serializer.mjs +3 -0
- package/evals/long/fixtures/long-horizon/src/user-validation.mjs +3 -0
- package/evals/long/graders/long-horizon.mjs +209 -0
- package/evals/long/tasks.json +36 -0
- package/evals/router-triggers.json +1238 -0
- package/evals/routing.json +321 -0
- package/global-config/AGENTS.md +212 -0
- package/global-config/agents/architect.md +38 -0
- package/global-config/agents/codebase-mapper.md +41 -0
- package/global-config/agents/critic.md +37 -0
- package/global-config/agents/debugger.md +36 -0
- package/global-config/agents/executor.md +48 -0
- package/global-config/agents/integration-verifier.md +43 -0
- package/global-config/agents/plan-checker.md +43 -0
- package/global-config/agents/researcher.md +30 -0
- package/global-config/agents/reviewer.md +30 -0
- package/global-config/agents/verifier.md +33 -0
- package/global-config/commands/audit.md +8 -0
- package/global-config/commands/critique.md +8 -0
- package/global-config/commands/debug.md +8 -0
- package/global-config/commands/feature.md +8 -0
- package/global-config/commands/fix.md +8 -0
- package/global-config/commands/plan.md +8 -0
- package/global-config/commands/research.md +8 -0
- package/global-config/commands/resume.md +22 -0
- package/global-config/commands/review.md +8 -0
- package/global-config/commands/run.md +29 -0
- package/global-config/commands/verify.md +8 -0
- package/global-config/plugins/ues-router/capabilities.js +20 -0
- package/global-config/plugins/ues-router/index.js +277 -0
- package/global-config/plugins/ues-router/router.js +74 -0
- package/global-config/plugins/ues-router/safety.js +16 -0
- package/global-config/skills/accessibility/SKILL.md +10 -0
- package/global-config/skills/accessibility/references/workflow.md +17 -0
- package/global-config/skills/api-contract/SKILL.md +12 -0
- package/global-config/skills/api-contract/references/workflow.md +17 -0
- package/global-config/skills/auth-security/SKILL.md +12 -0
- package/global-config/skills/auth-security/references/workflow.md +15 -0
- package/global-config/skills/bug-diagnosis/SKILL.md +28 -0
- package/global-config/skills/change-impact-analysis/SKILL.md +23 -0
- package/global-config/skills/code-review/SKILL.md +19 -0
- package/global-config/skills/context-engineering/SKILL.md +18 -0
- package/global-config/skills/context-engineering/references/large-repo.md +16 -0
- package/global-config/skills/database-engineering/SKILL.md +12 -0
- package/global-config/skills/database-engineering/references/workflow.md +15 -0
- package/global-config/skills/dependency-management/SKILL.md +12 -0
- package/global-config/skills/dependency-management/references/workflow.md +14 -0
- package/global-config/skills/devops-engineering/SKILL.md +10 -0
- package/global-config/skills/devops-engineering/references/workflow.md +11 -0
- package/global-config/skills/django-engineering/SKILL.md +10 -0
- package/global-config/skills/django-engineering/references/workflow.md +11 -0
- package/global-config/skills/documentation-engineering/SKILL.md +10 -0
- package/global-config/skills/documentation-engineering/references/workflow.md +15 -0
- package/global-config/skills/dotnet-engineering/SKILL.md +10 -0
- package/global-config/skills/dotnet-engineering/references/workflow.md +11 -0
- package/global-config/skills/ecommerce-engineering/SKILL.md +10 -0
- package/global-config/skills/ecommerce-engineering/references/workflow.md +17 -0
- package/global-config/skills/engineering-orchestrator/SKILL.md +29 -0
- package/global-config/skills/engineering-orchestrator/references/delegation.md +22 -0
- package/global-config/skills/engineering-orchestrator/references/evaluator-loop.md +18 -0
- package/global-config/skills/engineering-orchestrator/references/long-horizon.md +57 -0
- package/global-config/skills/engineering-orchestrator/references/model-escalation.md +19 -0
- package/global-config/skills/engineering-orchestrator/references/retry-policy.md +12 -0
- package/global-config/skills/engineering-orchestrator/references/routing.md +39 -0
- package/global-config/skills/engineering-orchestrator/references/verification-matrix.md +18 -0
- package/global-config/skills/fastapi-engineering/SKILL.md +10 -0
- package/global-config/skills/fastapi-engineering/references/workflow.md +13 -0
- package/global-config/skills/file-upload-engineering/SKILL.md +10 -0
- package/global-config/skills/file-upload-engineering/references/workflow.md +17 -0
- package/global-config/skills/flutter-engineering/SKILL.md +10 -0
- package/global-config/skills/flutter-engineering/references/workflow.md +13 -0
- package/global-config/skills/git-safety/SKILL.md +10 -0
- package/global-config/skills/git-safety/references/workflow.md +15 -0
- package/global-config/skills/implementation-engineer/SKILL.md +10 -0
- package/global-config/skills/implementation-engineer/references/workflow.md +14 -0
- package/global-config/skills/java-spring-engineering/SKILL.md +10 -0
- package/global-config/skills/java-spring-engineering/references/workflow.md +13 -0
- package/global-config/skills/long-task-state/SKILL.md +28 -0
- package/global-config/skills/long-task-state/references/context-ledger.md +27 -0
- package/global-config/skills/long-task-state/templates/STATE.md +48 -0
- package/global-config/skills/nestjs-engineering/SKILL.md +10 -0
- package/global-config/skills/nestjs-engineering/references/workflow.md +11 -0
- package/global-config/skills/nextjs-engineering/SKILL.md +12 -0
- package/global-config/skills/nextjs-engineering/references/workflow.md +15 -0
- package/global-config/skills/nodejs-engineering/SKILL.md +12 -0
- package/global-config/skills/nodejs-engineering/references/workflow.md +11 -0
- package/global-config/skills/payment-engineering/SKILL.md +12 -0
- package/global-config/skills/payment-engineering/references/workflow.md +19 -0
- package/global-config/skills/performance-engineering/SKILL.md +10 -0
- package/global-config/skills/performance-engineering/references/workflow.md +17 -0
- package/global-config/skills/python-engineering/SKILL.md +10 -0
- package/global-config/skills/python-engineering/references/workflow.md +11 -0
- package/global-config/skills/react-engineering/SKILL.md +14 -0
- package/global-config/skills/react-engineering/references/workflow.md +16 -0
- package/global-config/skills/react-native-engineering/SKILL.md +14 -0
- package/global-config/skills/react-native-engineering/references/workflow.md +16 -0
- package/global-config/skills/repo-explorer/SKILL.md +17 -0
- package/global-config/skills/research-verification/SKILL.md +18 -0
- package/global-config/skills/research-verification/references/source-hierarchy.md +12 -0
- package/global-config/skills/rest-api-design/SKILL.md +10 -0
- package/global-config/skills/rest-api-design/references/workflow.md +18 -0
- package/global-config/skills/software-architect/SKILL.md +10 -0
- package/global-config/skills/software-architect/references/workflow.md +18 -0
- package/global-config/skills/task-planner/SKILL.md +21 -0
- package/global-config/skills/task-planner/references/plan-schema.md +56 -0
- package/global-config/skills/test-driven-development/SKILL.md +22 -0
- package/global-config/skills/test-driven-development/references/writing-good-tests.md +21 -0
- package/global-config/skills/test-verification/SKILL.md +20 -0
- package/global-config/skills/ui-ux-engineering/SKILL.md +10 -0
- package/global-config/skills/ui-ux-engineering/references/workflow.md +17 -0
- package/global-config/skills/web-security-review/SKILL.md +10 -0
- package/global-config/skills/web-security-review/references/workflow.md +22 -0
- package/lib/context-manifest.mjs +101 -0
- package/lib/control-center.mjs +148 -0
- package/lib/eval-auth.mjs +21 -0
- package/lib/eval-report.mjs +72 -0
- package/lib/eval-telemetry.mjs +155 -0
- package/lib/evidence-receipt.mjs +50 -0
- package/lib/hermes-bridge.mjs +28 -0
- package/lib/ids.mjs +21 -0
- package/lib/installer.mjs +636 -0
- package/lib/learning-engine.mjs +161 -0
- package/lib/model-config.mjs +88 -0
- package/lib/model-policy.mjs +71 -0
- package/lib/opencode-compat.mjs +86 -0
- package/lib/orchestrator-policy.mjs +35 -0
- package/lib/process-runner.mjs +117 -0
- package/lib/repo-graph.mjs +135 -0
- package/lib/repo-inspect.mjs +264 -0
- package/lib/review-scope.mjs +98 -0
- package/lib/router-config.mjs +40 -0
- package/lib/task-engine.mjs +777 -0
- package/lib/task-graph.mjs +262 -0
- package/lib/update-resolver.mjs +43 -0
- package/lib/verification-plan.mjs +52 -0
- package/lib/version.mjs +47 -0
- package/lib/workspace-snapshot.mjs +45 -0
- package/lib/worktree-sandbox.mjs +59 -0
- package/package.json +69 -0
- package/scripts/check-working-tree.mjs +2 -0
- package/scripts/collect-evidence.mjs +2 -0
- package/scripts/control-center.mjs +37 -0
- package/scripts/detect-stack.mjs +2 -0
- package/scripts/detect-test-commands.mjs +2 -0
- package/scripts/eval-live.mjs +437 -0
- package/scripts/eval-report.mjs +49 -0
- package/scripts/eval-router.mjs +45 -0
- package/scripts/eval-skills.mjs +71 -0
- package/scripts/impact-map.mjs +4 -0
- package/scripts/install.mjs +61 -0
- package/scripts/repo-map.mjs +2 -0
- package/scripts/smoke-packed-install.mjs +326 -0
- package/scripts/syntax-check.mjs +35 -0
- package/scripts/uninstall.mjs +21 -0
- package/scripts/validate-live-suite.mjs +65 -0
- package/scripts/validate.mjs +133 -0
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
# Deterministic evidence and execution tools
|
|
2
|
+
|
|
3
|
+
UES 6 uses dependency-light Node helpers for work that should not rely on a model guessing or remembering it.
|
|
4
|
+
|
|
5
|
+
## Repository evidence
|
|
6
|
+
|
|
7
|
+
```cmd
|
|
8
|
+
ocskill inspect .
|
|
9
|
+
ocskill detect-stack .
|
|
10
|
+
ocskill detect-tests .
|
|
11
|
+
ocskill impact calculateOrderTotal .
|
|
12
|
+
ocskill evidence .
|
|
13
|
+
ocskill working-tree .
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
These identify stack/package manager, project-native checks, bounded impact hits and Git state.
|
|
17
|
+
|
|
18
|
+
## Repository graph
|
|
19
|
+
|
|
20
|
+
```cmd
|
|
21
|
+
ocskill repo-graph .
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Builds a bounded import graph, local edges, external import frequencies and coupling hotspots. It is not a full language server/call graph.
|
|
25
|
+
|
|
26
|
+
## Review scope
|
|
27
|
+
|
|
28
|
+
```cmd
|
|
29
|
+
ocskill review-scope main .
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
Enumerates changed files and deterministic risk hints for persistence/schema, auth/security, payments, public interfaces, dependencies and delivery/infrastructure.
|
|
33
|
+
|
|
34
|
+
`coverageRequired` lets a reviewer account for every changed file instead of relying on memory.
|
|
35
|
+
|
|
36
|
+
## Verification plan
|
|
37
|
+
|
|
38
|
+
```cmd
|
|
39
|
+
ocskill verification-plan .
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Combines project-native commands, working-tree evidence and changed-file risk into recommended checks plus risk-specific acceptance prompts.
|
|
43
|
+
|
|
44
|
+
## Task graph
|
|
45
|
+
|
|
46
|
+
```cmd
|
|
47
|
+
ocskill task-graph PLAN.json
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
Validates plan shape/dependencies/cycles and computes topological and safe waves. Same-wave tasks with overlapping/unknown declared files are serialized.
|
|
51
|
+
|
|
52
|
+
## Durable work state
|
|
53
|
+
|
|
54
|
+
```cmd
|
|
55
|
+
ocskill work init <slug> . --goal "..."
|
|
56
|
+
ocskill work plan <slug> PLAN.json .
|
|
57
|
+
ocskill work approve-plan <slug> . --evidence "..."
|
|
58
|
+
ocskill work start <slug> T1 .
|
|
59
|
+
ocskill work complete <slug> T1 . --evidence "..."
|
|
60
|
+
ocskill work fail <slug> T1 . --reason "..."
|
|
61
|
+
ocskill work verify-integration <slug> . --verdict PASS --evidence "..."
|
|
62
|
+
ocskill work finalize <slug> . --evidence "..."
|
|
63
|
+
ocskill work resume <slug> .
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
State/evidence writes use a per-item lock and atomic replacement.
|
|
67
|
+
|
|
68
|
+
## Context pack
|
|
69
|
+
|
|
70
|
+
```cmd
|
|
71
|
+
ocskill context-pack <slug> <task> .
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Returns only the task, bounded spec, dependency reports, decisions, blockers and current task state required for a fresh executor.
|
|
75
|
+
|
|
76
|
+
## Runtime dispatch on OpenCode V2
|
|
77
|
+
|
|
78
|
+
The managed plugin exposes `ues.dispatch_task`, which combines `work start`, context pack, model policy and a fresh OpenCode executor session.
|
|
79
|
+
|
|
80
|
+
## Constraints
|
|
81
|
+
|
|
82
|
+
These helpers:
|
|
83
|
+
|
|
84
|
+
- do not replace reading exact affected code
|
|
85
|
+
- do not pretend text/import scans are complete semantic analysis
|
|
86
|
+
- do not auto-merge/push/publish/deploy
|
|
87
|
+
- preserve unrelated user work
|
|
88
|
+
- use JSON outputs where machine consumption matters
|
|
@@ -0,0 +1,176 @@
|
|
|
1
|
+
# UES engineering design
|
|
2
|
+
|
|
3
|
+
UES 6 evolves the project from an engineering workflow harness into a **long-horizon execution engine** designed to reduce context pressure on coding models.
|
|
4
|
+
|
|
5
|
+
The selected model remains the selected model. UES improves orchestration, evidence, task boundaries, state persistence and verification; it does not claim model equivalence.
|
|
6
|
+
|
|
7
|
+
## Core design
|
|
8
|
+
|
|
9
|
+
### Thin orchestrator, durable artifacts, fresh workers
|
|
10
|
+
|
|
11
|
+
For long work:
|
|
12
|
+
|
|
13
|
+
```text
|
|
14
|
+
main/orchestrator
|
|
15
|
+
↓
|
|
16
|
+
SPEC + PLAN + STATE
|
|
17
|
+
↓
|
|
18
|
+
fresh executor per approved task
|
|
19
|
+
↓
|
|
20
|
+
task report + evidence
|
|
21
|
+
↓
|
|
22
|
+
integration verifier
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
Conversation history is not the source of truth. Durable artifacts are.
|
|
26
|
+
|
|
27
|
+
### Deterministic facts before probabilistic reasoning
|
|
28
|
+
|
|
29
|
+
UES moves cheap/reliable work into code:
|
|
30
|
+
|
|
31
|
+
- stack/test-command discovery
|
|
32
|
+
- repository import graph
|
|
33
|
+
- bounded impact search
|
|
34
|
+
- Git state
|
|
35
|
+
- changed-file review coverage
|
|
36
|
+
- risk hints
|
|
37
|
+
- verification recommendations
|
|
38
|
+
- task dependency validation
|
|
39
|
+
- safe-wave scheduling
|
|
40
|
+
- persistent work state
|
|
41
|
+
- completion gates
|
|
42
|
+
|
|
43
|
+
Models still reason about semantics and read affected code.
|
|
44
|
+
|
|
45
|
+
### Progressive disclosure
|
|
46
|
+
|
|
47
|
+
The catalog remains 39 skills. UES prefers a small active skill set and loads deeper references only when needed.
|
|
48
|
+
|
|
49
|
+
### Hard gates, not reminders
|
|
50
|
+
|
|
51
|
+
Three V6 gates are machine-enforced:
|
|
52
|
+
|
|
53
|
+
1. imported plans are not executable until plan-checker PASS is recorded;
|
|
54
|
+
2. durable mutations are serialized with a per-work-item lock and atomic replacement;
|
|
55
|
+
3. finalization requires integration PASS and an unchanged workspace fingerprint.
|
|
56
|
+
|
|
57
|
+
These checks do not depend on a model remembering an instruction.
|
|
58
|
+
|
|
59
|
+
## Long-task state model
|
|
60
|
+
|
|
61
|
+
```text
|
|
62
|
+
.ues-work/<slug>/
|
|
63
|
+
SPEC.md
|
|
64
|
+
PLAN.json
|
|
65
|
+
STATE.json
|
|
66
|
+
EVIDENCE.json
|
|
67
|
+
tasks/
|
|
68
|
+
reports/
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
The state contains operational facts only: task status, attempts, decisions, blockers, approvals and evidence. It is not chain-of-thought.
|
|
72
|
+
|
|
73
|
+
`PLAN.json` is validated for task IDs, dependencies, cycles, acceptance criteria, verification and risk. Safe waves serialize overlapping or unknown declared file scopes.
|
|
74
|
+
|
|
75
|
+
## Fresh-context execution
|
|
76
|
+
|
|
77
|
+
On OpenCode V2, the managed plugin exposes `ues.dispatch_task`.
|
|
78
|
+
|
|
79
|
+
It:
|
|
80
|
+
|
|
81
|
+
1. calls the state engine to start one ready task;
|
|
82
|
+
2. obtains the bounded context pack;
|
|
83
|
+
3. resolves the configured model tier for the executor attempt;
|
|
84
|
+
4. creates a fresh OpenCode session;
|
|
85
|
+
5. switches to `ues-executor`;
|
|
86
|
+
6. optionally switches to the configured model;
|
|
87
|
+
7. prompts exactly the approved task;
|
|
88
|
+
8. waits and returns child-session context.
|
|
89
|
+
|
|
90
|
+
The parent is still responsible for inspecting the child diff and recording completion/failure evidence.
|
|
91
|
+
|
|
92
|
+
## Model escalation
|
|
93
|
+
|
|
94
|
+
Roles map to `light`, `standard`, or `heavy`. Attempt number can raise a role one tier up to the configured cap. Model IDs are always user-configured; UES never invents provider/model identifiers.
|
|
95
|
+
|
|
96
|
+
Escalation follows evidence, not panic:
|
|
97
|
+
|
|
98
|
+
```text
|
|
99
|
+
attempt 1 fails
|
|
100
|
+
→ diagnose
|
|
101
|
+
→ fresh retry
|
|
102
|
+
→ stronger tier if configured
|
|
103
|
+
→ repeated causal failure
|
|
104
|
+
→ re-plan / architecture review
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
## OpenCode runtime
|
|
108
|
+
|
|
109
|
+
### V1
|
|
110
|
+
|
|
111
|
+
Uses compatible file-based skills, commands and agents. V2-only plugin behavior is not installed.
|
|
112
|
+
|
|
113
|
+
### V2
|
|
114
|
+
|
|
115
|
+
The managed plugin uses current V2 domains for:
|
|
116
|
+
|
|
117
|
+
- prompt admission skill routing
|
|
118
|
+
- model-context guardrails
|
|
119
|
+
- permission evaluation
|
|
120
|
+
- custom tools
|
|
121
|
+
- fresh session creation
|
|
122
|
+
- agent/model switching
|
|
123
|
+
- session waiting
|
|
124
|
+
|
|
125
|
+
Only UES-managed resources are rewritten/removed.
|
|
126
|
+
|
|
127
|
+
## Evaluation architecture
|
|
128
|
+
|
|
129
|
+
UES separates:
|
|
130
|
+
|
|
131
|
+
1. **static skill contract** — 34 scenarios covering the 39-skill catalog;
|
|
132
|
+
2. **V2 router precision matrix** — 120 required-route/negative-guard cases;
|
|
133
|
+
3. **standard live benchmark** — 20 executable hidden-graded tasks;
|
|
134
|
+
4. **long-horizon benchmark** — 5 tasks, including one 15-source-file integration workload.
|
|
135
|
+
|
|
136
|
+
For long-suite UES mode, final behavior alone is insufficient. A PASS also requires a completed durable work item with plan approval, at least two tasks, attempted/completed task records, integration PASS and finalization evidence.
|
|
137
|
+
|
|
138
|
+
This prevents a strong model from bypassing the architecture and still being counted as proof that the long-horizon engine worked.
|
|
139
|
+
|
|
140
|
+
## Package lifecycle and release safety
|
|
141
|
+
|
|
142
|
+
npm owns package installation. UES owns only its marked/namespaced OpenCode resources.
|
|
143
|
+
|
|
144
|
+
CI validates syntax, resource contracts, routing, hidden graders, unit tests, package contents and packed global installation. Update logic resolves npm's explicit `latest` tag and refuses accidental downgrade.
|
|
145
|
+
|
|
146
|
+
## Deliberate limits
|
|
147
|
+
|
|
148
|
+
UES deliberately avoids:
|
|
149
|
+
|
|
150
|
+
- hundreds of agents/tools
|
|
151
|
+
- loading every skill
|
|
152
|
+
- treating keyword routing as truth
|
|
153
|
+
- hidden chain-of-thought storage
|
|
154
|
+
- automatic merge/push/publish/deploy
|
|
155
|
+
- claiming success without fresh evidence
|
|
156
|
+
- claiming that one benchmark proves general model equivalence
|
|
157
|
+
|
|
158
|
+
The target is a small number of strong control loops: correct context, small tasks, durable state, deterministic checks and independent verification.
|
|
159
|
+
|
|
160
|
+
|
|
161
|
+
## V7.7 intelligence runtime additions
|
|
162
|
+
|
|
163
|
+
V7 adds four control loops around the V6 state machine:
|
|
164
|
+
|
|
165
|
+
1. **runtime reliability** — task attempts carry run fencing, heartbeats and leases; expired `running` state can be recovered after process/session interruption;
|
|
166
|
+
2. **evidence binding** — verification commands can emit structured receipts containing exit status, output digests and before/after workspace fingerprints;
|
|
167
|
+
3. **context intelligence** — fresh executors receive a bounded manifest of declared files, import neighbors, likely tests, instruction/manifests and accepted learnings;
|
|
168
|
+
4. **adaptive policy + learning** — deterministic task risk/complexity influences workflow/model tier and eval traces can produce explicit learning proposals.
|
|
169
|
+
|
|
170
|
+
Parallelism is now read/write aware. Read/read overlap can share a wave; writers serialize against readers/writers unless the parent intentionally moves them into isolated Git worktree sandboxes.
|
|
171
|
+
|
|
172
|
+
The V2 plugin probes actual session capabilities before dispatch rather than treating a major version number as sufficient proof that every runtime API exists.
|
|
173
|
+
|
|
174
|
+
Hermes is deliberately adapter-only. UES can detect Hermes and generate a bounded task handoff, but does not embed Hermes' runtime, memory, scheduler or gateway into core.
|
|
175
|
+
|
|
176
|
+
The local Control Center is observational. It reads durable artifacts and eval summaries; it does not bypass plan, verification, safety or finalization gates.
|
package/docs/EVALS.md
ADDED
|
@@ -0,0 +1,136 @@
|
|
|
1
|
+
# UES evaluations
|
|
2
|
+
|
|
3
|
+
UES 6 separates catalog correctness, routing precision, benchmark integrity, final behavior and long-horizon orchestration.
|
|
4
|
+
|
|
5
|
+
## 1. Static skill-routing contract
|
|
6
|
+
|
|
7
|
+
`evals/routing.json` keeps 34 representative scenarios and covers all installed skills.
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
npm run evals
|
|
11
|
+
# or
|
|
12
|
+
ocskill eval
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
This checks catalog consistency, not model behavior.
|
|
16
|
+
|
|
17
|
+
## 2. V2 router trigger matrix
|
|
18
|
+
|
|
19
|
+
`evals/router-triggers.json` contains **120 cases** spanning positive routes, negative guards and wording variations.
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
npm run evals:router
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
The evaluator reports required-route recall and negative-guard success. It exists to catch deterministic router drift separately from LLM behavior.
|
|
26
|
+
|
|
27
|
+
## 3. Standard live-suite integrity
|
|
28
|
+
|
|
29
|
+
The standard live suite contains **20 executable hidden-graded tasks**.
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
npm run evals:live:validate
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Each grader must reject its intentionally broken fixture with an assertion failure. A grader that already passes or fails for unrelated setup reasons invalidates the suite.
|
|
36
|
+
|
|
37
|
+
## 4. Long-horizon suite integrity
|
|
38
|
+
|
|
39
|
+
The long suite contains **5 tasks**. Four exercise coordinated 3–4 file domains; one combines all four domains into a **15-source-file** integration workload.
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
npm run evals:long:validate
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
The same broken-fixture rule applies.
|
|
46
|
+
|
|
47
|
+
## 5. Live baseline vs UES
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
ocskill eval-live --model provider/model --trials 3
|
|
51
|
+
ocskill eval-live --suite long --model provider/model --trials 3
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Each task runs as:
|
|
55
|
+
|
|
56
|
+
- **baseline** — isolated empty OpenCode config
|
|
57
|
+
- **ues** — same model/task with repository UES resources installed
|
|
58
|
+
|
|
59
|
+
Use multiple trials because coding-agent behavior is nondeterministic.
|
|
60
|
+
|
|
61
|
+
### Long-suite orchestration gate
|
|
62
|
+
|
|
63
|
+
For `--suite long`, a UES-mode result is PASS only if:
|
|
64
|
+
|
|
65
|
+
1. OpenCode agent process exits successfully;
|
|
66
|
+
2. hidden behavior grader passes;
|
|
67
|
+
3. at least one `.ues-work/<slug>/` item is valid;
|
|
68
|
+
4. the plan contains at least two tasks;
|
|
69
|
+
5. plan approval status is `passed`;
|
|
70
|
+
6. every planned task has an attempt and ends `completed`;
|
|
71
|
+
7. integration verification is `PASS`;
|
|
72
|
+
8. integration and finalization evidence exist;
|
|
73
|
+
9. work item status is `completed`.
|
|
74
|
+
|
|
75
|
+
Therefore a model that directly patches all files in its main context but bypasses the long-task engine is not counted as a successful UES long-horizon run.
|
|
76
|
+
|
|
77
|
+
## Authentication isolation
|
|
78
|
+
|
|
79
|
+
Default mode:
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
--auth env-only
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
The harness isolates HOME, USERPROFILE, XDG config/data/cache/state and `OPENCODE_CONFIG_DIR`.
|
|
86
|
+
|
|
87
|
+
If provider auth was established through OpenCode itself:
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
ocskill eval-live --model provider/model --auth current --trials 3
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
`current` copies only the current auth file, not the user's global UES configuration.
|
|
94
|
+
|
|
95
|
+
## Telemetry
|
|
96
|
+
|
|
97
|
+
Results may include:
|
|
98
|
+
|
|
99
|
+
- pass/fail
|
|
100
|
+
- process exit status
|
|
101
|
+
- duration
|
|
102
|
+
- changed files
|
|
103
|
+
- bounded stdout/stderr
|
|
104
|
+
- best-effort tool calls
|
|
105
|
+
- loaded skills/subagent targets
|
|
106
|
+
- token/cost data when exposed
|
|
107
|
+
- long-suite orchestration inspection
|
|
108
|
+
|
|
109
|
+
No hidden chain-of-thought is collected.
|
|
110
|
+
|
|
111
|
+
## Report aggregation
|
|
112
|
+
|
|
113
|
+
```bash
|
|
114
|
+
ocskill eval-report .ues-evals
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
Compare the same model, variant, prompt, fixture, grader and environment. Report multiple trials.
|
|
118
|
+
|
|
119
|
+
A benchmark result is evidence only for the measured workload. UES does not claim to turn one base model into another.
|
|
120
|
+
|
|
121
|
+
|
|
122
|
+
## V7 live-run observability and evidence gate
|
|
123
|
+
|
|
124
|
+
Live runs accept:
|
|
125
|
+
|
|
126
|
+
```bash
|
|
127
|
+
--heartbeat-ms 30000
|
|
128
|
+
--idle-timeout-ms 300000
|
|
129
|
+
--timeout-ms 900000
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
The harness prints a start line and heartbeat for an active model run. Hard timeout and idle timeout are recorded separately. Ctrl+C aborts the active OpenCode process tree and sets exit code 130 after the current result is recorded.
|
|
133
|
+
|
|
134
|
+
V7 long-suite UES mode additionally requires **receipt-backed verification for every planned task**. A narrative completion entry is still preserved for backwards compatibility, but it does not satisfy the V7 long benchmark unless at least one successful structured verification receipt is bound to that task.
|
|
135
|
+
|
|
136
|
+
This intentionally raises the benchmark bar: final code correctness + durable orchestration + machine-observable verification are all required.
|
|
@@ -0,0 +1,102 @@
|
|
|
1
|
+
# Publishing to npm
|
|
2
|
+
|
|
3
|
+
The package is published as:
|
|
4
|
+
|
|
5
|
+
```text
|
|
6
|
+
opencode-agent-skill
|
|
7
|
+
```
|
|
8
|
+
|
|
9
|
+
## Release prerequisites
|
|
10
|
+
|
|
11
|
+
1. The npm account must have publish rights to the unscoped `opencode-agent-skill` package name.
|
|
12
|
+
2. `package.json` and `package-lock.json` versions must match.
|
|
13
|
+
3. `CHANGELOG.md` must contain the release.
|
|
14
|
+
4. Run the complete local validation:
|
|
15
|
+
|
|
16
|
+
```cmd
|
|
17
|
+
npm run ci
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
V7.7 CI includes syntax validation, resource validation, static skill routing, the 120-case V2 router matrix, standard and long hidden-grader integrity checks, unit/integration tests, package dry-run, and an isolated packed global-install smoke.
|
|
21
|
+
|
|
22
|
+
## Manual release-like test
|
|
23
|
+
|
|
24
|
+
Do not use `npm install -g .` as a release simulation because npm may create a symlink/junction back to the checkout.
|
|
25
|
+
|
|
26
|
+
Use:
|
|
27
|
+
|
|
28
|
+
```cmd
|
|
29
|
+
npm pack
|
|
30
|
+
npm install -g .\opencode-agent-skill-7.7.0.tgz --allow-scripts=opencode-agent-skill
|
|
31
|
+
ocskill status
|
|
32
|
+
ocskill doctor
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
For routine development, the automated `smoke:pack` test uses an isolated npm prefix/OpenCode config so it does not replace the developer's currently installed UES.
|
|
36
|
+
|
|
37
|
+
## Manual publish
|
|
38
|
+
|
|
39
|
+
```cmd
|
|
40
|
+
npm login
|
|
41
|
+
npm whoami
|
|
42
|
+
npm run ci
|
|
43
|
+
npm publish --access public
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
After publication verify:
|
|
47
|
+
|
|
48
|
+
```cmd
|
|
49
|
+
npm view opencode-agent-skill versions --json
|
|
50
|
+
npm view opencode-agent-skill@7.7.0 version
|
|
51
|
+
npm dist-tag ls opencode-agent-skill
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
The expected release tag is:
|
|
55
|
+
|
|
56
|
+
```text
|
|
57
|
+
latest: 7.7.0
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
## GitHub Actions publishing
|
|
61
|
+
|
|
62
|
+
The repository's publish workflow is release-ready for token-based publishing and provenance. It runs the same package validation before `npm publish`.
|
|
63
|
+
|
|
64
|
+
For stronger long-term supply-chain security, configure npm Trusted Publishing for:
|
|
65
|
+
|
|
66
|
+
```text
|
|
67
|
+
GitHub owner: laivannha0202
|
|
68
|
+
Repository: opencode-agent-skill-
|
|
69
|
+
Workflow: publish.yml
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Then the GitHub-hosted workflow can authenticate through OIDC instead of a long-lived npm publish token. npm Trusted Publishing requires the corresponding publisher relationship to be configured on npm; repository code alone cannot create that account-side trust relationship.
|
|
73
|
+
|
|
74
|
+
Until that npm-side setup is complete, keep a valid publish credential configured as `NPM_TOKEN`.
|
|
75
|
+
|
|
76
|
+
## Release checklist
|
|
77
|
+
|
|
78
|
+
1. Confirm version/changelog/package-lock consistency.
|
|
79
|
+
2. Run `npm run ci` locally on Windows and at least one Unix-like environment when practical.
|
|
80
|
+
3. Run `npm pack` and inspect the tarball contents.
|
|
81
|
+
4. Verify packed install/state/resource counts.
|
|
82
|
+
5. When installer compatibility changed, exercise both forced V1 and V2 paths through tests.
|
|
83
|
+
6. When updater behavior changed, test explicit latest-tag resolution, equal-version behavior, and downgrade refusal.
|
|
84
|
+
7. Commit and push the release branch.
|
|
85
|
+
8. Merge only after review/local validation is clean.
|
|
86
|
+
9. Create/push the matching `vX.Y.Z` tag or run the publish workflow.
|
|
87
|
+
10. Verify registry version and `latest` dist-tag.
|
|
88
|
+
11. Install the published package on a clean environment before announcing it.
|
|
89
|
+
|
|
90
|
+
## One-command user install
|
|
91
|
+
|
|
92
|
+
After publication:
|
|
93
|
+
|
|
94
|
+
```cmd
|
|
95
|
+
npm install -g opencode-agent-skill
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
If lifecycle execution is blocked by local npm policy:
|
|
99
|
+
|
|
100
|
+
```cmd
|
|
101
|
+
ocskill install
|
|
102
|
+
```
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# OpenCode compatibility
|
|
2
|
+
|
|
3
|
+
UES 6 ships one npm package for OpenCode 1.x and 2.x, while only enabling V2-native runtime features when V2 is detected.
|
|
4
|
+
|
|
5
|
+
## Detection
|
|
6
|
+
|
|
7
|
+
During resource sync, UES reads:
|
|
8
|
+
|
|
9
|
+
```text
|
|
10
|
+
opencode --version
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Tests may override with:
|
|
14
|
+
|
|
15
|
+
```text
|
|
16
|
+
UES_OPENCODE_MAJOR=1
|
|
17
|
+
UES_OPENCODE_MAJOR=2
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
The detected major is recorded in the managed state.
|
|
21
|
+
|
|
22
|
+
## OpenCode 1.x
|
|
23
|
+
|
|
24
|
+
UES installs:
|
|
25
|
+
|
|
26
|
+
- 39 namespaced skills
|
|
27
|
+
- 11 namespaced commands
|
|
28
|
+
- 10 namespaced subagents using compatible V1 `permission` frontmatter
|
|
29
|
+
- managed global `AGENTS.md` block
|
|
30
|
+
|
|
31
|
+
The V2 runtime plugin is not installed.
|
|
32
|
+
|
|
33
|
+
Durable CLI state, task DAG, plan/integration gates and model-policy configuration remain available. V2-specific automatic fresh-session dispatch is unavailable.
|
|
34
|
+
|
|
35
|
+
## OpenCode 2.x
|
|
36
|
+
|
|
37
|
+
UES converts managed agent permission frontmatter to V2 ordered `permissions` and installs:
|
|
38
|
+
|
|
39
|
+
```text
|
|
40
|
+
<global-config>/plugins/ues-router/
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
The plugin provides:
|
|
44
|
+
|
|
45
|
+
- prompt-admission skill routing
|
|
46
|
+
- long-task context guardrails
|
|
47
|
+
- permission safety evaluation
|
|
48
|
+
- read-only durable-state/task-graph/context-pack tools
|
|
49
|
+
- `ues.dispatch_task` fresh executor runtime
|
|
50
|
+
|
|
51
|
+
`ues.dispatch_task` uses V2 session APIs to create a fresh session, select `ues-executor`, optionally switch to the configured model tier, prompt one approved task and wait for completion.
|
|
52
|
+
|
|
53
|
+
## Router control
|
|
54
|
+
|
|
55
|
+
```cmd
|
|
56
|
+
ocskill router status
|
|
57
|
+
ocskill router on
|
|
58
|
+
ocskill router on --max 3
|
|
59
|
+
ocskill router off
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Default maximum is 4 selected skills; supported range is 1–6.
|
|
63
|
+
|
|
64
|
+
## Model routing
|
|
65
|
+
|
|
66
|
+
```cmd
|
|
67
|
+
ocskill models status
|
|
68
|
+
ocskill models on
|
|
69
|
+
ocskill models set standard provider/model
|
|
70
|
+
ocskill models set heavy provider/strong-model
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
Re-run `ocskill install` after changing static managed-agent model frontmatter. Runtime `ues.dispatch_task` also reads the current model policy for attempt-based executor escalation.
|
|
74
|
+
|
|
75
|
+
## Upgrading OpenCode
|
|
76
|
+
|
|
77
|
+
After moving between V1/V2 lines:
|
|
78
|
+
|
|
79
|
+
```cmd
|
|
80
|
+
ocskill install
|
|
81
|
+
ocskill status
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Only UES-managed resources are rewritten/removed. Unrelated user plugins/resources are preserved.
|
|
85
|
+
|
|
86
|
+
## Primary references
|
|
87
|
+
|
|
88
|
+
- https://opencode.ai/v2/docs/build/plugins
|
|
89
|
+
- https://opencode.ai/v2/docs/build/plugins/migrate-v1
|
|
90
|
+
- https://opencode.ai/v2/docs/permissions
|
|
91
|
+
- https://opencode.ai/v2/docs/plugins
|
|
92
|
+
- https://opencode.ai/v2/docs/skills
|
|
93
|
+
|
|
94
|
+
|
|
95
|
+
## V7 capability probing
|
|
96
|
+
|
|
97
|
+
Version detection remains useful for install-time compatibility, but V7 runtime dispatch does not assume that a major version proves the availability of every session API.
|
|
98
|
+
|
|
99
|
+
The managed V2 plugin probes for:
|
|
100
|
+
|
|
101
|
+
- session creation
|
|
102
|
+
- prompting
|
|
103
|
+
- waiting
|
|
104
|
+
- context retrieval
|
|
105
|
+
- agent switching
|
|
106
|
+
- model switching
|
|
107
|
+
- session hooks
|
|
108
|
+
|
|
109
|
+
`ues.capabilities` exposes the observed surface. `ues.dispatch_task` fails closed when the minimum fresh-dispatch capability set is unavailable instead of attempting a partially supported execution path.
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Research sources
|
|
2
|
+
|
|
3
|
+
UES is informed by public engineering-agent patterns and current primary platform documentation.
|
|
4
|
+
|
|
5
|
+
## Engineering-agent and skill-system references
|
|
6
|
+
|
|
7
|
+
- Alibaba OpenCodeReview: https://github.com/alibaba/open-code-review
|
|
8
|
+
- Open GSD Core: https://github.com/open-gsd/gsd-core
|
|
9
|
+
- Superpowers: https://github.com/obra/superpowers
|
|
10
|
+
- Agent Skills open specification: https://github.com/agentskills/agentskills
|
|
11
|
+
- Anthropic Skills: https://github.com/anthropics/skills
|
|
12
|
+
- OpenAI Agents SDK: https://github.com/openai/openai-agents-python
|
|
13
|
+
- NVIDIA Skills: https://github.com/NVIDIA/skills
|
|
14
|
+
- Ruflo / Claude Flow: https://github.com/ruvnet/ruflo
|
|
15
|
+
|
|
16
|
+
## OpenCode references
|
|
17
|
+
|
|
18
|
+
V1/current installed-system references:
|
|
19
|
+
- Skills: https://opencode.ai/docs/skills
|
|
20
|
+
- Agents: https://opencode.ai/docs/agents
|
|
21
|
+
- Commands: https://opencode.ai/docs/commands
|
|
22
|
+
|
|
23
|
+
OpenCode V2 compatibility/runtime references:
|
|
24
|
+
- Migration from V1: https://opencode.ai/v2/docs/migrate-v1
|
|
25
|
+
- Agent permissions: https://opencode.ai/v2/docs/permissions
|
|
26
|
+
- Agents: https://opencode.ai/v2/docs/agents
|
|
27
|
+
- Skills: https://opencode.ai/v2/docs/skills
|
|
28
|
+
- Plugin discovery/configuration: https://opencode.ai/v2/docs/plugins
|
|
29
|
+
- Plugin API and prompt hooks: https://opencode.ai/v2/docs/build/plugins
|
|
30
|
+
- V1 plugin migration: https://opencode.ai/v2/docs/build/plugins/migrate-v1
|
|
31
|
+
|
|
32
|
+
## npm release references
|
|
33
|
+
|
|
34
|
+
- Trusted Publishing: https://docs.npmjs.com/trusted-publishers/
|
|
35
|
+
- Provenance: https://docs.npmjs.com/generating-provenance-statements/
|
|
36
|
+
|
|
37
|
+
The project does not vendor or copy these projects. External documentation is used to verify platform behavior and inform UES's own implementation.
|