@ngockhoale/ukit 2.2.16 → 2.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +104 -0
- package/manifests/platform.full.yaml +11 -0
- package/package.json +1 -1
- package/src/core/executionContracts.js +130 -0
- package/src/core/runtimeConfig.js +13 -50
- package/src/index/taskRouting.js +10 -103
- package/templates/.claude/agents/ukit-vision-analyst.md +32 -21
- package/templates/.claude/hooks/context-window-guard.sh +66 -14
- package/templates/.claude/hooks/protect-files.sh +1 -0
- package/templates/.claude/hooks/sensitive-data-guard.sh +269 -0
- package/templates/.claude/hooks/skill-router.sh +33 -0
- package/templates/.claude/hooks/vision-router.sh +63 -37
- package/templates/.claude/settings.json +17 -1
- package/templates/.claude/ukit/index/extract-image.mjs +18 -8
- package/templates/.claude/ukit/index/route-task.mjs +53 -2
- package/templates/.claude/ukit/runtime/execution-ledger.mjs +180 -22
- package/templates/.claude/ukit/runtime/hook-chain-runner.mjs +8 -4
- package/templates/.omp/hooks/pre/ukit-bridge.js +15 -5
- package/templates/AGENTS.md +2 -2
- package/templates/CLAUDE.md +1 -1
- package/templates/ukit/storage/config.json +8 -7
- package/src/core/router/advisor.js +0 -42
- package/src/core/router/router.js +0 -180
- package/src/core/validation/confidence.js +0 -89
- package/src/core/validation/validator.js +0 -165
- package/templates/docs/INSTALL.md +0 -115
- package/templates/docs/STATUS.md +0 -81
- package/templates/docs/TASKS.md +0 -79
- package/templates/docs/UKIT_USAGE_GUIDE.md +0 -163
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,110 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.3.1 - 2026-09-09
|
|
6
|
+
|
|
7
|
+
"1 prompt mà đứng 10 lần" is fixed at the root. Four stall vectors in the completion
|
|
8
|
+
gate and router are gone, no stop is silent anymore, and routine backend model swaps
|
|
9
|
+
are now worked through instead of stalled on. All found by root-cause debugging with a
|
|
10
|
+
failing regression test per fix.
|
|
11
|
+
|
|
12
|
+
1) Stop-hook recursion: `execution-ledger.mjs --evaluate-stop` re-blocked the reentrant
|
|
13
|
+
Stop event (Claude Code re-invokes Stop with `stop_hook_active: true` after a block), so
|
|
14
|
+
one block became a self-sustaining block/continue chain ending in a cap-forced mid-task
|
|
15
|
+
stop. The gate now honors `stop_hook_active` and, when evidence is still missing, emits
|
|
16
|
+
a user-visible `systemMessage` instead of silently ending.
|
|
17
|
+
|
|
18
|
+
2) requestKey churn wiped evidence: `skill-router.sh` rebuilds requestKey per tool call
|
|
19
|
+
(commandText/targetFile are hashed in), so the first verification command re-keyed the
|
|
20
|
+
route mid-request and `--record` took the freshLedger path, erasing the write/verification
|
|
21
|
+
evidence the same request had produced — the gate then re-blocked until the cap forced a
|
|
22
|
+
silent stop. The ledger now keys evidence identity by the user prompt text
|
|
23
|
+
(`promptKey`); a re-key within one logical request carries evidence forward, a genuinely
|
|
24
|
+
new prompt starts clean.
|
|
25
|
+
|
|
26
|
+
3) Subagent stalls: subagent tool calls run the same hooks against the same session
|
|
27
|
+
ledger and route state (docs-confirmed; sidechain payloads carry `agent_id`/
|
|
28
|
+
`agent_type`). A subagent's Edit could re-route the request to a different
|
|
29
|
+
executionMode, which both broke the old promptKey identity (wiping main evidence) and
|
|
30
|
+
replaced the main route's completionEvidence. Fixes: promptKey is the prompt text alone
|
|
31
|
+
(null prompt never carries), and `skill-router.sh` early-exits on sidechain payloads so
|
|
32
|
+
subagent tool calls never touch the main route state. omp's bridge invokes the same
|
|
33
|
+
script, so every harness is covered.
|
|
34
|
+
|
|
35
|
+
4) No silent endings: the non-gated modes' `notify` and the post-cap `capped` paths now
|
|
36
|
+
emit a `systemMessage` alongside their stderr line — previously these ended turns with
|
|
37
|
+
the user seeing nothing.
|
|
38
|
+
|
|
39
|
+
Gateway model swaps (limit hit → a different vendor model behind the same alias) are
|
|
40
|
+
detected from the transcript's per-assistant `message.model` and surfaced in
|
|
41
|
+
`context-window-guard.sh` as ONE routine line per hour: keep working, do not restart or
|
|
42
|
+
re-plan; if a tool call errors after a swap, re-check the tool and call it again. The
|
|
43
|
+
swap itself cold-reads the prompt cache once — that latency is transport-level and
|
|
44
|
+
cannot be removed by hooks; everything after it now continues automatically.
|
|
45
|
+
|
|
46
|
+
Integrated on top of 2.3.0 (rebased; not published as 2.2.17): 2.3.0's stricter gates
|
|
47
|
+
made the re-key carry load-bearing for `targetedVerificationSucceeded`/`sourceFiles`
|
|
48
|
+
too — without carrying them, a mid-request re-key re-demanded targeted evidence and
|
|
49
|
+
re-introduced the stall this release fixes (regression-tested RED→GREEN). The 2.3.0
|
|
50
|
+
notify test that expected a silent stdout was updated to the no-silent-stops contract.
|
|
51
|
+
|
|
52
|
+
Verified: executionLedgerCli 29/29, skillRouterHook 33/33, contextWindowGuard 18/18;
|
|
53
|
+
full suite 1280/1280 (74 files) and `release:verify` green on the integrated tree;
|
|
54
|
+
template↔live mirrors byte-identical (live settings.json re-rendered through the install
|
|
55
|
+
pipeline, not copied raw).
|
|
56
|
+
|
|
57
|
+
## 2.3.0 - 2026-09-08
|
|
58
|
+
|
|
59
|
+
Model-agnostic hardening release. Every configured model name is treated as a gateway
|
|
60
|
+
mapping whose backend can change at any time, so correctness now comes from stable role
|
|
61
|
+
contracts, explicit evidence, and regression fixtures — never provider identity.
|
|
62
|
+
|
|
63
|
+
**Sensitive-data gate (fail-closed, default ON).** NEW `sensitive-data-guard.sh` scans
|
|
64
|
+
PreToolUse Read/Grep paths, Bash commands (secret-dump shapes, inline credentials,
|
|
65
|
+
high-confidence token patterns like `sk-…`), and UserPromptSubmit text; on suspicion it
|
|
66
|
+
blocks with exit 2 and shows only redacted previews (3-char prefix + length + sha256
|
|
67
|
+
prefix) — never the value. The only approval path is the USER editing the protected
|
|
68
|
+
allowlist `.ukit/storage/security/allowlist.json` (sha256 of the value, or an allowed file
|
|
69
|
+
path); `security.sensitiveDataGate: true` ships as the default and the gate can be disabled
|
|
70
|
+
in `.ukit/storage/config.json` for debugging. Wired into Claude settings hook chains, the
|
|
71
|
+
omp bridge, and the install manifest. Live-fire verified across all four channels.
|
|
72
|
+
|
|
73
|
+
**Contract tables unified to one source.** NEW `src/core/executionContracts.js` is the
|
|
74
|
+
canonical table set (execution modes, contracts, contract→tier mapping, risk profiles).
|
|
75
|
+
`taskRouting.js` imports it, runtime config defaults derive from it, and the shipped
|
|
76
|
+
`route-task.mjs` keeps a literal copy locked by a new sync test — fixing a real drift where
|
|
77
|
+
route-task's `local-build` contract was missing verification evidence and the post-edit
|
|
78
|
+
review policy. `config.json` no longer duplicates `modelTier` per contract.
|
|
79
|
+
|
|
80
|
+
**Tier-role wiring.** Route state now carries a `tierLane` hand-off: silent when the
|
|
81
|
+
contract tier is `code` (the default implementation lane), an explicit "hand this to an
|
|
82
|
+
agent bound to the lite/smart tier" instruction otherwise, and it follows `escalatedTier`
|
|
83
|
+
on repeat failures. This surfaced and fixed a real gap: `compactRouteSummary` is a
|
|
84
|
+
whitelist and was silently dropping `expectedSourceFiles`, which made target-aware
|
|
85
|
+
completion evidence inert on the real CLI path.
|
|
86
|
+
|
|
87
|
+
**Evidence tightening.** Completion-gate `impact-evidence` now requires reading one of the
|
|
88
|
+
routed target files (any random read no longer counts), and verification receipts record
|
|
89
|
+
`targeted` vs `broad` scope against the routed verification plan; recovery instructions
|
|
90
|
+
name the exact files expected.
|
|
91
|
+
|
|
92
|
+
**Vision flow smoothing.** Already-analysed images no longer re-trigger route hints
|
|
93
|
+
(`alreadyAnalyzed[]` from `extract-image.mjs --json`); unresolvable image paths get a short
|
|
94
|
+
note instead of a wrong remedy; the vision analyst gets a capability self-check (first
|
|
95
|
+
image Read is the probe), a task envelope carrying the ORIGINAL user prompt verbatim, and
|
|
96
|
+
an observations-vs-inference output block; hook-chain budgets now scale with chain size so
|
|
97
|
+
growing Edit/Write chains never orphan the runner.
|
|
98
|
+
|
|
99
|
+
**Mapping fingerprint fixtures.** New consistency tests lock agent↔tier bindings across
|
|
100
|
+
Claude Code and omp, the omp modelRoles → `unic-*` labels, and the tier structure
|
|
101
|
+
(vision stays a capability tier outside the escalation ladder) — provider identities are
|
|
102
|
+
deliberately left unpinned so gateway mappings can change freely without breaking
|
|
103
|
+
behavior contracts.
|
|
104
|
+
|
|
105
|
+
Verified: `yarn test` 1277/1277 (74 files), `test:release-core` 1270, `test:artifact` 7,
|
|
106
|
+
`release:verify` all checks passed, 13/13 handoff node suites, sensitive-data-guard
|
|
107
|
+
live-fire 5/5 as designed.
|
|
108
|
+
|
|
5
109
|
## 2.2.16 - 2026-09-06
|
|
6
110
|
|
|
7
111
|
The vision gate is gone — image work is advisory-routed, never blocked. Since TASK-008, a
|
|
@@ -1094,6 +1094,17 @@ items:
|
|
|
1094
1094
|
packs:
|
|
1095
1095
|
- core
|
|
1096
1096
|
|
|
1097
|
+
- id: hook-sensitive-data-guard
|
|
1098
|
+
type: hook
|
|
1099
|
+
sourceTemplate: .claude/hooks/sensitive-data-guard.sh
|
|
1100
|
+
targetPath: .claude/hooks/sensitive-data-guard.sh
|
|
1101
|
+
requires: []
|
|
1102
|
+
mergeStrategy: overwrite_with_backup
|
|
1103
|
+
variables: []
|
|
1104
|
+
enabledByDefault: true
|
|
1105
|
+
packs:
|
|
1106
|
+
- core
|
|
1107
|
+
|
|
1097
1108
|
- id: hook-handoff-model-guard
|
|
1098
1109
|
type: hook
|
|
1099
1110
|
sourceTemplate: .claude/hooks/handoff-model-guard.sh
|
package/package.json
CHANGED
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
// Single source of truth for UKit execution contracts (WS-D contract table unification).
|
|
2
|
+
//
|
|
3
|
+
// Every routing decision that is keyed on an execution contract derives from the tables in
|
|
4
|
+
// this file:
|
|
5
|
+
// - src/index/taskRouting.js (in-process routing: contracts + risk profiles + mode order)
|
|
6
|
+
// - src/core/runtimeConfig.js (config defaults: contracts minus runtime-only keys)
|
|
7
|
+
// - templates/ukit/storage/config.json (literal mirror of the config defaults)
|
|
8
|
+
// - templates/.claude/ukit/index/route-task.mjs (shipped standalone CLI: contracts + tiers)
|
|
9
|
+
//
|
|
10
|
+
// route-task.mjs runs inside target projects and cannot import from src/, so it keeps a
|
|
11
|
+
// literal copy — tests/consistency/executionContractSync.test.js locks that copy to these
|
|
12
|
+
// tables so the four can never drift again (the local-build completionEvidence drift that
|
|
13
|
+
// motivated this module is the failure mode it exists to prevent).
|
|
14
|
+
|
|
15
|
+
export const EXECUTION_MODE_ORDER = [
|
|
16
|
+
'tiny-fix',
|
|
17
|
+
'local-fix',
|
|
18
|
+
'local-build',
|
|
19
|
+
'find-cause',
|
|
20
|
+
'shared-edit',
|
|
21
|
+
'map-impact',
|
|
22
|
+
'review-release',
|
|
23
|
+
];
|
|
24
|
+
|
|
25
|
+
export const EXECUTION_CONTRACTS = {
|
|
26
|
+
'tiny-fix': {
|
|
27
|
+
maxReadPasses: 0,
|
|
28
|
+
maxContextPulls: 0,
|
|
29
|
+
verificationPolicy: 'minimal-or-targeted',
|
|
30
|
+
completionRule: 'never-claim-done-without-write',
|
|
31
|
+
delegationPolicy: 'disallow',
|
|
32
|
+
completionEvidence: ['write-evidence'],
|
|
33
|
+
},
|
|
34
|
+
'local-fix': {
|
|
35
|
+
maxReadPasses: 1,
|
|
36
|
+
maxContextPulls: 1,
|
|
37
|
+
verificationPolicy: 'targeted-if-covered',
|
|
38
|
+
completionRule: 'require-write',
|
|
39
|
+
delegationPolicy: 'disallow',
|
|
40
|
+
completionEvidence: ['write-evidence'],
|
|
41
|
+
},
|
|
42
|
+
'local-build': {
|
|
43
|
+
maxReadPasses: 2,
|
|
44
|
+
maxContextPulls: 1,
|
|
45
|
+
verificationPolicy: 'targeted-if-covered',
|
|
46
|
+
completionRule: 'require-write-and-verification',
|
|
47
|
+
delegationPolicy: 'disallow-by-default',
|
|
48
|
+
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
49
|
+
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
50
|
+
},
|
|
51
|
+
'find-cause': {
|
|
52
|
+
maxReadPassesBeforeReassess: 3,
|
|
53
|
+
verificationPolicy: 'root-cause-then-targeted',
|
|
54
|
+
completionRule: 'never-claim-fixed-without-write-and-verification',
|
|
55
|
+
delegationPolicy: 'allow-specialized-debug-lane',
|
|
56
|
+
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
57
|
+
},
|
|
58
|
+
'shared-edit': {
|
|
59
|
+
maxReadPasses: 2,
|
|
60
|
+
maxContextPulls: 2,
|
|
61
|
+
verificationPolicy: 'targeted-then-widen-on-risk',
|
|
62
|
+
completionRule: 'require-write-and-verification',
|
|
63
|
+
delegationPolicy: 'allow-qualified-sidecar',
|
|
64
|
+
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
65
|
+
mirrorConsistencyRequired: true,
|
|
66
|
+
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
67
|
+
},
|
|
68
|
+
'map-impact': {
|
|
69
|
+
maxReadPasses: 3,
|
|
70
|
+
maxContextPulls: 3,
|
|
71
|
+
verificationPolicy: 'impact-first-then-targeted-then-widen-on-risk',
|
|
72
|
+
completionRule: 'require-impact-evidence-before-edit-claim',
|
|
73
|
+
delegationPolicy: 'allow-impact-sidecar',
|
|
74
|
+
completionEvidence: ['impact-evidence', 'write-evidence', 'verification-evidence'],
|
|
75
|
+
mirrorConsistencyRequired: true,
|
|
76
|
+
},
|
|
77
|
+
'review-release': {
|
|
78
|
+
verificationPolicy: 'evidence-first',
|
|
79
|
+
completionRule: 'report-findings-not-implementation',
|
|
80
|
+
delegationPolicy: 'allow-review-sidecar',
|
|
81
|
+
completionEvidence: ['verification-evidence'],
|
|
82
|
+
},
|
|
83
|
+
};
|
|
84
|
+
|
|
85
|
+
export const MODEL_TIER_BY_CONTRACT = {
|
|
86
|
+
'tiny-fix': 'lite',
|
|
87
|
+
'local-fix': 'code',
|
|
88
|
+
'local-build': 'code',
|
|
89
|
+
'find-cause': 'code',
|
|
90
|
+
'shared-edit': 'code',
|
|
91
|
+
'map-impact': 'code',
|
|
92
|
+
'review-release': 'smart',
|
|
93
|
+
};
|
|
94
|
+
|
|
95
|
+
export const CONTRACT_RISK_PROFILES = {
|
|
96
|
+
'tiny-fix': { riskLevel: 'minimal', contextPolicy: 'confirm-target-only' },
|
|
97
|
+
'local-fix': { riskLevel: 'local', contextPolicy: 'bounded-local' },
|
|
98
|
+
'local-build': { riskLevel: 'local', contextPolicy: 'bounded-with-related-files' },
|
|
99
|
+
'find-cause': { riskLevel: 'investigation', contextPolicy: 'trace-then-bounded-context' },
|
|
100
|
+
'shared-edit': { riskLevel: 'shared', contextPolicy: 'bounded-with-related-tests' },
|
|
101
|
+
'map-impact': { riskLevel: 'wide', contextPolicy: 'impact-map-first' },
|
|
102
|
+
'review-release': { riskLevel: 'release', contextPolicy: 'evidence-review-only' },
|
|
103
|
+
};
|
|
104
|
+
|
|
105
|
+
const CONFIG_EXCLUDED_KEYS = new Set(['modelTier', 'completionEvidence', 'mirrorConsistencyRequired']);
|
|
106
|
+
|
|
107
|
+
/** The user-facing config subset of a contract: runtime-only bookkeeping keys stripped. */
|
|
108
|
+
export function stripContractForConfig(contract = {}) {
|
|
109
|
+
return Object.fromEntries(
|
|
110
|
+
Object.entries(contract).filter(([key]) => !CONFIG_EXCLUDED_KEYS.has(key))
|
|
111
|
+
);
|
|
112
|
+
}
|
|
113
|
+
|
|
114
|
+
export function buildExecutionContract(executionMode = null) {
|
|
115
|
+
if (!executionMode) {
|
|
116
|
+
return null;
|
|
117
|
+
}
|
|
118
|
+
const contract = EXECUTION_CONTRACTS[executionMode];
|
|
119
|
+
if (!contract) {
|
|
120
|
+
return null;
|
|
121
|
+
}
|
|
122
|
+
return { ...contract, modelTier: MODEL_TIER_BY_CONTRACT[executionMode] };
|
|
123
|
+
}
|
|
124
|
+
|
|
125
|
+
/** Config-default shape: contracts without modelTier/completionEvidence/mirror keys. */
|
|
126
|
+
export function buildConfigContracts() {
|
|
127
|
+
return Object.fromEntries(
|
|
128
|
+
Object.entries(EXECUTION_CONTRACTS).map(([mode, contract]) => [mode, stripContractForConfig(contract)])
|
|
129
|
+
);
|
|
130
|
+
}
|
|
@@ -3,6 +3,7 @@ import path from 'node:path';
|
|
|
3
3
|
import { createRequire } from 'node:module';
|
|
4
4
|
import { buildRuntimePaths } from './runtimePaths.js';
|
|
5
5
|
import { loadShippedCompactBudget } from './compact/contextBudget.js';
|
|
6
|
+
import { buildConfigContracts } from './executionContracts.js';
|
|
6
7
|
|
|
7
8
|
const require = createRequire(import.meta.url);
|
|
8
9
|
const { version: PACKAGE_VERSION } = require('../../package.json');
|
|
@@ -69,6 +70,10 @@ export function buildDefaultRuntimeConfig(overrides = {}) {
|
|
|
69
70
|
affectVerification: true,
|
|
70
71
|
affectDelegation: true,
|
|
71
72
|
},
|
|
73
|
+
security: {
|
|
74
|
+
sensitiveDataGate: true,
|
|
75
|
+
allowlistPath: '.ukit/storage/security/allowlist.json',
|
|
76
|
+
},
|
|
72
77
|
compact: {
|
|
73
78
|
enabled: true,
|
|
74
79
|
tokenThreshold: loadShippedCompactBudget().tokenThreshold,
|
|
@@ -132,56 +137,7 @@ export function buildDefaultRuntimeConfig(overrides = {}) {
|
|
|
132
137
|
enabled: true,
|
|
133
138
|
orchestratorModel: 'claude-sonnet-5',
|
|
134
139
|
advisorEnabled: true,
|
|
135
|
-
contracts:
|
|
136
|
-
'tiny-fix': {
|
|
137
|
-
maxReadPasses: 0,
|
|
138
|
-
maxContextPulls: 0,
|
|
139
|
-
verificationPolicy: 'minimal-or-targeted',
|
|
140
|
-
completionRule: 'never-claim-done-without-write',
|
|
141
|
-
delegationPolicy: 'disallow',
|
|
142
|
-
},
|
|
143
|
-
'local-fix': {
|
|
144
|
-
maxReadPasses: 1,
|
|
145
|
-
maxContextPulls: 1,
|
|
146
|
-
verificationPolicy: 'targeted-if-covered',
|
|
147
|
-
completionRule: 'require-write',
|
|
148
|
-
delegationPolicy: 'disallow',
|
|
149
|
-
},
|
|
150
|
-
'local-build': {
|
|
151
|
-
maxReadPasses: 2,
|
|
152
|
-
maxContextPulls: 1,
|
|
153
|
-
verificationPolicy: 'targeted-if-covered',
|
|
154
|
-
completionRule: 'require-write-and-verification',
|
|
155
|
-
delegationPolicy: 'disallow-by-default',
|
|
156
|
-
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
157
|
-
},
|
|
158
|
-
'find-cause': {
|
|
159
|
-
maxReadPassesBeforeReassess: 3,
|
|
160
|
-
verificationPolicy: 'root-cause-then-targeted',
|
|
161
|
-
completionRule: 'never-claim-fixed-without-write-and-verification',
|
|
162
|
-
delegationPolicy: 'allow-specialized-debug-lane',
|
|
163
|
-
},
|
|
164
|
-
'shared-edit': {
|
|
165
|
-
maxReadPasses: 2,
|
|
166
|
-
maxContextPulls: 2,
|
|
167
|
-
verificationPolicy: 'targeted-then-widen-on-risk',
|
|
168
|
-
completionRule: 'require-write-and-verification',
|
|
169
|
-
delegationPolicy: 'allow-qualified-sidecar',
|
|
170
|
-
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
171
|
-
},
|
|
172
|
-
'map-impact': {
|
|
173
|
-
maxReadPasses: 3,
|
|
174
|
-
maxContextPulls: 3,
|
|
175
|
-
verificationPolicy: 'impact-first-then-targeted-then-widen-on-risk',
|
|
176
|
-
completionRule: 'require-impact-evidence-before-edit-claim',
|
|
177
|
-
delegationPolicy: 'allow-impact-sidecar',
|
|
178
|
-
},
|
|
179
|
-
'review-release': {
|
|
180
|
-
verificationPolicy: 'evidence-first',
|
|
181
|
-
completionRule: 'report-findings-not-implementation',
|
|
182
|
-
delegationPolicy: 'allow-review-sidecar',
|
|
183
|
-
},
|
|
184
|
-
},
|
|
140
|
+
contracts: buildConfigContracts(),
|
|
185
141
|
modelTiers: {
|
|
186
142
|
lite: { claudeModel: 'claude-haiku-4-5', genericModel: 'unic-lite' },
|
|
187
143
|
code: { claudeModel: 'claude-sonnet-5', genericModel: 'unic-code' },
|
|
@@ -312,6 +268,13 @@ export function validateRuntimeConfig(config) {
|
|
|
312
268
|
pushBooleanError(errors, config.autonomy.affectDelegation, 'autonomy.affectDelegation');
|
|
313
269
|
}
|
|
314
270
|
|
|
271
|
+
if (!isPlainObject(config.security)) {
|
|
272
|
+
errors.push('security must be an object.');
|
|
273
|
+
} else {
|
|
274
|
+
pushBooleanError(errors, config.security.sensitiveDataGate, 'security.sensitiveDataGate');
|
|
275
|
+
pushNonEmptyStringError(errors, config.security.allowlistPath, 'security.allowlistPath');
|
|
276
|
+
}
|
|
277
|
+
|
|
315
278
|
if (!VALID_AGENTS.has(config.agent)) {
|
|
316
279
|
errors.push(`agent must be one of: ${[...VALID_AGENTS].join(', ')}.`);
|
|
317
280
|
}
|
package/src/index/taskRouting.js
CHANGED
|
@@ -6,6 +6,11 @@ import { buildRouteSignalText } from './languageTools.js';
|
|
|
6
6
|
import { resolveContext } from './resolveContext.js';
|
|
7
7
|
import { deriveVerificationPlan } from './verificationPlan.js';
|
|
8
8
|
import { isSharedImpactFile } from './impactCatalog.js';
|
|
9
|
+
import {
|
|
10
|
+
CONTRACT_RISK_PROFILES,
|
|
11
|
+
EXECUTION_CONTRACTS,
|
|
12
|
+
EXECUTION_MODE_ORDER,
|
|
13
|
+
} from '../core/executionContracts.js';
|
|
9
14
|
|
|
10
15
|
const MAX_ACTIVE_ROUTE_SKILLS = 2;
|
|
11
16
|
|
|
@@ -545,18 +550,8 @@ function applySafeUpwardBias(candidates = [], fallbackMode = 'local-build') {
|
|
|
545
550
|
}
|
|
546
551
|
|
|
547
552
|
function executionModeRank(mode = '') {
|
|
548
|
-
const
|
|
549
|
-
|
|
550
|
-
'local-fix',
|
|
551
|
-
'local-build',
|
|
552
|
-
'find-cause',
|
|
553
|
-
'shared-edit',
|
|
554
|
-
'map-impact',
|
|
555
|
-
'review-release',
|
|
556
|
-
];
|
|
557
|
-
|
|
558
|
-
const index = orderedModes.indexOf(mode);
|
|
559
|
-
return index >= 0 ? index : orderedModes.length;
|
|
553
|
+
const index = EXECUTION_MODE_ORDER.indexOf(mode);
|
|
554
|
+
return index >= 0 ? index : EXECUTION_MODE_ORDER.length;
|
|
560
555
|
}
|
|
561
556
|
|
|
562
557
|
function buildExecutionContract(executionMode = null) {
|
|
@@ -564,67 +559,8 @@ function buildExecutionContract(executionMode = null) {
|
|
|
564
559
|
return null;
|
|
565
560
|
}
|
|
566
561
|
|
|
567
|
-
const
|
|
568
|
-
|
|
569
|
-
maxReadPasses: 0,
|
|
570
|
-
maxContextPulls: 0,
|
|
571
|
-
verificationPolicy: 'minimal-or-targeted',
|
|
572
|
-
completionRule: 'never-claim-done-without-write',
|
|
573
|
-
delegationPolicy: 'disallow',
|
|
574
|
-
completionEvidence: ['write-evidence'],
|
|
575
|
-
},
|
|
576
|
-
'local-fix': {
|
|
577
|
-
maxReadPasses: 1,
|
|
578
|
-
maxContextPulls: 1,
|
|
579
|
-
verificationPolicy: 'targeted-if-covered',
|
|
580
|
-
completionRule: 'require-write',
|
|
581
|
-
delegationPolicy: 'disallow',
|
|
582
|
-
completionEvidence: ['write-evidence'],
|
|
583
|
-
},
|
|
584
|
-
'local-build': {
|
|
585
|
-
maxReadPasses: 2,
|
|
586
|
-
maxContextPulls: 1,
|
|
587
|
-
verificationPolicy: 'targeted-if-covered',
|
|
588
|
-
completionRule: 'require-write-and-verification',
|
|
589
|
-
delegationPolicy: 'disallow-by-default',
|
|
590
|
-
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
591
|
-
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
592
|
-
},
|
|
593
|
-
'find-cause': {
|
|
594
|
-
maxReadPassesBeforeReassess: 3,
|
|
595
|
-
verificationPolicy: 'root-cause-then-targeted',
|
|
596
|
-
completionRule: 'never-claim-fixed-without-write-and-verification',
|
|
597
|
-
delegationPolicy: 'allow-specialized-debug-lane',
|
|
598
|
-
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
599
|
-
},
|
|
600
|
-
'shared-edit': {
|
|
601
|
-
maxReadPasses: 2,
|
|
602
|
-
maxContextPulls: 2,
|
|
603
|
-
verificationPolicy: 'targeted-then-widen-on-risk',
|
|
604
|
-
completionRule: 'require-write-and-verification',
|
|
605
|
-
delegationPolicy: 'allow-qualified-sidecar',
|
|
606
|
-
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
607
|
-
mirrorConsistencyRequired: true,
|
|
608
|
-
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
609
|
-
},
|
|
610
|
-
'map-impact': {
|
|
611
|
-
maxReadPasses: 3,
|
|
612
|
-
maxContextPulls: 3,
|
|
613
|
-
verificationPolicy: 'impact-first-then-targeted-then-widen-on-risk',
|
|
614
|
-
completionRule: 'require-impact-evidence-before-edit-claim',
|
|
615
|
-
delegationPolicy: 'allow-impact-sidecar',
|
|
616
|
-
completionEvidence: ['impact-evidence', 'write-evidence', 'verification-evidence'],
|
|
617
|
-
mirrorConsistencyRequired: true,
|
|
618
|
-
},
|
|
619
|
-
'review-release': {
|
|
620
|
-
verificationPolicy: 'evidence-first',
|
|
621
|
-
completionRule: 'report-findings-not-implementation',
|
|
622
|
-
delegationPolicy: 'allow-review-sidecar',
|
|
623
|
-
completionEvidence: ['verification-evidence'],
|
|
624
|
-
},
|
|
625
|
-
};
|
|
626
|
-
|
|
627
|
-
return contracts[executionMode] ? { ...contracts[executionMode] } : null;
|
|
562
|
+
const contract = EXECUTION_CONTRACTS[executionMode];
|
|
563
|
+
return contract ? { ...contract } : null;
|
|
628
564
|
}
|
|
629
565
|
|
|
630
566
|
function buildApproachSelectorResult({
|
|
@@ -637,36 +573,7 @@ function buildApproachSelectorResult({
|
|
|
637
573
|
}
|
|
638
574
|
|
|
639
575
|
const executionContract = buildExecutionContract(executionMode);
|
|
640
|
-
const policyByMode =
|
|
641
|
-
'tiny-fix': {
|
|
642
|
-
riskLevel: 'minimal',
|
|
643
|
-
contextPolicy: 'confirm-target-only',
|
|
644
|
-
},
|
|
645
|
-
'local-fix': {
|
|
646
|
-
riskLevel: 'local',
|
|
647
|
-
contextPolicy: 'bounded-local',
|
|
648
|
-
},
|
|
649
|
-
'local-build': {
|
|
650
|
-
riskLevel: 'local',
|
|
651
|
-
contextPolicy: 'bounded-with-related-files',
|
|
652
|
-
},
|
|
653
|
-
'find-cause': {
|
|
654
|
-
riskLevel: 'investigation',
|
|
655
|
-
contextPolicy: 'trace-then-bounded-context',
|
|
656
|
-
},
|
|
657
|
-
'shared-edit': {
|
|
658
|
-
riskLevel: 'shared',
|
|
659
|
-
contextPolicy: 'bounded-with-related-tests',
|
|
660
|
-
},
|
|
661
|
-
'map-impact': {
|
|
662
|
-
riskLevel: 'wide',
|
|
663
|
-
contextPolicy: 'impact-map-first',
|
|
664
|
-
},
|
|
665
|
-
'review-release': {
|
|
666
|
-
riskLevel: 'release',
|
|
667
|
-
contextPolicy: 'evidence-review-only',
|
|
668
|
-
},
|
|
669
|
-
};
|
|
576
|
+
const policyByMode = CONTRACT_RISK_PROFILES;
|
|
670
577
|
const selectedPolicy = policyByMode[executionMode] ?? {
|
|
671
578
|
riskLevel: 'local',
|
|
672
579
|
contextPolicy: 'bounded-local',
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: ukit-vision-analyst
|
|
3
|
-
description: "The
|
|
4
|
-
model: unic-vision #
|
|
3
|
+
description: "The vision specialist lane for this repo. Use whenever a prompt references, attaches, or points at an image (screenshot, mockup, diagram, photo of an error) and the caller needs to know what is actually in it. Non-vision mappings (unic-code/unic-smart) must never guess at image contents — route image analysis here whenever the active model has not verified native vision. Reports findings only; never writes product code."
|
|
4
|
+
model: unic-vision # capability-lane gateway mapping (image-capable backend). Not a cost tier; never "fix" this to a lite/code/smart label.
|
|
5
5
|
color: magenta
|
|
6
6
|
tools: ["Read", "Glob", "Bash"]
|
|
7
7
|
---
|
|
@@ -18,19 +18,20 @@ actually in them. You never write, edit, or refactor product code — you analys
|
|
|
18
18
|
- Stay end-user-invisible: this lane is internal UKit orchestration, not something end users invoke
|
|
19
19
|
by name.
|
|
20
20
|
|
|
21
|
-
## 2.
|
|
21
|
+
## 2. Capability self-check — the first image Read is the probe
|
|
22
22
|
|
|
23
|
-
|
|
23
|
+
Model names are gateway mappings whose backend can change at any time, so capability is
|
|
24
|
+
demonstrated, never assumed from a name:
|
|
24
25
|
|
|
25
|
-
-
|
|
26
|
-
|
|
27
|
-
`
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
26
|
+
- Your first `Read` of an image file doubles as the capability probe. If the tool result does not
|
|
27
|
+
actually give you the image (error, empty result, or you can only see text you were told about),
|
|
28
|
+
you are not vision-capable right now: emit `STATUS: WRONG_MODEL` and stop immediately. Do not
|
|
29
|
+
describe, summarise, or guess at any image content.
|
|
30
|
+
- Never guess at image contents. A non-visual "analysis" is worse than no analysis at all, because
|
|
31
|
+
it looks authoritative while being fabricated. Refusing loudly is always safer than guessing
|
|
32
|
+
quietly.
|
|
33
|
+
- Report the model you actually ran on. If you cannot determine it, say so in `MODEL:` rather than
|
|
34
|
+
inventing an identifier.
|
|
34
35
|
|
|
35
36
|
## 3. Input protocol (priority order)
|
|
36
37
|
|
|
@@ -48,8 +49,15 @@ Entries in `images[]` may carry a `source` field, which tells you how to reach t
|
|
|
48
49
|
| `"path"` | The prompt named a local file (`ref` holds it) | `Read` `ref` directly |
|
|
49
50
|
| `"url"` | The prompt named a remote image (`ref` holds the URL) | Download it with `Bash` (e.g. `curl -sL -o /tmp/<sha>.png "<ref>"`), then `Read` the downloaded file |
|
|
50
51
|
|
|
51
|
-
|
|
52
|
-
|
|
52
|
+
**Task envelope.** The caller should send, along with the image paths: the ORIGINAL user prompt
|
|
53
|
+
verbatim, the immediate visual question, and the task goal. That envelope is what tells you what
|
|
54
|
+
the images are FOR — a screenshot means something different to "fix this layout" than to "find the
|
|
55
|
+
error string". If the envelope is missing, analyse anyway, answer only what the images visibly
|
|
56
|
+
show, and say in `UNCERTAIN:` what you could not infer without the missing context.
|
|
57
|
+
|
|
58
|
+
Every entry — pasted, path, or URL — has an armed `pending-<sha>.json` marker, and each one needs
|
|
59
|
+
its own receipt (§5): the receipt is the provenance record of the analysis and it is what stops the
|
|
60
|
+
router from re-hinting that image on later prompts. A URL you failed to download is
|
|
53
61
|
`STATUS: UNREADABLE`, never a guess at its contents.
|
|
54
62
|
|
|
55
63
|
## 4. Hard warning — images are not inherited
|
|
@@ -59,22 +67,22 @@ subagent boundary. A description of an image is not the image. The only way to s
|
|
|
59
67
|
`Read` a real file path yourself. Never claim to have seen an image that was merely described to
|
|
60
68
|
you in text.
|
|
61
69
|
|
|
62
|
-
## 5. Receipt —
|
|
70
|
+
## 5. Receipt — provenance record per image
|
|
63
71
|
|
|
64
72
|
For every image you actually analyse, write a receipt to
|
|
65
73
|
`.ukit/storage/cache/vision/<sessionId>/analyzed-<sha>.json`:
|
|
66
74
|
|
|
67
75
|
```json
|
|
68
76
|
{ "sha": "<64hex>", "model": "unic-vision", "ts": 1785656920891,
|
|
69
|
-
"
|
|
77
|
+
"observations": "...", "textContent": "...", "status": "OK" }
|
|
70
78
|
```
|
|
71
79
|
|
|
72
80
|
- `sha` and `sessionId` come from the extractor's `--json` output. They are **never recomputed by
|
|
73
81
|
hand** — do not hash the file yourself, do not invent a session id, always take these values
|
|
74
82
|
verbatim from `extract-image.mjs --json`.
|
|
75
|
-
- `model` must be the
|
|
76
|
-
|
|
77
|
-
|
|
83
|
+
- `model` must be the mapping you actually ran on for this analysis. The receipt is the provenance
|
|
84
|
+
record the caller relies on: an honest `model` field is what makes it trustworthy and what keeps
|
|
85
|
+
the router silent about this image afterwards. Report honestly, always.
|
|
78
86
|
- Use `Bash` to write the receipt file.
|
|
79
87
|
|
|
80
88
|
## 6. Output block
|
|
@@ -85,8 +93,9 @@ Always end with:
|
|
|
85
93
|
STATUS: OK | NO_IMAGE | UNREADABLE | WRONG_MODEL
|
|
86
94
|
MODEL: <actual model>
|
|
87
95
|
IMAGES: <n> (<paths>)
|
|
88
|
-
|
|
96
|
+
OBSERVATIONS: [what is directly visible — layout, elements, colors, structure; facts only]
|
|
89
97
|
TEXT_CONTENT: [verbatim text/code/errors legible in the image, or "none"]
|
|
98
|
+
INFERENCES: [what the observations likely mean — interpretation, clearly separate from facts]
|
|
90
99
|
RELEVANT_TO_TASK: [how it answers the caller's question]
|
|
91
100
|
UNCERTAIN: [anything ambiguous or illegible, or "none"]
|
|
92
101
|
```
|
|
@@ -94,6 +103,8 @@ UNCERTAIN: [anything ambiguous or illegible, or "none"]
|
|
|
94
103
|
## 7. Guardrails
|
|
95
104
|
|
|
96
105
|
- Transcribe error text and code verbatim rather than paraphrasing it.
|
|
106
|
+
- Keep `OBSERVATIONS:` strictly factual; every interpretation goes in `INFERENCES:`. The caller
|
|
107
|
+
must be able to tell what you saw from what you concluded.
|
|
97
108
|
- State uncertainty explicitly instead of guessing — use `UNCERTAIN:` for anything ambiguous or
|
|
98
109
|
illegible.
|
|
99
110
|
- Do not read unrelated repo files; stay scoped to the image(s) and the immediate task question.
|