@ngockhoale/ukit 2.3.8 → 2.3.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +28 -0
- package/package.json +1 -1
- package/src/core/compact/threshold.js +13 -0
- package/templates/.claude/agents/bug-debugger.md +2 -2
- package/templates/.claude/agents/feature-implementer.md +5 -3
- package/templates/.claude/hooks/context-hardcap-gate.sh +15 -3
- package/templates/.claude/hooks/context-window-guard.sh +1 -1
- package/templates/.claude/ukit/runtime/compact-threshold.mjs +10 -1
- package/templates/.claude/ukit/runtime/execution-ledger.mjs +22 -0
- package/templates/.omp/agents/bug-debugger.md +2 -2
- package/templates/.omp/agents/feature-implementer.md +5 -3
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,34 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.3.9 - 2026-09-10
|
|
6
|
+
|
|
7
|
+
Freeze-sweep wave 5: fixes for completion-gate loops, worker-agent dead ends, and malformed
|
|
8
|
+
context-cap configuration that could otherwise halt all mutation progress.
|
|
9
|
+
|
|
10
|
+
**P1 — clean `find-cause` audits were forced into a Stop-hook loop.** A recommend-only
|
|
11
|
+
investigation that found no actionable bug still required write and verification evidence, so
|
|
12
|
+
the completion gate blocked the honest outcome and demanded a fabricated edit. The execution
|
|
13
|
+
ledger now permits that narrow clean-audit outcome while retaining normal recovery blocks after
|
|
14
|
+
a failed verification or incomplete mutation.
|
|
15
|
+
|
|
16
|
+
**P1 — routed workers could wait for a user they cannot contact.** The Claude Code and omp
|
|
17
|
+
`bug-debugger` / `feature-implementer` agents instructed workers to “ask the user” on unclear
|
|
18
|
+
or non-reproducible work. They now perform one bounded next diagnostic step and return a
|
|
19
|
+
structured `STATUS: BLOCKED` report with the exact missing artifact or decision to the parent
|
|
20
|
+
agent, which owns user communication.
|
|
21
|
+
|
|
22
|
+
**P1 — malformed context-cap configuration could brick mutations.** Invalid
|
|
23
|
+
`compact.hardCapTokens` could collapse the cap to one token, invalid
|
|
24
|
+
`hardCapGraceCalls` could eliminate the landing window, and wrong-shaped present
|
|
25
|
+
`hardCapBlock` values could unexpectedly leave the hard gate enabled. The shared source/runtime
|
|
26
|
+
threshold logic and Claude hook now accept only positive integer budgets, while malformed
|
|
27
|
+
non-boolean `hardCapBlock` values fail open; the documented absent-key default remains enabled.
|
|
28
|
+
The context-window advisory uses the same hard-cap validation. Regression coverage includes
|
|
29
|
+
negative/zero/fraction/string/boolean/null budgets and malformed toggle values.
|
|
30
|
+
|
|
31
|
+
Full suite: 81 files, 1,396 tests green.
|
|
32
|
+
|
|
5
33
|
## 2.3.8 - 2026-09-10
|
|
6
34
|
|
|
7
35
|
Freeze-sweep wave 4: every fix in this release targets the "agent silently stops working"
|
package/package.json
CHANGED
|
@@ -35,6 +35,15 @@ function finiteNumber(value, fallback = 0) {
|
|
|
35
35
|
return Number.isFinite(number) ? number : fallback;
|
|
36
36
|
}
|
|
37
37
|
|
|
38
|
+
// A malformed hard cap must fail back to a safe ceiling, never collapse to one token and
|
|
39
|
+
// brick every Edit/Write/Bash call. Unlike token estimates, configuration is an operator
|
|
40
|
+
// contract: accept only whole positive token counts.
|
|
41
|
+
function positiveInteger(value, fallback) {
|
|
42
|
+
return typeof value === 'number' && Number.isFinite(value) && Number.isInteger(value) && value > 0
|
|
43
|
+
? value
|
|
44
|
+
: fallback;
|
|
45
|
+
}
|
|
46
|
+
|
|
38
47
|
function uniqueStrings(values) {
|
|
39
48
|
const seen = new Set();
|
|
40
49
|
const unique = [];
|
|
@@ -434,11 +443,15 @@ export function buildCompactThresholds(config = {}) {
|
|
|
434
443
|
);
|
|
435
444
|
const hardThreshold = Math.max(softThreshold + 1, Math.round(softThreshold * 1.6));
|
|
436
445
|
const baselineTokens = Math.max(120, Math.min(18_000, Math.round(softThreshold * 0.18)));
|
|
446
|
+
// Keep the source mirror's config contract aligned with the installed runtime: malformed
|
|
447
|
+
// caps fall back to a safe ceiling rather than turning every mutation into an over-cap call.
|
|
448
|
+
const hardCapTokens = positiveInteger(config?.compact?.hardCapTokens, loadShippedCompactBudget().hardCapTokens);
|
|
437
449
|
|
|
438
450
|
return {
|
|
439
451
|
softThreshold,
|
|
440
452
|
hardThreshold,
|
|
441
453
|
baselineTokens,
|
|
454
|
+
hardCapTokens,
|
|
442
455
|
};
|
|
443
456
|
}
|
|
444
457
|
|
|
@@ -18,7 +18,7 @@ Systematic debugging — understand before fixing.
|
|
|
18
18
|
|
|
19
19
|
- Run the failing command/action.
|
|
20
20
|
- Capture exact error message and stack trace.
|
|
21
|
-
- If not reproducible → document conditions
|
|
21
|
+
- If not reproducible → document conditions, run the next most-discriminating bounded repro, then report any exact missing user-only artifact/permission to the parent agent. Never wait for or ask a user directly — you are a worker.
|
|
22
22
|
|
|
23
23
|
### 2. Trace Root Cause
|
|
24
24
|
|
|
@@ -83,4 +83,4 @@ Daily mode: skip. Handoff mode: set task status `pending_review` in `INDEX.md`;
|
|
|
83
83
|
- non-trivial bug: `docs/MEMORY.md` + `docs/PROJECT.md` + `docs/CODE_MAP.md`
|
|
84
84
|
- read `docs/WORKLOG.md` only recent relevant entries
|
|
85
85
|
- Keep fix scope minimal — no drive-by refactors.
|
|
86
|
-
- If root cause is unclear after 5 minutes of tracing → ask user
|
|
86
|
+
- If root cause is unclear after 5 minutes of tracing → try one bounded alternative hypothesis/repro, then report precise evidence plus the smallest needed missing context to the parent agent. Never wait for or ask a user directly — the parent owns user communication.
|
|
@@ -12,7 +12,9 @@ Implement requested behavior with minimal scope drift.
|
|
|
12
12
|
- **Daily/ad-hoc mode** (DEFAULT): task didn't come from `docs/AI_HANDOFF/` → use the original lightweight workflow. Tests only when touched code already has coverage. No reviewer trigger.
|
|
13
13
|
- **Handoff mode**: task file is `docs/AI_HANDOFF/tasks/TASK-xxx.md` OR user explicitly invokes handoff (e.g. "execute task TASK-001") → activate full Quality Gate: test-first → green → reviewer.
|
|
14
14
|
|
|
15
|
-
If unsure which mode applies,
|
|
15
|
+
If unsure which mode applies, default to Daily/ad-hoc mode unless the task path or prompt
|
|
16
|
+
explicitly selects Handoff. Do not ask a user — you are a worker; report the ambiguity and
|
|
17
|
+
reasoning to the parent agent so it can decide whether to re-route.
|
|
16
18
|
|
|
17
19
|
**In Handoff mode you are running unattended — ask nothing.** You were spawned by an
|
|
18
20
|
orchestrator driving a pipeline; there is no human in your conversation to answer, and a
|
|
@@ -33,7 +35,7 @@ and even then, report it, don't ask about it.
|
|
|
33
35
|
- non-trivial: `docs/MEMORY.md` + `docs/PROJECT.md` + `docs/CODE_MAP.md`
|
|
34
36
|
- Identify target files and existing patterns.
|
|
35
37
|
- If task came from handoff, read `tasks/TASK-xxx.md` and locate its **Test Plan** + **Verification Commands**.
|
|
36
|
-
- Daily mode: if confidence is low or risk is high,
|
|
38
|
+
- Daily mode: if confidence is low or risk is high, inspect the next bounded source/context signal and hand back a concise `STATUS: BLOCKED` report with the exact missing decision or artifact if confidence remains low. Do not ask a user directly. Handoff mode: do not ask — decide and record the decision (see above).
|
|
37
39
|
|
|
38
40
|
### 2. Plan Approach (< 1 minute)
|
|
39
41
|
|
|
@@ -44,7 +46,7 @@ and even then, report it, don't ask about it.
|
|
|
44
46
|
|
|
45
47
|
- Write the test(s) from §2 / from task Test Plan.
|
|
46
48
|
- Run them: must FAIL for the expected reason. Capture output.
|
|
47
|
-
- If test passes immediately → test is wrong or behavior already exists. Fix the test
|
|
49
|
+
- If test passes immediately → test is wrong or behavior already exists. Fix the test; if the intended behavior cannot be inferred, hand back `STATUS: BLOCKED` with the observed behavior and the smallest decision the parent must resolve. Never silently stop.
|
|
48
50
|
- **Daily mode**: skip this step unless touched code already has tests (then follow original rule).
|
|
49
51
|
|
|
50
52
|
|
|
@@ -96,7 +96,12 @@ function readRunCursor() {
|
|
|
96
96
|
? payload.session_id.trim()
|
|
97
97
|
: undefined,
|
|
98
98
|
};
|
|
99
|
-
|
|
99
|
+
// This is a liveness backstop, not a security control. A present value with the wrong
|
|
100
|
+
// JSON shape must never leave the gate unexpectedly enabled and brick every mutation once
|
|
101
|
+
// the cap is reached. Only an explicitly configured boolean `true` enables the gate;
|
|
102
|
+
// missing preserves the documented default-on behavior for existing installations.
|
|
103
|
+
const rawHardCapBlock = config?.compact?.hardCapBlock;
|
|
104
|
+
if (rawHardCapBlock === false || (rawHardCapBlock !== undefined && typeof rawHardCapBlock !== 'boolean')) {
|
|
100
105
|
process.exit(0);
|
|
101
106
|
return;
|
|
102
107
|
}
|
|
@@ -175,8 +180,15 @@ function readRunCursor() {
|
|
|
175
180
|
}
|
|
176
181
|
const resumable = run || ordinaryTask;
|
|
177
182
|
if (resumable) {
|
|
178
|
-
|
|
179
|
-
|
|
183
|
+
// Invalid config must not turn the landing allowance negative/zero (which makes every
|
|
184
|
+
// unfinished run hard-block immediately). This hook is a liveness backstop, so malformed
|
|
185
|
+
// values deliberately fall back to the documented default rather than fail closed.
|
|
186
|
+
const configuredGraceCalls = config?.compact?.hardCapGraceCalls;
|
|
187
|
+
const graceCalls = typeof configuredGraceCalls === 'number'
|
|
188
|
+
&& Number.isFinite(configuredGraceCalls)
|
|
189
|
+
&& Number.isInteger(configuredGraceCalls)
|
|
190
|
+
&& configuredGraceCalls > 0
|
|
191
|
+
? configuredGraceCalls
|
|
180
192
|
: 10;
|
|
181
193
|
|
|
182
194
|
let slots = readSlots();
|
|
@@ -73,7 +73,7 @@ function loadHardCap() {
|
|
|
73
73
|
try {
|
|
74
74
|
const raw = fs.readFileSync(path.join(projectRoot, '.ukit', 'storage', 'config.json'), 'utf8');
|
|
75
75
|
const value = JSON.parse(raw)?.compact?.hardCapTokens;
|
|
76
|
-
if (Number.isFinite(value) && value > 0) return value;
|
|
76
|
+
if (typeof value === 'number' && Number.isFinite(value) && Number.isInteger(value) && value > 0) return value;
|
|
77
77
|
} catch { /* fall through to the default */ }
|
|
78
78
|
return 500_000;
|
|
79
79
|
}
|
|
@@ -43,6 +43,15 @@ function finiteNumber(value, fallback = 0) {
|
|
|
43
43
|
return Number.isFinite(number) ? number : fallback;
|
|
44
44
|
}
|
|
45
45
|
|
|
46
|
+
// A malformed hard cap must fail back to a safe ceiling, never collapse to one token and
|
|
47
|
+
// brick every Edit/Write/Bash call. Unlike token estimates, configuration is an operator
|
|
48
|
+
// contract: accept only whole positive token counts.
|
|
49
|
+
function positiveInteger(value, fallback) {
|
|
50
|
+
return typeof value === 'number' && Number.isFinite(value) && Number.isInteger(value) && value > 0
|
|
51
|
+
? value
|
|
52
|
+
: fallback;
|
|
53
|
+
}
|
|
54
|
+
|
|
46
55
|
function uniqueStrings(values) {
|
|
47
56
|
const seen = new Set();
|
|
48
57
|
const unique = [];
|
|
@@ -452,7 +461,7 @@ export function buildCompactThresholds(config = {}) {
|
|
|
452
461
|
// env.CLAUDE_CODE_AUTO_COMPACT_WINDOW and omp's compaction.thresholdTokens are rendered
|
|
453
462
|
// from this number at install time (70% of it), so the client auto-compacts before the
|
|
454
463
|
// gate blocks tools without anyone hand-syncing a second value.
|
|
455
|
-
const hardCapTokens =
|
|
464
|
+
const hardCapTokens = positiveInteger(config?.compact?.hardCapTokens, 500_000);
|
|
456
465
|
|
|
457
466
|
return {
|
|
458
467
|
softThreshold,
|
|
@@ -513,6 +513,28 @@ export function evaluateCompletion({ state = {}, ledger = {} } = {}) {
|
|
|
513
513
|
return { continue: false, notify: false, missingEvidence: [] };
|
|
514
514
|
}
|
|
515
515
|
|
|
516
|
+
// `find-cause` can validly end clean: an investigation may establish that no actionable
|
|
517
|
+
// defect exists. Its contract says a FIX cannot be claimed without write + verification;
|
|
518
|
+
// it does not make a mutation mandatory. Treating missing evidence as an unconditional
|
|
519
|
+
// command to edit made a clean audit self-block forever (Stop -> "make an Edit" -> no
|
|
520
|
+
// honest edit exists -> Stop again). Once a mutation was attempted, retain the normal
|
|
521
|
+
// recovery gate — only a no-mutation, recommend-only investigation that has not already
|
|
522
|
+
// observed a failed verification gets this release valve. A route that named an actionable
|
|
523
|
+
// command, or a failing check, has concrete unfinished work and must keep recovering.
|
|
524
|
+
if (
|
|
525
|
+
mode === 'find-cause'
|
|
526
|
+
&& routeSummary.policyMode === 'recommend-only'
|
|
527
|
+
&& !effectiveLedger.writeAttempted
|
|
528
|
+
&& !effectiveLedger.verificationFailed
|
|
529
|
+
) {
|
|
530
|
+
return {
|
|
531
|
+
continue: false,
|
|
532
|
+
notify: true,
|
|
533
|
+
missingEvidence,
|
|
534
|
+
reason: 'UKit investigation ended without a mutation. A clean audit is valid; report whether no actionable defect was found or a concrete blocker remains. Do not claim a bug was fixed without write and verification evidence.',
|
|
535
|
+
};
|
|
536
|
+
}
|
|
537
|
+
|
|
516
538
|
const gated = IMPLEMENT_MODES.has(mode);
|
|
517
539
|
if (!gated) {
|
|
518
540
|
return {
|
|
@@ -17,7 +17,7 @@ Systematic debugging — understand before fixing.
|
|
|
17
17
|
|
|
18
18
|
- Run the failing command/action.
|
|
19
19
|
- Capture exact error message and stack trace.
|
|
20
|
-
- If not reproducible → document conditions
|
|
20
|
+
- If not reproducible → document conditions, run the next most-discriminating bounded repro, then report any exact missing user-only artifact/permission to the parent agent. Never wait for or ask a user directly — you are a worker.
|
|
21
21
|
|
|
22
22
|
### 2. Trace Root Cause
|
|
23
23
|
|
|
@@ -82,4 +82,4 @@ Daily mode: skip. Handoff mode: set task status `pending_review` in `INDEX.md`;
|
|
|
82
82
|
- non-trivial bug: `docs/MEMORY.md` + `docs/PROJECT.md` + `docs/CODE_MAP.md`
|
|
83
83
|
- read `docs/WORKLOG.md` only recent relevant entries
|
|
84
84
|
- Keep fix scope minimal — no drive-by refactors.
|
|
85
|
-
- If root cause is unclear after 5 minutes of tracing → ask user
|
|
85
|
+
- If root cause is unclear after 5 minutes of tracing → try one bounded alternative hypothesis/repro, then report precise evidence plus the smallest needed missing context to the parent agent. Never wait for or ask a user directly — the parent owns user communication.
|
|
@@ -11,7 +11,9 @@ Implement requested behavior with minimal scope drift.
|
|
|
11
11
|
- **Daily/ad-hoc mode** (DEFAULT): task didn't come from `docs/AI_HANDOFF/` → use the original lightweight workflow. Tests only when touched code already has coverage. No reviewer trigger.
|
|
12
12
|
- **Handoff mode**: task file is `docs/AI_HANDOFF/tasks/TASK-xxx.md` OR user explicitly invokes handoff (e.g. "execute task TASK-001") → activate full Quality Gate: test-first → green → reviewer.
|
|
13
13
|
|
|
14
|
-
If unsure which mode applies,
|
|
14
|
+
If unsure which mode applies, default to Daily/ad-hoc mode unless the task path or prompt
|
|
15
|
+
explicitly selects Handoff. Do not ask a user — you are a worker; report the ambiguity and
|
|
16
|
+
reasoning to the parent agent so it can decide whether to re-route.
|
|
15
17
|
|
|
16
18
|
**In Handoff mode you are running unattended — ask nothing.** You were spawned by an
|
|
17
19
|
orchestrator driving a pipeline; there is no human in your conversation to answer, and a
|
|
@@ -32,7 +34,7 @@ and even then, report it, don't ask about it.
|
|
|
32
34
|
- non-trivial: `docs/MEMORY.md` + `docs/PROJECT.md` + `docs/CODE_MAP.md`
|
|
33
35
|
- Identify target files and existing patterns.
|
|
34
36
|
- If task came from handoff, read `tasks/TASK-xxx.md` and locate its **Test Plan** + **Verification Commands**.
|
|
35
|
-
- Daily mode: if confidence is low or risk is high,
|
|
37
|
+
- Daily mode: if confidence is low or risk is high, inspect the next bounded source/context signal and hand back a concise `STATUS: BLOCKED` report with the exact missing decision or artifact if confidence remains low. Do not ask a user directly. Handoff mode: do not ask — decide and record the decision (see above).
|
|
36
38
|
|
|
37
39
|
### 2. Plan Approach (< 1 minute)
|
|
38
40
|
|
|
@@ -43,7 +45,7 @@ and even then, report it, don't ask about it.
|
|
|
43
45
|
|
|
44
46
|
- Write the test(s) from §2 / from task Test Plan.
|
|
45
47
|
- Run them: must FAIL for the expected reason. Capture output.
|
|
46
|
-
- If test passes immediately → test is wrong or behavior already exists. Fix the test
|
|
48
|
+
- If test passes immediately → test is wrong or behavior already exists. Fix the test; if the intended behavior cannot be inferred, hand back `STATUS: BLOCKED` with the observed behavior and the smallest decision the parent must resolve. Never silently stop.
|
|
47
49
|
- **Daily mode**: skip this step unless touched code already has tests (then follow original rule).
|
|
48
50
|
|
|
49
51
|
|