@ngockhoale/ukit 2.3.24 → 2.4.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,63 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.4.1 - 2026-09-14
6
+
7
+ Fix: the skill-router hook could hang the session and leak an orphaned process. `skill-router.sh`
8
+ runs on the `UserPromptSubmit`/`PreToolUse` hot path. Unlike its sibling hooks
9
+ (`task-watchdog.sh`, `handoff-resume.sh`), its node heredoc had no self-deadline watchdog: when
10
+ indexed-context resolution hung (a wedged import, a stalled mount), the host killed only the bash
11
+ wrapper with the node grandchild reparented to launchd and spinning forever — one orphan per run,
12
+ hundreds accumulated over days, adding sustained RSS pressure that presented as UKit freezing with
13
+ no message. The hook now arms the same wall-clock watchdog its siblings carry
14
+ (`setTimeout(() => process.exit(0), HOOK_DEADLINE_MS).unref()`, default 3000 ms, overridable via
15
+ `UKIT_HOOK_DEADLINE_MS`), so a hang can never outlive the hook. Early exit only drops the routing
16
+ hint — the same posture as a cache miss, and advisory-only as the hook already is. Applied
17
+ byte-identically to the live and shipped (template) copies.
18
+
19
+ ## 2.4.0 - 2026-09-14
20
+
21
+ Solution-discipline wording + offline instruments. This release is **additive**: it changes
22
+ instruction text that is already loaded in daily sessions and adds offline test instruments
23
+ under `tests/`, which the published tarball does not ship. It changes no runtime, config,
24
+ route, hook, or gateway behaviour, and it removes no existing behaviour.
25
+
26
+ - **Solution ladder + carve-outs in the delivery lane.** `templates/.claude/skills/delivery/SKILL.md`
27
+ (and its active mirror) now carry the canonical ladder: choose the smallest solution that fully
28
+ meets the stated requirements and project conventions; prefer reuse of an equivalent codebase
29
+ helper, then the standard library / native API / an already-installed dependency, and treat new
30
+ custom code or a new dependency as the last resort. The carve-out is explicit: never drop
31
+ validation, security, accessibility, error handling, compatibility, or required tests to reduce
32
+ code, and never ship a weaker version and ask. The delivery "Tier 2 — Structure Scan" step now
33
+ prefers the project index/resolver and falls back to raw Glob/Grep only when the index is stale,
34
+ missing, or returns no match — understanding source before editing is never skipped.
35
+ - **Feature-implementer ladder (Claude Code + omp).** `templates/.claude/agents/feature-implementer.md`,
36
+ its active byte-twin, and `templates/.omp/agents/feature-implementer.md` gain the same ladder plus
37
+ a guard reminder that fewer lines are not a reason to drop a guard. The daily-mode test rule is
38
+ left verbatim: the ladder must not become an excuse to cut tests.
39
+ - **Reviewer rubric (Claude Code + omp).** The `code-reviewer` copies gain 5 evidence-first
40
+ solution-fit questions (duplicate semantics / standard library suffices / guard-dropping
41
+ simplification / shared root cause / speculative abstraction) and a rule that says
42
+ **never grade brevity or line count as a quality win**. The sidecar diff lane stays
43
+ non-blocking with no auto-delete; verdict and model-isolation contracts are unchanged.
44
+ - **Offline instruments (not shipped).** A pure metric protocol
45
+ (`tests/benchmarks/solution-discipline/protocol.mjs`), four acceptance fixtures under
46
+ `tests/fixtures/solution-discipline/` (F01 reuse-helper, F06 trust-boundary, F09 keep-failure-path,
47
+ F11 required-interface) with good/bad reference solutions, and a host driver core exercised through
48
+ fake hosts (nonzero exit, timeout, malformed events, duplicate usage, missing terminal event).
49
+ These live under `tests/`, which `package.json.files` does not publish, so they stay out of the
50
+ tarball by construction; the release check confirms the packed file list contains no `tests/` path.
51
+ - **Deferred, with reasons.** The paid pilot (roadmap PT-06) and confirmatory benchmark (PT-09) are
52
+ deferred: no spending cap is approved and no billable model call was made this cycle. The runtime
53
+ lifecycle work (PT-07/PT-08 — shared renderer plus route/lifecycle/omp integration) is deferred
54
+ because the evidence for instruction-only wording is sufficient to start and no runtime drift has
55
+ been observed yet. The remaining fixtures and live-host probing are queued in
56
+ `docs/AI_HANDOFF/PROPOSAL-C18-benchmark-fixtures.md` for the funded pilot cycle.
57
+ - **Evidence record.** The preflight inventory this release builds on is pinned in
58
+ `docs/plans/ponytail-evidence.md` (revision pin, existing/proposed inventory, command map, blockers).
59
+ - **Truthfulness.** Savings are unverified until a paid benchmark pilot runs; no token/LOC reduction
60
+ is claimed for this release.
61
+
5
62
  ## 2.3.24 - 2026-09-13
6
63
 
7
64
  Post-2.3.23 routing and self-install audit. Two real defects fixed, plus the repo-local
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.24",
3
+ "version": "2.4.1",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -35,6 +35,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
35
35
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
36
36
  5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
37
37
  6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
38
+ 7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
39
+ - Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
40
+ - Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
41
+ - Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
42
+ - Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
43
+ - Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
44
+ **Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
38
45
 
39
46
  ### Severity ladder
40
47
 
@@ -171,13 +178,14 @@ the main task already has write + verification evidence; the caller is not waiti
171
178
 
172
179
  ### Review order
173
180
 
174
- Apply the same lenses as Code Review's steps 2-6, scoped to what the diff actually touches:
181
+ Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
175
182
 
176
183
  1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
177
184
  2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
178
185
  3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
179
186
  4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
180
187
  5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
188
+ 6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
181
189
 
182
190
  Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
183
191
  You may read files for context but this mode never edits anything.
@@ -41,6 +41,7 @@ and even then, report it, don't ask about it.
41
41
 
42
42
  - List files to create/modify (max diff).
43
43
  - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
44
+ - **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
44
45
 
45
46
  ### 3. Test First (RED) — Handoff mode
46
47
 
@@ -56,6 +57,7 @@ and even then, report it, don't ask about it.
56
57
  - Reuse existing code before creating new.
57
58
  - No unrelated changes or speculative refactors.
58
59
  - Follow project conventions (check `.claude/skills/` for patterns).
60
+ - **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
59
61
 
60
62
  ### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
61
63
 
@@ -102,6 +102,17 @@ const path = require('path');
102
102
  const crypto = require('crypto');
103
103
  const { pathToFileURL } = require('url');
104
104
 
105
+ // Wall-clock watchdog. This hook runs on the UserPromptSubmit/PreToolUse hot path but has
106
+ // no host-side process-group kill: when indexed-context resolution hangs (a wedged import,
107
+ // a stalled mount), the host kills only THIS bash wrapper, the node grandchild below is
108
+ // reparented to launchd, and it spins forever. Verified: exactly one orphan leaked per run
109
+ // of the router timeout test, 200+ accumulated over days. Self-exit at the deadline so a
110
+ // hang can never outlive the hook. Advisory only — an early exit just drops the routing
111
+ // hint, the same posture as a cache miss. Same watchdog task-watchdog.sh/handoff-resume.sh
112
+ // already carry; keep this one aligned with them.
113
+ const HOOK_DEADLINE_MS = Number.parseInt(process.env.UKIT_HOOK_DEADLINE_MS || '', 10) || 3000;
114
+ setTimeout(() => process.exit(0), HOOK_DEADLINE_MS).unref();
115
+
105
116
  (async () => {
106
117
  const STOPWORDS = new Set([
107
118
  'the', 'a', 'an', 'and', 'or', 'to', 'for', 'of', 'with', 'in', 'on', 'is', 'are',
@@ -13,6 +13,16 @@ Project: {{project.name}} | Stack: {{project.stack}}
13
13
  - Review gate for non-trivial edits (>3 files or >50 lines).
14
14
  - Reuse existing code. No drive-by refactors.
15
15
 
16
+ ## Solution Selection Ladder
17
+
18
+ Choose the smallest solution that fully meets the stated requirements and project conventions. When more than one option qualifies, prefer, in order:
19
+
20
+ 1. reuse an existing codebase helper with the same semantics;
21
+ 2. standard library, native platform API, or an already-installed dependency;
22
+ 3. custom code or a new dependency (last resort).
23
+
24
+ Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
25
+
16
26
  ## Docs-Centric Principle
17
27
 
18
28
  **Docs are the center.** Read docs first, write docs last.
@@ -35,9 +45,9 @@ Read in order. Never skip to Tier 3 without Tier 1+2.
35
45
  - Recent `docs/WORKLOG.md` → when continuing prior work or debugging
36
46
 
37
47
  **Tier 2 — Structure Scan** (before opening source files):
38
- - Glob file tree → know what exists where
39
- - Grep function/class/export signatures → understand shape without full reads
40
- - Narrows exactly which files to open in Tier 3
48
+ - Prefer the project index/resolver to find what exists where and the shape of symbols.
49
+ - Fall back to raw Glob/Grep only when the index is stale, missing, or returns no match.
50
+ - Narrows exactly which files to open in Tier 3; understanding source before editing is never skipped.
41
51
 
42
52
  **Tier 3 — Targeted Source Reads** (only what is needed):
43
53
  - Read only the specific files relevant to the task
@@ -57,7 +67,7 @@ After any non-trivial task, AI must update docs proactively — **do not wait to
57
67
  ## Workflow
58
68
 
59
69
  1. **Orient** — Tier 1: read docs, form hypothesis about the system
60
- 2. **Scan** — Tier 2: Glob + Grep to verify hypothesis, find relevant files
70
+ 2. **Scan** — Tier 2: use the project index/resolver to verify the hypothesis and find relevant files; raw Glob/Grep only when the index is stale, missing, or returns no match
61
71
  3. **Verify** — if source contradicts docs, update docs first
62
72
  4. **Execute** — smallest change, follow existing patterns
63
73
  5. **Test** — run tests, check regressions
@@ -34,6 +34,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
34
34
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
35
35
  5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
36
36
  6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
37
+ 7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
38
+ - Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
39
+ - Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
40
+ - Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
41
+ - Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
42
+ - Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
43
+ **Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
37
44
 
38
45
  ### Severity ladder
39
46
 
@@ -170,13 +177,14 @@ the main task already has write + verification evidence; the caller is not waiti
170
177
 
171
178
  ### Review order
172
179
 
173
- Apply the same lenses as Code Review's steps 2-6, scoped to what the diff actually touches:
180
+ Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
174
181
 
175
182
  1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
176
183
  2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
177
184
  3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
178
185
  4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
179
186
  5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
187
+ 6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
180
188
 
181
189
  Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
182
190
  You may read files for context but this mode never edits anything.
@@ -40,6 +40,7 @@ and even then, report it, don't ask about it.
40
40
 
41
41
  - List files to create/modify (max diff).
42
42
  - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
43
+ - **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
43
44
 
44
45
  ### 3. Test First (RED) — Handoff mode
45
46
 
@@ -55,6 +56,7 @@ and even then, report it, don't ask about it.
55
56
  - Reuse existing code before creating new.
56
57
  - No unrelated changes or speculative refactors.
57
58
  - Follow project conventions (check `.claude/skills/` for patterns).
59
+ - **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
58
60
 
59
61
  ### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
60
62