@ngockhoale/ukit 2.3.24 → 2.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +57 -0
- package/package.json +1 -1
- package/templates/.claude/agents/code-reviewer.md +9 -1
- package/templates/.claude/agents/feature-implementer.md +2 -0
- package/templates/.claude/hooks/skill-router.sh +11 -0
- package/templates/.claude/skills/delivery/SKILL.md +14 -4
- package/templates/.omp/agents/code-reviewer.md +9 -1
- package/templates/.omp/agents/feature-implementer.md +2 -0
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,63 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.4.1 - 2026-09-14
|
|
6
|
+
|
|
7
|
+
Fix: the skill-router hook could hang the session and leak an orphaned process. `skill-router.sh`
|
|
8
|
+
runs on the `UserPromptSubmit`/`PreToolUse` hot path. Unlike its sibling hooks
|
|
9
|
+
(`task-watchdog.sh`, `handoff-resume.sh`), its node heredoc had no self-deadline watchdog: when
|
|
10
|
+
indexed-context resolution hung (a wedged import, a stalled mount), the host killed only the bash
|
|
11
|
+
wrapper with the node grandchild reparented to launchd and spinning forever — one orphan per run,
|
|
12
|
+
hundreds accumulated over days, adding sustained RSS pressure that presented as UKit freezing with
|
|
13
|
+
no message. The hook now arms the same wall-clock watchdog its siblings carry
|
|
14
|
+
(`setTimeout(() => process.exit(0), HOOK_DEADLINE_MS).unref()`, default 3000 ms, overridable via
|
|
15
|
+
`UKIT_HOOK_DEADLINE_MS`), so a hang can never outlive the hook. Early exit only drops the routing
|
|
16
|
+
hint — the same posture as a cache miss, and advisory-only as the hook already is. Applied
|
|
17
|
+
byte-identically to the live and shipped (template) copies.
|
|
18
|
+
|
|
19
|
+
## 2.4.0 - 2026-09-14
|
|
20
|
+
|
|
21
|
+
Solution-discipline wording + offline instruments. This release is **additive**: it changes
|
|
22
|
+
instruction text that is already loaded in daily sessions and adds offline test instruments
|
|
23
|
+
under `tests/`, which the published tarball does not ship. It changes no runtime, config,
|
|
24
|
+
route, hook, or gateway behaviour, and it removes no existing behaviour.
|
|
25
|
+
|
|
26
|
+
- **Solution ladder + carve-outs in the delivery lane.** `templates/.claude/skills/delivery/SKILL.md`
|
|
27
|
+
(and its active mirror) now carry the canonical ladder: choose the smallest solution that fully
|
|
28
|
+
meets the stated requirements and project conventions; prefer reuse of an equivalent codebase
|
|
29
|
+
helper, then the standard library / native API / an already-installed dependency, and treat new
|
|
30
|
+
custom code or a new dependency as the last resort. The carve-out is explicit: never drop
|
|
31
|
+
validation, security, accessibility, error handling, compatibility, or required tests to reduce
|
|
32
|
+
code, and never ship a weaker version and ask. The delivery "Tier 2 — Structure Scan" step now
|
|
33
|
+
prefers the project index/resolver and falls back to raw Glob/Grep only when the index is stale,
|
|
34
|
+
missing, or returns no match — understanding source before editing is never skipped.
|
|
35
|
+
- **Feature-implementer ladder (Claude Code + omp).** `templates/.claude/agents/feature-implementer.md`,
|
|
36
|
+
its active byte-twin, and `templates/.omp/agents/feature-implementer.md` gain the same ladder plus
|
|
37
|
+
a guard reminder that fewer lines are not a reason to drop a guard. The daily-mode test rule is
|
|
38
|
+
left verbatim: the ladder must not become an excuse to cut tests.
|
|
39
|
+
- **Reviewer rubric (Claude Code + omp).** The `code-reviewer` copies gain 5 evidence-first
|
|
40
|
+
solution-fit questions (duplicate semantics / standard library suffices / guard-dropping
|
|
41
|
+
simplification / shared root cause / speculative abstraction) and a rule that says
|
|
42
|
+
**never grade brevity or line count as a quality win**. The sidecar diff lane stays
|
|
43
|
+
non-blocking with no auto-delete; verdict and model-isolation contracts are unchanged.
|
|
44
|
+
- **Offline instruments (not shipped).** A pure metric protocol
|
|
45
|
+
(`tests/benchmarks/solution-discipline/protocol.mjs`), four acceptance fixtures under
|
|
46
|
+
`tests/fixtures/solution-discipline/` (F01 reuse-helper, F06 trust-boundary, F09 keep-failure-path,
|
|
47
|
+
F11 required-interface) with good/bad reference solutions, and a host driver core exercised through
|
|
48
|
+
fake hosts (nonzero exit, timeout, malformed events, duplicate usage, missing terminal event).
|
|
49
|
+
These live under `tests/`, which `package.json.files` does not publish, so they stay out of the
|
|
50
|
+
tarball by construction; the release check confirms the packed file list contains no `tests/` path.
|
|
51
|
+
- **Deferred, with reasons.** The paid pilot (roadmap PT-06) and confirmatory benchmark (PT-09) are
|
|
52
|
+
deferred: no spending cap is approved and no billable model call was made this cycle. The runtime
|
|
53
|
+
lifecycle work (PT-07/PT-08 — shared renderer plus route/lifecycle/omp integration) is deferred
|
|
54
|
+
because the evidence for instruction-only wording is sufficient to start and no runtime drift has
|
|
55
|
+
been observed yet. The remaining fixtures and live-host probing are queued in
|
|
56
|
+
`docs/AI_HANDOFF/PROPOSAL-C18-benchmark-fixtures.md` for the funded pilot cycle.
|
|
57
|
+
- **Evidence record.** The preflight inventory this release builds on is pinned in
|
|
58
|
+
`docs/plans/ponytail-evidence.md` (revision pin, existing/proposed inventory, command map, blockers).
|
|
59
|
+
- **Truthfulness.** Savings are unverified until a paid benchmark pilot runs; no token/LOC reduction
|
|
60
|
+
is claimed for this release.
|
|
61
|
+
|
|
5
62
|
## 2.3.24 - 2026-09-13
|
|
6
63
|
|
|
7
64
|
Post-2.3.23 routing and self-install audit. Two real defects fixed, plus the repo-local
|
package/package.json
CHANGED
|
@@ -35,6 +35,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
35
35
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
36
36
|
5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
|
|
37
37
|
6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
38
|
+
7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
|
|
39
|
+
- Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
|
|
40
|
+
- Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
|
|
41
|
+
- Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
|
|
42
|
+
- Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
|
|
43
|
+
- Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
|
|
44
|
+
**Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
|
|
38
45
|
|
|
39
46
|
### Severity ladder
|
|
40
47
|
|
|
@@ -171,13 +178,14 @@ the main task already has write + verification evidence; the caller is not waiti
|
|
|
171
178
|
|
|
172
179
|
### Review order
|
|
173
180
|
|
|
174
|
-
Apply the same lenses as Code Review's steps 2-
|
|
181
|
+
Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
|
|
175
182
|
|
|
176
183
|
1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
|
|
177
184
|
2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
|
|
178
185
|
3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
179
186
|
4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
|
|
180
187
|
5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
188
|
+
6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
|
|
181
189
|
|
|
182
190
|
Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
|
|
183
191
|
You may read files for context but this mode never edits anything.
|
|
@@ -41,6 +41,7 @@ and even then, report it, don't ask about it.
|
|
|
41
41
|
|
|
42
42
|
- List files to create/modify (max diff).
|
|
43
43
|
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
|
|
44
|
+
- **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
|
|
44
45
|
|
|
45
46
|
### 3. Test First (RED) — Handoff mode
|
|
46
47
|
|
|
@@ -56,6 +57,7 @@ and even then, report it, don't ask about it.
|
|
|
56
57
|
- Reuse existing code before creating new.
|
|
57
58
|
- No unrelated changes or speculative refactors.
|
|
58
59
|
- Follow project conventions (check `.claude/skills/` for patterns).
|
|
60
|
+
- **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
|
|
59
61
|
|
|
60
62
|
### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
|
|
61
63
|
|
|
@@ -102,6 +102,17 @@ const path = require('path');
|
|
|
102
102
|
const crypto = require('crypto');
|
|
103
103
|
const { pathToFileURL } = require('url');
|
|
104
104
|
|
|
105
|
+
// Wall-clock watchdog. This hook runs on the UserPromptSubmit/PreToolUse hot path but has
|
|
106
|
+
// no host-side process-group kill: when indexed-context resolution hangs (a wedged import,
|
|
107
|
+
// a stalled mount), the host kills only THIS bash wrapper, the node grandchild below is
|
|
108
|
+
// reparented to launchd, and it spins forever. Verified: exactly one orphan leaked per run
|
|
109
|
+
// of the router timeout test, 200+ accumulated over days. Self-exit at the deadline so a
|
|
110
|
+
// hang can never outlive the hook. Advisory only — an early exit just drops the routing
|
|
111
|
+
// hint, the same posture as a cache miss. Same watchdog task-watchdog.sh/handoff-resume.sh
|
|
112
|
+
// already carry; keep this one aligned with them.
|
|
113
|
+
const HOOK_DEADLINE_MS = Number.parseInt(process.env.UKIT_HOOK_DEADLINE_MS || '', 10) || 3000;
|
|
114
|
+
setTimeout(() => process.exit(0), HOOK_DEADLINE_MS).unref();
|
|
115
|
+
|
|
105
116
|
(async () => {
|
|
106
117
|
const STOPWORDS = new Set([
|
|
107
118
|
'the', 'a', 'an', 'and', 'or', 'to', 'for', 'of', 'with', 'in', 'on', 'is', 'are',
|
|
@@ -13,6 +13,16 @@ Project: {{project.name}} | Stack: {{project.stack}}
|
|
|
13
13
|
- Review gate for non-trivial edits (>3 files or >50 lines).
|
|
14
14
|
- Reuse existing code. No drive-by refactors.
|
|
15
15
|
|
|
16
|
+
## Solution Selection Ladder
|
|
17
|
+
|
|
18
|
+
Choose the smallest solution that fully meets the stated requirements and project conventions. When more than one option qualifies, prefer, in order:
|
|
19
|
+
|
|
20
|
+
1. reuse an existing codebase helper with the same semantics;
|
|
21
|
+
2. standard library, native platform API, or an already-installed dependency;
|
|
22
|
+
3. custom code or a new dependency (last resort).
|
|
23
|
+
|
|
24
|
+
Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
|
|
25
|
+
|
|
16
26
|
## Docs-Centric Principle
|
|
17
27
|
|
|
18
28
|
**Docs are the center.** Read docs first, write docs last.
|
|
@@ -35,9 +45,9 @@ Read in order. Never skip to Tier 3 without Tier 1+2.
|
|
|
35
45
|
- Recent `docs/WORKLOG.md` → when continuing prior work or debugging
|
|
36
46
|
|
|
37
47
|
**Tier 2 — Structure Scan** (before opening source files):
|
|
38
|
-
-
|
|
39
|
-
-
|
|
40
|
-
- Narrows exactly which files to open in Tier 3
|
|
48
|
+
- Prefer the project index/resolver to find what exists where and the shape of symbols.
|
|
49
|
+
- Fall back to raw Glob/Grep only when the index is stale, missing, or returns no match.
|
|
50
|
+
- Narrows exactly which files to open in Tier 3; understanding source before editing is never skipped.
|
|
41
51
|
|
|
42
52
|
**Tier 3 — Targeted Source Reads** (only what is needed):
|
|
43
53
|
- Read only the specific files relevant to the task
|
|
@@ -57,7 +67,7 @@ After any non-trivial task, AI must update docs proactively — **do not wait to
|
|
|
57
67
|
## Workflow
|
|
58
68
|
|
|
59
69
|
1. **Orient** — Tier 1: read docs, form hypothesis about the system
|
|
60
|
-
2. **Scan** — Tier 2:
|
|
70
|
+
2. **Scan** — Tier 2: use the project index/resolver to verify the hypothesis and find relevant files; raw Glob/Grep only when the index is stale, missing, or returns no match
|
|
61
71
|
3. **Verify** — if source contradicts docs, update docs first
|
|
62
72
|
4. **Execute** — smallest change, follow existing patterns
|
|
63
73
|
5. **Test** — run tests, check regressions
|
|
@@ -34,6 +34,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
34
34
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
35
35
|
5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
|
|
36
36
|
6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
37
|
+
7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
|
|
38
|
+
- Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
|
|
39
|
+
- Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
|
|
40
|
+
- Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
|
|
41
|
+
- Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
|
|
42
|
+
- Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
|
|
43
|
+
**Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
|
|
37
44
|
|
|
38
45
|
### Severity ladder
|
|
39
46
|
|
|
@@ -170,13 +177,14 @@ the main task already has write + verification evidence; the caller is not waiti
|
|
|
170
177
|
|
|
171
178
|
### Review order
|
|
172
179
|
|
|
173
|
-
Apply the same lenses as Code Review's steps 2-
|
|
180
|
+
Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
|
|
174
181
|
|
|
175
182
|
1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
|
|
176
183
|
2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
|
|
177
184
|
3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
178
185
|
4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
|
|
179
186
|
5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
187
|
+
6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
|
|
180
188
|
|
|
181
189
|
Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
|
|
182
190
|
You may read files for context but this mode never edits anything.
|
|
@@ -40,6 +40,7 @@ and even then, report it, don't ask about it.
|
|
|
40
40
|
|
|
41
41
|
- List files to create/modify (max diff).
|
|
42
42
|
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
|
|
43
|
+
- **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
|
|
43
44
|
|
|
44
45
|
### 3. Test First (RED) — Handoff mode
|
|
45
46
|
|
|
@@ -55,6 +56,7 @@ and even then, report it, don't ask about it.
|
|
|
55
56
|
- Reuse existing code before creating new.
|
|
56
57
|
- No unrelated changes or speculative refactors.
|
|
57
58
|
- Follow project conventions (check `.claude/skills/` for patterns).
|
|
59
|
+
- **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
|
|
58
60
|
|
|
59
61
|
### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
|
|
60
62
|
|