@ngockhoale/ukit 2.3.23 → 2.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,83 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.4.0 - 2026-09-14
6
+
7
+ Solution-discipline wording + offline instruments. This release is **additive**: it changes
8
+ instruction text that is already loaded in daily sessions and adds offline test instruments
9
+ under `tests/`, which the published tarball does not ship. It changes no runtime, config,
10
+ route, hook, or gateway behaviour, and it removes no existing behaviour.
11
+
12
+ - **Solution ladder + carve-outs in the delivery lane.** `templates/.claude/skills/delivery/SKILL.md`
13
+ (and its active mirror) now carry the canonical ladder: choose the smallest solution that fully
14
+ meets the stated requirements and project conventions; prefer reuse of an equivalent codebase
15
+ helper, then the standard library / native API / an already-installed dependency, and treat new
16
+ custom code or a new dependency as the last resort. The carve-out is explicit: never drop
17
+ validation, security, accessibility, error handling, compatibility, or required tests to reduce
18
+ code, and never ship a weaker version and ask. The delivery "Tier 2 — Structure Scan" step now
19
+ prefers the project index/resolver and falls back to raw Glob/Grep only when the index is stale,
20
+ missing, or returns no match — understanding source before editing is never skipped.
21
+ - **Feature-implementer ladder (Claude Code + omp).** `templates/.claude/agents/feature-implementer.md`,
22
+ its active byte-twin, and `templates/.omp/agents/feature-implementer.md` gain the same ladder plus
23
+ a guard reminder that fewer lines are not a reason to drop a guard. The daily-mode test rule is
24
+ left verbatim: the ladder must not become an excuse to cut tests.
25
+ - **Reviewer rubric (Claude Code + omp).** The `code-reviewer` copies gain 5 evidence-first
26
+ solution-fit questions (duplicate semantics / standard library suffices / guard-dropping
27
+ simplification / shared root cause / speculative abstraction) and a rule that says
28
+ **never grade brevity or line count as a quality win**. The sidecar diff lane stays
29
+ non-blocking with no auto-delete; verdict and model-isolation contracts are unchanged.
30
+ - **Offline instruments (not shipped).** A pure metric protocol
31
+ (`tests/benchmarks/solution-discipline/protocol.mjs`), four acceptance fixtures under
32
+ `tests/fixtures/solution-discipline/` (F01 reuse-helper, F06 trust-boundary, F09 keep-failure-path,
33
+ F11 required-interface) with good/bad reference solutions, and a host driver core exercised through
34
+ fake hosts (nonzero exit, timeout, malformed events, duplicate usage, missing terminal event).
35
+ These live under `tests/`, which `package.json.files` does not publish, so they stay out of the
36
+ tarball by construction; the release check confirms the packed file list contains no `tests/` path.
37
+ - **Deferred, with reasons.** The paid pilot (roadmap PT-06) and confirmatory benchmark (PT-09) are
38
+ deferred: no spending cap is approved and no billable model call was made this cycle. The runtime
39
+ lifecycle work (PT-07/PT-08 — shared renderer plus route/lifecycle/omp integration) is deferred
40
+ because the evidence for instruction-only wording is sufficient to start and no runtime drift has
41
+ been observed yet. The remaining fixtures and live-host probing are queued in
42
+ `docs/AI_HANDOFF/PROPOSAL-C18-benchmark-fixtures.md` for the funded pilot cycle.
43
+ - **Evidence record.** The preflight inventory this release builds on is pinned in
44
+ `docs/plans/ponytail-evidence.md` (revision pin, existing/proposed inventory, command map, blockers).
45
+ - **Truthfulness.** Savings are unverified until a paid benchmark pilot runs; no token/LOC reduction
46
+ is claimed for this release.
47
+
48
+ ## 2.3.24 - 2026-09-13
49
+
50
+ Post-2.3.23 routing and self-install audit. Two real defects fixed, plus the repo-local
51
+ Release Policy mirror restored and guarded.
52
+
53
+ - **`isInformationalPrompt` ported into the src canonical copy.** `src/index/taskRouting.js`
54
+ was missing the informational-prompt guard that the shipped copies in
55
+ `templates/.claude/ukit/index/route-task.mjs` and `templates/.claude/hooks/skill-router.sh`
56
+ received in 2.2.14. The 2.3.22 `consultationOnlySignal` only catches question-shaped
57
+ prompts (trailing `?`, a leading interrogative, or Vietnamese `…là gì / được không`), so a
58
+ prompt with an **embedded** interrogative and no leading marker — `"please tell me what the
59
+ auth module does"`, `"explain what files the installer writes"` — routed to `shared-edit`
60
+ with a write+verification contract in src while the shipped templates routed it
61
+ `informational`. That is exactly the fabricated-mutation-debt class 2.3.22 set out to fix,
62
+ still live in the copy that drives `ukit index route`. The helper is now byte-equivalent
63
+ across all three copies; implement/debug-worded questions still keep their lanes. Locked by
64
+ a new `tests/index/taskRouting.test.js` regression (verified RED without the call site).
65
+ - **Self-install detection no longer misfires on a local install.** The 2.3.23 guard treated
66
+ any `templates/` directory inside the install root as a self-install. A downstream project
67
+ that installs UKit locally (`node_modules/@ngockhoale/ukit/templates`) therefore matched,
68
+ and every `docs/` entry was downgraded to `skip` — so `ukit update` silently stopped
69
+ refreshing that project's docs. `isSelfInstall` now requires the package root itself to be
70
+ the project root (`<projectRoot>/templates` exactly) or the target `package.json` to name
71
+ `@ngockhoale/ukit`. Locked by a new `buildPlan` regression for the `node_modules` shape
72
+ (verified RED under the old containment test).
73
+ - **Release Policy mirror restored.** A `ukit install` run inside this repo (self-install)
74
+ had overwritten root `CLAUDE.md`/`AGENTS.md` from `templates/`, wiping the repo-local
75
+ Release Policy section that `docs/PROJECT.md` mandates at the top of both files. Restored
76
+ from the byte-identical pre-wipe copy; a new `tests/consistency/configDocsSync.test.js`
77
+ guard fails loudly if a future self-install wipes either copy again.
78
+ - **Verification.** `yarn test` 87 files / 1,553 tests green (also under `--sequence.shuffle`);
79
+ `yarn release:verify` exit 0; artifact unpacked ceiling raised once to 5,960,000 with
80
+ rationale in `tests/integration/packageArtifact.test.js`.
81
+
5
82
  ## 2.3.23 - 2026-09-13
6
83
 
7
84
  Self-install no longer clobbers UKit's own canonical docs. Running `ukit install` inside the UKit
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.23",
3
+ "version": "2.4.0",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -45,22 +45,25 @@ function safeResolve(basePath, relativePath) {
45
45
  return resolvedPath;
46
46
  }
47
47
 
48
- function isPathInside(parentPath, childPath) {
49
- const relative = path.relative(parentPath, childPath);
50
- return relative === '' || (!relative.startsWith('..') && !path.isAbsolute(relative));
51
- }
52
-
53
48
  function isDocsTarget(targetPath, projectRoot) {
54
49
  const relative = path.relative(projectRoot, targetPath);
55
50
  return relative === 'docs' || relative.startsWith(`docs${path.sep}`);
56
51
  }
57
52
 
58
53
  // Detect `ukit install` running against the UKit source repository itself. Two independent
59
- // signals: the templates resolve inside the install target (the repo has a templates/ dir, the
60
- // global CLI does not), or the target's package.json declares UKIT_PACKAGE_NAME (the path check
61
- // misses the common case where the global CLI runs inside the repo).
54
+ // signals: the templates resolve directly under the install target (`<projectRoot>/templates`
55
+ // — the repo layout, which the global CLI does not have), or the target's package.json
56
+ // declares UKIT_PACKAGE_NAME (the path check misses the common case where the global CLI runs
57
+ // inside the repo).
58
+ //
59
+ // The path signal deliberately requires `<projectRoot>/templates` exactly, not merely
60
+ // "templates somewhere inside projectRoot". A downstream project that installs UKit locally
61
+ // (`node_modules/@ngockhoale/ukit/templates`) also has templates inside its root; treating
62
+ // that as a self-install downgraded every docs entry to `skip`, so `ukit update` never
63
+ // refreshed that project's docs. Only the package root itself being the project root is a
64
+ // self-install.
62
65
  async function isSelfInstall({ templatesRoot, projectRoot }) {
63
- if (isPathInside(projectRoot, templatesRoot)) {
66
+ if (path.dirname(path.resolve(templatesRoot)) === path.resolve(projectRoot)) {
64
67
  return true;
65
68
  }
66
69
  try {
@@ -406,6 +406,56 @@ function deriveExecutionScores({
406
406
  };
407
407
  }
408
408
 
409
+ // A prompt that only asks a question (no command, no action verb, no error report) is
410
+ // asking for an explanation, not a repository change. Routing it to an investigation or
411
+ // edit lane fabricates write/verification debt the completion gate then enforces against an
412
+ // answer turn. Kept byte-equivalent to the shipped copies in
413
+ // templates/.claude/ukit/index/route-task.mjs and templates/.claude/hooks/skill-router.sh;
414
+ // this is the canonical copy those are mirror-tested against. It deliberately ignores
415
+ // targetFile: an edit order phrased as a question is caught by implementWords below, while a
416
+ // pure targeted question ("how does X work?") is still just a question.
417
+ function isInformationalPrompt({
418
+ promptText = '',
419
+ commandText = '',
420
+ scores = null,
421
+ } = {}) {
422
+ if (String(commandText || '').trim()) return false;
423
+ const raw = String(promptText || '').toLowerCase().trim();
424
+ if (!raw) return false;
425
+ const questionSignal = raw.includes('?')
426
+ || /\b(what|how|when|which|where|who|is|are|does|do|did|can|could|would|should|will)\b/.test(raw)
427
+ || /(là\s+(?:[^\s]+\s+){0,2}gì|thế nào|như thế nào|bao nhiêu|khi nào|bao giờ|ở đâu|nghĩa là|được không|không\b)/.test(raw);
428
+ // '?' or a leading/Vietnamese interrogative — required to overrule error wording.
429
+ const strongQuestion = raw.includes('?')
430
+ || /^(what|how|when|which|where|who|is|are|does|do|did|can|could|would|should|will)\b/.test(raw)
431
+ || /(là\s+(?:[^\s]+\s+){0,2}gì|thế nào|như thế nào|bao nhiêu|khi nào|bao giờ|ở đâu)/.test(raw);
432
+ if (scores) {
433
+ if (
434
+ scores.implementSignal
435
+ || scores.reviewSignal
436
+ || scores.debugSignal
437
+ || scores.impactSignal
438
+ || scores.smallFixSignal
439
+ || scores.directTransformSignal
440
+ ) {
441
+ return false;
442
+ }
443
+ if (scores.failureSignal && !strongQuestion) {
444
+ return false;
445
+ }
446
+ }
447
+ const implementWords = /(?<![A-Za-z0-9_])(implement|apply|update|modify|add|create|ship|deliver|fix|refactor|remove|delete|rename|change|write|build|make|install|run|deploy|execute|sửa|thêm|tạo|xóa|đổi|thay thế|cập nhật|viết|build|chạy|cài)(?![A-Za-z0-9_])/.test(raw);
448
+ if (implementWords) return false;
449
+ const investigationWords = /\b(why|debug|triage|root cause|investigate|tại sao)\b/.test(raw);
450
+ if (investigationWords) return false;
451
+ // \b never matches around 'lỗi' (diacritics are not \w), so test it as a plain
452
+ // substring: an error report is informational only when phrased as a question.
453
+ if (raw.includes('lỗi') && !strongQuestion) return false;
454
+ const reviewWords = /\b(review|audit|verify)\b/.test(raw);
455
+ if (reviewWords) return false;
456
+ return questionSignal;
457
+ }
458
+
409
459
  function deriveExecutionMode({
410
460
  promptText = '',
411
461
  commandText = '',
@@ -494,6 +544,10 @@ function deriveExecutionMode({
494
544
  && !scores.sharedRisk
495
545
  && !targetFile;
496
546
 
547
+ if (isInformationalPrompt({ promptText, commandText, scores })) {
548
+ return 'informational';
549
+ }
550
+
497
551
  if (
498
552
  releaseVerificationContinuation
499
553
  || ((intentMode === 'review-specific' || explicitReviewLead) && !scores.implementSignal)
@@ -35,6 +35,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
35
35
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
36
36
  5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
37
37
  6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
38
+ 7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
39
+ - Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
40
+ - Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
41
+ - Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
42
+ - Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
43
+ - Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
44
+ **Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
38
45
 
39
46
  ### Severity ladder
40
47
 
@@ -171,13 +178,14 @@ the main task already has write + verification evidence; the caller is not waiti
171
178
 
172
179
  ### Review order
173
180
 
174
- Apply the same lenses as Code Review's steps 2-6, scoped to what the diff actually touches:
181
+ Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
175
182
 
176
183
  1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
177
184
  2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
178
185
  3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
179
186
  4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
180
187
  5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
188
+ 6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
181
189
 
182
190
  Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
183
191
  You may read files for context but this mode never edits anything.
@@ -41,6 +41,7 @@ and even then, report it, don't ask about it.
41
41
 
42
42
  - List files to create/modify (max diff).
43
43
  - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
44
+ - **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
44
45
 
45
46
  ### 3. Test First (RED) — Handoff mode
46
47
 
@@ -56,6 +57,7 @@ and even then, report it, don't ask about it.
56
57
  - Reuse existing code before creating new.
57
58
  - No unrelated changes or speculative refactors.
58
59
  - Follow project conventions (check `.claude/skills/` for patterns).
60
+ - **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
59
61
 
60
62
  ### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
61
63
 
@@ -13,6 +13,16 @@ Project: {{project.name}} | Stack: {{project.stack}}
13
13
  - Review gate for non-trivial edits (>3 files or >50 lines).
14
14
  - Reuse existing code. No drive-by refactors.
15
15
 
16
+ ## Solution Selection Ladder
17
+
18
+ Choose the smallest solution that fully meets the stated requirements and project conventions. When more than one option qualifies, prefer, in order:
19
+
20
+ 1. reuse an existing codebase helper with the same semantics;
21
+ 2. standard library, native platform API, or an already-installed dependency;
22
+ 3. custom code or a new dependency (last resort).
23
+
24
+ Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
25
+
16
26
  ## Docs-Centric Principle
17
27
 
18
28
  **Docs are the center.** Read docs first, write docs last.
@@ -35,9 +45,9 @@ Read in order. Never skip to Tier 3 without Tier 1+2.
35
45
  - Recent `docs/WORKLOG.md` → when continuing prior work or debugging
36
46
 
37
47
  **Tier 2 — Structure Scan** (before opening source files):
38
- - Glob file tree → know what exists where
39
- - Grep function/class/export signatures → understand shape without full reads
40
- - Narrows exactly which files to open in Tier 3
48
+ - Prefer the project index/resolver to find what exists where and the shape of symbols.
49
+ - Fall back to raw Glob/Grep only when the index is stale, missing, or returns no match.
50
+ - Narrows exactly which files to open in Tier 3; understanding source before editing is never skipped.
41
51
 
42
52
  **Tier 3 — Targeted Source Reads** (only what is needed):
43
53
  - Read only the specific files relevant to the task
@@ -57,7 +67,7 @@ After any non-trivial task, AI must update docs proactively — **do not wait to
57
67
  ## Workflow
58
68
 
59
69
  1. **Orient** — Tier 1: read docs, form hypothesis about the system
60
- 2. **Scan** — Tier 2: Glob + Grep to verify hypothesis, find relevant files
70
+ 2. **Scan** — Tier 2: use the project index/resolver to verify the hypothesis and find relevant files; raw Glob/Grep only when the index is stale, missing, or returns no match
61
71
  3. **Verify** — if source contradicts docs, update docs first
62
72
  4. **Execute** — smallest change, follow existing patterns
63
73
  5. **Test** — run tests, check regressions
@@ -34,6 +34,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
34
34
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
35
35
  5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
36
36
  6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
37
+ 7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
38
+ - Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
39
+ - Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
40
+ - Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
41
+ - Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
42
+ - Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
43
+ **Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
37
44
 
38
45
  ### Severity ladder
39
46
 
@@ -170,13 +177,14 @@ the main task already has write + verification evidence; the caller is not waiti
170
177
 
171
178
  ### Review order
172
179
 
173
- Apply the same lenses as Code Review's steps 2-6, scoped to what the diff actually touches:
180
+ Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
174
181
 
175
182
  1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
176
183
  2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
177
184
  3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
178
185
  4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
179
186
  5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
187
+ 6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
180
188
 
181
189
  Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
182
190
  You may read files for context but this mode never edits anything.
@@ -40,6 +40,7 @@ and even then, report it, don't ask about it.
40
40
 
41
41
  - List files to create/modify (max diff).
42
42
  - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
43
+ - **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
43
44
 
44
45
  ### 3. Test First (RED) — Handoff mode
45
46
 
@@ -55,6 +56,7 @@ and even then, report it, don't ask about it.
55
56
  - Reuse existing code before creating new.
56
57
  - No unrelated changes or speculative refactors.
57
58
  - Follow project conventions (check `.claude/skills/` for patterns).
59
+ - **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
58
60
 
59
61
  ### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
60
62