@ngockhoale/ukit 2.3.23 → 2.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +77 -0
- package/package.json +1 -1
- package/src/core/buildPlan.js +12 -9
- package/src/index/taskRouting.js +54 -0
- package/templates/.claude/agents/code-reviewer.md +9 -1
- package/templates/.claude/agents/feature-implementer.md +2 -0
- package/templates/.claude/skills/delivery/SKILL.md +14 -4
- package/templates/.omp/agents/code-reviewer.md +9 -1
- package/templates/.omp/agents/feature-implementer.md +2 -0
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,83 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.4.0 - 2026-09-14
|
|
6
|
+
|
|
7
|
+
Solution-discipline wording + offline instruments. This release is **additive**: it changes
|
|
8
|
+
instruction text that is already loaded in daily sessions and adds offline test instruments
|
|
9
|
+
under `tests/`, which the published tarball does not ship. It changes no runtime, config,
|
|
10
|
+
route, hook, or gateway behaviour, and it removes no existing behaviour.
|
|
11
|
+
|
|
12
|
+
- **Solution ladder + carve-outs in the delivery lane.** `templates/.claude/skills/delivery/SKILL.md`
|
|
13
|
+
(and its active mirror) now carry the canonical ladder: choose the smallest solution that fully
|
|
14
|
+
meets the stated requirements and project conventions; prefer reuse of an equivalent codebase
|
|
15
|
+
helper, then the standard library / native API / an already-installed dependency, and treat new
|
|
16
|
+
custom code or a new dependency as the last resort. The carve-out is explicit: never drop
|
|
17
|
+
validation, security, accessibility, error handling, compatibility, or required tests to reduce
|
|
18
|
+
code, and never ship a weaker version and ask. The delivery "Tier 2 — Structure Scan" step now
|
|
19
|
+
prefers the project index/resolver and falls back to raw Glob/Grep only when the index is stale,
|
|
20
|
+
missing, or returns no match — understanding source before editing is never skipped.
|
|
21
|
+
- **Feature-implementer ladder (Claude Code + omp).** `templates/.claude/agents/feature-implementer.md`,
|
|
22
|
+
its active byte-twin, and `templates/.omp/agents/feature-implementer.md` gain the same ladder plus
|
|
23
|
+
a guard reminder that fewer lines are not a reason to drop a guard. The daily-mode test rule is
|
|
24
|
+
left verbatim: the ladder must not become an excuse to cut tests.
|
|
25
|
+
- **Reviewer rubric (Claude Code + omp).** The `code-reviewer` copies gain 5 evidence-first
|
|
26
|
+
solution-fit questions (duplicate semantics / standard library suffices / guard-dropping
|
|
27
|
+
simplification / shared root cause / speculative abstraction) and a rule that says
|
|
28
|
+
**never grade brevity or line count as a quality win**. The sidecar diff lane stays
|
|
29
|
+
non-blocking with no auto-delete; verdict and model-isolation contracts are unchanged.
|
|
30
|
+
- **Offline instruments (not shipped).** A pure metric protocol
|
|
31
|
+
(`tests/benchmarks/solution-discipline/protocol.mjs`), four acceptance fixtures under
|
|
32
|
+
`tests/fixtures/solution-discipline/` (F01 reuse-helper, F06 trust-boundary, F09 keep-failure-path,
|
|
33
|
+
F11 required-interface) with good/bad reference solutions, and a host driver core exercised through
|
|
34
|
+
fake hosts (nonzero exit, timeout, malformed events, duplicate usage, missing terminal event).
|
|
35
|
+
These live under `tests/`, which `package.json.files` does not publish, so they stay out of the
|
|
36
|
+
tarball by construction; the release check confirms the packed file list contains no `tests/` path.
|
|
37
|
+
- **Deferred, with reasons.** The paid pilot (roadmap PT-06) and confirmatory benchmark (PT-09) are
|
|
38
|
+
deferred: no spending cap is approved and no billable model call was made this cycle. The runtime
|
|
39
|
+
lifecycle work (PT-07/PT-08 — shared renderer plus route/lifecycle/omp integration) is deferred
|
|
40
|
+
because the evidence for instruction-only wording is sufficient to start and no runtime drift has
|
|
41
|
+
been observed yet. The remaining fixtures and live-host probing are queued in
|
|
42
|
+
`docs/AI_HANDOFF/PROPOSAL-C18-benchmark-fixtures.md` for the funded pilot cycle.
|
|
43
|
+
- **Evidence record.** The preflight inventory this release builds on is pinned in
|
|
44
|
+
`docs/plans/ponytail-evidence.md` (revision pin, existing/proposed inventory, command map, blockers).
|
|
45
|
+
- **Truthfulness.** Savings are unverified until a paid benchmark pilot runs; no token/LOC reduction
|
|
46
|
+
is claimed for this release.
|
|
47
|
+
|
|
48
|
+
## 2.3.24 - 2026-09-13
|
|
49
|
+
|
|
50
|
+
Post-2.3.23 routing and self-install audit. Two real defects fixed, plus the repo-local
|
|
51
|
+
Release Policy mirror restored and guarded.
|
|
52
|
+
|
|
53
|
+
- **`isInformationalPrompt` ported into the src canonical copy.** `src/index/taskRouting.js`
|
|
54
|
+
was missing the informational-prompt guard that the shipped copies in
|
|
55
|
+
`templates/.claude/ukit/index/route-task.mjs` and `templates/.claude/hooks/skill-router.sh`
|
|
56
|
+
received in 2.2.14. The 2.3.22 `consultationOnlySignal` only catches question-shaped
|
|
57
|
+
prompts (trailing `?`, a leading interrogative, or Vietnamese `…là gì / được không`), so a
|
|
58
|
+
prompt with an **embedded** interrogative and no leading marker — `"please tell me what the
|
|
59
|
+
auth module does"`, `"explain what files the installer writes"` — routed to `shared-edit`
|
|
60
|
+
with a write+verification contract in src while the shipped templates routed it
|
|
61
|
+
`informational`. That is exactly the fabricated-mutation-debt class 2.3.22 set out to fix,
|
|
62
|
+
still live in the copy that drives `ukit index route`. The helper is now byte-equivalent
|
|
63
|
+
across all three copies; implement/debug-worded questions still keep their lanes. Locked by
|
|
64
|
+
a new `tests/index/taskRouting.test.js` regression (verified RED without the call site).
|
|
65
|
+
- **Self-install detection no longer misfires on a local install.** The 2.3.23 guard treated
|
|
66
|
+
any `templates/` directory inside the install root as a self-install. A downstream project
|
|
67
|
+
that installs UKit locally (`node_modules/@ngockhoale/ukit/templates`) therefore matched,
|
|
68
|
+
and every `docs/` entry was downgraded to `skip` — so `ukit update` silently stopped
|
|
69
|
+
refreshing that project's docs. `isSelfInstall` now requires the package root itself to be
|
|
70
|
+
the project root (`<projectRoot>/templates` exactly) or the target `package.json` to name
|
|
71
|
+
`@ngockhoale/ukit`. Locked by a new `buildPlan` regression for the `node_modules` shape
|
|
72
|
+
(verified RED under the old containment test).
|
|
73
|
+
- **Release Policy mirror restored.** A `ukit install` run inside this repo (self-install)
|
|
74
|
+
had overwritten root `CLAUDE.md`/`AGENTS.md` from `templates/`, wiping the repo-local
|
|
75
|
+
Release Policy section that `docs/PROJECT.md` mandates at the top of both files. Restored
|
|
76
|
+
from the byte-identical pre-wipe copy; a new `tests/consistency/configDocsSync.test.js`
|
|
77
|
+
guard fails loudly if a future self-install wipes either copy again.
|
|
78
|
+
- **Verification.** `yarn test` 87 files / 1,553 tests green (also under `--sequence.shuffle`);
|
|
79
|
+
`yarn release:verify` exit 0; artifact unpacked ceiling raised once to 5,960,000 with
|
|
80
|
+
rationale in `tests/integration/packageArtifact.test.js`.
|
|
81
|
+
|
|
5
82
|
## 2.3.23 - 2026-09-13
|
|
6
83
|
|
|
7
84
|
Self-install no longer clobbers UKit's own canonical docs. Running `ukit install` inside the UKit
|
package/package.json
CHANGED
package/src/core/buildPlan.js
CHANGED
|
@@ -45,22 +45,25 @@ function safeResolve(basePath, relativePath) {
|
|
|
45
45
|
return resolvedPath;
|
|
46
46
|
}
|
|
47
47
|
|
|
48
|
-
function isPathInside(parentPath, childPath) {
|
|
49
|
-
const relative = path.relative(parentPath, childPath);
|
|
50
|
-
return relative === '' || (!relative.startsWith('..') && !path.isAbsolute(relative));
|
|
51
|
-
}
|
|
52
|
-
|
|
53
48
|
function isDocsTarget(targetPath, projectRoot) {
|
|
54
49
|
const relative = path.relative(projectRoot, targetPath);
|
|
55
50
|
return relative === 'docs' || relative.startsWith(`docs${path.sep}`);
|
|
56
51
|
}
|
|
57
52
|
|
|
58
53
|
// Detect `ukit install` running against the UKit source repository itself. Two independent
|
|
59
|
-
// signals: the templates resolve
|
|
60
|
-
// global CLI does not), or the target's package.json
|
|
61
|
-
// misses the common case where the global CLI runs
|
|
54
|
+
// signals: the templates resolve directly under the install target (`<projectRoot>/templates`
|
|
55
|
+
// — the repo layout, which the global CLI does not have), or the target's package.json
|
|
56
|
+
// declares UKIT_PACKAGE_NAME (the path check misses the common case where the global CLI runs
|
|
57
|
+
// inside the repo).
|
|
58
|
+
//
|
|
59
|
+
// The path signal deliberately requires `<projectRoot>/templates` exactly, not merely
|
|
60
|
+
// "templates somewhere inside projectRoot". A downstream project that installs UKit locally
|
|
61
|
+
// (`node_modules/@ngockhoale/ukit/templates`) also has templates inside its root; treating
|
|
62
|
+
// that as a self-install downgraded every docs entry to `skip`, so `ukit update` never
|
|
63
|
+
// refreshed that project's docs. Only the package root itself being the project root is a
|
|
64
|
+
// self-install.
|
|
62
65
|
async function isSelfInstall({ templatesRoot, projectRoot }) {
|
|
63
|
-
if (
|
|
66
|
+
if (path.dirname(path.resolve(templatesRoot)) === path.resolve(projectRoot)) {
|
|
64
67
|
return true;
|
|
65
68
|
}
|
|
66
69
|
try {
|
package/src/index/taskRouting.js
CHANGED
|
@@ -406,6 +406,56 @@ function deriveExecutionScores({
|
|
|
406
406
|
};
|
|
407
407
|
}
|
|
408
408
|
|
|
409
|
+
// A prompt that only asks a question (no command, no action verb, no error report) is
|
|
410
|
+
// asking for an explanation, not a repository change. Routing it to an investigation or
|
|
411
|
+
// edit lane fabricates write/verification debt the completion gate then enforces against an
|
|
412
|
+
// answer turn. Kept byte-equivalent to the shipped copies in
|
|
413
|
+
// templates/.claude/ukit/index/route-task.mjs and templates/.claude/hooks/skill-router.sh;
|
|
414
|
+
// this is the canonical copy those are mirror-tested against. It deliberately ignores
|
|
415
|
+
// targetFile: an edit order phrased as a question is caught by implementWords below, while a
|
|
416
|
+
// pure targeted question ("how does X work?") is still just a question.
|
|
417
|
+
function isInformationalPrompt({
|
|
418
|
+
promptText = '',
|
|
419
|
+
commandText = '',
|
|
420
|
+
scores = null,
|
|
421
|
+
} = {}) {
|
|
422
|
+
if (String(commandText || '').trim()) return false;
|
|
423
|
+
const raw = String(promptText || '').toLowerCase().trim();
|
|
424
|
+
if (!raw) return false;
|
|
425
|
+
const questionSignal = raw.includes('?')
|
|
426
|
+
|| /\b(what|how|when|which|where|who|is|are|does|do|did|can|could|would|should|will)\b/.test(raw)
|
|
427
|
+
|| /(là\s+(?:[^\s]+\s+){0,2}gì|thế nào|như thế nào|bao nhiêu|khi nào|bao giờ|ở đâu|nghĩa là|được không|không\b)/.test(raw);
|
|
428
|
+
// '?' or a leading/Vietnamese interrogative — required to overrule error wording.
|
|
429
|
+
const strongQuestion = raw.includes('?')
|
|
430
|
+
|| /^(what|how|when|which|where|who|is|are|does|do|did|can|could|would|should|will)\b/.test(raw)
|
|
431
|
+
|| /(là\s+(?:[^\s]+\s+){0,2}gì|thế nào|như thế nào|bao nhiêu|khi nào|bao giờ|ở đâu)/.test(raw);
|
|
432
|
+
if (scores) {
|
|
433
|
+
if (
|
|
434
|
+
scores.implementSignal
|
|
435
|
+
|| scores.reviewSignal
|
|
436
|
+
|| scores.debugSignal
|
|
437
|
+
|| scores.impactSignal
|
|
438
|
+
|| scores.smallFixSignal
|
|
439
|
+
|| scores.directTransformSignal
|
|
440
|
+
) {
|
|
441
|
+
return false;
|
|
442
|
+
}
|
|
443
|
+
if (scores.failureSignal && !strongQuestion) {
|
|
444
|
+
return false;
|
|
445
|
+
}
|
|
446
|
+
}
|
|
447
|
+
const implementWords = /(?<![A-Za-z0-9_])(implement|apply|update|modify|add|create|ship|deliver|fix|refactor|remove|delete|rename|change|write|build|make|install|run|deploy|execute|sửa|thêm|tạo|xóa|đổi|thay thế|cập nhật|viết|build|chạy|cài)(?![A-Za-z0-9_])/.test(raw);
|
|
448
|
+
if (implementWords) return false;
|
|
449
|
+
const investigationWords = /\b(why|debug|triage|root cause|investigate|tại sao)\b/.test(raw);
|
|
450
|
+
if (investigationWords) return false;
|
|
451
|
+
// \b never matches around 'lỗi' (diacritics are not \w), so test it as a plain
|
|
452
|
+
// substring: an error report is informational only when phrased as a question.
|
|
453
|
+
if (raw.includes('lỗi') && !strongQuestion) return false;
|
|
454
|
+
const reviewWords = /\b(review|audit|verify)\b/.test(raw);
|
|
455
|
+
if (reviewWords) return false;
|
|
456
|
+
return questionSignal;
|
|
457
|
+
}
|
|
458
|
+
|
|
409
459
|
function deriveExecutionMode({
|
|
410
460
|
promptText = '',
|
|
411
461
|
commandText = '',
|
|
@@ -494,6 +544,10 @@ function deriveExecutionMode({
|
|
|
494
544
|
&& !scores.sharedRisk
|
|
495
545
|
&& !targetFile;
|
|
496
546
|
|
|
547
|
+
if (isInformationalPrompt({ promptText, commandText, scores })) {
|
|
548
|
+
return 'informational';
|
|
549
|
+
}
|
|
550
|
+
|
|
497
551
|
if (
|
|
498
552
|
releaseVerificationContinuation
|
|
499
553
|
|| ((intentMode === 'review-specific' || explicitReviewLead) && !scores.implementSignal)
|
|
@@ -35,6 +35,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
35
35
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
36
36
|
5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
|
|
37
37
|
6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
38
|
+
7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
|
|
39
|
+
- Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
|
|
40
|
+
- Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
|
|
41
|
+
- Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
|
|
42
|
+
- Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
|
|
43
|
+
- Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
|
|
44
|
+
**Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
|
|
38
45
|
|
|
39
46
|
### Severity ladder
|
|
40
47
|
|
|
@@ -171,13 +178,14 @@ the main task already has write + verification evidence; the caller is not waiti
|
|
|
171
178
|
|
|
172
179
|
### Review order
|
|
173
180
|
|
|
174
|
-
Apply the same lenses as Code Review's steps 2-
|
|
181
|
+
Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
|
|
175
182
|
|
|
176
183
|
1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
|
|
177
184
|
2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
|
|
178
185
|
3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
179
186
|
4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
|
|
180
187
|
5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
188
|
+
6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
|
|
181
189
|
|
|
182
190
|
Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
|
|
183
191
|
You may read files for context but this mode never edits anything.
|
|
@@ -41,6 +41,7 @@ and even then, report it, don't ask about it.
|
|
|
41
41
|
|
|
42
42
|
- List files to create/modify (max diff).
|
|
43
43
|
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
|
|
44
|
+
- **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
|
|
44
45
|
|
|
45
46
|
### 3. Test First (RED) — Handoff mode
|
|
46
47
|
|
|
@@ -56,6 +57,7 @@ and even then, report it, don't ask about it.
|
|
|
56
57
|
- Reuse existing code before creating new.
|
|
57
58
|
- No unrelated changes or speculative refactors.
|
|
58
59
|
- Follow project conventions (check `.claude/skills/` for patterns).
|
|
60
|
+
- **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
|
|
59
61
|
|
|
60
62
|
### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
|
|
61
63
|
|
|
@@ -13,6 +13,16 @@ Project: {{project.name}} | Stack: {{project.stack}}
|
|
|
13
13
|
- Review gate for non-trivial edits (>3 files or >50 lines).
|
|
14
14
|
- Reuse existing code. No drive-by refactors.
|
|
15
15
|
|
|
16
|
+
## Solution Selection Ladder
|
|
17
|
+
|
|
18
|
+
Choose the smallest solution that fully meets the stated requirements and project conventions. When more than one option qualifies, prefer, in order:
|
|
19
|
+
|
|
20
|
+
1. reuse an existing codebase helper with the same semantics;
|
|
21
|
+
2. standard library, native platform API, or an already-installed dependency;
|
|
22
|
+
3. custom code or a new dependency (last resort).
|
|
23
|
+
|
|
24
|
+
Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
|
|
25
|
+
|
|
16
26
|
## Docs-Centric Principle
|
|
17
27
|
|
|
18
28
|
**Docs are the center.** Read docs first, write docs last.
|
|
@@ -35,9 +45,9 @@ Read in order. Never skip to Tier 3 without Tier 1+2.
|
|
|
35
45
|
- Recent `docs/WORKLOG.md` → when continuing prior work or debugging
|
|
36
46
|
|
|
37
47
|
**Tier 2 — Structure Scan** (before opening source files):
|
|
38
|
-
-
|
|
39
|
-
-
|
|
40
|
-
- Narrows exactly which files to open in Tier 3
|
|
48
|
+
- Prefer the project index/resolver to find what exists where and the shape of symbols.
|
|
49
|
+
- Fall back to raw Glob/Grep only when the index is stale, missing, or returns no match.
|
|
50
|
+
- Narrows exactly which files to open in Tier 3; understanding source before editing is never skipped.
|
|
41
51
|
|
|
42
52
|
**Tier 3 — Targeted Source Reads** (only what is needed):
|
|
43
53
|
- Read only the specific files relevant to the task
|
|
@@ -57,7 +67,7 @@ After any non-trivial task, AI must update docs proactively — **do not wait to
|
|
|
57
67
|
## Workflow
|
|
58
68
|
|
|
59
69
|
1. **Orient** — Tier 1: read docs, form hypothesis about the system
|
|
60
|
-
2. **Scan** — Tier 2:
|
|
70
|
+
2. **Scan** — Tier 2: use the project index/resolver to verify the hypothesis and find relevant files; raw Glob/Grep only when the index is stale, missing, or returns no match
|
|
61
71
|
3. **Verify** — if source contradicts docs, update docs first
|
|
62
72
|
4. **Execute** — smallest change, follow existing patterns
|
|
63
73
|
5. **Test** — run tests, check regressions
|
|
@@ -34,6 +34,13 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
34
34
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
35
35
|
5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
|
|
36
36
|
6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
37
|
+
7. **Solution fit** — For each added block, ask with evidence and cite the construct + why (never the prose length):
|
|
38
|
+
- Does new code duplicate an existing helper with the same semantics? A same-semantics duplicate is `CHANGES-REQUESTED` — reuse the existing helper.
|
|
39
|
+
- Would the standard library or a native platform API satisfy every stated requirement? If yes, the custom implementation is avoidable.
|
|
40
|
+
- Does the simplification drop validation, security, accessibility, error handling, compatibility, or a required test? If yes it is a regression, not a simplification.
|
|
41
|
+
- Does a short diff fix the true shared root cause, or only one symptom or caller? Fix the root cause and its affected callers.
|
|
42
|
+
- Is there an abstraction serving only a hypothetical, speculative need? Remove it unless a stated requirement needs it.
|
|
43
|
+
**Never grade brevity or line count as a quality win** — a shorter diff is not evidence of a better change, and a longer one that fully meets requirements is not a finding. Every solution-fit finding must cite `file:line` + the construct + the reason, or it stays a hypothesis and must not drive code deletion.
|
|
37
44
|
|
|
38
45
|
### Severity ladder
|
|
39
46
|
|
|
@@ -170,13 +177,14 @@ the main task already has write + verification evidence; the caller is not waiti
|
|
|
170
177
|
|
|
171
178
|
### Review order
|
|
172
179
|
|
|
173
|
-
Apply the same lenses as Code Review's steps 2-
|
|
180
|
+
Apply the same lenses as Code Review's steps 2-7, scoped to what the diff actually touches:
|
|
174
181
|
|
|
175
182
|
1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
|
|
176
183
|
2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
|
|
177
184
|
3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
178
185
|
4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
|
|
179
186
|
5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
187
|
+
6. **Solution fit** — Apply the same 5 evidence-first questions as Code Review step 7 (duplicate semantics / standard library or native suffices / simplification dropping validation, security, accessibility or tests / shared root cause / speculative abstraction). **Never grade brevity or line count as a quality win.** Findings must cite the construct + reason, and here they remain advisory hypotheses only — never a delete-list, never an edit.
|
|
180
188
|
|
|
181
189
|
Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
|
|
182
190
|
You may read files for context but this mode never edits anything.
|
|
@@ -40,6 +40,7 @@ and even then, report it, don't ask about it.
|
|
|
40
40
|
|
|
41
41
|
- List files to create/modify (max diff).
|
|
42
42
|
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
|
|
43
|
+
- **Choose the smallest solution that fully meets the stated requirements and project conventions.** When more than one option qualifies, prefer, in order: (1) reuse an existing codebase helper with the same semantics; (2) standard library, native platform API, or an already-installed dependency; (3) custom code or a new dependency (last resort). Never drop validation, security, accessibility, error handling, compatibility, or required tests to reduce code. A bugfix fixes the root cause and the affected callers. If the small option is insufficient, build the full option — never ship a weaker version and ask. No speculative abstraction for imagined needs.
|
|
43
44
|
|
|
44
45
|
### 3. Test First (RED) — Handoff mode
|
|
45
46
|
|
|
@@ -55,6 +56,7 @@ and even then, report it, don't ask about it.
|
|
|
55
56
|
- Reuse existing code before creating new.
|
|
56
57
|
- No unrelated changes or speculative refactors.
|
|
57
58
|
- Follow project conventions (check `.claude/skills/` for patterns).
|
|
59
|
+
- **Guard the smallest solution:** never weaken validation, security, accessibility, error handling, or required tests just to shorten the diff — fewer lines is not a reason to drop a guard.
|
|
58
60
|
|
|
59
61
|
### 5. Verify (REQUIRED before DONE in Handoff mode; targeted in Daily mode)
|
|
60
62
|
|