peaks-loop 4.0.42 → 4.0.44
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +59 -0
- package/README-en.md +1 -1
- package/README.md +1 -1
- package/dist/cli/commands/_register.js +4 -0
- package/dist/cli/commands/api-diff-commands.d.ts +16 -0
- package/dist/cli/commands/api-diff-commands.js +55 -0
- package/dist/cli/commands/audit-commands.d.ts +16 -3
- package/dist/cli/commands/audit-commands.js +84 -31
- package/dist/cli/commands/codegraph-commands.js +191 -6
- package/dist/cli/commands/final-review-commands.d.ts +34 -10
- package/dist/cli/commands/final-review-commands.js +130 -34
- package/dist/cli/commands/job-commands.js +4 -2
- package/dist/cli/commands/scan-commands.js +1 -1
- package/dist/cli/commands/share-commands.d.ts +49 -0
- package/dist/cli/commands/share-commands.js +114 -14
- package/dist/cli/commands/test-commands.d.ts +60 -3
- package/dist/cli/commands/test-commands.js +125 -7
- package/dist/services/audit/audit-goal-service.js +38 -3
- package/dist/services/codegraph/codegraph-autorefresh.js +12 -0
- package/dist/services/codegraph/codegraph-exclude-integrity.d.ts +61 -0
- package/dist/services/codegraph/codegraph-exclude-integrity.js +98 -0
- package/dist/services/codegraph/codegraph-exclude-reconciler.d.ts +26 -0
- package/dist/services/codegraph/codegraph-exclude-reconciler.js +217 -0
- package/dist/services/codegraph/codegraph-exclude-repair.d.ts +102 -0
- package/dist/services/codegraph/codegraph-exclude-repair.js +266 -0
- package/dist/services/codegraph/codegraph-preflight-service.js +12 -0
- package/dist/services/codegraph/codegraph-service.d.ts +0 -1
- package/dist/services/codegraph/codegraph-service.js +5 -4
- package/dist/services/doctor/doctor-service/checks/codegraph-exclude-integrity.d.ts +29 -0
- package/dist/services/doctor/doctor-service/checks/codegraph-exclude-integrity.js +88 -0
- package/dist/services/doctor/doctor-service/checks/ecc-hooks-schema-drift.d.ts +65 -0
- package/dist/services/doctor/doctor-service/checks/ecc-hooks-schema-drift.js +186 -0
- package/dist/services/doctor/doctor-service/plugin-registry.js +4 -0
- package/dist/services/doctor/doctor-service/types.d.ts +47 -0
- package/dist/services/final-review/final-review-service.d.ts +154 -0
- package/dist/services/final-review/final-review-service.js +621 -7
- package/dist/services/final-review/index.d.ts +1 -1
- package/dist/services/final-review/index.js +1 -1
- package/dist/services/llm/anthropic-runner.d.ts +87 -0
- package/dist/services/llm/anthropic-runner.js +171 -0
- package/dist/services/llm/stub-runner.d.ts +11 -0
- package/dist/services/llm/stub-runner.js +33 -0
- package/dist/services/prd/handoff-auto-regen.js +0 -1
- package/dist/services/prd/handoff-service.d.ts +9 -1
- package/dist/services/prd/handoff-service.js +48 -6
- package/dist/services/prd/project-scan-bootstrap-service.js +7 -7
- package/dist/services/scan/api-diff-openapi.d.ts +32 -0
- package/dist/services/scan/api-diff-openapi.js +359 -0
- package/dist/services/scan/api-diff-recorded.d.ts +96 -0
- package/dist/services/scan/api-diff-recorded.js +577 -0
- package/dist/services/scan/api-diff-service.d.ts +34 -0
- package/dist/services/scan/api-diff-service.js +407 -0
- package/dist/services/scan/api-diff-types.d.ts +116 -0
- package/dist/services/scan/api-diff-types.js +46 -0
- package/dist/services/scan/archetype-service.js +27 -1
- package/dist/services/scan/existing-system-service.js +17 -4
- package/dist/services/scan/hook-convention-service.d.ts +26 -0
- package/dist/services/scan/hook-convention-service.js +562 -0
- package/dist/services/scan/scan-types.d.ts +47 -0
- package/dist/services/session/caller-binding-service.d.ts +28 -0
- package/dist/services/session/caller-binding-service.js +10 -2
- package/dist/services/session/caller-id-types.d.ts +12 -2
- package/dist/services/session/index.d.ts +2 -2
- package/dist/services/session/index.js +2 -2
- package/dist/services/session/session-binding-bridge.js +11 -6
- package/dist/services/session/session-manager.d.ts +33 -1
- package/dist/services/session/session-manager.js +84 -25
- package/dist/services/skills/skill-presence-service.d.ts +17 -3
- package/dist/services/skills/skill-presence-service.js +23 -3
- package/package.json +7 -5
- package/skills/bee/peaks-rd/SKILL.md +11 -3
- package/skills/peaks-code/references/existing-system-extraction.md +5 -1
- package/skills/peaks-code/references/frontend-only-mode.md +48 -6
- package/skills/peaks-code/references/project-scan-checklist.md +20 -1
- package/skills/peaks-doctor/references/doctor-check-catalog.md +1 -0
- package/skills/peaks-final-review/SKILL.md +43 -32
|
@@ -21,6 +21,7 @@ Slice L3.2 ships 69 doctor checks. The most user-relevant ones:
|
|
|
21
21
|
## Integration (third-party hooks)
|
|
22
22
|
|
|
23
23
|
- **`integration:gateguard-peaks-conflict`** — warns when `gateguard-fact-force` is installed without a `.peaks/**` skip pattern (the 3rd-party hook would block all peaks-qa .peaks/ artifact writes)
|
|
24
|
+
- **`integration:ecc-hooks-schema-drift`** — warns (cosmetic, never fails the run) when the 3rd-party ECC plugin's `hooks/hooks.json` carries keys Claude Code ignores (`$schema` at the root; `description` + `id` per matcher group), which is what prints `ecc: hooks.json: unknown keys ... ignored` at startup
|
|
24
25
|
|
|
25
26
|
## Skills
|
|
26
27
|
|
|
@@ -69,7 +69,9 @@ interface PrepareFinalReviewOptions {
|
|
|
69
69
|
}
|
|
70
70
|
```
|
|
71
71
|
|
|
72
|
-
The service is the **gate primitive** that closes the 10% human / 90% LLM loop. It reads the approved audit-goal JSON from `.peaks/_runtime/<sessionId>/audit-goal/<rid>.json`,
|
|
72
|
+
The service is the **gate primitive** that closes the 10% human / 90% LLM loop. It reads the approved audit-goal JSON from `.peaks/_runtime/<sessionId>/audit-goal/<rid>.json`, **collects real evidence from disk** (`qa/test-reports`, `qa/test-cases`, `qa/security-findings`, `qa/performance-findings`, `rd/{tech-doc,bug-analysis,code-review,security-review}.md`, `prd/handoff.md`), inlines it into the 4-dim review prompt under byte bounds, calls an injected `LlmRunner` exactly once, parses the response, and validates that all four required dimensions are present. It throws `IncompleteFinalReviewError` on malformed JSON or missing dimensions — callers MUST treat that as a gate failure (return to human for re-prompting) and never let autonomous work proceed on a partial review.
|
|
73
|
+
|
|
74
|
+
> **Evidence-backed verdicts (added 2026-09-12).** The service previously sent the model only `successCriteria` and no evidence at all, so it could not honestly grade anything — a real run returned 4/4 `inconclusive`, and a model willing to confabulate could have returned `pass`. Verdicts are now gated **structurally**, not by prompt wording: any `pass` whose supporting sources were all missing/empty is rewritten to `inconclusive` + `confidence: low` (`fail` is never softened), and `allPass` is derived from the gated verdicts so it can only narrow. Every non-`found` evidence source renders an explicit `STATUS: MISSING (<why>)` in the prompt, so the model always knows what it does not know.
|
|
73
75
|
|
|
74
76
|
> The `LlmRunner` interface is intentionally minimal so this service reuses the same provider injection seam as `audit-goal-service` and the slice LLMArbitrator (`src/services/audit/audit-goal-service.ts:16`). No provider implementation is baked in at this layer.
|
|
75
77
|
|
|
@@ -85,41 +87,32 @@ All of the following MUST be true before invoking this skill:
|
|
|
85
87
|
|
|
86
88
|
If any precondition is missing, **STOP** and route back to the responsible skill. Do not paper over a missing artifact with a hand-written successCriteria — the review must reflect what the human originally approved.
|
|
87
89
|
|
|
88
|
-
## Invocation
|
|
90
|
+
## Invocation
|
|
89
91
|
|
|
90
|
-
> **
|
|
91
|
-
>
|
|
92
|
-
>
|
|
93
|
-
>
|
|
94
|
-
>
|
|
95
|
-
>
|
|
96
|
-
> **This CLI command does NOT yet exist.** `prepareFinalReview()` is implemented as a service in `src/services/final-review/final-review-service.ts` and is unit-tested in `tests/unit/final-review/final-review-service.test.ts`, but it is **not wired to a CLI subcommand**. A `peaks prepare-final-review` subcommand is the planned forward-looking surface; the integration sits at the service layer, not the CLI layer, today.
|
|
97
|
-
>
|
|
98
|
-
> **Pick for the future CLI wrapper file (audit recommendation):** create a new `src/cli/commands/final-review-commands.ts` matching the `peaks-<group>-commands.ts` naming convention (`qa-commands.ts`, `code-review-commands.ts`, `audit-commands.ts`). A grep of `src/cli/commands/` confirms NO `final-review-commands.ts` and NO `prepareFinalReview` import in any CLI file. Rationale: the 4-dim business review is conceptually distinct from `peaks qa *` (which is autonomous gate verification) — it is the human-acceptance terminator, not an internal gate. A separate command group preserves that boundary.
|
|
99
|
-
|
|
100
|
-
### Current call path (today, until a CLI wrapper is built)
|
|
101
|
-
|
|
102
|
-
```ts
|
|
103
|
-
// peaks-code end-of-workflow, peaks-txt, or any other hand-rolled caller
|
|
104
|
-
import { prepareFinalReview } from './src/services/final-review/final-review-service.js';
|
|
105
|
-
import { auditGoalLlmRunner } from './src/services/audit/llm-runner.js'; // or your provider
|
|
106
|
-
|
|
107
|
-
const out = await prepareFinalReview('<rid>', {
|
|
108
|
-
projectRoot: '<absolute path>',
|
|
109
|
-
sessionId: '<sessionId>',
|
|
110
|
-
llmRunner: auditGoalLlmRunner
|
|
111
|
-
});
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
### Planned call path (after the CLI wrapper lands)
|
|
92
|
+
> **Correction (2026-09-12).** An earlier revision of this file claimed `peaks prepare-final-review`
|
|
93
|
+
> "does NOT yet exist" and told callers to hand-roll a `prepareFinalReview()` caller instead.
|
|
94
|
+
> **That was wrong.** The CLI wrapper exists and is registered —
|
|
95
|
+
> `src/cli/commands/final-review-commands.ts` (`W5 Fix M2`), command registered at its line ~130.
|
|
96
|
+
> The hand-rolled snippet that used to sit here also had the wrong flag shape (`--rid <rid>`);
|
|
97
|
+
> **the rid is positional**. Use the CLI.
|
|
115
98
|
|
|
116
99
|
```bash
|
|
117
|
-
|
|
118
|
-
# src/cli/commands/final-review-commands.ts
|
|
119
|
-
peaks prepare-final-review --rid <rid> [--session-id <sid>] --json
|
|
100
|
+
peaks prepare-final-review <rid> [--session-id <sid>] [--project <path>] [--llm-provider <name>] [--json]
|
|
120
101
|
```
|
|
121
102
|
|
|
122
|
-
|
|
103
|
+
- `<rid>` is **positional** (not `--rid`). It resolves to `.peaks/_runtime/<sessionId>/audit-goal/<rid>.json`.
|
|
104
|
+
- `--session-id` defaults to the active workspace binding when omitted.
|
|
105
|
+
- **`--llm-provider` defaults to `stub`, and `stub` is the only provider that works today.**
|
|
106
|
+
`stub` runs no real review — it returns a scaffold envelope so CI can prove the route is reachable.
|
|
107
|
+
Verified 2026-09-12: passing a real provider (`--llm-provider anthropic`) returns
|
|
108
|
+
`status: not-applicable` / `serviceWired: false` / `providerBinding: unknown` and tells you to re-run
|
|
109
|
+
with `stub`; the CLI's own nextAction calls real-provider binding "a follow-up slice". So there is
|
|
110
|
+
currently **no reachable path to a real 4-dim review** — confirming the route works is all `stub`
|
|
111
|
+
can do.
|
|
112
|
+
- `--json` is required for machine consumption (peaks-code, peaks-txt, downstream CI).
|
|
113
|
+
|
|
114
|
+
Calling the service directly (`prepareFinalReview(rid, { projectRoot, sessionId, llmRunner })`) remains
|
|
115
|
+
valid for callers that need a custom `LlmRunner` injection seam, but it is no longer the only path.
|
|
123
116
|
|
|
124
117
|
## Output
|
|
125
118
|
|
|
@@ -155,6 +148,24 @@ Full evidence contract per dimension: `references/4-dimensions.md`.
|
|
|
155
148
|
3. **no-new-bugs** — the regression suite is green AND the LLM surfaces 0 net-new failures (`evidence.kind === 'regression-suite'` + `manual-spot-check`).
|
|
156
149
|
4. **existing-functionality-intact** — a pre/post baseline diff (test count, public API surface, key behavior) shows no unintended drift (`evidence.kind === 'pre-post-diff'`).
|
|
157
150
|
|
|
151
|
+
> **⚠️ This dimension currently cannot pass — read before acting on it (verified 2026-09-12).**
|
|
152
|
+
> `pre-post-diff` is a declared `EvidenceKind`, but **nothing in peaks-loop produces that artifact**.
|
|
153
|
+
> The evidence actually mapped to this dimension is `rd/tech-doc.md` (design intent) and
|
|
154
|
+
> `prd/handoff.md` (scope / non-goals) — neither is a baseline diff, and the model correctly
|
|
155
|
+
> reports that ("the only FOUND source… is a design-intent document, not a regression assessment").
|
|
156
|
+
> `peaks scan api-diff <doc>` is *not* a producer: it diffs an API **document**, not the code surface.
|
|
157
|
+
>
|
|
158
|
+
> **Consequences:** `allPass === true` is **unreachable by construction**, for every workflow.
|
|
159
|
+
> This dimension will return `inconclusive` with an empty `evidence[]` even when the work is
|
|
160
|
+
> perfect. Treat that as a **tooling** state, not as evidence of a regression — and do **not**
|
|
161
|
+
> "fix" it by re-mapping `qa/test-reports` into this dimension's `supports`, which would turn the
|
|
162
|
+
> gate green without producing the baseline diff the definition above requires.
|
|
163
|
+
>
|
|
164
|
+
> **Real fix (unbuilt):** a producer for the pre/post baseline diff (test-count delta, public-API
|
|
165
|
+
> surface snapshot) written to `.peaks/_runtime/<sessionId>/final-review/api-diff.txt`, then mapped
|
|
166
|
+
> into this dimension's `supports`. Until that ships, `needsAttention` always contains this
|
|
167
|
+
> dimension — a permanently-red gate that reviewers will otherwise learn to ignore.
|
|
168
|
+
|
|
158
169
|
## Human's role
|
|
159
170
|
|
|
160
171
|
The human reviews evidence, **judges business outcomes (NOT code)**. The LLM produces structured evidence; the human's job is to:
|
|
@@ -185,7 +196,7 @@ When handing off, emit: rid, `allPass`, `needsAttention[]`, output path, source
|
|
|
185
196
|
| `src/services/final-review/final-review-service.ts` | Authoritative service implementation. |
|
|
186
197
|
| `src/services/final-review/final-review-types.ts` | `FinalReviewOutput`, `DimensionEvidence`, verdict/evidence/confidence enums. |
|
|
187
198
|
| `src/services/audit/audit-goal-service.ts:16` | Line of evidence that `LlmRunner` is reusable across audit + final-review (service-level integration). |
|
|
188
|
-
| `tests/unit/final-review/final-review-service.test.ts` |
|
|
199
|
+
| `tests/unit/final-review/final-review-service.test.ts` | Service-level unit tests (8 cases: evidence inlining, the no-evidence⇒no-`pass` gate, prompt bounds, plus contract guards). |
|
|
189
200
|
| `docs/superpowers/plans/2026-06-25-slice-topology-multipass-phase-4.md:127` | Phase-4 plan prose (Task 14). |
|
|
190
201
|
| `skills/peaks-qa/SKILL.md` | Upstream QA skill — 4-dim review is downstream of all QA gates. |
|
|
191
202
|
| `skills/peaks-audit/SKILL.md` | Sibling skill — produces the `audit-goal` JSON that this skill consumes. |
|