peaks-loop 4.0.44 → 4.0.45
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +41 -0
- package/README-en.md +1 -1
- package/README.md +1 -1
- package/dist/cli/commands/codegraph-commands.d.ts +1 -0
- package/dist/cli/commands/codegraph-commands.js +49 -3
- package/dist/cli/commands/final-review-commands.js +3 -1
- package/dist/services/final-review/final-review-service.d.ts +224 -44
- package/dist/services/final-review/final-review-service.js +934 -97
- package/dist/services/final-review/index.d.ts +2 -1
- package/dist/services/final-review/index.js +2 -1
- package/dist/services/final-review/pre-post-diff.d.ts +137 -0
- package/dist/services/final-review/pre-post-diff.js +657 -0
- package/package.json +5 -5
- package/skills/peaks-final-review/SKILL.md +51 -18
- package/skills/peaks-final-review/references/4-dimensions.md +42 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,46 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 4.0.45 — 2026-09-12 (一个永远无法通过的维度 + 一个验代理的守卫)
|
|
4
|
+
|
|
5
|
+
**Highlights**:
|
|
6
|
+
|
|
7
|
+
1. **4 维验收闸里有一维从构造上就不可能通过。** `existing-functionality-intact` 的契约要求一份 `pre-post-diff`(前后结构对比:测试面 / 公开 API 面),而**全仓没有任何东西生产它**。于是它被喂的是 `rd/tech-doc.md`(设计意图)与 `prd/handoff.md`(批准范围)—— 而这个模块自己的文件头注释,把"拿设计意图文档当回归评估"列为它存在的理由。结果:`allPass === true` 对**任何**工作流都不可达。
|
|
8
|
+
|
|
9
|
+
本版补上生产者(`pre-post-diff.ts`):以 base ref 对比工作树,报**测试面**(文件数 / 用例数)、**源文件清单**、**公开 API 面**(顶层 `export` 语句数与名字集合),写入 `.peaks/_runtime/<sessionId>/final-review/api-diff.txt`。**诚实边界写在产物里**:export 检测是行锚定的近似,能发现"少了一个导出"这类结构丢失,**发现不了签名变更**。算不出 baseline 时**不产出 artifact、不编造 diff** —— 该维保持 `inconclusive`。
|
|
10
|
+
|
|
11
|
+
2. **消费它的那个闸,键在"送达",不在"产物存在"。** 独立复核连做七轮,每轮都在下一层找到同一个形状 —— 一个接受**代理**的闸("评审者看到了吗"被替换成"磁盘上有吗"):
|
|
12
|
+
|
|
13
|
+
| 轮 | 被判为"检查过了"的东西 |
|
|
14
|
+
|---|---|
|
|
15
|
+
| 1 | `status === 'computed'`(磁盘上有文件) |
|
|
16
|
+
| 2 | `status === 'found'`(≥1 字节送达) |
|
|
17
|
+
| 3 | `totalBytes === 0`(把 missing 与 empty 混为一谈) |
|
|
18
|
+
|
|
19
|
+
共同根因是:一个源**可以被部分内联**,于是"送达了吗"永远有一个律师式答案。现在分配器是 **per-source 全有或全无**(要么整份、要么 omitted,不存在截在中间),且 `isDelivered` 是这个问题**唯一的家** —— 每个源声明自己的送达规则:`whole`(整份文档;被截断过的文档是另一份文档)或 `conclusion`(一个**就是结论本身**的字面量)。
|
|
20
|
+
|
|
21
|
+
3. **随之而来的两处实质变化,都要求诚实。**
|
|
22
|
+
- 证据预算改为**推导**(`10,240 = 总额 / 4`)而不是常量 —— 因为走 `whole` 规则的源一旦超过它,就**永远送不到**。
|
|
23
|
+
- 一个维度若其全部支撑源都装不下,现在会被**响亮报出**(prompt 里出 `## Evidence delivery reachability` 段、维度 summary 挂 `[delivery-reachability: …]` 标记、并进 `needsAttention`),而不是**静默地永远红**。**故意不抛错** —— 抛错会连带毁掉那三个本来可送达的维度。
|
|
24
|
+
|
|
25
|
+
4. **反回归守卫本身,就是同一个形状的第 7 个实例 —— 且它一开始是坏的。** 初版守卫按**语法形状**切分源码(只认行首 `function name`),于是**箭头函数、生成器、`export default function`、类方法**一律不被扫描:**8 种等价写法全部放行**,端到端实测"顶层箭头里放 2 个代理字面量"仍 `3 passed EXIT=0`。**一个能放行 8/8 的守卫比没有守卫更糟** —— 它把"没人在看"变成"有东西在看"。
|
|
26
|
+
|
|
27
|
+
现在它用仓库自带的 `typescript` 解析 AST、读**函数体内**的 token,于是那 4 种形状绕过全部关闭(实测:往真实源码注入箭头式代理 → 守卫转红)。剩下 4 种**不含该字面量**的改写(拆变量 `const F='found'`、`==`、`item['status']`、字符串拼接)**仍然放行** —— 这一点被**固定成一条通过的测试**,让守卫**陈述自己的边界**而不是暗示一个。
|
|
28
|
+
|
|
29
|
+
5. **`peaks codegraph status` 在真有缺口时会同时印出两行互相矛盾的话。** 退出码(74)与 `[FAIL]` 细节都是对的,但人第一眼看到的是:
|
|
30
|
+
|
|
31
|
+
```
|
|
32
|
+
[OK] Index is up to date
|
|
33
|
+
[FAIL] codegraph index is incomplete: 2 of 1133 tracked source files are excluded …
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
两个闸在回答不同的问题:上游答"图与上次扫描一致"(确实一致),peaks 答"图覆盖了仓库吗"(没有)。**都对,并排印出来就是矛盾。** 现在上游那行被**归因而非抑制**(措辞与统计保留,只是把它的适用范围说明白),只在这一行、且只在缺口门**确实触发**时改动;无缺口路径逐字节不变(含 ANSI),机器可读 envelope 未动。
|
|
37
|
+
|
|
38
|
+
**验证**:三个版本常量一致(**4.0.45**);`tsc -p tsconfig.build.json` exit 0;宽 `tsconfig.json` 保持 **142** 基线;`tests/unit` **214 files / 2165 passed / 3 skipped / 0 failed**;`pnpm build` 的 `build-integrity` OK。
|
|
39
|
+
|
|
40
|
+
**已知未修 / 边界**:`conclusion` 规则下超大的源不算"不可送达"(结论标记可能在截断点之后);单一可送达源的**穿越前余量**(本 session 实测 748 B)**没有**告警,只报"已经不可能"的状态;4 维中的 `existing-functionality-intact` 在**非 git 项目**里永久 `inconclusive`(这是诚实结果,已写进 SKILL.md 与 `4-dimensions.md`)。
|
|
41
|
+
|
|
42
|
+
**备注**:本版的 `prepare-final-review` 闸**未经过独立复核**(前六轮每一轮都经过并各找到 1 个 HIGH;第 7 轮由 orchestrator 自行验证,含箭头注入实测),这一点如实记录以便日后回溯。
|
|
43
|
+
|
|
3
44
|
## 4.0.44 — 2026-09-12 (被静默排除的源文件 + 一个验错东西的验收闸)
|
|
4
45
|
|
|
5
46
|
**Highlights**:
|
package/README-en.md
CHANGED
|
@@ -140,7 +140,7 @@ Every lane opens with **one slash command**.
|
|
|
140
140
|
|
|
141
141
|
| | |
|
|
142
142
|
| --- | --- |
|
|
143
|
-
| **Latest** | [](https://www.npmjs.com/package/peaks-loop) — 4.0.
|
|
143
|
+
| **Latest** | [](https://www.npmjs.com/package/peaks-loop) — 4.0.45 (2026-09-12) |
|
|
144
144
|
| **Domains** | Code (`peaks-code`) · Content (`peaks-content`) · Project health (`peaks-doctor`) · Issue sweep (`peaks-issue-fix-orchestrator`) · Custom SOP (`peaks-sop`) · Cross-domain primitives (`peaks-solo` dispatcher · `peaks-resume` · `peaks-status` · `peaks-test` · `peaks-slice-decompose`) |
|
|
145
145
|
| **Sediment pool** | `~/.peaks/` local pool · twice-clean runs auto-promote to a bee · broken runs come back for you to redefine · the bee grows with your taste |
|
|
146
146
|
| **Test suite** | 285+ cases · 4 packages (peaks-loop / peaks-loop-mut / peaks-loop-shared-channel / peaks-loop-shared) · **0 timeouts** · 14 BDD caller-binding edge cases |
|
package/README.md
CHANGED
|
@@ -140,7 +140,7 @@ npm i -g peaks-loop
|
|
|
140
140
|
|
|
141
141
|
| | |
|
|
142
142
|
| --- | --- |
|
|
143
|
-
| **最新版本** | [](https://www.npmjs.com/package/peaks-loop) — 4.0.
|
|
143
|
+
| **最新版本** | [](https://www.npmjs.com/package/peaks-loop) — 4.0.45(2026-09-12) |
|
|
144
144
|
| **覆盖域** | 代码(`peaks-code`) · 内容(`peaks-content`) · 项目健康(`peaks-doctor`) · 批量修 issue(`peaks-issue-fix-orchestrator`) · 自定义 SOP(`peaks-sop`) · 通用原语(`peaks-solo` 分诊 / `peaks-resume` 续 / `peaks-status` 看 / `peaks-test` 测 / `peaks-slice-decompose` 切片) |
|
|
145
145
|
| **沉淀池** | `~/.peaks/` 本地池 · 跑两次自动晋升成 bee · 跑翻车让你重定义 · bee 跟着你的口味长 |
|
|
146
146
|
| **测试套件** | 1096 cases · 4 packages (peaks-loop 1015 / runtime 39 / mut 22 / shared-channel 20) · **CI 首次全绿**(ubuntu + windows) · 14 BDD caller-binding coverage |
|
|
@@ -10,4 +10,5 @@ import { type ProgramIO } from '../cli-helpers.js';
|
|
|
10
10
|
* untouched.
|
|
11
11
|
*/
|
|
12
12
|
export declare function rewriteBareCodegraphHints(text: string): string;
|
|
13
|
+
export declare function attributeUpstreamUpToDateLine(stdout: string): string;
|
|
13
14
|
export declare function registerCodegraphCommands(program: Command, io: ProgramIO): void;
|
|
@@ -38,7 +38,48 @@ function printCodegraphFailure(io, command, error, asJson, exitCode = 1) {
|
|
|
38
38
|
export function rewriteBareCodegraphHints(text) {
|
|
39
39
|
return text.replace(/(?<![\w-])(?<!peaks\s)codegraph(?=\s+(?:status|init|index|query|files|context|affected)\b)/g, 'peaks codegraph');
|
|
40
40
|
}
|
|
41
|
-
|
|
41
|
+
const ANSI_SGR_PATTERN = /\x1b\[[0-9;]*m/g;
|
|
42
|
+
/**
|
|
43
|
+
* Upstream `status` answers a different question than peaks-loop's
|
|
44
|
+
* integrity gate: upstream says "the on-disk graph matches the last scan"
|
|
45
|
+
* (true), peaks says "that graph covers the repository" (false when rules
|
|
46
|
+
* exclude tracked files). Both verdicts are correct, but an unqualified
|
|
47
|
+
* `[OK] Index is up to date` printed above our `[FAIL] ...` reads as
|
|
48
|
+
* "nothing to see here" — and the OK is the line the eye lands on first.
|
|
49
|
+
* The exit code and the JSON envelope are already right; only this line
|
|
50
|
+
* lies by juxtaposition.
|
|
51
|
+
*
|
|
52
|
+
* So: keep upstream's wording — the line stays recognizable, and the
|
|
53
|
+
* Files/Nodes counts around it are untouched — but drop the bare OK
|
|
54
|
+
* marker and name the only question it answers. Clean runs never reach
|
|
55
|
+
* this, so their output stays byte-identical.
|
|
56
|
+
*
|
|
57
|
+
* The match is anchored to THAT line. An earlier version keyed on
|
|
58
|
+
* `includes('up to date')`, which is content-blind: upstream prints other
|
|
59
|
+
* `[OK] ... are up to date` lines (a language-server or watcher line is the
|
|
60
|
+
* observed one), and each of those was rewritten into a claim about the
|
|
61
|
+
* INDEX — a misattribution introduced by a change whose entire purpose was
|
|
62
|
+
* to stop misleading output. Anything that is not the index line is passed
|
|
63
|
+
* through byte-for-byte, tail note and all.
|
|
64
|
+
*/
|
|
65
|
+
const INDEX_UP_TO_DATE_RE = /^\[OK\]\s+Index is up to date\b/i;
|
|
66
|
+
export function attributeUpstreamUpToDateLine(stdout) {
|
|
67
|
+
return stdout
|
|
68
|
+
.split('\n')
|
|
69
|
+
.map((line) => {
|
|
70
|
+
const visible = line.replace(ANSI_SGR_PATTERN, '').trim();
|
|
71
|
+
if (!INDEX_UP_TO_DATE_RE.test(visible)) {
|
|
72
|
+
return line;
|
|
73
|
+
}
|
|
74
|
+
// Only the OK marker is downgraded and the attribution appended: the
|
|
75
|
+
// rest of the line — including whatever upstream wrote after it — is
|
|
76
|
+
// preserved, so nothing upstream actually said is replaced.
|
|
77
|
+
const withoutOk = visible.replace(/^\[OK\]\s*/, '');
|
|
78
|
+
return `[i] ${withoutOk} (upstream: matches the last scan only; repository coverage is answered below)`;
|
|
79
|
+
})
|
|
80
|
+
.join('\n');
|
|
81
|
+
}
|
|
82
|
+
async function runCodegraphCommand(io, command, options, asJson, attributeStdout) {
|
|
42
83
|
try {
|
|
43
84
|
const invocation = createCodegraphInvocation(options);
|
|
44
85
|
const result = await executeCodegraphInvocation(invocation);
|
|
@@ -47,7 +88,8 @@ async function runCodegraphCommand(io, command, options, asJson) {
|
|
|
47
88
|
return;
|
|
48
89
|
}
|
|
49
90
|
const didFail = result.exitCode !== null && result.exitCode !== 0;
|
|
50
|
-
const
|
|
91
|
+
const rewritten = rewriteBareCodegraphHints(result.stdout);
|
|
92
|
+
const stdout = attributeStdout === undefined ? rewritten : attributeStdout(rewritten);
|
|
51
93
|
const stderr = rewriteBareCodegraphHints(result.stderr);
|
|
52
94
|
if (stdout.length > 0) {
|
|
53
95
|
io.stdout((didFail ? redactSensitiveErrorMessage(stdout) : stdout).trimEnd());
|
|
@@ -132,7 +174,11 @@ async function runCodegraphStatusCommand(io, options, asJson) {
|
|
|
132
174
|
await runCodegraphStatusJson(io, options, integrity, integrityWarning);
|
|
133
175
|
}
|
|
134
176
|
else {
|
|
135
|
-
|
|
177
|
+
// Only when the gate found a gap: upstream's `[OK] Index is up to
|
|
178
|
+
// date` answers "consistent with the last scan", and printing it
|
|
179
|
+
// unqualified right above our `[FAIL]` tells the reader two opposite
|
|
180
|
+
// things at once. Clean runs get no transform and stay byte-identical.
|
|
181
|
+
await runCodegraphCommand(io, 'codegraph.status', { subcommand: 'status', project: options.project }, false, integrity?.gap === true ? attributeUpstreamUpToDateLine : undefined);
|
|
136
182
|
if (integrityWarning !== null) {
|
|
137
183
|
io.stdout(`[WARN] codegraph exclude integrity not evaluated: ${integrityWarning}`);
|
|
138
184
|
}
|
|
@@ -113,7 +113,8 @@ export function registerFinalReviewCommands(program, io) {
|
|
|
113
113
|
.description('Prepare the 4-dimension business review (final-review primitive) for human acceptance (W2 T9 service; CLI surface in W5 M2)')
|
|
114
114
|
.requiredOption('--project <path>', 'target project root')
|
|
115
115
|
.requiredOption('--session-id <sid>', 'session id whose .peaks/_runtime/<sid>/audit-goal/<rid>.json is the approved goal source')
|
|
116
|
-
.option('--llm-provider <name>', `LLM provider name: ${SUPPORTED_LLM_PROVIDERS.join(' | ')} (default: ${DEFAULT_LLM_PROVIDER} — performs no review)`, DEFAULT_LLM_PROVIDER)
|
|
116
|
+
.option('--llm-provider <name>', `LLM provider name: ${SUPPORTED_LLM_PROVIDERS.join(' | ')} (default: ${DEFAULT_LLM_PROVIDER} — performs no review)`, DEFAULT_LLM_PROVIDER)
|
|
117
|
+
.option('--base <ref>', 'base ref for the pre/post baseline diff (`existing-functionality-intact`); default: merge-base with origin/HEAD, then origin/main, then origin/master, then HEAD~1 — pass this explicitly when none of those resolve')).action(async (rid, options) => {
|
|
117
118
|
// 1. Project root must exist and be a directory.
|
|
118
119
|
const projectValidation = validateProjectRoot(options.project);
|
|
119
120
|
if (!projectValidation.ok) {
|
|
@@ -192,6 +193,7 @@ export function registerFinalReviewCommands(program, io) {
|
|
|
192
193
|
projectRoot: projectValidation.projectRoot,
|
|
193
194
|
sessionId: sessionValidation.sessionId,
|
|
194
195
|
llmRunner,
|
|
196
|
+
...(options.base === undefined ? {} : { baseRef: options.base }),
|
|
195
197
|
});
|
|
196
198
|
const data = {
|
|
197
199
|
status: 'review-complete',
|
|
@@ -9,12 +9,18 @@ export interface LlmRunner {
|
|
|
9
9
|
};
|
|
10
10
|
}>;
|
|
11
11
|
}
|
|
12
|
-
import type { FinalReviewOutput } from './final-review-types.js';
|
|
12
|
+
import type { DimensionKind, FinalReviewOutput } from './final-review-types.js';
|
|
13
13
|
import type { CapabilityAuditResult } from '../capability-audit-service/types.js';
|
|
14
14
|
export interface PrepareFinalReviewOptions {
|
|
15
15
|
readonly projectRoot: string;
|
|
16
16
|
readonly sessionId: string;
|
|
17
17
|
readonly llmRunner: LlmRunner;
|
|
18
|
+
/**
|
|
19
|
+
* Explicit base ref for the pre/post baseline diff. Unset ⇒ the producer
|
|
20
|
+
* resolves the merge-base with the upstream default branch itself, and
|
|
21
|
+
* reports the dimension `unavailable` when it cannot.
|
|
22
|
+
*/
|
|
23
|
+
readonly baseRef?: string;
|
|
18
24
|
}
|
|
19
25
|
export declare class IncompleteFinalReviewError extends Error {
|
|
20
26
|
readonly code: "INCOMPLETE_FINAL_REVIEW";
|
|
@@ -43,52 +49,91 @@ export declare class EmptyReviewReplyError extends Error {
|
|
|
43
49
|
*/
|
|
44
50
|
export declare const MAX_EMPTY_REPLY_ATTEMPTS = 3;
|
|
45
51
|
/**
|
|
46
|
-
*
|
|
47
|
-
* (`rd/tech-doc.md`, `rd/code-review.md`, `qa/*-findings-*.md`); 8 KB keeps the
|
|
48
|
-
* head of every file (header + verdict + first tables) without letting one
|
|
49
|
-
* verbose artifact crowd out the other sources. Enforced in BYTES against the
|
|
50
|
-
* raw buffer, so multi-byte (CJK) content cannot slip past the cap.
|
|
51
|
-
*/
|
|
52
|
-
export declare const MAX_EVIDENCE_BYTES_PER_FILE: number;
|
|
53
|
-
/**
|
|
54
|
-
* Total evidence budget across all sources. 32 KB ≈ 8k tokens of input, which
|
|
52
|
+
* Total evidence budget across all sources. 40 KB ≈ 10k tokens of input, which
|
|
55
53
|
* keeps the prompt far inside any modern context window. Sources that do not
|
|
56
54
|
* fit are reported as OMITTED — never dropped silently.
|
|
57
55
|
*
|
|
58
|
-
* This is an INPUT cap and stays fixed. The output ceiling that has to
|
|
59
|
-
* opposite it is derived per call by `outputBudgetForEvidence()` below —
|
|
60
|
-
* two used to drift apart, and that drift was the defect.
|
|
56
|
+
* This is an INPUT cap and stays fixed per run. The output ceiling that has to
|
|
57
|
+
* sit opposite it is derived per call by `outputBudgetForEvidence()` below —
|
|
58
|
+
* the two used to drift apart, and that drift was the defect.
|
|
59
|
+
*
|
|
60
|
+
* RE-EVALUATED 2026-09-12 (F-BLOCK) — 32 KiB did NOT hold once the tenth
|
|
61
|
+
* source (`final-review-pre-post-diff`) was appended. Measured on this repo's
|
|
62
|
+
* own run (`2026-09-12-session-e37ef0`, rid
|
|
63
|
+
* `2026-09-12-codegraph-exclude-integrity`), the ten sources are 2,621 / 8,164
|
|
64
|
+
* / 9,492 / 10,839 / 11,422 / 13,050 / 13,462 / 13,852 / 17,745 / 20,543
|
|
65
|
+
* bytes — 121,190 bytes on disk, 76,321 bytes once the 8 KiB per-file cap is
|
|
66
|
+
* applied. Four sources at the cap spent the old 32,768 to the byte, so the
|
|
67
|
+
* appended tenth (added LAST, by design) was the first block the budget
|
|
68
|
+
* dropped — on every run, and on the compact set QA measured too (9 x 3.8 KB
|
|
69
|
+
* ≈ 34 KB, which the old cap already could not hold).
|
|
70
|
+
*
|
|
71
|
+
* 40 KiB is a budget that fits the measured ten-source set at the allocator's
|
|
72
|
+
* per-source unit (5 x 8 KiB units + the smaller ones), and it is the largest
|
|
73
|
+
* this budget may grow to today: the derived output ceiling is
|
|
74
|
+
* `3000 + 12288 + bytes/4`, so it reaches MAX_OUTPUT_TOKENS (32,000) at exactly
|
|
75
|
+
* 66,848 bytes of inlined evidence and is clamped from there on — a cap at or
|
|
76
|
+
* above that point buys the reviewer no more output room at all.
|
|
77
|
+
*
|
|
78
|
+
* F-CAP — an earlier version of this comment claimed 40 KiB (40,960 =
|
|
79
|
+
* 10 x 4,096) was "the smallest cap that affords one floor-sized slice to EVERY
|
|
80
|
+
* source in the ten-source set". That was arithmetic about a budget, not a
|
|
81
|
+
* statement about the allocator: sources are capped at
|
|
82
|
+
* MAX_EVIDENCE_BYTES_PER_FILE (10 KiB each, derived — see below) while a floor
|
|
83
|
+
* is per DIMENSION (four of them), so 10 x 4,096 was never what any source
|
|
84
|
+
* received. Re-measured on
|
|
85
|
+
* the same two saturated fixtures (QA round 4: the 9 x 40 KB fixture and this
|
|
86
|
+
* repo's own file sizes), raising the cap from 32 KiB to 40 KiB left **4 of the
|
|
87
|
+
* 10 sources inlined with zero bytes** — `rd/security-review`,
|
|
88
|
+
* `rd/bug-analysis`, `prd/handoff` and the appended
|
|
89
|
+
* `final-review-pre-post-diff` — and it did not deliver the baseline either;
|
|
90
|
+
* all it did was change WHICH source was starved (at 32 KiB that source was
|
|
91
|
+
* `rd/code-review`).
|
|
92
|
+
*
|
|
93
|
+
* So read this constant as "how much evidence fits", never as "which sources
|
|
94
|
+
* arrive". Arrival is decided by the allocator below (one whole unit per source,
|
|
95
|
+
* with a dimension's floor reserving its holder's unit) and, for the two sources
|
|
96
|
+
* a dimension's verdict may not outlive, by the delivery gates on top of it.
|
|
97
|
+
* And note the honest consequence, stated rather than papered over: on a
|
|
98
|
+
* saturated run the allocator still omits sources, so `prd/handoff.md` and the
|
|
99
|
+
* pre/post diff can be dropped — which is exactly why
|
|
100
|
+
* `enforceScopeContractDelivery()` and `enforcePrePostDiffAvailability()` exist
|
|
101
|
+
* and why a `pass` on those two dimensions is keyed on delivery, not on
|
|
102
|
+
* existence.
|
|
61
103
|
*/
|
|
62
104
|
export declare const MAX_EVIDENCE_BYTES_TOTAL: number;
|
|
63
105
|
/**
|
|
64
|
-
*
|
|
65
|
-
*
|
|
66
|
-
*
|
|
67
|
-
*
|
|
68
|
-
*
|
|
69
|
-
*
|
|
70
|
-
*
|
|
71
|
-
*
|
|
72
|
-
*
|
|
73
|
-
*
|
|
74
|
-
*
|
|
75
|
-
*
|
|
76
|
-
*
|
|
77
|
-
*
|
|
78
|
-
*
|
|
79
|
-
*
|
|
80
|
-
*
|
|
81
|
-
*
|
|
82
|
-
*
|
|
83
|
-
*
|
|
84
|
-
*
|
|
85
|
-
*
|
|
86
|
-
*
|
|
87
|
-
*
|
|
88
|
-
*
|
|
89
|
-
*
|
|
90
|
-
|
|
91
|
-
|
|
106
|
+
* Per-file evidence cap — and the allocator's UNIT. DERIVED, never typed.
|
|
107
|
+
*
|
|
108
|
+
* F5 — this used to be the literal `8 * 1024`, and the whole floor guarantee
|
|
109
|
+
* below was silently conditioned on `4 x perFile <= total`: the reservation
|
|
110
|
+
* promises every pending holder its own unit, and that promise holds only while
|
|
111
|
+
* the four units fit the budget together. At `8 * 1024` the inequality held by
|
|
112
|
+
* coincidence (4 x 8,192 = 32,768 <= 40,960) and nothing in the code said so —
|
|
113
|
+
* raising the cap past 10,240 would have starved all four dimensions in the
|
|
114
|
+
* same run, which reads as four independent red gates rather than one broken
|
|
115
|
+
* constant. The cap is therefore DIVIDED OUT OF the total: the relation is
|
|
116
|
+
* definitional instead of remembered, and `assertFloorReservationAffordable()`
|
|
117
|
+
* still checks it at the point of use (the division must also be exact).
|
|
118
|
+
*
|
|
119
|
+
* The value is 10,240 because that is `40,960 / 4` — the largest unit the
|
|
120
|
+
* four-dimension reservation can afford. It is also the smallest cap at which
|
|
121
|
+
* the decisive evidence source of this repo's own run can be delivered WHOLE:
|
|
122
|
+
* `qa/test-reports/<rid>.md` measured 9,492 bytes there, and `whole` is the
|
|
123
|
+
* delivery rule for every source whose producer publishes no conclusion literal
|
|
124
|
+
* (see `isDelivered`), so a cap under 9,492 makes `problem-resolution` and
|
|
125
|
+
* `no-new-bugs` undeliverable on every run — the always-red gate this module
|
|
126
|
+
* refuses to ship. Enforced in BYTES against the raw buffer, so multi-byte
|
|
127
|
+
* (CJK) content cannot slip past the cap.
|
|
128
|
+
*
|
|
129
|
+
* A source is inlined as `min(bytes, this cap)` — its whole unit — or not at
|
|
130
|
+
* all. So a `TRUNCATED` source block can only ever mean "the FILE is bigger
|
|
131
|
+
* than this cap"; it can no longer mean "the budget ran out while this file was
|
|
132
|
+
* being copied in". That distinction is the point of the all-or-nothing
|
|
133
|
+
* allocator below: the second reading is what let one byte of a 4,226-byte
|
|
134
|
+
* artifact be counted as delivered evidence (F-BLOCK-1BYTE).
|
|
135
|
+
*/
|
|
136
|
+
export declare const MAX_EVIDENCE_BYTES_PER_FILE: number;
|
|
92
137
|
/**
|
|
93
138
|
* Floor — also the value that shipped before this fix, so no evidence set can
|
|
94
139
|
* end up with a smaller budget than it had. ~3000 tokens is enough for the
|
|
@@ -110,9 +155,12 @@ export declare const MIN_OUTPUT_TOKENS = 3000;
|
|
|
110
155
|
* truncated on 2 of 10 real runs.
|
|
111
156
|
*
|
|
112
157
|
* 12288 (12 KiB) is sized so the largest evidence pack the input caps allow
|
|
113
|
-
* lands at
|
|
114
|
-
*
|
|
115
|
-
*
|
|
158
|
+
* lands at 25528 (see the formula below; the input cap is 40 KiB as of the
|
|
159
|
+
* F-BLOCK re-evaluation, and this floor was checked against the raised cap, not
|
|
160
|
+
* against the 32 KiB the measurements above were taken at — a 12 KiB headroom
|
|
161
|
+
* under a bigger cap is the conservative direction) — about 1.9x the largest
|
|
162
|
+
* value ever OBSERVED to truncate (13240), which is the margin the observed
|
|
163
|
+
* variance asks for. The numbers are in the block comment above.
|
|
116
164
|
*/
|
|
117
165
|
export declare const REASONING_HEADROOM_TOKENS: number;
|
|
118
166
|
/**
|
|
@@ -174,6 +222,137 @@ export declare function outputBudgetForEvidence(includedEvidenceBytes: number):
|
|
|
174
222
|
* still gets a call made.
|
|
175
223
|
*/
|
|
176
224
|
export declare function resolveOutputBudget(includedEvidenceBytes: number, env?: NodeJS.ProcessEnv): number;
|
|
225
|
+
/**
|
|
226
|
+
* What DELIVERED means for one source — the module's ONE delivery definition,
|
|
227
|
+
* declared per source and read in exactly one place (`isDelivered`).
|
|
228
|
+
*
|
|
229
|
+
* whole the reviewer must have received the whole document. A
|
|
230
|
+
* truncated slice is not a weaker version of a document, it is a
|
|
231
|
+
* DIFFERENT document, and the conclusion may be in the part that
|
|
232
|
+
* was cut — which is precisely the state `found` used to call
|
|
233
|
+
* "delivered".
|
|
234
|
+
* conclusion the reviewer must have received the literal that IS the
|
|
235
|
+
* document's conclusion. Used where the producer publishes one
|
|
236
|
+
* (the pre/post diff opens with its `VERDICT:` line).
|
|
237
|
+
*
|
|
238
|
+
* `whole` is the rule wherever the producer is another role's skill and
|
|
239
|
+
* publishes no conclusion literal: the module may not GUESS where a document's
|
|
240
|
+
* conclusion lives. The two are the same judgement — "did the reviewer receive
|
|
241
|
+
* the conclusion" — checked at the only place the module can check it.
|
|
242
|
+
*/
|
|
243
|
+
type DeliveryRule = {
|
|
244
|
+
readonly kind: 'whole';
|
|
245
|
+
} | {
|
|
246
|
+
readonly kind: 'conclusion';
|
|
247
|
+
readonly marker: string;
|
|
248
|
+
};
|
|
249
|
+
interface EvidenceSource {
|
|
250
|
+
/** Stable id quoted by the model in its citations. */
|
|
251
|
+
readonly key: string;
|
|
252
|
+
readonly label: string;
|
|
253
|
+
/** Path segments under `.peaks/_runtime/<sessionId>/`. */
|
|
254
|
+
readonly segments: readonly string[];
|
|
255
|
+
/** Dimensions this source can supply evidence for. */
|
|
256
|
+
readonly supports: readonly DimensionKind[];
|
|
257
|
+
/** What DELIVERED means for this source. See `isDelivered()` — every source
|
|
258
|
+
* must declare one, and the declaration is the only thing the module's
|
|
259
|
+
* delivery judgement reads. */
|
|
260
|
+
readonly delivery: DeliveryRule;
|
|
261
|
+
}
|
|
262
|
+
/**
|
|
263
|
+
* `found` is the only status that carries bytes. The other four exist so the
|
|
264
|
+
* prompt can name *why* a source carries nothing: an absent file, one that
|
|
265
|
+
* exists but is empty, one that exists but could not be READ, and one that did
|
|
266
|
+
* not fit the byte budget are four different facts, and the model is told all
|
|
267
|
+
* four explicitly.
|
|
268
|
+
*
|
|
269
|
+
* F4 — `missing` and `unreadable` used to be one status. `readFileSync`'s catch
|
|
270
|
+
* collapsed EACCES / EBUSY / EPERM into the same `raw === null` as ENOENT, so a
|
|
271
|
+
* contract file that EXISTS and could not be opened was reported to the
|
|
272
|
+
* reviewer — and, worse, to the delivery gate — as "there was no PRD phase".
|
|
273
|
+
* Those are opposite facts about a run: one says "nothing to deliver", the
|
|
274
|
+
* other says "there is something to deliver and it did not arrive".
|
|
275
|
+
*/
|
|
276
|
+
type EvidenceStatus = 'found' | 'empty' | 'missing' | 'unreadable' | 'omitted';
|
|
277
|
+
interface CollectedEvidence {
|
|
278
|
+
readonly source: EvidenceSource;
|
|
279
|
+
readonly relativePath: string;
|
|
280
|
+
readonly absolutePath: string;
|
|
281
|
+
readonly status: EvidenceStatus;
|
|
282
|
+
/** Full size on disk (0 when nothing could be read). */
|
|
283
|
+
readonly totalBytes: number;
|
|
284
|
+
/** Bytes actually inlined into the prompt. */
|
|
285
|
+
readonly includedBytes: number;
|
|
286
|
+
readonly content: string;
|
|
287
|
+
/** Why this source carries no evidence (non-`found` statuses only). */
|
|
288
|
+
readonly reason: string;
|
|
289
|
+
}
|
|
290
|
+
/**
|
|
291
|
+
* F5 — the floor promise, checked where it is relied on instead of remembered.
|
|
292
|
+
*
|
|
293
|
+
* The allocator's loop invariant is `budgetLeft >= sum(units of pending
|
|
294
|
+
* holders)`, and its INITIAL condition is
|
|
295
|
+
* `REQUIRED_DIMENSIONS.length x MAX_EVIDENCE_BYTES_PER_FILE <=
|
|
296
|
+
* MAX_EVIDENCE_BYTES_TOTAL`. While that holds, every holder is served when it
|
|
297
|
+
* is reached and the reservation is a guarantee; the moment it fails, all four
|
|
298
|
+
* dimensions are starved in the same run — which surfaces as four independent
|
|
299
|
+
* red gates rather than as one broken constant, and is therefore the kind of
|
|
300
|
+
* breakage nobody diagnoses correctly.
|
|
301
|
+
*
|
|
302
|
+
* The cap is derived from the total so the inequality cannot be typed wrong,
|
|
303
|
+
* and this check covers the two ways a derivation can still go bad: a
|
|
304
|
+
* non-integer quotient (a fifth dimension, say) and a future re-typing of
|
|
305
|
+
* either constant. It is exported because a test asserts it directly.
|
|
306
|
+
*/
|
|
307
|
+
export declare function assertFloorReservationAffordable(): void;
|
|
308
|
+
/**
|
|
309
|
+
* H2 — the reason a dimension is red PERMANENTLY, as opposed to merely unfed on
|
|
310
|
+
* this run. Reported per source, so the arithmetic is checkable by the reader.
|
|
311
|
+
*/
|
|
312
|
+
export interface UndeliverableDimensionEvidence {
|
|
313
|
+
readonly dimension: DimensionKind;
|
|
314
|
+
readonly sources: readonly {
|
|
315
|
+
readonly key: string;
|
|
316
|
+
readonly relativePath: string;
|
|
317
|
+
readonly totalBytes: number;
|
|
318
|
+
}[];
|
|
319
|
+
}
|
|
320
|
+
/**
|
|
321
|
+
* H2 — every required dimension whose verdict is locked to `inconclusive` by
|
|
322
|
+
* BYTE ARITHMETIC rather than by the reviewer's judgement.
|
|
323
|
+
*
|
|
324
|
+
* The defect this exists for: `qa/test-reports/<rid>.md` measured 9,492 bytes
|
|
325
|
+
* on this repo's own run and is the ONLY source on disk for `problem-resolution`
|
|
326
|
+
* and `no-new-bugs` (the other sources that support them are absent), against a
|
|
327
|
+
* per-file cap of 10,240 — a 748-byte margin on a file that is REWRITTEN every
|
|
328
|
+
* round and only grows. The moment it crosses 10,240 both dimensions go
|
|
329
|
+
* permanently red, and NOTHING said so: `assertFloorReservationAffordable()`
|
|
330
|
+
* only checks the constant-level relation (`4 x cap <= total`), never whether
|
|
331
|
+
* any actual source fits the cap it must live under, so the red handoff read as
|
|
332
|
+
* "the reviewer was unsure" instead of "no evidence can ever reach the
|
|
333
|
+
* reviewer". A gate that is always red and never explains itself is noise, and
|
|
334
|
+
* this is the same "always red" harm the floor reservation was built to remove
|
|
335
|
+
* — one layer down.
|
|
336
|
+
*
|
|
337
|
+
* The conditions, all three of which must hold, are chosen so the report cannot
|
|
338
|
+
* be noise:
|
|
339
|
+
* 1. NOTHING on disk can back the dimension (no `isDelivered`), and
|
|
340
|
+
* 2. there IS something on disk to deliver (a source that was never written
|
|
341
|
+
* is not a delivery failure — same reasoning as the scope-contract gate's
|
|
342
|
+
* `missing` exemption: a run with no QA phase has no report to lose), and
|
|
343
|
+
* 3. EVERY one of those on-disk sources is structurally undeliverable — a
|
|
344
|
+
* single source that merely did not fit TODAY (budget-exhausted `omitted`)
|
|
345
|
+
* is a different, self-correcting state and is left to the allocator.
|
|
346
|
+
*
|
|
347
|
+
* The report is consumed by the prompt (stated to the reviewer), by the
|
|
348
|
+
* envelope (a marker on each dimension's summary) and by `needsAttention`
|
|
349
|
+
* (which also clears `allPass`) — see `enforceDeliveryReachability` and
|
|
350
|
+
* `renderDeliveryReachabilityStatus`. It is a LOUD STATEMENT, not a silent
|
|
351
|
+
* downgrade, and it is deliberately not a throw: a crash would destroy the
|
|
352
|
+
* evidence for the three dimensions that ARE deliverable, and the honest fact
|
|
353
|
+
* here is per-dimension, so it is reported per-dimension.
|
|
354
|
+
*/
|
|
355
|
+
export declare function undeliverableDimensions(collected: readonly CollectedEvidence[]): readonly UndeliverableDimensionEvidence[];
|
|
177
356
|
export declare function prepareFinalReview(rid: string, opts: PrepareFinalReviewOptions): Promise<FinalReviewOutput>;
|
|
178
357
|
export declare function decideFifthDimension(input: {
|
|
179
358
|
readonly audit: CapabilityAuditResult | null;
|
|
@@ -182,3 +361,4 @@ export declare function decideFifthDimension(input: {
|
|
|
182
361
|
readonly verdict: 'pass' | 'fail' | 'inconclusive';
|
|
183
362
|
readonly reason: string;
|
|
184
363
|
};
|
|
364
|
+
export {};
|