codecartographer-pi 0.23.0 → 0.24.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -26,7 +26,21 @@ not replace any phase; it tells phases where to look.
26
26
 
27
27
  ## Reading a Broad-Side run
28
28
 
29
- 1. Read `synthesis.md` first. It carries the executive summary, severity counts,
29
+ 0. If `verified.md` exists, read it first. A verification pass
30
+ (`codecarto_broadside {action: "verify"}`, `/codecarto-broadside verify`)
31
+ has read the top defect and security findings against the real source with
32
+ read-only tools and given each a verdict: **confirmed** (a reachable
33
+ failure, with the input or call site that triggers it), **not-a-defect**
34
+ (the claim is literally true of the code but nothing can reach the failure
35
+ it describes), **discarded** (the claim is wrong about the code, with the
36
+ guard or line that shows it), or **unclear**. Start from the confirmed
37
+ ones; treat a discarded one as answered unless the reasoning is thin.
38
+ Measured on this repository, the top twelve findings by severity were two
39
+ real defects and ten that a look at the guard, the caller, or the tsconfig
40
+ dismissed — the pass agreed with a reviewer on all twelve for about a cent
41
+ a finding. A confirmed verdict is still a model's reading: a strong lead for
42
+ a human's next look, not a validated claim.
43
+ 1. Read `synthesis.md`. It carries the executive summary, severity counts,
30
44
  the top cross-lens findings, and per-module risk levels.
31
45
  2. Read `triage.md` for the work order: each lead scored by impact ×
32
46
  difficulty with a P0–P3 priority and an effort estimate. It is a starting
@@ -66,6 +80,7 @@ Broad-Side is an executable-surface feature. On the Pi extension:
66
80
  /codecarto-broadside collect # poll, save, synthesize
67
81
  /codecarto-broadside status # show recorded runs
68
82
  /codecarto-broadside models # compare batch models
83
+ /codecarto-broadside verify --top=10 # read the top findings against the source
69
84
  ```
70
85
 
71
86
  On the MCP server:
@@ -75,6 +90,7 @@ codecarto_broadside {cwd, action: "submit", lenses: [...]} # fire the batches
75
90
  codecarto_broadside {cwd, action: "collect"} # poll, save, synthesize
76
91
  codecarto_broadside {cwd, action: "status"} # show recorded runs
77
92
  codecarto_broadside {cwd, action: "models"} # compare batch models
93
+ codecarto_broadside {cwd, action: "verify", top: 10} # read the top findings against the source
78
94
  ```
79
95
 
80
96
  The `models` action lists every `:batch` variant on OpenRouter — pricing per
@@ -105,6 +121,18 @@ Collect runs two cross-lens post-passes by default: **synthesis** (the
105
121
  executive report) and **triage** (the prioritized work order). Pass
106
122
  `include_synthesis: false` or `include_triage: false` on collect to skip one.
107
123
 
124
+ `verify` is the third pass, run separately after collect because it is
125
+ sync-priced rather than batch-priced: for each of the top `top` findings
126
+ (default 10, most severe first) of the defect and security lenses, one chat
127
+ completion on the run's model without its `:batch` suffix (`model` overrides),
128
+ with three read-only tools confined to the repository's source files —
129
+ `read_file` by line range, `grep`, `list_dir` — at most eight tool calls, low
130
+ reasoning effort, then a verdict. It writes `verified.md` and `verified.json`
131
+ beside `triage.md` and records the pass on the run. `max_cost` is a running
132
+ cap here, since a sync call's cost is known only when it returns: the pass
133
+ stops before the next finding once the calls so far have reached it and
134
+ reports `partial`. About a cent a finding on the default model.
135
+
108
136
  Two caveats apply to any model you pick. The `models` action lists every id
109
137
  OpenRouter advertises a `:batch` variant for, and many of those variants do not
110
138
  exist — submitting one returns `does not have a :batch endpoint`, with nothing in
@@ -3,4 +3,4 @@
3
3
  # workspace's framework-owned files (GUIDE.md, templates/, workflow/ pipelines
4
4
  # and VALIDATE.md) predate the running release. Written at release time and
5
5
  # copied verbatim by init — never edit by hand.
6
- scaffold_version: 0.23.0
6
+ scaffold_version: 0.24.0
package/README.md CHANGED
@@ -394,13 +394,16 @@ codecarto_broadside {cwd, action: "models"} # compare batch mo
394
394
  codecarto_broadside {cwd, action: "submit", lenses: [...]} # fire the batches, priced first
395
395
  codecarto_broadside {cwd, action: "status"} # what is in flight
396
396
  codecarto_broadside {cwd, action: "collect"} # poll, save, synthesize, triage
397
+ codecarto_broadside {cwd, action: "verify", top: 10} # read the top findings against the source
397
398
  ```
398
399
 
400
+ A third action, `verify`, reads the top defect and security findings of a collected run against the real source — one sync-priced call each with read-only tools (`read_file`, `grep`, `list_dir`) confined to the repository — and writes `verified.md` with a verdict per finding: **confirmed** (a reachable failure, with the trigger), **not-a-defect**, **discarded**, or **unclear**. Measured on this repository, the top twelve findings by severity were two real defects and ten claims a look at the guard or the caller dismissed; the pass agreed with a reviewer on all twelve for a cent a finding. Read `verified.md` before `triage.md`.
401
+
399
402
  Submit and collect are separate because batch jobs routinely take tens of minutes; collect is resumable and picks up whatever is still in flight (`wait_seconds: 0`, the default, polls once and returns), and two collects on one run — a retried tool call, a second session — never pay for the synthesis, triage, or truncation retry twice: each is claimed in the run's state before it is submitted, and a collect whose client has gone away stops polling and submits nothing further. Submit prices the run from the collected file sizes against the model's live per-token pricing (cached 24h) and refuses when the estimate exceeds `max_cost` — $1.00 unless the config or the call sets another value, `0` for no limit — unless `force: true` is passed — a pre-flight estimate, not a runtime stop. Actual spend lands in each run's `run-meta.json`.
400
403
 
401
404
  Repository defaults live in `.codecarto/broadside/config.yaml` (`model`, `api_key`, `default_lenses`, `max_cost`, `pricing` overrides, `lens_models`, `incremental`, `retry_truncated`, `include_synthesis`, `include_triage`, `wait_seconds`); an explicit tool parameter always wins. `lens_models` routes individual lenses to their own batch model — a stronger model changes security and defect findings far more than it changes an architecture map — and each override is priced, capability-checked, and clamped exactly like the default, with the estimate broken out per lens so a mixed-model run cannot be approved without seeing which lens costs what. CodeCartographer ships no stronger default: which model earns its price depends on your repository and budget, so compare candidates with the `models` action and choose — for one run with the `model` and `lens_models` parameters (Pi: `--model=ID`, `--lens-model=LENS:ID`), or for the repository in `config.yaml`. The `models` listing is advisory: OpenRouter's catalog returns a `:batch` id for some models its Batch API then refuses (`does not have a :batch endpoint`), at no cost, and nothing in the catalog tells them apart — so the listing tags the ids this repository's own submits have seen accepted or refused, and a refused lens says why in the submit report. `codecarto_skill {cwd, name: "broadside"}` returns the reading guide for a completed run, and unlike post-pipeline skills it is not gated on a finished pipeline.
402
405
 
403
- On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID]`, with tab-completion for actions, lens names, and flags and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
406
+ On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models|verify] [lenses…] [--model=ID] [--lens-model=LENS:ID] [--top=N]`, with tab-completion for actions, lens names, and flags and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
404
407
 
405
408
  Broad-Side needs runtime code, so firing a run is an executable-surface feature: Pi and MCP have it, the pure drop-in template does not (it carries only the reading guide).
406
409
 
@@ -95,7 +95,11 @@ Two more economies worth knowing:
95
95
 
96
96
  ## Reading a run
97
97
 
98
- Results land in `.codecarto/broadside/<run>/`. Read them in this order:
98
+ Results land in `.codecarto/broadside/<run>/`. If `verified.md` is there, read
99
+ it before anything else: `action: "verify"` has read the top defect and
100
+ security findings against the source with read-only tools and given each a
101
+ verdict (confirmed with its trigger, not-a-defect, discarded with the guard
102
+ that shows it, unclear). Then, in this order:
99
103
 
100
104
  1. `synthesis.md` — executive summary, severity counts, top cross-lens
101
105
  findings, per-module risk.
@@ -0,0 +1,95 @@
1
+ import { type BroadsideLensId, type BroadsideVerifyEntry, type FetchLike, type StoredLensResult } from "./broadside.ts";
2
+ export declare const BROADSIDE_CHAT_URL = "https://openrouter.ai/api/v1/chat/completions";
3
+ /** How many findings `verify` reads by default, most severe first. */
4
+ export declare const BROADSIDE_VERIFY_DEFAULT_TOP = 10;
5
+ /** Tool calls one finding may spend before it must answer. */
6
+ export declare const BROADSIDE_VERIFY_MAX_TOOL_CALLS = 8;
7
+ /** The lenses whose findings carry a file:line and a claim to check. */
8
+ export declare const BROADSIDE_VERIFIABLE_LENSES: readonly BroadsideLensId[];
9
+ export type VerifyVerdict = "confirmed" | "not-a-defect" | "discarded" | "unclear";
10
+ export type VerifiedFinding = {
11
+ index: number;
12
+ lensId: BroadsideLensId;
13
+ customId: string;
14
+ severity: string;
15
+ title: string;
16
+ location: string;
17
+ verdict: VerifyVerdict | "error";
18
+ confidence: string;
19
+ evidence: Array<{
20
+ file: string;
21
+ lines: string;
22
+ note: string;
23
+ }>;
24
+ reasoning: string;
25
+ toolCalls: number;
26
+ cost: number;
27
+ };
28
+ export type BroadsideVerifyResult = {
29
+ runId: string;
30
+ outputDir: string;
31
+ model: string;
32
+ status: BroadsideVerifyEntry["status"];
33
+ /** How many findings the run had in the verifiable lenses. */
34
+ candidates: number;
35
+ findings: VerifiedFinding[];
36
+ totalCost: number;
37
+ /** Set when the cost cap stopped the pass before every selected finding was read. */
38
+ stoppedByCost?: boolean;
39
+ };
40
+ type CandidateFinding = {
41
+ lensId: BroadsideLensId;
42
+ customId: string;
43
+ severity: string;
44
+ title: string;
45
+ location: string;
46
+ description: string;
47
+ pattern: string;
48
+ };
49
+ /** The sync-priced id behind a `:batch` model id (`vendor/name:batch` → `vendor/name`). */
50
+ export declare function syncModelFor(batchModel: string): string;
51
+ /** The findings a run's saved lens results carry, most severe first. */
52
+ export declare function rankVerifiableFindings(stored: StoredLensResult[]): CandidateFinding[];
53
+ export type RepoReader = {
54
+ readFile(path: string, startLine?: number, endLine?: number): Promise<string>;
55
+ grep(pattern: string, pathPrefix?: string): Promise<string>;
56
+ listDir(path: string): Promise<string>;
57
+ };
58
+ /**
59
+ * Three read-only tools over the repository's own file list — the same
60
+ * listing the lenses scan (tracked and untracked, ignore rules applied) minus
61
+ * everything {@link isSlurpable} keeps out of a lens: credential stores, build
62
+ * output, binaries. A path outside the repository, or one the listing does
63
+ * not contain, is an error the model sees, not a read.
64
+ */
65
+ export declare function createRepoReader(cwd: string): Promise<RepoReader>;
66
+ /**
67
+ * The rubric. The order is deliberate — it is the order a reviewer settles a
68
+ * claim in — and `not-a-defect` is the verdict that separates "the code does
69
+ * what the claim says" from "and that is a bug": without it, a cast every
70
+ * caller satisfies gets confirmed because it is literally there.
71
+ */
72
+ export declare const BROADSIDE_VERIFY_SYSTEM_PROMPT: string;
73
+ /** Verify one finding: up to the tool budget, then a verdict. */
74
+ export declare function verifyFinding(finding: CandidateFinding, index: number, reader: RepoReader, apiKey: string, model: string, fetcher: FetchLike): Promise<Omit<VerifiedFinding, "index" | "lensId" | "customId" | "severity" | "title" | "location">>;
75
+ /**
76
+ * Verify the top findings of a collected run against the repository.
77
+ *
78
+ * `maxCost` is a running cap, not a pre-flight estimate: a sync call's cost
79
+ * is only known when it returns, so the pass stops *before* starting the next
80
+ * finding once the cap is reached and reports `partial`. On the default
81
+ * model a finding costs about a cent.
82
+ */
83
+ export declare function runBroadsideVerify(cwd: string, apiKey: string, opts?: {
84
+ runId?: string;
85
+ top?: number;
86
+ model?: string;
87
+ /** USD; 0 means no limit. */
88
+ maxCost?: number;
89
+ fetcher?: FetchLike;
90
+ signal?: AbortSignal;
91
+ onProgress?: (finding: VerifiedFinding) => void;
92
+ }): Promise<BroadsideVerifyResult>;
93
+ export declare function renderVerifiedMarkdown(result: BroadsideVerifyResult): string;
94
+ export declare function verifyResultText(result: BroadsideVerifyResult): string;
95
+ export {};
@@ -0,0 +1,433 @@
1
+ // Broad-Side verification pass (#143): the hybrid the roadmap kept coming
2
+ // back to. The batch sweep is cheap and reads files without being able to
3
+ // look anything up, and its measured weakness is precision, not coverage —
4
+ // on this repository the top twelve findings by severity were two real
5
+ // defects and ten claims that a look at the guard, the caller, or the
6
+ // tsconfig would have dismissed. So after collect, one sync-priced call per
7
+ // finding, with three read-only tools confined to the repository, reads the
8
+ // cited code and says whether the failure is reachable.
9
+ //
10
+ // Measured on 2026-09-13 (the #143 comparison run): twelve findings, the
11
+ // rubric below, `google/gemini-3.7-flash` at low reasoning effort —
12
+ // twelve-for-twelve agreement with a reviewer's ground truth, precision of
13
+ // the confirmed set from 17% to 100%, both true findings kept, $0.13 in all
14
+ // (about a cent a finding). The rubric mattered more than the model: a first
15
+ // draft without the `not-a-defect` verdict "confirmed" two type casts that
16
+ // every caller satisfies, because they were literally true of the code.
17
+ //
18
+ // Read-only on purpose. A tool-using pass that could mutate would need the
19
+ // headless-agent retry rule (retry only before the first tool call); one
20
+ // that only reads keeps the batch property that re-running is always safe.
21
+ import { readdir, readFile, stat, writeFile } from "node:fs/promises";
22
+ import { isAbsolute, join, relative, resolve } from "node:path";
23
+ import { BROADSIDE_DIR, BROADSIDE_LENS_IDS, broadsideDirFor, defaultReasoningFor, isSlurpable, listRepoFiles, loadBroadsideState, loadSavedLensResults, parseLensJson, persistBroadsideRunMerging, } from "./broadside.js";
24
+ export const BROADSIDE_CHAT_URL = "https://openrouter.ai/api/v1/chat/completions";
25
+ /** How many findings `verify` reads by default, most severe first. */
26
+ export const BROADSIDE_VERIFY_DEFAULT_TOP = 10;
27
+ /** Tool calls one finding may spend before it must answer. */
28
+ export const BROADSIDE_VERIFY_MAX_TOOL_CALLS = 8;
29
+ /** The lenses whose findings carry a file:line and a claim to check. */
30
+ export const BROADSIDE_VERIFIABLE_LENSES = ["defect", "security"];
31
+ const READ_FILE_MAX_LINES = 200;
32
+ const GREP_MAX_MATCHES = 40;
33
+ const GREP_MAX_FILE_BYTES = 1_000_000;
34
+ const TOOL_OUTPUT_MAX_CHARS = 12_000;
35
+ const SEVERITY_RANK = { critical: 0, high: 1, medium: 2, low: 3 };
36
+ /** The sync-priced id behind a `:batch` model id (`vendor/name:batch` → `vendor/name`). */
37
+ export function syncModelFor(batchModel) {
38
+ return batchModel.replace(/:batch$/, "");
39
+ }
40
+ /** The findings a run's saved lens results carry, most severe first. */
41
+ export function rankVerifiableFindings(stored) {
42
+ const out = [];
43
+ for (const result of stored) {
44
+ if (!BROADSIDE_VERIFIABLE_LENSES.includes(result.lensId) || result.truncated)
45
+ continue;
46
+ const parsed = parseLensJson(result.content);
47
+ for (const finding of parsed?.findings ?? []) {
48
+ const location = typeof finding.location === "string" ? finding.location : "";
49
+ const title = typeof finding.title === "string" ? finding.title : "";
50
+ if (!location || !title)
51
+ continue;
52
+ out.push({
53
+ lensId: result.lensId,
54
+ customId: result.customId,
55
+ severity: typeof finding.severity === "string" ? finding.severity.toLowerCase() : "unknown",
56
+ title,
57
+ location,
58
+ description: typeof finding.description === "string" ? finding.description : "",
59
+ pattern: typeof finding.pattern === "string" ? finding.pattern : typeof finding.category === "string" ? finding.category : "",
60
+ });
61
+ }
62
+ }
63
+ // Stable: severity, then lens order, then the order the lens listed them.
64
+ return out
65
+ .map((finding, order) => ({ finding, order }))
66
+ .sort((a, b) => (SEVERITY_RANK[a.finding.severity] ?? 9) - (SEVERITY_RANK[b.finding.severity] ?? 9)
67
+ || BROADSIDE_LENS_IDS.indexOf(a.finding.lensId) - BROADSIDE_LENS_IDS.indexOf(b.finding.lensId)
68
+ || a.order - b.order)
69
+ .map(({ finding }) => finding);
70
+ }
71
+ /**
72
+ * Three read-only tools over the repository's own file list — the same
73
+ * listing the lenses scan (tracked and untracked, ignore rules applied) minus
74
+ * everything {@link isSlurpable} keeps out of a lens: credential stores, build
75
+ * output, binaries. A path outside the repository, or one the listing does
76
+ * not contain, is an error the model sees, not a read.
77
+ */
78
+ export async function createRepoReader(cwd) {
79
+ const { files } = await listRepoFiles(cwd);
80
+ const readable = new Set(files.filter(isSlurpable));
81
+ const confine = (path) => {
82
+ const abs = resolve(cwd, path);
83
+ const rel = relative(cwd, abs).split("\\").join("/");
84
+ if (!rel || rel.startsWith("..") || isAbsolute(rel))
85
+ throw new Error(`path is outside the repository: ${path}`);
86
+ return rel;
87
+ };
88
+ return {
89
+ async readFile(path, startLine, endLine) {
90
+ const rel = confine(path);
91
+ if (!readable.has(rel))
92
+ return `error: ${rel} is not a readable source file of this repository`;
93
+ const lines = (await readFile(join(cwd, rel), "utf8")).split("\n");
94
+ const start = Math.max(1, Math.floor(Number(startLine ?? 1)) || 1);
95
+ const requestedEnd = Math.floor(Number(endLine ?? start + READ_FILE_MAX_LINES - 1)) || start;
96
+ const end = Math.min(lines.length, requestedEnd, start + READ_FILE_MAX_LINES - 1);
97
+ if (start > lines.length)
98
+ return `error: ${rel} has ${lines.length} lines`;
99
+ return lines.slice(start - 1, end).map((line, i) => `${start + i}: ${line}`).join("\n");
100
+ },
101
+ async grep(pattern, pathPrefix) {
102
+ let regex;
103
+ try {
104
+ regex = new RegExp(pattern);
105
+ }
106
+ catch (error) {
107
+ return `error: invalid pattern (${error instanceof Error ? error.message : String(error)})`;
108
+ }
109
+ const prefix = pathPrefix ? confine(pathPrefix) : "";
110
+ const matches = [];
111
+ for (const rel of readable) {
112
+ if (prefix && rel !== prefix && !rel.startsWith(`${prefix}/`) && !rel.startsWith(prefix))
113
+ continue;
114
+ let text;
115
+ try {
116
+ if ((await stat(join(cwd, rel))).size > GREP_MAX_FILE_BYTES)
117
+ continue;
118
+ text = await readFile(join(cwd, rel), "utf8");
119
+ }
120
+ catch {
121
+ continue;
122
+ }
123
+ const lines = text.split("\n");
124
+ for (let i = 0; i < lines.length && matches.length < GREP_MAX_MATCHES; i++) {
125
+ if (regex.test(lines[i]))
126
+ matches.push(`${rel}:${i + 1}: ${lines[i]}`);
127
+ }
128
+ if (matches.length >= GREP_MAX_MATCHES)
129
+ break;
130
+ }
131
+ return matches.length > 0 ? matches.join("\n") : "(no matches)";
132
+ },
133
+ async listDir(path) {
134
+ const rel = path === "." || path === "" ? "" : confine(path);
135
+ try {
136
+ const entries = await readdir(join(cwd, rel), { withFileTypes: true });
137
+ return entries
138
+ .filter((entry) => {
139
+ const child = rel ? `${rel}/${entry.name}` : entry.name;
140
+ return entry.isDirectory() ? [...readable].some((f) => f.startsWith(`${child}/`)) : readable.has(child);
141
+ })
142
+ .map((entry) => (entry.isDirectory() ? `${entry.name}/` : entry.name))
143
+ .sort()
144
+ .join("\n") || "(empty)";
145
+ }
146
+ catch (error) {
147
+ return `error: ${error instanceof Error ? error.message : String(error)}`;
148
+ }
149
+ },
150
+ };
151
+ }
152
+ const TOOLS = [
153
+ { type: "function", function: { name: "read_file", description: "Read a range of lines (1-based, inclusive) from a source file of the repository; at most 200 lines per call.", parameters: { type: "object", properties: { path: { type: "string" }, start_line: { type: "integer" }, end_line: { type: "integer" } }, required: ["path"] } } },
154
+ { type: "function", function: { name: "grep", description: "Search the repository's source files for a regular expression; returns at most 40 matching lines as path:line: text.", parameters: { type: "object", properties: { pattern: { type: "string" }, path_prefix: { type: "string", description: "optional directory or file to search under" } }, required: ["pattern"] } } },
155
+ { type: "function", function: { name: "list_dir", description: "List a directory of the repository.", parameters: { type: "object", properties: { path: { type: "string" } }, required: ["path"] } } },
156
+ ];
157
+ const VERDICT_SCHEMA = {
158
+ name: "broadside_verification",
159
+ strict: true,
160
+ schema: {
161
+ type: "object",
162
+ properties: {
163
+ verdict: { type: "string", enum: ["confirmed", "not-a-defect", "discarded", "unclear"] },
164
+ confidence: { type: "string", enum: ["high", "medium", "low"] },
165
+ evidence: {
166
+ type: "array",
167
+ items: { type: "object", properties: { file: { type: "string" }, lines: { type: "string" }, note: { type: "string" } }, required: ["file", "lines", "note"], additionalProperties: false },
168
+ },
169
+ reasoning: { type: "string" },
170
+ },
171
+ required: ["verdict", "confidence", "evidence", "reasoning"],
172
+ additionalProperties: false,
173
+ },
174
+ };
175
+ /**
176
+ * The rubric. The order is deliberate — it is the order a reviewer settles a
177
+ * claim in — and `not-a-defect` is the verdict that separates "the code does
178
+ * what the claim says" from "and that is a bug": without it, a cast every
179
+ * caller satisfies gets confirmed because it is literally there.
180
+ */
181
+ export const BROADSIDE_VERIFY_SYSTEM_PROMPT = "You verify a scouting finding produced by a one-shot batch scan against the real source code. " +
182
+ "The scan saw files without being able to look anything up; you can. Use the tools to read the cited location " +
183
+ "and whatever else the claim depends on (callers, the definition of a helper, error handling around it). " +
184
+ "Then decide, in this order: " +
185
+ "'discarded' — the claim is wrong about the code (the guard exists, the value cannot be what the claim assumes, " +
186
+ "the cited line does something else, the condition it fears is ruled out by the project's config or runtime); " +
187
+ "'not-a-defect' — the claim is literally true of the code but no caller, input, or state can reach the failure it " +
188
+ "describes: a TypeScript cast every caller satisfies, a hypothetical about an environment the project does not " +
189
+ "target, a style or type-hygiene observation; " +
190
+ "'confirmed' — the failure is reachable: name the concrete input, call site, or sequence that triggers it, and what " +
191
+ "then goes wrong; " +
192
+ "'unclear' — settling it needs runtime behaviour or specification knowledge the code does not contain. " +
193
+ "Be strict: a real but different problem than the one claimed is 'discarded' with the difference noted, and " +
194
+ "'confirmed' without a trigger you found in the code is not allowed. " +
195
+ "Cite line ranges you actually read. When you are done, reply with only the JSON verdict object.";
196
+ function findingPrompt(index, finding) {
197
+ return (`Finding ${index} (${finding.lensId} lens, severity ${finding.severity}):\n` +
198
+ `Title: ${finding.title}\nLocation: ${finding.location}\nPattern/category: ${finding.pattern || "-"}\n` +
199
+ `Description: ${finding.description}\n\n` +
200
+ "Verify it. Reply with a JSON object {verdict, confidence, evidence:[{file,lines,note}], reasoning}.");
201
+ }
202
+ async function chat(fetcher, apiKey, model, body) {
203
+ const resp = await fetcher(BROADSIDE_CHAT_URL, {
204
+ method: "POST",
205
+ headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
206
+ body: JSON.stringify({ model, reasoning: defaultReasoningFor(), usage: { include: true }, ...body }),
207
+ signal: AbortSignal.timeout(120_000),
208
+ });
209
+ const data = (await resp.json());
210
+ if (!resp.ok || data.error) {
211
+ const detail = data.error?.message ?? JSON.stringify(data.error ?? data).slice(0, 300);
212
+ throw new Error(`OpenRouter chat: HTTP ${resp.status}: ${detail}`);
213
+ }
214
+ return data;
215
+ }
216
+ function parseVerdict(text) {
217
+ const trimmed = text.trim();
218
+ const fenced = /```(?:json)?\s*([\s\S]*?)```/.exec(trimmed);
219
+ try {
220
+ const parsed = JSON.parse(fenced ? fenced[1] : trimmed);
221
+ return typeof parsed?.verdict === "string" ? parsed : null;
222
+ }
223
+ catch {
224
+ return null;
225
+ }
226
+ }
227
+ /** Verify one finding: up to the tool budget, then a verdict. */
228
+ export async function verifyFinding(finding, index, reader, apiKey, model, fetcher) {
229
+ const messages = [
230
+ { role: "system", content: BROADSIDE_VERIFY_SYSTEM_PROMPT },
231
+ { role: "user", content: findingPrompt(index, finding) },
232
+ ];
233
+ let toolCalls = 0;
234
+ let cost = 0;
235
+ const usageCost = (data) => {
236
+ const usage = data.usage;
237
+ return typeof usage?.cost === "number" ? usage.cost : 0;
238
+ };
239
+ try {
240
+ // The +2 leaves room for the final schema-forced call after the budget.
241
+ for (let step = 0; step < BROADSIDE_VERIFY_MAX_TOOL_CALLS + 2; step++) {
242
+ const budgetSpent = toolCalls >= BROADSIDE_VERIFY_MAX_TOOL_CALLS;
243
+ if (budgetSpent)
244
+ messages.push({ role: "user", content: "Tool budget spent. Emit the verdict JSON object now from what you have read." });
245
+ const data = await chat(fetcher, apiKey, model, budgetSpent
246
+ ? { messages, response_format: { type: "json_schema", json_schema: VERDICT_SCHEMA }, max_tokens: 2000 }
247
+ : { messages, tools: TOOLS, tool_choice: "auto", max_tokens: 4000 });
248
+ cost += usageCost(data);
249
+ const choice = data.choices?.[0];
250
+ const message = choice?.message ?? { role: "assistant", content: "" };
251
+ messages.push(message);
252
+ const calls = message.tool_calls;
253
+ if (calls && calls.length > 0 && !budgetSpent) {
254
+ for (const call of calls) {
255
+ toolCalls += 1;
256
+ let args = {};
257
+ try {
258
+ args = JSON.parse(call.function.arguments || "{}");
259
+ }
260
+ catch {
261
+ // Malformed arguments: the model sees the error and can retry.
262
+ }
263
+ let output;
264
+ try {
265
+ output = call.function.name === "read_file"
266
+ ? await reader.readFile(String(args.path ?? ""), args.start_line, args.end_line)
267
+ : call.function.name === "grep"
268
+ ? await reader.grep(String(args.pattern ?? ""), typeof args.path_prefix === "string" ? args.path_prefix : undefined)
269
+ : call.function.name === "list_dir"
270
+ ? await reader.listDir(String(args.path ?? "."))
271
+ : `error: unknown tool ${call.function.name}`;
272
+ }
273
+ catch (error) {
274
+ output = `error: ${error instanceof Error ? error.message : String(error)}`;
275
+ }
276
+ messages.push({ role: "tool", tool_call_id: call.id, content: output.slice(0, TOOL_OUTPUT_MAX_CHARS) });
277
+ }
278
+ continue;
279
+ }
280
+ const verdict = parseVerdict(typeof message.content === "string" ? message.content : "");
281
+ if (verdict)
282
+ return finish(verdict, toolCalls, cost);
283
+ if (budgetSpent)
284
+ break;
285
+ // Prose instead of JSON: one schema-forced call, no tools.
286
+ const final = await chat(fetcher, apiKey, model, {
287
+ messages: [...messages, { role: "user", content: "Emit the verdict JSON object now." }],
288
+ response_format: { type: "json_schema", json_schema: VERDICT_SCHEMA },
289
+ max_tokens: 2000,
290
+ });
291
+ cost += usageCost(final);
292
+ const forced = parseVerdict(String((final.choices?.[0]?.message?.content) ?? ""));
293
+ if (forced)
294
+ return finish(forced, toolCalls, cost);
295
+ break;
296
+ }
297
+ return { verdict: "error", confidence: "low", evidence: [], reasoning: "no verdict within the tool budget", toolCalls, cost };
298
+ }
299
+ catch (error) {
300
+ return { verdict: "error", confidence: "low", evidence: [], reasoning: error instanceof Error ? error.message : String(error), toolCalls, cost };
301
+ }
302
+ }
303
+ function finish(verdict, toolCalls, cost) {
304
+ const known = ["confirmed", "not-a-defect", "discarded", "unclear"];
305
+ const value = String(verdict.verdict);
306
+ const evidence = Array.isArray(verdict.evidence)
307
+ ? verdict.evidence.map((e) => ({ file: String(e.file ?? ""), lines: String(e.lines ?? ""), note: String(e.note ?? "") }))
308
+ : [];
309
+ return {
310
+ verdict: (known.includes(value) ? value : "unclear"),
311
+ confidence: typeof verdict.confidence === "string" ? verdict.confidence : "low",
312
+ evidence,
313
+ reasoning: typeof verdict.reasoning === "string" ? verdict.reasoning : "",
314
+ toolCalls,
315
+ cost,
316
+ };
317
+ }
318
+ /**
319
+ * Verify the top findings of a collected run against the repository.
320
+ *
321
+ * `maxCost` is a running cap, not a pre-flight estimate: a sync call's cost
322
+ * is only known when it returns, so the pass stops *before* starting the next
323
+ * finding once the cap is reached and reports `partial`. On the default
324
+ * model a finding costs about a cent.
325
+ */
326
+ export async function runBroadsideVerify(cwd, apiKey, opts = {}) {
327
+ const broadsideDir = broadsideDirFor(cwd);
328
+ const state = await loadBroadsideState(broadsideDir);
329
+ const run = opts.runId ? state.runs.find((candidate) => candidate.id === opts.runId) : state.runs[state.runs.length - 1];
330
+ if (!run) {
331
+ throw new Error(opts.runId ? `No Broad-Side run with id ${opts.runId}.` : "No Broad-Side run recorded. Submit and collect one first.");
332
+ }
333
+ const runDir = join(broadsideDir, run.outputDir);
334
+ const stored = await loadSavedLensResults(runDir, run.lenses);
335
+ const ranked = rankVerifiableFindings(stored);
336
+ if (ranked.length === 0) {
337
+ throw new Error(`Run ${run.id} has no verifiable findings on disk: the defect and security lenses either did not run, are not collected yet, or found nothing. Collect the run first.`);
338
+ }
339
+ const top = Math.max(1, Math.floor(opts.top ?? BROADSIDE_VERIFY_DEFAULT_TOP));
340
+ // The run's model is a `:batch` variant; the sync endpoint wants the base id.
341
+ const model = opts.model ?? syncModelFor(run.model);
342
+ const maxCost = opts.maxCost ?? 0;
343
+ const fetcher = opts.fetcher ?? fetch;
344
+ const reader = await createRepoReader(cwd);
345
+ const selected = ranked.slice(0, top);
346
+ const findings = [];
347
+ let totalCost = 0;
348
+ let stoppedByCost = false;
349
+ for (const [i, candidate] of selected.entries()) {
350
+ if (opts.signal?.aborted)
351
+ break;
352
+ if (maxCost > 0 && totalCost >= maxCost) {
353
+ stoppedByCost = true;
354
+ break;
355
+ }
356
+ const outcome = await verifyFinding(candidate, i + 1, reader, apiKey, model, fetcher);
357
+ const finding = {
358
+ index: i + 1,
359
+ lensId: candidate.lensId,
360
+ customId: candidate.customId,
361
+ severity: candidate.severity,
362
+ title: candidate.title,
363
+ location: candidate.location,
364
+ ...outcome,
365
+ };
366
+ findings.push(finding);
367
+ totalCost += outcome.cost;
368
+ opts.onProgress?.(finding);
369
+ }
370
+ const status = findings.length === selected.length ? "completed" : "partial";
371
+ const entry = {
372
+ status,
373
+ model,
374
+ top,
375
+ verified: findings.length,
376
+ confirmed: findings.filter((f) => f.verdict === "confirmed").length,
377
+ cost: totalCost,
378
+ at: new Date().toISOString(),
379
+ };
380
+ const result = {
381
+ runId: run.id,
382
+ outputDir: join(".codecarto", BROADSIDE_DIR, run.id),
383
+ model,
384
+ status,
385
+ candidates: ranked.length,
386
+ findings,
387
+ totalCost,
388
+ ...(stoppedByCost && { stoppedByCost: true }),
389
+ };
390
+ await writeFile(join(runDir, "verified.json"), `${JSON.stringify({ ...entry, run_id: run.id, candidates: ranked.length, findings }, null, "\t")}\n`, "utf8");
391
+ await writeFile(join(runDir, "verified.md"), renderVerifiedMarkdown(result), "utf8");
392
+ run.verify = entry;
393
+ await persistBroadsideRunMerging(broadsideDir, run);
394
+ return result;
395
+ }
396
+ const VERDICT_MARK = { confirmed: "✓", "not-a-defect": "–", discarded: "✗", unclear: "?", error: "!" };
397
+ export function renderVerifiedMarkdown(result) {
398
+ const lines = [
399
+ `# Verified findings — run ${result.runId}`,
400
+ "",
401
+ `${result.findings.length} of ${result.candidates} verifiable finding(s) read against the source on \`${result.model}\` (most severe first), $${result.totalCost.toFixed(4)}.` +
402
+ (result.stoppedByCost ? " Stopped by the cost cap before the rest." : ""),
403
+ "",
404
+ "A **confirmed** finding names the input, call site, or sequence that reaches the failure. **not-a-defect** means the claim is",
405
+ "literally true of the code but nothing can reach the failure it describes; **discarded** means the claim is wrong about the",
406
+ "code; **unclear** needs runtime or specification knowledge. Every verdict is still a model's reading — a confirmed finding",
407
+ "is a lead worth a human's next look, not a validated claim.",
408
+ "",
409
+ ];
410
+ for (const f of result.findings) {
411
+ lines.push(`## ${VERDICT_MARK[f.verdict] ?? "?"} ${f.index}. [${f.severity}] ${f.title}`, "", `- **verdict**: ${f.verdict} (${f.confidence})`, `- **location**: ${f.location}`, `- **lens**: ${f.lensId} (${f.customId})`);
412
+ if (f.evidence.length > 0)
413
+ lines.push(`- **evidence**: ${f.evidence.map((e) => `${e.file}:${e.lines} — ${e.note}`).join("; ")}`);
414
+ lines.push(`- **reasoning**: ${f.reasoning.replace(/\s+/g, " ").trim()}`, `- **cost**: $${f.cost.toFixed(4)} (${f.toolCalls} tool call(s))`, "");
415
+ }
416
+ return lines.join("\n");
417
+ }
418
+ export function verifyResultText(result) {
419
+ const counts = { confirmed: 0, "not-a-defect": 0, discarded: 0, unclear: 0, error: 0 };
420
+ for (const f of result.findings)
421
+ counts[f.verdict] = (counts[f.verdict] ?? 0) + 1;
422
+ const lines = [
423
+ `Broad-Side verify — run ${result.runId}: ${result.status}`,
424
+ ` ${result.findings.length} of ${result.candidates} verifiable finding(s) read on ${result.model} | cost: $${result.totalCost.toFixed(4)}` +
425
+ (result.stoppedByCost ? " (stopped by the cost cap)" : ""),
426
+ ` confirmed ${counts.confirmed} · not-a-defect ${counts["not-a-defect"]} · discarded ${counts.discarded} · unclear ${counts.unclear}` + (counts.error ? ` · error ${counts.error}` : ""),
427
+ ];
428
+ for (const f of result.findings) {
429
+ lines.push(` ${VERDICT_MARK[f.verdict] ?? "?"} [${f.severity}] ${f.title} @ ${f.location} — ${f.verdict}`);
430
+ }
431
+ lines.push(`Details in ${result.outputDir}/verified.md. A confirmed finding is a lead for a human's next look, not a validated claim.`);
432
+ return lines.join("\n");
433
+ }
@@ -244,6 +244,17 @@ export type BroadsideTriageEntry = {
244
244
  cost?: number;
245
245
  error?: string;
246
246
  };
247
+ /** Recorded on the run once a verification pass has run (#143); see core/broadside-verify.ts. */
248
+ export type BroadsideVerifyEntry = {
249
+ /** `completed`: every selected finding got a verdict; `partial`: the cost cap or an abort stopped it early. */
250
+ status: "completed" | "partial";
251
+ model: string;
252
+ top: number;
253
+ verified: number;
254
+ confirmed: number;
255
+ cost: number;
256
+ at: string;
257
+ };
247
258
  /** The truncation retry pass of one run: one batch per model (#206). */
248
259
  export type BroadsideRetryEntry = {
249
260
  status: "submitted" | "completed" | "failed";
@@ -275,6 +286,8 @@ export type BroadsideRun = {
275
286
  * run cannot both submit it (#322). Absent until a collect claims it.
276
287
  */
277
288
  retry?: BroadsideRetryEntry;
289
+ /** The verification pass over the top findings, when one has run (#143). */
290
+ verify?: BroadsideVerifyEntry;
278
291
  totalCost?: number;
279
292
  pricing?: ModelPricing;
280
293
  maxCost?: number;
@@ -520,9 +533,22 @@ export declare function getLens(lensId: BroadsideLensId): LensDefinition;
520
533
  export declare function listLenses(): LensDefinition[];
521
534
  /** The languages Broad-Side can scan; anything else is refused at submit. */
522
535
  export declare const BROADSIDE_LANGUAGES: readonly ["go", "python", "rust", "typescript", "javascript"];
536
+ /**
537
+ * The files a run scans, and where they came from. Contents are always read
538
+ * from the working tree, so the list is the working tree's too: tracked files
539
+ * plus untracked ones git does not ignore, minus files deleted on disk. The
540
+ * list used to come from `git ls-tree HEAD`, so a run mixed the committed
541
+ * file list with uncommitted contents and never saw an untracked file (#248).
542
+ * A target that is not a git repository gets a bounded walk.
543
+ */
544
+ export declare function listRepoFiles(targetDir: string): Promise<{
545
+ files: string[];
546
+ snapshot: RepoSnapshotSource;
547
+ }>;
523
548
  export declare function collectRepoInfo(targetDir: string, opts?: {
524
549
  redact?: boolean;
525
550
  }): Promise<RepoInfo>;
551
+ export declare function isSlurpable(relPath: string): boolean;
526
552
  type CollectedFile = {
527
553
  relPath: string;
528
554
  moduleName: string;
@@ -920,7 +920,7 @@ const SOURCE_SPECS = {
920
920
  * file list with uncommitted contents and never saw an untracked file (#248).
921
921
  * A target that is not a git repository gets a bounded walk.
922
922
  */
923
- async function listRepoFiles(targetDir) {
923
+ export async function listRepoFiles(targetDir) {
924
924
  try {
925
925
  const listed = await execFileAsync("git", ["-C", targetDir, "ls-files", "-z", "--cached", "--others", "--exclude-standard"], { maxBuffer: 64 * 1024 * 1024, timeout: GIT_TIMEOUT_MS });
926
926
  const deleted = await execFileAsync("git", ["-C", targetDir, "ls-files", "-z", "--deleted"], {
@@ -1200,7 +1200,7 @@ function matchesAnyGlob(path, globs) {
1200
1200
  }
1201
1201
  return false;
1202
1202
  }
1203
- function isSlurpable(relPath) {
1203
+ export function isSlurpable(relPath) {
1204
1204
  // A credential store is never a lens input, whatever its globs say (#252).
1205
1205
  if (isSecretFile(relPath))
1206
1206
  return false;
@@ -1624,6 +1624,10 @@ export async function persistBroadsideRunMerging(broadsideDir, run) {
1624
1624
  run.triage = onDisk.triage;
1625
1625
  if (retryEntryRank(onDisk.retry) > retryEntryRank(run.retry))
1626
1626
  run.retry = onDisk.retry;
1627
+ // A verification pass another process recorded is never dropped by
1628
+ // a collect that never knew about it; a newer pass replaces an older.
1629
+ if (onDisk.verify && (!run.verify || onDisk.verify.at > run.verify.at))
1630
+ run.verify = onDisk.verify;
1627
1631
  for (const [lensId, theirs] of Object.entries(onDisk.batches)) {
1628
1632
  if (theirs && batchEntryRank(theirs) > batchEntryRank(run.batches[lensId]))
1629
1633
  run.batches[lensId] = theirs;
@@ -3545,6 +3549,9 @@ export function statusText(state) {
3545
3549
  }
3546
3550
  lines.push(` synthesis: ${run.synthesis.status}`);
3547
3551
  lines.push(` triage: ${run.triage?.status ?? "pending"}`);
3552
+ if (run.verify) {
3553
+ lines.push(` verify: ${run.verify.status} — ${run.verify.confirmed} confirmed of ${run.verify.verified} read on ${run.verify.model}, $${run.verify.cost.toFixed(4)}`);
3554
+ }
3548
3555
  if (run.totalCost !== undefined)
3549
3556
  lines.push(` total cost: $${run.totalCost.toFixed(6)}`);
3550
3557
  }
@@ -16,5 +16,6 @@ export * from "./dashboard.ts";
16
16
  export * from "./library.ts";
17
17
  export * from "./synthesis.ts";
18
18
  export * from "./broadside.ts";
19
+ export * from "./broadside-verify.ts";
19
20
  export * from "./secrets.ts";
20
21
  export * from "./dashboard-writer.ts";
@@ -19,5 +19,6 @@ export * from "./dashboard.js";
19
19
  export * from "./library.js";
20
20
  export * from "./synthesis.js";
21
21
  export * from "./broadside.js";
22
+ export * from "./broadside-verify.js";
22
23
  export * from "./secrets.js";
23
24
  export * from "./dashboard-writer.js";
@@ -1,5 +1,5 @@
1
1
  import { type BroadsideLensId } from "../../core/index.ts";
2
- export type BroadsideAction = "submit" | "collect" | "status" | "models";
2
+ export type BroadsideAction = "submit" | "collect" | "status" | "models" | "verify";
3
3
  export interface BroadsideFlags {
4
4
  action: BroadsideAction;
5
5
  /** Empty means "the repository's default lens set". */
@@ -22,11 +22,13 @@ export interface BroadsideFlags {
22
22
  model?: string;
23
23
  /** For submit: per-lens model overrides, layered over config.yaml's (#141). */
24
24
  lensModels?: Partial<Record<BroadsideLensId, string>>;
25
+ /** For verify: how many findings to read (#143). */
26
+ top?: number;
25
27
  benchmarks: boolean;
26
28
  unknown: string[];
27
29
  /** Set on an invalid combination. The caller surfaces it as an error. */
28
30
  error?: string;
29
31
  }
30
32
  /** Every token the completer offers, in the order it offers them. */
31
- export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--model=", "--lens-model=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
33
+ export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "verify", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--model=", "--lens-model=", "--top=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
32
34
  export declare function parseBroadsideFlags(args: string): BroadsideFlags;
@@ -6,6 +6,7 @@
6
6
  // /codecarto-broadside collect --wait=900
7
7
  // /codecarto-broadside status
8
8
  // /codecarto-broadside models --benchmarks
9
+ // /codecarto-broadside verify --top=10 → read the top findings against the source
9
10
  //
10
11
  // Flags mirror the codecarto_broadside tool parameters, with the negative
11
12
  // forms spelled out because a slash command has no place to pass `false`:
@@ -14,8 +15,9 @@
14
15
  // --max-cost=N --no-retry-truncated
15
16
  // --wait=SECONDS --benchmarks (models only)
16
17
  // --run=ID (collect only: an older run, as listed by status)
17
- // --model=ID (submit only: the run's batch model, as listed by models)
18
+ // --model=ID (submit: the run's batch model, as listed by models; verify: the sync model to read with)
18
19
  // --lens-model=LENS:ID (submit only, repeatable: one lens on its own model)
20
+ // --top=N (verify only: how many findings to read, most severe first)
19
21
  //
20
22
  // A model id itself contains a colon (`vendor/name:batch`), so --lens-model
21
23
  // splits on the first colon only: `security:deepseek/deepseek-v4-pro:batch`.
@@ -28,13 +30,14 @@
28
30
  // The parser never throws. index.ts decides how to surface unknown tokens and
29
31
  // invalid combinations, matching parseNextFlags.
30
32
  import { BROADSIDE_LENS_IDS } from "../../core/index.js";
31
- const ACTIONS = new Set(["submit", "collect", "status", "models"]);
33
+ const ACTIONS = new Set(["submit", "collect", "status", "models", "verify"]);
32
34
  /** Every token the completer offers, in the order it offers them. */
33
35
  export const KNOWN_BROADSIDE_TOKENS = [
34
36
  "submit",
35
37
  "collect",
36
38
  "status",
37
39
  "models",
40
+ "verify",
38
41
  ...BROADSIDE_LENS_IDS,
39
42
  "--incremental",
40
43
  "--no-incremental",
@@ -43,6 +46,7 @@ export const KNOWN_BROADSIDE_TOKENS = [
43
46
  "--run=",
44
47
  "--model=",
45
48
  "--lens-model=",
49
+ "--top=",
46
50
  "--no-synthesis",
47
51
  "--no-triage",
48
52
  "--no-retry-truncated",
@@ -122,6 +126,14 @@ export function parseBroadsideFlags(args) {
122
126
  result.runId = value || undefined;
123
127
  continue;
124
128
  }
129
+ if (token.startsWith("--top=")) {
130
+ const value = parseNumeric(token, "--top", result);
131
+ if (value !== undefined && (!Number.isInteger(value) || value < 1))
132
+ result.error ??= `--top needs a positive whole number (got "${token.slice("--top=".length)}").`;
133
+ else if (value !== undefined)
134
+ result.top = value;
135
+ continue;
136
+ }
125
137
  if (token.startsWith("--model=")) {
126
138
  const value = token.slice("--model=".length).trim();
127
139
  // An empty value is a mistyped selection, not "use the default":
@@ -167,11 +179,17 @@ export function parseBroadsideFlags(args) {
167
179
  if (result.action === "status" && result.waitSeconds !== undefined) {
168
180
  result.error ??= "--wait is only meaningful for submit and collect; status reads recorded state.";
169
181
  }
170
- if (result.runId !== undefined && result.action !== "collect") {
171
- result.error ??= `--run is only meaningful for collect (got action "${result.action}").`;
182
+ if (result.runId !== undefined && result.action !== "collect" && result.action !== "verify") {
183
+ result.error ??= `--run is only meaningful for collect and verify (got action "${result.action}").`;
184
+ }
185
+ if (result.model !== undefined && result.action !== "submit" && result.action !== "verify") {
186
+ result.error ??= `--model is only meaningful for submit and verify (got action "${result.action}").`;
187
+ }
188
+ if (result.top !== undefined && result.action !== "verify") {
189
+ result.error ??= `--top is only meaningful for verify (got action "${result.action}").`;
172
190
  }
173
- if (result.model !== undefined && result.action !== "submit") {
174
- result.error ??= `--model is only meaningful for submit (got action "${result.action}").`;
191
+ if (result.action === "verify" && result.waitSeconds !== undefined) {
192
+ result.error ??= "--wait is only meaningful for submit and collect; verify runs to completion.";
175
193
  }
176
194
  if (result.lensModels !== undefined && result.action !== "submit") {
177
195
  result.error ??= `--lens-model is only meaningful for submit (got action "${result.action}").`;
@@ -10,7 +10,7 @@ import { completeLastToken } from "./completions.js";
10
10
  import { buildPiGuideMessage } from "./guide-framing.js";
11
11
  import { isCtxLive, notifyCtx } from "./notify.js";
12
12
  import { phaseCompactionExtension } from "./phase-compaction.js";
13
- import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, statusText, describeConfigProblems, describeIncrementalFallback, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
13
+ import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, runBroadsideVerify, verifyResultText, statusText, describeConfigProblems, describeIncrementalFallback, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
14
14
  import { initLibrary } from "../../core/library.js";
15
15
  import { resolveUserConfigPath } from "../../core/orchestrator-config.js";
16
16
  const STATUS_WIDGET_ID = "codecarto-widget";
@@ -995,7 +995,7 @@ export default function codeCartographerExtension(pi) {
995
995
  },
996
996
  });
997
997
  pi.registerCommand("codecarto-broadside", {
998
- description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID] [flags]",
998
+ description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models|verify] [lenses…] [--model=ID] [--lens-model=LENS:ID] [--top=N] [flags]",
999
999
  // Completes the token under the cursor, so lens names and flags are
1000
1000
  // offered after the action too, and keeps everything typed before it.
1001
1001
  getArgumentCompletions: (prefix) => completeLastToken(prefix, KNOWN_BROADSIDE_TOKENS.map((value) => ({ value }))),
@@ -1181,6 +1181,35 @@ export default function codeCartographerExtension(pi) {
1181
1181
  }
1182
1182
  return;
1183
1183
  }
1184
+ if (flags.action === "verify") {
1185
+ // The verification pass (#143): one sync call per finding with
1186
+ // read-only tools; the widget counts verdicts as they land.
1187
+ const tally = { done: 0 };
1188
+ renderProgress("Reading the top findings against the source…");
1189
+ try {
1190
+ const verified = await runBroadsideVerify(ctx.cwd, apiKey, {
1191
+ ...(flags.runId && { runId: flags.runId }),
1192
+ ...(flags.top !== undefined && { top: flags.top }),
1193
+ ...(flags.model && { model: flags.model }),
1194
+ maxCost: flags.maxCost ?? config.maxCost,
1195
+ signal: ctx.signal,
1196
+ onProgress: (finding) => {
1197
+ tally.done += 1;
1198
+ progress.set(`#${finding.index}`, `${finding.verdict} — ${finding.title}`);
1199
+ renderProgress(`Verifying findings… ${tally.done} read`);
1200
+ },
1201
+ });
1202
+ const lines = verifyResultText(verified).split("\n");
1203
+ const confirmed = verified.findings.filter((f) => f.verdict === "confirmed").length;
1204
+ finish(lines, `Broad-Side verify: ${confirmed} confirmed of ${verified.findings.length} read`, verified.status === "completed" ? "info" : "warning");
1205
+ }
1206
+ catch (error) {
1207
+ if (ctx.hasUI)
1208
+ ctx.ui.setWidget(BROADSIDE_WIDGET_ID, undefined);
1209
+ notifyCtx(ctx, `Broad-Side verify failed: ${error instanceof Error ? error.message : String(error)}`, "error");
1210
+ }
1211
+ return;
1212
+ }
1184
1213
  // action === "collect"
1185
1214
  renderProgress("Polling batches…");
1186
1215
  try {
@@ -233,7 +233,7 @@ export declare function handleAmend(args: {
233
233
  }>;
234
234
  export declare function handleBroadside(args: {
235
235
  cwd: string;
236
- action: "submit" | "collect" | "status" | "models";
236
+ action: "submit" | "collect" | "status" | "models" | "verify";
237
237
  lenses?: string[];
238
238
  api_key?: string;
239
239
  wait_seconds?: number;
@@ -247,6 +247,7 @@ export declare function handleBroadside(args: {
247
247
  incremental?: boolean;
248
248
  model?: string;
249
249
  lens_models?: Record<string, string>;
250
+ top?: number;
250
251
  }): Promise<{
251
252
  content: {
252
253
  type: "text";
@@ -17,7 +17,7 @@ import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"
17
17
  import { CallToolRequestSchema, ErrorCode, ListToolsRequestSchema, McpError, } from "@modelcontextprotocol/sdk/types.js";
18
18
  import { mkdir, readFile, readdir, rename, writeFile } from "node:fs/promises";
19
19
  import { basename, isAbsolute, join } from "node:path";
20
- import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, seedOrchestratorFiles, statusLineWriter, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
20
+ import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, runBroadsideVerify, verifyResultText, seedOrchestratorFiles, statusLineWriter, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
21
21
  import { applyAmendment } from "../core/amendment.js";
22
22
  import { appendUsageRun } from "../core/usage.js";
23
23
  import { initLibrary } from "../core/library.js";
@@ -1037,8 +1037,8 @@ function resolveBroadsideApiKey(explicit, config) {
1037
1037
  export async function handleBroadside(args) {
1038
1038
  const cwd = await validateCwd(args.cwd);
1039
1039
  const action = args.action ?? "submit";
1040
- if (!["submit", "collect", "status", "models"].includes(action)) {
1041
- throw new McpError(ErrorCode.InvalidParams, `Unknown action: ${action}. Valid actions: submit, collect, status, models.`);
1040
+ if (!["submit", "collect", "status", "models", "verify"].includes(action)) {
1041
+ throw new McpError(ErrorCode.InvalidParams, `Unknown action: ${action}. Valid actions: submit, collect, status, models, verify.`);
1042
1042
  }
1043
1043
  // A config.yaml that exists but cannot be read refuses every action that
1044
1044
  // would act on it (#232); status only reads recorded runs, so it answers
@@ -1173,8 +1173,40 @@ export async function handleBroadside(args) {
1173
1173
  maxCost: result.maxCost,
1174
1174
  });
1175
1175
  }
1176
- // action === "collect"
1177
1176
  const runId = typeof args.run_id === "string" && args.run_id.trim() ? args.run_id.trim() : undefined;
1177
+ if (action === "verify") {
1178
+ // The verification pass (#143): one sync call per finding with read-only
1179
+ // tools, most severe first. `max_cost` is a running cap here — a sync
1180
+ // call's cost is known only when it returns — so the pass stops before
1181
+ // the next finding once reached; absent, config.yaml's cap applies.
1182
+ if (args.top !== undefined && !(typeof args.top === "number" && Number.isInteger(args.top) && args.top >= 1)) {
1183
+ throw new McpError(ErrorCode.InvalidParams, "top must be a positive integer.");
1184
+ }
1185
+ if (args.model !== undefined && !(typeof args.model === "string" && args.model.trim())) {
1186
+ throw new McpError(ErrorCode.InvalidParams, "model must be a non-empty OpenRouter model id.");
1187
+ }
1188
+ const maxCost = typeof args.max_cost === "number" && args.max_cost >= 0 ? args.max_cost : config.maxCost;
1189
+ const verified = await runBroadsideVerify(cwd, apiKey, {
1190
+ ...(runId && { runId }),
1191
+ ...(args.top !== undefined && { top: args.top }),
1192
+ ...(typeof args.model === "string" && args.model.trim() && { model: args.model.trim() }),
1193
+ maxCost,
1194
+ signal: serverLifetime?.signal,
1195
+ }).catch((error) => {
1196
+ throw new McpError(ErrorCode.InvalidRequest, error instanceof Error ? error.message : String(error));
1197
+ });
1198
+ return textResult(verifyResultText(verified), {
1199
+ runId: verified.runId,
1200
+ outputDir: verified.outputDir,
1201
+ status: verified.status,
1202
+ model: verified.model,
1203
+ candidates: verified.candidates,
1204
+ totalCost: verified.totalCost,
1205
+ ...(verified.stoppedByCost && { stoppedByCost: true }),
1206
+ findings: verified.findings,
1207
+ });
1208
+ }
1209
+ // action === "collect"
1178
1210
  const collect = await runBroadsideCollect(cwd, apiKey, {
1179
1211
  waitMs,
1180
1212
  includeSynthesis,
@@ -1502,8 +1534,8 @@ const TOOLS = [
1502
1534
  cwd: { type: "string", description: "Absolute path to the target repository." },
1503
1535
  action: {
1504
1536
  type: "string",
1505
- enum: ["submit", "collect", "status", "models"],
1506
- description: "submit fires all lens batches and returns batch ids; collect polls submitted batches, saves results, and optionally runs the synthesis pass; status shows recorded runs; models lists batch-capable models with pricing and capabilities.",
1537
+ enum: ["submit", "collect", "status", "models", "verify"],
1538
+ description: "submit fires all lens batches and returns batch ids; collect polls submitted batches, saves results, and optionally runs the synthesis pass; status shows recorded runs; models lists batch-capable models with pricing and capabilities; verify reads a collected run's top defect and security findings against the repository with read-only tools (one sync-priced call each, about a cent on the default model) and writes verified.md/verified.json beside triage.md with a verdict per finding: confirmed (a reachable failure, with the trigger), not-a-defect, discarded, or unclear.",
1507
1539
  },
1508
1540
  lenses: {
1509
1541
  type: "array",
@@ -1516,7 +1548,11 @@ const TOOLS = [
1516
1548
  },
1517
1549
  run_id: {
1518
1550
  type: "string",
1519
- description: "For collect: the run to collect, as listed by the status action. Defaults to the most recent run; pass this to collect an older run that is still in flight after a newer submit.",
1551
+ description: "For collect and verify: the run to act on, as listed by the status action. Defaults to the most recent run; pass this to collect an older run that is still in flight after a newer submit.",
1552
+ },
1553
+ top: {
1554
+ type: "integer",
1555
+ description: "For verify: how many findings to read, most severe first (default 10). Each costs one sync call; max_cost caps the pass as a running total.",
1520
1556
  },
1521
1557
  wait_seconds: {
1522
1558
  type: "number",
@@ -1536,7 +1572,7 @@ const TOOLS = [
1536
1572
  },
1537
1573
  max_cost: {
1538
1574
  type: "number",
1539
- description: "Approximate run expense limit in USD. The submit action estimates the run cost from slice sizes and the configured model's per-token pricing (live OpenRouter lookup, cached 24h) and refuses to submit when the estimate exceeds the limit unless force is true. Falls back to max_cost in .codecarto/broadside/config.yaml, whose default is $1.00; pass 0 for no limit.",
1575
+ description: "Approximate run expense limit in USD. The submit action estimates the run cost from slice sizes and the configured model's per-token pricing (live OpenRouter lookup, cached 24h) and refuses to submit when the estimate exceeds the limit unless force is true. For verify it is a running cap: the pass stops before the next finding once the calls so far have reached it. Falls back to max_cost in .codecarto/broadside/config.yaml, whose default is $1.00; pass 0 for no limit.",
1540
1576
  },
1541
1577
  force: {
1542
1578
  type: "boolean",
@@ -1552,7 +1588,7 @@ const TOOLS = [
1552
1588
  },
1553
1589
  model: {
1554
1590
  type: "string",
1555
- description: "For submit: the OpenRouter batch model for this run (an id ending in :batch, as listed by action 'models'). Falls back to model in .codecarto/broadside/config.yaml, then the shipped default. Pre-flighted like the configured model: priced from the live catalog, refused without structured-output support, clamped to its completion ceiling. The models listing is advisory — some catalog ids have no batch endpoint and are refused at submit, at no cost; the listing tags ids this repository has already seen accepted or refused.",
1591
+ description: "For verify: the sync (non-batch) OpenRouter model to read with; defaults to the run's model without its :batch suffix. For submit: the OpenRouter batch model for this run (an id ending in :batch, as listed by action 'models'). Falls back to model in .codecarto/broadside/config.yaml, then the shipped default. Pre-flighted like the configured model: priced from the live catalog, refused without structured-output support, clamped to its completion ceiling. The models listing is advisory — some catalog ids have no batch endpoint and are refused at submit, at no cost; the listing tags ids this repository has already seen accepted or refused.",
1556
1592
  },
1557
1593
  lens_models: {
1558
1594
  type: "object",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "codecartographer-pi",
3
- "version": "0.23.0",
3
+ "version": "0.24.0",
4
4
  "mcpName": "io.github.HuginnIndustries/codecartographer",
5
5
  "description": "Turn an unfamiliar codebase into a validated reimplementation spec, then synthesize confirmed specs and a product vision into a traceable plan.",
6
6
  "type": "module",