codecartographer-pi 0.23.0 → 0.24.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codecarto/broadside/SKILL.md +29 -1
- package/.codecarto/workflow/scaffold-version.yaml +1 -1
- package/README.md +4 -1
- package/agent-skill/codecartographer/references/broadside.md +5 -1
- package/dist/core/broadside-verify.d.ts +95 -0
- package/dist/core/broadside-verify.js +433 -0
- package/dist/core/broadside.d.ts +26 -0
- package/dist/core/broadside.js +9 -2
- package/dist/core/index.d.ts +1 -0
- package/dist/core/index.js +1 -0
- package/dist/extensions/codecarto/broadside-flags.d.ts +4 -2
- package/dist/extensions/codecarto/broadside-flags.js +24 -6
- package/dist/extensions/codecarto/index.js +31 -2
- package/dist/mcp-server/server.d.ts +2 -1
- package/dist/mcp-server/server.js +45 -9
- package/package.json +1 -1
|
@@ -26,7 +26,21 @@ not replace any phase; it tells phases where to look.
|
|
|
26
26
|
|
|
27
27
|
## Reading a Broad-Side run
|
|
28
28
|
|
|
29
|
-
|
|
29
|
+
0. If `verified.md` exists, read it first. A verification pass
|
|
30
|
+
(`codecarto_broadside {action: "verify"}`, `/codecarto-broadside verify`)
|
|
31
|
+
has read the top defect and security findings against the real source with
|
|
32
|
+
read-only tools and given each a verdict: **confirmed** (a reachable
|
|
33
|
+
failure, with the input or call site that triggers it), **not-a-defect**
|
|
34
|
+
(the claim is literally true of the code but nothing can reach the failure
|
|
35
|
+
it describes), **discarded** (the claim is wrong about the code, with the
|
|
36
|
+
guard or line that shows it), or **unclear**. Start from the confirmed
|
|
37
|
+
ones; treat a discarded one as answered unless the reasoning is thin.
|
|
38
|
+
Measured on this repository, the top twelve findings by severity were two
|
|
39
|
+
real defects and ten that a look at the guard, the caller, or the tsconfig
|
|
40
|
+
dismissed — the pass agreed with a reviewer on all twelve for about a cent
|
|
41
|
+
a finding. A confirmed verdict is still a model's reading: a strong lead for
|
|
42
|
+
a human's next look, not a validated claim.
|
|
43
|
+
1. Read `synthesis.md`. It carries the executive summary, severity counts,
|
|
30
44
|
the top cross-lens findings, and per-module risk levels.
|
|
31
45
|
2. Read `triage.md` for the work order: each lead scored by impact ×
|
|
32
46
|
difficulty with a P0–P3 priority and an effort estimate. It is a starting
|
|
@@ -66,6 +80,7 @@ Broad-Side is an executable-surface feature. On the Pi extension:
|
|
|
66
80
|
/codecarto-broadside collect # poll, save, synthesize
|
|
67
81
|
/codecarto-broadside status # show recorded runs
|
|
68
82
|
/codecarto-broadside models # compare batch models
|
|
83
|
+
/codecarto-broadside verify --top=10 # read the top findings against the source
|
|
69
84
|
```
|
|
70
85
|
|
|
71
86
|
On the MCP server:
|
|
@@ -75,6 +90,7 @@ codecarto_broadside {cwd, action: "submit", lenses: [...]} # fire the batches
|
|
|
75
90
|
codecarto_broadside {cwd, action: "collect"} # poll, save, synthesize
|
|
76
91
|
codecarto_broadside {cwd, action: "status"} # show recorded runs
|
|
77
92
|
codecarto_broadside {cwd, action: "models"} # compare batch models
|
|
93
|
+
codecarto_broadside {cwd, action: "verify", top: 10} # read the top findings against the source
|
|
78
94
|
```
|
|
79
95
|
|
|
80
96
|
The `models` action lists every `:batch` variant on OpenRouter — pricing per
|
|
@@ -105,6 +121,18 @@ Collect runs two cross-lens post-passes by default: **synthesis** (the
|
|
|
105
121
|
executive report) and **triage** (the prioritized work order). Pass
|
|
106
122
|
`include_synthesis: false` or `include_triage: false` on collect to skip one.
|
|
107
123
|
|
|
124
|
+
`verify` is the third pass, run separately after collect because it is
|
|
125
|
+
sync-priced rather than batch-priced: for each of the top `top` findings
|
|
126
|
+
(default 10, most severe first) of the defect and security lenses, one chat
|
|
127
|
+
completion on the run's model without its `:batch` suffix (`model` overrides),
|
|
128
|
+
with three read-only tools confined to the repository's source files —
|
|
129
|
+
`read_file` by line range, `grep`, `list_dir` — at most eight tool calls, low
|
|
130
|
+
reasoning effort, then a verdict. It writes `verified.md` and `verified.json`
|
|
131
|
+
beside `triage.md` and records the pass on the run. `max_cost` is a running
|
|
132
|
+
cap here, since a sync call's cost is known only when it returns: the pass
|
|
133
|
+
stops before the next finding once the calls so far have reached it and
|
|
134
|
+
reports `partial`. About a cent a finding on the default model.
|
|
135
|
+
|
|
108
136
|
Two caveats apply to any model you pick. The `models` action lists every id
|
|
109
137
|
OpenRouter advertises a `:batch` variant for, and many of those variants do not
|
|
110
138
|
exist — submitting one returns `does not have a :batch endpoint`, with nothing in
|
package/README.md
CHANGED
|
@@ -394,13 +394,16 @@ codecarto_broadside {cwd, action: "models"} # compare batch mo
|
|
|
394
394
|
codecarto_broadside {cwd, action: "submit", lenses: [...]} # fire the batches, priced first
|
|
395
395
|
codecarto_broadside {cwd, action: "status"} # what is in flight
|
|
396
396
|
codecarto_broadside {cwd, action: "collect"} # poll, save, synthesize, triage
|
|
397
|
+
codecarto_broadside {cwd, action: "verify", top: 10} # read the top findings against the source
|
|
397
398
|
```
|
|
398
399
|
|
|
400
|
+
A third action, `verify`, reads the top defect and security findings of a collected run against the real source — one sync-priced call each with read-only tools (`read_file`, `grep`, `list_dir`) confined to the repository — and writes `verified.md` with a verdict per finding: **confirmed** (a reachable failure, with the trigger), **not-a-defect**, **discarded**, or **unclear**. Measured on this repository, the top twelve findings by severity were two real defects and ten claims a look at the guard or the caller dismissed; the pass agreed with a reviewer on all twelve for a cent a finding. Read `verified.md` before `triage.md`.
|
|
401
|
+
|
|
399
402
|
Submit and collect are separate because batch jobs routinely take tens of minutes; collect is resumable and picks up whatever is still in flight (`wait_seconds: 0`, the default, polls once and returns), and two collects on one run — a retried tool call, a second session — never pay for the synthesis, triage, or truncation retry twice: each is claimed in the run's state before it is submitted, and a collect whose client has gone away stops polling and submits nothing further. Submit prices the run from the collected file sizes against the model's live per-token pricing (cached 24h) and refuses when the estimate exceeds `max_cost` — $1.00 unless the config or the call sets another value, `0` for no limit — unless `force: true` is passed — a pre-flight estimate, not a runtime stop. Actual spend lands in each run's `run-meta.json`.
|
|
400
403
|
|
|
401
404
|
Repository defaults live in `.codecarto/broadside/config.yaml` (`model`, `api_key`, `default_lenses`, `max_cost`, `pricing` overrides, `lens_models`, `incremental`, `retry_truncated`, `include_synthesis`, `include_triage`, `wait_seconds`); an explicit tool parameter always wins. `lens_models` routes individual lenses to their own batch model — a stronger model changes security and defect findings far more than it changes an architecture map — and each override is priced, capability-checked, and clamped exactly like the default, with the estimate broken out per lens so a mixed-model run cannot be approved without seeing which lens costs what. CodeCartographer ships no stronger default: which model earns its price depends on your repository and budget, so compare candidates with the `models` action and choose — for one run with the `model` and `lens_models` parameters (Pi: `--model=ID`, `--lens-model=LENS:ID`), or for the repository in `config.yaml`. The `models` listing is advisory: OpenRouter's catalog returns a `:batch` id for some models its Batch API then refuses (`does not have a :batch endpoint`), at no cost, and nothing in the catalog tells them apart — so the listing tags the ids this repository's own submits have seen accepted or refused, and a refused lens says why in the submit report. `codecarto_skill {cwd, name: "broadside"}` returns the reading guide for a completed run, and unlike post-pipeline skills it is not gated on a finished pipeline.
|
|
402
405
|
|
|
403
|
-
On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID]`, with tab-completion for actions, lens names, and flags and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
|
|
406
|
+
On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models|verify] [lenses…] [--model=ID] [--lens-model=LENS:ID] [--top=N]`, with tab-completion for actions, lens names, and flags and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
|
|
404
407
|
|
|
405
408
|
Broad-Side needs runtime code, so firing a run is an executable-surface feature: Pi and MCP have it, the pure drop-in template does not (it carries only the reading guide).
|
|
406
409
|
|
|
@@ -95,7 +95,11 @@ Two more economies worth knowing:
|
|
|
95
95
|
|
|
96
96
|
## Reading a run
|
|
97
97
|
|
|
98
|
-
Results land in `.codecarto/broadside/<run>/`.
|
|
98
|
+
Results land in `.codecarto/broadside/<run>/`. If `verified.md` is there, read
|
|
99
|
+
it before anything else: `action: "verify"` has read the top defect and
|
|
100
|
+
security findings against the source with read-only tools and given each a
|
|
101
|
+
verdict (confirmed with its trigger, not-a-defect, discarded with the guard
|
|
102
|
+
that shows it, unclear). Then, in this order:
|
|
99
103
|
|
|
100
104
|
1. `synthesis.md` — executive summary, severity counts, top cross-lens
|
|
101
105
|
findings, per-module risk.
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
import { type BroadsideLensId, type BroadsideVerifyEntry, type FetchLike, type StoredLensResult } from "./broadside.ts";
|
|
2
|
+
export declare const BROADSIDE_CHAT_URL = "https://openrouter.ai/api/v1/chat/completions";
|
|
3
|
+
/** How many findings `verify` reads by default, most severe first. */
|
|
4
|
+
export declare const BROADSIDE_VERIFY_DEFAULT_TOP = 10;
|
|
5
|
+
/** Tool calls one finding may spend before it must answer. */
|
|
6
|
+
export declare const BROADSIDE_VERIFY_MAX_TOOL_CALLS = 8;
|
|
7
|
+
/** The lenses whose findings carry a file:line and a claim to check. */
|
|
8
|
+
export declare const BROADSIDE_VERIFIABLE_LENSES: readonly BroadsideLensId[];
|
|
9
|
+
export type VerifyVerdict = "confirmed" | "not-a-defect" | "discarded" | "unclear";
|
|
10
|
+
export type VerifiedFinding = {
|
|
11
|
+
index: number;
|
|
12
|
+
lensId: BroadsideLensId;
|
|
13
|
+
customId: string;
|
|
14
|
+
severity: string;
|
|
15
|
+
title: string;
|
|
16
|
+
location: string;
|
|
17
|
+
verdict: VerifyVerdict | "error";
|
|
18
|
+
confidence: string;
|
|
19
|
+
evidence: Array<{
|
|
20
|
+
file: string;
|
|
21
|
+
lines: string;
|
|
22
|
+
note: string;
|
|
23
|
+
}>;
|
|
24
|
+
reasoning: string;
|
|
25
|
+
toolCalls: number;
|
|
26
|
+
cost: number;
|
|
27
|
+
};
|
|
28
|
+
export type BroadsideVerifyResult = {
|
|
29
|
+
runId: string;
|
|
30
|
+
outputDir: string;
|
|
31
|
+
model: string;
|
|
32
|
+
status: BroadsideVerifyEntry["status"];
|
|
33
|
+
/** How many findings the run had in the verifiable lenses. */
|
|
34
|
+
candidates: number;
|
|
35
|
+
findings: VerifiedFinding[];
|
|
36
|
+
totalCost: number;
|
|
37
|
+
/** Set when the cost cap stopped the pass before every selected finding was read. */
|
|
38
|
+
stoppedByCost?: boolean;
|
|
39
|
+
};
|
|
40
|
+
type CandidateFinding = {
|
|
41
|
+
lensId: BroadsideLensId;
|
|
42
|
+
customId: string;
|
|
43
|
+
severity: string;
|
|
44
|
+
title: string;
|
|
45
|
+
location: string;
|
|
46
|
+
description: string;
|
|
47
|
+
pattern: string;
|
|
48
|
+
};
|
|
49
|
+
/** The sync-priced id behind a `:batch` model id (`vendor/name:batch` → `vendor/name`). */
|
|
50
|
+
export declare function syncModelFor(batchModel: string): string;
|
|
51
|
+
/** The findings a run's saved lens results carry, most severe first. */
|
|
52
|
+
export declare function rankVerifiableFindings(stored: StoredLensResult[]): CandidateFinding[];
|
|
53
|
+
export type RepoReader = {
|
|
54
|
+
readFile(path: string, startLine?: number, endLine?: number): Promise<string>;
|
|
55
|
+
grep(pattern: string, pathPrefix?: string): Promise<string>;
|
|
56
|
+
listDir(path: string): Promise<string>;
|
|
57
|
+
};
|
|
58
|
+
/**
|
|
59
|
+
* Three read-only tools over the repository's own file list — the same
|
|
60
|
+
* listing the lenses scan (tracked and untracked, ignore rules applied) minus
|
|
61
|
+
* everything {@link isSlurpable} keeps out of a lens: credential stores, build
|
|
62
|
+
* output, binaries. A path outside the repository, or one the listing does
|
|
63
|
+
* not contain, is an error the model sees, not a read.
|
|
64
|
+
*/
|
|
65
|
+
export declare function createRepoReader(cwd: string): Promise<RepoReader>;
|
|
66
|
+
/**
|
|
67
|
+
* The rubric. The order is deliberate — it is the order a reviewer settles a
|
|
68
|
+
* claim in — and `not-a-defect` is the verdict that separates "the code does
|
|
69
|
+
* what the claim says" from "and that is a bug": without it, a cast every
|
|
70
|
+
* caller satisfies gets confirmed because it is literally there.
|
|
71
|
+
*/
|
|
72
|
+
export declare const BROADSIDE_VERIFY_SYSTEM_PROMPT: string;
|
|
73
|
+
/** Verify one finding: up to the tool budget, then a verdict. */
|
|
74
|
+
export declare function verifyFinding(finding: CandidateFinding, index: number, reader: RepoReader, apiKey: string, model: string, fetcher: FetchLike): Promise<Omit<VerifiedFinding, "index" | "lensId" | "customId" | "severity" | "title" | "location">>;
|
|
75
|
+
/**
|
|
76
|
+
* Verify the top findings of a collected run against the repository.
|
|
77
|
+
*
|
|
78
|
+
* `maxCost` is a running cap, not a pre-flight estimate: a sync call's cost
|
|
79
|
+
* is only known when it returns, so the pass stops *before* starting the next
|
|
80
|
+
* finding once the cap is reached and reports `partial`. On the default
|
|
81
|
+
* model a finding costs about a cent.
|
|
82
|
+
*/
|
|
83
|
+
export declare function runBroadsideVerify(cwd: string, apiKey: string, opts?: {
|
|
84
|
+
runId?: string;
|
|
85
|
+
top?: number;
|
|
86
|
+
model?: string;
|
|
87
|
+
/** USD; 0 means no limit. */
|
|
88
|
+
maxCost?: number;
|
|
89
|
+
fetcher?: FetchLike;
|
|
90
|
+
signal?: AbortSignal;
|
|
91
|
+
onProgress?: (finding: VerifiedFinding) => void;
|
|
92
|
+
}): Promise<BroadsideVerifyResult>;
|
|
93
|
+
export declare function renderVerifiedMarkdown(result: BroadsideVerifyResult): string;
|
|
94
|
+
export declare function verifyResultText(result: BroadsideVerifyResult): string;
|
|
95
|
+
export {};
|
|
@@ -0,0 +1,433 @@
|
|
|
1
|
+
// Broad-Side verification pass (#143): the hybrid the roadmap kept coming
|
|
2
|
+
// back to. The batch sweep is cheap and reads files without being able to
|
|
3
|
+
// look anything up, and its measured weakness is precision, not coverage —
|
|
4
|
+
// on this repository the top twelve findings by severity were two real
|
|
5
|
+
// defects and ten claims that a look at the guard, the caller, or the
|
|
6
|
+
// tsconfig would have dismissed. So after collect, one sync-priced call per
|
|
7
|
+
// finding, with three read-only tools confined to the repository, reads the
|
|
8
|
+
// cited code and says whether the failure is reachable.
|
|
9
|
+
//
|
|
10
|
+
// Measured on 2026-09-13 (the #143 comparison run): twelve findings, the
|
|
11
|
+
// rubric below, `google/gemini-3.7-flash` at low reasoning effort —
|
|
12
|
+
// twelve-for-twelve agreement with a reviewer's ground truth, precision of
|
|
13
|
+
// the confirmed set from 17% to 100%, both true findings kept, $0.13 in all
|
|
14
|
+
// (about a cent a finding). The rubric mattered more than the model: a first
|
|
15
|
+
// draft without the `not-a-defect` verdict "confirmed" two type casts that
|
|
16
|
+
// every caller satisfies, because they were literally true of the code.
|
|
17
|
+
//
|
|
18
|
+
// Read-only on purpose. A tool-using pass that could mutate would need the
|
|
19
|
+
// headless-agent retry rule (retry only before the first tool call); one
|
|
20
|
+
// that only reads keeps the batch property that re-running is always safe.
|
|
21
|
+
import { readdir, readFile, stat, writeFile } from "node:fs/promises";
|
|
22
|
+
import { isAbsolute, join, relative, resolve } from "node:path";
|
|
23
|
+
import { BROADSIDE_DIR, BROADSIDE_LENS_IDS, broadsideDirFor, defaultReasoningFor, isSlurpable, listRepoFiles, loadBroadsideState, loadSavedLensResults, parseLensJson, persistBroadsideRunMerging, } from "./broadside.js";
|
|
24
|
+
export const BROADSIDE_CHAT_URL = "https://openrouter.ai/api/v1/chat/completions";
|
|
25
|
+
/** How many findings `verify` reads by default, most severe first. */
|
|
26
|
+
export const BROADSIDE_VERIFY_DEFAULT_TOP = 10;
|
|
27
|
+
/** Tool calls one finding may spend before it must answer. */
|
|
28
|
+
export const BROADSIDE_VERIFY_MAX_TOOL_CALLS = 8;
|
|
29
|
+
/** The lenses whose findings carry a file:line and a claim to check. */
|
|
30
|
+
export const BROADSIDE_VERIFIABLE_LENSES = ["defect", "security"];
|
|
31
|
+
const READ_FILE_MAX_LINES = 200;
|
|
32
|
+
const GREP_MAX_MATCHES = 40;
|
|
33
|
+
const GREP_MAX_FILE_BYTES = 1_000_000;
|
|
34
|
+
const TOOL_OUTPUT_MAX_CHARS = 12_000;
|
|
35
|
+
const SEVERITY_RANK = { critical: 0, high: 1, medium: 2, low: 3 };
|
|
36
|
+
/** The sync-priced id behind a `:batch` model id (`vendor/name:batch` → `vendor/name`). */
|
|
37
|
+
export function syncModelFor(batchModel) {
|
|
38
|
+
return batchModel.replace(/:batch$/, "");
|
|
39
|
+
}
|
|
40
|
+
/** The findings a run's saved lens results carry, most severe first. */
|
|
41
|
+
export function rankVerifiableFindings(stored) {
|
|
42
|
+
const out = [];
|
|
43
|
+
for (const result of stored) {
|
|
44
|
+
if (!BROADSIDE_VERIFIABLE_LENSES.includes(result.lensId) || result.truncated)
|
|
45
|
+
continue;
|
|
46
|
+
const parsed = parseLensJson(result.content);
|
|
47
|
+
for (const finding of parsed?.findings ?? []) {
|
|
48
|
+
const location = typeof finding.location === "string" ? finding.location : "";
|
|
49
|
+
const title = typeof finding.title === "string" ? finding.title : "";
|
|
50
|
+
if (!location || !title)
|
|
51
|
+
continue;
|
|
52
|
+
out.push({
|
|
53
|
+
lensId: result.lensId,
|
|
54
|
+
customId: result.customId,
|
|
55
|
+
severity: typeof finding.severity === "string" ? finding.severity.toLowerCase() : "unknown",
|
|
56
|
+
title,
|
|
57
|
+
location,
|
|
58
|
+
description: typeof finding.description === "string" ? finding.description : "",
|
|
59
|
+
pattern: typeof finding.pattern === "string" ? finding.pattern : typeof finding.category === "string" ? finding.category : "",
|
|
60
|
+
});
|
|
61
|
+
}
|
|
62
|
+
}
|
|
63
|
+
// Stable: severity, then lens order, then the order the lens listed them.
|
|
64
|
+
return out
|
|
65
|
+
.map((finding, order) => ({ finding, order }))
|
|
66
|
+
.sort((a, b) => (SEVERITY_RANK[a.finding.severity] ?? 9) - (SEVERITY_RANK[b.finding.severity] ?? 9)
|
|
67
|
+
|| BROADSIDE_LENS_IDS.indexOf(a.finding.lensId) - BROADSIDE_LENS_IDS.indexOf(b.finding.lensId)
|
|
68
|
+
|| a.order - b.order)
|
|
69
|
+
.map(({ finding }) => finding);
|
|
70
|
+
}
|
|
71
|
+
/**
|
|
72
|
+
* Three read-only tools over the repository's own file list — the same
|
|
73
|
+
* listing the lenses scan (tracked and untracked, ignore rules applied) minus
|
|
74
|
+
* everything {@link isSlurpable} keeps out of a lens: credential stores, build
|
|
75
|
+
* output, binaries. A path outside the repository, or one the listing does
|
|
76
|
+
* not contain, is an error the model sees, not a read.
|
|
77
|
+
*/
|
|
78
|
+
export async function createRepoReader(cwd) {
|
|
79
|
+
const { files } = await listRepoFiles(cwd);
|
|
80
|
+
const readable = new Set(files.filter(isSlurpable));
|
|
81
|
+
const confine = (path) => {
|
|
82
|
+
const abs = resolve(cwd, path);
|
|
83
|
+
const rel = relative(cwd, abs).split("\\").join("/");
|
|
84
|
+
if (!rel || rel.startsWith("..") || isAbsolute(rel))
|
|
85
|
+
throw new Error(`path is outside the repository: ${path}`);
|
|
86
|
+
return rel;
|
|
87
|
+
};
|
|
88
|
+
return {
|
|
89
|
+
async readFile(path, startLine, endLine) {
|
|
90
|
+
const rel = confine(path);
|
|
91
|
+
if (!readable.has(rel))
|
|
92
|
+
return `error: ${rel} is not a readable source file of this repository`;
|
|
93
|
+
const lines = (await readFile(join(cwd, rel), "utf8")).split("\n");
|
|
94
|
+
const start = Math.max(1, Math.floor(Number(startLine ?? 1)) || 1);
|
|
95
|
+
const requestedEnd = Math.floor(Number(endLine ?? start + READ_FILE_MAX_LINES - 1)) || start;
|
|
96
|
+
const end = Math.min(lines.length, requestedEnd, start + READ_FILE_MAX_LINES - 1);
|
|
97
|
+
if (start > lines.length)
|
|
98
|
+
return `error: ${rel} has ${lines.length} lines`;
|
|
99
|
+
return lines.slice(start - 1, end).map((line, i) => `${start + i}: ${line}`).join("\n");
|
|
100
|
+
},
|
|
101
|
+
async grep(pattern, pathPrefix) {
|
|
102
|
+
let regex;
|
|
103
|
+
try {
|
|
104
|
+
regex = new RegExp(pattern);
|
|
105
|
+
}
|
|
106
|
+
catch (error) {
|
|
107
|
+
return `error: invalid pattern (${error instanceof Error ? error.message : String(error)})`;
|
|
108
|
+
}
|
|
109
|
+
const prefix = pathPrefix ? confine(pathPrefix) : "";
|
|
110
|
+
const matches = [];
|
|
111
|
+
for (const rel of readable) {
|
|
112
|
+
if (prefix && rel !== prefix && !rel.startsWith(`${prefix}/`) && !rel.startsWith(prefix))
|
|
113
|
+
continue;
|
|
114
|
+
let text;
|
|
115
|
+
try {
|
|
116
|
+
if ((await stat(join(cwd, rel))).size > GREP_MAX_FILE_BYTES)
|
|
117
|
+
continue;
|
|
118
|
+
text = await readFile(join(cwd, rel), "utf8");
|
|
119
|
+
}
|
|
120
|
+
catch {
|
|
121
|
+
continue;
|
|
122
|
+
}
|
|
123
|
+
const lines = text.split("\n");
|
|
124
|
+
for (let i = 0; i < lines.length && matches.length < GREP_MAX_MATCHES; i++) {
|
|
125
|
+
if (regex.test(lines[i]))
|
|
126
|
+
matches.push(`${rel}:${i + 1}: ${lines[i]}`);
|
|
127
|
+
}
|
|
128
|
+
if (matches.length >= GREP_MAX_MATCHES)
|
|
129
|
+
break;
|
|
130
|
+
}
|
|
131
|
+
return matches.length > 0 ? matches.join("\n") : "(no matches)";
|
|
132
|
+
},
|
|
133
|
+
async listDir(path) {
|
|
134
|
+
const rel = path === "." || path === "" ? "" : confine(path);
|
|
135
|
+
try {
|
|
136
|
+
const entries = await readdir(join(cwd, rel), { withFileTypes: true });
|
|
137
|
+
return entries
|
|
138
|
+
.filter((entry) => {
|
|
139
|
+
const child = rel ? `${rel}/${entry.name}` : entry.name;
|
|
140
|
+
return entry.isDirectory() ? [...readable].some((f) => f.startsWith(`${child}/`)) : readable.has(child);
|
|
141
|
+
})
|
|
142
|
+
.map((entry) => (entry.isDirectory() ? `${entry.name}/` : entry.name))
|
|
143
|
+
.sort()
|
|
144
|
+
.join("\n") || "(empty)";
|
|
145
|
+
}
|
|
146
|
+
catch (error) {
|
|
147
|
+
return `error: ${error instanceof Error ? error.message : String(error)}`;
|
|
148
|
+
}
|
|
149
|
+
},
|
|
150
|
+
};
|
|
151
|
+
}
|
|
152
|
+
const TOOLS = [
|
|
153
|
+
{ type: "function", function: { name: "read_file", description: "Read a range of lines (1-based, inclusive) from a source file of the repository; at most 200 lines per call.", parameters: { type: "object", properties: { path: { type: "string" }, start_line: { type: "integer" }, end_line: { type: "integer" } }, required: ["path"] } } },
|
|
154
|
+
{ type: "function", function: { name: "grep", description: "Search the repository's source files for a regular expression; returns at most 40 matching lines as path:line: text.", parameters: { type: "object", properties: { pattern: { type: "string" }, path_prefix: { type: "string", description: "optional directory or file to search under" } }, required: ["pattern"] } } },
|
|
155
|
+
{ type: "function", function: { name: "list_dir", description: "List a directory of the repository.", parameters: { type: "object", properties: { path: { type: "string" } }, required: ["path"] } } },
|
|
156
|
+
];
|
|
157
|
+
const VERDICT_SCHEMA = {
|
|
158
|
+
name: "broadside_verification",
|
|
159
|
+
strict: true,
|
|
160
|
+
schema: {
|
|
161
|
+
type: "object",
|
|
162
|
+
properties: {
|
|
163
|
+
verdict: { type: "string", enum: ["confirmed", "not-a-defect", "discarded", "unclear"] },
|
|
164
|
+
confidence: { type: "string", enum: ["high", "medium", "low"] },
|
|
165
|
+
evidence: {
|
|
166
|
+
type: "array",
|
|
167
|
+
items: { type: "object", properties: { file: { type: "string" }, lines: { type: "string" }, note: { type: "string" } }, required: ["file", "lines", "note"], additionalProperties: false },
|
|
168
|
+
},
|
|
169
|
+
reasoning: { type: "string" },
|
|
170
|
+
},
|
|
171
|
+
required: ["verdict", "confidence", "evidence", "reasoning"],
|
|
172
|
+
additionalProperties: false,
|
|
173
|
+
},
|
|
174
|
+
};
|
|
175
|
+
/**
|
|
176
|
+
* The rubric. The order is deliberate — it is the order a reviewer settles a
|
|
177
|
+
* claim in — and `not-a-defect` is the verdict that separates "the code does
|
|
178
|
+
* what the claim says" from "and that is a bug": without it, a cast every
|
|
179
|
+
* caller satisfies gets confirmed because it is literally there.
|
|
180
|
+
*/
|
|
181
|
+
export const BROADSIDE_VERIFY_SYSTEM_PROMPT = "You verify a scouting finding produced by a one-shot batch scan against the real source code. " +
|
|
182
|
+
"The scan saw files without being able to look anything up; you can. Use the tools to read the cited location " +
|
|
183
|
+
"and whatever else the claim depends on (callers, the definition of a helper, error handling around it). " +
|
|
184
|
+
"Then decide, in this order: " +
|
|
185
|
+
"'discarded' — the claim is wrong about the code (the guard exists, the value cannot be what the claim assumes, " +
|
|
186
|
+
"the cited line does something else, the condition it fears is ruled out by the project's config or runtime); " +
|
|
187
|
+
"'not-a-defect' — the claim is literally true of the code but no caller, input, or state can reach the failure it " +
|
|
188
|
+
"describes: a TypeScript cast every caller satisfies, a hypothetical about an environment the project does not " +
|
|
189
|
+
"target, a style or type-hygiene observation; " +
|
|
190
|
+
"'confirmed' — the failure is reachable: name the concrete input, call site, or sequence that triggers it, and what " +
|
|
191
|
+
"then goes wrong; " +
|
|
192
|
+
"'unclear' — settling it needs runtime behaviour or specification knowledge the code does not contain. " +
|
|
193
|
+
"Be strict: a real but different problem than the one claimed is 'discarded' with the difference noted, and " +
|
|
194
|
+
"'confirmed' without a trigger you found in the code is not allowed. " +
|
|
195
|
+
"Cite line ranges you actually read. When you are done, reply with only the JSON verdict object.";
|
|
196
|
+
function findingPrompt(index, finding) {
|
|
197
|
+
return (`Finding ${index} (${finding.lensId} lens, severity ${finding.severity}):\n` +
|
|
198
|
+
`Title: ${finding.title}\nLocation: ${finding.location}\nPattern/category: ${finding.pattern || "-"}\n` +
|
|
199
|
+
`Description: ${finding.description}\n\n` +
|
|
200
|
+
"Verify it. Reply with a JSON object {verdict, confidence, evidence:[{file,lines,note}], reasoning}.");
|
|
201
|
+
}
|
|
202
|
+
async function chat(fetcher, apiKey, model, body) {
|
|
203
|
+
const resp = await fetcher(BROADSIDE_CHAT_URL, {
|
|
204
|
+
method: "POST",
|
|
205
|
+
headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
|
|
206
|
+
body: JSON.stringify({ model, reasoning: defaultReasoningFor(), usage: { include: true }, ...body }),
|
|
207
|
+
signal: AbortSignal.timeout(120_000),
|
|
208
|
+
});
|
|
209
|
+
const data = (await resp.json());
|
|
210
|
+
if (!resp.ok || data.error) {
|
|
211
|
+
const detail = data.error?.message ?? JSON.stringify(data.error ?? data).slice(0, 300);
|
|
212
|
+
throw new Error(`OpenRouter chat: HTTP ${resp.status}: ${detail}`);
|
|
213
|
+
}
|
|
214
|
+
return data;
|
|
215
|
+
}
|
|
216
|
+
function parseVerdict(text) {
|
|
217
|
+
const trimmed = text.trim();
|
|
218
|
+
const fenced = /```(?:json)?\s*([\s\S]*?)```/.exec(trimmed);
|
|
219
|
+
try {
|
|
220
|
+
const parsed = JSON.parse(fenced ? fenced[1] : trimmed);
|
|
221
|
+
return typeof parsed?.verdict === "string" ? parsed : null;
|
|
222
|
+
}
|
|
223
|
+
catch {
|
|
224
|
+
return null;
|
|
225
|
+
}
|
|
226
|
+
}
|
|
227
|
+
/** Verify one finding: up to the tool budget, then a verdict. */
|
|
228
|
+
export async function verifyFinding(finding, index, reader, apiKey, model, fetcher) {
|
|
229
|
+
const messages = [
|
|
230
|
+
{ role: "system", content: BROADSIDE_VERIFY_SYSTEM_PROMPT },
|
|
231
|
+
{ role: "user", content: findingPrompt(index, finding) },
|
|
232
|
+
];
|
|
233
|
+
let toolCalls = 0;
|
|
234
|
+
let cost = 0;
|
|
235
|
+
const usageCost = (data) => {
|
|
236
|
+
const usage = data.usage;
|
|
237
|
+
return typeof usage?.cost === "number" ? usage.cost : 0;
|
|
238
|
+
};
|
|
239
|
+
try {
|
|
240
|
+
// The +2 leaves room for the final schema-forced call after the budget.
|
|
241
|
+
for (let step = 0; step < BROADSIDE_VERIFY_MAX_TOOL_CALLS + 2; step++) {
|
|
242
|
+
const budgetSpent = toolCalls >= BROADSIDE_VERIFY_MAX_TOOL_CALLS;
|
|
243
|
+
if (budgetSpent)
|
|
244
|
+
messages.push({ role: "user", content: "Tool budget spent. Emit the verdict JSON object now from what you have read." });
|
|
245
|
+
const data = await chat(fetcher, apiKey, model, budgetSpent
|
|
246
|
+
? { messages, response_format: { type: "json_schema", json_schema: VERDICT_SCHEMA }, max_tokens: 2000 }
|
|
247
|
+
: { messages, tools: TOOLS, tool_choice: "auto", max_tokens: 4000 });
|
|
248
|
+
cost += usageCost(data);
|
|
249
|
+
const choice = data.choices?.[0];
|
|
250
|
+
const message = choice?.message ?? { role: "assistant", content: "" };
|
|
251
|
+
messages.push(message);
|
|
252
|
+
const calls = message.tool_calls;
|
|
253
|
+
if (calls && calls.length > 0 && !budgetSpent) {
|
|
254
|
+
for (const call of calls) {
|
|
255
|
+
toolCalls += 1;
|
|
256
|
+
let args = {};
|
|
257
|
+
try {
|
|
258
|
+
args = JSON.parse(call.function.arguments || "{}");
|
|
259
|
+
}
|
|
260
|
+
catch {
|
|
261
|
+
// Malformed arguments: the model sees the error and can retry.
|
|
262
|
+
}
|
|
263
|
+
let output;
|
|
264
|
+
try {
|
|
265
|
+
output = call.function.name === "read_file"
|
|
266
|
+
? await reader.readFile(String(args.path ?? ""), args.start_line, args.end_line)
|
|
267
|
+
: call.function.name === "grep"
|
|
268
|
+
? await reader.grep(String(args.pattern ?? ""), typeof args.path_prefix === "string" ? args.path_prefix : undefined)
|
|
269
|
+
: call.function.name === "list_dir"
|
|
270
|
+
? await reader.listDir(String(args.path ?? "."))
|
|
271
|
+
: `error: unknown tool ${call.function.name}`;
|
|
272
|
+
}
|
|
273
|
+
catch (error) {
|
|
274
|
+
output = `error: ${error instanceof Error ? error.message : String(error)}`;
|
|
275
|
+
}
|
|
276
|
+
messages.push({ role: "tool", tool_call_id: call.id, content: output.slice(0, TOOL_OUTPUT_MAX_CHARS) });
|
|
277
|
+
}
|
|
278
|
+
continue;
|
|
279
|
+
}
|
|
280
|
+
const verdict = parseVerdict(typeof message.content === "string" ? message.content : "");
|
|
281
|
+
if (verdict)
|
|
282
|
+
return finish(verdict, toolCalls, cost);
|
|
283
|
+
if (budgetSpent)
|
|
284
|
+
break;
|
|
285
|
+
// Prose instead of JSON: one schema-forced call, no tools.
|
|
286
|
+
const final = await chat(fetcher, apiKey, model, {
|
|
287
|
+
messages: [...messages, { role: "user", content: "Emit the verdict JSON object now." }],
|
|
288
|
+
response_format: { type: "json_schema", json_schema: VERDICT_SCHEMA },
|
|
289
|
+
max_tokens: 2000,
|
|
290
|
+
});
|
|
291
|
+
cost += usageCost(final);
|
|
292
|
+
const forced = parseVerdict(String((final.choices?.[0]?.message?.content) ?? ""));
|
|
293
|
+
if (forced)
|
|
294
|
+
return finish(forced, toolCalls, cost);
|
|
295
|
+
break;
|
|
296
|
+
}
|
|
297
|
+
return { verdict: "error", confidence: "low", evidence: [], reasoning: "no verdict within the tool budget", toolCalls, cost };
|
|
298
|
+
}
|
|
299
|
+
catch (error) {
|
|
300
|
+
return { verdict: "error", confidence: "low", evidence: [], reasoning: error instanceof Error ? error.message : String(error), toolCalls, cost };
|
|
301
|
+
}
|
|
302
|
+
}
|
|
303
|
+
function finish(verdict, toolCalls, cost) {
|
|
304
|
+
const known = ["confirmed", "not-a-defect", "discarded", "unclear"];
|
|
305
|
+
const value = String(verdict.verdict);
|
|
306
|
+
const evidence = Array.isArray(verdict.evidence)
|
|
307
|
+
? verdict.evidence.map((e) => ({ file: String(e.file ?? ""), lines: String(e.lines ?? ""), note: String(e.note ?? "") }))
|
|
308
|
+
: [];
|
|
309
|
+
return {
|
|
310
|
+
verdict: (known.includes(value) ? value : "unclear"),
|
|
311
|
+
confidence: typeof verdict.confidence === "string" ? verdict.confidence : "low",
|
|
312
|
+
evidence,
|
|
313
|
+
reasoning: typeof verdict.reasoning === "string" ? verdict.reasoning : "",
|
|
314
|
+
toolCalls,
|
|
315
|
+
cost,
|
|
316
|
+
};
|
|
317
|
+
}
|
|
318
|
+
/**
|
|
319
|
+
* Verify the top findings of a collected run against the repository.
|
|
320
|
+
*
|
|
321
|
+
* `maxCost` is a running cap, not a pre-flight estimate: a sync call's cost
|
|
322
|
+
* is only known when it returns, so the pass stops *before* starting the next
|
|
323
|
+
* finding once the cap is reached and reports `partial`. On the default
|
|
324
|
+
* model a finding costs about a cent.
|
|
325
|
+
*/
|
|
326
|
+
export async function runBroadsideVerify(cwd, apiKey, opts = {}) {
|
|
327
|
+
const broadsideDir = broadsideDirFor(cwd);
|
|
328
|
+
const state = await loadBroadsideState(broadsideDir);
|
|
329
|
+
const run = opts.runId ? state.runs.find((candidate) => candidate.id === opts.runId) : state.runs[state.runs.length - 1];
|
|
330
|
+
if (!run) {
|
|
331
|
+
throw new Error(opts.runId ? `No Broad-Side run with id ${opts.runId}.` : "No Broad-Side run recorded. Submit and collect one first.");
|
|
332
|
+
}
|
|
333
|
+
const runDir = join(broadsideDir, run.outputDir);
|
|
334
|
+
const stored = await loadSavedLensResults(runDir, run.lenses);
|
|
335
|
+
const ranked = rankVerifiableFindings(stored);
|
|
336
|
+
if (ranked.length === 0) {
|
|
337
|
+
throw new Error(`Run ${run.id} has no verifiable findings on disk: the defect and security lenses either did not run, are not collected yet, or found nothing. Collect the run first.`);
|
|
338
|
+
}
|
|
339
|
+
const top = Math.max(1, Math.floor(opts.top ?? BROADSIDE_VERIFY_DEFAULT_TOP));
|
|
340
|
+
// The run's model is a `:batch` variant; the sync endpoint wants the base id.
|
|
341
|
+
const model = opts.model ?? syncModelFor(run.model);
|
|
342
|
+
const maxCost = opts.maxCost ?? 0;
|
|
343
|
+
const fetcher = opts.fetcher ?? fetch;
|
|
344
|
+
const reader = await createRepoReader(cwd);
|
|
345
|
+
const selected = ranked.slice(0, top);
|
|
346
|
+
const findings = [];
|
|
347
|
+
let totalCost = 0;
|
|
348
|
+
let stoppedByCost = false;
|
|
349
|
+
for (const [i, candidate] of selected.entries()) {
|
|
350
|
+
if (opts.signal?.aborted)
|
|
351
|
+
break;
|
|
352
|
+
if (maxCost > 0 && totalCost >= maxCost) {
|
|
353
|
+
stoppedByCost = true;
|
|
354
|
+
break;
|
|
355
|
+
}
|
|
356
|
+
const outcome = await verifyFinding(candidate, i + 1, reader, apiKey, model, fetcher);
|
|
357
|
+
const finding = {
|
|
358
|
+
index: i + 1,
|
|
359
|
+
lensId: candidate.lensId,
|
|
360
|
+
customId: candidate.customId,
|
|
361
|
+
severity: candidate.severity,
|
|
362
|
+
title: candidate.title,
|
|
363
|
+
location: candidate.location,
|
|
364
|
+
...outcome,
|
|
365
|
+
};
|
|
366
|
+
findings.push(finding);
|
|
367
|
+
totalCost += outcome.cost;
|
|
368
|
+
opts.onProgress?.(finding);
|
|
369
|
+
}
|
|
370
|
+
const status = findings.length === selected.length ? "completed" : "partial";
|
|
371
|
+
const entry = {
|
|
372
|
+
status,
|
|
373
|
+
model,
|
|
374
|
+
top,
|
|
375
|
+
verified: findings.length,
|
|
376
|
+
confirmed: findings.filter((f) => f.verdict === "confirmed").length,
|
|
377
|
+
cost: totalCost,
|
|
378
|
+
at: new Date().toISOString(),
|
|
379
|
+
};
|
|
380
|
+
const result = {
|
|
381
|
+
runId: run.id,
|
|
382
|
+
outputDir: join(".codecarto", BROADSIDE_DIR, run.id),
|
|
383
|
+
model,
|
|
384
|
+
status,
|
|
385
|
+
candidates: ranked.length,
|
|
386
|
+
findings,
|
|
387
|
+
totalCost,
|
|
388
|
+
...(stoppedByCost && { stoppedByCost: true }),
|
|
389
|
+
};
|
|
390
|
+
await writeFile(join(runDir, "verified.json"), `${JSON.stringify({ ...entry, run_id: run.id, candidates: ranked.length, findings }, null, "\t")}\n`, "utf8");
|
|
391
|
+
await writeFile(join(runDir, "verified.md"), renderVerifiedMarkdown(result), "utf8");
|
|
392
|
+
run.verify = entry;
|
|
393
|
+
await persistBroadsideRunMerging(broadsideDir, run);
|
|
394
|
+
return result;
|
|
395
|
+
}
|
|
396
|
+
const VERDICT_MARK = { confirmed: "✓", "not-a-defect": "–", discarded: "✗", unclear: "?", error: "!" };
|
|
397
|
+
export function renderVerifiedMarkdown(result) {
|
|
398
|
+
const lines = [
|
|
399
|
+
`# Verified findings — run ${result.runId}`,
|
|
400
|
+
"",
|
|
401
|
+
`${result.findings.length} of ${result.candidates} verifiable finding(s) read against the source on \`${result.model}\` (most severe first), $${result.totalCost.toFixed(4)}.` +
|
|
402
|
+
(result.stoppedByCost ? " Stopped by the cost cap before the rest." : ""),
|
|
403
|
+
"",
|
|
404
|
+
"A **confirmed** finding names the input, call site, or sequence that reaches the failure. **not-a-defect** means the claim is",
|
|
405
|
+
"literally true of the code but nothing can reach the failure it describes; **discarded** means the claim is wrong about the",
|
|
406
|
+
"code; **unclear** needs runtime or specification knowledge. Every verdict is still a model's reading — a confirmed finding",
|
|
407
|
+
"is a lead worth a human's next look, not a validated claim.",
|
|
408
|
+
"",
|
|
409
|
+
];
|
|
410
|
+
for (const f of result.findings) {
|
|
411
|
+
lines.push(`## ${VERDICT_MARK[f.verdict] ?? "?"} ${f.index}. [${f.severity}] ${f.title}`, "", `- **verdict**: ${f.verdict} (${f.confidence})`, `- **location**: ${f.location}`, `- **lens**: ${f.lensId} (${f.customId})`);
|
|
412
|
+
if (f.evidence.length > 0)
|
|
413
|
+
lines.push(`- **evidence**: ${f.evidence.map((e) => `${e.file}:${e.lines} — ${e.note}`).join("; ")}`);
|
|
414
|
+
lines.push(`- **reasoning**: ${f.reasoning.replace(/\s+/g, " ").trim()}`, `- **cost**: $${f.cost.toFixed(4)} (${f.toolCalls} tool call(s))`, "");
|
|
415
|
+
}
|
|
416
|
+
return lines.join("\n");
|
|
417
|
+
}
|
|
418
|
+
export function verifyResultText(result) {
|
|
419
|
+
const counts = { confirmed: 0, "not-a-defect": 0, discarded: 0, unclear: 0, error: 0 };
|
|
420
|
+
for (const f of result.findings)
|
|
421
|
+
counts[f.verdict] = (counts[f.verdict] ?? 0) + 1;
|
|
422
|
+
const lines = [
|
|
423
|
+
`Broad-Side verify — run ${result.runId}: ${result.status}`,
|
|
424
|
+
` ${result.findings.length} of ${result.candidates} verifiable finding(s) read on ${result.model} | cost: $${result.totalCost.toFixed(4)}` +
|
|
425
|
+
(result.stoppedByCost ? " (stopped by the cost cap)" : ""),
|
|
426
|
+
` confirmed ${counts.confirmed} · not-a-defect ${counts["not-a-defect"]} · discarded ${counts.discarded} · unclear ${counts.unclear}` + (counts.error ? ` · error ${counts.error}` : ""),
|
|
427
|
+
];
|
|
428
|
+
for (const f of result.findings) {
|
|
429
|
+
lines.push(` ${VERDICT_MARK[f.verdict] ?? "?"} [${f.severity}] ${f.title} @ ${f.location} — ${f.verdict}`);
|
|
430
|
+
}
|
|
431
|
+
lines.push(`Details in ${result.outputDir}/verified.md. A confirmed finding is a lead for a human's next look, not a validated claim.`);
|
|
432
|
+
return lines.join("\n");
|
|
433
|
+
}
|
package/dist/core/broadside.d.ts
CHANGED
|
@@ -244,6 +244,17 @@ export type BroadsideTriageEntry = {
|
|
|
244
244
|
cost?: number;
|
|
245
245
|
error?: string;
|
|
246
246
|
};
|
|
247
|
+
/** Recorded on the run once a verification pass has run (#143); see core/broadside-verify.ts. */
|
|
248
|
+
export type BroadsideVerifyEntry = {
|
|
249
|
+
/** `completed`: every selected finding got a verdict; `partial`: the cost cap or an abort stopped it early. */
|
|
250
|
+
status: "completed" | "partial";
|
|
251
|
+
model: string;
|
|
252
|
+
top: number;
|
|
253
|
+
verified: number;
|
|
254
|
+
confirmed: number;
|
|
255
|
+
cost: number;
|
|
256
|
+
at: string;
|
|
257
|
+
};
|
|
247
258
|
/** The truncation retry pass of one run: one batch per model (#206). */
|
|
248
259
|
export type BroadsideRetryEntry = {
|
|
249
260
|
status: "submitted" | "completed" | "failed";
|
|
@@ -275,6 +286,8 @@ export type BroadsideRun = {
|
|
|
275
286
|
* run cannot both submit it (#322). Absent until a collect claims it.
|
|
276
287
|
*/
|
|
277
288
|
retry?: BroadsideRetryEntry;
|
|
289
|
+
/** The verification pass over the top findings, when one has run (#143). */
|
|
290
|
+
verify?: BroadsideVerifyEntry;
|
|
278
291
|
totalCost?: number;
|
|
279
292
|
pricing?: ModelPricing;
|
|
280
293
|
maxCost?: number;
|
|
@@ -520,9 +533,22 @@ export declare function getLens(lensId: BroadsideLensId): LensDefinition;
|
|
|
520
533
|
export declare function listLenses(): LensDefinition[];
|
|
521
534
|
/** The languages Broad-Side can scan; anything else is refused at submit. */
|
|
522
535
|
export declare const BROADSIDE_LANGUAGES: readonly ["go", "python", "rust", "typescript", "javascript"];
|
|
536
|
+
/**
|
|
537
|
+
* The files a run scans, and where they came from. Contents are always read
|
|
538
|
+
* from the working tree, so the list is the working tree's too: tracked files
|
|
539
|
+
* plus untracked ones git does not ignore, minus files deleted on disk. The
|
|
540
|
+
* list used to come from `git ls-tree HEAD`, so a run mixed the committed
|
|
541
|
+
* file list with uncommitted contents and never saw an untracked file (#248).
|
|
542
|
+
* A target that is not a git repository gets a bounded walk.
|
|
543
|
+
*/
|
|
544
|
+
export declare function listRepoFiles(targetDir: string): Promise<{
|
|
545
|
+
files: string[];
|
|
546
|
+
snapshot: RepoSnapshotSource;
|
|
547
|
+
}>;
|
|
523
548
|
export declare function collectRepoInfo(targetDir: string, opts?: {
|
|
524
549
|
redact?: boolean;
|
|
525
550
|
}): Promise<RepoInfo>;
|
|
551
|
+
export declare function isSlurpable(relPath: string): boolean;
|
|
526
552
|
type CollectedFile = {
|
|
527
553
|
relPath: string;
|
|
528
554
|
moduleName: string;
|
package/dist/core/broadside.js
CHANGED
|
@@ -920,7 +920,7 @@ const SOURCE_SPECS = {
|
|
|
920
920
|
* file list with uncommitted contents and never saw an untracked file (#248).
|
|
921
921
|
* A target that is not a git repository gets a bounded walk.
|
|
922
922
|
*/
|
|
923
|
-
async function listRepoFiles(targetDir) {
|
|
923
|
+
export async function listRepoFiles(targetDir) {
|
|
924
924
|
try {
|
|
925
925
|
const listed = await execFileAsync("git", ["-C", targetDir, "ls-files", "-z", "--cached", "--others", "--exclude-standard"], { maxBuffer: 64 * 1024 * 1024, timeout: GIT_TIMEOUT_MS });
|
|
926
926
|
const deleted = await execFileAsync("git", ["-C", targetDir, "ls-files", "-z", "--deleted"], {
|
|
@@ -1200,7 +1200,7 @@ function matchesAnyGlob(path, globs) {
|
|
|
1200
1200
|
}
|
|
1201
1201
|
return false;
|
|
1202
1202
|
}
|
|
1203
|
-
function isSlurpable(relPath) {
|
|
1203
|
+
export function isSlurpable(relPath) {
|
|
1204
1204
|
// A credential store is never a lens input, whatever its globs say (#252).
|
|
1205
1205
|
if (isSecretFile(relPath))
|
|
1206
1206
|
return false;
|
|
@@ -1624,6 +1624,10 @@ export async function persistBroadsideRunMerging(broadsideDir, run) {
|
|
|
1624
1624
|
run.triage = onDisk.triage;
|
|
1625
1625
|
if (retryEntryRank(onDisk.retry) > retryEntryRank(run.retry))
|
|
1626
1626
|
run.retry = onDisk.retry;
|
|
1627
|
+
// A verification pass another process recorded is never dropped by
|
|
1628
|
+
// a collect that never knew about it; a newer pass replaces an older.
|
|
1629
|
+
if (onDisk.verify && (!run.verify || onDisk.verify.at > run.verify.at))
|
|
1630
|
+
run.verify = onDisk.verify;
|
|
1627
1631
|
for (const [lensId, theirs] of Object.entries(onDisk.batches)) {
|
|
1628
1632
|
if (theirs && batchEntryRank(theirs) > batchEntryRank(run.batches[lensId]))
|
|
1629
1633
|
run.batches[lensId] = theirs;
|
|
@@ -3545,6 +3549,9 @@ export function statusText(state) {
|
|
|
3545
3549
|
}
|
|
3546
3550
|
lines.push(` synthesis: ${run.synthesis.status}`);
|
|
3547
3551
|
lines.push(` triage: ${run.triage?.status ?? "pending"}`);
|
|
3552
|
+
if (run.verify) {
|
|
3553
|
+
lines.push(` verify: ${run.verify.status} — ${run.verify.confirmed} confirmed of ${run.verify.verified} read on ${run.verify.model}, $${run.verify.cost.toFixed(4)}`);
|
|
3554
|
+
}
|
|
3548
3555
|
if (run.totalCost !== undefined)
|
|
3549
3556
|
lines.push(` total cost: $${run.totalCost.toFixed(6)}`);
|
|
3550
3557
|
}
|
package/dist/core/index.d.ts
CHANGED
package/dist/core/index.js
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import { type BroadsideLensId } from "../../core/index.ts";
|
|
2
|
-
export type BroadsideAction = "submit" | "collect" | "status" | "models";
|
|
2
|
+
export type BroadsideAction = "submit" | "collect" | "status" | "models" | "verify";
|
|
3
3
|
export interface BroadsideFlags {
|
|
4
4
|
action: BroadsideAction;
|
|
5
5
|
/** Empty means "the repository's default lens set". */
|
|
@@ -22,11 +22,13 @@ export interface BroadsideFlags {
|
|
|
22
22
|
model?: string;
|
|
23
23
|
/** For submit: per-lens model overrides, layered over config.yaml's (#141). */
|
|
24
24
|
lensModels?: Partial<Record<BroadsideLensId, string>>;
|
|
25
|
+
/** For verify: how many findings to read (#143). */
|
|
26
|
+
top?: number;
|
|
25
27
|
benchmarks: boolean;
|
|
26
28
|
unknown: string[];
|
|
27
29
|
/** Set on an invalid combination. The caller surfaces it as an error. */
|
|
28
30
|
error?: string;
|
|
29
31
|
}
|
|
30
32
|
/** Every token the completer offers, in the order it offers them. */
|
|
31
|
-
export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--model=", "--lens-model=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
|
|
33
|
+
export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "verify", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--model=", "--lens-model=", "--top=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
|
|
32
34
|
export declare function parseBroadsideFlags(args: string): BroadsideFlags;
|
|
@@ -6,6 +6,7 @@
|
|
|
6
6
|
// /codecarto-broadside collect --wait=900
|
|
7
7
|
// /codecarto-broadside status
|
|
8
8
|
// /codecarto-broadside models --benchmarks
|
|
9
|
+
// /codecarto-broadside verify --top=10 → read the top findings against the source
|
|
9
10
|
//
|
|
10
11
|
// Flags mirror the codecarto_broadside tool parameters, with the negative
|
|
11
12
|
// forms spelled out because a slash command has no place to pass `false`:
|
|
@@ -14,8 +15,9 @@
|
|
|
14
15
|
// --max-cost=N --no-retry-truncated
|
|
15
16
|
// --wait=SECONDS --benchmarks (models only)
|
|
16
17
|
// --run=ID (collect only: an older run, as listed by status)
|
|
17
|
-
// --model=ID (submit
|
|
18
|
+
// --model=ID (submit: the run's batch model, as listed by models; verify: the sync model to read with)
|
|
18
19
|
// --lens-model=LENS:ID (submit only, repeatable: one lens on its own model)
|
|
20
|
+
// --top=N (verify only: how many findings to read, most severe first)
|
|
19
21
|
//
|
|
20
22
|
// A model id itself contains a colon (`vendor/name:batch`), so --lens-model
|
|
21
23
|
// splits on the first colon only: `security:deepseek/deepseek-v4-pro:batch`.
|
|
@@ -28,13 +30,14 @@
|
|
|
28
30
|
// The parser never throws. index.ts decides how to surface unknown tokens and
|
|
29
31
|
// invalid combinations, matching parseNextFlags.
|
|
30
32
|
import { BROADSIDE_LENS_IDS } from "../../core/index.js";
|
|
31
|
-
const ACTIONS = new Set(["submit", "collect", "status", "models"]);
|
|
33
|
+
const ACTIONS = new Set(["submit", "collect", "status", "models", "verify"]);
|
|
32
34
|
/** Every token the completer offers, in the order it offers them. */
|
|
33
35
|
export const KNOWN_BROADSIDE_TOKENS = [
|
|
34
36
|
"submit",
|
|
35
37
|
"collect",
|
|
36
38
|
"status",
|
|
37
39
|
"models",
|
|
40
|
+
"verify",
|
|
38
41
|
...BROADSIDE_LENS_IDS,
|
|
39
42
|
"--incremental",
|
|
40
43
|
"--no-incremental",
|
|
@@ -43,6 +46,7 @@ export const KNOWN_BROADSIDE_TOKENS = [
|
|
|
43
46
|
"--run=",
|
|
44
47
|
"--model=",
|
|
45
48
|
"--lens-model=",
|
|
49
|
+
"--top=",
|
|
46
50
|
"--no-synthesis",
|
|
47
51
|
"--no-triage",
|
|
48
52
|
"--no-retry-truncated",
|
|
@@ -122,6 +126,14 @@ export function parseBroadsideFlags(args) {
|
|
|
122
126
|
result.runId = value || undefined;
|
|
123
127
|
continue;
|
|
124
128
|
}
|
|
129
|
+
if (token.startsWith("--top=")) {
|
|
130
|
+
const value = parseNumeric(token, "--top", result);
|
|
131
|
+
if (value !== undefined && (!Number.isInteger(value) || value < 1))
|
|
132
|
+
result.error ??= `--top needs a positive whole number (got "${token.slice("--top=".length)}").`;
|
|
133
|
+
else if (value !== undefined)
|
|
134
|
+
result.top = value;
|
|
135
|
+
continue;
|
|
136
|
+
}
|
|
125
137
|
if (token.startsWith("--model=")) {
|
|
126
138
|
const value = token.slice("--model=".length).trim();
|
|
127
139
|
// An empty value is a mistyped selection, not "use the default":
|
|
@@ -167,11 +179,17 @@ export function parseBroadsideFlags(args) {
|
|
|
167
179
|
if (result.action === "status" && result.waitSeconds !== undefined) {
|
|
168
180
|
result.error ??= "--wait is only meaningful for submit and collect; status reads recorded state.";
|
|
169
181
|
}
|
|
170
|
-
if (result.runId !== undefined && result.action !== "collect") {
|
|
171
|
-
result.error ??= `--run is only meaningful for collect (got action "${result.action}").`;
|
|
182
|
+
if (result.runId !== undefined && result.action !== "collect" && result.action !== "verify") {
|
|
183
|
+
result.error ??= `--run is only meaningful for collect and verify (got action "${result.action}").`;
|
|
184
|
+
}
|
|
185
|
+
if (result.model !== undefined && result.action !== "submit" && result.action !== "verify") {
|
|
186
|
+
result.error ??= `--model is only meaningful for submit and verify (got action "${result.action}").`;
|
|
187
|
+
}
|
|
188
|
+
if (result.top !== undefined && result.action !== "verify") {
|
|
189
|
+
result.error ??= `--top is only meaningful for verify (got action "${result.action}").`;
|
|
172
190
|
}
|
|
173
|
-
if (result.
|
|
174
|
-
result.error ??=
|
|
191
|
+
if (result.action === "verify" && result.waitSeconds !== undefined) {
|
|
192
|
+
result.error ??= "--wait is only meaningful for submit and collect; verify runs to completion.";
|
|
175
193
|
}
|
|
176
194
|
if (result.lensModels !== undefined && result.action !== "submit") {
|
|
177
195
|
result.error ??= `--lens-model is only meaningful for submit (got action "${result.action}").`;
|
|
@@ -10,7 +10,7 @@ import { completeLastToken } from "./completions.js";
|
|
|
10
10
|
import { buildPiGuideMessage } from "./guide-framing.js";
|
|
11
11
|
import { isCtxLive, notifyCtx } from "./notify.js";
|
|
12
12
|
import { phaseCompactionExtension } from "./phase-compaction.js";
|
|
13
|
-
import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, statusText, describeConfigProblems, describeIncrementalFallback, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
|
|
13
|
+
import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, runBroadsideVerify, verifyResultText, statusText, describeConfigProblems, describeIncrementalFallback, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
|
|
14
14
|
import { initLibrary } from "../../core/library.js";
|
|
15
15
|
import { resolveUserConfigPath } from "../../core/orchestrator-config.js";
|
|
16
16
|
const STATUS_WIDGET_ID = "codecarto-widget";
|
|
@@ -995,7 +995,7 @@ export default function codeCartographerExtension(pi) {
|
|
|
995
995
|
},
|
|
996
996
|
});
|
|
997
997
|
pi.registerCommand("codecarto-broadside", {
|
|
998
|
-
description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID] [flags]",
|
|
998
|
+
description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models|verify] [lenses…] [--model=ID] [--lens-model=LENS:ID] [--top=N] [flags]",
|
|
999
999
|
// Completes the token under the cursor, so lens names and flags are
|
|
1000
1000
|
// offered after the action too, and keeps everything typed before it.
|
|
1001
1001
|
getArgumentCompletions: (prefix) => completeLastToken(prefix, KNOWN_BROADSIDE_TOKENS.map((value) => ({ value }))),
|
|
@@ -1181,6 +1181,35 @@ export default function codeCartographerExtension(pi) {
|
|
|
1181
1181
|
}
|
|
1182
1182
|
return;
|
|
1183
1183
|
}
|
|
1184
|
+
if (flags.action === "verify") {
|
|
1185
|
+
// The verification pass (#143): one sync call per finding with
|
|
1186
|
+
// read-only tools; the widget counts verdicts as they land.
|
|
1187
|
+
const tally = { done: 0 };
|
|
1188
|
+
renderProgress("Reading the top findings against the source…");
|
|
1189
|
+
try {
|
|
1190
|
+
const verified = await runBroadsideVerify(ctx.cwd, apiKey, {
|
|
1191
|
+
...(flags.runId && { runId: flags.runId }),
|
|
1192
|
+
...(flags.top !== undefined && { top: flags.top }),
|
|
1193
|
+
...(flags.model && { model: flags.model }),
|
|
1194
|
+
maxCost: flags.maxCost ?? config.maxCost,
|
|
1195
|
+
signal: ctx.signal,
|
|
1196
|
+
onProgress: (finding) => {
|
|
1197
|
+
tally.done += 1;
|
|
1198
|
+
progress.set(`#${finding.index}`, `${finding.verdict} — ${finding.title}`);
|
|
1199
|
+
renderProgress(`Verifying findings… ${tally.done} read`);
|
|
1200
|
+
},
|
|
1201
|
+
});
|
|
1202
|
+
const lines = verifyResultText(verified).split("\n");
|
|
1203
|
+
const confirmed = verified.findings.filter((f) => f.verdict === "confirmed").length;
|
|
1204
|
+
finish(lines, `Broad-Side verify: ${confirmed} confirmed of ${verified.findings.length} read`, verified.status === "completed" ? "info" : "warning");
|
|
1205
|
+
}
|
|
1206
|
+
catch (error) {
|
|
1207
|
+
if (ctx.hasUI)
|
|
1208
|
+
ctx.ui.setWidget(BROADSIDE_WIDGET_ID, undefined);
|
|
1209
|
+
notifyCtx(ctx, `Broad-Side verify failed: ${error instanceof Error ? error.message : String(error)}`, "error");
|
|
1210
|
+
}
|
|
1211
|
+
return;
|
|
1212
|
+
}
|
|
1184
1213
|
// action === "collect"
|
|
1185
1214
|
renderProgress("Polling batches…");
|
|
1186
1215
|
try {
|
|
@@ -233,7 +233,7 @@ export declare function handleAmend(args: {
|
|
|
233
233
|
}>;
|
|
234
234
|
export declare function handleBroadside(args: {
|
|
235
235
|
cwd: string;
|
|
236
|
-
action: "submit" | "collect" | "status" | "models";
|
|
236
|
+
action: "submit" | "collect" | "status" | "models" | "verify";
|
|
237
237
|
lenses?: string[];
|
|
238
238
|
api_key?: string;
|
|
239
239
|
wait_seconds?: number;
|
|
@@ -247,6 +247,7 @@ export declare function handleBroadside(args: {
|
|
|
247
247
|
incremental?: boolean;
|
|
248
248
|
model?: string;
|
|
249
249
|
lens_models?: Record<string, string>;
|
|
250
|
+
top?: number;
|
|
250
251
|
}): Promise<{
|
|
251
252
|
content: {
|
|
252
253
|
type: "text";
|
|
@@ -17,7 +17,7 @@ import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"
|
|
|
17
17
|
import { CallToolRequestSchema, ErrorCode, ListToolsRequestSchema, McpError, } from "@modelcontextprotocol/sdk/types.js";
|
|
18
18
|
import { mkdir, readFile, readdir, rename, writeFile } from "node:fs/promises";
|
|
19
19
|
import { basename, isAbsolute, join } from "node:path";
|
|
20
|
-
import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, seedOrchestratorFiles, statusLineWriter, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
|
|
20
|
+
import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, runBroadsideVerify, verifyResultText, seedOrchestratorFiles, statusLineWriter, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
|
|
21
21
|
import { applyAmendment } from "../core/amendment.js";
|
|
22
22
|
import { appendUsageRun } from "../core/usage.js";
|
|
23
23
|
import { initLibrary } from "../core/library.js";
|
|
@@ -1037,8 +1037,8 @@ function resolveBroadsideApiKey(explicit, config) {
|
|
|
1037
1037
|
export async function handleBroadside(args) {
|
|
1038
1038
|
const cwd = await validateCwd(args.cwd);
|
|
1039
1039
|
const action = args.action ?? "submit";
|
|
1040
|
-
if (!["submit", "collect", "status", "models"].includes(action)) {
|
|
1041
|
-
throw new McpError(ErrorCode.InvalidParams, `Unknown action: ${action}. Valid actions: submit, collect, status, models.`);
|
|
1040
|
+
if (!["submit", "collect", "status", "models", "verify"].includes(action)) {
|
|
1041
|
+
throw new McpError(ErrorCode.InvalidParams, `Unknown action: ${action}. Valid actions: submit, collect, status, models, verify.`);
|
|
1042
1042
|
}
|
|
1043
1043
|
// A config.yaml that exists but cannot be read refuses every action that
|
|
1044
1044
|
// would act on it (#232); status only reads recorded runs, so it answers
|
|
@@ -1173,8 +1173,40 @@ export async function handleBroadside(args) {
|
|
|
1173
1173
|
maxCost: result.maxCost,
|
|
1174
1174
|
});
|
|
1175
1175
|
}
|
|
1176
|
-
// action === "collect"
|
|
1177
1176
|
const runId = typeof args.run_id === "string" && args.run_id.trim() ? args.run_id.trim() : undefined;
|
|
1177
|
+
if (action === "verify") {
|
|
1178
|
+
// The verification pass (#143): one sync call per finding with read-only
|
|
1179
|
+
// tools, most severe first. `max_cost` is a running cap here — a sync
|
|
1180
|
+
// call's cost is known only when it returns — so the pass stops before
|
|
1181
|
+
// the next finding once reached; absent, config.yaml's cap applies.
|
|
1182
|
+
if (args.top !== undefined && !(typeof args.top === "number" && Number.isInteger(args.top) && args.top >= 1)) {
|
|
1183
|
+
throw new McpError(ErrorCode.InvalidParams, "top must be a positive integer.");
|
|
1184
|
+
}
|
|
1185
|
+
if (args.model !== undefined && !(typeof args.model === "string" && args.model.trim())) {
|
|
1186
|
+
throw new McpError(ErrorCode.InvalidParams, "model must be a non-empty OpenRouter model id.");
|
|
1187
|
+
}
|
|
1188
|
+
const maxCost = typeof args.max_cost === "number" && args.max_cost >= 0 ? args.max_cost : config.maxCost;
|
|
1189
|
+
const verified = await runBroadsideVerify(cwd, apiKey, {
|
|
1190
|
+
...(runId && { runId }),
|
|
1191
|
+
...(args.top !== undefined && { top: args.top }),
|
|
1192
|
+
...(typeof args.model === "string" && args.model.trim() && { model: args.model.trim() }),
|
|
1193
|
+
maxCost,
|
|
1194
|
+
signal: serverLifetime?.signal,
|
|
1195
|
+
}).catch((error) => {
|
|
1196
|
+
throw new McpError(ErrorCode.InvalidRequest, error instanceof Error ? error.message : String(error));
|
|
1197
|
+
});
|
|
1198
|
+
return textResult(verifyResultText(verified), {
|
|
1199
|
+
runId: verified.runId,
|
|
1200
|
+
outputDir: verified.outputDir,
|
|
1201
|
+
status: verified.status,
|
|
1202
|
+
model: verified.model,
|
|
1203
|
+
candidates: verified.candidates,
|
|
1204
|
+
totalCost: verified.totalCost,
|
|
1205
|
+
...(verified.stoppedByCost && { stoppedByCost: true }),
|
|
1206
|
+
findings: verified.findings,
|
|
1207
|
+
});
|
|
1208
|
+
}
|
|
1209
|
+
// action === "collect"
|
|
1178
1210
|
const collect = await runBroadsideCollect(cwd, apiKey, {
|
|
1179
1211
|
waitMs,
|
|
1180
1212
|
includeSynthesis,
|
|
@@ -1502,8 +1534,8 @@ const TOOLS = [
|
|
|
1502
1534
|
cwd: { type: "string", description: "Absolute path to the target repository." },
|
|
1503
1535
|
action: {
|
|
1504
1536
|
type: "string",
|
|
1505
|
-
enum: ["submit", "collect", "status", "models"],
|
|
1506
|
-
description: "submit fires all lens batches and returns batch ids; collect polls submitted batches, saves results, and optionally runs the synthesis pass; status shows recorded runs; models lists batch-capable models with pricing and capabilities.",
|
|
1537
|
+
enum: ["submit", "collect", "status", "models", "verify"],
|
|
1538
|
+
description: "submit fires all lens batches and returns batch ids; collect polls submitted batches, saves results, and optionally runs the synthesis pass; status shows recorded runs; models lists batch-capable models with pricing and capabilities; verify reads a collected run's top defect and security findings against the repository with read-only tools (one sync-priced call each, about a cent on the default model) and writes verified.md/verified.json beside triage.md with a verdict per finding: confirmed (a reachable failure, with the trigger), not-a-defect, discarded, or unclear.",
|
|
1507
1539
|
},
|
|
1508
1540
|
lenses: {
|
|
1509
1541
|
type: "array",
|
|
@@ -1516,7 +1548,11 @@ const TOOLS = [
|
|
|
1516
1548
|
},
|
|
1517
1549
|
run_id: {
|
|
1518
1550
|
type: "string",
|
|
1519
|
-
description: "For collect: the run to
|
|
1551
|
+
description: "For collect and verify: the run to act on, as listed by the status action. Defaults to the most recent run; pass this to collect an older run that is still in flight after a newer submit.",
|
|
1552
|
+
},
|
|
1553
|
+
top: {
|
|
1554
|
+
type: "integer",
|
|
1555
|
+
description: "For verify: how many findings to read, most severe first (default 10). Each costs one sync call; max_cost caps the pass as a running total.",
|
|
1520
1556
|
},
|
|
1521
1557
|
wait_seconds: {
|
|
1522
1558
|
type: "number",
|
|
@@ -1536,7 +1572,7 @@ const TOOLS = [
|
|
|
1536
1572
|
},
|
|
1537
1573
|
max_cost: {
|
|
1538
1574
|
type: "number",
|
|
1539
|
-
description: "Approximate run expense limit in USD. The submit action estimates the run cost from slice sizes and the configured model's per-token pricing (live OpenRouter lookup, cached 24h) and refuses to submit when the estimate exceeds the limit unless force is true. Falls back to max_cost in .codecarto/broadside/config.yaml, whose default is $1.00; pass 0 for no limit.",
|
|
1575
|
+
description: "Approximate run expense limit in USD. The submit action estimates the run cost from slice sizes and the configured model's per-token pricing (live OpenRouter lookup, cached 24h) and refuses to submit when the estimate exceeds the limit unless force is true. For verify it is a running cap: the pass stops before the next finding once the calls so far have reached it. Falls back to max_cost in .codecarto/broadside/config.yaml, whose default is $1.00; pass 0 for no limit.",
|
|
1540
1576
|
},
|
|
1541
1577
|
force: {
|
|
1542
1578
|
type: "boolean",
|
|
@@ -1552,7 +1588,7 @@ const TOOLS = [
|
|
|
1552
1588
|
},
|
|
1553
1589
|
model: {
|
|
1554
1590
|
type: "string",
|
|
1555
|
-
description: "For submit: the OpenRouter batch model for this run (an id ending in :batch, as listed by action 'models'). Falls back to model in .codecarto/broadside/config.yaml, then the shipped default. Pre-flighted like the configured model: priced from the live catalog, refused without structured-output support, clamped to its completion ceiling. The models listing is advisory — some catalog ids have no batch endpoint and are refused at submit, at no cost; the listing tags ids this repository has already seen accepted or refused.",
|
|
1591
|
+
description: "For verify: the sync (non-batch) OpenRouter model to read with; defaults to the run's model without its :batch suffix. For submit: the OpenRouter batch model for this run (an id ending in :batch, as listed by action 'models'). Falls back to model in .codecarto/broadside/config.yaml, then the shipped default. Pre-flighted like the configured model: priced from the live catalog, refused without structured-output support, clamped to its completion ceiling. The models listing is advisory — some catalog ids have no batch endpoint and are refused at submit, at no cost; the listing tags ids this repository has already seen accepted or refused.",
|
|
1556
1592
|
},
|
|
1557
1593
|
lens_models: {
|
|
1558
1594
|
type: "object",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "codecartographer-pi",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.24.0",
|
|
4
4
|
"mcpName": "io.github.HuginnIndustries/codecartographer",
|
|
5
5
|
"description": "Turn an unfamiliar codebase into a validated reimplementation spec, then synthesize confirmed specs and a product vision into a traceable plan.",
|
|
6
6
|
"type": "module",
|