codecartographer-pi 0.21.0 → 0.22.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codecarto/broadside/SKILL.md +16 -6
- package/.codecarto/broadside/config.yaml +49 -20
- package/.codecarto/workflow/scaffold-version.yaml +1 -1
- package/README.md +2 -2
- package/agent-skill/codecartographer/references/broadside.md +10 -0
- package/dist/core/broadside.d.ts +91 -20
- package/dist/core/broadside.js +312 -59
- package/dist/extensions/codecarto/broadside-flags.d.ts +5 -1
- package/dist/extensions/codecarto/broadside-flags.js +39 -0
- package/dist/extensions/codecarto/index.js +40 -12
- package/dist/mcp-server/server.d.ts +2 -0
- package/dist/mcp-server/server.js +46 -5
- package/package.json +1 -1
|
@@ -95,11 +95,14 @@ OpenRouter advertises a `:batch` variant for, and many of those variants do not
|
|
|
95
95
|
exist — submitting one returns `does not have a :batch endpoint`, with nothing in
|
|
96
96
|
the catalog to distinguish it beforehand. A rejected batch costs nothing, so
|
|
97
97
|
probe a candidate on a single lens first. And reasoning competes with the answer for
|
|
98
|
-
`max_tokens`: Broad-Side
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
98
|
+
`max_tokens`: Broad-Side asks every model for low reasoning effort so the output
|
|
99
|
+
budget stays with the JSON, which is the split the cost estimate already
|
|
100
|
+
assumes. An effort level rather than a token cap, because Gemini 3.x ignores a
|
|
101
|
+
cap (measured: 11,518 thinking tokens under a 5,800 cap) and honours the level;
|
|
102
|
+
a level rather than off, because some endpoints refuse to be switched off. A
|
|
103
|
+
slice that still truncates is retried once with a doubled budget at low effort.
|
|
104
|
+
Override with `reasoning:` in `config.yaml` only alongside a raised output
|
|
105
|
+
budget, and set `effort` or `max_tokens`, never both.
|
|
103
106
|
|
|
104
107
|
Lenses do not all have to run on the same model. `lens_models` in `config.yaml`
|
|
105
108
|
routes individual lenses to their own batch model — the usual reason being that
|
|
@@ -108,7 +111,14 @@ architecture map. Each override is priced, capability-checked, and clamped like
|
|
|
108
111
|
the default, the submit estimate breaks cost out per lens, and `run-meta.json`
|
|
109
112
|
records which lens ran on what. No stronger default is shipped: which model is
|
|
110
113
|
worth the money depends on the repository and the budget, so compare with the
|
|
111
|
-
`models` action and decide
|
|
114
|
+
`models` action and decide — for one run with the `model` and `lens_models`
|
|
115
|
+
parameters (Pi: `--model=ID`, `--lens-model=LENS:ID`), or here for the
|
|
116
|
+
repository. The `models` listing is advisory: OpenRouter's catalog returns a
|
|
117
|
+
`:batch` id for some models its Batch API refuses (`does not have a :batch
|
|
118
|
+
endpoint`), at no cost, and nothing in the catalog tells them apart. The
|
|
119
|
+
listing tags the ids this repository's own submits have seen accepted or
|
|
120
|
+
refused (`broadside/batch-endpoints.json`), and a refused lens says why in the
|
|
121
|
+
submit report.
|
|
112
122
|
|
|
113
123
|
Every run knob — `incremental`, `retry_truncated`, `include_synthesis`,
|
|
114
124
|
`include_triage`, `wait_seconds` — also has a repository default under the same
|
|
@@ -54,39 +54,68 @@
|
|
|
54
54
|
# Nothing in the catalog distinguishes them. Every Anthropic and OpenAI
|
|
55
55
|
# batch id tried so far is rejected this way; Google and DeepSeek work.
|
|
56
56
|
# A rejected batch costs nothing, so probe a candidate on one lens before
|
|
57
|
-
# relying on it
|
|
57
|
+
# relying on it: `codecarto_broadside {action: "submit", lenses:
|
|
58
|
+
# ["architecture"], model: "<id>"}` (Pi: `/codecarto-broadside architecture
|
|
59
|
+
# --model=<id>`). Submits remember the answer in batch-endpoints.json next
|
|
60
|
+
# to this file, and the `models` listing tags each id accordingly.
|
|
58
61
|
# 2. A reasoning-capable model spends its output budget thinking, and the
|
|
59
62
|
# thinking is billed at the full output rate. See `reasoning:` below.
|
|
60
63
|
#
|
|
64
|
+
# The same two keys are submit parameters for one run — `model` and
|
|
65
|
+
# `lens_models` on codecarto_broadside, `--model=ID` and `--lens-model=LENS:ID`
|
|
66
|
+
# on /codecarto-broadside — and a lens set both there and here takes the
|
|
67
|
+
# parameter's.
|
|
68
|
+
#
|
|
61
69
|
# lens_models:
|
|
62
70
|
# security: deepseek/deepseek-v4-pro-0813:batch
|
|
63
71
|
# defect: deepseek/deepseek-v4-pro-0813:batch
|
|
64
72
|
|
|
65
73
|
# Reasoning control, sent on every lens request.
|
|
66
74
|
#
|
|
67
|
-
# By default Broad-Side
|
|
68
|
-
# leaving the
|
|
69
|
-
# the cost estimate already assumes.
|
|
70
|
-
#
|
|
71
|
-
# The
|
|
72
|
-
# One measured run spent 5,758 of a 6,000-token budget thinking
|
|
73
|
-
# tokens for the JSON, which truncated mid-structure on 11 of 13
|
|
74
|
-
# tokens bill at the full output rate, so it paid for ~6,000
|
|
75
|
-
# slice to receive ~230 usable ones. The shipped default
|
|
76
|
-
# thing less consistently — reasoning from 0 to 5,757
|
|
77
|
-
# three of them cut off — so this is not something
|
|
75
|
+
# By default Broad-Side asks every model for low reasoning effort
|
|
76
|
+
# (`effort: low`), leaving the output budget for the answer — which is the
|
|
77
|
+
# split the cost estimate already assumes.
|
|
78
|
+
#
|
|
79
|
+
# The control exists because reasoning competes with the answer for
|
|
80
|
+
# `max_tokens`. One measured run spent 5,758 of a 6,000-token budget thinking
|
|
81
|
+
# and left ~230 tokens for the JSON, which truncated mid-structure on 11 of 13
|
|
82
|
+
# slices; those tokens bill at the full output rate, so it paid for ~6,000
|
|
83
|
+
# output tokens per slice to receive ~230 usable ones. The shipped default
|
|
84
|
+
# model does the same thing less consistently — reasoning from 0 to 5,757
|
|
85
|
+
# tokens across 13 slices, three of them cut off — so this is not something
|
|
86
|
+
# only exotic models do.
|
|
87
|
+
#
|
|
88
|
+
# It is an effort level rather than a token cap because a cap is not honoured
|
|
89
|
+
# everywhere. Gemini 3.x models take a thinking *level*, not a budget: under
|
|
90
|
+
# `max_tokens: 5800`, google/gemini-3.8-flash:batch reasoned 5,218 tokens on
|
|
91
|
+
# the first pass and 11,518 on the doubled-budget retry — the cap changed
|
|
92
|
+
# nothing, both results truncated, and the retry cost twice the original for
|
|
93
|
+
# no JSON. The same lens at `effort: low` reasoned 0 tokens, returned valid
|
|
94
|
+
# JSON, and cost a twelfth as much. `effort` is what OpenRouter can translate
|
|
95
|
+
# for every provider; a token cap reaches only the providers that take one.
|
|
96
|
+
#
|
|
97
|
+
# And it is a level rather than an off switch on purpose. Some endpoints refuse
|
|
98
|
+
# to be switched off: `google/gemini-3.8-flash:batch` rejects the entire batch
|
|
99
|
+
# with "Reasoning is mandatory for this endpoint and cannot be disabled", which
|
|
100
|
+
# turns a partial result into none at all. Low effort works either way.
|
|
101
|
+
#
|
|
102
|
+
# A slice that still truncates is re-submitted once with a doubled `max_tokens`
|
|
103
|
+
# and, whatever this block says, low effort — the cutoff is usually thinking.
|
|
104
|
+
#
|
|
105
|
+
# Set `effort` OR `max_tokens`, not both: OpenRouter refuses a request that
|
|
106
|
+
# carries both ("Only one of reasoning.effort and reasoning.max_tokens can be
|
|
107
|
+
# specified") — per request, after the batch is accepted, so every lens fails
|
|
108
|
+
# at $0. Submit refuses a config.yaml that sets both. Raise the effort only
|
|
109
|
+
# together with a raised lens output budget, or the JSON truncates exactly as
|
|
110
|
+
# above.
|
|
78
111
|
#
|
|
79
|
-
#
|
|
80
|
-
#
|
|
81
|
-
# "Reasoning is mandatory for this endpoint and cannot be disabled", which turns
|
|
82
|
-
# a partial result into none at all. Capping works either way.
|
|
112
|
+
# reasoning:
|
|
113
|
+
# effort: low # minimal | low | medium | high
|
|
83
114
|
#
|
|
84
|
-
#
|
|
85
|
-
#
|
|
115
|
+
# reasoning:
|
|
116
|
+
# max_tokens: 2000 # a thinking budget, where the provider honours one
|
|
86
117
|
#
|
|
87
118
|
# reasoning:
|
|
88
|
-
# effort: low # minimal | low | medium | high
|
|
89
|
-
# max_tokens: 2000 # or set the thinking budget directly
|
|
90
119
|
# enabled: false # only where the provider allows it
|
|
91
120
|
|
|
92
121
|
# Approximate run expense limit in USD. The default is 1.00, and it applies
|
package/README.md
CHANGED
|
@@ -398,9 +398,9 @@ codecarto_broadside {cwd, action: "collect"} # poll, save, synt
|
|
|
398
398
|
|
|
399
399
|
Submit and collect are separate because batch jobs routinely take tens of minutes; collect is resumable and picks up whatever is still in flight (`wait_seconds: 0`, the default, polls once and returns). Submit prices the run from the collected file sizes against the model's live per-token pricing (cached 24h) and refuses when the estimate exceeds `max_cost` — $1.00 unless the config or the call sets another value, `0` for no limit — unless `force: true` is passed — a pre-flight estimate, not a runtime stop. Actual spend lands in each run's `run-meta.json`.
|
|
400
400
|
|
|
401
|
-
Repository defaults live in `.codecarto/broadside/config.yaml` (`model`, `api_key`, `default_lenses`, `max_cost`, `pricing` overrides, `lens_models`, `incremental`, `retry_truncated`, `include_synthesis`, `include_triage`, `wait_seconds`); an explicit tool parameter always wins. `lens_models` routes individual lenses to their own batch model — a stronger model changes security and defect findings far more than it changes an architecture map — and each override is priced, capability-checked, and clamped exactly like the default, with the estimate broken out per lens so a mixed-model run cannot be approved without seeing which lens costs what. CodeCartographer ships no stronger default: which model earns its price depends on your repository and budget, so compare candidates with the `models` action and choose. `codecarto_skill {cwd, name: "broadside"}` returns the reading guide for a completed run, and unlike post-pipeline skills it is not gated on a finished pipeline.
|
|
401
|
+
Repository defaults live in `.codecarto/broadside/config.yaml` (`model`, `api_key`, `default_lenses`, `max_cost`, `pricing` overrides, `lens_models`, `incremental`, `retry_truncated`, `include_synthesis`, `include_triage`, `wait_seconds`); an explicit tool parameter always wins. `lens_models` routes individual lenses to their own batch model — a stronger model changes security and defect findings far more than it changes an architecture map — and each override is priced, capability-checked, and clamped exactly like the default, with the estimate broken out per lens so a mixed-model run cannot be approved without seeing which lens costs what. CodeCartographer ships no stronger default: which model earns its price depends on your repository and budget, so compare candidates with the `models` action and choose — for one run with the `model` and `lens_models` parameters (Pi: `--model=ID`, `--lens-model=LENS:ID`), or for the repository in `config.yaml`. The `models` listing is advisory: OpenRouter's catalog returns a `:batch` id for some models its Batch API then refuses (`does not have a :batch endpoint`), at no cost, and nothing in the catalog tells them apart — so the listing tags the ids this repository's own submits have seen accepted or refused, and a refused lens says why in the submit report. `codecarto_skill {cwd, name: "broadside"}` returns the reading guide for a completed run, and unlike post-pipeline skills it is not gated on a finished pipeline.
|
|
402
402
|
|
|
403
|
-
On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models] [lenses…]`, with tab-completion for actions
|
|
403
|
+
On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID]`, with tab-completion for actions, lens names, and flags and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
|
|
404
404
|
|
|
405
405
|
Broad-Side needs runtime code, so firing a run is an executable-surface feature: Pi and MCP have it, the pure drop-in template does not (it carries only the reading guide).
|
|
406
406
|
|
|
@@ -82,6 +82,16 @@ Two more economies worth knowing:
|
|
|
82
82
|
conventions is usually a better trade than raising the model for everything.
|
|
83
83
|
Overrides are priced and capability-checked individually, and the estimate
|
|
84
84
|
breaks cost out per lens.
|
|
85
|
+
- `model` and `lens_models` are also submit parameters (Pi: `--model=ID`,
|
|
86
|
+
`--lens-model=LENS:ID`), for one run without editing the file. Choose from
|
|
87
|
+
`action: "models"`, and read that listing as advisory: OpenRouter's catalog
|
|
88
|
+
returns a `:batch` id for some models its Batch API then refuses (`does not
|
|
89
|
+
have a :batch endpoint`) — free, reported on the lens with the reason, and
|
|
90
|
+
remembered, so the listing tags ids this repository has seen accepted or
|
|
91
|
+
refused. Probe an untried model on one lens before a six-lens run. A
|
|
92
|
+
`job-submission-count` refusal is the account's concurrent-job quota (one
|
|
93
|
+
job per lens fills it fast across runs); collect or wait out what is in
|
|
94
|
+
flight, then re-submit.
|
|
85
95
|
|
|
86
96
|
## Reading a run
|
|
87
97
|
|
package/dist/core/broadside.d.ts
CHANGED
|
@@ -11,6 +11,13 @@ export declare const BROADSIDE_OUTPUT_PRICE_PER_M = 1.875;
|
|
|
11
11
|
export declare const BROADSIDE_MODELS_URL = "https://openrouter.ai/api/v1/models";
|
|
12
12
|
export declare const BROADSIDE_BENCHMARKS_URL = "https://openrouter.ai/api/v1/benchmarks";
|
|
13
13
|
export declare const BROADSIDE_CATALOG_CACHE_FILE = "model-catalog.json";
|
|
14
|
+
/**
|
|
15
|
+
* What this repository's own submits learned about batch endpoints: which
|
|
16
|
+
* `:batch` ids OpenRouter accepted a job for and which it refused with
|
|
17
|
+
* "does not have a :batch endpoint". The catalog cannot tell the two apart
|
|
18
|
+
* (#141), so the `models` action annotates its rows from this file.
|
|
19
|
+
*/
|
|
20
|
+
export declare const BROADSIDE_ENDPOINTS_FILE = "batch-endpoints.json";
|
|
14
21
|
export declare const BROADSIDE_CATALOG_CACHE_TTL_MS: number;
|
|
15
22
|
export declare const BROADSIDE_LENS_IDS: readonly ["architecture", "api", "security", "defect", "conventions", "porting"];
|
|
16
23
|
export type BroadsideLensId = (typeof BROADSIDE_LENS_IDS)[number];
|
|
@@ -130,31 +137,43 @@ export type BroadsideReasoning = {
|
|
|
130
137
|
max_tokens?: number;
|
|
131
138
|
};
|
|
132
139
|
/**
|
|
133
|
-
* The
|
|
140
|
+
* The reasoning control every lens request carries: low effort.
|
|
141
|
+
*
|
|
142
|
+
* It used to be a token cap — `max_tokens` at a quarter of the lens's output
|
|
143
|
+
* budget, so three quarters stayed for the answer. Measured live on
|
|
144
|
+
* `google/gemini-3.8-flash:batch` (0.22.0 verification, defect lens, cap
|
|
145
|
+
* 5,800 of a 6,000 budget): the model reasoned 5,218 tokens on the first
|
|
146
|
+
* pass and **11,518 under the same cap** on the doubled-budget retry —
|
|
147
|
+
* thinking scaled with `max_tokens` and the cap changed nothing, both
|
|
148
|
+
* results truncated, and the retry cost twice the original for no JSON.
|
|
149
|
+
* The same lens with `effort: "low"` reasoned 0 tokens, finished with
|
|
150
|
+
* `stop`, returned valid JSON, and cost a twelfth as much. Gemini 3.x
|
|
151
|
+
* models take a thinking *level*, not a budget, and OpenRouter forwards a
|
|
152
|
+
* `max_tokens` cap to them as nothing at all; `effort` is what it can
|
|
153
|
+
* translate for every provider (a level where the provider has levels, a
|
|
154
|
+
* fraction of the budget where it takes a budget). So the default asks for
|
|
155
|
+
* little thinking in the one vocabulary that reaches everyone.
|
|
134
156
|
*
|
|
135
|
-
*
|
|
136
|
-
*
|
|
137
|
-
*
|
|
157
|
+
* Deliberately not `enabled: false`: `google/gemini-3.8-flash:batch` refuses
|
|
158
|
+
* the whole batch with *"Reasoning is mandatory for this endpoint and cannot
|
|
159
|
+
* be disabled"*, turning a partial result into none at all. Low effort works
|
|
160
|
+
* whether or not a provider allows reasoning to be switched off.
|
|
138
161
|
*/
|
|
139
|
-
export declare const
|
|
140
|
-
|
|
162
|
+
export declare const BROADSIDE_DEFAULT_REASONING: Readonly<BroadsideReasoning>;
|
|
163
|
+
/** The reasoning control a lens request carries when config.yaml sets none. */
|
|
164
|
+
export declare function defaultReasoningFor(): BroadsideReasoning;
|
|
141
165
|
/**
|
|
142
|
-
*
|
|
166
|
+
* The reasoning control a truncated slice is re-submitted with.
|
|
143
167
|
*
|
|
144
|
-
*
|
|
145
|
-
*
|
|
146
|
-
*
|
|
147
|
-
*
|
|
148
|
-
*
|
|
149
|
-
*
|
|
150
|
-
*
|
|
151
|
-
* reasoning, leaving ~230 tokens for JSON that truncated mid-structure — and
|
|
152
|
-
* those tokens bill at the full output rate. The shipped default model does the
|
|
153
|
-
* same thing less consistently (reasoning tokens from 0 to 5,757 across 13
|
|
154
|
-
* slices, three of them cut off at `finish_reason: length`), so this is not a
|
|
155
|
-
* multi-model concern.
|
|
168
|
+
* A truncation on a reasoning-capable model is usually thinking that ate the
|
|
169
|
+
* answer's budget, and doubling `max_tokens` doubles the thinking where the
|
|
170
|
+
* provider ignores a token cap (see {@link BROADSIDE_DEFAULT_REASONING}). The
|
|
171
|
+
* retry therefore asks for low effort as well, replacing a `max_tokens` cap
|
|
172
|
+
* (OpenRouter refuses a request carrying both) and lowering a higher effort.
|
|
173
|
+
* An explicit `enabled: false` and an effort already at or below low are left
|
|
174
|
+
* as they are.
|
|
156
175
|
*/
|
|
157
|
-
export declare function
|
|
176
|
+
export declare function retryReasoningFor(original: BroadsideReasoning | undefined): BroadsideReasoning;
|
|
158
177
|
export type BatchRequest = {
|
|
159
178
|
custom_id: string;
|
|
160
179
|
body: {
|
|
@@ -182,6 +201,8 @@ export type BroadsideBatchEntry = {
|
|
|
182
201
|
cost?: number;
|
|
183
202
|
resultCount?: number;
|
|
184
203
|
error?: unknown;
|
|
204
|
+
/** Why a `skipped` lens had nothing to submit: the globs that matched no file. */
|
|
205
|
+
reason?: string;
|
|
185
206
|
/** Set when this lens used a model other than the run default. */
|
|
186
207
|
model?: string;
|
|
187
208
|
/** The completion ceiling of this lens's model; bounds the truncation retry. */
|
|
@@ -539,6 +560,26 @@ export declare function loadBroadsideConfig(broadsideDir: string): Promise<Broad
|
|
|
539
560
|
export declare function defaultBroadsideConfig(): BroadsideConfig;
|
|
540
561
|
/** The catalog cache schema this build writes; a file from another is not read. */
|
|
541
562
|
export declare const BROADSIDE_CATALOG_CACHE_SCHEMA = 3;
|
|
563
|
+
/** One model's most recent submit outcome, as remembered in {@link BROADSIDE_ENDPOINTS_FILE}. */
|
|
564
|
+
export type BatchEndpointRecord = {
|
|
565
|
+
status: "accepted" | "rejected";
|
|
566
|
+
/** ISO timestamp of the submit that produced this record. */
|
|
567
|
+
at: string;
|
|
568
|
+
/** The provider's refusal, for a rejected endpoint. */
|
|
569
|
+
error?: string;
|
|
570
|
+
};
|
|
571
|
+
export declare function readBatchEndpoints(broadsideDir: string): Promise<Record<string, BatchEndpointRecord>>;
|
|
572
|
+
/**
|
|
573
|
+
* Remember what a submit learned about each model it posted to. An accepted
|
|
574
|
+
* job proves the endpoint exists; a "does not have a :batch endpoint"
|
|
575
|
+
* refusal proves it does not. Any other rejection (quota, malformed request,
|
|
576
|
+
* auth) says nothing about the endpoint and leaves the record alone.
|
|
577
|
+
*/
|
|
578
|
+
export declare function recordBatchEndpoints(broadsideDir: string, outcomes: Array<{
|
|
579
|
+
model: string;
|
|
580
|
+
batchId: string;
|
|
581
|
+
error?: unknown;
|
|
582
|
+
}>): Promise<void>;
|
|
542
583
|
export declare function builtInCatalogEntry(model: string): CatalogEntry | null;
|
|
543
584
|
export declare function builtInPricing(model: string): ModelPricing | null;
|
|
544
585
|
export declare function resolveCatalogEntry(broadsideDir: string, config: BroadsideConfig, model: string, apiKey: string, fetcher?: FetchLike): Promise<BroadsideCatalogResult>;
|
|
@@ -552,6 +593,8 @@ export declare function listBatchModels(broadsideDir: string, config: BroadsideC
|
|
|
552
593
|
source: string;
|
|
553
594
|
benchmarks: CodingBenchmarks | null;
|
|
554
595
|
defaultModel: string;
|
|
596
|
+
/** This repository's remembered submit outcomes per model, from {@link BROADSIDE_ENDPOINTS_FILE}. */
|
|
597
|
+
endpoints: Record<string, BatchEndpointRecord>;
|
|
555
598
|
}>;
|
|
556
599
|
export type FetchLike = (url: string, init: Record<string, unknown>) => Promise<Response>;
|
|
557
600
|
export declare function submitBatch(batchRequests: BatchRequest[], apiKey: string, fetcher?: FetchLike, model?: string): Promise<{
|
|
@@ -602,6 +645,12 @@ export declare function runBroadsideSubmit(cwd: string, apiKey: string, opts?: {
|
|
|
602
645
|
lenses?: BroadsideLensId[];
|
|
603
646
|
fetcher?: FetchLike;
|
|
604
647
|
model?: string;
|
|
648
|
+
/**
|
|
649
|
+
* Per-lens model overrides for this run, layered over config.yaml's
|
|
650
|
+
* `lens_models`: a lens named here runs on this model, a lens named only
|
|
651
|
+
* in the file runs on the file's, and the rest run on `model` (#141).
|
|
652
|
+
*/
|
|
653
|
+
lensModels?: Partial<Record<BroadsideLensId, string>>;
|
|
605
654
|
/** Approximate run expense limit in USD; 0 means no limit. */
|
|
606
655
|
maxCost?: number;
|
|
607
656
|
/** Submit even when the estimate exceeds maxCost. */
|
|
@@ -669,11 +718,33 @@ export declare function runBroadsideStatus(cwd: string): Promise<{
|
|
|
669
718
|
state: BroadsideStateFile;
|
|
670
719
|
}>;
|
|
671
720
|
export declare function renderFindingsMarkdown(content: string): string;
|
|
721
|
+
export declare function describeIncrementalFallback(reason: BroadsideIncrementalOutcome["reason"]): string;
|
|
672
722
|
export declare function estimateSubmitText(result: BroadsideSubmitResult, lenses: LensDefinition[]): string;
|
|
673
723
|
export declare function modelsText(entries: CatalogEntry[], opts: {
|
|
674
724
|
benchmarks: CodingBenchmarks | null;
|
|
675
725
|
defaultModel: string;
|
|
726
|
+
endpoints?: Record<string, BatchEndpointRecord>;
|
|
676
727
|
}): string;
|
|
728
|
+
/**
|
|
729
|
+
* A provider refusal plus what to do about it, for the two refusals a batch
|
|
730
|
+
* run meets in practice and cannot fix by itself (#141):
|
|
731
|
+
*
|
|
732
|
+
* - `Model '<id>' does not have a :batch endpoint.` — the catalog advertises a
|
|
733
|
+
* `:batch` id that OpenRouter runs no batch endpoint for. Nothing in the
|
|
734
|
+
* catalog distinguishes these; the `models` action marks ids this
|
|
735
|
+
* repository has seen refused.
|
|
736
|
+
* - `job-submission-count … in use: 16, quota: 16` — the per-account limit
|
|
737
|
+
* on concurrent batch jobs. Broad-Side submits one job per lens, so a few
|
|
738
|
+
* runs in flight on the same key fill it; the refusal costs nothing.
|
|
739
|
+
*/
|
|
740
|
+
export declare function explainBatchError(error: unknown): string | null;
|
|
677
741
|
export declare function collectResultText(result: BroadsideCollectResult): string;
|
|
742
|
+
/**
|
|
743
|
+
* An `onStatus` callback that appends one line to `lines` per *change* of a
|
|
744
|
+
* lens's polled status. Every poll used to append a line, so a four-minute
|
|
745
|
+
* wait returned twenty-six identical "in_progress (0/1)" lines per lens
|
|
746
|
+
* before the result (0.22.0 live run).
|
|
747
|
+
*/
|
|
748
|
+
export declare function statusLineWriter(lines: string[]): (lensId: string, status: string, counts: Record<string, unknown>) => void;
|
|
678
749
|
export declare function statusText(state: BroadsideStateFile): string;
|
|
679
750
|
export {};
|
package/dist/core/broadside.js
CHANGED
|
@@ -65,6 +65,13 @@ export const BROADSIDE_OUTPUT_PRICE_PER_M = 1.875;
|
|
|
65
65
|
export const BROADSIDE_MODELS_URL = "https://openrouter.ai/api/v1/models";
|
|
66
66
|
export const BROADSIDE_BENCHMARKS_URL = "https://openrouter.ai/api/v1/benchmarks";
|
|
67
67
|
export const BROADSIDE_CATALOG_CACHE_FILE = "model-catalog.json";
|
|
68
|
+
/**
|
|
69
|
+
* What this repository's own submits learned about batch endpoints: which
|
|
70
|
+
* `:batch` ids OpenRouter accepted a job for and which it refused with
|
|
71
|
+
* "does not have a :batch endpoint". The catalog cannot tell the two apart
|
|
72
|
+
* (#141), so the `models` action annotates its rows from this file.
|
|
73
|
+
*/
|
|
74
|
+
export const BROADSIDE_ENDPOINTS_FILE = "batch-endpoints.json";
|
|
68
75
|
export const BROADSIDE_CATALOG_CACHE_TTL_MS = 24 * 60 * 60 * 1000;
|
|
69
76
|
export const BROADSIDE_LENS_IDS = [
|
|
70
77
|
"architecture",
|
|
@@ -87,32 +94,51 @@ export const BROADSIDE_DEFAULT_POLL_BUDGET_MS = 25 * 60 * 1000;
|
|
|
87
94
|
*/
|
|
88
95
|
export const BROADSIDE_DEFAULT_MAX_COST = 1;
|
|
89
96
|
/**
|
|
90
|
-
* The
|
|
97
|
+
* The reasoning control every lens request carries: low effort.
|
|
98
|
+
*
|
|
99
|
+
* It used to be a token cap — `max_tokens` at a quarter of the lens's output
|
|
100
|
+
* budget, so three quarters stayed for the answer. Measured live on
|
|
101
|
+
* `google/gemini-3.8-flash:batch` (0.22.0 verification, defect lens, cap
|
|
102
|
+
* 5,800 of a 6,000 budget): the model reasoned 5,218 tokens on the first
|
|
103
|
+
* pass and **11,518 under the same cap** on the doubled-budget retry —
|
|
104
|
+
* thinking scaled with `max_tokens` and the cap changed nothing, both
|
|
105
|
+
* results truncated, and the retry cost twice the original for no JSON.
|
|
106
|
+
* The same lens with `effort: "low"` reasoned 0 tokens, finished with
|
|
107
|
+
* `stop`, returned valid JSON, and cost a twelfth as much. Gemini 3.x
|
|
108
|
+
* models take a thinking *level*, not a budget, and OpenRouter forwards a
|
|
109
|
+
* `max_tokens` cap to them as nothing at all; `effort` is what it can
|
|
110
|
+
* translate for every provider (a level where the provider has levels, a
|
|
111
|
+
* fraction of the budget where it takes a budget). So the default asks for
|
|
112
|
+
* little thinking in the one vocabulary that reaches everyone.
|
|
91
113
|
*
|
|
92
|
-
*
|
|
93
|
-
*
|
|
94
|
-
*
|
|
114
|
+
* Deliberately not `enabled: false`: `google/gemini-3.8-flash:batch` refuses
|
|
115
|
+
* the whole batch with *"Reasoning is mandatory for this endpoint and cannot
|
|
116
|
+
* be disabled"*, turning a partial result into none at all. Low effort works
|
|
117
|
+
* whether or not a provider allows reasoning to be switched off.
|
|
95
118
|
*/
|
|
96
|
-
export const
|
|
97
|
-
|
|
119
|
+
export const BROADSIDE_DEFAULT_REASONING = Object.freeze({ effort: "low" });
|
|
120
|
+
/** The reasoning control a lens request carries when config.yaml sets none. */
|
|
121
|
+
export function defaultReasoningFor() {
|
|
122
|
+
return { ...BROADSIDE_DEFAULT_REASONING };
|
|
123
|
+
}
|
|
98
124
|
/**
|
|
99
|
-
*
|
|
100
|
-
*
|
|
101
|
-
* Disabling outright is not portable: `google/gemini-3.8-flash:batch` refuses
|
|
102
|
-
* the whole batch with *"Reasoning is mandatory for this endpoint and cannot be
|
|
103
|
-
* disabled"*, turning a partial result into none at all. Capping works whether
|
|
104
|
-
* or not a provider allows reasoning to be switched off.
|
|
125
|
+
* The reasoning control a truncated slice is re-submitted with.
|
|
105
126
|
*
|
|
106
|
-
*
|
|
107
|
-
*
|
|
108
|
-
*
|
|
109
|
-
*
|
|
110
|
-
*
|
|
111
|
-
*
|
|
112
|
-
*
|
|
127
|
+
* A truncation on a reasoning-capable model is usually thinking that ate the
|
|
128
|
+
* answer's budget, and doubling `max_tokens` doubles the thinking where the
|
|
129
|
+
* provider ignores a token cap (see {@link BROADSIDE_DEFAULT_REASONING}). The
|
|
130
|
+
* retry therefore asks for low effort as well, replacing a `max_tokens` cap
|
|
131
|
+
* (OpenRouter refuses a request carrying both) and lowering a higher effort.
|
|
132
|
+
* An explicit `enabled: false` and an effort already at or below low are left
|
|
133
|
+
* as they are.
|
|
113
134
|
*/
|
|
114
|
-
export function
|
|
115
|
-
|
|
135
|
+
export function retryReasoningFor(original) {
|
|
136
|
+
if (original?.enabled === false)
|
|
137
|
+
return { ...original };
|
|
138
|
+
if (original?.effort === "minimal" || original?.effort === "low")
|
|
139
|
+
return { ...original };
|
|
140
|
+
const { max_tokens: _cap, effort: _effort, ...rest } = original ?? {};
|
|
141
|
+
return { ...rest, effort: "low" };
|
|
116
142
|
}
|
|
117
143
|
/**
|
|
118
144
|
* OpenRouter rejected the API key (HTTP 401/403). Thrown from the catalog
|
|
@@ -1341,7 +1367,7 @@ export function buildBatchRequest(lens, info, slice, index, sliceCount, model =
|
|
|
1341
1367
|
max_tokens: maxTokensOverride ?? lens.maxTokens,
|
|
1342
1368
|
// Always sent, never inherited: an absent field means the model's
|
|
1343
1369
|
// own default, and that default is what truncated the JSON.
|
|
1344
|
-
reasoning: reasoningOverride ?? lens.reasoning ?? defaultReasoningFor(
|
|
1370
|
+
reasoning: reasoningOverride ?? lens.reasoning ?? defaultReasoningFor(),
|
|
1345
1371
|
},
|
|
1346
1372
|
};
|
|
1347
1373
|
}
|
|
@@ -1542,6 +1568,21 @@ export async function loadBroadsideConfig(broadsideDir) {
|
|
|
1542
1568
|
throw new BroadsideConfigError(configPath, "is not a YAML mapping");
|
|
1543
1569
|
raw = parsed;
|
|
1544
1570
|
}
|
|
1571
|
+
// OpenRouter accepts `reasoning.effort` or `reasoning.max_tokens`, not
|
|
1572
|
+
// both: a request carrying both is refused per request *after* the batch
|
|
1573
|
+
// is accepted, so every lens fails at $0 with the reason in each
|
|
1574
|
+
// result's error. Seen live on 0.22.0 with the two keys set together.
|
|
1575
|
+
// Refuse here, where the file can be fixed, rather than submit a run
|
|
1576
|
+
// that cannot produce a result.
|
|
1577
|
+
const reasoning = raw.reasoning;
|
|
1578
|
+
if (reasoning && typeof reasoning === "object" && !Array.isArray(reasoning)) {
|
|
1579
|
+
const value = reasoning;
|
|
1580
|
+
const hasEffort = typeof value.effort === "string";
|
|
1581
|
+
const hasBudget = typeof value.max_tokens === "number" && value.max_tokens > 0;
|
|
1582
|
+
if (hasEffort && hasBudget) {
|
|
1583
|
+
throw new BroadsideConfigError(configPath, 'sets both reasoning.effort and reasoning.max_tokens; OpenRouter accepts one or the other ("Only one of reasoning.effort and reasoning.max_tokens can be specified"), and every lens request would fail after the batch is accepted. Keep one');
|
|
1584
|
+
}
|
|
1585
|
+
}
|
|
1545
1586
|
}
|
|
1546
1587
|
return buildBroadsideConfig(raw);
|
|
1547
1588
|
}
|
|
@@ -1621,6 +1662,67 @@ async function writeCatalogCache(broadsideDir, cache) {
|
|
|
1621
1662
|
await mkdir(broadsideDir, { recursive: true });
|
|
1622
1663
|
await writeFile(join(broadsideDir, BROADSIDE_CATALOG_CACHE_FILE), `${JSON.stringify(cache, null, "\t")}\n`, "utf8");
|
|
1623
1664
|
}
|
|
1665
|
+
const BROADSIDE_ENDPOINTS_SCHEMA = 1;
|
|
1666
|
+
export async function readBatchEndpoints(broadsideDir) {
|
|
1667
|
+
const path = join(broadsideDir, BROADSIDE_ENDPOINTS_FILE);
|
|
1668
|
+
if (!(await pathExists(path)))
|
|
1669
|
+
return {};
|
|
1670
|
+
try {
|
|
1671
|
+
const parsed = JSON.parse(await readFile(path, "utf8"));
|
|
1672
|
+
if (!parsed || typeof parsed !== "object" || parsed.schema_version !== BROADSIDE_ENDPOINTS_SCHEMA)
|
|
1673
|
+
return {};
|
|
1674
|
+
if (!parsed.models || typeof parsed.models !== "object")
|
|
1675
|
+
return {};
|
|
1676
|
+
const out = {};
|
|
1677
|
+
for (const [model, record] of Object.entries(parsed.models)) {
|
|
1678
|
+
if (!record || typeof record !== "object")
|
|
1679
|
+
continue;
|
|
1680
|
+
if (record.status !== "accepted" && record.status !== "rejected")
|
|
1681
|
+
continue;
|
|
1682
|
+
if (typeof record.at !== "string")
|
|
1683
|
+
continue;
|
|
1684
|
+
out[model] = { status: record.status, at: record.at, ...(typeof record.error === "string" && { error: record.error }) };
|
|
1685
|
+
}
|
|
1686
|
+
return out;
|
|
1687
|
+
}
|
|
1688
|
+
catch {
|
|
1689
|
+
// An unreadable memory is an empty one: it only annotates a listing.
|
|
1690
|
+
return {};
|
|
1691
|
+
}
|
|
1692
|
+
}
|
|
1693
|
+
/**
|
|
1694
|
+
* The refusal OpenRouter returns for a catalog id that has no batch endpoint
|
|
1695
|
+
* behind it. Matched loosely: the message is the only signal there is.
|
|
1696
|
+
*/
|
|
1697
|
+
const NO_BATCH_ENDPOINT_RE = /does not have a :batch endpoint/i;
|
|
1698
|
+
/** The refusal for a full per-account concurrent batch-job quota. */
|
|
1699
|
+
const BATCH_QUOTA_RE = /job-submission-count/i;
|
|
1700
|
+
/**
|
|
1701
|
+
* Remember what a submit learned about each model it posted to. An accepted
|
|
1702
|
+
* job proves the endpoint exists; a "does not have a :batch endpoint"
|
|
1703
|
+
* refusal proves it does not. Any other rejection (quota, malformed request,
|
|
1704
|
+
* auth) says nothing about the endpoint and leaves the record alone.
|
|
1705
|
+
*/
|
|
1706
|
+
export async function recordBatchEndpoints(broadsideDir, outcomes) {
|
|
1707
|
+
const at = new Date().toISOString();
|
|
1708
|
+
const updates = {};
|
|
1709
|
+
for (const { model, batchId, error } of outcomes) {
|
|
1710
|
+
if (batchId) {
|
|
1711
|
+
updates[model] = { status: "accepted", at };
|
|
1712
|
+
continue;
|
|
1713
|
+
}
|
|
1714
|
+
const message = describeBatchError(error);
|
|
1715
|
+
if (message && NO_BATCH_ENDPOINT_RE.test(message)) {
|
|
1716
|
+
updates[model] = { status: "rejected", at, error: message };
|
|
1717
|
+
}
|
|
1718
|
+
}
|
|
1719
|
+
if (Object.keys(updates).length === 0)
|
|
1720
|
+
return;
|
|
1721
|
+
const models = { ...(await readBatchEndpoints(broadsideDir)), ...updates };
|
|
1722
|
+
await mkdir(broadsideDir, { recursive: true });
|
|
1723
|
+
const file = { schema_version: BROADSIDE_ENDPOINTS_SCHEMA, models };
|
|
1724
|
+
await atomicWriteFile(join(broadsideDir, BROADSIDE_ENDPOINTS_FILE), `${JSON.stringify(file, null, "\t")}\n`);
|
|
1725
|
+
}
|
|
1624
1726
|
function parseCatalogEntry(raw) {
|
|
1625
1727
|
const id = String(raw.id ?? "");
|
|
1626
1728
|
if (!id)
|
|
@@ -1843,7 +1945,8 @@ export async function listBatchModels(broadsideDir, config, apiKey, opts = {}) {
|
|
|
1843
1945
|
cache.models[entry.id] = { ...entry, fetched_at: fetchedAt };
|
|
1844
1946
|
await writeCatalogCache(broadsideDir, cache);
|
|
1845
1947
|
const benchmarks = opts.includeBenchmarks ? await fetchCodingBenchmarks(apiKey, fetcher) : null;
|
|
1846
|
-
|
|
1948
|
+
const endpoints = await readBatchEndpoints(broadsideDir);
|
|
1949
|
+
return { entries, source: "live", benchmarks, defaultModel: config.model, endpoints };
|
|
1847
1950
|
}
|
|
1848
1951
|
export async function submitBatch(batchRequests, apiKey, fetcher = fetch, model = BROADSIDE_MODEL) {
|
|
1849
1952
|
// The OpenRouter batch endpoint stream-parses the body and requires
|
|
@@ -2001,7 +2104,8 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
|
|
|
2001
2104
|
throw new Error(`Broad-Side found no ${info.language} source files to scan (detected from ${info.manifest?.path ?? "the file counts"}; ` +
|
|
2002
2105
|
`the lenses look for ${info.sourceExts.join(", ")}). Nothing was submitted.`);
|
|
2003
2106
|
}
|
|
2004
|
-
const
|
|
2107
|
+
const lensModels = { ...config.lensModels, ...opts.lensModels };
|
|
2108
|
+
const modelForLens = (lensId) => lensModels[lensId] ?? model;
|
|
2005
2109
|
const resolved = new Map();
|
|
2006
2110
|
for (const candidate of new Set([model, ...lensIds.map(modelForLens)])) {
|
|
2007
2111
|
const catalog = await resolveCatalogEntry(broadsideDir, config, candidate, apiKey, opts.fetcher);
|
|
@@ -2062,6 +2166,8 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
|
|
|
2062
2166
|
}
|
|
2063
2167
|
// Slice offline first so the estimate covers every request we would send.
|
|
2064
2168
|
const slicesByLens = new Map();
|
|
2169
|
+
// Why a lens ended up with nothing to submit, for the report (see below).
|
|
2170
|
+
const skipReasons = new Map();
|
|
2065
2171
|
let estimatedInputTokens = 0;
|
|
2066
2172
|
let estimatedOutputTokens = 0;
|
|
2067
2173
|
let estimatedTotalCost = 0;
|
|
@@ -2079,12 +2185,21 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
|
|
|
2079
2185
|
for (const file of slice.redactedFiles ?? [])
|
|
2080
2186
|
redactedFiles.add(file);
|
|
2081
2187
|
}
|
|
2188
|
+
const matchedBeforeIncremental = slices.length;
|
|
2082
2189
|
if (changed) {
|
|
2083
2190
|
// Repo-info slices (empty files, e.g. architecture) always run;
|
|
2084
2191
|
// file-backed slices run only when one of their files changed.
|
|
2085
2192
|
slices = slices.filter((s) => s.files.length === 0 || s.files.some((f) => changed.has(f)));
|
|
2086
2193
|
}
|
|
2087
2194
|
slicesByLens.set(lensId, slices);
|
|
2195
|
+
if (slices.length === 0) {
|
|
2196
|
+
const globs = lens.globsFor(info).filter(Boolean);
|
|
2197
|
+
skipReasons.set(lensId, globs.length === 0
|
|
2198
|
+
? "the lens has no file patterns for this language"
|
|
2199
|
+
: matchedBeforeIncremental > 0
|
|
2200
|
+
? "incremental: none of this lens's files changed since the previous run"
|
|
2201
|
+
: `no files matched ${globs.join(", ")}${lens.skipTestFiles ? " (test files excluded)" : ""}`);
|
|
2202
|
+
}
|
|
2088
2203
|
const lensModel = modelForLens(lensId);
|
|
2089
2204
|
const { pricing: lensPricing, outputCap: lensOutputCap } = resolved.get(lensModel);
|
|
2090
2205
|
const maxTokens = lensOutputCap ? Math.min(lens.maxTokens, lensOutputCap) : lens.maxTokens;
|
|
@@ -2193,7 +2308,15 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
|
|
|
2193
2308
|
if (requests.length === 0) {
|
|
2194
2309
|
// No files matched the lens's globs. That is a coverage gap to
|
|
2195
2310
|
// report, not a batch to submit — the API rejects empty batches.
|
|
2311
|
+
// Name the globs: a JavaScript service whose server lives at
|
|
2312
|
+
// src/server.js gets no security review (that lens reads server/**,
|
|
2313
|
+
// **/auth*, **/middleware/**), and "skipped (0 request(s))" alone
|
|
2314
|
+
// read as an empty repository rather than a lens that looked in
|
|
2315
|
+
// the wrong place.
|
|
2196
2316
|
entry.status = "skipped";
|
|
2317
|
+
const reason = skipReasons.get(lensId);
|
|
2318
|
+
if (reason)
|
|
2319
|
+
entry.reason = reason;
|
|
2197
2320
|
continue;
|
|
2198
2321
|
}
|
|
2199
2322
|
submissions.push((async () => {
|
|
@@ -2214,7 +2337,19 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
|
|
|
2214
2337
|
})());
|
|
2215
2338
|
}
|
|
2216
2339
|
await Promise.allSettled(submissions);
|
|
2340
|
+
// A run with no batch behind it has nothing in flight. Every lens was
|
|
2341
|
+
// skipped or refused, so no poll will ever complete it; leaving it
|
|
2342
|
+
// "in-flight" had status listing a refused run above the completed ones
|
|
2343
|
+
// with synthesis and triage "pending" forever.
|
|
2344
|
+
if (!Object.values(run.batches).some((entry) => entry.batchId))
|
|
2345
|
+
run.status = "failed";
|
|
2217
2346
|
await persistBroadsideRun(broadsideDir, run);
|
|
2347
|
+
// What the provider just said about each model's batch endpoint outlives
|
|
2348
|
+
// the run: the `models` action reads it back (#141).
|
|
2349
|
+
await recordBatchEndpoints(broadsideDir, lensIds
|
|
2350
|
+
.map((lensId) => run.batches[lensId])
|
|
2351
|
+
.filter((entry) => Boolean(entry) && entry.status !== "skipped")
|
|
2352
|
+
.map((entry) => ({ model: entry.model ?? model, batchId: entry.batchId, error: entry.error })));
|
|
2218
2353
|
// Persist the exact request bodies so collect can re-submit a truncated
|
|
2219
2354
|
// slice (bumped output cap) without re-walking the repo (#133). The run
|
|
2220
2355
|
// dir is created here rather than waiting for collect so a crash between
|
|
@@ -2520,7 +2655,19 @@ export async function runBroadsideCollect(cwd, apiKey, opts = {}) {
|
|
|
2520
2655
|
truncatedCount += truncated;
|
|
2521
2656
|
totalCost += cost ?? 0;
|
|
2522
2657
|
await writeFile(join(runDir, `raw-${lensId}.json`), `${JSON.stringify(batch, null, "\t")}\n`, "utf8");
|
|
2523
|
-
|
|
2658
|
+
// A batch can complete with every request failed — the account's
|
|
2659
|
+
// concurrent-job quota filling after acceptance does exactly this.
|
|
2660
|
+
// The per-request errors are on disk as `<id>.error.json`, but a
|
|
2661
|
+
// lens reporting "completed, 0 result(s)" with the reason buried
|
|
2662
|
+
// there read as an empty repository rather than a refused run.
|
|
2663
|
+
const results = Array.isArray(batch.results) ? batch.results : [];
|
|
2664
|
+
const failed = results.filter((r) => r.error && extractContent(r) === null);
|
|
2665
|
+
const allFailed = stored.length === 0 && failed.length > 0
|
|
2666
|
+
? `all ${failed.length} request(s) failed: ${explainBatchError(failed[0].error)}`
|
|
2667
|
+
: null;
|
|
2668
|
+
if (allFailed)
|
|
2669
|
+
entry.error = allFailed;
|
|
2670
|
+
lensOutcomes[lensId] = { status, cost: entry.cost, resultCount: entry.resultCount, truncated, ...(allFailed && { error: allFailed }) };
|
|
2524
2671
|
}
|
|
2525
2672
|
else {
|
|
2526
2673
|
// Every non-completed outcome still has to reach the report.
|
|
@@ -2532,17 +2679,30 @@ export async function runBroadsideCollect(cwd, apiKey, opts = {}) {
|
|
|
2532
2679
|
// indistinguishable in the output from one that was never requested.
|
|
2533
2680
|
if (batch.error)
|
|
2534
2681
|
entry.error = batch.error;
|
|
2535
|
-
const error =
|
|
2682
|
+
const error = explainBatchError(batch.error);
|
|
2536
2683
|
lensOutcomes[lensId] = { status, cost: entry.cost, resultCount: entry.resultCount, ...(error && { error }) };
|
|
2537
2684
|
}
|
|
2538
2685
|
await persistBroadsideRun(broadsideDir, run);
|
|
2539
2686
|
}
|
|
2540
|
-
// #133: re-submit truncated slices once with a bumped output cap
|
|
2541
|
-
// requests are pure, so re-running is always safe;
|
|
2542
|
-
// coverage the first pass lost to a max_tokens
|
|
2687
|
+
// #133: re-submit truncated slices once with a bumped output cap and low
|
|
2688
|
+
// reasoning effort. Batch requests are pure, so re-running is always safe;
|
|
2689
|
+
// the aim is to recover coverage the first pass lost to a max_tokens
|
|
2690
|
+
// cutoff, not to loop forever. Low effort because the cutoff is usually
|
|
2691
|
+
// thinking, and a doubled budget doubled the thinking where a token cap
|
|
2692
|
+
// was ignored (see retryReasoningFor).
|
|
2693
|
+
//
|
|
2694
|
+
// All bumped requests for one model go out as ONE batch, and the batches
|
|
2695
|
+
// (one per model, since a batch carries a single model) are polled
|
|
2696
|
+
// together against the shared deadline. Each truncated slice used to be
|
|
2697
|
+
// submitted and polled to terminal before the next was submitted, so a
|
|
2698
|
+
// model that truncated 11 of 13 slices turned a five-minute collect into
|
|
2699
|
+
// eleven sequential round trips — the serialization #136 removed from the
|
|
2700
|
+
// lens pass, still present here (#206). Grouping also keeps the retry to
|
|
2701
|
+
// one job per model against OpenRouter's 16-concurrent-job quota.
|
|
2543
2702
|
let retriedCount = 0;
|
|
2544
2703
|
if (opts.retryTruncated !== false && truncatedCount > 0) {
|
|
2545
2704
|
const requestsByCustomId = await loadStoredRequests(runDir);
|
|
2705
|
+
const byModel = new Map();
|
|
2546
2706
|
for (const stored of allLensResults) {
|
|
2547
2707
|
if (!stored.truncated)
|
|
2548
2708
|
continue;
|
|
@@ -2559,42 +2719,60 @@ export async function runBroadsideCollect(cwd, apiKey, opts = {}) {
|
|
|
2559
2719
|
const bumpedMax = lensCap ? Math.min(previousMax * 2, lensCap) : previousMax * 2;
|
|
2560
2720
|
if (bumpedMax <= previousMax)
|
|
2561
2721
|
continue; // already at the ceiling
|
|
2562
|
-
const
|
|
2722
|
+
const group = byModel.get(lensModel) ?? { requests: [], slices: new Map() };
|
|
2723
|
+
group.requests.push({
|
|
2563
2724
|
...original,
|
|
2564
|
-
body: { ...original.body, max_tokens: bumpedMax },
|
|
2565
|
-
};
|
|
2725
|
+
body: { ...original.body, max_tokens: bumpedMax, reasoning: retryReasoningFor(original.body.reasoning) },
|
|
2726
|
+
});
|
|
2727
|
+
group.slices.set(stored.customId, stored);
|
|
2728
|
+
byModel.set(lensModel, group);
|
|
2729
|
+
}
|
|
2730
|
+
// Submit every group, then poll whatever was accepted, together.
|
|
2731
|
+
const submitted = [];
|
|
2732
|
+
for (const [model, group] of byModel) {
|
|
2566
2733
|
try {
|
|
2567
|
-
const { batchId, error } = await submitBatch(
|
|
2568
|
-
if (error)
|
|
2734
|
+
const { batchId, error } = await submitBatch(group.requests, apiKey, opts.fetcher, model);
|
|
2735
|
+
if (!error && batchId)
|
|
2736
|
+
submitted.push({ model, batchId });
|
|
2737
|
+
}
|
|
2738
|
+
catch {
|
|
2739
|
+
// A retry batch that fails to submit leaves its slices' original
|
|
2740
|
+
// truncated results in place — nothing is lost.
|
|
2741
|
+
}
|
|
2742
|
+
}
|
|
2743
|
+
const polled = await pollBatchesConcurrently(submitted.map(({ model, batchId }) => ({ lensId: `retry:${model}`, batchId })), apiKey, {
|
|
2744
|
+
// Share the caller's deadline. Each of these polls used to start a
|
|
2745
|
+
// fresh 25-minute budget, so `wait_seconds` bounded only the lens
|
|
2746
|
+
// poll and a collect could run for the caller's budget plus fifty
|
|
2747
|
+
// minutes.
|
|
2748
|
+
deadlineMs: Math.max(0, deadline - Date.now()),
|
|
2749
|
+
fetcher: opts.fetcher,
|
|
2750
|
+
onStatus: opts.onStatus,
|
|
2751
|
+
});
|
|
2752
|
+
for (const { model, batchId } of submitted) {
|
|
2753
|
+
const batch = polled.get(batchId);
|
|
2754
|
+
if (!batch || batch.status !== "completed")
|
|
2755
|
+
continue;
|
|
2756
|
+
const group = byModel.get(model);
|
|
2757
|
+
const usage = (batch.usage ?? {});
|
|
2758
|
+
totalCost += typeof usage.cost === "number" ? usage.cost : 0;
|
|
2759
|
+
const results = Array.isArray(batch.results) ? batch.results : [];
|
|
2760
|
+
for (const result of results) {
|
|
2761
|
+
const stored = group.slices.get(String(result.custom_id ?? ""));
|
|
2762
|
+
if (!stored)
|
|
2569
2763
|
continue;
|
|
2570
|
-
const
|
|
2571
|
-
|
|
2572
|
-
// start a fresh 25-minute budget, so `wait_seconds` bounded
|
|
2573
|
-
// only the lens poll and a collect could run for the caller's
|
|
2574
|
-
// budget plus fifty minutes.
|
|
2575
|
-
deadlineMs: Math.max(0, deadline - Date.now()),
|
|
2576
|
-
onStatus: (status, counts) => opts.onStatus?.(`${stored.lensId}:retry`, status, counts),
|
|
2577
|
-
fetcher: opts.fetcher,
|
|
2578
|
-
});
|
|
2579
|
-
if (batch.status !== "completed")
|
|
2764
|
+
const content = extractContent(result);
|
|
2765
|
+
if (content === null)
|
|
2580
2766
|
continue;
|
|
2581
|
-
const results = Array.isArray(batch.results) ? batch.results : [];
|
|
2582
|
-
const content = results.length > 0 ? extractContent(results[0]) : null;
|
|
2583
|
-
if (content === null || parseLensJson(content) === null)
|
|
2584
|
-
continue; // still no good
|
|
2585
|
-
const usage = (batch.usage ?? {});
|
|
2586
|
-
totalCost += typeof usage.cost === "number" ? usage.cost : 0;
|
|
2587
2767
|
const parsed = parseLensJson(content);
|
|
2768
|
+
if (parsed === null)
|
|
2769
|
+
continue; // still no good
|
|
2588
2770
|
await writeFile(join(runDir, `${sanitizeId(stored.customId)}.json`), `${JSON.stringify(parsed, null, "\t")}\n`, "utf8");
|
|
2589
2771
|
await writeFile(join(runDir, `${sanitizeId(stored.customId)}.md`), renderFindingsMarkdown(content), "utf8");
|
|
2590
2772
|
stored.content = content;
|
|
2591
2773
|
stored.truncated = false;
|
|
2592
2774
|
retriedCount += 1;
|
|
2593
2775
|
}
|
|
2594
|
-
catch {
|
|
2595
|
-
// A retry that fails to submit/poll leaves the original
|
|
2596
|
-
// truncated result in place — nothing is lost.
|
|
2597
|
-
}
|
|
2598
2776
|
}
|
|
2599
2777
|
truncatedCount = allLensResults.filter((s) => s.truncated).length;
|
|
2600
2778
|
for (const [lensId, outcome] of Object.entries(lensOutcomes)) {
|
|
@@ -2846,7 +3024,7 @@ function parseSynthesisTopFindings(content) {
|
|
|
2846
3024
|
}
|
|
2847
3025
|
}
|
|
2848
3026
|
// ---------- formatting helpers for tool output ----------
|
|
2849
|
-
function describeIncrementalFallback(reason) {
|
|
3027
|
+
export function describeIncrementalFallback(reason) {
|
|
2850
3028
|
switch (reason) {
|
|
2851
3029
|
case "dirty-worktree":
|
|
2852
3030
|
return "the working tree has uncommitted changes, so there is no committed state to diff against";
|
|
@@ -2878,7 +3056,16 @@ export function estimateSubmitText(result, lenses) {
|
|
|
2878
3056
|
continue;
|
|
2879
3057
|
const status = entry.batchId ? `batch ${entry.batchId}` : entry.status;
|
|
2880
3058
|
const override = entry.model ? ` on ${entry.model}` : "";
|
|
2881
|
-
|
|
3059
|
+
// A rejected lens says why: the message is the only way to tell a
|
|
3060
|
+
// catalog id with no batch endpoint from a full job quota, and both
|
|
3061
|
+
// used to read as a bare "rejected". A skipped lens names the globs
|
|
3062
|
+
// that matched nothing.
|
|
3063
|
+
const reason = !entry.batchId && entry.error
|
|
3064
|
+
? ` — ${explainBatchError(entry.error)}`
|
|
3065
|
+
: !entry.batchId && entry.reason
|
|
3066
|
+
? ` — ${entry.reason}`
|
|
3067
|
+
: "";
|
|
3068
|
+
lines.push(` ${lens.name}: ${status} (${entry.requests} request(s), ~$${entry.estimatedCost.toFixed(4)})${override}${reason}`);
|
|
2882
3069
|
}
|
|
2883
3070
|
if (result.repo) {
|
|
2884
3071
|
const head = result.repo.sourceHead ? ` at ${result.repo.sourceHead.slice(0, 8)}${result.repo.sourceDirty ? " (dirty)" : ""}` : "";
|
|
@@ -2915,8 +3102,15 @@ export function estimateSubmitText(result, lenses) {
|
|
|
2915
3102
|
return lines.join("\n");
|
|
2916
3103
|
}
|
|
2917
3104
|
export function modelsText(entries, opts) {
|
|
3105
|
+
const endpoints = opts.endpoints ?? {};
|
|
2918
3106
|
const lines = [
|
|
2919
3107
|
`Batch models on OpenRouter (${entries.length}, cheapest first).`,
|
|
3108
|
+
// The catalog over-reports: it returns a `:batch` id for models whose
|
|
3109
|
+
// Batch API refuses the job, with nothing in the entry to tell them
|
|
3110
|
+
// apart (#141). Say so before the table, not after it.
|
|
3111
|
+
"Advisory: this is the catalog's list of :batch ids, not a list of working batch endpoints. Some ids are refused at submit " +
|
|
3112
|
+
"(\"does not have a :batch endpoint\"), at no cost. Rows tagged [no batch endpoint …] or [batch OK …] carry what this " +
|
|
3113
|
+
"repository's own submits found; an untagged row has not been tried here.",
|
|
2920
3114
|
"",
|
|
2921
3115
|
"id | $/M in | $/M out | ctx | max out | structured | coding idx",
|
|
2922
3116
|
];
|
|
@@ -2936,12 +3130,21 @@ export function modelsText(entries, opts) {
|
|
|
2936
3130
|
const out = entry.maxCompletionTokens ? `${(entry.maxCompletionTokens / 1024).toFixed(0)}k` : "?";
|
|
2937
3131
|
const tag = entry.id === opts.defaultModel ? " (default)" : "";
|
|
2938
3132
|
const exp = entry.expirationDate ? " [deprecated]" : "";
|
|
2939
|
-
|
|
3133
|
+
const record = endpoints[entry.id];
|
|
3134
|
+
const seen = record
|
|
3135
|
+
? record.status === "rejected"
|
|
3136
|
+
? ` [no batch endpoint, refused ${record.at.slice(0, 10)}]`
|
|
3137
|
+
: ` [batch OK ${record.at.slice(0, 10)}]`
|
|
3138
|
+
: "";
|
|
3139
|
+
lines.push(`${entry.id}${tag}${exp}${seen} | ${entry.inputPerM.toFixed(3)} | ${entry.outputPerM.toFixed(3)} | ${ctx} | ${out} | ${structured} | ${coding}`);
|
|
2940
3140
|
}
|
|
2941
3141
|
if (opts.benchmarks?.meta.as_of) {
|
|
2942
3142
|
lines.push("", `Benchmarks: Artificial Analysis coding index (as of ${String(opts.benchmarks.meta.as_of)}).`);
|
|
2943
3143
|
}
|
|
2944
|
-
lines.push("", "
|
|
3144
|
+
lines.push("", "Choose with the model parameter (--model= on Pi) for one run, lens_models (--lens-model=LENS:ID) per lens, or the model key in " +
|
|
3145
|
+
".codecarto/broadside/config.yaml for the repository. Higher coding index ≠ better scout: precision, context, structured-output " +
|
|
3146
|
+
"support, and whether the model spends its output budget reasoning (see reasoning: in config.yaml) matter most here. " +
|
|
3147
|
+
"A refused submit costs nothing, so probe an untried model on one lens first.");
|
|
2945
3148
|
return lines.join("\n");
|
|
2946
3149
|
}
|
|
2947
3150
|
/** One line of a batch's error field, whatever shape the provider gave it. */
|
|
@@ -2954,6 +3157,15 @@ function describeBatchError(error) {
|
|
|
2954
3157
|
const message = error.message;
|
|
2955
3158
|
if (typeof message === "string" && message)
|
|
2956
3159
|
return message.slice(0, 300);
|
|
3160
|
+
// OpenRouter wraps a submit refusal as `{ error: { message } }`.
|
|
3161
|
+
const nested = error.error;
|
|
3162
|
+
if (nested && typeof nested === "object") {
|
|
3163
|
+
const inner = nested.message;
|
|
3164
|
+
if (typeof inner === "string" && inner)
|
|
3165
|
+
return inner.slice(0, 300);
|
|
3166
|
+
}
|
|
3167
|
+
if (typeof nested === "string" && nested)
|
|
3168
|
+
return nested.slice(0, 300);
|
|
2957
3169
|
try {
|
|
2958
3170
|
return JSON.stringify(error).slice(0, 300);
|
|
2959
3171
|
}
|
|
@@ -2963,6 +3175,30 @@ function describeBatchError(error) {
|
|
|
2963
3175
|
}
|
|
2964
3176
|
return String(error);
|
|
2965
3177
|
}
|
|
3178
|
+
/**
|
|
3179
|
+
* A provider refusal plus what to do about it, for the two refusals a batch
|
|
3180
|
+
* run meets in practice and cannot fix by itself (#141):
|
|
3181
|
+
*
|
|
3182
|
+
* - `Model '<id>' does not have a :batch endpoint.` — the catalog advertises a
|
|
3183
|
+
* `:batch` id that OpenRouter runs no batch endpoint for. Nothing in the
|
|
3184
|
+
* catalog distinguishes these; the `models` action marks ids this
|
|
3185
|
+
* repository has seen refused.
|
|
3186
|
+
* - `job-submission-count … in use: 16, quota: 16` — the per-account limit
|
|
3187
|
+
* on concurrent batch jobs. Broad-Side submits one job per lens, so a few
|
|
3188
|
+
* runs in flight on the same key fill it; the refusal costs nothing.
|
|
3189
|
+
*/
|
|
3190
|
+
export function explainBatchError(error) {
|
|
3191
|
+
const message = describeBatchError(error);
|
|
3192
|
+
if (!message)
|
|
3193
|
+
return null;
|
|
3194
|
+
if (NO_BATCH_ENDPOINT_RE.test(message)) {
|
|
3195
|
+
return `${message} — the catalog lists this id, but OpenRouter runs no batch endpoint for it. Nothing was charged; pick another model (the models action marks ids this repository has seen refused).`;
|
|
3196
|
+
}
|
|
3197
|
+
if (BATCH_QUOTA_RE.test(message)) {
|
|
3198
|
+
return `${message} — OpenRouter's per-account limit on concurrent batch jobs is full. Broad-Side submits one job per lens, so a few runs in flight on this key (in any repository) fill it. Nothing was charged; collect or wait out the runs in flight, then re-submit.`;
|
|
3199
|
+
}
|
|
3200
|
+
return message;
|
|
3201
|
+
}
|
|
2966
3202
|
export function collectResultText(result) {
|
|
2967
3203
|
const lines = [
|
|
2968
3204
|
`Broad-Side run ${result.runId}: ${result.status}`,
|
|
@@ -3012,6 +3248,22 @@ export function collectResultText(result) {
|
|
|
3012
3248
|
lines.push("", "Disclaimer: Broad-Side findings are unverified scouting signals from a batch model, not validated claims.");
|
|
3013
3249
|
return lines.join("\n");
|
|
3014
3250
|
}
|
|
3251
|
+
/**
|
|
3252
|
+
* An `onStatus` callback that appends one line to `lines` per *change* of a
|
|
3253
|
+
* lens's polled status. Every poll used to append a line, so a four-minute
|
|
3254
|
+
* wait returned twenty-six identical "in_progress (0/1)" lines per lens
|
|
3255
|
+
* before the result (0.22.0 live run).
|
|
3256
|
+
*/
|
|
3257
|
+
export function statusLineWriter(lines) {
|
|
3258
|
+
const last = new Map();
|
|
3259
|
+
return (lensId, status, counts) => {
|
|
3260
|
+
const line = ` ${lensId}: ${status} (${counts.completed ?? 0}/${counts.total ?? "?"})`;
|
|
3261
|
+
if (last.get(lensId) === line)
|
|
3262
|
+
return;
|
|
3263
|
+
last.set(lensId, line);
|
|
3264
|
+
lines.push(line);
|
|
3265
|
+
};
|
|
3266
|
+
}
|
|
3015
3267
|
export function statusText(state) {
|
|
3016
3268
|
if (state.runs.length === 0) {
|
|
3017
3269
|
return "No Broad-Side runs recorded. Call codecarto_broadside with action 'submit' first.";
|
|
@@ -3029,7 +3281,8 @@ export function statusText(state) {
|
|
|
3029
3281
|
const entry = run.batches[lensId];
|
|
3030
3282
|
if (!entry)
|
|
3031
3283
|
continue;
|
|
3032
|
-
lines.push(` ${lensId}: ${entry.status}${entry.batchId ? ` (${entry.batchId})` : ""}${entry.cost !== undefined ? `, $${entry.cost.toFixed(6)}` : ""}`
|
|
3284
|
+
lines.push(` ${lensId}: ${entry.status}${entry.batchId ? ` (${entry.batchId})` : ""}${entry.cost !== undefined ? `, $${entry.cost.toFixed(6)}` : ""}` +
|
|
3285
|
+
(entry.status === "skipped" && entry.reason ? ` — ${entry.reason}` : ""));
|
|
3033
3286
|
}
|
|
3034
3287
|
lines.push(` synthesis: ${run.synthesis.status}`);
|
|
3035
3288
|
lines.push(` triage: ${run.triage?.status ?? "pending"}`);
|
|
@@ -18,11 +18,15 @@ export interface BroadsideFlags {
|
|
|
18
18
|
waitSeconds?: number;
|
|
19
19
|
/** For collect: the run to collect instead of the most recent (#268). */
|
|
20
20
|
runId?: string;
|
|
21
|
+
/** For submit: the run's batch model, replacing config.yaml's (#141). */
|
|
22
|
+
model?: string;
|
|
23
|
+
/** For submit: per-lens model overrides, layered over config.yaml's (#141). */
|
|
24
|
+
lensModels?: Partial<Record<BroadsideLensId, string>>;
|
|
21
25
|
benchmarks: boolean;
|
|
22
26
|
unknown: string[];
|
|
23
27
|
/** Set on an invalid combination. The caller surfaces it as an error. */
|
|
24
28
|
error?: string;
|
|
25
29
|
}
|
|
26
30
|
/** Every token the completer offers, in the order it offers them. */
|
|
27
|
-
export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
|
|
31
|
+
export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--model=", "--lens-model=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
|
|
28
32
|
export declare function parseBroadsideFlags(args: string): BroadsideFlags;
|
|
@@ -14,6 +14,11 @@
|
|
|
14
14
|
// --max-cost=N --no-retry-truncated
|
|
15
15
|
// --wait=SECONDS --benchmarks (models only)
|
|
16
16
|
// --run=ID (collect only: an older run, as listed by status)
|
|
17
|
+
// --model=ID (submit only: the run's batch model, as listed by models)
|
|
18
|
+
// --lens-model=LENS:ID (submit only, repeatable: one lens on its own model)
|
|
19
|
+
//
|
|
20
|
+
// A model id itself contains a colon (`vendor/name:batch`), so --lens-model
|
|
21
|
+
// splits on the first colon only: `security:deepseek/deepseek-v4-pro:batch`.
|
|
17
22
|
//
|
|
18
23
|
// --incremental has a spelled-out negative because the value is tri-state:
|
|
19
24
|
// absent defers to config.yaml, so a repository that set `incremental: true`
|
|
@@ -36,6 +41,8 @@ export const KNOWN_BROADSIDE_TOKENS = [
|
|
|
36
41
|
"--max-cost=",
|
|
37
42
|
"--wait=",
|
|
38
43
|
"--run=",
|
|
44
|
+
"--model=",
|
|
45
|
+
"--lens-model=",
|
|
39
46
|
"--no-synthesis",
|
|
40
47
|
"--no-triage",
|
|
41
48
|
"--no-retry-truncated",
|
|
@@ -115,6 +122,32 @@ export function parseBroadsideFlags(args) {
|
|
|
115
122
|
result.runId = value || undefined;
|
|
116
123
|
continue;
|
|
117
124
|
}
|
|
125
|
+
if (token.startsWith("--model=")) {
|
|
126
|
+
const value = token.slice("--model=".length).trim();
|
|
127
|
+
// An empty value is a mistyped selection, not "use the default":
|
|
128
|
+
// the command is about to spend money on whichever model wins.
|
|
129
|
+
if (!value)
|
|
130
|
+
result.error ??= "--model= needs an OpenRouter batch model id (see /codecarto-broadside models).";
|
|
131
|
+
result.model = value || undefined;
|
|
132
|
+
continue;
|
|
133
|
+
}
|
|
134
|
+
if (token.startsWith("--lens-model=")) {
|
|
135
|
+
const value = token.slice("--lens-model=".length).trim();
|
|
136
|
+
const colon = value.indexOf(":");
|
|
137
|
+
const lensId = colon > 0 ? value.slice(0, colon).trim() : "";
|
|
138
|
+
const modelId = colon > 0 ? value.slice(colon + 1).trim() : "";
|
|
139
|
+
if (!lensId || !modelId) {
|
|
140
|
+
result.error ??= `--lens-model needs LENS:MODEL, e.g. --lens-model=security:vendor/name:batch (got "${value}").`;
|
|
141
|
+
}
|
|
142
|
+
else if (!BROADSIDE_LENS_IDS.includes(lensId)) {
|
|
143
|
+
result.error ??= `--lens-model: unknown lens "${lensId}". Lenses: ${BROADSIDE_LENS_IDS.join(", ")}.`;
|
|
144
|
+
}
|
|
145
|
+
else {
|
|
146
|
+
result.lensModels ??= {};
|
|
147
|
+
result.lensModels[lensId] = modelId;
|
|
148
|
+
}
|
|
149
|
+
continue;
|
|
150
|
+
}
|
|
118
151
|
result.unknown.push(token);
|
|
119
152
|
}
|
|
120
153
|
// Flags that only mean something for one action are refused rather than
|
|
@@ -137,5 +170,11 @@ export function parseBroadsideFlags(args) {
|
|
|
137
170
|
if (result.runId !== undefined && result.action !== "collect") {
|
|
138
171
|
result.error ??= `--run is only meaningful for collect (got action "${result.action}").`;
|
|
139
172
|
}
|
|
173
|
+
if (result.model !== undefined && result.action !== "submit") {
|
|
174
|
+
result.error ??= `--model is only meaningful for submit (got action "${result.action}").`;
|
|
175
|
+
}
|
|
176
|
+
if (result.lensModels !== undefined && result.action !== "submit") {
|
|
177
|
+
result.error ??= `--lens-model is only meaningful for submit (got action "${result.action}").`;
|
|
178
|
+
}
|
|
140
179
|
return result;
|
|
141
180
|
}
|
|
@@ -9,7 +9,7 @@ import { parseNextFlags } from "./next-flags.js";
|
|
|
9
9
|
import { buildPiGuideMessage } from "./guide-framing.js";
|
|
10
10
|
import { isCtxLive, notifyCtx } from "./notify.js";
|
|
11
11
|
import { phaseCompactionExtension } from "./phase-compaction.js";
|
|
12
|
-
import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, statusText, describeConfigProblems, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
|
|
12
|
+
import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, statusText, describeConfigProblems, describeIncrementalFallback, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
|
|
13
13
|
import { initLibrary } from "../../core/library.js";
|
|
14
14
|
import { resolveUserConfigPath } from "../../core/orchestrator-config.js";
|
|
15
15
|
const STATUS_WIDGET_ID = "codecarto-widget";
|
|
@@ -144,11 +144,14 @@ function describeBroadsideEstimate(estimate) {
|
|
|
144
144
|
? `This EXCEEDS the configured max_cost of $${estimate.maxCost.toFixed(2)}. Approving here overrides it for this run.`
|
|
145
145
|
: `Within the configured max_cost of $${estimate.maxCost.toFixed(2)}.`);
|
|
146
146
|
}
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
147
|
+
// Only a requested incremental run has anything to say here. The dialog
|
|
148
|
+
// used to print "Incremental was requested but the tree is dirty" on every
|
|
149
|
+
// dirty tree, requested or not — a full scan that nobody asked to shrink
|
|
150
|
+
// read as a fallback.
|
|
151
|
+
if (estimate.incremental?.requested) {
|
|
152
|
+
lines.push(estimate.incremental.applied
|
|
153
|
+
? `Incremental: only modules changed since ${(estimate.baseHead ?? "").slice(0, 8)} are included.`
|
|
154
|
+
: `Incremental was requested but NOT applied — ${describeIncrementalFallback(estimate.incremental.reason)}. This is a full scan.`);
|
|
152
155
|
}
|
|
153
156
|
lines.push("", "The estimate is a pre-flight prediction from file sizes; OpenRouter bills actual usage.");
|
|
154
157
|
return lines.join("\n");
|
|
@@ -1000,7 +1003,7 @@ export default function codeCartographerExtension(pi) {
|
|
|
1000
1003
|
},
|
|
1001
1004
|
});
|
|
1002
1005
|
pi.registerCommand("codecarto-broadside", {
|
|
1003
|
-
description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models] [lenses…] [flags]",
|
|
1006
|
+
description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID] [flags]",
|
|
1004
1007
|
getArgumentCompletions: (prefix) => {
|
|
1005
1008
|
const items = KNOWN_BROADSIDE_TOKENS
|
|
1006
1009
|
.filter((value) => value.startsWith(prefix))
|
|
@@ -1077,10 +1080,10 @@ export default function codeCartographerExtension(pi) {
|
|
|
1077
1080
|
if (flags.action === "models") {
|
|
1078
1081
|
notifyCtx(ctx, "Fetching the OpenRouter batch-model catalog…", "info");
|
|
1079
1082
|
try {
|
|
1080
|
-
const { entries, benchmarks } = await listBatchModels(broadsideDir, config, apiKey, {
|
|
1083
|
+
const { entries, benchmarks, endpoints } = await listBatchModels(broadsideDir, config, apiKey, {
|
|
1081
1084
|
includeBenchmarks: flags.benchmarks,
|
|
1082
1085
|
});
|
|
1083
|
-
finish(modelsText(entries, { benchmarks, defaultModel: config.model }).split("\n"), `Broad-Side: ${entries.length} batch model${entries.length === 1 ? "" : "s"} listed`);
|
|
1086
|
+
finish(modelsText(entries, { benchmarks, defaultModel: config.model, endpoints }).split("\n"), `Broad-Side: ${entries.length} batch model${entries.length === 1 ? "" : "s"} listed`);
|
|
1084
1087
|
}
|
|
1085
1088
|
catch (error) {
|
|
1086
1089
|
notifyCtx(ctx, `Model catalog lookup failed: ${error instanceof Error ? error.message : String(error)}`, "error");
|
|
@@ -1114,24 +1117,49 @@ export default function codeCartographerExtension(pi) {
|
|
|
1114
1117
|
const lenses = flags.lenses.length > 0 ? flags.lenses : config.defaultLenses;
|
|
1115
1118
|
renderProgress("Slicing the repository and pricing the run…");
|
|
1116
1119
|
let submit;
|
|
1120
|
+
// Set when a headless run was refused over max_cost, so the cancel
|
|
1121
|
+
// message says so instead of reading as a user's "no".
|
|
1122
|
+
let headlessRefusal = null;
|
|
1117
1123
|
try {
|
|
1118
1124
|
submit = await runBroadsideSubmit(ctx.cwd, apiKey, {
|
|
1119
1125
|
lenses,
|
|
1120
|
-
model
|
|
1126
|
+
// --model= and --lens-model= select for this run; the file's
|
|
1127
|
+
// values are the fallback, and core pre-flights either the
|
|
1128
|
+
// same way (#141).
|
|
1129
|
+
model: flags.model ?? config.model,
|
|
1130
|
+
lensModels: flags.lensModels,
|
|
1121
1131
|
maxCost: flags.maxCost ?? config.maxCost,
|
|
1122
1132
|
// `??`, not `||`: --no-incremental parses to false and must beat a
|
|
1123
1133
|
// config-set true, exactly as MCP's `incremental: false` does (#163).
|
|
1124
1134
|
incremental: flags.incremental ?? config.incremental,
|
|
1125
1135
|
// Pi can ask, so it asks instead of refusing over max_cost the
|
|
1126
1136
|
// way MCP has to. An approval here IS the force flag.
|
|
1127
|
-
confirm: (estimate) =>
|
|
1137
|
+
confirm: (estimate) => {
|
|
1138
|
+
if (ctx.hasUI) {
|
|
1139
|
+
return ctx.ui.confirm(`Broad-Side will spend about $${estimate.totalCost.toFixed(4)}`, describeBroadsideEstimate(estimate));
|
|
1140
|
+
}
|
|
1141
|
+
// No dialog under `pi -p`: the stub answered "no" to every
|
|
1142
|
+
// estimate, so a headless submit could never fire. Behave as
|
|
1143
|
+
// the MCP surface does — an estimate within max_cost is
|
|
1144
|
+
// approved by the cap itself; one over it is refused, since
|
|
1145
|
+
// nobody is here to say yes — and print the breakdown either
|
|
1146
|
+
// way, because the dialog was the only place it showed.
|
|
1147
|
+
notifyCtx(ctx, describeBroadsideEstimate(estimate), "info");
|
|
1148
|
+
if (estimate.exceedsLimit) {
|
|
1149
|
+
headlessRefusal =
|
|
1150
|
+
`Broad-Side refused: the estimate ~$${estimate.totalCost.toFixed(4)} exceeds max_cost $${estimate.maxCost.toFixed(2)} ` +
|
|
1151
|
+
"and there is no dialog to approve it in a headless run. Raise --max-cost (0 for no limit) or run interactively.";
|
|
1152
|
+
return false;
|
|
1153
|
+
}
|
|
1154
|
+
return true;
|
|
1155
|
+
},
|
|
1128
1156
|
});
|
|
1129
1157
|
}
|
|
1130
1158
|
catch (error) {
|
|
1131
1159
|
if (ctx.hasUI)
|
|
1132
1160
|
ctx.ui.setWidget(BROADSIDE_WIDGET_ID, undefined);
|
|
1133
1161
|
if (error instanceof BroadsideCancelledError) {
|
|
1134
|
-
notifyCtx(ctx, "Broad-Side cancelled. Nothing was submitted.", "info");
|
|
1162
|
+
notifyCtx(ctx, headlessRefusal ?? "Broad-Side cancelled. Nothing was submitted.", headlessRefusal ? "error" : "info");
|
|
1135
1163
|
return;
|
|
1136
1164
|
}
|
|
1137
1165
|
notifyCtx(ctx, `Broad-Side submit failed: ${error instanceof Error ? error.message : String(error)}`, "error");
|
|
@@ -17,7 +17,7 @@ import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"
|
|
|
17
17
|
import { CallToolRequestSchema, ErrorCode, ListToolsRequestSchema, McpError, } from "@modelcontextprotocol/sdk/types.js";
|
|
18
18
|
import { mkdir, readFile, readdir, rename, writeFile } from "node:fs/promises";
|
|
19
19
|
import { basename, isAbsolute, join } from "node:path";
|
|
20
|
-
import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, seedOrchestratorFiles, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
|
|
20
|
+
import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, seedOrchestratorFiles, statusLineWriter, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
|
|
21
21
|
import { applyAmendment } from "../core/amendment.js";
|
|
22
22
|
import { appendUsageRun } from "../core/usage.js";
|
|
23
23
|
import { initLibrary } from "../core/library.js";
|
|
@@ -1079,15 +1079,19 @@ export async function handleBroadside(args) {
|
|
|
1079
1079
|
const retryTruncated = args.retry_truncated ?? config.retryTruncated;
|
|
1080
1080
|
const incremental = args.incremental ?? config.incremental;
|
|
1081
1081
|
if (action === "models") {
|
|
1082
|
-
const { entries, benchmarks } = await listBatchModels(broadsideDirFor(cwd), config, apiKey, {
|
|
1082
|
+
const { entries, benchmarks, endpoints } = await listBatchModels(broadsideDirFor(cwd), config, apiKey, {
|
|
1083
1083
|
includeBenchmarks: args.include_benchmarks === true,
|
|
1084
1084
|
}).catch((error) => {
|
|
1085
1085
|
throw new McpError(ErrorCode.InvalidRequest, error instanceof Error ? error.message : String(error));
|
|
1086
1086
|
});
|
|
1087
|
-
return textResult(modelsText(entries, { benchmarks, defaultModel: config.model }), {
|
|
1087
|
+
return textResult(modelsText(entries, { benchmarks, defaultModel: config.model, endpoints }), {
|
|
1088
1088
|
models: entries,
|
|
1089
1089
|
defaultModel: config.model,
|
|
1090
1090
|
benchmarkMeta: benchmarks?.meta ?? null,
|
|
1091
|
+
// The catalog is advisory (#141): what this repository's submits
|
|
1092
|
+
// learned about each id's batch endpoint rides alongside it.
|
|
1093
|
+
catalogAdvisory: true,
|
|
1094
|
+
endpoints,
|
|
1091
1095
|
});
|
|
1092
1096
|
}
|
|
1093
1097
|
if (action === "submit") {
|
|
@@ -1105,9 +1109,34 @@ export async function handleBroadside(args) {
|
|
|
1105
1109
|
// An explicit 0 is "no limit" (#231); absent falls back to config.yaml,
|
|
1106
1110
|
// whose own default is BROADSIDE_DEFAULT_MAX_COST.
|
|
1107
1111
|
const maxCost = typeof args.max_cost === "number" && args.max_cost >= 0 ? args.max_cost : config.maxCost;
|
|
1112
|
+
// Model selection for one run (#141): `model` replaces the run default,
|
|
1113
|
+
// `lens_models` layers per-lens overrides over config.yaml's. Both are
|
|
1114
|
+
// pre-flighted by core exactly like the file's values — priced from the
|
|
1115
|
+
// catalog, refused without structured-output support, clamped to the
|
|
1116
|
+
// model's ceiling — so a wrong id fails before anything is submitted.
|
|
1117
|
+
const model = typeof args.model === "string" && args.model.trim() ? args.model.trim() : config.model;
|
|
1118
|
+
if (args.model !== undefined && !(typeof args.model === "string" && args.model.trim())) {
|
|
1119
|
+
throw new McpError(ErrorCode.InvalidParams, "model must be a non-empty OpenRouter batch model id (see action 'models').");
|
|
1120
|
+
}
|
|
1121
|
+
const lensModels = {};
|
|
1122
|
+
if (args.lens_models !== undefined) {
|
|
1123
|
+
if (!args.lens_models || typeof args.lens_models !== "object" || Array.isArray(args.lens_models)) {
|
|
1124
|
+
throw new McpError(ErrorCode.InvalidParams, "lens_models must be an object mapping lens ids to batch model ids.");
|
|
1125
|
+
}
|
|
1126
|
+
for (const [lensId, value] of Object.entries(args.lens_models)) {
|
|
1127
|
+
if (!BROADSIDE_LENS_IDS.includes(lensId)) {
|
|
1128
|
+
throw new McpError(ErrorCode.InvalidParams, `lens_models: unknown lens "${lensId}". Valid: ${BROADSIDE_LENS_IDS.join(", ")}`);
|
|
1129
|
+
}
|
|
1130
|
+
if (typeof value !== "string" || !value.trim()) {
|
|
1131
|
+
throw new McpError(ErrorCode.InvalidParams, `lens_models.${lensId} must be a non-empty OpenRouter batch model id.`);
|
|
1132
|
+
}
|
|
1133
|
+
lensModels[lensId] = value.trim();
|
|
1134
|
+
}
|
|
1135
|
+
}
|
|
1108
1136
|
const result = await runBroadsideSubmit(cwd, apiKey, {
|
|
1109
1137
|
lenses,
|
|
1110
|
-
model
|
|
1138
|
+
model,
|
|
1139
|
+
lensModels,
|
|
1111
1140
|
maxCost,
|
|
1112
1141
|
force: args.force === true,
|
|
1113
1142
|
incremental,
|
|
@@ -1122,7 +1151,10 @@ export async function handleBroadside(args) {
|
|
|
1122
1151
|
includeSynthesis,
|
|
1123
1152
|
includeTriage,
|
|
1124
1153
|
retryTruncated,
|
|
1125
|
-
|
|
1154
|
+
// One line per *change* of a lens's status. Every poll used to
|
|
1155
|
+
// append a line, so a four-minute wait returned twenty-six
|
|
1156
|
+
// "in_progress (0/1)" lines before the result (0.22.0 live run).
|
|
1157
|
+
onStatus: statusLineWriter(lines),
|
|
1126
1158
|
}).catch((error) => {
|
|
1127
1159
|
// The `collect` action normalizes this same call; without it here,
|
|
1128
1160
|
// a failure during submit-with-wait reached the client as an
|
|
@@ -1516,6 +1548,15 @@ const TOOLS = [
|
|
|
1516
1548
|
type: "boolean",
|
|
1517
1549
|
description: "For action 'models': annotate each model with its Artificial Analysis coding index (extra API call; default false).",
|
|
1518
1550
|
},
|
|
1551
|
+
model: {
|
|
1552
|
+
type: "string",
|
|
1553
|
+
description: "For submit: the OpenRouter batch model for this run (an id ending in :batch, as listed by action 'models'). Falls back to model in .codecarto/broadside/config.yaml, then the shipped default. Pre-flighted like the configured model: priced from the live catalog, refused without structured-output support, clamped to its completion ceiling. The models listing is advisory — some catalog ids have no batch endpoint and are refused at submit, at no cost; the listing tags ids this repository has already seen accepted or refused.",
|
|
1554
|
+
},
|
|
1555
|
+
lens_models: {
|
|
1556
|
+
type: "object",
|
|
1557
|
+
additionalProperties: { type: "string" },
|
|
1558
|
+
description: "For submit: per-lens model overrides for this run, e.g. {\"security\": \"deepseek/deepseek-v4-pro-0813:batch\"}. Keys are lens ids; a lens named here runs on that model, others on `model`. Layered over lens_models in .codecarto/broadside/config.yaml (a lens set in both takes the parameter's). Each override is priced, capability-checked, and clamped individually, and the estimate breaks cost out per lens.",
|
|
1559
|
+
},
|
|
1519
1560
|
},
|
|
1520
1561
|
required: ["cwd", "action"],
|
|
1521
1562
|
},
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "codecartographer-pi",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.22.1",
|
|
4
4
|
"mcpName": "io.github.HuginnIndustries/codecartographer",
|
|
5
5
|
"description": "Turn an unfamiliar codebase into a validated reimplementation spec, then synthesize confirmed specs and a product vision into a traceable plan.",
|
|
6
6
|
"type": "module",
|