codecartographer-pi 0.21.0 → 0.22.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -95,11 +95,14 @@ OpenRouter advertises a `:batch` variant for, and many of those variants do not
95
95
  exist — submitting one returns `does not have a :batch endpoint`, with nothing in
96
96
  the catalog to distinguish it beforehand. A rejected batch costs nothing, so
97
97
  probe a candidate on a single lens first. And reasoning competes with the answer for
98
- `max_tokens`: Broad-Side caps thinking at a quarter of each lens's output budget
99
- so three quarters remain for the JSON, which is the split the cost estimate
100
- already assumes. It caps rather than disables because some endpoints refuse to
101
- be switched off entirely. Override with `reasoning:` in `config.yaml` only
102
- alongside a raised output budget.
98
+ `max_tokens`: Broad-Side asks every model for low reasoning effort so the output
99
+ budget stays with the JSON, which is the split the cost estimate already
100
+ assumes. An effort level rather than a token cap, because Gemini 3.x ignores a
101
+ cap (measured: 11,518 thinking tokens under a 5,800 cap) and honours the level;
102
+ a level rather than off, because some endpoints refuse to be switched off. A
103
+ slice that still truncates is retried once with a doubled budget at low effort.
104
+ Override with `reasoning:` in `config.yaml` only alongside a raised output
105
+ budget, and set `effort` or `max_tokens`, never both.
103
106
 
104
107
  Lenses do not all have to run on the same model. `lens_models` in `config.yaml`
105
108
  routes individual lenses to their own batch model — the usual reason being that
@@ -108,7 +111,14 @@ architecture map. Each override is priced, capability-checked, and clamped like
108
111
  the default, the submit estimate breaks cost out per lens, and `run-meta.json`
109
112
  records which lens ran on what. No stronger default is shipped: which model is
110
113
  worth the money depends on the repository and the budget, so compare with the
111
- `models` action and decide.
114
+ `models` action and decide — for one run with the `model` and `lens_models`
115
+ parameters (Pi: `--model=ID`, `--lens-model=LENS:ID`), or here for the
116
+ repository. The `models` listing is advisory: OpenRouter's catalog returns a
117
+ `:batch` id for some models its Batch API refuses (`does not have a :batch
118
+ endpoint`), at no cost, and nothing in the catalog tells them apart. The
119
+ listing tags the ids this repository's own submits have seen accepted or
120
+ refused (`broadside/batch-endpoints.json`), and a refused lens says why in the
121
+ submit report.
112
122
 
113
123
  Every run knob — `incremental`, `retry_truncated`, `include_synthesis`,
114
124
  `include_triage`, `wait_seconds` — also has a repository default under the same
@@ -54,39 +54,68 @@
54
54
  # Nothing in the catalog distinguishes them. Every Anthropic and OpenAI
55
55
  # batch id tried so far is rejected this way; Google and DeepSeek work.
56
56
  # A rejected batch costs nothing, so probe a candidate on one lens before
57
- # relying on it.
57
+ # relying on it: `codecarto_broadside {action: "submit", lenses:
58
+ # ["architecture"], model: "<id>"}` (Pi: `/codecarto-broadside architecture
59
+ # --model=<id>`). Submits remember the answer in batch-endpoints.json next
60
+ # to this file, and the `models` listing tags each id accordingly.
58
61
  # 2. A reasoning-capable model spends its output budget thinking, and the
59
62
  # thinking is billed at the full output rate. See `reasoning:` below.
60
63
  #
64
+ # The same two keys are submit parameters for one run — `model` and
65
+ # `lens_models` on codecarto_broadside, `--model=ID` and `--lens-model=LENS:ID`
66
+ # on /codecarto-broadside — and a lens set both there and here takes the
67
+ # parameter's.
68
+ #
61
69
  # lens_models:
62
70
  # security: deepseek/deepseek-v4-pro-0813:batch
63
71
  # defect: deepseek/deepseek-v4-pro-0813:batch
64
72
 
65
73
  # Reasoning control, sent on every lens request.
66
74
  #
67
- # By default Broad-Side caps thinking at a quarter of the lens's output budget,
68
- # leaving the other three quarters for the answer — which is exactly the split
69
- # the cost estimate already assumes.
70
- #
71
- # The cap exists because reasoning competes with the answer for `max_tokens`.
72
- # One measured run spent 5,758 of a 6,000-token budget thinking and left ~230
73
- # tokens for the JSON, which truncated mid-structure on 11 of 13 slices; those
74
- # tokens bill at the full output rate, so it paid for ~6,000 output tokens per
75
- # slice to receive ~230 usable ones. The shipped default model does the same
76
- # thing less consistently — reasoning from 0 to 5,757 tokens across 13 slices,
77
- # three of them cut off — so this is not something only exotic models do.
75
+ # By default Broad-Side asks every model for low reasoning effort
76
+ # (`effort: low`), leaving the output budget for the answer — which is the
77
+ # split the cost estimate already assumes.
78
+ #
79
+ # The control exists because reasoning competes with the answer for
80
+ # `max_tokens`. One measured run spent 5,758 of a 6,000-token budget thinking
81
+ # and left ~230 tokens for the JSON, which truncated mid-structure on 11 of 13
82
+ # slices; those tokens bill at the full output rate, so it paid for ~6,000
83
+ # output tokens per slice to receive ~230 usable ones. The shipped default
84
+ # model does the same thing less consistently — reasoning from 0 to 5,757
85
+ # tokens across 13 slices, three of them cut off — so this is not something
86
+ # only exotic models do.
87
+ #
88
+ # It is an effort level rather than a token cap because a cap is not honoured
89
+ # everywhere. Gemini 3.x models take a thinking *level*, not a budget: under
90
+ # `max_tokens: 5800`, google/gemini-3.8-flash:batch reasoned 5,218 tokens on
91
+ # the first pass and 11,518 on the doubled-budget retry — the cap changed
92
+ # nothing, both results truncated, and the retry cost twice the original for
93
+ # no JSON. The same lens at `effort: low` reasoned 0 tokens, returned valid
94
+ # JSON, and cost a twelfth as much. `effort` is what OpenRouter can translate
95
+ # for every provider; a token cap reaches only the providers that take one.
96
+ #
97
+ # And it is a level rather than an off switch on purpose. Some endpoints refuse
98
+ # to be switched off: `google/gemini-3.8-flash:batch` rejects the entire batch
99
+ # with "Reasoning is mandatory for this endpoint and cannot be disabled", which
100
+ # turns a partial result into none at all. Low effort works either way.
101
+ #
102
+ # A slice that still truncates is re-submitted once with a doubled `max_tokens`
103
+ # and, whatever this block says, low effort — the cutoff is usually thinking.
104
+ #
105
+ # Set `effort` OR `max_tokens`, not both: OpenRouter refuses a request that
106
+ # carries both ("Only one of reasoning.effort and reasoning.max_tokens can be
107
+ # specified") — per request, after the batch is accepted, so every lens fails
108
+ # at $0. Submit refuses a config.yaml that sets both. Raise the effort only
109
+ # together with a raised lens output budget, or the JSON truncates exactly as
110
+ # above.
78
111
  #
79
- # It is a cap rather than an off switch on purpose. Some endpoints refuse to be
80
- # switched off: `google/gemini-3.8-flash:batch` rejects the entire batch with
81
- # "Reasoning is mandatory for this endpoint and cannot be disabled", which turns
82
- # a partial result into none at all. Capping works either way.
112
+ # reasoning:
113
+ # effort: low # minimal | low | medium | high
83
114
  #
84
- # Override only with a raised lens output budget, or the JSON truncates exactly
85
- # as above.
115
+ # reasoning:
116
+ # max_tokens: 2000 # a thinking budget, where the provider honours one
86
117
  #
87
118
  # reasoning:
88
- # effort: low # minimal | low | medium | high
89
- # max_tokens: 2000 # or set the thinking budget directly
90
119
  # enabled: false # only where the provider allows it
91
120
 
92
121
  # Approximate run expense limit in USD. The default is 1.00, and it applies
@@ -3,4 +3,4 @@
3
3
  # workspace's framework-owned files (GUIDE.md, templates/, workflow/ pipelines
4
4
  # and VALIDATE.md) predate the running release. Written at release time and
5
5
  # copied verbatim by init — never edit by hand.
6
- scaffold_version: 0.21.0
6
+ scaffold_version: 0.22.1
package/README.md CHANGED
@@ -398,9 +398,9 @@ codecarto_broadside {cwd, action: "collect"} # poll, save, synt
398
398
 
399
399
  Submit and collect are separate because batch jobs routinely take tens of minutes; collect is resumable and picks up whatever is still in flight (`wait_seconds: 0`, the default, polls once and returns). Submit prices the run from the collected file sizes against the model's live per-token pricing (cached 24h) and refuses when the estimate exceeds `max_cost` — $1.00 unless the config or the call sets another value, `0` for no limit — unless `force: true` is passed — a pre-flight estimate, not a runtime stop. Actual spend lands in each run's `run-meta.json`.
400
400
 
401
- Repository defaults live in `.codecarto/broadside/config.yaml` (`model`, `api_key`, `default_lenses`, `max_cost`, `pricing` overrides, `lens_models`, `incremental`, `retry_truncated`, `include_synthesis`, `include_triage`, `wait_seconds`); an explicit tool parameter always wins. `lens_models` routes individual lenses to their own batch model — a stronger model changes security and defect findings far more than it changes an architecture map — and each override is priced, capability-checked, and clamped exactly like the default, with the estimate broken out per lens so a mixed-model run cannot be approved without seeing which lens costs what. CodeCartographer ships no stronger default: which model earns its price depends on your repository and budget, so compare candidates with the `models` action and choose. `codecarto_skill {cwd, name: "broadside"}` returns the reading guide for a completed run, and unlike post-pipeline skills it is not gated on a finished pipeline.
401
+ Repository defaults live in `.codecarto/broadside/config.yaml` (`model`, `api_key`, `default_lenses`, `max_cost`, `pricing` overrides, `lens_models`, `incremental`, `retry_truncated`, `include_synthesis`, `include_triage`, `wait_seconds`); an explicit tool parameter always wins. `lens_models` routes individual lenses to their own batch model — a stronger model changes security and defect findings far more than it changes an architecture map — and each override is priced, capability-checked, and clamped exactly like the default, with the estimate broken out per lens so a mixed-model run cannot be approved without seeing which lens costs what. CodeCartographer ships no stronger default: which model earns its price depends on your repository and budget, so compare candidates with the `models` action and choose — for one run with the `model` and `lens_models` parameters (Pi: `--model=ID`, `--lens-model=LENS:ID`), or for the repository in `config.yaml`. The `models` listing is advisory: OpenRouter's catalog returns a `:batch` id for some models its Batch API then refuses (`does not have a :batch endpoint`), at no cost, and nothing in the catalog tells them apart — so the listing tags the ids this repository's own submits have seen accepted or refused, and a refused lens says why in the submit report. `codecarto_skill {cwd, name: "broadside"}` returns the reading guide for a completed run, and unlike post-pipeline skills it is not gated on a finished pipeline.
402
402
 
403
- On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models] [lenses…]`, with tab-completion for actions and lens names and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
403
+ On the Pi extension the same run is `/codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID]`, with tab-completion for actions, lens names, and flags and live per-lens progress while batches poll. The two surfaces differ in one deliberate place: MCP cannot ask a human, so it refuses a run over `max_cost` until you pass `force`; Pi shows the per-lens breakdown and asks, and your approval *is* the force flag. Neither surface takes an API key as a command argument — a key typed into a slash command lands in the session transcript.
404
404
 
405
405
  Broad-Side needs runtime code, so firing a run is an executable-surface feature: Pi and MCP have it, the pure drop-in template does not (it carries only the reading guide).
406
406
 
@@ -82,6 +82,16 @@ Two more economies worth knowing:
82
82
  conventions is usually a better trade than raising the model for everything.
83
83
  Overrides are priced and capability-checked individually, and the estimate
84
84
  breaks cost out per lens.
85
+ - `model` and `lens_models` are also submit parameters (Pi: `--model=ID`,
86
+ `--lens-model=LENS:ID`), for one run without editing the file. Choose from
87
+ `action: "models"`, and read that listing as advisory: OpenRouter's catalog
88
+ returns a `:batch` id for some models its Batch API then refuses (`does not
89
+ have a :batch endpoint`) — free, reported on the lens with the reason, and
90
+ remembered, so the listing tags ids this repository has seen accepted or
91
+ refused. Probe an untried model on one lens before a six-lens run. A
92
+ `job-submission-count` refusal is the account's concurrent-job quota (one
93
+ job per lens fills it fast across runs); collect or wait out what is in
94
+ flight, then re-submit.
85
95
 
86
96
  ## Reading a run
87
97
 
@@ -11,6 +11,13 @@ export declare const BROADSIDE_OUTPUT_PRICE_PER_M = 1.875;
11
11
  export declare const BROADSIDE_MODELS_URL = "https://openrouter.ai/api/v1/models";
12
12
  export declare const BROADSIDE_BENCHMARKS_URL = "https://openrouter.ai/api/v1/benchmarks";
13
13
  export declare const BROADSIDE_CATALOG_CACHE_FILE = "model-catalog.json";
14
+ /**
15
+ * What this repository's own submits learned about batch endpoints: which
16
+ * `:batch` ids OpenRouter accepted a job for and which it refused with
17
+ * "does not have a :batch endpoint". The catalog cannot tell the two apart
18
+ * (#141), so the `models` action annotates its rows from this file.
19
+ */
20
+ export declare const BROADSIDE_ENDPOINTS_FILE = "batch-endpoints.json";
14
21
  export declare const BROADSIDE_CATALOG_CACHE_TTL_MS: number;
15
22
  export declare const BROADSIDE_LENS_IDS: readonly ["architecture", "api", "security", "defect", "conventions", "porting"];
16
23
  export type BroadsideLensId = (typeof BROADSIDE_LENS_IDS)[number];
@@ -130,31 +137,43 @@ export type BroadsideReasoning = {
130
137
  max_tokens?: number;
131
138
  };
132
139
  /**
133
- * The share of a lens's output budget reasoning may spend.
140
+ * The reasoning control every lens request carries: low effort.
141
+ *
142
+ * It used to be a token cap — `max_tokens` at a quarter of the lens's output
143
+ * budget, so three quarters stayed for the answer. Measured live on
144
+ * `google/gemini-3.8-flash:batch` (0.22.0 verification, defect lens, cap
145
+ * 5,800 of a 6,000 budget): the model reasoned 5,218 tokens on the first
146
+ * pass and **11,518 under the same cap** on the doubled-budget retry —
147
+ * thinking scaled with `max_tokens` and the cap changed nothing, both
148
+ * results truncated, and the retry cost twice the original for no JSON.
149
+ * The same lens with `effort: "low"` reasoned 0 tokens, finished with
150
+ * `stop`, returned valid JSON, and cost a twelfth as much. Gemini 3.x
151
+ * models take a thinking *level*, not a budget, and OpenRouter forwards a
152
+ * `max_tokens` cap to them as nothing at all; `effort` is what it can
153
+ * translate for every provider (a level where the provider has levels, a
154
+ * fraction of the budget where it takes a budget). So the default asks for
155
+ * little thinking in the one vocabulary that reaches everyone.
134
156
  *
135
- * `estimateCost` already budgets output at 75% of `maxTokens`; capping thinking
136
- * at the remaining quarter makes that assumption true by construction and
137
- * guarantees the answer has room. A floor keeps the cap sane for a small lens.
157
+ * Deliberately not `enabled: false`: `google/gemini-3.8-flash:batch` refuses
158
+ * the whole batch with *"Reasoning is mandatory for this endpoint and cannot
159
+ * be disabled"*, turning a partial result into none at all. Low effort works
160
+ * whether or not a provider allows reasoning to be switched off.
138
161
  */
139
- export declare const BROADSIDE_REASONING_BUDGET_FRACTION = 0.25;
140
- export declare const BROADSIDE_MIN_REASONING_TOKENS = 512;
162
+ export declare const BROADSIDE_DEFAULT_REASONING: Readonly<BroadsideReasoning>;
163
+ /** The reasoning control a lens request carries when config.yaml sets none. */
164
+ export declare function defaultReasoningFor(): BroadsideReasoning;
141
165
  /**
142
- * Cap reasoning for a lens request — deliberately a cap, not an off switch.
166
+ * The reasoning control a truncated slice is re-submitted with.
143
167
  *
144
- * Disabling outright is not portable: `google/gemini-3.8-flash:batch` refuses
145
- * the whole batch with *"Reasoning is mandatory for this endpoint and cannot be
146
- * disabled"*, turning a partial result into none at all. Capping works whether
147
- * or not a provider allows reasoning to be switched off.
148
- *
149
- * The failure this prevents is the budget being spent thinking rather than
150
- * answering. Measured on one run: 5,758 of a 6,000-token budget went to
151
- * reasoning, leaving ~230 tokens for JSON that truncated mid-structure — and
152
- * those tokens bill at the full output rate. The shipped default model does the
153
- * same thing less consistently (reasoning tokens from 0 to 5,757 across 13
154
- * slices, three of them cut off at `finish_reason: length`), so this is not a
155
- * multi-model concern.
168
+ * A truncation on a reasoning-capable model is usually thinking that ate the
169
+ * answer's budget, and doubling `max_tokens` doubles the thinking where the
170
+ * provider ignores a token cap (see {@link BROADSIDE_DEFAULT_REASONING}). The
171
+ * retry therefore asks for low effort as well, replacing a `max_tokens` cap
172
+ * (OpenRouter refuses a request carrying both) and lowering a higher effort.
173
+ * An explicit `enabled: false` and an effort already at or below low are left
174
+ * as they are.
156
175
  */
157
- export declare function defaultReasoningFor(maxTokens: number): BroadsideReasoning;
176
+ export declare function retryReasoningFor(original: BroadsideReasoning | undefined): BroadsideReasoning;
158
177
  export type BatchRequest = {
159
178
  custom_id: string;
160
179
  body: {
@@ -182,6 +201,8 @@ export type BroadsideBatchEntry = {
182
201
  cost?: number;
183
202
  resultCount?: number;
184
203
  error?: unknown;
204
+ /** Why a `skipped` lens had nothing to submit: the globs that matched no file. */
205
+ reason?: string;
185
206
  /** Set when this lens used a model other than the run default. */
186
207
  model?: string;
187
208
  /** The completion ceiling of this lens's model; bounds the truncation retry. */
@@ -539,6 +560,26 @@ export declare function loadBroadsideConfig(broadsideDir: string): Promise<Broad
539
560
  export declare function defaultBroadsideConfig(): BroadsideConfig;
540
561
  /** The catalog cache schema this build writes; a file from another is not read. */
541
562
  export declare const BROADSIDE_CATALOG_CACHE_SCHEMA = 3;
563
+ /** One model's most recent submit outcome, as remembered in {@link BROADSIDE_ENDPOINTS_FILE}. */
564
+ export type BatchEndpointRecord = {
565
+ status: "accepted" | "rejected";
566
+ /** ISO timestamp of the submit that produced this record. */
567
+ at: string;
568
+ /** The provider's refusal, for a rejected endpoint. */
569
+ error?: string;
570
+ };
571
+ export declare function readBatchEndpoints(broadsideDir: string): Promise<Record<string, BatchEndpointRecord>>;
572
+ /**
573
+ * Remember what a submit learned about each model it posted to. An accepted
574
+ * job proves the endpoint exists; a "does not have a :batch endpoint"
575
+ * refusal proves it does not. Any other rejection (quota, malformed request,
576
+ * auth) says nothing about the endpoint and leaves the record alone.
577
+ */
578
+ export declare function recordBatchEndpoints(broadsideDir: string, outcomes: Array<{
579
+ model: string;
580
+ batchId: string;
581
+ error?: unknown;
582
+ }>): Promise<void>;
542
583
  export declare function builtInCatalogEntry(model: string): CatalogEntry | null;
543
584
  export declare function builtInPricing(model: string): ModelPricing | null;
544
585
  export declare function resolveCatalogEntry(broadsideDir: string, config: BroadsideConfig, model: string, apiKey: string, fetcher?: FetchLike): Promise<BroadsideCatalogResult>;
@@ -552,6 +593,8 @@ export declare function listBatchModels(broadsideDir: string, config: BroadsideC
552
593
  source: string;
553
594
  benchmarks: CodingBenchmarks | null;
554
595
  defaultModel: string;
596
+ /** This repository's remembered submit outcomes per model, from {@link BROADSIDE_ENDPOINTS_FILE}. */
597
+ endpoints: Record<string, BatchEndpointRecord>;
555
598
  }>;
556
599
  export type FetchLike = (url: string, init: Record<string, unknown>) => Promise<Response>;
557
600
  export declare function submitBatch(batchRequests: BatchRequest[], apiKey: string, fetcher?: FetchLike, model?: string): Promise<{
@@ -602,6 +645,12 @@ export declare function runBroadsideSubmit(cwd: string, apiKey: string, opts?: {
602
645
  lenses?: BroadsideLensId[];
603
646
  fetcher?: FetchLike;
604
647
  model?: string;
648
+ /**
649
+ * Per-lens model overrides for this run, layered over config.yaml's
650
+ * `lens_models`: a lens named here runs on this model, a lens named only
651
+ * in the file runs on the file's, and the rest run on `model` (#141).
652
+ */
653
+ lensModels?: Partial<Record<BroadsideLensId, string>>;
605
654
  /** Approximate run expense limit in USD; 0 means no limit. */
606
655
  maxCost?: number;
607
656
  /** Submit even when the estimate exceeds maxCost. */
@@ -669,11 +718,33 @@ export declare function runBroadsideStatus(cwd: string): Promise<{
669
718
  state: BroadsideStateFile;
670
719
  }>;
671
720
  export declare function renderFindingsMarkdown(content: string): string;
721
+ export declare function describeIncrementalFallback(reason: BroadsideIncrementalOutcome["reason"]): string;
672
722
  export declare function estimateSubmitText(result: BroadsideSubmitResult, lenses: LensDefinition[]): string;
673
723
  export declare function modelsText(entries: CatalogEntry[], opts: {
674
724
  benchmarks: CodingBenchmarks | null;
675
725
  defaultModel: string;
726
+ endpoints?: Record<string, BatchEndpointRecord>;
676
727
  }): string;
728
+ /**
729
+ * A provider refusal plus what to do about it, for the two refusals a batch
730
+ * run meets in practice and cannot fix by itself (#141):
731
+ *
732
+ * - `Model '<id>' does not have a :batch endpoint.` — the catalog advertises a
733
+ * `:batch` id that OpenRouter runs no batch endpoint for. Nothing in the
734
+ * catalog distinguishes these; the `models` action marks ids this
735
+ * repository has seen refused.
736
+ * - `job-submission-count … in use: 16, quota: 16` — the per-account limit
737
+ * on concurrent batch jobs. Broad-Side submits one job per lens, so a few
738
+ * runs in flight on the same key fill it; the refusal costs nothing.
739
+ */
740
+ export declare function explainBatchError(error: unknown): string | null;
677
741
  export declare function collectResultText(result: BroadsideCollectResult): string;
742
+ /**
743
+ * An `onStatus` callback that appends one line to `lines` per *change* of a
744
+ * lens's polled status. Every poll used to append a line, so a four-minute
745
+ * wait returned twenty-six identical "in_progress (0/1)" lines per lens
746
+ * before the result (0.22.0 live run).
747
+ */
748
+ export declare function statusLineWriter(lines: string[]): (lensId: string, status: string, counts: Record<string, unknown>) => void;
678
749
  export declare function statusText(state: BroadsideStateFile): string;
679
750
  export {};
@@ -65,6 +65,13 @@ export const BROADSIDE_OUTPUT_PRICE_PER_M = 1.875;
65
65
  export const BROADSIDE_MODELS_URL = "https://openrouter.ai/api/v1/models";
66
66
  export const BROADSIDE_BENCHMARKS_URL = "https://openrouter.ai/api/v1/benchmarks";
67
67
  export const BROADSIDE_CATALOG_CACHE_FILE = "model-catalog.json";
68
+ /**
69
+ * What this repository's own submits learned about batch endpoints: which
70
+ * `:batch` ids OpenRouter accepted a job for and which it refused with
71
+ * "does not have a :batch endpoint". The catalog cannot tell the two apart
72
+ * (#141), so the `models` action annotates its rows from this file.
73
+ */
74
+ export const BROADSIDE_ENDPOINTS_FILE = "batch-endpoints.json";
68
75
  export const BROADSIDE_CATALOG_CACHE_TTL_MS = 24 * 60 * 60 * 1000;
69
76
  export const BROADSIDE_LENS_IDS = [
70
77
  "architecture",
@@ -87,32 +94,51 @@ export const BROADSIDE_DEFAULT_POLL_BUDGET_MS = 25 * 60 * 1000;
87
94
  */
88
95
  export const BROADSIDE_DEFAULT_MAX_COST = 1;
89
96
  /**
90
- * The share of a lens's output budget reasoning may spend.
97
+ * The reasoning control every lens request carries: low effort.
98
+ *
99
+ * It used to be a token cap — `max_tokens` at a quarter of the lens's output
100
+ * budget, so three quarters stayed for the answer. Measured live on
101
+ * `google/gemini-3.8-flash:batch` (0.22.0 verification, defect lens, cap
102
+ * 5,800 of a 6,000 budget): the model reasoned 5,218 tokens on the first
103
+ * pass and **11,518 under the same cap** on the doubled-budget retry —
104
+ * thinking scaled with `max_tokens` and the cap changed nothing, both
105
+ * results truncated, and the retry cost twice the original for no JSON.
106
+ * The same lens with `effort: "low"` reasoned 0 tokens, finished with
107
+ * `stop`, returned valid JSON, and cost a twelfth as much. Gemini 3.x
108
+ * models take a thinking *level*, not a budget, and OpenRouter forwards a
109
+ * `max_tokens` cap to them as nothing at all; `effort` is what it can
110
+ * translate for every provider (a level where the provider has levels, a
111
+ * fraction of the budget where it takes a budget). So the default asks for
112
+ * little thinking in the one vocabulary that reaches everyone.
91
113
  *
92
- * `estimateCost` already budgets output at 75% of `maxTokens`; capping thinking
93
- * at the remaining quarter makes that assumption true by construction and
94
- * guarantees the answer has room. A floor keeps the cap sane for a small lens.
114
+ * Deliberately not `enabled: false`: `google/gemini-3.8-flash:batch` refuses
115
+ * the whole batch with *"Reasoning is mandatory for this endpoint and cannot
116
+ * be disabled"*, turning a partial result into none at all. Low effort works
117
+ * whether or not a provider allows reasoning to be switched off.
95
118
  */
96
- export const BROADSIDE_REASONING_BUDGET_FRACTION = 0.25;
97
- export const BROADSIDE_MIN_REASONING_TOKENS = 512;
119
+ export const BROADSIDE_DEFAULT_REASONING = Object.freeze({ effort: "low" });
120
+ /** The reasoning control a lens request carries when config.yaml sets none. */
121
+ export function defaultReasoningFor() {
122
+ return { ...BROADSIDE_DEFAULT_REASONING };
123
+ }
98
124
  /**
99
- * Cap reasoning for a lens request — deliberately a cap, not an off switch.
100
- *
101
- * Disabling outright is not portable: `google/gemini-3.8-flash:batch` refuses
102
- * the whole batch with *"Reasoning is mandatory for this endpoint and cannot be
103
- * disabled"*, turning a partial result into none at all. Capping works whether
104
- * or not a provider allows reasoning to be switched off.
125
+ * The reasoning control a truncated slice is re-submitted with.
105
126
  *
106
- * The failure this prevents is the budget being spent thinking rather than
107
- * answering. Measured on one run: 5,758 of a 6,000-token budget went to
108
- * reasoning, leaving ~230 tokens for JSON that truncated mid-structure — and
109
- * those tokens bill at the full output rate. The shipped default model does the
110
- * same thing less consistently (reasoning tokens from 0 to 5,757 across 13
111
- * slices, three of them cut off at `finish_reason: length`), so this is not a
112
- * multi-model concern.
127
+ * A truncation on a reasoning-capable model is usually thinking that ate the
128
+ * answer's budget, and doubling `max_tokens` doubles the thinking where the
129
+ * provider ignores a token cap (see {@link BROADSIDE_DEFAULT_REASONING}). The
130
+ * retry therefore asks for low effort as well, replacing a `max_tokens` cap
131
+ * (OpenRouter refuses a request carrying both) and lowering a higher effort.
132
+ * An explicit `enabled: false` and an effort already at or below low are left
133
+ * as they are.
113
134
  */
114
- export function defaultReasoningFor(maxTokens) {
115
- return { max_tokens: Math.max(BROADSIDE_MIN_REASONING_TOKENS, Math.floor(maxTokens * BROADSIDE_REASONING_BUDGET_FRACTION)) };
135
+ export function retryReasoningFor(original) {
136
+ if (original?.enabled === false)
137
+ return { ...original };
138
+ if (original?.effort === "minimal" || original?.effort === "low")
139
+ return { ...original };
140
+ const { max_tokens: _cap, effort: _effort, ...rest } = original ?? {};
141
+ return { ...rest, effort: "low" };
116
142
  }
117
143
  /**
118
144
  * OpenRouter rejected the API key (HTTP 401/403). Thrown from the catalog
@@ -1341,7 +1367,7 @@ export function buildBatchRequest(lens, info, slice, index, sliceCount, model =
1341
1367
  max_tokens: maxTokensOverride ?? lens.maxTokens,
1342
1368
  // Always sent, never inherited: an absent field means the model's
1343
1369
  // own default, and that default is what truncated the JSON.
1344
- reasoning: reasoningOverride ?? lens.reasoning ?? defaultReasoningFor(maxTokensOverride ?? lens.maxTokens),
1370
+ reasoning: reasoningOverride ?? lens.reasoning ?? defaultReasoningFor(),
1345
1371
  },
1346
1372
  };
1347
1373
  }
@@ -1542,6 +1568,21 @@ export async function loadBroadsideConfig(broadsideDir) {
1542
1568
  throw new BroadsideConfigError(configPath, "is not a YAML mapping");
1543
1569
  raw = parsed;
1544
1570
  }
1571
+ // OpenRouter accepts `reasoning.effort` or `reasoning.max_tokens`, not
1572
+ // both: a request carrying both is refused per request *after* the batch
1573
+ // is accepted, so every lens fails at $0 with the reason in each
1574
+ // result's error. Seen live on 0.22.0 with the two keys set together.
1575
+ // Refuse here, where the file can be fixed, rather than submit a run
1576
+ // that cannot produce a result.
1577
+ const reasoning = raw.reasoning;
1578
+ if (reasoning && typeof reasoning === "object" && !Array.isArray(reasoning)) {
1579
+ const value = reasoning;
1580
+ const hasEffort = typeof value.effort === "string";
1581
+ const hasBudget = typeof value.max_tokens === "number" && value.max_tokens > 0;
1582
+ if (hasEffort && hasBudget) {
1583
+ throw new BroadsideConfigError(configPath, 'sets both reasoning.effort and reasoning.max_tokens; OpenRouter accepts one or the other ("Only one of reasoning.effort and reasoning.max_tokens can be specified"), and every lens request would fail after the batch is accepted. Keep one');
1584
+ }
1585
+ }
1545
1586
  }
1546
1587
  return buildBroadsideConfig(raw);
1547
1588
  }
@@ -1621,6 +1662,67 @@ async function writeCatalogCache(broadsideDir, cache) {
1621
1662
  await mkdir(broadsideDir, { recursive: true });
1622
1663
  await writeFile(join(broadsideDir, BROADSIDE_CATALOG_CACHE_FILE), `${JSON.stringify(cache, null, "\t")}\n`, "utf8");
1623
1664
  }
1665
+ const BROADSIDE_ENDPOINTS_SCHEMA = 1;
1666
+ export async function readBatchEndpoints(broadsideDir) {
1667
+ const path = join(broadsideDir, BROADSIDE_ENDPOINTS_FILE);
1668
+ if (!(await pathExists(path)))
1669
+ return {};
1670
+ try {
1671
+ const parsed = JSON.parse(await readFile(path, "utf8"));
1672
+ if (!parsed || typeof parsed !== "object" || parsed.schema_version !== BROADSIDE_ENDPOINTS_SCHEMA)
1673
+ return {};
1674
+ if (!parsed.models || typeof parsed.models !== "object")
1675
+ return {};
1676
+ const out = {};
1677
+ for (const [model, record] of Object.entries(parsed.models)) {
1678
+ if (!record || typeof record !== "object")
1679
+ continue;
1680
+ if (record.status !== "accepted" && record.status !== "rejected")
1681
+ continue;
1682
+ if (typeof record.at !== "string")
1683
+ continue;
1684
+ out[model] = { status: record.status, at: record.at, ...(typeof record.error === "string" && { error: record.error }) };
1685
+ }
1686
+ return out;
1687
+ }
1688
+ catch {
1689
+ // An unreadable memory is an empty one: it only annotates a listing.
1690
+ return {};
1691
+ }
1692
+ }
1693
+ /**
1694
+ * The refusal OpenRouter returns for a catalog id that has no batch endpoint
1695
+ * behind it. Matched loosely: the message is the only signal there is.
1696
+ */
1697
+ const NO_BATCH_ENDPOINT_RE = /does not have a :batch endpoint/i;
1698
+ /** The refusal for a full per-account concurrent batch-job quota. */
1699
+ const BATCH_QUOTA_RE = /job-submission-count/i;
1700
+ /**
1701
+ * Remember what a submit learned about each model it posted to. An accepted
1702
+ * job proves the endpoint exists; a "does not have a :batch endpoint"
1703
+ * refusal proves it does not. Any other rejection (quota, malformed request,
1704
+ * auth) says nothing about the endpoint and leaves the record alone.
1705
+ */
1706
+ export async function recordBatchEndpoints(broadsideDir, outcomes) {
1707
+ const at = new Date().toISOString();
1708
+ const updates = {};
1709
+ for (const { model, batchId, error } of outcomes) {
1710
+ if (batchId) {
1711
+ updates[model] = { status: "accepted", at };
1712
+ continue;
1713
+ }
1714
+ const message = describeBatchError(error);
1715
+ if (message && NO_BATCH_ENDPOINT_RE.test(message)) {
1716
+ updates[model] = { status: "rejected", at, error: message };
1717
+ }
1718
+ }
1719
+ if (Object.keys(updates).length === 0)
1720
+ return;
1721
+ const models = { ...(await readBatchEndpoints(broadsideDir)), ...updates };
1722
+ await mkdir(broadsideDir, { recursive: true });
1723
+ const file = { schema_version: BROADSIDE_ENDPOINTS_SCHEMA, models };
1724
+ await atomicWriteFile(join(broadsideDir, BROADSIDE_ENDPOINTS_FILE), `${JSON.stringify(file, null, "\t")}\n`);
1725
+ }
1624
1726
  function parseCatalogEntry(raw) {
1625
1727
  const id = String(raw.id ?? "");
1626
1728
  if (!id)
@@ -1843,7 +1945,8 @@ export async function listBatchModels(broadsideDir, config, apiKey, opts = {}) {
1843
1945
  cache.models[entry.id] = { ...entry, fetched_at: fetchedAt };
1844
1946
  await writeCatalogCache(broadsideDir, cache);
1845
1947
  const benchmarks = opts.includeBenchmarks ? await fetchCodingBenchmarks(apiKey, fetcher) : null;
1846
- return { entries, source: "live", benchmarks, defaultModel: config.model };
1948
+ const endpoints = await readBatchEndpoints(broadsideDir);
1949
+ return { entries, source: "live", benchmarks, defaultModel: config.model, endpoints };
1847
1950
  }
1848
1951
  export async function submitBatch(batchRequests, apiKey, fetcher = fetch, model = BROADSIDE_MODEL) {
1849
1952
  // The OpenRouter batch endpoint stream-parses the body and requires
@@ -2001,7 +2104,8 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
2001
2104
  throw new Error(`Broad-Side found no ${info.language} source files to scan (detected from ${info.manifest?.path ?? "the file counts"}; ` +
2002
2105
  `the lenses look for ${info.sourceExts.join(", ")}). Nothing was submitted.`);
2003
2106
  }
2004
- const modelForLens = (lensId) => config.lensModels[lensId] ?? model;
2107
+ const lensModels = { ...config.lensModels, ...opts.lensModels };
2108
+ const modelForLens = (lensId) => lensModels[lensId] ?? model;
2005
2109
  const resolved = new Map();
2006
2110
  for (const candidate of new Set([model, ...lensIds.map(modelForLens)])) {
2007
2111
  const catalog = await resolveCatalogEntry(broadsideDir, config, candidate, apiKey, opts.fetcher);
@@ -2062,6 +2166,8 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
2062
2166
  }
2063
2167
  // Slice offline first so the estimate covers every request we would send.
2064
2168
  const slicesByLens = new Map();
2169
+ // Why a lens ended up with nothing to submit, for the report (see below).
2170
+ const skipReasons = new Map();
2065
2171
  let estimatedInputTokens = 0;
2066
2172
  let estimatedOutputTokens = 0;
2067
2173
  let estimatedTotalCost = 0;
@@ -2079,12 +2185,21 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
2079
2185
  for (const file of slice.redactedFiles ?? [])
2080
2186
  redactedFiles.add(file);
2081
2187
  }
2188
+ const matchedBeforeIncremental = slices.length;
2082
2189
  if (changed) {
2083
2190
  // Repo-info slices (empty files, e.g. architecture) always run;
2084
2191
  // file-backed slices run only when one of their files changed.
2085
2192
  slices = slices.filter((s) => s.files.length === 0 || s.files.some((f) => changed.has(f)));
2086
2193
  }
2087
2194
  slicesByLens.set(lensId, slices);
2195
+ if (slices.length === 0) {
2196
+ const globs = lens.globsFor(info).filter(Boolean);
2197
+ skipReasons.set(lensId, globs.length === 0
2198
+ ? "the lens has no file patterns for this language"
2199
+ : matchedBeforeIncremental > 0
2200
+ ? "incremental: none of this lens's files changed since the previous run"
2201
+ : `no files matched ${globs.join(", ")}${lens.skipTestFiles ? " (test files excluded)" : ""}`);
2202
+ }
2088
2203
  const lensModel = modelForLens(lensId);
2089
2204
  const { pricing: lensPricing, outputCap: lensOutputCap } = resolved.get(lensModel);
2090
2205
  const maxTokens = lensOutputCap ? Math.min(lens.maxTokens, lensOutputCap) : lens.maxTokens;
@@ -2193,7 +2308,15 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
2193
2308
  if (requests.length === 0) {
2194
2309
  // No files matched the lens's globs. That is a coverage gap to
2195
2310
  // report, not a batch to submit — the API rejects empty batches.
2311
+ // Name the globs: a JavaScript service whose server lives at
2312
+ // src/server.js gets no security review (that lens reads server/**,
2313
+ // **/auth*, **/middleware/**), and "skipped (0 request(s))" alone
2314
+ // read as an empty repository rather than a lens that looked in
2315
+ // the wrong place.
2196
2316
  entry.status = "skipped";
2317
+ const reason = skipReasons.get(lensId);
2318
+ if (reason)
2319
+ entry.reason = reason;
2197
2320
  continue;
2198
2321
  }
2199
2322
  submissions.push((async () => {
@@ -2214,7 +2337,19 @@ export async function runBroadsideSubmit(cwd, apiKey, opts = {}) {
2214
2337
  })());
2215
2338
  }
2216
2339
  await Promise.allSettled(submissions);
2340
+ // A run with no batch behind it has nothing in flight. Every lens was
2341
+ // skipped or refused, so no poll will ever complete it; leaving it
2342
+ // "in-flight" had status listing a refused run above the completed ones
2343
+ // with synthesis and triage "pending" forever.
2344
+ if (!Object.values(run.batches).some((entry) => entry.batchId))
2345
+ run.status = "failed";
2217
2346
  await persistBroadsideRun(broadsideDir, run);
2347
+ // What the provider just said about each model's batch endpoint outlives
2348
+ // the run: the `models` action reads it back (#141).
2349
+ await recordBatchEndpoints(broadsideDir, lensIds
2350
+ .map((lensId) => run.batches[lensId])
2351
+ .filter((entry) => Boolean(entry) && entry.status !== "skipped")
2352
+ .map((entry) => ({ model: entry.model ?? model, batchId: entry.batchId, error: entry.error })));
2218
2353
  // Persist the exact request bodies so collect can re-submit a truncated
2219
2354
  // slice (bumped output cap) without re-walking the repo (#133). The run
2220
2355
  // dir is created here rather than waiting for collect so a crash between
@@ -2520,7 +2655,19 @@ export async function runBroadsideCollect(cwd, apiKey, opts = {}) {
2520
2655
  truncatedCount += truncated;
2521
2656
  totalCost += cost ?? 0;
2522
2657
  await writeFile(join(runDir, `raw-${lensId}.json`), `${JSON.stringify(batch, null, "\t")}\n`, "utf8");
2523
- lensOutcomes[lensId] = { status, cost: entry.cost, resultCount: entry.resultCount, truncated };
2658
+ // A batch can complete with every request failed — the account's
2659
+ // concurrent-job quota filling after acceptance does exactly this.
2660
+ // The per-request errors are on disk as `<id>.error.json`, but a
2661
+ // lens reporting "completed, 0 result(s)" with the reason buried
2662
+ // there read as an empty repository rather than a refused run.
2663
+ const results = Array.isArray(batch.results) ? batch.results : [];
2664
+ const failed = results.filter((r) => r.error && extractContent(r) === null);
2665
+ const allFailed = stored.length === 0 && failed.length > 0
2666
+ ? `all ${failed.length} request(s) failed: ${explainBatchError(failed[0].error)}`
2667
+ : null;
2668
+ if (allFailed)
2669
+ entry.error = allFailed;
2670
+ lensOutcomes[lensId] = { status, cost: entry.cost, resultCount: entry.resultCount, truncated, ...(allFailed && { error: allFailed }) };
2524
2671
  }
2525
2672
  else {
2526
2673
  // Every non-completed outcome still has to reach the report.
@@ -2532,17 +2679,30 @@ export async function runBroadsideCollect(cwd, apiKey, opts = {}) {
2532
2679
  // indistinguishable in the output from one that was never requested.
2533
2680
  if (batch.error)
2534
2681
  entry.error = batch.error;
2535
- const error = describeBatchError(batch.error);
2682
+ const error = explainBatchError(batch.error);
2536
2683
  lensOutcomes[lensId] = { status, cost: entry.cost, resultCount: entry.resultCount, ...(error && { error }) };
2537
2684
  }
2538
2685
  await persistBroadsideRun(broadsideDir, run);
2539
2686
  }
2540
- // #133: re-submit truncated slices once with a bumped output cap. Batch
2541
- // requests are pure, so re-running is always safe; the aim is to recover
2542
- // coverage the first pass lost to a max_tokens cutoff, not to loop forever.
2687
+ // #133: re-submit truncated slices once with a bumped output cap and low
2688
+ // reasoning effort. Batch requests are pure, so re-running is always safe;
2689
+ // the aim is to recover coverage the first pass lost to a max_tokens
2690
+ // cutoff, not to loop forever. Low effort because the cutoff is usually
2691
+ // thinking, and a doubled budget doubled the thinking where a token cap
2692
+ // was ignored (see retryReasoningFor).
2693
+ //
2694
+ // All bumped requests for one model go out as ONE batch, and the batches
2695
+ // (one per model, since a batch carries a single model) are polled
2696
+ // together against the shared deadline. Each truncated slice used to be
2697
+ // submitted and polled to terminal before the next was submitted, so a
2698
+ // model that truncated 11 of 13 slices turned a five-minute collect into
2699
+ // eleven sequential round trips — the serialization #136 removed from the
2700
+ // lens pass, still present here (#206). Grouping also keeps the retry to
2701
+ // one job per model against OpenRouter's 16-concurrent-job quota.
2543
2702
  let retriedCount = 0;
2544
2703
  if (opts.retryTruncated !== false && truncatedCount > 0) {
2545
2704
  const requestsByCustomId = await loadStoredRequests(runDir);
2705
+ const byModel = new Map();
2546
2706
  for (const stored of allLensResults) {
2547
2707
  if (!stored.truncated)
2548
2708
  continue;
@@ -2559,42 +2719,60 @@ export async function runBroadsideCollect(cwd, apiKey, opts = {}) {
2559
2719
  const bumpedMax = lensCap ? Math.min(previousMax * 2, lensCap) : previousMax * 2;
2560
2720
  if (bumpedMax <= previousMax)
2561
2721
  continue; // already at the ceiling
2562
- const bumped = {
2722
+ const group = byModel.get(lensModel) ?? { requests: [], slices: new Map() };
2723
+ group.requests.push({
2563
2724
  ...original,
2564
- body: { ...original.body, max_tokens: bumpedMax },
2565
- };
2725
+ body: { ...original.body, max_tokens: bumpedMax, reasoning: retryReasoningFor(original.body.reasoning) },
2726
+ });
2727
+ group.slices.set(stored.customId, stored);
2728
+ byModel.set(lensModel, group);
2729
+ }
2730
+ // Submit every group, then poll whatever was accepted, together.
2731
+ const submitted = [];
2732
+ for (const [model, group] of byModel) {
2566
2733
  try {
2567
- const { batchId, error } = await submitBatch([bumped], apiKey, opts.fetcher, lensModel);
2568
- if (error)
2734
+ const { batchId, error } = await submitBatch(group.requests, apiKey, opts.fetcher, model);
2735
+ if (!error && batchId)
2736
+ submitted.push({ model, batchId });
2737
+ }
2738
+ catch {
2739
+ // A retry batch that fails to submit leaves its slices' original
2740
+ // truncated results in place — nothing is lost.
2741
+ }
2742
+ }
2743
+ const polled = await pollBatchesConcurrently(submitted.map(({ model, batchId }) => ({ lensId: `retry:${model}`, batchId })), apiKey, {
2744
+ // Share the caller's deadline. Each of these polls used to start a
2745
+ // fresh 25-minute budget, so `wait_seconds` bounded only the lens
2746
+ // poll and a collect could run for the caller's budget plus fifty
2747
+ // minutes.
2748
+ deadlineMs: Math.max(0, deadline - Date.now()),
2749
+ fetcher: opts.fetcher,
2750
+ onStatus: opts.onStatus,
2751
+ });
2752
+ for (const { model, batchId } of submitted) {
2753
+ const batch = polled.get(batchId);
2754
+ if (!batch || batch.status !== "completed")
2755
+ continue;
2756
+ const group = byModel.get(model);
2757
+ const usage = (batch.usage ?? {});
2758
+ totalCost += typeof usage.cost === "number" ? usage.cost : 0;
2759
+ const results = Array.isArray(batch.results) ? batch.results : [];
2760
+ for (const result of results) {
2761
+ const stored = group.slices.get(String(result.custom_id ?? ""));
2762
+ if (!stored)
2569
2763
  continue;
2570
- const batch = await pollBatchUntilTerminal(batchId, apiKey, {
2571
- // Share the caller's deadline. Each of these polls used to
2572
- // start a fresh 25-minute budget, so `wait_seconds` bounded
2573
- // only the lens poll and a collect could run for the caller's
2574
- // budget plus fifty minutes.
2575
- deadlineMs: Math.max(0, deadline - Date.now()),
2576
- onStatus: (status, counts) => opts.onStatus?.(`${stored.lensId}:retry`, status, counts),
2577
- fetcher: opts.fetcher,
2578
- });
2579
- if (batch.status !== "completed")
2764
+ const content = extractContent(result);
2765
+ if (content === null)
2580
2766
  continue;
2581
- const results = Array.isArray(batch.results) ? batch.results : [];
2582
- const content = results.length > 0 ? extractContent(results[0]) : null;
2583
- if (content === null || parseLensJson(content) === null)
2584
- continue; // still no good
2585
- const usage = (batch.usage ?? {});
2586
- totalCost += typeof usage.cost === "number" ? usage.cost : 0;
2587
2767
  const parsed = parseLensJson(content);
2768
+ if (parsed === null)
2769
+ continue; // still no good
2588
2770
  await writeFile(join(runDir, `${sanitizeId(stored.customId)}.json`), `${JSON.stringify(parsed, null, "\t")}\n`, "utf8");
2589
2771
  await writeFile(join(runDir, `${sanitizeId(stored.customId)}.md`), renderFindingsMarkdown(content), "utf8");
2590
2772
  stored.content = content;
2591
2773
  stored.truncated = false;
2592
2774
  retriedCount += 1;
2593
2775
  }
2594
- catch {
2595
- // A retry that fails to submit/poll leaves the original
2596
- // truncated result in place — nothing is lost.
2597
- }
2598
2776
  }
2599
2777
  truncatedCount = allLensResults.filter((s) => s.truncated).length;
2600
2778
  for (const [lensId, outcome] of Object.entries(lensOutcomes)) {
@@ -2846,7 +3024,7 @@ function parseSynthesisTopFindings(content) {
2846
3024
  }
2847
3025
  }
2848
3026
  // ---------- formatting helpers for tool output ----------
2849
- function describeIncrementalFallback(reason) {
3027
+ export function describeIncrementalFallback(reason) {
2850
3028
  switch (reason) {
2851
3029
  case "dirty-worktree":
2852
3030
  return "the working tree has uncommitted changes, so there is no committed state to diff against";
@@ -2878,7 +3056,16 @@ export function estimateSubmitText(result, lenses) {
2878
3056
  continue;
2879
3057
  const status = entry.batchId ? `batch ${entry.batchId}` : entry.status;
2880
3058
  const override = entry.model ? ` on ${entry.model}` : "";
2881
- lines.push(` ${lens.name}: ${status} (${entry.requests} request(s), ~$${entry.estimatedCost.toFixed(4)})${override}`);
3059
+ // A rejected lens says why: the message is the only way to tell a
3060
+ // catalog id with no batch endpoint from a full job quota, and both
3061
+ // used to read as a bare "rejected". A skipped lens names the globs
3062
+ // that matched nothing.
3063
+ const reason = !entry.batchId && entry.error
3064
+ ? ` — ${explainBatchError(entry.error)}`
3065
+ : !entry.batchId && entry.reason
3066
+ ? ` — ${entry.reason}`
3067
+ : "";
3068
+ lines.push(` ${lens.name}: ${status} (${entry.requests} request(s), ~$${entry.estimatedCost.toFixed(4)})${override}${reason}`);
2882
3069
  }
2883
3070
  if (result.repo) {
2884
3071
  const head = result.repo.sourceHead ? ` at ${result.repo.sourceHead.slice(0, 8)}${result.repo.sourceDirty ? " (dirty)" : ""}` : "";
@@ -2915,8 +3102,15 @@ export function estimateSubmitText(result, lenses) {
2915
3102
  return lines.join("\n");
2916
3103
  }
2917
3104
  export function modelsText(entries, opts) {
3105
+ const endpoints = opts.endpoints ?? {};
2918
3106
  const lines = [
2919
3107
  `Batch models on OpenRouter (${entries.length}, cheapest first).`,
3108
+ // The catalog over-reports: it returns a `:batch` id for models whose
3109
+ // Batch API refuses the job, with nothing in the entry to tell them
3110
+ // apart (#141). Say so before the table, not after it.
3111
+ "Advisory: this is the catalog's list of :batch ids, not a list of working batch endpoints. Some ids are refused at submit " +
3112
+ "(\"does not have a :batch endpoint\"), at no cost. Rows tagged [no batch endpoint …] or [batch OK …] carry what this " +
3113
+ "repository's own submits found; an untagged row has not been tried here.",
2920
3114
  "",
2921
3115
  "id | $/M in | $/M out | ctx | max out | structured | coding idx",
2922
3116
  ];
@@ -2936,12 +3130,21 @@ export function modelsText(entries, opts) {
2936
3130
  const out = entry.maxCompletionTokens ? `${(entry.maxCompletionTokens / 1024).toFixed(0)}k` : "?";
2937
3131
  const tag = entry.id === opts.defaultModel ? " (default)" : "";
2938
3132
  const exp = entry.expirationDate ? " [deprecated]" : "";
2939
- lines.push(`${entry.id}${tag}${exp} | ${entry.inputPerM.toFixed(3)} | ${entry.outputPerM.toFixed(3)} | ${ctx} | ${out} | ${structured} | ${coding}`);
3133
+ const record = endpoints[entry.id];
3134
+ const seen = record
3135
+ ? record.status === "rejected"
3136
+ ? ` [no batch endpoint, refused ${record.at.slice(0, 10)}]`
3137
+ : ` [batch OK ${record.at.slice(0, 10)}]`
3138
+ : "";
3139
+ lines.push(`${entry.id}${tag}${exp}${seen} | ${entry.inputPerM.toFixed(3)} | ${entry.outputPerM.toFixed(3)} | ${ctx} | ${out} | ${structured} | ${coding}`);
2940
3140
  }
2941
3141
  if (opts.benchmarks?.meta.as_of) {
2942
3142
  lines.push("", `Benchmarks: Artificial Analysis coding index (as of ${String(opts.benchmarks.meta.as_of)}).`);
2943
3143
  }
2944
- lines.push("", "Set the batch model in .codecarto/broadside/config.yaml (model key). Higher coding index ≠ better scout: precision, context, and structured-output support matter most here.");
3144
+ lines.push("", "Choose with the model parameter (--model= on Pi) for one run, lens_models (--lens-model=LENS:ID) per lens, or the model key in " +
3145
+ ".codecarto/broadside/config.yaml for the repository. Higher coding index ≠ better scout: precision, context, structured-output " +
3146
+ "support, and whether the model spends its output budget reasoning (see reasoning: in config.yaml) matter most here. " +
3147
+ "A refused submit costs nothing, so probe an untried model on one lens first.");
2945
3148
  return lines.join("\n");
2946
3149
  }
2947
3150
  /** One line of a batch's error field, whatever shape the provider gave it. */
@@ -2954,6 +3157,15 @@ function describeBatchError(error) {
2954
3157
  const message = error.message;
2955
3158
  if (typeof message === "string" && message)
2956
3159
  return message.slice(0, 300);
3160
+ // OpenRouter wraps a submit refusal as `{ error: { message } }`.
3161
+ const nested = error.error;
3162
+ if (nested && typeof nested === "object") {
3163
+ const inner = nested.message;
3164
+ if (typeof inner === "string" && inner)
3165
+ return inner.slice(0, 300);
3166
+ }
3167
+ if (typeof nested === "string" && nested)
3168
+ return nested.slice(0, 300);
2957
3169
  try {
2958
3170
  return JSON.stringify(error).slice(0, 300);
2959
3171
  }
@@ -2963,6 +3175,30 @@ function describeBatchError(error) {
2963
3175
  }
2964
3176
  return String(error);
2965
3177
  }
3178
+ /**
3179
+ * A provider refusal plus what to do about it, for the two refusals a batch
3180
+ * run meets in practice and cannot fix by itself (#141):
3181
+ *
3182
+ * - `Model '<id>' does not have a :batch endpoint.` — the catalog advertises a
3183
+ * `:batch` id that OpenRouter runs no batch endpoint for. Nothing in the
3184
+ * catalog distinguishes these; the `models` action marks ids this
3185
+ * repository has seen refused.
3186
+ * - `job-submission-count … in use: 16, quota: 16` — the per-account limit
3187
+ * on concurrent batch jobs. Broad-Side submits one job per lens, so a few
3188
+ * runs in flight on the same key fill it; the refusal costs nothing.
3189
+ */
3190
+ export function explainBatchError(error) {
3191
+ const message = describeBatchError(error);
3192
+ if (!message)
3193
+ return null;
3194
+ if (NO_BATCH_ENDPOINT_RE.test(message)) {
3195
+ return `${message} — the catalog lists this id, but OpenRouter runs no batch endpoint for it. Nothing was charged; pick another model (the models action marks ids this repository has seen refused).`;
3196
+ }
3197
+ if (BATCH_QUOTA_RE.test(message)) {
3198
+ return `${message} — OpenRouter's per-account limit on concurrent batch jobs is full. Broad-Side submits one job per lens, so a few runs in flight on this key (in any repository) fill it. Nothing was charged; collect or wait out the runs in flight, then re-submit.`;
3199
+ }
3200
+ return message;
3201
+ }
2966
3202
  export function collectResultText(result) {
2967
3203
  const lines = [
2968
3204
  `Broad-Side run ${result.runId}: ${result.status}`,
@@ -3012,6 +3248,22 @@ export function collectResultText(result) {
3012
3248
  lines.push("", "Disclaimer: Broad-Side findings are unverified scouting signals from a batch model, not validated claims.");
3013
3249
  return lines.join("\n");
3014
3250
  }
3251
+ /**
3252
+ * An `onStatus` callback that appends one line to `lines` per *change* of a
3253
+ * lens's polled status. Every poll used to append a line, so a four-minute
3254
+ * wait returned twenty-six identical "in_progress (0/1)" lines per lens
3255
+ * before the result (0.22.0 live run).
3256
+ */
3257
+ export function statusLineWriter(lines) {
3258
+ const last = new Map();
3259
+ return (lensId, status, counts) => {
3260
+ const line = ` ${lensId}: ${status} (${counts.completed ?? 0}/${counts.total ?? "?"})`;
3261
+ if (last.get(lensId) === line)
3262
+ return;
3263
+ last.set(lensId, line);
3264
+ lines.push(line);
3265
+ };
3266
+ }
3015
3267
  export function statusText(state) {
3016
3268
  if (state.runs.length === 0) {
3017
3269
  return "No Broad-Side runs recorded. Call codecarto_broadside with action 'submit' first.";
@@ -3029,7 +3281,8 @@ export function statusText(state) {
3029
3281
  const entry = run.batches[lensId];
3030
3282
  if (!entry)
3031
3283
  continue;
3032
- lines.push(` ${lensId}: ${entry.status}${entry.batchId ? ` (${entry.batchId})` : ""}${entry.cost !== undefined ? `, $${entry.cost.toFixed(6)}` : ""}`);
3284
+ lines.push(` ${lensId}: ${entry.status}${entry.batchId ? ` (${entry.batchId})` : ""}${entry.cost !== undefined ? `, $${entry.cost.toFixed(6)}` : ""}` +
3285
+ (entry.status === "skipped" && entry.reason ? ` — ${entry.reason}` : ""));
3033
3286
  }
3034
3287
  lines.push(` synthesis: ${run.synthesis.status}`);
3035
3288
  lines.push(` triage: ${run.triage?.status ?? "pending"}`);
@@ -18,11 +18,15 @@ export interface BroadsideFlags {
18
18
  waitSeconds?: number;
19
19
  /** For collect: the run to collect instead of the most recent (#268). */
20
20
  runId?: string;
21
+ /** For submit: the run's batch model, replacing config.yaml's (#141). */
22
+ model?: string;
23
+ /** For submit: per-lens model overrides, layered over config.yaml's (#141). */
24
+ lensModels?: Partial<Record<BroadsideLensId, string>>;
21
25
  benchmarks: boolean;
22
26
  unknown: string[];
23
27
  /** Set on an invalid combination. The caller surfaces it as an error. */
24
28
  error?: string;
25
29
  }
26
30
  /** Every token the completer offers, in the order it offers them. */
27
- export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
31
+ export declare const KNOWN_BROADSIDE_TOKENS: readonly ["submit", "collect", "status", "models", "architecture", "api", "security", "defect", "conventions", "porting", "--incremental", "--no-incremental", "--max-cost=", "--wait=", "--run=", "--model=", "--lens-model=", "--no-synthesis", "--no-triage", "--no-retry-truncated", "--benchmarks"];
28
32
  export declare function parseBroadsideFlags(args: string): BroadsideFlags;
@@ -14,6 +14,11 @@
14
14
  // --max-cost=N --no-retry-truncated
15
15
  // --wait=SECONDS --benchmarks (models only)
16
16
  // --run=ID (collect only: an older run, as listed by status)
17
+ // --model=ID (submit only: the run's batch model, as listed by models)
18
+ // --lens-model=LENS:ID (submit only, repeatable: one lens on its own model)
19
+ //
20
+ // A model id itself contains a colon (`vendor/name:batch`), so --lens-model
21
+ // splits on the first colon only: `security:deepseek/deepseek-v4-pro:batch`.
17
22
  //
18
23
  // --incremental has a spelled-out negative because the value is tri-state:
19
24
  // absent defers to config.yaml, so a repository that set `incremental: true`
@@ -36,6 +41,8 @@ export const KNOWN_BROADSIDE_TOKENS = [
36
41
  "--max-cost=",
37
42
  "--wait=",
38
43
  "--run=",
44
+ "--model=",
45
+ "--lens-model=",
39
46
  "--no-synthesis",
40
47
  "--no-triage",
41
48
  "--no-retry-truncated",
@@ -115,6 +122,32 @@ export function parseBroadsideFlags(args) {
115
122
  result.runId = value || undefined;
116
123
  continue;
117
124
  }
125
+ if (token.startsWith("--model=")) {
126
+ const value = token.slice("--model=".length).trim();
127
+ // An empty value is a mistyped selection, not "use the default":
128
+ // the command is about to spend money on whichever model wins.
129
+ if (!value)
130
+ result.error ??= "--model= needs an OpenRouter batch model id (see /codecarto-broadside models).";
131
+ result.model = value || undefined;
132
+ continue;
133
+ }
134
+ if (token.startsWith("--lens-model=")) {
135
+ const value = token.slice("--lens-model=".length).trim();
136
+ const colon = value.indexOf(":");
137
+ const lensId = colon > 0 ? value.slice(0, colon).trim() : "";
138
+ const modelId = colon > 0 ? value.slice(colon + 1).trim() : "";
139
+ if (!lensId || !modelId) {
140
+ result.error ??= `--lens-model needs LENS:MODEL, e.g. --lens-model=security:vendor/name:batch (got "${value}").`;
141
+ }
142
+ else if (!BROADSIDE_LENS_IDS.includes(lensId)) {
143
+ result.error ??= `--lens-model: unknown lens "${lensId}". Lenses: ${BROADSIDE_LENS_IDS.join(", ")}.`;
144
+ }
145
+ else {
146
+ result.lensModels ??= {};
147
+ result.lensModels[lensId] = modelId;
148
+ }
149
+ continue;
150
+ }
118
151
  result.unknown.push(token);
119
152
  }
120
153
  // Flags that only mean something for one action are refused rather than
@@ -137,5 +170,11 @@ export function parseBroadsideFlags(args) {
137
170
  if (result.runId !== undefined && result.action !== "collect") {
138
171
  result.error ??= `--run is only meaningful for collect (got action "${result.action}").`;
139
172
  }
173
+ if (result.model !== undefined && result.action !== "submit") {
174
+ result.error ??= `--model is only meaningful for submit (got action "${result.action}").`;
175
+ }
176
+ if (result.lensModels !== undefined && result.action !== "submit") {
177
+ result.error ??= `--lens-model is only meaningful for submit (got action "${result.action}").`;
178
+ }
140
179
  return result;
141
180
  }
@@ -9,7 +9,7 @@ import { parseNextFlags } from "./next-flags.js";
9
9
  import { buildPiGuideMessage } from "./guide-framing.js";
10
10
  import { isCtxLive, notifyCtx } from "./notify.js";
11
11
  import { phaseCompactionExtension } from "./phase-compaction.js";
12
- import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, statusText, describeConfigProblems, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
12
+ import { applyAmendment, buildPhasePrompt, buildSkillPrompt, buildValidationSummary, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, computePerPhaseTotals, computeTotals, ConfidentialityMismatchError, createEmptyStatus, DEFAULT_PIPELINE_PATH, describeScaffoldStaleness, deriveSlug, discoverLibrary, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, expandTilde, getPipelineLabel, getWorkspaceState, isWithinPath, resolveExistingPrefix, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, BroadsideCancelledError, broadsideDirFor, collectResultText, estimateSubmitText, getLens, listAmendmentNames, listBatchModels, listGuideTopics, listScaffoldRefreshFiles, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadAmendmentFile, loadBroadsideConfig, modelsText, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, statusText, describeConfigProblems, describeIncrementalFallback, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, packagedWorkspaceDir, pathExists, PACKAGE_VERSION, readBroadsideSkill, readGuide, refreshScaffold, PhasePreflightError, PIPELINE_ALIASES, publishEntry, resolvePhase, resolvePipelineChoice, resolvePublishSourceRepo, SourceRepoMismatchError, runPhasePreflight, SCAFFOLD_REFRESH_PROTECTED, seedOrchestratorFiles, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../../core/index.js";
13
13
  import { initLibrary } from "../../core/library.js";
14
14
  import { resolveUserConfigPath } from "../../core/orchestrator-config.js";
15
15
  const STATUS_WIDGET_ID = "codecarto-widget";
@@ -144,11 +144,14 @@ function describeBroadsideEstimate(estimate) {
144
144
  ? `This EXCEEDS the configured max_cost of $${estimate.maxCost.toFixed(2)}. Approving here overrides it for this run.`
145
145
  : `Within the configured max_cost of $${estimate.maxCost.toFixed(2)}.`);
146
146
  }
147
- if (estimate.baseHead) {
148
- lines.push(`Incremental: only modules changed since ${estimate.baseHead.slice(0, 8)} are included.`);
149
- }
150
- else if (estimate.sourceDirty) {
151
- lines.push("Incremental was requested but the tree is dirty — this is a full scan.");
147
+ // Only a requested incremental run has anything to say here. The dialog
148
+ // used to print "Incremental was requested but the tree is dirty" on every
149
+ // dirty tree, requested or not — a full scan that nobody asked to shrink
150
+ // read as a fallback.
151
+ if (estimate.incremental?.requested) {
152
+ lines.push(estimate.incremental.applied
153
+ ? `Incremental: only modules changed since ${(estimate.baseHead ?? "").slice(0, 8)} are included.`
154
+ : `Incremental was requested but NOT applied — ${describeIncrementalFallback(estimate.incremental.reason)}. This is a full scan.`);
152
155
  }
153
156
  lines.push("", "The estimate is a pre-flight prediction from file sizes; OpenRouter bills actual usage.");
154
157
  return lines.join("\n");
@@ -1000,7 +1003,7 @@ export default function codeCartographerExtension(pi) {
1000
1003
  },
1001
1004
  });
1002
1005
  pi.registerCommand("codecarto-broadside", {
1003
- description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models] [lenses…] [flags]",
1006
+ description: "Batch reconnaissance (Broad-Side): /codecarto-broadside [submit|collect|status|models] [lenses…] [--model=ID] [--lens-model=LENS:ID] [flags]",
1004
1007
  getArgumentCompletions: (prefix) => {
1005
1008
  const items = KNOWN_BROADSIDE_TOKENS
1006
1009
  .filter((value) => value.startsWith(prefix))
@@ -1077,10 +1080,10 @@ export default function codeCartographerExtension(pi) {
1077
1080
  if (flags.action === "models") {
1078
1081
  notifyCtx(ctx, "Fetching the OpenRouter batch-model catalog…", "info");
1079
1082
  try {
1080
- const { entries, benchmarks } = await listBatchModels(broadsideDir, config, apiKey, {
1083
+ const { entries, benchmarks, endpoints } = await listBatchModels(broadsideDir, config, apiKey, {
1081
1084
  includeBenchmarks: flags.benchmarks,
1082
1085
  });
1083
- finish(modelsText(entries, { benchmarks, defaultModel: config.model }).split("\n"), `Broad-Side: ${entries.length} batch model${entries.length === 1 ? "" : "s"} listed`);
1086
+ finish(modelsText(entries, { benchmarks, defaultModel: config.model, endpoints }).split("\n"), `Broad-Side: ${entries.length} batch model${entries.length === 1 ? "" : "s"} listed`);
1084
1087
  }
1085
1088
  catch (error) {
1086
1089
  notifyCtx(ctx, `Model catalog lookup failed: ${error instanceof Error ? error.message : String(error)}`, "error");
@@ -1114,24 +1117,49 @@ export default function codeCartographerExtension(pi) {
1114
1117
  const lenses = flags.lenses.length > 0 ? flags.lenses : config.defaultLenses;
1115
1118
  renderProgress("Slicing the repository and pricing the run…");
1116
1119
  let submit;
1120
+ // Set when a headless run was refused over max_cost, so the cancel
1121
+ // message says so instead of reading as a user's "no".
1122
+ let headlessRefusal = null;
1117
1123
  try {
1118
1124
  submit = await runBroadsideSubmit(ctx.cwd, apiKey, {
1119
1125
  lenses,
1120
- model: config.model,
1126
+ // --model= and --lens-model= select for this run; the file's
1127
+ // values are the fallback, and core pre-flights either the
1128
+ // same way (#141).
1129
+ model: flags.model ?? config.model,
1130
+ lensModels: flags.lensModels,
1121
1131
  maxCost: flags.maxCost ?? config.maxCost,
1122
1132
  // `??`, not `||`: --no-incremental parses to false and must beat a
1123
1133
  // config-set true, exactly as MCP's `incremental: false` does (#163).
1124
1134
  incremental: flags.incremental ?? config.incremental,
1125
1135
  // Pi can ask, so it asks instead of refusing over max_cost the
1126
1136
  // way MCP has to. An approval here IS the force flag.
1127
- confirm: (estimate) => ctx.ui.confirm(`Broad-Side will spend about $${estimate.totalCost.toFixed(4)}`, describeBroadsideEstimate(estimate)),
1137
+ confirm: (estimate) => {
1138
+ if (ctx.hasUI) {
1139
+ return ctx.ui.confirm(`Broad-Side will spend about $${estimate.totalCost.toFixed(4)}`, describeBroadsideEstimate(estimate));
1140
+ }
1141
+ // No dialog under `pi -p`: the stub answered "no" to every
1142
+ // estimate, so a headless submit could never fire. Behave as
1143
+ // the MCP surface does — an estimate within max_cost is
1144
+ // approved by the cap itself; one over it is refused, since
1145
+ // nobody is here to say yes — and print the breakdown either
1146
+ // way, because the dialog was the only place it showed.
1147
+ notifyCtx(ctx, describeBroadsideEstimate(estimate), "info");
1148
+ if (estimate.exceedsLimit) {
1149
+ headlessRefusal =
1150
+ `Broad-Side refused: the estimate ~$${estimate.totalCost.toFixed(4)} exceeds max_cost $${estimate.maxCost.toFixed(2)} ` +
1151
+ "and there is no dialog to approve it in a headless run. Raise --max-cost (0 for no limit) or run interactively.";
1152
+ return false;
1153
+ }
1154
+ return true;
1155
+ },
1128
1156
  });
1129
1157
  }
1130
1158
  catch (error) {
1131
1159
  if (ctx.hasUI)
1132
1160
  ctx.ui.setWidget(BROADSIDE_WIDGET_ID, undefined);
1133
1161
  if (error instanceof BroadsideCancelledError) {
1134
- notifyCtx(ctx, "Broad-Side cancelled. Nothing was submitted.", "info");
1162
+ notifyCtx(ctx, headlessRefusal ?? "Broad-Side cancelled. Nothing was submitted.", headlessRefusal ? "error" : "info");
1135
1163
  return;
1136
1164
  }
1137
1165
  notifyCtx(ctx, `Broad-Side submit failed: ${error instanceof Error ? error.message : String(error)}`, "error");
@@ -245,6 +245,8 @@ export declare function handleBroadside(args: {
245
245
  force?: boolean;
246
246
  include_benchmarks?: boolean;
247
247
  incremental?: boolean;
248
+ model?: string;
249
+ lens_models?: Record<string, string>;
248
250
  }): Promise<{
249
251
  content: {
250
252
  type: "text";
@@ -17,7 +17,7 @@ import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"
17
17
  import { CallToolRequestSchema, ErrorCode, ListToolsRequestSchema, McpError, } from "@modelcontextprotocol/sdk/types.js";
18
18
  import { mkdir, readFile, readdir, rename, writeFile } from "node:fs/promises";
19
19
  import { basename, isAbsolute, join } from "node:path";
20
- import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, seedOrchestratorFiles, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
20
+ import { buildPhasePrompt, buildSkillPrompt, buildValidationSummary, BROADSIDE_DIR, broadsideDirFor, BROADSIDE_LENS_IDS, BROADSIDE_SKILL_NAME, BroadsideConfigError, defaultBroadsideConfig, backupWorkspaceState, canonicalPath, copyPackagedWorkspace, collectResultText, completeValidatedPhase, computePerPhaseTotals, computeTotals, createEmptyStatus, DEFAULT_PIPELINE_PATH, deriveSlug, discoverLibrary, describeScaffoldStaleness, detectProvenanceConflicts, estimateSubmitText, getLens, describeDanglingCarryForward, describeMissingCompletedOutputs, describeStuckPipeline, getPipelineLabel, getWorkspaceState, isValidSlug, isWithinPathResolved, listBatchModels, listEntries, listGuideTopics, loadBroadsideConfig, modelsText, readGuide, listMissingCompletedOutputs, listSkillNames, resolvePipelineOutcome, resolveSkillName, loadCodecartoConfig, loadUsage, loadYamlFile, normalizeForComparison, PACKAGE_VERSION, packagedWorkspaceDir, pathExists, PhasePreflightError, previewPublishVersion, publishEntry, reindex as libraryReindex, refreshScaffold, resolvePhase, resolvePipelineChoice, readBroadsideSkill, runBroadsideCollect, runBroadsideStatus, runBroadsideSubmit, seedOrchestratorFiles, statusLineWriter, statusText, stringifySimpleYaml, switchPipeline, validatePhaseOutput, writeLibraryConfig, writeDashboard, } from "../core/index.js";
21
21
  import { applyAmendment } from "../core/amendment.js";
22
22
  import { appendUsageRun } from "../core/usage.js";
23
23
  import { initLibrary } from "../core/library.js";
@@ -1079,15 +1079,19 @@ export async function handleBroadside(args) {
1079
1079
  const retryTruncated = args.retry_truncated ?? config.retryTruncated;
1080
1080
  const incremental = args.incremental ?? config.incremental;
1081
1081
  if (action === "models") {
1082
- const { entries, benchmarks } = await listBatchModels(broadsideDirFor(cwd), config, apiKey, {
1082
+ const { entries, benchmarks, endpoints } = await listBatchModels(broadsideDirFor(cwd), config, apiKey, {
1083
1083
  includeBenchmarks: args.include_benchmarks === true,
1084
1084
  }).catch((error) => {
1085
1085
  throw new McpError(ErrorCode.InvalidRequest, error instanceof Error ? error.message : String(error));
1086
1086
  });
1087
- return textResult(modelsText(entries, { benchmarks, defaultModel: config.model }), {
1087
+ return textResult(modelsText(entries, { benchmarks, defaultModel: config.model, endpoints }), {
1088
1088
  models: entries,
1089
1089
  defaultModel: config.model,
1090
1090
  benchmarkMeta: benchmarks?.meta ?? null,
1091
+ // The catalog is advisory (#141): what this repository's submits
1092
+ // learned about each id's batch endpoint rides alongside it.
1093
+ catalogAdvisory: true,
1094
+ endpoints,
1091
1095
  });
1092
1096
  }
1093
1097
  if (action === "submit") {
@@ -1105,9 +1109,34 @@ export async function handleBroadside(args) {
1105
1109
  // An explicit 0 is "no limit" (#231); absent falls back to config.yaml,
1106
1110
  // whose own default is BROADSIDE_DEFAULT_MAX_COST.
1107
1111
  const maxCost = typeof args.max_cost === "number" && args.max_cost >= 0 ? args.max_cost : config.maxCost;
1112
+ // Model selection for one run (#141): `model` replaces the run default,
1113
+ // `lens_models` layers per-lens overrides over config.yaml's. Both are
1114
+ // pre-flighted by core exactly like the file's values — priced from the
1115
+ // catalog, refused without structured-output support, clamped to the
1116
+ // model's ceiling — so a wrong id fails before anything is submitted.
1117
+ const model = typeof args.model === "string" && args.model.trim() ? args.model.trim() : config.model;
1118
+ if (args.model !== undefined && !(typeof args.model === "string" && args.model.trim())) {
1119
+ throw new McpError(ErrorCode.InvalidParams, "model must be a non-empty OpenRouter batch model id (see action 'models').");
1120
+ }
1121
+ const lensModels = {};
1122
+ if (args.lens_models !== undefined) {
1123
+ if (!args.lens_models || typeof args.lens_models !== "object" || Array.isArray(args.lens_models)) {
1124
+ throw new McpError(ErrorCode.InvalidParams, "lens_models must be an object mapping lens ids to batch model ids.");
1125
+ }
1126
+ for (const [lensId, value] of Object.entries(args.lens_models)) {
1127
+ if (!BROADSIDE_LENS_IDS.includes(lensId)) {
1128
+ throw new McpError(ErrorCode.InvalidParams, `lens_models: unknown lens "${lensId}". Valid: ${BROADSIDE_LENS_IDS.join(", ")}`);
1129
+ }
1130
+ if (typeof value !== "string" || !value.trim()) {
1131
+ throw new McpError(ErrorCode.InvalidParams, `lens_models.${lensId} must be a non-empty OpenRouter batch model id.`);
1132
+ }
1133
+ lensModels[lensId] = value.trim();
1134
+ }
1135
+ }
1108
1136
  const result = await runBroadsideSubmit(cwd, apiKey, {
1109
1137
  lenses,
1110
- model: config.model,
1138
+ model,
1139
+ lensModels,
1111
1140
  maxCost,
1112
1141
  force: args.force === true,
1113
1142
  incremental,
@@ -1122,7 +1151,10 @@ export async function handleBroadside(args) {
1122
1151
  includeSynthesis,
1123
1152
  includeTriage,
1124
1153
  retryTruncated,
1125
- onStatus: (lensId, status, counts) => lines.push(` ${lensId}: ${status} (${counts.completed ?? 0}/${counts.total ?? "?"})`),
1154
+ // One line per *change* of a lens's status. Every poll used to
1155
+ // append a line, so a four-minute wait returned twenty-six
1156
+ // "in_progress (0/1)" lines before the result (0.22.0 live run).
1157
+ onStatus: statusLineWriter(lines),
1126
1158
  }).catch((error) => {
1127
1159
  // The `collect` action normalizes this same call; without it here,
1128
1160
  // a failure during submit-with-wait reached the client as an
@@ -1516,6 +1548,15 @@ const TOOLS = [
1516
1548
  type: "boolean",
1517
1549
  description: "For action 'models': annotate each model with its Artificial Analysis coding index (extra API call; default false).",
1518
1550
  },
1551
+ model: {
1552
+ type: "string",
1553
+ description: "For submit: the OpenRouter batch model for this run (an id ending in :batch, as listed by action 'models'). Falls back to model in .codecarto/broadside/config.yaml, then the shipped default. Pre-flighted like the configured model: priced from the live catalog, refused without structured-output support, clamped to its completion ceiling. The models listing is advisory — some catalog ids have no batch endpoint and are refused at submit, at no cost; the listing tags ids this repository has already seen accepted or refused.",
1554
+ },
1555
+ lens_models: {
1556
+ type: "object",
1557
+ additionalProperties: { type: "string" },
1558
+ description: "For submit: per-lens model overrides for this run, e.g. {\"security\": \"deepseek/deepseek-v4-pro-0813:batch\"}. Keys are lens ids; a lens named here runs on that model, others on `model`. Layered over lens_models in .codecarto/broadside/config.yaml (a lens set in both takes the parameter's). Each override is priced, capability-checked, and clamped individually, and the estimate breaks cost out per lens.",
1559
+ },
1519
1560
  },
1520
1561
  required: ["cwd", "action"],
1521
1562
  },
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "codecartographer-pi",
3
- "version": "0.21.0",
3
+ "version": "0.22.1",
4
4
  "mcpName": "io.github.HuginnIndustries/codecartographer",
5
5
  "description": "Turn an unfamiliar codebase into a validated reimplementation spec, then synthesize confirmed specs and a product vision into a traceable plan.",
6
6
  "type": "module",