codecartographer-pi 0.19.0 → 0.19.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -90,6 +90,17 @@ Collect runs two cross-lens post-passes by default: **synthesis** (the
90
90
  executive report) and **triage** (the prioritized work order). Pass
91
91
  `include_synthesis: false` or `include_triage: false` on collect to skip one.
92
92
 
93
+ Two caveats apply to any model you pick. The `models` action lists every id
94
+ OpenRouter advertises a `:batch` variant for, and many of those variants do not
95
+ exist — submitting one returns `does not have a :batch endpoint`, with nothing in
96
+ the catalog to distinguish it beforehand. A rejected batch costs nothing, so
97
+ probe a candidate on a single lens first. And reasoning competes with the answer for
98
+ `max_tokens`: Broad-Side caps thinking at a quarter of each lens's output budget
99
+ so three quarters remain for the JSON, which is the split the cost estimate
100
+ already assumes. It caps rather than disables because some endpoints refuse to
101
+ be switched off entirely. Override with `reasoning:` in `config.yaml` only
102
+ alongside a raised output budget.
103
+
93
104
  Lenses do not all have to run on the same model. `lens_models` in `config.yaml`
94
105
  routes individual lenses to their own batch model — the usual reason being that
95
106
  a stronger model changes security and defect findings more than it changes an
@@ -125,6 +136,13 @@ from OpenRouter's runtime cost tracking: it predicts from file sizes before
125
136
  spend, it does not stop a batch mid-flight. Actual spend appears in
126
137
  `run-meta.json` after collect.
127
138
 
139
+ Budget for the **ceiling, not the estimate**. OpenRouter charges the worst
140
+ case when it accepts a batch — every request's full `max_tokens` — and
141
+ refunds the unused output as each request settles. A six-lens scan of this
142
+ repository held $0.6972 and settled at $0.3806, so the balance a run needs to
143
+ *start* is roughly double what it ends up costing. The submit estimate sits
144
+ between the two: above what the run should settle at, below what it can hold.
145
+
128
146
  ## Resilience notes
129
147
 
130
148
  - **Truncation is spoken.** A lens output whose JSON does not parse — even
@@ -46,9 +46,48 @@
46
46
  # and picking one for you would spend your money on our guess. Compare
47
47
  # candidates with the `models` action first.
48
48
  #
49
+ # Two things to check before committing to a model, both learned the hard way:
50
+ #
51
+ # 1. The `models` action lists whatever OpenRouter advertises a `:batch`
52
+ # variant for, and a good number of those variants do not actually exist —
53
+ # submitting one comes back `Model '<id>' does not have a :batch endpoint.`
54
+ # Nothing in the catalog distinguishes them. Every Anthropic and OpenAI
55
+ # batch id tried so far is rejected this way; Google and DeepSeek work.
56
+ # A rejected batch costs nothing, so probe a candidate on one lens before
57
+ # relying on it.
58
+ # 2. A reasoning-capable model spends its output budget thinking, and the
59
+ # thinking is billed at the full output rate. See `reasoning:` below.
60
+ #
49
61
  # lens_models:
50
- # security: anthropic/claude-opus-4.5:batch
51
- # defect: anthropic/claude-opus-4.5:batch
62
+ # security: deepseek/deepseek-v4-pro-0813:batch
63
+ # defect: deepseek/deepseek-v4-pro-0813:batch
64
+
65
+ # Reasoning control, sent on every lens request.
66
+ #
67
+ # By default Broad-Side caps thinking at a quarter of the lens's output budget,
68
+ # leaving the other three quarters for the answer — which is exactly the split
69
+ # the cost estimate already assumes.
70
+ #
71
+ # The cap exists because reasoning competes with the answer for `max_tokens`.
72
+ # One measured run spent 5,758 of a 6,000-token budget thinking and left ~230
73
+ # tokens for the JSON, which truncated mid-structure on 11 of 13 slices; those
74
+ # tokens bill at the full output rate, so it paid for ~6,000 output tokens per
75
+ # slice to receive ~230 usable ones. The shipped default model does the same
76
+ # thing less consistently — reasoning from 0 to 5,757 tokens across 13 slices,
77
+ # three of them cut off — so this is not something only exotic models do.
78
+ #
79
+ # It is a cap rather than an off switch on purpose. Some endpoints refuse to be
80
+ # switched off: `google/gemini-3.8-flash:batch` rejects the entire batch with
81
+ # "Reasoning is mandatory for this endpoint and cannot be disabled", which turns
82
+ # a partial result into none at all. Capping works either way.
83
+ #
84
+ # Override only with a raised lens output budget, or the JSON truncates exactly
85
+ # as above.
86
+ #
87
+ # reasoning:
88
+ # effort: low # minimal | low | medium | high
89
+ # max_tokens: 2000 # or set the thinking budget directly
90
+ # enabled: false # only where the provider allows it
52
91
 
53
92
  # Approximate run expense limit in USD (0 = no limit). Before submitting,
54
93
  # Broad-Side estimates the run cost from the collected file sizes and the
@@ -68,8 +107,8 @@
68
107
  # lookup fails (offline, private model) or you want to assert a ceiling.
69
108
  #
70
109
  # pricing:
71
- # input_per_m: 0.1875
72
- # output_per_m: 0.9375
110
+ # input_per_m: 0.375
111
+ # output_per_m: 1.875
73
112
  # ---------------------------------------------------------------------------
74
113
  # Run defaults. Each key below mirrors a codecarto_broadside parameter of the
75
114
  # same name and sets this repository's default for it; an explicit parameter on
@@ -3,4 +3,4 @@
3
3
  # workspace's framework-owned files (GUIDE.md, templates/, workflow/ pipelines
4
4
  # and VALIDATE.md) predate the running release. Written at release time and
5
5
  # copied verbatim by init — never edit by hand.
6
- scaffold_version: 0.19.0
6
+ scaffold_version: 0.19.2
@@ -6,8 +6,8 @@ export declare const BROADSIDE_SKILL_NAME = "broadside";
6
6
  export declare const BROADSIDE_STATE_FILE = "state.json";
7
7
  export declare const BROADSIDE_CONFIG_FILE = "config.yaml";
8
8
  export declare const BROADSIDE_STATE_SCHEMA_VERSION = 1;
9
- export declare const BROADSIDE_INPUT_PRICE_PER_M = 0.1875;
10
- export declare const BROADSIDE_OUTPUT_PRICE_PER_M = 0.9375;
9
+ export declare const BROADSIDE_INPUT_PRICE_PER_M = 0.375;
10
+ export declare const BROADSIDE_OUTPUT_PRICE_PER_M = 1.875;
11
11
  export declare const BROADSIDE_MODELS_URL = "https://openrouter.ai/api/v1/models";
12
12
  export declare const BROADSIDE_BENCHMARKS_URL = "https://openrouter.ai/api/v1/benchmarks";
13
13
  export declare const BROADSIDE_CATALOG_CACHE_FILE = "model-catalog.json";
@@ -49,7 +49,13 @@ export type CodingBenchmarks = {
49
49
  export type BroadsideCatalogResult = {
50
50
  model: string;
51
51
  source: "built-in" | "config" | "live" | "cache";
52
- entry: CatalogEntry | null;
52
+ /**
53
+ * Always resolved. `resolveCatalogEntry` either returns an entry — from
54
+ * config, cache, the live catalog, or the compile-time fallback — or throws
55
+ * naming the model it could not price. This was declared nullable, which is
56
+ * the only reason the single consumer needed a non-null assertion to read it.
57
+ */
58
+ entry: CatalogEntry;
53
59
  benchmarks?: CodingBenchmarks;
54
60
  };
55
61
  export type JsonSchemaDef = {
@@ -80,6 +86,46 @@ export type FileSlice = {
80
86
  /** Repo-relative paths of the files folded into this slice. */
81
87
  files: string[];
82
88
  };
89
+ /**
90
+ * OpenRouter's unified `reasoning` control, as sent on a lens request.
91
+ *
92
+ * Left unsent, each model applies its own default — which is how a
93
+ * reasoning-capable model came to spend 5,758 of a 6,000-token output budget
94
+ * thinking, leaving ~230 tokens for JSON that then truncated mid-structure. The
95
+ * thinking is billed at the full *output* rate, so the run paid for roughly
96
+ * 6,000 output tokens per slice to receive 230 usable ones.
97
+ */
98
+ export type BroadsideReasoning = {
99
+ enabled?: boolean;
100
+ effort?: "minimal" | "low" | "medium" | "high";
101
+ max_tokens?: number;
102
+ };
103
+ /**
104
+ * The share of a lens's output budget reasoning may spend.
105
+ *
106
+ * `estimateCost` already budgets output at 75% of `maxTokens`; capping thinking
107
+ * at the remaining quarter makes that assumption true by construction and
108
+ * guarantees the answer has room. A floor keeps the cap sane for a small lens.
109
+ */
110
+ export declare const BROADSIDE_REASONING_BUDGET_FRACTION = 0.25;
111
+ export declare const BROADSIDE_MIN_REASONING_TOKENS = 512;
112
+ /**
113
+ * Cap reasoning for a lens request — deliberately a cap, not an off switch.
114
+ *
115
+ * Disabling outright is not portable: `google/gemini-3.8-flash:batch` refuses
116
+ * the whole batch with *"Reasoning is mandatory for this endpoint and cannot be
117
+ * disabled"*, turning a partial result into none at all. Capping works whether
118
+ * or not a provider allows reasoning to be switched off.
119
+ *
120
+ * The failure this prevents is the budget being spent thinking rather than
121
+ * answering. Measured on one run: 5,758 of a 6,000-token budget went to
122
+ * reasoning, leaving ~230 tokens for JSON that truncated mid-structure — and
123
+ * those tokens bill at the full output rate. The shipped default model does the
124
+ * same thing less consistently (reasoning tokens from 0 to 5,757 across 13
125
+ * slices, three of them cut off at `finish_reason: length`), so this is not a
126
+ * multi-model concern.
127
+ */
128
+ export declare function defaultReasoningFor(maxTokens: number): BroadsideReasoning;
83
129
  export type BatchRequest = {
84
130
  custom_id: string;
85
131
  body: {
@@ -93,6 +139,7 @@ export type BatchRequest = {
93
139
  json_schema: JsonSchemaDef;
94
140
  };
95
141
  max_tokens: number;
142
+ reasoning?: BroadsideReasoning;
96
143
  };
97
144
  };
98
145
  export type BatchTerminalStatus = "completed" | "failed" | "expired" | "cancelled";
@@ -115,6 +162,8 @@ export type BroadsideSynthesisEntry = {
115
162
  batchId?: string;
116
163
  status: "pending" | "submitted" | "completed" | "failed";
117
164
  cost?: number;
165
+ /** Why the pass was retired, when the batch reported one. */
166
+ error?: string;
118
167
  };
119
168
  /** One triage item — a scouting lead turned into a work-order entry. */
120
169
  export type TriageItem = {
@@ -177,6 +226,8 @@ export type BroadsideConfig = {
177
226
  * trade-off is not the same for every lens.
178
227
  */
179
228
  lensModels: Partial<Record<BroadsideLensId, string>>;
229
+ /** Overrides every lens's reasoning setting when present. */
230
+ reasoning: BroadsideReasoning | null;
180
231
  /**
181
232
  * Repo defaults for the per-call run knobs. Each mirrors a tool parameter
182
233
  * of the same name; an explicit parameter always wins. They live here so a
@@ -220,6 +271,8 @@ export type BroadsideEstimate = {
220
271
  /** Set when incremental scouting found a baseline to diff against. */
221
272
  baseHead: string | null;
222
273
  sourceDirty: boolean;
274
+ /** Whether a requested incremental run actually narrowed this estimate. */
275
+ incremental: BroadsideIncrementalOutcome;
223
276
  /** The provider's completion ceiling, when the catalog advertises one. */
224
277
  outputCap?: number;
225
278
  };
@@ -227,6 +280,22 @@ export type BroadsideEstimate = {
227
280
  export declare class BroadsideCancelledError extends Error {
228
281
  constructor(message?: string);
229
282
  }
283
+ /**
284
+ * Whether incremental scouting actually narrowed the run.
285
+ *
286
+ * A request for incremental falls back to a full scan whenever there is nothing
287
+ * to diff against, and that fallback costs real money — the caller asked for the
288
+ * cheap mode and gets the expensive one. It must be reported, not inferred from
289
+ * the request counts.
290
+ */
291
+ export type BroadsideIncrementalOutcome = {
292
+ requested: boolean;
293
+ applied: boolean;
294
+ /** The commit the run diffed against, when one was found. */
295
+ baseHead: string | null;
296
+ /** Why a requested incremental run did not apply. */
297
+ reason?: "dirty-worktree" | "no-baseline" | "diff-failed";
298
+ };
230
299
  export type BroadsideSubmitResult = {
231
300
  runId: string;
232
301
  outputDir: string;
@@ -242,6 +311,7 @@ export type BroadsideSubmitResult = {
242
311
  supportsStructuredOutputs?: boolean;
243
312
  expirationDate?: string | null;
244
313
  };
314
+ incremental: BroadsideIncrementalOutcome;
245
315
  };
246
316
  export type BroadsideCollectResult = {
247
317
  runId: string;
@@ -277,6 +347,7 @@ type LensDefinition = {
277
347
  sliceBy: "none" | "directory" | "auto";
278
348
  maxChars: number;
279
349
  maxTokens: number;
350
+ reasoning?: BroadsideReasoning;
280
351
  skipTestFiles?: boolean;
281
352
  globsFor: (info: RepoInfo) => string[];
282
353
  systemPrompt: (info: RepoInfo) => string;
@@ -286,8 +357,23 @@ export declare function getLens(lensId: BroadsideLensId): LensDefinition;
286
357
  export declare function listLenses(): LensDefinition[];
287
358
  export declare function collectRepoInfo(targetDir: string): Promise<RepoInfo>;
288
359
  export declare function gatherSlices(targetDir: string, lens: LensDefinition, info: RepoInfo): Promise<FileSlice[]>;
289
- export declare function buildBatchRequest(lens: LensDefinition, info: RepoInfo, slice: FileSlice, index: number, sliceCount: number, model?: string, maxTokensOverride?: number): BatchRequest;
290
- export declare function estimateCost(lens: LensDefinition, slices: FileSlice[], pricing: ModelPricing, maxTokensOverride?: number): {
360
+ export declare function buildBatchRequest(lens: LensDefinition, info: RepoInfo, slice: FileSlice, index: number, sliceCount: number, model?: string, maxTokensOverride?: number, reasoningOverride?: BroadsideReasoning): BatchRequest;
361
+ /**
362
+ * Pre-flight cost estimate for one lens.
363
+ *
364
+ * Every slice is its own batch request, so both halves scale with the slice
365
+ * count. The output half used to be a single `maxTokens * 0.75` for the whole
366
+ * lens no matter how many requests it sent — on a repository that sliced into
367
+ * 13 modules that budgeted one request's output and shipped thirteen, and a
368
+ * live run came in at roughly 3x its estimate. Since this number is what
369
+ * `max_cost` binds against, under-counting it lets a run outspend the cap the
370
+ * user set.
371
+ *
372
+ * @param info - Repo info, when the caller has it: lets the estimate include
373
+ * the system prompt and JSON schema each request carries. Omitted, the
374
+ * estimate covers slice content only, which is what the old signature did.
375
+ */
376
+ export declare function estimateCost(lens: LensDefinition, slices: FileSlice[], pricing: ModelPricing, maxTokensOverride?: number, info?: RepoInfo): {
291
377
  inputTokens: number;
292
378
  outputTokens: number;
293
379
  cost: number;
@@ -313,7 +399,40 @@ export declare function readBroadsideSkill(cwd: string): Promise<{
313
399
  }>;
314
400
  export declare function defaultBroadsideState(): BroadsideStateFile;
315
401
  export declare function loadBroadsideState(broadsideDir: string): Promise<BroadsideStateFile>;
402
+ /**
403
+ * Overwrite `state.json` wholesale with `state`.
404
+ *
405
+ * Prefer {@link persistBroadsideRun} anywhere a live operation is recording its
406
+ * own progress — this entry point replaces the file, so any run a concurrent
407
+ * process recorded in the meantime is erased. It remains the right call for
408
+ * seeding a fresh workspace and for test fixtures, where "make the file exactly
409
+ * this" is the intent.
410
+ */
316
411
  export declare function saveBroadsideState(broadsideDir: string, state: BroadsideStateFile): Promise<void>;
412
+ /**
413
+ * Read-modify-write `state.json` under a lock.
414
+ *
415
+ * The lock is held only for the read-modify-write, never for the surrounding
416
+ * operation: a `collect` can poll for the better part of an hour, and holding
417
+ * the lock across that would push every concurrent caller past the 5s lock
418
+ * timeout.
419
+ */
420
+ export declare function updateBroadsideStateAtomically(broadsideDir: string, mutate: (state: BroadsideStateFile) => void | Promise<void>): Promise<BroadsideStateFile>;
421
+ /**
422
+ * Record one run's current shape, merged into whatever is on disk *now*.
423
+ *
424
+ * Broad-Side operations are long-lived and hold their state in memory while
425
+ * they poll. Writing that snapshot back wholesale silently erased any run a
426
+ * concurrent operation had recorded since it was loaded, orphaning that run's
427
+ * paid results on disk — present as files, invisible to `list`, and unreachable
428
+ * by `collect`, which finds its run by position in `state.runs`. Observed live:
429
+ * a submit at 23:35 was erased by a collect that had loaded state before it and
430
+ * wrote back at 00:08.
431
+ *
432
+ * Merging by run id also self-heals: a run erased by an older writer is
433
+ * restored the next time its own operation checkpoints.
434
+ */
435
+ export declare function persistBroadsideRun(broadsideDir: string, run: BroadsideRun): Promise<BroadsideStateFile>;
317
436
  export declare function loadBroadsideConfig(broadsideDir: string): Promise<BroadsideConfig>;
318
437
  export declare function builtInCatalogEntry(model: string): CatalogEntry | null;
319
438
  export declare function builtInPricing(model: string): ModelPricing | null;
@@ -336,6 +455,14 @@ export declare function submitBatch(batchRequests: BatchRequest[], apiKey: strin
336
455
  error?: unknown;
337
456
  }>;
338
457
  export declare function fetchBatch(batchId: string, apiKey: string, fetcher?: FetchLike): Promise<Record<string, unknown>>;
458
+ /**
459
+ * Batch statuses that will never produce a result.
460
+ *
461
+ * Deliberately excludes the synthetic `timeout` this module returns when a poll
462
+ * budget expires: that batch is still running server-side and has already been
463
+ * charged, so callers must come back for it rather than retire it.
464
+ */
465
+ export declare const BROADSIDE_DEAD_BATCH_STATUSES: string[];
339
466
  export declare function pollBatchUntilTerminal(batchId: string, apiKey: string, opts?: {
340
467
  deadlineMs?: number;
341
468
  onStatus?: (status: string, counts: Record<string, unknown>) => void;
@@ -398,6 +525,18 @@ export type StoredLensResult = {
398
525
  */
399
526
  export declare function parseLensJson(content: string): unknown | null;
400
527
  export declare function saveLensResults(runDir: string, lensId: BroadsideLensId, batch: Record<string, unknown>): Promise<StoredLensResult[]>;
528
+ /**
529
+ * Rebuild lens results from what a previous collect already wrote to disk.
530
+ *
531
+ * The post-passes are gated on having lens findings in hand, and a collect
532
+ * only holds the ones *it* polled. When an earlier collect saved every lens
533
+ * and then died before synthesis and triage ran — the batch window is long
534
+ * and a poll can easily be interrupted — the next collect finds every lens
535
+ * already terminal, skips them all, and would otherwise reach the post-pass
536
+ * gate with nothing to hand it. Reading the saved results back is what makes
537
+ * "a resumed collect can finish whichever is still pending" true.
538
+ */
539
+ export declare function loadSavedLensResults(runDir: string, lenses: BroadsideLensId[]): Promise<StoredLensResult[]>;
401
540
  export declare function runBroadsideCollect(cwd: string, apiKey: string, opts?: {
402
541
  waitMs?: number;
403
542
  includeSynthesis?: boolean;