codecartographer-pi 0.18.0 → 0.19.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -215,6 +215,13 @@ post_pipeline:
215
215
 
216
216
  `kind` is one of: `needs-runtime-test`, `needs-maintainer-decision`, `needs-spec-ruling`, `defer-to-phase`, `needs-fixture-capture`, or a post-pipeline work kind such as `spike` or `amendment`. Every new `carry_forward.target_phase` must be an ID in the active pipeline. Every `post_pipeline` entry requires a stable ID. Open questions should carry a stable `id` (e.g. `q-loadconfig-ambiguity`); if omitted, the framework auto-assigns one. When a later phase resolves an open question, list its id in `open_question_closures` to remove it from all phases. The downstream phase records resolved carry-forward IDs in `carry_forward_closures`; completion removes those entries atomically.
217
217
 
218
+ **Closing a routed item does not settle the question it came from.** A phase routinely registers an open question and routes one of its candidate answers onward in the same handoff; addressing the routed item later is not the same as answering the question. Two optional fields make that distinction enforceable:
219
+
220
+ - A `carry_forward` entry may name `derives_from: <open-question-id>`, meaning "this routed item is one candidate answer to that question." Fill it in the handoff that registers the question, when both entries are in front of you. Completion then refuses a `carry_forward_closures` entry for that item while its question is still open and is not closed by the same handoff, naming both ids. Close the question with evidence in that handoff, or leave the item routed and give the finding an unsettled action (`verify at runtime`). Entirely opt-in: a `carry_forward` entry without `derives_from` closes exactly as it always has.
221
+ - An `open_question_closures` entry may be written as `{ id, evidence }` instead of a bare id, where `evidence` names what settled the question. For a question whose `kind` is `needs-runtime-test`, non-empty `evidence` is **required** — such a question closes on a spike report or an observation against the running system, not on another read of the same source — and that requirement holds however the closure is written, so a bare id no longer closes a runtime question. Questions of every other kind are unaffected, and a bare id stays valid for them. The requirement follows the scaffold: a workspace at scaffold version 0.19.0 or newer has the closure refused, while an older scaffold — which never documented the rule — gets a non-gating note, and refreshing the scaffold opts it in.
222
+
223
+ Completed phases' declared coverage gaps travel too: the `Skipped scope` and `Known blind spots` bullets of every completed phase's `## Coverage and limits` section are surfaced in the next phase's orchestrator duties. A finding that lands inside one of those gaps must either close it with cited new evidence of its own or inherit its uncertainty.
224
+
218
225
  After the pipeline completes, the handoff channel closes with it. Post-pipeline resolutions — an open question answered on evidence, a finished `post_pipeline` backlog item — are applied with an **amendment**: write `scratch/amendments/<slug>.yaml` (see `templates/amendment.yaml`) and run `codecarto_amend` (MCP) or `/codecarto-amend <slug>` (Pi). It updates `workflow/status.yaml` under the same lock completion uses and writes an amendment closeout plus THREAD_LOG entry. Never hand-edit `status.yaml` for this; amendments are refused while the pipeline is still running, so the two channels cannot race.
219
226
 
220
227
  ## Phase Selection Logic
@@ -125,6 +125,13 @@ from OpenRouter's runtime cost tracking: it predicts from file sizes before
125
125
  spend, it does not stop a batch mid-flight. Actual spend appears in
126
126
  `run-meta.json` after collect.
127
127
 
128
+ Budget for the **ceiling, not the estimate**. OpenRouter charges the worst
129
+ case when it accepts a batch — every request's full `max_tokens` — and
130
+ refunds the unused output as each request settles. A six-lens scan of this
131
+ repository held $0.6972 and settled at $0.3806, so the balance a run needs to
132
+ *start* is roughly double what it ends up costing. The submit estimate sits
133
+ between the two: above what the run should settle at, below what it can hold.
134
+
128
135
  ## Resilience notes
129
136
 
130
137
  - **Truncation is spoken.** A lens output whose JSON does not parse — even
@@ -68,8 +68,8 @@
68
68
  # lookup fails (offline, private model) or you want to assert a ceiling.
69
69
  #
70
70
  # pricing:
71
- # input_per_m: 0.1875
72
- # output_per_m: 0.9375
71
+ # input_per_m: 0.375
72
+ # output_per_m: 1.875
73
73
  # ---------------------------------------------------------------------------
74
74
  # Run defaults. Each key below mirrors a codecarto_broadside parameter of the
75
75
  # same name and sets this repository's default for it; an explicit parameter on
@@ -4,9 +4,28 @@ schema_version: 1
4
4
  phase_id: <phase-id>
5
5
  owner_notes: []
6
6
  open_questions: []
7
+ # Items a specific later phase closes. Every entry needs a target_phase naming a
8
+ # downstream phase in the active pipeline. Optional derives_from names the
9
+ # open_questions id this item is one candidate answer to:
10
+ # - id: mech-CF3
11
+ # kind: defer-to-phase
12
+ # target_phase: defect-scan-semantic
13
+ # derives_from: q-logit-bias-root-cause
14
+ # description: <the routed item>
15
+ # Fill derives_from in the handoff that registers the question — that is the one
16
+ # moment both entries are in front of you. Closing the routed item later does not
17
+ # settle the question it came from, and completion refuses such a closure while
18
+ # the question is still open and unclosed by the same handoff.
7
19
  carry_forward: []
8
20
  carry_forward_closures: []
9
- # Resolved open questions: their IDs are removed from all phases.
21
+ # Resolved open questions: their IDs are removed from all phases. Either a bare
22
+ # id, or an object naming the evidence that settled it:
23
+ # - q-loadconfig-ambiguity
24
+ # - id: q-logit-bias-root-cause
25
+ # evidence: <the spike report or runtime observation that answered it>
26
+ # A question of kind needs-runtime-test requires non-empty evidence — it closes
27
+ # on runtime evidence, not on another read of the same source. Refused from
28
+ # scaffold version 0.19.0; older scaffolds get a non-gating note instead.
10
29
  open_question_closures: []
11
30
  # Work after the active pipeline: spikes, amendments, deltas, maintainer rulings,
12
31
  # or opinionated reruns. Every entry requires a stable id.
@@ -3,4 +3,4 @@
3
3
  # workspace's framework-owned files (GUIDE.md, templates/, workflow/ pipelines
4
4
  # and VALIDATE.md) predate the running release. Written at release time and
5
5
  # copied verbatim by init — never edit by hand.
6
- scaffold_version: 0.18.0
6
+ scaffold_version: 0.19.1
@@ -17,7 +17,7 @@ owner_notes: [] # 2-3 durable observations; appended to the phas
17
17
  open_questions: [] # genuinely unknown, no later phase will close them
18
18
  carry_forward: [] # deferred to a specific later phase in this pipeline
19
19
  carry_forward_closures: [] # ids of carry_forward entries this phase resolved
20
- open_question_closures: [] # ids of open questions this phase resolved, removed everywhere
20
+ open_question_closures: [] # open questions this phase resolved, removed everywhere; bare id or {id, evidence}
21
21
  post_pipeline: [] # work after the pipeline; every entry needs a stable id
22
22
  decisions: [] # choices made beyond what the prompt specified; completion appends them to DECISIONS.md
23
23
  proposed_conventions: [] # patterns proposed for promotion; completion stages them in CONVENTIONS.md
@@ -41,7 +41,7 @@ Omitted arrays default to empty. A malformed collection fails completion without
41
41
  deferred_reason: Distinguishing them needs a runtime probe this phase cannot run.
42
42
  ```
43
43
 
44
- `carry_forward` entries add `target_phase`:
44
+ `carry_forward` entries add `target_phase`, and optionally `derives_from`:
45
45
 
46
46
  ```yaml
47
47
  - id: arch-CF2
@@ -49,10 +49,18 @@ Omitted arrays default to empty. A malformed collection fails completion without
49
49
  target_phase: protocols
50
50
  description: MCP endpoints listed by name only; schemas not extracted.
51
51
  deferred_reason: Wire-format extraction is the protocols phase's rubric.
52
+
53
+ - id: mech-CF3
54
+ kind: defer-to-phase
55
+ target_phase: defect-scan-semantic
56
+ derives_from: q-logit-bias-root-cause # optional: the open question this is one candidate answer to
57
+ description: The client sends logit_bias as a map; the documented shape is an array.
52
58
  ```
53
59
 
54
60
  Allowed `kind` values: `needs-runtime-test`, `needs-maintainer-decision`, `needs-spec-ruling`, `defer-to-phase`, `needs-fixture-capture`.
55
61
 
62
+ `derives_from` names an `open_questions` id. Fill it in the handoff that registers the question — a phase that routes a candidate answer onward usually writes both entries at once, which is the one moment both are in front of you. It is optional and additive: a handoff that omits it behaves exactly as before.
63
+
56
64
  `proposed_conventions` entries (optional; omitted defaults to empty):
57
65
 
58
66
  ```yaml
@@ -81,6 +89,26 @@ Completion then removes the entry atomically. Resolving an open question works t
81
89
 
82
90
  Re-deferring instead of closing means writing a fresh `carry_forward` entry naming a later `target_phase`.
83
91
 
92
+ ### Closing a routed item does not settle the question it came from
93
+
94
+ Addressing what was routed to you is not the same as answering the question that produced it. Two rules make that enforceable:
95
+
96
+ - **A closure whose `derives_from` question is still open is refused.** If the item you are closing declares `derives_from: <question-id>`, that question is still in `status.yaml`, and this same handoff does not close it, completion refuses and names both ids. Either close the question here with the evidence that settles it, or leave the item routed and give the finding an unsettled action (`verify at runtime`) so it inherits the question's uncertainty. This rule is entirely opt-in — an entry without `derives_from` closes as it always has.
97
+ - **A `needs-runtime-test` question closes on runtime evidence.** Write the closure as an object and say where that evidence lives:
98
+
99
+ ```yaml
100
+ open_question_closures:
101
+ - q-loadconfig-ambiguity # bare id: still valid for any other kind
102
+ - id: q-logit-bias-root-cause
103
+ evidence: scratch/spikes/logit-bias.md — probe against llama-server b4321
104
+ ```
105
+
106
+ Non-empty `evidence` is required when the question's `kind` is `needs-runtime-test`, whether the closure is written as a bare id or as an object. It is checked for presence, not judged — a spike report or an observation against the running system is what belongs there, and another read of the same source is not. Questions of every other kind close on a bare id exactly as before.
107
+
108
+ The requirement is scoped to the scaffold that documents it. A workspace whose `workflow/scaffold-version.yaml` is 0.19.0 or newer has completion refuse such a closure; an older or unversioned scaffold — whose own templates never stated the rule — gets a non-gating `NOTE:` instead, so an in-flight run written against the older contract cannot be stopped by a rule it was never told. Refreshing the scaffold (`codecarto_refresh_scaffold` on MCP, `/codecarto-refresh-scaffold` on Pi) opts a workspace in.
109
+
110
+ Upstream coverage gaps travel the same way, without gating: the `Skipped scope` and `Known blind spots` bullets of every completed phase's `## Coverage and limits` section appear in your phase prompt's orchestrator duties. A finding of yours inside one of those gaps must either close it with cited new evidence or inherit its uncertainty.
111
+
84
112
  ## The failure this prevents
85
113
 
86
114
  Before completion required a handoff, a phase could finish with empty state and no signal. A real seven-phase run documented five cross-phase routings in its report prose, wrote no handoffs, and completed all seven phases with `carry_forward: []` throughout. Every downstream phase's routed-item intake was empty. The findings survived only because each phase happened to re-read the previous phase's full markdown.
@@ -6,8 +6,8 @@ export declare const BROADSIDE_SKILL_NAME = "broadside";
6
6
  export declare const BROADSIDE_STATE_FILE = "state.json";
7
7
  export declare const BROADSIDE_CONFIG_FILE = "config.yaml";
8
8
  export declare const BROADSIDE_STATE_SCHEMA_VERSION = 1;
9
- export declare const BROADSIDE_INPUT_PRICE_PER_M = 0.1875;
10
- export declare const BROADSIDE_OUTPUT_PRICE_PER_M = 0.9375;
9
+ export declare const BROADSIDE_INPUT_PRICE_PER_M = 0.375;
10
+ export declare const BROADSIDE_OUTPUT_PRICE_PER_M = 1.875;
11
11
  export declare const BROADSIDE_MODELS_URL = "https://openrouter.ai/api/v1/models";
12
12
  export declare const BROADSIDE_BENCHMARKS_URL = "https://openrouter.ai/api/v1/benchmarks";
13
13
  export declare const BROADSIDE_CATALOG_CACHE_FILE = "model-catalog.json";
@@ -49,7 +49,13 @@ export type CodingBenchmarks = {
49
49
  export type BroadsideCatalogResult = {
50
50
  model: string;
51
51
  source: "built-in" | "config" | "live" | "cache";
52
- entry: CatalogEntry | null;
52
+ /**
53
+ * Always resolved. `resolveCatalogEntry` either returns an entry — from
54
+ * config, cache, the live catalog, or the compile-time fallback — or throws
55
+ * naming the model it could not price. This was declared nullable, which is
56
+ * the only reason the single consumer needed a non-null assertion to read it.
57
+ */
58
+ entry: CatalogEntry;
53
59
  benchmarks?: CodingBenchmarks;
54
60
  };
55
61
  export type JsonSchemaDef = {
@@ -115,6 +121,8 @@ export type BroadsideSynthesisEntry = {
115
121
  batchId?: string;
116
122
  status: "pending" | "submitted" | "completed" | "failed";
117
123
  cost?: number;
124
+ /** Why the pass was retired, when the batch reported one. */
125
+ error?: string;
118
126
  };
119
127
  /** One triage item — a scouting lead turned into a work-order entry. */
120
128
  export type TriageItem = {
@@ -220,6 +228,8 @@ export type BroadsideEstimate = {
220
228
  /** Set when incremental scouting found a baseline to diff against. */
221
229
  baseHead: string | null;
222
230
  sourceDirty: boolean;
231
+ /** Whether a requested incremental run actually narrowed this estimate. */
232
+ incremental: BroadsideIncrementalOutcome;
223
233
  /** The provider's completion ceiling, when the catalog advertises one. */
224
234
  outputCap?: number;
225
235
  };
@@ -227,6 +237,22 @@ export type BroadsideEstimate = {
227
237
  export declare class BroadsideCancelledError extends Error {
228
238
  constructor(message?: string);
229
239
  }
240
+ /**
241
+ * Whether incremental scouting actually narrowed the run.
242
+ *
243
+ * A request for incremental falls back to a full scan whenever there is nothing
244
+ * to diff against, and that fallback costs real money — the caller asked for the
245
+ * cheap mode and gets the expensive one. It must be reported, not inferred from
246
+ * the request counts.
247
+ */
248
+ export type BroadsideIncrementalOutcome = {
249
+ requested: boolean;
250
+ applied: boolean;
251
+ /** The commit the run diffed against, when one was found. */
252
+ baseHead: string | null;
253
+ /** Why a requested incremental run did not apply. */
254
+ reason?: "dirty-worktree" | "no-baseline" | "diff-failed";
255
+ };
230
256
  export type BroadsideSubmitResult = {
231
257
  runId: string;
232
258
  outputDir: string;
@@ -242,6 +268,7 @@ export type BroadsideSubmitResult = {
242
268
  supportsStructuredOutputs?: boolean;
243
269
  expirationDate?: string | null;
244
270
  };
271
+ incremental: BroadsideIncrementalOutcome;
245
272
  };
246
273
  export type BroadsideCollectResult = {
247
274
  runId: string;
@@ -287,7 +314,22 @@ export declare function listLenses(): LensDefinition[];
287
314
  export declare function collectRepoInfo(targetDir: string): Promise<RepoInfo>;
288
315
  export declare function gatherSlices(targetDir: string, lens: LensDefinition, info: RepoInfo): Promise<FileSlice[]>;
289
316
  export declare function buildBatchRequest(lens: LensDefinition, info: RepoInfo, slice: FileSlice, index: number, sliceCount: number, model?: string, maxTokensOverride?: number): BatchRequest;
290
- export declare function estimateCost(lens: LensDefinition, slices: FileSlice[], pricing: ModelPricing, maxTokensOverride?: number): {
317
+ /**
318
+ * Pre-flight cost estimate for one lens.
319
+ *
320
+ * Every slice is its own batch request, so both halves scale with the slice
321
+ * count. The output half used to be a single `maxTokens * 0.75` for the whole
322
+ * lens no matter how many requests it sent — on a repository that sliced into
323
+ * 13 modules that budgeted one request's output and shipped thirteen, and a
324
+ * live run came in at roughly 3x its estimate. Since this number is what
325
+ * `max_cost` binds against, under-counting it lets a run outspend the cap the
326
+ * user set.
327
+ *
328
+ * @param info - Repo info, when the caller has it: lets the estimate include
329
+ * the system prompt and JSON schema each request carries. Omitted, the
330
+ * estimate covers slice content only, which is what the old signature did.
331
+ */
332
+ export declare function estimateCost(lens: LensDefinition, slices: FileSlice[], pricing: ModelPricing, maxTokensOverride?: number, info?: RepoInfo): {
291
333
  inputTokens: number;
292
334
  outputTokens: number;
293
335
  cost: number;
@@ -313,7 +355,40 @@ export declare function readBroadsideSkill(cwd: string): Promise<{
313
355
  }>;
314
356
  export declare function defaultBroadsideState(): BroadsideStateFile;
315
357
  export declare function loadBroadsideState(broadsideDir: string): Promise<BroadsideStateFile>;
358
+ /**
359
+ * Overwrite `state.json` wholesale with `state`.
360
+ *
361
+ * Prefer {@link persistBroadsideRun} anywhere a live operation is recording its
362
+ * own progress — this entry point replaces the file, so any run a concurrent
363
+ * process recorded in the meantime is erased. It remains the right call for
364
+ * seeding a fresh workspace and for test fixtures, where "make the file exactly
365
+ * this" is the intent.
366
+ */
316
367
  export declare function saveBroadsideState(broadsideDir: string, state: BroadsideStateFile): Promise<void>;
368
+ /**
369
+ * Read-modify-write `state.json` under a lock.
370
+ *
371
+ * The lock is held only for the read-modify-write, never for the surrounding
372
+ * operation: a `collect` can poll for the better part of an hour, and holding
373
+ * the lock across that would push every concurrent caller past the 5s lock
374
+ * timeout.
375
+ */
376
+ export declare function updateBroadsideStateAtomically(broadsideDir: string, mutate: (state: BroadsideStateFile) => void | Promise<void>): Promise<BroadsideStateFile>;
377
+ /**
378
+ * Record one run's current shape, merged into whatever is on disk *now*.
379
+ *
380
+ * Broad-Side operations are long-lived and hold their state in memory while
381
+ * they poll. Writing that snapshot back wholesale silently erased any run a
382
+ * concurrent operation had recorded since it was loaded, orphaning that run's
383
+ * paid results on disk — present as files, invisible to `list`, and unreachable
384
+ * by `collect`, which finds its run by position in `state.runs`. Observed live:
385
+ * a submit at 23:35 was erased by a collect that had loaded state before it and
386
+ * wrote back at 00:08.
387
+ *
388
+ * Merging by run id also self-heals: a run erased by an older writer is
389
+ * restored the next time its own operation checkpoints.
390
+ */
391
+ export declare function persistBroadsideRun(broadsideDir: string, run: BroadsideRun): Promise<BroadsideStateFile>;
317
392
  export declare function loadBroadsideConfig(broadsideDir: string): Promise<BroadsideConfig>;
318
393
  export declare function builtInCatalogEntry(model: string): CatalogEntry | null;
319
394
  export declare function builtInPricing(model: string): ModelPricing | null;
@@ -336,6 +411,14 @@ export declare function submitBatch(batchRequests: BatchRequest[], apiKey: strin
336
411
  error?: unknown;
337
412
  }>;
338
413
  export declare function fetchBatch(batchId: string, apiKey: string, fetcher?: FetchLike): Promise<Record<string, unknown>>;
414
+ /**
415
+ * Batch statuses that will never produce a result.
416
+ *
417
+ * Deliberately excludes the synthetic `timeout` this module returns when a poll
418
+ * budget expires: that batch is still running server-side and has already been
419
+ * charged, so callers must come back for it rather than retire it.
420
+ */
421
+ export declare const BROADSIDE_DEAD_BATCH_STATUSES: string[];
339
422
  export declare function pollBatchUntilTerminal(batchId: string, apiKey: string, opts?: {
340
423
  deadlineMs?: number;
341
424
  onStatus?: (status: string, counts: Record<string, unknown>) => void;
@@ -398,6 +481,18 @@ export type StoredLensResult = {
398
481
  */
399
482
  export declare function parseLensJson(content: string): unknown | null;
400
483
  export declare function saveLensResults(runDir: string, lensId: BroadsideLensId, batch: Record<string, unknown>): Promise<StoredLensResult[]>;
484
+ /**
485
+ * Rebuild lens results from what a previous collect already wrote to disk.
486
+ *
487
+ * The post-passes are gated on having lens findings in hand, and a collect
488
+ * only holds the ones *it* polled. When an earlier collect saved every lens
489
+ * and then died before synthesis and triage ran — the batch window is long
490
+ * and a poll can easily be interrupted — the next collect finds every lens
491
+ * already terminal, skips them all, and would otherwise reach the post-pass
492
+ * gate with nothing to hand it. Reading the saved results back is what makes
493
+ * "a resumed collect can finish whichever is still pending" true.
494
+ */
495
+ export declare function loadSavedLensResults(runDir: string, lenses: BroadsideLensId[]): Promise<StoredLensResult[]>;
401
496
  export declare function runBroadsideCollect(cwd: string, apiKey: string, opts?: {
402
497
  waitMs?: number;
403
498
  includeSynthesis?: boolean;