codecartographer-pi 0.18.0 → 0.19.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codecarto/GUIDE.md +7 -0
- package/.codecarto/templates/phase-handoff.yaml +20 -1
- package/.codecarto/workflow/scaffold-version.yaml +1 -1
- package/agent-skill/codecartographer/references/handoff-contract.md +30 -2
- package/dist/core/completion.d.ts +15 -0
- package/dist/core/completion.js +80 -3
- package/dist/core/coverage.d.ts +45 -0
- package/dist/core/coverage.js +131 -0
- package/dist/core/index.d.ts +1 -0
- package/dist/core/index.js +1 -0
- package/dist/core/prompts.js +19 -1
- package/dist/core/status.d.ts +8 -1
- package/dist/core/status.js +37 -2
- package/dist/core/types.d.ts +24 -1
- package/package.json +1 -1
package/.codecarto/GUIDE.md
CHANGED
|
@@ -215,6 +215,13 @@ post_pipeline:
|
|
|
215
215
|
|
|
216
216
|
`kind` is one of: `needs-runtime-test`, `needs-maintainer-decision`, `needs-spec-ruling`, `defer-to-phase`, `needs-fixture-capture`, or a post-pipeline work kind such as `spike` or `amendment`. Every new `carry_forward.target_phase` must be an ID in the active pipeline. Every `post_pipeline` entry requires a stable ID. Open questions should carry a stable `id` (e.g. `q-loadconfig-ambiguity`); if omitted, the framework auto-assigns one. When a later phase resolves an open question, list its id in `open_question_closures` to remove it from all phases. The downstream phase records resolved carry-forward IDs in `carry_forward_closures`; completion removes those entries atomically.
|
|
217
217
|
|
|
218
|
+
**Closing a routed item does not settle the question it came from.** A phase routinely registers an open question and routes one of its candidate answers onward in the same handoff; addressing the routed item later is not the same as answering the question. Two optional fields make that distinction enforceable:
|
|
219
|
+
|
|
220
|
+
- A `carry_forward` entry may name `derives_from: <open-question-id>`, meaning "this routed item is one candidate answer to that question." Fill it in the handoff that registers the question, when both entries are in front of you. Completion then refuses a `carry_forward_closures` entry for that item while its question is still open and is not closed by the same handoff, naming both ids. Close the question with evidence in that handoff, or leave the item routed and give the finding an unsettled action (`verify at runtime`). Entirely opt-in: a `carry_forward` entry without `derives_from` closes exactly as it always has.
|
|
221
|
+
- An `open_question_closures` entry may be written as `{ id, evidence }` instead of a bare id, where `evidence` names what settled the question. For a question whose `kind` is `needs-runtime-test`, non-empty `evidence` is **required** — such a question closes on a spike report or an observation against the running system, not on another read of the same source — and that requirement holds however the closure is written, so a bare id no longer closes a runtime question. Questions of every other kind are unaffected, and a bare id stays valid for them. The requirement follows the scaffold: a workspace at scaffold version 0.19.0 or newer has the closure refused, while an older scaffold — which never documented the rule — gets a non-gating note, and refreshing the scaffold opts it in.
|
|
222
|
+
|
|
223
|
+
Completed phases' declared coverage gaps travel too: the `Skipped scope` and `Known blind spots` bullets of every completed phase's `## Coverage and limits` section are surfaced in the next phase's orchestrator duties. A finding that lands inside one of those gaps must either close it with cited new evidence of its own or inherit its uncertainty.
|
|
224
|
+
|
|
218
225
|
After the pipeline completes, the handoff channel closes with it. Post-pipeline resolutions — an open question answered on evidence, a finished `post_pipeline` backlog item — are applied with an **amendment**: write `scratch/amendments/<slug>.yaml` (see `templates/amendment.yaml`) and run `codecarto_amend` (MCP) or `/codecarto-amend <slug>` (Pi). It updates `workflow/status.yaml` under the same lock completion uses and writes an amendment closeout plus THREAD_LOG entry. Never hand-edit `status.yaml` for this; amendments are refused while the pipeline is still running, so the two channels cannot race.
|
|
219
226
|
|
|
220
227
|
## Phase Selection Logic
|
|
@@ -4,9 +4,28 @@ schema_version: 1
|
|
|
4
4
|
phase_id: <phase-id>
|
|
5
5
|
owner_notes: []
|
|
6
6
|
open_questions: []
|
|
7
|
+
# Items a specific later phase closes. Every entry needs a target_phase naming a
|
|
8
|
+
# downstream phase in the active pipeline. Optional derives_from names the
|
|
9
|
+
# open_questions id this item is one candidate answer to:
|
|
10
|
+
# - id: mech-CF3
|
|
11
|
+
# kind: defer-to-phase
|
|
12
|
+
# target_phase: defect-scan-semantic
|
|
13
|
+
# derives_from: q-logit-bias-root-cause
|
|
14
|
+
# description: <the routed item>
|
|
15
|
+
# Fill derives_from in the handoff that registers the question — that is the one
|
|
16
|
+
# moment both entries are in front of you. Closing the routed item later does not
|
|
17
|
+
# settle the question it came from, and completion refuses such a closure while
|
|
18
|
+
# the question is still open and unclosed by the same handoff.
|
|
7
19
|
carry_forward: []
|
|
8
20
|
carry_forward_closures: []
|
|
9
|
-
# Resolved open questions: their IDs are removed from all phases.
|
|
21
|
+
# Resolved open questions: their IDs are removed from all phases. Either a bare
|
|
22
|
+
# id, or an object naming the evidence that settled it:
|
|
23
|
+
# - q-loadconfig-ambiguity
|
|
24
|
+
# - id: q-logit-bias-root-cause
|
|
25
|
+
# evidence: <the spike report or runtime observation that answered it>
|
|
26
|
+
# A question of kind needs-runtime-test requires non-empty evidence — it closes
|
|
27
|
+
# on runtime evidence, not on another read of the same source. Refused from
|
|
28
|
+
# scaffold version 0.19.0; older scaffolds get a non-gating note instead.
|
|
10
29
|
open_question_closures: []
|
|
11
30
|
# Work after the active pipeline: spikes, amendments, deltas, maintainer rulings,
|
|
12
31
|
# or opinionated reruns. Every entry requires a stable id.
|
|
@@ -17,7 +17,7 @@ owner_notes: [] # 2-3 durable observations; appended to the phas
|
|
|
17
17
|
open_questions: [] # genuinely unknown, no later phase will close them
|
|
18
18
|
carry_forward: [] # deferred to a specific later phase in this pipeline
|
|
19
19
|
carry_forward_closures: [] # ids of carry_forward entries this phase resolved
|
|
20
|
-
open_question_closures: [] #
|
|
20
|
+
open_question_closures: [] # open questions this phase resolved, removed everywhere; bare id or {id, evidence}
|
|
21
21
|
post_pipeline: [] # work after the pipeline; every entry needs a stable id
|
|
22
22
|
decisions: [] # choices made beyond what the prompt specified; completion appends them to DECISIONS.md
|
|
23
23
|
proposed_conventions: [] # patterns proposed for promotion; completion stages them in CONVENTIONS.md
|
|
@@ -41,7 +41,7 @@ Omitted arrays default to empty. A malformed collection fails completion without
|
|
|
41
41
|
deferred_reason: Distinguishing them needs a runtime probe this phase cannot run.
|
|
42
42
|
```
|
|
43
43
|
|
|
44
|
-
`carry_forward` entries add `target_phase`:
|
|
44
|
+
`carry_forward` entries add `target_phase`, and optionally `derives_from`:
|
|
45
45
|
|
|
46
46
|
```yaml
|
|
47
47
|
- id: arch-CF2
|
|
@@ -49,10 +49,18 @@ Omitted arrays default to empty. A malformed collection fails completion without
|
|
|
49
49
|
target_phase: protocols
|
|
50
50
|
description: MCP endpoints listed by name only; schemas not extracted.
|
|
51
51
|
deferred_reason: Wire-format extraction is the protocols phase's rubric.
|
|
52
|
+
|
|
53
|
+
- id: mech-CF3
|
|
54
|
+
kind: defer-to-phase
|
|
55
|
+
target_phase: defect-scan-semantic
|
|
56
|
+
derives_from: q-logit-bias-root-cause # optional: the open question this is one candidate answer to
|
|
57
|
+
description: The client sends logit_bias as a map; the documented shape is an array.
|
|
52
58
|
```
|
|
53
59
|
|
|
54
60
|
Allowed `kind` values: `needs-runtime-test`, `needs-maintainer-decision`, `needs-spec-ruling`, `defer-to-phase`, `needs-fixture-capture`.
|
|
55
61
|
|
|
62
|
+
`derives_from` names an `open_questions` id. Fill it in the handoff that registers the question — a phase that routes a candidate answer onward usually writes both entries at once, which is the one moment both are in front of you. It is optional and additive: a handoff that omits it behaves exactly as before.
|
|
63
|
+
|
|
56
64
|
`proposed_conventions` entries (optional; omitted defaults to empty):
|
|
57
65
|
|
|
58
66
|
```yaml
|
|
@@ -81,6 +89,26 @@ Completion then removes the entry atomically. Resolving an open question works t
|
|
|
81
89
|
|
|
82
90
|
Re-deferring instead of closing means writing a fresh `carry_forward` entry naming a later `target_phase`.
|
|
83
91
|
|
|
92
|
+
### Closing a routed item does not settle the question it came from
|
|
93
|
+
|
|
94
|
+
Addressing what was routed to you is not the same as answering the question that produced it. Two rules make that enforceable:
|
|
95
|
+
|
|
96
|
+
- **A closure whose `derives_from` question is still open is refused.** If the item you are closing declares `derives_from: <question-id>`, that question is still in `status.yaml`, and this same handoff does not close it, completion refuses and names both ids. Either close the question here with the evidence that settles it, or leave the item routed and give the finding an unsettled action (`verify at runtime`) so it inherits the question's uncertainty. This rule is entirely opt-in — an entry without `derives_from` closes as it always has.
|
|
97
|
+
- **A `needs-runtime-test` question closes on runtime evidence.** Write the closure as an object and say where that evidence lives:
|
|
98
|
+
|
|
99
|
+
```yaml
|
|
100
|
+
open_question_closures:
|
|
101
|
+
- q-loadconfig-ambiguity # bare id: still valid for any other kind
|
|
102
|
+
- id: q-logit-bias-root-cause
|
|
103
|
+
evidence: scratch/spikes/logit-bias.md — probe against llama-server b4321
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
Non-empty `evidence` is required when the question's `kind` is `needs-runtime-test`, whether the closure is written as a bare id or as an object. It is checked for presence, not judged — a spike report or an observation against the running system is what belongs there, and another read of the same source is not. Questions of every other kind close on a bare id exactly as before.
|
|
107
|
+
|
|
108
|
+
The requirement is scoped to the scaffold that documents it. A workspace whose `workflow/scaffold-version.yaml` is 0.19.0 or newer has completion refuse such a closure; an older or unversioned scaffold — whose own templates never stated the rule — gets a non-gating `NOTE:` instead, so an in-flight run written against the older contract cannot be stopped by a rule it was never told. Refreshing the scaffold (`codecarto_refresh_scaffold` on MCP, `/codecarto-refresh-scaffold` on Pi) opts a workspace in.
|
|
109
|
+
|
|
110
|
+
Upstream coverage gaps travel the same way, without gating: the `Skipped scope` and `Known blind spots` bullets of every completed phase's `## Coverage and limits` section appear in your phase prompt's orchestrator duties. A finding of yours inside one of those gaps must either close it with cited new evidence or inherit its uncertainty.
|
|
111
|
+
|
|
84
112
|
## The failure this prevents
|
|
85
113
|
|
|
86
114
|
Before completion required a handoff, a phase could finish with empty state and no signal. A real seven-phase run documented five cross-phase routings in its report prose, wrote no handoffs, and completed all seven phases with `carry_forward: []` throughout. Every downstream phase's routed-item intake was empty. The findings survived only because each phase happened to re-read the previous phase's full markdown.
|
|
@@ -24,4 +24,19 @@ export declare const CONVENTIONS_PENDING_HEADING = "## Pending proposals";
|
|
|
24
24
|
* @param content - the file content, or null to read from disk (null when the file is absent).
|
|
25
25
|
*/
|
|
26
26
|
export declare function countPendingProposals(workspaceDir: string, content?: string): Promise<number>;
|
|
27
|
+
/**
|
|
28
|
+
* First scaffold version whose handoff template documents the `{id, evidence}`
|
|
29
|
+
* closure shape and the runtime-evidence rule (#122, #186).
|
|
30
|
+
*
|
|
31
|
+
* A workspace scaffolded before it was never told that closing a
|
|
32
|
+
* `needs-runtime-test` question requires evidence, so refusing its completion
|
|
33
|
+
* would apply a rule its own templates do not carry — and could stop an
|
|
34
|
+
* in-flight `--auto` run on a handoff written against the older contract.
|
|
35
|
+
* Those workspaces get a warning instead; refreshing the scaffold
|
|
36
|
+
* (`codecarto_refresh_scaffold` on MCP, `/codecarto-refresh-scaffold` on Pi)
|
|
37
|
+
* opts them in. Same treatment Stage 2's findings pairing check uses.
|
|
38
|
+
*/
|
|
39
|
+
export declare const CLOSURE_EVIDENCE_GATE_SCAFFOLD_VERSION = "0.19.0";
|
|
40
|
+
/** Whether the runtime-evidence requirement refuses (current scaffold) or warns (older). */
|
|
41
|
+
export declare function closureEvidenceGateActive(scaffoldVersion: string | undefined | null): boolean;
|
|
27
42
|
export declare function completeValidatedPhase(cwd: string, validation: ValidationResult, sourceLabel: string): Promise<CompletionResult>;
|
package/dist/core/completion.js
CHANGED
|
@@ -2,7 +2,7 @@ import { appendFile, copyFile, mkdir, readFile, readdir, writeFile } from "node:
|
|
|
2
2
|
import { join } from "node:path";
|
|
3
3
|
import { getNextEligiblePhase, resolvePhase, validatePhaseOutput } from "./pipeline.js";
|
|
4
4
|
import { applyHandoff, autoAssignIds, buildTerminalNextActions, loadHandoffFile, normalizeStatus } from "./status.js";
|
|
5
|
-
import { dateOnly, newlineIfUnterminated, pathExists, uniqueStrings } from "./utils.js";
|
|
5
|
+
import { compareDottedVersions, dateOnly, newlineIfUnterminated, pathExists, uniqueStrings } from "./utils.js";
|
|
6
6
|
import { getWorkspaceState, updateStatusAtomically } from "./workspace.js";
|
|
7
7
|
/**
|
|
8
8
|
* The Markdown a reader sees: content inside `<!-- -->` blocks removed by a
|
|
@@ -233,6 +233,26 @@ async function writeCompletionArtifacts(workspaceDir, phaseId, validation, times
|
|
|
233
233
|
const { totalPending: totalPendingProposals } = await stageProposedConventions(workspaceDir, phaseId, timestamp, handoff?.proposed_conventions ?? []);
|
|
234
234
|
return { closeoutPath: `.codecarto/closeouts/${closeoutFile}`, decisionsAppended, totalPendingProposals };
|
|
235
235
|
}
|
|
236
|
+
/**
|
|
237
|
+
* First scaffold version whose handoff template documents the `{id, evidence}`
|
|
238
|
+
* closure shape and the runtime-evidence rule (#122, #186).
|
|
239
|
+
*
|
|
240
|
+
* A workspace scaffolded before it was never told that closing a
|
|
241
|
+
* `needs-runtime-test` question requires evidence, so refusing its completion
|
|
242
|
+
* would apply a rule its own templates do not carry — and could stop an
|
|
243
|
+
* in-flight `--auto` run on a handoff written against the older contract.
|
|
244
|
+
* Those workspaces get a warning instead; refreshing the scaffold
|
|
245
|
+
* (`codecarto_refresh_scaffold` on MCP, `/codecarto-refresh-scaffold` on Pi)
|
|
246
|
+
* opts them in. Same treatment Stage 2's findings pairing check uses.
|
|
247
|
+
*/
|
|
248
|
+
export const CLOSURE_EVIDENCE_GATE_SCAFFOLD_VERSION = "0.19.0";
|
|
249
|
+
/** Whether the runtime-evidence requirement refuses (current scaffold) or warns (older). */
|
|
250
|
+
export function closureEvidenceGateActive(scaffoldVersion) {
|
|
251
|
+
if (!scaffoldVersion)
|
|
252
|
+
return false;
|
|
253
|
+
const comparison = compareDottedVersions(scaffoldVersion, CLOSURE_EVIDENCE_GATE_SCAFFOLD_VERSION);
|
|
254
|
+
return comparison !== null && comparison >= 0;
|
|
255
|
+
}
|
|
236
256
|
export async function completeValidatedPhase(cwd, validation, sourceLabel) {
|
|
237
257
|
const initialState = await getWorkspaceState(cwd);
|
|
238
258
|
if (!initialState)
|
|
@@ -254,6 +274,7 @@ export async function completeValidatedPhase(cwd, validation, sourceLabel) {
|
|
|
254
274
|
+ `plus closeout_summary and optional closeout_content. Then re-run completion.`);
|
|
255
275
|
}
|
|
256
276
|
}
|
|
277
|
+
const warnings = [];
|
|
257
278
|
if (handoff) {
|
|
258
279
|
const activePhases = new Set(initialState.pipeline.phase_order);
|
|
259
280
|
const sourceIndex = initialState.pipeline.phase_order.indexOf(validation.phaseId);
|
|
@@ -267,14 +288,70 @@ export async function completeValidatedPhase(cwd, validation, sourceLabel) {
|
|
|
267
288
|
if (!entry.id?.trim())
|
|
268
289
|
throw new Error("Invalid handoff: post_pipeline entries require a canonical id");
|
|
269
290
|
}
|
|
291
|
+
// Closure integrity, gating (#122, #186). Both checks are deterministic
|
|
292
|
+
// reads of ids the model wrote itself, so neither can wedge an --auto run
|
|
293
|
+
// on a heuristic; both sit here, before the lock, alongside the
|
|
294
|
+
// target_phase check, so a refusal mutates nothing.
|
|
295
|
+
const questionsById = new Map();
|
|
296
|
+
const derivesFromById = new Map();
|
|
297
|
+
for (const phaseState of Object.values(initialState.status.phases)) {
|
|
298
|
+
for (const entry of phaseState.open_questions ?? []) {
|
|
299
|
+
if (entry.id)
|
|
300
|
+
questionsById.set(entry.id, entry);
|
|
301
|
+
}
|
|
302
|
+
for (const entry of phaseState.carry_forward ?? []) {
|
|
303
|
+
if (entry.id && entry.derives_from)
|
|
304
|
+
derivesFromById.set(entry.id, entry.derives_from);
|
|
305
|
+
}
|
|
306
|
+
}
|
|
307
|
+
const closingQuestionIds = new Set(handoff.open_question_closures.map((closure) => closure.id).filter(Boolean));
|
|
308
|
+
// D1: a routed item that declares `derives_from` is one candidate answer
|
|
309
|
+
// to that question. Closing it while the question stands is exactly the
|
|
310
|
+
// contradiction #122 reported — the routed candidate shipped as settled
|
|
311
|
+
// while the question that said "source alone cannot determine which" was
|
|
312
|
+
// still open. A derives_from naming an id that no longer exists is fine:
|
|
313
|
+
// the question was already resolved.
|
|
314
|
+
for (const closureId of handoff.carry_forward_closures) {
|
|
315
|
+
const questionId = derivesFromById.get(closureId);
|
|
316
|
+
if (!questionId || !questionsById.has(questionId))
|
|
317
|
+
continue;
|
|
318
|
+
if (closingQuestionIds.has(questionId))
|
|
319
|
+
continue;
|
|
320
|
+
throw new Error(`Refusing to complete ${validation.phaseId}: the handoff closes carry_forward ${closureId}, which derives_from open question ${questionId} — and ${questionId} is still unresolved and is not in this handoff's open_question_closures. `
|
|
321
|
+
+ `A routed item is one candidate answer to the question it came from; closing it does not settle the question. `
|
|
322
|
+
+ `Either close ${questionId} in this same handoff with the evidence that settles it, or leave ${closureId} routed and give the finding an unsettled action ("verify at runtime") instead.`);
|
|
323
|
+
}
|
|
324
|
+
// D3: a `needs-runtime-test` question closes on runtime evidence, not on
|
|
325
|
+
// another source read. Requiring the evidence string to be non-empty is
|
|
326
|
+
// the whole gate — judging what it says stays prose guidance.
|
|
327
|
+
//
|
|
328
|
+
// Unlike D1, this one can fire on a handoff that uses none of the new
|
|
329
|
+
// fields: a bare-string closure was the only shape before this release.
|
|
330
|
+
// Gating it unconditionally would apply a rule to workspaces whose own
|
|
331
|
+
// templates never state it, so it is version-gated exactly like Stage
|
|
332
|
+
// 2's pairing check — refuse on a scaffold that documents the rule, warn
|
|
333
|
+
// on one that predates it.
|
|
334
|
+
const evidenceGateActive = closureEvidenceGateActive(initialState.scaffoldVersion);
|
|
335
|
+
for (const closure of handoff.open_question_closures) {
|
|
336
|
+
if (questionsById.get(closure.id)?.kind !== "needs-runtime-test")
|
|
337
|
+
continue;
|
|
338
|
+
if (closure.evidence?.trim())
|
|
339
|
+
continue;
|
|
340
|
+
const detail = `open_question_closures closes ${closure.id}, whose kind is needs-runtime-test, without evidence. `
|
|
341
|
+
+ `A runtime question closes on runtime evidence — a spike report or an observation against the running system — not on a source read. `
|
|
342
|
+
+ `Write the closure as an object: { id: ${closure.id}, evidence: <where that evidence lives> }. If you do not have it, leave the question open.`;
|
|
343
|
+
if (evidenceGateActive)
|
|
344
|
+
throw new Error(`Refusing to complete ${validation.phaseId}: ${detail}`);
|
|
345
|
+
warnings.push(`${detail} Warning only: this workspace's scaffold predates the requirement — refresh it `
|
|
346
|
+
+ `(codecarto_refresh_scaffold on MCP, /codecarto-refresh-scaffold on Pi) to make this gating.`);
|
|
347
|
+
}
|
|
270
348
|
}
|
|
271
349
|
// Closure integrity (#122, warning only): a handoff can close a carry-forward
|
|
272
350
|
// or open question the report never addressed — "closed in the handoff,
|
|
273
351
|
// resolved nowhere." The id of every claimed closure should appear somewhere
|
|
274
352
|
// in the primary output that claims to resolve it.
|
|
275
|
-
const warnings = [];
|
|
276
353
|
if (handoff && validation.outputPath) {
|
|
277
|
-
const closures = [...handoff.carry_forward_closures, ...handoff.open_question_closures].filter((id) => id?.trim());
|
|
354
|
+
const closures = [...handoff.carry_forward_closures, ...handoff.open_question_closures.map((closure) => closure.id)].filter((id) => id?.trim());
|
|
278
355
|
if (closures.length > 0) {
|
|
279
356
|
const output = await readFile(validation.outputPath, "utf8").catch(() => "");
|
|
280
357
|
const unmentioned = closures.filter((id) => !output.includes(id));
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
import type { WorkspaceState } from "./types.ts";
|
|
2
|
+
/** The section heading whose bullets this module reads. */
|
|
3
|
+
export declare const COVERAGE_SECTION_HEADING = "Coverage and limits";
|
|
4
|
+
/** The ledger's five fixed bullets, verbatim values with the label stripped. */
|
|
5
|
+
export type CoverageLedger = {
|
|
6
|
+
inspected_scope: string;
|
|
7
|
+
skipped_scope: string;
|
|
8
|
+
evidence_basis: string;
|
|
9
|
+
known_blind_spots: string;
|
|
10
|
+
coverage_disposition: string;
|
|
11
|
+
};
|
|
12
|
+
/** One declared gap from one completed phase's ledger. */
|
|
13
|
+
export type CoverageGap = {
|
|
14
|
+
/** The phase that declared it. */
|
|
15
|
+
phaseId: string;
|
|
16
|
+
/** Which bullet it came from: "skipped scope" or "known blind spots". */
|
|
17
|
+
label: string;
|
|
18
|
+
/** The bullet's text, sub-bullets folded onto one line. */
|
|
19
|
+
detail: string;
|
|
20
|
+
/** `.codecarto/`-relative path of the output the ledger was read from. */
|
|
21
|
+
output: string;
|
|
22
|
+
};
|
|
23
|
+
/**
|
|
24
|
+
* Read a phase output's `## Coverage and limits` ledger.
|
|
25
|
+
*
|
|
26
|
+
* Returns null when the document has no such section. A bullet the author left
|
|
27
|
+
* blank comes back as an empty string, as does a label the section omits, so a
|
|
28
|
+
* caller never has to distinguish "absent" from "empty" — both mean nothing to
|
|
29
|
+
* carry forward.
|
|
30
|
+
*
|
|
31
|
+
* Sub-bullets and wrapped continuation lines under a label belong to that
|
|
32
|
+
* label: a real report writes its blind spots as a nested list, and dropping
|
|
33
|
+
* them would silence exactly the case this exists for. They fold onto one line
|
|
34
|
+
* (sub-bullets joined with "; ") because the consumer is a prompt bullet.
|
|
35
|
+
*/
|
|
36
|
+
export declare function parseCoverageAndLimits(content: string): CoverageLedger | null;
|
|
37
|
+
/**
|
|
38
|
+
* Every declared gap from every completed phase whose primary output exists.
|
|
39
|
+
*
|
|
40
|
+
* Walks `phase_order` so the list is deterministic and reads upstream-first.
|
|
41
|
+
* Only `Skipped scope` and `Known blind spots` are collected: those are the
|
|
42
|
+
* two bullets that bind a later phase's claims. Unreadable or unparsable
|
|
43
|
+
* outputs contribute nothing.
|
|
44
|
+
*/
|
|
45
|
+
export declare function collectCoverageGaps(state: WorkspaceState): Promise<CoverageGap[]>;
|
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
// The coverage-gap ledger every primary output carries (#122, #186).
|
|
2
|
+
//
|
|
3
|
+
// Each phase output ends with a `## Coverage and limits` section whose bullet
|
|
4
|
+
// labels are fixed across every template — `- Inspected scope:`, `- Skipped
|
|
5
|
+
// scope:`, `- Evidence basis:`, `- Known blind spots:`, `- Coverage
|
|
6
|
+
// disposition:` — which is what makes reading it a parse and not a guess.
|
|
7
|
+
//
|
|
8
|
+
// The ledger was carried nowhere: not into the next phase's prompt, not into
|
|
9
|
+
// status.yaml, not into validation. So an upstream phase could declare "the
|
|
10
|
+
// encoded search-proxy command was not fully decoded" and the next phase could
|
|
11
|
+
// assert an `observed fact` about that exact component, with nothing in the
|
|
12
|
+
// framework comparing the two. The contradiction sweep could not have caught
|
|
13
|
+
// it either: that sweep compares against `owner_notes`, and a declared blind
|
|
14
|
+
// spot is not an owner note.
|
|
15
|
+
//
|
|
16
|
+
// Everything here is non-gating and read-only. A missing file, a missing
|
|
17
|
+
// section, or an empty bullet yields nothing; nothing throws.
|
|
18
|
+
import { readFile } from "node:fs/promises";
|
|
19
|
+
import { join } from "node:path";
|
|
20
|
+
import { pathExists } from "./utils.js";
|
|
21
|
+
/** The section heading whose bullets this module reads. */
|
|
22
|
+
export const COVERAGE_SECTION_HEADING = "Coverage and limits";
|
|
23
|
+
/** Bullet label (normalized) to ledger field. */
|
|
24
|
+
const LEDGER_LABELS = new Map([
|
|
25
|
+
["inspected scope", "inspected_scope"],
|
|
26
|
+
["skipped scope", "skipped_scope"],
|
|
27
|
+
["evidence basis", "evidence_basis"],
|
|
28
|
+
["known blind spots", "known_blind_spots"],
|
|
29
|
+
["coverage disposition", "coverage_disposition"],
|
|
30
|
+
]);
|
|
31
|
+
/** The two ledger bullets a downstream phase is bound by, in render order. */
|
|
32
|
+
const GAP_FIELDS = [
|
|
33
|
+
{ field: "skipped_scope", label: "skipped scope" },
|
|
34
|
+
{ field: "known_blind_spots", label: "known blind spots" },
|
|
35
|
+
];
|
|
36
|
+
function emptyLedger() {
|
|
37
|
+
return {
|
|
38
|
+
inspected_scope: "",
|
|
39
|
+
skipped_scope: "",
|
|
40
|
+
evidence_basis: "",
|
|
41
|
+
known_blind_spots: "",
|
|
42
|
+
coverage_disposition: "",
|
|
43
|
+
};
|
|
44
|
+
}
|
|
45
|
+
function normalizeLabel(label) {
|
|
46
|
+
return label.replace(/[`*_]/g, "").replace(/\s+/g, " ").trim().toLowerCase();
|
|
47
|
+
}
|
|
48
|
+
/**
|
|
49
|
+
* Read a phase output's `## Coverage and limits` ledger.
|
|
50
|
+
*
|
|
51
|
+
* Returns null when the document has no such section. A bullet the author left
|
|
52
|
+
* blank comes back as an empty string, as does a label the section omits, so a
|
|
53
|
+
* caller never has to distinguish "absent" from "empty" — both mean nothing to
|
|
54
|
+
* carry forward.
|
|
55
|
+
*
|
|
56
|
+
* Sub-bullets and wrapped continuation lines under a label belong to that
|
|
57
|
+
* label: a real report writes its blind spots as a nested list, and dropping
|
|
58
|
+
* them would silence exactly the case this exists for. They fold onto one line
|
|
59
|
+
* (sub-bullets joined with "; ") because the consumer is a prompt bullet.
|
|
60
|
+
*/
|
|
61
|
+
export function parseCoverageAndLimits(content) {
|
|
62
|
+
const lines = content.split(/\r?\n/);
|
|
63
|
+
const start = lines.findIndex((line) => /^##\s+Coverage and limits\s*$/i.test(line.replace(/[`*_]/g, "")));
|
|
64
|
+
if (start < 0)
|
|
65
|
+
return null;
|
|
66
|
+
const ledger = emptyLedger();
|
|
67
|
+
let current = null;
|
|
68
|
+
for (let i = start + 1; i < lines.length; i++) {
|
|
69
|
+
const line = lines[i];
|
|
70
|
+
if (/^##\s/.test(line))
|
|
71
|
+
break;
|
|
72
|
+
const bullet = /^[-*]\s+(.*)$/.exec(line);
|
|
73
|
+
if (bullet) {
|
|
74
|
+
const separator = bullet[1].indexOf(":");
|
|
75
|
+
const field = separator >= 0 ? LEDGER_LABELS.get(normalizeLabel(bullet[1].slice(0, separator))) : undefined;
|
|
76
|
+
// An unrecognized top-level bullet ends the previous label's block
|
|
77
|
+
// rather than absorbing text that belongs to neither.
|
|
78
|
+
current = field ?? null;
|
|
79
|
+
if (field)
|
|
80
|
+
ledger[field] = bullet[1].slice(separator + 1).trim();
|
|
81
|
+
continue;
|
|
82
|
+
}
|
|
83
|
+
if (!line.trim())
|
|
84
|
+
continue; // a blank line does not end a label's block
|
|
85
|
+
if (!current || !/^\s/.test(line)) {
|
|
86
|
+
current = null; // unindented prose is not part of any bullet
|
|
87
|
+
continue;
|
|
88
|
+
}
|
|
89
|
+
const continuation = line.trim();
|
|
90
|
+
const nested = /^[-*]\s+(.*)$/.exec(continuation);
|
|
91
|
+
const text = (nested ? nested[1] : continuation).trim();
|
|
92
|
+
if (!text)
|
|
93
|
+
continue;
|
|
94
|
+
ledger[current] = ledger[current] ? `${ledger[current]}${nested ? "; " : " "}${text}` : text;
|
|
95
|
+
}
|
|
96
|
+
return ledger;
|
|
97
|
+
}
|
|
98
|
+
/**
|
|
99
|
+
* Every declared gap from every completed phase whose primary output exists.
|
|
100
|
+
*
|
|
101
|
+
* Walks `phase_order` so the list is deterministic and reads upstream-first.
|
|
102
|
+
* Only `Skipped scope` and `Known blind spots` are collected: those are the
|
|
103
|
+
* two bullets that bind a later phase's claims. Unreadable or unparsable
|
|
104
|
+
* outputs contribute nothing.
|
|
105
|
+
*/
|
|
106
|
+
export async function collectCoverageGaps(state) {
|
|
107
|
+
const configs = new Map(state.pipeline.phases.map((phase) => [phase.id, phase]));
|
|
108
|
+
const gaps = [];
|
|
109
|
+
for (const phaseId of state.pipeline.phase_order ?? []) {
|
|
110
|
+
if (state.status.phases[phaseId]?.status !== "complete")
|
|
111
|
+
continue;
|
|
112
|
+
const output = configs.get(phaseId)?.primary_output;
|
|
113
|
+
if (!output)
|
|
114
|
+
continue;
|
|
115
|
+
const outputPath = join(state.workspaceDir, output);
|
|
116
|
+
if (!(await pathExists(outputPath)))
|
|
117
|
+
continue;
|
|
118
|
+
const content = await readFile(outputPath, "utf8").catch(() => null);
|
|
119
|
+
if (content === null)
|
|
120
|
+
continue;
|
|
121
|
+
const ledger = parseCoverageAndLimits(content);
|
|
122
|
+
if (!ledger)
|
|
123
|
+
continue;
|
|
124
|
+
for (const { field, label } of GAP_FIELDS) {
|
|
125
|
+
const detail = ledger[field];
|
|
126
|
+
if (detail)
|
|
127
|
+
gaps.push({ phaseId, label, detail, output });
|
|
128
|
+
}
|
|
129
|
+
}
|
|
130
|
+
return gaps;
|
|
131
|
+
}
|
package/dist/core/index.d.ts
CHANGED
package/dist/core/index.js
CHANGED
|
@@ -8,6 +8,7 @@ export * from "./status.js";
|
|
|
8
8
|
export * from "./amendment.js";
|
|
9
9
|
export * from "./pipeline.js";
|
|
10
10
|
export * from "./findings.js";
|
|
11
|
+
export * from "./coverage.js";
|
|
11
12
|
export * from "./prompts.js";
|
|
12
13
|
export * from "./workspace.js";
|
|
13
14
|
export * from "./completion.js";
|
package/dist/core/prompts.js
CHANGED
|
@@ -4,12 +4,17 @@
|
|
|
4
4
|
import { readdir } from "node:fs/promises";
|
|
5
5
|
import { join } from "node:path";
|
|
6
6
|
import { pathExists } from "./utils.js";
|
|
7
|
+
import { collectCoverageGaps } from "./coverage.js";
|
|
7
8
|
import { countPendingProposals } from "./completion.js";
|
|
8
9
|
import { describeScaffoldStaleness } from "./workspace.js";
|
|
9
10
|
import { runPhasePreflight } from "./synthesis.js";
|
|
10
11
|
/** Open-question kinds whose label the orchestrator re-tests at each phase boundary. */
|
|
11
12
|
const RETRIAGE_KINDS = new Set(["needs-maintainer-decision", "needs-runtime-test"]);
|
|
12
|
-
/**
|
|
13
|
+
/**
|
|
14
|
+
* Cap on the individually listed entries of a mechanically surfaced duty list
|
|
15
|
+
* — re-triage questions and upstream coverage gaps alike; the rest collapse to
|
|
16
|
+
* a count naming where the full list lives.
|
|
17
|
+
*/
|
|
13
18
|
const RETRIAGE_LIST_LIMIT = 10;
|
|
14
19
|
/**
|
|
15
20
|
* Build the "Orchestrator duties" prompt block (issue #98): the cross-phase
|
|
@@ -47,6 +52,19 @@ async function buildOrchestratorDuties(state, phase, auto) {
|
|
|
47
52
|
lines.push(` - .codecarto/${output.path} (${exists ? "exists" : "missing"})`);
|
|
48
53
|
}
|
|
49
54
|
}
|
|
55
|
+
// The coverage-gap ledger (#122, #186). The contradiction sweep below
|
|
56
|
+
// compares against `owner_notes` only, and a declared blind spot is not an
|
|
57
|
+
// owner note — so a phase could assert an observed fact about a component
|
|
58
|
+
// the upstream phase had recorded as not decoded, and nothing compared the
|
|
59
|
+
// two. Non-gating: this is a duty in the prompt, not a validation rule.
|
|
60
|
+
const coverageGaps = await collectCoverageGaps(state);
|
|
61
|
+
if (coverageGaps.length > 0) {
|
|
62
|
+
lines.push("- Upstream phases declared these coverage gaps in their `## Coverage and limits` sections. A finding of yours that lands inside one must either close the gap with cited new evidence of its own or inherit its uncertainty — an upstream `not inspected` or `not decoded` does not license an `observed fact` about that scope:");
|
|
63
|
+
for (const gap of coverageGaps.slice(0, RETRIAGE_LIST_LIMIT))
|
|
64
|
+
lines.push(` - ${gap.phaseId} (${gap.label}): ${gap.detail}`);
|
|
65
|
+
if (coverageGaps.length > RETRIAGE_LIST_LIMIT)
|
|
66
|
+
lines.push(` - (+${coverageGaps.length - RETRIAGE_LIST_LIMIT} more in completed phases' Coverage and limits sections)`);
|
|
67
|
+
}
|
|
50
68
|
const anyCompleted = Object.values(state.status.phases).some((phaseState) => phaseState.status === "complete");
|
|
51
69
|
if (anyCompleted) {
|
|
52
70
|
lines.push("- Contradiction sweep: compare this phase's required reads against completed phases' owner_notes; a measured fact that contradicts a summarized claim is a gap to route through the handoff, not a nuance to smooth over.");
|
package/dist/core/status.d.ts
CHANGED
|
@@ -1,9 +1,16 @@
|
|
|
1
|
-
import type { NormalizedStatus, OpenQuestionEntry, PostPipelineEntry, PhaseHandoff, PipelineFile, ProposedConventionEntry, StatusFile, StatusPhase } from "./types.ts";
|
|
1
|
+
import type { ClosureEntry, NormalizedStatus, OpenQuestionEntry, PostPipelineEntry, PhaseHandoff, PipelineFile, ProposedConventionEntry, StatusFile, StatusPhase } from "./types.ts";
|
|
2
2
|
export declare const LOCK_RETRY_MS = 125;
|
|
3
3
|
export declare const LOCK_TIMEOUT_MS = 5000;
|
|
4
4
|
export declare const STALE_LOCK_MS = 60000;
|
|
5
5
|
export declare function assertSafePhaseId(phaseId: string): void;
|
|
6
6
|
export declare function ensureArray(value: unknown): string[];
|
|
7
|
+
/**
|
|
8
|
+
* Normalize a handoff's `open_question_closures` (#122, #186). Accepts both
|
|
9
|
+
* the original bare-string shape and `{ id, evidence }`; a string becomes
|
|
10
|
+
* `{ id }`, and an entry with no usable id is dropped rather than resolving
|
|
11
|
+
* nothing under the lock. Values are trimmed.
|
|
12
|
+
*/
|
|
13
|
+
export declare function ensureClosureArray(value: unknown): ClosureEntry[];
|
|
7
14
|
export declare function ensureEntryArray<T extends OpenQuestionEntry>(value: unknown, allowTargetPhase?: boolean): T[];
|
|
8
15
|
export declare function autoAssignIds(entries: OpenQuestionEntry[], prefix: string, phaseId: string): void;
|
|
9
16
|
export declare function ensurePostPipelineArray(value: unknown): PostPipelineEntry[];
|
package/dist/core/status.js
CHANGED
|
@@ -36,8 +36,42 @@ function coerceEntry(value, allowTargetPhase) {
|
|
|
36
36
|
entry.deferred_reason = raw.deferred_reason.trim();
|
|
37
37
|
if (allowTargetPhase && typeof raw.target_phase === "string" && raw.target_phase.trim())
|
|
38
38
|
entry.target_phase = raw.target_phase.trim();
|
|
39
|
+
// derives_from rides the same flag as target_phase: it is a carry-forward
|
|
40
|
+
// concept only — the id of the open question this routed item answers one
|
|
41
|
+
// candidate of (#122, #186). An open_questions entry has nothing to derive
|
|
42
|
+
// from, so the field is dropped there rather than silently carried.
|
|
43
|
+
if (allowTargetPhase && typeof raw.derives_from === "string" && raw.derives_from.trim())
|
|
44
|
+
entry.derives_from = raw.derives_from.trim();
|
|
39
45
|
return Object.keys(entry).length > 0 ? entry : null;
|
|
40
46
|
}
|
|
47
|
+
/**
|
|
48
|
+
* Normalize a handoff's `open_question_closures` (#122, #186). Accepts both
|
|
49
|
+
* the original bare-string shape and `{ id, evidence }`; a string becomes
|
|
50
|
+
* `{ id }`, and an entry with no usable id is dropped rather than resolving
|
|
51
|
+
* nothing under the lock. Values are trimmed.
|
|
52
|
+
*/
|
|
53
|
+
export function ensureClosureArray(value) {
|
|
54
|
+
if (!Array.isArray(value))
|
|
55
|
+
return [];
|
|
56
|
+
const result = [];
|
|
57
|
+
for (const item of value) {
|
|
58
|
+
if (typeof item === "string") {
|
|
59
|
+
const id = item.trim();
|
|
60
|
+
if (id)
|
|
61
|
+
result.push({ id });
|
|
62
|
+
continue;
|
|
63
|
+
}
|
|
64
|
+
if (!item || typeof item !== "object" || Array.isArray(item))
|
|
65
|
+
continue;
|
|
66
|
+
const raw = item;
|
|
67
|
+
const id = typeof raw.id === "string" ? raw.id.trim() : "";
|
|
68
|
+
if (!id)
|
|
69
|
+
continue;
|
|
70
|
+
const evidence = typeof raw.evidence === "string" && raw.evidence.trim() ? raw.evidence.trim() : undefined;
|
|
71
|
+
result.push({ id, ...(evidence !== undefined && { evidence }) });
|
|
72
|
+
}
|
|
73
|
+
return result;
|
|
74
|
+
}
|
|
41
75
|
export function ensureEntryArray(value, allowTargetPhase = false) {
|
|
42
76
|
if (!Array.isArray(value))
|
|
43
77
|
return [];
|
|
@@ -218,7 +252,7 @@ export function parseHandoff(value) {
|
|
|
218
252
|
open_questions: openQuestions,
|
|
219
253
|
carry_forward: carryForward,
|
|
220
254
|
carry_forward_closures: ensureArray(raw.carry_forward_closures),
|
|
221
|
-
open_question_closures:
|
|
255
|
+
open_question_closures: ensureClosureArray(raw.open_question_closures),
|
|
222
256
|
post_pipeline: ensurePostPipelineArray(raw.post_pipeline),
|
|
223
257
|
decisions: ensureArray(raw.decisions),
|
|
224
258
|
proposed_conventions: ensureProposedConventionArray(raw.proposed_conventions),
|
|
@@ -335,7 +369,8 @@ export function applyHandoff(status, handoff) {
|
|
|
335
369
|
}
|
|
336
370
|
}
|
|
337
371
|
// Apply open_question_closures: remove resolved questions from ALL phases by id
|
|
338
|
-
for (const
|
|
372
|
+
for (const closure of handoff.open_question_closures) {
|
|
373
|
+
const closureId = closure?.id;
|
|
339
374
|
if (!closureId)
|
|
340
375
|
continue;
|
|
341
376
|
for (const ph of Object.values(status.phases)) {
|
package/dist/core/types.d.ts
CHANGED
|
@@ -9,6 +9,16 @@ export type OpenQuestionEntry = {
|
|
|
9
9
|
};
|
|
10
10
|
export type CarryForwardEntry = OpenQuestionEntry & {
|
|
11
11
|
target_phase?: string;
|
|
12
|
+
/**
|
|
13
|
+
* Optional id of the `open_questions` entry this routed item is one
|
|
14
|
+
* candidate answer to (#122, #186). The upstream phase usually registers
|
|
15
|
+
* the question and routes the candidate onward in the same handoff, which
|
|
16
|
+
* is the cheapest moment to record the link. Completion refuses a closure
|
|
17
|
+
* of this entry while that question is still open and is not closed by the
|
|
18
|
+
* same handoff: closing a routed item does not settle the question it came
|
|
19
|
+
* from. Omitted on every entry that predates the field.
|
|
20
|
+
*/
|
|
21
|
+
derives_from?: string;
|
|
12
22
|
};
|
|
13
23
|
export type PostPipelineEntry = OpenQuestionEntry & {
|
|
14
24
|
source_phase?: string;
|
|
@@ -114,6 +124,18 @@ export type ProposedConventionEntry = {
|
|
|
114
124
|
/** Optional: where the pattern showed up (file, phase, incident). */
|
|
115
125
|
evidence?: string;
|
|
116
126
|
};
|
|
127
|
+
/**
|
|
128
|
+
* One claimed closure in a handoff's `open_question_closures` (#122, #186).
|
|
129
|
+
* A bare string — the only shape before this field grew — normalizes to
|
|
130
|
+
* `{ id }`; the object form adds the evidence that settles the question.
|
|
131
|
+
* Completion requires non-empty `evidence` when the closed question's `kind`
|
|
132
|
+
* is `needs-runtime-test`, because such a question closes on runtime evidence
|
|
133
|
+
* rather than on another source read.
|
|
134
|
+
*/
|
|
135
|
+
export type ClosureEntry = {
|
|
136
|
+
id: string;
|
|
137
|
+
evidence?: string;
|
|
138
|
+
};
|
|
117
139
|
export type PhaseHandoff = {
|
|
118
140
|
phase_id: string;
|
|
119
141
|
/**
|
|
@@ -126,7 +148,8 @@ export type PhaseHandoff = {
|
|
|
126
148
|
open_questions: OpenQuestionEntry[];
|
|
127
149
|
carry_forward: CarryForwardEntry[];
|
|
128
150
|
carry_forward_closures: string[];
|
|
129
|
-
|
|
151
|
+
/** ids to resolve and remove from all phases; a bare string parses to `{ id }`. */
|
|
152
|
+
open_question_closures: ClosureEntry[];
|
|
130
153
|
post_pipeline: PostPipelineEntry[];
|
|
131
154
|
decisions: string[];
|
|
132
155
|
/** Conventions proposed for promotion; completion stages them in CONVENTIONS.md. Omitted defaults to empty. */
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "codecartographer-pi",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.19.0",
|
|
4
4
|
"mcpName": "io.github.HuginnIndustries/codecartographer",
|
|
5
5
|
"description": "Turn an unfamiliar codebase into a validated reimplementation spec, then synthesize confirmed specs and a product vision into a traceable plan.",
|
|
6
6
|
"type": "module",
|