codecartographer-pi 0.17.0 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (66) hide show
  1. package/.codecarto/GUIDE.md +3 -3
  2. package/.codecarto/findings/architecture/SKILL.md +1 -0
  3. package/.codecarto/findings/contracts/SKILL.md +1 -0
  4. package/.codecarto/findings/defect-scan/SKILL.md +15 -1
  5. package/.codecarto/findings/defect-scan/passes/01-logic-and-correctness.md +1 -1
  6. package/.codecarto/findings/defect-scan/passes/02-error-handling.md +1 -1
  7. package/.codecarto/findings/defect-scan/passes/03-concurrency-and-resources.md +1 -1
  8. package/.codecarto/findings/defect-scan/passes/04-security-and-trust.md +1 -1
  9. package/.codecarto/findings/defect-scan/passes/05-api-contract-violations.md +2 -1
  10. package/.codecarto/findings/defect-scan/passes/06-config-and-environment.md +1 -1
  11. package/.codecarto/findings/defect-scan-mechanical/SKILL.md +3 -2
  12. package/.codecarto/findings/defect-scan-semantic/SKILL.md +2 -0
  13. package/.codecarto/findings/porting/SKILL.md +2 -1
  14. package/.codecarto/findings/protocols/SKILL.md +1 -0
  15. package/.codecarto/findings/reimplementation-spec/SKILL.md +1 -1
  16. package/.codecarto/skills/spec-delta-application/SKILL.md +1 -1
  17. package/.codecarto/templates/amendment.yaml +3 -3
  18. package/.codecarto/templates/architecture-map.md +1 -1
  19. package/.codecarto/templates/defect-report.md +23 -0
  20. package/.codecarto/templates/mechanical-defects.md +22 -0
  21. package/.codecarto/templates/reverse-engineering-bundle.md +2 -2
  22. package/.codecarto/templates/semantic-defects.md +26 -0
  23. package/.codecarto/templates/spike-report.md +2 -2
  24. package/.codecarto/workflow/VALIDATE.md +1 -1
  25. package/.codecarto/workflow/pipeline-defect-scan.yaml +2 -0
  26. package/.codecarto/workflow/pipeline-full-with-audit.yaml +3 -1
  27. package/.codecarto/workflow/pipeline-full-with-deep-audit.yaml +5 -1
  28. package/.codecarto/workflow/pipeline-scout-first.yaml +5 -1
  29. package/.codecarto/workflow/scaffold-version.yaml +1 -1
  30. package/README.md +14 -10
  31. package/agent-skill/codecartographer/SKILL.md +1 -1
  32. package/agent-skill/codecartographer/references/deep-audit-synthesis.md +4 -1
  33. package/agent-skill/codecartographer/references/library.md +3 -3
  34. package/agent-skill/codecartographer/references/orchestration.md +1 -1
  35. package/agent-skill/codecartographer/references/phase-recovery.md +1 -1
  36. package/dist/core/amendment.d.ts +5 -0
  37. package/dist/core/amendment.js +23 -4
  38. package/dist/core/completion.d.ts +5 -0
  39. package/dist/core/completion.js +18 -2
  40. package/dist/core/dashboard.js +5 -3
  41. package/dist/core/findings.d.ts +59 -0
  42. package/dist/core/findings.js +145 -0
  43. package/dist/core/index.d.ts +1 -0
  44. package/dist/core/index.js +1 -0
  45. package/dist/core/library.d.ts +139 -1
  46. package/dist/core/library.js +291 -40
  47. package/dist/core/orchestrator-config.d.ts +6 -0
  48. package/dist/core/orchestrator-config.js +2 -0
  49. package/dist/core/pipeline.js +15 -0
  50. package/dist/core/prompts.js +1 -1
  51. package/dist/core/status.d.ts +2 -1
  52. package/dist/core/status.js +33 -11
  53. package/dist/core/types.d.ts +6 -0
  54. package/dist/core/utils.d.ts +14 -0
  55. package/dist/core/utils.js +30 -0
  56. package/dist/core/workspace.d.ts +16 -0
  57. package/dist/core/workspace.js +42 -22
  58. package/dist/core/yaml.js +19 -4
  59. package/dist/extensions/codecarto/auto-runner.d.ts +2 -0
  60. package/dist/extensions/codecarto/broadside-flags.d.ts +7 -2
  61. package/dist/extensions/codecarto/broadside-flags.js +22 -9
  62. package/dist/extensions/codecarto/dashboard-narrator.js +1 -1
  63. package/dist/extensions/codecarto/index.js +379 -20
  64. package/dist/extensions/codecarto/phase-compaction.js +4 -0
  65. package/dist/mcp-server/server.js +127 -10
  66. package/package.json +2 -2
@@ -17,7 +17,7 @@ CodeCartographer distinguishes two roles. They are different jobs — but they a
17
17
  **Orchestrator.** The persistent chat driving the run — normally the very session reading this guide. The role is defined by its **duties**, not by who executes phases:
18
18
 
19
19
  - **Curate `CONVENTIONS.md` and `DECISIONS.md`**: promote patterns when they recur; append cross-cutting decisions as they land.
20
- - **Re-triage open questions at every phase boundary.** A question's `kind` label is itself a claim that needs evidence: before accepting `needs-maintainer-decision` or `needs-runtime-test`, re-test whether the question has become answerable by reading. A mislabeled question suppresses verification for every later phase (the `orchestration` guide topic records a real four-phase failure).
20
+ - **Re-triage open questions at every phase boundary.** A question's `kind` label is itself a claim that needs evidence: before accepting `needs-maintainer-decision` or `needs-runtime-test`, re-test whether the question has become answerable by reading. A mislabeled question suppresses verification for every later phase (the `orchestration` guide topic records a real four-phase failure). The duty has a counterpart: when re-triage concludes a question genuinely still needs a runtime test, no finding in the next phase may assert one of that question's candidate answers with a settled action (`fix before porting`, `fix now`) — it inherits the question's uncertainty (`verify at runtime`) until runtime evidence closes the question.
21
21
  - **Sweep for contradictions** between the incoming phase's required reads and earlier phases' `owner_notes`. A measured fact that contradicts a summarized claim is a gap to route, not a nuance to smooth over.
22
22
  - **Route gaps**: confirm each completed phase's declared secondary outputs were written or explicitly routed, and that handoff `decisions` deferring work landed somewhere a later phase will actually see.
23
23
  - **Gate strategic forks** with the user: pipeline switches, opinionated-vs-agnostic specs, scope changes.
@@ -92,7 +92,7 @@ If you are uncertain whether a file should be modified, treat it as read-only.
92
92
  ### Two things named "backlog", and neither is the other
93
93
 
94
94
  - **`BACKLOG.md` in this workspace** is *this project's* deferrals: work the project chose not to do yet. `DECISIONS.md` records what the project decided to **do**; `BACKLOG.md` records what it decided to **defer**. Deferrals get no `D` number. When a deferred item is later picked up, remove its entry here and record the decision in `DECISIONS.md`.
95
- - **`status.yaml`'s `post_pipeline` list** is framework-owned lifecycle state, not this file. Items there are retired by `codecarto_amend`, never by hand.
95
+ - **`status.yaml`'s `post_pipeline` list** is framework-owned lifecycle state, not this file. Items there are retired by an amendment (`codecarto_amend` on MCP, `/codecarto-amend` on Pi), never by hand.
96
96
  - **CodeCartographer's own backlog** — deferred improvements to the *framework* — lives in the CodeCartographer repository, not in your workspace. If a phase prompt misled you or a validation criterion did not fit, that is feedback to the framework; it does not belong in this file.
97
97
 
98
98
  ## Pipeline Selection
@@ -215,7 +215,7 @@ post_pipeline:
215
215
 
216
216
  `kind` is one of: `needs-runtime-test`, `needs-maintainer-decision`, `needs-spec-ruling`, `defer-to-phase`, `needs-fixture-capture`, or a post-pipeline work kind such as `spike` or `amendment`. Every new `carry_forward.target_phase` must be an ID in the active pipeline. Every `post_pipeline` entry requires a stable ID. Open questions should carry a stable `id` (e.g. `q-loadconfig-ambiguity`); if omitted, the framework auto-assigns one. When a later phase resolves an open question, list its id in `open_question_closures` to remove it from all phases. The downstream phase records resolved carry-forward IDs in `carry_forward_closures`; completion removes those entries atomically.
217
217
 
218
- After the pipeline completes, the handoff channel closes with it. Post-pipeline resolutions — an open question answered on evidence, a finished `post_pipeline` backlog item — are applied with an **amendment**: write `scratch/amendments/<slug>.yaml` (see `templates/amendment.yaml`) and run `codecarto_amend`. It updates `workflow/status.yaml` under the same lock completion uses and writes an amendment closeout plus THREAD_LOG entry. Never hand-edit `status.yaml` for this; amendments are refused while the pipeline is still running, so the two channels cannot race.
218
+ After the pipeline completes, the handoff channel closes with it. Post-pipeline resolutions — an open question answered on evidence, a finished `post_pipeline` backlog item — are applied with an **amendment**: write `scratch/amendments/<slug>.yaml` (see `templates/amendment.yaml`) and run `codecarto_amend` (MCP) or `/codecarto-amend <slug>` (Pi). It updates `workflow/status.yaml` under the same lock completion uses and writes an amendment closeout plus THREAD_LOG entry. Never hand-edit `status.yaml` for this; amendments are refused while the pipeline is still running, so the two channels cannot race.
219
219
 
220
220
  ## Phase Selection Logic
221
221
 
@@ -84,6 +84,7 @@ Write the output in seven sections:
84
84
  Mark every conclusion with one of these evidence levels:
85
85
  - `observed fact`: direct statement from docs, tests, schemas, types, or code.
86
86
  - `strong inference`: architectural conclusion drawn from multiple facts.
87
+ - `external-behavior claim`: a claim about what a system outside this source tree does (a server, engine, driver, third-party API, OS); unverifiable by reading this code, so it stays unsettled until a runtime probe or that system's own source confirms it.
87
88
  - `portability hazard`: assumption tied to the source language, runtime, terminal, OS, or third-party SDKs.
88
89
  - `open question`: missing or conflicting behavior that still needs evidence.
89
90
 
@@ -79,6 +79,7 @@ End with a black-box acceptance list:
79
79
  Mark every finding with one of these evidence levels:
80
80
  - `observed fact`: direct statement from docs, tests, schemas, types, or code.
81
81
  - `strong inference`: behavioral conclusion drawn from multiple facts.
82
+ - `external-behavior claim`: a claim about what a system outside this source tree does (a server, engine, driver, third-party API, OS); unverifiable by reading this code, so it stays unsettled until a runtime probe or that system's own source confirms it.
82
83
  - `portability hazard`: assumption tied to the source language, runtime, terminal, OS, or third-party SDKs.
83
84
  - `open question`: missing or conflicting behavior that still needs evidence.
84
85
 
@@ -45,8 +45,11 @@ Not every pass applies to every codebase. Use the architecture map to decide emp
45
45
  Mark every finding with one of these evidence levels:
46
46
  - `observed fact`: the defect is directly visible in the code.
47
47
  - `strong inference`: the defect is highly likely based on multiple code observations.
48
+ - `external-behavior claim`: the claim is about a component outside the analyzed source tree — a server, engine, driver, third-party API, or OS — including how it parses a payload, what it silently ignores, and version-dependent behavior. Not obtainable at any read depth of this source: settling it needs a runtime probe against the pinned version, or that system's own source at that version. This is not a weaker `strong inference`; it is a claim about a different artifact. "The client sends a map where the server's API documents an array" is an `external-behavior claim` about the server, even when both halves of the sentence are observed facts.
48
49
  - `open question`: the code is suspicious but would need runtime testing to confirm.
49
50
 
51
+ **Cite or hedge quantities.** Any number in a finding — a file size, a default value, a call-site count, a version, a timeout — cites the file and line or the command output it was read from, or is written as an explicit estimate ("~300 MB, not measured"). A plausible specific stated in the same register as a read one is how a 32.8 MB artifact gets recorded as ~300 MB and a default of -20 as -5.
52
+
50
53
  ## Severity Classification
51
54
 
52
55
  Assign one severity per finding:
@@ -59,10 +62,11 @@ Assign one severity per finding:
59
62
 
60
63
  Tag each finding with a recommended action. Use the set that matches your pipeline:
61
64
 
62
- **Pre-porting pipelines** (full-with-audit, full-with-deep-audit):
65
+ **Pre-porting pipelines** (full-with-audit, full-with-deep-audit, scout-first):
63
66
  - `fix before porting`: the defect would carry into a new implementation if not addressed first.
64
67
  - `port differently`: the new implementation should handle this case differently by design.
65
68
  - `leave behind`: the defect is specific to the source implementation and won't survive porting.
69
+ - `verify at runtime`: the diagnosis names behavior of a system outside the analyzed source, or otherwise cannot be settled by reading; a runtime probe must confirm it before any fix is designed. Its destination is the spec's Spike List plus a `post_pipeline` entry of `kind: spike` — not a design change.
66
70
 
67
71
  **Maintenance pipelines** (defect-scan):
68
72
  - `fix now`: the defect is actively causing or risking problems.
@@ -72,6 +76,16 @@ Tag each finding with a recommended action. Use the set that matches your pipeli
72
76
 
73
77
  Determine which action set to use by checking the `pipeline` field in `workflow/status.yaml`.
74
78
 
79
+ ### Pairing rules — the evidence level bounds the action
80
+
81
+ A finding's action may not assert more certainty than its evidence level carries:
82
+
83
+ - Evidence `open question` or `external-behavior claim` ⇒ action is `verify at runtime` or `port differently` (pre-porting), or `investigate` (maintenance). **Never `fix before porting` or `fix now`.** A fix designed from an unsettled diagnosis can convert working behavior into the one shape the external system ignores — a real run recommended exactly that reshape, and runtime testing inverted it.
84
+ - Evidence `observed fact` ⇒ action is not `verify at runtime`. A settled label with an unsettled action contradicts itself; pick the one that is true.
85
+ - Every finding whose evidence is `open question` or `external-behavior claim` also appears in the report's **Open Questions** table, so the hedge travels with the finding into every document a later phase or a human reads — not only into the handoff.
86
+
87
+ Validation checks the first rule mechanically on the findings tables; on a current scaffold a violation fails the phase. If a routed `carry_forward` item derives from an open question of `kind: needs-runtime-test` that is still unresolved, the finding that addresses it inherits that uncertainty: it closes the carry-forward with `verify at runtime`, not with a settled action, unless the question itself is closed with runtime evidence in the same handoff.
88
+
75
89
  ## Output
76
90
 
77
91
  Write findings to the primary output using the template at `templates/defect-report.md`.
@@ -41,7 +41,7 @@ For each finding, record:
41
41
  - **Defect**: what is wrong, in one sentence.
42
42
  - **Evidence**: what you observed that proves or strongly suggests the defect.
43
43
  - **Severity**: critical (incorrect results in normal use), high (incorrect results in edge cases), medium (dead code or latent risk), low (style issue with correctness implications).
44
- - **Evidence level**: observed fact / strong inference / open question.
44
+ - **Evidence level**: observed fact / strong inference / external-behavior claim / open question. A claim about what another system does with this code's output (a server, engine, driver, third-party API, OS) is an `external-behavior claim`, not a `strong inference` — see findings/defect-scan/SKILL.md §Evidence Classification.
45
45
 
46
46
  ## What to skip
47
47
 
@@ -47,7 +47,7 @@ For each finding, record:
47
47
  - **Defect**: what error scenario is mishandled.
48
48
  - **Evidence**: the specific code pattern or path that demonstrates the gap.
49
49
  - **Severity**: critical (data loss or corruption on failure), high (silent failure in normal operations), medium (poor error messages or missing cleanup), low (observability gap).
50
- - **Evidence level**: observed fact / strong inference / open question.
50
+ - **Evidence level**: observed fact / strong inference / external-behavior claim / open question. A claim about what another system does with this code's output (a server, engine, driver, third-party API, OS) is an `external-behavior claim`, not a `strong inference` — see findings/defect-scan/SKILL.md §Evidence Classification.
51
51
 
52
52
  ## What to skip
53
53
 
@@ -46,7 +46,7 @@ For each finding, record:
46
46
  - **Defect**: what concurrent or resource scenario is mishandled.
47
47
  - **Evidence**: the specific shared state, missing synchronization, or unclosed resource.
48
48
  - **Severity**: critical (data corruption or deadlock in normal operation), high (race condition in common paths), medium (resource leak under error conditions), low (theoretical race in rarely-exercised path).
49
- - **Evidence level**: observed fact / strong inference / open question.
49
+ - **Evidence level**: observed fact / strong inference / external-behavior claim / open question. A claim about what another system does with this code's output (a server, engine, driver, third-party API, OS) is an `external-behavior claim`, not a `strong inference` — see findings/defect-scan/SKILL.md §Evidence Classification.
50
50
 
51
51
  ## What to skip
52
52
 
@@ -54,7 +54,7 @@ For each finding, record:
54
54
  - **Defect**: what security property is violated.
55
55
  - **Evidence**: the specific code path, missing check, or exposed secret.
56
56
  - **Severity**: critical (actively exploitable, data exposure, auth bypass), high (exploitable with some effort or preconditions), medium (defense-in-depth gap, hardcoded non-production secret), low (missing header, informational disclosure).
57
- - **Evidence level**: observed fact / strong inference / open question.
57
+ - **Evidence level**: observed fact / strong inference / external-behavior claim / open question. A claim about what another system does with this code's output (a server, engine, driver, third-party API, OS) is an `external-behavior claim`, not a `strong inference` — see findings/defect-scan/SKILL.md §Evidence Classification.
58
58
 
59
59
  ## What to skip
60
60
 
@@ -47,9 +47,10 @@ For each finding, record:
47
47
  - **Location**: file path and function/method name.
48
48
  - **Defect**: what contract is violated and how.
49
49
  - **Spec source**: where the expected behavior is documented (docstring, type signature, contracts phase, protocols phase, README).
50
+ - **Which side implements the contract**: the analyzed code is often the *caller* of a contract another system implements (an HTTP API, an inference engine, a driver). A divergence between what this code sends and what the other side documents is only an `observed fact` about this code; what the other side actually does with it — accepts, ignores, rejects — is an `external-behavior claim` until a runtime probe against the pinned version says otherwise, and its action is `verify at runtime`, never `fix before porting`.
50
51
  - **Evidence**: the specific divergence between spec and implementation.
51
52
  - **Severity**: critical (public API returns wrong results), high (documented behavior incorrect in edge cases), medium (internal API inconsistency), low (stale docs, minor parameter mismatch).
52
- - **Evidence level**: observed fact / strong inference / open question.
53
+ - **Evidence level**: observed fact / strong inference / external-behavior claim / open question. A claim about what another system does with this code's output (a server, engine, driver, third-party API, OS) is an `external-behavior claim`, not a `strong inference` — see findings/defect-scan/SKILL.md §Evidence Classification.
53
54
 
54
55
  ## What to skip
55
56
 
@@ -50,7 +50,7 @@ For each finding, record:
50
50
  - **Defect**: what configuration or environment hazard exists.
51
51
  - **Evidence**: the specific hardcoded value, missing validation, or dangerous default.
52
52
  - **Severity**: critical (security exposure in default config, data loss on misconfiguration), high (production failure from missing validation), medium (hardcoded value that will break in a different environment), low (undocumented config behavior).
53
- - **Evidence level**: observed fact / strong inference / open question.
53
+ - **Evidence level**: observed fact / strong inference / external-behavior claim / open question. A claim about what another system does with this code's output (a server, engine, driver, third-party API, OS) is an `external-behavior claim`, not a `strong inference` — see findings/defect-scan/SKILL.md §Evidence Classification.
54
54
 
55
55
  ## What to skip
56
56
 
@@ -39,9 +39,10 @@ Use the architecture map to decide emphasis:
39
39
 
40
40
  Use the same scheme as the legacy defect-scan SKILL:
41
41
 
42
- - **Evidence levels:** `observed fact`, `strong inference`, `open question`.
42
+ - **Evidence levels:** `observed fact`, `strong inference`, `external-behavior claim`, `open question`.
43
43
  - **Severity:** `critical`, `high`, `medium`, `low`.
44
- - **Action (pre-porting pipelines):** `fix before porting`, `port differently`, `leave behind`.
44
+ - **Action (pre-porting pipelines):** `fix before porting`, `port differently`, `leave behind`, `verify at runtime`.
45
+ - **Pairing rule:** `open question` or `external-behavior claim` evidence takes `verify at runtime` or `port differently`, never `fix before porting`, and the finding also appears in the report's Open Questions table. Validation checks this on the findings tables.
45
46
 
46
47
  See `findings/defect-scan/SKILL.md` for the full criteria; this phase intentionally does not duplicate them.
47
48
 
@@ -47,6 +47,8 @@ When citing a contract or protocol violation, include the contract ID or state-m
47
47
 
48
48
  Read the `carry_forward` entries in `workflow/status.yaml` whose `target_phase` is `defect-scan-semantic`. The mechanical phase may have routed semantic-flavored sightings here for closure.
49
49
 
50
+ **Closing a routed item does not settle the question it came from.** Before closing a carry-forward, check whether it derives from an `open_questions` entry of `kind: needs-runtime-test` that is still unresolved — the mechanical phase typically registers the question and routes one of its candidate explanations onward in the same handoff. If the question stands, the finding that addresses the carry-forward inherits its uncertainty: evidence `external-behavior claim` or `open question`, action `verify at runtime`, and a row in this report's Open Questions table naming the question's id. Asserting one of the question's candidates as `strong inference` / `fix before porting` while the question remains open is the contradiction this rule exists to prevent — a real run did exactly that, and runtime testing inverted the finding. Only runtime evidence (or the external system's own source at the pinned version) closes such a question; when you have it, list the question in `open_question_closures` and cite the evidence in the finding.
51
+
50
52
  ## Output
51
53
 
52
54
  Write findings to the primary output using `templates/semantic-defects.md`. Organize by pass, then by severity within each pass. End with a summary table covering only passes 3, 4, and 5.
@@ -19,6 +19,7 @@ Keep four classes of findings separate throughout:
19
19
  - `observed fact`: direct statements from docs, tests, schemas, types, and code.
20
20
  - `strong inference`: architectural conclusions drawn from multiple facts.
21
21
  - `portability hazard`: assumptions tied to the source language, runtime, terminal, OS, or third-party SDKs.
22
+ - `external-behavior claim`: a claim about what a system outside this source tree does (a server, engine, driver, third-party API, OS) — unverifiable at any read depth here; carry it forward as unsettled, never promote it to a fact.
22
23
  - `open question`: missing or conflicting behavior that still needs evidence.
23
24
 
24
25
  Prefer concept names over source names:
@@ -39,7 +40,7 @@ Sort features by porting importance:
39
40
 
40
41
  If the defect report is available, integrate defect findings into the porting bundle:
41
42
  - Reference relevant defects in the feature contract table.
42
- - Tag each referenced defect with a porting recommendation: `fix before porting` (the defect would carry into a new implementation), `port differently` (the new implementation should handle this case differently by design), or `leave behind` (the defect is specific to the source implementation and won't survive porting).
43
+ - Tag each referenced defect with a porting recommendation: `fix before porting` (the defect would carry into a new implementation), `port differently` (the new implementation should handle this case differently by design), `leave behind` (the defect is specific to the source implementation and won't survive porting), or `verify at runtime` (the diagnosis is an `external-behavior claim` or `open question` — carry it as a spike for the spec, and do not design around an unverified diagnosis). Preserve `verify at runtime` as written: flattening it into one of the settled three is how a hedge stops traveling.
43
44
  - Consolidate defect-related portability hazards alongside hazards from other phases.
44
45
 
45
46
  Use the output template at `templates/reverse-engineering-bundle.md`. Produce:
@@ -77,6 +77,7 @@ Produce four outputs:
77
77
  Mark every finding with one of these evidence levels:
78
78
  - `observed fact`: direct statement from docs, tests, schemas, types, or code.
79
79
  - `strong inference`: protocol conclusion drawn from multiple facts.
80
+ - `external-behavior claim`: a claim about what the peer or a system outside this source tree does with a message (how it parses, what it ignores, version-dependent behavior); unverifiable by reading this code, so it stays unsettled until a runtime capture or that system's own source confirms it.
80
81
  - `portability hazard`: assumption tied to the source language, runtime, terminal, OS, or third-party SDKs.
81
82
  - `open question`: missing or conflicting behavior that still needs evidence.
82
83
 
@@ -69,7 +69,7 @@ End with a spike list:
69
69
  - risky performance assumptions
70
70
  - platform-sensitive areas that need targeted tests
71
71
 
72
- For every defect in the bundle, preserve its disposition (`fix before porting`, `port differently`, or `leave behind`) and convert it into an explicit design consequence or acceptance check.
72
+ For every defect in the bundle, preserve its disposition (`fix before porting`, `port differently`, `leave behind`, or `verify at runtime`) and convert it into an explicit design consequence or acceptance check — except `verify at runtime`, which becomes a Spike List entry and a `post_pipeline` entry of `kind: spike` in your handoff, never a design consequence: the diagnosis has not been confirmed, and designing around it would build the unverified claim into the new system.
73
73
 
74
74
  Use the output template at `templates/reimplementation-spec.md`.
75
75
 
@@ -88,7 +88,7 @@ Per the standard closeout ritual:
88
88
 
89
89
  - Append a one-line entry to `THREAD_LOG.md` pointing at the closeout file.
90
90
  - Write `closeouts/<YYYY-MM-DD>-spec-deltas.md` using `templates/closeout-template.md`.
91
- - If a delta resolved an `open_questions` entry (or finished a `post_pipeline` backlog item), apply it: write `scratch/amendments/<slug>.yaml` (see `templates/amendment.yaml`) listing the closures, then run `codecarto_amend`. It updates `status.yaml` under the completion lock and writes the amendment closeout — never hand-edit `status.yaml`, which is framework-owned. Without the MCP server, record the intended amendment file in the closeout for the next MCP-capable session to apply.
91
+ - If a delta resolved an `open_questions` entry (or finished a `post_pipeline` backlog item), apply it: write `scratch/amendments/<slug>.yaml` (see `templates/amendment.yaml`) listing the closures, then run `codecarto_amend` (MCP) or `/codecarto-amend <slug>` (Pi). It updates `status.yaml` under the completion lock and writes the amendment closeout — never hand-edit `status.yaml`, which is framework-owned. Without Pi or the MCP server (drop-in mode), record the intended amendment file in the closeout for the next Pi or MCP session to apply.
92
92
  - Append numbered entries to `DECISIONS.md` for any decisions made during triage that weren't already in the deltas (e.g., "rejected Δ7 because the spec already covered the case at §X").
93
93
 
94
94
  ## What to avoid
@@ -1,7 +1,7 @@
1
1
  # Post-pipeline amendment schema v1. Copy to scratch/amendments/<slug>.yaml,
2
- # then run codecarto_amend (MCP) with the slug. Refused while the pipeline is
3
- # incomplete — mid-pipeline resolutions belong in the phase handoff
4
- # (open_question_closures / carry_forward_closures).
2
+ # then run codecarto_amend (MCP) or /codecarto-amend <slug> (Pi). Refused while
3
+ # the pipeline is incomplete — mid-pipeline resolutions belong in the phase
4
+ # handoff (open_question_closures / carry_forward_closures).
5
5
  schema_version: 1
6
6
  # Open-question ids resolved on evidence after the pipeline completed;
7
7
  # removed from every phase in workflow/status.yaml.
@@ -3,7 +3,7 @@
3
3
  <!--
4
4
  Output template for the architecture phase.
5
5
  Fill in each section. Remove placeholder text. Keep the section headers.
6
- Mark every conclusion as: fact / strong inference / open question.
6
+ Mark every conclusion as: fact / strong inference / external-behavior claim / open question.
7
7
  -->
8
8
 
9
9
  ## System Intent
@@ -16,6 +16,12 @@
16
16
  <!-- For each finding: location, defect, evidence, severity, evidence level, action. -->
17
17
  <!-- Sort by severity: critical → high → medium → low. -->
18
18
  <!-- If no findings, write "No defects found in this category." -->
19
+ <!-- Evidence Level: observed fact / strong inference / external-behavior claim / open question.
20
+ Action — pre-porting pipelines: fix before porting / port differently / leave behind / verify at runtime;
21
+ maintenance pipelines: fix now / track / accept / investigate.
22
+ Pairing rule (validated mechanically): open question or external-behavior claim ⇒ verify at
23
+ runtime or port differently (pre-porting) / investigate (maintenance), never fix before porting
24
+ or fix now; and list the finding in ## Open Questions below. -->
19
25
 
20
26
  | # | Location | Defect | Severity | Evidence Level | Action |
21
27
  |---|----------|--------|----------|----------------|--------|
@@ -99,6 +105,21 @@
99
105
 
100
106
  ---
101
107
 
108
+ ## Open Questions
109
+
110
+ <!-- Every finding whose Evidence Level is open question or external-behavior claim gets a row
111
+ here, so the hedge travels with the finding into this document — not only into the handoff.
112
+ Mirror each row into your phase handoff's open_questions (kind: needs-runtime-test unless a
113
+ maintainer decision or spec ruling is what is missing). Derived findings lists the finding
114
+ numbers (e.g. "5.2, 4.1") that depend on this question; none of them may carry a settled
115
+ action while the question stands. -->
116
+
117
+ | ID | Kind | Question | Why source cannot settle it | Derived findings |
118
+ |----|------|----------|-----------------------------|------------------|
119
+ | | | | | |
120
+
121
+ ---
122
+
102
123
  ## Coverage and limits
103
124
 
104
125
  - Inspected scope:
@@ -120,6 +141,8 @@
120
141
  | 4 | Summary tables are complete and counts match the detailed findings. | PASS / PARTIAL / FAIL | |
121
142
  | 5 | Findings are marked with evidence levels. | PASS / PARTIAL / FAIL | |
122
143
  | 6 | Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots. | PASS / PARTIAL / FAIL | |
144
+ | 7 | Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table. | PASS / PARTIAL / FAIL | |
145
+ | 8 | Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate. | PASS / PARTIAL / FAIL | |
123
146
 
124
147
  **Validated by:** [session identifier or date]
125
148
  **Overall:** PASS / PASS WITH GAPS / FAIL
@@ -22,6 +22,11 @@
22
22
  <!-- For each finding: location, defect, evidence, severity, evidence level, action. -->
23
23
  <!-- Sort by severity: critical → high → medium → low. -->
24
24
  <!-- If no findings, write "No defects found in this category." -->
25
+ <!-- Evidence Level: observed fact / strong inference / external-behavior claim / open question.
26
+ Action: fix before porting / port differently / leave behind / verify at runtime.
27
+ Pairing rule (validated mechanically): open question or external-behavior claim ⇒
28
+ verify at runtime or port differently, never fix before porting; and list the finding
29
+ in ## Open Questions below. -->
25
30
 
26
31
  | # | Location | Defect | Severity | Evidence Level | Action |
27
32
  |---|----------|--------|----------|----------------|--------|
@@ -90,6 +95,21 @@
90
95
 
91
96
  ---
92
97
 
98
+ ## Open Questions
99
+
100
+ <!-- Every finding whose Evidence Level is open question or external-behavior claim gets a row
101
+ here, so the hedge travels with the finding into this document — not only into the handoff.
102
+ Mirror each row into your phase handoff's open_questions (kind: needs-runtime-test unless a
103
+ maintainer decision or spec ruling is what is missing). Derived findings lists the finding
104
+ numbers (e.g. "1.3, 6.1") that depend on this question; none of them may carry a settled
105
+ action while the question stands. -->
106
+
107
+ | ID | Kind | Question | Why source cannot settle it | Derived findings |
108
+ |----|------|----------|-----------------------------|------------------|
109
+ | | | | | |
110
+
111
+ ---
112
+
93
113
  ## Coverage and limits
94
114
 
95
115
  - Inspected scope:
@@ -110,6 +130,8 @@
110
130
  | 4 | Summary tables are complete and counts match the detailed findings. | PASS / PARTIAL / FAIL | |
111
131
  | 5 | Findings are marked with evidence levels. | PASS / PARTIAL / FAIL | |
112
132
  | 6 | Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots. | PASS / PARTIAL / FAIL | |
133
+ | 7 | Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table. | PASS / PARTIAL / FAIL | |
134
+ | 8 | Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate. | PASS / PARTIAL / FAIL | |
113
135
 
114
136
  **Validated by:** [session identifier or date]
115
137
  **Overall:** PASS / PASS WITH GAPS / FAIL
@@ -86,7 +86,7 @@
86
86
 
87
87
  | Defect ID | Source Report | One-line Description | Severity | Disposition | Required design consequence |
88
88
  |-----------|---------------|----------------------|----------|-------------|-----------------------------|
89
- | | | | | fix before porting / port differently / leave behind | |
89
+ | | | | | fix before porting / port differently / leave behind / verify at runtime | |
90
90
 
91
91
  ## Observed Facts vs. Inferred Structure
92
92
 
@@ -159,7 +159,7 @@
159
159
  | 1 | The system summary, layer map, contract table, protocol notes, and porting findings are synthesized. | PASS / PARTIAL / FAIL | |
160
160
  | 2 | Portability hazards and open questions are separated from facts. | PASS / PARTIAL / FAIL | |
161
161
  | 3 | Feature importance is sorted for porting. | PASS / PARTIAL / FAIL | |
162
- | 4 | Known defects are referenced in the Defect Synthesis with porting recommendations (fix before porting / port differently / leave behind), or the section explicitly notes that no defect scan ran. | PASS / PARTIAL / FAIL | |
162
+ | 4 | Known defects are referenced in the Defect Synthesis with porting recommendations (fix before porting / port differently / leave behind / verify at runtime), or the section explicitly notes that no defect scan ran. | PASS / PARTIAL / FAIL | |
163
163
  | 5 | Findings are marked with evidence levels. | PASS / PARTIAL / FAIL | |
164
164
  | 6 | Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots. | PASS / PARTIAL / FAIL | |
165
165
  | 7 | The Source Index makes the bundle a self-contained compression boundary and identifies targeted deep-read triggers. | PASS / PARTIAL / FAIL | |
@@ -25,6 +25,12 @@
25
25
  <!-- For each finding: location, defect, evidence, severity, evidence level, action. -->
26
26
  <!-- Cite the protocol or state-machine entry that the finding violates, when relevant. -->
27
27
  <!-- If no findings, write "No defects found in this category." -->
28
+ <!-- Evidence Level: observed fact / strong inference / external-behavior claim / open question.
29
+ Action: fix before porting / port differently / leave behind / verify at runtime.
30
+ Pairing rule (validated mechanically): open question or external-behavior claim ⇒
31
+ verify at runtime or port differently, never fix before porting; and list the finding
32
+ in ## Open Questions below. A finding that closes a routed carry-forward derived from a
33
+ still-open needs-runtime-test question inherits that uncertainty. -->
28
34
 
29
35
  | # | Location | Defect | Severity | Evidence Level | Action |
30
36
  |---|----------|--------|----------|----------------|--------|
@@ -45,6 +51,9 @@
45
51
  ## Pass 5: API Contract Violations
46
52
 
47
53
  <!-- Each finding should pair the source contract/protocol reference with the diverging code location. -->
54
+ <!-- When the analyzed code is the CALLER of a contract another system implements, what that system
55
+ does with the payload is an external-behavior claim (action: verify at runtime) until a runtime
56
+ probe against the pinned version says otherwise — see passes/05 "Which side implements the contract". -->
48
57
 
49
58
  | # | Location | Defect | Severity | Evidence Level | Action | Spec Reference |
50
59
  |---|----------|--------|----------|----------------|--------|----------------|
@@ -95,6 +104,21 @@
95
104
 
96
105
  ---
97
106
 
107
+ ## Open Questions
108
+
109
+ <!-- Every finding whose Evidence Level is open question or external-behavior claim gets a row
110
+ here, so the hedge travels with the finding into this document — not only into the handoff.
111
+ Include any still-open needs-runtime-test question a closed carry-forward derived from.
112
+ Mirror each row into your phase handoff's open_questions. Derived findings lists the finding
113
+ numbers (e.g. "5.2") that depend on this question; none of them may carry a settled action
114
+ while the question stands. -->
115
+
116
+ | ID | Kind | Question | Why source cannot settle it | Derived findings |
117
+ |----|------|----------|-----------------------------|------------------|
118
+ | | | | | |
119
+
120
+ ---
121
+
98
122
  ## Coverage and limits
99
123
 
100
124
  - Inspected scope:
@@ -115,6 +139,8 @@
115
139
  | 4 | Findings are organized by pass and sorted by severity; summary tables match the detailed findings. | PASS / PARTIAL / FAIL | |
116
140
  | 5 | Findings are marked with evidence levels. | PASS / PARTIAL / FAIL | |
117
141
  | 6 | Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots. | PASS / PARTIAL / FAIL | |
142
+ | 7 | Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table. | PASS / PARTIAL / FAIL | |
143
+ | 8 | Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate. | PASS / PARTIAL / FAIL | |
118
144
 
119
145
  **Validated by:** [session identifier or date]
120
146
  **Overall:** PASS / PASS WITH GAPS / FAIL
@@ -12,8 +12,8 @@
12
12
  phase handoff. When a spike's findings change the reimplementation spec, write
13
13
  the deltas as Recommended Deltas below and apply them with the
14
14
  spec-delta-application skill; when a spike resolves an open question after the
15
- pipeline completed, close it with an amendment (templates/amendment.yaml +
16
- codecarto_amend), citing this report.
15
+ pipeline completed, close it with an amendment (templates/amendment.yaml, then
16
+ codecarto_amend on MCP or /codecarto-amend on Pi), citing this report.
17
17
 
18
18
  Keep it honest: a spike that failed to answer its question is a valid result —
19
19
  record what was tried and what blocked it.
@@ -50,7 +50,7 @@ Below is what a real PASS WITH GAPS block looks like — useful when a phase fin
50
50
  | 2 | The layer map and dependency direction are documented. | PASS | §Layer Map; dependency direction in §Layer Map → "Dependency Direction." |
51
51
  | 3 | Public surfaces are identified. | PARTIAL | CLI commands and HTTP routes enumerated (§Public Surfaces). MCP server endpoints and the websocket subscription channel are listed by name only — schemas not extracted. Routed to `carry_forward` as `arch-CF2` with `target_phase: protocols`. |
52
52
  | 4 | Runtime lifecycle, concurrency model, and porting priorities are summarized. | PASS | §Runtime Lifecycle, §Concurrency Model, §Porting Priorities (table). |
53
- | 5 | Findings are marked with evidence levels. | PASS | All inferences marked `observed fact` / `strong inference` / `portability hazard` / `open question`. |
53
+ | 5 | Findings are marked with evidence levels. | PASS | All inferences marked `observed fact` / `strong inference` / `portability hazard` / `external-behavior claim` / `open question`. |
54
54
 
55
55
  **Validated by:** 2026-05-02 (architecture phase, session 1)
56
56
  **Overall:** PASS WITH GAPS
@@ -56,6 +56,8 @@ phases:
56
56
  - Findings are organized by pass and sorted by severity.
57
57
  - Summary tables are complete and counts match the detailed findings.
58
58
  - Findings are marked with evidence levels.
59
+ - Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table.
60
+ - Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate.
59
61
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
60
62
  handoff_requirements:
61
63
  - Run validation per workflow/VALIDATE.md. Append validation block to primary output.
@@ -60,6 +60,8 @@ phases:
60
60
  - Findings are organized by pass and sorted by severity.
61
61
  - Summary tables are complete and counts match the detailed findings.
62
62
  - Findings are marked with evidence levels.
63
+ - Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table.
64
+ - Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate.
63
65
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
64
66
  handoff_requirements:
65
67
  - Run validation per workflow/VALIDATE.md. Append validation block to primary output.
@@ -159,7 +161,7 @@ phases:
159
161
  - The system summary, layer map, contract table, protocol notes, and porting findings are synthesized.
160
162
  - Portability hazards and open questions are separated from facts.
161
163
  - Feature importance is sorted for porting.
162
- - Known defects are referenced with porting recommendations (fix before porting / port differently / leave behind).
164
+ - Known defects are referenced with porting recommendations (fix before porting / port differently / leave behind / verify at runtime).
163
165
  - Findings are marked with evidence levels.
164
166
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
165
167
  - The Source Index makes the bundle a self-contained compression boundary and identifies targeted deep-read triggers.
@@ -62,6 +62,8 @@ phases:
62
62
  - Summary tables are complete and counts match the detailed findings.
63
63
  - Items spotted that are actually semantic in nature are routed onward via a carry_forward entry in the phase handoff targeting defect-scan-semantic.
64
64
  - Findings are marked with evidence levels.
65
+ - Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table.
66
+ - Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate.
65
67
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
66
68
  handoff_requirements:
67
69
  - Run validation per workflow/VALIDATE.md. Append validation block to primary output.
@@ -157,6 +159,8 @@ phases:
157
159
  - Findings are organized by pass and sorted by severity; summary tables match the detailed findings.
158
160
  - Any carry_forward entries that targeted defect-scan-semantic have been resolved or explicitly re-routed.
159
161
  - Findings are marked with evidence levels.
162
+ - Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table.
163
+ - Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate.
160
164
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
161
165
  handoff_requirements:
162
166
  - Run validation per workflow/VALIDATE.md. Append validation block to primary output.
@@ -196,7 +200,7 @@ phases:
196
200
  - The system summary, layer map, contract table, protocol notes, and porting findings are synthesized.
197
201
  - Portability hazards and open questions are separated from facts.
198
202
  - Feature importance is sorted for porting.
199
- - Defect Synthesis consolidates mechanical-defects.md and semantic-defects.md with porting recommendations (fix before porting / port differently / leave behind).
203
+ - Defect Synthesis consolidates mechanical-defects.md and semantic-defects.md with porting recommendations (fix before porting / port differently / leave behind / verify at runtime).
200
204
  - Findings are marked with evidence levels.
201
205
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
202
206
  - The Source Index makes the bundle a self-contained compression boundary and identifies targeted deep-read triggers.
@@ -90,6 +90,8 @@ phases:
90
90
  - Summary tables are complete and counts match the detailed findings.
91
91
  - Items spotted that are actually semantic in nature are routed onward via a carry_forward entry in the phase handoff targeting defect-scan-semantic.
92
92
  - Findings are marked with evidence levels.
93
+ - Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table.
94
+ - Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate.
93
95
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
94
96
  handoff_requirements:
95
97
  - Run validation per workflow/VALIDATE.md. Append validation block to primary output.
@@ -194,6 +196,8 @@ phases:
194
196
  - Findings are organized by pass and sorted by severity; summary tables match the detailed findings.
195
197
  - Any carry_forward entries that targeted defect-scan-semantic have been resolved or explicitly re-routed.
196
198
  - Findings are marked with evidence levels.
199
+ - Unsettled findings (evidence level open question or external-behavior claim) carry an unsettled action (verify at runtime or port differently on pre-porting pipelines; investigate on maintenance pipelines), never a settled one, and each appears in the Open Questions table.
200
+ - Every quantitative specific in a finding (size, count, default, version, timeout) cites the file and line or command output it was read from, or is marked as an estimate.
197
201
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
198
202
  handoff_requirements:
199
203
  - Run validation per workflow/VALIDATE.md. Append validation block to primary output.
@@ -236,7 +240,7 @@ phases:
236
240
  - The system summary, layer map, contract table, protocol notes, and porting findings are synthesized.
237
241
  - Portability hazards and open questions are separated from facts.
238
242
  - Feature importance is sorted for porting.
239
- - Defect Synthesis consolidates mechanical-defects.md and semantic-defects.md with porting recommendations (fix before porting / port differently / leave behind).
243
+ - Defect Synthesis consolidates mechanical-defects.md and semantic-defects.md with porting recommendations (fix before porting / port differently / leave behind / verify at runtime).
240
244
  - Findings are marked with evidence levels.
241
245
  - Coverage and limits name inspected scope, skipped scope, evidence basis, and blind spots.
242
246
  - The Source Index makes the bundle a self-contained compression boundary and identifies targeted deep-read triggers.
@@ -3,4 +3,4 @@
3
3
  # workspace's framework-owned files (GUIDE.md, templates/, workflow/ pipelines
4
4
  # and VALIDATE.md) predate the running release. Written at release time and
5
5
  # copied verbatim by init — never edit by hand.
6
- scaffold_version: 0.17.0
6
+ scaffold_version: 0.18.0