sourcecode 5.8.18__py3-none-any.whl → 5.8.20__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Potentially problematic release.


This version of sourcecode might be problematic. Click here for more details.

sourcecode/__init__.py CHANGED
@@ -4,4 +4,4 @@ ASK Engine is the product. ``ask`` is the canonical CLI command; ``sourcecode``
4
4
  the legacy compatibility alias and the Python/PyPI package name. See
5
5
  docs/PRODUCT_IDENTITY.md (normative)."""
6
6
 
7
- __version__ = "5.8.18"
7
+ __version__ = "5.8.20"
@@ -16,28 +16,36 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
16
16
 
17
17
  ## Current Synchronization
18
18
 
19
- **Latest attached audits — release `5.8.18`, 2026-08-21, two independent rounds, and the
20
- first pass since 5.8.2 that opens with a regression.** Round A (`saint-server`/MSAS,
21
- 3 342 Java files, 17 commands, 29 invocations, same HEAD as the previous round bit for
22
- bit) scores **8.77/10** — up from 8.21 — with **zero false statements about the
23
- repository for the second consecutive run**, 6 of 6 reported defects closed and verified
24
- from outside, and **five new contract defects, four of them the same class as the fix that
25
- preceded them**. Round B (a 12-repository bank, ~150 invocations, the §7 measurement
26
- protocol applied: idle host, interleaved, ×3, median, cache state declared per row) closes
27
- **8 of the 10 new rows and 2 of the 9 inherited ones**, records **zero content regressions
28
- and no performance row above the ×1.25 threshold**, and finds **two regressions, one of
29
- them a P0**: `compare` is dead in every valid invocation with an unhandled `NameError`,
30
- and `project_summary` publishes an Apache licence header as the description of the project
31
- on 3 of 8 repositories.
32
-
33
- **The two rounds agree on the shape of what is left, and it has moved again.** The
34
- previous pass ended with contracts that were *absent*; this one ends with contracts that
35
- *contradict each other across authorities*: two homonymous fields declaring the same unit
36
- and the same population and differing 16×, one budget published three times with three
37
- values, a help page denying a scope the fix beside it granted, and two ledger rows marked
38
- `closed` that do not reproduce as closed in the field. **A row marked closed that is not
39
- closed is worse than an open one — nobody looks at it again** — so those are carried at P1
40
- below rather than left in their original sections.
19
+ **Latest attached audits — audit of release `5.8.18`, 2026-08-21, two independent rounds,
20
+ the first pass in this record that closes with no P0 and no P1 open, and **closed end to
21
+ end in `5.8.19`**.** Round A (MSAS,
22
+ 3 342 Java files, the same HEAD bit for bit for the third consecutive round, 18 commands,
23
+ ~30 invocations) scores **8.95/10** — 8.21 → 8.77 → 8.95 — with **zero false statements
24
+ about the repository for the third consecutive run**, 4 of its 6 previous findings closed
25
+ clean, one closed by half and one open. Round B (a 12-repository bank, ~200 invocations,
26
+ the §7 protocol) scores **8.8/10** — up from 7.7 — closes **8 of its 10 previous rows**,
27
+ and calls the release adoptable without reservation. **Both rounds record zero content
28
+ regressions, zero unhandled exceptions and zero unauthorised writes; the two performance
29
+ regressions Round B initially published were retracted by its own re-measurement.** The
30
+ whole eighth-pass queue is verified closed from outside, including the two regressions the
31
+ shipped ledger still marked `open` — which is itself a row (`AUD-594-X04`) and is corrected
32
+ in place below. **The ninth-pass queue is closed in `5.8.19`: twelve rows, one atomic
33
+ commit each, with every regression assertion written over the catalogue that already
34
+ exists rather than over the reported symbol — the rule three passes of evidence now
35
+ support. There is no open queue in this ledger.**
36
+
37
+ **What is left has moved once more, and it is no longer a defect of design.** The seventh
38
+ pass ended with contracts that were *absent*; the eighth with contracts that *contradicted
39
+ each other across authorities*; this one ends with contracts that are nearly coherent and
40
+ leak at the places a fix did not sweep to: a help header the generator never reached, a
41
+ missing `*_population` key, a `sibling_view` that names no block, a `[:5]` beside a `[:30]`
42
+ that was just declared, and a `selftest` that does not expose the invariants the last three
43
+ releases repaired. Two rows carry more weight than their P2 severity suggests and lead the
44
+ queue for that reason: `AUD-595-A03`, because a budget-exhausted audit exits 0 and a CI gate
45
+ reads it as a pass over a payload that declares itself a floor; and `AUD-594-X01`, because
46
+ one of its three stale literals tells the reader the root analysis **must not be CI-gated**,
47
+ which is the opposite of what the binary now does.
48
+
41
49
 
42
50
  **Previous attached audits — release `5.8.16`, 2026-08-21, two independent rounds.**
43
51
  Round A (`saint-server`, 3 342 Java files, 17 commands, 25 invocations) scores **8.21/10**
@@ -51,9 +59,9 @@ previously open defects by correction with **zero regressions among the inherite
51
59
 
52
60
  Both rounds agree on where the product now fails, and it is not the Java/Spring analysis:
53
61
  it is the contract an agent consumes programmatically — the write guard, repository
54
- identity, list units, flag scope, and published budgets. Its queue is closed in `5.8.17`;
55
- **the eighth-pass section below is the only current queue**, and older rounds remain for
56
- traceability.
62
+ identity, list units, flag scope, and published budgets. Its queue closed in `5.8.17` and the eighth-pass queue in `5.8.18`;
63
+ **the ninth-pass section below is the most recent queue and it is closed**, and older rounds
64
+ remain for traceability.
57
65
 
58
66
  **Mechanisms confirmed in source during the seventh-pass intake** (not taken on the reports' word):
59
67
  `verify --init` writes with no `readonly.guard` (`cli.py:10706-10708`) while five sibling
@@ -76,8 +84,9 @@ itself rather than merely under-declaring.
76
84
  **Release status:** entries marked “pending `5.8.14`” in their historical wording are
77
85
  shipped in `5.8.14`; the correction battery fixes listed above are included in `5.8.15`.
78
86
 
79
- **Current release:** `5.8.18` closes the whole eighth-pass queue, one commit per row, and
80
- includes the targeted release battery. `5.8.17` before it carried the seventh-pass queue;
87
+ **Current release:** `5.8.19` closes the whole ninth-pass queue, one commit per row, with
88
+ the suite green end to end — 11 008 passing, and the pre-existing `B29` failure the
89
+ previous release shipped with is gone. `5.8.18` before it closed the eighth-pass queue. `5.8.17` before it carried the seventh-pass queue;
81
90
  `5.8.16` before that carried two post-audit
82
91
  corrections: compact cold-start identity is preserved for `no_ris`, and manifest-cache keys
83
92
  include Python requirements files. The testing-repository hygiene exercise changed no
@@ -100,9 +109,110 @@ with neither parseable stdout nor its requested artifact) or `BUG-6`
100
109
  queue position. `BUG-1` and `BUG-5` gain witnesses on public OSS repositories, recorded
101
110
  under their own rows rather than as new ones.
102
111
 
112
+ ### Ninth Audit Pass: `5.8.18` findings, closed in `5.8.19` / MSAS + 12-repository bank / 2026-08-21
113
+
114
+ **Closed queue — all twelve rows are `closed 5.8.19`, one atomic commit each.** Kept in full because the mechanisms are the evidence and the direction column is what the fixes were written against.
115
+
116
+ **How the pass read when it was filed.** Two independent rounds re-ran their protocols against the release
117
+ that closed the eighth-pass queue. Round A (MSAS, `banyan-v2` @ `3dde0376`, 3 342 Java
118
+ files — the same subject bit for bit for the third consecutive round — 18 commands, ~30
119
+ invocations) scores **8.95/10**, up from 8.77 and 8.21. Round B (a 12-repository bank,
120
+ ~200 invocations, the §7 measurement protocol) scores **8.8/10**, up from 7.7, and calls
121
+ the release **adoptable without reservation**.
122
+
123
+ **No regressions, of either kind.** Round A compared eleven payloads byte for byte against
124
+ 5.8.17 and found nine identical, with `root --compact`, `root --agent` and `selftest`
125
+ moving by the 1–2 bytes of their own version string; the analysis figures are unchanged
126
+ (`spring-audit` 373/94, `posture` 1 908/5, `migrate-check` 0/532/97, `endpoints` 3 763 with
127
+ no drift where 5.8.17 drifted +21, `no_security_signal` 725 = `undocumented` 725, `repo_id`
128
+ cardinality 1). Round B records **zero content regressions, zero unhandled exceptions in
129
+ ~200 invocations, zero unauthorised writes across 10 of 10 trees, and — after the
130
+ retraction in the next paragraph — zero timing rows above the ×1.25 threshold**. This is
131
+ the first pass in this record with no P0 and no P1 open at its close.
132
+
133
+ **A reported performance regression is retracted, and the retraction is a product finding.**
134
+ Round B published `spring-audit` on `openmrs-core` as ×2.30, then ×1.39, then withdrew both:
135
+ with an explicit `cache warm` and a settled host the figure is **×1.07** (2.48 s against a
136
+ 5.8.17 anchor of 3.00 s at the same protocol). The diagnostic is the product's own
137
+ `timings.unaccounted_pct` — 55 % in the state that *reports* warm but is not, 11.8 % fully
138
+ warm, 1.5 % genuinely cold — and it is the reason `AUD-595-B01` below is filed as a defect
139
+ rather than as an auditor's note. Round A's `impact` 7 s → 101 s is likewise not a
140
+ regression: the version bump emptied the Shared CIR (19.18 MB → 0) and retired 58 056 parse
141
+ entries, and `impact-chain` immediately after, on the now-warm CIR, took 11 s.
142
+
143
+ **Two figures the previous pass carried are corrected here, both downward in severity.**
144
+ `B-3`'s *"`--agent` is +61 % over its published band"* was measured against the guide's
145
+ hand-written band, which `AUD-593-N02` correctly deleted; against the ceiling the code
146
+ actually enforces (`BUDGET_AGENT = 40 000`, `output_budget.py:292`, applied at
147
+ `cli.py:4585-4589`) the measured 35 358 B is **inside budget**. `--agent` is not over
148
+ budget; it is unpublished, which is `AUD-594-X02`. And `AUD-592-A02`'s `endpoints` half was
149
+ not a truncation at all — see that row.
150
+
151
+ | ID | Severity | Current status | Required direction |
152
+ |---|---|---|---|
153
+ | `AUD-594-X01` / `A-1`, residual of `AUD-593-N03` | **P2** | **closed `5.8.19`** (`dd222b9`) — the tuple moved to `analysis_budget`, a module neither `cli` nor `cache_model` owns, both render from it, and the guide now defers to `ask --help` rather than keeping a copy; the `must not be CI-gated` sentence is corrected, not trimmed. Regression is parametric over the tuple: any paragraph of any shipped authority that names the variable and enumerates commands must enumerate all of them. Mechanism as found: open, confirmed in source, and one of the three literals is an inverted operating instruction, not a stale number. The generated Options line enumerates eight commands (`cli.py:589-596` → `:632`); three hand-written literals in two other shipped authorities still name three. `cache_model.py:679` — the header paragraph of the same help page — ends *"ASK_MAX_ANALYSIS_SECONDS bounds only the phase-runner commands (spring-audit, risk and audit-report)"*. `docs/USER_GUIDE.md:1533` repeats it, and **`:1558` goes further: commands outside "that phase-runner set are not bounded by this variable and must not be CI-gated"** — it tells a reader they cannot gate the root analysis in CI, which they now can. Behaviour settles it: with the variable at 90 the root prints *"ask is a repo-wide analysis and the configured budget is 90s"*. Same shape as the closed `AUD-590-B03`: two incompatible statements on one page, half generated and half not. **`D-09`/`BUG-5`'s stale cost anchors live in the same `return`** (`cache_model.py:671-680`, `FIELD_ANCHOR_MEASURED_VERSION`), so one function carries two of the three oldest rows in this ledger. | Replace the three literals with `_budgeted_analysis_commands_phrase()`, moving `_BUDGETED_ANALYSIS_COMMANDS` into a module both `cli` and `cache_model` may import rather than creating a `cache_model → cli` cycle. Delete the `:1558` sentence outright — it is wrong in the direction that costs a user a CI gate. Assertion, parametric over the tuple: no string in any shipped authority (rendered `--help`, `cache_model`, `docs/*.md`) enumerates a **proper subset** of `_BUDGETED_ANALYSIS_COMMANDS` next to the variable name. |
154
+ | `AUD-595-A03` / `B-6` exit-code half | **P2** | **closed `5.8.19`** (`ab7285f`), **and the reported premise is half refuted by measurement.** Measured on dubbo (4 049 files, budget 1 s): the default exits 0, but **`--ci` already exits 75** — `F-AV`'s code for *"the gate did not finish"* — so the product does not read a cut run as a pass wherever it was asked to gate, and 0 without `--ci` is the same contract findings have had all along. What was genuinely missing is the other half of `AUD-591-A09`: the block said what did not run and never what the process would do about it. Now `_partial` carries `gate_exit_code` (75), `exit_code` (what this process exits with) and `gate_exit_code_basis`, from one decision function the exits themselves read, and `--allow-partial` lets a gating pipeline accept a cut answer deliberately. Field-verified in three states: default 0/0, `--ci` 75/75, `--ci --allow-partial` 0/0 with the gate still publishing 75. Mechanism as found: open. `spring-audit` exhausting its budget exits **0** on 3 of 3 runs, over a payload that declares itself a floor: `_partial: {partial: true, why_stopped: "budget_exhausted", phases_completed: ["ir"], phases_pending: ["tx_audit", "security_audit"], how_to_read: "…every count here is a FLOOR over the complete answer…"}`. The block is exemplary; the exit code contradicts it, and a pipeline reading `$?` goes green with two of three phases unexecuted. The overrun half of `B-6` closed in `17c6e06`; this half has never had a row of its own, which is part of why it has survived three passes. Note for the next measurement: `spring-boot` no longer exhausts a 15 s budget (~9 s now), so the repro needs `ASK_MAX_ANALYSIS_SECONDS=3`. | Apply `AUD-591-A09` verbatim — it is already shipped and re-verified in this pass: `verify --no-ci` → process 0 / `exit_code` 0 / `gate_exit_code` 2 / `gate_exit_code_basis`; `verify --ci` → process 2. Emit a non-zero exit whenever `_partial.partial` is true unless `--allow-partial` is passed, with the basis string saying which it was. Site: `cli.py:10650-10653`. |
155
+ | `AUD-595-A02` / `N-5`, advisory half of `AUD-588-B11` | **P2** | **closed `5.8.19`** (`3e15227`), and the six-report survival is explained by what the fix had to be: **the two figures count different populations**, so labelling the advisory with the payload's name would have been a second falsehood. The cheap capped walk now names its own — `all_java_files_excluding_build_output`, a constant with a `Scope.population` property so a consumer need not parse the sentence — and the payload publishes `java_population_count` beside `java_population`, which no surface had ever carried together. Field-verified on halo: advisory `1 349 Java files [population=all_java_files_excluding_build_output]`, payload `production_java_sources 990`. Two populations, both named, both numbered. Mechanism as found: open, and it is one line. One run publishes four figures for one population and labels none of them: the advisory says *"This repository: 8 372 Java files"*, the payload carries `java_population: "production_java_sources"` **with no number**, the help says *"at least 8 400"*, and the auditor's controls give 8 667 (all `*.java`) and 5 135 (`src/main` only) — 8 372 matches neither, so from outside the population cannot be named. The label exists and works on three of four surfaces: built at `cli.py:10301`, emitted as `[population=…]` at `:10348`, `:10400`, `:10412`, attached to the payload at `:10518`. The fourth is `phased_run.py:367`, `lines.append(f" This repository: {scope}.")`. | Emit the same `[population=…]` from `phased_run.py:367`, and publish the **count** beside `java_population` in the payload — today no single surface carries number and label together, which is why six reports have not closed it. Assertion: every Java-file count printed or published carries its population name, and two surfaces of one run never publish two populations without declaring the difference. |
156
+ | `AUD-594-X02` / `A-1` second half | **P2** | **closed `5.8.19`** (`1193937`) — `agent_budget_phrase()` is derived from `BUDGET_AGENT`, the constant `_apply_budget` enforces, and both phrases now share one renderer so two formats for one fact cannot appear. The rendered help publishes `bounded at ~10K tokens (40 KB, chars/4)` on the `--agent` entry and on its option, beside `--compact`'s `~7K`. `BUDGET_COMPACT`'s misleading `# compact/agent main cmd` comment is corrected. The sweep is over the module: every `BUDGET_*` reaching an `_apply_budget` call and bounding a documented view must have a `*_budget_phrase()` that appears in the rendered help; the interior budgets carry a stated exemption. `AUD-590-B03`'s own assertion was re-scoped from *one figure per page* to **one figure per flag**, with the page held to figures the budget module generates — the page-wide form would have failed on the correct fix. Mechanism as found: open. `BUDGET_AGENT = 40 000` exists (`output_budget.py:292`) and is enforced (`cli.py:4585-4589`), and **no authority publishes it**: the packaged guide now defers to `ask --help` (the `AUD-593-N02` fix, correct), and the help's `--agent` entry describes content — *"identity, entry points, dependencies, confidence, gaps"* — and no budget. The only *"~7K tokens"* in the help belongs to `--compact`. `--agent` went from four contradictory figures to none; the deference works for `--compact` because the help publishes, and for `--agent` it points at a blank. Measured: 35 358 B ≈ 8 839 est. tokens against an unpublished 40 000 B ceiling. Adjacent and misleading: `BUDGET_COMPACT`'s comment reads `# compact/agent main cmd`. | Add `agent_budget_phrase()` beside the existing `compact_budget_phrase()` (`output_budget.py:302-320`), derived from `BUDGET_AGENT` and `TOKEN_MODEL` exactly as its sibling is, and consume it in the help's `--agent` entry as `cli.py:583` consumes the compact one. Fix the `BUDGET_COMPACT` comment. Assertion, parametric over the module: for every `BUDGET_*` constant reaching an `_apply_budget` call there is a `*_budget_phrase()` and it appears in the rendered help. |
157
+ | `AUD-594-X03` / `X-03`, residual of `AUD-593-N01` | **P2** | **closed `5.8.19`** (`e89e8f0`) **as a class, not as the reported key.** `direct_caller_count_population` now ships from a single constant shared with its twin, and the sweep written to prove it found **four more catalogued figures shipping a unit and no population** — `chain_classes_total`, `impact.stats.direct_caller_symbol_count`, `explain.incoming_callers_count` and `modernize.in_degree` — none of them reported by either round. All five emit the catalogue's own tags through `declared_tags()`, which refuses an ambiguous tail (`direct_caller_count` is carried by two rows with different populations, which is the whole of `AUD-593-N01`) and leaves an already-published prose `_unit` untouched, so closing a population gap cannot silently change a shipped unit. Field-verified on halo. Mechanism as found: open, confirmed in source. `impact-chain` emits `direct_caller_count`, `direct_caller_count_unit: "distinct caller classes"` and a prose `direct_caller_count_note`, and **no `direct_caller_count_population`** (`spring_impact.py:1640-1654`). The twin beside it carries all three (`expanded_seed_caller_count{,_unit,_population}`), and the catalogue is already right: `caller_metrics.py:126-135` declares the row's population as `POP_EXPANDED_SEED_REFERENCES`. The product's published comparability rule is *"two figures are comparable only when both match"*; a consumer applying it over `_unit` still sees a match between figures that differ 16×. The separation exists only in prose and in a neighbouring key. | Emit `direct_caller_count_population` from the catalogue row itself. Assertion, derived from `FAN_IN_FIGURES` rather than from a hand-written list: for every catalogued key, the emitted payload carries `*_unit` **and** `*_population`, equal to `row["unit"]` and `row["population"]`. That single sweep covers this residual and the class `AUD-593-N05` closed by instance. |
158
+ | `AUD-594-N04` / `D-5`, disclosure half of `B-3` | **P2** | **closed `5.8.19`** (`a0060b7`). `sibling_view.omitted_blocks` is computed by running the other view constructor over the same `SourceMap` — never a hand-written list — with `omitted_blocks_basis` stating that an empty list means the two views agree rather than that the delta was not measured, and `blocks_only_in_this_view` publishing the other direction, because the two views are not a subset relation. `--compact` gains the reciprocal pointer it never had, so a reader who starts there learns the agent view exists. Recursion is prevented by `_emit_sibling=False` on the probe, and the probe is wrapped: a sibling pointer must never be why an answer fails to render. Cost measured: 7,5 µs per assembly, 0,40–0,45 s end to end on halo. Field-verified on halo: 13 omitted, 8 agent-only, published list equal to the measured set. Mechanism as found: open. `serializer.py:2688-2698` carries the comment *"Make the complete sibling view discoverable instead of making consumers infer where omitted structural blocks went"* and then emits `{view, command, reason}` with no block named. Measured both rounds: MSAS 16 blocks in `--compact` and absent from `--agent`, the bank 15 — `mybatis`, `transactional_boundaries`, `stacks`, `spring_profiles`, `deployment`, `env_map`, `project_summary` among them. `sibling_view` is also absent from `--compact`, so the pointer between views is one-way. Substantive rather than cosmetic: `security_surface` is missing from **both** channels on a repository with 368 security findings and 725 routes open by rule (disclosed under `AUD-590-B03`, not closed), and the channel the docs designate for agents is the poorer of the two in security signal. | Publish `omitted_blocks` inside `sibling_view` as `set(compact_keys) - set(agent_keys)`, computed from the two view constructors (`serializer.py:1708` and `:2350`) against the section registry at `:932`, never from a hand-written list, and emit the reciprocal `sibling_view` on the `--compact` side. Assertion: set equality between the computed delta and the published list. |
159
+ | `AUD-595-B01` / §B | **P2** | **closed `5.8.19`** (`b77a1ad`) — **each layer now says which question its word answers**, and the parse store publishes the measurement rather than the word. `cache_layers.basis` marks `cir` as *validated reuse* (its `warm` means this run's key hit) and `parse_store`/`snapshot`/`ris` as *presence before the run, not reuse*. `parse_store_reuse` publishes counted lookups — `hits`, `misses`, `reuse_ratio` — from counters in `parse_cache.get()` that are process-global rather than thread-local, because `--jobs` parses concurrently and the question is about the run. A corrupt entry counts as a miss: it costs a re-parse like any other. The block is **omitted when a run made no lookups** — a ratio over zero lookups is a number invented from nothing. Field-verified on halo: `parse_store: warm` with `990 hits / 0 misses / 1.0`, and on a run whose CIR hit, no `parse_store_reuse` at all. An existing assertion pinning `cache_layers` to exactly four keys was relaxed to a subset, in place, with the reason. Mechanism as found: open, and it has now caused three retracted measurements in three passes. `metadata.cache_layers` reports `warm` in states whose cost differs ×2.7: `spring-audit` on `openmrs-core` measured 4.17 s at `unaccounted_pct` 55 % and 2.48 s at 11.8 %, both with all four layers reporting `warm`, against 1.5 % genuinely cold. The remainder the product publishes is larger than every phase it measures put together, and `B-7`'s own `unit` string enumerates what it should contain — *"serialization, payload assembly, ceiling estimation, interpreter start"* — none of which explains 1.9 s. What is missing is cache **validation** against the tree. The instrumentation is right and nobody reads it. | Either distinguish *present* from *valid for this invocation* in `cache_layers` — a warm that re-validates 8 349 files is not the warm that serves from the snapshot — or instrument the validation as its own phase so it leaves `unaccounted`. Also worth a threshold: an `unaccounted_pct` above a declared bound is a measurement the product should refuse to present as clean. |
160
+ | `AUD-595-A05` | **P3 — new** | **closed `5.8.19`** (`e0b3d10`). `endpoints` publishes `scope` (path, `is_repository_root`, `nested_repositories`, basis), `repo_id` with its `repo_id_basis` — the same cache key `AUD-590-B01` established — and the **noun in `total_unit` now follows the tree that was walked**: *"this repository"* over a checkout, *"the analysed tree"* over a directory whose children are checkouts. The probe is depth-1 and structural (a checkout marks itself; build output and vendored trees are skipped) and branches on no name. Field-verified: halo → `is_repository_root: true`, 158 routes, unit unchanged; the 36-repository workspace → `is_repository_root: false, nested_repositories: 36`. The counts do not move — a scope fix that changed a number would be a different defect, and there is an assertion for that. Mechanism as found: open. `ask endpoints .` over a tree holding 16 repositories and 52 634 `.java` files publishes `total: 3 575` under `total_unit: "handler mappings declared in this repository — the whole route population"`, with `scope: null` and `repo_id: null`. `by_source` decomposes by technology, nothing decomposes by tree, and the unit asserts a repository. **It is not a truncation** — the auditor summed 2 802 endpoints over 10 of the 16 repositories against the 3 575 of the CWD, plausible without a cap and consistent with the uncapped `rglob` this ledger documents — so it is a scope-declaration defect only. Same class as the `scope` that `compare` gained in `5.8.18`: the fix exists in the product and did not reach the most-used core-tier command. | Reuse `data.setdefault("scope", {"path": str(target_dir)})` (`cli.py:8896`, `:9000`) at the `endpoints` call site, and either publish `repo_id` with its `repo_id_basis` (`AUD-590-B01`) or say *"in the analysed tree"* where the tree is not one repository. Sweep, not instance: every census that describes itself as *"this repository"* declares the scope it counted. |
161
+ | `AUD-595-A07`, narrow half of `AUD-592-A02` | **P3** | **closed `5.8.19`** (`cc15177`) — the same shape as `AUD-595-A02` and closed the same way: **two walks, two units, each naming which it is.** `scan_truncated` gains `cap_unit: "java files admitted to the analysis"` and a `cap_note` saying the help's directory-entry cap bounds the cheap sizing scan and is neither comparable nor the same limit; the sizing sentence now reads *"this sizing scan stopped at 25 000 directory entries … the analysis walk has its own cap, in files, published as `scan_truncated`"*. No measured figure moves, and there is an assertion for that. `ASK-13` still holds: a cap that did not stop the walk is not named. Mechanism as found: open. Two scan limits, two units, no cross-reference: the payload publishes `scan_truncated: {truncated: true, files_scanned: 8000, cap: 8000}` and the root `--help` describes *"the scan stopped at 25 000 directory entries"*. A reader of the help cannot relate `files_scanned: 8000` to 25 000 directory entries. | Declare both with their units and their relation, or name only the one that cuts first. |
162
+ | `E-39`, class residual of `AUD-593-N06` | **P3** | **closed `5.8.19`** (`25fc11f`), **and it closes criterion 3 of `B-1` with it — the fixture that was missing for four reports now exists.** Both lists route through `_cap_effect("display_list", …)` above five, and both keys are **published empty rather than omitted**: an absent key was indistinguishable from a field that was never computed, which is exactly why `mall`'s `{mapper_interfaces: 76, xml_files: 76}` could not be read from outside. Field-verified on `mall`: same 76/76, now with `orphan_xml: []` and `missing_xml: []` beside them. The `B-1` fixture is a generated repository with an XML mapper whose `namespace` names an interface that does not exist; it is asserted to appear in `orphan_xml` and **not** to be counted in `mapper_interfaces`. Two older assertions that required the keys to be *absent* when empty were updated to the new contract, with the reason recorded in place. Mechanism as found: open, found in source during this pass's triage, not reported by either round. `serializer.py:480-489` emits `orphan_xml = orphan_xml[:5]` and `missing_xml = missing_xml[:5]` with no `*_cap` block, in the same payload where the caps closed by `AUD-593-N06` now declare `total`/`shown`/`omitted`. Both keys are also emitted **only when non-empty**, so their absence is indistinguishable from *"not computed"* — which is exactly how this pass read `mall`'s `{mapper_interfaces: 76, xml_files: 76}` before reading the source. It is also why criterion 3 of `B-1` stays unverifiable from the field. | Route both through `_cap_effect(...)`, and publish the empty case explicitly rather than omitting the key — a declared zero is the contract this product enforces everywhere else. Fixture for the `B-1` criterion, still absent after four reports: a `Foo.xml` whose `namespace` names a mapper interface that does not exist, in a module under `@MapperScan`; it must appear in `orphan_xml` and must not be counted in `mapper_interfaces`. |
163
+ | `AUD-594-X04` | **P3** | **closed `5.8.19`** (`05b8cc9`) — the rows were corrected in place by this pass, and the durable half now ships as a release assertion: a row filed under a section headed *"closed in `X`"* must carry a closed verdict, **or the section's own preamble must name it as carrying forward**, so a half-closed row stays possible and a stale one does not. The preamble is read only from the lines *above* the table — prose after it names half the ledger, and counting that would make the assertion pass on anything, which was verified by mutation: reopening one closed row fails the battery. Its precondition `AUD-592-D04` is closed in the same release. Mechanism as found: open, and its precondition is `AUD-592-D04`. The packaged `DEFECT-LEDGER.md` in the 5.8.18 wheel marks `AUD-593-N01/N02/N05/N06` and both regressions `AUD-592-R01/R02` as `Current status: open`, and cites lines the shipped code no longer matches (`repository_ir.py:9248` → `:9268`). All six are fixed in that same binary, verified by behaviour in both rounds. Mitigating and real: the section header does say *"5.8.17 findings, closed in 5.8.18"*, so it is a header↔row incoherence rather than a flat falsehood, and `ask version` declares *"Rows closed after it shipped are not in this copy"* — but these closed **before** it shipped, so the disclaimer does not cover them. An auditor reads the row. | Derive each row's `Current status` from the same state as its section header, or drop the per-row field and keep the section's. This pass's own hygiene edit — closing the eighth-pass rows in place, above — is the manual form of that fix; the durable form is a release assertion that no row filed under a *"closed in `X`"* header ships with status `open`. |
164
+ | `AUD-595-Q01` | **P3 — process** | **closed `5.8.19`** (`91a6566`) — 7 rows → 10. `AUD-590-B01` (every command publishing `repo_id` publishes the same one, and it equals the cache key), `AUD-590-B02` (figures sharing a `(unit, population)` are equal **on this repository's own highest-fan-in symbol**, chosen from its endpoint census rather than from a fixture — which is the exact failure mode that closed it falsely twice) and `AUD-594-X03` (every catalogued figure emits both tags). Field-verified on halo: 3 of 3 pass, the fan-in row resolving to `ThemeEndpoint` and the tag row checking 5 catalogued figures. All three refuse with `not_applicable` and a reason where the repository offers nothing to query — a green run over nothing is how a defect gets closed twice, and there is an assertion for that over every row, not only the new ones. Mechanism as found: open. `selftest` still runs the same 7 rows (`selftest.py:415-442`: `ASK-09`, `R2`, `ASK-16`, `B6`, `E-37`, `B20`, `AS-16`), and **not one covers an invariant repaired in the last three releases**: `repo_id` cardinality 1 and its equality with the cache-dir key (`AUD-590-B01`), cross-command agreement between fan-in figures sharing `(unit, population)` (`AUD-590-B02`, `AUD-593-N01`), or the rule that every catalogued key emits `*_unit` and `*_population` (`AUD-594-X03`, `AUD-593-N05`). The command's own basis says *"Each row is the acceptance criterion its ledger entry publishes, run against this repository"*. What is not a `selftest` row is not verifiable in the field, and with `build_commit: unrecorded` it is not attributable to a commit either. Those two together are the mechanical reason this ledger's oldest classes keep returning. | Add one row per invariant above. They are the cheapest rows to add — the assertions already exist as internal regressions; what is missing is the field-facing surface. |
165
+
166
+ **Mapped, not new.** `AUD-592-D04` (`build_commit`) is re-witnessed for the fifth time and
167
+ its row above now carries the corrected mechanism: the reader shipped, the stamp never did.
168
+ `D-09` (`BUG-5`) cost anchors are re-counted at **17 on 5.1.0, one each on 4.18.0 and
169
+ 5.0.0** in a 5.8.18 build, and they are now materially misleading rather than merely stale —
170
+ `./mall --agent` measures ×0.09 of its anchor and `spring-audit ./spring-boot` under budget
171
+ ×0.43, so a reader sizing CI from `cache model` overestimates by a large factor; the row
172
+ belongs to `BUG-5` and its literal lives in the `AUD-594-X01` return. `AUD-593-N05` and
173
+ `AUD-593-N06` could not be re-measured on the bank (`via_interface_resolution` is not
174
+ emitted for any of the three symbols tried, including a real interface) and are closed on
175
+ Round A's MSAS evidence alone.
176
+
177
+ **Verified and not defects — do not spend effort.** `interface_mediated_callers_cap: null`
178
+ over a 23-element list (the cap emits above 30). `compare` without `scan_truncated` on a
179
+ 3 342-file subject (below the 8 000 cap). `indirect_callers: []` in `impact-chain` (hub
180
+ guard, declared, with `depth_capped_by_guard` and `closure_complete: false`). `pr-impact` on
181
+ a clean tree (typed `INVALID_INPUT`, exit 1). `data-exposure` answering `answered: false`
182
+ without `dataLabels`. `endpoints` publishing `scan_truncated: null` over the 16-repository
183
+ tree (verified by summation, not assumed). `compare` returning no verdict — *"ranks by
184
+ measurable architectural cost and STOPS"*, and the payload holds to it: `cost_dimensions`,
185
+ `cost_units`, a `why` per candidate, no recommendation. `compare` without `--path` costing
186
+ 89.71 s against 62.44 s in 5.8.16 is not a hidden regression: the earlier run returned
187
+ cross-repository candidates without declaring them and the 5.8.17 run crashed at 99.89 s,
188
+ where this one completes and publishes `scope` and `scan_truncated`; the canonical form with
189
+ `--path` costs 0.72 s.
190
+
191
+ **Process findings of this pass.**
192
+
193
+ 1. **The sweep rule is holding where it was applied, and only there.** Four classes closed
194
+ as parametric sweeps have not recurred across three versions — write guards 8/8, numeric
195
+ bounds 7/7, error-path cost 8/8 at ≤0.79 s, selection flags 3/3 × 4 commands — and scan
196
+ truncation moved 1/4 → 4/4 by being read at each call site. Every row opened in this pass
197
+ is the residue of a *correct* fix that stopped at the instance: a help header the
198
+ generator did not reach, a `*_population` key, a `sibling_view` that names no block, a
199
+ `[:5]` beside a `[:30]` that was just fixed.
200
+ 2. **Deleting a figure can open a hole.** `AUD-593-N02` was closed the right way — four
201
+ hand-written numbers replaced by a deference to the generated authority — and that is
202
+ exactly how `AUD-594-X02` was born, because the authority publishes nothing for
203
+ `--agent`. When a hand-written figure is deleted, assert that the authority it defers to
204
+ actually publishes one.
205
+ 3. **A closure is only as good as the surface that can see it.** Six of this pass's rows
206
+ are invisible to `selftest` (`AUD-595-Q01`) and none of the ~200 measurements is
207
+ attributable to a commit (`AUD-592-D04`). Those two rows are cheap and they gate the
208
+ falsifiability of everything else in this ledger.
209
+
103
210
  ### Eighth Audit Pass: `5.8.17` findings, closed in `5.8.18` / MSAS + 12-repository bank / 2026-08-21
104
211
 
105
- The current queue. Ordered as the correction order: **regressions first, then the rows a
212
+ **Closed queue — every row below is `closed 5.8.18` and re-verified from outside by the
213
+ ninth pass, except `AUD-592-D02` (half) and `AUD-592-D04`, which carry forward.** Kept in
214
+ full because the mechanisms are the evidence, and because the direction column is what the
215
+ fixes were written against. Ordered as the correction order was: **regressions first, then the rows a
106
216
  previous pass marked closed that the field does not reproduce as closed, then new contract
107
217
  defects, then residuals.** Canonical IDs are `AUD-592-*` (bank round, reporter IDs `A-*`,
108
218
  `B-*`, `N-*`, `D-*`) and `AUD-593-*` (MSAS round, reporter IDs `N-0*`). Rows held elsewhere
@@ -110,19 +220,19 @@ in this ledger are listed after the table as mappings, not as new defects.
110
220
 
111
221
  | ID | Severity | Current status | Required direction |
112
222
  |---|---|---|---|
113
- | `AUD-592-R01` / `A-1` | **P0 — regression** | **open, confirmed in source, and it is two defects from one error.** `compare` raises `NameError: name '_compare_scan_trunc' is not defined` on **5 of 5 valid invocations** — Rich traceback on stderr, zero bytes on stdout, no error envelope, exit 1 — ending a 240-invocation streak with no traceback. The assignment introduced by `AUD-591-A06` landed **inside `plan_cmd`** (`cli.py:8874`, beside `plan`'s own `find_java_files` at `:8873`) while the read is in `compare_cmd` (`cli.py:8972`); the two functions are `8818-8901` and `8903-9002`. Verified by reading the source at HEAD, not from the report. Second effect: `plan` computes its truncation state and **discards it** — nothing in the payload carries it, so `plan` answers `resolution: not_found` over a walk that stopped at its cap, offering candidates from a different repository. Third effect: `compare`'s other `AUD-591-A06` fix, `data.setdefault("scope", ...)` at `cli.py:8975`, is **after the line that raises** and has never executed, so the cross-repository candidate leak is still undeclared. | Read `_last_scan_truncation()` at `compare`'s own call site, immediately after `cli.py:8965`; rename the `plan_cmd` variable to `_plan_scan_trunc` and **bind it to the payload** — an unused walk-state variable is the defect, not the fix. Regression: an AST assertion that no `_*_scan_trunc` name is read in a `FunctionDef` that does not assign it, plus one execution test per command that owns a walk. A unit test that imported `compare_cmd` would not have caught this; one that *calls* it would. |
114
- | `AUD-592-R02` / `N-4`, reopens `AUD-513-N04` | **P1 — regression** | **open, confirmed in source, on 3 of 8 repositories (37,5 %), in `--compact` and `--agent` alike.** `project_summary` publishes an Apache-2.0 licence header as the description of the project (struts), a 30-character horizontal rule (halo) and a setext underline (shopizer) — and `summary_basis` asserts `"README descriptive prose"` over all three. In 5.8.16 this field was `null` with `"no descriptive section found in README"`, the literal acceptance text of `AUD-513-N04`; the field went from honest to false. Mechanism, four separate holes, all read at HEAD: HTML comments are detected **only on their opening line** (`summarizer.py:335-337`) with no `in_html_comment` state — three lines above, `in_code_block` does exactly that for fences (`:326-330`) — so `<!---` on line 1 flushes and lines 2-15 of the licence become the first paragraph; headings are ATX-only (`:331`), so setext titles underlined with `===` are content; the short-fragment floor is `len(paragraph) < 30` **strictly** (`:359`), so a rule of exactly 30 dashes survives; `_LICENSE_MARKETING_RE` (`:266-281`) matches product-tier and marketing phrasing and **carries no Apache/MIT/GPL/BSD boilerplate pattern at all**; and `summary_basis` is assigned unconditionally to whatever paragraph survives the filters (`:382-385`) — a restatement of the code path, not a property that was checked. | Add `in_html_comment` symmetric to `in_code_block`; recognise setext underlines as headings, not content; extend the licence regex with the four standard boilerplates; and make `summary_basis` describe the path actually taken — where nothing is verifiably descriptive, `null` plus the 5.8.14 text. **The mechanical cause of the reopening is that no README fixtures exist**: add the three witnesses from the bank as fixtures (HTML licence comment, 30-character rule, setext `===`) in the same commit as the fix. |
115
- | `AUD-593-N01` / `N-01` | **P1** | **open, and it falsifies the closure assertion of `AUD-590-B02`.** For one symbol, one tree, one HEAD, consecutive invocations, both at `depth=4`: `impact` publishes `direct_caller_count: 34` (210 reference sites) and `impact-chain` publishes `554` (3 103), both under the literal unit `"distinct caller classes"`. The product's own comparability rule (`CALLER_METRIC_RECONCILIATION`) is *"two figures are comparable only when both match"* — they match, and differ 16×. **The authority is what is wrong, not only the number**: `FAN_IN_FIGURES` (`caller_metrics.py:80-110`) declares both rows `unit=distinct_classes, population=references_excluding_imports`, and the populations are not the same. Measured mechanism: `impact-chain` expands the seed set with every member of the interfaces the target implements (CH-001b, `spring_impact.py:1110-1135`) and then takes depth 1 **from the expanded seeds** (`caller_reach.py:309-336`), so its "direct" includes callers that name a sibling implementation; `impact` admits only references that name the target and keeps interface-mediated callers on a separate axis (`repository_ir.py:9263`, `:9596-9606`). Corroborating symptom: `chain_classes_total`, declared *"unbounded above by any direct-caller count"*, equals `direct_caller_count` exactly (554), and `indirect_callers` is `[]` — that emptiness is the hub guard capping the walk to depth 1 (`caller_reach.py:319-321`) and **is** declared, so it is not a confident zero. The published closure assertion `len(impact.direct_callers) == chain.metadata.direct_caller_count` measures 30 (list, capped from 34) against 554: it passes on a fixture where the two coincide by accident. | Decide which figure the `direct_caller_*` namespace owns, and fix the **catalogue** either way: either `impact-chain` measures depth-1 fan-in on the unexpanded seed (≈34), or the figure is renamed out of that namespace and its `population` becomes its own tag — an interface-expanded seed set is not `references_excluding_imports`. Regression is parametric over the catalogue, not over `CommonService`: for every pair of `FAN_IN_FIGURES` rows sharing `(unit, population)`, the emitted values must be equal, run over ≥3 symbols of different fan-in **including one that trips the hub guard** — the condition under which the divergence appears, and the reason two previous corrections of this class (`callers_total`, `P0-2`) closed a name and left the divergence. `chain_classes_total ≥ direct_caller_count` with strict inequality wherever uncapped transitive reach exists. |
116
- | `AUD-592-B01` / `B-8` | **P1 — closed row that does not reproduce as closed** | **open; the ledger row overstates its own fix.** `B-8` is recorded as *"compact posture retains one unresolved sample when the total is non-zero"* (`52593b6`). The code retains a sample only when the total is **exactly one**: `_posture_limit = 1 if summary["unresolved"] == 1 else 0` (`cli.py:12653-12654`). Field measurement on `mall`: `summary.unresolved: 3`, `unresolved: []`, `unresolved_cap: {total: 3, shown: 0, omitted: 3, limit: 0}` — and it is the only cap in that payload with `limit: 0` while `undecided_cap` and `evidence_cap` carry 200. `--limit 10` returns all three, so the analysis is intact and only the compact contract is wrong. | Either make the code do what the row says (retain ≥1 sample whenever `summary.unresolved > 0`) or correct the row to the narrower promise and reopen the defect it leaves. Do not leave `closed` standing over `shown: 0`. Add an assertion that no `cap_effect` receives `limit=0` by default on a path where `--limit 0` is published as *"no cap"* (`AUD-513-N08`), so the published semantics and the internal default cannot invert each other. |
117
- | `AUD-592-A02` / `A-1` second half, `§12.5` | **P2** | **open — scan truncation is declared by 1 of 4 commands that own a walk.** Over a truncated scan (cap observed at **8 000 files**), `impact` publishes `scan_truncated`; `plan` and `impact-chain` answer `not_found` with `scan_truncated: null`, and `endpoints` exits 1 with no payload — every one of them an exclusion derived from an incomplete population. `compare` cannot be measured (`AUD-592-R01`). Second, narrower defect in the same row: the root `--help` describes the cap as *"the scan stopped at 25 000 directory entries"* — a different limit in a different unit from the 8 000-file cap the field observes. | Publish the truncation beside every walk that can end in `not_found`, `candidates` or a census — the seventh pass already recorded that `extract_java_endpoints` does its own uncapped `rglob`, so this must be read at each call site, never centrally. Reconcile the two caps in the help text or declare both with their units. |
118
- | `AUD-593-N02` / `N-02` | **P2** | **open — the `--compact` budget is published three times, in two shipped authorities, with three values.** `ask --help` generates *"~7K tokens (30 KB, chars/4)"* from `output_budget`, the module that enforces the limit — this half is the `AUD-590-B03` fix and it is correct. The packaged `docs/USER_GUIDE.md` is still hand-maintained and disagrees with it twice: `:111` *"A ~10K-token subset"* and `:243` / `:1366` *"~2,500–4,000 tokens"*; `--agent` carries a fourth figure at `:254` / `:1367` (*4,500–5,500*). Measured payload: 22 390 B ≈ 5 597 est. tokens — inside the help ceiling, **+40 % over the guide's table**. `grep '7K tokens' *.py` finds no literal: the help is generated and the guide is not, and that asymmetry is the defect. | Generate every token figure in the packaged docs from `output_budget.token_estimate`, or delete them and point at `--help`. Extend the existing release assertion (today help ↔ payload ↔ section registry) to cover `docs/*.md`: a token band in a shipped `.md` that does not come from that module fails the battery. Packaged documentation is a published authority and belongs inside the same assertion as the help text. |
119
- | `AUD-593-N03` / `N-03` | **P2** | **open — `--help` denies the scope that `AUD-590-B04` had just granted.** With `ASK_MAX_ANALYSIS_SECONDS=90` the root analysis prints *"ask is a repo-wide analysis and the configured budget is 90s. Budget read from ASK_MAX_ANALYSIS_SECONDS=90 (process environment)"* — the fix works. The help page beside it still reads *"deadline only for spring-audit, risk, audit-report"* and *"Other commands are not time-bounded by this variable"* (`cli.py:621-622`). The root analysis is the most expensive command in the product and the one an agent most needs to bound; a reader of the help will not set the variable there. | Generate that line from the same registry `_build_analysis_classes` populates (the root is keyed there as `B14`), never by hand. Assertion: the set of commands the help names equals the set the preamble emitter covers. |
120
- | `AUD-593-N05` / `N-05` | **P3** | **open — a figure emitted outside the catalogue that forbids it.** `via_interface_resolution[].caller_count` publishes 37 where every other surface of the same fact says 23 (`interface_mediated_caller_count`, the list length, and the `explanation` prose), and it carries no `_unit` while the rest of the payload does. Both are legitimate and neither is declared: `caller_count` is `len(_iface_callers)` — raw caller symbols of the interface, before the `not in all_affected` filter and before normalisation to classes (`repository_ir.py:9248-9251`); the 23 is `caller_classes(_iface_mediated_callers, ...)` (`:9603`, `:10173`). It is also **not in `FAN_IN_FIGURES`**, which states that a figure absent from it may not be emitted — so `tests/test_fan_in_authority.py` did not catch a fan-in figure added in this release. | Give it a unit and a `FAN_IN_FIGURES` row, or rename it to `interface_caller_symbols`. Make the authority test sweep the **emitted** keys against the catalogue rather than the catalogue against itself; a catalogue that only validates its own rows cannot see an unregistered emission. |
121
- | `AUD-593-N06` / `N-06` | **P3** | **open, latent — a silent `[:30]` where every sibling list declares its cap.** `out["interface_mediated_callers"] = _iface_mediated_classes[:30]` (`repository_ir.py:10172`) with no `*_cap` block, in a payload where `direct_callers_cap`, `indirect_callers_cap` and `security_surface_affected_omitted` all carry `total`/`shown`/`omitted`/`direction`/`how_to_read`. On MSAS the list is 23, so nothing is visible; the first repository with more than 30 interface-mediated callers loses the excess without a word. | Route it through the same `_cap_effect(...)` used 27 lines above (`repository_ir.py:10145`). |
122
- | `AUD-592-D01` / `D-1` | **P3 — residual of `AUD-591-A08`** | **open, confirmed in source.** The write inventory in the root help says `migrate-check` writes *"only with `--history-dir`"* (`cli.py:296`). `--history-dir` only relocates the destination; the flag that writes is `--snapshot`, as `migrate-check --help` states correctly and as the code comment at `cli.py:12980-12986` states explicitly. Verified in the field: `migrate-check --history-dir <path>` exits 0 with zero writes and no guard message, while `--snapshot` fires the guard correctly. `AUD-591-A08`'s assertion checks that each cited flag **exists** in that command's parser — `--history-dir` does — not that it is the flag that triggers the write. | Key the inventory on `(command, triggering flag, path)` taken from the guard site itself, as `AUD-591-A08`'s own acceptance text required. An auditor tests what the inventory lists: this is the same mechanism by which `AUD-591-A01` went unfound. |
123
- | `AUD-592-D02` / `D-2` | **P3 — residual of `AUD-591-A02`** | **open.** `--min-band` filters and validates (`cli.py:11294-11308` in `risk`, `:11563-11568` in `enrich`) but publishes no `_filter` block, while its sibling `--band` publishes one through `_apply_selection`. A consumer cannot tell a filtered answer from a small one. | Route `--min-band` through `_apply_selection` like every other selection flag, or publish the same `_filter` shape from wherever it is applied. |
124
- | `AUD-592-D03` / `D-3` | **P3 — residual of `AUD-591-A02`** | **open.** `enrich --rule NOPE-999` exits 0 with an empty result; `spring-audit --rule NOPE-999` exits 1 against the rule catalogue. One flag name, two contracts. Measured with a synthetic SARIF 2.1.0 (2 results): `--rule`, `--band` and an invalid `--band` all behave correctly in `enrich`; only unknown-rule validation is missing. | Validate against the same catalogue in both commands. `enrich` is experimental-tier, which sets the severity, not the contract. |
125
- | `AUD-592-D04` / `build_commit` | **P3 — process** | **open, fourth consecutive report.** `ask version` publishes `build_commit: unrecorded`, so no measurement in these ~150 invocations is attributable to a commit and *"fixed in 5.8.17"* is not falsifiable from outside. | Stamp the commit at build time. It is the cheapest row in this table and it is the precondition for every other row's verification. |
223
+ | `AUD-592-R01` / `A-1` | **P0 — regression** | **closed `5.8.18`** (`216025f`) — re-verified from outside on the ninth pass: **3 of 3 valid invocations exit 0**, empty stderr, no traceback, in all three invocation forms (simple name, `-p`, FQN); `scope` is published, and `_compare_scan_trunc` now lives inside `compare_cmd` while `plan_cmd` owns `_plan_scan_trunc` bound to its payload. Mechanism as found, kept for traceability: open, confirmed in source, and it is two defects from one error. `compare` raises `NameError: name '_compare_scan_trunc' is not defined` on **5 of 5 valid invocations** — Rich traceback on stderr, zero bytes on stdout, no error envelope, exit 1 — ending a 240-invocation streak with no traceback. The assignment introduced by `AUD-591-A06` landed **inside `plan_cmd`** (`cli.py:8874`, beside `plan`'s own `find_java_files` at `:8873`) while the read is in `compare_cmd` (`cli.py:8972`); the two functions are `8818-8901` and `8903-9002`. Verified by reading the source at HEAD, not from the report. Second effect: `plan` computes its truncation state and **discards it** — nothing in the payload carries it, so `plan` answers `resolution: not_found` over a walk that stopped at its cap, offering candidates from a different repository. Third effect: `compare`'s other `AUD-591-A06` fix, `data.setdefault("scope", ...)` at `cli.py:8975`, is **after the line that raises** and has never executed, so the cross-repository candidate leak is still undeclared. | Read `_last_scan_truncation()` at `compare`'s own call site, immediately after `cli.py:8965`; rename the `plan_cmd` variable to `_plan_scan_trunc` and **bind it to the payload** — an unused walk-state variable is the defect, not the fix. Regression: an AST assertion that no `_*_scan_trunc` name is read in a `FunctionDef` that does not assign it, plus one execution test per command that owns a walk. A unit test that imported `compare_cmd` would not have caught this; one that *calls* it would. |
224
+ | `AUD-592-R02` / `N-4`, reopens `AUD-513-N04` | **P1 — regression** | **closed `5.8.18`** (`ec3050d`) — re-verified from outside on the ninth pass by driving the summariser directly: **8 of 8 repositories and 4 of 4 witnesses correct** (HTML licence comment, 30-dash rule, setext `===`, plus a bare Apache-2.0 boilerplate added by the auditor), and `summary_basis` moved from asserting prose to naming the section and the five filters it passed. Mechanism as found, kept for traceability: open, confirmed in source, on 3 of 8 repositories (37,5 %), in `--compact` and `--agent` alike. `project_summary` publishes an Apache-2.0 licence header as the description of the project (struts), a 30-character horizontal rule (halo) and a setext underline (shopizer) — and `summary_basis` asserts `"README descriptive prose"` over all three. In 5.8.16 this field was `null` with `"no descriptive section found in README"`, the literal acceptance text of `AUD-513-N04`; the field went from honest to false. Mechanism, four separate holes, all read at HEAD: HTML comments are detected **only on their opening line** (`summarizer.py:335-337`) with no `in_html_comment` state — three lines above, `in_code_block` does exactly that for fences (`:326-330`) — so `<!---` on line 1 flushes and lines 2-15 of the licence become the first paragraph; headings are ATX-only (`:331`), so setext titles underlined with `===` are content; the short-fragment floor is `len(paragraph) < 30` **strictly** (`:359`), so a rule of exactly 30 dashes survives; `_LICENSE_MARKETING_RE` (`:266-281`) matches product-tier and marketing phrasing and **carries no Apache/MIT/GPL/BSD boilerplate pattern at all**; and `summary_basis` is assigned unconditionally to whatever paragraph survives the filters (`:382-385`) — a restatement of the code path, not a property that was checked. | Add `in_html_comment` symmetric to `in_code_block`; recognise setext underlines as headings, not content; extend the licence regex with the four standard boilerplates; and make `summary_basis` describe the path actually taken — where nothing is verifiably descriptive, `null` plus the 5.8.14 text. **The mechanical cause of the reopening is that no README fixtures exist**: add the three witnesses from the bank as fixtures (HTML licence comment, 30-character rule, setext `===`) in the same commit as the fix. |
225
+ | `AUD-593-N01` / `N-01` | **P1** | **closed `5.8.18`** (`dd6a706`) **by declaration, not by making the numbers agree** — the right answer to a contract defect, and the ninth pass records it as the model: the figures still read 34 and 554 and did not have to change; `impact-chain`'s row gained its own population (`POP_EXPANDED_SEED_REFERENCES`), the payload carries `direct_caller_count_note` naming CH-001b and stating the two are not peers, and an honestly-named twin `expanded_seed_caller_count` ships with `_unit` and `_population`. **Residual for the programmatic consumer: `AUD-594-X03`** — the emitter publishes no `direct_caller_count_population`. Mechanism as found, kept for traceability: open, and it falsifies the closure assertion of `AUD-590-B02`. For one symbol, one tree, one HEAD, consecutive invocations, both at `depth=4`: `impact` publishes `direct_caller_count: 34` (210 reference sites) and `impact-chain` publishes `554` (3 103), both under the literal unit `"distinct caller classes"`. The product's own comparability rule (`CALLER_METRIC_RECONCILIATION`) is *"two figures are comparable only when both match"* — they match, and differ 16×. **The authority is what is wrong, not only the number**: `FAN_IN_FIGURES` (`caller_metrics.py:80-110`) declares both rows `unit=distinct_classes, population=references_excluding_imports`, and the populations are not the same. Measured mechanism: `impact-chain` expands the seed set with every member of the interfaces the target implements (CH-001b, `spring_impact.py:1110-1135`) and then takes depth 1 **from the expanded seeds** (`caller_reach.py:309-336`), so its "direct" includes callers that name a sibling implementation; `impact` admits only references that name the target and keeps interface-mediated callers on a separate axis (`repository_ir.py:9263`, `:9596-9606`). Corroborating symptom: `chain_classes_total`, declared *"unbounded above by any direct-caller count"*, equals `direct_caller_count` exactly (554), and `indirect_callers` is `[]` — that emptiness is the hub guard capping the walk to depth 1 (`caller_reach.py:319-321`) and **is** declared, so it is not a confident zero. The published closure assertion `len(impact.direct_callers) == chain.metadata.direct_caller_count` measures 30 (list, capped from 34) against 554: it passes on a fixture where the two coincide by accident. | Decide which figure the `direct_caller_*` namespace owns, and fix the **catalogue** either way: either `impact-chain` measures depth-1 fan-in on the unexpanded seed (≈34), or the figure is renamed out of that namespace and its `population` becomes its own tag — an interface-expanded seed set is not `references_excluding_imports`. Regression is parametric over the catalogue, not over `CommonService`: for every pair of `FAN_IN_FIGURES` rows sharing `(unit, population)`, the emitted values must be equal, run over ≥3 symbols of different fan-in **including one that trips the hub guard** — the condition under which the divergence appears, and the reason two previous corrections of this class (`callers_total`, `P0-2`) closed a name and left the divergence. `chain_classes_total ≥ direct_caller_count` with strict inequality wherever uncapped transitive reach exists. |
226
+ | `AUD-592-B01` / `B-8` | **P1 — closed row that does not reproduce as closed** | **closed `5.8.18`** (`8329614`) — re-verified from outside: `total: 3, shown: 1, omitted: 2`, and `--limit 10` still returns all three. Mechanism as found, kept for traceability: open; the ledger row overstates its own fix. `B-8` is recorded as *"compact posture retains one unresolved sample when the total is non-zero"* (`52593b6`). The code retains a sample only when the total is **exactly one**: `_posture_limit = 1 if summary["unresolved"] == 1 else 0` (`cli.py:12653-12654`). Field measurement on `mall`: `summary.unresolved: 3`, `unresolved: []`, `unresolved_cap: {total: 3, shown: 0, omitted: 3, limit: 0}` — and it is the only cap in that payload with `limit: 0` while `undecided_cap` and `evidence_cap` carry 200. `--limit 10` returns all three, so the analysis is intact and only the compact contract is wrong. | Either make the code do what the row says (retain ≥1 sample whenever `summary.unresolved > 0`) or correct the row to the narrower promise and reopen the defect it leaves. Do not leave `closed` standing over `shown: 0`. Add an assertion that no `cap_effect` receives `limit=0` by default on a path where `--limit 0` is published as *"no cap"* (`AUD-513-N08`), so the published semantics and the internal default cannot invert each other. |
227
+ | `AUD-592-A02` / `A-1` second half, `§12.5` | **P2** | **closed `5.8.18`** (`63f3d61`) in its wide half — **4 of 4**: `plan`, `impact-chain` and `compare` publish `scan_truncated`, and `endpoints` answering `null` is correct rather than missing (verified, not assumed: the CWD census of 3 575 against 2 802 summed over 10 of 16 repositories is plausible without truncation, consistent with the uncapped `rglob` this ledger already documents). **The narrow half stays open as `AUD-595-A07`**: the root help's 25 000 directory-entry cap is still unreconciled with the 8 000-file cap the payload publishes. Mechanism as found, kept for traceability: open — scan truncation is declared by 1 of 4 commands that own a walk. Over a truncated scan (cap observed at **8 000 files**), `impact` publishes `scan_truncated`; `plan` and `impact-chain` answer `not_found` with `scan_truncated: null`, and `endpoints` exits 1 with no payload — every one of them an exclusion derived from an incomplete population. `compare` cannot be measured (`AUD-592-R01`). Second, narrower defect in the same row: the root `--help` describes the cap as *"the scan stopped at 25 000 directory entries"* — a different limit in a different unit from the 8 000-file cap the field observes. | Publish the truncation beside every walk that can end in `not_found`, `candidates` or a census — the seventh pass already recorded that `extract_java_endpoints` does its own uncapped `rglob`, so this must be read at each call site, never centrally. Reconcile the two caps in the help text or declare both with their units. |
228
+ | `AUD-593-N02` / `N-02` | **P2** | **closed `5.8.18`** (`f539787`) **by deleting figures, not adding them** — the four hand-maintained numbers in the packaged guide became *"See `ask --help`"*; zero token bands remain in the shipped `.md`, and the measured payload (22 391 B ≈ 5 597 est. tokens) sits inside the generated ceiling. **The fix opened `AUD-594-X02`**: `--agent` deferred to an authority that publishes no figure for it. Mechanism as found, kept for traceability: open — the `--compact` budget is published three times, in two shipped authorities, with three values. `ask --help` generates *"~7K tokens (30 KB, chars/4)"* from `output_budget`, the module that enforces the limit — this half is the `AUD-590-B03` fix and it is correct. The packaged `docs/USER_GUIDE.md` is still hand-maintained and disagrees with it twice: `:111` *"A ~10K-token subset"* and `:243` / `:1366` *"~2,500–4,000 tokens"*; `--agent` carries a fourth figure at `:254` / `:1367` (*4,500–5,500*). Measured payload: 22 390 B ≈ 5 597 est. tokens — inside the help ceiling, **+40 % over the guide's table**. `grep '7K tokens' *.py` finds no literal: the help is generated and the guide is not, and that asymmetry is the defect. | Generate every token figure in the packaged docs from `output_budget.token_estimate`, or delete them and point at `--help`. Extend the existing release assertion (today help ↔ payload ↔ section registry) to cover `docs/*.md`: a token band in a shipped `.md` that does not come from that module fails the battery. Packaged documentation is a published authority and belongs inside the same assertion as the help text. |
229
+ | `AUD-593-N03` / `N-03` | **P2** | **closed `5.8.18`** (`1dd0b1c`) **in the generated half only** — the Options line is now built from `_BUDGETED_ANALYSIS_COMMANDS` (`cli.py:589-596`, read at `:632`) and enumerates all eight commands, wider than this row itself described. **The hand-written half is open as `AUD-594-X01`**: three literals in two other authorities still name three. Mechanism as found, kept for traceability: open — `--help` denies the scope that `AUD-590-B04` had just granted. With `ASK_MAX_ANALYSIS_SECONDS=90` the root analysis prints *"ask is a repo-wide analysis and the configured budget is 90s. Budget read from ASK_MAX_ANALYSIS_SECONDS=90 (process environment)"* — the fix works. The help page beside it still reads *"deadline only for spring-audit, risk, audit-report"* and *"Other commands are not time-bounded by this variable"* (`cli.py:621-622`). The root analysis is the most expensive command in the product and the one an agent most needs to bound; a reader of the help will not set the variable there. | Generate that line from the same registry `_build_analysis_classes` populates (the root is keyed there as `B14`), never by hand. Assertion: the set of commands the help names equals the set the preamble emitter covers. |
230
+ | `AUD-593-N05` / `N-05` | **P3** | **closed `5.8.18`** (`aa89388`) — the payload now carries `caller_count_unit: "caller_symbols"` and `caller_count_population: "references_to_target_or_its_declared_interfaces"` with a `FAN_IN_FIGURES` row behind them. Mechanism as found, kept for traceability: open — a figure emitted outside the catalogue that forbids it. `via_interface_resolution[].caller_count` publishes 37 where every other surface of the same fact says 23 (`interface_mediated_caller_count`, the list length, and the `explanation` prose), and it carries no `_unit` while the rest of the payload does. Both are legitimate and neither is declared: `caller_count` is `len(_iface_callers)` — raw caller symbols of the interface, before the `not in all_affected` filter and before normalisation to classes (`repository_ir.py:9248-9251`); the 23 is `caller_classes(_iface_mediated_callers, ...)` (`:9603`, `:10173`). It is also **not in `FAN_IN_FIGURES`**, which states that a figure absent from it may not be emitted — so `tests/test_fan_in_authority.py` did not catch a fan-in figure added in this release. | Give it a unit and a `FAN_IN_FIGURES` row, or rename it to `interface_caller_symbols`. Make the authority test sweep the **emitted** keys against the catalogue rather than the catalogue against itself; a catalogue that only validates its own rows cannot see an unregistered emission. |
231
+ | `AUD-593-N06` / `N-06` | **P3** | **closed `5.8.18`** (`aa584a6`) — `interface_mediated_callers_cap` is routed through `_cap_effect(...)` and emits `null` below the threshold, which is the correct behaviour at MSAS's 23. **Closed by instance, and the class survived it: `E-39`** — `serializer.py:480-489` still truncates `orphan_xml` and `missing_xml` at `[:5]` with no cap block. Mechanism as found, kept for traceability: open, latent — a silent `[:30]` where every sibling list declares its cap. `out["interface_mediated_callers"] = _iface_mediated_classes[:30]` (`repository_ir.py:10172`) with no `*_cap` block, in a payload where `direct_callers_cap`, `indirect_callers_cap` and `security_surface_affected_omitted` all carry `total`/`shown`/`omitted`/`direction`/`how_to_read`. On MSAS the list is 23, so nothing is visible; the first repository with more than 30 interface-mediated callers loses the excess without a word. | Route it through the same `_cap_effect(...)` used 27 lines above (`repository_ir.py:10145`). |
232
+ | `AUD-592-D01` / `D-1` | **P3 — residual of `AUD-591-A08`** | **closed `5.8.18`** (`0f8e03f`) — the root inventory keys on the triggering flag: `<repo>/.ask/readiness-history/` — only with `--snapshot` (optionally relocated by `--history-dir`). Mechanism as found, kept for traceability: open, confirmed in source. The write inventory in the root help says `migrate-check` writes *"only with `--history-dir`"* (`cli.py:296`). `--history-dir` only relocates the destination; the flag that writes is `--snapshot`, as `migrate-check --help` states correctly and as the code comment at `cli.py:12980-12986` states explicitly. Verified in the field: `migrate-check --history-dir <path>` exits 0 with zero writes and no guard message, while `--snapshot` fires the guard correctly. `AUD-591-A08`'s assertion checks that each cited flag **exists** in that command's parser — `--history-dir` does — not that it is the flag that triggers the write. | Key the inventory on `(command, triggering flag, path)` taken from the guard site itself, as `AUD-591-A08`'s own acceptance text required. An auditor tests what the inventory lists: this is the same mechanism by which `AUD-591-A01` went unfound. |
233
+ | `AUD-592-D02` / `D-2` | **P3 — residual of `AUD-591-A02`** | **closed `5.8.19`** (`1d3db4d`) **for both halves, from one module.** `--min-band` is a floor applied while the answer is composed, not a post-filter over rows, which is why it was left out of `_apply_selection` and why the exception cost what the rule was written to prevent. `selection_contract.min_band_filter_block()` now owns the disclosure: it merges rather than replaces (so `--band` and `--min-band` can both be in play), it returns nothing at the permissive floor (a payload that excluded nothing must not claim a selection), and it carries a basis saying the floor is not a row filter. `risk` publishes it from `_assemble_payload`, the one construction both its callers use. Field-verified on halo: `--min-band high` → `{min_band: high, total_before_filter: 0, matched: 0, shown: 0}`; default → `null`. Regression sweeps the commands offering the flag and fails any module that hand-builds the block. Mechanism as found: open in half, and the half that closed shows the fix the other needs. `f78522a` gave `enrich --min-band` its `_filter` block (`{min_band, total_before_filter, matched, shown}`, `cli.py:11666-11673`); `risk --min-band` still publishes `_filter: null` while `risk --band` publishes one — measured on `mall`: `--band high` → `_filter` with `total_before_filter: 50`, `--min-band medium` → `null`. The two call the same helper 250 lines apart and only one adds the block afterwards (`cli.py:11406-11413` against `:11658-11673`). Mechanism as found, kept for traceability: `--min-band` filters and validates (`cli.py:11294-11308` in `risk`, `:11563-11568` in `enrich`) but publishes no `_filter` block, while its sibling `--band` publishes one through `_apply_selection`. A consumer cannot tell a filtered answer from a small one. | Route `--min-band` through `_apply_selection` like every other selection flag, or publish the same `_filter` shape from wherever it is applied. |
234
+ | `AUD-592-D03` / `D-3` | **P3 — residual of `AUD-591-A02`** | **closed `5.8.18`** (`f28fe40`) — `enrich --rule NOPE-999` exits 1 against the same catalogue `spring-audit` validates against. Mechanism as found, kept for traceability: open. `enrich --rule NOPE-999` exits 0 with an empty result; `spring-audit --rule NOPE-999` exits 1 against the rule catalogue. One flag name, two contracts. Measured with a synthetic SARIF 2.1.0 (2 results): `--rule`, `--band` and an invalid `--band` all behave correctly in `enrich`; only unknown-rule validation is missing. | Validate against the same catalogue in both commands. `enrich` is experimental-tier, which sets the severity, not the contract. |
235
+ | `AUD-592-D04` / `build_commit` | **P3 — process** | **closed `5.8.19`** (`6eba05a`) **on the writer side, which is the side that was missing.** A `hatchling` build hook now writes `src/sourcecode/_build_commit.py` inside the build, where the checkout still exists, and the reader consults three sources narrowest first: `ASK_BUILD_COMMIT`, the stamped module, then the checkout — with `unrecorded` in the stamp treated as an absence that falls through rather than as an answer. The logic lives in `sourcecode.build_stamp`, importable without the build backend, because **a build-time behaviour the suite cannot exercise is how this row survived five reports**: the previous fix was asserted through a monkeypatched environment variable and nothing else. The hook never fails a build — an exported tarball has no git, and a release that cannot be built is worse than one that cannot name its commit. Verified end to end: stamp written, `ask version` publishes the commit, generated file gitignored so no checked-in copy can drift. Mechanism as found: open, fifth consecutive report, and the fix landed on the wrong side of the build boundary. `6d8eb1d` closed the *reader*: `product_info._build_commit()` takes `ASK_BUILD_COMMIT` and falls back to `git rev-parse HEAD` of the package checkout. **Nothing stamps it.** No `hatchling` hook, no release script and no CI step sets `ASK_BUILD_COMMIT` (`grep -rn ASK_BUILD_COMMIT` finds the reader, and a test that monkeypatches the variable — nothing that writes it), and an installed wheel has no git checkout, so the fallback returns `unrecorded` on every field install. The battery passed because it asserted the reader. Field verdict on 5.8.18 is unchanged: `build_commit: unrecorded`. Mechanism as found, kept for traceability: open, fourth consecutive report. `ask version` publishes `build_commit: unrecorded`, so no measurement in these ~150 invocations is attributable to a commit and *"fixed in 5.8.17"* is not falsifiable from outside. | Stamp the commit at build time. It is the cheapest row in this table and it is the precondition for every other row's verification. |
126
236
 
127
237
  **Mapped, not new — inherited rows this pass re-witnesses.** `B-6` (partial answers exit 0)
128
238
  keeps its row and gains its best measurement yet: overrun ×5.4 → ×4.2 → **×1.39** (median
@@ -535,8 +645,8 @@ retained only as evidence of what the auditor observed before the fix.
535
645
  | ASK-17 | **The parse store's default budget cannot hold a multi-repository workspace warm, and the cold cost of the largest repository is 7,6 minutes.** Measured on the 8-repo corpus (43 986 `.java`): `cache status` reports 56 730 entries at **511,12 MB against a 512 MB budget** — i.e. permanently sweeping by LRU — so `tutorials` (24 073 `.java`, the largest contributor) is the first candidate for eviction. Cold `--compact` on it: **457,6 s**, independently reproduced at 449 s, against 2,0 s on the second pass. | 5.8.4 corpus re-audit | Medium — warm plus `--compact` is still the answer (2,0 s), but a workspace this size cannot keep every repository warm at the default budget, and nothing tells the caller which repository is cold before it pays for it | **closed 5.8.5 — the disclosure shipped and the owed measurement taken, under our own control.** ⚠ **The +66 % does not reproduce, and the mechanism the row suspected is worth ~3 %, not 66 %.** Protocol: `tutorials` (24 073 `.java`), root `--compact`, both cache bases isolated **and emptied between runs** (`SOURCECODE_CONTEXT_CACHE_DIR` *and* `SOURCECODE_CACHE_DIR` — the first attempt isolated only the first, and the second pass of each version answered from the L2 view its own first pass had written: 2,4 s with an empty parse store, which is how a measurement of a cold path becomes a measurement of a warm one), `ASK_PARSE_CACHE_MAX_MB` fixed at 4 096 MB so no sweep can confound the comparison, 2 passes per version. **5.7.2 (`e803c8b`): 112,0 s / 112,6 s. This tree: 113,8 s / 115,2 s — +2,3 %**, with per-version dispersion of 1,005× and 1,012× and a payload 2,2 % larger. Then the audit's own condition, isolated as the only variable — the store pre-filled to 489 MB against the **default** 512 MB budget, so every 32 MB written sweeps: **117,0 s, +2,7 %.** So the LRU sweep is not where 457,6 s comes from, and neither is the analysis path: the auditor was right to hold the regression, and the held figure is now retracted **with a number** rather than on suspicion. ⚠ Scope, stated because the difference is unexplained rather than explained away: 457,6 s on Windows 11 / pipx is **not reproduced here** (113 s on 20 cores), and that gap is not assertable in either direction from this measurement — what is assertable is that 5.7.2 → this tree did not get slower and that a saturated store costs ~3 %. ✅ **What the measurement does confirm is the row's own claim, with our number: one repository of 24 073 `.java` leaves 46 840 entries and 470 MB in the store — 92 % of the 512 MB default budget** — so a workspace with a second repository of any size is permanently sweeping by construction, exactly as reported. Sizing rule, measured rather than guessed: ~20 KB of store per Java file, so ~500 MB per 24 000-file repository. Collateral confirmation of `AS-18`: after the saturated run the store rests at 669 MB — 157 MB over the budget and **under** the 736 MB effective ceiling it publishes at 7 writers — so the ceiling holds under the condition that produced the complaint. **The disclosure half:** `cache status` now publishes **coverage per repository**, which is the half the row itself recommended and the half that changes a decision: *«tutorials: 3 100 of 24 073 files cached (13 %)»* replaces a blind guess about whether to raise the budget, and an LRU eviction becomes visible **before** somebody pays 457,6 s to discover it. Both halves of the attribution were already held and thrown away — the walk knows the repository and it computes the store key for every file — so the run records the pairs (`parse_cache.record_repository_index`, from **both** readers of the store: `build_repo_ir` and the route-surface extractor, so the figure does not depend on which command was typed) and `store_stats` intersects them with the keys it collects **in the entry walk it already performs**: coverage costs an intersection, never a second scan of the store and never a scan of the repository. Three rules keep it honest. The index is **merged, never replaced**, because a `--changed-only` or `--since` run would otherwise shrink a 24 000-file population to the twelve files it read and publish *«12 of 12 cached (100 %)»* about a repository that is cold. The number travels with its **basis** — a file deleted since its last analysis still counts as recorded and reads as uncached, which errs toward *colder than it is* and says so. And the index is **neither an entry nor evictable**: the entry walk, the byte accounting and the LRU sweep all glob `*.json`, so its bytes are published under their own name (`repository_index_bytes`) rather than folded into a total that means entries — an index swept away with the entries it describes cannot report the eviction, which is the one moment it exists for. It lives inside the generation root, so retiring a generation retires its indexes with it: the keys are only readable by the build that wrote them. Regression `tests/test_parse_store_repository_coverage_ask17.py`, 11 assertions, including the one the row is about — entries deleted underneath a recorded repository make coverage **fall** while the population holds. ⚠ **Still owed, and unchanged:** the controlled cold measurement (5.7.2 against this tree, `ASK_PARSE_CACHE_MAX_MB` fixed, store emptied between runs, on a >20 000-file repository). Until it exists neither the +66 % nor its absence is assertable, and the row stays open on that half alone. **History:** **open** — ⚠ **not filed as a regression, deliberately**: the same cold figure was 274,8 s in 5.7.2 (+66 %), and the auditor refuses to call it one because the conditions are not comparable — in 5.7.2 the store had been retired by a version change, here it was mid-LRU-sweep. Seven false positives of exactly this class have been retracted over seven cycles; this is the eighth candidate and it is being held. What is owed is a measurement **we** control: cold `--compact` on a >20 000-file repository, 5.7.2 against 5.8.4, with `ASK_PARSE_CACHE_MAX_MB` fixed and the store emptied between runs — until that exists, neither the +66 % nor its absence is assertable. Recommended beside it, and cheap because both halves already exist: `cache status` should publish **coverage per repository** — how many entries belong to each analysed repository and what fraction of its files are covered — so *tutorials: 3 100 of 24 073 files cached (13 %)* replaces a blind decision about whether to raise the budget. Entries are content-addressed and the walk knows the repository, so this is a projection of facts we hold, not new analysis. **Refutation reproduced independently in cycle 8, under the reporter's own protocol**: both stores isolated and emptied, `ASK_PARSE_CACHE_MAX_MB=4096`, 2 passes per version — 5.7.2 at 112,0 / 112,6 s against 5.8.5 at 113,8 / 115,2 s (~3 %), and an A/B of the budget itself (512 MB sweeping by LRU against 4 096 MB that cannot sweep) at 9,1–10,1 s against 9,0–9,7 s: **the LRU sweep costs nothing measurable**. The +66 % is retired as the reporter's eighth false positive, with `E-37` named as its confounder. |
536
646
  | C3-128 | **`timeline` is the only gate-shaped command above 120 s, and it has not come back to its record.** Steady state, 4 runs, audited corpus: `timeline --since HEAD~5` costs **198 469 ms**, +27,3 % over its 5.7.2 record of 155 950 ms. The two other commands over 120 s are there structurally — `delta` (234 s) and `contract-diff` (136 s) analyse two whole trees — while `timeline` analyses **five** and costs less than `delta` does with two, so its cost is not explained by the number of trees it walks. | 5.8.2 → 5.8.4 re-audits | Low-Medium — an investigation command rather than a gate, and the last performance figure of the round sitting outside its own best | **closed 5.8.5 by measurement — the clock is attributed, and there is nothing left unaccounted to tune against.** The row's own instruction was ASK-16's: instrument before tuning. `timeline` published a per-sample total over what is really four costs — materialising a tree, measuring each watched metric, releasing the tree, and the remainder — so *«198 469 ms»* named a command rather than a phase. `perf.PhaseTimings` (the ASK-16 authority, always on) now splits it, with the phase names taken from `--watch` so the split cannot drift from the population, and the tree materialisation kept as its own phase because git's work must not be attributed to an analysis. **Measured, BroadleafCommerce (2 985 `.java`), `--since HEAD~5 --watch posture`, 5 samples: wall 38 921 ms — `measure:posture` 32 273 (82,9 %), `materialise_tree` 5 376 (13,8 %, ~1 075 ms per tree), `release_tree` 1 272 (3,3 %), unaccounted 0,45 ms (0,0 %).** So the answer to the row's premise — *«it analyses five trees and costs less than `delta` does with two»* — is that five sixths of the cost **is** the analysis, re-run per tree by construction, and the git work is a sixth of it: `timeline` is N × one analysis and there is no timeline-specific overhead to remove. Any future gain belongs to the metric being sampled (`C3-84`, `C3-127`), which is where it would also help every other command, and the payload now says so per run instead of per audit. Regression `tests/test_timeline_timings_c3_128.py`, 9 assertions, including that the series itself is byte-identical across two runs once the clock readings are removed — instrumentation that moved an answer would be a worse defect than the row. **Original note:** **open** — ⚠ the cache-reuse half of `B7` must **not** be reopened on this evidence: the v3 assertion (`max(sample) < 8 × posture_warm`) passes at 6,8 in both 5.8.2 and 5.8.4, against 14,8 / 13,4 / 12,4 / 8,5 in the four versions that genuinely failed it. What has not returned is the absolute cost. Attribution comes before tuning and is now cheap: `metadata.timings` (`ASK-16`) exists, so the per-phase split across the five trees can be published before anything is changed. |
537
647
  | B20 | **Thirteen declared renames with no cut-off date, and two distinct `1.0` identifiers meanwhile.** The registry publishes `pending_renames` for the 13 non-conforming `schema_version` values with their canonical name — exactly the policy `ASK-11` exists to enforce: the rename is an incompatible change, declared before it is made, never applied in silence. The consequence is published by the registry itself as `ambiguous_identifiers: 1` — `spring-audit` emits `1.0` for `core-analysis-v1` and `impact-chain` emits `1.0` for `impact-chain-v1`, so a consumer dispatching on the emitted version cannot tell them apart. | 5.8.4 re-audit (residual of `B19` / `B6`) | Low — declared debt rather than a defect, and the audit says so in as many words | **closed 5.8.5 — the window is declared where every other incompatible change is, and the reported ambiguity was understated by five shapes.** `BC-002` in `sourcecode.breaking_changes`: **announced 5.8.5, takes effect 6.0.0**, printed by `ask schema breaking-changes-v1`, carried in the CHANGELOG's `## Upgrading` section above the release history, and projected into `ask schema schemas-v1` as `pending_renames_window` — **one fact, two surfaces**, with a structural assertion that the registry source holds no second copy of the date. The registry's policy is generalised rather than loosened: a declared change is now `kind: exit_code` (must move an exit code) **or** `kind: contract` (must move a **published value** *and* name the version it takes effect in), because a rename breaks a consumer with every exit code still 0 and a registry that only knew about exit codes had nowhere to put it. A contract change publishes no `exit_code_before`/`after` at all — inviting a reader to check a field that cannot move is how a disclosure becomes noise. The 13 affected shapes are **read from `schema_registry.canonical_migrations()` at call time**, never copied: a shape that starts conforming leaves the declaration by itself (asserted by swapping the registry for a conforming one and watching the list empty). ⚠ **Correction to the reported cause, measured**: the audit named *two* shapes spelling their version `1.0`; there are **seven** — `core-analysis-v1`, `impact-chain-v1`, `pr-impact-v1`, `migration-blast-v1`, `spring-impact-v1`, `event-topology-v1`, `test-gap-ranking-v1` — so `ambiguous_identifiers: 1` was counting one *identifier* over seven shapes, not two. `ask schema 1.0` already resolved to all seven with each canonical name; the count is now asserted from the registry so it cannot be quoted from prose again. The rename itself is deliberately **not** made early: that is the incompatible change this policy exists to prevent. Regression `tests/test_schema_rename_window_b20.py`, 11 assertions, plus the ASK-11 battery generalised to both kinds. **Original note:** **open** — the missing half is a date, not a decision. Announce the cut-off window in `breaking-changes-v1` with the target version, the way every other incompatible change is announced, so a consumer can pin `core-analysis-v1` today and know when the bare `1.0` stops being emitted. Until then `ask schema 1.0` resolving to all seven shapes, each offering its canonical name, is the correct behaviour and must not be *fixed* by renaming an emitted value early — that is the incompatible change this policy exists to prevent. ⚠ **Round 10 ran on the 5.8.5 build and still reports the renames as *«deuda bien declarada, pero sin fecha»*, asking for precisely what `BC-002` already ships.** The row stays closed — the window exists, is announced 5.8.5 / effective 6.0.0, and is printed by `ask schema breaking-changes-v1`, carried in the CHANGELOG's `## Upgrading` and projected into `schemas-v1` as `pending_renames_window`. What it leaves behind is a **discoverability check, not a defect**: the reporter quoted `pending_renames` and `counts` out of the registry payload and did not see the window beside them, so verify that `pending_renames_window` travels in the same payload those two keys do — and if it does not, that is where it belongs. A declaration a ten-round auditor cannot find is not yet declared to a consumer. |
538
- | E-38 | **A class-level route prefix carried by a meta-annotation is dropped, and the route is published without it.** `_build_route_surface` reads the class prefix only when the class symbol carries `@RequestMapping` or `@Path` **literally** (`repository_ir.py:5882-5887`); a repository that declares its own composed annotation gets `prefixes = [""]`, and the published `path` is the method suffix alone. **Measured on shenyu (5.8.17): 35 of 365 published routes — 9,6 %, over 21 controllers — carry `path: "/"`.** `AiProxyApiKeyController` is annotated `@RestApi("/selector/{selectorId}/ai-proxy-apikey")`, a Shenyu annotation meta-annotated with `@RestController` + `@RequestMapping`; its five handlers are published at `/`, `/`, `/batchDelete`, `/{id}`, `/page`. The served URLs do not exist as published. This is a **confident falsehood**, not a gap: no `path_resolution: "unresolved"`, no `path_expression`, no warning, and `route_census.distinct_routes` (316 of 365) reads the collisions as if two controllers genuinely shared a path. It also propagates to every consumer keyed on route text — `explain-endpoint`, `data-exposure --path-prefix`, `validation --path-prefix`, `pr-impact` route matching — where a correct query returns nothing. **The machinery already exists on the other axis**: `E-12`/`5.7.1` closed exactly this class for *beans* by resolving the spelling against the graph's `imports` edges, and `BeanGraph.build` already builds a meta-annotation map in its first pass. The route axis never asked. Vendor-agnostic by construction: the rule is "an annotation this repository declares that itself carries `@RequestMapping`/`@Path`", never a proprietary name. | found 2026-08-21 during the `AUD-590-R01` census triage, on the corrected 5.8.17 build; **not part of the seventh-pass queue and deliberately not fixed in it** — it moves route counts on any repository using composed annotations and needs its own measured pass with the golden set re-baselined | **High** — it is the rule this repository enforces most loudly, inverted on the route axis: a published route that is not served, with no affordance saying so. Blast radius is every repository with a framework-level or in-house composed controller annotation; shenyu is the witness, dubbo/sa-token/halo are unmeasured | **open** — direction: resolve the class-level prefix through the meta-annotation map `BeanGraph.build` already computes, in the order `E-12` established (explicit single-type import decides; wildcard leaves it admissible; a repository-declared annotation in the owner's package is the owner's own; **nothing resolved stays unresolved rather than acquiring a verdict from silence**). Where it cannot be resolved, publish `path_resolution: "unresolved"` with the annotation as `path_expression` — the contract the method-level path already honours — never a bare `/`. Regression on a fixture declaring its own composed annotation, plus a shenyu assertion that no published route is `/` |
539
- | B29 | **A cache-freshness test asserts a cold cache without ensuring one, so it fails on every run but the first after a purge.** `tests/test_cli.py::test_compact` asserts `_meta.timing_ms.served_from_cache is False` and does nothing to make that true: the first `--compact` run over its fixture populates the cache the second one then hits. Verified symmetrically at `d9071d4` (clean) and at every commit of the seventh-pass queue — purge, run, pass; run again, fail — so it is **state-dependent, not a regression**, and it has been reported as "1 pre-existing failure" for the whole pass. It is still a red line in a suite whose value is that red means something. | found 2026-08-21 while establishing the baseline for the seventh-pass queue | **Low** as a defect, **medium** as an eroder — a permanently-red test trains a reader to skip the summary line, which is how a real regression gets shipped | **open** — direction: the test must create the state it asserts (purge this fixture's cache dir in the fixture, or assert the *transition* cold→warm across two runs rather than the absolute). Do not "fix" it by deleting the assertion: `served_from_cache` is the field `is_stale`/cache-freshness rows are keyed to, and a cold-start claim is worth holding |
648
+ | E-38 | **A class-level route prefix carried by a meta-annotation is dropped, and the route is published without it.** `_build_route_surface` reads the class prefix only when the class symbol carries `@RequestMapping` or `@Path` **literally** (`repository_ir.py:5882-5887`); a repository that declares its own composed annotation gets `prefixes = [""]`, and the published `path` is the method suffix alone. **Measured on shenyu (5.8.17): 35 of 365 published routes — 9,6 %, over 21 controllers — carry `path: "/"`.** `AiProxyApiKeyController` is annotated `@RestApi("/selector/{selectorId}/ai-proxy-apikey")`, a Shenyu annotation meta-annotated with `@RestController` + `@RequestMapping`; its five handlers are published at `/`, `/`, `/batchDelete`, `/{id}`, `/page`. The served URLs do not exist as published. This is a **confident falsehood**, not a gap: no `path_resolution: "unresolved"`, no `path_expression`, no warning, and `route_census.distinct_routes` (316 of 365) reads the collisions as if two controllers genuinely shared a path. It also propagates to every consumer keyed on route text — `explain-endpoint`, `data-exposure --path-prefix`, `validation --path-prefix`, `pr-impact` route matching — where a correct query returns nothing. **The machinery already exists on the other axis**: `E-12`/`5.7.1` closed exactly this class for *beans* by resolving the spelling against the graph's `imports` edges, and `BeanGraph.build` already builds a meta-annotation map in its first pass. The route axis never asked. Vendor-agnostic by construction: the rule is "an annotation this repository declares that itself carries `@RequestMapping`/`@Path`", never a proprietary name. | found 2026-08-21 during the `AUD-590-R01` census triage, on the corrected 5.8.17 build; **not part of the seventh-pass queue and deliberately not fixed in it** — it moves route counts on any repository using composed annotations and needs its own measured pass with the golden set re-baselined | **High** — it is the rule this repository enforces most loudly, inverted on the route axis: a published route that is not served, with no affordance saying so. Blast radius is every repository with a framework-level or in-house composed controller annotation; shenyu is the witness, dubbo/sa-token/halo are unmeasured | **closed `5.8.18`** (`096c7c3`) — direction taken as written: resolve the class-level prefix through the meta-annotation map `BeanGraph.build` already computes, in the order `E-12` established (explicit single-type import decides; wildcard leaves it admissible; a repository-declared annotation in the owner's package is the owner's own; **nothing resolved stays unresolved rather than acquiring a verdict from silence**). Where it cannot be resolved, publish `path_resolution: "unresolved"` with the annotation as `path_expression` — the contract the method-level path already honours — never a bare `/`. Regression on a fixture declaring its own composed annotation, plus a shenyu assertion that no published route is `/` |
649
+ | B29 | **A cache-freshness test asserts a cold cache without ensuring one, so it fails on every run but the first after a purge.** `tests/test_cli.py::test_compact` asserts `_meta.timing_ms.served_from_cache is False` and does nothing to make that true: the first `--compact` run over its fixture populates the cache the second one then hits. Verified symmetrically at `d9071d4` (clean) and at every commit of the seventh-pass queue — purge, run, pass; run again, fail — so it is **state-dependent, not a regression**, and it has been reported as "1 pre-existing failure" for the whole pass. It is still a red line in a suite whose value is that red means something. | found 2026-08-21 while establishing the baseline for the seventh-pass queue | **Low** as a defect, **medium** as an eroder — a permanently-red test trains a reader to skip the summary line, which is how a real regression gets shipped | **closed `5.8.18`** (`684a9df`) — direction taken as written: the test must create the state it asserts (purge this fixture's cache dir in the fixture, or assert the *transition* cold→warm across two runs rather than the absolute). Do not "fix" it by deleting the assertion: `served_from_cache` is the field `is_stale`/cache-freshness rows are keyed to, and a cold-start claim is worth holding |
540
650
  | E-37 | **Two environment variables move three cache stores, and no surface said which moves which.** `SOURCECODE_CACHE_DIR` relocates the per-repository snapshot store (core snapshots, the rendered L2 view, the RIS); `SOURCECODE_CONTEXT_CACHE_DIR` relocates the shared ones (the Canonical IR and the per-file parse store). Their defaults are siblings under `~/.sourcecode`, so the split is not derivable from the paths, and `cache status` published a path per store with no variable beside it while `docs/CACHE.md` named both variables in one sentence after the table. | found here 2026-08-18 while taking the controlled A/B `ASK-17` owed, not reported by the field | Medium — it does not make an answer wrong, it makes a **measurement** wrong, silently: an operator who redirects or empties *the cache* moves one store and is served answers out of the other | **closed 5.8.5** — measured first: with only `SOURCECODE_CONTEXT_CACHE_DIR` redirected, the second cold pass of each version answered from the L2 view its own first pass had written — **2,4 s against 113 s, with an empty parse store** — and the run reported `cache_source: L2_view` without anything naming the store it came out of. Now every store in `cache status` publishes `base_env`, from one authority (`cache._STORE_BASE_ENV`), the text answer prints *«<path> ($VAR moves it)»* beside each one, `stores.base_env_note` states the consequence rather than the layout, and `docs/CACHE.md` carries a **column** instead of a sentence covering both. Regression `tests/test_cache_store_base_env_e37.py`, 8 assertions, and the ones that matter are **behavioural**: setting the variable a store names moves *that* store and leaves the others where they were — a published mapping nobody exercises is the class of claim this repository refuses everywhere else. The CLI is forbidden a second copy of the mapping (structural assertion). **Confirmed from the outside in cycle 8, by the reporter falling into it first**: isolating only `SOURCECODE_CONTEXT_CACHE_DIR`, their second *cold* pass answered from the L2 view their first pass had written — **2,4 s against 113 s with an empty parse store** — and their write-up now quotes the shipped `base_env` mapping back as the fix (*"dos variables mueven tres almacenes y ninguna superficie decía cuál"*). A defect found while measuring, that was itself corrupting the measurement. |
541
651
  | B21 | **Six of the thirteen shapes being renamed name no command that emits them.** `schemas-v1` publishes `emitted_by` per shape; for `migration-blast-v1`, `spring-impact-v1`, `event-topology-v1`, `test-gap-ranking-v1`, `canonical-ir-v1` and `hibernate-strategy-v2` it is `[]`, beside seven shapes where it is populated. Read the way a consumer reads a list, an empty `emitted_by` says *nothing emits this shape* — which tells the one consumer that **is** affected by `BC-002` that it is not. | found here while declaring `BC-002` (5.8.5), not reported by the field | Low — a disclosure gap in a declaration whose whole purpose is to let a consumer decide whether it is affected | **closed 5.8.6 — derived, not typed out.** `schema_producers` reads the two authorities the CLI already holds — the registered command tree and the import graph of the module that defines each callback — and answers with the evidence it derived on: `command_import` (the command's own body imports the producing module), `module_import` (one hop, through an intermediate narrow enough to attribute), `substrate` (reached from more commands than a producer would be) or `unresolved` (nothing found, said in a sentence). Five of the six now name a command — `migration-blast-v1` → `migrate-check`, `spring-impact-v1` and `event-topology-v1` → `impact-chain`, `test-gap-ranking-v1` → `impact`, `hibernate-strategy-v2` → `migrate-check`/`migrate-recipe` — and the sixth, `canonical-ir-v1`, is reached from **28 modules**: naming every command would be less true than naming none, so it publishes `substrate` with the count and the instruction to match on `subject`. A declared producer still wins, because some producers are not one registered command (`ask (root)`, `baseline capture`). Regression `tests/test_schema_producers_derived_b21.py`, 33 assertions, parametrised by the registry so a fourteenth unnamed shape fails the battery. **History:** open, and named rather than shipped as a fact: every `BC-002` row with an empty producer list carries `emitted_by_basis` — *«not declared in the registry: this shape's producing command is not named yet, so match on `subject`»* — so the gap is visible instead of being read as a negative claim. What remains is to fill the six, and the honest way is derivation rather than six more hand-written strings: the producer is discoverable from the command that constructs each payload, which is the same authority every other population in this repository comes from. Until then the assertion in `tests/test_schema_rename_window_b20.py` holds the weaker invariant a reader can rely on: a row either names its emitters or says why it cannot. |
542
652
  | C3-129 | **`timeline`'s absolute cost, now fully attributed and still 27 % above its own record.** `C3-128` closed by instrumenting rather than tuning, which was the right order and left this behind: the clock is split, nothing is unaccounted, and the number has not moved. Round 10 on 5.8.5 measures `timeline --since HEAD~5` at **~198 000 ms against the 5.7.2 record of 155 950 ms (+27 %)** — the only gate-shaped command over 120 s that is not there structurally (`delta` 191 s and `contract-diff` 115 s walk two whole trees, `compare` 46 s walks N candidates; `timeline` walks **five** trees and costs more than `delta` does with two). | 5.8.2 → 5.8.5 re-audits (residual of `C3-128`, which closed the attribution half in 5.8.5) | Low-Medium — an investigation command rather than a CI gate, and the last performance figure of the round sitting outside its own best | **closed 5.8.6 — by not paying twice for what does not change.** A metric at a commit is a pure function of that tree, and the tree at a sha never changes: the only cache in this product whose key can be exact rather than heuristic (analyzer fingerprint + metric + commit sha). `timeline_cache.SampleCache` stores measured values under the per-repository core store, and when every watched metric for a sample is a hit the tree is **never materialised**, which is the 13,8 % the split attributes to git. Measured on BroadleafCommerce over the same range: **26 845 ms → 12 ms**, samples and transitions byte-identical. What makes it admissible rather than merely fast, all asserted: a failed sample is never stored (one bad run cannot become permanent), the fingerprint is in the key (corrected analysis is never served from a previous release), the hit is published (`sample_cache`, `from_cache` per sample, and `cost` separates the cached population from the measured one instead of reporting a median over both), and `--no-cache` / `ASK_TIMELINE_NO_SAMPLE_CACHE=1` measure everything again so the command stays measurable without emptying a store (E-37). `cache_model` carries the layer, so `cache model` and `docs/CACHE.md` state it from one authority. Regression `tests/test_timeline_sample_cache_c3_129.py`, 11 assertions. **History:** open — and it opened with the answer already published, which is why it is cheap to attack and expensive to leave: `timeline`'s own `timings` say **82,9 % `measure:posture`, 13,8 % `materialise_tree`, 3,3 % `release_tree`, 0,0 % unattributed**. So this is not a `timeline` defect at all: it is `posture` paid five times over five materialised trees, and it makes the same substrate question `C3-127` asks — with the difference that here the five payers are one command, so the cache is intra-invocation and needs no cross-command contract. ⚠ **Do not reopen `B7`**: cross-tree cache reuse is healthy, the v3 assertion passes at 6,8×. Acceptance: `timeline --since HEAD~5` back inside +10 % of 155 950 ms on the audited corpus with `timings` still summing to 100 %, and `measure:posture` falling as a share rather than the total falling for an unnamed reason. |
@@ -587,7 +697,7 @@ retained only as evidence of what the auditor observed before the fix.
587
697
  | C1-27 | **Two definitions of "the working tree changed" inside one command.** `cache freshness` reports STALE with `RIS HEAD == current HEAD` and `Delta: 0` because **one untracked file no analyser reads** (an editor settings file) is a porcelain line. C3-40 (4.5.3) replaced a false *fresh* with a false *stale*: its predicate is *any* `git status` entry, while the fact the snapshot describes is *the files the analysis admits*. Reproduced on a battery repository with a clean tree: `FRESH` → write `.claude/settings.json` → `STALE` | 4.5.3 (eval #11) | Medium — a freshness signal that fires on a file the analysis cannot see teaches a reader to ignore the one that matters, and it invalidates a warm cache for nothing on a repository where a cold run costs minutes | **closed 4.6.0 — verified in the field (eval #12).** One authority decides which paths qualify (`path_filters.analysis_reads`) and both consumers ask it: the freshness boolean and the **tree signature**, which is also the cache key — so an agent writing its own settings file was discarding every warm answer. `.github` stays a change (tooling evidence); the default is inclusive, because a missed change is a false *fresh*. Field: `FRESH` / `Uncommitted: False` with the same untracked settings file still present |
588
698
  | C1-28 | **The test/production split still has two authorities.** C1-2 made `test_sources` the authority in 3.2.1 and rebound five consumers; `path_filters.is_test_path` was never retired and still answers for its own callers. They disagree on five path shapes, in both directions — `src/main/java/**/Test*.java` (`False` / **`True`**), `src/main/java/**/*IT.java` (**`True`** / `False`), `src/main/java/**/test/Helper.java` (`False` / **`True`**). Field symptom: `--compact` counts **1 test file that lives in `src/main`**, with the verdict (0.0 %, critical) correct and the numerator wrong | 4.5.3 (eval #11) | Medium — it is the denominator of every coverage statement, and the direction that inflates is the one a reader must never be misled about (C1-2's own rule) | **closed 4.6.0 — verified in the field (eval #12).** `path_filters.is_test_path` is now a passthrough to the authority, after absorbing everything it knew that the authority did not, and a declared **main** source root outranks a naming convention. Fleet effect, cold: keycloak 5 525 → 5 486 Java files with endpoints **676 → 678**, openmrs-core 866 → 861, Broadleaf / alfresco / petclinic unchanged. Field: `0 test files for 3337` **with the reason stated** (no declared test source root, no file matching a test naming convention) — the numerator and its basis, which is what made it verifiable from outside |
589
699
  | C1-29 | **Candidate, not reproduced: three commands publish three severities for one `defect_id`.** Eval #12 states it and could not measure it — verifying it costs ~45 min of `risk` (C3-42). Recorded so the complaint is not lost, and recorded as a **candidate** because the standing rule since audit #8 is that no row is opened on an unreproduced complaint. What must be measured, in this order: (1) do `spring-audit`'s rule severity, `risk`'s `severity_declared` / `severity_effective` band and `impact-chain`'s per-finding severity (beside its own `risk_level`) ever publish **the same field name with different values** for one `defect_id` — that is the defect; (2) or do they publish **different names for different units** — that is the design, and `risk` already does it correctly today (`severity_declared` carries `severity_authority: spring-audit`, `severity_effective` is the composed product). The likely real row, if one survives, is the **third** publication: `impact-chain` emits findings with a severity and a `risk_level` of its own, computed by a scale that is not `RiskComposer`'s | 4.6.0 (eval #12) | Unknown until reproduced — potentially **High** (it is the composition axis, scored 4/10), potentially not a defect at all | **does not reproduce (measured 4.10.0)** — on the battery fixture, the same `defect_id` carries `severity: medium` in `spring-audit`, `impact-chain.impact_findings[].severity: medium`, and `risk.severity_declared: medium`; the composed result is deliberately a different field (`severity_effective` plus `band`) with factor authority `spring-audit`. That is the intended naming case, not a contradiction. Locked by `tests/test_risk.py::test_c129_declared_and_effective_severity_are_different_units` |
590
- | C1-36 | **A container-wired component has a call-graph fan-in of ~0, so every risk surface publishes its blast radius as zero.** Reproduced twice, on two builds, by two evaluations. Eval #14 (4.10.4): `pr-impact . --files -` over `SecurityConfig.java` + `M3FiltroSeguridadAspect.java` + `LegacySha1PasswordEncoder.java` — the application's whole filter chain, the aspect that gates 935 routes, and the production password encoder — returns `risk_level: LOW`, `risk_reason: "No high-risk signals detected"`, `affected_endpoints: []`, `direct_callers: []`. Positive control, same command, same cache: a single POJO (`Persona.java`) returns `CRITICAL`, *"Public API + High Call Fan-in"*, 3 382 endpoints, 48 callers. Eval #15 (4.10.5) reproduced it on a two-file diff (`SecurityConfig` + `LegacySha1PasswordEncoder`): `LOW`, 2 classes, **0 endpoints** — *"un PR que borrase `.antMatchers("/privado/**","/api/**").authenticated()` pasa el gate en verde"*. Same root in three more commands: `impact-chain M3FiltroSeguridadAspect` → `risk: low`, 0 callers, 0 endpoints; `impact-chain LegacySha1PasswordEncoder` → `low`, 1 caller, 0 endpoints; `plan M3FiltroSeguridadAspect` → `affected_endpoints.count: 0`, `rollback_surface: 1 file`; `compare SecurityConfig M3FiltroSeguridadAspect` → both `blast_radius: 0`, `secured_endpoints: 0`, ranked 2nd and 3rd behind a directory (C3-67). **The knowledge is already in the binary and no surface consults it**: `endpoints` resolves `@M3FiltroSeguridad` → 935 routes (`custom_gate_inferred`); `explain SecurityConfig` emits `request chain: /privado/**, /api/** → authenticated (SecurityConfig.java:47)`; `posture` computed `open_by_rule: 2635` for `default` against `1` for `m3` **from that same class**. Two commands of one binary disagree about whether the aspect matters | 4.10.4 (eval #14), 4.10.5 (eval #15) | **Critical** — it is a security false negative on a command the `--help` sells as *"contract stable within a major — safe to gate CI on"*, and it is the one important blind spot the product does **not** declare: every other gap has an `NC-0xx`, this one has none. Eval #14 scores risk scoring **3/10** for it and states 7 → 8,5-9 on this row plus C3-53 alone | **remedies (1), (2) and (3) closed (4.10.6); (4) is F-Y and stays open.** The three cheap ones landed together because separating them would have shipped a warning nobody could read next to a verdict that still said `LOW`. One authority, `container_wiring.py`, answers *is this class reached by the container rather than by the code?* from facts the CIR already carries — the role it already assigns (`config`), the interception annotations and supertypes fixed by published specifications, and one wholly structural route that needs no name at all: **the class implements an interface something injects and has zero incoming call edges**, which is how the password encoder is caught without anyone naming a password encoder. It is wired at the *one* blast-radius authority (`_compute_blast_radius`, which `impact`, `plan` and `compare` all project from) plus `pr_impact` and `spring_impact`, so the four commands that reproduced the defect answer from one place rather than four. The verdict is **floored, never lowered** — a `CRITICAL` stays `CRITICAL` — the payload publishes the population (`container_wired`) and the reason, and `NC-010` states the limit in the payload, the README and the help. Locked by `tests/test_container_wiring.py` (16), which asserts the observable verdict on real sources — an `@Aspect` with fan-in 0 rates `high` and says why, an ordinary service with callers is untouched — and by the negative controls that keep the floor from becoming a blanket (`pom.xml` and a CI workflow are not wiring descriptors). Field-shaped remedies, in ascending cost. (1) **Half an hour, today: declare the non-coverage.** An `NC-0xx` stating that blast radius is projected over the call graph, that container-wired components (`@Configuration`, `@Aspect`, `Filter`, `web.xml`) have fan-in ≈ 0 and therefore appear with null impact, and that their real risk is not measured here. This converts a silent failure into an honest limit, which is the whole promise of the product. (2) **`analysis_warnings` when ≥1 analysed class is container-wired and its fan-in is 0** — today the only warning on that payload is the self-referential exclusion, cosmetic beside this. (3) **A floor by file class**: a diff touching `@Configuration`, `@Aspect`, `Filter`, `@ControllerAdvice`, `PasswordEncoder`, `web.xml` or `application.y*ml` cannot publish below `HIGH`, with `risk_reason: "container-wired component — call-graph fan-in is not a risk proxy here"`. (4) **The real fix, F-Y**: pointcut → surface (an `@Aspect` whose pointcut is `@annotation(X)` inherits the routes carrying `X` — 935, not 0) and matcher → routes (a modified `SecurityFilterChain` inherits the population its matchers cover). Also see CL-16: while this is open, `pr-impact` cannot honestly stay in `core` |
700
+ | C1-36 | **A container-wired component has a call-graph fan-in of ~0, so every risk surface publishes its blast radius as zero.** Reproduced twice, on two builds, by two evaluations. Eval #14 (4.10.4): `pr-impact . --files -` over `SecurityConfig.java` + `M3FiltroSeguridadAspect.java` + `LegacySha1PasswordEncoder.java` — the application's whole filter chain, the aspect that gates 935 routes, and the production password encoder — returns `risk_level: LOW`, `risk_reason: "No high-risk signals detected"`, `affected_endpoints: []`, `direct_callers: []`. Positive control, same command, same cache: a single POJO (`Persona.java`) returns `CRITICAL`, *"Public API + High Call Fan-in"*, 3 382 endpoints, 48 callers. Eval #15 (4.10.5) reproduced it on a two-file diff (`SecurityConfig` + `LegacySha1PasswordEncoder`): `LOW`, 2 classes, **0 endpoints** — *"un PR que borrase `.antMatchers("/privado/**","/api/**").authenticated()` pasa el gate en verde"*. Same root in three more commands: `impact-chain M3FiltroSeguridadAspect` → `risk: low`, 0 callers, 0 endpoints; `impact-chain LegacySha1PasswordEncoder` → `low`, 1 caller, 0 endpoints; `plan M3FiltroSeguridadAspect` → `affected_endpoints.count: 0`, `rollback_surface: 1 file`; `compare SecurityConfig M3FiltroSeguridadAspect` → both `blast_radius: 0`, `secured_endpoints: 0`, ranked 2nd and 3rd behind a directory (C3-67). **The knowledge is already in the binary and no surface consults it**: `endpoints` resolves `@M3FiltroSeguridad` → 935 routes (`custom_gate_inferred`); `explain SecurityConfig` emits `request chain: /privado/**, /api/** → authenticated (SecurityConfig.java:47)`; `posture` computed `open_by_rule: 2635` for `default` against `1` for `m3` **from that same class**. Two commands of one binary disagree about whether the aspect matters | 4.10.4 (eval #14), 4.10.5 (eval #15) | **Critical** — it is a security false negative on a command the `--help` sells as *"contract stable within a major — safe to gate CI on"*, and it is the one important blind spot the product does **not** declare: every other gap has an `NC-0xx`, this one has none. Eval #14 scores risk scoring **3/10** for it and states 7 → 8,5-9 on this row plus C3-53 alone | **closed 5.8.20 — all four remedies. (1), (2) and (3) landed together in 4.10.6; (4), F-Y, landed here.** The floor said the fan-in cannot describe this component and nothing said what can, so the honest `HIGH` shipped beside an empty endpoint list for nine releases. F-Y measures the reach from what the component itself declares, in one authority (`wired_surface`) that all three derivations consume — `compute_blast_radius` (impact / plan / compare), `spring_impact` (impact-chain) and `pr_impact`. **pointcut → surface**: an `@Aspect` advises the members its pointcut selects, so `@Around("@annotation(X)")` reaches every handler carrying `X`. The expression is the fact and the IR did not carry it — `_ADVICE_ANN_ARGS` joins the annotations whose argument already is one, so nothing re-parses Java to recover it. **matcher → routes**: a chain covers the routes its matchers select, read through `chain_rules`, the authority `posture` already uses, rather than a second extractor. Measured on a fixture reproducing the field's shape: `impact` and `impact-chain` both publish 2 of 3 routes for the aspect (the handler without the annotation is excluded) and `pr-impact` publishes 3 via the chain's `/orders/**` matcher, each figure carrying its `_unit`, `_population`, `_basis` and `direction: under_reports`. **The floor is deliberately untouched**: `risk_level` stays `high` and `risk_score` stays `null` (C1-41) — the band was established *without* the fan-in and a measurement of reach is not fan-in arithmetic. A component declaring neither mechanism gets `null`, never `0`; a matcher rule `chain_rules` records as `not_evaluated` is counted `undecided`, never covered. `NC-010` now states what is measured and what is still not — a pointcut this analysis cannot resolve, a matcher built by a helper method, a filter registered programmatically — because *"nothing here measures this"* printed beside a measurement is the contradiction this product is audited for. Locked by `tests/test_wired_surface_fy.py` (15), which asserts through the commands on real sources, including that **`impact` and `impact-chain` publish the same count for one symbol** — this row's own sentence was *two commands of one binary disagree about whether the aspect matters*. Registered in `defect-classes-v1`: `container_reachers` flips to `swept: true` on a sweep that iterates `KINDS`, and `wired_component_reach` joins over `REACHES`. **What this closes for `CL-16`**: `pr-impact` keeps its `core` tier on the evidence the tier promises, not on a floor. Previously — The three cheap ones landed together because separating them would have shipped a warning nobody could read next to a verdict that still said `LOW`. One authority, `container_wiring.py`, answers *is this class reached by the container rather than by the code?* from facts the CIR already carries — the role it already assigns (`config`), the interception annotations and supertypes fixed by published specifications, and one wholly structural route that needs no name at all: **the class implements an interface something injects and has zero incoming call edges**, which is how the password encoder is caught without anyone naming a password encoder. It is wired at the *one* blast-radius authority (`_compute_blast_radius`, which `impact`, `plan` and `compare` all project from) plus `pr_impact` and `spring_impact`, so the four commands that reproduced the defect answer from one place rather than four. The verdict is **floored, never lowered** — a `CRITICAL` stays `CRITICAL` — the payload publishes the population (`container_wired`) and the reason, and `NC-010` states the limit in the payload, the README and the help. Locked by `tests/test_container_wiring.py` (16), which asserts the observable verdict on real sources — an `@Aspect` with fan-in 0 rates `high` and says why, an ordinary service with callers is untouched — and by the negative controls that keep the floor from becoming a blanket (`pom.xml` and a CI workflow are not wiring descriptors). Field-shaped remedies, in ascending cost. (1) **Half an hour, today: declare the non-coverage.** An `NC-0xx` stating that blast radius is projected over the call graph, that container-wired components (`@Configuration`, `@Aspect`, `Filter`, `web.xml`) have fan-in ≈ 0 and therefore appear with null impact, and that their real risk is not measured here. This converts a silent failure into an honest limit, which is the whole promise of the product. (2) **`analysis_warnings` when ≥1 analysed class is container-wired and its fan-in is 0** — today the only warning on that payload is the self-referential exclusion, cosmetic beside this. (3) **A floor by file class**: a diff touching `@Configuration`, `@Aspect`, `Filter`, `@ControllerAdvice`, `PasswordEncoder`, `web.xml` or `application.y*ml` cannot publish below `HIGH`, with `risk_reason: "container-wired component — call-graph fan-in is not a risk proxy here"`. (4) **The real fix, F-Y**: pointcut → surface (an `@Aspect` whose pointcut is `@annotation(X)` inherits the routes carrying `X` — 935, not 0) and matcher → routes (a modified `SecurityFilterChain` inherits the population its matchers cover). |
591
701
  | C1-37 | **Four incompatible figures for "endpoints" at one commit, and a baseline that disagrees with itself.** `impact Persona . → stats.endpoints_affected_count` = **3 364**; `pr-impact . --files -` over the same class → `affected_endpoints` = **3 382**; `data-exposure . → summary.endpoints_exposed` = **3 382**; `endpoints .` total = **3 574**. Inside one artefact, `.ask/baselines/3dde0376.json`: `totals.endpoints` = **3 574** against `len(endpoint_surface)` = **3 538** — a file whose declared purpose is to be a replayable fingerprint differs from itself by 36 endpoints | 4.10.4 (eval #14) | **High** — it breaks comparability between commands and the credibility of every ratio keyed on an endpoint count. Eval #14 scores numeric consistency **5/10** and names this first | **closed (4.10.6).** The baseline half had a cause, not just a discrepancy: `totals.endpoints` counts **handler mappings** and `endpoint_surface` is the set of distinct **(METHOD, path) contract pairs**, so the 36 are routes declared by more than one handler. Both facts are now named — `totals.endpoint_contract_pairs` beside `totals.endpoints`, with `endpoints_unit` saying which is which — and the battery holds `len(endpoint_surface) == totals.endpoint_contract_pairs`, an invariant that is checkable because the two figures finally describe two things. The cross-command half applies the `direct_callers_note` pattern to all four: `endpoints.total_unit` (the declared population), `impact.stats.endpoints_affected_count_unit` (one blast cone), `pr-impact.metadata.affected_endpoints_unit` (a UNION over the changed classes, keyed on endpoint id — which is exactly why it can exceed a single cone) and `data-exposure.summary.endpoints_exposed_unit` (reach-selected subset). Each names the others, so a reader who meets one figure learns the rest exist. Original remedy note: the pattern is already solved in this product: `direct_callers_note` declares the unit of its figure and what references it admits, and says so after measurement falsified an earlier promise. Apply it to the endpoint counts: every published count carries its `unit`; `impact` and `pr-impact` either declare why they differ by 18 or converge; and the battery gets the invariant `len(endpoint_surface) == totals.endpoints` unless a declared cap says otherwise — `onboard` already has the shape for that (`_truncation_summary`) |
592
702
  | C1-38 | **`secured_endpoints` names four different populations, and one of them is reconcilable with nothing.** At one commit: `endpoints.exposure.by_security_policy.custom_gate_inferred` = **935**; `data-exposure.by_security_verdict.protected_custom` = **939**; `impact Persona.stats.security_surface_count` = **15**; `compare` (whole-repo candidate) `secured_endpoints` = **15**. 935 vs 939 is an axis difference and defensible — but it is not declared. **15 matches nothing**: not 4 (`programmatic`), not 935, not 939 | 4.10.4 (eval #14) | **Medium-High** — it is the field an auditor reads first, and *"un campo llamado `secured_endpoints` que devuelve 15 sobre un repo con 935 rutas anotadas es una trampa para el lector"* | **closed (4.10.6), and the cause was not naming.** The `15` reconciled with nothing because it was **`min(real, 15)`**: `security_surface_affected` was truncated to a 15-row display sample and `stats.security_surface_count` was `len()` of the truncated list — a rendering choice that had become the measurement, the exact R5 defect the rest of that file already fixes. The population is now counted before the cut, the list is still sampled, and the omission is published with its `cap_effect`. The naming half shipped with it: all four figures now carry their unit beside them in the `direct_callers_note` style — `by_security_policy_unit` (annotated routes), `by_security_verdict_unit` (posture verdicts over the same routes, a different axis), `security_surface_count_unit` (declarations inside one blast cone) and `compare.cost_units.secured_endpoints` — each naming the other two so a reader who meets one figure learns the other two exist. Original remedy note: rename to what each one measures (`security_annotated_routes` / `programmatic_gate_routes` / `security_gated_in_blast_cone`) and attach the unit note in the `direct_callers_note` style. Same remedy family as C1-37 and C1-32 |
593
703
  | C1-39 | **A `not_found` resolution publishes zeros where the product's own best pattern publishes null.** `ask plan "añadir autorización por endpoint a los controladores sin @M3FiltroSeguridad" .` returns a well-formed `change-plan-v1`: `resolution: "not_found"` **and** `affected_components.count: 0`, `affected_endpoints.count: 0`, `rollback_surface.file_count: 0`, `review_checklist: []`. It is not an error envelope — an agent reading `affected_endpoints.count` concludes *"no impact"* where the truth is *"not measured"*. `--agent` is a declared use case. Same shape elsewhere: on a repository with **0 test files for 3 337 non-test Java files**, every `tests_at_risk: 0` and `covering_tests: 0` reads as *"the change is safe"* when it means *"there are no tests"* | 4.10.4 (eval #14) | **Medium** — dangerous specifically for agent consumption, and it is the one place the product breaks its own best habit: `data-exposure` handles exactly this with `answered: false` and measures nothing rather than publishing a confident zero | **closed (4.10.6).** With an unresolved target every cost field is `null` and carries `answered: false` with the reason, plus a `how_to_read` that says why in one line — *a zero would say the change is free*. Same rule for the second half: on a repository with **no test source root**, `covering_tests` answers `null` / `answered: false` / *no test source root in this repository*, and the review checklist says the axis was not measured instead of reporting no covering tests. `compare` reads a 0 for the ranking (a total order needs a number) and its `why` trail carries the plan's own reason, so the placeholder is never narrated as a measurement. This is the shape `data-exposure` already had; it is now the same shape one level up. Original remedy note: with `resolution: "not_found"`, cost fields become `null`, never `0`. With no test source root, `tests_at_risk` and `covering_tests` become `null` with `reason: "no test source root"`. This is the C5 silent-failure shape, one level up: the envelope is valid, so nothing looks wrong |
@@ -46,7 +46,7 @@ CLI commands — impact, endpoints, spring-audit, explain, … each a pro
46
46
  The key idea: the extraction is **content-addressed**. Commands reuse the parse cache and,
47
47
  where their analysed scope matches, the shared Canonical IR; `ask cache model` names what a
48
48
  warm buys for each command rather than implying that every projection costs the same. In
49
- 5.8.18, `validation` enters through that shared CIR and `data-exposure` reuses one semantic
49
+ 5.8.20, `validation` enters through that shared CIR and `data-exposure` reuses one semantic
50
50
  model across all declared label seeds. (The extraction and consumption contract is fixed in
51
51
  the architecture ADRs 0001–0004 under `docs/architecture/`.)
52
52
 
@@ -142,7 +142,7 @@ pipx install sourcecode # isolated install, no venv needed
142
142
 
143
143
  # Verify
144
144
  ask version
145
- # ask 5.8.18
145
+ # ask 5.8.20
146
146
  ```
147
147
 
148
148
  Requires Python 3.9+.
@@ -1530,7 +1530,9 @@ repository's own artefacts, or `ASK_RUNS_DIR=/path/to/runs` to choose an externa
1530
1530
  location explicitly.
1531
1531
 
1532
1532
  **Large-repo budgets.** Set `ASK_MAX_ANALYSIS_SECONDS=<seconds>` to bound the
1533
- phase-runner commands `spring-audit`, `risk` and `audit-report`. It is not a
1533
+ commands that consume it — including the root repository-wide analysis. `ask
1534
+ --help` publishes the current set, generated from the same authority the
1535
+ budget-preamble emitter reads; this guide does not keep a second copy. It is not a
1534
1536
  process-wide timeout for every repo-wide or deep command. The budget is
1535
1537
  **honoured, not vetoed**: if the value is below the class floor ASK says so on
1536
1538
  stderr and runs anyway, and a run that spends its budget publishes what it
@@ -1554,9 +1556,9 @@ inventories `rules_run` / `rules_not_run` with `rules_run_count` and
1554
1556
  stops at 1 408 files inside the `SEC-004` group and ends at 5.1s.
1555
1557
 
1556
1558
  Every count in such an answer is a floor over what ran — a family or phase that
1557
- never ran is named, never reported as an absence of findings. Commands outside
1558
- that phase-runner set are not bounded by this variable and must not be CI-gated
1559
- on it. The mark lives
1559
+ never ran is named, never reported as an absence of findings. A command outside
1560
+ the set `ask --help` publishes does not consume this budget; bound it with
1561
+ `--progress` or `--detach` where they are offered. The mark lives
1560
1562
  where the numbers are as well as at the root: `summary.partial`,
1561
1563
  `summary.counts_are_floor` and `summary.counts_basis`, with
1562
1564
  `confidence_level` capped at `low` and `confidence_basis` saying the cut is the
@@ -0,0 +1,27 @@
1
+ """Which commands `ASK_MAX_ANALYSIS_SECONDS` bounds — the one authority for it.
2
+
3
+ `AUD-590-B04` extended the deadline to the root analysis and four more commands,
4
+ and `AUD-593-N03` closed the generated half of the help against that change. Three
5
+ hand-written sentences kept naming the original three (`AUD-594-X01`): the help's
6
+ own header paragraph in `cache_model`, and two passages of the packaged user guide
7
+ — one of which told the reader the root analysis *"must not be CI-gated"*, which is
8
+ the opposite of what the release beside it made possible.
9
+
10
+ The tuple lives here rather than in `cli` because `cache_model` renders the header
11
+ and `cli` renders the options line: a shared fact with two emitters needs a module
12
+ neither of them owns.
13
+ """
14
+ from __future__ import annotations
15
+
16
+ #: Every command whose run consumes `ASK_MAX_ANALYSIS_SECONDS`, in help order.
17
+ #: The budget-preamble emitter is keyed to the same set; the release battery
18
+ #: asserts that no shipped authority names a proper subset of it.
19
+ BUDGETED_ANALYSIS_COMMANDS: tuple[str, ...] = (
20
+ "ask (root)", "spring-audit", "verify --init", "verify", "risk",
21
+ "audit-report", "migrate-check", "migrate-apply",
22
+ )
23
+
24
+
25
+ def budgeted_analysis_commands_phrase() -> str:
26
+ """The command list as one sentence fragment, for any surface that names it."""
27
+ return ", ".join(BUDGETED_ANALYSIS_COMMANDS)