sourcecode 5.8.16__py3-none-any.whl → 5.8.18__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Potentially problematic release.


This version of sourcecode might be problematic. Click here for more details.

sourcecode/__init__.py CHANGED
@@ -4,4 +4,4 @@ ASK Engine is the product. ``ask`` is the canonical CLI command; ``sourcecode``
4
4
  the legacy compatibility alias and the Python/PyPI package name. See
5
5
  docs/PRODUCT_IDENTITY.md (normative)."""
6
6
 
7
- __version__ = "5.8.16"
7
+ __version__ = "5.8.18"
@@ -16,20 +16,73 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
16
16
 
17
17
  ## Current Synchronization
18
18
 
19
- **Latest attached audit of release `5.8.15`:** score **90/100**, coverage **54/62
20
- invocables (87%)**, byte-identical CIR across six versions, and no longitudinal functional
21
- regression. A new P1 affects dirty-tree `verify-edit`; BUG-1 is reproduced on the
22
- Windows/NTFS audit host. The historical `risk` P0 remains open because nine successful
23
- runs do not disprove two earlier exit-0/no-output failures. Only the sixth-pass section
24
- below declares present state; older audit rounds remain for traceability.
19
+ **Latest attached audits — release `5.8.18`, 2026-08-21, two independent rounds, and the
20
+ first pass since 5.8.2 that opens with a regression.** Round A (`saint-server`/MSAS,
21
+ 3 342 Java files, 17 commands, 29 invocations, same HEAD as the previous round bit for
22
+ bit) scores **8.77/10** — up from 8.21 — with **zero false statements about the
23
+ repository for the second consecutive run**, 6 of 6 reported defects closed and verified
24
+ from outside, and **five new contract defects, four of them the same class as the fix that
25
+ preceded them**. Round B (a 12-repository bank, ~150 invocations, the §7 measurement
26
+ protocol applied: idle host, interleaved, ×3, median, cache state declared per row) closes
27
+ **8 of the 10 new rows and 2 of the 9 inherited ones**, records **zero content regressions
28
+ and no performance row above the ×1.25 threshold**, and finds **two regressions, one of
29
+ them a P0**: `compare` is dead in every valid invocation with an unhandled `NameError`,
30
+ and `project_summary` publishes an Apache licence header as the description of the project
31
+ on 3 of 8 repositories.
32
+
33
+ **The two rounds agree on the shape of what is left, and it has moved again.** The
34
+ previous pass ended with contracts that were *absent*; this one ends with contracts that
35
+ *contradict each other across authorities*: two homonymous fields declaring the same unit
36
+ and the same population and differing 16×, one budget published three times with three
37
+ values, a help page denying a scope the fix beside it granted, and two ledger rows marked
38
+ `closed` that do not reproduce as closed in the field. **A row marked closed that is not
39
+ closed is worse than an open one — nobody looks at it again** — so those are carried at P1
40
+ below rather than left in their original sections.
41
+
42
+ **Previous attached audits — release `5.8.16`, 2026-08-21, two independent rounds.**
43
+ Round A (`saint-server`, 3 342 Java files, 17 commands, 25 invocations) scores **8.21/10**
44
+ with **zero false statements about the repository** and two contract defects. Round B (a
45
+ 16-repository bank, 49 to 24 073 Java files, **240 invocations, 0 crashes, 0 tracebacks**)
46
+ confirms **30 of 30 claims about source** — 19 exact to file and line, two of them where
47
+ ASK is right and a naive `grep` is wrong — and **refutes 8 of 8 claims ASK makes about its
48
+ own execution**. It scores **6.7/10** because one unauthorised write triggers its rubric's
49
+ elimination rule; its own counterfactual without that rule is 7.6, and it closes **8 of 18**
50
+ previously open defects by correction with **zero regressions among the inherited ones**.
51
+
52
+ Both rounds agree on where the product now fails, and it is not the Java/Spring analysis:
53
+ it is the contract an agent consumes programmatically — the write guard, repository
54
+ identity, list units, flag scope, and published budgets. Its queue is closed in `5.8.17`;
55
+ **the eighth-pass section below is the only current queue**, and older rounds remain for
56
+ traceability.
57
+
58
+ **Mechanisms confirmed in source during the seventh-pass intake** (not taken on the reports' word):
59
+ `verify --init` writes with no `readonly.guard` (`cli.py:10706-10708`) while five sibling
60
+ writers guard; `spring-audit` derives `repo_id` from the CIR content hash
61
+ (`spring_security_audit.py:1416`, `spring_tx_analyzer.py:1081`) while every other surface
62
+ derives it from the resolved path (`cache.repo_id`); `_resolve_target`'s last resort is an
63
+ unbounded case-insensitive substring match (`repository_ir.py:10205-10214`) under a constant
64
+ `matched_fqns_basis` string (`repository_ir.py:9788`); no payload publishes
65
+ `direct_callers_unit`; and the two `--compact` token figures differ in the same help page
66
+ (`cli.py:587` versus `cli.py:3613`).
67
+
68
+ **Two reported mechanisms are corrected here, and neither correction dismisses the defect.**
69
+ `--rule`/`--band` are not unwired: they are declared `--table`-only render filters and are
70
+ applied inside the table path alone, so the defect is that they are accepted and silently
71
+ ignored outside it and that `--rule` validates nothing (`--band` does, in `risk`).
72
+ `validation --path-prefix` is *worse* than reported: it filters `endpoints` and `gaps`
73
+ (`cli.py:8172-8180`) and leaves `summary` at its pre-filter values, so the answer contradicts
74
+ itself rather than merely under-declaring.
25
75
 
26
76
  **Release status:** entries marked “pending `5.8.14`” in their historical wording are
27
77
  shipped in `5.8.14`; the correction battery fixes listed above are included in `5.8.15`.
28
78
 
29
- **Current release:** `5.8.16` includes two post-audit corrections: compact cold-start
30
- identity is preserved for `no_ris`, and manifest-cache keys include Python requirements
31
- files. The testing-repository hygiene exercise changed no external repository; it
32
- completed 108/108 read-only workflow runs with valid JSON and exit code 0.
79
+ **Current release:** `5.8.18` closes the whole eighth-pass queue, one commit per row, and
80
+ includes the targeted release battery. `5.8.17` before it carried the seventh-pass queue;
81
+ `5.8.16` before that carried two post-audit
82
+ corrections: compact cold-start identity is preserved for `no_ris`, and manifest-cache keys
83
+ include Python requirements files. The testing-repository hygiene exercise changed no
84
+ external repository; it completed 108/108 read-only workflow runs with valid JSON and exit
85
+ code 0.
33
86
 
34
87
  The battery also closed the following reproducible contract defects in `5.8.15`:
35
88
  `migrate-apply` is registered across help, progress, path-admission, output-purity,
@@ -37,9 +90,181 @@ cache and schema authorities (`1d00ba2`); endpoint census output distinguishes e
37
90
  test modules from retained test-fixture routes (`f65336d`, `014dec7`, `1c6cd0b`);
38
91
  and compact posture output retains counts while omitting unresolved bodies (`dad5dad`).
39
92
 
40
- The sixth audit supersedes the previous queue where it has stronger evidence: BUG-3 is
41
- closed in stable-cache regime, BUG-1 is reproduced on the audit host, and BUG-4a is
42
- partial because `onboard` still has no machine-readable population block.
93
+ **Historical, sixth pass:** it superseded the queue before it where its evidence was
94
+ stronger — BUG-3 closed in stable-cache regime, BUG-1 reproduced on the audit host, BUG-4a
95
+ partial because `onboard` still had no machine-readable population block.
96
+
97
+ **What the seventh pass does not touch.** Neither round exercised `BUG-2` (`risk` exiting 0
98
+ with neither parseable stdout nor its requested artifact) or `BUG-6`
99
+ (`verify-edit.security_delta` missing access opening): both rows keep their status and their
100
+ queue position. `BUG-1` and `BUG-5` gain witnesses on public OSS repositories, recorded
101
+ under their own rows rather than as new ones.
102
+
103
+ ### Eighth Audit Pass: `5.8.17` findings, closed in `5.8.18` / MSAS + 12-repository bank / 2026-08-21
104
+
105
+ The current queue. Ordered as the correction order: **regressions first, then the rows a
106
+ previous pass marked closed that the field does not reproduce as closed, then new contract
107
+ defects, then residuals.** Canonical IDs are `AUD-592-*` (bank round, reporter IDs `A-*`,
108
+ `B-*`, `N-*`, `D-*`) and `AUD-593-*` (MSAS round, reporter IDs `N-0*`). Rows held elsewhere
109
+ in this ledger are listed after the table as mappings, not as new defects.
110
+
111
+ | ID | Severity | Current status | Required direction |
112
+ |---|---|---|---|
113
+ | `AUD-592-R01` / `A-1` | **P0 — regression** | **open, confirmed in source, and it is two defects from one error.** `compare` raises `NameError: name '_compare_scan_trunc' is not defined` on **5 of 5 valid invocations** — Rich traceback on stderr, zero bytes on stdout, no error envelope, exit 1 — ending a 240-invocation streak with no traceback. The assignment introduced by `AUD-591-A06` landed **inside `plan_cmd`** (`cli.py:8874`, beside `plan`'s own `find_java_files` at `:8873`) while the read is in `compare_cmd` (`cli.py:8972`); the two functions are `8818-8901` and `8903-9002`. Verified by reading the source at HEAD, not from the report. Second effect: `plan` computes its truncation state and **discards it** — nothing in the payload carries it, so `plan` answers `resolution: not_found` over a walk that stopped at its cap, offering candidates from a different repository. Third effect: `compare`'s other `AUD-591-A06` fix, `data.setdefault("scope", ...)` at `cli.py:8975`, is **after the line that raises** and has never executed, so the cross-repository candidate leak is still undeclared. | Read `_last_scan_truncation()` at `compare`'s own call site, immediately after `cli.py:8965`; rename the `plan_cmd` variable to `_plan_scan_trunc` and **bind it to the payload** — an unused walk-state variable is the defect, not the fix. Regression: an AST assertion that no `_*_scan_trunc` name is read in a `FunctionDef` that does not assign it, plus one execution test per command that owns a walk. A unit test that imported `compare_cmd` would not have caught this; one that *calls* it would. |
114
+ | `AUD-592-R02` / `N-4`, reopens `AUD-513-N04` | **P1 — regression** | **open, confirmed in source, on 3 of 8 repositories (37,5 %), in `--compact` and `--agent` alike.** `project_summary` publishes an Apache-2.0 licence header as the description of the project (struts), a 30-character horizontal rule (halo) and a setext underline (shopizer) — and `summary_basis` asserts `"README descriptive prose"` over all three. In 5.8.16 this field was `null` with `"no descriptive section found in README"`, the literal acceptance text of `AUD-513-N04`; the field went from honest to false. Mechanism, four separate holes, all read at HEAD: HTML comments are detected **only on their opening line** (`summarizer.py:335-337`) with no `in_html_comment` state — three lines above, `in_code_block` does exactly that for fences (`:326-330`) — so `<!---` on line 1 flushes and lines 2-15 of the licence become the first paragraph; headings are ATX-only (`:331`), so setext titles underlined with `===` are content; the short-fragment floor is `len(paragraph) < 30` **strictly** (`:359`), so a rule of exactly 30 dashes survives; `_LICENSE_MARKETING_RE` (`:266-281`) matches product-tier and marketing phrasing and **carries no Apache/MIT/GPL/BSD boilerplate pattern at all**; and `summary_basis` is assigned unconditionally to whatever paragraph survives the filters (`:382-385`) — a restatement of the code path, not a property that was checked. | Add `in_html_comment` symmetric to `in_code_block`; recognise setext underlines as headings, not content; extend the licence regex with the four standard boilerplates; and make `summary_basis` describe the path actually taken — where nothing is verifiably descriptive, `null` plus the 5.8.14 text. **The mechanical cause of the reopening is that no README fixtures exist**: add the three witnesses from the bank as fixtures (HTML licence comment, 30-character rule, setext `===`) in the same commit as the fix. |
115
+ | `AUD-593-N01` / `N-01` | **P1** | **open, and it falsifies the closure assertion of `AUD-590-B02`.** For one symbol, one tree, one HEAD, consecutive invocations, both at `depth=4`: `impact` publishes `direct_caller_count: 34` (210 reference sites) and `impact-chain` publishes `554` (3 103), both under the literal unit `"distinct caller classes"`. The product's own comparability rule (`CALLER_METRIC_RECONCILIATION`) is *"two figures are comparable only when both match"* — they match, and differ 16×. **The authority is what is wrong, not only the number**: `FAN_IN_FIGURES` (`caller_metrics.py:80-110`) declares both rows `unit=distinct_classes, population=references_excluding_imports`, and the populations are not the same. Measured mechanism: `impact-chain` expands the seed set with every member of the interfaces the target implements (CH-001b, `spring_impact.py:1110-1135`) and then takes depth 1 **from the expanded seeds** (`caller_reach.py:309-336`), so its "direct" includes callers that name a sibling implementation; `impact` admits only references that name the target and keeps interface-mediated callers on a separate axis (`repository_ir.py:9263`, `:9596-9606`). Corroborating symptom: `chain_classes_total`, declared *"unbounded above by any direct-caller count"*, equals `direct_caller_count` exactly (554), and `indirect_callers` is `[]` — that emptiness is the hub guard capping the walk to depth 1 (`caller_reach.py:319-321`) and **is** declared, so it is not a confident zero. The published closure assertion `len(impact.direct_callers) == chain.metadata.direct_caller_count` measures 30 (list, capped from 34) against 554: it passes on a fixture where the two coincide by accident. | Decide which figure the `direct_caller_*` namespace owns, and fix the **catalogue** either way: either `impact-chain` measures depth-1 fan-in on the unexpanded seed (≈34), or the figure is renamed out of that namespace and its `population` becomes its own tag — an interface-expanded seed set is not `references_excluding_imports`. Regression is parametric over the catalogue, not over `CommonService`: for every pair of `FAN_IN_FIGURES` rows sharing `(unit, population)`, the emitted values must be equal, run over ≥3 symbols of different fan-in **including one that trips the hub guard** — the condition under which the divergence appears, and the reason two previous corrections of this class (`callers_total`, `P0-2`) closed a name and left the divergence. `chain_classes_total ≥ direct_caller_count` with strict inequality wherever uncapped transitive reach exists. |
116
+ | `AUD-592-B01` / `B-8` | **P1 — closed row that does not reproduce as closed** | **open; the ledger row overstates its own fix.** `B-8` is recorded as *"compact posture retains one unresolved sample when the total is non-zero"* (`52593b6`). The code retains a sample only when the total is **exactly one**: `_posture_limit = 1 if summary["unresolved"] == 1 else 0` (`cli.py:12653-12654`). Field measurement on `mall`: `summary.unresolved: 3`, `unresolved: []`, `unresolved_cap: {total: 3, shown: 0, omitted: 3, limit: 0}` — and it is the only cap in that payload with `limit: 0` while `undecided_cap` and `evidence_cap` carry 200. `--limit 10` returns all three, so the analysis is intact and only the compact contract is wrong. | Either make the code do what the row says (retain ≥1 sample whenever `summary.unresolved > 0`) or correct the row to the narrower promise and reopen the defect it leaves. Do not leave `closed` standing over `shown: 0`. Add an assertion that no `cap_effect` receives `limit=0` by default on a path where `--limit 0` is published as *"no cap"* (`AUD-513-N08`), so the published semantics and the internal default cannot invert each other. |
117
+ | `AUD-592-A02` / `A-1` second half, `§12.5` | **P2** | **open — scan truncation is declared by 1 of 4 commands that own a walk.** Over a truncated scan (cap observed at **8 000 files**), `impact` publishes `scan_truncated`; `plan` and `impact-chain` answer `not_found` with `scan_truncated: null`, and `endpoints` exits 1 with no payload — every one of them an exclusion derived from an incomplete population. `compare` cannot be measured (`AUD-592-R01`). Second, narrower defect in the same row: the root `--help` describes the cap as *"the scan stopped at 25 000 directory entries"* — a different limit in a different unit from the 8 000-file cap the field observes. | Publish the truncation beside every walk that can end in `not_found`, `candidates` or a census — the seventh pass already recorded that `extract_java_endpoints` does its own uncapped `rglob`, so this must be read at each call site, never centrally. Reconcile the two caps in the help text or declare both with their units. |
118
+ | `AUD-593-N02` / `N-02` | **P2** | **open — the `--compact` budget is published three times, in two shipped authorities, with three values.** `ask --help` generates *"~7K tokens (30 KB, chars/4)"* from `output_budget`, the module that enforces the limit — this half is the `AUD-590-B03` fix and it is correct. The packaged `docs/USER_GUIDE.md` is still hand-maintained and disagrees with it twice: `:111` *"A ~10K-token subset"* and `:243` / `:1366` *"~2,500–4,000 tokens"*; `--agent` carries a fourth figure at `:254` / `:1367` (*4,500–5,500*). Measured payload: 22 390 B ≈ 5 597 est. tokens — inside the help ceiling, **+40 % over the guide's table**. `grep '7K tokens' *.py` finds no literal: the help is generated and the guide is not, and that asymmetry is the defect. | Generate every token figure in the packaged docs from `output_budget.token_estimate`, or delete them and point at `--help`. Extend the existing release assertion (today help ↔ payload ↔ section registry) to cover `docs/*.md`: a token band in a shipped `.md` that does not come from that module fails the battery. Packaged documentation is a published authority and belongs inside the same assertion as the help text. |
119
+ | `AUD-593-N03` / `N-03` | **P2** | **open — `--help` denies the scope that `AUD-590-B04` had just granted.** With `ASK_MAX_ANALYSIS_SECONDS=90` the root analysis prints *"ask is a repo-wide analysis and the configured budget is 90s. Budget read from ASK_MAX_ANALYSIS_SECONDS=90 (process environment)"* — the fix works. The help page beside it still reads *"deadline only for spring-audit, risk, audit-report"* and *"Other commands are not time-bounded by this variable"* (`cli.py:621-622`). The root analysis is the most expensive command in the product and the one an agent most needs to bound; a reader of the help will not set the variable there. | Generate that line from the same registry `_build_analysis_classes` populates (the root is keyed there as `B14`), never by hand. Assertion: the set of commands the help names equals the set the preamble emitter covers. |
120
+ | `AUD-593-N05` / `N-05` | **P3** | **open — a figure emitted outside the catalogue that forbids it.** `via_interface_resolution[].caller_count` publishes 37 where every other surface of the same fact says 23 (`interface_mediated_caller_count`, the list length, and the `explanation` prose), and it carries no `_unit` while the rest of the payload does. Both are legitimate and neither is declared: `caller_count` is `len(_iface_callers)` — raw caller symbols of the interface, before the `not in all_affected` filter and before normalisation to classes (`repository_ir.py:9248-9251`); the 23 is `caller_classes(_iface_mediated_callers, ...)` (`:9603`, `:10173`). It is also **not in `FAN_IN_FIGURES`**, which states that a figure absent from it may not be emitted — so `tests/test_fan_in_authority.py` did not catch a fan-in figure added in this release. | Give it a unit and a `FAN_IN_FIGURES` row, or rename it to `interface_caller_symbols`. Make the authority test sweep the **emitted** keys against the catalogue rather than the catalogue against itself; a catalogue that only validates its own rows cannot see an unregistered emission. |
121
+ | `AUD-593-N06` / `N-06` | **P3** | **open, latent — a silent `[:30]` where every sibling list declares its cap.** `out["interface_mediated_callers"] = _iface_mediated_classes[:30]` (`repository_ir.py:10172`) with no `*_cap` block, in a payload where `direct_callers_cap`, `indirect_callers_cap` and `security_surface_affected_omitted` all carry `total`/`shown`/`omitted`/`direction`/`how_to_read`. On MSAS the list is 23, so nothing is visible; the first repository with more than 30 interface-mediated callers loses the excess without a word. | Route it through the same `_cap_effect(...)` used 27 lines above (`repository_ir.py:10145`). |
122
+ | `AUD-592-D01` / `D-1` | **P3 — residual of `AUD-591-A08`** | **open, confirmed in source.** The write inventory in the root help says `migrate-check` writes *"only with `--history-dir`"* (`cli.py:296`). `--history-dir` only relocates the destination; the flag that writes is `--snapshot`, as `migrate-check --help` states correctly and as the code comment at `cli.py:12980-12986` states explicitly. Verified in the field: `migrate-check --history-dir <path>` exits 0 with zero writes and no guard message, while `--snapshot` fires the guard correctly. `AUD-591-A08`'s assertion checks that each cited flag **exists** in that command's parser — `--history-dir` does — not that it is the flag that triggers the write. | Key the inventory on `(command, triggering flag, path)` taken from the guard site itself, as `AUD-591-A08`'s own acceptance text required. An auditor tests what the inventory lists: this is the same mechanism by which `AUD-591-A01` went unfound. |
123
+ | `AUD-592-D02` / `D-2` | **P3 — residual of `AUD-591-A02`** | **open.** `--min-band` filters and validates (`cli.py:11294-11308` in `risk`, `:11563-11568` in `enrich`) but publishes no `_filter` block, while its sibling `--band` publishes one through `_apply_selection`. A consumer cannot tell a filtered answer from a small one. | Route `--min-band` through `_apply_selection` like every other selection flag, or publish the same `_filter` shape from wherever it is applied. |
124
+ | `AUD-592-D03` / `D-3` | **P3 — residual of `AUD-591-A02`** | **open.** `enrich --rule NOPE-999` exits 0 with an empty result; `spring-audit --rule NOPE-999` exits 1 against the rule catalogue. One flag name, two contracts. Measured with a synthetic SARIF 2.1.0 (2 results): `--rule`, `--band` and an invalid `--band` all behave correctly in `enrich`; only unknown-rule validation is missing. | Validate against the same catalogue in both commands. `enrich` is experimental-tier, which sets the severity, not the contract. |
125
+ | `AUD-592-D04` / `build_commit` | **P3 — process** | **open, fourth consecutive report.** `ask version` publishes `build_commit: unrecorded`, so no measurement in these ~150 invocations is attributable to a commit and *"fixed in 5.8.17"* is not falsifiable from outside. | Stamp the commit at build time. It is the cheapest row in this table and it is the precondition for every other row's verification. |
126
+
127
+ **Mapped, not new — inherited rows this pass re-witnesses.** `B-6` (partial answers exit 0)
128
+ keeps its row and gains its best measurement yet: overrun ×5.4 → ×4.2 → **×1.39** (median
129
+ 20.90 s of 70.74/20.90/19.70 against a 15 s budget), with a `_partial` block that is
130
+ exemplary (`phases_completed`, `phases_pending`, `why_stopped`, and *"every count here is a
131
+ FLOOR"*) — and `exit 0` on all three runs, so a pipeline reading `$?` passes with two of
132
+ three phases unexecuted. **`AUD-591-A09` is the template**: `exit_code` / `gate_exit_code` /
133
+ `gate_exit_code_basis`, or an explicit `--allow-partial`. `B-3` (`--agent` versus
134
+ `--compact`) gains a witness at 3 342 files: 35 357 B ≈ 8 839 est. tokens, **+61 % over the
135
+ only published band**, 16 blocks present in `--compact` and absent from `--agent`
136
+ (`transactional_boundaries`, `mybatis`, `spring_profiles`, `deployment`, `env_map`,
137
+ `stacks` among them), and `sibling_view` still names no block; on the bank round the two
138
+ views are no longer a subset relation at all (9 blocks agent-only, 15 compact-only) and the
139
+ ×2.25 time overhead is gone (×1.02). Substantively, `security_surface` is absent from
140
+ **both** channels on a repository with 368 security findings and 725 routes open by rule —
141
+ the predicate is disclosed under `AUD-590-B03`, not closed. `N-5` (an advisory saying
142
+ *"8 372 Java files"* with no population label while the progress line says
143
+ `[population=production_java_sources]`) remains the unclosed half of `AUD-588-B11`; the
144
+ label exists (`cli.py:10278-10280`) and reaches the progress lines and the payload but not
145
+ the advisory emitted from `phased_run.py:367`. `D-09` (cost anchors stamped *"on 5.1.0"* in
146
+ a 5.8.17 build — now counted: **17 anchors on 5.1.0, one each on 4.18.0, 5.0.0 and 3.2.2**,
147
+ in a `cache model` output presented as operational guidance) belongs to `BUG-5`. `N-6`
148
+ (cold `spring-audit` on `openmrs-core`) belongs to `BUG-1` and **improved**: 8.86 s → 8.06 s.
149
+
150
+ **Closed by correction and re-verified from outside in this pass — do not re-derive.**
151
+ `AUD-591-A01` (8 of 8 write paths refuse, verified by filesystem footprint rather than
152
+ `git status`, `ASK_RUNS_IN_REPO=1` included), `AUD-591-A02`, `AUD-591-A03` (and the
153
+ proposed highest-priority probe is retired: `rename-class --from O` exits 1, as the seventh
154
+ pass had already established), `AUD-591-A05`, `AUD-591-A06` (50.30 s → 0.43 s, and the
155
+ sweep it asked for holds: **6 of 6 error paths ≤0.54 s**), `AUD-591-A08` (modulo
156
+ `AUD-592-D01`), `AUD-591-A09`, `AUD-591-A10`, `AUD-590-B01` (`repo_id` cardinality 1 with
157
+ `repo_id_basis`; the content hash lives on as `content_id`), `AUD-590-B02` (the *lists* —
158
+ the counters above it are `AUD-593-N01`), `AUD-590-B03` (in `--help`; the packaged guide is
159
+ `AUD-593-N02`), `AUD-590-B04` (both halves: a `partial-analysis-v1` stub with `partial:
160
+ true` where 5.8.16 left zero bytes, and an 807 B preamble byte-compatible with
161
+ `spring-audit`'s 813 B), `AUD-590-R01` (`route_census` — and it **refuted the auditor's own
162
+ double-counting hypothesis with the product's data**: `mappings 3763` against
163
+ `on_rows 1120`, with `distinct_routes ≥ mappings`), `AUD-591-Q01`, `B-7` (closed by
164
+ deletion, which was the right answer: `total_ms` is gone, replaced by
165
+ `wall_ms`/`measured_ms`/`unaccounted_ms`/`unaccounted_pct` with arithmetic exact to the
166
+ decimal and a `unit` explaining why the remainder is published rather than distributed).
167
+
168
+ **Verified and not defects — do not spend effort.** `posture --profile` accepting an
169
+ undeclared profile is **by design and correctly declared**: it is a hypothesis input, not a
170
+ catalogue name, and the payload publishes `profiles_requested` beside `profiles_declared`
171
+ (the bank round withdrew this finding itself after reading the payload). `pr-impact` on a
172
+ clean tree exits 1 with a typed `INVALID_INPUT` naming the empty diff. `pr-impact` over
173
+ three classes refuses with `OUTPUT_TOO_LARGE` and names `--output` — a ceiling as contract,
174
+ not a truncation. `verify` publishes `exit_code: 2` matching the process (`AUD-591-A09`).
175
+ `data-exposure` still answers `answered: false` rather than inferring labels. Cache
176
+ freshness reporting `STALE` at delta 0 commits is `Uncommitted: True` over the auditor's own
177
+ untracked files. `repo_id` absent from `endpoints`, `posture` and `--compact` is coverage,
178
+ not a defect — the criterion was cardinality 1 among emitters, and the help no longer
179
+ promises it.
180
+
181
+ **Process findings of this pass, and they are the same finding twice.**
182
+
183
+ 1. **The class survives the fix, four times out of five.** `AUD-590-B02` (unit absent from
184
+ the lists) returned as `AUD-593-N01` (unit present and false in the counters) and
185
+ `AUD-593-N05` (a new field shipped without one); `AUD-590-B03` (two figures in one help
186
+ page) returned as `AUD-593-N02` (three figures once the packaged guide is counted) and
187
+ as the `--agent` band in `B-3`; `AUD-590-B04` (a variable invisible to the root) returned
188
+ as `AUD-593-N03` (a help page denying the scope the fix granted). The seventh pass named
189
+ this pattern its highest-value finding and wrote the rule — *every acceptance criterion
190
+ carrying a sweep clause is closed with a parametric test over the catalogue* — and the
191
+ two classes closed **as sweeps** in that pass (write guards 8/8, numeric bounds 7/7, and
192
+ now error-path cost 6/6) did not recur, while three of three closed by instance did.
193
+ `FAN_IN_FIGURES`, `COMMANDS_THAT_WRITE` and the serializer's section registry are already
194
+ catalogues: the test goes over the catalogue, never over the reported symbol.
195
+ 2. **Two rows in this ledger were marked `closed` while the field reproduces them open**
196
+ (`AUD-592-B01`, and `AUD-588-B11`'s advisory half). Both were closed against a surface
197
+ narrower than the row's own wording. A closure whose assertion cannot fail on the tree
198
+ that produced the defect is not a closure: `AUD-593-N01`'s assertion passes today only
199
+ because the fixture makes both figures coincide.
200
+ 3. **Packaged documentation is a published authority.** `AUD-593-N02` exists because the
201
+ release assertion covers help ↔ payload and stops at the package boundary. Two stale
202
+ figures in a shipped `.md` are enough to contradict the correct one in the help.
203
+
204
+ ### Seventh Audit Pass: `5.8.16` / 16-repository bank + `saint-server` / 2026-08-21
205
+
206
+ Ordered as the correction queue, regressions and permissive-direction failures first. Each
207
+ row carries the canonical ledger ID; the reporters' own IDs are given for traceability
208
+ (`A-*` and `B-*`/`N-*` from the bank round, `F-*`/`B-0*`/`R-01` from the `saint-server`
209
+ round). Rows already held elsewhere in this ledger are listed at the end as mappings, not
210
+ as new defects.
211
+
212
+ | ID | Severity | Current status | Required direction |
213
+ |---|---|---|---|
214
+ | `AUD-591-A01` / `A-1` | P0 | **closed 5.8.17**: Guarded at the emitter. The sweep the row asked for found **two more** unguarded paths a per-module audit could not see: `migrate-recipe --write` (writes `<repo>/rewrite.yml` from the command body) and `verify-edit --install-hook`, whose guard sat on the *hooks directory* — and `guard()` never refuses a path that already exists, so `git init` made that guard unfireable on every repository; it is `guard_mutation` on the hook file now. The proposed criterion `git status --porcelain == ""` **cannot see this defect**: the first thing written into `.ask/` is a `.gitignore` containing `*`, so every artefact after it is invisible to git. The sweep compares a filesystem footprint of the whole tree including `.git/hooks/`, both halves, over all ten write combinations, `--dir`/`--history-dir` pointed inside the repository. Previously — **open, reproduced in source**: `verify <repo> --init` writes `<repo>/.ask/contracts.yml` and exits 0 under `ASK_READONLY=1` and under `--no-write`. Six of seven write paths guard; this one calls `target.write_text` directly (`cli.py:10706-10708`). Second instance in two releases of the class `AUD-513-N01` closed. | Put the guard on the emitter, not the call site: `written_to != null` and an active guard are mutually exclusive states. Then sweep — for every command and flag combination that can produce a non-null `written_to`, a subprocess test with a clean git tree asserting `git status --porcelain == ""`. |
215
+ | `AUD-590-R01` / `R-01` | P0 (triage) | **closed 5.8.17**: Triaged and both halves answered. `no_security_signal` and `undocumented` are **one number twice** — the same variable assigned to both keys, in both endpoint authorities — and the payload now says the second is an alias and never an independent measure. `total` is decomposed by `route_census`: `distinct_routes` (distinct `(effective_path, method)`, every verb in `methods`) and `expansion_rows`, from one helper both authorities call. Measured on shenyu: 359 mappings, 235 distinct routes, 6 expansion rows. The triage also found the census **under**-counting: `_parse_route_path` returned the first literal only, so `@GetMapping({"/a","/b"})` published `/a` and dropped `/b` while the class-level array had always expanded. shenyu 359 → 365 published routes. Previously — **open, unexplained movement**: the endpoint census on `saint-server` moved 2 635 → 3 742 (+42%) between rounds at the same HEAD, published as `total_unit: "handler mappings declared … the whole route population"`. The reporter did not record the earlier version, so this is not yet attributable to `5.8.16`. | Decide between annotation expansion counted as population (multi-path or multi-verb mappings, class-level mappings counted beside method-level) and a real widening of detection. Compare `total` against distinct `(effective_path, method)` on the same tree. If it is expansion, it contaminates `data-exposure` and `pr-impact` gating; if it is detection, it needs a change note. Also verify that `no_security_signal` and `undocumented` — both 718 — are computed separately and are not one number under two names. |
216
+ | `AUD-591-A03` / `A-3` | P1 | **closed 5.8.17**: `partial` publishes what it is: a basis naming the substring match, `risk_score`/`risk_level` null, `risk_reason` and `candidates_total`, and prose that no longer opens with a verdict the payload withheld. The `not_found` half was the same defect with the opposite sign — `risk_score: 0.0` over a symbol never measured — and is null with its basis on both `not_found` and `ambiguous_path`. Resolution policy published in `impact --help`. Release assertion over every `*_basis`, derived from source: each must explain a field some payload publishes, a basis describing a null must say null, and the two paths producing a null score must not share one string. **The report is wrong on one point:** `rename-class --from` does not share this resolver — it requires a file declaring the class and refuses `--from O` outright, so the one path that rewrites source was never reachable from the substring match. Previously — **open, reproduced in source**: `_resolve_target`'s last resort is an unbounded case-insensitive substring match (`repository_ir.py:10205-10214`), so `ask impact O` answers over 129 symbols with `exit 0`, `risk_level: critical` and `matched_fqns_basis: "symbols resolved directly from the requested target"` — a constant string (`repository_ir.py:9788`) emitted on a path that resolved nothing. `PetControler` and `OwnerRepositor` are the same class of input and get opposite verdicts. Same class as `E-17` and as `B-5`, both closed. | Either `not_found` with the `candidates` list that already works, or a basis that names the substring match with `risk_level: null` and a `risk_reason`. Publish the resolution policy in `impact --help` with its minimum length and specificity. Add the release assertion that every `*_basis` describes the path actually taken. The shared resolver reaches `impact-chain`, `explain`, `plan`, `compare`, `fix-bug` and `rename-class --from`; `rename-class --from O` must be exercised on a disposable fixture before anything else. |
217
+ | `AUD-591-A10` / `A-10` | P1 | **closed 5.8.17**: Route-derived summary counts (`endpoints_with_body`, `validated_fields`, `gaps`, `endpoints_with_declared_constraints`, `declared_constraint_routes`) move with the filtered lists; `_filter` carries `total_before_filter` and `summary_before_filter`. `source_derived_routes`, `body_endpoints_in_code` and the validator catalogue stay whole and the block says why — recomputing them from the filtered list would replace an unscoped number with a wrong one. Both commands reject a prefix that is not a route path, echo the value that arrived, and name the MSYS rewrite with its two workarounds. The prefix matches `effective_path` as well as `path`, so a deployment prefix no longer empties every selection. Previously — **open, worse than reported**: `validation --path-prefix` filters `endpoints` and `gaps` and leaves `summary` at its whole-repository values (`cli.py:8172-8180`), and publishes no `_filter` block at all. `endpoints --path-prefix` filters and echoes correctly but validates nothing, so a Git Bash/MSYS path rewrite turns `/owners` into `C:/Program Files/Git/owners` and the answer is `total: 0` with `exit 0` — a silent false negative on a core-tier command in the audit host's own shell. | Reject a prefix that is not a route path, echoing the value received and naming the MSYS rewrite only when the value looks like an absolute Windows path. Recompute or scope `summary` under a filter, and reuse the `endpoints._filter` block (`path_prefix`, `total_before_filter`, `note`) rather than reimplementing it. The regression must run in Git Bash: a Linux-only test cannot see this. |
218
+ | `AUD-590-B01` / `B-01`, `F-01` | P1 | **closed 5.8.17**: One `repo_id` per tree, equal to the cache directory key, with `repo_id_basis`. The CIR fingerprint keeps its own name (`content_id`). `TransactionBoundaryIndex.repo_id` was a third site holding the content hash and is `content_id` too. Regression asserts cardinality 1 across the emitters **and** equality with the cache key — not merely that they agree with each other, but that they agree with the thing they name. Previously — **open, mechanism identified in source**: one field name, two identity functions. `spring-audit` sets `repo_id` from the CIR content hash (`spring_security_audit.py:1416`, `spring_tx_analyzer.py:1081`), while `migrate-check`, the cache directories and every other surface use the path hash (`cache.repo_id`). Same tree, same HEAD, same argv: `0123aa26a26297da` versus `7fced5c877cacfc5`. | One identity per tree, equal to the key of `~/.sourcecode/cache/<hash>` and `~/.sourcecode/context-cache/<hash>`; if a content hash is wanted it needs its own field name. Regression: invoke every command that emits `repo_id` on one path and assert cardinality 1. `audit-report` and `regress` correlate on this field and drop evidence silently when it diverges. |
219
+ | `AUD-591-A02` / `A-2`, `A-4` | P1 | **closed 5.8.17**: One `_apply_selection` helper on both sides of the format branch in `spring-audit`, `risk`, `enrich` and `migrate-check`, publishing `_filter` with `total_before_filter`/`matched`/`shown`; no selection, no block. `--rule` is rejected against the catalogue the analysis iterates (`rule_catalog.ids()`, and a new `migration_rule_ids()` read from `_ALL_RULES`), before the analysis runs. `enrich` is the declared exception: its rule ids come from the scanner's SARIF, so there is nothing to validate against and its help says so. One parametric battery, three commands × valid-with-matches / valid-without / invalid. Previously — **open, reported mechanism corrected**: `--rule` and `--band` are declared `--table`-only render filters and are applied only inside the table path (`cli.py:10992`, `11204`, `12584`, `10343`). In JSON they are accepted, ignored and never echoed, so `spring-audit --rule SEC-008`, `--rule NOPE-999` and no flag produce byte-identical payloads with `total_findings: 86`. `--rule` validates nothing anywhere; `--band` does validate in `risk` (`cli.py:10932-10943`); `--min-band` applies to the payload while its sibling `--band` does not. | A selection flag either selects or refuses. Outside `--table`, apply it or reject it with `flag`/`value`/`valid_values` taken from the rule catalogue, and echo the selection in a `_filter` block with `total_before_filter`. One parametric test over the selection-flag catalogue, three cases each — valid with matches, valid without matches, invalid — rather than two patches. Barrido: `--rule` also exists on `risk` and `enrich`. |
220
+ | `AUD-590-B02` / `B-02`, `F-02` | P2 | **closed 5.8.17**: Both lists publish `direct_callers_unit`, and each unit names where the other command keeps the same figure. `impact-chain` gained `indirect_callers_unit`. `resolution` is **not** merged — the commands resolve different things — so each publishes `resolution_vocabulary` with its values and a note mapping them onto the other's. Assertion: `len(impact.direct_callers) == chain.metadata.direct_caller_count`, so the declared units account for the cardinality difference. Previously — **open**: `direct_callers` names caller *classes* in `impact` and caller *methods* in `impact-chain` — two core-tier commands, one field name, aggregable by a consumer, and no payload publishes `direct_callers_unit`. `impact` publishes `indirect_callers_unit` (`repository_ir.py:9842`) and `impact-chain` publishes `direct_caller_count_unit` for the count but not for the list. Adjacent: the same symbol resolves `exact` in one and `class_expanded` in the other. | Declare the unit on both lists, reconcile the `resolution` vocabulary across the two resolvers, and assert in one test that the declared units explain the cardinality difference. `impact.implementation_fqns: []` beside an explanation naming 23 interface-reached callers needs either a `_basis` or a fix. |
221
+ | `AUD-590-B04` / `B-04`, `F-04` | P2 | **closed 5.8.17**: The root analysis was the single row `_build_analysis_classes` skipped, so its budget fell to an unnamed default and `ASK_MAX_ANALYSIS_SECONDS` was invisible on it; keyed now under the name the envelope gives it (B14) and calling the same preamble emitter `spring-audit` uses. `--output` is created before the analysis with a `partial-analysis-v1` stub (`partial: true`, `status: running`, and a reason saying it is a trace and not an answer), replaced atomically on success, guarded by `--no-write` and best-effort throughout. Previously — **open**: the one command `cache model` classifies as *"not a foreground run here"* is the one that emits nothing while it runs. `ask . --agent -o file` was killed at 140 s having written zero bytes to stdout, stderr and the output file, while `spring-audit` prints a budget preamble before starting and `ASK_MAX_ANALYSIS_SECONDS` explicitly does not bound the root command. | Reuse the `spring-audit` preamble emitter for every command whose `cache model` classification is not-a-foreground-run — the classification is already a product fact. Flush it to stderr so a killed run leaves a trace, and create the `-o` file at start or write a `partial: true` stub on deadline. |
222
+ | `AUD-590-B03` / `B-03`, `F-03` | P2 | **closed 5.8.17**: One phrase from `output_budget`, the module that enforces the limit, converted by the `chars/4` model `token_estimate` declares — a ceiling, not a typical size. Help field names equal payload keys, checked against the payload plus the serializer's section registry. **The predicate question is answered and the answer is the feared one:** `security_surface` needs a custom annotation that *names a resource*, while `GATE-001..004` fire on any custom authorization mechanism, so a gate taking no resource argument produces audit findings and no section. The section carries a `basis` naming the wider predicate. Previously — **open**: `--compact` publishes two different budgets in one help page — `~2,500–4,000 tokens` (`cli.py:587`) and `typically 1000–3000 tokens` (`cli.py:3613`) — and measures 22 390 B ≈ 5 597 tokens under the product's own `UTF-8 bytes ÷ 4` estimator, over both. The same help promises `confidence` and `gaps`; the payload has `confidence_summary` and `analysis_gaps`, and no `repo_id` at all. `security_surface` is absent on a repository where `spring-audit` reports 368 security findings including the GATE-* family. | One figure, derived from the same estimator or expressed as a function of repository size, and help field names equal to payload keys. Check whether the predicate that emits `security_surface` in `--compact` is the one that fires the `GATE-*` rules; if they diverge, the agent channel is hiding security surface. |
223
+ | `AUD-591-A06` / `A-6`, `A-7` | P2 | **closed 5.8.17**: `build_repo_ir(...)` was an argument expression, evaluated to be passed into a function whose first act is a syntactic parse: the repository was analysed so a regex could reject a string. `parse_query` runs before the repository is touched — **50.3 s → 0.31 s measured** — and a syntactic failure builds no `near_matches`. `find_java_files` records its own truncation (thread-local, because `--jobs` runs walks concurrently) and `last_scan_truncation()` is read *at the call site*, beside the walk that produced the list: read anywhere else it is a previous walk's state, and `extract_java_endpoints` does its own uncapped walk — a test holds that its census declares nothing. Release ceiling: four malformed routes rejected under 10 s each. `compare` also publishes the tree its candidates came from. Previously — **open**: the error path costs more than the success path. `explain-endpoint ./repo` — a purely syntactic rejection, *"a route path starts with `/`"* — takes 50.3 s against 0.67 s for a valid route, on the 49-file repository, and returns `near_matches: []` after building them. `compare` without `--path` spends 62–111 s scanning the CWD, offers closest candidates from a *different* repository under it, declares `not_found` for a symbol that is in the analysed tree, and never publishes that the scan stopped at 25 000 directory entries — a truncation the root `--help` already reports. | Validate the shape of a typed positional before touching the repository, and do not build suggestions for a syntactic failure. Publish `scan_truncated` with its limit and unit wherever a `not_found`, a `candidates` list or a census is derived from a truncated scan; warn on stderr before `compare` scans a CWD holding several repositories. Release assertion: no `INVALID_INPUT` rejection costs more than a small fixed ceiling on the reference repository. |
224
+ | `AUD-591-A09` / `A-9` | P3 | **closed 5.8.17**: `exit_code` is what the process returns; the gate verdict keeps `gate_exit_code` with its basis. Asserted against the observed return code on both `--ci` and `--no-ci`. Previously — **open**: `verify --no-ci` publishes `exit_code: 2` in a payload while the process exits 0. A pipeline reading the field rather than `$?` inverts the decision `--no-ci` just took. | Make the field describe the process exit, or rename it `would_exit_code` and say that `--no-ci` suppresses it. Adjacent, not a defect: `--fail-on never` exiting 2 comes from the `unverified` axis, not the violations axis, and the payload names `--allow-unverified` correctly — only the help example's comment oversells it. |
225
+ | `AUD-591-A08` / `A-8` | P3 | **closed 5.8.17**: Rows corrected (`--history-dir`, `--capture-baseline`) and three writing commands added, `verify --init` among them. Every backticked flag in the inventory is resolved against that command's own click parameters, so the table cannot cite an option the parser rejects. Previously — **open, and the reason `AUD-591-A01` was not found earlier**: the root help's *"commands that modify files inside the repository"* table cites `verify --update-baseline` (the real flags are `--baseline` and `--capture-baseline`) and `migrate-check --history` (the real flag is `--history-dir`), and omits `verify --init`, which is the one that writes without a guard. An auditor probes what that inventory lists. | Correct both lines and add `--init`. Then generate the inventory from the same registry the write guard consults — while they are two hand-kept sources they will diverge again — and extend the existing backtick assertion (closed for payload `message`/`hint` under `N-7`) to the help text. |
226
+ | `AUD-591-A05` / `A-5` | P3 | **closed 5.8.17**: The probe answers per layer (`_cache_probe`); `Scope` carries `snapshot_warm` and `ris_warm`; the advisory names each with the state actually probed for it. `warm` survives as their OR, for the coarse question it answers. Finishes `N-3`'s parity criterion. Previously — **open, `N-3` half-corrected**: the advisory no longer collapses four cache layers into one boolean and honestly says context/parse are unknown before analysis — but it asserts *"snapshot/RIS present"* while `metadata.cache_layers` in the same run publishes `snapshot: "cold"`, and `cache freshness` reports `RIS HEAD: (none), STALE` while the payload calls RIS `warm`. Three authorities, one fact. | The advisory prints what the payload will publish or says `unknown before analysis` for every layer it cannot know, `snapshot` included. Parity test across stderr, `metadata.cache_layers` and `cache freshness` in a single invocation. |
227
+ | `AUD-591-Q01` / `Q-1` | Question | **answered 5.8.17**: Reachable, and only from the root analysis: `ask <repo>` and `ask cache warm` write `core-*.json.gz`; no other command does. All 240 field invocations were of other commands, so `cold` was correct on every one of them and said nothing about the cache. `Layer` gained `written_by`, `cache model` prints it, and a test runs `spring-audit` twice from a cleared cache and asserts the core files stay absent — the answer expires if the mechanism changes. Previously — **open, not classified as a defect**: `metadata.cache_layers.snapshot` was `"cold"` in all 240 invocations, on 16 repositories, including after `cache clear --all -y` followed by two runs of the same command. From outside, an unreachable state published as a layer cannot be distinguished from a broken cache write or from a layer only `cache warm` fills. | Answer it internally: if the state is reachable, `cache model` should say which command reaches it; if it is not, it should not be published as a layer. |
228
+
229
+ **Mapped, not new — inherited rows this pass re-witnesses.** `N-6` (cold `spring-audit`
230
+ on `openmrs-core` 8.86 s against the 5.8.7 anchor of 5.6 s, ×1.58, criterion ≤6.5 s) and
231
+ the Windows cost family belong to `BUG-1` / `AUD-511-R02` / `AUD-589-B02`. `D-09` (cost
232
+ anchors still stamped *"on 5.1.0"* in a 5.8.16 build, fourth witness) and `B-5` belong to
233
+ `BUG-5` / `AUD-588-B12` / `AUD-588-F03`. `N-5` (the advisory says *"136 Java files"* with
234
+ no population label while the progress line correctly says `82 [population=production_java_sources]`)
235
+ is the unclosed half of `AUD-588-B11`. `B-6` (budget overrun ×4.2 with `exit 0` on a
236
+ partial answer, and an `overrun_bound` whose basis measures this run rather than bounding
237
+ the next), `B-7` (`total_ms` covers 17.5–87% of wall clock, third consecutive version, and
238
+ its published summands exceed the total by 23% and 40%), `B-3` (`--agent` omits 15 blocks
239
+ `--compact` publishes, `mybatis` among them, at ×2.17 bytes and ×2.25 time, and
240
+ `sibling_view` exists only on one side and names no block) and `B-8` (`posture --compact`
241
+ caps `unresolved` at `limit: 0`, so `total: 3, shown: 0`) remain open exactly as recorded;
242
+ this pass adds witnesses, not rows.
243
+
244
+ **Closed by correction and re-verified from outside in this pass** — do not re-derive:
245
+ `N-1` (`rename-class` write guard, and it is the model for `AUD-591-A01`), `N-2` (posture counts
246
+ Spring IoC beans, verified against source on three repositories: struts 5, mall 155 of 161
247
+ countable, petclinic 12 exactly), `B-1` (`mybatis.mapper_interfaces` 76, matching the tree),
248
+ `B-4` (posture security population), `B-5` (`imports_found` renamed to `evidence`), `N-4`
249
+ (`summary_basis: "no descriptive section found in README"`, the literal acceptance text),
250
+ `N-7` (`impact` emits `candidates`), `N-8` (numeric bounds validated, and the only criterion
251
+ whose sweep was executed as a sweep — seven numeric flags), `N-9` (cache hints out of
252
+ `limitations`). `B-2`/`D-05` remains closed by refutation, re-verified independently: four
253
+ `<servlet-mapping>` elements in the struts `web.xml`, two of them inside the comment at
254
+ lines 143-153, and ASK reports 2.
255
+
256
+ ⚠ **`B-1` is not verifiable from outside.** Its criterion 3 required a negative fixture — an
257
+ XML whose namespace resolves to no interface must stay in `orphan_xml`. The field cannot
258
+ distinguish *"resolved correctly"* from *"closed by emptying the list"* without writing a
259
+ file into the audited tree. Verify it here, with a mapper XML naming a non-existent
260
+ interface under `@MapperScan`, before treating the row as closed.
261
+
262
+ **Process finding, and the highest-value item of this pass.** In four defect classes out of
263
+ four, the reported instance was closed and the class was not: `N-1` → `AUD-591-A01`, `B-5` →
264
+ `AUD-591-A03`, `N-7` → `AUD-591-A08`, `N-8` → `AUD-591-A02`/`AUD-591-A10`. The single exception is the one whose
265
+ acceptance criterion contained a sweep clause *and* was executed as a sweep (`N-8`, seven
266
+ numeric flags), and it did not recur. Every acceptance criterion carrying a sweep clause is
267
+ to be closed with a parametric test over the catalogue, not with the reported case.
43
268
 
44
269
  ### Sixth Audit Pass: `5.8.15` / `saint-server` / 2026-08-20
45
270
 
@@ -87,7 +312,7 @@ partial, `F03` open, `F04` explicit opt-in and `F05` locally verified only.
87
312
  | `N-6` | Medium | **not reproduced after cache fixes** | Isolated cold `openmrs-core`: 5.18 s wall / 4.76 s internal, all four layers cold. Retain the field report as an environment-specific reproduction request. |
88
313
  | `B-3` | Medium, core/supported | **closed by `773268d`** | Help no longer claims maximum signal; `sibling_view` points to compact. |
89
314
  | `B-6` | Medium, core | **closed by `17c6e06`** | Partial answers publish `overrun_bound` and its wall-clock basis. |
90
- | `B-8` | Low | **closed by `52593b6`** | Compact posture retains one unresolved sample when the total is non-zero. |
315
+ | `B-8` | Low | **reopened 5.8.17 — see `AUD-592-B01`** (was: closed by `52593b6`) | The fix retains a sample only when the total is exactly 1, not whenever it is non-zero, so `total: 3, shown: 0` still reproduces. |
91
316
  | `N-5` | Low, core | **closed by `c68083f`** | Progress and payload name the Java population. |
92
317
  | `N-7` | Low, core | **closed by `082e9de`** | Bounded fuzzy candidates and matching message hints handle close typos. |
93
318
 
@@ -110,7 +335,7 @@ automation remains deferred; provider non-coverage is the current product behavi
110
335
  | `AUD-513-N01` (`N-1`) | **closed in 5.8.14** | `rename-class` rewrote three Petclinic files under both `ASK_READONLY=1` and `--no-write`, then exited 0. The audit restored the tree; this is a write-policy breach, not an artifact-directory exception. `readonly.guard_mutation()` now refuses before any planned source write or physical rename; help names the command as mutating and regressions cover both readonly controls plus `--dry-run`. |
111
336
  | `AUD-513-N02` / `B-4` / `N-5` | **N02/B-4 implemented in 5.8.14; N-5 open** | `posture.summary.unconditional` counted 1,697 on a non-Spring Struts tree and 1,573 on Mall, and `security` was `0/0/0` with no population. The posture projection now uses the canonical Spring IoC bean population, publishes its population/unit and separates unconditional security beans. Regressions cover non-Spring, mixed-bean and conditional/unconditional security fixtures. Progress/advice still use incompatible “Java files” scopes (`N-5`), so the `AUD-588-F02` family remains open. |
112
337
  | `AUD-513-B01` / `B-3` / `B-5` | **B01/B-5 implemented in 5.8.14; B-3 closed as revalidated** | A MyBatis XML paired with a `*Mapper.java` interface now takes precedence over an absent `@Mapper` annotation, preserving XML-backed/`@MapperScan` interfaces in `mapper_interfaces` rather than misclassifying them as DTO mappers. `imports_found` now contains only actual Java imports; code, XML and build matches are emitted as `evidence`, including in compact output. `B-3` was stale against the current source: both views consume `JAVA_SPRING_SECTIONS`; a real MyBatis Mapper/XML regression proves the section is retained by `--agent`. |
113
- | `AUD-513-N03/N04/N07/N08/N09`, `B-6/B-7/B-8`, `D-09` | **N03/N04/N07/N08/N09 implemented in 5.8.14; remainder open** | `impact` now always publishes the `candidates` field its not-found message names, including `[]` when no close symbol exists. Budget advice no longer calls partial snapshot/RIS evidence a warm cache: it identifies that presence and says context/parse are unknown before analysis, whose payload then reports all four actual layers. README operational sections (`clone`, install, build, getting started, troubleshooting and related headings) are never promoted into a project summary; with no descriptive prose compact output publishes `project_summary: null` and `summary_basis`. `security_posture.limitations` is now cache-state invariant; dynamic execution advice travels in `security_posture.operational_hints`. Every numeric output bound rejects negative values at parse time with a structured `INVALID_INPUT` envelope (`flag`, `value`, `expected`); `--limit 0` still means no cap wherever that behavior was published. Budgets overrun, timing coverage is inconsistent, compact hides unresolved identities and cost anchors remain stale. |
338
+ | `AUD-513-N03/N04/N07/N08/N09`, `B-6/B-7/B-8`, `D-09` | **N03/N04/N07/N08/N09 implemented in 5.8.14; remainder open** | `impact` now always publishes the `candidates` field its not-found message names, including `[]` when no close symbol exists. Budget advice no longer calls partial snapshot/RIS evidence a warm cache: it identifies that presence and says context/parse are unknown before analysis, whose payload then reports all four actual layers. README operational sections (`clone`, install, build, getting started, troubleshooting and related headings) are never promoted into a project summary; with no descriptive prose compact output publishes `project_summary: null` and `summary_basis` — **`N04` regressed in 5.8.17 and is `AUD-592-R02`: licence headers and layout rules are promoted again, under a basis that asserts they are prose.** `security_posture.limitations` is now cache-state invariant; dynamic execution advice travels in `security_posture.operational_hints`. Every numeric output bound rejects negative values at parse time with a structured `INVALID_INPUT` envelope (`flag`, `value`, `expected`); `--limit 0` still means no cap wherever that behavior was published. Budgets overrun, timing coverage is inconsistent, compact hides unresolved identities and cost anchors remain stale. |
114
339
  | `AUD-588-B11` (container/units subset) | **partially implemented in 5.8.14** | `d745d31` discovers nested Maven modules; `273ae8e` suppresses unsupported `pr-impact.unaffected_basis`. `e53507d` separates direct impact matches from analysis seeds and implementation classes; `8386744` carries the established container-wiring fact into plan checklists and compare cost dimensions. Source population units and risk-tier wording remain open. |
115
340
  | `AUD-511-R02` / `AUD-589-B02` / `AUD-513-N06` | **partially corrected; `data-exposure` open, P1** | Nine of ten measured commands recovered, many to their best historical timings. `data-exposure --config` did not: 42.7 s versus 16.2 s in 5.8.8 and a 28.9 s model, with byte-equivalent output. Measure it with the idle-host, isolated-cache, interleaved-control protocol before changing code, then compare its reuse path with `validation`. |
116
341
  | `AUD-588-B02`, `B06`, B11 residuals, `B12` | **B02/B06 closed; B11/B12 open, P2** | `3869c87` excludes explicit EclipseLink from Hibernate applicability and effort; `037223e` publishes provider non-coverage. The re-audit verifies both and the payload reduction; `cf57457` makes aggregate wording name only applicable dimensions. Complete B11 consumer parity and calibrate B12 only from provenance-bearing controlled samples. |
@@ -310,6 +535,8 @@ retained only as evidence of what the auditor observed before the fix.
310
535
  | ASK-17 | **The parse store's default budget cannot hold a multi-repository workspace warm, and the cold cost of the largest repository is 7,6 minutes.** Measured on the 8-repo corpus (43 986 `.java`): `cache status` reports 56 730 entries at **511,12 MB against a 512 MB budget** — i.e. permanently sweeping by LRU — so `tutorials` (24 073 `.java`, the largest contributor) is the first candidate for eviction. Cold `--compact` on it: **457,6 s**, independently reproduced at 449 s, against 2,0 s on the second pass. | 5.8.4 corpus re-audit | Medium — warm plus `--compact` is still the answer (2,0 s), but a workspace this size cannot keep every repository warm at the default budget, and nothing tells the caller which repository is cold before it pays for it | **closed 5.8.5 — the disclosure shipped and the owed measurement taken, under our own control.** ⚠ **The +66 % does not reproduce, and the mechanism the row suspected is worth ~3 %, not 66 %.** Protocol: `tutorials` (24 073 `.java`), root `--compact`, both cache bases isolated **and emptied between runs** (`SOURCECODE_CONTEXT_CACHE_DIR` *and* `SOURCECODE_CACHE_DIR` — the first attempt isolated only the first, and the second pass of each version answered from the L2 view its own first pass had written: 2,4 s with an empty parse store, which is how a measurement of a cold path becomes a measurement of a warm one), `ASK_PARSE_CACHE_MAX_MB` fixed at 4 096 MB so no sweep can confound the comparison, 2 passes per version. **5.7.2 (`e803c8b`): 112,0 s / 112,6 s. This tree: 113,8 s / 115,2 s — +2,3 %**, with per-version dispersion of 1,005× and 1,012× and a payload 2,2 % larger. Then the audit's own condition, isolated as the only variable — the store pre-filled to 489 MB against the **default** 512 MB budget, so every 32 MB written sweeps: **117,0 s, +2,7 %.** So the LRU sweep is not where 457,6 s comes from, and neither is the analysis path: the auditor was right to hold the regression, and the held figure is now retracted **with a number** rather than on suspicion. ⚠ Scope, stated because the difference is unexplained rather than explained away: 457,6 s on Windows 11 / pipx is **not reproduced here** (113 s on 20 cores), and that gap is not assertable in either direction from this measurement — what is assertable is that 5.7.2 → this tree did not get slower and that a saturated store costs ~3 %. ✅ **What the measurement does confirm is the row's own claim, with our number: one repository of 24 073 `.java` leaves 46 840 entries and 470 MB in the store — 92 % of the 512 MB default budget** — so a workspace with a second repository of any size is permanently sweeping by construction, exactly as reported. Sizing rule, measured rather than guessed: ~20 KB of store per Java file, so ~500 MB per 24 000-file repository. Collateral confirmation of `AS-18`: after the saturated run the store rests at 669 MB — 157 MB over the budget and **under** the 736 MB effective ceiling it publishes at 7 writers — so the ceiling holds under the condition that produced the complaint. **The disclosure half:** `cache status` now publishes **coverage per repository**, which is the half the row itself recommended and the half that changes a decision: *«tutorials: 3 100 of 24 073 files cached (13 %)»* replaces a blind guess about whether to raise the budget, and an LRU eviction becomes visible **before** somebody pays 457,6 s to discover it. Both halves of the attribution were already held and thrown away — the walk knows the repository and it computes the store key for every file — so the run records the pairs (`parse_cache.record_repository_index`, from **both** readers of the store: `build_repo_ir` and the route-surface extractor, so the figure does not depend on which command was typed) and `store_stats` intersects them with the keys it collects **in the entry walk it already performs**: coverage costs an intersection, never a second scan of the store and never a scan of the repository. Three rules keep it honest. The index is **merged, never replaced**, because a `--changed-only` or `--since` run would otherwise shrink a 24 000-file population to the twelve files it read and publish *«12 of 12 cached (100 %)»* about a repository that is cold. The number travels with its **basis** — a file deleted since its last analysis still counts as recorded and reads as uncached, which errs toward *colder than it is* and says so. And the index is **neither an entry nor evictable**: the entry walk, the byte accounting and the LRU sweep all glob `*.json`, so its bytes are published under their own name (`repository_index_bytes`) rather than folded into a total that means entries — an index swept away with the entries it describes cannot report the eviction, which is the one moment it exists for. It lives inside the generation root, so retiring a generation retires its indexes with it: the keys are only readable by the build that wrote them. Regression `tests/test_parse_store_repository_coverage_ask17.py`, 11 assertions, including the one the row is about — entries deleted underneath a recorded repository make coverage **fall** while the population holds. ⚠ **Still owed, and unchanged:** the controlled cold measurement (5.7.2 against this tree, `ASK_PARSE_CACHE_MAX_MB` fixed, store emptied between runs, on a >20 000-file repository). Until it exists neither the +66 % nor its absence is assertable, and the row stays open on that half alone. **History:** **open** — ⚠ **not filed as a regression, deliberately**: the same cold figure was 274,8 s in 5.7.2 (+66 %), and the auditor refuses to call it one because the conditions are not comparable — in 5.7.2 the store had been retired by a version change, here it was mid-LRU-sweep. Seven false positives of exactly this class have been retracted over seven cycles; this is the eighth candidate and it is being held. What is owed is a measurement **we** control: cold `--compact` on a >20 000-file repository, 5.7.2 against 5.8.4, with `ASK_PARSE_CACHE_MAX_MB` fixed and the store emptied between runs — until that exists, neither the +66 % nor its absence is assertable. Recommended beside it, and cheap because both halves already exist: `cache status` should publish **coverage per repository** — how many entries belong to each analysed repository and what fraction of its files are covered — so *tutorials: 3 100 of 24 073 files cached (13 %)* replaces a blind decision about whether to raise the budget. Entries are content-addressed and the walk knows the repository, so this is a projection of facts we hold, not new analysis. **Refutation reproduced independently in cycle 8, under the reporter's own protocol**: both stores isolated and emptied, `ASK_PARSE_CACHE_MAX_MB=4096`, 2 passes per version — 5.7.2 at 112,0 / 112,6 s against 5.8.5 at 113,8 / 115,2 s (~3 %), and an A/B of the budget itself (512 MB sweeping by LRU against 4 096 MB that cannot sweep) at 9,1–10,1 s against 9,0–9,7 s: **the LRU sweep costs nothing measurable**. The +66 % is retired as the reporter's eighth false positive, with `E-37` named as its confounder. |
311
536
  | C3-128 | **`timeline` is the only gate-shaped command above 120 s, and it has not come back to its record.** Steady state, 4 runs, audited corpus: `timeline --since HEAD~5` costs **198 469 ms**, +27,3 % over its 5.7.2 record of 155 950 ms. The two other commands over 120 s are there structurally — `delta` (234 s) and `contract-diff` (136 s) analyse two whole trees — while `timeline` analyses **five** and costs less than `delta` does with two, so its cost is not explained by the number of trees it walks. | 5.8.2 → 5.8.4 re-audits | Low-Medium — an investigation command rather than a gate, and the last performance figure of the round sitting outside its own best | **closed 5.8.5 by measurement — the clock is attributed, and there is nothing left unaccounted to tune against.** The row's own instruction was ASK-16's: instrument before tuning. `timeline` published a per-sample total over what is really four costs — materialising a tree, measuring each watched metric, releasing the tree, and the remainder — so *«198 469 ms»* named a command rather than a phase. `perf.PhaseTimings` (the ASK-16 authority, always on) now splits it, with the phase names taken from `--watch` so the split cannot drift from the population, and the tree materialisation kept as its own phase because git's work must not be attributed to an analysis. **Measured, BroadleafCommerce (2 985 `.java`), `--since HEAD~5 --watch posture`, 5 samples: wall 38 921 ms — `measure:posture` 32 273 (82,9 %), `materialise_tree` 5 376 (13,8 %, ~1 075 ms per tree), `release_tree` 1 272 (3,3 %), unaccounted 0,45 ms (0,0 %).** So the answer to the row's premise — *«it analyses five trees and costs less than `delta` does with two»* — is that five sixths of the cost **is** the analysis, re-run per tree by construction, and the git work is a sixth of it: `timeline` is N × one analysis and there is no timeline-specific overhead to remove. Any future gain belongs to the metric being sampled (`C3-84`, `C3-127`), which is where it would also help every other command, and the payload now says so per run instead of per audit. Regression `tests/test_timeline_timings_c3_128.py`, 9 assertions, including that the series itself is byte-identical across two runs once the clock readings are removed — instrumentation that moved an answer would be a worse defect than the row. **Original note:** **open** — ⚠ the cache-reuse half of `B7` must **not** be reopened on this evidence: the v3 assertion (`max(sample) < 8 × posture_warm`) passes at 6,8 in both 5.8.2 and 5.8.4, against 14,8 / 13,4 / 12,4 / 8,5 in the four versions that genuinely failed it. What has not returned is the absolute cost. Attribution comes before tuning and is now cheap: `metadata.timings` (`ASK-16`) exists, so the per-phase split across the five trees can be published before anything is changed. |
312
537
  | B20 | **Thirteen declared renames with no cut-off date, and two distinct `1.0` identifiers meanwhile.** The registry publishes `pending_renames` for the 13 non-conforming `schema_version` values with their canonical name — exactly the policy `ASK-11` exists to enforce: the rename is an incompatible change, declared before it is made, never applied in silence. The consequence is published by the registry itself as `ambiguous_identifiers: 1` — `spring-audit` emits `1.0` for `core-analysis-v1` and `impact-chain` emits `1.0` for `impact-chain-v1`, so a consumer dispatching on the emitted version cannot tell them apart. | 5.8.4 re-audit (residual of `B19` / `B6`) | Low — declared debt rather than a defect, and the audit says so in as many words | **closed 5.8.5 — the window is declared where every other incompatible change is, and the reported ambiguity was understated by five shapes.** `BC-002` in `sourcecode.breaking_changes`: **announced 5.8.5, takes effect 6.0.0**, printed by `ask schema breaking-changes-v1`, carried in the CHANGELOG's `## Upgrading` section above the release history, and projected into `ask schema schemas-v1` as `pending_renames_window` — **one fact, two surfaces**, with a structural assertion that the registry source holds no second copy of the date. The registry's policy is generalised rather than loosened: a declared change is now `kind: exit_code` (must move an exit code) **or** `kind: contract` (must move a **published value** *and* name the version it takes effect in), because a rename breaks a consumer with every exit code still 0 and a registry that only knew about exit codes had nowhere to put it. A contract change publishes no `exit_code_before`/`after` at all — inviting a reader to check a field that cannot move is how a disclosure becomes noise. The 13 affected shapes are **read from `schema_registry.canonical_migrations()` at call time**, never copied: a shape that starts conforming leaves the declaration by itself (asserted by swapping the registry for a conforming one and watching the list empty). ⚠ **Correction to the reported cause, measured**: the audit named *two* shapes spelling their version `1.0`; there are **seven** — `core-analysis-v1`, `impact-chain-v1`, `pr-impact-v1`, `migration-blast-v1`, `spring-impact-v1`, `event-topology-v1`, `test-gap-ranking-v1` — so `ambiguous_identifiers: 1` was counting one *identifier* over seven shapes, not two. `ask schema 1.0` already resolved to all seven with each canonical name; the count is now asserted from the registry so it cannot be quoted from prose again. The rename itself is deliberately **not** made early: that is the incompatible change this policy exists to prevent. Regression `tests/test_schema_rename_window_b20.py`, 11 assertions, plus the ASK-11 battery generalised to both kinds. **Original note:** **open** — the missing half is a date, not a decision. Announce the cut-off window in `breaking-changes-v1` with the target version, the way every other incompatible change is announced, so a consumer can pin `core-analysis-v1` today and know when the bare `1.0` stops being emitted. Until then `ask schema 1.0` resolving to all seven shapes, each offering its canonical name, is the correct behaviour and must not be *fixed* by renaming an emitted value early — that is the incompatible change this policy exists to prevent. ⚠ **Round 10 ran on the 5.8.5 build and still reports the renames as *«deuda bien declarada, pero sin fecha»*, asking for precisely what `BC-002` already ships.** The row stays closed — the window exists, is announced 5.8.5 / effective 6.0.0, and is printed by `ask schema breaking-changes-v1`, carried in the CHANGELOG's `## Upgrading` and projected into `schemas-v1` as `pending_renames_window`. What it leaves behind is a **discoverability check, not a defect**: the reporter quoted `pending_renames` and `counts` out of the registry payload and did not see the window beside them, so verify that `pending_renames_window` travels in the same payload those two keys do — and if it does not, that is where it belongs. A declaration a ten-round auditor cannot find is not yet declared to a consumer. |
538
+ | E-38 | **A class-level route prefix carried by a meta-annotation is dropped, and the route is published without it.** `_build_route_surface` reads the class prefix only when the class symbol carries `@RequestMapping` or `@Path` **literally** (`repository_ir.py:5882-5887`); a repository that declares its own composed annotation gets `prefixes = [""]`, and the published `path` is the method suffix alone. **Measured on shenyu (5.8.17): 35 of 365 published routes — 9,6 %, over 21 controllers — carry `path: "/"`.** `AiProxyApiKeyController` is annotated `@RestApi("/selector/{selectorId}/ai-proxy-apikey")`, a Shenyu annotation meta-annotated with `@RestController` + `@RequestMapping`; its five handlers are published at `/`, `/`, `/batchDelete`, `/{id}`, `/page`. The served URLs do not exist as published. This is a **confident falsehood**, not a gap: no `path_resolution: "unresolved"`, no `path_expression`, no warning, and `route_census.distinct_routes` (316 of 365) reads the collisions as if two controllers genuinely shared a path. It also propagates to every consumer keyed on route text — `explain-endpoint`, `data-exposure --path-prefix`, `validation --path-prefix`, `pr-impact` route matching — where a correct query returns nothing. **The machinery already exists on the other axis**: `E-12`/`5.7.1` closed exactly this class for *beans* by resolving the spelling against the graph's `imports` edges, and `BeanGraph.build` already builds a meta-annotation map in its first pass. The route axis never asked. Vendor-agnostic by construction: the rule is "an annotation this repository declares that itself carries `@RequestMapping`/`@Path`", never a proprietary name. | found 2026-08-21 during the `AUD-590-R01` census triage, on the corrected 5.8.17 build; **not part of the seventh-pass queue and deliberately not fixed in it** — it moves route counts on any repository using composed annotations and needs its own measured pass with the golden set re-baselined | **High** — it is the rule this repository enforces most loudly, inverted on the route axis: a published route that is not served, with no affordance saying so. Blast radius is every repository with a framework-level or in-house composed controller annotation; shenyu is the witness, dubbo/sa-token/halo are unmeasured | **open** — direction: resolve the class-level prefix through the meta-annotation map `BeanGraph.build` already computes, in the order `E-12` established (explicit single-type import decides; wildcard leaves it admissible; a repository-declared annotation in the owner's package is the owner's own; **nothing resolved stays unresolved rather than acquiring a verdict from silence**). Where it cannot be resolved, publish `path_resolution: "unresolved"` with the annotation as `path_expression` — the contract the method-level path already honours — never a bare `/`. Regression on a fixture declaring its own composed annotation, plus a shenyu assertion that no published route is `/` |
539
+ | B29 | **A cache-freshness test asserts a cold cache without ensuring one, so it fails on every run but the first after a purge.** `tests/test_cli.py::test_compact` asserts `_meta.timing_ms.served_from_cache is False` and does nothing to make that true: the first `--compact` run over its fixture populates the cache the second one then hits. Verified symmetrically at `d9071d4` (clean) and at every commit of the seventh-pass queue — purge, run, pass; run again, fail — so it is **state-dependent, not a regression**, and it has been reported as "1 pre-existing failure" for the whole pass. It is still a red line in a suite whose value is that red means something. | found 2026-08-21 while establishing the baseline for the seventh-pass queue | **Low** as a defect, **medium** as an eroder — a permanently-red test trains a reader to skip the summary line, which is how a real regression gets shipped | **open** — direction: the test must create the state it asserts (purge this fixture's cache dir in the fixture, or assert the *transition* cold→warm across two runs rather than the absolute). Do not "fix" it by deleting the assertion: `served_from_cache` is the field `is_stale`/cache-freshness rows are keyed to, and a cold-start claim is worth holding |
313
540
  | E-37 | **Two environment variables move three cache stores, and no surface said which moves which.** `SOURCECODE_CACHE_DIR` relocates the per-repository snapshot store (core snapshots, the rendered L2 view, the RIS); `SOURCECODE_CONTEXT_CACHE_DIR` relocates the shared ones (the Canonical IR and the per-file parse store). Their defaults are siblings under `~/.sourcecode`, so the split is not derivable from the paths, and `cache status` published a path per store with no variable beside it while `docs/CACHE.md` named both variables in one sentence after the table. | found here 2026-08-18 while taking the controlled A/B `ASK-17` owed, not reported by the field | Medium — it does not make an answer wrong, it makes a **measurement** wrong, silently: an operator who redirects or empties *the cache* moves one store and is served answers out of the other | **closed 5.8.5** — measured first: with only `SOURCECODE_CONTEXT_CACHE_DIR` redirected, the second cold pass of each version answered from the L2 view its own first pass had written — **2,4 s against 113 s, with an empty parse store** — and the run reported `cache_source: L2_view` without anything naming the store it came out of. Now every store in `cache status` publishes `base_env`, from one authority (`cache._STORE_BASE_ENV`), the text answer prints *«<path> ($VAR moves it)»* beside each one, `stores.base_env_note` states the consequence rather than the layout, and `docs/CACHE.md` carries a **column** instead of a sentence covering both. Regression `tests/test_cache_store_base_env_e37.py`, 8 assertions, and the ones that matter are **behavioural**: setting the variable a store names moves *that* store and leaves the others where they were — a published mapping nobody exercises is the class of claim this repository refuses everywhere else. The CLI is forbidden a second copy of the mapping (structural assertion). **Confirmed from the outside in cycle 8, by the reporter falling into it first**: isolating only `SOURCECODE_CONTEXT_CACHE_DIR`, their second *cold* pass answered from the L2 view their first pass had written — **2,4 s against 113 s with an empty parse store** — and their write-up now quotes the shipped `base_env` mapping back as the fix (*"dos variables mueven tres almacenes y ninguna superficie decía cuál"*). A defect found while measuring, that was itself corrupting the measurement. |
314
541
  | B21 | **Six of the thirteen shapes being renamed name no command that emits them.** `schemas-v1` publishes `emitted_by` per shape; for `migration-blast-v1`, `spring-impact-v1`, `event-topology-v1`, `test-gap-ranking-v1`, `canonical-ir-v1` and `hibernate-strategy-v2` it is `[]`, beside seven shapes where it is populated. Read the way a consumer reads a list, an empty `emitted_by` says *nothing emits this shape* — which tells the one consumer that **is** affected by `BC-002` that it is not. | found here while declaring `BC-002` (5.8.5), not reported by the field | Low — a disclosure gap in a declaration whose whole purpose is to let a consumer decide whether it is affected | **closed 5.8.6 — derived, not typed out.** `schema_producers` reads the two authorities the CLI already holds — the registered command tree and the import graph of the module that defines each callback — and answers with the evidence it derived on: `command_import` (the command's own body imports the producing module), `module_import` (one hop, through an intermediate narrow enough to attribute), `substrate` (reached from more commands than a producer would be) or `unresolved` (nothing found, said in a sentence). Five of the six now name a command — `migration-blast-v1` → `migrate-check`, `spring-impact-v1` and `event-topology-v1` → `impact-chain`, `test-gap-ranking-v1` → `impact`, `hibernate-strategy-v2` → `migrate-check`/`migrate-recipe` — and the sixth, `canonical-ir-v1`, is reached from **28 modules**: naming every command would be less true than naming none, so it publishes `substrate` with the count and the instruction to match on `subject`. A declared producer still wins, because some producers are not one registered command (`ask (root)`, `baseline capture`). Regression `tests/test_schema_producers_derived_b21.py`, 33 assertions, parametrised by the registry so a fourteenth unnamed shape fails the battery. **History:** open, and named rather than shipped as a fact: every `BC-002` row with an empty producer list carries `emitted_by_basis` — *«not declared in the registry: this shape's producing command is not named yet, so match on `subject`»* — so the gap is visible instead of being read as a negative claim. What remains is to fill the six, and the honest way is derivation rather than six more hand-written strings: the producer is discoverable from the command that constructs each payload, which is the same authority every other population in this repository comes from. Until then the assertion in `tests/test_schema_rename_window_b20.py` holds the weaker invariant a reader can rely on: a row either names its emitters or says why it cannot. |
315
542
  | C3-129 | **`timeline`'s absolute cost, now fully attributed and still 27 % above its own record.** `C3-128` closed by instrumenting rather than tuning, which was the right order and left this behind: the clock is split, nothing is unaccounted, and the number has not moved. Round 10 on 5.8.5 measures `timeline --since HEAD~5` at **~198 000 ms against the 5.7.2 record of 155 950 ms (+27 %)** — the only gate-shaped command over 120 s that is not there structurally (`delta` 191 s and `contract-diff` 115 s walk two whole trees, `compare` 46 s walks N candidates; `timeline` walks **five** trees and costs more than `delta` does with two). | 5.8.2 → 5.8.5 re-audits (residual of `C3-128`, which closed the attribution half in 5.8.5) | Low-Medium — an investigation command rather than a CI gate, and the last performance figure of the round sitting outside its own best | **closed 5.8.6 — by not paying twice for what does not change.** A metric at a commit is a pure function of that tree, and the tree at a sha never changes: the only cache in this product whose key can be exact rather than heuristic (analyzer fingerprint + metric + commit sha). `timeline_cache.SampleCache` stores measured values under the per-repository core store, and when every watched metric for a sample is a hit the tree is **never materialised**, which is the 13,8 % the split attributes to git. Measured on BroadleafCommerce over the same range: **26 845 ms → 12 ms**, samples and transitions byte-identical. What makes it admissible rather than merely fast, all asserted: a failed sample is never stored (one bad run cannot become permanent), the fingerprint is in the key (corrected analysis is never served from a previous release), the hit is published (`sample_cache`, `from_cache` per sample, and `cost` separates the cached population from the measured one instead of reporting a median over both), and `--no-cache` / `ASK_TIMELINE_NO_SAMPLE_CACHE=1` measure everything again so the command stays measurable without emptying a store (E-37). `cache_model` carries the layer, so `cache model` and `docs/CACHE.md` state it from one authority. Regression `tests/test_timeline_sample_cache_c3_129.py`, 11 assertions. **History:** open — and it opened with the answer already published, which is why it is cheap to attack and expensive to leave: `timeline`'s own `timings` say **82,9 % `measure:posture`, 13,8 % `materialise_tree`, 3,3 % `release_tree`, 0,0 % unattributed**. So this is not a `timeline` defect at all: it is `posture` paid five times over five materialised trees, and it makes the same substrate question `C3-127` asks — with the difference that here the five payers are one command, so the cache is intra-invocation and needs no cross-command contract. ⚠ **Do not reopen `B7`**: cross-tree cache reuse is healthy, the v3 assertion passes at 6,8×. Acceptance: `timeline --since HEAD~5` back inside +10 % of 155 950 ms on the audited corpus with `timings` still summing to 100 %, and `measure:posture` falling as a share rather than the total falling for an unnamed reason. |
@@ -46,7 +46,7 @@ CLI commands — impact, endpoints, spring-audit, explain, … each a pro
46
46
  The key idea: the extraction is **content-addressed**. Commands reuse the parse cache and,
47
47
  where their analysed scope matches, the shared Canonical IR; `ask cache model` names what a
48
48
  warm buys for each command rather than implying that every projection costs the same. In
49
- 5.8.16, `validation` enters through that shared CIR and `data-exposure` reuses one semantic
49
+ 5.8.18, `validation` enters through that shared CIR and `data-exposure` reuses one semantic
50
50
  model across all declared label seeds. (The extraction and consumption contract is fixed in
51
51
  the architecture ADRs 0001–0004 under `docs/architecture/`.)
52
52
 
@@ -108,7 +108,7 @@ human. They project the same semantic model into agent-ready context:
108
108
  | `ask prepare-context <task>` | Task-shaped context bundle (`onboard`, `review-pr`, `fix-bug`, `refactor`, `generate-tests`) sized for an LLM prompt. |
109
109
  | `ask repo-ir` | The low-level symbol IR itself — the raw semantic representation, for tools that want the graph, not a report. |
110
110
  | `ask --agent` | Agent-optimized output envelope for any command. |
111
- | `ask --compact` | A ~10K-token subset (status, stacks, entry points, key dependencies) for cheap orientation. |
111
+ | `ask --compact` | A bounded subset (status, stacks, entry points, key dependencies) for cheap orientation. See `ask --help` for the current output budget. |
112
112
  | `ask mcp` | Runs ASK as an MCP server so an agent (Claude Desktop / Cursor) calls it as a tool. |
113
113
 
114
114
  Rule of thumb: **developer-facing** commands (`impact`, `spring-audit`, `endpoints`,
@@ -142,7 +142,7 @@ pipx install sourcecode # isolated install, no venv needed
142
142
 
143
143
  # Verify
144
144
  ask version
145
- # ask 5.8.16
145
+ # ask 5.8.18
146
146
  ```
147
147
 
148
148
  Requires Python 3.9+.
@@ -240,7 +240,7 @@ ask . --compact --git-context # includes commit hotspots
240
240
  ask . --compact --copy # copy to clipboard
241
241
  ```
242
242
 
243
- Output (~2,500–4,000 tokens): detected stacks, entry points, dependencies, transactional boundaries (Spring), env vars, confidence level, analysis gaps.
243
+ Output: detected stacks, entry points, dependencies, transactional boundaries (Spring), env vars, confidence level, analysis gaps. The current output budget is published by `ask --help`.
244
244
 
245
245
  ### `ask --agent`
246
246
 
@@ -251,7 +251,7 @@ ask /repo --agent --output context.json
251
251
  cat context.json | claude -p "Explain the architecture"
252
252
  ```
253
253
 
254
- Output (~4,500–5,500 tokens): project identity, entry points, file relevance ranking, architecture classification, confidence.
254
+ Output: project identity, entry points, file relevance ranking, architecture classification, confidence. The current output budget is published by `ask --help`.
255
255
 
256
256
  ### `ask onboard` *(supported)*
257
257
 
@@ -1363,8 +1363,8 @@ reason to warm before the first prompt (`ask cache warm <path>`).
1363
1363
 
1364
1364
  | Mode | When to use | Token size |
1365
1365
  |------|-------------|-----------|
1366
- | `--compact` | AI session start, high-level overview | 2,500–4,000 |
1367
- | `--agent` | AI agent system prompt injection | 4,500–5,500 |
1366
+ | `--compact` | AI session start, high-level overview | See `ask --help` |
1367
+ | `--agent` | AI agent system prompt injection | See `ask --help` |
1368
1368
  | `onboard` | Task-structured agent/developer onboarding | ~2,600 |
1369
1369
  | `fix-bug` (trimmed) | Bug triage with token budget | ~4,600 |
1370
1370
  | `--full` | Raise the display caps (boundaries and DTO mappers in full, `file_relevance` to 40) | bounded, larger |
sourcecode/cache_model.py CHANGED
@@ -39,6 +39,11 @@ class Layer:
39
39
  location: str
40
40
  invalidated_by: str
41
41
  warmed: str # what `ask cache warm` does to this layer
42
+ #: Which commands populate it. Empty means every command that reads it also
43
+ #: fills it — the ordinary case. Named where it is not (AUD-591-Q01): a layer
44
+ #: only one command writes reads `cold` on every other command's payload, and
45
+ #: a reader with no way to know that reads it as a cache that is not working.
46
+ written_by: str = ""
42
47
 
43
48
 
44
49
  @dataclass(frozen=True)
@@ -191,6 +196,15 @@ LAYERS: tuple[Layer, ...] = (
191
196
  "presentation flags (--compact, --agent, --full, --format, …) for the view"
192
197
  ),
193
198
  warmed="built for the compact view (`--agent` also builds the agent view)",
199
+ written_by=(
200
+ "the root analysis alone — `ask <repo>` and `ask cache warm`. "
201
+ "AUD-591-Q01: `cache_layers.snapshot` read `cold` in all 240 field "
202
+ "invocations across 16 repositories, including after `cache clear "
203
+ "--all -y` plus two runs, because every one of those runs was of "
204
+ "another command. No other command writes `core-*.json.gz`, so on "
205
+ "any run but the root analysis this layer is `cold` by construction "
206
+ "and its being cold says nothing about the cache."
207
+ ),
194
208
  ),
195
209
  Layer(
196
210
  id="ris",
@@ -941,10 +955,13 @@ def as_dict(here: "Optional[Conditioning]" = None) -> dict:
941
955
  def render_markdown() -> str:
942
956
  """The model as the tables published in the user guide."""
943
957
  out: list[str] = []
944
- out.append("| Layer | What it stores | What invalidates it | `cache warm` |")
945
- out.append("|---|---|---|---|")
958
+ out.append("| Layer | What it stores | What invalidates it | `cache warm` | Written by |")
959
+ out.append("|---|---|---|---|---|")
946
960
  for lyr in LAYERS:
947
- out.append(f"| `{lyr.id}` | {lyr.stores} | {lyr.invalidated_by} | {lyr.warmed} |")
961
+ out.append(
962
+ f"| `{lyr.id}` | {lyr.stores} | {lyr.invalidated_by} | {lyr.warmed} | "
963
+ f"{lyr.written_by or 'every command that reads it'} |"
964
+ )
948
965
  out.append("")
949
966
  out.append(f"Measured on {REFERENCE_REPOSITORY}, each command in isolation.")
950
967
  out.append("")
@@ -1038,6 +1055,10 @@ def render_text(here: "Optional[Conditioning]" = None) -> str:
1038
1055
  lines.append(f" location {lyr.location}")
1039
1056
  lines.append(f" invalidated {lyr.invalidated_by}")
1040
1057
  lines.append(f" cache warm {lyr.warmed}")
1058
+ if lyr.written_by:
1059
+ # AUD-591-Q01: `cache model` is where a reader goes to find out why a
1060
+ # layer reads `cold`, so the answer belongs on the layer.
1061
+ lines.append(f" written by {lyr.written_by}")
1041
1062
  lines.append("")
1042
1063
  lines.append("What a warm gives each command")
1043
1064
  lines.append(f" Timings: {REFERENCE_REPOSITORY}, each command measured in isolation.")