sourcecode 5.8.17__py3-none-any.whl → 5.8.20__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Potentially problematic release.
This version of sourcecode might be problematic. Click here for more details.
- sourcecode/__init__.py +1 -1
- sourcecode/_docs/DEFECT-LEDGER.md +251 -11
- sourcecode/_docs/USER_GUIDE.md +13 -11
- sourcecode/analysis_budget.py +27 -0
- sourcecode/build_stamp.py +49 -0
- sourcecode/cache_model.py +4 -2
- sourcecode/caller_metrics.py +92 -3
- sourcecode/cli.py +238 -70
- sourcecode/container_wiring.py +66 -8
- sourcecode/defect_classes/__init__.py +160 -0
- sourcecode/defect_classes/registry.json +158 -0
- sourcecode/execution_plan.py +32 -5
- sourcecode/explain.py +5 -1
- sourcecode/non_coverage.py +13 -7
- sourcecode/output_budget.py +27 -8
- sourcecode/parse_cache.py +42 -0
- sourcecode/partial_contract.py +27 -0
- sourcecode/pr_impact.py +18 -1
- sourcecode/product_info.py +38 -1
- sourcecode/repository_ir.py +148 -4
- sourcecode/risk.py +22 -0
- sourcecode/sarif.py +2 -0
- sourcecode/schema_registry.py +15 -0
- sourcecode/selection_contract.py +52 -0
- sourcecode/selftest.py +163 -0
- sourcecode/serializer.py +63 -7
- sourcecode/spring_impact.py +45 -0
- sourcecode/summarizer.py +58 -4
- sourcecode/wired_surface.py +312 -0
- {sourcecode-5.8.17.dist-info → sourcecode-5.8.20.dist-info}/METADATA +4 -4
- {sourcecode-5.8.17.dist-info → sourcecode-5.8.20.dist-info}/RECORD +35 -29
- {sourcecode-5.8.17.dist-info → sourcecode-5.8.20.dist-info}/WHEEL +0 -0
- {sourcecode-5.8.17.dist-info → sourcecode-5.8.20.dist-info}/entry_points.txt +0 -0
- {sourcecode-5.8.17.dist-info → sourcecode-5.8.20.dist-info}/licenses/LICENSE +0 -0
- {sourcecode-5.8.17.dist-info → sourcecode-5.8.20.dist-info}/licenses/NOTICE +0 -0
sourcecode/__init__.py
CHANGED
|
@@ -16,7 +16,38 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
|
|
|
16
16
|
|
|
17
17
|
## Current Synchronization
|
|
18
18
|
|
|
19
|
-
**Latest attached audits — release `5.8.
|
|
19
|
+
**Latest attached audits — audit of release `5.8.18`, 2026-08-21, two independent rounds,
|
|
20
|
+
the first pass in this record that closes with no P0 and no P1 open, and **closed end to
|
|
21
|
+
end in `5.8.19`**.** Round A (MSAS,
|
|
22
|
+
3 342 Java files, the same HEAD bit for bit for the third consecutive round, 18 commands,
|
|
23
|
+
~30 invocations) scores **8.95/10** — 8.21 → 8.77 → 8.95 — with **zero false statements
|
|
24
|
+
about the repository for the third consecutive run**, 4 of its 6 previous findings closed
|
|
25
|
+
clean, one closed by half and one open. Round B (a 12-repository bank, ~200 invocations,
|
|
26
|
+
the §7 protocol) scores **8.8/10** — up from 7.7 — closes **8 of its 10 previous rows**,
|
|
27
|
+
and calls the release adoptable without reservation. **Both rounds record zero content
|
|
28
|
+
regressions, zero unhandled exceptions and zero unauthorised writes; the two performance
|
|
29
|
+
regressions Round B initially published were retracted by its own re-measurement.** The
|
|
30
|
+
whole eighth-pass queue is verified closed from outside, including the two regressions the
|
|
31
|
+
shipped ledger still marked `open` — which is itself a row (`AUD-594-X04`) and is corrected
|
|
32
|
+
in place below. **The ninth-pass queue is closed in `5.8.19`: twelve rows, one atomic
|
|
33
|
+
commit each, with every regression assertion written over the catalogue that already
|
|
34
|
+
exists rather than over the reported symbol — the rule three passes of evidence now
|
|
35
|
+
support. There is no open queue in this ledger.**
|
|
36
|
+
|
|
37
|
+
**What is left has moved once more, and it is no longer a defect of design.** The seventh
|
|
38
|
+
pass ended with contracts that were *absent*; the eighth with contracts that *contradicted
|
|
39
|
+
each other across authorities*; this one ends with contracts that are nearly coherent and
|
|
40
|
+
leak at the places a fix did not sweep to: a help header the generator never reached, a
|
|
41
|
+
missing `*_population` key, a `sibling_view` that names no block, a `[:5]` beside a `[:30]`
|
|
42
|
+
that was just declared, and a `selftest` that does not expose the invariants the last three
|
|
43
|
+
releases repaired. Two rows carry more weight than their P2 severity suggests and lead the
|
|
44
|
+
queue for that reason: `AUD-595-A03`, because a budget-exhausted audit exits 0 and a CI gate
|
|
45
|
+
reads it as a pass over a payload that declares itself a floor; and `AUD-594-X01`, because
|
|
46
|
+
one of its three stale literals tells the reader the root analysis **must not be CI-gated**,
|
|
47
|
+
which is the opposite of what the binary now does.
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
**Previous attached audits — release `5.8.16`, 2026-08-21, two independent rounds.**
|
|
20
51
|
Round A (`saint-server`, 3 342 Java files, 17 commands, 25 invocations) scores **8.21/10**
|
|
21
52
|
with **zero false statements about the repository** and two contract defects. Round B (a
|
|
22
53
|
16-repository bank, 49 to 24 073 Java files, **240 invocations, 0 crashes, 0 tracebacks**)
|
|
@@ -28,10 +59,11 @@ previously open defects by correction with **zero regressions among the inherite
|
|
|
28
59
|
|
|
29
60
|
Both rounds agree on where the product now fails, and it is not the Java/Spring analysis:
|
|
30
61
|
it is the contract an agent consumes programmatically — the write guard, repository
|
|
31
|
-
identity, list units, flag scope, and published budgets.
|
|
32
|
-
the
|
|
62
|
+
identity, list units, flag scope, and published budgets. Its queue closed in `5.8.17` and the eighth-pass queue in `5.8.18`;
|
|
63
|
+
**the ninth-pass section below is the most recent queue and it is closed**, and older rounds
|
|
64
|
+
remain for traceability.
|
|
33
65
|
|
|
34
|
-
**Mechanisms confirmed in source during intake** (not taken on the reports' word):
|
|
66
|
+
**Mechanisms confirmed in source during the seventh-pass intake** (not taken on the reports' word):
|
|
35
67
|
`verify --init` writes with no `readonly.guard` (`cli.py:10706-10708`) while five sibling
|
|
36
68
|
writers guard; `spring-audit` derives `repo_id` from the CIR content hash
|
|
37
69
|
(`spring_security_audit.py:1416`, `spring_tx_analyzer.py:1081`) while every other surface
|
|
@@ -52,10 +84,14 @@ itself rather than merely under-declaring.
|
|
|
52
84
|
**Release status:** entries marked “pending `5.8.14`” in their historical wording are
|
|
53
85
|
shipped in `5.8.14`; the correction battery fixes listed above are included in `5.8.15`.
|
|
54
86
|
|
|
55
|
-
**Current release:** `5.8.
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
87
|
+
**Current release:** `5.8.19` closes the whole ninth-pass queue, one commit per row, with
|
|
88
|
+
the suite green end to end — 11 008 passing, and the pre-existing `B29` failure the
|
|
89
|
+
previous release shipped with is gone. `5.8.18` before it closed the eighth-pass queue. `5.8.17` before it carried the seventh-pass queue;
|
|
90
|
+
`5.8.16` before that carried two post-audit
|
|
91
|
+
corrections: compact cold-start identity is preserved for `no_ris`, and manifest-cache keys
|
|
92
|
+
include Python requirements files. The testing-repository hygiene exercise changed no
|
|
93
|
+
external repository; it completed 108/108 read-only workflow runs with valid JSON and exit
|
|
94
|
+
code 0.
|
|
59
95
|
|
|
60
96
|
The battery also closed the following reproducible contract defects in `5.8.15`:
|
|
61
97
|
`migrate-apply` is registered across help, progress, path-admission, output-purity,
|
|
@@ -73,6 +109,208 @@ with neither parseable stdout nor its requested artifact) or `BUG-6`
|
|
|
73
109
|
queue position. `BUG-1` and `BUG-5` gain witnesses on public OSS repositories, recorded
|
|
74
110
|
under their own rows rather than as new ones.
|
|
75
111
|
|
|
112
|
+
### Ninth Audit Pass: `5.8.18` findings, closed in `5.8.19` / MSAS + 12-repository bank / 2026-08-21
|
|
113
|
+
|
|
114
|
+
**Closed queue — all twelve rows are `closed 5.8.19`, one atomic commit each.** Kept in full because the mechanisms are the evidence and the direction column is what the fixes were written against.
|
|
115
|
+
|
|
116
|
+
**How the pass read when it was filed.** Two independent rounds re-ran their protocols against the release
|
|
117
|
+
that closed the eighth-pass queue. Round A (MSAS, `banyan-v2` @ `3dde0376`, 3 342 Java
|
|
118
|
+
files — the same subject bit for bit for the third consecutive round — 18 commands, ~30
|
|
119
|
+
invocations) scores **8.95/10**, up from 8.77 and 8.21. Round B (a 12-repository bank,
|
|
120
|
+
~200 invocations, the §7 measurement protocol) scores **8.8/10**, up from 7.7, and calls
|
|
121
|
+
the release **adoptable without reservation**.
|
|
122
|
+
|
|
123
|
+
**No regressions, of either kind.** Round A compared eleven payloads byte for byte against
|
|
124
|
+
5.8.17 and found nine identical, with `root --compact`, `root --agent` and `selftest`
|
|
125
|
+
moving by the 1–2 bytes of their own version string; the analysis figures are unchanged
|
|
126
|
+
(`spring-audit` 373/94, `posture` 1 908/5, `migrate-check` 0/532/97, `endpoints` 3 763 with
|
|
127
|
+
no drift where 5.8.17 drifted +21, `no_security_signal` 725 = `undocumented` 725, `repo_id`
|
|
128
|
+
cardinality 1). Round B records **zero content regressions, zero unhandled exceptions in
|
|
129
|
+
~200 invocations, zero unauthorised writes across 10 of 10 trees, and — after the
|
|
130
|
+
retraction in the next paragraph — zero timing rows above the ×1.25 threshold**. This is
|
|
131
|
+
the first pass in this record with no P0 and no P1 open at its close.
|
|
132
|
+
|
|
133
|
+
**A reported performance regression is retracted, and the retraction is a product finding.**
|
|
134
|
+
Round B published `spring-audit` on `openmrs-core` as ×2.30, then ×1.39, then withdrew both:
|
|
135
|
+
with an explicit `cache warm` and a settled host the figure is **×1.07** (2.48 s against a
|
|
136
|
+
5.8.17 anchor of 3.00 s at the same protocol). The diagnostic is the product's own
|
|
137
|
+
`timings.unaccounted_pct` — 55 % in the state that *reports* warm but is not, 11.8 % fully
|
|
138
|
+
warm, 1.5 % genuinely cold — and it is the reason `AUD-595-B01` below is filed as a defect
|
|
139
|
+
rather than as an auditor's note. Round A's `impact` 7 s → 101 s is likewise not a
|
|
140
|
+
regression: the version bump emptied the Shared CIR (19.18 MB → 0) and retired 58 056 parse
|
|
141
|
+
entries, and `impact-chain` immediately after, on the now-warm CIR, took 11 s.
|
|
142
|
+
|
|
143
|
+
**Two figures the previous pass carried are corrected here, both downward in severity.**
|
|
144
|
+
`B-3`'s *"`--agent` is +61 % over its published band"* was measured against the guide's
|
|
145
|
+
hand-written band, which `AUD-593-N02` correctly deleted; against the ceiling the code
|
|
146
|
+
actually enforces (`BUDGET_AGENT = 40 000`, `output_budget.py:292`, applied at
|
|
147
|
+
`cli.py:4585-4589`) the measured 35 358 B is **inside budget**. `--agent` is not over
|
|
148
|
+
budget; it is unpublished, which is `AUD-594-X02`. And `AUD-592-A02`'s `endpoints` half was
|
|
149
|
+
not a truncation at all — see that row.
|
|
150
|
+
|
|
151
|
+
| ID | Severity | Current status | Required direction |
|
|
152
|
+
|---|---|---|---|
|
|
153
|
+
| `AUD-594-X01` / `A-1`, residual of `AUD-593-N03` | **P2** | **closed `5.8.19`** (`dd222b9`) — the tuple moved to `analysis_budget`, a module neither `cli` nor `cache_model` owns, both render from it, and the guide now defers to `ask --help` rather than keeping a copy; the `must not be CI-gated` sentence is corrected, not trimmed. Regression is parametric over the tuple: any paragraph of any shipped authority that names the variable and enumerates commands must enumerate all of them. Mechanism as found: open, confirmed in source, and one of the three literals is an inverted operating instruction, not a stale number. The generated Options line enumerates eight commands (`cli.py:589-596` → `:632`); three hand-written literals in two other shipped authorities still name three. `cache_model.py:679` — the header paragraph of the same help page — ends *"ASK_MAX_ANALYSIS_SECONDS bounds only the phase-runner commands (spring-audit, risk and audit-report)"*. `docs/USER_GUIDE.md:1533` repeats it, and **`:1558` goes further: commands outside "that phase-runner set are not bounded by this variable and must not be CI-gated"** — it tells a reader they cannot gate the root analysis in CI, which they now can. Behaviour settles it: with the variable at 90 the root prints *"ask is a repo-wide analysis and the configured budget is 90s"*. Same shape as the closed `AUD-590-B03`: two incompatible statements on one page, half generated and half not. **`D-09`/`BUG-5`'s stale cost anchors live in the same `return`** (`cache_model.py:671-680`, `FIELD_ANCHOR_MEASURED_VERSION`), so one function carries two of the three oldest rows in this ledger. | Replace the three literals with `_budgeted_analysis_commands_phrase()`, moving `_BUDGETED_ANALYSIS_COMMANDS` into a module both `cli` and `cache_model` may import rather than creating a `cache_model → cli` cycle. Delete the `:1558` sentence outright — it is wrong in the direction that costs a user a CI gate. Assertion, parametric over the tuple: no string in any shipped authority (rendered `--help`, `cache_model`, `docs/*.md`) enumerates a **proper subset** of `_BUDGETED_ANALYSIS_COMMANDS` next to the variable name. |
|
|
154
|
+
| `AUD-595-A03` / `B-6` exit-code half | **P2** | **closed `5.8.19`** (`ab7285f`), **and the reported premise is half refuted by measurement.** Measured on dubbo (4 049 files, budget 1 s): the default exits 0, but **`--ci` already exits 75** — `F-AV`'s code for *"the gate did not finish"* — so the product does not read a cut run as a pass wherever it was asked to gate, and 0 without `--ci` is the same contract findings have had all along. What was genuinely missing is the other half of `AUD-591-A09`: the block said what did not run and never what the process would do about it. Now `_partial` carries `gate_exit_code` (75), `exit_code` (what this process exits with) and `gate_exit_code_basis`, from one decision function the exits themselves read, and `--allow-partial` lets a gating pipeline accept a cut answer deliberately. Field-verified in three states: default 0/0, `--ci` 75/75, `--ci --allow-partial` 0/0 with the gate still publishing 75. Mechanism as found: open. `spring-audit` exhausting its budget exits **0** on 3 of 3 runs, over a payload that declares itself a floor: `_partial: {partial: true, why_stopped: "budget_exhausted", phases_completed: ["ir"], phases_pending: ["tx_audit", "security_audit"], how_to_read: "…every count here is a FLOOR over the complete answer…"}`. The block is exemplary; the exit code contradicts it, and a pipeline reading `$?` goes green with two of three phases unexecuted. The overrun half of `B-6` closed in `17c6e06`; this half has never had a row of its own, which is part of why it has survived three passes. Note for the next measurement: `spring-boot` no longer exhausts a 15 s budget (~9 s now), so the repro needs `ASK_MAX_ANALYSIS_SECONDS=3`. | Apply `AUD-591-A09` verbatim — it is already shipped and re-verified in this pass: `verify --no-ci` → process 0 / `exit_code` 0 / `gate_exit_code` 2 / `gate_exit_code_basis`; `verify --ci` → process 2. Emit a non-zero exit whenever `_partial.partial` is true unless `--allow-partial` is passed, with the basis string saying which it was. Site: `cli.py:10650-10653`. |
|
|
155
|
+
| `AUD-595-A02` / `N-5`, advisory half of `AUD-588-B11` | **P2** | **closed `5.8.19`** (`3e15227`), and the six-report survival is explained by what the fix had to be: **the two figures count different populations**, so labelling the advisory with the payload's name would have been a second falsehood. The cheap capped walk now names its own — `all_java_files_excluding_build_output`, a constant with a `Scope.population` property so a consumer need not parse the sentence — and the payload publishes `java_population_count` beside `java_population`, which no surface had ever carried together. Field-verified on halo: advisory `1 349 Java files [population=all_java_files_excluding_build_output]`, payload `production_java_sources 990`. Two populations, both named, both numbered. Mechanism as found: open, and it is one line. One run publishes four figures for one population and labels none of them: the advisory says *"This repository: 8 372 Java files"*, the payload carries `java_population: "production_java_sources"` **with no number**, the help says *"at least 8 400"*, and the auditor's controls give 8 667 (all `*.java`) and 5 135 (`src/main` only) — 8 372 matches neither, so from outside the population cannot be named. The label exists and works on three of four surfaces: built at `cli.py:10301`, emitted as `[population=…]` at `:10348`, `:10400`, `:10412`, attached to the payload at `:10518`. The fourth is `phased_run.py:367`, `lines.append(f" This repository: {scope}.")`. | Emit the same `[population=…]` from `phased_run.py:367`, and publish the **count** beside `java_population` in the payload — today no single surface carries number and label together, which is why six reports have not closed it. Assertion: every Java-file count printed or published carries its population name, and two surfaces of one run never publish two populations without declaring the difference. |
|
|
156
|
+
| `AUD-594-X02` / `A-1` second half | **P2** | **closed `5.8.19`** (`1193937`) — `agent_budget_phrase()` is derived from `BUDGET_AGENT`, the constant `_apply_budget` enforces, and both phrases now share one renderer so two formats for one fact cannot appear. The rendered help publishes `bounded at ~10K tokens (40 KB, chars/4)` on the `--agent` entry and on its option, beside `--compact`'s `~7K`. `BUDGET_COMPACT`'s misleading `# compact/agent main cmd` comment is corrected. The sweep is over the module: every `BUDGET_*` reaching an `_apply_budget` call and bounding a documented view must have a `*_budget_phrase()` that appears in the rendered help; the interior budgets carry a stated exemption. `AUD-590-B03`'s own assertion was re-scoped from *one figure per page* to **one figure per flag**, with the page held to figures the budget module generates — the page-wide form would have failed on the correct fix. Mechanism as found: open. `BUDGET_AGENT = 40 000` exists (`output_budget.py:292`) and is enforced (`cli.py:4585-4589`), and **no authority publishes it**: the packaged guide now defers to `ask --help` (the `AUD-593-N02` fix, correct), and the help's `--agent` entry describes content — *"identity, entry points, dependencies, confidence, gaps"* — and no budget. The only *"~7K tokens"* in the help belongs to `--compact`. `--agent` went from four contradictory figures to none; the deference works for `--compact` because the help publishes, and for `--agent` it points at a blank. Measured: 35 358 B ≈ 8 839 est. tokens against an unpublished 40 000 B ceiling. Adjacent and misleading: `BUDGET_COMPACT`'s comment reads `# compact/agent main cmd`. | Add `agent_budget_phrase()` beside the existing `compact_budget_phrase()` (`output_budget.py:302-320`), derived from `BUDGET_AGENT` and `TOKEN_MODEL` exactly as its sibling is, and consume it in the help's `--agent` entry as `cli.py:583` consumes the compact one. Fix the `BUDGET_COMPACT` comment. Assertion, parametric over the module: for every `BUDGET_*` constant reaching an `_apply_budget` call there is a `*_budget_phrase()` and it appears in the rendered help. |
|
|
157
|
+
| `AUD-594-X03` / `X-03`, residual of `AUD-593-N01` | **P2** | **closed `5.8.19`** (`e89e8f0`) **as a class, not as the reported key.** `direct_caller_count_population` now ships from a single constant shared with its twin, and the sweep written to prove it found **four more catalogued figures shipping a unit and no population** — `chain_classes_total`, `impact.stats.direct_caller_symbol_count`, `explain.incoming_callers_count` and `modernize.in_degree` — none of them reported by either round. All five emit the catalogue's own tags through `declared_tags()`, which refuses an ambiguous tail (`direct_caller_count` is carried by two rows with different populations, which is the whole of `AUD-593-N01`) and leaves an already-published prose `_unit` untouched, so closing a population gap cannot silently change a shipped unit. Field-verified on halo. Mechanism as found: open, confirmed in source. `impact-chain` emits `direct_caller_count`, `direct_caller_count_unit: "distinct caller classes"` and a prose `direct_caller_count_note`, and **no `direct_caller_count_population`** (`spring_impact.py:1640-1654`). The twin beside it carries all three (`expanded_seed_caller_count{,_unit,_population}`), and the catalogue is already right: `caller_metrics.py:126-135` declares the row's population as `POP_EXPANDED_SEED_REFERENCES`. The product's published comparability rule is *"two figures are comparable only when both match"*; a consumer applying it over `_unit` still sees a match between figures that differ 16×. The separation exists only in prose and in a neighbouring key. | Emit `direct_caller_count_population` from the catalogue row itself. Assertion, derived from `FAN_IN_FIGURES` rather than from a hand-written list: for every catalogued key, the emitted payload carries `*_unit` **and** `*_population`, equal to `row["unit"]` and `row["population"]`. That single sweep covers this residual and the class `AUD-593-N05` closed by instance. |
|
|
158
|
+
| `AUD-594-N04` / `D-5`, disclosure half of `B-3` | **P2** | **closed `5.8.19`** (`a0060b7`). `sibling_view.omitted_blocks` is computed by running the other view constructor over the same `SourceMap` — never a hand-written list — with `omitted_blocks_basis` stating that an empty list means the two views agree rather than that the delta was not measured, and `blocks_only_in_this_view` publishing the other direction, because the two views are not a subset relation. `--compact` gains the reciprocal pointer it never had, so a reader who starts there learns the agent view exists. Recursion is prevented by `_emit_sibling=False` on the probe, and the probe is wrapped: a sibling pointer must never be why an answer fails to render. Cost measured: 7,5 µs per assembly, 0,40–0,45 s end to end on halo. Field-verified on halo: 13 omitted, 8 agent-only, published list equal to the measured set. Mechanism as found: open. `serializer.py:2688-2698` carries the comment *"Make the complete sibling view discoverable instead of making consumers infer where omitted structural blocks went"* and then emits `{view, command, reason}` with no block named. Measured both rounds: MSAS 16 blocks in `--compact` and absent from `--agent`, the bank 15 — `mybatis`, `transactional_boundaries`, `stacks`, `spring_profiles`, `deployment`, `env_map`, `project_summary` among them. `sibling_view` is also absent from `--compact`, so the pointer between views is one-way. Substantive rather than cosmetic: `security_surface` is missing from **both** channels on a repository with 368 security findings and 725 routes open by rule (disclosed under `AUD-590-B03`, not closed), and the channel the docs designate for agents is the poorer of the two in security signal. | Publish `omitted_blocks` inside `sibling_view` as `set(compact_keys) - set(agent_keys)`, computed from the two view constructors (`serializer.py:1708` and `:2350`) against the section registry at `:932`, never from a hand-written list, and emit the reciprocal `sibling_view` on the `--compact` side. Assertion: set equality between the computed delta and the published list. |
|
|
159
|
+
| `AUD-595-B01` / §B | **P2** | **closed `5.8.19`** (`b77a1ad`) — **each layer now says which question its word answers**, and the parse store publishes the measurement rather than the word. `cache_layers.basis` marks `cir` as *validated reuse* (its `warm` means this run's key hit) and `parse_store`/`snapshot`/`ris` as *presence before the run, not reuse*. `parse_store_reuse` publishes counted lookups — `hits`, `misses`, `reuse_ratio` — from counters in `parse_cache.get()` that are process-global rather than thread-local, because `--jobs` parses concurrently and the question is about the run. A corrupt entry counts as a miss: it costs a re-parse like any other. The block is **omitted when a run made no lookups** — a ratio over zero lookups is a number invented from nothing. Field-verified on halo: `parse_store: warm` with `990 hits / 0 misses / 1.0`, and on a run whose CIR hit, no `parse_store_reuse` at all. An existing assertion pinning `cache_layers` to exactly four keys was relaxed to a subset, in place, with the reason. Mechanism as found: open, and it has now caused three retracted measurements in three passes. `metadata.cache_layers` reports `warm` in states whose cost differs ×2.7: `spring-audit` on `openmrs-core` measured 4.17 s at `unaccounted_pct` 55 % and 2.48 s at 11.8 %, both with all four layers reporting `warm`, against 1.5 % genuinely cold. The remainder the product publishes is larger than every phase it measures put together, and `B-7`'s own `unit` string enumerates what it should contain — *"serialization, payload assembly, ceiling estimation, interpreter start"* — none of which explains 1.9 s. What is missing is cache **validation** against the tree. The instrumentation is right and nobody reads it. | Either distinguish *present* from *valid for this invocation* in `cache_layers` — a warm that re-validates 8 349 files is not the warm that serves from the snapshot — or instrument the validation as its own phase so it leaves `unaccounted`. Also worth a threshold: an `unaccounted_pct` above a declared bound is a measurement the product should refuse to present as clean. |
|
|
160
|
+
| `AUD-595-A05` | **P3 — new** | **closed `5.8.19`** (`e0b3d10`). `endpoints` publishes `scope` (path, `is_repository_root`, `nested_repositories`, basis), `repo_id` with its `repo_id_basis` — the same cache key `AUD-590-B01` established — and the **noun in `total_unit` now follows the tree that was walked**: *"this repository"* over a checkout, *"the analysed tree"* over a directory whose children are checkouts. The probe is depth-1 and structural (a checkout marks itself; build output and vendored trees are skipped) and branches on no name. Field-verified: halo → `is_repository_root: true`, 158 routes, unit unchanged; the 36-repository workspace → `is_repository_root: false, nested_repositories: 36`. The counts do not move — a scope fix that changed a number would be a different defect, and there is an assertion for that. Mechanism as found: open. `ask endpoints .` over a tree holding 16 repositories and 52 634 `.java` files publishes `total: 3 575` under `total_unit: "handler mappings declared in this repository — the whole route population"`, with `scope: null` and `repo_id: null`. `by_source` decomposes by technology, nothing decomposes by tree, and the unit asserts a repository. **It is not a truncation** — the auditor summed 2 802 endpoints over 10 of the 16 repositories against the 3 575 of the CWD, plausible without a cap and consistent with the uncapped `rglob` this ledger documents — so it is a scope-declaration defect only. Same class as the `scope` that `compare` gained in `5.8.18`: the fix exists in the product and did not reach the most-used core-tier command. | Reuse `data.setdefault("scope", {"path": str(target_dir)})` (`cli.py:8896`, `:9000`) at the `endpoints` call site, and either publish `repo_id` with its `repo_id_basis` (`AUD-590-B01`) or say *"in the analysed tree"* where the tree is not one repository. Sweep, not instance: every census that describes itself as *"this repository"* declares the scope it counted. |
|
|
161
|
+
| `AUD-595-A07`, narrow half of `AUD-592-A02` | **P3** | **closed `5.8.19`** (`cc15177`) — the same shape as `AUD-595-A02` and closed the same way: **two walks, two units, each naming which it is.** `scan_truncated` gains `cap_unit: "java files admitted to the analysis"` and a `cap_note` saying the help's directory-entry cap bounds the cheap sizing scan and is neither comparable nor the same limit; the sizing sentence now reads *"this sizing scan stopped at 25 000 directory entries … the analysis walk has its own cap, in files, published as `scan_truncated`"*. No measured figure moves, and there is an assertion for that. `ASK-13` still holds: a cap that did not stop the walk is not named. Mechanism as found: open. Two scan limits, two units, no cross-reference: the payload publishes `scan_truncated: {truncated: true, files_scanned: 8000, cap: 8000}` and the root `--help` describes *"the scan stopped at 25 000 directory entries"*. A reader of the help cannot relate `files_scanned: 8000` to 25 000 directory entries. | Declare both with their units and their relation, or name only the one that cuts first. |
|
|
162
|
+
| `E-39`, class residual of `AUD-593-N06` | **P3** | **closed `5.8.19`** (`25fc11f`), **and it closes criterion 3 of `B-1` with it — the fixture that was missing for four reports now exists.** Both lists route through `_cap_effect("display_list", …)` above five, and both keys are **published empty rather than omitted**: an absent key was indistinguishable from a field that was never computed, which is exactly why `mall`'s `{mapper_interfaces: 76, xml_files: 76}` could not be read from outside. Field-verified on `mall`: same 76/76, now with `orphan_xml: []` and `missing_xml: []` beside them. The `B-1` fixture is a generated repository with an XML mapper whose `namespace` names an interface that does not exist; it is asserted to appear in `orphan_xml` and **not** to be counted in `mapper_interfaces`. Two older assertions that required the keys to be *absent* when empty were updated to the new contract, with the reason recorded in place. Mechanism as found: open, found in source during this pass's triage, not reported by either round. `serializer.py:480-489` emits `orphan_xml = orphan_xml[:5]` and `missing_xml = missing_xml[:5]` with no `*_cap` block, in the same payload where the caps closed by `AUD-593-N06` now declare `total`/`shown`/`omitted`. Both keys are also emitted **only when non-empty**, so their absence is indistinguishable from *"not computed"* — which is exactly how this pass read `mall`'s `{mapper_interfaces: 76, xml_files: 76}` before reading the source. It is also why criterion 3 of `B-1` stays unverifiable from the field. | Route both through `_cap_effect(...)`, and publish the empty case explicitly rather than omitting the key — a declared zero is the contract this product enforces everywhere else. Fixture for the `B-1` criterion, still absent after four reports: a `Foo.xml` whose `namespace` names a mapper interface that does not exist, in a module under `@MapperScan`; it must appear in `orphan_xml` and must not be counted in `mapper_interfaces`. |
|
|
163
|
+
| `AUD-594-X04` | **P3** | **closed `5.8.19`** (`05b8cc9`) — the rows were corrected in place by this pass, and the durable half now ships as a release assertion: a row filed under a section headed *"closed in `X`"* must carry a closed verdict, **or the section's own preamble must name it as carrying forward**, so a half-closed row stays possible and a stale one does not. The preamble is read only from the lines *above* the table — prose after it names half the ledger, and counting that would make the assertion pass on anything, which was verified by mutation: reopening one closed row fails the battery. Its precondition `AUD-592-D04` is closed in the same release. Mechanism as found: open, and its precondition is `AUD-592-D04`. The packaged `DEFECT-LEDGER.md` in the 5.8.18 wheel marks `AUD-593-N01/N02/N05/N06` and both regressions `AUD-592-R01/R02` as `Current status: open`, and cites lines the shipped code no longer matches (`repository_ir.py:9248` → `:9268`). All six are fixed in that same binary, verified by behaviour in both rounds. Mitigating and real: the section header does say *"5.8.17 findings, closed in 5.8.18"*, so it is a header↔row incoherence rather than a flat falsehood, and `ask version` declares *"Rows closed after it shipped are not in this copy"* — but these closed **before** it shipped, so the disclaimer does not cover them. An auditor reads the row. | Derive each row's `Current status` from the same state as its section header, or drop the per-row field and keep the section's. This pass's own hygiene edit — closing the eighth-pass rows in place, above — is the manual form of that fix; the durable form is a release assertion that no row filed under a *"closed in `X`"* header ships with status `open`. |
|
|
164
|
+
| `AUD-595-Q01` | **P3 — process** | **closed `5.8.19`** (`91a6566`) — 7 rows → 10. `AUD-590-B01` (every command publishing `repo_id` publishes the same one, and it equals the cache key), `AUD-590-B02` (figures sharing a `(unit, population)` are equal **on this repository's own highest-fan-in symbol**, chosen from its endpoint census rather than from a fixture — which is the exact failure mode that closed it falsely twice) and `AUD-594-X03` (every catalogued figure emits both tags). Field-verified on halo: 3 of 3 pass, the fan-in row resolving to `ThemeEndpoint` and the tag row checking 5 catalogued figures. All three refuse with `not_applicable` and a reason where the repository offers nothing to query — a green run over nothing is how a defect gets closed twice, and there is an assertion for that over every row, not only the new ones. Mechanism as found: open. `selftest` still runs the same 7 rows (`selftest.py:415-442`: `ASK-09`, `R2`, `ASK-16`, `B6`, `E-37`, `B20`, `AS-16`), and **not one covers an invariant repaired in the last three releases**: `repo_id` cardinality 1 and its equality with the cache-dir key (`AUD-590-B01`), cross-command agreement between fan-in figures sharing `(unit, population)` (`AUD-590-B02`, `AUD-593-N01`), or the rule that every catalogued key emits `*_unit` and `*_population` (`AUD-594-X03`, `AUD-593-N05`). The command's own basis says *"Each row is the acceptance criterion its ledger entry publishes, run against this repository"*. What is not a `selftest` row is not verifiable in the field, and with `build_commit: unrecorded` it is not attributable to a commit either. Those two together are the mechanical reason this ledger's oldest classes keep returning. | Add one row per invariant above. They are the cheapest rows to add — the assertions already exist as internal regressions; what is missing is the field-facing surface. |
|
|
165
|
+
|
|
166
|
+
**Mapped, not new.** `AUD-592-D04` (`build_commit`) is re-witnessed for the fifth time and
|
|
167
|
+
its row above now carries the corrected mechanism: the reader shipped, the stamp never did.
|
|
168
|
+
`D-09` (`BUG-5`) cost anchors are re-counted at **17 on 5.1.0, one each on 4.18.0 and
|
|
169
|
+
5.0.0** in a 5.8.18 build, and they are now materially misleading rather than merely stale —
|
|
170
|
+
`./mall --agent` measures ×0.09 of its anchor and `spring-audit ./spring-boot` under budget
|
|
171
|
+
×0.43, so a reader sizing CI from `cache model` overestimates by a large factor; the row
|
|
172
|
+
belongs to `BUG-5` and its literal lives in the `AUD-594-X01` return. `AUD-593-N05` and
|
|
173
|
+
`AUD-593-N06` could not be re-measured on the bank (`via_interface_resolution` is not
|
|
174
|
+
emitted for any of the three symbols tried, including a real interface) and are closed on
|
|
175
|
+
Round A's MSAS evidence alone.
|
|
176
|
+
|
|
177
|
+
**Verified and not defects — do not spend effort.** `interface_mediated_callers_cap: null`
|
|
178
|
+
over a 23-element list (the cap emits above 30). `compare` without `scan_truncated` on a
|
|
179
|
+
3 342-file subject (below the 8 000 cap). `indirect_callers: []` in `impact-chain` (hub
|
|
180
|
+
guard, declared, with `depth_capped_by_guard` and `closure_complete: false`). `pr-impact` on
|
|
181
|
+
a clean tree (typed `INVALID_INPUT`, exit 1). `data-exposure` answering `answered: false`
|
|
182
|
+
without `dataLabels`. `endpoints` publishing `scan_truncated: null` over the 16-repository
|
|
183
|
+
tree (verified by summation, not assumed). `compare` returning no verdict — *"ranks by
|
|
184
|
+
measurable architectural cost and STOPS"*, and the payload holds to it: `cost_dimensions`,
|
|
185
|
+
`cost_units`, a `why` per candidate, no recommendation. `compare` without `--path` costing
|
|
186
|
+
89.71 s against 62.44 s in 5.8.16 is not a hidden regression: the earlier run returned
|
|
187
|
+
cross-repository candidates without declaring them and the 5.8.17 run crashed at 99.89 s,
|
|
188
|
+
where this one completes and publishes `scope` and `scan_truncated`; the canonical form with
|
|
189
|
+
`--path` costs 0.72 s.
|
|
190
|
+
|
|
191
|
+
**Process findings of this pass.**
|
|
192
|
+
|
|
193
|
+
1. **The sweep rule is holding where it was applied, and only there.** Four classes closed
|
|
194
|
+
as parametric sweeps have not recurred across three versions — write guards 8/8, numeric
|
|
195
|
+
bounds 7/7, error-path cost 8/8 at ≤0.79 s, selection flags 3/3 × 4 commands — and scan
|
|
196
|
+
truncation moved 1/4 → 4/4 by being read at each call site. Every row opened in this pass
|
|
197
|
+
is the residue of a *correct* fix that stopped at the instance: a help header the
|
|
198
|
+
generator did not reach, a `*_population` key, a `sibling_view` that names no block, a
|
|
199
|
+
`[:5]` beside a `[:30]` that was just fixed.
|
|
200
|
+
2. **Deleting a figure can open a hole.** `AUD-593-N02` was closed the right way — four
|
|
201
|
+
hand-written numbers replaced by a deference to the generated authority — and that is
|
|
202
|
+
exactly how `AUD-594-X02` was born, because the authority publishes nothing for
|
|
203
|
+
`--agent`. When a hand-written figure is deleted, assert that the authority it defers to
|
|
204
|
+
actually publishes one.
|
|
205
|
+
3. **A closure is only as good as the surface that can see it.** Six of this pass's rows
|
|
206
|
+
are invisible to `selftest` (`AUD-595-Q01`) and none of the ~200 measurements is
|
|
207
|
+
attributable to a commit (`AUD-592-D04`). Those two rows are cheap and they gate the
|
|
208
|
+
falsifiability of everything else in this ledger.
|
|
209
|
+
|
|
210
|
+
### Eighth Audit Pass: `5.8.17` findings, closed in `5.8.18` / MSAS + 12-repository bank / 2026-08-21
|
|
211
|
+
|
|
212
|
+
**Closed queue — every row below is `closed 5.8.18` and re-verified from outside by the
|
|
213
|
+
ninth pass, except `AUD-592-D02` (half) and `AUD-592-D04`, which carry forward.** Kept in
|
|
214
|
+
full because the mechanisms are the evidence, and because the direction column is what the
|
|
215
|
+
fixes were written against. Ordered as the correction order was: **regressions first, then the rows a
|
|
216
|
+
previous pass marked closed that the field does not reproduce as closed, then new contract
|
|
217
|
+
defects, then residuals.** Canonical IDs are `AUD-592-*` (bank round, reporter IDs `A-*`,
|
|
218
|
+
`B-*`, `N-*`, `D-*`) and `AUD-593-*` (MSAS round, reporter IDs `N-0*`). Rows held elsewhere
|
|
219
|
+
in this ledger are listed after the table as mappings, not as new defects.
|
|
220
|
+
|
|
221
|
+
| ID | Severity | Current status | Required direction |
|
|
222
|
+
|---|---|---|---|
|
|
223
|
+
| `AUD-592-R01` / `A-1` | **P0 — regression** | **closed `5.8.18`** (`216025f`) — re-verified from outside on the ninth pass: **3 of 3 valid invocations exit 0**, empty stderr, no traceback, in all three invocation forms (simple name, `-p`, FQN); `scope` is published, and `_compare_scan_trunc` now lives inside `compare_cmd` while `plan_cmd` owns `_plan_scan_trunc` bound to its payload. Mechanism as found, kept for traceability: open, confirmed in source, and it is two defects from one error. `compare` raises `NameError: name '_compare_scan_trunc' is not defined` on **5 of 5 valid invocations** — Rich traceback on stderr, zero bytes on stdout, no error envelope, exit 1 — ending a 240-invocation streak with no traceback. The assignment introduced by `AUD-591-A06` landed **inside `plan_cmd`** (`cli.py:8874`, beside `plan`'s own `find_java_files` at `:8873`) while the read is in `compare_cmd` (`cli.py:8972`); the two functions are `8818-8901` and `8903-9002`. Verified by reading the source at HEAD, not from the report. Second effect: `plan` computes its truncation state and **discards it** — nothing in the payload carries it, so `plan` answers `resolution: not_found` over a walk that stopped at its cap, offering candidates from a different repository. Third effect: `compare`'s other `AUD-591-A06` fix, `data.setdefault("scope", ...)` at `cli.py:8975`, is **after the line that raises** and has never executed, so the cross-repository candidate leak is still undeclared. | Read `_last_scan_truncation()` at `compare`'s own call site, immediately after `cli.py:8965`; rename the `plan_cmd` variable to `_plan_scan_trunc` and **bind it to the payload** — an unused walk-state variable is the defect, not the fix. Regression: an AST assertion that no `_*_scan_trunc` name is read in a `FunctionDef` that does not assign it, plus one execution test per command that owns a walk. A unit test that imported `compare_cmd` would not have caught this; one that *calls* it would. |
|
|
224
|
+
| `AUD-592-R02` / `N-4`, reopens `AUD-513-N04` | **P1 — regression** | **closed `5.8.18`** (`ec3050d`) — re-verified from outside on the ninth pass by driving the summariser directly: **8 of 8 repositories and 4 of 4 witnesses correct** (HTML licence comment, 30-dash rule, setext `===`, plus a bare Apache-2.0 boilerplate added by the auditor), and `summary_basis` moved from asserting prose to naming the section and the five filters it passed. Mechanism as found, kept for traceability: open, confirmed in source, on 3 of 8 repositories (37,5 %), in `--compact` and `--agent` alike. `project_summary` publishes an Apache-2.0 licence header as the description of the project (struts), a 30-character horizontal rule (halo) and a setext underline (shopizer) — and `summary_basis` asserts `"README descriptive prose"` over all three. In 5.8.16 this field was `null` with `"no descriptive section found in README"`, the literal acceptance text of `AUD-513-N04`; the field went from honest to false. Mechanism, four separate holes, all read at HEAD: HTML comments are detected **only on their opening line** (`summarizer.py:335-337`) with no `in_html_comment` state — three lines above, `in_code_block` does exactly that for fences (`:326-330`) — so `<!---` on line 1 flushes and lines 2-15 of the licence become the first paragraph; headings are ATX-only (`:331`), so setext titles underlined with `===` are content; the short-fragment floor is `len(paragraph) < 30` **strictly** (`:359`), so a rule of exactly 30 dashes survives; `_LICENSE_MARKETING_RE` (`:266-281`) matches product-tier and marketing phrasing and **carries no Apache/MIT/GPL/BSD boilerplate pattern at all**; and `summary_basis` is assigned unconditionally to whatever paragraph survives the filters (`:382-385`) — a restatement of the code path, not a property that was checked. | Add `in_html_comment` symmetric to `in_code_block`; recognise setext underlines as headings, not content; extend the licence regex with the four standard boilerplates; and make `summary_basis` describe the path actually taken — where nothing is verifiably descriptive, `null` plus the 5.8.14 text. **The mechanical cause of the reopening is that no README fixtures exist**: add the three witnesses from the bank as fixtures (HTML licence comment, 30-character rule, setext `===`) in the same commit as the fix. |
|
|
225
|
+
| `AUD-593-N01` / `N-01` | **P1** | **closed `5.8.18`** (`dd6a706`) **by declaration, not by making the numbers agree** — the right answer to a contract defect, and the ninth pass records it as the model: the figures still read 34 and 554 and did not have to change; `impact-chain`'s row gained its own population (`POP_EXPANDED_SEED_REFERENCES`), the payload carries `direct_caller_count_note` naming CH-001b and stating the two are not peers, and an honestly-named twin `expanded_seed_caller_count` ships with `_unit` and `_population`. **Residual for the programmatic consumer: `AUD-594-X03`** — the emitter publishes no `direct_caller_count_population`. Mechanism as found, kept for traceability: open, and it falsifies the closure assertion of `AUD-590-B02`. For one symbol, one tree, one HEAD, consecutive invocations, both at `depth=4`: `impact` publishes `direct_caller_count: 34` (210 reference sites) and `impact-chain` publishes `554` (3 103), both under the literal unit `"distinct caller classes"`. The product's own comparability rule (`CALLER_METRIC_RECONCILIATION`) is *"two figures are comparable only when both match"* — they match, and differ 16×. **The authority is what is wrong, not only the number**: `FAN_IN_FIGURES` (`caller_metrics.py:80-110`) declares both rows `unit=distinct_classes, population=references_excluding_imports`, and the populations are not the same. Measured mechanism: `impact-chain` expands the seed set with every member of the interfaces the target implements (CH-001b, `spring_impact.py:1110-1135`) and then takes depth 1 **from the expanded seeds** (`caller_reach.py:309-336`), so its "direct" includes callers that name a sibling implementation; `impact` admits only references that name the target and keeps interface-mediated callers on a separate axis (`repository_ir.py:9263`, `:9596-9606`). Corroborating symptom: `chain_classes_total`, declared *"unbounded above by any direct-caller count"*, equals `direct_caller_count` exactly (554), and `indirect_callers` is `[]` — that emptiness is the hub guard capping the walk to depth 1 (`caller_reach.py:319-321`) and **is** declared, so it is not a confident zero. The published closure assertion `len(impact.direct_callers) == chain.metadata.direct_caller_count` measures 30 (list, capped from 34) against 554: it passes on a fixture where the two coincide by accident. | Decide which figure the `direct_caller_*` namespace owns, and fix the **catalogue** either way: either `impact-chain` measures depth-1 fan-in on the unexpanded seed (≈34), or the figure is renamed out of that namespace and its `population` becomes its own tag — an interface-expanded seed set is not `references_excluding_imports`. Regression is parametric over the catalogue, not over `CommonService`: for every pair of `FAN_IN_FIGURES` rows sharing `(unit, population)`, the emitted values must be equal, run over ≥3 symbols of different fan-in **including one that trips the hub guard** — the condition under which the divergence appears, and the reason two previous corrections of this class (`callers_total`, `P0-2`) closed a name and left the divergence. `chain_classes_total ≥ direct_caller_count` with strict inequality wherever uncapped transitive reach exists. |
|
|
226
|
+
| `AUD-592-B01` / `B-8` | **P1 — closed row that does not reproduce as closed** | **closed `5.8.18`** (`8329614`) — re-verified from outside: `total: 3, shown: 1, omitted: 2`, and `--limit 10` still returns all three. Mechanism as found, kept for traceability: open; the ledger row overstates its own fix. `B-8` is recorded as *"compact posture retains one unresolved sample when the total is non-zero"* (`52593b6`). The code retains a sample only when the total is **exactly one**: `_posture_limit = 1 if summary["unresolved"] == 1 else 0` (`cli.py:12653-12654`). Field measurement on `mall`: `summary.unresolved: 3`, `unresolved: []`, `unresolved_cap: {total: 3, shown: 0, omitted: 3, limit: 0}` — and it is the only cap in that payload with `limit: 0` while `undecided_cap` and `evidence_cap` carry 200. `--limit 10` returns all three, so the analysis is intact and only the compact contract is wrong. | Either make the code do what the row says (retain ≥1 sample whenever `summary.unresolved > 0`) or correct the row to the narrower promise and reopen the defect it leaves. Do not leave `closed` standing over `shown: 0`. Add an assertion that no `cap_effect` receives `limit=0` by default on a path where `--limit 0` is published as *"no cap"* (`AUD-513-N08`), so the published semantics and the internal default cannot invert each other. |
|
|
227
|
+
| `AUD-592-A02` / `A-1` second half, `§12.5` | **P2** | **closed `5.8.18`** (`63f3d61`) in its wide half — **4 of 4**: `plan`, `impact-chain` and `compare` publish `scan_truncated`, and `endpoints` answering `null` is correct rather than missing (verified, not assumed: the CWD census of 3 575 against 2 802 summed over 10 of 16 repositories is plausible without truncation, consistent with the uncapped `rglob` this ledger already documents). **The narrow half stays open as `AUD-595-A07`**: the root help's 25 000 directory-entry cap is still unreconciled with the 8 000-file cap the payload publishes. Mechanism as found, kept for traceability: open — scan truncation is declared by 1 of 4 commands that own a walk. Over a truncated scan (cap observed at **8 000 files**), `impact` publishes `scan_truncated`; `plan` and `impact-chain` answer `not_found` with `scan_truncated: null`, and `endpoints` exits 1 with no payload — every one of them an exclusion derived from an incomplete population. `compare` cannot be measured (`AUD-592-R01`). Second, narrower defect in the same row: the root `--help` describes the cap as *"the scan stopped at 25 000 directory entries"* — a different limit in a different unit from the 8 000-file cap the field observes. | Publish the truncation beside every walk that can end in `not_found`, `candidates` or a census — the seventh pass already recorded that `extract_java_endpoints` does its own uncapped `rglob`, so this must be read at each call site, never centrally. Reconcile the two caps in the help text or declare both with their units. |
|
|
228
|
+
| `AUD-593-N02` / `N-02` | **P2** | **closed `5.8.18`** (`f539787`) **by deleting figures, not adding them** — the four hand-maintained numbers in the packaged guide became *"See `ask --help`"*; zero token bands remain in the shipped `.md`, and the measured payload (22 391 B ≈ 5 597 est. tokens) sits inside the generated ceiling. **The fix opened `AUD-594-X02`**: `--agent` deferred to an authority that publishes no figure for it. Mechanism as found, kept for traceability: open — the `--compact` budget is published three times, in two shipped authorities, with three values. `ask --help` generates *"~7K tokens (30 KB, chars/4)"* from `output_budget`, the module that enforces the limit — this half is the `AUD-590-B03` fix and it is correct. The packaged `docs/USER_GUIDE.md` is still hand-maintained and disagrees with it twice: `:111` *"A ~10K-token subset"* and `:243` / `:1366` *"~2,500–4,000 tokens"*; `--agent` carries a fourth figure at `:254` / `:1367` (*4,500–5,500*). Measured payload: 22 390 B ≈ 5 597 est. tokens — inside the help ceiling, **+40 % over the guide's table**. `grep '7K tokens' *.py` finds no literal: the help is generated and the guide is not, and that asymmetry is the defect. | Generate every token figure in the packaged docs from `output_budget.token_estimate`, or delete them and point at `--help`. Extend the existing release assertion (today help ↔ payload ↔ section registry) to cover `docs/*.md`: a token band in a shipped `.md` that does not come from that module fails the battery. Packaged documentation is a published authority and belongs inside the same assertion as the help text. |
|
|
229
|
+
| `AUD-593-N03` / `N-03` | **P2** | **closed `5.8.18`** (`1dd0b1c`) **in the generated half only** — the Options line is now built from `_BUDGETED_ANALYSIS_COMMANDS` (`cli.py:589-596`, read at `:632`) and enumerates all eight commands, wider than this row itself described. **The hand-written half is open as `AUD-594-X01`**: three literals in two other authorities still name three. Mechanism as found, kept for traceability: open — `--help` denies the scope that `AUD-590-B04` had just granted. With `ASK_MAX_ANALYSIS_SECONDS=90` the root analysis prints *"ask is a repo-wide analysis and the configured budget is 90s. Budget read from ASK_MAX_ANALYSIS_SECONDS=90 (process environment)"* — the fix works. The help page beside it still reads *"deadline only for spring-audit, risk, audit-report"* and *"Other commands are not time-bounded by this variable"* (`cli.py:621-622`). The root analysis is the most expensive command in the product and the one an agent most needs to bound; a reader of the help will not set the variable there. | Generate that line from the same registry `_build_analysis_classes` populates (the root is keyed there as `B14`), never by hand. Assertion: the set of commands the help names equals the set the preamble emitter covers. |
|
|
230
|
+
| `AUD-593-N05` / `N-05` | **P3** | **closed `5.8.18`** (`aa89388`) — the payload now carries `caller_count_unit: "caller_symbols"` and `caller_count_population: "references_to_target_or_its_declared_interfaces"` with a `FAN_IN_FIGURES` row behind them. Mechanism as found, kept for traceability: open — a figure emitted outside the catalogue that forbids it. `via_interface_resolution[].caller_count` publishes 37 where every other surface of the same fact says 23 (`interface_mediated_caller_count`, the list length, and the `explanation` prose), and it carries no `_unit` while the rest of the payload does. Both are legitimate and neither is declared: `caller_count` is `len(_iface_callers)` — raw caller symbols of the interface, before the `not in all_affected` filter and before normalisation to classes (`repository_ir.py:9248-9251`); the 23 is `caller_classes(_iface_mediated_callers, ...)` (`:9603`, `:10173`). It is also **not in `FAN_IN_FIGURES`**, which states that a figure absent from it may not be emitted — so `tests/test_fan_in_authority.py` did not catch a fan-in figure added in this release. | Give it a unit and a `FAN_IN_FIGURES` row, or rename it to `interface_caller_symbols`. Make the authority test sweep the **emitted** keys against the catalogue rather than the catalogue against itself; a catalogue that only validates its own rows cannot see an unregistered emission. |
|
|
231
|
+
| `AUD-593-N06` / `N-06` | **P3** | **closed `5.8.18`** (`aa584a6`) — `interface_mediated_callers_cap` is routed through `_cap_effect(...)` and emits `null` below the threshold, which is the correct behaviour at MSAS's 23. **Closed by instance, and the class survived it: `E-39`** — `serializer.py:480-489` still truncates `orphan_xml` and `missing_xml` at `[:5]` with no cap block. Mechanism as found, kept for traceability: open, latent — a silent `[:30]` where every sibling list declares its cap. `out["interface_mediated_callers"] = _iface_mediated_classes[:30]` (`repository_ir.py:10172`) with no `*_cap` block, in a payload where `direct_callers_cap`, `indirect_callers_cap` and `security_surface_affected_omitted` all carry `total`/`shown`/`omitted`/`direction`/`how_to_read`. On MSAS the list is 23, so nothing is visible; the first repository with more than 30 interface-mediated callers loses the excess without a word. | Route it through the same `_cap_effect(...)` used 27 lines above (`repository_ir.py:10145`). |
|
|
232
|
+
| `AUD-592-D01` / `D-1` | **P3 — residual of `AUD-591-A08`** | **closed `5.8.18`** (`0f8e03f`) — the root inventory keys on the triggering flag: `<repo>/.ask/readiness-history/` — only with `--snapshot` (optionally relocated by `--history-dir`). Mechanism as found, kept for traceability: open, confirmed in source. The write inventory in the root help says `migrate-check` writes *"only with `--history-dir`"* (`cli.py:296`). `--history-dir` only relocates the destination; the flag that writes is `--snapshot`, as `migrate-check --help` states correctly and as the code comment at `cli.py:12980-12986` states explicitly. Verified in the field: `migrate-check --history-dir <path>` exits 0 with zero writes and no guard message, while `--snapshot` fires the guard correctly. `AUD-591-A08`'s assertion checks that each cited flag **exists** in that command's parser — `--history-dir` does — not that it is the flag that triggers the write. | Key the inventory on `(command, triggering flag, path)` taken from the guard site itself, as `AUD-591-A08`'s own acceptance text required. An auditor tests what the inventory lists: this is the same mechanism by which `AUD-591-A01` went unfound. |
|
|
233
|
+
| `AUD-592-D02` / `D-2` | **P3 — residual of `AUD-591-A02`** | **closed `5.8.19`** (`1d3db4d`) **for both halves, from one module.** `--min-band` is a floor applied while the answer is composed, not a post-filter over rows, which is why it was left out of `_apply_selection` and why the exception cost what the rule was written to prevent. `selection_contract.min_band_filter_block()` now owns the disclosure: it merges rather than replaces (so `--band` and `--min-band` can both be in play), it returns nothing at the permissive floor (a payload that excluded nothing must not claim a selection), and it carries a basis saying the floor is not a row filter. `risk` publishes it from `_assemble_payload`, the one construction both its callers use. Field-verified on halo: `--min-band high` → `{min_band: high, total_before_filter: 0, matched: 0, shown: 0}`; default → `null`. Regression sweeps the commands offering the flag and fails any module that hand-builds the block. Mechanism as found: open in half, and the half that closed shows the fix the other needs. `f78522a` gave `enrich --min-band` its `_filter` block (`{min_band, total_before_filter, matched, shown}`, `cli.py:11666-11673`); `risk --min-band` still publishes `_filter: null` while `risk --band` publishes one — measured on `mall`: `--band high` → `_filter` with `total_before_filter: 50`, `--min-band medium` → `null`. The two call the same helper 250 lines apart and only one adds the block afterwards (`cli.py:11406-11413` against `:11658-11673`). Mechanism as found, kept for traceability: `--min-band` filters and validates (`cli.py:11294-11308` in `risk`, `:11563-11568` in `enrich`) but publishes no `_filter` block, while its sibling `--band` publishes one through `_apply_selection`. A consumer cannot tell a filtered answer from a small one. | Route `--min-band` through `_apply_selection` like every other selection flag, or publish the same `_filter` shape from wherever it is applied. |
|
|
234
|
+
| `AUD-592-D03` / `D-3` | **P3 — residual of `AUD-591-A02`** | **closed `5.8.18`** (`f28fe40`) — `enrich --rule NOPE-999` exits 1 against the same catalogue `spring-audit` validates against. Mechanism as found, kept for traceability: open. `enrich --rule NOPE-999` exits 0 with an empty result; `spring-audit --rule NOPE-999` exits 1 against the rule catalogue. One flag name, two contracts. Measured with a synthetic SARIF 2.1.0 (2 results): `--rule`, `--band` and an invalid `--band` all behave correctly in `enrich`; only unknown-rule validation is missing. | Validate against the same catalogue in both commands. `enrich` is experimental-tier, which sets the severity, not the contract. |
|
|
235
|
+
| `AUD-592-D04` / `build_commit` | **P3 — process** | **closed `5.8.19`** (`6eba05a`) **on the writer side, which is the side that was missing.** A `hatchling` build hook now writes `src/sourcecode/_build_commit.py` inside the build, where the checkout still exists, and the reader consults three sources narrowest first: `ASK_BUILD_COMMIT`, the stamped module, then the checkout — with `unrecorded` in the stamp treated as an absence that falls through rather than as an answer. The logic lives in `sourcecode.build_stamp`, importable without the build backend, because **a build-time behaviour the suite cannot exercise is how this row survived five reports**: the previous fix was asserted through a monkeypatched environment variable and nothing else. The hook never fails a build — an exported tarball has no git, and a release that cannot be built is worse than one that cannot name its commit. Verified end to end: stamp written, `ask version` publishes the commit, generated file gitignored so no checked-in copy can drift. Mechanism as found: open, fifth consecutive report, and the fix landed on the wrong side of the build boundary. `6d8eb1d` closed the *reader*: `product_info._build_commit()` takes `ASK_BUILD_COMMIT` and falls back to `git rev-parse HEAD` of the package checkout. **Nothing stamps it.** No `hatchling` hook, no release script and no CI step sets `ASK_BUILD_COMMIT` (`grep -rn ASK_BUILD_COMMIT` finds the reader, and a test that monkeypatches the variable — nothing that writes it), and an installed wheel has no git checkout, so the fallback returns `unrecorded` on every field install. The battery passed because it asserted the reader. Field verdict on 5.8.18 is unchanged: `build_commit: unrecorded`. Mechanism as found, kept for traceability: open, fourth consecutive report. `ask version` publishes `build_commit: unrecorded`, so no measurement in these ~150 invocations is attributable to a commit and *"fixed in 5.8.17"* is not falsifiable from outside. | Stamp the commit at build time. It is the cheapest row in this table and it is the precondition for every other row's verification. |
|
|
236
|
+
|
|
237
|
+
**Mapped, not new — inherited rows this pass re-witnesses.** `B-6` (partial answers exit 0)
|
|
238
|
+
keeps its row and gains its best measurement yet: overrun ×5.4 → ×4.2 → **×1.39** (median
|
|
239
|
+
20.90 s of 70.74/20.90/19.70 against a 15 s budget), with a `_partial` block that is
|
|
240
|
+
exemplary (`phases_completed`, `phases_pending`, `why_stopped`, and *"every count here is a
|
|
241
|
+
FLOOR"*) — and `exit 0` on all three runs, so a pipeline reading `$?` passes with two of
|
|
242
|
+
three phases unexecuted. **`AUD-591-A09` is the template**: `exit_code` / `gate_exit_code` /
|
|
243
|
+
`gate_exit_code_basis`, or an explicit `--allow-partial`. `B-3` (`--agent` versus
|
|
244
|
+
`--compact`) gains a witness at 3 342 files: 35 357 B ≈ 8 839 est. tokens, **+61 % over the
|
|
245
|
+
only published band**, 16 blocks present in `--compact` and absent from `--agent`
|
|
246
|
+
(`transactional_boundaries`, `mybatis`, `spring_profiles`, `deployment`, `env_map`,
|
|
247
|
+
`stacks` among them), and `sibling_view` still names no block; on the bank round the two
|
|
248
|
+
views are no longer a subset relation at all (9 blocks agent-only, 15 compact-only) and the
|
|
249
|
+
×2.25 time overhead is gone (×1.02). Substantively, `security_surface` is absent from
|
|
250
|
+
**both** channels on a repository with 368 security findings and 725 routes open by rule —
|
|
251
|
+
the predicate is disclosed under `AUD-590-B03`, not closed. `N-5` (an advisory saying
|
|
252
|
+
*"8 372 Java files"* with no population label while the progress line says
|
|
253
|
+
`[population=production_java_sources]`) remains the unclosed half of `AUD-588-B11`; the
|
|
254
|
+
label exists (`cli.py:10278-10280`) and reaches the progress lines and the payload but not
|
|
255
|
+
the advisory emitted from `phased_run.py:367`. `D-09` (cost anchors stamped *"on 5.1.0"* in
|
|
256
|
+
a 5.8.17 build — now counted: **17 anchors on 5.1.0, one each on 4.18.0, 5.0.0 and 3.2.2**,
|
|
257
|
+
in a `cache model` output presented as operational guidance) belongs to `BUG-5`. `N-6`
|
|
258
|
+
(cold `spring-audit` on `openmrs-core`) belongs to `BUG-1` and **improved**: 8.86 s → 8.06 s.
|
|
259
|
+
|
|
260
|
+
**Closed by correction and re-verified from outside in this pass — do not re-derive.**
|
|
261
|
+
`AUD-591-A01` (8 of 8 write paths refuse, verified by filesystem footprint rather than
|
|
262
|
+
`git status`, `ASK_RUNS_IN_REPO=1` included), `AUD-591-A02`, `AUD-591-A03` (and the
|
|
263
|
+
proposed highest-priority probe is retired: `rename-class --from O` exits 1, as the seventh
|
|
264
|
+
pass had already established), `AUD-591-A05`, `AUD-591-A06` (50.30 s → 0.43 s, and the
|
|
265
|
+
sweep it asked for holds: **6 of 6 error paths ≤0.54 s**), `AUD-591-A08` (modulo
|
|
266
|
+
`AUD-592-D01`), `AUD-591-A09`, `AUD-591-A10`, `AUD-590-B01` (`repo_id` cardinality 1 with
|
|
267
|
+
`repo_id_basis`; the content hash lives on as `content_id`), `AUD-590-B02` (the *lists* —
|
|
268
|
+
the counters above it are `AUD-593-N01`), `AUD-590-B03` (in `--help`; the packaged guide is
|
|
269
|
+
`AUD-593-N02`), `AUD-590-B04` (both halves: a `partial-analysis-v1` stub with `partial:
|
|
270
|
+
true` where 5.8.16 left zero bytes, and an 807 B preamble byte-compatible with
|
|
271
|
+
`spring-audit`'s 813 B), `AUD-590-R01` (`route_census` — and it **refuted the auditor's own
|
|
272
|
+
double-counting hypothesis with the product's data**: `mappings 3763` against
|
|
273
|
+
`on_rows 1120`, with `distinct_routes ≥ mappings`), `AUD-591-Q01`, `B-7` (closed by
|
|
274
|
+
deletion, which was the right answer: `total_ms` is gone, replaced by
|
|
275
|
+
`wall_ms`/`measured_ms`/`unaccounted_ms`/`unaccounted_pct` with arithmetic exact to the
|
|
276
|
+
decimal and a `unit` explaining why the remainder is published rather than distributed).
|
|
277
|
+
|
|
278
|
+
**Verified and not defects — do not spend effort.** `posture --profile` accepting an
|
|
279
|
+
undeclared profile is **by design and correctly declared**: it is a hypothesis input, not a
|
|
280
|
+
catalogue name, and the payload publishes `profiles_requested` beside `profiles_declared`
|
|
281
|
+
(the bank round withdrew this finding itself after reading the payload). `pr-impact` on a
|
|
282
|
+
clean tree exits 1 with a typed `INVALID_INPUT` naming the empty diff. `pr-impact` over
|
|
283
|
+
three classes refuses with `OUTPUT_TOO_LARGE` and names `--output` — a ceiling as contract,
|
|
284
|
+
not a truncation. `verify` publishes `exit_code: 2` matching the process (`AUD-591-A09`).
|
|
285
|
+
`data-exposure` still answers `answered: false` rather than inferring labels. Cache
|
|
286
|
+
freshness reporting `STALE` at delta 0 commits is `Uncommitted: True` over the auditor's own
|
|
287
|
+
untracked files. `repo_id` absent from `endpoints`, `posture` and `--compact` is coverage,
|
|
288
|
+
not a defect — the criterion was cardinality 1 among emitters, and the help no longer
|
|
289
|
+
promises it.
|
|
290
|
+
|
|
291
|
+
**Process findings of this pass, and they are the same finding twice.**
|
|
292
|
+
|
|
293
|
+
1. **The class survives the fix, four times out of five.** `AUD-590-B02` (unit absent from
|
|
294
|
+
the lists) returned as `AUD-593-N01` (unit present and false in the counters) and
|
|
295
|
+
`AUD-593-N05` (a new field shipped without one); `AUD-590-B03` (two figures in one help
|
|
296
|
+
page) returned as `AUD-593-N02` (three figures once the packaged guide is counted) and
|
|
297
|
+
as the `--agent` band in `B-3`; `AUD-590-B04` (a variable invisible to the root) returned
|
|
298
|
+
as `AUD-593-N03` (a help page denying the scope the fix granted). The seventh pass named
|
|
299
|
+
this pattern its highest-value finding and wrote the rule — *every acceptance criterion
|
|
300
|
+
carrying a sweep clause is closed with a parametric test over the catalogue* — and the
|
|
301
|
+
two classes closed **as sweeps** in that pass (write guards 8/8, numeric bounds 7/7, and
|
|
302
|
+
now error-path cost 6/6) did not recur, while three of three closed by instance did.
|
|
303
|
+
`FAN_IN_FIGURES`, `COMMANDS_THAT_WRITE` and the serializer's section registry are already
|
|
304
|
+
catalogues: the test goes over the catalogue, never over the reported symbol.
|
|
305
|
+
2. **Two rows in this ledger were marked `closed` while the field reproduces them open**
|
|
306
|
+
(`AUD-592-B01`, and `AUD-588-B11`'s advisory half). Both were closed against a surface
|
|
307
|
+
narrower than the row's own wording. A closure whose assertion cannot fail on the tree
|
|
308
|
+
that produced the defect is not a closure: `AUD-593-N01`'s assertion passes today only
|
|
309
|
+
because the fixture makes both figures coincide.
|
|
310
|
+
3. **Packaged documentation is a published authority.** `AUD-593-N02` exists because the
|
|
311
|
+
release assertion covers help ↔ payload and stops at the package boundary. Two stale
|
|
312
|
+
figures in a shipped `.md` are enough to contradict the correct one in the help.
|
|
313
|
+
|
|
76
314
|
### Seventh Audit Pass: `5.8.16` / 16-repository bank + `saint-server` / 2026-08-21
|
|
77
315
|
|
|
78
316
|
Ordered as the correction queue, regressions and permissive-direction failures first. Each
|
|
@@ -184,7 +422,7 @@ partial, `F03` open, `F04` explicit opt-in and `F05` locally verified only.
|
|
|
184
422
|
| `N-6` | Medium | **not reproduced after cache fixes** | Isolated cold `openmrs-core`: 5.18 s wall / 4.76 s internal, all four layers cold. Retain the field report as an environment-specific reproduction request. |
|
|
185
423
|
| `B-3` | Medium, core/supported | **closed by `773268d`** | Help no longer claims maximum signal; `sibling_view` points to compact. |
|
|
186
424
|
| `B-6` | Medium, core | **closed by `17c6e06`** | Partial answers publish `overrun_bound` and its wall-clock basis. |
|
|
187
|
-
| `B-8` | Low | **closed by `52593b6
|
|
425
|
+
| `B-8` | Low | **reopened 5.8.17 — see `AUD-592-B01`** (was: closed by `52593b6`) | The fix retains a sample only when the total is exactly 1, not whenever it is non-zero, so `total: 3, shown: 0` still reproduces. |
|
|
188
426
|
| `N-5` | Low, core | **closed by `c68083f`** | Progress and payload name the Java population. |
|
|
189
427
|
| `N-7` | Low, core | **closed by `082e9de`** | Bounded fuzzy candidates and matching message hints handle close typos. |
|
|
190
428
|
|
|
@@ -207,7 +445,7 @@ automation remains deferred; provider non-coverage is the current product behavi
|
|
|
207
445
|
| `AUD-513-N01` (`N-1`) | **closed in 5.8.14** | `rename-class` rewrote three Petclinic files under both `ASK_READONLY=1` and `--no-write`, then exited 0. The audit restored the tree; this is a write-policy breach, not an artifact-directory exception. `readonly.guard_mutation()` now refuses before any planned source write or physical rename; help names the command as mutating and regressions cover both readonly controls plus `--dry-run`. |
|
|
208
446
|
| `AUD-513-N02` / `B-4` / `N-5` | **N02/B-4 implemented in 5.8.14; N-5 open** | `posture.summary.unconditional` counted 1,697 on a non-Spring Struts tree and 1,573 on Mall, and `security` was `0/0/0` with no population. The posture projection now uses the canonical Spring IoC bean population, publishes its population/unit and separates unconditional security beans. Regressions cover non-Spring, mixed-bean and conditional/unconditional security fixtures. Progress/advice still use incompatible “Java files” scopes (`N-5`), so the `AUD-588-F02` family remains open. |
|
|
209
447
|
| `AUD-513-B01` / `B-3` / `B-5` | **B01/B-5 implemented in 5.8.14; B-3 closed as revalidated** | A MyBatis XML paired with a `*Mapper.java` interface now takes precedence over an absent `@Mapper` annotation, preserving XML-backed/`@MapperScan` interfaces in `mapper_interfaces` rather than misclassifying them as DTO mappers. `imports_found` now contains only actual Java imports; code, XML and build matches are emitted as `evidence`, including in compact output. `B-3` was stale against the current source: both views consume `JAVA_SPRING_SECTIONS`; a real MyBatis Mapper/XML regression proves the section is retained by `--agent`. |
|
|
210
|
-
| `AUD-513-N03/N04/N07/N08/N09`, `B-6/B-7/B-8`, `D-09` | **N03/N04/N07/N08/N09 implemented in 5.8.14; remainder open** | `impact` now always publishes the `candidates` field its not-found message names, including `[]` when no close symbol exists. Budget advice no longer calls partial snapshot/RIS evidence a warm cache: it identifies that presence and says context/parse are unknown before analysis, whose payload then reports all four actual layers. README operational sections (`clone`, install, build, getting started, troubleshooting and related headings) are never promoted into a project summary; with no descriptive prose compact output publishes `project_summary: null` and `summary_basis
|
|
448
|
+
| `AUD-513-N03/N04/N07/N08/N09`, `B-6/B-7/B-8`, `D-09` | **N03/N04/N07/N08/N09 implemented in 5.8.14; remainder open** | `impact` now always publishes the `candidates` field its not-found message names, including `[]` when no close symbol exists. Budget advice no longer calls partial snapshot/RIS evidence a warm cache: it identifies that presence and says context/parse are unknown before analysis, whose payload then reports all four actual layers. README operational sections (`clone`, install, build, getting started, troubleshooting and related headings) are never promoted into a project summary; with no descriptive prose compact output publishes `project_summary: null` and `summary_basis` — **`N04` regressed in 5.8.17 and is `AUD-592-R02`: licence headers and layout rules are promoted again, under a basis that asserts they are prose.** `security_posture.limitations` is now cache-state invariant; dynamic execution advice travels in `security_posture.operational_hints`. Every numeric output bound rejects negative values at parse time with a structured `INVALID_INPUT` envelope (`flag`, `value`, `expected`); `--limit 0` still means no cap wherever that behavior was published. Budgets overrun, timing coverage is inconsistent, compact hides unresolved identities and cost anchors remain stale. |
|
|
211
449
|
| `AUD-588-B11` (container/units subset) | **partially implemented in 5.8.14** | `d745d31` discovers nested Maven modules; `273ae8e` suppresses unsupported `pr-impact.unaffected_basis`. `e53507d` separates direct impact matches from analysis seeds and implementation classes; `8386744` carries the established container-wiring fact into plan checklists and compare cost dimensions. Source population units and risk-tier wording remain open. |
|
|
212
450
|
| `AUD-511-R02` / `AUD-589-B02` / `AUD-513-N06` | **partially corrected; `data-exposure` open, P1** | Nine of ten measured commands recovered, many to their best historical timings. `data-exposure --config` did not: 42.7 s versus 16.2 s in 5.8.8 and a 28.9 s model, with byte-equivalent output. Measure it with the idle-host, isolated-cache, interleaved-control protocol before changing code, then compare its reuse path with `validation`. |
|
|
213
451
|
| `AUD-588-B02`, `B06`, B11 residuals, `B12` | **B02/B06 closed; B11/B12 open, P2** | `3869c87` excludes explicit EclipseLink from Hibernate applicability and effort; `037223e` publishes provider non-coverage. The re-audit verifies both and the payload reduction; `cf57457` makes aggregate wording name only applicable dimensions. Complete B11 consumer parity and calibrate B12 only from provenance-bearing controlled samples. |
|
|
@@ -407,6 +645,8 @@ retained only as evidence of what the auditor observed before the fix.
|
|
|
407
645
|
| ASK-17 | **The parse store's default budget cannot hold a multi-repository workspace warm, and the cold cost of the largest repository is 7,6 minutes.** Measured on the 8-repo corpus (43 986 `.java`): `cache status` reports 56 730 entries at **511,12 MB against a 512 MB budget** — i.e. permanently sweeping by LRU — so `tutorials` (24 073 `.java`, the largest contributor) is the first candidate for eviction. Cold `--compact` on it: **457,6 s**, independently reproduced at 449 s, against 2,0 s on the second pass. | 5.8.4 corpus re-audit | Medium — warm plus `--compact` is still the answer (2,0 s), but a workspace this size cannot keep every repository warm at the default budget, and nothing tells the caller which repository is cold before it pays for it | **closed 5.8.5 — the disclosure shipped and the owed measurement taken, under our own control.** ⚠ **The +66 % does not reproduce, and the mechanism the row suspected is worth ~3 %, not 66 %.** Protocol: `tutorials` (24 073 `.java`), root `--compact`, both cache bases isolated **and emptied between runs** (`SOURCECODE_CONTEXT_CACHE_DIR` *and* `SOURCECODE_CACHE_DIR` — the first attempt isolated only the first, and the second pass of each version answered from the L2 view its own first pass had written: 2,4 s with an empty parse store, which is how a measurement of a cold path becomes a measurement of a warm one), `ASK_PARSE_CACHE_MAX_MB` fixed at 4 096 MB so no sweep can confound the comparison, 2 passes per version. **5.7.2 (`e803c8b`): 112,0 s / 112,6 s. This tree: 113,8 s / 115,2 s — +2,3 %**, with per-version dispersion of 1,005× and 1,012× and a payload 2,2 % larger. Then the audit's own condition, isolated as the only variable — the store pre-filled to 489 MB against the **default** 512 MB budget, so every 32 MB written sweeps: **117,0 s, +2,7 %.** So the LRU sweep is not where 457,6 s comes from, and neither is the analysis path: the auditor was right to hold the regression, and the held figure is now retracted **with a number** rather than on suspicion. ⚠ Scope, stated because the difference is unexplained rather than explained away: 457,6 s on Windows 11 / pipx is **not reproduced here** (113 s on 20 cores), and that gap is not assertable in either direction from this measurement — what is assertable is that 5.7.2 → this tree did not get slower and that a saturated store costs ~3 %. ✅ **What the measurement does confirm is the row's own claim, with our number: one repository of 24 073 `.java` leaves 46 840 entries and 470 MB in the store — 92 % of the 512 MB default budget** — so a workspace with a second repository of any size is permanently sweeping by construction, exactly as reported. Sizing rule, measured rather than guessed: ~20 KB of store per Java file, so ~500 MB per 24 000-file repository. Collateral confirmation of `AS-18`: after the saturated run the store rests at 669 MB — 157 MB over the budget and **under** the 736 MB effective ceiling it publishes at 7 writers — so the ceiling holds under the condition that produced the complaint. **The disclosure half:** `cache status` now publishes **coverage per repository**, which is the half the row itself recommended and the half that changes a decision: *«tutorials: 3 100 of 24 073 files cached (13 %)»* replaces a blind guess about whether to raise the budget, and an LRU eviction becomes visible **before** somebody pays 457,6 s to discover it. Both halves of the attribution were already held and thrown away — the walk knows the repository and it computes the store key for every file — so the run records the pairs (`parse_cache.record_repository_index`, from **both** readers of the store: `build_repo_ir` and the route-surface extractor, so the figure does not depend on which command was typed) and `store_stats` intersects them with the keys it collects **in the entry walk it already performs**: coverage costs an intersection, never a second scan of the store and never a scan of the repository. Three rules keep it honest. The index is **merged, never replaced**, because a `--changed-only` or `--since` run would otherwise shrink a 24 000-file population to the twelve files it read and publish *«12 of 12 cached (100 %)»* about a repository that is cold. The number travels with its **basis** — a file deleted since its last analysis still counts as recorded and reads as uncached, which errs toward *colder than it is* and says so. And the index is **neither an entry nor evictable**: the entry walk, the byte accounting and the LRU sweep all glob `*.json`, so its bytes are published under their own name (`repository_index_bytes`) rather than folded into a total that means entries — an index swept away with the entries it describes cannot report the eviction, which is the one moment it exists for. It lives inside the generation root, so retiring a generation retires its indexes with it: the keys are only readable by the build that wrote them. Regression `tests/test_parse_store_repository_coverage_ask17.py`, 11 assertions, including the one the row is about — entries deleted underneath a recorded repository make coverage **fall** while the population holds. ⚠ **Still owed, and unchanged:** the controlled cold measurement (5.7.2 against this tree, `ASK_PARSE_CACHE_MAX_MB` fixed, store emptied between runs, on a >20 000-file repository). Until it exists neither the +66 % nor its absence is assertable, and the row stays open on that half alone. **History:** **open** — ⚠ **not filed as a regression, deliberately**: the same cold figure was 274,8 s in 5.7.2 (+66 %), and the auditor refuses to call it one because the conditions are not comparable — in 5.7.2 the store had been retired by a version change, here it was mid-LRU-sweep. Seven false positives of exactly this class have been retracted over seven cycles; this is the eighth candidate and it is being held. What is owed is a measurement **we** control: cold `--compact` on a >20 000-file repository, 5.7.2 against 5.8.4, with `ASK_PARSE_CACHE_MAX_MB` fixed and the store emptied between runs — until that exists, neither the +66 % nor its absence is assertable. Recommended beside it, and cheap because both halves already exist: `cache status` should publish **coverage per repository** — how many entries belong to each analysed repository and what fraction of its files are covered — so *tutorials: 3 100 of 24 073 files cached (13 %)* replaces a blind decision about whether to raise the budget. Entries are content-addressed and the walk knows the repository, so this is a projection of facts we hold, not new analysis. **Refutation reproduced independently in cycle 8, under the reporter's own protocol**: both stores isolated and emptied, `ASK_PARSE_CACHE_MAX_MB=4096`, 2 passes per version — 5.7.2 at 112,0 / 112,6 s against 5.8.5 at 113,8 / 115,2 s (~3 %), and an A/B of the budget itself (512 MB sweeping by LRU against 4 096 MB that cannot sweep) at 9,1–10,1 s against 9,0–9,7 s: **the LRU sweep costs nothing measurable**. The +66 % is retired as the reporter's eighth false positive, with `E-37` named as its confounder. |
|
|
408
646
|
| C3-128 | **`timeline` is the only gate-shaped command above 120 s, and it has not come back to its record.** Steady state, 4 runs, audited corpus: `timeline --since HEAD~5` costs **198 469 ms**, +27,3 % over its 5.7.2 record of 155 950 ms. The two other commands over 120 s are there structurally — `delta` (234 s) and `contract-diff` (136 s) analyse two whole trees — while `timeline` analyses **five** and costs less than `delta` does with two, so its cost is not explained by the number of trees it walks. | 5.8.2 → 5.8.4 re-audits | Low-Medium — an investigation command rather than a gate, and the last performance figure of the round sitting outside its own best | **closed 5.8.5 by measurement — the clock is attributed, and there is nothing left unaccounted to tune against.** The row's own instruction was ASK-16's: instrument before tuning. `timeline` published a per-sample total over what is really four costs — materialising a tree, measuring each watched metric, releasing the tree, and the remainder — so *«198 469 ms»* named a command rather than a phase. `perf.PhaseTimings` (the ASK-16 authority, always on) now splits it, with the phase names taken from `--watch` so the split cannot drift from the population, and the tree materialisation kept as its own phase because git's work must not be attributed to an analysis. **Measured, BroadleafCommerce (2 985 `.java`), `--since HEAD~5 --watch posture`, 5 samples: wall 38 921 ms — `measure:posture` 32 273 (82,9 %), `materialise_tree` 5 376 (13,8 %, ~1 075 ms per tree), `release_tree` 1 272 (3,3 %), unaccounted 0,45 ms (0,0 %).** So the answer to the row's premise — *«it analyses five trees and costs less than `delta` does with two»* — is that five sixths of the cost **is** the analysis, re-run per tree by construction, and the git work is a sixth of it: `timeline` is N × one analysis and there is no timeline-specific overhead to remove. Any future gain belongs to the metric being sampled (`C3-84`, `C3-127`), which is where it would also help every other command, and the payload now says so per run instead of per audit. Regression `tests/test_timeline_timings_c3_128.py`, 9 assertions, including that the series itself is byte-identical across two runs once the clock readings are removed — instrumentation that moved an answer would be a worse defect than the row. **Original note:** **open** — ⚠ the cache-reuse half of `B7` must **not** be reopened on this evidence: the v3 assertion (`max(sample) < 8 × posture_warm`) passes at 6,8 in both 5.8.2 and 5.8.4, against 14,8 / 13,4 / 12,4 / 8,5 in the four versions that genuinely failed it. What has not returned is the absolute cost. Attribution comes before tuning and is now cheap: `metadata.timings` (`ASK-16`) exists, so the per-phase split across the five trees can be published before anything is changed. |
|
|
409
647
|
| B20 | **Thirteen declared renames with no cut-off date, and two distinct `1.0` identifiers meanwhile.** The registry publishes `pending_renames` for the 13 non-conforming `schema_version` values with their canonical name — exactly the policy `ASK-11` exists to enforce: the rename is an incompatible change, declared before it is made, never applied in silence. The consequence is published by the registry itself as `ambiguous_identifiers: 1` — `spring-audit` emits `1.0` for `core-analysis-v1` and `impact-chain` emits `1.0` for `impact-chain-v1`, so a consumer dispatching on the emitted version cannot tell them apart. | 5.8.4 re-audit (residual of `B19` / `B6`) | Low — declared debt rather than a defect, and the audit says so in as many words | **closed 5.8.5 — the window is declared where every other incompatible change is, and the reported ambiguity was understated by five shapes.** `BC-002` in `sourcecode.breaking_changes`: **announced 5.8.5, takes effect 6.0.0**, printed by `ask schema breaking-changes-v1`, carried in the CHANGELOG's `## Upgrading` section above the release history, and projected into `ask schema schemas-v1` as `pending_renames_window` — **one fact, two surfaces**, with a structural assertion that the registry source holds no second copy of the date. The registry's policy is generalised rather than loosened: a declared change is now `kind: exit_code` (must move an exit code) **or** `kind: contract` (must move a **published value** *and* name the version it takes effect in), because a rename breaks a consumer with every exit code still 0 and a registry that only knew about exit codes had nowhere to put it. A contract change publishes no `exit_code_before`/`after` at all — inviting a reader to check a field that cannot move is how a disclosure becomes noise. The 13 affected shapes are **read from `schema_registry.canonical_migrations()` at call time**, never copied: a shape that starts conforming leaves the declaration by itself (asserted by swapping the registry for a conforming one and watching the list empty). ⚠ **Correction to the reported cause, measured**: the audit named *two* shapes spelling their version `1.0`; there are **seven** — `core-analysis-v1`, `impact-chain-v1`, `pr-impact-v1`, `migration-blast-v1`, `spring-impact-v1`, `event-topology-v1`, `test-gap-ranking-v1` — so `ambiguous_identifiers: 1` was counting one *identifier* over seven shapes, not two. `ask schema 1.0` already resolved to all seven with each canonical name; the count is now asserted from the registry so it cannot be quoted from prose again. The rename itself is deliberately **not** made early: that is the incompatible change this policy exists to prevent. Regression `tests/test_schema_rename_window_b20.py`, 11 assertions, plus the ASK-11 battery generalised to both kinds. **Original note:** **open** — the missing half is a date, not a decision. Announce the cut-off window in `breaking-changes-v1` with the target version, the way every other incompatible change is announced, so a consumer can pin `core-analysis-v1` today and know when the bare `1.0` stops being emitted. Until then `ask schema 1.0` resolving to all seven shapes, each offering its canonical name, is the correct behaviour and must not be *fixed* by renaming an emitted value early — that is the incompatible change this policy exists to prevent. ⚠ **Round 10 ran on the 5.8.5 build and still reports the renames as *«deuda bien declarada, pero sin fecha»*, asking for precisely what `BC-002` already ships.** The row stays closed — the window exists, is announced 5.8.5 / effective 6.0.0, and is printed by `ask schema breaking-changes-v1`, carried in the CHANGELOG's `## Upgrading` and projected into `schemas-v1` as `pending_renames_window`. What it leaves behind is a **discoverability check, not a defect**: the reporter quoted `pending_renames` and `counts` out of the registry payload and did not see the window beside them, so verify that `pending_renames_window` travels in the same payload those two keys do — and if it does not, that is where it belongs. A declaration a ten-round auditor cannot find is not yet declared to a consumer. |
|
|
648
|
+
| E-38 | **A class-level route prefix carried by a meta-annotation is dropped, and the route is published without it.** `_build_route_surface` reads the class prefix only when the class symbol carries `@RequestMapping` or `@Path` **literally** (`repository_ir.py:5882-5887`); a repository that declares its own composed annotation gets `prefixes = [""]`, and the published `path` is the method suffix alone. **Measured on shenyu (5.8.17): 35 of 365 published routes — 9,6 %, over 21 controllers — carry `path: "/"`.** `AiProxyApiKeyController` is annotated `@RestApi("/selector/{selectorId}/ai-proxy-apikey")`, a Shenyu annotation meta-annotated with `@RestController` + `@RequestMapping`; its five handlers are published at `/`, `/`, `/batchDelete`, `/{id}`, `/page`. The served URLs do not exist as published. This is a **confident falsehood**, not a gap: no `path_resolution: "unresolved"`, no `path_expression`, no warning, and `route_census.distinct_routes` (316 of 365) reads the collisions as if two controllers genuinely shared a path. It also propagates to every consumer keyed on route text — `explain-endpoint`, `data-exposure --path-prefix`, `validation --path-prefix`, `pr-impact` route matching — where a correct query returns nothing. **The machinery already exists on the other axis**: `E-12`/`5.7.1` closed exactly this class for *beans* by resolving the spelling against the graph's `imports` edges, and `BeanGraph.build` already builds a meta-annotation map in its first pass. The route axis never asked. Vendor-agnostic by construction: the rule is "an annotation this repository declares that itself carries `@RequestMapping`/`@Path`", never a proprietary name. | found 2026-08-21 during the `AUD-590-R01` census triage, on the corrected 5.8.17 build; **not part of the seventh-pass queue and deliberately not fixed in it** — it moves route counts on any repository using composed annotations and needs its own measured pass with the golden set re-baselined | **High** — it is the rule this repository enforces most loudly, inverted on the route axis: a published route that is not served, with no affordance saying so. Blast radius is every repository with a framework-level or in-house composed controller annotation; shenyu is the witness, dubbo/sa-token/halo are unmeasured | **closed `5.8.18`** (`096c7c3`) — direction taken as written: resolve the class-level prefix through the meta-annotation map `BeanGraph.build` already computes, in the order `E-12` established (explicit single-type import decides; wildcard leaves it admissible; a repository-declared annotation in the owner's package is the owner's own; **nothing resolved stays unresolved rather than acquiring a verdict from silence**). Where it cannot be resolved, publish `path_resolution: "unresolved"` with the annotation as `path_expression` — the contract the method-level path already honours — never a bare `/`. Regression on a fixture declaring its own composed annotation, plus a shenyu assertion that no published route is `/` |
|
|
649
|
+
| B29 | **A cache-freshness test asserts a cold cache without ensuring one, so it fails on every run but the first after a purge.** `tests/test_cli.py::test_compact` asserts `_meta.timing_ms.served_from_cache is False` and does nothing to make that true: the first `--compact` run over its fixture populates the cache the second one then hits. Verified symmetrically at `d9071d4` (clean) and at every commit of the seventh-pass queue — purge, run, pass; run again, fail — so it is **state-dependent, not a regression**, and it has been reported as "1 pre-existing failure" for the whole pass. It is still a red line in a suite whose value is that red means something. | found 2026-08-21 while establishing the baseline for the seventh-pass queue | **Low** as a defect, **medium** as an eroder — a permanently-red test trains a reader to skip the summary line, which is how a real regression gets shipped | **closed `5.8.18`** (`684a9df`) — direction taken as written: the test must create the state it asserts (purge this fixture's cache dir in the fixture, or assert the *transition* cold→warm across two runs rather than the absolute). Do not "fix" it by deleting the assertion: `served_from_cache` is the field `is_stale`/cache-freshness rows are keyed to, and a cold-start claim is worth holding |
|
|
410
650
|
| E-37 | **Two environment variables move three cache stores, and no surface said which moves which.** `SOURCECODE_CACHE_DIR` relocates the per-repository snapshot store (core snapshots, the rendered L2 view, the RIS); `SOURCECODE_CONTEXT_CACHE_DIR` relocates the shared ones (the Canonical IR and the per-file parse store). Their defaults are siblings under `~/.sourcecode`, so the split is not derivable from the paths, and `cache status` published a path per store with no variable beside it while `docs/CACHE.md` named both variables in one sentence after the table. | found here 2026-08-18 while taking the controlled A/B `ASK-17` owed, not reported by the field | Medium — it does not make an answer wrong, it makes a **measurement** wrong, silently: an operator who redirects or empties *the cache* moves one store and is served answers out of the other | **closed 5.8.5** — measured first: with only `SOURCECODE_CONTEXT_CACHE_DIR` redirected, the second cold pass of each version answered from the L2 view its own first pass had written — **2,4 s against 113 s, with an empty parse store** — and the run reported `cache_source: L2_view` without anything naming the store it came out of. Now every store in `cache status` publishes `base_env`, from one authority (`cache._STORE_BASE_ENV`), the text answer prints *«<path> ($VAR moves it)»* beside each one, `stores.base_env_note` states the consequence rather than the layout, and `docs/CACHE.md` carries a **column** instead of a sentence covering both. Regression `tests/test_cache_store_base_env_e37.py`, 8 assertions, and the ones that matter are **behavioural**: setting the variable a store names moves *that* store and leaves the others where they were — a published mapping nobody exercises is the class of claim this repository refuses everywhere else. The CLI is forbidden a second copy of the mapping (structural assertion). **Confirmed from the outside in cycle 8, by the reporter falling into it first**: isolating only `SOURCECODE_CONTEXT_CACHE_DIR`, their second *cold* pass answered from the L2 view their first pass had written — **2,4 s against 113 s with an empty parse store** — and their write-up now quotes the shipped `base_env` mapping back as the fix (*"dos variables mueven tres almacenes y ninguna superficie decía cuál"*). A defect found while measuring, that was itself corrupting the measurement. |
|
|
411
651
|
| B21 | **Six of the thirteen shapes being renamed name no command that emits them.** `schemas-v1` publishes `emitted_by` per shape; for `migration-blast-v1`, `spring-impact-v1`, `event-topology-v1`, `test-gap-ranking-v1`, `canonical-ir-v1` and `hibernate-strategy-v2` it is `[]`, beside seven shapes where it is populated. Read the way a consumer reads a list, an empty `emitted_by` says *nothing emits this shape* — which tells the one consumer that **is** affected by `BC-002` that it is not. | found here while declaring `BC-002` (5.8.5), not reported by the field | Low — a disclosure gap in a declaration whose whole purpose is to let a consumer decide whether it is affected | **closed 5.8.6 — derived, not typed out.** `schema_producers` reads the two authorities the CLI already holds — the registered command tree and the import graph of the module that defines each callback — and answers with the evidence it derived on: `command_import` (the command's own body imports the producing module), `module_import` (one hop, through an intermediate narrow enough to attribute), `substrate` (reached from more commands than a producer would be) or `unresolved` (nothing found, said in a sentence). Five of the six now name a command — `migration-blast-v1` → `migrate-check`, `spring-impact-v1` and `event-topology-v1` → `impact-chain`, `test-gap-ranking-v1` → `impact`, `hibernate-strategy-v2` → `migrate-check`/`migrate-recipe` — and the sixth, `canonical-ir-v1`, is reached from **28 modules**: naming every command would be less true than naming none, so it publishes `substrate` with the count and the instruction to match on `subject`. A declared producer still wins, because some producers are not one registered command (`ask (root)`, `baseline capture`). Regression `tests/test_schema_producers_derived_b21.py`, 33 assertions, parametrised by the registry so a fourteenth unnamed shape fails the battery. **History:** open, and named rather than shipped as a fact: every `BC-002` row with an empty producer list carries `emitted_by_basis` — *«not declared in the registry: this shape's producing command is not named yet, so match on `subject`»* — so the gap is visible instead of being read as a negative claim. What remains is to fill the six, and the honest way is derivation rather than six more hand-written strings: the producer is discoverable from the command that constructs each payload, which is the same authority every other population in this repository comes from. Until then the assertion in `tests/test_schema_rename_window_b20.py` holds the weaker invariant a reader can rely on: a row either names its emitters or says why it cannot. |
|
|
412
652
|
| C3-129 | **`timeline`'s absolute cost, now fully attributed and still 27 % above its own record.** `C3-128` closed by instrumenting rather than tuning, which was the right order and left this behind: the clock is split, nothing is unaccounted, and the number has not moved. Round 10 on 5.8.5 measures `timeline --since HEAD~5` at **~198 000 ms against the 5.7.2 record of 155 950 ms (+27 %)** — the only gate-shaped command over 120 s that is not there structurally (`delta` 191 s and `contract-diff` 115 s walk two whole trees, `compare` 46 s walks N candidates; `timeline` walks **five** trees and costs more than `delta` does with two). | 5.8.2 → 5.8.5 re-audits (residual of `C3-128`, which closed the attribution half in 5.8.5) | Low-Medium — an investigation command rather than a CI gate, and the last performance figure of the round sitting outside its own best | **closed 5.8.6 — by not paying twice for what does not change.** A metric at a commit is a pure function of that tree, and the tree at a sha never changes: the only cache in this product whose key can be exact rather than heuristic (analyzer fingerprint + metric + commit sha). `timeline_cache.SampleCache` stores measured values under the per-repository core store, and when every watched metric for a sample is a hit the tree is **never materialised**, which is the 13,8 % the split attributes to git. Measured on BroadleafCommerce over the same range: **26 845 ms → 12 ms**, samples and transitions byte-identical. What makes it admissible rather than merely fast, all asserted: a failed sample is never stored (one bad run cannot become permanent), the fingerprint is in the key (corrected analysis is never served from a previous release), the hit is published (`sample_cache`, `from_cache` per sample, and `cost` separates the cached population from the measured one instead of reporting a median over both), and `--no-cache` / `ASK_TIMELINE_NO_SAMPLE_CACHE=1` measure everything again so the command stays measurable without emptying a store (E-37). `cache_model` carries the layer, so `cache model` and `docs/CACHE.md` state it from one authority. Regression `tests/test_timeline_sample_cache_c3_129.py`, 11 assertions. **History:** open — and it opened with the answer already published, which is why it is cheap to attack and expensive to leave: `timeline`'s own `timings` say **82,9 % `measure:posture`, 13,8 % `materialise_tree`, 3,3 % `release_tree`, 0,0 % unattributed**. So this is not a `timeline` defect at all: it is `posture` paid five times over five materialised trees, and it makes the same substrate question `C3-127` asks — with the difference that here the five payers are one command, so the cache is intra-invocation and needs no cross-command contract. ⚠ **Do not reopen `B7`**: cross-tree cache reuse is healthy, the v3 assertion passes at 6,8×. Acceptance: `timeline --since HEAD~5` back inside +10 % of 155 950 ms on the audited corpus with `timings` still summing to 100 %, and `measure:posture` falling as a share rather than the total falling for an unnamed reason. |
|
|
@@ -457,7 +697,7 @@ retained only as evidence of what the auditor observed before the fix.
|
|
|
457
697
|
| C1-27 | **Two definitions of "the working tree changed" inside one command.** `cache freshness` reports STALE with `RIS HEAD == current HEAD` and `Delta: 0` because **one untracked file no analyser reads** (an editor settings file) is a porcelain line. C3-40 (4.5.3) replaced a false *fresh* with a false *stale*: its predicate is *any* `git status` entry, while the fact the snapshot describes is *the files the analysis admits*. Reproduced on a battery repository with a clean tree: `FRESH` → write `.claude/settings.json` → `STALE` | 4.5.3 (eval #11) | Medium — a freshness signal that fires on a file the analysis cannot see teaches a reader to ignore the one that matters, and it invalidates a warm cache for nothing on a repository where a cold run costs minutes | **closed 4.6.0 — verified in the field (eval #12).** One authority decides which paths qualify (`path_filters.analysis_reads`) and both consumers ask it: the freshness boolean and the **tree signature**, which is also the cache key — so an agent writing its own settings file was discarding every warm answer. `.github` stays a change (tooling evidence); the default is inclusive, because a missed change is a false *fresh*. Field: `FRESH` / `Uncommitted: False` with the same untracked settings file still present |
|
|
458
698
|
| C1-28 | **The test/production split still has two authorities.** C1-2 made `test_sources` the authority in 3.2.1 and rebound five consumers; `path_filters.is_test_path` was never retired and still answers for its own callers. They disagree on five path shapes, in both directions — `src/main/java/**/Test*.java` (`False` / **`True`**), `src/main/java/**/*IT.java` (**`True`** / `False`), `src/main/java/**/test/Helper.java` (`False` / **`True`**). Field symptom: `--compact` counts **1 test file that lives in `src/main`**, with the verdict (0.0 %, critical) correct and the numerator wrong | 4.5.3 (eval #11) | Medium — it is the denominator of every coverage statement, and the direction that inflates is the one a reader must never be misled about (C1-2's own rule) | **closed 4.6.0 — verified in the field (eval #12).** `path_filters.is_test_path` is now a passthrough to the authority, after absorbing everything it knew that the authority did not, and a declared **main** source root outranks a naming convention. Fleet effect, cold: keycloak 5 525 → 5 486 Java files with endpoints **676 → 678**, openmrs-core 866 → 861, Broadleaf / alfresco / petclinic unchanged. Field: `0 test files for 3337` **with the reason stated** (no declared test source root, no file matching a test naming convention) — the numerator and its basis, which is what made it verifiable from outside |
|
|
459
699
|
| C1-29 | **Candidate, not reproduced: three commands publish three severities for one `defect_id`.** Eval #12 states it and could not measure it — verifying it costs ~45 min of `risk` (C3-42). Recorded so the complaint is not lost, and recorded as a **candidate** because the standing rule since audit #8 is that no row is opened on an unreproduced complaint. What must be measured, in this order: (1) do `spring-audit`'s rule severity, `risk`'s `severity_declared` / `severity_effective` band and `impact-chain`'s per-finding severity (beside its own `risk_level`) ever publish **the same field name with different values** for one `defect_id` — that is the defect; (2) or do they publish **different names for different units** — that is the design, and `risk` already does it correctly today (`severity_declared` carries `severity_authority: spring-audit`, `severity_effective` is the composed product). The likely real row, if one survives, is the **third** publication: `impact-chain` emits findings with a severity and a `risk_level` of its own, computed by a scale that is not `RiskComposer`'s | 4.6.0 (eval #12) | Unknown until reproduced — potentially **High** (it is the composition axis, scored 4/10), potentially not a defect at all | **does not reproduce (measured 4.10.0)** — on the battery fixture, the same `defect_id` carries `severity: medium` in `spring-audit`, `impact-chain.impact_findings[].severity: medium`, and `risk.severity_declared: medium`; the composed result is deliberately a different field (`severity_effective` plus `band`) with factor authority `spring-audit`. That is the intended naming case, not a contradiction. Locked by `tests/test_risk.py::test_c129_declared_and_effective_severity_are_different_units` |
|
|
460
|
-
| C1-36 | **A container-wired component has a call-graph fan-in of ~0, so every risk surface publishes its blast radius as zero.** Reproduced twice, on two builds, by two evaluations. Eval #14 (4.10.4): `pr-impact . --files -` over `SecurityConfig.java` + `M3FiltroSeguridadAspect.java` + `LegacySha1PasswordEncoder.java` — the application's whole filter chain, the aspect that gates 935 routes, and the production password encoder — returns `risk_level: LOW`, `risk_reason: "No high-risk signals detected"`, `affected_endpoints: []`, `direct_callers: []`. Positive control, same command, same cache: a single POJO (`Persona.java`) returns `CRITICAL`, *"Public API + High Call Fan-in"*, 3 382 endpoints, 48 callers. Eval #15 (4.10.5) reproduced it on a two-file diff (`SecurityConfig` + `LegacySha1PasswordEncoder`): `LOW`, 2 classes, **0 endpoints** — *"un PR que borrase `.antMatchers("/privado/**","/api/**").authenticated()` pasa el gate en verde"*. Same root in three more commands: `impact-chain M3FiltroSeguridadAspect` → `risk: low`, 0 callers, 0 endpoints; `impact-chain LegacySha1PasswordEncoder` → `low`, 1 caller, 0 endpoints; `plan M3FiltroSeguridadAspect` → `affected_endpoints.count: 0`, `rollback_surface: 1 file`; `compare SecurityConfig M3FiltroSeguridadAspect` → both `blast_radius: 0`, `secured_endpoints: 0`, ranked 2nd and 3rd behind a directory (C3-67). **The knowledge is already in the binary and no surface consults it**: `endpoints` resolves `@M3FiltroSeguridad` → 935 routes (`custom_gate_inferred`); `explain SecurityConfig` emits `request chain: /privado/**, /api/** → authenticated (SecurityConfig.java:47)`; `posture` computed `open_by_rule: 2635` for `default` against `1` for `m3` **from that same class**. Two commands of one binary disagree about whether the aspect matters | 4.10.4 (eval #14), 4.10.5 (eval #15) | **Critical** — it is a security false negative on a command the `--help` sells as *"contract stable within a major — safe to gate CI on"*, and it is the one important blind spot the product does **not** declare: every other gap has an `NC-0xx`, this one has none. Eval #14 scores risk scoring **3/10** for it and states 7 → 8,5-9 on this row plus C3-53 alone | **remedies (1), (2) and (3)
|
|
700
|
+
| C1-36 | **A container-wired component has a call-graph fan-in of ~0, so every risk surface publishes its blast radius as zero.** Reproduced twice, on two builds, by two evaluations. Eval #14 (4.10.4): `pr-impact . --files -` over `SecurityConfig.java` + `M3FiltroSeguridadAspect.java` + `LegacySha1PasswordEncoder.java` — the application's whole filter chain, the aspect that gates 935 routes, and the production password encoder — returns `risk_level: LOW`, `risk_reason: "No high-risk signals detected"`, `affected_endpoints: []`, `direct_callers: []`. Positive control, same command, same cache: a single POJO (`Persona.java`) returns `CRITICAL`, *"Public API + High Call Fan-in"*, 3 382 endpoints, 48 callers. Eval #15 (4.10.5) reproduced it on a two-file diff (`SecurityConfig` + `LegacySha1PasswordEncoder`): `LOW`, 2 classes, **0 endpoints** — *"un PR que borrase `.antMatchers("/privado/**","/api/**").authenticated()` pasa el gate en verde"*. Same root in three more commands: `impact-chain M3FiltroSeguridadAspect` → `risk: low`, 0 callers, 0 endpoints; `impact-chain LegacySha1PasswordEncoder` → `low`, 1 caller, 0 endpoints; `plan M3FiltroSeguridadAspect` → `affected_endpoints.count: 0`, `rollback_surface: 1 file`; `compare SecurityConfig M3FiltroSeguridadAspect` → both `blast_radius: 0`, `secured_endpoints: 0`, ranked 2nd and 3rd behind a directory (C3-67). **The knowledge is already in the binary and no surface consults it**: `endpoints` resolves `@M3FiltroSeguridad` → 935 routes (`custom_gate_inferred`); `explain SecurityConfig` emits `request chain: /privado/**, /api/** → authenticated (SecurityConfig.java:47)`; `posture` computed `open_by_rule: 2635` for `default` against `1` for `m3` **from that same class**. Two commands of one binary disagree about whether the aspect matters | 4.10.4 (eval #14), 4.10.5 (eval #15) | **Critical** — it is a security false negative on a command the `--help` sells as *"contract stable within a major — safe to gate CI on"*, and it is the one important blind spot the product does **not** declare: every other gap has an `NC-0xx`, this one has none. Eval #14 scores risk scoring **3/10** for it and states 7 → 8,5-9 on this row plus C3-53 alone | **closed 5.8.20 — all four remedies. (1), (2) and (3) landed together in 4.10.6; (4), F-Y, landed here.** The floor said the fan-in cannot describe this component and nothing said what can, so the honest `HIGH` shipped beside an empty endpoint list for nine releases. F-Y measures the reach from what the component itself declares, in one authority (`wired_surface`) that all three derivations consume — `compute_blast_radius` (impact / plan / compare), `spring_impact` (impact-chain) and `pr_impact`. **pointcut → surface**: an `@Aspect` advises the members its pointcut selects, so `@Around("@annotation(X)")` reaches every handler carrying `X`. The expression is the fact and the IR did not carry it — `_ADVICE_ANN_ARGS` joins the annotations whose argument already is one, so nothing re-parses Java to recover it. **matcher → routes**: a chain covers the routes its matchers select, read through `chain_rules`, the authority `posture` already uses, rather than a second extractor. Measured on a fixture reproducing the field's shape: `impact` and `impact-chain` both publish 2 of 3 routes for the aspect (the handler without the annotation is excluded) and `pr-impact` publishes 3 via the chain's `/orders/**` matcher, each figure carrying its `_unit`, `_population`, `_basis` and `direction: under_reports`. **The floor is deliberately untouched**: `risk_level` stays `high` and `risk_score` stays `null` (C1-41) — the band was established *without* the fan-in and a measurement of reach is not fan-in arithmetic. A component declaring neither mechanism gets `null`, never `0`; a matcher rule `chain_rules` records as `not_evaluated` is counted `undecided`, never covered. `NC-010` now states what is measured and what is still not — a pointcut this analysis cannot resolve, a matcher built by a helper method, a filter registered programmatically — because *"nothing here measures this"* printed beside a measurement is the contradiction this product is audited for. Locked by `tests/test_wired_surface_fy.py` (15), which asserts through the commands on real sources, including that **`impact` and `impact-chain` publish the same count for one symbol** — this row's own sentence was *two commands of one binary disagree about whether the aspect matters*. Registered in `defect-classes-v1`: `container_reachers` flips to `swept: true` on a sweep that iterates `KINDS`, and `wired_component_reach` joins over `REACHES`. **What this closes for `CL-16`**: `pr-impact` keeps its `core` tier on the evidence the tier promises, not on a floor. Previously — The three cheap ones landed together because separating them would have shipped a warning nobody could read next to a verdict that still said `LOW`. One authority, `container_wiring.py`, answers *is this class reached by the container rather than by the code?* from facts the CIR already carries — the role it already assigns (`config`), the interception annotations and supertypes fixed by published specifications, and one wholly structural route that needs no name at all: **the class implements an interface something injects and has zero incoming call edges**, which is how the password encoder is caught without anyone naming a password encoder. It is wired at the *one* blast-radius authority (`_compute_blast_radius`, which `impact`, `plan` and `compare` all project from) plus `pr_impact` and `spring_impact`, so the four commands that reproduced the defect answer from one place rather than four. The verdict is **floored, never lowered** — a `CRITICAL` stays `CRITICAL` — the payload publishes the population (`container_wired`) and the reason, and `NC-010` states the limit in the payload, the README and the help. Locked by `tests/test_container_wiring.py` (16), which asserts the observable verdict on real sources — an `@Aspect` with fan-in 0 rates `high` and says why, an ordinary service with callers is untouched — and by the negative controls that keep the floor from becoming a blanket (`pom.xml` and a CI workflow are not wiring descriptors). Field-shaped remedies, in ascending cost. (1) **Half an hour, today: declare the non-coverage.** An `NC-0xx` stating that blast radius is projected over the call graph, that container-wired components (`@Configuration`, `@Aspect`, `Filter`, `web.xml`) have fan-in ≈ 0 and therefore appear with null impact, and that their real risk is not measured here. This converts a silent failure into an honest limit, which is the whole promise of the product. (2) **`analysis_warnings` when ≥1 analysed class is container-wired and its fan-in is 0** — today the only warning on that payload is the self-referential exclusion, cosmetic beside this. (3) **A floor by file class**: a diff touching `@Configuration`, `@Aspect`, `Filter`, `@ControllerAdvice`, `PasswordEncoder`, `web.xml` or `application.y*ml` cannot publish below `HIGH`, with `risk_reason: "container-wired component — call-graph fan-in is not a risk proxy here"`. (4) **The real fix, F-Y**: pointcut → surface (an `@Aspect` whose pointcut is `@annotation(X)` inherits the routes carrying `X` — 935, not 0) and matcher → routes (a modified `SecurityFilterChain` inherits the population its matchers cover). |
|
|
461
701
|
| C1-37 | **Four incompatible figures for "endpoints" at one commit, and a baseline that disagrees with itself.** `impact Persona . → stats.endpoints_affected_count` = **3 364**; `pr-impact . --files -` over the same class → `affected_endpoints` = **3 382**; `data-exposure . → summary.endpoints_exposed` = **3 382**; `endpoints .` total = **3 574**. Inside one artefact, `.ask/baselines/3dde0376.json`: `totals.endpoints` = **3 574** against `len(endpoint_surface)` = **3 538** — a file whose declared purpose is to be a replayable fingerprint differs from itself by 36 endpoints | 4.10.4 (eval #14) | **High** — it breaks comparability between commands and the credibility of every ratio keyed on an endpoint count. Eval #14 scores numeric consistency **5/10** and names this first | **closed (4.10.6).** The baseline half had a cause, not just a discrepancy: `totals.endpoints` counts **handler mappings** and `endpoint_surface` is the set of distinct **(METHOD, path) contract pairs**, so the 36 are routes declared by more than one handler. Both facts are now named — `totals.endpoint_contract_pairs` beside `totals.endpoints`, with `endpoints_unit` saying which is which — and the battery holds `len(endpoint_surface) == totals.endpoint_contract_pairs`, an invariant that is checkable because the two figures finally describe two things. The cross-command half applies the `direct_callers_note` pattern to all four: `endpoints.total_unit` (the declared population), `impact.stats.endpoints_affected_count_unit` (one blast cone), `pr-impact.metadata.affected_endpoints_unit` (a UNION over the changed classes, keyed on endpoint id — which is exactly why it can exceed a single cone) and `data-exposure.summary.endpoints_exposed_unit` (reach-selected subset). Each names the others, so a reader who meets one figure learns the rest exist. Original remedy note: the pattern is already solved in this product: `direct_callers_note` declares the unit of its figure and what references it admits, and says so after measurement falsified an earlier promise. Apply it to the endpoint counts: every published count carries its `unit`; `impact` and `pr-impact` either declare why they differ by 18 or converge; and the battery gets the invariant `len(endpoint_surface) == totals.endpoints` unless a declared cap says otherwise — `onboard` already has the shape for that (`_truncation_summary`) |
|
|
462
702
|
| C1-38 | **`secured_endpoints` names four different populations, and one of them is reconcilable with nothing.** At one commit: `endpoints.exposure.by_security_policy.custom_gate_inferred` = **935**; `data-exposure.by_security_verdict.protected_custom` = **939**; `impact Persona.stats.security_surface_count` = **15**; `compare` (whole-repo candidate) `secured_endpoints` = **15**. 935 vs 939 is an axis difference and defensible — but it is not declared. **15 matches nothing**: not 4 (`programmatic`), not 935, not 939 | 4.10.4 (eval #14) | **Medium-High** — it is the field an auditor reads first, and *"un campo llamado `secured_endpoints` que devuelve 15 sobre un repo con 935 rutas anotadas es una trampa para el lector"* | **closed (4.10.6), and the cause was not naming.** The `15` reconciled with nothing because it was **`min(real, 15)`**: `security_surface_affected` was truncated to a 15-row display sample and `stats.security_surface_count` was `len()` of the truncated list — a rendering choice that had become the measurement, the exact R5 defect the rest of that file already fixes. The population is now counted before the cut, the list is still sampled, and the omission is published with its `cap_effect`. The naming half shipped with it: all four figures now carry their unit beside them in the `direct_callers_note` style — `by_security_policy_unit` (annotated routes), `by_security_verdict_unit` (posture verdicts over the same routes, a different axis), `security_surface_count_unit` (declarations inside one blast cone) and `compare.cost_units.secured_endpoints` — each naming the other two so a reader who meets one figure learns the other two exist. Original remedy note: rename to what each one measures (`security_annotated_routes` / `programmatic_gate_routes` / `security_gated_in_blast_cone`) and attach the unit note in the `direct_callers_note` style. Same remedy family as C1-37 and C1-32 |
|
|
463
703
|
| C1-39 | **A `not_found` resolution publishes zeros where the product's own best pattern publishes null.** `ask plan "añadir autorización por endpoint a los controladores sin @M3FiltroSeguridad" .` returns a well-formed `change-plan-v1`: `resolution: "not_found"` **and** `affected_components.count: 0`, `affected_endpoints.count: 0`, `rollback_surface.file_count: 0`, `review_checklist: []`. It is not an error envelope — an agent reading `affected_endpoints.count` concludes *"no impact"* where the truth is *"not measured"*. `--agent` is a declared use case. Same shape elsewhere: on a repository with **0 test files for 3 337 non-test Java files**, every `tests_at_risk: 0` and `covering_tests: 0` reads as *"the change is safe"* when it means *"there are no tests"* | 4.10.4 (eval #14) | **Medium** — dangerous specifically for agent consumption, and it is the one place the product breaks its own best habit: `data-exposure` handles exactly this with `answered: false` and measures nothing rather than publishing a confident zero | **closed (4.10.6).** With an unresolved target every cost field is `null` and carries `answered: false` with the reason, plus a `how_to_read` that says why in one line — *a zero would say the change is free*. Same rule for the second half: on a repository with **no test source root**, `covering_tests` answers `null` / `answered: false` / *no test source root in this repository*, and the review checklist says the axis was not measured instead of reporting no covering tests. `compare` reads a 0 for the ranking (a total order needs a number) and its `why` trail carries the plan's own reason, so the placeholder is never narrated as a measurement. This is the shape `data-exposure` already had; it is now the same shape one level up. Original remedy note: with `resolution: "not_found"`, cost fields become `null`, never `0`. With no test source root, `tests_at_risk` and `covering_tests` become `null` with `reason: "no test source root"`. This is the C5 silent-failure shape, one level up: the envelope is valid, so nothing looks wrong |
|