sourcecode 5.8.6__py3-none-any.whl → 5.8.7__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Potentially problematic release.


This version of sourcecode might be problematic. Click here for more details.

sourcecode/__init__.py CHANGED
@@ -4,4 +4,4 @@ ASK Engine is the product. ``ask`` is the canonical CLI command; ``sourcecode``
4
4
  the legacy compatibility alias and the Python/PyPI package name. See
5
5
  docs/PRODUCT_IDENTITY.md (normative)."""
6
6
 
7
- __version__ = "5.8.6"
7
+ __version__ = "5.8.7"
@@ -89,7 +89,7 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
89
89
  | C1-13 | **Hub-guard publishes `endpoints_total: 0` instead of `unknown`.** A service with 780 callers is capped at `depth=1` for cost — defensible — but reports 0 endpoints reached, contradicting the epistemic standard the tool imposes on itself everywhere else. Hubs are precisely where blast radius is asked for | 3.2.1 (eval #4) | Medium | **closed 3.6.0** — no battery repository crosses the guard (openmrs-core's largest hub, `Context`, has 215 caller classes against a threshold of 500), so it was reproduced on a generated fixture: `bfs_truncated: true`, `depth_reached: 1`, `endpoints_affected_count: 0` beside four more zeros nothing looked for. A reach-derived count of 0 under the cap is now `null`, named in `stats.unknown_under_cap`. The distinction is the point and is asserted both ways: a **non-zero** count under the cap stays a number (it is a floor the effect block already describes), and an **uncapped** run keeps its measured zero, because converting that would be the mirror-image defect. The hub threshold moved to module level so the fixture crosses the real guard instead of a copy of its value |
90
90
  | C1-14 | **File-count drift across commands**: 3 334 / 3 337 / 3 342 / 3 374 files for one repository in one session. No command is wrong in isolation; together they cost the one thing a deterministic tool cannot lose — that its numbers agree with themselves | 3.2.1 (eval #4) | Medium | **within-document half closed 3.6.0 · cross-command half does not reproduce.** Of six commands measured on openmrs-core (`--compact`, `endpoints`, `spring-audit`, `migrate-check`, `modernize`, `validation`) only `migrate-check` publishes a scanned-file count at all, so there are no two figures left to disagree. What does reproduce is the same defect *inside one document*: `affected_files` named two populations — 265 under `summary` (every finding, any scope) against 264 under `effort_breakdown` (blocking findings only). Both correct; the shared name was the contradiction. Same remedy as C1-5: the narrower becomes `blocking_affected_files`, both carry a `_basis`, the old key stays an alias, and the battery asserts the two diverge exactly when a finding is non-blocking |
91
91
  | C1-15 | **Analysis assumes a single destination.** Asked about a JDK move, `migrate-check` answers `headline_blocker: boot3_migration` and emits fix hints that send a developer to migrate namespaces they never planned to touch (481 `javax.persistence` under MIG-001). There is no way to say "JDK 11, keep Boot 2.7" | 3.2.1 (eval #4) | **High (product)** | **closed 3.6.0** — `--target-jdk N` and `--keep-boot` say which migration is being planned, filtered over the `migration_target` vocabulary every rule already carried rather than a parallel one. Measured on openmrs-core with `--target-jdk 11 --keep-boot`, the row's own acceptance: **MIG-001 85 → 0**, blocking 114 → 0, `headline_blocker` `hibernate_rewrite` → `none_applicable`, and the Hibernate 5→6 rewrite drops with it because Hibernate 6 arrives with Boot 3. Narrowing is not silencing: `metadata.route` publishes the target, the flag and every excluded axis, a `limitations` entry states how many findings were left out and that they are unresolved rather than fixed, and the route's own blockers stay (a JDK 11 route still reports MIG-021, asserted). The default is untouched, verified identical on three repositories |
92
- | C1-16 | **Four fan-in figures for one class, and the note that promises the relationship is itself false.** Field: `modernize.in_degree` 1 026 / `impact-chain.direct_callers` 611 / `explain.incoming_callers` 519 / `impact.stats.direct_caller_count` 484 for one service, plus 3 222 vs 3 458 endpoints affected. **Re-measured on the battery, 3.4.0, `Money` @ BroadleafCommerce: 673 / 1 868 / 151 / 418** — `impact-chain` is absent from the published relationship and is **2,8× larger than the figure the note declares "always the largest of the three"**, while `explain` vs `impact` differ by **177 %**, not the *"small margin"* the same note promises. In the same runs `impact.stats.endpoints_affected_count: 0` against a populated `endpoints_affected` from `impact-chain` — a **false zero** on the blast-radius axis | 3.2.1 (eval #5), reopened 5.5.4 | **High** — this is the number a risk decision is taken with: *"¿el blast radius son 484 o 611 clases?"* | **reopened 5.5.4.** The DI and subtype fixtures remain closed, but the claimed one-authority implementation was false for service locators: `impact` kept a local BFS that elevated `Locator#accessor` to every caller of `Locator`, while `impact-chain` used `caller_reach` and did not. 5.5.4 recovery removes that elevation and adds a two-accessor regression fixture; authority unification and cross-command parity remain open as Q-03. **Three further field witnesses, 2026-08-16, all on Q-03 and none needing a row of their own.** (a) `keycloak-config-cli` F-5: `KeycloakProvider` fan-in published as **10** (`impact.stats.direct_caller_count`), **12** (`explain.incoming_callers_count`) and **22** (`impact-chain.metadata.direct_caller_symbol_count`) in one session — every figure carries its declared unit and `caller_metric_note` reconciles them, but `impact-chain`'s *prose* prints the symbol count under the words "22 direct callers", contradicting its own metadata two lines below. The unit belongs in the sentence, not only in the field. (b) `keycloak-config-cli` F-4 → **closed `ea4d791`**: `impact` and `impact-chain` published 77.26 vs 74.19 and 79.81 vs 76.79 for the same symbols while `impact` shipped no `risk_model`, so the divergence could not be read off the payload; `impact` now publishes the basis it already shares. (c) BroadleafCommerce D7: `impact-chain` 24 direct callers / 227 indirect / 37 endpoints against `impact` 12 classes / 34 symbols / 249 / 38 for `ExploitProtectionService`, plus `project_summary: "41 transactional boundaries"` beside `transactional_boundaries.class_count: 29`. The evaluator's verdict is the one that matters commercially: *"un cliente que compare dos comandos verá una contradicción antes de leer la nota."* |
92
+ | C1-16 | **Four fan-in figures for one class, and the note that promises the relationship is itself false.** Field: `modernize.in_degree` 1 026 / `impact-chain.direct_callers` 611 / `explain.incoming_callers` 519 / `impact.stats.direct_caller_count` 484 for one service, plus 3 222 vs 3 458 endpoints affected. **Re-measured on the battery, 3.4.0, `Money` @ BroadleafCommerce: 673 / 1 868 / 151 / 418** — `impact-chain` is absent from the published relationship and is **2,8× larger than the figure the note declares "always the largest of the three"**, while `explain` vs `impact` differ by **177 %**, not the *"small margin"* the same note promises. In the same runs `impact.stats.endpoints_affected_count: 0` against a populated `endpoints_affected` from `impact-chain` — a **false zero** on the blast-radius axis | 3.2.1 (eval #5), reopened 5.5.4 | **High** — this is the number a risk decision is taken with: *"¿el blast radius son 484 o 611 clases?"* | **reopened 5.5.4.** The DI and subtype fixtures remain closed, but the claimed one-authority implementation was false for service locators: `impact` kept a local BFS that elevated `Locator#accessor` to every caller of `Locator`, while `impact-chain` used `caller_reach` and did not. 5.5.4 recovery removes that elevation and adds a two-accessor regression fixture; authority unification and cross-command parity remain open as Q-03. **Three further field witnesses, 2026-08-16, all on Q-03 and none needing a row of their own.** (a) `keycloak-config-cli` F-5: `KeycloakProvider` fan-in published as **10** (`impact.stats.direct_caller_count`), **12** (`explain.incoming_callers_count`) and **22** (`impact-chain.metadata.direct_caller_symbol_count`) in one session — every figure carries its declared unit and `caller_metric_note` reconciles them, but `impact-chain`'s *prose* prints the symbol count under the words "22 direct callers", contradicting its own metadata two lines below. The unit belongs in the sentence, not only in the field. **Witness (a) closed 5.8.7**: `impact` had already been corrected this way (*"433 direct callers" beside `direct_caller_count: 134`*) and `impact-chain` was left phrasing the entry count as *"22 direct callers"*. Its explanation now reads `N direct caller classes (M reference sites)`, with N taken from the same `caller_classes` call `metadata.direct_caller_count` publishes — one authority, so the sentence and the field cannot drift apart. The parity is asserted against the payload rather than against a literal string. Witnesses (b) and (c) are unchanged: (b) is closed, and (c) is the cross-command reconciliation that remains open as Q-03. (b) `keycloak-config-cli` F-4 → **closed `ea4d791`**: `impact` and `impact-chain` published 77.26 vs 74.19 and 79.81 vs 76.79 for the same symbols while `impact` shipped no `risk_model`, so the divergence could not be read off the payload; `impact` now publishes the basis it already shares. (c) BroadleafCommerce D7: `impact-chain` 24 direct callers / 227 indirect / 37 endpoints against `impact` 12 classes / 34 symbols / 249 / 38 for `ExploitProtectionService`, plus `project_summary: "41 transactional boundaries"` beside `transactional_boundaries.class_count: 29`. The evaluator's verdict is the one that matters commercially: *"un cliente que compare dos comandos verá una contradicción antes de leer la nota."* |
93
93
  | C1-17 | **A fact with an authority still contradicts itself across two commands**: `--compact.analysis_gaps` says *"1 test file for 3 336 Java files"* while `review-pr.test_coverage_risk.repository_has_test_sources` says `false`, in the same repository. C1-2 was closed in 3.2.1 with five consumers rebound, and this evaluation ran **on 3.2.1** | 3.2.1 (eval #5) | Medium | **does not reproduce (measured 3.6.0)** — re-measured as the row instructed. On openmrs-core, keycloak and spring-petclinic neither statement is emitted, and both consumers reach one authority: `--compact` through `analyze_test_sources`, `review-pr` through `is_test_path`. What was missing is now asserted: the authority has **two entry points** and nothing held them to each other, which is exactly the shape a sixth consumer would take. Zero mismatches across a case table including the C1-2 discriminator (a `test` package inside a *main* source root is production code), plus the literal contradiction — "N test files" and "no test sources" — asserted unreachable in both directions |
94
94
  | C1-18 | **The profile population differs per command, and the narrower one is what an agent reads.** Field: `--agent` reports a two-name profile set while `prepare-context fix-bug` reports four for the same repository in the same session — and the two names missing from the agent payload are the ones the *build* names for its deployable artefacts. Every finding about conditional security is keyed on that set, so the default agent surface answers the profile question with the population that lost the environments | 3.5.0 (eval #6) | **High** — it is the input to the one capability nothing else in the market has | **closed (unreleased)** — **fixture first, as C1-10 required**, and it reproduced worse than reported: on `tests/fixtures/spring_profile_population` the same repository in one session answered **`["qa","dev"]`** (default view — filenames + overlay directories), **`["dev","prod","staging"]`** (`posture` — config documents + `@Profile`) and **`["dev","prod","qa","staging"]`** (`--agent` / `fix-bug`, the union). Three populations under one name; every conditional-security finding is keyed on it. Authority `spring_profiles.profile_population`: four named signals (`file_naming`, `overlay_directory`, `config_documents`, `annotations`, plus `caller_supplied`), the union, and `by_signal` published so a name's provenance is auditable. The stack detector, `posture` and the serializer section all read it — asserted per surface, not claimed. Closing it exposed two false positives in the signals it inherited, both fixed structurally (never by name — VAI): an `application-{x}` file counts only **beside its base config** (alfresco's `alfresco/messages/application-model_de.properties` + 19 siblings invented twenty environments; now 0) and an overlay directory counts only when it **holds a config file** (spring-petclinic's `db/ messages/ static/ templates/` reported four; now `mysql, postgres`, which are its real profiles) |
95
95
  | C1-19 | **File-count drift across commands, with witnesses this time.** 3 337 (`compact.project_summary`) / 3 336 (`analysis_gaps.testing`) / 3 342 (`migrate-check.metadata.java_files_scanned`) for one repository in one session | 3.8.0 (eval #7) | Medium | **closed 3.9.0** — measured, not equalised: the three figures count code-extension files in any language, the non-test denominator of a test ratio, and the .java files handed to the migration rules. Each names its population now, and `migrate-check.files_scanned_basis` states the two it is not. Original note: open — **C1-14's cross-command half, reopened.** 3.6.0 recorded that half as *not reproducing* because on openmrs-core only `migrate-check` published a scanned-file count at all; eval #7 shows three commands publishing one on a repository we do not have. Three populations are plausible and each may be right in isolation (all Java files / files admitted to the graph / files the migration rules scanned), which is exactly the C1-5 remedy: name the unit per figure, or bind them to one authority. Do not fix by making the numbers equal — measure which population each counts first |
@@ -111,7 +111,7 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
111
111
  | C1-41 | **`risk_level: high` arrives with `risk_score: 0.0` — the band and the number are two authorities for one verdict.** Third round on saint-server, verbatim from 4.11.0: `risk_level=high risk_score=0.0 raw=0.0 chain_classes_total=0 endpoints_total=0 blind_spots=container_wired findings_status=not_computed`. The symbol is `M3FiltroSeguridadAspect`, which the same build's `spring-audit` shows covering 713 of 936 handler declarations (76 %) and the same build's `posture` shows as the only access control over 939 endpoints under `default`. C1-36 floored the *band* to `high` when the container-wiring blind spot fires and left the *score* on the fan-in arithmetic, which is 0 over an empty chain. The residue also **got worse across the window**: 4.10.7 published `2.0` (a GATE-004 finding in the chain), and 4.11.0 publishes `0.0` because without `--with-findings` the enrichment is `not_computed` | 4.11.0 (eval #16) | High — a CI gate that reads the numeric field gets the *worst possible* value for the most critical symbol in the system, from the same object whose band says `high` | **closed 4.12.0** — `risk_score` and `risk_score_raw` are `null` on this population, in the payload and in the `metadata` copy, with `risk_score_basis` stating why and naming the three surfaces that do measure such a component. The sentence lives in `container_wiring`, the module that owns the fact, and **both** derivations read it: `spring_impact` (impact-chain) and `_compute_blast_radius` (impact / plan / compare), where the same defect was live and had not been reported. Only this blind spot voids the number — endpoint coverage, method scope and parse coverage leave the callers trustworthy and the score a floor, which is a claim we can defend — and the two consumers that read the field with a `0.0` default now carry the null instead of inventing a zero |
112
112
  | C1-42 | **The rule inventory disagrees with the findings printed beside it, in both directions.** Fourth round on saint-server, from a 4.12.0 run the budget cut: `metadata.rules_run: ["SEC-001","SEC-002","SEC-003","SEC-004"]` and `metadata.rules_not_run: ["GATE-001"]`, in a payload whose own findings carry `DEAD-001` (254), `SEC-008` (16), `TX-006` (8), `TX-001` (5), `SEC-004` (3), `SEC-003`/`SEC-005`/`SEC-006` (1 each). **Six families that demonstrably ran are missing from `rules_run`**, and the arithmetic against the complete 4.11.0 run shows a second family missing from `rules_not_run`: `395 − 289 = 106 = GATE-001 (105) + GATE-004 (1)` findings, `98 − 15 = 83 = 82 + 1` defects. Two causes, both structural: the lists are keyed on the **pattern object's** `pattern_id` while a pattern that emits a whole family declares `pattern_ids` (the `SEC-004` group emits SEC-004…008 + DEAD-001; the `GATE-001` group emits GATE-001…004), and `SpringAuditResult.merge` carries `metadata` with `dict.update`, so the security scope's inventory overwrites the TX scope's | 4.12.0 (eval #17) | **High** — the differential asset of this product is declaring what it did not measure, and on a truncated run `rules_not_run` is the one field a consumer reads to learn what is missing. *"En un producto cuyo valor diferencial es declarar sus huecos, un `rules_not_run` incompleto es peor que en cualquier otro sitio."* | **closed 4.14.0** — the inventory is expressed at **rule-id** granularity, from the pattern registry rather than from a pattern's identity: `rule_pass.rule_ids_of()` reads `pattern_ids` when a family declares one and `pattern_id` when it emits its own, so a group can never again stand for its members. The two lists are unioned across scopes in `merge` (order-preserving) instead of overwritten, so a `--scope all` run publishes TX and SEC together, and `rules_run_count` / `rules_total` sit beside them so a mismatch is visible without counting. Three invariants are asserted over a real audit, not over a fixture: every `pattern_id` present in `findings[]` is in `rules_run`; `rules_run ∩ rules_not_run` is empty; and on a complete `--scope all` run `rules_run` **is** `rule_catalog.ids()` — the same catalogue `--help` renders from, so a twentieth rule fails the suite until the inventory can see it. **Verified in the field (eval #18):** `rules_run` (9) ∪ `rules_partially_run` (6) ∪ `rules_not_run` (4) covers the catalogue and every `rule_id` in `findings[]` is inside the first two — *"introdujeron la categoría intermedia que faltaba en vez de forzar una clasificación binaria"* |
113
113
  | C1-40 | **`cache model` promises `onboard` an answer-level cache hit and the second run recomputes.** `ask cache model .` publishes `onboard` → *"the answer (repeat cached)"* `[task, ris]`. Measured: first run 56 s (cold, expected), **second run on a warm cache 23,5 s** — not an answer hit, a partial recomputation. Corroborated by `cache context-stats .`: `Contexts: 1 · Hits: 1 · Misses: 1 · Hit ratio: 0.5 · Avg lookup 1029,1 ms (p95 1955,2 ms) · Avg build 18332,5 ms · Bytes stored 9 969 057`. A **full second to decide whether there is a hit**, over a single 10 MB entry, is the signature of deserialising the whole payload to read its key | 4.10.4 (eval #14) | **Medium** — `cache model` exists so a user can plan instead of guess; a row that overstates what a warm buys is the one defect that command cannot afford. The lookup half is the second field measurement of C3-47 | **closed (4.10.6), and the claim was the false half.** Reproduced at home on keycloak (5 486 sources): **with** a `cache warm`, `onboard` costs 0,7 s — the row's promise holds. **Without** one, a second identical run costs 9,4 s against 10,2 s cold: the warm is what stores the answer, and the command does not store its own. The row said `repeat cached`, the field ran it twice with no warm, and got 56 s → 23,5 s. `repeat` is now `false` for `onboard` and the measured figures are published at two scales, with and without a warm, so the reader can tell which run they are about to make. Half (b) shipped too: the validity keys (`ev`/`kl`/`sig`) live in an uncompressed sidecar beside the blob, so deciding *is this entry still valid?* is a small file read instead of decompressing and parsing a 10 MB envelope — the case that matters most, an entry invalidated by a changed signature, used to pay all of it and then discard the result. A genuine hit still pays for the payload it returns; that is the answer, not overhead. Entries written before the sidecar existed fall back to the old full read. Original remedy note: two halves. (a) Reconcile the claim: either `onboard` really caches its answer, or its row says `shared work`, not `the answer`. (b) Split key metadata (hash + validity) from the payload into a sidecar index so a hit is decided without deserialising the blob; target < 5 ms against the ~1 029 ms measured. Related F-AA |
114
- | C1-43 | **"Does this repository configure a request filter chain?" has two authorities, and the one that decides whether a `high` finding family runs is the older, weaker one.** `posture._security_config_types` reads the question off the IR's own edges — a type that `extends`/`implements` a chain SPI type, **or a member that `returns` one**, i.e. the `@Bean SecurityFilterChain` shape — and the docstring beside it records why: *"a repository that wires its chain with a `@Bean SecurityFilterChain` and no `@EnableWebSecurity` reported `security.active: 0` — a false zero on the axis this command exists for."* `repository_ir._filter_based`, which decides `security_model` and therefore whether **SEC-001 runs at all**, still tests only `@EnableWebSecurity` (`_FILTER_SECURITY_ANNOTATIONS`, deliberately one entry) and `extends WebSecurityConfigurerAdapter`. A Spring Security 6 repository that declares its chain as a bean is therefore classified `annotation_based`, and SEC-001 then asserts, per unannotated handler, *"In an annotation_based security model there is no centralized filter — any caller can reach this endpoint without authentication"* | 4.18.0 (eval #24) | **High** — a confident falsehood on the strongest axis, and the only one of the series: measured on a greenfield Boot 3 module, the single SEC-001 finding is `POST /api/v1/auth/login`, which the repository's `SecurityConfig` lists in `PUBLIC_ENDPOINTS` with `permitAll()` under a closing `anyRequest().authenticated()` — the chain the model says does not exist both admits that route *and* protects the other 24. The evaluator adjudicated it by reading the source: *"la herramienta lo marca `high` porque solo lee la anotación del método, no la cadena de filtros"* | **in-progress** — the premise half shipped on `fix/sec-001-chain-authority`. One authority, `security_chain`, holds the request-chain SPI vocabulary, the member rule (*a member returning the SPI type makes its declaring class the configuration*) and the name normalisation; `posture` and both `repository_ir` sites read it, so the mechanism can no longer differ. **Two questions are named apart there rather than merged, because only one of them may silence a finding:** `declares_request_chain` (is there a chain at all — the annotation, the pre-5.7 adapter, or a member producing a `SecurityFilterChain`) is SEC-001's premise and is deliberately narrow, while `chain_participant_types` (what takes part in configuring it, filters included) is `posture`'s question. Widening the first to the second would silence SEC-001 on repositories with no centralized authorization — the P1-A failure mode, and the one with no symptom. Measured end-to-end on a fixture of the reported shape: `security_model` `annotation_based` → `mixed`, and the SEC-001 finding on `POST /api/v1/auth/login` disappears; recall asserted in the same battery (no chain → still reported; `@EnableMethodSecurity` alone → still reported). Fleet A/B over 5 repositories × 3 commands: 15/15 payloads identical in content, which is the expected result and the point of running it — no fleet repository declares a chain in either shape, so nothing could move. **Open: the verdict half.** Where the chain resolves, the per-endpoint answer should come from the resolved chain rather than from the model classification (C1-8's `matched_path` machinery already exists), and where it does not resolve, SEC-001 abstains — never *"any caller can reach this endpoint"* over a chain the run did not read |
114
+ | C1-43 | **"Does this repository configure a request filter chain?" has two authorities, and the one that decides whether a `high` finding family runs is the older, weaker one.** `posture._security_config_types` reads the question off the IR's own edges — a type that `extends`/`implements` a chain SPI type, **or a member that `returns` one**, i.e. the `@Bean SecurityFilterChain` shape — and the docstring beside it records why: *"a repository that wires its chain with a `@Bean SecurityFilterChain` and no `@EnableWebSecurity` reported `security.active: 0` — a false zero on the axis this command exists for."* `repository_ir._filter_based`, which decides `security_model` and therefore whether **SEC-001 runs at all**, still tests only `@EnableWebSecurity` (`_FILTER_SECURITY_ANNOTATIONS`, deliberately one entry) and `extends WebSecurityConfigurerAdapter`. A Spring Security 6 repository that declares its chain as a bean is therefore classified `annotation_based`, and SEC-001 then asserts, per unannotated handler, *"In an annotation_based security model there is no centralized filter — any caller can reach this endpoint without authentication"* | 4.18.0 (eval #24) | **High** — a confident falsehood on the strongest axis, and the only one of the series: measured on a greenfield Boot 3 module, the single SEC-001 finding is `POST /api/v1/auth/login`, which the repository's `SecurityConfig` lists in `PUBLIC_ENDPOINTS` with `permitAll()` under a closing `anyRequest().authenticated()` — the chain the model says does not exist both admits that route *and* protects the other 24. The evaluator adjudicated it by reading the source: *"la herramienta lo marca `high` porque solo lee la anotación del método, no la cadena de filtros"* | **closed 5.8.7** — the premise half shipped on `fix/sec-001-chain-authority`. One authority, `security_chain`, holds the request-chain SPI vocabulary, the member rule (*a member returning the SPI type makes its declaring class the configuration*) and the name normalisation; `posture` and both `repository_ir` sites read it, so the mechanism can no longer differ. **Two questions are named apart there rather than merged, because only one of them may silence a finding:** `declares_request_chain` (is there a chain at all — the annotation, the pre-5.7 adapter, or a member producing a `SecurityFilterChain`) is SEC-001's premise and is deliberately narrow, while `chain_participant_types` (what takes part in configuring it, filters included) is `posture`'s question. Widening the first to the second would silence SEC-001 on repositories with no centralized authorization — the P1-A failure mode, and the one with no symptom. Measured end-to-end on a fixture of the reported shape: `security_model` `annotation_based` → `mixed`, and the SEC-001 finding on `POST /api/v1/auth/login` disappears; recall asserted in the same battery (no chain → still reported; `@EnableMethodSecurity` alone → still reported). Fleet A/B over 5 repositories × 3 commands: 15/15 payloads identical in content, which is the expected result and the point of running it — no fleet repository declares a chain in either shape, so nothing could move. **Verdict half closed 5.8.7**, and one of its two branches turned out to be already held. *"Where the chain resolves, the per-endpoint answer comes from the resolved chain"* is what the premise half did: a resolved chain makes the model `filter_based`/`mixed`, and SEC-001 does not run there at all — `posture.endpoint_access` is the resolved-chain answer and it is published under its own name. So the branch that was genuinely open is the other one, and it is the one with the confident falsehood in it: where **no** chain was found, the rule narrated an absence as a verdict — *"there is no centralized filter — any caller can reach this endpoint without authentication"*, at `confidence: high`, derived from not finding the two shapes it knows how to look for. It now states what was measured (no guard on the handler, and no `@EnableWebSecurity`, no `WebSecurityConfigurerAdapter` subtype, no member returning a `SecurityFilterChain`) and says plainly that this is the absence of a guard rather than a demonstration that the route is reachable; `evidence.chain_evidence_tested` publishes the premise from `security_chain`'s vocabulary so a reader can check it instead of taking it. And the coverage question is bound to the authority that owns it (ADR-0008 R1): every SEC-001 finding carries `security_verdict` / `security_verdict_confidence` / `security_verdict_basis` from `security_posture.endpoint_security_surface`, and where that authority answers anything other than `probably_exposed` the finding drops to `confidence: low` and says which verdict it is standing beside — the rule stops outranking the authority it is supposed to read. Severity is untouched: it is a published contract, and moving it belongs in a release note (the C3-101 (c) precedent). |
115
115
  | C1-44 | **`has_uncommitted_changes: false` sits in the same document as `git_context.uncommitted_files: 1`, and it is the field cache freshness is decided on.** `ask <repo> --compact --git-context` publishes both, one run, one payload: `git_context.uncommitted_files` counts the file, `_cache.has_uncommitted_changes` says there is none, and `git status --porcelain` on the same tree prints `?? .claude/settings.json`. The third authority in the build agrees with the first — `--changed-only` reports `changed_files_count: 1` for that file and its `--help` defines the set it counts as *"staged, unstaged, untracked"* — so two of three surfaces admit the file and the one a consumer trusts for staleness does not. Under that published definition the boolean is simply false | 5.3.1 (audit #28) | Medium | **closed 5.4.0** — one predicate (`baseline_autocapture.counts_as_dirty`), extracted from `worktree_dirty` and read by the count in both the fresh path and the cache-hit patch, so the boolean cannot drift from the function that defines it. What it excludes is published (`uncommitted_files_not_read`, `uncommitted_files_basis`) rather than subtracted in silence, and only when something was actually excluded. Original remedy note: one predicate over one definition of *the working set*, read by both fields; the boolean is derived from the count rather than measured a second time, and where the two cannot be derived from one another the document may not publish both |
116
116
  | C1-46 | **`existing_test_count: 6` sits beside `test_files: 0` and `has_test_sources: false` in one `prepare-context generate-tests` payload, and the 6 are five production files in a business package.** C3-115 closed the half that was reported — `test_gaps` now declares its population — and left a second authority for *"how many tests are there"* answering a different number in the same document. Re-measured in the field: 0 Java tests, 1 `.spec.ts`, 3 karma/jest configs; no population is worth 6. Root cause is in `test_sources.declared_test_root`: `_TEST_DIR_NAMES` carries the bare segment `it`, the final loop accepts it **at any depth**, and the only disarm — `_under_main_root` — knows `src/main` alone, which is a Maven layout. An Angular app lives under `src/app`, so `saint-client/src/app/shared/akita/it/` — *Incapacidad Temporal*, a business entity whose `it.model.ts` declares `ENFERMEDAD COMÚN` and `ACCIDENTE DE TRABAJO` — is read as a test source root and contributes 5 files; the 6th is the one real `.spec.ts`, and the reconciliation is exact. **It generalizes off this repository:** `it` is also the ISO 639-1 code for Italian, so `web/src/assets/i18n/it/messages.json` and `app/locales/it/common.ts` both answer `is_test_path: true` on the shipped build. Sibling of C1-28, in the direction C1-28 was written to prevent | 5.4.0 (audit #29) | **Medium** — the path-classification half is closed in 5.4.1; `existing_test_count_unit` now declares its all-path, multi-language task population in 5.4.2, distinct from the Java IR population | **closed 5.4.2** |
117
117
  | C1-45 | **The SQL taint join keys a sink on a bare statement id, so a mapper method in one namespace is joined to a call site that resolves to another — and the result is published as the repository's #1 critical risk with a call trace that does not exist.** `risk`'s top row on a 3 342-file monolith is `severity_effective: 29.57` at `ProcesosGerenciaMapper.xml:74`, `factors.query_construction = http_input_reaches_sql_interpolation`, with evidence naming `AutocoberturasRestController#actualizar(id, obj)` *calls* `actualizar(…)` in `ProcesosGerenciaMapper`. The controller calls its own service's `actualizar`; the join matched on the method name alone. **The correct key already exists in the code and is not used** — `sql_taint` line 76 returns `f"{self.namespace}#{self.statement_id}" if self.namespace else self.statement_id`, and `namespace` is captured (141, 151) and then spent only on an explanatory sentence (313-314), while the published basis claims the sinks are *"keyed by namespace and statement id"*. **Accompanying defect, same row:** the interpolated expression is stored (72, 144) and emitted (86) and never inspected, so `bloqueada = ${dto.bloqueada ? "'S'" : "'N'"}` — an OGNL ternary whose two branches are string literals, a closed allowlist by construction and precisely the mitigation SEC-008's own `fix_hint` recommends — is ranked as a text-controllable sink | 5.3.1 (audit #28) | **High** — F-BF's whole argument is that this axis fires on evidence rather than on multiplication by 1,0; a name-collision join inverts the first row of the ranking, which is the product | **closed 5.4.0** — both sides keyed on `namespace#statement_id`: a call carries the type its receiver resolves to and is that statement's only when the receiver resolves to the mapper it lives in; an unresolved receiver is not the same answer as any type name, so the sink stays `undecidable` rather than `reaches`. `does_not_reach` carries the same qualification, because it is a confident no. `splices_free_text` reads the interpolation at last — anything it cannot prove closed is free text, and what it closes leaves the population into a published `bounded_interpolations` list. The regression fixture this row specified is in the suite. Original remedy note: key both sides on `namespace#statement_id` (the property is already written, the call sites read the wrong half), and give the interpolation a three-state read: an expression whose reachable values are all literals is `does_not_reach`, an unresolvable one stays `undecidable`, never `reaches` by default. Regression fixture: two mappers in different namespaces sharing a `statement_id`, one reached by HTTP input and one not — expected one `reaches`, one `undecidable`; today, two `reaches` |
@@ -245,7 +245,7 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
245
245
  | C3-81 | **A family the budget cut in flight publishes `units_done == units_total`, so its own entry says it finished.** Fifth round on saint-server, verbatim from 4.14.0: `rules_partially_run: [{"rule_ids": ["SEC-004","SEC-005","SEC-006","SEC-007","SEC-008","DEAD-001"], "units_done": 3342, "units_total": 3342, "unit": "java files"}]` — a complete population, in a list whose meaning is *"this family did not finish"*. It is not a display artefact: `DEAD-001` yields **248** findings against **254** on the complete run, so six findings really are missing. The cause is that the family walks **four populations in sequence** (`security_config_scan`: Java sources, then configuration files, then deployment descriptors, then MyBatis mappers) and `FamilyBudget` records the counts of the stage it happened to be in. The deadline fell on the **last** Java file, so the entry names the one population that did finish and says nothing about the three that never started — the same shape as C1-42 one level down: a container standing for its members | 4.14.0 (eval #18) | Low — the numbers beside it are honest (`partial`, `counts_are_floor`, `confidence: low`), and this is the one field a reader consults to decide *how much* more time to give the run | **closed 4.15.0** — the unit is the stage. Every walk a family makes now ticks with what it has **finished** and names itself (`scanning java sources`, `scanning configuration files`, `scanning deployment descriptors`, `scanning mapper files`, `walking endpoints`, `resolving gate annotations`, and on the TX side the declared-boundary and method-body walks), so a stage in flight can never report its own population complete — the literal field contradiction is asserted unreachable at every cut point, not at one. The entry carries `stage_stopped_in` and a `units_basis` stating that the two numbers count the stage rather than the family, and the `PARTIAL:` sentence a human reads carries the same fact as the payload. Off-by-one was the second half and the invisible one: the counter reported the unit it was *about to* process, so a walk stopped at its own last file published a finished population |
246
246
  | C3-82 | **The heartbeat is silent for 13,6 minutes *between* batch boundaries, and the batch is the sampling constant a different row chose.** C3-80 made the work emit in-band and the field confirms the mechanism arrived — the line now carries the family, a per-file count and an ETA. It also measures what it did not fix: `elapsed=5.0s stage=linking 3206/3342` · `elapsed=10.0s stage=security rules (SEC-004) 1856/3342 java files` · **13,6 minutes of nothing** · `elapsed=13.6m … 3328/3342 eta=3.4s` · `done (13.6m)`. The emitter is reached only from `RulePassProgress.tick()`, which returns early on `done % sample_every` — so between two boundaries 64 units apart nothing can print, and a stretch of expensive units (or one pathological file, which the 4 files/s average implies) is invisible for as long as it lasts. The 39-minute hole became a 13,6-minute hole because the *work* got shorter (C3-79), not because the emission got denser | 4.14.0 (eval #18, partial since 4.10.7) | Medium — fifth version of the same question (*"is it working or is it hung?"*) going unanswered on the most expensive command, now with the content already correct | **closed 4.15.0** — one number was doing two jobs, and now each has its own. The **clock** is still read once per `sample_every` units, which is C3-79's reason unchanged and what took the overshoot from 5,66× to 1,94×; the **counter** reports on every unit, because `Progress` already throttles emission on an interval, so a per-unit report costs one clock read and produces a line only when one is due. Asserted both ways: a thousand cheap units print at most one line, and five units that each outlast the interval print five — the shape the field watched go silent |
247
247
  | C3-83 | **A budget inherited from the environment truncates the audit and nothing in the answer says where the limit came from.** Measured both ways on the same build: without `ASK_MAX_ANALYSIS_SECONDS`, **395 findings / 98 defects**; with `=420`, **267 / 11** — GATE-001's 82 gate-bypass defects among what is missing. Everything about the truncation is declared impeccably (`_partial`, `summary.partial`, `counts_are_floor`, `confidence_level: low` with its basis — C2-31 and C3-71 closed that), and the residue is provenance: the variable is inherited from a shell or a runner, so the person reading `total_defects: 11` is often not the person who set it, and nothing in the payload names it. *"Quien fije la variable en CI y solo lea `total_defects` verá 11 donde hay 98"* | 4.14.0 (eval #18) | Medium — a configuration trap rather than a false claim, and the trap is on the axis (audit completeness) the product is bought for | **provenance half closed 4.15.0; the contract halves stay F-AV.** The budget carries its origin verbatim (`ASK_MAX_ANALYSIS_SECONDS=420 (process environment)`) from the place it is read to the places it is published: under the class floor the measured warning C3-72 shipped gains a line naming the variable and its value, over the floor — where there is no advice to give — a single provenance line replaces silence and says what a run that hits the limit publishes and how to get a complete audit, and `_partial` carries `budget_source` beside `budget_seconds`. A value that does not parse is not a budget and invents no source, and a budget that was not inherited publishes no origin. Verified end to end: `spring-audit` under a 0,01 s budget publishes `budget_source` in `_partial`. The exit code and any renaming of a truncated total remain **F-AV** |
248
- | C3-84 | **`--jobs` parallelises the phase that is not the cost, and the help implies otherwise.** The evaluator corrected their own four-round diagnosis this round (*"reporté 'sin paralelismo'. Es inexacto"*): `--jobs` / `ASK_JOBS` exist and default to `cpu_count()-1` (=19 there). What the OS sampling shows is which phase they cover — `t=7,6s threads=4` during parsing, `t=441,1s threads=1 CPU/wall 0,97` for everything after — matching the run's own stages exactly: `linking` ~10 s parallel, `security rules … 3342 java files` ~13,6 min serial. On a **warm** cache, which is the normal case and the one the field measures, there is nothing left to parse, so `--jobs` buys nothing at any value while 19 workers sit idle. The flag's help is honest about *what* it parallelises (*"Parallelism only warms the content-addressed parse cache — the answer is byte-identical at any value"*) and says nothing about the phase that holds ~99 % of a repo-wide run | 4.14.0 (eval #18, fifth round on the cost) | Medium as a claim, **Critical as a capability** — it is the single item the field has priced the CI verdict on for five rounds: *"con 16 workers, `spring-audit` bajaría de 13,6 min a ~1 min y este producto pasaría a otra categoría"* | **claim half closed 4.15.0; the capability is F-AT.** `JOBS_OPTION_HELP` — the single string every command that publishes the flag declares it with — now states that the rule pass is **not** parallelised and that a warm cache leaves the flag nothing to parse, with F-AQ's promise (byte-identical at any value) untouched. The README and the user guide are qualifications of that one claim rather than three versions of it, and the battery asserts every `--jobs` parameter in the registry carries it, so a second copy cannot keep the old sentence. The **capability** is **F-AT** (parallel rule evaluation, with the map-reduce shape the field spells out: per-file families map and reduce on the `(category, defect_kind, symbol)` identity `defect_identity` already publishes; graph-scope families such as GATE-001 stay in the serial tail) and does not start before this queue finishes ⚠ **The remaining half — parallelise the rule pass — is refuted by measurement (5.8.6), and the figure is published so nobody re-derives it.** Instrumented with `metadata.timings` on BroadleafCommerce (2 766 `.java`): `cir_build` **8 711 ms of a 9 915 ms run (88 %)**, `security_rules` 740 ms (7,5 %), `tx_rules` 53 ms (0,5 %), `semantic_model` 61 ms, `waivers` 0,04 ms. **Perfect parallelism of every rule would save at most ~8 % of the clock**, not the *«mayor palanca absoluta»* two field rounds estimated — which is ASK-16's finding again, one level down: the cost is the build, not the rules. The lever that remains is the CIR build, and `--jobs` already parallelises the parse inside it. Kept open only as the *disclosure* half: the help still implies the flag parallelises the audit. |
248
+ | C3-84 | **`--jobs` parallelises the phase that is not the cost, and the help implies otherwise.** The evaluator corrected their own four-round diagnosis this round (*"reporté 'sin paralelismo'. Es inexacto"*): `--jobs` / `ASK_JOBS` exist and default to `cpu_count()-1` (=19 there). What the OS sampling shows is which phase they cover — `t=7,6s threads=4` during parsing, `t=441,1s threads=1 CPU/wall 0,97` for everything after — matching the run's own stages exactly: `linking` ~10 s parallel, `security rules … 3342 java files` ~13,6 min serial. On a **warm** cache, which is the normal case and the one the field measures, there is nothing left to parse, so `--jobs` buys nothing at any value while 19 workers sit idle. The flag's help is honest about *what* it parallelises (*"Parallelism only warms the content-addressed parse cache — the answer is byte-identical at any value"*) and says nothing about the phase that holds ~99 % of a repo-wide run | 4.14.0 (eval #18, fifth round on the cost) | Medium as a claim, **Critical as a capability** — it is the single item the field has priced the CI verdict on for five rounds: *"con 16 workers, `spring-audit` bajaría de 13,6 min a ~1 min y este producto pasaría a otra categoría"* | **claim half closed 4.15.0; the capability is F-AT.** `JOBS_OPTION_HELP` — the single string every command that publishes the flag declares it with — now states that the rule pass is **not** parallelised and that a warm cache leaves the flag nothing to parse, with F-AQ's promise (byte-identical at any value) untouched. The README and the user guide are qualifications of that one claim rather than three versions of it, and the battery asserts every `--jobs` parameter in the registry carries it, so a second copy cannot keep the old sentence. The **capability** is **F-AT** (parallel rule evaluation, with the map-reduce shape the field spells out: per-file families map and reduce on the `(category, defect_kind, symbol)` identity `defect_identity` already publishes; graph-scope families such as GATE-001 stay in the serial tail) and does not start before this queue finishes ⚠ **The remaining half — parallelise the rule pass — is refuted by measurement (5.8.6), and the figure is published so nobody re-derives it.** Instrumented with `metadata.timings` on BroadleafCommerce (2 766 `.java`): `cir_build` **8 711 ms of a 9 915 ms run (88 %)**, `security_rules` 740 ms (7,5 %), `tx_rules` 53 ms (0,5 %), `semantic_model` 61 ms, `waivers` 0,04 ms. **Perfect parallelism of every rule would save at most ~8 % of the clock**, not the *«mayor palanca absoluta»* two field rounds estimated — which is ASK-16's finding again, one level down: the cost is the build, not the rules. The lever that remains is the CIR build, and `--jobs` already parallelises the parse inside it. **Disclosure half closed 5.8.7**, and the correction went the opposite way to the one the row predicted. The 4.15.0 help did not imply the flag parallelises the audit — it said the rule pass is *"where a repository-wide audit spends most of its time"*, which was the estimate everyone held, us included, and which the 5.8.6 measurement refutes. So the falsehood in the surface was **our own qualifier**, not the original promise: the phase `--jobs` covers is the one that holds 88 % of the clock. All three surfaces now carry the measurement instead of the estimate (`JOBS_OPTION_HELP` — the one string every command declares the flag with — the README and the user guide), and `JH-06` fails on any of them restating the refuted claim or dropping the number. What limits the flag is stated as what it is: a warm cache leaves nothing to parse. The same refuted sentence was also carried by the rule-pass observer's docstring and is corrected there. |
249
249
  | C3-69 | **Mojibake on the Windows console: we write correct UTF-8 bytes and the console decodes them in the ANSI codepage.** Field payload, verbatim: `"analysis_warnings": ["Self-referential exclusion: 1 member(s) ... â€" a class's own methods ..."]` — U+2014 arriving as `â€"`. The same corruption appeared in the baseline when read back from Python. Verified in source: the shared emit seam (`cli.py:1219`) writes `content.encode("utf-8")` to `sys.stdout.buffer`, so **our bytes are right**; every `json.dumps` in the package uses `ensure_ascii=False`. The defect is real to the consumer regardless of where it is produced: the JSON stops being valid UTF-8 *as received* | 4.10.4 (eval #14) | **Low** — it breaks downstream parsing, and the diagnosis matters more than the severity: this is the output-side mirror of C3-58, which was closed on the input side by letting bytes name the encoding instead of a locale | **closed (4.10.6), and not by escaping everything.** Encoding the stream harder cannot fix a *decoder*, so the first move is to fix the decoder: `SetConsoleOutputCP(65001)` at entry tells the Windows console to read UTF-8, and when it works nothing else changes — the payload keeps its em-dashes and accents on every platform, byte-identical to today. Only when that call fails does the emit seam escape the JSON to pure ASCII, which is the subset every codepage agrees on, so `\u2014` arrives intact and `json.loads` gives the em-dash back. The decision is made **once**, at entry, in `output_encoding.py`, and every emit seam reads it (`_serialize_dict`, `serializer.to_json`, the error and `_meta` envelopes) — no command can make a different one, and redirected output never involves a console codepage so it never takes the fallback. Original remedy note: the whole class dies with `ensure_ascii=True` on the output `json.dumps` (ASCII is a subset of every codepage the console might choose), at the cost of `\uXXXX` escapes in a payload the field also reads by eye. Alternative: set the console codepage on Windows at entry, and document it. Decide once, for every emit seam, and lock it in the battery |
250
250
  | C3-85 | **The stretches between counted stages report nothing, and the frozen `n/n` left on screen reads as the stage still running.** Sixth round on saint-server, verbatim from 4.15.0: `elapsed=10.0s stage=security rules (SEC-004) 37/3342 java files`, then **14,5 minutes with no line at all**, then `elapsed=14.5m stage=security rules (SEC-004) 3340/3342 java files eta=0.5s`. The field reads this as the periodic emitter being dead and proposes re-arming it. **Measured here, the emitter is alive and the diagnosis is one seam further in.** A per-unit tick loop driven through the real seam (`FamilyBudget.tick` → `RulePassProgress.detail_sink` → `Progress.work`) emits one line per interval, exactly as designed: 100 units over 5,4 s at a 1 s interval produced **5 lines, evenly spaced**. What produces the silence is the stretches no counter covers. Instrumented on keycloak (5 486 files, warm, every `work`/`step`/`update` call timestamped): **8,2 s of a 24 s run — 34 % — pass between the last `linking` tick and the first rule tick**, with `linking 5486/5486` frozen on screen throughout; 2,7 s more inside `GATE-001` between two of its own ticks; and on a cold run **24,7 s after `parsing 50/5486`**. Scaled to the field's repository and its per-file cost, those are the minutes they watch | 4.15.0 (eval #19, sixth round on the emission) | Medium — it is the question *"is this working or is it hung"* on the command that costs the most, and the answer currently on screen is a count that finished | **closed 4.16.0** — every stage of an audit now either counts its population or names itself, and no run ends a counted pass without announcing what follows it. Named: the IR build's tail (`assembling the IR`, `recovering spec-declared routes`, `scanning XML security configuration`), the semantic-model build, the whole-tree walk *inside* the configuration family (through `FamilyBudget.stage`, so the line says which family it belongs to), the rule pass's dedup-and-order tail, and each of the five repository-wide walks `risk` performs before it composes. Inventing a denominator for a single walk is the fabrication C3-56 refused, so none is invented. Re-measured on keycloak after the fix: the 8,2 s that read `linking 5486/5486` now read `assembling the IR` and `scanning XML security configuration`, and NS-05 asserts the invariant directly — a build may not end on a completed counter. The naming is also what lets the *next* round say which walk the wall time is in, which is C3-88 |
251
251
  | C3-86 | **A count that has not moved is re-emitted as live progress, with an ETA extrapolated across the stall.** The same field line is the evidence: `3340/3342 java files eta=0.5s`, printed **14,5 minutes** after the counter last advanced. Both halves are false claims of the kind this product does not otherwise make — the line asserts the stage is 99,9 % done *now*, and the ETA is a rate computed over a window in which nothing happened, so the longer the stall lasts the more imminent the finish looks. `_eta_seconds` divides the remaining units by `done / (now − stage_t0)`, and `now` keeps advancing while `done` does not | 4.15.0 (eval #19) | Medium — C2 in nature: presentation asserting something measurement did not establish. It is the one place in six rounds where the terminal states a fact the run cannot support | **closed 4.16.0** — a count that has not advanced for a whole interval publishes `unchanged_for=<duration>` and withholds the ETA, in both line mode and on the spinner. The staleness clock is deliberately **not** `_stage_t0`: that one answers *"when did this stage start"*, and conflating the two would mark a long, healthy, advancing stage as frozen — asserted. The count itself is still printed, because it is true and it is what a reader wants; what is withheld is the claim that it is *current*, and what is added is how long it has stood there. An uncounted stage is never called stale: it never claimed a denominator, which is C3-56's refusal to fabricate one |
@@ -265,7 +265,7 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
265
265
  | C3-98 | **The release performance gate has been armed since 4.13.0 and has never had a baseline to compare against.** F-AP shipped the detector — `.github/workflows/perf-gate.yml` (nightly `cron: 0 3 * * *`), `scripts/perf_gate.py`, and one authority for every constant in `sourcecode.perf` (`RELEASE_GATE_THRESHOLD = 0.20`, `RELEASE_GATE_REPO = "broadleaf"`, `RELEASE_GATE_MODE = "warm"`, `RELEASE_GATE_MIN_RUNS = 5`, six gated commands, three published exclusions) — and the job even proves itself armed before comparing (`--self-test` fails if a synthetic 3× slower cell would pass). The comparison then reads `BASELINE_DIR: docs/perf/baselines/${{ inputs.baseline_ref || 'gate-latest' }}`, and **`docs/perf/baselines/gate-latest/` does not exist**: the directory holds `2.5.15` and `2.5.16` only, both root-`ask` cells, measured nine minors before the gated population was defined. The field has now written *"sigo sin evidencia de que exista un gate que gatee releases"* in three consecutive rounds and is right for a reason no report could see: the gate exists, the baseline does not, so no release between 4.10 and 4.18 was gated by it | 4.18.0 (eval #22, found in-house while dimensioning their B10) | High — it is the control that would have stopped 4.11.0 and 4.16.0, and "the gate is green" is unfalsifiable while it has nothing to compare | **in-progress** — the baseline exists as of 2026-08-11. Confirmed from the field first: the manually triggered run [31464236961](https://github.com/HarounDominique/ASK-Engine/actions/runs/31464236961) at `0305171` measured for 4m11s and then failed with the gate's own words — *"no baseline cells in `docs/perf/baselines/gate-latest` — every cell would be recorded as `new` and the run would pass without comparing anything. A gate that cannot be red is not evidence."* The hard-failure half was therefore **already correct** (`--require-baseline`); what the run also proves is that it measures **before** it compares, so the red run uploaded the six cells as an artifact. Those cells — 4.18.0, warm, 5 runs each, `env.host_id_declared: true` for the `github-ubuntu-latest` class — are now committed as `docs/perf/baselines/gate-latest/` with their provenance (`README.md` in that directory), and the gate run on that branch ([31465902639](https://github.com/HarounDominique/ASK-Engine/actions/runs/31465902639)) **passed with 6 cells compared, `regressions: []`, `missing_cells: []`, `under_sampled_cells: []`** — the first run in the gate's existence that measured, compared and returned a verdict. **On `master` since `e9696fa` (2026-08-11)**, so the next scheduled run (03:00 UTC) is the first nightly in the gate's history with both halves in place; the row closes on that verdict rather than on the merge. **One measured fact to carry forward:** with nothing changed but the runner, the six deltas spread −5,1 % … +5,0 %, widest on the shortest cell, so ±5 % is this environment's noise floor and the 20 % threshold is set against it — it catches the 2–4× swings the field reported and would not catch a 15 % regression on a half-second command |
266
266
  | C3-99 | **The gate's own exclusion list quotes a cost that a later release deleted.** `perf.RELEASE_GATE_EXCLUSIONS` keeps `risk` out of the gated population because *"54 minutes on the gate repository (F-AQ); a nightly cannot hold it, and the cost is structural rather than a regression to detect"*. C3-96 closed in 4.18.0 and `risk` on that same repository (BroadleafCommerce, 2 985 Java files) now takes **68 s** — the premise of the exclusion no longer exists, and the excluded command is the most volatile cell in the whole series (947,5 → 3 232,2 → 3 296,6 → 2 014,5 → 2 927,1 → 32,1 s across eight field rounds). Same shape as C3-97: a measurement outliving the build it was taken on and continuing to drive a decision | 4.18.0 (in-house, from the same sweep) | Medium — the population the gate publishes as *"what the gate does not cover, and why"* is honest about the omission and wrong about the reason | **closed 5.0.0** — the reason states the measured 68 s, that C3-96 is what deleted the 54 minutes, and that the only thing still keeping `risk` out is the absence of a baseline cell for it — the first capture that includes it makes it gated. `docs/perf/REGRESSION-GATE.md` §3 says the same in prose and records the general rule: an exclusion is a claim about cost and ages exactly like an anchor |
267
267
  | C3-100 | **Four cost anchors were left behind by the 5.0.0 refresh, and two of them publish *"the field could not run this"* about commands the field has now run twice.** C3-97 re-measured what eval #22 measured and stamped `field_measured_version="4.18.0"` on it. What eval #22 did not run kept its 4.10.4 cell: `ask` (root) `field_seconds=115.0`, `baseline` `11.0`, and — worse, because the claim is categorical rather than numeric — `posture` and `repo-ir` still carry **`field_blocked=True`**, which the model publishes as a command the field was unable to complete. Eval #23 ran all four from a purged cache on the same machine: root `--compact --git-context --env-map` **72,3 s cold**, root `--agent` **71,7 s**, `repo-ir` **9,8 s**, `baseline capture` ≈ **2,2 s**, and `posture` three ways inside a 97,9 s Java/Spring phase (eval #22 had already measured `posture --diff` at 6,3 s and `--resolve-env` at 8,0 s, so `field_blocked` was false at 5.0.0 too) | 5.0.0 (eval #23) | **High** — same class as C3-97 and the same authority, but the failure is worse in kind: a *number* that is stale is a stale measurement, while `field_blocked` on a command that runs in 9,8 s is the product asserting an incapacity it does not have, on the axis (epistemic honesty) the field scores 10/10 | **closed 5.0.1.** ⚠**One quarter of the row's premise did not survive the source check:** `posture` was already refreshed by C3-97 and carries `field_seconds=8.0, field_measured_version="4.18.0"` at HEAD — the pair still publishing `field_blocked` was `repo-ir` and **`verify-edit`**, which the report did not name. The four cells refreshed from eval #23 and stamped `5.0.0` are therefore `ask (root)` 115 → **72,3 s**, `baseline` 11 → **2,2 s**, `repo-ir` *did not finish* → **9,8 s**, `verify-edit` *did not finish* → **76,6 s**; two verdicts move with them (`repo-ir` and `baseline capture` become `runs_now` at field scale, `verify-edit` stays `detach` on a figure rather than on an outcome, which is C3-102's row). **The mechanism, because refreshing the cells alone reproduces the row at the next evaluation:** (1) `field_blocked` is dated wherever it is published — the verdict now reads *"did not finish in an interactive session **on 4.10.4** … and nothing has re-measured it since"* — and the battery rejects an undated one, so an outcome expires with its build exactly as a duration does; (2) every field figure states the **cache state** it was taken in (`warm`/`cold`/`mixed`), because eval #22 measured a warm machine and eval #23 purged everything first, and 72,3 s cold against 0,3 s warm are two facts about one command; (3) `FIELD_ANCHOR`, `FIELD_ANCHOR_SECONDS` and `FIELD_ANCHOR_MEASURED_VERSION` are **derived from the `spring-audit` row** instead of written out beside it — the missed copy is precisely what this row is — and `field_currency()` publishes `rows_behind_the_newest_field_measurement`, naming every cell older than the newest measurement *in the same table*, which is the staleness nothing could see: not "older than the build you are running" (every row is, the day after a release) but "older than the refresh that reached the row next to it". Nine assertions in `TestEveryFieldMeasurementCanBeRead` and `TestTheMeasurementDecidesNotTheClass` fail on the parent commit, verified with `git stash`; no analysis path is touched. ⚠Side effect worth naming: the cache-state axis closes a clause **C4-19 left open in 4.10.7** — *"4 925 tokens in 115 s against ~15 s warm, so the figure is cache-state-dependent and the help says nothing about that either"* — which had been true of every field figure since |
268
- | C3-101 | **DEAD-001 cannot tell a control that was switched off from a control that is documented, and on a well-documented repository it is the documentation that is reported.** `_scan_commented_controls` walks every span `source_text.commented_spans` returns and reports any `@PreAuthorize`/`@Secured`/… inside it as *"a security annotation is present only in a Java comment … the running code will not enforce"*, `severity: high`, `confidence: high`. `commented_spans` makes no distinction between `//`, `/* */` and `/** */`: a Javadoc block that names the annotation it is about — which is what good documentation of an authorization decision looks like — is scanned identically to a commented-out annotation. Neither of the two discriminators the file already has access to is applied: **Javadoc is documentation by construction** (a disabled control is not written in `/** */`), and **the annotation is live a few lines below** in the same declaration | 4.18.0 (eval #24) | **High** — 8 of 8 `high` findings on the subject module were this, 0 true positives. The evaluator verified each against the source: in `PersonaRestController` the warning is line 23 and the live `@PreAuthorize`s are lines 87, 102 and 108; the same shape in `MethodSecurityConfig`, `RbacPermissionEvaluator`, `SecurityConfig`, `GlobalExceptionHandler`, `Recurso`; and in `UpsertPersonaService` the flagged comment *states* that the check is programmatic via `RbacPermissionEvaluator`, i.e. the exact opposite of a dead control. Their conclusion is the cost: *"un repo que documenta bien su seguridad dispara esta regla precisamente por documentarla bien"*, and *"entrena al equipo a ignorarla — la próxima vez que acierte, nadie mirará"* | **in-progress** — (a) and (b) shipped on `fix/dead-001-javadoc`: comment lexis stays one authority (`source_text.is_javadoc`, `/**/` excluded the way `javadoc` excludes it) and DEAD-001 consumes it, plus a second discriminator that reads the declaration head the comment sits on — window ends at the first `{` or `;`, on comment-blanked source, so a control commented out on one member is never excused by a live one on the next. All four field shapes are regression-asserted, including the two that must keep firing (a bare commented-out annotation, and `/* */` as against `/** */`); the three new negative assertions fail on the parent commit, so the battery is not vacuous. **(c) is deliberately not in this change:** `DEAD-001: high` is a published severity, and moving it is a contract change that belongs in a release note, not in a precision fix. It stays open here so the decision is visible rather than absorbed |
268
+ | C3-101 | **DEAD-001 cannot tell a control that was switched off from a control that is documented, and on a well-documented repository it is the documentation that is reported.** `_scan_commented_controls` walks every span `source_text.commented_spans` returns and reports any `@PreAuthorize`/`@Secured`/… inside it as *"a security annotation is present only in a Java comment … the running code will not enforce"*, `severity: high`, `confidence: high`. `commented_spans` makes no distinction between `//`, `/* */` and `/** */`: a Javadoc block that names the annotation it is about — which is what good documentation of an authorization decision looks like — is scanned identically to a commented-out annotation. Neither of the two discriminators the file already has access to is applied: **Javadoc is documentation by construction** (a disabled control is not written in `/** */`), and **the annotation is live a few lines below** in the same declaration | 4.18.0 (eval #24) | **High** — 8 of 8 `high` findings on the subject module were this, 0 true positives. The evaluator verified each against the source: in `PersonaRestController` the warning is line 23 and the live `@PreAuthorize`s are lines 87, 102 and 108; the same shape in `MethodSecurityConfig`, `RbacPermissionEvaluator`, `SecurityConfig`, `GlobalExceptionHandler`, `Recurso`; and in `UpsertPersonaService` the flagged comment *states* that the check is programmatic via `RbacPermissionEvaluator`, i.e. the exact opposite of a dead control. Their conclusion is the cost: *"un repo que documenta bien su seguridad dispara esta regla precisamente por documentarla bien"*, and *"entrena al equipo a ignorarla — la próxima vez que acierte, nadie mirará"* | **closed 5.8.7** — (a) and (b) shipped on `fix/dead-001-javadoc`: comment lexis stays one authority (`source_text.is_javadoc`, `/**/` excluded the way `javadoc` excludes it) and DEAD-001 consumes it, plus a second discriminator that reads the declaration head the comment sits on — window ends at the first `{` or `;`, on comment-blanked source, so a control commented out on one member is never excused by a live one on the next. All four field shapes are regression-asserted, including the two that must keep firing (a bare commented-out annotation, and `/* */` as against `/** */`); the three new negative assertions fail on the parent commit, so the battery is not vacuous. **(c) decided and closed 5.8.7 — the severity stays `high`.** It was left open on purpose so the decision would be visible, and the decision is that **the severity was never the defect; the precision was**. What fires after (a) and (b) is a control commented out with `//` or `/* */` on a declaration carrying no live control of its own — a security decision somebody reversed without removing the evidence, which is what `high` is for. Lowering it would weaken the true positive to compensate for a false positive that no longer occurs, and it is a published contract a consumer gates CI on. Pinned by a regression beside the reason, so the next change to it is a decision rather than a drift. Row closed: (a), (b) and (c) all resolved |
269
269
  | C3-102 | **`verify-edit` costs 76,6 s to answer *"you have not edited anything"*, and it is the one command whose value is bounded by its latency.** C3-32 closed the second half of this — on a tree git reports unmodified the working CIR is the HEAD CIR by construction (`working_cir_reused`) — and what is left is the first half: `_head_cir` builds the HEAD side by materialising HEAD in a **throwaway detached worktree** (`_build_head_cir_via_worktree`) and parsing it there, under a cache keyed on the HEAD sha in a deliberately separate lane from the shared repo-wide CIR. On a clean tree the two trees are provably identical content, and the shared CIR for that content already exists — built by whatever else ran on the repository — yet nothing joins them | 5.0.0 (eval #23) | **Medium** — the most expensive command of the 45, above root `--agent` (71,7 s) and `contract-diff` (51,7 s), for 807 bytes of output, and it is the command built for the edit loop (`--install-hook`, pre-commit). A short-loop gate that costs 76 s is not run: *"un gate de bucle corto que tarda 76 s no se usa"* | **closed (unreleased)** — when the tree is clean and git answered, the HEAD side reads the shared entry instead of checking out a throwaway worktree, and writes what it builds into **both** lanes (shared, so the analysis is not a private copy; HEAD-sha, so the first edit after a clean run is still warm). ⚠**Measured, and the number is small:** BroadleafCommerce, clean tree, every layer purged, 10,32 s → **10,03 s**; warm 1,51 s either way. The two lanes already collided on a clean tree — the worktree signature there *is* the HEAD sha — so the field's 76,6 s is one build of a 3 342-file repository, not a duplicated one. What changes is that the build is no longer written where only `verify-edit` can read it: the HEAD-sha entry carried no `analysed_root`, so `peek_cir` refused it and every other command rebuilt what this one had just paid for. A clean-tree run also stops writing `.git/worktrees` in the user's repository. **The second remedy stays open** — scoping the working side to the touched files on a dirty tree is the lever that would make this an edit-loop gate, and it is not in this change. Original plan: when the tree is clean *and* git answered (the `git_answered` guard C3-32 already established), the HEAD CIR **is** the worktree CIR: reuse the shared entry instead of checking out and re-parsing. Second, on a dirty tree, scope the working side to the touched files the way `pr-impact --files` already does (~12 s). If neither is viable the command is repo-wide analysis and must stop being labelled an edit gate **✅ closed 5.2.0** — a clean tree builds no model at all: the verdict every axis would compute is the one an empty change set has by construction, and a model already cached for that exact tree state is taken because it costs a lookup. The verdict publishes what it read (`analysis.model_built`, the basis, and the claim it does not make — `pass` says the tree equals HEAD, not that HEAD was audited). BroadleafCommerce, clean detached worktree, every layer purged: **9,94 s → 0,62 s**. The dirty-tree half (scope the working side to the touched files) stays open and is now the whole of the row. **Eval #27 measures 68,3 s on a tree it reports as clean, and the fix is not what came undone** — the short-circuit is correct and reproduces; what defeats it is the *condition* it is asked, which admits every path git prints rather than every path an axis can read. That is **C3-112**, a distinct hole, and it is why B19 has now been reported four rounds running. |
270
270
  | C3-103 | **A cached core that is richer than the one a view needs cannot serve it, so the two root views of one analysis are two analyses.** The L1/L2 split exists precisely so *"different views of the same core share a common ancestry without a full re-analysis"*, and the overlay path already generalises it in one direction: with `--git-context`/`--env-map` the lookup tries candidate bases *with those overlays off*, injecting them afterwards, *"fewest flips first — so the RICHEST already-cached core wins"*. The reverse is never tried. `--agent` sets `dependencies/env_map/code_notes/architecture/graph_modules`, so its core key differs from the core just written by `--compact --git-context --env-map` in `gc` — a flag the cached core **has** and the request does not — and no candidate is generated for flipping an overlay *off*, giving a full L1 miss | 5.0.0 (eval #23) | **Low** — measured: `--compact --git-context --env-map` 72,3 s cold, then `--agent` **71,7 s with the RIS already built**, then `--agent --full` 0,8 s from the second run's snapshot. 144 s for two views of one repository state; a user who wants a human summary and an agent context pays the analysis twice | **closed (unreleased)** — the candidate set is symmetric now: a base whose overlays are a superset of the request is reused by *dropping* the blocks the request did not ask for, with its own view-key suffix (it shares the base's core hash and is a different answer). Measured on jobrunr with isolated caches: `--compact --git-context` then `--compact` goes 1,15 s → **0,32 s**, and the projected payload equals a cold run's block for block. ⚠**A quarter of the row's premise did not survive the source check:** the field's `--compact --git-context --env-map` → `--agent` sequence does not differ in `gc` alone — `--agent` also sets `graph_modules`, which the compact core lacks — so that request needs a block *added* as well as one dropped, and adding is analysis, not projection. That particular 71,7 s is still a full run. Original plan: the candidate set is symmetric: a base whose overlays are a **superset** of the request's is also reusable, because an overlay only attaches a block (the same additivity the inject path relies on). Projecting the richer core down is a drop, not a rescan **Measured again at 5.1.0 (eval #26 re-report), and the delta is now named**: on a clean detached worktree of BroadleafCommerce (2 985 files, every layer purged), `--compact --git-context --env-map` costs **14,1 s**, `--agent` after it **20,5 s**, and `--agent` with nothing before it **25,3 s** — so the shared parse buys the second view 4,8 s and the analysis is otherwise repeated whole. With the shared CIR and parse store warm and only the snapshot layer purged, the `--agent` run still spends **13,5 s** (detection 1,9 · analysis 6,3 · serialize 5,2). The *analysis* difference between the two cores is exactly one flag — `graph_modules`, which `--agent` sets for IC-003 and `--compact` does not — so the increment this row needs is augmenting a cached core with one analyzer rather than re-running the pass, and the acceptance test is byte-identity with a cold `--agent` run. Not attempted under a release deadline: a core that is *almost* the cold one is the C3-94 class of defect, and this row is a cost, not a falsehood. **Re-measured in the field at 5.2.0 (eval #27): 71,2 s (`--compact --git-context --env-map`, cold) + 64,0 s (`--agent`, after it) = 135 s for two views of one state, down from 347 s — and the third invocation, `--agent --full`, costs 0,7 s.** The row is unchanged by that: the drop is the general performance recovery, not sharing, and the 0,7 s run is the proof that caching works *at the view layer and not below it*. The field's own reading, kept because it is the acceptance test in one line: *"quien quiera resumen humano y contexto para agente paga el análisis dos veces"*. |
271
271
  | C3-104 | **`cache clear` cannot clear a layer `cache status` publishes.** `cache.stats()` reports the parse store — *"per-file parses, content-addressed across every repository"*, `~/.sourcecode/parse-cache-v1/` per `cache_model` — and `cache.clear()` deletes only the per-repository core/view/snapshot files; `parse_cache` exposes `purge_stale_generations()` and no eviction surface at all. There are seven `cache` subcommands and none of them empties it | 5.0.0 (eval #23) | **Medium** — the field could only purge for a from-scratch audit with `rm -rf ~/.sourcecode/parse-cache-v1` (**6 684 files / 142,5 MB**), a path they had to learn from our docs rather than from the CLI; the second evaluator counted the same layer as *"412 entradas irreversibles en caché global compartida"* under residue left outside the repository. Clearing cache between versions is what we recommend doing, and it is the one thing `cache clear` does not finish | **closed 5.1.0** — `parse_cache.clear_store()` empties the store (every generation, entries and bytes reported) and `ask cache clear --global` is the flag that calls it. Global on purpose and named for it: the store is keyed on content **across** repositories, so "this repository's entries" is not a set it can produce, and deleting it under `--all` would quietly take parses other repositories are using. The default run now *says* the layer is still there, with its size and the flag that empties it — a layer that can be evicted and nobody knows how is the same defect one step further away. Original plan: `cache clear --all` clears every layer it reports for that repository, including this repository's entries in the shared store (they are content-addressed, so the set is derivable from the file list), plus a `cache clear --global` for the store as a whole. A layer that is published and cannot be evicted is a fact with no authority over it |
@@ -286,7 +286,7 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
286
286
  | C3-119 | **A client call whose URL is an expression is promoted to `broken_integration` — the one verdict this command presents as unconditional.** `ask endpoints <repo> --consumer <repo>/saint-client` reports **20 calls resolved, 115 unresolved** with a single reason (*"the URL is an expression, not a literal at the call site"*), `routes consumed: 0 / 3 574`, `delete_candidates: 2 637` — and 16 of the unresolved arrive as `broken_integrations`. The consumer centralises its base URL in a service (``this.httpGet(`${this.END_POINT}/porAgrupacion/${id}/${fecha}`)``, `planificacion-especial.service.ts:109`) and only the literal tail is captured, so the join key is a fragment and the route is then declared unserved. F-BH's own framing is that this direction *"needs no assumption about coverage to be true"*; it is the one verdict presented as incondicional and the one that is wrong. **Second facet, same extractor:** the TypeScript scan does not strip comments — `"GET /to/your/validation/service"` at `seguro-grid.component.ts:204` is a commented-out DevExpress documentation placeholder admitted as a live call. That is the E-1/E-2 predicate, applied in Java and not in TypeScript | 5.3.1 (audit #28) | Medium-High — a team acting on `broken_integrations` chases 16 integrations that are not broken, and the `delete_candidates` list inherits the same error at 2 637 rows | **closed 5.4.0** — a call whose head was an interpolation travels as `path_resolution: prefix_stripped`; it still matches where it matches, and where it does not it leaves `broken_integrations` for `partial_client_calls` with the reason and a count, on both join surfaces. `broken_integrations`' own `meaning` now states the condition it holds under. The comment model is `blank_comments`, offsets and line count preserved so every call still points at a line a reader can open, and string/template literals are state. Original remedy note: an unresolvable URL degrades to `unresolved_client_call` and never to `broken_integration` (absence of evidence is not evidence — INV-F1-1); the comment model the Java side has runs on the TypeScript side too. The resolution half is **F-BO**, not this row |
287
287
  | C3-120 | **The composition engine became unusable at 3 342 files, and one line of C1-45's fix is why.** Measured against 5.3.1 artefacts on the same subject: `risk` 38,37 s → **601,74 s** under a 300 s budget, returning `partial: true` with **0 of 98** defects composed; `risk` under a 3 000 s budget → 444 s complete (×11,6); `enrich` 36,30 s → **>600 s**, stopped without completing; `audit-report` not run at all after two hangs. Median ratio over 48 comparable pairs is **1,00x** — the regression is confined to this family. Root cause is `sql_taint._receiver_type`: `_RECEIVER_RE` is anchored `…\s*$` and is applied as `search(method.body[:body_index])`, which copies an O(body_index) prefix **and** makes the regex engine try every start position in it before concluding. `_called_names` pays that for **every call of every method in the repository**, unfiltered and unbounded, where 5.3.1 did `names.add(call.group(1))` at O(1); a service body of 1 650+ lines makes the pass quadratic. Auditor's microbenchmark on the shipped regex, results byte-identical to a windowed `search(body, start, endpos)`: 8,6 KB/200 calls **36,9×**, 43 KB/1 000 calls **186,9×**, 215 KB/5 000 calls **916,5×** — 5× the calls, 25× the time | 5.4.0 (audit #29) | **Critical** — three of the seven experimental commands, and the ones that carry the price premium, do not return on the repository this product is sold for | **closed 5.4.1** — the read is windowed (`search(body, start, body_index)`: the `$` still ends at the call, no prefix is copied and the engine is offered a bounded number of start positions), and the window start retreats over any identifier it lands inside so a receiver wider than 128 characters is resolved whole rather than by its suffix — `""` is *unresolved* here and must never be reached by truncation. Second half: `_called_names` takes the statement ids being decided as its scope, because the only question asked of that set is `(mapper, statement id) in called`, so a name no statement carries can change no verdict. Re-measured on the shipped regex: 5 000 calls over a 159 KB body, **18 079 ms → 28 ms**, results identical at every call site. Four regression assertions, including a wall-clock gate on the pass |
288
288
  | C3-121 | **The budget does not bind in the composition family, and one invocation froze the process.** `ASK_MAX_ANALYSIS_SECONDS=300 ask risk .` ran **601,74 s** of wall clock and wrote `elapsed_seconds: 600.92` into its own `_partial`; `ASK_MAX_ANALYSIS_SECONDS=150 ask enrich . --sarif` passed 600 s and was stopped by hand; a third `risk` invocation stopped responding until the working session was killed. This is the failure mode C3-82/C3-118 closed for `generate-tests` — *the budget must be checked where the walk can be cut, not only between phases* — reappearing in another family, and the lesson did not travel. Second half, independent of the overrun: a **respected** budget yields no answer — at 300 s the payload is `phases_completed: ["audit"]`, `phases_pending: ["compose"]`, `composition_stopped_after.defects_composed: 0`, which in CI is an empty envelope after ten minutes. The degradation shape itself is exemplary and is not what this row is against: `_partial` names the completed and pending phases, `why_stopped`, the repo-wide walks it skipped and the axes left `unknown`. ⚠ Re-measure after C3-120 before investigating separately — same subsystem, same triage C3-114 got against E-3 | 5.4.0 (audit #29) | **High** — a budget that does not cut is worse than no budget: it sells CI a guarantee it does not keep, and a hard freeze consumes the runner | **closed 5.4.1** — the deadline is asked *inside* the two repository-wide walks that hold the wall time, not only at their doors: the method parse samples the clock every 16 files and the two propagation walks sample it on their own units. A walk cut while running raises rather than returning what it had — every verdict these walks publish below `reaches` is a statement about the **whole** repository, and a negative over a partial universe is a confident no built on an unfinished read — so the measurement is dropped, the axis reads `unknown`, and the walk is named in `_partial.composition_walks_not_run` exactly as a walk that never started is. Nothing partial is cached on the CIR, so a later caller cannot be served a repository with files missing. Two assertions, both non-vacuous: the walk cut inside, and the composer naming it. ⚠ The residue this row does not close: the walks that are still gated only at the door (`inferring the security posture`, `indexing the validation surface`, `resolving the conditional bean graph`) — none of them was measured holding the overrun, and giving each a sampling seam is its own increment |
289
- | C3-122 | **`ask risk <repo> -o <path>` spent 444 s, exited 0, printed nothing and wrote no file.** One occurrence, from bash with a Windows path in forward slashes; a later recursive search found the file nowhere in the user tree. The same path shapes worked in the same session for `ask endpoints -o`, `ask config` and `ask baseline capture --dir`. Not reproduced — the auditor stopped invoking `risk` after the freeze in C3-121 — and recorded with that limitation. An analysis command that consumes 444 s and reports success without producing its output owes the reader a line on stderr at minimum | 5.4.0 (audit #29) | **Low** (unreproduced) — but it is a C5 shape: the exit code is the published fact and it is wrong | **open** |
289
+ | C3-122 | **`ask risk <repo> -o <path>` spent 444 s, exited 0, printed nothing and wrote no file.** One occurrence, from bash with a Windows path in forward slashes; a later recursive search found the file nowhere in the user tree. The same path shapes worked in the same session for `ask endpoints -o`, `ask config` and `ask baseline capture --dir`. Not reproduced — the auditor stopped invoking `risk` after the freeze in C3-121 — and recorded with that limitation. An analysis command that consumes 444 s and reports success without producing its output owes the reader a line on stderr at minimum | 5.4.0 (audit #29) | **Low** (unreproduced) — but it is a C5 shape: the exit code is the published fact and it is wrong | **closed 5.8.7** — fixed as the shape, not as the occurrence, because the occurrence is one report and the shape is checkable on every run. **A run that promised a file and produced none can no longer exit 0**: the requested path is noted from the same vector the parser is about to read (so the guard covers all 65 commands that publish `--output`, including one that returns before reaching its emit seam — which is what the report describes), and the promise is discharged at `_progress_output_emitted`, the seam every answer already passes through, the three bespoke file writers included. At the end of a run that exited 0 with the promise still standing, the `EXECUTION_FAILED` envelope names the path and the run exits 1. The guard is conservative on both sides so it can only fire on the measured state: a discard sink was never a promise of a file, a file that exists clears the promise however it got there, and a run that already failed is not second-guessed. Second half, the row's own minimum (*"owes the reader a line on stderr at minimum"*): a file write is never silent — where the command has no sentence of its own, the byte count and the path go out as a notice, which keeps a pipe and a PowerShell capture byte-identical (C3-26). |
290
290
  | C3-123 | **A static asset reaches `broken_integrations`, the one verdict this command presents as unconditional.** C3-119 did the main work — 19 → 1, with 14 relocated to `partial_client_calls` and the `meaning` stating its condition. The survivor is `GET /assets/i18n/{}.json` from `transloco-root.module.ts:18`: a translation bundle served by the container, not a Spring route. True under the published definition and not a broken integration in any sense a reader means. Remedy: exclude paths under known asset roots (`/assets/`, `/static/`, `/public/`) or ending in a static file extension, and say so in the `meaning` | 5.4.0 (audit #29) | **Low** | **closed 5.4.1** — static asset roots and file extensions are excluded and the meaning names the exclusion |
291
291
  | C3-124 | **`risk` renamed three published keys inside a minor with no alias.** 5.3.1 emitted `total_defects`, `total_findings` and `risks_cap`; 5.4.0 emits `total_defects_floor`, `total_findings_floor` (plus `counts_basis`, `counts_are_floor`, `partial`, `_partial`) and dropped `risks_cap`. Permitted by the command's own experimental tier — *shape may change in a minor* — and the new names are better; the payload still carries no notice, so a consumer bound to the old keys reads `None` with nothing telling it why. `auth status` already publishes a `deprecated_fields` block: the channel exists and this command did not use it | 5.4.0 (audit #29) | **Informative** | **closed 5.4.1 as not a defect** — a complete 5.4.1 run retains the legacy keys; the renamed floor fields are emitted only for budget-cut payloads and are explicitly marked as floors |
292
292
  | C3-125 | **`ask <repo> --compact` writes its complete, valid answer to stdout and then does not exit.** On tutorials (24 074 java files, hundreds of Maven modules — the largest repository this battery has run, first time past ~9k files), the documented `OUTPUT_TOO_LARGE` fallback (`ask . --compact`) finishes writing full valid JSON by roughly 100-116 s wall clock, but the process stays alive — polling confirmed stdout stable at t=103 s while the process remained running and consuming CPU past 180 s, when it was killed by hand. Reproduced twice independently in the same session. A caller that waits on process exit rather than reading stdout and moving on (a CI step, a subprocess wrapper with a completion callback) hangs indefinitely on a repository at this scale; suspected an unshutdown worker pool (`parallel.py`/`ast_extractor.py`), not confirmed by reading source this session | 5.7.0 (golden-repo battery #2 2026-08-15) | High — a fast path completing correctly and then hanging is worse for automation than a slow path that returns, because nothing downstream distinguishes "still working" from "done, stuck" | **closed 5.7.1** — reproduced, and the suspicion this row carried is **refuted**: nothing leaks and the process does exit. A cold `ask tutorials --compact` (24 074 java files) ran **305,26 s** wall / 276,74 s user and exited 0 — the field's kill at 180 s landed inside a run that had 125 s still to go. What looked like a hang is the L1 cache write: `core_view` renders compact, agent **and** standard again, after `_emit_command_output` has already written the answer, and it did so with nothing on stderr. Sampled stacks put **57 %** of the process inside one function — `_spring_profiles_context`, a pure derivation of the snapshot, derived five times at ~50 s each. Two costs inside it: per (profile, file) pair it built three `pathlib.Path` objects (40 profiles × 35 142 paths × 3 = **4,2 million** constructions against a 24 000-file live heap), and `agent_view`'s `_spring_event_signal` re-parsed the whole repository into a private CIR. Three changes: the stems are derived once from the file list (same matching rule, same order), the context is memoized per snapshot and hands back a copy, and the event signal goes through `context_cache.shared_cir` — the one door C3-94 established. Measured after on the same repository, same 17 426-byte answer (byte-identical but for run id, timings and the analyzer fingerprint): **102,31 s** wall / 78,17 s user — **2,98× wall, 3,54× CPU**; the answer's own `total_ms` fell 97 035 → 45 816. The tail is also no longer silent: the post-answer build runs under a spinner reading *building cache entry (the answer above is complete)*, so "still working" and "done, stuck" are distinguishable from outside. Regression: `tests/test_post_answer_recompute_c3_125.py`, 7 assertions |
@@ -402,12 +402,12 @@ class's subject, not its provenance.
402
402
  | E-6 | **A module-name heuristic meant for `test-framework`-shaped modules excludes a real, deployed Spring Boot surface, and the top-level fact is published with no gap flag.** `repository_ir.py:6363-6382` (`is_test_source_file`): a file under `.../src/main/...` is correctly not a test by the file-level authority (`test_sources.is_test_path`), but the function separately excludes it when the enclosing module's path prefix has any segment `p.lower().startswith("test")` (line 6375) — written for modules like `test-framework`/`test-providers`/`testsuite`, which the docstring says "have no product of their own." Apache SkyWalking's own e2e harness is rooted at a top-level directory literally named `test/` (`test/e2e-v2/java-test-service/*/src/main/java/...`), which matches the rule on the first path segment even though the files under it are real `@RestController`/`@GetMapping` classes (`HealthController`, `AlarmController`, `LogController`) — not test infrastructure with no product. Effect: `ask endpoints` on the full repository returns `total: 0`, `spring_mvc_annotations: 0`; `ask spring-audit` returns `spring_detected: false`, for a repository with confirmed (narrow, 23-file) real Spring usage. `--include-tests` recovers all 20 routes, and running `ask endpoints` directly on the submodule finds them without any flag — proving detection works once the module-exclusion is bypassed. `coverage.gaps` stays empty on the default run; the exclusion note (`test_source_excluded_detail`, "29 route annotation(s) from test source trees excluded") reads as routine test filtering, not as "an entire production-shaped module was treated as test infrastructure" | 5.7.0 (golden-repo battery #2 2026-08-15) | Medium — a judgment-call heuristic operating exactly as designed for its intended target, overreaching on a repository whose harness happens to be named `test/`, but landing on the same headline field (`spring_detected`) E-5 does, with no hedge published | **closed 5.7.2 — the disclosure half, and the admission left standing on purpose.** Whether an e2e harness rooted at `test/` is product or infrastructure is a judgement, and overturning it on one repository would move surface on twenty-three; what was indefensible was publishing it as routine test filtering. `test_source_excluded_detail.excluded_by_module_name_rule` now names the modules, counts the route annotations, states that the rule judges the **module** rather than the file, and gives both ways back (`--include-tests`, or run the command on the module path); `coverage.gaps` carries the entry, because that list is what a reader checks before trusting a zero. Only the module-name branch is reported — `is_test_path` never fires inside a declared main source root, so it cannot remove a deployable module. Measured: skywalking names 3 modules and 29 route annotations where it named none; spring-petclinic (17) and jobrunr (12) gain nothing. ⚠ Residue, deliberate: `spring-audit`'s `spring_detected: false` still carries no hedge on this repository — the exclusion fact does not reach that surface, and wiring it there is its own row |
403
403
  | E-7 | **A security-policy classifier regex matches a method's own declaration text, not a call site, and the resulting false flag is cached per-file and leaks onto a handler explicitly marked bypass-all.** On sa-token (a non-Spring-Security Java auth framework with its own `@SaCheckLogin`/`@SaCheckPermission`/`@SaIgnore` annotations), `ask endpoints` labels `AtCheckController`'s handlers `policy: "programmatic"` although no `sourcecode.config.json` custom-security declaration exists (`ask config` confirms "Declaration: none"). The only evidence found is a regex matching the literal text `checkPermission(` in the file — which matches the demo's own method **declaration**, `public SaResult checkPermission()`, not a call to a security check. Reported at `repository_ir.py:394-408`/`:5064-5076` (field-agent-located; not independently re-read this session, one level less verified than E-5/E-6/C1-47/C1-48). The flag, once set, is applied to every handler in the file lacking its own recognized policy — reaching `ignore()`, a handler annotated `@SaIgnore` (sa-token's explicit "skip all checks" marker), which is reported as `programmatic`-protected, the opposite of what the annotation on it says. Contamination reproduces in `impact-chain`'s `endpoints_affected` for the same controller | 5.7.0 (golden-repo battery #2 2026-08-15) | **High** — this is an auth-verdict false positive on the exact axis `ask risk`'s own documentation says its composition depends on (`severity_effective = ... × auth_verdict × ...`); a wrongly-`programmatic` handler would suppress severity on a genuinely open endpoint | **closed 5.7.1** — **re-verified at source first, as this row required**, and the field report was exact. `_PROGRAMMATIC_SECURITY_RE` matched `\b(?:…\|checkPermission\|…)\s*\(`, which is the shape of a *call* and equally the shape of the *method that declares one*. On `AtCheckController` the regex finds **one** match in the entire file — line 46, `public SaResult checkPermission() {` — and on that evidence all **seven** handlers were published `policy: programmatic`, `ignore()` among them, which carries `@SaIgnore`, sa-token's explicit skip-all-checks marker: the verdict stated the opposite of the annotation on the method. Repo-wide the false population was **23 endpoints (programmatic 41 → 18)**, and every one of the 15 surviving files was re-read and carries a real call site (`StpUtil.checkPermission("")`, `stpLogic.hasRole(role)`, `SaRouter.match(…)`). The discriminator is what precedes the name on its own line: a call is led by a receiver, an operator, a bracket, a statement boundary or one of `return`/`new`/`throw`/…; a declaration is led by its return type — an identifier, a closing generic, or a closing array bracket. ⚠ **The first attempt introduced a false negative and the cross-repository A/B is what caught it**, not the suite: applied to the whole pattern the filter also ran on the receiver alternatives, and `ReactiveSecurityContextHolder.getContext()` reads as identifier-then-match — exactly what a declaration looks like — so spring-security-samples' reactive `MeController` silently lost a correct verdict (that repository ships **two** `example.MeController` classes, which is how it surfaced). The pattern is now split: the ambiguous `name(` half is filtered, the unambiguous receiver/member/`throw` half never is, and that negative control is the sixth assertion in the battery. **Measured across all 23 golden repositories: 22 of 23 policy censuses byte-identical**, sa-token the only mover — BroadleafCommerce (12), keycloak (147) and spring-security-samples (3) unchanged. 6 assertions, **3 red on the previous build**, the 3 green ones being the negative controls. Suite 8 761. ⚠ Residue this row does not close: sa-token's handlers now read `none_detected`, which is honest but incomplete — `@SaCheckPermission` is a vendor annotation this analyzer models only when declared in `sourcecode.config.json`, and that is the documented custom-security path, not a new gap |
404
404
  | E-8 | **`impact-chain` silently merges unrelated same-named classes from different modules into one answer; `impact` on the identical symbol correctly refuses.** tutorials (Baeldung's mega-repo, hundreds of independent modules) has 4 distinct classes named `PersonService` in unrelated modules. `ask impact PersonService <repo>` returns `resolution: "ambiguous"` and lists all 4 `matched_fqns`. `ask impact-chain PersonService <repo>` returns `resolution: "partial"`, but its `direct_callers` mix classes from two different, unrelated modules (`com.baeldung.activej.*` and `com.baeldung.hibernatejfr.*`) into one merged answer and a single risk score (25.0, `critical`), with no field disclosing which of the 4 classes contributed which caller. The command that should have said "ambiguous" said "partial" and guessed | 5.7.0 (golden-repo battery #2 2026-08-15) | Medium-High — only demonstrated on a repository large and diverse enough to carry real name collisions across modules (tutorials, 24 074 java files — the largest repository this battery has run against); a confident composite risk score built from an admittedly-ambiguous symbol resolution is the opposite of `impact`'s own behavior on the same input | **closed 5.7.2** — on C1-41's precedent rather than a second one of its own. The facts stay (every caller listed is a real caller of *some* candidate, and dropping them would hide reach the reader asked about); the figure that reads as *this symbol's risk* is void — `risk_score` and `risk_score_raw` null, `risk_level` `unknown` — with the reason published beside the null in `risk_score_basis`. The candidates are named in `metadata.matched_classes` / `matched_classes_count`, and both a warning and the explanation say so, because an agent quotes the explanation and a script reads the metadata. `impact-chain` is `core` tier, so no published enum value is invented: `resolution` stays `partial` and `risk_level` uses `unknown`, which its schema already documents, and every new key appears only on an ambiguous match — an unambiguous answer is byte-identical. Measured on tutorials: 4 candidates named where the payload named none, `high`/25.0 replaced by `unknown`/null, the six direct callers still listed. Regression: `tests/test_ambiguous_chain_e8.py`, 8 assertions |
405
- | E-9 | **An unresolved JAX-RS path's caveat is dropped between `endpoints` and `impact-chain`, and the path's embedded regex syntax collides with the `endpoint_id` field separator.** killbill's `AccountResource` declares `@Path("/{accountId:" + UUID_PATTERN + "}")` (string-concatenated). `ask endpoints` correctly marks it `path_resolution: "unresolved"`. `ask impact-chain AccountResource <repo>` shows the same route without that caveat and truncates the path to `"/{accountId:"`, and because the embedded `:` collides with the `endpoint_id` separator the emitted id is malformed: `DELETE:/{accountId::org.killbill.billing.jaxrs.resources.AccountResource:...#closeAccount` | 5.7.0 (golden-repo battery #2 2026-08-15) | Low | open |
405
+ | E-9 | **An unresolved JAX-RS path's caveat is dropped between `endpoints` and `impact-chain`, and the path's embedded regex syntax collides with the `endpoint_id` field separator.** killbill's `AccountResource` declares `@Path("/{accountId:" + UUID_PATTERN + "}")` (string-concatenated). `ask endpoints` correctly marks it `path_resolution: "unresolved"`. `ask impact-chain AccountResource <repo>` shows the same route without that caveat and truncates the path to `"/{accountId:"`, and because the embedded `:` collides with the `endpoint_id` separator the emitted id is malformed: `DELETE:/{accountId::org.killbill.billing.jaxrs.resources.AccountResource:...#closeAccount` | 5.7.0 (golden-repo battery #2 2026-08-15) | Low | **closed 5.8.7** — both halves, and the first one was not where the report put it. The caveat was not *dropped* between the two commands: on this shape it was never produced. `_route_path_expression` returned "" for any argument list containing a string literal, and `"/{accountId:" + UUID_PATTERN + "}"` contains one — so the route was published as fully known and the first literal (`/{accountId:`) as the path, silently. **A partial resolution is not a resolution**: the presence of one literal says nothing about the operands beside it, and a concatenation whose non-literal operand does not fold against the repository's constants is now `path_resolution: unresolved` with its expression. The caveat then travels rather than being re-derived: it is a field of `CanonicalEndpoint` (not of one projection), so `endpoints`, `impact-chain` (`AffectedEndpoint`) and `impact` (`endpoints_affected`) publish the same two keys from the same authority. Second half: `make_id` percent-escapes the separator in all four components, so an endpoint id has four fields for every route and never five — `parse_id` is the published inverse and refuses anything else rather than returning a shorter tuple a caller would read as data. Measured on the reported shape: `PUT /1.0/kb/other/{id:` → `unresolved` with `path_expression`, id `PUT:/1.0/kb/other/{id%3A:…`, and the resolved route beside it unchanged and un-annotated. |
406
406
  | E-10 | **Malformed Java is silently absorbed into the symbol export with full confidence, and the parse failure is not named anywhere in the payload.** spaghetti-api's `BrokenSyntax.java` has 2 open braces and 0 close braces — invalid Java, confirmed by brace count and by reading the file. The root command's `contracts` export nonetheless reports it with a clean `class BrokenSyntax` export plus a `method` export (signature `()->void`), no `parse_error`, absent from `analysis_gaps` — and it never surfaces in `spring-audit`, `migrate-check` or `risk` output either, no "N files failed to parse" anywhere in any of them | 5.7.0 (golden-repo battery #1 2026-08-14) | Medium — a file that cannot be Java is exported as if it parsed cleanly, on the exact axis this repository's own CLAUDE.md is most explicit about ("never a confident falsehood... unknown, never 0") | **closed 5.7.3** — the recovery stays (a file that fails to compile still tells us its type exists); what changes is that it is qualified. A `parse_structure` gap names every file whose braces do not balance **once comments and strings are masked**, and states what is missing from the model: the members after the unclosed block and the type's own end. Distinct from `parse_coverage`, which fires on *total* extraction absence — this one fires when extraction succeeded over source that cannot be complete, and both readings of "this file was not fully read" now live in `reconciliation`. Cost is a `str.count` per file; the mask, which is not cheap, runs only for the rare file that already looks unbalanced, so a brace inside a string or a comment is not a false positive (asserted). Measured on spaghetti-api: `BrokenSyntax.java` named, its symbols retained. ⚠ Residue: the gap reaches every CIR consumer through `analysis_gaps`, but the root command builds its own list from `ConfidenceAnalyzer`, so `ask <repo> --compact` still does not carry it. Regression: `tests/test_parse_structure_gap_e10.py`, 7 assertions |
407
407
  | E-11 | **`impact`'s DI-interface caller resolution fabricates direct callers that never reference the target, inflating the risk band; `impact-chain` on the identical symbol does not.** On examples, `ask impact OrdersService` reports `critical` (score 35.5, 4 direct + 35 indirect callers); `ask impact-chain OrdersService` reports `medium` (score 11.0, 1 direct caller) for the same repository state. The extra "direct callers" `impact` adds — `MicroserviceUtils` and its two nested classes — never reference `OrdersService` anywhere (grepped, 0 hits); they only import the shared `Service` interface `OrdersService` implements. `impact`'s own explanation names the mechanism: `"callers resolved via interface (Service) — Spring/CDI/Guice DI pattern"` — the resolver treats any class touching a common interface as a caller of every implementer of it. This compounds C1-48 (the two commands already score from different formulas) with a caller **set** that disagrees on its own, independent of which formula is applied to it | 5.7.0 (golden-repo battery #1 2026-08-14) | Medium-High — same family as C1-21 (a fabricated identity reported as protection): here a fabricated call edge is reported as reach, on the exact factor `ask risk`'s reachability axis depends on | **closed 5.7.3** — the DI recovery is right and stays: a container binds implementations to interfaces, and a walk that stops at the interface loses every real production dependent (C1-16). What was wrong is the name it was filed under. A class that reaches the target only if the container binds *this* implementation sits at the same conditional distance as a transitive caller, not at the distance of a class that names the target. `interface_mediated_callers` / `interface_mediated_caller_count` publish the population, `via_interface_note` says what it is and is not, the explanation carries both figures, and they are scored on the **indirect** axis (0,5) rather than the direct one (2,0). Reach is untouched — they remain BFS seeds and members of the blast cone, so endpoints, transactions, mappers and cross-module reach are unchanged. Measured across 8 symbols on 7 repositories: **7 byte-identical**, including BroadleafCommerce `Money` (134 direct, 1 584 indirect) and keycloak `UserResource`; `examples/OrdersService` moves 4 → 1 direct, 35 → 38 indirect, 35,5 → 31,0, `endpoints_affected` unchanged at 3. Regression: `tests/test_interface_mediated_callers_e11.py`, 7 assertions |
408
408
  | E-12 | **Spring bean detection matches an annotation's *simple name* with no check that it is Spring's, so a repository that declares its own `@Service` is published `spring_detected: true`.** `spring_model.py:41-46` (`_BEAN_ANNOTATIONS`) holds bare spellings — `@Component`, `@Service`, `@Repository`, `@Controller`, `@RestController`, `@Configuration`, `@Bean` — and `BeanGraph.build` (`:184-202`) admits a node whose `annotations` list intersects that set. The list carries simple names only, so the match is on the token, not on what it resolves to. `spring_detected` (`spring_model.py:455`, the F-AY single authority) is then `has_spring_beans() or tx_total > 0`, and the bean half decides alone. **Measured on neo4j**: `grep -r org.springframework` over the whole tree returns **0** occurrences and the graph holds **0** edges mentioning it, while `import org.neo4j.annotations.service.Service` appears **199** times — neo4j's own SPI marker, declared at `annotations/src/main/java/org/neo4j/annotations/service/Service.java`. `BeanGraph.build` over the CLI's own CIR (5 598 files) returns **34 beans, every one stereotype `service`**, all of them that annotation; `tx_index.stats()["total"]` is **0**. So `ask spring-audit neo4j` publishes `spring_detected: true` in the same payload as `tx_stats: {total: 0}`, under both `--scope security` and `--scope tx`. The discriminating evidence is **already in the graph and already walked**: pass 1 of the same function iterates annotation-type nodes to build the meta-annotation map, and one of those nodes is `org.neo4j.annotations.service.Service` — a repository that declares the annotation itself is the witness that the token is not Spring's. Negative control: killbill (0 `org.springframework`, 0 bare bean annotations) correctly publishes `spring_detected: false`. **Not observed**: no false finding fires on neo4j today (`total_findings: 0`), so this row is against the published fact and what reads it, not against a finding — but `has_spring_beans()` is also the AOP-premise witness (`spring_security_audit.py:638-642`), the one C2-era fix installed so a self-invocation finding cannot fire where nothing proxies, and on this repository that premise is satisfied by a non-Spring annotation. **Propagates to every surface that reads the authority**: `spring-audit` and `spring-audit --scope tx` (`spring_tx_analyzer.py:1039`), `posture` (`posture.py:1207`), `risk` (`risk.py:1593`), `ris.py:355`, and the MCP orchestrator's `repo_type: "java_spring"` (`mcp/orchestrator.py:360`). ⚠ Unquantified exposure: sa-token (134) and dubbo (164) also carry bare `@Service`/`@Component` spellings — dubbo does use Spring, so its verdict may be right for the wrong reason; neither was measured this round | 5.7.0 (split out of E-5 while closing it, 2026-08-15) | **High** — this is the vendor-agnostic rule this repository enforces, inverted: the logic branches on a proprietary name and treats it as the rule rather than as evidence, so the blast radius is every repository that spells an annotation the way Spring does, and the field it decides is the one that says whether the Spring axes are measuring this repository at all | **closed 5.7.1** — the annotation is **resolved**, not merely spelled. `BeanGraph.build` walks the `imports` edges the graph already carries (109 475 of them on neo4j, indexed by the same class FQN the bean nodes use) and asks where the spelling binds, in the order the compiler would: an explicit single-type import **decides it either way**; a wildcard import of a Spring package leaves Spring reachable, so the spelling stays admissible; failing that, an annotation this repository declares in the owner's own package is the owner's own; and **nothing resolved stays `None`** — an absent import edge is not evidence, so the node keeps exactly the behaviour it had rather than acquiring a verdict from silence. The node is not deleted: `BeanNode.spring` records the verdict and `annotation_fqn` records what it bound to, so `get_stereotype` still answers for `explain` and only the **Spring** claim is withdrawn — `has_spring_beans()` is what filters. A meta-annotation is decided by what it *carries*, not by its own name: `@DomainService` is repository-declared by definition, so resolving its own spelling would have rejected every meta-bean the codebase deliberately supports — the question asked is where the `@Service` **on it** binds. The two annotation passes were split so the same-package map is complete before any resolution reads it, and the import walk shares the edge traversal that was already being made for the injection edges. **Measured**: neo4j 34 beans → **0 Spring** (all 34 bind to `org.neo4j.annotations.service.Service`), `spring_detected` **true → false**, and the `stack_fit` block — the honesty affordance that tells the reader to read the Spring axes as `unknown` rather than `none` — is **now published on that repository, where it was suppressed before** (verified by stashing the fix and re-running). Two more repos were silently wrong and are now right: sa-token had **35** beans bound to Solon (`org.noear.solon.annotation.*`) and loveqq, jobrunr **1** bound to Micronaut (`io.micronaut.http.annotation.Controller`) — both repositories still read `spring_detected: true` on their genuine Spring beans, so the correction shows up in the bean census rather than the headline. **Negative controls, unchanged**: spring-petclinic (12 Spring beans, 9 correctly withheld as `jakarta.persistence.*` entities the stereotype rule already excluded), dubbo 133/133, mall 158/158, spring-boot-admin 199/199, sagan 76 with 12 JPA withheld, killbill and eureka still `false`. 9 assertions, **7 red on the previous build**. Suite 8 360, the same 21 previous reds. **Re-measured on the full golden set before the bump**: in the 23-repository A/B battery neo4j's `spring_detected` is the only headline that moved, and the two independent authorities now **agree on all five repositories checked** — `migrate-check`'s `spring_present` and `spring-audit`'s `spring_detected` read false/false on neo4j, killbill and eureka, true/true on examples and spring-petclinic; before this fix neo4j was the one disagreement. Cost: `BeanGraph.build` goes from 7,2 → 17,7 ms on skywalking, 28,2 → 58,8 ms on keycloak and 41,7 → 82,4 ms on neo4j (284 338 edges), best-of-3 — the import walk roughly doubles a step that is tens of milliseconds inside a multi-second command. The three apparent `spring-audit` slowdowns in the battery log (skywalking +259 %, open-banking-gateway +200 %, mall +162 %) were **contention noise, not this fix**: re-timed in isolation they are 2,5 s / 1,2 s / 1,1 s against baselines of 2,7 s / 1,3 s / 1,3 s. ⚠ Residue this row does not close: `@Bean`/`@Configuration` on a class with no import edge for the name still falls to the unresolved branch (137 of spring-boot-admin's 199, 87 of dubbo's 133) — admitted, which is the direction that preserves behaviour but is not a resolution; and `next(iter(match))` still picks arbitrarily when a node carries two bean annotations |
409
409
  | E-13 | **No field in any repository carries a `contained_in` edge, so `ask verify` — the `core`-tier CI gate — returns `pass` on a repository that violates its own declared contract.** `repository_ir.py:3041-3052` emits `contained_in` for `sym.type in ("method", "field")`, and the owner is computed by `_enclosing_class` (`:3833`), which splits on `#` only. A method is spelled `Type#member` and resolves; a **field is spelled `Type.member`** and is handed back unchanged, so the `enclosing != sym.symbol` guard on the next line drops it. Measured on four repositories: **0 of 14** fields on spring-petclinic, **0 of 821** on mall, **0 of 1 086** on neo4j, **0 of 2 188** on keycloak. The consequence is a silent gate. `verify_rules.ForbiddenEdgeRule._endpoint_matches` is correct by design — it matches a selector against the edge endpoint **or the type that declares it**, resolved through `_owner_type_map`, precisely because *"edges are recorded at member granularity but client invariants are stated at type granularity"* — and it starves: for the shape `@RestController class Ctrl { @Autowired OrderDaoJpa dao; }` the graph holds `injects com.example.Ctrl.dao → com.example.OrderDaoJpa` and the `@RestController` lives on `Ctrl`, which the matcher can never reach. On mall, **176 of 181** `injects` edges originate at a field and are orphaned from their class this way. **How it survived is the second half of this row**: a test named `test_field_has_contained_in_edge` existed for exactly this property and asserted nothing — its filter (`"." in e["from"].split(".")[-1]`) cannot match, because the last dot-segment of an FQN never contains a dot, and `assert len(field_edges) >= 0` is true of every list. It passed green for as long as the defect existed | 5.7.0 (triage of the 21 standing red tests, 2026-08-15) | **Critical** — a gate that cannot fail is worse than no gate: a team switches CI on, sees green, and concludes the contract holds. `ask verify` is `core` tier, which this repository defines as *"safe to gate CI on"* | **closed 5.7.1** — one condition at the emission site. The dot is ambiguous with a package segment, so the last segment is stripped **only when what remains is a type this repository declares** — membership in `_local_classes` is the evidence, never the spelling, and `_enclosing_class` itself is untouched because its `#` behaviour is right for every other caller. **11 of the 21 standing red tests went green with this one edit** (`test_verify_repo.py` ×9, `test_contracts_file.py`, `test_verify_edit_v5.py::TestForbiddenEdge`) — they were never stale tests, they were an unread bug report. Collateral measured and nil: the 23-repository battery is **69 of 69 pairs byte-identical** across `endpoints`/`spring-audit`/`posture`, and `impact` — the one consumer that walks `contained_in` — returns identical band, score, caller and endpoint counts on spring-petclinic (×2), mall and sagan. Two assertions replace the vacuous one: the witness edge, and the general property that **no** member symbol is left without an owner. Suite 8 756 (21 → 10 reds). ⚠ Residue, deliberately not fixed here and filed as its own row: `BeanGraph.injections` is empty on all four repositories measured, because it keys on `frm in beans` where beans are class FQNs and the edges originate at fields and constructors. It has **no consumer in the codebase** — fixing dead structure to move a score is not a fix |
410
- | E-14 | **A 3-level Python re-export chain resolves through a documented 2-level limit, and nothing is reported.** `test_reexport_chain_limit`'s own docstring states the contract — *"El tercer nivel NO debe resolverse (chain limit = 2). limitations debe contener algun indicador del limite alcanzado"*. Measured on its fixture (`a/__init__.py` → `b/__init__.py` → `c/module.py`, consumed from the root): the `SymbolLink` for `deep_func` resolves to **`a/b/c/module.py`** with `is_external: false`, and `summary.limitations` carries only `namespace_package:a/b/c` — no chain indicator. Either the limit is not enforced or it is enforced and unreported; both are the shape this ledger calls a confident answer past the edge of what was measured. **Found the same way as E-13 and by the same defect in the suite**: the test ended in `assert ... or True`, which accepted every outcome including this one | 5.7.0 (vacuous-assertion sweep, 2026-08-15) | Low-Medium — Python import resolution is not the Java/Spring core this product is bought for, and no headline field is built on it; the row exists because the property is declared in the code and is not held | open — **tracked, not forgotten**: `test_reexport_chain_beyond_the_limit_is_reported` asserts the intended property under `xfail(strict=True)`, so the day it is fixed the suite turns red and forces this row closed rather than letting the fix land unnoticed |
410
+ | E-14 | **A 3-level Python re-export chain resolves through a documented 2-level limit, and nothing is reported.** `test_reexport_chain_limit`'s own docstring states the contract — *"El tercer nivel NO debe resolverse (chain limit = 2). limitations debe contener algun indicador del limite alcanzado"*. Measured on its fixture (`a/__init__.py` → `b/__init__.py` → `c/module.py`, consumed from the root): the `SymbolLink` for `deep_func` resolves to **`a/b/c/module.py`** with `is_external: false`, and `summary.limitations` carries only `namespace_package:a/b/c` — no chain indicator. Either the limit is not enforced or it is enforced and unreported; both are the shape this ledger calls a confident answer past the edge of what was measured. **Found the same way as E-13 and by the same defect in the suite**: the test ended in `assert ... or True`, which accepted every outcome including this one | 5.7.0 (vacuous-assertion sweep, 2026-08-15) | Low-Medium — Python import resolution is not the Java/Spring core this product is bought for, and no headline field is built on it; the row exists because the property is declared in the code and is not held | **closed 5.8.7** — the row was *"either the limit is not enforced or it is enforced and unreported"*, and it was neither: **two authorities disagreed about the same number**. The test's docstring said a three-level chain must not resolve; the walk's own comment said *"one level of chaining (depth=2 total)"* and resolved it. The walk wins, and the reason is measured: two hops land on `a/b/c/module.py`, which is where `deep_func` is actually declared — refusing a correct answer buys nothing, and the `xfail` encoded the losing reading. What was genuinely missing is the other side, plus two defects found while adjudicating it. (1) **A chain still going when the budget ran out was published as if it had ended** — the package `__init__.py` became the symbol's source and `limitations` said nothing; a four-level fixture now resolves to `a/b/c/__init__.py` (the last package that exports the name, never the module the walk did not reach) and declares `reexport_chain_limit:a.deep_func (resolved 2 of a chain longer than the 2-hop limit …)`. (2) **A chain that simply ends inside a package reported a limit it never reached** — `from .b import f` where `b/__init__.py` declares `f` emitted `reexport_chain_limit:f` on every such symbol, a false caveat, which is the same defect class pointing the other way. (3) **The walk read the map it was mutating**, so whether a chain resolved one hop or two was decided by `source_files` order; it now walks the first pass's own entries. `_REEXPORT_MAX_HOPS` is the one authority for the budget, and the five assertions that replace the `xfail` fail on the parent commit. |
411
411
  | E-15 | **No repository has ever published a `field_type`: the class-scope type-reference surface keys its output by the field's own FQN, so the first context named in its own acceptance claim is empty everywhere.** Same root as E-13, second consumer, found by triaging the reds E-13 left standing. `repository_ir.py:2381` (`_build_class_type_refs`) does `cls = _enclosing_class(s.symbol)` for a field, and that resolver split on `#` only — so a field indexed under `com.x.Svc.repo` instead of `com.x.Svc`, and `class_type_surface` querying by class FQN found nothing. Measured on the surface's own fixture (`@Autowired private OrderRepository repo;` plus `private java.util.Map<CustomerId, Order> ledger;`): the published contexts are `param_type`, `ctor_param_type`, `return_type` — and **`field_type` and `field_type_arg` are absent entirely**. On real repositories, spring-petclinic published `{ctor_param_type: 3, param_type: 24, return_type: 18}` and sagan `{ctor_param_type: 59, param_type: 103, return_type: 95}`, both with **zero** field contexts. The module's docstring states the claim this falsifies — *"the class's declaration structure — field, field generic arg, constructor param, method param, method return — is recovered as a pure ContextGraph query"* — and the field is named first | 5.7.0 (triage of the reds left standing by E-13, 2026-08-15) | **High** — a shipped acceptance claim for a Semantic IR capability, false in two of its five named contexts, on the two contexts that describe what a class *holds* rather than what it passes | **closed 5.7.1** — E-13's patch was at one emission site; a second consumer with the same root made that the wrong shape, so the resolution is now **one authority**, `_owner_type_of(symbol, symbol_kind)`, and both sites read it. The kind is the evidence: a field symbol is built as `f"{class_fqn}.{fname}"` at the single site that makes one, so the last dot-segment is the member and nothing is inferred from spelling; a caller that does not know the kind gets the `#` behaviour alone, because a bare dotted FQN is ambiguous with a package segment and this function may not guess. `_enclosing_class` survives as a kind-less delegate so its ten existing method-side callers are untouched. Measured after: spring-petclinic gains `field_type: 6, field_type_arg: 3` and sagan `field_type: 22, field_type_arg: 10`, with the other three contexts **byte-identical** — recovered, not double-counted. Collateral nil on the 23-repository battery (**69 of 69 pairs identical**) and on `impact` (4 symbols, 3 repositories, identical band/score/callers/endpoints). **21 standing red tests are now 7**, and 14 of the 21 were this one root |
412
412
  | E-16 | **A second, hand-written test-source filter stands beside the authority and removes production code from the model — the exact substring shape C1-2 closed in 3.2.1.** `find_java_files` asks `is_test_source_file`, which knows that a Java package named `test` under a *declared main source root* is production. Three command entry points then asked again as `[f for f in find_java_files(root) if "/test/" not in f and "/tests/" not in f]` (`repo-ir`, `impact`, `export`), and three more populations did the same by hand (`serializer._mybatis_pairing`, `serializer._bootstrap_structured`, `detectors/java._collect_transactional_classes`). Measured over the 24 golden repositories the second filter removed production sources from four of them — **sa-token 82 of 919 files (8,9 %)**, ofbiz-framework 41 of 1 075 (3,8 %), BroadleafCommerce 4, tutorials 1 — and removed them *silently*: `analysis_gaps` was empty and no payload named the loss. What it cost, on sa-token: `repo-ir` read **837 of 919 files**, published **5 807 symbols** and **154 endpoint symbols**; `ask impact com.pj.test.TestController` — a fully qualified name for a class that is on disk — answered `resolution: ambiguous` over `com.lym.controller.TestController` and `com.pj.controller.TestController`, published `risk_score: 26.0, risk_level: high` for that substitution, and reported `endpoints_affected_count: 8` against `impact-chain`'s **53** for the same symbol. Second defect in the same lines: `--include-tests` was a **no-op** on `repo-ir` and `impact`, because the flag only skipped the redundant pass while `find_java_files` was still called with its default `include_tests=False` | 5.7.3 (cross-command coherence battery, 2026-08-16) | **High** — a repository that names a package `test` is analysed with 9 % of its production classes missing from every CIR consumer, and the blast radius of a symbol is 6,6× smaller than the same product's other answer for it | **closed — on `master`, unreleased.** One authority at all six sites: the three CLI populations read `find_java_files(root, include_tests=include_tests)` and nothing else, the three list filters read `path_filters.is_test_path`. Measured after on sa-token: `repo-ir` **919 files, 6 248 symbols, 381 endpoint symbols, 902 classes** (was 837 / 5 807 / 154 / 820); `impact com.pj.test.TestController` → `resolution: exact`, one matched FQN, **53 endpoints on both commands** — the C1-16 endpoints-affected divergence on this repository closes with it. `--include-tests` now widens the population it names. Collateral nil: the coherence battery over spring-petclinic, spring-boot-admin, sagan, eureka, examples, spaghetti-api and jobrunr returns identical endpoint totals and identical `impact` scores. Regression: `tests/test_one_test_source_authority_e16.py`, 10 assertions — including one that greps `src/` for the substring shape, because the defect is a shape and it has now grown back twice |
413
413
  | E-17 | **A fully qualified target that is absent from the model is answered by every other class that shares its class name, with a confident band.** `_resolve_target` reduced every target to its last dotted segment before matching, so `ask impact com.acme.nonexistent.TestController` on sa-token — a package that exists in no repository on disk — returned `resolution: ambiguous`, `matched_fqns` naming **four classes across three unrelated packages**, and `risk_score: 69.87, risk_level: critical` for their union. Nothing in the payload said the name asked about was not there. This is C3-35 — *being more specific made the answer worse* — which was closed for **paths** and never applied to **names** | 5.7.3 (cross-command coherence battery, 2026-08-16) | **High** — the answer is about symbols the caller did not name, and it carries the band a CI gate reads; a typo in a package returns `critical` instead of an error | **closed — on `master`, unreleased.** A target carrying a package is matched as a *qualified* suffix (`fqn == t or fqn.endswith("." + t)`), so `test.TestController` still narrows to the two classes that carry that suffix and `com.a.Widget` resolves to itself alone. When nothing carries the qualified suffix the answer is `not_found` — never another class's numbers — and the near names are offered under `candidates`, which the existing ranking already orders exact-simple-name first, so the answer stays actionable. A **bare** class name keeps the suffix affordance the CLI documents (`ask impact OwnerRepository`), because there the caller has not said which package and ambiguity is the honest reading. Measured after on sa-token: `com.acme.nonexistent.TestController` → `not_found`, `risk_level: unknown`, candidates `[com.pj.test.TestController, com.pj.controller.TestController, …]`; `TestController` → `ambiguous` over the same 4; `com.pj.test.TestController` → `exact`. Regression: `tests/test_qualified_target_resolution_e17.py`, 9 assertions, including the negative control that the bare-name affordance is unchanged |
@@ -471,7 +471,7 @@ class's subject, not its provenance.
471
471
  | CL-15 | The `spring-audit` section is named **security surface**, and the catalogue is presented as the security answer | Eval #13 found the repository's **critical SQL-injection** by hand — `grep '\$\{'` plus four file reads — and ASK did not report it, because there is no taint analysis (NC-001, declared out of scope and correctly so). What ASK **did** do, and did well, was quantify its blast radius: 1 144 endpoints reached, `confidence: high`. The evaluator's framing is the one to adopt: *"es un mapa de alcance, no un buscador de defectos — pero `spring-audit` se anuncia como security surface y eso invita a confundirlo."* Vulnerability detection scored **4/10**, the axis where the named competitors (Semgrep, CodeQL, Snyk Code) win outright | open — **this is a naming defect, not a coverage one**, and the remedy is not a new rule. NC-001 already says we do no dataflow; the section's *name* out-claims its own non-coverage block, which is the one asset five evaluations agree is the moat. Two moves, both cheap: state in the section header and the `--help` one-liner what the surface *is* (configuration, wiring and effective exposure — reach, not injection), and make the reach half a **positive claim** where it is strongest, since quantifying a defect somebody else found is a capability none of those competitors offers. Related: CL-7 (the four declared misses), CL-9 (the table-stakes rules), CL-13 (the closed catalogue). **Closed 4.11.0 — the name stopped over-promising and the strength became a claim.** `non_coverage.SURFACE_CLAIMS` is the one table that says what a surface *is*, beside the rows that say what it is not, and it is rendered into both readers: the `spring-audit --help` screen and the payload's `non_coverage` block (`this_surface_is`, `this_surface_is_not`, `where_it_is_strongest`). What it reads is *the security configuration and exposure of this repository* — the rules the request chain declares, the endpoints they leave open, the controls present but disabled, the custom gates that can be bypassed, each with file and line. What it is not is *a vulnerability scanner*, in the same block, pointing at NC-001 rather than restating it. The rule family label follows (`rule_catalog`: `security surface` → `security configuration and exposure`), so the one-liner the root panel prints, the command's own screen and the README section are one sentence rather than three. The positive half is the part the row asked for and it is now published rather than implied: *given a defect somebody else found — a scanner finding, a line from a review — `ask impact <symbol>` and `ask risk` quantify what it reaches, with the confidence of each hop*, which is the question a taint scanner does not answer. No rule changed and no finding changed: this row was always a naming defect |
472
472
  | CL-16 | The stability tiers are **inverted against measured behaviour**. `pr-impact` sits in `core` — *"contract stable within a major — safe to gate CI on"* — while it publishes `LOW` over the most dangerous diff constructible in the repository (C1-36). `posture` sits in `experimental` — *"do not gate CI on it"* — while it is the one command neither evaluation found a competitor for, and it produced 2 of eval #15's 3 CRITICAL findings | A user who follows the tiers literally uses what fails and avoids what works. The tiers are a promise of stability, not a value ranking — but a `core` promise the command does not keep spends the exact asset (coherence between what is claimed and what is delivered) this product cannot afford to spend | **closed (4.10.6) — by the first branch.** C1-36's silent half is closed: `pr-impact` no longer publishes a fan-in verdict over a component the container invokes; it floors, warns, publishes the population and declares NC-010. So `core` — *safe to gate CI on* — is a promise it can keep, and a gate on that diff now fails instead of passing green. The cheaper half shipped with it: **`posture` moves `experimental` → `supported`**, and the `[EXPERIMENTAL]` banner leaves its `--help`. It is the one command neither evaluation found a competitor for, it produced 2 of eval #15's 3 CRITICAL findings, and refusing to promise stability about something that has been stable for four releases spends the same asset an overstated promise does |
473
473
  | CL-17 | **Refuted, and recorded so it is not re-opened.** Eval #14 reports `review-pr` as *"no documentado en `--help`, ni en la lista de tiers"* — *"un usuario que audita PRs no puede saber que existe"* | **False against 4.10.4 and 4.10.5.** Verified three ways in the shipped source: the command is registered (`cli.py:9807`, `@app.command("review-pr")`), it is listed in `HELP_PANELS` under *Change and risk* (`cli.py:12382`) and it renders there in `ask --help`, and it is in `COMMAND_TIERS` under `supported` (`cli.py:181`) since 3.2.2. The related half is false too: `prepare-context --help` already enumerates the closed enum (*"Task: explain \| fix-bug \| refactor \| generate-tests \| onboard \| review-pr \| delta"*, `cli.py:3879`) | **closed — no work.** One real residue, and it is C4-19's cause rather than its own row: the **prose** *Change and risk* header near the top of `--help` names eleven commands and omits `review-pr`, while the panel 170 lines further down includes it. A reader who scans the header and stops concludes the command does not exist. That is an argument for a shorter, scope-aware front page, not for a second catalogue |
474
- | CL-18 | `--jobs` / `ASK_JOBS` are published as the parallelism control, default `cpu_count()-1`, with the honest qualifier *"parallelism only warms the content-addressed parse cache — the answer is byte-identical at any value"* | True, and it does not say the part that decides whether the flag is worth setting: **parsing is ~10 s of a 13,6-minute repo-wide run, and on a warm cache it is 0 s**. Eval #18 sampled the process and measured `threads=4` for the first seconds and `threads=1` with CPU/wall **0,97** for the remaining ~99 %, then corrected its own four-round complaint from *"no parallelism"* to *"the parallelism covers the wrong phase"*. An operator who reads the flag and raises it on a warm repository buys nothing and has no way to learn that from the surface | **claim corrected 4.15.0** — the help, the README and the user guide state that the rule pass is not parallelised and that a warm cache leaves the flag nothing to parse, from the one help string the flag is declared with. The capability half is **F-AT**. Row C3-84. **Eval #19 re-raised this as an open item and it does not reproduce (measured 4.16.0):** the report's *"corregir la documentación del flag"* quotes `--jobs` without the sentence, while the shipped `JOBS_OPTION_HELP` reads *"It does NOT parallelise the rule pass, which is where a repository-wide audit spends most of its time, and on a warm cache there is nothing left to parse, so raising it buys nothing there."* Verified against the string every command declares the flag with. Recorded here so it is not re-opened a third time — the CL-17 precedent |
474
+ | CL-18 | `--jobs` / `ASK_JOBS` are published as the parallelism control, default `cpu_count()-1`, with the honest qualifier *"parallelism only warms the content-addressed parse cache — the answer is byte-identical at any value"* | True, and it does not say the part that decides whether the flag is worth setting: **parsing is ~10 s of a 13,6-minute repo-wide run, and on a warm cache it is 0 s**. Eval #18 sampled the process and measured `threads=4` for the first seconds and `threads=1` with CPU/wall **0,97** for the remaining ~99 %, then corrected its own four-round complaint from *"no parallelism"* to *"the parallelism covers the wrong phase"*. An operator who reads the flag and raises it on a warm repository buys nothing and has no way to learn that from the surface | **claim corrected 4.15.0** — the help, the README and the user guide state that the rule pass is not parallelised and that a warm cache leaves the flag nothing to parse, from the one help string the flag is declared with. The capability half is **F-AT**. Row C3-84. **Eval #19 re-raised this as an open item and it does not reproduce (measured 4.16.0):** the report's *"corregir la documentación del flag"* quotes `--jobs` without the sentence, while the shipped `JOBS_OPTION_HELP` reads *"It does NOT parallelise the rule pass, which is where a repository-wide audit spends most of its time, and on a warm cache there is nothing left to parse, so raising it buys nothing there."* Verified against the string every command declares the flag with. Recorded here so it is not re-opened a third time — the CL-17 precedent. ⚠ **The qualifier itself was corrected in 5.8.7**: *"where a repo-wide audit spends most of its time"* was an estimate, and `metadata.timings` measures the rule pass at 8 % against 88 % for the CIR build. The three surfaces carry the measurement now. Row C3-84. |
475
475
  | CL-19 | Every axis of the product publishes its own `non_coverage`, its own units and its own confidence, and the run states `spring_detected` inside each payload | True per axis, and **no surface adds them up**. On a repository where the answer is `false`, eval #20 measured what the operator actually receives: 614 endpoints `none_detected`, `access_not_decided` **3 152/3 152**, `security_model: "unknown"`, `TX-001..006` at zero, `boot3.applicable: false` — six axes returning the shape of an answer, none of them saying *this axis does not model your stack*. The evaluator had to derive it: *"tuve que deducirlo de `spring_detected: false` enterrado en el JSON"*, and their verdict on the product is a direct consequence — *"una herramienta de inventario estructural muy buena vendida como auditor de seguridad"*. The facts are all present; what is missing is the one sentence at the top that composes them | **claim corrected 4.17.0** — `non_coverage.stack_fit()` does the arithmetic from the flag the payload already carries, naming the Spring-shaped axes **and what each reports on such a repository**, beside the structural axes that answer here as they do anywhere. Published as `stack_fit` in the audit payload and said once through `_notice` (terminal-gated, so piped output stays byte-identical). A repository this product models gains **no field**. Extending the statement to the other commands' payloads is **F-AY** |
476
476
  | CL-21 | The root `--help` narrative states *"repo-wide/deep compositions can take minutes on multi-thousand-endpoint repositories. Use deep jobs nightly or with `ASK_MAX_ANALYSIS_SECONDS`/`ASK_PROGRESS` when CI needs an explicit budget"* (`cli.py`) | It was true when it was written and 4.18.0 deleted the premise. Eighth field round, same 3 342-file / 3 574-endpoint subject: the whole security battery — `spring-audit` 8,4 s, `posture --resolve-env` 8,0 s, `impact-chain` 5,5 s, `pr-impact` on security files 11,8 s — fits in **under 40 s of pipeline**, and the field's operational note is the inverse of ours: *"ya no hace falta `ASK_MAX_ANALYSIS_SECONDS` en este repo … los repo-wide ya se pueden lanzar en foreground"*. This is the prose sibling of C3-97's data, and it must be **re-derived from the same authority**, not re-written by hand: a sentence maintained separately from `cache_model` is the second copy of a fact, which is the rule this ledger exists to enforce | **corrected 5.0.0** — generated, not rewritten. `cache_model.cost_sentence()` builds the paragraph from the anchors, `cli._cost_sentence()` prints it, and a battery asserts the front page contains exactly what the model generates, so the prose cannot age separately from the table again. The README, the guide's "Start here" framing and the spinner paragraph were corrected against the same measurements |
477
477
  | CL-22 | *(non-reproducers from eval #22, recorded so they are not raised a fourth time — the CL-17/CL-20 precedent)* | **B16 (*"a partial answer is invisible to a programmatic consumer"*) shipped in 4.17.0 as F-AV**, in both halves the report asks for: `partial_contract.PARTIAL_EXIT_CODE = 75` for a gating command whose evidence was truncated, and `floor_counts()` renaming `total_defects` → `total_defects_floor` when `partial` is true, wired at `risk.py`, `spring_findings.py` and `audit_report.py`. The field read a **complete** 4.18.0 run — no `_partial`, `rules_not_run` empty — so it saw the unrenamed key correctly: the rename is conditional by design, because a full answer's total is not a floor. · **B3c (progress emission) is C3-82 + C3-92, closed in 4.15.0 and 4.17.0**; the report agrees it is no longer observable here. · **B8 (serial rule pass) is F-AT and the field itself downgraded it** from critical to informational after measuring `--jobs 1` ≡ `--jobs 19` ≡ 7,8 s with byte-identical output — parallelising eight seconds buys ~7 s, and `JOBS_OPTION_HELP` already states the rule pass is not parallelised (CL-18) | **no action on B16/B3c/B8** — verified against source at 4.18.0 before opening anything. The rule that produced this row is the audit-#8 lesson: check the build against the source before accepting a re-derived complaint |
@@ -136,7 +136,7 @@ pipx install sourcecode # isolated install, no venv needed
136
136
 
137
137
  # Verify
138
138
  ask version
139
- # ask 5.8.6
139
+ # ask 5.8.7
140
140
  ```
141
141
 
142
142
  Requires Python 3.9+.
@@ -177,7 +177,7 @@ estimate scaled by file count would be wrong in the direction that costs you a s
177
177
  and an anchor printed without its release goes on recommending a nightly job for a command
178
178
  that has come to finish in seconds (C3-97).
179
179
 
180
- Thirty-nine commands and six command groups exist. Four of them carry most of the measured
180
+ Forty commands and six command groups exist. Four of them carry most of the measured
181
181
  value in field use, and they are the ones to learn first:
182
182
 
183
183
  | Start with | Because |
@@ -337,8 +337,45 @@ ask spring-audit . --min-severity high
337
337
  ask spring-audit . --table --rule GATE-001 --top-n 20
338
338
  ask spring-audit . --compact # bounded summary (top 5 per list; counts intact)
339
339
  ask spring-audit . --output audit.json
340
+ ask spring-audit . --since origin/main --fail-on-new # gate on what THIS change added
341
+ ask spring-audit . --remediation-diff # the patches a reviewer applies
342
+ ask spring-audit . --clusters # the work items behind the findings
340
343
  ```
341
344
 
345
+ **`--since <ref>` / `--fail-on-new` — the gate a repository with debt can turn on.**
346
+ `--ci` is absolute: any finding fails, debt included. `--since` audits the tree at the
347
+ ref too and joins the two on the content-stable `defect_id`, so `since.new`,
348
+ `since.resolved` and `since.counts` answer *"what did this change add?"* — and
349
+ `--fail-on-new` exits 1 only on that. Both sides are audited under the same `--scope`,
350
+ `--min-severity` and test-source decision, published as `since.compared_under`: a delta
351
+ between two different questions is not a delta. A renamed symbol reads as one resolved
352
+ plus one new, and the payload says so rather than guessing. A comparison that could not
353
+ be made publishes `verdict: unknown` and exits non-zero — a run that did not finish is
354
+ not a pass. The ref is read with `git archive`: nothing is written to the repository,
355
+ and that side is parsed cold, so it roughly doubles the cost.
356
+
357
+ **`--clusters` — 386 rows as the handful of things somebody has to do.** A work
358
+ item is `(defect_kind, repair_shape, locus)`: what is wrong, whether fixing it is a
359
+ mechanical edit somebody reviews or a decision somebody makes, and the package it lives
360
+ in. The axis is declared in the payload because the whole answer depends on it — the
361
+ same repair in two unrelated packages is two work items, and merging them would produce
362
+ a number nobody can assign. Each item carries its size (defects, witnesses, files,
363
+ symbols), its severity ceiling, the repair, and its owner **where CODEOWNERS declares
364
+ one**; commit history says who touched a file, which is not who owns the decision, so it
365
+ is not published as one. The effort figure is an assumption and says so: the rate, the
366
+ count and their product are all in the payload, so substituting your team's rate is
367
+ multiplication and the clustering does not move.
368
+
369
+ **`--remediation-diff` — the findings as patches, not as prose.** For every defect whose
370
+ repair is the **lexical inverse** of the defect — today: a security control commented
371
+ out with `//`, which on one field subject was 192 endpoints one `//` away from live —
372
+ this emits a unified diff `git apply` takes as written. Every other defect appears under
373
+ `not_patched` with the reason, read from the same authority `--recipe` states it from.
374
+ Every patch carries `decides`: what applying it decides. That distinction is the design
375
+ — `--recipe` is a document an executor *runs*, and re-enabling a control somebody
376
+ disabled changes runtime behaviour nobody measured; a diff is how a person makes that
377
+ call with the evidence in front of them.
378
+
342
379
  **TX patterns (TX-001..TX-006):** proxy bypass, nested transactions, readOnly propagation, NOT_SUPPORTED in active TX, exception swallowing, self-invocation of a `@Transactional` sibling.
343
380
  **SEC patterns (SEC-001..SEC-008):** unsecured endpoints, CVE-2025-41248 `@PreAuthorize` inheritance bypass, `@Transactional` on controllers, passwords stored under a fast unsalted digest, CSRF disabled under session-bearing authentication, cookies created without the Secure attribute, credentials stored in a deployment descriptor, and SQL built by MyBatis string interpolation. **DEAD-001** reports a security control that is present only in comments and therefore disabled.
344
381
  **GATE patterns (GATE-001..GATE-004):** custom AOP/security-gate risks over the same custom-gate authority published in `security_posture`: proxy bypass (`private`/`final`/self-invocation), advice that appears to fail open, disabled or tautological annotation attributes, and fragile `args[0]`/reflection binding.
@@ -886,6 +923,38 @@ Three properties make the series usable as evidence rather than as a chart:
886
923
 
887
924
  ---
888
925
 
926
+ ### `ask explain-endpoint '<VERB> <path>' [repo]` *(supported)* — show your work for one route
927
+
928
+ ```bash
929
+ ask explain-endpoint 'POST /api/v1/auth/login'
930
+ ask explain-endpoint 'GET /admin/users/{id}' /path/to/repo
931
+ ask explain-endpoint '/health' --profiles prod # every verb on that path
932
+ ```
933
+
934
+ A row of the endpoint census raises exactly one question — **why does this route have
935
+ this verdict?** — and the census answers with the verdict alone. This is the
936
+ derivation behind it, in six steps, each naming the authority that produced it and the
937
+ file it can be checked against:
938
+
939
+ | step | what it answers | authority |
940
+ |---|---|---|
941
+ | `match` | which endpoint the query names | `canonical_ir.request_identity` |
942
+ | `declaration` | the mapping that mounts it, inherited or not, with `file:line` | `repository_ir.route_surface` |
943
+ | `handler` | the guard the handler itself declares, and its scope | `CanonicalEndpoint.security` |
944
+ | `chain` | the filter-chain rule that decides the request, with the rule's own line | `posture.access_projection` |
945
+ | `profiles` | which chain configuration is active, and what is undecided | `posture` |
946
+ | `verdict` | the coverage answer | `security_posture.endpoint_security_surface` |
947
+
948
+ **A step that could not be derived says so** — `undetermined`, with the reason. None
949
+ of the six is ever absent in silence: a derivation with a hole in it is the thing this
950
+ command exists to expose, so it may not hide one of its own.
951
+
952
+ **Lines are located, never guessed.** The locator reads comment-blanked source, so a
953
+ mapping inside a comment is never cited as the declaration; where the declaration is
954
+ not located, `line` is `null` and `line_basis` says why.
955
+
956
+ ---
957
+
889
958
  ### `ask selftest <repo>` *(supported)* — this product's defect ledger, run against your repository
890
959
 
891
960
  ```bash
@@ -1316,7 +1385,7 @@ What invalidates what:
1316
1385
  | `--copy` | `-c` | Copy output to clipboard |
1317
1386
  | `--no-redact` | | Disable automatic secret redaction |
1318
1387
  | `--exclude PATTERN` | | Skip directories matching pattern |
1319
- | `--jobs N` | `-j` | Worker processes that parse files (default: CPU count − 1). Only warms the parse cache — the answer is byte-identical at any value. Not the rule pass, where a repo-wide audit spends most of its time, and nothing at all on a warm cache. Env: `ASK_JOBS` |
1388
+ | `--jobs N` | `-j` | Worker processes that parse files (default: CPU count − 1). Only warms the parse cache — the answer is byte-identical at any value. Not the rule pass — measured, the CIR build is 88 % of a repo-wide audit's clock and every rule together 8 %, so this is the phase with the cost; nothing at all on a warm cache, where there is nothing left to parse. Env: `ASK_JOBS` |
1320
1389
  | `--no-write` | | Create nothing inside the analysed repository; a command that cannot answer without writing says so, naming the path on stderr. Env: `ASK_READONLY=1` |
1321
1390
  | `--progress MODE` | | Heartbeat lines for a non-TTY supervised run: `1`/`stderr`, `off`, `file:<path>`. Env: `ASK_PROGRESS` |
1322
1391
  | `--version` | `-v` | Show version and exit |