sourcecode 5.8.11__py3-none-any.whl → 5.8.13__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Potentially problematic release.


This version of sourcecode might be problematic. Click here for more details.

sourcecode/__init__.py CHANGED
@@ -4,4 +4,4 @@ ASK Engine is the product. ``ask`` is the canonical CLI command; ``sourcecode``
4
4
  the legacy compatibility alias and the Python/PyPI package name. See
5
5
  docs/PRODUCT_IDENTITY.md (normative)."""
6
6
 
7
- __version__ = "5.8.11"
7
+ __version__ = "5.8.13"
@@ -16,20 +16,57 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
16
16
 
17
17
  ## Current Synchronization
18
18
 
19
- **As of 2026-08-19, the `5.8.11` release commit:** this is the current status of
20
- the 5.8.9 re-audit queue. The audit observations below remain historical evidence;
21
- only the table here declares present state.
19
+ **As of 2026-08-20, release `5.8.13` and the external `5.8.11` re-audit:** this is the current status of
20
+ the 5.8.8/5.8.9 audit queue. The subject remained `saint-server` at `3dde0376`, with a
21
+ byte-identical CIR and the same 53/61 invocables exercised. The score is **78/100**
22
+ (75 in 5.8.8, 80 in 5.8.9). The observations below remain historical evidence; only
23
+ this table declares present state.
22
24
 
23
25
  | Rows | Current status | Evidence |
24
26
  |---|---|---|
25
- | `AUD-588-B01`, `AUD-589-B01`, `C3-122` | **closed** | `20596c8`; failed writes cannot suppress the global output-promise guard, covering both `risk` and `audit-report` emit paths. |
26
- | `AUD-588-B03` | **closed** | `16a52cc`; `SEC-009`, static `@Profile` evidence and `permit_all_rules` are emitted with regressions. |
27
- | `AUD-588-B04` | **closed** | `fd34eb1`; `data-exposure` preserves the chain decision for every route. |
27
+ | `AUD-588-B01`, `AUD-589-B01`, `C3-122` | **implemented in 5.8.12** | `9884488` requires `risk` and `audit-report` to deliver stdout or their requested artifact before exit 0; the remaining closure evidence is a field-scale subprocess reproduction. |
28
+ | `AUD-511-R01` / `AUD-588-B03` | **implemented in 5.8.12** | `13362ec` makes SEC-009 read comment-blanked Java, so a Javadoc profile cannot override the live `@Profile("!m3")` declaration. |
29
+ | `AUD-588-B04` | **retracted audit premise** | `data-exposure` correctly preserves `chain_decision`; the reported `coverage_unknown` routes were custom-gate annotations whose application was not configured. `fd34eb1` remains a valid additive disclosure, but no further B04 fix is due. |
28
30
  | `AUD-588-B05` | **closed** | `c15f5b6`; changed container-wired components block a false `unaffected` disposition. |
29
- | `AUD-588-B07` | **closed** | `5cc6ba7`; ASK-09 reads the envelope error contract and covers the wrapped shape. |
31
+ | `AUD-511-R03` / `AUD-588-B07` | **implemented in 5.8.12** | `47aee1e` gives the published `OUTPUT_TOO_LARGE` envelope precedence over an inconsistent process status and preserves the excerpt in every branch. |
30
32
  | `AUD-588-B08` | **closed** | `c084555`; `impact` publishes its independent stdout-byte budget. |
31
- | `AUD-588-B11` (container/units subset) | **closed** | `ca72b61`, `a49cf6a`; plan/compare publish container wiring and impact-chain names security-surface units. |
32
- | `AUD-588-B02`, `B06`, B11 residuals, `B12` | **open** | Require provider non-coverage verification, remaining contract parity and controlled cost calibration; no closure is inferred from the release. |
33
+ | `MCP-001` | **closed 5.8.13** | Deep corpus testing reproduced concurrent `CliRunner` stream capture leaking a tool response to the host stdout. `b2c9660` serializes the narrow process-global capture seam; concurrent real MCP calls and a regression test verify response isolation. |
34
+ | `AUD-588-B11` (container/units subset) | **partially implemented in 5.8.12** | `d745d31` discovers nested Maven modules; `273ae8e` suppresses unsupported `pr-impact.unaffected_basis`. `impact` naming, source population units, risk tiers and plan narrative remain open. |
35
+ | `AUD-511-R02` / `AUD-589-B02` | **partially mitigated in 5.8.12, still under investigation** | `b93f230` reuses one `SpringSemanticModel` across every declared `data-exposure` seed; `473f47a` routes `validation` through the shared CIR cache, and `91bd0c0` retains subcommand `--no-cache` compatibility. On the local BroadleafCommerce corpus at the final analyzer fingerprint, a first shared-CIR fill took 10.46 s and its warm `validation --compact` hit took 4.01 s with byte-identical output. This is a same-host cache-state observation, not the 5.8.9/5.8.11 external A/B: the six-run C3-120 comparison and the configured `saint-server` reproduction remain required. |
36
+ | `AUD-588-B02`, `B06`, B11 residuals, `B12` | **partially open in 5.8.12** | `3869c87` excludes explicit EclipseLink from Hibernate applicability and effort; `037223e` publishes provider non-coverage. B11 parity and controlled cost calibration remain required. |
37
+
38
+ ### Re-audit Intake: 5.8.11
39
+
40
+ The third external pass uses the same `saint-server` snapshot (`3dde0376`), read-only
41
+ protocol and warm-cache preparation as 5.8.8 and 5.8.9. The CIR hash is identical across
42
+ all three versions. It exercises 53/61 invocables (87%) and reports 78/100. That makes
43
+ rule, output and execution changes attributable without claiming that the application was
44
+ run or that the eight blocked invocables were covered.
45
+
46
+ | ID | Priority | Finding | Required correction / verification |
47
+ |---|---|---|---|
48
+ | `AUD-588-B01`, `AUD-589-B01`, `C3-122` | P0 | `risk -o` can exit 0 with no file; `risk` and `audit-report` can finish with empty stdout. The failure has fast and slow composition paths. | Add a real-subprocess invariant: exit 0 must produce parseable stdout or the requested artifact. Exercise `risk` and `audit-report` with and without `-o`, plus budget-cut partial output. |
49
+ | `AUD-511-R01` | P1 regression | SEC-009 takes `@Profile` from comments and reports an active `@Profile("!m3")` request chain as inactive. | Reuse comment-blanked source and declaration-local profile evidence, or consume posture's authority. Regress a misleading javadoc and both positive/negative profile forms. |
50
+ | `AUD-511-R02` / `AUD-589-B02` | P1 investigation | `endpoints`, `impact`, `spring-audit`, `modernize`, `validation` and `timeline` measure 1.4-2.4x slower than 5.8.9; `data-exposure --config` remains high. | **Partial implementation:** `b93f230` shares the data-exposure semantic model across seeds and `473f47a` makes `validation` reuse the shared CIR; `91bd0c0` confirms `--no-cache` remains a compatibility no-op. Re-run the idle-host, isolated-cache, interleaved-control protocol against the audited `saint-server` configuration before attributing or closing the field regression. |
51
+ | `AUD-511-R03`, `AUD-588-B07` | P1 regression | ASK-09 reports `not_applicable` even though its literal probe emits `OUTPUT_TOO_LARGE` 12x above the ceiling. | Derive branch selection exclusively from the output envelope. Publish `observed_payload_excerpt` for pass, fail, error and not-applicable rows. |
52
+ | `AUD-588-B02` + `B06` | P1 | The EclipseLink repository still receives Hibernate readiness, rewrite-zone classification and 32.9-66.2 person-days; `non_coverage` is null. | A declared competing provider must remove the Hibernate axis from readiness, effort and classification and add provider non-coverage. EclipseLink automation remains out of scope. |
53
+ | `AUD-588-B11` residual | P2 | `impact.matched_fqns` mixes exact resolution with 235 implementations; `migrate-recipe` misses nested Maven manifests; source-file populations lack units; `risk` tier labels diverge; plan/pr-impact prose trails their new container facts. | Bind consumers to established authorities, name populations and units, and add cross-command parity tests. Do not add parallel detection. |
54
+ | `AUD-588-B12` | P2 | 5.1.0 anchors disagree in both directions with 5.8.11 measurements; `timeline` is advertised as foreground despite a 404s sample. | Recalibrate only after controlled measurements and align foreground advice with the explicit opt-in progress policy. |
55
+
56
+ ### Feature Recommendations: 5.8.11
57
+
58
+ | ID | Recommendation | Current decision |
59
+ |---|---|---|
60
+ | `AUD-588-F01` | EclipseLink/JPA-provider migration axis. | **Deferred product feature.** First close B02/B06 so the current Hibernate answer is honest about non-coverage. |
61
+ | `AUD-588-F02` | Shared answer-contract authority for profiles, units, populations, caveats and container reach. | **Partial.** The SEC-009 regression and B11 residuals show that new consumers are still re-deriving facts. |
62
+ | `AUD-588-F03` | Versioned cost model by repository size and cache state. | **Partial.** Existing dimensions do not make stale 5.1.0 anchors operationally safe; controlled remeasurement is prerequisite. |
63
+ | `AUD-588-F04` | Default foreground heartbeat without corrupting machine output. | **Explicit opt-in by decision.** Until revised, cost advice must not imply that long commands emit progress by default. |
64
+ | `AUD-588-F05` | Raise audit coverage through safe operational paths and dirty-tree `verify-edit`. | **Locally verified only.** The external run remains 53/61; treat dirty-tree behavior as a validation gap, not a closed product claim. |
65
+
66
+ **Recommended order:** composition-output invariant → SEC-009 profile regression → controlled
67
+ performance measurement → EclipseLink applicability/non-coverage → ASK-09 → B11 narrative
68
+ and parity residuals → cost-model calibration. Keep B04 closed as a retracted premise, and
69
+ do not implement EclipseLink migration automation as part of this bug battery.
33
70
 
34
71
  **As of 2026-08-19, the `5.8.9` release commit:** the latest audit queue was implemented in
35
72
  committed post-5.8.7 fixes. This table is
@@ -558,7 +595,7 @@ class's subject, not its provenance.
558
595
 
559
596
  | ID | Item | Found | Severity | Status |
560
597
  |---|---|---|---|---|
561
- | P-1 | Telemetry is **opt-out**. Auditing a public-sector health client's code, it must be disabled *before* the first run, not after | 3.2.0 | **High (procurement)** | **closed 3.3.0** — opt-in: with no explicit choice on record, `is_enabled()` is False everywhere, so a fresh install transmits nothing on the run that prints the notice. The CI special case is gone with the default it existed to override; CI is still detected for one purpose, skipping a notice that has no terminal to appear on. The first-run notice became an invitation that states what *would* be collected. `telemetry status` distinguishes "no choice recorded" from "you chose disabled" (`stored_choice()` returns `None`, not `False`) — a buyer auditing this needs to know which one they have. An explicit choice made before the change is honoured unchanged |
598
+ | P-1 | Telemetry default and procurement disclosure | 3.2.0 | **High (procurement)** | **reopened and corrected 5.8.12** — telemetry is on by default by explicit product decision, with `ask telemetry disable`, `SOURCECODE_TELEMETRY=0` and `DO_NOT_TRACK=1` honored before any event is sent. The first-run notice discloses the fields and the disable path. `telemetry status` distinguishes the default from an explicit disabled choice; the privacy policy documents the endpoint, 90-day retention and bounded anonymous fields. This supersedes the historical 3.3.0 opt-in closure for the current release |
562
599
  | P-2 | Licence state is ambiguous: `auth status` → `{"status":"unauthenticated","pro":true,"pro_reason":"early-adoption unlock"}`. The buyer cannot tell what it will cost or what they lose when the door closes | 3.2.0 | Medium | **closed 3.4.0** — one authority, `license.entitlement()`: `entitlement` (what runs), `source` (why: `license_key` / `early_adoption_unlock` / `free_tier`), `authenticated` (a separate fact — today one can be entitled without a credential), `paywall_active`, and `when_it_changes`, which states what the user loses when that source stops applying, in the terms the gate uses and **never as a price**. `is_pro` — what every gate reads — is now derived from it rather than computed beside it, which is how the status page came to contradict the commands. The legacy keys stay for consumers, are derived from the authority, and are listed in `deprecated_fields` |
563
600
 
564
601
  | P-3 | **The price is not published.** With P-2 closed the entitlement is unambiguous, but a buyer still cannot see what Pro costs or what a team of 25 pays. Eval #5 lists *"precio Pro público"* among the conditions for the top of its price band, beside the robustness fixes | 3.2.1 (eval #5) | Medium (procurement) | open — packaging question, not a page: the field's own reading is that **the value is frontal** (a first audit replaces 5–15 person-days) **and the recurring half is the CI gate**, so a full seat subscription on a product whose value peaks on run one *"genera churn en el mes 4"*. Model to test: one-off assessment **3–5 k€ per repo** + **8–12 €/dev/month** for the gate. Depends on C3-22: without the gate there is nothing recurring to sell |
@@ -39,10 +39,12 @@ Evidence — every answer is backed by auditable evidence, not a gues
39
39
  CLI commands — impact, endpoints, spring-audit, explain, … each a projection
40
40
  ```
41
41
 
42
- The key idea: **the source is parsed once** into the Semantic IR; everything above that is
43
- computed from the IR, cached, and shared. `impact`, `endpoints`, and `spring-audit` are not
44
- three analyzers — they are three questions asked of the same graph. (The extraction and
45
- consumption contract is fixed in the architecture ADRs 0001–0004 under `docs/architecture/`.)
42
+ The key idea: the extraction is **content-addressed**. Commands reuse the parse cache and,
43
+ where their analysed scope matches, the shared Canonical IR; `ask cache model` names what a
44
+ warm buys for each command rather than implying that every projection costs the same. In
45
+ 5.8.12, `validation` enters through that shared CIR and `data-exposure` reuses one semantic
46
+ model across all declared label seeds. (The extraction and consumption contract is fixed in
47
+ the architecture ADRs 0001–0004 under `docs/architecture/`.)
46
48
 
47
49
  ---
48
50
 
@@ -136,7 +138,7 @@ pipx install sourcecode # isolated install, no venv needed
136
138
 
137
139
  # Verify
138
140
  ask version
139
- # ask 5.8.11
141
+ # ask 5.8.13
140
142
  ```
141
143
 
142
144
  Requires Python 3.9+.
@@ -1197,7 +1199,7 @@ explaining. The rest do one thing; `ask <command> --help` is the whole story.
1197
1199
  | `ask retrieve` | parked | Typed knowledge queries over the model — kept working, no longer developed |
1198
1200
  | `ask baseline capture\|diff\|trend` | supported | Versioned architectural metrics over time; see [`ask trend`](#ask-trend-dir--how-the-architecture-moved) |
1199
1201
  | `ask cache status\|warm\|model\|clear\|freshness` | supported | Cache inspection; `ask cache model` states what a warm buys each command |
1200
- | `ask config` · `ask version` · `ask activate` · `ask auth` · `ask telemetry` · `ask mcp` | supported | Configuration, version, licence, authentication, telemetry (off by default), MCP integration |
1202
+ | `ask config` · `ask version` · `ask activate` · `ask auth` · `ask telemetry` · `ask mcp` | supported | Configuration, version, licence, authentication, telemetry (on by default), MCP integration |
1201
1203
 
1202
1204
  ## Typical workflows
1203
1205
 
@@ -1285,7 +1287,7 @@ ask mcp serve
1285
1287
  ask mcp remove
1286
1288
  ```
1287
1289
 
1288
- The MCP server exposes structural analysis tools to AI agents without requiring the agent to call the CLI directly. Claude Desktop and Cursor can query impact, endpoints, and context through the MCP protocol.
1290
+ The MCP server exposes structural analysis tools to AI agents without requiring the agent to call the CLI directly. Claude Desktop and Cursor can query impact, endpoints, and context through the MCP protocol. Concurrent requests are accepted safely: ASK serializes only the in-process CLI capture seam, because its stdout/stderr redirection is process-global, so one tool response cannot contaminate another.
1289
1291
 
1290
1292
  ---
1291
1293
 
@@ -1482,10 +1484,12 @@ the tree. Set `ASK_RUNS_IN_REPO=1` only when `.ask/runs` belongs in that
1482
1484
  repository's own artefacts, or `ASK_RUNS_DIR=/path/to/runs` to choose an external
1483
1485
  location explicitly.
1484
1486
 
1485
- **Large-repo budgets.** Set `ASK_MAX_ANALYSIS_SECONDS=<seconds>` to bound a
1486
- repo-wide or deep command. The budget is **honoured, not vetoed**: if the value
1487
- is below the class floor ASK says so on stderr and runs anyway, and a run that
1488
- spends its budget publishes what it measured with
1487
+ **Large-repo budgets.** Set `ASK_MAX_ANALYSIS_SECONDS=<seconds>` to bound the
1488
+ phase-runner commands `spring-audit`, `risk` and `audit-report`. It is not a
1489
+ process-wide timeout for every repo-wide or deep command. The budget is
1490
+ **honoured, not vetoed**: if the value is below the class floor ASK says so on
1491
+ stderr and runs anyway, and a run that spends its budget publishes what it
1492
+ measured with
1489
1493
 
1490
1494
  ```json
1491
1495
  "_partial": {"partial": true,
@@ -1505,7 +1509,9 @@ inventories `rules_run` / `rules_not_run` with `rules_run_count` and
1505
1509
  stops at 1 408 files inside the `SEC-004` group and ends at 5.1s.
1506
1510
 
1507
1511
  Every count in such an answer is a floor over what ran — a family or phase that
1508
- never ran is named, never reported as an absence of findings. The mark lives
1512
+ never ran is named, never reported as an absence of findings. Commands outside
1513
+ that phase-runner set are not bounded by this variable and must not be CI-gated
1514
+ on it. The mark lives
1509
1515
  where the numbers are as well as at the root: `summary.partial`,
1510
1516
  `summary.counts_are_floor` and `summary.counts_basis`, with
1511
1517
  `confidence_level` capped at `low` and `confidence_basis` saying the cut is the
@@ -127,6 +127,11 @@ MANUAL_BASIS: "dict[str, str]" = {
127
127
  "clause) and an injection sink for a value, and an executor cannot tell "
128
128
  "them apart without the caller's intent"
129
129
  ),
130
+ "SEC-009": (
131
+ "whether an anonymous request matcher is intentional depends on the "
132
+ "deployment profile and authorization policy, so removing it would be "
133
+ "an unreviewed behaviour change rather than a safe rewrite"
134
+ ),
130
135
  "DEAD-001": (
131
136
  "a disabled security control is a decision somebody made; re-enabling it "
132
137
  "changes runtime behaviour nobody has measured here"
sourcecode/cache_model.py CHANGED
@@ -653,9 +653,9 @@ def cost_sentence() -> str:
653
653
  f"files, and a repository-wide audit measured {seconds:g} s on a "
654
654
  f"{_thousands(FIELD_ANCHOR_JAVA_FILES)}-file repository "
655
655
  f"({FIELD_ANCHOR_MEASURED_VERSION}{stale}).{gate} `ask cache model` prints "
656
- f"the figure and the build behind it for every command; "
657
- f"ASK_MAX_ANALYSIS_SECONDS/ASK_PROGRESS bound and narrate a run where CI "
658
- f"wants an explicit budget."
656
+ f"the figure and the build behind it for every command; ASK_PROGRESS "
657
+ f"narrates supervised work, while ASK_MAX_ANALYSIS_SECONDS bounds only "
658
+ f"the phase-runner commands (spring-audit, risk and audit-report)."
659
659
  )
660
660
 
661
661
 
sourcecode/cli.py CHANGED
@@ -572,7 +572,8 @@ def _build_help_text(scope: "Optional[Any]" = None) -> str:
572
572
 
573
573
  Deterministic Java/Spring semantics and reusable structural context for AI coding agents.
574
574
 
575
- Cache warms on first scan; later calls reuse pre-built context instead of rescanning.
575
+ Content-addressed parse and context caches reuse prior knowledge where a command's
576
+ scope matches; `ask cache model` says exactly what a warm buys per command.
576
577
  {_cost_sentence()}
577
578
 
578
579
  {_start_here_block(scope)}
@@ -584,6 +585,12 @@ Cache warms on first scan; later calls reuse pre-built context instead of rescan
584
585
  ask --compact --git-context [dim]# + git hotspots and uncommitted files[/dim]
585
586
  ask --agent [dim]# full structured JSON for AI agents[/dim]
586
587
 
588
+ [bold]Performance / CI:[/bold]
589
+ validation . --compact [dim]# reuses the shared CIR on a warm matching scope[/dim]
590
+ data-exposure . [dim]# one semantic model is shared across declared label seeds[/dim]
591
+ ASK_MAX_ANALYSIS_SECONDS=N [dim]# deadline only for spring-audit, risk, audit-report[/dim]
592
+ [dim]Other commands are not time-bounded by this variable; use `--progress` or `--detach` where offered.[/dim]
593
+
587
594
  [bold]Change and risk:[/bold]
588
595
  impact-chain <Class> . [dim]# blast radius; TX/SEC findings opt-in[/dim]
589
596
  impact <Class> . [dim]# reverse deps → endpoints reached[/dim]
@@ -1954,6 +1961,7 @@ def _output_ceiling_error(content: str, *, to_file: bool) -> "Optional[dict[str,
1954
1961
  #: that cannot produce the file it was asked for must not be able to say it did.
1955
1962
  _OUTPUT_PROMISED: "Optional[str]" = None
1956
1963
  _OUTPUT_DELIVERED: bool = False
1964
+ _ANSWER_REQUIRED_BY: "Optional[str]" = None
1957
1965
 
1958
1966
 
1959
1967
  def _record_output_promise(argv: "list[str]") -> None:
@@ -1966,9 +1974,10 @@ def _record_output_promise(argv: "list[str]") -> None:
1966
1974
  parser owns that. It is used only to ask, at the end, whether a file the run
1967
1975
  promised exists.
1968
1976
  """
1969
- global _OUTPUT_PROMISED, _OUTPUT_DELIVERED
1977
+ global _OUTPUT_PROMISED, _OUTPUT_DELIVERED, _ANSWER_REQUIRED_BY
1970
1978
  _OUTPUT_PROMISED = None
1971
1979
  _OUTPUT_DELIVERED = False
1980
+ _ANSWER_REQUIRED_BY = None
1972
1981
  for i, token in enumerate(argv):
1973
1982
  value: "Optional[str]" = None
1974
1983
  if token in ("--output", "-o"):
@@ -1987,6 +1996,36 @@ def _mark_output_delivered() -> None:
1987
1996
  _OUTPUT_DELIVERED = True
1988
1997
 
1989
1998
 
1999
+ def _require_answer_delivery(command: str) -> None:
2000
+ """Require a composed-report command to publish an answer before exit 0.
2001
+
2002
+ A file promise catches only ``-o``. Risk composition also has a stdout
2003
+ contract: a successful invocation must leave one parseable document behind.
2004
+ Keeping that obligation opt-in avoids changing intentionally quiet commands
2005
+ while covering both expensive composition surfaces.
2006
+ """
2007
+ global _ANSWER_REQUIRED_BY
2008
+ _ANSWER_REQUIRED_BY = command
2009
+
2010
+
2011
+ def _fail_on_missing_required_answer(exit_code: object) -> None:
2012
+ """Turn a silent successful composed run into a typed failure (C3-122)."""
2013
+ if exit_code not in (0, None) or _OUTPUT_DELIVERED or not _ANSWER_REQUIRED_BY:
2014
+ return
2015
+ _emit_error_json(
2016
+ EXECUTION_FAILED_CODE,
2017
+ f"{_ANSWER_REQUIRED_BY} finished without publishing an answer.",
2018
+ hint=(
2019
+ "The command returned successfully before its output was delivered. "
2020
+ "Re-run with progress enabled and report the invocation; a composed "
2021
+ "report must produce parseable stdout or its requested file."
2022
+ ),
2023
+ command=_ANSWER_REQUIRED_BY,
2024
+ expected="parseable stdout or the requested output artifact",
2025
+ )
2026
+ raise SystemExit(1)
2027
+
2028
+
1990
2029
  def _fail_on_unkept_output_promise(exit_code: object) -> None:
1991
2030
  """Refuse to report success for an answer that was never produced (C3-122).
1992
2031
 
@@ -2994,13 +3033,59 @@ def _get_command_with_preprocessing(typer_instance: Any) -> Any:
2994
3033
 
2995
3034
  _orig_cmd_main = cmd.main
2996
3035
 
3036
+ _TELEMETRY_COMMANDS = frozenset({
3037
+ "repo-ir", "impact", "endpoints", "export", "validation", "delta", "contract-diff",
3038
+ "plan", "compare", "trend", "timeline", "spring-audit", "verify-edit", "verify", "risk",
3039
+ "enrich", "audit-report", "migrate-recipe", "data-exposure", "posture", "migrate-check",
3040
+ "impact-chain", "pr-impact", "explain", "onboard", "review-pr", "fix-bug", "modernize",
3041
+ "rename-class", "chunk-file", "activate", "selftest", "explain-endpoint", "regress", "version",
3042
+ "schema", "config", "archetype", "cold-start", "telemetry", "mcp", "cache", "baseline",
3043
+ "auth", "retrieve", "prepare-context",
3044
+ })
3045
+
3046
+ def _telemetry_command(args_for_command: Optional[list[str]]) -> str:
3047
+ """Extract only a registered CLI command, never a path or option value."""
3048
+ tokens = list(args_for_command) if args_for_command is not None else sys.argv[1:]
3049
+ for token in tokens:
3050
+ if token in _TELEMETRY_COMMANDS:
3051
+ return token
3052
+ return "analyze"
3053
+
3054
+ def _telemetry_flags(args_for_command: Optional[list[str]]) -> list[str]:
3055
+ tokens = list(args_for_command) if args_for_command is not None else sys.argv[1:]
3056
+ try:
3057
+ from sourcecode.telemetry.filters import _SAFE_FLAGS
3058
+ return sorted({token for token in tokens if token in _SAFE_FLAGS})
3059
+ except Exception:
3060
+ return []
3061
+
2997
3062
  def _cmd_main(args: Optional[list[str]] = None, **kwargs: Any) -> Any:
2998
3063
  if args is not None:
2999
3064
  # CliRunner / programmatic call: preprocess the explicit args list.
3000
3065
  _set_detected_path(".")
3001
3066
  args = _preprocess_args(list(args))
3002
3067
  # args=None → Click reads sys.argv; _preprocess_argv() in main_entry handled it.
3003
- return _orig_cmd_main(args=args, **kwargs)
3068
+ started = time.monotonic()
3069
+ command = _telemetry_command(args)
3070
+ success = True
3071
+ try:
3072
+ return _orig_cmd_main(args=args, **kwargs)
3073
+ except BaseException:
3074
+ success = False
3075
+ raise
3076
+ finally:
3077
+ try:
3078
+ from sourcecode import telemetry as _tel
3079
+ _tel.record(
3080
+ "execution_completed",
3081
+ cmd=("mcp" if command == "mcp" else "telemetry" if command == "telemetry" else "analyze"),
3082
+ command=command,
3083
+ flags=_telemetry_flags(args),
3084
+ duration_s=time.monotonic() - started,
3085
+ success=success,
3086
+ )
3087
+ except Exception:
3088
+ pass
3004
3089
 
3005
3090
  cmd.main = _cmd_main
3006
3091
  return cmd
@@ -3016,7 +3101,7 @@ try:
3016
3101
  except Exception:
3017
3102
  pass
3018
3103
 
3019
- telemetry_app = typer.Typer(help="Manage anonymous telemetry (off by default; opt-in).", rich_markup_mode="rich")
3104
+ telemetry_app = typer.Typer(help="Manage anonymous telemetry (on by default; disable any time).", rich_markup_mode="rich")
3020
3105
  app.add_typer(telemetry_app, name="telemetry")
3021
3106
 
3022
3107
  mcp_app = typer.Typer(help="MCP integration: setup, status, serve, remove.", rich_markup_mode="rich")
@@ -3044,9 +3129,8 @@ app.add_typer(retrieve_app, name="retrieve")
3044
3129
  def _maybe_show_telemetry_notice() -> None:
3045
3130
  """Show first-run telemetry notice once, on interactive TTYs only.
3046
3131
 
3047
- Telemetry is off by default (opt-in). The notice is an invitation, not a
3048
- disclosure: nothing has been collected when it appears. Marked as shown so it
3049
- appears only once.
3132
+ Telemetry is on by default. The notice is a one-time disclosure of what is
3133
+ collected and how to disable it.
3050
3134
  """
3051
3135
  try:
3052
3136
  from sourcecode.telemetry.config import has_been_asked, mark_asked
@@ -5275,26 +5359,7 @@ def main(
5275
5359
  perf.stop("serialize", _perf_serialize)
5276
5360
  perf.flush_recorder()
5277
5361
 
5278
- # 5. Telemetry (fire-and-forget, never blocks)
5279
- try:
5280
- from sourcecode import telemetry as _tel
5281
- _tel.record(
5282
- "execution_completed",
5283
- cmd="analyze",
5284
- flags=_active_flags(
5285
- dependencies, graph_modules, docs, full_metrics,
5286
- semantics, architecture, git_context, env_map,
5287
- code_notes, agent, compact, tree, no_redact, format,
5288
- ),
5289
- output_fmt=format,
5290
- file_count=len(sm.file_paths),
5291
- duration_s=time.monotonic() - _t0,
5292
- success=True,
5293
- )
5294
- except Exception:
5295
- pass
5296
-
5297
- # 6. Write output (CLI-04)
5362
+ # 5. Write output (CLI-04)
5298
5363
  _progress.finish()
5299
5364
  if format == "json":
5300
5365
  _uncommitted_fresh, _uncommitted_fresh_basis = _uncommitted_fact(
@@ -5887,17 +5952,6 @@ def prepare_context_cmd(
5887
5952
  except Exception:
5888
5953
  pass # not JSON (or unreadable) → serve exactly what was stored
5889
5954
  _emit_command_output(_cached_pctx, output_path, copy)
5890
- try:
5891
- from sourcecode import telemetry as _tel
5892
- _tel.record(
5893
- "execution_completed",
5894
- cmd="prepare-context",
5895
- feature=task,
5896
- output_fmt=format,
5897
- duration_s=0.0,
5898
- )
5899
- except Exception:
5900
- pass
5901
5955
  return
5902
5956
 
5903
5957
  _scope_files = None
@@ -6424,18 +6478,6 @@ def prepare_context_cmd(
6424
6478
  fmt=("yaml" if format == "yaml" else "json")
6425
6479
  )
6426
6480
 
6427
- try:
6428
- from sourcecode import telemetry as _tel
6429
- _tel.record(
6430
- "execution_completed",
6431
- cmd="prepare-context",
6432
- feature=task,
6433
- output_fmt=format,
6434
- duration_s=_time.perf_counter() - _t0,
6435
- )
6436
- except Exception:
6437
- pass
6438
-
6439
6481
  from sourcecode.mcp_nudge import nudge_mcp_if_needed as _nudge
6440
6482
  _nudge()
6441
6483
 
@@ -6452,11 +6494,11 @@ def telemetry_status(
6452
6494
  enabled = is_enabled()
6453
6495
  choice = stored_choice()
6454
6496
  status = "enabled" if enabled else "disabled"
6455
- lines = [f"Telemetry: {status} (off by default; opt-in)"]
6456
- # "off because you said so" and "off because nobody asked you" are different
6457
- # answers, and a buyer auditing this needs to be told which one they have.
6497
+ lines = [f"Telemetry: {status} (on by default; disable with `ask telemetry disable`)"]
6498
+ # The default and an explicit disabled choice are different answers, and a
6499
+ # buyer auditing this needs to be told which one they have.
6458
6500
  if choice is None:
6459
- lines.append(" No choice recorded — nothing has been collected or sent.")
6501
+ lines.append(" No choice recorded — the on-by-default setting applies.")
6460
6502
  else:
6461
6503
  lines.append(f" Your recorded choice: {'enabled' if choice else 'disabled'}.")
6462
6504
  lines.append(f" Config: {config_file_path()}")
@@ -6469,7 +6511,7 @@ def telemetry_status(
6469
6511
 
6470
6512
  @telemetry_app.command("enable")
6471
6513
  def telemetry_enable() -> None:
6472
- """Opt in to anonymous telemetry."""
6514
+ """Enable anonymous telemetry and remember the choice."""
6473
6515
  from sourcecode.telemetry.config import set_enabled
6474
6516
  from sourcecode import telemetry as _tel
6475
6517
  set_enabled(True)
@@ -6482,11 +6524,11 @@ def telemetry_enable() -> None:
6482
6524
 
6483
6525
  @telemetry_app.command("disable")
6484
6526
  def telemetry_disable() -> None:
6485
- """Opt out of anonymous telemetry."""
6527
+ """Disable anonymous telemetry and remember the choice."""
6486
6528
  from sourcecode.telemetry.config import set_enabled
6487
6529
  set_enabled(False)
6488
6530
  typer.echo("Telemetry disabled. No data will be collected or sent.")
6489
- typer.echo("Telemetry is off by default; this choice is recorded so the notice stops asking.")
6531
+ typer.echo("Telemetry is enabled by default; this choice is recorded so the notice stops appearing.")
6490
6532
  typer.echo("Re-enable at any time: ask telemetry enable")
6491
6533
 
6492
6534
 
@@ -7946,11 +7988,22 @@ def validation_cmd(
7946
7988
  target = _admit_path(path)
7947
7989
 
7948
7990
  from sourcecode.context_graph import ContextGraph
7991
+ from sourcecode.repository_ir import find_java_files
7949
7992
  from sourcecode.validation_surface import build_validation_surface
7950
7993
  _prog = Progress()
7951
7994
  _prog.start("mapping validation surface")
7952
- # Structural facts come from the ContextGraph — the single access layer.
7953
- _graph = ContextGraph.build_from_root(target)
7995
+ # Structural facts still come from the ContextGraph — the single access
7996
+ # layer — but its CIR now follows the shared knowledge-cache door used by
7997
+ # the other repository-wide commands. A validation run therefore does not
7998
+ # parse the same tree again after `cache warm`, `impact`, or `explain`.
7999
+ # `--no-cache` remains the published no-op for subcommands; the direct build
8000
+ # is strictly the best-effort fallback if the shared layer cannot answer.
8001
+ try:
8002
+ from sourcecode.context_cache import shared_cir
8003
+
8004
+ _graph = ContextGraph.from_cir(shared_cir(target, find_java_files(target)))
8005
+ except Exception:
8006
+ _graph = ContextGraph.build_from_root(target)
7954
8007
  data = build_validation_surface(target, graph=_graph)
7955
8008
 
7956
8009
  # P1-C: classify the request-body validation *pattern* so a repo with a
@@ -10666,6 +10719,7 @@ def risk_cmd(
10666
10719
  ask risk . --table --rule SEC-008 --band high --top-n 20
10667
10720
  ask risk . --limit 10 -o risk.json
10668
10721
  """
10722
+ _require_answer_delivery("risk")
10669
10723
  _apply_jobs(jobs)
10670
10724
  from sourcecode.risk import build_risk
10671
10725
 
@@ -11023,6 +11077,7 @@ def audit_report_cmd(
11023
11077
  ask audit-report . --profile prod --format markdown
11024
11078
  ask audit-report . --sign-key audit.key -o audit-report.json
11025
11079
  """
11080
+ _require_answer_delivery("audit-report")
11026
11081
  from sourcecode.audit_report import build_audit_report, render_markdown
11027
11082
 
11028
11083
  path = _admit_path(path)
@@ -14138,7 +14193,7 @@ def config_cmd(
14138
14193
  _answer = _TextAnswer(output_path)
14139
14194
  _answer.say(f"ask {__version__}")
14140
14195
  _answer.say(f"Config: {config_file_path()}")
14141
- _answer.say(f"Telemetry: {'enabled' if is_enabled() else 'disabled'} (off by default; opt-in)")
14196
+ _answer.say(f"Telemetry: {'enabled' if is_enabled() else 'disabled'} (on by default; disable with `ask telemetry disable`)")
14142
14197
  _answer.say("")
14143
14198
 
14144
14199
  # F-AX: the declaration may sit outside the analysed tree, and *which* file
@@ -16406,8 +16461,10 @@ def main_entry() -> None:
16406
16461
  try:
16407
16462
  # prog_name pins usage/help to the canonical `ask`, whatever the alias.
16408
16463
  app(prog_name="ask")
16464
+ _fail_on_missing_required_answer(0)
16409
16465
  _fail_on_unkept_output_promise(0)
16410
16466
  except SystemExit as _sysexit:
16467
+ _fail_on_missing_required_answer(_sysexit.code)
16411
16468
  _fail_on_unkept_output_promise(_sysexit.code)
16412
16469
  raise
16413
16470
  except WriteRefused as _refused:
@@ -253,10 +253,30 @@ def build_data_exposure(
253
253
  rows: dict[str, dict] = {}
254
254
  seeds_report: list[dict] = []
255
255
 
256
+ # Every seed answers a different reachability question over the same CIR.
257
+ # `run_impact_chain` accepts a prepared model precisely so that a declaration
258
+ # with several types does not rebuild the repository-wide semantic indexes
259
+ # once per type. Build lazily: declarations whose labels do not resolve keep
260
+ # the existing failure isolation in `run_impact_chain`.
261
+ impact_model = model
262
+
256
263
  for label in decl.labels:
257
264
  for seed in label.seed_types:
258
265
  found_signature = _signature_routes(graph, seed, handlers)
259
- chain = run_impact_chain(cir, seed, root=root, model=model, depth=depth)
266
+ if impact_model is None:
267
+ try:
268
+ from sourcecode.spring_model import SpringSemanticModel
269
+
270
+ impact_model = SpringSemanticModel.build(cir)
271
+ except Exception:
272
+ # `run_impact_chain` owns its error envelope. Retaining None
273
+ # here preserves that per-seed fallback if the model cannot be
274
+ # assembled instead of turning a partial exposure answer into
275
+ # a command failure.
276
+ pass
277
+ chain = run_impact_chain(
278
+ cir, seed, root=root, model=impact_model, depth=depth,
279
+ )
260
280
  resolution = str(getattr(chain, "resolution", "not_found"))
261
281
  reached = list(getattr(chain, "endpoints_affected", []) or [])
262
282
 
@@ -220,7 +220,7 @@ class DispositionSummary:
220
220
  return summary
221
221
 
222
222
  def to_dict(self) -> dict:
223
- return {
223
+ out = {
224
224
  "affected": self.affected,
225
225
  "unknown": self.unknown,
226
226
  "unaffected": self.unaffected,
@@ -230,6 +230,11 @@ class DispositionSummary:
230
230
  "three counts sum to `total`, so the negative set is published "
231
231
  "rather than left to subtraction (P0-3)"
232
232
  ),
233
- "unaffected_basis": UNAFFECTED_BASIS,
234
233
  "blockers": [b.to_dict() for b in self.blockers],
235
234
  }
235
+ # A basis is evidence for the published negative set, not boilerplate.
236
+ # When no route is excluded (or a blocker demotes that whole set), showing
237
+ # it reads like a certification the report did not earn.
238
+ if self.unaffected and not self.blockers:
239
+ out["unaffected_basis"] = UNAFFECTED_BASIS
240
+ return out
@@ -1108,7 +1108,14 @@ def analyze_hibernate(file_paths: list[str], root: Path) -> HibernateStratificat
1108
1108
  # right (it is the Spring Boot default) and was wrong in the field case that
1109
1109
  # billed 45.6 rewrite days to an EclipseLink repository.
1110
1110
  hibernate_evidence = dep_present or any_hibernate_import
1111
- transitive_starter_evidence = jpa_starter_present and not hibernate_evidence
1111
+ # Spring Boot defaults the JPA starter to Hibernate, but an explicit provider
1112
+ # declaration overrides that default. Treating both as evidence made an
1113
+ # EclipseLink project inherit a fictional Boot-2 Hibernate-5 rewrite axis.
1114
+ transitive_starter_evidence = (
1115
+ jpa_starter_present
1116
+ and not hibernate_evidence
1117
+ and not competing_providers
1118
+ )
1112
1119
  inferred_from_jpa_only = any_jpa_import and not hibernate_evidence and not transitive_starter_evidence
1113
1120
  # Another declared provider refutes the inference outright. It cannot refute
1114
1121
  # direct evidence: a repository may legitimately carry both.
sourcecode/mcp/runner.py CHANGED
@@ -7,11 +7,15 @@ lookup, no process fork, no stdout encoding issues.
7
7
  from __future__ import annotations
8
8
 
9
9
  import json
10
+ import threading
10
11
  from typing import Any
11
12
 
12
13
  from typer.testing import CliRunner
13
14
 
14
15
  _runner = CliRunner()
16
+ # CliRunner temporarily replaces process-global stdout/stderr. MCP dispatches
17
+ # tools concurrently, so every in-process invocation must share this lock.
18
+ _runner_lock = threading.RLock()
15
19
 
16
20
 
17
21
  class CommandError(RuntimeError):
@@ -44,7 +48,8 @@ def run_command(args: list[str]) -> Any:
44
48
  # Pass raw args to invoke — the _cmd_main hook inside cli.py handles path
45
49
  # extraction via _preprocess_args. Pre-processing here would strip the path
46
50
  # from args, then _cmd_main would re-process the stripped list and lose it.
47
- result = _runner.invoke(app, list(args))
51
+ with _runner_lock:
52
+ result = _runner.invoke(app, list(args))
48
53
 
49
54
  if result.exit_code != 0:
50
55
  stdout_raw = getattr(result, "output", "")
sourcecode/mcp/server.py CHANGED
@@ -51,9 +51,8 @@ def _record_tool_invocation(name: Any, success: bool, started: float) -> None:
51
51
 
52
52
  Aggregate only: which of our own tools ran, whether it succeeded, and a
53
53
  duration bucket. Never the arguments — those carry repository paths — and
54
- never any result content. Honours the same opt-in as every other event
55
- (off until `ask telemetry enable` or SOURCECODE_TELEMETRY=1); when
56
- telemetry is off, `record` returns before building anything.
54
+ never any result content. Honours the same default-on setting as every other
55
+ event; `ask telemetry disable` or the documented environment controls stop it.
57
56
  """
58
57
  try:
59
58
  import time as _time
@@ -1480,7 +1479,7 @@ def telemetry(action: str) -> dict:
1480
1479
  """Manage telemetry settings.
1481
1480
 
1482
1481
  Maps to: ask telemetry <action>
1483
- action: one of "status" (show current state), "enable" (opt in), "disable" (opt out).
1482
+ action: one of "status" (show current state), "enable", or "disable".
1484
1483
  Valid values: "status" | "enable" | "disable"
1485
1484
  """
1486
1485
  # FIX-P2-10: enumerate valid actions in docstring so agents don't guess.