sourcecode 5.8.1__py3-none-any.whl → 5.8.2__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
sourcecode/__init__.py CHANGED
@@ -4,4 +4,4 @@ ASK Engine is the product. ``ask`` is the canonical CLI command; ``sourcecode``
4
4
  the legacy compatibility alias and the Python/PyPI package name. See
5
5
  docs/PRODUCT_IDENTITY.md (normative)."""
6
6
 
7
- __version__ = "5.8.1"
7
+ __version__ = "5.8.2"
@@ -45,6 +45,11 @@ The facts these rows are keyed to are published: `ask schema facts-v1` prints th
45
45
  | ASK-12 | **A field describing an exclusion named three production trees as excluded.** `spring-audit spring-boot` published `patterns_matched: [… "core/", "module/", "cli/" …]` — 4 254 + 1 625 + 76 production files. Read literally, whole modules had left the audit population; measured, the exclusion was *conservative* (3 650 excluded against 4 214 files actually under a test root). The auditor filed an over-exclusion false positive on the strength of the field and then withdrew it. Cause: the constructor's fallback named `parts[0]`, the first path segment, when no source-root marker matched — and the same published field had **two constructors** with different shapes and different `basis` strings (`cli.py` for `spring-audit`, `repository_ir.py` for `endpoints`) over the one authority both already shared. | 5.7.2 audits | Low — a reporting defect, not an exclusion one, which is why it cost an auditor a round instead of costing a finding | **closed 5.8.1** — one producer, `repository_ir.excluded_test_source_detail`, beside `is_test_source_file`, which is the authority that decides the fact. `patterns_matched` can now only hold a **declared test root** — a directory whose files are excluded in their entirety — and everything else is published as a `rules_matched` census with a `rules_meaning` sentence per rule, so a per-file rule can no longer be read as a whole-directory claim. A surface counting something other than files says so (`unit`, `files_matched`), because `endpoints` counts route annotations. Regression `tests/test_test_source_exclusion_one_producer_ask12.py`, 14 assertions, two of them structural: `cli.py` may not compose the key at all, and exactly one function in `repository_ir.py` may (AST-checked). |
46
46
  | B14 | **`--compact` made the answer smaller by making it unattributable.** `cold-start --compact` took 1 968 807 characters to 16 452 (−99,2 %) and removed **11 of 19 keys** on the way — `schema_version`, `repo_id`, `current_git_head` and the 331-character `hotspots` among them — in a document whose weight was **97,7 % `endpoints`**, and cut the endpoint list to 30 with no `*_cap`. The view an agent is told to inject could not say which repository, which commit or which schema it described. | 5.6.1 audit; AS-15 closed the root `--compact` half in 5.7.0, this is the `cold-start` half, re-reported against 5.8.0 | High | **closed 5.8.1** — three rules, in the shared bound rather than invented per command. `COMPACT_IMMUTABLE_KEYS` is the minimum no `--compact` may remove (provenance is contract, not content). Every collection that was cut declares `{total, shown, omitted}` with the registered effect — measured on BroadleafCommerce, `endpoints_cap` now publishes `total: 272, shown: 30, omitted: 242` where the cut was silent. And trimming goes by **weight**: a section already under 2 KB is kept rather than dropped by a fixed list, which is what recovers `hotspots`; everything actually removed is named with its size in `compact_omitted`. Regression `tests/test_compact_keeps_identity_b14.py`, 12 assertions, including that nothing removed claims nothing. |
47
47
  | B2 | **A risk band and a rule severity share a vocabulary and measure different things.** They disagreed on **34 of 50 rows (68 %)** — 5 defects rated `medium` band as `critical`, 29 rated `high` band as `medium` — and both numbers were right: severity is what the rule says about the defect, the band is what the composition says about it in *this* repository (reach, access, write effect). One surface on, the collision became a contradiction a reader could not resolve: `audit-report` announced **14 critical** over bands while `--min-severity critical` filtered severities and returned **0**. | 5.6.1 audit, reproduced through 5.8.0 | Medium | **closed 5.8.1** — the band now also has a name that cannot be mistaken for a severity: `risk_tier: R1…R5`, published **beside** `band` and never instead of it, so no consumer loses a field it reads. One mapping (`risk_model.tier`), one sentence (`TIER_CROSSWALK_NOTE`) travelling with every count — it names both axes, states that they disagree by construction, gives the crosswalk, and points at `--min-band` as the flag that filters tiers rather than `--min-severity`. `risk` publishes `by_tier` beside `by_band`; `audit-report` leads with tiers, keeps bands under the legacy spelling, renders the note, and derives tiers from bands when it is handed an older bundle. Regression `tests/test_risk_tier_is_not_a_severity_b2.py`, 12 assertions, including that no tier shares a spelling with any severity and that a summary with no census still says *not measured* rather than zero. ⚠ The row's `reach`/`access`/`score_raw` per row was already shipped (`factors`, `endpoints_reached`, `reach_scope`); `--min-band` already existed on `risk` and `enrich`. |
48
+ | B11 | **A boolean about the working tree was published without the unit it was measured in.** The field read `has_uncommitted_changes: false` in the same terminal where `git status --porcelain` had just printed `?? .claude/settings.json`, with no `_basis` and no unit beside it. The rule is right and stays: C1-27 decided that a change under an editor or agent state directory is not a change to the code, because no analyser opens those files, and counting it reported STALE on a snapshot that described the tree exactly with `Delta: 0` next to it. What the payload could not do was *say* it — with no basis, "no changes" and "changes I did not count" are the same three characters, and the only way to tell them apart was to read the source. | 5.6.1 audit, inherited | Low — a disclosure defect over a correct rule, which is why it cost a reader's trust rather than a wrong answer | **closed 5.8.1** — the boolean now travels with the census that explains it. One producer, `baseline_autocapture.worktree_dirty_detail`: one `git status --porcelain -uall`, every path it prints classified once by `dirty_exclusion_reason`, and `{dirty, counted, excluded, excluded_by, basis}` out. `worktree_dirty` is that dict read as a boolean and `counts_as_dirty` is the same rule read as a predicate, so the three cannot drift — the C1-44 shape, kept. The `basis` names the command a reader can check it against and reconciles with it: `counted + excluded` is exactly what `git status` printed (asserted), and each exclusion carries its sentence from `EXCLUSION_REASONS` rather than a bare key. **`false` and "not checked" stopped being the same answer**: a non-git directory publishes *"not a git worktree: the working tree could not be compared to HEAD"* instead of a confident `false`, and a run whose own scan measured the dirtiness says which source answered. Published as `has_uncommitted_changes_basis` at every site that prints the boolean — `cold-start` and its `_cache` block (both CLI paths through `_cache_meta_block`, plus the RIS bootstrap), and the MCP `check_freshness` tool with its docstring corrected. Regression `tests/test_uncommitted_changes_basis_b11.py`, 14 assertions, including the reported pair end-to-end and the reconciliation against `git status`. |
49
+ | P1 / F-BR | **Every comparison the release gate made was a ratio, and a cost that was always too high never regresses.** `validation` went 13,8 → 34,1 → 39,5 s across three releases with no single step over the 20 % threshold, so the gate was green at every step and the product shipped a staircase. The row also carried two prescribed cache-reuse assertions that the field **refuted with its own measurements** before either was implemented: `samples[0]/samples[-1] > 1,5` reads only the ends of the series and passes 5.6.1, and `median(samples)/min(samples) < 1,5` scores shape rather than cost and therefore awards **5.5.5, the worst release of the six, a perfect mark** — flat-and-slow (nothing cached, 80 s) and flat-and-fast (everything cached, 30 s) are the same ratio. | 5.7.2 audits (`P1`, `B7`), with the third formulation validated by the field against six measured releases | Medium — the detector gap that let a ×15,7 regression ship (C3-120), compounded by a documented instruction to build the wrong detector | **closed 5.8.1** — a second axis, `perf.absolute_cost_report`, published under `absolute_cost` in the gate report and able to turn the verdict red on its own. Two assertions: a **declared ceiling per command** (`RELEASE_GATE_ABSOLUTE_WARM_MS`, each published beside the mark it was calibrated against at 5.7.2, because a limit nobody can audit against a measurement is just a number), and the **anchored** one the field validated — `max(sample) < 8 × posture_warm`, with the anchor measured *in the same run* so a slower runner moves the limit with it and the number stays a claim about the product rather than about the machine CI allocated. Scope is enforced rather than assumed: a ceiling calibrated on ~3 300 files takes no part on another corpus, and cold cells still do not decide (C3-108). **The two refuted formulations are carried in the source** (`REFUTED_REUSE_FORMULATIONS`) with the counter-example that killed each, and the battery replays both refutations plus the six-release validation table (`ANCHOR_MULTIPLE_VALIDATION`), so a later change to the multiple cannot quietly stop discriminating — the multiple is asserted to sit between the worst pass (6,2×) and the best fail (8,5×). `--self-test` gained two reds, because a green gate is evidence only when a red one is reachable on the same constants. Regression `tests/test_absolute_cost_gate_p1.py`, 25 assertions, including the staircase end-to-end: ratio axis green, release red. ⚠ The rest of F-BR — a synthetic ~200 KB / ~5 000-call fixture timing `build_sql_taint`, and warm cells for the composition family in the gate baseline — is **not** in this row and stays open; and C3-113 still applies, since the gate's release path is triggered by a tag and no release since `v5.2.0` carries one. |
50
+ | ASK-11 | **A correct decision announced only by the error message of the run it broke.** The unified output ceiling turned `exit 0` into `exit 1` for invocations that returned data one minor earlier — `endpoints tutorials` (987 141 B ≈ 247 K tokens), `endpoints thingsboard` (318 380 B), `posture spring-boot` (355 032 B on 8 667 files). The decision is not in dispute and is not being revisited: 247 K tokens printed into an agent's context window was worse than a refusal. The defect is the **mode of disclosure** — the trigger is repository size rather than a flag, so an invocation that fits today stops fitting on a bigger repository or after a release that detects more, and the failure lands as `exit 1`, which stops a CI build. | 5.7.2 audits — filed by the field as *"cambio incompatible, no defecto"* | Medium — a release-policy defect: nothing was wrong with the run, only with what could be known before it | **closed 5.8.1** — treated as the policy it is, not as a changelog line. `sourcecode.breaking_changes` declares the change with what a reader planning an upgrade actually needs: the exit code on both sides, the trigger, the **compatibility route** (`ASK_MAX_OUTPUT_TOKENS=0`, which restores pre-5.7.0 behaviour exactly, for the pipeline that is red right now) and the **remedy** (a bounding flag, or `--output FILE`, for when there is time) — in that order, because an upgrade note offering only the principled fix is one a reader skips. "A large repository" is given its sizes: the three field invocations are carried with their measured bytes. Three surfaces, all derived from that one declaration: `ask schema breaking-changes-v1` prints it, a `## Upgrading` section sits above the release history in the CHANGELOG, and **every affected command names the ceiling in its `--help`** — the population derived from the CLI registry (all commands minus `COMMANDS_WITHOUT_COPY`, the existing authority for "stdout is a confirmation, not an answer"; a confirmation cannot reach a 50K-token ceiling), never a hand-written list, since the list is exactly what a reader is told to trust. The battery enforces the **policy**: a declared change must move an exit code, carry both routes, name its invocations with measurements, and reach the changelog. Regression `tests/test_breaking_change_disclosure_ask11.py`, 95 assertions (one per affected command, parametrised from the live registry). ⚠ Found and fixed on the way: `version` published `registries: ["facts-v1"]` as a literal while `schema` derived its list, under a comment promising the two could not disagree — the second registry made them disagree. Both read `published_registries()` now. |
51
+ | ASK-13 | **A floor that could be anything, in the sentence whose job is to size the work.** `cache warm tutorials` reported *"This repository: at least 8 418 Java files, cache warm"* about a repository holding **24 073** — a third of the truth — and *"at least 8 401"* on `spring-boot` against 8 667. The two land near each other for one reason: the walk stops after 25 000 **directory entries**, not after 25 000 Java files, so a tree with ordinary amounts of non-Java content exhausts the budget at roughly the same place however much Java it holds. `at least` was honest about the direction of the error and silent about its size, which is exactly what made it useless here. | 5.7.2 audits | Low — the count is a floor and says so; what it could not say was how far a floor it is | **closed 5.8.1** — the second of the two outcomes the acceptance allows: the cap stays (this count runs while somebody waits for `--help` to render) and **names itself**. The message now reads *"at least 8 418 Java files (the scan stopped at 25 000 directory entries, so the repository holds more)"*, with the unit named, because "entries" and "Java files" being different units is the whole reason the number saturates where it does. A cap that did **not** bind is not mentioned at all — it is not part of that count's basis, and naming it would suggest it was. ⚠ **A second half that is not a message, found while measuring the first**: `Scope.large` answered `False` from a saturated count, so a repository whose first 25 000 entries held few sources was advised as *small* — and the class had already written the correct rule in prose (*"a capped count below it is not [enough]"*) while the code ignored it. `size_decided` now separates the two cases a floor supports: above the threshold a capped count still decides (more files only make `large` more true), below it decides nothing, and the cost verdict answers `unknown` with its reason rather than guessing the cheap regime. Also fixed on the way: `count_java_files` takes its cap as a **default argument**, bound at definition time, so `measure_scope` now passes it explicitly — otherwise the cap that bound the walk and the cap the message names could drift into a published falsehood. Regression `tests/test_scan_cap_names_itself_ask13.py`, 13 assertions, including a tree shaped like the field's (many entries, few sources). |
52
+ | B6 | **Fourteen `schema_version` values in payloads and not one of them resolvable.** `ask schema verify-v1` answered *invalid* about a string this tool had just printed; `endpoints` emitted no version at all; and the spellings did not agree — `<domain>-v1` beside a bare `1.0`, a `1.1`, a `1.2`, a `1.0.0`, a `2.1` and a `0.1-experimental`. A version identifier a consumer cannot look up is a string, not a contract. | 5.6.1 audits, inherited | Medium — every consumer that pins a shape pins it on one of these | **closed 5.8.1** — `sourcecode.schema_registry` names **every** `schema_version` this release can emit (50 entries over 44 distinct identifiers, 13 of them non-conforming) with what it versions, which module defines it, which commands carry it, and whether a JSON Schema document exists behind it. `ask schema <value>` resolves all of them: a documented identifier still prints its schema, and one without a document resolves to its entry rather than to a rejection — which is what it honestly is, a named versioned shape with no formal document yet, said plainly instead of implied away. `endpoints` gained `endpoints-v1`, the version it never had. **The population is checked against the source, not maintained by hand**: the battery walks `src/sourcecode` by AST and fails when a `schema_version` literal or a `*SCHEMA*` constant is not named by the registry, and it also verifies the registry's own claims — the module and constant each entry cites must really hold that value (this caught one wrong `source` while it was being written). **The finding the field could not have made from outside**: `1.0` is not a schema. **Seven** unrelated shapes spell their version that way — the root analysis, `impact-chain`, `pr-impact`, migration blast, Spring impact, event topology and test-gap ranking — so `resolve()` returns all seven and publishes the identifier as ambiguous rather than picking one, and every ambiguous entry must offer a canonical name to pin instead (asserted). **The `<domain>-v<N>` migration is declared, not applied**: renaming an emitted `schema_version` changes a payload consumers pin on, which is precisely the incompatible change `breaking-changes-v1` exists to stop us making in silence, so each non-conforming identifier carries the name it should become and the battery asserts none of them was quietly renamed in this release. `schema --help` documents the mapping and the `1.0` trap. Regression `tests/test_schema_registry_resolves_b6.py`, 66 assertions. ⚠ Writing the 40-odd JSON Schema documents themselves is **not** in this row and stays open; what closed is resolvability, which is what the row was filed for. |
48
53
 
49
54
  ---
50
55
 
@@ -136,7 +136,7 @@ pipx install sourcecode # isolated install, no venv needed
136
136
 
137
137
  # Verify
138
138
  ask version
139
- # ask 5.8.1
139
+ # ask 5.8.2
140
140
  ```
141
141
 
142
142
  Requires Python 3.9+.
@@ -36,7 +36,7 @@ import os
36
36
  import subprocess
37
37
  from dataclasses import dataclass
38
38
  from pathlib import Path
39
- from typing import TYPE_CHECKING, Iterable, Optional
39
+ from typing import TYPE_CHECKING, Any, Iterable, Optional
40
40
 
41
41
  if TYPE_CHECKING: # pragma: no cover - typing only
42
42
  from sourcecode.canonical_ir import CanonicalRepositoryIR
@@ -198,6 +198,62 @@ def worktree_dirty(
198
198
 
199
199
  Nothing else is excused, because everything else is what the baseline measures.
200
200
  """
201
+ detail = worktree_dirty_detail(root, ignore=ignore)
202
+ return None if detail is None else detail["dirty"]
203
+
204
+
205
+ #: The reason a path git reports as changed does not make the tree dirty. Keys are
206
+ #: stable (they are published); the sentences are what a reader is owed.
207
+ EXCLUSION_REASONS: "dict[str, str]" = {
208
+ "tool-state-directory": (
209
+ "under an editor or agent state directory (.claude/, .idea/, .vscode/, …) "
210
+ "that no analyser opens"
211
+ ),
212
+ "own-footprint": "this tool's own output, which the run wrote itself",
213
+ }
214
+
215
+
216
+ def dirty_exclusion_reason(
217
+ path: str, *, own_prefixes: "Iterable[str]" = ()
218
+ ) -> "Optional[str]":
219
+ """Why *path* does not count as a modification, or None when it does count.
220
+
221
+ The one place the rule lives. `counts_as_dirty` is this predicate read as a
222
+ boolean, and `worktree_dirty_detail` is it read as a census — B11 was filed
223
+ because the boolean was published with neither its unit nor this reason, so
224
+ `has_uncommitted_changes: false` sat beside a `?? .claude/settings.json` that
225
+ `git status` had just printed and the payload could not explain the pair.
226
+ """
227
+ from sourcecode.path_filters import analysis_reads
228
+
229
+ if not analysis_reads(path):
230
+ return "tool-state-directory"
231
+ probe = str(path).replace("\\", "/").rstrip("/") + "/"
232
+ if any(probe.startswith(prefix) for prefix in own_prefixes):
233
+ return "own-footprint"
234
+ return None
235
+
236
+
237
+ def worktree_dirty_detail(
238
+ root: Path, ignore: "Optional[Path | Iterable[Path]]" = None
239
+ ) -> "Optional[dict[str, Any]]":
240
+ """The boolean `worktree_dirty` publishes, with the census that explains it.
241
+
242
+ One `git status` call, every path it prints classified once. Returns None when
243
+ the state could not be determined, which is the same contract as the boolean.
244
+
245
+ C1-27. "Dirty" means the tree *the analysis reads* differs from HEAD. A change
246
+ under an editor or agent state directory is not one — no analyser opens those
247
+ files — and counting it reported STALE on a snapshot that described the tree
248
+ exactly, with `Delta: 0` beside it. One authority answers which paths qualify
249
+ (`path_filters.analysis_reads`) and its default is inclusive: anything not
250
+ named there still counts, because a missed change is a false *fresh* and a
251
+ spurious one only costs a rebuild.
252
+
253
+ B11: that rule is right and was invisible. The `excluded` census and `basis`
254
+ below are what makes `false` checkable against the reader's own `git status`
255
+ without running the analyser — the count is named, and so is the reason.
256
+ """
201
257
  try:
202
258
  out = subprocess.run(
203
259
  # `-uall` because git collapses an untracked directory to one entry
@@ -210,13 +266,6 @@ def worktree_dirty(
210
266
  return None
211
267
  if out.returncode != 0:
212
268
  return None
213
- # C1-27. "Dirty" means the tree *the analysis reads* differs from HEAD. A
214
- # change under an editor or agent state directory is not one — no analyser
215
- # opens those files — and counting it reported STALE on a snapshot that
216
- # described the tree exactly, with `Delta: 0` beside it. One authority answers
217
- # which paths qualify (`path_filters.analysis_reads`) and its default is
218
- # inclusive: anything not named there still counts, because a missed change
219
- # is a false *fresh* and a spurious one only costs a rebuild.
220
269
  prefixes: list[str] = []
221
270
  if ignore is not None:
222
271
  candidates = [ignore] if isinstance(ignore, (str, Path)) else list(ignore)
@@ -229,10 +278,45 @@ def worktree_dirty(
229
278
  except ValueError: # outside the repo (`--dir`): nothing of it to drop
230
279
  continue
231
280
  prefixes.append(rel.rstrip("/") + "/")
232
- return any(
233
- counts_as_dirty(_porcelain_path(ln), own_prefixes=prefixes)
234
- for ln in out.stdout.splitlines() if ln.strip()
281
+
282
+ counted = 0
283
+ excluded_by: "dict[str, int]" = {}
284
+ for line in out.stdout.splitlines():
285
+ if not line.strip():
286
+ continue
287
+ reason = dirty_exclusion_reason(_porcelain_path(line), own_prefixes=prefixes)
288
+ if reason is None:
289
+ counted += 1
290
+ else:
291
+ excluded_by[reason] = excluded_by.get(reason, 0) + 1
292
+
293
+ excluded = sum(excluded_by.values())
294
+ return {
295
+ "dirty": counted > 0,
296
+ "counted": counted,
297
+ "excluded": excluded,
298
+ "excluded_by": excluded_by,
299
+ "basis": _dirty_basis(counted, excluded_by),
300
+ }
301
+
302
+
303
+ def _dirty_basis(counted: int, excluded_by: "dict[str, int]") -> str:
304
+ """The sentence that names the unit of the boolean, and the paths it dropped.
305
+
306
+ Always names the unit, because "0 excluded" and "the field means every path
307
+ git prints" are different claims and only one of them is true.
308
+ """
309
+ unit = (
310
+ f"{counted} of {counted + sum(excluded_by.values())} path(s) reported by "
311
+ "`git status --porcelain -uall` are changes the analysis reads"
312
+ )
313
+ if not excluded_by:
314
+ return unit + "; none were excluded."
315
+ dropped = "; ".join(
316
+ f"{n} {EXCLUSION_REASONS.get(reason, reason)}"
317
+ for reason, n in sorted(excluded_by.items())
235
318
  )
319
+ return f"{unit}. Excluded: {dropped}."
236
320
 
237
321
 
238
322
  def counts_as_dirty(path: str, *, own_prefixes: "Iterable[str]" = ()) -> bool:
@@ -248,12 +332,7 @@ def counts_as_dirty(path: str, *, own_prefixes: "Iterable[str]" = ()) -> bool:
248
332
 
249
333
  `own_prefixes` are repository-relative directory prefixes ending in `/`.
250
334
  """
251
- from sourcecode.path_filters import analysis_reads
252
-
253
- if not analysis_reads(path):
254
- return False
255
- probe = str(path).replace("\\", "/").rstrip("/") + "/"
256
- return not any(probe.startswith(prefix) for prefix in own_prefixes)
335
+ return dirty_exclusion_reason(path, own_prefixes=own_prefixes) is None
257
336
 
258
337
 
259
338
  def own_footprint_prefixes(root: Path) -> "list[str]":
@@ -0,0 +1,163 @@
1
+ """Changes that turn `exit 0` into `exit 1`, declared where a reader can find them.
2
+
3
+ ASK-11. The unified output ceiling was the right call — 247 K tokens of silence
4
+ was worse than a refusal — but it changed the exit code of invocations that had
5
+ returned data one minor earlier, and `exit 1` in CI stops a build. A correct
6
+ decision announced only in a changelog line is still a script that breaks at 3am
7
+ on somebody else's pipeline.
8
+
9
+ The rule this module exists to enforce is a policy, not a fix:
10
+
11
+ a change that turns `exit 0` into `exit 1` is a **breaking change**, and it
12
+ is discoverable *before* it fires — from `--help`, from a command, and from
13
+ the release notes — never only from the error message of the run it broke.
14
+
15
+ That is why the population is derived rather than listed. A hand-written
16
+ catalogue of affected commands beside a registry drifts the moment a command is
17
+ added, and a reader who trusts it is worse off than one who trusts nothing.
18
+ """
19
+
20
+ from __future__ import annotations
21
+
22
+ from dataclasses import dataclass, field
23
+
24
+
25
+ @dataclass(frozen=True)
26
+ class BreakingChange:
27
+ """One change in behaviour a caller can be broken by, and the way out.
28
+
29
+ `compatibility` is the switch that restores the old behaviour *now*, for
30
+ somebody whose pipeline is red while they read this. `remedy` is what to do
31
+ instead, once there is time to do it properly. Both, in that order: an
32
+ upgrade note that only offers the principled fix is one a reader skips.
33
+ """
34
+
35
+ id: str
36
+ released_in: str
37
+ title: str
38
+ what_changed: str
39
+ exit_code_before: int
40
+ exit_code_after: int
41
+ trigger: str
42
+ compatibility: str
43
+ remedy: str
44
+ #: What the field actually measured, so "large repository" has a size.
45
+ observed: tuple[tuple[str, str], ...] = field(default_factory=tuple)
46
+
47
+ def as_payload(self) -> "dict[str, object]":
48
+ return {
49
+ "id": self.id,
50
+ "released_in": self.released_in,
51
+ "title": self.title,
52
+ "what_changed": self.what_changed,
53
+ "exit_code_before": self.exit_code_before,
54
+ "exit_code_after": self.exit_code_after,
55
+ "trigger": self.trigger,
56
+ "compatibility": self.compatibility,
57
+ "remedy": self.remedy,
58
+ "observed": [
59
+ {"invocation": inv, "measured": measured}
60
+ for inv, measured in self.observed
61
+ ],
62
+ }
63
+
64
+
65
+ #: The output ceiling. Filed by the field as ASK-11 against 5.7.0, which is the
66
+ #: release that unified the guard; the report upgraded from 5.6.1, where these
67
+ #: three invocations returned their payload with `exit 0`.
68
+ OUTPUT_CEILING = BreakingChange(
69
+ id="BC-001",
70
+ released_in="5.7.0",
71
+ title="stdout over the output ceiling exits 1 instead of printing",
72
+ what_changed=(
73
+ "Before 5.7.0 the ~50K-token ceiling existed on `migrate-check` and "
74
+ "`repo-ir` only, so the same size was refused by one command and printed "
75
+ "in silence by another (AS-09). The guard is now one policy applied at "
76
+ "the single emit seam, which is the correct behaviour and a different "
77
+ "one: an answer that does not fit is refused with a hint instead of "
78
+ "flooding the caller."
79
+ ),
80
+ exit_code_before=0,
81
+ exit_code_after=1,
82
+ trigger=(
83
+ "any command whose stdout payload exceeds the ceiling — reached by "
84
+ "repository size, not by a flag, so the same command that fits today "
85
+ "can stop fitting on a larger repository or after a release that "
86
+ "detects more"
87
+ ),
88
+ compatibility=(
89
+ "ASK_MAX_OUTPUT_TOKENS=0 (or --max-output-tokens 0) restores the "
90
+ "pre-5.7.0 behaviour exactly: the guard is off and the whole payload is "
91
+ "printed. Set it in the pipeline that broke, then take the remedy at "
92
+ "leisure. Any other value moves the ceiling rather than removing it."
93
+ ),
94
+ remedy=(
95
+ "a bounding flag keeps the answer where the reader is — `--compact`, "
96
+ "`--limit N`, `--top-n N` or `--min-severity`, whichever the command "
97
+ "declares; `--output FILE` writes the whole payload to disk and is not "
98
+ "subject to the stdout ceiling at all. The error message names the ones "
99
+ "this invocation has not already used."
100
+ ),
101
+ observed=(
102
+ ("endpoints tutorials", "987 141 B — ~247 K estimated tokens"),
103
+ ("endpoints thingsboard", "318 380 B"),
104
+ ("posture spring-boot", "355 032 B on 8 667 Java files"),
105
+ ),
106
+ )
107
+
108
+ #: Every declared breaking change, newest first. A release that turns `exit 0`
109
+ #: into `exit 1` and does not add a row here fails `tests/test_breaking_change_
110
+ #: disclosure_ask11.py` — the policy is enforced, not remembered.
111
+ BREAKING_CHANGES: tuple[BreakingChange, ...] = (OUTPUT_CEILING,)
112
+
113
+
114
+ def help_note(change: BreakingChange = OUTPUT_CEILING) -> str:
115
+ """The `--help` line for a command the change can fire on.
116
+
117
+ Short on purpose: a help epilog that explains the whole policy is one nobody
118
+ reads. It names the limit, the exit code and the switch — the three facts
119
+ somebody scanning `--help` before writing a pipeline needs — and points at
120
+ the command that prints the rest.
121
+ """
122
+ from sourcecode.output_ceiling import DEFAULT_CEILING_TOKENS, ENV_CEILING
123
+
124
+ return (
125
+ f"Output ceiling: a stdout answer over ~{DEFAULT_CEILING_TOKENS // 1000}K "
126
+ f"estimated tokens exits {change.exit_code_after} with a hint instead of "
127
+ f"printing (changed in {change.released_in}). {ENV_CEILING}=0 restores "
128
+ f"the old behaviour; `ask schema breaking-changes-v1` explains it."
129
+ )
130
+
131
+
132
+ def affected_commands(
133
+ all_commands: "set[str]", not_an_answer: "set[str]"
134
+ ) -> "list[str]":
135
+ """The commands the output ceiling can fire on, derived rather than listed.
136
+
137
+ `not_an_answer` is the population whose stdout is a confirmation that an
138
+ effect happened rather than an answer worth keeping — the same authority
139
+ `--copy` is decided by. A confirmation cannot reach a 50K-token ceiling, and
140
+ everything else can, because the guard sits at the one emit seam they all
141
+ pass through.
142
+
143
+ Deriving it is the point. A hand-written list of affected commands beside
144
+ the registry drifts on the next command added, and this list is what a
145
+ reader is told to trust.
146
+ """
147
+ return sorted(all_commands - not_an_answer)
148
+
149
+
150
+ def as_payload(changes: "tuple[BreakingChange, ...]" = BREAKING_CHANGES) -> dict:
151
+ """The machine-readable form, for the CLI surface and for an agent."""
152
+ return {
153
+ "schema_version": "breaking-changes-v1",
154
+ "policy": (
155
+ "A change that turns `exit 0` into `exit 1` is a breaking change. It "
156
+ "is declared here, named in the `--help` of every command it can fire "
157
+ "on, and carried in the release notes with the invocations that "
158
+ "change and the size at which they change — so an upgrade can be "
159
+ "planned rather than discovered by a red pipeline."
160
+ ),
161
+ "count": len(changes),
162
+ "changes": [c.as_payload() for c in changes],
163
+ }