pi-smart-compact 10.0.1 → 10.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (53) hide show
  1. package/ARCHITECTURE.md +71 -28
  2. package/CHANGELOG.md +124 -0
  3. package/README.md +12 -7
  4. package/assets/skills/context-management/SKILL.md +5 -5
  5. package/dist/app/background-preparation.d.ts +13 -0
  6. package/dist/app/background-preparation.d.ts.map +1 -1
  7. package/dist/app/compaction-commit-store.d.ts +1 -1
  8. package/dist/app/compaction-commit-store.d.ts.map +1 -1
  9. package/dist/app/context-operations.d.ts +8 -2
  10. package/dist/app/context-operations.d.ts.map +1 -1
  11. package/dist/app/effective-state.d.ts.map +1 -1
  12. package/dist/app/hindsight-memory.d.ts.map +1 -1
  13. package/dist/app/lazy-tools.d.ts.map +1 -1
  14. package/dist/app/pending-slot.d.ts +2 -2
  15. package/dist/app/pending-slot.d.ts.map +1 -1
  16. package/dist/app/register-context-tools.d.ts.map +1 -1
  17. package/dist/app/register-navigation.d.ts +3 -1
  18. package/dist/app/register-navigation.d.ts.map +1 -1
  19. package/dist/app/register-smart-compact-command.d.ts.map +1 -1
  20. package/dist/app/register-smart-compact-tool.d.ts.map +1 -1
  21. package/dist/app/register-smart-context-tool.d.ts +1 -0
  22. package/dist/app/register-smart-context-tool.d.ts.map +1 -1
  23. package/dist/app/settled-auto-trigger.d.ts.map +1 -1
  24. package/dist/app/stage-auth.d.ts.map +1 -1
  25. package/dist/app/steps/synthesize.d.ts.map +1 -1
  26. package/dist/constants.d.ts +5 -4
  27. package/dist/constants.d.ts.map +1 -1
  28. package/dist/index.d.ts.map +1 -1
  29. package/dist/index.js +1424 -357
  30. package/dist/phases/explore.d.ts.map +1 -1
  31. package/dist/phases/synthesize.d.ts +1 -1
  32. package/dist/phases/synthesize.d.ts.map +1 -1
  33. package/dist/phases/verify.d.ts +13 -5
  34. package/dist/phases/verify.d.ts.map +1 -1
  35. package/dist/rtk.js +8 -7
  36. package/dist/types.d.ts +5 -2
  37. package/dist/types.d.ts.map +1 -1
  38. package/dist/ui/profiles.d.ts +1 -1
  39. package/dist/ui/profiles.d.ts.map +1 -1
  40. package/dist/ui/settings-overlay.d.ts.map +1 -1
  41. package/dist/ui/tool-rows.d.ts +76 -0
  42. package/dist/ui/tool-rows.d.ts.map +1 -0
  43. package/dist/utils/cache.d.ts.map +1 -1
  44. package/dist/utils/config.d.ts.map +1 -1
  45. package/dist/utils/extraction.d.ts +22 -0
  46. package/dist/utils/extraction.d.ts.map +1 -1
  47. package/dist/utils/state.d.ts.map +1 -1
  48. package/docs/README.md +8 -13
  49. package/docs/RELEASE.md +4 -3
  50. package/docs/configuration.md +101 -38
  51. package/docs/evaluation.md +29 -24
  52. package/docs/guide.md +58 -32
  53. package/package.json +1 -1
@@ -7,9 +7,9 @@ Pi Continuity is the product name only. Settings keep their technical names:
7
7
  everything lives under the `smartCompact` key, and the settings screen is opened
8
8
  with `/smart-compact settings`.
9
9
 
10
- Defaults on this page were checked against Pi Continuity `10.0.1`. For an older
11
- installation, use [the changelog](../CHANGELOG.md) and the
12
- [upgrade notes](./guide.md#upgrade-from-9x) before adopting these settings.
10
+ Defaults describe the shipped `10.1.0` working tree, including the pressure-first
11
+ changes. See [the changelog](../CHANGELOG.md) for the release scope and
12
+ evidence limits before adopting these settings.
13
13
 
14
14
  ## Contents
15
15
 
@@ -75,8 +75,8 @@ means all defaults are used.
75
75
  | Hand edits to `settings.json` | Next operation (the file's modification time is checked). Active tools and footer refresh on `/reload` or session restore; no file watcher runs. |
76
76
 
77
77
  Changing agent-tool settings from the TUI re-applies the tool list at once.
78
- Groups the agent loaded through `smart_tools` are forgotten at session start,
79
- on a branch change and after a compaction; see
78
+ In optional lazy mode, loaded groups are forgotten at session start and branch
79
+ changes, but retained across compaction to preserve the tool prefix; see
80
80
  [agent tools](./guide.md#agent-tools).
81
81
 
82
82
  ## How settings combine
@@ -114,14 +114,14 @@ nothing new is stored.
114
114
 
115
115
  | Preset | `autoTrigger` | `autoTriggerStrategy` | `contextHygieneEnabled` | `agentToolAccess` |
116
116
  | --- | --- | --- | --- | --- |
117
- | `With Pi (default)` | `true` | `native-hook` | `false` | `inherit` |
117
+ | `Pressure-first (default)` | `true` | `settled` | `true` | `inherit` |
118
118
  | `Manual only` | `false` | unchanged | `false` | `disabled` |
119
119
  | `Manual + agent` | `false` | unchanged | `false` | `enabled` |
120
120
  | `Cleanup only` | `false` | unchanged | `true` | `disabled` |
121
121
  | `Fully automatic` | `true` | `settled` | `true` | `disabled` |
122
122
 
123
- `With Pi (default)` is the unchanged built-in default, not a preset you pick.
124
- Any other combination is shown as `Custom` with a one-line description. Existing
123
+ `Pressure-first (default)` is the built-in default, not a preset you pick.
124
+ Cleanup timing defaults to pressure-only. Other combinations are shown as `Custom`. Existing
125
125
  `native-hook` configurations are not migrated to `settled`.
126
126
 
127
127
  ### Summary format
@@ -151,8 +151,8 @@ compaction. `autoTriggerStrategy` chooses how it starts.
151
151
 
152
152
  | `autoTriggerStrategy` | TUI label (`Start when`) | Who decides when | Needs Pi auto-compaction |
153
153
  | --- | --- | --- | --- |
154
- | `native-hook` (default) | `Before Pi's compaction` | Pi | Yes |
155
- | `settled` | `When idle` | Smart Compact, at an idle boundary | No |
154
+ | `native-hook` | `Before Pi's compaction` | Pi | Yes |
155
+ | `settled` (default) | `When idle` | Smart Compact, at an idle boundary | No |
156
156
  | `background` | `Prepare in background` | Smart Compact; also prepares early | No |
157
157
 
158
158
  Rules that hold for all strategies:
@@ -176,9 +176,11 @@ not as on. If it is off, nothing compacts automatically.
176
176
 
177
177
  At the next idle, queue-empty boundary where context is at least
178
178
  `minContextPercent` of the active model's window (and at least 5,000 tokens),
179
- Smart Compact asks Pi to compact. A running tool loop is not interrupted. After
180
- any confirmed compaction there is a 10-minute cooldown. Pi's own threshold and
181
- overflow triggers stay active.
179
+ Smart Compact asks Pi to compact. A running tool loop is not interrupted. Every
180
+ finished attempt starts a 10-minute cooldown, including errors, synchronous
181
+ host failures and missing-callback timeouts, so repeated idle events do not
182
+ spend another compaction budget. Confirmed compactions also start the cooldown.
183
+ Manual requests and Pi's own threshold and overflow triggers stay available.
182
184
 
183
185
  ### background: early preparation
184
186
 
@@ -192,15 +194,15 @@ hides latency, not cost: an unused summary still spends its model budget.
192
194
  | `prepareContextPercent: null` (Auto) | Starts 12.5% of the apply-token threshold earlier, bounded to 8,192–32,000 tokens |
193
195
  | Concurrency | One speculative task per extension |
194
196
  | Retry cooldown | 10 minutes |
195
- | Ready result lifetime | 5 minutes |
196
- | Context hygiene | Pressure-gated trims run under this strategy even if `contextHygieneEnabled` is off (needs active `smart_context`); break-even and cold-cache trims need `contextHygieneEnabled` |
197
+ | Ready result lifetime | `pendingTtlMs` (default 5 minutes) |
198
+ | Context hygiene | Pressure-gated trims run even if `contextHygieneEnabled` is off; economic/cold-cache timing additionally requires `contextPressureOnly: false` |
197
199
 
198
200
  Example: with `prepareContextPercent: 60` and `minContextPercent: 70` in a 200k
199
201
  window, preparation starts at 120k and applies at 140k. With Auto and an 80%
200
202
  apply gate in a 200k window, preparation starts at 140k (160k minus a 20k lead).
201
203
 
202
204
  At apply, the prepared summary is revalidated against session, branch,
203
- projected content, model, configuration, target and response headroom. A
205
+ projected content, model, effective system/tool definitions, configuration, target and response headroom. A
204
206
  changed, expired or unfinished preparation is discarded, and Smart Compact
205
207
  falls back to a normal run or Pi's own compactor. New messages after the
206
208
  snapshot stay verbatim. When cleanup changes the context first, preparation
@@ -226,10 +228,10 @@ compaction is on.
226
228
 
227
229
  ### Automatic runs cap
228
230
 
229
- Automatic runs (native-hook, settled and background) are capped at **60
230
- seconds** and **four model calls**. `autoTriggerTimeoutMs` defaults to 120,000
231
- ms and accepts up to 300,000 ms, but the effective automatic deadline is
232
- `min(autoTriggerTimeoutMs × provider timeout multiplier, 60 s)`. The call cap is
231
+ Automatic runs (native-hook, settled and background) are capped at **300
232
+ seconds** and **four model calls**. `autoTriggerTimeoutMs` defaults to 300,000
233
+ ms and accepts up to 300,000 ms; the effective automatic deadline is
234
+ `min(autoTriggerTimeoutMs × provider timeout multiplier, 300 s)`. The call cap is
233
235
  `min(mode call budget, 4)`. When an automatic run times out or fails, it unwinds
234
236
  so Pi's own compactor can run.
235
237
 
@@ -267,7 +269,7 @@ For calls (`maxLlmCalls`) and prompt tokens (`maxLlmInputTokens`):
267
269
  | Per-run option (`--max-calls`, `max_calls`, ...) | Used as given, even above the mode limit |
268
270
  | Config value `0` | The mode limit |
269
271
  | Config value above `0` | `min(config value, mode limit)` |
270
- | Automatic run | Additionally capped at 4 calls and 60 s |
272
+ | Automatic run | Additionally capped at 4 calls and 300 s |
271
273
 
272
274
  With the default `maxLlmCalls: 8`, every mode keeps its own limit, because 8 is
273
275
  not below any mode's call limit. A config value can only lower the limit.
@@ -280,6 +282,11 @@ Per-run options accept narrower ranges than the settings:
280
282
  | Prompt tokens | `--max-input-tokens` / `max_input_tokens`: 10,000–1,000,000 | `maxLlmInputTokens`: 0–1,000,000 |
281
283
  | Deadline (ms) | `--max-latency` / `max_latency_ms`: 5,000–600,000 | `maxLatencyMs`: 0 or 5,000–7,200,000 |
282
284
 
285
+ Optional exploration leaves two calls for batch synthesis and final assembly;
286
+ with two or fewer calls remaining it is skipped. Batch retries and queued work
287
+ also preserve the final assembly call. This does not raise call, token or time
288
+ limits, and cannot guarantee provider success.
289
+
283
290
  Running out of calls or tokens falls back to a deterministic summary. A deadline
284
291
  or cancellation stops the run: no staged summary, no apply, and a manual
285
292
  timeout does not start Pi's compactor.
@@ -353,7 +360,8 @@ choice because a rejected provider request falls back to the smart summary.
353
360
 
354
361
  ## Context hygiene and archives
355
362
 
356
- All three switches are off by default. Automatic trimming and offload run only
363
+ Pressure-gated automatic trimming is on by default; offload and image snapshots
364
+ remain opt-in. Automatic trimming and offload run only
357
365
  while the model can reach `smart_context`: active, or loadable through
358
366
  `smart_tools` in On demand mode. With agent tools `off`, or `smart_context`
359
367
  hidden with `/tools`, new automatic trims and offloads stop; existing archives
@@ -361,7 +369,8 @@ stay on disk and manual `/smart-compact trim` still works.
361
369
 
362
370
  | Key | TUI label | Default | Effect |
363
371
  | --- | --- | --- | --- |
364
- | `contextHygieneEnabled` | `Automatic cleanup` | `false` | Batched, recoverable trimming. Needs 16,384 characters of net savings and eight assistant turns since the last trim, rewind or compaction; commits under pressure, at break-even, or once the prompt cache is cold (below). Works with `autoTrigger: false`. |
372
+ | `contextHygieneEnabled` | `Automatic cleanup` | `true` | Batched, recoverable trimming. Needs 16,384 characters of net savings and eight assistant turns since the last trim, rewind or compaction. Works with `autoTrigger: false`. |
373
+ | `contextPressureOnly` | `Cleanup timing` | `true` | Automatic cleanup and agent trim/rewind/anchor requests require the early pressure gate. `false` opts into the legacy economic/cold-cache timing below. Human commands bypass pressure, never safety checks. |
365
374
  | `artifactOffloadEnabled` | `Offload huge outputs` | `false` | Saves eligible read-only text results of 16,384+ characters before the model sees them. Independent of pressure gates. |
366
375
  | `visualArchiveEnabled` | `Image snapshots` | `false` | Experimental image snapshots beside the verified text. Adds image tokens; needs a vision model with a validated cost rule and the optional `@resvg/resvg-js` component (not installed with the extension; `Readiness & details` shows the install command). Without it, output falls back to text. |
367
376
  | `pinPaths` | `Always-kept files` | `[]` | Paths every summary must keep |
@@ -386,6 +395,13 @@ These limits are not configurable. Behavior and examples are in the guide's
386
395
 
387
396
  ### Automatic trim timing
388
397
 
398
+ Default: only `pressure` can trigger automatic cleanup. The shared early gate is
399
+ `prepareContextPercent`, or the adaptive lead when null. A 400k policy window
400
+ with the default 80% apply gate cleans from 288k and compacts from 320k.
401
+ Unknown usage does not authorize cleanup. First-delivery offload and metadata
402
+ checkpoints do not rewrite cached history and remain independent of pressure.
403
+
404
+ The following economics apply **only with `contextPressureOnly: false`**.
389
405
  Let `X` be the estimated tokens a batch removes (net of its markers) and `T`
390
406
  the estimated tokens of every message from the first trimmed output to the
391
407
  end, the part of the prompt cache a trim rewrites. With the active model's
@@ -405,8 +421,8 @@ At a completed turn boundary a ready batch commits with cause:
405
421
  requests keep them, and the edits commit at the next completed, uncontested
406
422
  turn.
407
423
 
408
- An unknown price only allows `pressure` and `cold`. Manual and agent trims
409
- commit at the next boundary as before (causes `manual`, `agent`).
424
+ An unknown price only allows `pressure` and `cold`. Manual and permitted agent trims
425
+ commit at the next boundary (causes `manual`, `agent`).
410
426
 
411
427
  A Pi cache-warming refresh counts as a response for this expiry. While a
412
428
  batch is held, warming stops once
@@ -415,8 +431,16 @@ miss cost net of the removed output's cache write; `w'` is the cache-write
415
431
  price per million tokens, or the input price when none is listed).
416
432
 
417
433
  A newer compaction, context edit, session change, queued manual/agent request,
418
- or turning `contextHygieneEnabled` off drops the held batch. Prices are
419
- catalog ratios, not measured cache behavior.
434
+ or turning `contextHygieneEnabled` off drops the held batch. Re-enabling
435
+ pressure-only cancels unconsumed economic plans. Prices and cache lifetimes
436
+ here are heuristics, not measured provider cache behavior.
437
+
438
+ Trim/rewind refuse edits that would invalidate retained signed Anthropic
439
+ thinking. Keeping thinking bytes unchanged alone is insufficient. Use a
440
+ supported provider-native compaction route instead; native support remains opt-in.
441
+ A new anchor can request one safe cleanup of its new region before first replay;
442
+ previous anchor prefixes and recent turns remain protected. Anchor creation alone
443
+ does not invalidate an append-only background snapshot.
420
444
 
421
445
  ## Agent tools and session navigation
422
446
 
@@ -424,7 +448,7 @@ Settings → **Agent tools & navigation**.
424
448
 
425
449
  | Key | TUI label | Default | Effect |
426
450
  | --- | --- | --- | --- |
427
- | `toolLoading` | `Agent tools` | `lazy` | `lazy` (On demand): the agent sees the `smart_tools` loader and loads the `navigation`, `history`, `memory` or `compaction` group when needed; loaded groups reset at session start, branch change and compaction. `eager` (Always available): every permitted tool is active from the start. `off`: no context tools; Home, `/smart-compact` and the navigation panel keep working. |
451
+ | `toolLoading` | `Agent tools` | `eager` | `eager` (Always available): permitted tools are present from the start, keeping their prefix stable. `lazy` (On demand): load groups through `smart_tools`; late loading can rebuild the cache. Loaded groups reset at session start/branch change, not compaction. `off`: no agent context tools; human commands and navigation keep working. |
428
452
  | `contextNavigationEnabled` | `Session navigation` | `true` | Anchors, search and returning to an anchor. Off hides the panel and `smart_navigation`; recorded anchors and the keys below are kept. |
429
453
  | `contextRecallEnabled` | `Search other sessions` | `true` | Read-only search of anchors saved by earlier sessions, this project by default. |
430
454
  | `contextPivotEnabled` | `Return to an anchor` | `true` | Returning to an anchor on a new branch with a required carryover. |
@@ -487,13 +511,13 @@ Hindsight setup, data flow and troubleshooting:
487
511
  | `summaryThinkingLevel` | level \| `null` | `minimal` | `Summary thinking` |
488
512
  | `segmentationThinkingLevel` | level \| `null` | `minimal` | `Topic split thinking` |
489
513
  | `agentToolAccess` | `inherit` \| `enabled` \| `disabled` | `inherit` | `Agent can compact` |
490
- | `toolLoading` | `lazy` \| `eager` \| `off` | `lazy` | `Agent tools` |
514
+ | `toolLoading` | `lazy` \| `eager` \| `off` | `eager` | `Agent tools` |
491
515
  | `autoTrigger` | boolean | `true` | `Automatic compaction` |
492
- | `autoTriggerStrategy` | `native-hook` \| `settled` \| `background` | `native-hook` | `Start when` |
493
- | `minContextPercent` | 0–100 | `60` | `Start at context %` |
494
- | `prepareContextPercent` | `null` or 0–100, below `minContextPercent` | `null` | `Prepare at context %` |
516
+ | `autoTriggerStrategy` | `native-hook` \| `settled` \| `background` | `settled` | `Start when` |
517
+ | `minContextPercent` | 0–100 | `80` | `Start at context %` |
518
+ | `prepareContextPercent` | `null` or 0–100, below `minContextPercent` | `null` | `Cleanup / prepare at context %` |
495
519
  | `maxContextTokens` | `0` (off) or integer 16,384–2,000,000 | `0` | `Context cap for start % (tokens)` |
496
- | `autoTriggerTimeoutMs` | integer 1,000–300,000 | `120000` | `Automatic run time limit (ms)`; capped at 60 s |
520
+ | `autoTriggerTimeoutMs` | integer 1,000–300,000 | `300000` | `Automatic run time limit (ms)`; capped at 300 s |
497
521
  | `compactionEngines` | ordered list of `eesv`, `native` | `["eesv"]` | `Engine` |
498
522
  | `requireApproval` | boolean | `true` | `Ask before applying` |
499
523
  | `showStatus` | boolean | `true` | `Footer status` |
@@ -502,7 +526,8 @@ Hindsight setup, data flow and troubleshooting:
502
526
  | `maxLatencyMs` | `0` or integer 5,000–7,200,000 | `0` (no limit) | `Run time limit (ms)` |
503
527
  | `codexMaxCallMs` | `0` or integer 5,000–3,600,000 | `0` (auto 15–90 s) | `Stuck-call timeout (ms)` |
504
528
  | `pendingTtlMs` | integer 1,000–3,600,000 | `300000` | `Prepared summary lifetime (ms)` |
505
- | `contextHygieneEnabled` | boolean | `false` | `Automatic cleanup` |
529
+ | `contextHygieneEnabled` | boolean | `true` | `Automatic cleanup` |
530
+ | `contextPressureOnly` | boolean | `true` | `Cleanup timing` |
506
531
  | `artifactOffloadEnabled` | boolean | `false` | `Offload huge outputs` |
507
532
  | `visualArchiveEnabled` | boolean | `false` | `Image snapshots` |
508
533
  | `pinPaths` | string array | `[]` | `Always-kept files` |
@@ -537,10 +562,9 @@ Notes on specific keys:
537
562
  apply gate for automatic and agent runs and the replacement gate for
538
563
  `native-hook`. Manual `/smart-compact` shows a warning and ignores it. A
539
564
  5,000-token floor always applies.
540
- - `pendingTtlMs` is how long a summary returned to Pi waits for Pi's matching
541
- compaction confirmation. It does not change the fixed 5-minute window in
542
- which a summary staged by the `smart_compact` tool must be applied with
543
- `/compact`, or the 5-minute lifetime of a background preparation.
565
+ - `pendingTtlMs` governs staged and ready background summaries (default five
566
+ minutes), as well as the host-confirmation retention budget. Changing it
567
+ never permits reuse of a changed prefix or an unsafe retained tail.
544
568
  - `showStatus` adds a footer note only when compaction is manual-only or
545
569
  disabled. Nothing is shown while healthy, and background preparation is not
546
570
  shown in the footer.
@@ -570,6 +594,45 @@ Notes on specific keys:
570
594
 
571
595
  ## Examples
572
596
 
597
+ ### Cache-first automatic compaction
598
+
599
+ Keep roomy cached history stable, clean under pressure, and avoid paying for
600
+ speculative summaries that may expire unused:
601
+
602
+ ```json
603
+ {
604
+ "smartCompact": {
605
+ "autoTrigger": true,
606
+ "autoTriggerStrategy": "settled",
607
+ "toolLoading": "eager",
608
+ "contextHygieneEnabled": true,
609
+ "contextPressureOnly": true,
610
+ "minContextPercent": 80,
611
+ "maxContextTokens": 0,
612
+ "artifactOffloadEnabled": true
613
+ }
614
+ }
615
+ ```
616
+
617
+ With the optional cap off, thresholds follow Pi's active model `contextWindow`,
618
+ including `models.json` overrides: a 400k window cleans at 288k and compacts at
619
+ 320k; a 1M window cleans at 768k and compacts at 800k. The adaptive lead is bounded,
620
+ not a fixed 72% cleanup threshold. After editing model metadata, reselect the
621
+ model through Pi so its active model reflects the override. Overrides are scoped
622
+ to a provider/model pair and cannot enlarge the backend's actual capacity.
623
+
624
+ Existing summary budgets remain unchanged. Compaction waits
625
+ for an idle, queue-empty boundary; long tool loops can still grow before that
626
+ boundary, and preparing the summary adds latency when it is needed. Offload
627
+ shortens eligible large search/web results before their first model request,
628
+ not cached history; file reads and shell output stay inline. Bound those at
629
+ the tool call with ranges, symbols or concise command output. Retrieval can
630
+ spend tokens too, so smaller context alone is not proof of lower billed cost.
631
+ For a lower automatic trigger threshold on large-window models, set
632
+ `maxContextTokens`; it is not a hard per-request token limit.
633
+
634
+ ### Other configurations
635
+
573
636
  Cleanup and offload without automatic compaction, with every permitted tool
574
637
  visible to the agent from the start:
575
638
 
@@ -133,9 +133,8 @@ receives the same three bounded coding-continuity scenarios (`implementation`,
133
133
  `debugging`, `continuity`) sequentially, with a 1,500-token output cap and a
134
134
  60-second timeout per call, scored by the deterministic verifier. It reports
135
135
  score, latency and reported usage. Apply a route manually, and only after
136
- representative evidence; one probe is not that evidence. The dated
137
- [2026-08-06 baseline](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/reports/provider-evaluation-2026-08-06.md)
138
- (repository only) is an example of this output, not a current ranking.
136
+ representative evidence; one probe is not that evidence. Historical probe
137
+ results are not a current model ranking.
139
138
 
140
139
  ## Paired continuation and memory evaluation
141
140
 
@@ -160,8 +159,21 @@ Options (from `--help`): `--arms`, `--repeats=1..8` (default five rounds, probes
160
159
  after rounds two and five), `--out` (absolute directory; defaults to
161
160
  `./task-eval-reports/<timestamp>`), `--json`, and two offline-only
162
161
  cost-accounting fixtures, `--cache-warming=off|streaming|idle` and
163
- `--background-prep`. Every arm runs with `toolLoading: "eager"`, so the arms
164
- differ only in hygiene/offload knobs, not in on-demand tool discovery. With
162
+ `--background-prep`. `--tool-loading=eager|lazy` selects the same exposure mode
163
+ for every arm (default eager). Offline lazy frames explicitly load missing groups;
164
+ this measures host/discovery mechanics, not autonomous tool choice. Prefix
165
+ comparisons include active tool schemas and the SDK-collapsed system prompt,
166
+ not just conversation text. They model providers without in-place additions,
167
+ not an exact provider wire payload or billed cache-hit rate.
168
+
169
+ For a paired exposure comparison:
170
+
171
+ ```bash
172
+ bun run task-eval --arms=hybrid --repeats=2 --tool-loading=eager --out=/tmp/psc-eager
173
+ bun run task-eval --arms=hybrid --repeats=2 --tool-loading=lazy --out=/tmp/psc-lazy
174
+ ```
175
+
176
+ With
165
177
  `--cache-warming=idle`, arms that have automatic cleanup on may commit a
166
178
  break-even trim at an idle boundary; the rebuilt context ends Pi's warming for
167
179
  that entry, so zero warm replays there is expected, while the `no-compaction`
@@ -387,22 +399,15 @@ no provider request.
387
399
  | `PSC_HINDSIGHT_LIVE=1 … bun run test/hindsight-live.canary.ts` | Not run by `bun test` | Live Hindsight contract with a hard call budget; see [Hindsight memory](./hindsight-memory.md#tests) |
388
400
 
389
401
  Dated reports record what was measured on their date, with the code and host
390
- versions stated inside. They are kept for provenance and are not updated
391
- retroactively; later findings are added as dated addenda. Current behavior is
392
- described in the [guide](./guide.md), [configuration](./configuration.md) and
393
- [architecture](../ARCHITECTURE.md). Reports live in the repository only; the
394
- npm package does not include `docs/reports/` or `docs/findings/`. Pilot
395
- reports have a machine-readable `.json` companion beside them.
396
-
397
- | Report | Scope |
398
- | --- | --- |
399
- | [Provider evaluation baseline, 2026-08-06](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/reports/provider-evaluation-2026-08-06.md) | Live three-scenario probe across five models; advisory; 2026-09-25 addendum on route-report token semantics |
400
- | [Context hygiene and continuity, 2026-09-24](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/reports/context-hygiene-2026-09-24.md) | Hygiene design and offline experiments |
401
- | [Hindsight and provider-native compaction research, 2026-09-24](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/reports/hindsight-native-compaction-research-2026-09-24.md) | Pre-implementation research plus later measured results |
402
- | [Full AgentSession offline pilot, 2026-09-24](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/reports/session-pilot-2026-09-24.md) | Scripted-transport lifecycle pilot; 2026-09-25 candidate and `9.8.0-canary.1` follow-ups |
403
- | [Visual evidence pilot, 2026-09-24](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/reports/visual-pilot-2026-09-24.md) | Live synthetic bitmap-versus-text reading pilot on one model |
404
-
405
- External review findings, one folder per reviewer, are indexed in
406
- [`docs/findings/`](https://github.com/alpertarhan/pi-smart-compact/blob/main/docs/findings/README.md).
407
- They are advisory analyses of a specific revision range, not release gates or
408
- evidence of current behavior.
402
+ versions stated inside. Preserve their provenance; add dated addenda rather
403
+ than updating measurements retroactively. Current behavior is described in the
404
+ [guide](./guide.md), [configuration](./configuration.md) and
405
+ [architecture](../ARCHITECTURE.md).
406
+
407
+ Research reports in `docs/reports/` and review findings in `docs/findings/`
408
+ are local, ignored artifacts, not tracked repository or npm content. Pilot
409
+ reports can keep machine-readable `.json` companions beside them. Earlier
410
+ copies may remain in Git history; this policy does not erase that history.
411
+
412
+ External review findings are advisory analyses of a specific revision range,
413
+ not release gates or evidence of current behavior.
package/docs/guide.md CHANGED
@@ -15,8 +15,8 @@ This guide is task oriented. For every setting, default and range, see the
15
15
  release evidence, see [evaluation](./evaluation.md).
16
16
 
17
17
  > [!NOTE]
18
- > **Which version this describes.** This guide targets Pi Continuity `10.0.1`,
19
- > the stable release of the context-hygiene and continuity rework. The earlier
18
+ > **Which version this describes.** This guide describes the `10.1.0` shipped
19
+ > defaults, including the pressure-first changes and native tool rows. The earlier
20
20
  > `9.8.0-canary.*` entries in the changelog are historical local candidates,
21
21
  > not npm releases. See [upgrade notes](#upgrade-from-9x) when coming from 9.x
22
22
  > and [the changelog](../CHANGELOG.md) for the release scope and evidence limits.
@@ -65,8 +65,8 @@ default.
65
65
 
66
66
  | Status | Features |
67
67
  | --- | --- |
68
- | On by default | Replacing Pi's summary when Pi compacts, **Compact now**, `/smart-compact trim`, session navigation, agent tools on demand, local project memory, backups, secret scrubbing |
69
- | Optional, off until selected | **Automatic cleanup**, **Offload huge outputs**, the `When idle` and `Prepare in background` strategies, the Hindsight and Mnemopi memory stores, personal-data scrubbing |
68
+ | On by default | Pressure-gated batched cleanup, idle compaction at 80%, stable permitted agent tools, **Compact now**, `/smart-compact trim`, session navigation, local project memory, backups, secret scrubbing |
69
+ | Optional, off until selected | **Offload huge outputs**, economic/cold-cache cleanup, `Prepare in background`, the Hindsight and Mnemopi memory stores, personal-data scrubbing |
70
70
  | Experimental | [Provider compaction, image snapshots](#summary-format-provider-compaction-and-images) and the [RTK companion](#rtk-companion-optional-experimental) |
71
71
 
72
72
  Settings and defaults are listed in the
@@ -86,9 +86,10 @@ Then, inside Pi:
86
86
  /smart-compact
87
87
  ```
88
88
 
89
- Opening Home changes nothing. With the built-in defaults, automatic compaction
90
- depends on Pi starting it. Automatic local cleanup, early output offload and
91
- speculative background preparation are off until selected.
89
+ Opening Home changes nothing. With the new defaults, batched cleanup starts
90
+ under early pressure and compaction starts at an idle 80% boundary. A 400k
91
+ policy window means 288k cleanup / 320k compaction. Offload and speculative
92
+ background preparation remain optional; no extra model calls run while roomy.
92
93
 
93
94
  ## Upgrade from 9.x
94
95
 
@@ -97,20 +98,20 @@ speculative background preparation are off until selected.
97
98
  and stored paths keep their names; do not rename existing data directories.
98
99
 
99
100
  1. Update Pi to **0.87.1+** and use **Node.js 22.19+**. Install the release with
100
- `pi install npm:pi-smart-compact@10.0.1`, then reload or restart Pi.
101
+ `pi install npm:pi-smart-compact@10.1.0`, then reload or restart Pi.
101
102
  2. Open `/smart-compact` in the TUI. The bare command now opens Home; **Compact
102
103
  now** starts the interactive compaction flow. Print/RPC/SDK use still runs
103
104
  compaction directly. Review requires **A** to apply, not Enter.
104
- 3. Expect the agent to see `smart_tools` first. It loads navigation, history,
105
- memory and compaction tools on demand. Choose **Always available** only if
106
- you want all permitted tools exposed from the start.
105
+ 3. 10.1.0 uses the pressure-first default: **Always available** tool exposure
106
+ to keep tool definitions stable. Existing explicit `toolLoading: "lazy"`
107
+ settings remain respected.
107
108
  4. Review **Memory store** if you use project memory. Exactly one backend is
108
109
  used, with no silent fallback. Mnemopi and image snapshots now require
109
110
  [separately installed optional components](../README.md#optional-components);
110
111
  the default local memory store does not.
111
112
 
112
- Automatic cleanup, early offload, background preparation, provider-native
113
- compaction and image snapshots remain opt-in. Existing compaction permissions
113
+ In the 10.1.0 defaults, cleanup is pressure-gated and enabled. Early
114
+ offload, background preparation, provider-native compaction and images remain opt-in. Existing compaction permissions
114
115
  are preserved. Claude subscription routes still need the compatible separate
115
116
  adapter described under [provider compaction](#summary-format-provider-compaction-and-images).
116
117
 
@@ -207,15 +208,16 @@ changes only the order and the note, never which outputs are eligible.
207
208
 
208
209
  ### Automatic cleanup (optional)
209
210
 
210
- Automatic trimming is separate and off by default: turn on
211
- **Automatic cleanup** (`contextHygieneEnabled`), or pick the `Cleanup only` or
212
- `Fully automatic` preset. It uses the same eligibility rules as the manual
211
+ Automatic trimming is separate and enabled under pressure by default. Control
212
+ it with **Automatic cleanup** (`contextHygieneEnabled`); **Cleanup timing**
213
+ (`contextPressureOnly`) defaults to pressure-only. It uses the same eligibility rules as the manual
213
214
  command and needs `smart_context` reachable by the model, since the digest
214
215
  points it there (any **Agent tools** choice except **Off**).
215
216
 
216
217
  A batch needs at least 16,384 characters of net savings and eight assistant
217
218
  turns after the last trim, rewind or compaction. It then commits at the turn
218
- boundary for one of three causes, recorded on the trim entry:
219
+ boundary under pressure. The other two causes below require the explicit
220
+ **Economic (opt-in)** timing choice (`contextPressureOnly: false`):
219
221
 
220
222
  - `pressure`: context usage reached the early pressure gate.
221
223
  - `break-even`: the model's catalog prices say the trim pays back its prompt
@@ -290,7 +292,7 @@ to existing settings.
290
292
 
291
293
  | Preset | Automatic compaction | Automatic cleanup | Agent can call `smart_compact` |
292
294
  | --- | --- | --- | --- |
293
- | `With Pi (default)` | Only when Pi's own auto-compaction fires | Off | Follows Pi's tool settings |
295
+ | `Pressure-first (default)` | Starts when idle at the 80% apply gate | Under early pressure | Follows Pi's tool settings |
294
296
  | `Manual only` | Off | Off | No |
295
297
  | `Manual + agent` | Off | Off | Yes |
296
298
  | `Cleanup only` | Off | On | No |
@@ -298,13 +300,13 @@ to existing settings.
298
300
 
299
301
  Two points matter most:
300
302
 
301
- - **`With Pi (default)` is passive.** It replaces the summary only when Pi
302
- starts a compaction. If Pi's auto-compaction is off, nothing compacts
303
- automatically. Extensions cannot read Pi's setting, so readiness reports it
304
- as `unknown`.
303
+ - **The default is pressure-first.** Roomy history stays unchanged; cleanup
304
+ and compaction share one policy window. The optional `native-hook` strategy
305
+ remains passive and requires Pi's auto-compaction to be enabled.
305
306
  - **`Fully automatic` works with Pi's auto-compaction off.** It asks Pi to
306
307
  compact at the next idle, queue-empty boundary once context reaches
307
- `Start at context %`, with a 10-minute cooldown after a compaction.
308
+ `Start at context %`, with a 10-minute cooldown after every completed
309
+ attempt, including failures and missing-callback timeouts.
308
310
 
309
311
  `Start at context %` counts against the model's full window. On a large-window
310
312
  model (Home warns above 400k tokens), set `Context cap for start % (tokens)`
@@ -312,12 +314,17 @@ model (Home warns above 400k tokens), set `Context cap for start % (tokens)`
312
314
  and safety headroom still use the real window.
313
315
 
314
316
  Turning `Automatic compaction` off disables both Smart Compact strategies. It
315
- does not turn off Pi's own compactor. Automatic runs are capped at 60 seconds
317
+ does not turn off Pi's own compactor. Automatic runs are capped at 300 seconds
316
318
  and four model calls, whatever the configured limits are. When they fail, Pi
317
319
  may still use its own compactor. See
318
320
  [automatic strategies](./configuration.md#automatic-strategies) for the
319
321
  background strategy and the gate arithmetic.
320
322
 
323
+ For cost-sensitive sessions, use the [cache-first configuration](./configuration.md#cache-first-automatic-compaction):
324
+ stable tools, idle compaction, pressure-only cleanup, no speculative preparation,
325
+ and first-delivery offload for eligible large outputs. Existing mode budgets and
326
+ summary quality checks remain in place.
327
+
321
328
  ## Retrieve archived output
322
329
 
323
330
  The agent retrieves trimmed, rewound, offloaded or image-archived output with
@@ -393,6 +400,11 @@ Example inputs (one JSON object per call; explore between checkpoint and rewind)
393
400
  {"action":"trim"}
394
401
  ```
395
402
 
403
+ - Check `status` for actual usage and gates. Checkpoints are permitted below
404
+ pressure; agent rewind/trim/anchor requests are not, by default. Human
405
+ commands can request early work. Unsafe edits before retained signed
406
+ Anthropic thinking are refused, even when thinking bytes themselves would
407
+ stay unchanged.
396
408
  - `checkpoint`, `rewind` and `trim` return **queued**. Pi commits them at the
397
409
  end of the current tool batch.
398
410
  - One checkpoint is active at a time; a new one replaces it. It survives reload.
@@ -414,7 +426,11 @@ Example inputs (one JSON object per call; explore between checkpoint and rewind)
414
426
  An anchor is a named point in this conversation with a summary of what was
415
427
  true there: the goal, decisions and the state of the work. Anchors are stored
416
428
  as messages in the session, so the agent sees them too. Use them to find the
417
- way back after a long detour.
429
+ way back after a long detour. Under pressure, an agent anchor also queues one
430
+ safe cleanup of its new region through the shared trim controller. Previous
431
+ anchor prefixes stay protected. It is not a full compaction; the response says
432
+ whether cleanup was queued, blocked or unnecessary. A human may mark a milestone
433
+ earlier. Append-only anchors do not discard an already prepared summary.
418
434
 
419
435
  Home → **History & recovery** → **Session navigation**, or `/smart-compact context`:
420
436
 
@@ -488,8 +504,8 @@ decides what the agent sees:
488
504
 
489
505
  | Choice | What the agent sees |
490
506
  | --- | --- |
491
- | **On demand** (default) | One small loader, `smart_tools`. The agent loads a group when it needs it: `navigation`, `history`, `memory` or `compaction`. Loaded groups are forgotten at the next session start, branch change or compaction. |
492
- | **Always available** | Every permitted tool from the start. |
507
+ | **Always available** (default) | Every permitted tool from the start, keeping tool definitions stable. |
508
+ | **On demand** | `smart_tools` loads a missing group. Late loading can rebuild the provider cache. Loaded groups reset at session start/branch change, not compaction. |
493
509
  | **Off** | No context tools. The human UI keeps working. |
494
510
 
495
511
  `smart_tools` also answers `status`, `unload` and `guide`. The guide is the
@@ -508,11 +524,19 @@ earlier request bytes disappear.
508
524
  | `smart_context` | `history` | Status, search, read, plan, trim, checkpoint, rewind | Queued edits apply at the end of the current tool batch |
509
525
  | `smart_recall` | `memory` | Searches this project's memory in the selected store | Immediately; read-only |
510
526
  | `smart_save_memory` | `memory` | Saves or resolves one durable fact | Only after **you** approve the host confirmation dialog |
511
- | `smart_compact` | `compaction` | Prepares a verified summary and stages it | Never mid-turn. Staged for 5 minutes; applied by the next `/compact` or a compaction Pi starts. |
527
+ | `smart_compact` | `compaction` | Prepares a verified summary and stages it | Never mid-turn. Staged for `pendingTtlMs` (default 5 minutes); applied by the next `/compact` or a compaction Pi starts. |
528
+
529
+ **Tool rows.** Every Smart Compact tool renders a native, compact status row:
530
+ queued, staged, dry-run, skipped, pending, failed and cancelled states are
531
+ labeled as such and never displayed as done, applied or saved. Collapsed rows
532
+ show bounded previews and hide long identifiers behind the native expansion
533
+ key (the configured keybinding, shown as a hint); expanding shows the full
534
+ original result text. Rows are call-time snapshots — they never poll or refresh
535
+ on their own.
512
536
 
513
537
  Approval and gates:
514
538
 
515
- - `smart_compact` refuses below `Start at context %` (default 60%) or below
539
+ - `smart_compact` refuses below `Start at context %` (default 80%) or below
516
540
  5,000 tokens. `tool=XX%` in Pi's footer is the tool-output share, not
517
541
  context fullness. If a summary is already staged, it reports that and makes
518
542
  no calls. With automatic compaction off, nothing consumes the staged summary
@@ -795,9 +819,9 @@ What is not covered:
795
819
  | Symptom | Cause and action |
796
820
  | --- | --- |
797
821
  | Agent reports "Compaction skipped: context 38% (…) below the 60% agent-tool threshold" | `smart_compact` waits for `Start at context %` of the **active model's** window (`tool=XX%` in the footer is the tool-output share, not context fullness). Use **Compact now** for early compaction. |
798
- | Nothing compacts automatically | With the default `With Pi`, Pi's auto-compaction must be on (readiness shows `unknown`). Choose `Fully automatic` to start without it. Check `Automatic compaction` and `This branch only`. On a large-window model, see `Context cap for start % (tokens)` in [Let it run automatically](#let-it-run-automatically). |
822
+ | Nothing compacts automatically | Check the idle boundary and pressure gates. Only the optional `native-hook` strategy requires Pi's auto-compaction (readiness shows `unknown`). Check `Automatic compaction` and `This branch only`. On a large-window model, see `Context cap for start % (tokens)` in [Let it run automatically](#let-it-run-automatically). |
799
823
  | Cleanup did nothing on the next reply | Expected: the next request is sent untrimmed; cleanup applies at the next completed turn. |
800
- | **Clean up tool output** shows `held for a cold cache` | Expected with **Automatic cleanup**: the batch waits for the prompt cache to expire; see [automatic cleanup](#automatic-cleanup-optional). Select the row to apply it at the next completed turn instead. |
824
+ | **Clean up tool output** shows `held for a cold cache` | Expected only with **Economic (opt-in)** cleanup timing: the batch waits for the prompt cache to expire; see [automatic cleanup](#automatic-cleanup-optional). Select the row to apply it at the next completed turn instead. |
801
825
  | Automatic cleanup or offload never happens | Both need `smart_context` reachable: **Agent tools** must not be **Off**, and `smart_context` must not be hidden with `/tools`. Offload also needs **Offload huge outputs** on and applies only to read-only text results of 16,384+ characters. |
802
826
  | Agent says a summary is staged, but context did not shrink | Run `/compact` within 5 minutes. With automatic compaction off, nothing else consumes it. |
803
827
  | `smart_context`, `smart_recall`, `smart_save_memory` or `smart_navigation` missing | With **On demand**, the agent loads a group through `smart_tools`; with **Off**, no context tool is shown; see [agent tools](#agent-tools). Check `/tools`. |
@@ -810,7 +834,9 @@ What is not covered:
810
834
  | Warning line at the top of Settings | Some `settings.json` values were invalid or converted and are being ignored; the line names them. |
811
835
  | Provider compaction was skipped | The chat model's API is not a supported native route; verified text was used. In sessions smaller than Pi's `compaction.keepRecentTokens`, Pi refuses the result and the conversation stays unchanged. |
812
836
  | `Text + images` produced text only | Expected unless the chat model is direct Anthropic `claude-sonnet-5` and the optional `@resvg/resvg-js` component is installed; see [image snapshots](#summary-format-provider-compaction-and-images). |
813
- | Timeout or verification failure | A single `Smart Compact: ...` line explains the effect and next step. Details: `/smart-compact metrics`. Stack traces: `DEBUG=smart-compact`. |
837
+ | Synthesis fallback in progress or metrics | Intermediate recovery, not an apply result. The final `Smart compact applied` notice discloses fallback use. Generation error details remain in `/smart-compact metrics` and verbose output. |
838
+ | Timeout or verification failure | No Smart Compact summary was applied. During host-triggered compaction, Pi may then produce its own summary; that is a separate outcome. A single `Smart Compact: ...` line explains the failure. Details: `/smart-compact metrics`. Stack traces: `DEBUG=smart-compact`. |
839
+ | `Cache miss: … tokens re-billed (~$…)` | Pi's own notice, not a compaction failure. The amount is a token-price estimate, not a verified invoice. Smart Compact records rebuild attribution in Home → Readiness & details and metrics without a duplicate toast. |
814
840
 
815
841
  Without a UI, warnings and errors go to stderr.
816
842
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-smart-compact",
3
- "version": "10.0.1",
3
+ "version": "10.1.1",
4
4
  "description": "Pi Continuity — context hygiene and session continuity for Pi Coding Agent: recoverable cleanup, checkpoints, memory and verified compaction.",
5
5
  "type": "module",
6
6
  "packageManager": "bun@1.4.2",