pi-zip 0.2.8 → 0.3.0-rc.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -30,76 +30,91 @@ Known limits:
30
30
 
31
31
  - **GLM** (automatic prefix cache that outlives its declared 5 minutes): v0.1 cost about 1.2× BC live. v0.2 learns the real cache lifetime and folds the previous turn after the declared TTL (offline 0.95–0.97× v0.1), but this was not verified live.
32
32
  - **Long autonomous runs** (one prompt, hundreds of tool calls, e.g. sub-agents): roughly on par with BC, not better. Outputs are only folded inside a running turn once they are 60 requests old.
33
- - **Tool allowlists** (`pi --tools read,bash`, and sub-agent launchers that pass one): Pi then hides `zip_recall`, so pi-zip folds only outputs the model can re-read (files, read-only commands) and the placeholder says to re-read. Add `zip_recall` to the list to get full folding.
33
+ - **Tool allowlists** (`pi --tools read,bash`, and sub-agent launchers that pass one): Pi then hides `zip_recall`. If `read`, `grep` or `bash` is on the list, pi-zip still folds as usual and the placeholder names a file the model reads or greps instead (see [Recall through files](#recall-through-files)). Only when none of them is allowed does it fold just the outputs the model can re-read (files, read-only commands), with a placeholder that says to re-read. Add `zip_recall` to the list to get the tool itself.
34
34
  - If you always answer within the cache lifetime, there is little to save, by design.
35
35
 
36
36
  ## What you see
37
37
 
38
- At most one line per turn in the transcript, only when the context was folded or summarized. The numbers come first, and the bar shows how much is left:
38
+ One mark, ƶ, and three words: **zipped**, **unzip**, **summarized**. Only the new context size is bright. At most one line per turn, at the top of the turn, only when the context was zipped or summarized:
39
39
 
40
40
  ```
41
- ▸ pi-zip 74K → 43K ▰▰▰▰▰▰▱▱▱▱ folded 12 old outputs
42
- originals are kept; the model can recall any of them with zip_recall
43
- ▸ pi-zip 182K → 41K ▰▰▱▱▱▱▱▱▱▱ summarized 64 requests · ready while you were away
41
+ ƶ zipped 12 old outputs 74K → 43K 6 ms
42
+ originals are kept · the model can unzip any of them
43
+ ƶ summarized 64 requests 182K → 41K 38.2 s · ready while you were away
44
44
  ```
45
45
 
46
- The second line appears once per session. Click the notice (or expand tool output with ctrl+o) to see why it happened now and what was folded:
46
+ The second line appears once per session. A summary that made you wait says so (`waited 8.4 s`, highlighted from 2 s on). Click the line (or expand tool output with ctrl+o) to see what was zipped and why now (a cold return says `cache cold · away 47 min` or `cache cold · model switched`, a warm edit `cache warm · worth it at this size`):
47
47
 
48
48
  ```
49
- ▸ pi-zip 74K → 43K ▰▰▰▰▰▰▱▱▱▱ folded 12 old outputs
50
- cache cold (away 47 min): this request rewrites it anyway, so editing is free
51
- bash npm test turn 1 14K k3x9q2m7ab
52
- read src/payment.ts turn 1 9K p8d2x1qa0m
53
- … 10 more
49
+ ƶ zipped 12 old outputs 74K → 43K 6 ms
50
+ bash npm test 14K k3x9q2m7ab
51
+ read src/payment.ts 9.1K p8d2x1qa0m
52
+ … 10 more
53
+ cache cold · away 47 min
54
54
  ```
55
55
 
56
- These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the words go first, then the bar; the numbers always stay. A summary still shows up as Pi's own `[compaction]` block as well; the pi-zip line next to it tells you who made it. The pi-zip numbers are real tokens, the same scale as Pi's context meter (the "after" side is what the edited request actually carried); the `Compacted from N tokens` figure in Pi's block is Pi's own estimate taken when the entry is saved, so it will not match.
56
+ These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the time goes first, then the words; the numbers always stay. A summary still shows up as Pi's own `[compaction]` block as well; the ƶ line next to it tells you who made it. The pi-zip numbers are real tokens, the same scale as Pi's context meter (the meter adds the newest reply's output on top); the `Compacted from N tokens` figure in Pi's block is pi-zip's own size before for a summary prepared while you were away (that one goes in through Pi's compaction, before your prompt); for any other summary it is Pi's own estimate taken when the entry is saved, so it may not match.
57
57
 
58
- Everything else pi-zip shows uses the same one-line grammar, and only when something changed:
58
+ A zipped output keeps its place in the transcript; its tool row gets a dim mark, so you can see what the model no longer sees in full:
59
59
 
60
60
  ```
61
- ▸ pi-zip on folds old tool output after the prompt cache expires (5 min here) · nothing to set up · /zip for status
62
- ▸ pi-zip reread-only a tool allowlist hides zip_recall · only re-readable outputs fold · allow zip_recall to fold more
63
- ▸ pi-zip paused billion-context also manages context, so pi-zip only guards requests · to use pi-zip: pi remove the other one
61
+ $ npm test ƶ zipped
62
+ read src/payment.ts ƶ zipped
64
63
  ```
65
64
 
66
- The first one appears once per machine. A folded output keeps its place in the transcript; its tool row gets a dim mark on the right, so you can see what the model no longer sees in full:
65
+ When the model fetches an original back, it looks like any other tool row and names what came back (ctrl+o shows the text):
67
66
 
68
67
  ```
69
- $ npm test ▸ folded · k3x9q2m7ab
68
+ ƶ unzip bash npm test /Expected/
69
+ 3 of 812 lines
70
70
  ```
71
71
 
72
- A recall is one quiet row (ctrl+o shows the recalled text):
72
+ Everything else uses the same grammar, and only when something changed:
73
73
 
74
74
  ```
75
- ↺ recall k3x9q2m7ab grep "Expected"
76
- bash npm test · turn 1 3 of 812 lines
75
+ ƶ pi-zip on old tool output gets zipped once the prompt cache expires (5 min here) · /zip
76
+ ƶ pi-zip on a tool allowlist hides zip_recall · zipped outputs are read back from files in ~/.cache/pi-zip/recall
77
+ ƶ pi-zip reread-only a tool allowlist hides zip_recall · only re-readable outputs fold · allow zip_recall, read, grep or bash to fold more
78
+ ƶ pi-zip paused billion-context also manages context, so pi-zip only guards requests · to use pi-zip: pi remove the other one
77
79
  ```
78
80
 
79
- While a summary started in the background is being finished, Pi's working line says so (Esc skips the wait). `/zip` prints a small card into the transcript:
81
+ The first one appears once per machine. While a summary started in the background is being finished, Pi's working line says so (Esc there cancels the prompt, as it does whenever Pi is working; the summary goes on for your next try). `/zip` prints a small card into the transcript:
80
82
 
81
83
  ```
82
- ▸ pi-zip on
83
- cache anthropic/claude-sonnet-5-5 · explicit · lives ~5 min (declared)
84
- alive ██████▁▁▁▁▁▁▁▁▁▁ 30 s → 2 h
85
- session 56 folds ~310K · 2 summaries $0.41 · 3 recalls
86
- last 10:50 folded 12 old outputs · cache cold (away 47 min): this request rewrites it anyway, so editing is free
87
- mode full (zip_recall available)
84
+ ƶ pi-zip on
85
+ model zhipu/glm-5.3
86
+ cache each dot is one time you came back
87
+ 6 5
88
+ ● ● ● ●
89
+ ● ● ● ● ●
90
+ ● ● ● ● ● ● still there 19
91
+ ━━━━━━━━━━━━━━━━┅┅┅┅┅┅┅┅┅┅───────────────────────────────────
92
+ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ gone 74
93
+ ○ ○ ○ ○ ○ ○ ○ ○ ○
94
+ ○ ○ ○ ○ ○ ○ ○ ○
95
+ ○ 12 7 9 15 6 11 5
96
+ away 30s 1m 2m 5m 10m 30m 1h 2h
97
+ under 2 min: usually still there · 2–5 min: not sure yet · over 5 min: gone
98
+ session 30 zipped ~310K · 8 summaries $1.75 · 2 unzips
99
+ last 15:22 zipped 1 old output in 4 ms · cache cold · away 13 min
88
100
  ```
89
101
 
90
- `alive` is what pi-zip currently believes about the cache: how likely it is to be still warm after 30 s, 1 min, … 2 h away, learned from the provider's replies (see below).
102
+
103
+ ƶ (U+01B6) is a plain Latin letter, one column wide everywhere, CJK terminals included. Most coding fonts have it; where one does not (Hack, Source Code Pro), the terminal draws it from a fallback font.
104
+
105
+ The chart is the one thing worth knowing: come back sooner than the solid part of the axis and pi-zip touches nothing; come back later and the old outputs get zipped, at no extra cost. Across (`away`) is how long you were gone, from 30 s to 2 h on a log scale. Every dot is one time you came back after that long: a filled dot above the axis is a time the cache was still there, an empty one below it a time it was gone, so the picture is your own history with this model, and a few gone samples take only a row or two. A column shows at most 4 dots a side; beyond that it shows 3 dots and the number at the far end of the stack (`12` = twelve times in that gap range). The axis is solid up to the longest gap that was (nearly) always safe, dashed while it is not sure, thin after it is mostly gone; the sentence under it says the same in words (`usually still there` when even the safe part had some misses). The counts are the learned ones, kept with older evidence halved after 16 newer observations of the same gap range, so they can be a little below the number of times you actually came back. Until something is learned there is no chart, one line only (`expires after ~5 min idle · a chart appears once you come back after a break`, from the model's declared cache lifetime; Claude's is 5 min). The chart needs 72 columns; where the legend does not fit beside it, it moves to one line under it, and in a narrower terminal only the sentence is shown. Notices state facts (how long you were away), not guesses: whether the cache really had expired is read from the provider's reply afterwards, so a wrong guess is not repeated.
91
106
 
92
107
  ## The three rules
93
108
 
94
109
  1. **Nothing is lost.** The session file stays the single source of truth. pi-zip only changes the view sent to the model, never deletes anything. User messages, tool-call arguments, the system prompt, tool definitions and thinking are never rewritten. Every folded block carries a handle and shows its key lines (errors, ids, first and last line).
95
- 2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens in one step, and every request of that turn sends the same bytes. While the cache is warm it does nothing, with one exception, the warm valve: above the 40K target it applies that plan, minus the previous user turn (a warm edit never folds the turn you just finished, except when you come back after the declared TTL: there the previous turn is eligible exactly as at a cold return, so a provider whose cache outlives its TTL does not keep it at every return), when the edit pays for the rewrite it causes, and never on the request right after an edited one (no back-to-back warm rewrites). That is one inequality, r Δ²/(2g) + η Δ ≥ K with K = (w − r)(P T − (1 − P) Δ): the reads the removed Δ tokens would cost while the context grows back at g tokens per request (measured in the session), plus, near Pi's compaction trigger, what Pi would charge for the same room (η), against the rewrite of the T = A tokens left after the edit (pricing only the suffix after the earliest edit fires warm edits earlier and lost quality in the offline evaluation; the suffix is logged as `Tsuf` for measurement). r and w are the read and rewrite price ratios of the cache class, never the model's price table: explicit write premium 0.1 / 1.25 x input (2 x on the 1-hour tier), automatic prefix cache 0.2 / 1 x input; the class is read from the provider's usage reports, and until the first response the old fixed rule applies. P is the probability that the cache is still warm. A cold return is P = 0, so K < 0 and it always fires; a single small fold never pays at a warm cache, a large one does.
96
- 3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited, Esc stops the waiting, and the notice says so. If you return while the cache is still warm and the valve does not fire, the prepared summary is discarded (its cost is still counted). In non-interactive modes (`-p`, `--mode json`) nothing is ever started in the background: a cold return that needs a summary computes it right then.
110
+ 2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens of context in one step, system prompt and tool definitions included, and every request of that turn sends the same bytes. No edit can shrink the system prompt and the tools, so once they take more than half of the 40K (a large tool set) the target becomes them plus 20K of conversation: with a 50K tool set the context lands near 70K. It is never a 40K that no edit can reach (that folds everything and summarises at every return), and it does not keep a full 40K of conversation on top either (on a live bench with a 36K prefix that cost 13% more, with no quality difference measured). While the cache is warm it does nothing, with one exception, the warm valve: above that target it applies that plan, minus the previous user turn (a warm edit never folds the turn you just finished, except when you come back after the declared TTL: there the previous turn is eligible exactly as at a cold return, so a provider whose cache outlives its TTL does not keep it at every return; and it leaves alone what the last cold return had to protect, the previous turn's outputs of that return, until the next cold return, until the session has grown by more than about 11K tokens since, or until the context reaches Pi's compaction room: the window sliding by one prompt is not news, and folding them then would pay a warm rewrite for what the free rewrite at the cold return could not take), when the edit pays for the rewrite it causes, and never on the request right after an edited one (no back-to-back warm rewrites). That is one inequality, r Δ²/(2g) + η Δ ≥ K with K = (w − r)(P T − (1 − P) Δ): the reads the removed Δ tokens would cost while the context grows back at g tokens per request (measured in the session), plus, near Pi's compaction trigger, what Pi would charge for the same room (η), against the rewrite of the T = A tokens left after the edit (pricing only the suffix after the earliest edit fires warm edits earlier and lost quality in the offline evaluation; the suffix is logged as `Tsuf` for measurement). r and w are the read and rewrite price ratios of the cache class, never the model's price table: explicit write premium 0.1 / 1.25 x input (2 x on the 1-hour tier), automatic prefix cache 0.2 / 1 x input; the class is read from the provider's usage reports, and until the first response the old fixed rule applies. P is the probability that the cache is still warm. A cold return is P = 0, so K < 0 and it always fires; a single small fold never pays at a warm cache, a large one does.
111
+ 3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in, at a cold return and in the warm valve alike; the call is priced as what it is: an uncached read of the conversation since the previous summary, which code carries forward and the model never rewrites, plus the narrative it writes) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited (Esc cancels the prompt; the summary goes on and is used when you send it again), and the notice says so. A summary that finished while you were away goes into the session through Pi's own compaction before your prompt (Pi's `[compaction]` block then shows pi-zip's size before), in milliseconds; when Pi cannot take it there (busy, nothing it would compact, an older Pi, another extension taking part in Pi's compaction: see Coexistence), it rides along with the first request as before. If you return while the cache is still warm and the valve does not fire, the prepared summary is discarded (its cost is still counted). With Pi's own idle cache warming on (`"cacheWarming": "idle"` in the global settings, the only place Pi reads it), the summary waits for Pi's decision instead of the 0.8 mark: while Pi keeps refreshing the cache, nothing is prepared (you would come back to a warm cache and not need it); once Pi stops, or no refresh follows its decision, or the next decision (due one warming delay after the refresh finished) does not come, the summary is prepared as usual. pi-zip only watches those decisions, it never answers them. In non-interactive modes (`-p`, `--mode json`) nothing is ever started in the background: a cold return that needs a summary computes it right then.
97
112
 
98
- **The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip status` shows the class, the lifetime it currently believes, and the learned survival as the `alive` row.
113
+ **The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip` draws what it has learned as the dot chart above.
99
114
 
100
- Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves, are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
115
+ Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves (or reads of a recall file), are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
101
116
 
102
- **Token counts are calibrated, not guessed.** Sizes are estimated as chars/4, which undercounts real tokens (typically by about 1.7x in coding sessions). So the cold cap, the compaction room and the law's token counts are all compared against `k` x the estimate, where `k` = real tokens / estimated tokens for the newest assistant message that reports usage (input + cache read + cache write, over the estimate of the context that request carried; clamped to 1 to 2.5). `k` is read from the session itself on every decision, so a restart, `pi -p` or a resumed session calibrates exactly like a long-lived one, and nothing extra is stored. With no usage to read (a brand-new session, or a provider that reports none) `k` is 1.7. Sizes in the notices, `/zip status` and the ledger use the same scale.
117
+ **Token counts are calibrated, not guessed.** Sizes are estimated as chars/4 (Pi's rule), and real tokens are not proportional to that: every request carries a fixed prefix O (system prompt, tool definitions, the provider's own framing: 2-4K on a bare Pi, 30-55K with a large tool set) that no edit shrinks, and the conversation costs c real tokens per estimated one (about 0.7-1.1 on GLM and GPT, 1.5-2.2 on Claude, more on CJK-heavy text). So a size is O + c x the chars/4 estimate of the conversation. O is read from the session: what the provider still read from its cache on the request that first carried the newest summary (everything after the system prompt and tools had changed there), else the session's first request minus its small opening prompt, else 2 x the chars/4 of the system prompt and the tool definitions; a tool set that changed since moves it by the same factor. c = (real - O) / estimate for the newest reply that reports usage (input + cache read + cache write; clamped to 0.5-4, 1.5 before the first reply). Both depend on the provider's tokenizer, so only replies from the model the next request goes to count: right after a model switch O and c fall back to the defaults until its first reply, and when that model has no O of its own, the previous model's O carries over in conversation units (O / c). Both are read from the session itself on every decision, so a restart, `pi -p` or a resumed session calibrates exactly like a long-lived one, and nothing extra is stored. The conversation estimate also counts what every tool call puts on the wire besides its arguments and output (the call id, twice, and the tool-use markup), since that does not shrink when an output is folded. On recorded sessions the next request's size comes out within 2.8% (p90; 31% with the single ratio of 0.2.9). The size a plan predicts for the request that carries it is a few percent off in the median after a fold that removed over half of the context, 12-15% at worst one time in ten (what is left has not been measured yet); from the next reply on it is measured again (within 4%). The cold cap counts the whole context and leaves at least half of it to the conversation; the compaction room is the whole context too (it is Pi's trigger); sizes in the notices, `/zip` and the ledger use the same scale.
103
118
 
104
119
  The cold cap is kept below Pi's own compaction trigger (window minus `compaction.reserveTokens`), so on small windows Pi's lossy compaction does not get there first.
105
120
 
@@ -107,12 +122,12 @@ The cold cap is kept below Pi's own compaction trigger (window minus `compaction
107
122
 
108
123
  | Command | Effect |
109
124
  |---|---|
110
- | `/zip` or `/zip status` | the card above: state, cache class and lifetime, learned survival, this session's folds, summaries and recalls (counted from the session, so a restart does not reset them) |
125
+ | `/zip` | the card above: state, how long the cache survives you being away (a dot chart of your own returns), this session's folds, summaries and recalls (counted from the session, so a restart does not reset them) |
111
126
  | `/zip off` | strict no-op: no folds, no summaries, requests left untouched (earlier folds stay recallable) |
112
127
  | `/zip on` | resume |
113
128
  | `/zip quiet` | toggle the per-turn notice (folding continues) |
114
129
 
115
- `off` and `quiet` are remembered per session; each change is one line in the transcript. In `-p` / json mode `/zip status` prints plain text to stderr.
130
+ `off` and `quiet` are remembered per session; each change is one line in the transcript. In `-p` / json mode `/zip` prints plain text to stderr.
116
131
 
117
132
  ## Recall
118
133
 
@@ -127,7 +142,7 @@ key lines kept (original line numbers; up to 8):
127
142
  Original kept byte for byte, recallable even after summaries or compaction: zip_recall("k3f9a0x1qz") (optional grep/range) is instant, free, no side effects; prefer it to re-running or re-reading (output may differ). Do not guess its content.
128
143
  ```
129
144
 
130
- Placeholders are written once, when the fold is saved: sessions folded by an earlier version keep their old placeholder text unchanged. Summaries list each folded output with its handle, turn and outcome in the same way. A summary of an earlier summary merges it section by section (requests, files, commands, other calls, errors, handle table, handle index) instead of clipping its text: items are deduplicated, the oldest are dropped first under a per-section budget with a count of what was left out, and the previous narrative survives as a short tail excerpt. Pi's own free-text compaction summaries are carried as a head-and-tail excerpt.
145
+ Placeholders are written once, when the fold is saved: sessions folded by an earlier version keep their old placeholder text unchanged. Summaries list each folded output with its handle, turn and outcome in the same way, and say at the top that the 10-character code next to a command, file or call is its handle and that what an output said is to be recalled rather than re-run or re-read. A summary of an earlier summary merges it section by section (requests, files, commands, other calls, errors, handle table, handle index) instead of clipping its text: items are deduplicated, the oldest are dropped first under a per-section budget with a count of what was left out, and the previous narrative survives as a short tail excerpt. Pi's own free-text compaction summaries are carried as a head-and-tail excerpt.
131
146
 
132
147
  The model recalls by itself when it needs the content:
133
148
 
@@ -135,9 +150,10 @@ The model recalls by itself when it needs the content:
135
150
  zip_recall({ handles: ["k3f9a0x1qz", "m2b7c4d8ww"] })
136
151
  zip_recall({ handle: "k3f9a0x1qz", grep: "ERROR|FAIL" })
137
152
  zip_recall({ handle: "k3f9a0x1qz", range: "120-240" })
153
+ zip_recall({ grep: "ECONNREFUSED" }) // no handle: searches every earlier tool output
138
154
  ```
139
155
 
140
- Recall is batched, exact, and also works for outputs from before a compaction or a pi-zip summary (summaries carry a handle table and a budgeted index of older handles). Output comes back in pages of 20,000 characters; the page says how to continue:
156
+ Recall is batched, exact, and also works for outputs from before a compaction or a pi-zip summary (summaries carry a handle table and a budgeted index of older handles). Each recalled section is headed by the handle and the command or path that produced it, so a look-alike output (the same command run again, another slice of the same file) is easy to spot. Without a handle, `grep` searches every earlier tool output of the session (zip_recall results excepted, outputs from before a compaction included), newest output first, each match under its output's handle and command: this also reaches an output whose handle a long summary no longer lists. Output comes back in pages of 20,000 characters (8,000 for a search without a handle; a later page of a search covers the same outputs as its first page, so new output arriving in between does not shift it); the page says how to continue:
141
157
 
142
158
  ```
143
159
  zip_recall({ handle: "k3f9a0x1qz", offset: 20000 }) // next page, by characters
@@ -146,17 +162,31 @@ zip_recall({ handle: "k3f9a0x1qz", offset: 20000, limit: 50000 }) // up to 50,00
146
162
 
147
163
  Paging is by characters, so even one 45,000-character line can be read in full, and the pages add up to the original byte for byte. `offset` and `limit` also page a `grep` or `range` selection. `grep` is a case-insensitive regular expression; patterns that could backtrack badly (nested quantifiers, quantified alternation, back-references, very long patterns) are searched as literal text instead.
148
164
 
165
+ ### Recall through files
166
+
167
+ When a tool allowlist hides `zip_recall` (`pi --tools read,grep,bash`, sub-agent launchers) but `read`, `grep` or `bash` is active, pi-zip folds exactly as in full mode and points the model at a file instead of the tool:
168
+
169
+ ```
170
+ Original kept byte for byte, even after summaries or compaction, in ~/.cache/pi-zip/recall/01a120cb-3d40-75a6-aeac-320e71967606/k3f9a0x1qz.txt: read it (offset/limit for pages) or grep it; prefer that to re-running or re-reading the source (output may differ). Do not guess its content.
171
+ ```
172
+
173
+ The sentence names only the tools that are active (with bash alone: `sed -n` to page, `grep -n` to search), and summaries name the same files. A file does not exist until a tool asks for it: when `read`, `grep`, `find` or `ls` gets such a path (whatever comes before `pi-zip/recall/`: another machine, another user, a forked session), or a `bash` command contains one, pi-zip writes the original from the session (the current branch) and points the call at this machine's copy. A `grep`, `ls` or `find` of the session's directory writes every output the model no longer sees in full (folded or summarized away), so a grep over the directory searches them all. A read of a recall file is a recall: it is never folded and counts as an unzip, and the transcript draws it like zip_recall's row (`ƶ unzip …`). `/zip` says `on` (it is a full mode) and has one more row: `unzip from files in ~/.cache/pi-zip/recall (a tool allowlist hides zip_recall)`, naming where the copies really are (see below; a narrow terminal drops the part in brackets). In a bash command the path is recognised as its own word, also after `X=`, `--file=` or inside `{a,b}`.
174
+
175
+ The copies live in `$XDG_CACHE_HOME/pi-zip/recall/<session id>/` (by default `~/.cache/pi-zip/recall/`, a place backups and dotfile sync conventionally skip; a relative `XDG_CACHE_HOME` is ignored, as the XDG spec says, since it would put them in the project folder): directories 0700, files 0600. The state line and `/zip` show the place in use. The same text is already in the session file, which Pi writes 0644. Only what the model asks for is copied, and every copy can be written again from the session at any time, so copies older than 14 days are deleted at the next session start; that cleanup touches nothing but pi-zip's own copies (`<handle>.txt` files in session-id directories, never through a symlink). A copy that no longer holds the original byte for byte (a bash command appended to it or edited it) is written again the next time a tool names it. If you delete a session file to get rid of a secret it contained, delete `<that place>/<its session id>/` too, or it stays until then. Without a session file (`pi --no-session`) nothing is copied: the outputs are in no file, so pi-zip keeps them off the disk too, folds only re-readable outputs and says so (`reread-only`). Without pi-zip loaded nothing is written, just as zip_recall would be missing.
176
+
149
177
  ## Coexistence
150
178
 
151
- If another extension that manages context is loaded (for example billion-context, magic-context, pi-smart-compact, pi-hot-compact, pi-context-prune), pi-zip pauses folding, says so once, and keeps only its request guard. Two writers on the same view give unpredictable results. `/zip status` shows `paused`.
179
+ If another extension that manages context is loaded (for example billion-context, magic-context, pi-smart-compact, pi-hot-compact, pi-context-prune), pi-zip pauses folding, says so once, and keeps only its request guard. Two writers on the same view give unpredictable results. `/zip` shows `paused`.
180
+
181
+ An extension that only writes Pi's compaction summaries (a `session_before_compact` handler, like Pi's own `custom-compaction` example) does not pause pi-zip, but Pi runs its handler for every compaction, the one that puts pi-zip's prepared summary in at a cold return included. So pi-zip watches that compaction: if another extension's summary went in instead of pi-zip's, pi-zip adds none of its own on top (its prepared summary is discarded, its cost still counted); if the compaction is not done within 2 s (that handler is writing a summary of its own), pi-zip cancels it before your prompt goes on, so nothing can land in the middle of the run, and sends its summary with the first request instead; and if it was cancelled, or took longer than pi-zip's own part ever does (250 ms), pi-zip stops going through Pi's compaction for the rest of the process, so that handler's work is not paid for and waited for at every return. When such an extension's summary is already in the session, pi-zip does not go through Pi's compaction at all. Either way, summaries then ride along with the first request, as on an older Pi. A handler that cancels manual compactions shows Pi's red "Compaction cancelled" line at most once per process.
152
182
 
153
183
  The guard has two parts. Before saving a fold or a summary it checks that the edit itself would not leave a tool result without its tool call (failed or aborted assistant turns, which the provider layer drops anyway, are ignored, as is any oddity the session already had). And before a request is sent it repairs a tool result that lost its tool call, instead of letting the provider reject the request and brick the session. The repair understands the request shapes of Pi's Anthropic, OpenAI chat completions (and Mistral), OpenAI Responses (Azure, Codex), Google Gemini/Vertex and Bedrock Converse providers; a request of any other shape is passed through untouched.
154
184
 
155
185
  ## FAQ
156
186
 
157
- **Will it save money?** Mostly on cold returns, which is where a long session pays for a full cache rewrite. While the cache is warm it edits only when the inequality above says the rewrite pays back (large contexts, near Pi's compaction trigger, outputs 60+ requests old). `/zip status` shows what this session has folded (and roughly how many tokens that removed), its summaries with what their model calls cost, and its recalls. The numbers live in the session file, so a restart does not reset them.
187
+ **Will it save money?** Mostly on cold returns, which is where a long session pays for a full cache rewrite. While the cache is warm it edits only when the inequality above says the rewrite pays back (large contexts, near Pi's compaction trigger, outputs 60+ requests old). `/zip` shows what this session has folded (and roughly how many tokens that removed), its summaries with what their model calls cost, and its recalls. The numbers live in the session file, so a restart does not reset them.
158
188
 
159
- **Does it cost extra?** Planning is free. A summary is one extra model call (the current model, no tools, no prompt cache), shown in `/zip status`. A background summary you never use (you came back while the cache was warm) is counted there too. Recalled content re-enters the context at normal prices.
189
+ **Does it cost extra?** Planning is free. A summary is one extra model call (the current model, no tools, no prompt cache), shown in `/zip`. A background summary you never use (you came back while the cache was warm) is counted there too. Recalled content re-enters the context at normal prices.
160
190
 
161
191
  **Can the model lose information?** Folded outputs are replaced by a placeholder with key lines and a handle, and the placeholder tells the model not to guess. Summaries quote your requests verbatim and never paraphrase the model's reasoning. If the model ignores the handle and guesses, that is a model failure pi-zip cannot catch; this is why only re-readable outputs of the previous turn are folded at a cold return; other outputs of the protected turns wait until they are 60 requests old.
162
192
 
@@ -177,7 +207,7 @@ bun run build-check # bun
177
207
 
178
208
  `test/pi-integration.test.ts` runs the extension through Pi's real session manager, extension loader and `emitContext`, with a mid-conversation system update and an aborted tool-call turn in the session, and checks that the request is byte-identical before and after turn_end persists the edits.
179
209
 
180
- Environment overrides exist for tests only and are not part of the product surface: `PI_ZIP_TTL_SECS` (cache TTL), `PI_ZIP_COLD_CAP` (fold target in tokens, default 40000), `PI_ZIP_FOLD_MIN` (smallest output worth folding, default 500 tokens), `PI_ZIP_KEEP_LINES`, `PI_ZIP_MIN_GAIN` (summary gain floor of the legacy rule used only when nothing is known about prices, default 10000), `PI_ZIP_INTURN_AGE` (age in assistant requests from which any output of the protected turns may fold, default 60; 0 = never), `PI_ZIP_CACHE_STATS=<path>` (where the learned cache survival lives; benches isolate it), `PI_ZIP_OFF=1` (register nothing), `PI_ZIP_LEDGER=<path>` (append a JSON line per decision; `prompt` records `ttlMs` and `ttlSource`, `declared` or `observed`; including the summary call's token usage; `cold_plan` records the calibration as `k`, `calReal`, `calEst`, `calSource`, next to the scaled `ctxBefore` and `ctxAfter`; `prompt` also records `gapS`, `pWarm`, `survSrc` and `cls`; `law` records every evaluation with `g`, `wr` (w/r), `pWarm`, `survSrc`, `prSrc` and per step `B`, `A`, `T` (= A), `Tsuf` (the suffix after the earliest edit, measurement only), `phi`, `K`, `eta`; `b2b_skip` marks a warm edit skipped right after an edited request; `cache_sample` records each learned survival observation).
210
+ Environment overrides exist for tests only and are not part of the product surface: `PI_ZIP_TTL_SECS` (cache TTL), `PI_ZIP_COLD_CAP` (fold target in real tokens of the whole context, default 40000; at least half of it is always left to the conversation above the system prompt and tools), `PI_ZIP_FOLD_MIN` (smallest output worth folding, default 500 tokens), `PI_ZIP_KEEP_LINES`, `PI_ZIP_MIN_GAIN` (summary gain floor of the legacy rule used only when nothing is known about prices, default 10000), `PI_ZIP_INTURN_AGE` (age in assistant requests from which any output of the protected turns may fold, default 60; 0 = never), `PI_ZIP_CACHE_STATS=<path>` (where the learned cache survival lives; benches isolate it), `PI_ZIP_OFF=1` (register nothing), `PI_ZIP_NATIVE_MS` (how long pi-zip's own compaction may take before it is cancelled, default 2000), `PI_ZIP_LEDGER=<path>` (append a JSON line per decision; `prompt` records `ttlMs` and `ttlSource`, `declared` or `observed`; including the summary call's token usage; `cold_plan` records the calibration as `O`, `c`, `oSource`, `calReal`, `calEst`, `calSource`, next to the real `ctxBefore` and `ctxAfter`; `prompt` also records `gapS`, `pWarm`, `survSrc` and `cls`; `law` records every evaluation with `g`, `wr` (w/r), `pWarm`, `survSrc`, `prSrc` and per step `B`, `A`, `T` (= A), `Tsuf` (the suffix after the earliest edit, measurement only), `phi`, `K`, `eta`; `b2b_skip` marks a warm edit skipped right after an edited request; `cache_sample` records each learned survival observation).
181
211
 
182
212
  Internally, `RELAX_PREV_TURN` in `src/plan.ts` selects whether re-readable outputs of the previous user turn may be folded on a cold return or a return after the declared TTL (default `true`; other warm plans never fold them).
183
213
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-zip",
3
- "version": "0.2.8",
3
+ "version": "0.3.0-rc.2",
4
4
  "description": "Keeps long Pi sessions cheap without losing anything: folds old tool output only when the prompt cache has already gone cold, and every fold can be recalled byte for byte. Zero config.",
5
5
  "type": "module",
6
6
  "license": "MIT",
package/src/cache.ts CHANGED
@@ -44,12 +44,34 @@ export function observedTier(branch: Any[], key: string): "1h" | "5m" | undefine
44
44
 
45
45
  export interface TtlInfo { ms: number; source: "declared" | "observed"; note?: string }
46
46
 
47
- /** Declared TTL, corrected by what the provider was seen to write. PI_ZIP_TTL_SECS wins. `note` = the one-time mismatch notice text. */
48
- export function resolveTtl(model: Any, branch: Any[] = []): TtlInfo {
47
+ /** The cache tier a request asked for, read from the payload Pi built (null: the payload does not say). Anthropic: a 1-hour
48
+ * cache_control ttl is the long tier, cache_control without one the short tier; OpenAI: prompt_cache_retention "24h" is long. */
49
+ export function payloadRetention(p: Any): "short" | "long" | null {
50
+ if (!p || typeof p !== "object") return null;
51
+ if (typeof p.prompt_cache_retention === "string") return p.prompt_cache_retention === "24h" ? "long" : "short";
52
+ let seen = false;
53
+ const one = (x: Any): boolean => {
54
+ const cc = x?.cache_control;
55
+ if (!cc || typeof cc !== "object") return false;
56
+ seen = true;
57
+ return cc.ttl === "1h";
58
+ };
59
+ for (const list of [p.system, p.tools]) if (Array.isArray(list)) for (const b of list) if (one(b)) return "long";
60
+ if (Array.isArray(p.messages)) for (const m of p.messages) {
61
+ if (one(m)) return "long";
62
+ if (Array.isArray(m?.content)) for (const b of m.content) if (one(b)) return "long";
63
+ }
64
+ return seen ? "short" : null;
65
+ }
66
+
67
+ /** Declared TTL, corrected by what the provider was seen to write. PI_ZIP_TTL_SECS wins. `note` = the one-time mismatch notice text.
68
+ * `tier` = the retention the request used (the provider-scoped PI_CACHE_RETENTION, or the request payload itself; run.ts); without it,
69
+ * the process environment decides, as before. */
70
+ export function resolveTtl(model: Any, branch: Any[] = [], tier?: "short" | "long"): TtlInfo {
49
71
  if (process.env.PI_ZIP_TTL_SECS) return { ms: envInt("TTL_SECS", 300) * 1000, source: "declared" };
50
72
  const pc = model?.promptCache;
51
- const declared = cacheTtlMs(pc);
52
- const long = process.env.PI_CACHE_RETENTION === "long";
73
+ const long = tier ? tier === "long" : process.env.PI_CACHE_RETENTION === "long";
74
+ const declared = cacheTtlMs(pc, DEFAULT_TTL_MS, { cacheRetention: long ? "long" : "short" });
53
75
  const seen = declared > 0 ? observedTier(branch, modelKey(model)) : undefined;
54
76
  const mismatch = (req: string, got: string, ms: number): TtlInfo => ({ ms, source: "observed", note: `${PRODUCT} · requested ${req} prompt cache, provider wrote ${got} · using ${got}` });
55
77
  if (long && seen === "5m") return mismatch("1h", "5m", cacheTtlMs(pc, DEFAULT_TTL_MS, { cacheRetention: "short" }));
@@ -57,7 +79,7 @@ export function resolveTtl(model: Any, branch: Any[] = []): TtlInfo {
57
79
  return { ms: declared, source: "declared" };
58
80
  }
59
81
 
60
- export const ttlFor = (model: Any, branch: Any[] = []): number => resolveTtl(model, branch).ms;
82
+ export const ttlFor = (model: Any, branch: Any[] = [], tier?: "short" | "long"): number => resolveTtl(model, branch, tier).ms;
61
83
 
62
84
  /** Cold = a prior request exists and nothing touched the cache for longer than the TTL. No prior request counts as warm. */
63
85
  export const isColdByTtl = (lastActivityMs: number, nowMs: number, ttlMs: number): boolean => lastActivityMs > 0 && nowMs - lastActivityMs > ttlMs;
@@ -69,18 +91,33 @@ const tsOf = (en: Any, inner?: Any): number => {
69
91
  return Number.isFinite(t) ? t : 0;
70
92
  };
71
93
 
94
+ /** A response that reached no cache: an error that reports no usage (the provider refused the request). One that reports usage was
95
+ * read: it touched the cache. An abort (Esc) comes after the provider read and wrote the prompt: it is a touch, with or without usage. */
96
+ export const refused = (m: Any): boolean => {
97
+ if (m?.role !== "assistant" || m.stopReason !== "error") return false;
98
+ const u = m.usage;
99
+ return !((Number(u?.input) || 0) + (Number(u?.cacheRead) || 0) + (Number(u?.cacheWrite) || 0) > 0);
100
+ };
101
+
72
102
  /**
73
103
  * Last time anything touched the cache, according to the branch: the newest message, or the newest cache-warm usage entry
74
104
  * (Pi's own idle refresh persists `{ type: "usage", kind: "cache_warm" }`, and a refresh keeps the entry alive for another TTL).
105
+ * A request the provider refused (an error with no usage) touched nothing, and neither did the messages that made it (the prompt or
106
+ * the tool results it carried, written before it failed): the clock stays at the last request that was read.
75
107
  * Fresh process: err on the warm side (0 = unknown).
76
108
  */
77
109
  export function lastMessageMs(branch: Any[]): number {
78
- let last = 0;
110
+ let last = 0, pending = 0; // pending: the newest non-assistant message since the last response, counted once its response is known
79
111
  for (const en of branch) {
80
- if (en?.type === "message" && en.message) last = Math.max(last, tsOf(en, en.message));
81
- else if (en?.type === "usage" && en.kind === "cache_warm") last = Math.max(last, tsOf(en));
112
+ if (en?.type === "message" && en.message) {
113
+ const t = tsOf(en, en.message);
114
+ if (en.message.role === "assistant") {
115
+ if (!refused(en.message)) last = Math.max(last, pending, t);
116
+ pending = 0;
117
+ } else pending = Math.max(pending, t);
118
+ } else if (en?.type === "usage" && en.kind === "cache_warm") last = Math.max(last, tsOf(en));
82
119
  }
83
- return last;
120
+ return Math.max(last, pending); // trailing messages with no response yet (the run in progress) count, as before
84
121
  }
85
122
 
86
123
  /** The newest assistant message that really ran (errors and aborts may never have reached a cache): its provider/model and prompt size. */
@@ -102,11 +139,12 @@ export interface ColdInfo { cold: boolean; reason: string; ttl: TtlInfo; pWarm:
102
139
 
103
140
  /** Cold by time, or because the model changed: a cache entry belongs to one provider and model. `learned` = P(warm) after a gap from the
104
141
  * survival the provider was seen to have (learn.ts); without it, or under PI_ZIP_TTL_SECS (tests), the TTL decides. cold = P(warm) < 0.5. */
105
- export function detectCold(model: Any, lastReqMs: number, branch: Any[], nowMs = Date.now(), lastModel = "", learned?: Survival): ColdInfo {
106
- const ttl = resolveTtl(model, branch);
142
+ export function detectCold(model: Any, lastReqMs: number, branch: Any[], nowMs = Date.now(), lastModel = "", learned?: Survival, tier?: "short" | "long"): ColdInfo {
143
+ const ttl = resolveTtl(model, branch, tier);
107
144
  const last = Math.max(lastReqMs, lastMessageMs(branch));
108
145
  const prev = lastModel || lastModelInBranch(branch);
109
- if (last && prev && modelKey(model) && prev !== modelKey(model)) return { cold: true, reason: `model switch ${prev} -> ${modelKey(model)}`, ttl, pWarm: 0, src: "model switch", gapS: null, pastTtl: false };
146
+ // a request that went out, even one the provider refused, was made on another model: the new model's cache is not that one
147
+ if (prev && modelKey(model) && prev !== modelKey(model)) return { cold: true, reason: `model switch ${prev} -> ${modelKey(model)}`, ttl, pWarm: 0, src: "model switch", gapS: null, pastTtl: false };
110
148
  if (!last) return { cold: false, reason: "no prior request", ttl, pWarm: 1, src: "no prior request", gapS: null, pastTtl: false };
111
149
  const gapS = (nowMs - last) / 1000;
112
150
  const byTtl = isColdByTtl(last, nowMs, ttl.ms);
package/src/guard.ts CHANGED
@@ -1,5 +1,5 @@
1
1
  // Guard (F13, I5): validate an edit set before committing it, and repair orphan tool results in the outgoing payload.
2
- import { RELAX_PREV_TURN, settings, type Block } from "./plan.ts";
2
+ import { finalAnswerIdx, RELAX_PREV_TURN, settings, type Block } from "./plan.ts";
3
3
  import type { Any } from "./util.ts";
4
4
 
5
5
  const PROTECT_USER_TURNS = 2;
@@ -10,6 +10,7 @@ export interface EditSet {
10
10
  recover?: Map<number, string | undefined>; // block idx -> "rereadable" | "nonrereadable" (needed for folds in the previous user turn)
11
11
  relax?: boolean; // default RELAX_PREV_TURN
12
12
  inturnAge?: number; // default settings().inturnAge: an output (any class) at least this many assistant requests old may fold in a protected turn (0 = never)
13
+ ahead?: boolean; // a lookahead cut (plan.ts LOOKAHEAD): it may lie inside the previous user turn, never past that turn's final answer
13
14
  }
14
15
 
15
16
  /**
@@ -63,7 +64,7 @@ export function validateEdits(blocks: Block[], plan: EditSet, userTurns: number)
63
64
  }
64
65
  if (plan.cut !== null) {
65
66
  if (plan.cut < 1 || plan.cut >= blocks.length) return `bad cut ${plan.cut}`;
66
- if (limit >= 0 && plan.cut > limit) return `cut ${plan.cut} inside the last ${PROTECT_USER_TURNS} user turns`;
67
+ if (limit >= 0 && plan.cut > limit && !(plan.ahead && plan.cut <= finalAnswerIdx(blocks, userTurns - 1))) return `cut ${plan.cut} inside the last ${PROTECT_USER_TURNS} user turns`;
67
68
  const kb = blocks[plan.cut];
68
69
  if (!kb.entryId || (kb.kind !== "user" && kb.kind !== "assistant")) return `cut ${plan.cut} is not at a user/assistant message`;
69
70
  // only what the cut itself would create is our problem; a session that was already odd stays as odd as it was
package/src/index.ts CHANGED
@@ -4,8 +4,9 @@ import type { Any } from "./util.ts";
4
4
  import { type NoticeData, registerZipCommand, renderNotice } from "./notice.ts";
5
5
  import { RECALL_TOOL } from "./placeholder.ts";
6
6
  import { registerRecallTool } from "./recall.ts";
7
+ import { parseRecallPath } from "./recallfile.ts";
7
8
  import { NOTICE_CUSTOM, Zip } from "./run.ts";
8
- import { markFirstLine, measure, renderCard, renderState, setMeasure } from "./ui.ts";
9
+ import { markFirstLine, measure, recallCallLine, renderCard, renderState, setMeasure } from "./ui.ts";
9
10
 
10
11
  export default function piZip(pi: ExtensionAPI) {
11
12
  if (process.env.PI_ZIP_OFF === "1") return; // test only: behave exactly as if not installed (registers nothing)
@@ -14,7 +15,7 @@ export default function piZip(pi: ExtensionAPI) {
14
15
  registerZipCommand(pi, zip);
15
16
  try {
16
17
  // Everything pi-zip shows lives in the transcript as custom entries (never part of the model's context): fold/summary
17
- // notices, state lines and the /zip status card. Widths are measured with Pi's own pi-tui once it has loaded.
18
+ // notices, state lines and the /zip card. Widths are measured with Pi's own pi-tui once it has loaded.
18
19
  import("@earendil-works/pi-tui").then((m: Any) => {
19
20
  if (typeof m?.visibleWidth === "function" && typeof m?.truncateToWidth === "function")
20
21
  setMeasure({ vw: m.visibleWidth, cut: (s: string, w: number) => (m.visibleWidth(s) <= w ? s : w <= 0 ? "" : m.truncateToWidth(s, w, "…")) });
@@ -38,7 +39,9 @@ export default function piZip(pi: ExtensionAPI) {
38
39
  if (d.v !== 2) return [theme.fg("dim", measure.cut(d.text, width))];
39
40
  if (d.kind === "state") return renderState(d.word, d.reason ?? "", width, theme, measure);
40
41
  if (d.kind === "card" && d.card) return renderCard(d.card, width, theme, measure);
41
- return renderNotice(d as NoticeData, open(), width, theme, measure.vw);
42
+ // once the request that carried the plan has been answered, "after" is its real size (older notices keep their number)
43
+ const real = typeof d.reqTs === "number" ? zip.realSizeAfter(d.reqTs) : undefined;
44
+ return renderNotice((real ? { ...d, after: real } : d) as NoticeData, open(), width, theme, measure.vw);
42
45
  } catch {
43
46
  return [];
44
47
  }
@@ -52,13 +55,21 @@ export default function piZip(pi: ExtensionAPI) {
52
55
  };
53
56
  });
54
57
  zip.entryRenderer = typeof pi.registerEntryRenderer === "function";
55
- // a folded output keeps its place in the transcript; its call row gets a dim "▸ folded · <handle>" on the right
58
+ // a folded output keeps its place in the transcript; its call row gets "ƶ zipped" right after its own text
56
59
  pi.registerToolRenderer?.((toolName: string, next: () => Any) => {
57
60
  const base = next();
58
61
  if (toolName === RECALL_TOOL || typeof base?.renderCall !== "function") return base;
59
62
  return {
60
63
  ...base,
61
64
  renderCall: (args: Any, theme: Any, c: Any) => {
65
+ // file recall: a read or grep of a recall file (or of the session's recall directory) reads like zip_recall's row
66
+ const ref = toolName === "read" || toolName === "grep" ? parseRecallPath(args?.path) : null;
67
+ if (ref && (ref.handle || toolName === "grep")) {
68
+ const off = typeof args.offset === "number" ? args.offset : undefined, lim = typeof args.limit === "number" ? args.limit : undefined;
69
+ const range = toolName === "read" && (off !== undefined || lim !== undefined) ? `${off ?? 1}-${lim !== undefined ? (off ?? 1) + lim - 1 : ""}` : undefined;
70
+ const a = { handle: ref.handle, grep: toolName === "grep" && typeof args.pattern === "string" ? args.pattern : undefined, range };
71
+ return { render: (width: number) => recallCallLine(a, Math.max(1, width), theme, measure, (h) => zip.labelOf(h)), invalidate: () => {} };
72
+ }
62
73
  const comp = base.renderCall(args, theme, c);
63
74
  const id = c?.toolCallId;
64
75
  if (!comp || typeof comp.render !== "function" || typeof id !== "string") return comp;
@@ -83,11 +94,15 @@ export default function piZip(pi: ExtensionAPI) {
83
94
  /* older Pi: notices fall back to the status line */
84
95
  }
85
96
  // A tool allowlist (`pi --tools read,bash`, sub-agent launchers) replaces the whole selection and Pi then does not even register
86
- // zip_recall; folds of outputs the model could not get back would be lost, so only rereadable outputs fold then (plan rereadOnly).
97
+ // zip_recall; folds of outputs the model could not get back would be lost, so only rereadable outputs fold then (plan rereadOnly)...
98
+ // ... unless read, grep or bash is active: then folds are as in full mode and placeholders name a recall file (recallfile.ts) that the
99
+ // tool_call hook below writes from the session when a tool asks for it.
87
100
  const checkRecall = (ctx: Any) => {
88
101
  try {
89
102
  const active = pi.getActiveTools?.();
90
- if (Array.isArray(active)) zip.setRecallOk(active.includes(RECALL_TOOL), ctx);
103
+ if (!Array.isArray(active)) return;
104
+ const tools = { read: active.includes("read"), grep: active.includes("grep"), bash: active.includes("bash") };
105
+ zip.setRecallMode(active.includes(RECALL_TOOL) ? "tool" : tools.read || tools.grep || tools.bash ? "file" : "reread", tools, ctx);
91
106
  } catch {
92
107
  /* never in the way */
93
108
  }
@@ -97,17 +112,24 @@ export default function piZip(pi: ExtensionAPI) {
97
112
  checkRecall(ctx);
98
113
  zip.refreshFolded(ctx);
99
114
  zip.welcome(ctx);
115
+ zip.resumeAway(ctx); // a resumed session: arm the away timer from its last request, as the settle after it would have
100
116
  return r;
101
117
  });
102
- pi.on("before_agent_start", (_e, ctx) => {
118
+ pi.on("before_agent_start", async (e, ctx) => {
103
119
  checkRecall(ctx);
104
120
  zip.refreshFolded(ctx); // the branch may have changed (/tree, fork)
105
- return zip.beforeAgentStart(ctx);
121
+ await zip.beforeAgentStart(ctx, e); // may persist a summary prepared while the user was away through Pi's compaction (Pi is idle here)
122
+ return undefined;
106
123
  });
124
+ // answers only the compaction pi-zip itself just asked for; every other compaction (Pi's, /compact, other extensions') gets undefined
125
+ pi.on("session_before_compact", (e) => zip.sessionBeforeCompact(e));
107
126
  pi.on("context_with_system", (e, ctx) => zip.context(e, ctx)); // the complete transcript: system messages stay where Pi put them
108
127
  pi.on("turn_end", (e, ctx) => zip.turnEnd(e, ctx));
109
128
  pi.on("agent_before_settle", (e, ctx) => zip.settle(e, ctx));
110
129
  pi.on("before_provider_request", (e) => zip.providerRequest(e));
130
+ pi.on("tool_call", (e, ctx) => zip.toolCall(e, ctx)); // a recall file named by a built-in tool: this machine's copy, written on demand
131
+ // Pi's idle cache warming: observed only (the return value is undefined, so Pi's own decision stands); an older Pi never fires it
132
+ pi.on("cache_warming_decision", (e) => zip.cacheWarmingDecision(e));
111
133
  pi.on("message_end", (e) => zip.messageEnd(e.message));
112
134
  pi.on("session_shutdown", () => zip.shutdown());
113
135
  }
package/src/learn.ts CHANGED
@@ -26,8 +26,10 @@ export function lawPrices(cls: CacheClass | undefined, longTier: boolean): Price
26
26
  }
27
27
 
28
28
  // ---- cache survival ------------------------------------------------------------------------------------------------
29
- /** Gap bins (s): [GAP_EDGES[i], GAP_EDGES[i+1]). Edges sit on the known TTL tiers (300 s, 3600 s): a deterministic TTL never splits a bin. */
30
- export const GAP_EDGES = [30, 60, 120, 180, 240, 300, 330, 360, 420, 480, 600, 900, 1200, 1800, 2700, 3600, 5400];
29
+ /** Gap bins (s): [GAP_EDGES[i], GAP_EDGES[i+1]). Edges sit on the known TTL tiers (300 s, 3600 s): a deterministic TTL never splits a bin.
30
+ * The last edge closes the last bin: a gap at or beyond it (2 h) is outside what was ever observed, it is not recorded and it is planned
31
+ * from the declared TTL (a cache seen alive at 110 min says nothing about 156 min). */
32
+ export const GAP_EDGES = [30, 60, 120, 180, 240, 300, 330, 360, 420, 480, 600, 900, 1200, 1800, 2700, 3600, 5400, 7200];
31
33
  export const HALF_LIFE = 16; // observations per bin: old evidence counts half after 16 newer ones in the same bin (a provider may change its TTL)
32
34
  // The declared TTL as pseudo-observations. Asymmetric because the signal is: a read of the re-sent prefix cannot happen on a dead cache,
33
35
  // but warm misses do (GLM 4-9%, glm-flash ~22%, research round 5). Beyond the TTL one clean read overrides it (weight 1/4: one hit -> 0.8);
@@ -41,10 +43,17 @@ export const MIN_EXPECT = 8192;
41
43
  export interface Entry { cls?: CacheClass; bins: Record<string, [number, number]>; n: number } // bin index -> [alive, dead] (decayed counts)
42
44
  interface File { v: 1; models: Record<string, Entry> }
43
45
 
44
- export const statsPath = (): string =>
45
- process.env.PI_ZIP_CACHE_STATS || join(process.env.PI_CODING_AGENT_DIR || join(homedir(), ".pi", "agent"), "pi-zip", "cache-survival.json");
46
+ /** Pi's agent directory (settings.json, sessions): PI_CODING_AGENT_DIR, else ~/.pi/agent. */
47
+ export const agentDir = (): string => process.env.PI_CODING_AGENT_DIR || join(homedir(), ".pi", "agent");
46
48
 
47
- export const binOf = (gapS: number): number => GAP_EDGES.findLastIndex((e) => gapS >= e); // -1 = below 30 s: not recorded
49
+ export const statsPath = (): string => process.env.PI_ZIP_CACHE_STATS || join(agentDir(), "pi-zip", "cache-survival.json");
50
+
51
+ /** ES2023's Array.prototype.findLastIndex (Node >= 18 and Bun have it); tsconfig's lib is ES2022, so it is typed here (types only). */
52
+ interface FindLastIndex { findLastIndex(predicate: (value: number, index: number) => boolean): number }
53
+ export const binOf = (gapS: number): number => {
54
+ const i = (GAP_EDGES as number[] & FindLastIndex).findLastIndex((e) => gapS >= e);
55
+ return i >= GAP_EDGES.length - 1 ? -1 : i; // -1 = below the first edge or at/above the last: not recorded
56
+ };
48
57
  const fin = (x: Any) => typeof x === "number" && Number.isFinite(x) && x >= 0;
49
58
 
50
59
  /** The persisted stats; a missing, unreadable or malformed file is ignored (empty). */
@@ -55,7 +64,7 @@ export function loadStats(path = statsPath()): File {
55
64
  if (j?.v !== 1 || typeof j.models !== "object" || !j.models) return out;
56
65
  for (const [k, e] of Object.entries<Any>(j.models)) {
57
66
  const bins: Entry["bins"] = {};
58
- for (const [b, c] of Object.entries<Any>(e?.bins ?? {})) if (GAP_EDGES[Number(b)] !== undefined && Array.isArray(c) && fin(c[0]) && fin(c[1])) bins[b] = [c[0], c[1]];
67
+ for (const [b, c] of Object.entries<Any>(e?.bins ?? {})) if (Number(b) < GAP_EDGES.length - 1 && GAP_EDGES[Number(b)] !== undefined && Array.isArray(c) && fin(c[0]) && fin(c[1])) bins[b] = [c[0], c[1]];
59
68
  out.models[k] = { bins, n: fin(e?.n) ? e.n : 0, ...(e?.cls === "explicit" || e?.cls === "automatic" ? { cls: e.cls } : {}) };
60
69
  }
61
70
  return out;
@@ -125,7 +134,8 @@ export function curve(e: Entry | undefined, priorS: number): { bin: number; p: n
125
134
  * between the nearest observed bins below and above (an alive read at 365 s makes every shorter gap warm; a miss makes longer ones dead). */
126
135
  export function pWarm(e: Entry | undefined, gapS: number, priorS: number): { p: number; src: "prior" | "learned" } {
127
136
  const prior = gapS <= priorS ? 1 : 0;
128
- const c = curve(e, priorS), b = binOf(gapS);
137
+ // beyond the last edge: no bin, the declared TTL decides (clamped by the newest bin below, as for an unobserved bin)
138
+ const c = curve(e, priorS), b = gapS >= GAP_EDGES[GAP_EDGES.length - 1] ? GAP_EDGES.length - 1 : binOf(gapS);
129
139
  const at = c.find((x) => x.bin === b);
130
140
  if (at) return { p: at.p, src: "learned" };
131
141
  const hi = c.filter((x) => x.bin < b).at(-1)?.p ?? 1, lo = c.find((x) => x.bin > b)?.p ?? 0;
@@ -133,7 +143,7 @@ export function pWarm(e: Entry | undefined, gapS: number, priorS: number): { p:
133
143
  return { p, src: p === prior ? "prior" : "learned" };
134
144
  }
135
145
 
136
- /** /zip status text: class and the learned curve ("360-420s 0.80 n1"). */
146
+ /** /zip text: class and the learned curve ("360-420s 0.80 n1"). */
137
147
  export function describe(e: Entry | undefined, priorS: number): string {
138
148
  const c = curve(e, priorS);
139
149
  const bins = c.map((x) => `${GAP_EDGES[x.bin]}-${GAP_EDGES[x.bin + 1] ?? "∞"}s ${x.p.toFixed(2)} n${Math.round(x.n * 10) / 10}`).join(", ");