pi-zip 0.2.8 → 0.3.0-rc.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +72 -42
- package/package.json +1 -1
- package/src/cache.ts +50 -12
- package/src/guard.ts +3 -2
- package/src/index.ts +30 -8
- package/src/learn.ts +18 -8
- package/src/notice.ts +91 -33
- package/src/placeholder.ts +12 -4
- package/src/plan.ts +470 -107
- package/src/recall.ts +111 -27
- package/src/recallfile.ts +165 -0
- package/src/run.ts +807 -122
- package/src/state.ts +53 -0
- package/src/summary.ts +22 -10
- package/src/ui.ts +214 -46
- package/src/util.ts +62 -2
package/README.md
CHANGED
|
@@ -30,76 +30,91 @@ Known limits:
|
|
|
30
30
|
|
|
31
31
|
- **GLM** (automatic prefix cache that outlives its declared 5 minutes): v0.1 cost about 1.2× BC live. v0.2 learns the real cache lifetime and folds the previous turn after the declared TTL (offline 0.95–0.97× v0.1), but this was not verified live.
|
|
32
32
|
- **Long autonomous runs** (one prompt, hundreds of tool calls, e.g. sub-agents): roughly on par with BC, not better. Outputs are only folded inside a running turn once they are 60 requests old.
|
|
33
|
-
- **Tool allowlists** (`pi --tools read,bash`, and sub-agent launchers that pass one): Pi then hides `zip_recall`,
|
|
33
|
+
- **Tool allowlists** (`pi --tools read,bash`, and sub-agent launchers that pass one): Pi then hides `zip_recall`. If `read`, `grep` or `bash` is on the list, pi-zip still folds as usual and the placeholder names a file the model reads or greps instead (see [Recall through files](#recall-through-files)). Only when none of them is allowed does it fold just the outputs the model can re-read (files, read-only commands), with a placeholder that says to re-read. Add `zip_recall` to the list to get the tool itself.
|
|
34
34
|
- If you always answer within the cache lifetime, there is little to save, by design.
|
|
35
35
|
|
|
36
36
|
## What you see
|
|
37
37
|
|
|
38
|
-
At most one line per turn
|
|
38
|
+
One mark, ƶ, and three words: **zipped**, **unzip**, **summarized**. Only the new context size is bright. At most one line per turn, at the top of the turn, only when the context was zipped or summarized:
|
|
39
39
|
|
|
40
40
|
```
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
41
|
+
ƶ zipped 12 old outputs 74K → 43K 6 ms
|
|
42
|
+
originals are kept · the model can unzip any of them
|
|
43
|
+
ƶ summarized 64 requests 182K → 41K 38.2 s · ready while you were away
|
|
44
44
|
```
|
|
45
45
|
|
|
46
|
-
The second line appears once per session. Click the
|
|
46
|
+
The second line appears once per session. A summary that made you wait says so (`waited 8.4 s`, highlighted from 2 s on). Click the line (or expand tool output with ctrl+o) to see what was zipped and why now (a cold return says `cache cold · away 47 min` or `cache cold · model switched`, a warm edit `cache warm · worth it at this size`):
|
|
47
47
|
|
|
48
48
|
```
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
49
|
+
ƶ zipped 12 old outputs 74K → 43K 6 ms
|
|
50
|
+
bash npm test 14K k3x9q2m7ab
|
|
51
|
+
read src/payment.ts 9.1K p8d2x1qa0m
|
|
52
|
+
… 10 more
|
|
53
|
+
cache cold · away 47 min
|
|
54
54
|
```
|
|
55
55
|
|
|
56
|
-
These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the
|
|
56
|
+
These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the time goes first, then the words; the numbers always stay. A summary still shows up as Pi's own `[compaction]` block as well; the ƶ line next to it tells you who made it. The pi-zip numbers are real tokens, the same scale as Pi's context meter (the meter adds the newest reply's output on top); the `Compacted from N tokens` figure in Pi's block is pi-zip's own size before for a summary prepared while you were away (that one goes in through Pi's compaction, before your prompt); for any other summary it is Pi's own estimate taken when the entry is saved, so it may not match.
|
|
57
57
|
|
|
58
|
-
|
|
58
|
+
A zipped output keeps its place in the transcript; its tool row gets a dim mark, so you can see what the model no longer sees in full:
|
|
59
59
|
|
|
60
60
|
```
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
▸ pi-zip paused billion-context also manages context, so pi-zip only guards requests · to use pi-zip: pi remove the other one
|
|
61
|
+
$ npm test ƶ zipped
|
|
62
|
+
read src/payment.ts ƶ zipped
|
|
64
63
|
```
|
|
65
64
|
|
|
66
|
-
|
|
65
|
+
When the model fetches an original back, it looks like any other tool row and names what came back (ctrl+o shows the text):
|
|
67
66
|
|
|
68
67
|
```
|
|
69
|
-
|
|
68
|
+
ƶ unzip bash npm test /Expected/
|
|
69
|
+
3 of 812 lines
|
|
70
70
|
```
|
|
71
71
|
|
|
72
|
-
|
|
72
|
+
Everything else uses the same grammar, and only when something changed:
|
|
73
73
|
|
|
74
74
|
```
|
|
75
|
-
|
|
76
|
-
|
|
75
|
+
ƶ pi-zip on old tool output gets zipped once the prompt cache expires (5 min here) · /zip
|
|
76
|
+
ƶ pi-zip on a tool allowlist hides zip_recall · zipped outputs are read back from files in ~/.cache/pi-zip/recall
|
|
77
|
+
ƶ pi-zip reread-only a tool allowlist hides zip_recall · only re-readable outputs fold · allow zip_recall, read, grep or bash to fold more
|
|
78
|
+
ƶ pi-zip paused billion-context also manages context, so pi-zip only guards requests · to use pi-zip: pi remove the other one
|
|
77
79
|
```
|
|
78
80
|
|
|
79
|
-
While a summary started in the background is being finished, Pi's working line says so (Esc
|
|
81
|
+
The first one appears once per machine. While a summary started in the background is being finished, Pi's working line says so (Esc there cancels the prompt, as it does whenever Pi is working; the summary goes on for your next try). `/zip` prints a small card into the transcript:
|
|
80
82
|
|
|
81
83
|
```
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
84
|
+
ƶ pi-zip on
|
|
85
|
+
model zhipu/glm-5.3
|
|
86
|
+
cache each dot is one time you came back
|
|
87
|
+
6 5
|
|
88
|
+
● ● ● ●
|
|
89
|
+
● ● ● ● ●
|
|
90
|
+
● ● ● ● ● ● still there 19
|
|
91
|
+
━━━━━━━━━━━━━━━━┅┅┅┅┅┅┅┅┅┅───────────────────────────────────
|
|
92
|
+
○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ gone 74
|
|
93
|
+
○ ○ ○ ○ ○ ○ ○ ○ ○
|
|
94
|
+
○ ○ ○ ○ ○ ○ ○ ○
|
|
95
|
+
○ 12 7 9 15 6 11 5
|
|
96
|
+
away 30s 1m 2m 5m 10m 30m 1h 2h
|
|
97
|
+
under 2 min: usually still there · 2–5 min: not sure yet · over 5 min: gone
|
|
98
|
+
session 30 zipped ~310K · 8 summaries $1.75 · 2 unzips
|
|
99
|
+
last 15:22 zipped 1 old output in 4 ms · cache cold · away 13 min
|
|
88
100
|
```
|
|
89
101
|
|
|
90
|
-
|
|
102
|
+
|
|
103
|
+
ƶ (U+01B6) is a plain Latin letter, one column wide everywhere, CJK terminals included. Most coding fonts have it; where one does not (Hack, Source Code Pro), the terminal draws it from a fallback font.
|
|
104
|
+
|
|
105
|
+
The chart is the one thing worth knowing: come back sooner than the solid part of the axis and pi-zip touches nothing; come back later and the old outputs get zipped, at no extra cost. Across (`away`) is how long you were gone, from 30 s to 2 h on a log scale. Every dot is one time you came back after that long: a filled dot above the axis is a time the cache was still there, an empty one below it a time it was gone, so the picture is your own history with this model, and a few gone samples take only a row or two. A column shows at most 4 dots a side; beyond that it shows 3 dots and the number at the far end of the stack (`12` = twelve times in that gap range). The axis is solid up to the longest gap that was (nearly) always safe, dashed while it is not sure, thin after it is mostly gone; the sentence under it says the same in words (`usually still there` when even the safe part had some misses). The counts are the learned ones, kept with older evidence halved after 16 newer observations of the same gap range, so they can be a little below the number of times you actually came back. Until something is learned there is no chart, one line only (`expires after ~5 min idle · a chart appears once you come back after a break`, from the model's declared cache lifetime; Claude's is 5 min). The chart needs 72 columns; where the legend does not fit beside it, it moves to one line under it, and in a narrower terminal only the sentence is shown. Notices state facts (how long you were away), not guesses: whether the cache really had expired is read from the provider's reply afterwards, so a wrong guess is not repeated.
|
|
91
106
|
|
|
92
107
|
## The three rules
|
|
93
108
|
|
|
94
109
|
1. **Nothing is lost.** The session file stays the single source of truth. pi-zip only changes the view sent to the model, never deletes anything. User messages, tool-call arguments, the system prompt, tool definitions and thinking are never rewritten. Every folded block carries a handle and shows its key lines (errors, ids, first and last line).
|
|
95
|
-
2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens in one step, and every request of that turn sends the same bytes. While the cache is warm it does nothing, with one exception, the warm valve: above
|
|
96
|
-
3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited
|
|
110
|
+
2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens of context in one step, system prompt and tool definitions included, and every request of that turn sends the same bytes. No edit can shrink the system prompt and the tools, so once they take more than half of the 40K (a large tool set) the target becomes them plus 20K of conversation: with a 50K tool set the context lands near 70K. It is never a 40K that no edit can reach (that folds everything and summarises at every return), and it does not keep a full 40K of conversation on top either (on a live bench with a 36K prefix that cost 13% more, with no quality difference measured). While the cache is warm it does nothing, with one exception, the warm valve: above that target it applies that plan, minus the previous user turn (a warm edit never folds the turn you just finished, except when you come back after the declared TTL: there the previous turn is eligible exactly as at a cold return, so a provider whose cache outlives its TTL does not keep it at every return; and it leaves alone what the last cold return had to protect, the previous turn's outputs of that return, until the next cold return, until the session has grown by more than about 11K tokens since, or until the context reaches Pi's compaction room: the window sliding by one prompt is not news, and folding them then would pay a warm rewrite for what the free rewrite at the cold return could not take), when the edit pays for the rewrite it causes, and never on the request right after an edited one (no back-to-back warm rewrites). That is one inequality, r Δ²/(2g) + η Δ ≥ K with K = (w − r)(P T − (1 − P) Δ): the reads the removed Δ tokens would cost while the context grows back at g tokens per request (measured in the session), plus, near Pi's compaction trigger, what Pi would charge for the same room (η), against the rewrite of the T = A tokens left after the edit (pricing only the suffix after the earliest edit fires warm edits earlier and lost quality in the offline evaluation; the suffix is logged as `Tsuf` for measurement). r and w are the read and rewrite price ratios of the cache class, never the model's price table: explicit write premium 0.1 / 1.25 x input (2 x on the 1-hour tier), automatic prefix cache 0.2 / 1 x input; the class is read from the provider's usage reports, and until the first response the old fixed rule applies. P is the probability that the cache is still warm. A cold return is P = 0, so K < 0 and it always fires; a single small fold never pays at a warm cache, a large one does.
|
|
111
|
+
3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in, at a cold return and in the warm valve alike; the call is priced as what it is: an uncached read of the conversation since the previous summary, which code carries forward and the model never rewrites, plus the narrative it writes) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited (Esc cancels the prompt; the summary goes on and is used when you send it again), and the notice says so. A summary that finished while you were away goes into the session through Pi's own compaction before your prompt (Pi's `[compaction]` block then shows pi-zip's size before), in milliseconds; when Pi cannot take it there (busy, nothing it would compact, an older Pi, another extension taking part in Pi's compaction: see Coexistence), it rides along with the first request as before. If you return while the cache is still warm and the valve does not fire, the prepared summary is discarded (its cost is still counted). With Pi's own idle cache warming on (`"cacheWarming": "idle"` in the global settings, the only place Pi reads it), the summary waits for Pi's decision instead of the 0.8 mark: while Pi keeps refreshing the cache, nothing is prepared (you would come back to a warm cache and not need it); once Pi stops, or no refresh follows its decision, or the next decision (due one warming delay after the refresh finished) does not come, the summary is prepared as usual. pi-zip only watches those decisions, it never answers them. In non-interactive modes (`-p`, `--mode json`) nothing is ever started in the background: a cold return that needs a summary computes it right then.
|
|
97
112
|
|
|
98
|
-
**The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip
|
|
113
|
+
**The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip` draws what it has learned as the dot chart above.
|
|
99
114
|
|
|
100
|
-
Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves, are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
|
|
115
|
+
Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves (or reads of a recall file), are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
|
|
101
116
|
|
|
102
|
-
**Token counts are calibrated, not guessed.** Sizes are estimated as chars/4,
|
|
117
|
+
**Token counts are calibrated, not guessed.** Sizes are estimated as chars/4 (Pi's rule), and real tokens are not proportional to that: every request carries a fixed prefix O (system prompt, tool definitions, the provider's own framing: 2-4K on a bare Pi, 30-55K with a large tool set) that no edit shrinks, and the conversation costs c real tokens per estimated one (about 0.7-1.1 on GLM and GPT, 1.5-2.2 on Claude, more on CJK-heavy text). So a size is O + c x the chars/4 estimate of the conversation. O is read from the session: what the provider still read from its cache on the request that first carried the newest summary (everything after the system prompt and tools had changed there), else the session's first request minus its small opening prompt, else 2 x the chars/4 of the system prompt and the tool definitions; a tool set that changed since moves it by the same factor. c = (real - O) / estimate for the newest reply that reports usage (input + cache read + cache write; clamped to 0.5-4, 1.5 before the first reply). Both depend on the provider's tokenizer, so only replies from the model the next request goes to count: right after a model switch O and c fall back to the defaults until its first reply, and when that model has no O of its own, the previous model's O carries over in conversation units (O / c). Both are read from the session itself on every decision, so a restart, `pi -p` or a resumed session calibrates exactly like a long-lived one, and nothing extra is stored. The conversation estimate also counts what every tool call puts on the wire besides its arguments and output (the call id, twice, and the tool-use markup), since that does not shrink when an output is folded. On recorded sessions the next request's size comes out within 2.8% (p90; 31% with the single ratio of 0.2.9). The size a plan predicts for the request that carries it is a few percent off in the median after a fold that removed over half of the context, 12-15% at worst one time in ten (what is left has not been measured yet); from the next reply on it is measured again (within 4%). The cold cap counts the whole context and leaves at least half of it to the conversation; the compaction room is the whole context too (it is Pi's trigger); sizes in the notices, `/zip` and the ledger use the same scale.
|
|
103
118
|
|
|
104
119
|
The cold cap is kept below Pi's own compaction trigger (window minus `compaction.reserveTokens`), so on small windows Pi's lossy compaction does not get there first.
|
|
105
120
|
|
|
@@ -107,12 +122,12 @@ The cold cap is kept below Pi's own compaction trigger (window minus `compaction
|
|
|
107
122
|
|
|
108
123
|
| Command | Effect |
|
|
109
124
|
|---|---|
|
|
110
|
-
| `/zip`
|
|
125
|
+
| `/zip` | the card above: state, how long the cache survives you being away (a dot chart of your own returns), this session's folds, summaries and recalls (counted from the session, so a restart does not reset them) |
|
|
111
126
|
| `/zip off` | strict no-op: no folds, no summaries, requests left untouched (earlier folds stay recallable) |
|
|
112
127
|
| `/zip on` | resume |
|
|
113
128
|
| `/zip quiet` | toggle the per-turn notice (folding continues) |
|
|
114
129
|
|
|
115
|
-
`off` and `quiet` are remembered per session; each change is one line in the transcript. In `-p` / json mode `/zip
|
|
130
|
+
`off` and `quiet` are remembered per session; each change is one line in the transcript. In `-p` / json mode `/zip` prints plain text to stderr.
|
|
116
131
|
|
|
117
132
|
## Recall
|
|
118
133
|
|
|
@@ -127,7 +142,7 @@ key lines kept (original line numbers; up to 8):
|
|
|
127
142
|
Original kept byte for byte, recallable even after summaries or compaction: zip_recall("k3f9a0x1qz") (optional grep/range) is instant, free, no side effects; prefer it to re-running or re-reading (output may differ). Do not guess its content.
|
|
128
143
|
```
|
|
129
144
|
|
|
130
|
-
Placeholders are written once, when the fold is saved: sessions folded by an earlier version keep their old placeholder text unchanged. Summaries list each folded output with its handle, turn and outcome in the same way. A summary of an earlier summary merges it section by section (requests, files, commands, other calls, errors, handle table, handle index) instead of clipping its text: items are deduplicated, the oldest are dropped first under a per-section budget with a count of what was left out, and the previous narrative survives as a short tail excerpt. Pi's own free-text compaction summaries are carried as a head-and-tail excerpt.
|
|
145
|
+
Placeholders are written once, when the fold is saved: sessions folded by an earlier version keep their old placeholder text unchanged. Summaries list each folded output with its handle, turn and outcome in the same way, and say at the top that the 10-character code next to a command, file or call is its handle and that what an output said is to be recalled rather than re-run or re-read. A summary of an earlier summary merges it section by section (requests, files, commands, other calls, errors, handle table, handle index) instead of clipping its text: items are deduplicated, the oldest are dropped first under a per-section budget with a count of what was left out, and the previous narrative survives as a short tail excerpt. Pi's own free-text compaction summaries are carried as a head-and-tail excerpt.
|
|
131
146
|
|
|
132
147
|
The model recalls by itself when it needs the content:
|
|
133
148
|
|
|
@@ -135,9 +150,10 @@ The model recalls by itself when it needs the content:
|
|
|
135
150
|
zip_recall({ handles: ["k3f9a0x1qz", "m2b7c4d8ww"] })
|
|
136
151
|
zip_recall({ handle: "k3f9a0x1qz", grep: "ERROR|FAIL" })
|
|
137
152
|
zip_recall({ handle: "k3f9a0x1qz", range: "120-240" })
|
|
153
|
+
zip_recall({ grep: "ECONNREFUSED" }) // no handle: searches every earlier tool output
|
|
138
154
|
```
|
|
139
155
|
|
|
140
|
-
Recall is batched, exact, and also works for outputs from before a compaction or a pi-zip summary (summaries carry a handle table and a budgeted index of older handles). Output comes back in pages of 20,000 characters; the page says how to continue:
|
|
156
|
+
Recall is batched, exact, and also works for outputs from before a compaction or a pi-zip summary (summaries carry a handle table and a budgeted index of older handles). Each recalled section is headed by the handle and the command or path that produced it, so a look-alike output (the same command run again, another slice of the same file) is easy to spot. Without a handle, `grep` searches every earlier tool output of the session (zip_recall results excepted, outputs from before a compaction included), newest output first, each match under its output's handle and command: this also reaches an output whose handle a long summary no longer lists. Output comes back in pages of 20,000 characters (8,000 for a search without a handle; a later page of a search covers the same outputs as its first page, so new output arriving in between does not shift it); the page says how to continue:
|
|
141
157
|
|
|
142
158
|
```
|
|
143
159
|
zip_recall({ handle: "k3f9a0x1qz", offset: 20000 }) // next page, by characters
|
|
@@ -146,17 +162,31 @@ zip_recall({ handle: "k3f9a0x1qz", offset: 20000, limit: 50000 }) // up to 50,00
|
|
|
146
162
|
|
|
147
163
|
Paging is by characters, so even one 45,000-character line can be read in full, and the pages add up to the original byte for byte. `offset` and `limit` also page a `grep` or `range` selection. `grep` is a case-insensitive regular expression; patterns that could backtrack badly (nested quantifiers, quantified alternation, back-references, very long patterns) are searched as literal text instead.
|
|
148
164
|
|
|
165
|
+
### Recall through files
|
|
166
|
+
|
|
167
|
+
When a tool allowlist hides `zip_recall` (`pi --tools read,grep,bash`, sub-agent launchers) but `read`, `grep` or `bash` is active, pi-zip folds exactly as in full mode and points the model at a file instead of the tool:
|
|
168
|
+
|
|
169
|
+
```
|
|
170
|
+
Original kept byte for byte, even after summaries or compaction, in ~/.cache/pi-zip/recall/01a120cb-3d40-75a6-aeac-320e71967606/k3f9a0x1qz.txt: read it (offset/limit for pages) or grep it; prefer that to re-running or re-reading the source (output may differ). Do not guess its content.
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
The sentence names only the tools that are active (with bash alone: `sed -n` to page, `grep -n` to search), and summaries name the same files. A file does not exist until a tool asks for it: when `read`, `grep`, `find` or `ls` gets such a path (whatever comes before `pi-zip/recall/`: another machine, another user, a forked session), or a `bash` command contains one, pi-zip writes the original from the session (the current branch) and points the call at this machine's copy. A `grep`, `ls` or `find` of the session's directory writes every output the model no longer sees in full (folded or summarized away), so a grep over the directory searches them all. A read of a recall file is a recall: it is never folded and counts as an unzip, and the transcript draws it like zip_recall's row (`ƶ unzip …`). `/zip` says `on` (it is a full mode) and has one more row: `unzip from files in ~/.cache/pi-zip/recall (a tool allowlist hides zip_recall)`, naming where the copies really are (see below; a narrow terminal drops the part in brackets). In a bash command the path is recognised as its own word, also after `X=`, `--file=` or inside `{a,b}`.
|
|
174
|
+
|
|
175
|
+
The copies live in `$XDG_CACHE_HOME/pi-zip/recall/<session id>/` (by default `~/.cache/pi-zip/recall/`, a place backups and dotfile sync conventionally skip; a relative `XDG_CACHE_HOME` is ignored, as the XDG spec says, since it would put them in the project folder): directories 0700, files 0600. The state line and `/zip` show the place in use. The same text is already in the session file, which Pi writes 0644. Only what the model asks for is copied, and every copy can be written again from the session at any time, so copies older than 14 days are deleted at the next session start; that cleanup touches nothing but pi-zip's own copies (`<handle>.txt` files in session-id directories, never through a symlink). A copy that no longer holds the original byte for byte (a bash command appended to it or edited it) is written again the next time a tool names it. If you delete a session file to get rid of a secret it contained, delete `<that place>/<its session id>/` too, or it stays until then. Without a session file (`pi --no-session`) nothing is copied: the outputs are in no file, so pi-zip keeps them off the disk too, folds only re-readable outputs and says so (`reread-only`). Without pi-zip loaded nothing is written, just as zip_recall would be missing.
|
|
176
|
+
|
|
149
177
|
## Coexistence
|
|
150
178
|
|
|
151
|
-
If another extension that manages context is loaded (for example billion-context, magic-context, pi-smart-compact, pi-hot-compact, pi-context-prune), pi-zip pauses folding, says so once, and keeps only its request guard. Two writers on the same view give unpredictable results. `/zip
|
|
179
|
+
If another extension that manages context is loaded (for example billion-context, magic-context, pi-smart-compact, pi-hot-compact, pi-context-prune), pi-zip pauses folding, says so once, and keeps only its request guard. Two writers on the same view give unpredictable results. `/zip` shows `paused`.
|
|
180
|
+
|
|
181
|
+
An extension that only writes Pi's compaction summaries (a `session_before_compact` handler, like Pi's own `custom-compaction` example) does not pause pi-zip, but Pi runs its handler for every compaction, the one that puts pi-zip's prepared summary in at a cold return included. So pi-zip watches that compaction: if another extension's summary went in instead of pi-zip's, pi-zip adds none of its own on top (its prepared summary is discarded, its cost still counted); if the compaction is not done within 2 s (that handler is writing a summary of its own), pi-zip cancels it before your prompt goes on, so nothing can land in the middle of the run, and sends its summary with the first request instead; and if it was cancelled, or took longer than pi-zip's own part ever does (250 ms), pi-zip stops going through Pi's compaction for the rest of the process, so that handler's work is not paid for and waited for at every return. When such an extension's summary is already in the session, pi-zip does not go through Pi's compaction at all. Either way, summaries then ride along with the first request, as on an older Pi. A handler that cancels manual compactions shows Pi's red "Compaction cancelled" line at most once per process.
|
|
152
182
|
|
|
153
183
|
The guard has two parts. Before saving a fold or a summary it checks that the edit itself would not leave a tool result without its tool call (failed or aborted assistant turns, which the provider layer drops anyway, are ignored, as is any oddity the session already had). And before a request is sent it repairs a tool result that lost its tool call, instead of letting the provider reject the request and brick the session. The repair understands the request shapes of Pi's Anthropic, OpenAI chat completions (and Mistral), OpenAI Responses (Azure, Codex), Google Gemini/Vertex and Bedrock Converse providers; a request of any other shape is passed through untouched.
|
|
154
184
|
|
|
155
185
|
## FAQ
|
|
156
186
|
|
|
157
|
-
**Will it save money?** Mostly on cold returns, which is where a long session pays for a full cache rewrite. While the cache is warm it edits only when the inequality above says the rewrite pays back (large contexts, near Pi's compaction trigger, outputs 60+ requests old). `/zip
|
|
187
|
+
**Will it save money?** Mostly on cold returns, which is where a long session pays for a full cache rewrite. While the cache is warm it edits only when the inequality above says the rewrite pays back (large contexts, near Pi's compaction trigger, outputs 60+ requests old). `/zip` shows what this session has folded (and roughly how many tokens that removed), its summaries with what their model calls cost, and its recalls. The numbers live in the session file, so a restart does not reset them.
|
|
158
188
|
|
|
159
|
-
**Does it cost extra?** Planning is free. A summary is one extra model call (the current model, no tools, no prompt cache), shown in `/zip
|
|
189
|
+
**Does it cost extra?** Planning is free. A summary is one extra model call (the current model, no tools, no prompt cache), shown in `/zip`. A background summary you never use (you came back while the cache was warm) is counted there too. Recalled content re-enters the context at normal prices.
|
|
160
190
|
|
|
161
191
|
**Can the model lose information?** Folded outputs are replaced by a placeholder with key lines and a handle, and the placeholder tells the model not to guess. Summaries quote your requests verbatim and never paraphrase the model's reasoning. If the model ignores the handle and guesses, that is a model failure pi-zip cannot catch; this is why only re-readable outputs of the previous turn are folded at a cold return; other outputs of the protected turns wait until they are 60 requests old.
|
|
162
192
|
|
|
@@ -177,7 +207,7 @@ bun run build-check # bun
|
|
|
177
207
|
|
|
178
208
|
`test/pi-integration.test.ts` runs the extension through Pi's real session manager, extension loader and `emitContext`, with a mid-conversation system update and an aborted tool-call turn in the session, and checks that the request is byte-identical before and after turn_end persists the edits.
|
|
179
209
|
|
|
180
|
-
Environment overrides exist for tests only and are not part of the product surface: `PI_ZIP_TTL_SECS` (cache TTL), `PI_ZIP_COLD_CAP` (fold target in tokens, default 40000), `PI_ZIP_FOLD_MIN` (smallest output worth folding, default 500 tokens), `PI_ZIP_KEEP_LINES`, `PI_ZIP_MIN_GAIN` (summary gain floor of the legacy rule used only when nothing is known about prices, default 10000), `PI_ZIP_INTURN_AGE` (age in assistant requests from which any output of the protected turns may fold, default 60; 0 = never), `PI_ZIP_CACHE_STATS=<path>` (where the learned cache survival lives; benches isolate it), `PI_ZIP_OFF=1` (register nothing), `PI_ZIP_LEDGER=<path>` (append a JSON line per decision; `prompt` records `ttlMs` and `ttlSource`, `declared` or `observed`; including the summary call's token usage; `cold_plan` records the calibration as `
|
|
210
|
+
Environment overrides exist for tests only and are not part of the product surface: `PI_ZIP_TTL_SECS` (cache TTL), `PI_ZIP_COLD_CAP` (fold target in real tokens of the whole context, default 40000; at least half of it is always left to the conversation above the system prompt and tools), `PI_ZIP_FOLD_MIN` (smallest output worth folding, default 500 tokens), `PI_ZIP_KEEP_LINES`, `PI_ZIP_MIN_GAIN` (summary gain floor of the legacy rule used only when nothing is known about prices, default 10000), `PI_ZIP_INTURN_AGE` (age in assistant requests from which any output of the protected turns may fold, default 60; 0 = never), `PI_ZIP_CACHE_STATS=<path>` (where the learned cache survival lives; benches isolate it), `PI_ZIP_OFF=1` (register nothing), `PI_ZIP_NATIVE_MS` (how long pi-zip's own compaction may take before it is cancelled, default 2000), `PI_ZIP_LEDGER=<path>` (append a JSON line per decision; `prompt` records `ttlMs` and `ttlSource`, `declared` or `observed`; including the summary call's token usage; `cold_plan` records the calibration as `O`, `c`, `oSource`, `calReal`, `calEst`, `calSource`, next to the real `ctxBefore` and `ctxAfter`; `prompt` also records `gapS`, `pWarm`, `survSrc` and `cls`; `law` records every evaluation with `g`, `wr` (w/r), `pWarm`, `survSrc`, `prSrc` and per step `B`, `A`, `T` (= A), `Tsuf` (the suffix after the earliest edit, measurement only), `phi`, `K`, `eta`; `b2b_skip` marks a warm edit skipped right after an edited request; `cache_sample` records each learned survival observation).
|
|
181
211
|
|
|
182
212
|
Internally, `RELAX_PREV_TURN` in `src/plan.ts` selects whether re-readable outputs of the previous user turn may be folded on a cold return or a return after the declared TTL (default `true`; other warm plans never fold them).
|
|
183
213
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-zip",
|
|
3
|
-
"version": "0.2
|
|
3
|
+
"version": "0.3.0-rc.2",
|
|
4
4
|
"description": "Keeps long Pi sessions cheap without losing anything: folds old tool output only when the prompt cache has already gone cold, and every fold can be recalled byte for byte. Zero config.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
package/src/cache.ts
CHANGED
|
@@ -44,12 +44,34 @@ export function observedTier(branch: Any[], key: string): "1h" | "5m" | undefine
|
|
|
44
44
|
|
|
45
45
|
export interface TtlInfo { ms: number; source: "declared" | "observed"; note?: string }
|
|
46
46
|
|
|
47
|
-
/**
|
|
48
|
-
|
|
47
|
+
/** The cache tier a request asked for, read from the payload Pi built (null: the payload does not say). Anthropic: a 1-hour
|
|
48
|
+
* cache_control ttl is the long tier, cache_control without one the short tier; OpenAI: prompt_cache_retention "24h" is long. */
|
|
49
|
+
export function payloadRetention(p: Any): "short" | "long" | null {
|
|
50
|
+
if (!p || typeof p !== "object") return null;
|
|
51
|
+
if (typeof p.prompt_cache_retention === "string") return p.prompt_cache_retention === "24h" ? "long" : "short";
|
|
52
|
+
let seen = false;
|
|
53
|
+
const one = (x: Any): boolean => {
|
|
54
|
+
const cc = x?.cache_control;
|
|
55
|
+
if (!cc || typeof cc !== "object") return false;
|
|
56
|
+
seen = true;
|
|
57
|
+
return cc.ttl === "1h";
|
|
58
|
+
};
|
|
59
|
+
for (const list of [p.system, p.tools]) if (Array.isArray(list)) for (const b of list) if (one(b)) return "long";
|
|
60
|
+
if (Array.isArray(p.messages)) for (const m of p.messages) {
|
|
61
|
+
if (one(m)) return "long";
|
|
62
|
+
if (Array.isArray(m?.content)) for (const b of m.content) if (one(b)) return "long";
|
|
63
|
+
}
|
|
64
|
+
return seen ? "short" : null;
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
/** Declared TTL, corrected by what the provider was seen to write. PI_ZIP_TTL_SECS wins. `note` = the one-time mismatch notice text.
|
|
68
|
+
* `tier` = the retention the request used (the provider-scoped PI_CACHE_RETENTION, or the request payload itself; run.ts); without it,
|
|
69
|
+
* the process environment decides, as before. */
|
|
70
|
+
export function resolveTtl(model: Any, branch: Any[] = [], tier?: "short" | "long"): TtlInfo {
|
|
49
71
|
if (process.env.PI_ZIP_TTL_SECS) return { ms: envInt("TTL_SECS", 300) * 1000, source: "declared" };
|
|
50
72
|
const pc = model?.promptCache;
|
|
51
|
-
const
|
|
52
|
-
const
|
|
73
|
+
const long = tier ? tier === "long" : process.env.PI_CACHE_RETENTION === "long";
|
|
74
|
+
const declared = cacheTtlMs(pc, DEFAULT_TTL_MS, { cacheRetention: long ? "long" : "short" });
|
|
53
75
|
const seen = declared > 0 ? observedTier(branch, modelKey(model)) : undefined;
|
|
54
76
|
const mismatch = (req: string, got: string, ms: number): TtlInfo => ({ ms, source: "observed", note: `${PRODUCT} · requested ${req} prompt cache, provider wrote ${got} · using ${got}` });
|
|
55
77
|
if (long && seen === "5m") return mismatch("1h", "5m", cacheTtlMs(pc, DEFAULT_TTL_MS, { cacheRetention: "short" }));
|
|
@@ -57,7 +79,7 @@ export function resolveTtl(model: Any, branch: Any[] = []): TtlInfo {
|
|
|
57
79
|
return { ms: declared, source: "declared" };
|
|
58
80
|
}
|
|
59
81
|
|
|
60
|
-
export const ttlFor = (model: Any, branch: Any[] = []): number => resolveTtl(model, branch).ms;
|
|
82
|
+
export const ttlFor = (model: Any, branch: Any[] = [], tier?: "short" | "long"): number => resolveTtl(model, branch, tier).ms;
|
|
61
83
|
|
|
62
84
|
/** Cold = a prior request exists and nothing touched the cache for longer than the TTL. No prior request counts as warm. */
|
|
63
85
|
export const isColdByTtl = (lastActivityMs: number, nowMs: number, ttlMs: number): boolean => lastActivityMs > 0 && nowMs - lastActivityMs > ttlMs;
|
|
@@ -69,18 +91,33 @@ const tsOf = (en: Any, inner?: Any): number => {
|
|
|
69
91
|
return Number.isFinite(t) ? t : 0;
|
|
70
92
|
};
|
|
71
93
|
|
|
94
|
+
/** A response that reached no cache: an error that reports no usage (the provider refused the request). One that reports usage was
|
|
95
|
+
* read: it touched the cache. An abort (Esc) comes after the provider read and wrote the prompt: it is a touch, with or without usage. */
|
|
96
|
+
export const refused = (m: Any): boolean => {
|
|
97
|
+
if (m?.role !== "assistant" || m.stopReason !== "error") return false;
|
|
98
|
+
const u = m.usage;
|
|
99
|
+
return !((Number(u?.input) || 0) + (Number(u?.cacheRead) || 0) + (Number(u?.cacheWrite) || 0) > 0);
|
|
100
|
+
};
|
|
101
|
+
|
|
72
102
|
/**
|
|
73
103
|
* Last time anything touched the cache, according to the branch: the newest message, or the newest cache-warm usage entry
|
|
74
104
|
* (Pi's own idle refresh persists `{ type: "usage", kind: "cache_warm" }`, and a refresh keeps the entry alive for another TTL).
|
|
105
|
+
* A request the provider refused (an error with no usage) touched nothing, and neither did the messages that made it (the prompt or
|
|
106
|
+
* the tool results it carried, written before it failed): the clock stays at the last request that was read.
|
|
75
107
|
* Fresh process: err on the warm side (0 = unknown).
|
|
76
108
|
*/
|
|
77
109
|
export function lastMessageMs(branch: Any[]): number {
|
|
78
|
-
let last = 0;
|
|
110
|
+
let last = 0, pending = 0; // pending: the newest non-assistant message since the last response, counted once its response is known
|
|
79
111
|
for (const en of branch) {
|
|
80
|
-
if (en?.type === "message" && en.message)
|
|
81
|
-
|
|
112
|
+
if (en?.type === "message" && en.message) {
|
|
113
|
+
const t = tsOf(en, en.message);
|
|
114
|
+
if (en.message.role === "assistant") {
|
|
115
|
+
if (!refused(en.message)) last = Math.max(last, pending, t);
|
|
116
|
+
pending = 0;
|
|
117
|
+
} else pending = Math.max(pending, t);
|
|
118
|
+
} else if (en?.type === "usage" && en.kind === "cache_warm") last = Math.max(last, tsOf(en));
|
|
82
119
|
}
|
|
83
|
-
return last;
|
|
120
|
+
return Math.max(last, pending); // trailing messages with no response yet (the run in progress) count, as before
|
|
84
121
|
}
|
|
85
122
|
|
|
86
123
|
/** The newest assistant message that really ran (errors and aborts may never have reached a cache): its provider/model and prompt size. */
|
|
@@ -102,11 +139,12 @@ export interface ColdInfo { cold: boolean; reason: string; ttl: TtlInfo; pWarm:
|
|
|
102
139
|
|
|
103
140
|
/** Cold by time, or because the model changed: a cache entry belongs to one provider and model. `learned` = P(warm) after a gap from the
|
|
104
141
|
* survival the provider was seen to have (learn.ts); without it, or under PI_ZIP_TTL_SECS (tests), the TTL decides. cold = P(warm) < 0.5. */
|
|
105
|
-
export function detectCold(model: Any, lastReqMs: number, branch: Any[], nowMs = Date.now(), lastModel = "", learned?: Survival): ColdInfo {
|
|
106
|
-
const ttl = resolveTtl(model, branch);
|
|
142
|
+
export function detectCold(model: Any, lastReqMs: number, branch: Any[], nowMs = Date.now(), lastModel = "", learned?: Survival, tier?: "short" | "long"): ColdInfo {
|
|
143
|
+
const ttl = resolveTtl(model, branch, tier);
|
|
107
144
|
const last = Math.max(lastReqMs, lastMessageMs(branch));
|
|
108
145
|
const prev = lastModel || lastModelInBranch(branch);
|
|
109
|
-
|
|
146
|
+
// a request that went out, even one the provider refused, was made on another model: the new model's cache is not that one
|
|
147
|
+
if (prev && modelKey(model) && prev !== modelKey(model)) return { cold: true, reason: `model switch ${prev} -> ${modelKey(model)}`, ttl, pWarm: 0, src: "model switch", gapS: null, pastTtl: false };
|
|
110
148
|
if (!last) return { cold: false, reason: "no prior request", ttl, pWarm: 1, src: "no prior request", gapS: null, pastTtl: false };
|
|
111
149
|
const gapS = (nowMs - last) / 1000;
|
|
112
150
|
const byTtl = isColdByTtl(last, nowMs, ttl.ms);
|
package/src/guard.ts
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
// Guard (F13, I5): validate an edit set before committing it, and repair orphan tool results in the outgoing payload.
|
|
2
|
-
import { RELAX_PREV_TURN, settings, type Block } from "./plan.ts";
|
|
2
|
+
import { finalAnswerIdx, RELAX_PREV_TURN, settings, type Block } from "./plan.ts";
|
|
3
3
|
import type { Any } from "./util.ts";
|
|
4
4
|
|
|
5
5
|
const PROTECT_USER_TURNS = 2;
|
|
@@ -10,6 +10,7 @@ export interface EditSet {
|
|
|
10
10
|
recover?: Map<number, string | undefined>; // block idx -> "rereadable" | "nonrereadable" (needed for folds in the previous user turn)
|
|
11
11
|
relax?: boolean; // default RELAX_PREV_TURN
|
|
12
12
|
inturnAge?: number; // default settings().inturnAge: an output (any class) at least this many assistant requests old may fold in a protected turn (0 = never)
|
|
13
|
+
ahead?: boolean; // a lookahead cut (plan.ts LOOKAHEAD): it may lie inside the previous user turn, never past that turn's final answer
|
|
13
14
|
}
|
|
14
15
|
|
|
15
16
|
/**
|
|
@@ -63,7 +64,7 @@ export function validateEdits(blocks: Block[], plan: EditSet, userTurns: number)
|
|
|
63
64
|
}
|
|
64
65
|
if (plan.cut !== null) {
|
|
65
66
|
if (plan.cut < 1 || plan.cut >= blocks.length) return `bad cut ${plan.cut}`;
|
|
66
|
-
if (limit >= 0 && plan.cut > limit) return `cut ${plan.cut} inside the last ${PROTECT_USER_TURNS} user turns`;
|
|
67
|
+
if (limit >= 0 && plan.cut > limit && !(plan.ahead && plan.cut <= finalAnswerIdx(blocks, userTurns - 1))) return `cut ${plan.cut} inside the last ${PROTECT_USER_TURNS} user turns`;
|
|
67
68
|
const kb = blocks[plan.cut];
|
|
68
69
|
if (!kb.entryId || (kb.kind !== "user" && kb.kind !== "assistant")) return `cut ${plan.cut} is not at a user/assistant message`;
|
|
69
70
|
// only what the cut itself would create is our problem; a session that was already odd stays as odd as it was
|
package/src/index.ts
CHANGED
|
@@ -4,8 +4,9 @@ import type { Any } from "./util.ts";
|
|
|
4
4
|
import { type NoticeData, registerZipCommand, renderNotice } from "./notice.ts";
|
|
5
5
|
import { RECALL_TOOL } from "./placeholder.ts";
|
|
6
6
|
import { registerRecallTool } from "./recall.ts";
|
|
7
|
+
import { parseRecallPath } from "./recallfile.ts";
|
|
7
8
|
import { NOTICE_CUSTOM, Zip } from "./run.ts";
|
|
8
|
-
import { markFirstLine, measure, renderCard, renderState, setMeasure } from "./ui.ts";
|
|
9
|
+
import { markFirstLine, measure, recallCallLine, renderCard, renderState, setMeasure } from "./ui.ts";
|
|
9
10
|
|
|
10
11
|
export default function piZip(pi: ExtensionAPI) {
|
|
11
12
|
if (process.env.PI_ZIP_OFF === "1") return; // test only: behave exactly as if not installed (registers nothing)
|
|
@@ -14,7 +15,7 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
14
15
|
registerZipCommand(pi, zip);
|
|
15
16
|
try {
|
|
16
17
|
// Everything pi-zip shows lives in the transcript as custom entries (never part of the model's context): fold/summary
|
|
17
|
-
// notices, state lines and the /zip
|
|
18
|
+
// notices, state lines and the /zip card. Widths are measured with Pi's own pi-tui once it has loaded.
|
|
18
19
|
import("@earendil-works/pi-tui").then((m: Any) => {
|
|
19
20
|
if (typeof m?.visibleWidth === "function" && typeof m?.truncateToWidth === "function")
|
|
20
21
|
setMeasure({ vw: m.visibleWidth, cut: (s: string, w: number) => (m.visibleWidth(s) <= w ? s : w <= 0 ? "" : m.truncateToWidth(s, w, "…")) });
|
|
@@ -38,7 +39,9 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
38
39
|
if (d.v !== 2) return [theme.fg("dim", measure.cut(d.text, width))];
|
|
39
40
|
if (d.kind === "state") return renderState(d.word, d.reason ?? "", width, theme, measure);
|
|
40
41
|
if (d.kind === "card" && d.card) return renderCard(d.card, width, theme, measure);
|
|
41
|
-
|
|
42
|
+
// once the request that carried the plan has been answered, "after" is its real size (older notices keep their number)
|
|
43
|
+
const real = typeof d.reqTs === "number" ? zip.realSizeAfter(d.reqTs) : undefined;
|
|
44
|
+
return renderNotice((real ? { ...d, after: real } : d) as NoticeData, open(), width, theme, measure.vw);
|
|
42
45
|
} catch {
|
|
43
46
|
return [];
|
|
44
47
|
}
|
|
@@ -52,13 +55,21 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
52
55
|
};
|
|
53
56
|
});
|
|
54
57
|
zip.entryRenderer = typeof pi.registerEntryRenderer === "function";
|
|
55
|
-
// a folded output keeps its place in the transcript; its call row gets
|
|
58
|
+
// a folded output keeps its place in the transcript; its call row gets "ƶ zipped" right after its own text
|
|
56
59
|
pi.registerToolRenderer?.((toolName: string, next: () => Any) => {
|
|
57
60
|
const base = next();
|
|
58
61
|
if (toolName === RECALL_TOOL || typeof base?.renderCall !== "function") return base;
|
|
59
62
|
return {
|
|
60
63
|
...base,
|
|
61
64
|
renderCall: (args: Any, theme: Any, c: Any) => {
|
|
65
|
+
// file recall: a read or grep of a recall file (or of the session's recall directory) reads like zip_recall's row
|
|
66
|
+
const ref = toolName === "read" || toolName === "grep" ? parseRecallPath(args?.path) : null;
|
|
67
|
+
if (ref && (ref.handle || toolName === "grep")) {
|
|
68
|
+
const off = typeof args.offset === "number" ? args.offset : undefined, lim = typeof args.limit === "number" ? args.limit : undefined;
|
|
69
|
+
const range = toolName === "read" && (off !== undefined || lim !== undefined) ? `${off ?? 1}-${lim !== undefined ? (off ?? 1) + lim - 1 : ""}` : undefined;
|
|
70
|
+
const a = { handle: ref.handle, grep: toolName === "grep" && typeof args.pattern === "string" ? args.pattern : undefined, range };
|
|
71
|
+
return { render: (width: number) => recallCallLine(a, Math.max(1, width), theme, measure, (h) => zip.labelOf(h)), invalidate: () => {} };
|
|
72
|
+
}
|
|
62
73
|
const comp = base.renderCall(args, theme, c);
|
|
63
74
|
const id = c?.toolCallId;
|
|
64
75
|
if (!comp || typeof comp.render !== "function" || typeof id !== "string") return comp;
|
|
@@ -83,11 +94,15 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
83
94
|
/* older Pi: notices fall back to the status line */
|
|
84
95
|
}
|
|
85
96
|
// A tool allowlist (`pi --tools read,bash`, sub-agent launchers) replaces the whole selection and Pi then does not even register
|
|
86
|
-
// zip_recall; folds of outputs the model could not get back would be lost, so only rereadable outputs fold then (plan rereadOnly)
|
|
97
|
+
// zip_recall; folds of outputs the model could not get back would be lost, so only rereadable outputs fold then (plan rereadOnly)...
|
|
98
|
+
// ... unless read, grep or bash is active: then folds are as in full mode and placeholders name a recall file (recallfile.ts) that the
|
|
99
|
+
// tool_call hook below writes from the session when a tool asks for it.
|
|
87
100
|
const checkRecall = (ctx: Any) => {
|
|
88
101
|
try {
|
|
89
102
|
const active = pi.getActiveTools?.();
|
|
90
|
-
if (Array.isArray(active))
|
|
103
|
+
if (!Array.isArray(active)) return;
|
|
104
|
+
const tools = { read: active.includes("read"), grep: active.includes("grep"), bash: active.includes("bash") };
|
|
105
|
+
zip.setRecallMode(active.includes(RECALL_TOOL) ? "tool" : tools.read || tools.grep || tools.bash ? "file" : "reread", tools, ctx);
|
|
91
106
|
} catch {
|
|
92
107
|
/* never in the way */
|
|
93
108
|
}
|
|
@@ -97,17 +112,24 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
97
112
|
checkRecall(ctx);
|
|
98
113
|
zip.refreshFolded(ctx);
|
|
99
114
|
zip.welcome(ctx);
|
|
115
|
+
zip.resumeAway(ctx); // a resumed session: arm the away timer from its last request, as the settle after it would have
|
|
100
116
|
return r;
|
|
101
117
|
});
|
|
102
|
-
pi.on("before_agent_start", (
|
|
118
|
+
pi.on("before_agent_start", async (e, ctx) => {
|
|
103
119
|
checkRecall(ctx);
|
|
104
120
|
zip.refreshFolded(ctx); // the branch may have changed (/tree, fork)
|
|
105
|
-
|
|
121
|
+
await zip.beforeAgentStart(ctx, e); // may persist a summary prepared while the user was away through Pi's compaction (Pi is idle here)
|
|
122
|
+
return undefined;
|
|
106
123
|
});
|
|
124
|
+
// answers only the compaction pi-zip itself just asked for; every other compaction (Pi's, /compact, other extensions') gets undefined
|
|
125
|
+
pi.on("session_before_compact", (e) => zip.sessionBeforeCompact(e));
|
|
107
126
|
pi.on("context_with_system", (e, ctx) => zip.context(e, ctx)); // the complete transcript: system messages stay where Pi put them
|
|
108
127
|
pi.on("turn_end", (e, ctx) => zip.turnEnd(e, ctx));
|
|
109
128
|
pi.on("agent_before_settle", (e, ctx) => zip.settle(e, ctx));
|
|
110
129
|
pi.on("before_provider_request", (e) => zip.providerRequest(e));
|
|
130
|
+
pi.on("tool_call", (e, ctx) => zip.toolCall(e, ctx)); // a recall file named by a built-in tool: this machine's copy, written on demand
|
|
131
|
+
// Pi's idle cache warming: observed only (the return value is undefined, so Pi's own decision stands); an older Pi never fires it
|
|
132
|
+
pi.on("cache_warming_decision", (e) => zip.cacheWarmingDecision(e));
|
|
111
133
|
pi.on("message_end", (e) => zip.messageEnd(e.message));
|
|
112
134
|
pi.on("session_shutdown", () => zip.shutdown());
|
|
113
135
|
}
|
package/src/learn.ts
CHANGED
|
@@ -26,8 +26,10 @@ export function lawPrices(cls: CacheClass | undefined, longTier: boolean): Price
|
|
|
26
26
|
}
|
|
27
27
|
|
|
28
28
|
// ---- cache survival ------------------------------------------------------------------------------------------------
|
|
29
|
-
/** Gap bins (s): [GAP_EDGES[i], GAP_EDGES[i+1]). Edges sit on the known TTL tiers (300 s, 3600 s): a deterministic TTL never splits a bin.
|
|
30
|
-
|
|
29
|
+
/** Gap bins (s): [GAP_EDGES[i], GAP_EDGES[i+1]). Edges sit on the known TTL tiers (300 s, 3600 s): a deterministic TTL never splits a bin.
|
|
30
|
+
* The last edge closes the last bin: a gap at or beyond it (2 h) is outside what was ever observed, it is not recorded and it is planned
|
|
31
|
+
* from the declared TTL (a cache seen alive at 110 min says nothing about 156 min). */
|
|
32
|
+
export const GAP_EDGES = [30, 60, 120, 180, 240, 300, 330, 360, 420, 480, 600, 900, 1200, 1800, 2700, 3600, 5400, 7200];
|
|
31
33
|
export const HALF_LIFE = 16; // observations per bin: old evidence counts half after 16 newer ones in the same bin (a provider may change its TTL)
|
|
32
34
|
// The declared TTL as pseudo-observations. Asymmetric because the signal is: a read of the re-sent prefix cannot happen on a dead cache,
|
|
33
35
|
// but warm misses do (GLM 4-9%, glm-flash ~22%, research round 5). Beyond the TTL one clean read overrides it (weight 1/4: one hit -> 0.8);
|
|
@@ -41,10 +43,17 @@ export const MIN_EXPECT = 8192;
|
|
|
41
43
|
export interface Entry { cls?: CacheClass; bins: Record<string, [number, number]>; n: number } // bin index -> [alive, dead] (decayed counts)
|
|
42
44
|
interface File { v: 1; models: Record<string, Entry> }
|
|
43
45
|
|
|
44
|
-
|
|
45
|
-
|
|
46
|
+
/** Pi's agent directory (settings.json, sessions): PI_CODING_AGENT_DIR, else ~/.pi/agent. */
|
|
47
|
+
export const agentDir = (): string => process.env.PI_CODING_AGENT_DIR || join(homedir(), ".pi", "agent");
|
|
46
48
|
|
|
47
|
-
export const
|
|
49
|
+
export const statsPath = (): string => process.env.PI_ZIP_CACHE_STATS || join(agentDir(), "pi-zip", "cache-survival.json");
|
|
50
|
+
|
|
51
|
+
/** ES2023's Array.prototype.findLastIndex (Node >= 18 and Bun have it); tsconfig's lib is ES2022, so it is typed here (types only). */
|
|
52
|
+
interface FindLastIndex { findLastIndex(predicate: (value: number, index: number) => boolean): number }
|
|
53
|
+
export const binOf = (gapS: number): number => {
|
|
54
|
+
const i = (GAP_EDGES as number[] & FindLastIndex).findLastIndex((e) => gapS >= e);
|
|
55
|
+
return i >= GAP_EDGES.length - 1 ? -1 : i; // -1 = below the first edge or at/above the last: not recorded
|
|
56
|
+
};
|
|
48
57
|
const fin = (x: Any) => typeof x === "number" && Number.isFinite(x) && x >= 0;
|
|
49
58
|
|
|
50
59
|
/** The persisted stats; a missing, unreadable or malformed file is ignored (empty). */
|
|
@@ -55,7 +64,7 @@ export function loadStats(path = statsPath()): File {
|
|
|
55
64
|
if (j?.v !== 1 || typeof j.models !== "object" || !j.models) return out;
|
|
56
65
|
for (const [k, e] of Object.entries<Any>(j.models)) {
|
|
57
66
|
const bins: Entry["bins"] = {};
|
|
58
|
-
for (const [b, c] of Object.entries<Any>(e?.bins ?? {})) if (GAP_EDGES[Number(b)] !== undefined && Array.isArray(c) && fin(c[0]) && fin(c[1])) bins[b] = [c[0], c[1]];
|
|
67
|
+
for (const [b, c] of Object.entries<Any>(e?.bins ?? {})) if (Number(b) < GAP_EDGES.length - 1 && GAP_EDGES[Number(b)] !== undefined && Array.isArray(c) && fin(c[0]) && fin(c[1])) bins[b] = [c[0], c[1]];
|
|
59
68
|
out.models[k] = { bins, n: fin(e?.n) ? e.n : 0, ...(e?.cls === "explicit" || e?.cls === "automatic" ? { cls: e.cls } : {}) };
|
|
60
69
|
}
|
|
61
70
|
return out;
|
|
@@ -125,7 +134,8 @@ export function curve(e: Entry | undefined, priorS: number): { bin: number; p: n
|
|
|
125
134
|
* between the nearest observed bins below and above (an alive read at 365 s makes every shorter gap warm; a miss makes longer ones dead). */
|
|
126
135
|
export function pWarm(e: Entry | undefined, gapS: number, priorS: number): { p: number; src: "prior" | "learned" } {
|
|
127
136
|
const prior = gapS <= priorS ? 1 : 0;
|
|
128
|
-
|
|
137
|
+
// beyond the last edge: no bin, the declared TTL decides (clamped by the newest bin below, as for an unobserved bin)
|
|
138
|
+
const c = curve(e, priorS), b = gapS >= GAP_EDGES[GAP_EDGES.length - 1] ? GAP_EDGES.length - 1 : binOf(gapS);
|
|
129
139
|
const at = c.find((x) => x.bin === b);
|
|
130
140
|
if (at) return { p: at.p, src: "learned" };
|
|
131
141
|
const hi = c.filter((x) => x.bin < b).at(-1)?.p ?? 1, lo = c.find((x) => x.bin > b)?.p ?? 0;
|
|
@@ -133,7 +143,7 @@ export function pWarm(e: Entry | undefined, gapS: number, priorS: number): { p:
|
|
|
133
143
|
return { p, src: p === prior ? "prior" : "learned" };
|
|
134
144
|
}
|
|
135
145
|
|
|
136
|
-
/** /zip
|
|
146
|
+
/** /zip text: class and the learned curve ("360-420s 0.80 n1"). */
|
|
137
147
|
export function describe(e: Entry | undefined, priorS: number): string {
|
|
138
148
|
const c = curve(e, priorS);
|
|
139
149
|
const bins = c.map((x) => `${GAP_EDGES[x.bin]}-${GAP_EDGES[x.bin + 1] ?? "∞"}s ${x.p.toFixed(2)} n${Math.round(x.n * 10) / 10}`).join(", ");
|