pi-zip 0.2.7 → 0.3.0-rc.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +60 -42
- package/package.json +1 -1
- package/src/cache.ts +50 -12
- package/src/guard.ts +3 -2
- package/src/index.ts +45 -8
- package/src/learn.ts +18 -8
- package/src/notice.ts +91 -33
- package/src/placeholder.ts +12 -4
- package/src/plan.ts +434 -106
- package/src/recall.ts +111 -27
- package/src/recallfile.ts +165 -0
- package/src/run.ts +782 -116
- package/src/state.ts +53 -0
- package/src/summary.ts +22 -10
- package/src/ui.ts +107 -44
- package/src/util.ts +57 -2
package/README.md
CHANGED
|
@@ -30,76 +30,79 @@ Known limits:
|
|
|
30
30
|
|
|
31
31
|
- **GLM** (automatic prefix cache that outlives its declared 5 minutes): v0.1 cost about 1.2× BC live. v0.2 learns the real cache lifetime and folds the previous turn after the declared TTL (offline 0.95–0.97× v0.1), but this was not verified live.
|
|
32
32
|
- **Long autonomous runs** (one prompt, hundreds of tool calls, e.g. sub-agents): roughly on par with BC, not better. Outputs are only folded inside a running turn once they are 60 requests old.
|
|
33
|
-
- **Tool allowlists** (`pi --tools read,bash`, and sub-agent launchers that pass one): Pi then hides `zip_recall`,
|
|
33
|
+
- **Tool allowlists** (`pi --tools read,bash`, and sub-agent launchers that pass one): Pi then hides `zip_recall`. If `read`, `grep` or `bash` is on the list, pi-zip still folds as usual and the placeholder names a file the model reads or greps instead (see [Recall through files](#recall-through-files)). Only when none of them is allowed does it fold just the outputs the model can re-read (files, read-only commands), with a placeholder that says to re-read. Add `zip_recall` to the list to get the tool itself.
|
|
34
34
|
- If you always answer within the cache lifetime, there is little to save, by design.
|
|
35
35
|
|
|
36
36
|
## What you see
|
|
37
37
|
|
|
38
|
-
At most one line per turn
|
|
38
|
+
One mark, ƶ, and three words: **zipped**, **unzip**, **summarized**. Only the new context size is bright. At most one line per turn, at the top of the turn, only when the context was zipped or summarized:
|
|
39
39
|
|
|
40
40
|
```
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
41
|
+
ƶ zipped 12 old outputs 74K → 43K 6 ms
|
|
42
|
+
originals are kept · the model can unzip any of them
|
|
43
|
+
ƶ summarized 64 requests 182K → 41K 38.2 s · ready while you were away
|
|
44
44
|
```
|
|
45
45
|
|
|
46
|
-
The second line appears once per session.
|
|
46
|
+
The second line appears once per session. A summary that made you wait says so (`waited 8.4 s`, highlighted from 2 s on). Click the line (or expand tool output with ctrl+o) to see what was zipped and why now:
|
|
47
47
|
|
|
48
48
|
```
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
49
|
+
ƶ zipped 12 old outputs 74K → 43K 6 ms
|
|
50
|
+
bash npm test 14K k3x9q2m7ab
|
|
51
|
+
read src/payment.ts 9.1K p8d2x1qa0m
|
|
52
|
+
… 10 more
|
|
53
|
+
away 47 min
|
|
54
54
|
```
|
|
55
55
|
|
|
56
|
-
These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the
|
|
56
|
+
These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the time goes first, then the words; the numbers always stay. A summary still shows up as Pi's own `[compaction]` block as well; the ƶ line next to it tells you who made it. The pi-zip numbers are real tokens, the same scale as Pi's context meter (the meter adds the newest reply's output on top); the `Compacted from N tokens` figure in Pi's block is pi-zip's own size before for a summary prepared while you were away (that one goes in through Pi's compaction, before your prompt); for any other summary it is Pi's own estimate taken when the entry is saved, so it may not match.
|
|
57
57
|
|
|
58
|
-
|
|
58
|
+
A zipped output keeps its place in the transcript; its tool row gets a dim mark, so you can see what the model no longer sees in full:
|
|
59
59
|
|
|
60
60
|
```
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
▸ pi-zip paused billion-context also manages context, so pi-zip only guards requests · to use pi-zip: pi remove the other one
|
|
61
|
+
$ npm test ƶ zipped
|
|
62
|
+
read src/payment.ts ƶ zipped
|
|
64
63
|
```
|
|
65
64
|
|
|
66
|
-
|
|
65
|
+
When the model fetches an original back, it looks like any other tool row and names what came back (ctrl+o shows the text):
|
|
67
66
|
|
|
68
67
|
```
|
|
69
|
-
|
|
68
|
+
ƶ unzip bash npm test /Expected/
|
|
69
|
+
3 of 812 lines
|
|
70
70
|
```
|
|
71
71
|
|
|
72
|
-
|
|
72
|
+
Everything else uses the same grammar, and only when something changed:
|
|
73
73
|
|
|
74
74
|
```
|
|
75
|
-
|
|
76
|
-
|
|
75
|
+
ƶ pi-zip on old tool output gets zipped once the prompt cache expires (5 min here) · /zip
|
|
76
|
+
ƶ pi-zip on a tool allowlist hides zip_recall · zipped outputs are read back from files in ~/.cache/pi-zip/recall
|
|
77
|
+
ƶ pi-zip reread-only a tool allowlist hides zip_recall · only re-readable outputs fold · allow zip_recall, read, grep or bash to fold more
|
|
78
|
+
ƶ pi-zip paused billion-context also manages context, so pi-zip only guards requests · to use pi-zip: pi remove the other one
|
|
77
79
|
```
|
|
78
80
|
|
|
79
|
-
While a summary started in the background is being finished, Pi's working line says so (Esc
|
|
81
|
+
The first one appears once per machine. While a summary started in the background is being finished, Pi's working line says so (Esc there cancels the prompt, as it does whenever Pi is working; the summary goes on for your next try). `/zip` prints a small card into the transcript:
|
|
80
82
|
|
|
81
83
|
```
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
mode full (zip_recall available)
|
|
84
|
+
ƶ pi-zip on
|
|
85
|
+
model anthropic/claude-sonnet-5-5
|
|
86
|
+
cache ██████████████▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ expires after ~5 min idle
|
|
87
|
+
away 1m 5m 15m 1h 2h
|
|
88
|
+
session 56 zipped ~310K · 2 summaries $0.41 · 3 unzips
|
|
88
89
|
```
|
|
89
90
|
|
|
90
|
-
|
|
91
|
+
ƶ (U+01B6) is a plain Latin letter, one column wide everywhere, CJK terminals included. Most coding fonts have it; where one does not (Hack, Source Code Pro), the terminal draws it from a fallback font.
|
|
92
|
+
|
|
93
|
+
`expires after ~5 min idle` is the one thing worth knowing: come back sooner and pi-zip touches nothing; come back later and the old outputs get zipped, at no extra cost. The chart draws the same fact: across (`away`) is how long you were gone, from 30 s to 2 h on a log scale; up (`cache`) is how likely the cache still holds your conversation then. Bars in the accent colour are still there, dim ones are gone. It starts from the model's declared cache lifetime (a cliff at 5 min for Claude) and turns into `(measured)` once the provider's own replies have shown how long its cache really lasts (see below); a cache that fades out slowly draws a slope. In a narrow terminal the chart is left out and the sentence stays whole (the model name goes first). Notices state facts (how long you were away), not guesses: whether the cache really had expired is read from the provider's reply afterwards, so a wrong guess is not repeated.
|
|
91
94
|
|
|
92
95
|
## The three rules
|
|
93
96
|
|
|
94
97
|
1. **Nothing is lost.** The session file stays the single source of truth. pi-zip only changes the view sent to the model, never deletes anything. User messages, tool-call arguments, the system prompt, tool definitions and thinking are never rewritten. Every folded block carries a handle and shows its key lines (errors, ids, first and last line).
|
|
95
|
-
2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens in one step, and every request of that turn sends the same bytes. While the cache is warm it does nothing, with one exception, the warm valve: above
|
|
96
|
-
3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited
|
|
98
|
+
2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens of context in one step, system prompt and tool definitions included, and every request of that turn sends the same bytes. No edit can shrink the system prompt and the tools, so once they take more than half of the 40K (a large tool set) the target becomes them plus 20K of conversation: with a 50K tool set the context lands near 70K. It is never a 40K that no edit can reach (that folds everything and summarises at every return), and it does not keep a full 40K of conversation on top either (on a live bench with a 36K prefix that cost 13% more, with no quality difference measured). While the cache is warm it does nothing, with one exception, the warm valve: above that target it applies that plan, minus the previous user turn (a warm edit never folds the turn you just finished, except when you come back after the declared TTL: there the previous turn is eligible exactly as at a cold return, so a provider whose cache outlives its TTL does not keep it at every return), when the edit pays for the rewrite it causes, and never on the request right after an edited one (no back-to-back warm rewrites). That is one inequality, r Δ²/(2g) + η Δ ≥ K with K = (w − r)(P T − (1 − P) Δ): the reads the removed Δ tokens would cost while the context grows back at g tokens per request (measured in the session), plus, near Pi's compaction trigger, what Pi would charge for the same room (η), against the rewrite of the T = A tokens left after the edit (pricing only the suffix after the earliest edit fires warm edits earlier and lost quality in the offline evaluation; the suffix is logged as `Tsuf` for measurement). r and w are the read and rewrite price ratios of the cache class, never the model's price table: explicit write premium 0.1 / 1.25 x input (2 x on the 1-hour tier), automatic prefix cache 0.2 / 1 x input; the class is read from the provider's usage reports, and until the first response the old fixed rule applies. P is the probability that the cache is still warm. A cold return is P = 0, so K < 0 and it always fires; a single small fold never pays at a warm cache, a large one does.
|
|
99
|
+
3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in, at a cold return and in the warm valve alike; the call is priced as what it is: an uncached read of the conversation since the previous summary, which code carries forward and the model never rewrites, plus the narrative it writes) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited (Esc cancels the prompt; the summary goes on and is used when you send it again), and the notice says so. A summary that finished while you were away goes into the session through Pi's own compaction before your prompt (Pi's `[compaction]` block then shows pi-zip's size before), in milliseconds; when Pi cannot take it there (busy, nothing it would compact, an older Pi, another extension taking part in Pi's compaction: see Coexistence), it rides along with the first request as before. If you return while the cache is still warm and the valve does not fire, the prepared summary is discarded (its cost is still counted). With Pi's own idle cache warming on (`"cacheWarming": "idle"` in the global settings, the only place Pi reads it), the summary waits for Pi's decision instead of the 0.8 mark: while Pi keeps refreshing the cache, nothing is prepared (you would come back to a warm cache and not need it); once Pi stops, or no refresh follows its decision, or the next decision (due one warming delay after the refresh finished) does not come, the summary is prepared as usual. pi-zip only watches those decisions, it never answers them. In non-interactive modes (`-p`, `--mode json`) nothing is ever started in the background: a cold return that needs a summary computes it right then.
|
|
97
100
|
|
|
98
|
-
**The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip
|
|
101
|
+
**The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip` shows the lifetime it currently believes, marked `(measured)` once it is learned.
|
|
99
102
|
|
|
100
|
-
Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves, are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
|
|
103
|
+
Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves (or reads of a recall file), are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
|
|
101
104
|
|
|
102
|
-
**Token counts are calibrated, not guessed.** Sizes are estimated as chars/4,
|
|
105
|
+
**Token counts are calibrated, not guessed.** Sizes are estimated as chars/4 (Pi's rule), and real tokens are not proportional to that: every request carries a fixed prefix O (system prompt, tool definitions, the provider's own framing: 2-4K on a bare Pi, 30-55K with a large tool set) that no edit shrinks, and the conversation costs c real tokens per estimated one (about 0.7-1.1 on GLM and GPT, 1.5-2.2 on Claude, more on CJK-heavy text). So a size is O + c x the chars/4 estimate of the conversation. O is read from the session: what the provider still read from its cache on the request that first carried the newest summary (everything after the system prompt and tools had changed there), else the session's first request minus its small opening prompt, else 2 x the chars/4 of the system prompt and the tool definitions; a tool set that changed since moves it by the same factor. c = (real - O) / estimate for the newest reply that reports usage (input + cache read + cache write; clamped to 0.5-4, 1.5 before the first reply). Both depend on the provider's tokenizer, so only replies from the model the next request goes to count: right after a model switch O and c fall back to the defaults until its first reply, and when that model has no O of its own, the previous model's O carries over in conversation units (O / c). Both are read from the session itself on every decision, so a restart, `pi -p` or a resumed session calibrates exactly like a long-lived one, and nothing extra is stored. The conversation estimate also counts what every tool call puts on the wire besides its arguments and output (the call id, twice, and the tool-use markup), since that does not shrink when an output is folded. On recorded sessions the next request's size comes out within 2.8% (p90; 31% with the single ratio of 0.2.9). The size a plan predicts for the request that carries it is a few percent off in the median after a fold that removed over half of the context, 12-15% at worst one time in ten (what is left has not been measured yet); from the next reply on it is measured again (within 4%). The cold cap counts the whole context and leaves at least half of it to the conversation; the compaction room is the whole context too (it is Pi's trigger); sizes in the notices, `/zip` and the ledger use the same scale.
|
|
103
106
|
|
|
104
107
|
The cold cap is kept below Pi's own compaction trigger (window minus `compaction.reserveTokens`), so on small windows Pi's lossy compaction does not get there first.
|
|
105
108
|
|
|
@@ -107,12 +110,12 @@ The cold cap is kept below Pi's own compaction trigger (window minus `compaction
|
|
|
107
110
|
|
|
108
111
|
| Command | Effect |
|
|
109
112
|
|---|---|
|
|
110
|
-
| `/zip`
|
|
113
|
+
| `/zip` | the card above: state, how long the cache survives you being away (chart and sentence), this session's folds, summaries and recalls (counted from the session, so a restart does not reset them) |
|
|
111
114
|
| `/zip off` | strict no-op: no folds, no summaries, requests left untouched (earlier folds stay recallable) |
|
|
112
115
|
| `/zip on` | resume |
|
|
113
116
|
| `/zip quiet` | toggle the per-turn notice (folding continues) |
|
|
114
117
|
|
|
115
|
-
`off` and `quiet` are remembered per session; each change is one line in the transcript. In `-p` / json mode `/zip
|
|
118
|
+
`off` and `quiet` are remembered per session; each change is one line in the transcript. In `-p` / json mode `/zip` prints plain text to stderr.
|
|
116
119
|
|
|
117
120
|
## Recall
|
|
118
121
|
|
|
@@ -127,7 +130,7 @@ key lines kept (original line numbers; up to 8):
|
|
|
127
130
|
Original kept byte for byte, recallable even after summaries or compaction: zip_recall("k3f9a0x1qz") (optional grep/range) is instant, free, no side effects; prefer it to re-running or re-reading (output may differ). Do not guess its content.
|
|
128
131
|
```
|
|
129
132
|
|
|
130
|
-
Placeholders are written once, when the fold is saved: sessions folded by an earlier version keep their old placeholder text unchanged. Summaries list each folded output with its handle, turn and outcome in the same way. A summary of an earlier summary merges it section by section (requests, files, commands, other calls, errors, handle table, handle index) instead of clipping its text: items are deduplicated, the oldest are dropped first under a per-section budget with a count of what was left out, and the previous narrative survives as a short tail excerpt. Pi's own free-text compaction summaries are carried as a head-and-tail excerpt.
|
|
133
|
+
Placeholders are written once, when the fold is saved: sessions folded by an earlier version keep their old placeholder text unchanged. Summaries list each folded output with its handle, turn and outcome in the same way, and say at the top that the 10-character code next to a command, file or call is its handle and that what an output said is to be recalled rather than re-run or re-read. A summary of an earlier summary merges it section by section (requests, files, commands, other calls, errors, handle table, handle index) instead of clipping its text: items are deduplicated, the oldest are dropped first under a per-section budget with a count of what was left out, and the previous narrative survives as a short tail excerpt. Pi's own free-text compaction summaries are carried as a head-and-tail excerpt.
|
|
131
134
|
|
|
132
135
|
The model recalls by itself when it needs the content:
|
|
133
136
|
|
|
@@ -135,9 +138,10 @@ The model recalls by itself when it needs the content:
|
|
|
135
138
|
zip_recall({ handles: ["k3f9a0x1qz", "m2b7c4d8ww"] })
|
|
136
139
|
zip_recall({ handle: "k3f9a0x1qz", grep: "ERROR|FAIL" })
|
|
137
140
|
zip_recall({ handle: "k3f9a0x1qz", range: "120-240" })
|
|
141
|
+
zip_recall({ grep: "ECONNREFUSED" }) // no handle: searches every earlier tool output
|
|
138
142
|
```
|
|
139
143
|
|
|
140
|
-
Recall is batched, exact, and also works for outputs from before a compaction or a pi-zip summary (summaries carry a handle table and a budgeted index of older handles). Output comes back in pages of 20,000 characters; the page says how to continue:
|
|
144
|
+
Recall is batched, exact, and also works for outputs from before a compaction or a pi-zip summary (summaries carry a handle table and a budgeted index of older handles). Each recalled section is headed by the handle and the command or path that produced it, so a look-alike output (the same command run again, another slice of the same file) is easy to spot. Without a handle, `grep` searches every earlier tool output of the session (zip_recall results excepted, outputs from before a compaction included), newest output first, each match under its output's handle and command: this also reaches an output whose handle a long summary no longer lists. Output comes back in pages of 20,000 characters (8,000 for a search without a handle; a later page of a search covers the same outputs as its first page, so new output arriving in between does not shift it); the page says how to continue:
|
|
141
145
|
|
|
142
146
|
```
|
|
143
147
|
zip_recall({ handle: "k3f9a0x1qz", offset: 20000 }) // next page, by characters
|
|
@@ -146,17 +150,31 @@ zip_recall({ handle: "k3f9a0x1qz", offset: 20000, limit: 50000 }) // up to 50,00
|
|
|
146
150
|
|
|
147
151
|
Paging is by characters, so even one 45,000-character line can be read in full, and the pages add up to the original byte for byte. `offset` and `limit` also page a `grep` or `range` selection. `grep` is a case-insensitive regular expression; patterns that could backtrack badly (nested quantifiers, quantified alternation, back-references, very long patterns) are searched as literal text instead.
|
|
148
152
|
|
|
153
|
+
### Recall through files
|
|
154
|
+
|
|
155
|
+
When a tool allowlist hides `zip_recall` (`pi --tools read,grep,bash`, sub-agent launchers) but `read`, `grep` or `bash` is active, pi-zip folds exactly as in full mode and points the model at a file instead of the tool:
|
|
156
|
+
|
|
157
|
+
```
|
|
158
|
+
Original kept byte for byte, even after summaries or compaction, in ~/.cache/pi-zip/recall/01a120cb-3d40-75a6-aeac-320e71967606/k3f9a0x1qz.txt: read it (offset/limit for pages) or grep it; prefer that to re-running or re-reading the source (output may differ). Do not guess its content.
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
The sentence names only the tools that are active (with bash alone: `sed -n` to page, `grep -n` to search), and summaries name the same files. A file does not exist until a tool asks for it: when `read`, `grep`, `find` or `ls` gets such a path (whatever comes before `pi-zip/recall/`: another machine, another user, a forked session), or a `bash` command contains one, pi-zip writes the original from the session (the current branch) and points the call at this machine's copy. A `grep`, `ls` or `find` of the session's directory writes every output the model no longer sees in full (folded or summarized away), so a grep over the directory searches them all. A read of a recall file is a recall: it is never folded and counts as an unzip, and the transcript draws it like zip_recall's row (`ƶ unzip …`). `/zip` says `on` (it is a full mode) and has one more row: `unzip from files in ~/.cache/pi-zip/recall (a tool allowlist hides zip_recall)`, naming where the copies really are (see below; a narrow terminal drops the part in brackets). In a bash command the path is recognised as its own word, also after `X=`, `--file=` or inside `{a,b}`.
|
|
162
|
+
|
|
163
|
+
The copies live in `$XDG_CACHE_HOME/pi-zip/recall/<session id>/` (by default `~/.cache/pi-zip/recall/`, a place backups and dotfile sync conventionally skip; a relative `XDG_CACHE_HOME` is ignored, as the XDG spec says, since it would put them in the project folder): directories 0700, files 0600. The state line and `/zip` show the place in use. The same text is already in the session file, which Pi writes 0644. Only what the model asks for is copied, and every copy can be written again from the session at any time, so copies older than 14 days are deleted at the next session start; that cleanup touches nothing but pi-zip's own copies (`<handle>.txt` files in session-id directories, never through a symlink). A copy that no longer holds the original byte for byte (a bash command appended to it or edited it) is written again the next time a tool names it. If you delete a session file to get rid of a secret it contained, delete `<that place>/<its session id>/` too, or it stays until then. Without a session file (`pi --no-session`) nothing is copied: the outputs are in no file, so pi-zip keeps them off the disk too, folds only re-readable outputs and says so (`reread-only`). Without pi-zip loaded nothing is written, just as zip_recall would be missing.
|
|
164
|
+
|
|
149
165
|
## Coexistence
|
|
150
166
|
|
|
151
|
-
If another extension that manages context is loaded (for example billion-context, magic-context, pi-smart-compact, pi-hot-compact, pi-context-prune), pi-zip pauses folding, says so once, and keeps only its request guard. Two writers on the same view give unpredictable results. `/zip
|
|
167
|
+
If another extension that manages context is loaded (for example billion-context, magic-context, pi-smart-compact, pi-hot-compact, pi-context-prune), pi-zip pauses folding, says so once, and keeps only its request guard. Two writers on the same view give unpredictable results. `/zip` shows `paused`.
|
|
168
|
+
|
|
169
|
+
An extension that only writes Pi's compaction summaries (a `session_before_compact` handler, like Pi's own `custom-compaction` example) does not pause pi-zip, but Pi runs its handler for every compaction, the one that puts pi-zip's prepared summary in at a cold return included. So pi-zip watches that compaction: if another extension's summary went in instead of pi-zip's, pi-zip adds none of its own on top (its prepared summary is discarded, its cost still counted); if the compaction is not done within 2 s (that handler is writing a summary of its own), pi-zip cancels it before your prompt goes on, so nothing can land in the middle of the run, and sends its summary with the first request instead; and if it was cancelled, or took longer than pi-zip's own part ever does (250 ms), pi-zip stops going through Pi's compaction for the rest of the process, so that handler's work is not paid for and waited for at every return. When such an extension's summary is already in the session, pi-zip does not go through Pi's compaction at all. Either way, summaries then ride along with the first request, as on an older Pi. A handler that cancels manual compactions shows Pi's red "Compaction cancelled" line at most once per process.
|
|
152
170
|
|
|
153
171
|
The guard has two parts. Before saving a fold or a summary it checks that the edit itself would not leave a tool result without its tool call (failed or aborted assistant turns, which the provider layer drops anyway, are ignored, as is any oddity the session already had). And before a request is sent it repairs a tool result that lost its tool call, instead of letting the provider reject the request and brick the session. The repair understands the request shapes of Pi's Anthropic, OpenAI chat completions (and Mistral), OpenAI Responses (Azure, Codex), Google Gemini/Vertex and Bedrock Converse providers; a request of any other shape is passed through untouched.
|
|
154
172
|
|
|
155
173
|
## FAQ
|
|
156
174
|
|
|
157
|
-
**Will it save money?** Mostly on cold returns, which is where a long session pays for a full cache rewrite. While the cache is warm it edits only when the inequality above says the rewrite pays back (large contexts, near Pi's compaction trigger, outputs 60+ requests old). `/zip
|
|
175
|
+
**Will it save money?** Mostly on cold returns, which is where a long session pays for a full cache rewrite. While the cache is warm it edits only when the inequality above says the rewrite pays back (large contexts, near Pi's compaction trigger, outputs 60+ requests old). `/zip` shows what this session has folded (and roughly how many tokens that removed), its summaries with what their model calls cost, and its recalls. The numbers live in the session file, so a restart does not reset them.
|
|
158
176
|
|
|
159
|
-
**Does it cost extra?** Planning is free. A summary is one extra model call (the current model, no tools, no prompt cache), shown in `/zip
|
|
177
|
+
**Does it cost extra?** Planning is free. A summary is one extra model call (the current model, no tools, no prompt cache), shown in `/zip`. A background summary you never use (you came back while the cache was warm) is counted there too. Recalled content re-enters the context at normal prices.
|
|
160
178
|
|
|
161
179
|
**Can the model lose information?** Folded outputs are replaced by a placeholder with key lines and a handle, and the placeholder tells the model not to guess. Summaries quote your requests verbatim and never paraphrase the model's reasoning. If the model ignores the handle and guesses, that is a model failure pi-zip cannot catch; this is why only re-readable outputs of the previous turn are folded at a cold return; other outputs of the protected turns wait until they are 60 requests old.
|
|
162
180
|
|
|
@@ -177,7 +195,7 @@ bun run build-check # bun
|
|
|
177
195
|
|
|
178
196
|
`test/pi-integration.test.ts` runs the extension through Pi's real session manager, extension loader and `emitContext`, with a mid-conversation system update and an aborted tool-call turn in the session, and checks that the request is byte-identical before and after turn_end persists the edits.
|
|
179
197
|
|
|
180
|
-
Environment overrides exist for tests only and are not part of the product surface: `PI_ZIP_TTL_SECS` (cache TTL), `PI_ZIP_COLD_CAP` (fold target in tokens, default 40000), `PI_ZIP_FOLD_MIN` (smallest output worth folding, default 500 tokens), `PI_ZIP_KEEP_LINES`, `PI_ZIP_MIN_GAIN` (summary gain floor of the legacy rule used only when nothing is known about prices, default 10000), `PI_ZIP_INTURN_AGE` (age in assistant requests from which any output of the protected turns may fold, default 60; 0 = never), `PI_ZIP_CACHE_STATS=<path>` (where the learned cache survival lives; benches isolate it), `PI_ZIP_OFF=1` (register nothing), `PI_ZIP_LEDGER=<path>` (append a JSON line per decision; `prompt` records `ttlMs` and `ttlSource`, `declared` or `observed`; including the summary call's token usage; `cold_plan` records the calibration as `
|
|
198
|
+
Environment overrides exist for tests only and are not part of the product surface: `PI_ZIP_TTL_SECS` (cache TTL), `PI_ZIP_COLD_CAP` (fold target in real tokens of the whole context, default 40000; at least half of it is always left to the conversation above the system prompt and tools), `PI_ZIP_FOLD_MIN` (smallest output worth folding, default 500 tokens), `PI_ZIP_KEEP_LINES`, `PI_ZIP_MIN_GAIN` (summary gain floor of the legacy rule used only when nothing is known about prices, default 10000), `PI_ZIP_INTURN_AGE` (age in assistant requests from which any output of the protected turns may fold, default 60; 0 = never), `PI_ZIP_CACHE_STATS=<path>` (where the learned cache survival lives; benches isolate it), `PI_ZIP_OFF=1` (register nothing), `PI_ZIP_NATIVE_MS` (how long pi-zip's own compaction may take before it is cancelled, default 2000), `PI_ZIP_LEDGER=<path>` (append a JSON line per decision; `prompt` records `ttlMs` and `ttlSource`, `declared` or `observed`; including the summary call's token usage; `cold_plan` records the calibration as `O`, `c`, `oSource`, `calReal`, `calEst`, `calSource`, next to the real `ctxBefore` and `ctxAfter`; `prompt` also records `gapS`, `pWarm`, `survSrc` and `cls`; `law` records every evaluation with `g`, `wr` (w/r), `pWarm`, `survSrc`, `prSrc` and per step `B`, `A`, `T` (= A), `Tsuf` (the suffix after the earliest edit, measurement only), `phi`, `K`, `eta`; `b2b_skip` marks a warm edit skipped right after an edited request; `cache_sample` records each learned survival observation).
|
|
181
199
|
|
|
182
200
|
Internally, `RELAX_PREV_TURN` in `src/plan.ts` selects whether re-readable outputs of the previous user turn may be folded on a cold return or a return after the declared TTL (default `true`; other warm plans never fold them).
|
|
183
201
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-zip",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0-rc.1",
|
|
4
4
|
"description": "Keeps long Pi sessions cheap without losing anything: folds old tool output only when the prompt cache has already gone cold, and every fold can be recalled byte for byte. Zero config.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
package/src/cache.ts
CHANGED
|
@@ -44,12 +44,34 @@ export function observedTier(branch: Any[], key: string): "1h" | "5m" | undefine
|
|
|
44
44
|
|
|
45
45
|
export interface TtlInfo { ms: number; source: "declared" | "observed"; note?: string }
|
|
46
46
|
|
|
47
|
-
/**
|
|
48
|
-
|
|
47
|
+
/** The cache tier a request asked for, read from the payload Pi built (null: the payload does not say). Anthropic: a 1-hour
|
|
48
|
+
* cache_control ttl is the long tier, cache_control without one the short tier; OpenAI: prompt_cache_retention "24h" is long. */
|
|
49
|
+
export function payloadRetention(p: Any): "short" | "long" | null {
|
|
50
|
+
if (!p || typeof p !== "object") return null;
|
|
51
|
+
if (typeof p.prompt_cache_retention === "string") return p.prompt_cache_retention === "24h" ? "long" : "short";
|
|
52
|
+
let seen = false;
|
|
53
|
+
const one = (x: Any): boolean => {
|
|
54
|
+
const cc = x?.cache_control;
|
|
55
|
+
if (!cc || typeof cc !== "object") return false;
|
|
56
|
+
seen = true;
|
|
57
|
+
return cc.ttl === "1h";
|
|
58
|
+
};
|
|
59
|
+
for (const list of [p.system, p.tools]) if (Array.isArray(list)) for (const b of list) if (one(b)) return "long";
|
|
60
|
+
if (Array.isArray(p.messages)) for (const m of p.messages) {
|
|
61
|
+
if (one(m)) return "long";
|
|
62
|
+
if (Array.isArray(m?.content)) for (const b of m.content) if (one(b)) return "long";
|
|
63
|
+
}
|
|
64
|
+
return seen ? "short" : null;
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
/** Declared TTL, corrected by what the provider was seen to write. PI_ZIP_TTL_SECS wins. `note` = the one-time mismatch notice text.
|
|
68
|
+
* `tier` = the retention the request used (the provider-scoped PI_CACHE_RETENTION, or the request payload itself; run.ts); without it,
|
|
69
|
+
* the process environment decides, as before. */
|
|
70
|
+
export function resolveTtl(model: Any, branch: Any[] = [], tier?: "short" | "long"): TtlInfo {
|
|
49
71
|
if (process.env.PI_ZIP_TTL_SECS) return { ms: envInt("TTL_SECS", 300) * 1000, source: "declared" };
|
|
50
72
|
const pc = model?.promptCache;
|
|
51
|
-
const
|
|
52
|
-
const
|
|
73
|
+
const long = tier ? tier === "long" : process.env.PI_CACHE_RETENTION === "long";
|
|
74
|
+
const declared = cacheTtlMs(pc, DEFAULT_TTL_MS, { cacheRetention: long ? "long" : "short" });
|
|
53
75
|
const seen = declared > 0 ? observedTier(branch, modelKey(model)) : undefined;
|
|
54
76
|
const mismatch = (req: string, got: string, ms: number): TtlInfo => ({ ms, source: "observed", note: `${PRODUCT} · requested ${req} prompt cache, provider wrote ${got} · using ${got}` });
|
|
55
77
|
if (long && seen === "5m") return mismatch("1h", "5m", cacheTtlMs(pc, DEFAULT_TTL_MS, { cacheRetention: "short" }));
|
|
@@ -57,7 +79,7 @@ export function resolveTtl(model: Any, branch: Any[] = []): TtlInfo {
|
|
|
57
79
|
return { ms: declared, source: "declared" };
|
|
58
80
|
}
|
|
59
81
|
|
|
60
|
-
export const ttlFor = (model: Any, branch: Any[] = []): number => resolveTtl(model, branch).ms;
|
|
82
|
+
export const ttlFor = (model: Any, branch: Any[] = [], tier?: "short" | "long"): number => resolveTtl(model, branch, tier).ms;
|
|
61
83
|
|
|
62
84
|
/** Cold = a prior request exists and nothing touched the cache for longer than the TTL. No prior request counts as warm. */
|
|
63
85
|
export const isColdByTtl = (lastActivityMs: number, nowMs: number, ttlMs: number): boolean => lastActivityMs > 0 && nowMs - lastActivityMs > ttlMs;
|
|
@@ -69,18 +91,33 @@ const tsOf = (en: Any, inner?: Any): number => {
|
|
|
69
91
|
return Number.isFinite(t) ? t : 0;
|
|
70
92
|
};
|
|
71
93
|
|
|
94
|
+
/** A response that reached no cache: an error that reports no usage (the provider refused the request). One that reports usage was
|
|
95
|
+
* read: it touched the cache. An abort (Esc) comes after the provider read and wrote the prompt: it is a touch, with or without usage. */
|
|
96
|
+
export const refused = (m: Any): boolean => {
|
|
97
|
+
if (m?.role !== "assistant" || m.stopReason !== "error") return false;
|
|
98
|
+
const u = m.usage;
|
|
99
|
+
return !((Number(u?.input) || 0) + (Number(u?.cacheRead) || 0) + (Number(u?.cacheWrite) || 0) > 0);
|
|
100
|
+
};
|
|
101
|
+
|
|
72
102
|
/**
|
|
73
103
|
* Last time anything touched the cache, according to the branch: the newest message, or the newest cache-warm usage entry
|
|
74
104
|
* (Pi's own idle refresh persists `{ type: "usage", kind: "cache_warm" }`, and a refresh keeps the entry alive for another TTL).
|
|
105
|
+
* A request the provider refused (an error with no usage) touched nothing, and neither did the messages that made it (the prompt or
|
|
106
|
+
* the tool results it carried, written before it failed): the clock stays at the last request that was read.
|
|
75
107
|
* Fresh process: err on the warm side (0 = unknown).
|
|
76
108
|
*/
|
|
77
109
|
export function lastMessageMs(branch: Any[]): number {
|
|
78
|
-
let last = 0;
|
|
110
|
+
let last = 0, pending = 0; // pending: the newest non-assistant message since the last response, counted once its response is known
|
|
79
111
|
for (const en of branch) {
|
|
80
|
-
if (en?.type === "message" && en.message)
|
|
81
|
-
|
|
112
|
+
if (en?.type === "message" && en.message) {
|
|
113
|
+
const t = tsOf(en, en.message);
|
|
114
|
+
if (en.message.role === "assistant") {
|
|
115
|
+
if (!refused(en.message)) last = Math.max(last, pending, t);
|
|
116
|
+
pending = 0;
|
|
117
|
+
} else pending = Math.max(pending, t);
|
|
118
|
+
} else if (en?.type === "usage" && en.kind === "cache_warm") last = Math.max(last, tsOf(en));
|
|
82
119
|
}
|
|
83
|
-
return last;
|
|
120
|
+
return Math.max(last, pending); // trailing messages with no response yet (the run in progress) count, as before
|
|
84
121
|
}
|
|
85
122
|
|
|
86
123
|
/** The newest assistant message that really ran (errors and aborts may never have reached a cache): its provider/model and prompt size. */
|
|
@@ -102,11 +139,12 @@ export interface ColdInfo { cold: boolean; reason: string; ttl: TtlInfo; pWarm:
|
|
|
102
139
|
|
|
103
140
|
/** Cold by time, or because the model changed: a cache entry belongs to one provider and model. `learned` = P(warm) after a gap from the
|
|
104
141
|
* survival the provider was seen to have (learn.ts); without it, or under PI_ZIP_TTL_SECS (tests), the TTL decides. cold = P(warm) < 0.5. */
|
|
105
|
-
export function detectCold(model: Any, lastReqMs: number, branch: Any[], nowMs = Date.now(), lastModel = "", learned?: Survival): ColdInfo {
|
|
106
|
-
const ttl = resolveTtl(model, branch);
|
|
142
|
+
export function detectCold(model: Any, lastReqMs: number, branch: Any[], nowMs = Date.now(), lastModel = "", learned?: Survival, tier?: "short" | "long"): ColdInfo {
|
|
143
|
+
const ttl = resolveTtl(model, branch, tier);
|
|
107
144
|
const last = Math.max(lastReqMs, lastMessageMs(branch));
|
|
108
145
|
const prev = lastModel || lastModelInBranch(branch);
|
|
109
|
-
|
|
146
|
+
// a request that went out, even one the provider refused, was made on another model: the new model's cache is not that one
|
|
147
|
+
if (prev && modelKey(model) && prev !== modelKey(model)) return { cold: true, reason: `model switch ${prev} -> ${modelKey(model)}`, ttl, pWarm: 0, src: "model switch", gapS: null, pastTtl: false };
|
|
110
148
|
if (!last) return { cold: false, reason: "no prior request", ttl, pWarm: 1, src: "no prior request", gapS: null, pastTtl: false };
|
|
111
149
|
const gapS = (nowMs - last) / 1000;
|
|
112
150
|
const byTtl = isColdByTtl(last, nowMs, ttl.ms);
|
package/src/guard.ts
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
// Guard (F13, I5): validate an edit set before committing it, and repair orphan tool results in the outgoing payload.
|
|
2
|
-
import { RELAX_PREV_TURN, settings, type Block } from "./plan.ts";
|
|
2
|
+
import { finalAnswerIdx, RELAX_PREV_TURN, settings, type Block } from "./plan.ts";
|
|
3
3
|
import type { Any } from "./util.ts";
|
|
4
4
|
|
|
5
5
|
const PROTECT_USER_TURNS = 2;
|
|
@@ -10,6 +10,7 @@ export interface EditSet {
|
|
|
10
10
|
recover?: Map<number, string | undefined>; // block idx -> "rereadable" | "nonrereadable" (needed for folds in the previous user turn)
|
|
11
11
|
relax?: boolean; // default RELAX_PREV_TURN
|
|
12
12
|
inturnAge?: number; // default settings().inturnAge: an output (any class) at least this many assistant requests old may fold in a protected turn (0 = never)
|
|
13
|
+
ahead?: boolean; // a lookahead cut (plan.ts LOOKAHEAD): it may lie inside the previous user turn, never past that turn's final answer
|
|
13
14
|
}
|
|
14
15
|
|
|
15
16
|
/**
|
|
@@ -63,7 +64,7 @@ export function validateEdits(blocks: Block[], plan: EditSet, userTurns: number)
|
|
|
63
64
|
}
|
|
64
65
|
if (plan.cut !== null) {
|
|
65
66
|
if (plan.cut < 1 || plan.cut >= blocks.length) return `bad cut ${plan.cut}`;
|
|
66
|
-
if (limit >= 0 && plan.cut > limit) return `cut ${plan.cut} inside the last ${PROTECT_USER_TURNS} user turns`;
|
|
67
|
+
if (limit >= 0 && plan.cut > limit && !(plan.ahead && plan.cut <= finalAnswerIdx(blocks, userTurns - 1))) return `cut ${plan.cut} inside the last ${PROTECT_USER_TURNS} user turns`;
|
|
67
68
|
const kb = blocks[plan.cut];
|
|
68
69
|
if (!kb.entryId || (kb.kind !== "user" && kb.kind !== "assistant")) return `cut ${plan.cut} is not at a user/assistant message`;
|
|
69
70
|
// only what the cut itself would create is our problem; a session that was already odd stays as odd as it was
|
package/src/index.ts
CHANGED
|
@@ -4,8 +4,9 @@ import type { Any } from "./util.ts";
|
|
|
4
4
|
import { type NoticeData, registerZipCommand, renderNotice } from "./notice.ts";
|
|
5
5
|
import { RECALL_TOOL } from "./placeholder.ts";
|
|
6
6
|
import { registerRecallTool } from "./recall.ts";
|
|
7
|
+
import { parseRecallPath } from "./recallfile.ts";
|
|
7
8
|
import { NOTICE_CUSTOM, Zip } from "./run.ts";
|
|
8
|
-
import { markFirstLine, measure, renderCard, renderState, setMeasure } from "./ui.ts";
|
|
9
|
+
import { markFirstLine, measure, recallCallLine, renderCard, renderState, setMeasure } from "./ui.ts";
|
|
9
10
|
|
|
10
11
|
export default function piZip(pi: ExtensionAPI) {
|
|
11
12
|
if (process.env.PI_ZIP_OFF === "1") return; // test only: behave exactly as if not installed (registers nothing)
|
|
@@ -14,36 +15,61 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
14
15
|
registerZipCommand(pi, zip);
|
|
15
16
|
try {
|
|
16
17
|
// Everything pi-zip shows lives in the transcript as custom entries (never part of the model's context): fold/summary
|
|
17
|
-
// notices, state lines and the /zip
|
|
18
|
+
// notices, state lines and the /zip card. Widths are measured with Pi's own pi-tui once it has loaded.
|
|
18
19
|
import("@earendil-works/pi-tui").then((m: Any) => {
|
|
19
20
|
if (typeof m?.visibleWidth === "function" && typeof m?.truncateToWidth === "function")
|
|
20
21
|
setMeasure({ vw: m.visibleWidth, cut: (s: string, w: number) => (m.visibleWidth(s) <= w ? s : w <= 0 ? "" : m.truncateToWidth(s, w, "…")) });
|
|
21
22
|
}, () => {});
|
|
23
|
+
// A click on a notice toggles its detail, like Pi's own compaction row; ctrl+o (Pi's global expand) still wins: a click
|
|
24
|
+
// only overrides the global state it was made under. Keyed by entry id because Pi rebuilds the component on every toggle.
|
|
25
|
+
const clicked = new Map<string, { open: boolean; under: boolean }>();
|
|
22
26
|
pi.registerEntryRenderer?.(NOTICE_CUSTOM, (entry: Any, o: Any, theme: Any) => {
|
|
23
27
|
const d = entry?.data;
|
|
24
28
|
if (!d?.text) return undefined;
|
|
29
|
+
const global = !!o?.expanded;
|
|
30
|
+
const key = typeof entry?.id === "string" ? entry.id : "";
|
|
31
|
+
const open = () => {
|
|
32
|
+
const c = key ? clicked.get(key) : undefined;
|
|
33
|
+
return c && c.under === global ? c.open : global;
|
|
34
|
+
};
|
|
35
|
+
const expandable = d.v === 2 && d.kind !== "state" && d.kind !== "card" && !!(d.why || d.items?.length);
|
|
25
36
|
return {
|
|
26
37
|
render: (width: number) => {
|
|
27
38
|
try {
|
|
28
39
|
if (d.v !== 2) return [theme.fg("dim", measure.cut(d.text, width))];
|
|
29
40
|
if (d.kind === "state") return renderState(d.word, d.reason ?? "", width, theme, measure);
|
|
30
41
|
if (d.kind === "card" && d.card) return renderCard(d.card, width, theme, measure);
|
|
31
|
-
|
|
42
|
+
// once the request that carried the plan has been answered, "after" is its real size (older notices keep their number)
|
|
43
|
+
const real = typeof d.reqTs === "number" ? zip.realSizeAfter(d.reqTs) : undefined;
|
|
44
|
+
return renderNotice((real ? { ...d, after: real } : d) as NoticeData, open(), width, theme, measure.vw);
|
|
32
45
|
} catch {
|
|
33
46
|
return [];
|
|
34
47
|
}
|
|
35
48
|
},
|
|
49
|
+
handleMouse: (ev: Any) => {
|
|
50
|
+
if (!expandable || !key || ev?.type !== "click" || ev?.button !== "left") return undefined;
|
|
51
|
+
clicked.set(key, { open: !open(), under: global });
|
|
52
|
+
return { handled: true, render: true };
|
|
53
|
+
},
|
|
36
54
|
invalidate: () => {},
|
|
37
55
|
};
|
|
38
56
|
});
|
|
39
57
|
zip.entryRenderer = typeof pi.registerEntryRenderer === "function";
|
|
40
|
-
// a folded output keeps its place in the transcript; its call row gets
|
|
58
|
+
// a folded output keeps its place in the transcript; its call row gets "ƶ zipped" right after its own text
|
|
41
59
|
pi.registerToolRenderer?.((toolName: string, next: () => Any) => {
|
|
42
60
|
const base = next();
|
|
43
61
|
if (toolName === RECALL_TOOL || typeof base?.renderCall !== "function") return base;
|
|
44
62
|
return {
|
|
45
63
|
...base,
|
|
46
64
|
renderCall: (args: Any, theme: Any, c: Any) => {
|
|
65
|
+
// file recall: a read or grep of a recall file (or of the session's recall directory) reads like zip_recall's row
|
|
66
|
+
const ref = toolName === "read" || toolName === "grep" ? parseRecallPath(args?.path) : null;
|
|
67
|
+
if (ref && (ref.handle || toolName === "grep")) {
|
|
68
|
+
const off = typeof args.offset === "number" ? args.offset : undefined, lim = typeof args.limit === "number" ? args.limit : undefined;
|
|
69
|
+
const range = toolName === "read" && (off !== undefined || lim !== undefined) ? `${off ?? 1}-${lim !== undefined ? (off ?? 1) + lim - 1 : ""}` : undefined;
|
|
70
|
+
const a = { handle: ref.handle, grep: toolName === "grep" && typeof args.pattern === "string" ? args.pattern : undefined, range };
|
|
71
|
+
return { render: (width: number) => recallCallLine(a, Math.max(1, width), theme, measure, (h) => zip.labelOf(h)), invalidate: () => {} };
|
|
72
|
+
}
|
|
47
73
|
const comp = base.renderCall(args, theme, c);
|
|
48
74
|
const id = c?.toolCallId;
|
|
49
75
|
if (!comp || typeof comp.render !== "function" || typeof id !== "string") return comp;
|
|
@@ -68,11 +94,15 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
68
94
|
/* older Pi: notices fall back to the status line */
|
|
69
95
|
}
|
|
70
96
|
// A tool allowlist (`pi --tools read,bash`, sub-agent launchers) replaces the whole selection and Pi then does not even register
|
|
71
|
-
// zip_recall; folds of outputs the model could not get back would be lost, so only rereadable outputs fold then (plan rereadOnly)
|
|
97
|
+
// zip_recall; folds of outputs the model could not get back would be lost, so only rereadable outputs fold then (plan rereadOnly)...
|
|
98
|
+
// ... unless read, grep or bash is active: then folds are as in full mode and placeholders name a recall file (recallfile.ts) that the
|
|
99
|
+
// tool_call hook below writes from the session when a tool asks for it.
|
|
72
100
|
const checkRecall = (ctx: Any) => {
|
|
73
101
|
try {
|
|
74
102
|
const active = pi.getActiveTools?.();
|
|
75
|
-
if (Array.isArray(active))
|
|
103
|
+
if (!Array.isArray(active)) return;
|
|
104
|
+
const tools = { read: active.includes("read"), grep: active.includes("grep"), bash: active.includes("bash") };
|
|
105
|
+
zip.setRecallMode(active.includes(RECALL_TOOL) ? "tool" : tools.read || tools.grep || tools.bash ? "file" : "reread", tools, ctx);
|
|
76
106
|
} catch {
|
|
77
107
|
/* never in the way */
|
|
78
108
|
}
|
|
@@ -82,17 +112,24 @@ export default function piZip(pi: ExtensionAPI) {
|
|
|
82
112
|
checkRecall(ctx);
|
|
83
113
|
zip.refreshFolded(ctx);
|
|
84
114
|
zip.welcome(ctx);
|
|
115
|
+
zip.resumeAway(ctx); // a resumed session: arm the away timer from its last request, as the settle after it would have
|
|
85
116
|
return r;
|
|
86
117
|
});
|
|
87
|
-
pi.on("before_agent_start", (
|
|
118
|
+
pi.on("before_agent_start", async (e, ctx) => {
|
|
88
119
|
checkRecall(ctx);
|
|
89
120
|
zip.refreshFolded(ctx); // the branch may have changed (/tree, fork)
|
|
90
|
-
|
|
121
|
+
await zip.beforeAgentStart(ctx, e); // may persist a summary prepared while the user was away through Pi's compaction (Pi is idle here)
|
|
122
|
+
return undefined;
|
|
91
123
|
});
|
|
124
|
+
// answers only the compaction pi-zip itself just asked for; every other compaction (Pi's, /compact, other extensions') gets undefined
|
|
125
|
+
pi.on("session_before_compact", (e) => zip.sessionBeforeCompact(e));
|
|
92
126
|
pi.on("context_with_system", (e, ctx) => zip.context(e, ctx)); // the complete transcript: system messages stay where Pi put them
|
|
93
127
|
pi.on("turn_end", (e, ctx) => zip.turnEnd(e, ctx));
|
|
94
128
|
pi.on("agent_before_settle", (e, ctx) => zip.settle(e, ctx));
|
|
95
129
|
pi.on("before_provider_request", (e) => zip.providerRequest(e));
|
|
130
|
+
pi.on("tool_call", (e, ctx) => zip.toolCall(e, ctx)); // a recall file named by a built-in tool: this machine's copy, written on demand
|
|
131
|
+
// Pi's idle cache warming: observed only (the return value is undefined, so Pi's own decision stands); an older Pi never fires it
|
|
132
|
+
pi.on("cache_warming_decision", (e) => zip.cacheWarmingDecision(e));
|
|
96
133
|
pi.on("message_end", (e) => zip.messageEnd(e.message));
|
|
97
134
|
pi.on("session_shutdown", () => zip.shutdown());
|
|
98
135
|
}
|
package/src/learn.ts
CHANGED
|
@@ -26,8 +26,10 @@ export function lawPrices(cls: CacheClass | undefined, longTier: boolean): Price
|
|
|
26
26
|
}
|
|
27
27
|
|
|
28
28
|
// ---- cache survival ------------------------------------------------------------------------------------------------
|
|
29
|
-
/** Gap bins (s): [GAP_EDGES[i], GAP_EDGES[i+1]). Edges sit on the known TTL tiers (300 s, 3600 s): a deterministic TTL never splits a bin.
|
|
30
|
-
|
|
29
|
+
/** Gap bins (s): [GAP_EDGES[i], GAP_EDGES[i+1]). Edges sit on the known TTL tiers (300 s, 3600 s): a deterministic TTL never splits a bin.
|
|
30
|
+
* The last edge closes the last bin: a gap at or beyond it (2 h) is outside what was ever observed, it is not recorded and it is planned
|
|
31
|
+
* from the declared TTL (a cache seen alive at 110 min says nothing about 156 min). */
|
|
32
|
+
export const GAP_EDGES = [30, 60, 120, 180, 240, 300, 330, 360, 420, 480, 600, 900, 1200, 1800, 2700, 3600, 5400, 7200];
|
|
31
33
|
export const HALF_LIFE = 16; // observations per bin: old evidence counts half after 16 newer ones in the same bin (a provider may change its TTL)
|
|
32
34
|
// The declared TTL as pseudo-observations. Asymmetric because the signal is: a read of the re-sent prefix cannot happen on a dead cache,
|
|
33
35
|
// but warm misses do (GLM 4-9%, glm-flash ~22%, research round 5). Beyond the TTL one clean read overrides it (weight 1/4: one hit -> 0.8);
|
|
@@ -41,10 +43,17 @@ export const MIN_EXPECT = 8192;
|
|
|
41
43
|
export interface Entry { cls?: CacheClass; bins: Record<string, [number, number]>; n: number } // bin index -> [alive, dead] (decayed counts)
|
|
42
44
|
interface File { v: 1; models: Record<string, Entry> }
|
|
43
45
|
|
|
44
|
-
|
|
45
|
-
|
|
46
|
+
/** Pi's agent directory (settings.json, sessions): PI_CODING_AGENT_DIR, else ~/.pi/agent. */
|
|
47
|
+
export const agentDir = (): string => process.env.PI_CODING_AGENT_DIR || join(homedir(), ".pi", "agent");
|
|
46
48
|
|
|
47
|
-
export const
|
|
49
|
+
export const statsPath = (): string => process.env.PI_ZIP_CACHE_STATS || join(agentDir(), "pi-zip", "cache-survival.json");
|
|
50
|
+
|
|
51
|
+
/** ES2023's Array.prototype.findLastIndex (Node >= 18 and Bun have it); tsconfig's lib is ES2022, so it is typed here (types only). */
|
|
52
|
+
interface FindLastIndex { findLastIndex(predicate: (value: number, index: number) => boolean): number }
|
|
53
|
+
export const binOf = (gapS: number): number => {
|
|
54
|
+
const i = (GAP_EDGES as number[] & FindLastIndex).findLastIndex((e) => gapS >= e);
|
|
55
|
+
return i >= GAP_EDGES.length - 1 ? -1 : i; // -1 = below the first edge or at/above the last: not recorded
|
|
56
|
+
};
|
|
48
57
|
const fin = (x: Any) => typeof x === "number" && Number.isFinite(x) && x >= 0;
|
|
49
58
|
|
|
50
59
|
/** The persisted stats; a missing, unreadable or malformed file is ignored (empty). */
|
|
@@ -55,7 +64,7 @@ export function loadStats(path = statsPath()): File {
|
|
|
55
64
|
if (j?.v !== 1 || typeof j.models !== "object" || !j.models) return out;
|
|
56
65
|
for (const [k, e] of Object.entries<Any>(j.models)) {
|
|
57
66
|
const bins: Entry["bins"] = {};
|
|
58
|
-
for (const [b, c] of Object.entries<Any>(e?.bins ?? {})) if (GAP_EDGES[Number(b)] !== undefined && Array.isArray(c) && fin(c[0]) && fin(c[1])) bins[b] = [c[0], c[1]];
|
|
67
|
+
for (const [b, c] of Object.entries<Any>(e?.bins ?? {})) if (Number(b) < GAP_EDGES.length - 1 && GAP_EDGES[Number(b)] !== undefined && Array.isArray(c) && fin(c[0]) && fin(c[1])) bins[b] = [c[0], c[1]];
|
|
59
68
|
out.models[k] = { bins, n: fin(e?.n) ? e.n : 0, ...(e?.cls === "explicit" || e?.cls === "automatic" ? { cls: e.cls } : {}) };
|
|
60
69
|
}
|
|
61
70
|
return out;
|
|
@@ -125,7 +134,8 @@ export function curve(e: Entry | undefined, priorS: number): { bin: number; p: n
|
|
|
125
134
|
* between the nearest observed bins below and above (an alive read at 365 s makes every shorter gap warm; a miss makes longer ones dead). */
|
|
126
135
|
export function pWarm(e: Entry | undefined, gapS: number, priorS: number): { p: number; src: "prior" | "learned" } {
|
|
127
136
|
const prior = gapS <= priorS ? 1 : 0;
|
|
128
|
-
|
|
137
|
+
// beyond the last edge: no bin, the declared TTL decides (clamped by the newest bin below, as for an unobserved bin)
|
|
138
|
+
const c = curve(e, priorS), b = gapS >= GAP_EDGES[GAP_EDGES.length - 1] ? GAP_EDGES.length - 1 : binOf(gapS);
|
|
129
139
|
const at = c.find((x) => x.bin === b);
|
|
130
140
|
if (at) return { p: at.p, src: "learned" };
|
|
131
141
|
const hi = c.filter((x) => x.bin < b).at(-1)?.p ?? 1, lo = c.find((x) => x.bin > b)?.p ?? 0;
|
|
@@ -133,7 +143,7 @@ export function pWarm(e: Entry | undefined, gapS: number, priorS: number): { p:
|
|
|
133
143
|
return { p, src: p === prior ? "prior" : "learned" };
|
|
134
144
|
}
|
|
135
145
|
|
|
136
|
-
/** /zip
|
|
146
|
+
/** /zip text: class and the learned curve ("360-420s 0.80 n1"). */
|
|
137
147
|
export function describe(e: Entry | undefined, priorS: number): string {
|
|
138
148
|
const c = curve(e, priorS);
|
|
139
149
|
const bins = c.map((x) => `${GAP_EDGES[x.bin]}-${GAP_EDGES[x.bin + 1] ?? "∞"}s ${x.p.toFixed(2)} n${Math.round(x.n * 10) / 10}`).join(", ");
|