pi-zip 0.3.0-rc.1 → 0.3.0-rc.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -43,14 +43,14 @@ One mark, ƶ, and three words: **zipped**, **unzip**, **summarized**. Only the n
43
43
  ƶ summarized 64 requests 182K → 41K 38.2 s · ready while you were away
44
44
  ```
45
45
 
46
- The second line appears once per session. A summary that made you wait says so (`waited 8.4 s`, highlighted from 2 s on). Click the line (or expand tool output with ctrl+o) to see what was zipped and why now:
46
+ The second line appears once per session. A summary that made you wait says so (`waited 8.4 s`, highlighted from 2 s on). Click the line (or expand tool output with ctrl+o) to see what was zipped and why now (a cold return says `cache cold · away 47 min` or `cache cold · model switched`, a warm edit `cache warm · worth it at this size`):
47
47
 
48
48
  ```
49
49
  ƶ zipped 12 old outputs 74K → 43K 6 ms
50
50
  bash npm test 14K k3x9q2m7ab
51
51
  read src/payment.ts 9.1K p8d2x1qa0m
52
52
  … 10 more
53
- away 47 min
53
+ cache cold · away 47 min
54
54
  ```
55
55
 
56
56
  These lines are saved in the session, so they are still there after a restart, but they are never sent to the model. On narrow terminals the time goes first, then the words; the numbers always stay. A summary still shows up as Pi's own `[compaction]` block as well; the ƶ line next to it tells you who made it. The pi-zip numbers are real tokens, the same scale as Pi's context meter (the meter adds the newest reply's output on top); the `Compacted from N tokens` figure in Pi's block is pi-zip's own size before for a summary prepared while you were away (that one goes in through Pi's compaction, before your prompt); for any other summary it is Pi's own estimate taken when the entry is saved, so it may not match.
@@ -82,23 +82,35 @@ The first one appears once per machine. While a summary started in the backgroun
82
82
 
83
83
  ```
84
84
  ƶ pi-zip on
85
- model anthropic/claude-sonnet-5-5
86
- cache ██████████████▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ expires after ~5 min idle
87
- away 1m 5m 15m 1h 2h
88
- session 56 zipped ~310K · 2 summaries $0.41 · 3 unzips
85
+ model zhipu/glm-5.3
86
+ cache each dot is one time you came back
87
+ 6 5
88
+ ● ● ● ●
89
+ ● ● ● ● ●
90
+ ● ● ● ● ● ● still there 19
91
+ ━━━━━━━━━━━━━━━━┅┅┅┅┅┅┅┅┅┅───────────────────────────────────
92
+ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ ○ gone 74
93
+ ○ ○ ○ ○ ○ ○ ○ ○ ○
94
+ ○ ○ ○ ○ ○ ○ ○ ○
95
+ ○ 12 7 9 15 6 11 5
96
+ away 30s 1m 2m 5m 10m 30m 1h 2h
97
+ under 2 min: usually still there · 2–5 min: not sure yet · over 5 min: gone
98
+ session 30 zipped ~310K · 8 summaries $1.75 · 2 unzips
99
+ last 15:22 zipped 1 old output in 4 ms · cache cold · away 13 min
89
100
  ```
90
101
 
102
+
91
103
  ƶ (U+01B6) is a plain Latin letter, one column wide everywhere, CJK terminals included. Most coding fonts have it; where one does not (Hack, Source Code Pro), the terminal draws it from a fallback font.
92
104
 
93
- `expires after ~5 min idle` is the one thing worth knowing: come back sooner and pi-zip touches nothing; come back later and the old outputs get zipped, at no extra cost. The chart draws the same fact: across (`away`) is how long you were gone, from 30 s to 2 h on a log scale; up (`cache`) is how likely the cache still holds your conversation then. Bars in the accent colour are still there, dim ones are gone. It starts from the model's declared cache lifetime (a cliff at 5 min for Claude) and turns into `(measured)` once the provider's own replies have shown how long its cache really lasts (see below); a cache that fades out slowly draws a slope. In a narrow terminal the chart is left out and the sentence stays whole (the model name goes first). Notices state facts (how long you were away), not guesses: whether the cache really had expired is read from the provider's reply afterwards, so a wrong guess is not repeated.
105
+ The chart is the one thing worth knowing: come back sooner than the solid part of the axis and pi-zip touches nothing; come back later and the old outputs get zipped, at no extra cost. Across (`away`) is how long you were gone, from 30 s to 2 h on a log scale. Every dot is one time you came back after that long: a filled dot above the axis is a time the cache was still there, an empty one below it a time it was gone, so the picture is your own history with this model, and a few gone samples take only a row or two. A column shows at most 4 dots a side; beyond that it shows 3 dots and the number at the far end of the stack (`12` = twelve times in that gap range). The axis is solid up to the longest gap that was (nearly) always safe, dashed while it is not sure, thin after it is mostly gone; the sentence under it says the same in words (`usually still there` when even the safe part had some misses). The counts are the learned ones, kept with older evidence halved after 16 newer observations of the same gap range, so they can be a little below the number of times you actually came back. Until something is learned there is no chart, one line only (`expires after ~5 min idle · a chart appears once you come back after a break`, from the model's declared cache lifetime; Claude's is 5 min). The chart needs 72 columns; where the legend does not fit beside it, it moves to one line under it, and in a narrower terminal only the sentence is shown. Notices state facts (how long you were away), not guesses: whether the cache really had expired is read from the provider's reply afterwards, so a wrong guess is not repeated.
94
106
 
95
107
  ## The three rules
96
108
 
97
109
  1. **Nothing is lost.** The session file stays the single source of truth. pi-zip only changes the view sent to the model, never deletes anything. User messages, tool-call arguments, the system prompt, tool definitions and thinking are never rewritten. Every folded block carries a handle and shows its key lines (errors, ids, first and last line).
98
- 2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens of context in one step, system prompt and tool definitions included, and every request of that turn sends the same bytes. No edit can shrink the system prompt and the tools, so once they take more than half of the 40K (a large tool set) the target becomes them plus 20K of conversation: with a 50K tool set the context lands near 70K. It is never a 40K that no edit can reach (that folds everything and summarises at every return), and it does not keep a full 40K of conversation on top either (on a live bench with a 36K prefix that cost 13% more, with no quality difference measured). While the cache is warm it does nothing, with one exception, the warm valve: above that target it applies that plan, minus the previous user turn (a warm edit never folds the turn you just finished, except when you come back after the declared TTL: there the previous turn is eligible exactly as at a cold return, so a provider whose cache outlives its TTL does not keep it at every return), when the edit pays for the rewrite it causes, and never on the request right after an edited one (no back-to-back warm rewrites). That is one inequality, r Δ²/(2g) + η Δ ≥ K with K = (w − r)(P T − (1 − P) Δ): the reads the removed Δ tokens would cost while the context grows back at g tokens per request (measured in the session), plus, near Pi's compaction trigger, what Pi would charge for the same room (η), against the rewrite of the T = A tokens left after the edit (pricing only the suffix after the earliest edit fires warm edits earlier and lost quality in the offline evaluation; the suffix is logged as `Tsuf` for measurement). r and w are the read and rewrite price ratios of the cache class, never the model's price table: explicit write premium 0.1 / 1.25 x input (2 x on the 1-hour tier), automatic prefix cache 0.2 / 1 x input; the class is read from the provider's usage reports, and until the first response the old fixed rule applies. P is the probability that the cache is still warm. A cold return is P = 0, so K < 0 and it always fires; a single small fold never pays at a warm cache, a large one does.
110
+ 2. **Edit only when the cache is already gone.** Provider prompt caches expire (the model's declared TTL, usually minutes). Changing the context while the cache is warm means paying to rewrite it; changing it after it expired is free, because the whole context is rewritten anyway, and a smaller context makes that rewrite cheaper. So when you come back after the TTL (or after switching model, or when the session was last touched longer ago than the TTL), pi-zip folds old outputs down to about 40K real tokens of context in one step, system prompt and tool definitions included, and every request of that turn sends the same bytes. No edit can shrink the system prompt and the tools, so once they take more than half of the 40K (a large tool set) the target becomes them plus 20K of conversation: with a 50K tool set the context lands near 70K. It is never a 40K that no edit can reach (that folds everything and summarises at every return), and it does not keep a full 40K of conversation on top either (on a live bench with a 36K prefix that cost 13% more, with no quality difference measured). While the cache is warm it does nothing, with one exception, the warm valve: above that target it applies that plan, minus the previous user turn (a warm edit never folds the turn you just finished, except when you come back after the declared TTL: there the previous turn is eligible exactly as at a cold return, so a provider whose cache outlives its TTL does not keep it at every return; and it leaves alone what the last cold return had to protect, the previous turn's outputs of that return, until the next cold return, until the session has grown by more than about 11K tokens since, or until the context reaches Pi's compaction room: the window sliding by one prompt is not news, and folding them then would pay a warm rewrite for what the free rewrite at the cold return could not take), when the edit pays for the rewrite it causes, and never on the request right after an edited one (no back-to-back warm rewrites). That is one inequality, r Δ²/(2g) + η Δ ≥ K with K = (w − r)(P T − (1 − P) Δ): the reads the removed Δ tokens would cost while the context grows back at g tokens per request (measured in the session), plus, near Pi's compaction trigger, what Pi would charge for the same room (η), against the rewrite of the T = A tokens left after the edit (pricing only the suffix after the earliest edit fires warm edits earlier and lost quality in the offline evaluation; the suffix is logged as `Tsuf` for measurement). r and w are the read and rewrite price ratios of the cache class, never the model's price table: explicit write premium 0.1 / 1.25 x input (2 x on the 1-hour tier), automatic prefix cache 0.2 / 1 x input; the class is read from the provider's usage reports, and until the first response the old fixed rule applies. P is the probability that the cache is still warm. A cold return is P = 0, so K < 0 and it always fires; a single small fold never pays at a warm cache, a large one does.
99
111
  3. **Never in the way.** Planning is local and takes milliseconds. Anything that needs a model call (a summary, only when folding is not enough and the same inequality prices the extra model call in, at a cold return and in the warm valve alike; the call is priced as what it is: an uncached read of the conversation since the previous summary, which code carries forward and the model never rewrites, plus the narrative it writes) is prepared while you are away: if the cache is about to expire (0.8 x its lifetime after your last request) and you have not come back, a background timer writes the summary with a separate, uncached call. The timer is cancelled the moment you send a prompt. If you return before the summary finishes, only the remaining time is waited (Esc cancels the prompt; the summary goes on and is used when you send it again), and the notice says so. A summary that finished while you were away goes into the session through Pi's own compaction before your prompt (Pi's `[compaction]` block then shows pi-zip's size before), in milliseconds; when Pi cannot take it there (busy, nothing it would compact, an older Pi, another extension taking part in Pi's compaction: see Coexistence), it rides along with the first request as before. If you return while the cache is still warm and the valve does not fire, the prepared summary is discarded (its cost is still counted). With Pi's own idle cache warming on (`"cacheWarming": "idle"` in the global settings, the only place Pi reads it), the summary waits for Pi's decision instead of the 0.8 mark: while Pi keeps refreshing the cache, nothing is prepared (you would come back to a warm cache and not need it); once Pi stops, or no refresh follows its decision, or the next decision (due one warming delay after the refresh finished) does not come, the summary is prepared as usual. pi-zip only watches those decisions, it never answers them. In non-interactive modes (`-p`, `--mode json`) nothing is ever started in the background: a cold return that needs a summary computes it right then.
100
112
 
101
- **The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip` shows the lifetime it currently believes, marked `(measured)` once it is learned.
113
+ **The cache lifetime is learned, not configured.** Every response says how much of the prompt came from the cache. pi-zip compares that read with what the request re-sent unchanged (the previous prompt, or the untouched prefix before one of its own edits: on an automatic prefix cache every edit leaves the first 8K tokens alone, so even the response right after a fold says whether the cache survived) and so learns, per provider and model, whether the cache survived a gap of that length: a few counts per gap bin (30 s to 90 min, with bin edges on the 5-minute and 1-hour tiers), monotone in the gap, older evidence halved after 16 newer observations of the same bin, stored without any content in `~/.pi/agent/pi-zip/cache-survival.json`. Before any evidence the model's declared TTL decides, exactly as before (300 s when it declares none; too short a guess is cheaper than too long); beyond it one clean read overrides it, inside it a lone miss counts as noise (warm caches do miss now and then) and only repeated misses do. A GLM cache read in full after 365 s makes the next 365 s return warm; a Claude 5-minute cache that read nothing after 360 s stays dead. Whether the provider bills cache writes (explicit cache) or not (automatic prefix cache) is read from the first response too. `/zip` draws what it has learned as the dot chart above.
102
114
 
103
115
  Protected from folding: the current user turn and the previous one. When the context is above the target, re-readable outputs of the previous turn (an unchanged file, a read-only command) can still be folded at a cold return or at any return after the declared TTL (never on a warm request inside it), and any output in either turn can be folded once it is 60 assistant requests old (so a long agent run that is a single user turn with hundreds of tool calls is not exempt from folding; the newest 59 requests' outputs always stay, and every fold stays recallable). Messages you type while the agent is running (steering, follow-up) belong to that turn and do not start a new one. Outputs you have already recalled, and `zip_recall` results themselves (or reads of a recall file), are never folded again. "Read-only" is a conservative whitelist: `find -delete` or `-exec`, command substitution, redirects, background jobs, `git diff --output` and the like are not.
104
116
 
@@ -110,7 +122,7 @@ The cold cap is kept below Pi's own compaction trigger (window minus `compaction
110
122
 
111
123
  | Command | Effect |
112
124
  |---|---|
113
- | `/zip` | the card above: state, how long the cache survives you being away (chart and sentence), this session's folds, summaries and recalls (counted from the session, so a restart does not reset them) |
125
+ | `/zip` | the card above: state, how long the cache survives you being away (a dot chart of your own returns), this session's folds, summaries and recalls (counted from the session, so a restart does not reset them) |
114
126
  | `/zip off` | strict no-op: no folds, no summaries, requests left untouched (earlier folds stay recallable) |
115
127
  | `/zip on` | resume |
116
128
  | `/zip quiet` | toggle the per-turn notice (folding continues) |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-zip",
3
- "version": "0.3.0-rc.1",
3
+ "version": "0.3.0-rc.2",
4
4
  "description": "Keeps long Pi sessions cheap without losing anything: folds old tool output only when the prompt cache has already gone cold, and every fold can be recalled byte for byte. Zero config.",
5
5
  "type": "module",
6
6
  "license": "MIT",
package/src/plan.ts CHANGED
@@ -38,6 +38,16 @@ export const LOOKAHEAD = false;
38
38
  export const NARRATIVE_PRICE = true;
39
39
  /** The narrative's token budget (chars/4), summary.ts buildCut: clamp(planned - skeleton, 300, NARRATIVE_MAX). */
40
40
  export const NARRATIVE_MAX = 4000;
41
+ /** B2 for folds (formal TRIAGE (a) 6): the outputs a cold plan must protect (the previous user turn's, the mutating ones) are no longer
42
+ * protected one prompt later, when the window slides, and the warm plan then folded them at a warm cache: a rewrite the free cold one
43
+ * dominated. Now a warm plan holds back the blocks that were protected at the last cold return (PlanOpts.hold: that return's prompt)
44
+ * until the next cold return, or until the session has grown by more than HOLD_RELEASE real tokens since it (new information), or the
45
+ * context is at Pi's compaction room; and when even folding them all would not avoid a summary, it is the plan it was before. Chosen
46
+ * over folding them early at the cold return (the user would lose the previous turn's outputs at the prompt that asks about them):
47
+ * formal/TRIAGE.md, B2 for folds. */
48
+ export const FOLD_HOLD = true;
49
+ /** The spec's "no new information" is growth <= 10K real tokens (formal DELTA); the estimate is chars/4 x c, 15% off at worst (TOL_TARGET). */
50
+ export const HOLD_RELEASE = 11_500;
41
51
  const SUMMARY_FLOOR = 1000;
42
52
  const SUMMARY_CAP = 8000;
43
53
  const SUMMARY_RATIO = 0.1;
@@ -550,6 +560,8 @@ export interface PlanOpts {
550
560
  narrativePrice?: boolean; // F3, default NARRATIVE_PRICE
551
561
  slide?: number; // internal (F2): plan as if this many more user turns had started (the protected window slid)
552
562
  valveB?: number; // internal (F2): the real context the slid warm plan's law sees (the cold plan's folds already applied)
563
+ hold?: { anchor: string }; // B2 for folds: the entry id of the prompt of the last cold return (run.ts holdOf); undefined = nothing held
564
+ holdPass?: boolean; // internal: the held pass of a warm plan with a hold
553
565
  warmStyle?: boolean; // internal: a cold plan at 0 < P(warm) < 0.5 that the law refused is planned once more the way a warm plan is (no relax folds unless pastTtl, the summary gated)
554
566
  }
555
567
 
@@ -603,6 +615,14 @@ export const contentKeyOf = (content: Any): string => {
603
615
  * passes the same law (expected cost). The cap never exceeds Pi's compaction room. Returns null when there is nothing to plan on
604
616
  * or the law says no. */
605
617
  export function planContext(entries: Any[], o: PlanOpts): PlanResult | null {
618
+ if (o.hold && !o.holdPass && (o.mode ?? "cold") === "warm") {
619
+ // the hold gives way to a summary: when the plan without it summarises, that is the plan (a summary voids the folds before its
620
+ // cut anyway); otherwise the held plan, with no summary of its own (the next cold return writes it)
621
+ const full = planContext(entries, { ...o, hold: undefined, trace: undefined });
622
+ if (!full) return null;
623
+ if (full.cutIdx !== null) return planContext(entries, { ...o, hold: undefined });
624
+ return planContext(entries, { ...o, noSummary: true, holdPass: true });
625
+ }
606
626
  const s = settings();
607
627
  const mode = o.mode ?? "cold";
608
628
  const warmLike = mode === "warm" || o.warmStyle === true; // plans the way a warm plan does (the label of the trigger stays the mode's)
@@ -679,7 +699,22 @@ export function planContext(entries: Any[], o: PlanOpts): PlanResult | null {
679
699
  const foldable = (b: Block) => b.kind === "toolResult" && !b.edited && !!b.entryId && b.tokens > foldMin && !isRecall(b) && start[b.idx] >= head && (!o.rereadOnly || classify(b) === "rereadable");
680
700
  const protectedTurn = (b: Block) => b.userTurn >= userTurns - PROTECT_USER_TURNS + 1;
681
701
  const savings = () => folds.reduce((a, t) => a + t.entryTokens - t.phTokens, 0);
682
- const cands = blocks.filter((b) => foldable(b) && !protectedTurn(b));
702
+ // the blocks protected at the last cold return stay out of a warm plan's folds (see FOLD_HOLD)
703
+ const held: (b: Block) => boolean = (() => {
704
+ const NOT_HELD = () => false;
705
+ if (mode !== "warm" || !o.hold) return NOT_HELD;
706
+ const from = blocks.findIndex((b) => b.entryId === o.hold!.anchor); // gone (a later summary replaced it): nothing is held
707
+ if (from < 0) return NOT_HELD;
708
+ // what the session grew by before this decision: the blocks between the cold return's prompt and the newest response (that
709
+ // response, the tool results after it and a new prompt are what the decision is about, not yet something the cache has seen)
710
+ let lastAsst = blocks.length - 1;
711
+ while (lastAsst > 0 && blocks[lastAsst].kind !== "assistant") lastAsst--;
712
+ const grown = c * blocks.slice(from + 1, Math.max(from + 1, lastAsst)).reduce((a, b) => a + b.tokens, 0);
713
+ if (grown > HOLD_RELEASE || (room !== null && real(ctxEst) >= room)) return NOT_HELD;
714
+ const turn = blocks[from].userTurn;
715
+ return (b: Block) => b.userTurn >= turn - PROTECT_USER_TURNS + 1;
716
+ })();
717
+ const cands = blocks.filter((b) => foldable(b) && !protectedTurn(b) && !held(b));
683
718
  for (const b of cands) addFold(b, trig);
684
719
  if (relax && ((mode === "cold" && !o.warmStyle) || o.pastTtl === true)) {
685
720
  // the protected window = the new prompt + the previous user turn; that turn's big reads are what makes a cold return
package/src/run.ts CHANGED
@@ -4,17 +4,17 @@ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
4
4
  import { appendFileSync, existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs";
5
5
  import { dirname, join } from "node:path";
6
6
  import { detectCold, lastMessageMs, lastPrompt, modelKey, payloadRetention, refused, resolveTtl, ttlFor } from "./cache.ts";
7
- import { agentDir, describe, lawPrices, loadStats, pWarm, record, sample, statsPath, GAP_EDGES, type Entry } from "./learn.ts";
7
+ import { agentDir, curve, describe, lawPrices, loadStats, pWarm, record, sample, statsPath, GAP_EDGES, type Entry } from "./learn.ts";
8
8
  import { validateEdits, repairPayload } from "./guard.ts";
9
9
  import { fmtK, Stats, noticeDesc, noticeText, type NoticeAction, type NoticeData, type ZipControl } from "./notice.ts";
10
- import { applyPlanToMessages, buildBlocks, calibrate, countUserTurns, G0, keepRecentTokensFor, piCanCompact, planContext, rebaseFolds, reserveTokensFor, sizeOf, spliceNative, untouchedEst, UNIT, type Block, type Cut, type FoldTarget, type Law, type PlanOpts, type PlanResult, type RunPlan, type Scale, type SysInfo } from "./plan.ts";
10
+ import { applyPlanToMessages, buildBlocks, calibrate, countUserTurns, FOLD_HOLD, G0, keepRecentTokensFor, piCanCompact, planContext, rebaseFolds, reserveTokensFor, sizeOf, spliceNative, untouchedEst, UNIT, type Block, type Cut, type FoldTarget, type Law, type PlanOpts, type PlanResult, type RunPlan, type Scale, type SysInfo } from "./plan.ts";
11
11
  import { handleFor, PH_MARK, RECALL_TOOL } from "./placeholder.ts";
12
12
  import { branchOutputs, recalledHandlesFromBranch, resolveHandlesInBranch } from "./recall.ts";
13
13
  import { cleanupRecall, commandNamesRecall, fileRecallFor, materialize, parseRecallPath, recallRoot, recallRootShown, rewriteCommand, sidDir, type FileRecall } from "./recallfile.ts";
14
14
  import { buildCut } from "./summary.ts";
15
- import { type Any, COMMIT_CUSTOM, envInt, PRODUCT, textOf, tok4 } from "./util.ts";
15
+ import { type Any, COLD_CUSTOM, COMMIT_CUSTOM, envInt, PRODUCT, textOf, tok4 } from "./util.ts";
16
16
  import { replayGrowth } from "./state.ts";
17
- import { CHART_COLS, chartGap, fmtLife, type CardData, type StateWord } from "./ui.ts";
17
+ import { fmtLife, type CardData, type ChartBin, type StateWord } from "./ui.ts";
18
18
 
19
19
  export const PLAN_CUSTOM = "pi-zip/plan";
20
20
  export const STATE_CUSTOM = "pi-zip/state";
@@ -387,7 +387,17 @@ export class Zip implements ZipControl {
387
387
  const cal = this.scaleOf(ctx, entries);
388
388
  const fileRecall = this.fileRecall(ctx);
389
389
  const rereadOnly = this.recallMode === "reread" || (this.recallMode === "file" && !fileRecall);
390
- return { cal, o: { mode, scale: cal, cwd: (ctx?.cwd as string) ?? process.cwd(), recalled: this.recalled, model: ctx?.model, promptPending: pending, rereadOnly, ...(fileRecall ? { fileRecall } : {}), noSummary: mode === "warm" && this.summarizedWarm, reserve: this.reserve(ctx), steerIds: this.steerIds, law: this.law(ctx, pWarm), trace: [] } };
390
+ return { cal, o: { mode, scale: cal, cwd: (ctx?.cwd as string) ?? process.cwd(), recalled: this.recalled, model: ctx?.model, promptPending: pending, rereadOnly, ...(fileRecall ? { fileRecall } : {}), noSummary: mode === "warm" && this.summarizedWarm, ...(FOLD_HOLD && mode === "warm" ? this.holdOf(ctx) : {}), reserve: this.reserve(ctx), steerIds: this.steerIds, law: this.law(ctx, pWarm), trace: [] } };
391
+ }
392
+
393
+ /** The prompt of the session's last cold return: the first user message after the newest COLD_CUSTOM entry (none: nothing is held). */
394
+ private holdOf(ctx: Any): { hold?: { anchor: string } } {
395
+ const branch = this.branch(ctx);
396
+ let mark = -1;
397
+ branch.forEach((e: Any, i: number) => { if (e?.type === "custom" && e.customType === COLD_CUSTOM) mark = i; });
398
+ if (mark < 0) return {};
399
+ const at = branch.findIndex((e: Any, i: number) => i > mark && e?.type === "message" && e.message?.role === "user" && !this.steerIds.has(e.id));
400
+ return at < 0 ? {} : { hold: { anchor: branch[at].id } };
391
401
  }
392
402
 
393
403
  private law(ctx: Any, pWarm: number): Law {
@@ -630,13 +640,18 @@ export class Zip implements ZipControl {
630
640
  await this.resolveRetention(ctx);
631
641
  const { cold, reason, ttl, pWarm: p, src, gapS, pastTtl } = detectCold(ctx.model, this.lastReqMs, branch, Date.now(), this.lastModelKey, (g, prior) => pWarm(this.ent, g, prior), this.tier(ctx.model));
632
642
  this.cold = cold;
633
- if (cold) this.summarizedWarm = false; // a cold return may summarise again
643
+ if (cold) {
644
+ this.summarizedWarm = false; // a cold return may summarise again
645
+ try {
646
+ this.pi.appendEntry(COLD_CUSTOM, {}); // right before the prompt: the session says where its last cold return was (plan.ts FOLD_HOLD)
647
+ } catch {}
648
+ }
634
649
  this.pastTtl = pastTtl;
635
650
  this.pWarm = p;
636
651
  this.survSrc = src;
637
652
  this.runChecked = false; // the first request decides (cold: the cold plan; warm: only above the cap, if the law fires)
638
653
  this.coldReason = reason;
639
- this.coldWhy = src === "model switch" ? "model switched" : cold && gapS !== null ? `away ${fmtAway(gapS)}` : "";
654
+ this.coldWhy = src === "model switch" ? "cache cold · model switched" : cold && gapS !== null ? `cache cold · away ${fmtAway(gapS)}` : ""; // the warm form says "cache warm · worth it at this size"
640
655
  for (const h of recalledHandlesFromBranch(branch)) this.recalled.add(h);
641
656
  for (let i = branch.length - 1; i >= 0; i--) {
642
657
  const en = branch[i];
@@ -1222,7 +1237,7 @@ export class Zip implements ZipControl {
1222
1237
 
1223
1238
  /**
1224
1239
  * The transcript notice for one plan. Facts only, no claims: what changed (sizes), how long it took, and, expanded, why now
1225
- * ("away 47 min") and what was folded. Whether the cache really was cold is learned from the response, silently (/zip shows it).
1240
+ * ("cache cold · away 47 min") and what was folded. Whether the cache really was cold is learned from the response, silently (/zip shows it).
1226
1241
  * `realAfter` > 0: the size the provider reported for the request that carried the plan (shown instead of the estimate).
1227
1242
  */
1228
1243
  private noticeFor(ctx: Any, plan: RunPlan, live: FoldTarget[], cut: Cut | null, ctxTokens: number, realAfter: number) {
@@ -1433,9 +1448,18 @@ export class Zip implements ZipControl {
1433
1448
  const key = modelKey(model), ttlS = ttlFor(model, [], this.tier(model)) / 1000, ent = key ? loadStats().models[key] : undefined;
1434
1449
  let lifeS = 0;
1435
1450
  for (let g = 30; g <= 7200; g += 30) if (pWarm(ent, g, ttlS).p >= 0.5) lifeS = g;
1436
- const measured = !!ent && ent.n > 0 && pWarm(ent, Math.max(30, lifeS), ttlS).src === "learned";
1437
1451
  const life = key ? (lifeS >= 7200 ? "2 h+" : fmtLife(lifeS || ttlS)) : undefined;
1438
- const alive = key ? Array.from({ length: CHART_COLS }, (_, i) => pWarm(ent, chartGap(i), ttlS).p) : undefined;
1452
+ // the chart: every learned bin that still rounds to a return (the weights are decayed), and where the learned curve says
1453
+ // "still there" ends (s1) and "gone" starts (s2); never the declared TTL
1454
+ const bins: ChartBin[] = Object.entries(ent?.bins ?? {})
1455
+ .map(([b, [a, d]]) => ({ lo: GAP_EDGES[Number(b)], hi: GAP_EDGES[Number(b) + 1] ?? 7200, alive: Math.round(a), gone: Math.round(d) }))
1456
+ .filter((x) => x.lo !== undefined && x.alive + x.gone > 0)
1457
+ .sort((x, y) => x.lo - y.lo);
1458
+ const cur = (ent ? curve(ent, ttlS) : []).map((x) => ({ lo: GAP_EDGES[x.bin], hi: GAP_EDGES[x.bin + 1] ?? 7200, p: x.p }));
1459
+ let s1 = GAP_EDGES[0];
1460
+ for (const x of cur) { if (x.p >= 0.8) s1 = x.hi; else break; }
1461
+ const s2 = cur.find((x) => x.lo >= s1 && x.p <= 0.2)?.lo ?? 7200;
1462
+ const usually = cur.some((x) => x.hi <= s1 && x.p < 0.95);
1439
1463
  const ts = t.last?.timestamp ? new Date(t.last.timestamp) : null;
1440
1464
  const hhmm = ts && !Number.isNaN(ts.getTime()) ? `${String(ts.getHours()).padStart(2, "0")}:${String(ts.getMinutes()).padStart(2, "0")} ` : "";
1441
1465
  const last = t.last ? `${hhmm}${t.last.data.desc}${t.last.data.why ? " · " + t.last.data.why : ""}` : undefined;
@@ -1444,7 +1468,7 @@ export class Zip implements ZipControl {
1444
1468
  const rereadWhy = this.recallMode === "file" ? "a tool allowlist hides zip_recall · no session file, so no recall copies" : "a tool allowlist hides zip_recall · only re-readable outputs fold";
1445
1469
  return {
1446
1470
  state, stateNote: state === "paused" ? `${this.conflict} also manages context` : state === "off" ? "/zip on to resume" : state === "reread-only" ? rereadWhy : undefined,
1447
- model: key || undefined, life, measured, alive, folds: t.folds, foldedTokens: fmtK(t.folded), summaries: t.summaries, summaryUsd: t.usd, recalls: t.recalls, last, quiet: this.quiet,
1471
+ model: key || undefined, life, ...(bins.length ? { bins, s1, s2, usually } : {}), folds: t.folds, foldedTokens: fmtK(t.folded), summaries: t.summaries, summaryUsd: t.usd, recalls: t.recalls, last, quiet: this.quiet,
1448
1472
  ...(file ? { unzip: `from files in ${recallRootShown()} (a tool allowlist hides zip_recall)` } : {}),
1449
1473
  };
1450
1474
  }
package/src/ui.ts CHANGED
@@ -49,13 +49,19 @@ export function renderState(word: StateWord, reason: string, width: number, th:
49
49
 
50
50
  // ---- status card ------------------------------------------------------------------------------------------------------
51
51
 
52
+ /** One learned gap bin of the cache chart: how many times you came back after a gap in [lo, hi) s and the cache was still there / gone.
53
+ * Counts are rounded decayed weights (learn.ts keeps no raw per-bin count: old evidence halves after HALF_LIFE newer ones of the bin). */
54
+ export interface ChartBin { lo: number; hi: number; alive: number; gone: number }
55
+
52
56
  export interface CardData {
53
57
  state: "on" | "off" | "paused" | "reread-only";
54
58
  stateNote?: string; // why paused / reread-only
55
59
  model?: string; // provider/id
56
- life?: string; // "~5 min" | "~1 h" | "2 h+": how long the cache outlives the last reply
57
- measured?: boolean; // life comes from this provider's own replies, not the declared TTL
58
- alive?: number[]; // P(cache still there) after being away chartGap(i), one per chart column
60
+ life?: string; // "~5 min" | "~1 h" | "2 h+": how long the cache outlives the last reply (the declared TTL until something is learned)
61
+ bins?: ChartBin[]; // learned bins that round to at least one return; none = nothing learned yet: one line, no chart
62
+ s1?: number; // seconds: the cache is (nearly) always still there under this (learned curve, never the prior)
63
+ s2?: number; // seconds: it is mostly gone from here on; between s1 and s2 it is not sure yet
64
+ usually?: boolean; // some bin under s1 has P(still there) < 0.95: "usually still there" instead of "still there"
59
65
  folds: number;
60
66
  foldedTokens: string; // "310K"
61
67
  summaries: number;
@@ -73,34 +79,106 @@ export function fmtLife(s: number): string {
73
79
  return `~${+(s / 3600).toFixed(1)} h`;
74
80
  }
75
81
 
76
- /** The cache chart: one column per time away, log-spaced from 30 s to 2 h. The row labels are the axes ("cache" up, "away" across),
77
- * so it reads like any chart; the bars still there are in the accent colour, the dead ones dim. */
78
- export const CHART_COLS = 32;
79
- const CHART_T0 = 30, CHART_T1 = 7200;
80
- export const chartGap = (i: number): number => CHART_T0 * (CHART_T1 / CHART_T0) ** (i / (CHART_COLS - 1));
81
- const chartCol = (s: number): number => Math.round(((CHART_COLS - 1) * Math.log(s / CHART_T0)) / Math.log(CHART_T1 / CHART_T0));
82
- const BARS = "▁▂▃▄▅▆▇█";
83
- const TICKS: [number, string][] = [[60, "1m"], [300, "5m"], [900, "15m"], [3600, "1h"], [7200, "2h"]];
84
- const bar = (p: number): string => BARS[Math.max(0, Math.min(7, Math.round(p * 7)))];
85
- const ALIVE = 0.5; // a bar at or above this is drawn as still there
86
-
87
- /** "1m 5m 15m 1h 2h": tick labels start at their column. */
88
- export function chartAxis(): string {
89
- const s = Array<string>(CHART_COLS + 2).fill(" ");
90
- for (const [t, lab] of TICKS) {
91
- const c = Math.min(chartCol(t), CHART_COLS - lab.length + 1);
92
- if (s.slice(Math.max(0, c - 1), c + lab.length + 1).every((x) => x === " ")) s.splice(c, lab.length, ...lab);
82
+ // ---- the cache chart: one dot per time you came back ----------------------------------------------------------------------
83
+ // Across: how long you were away, 30 s to 2 h on a log scale. Above the axis (accent): the times the cache was still there, below
84
+ // (dim): the times it was gone. The axis itself says what to expect: solid while it is (nearly) always there, dashed while not sure,
85
+ // thin after it is gone. A column shows at most STACK_CAP marks; beyond that STACK_CAP - 1 dots and then the count.
86
+
87
+ export const CHART_W = 60; // columns 0..CHART_W
88
+ const CT0 = 30, CT1 = 7200;
89
+ export const STACK_CAP = 4;
90
+ const IND = 11; // the values column of the card: " " + the 9-wide label
91
+ export const colW = (s: number): number => Math.round((CHART_W * Math.log(s / CT0)) / Math.log(CT1 / CT0));
92
+ const gapW = (c: number): number => CT0 * (CT1 / CT0) ** (c / CHART_W);
93
+ const AXIS_TICKS: [number, string][] = [[30, "30s"], [60, "1m"], [120, "2m"], [300, "5m"], [600, "10m"], [1800, "30m"], [3600, "1h"], [7200, "2h"]];
94
+ const fmtGap = (s: number): string => (s < 60 ? `${Math.round(s)} s` : s < 3600 ? `${+(s / 60).toFixed(1)} min` : `${+(s / 3600).toFixed(1)} h`);
95
+
96
+ /** What the marks of one side of the axis look like, as rows from the axis outwards: row r (1-based) holds, per column, a dot, a digit of
97
+ * the count or a blank. `items` are (column, times); items on the same column merge. A count is drawn on the row of the stack's end,
98
+ * centred on its column; when that would touch another mark in its row (neighbouring columns, other counts) it moves, the nearest free
99
+ * place first, right before left, and keeps one blank on each side, so two counts never read as one number. Deterministic. */
100
+ export function stackRows(items: { col: number; n: number }[], dot: string = "\u25cf", cap: number = STACK_CAP, width: number = CHART_W + 1): { rows: string[][]; height: number } {
101
+ const merged = new Map<number, number>();
102
+ for (const it of items) {
103
+ const col = Math.max(0, Math.min(width - 1, it.col));
104
+ if (it.n > 0) merged.set(col, (merged.get(col) ?? 0) + it.n);
105
+ }
106
+ const cols = [...merged.entries()].sort((a, b) => a[0] - b[0]);
107
+ const height = Math.max(1, ...cols.map(([, n]) => Math.min(n, cap)));
108
+ const rows = Array.from({ length: height }, () => Array<string>(width).fill(" "));
109
+ const counts: { col: number; text: string; row: number }[] = [];
110
+ for (const [col, n] of cols) {
111
+ const dots = n <= cap ? n : cap - 1;
112
+ for (let r = 0; r < dots; r++) rows[r][col] = dot;
113
+ if (n > cap) counts.push({ col, text: String(n), row: cap - 1 });
114
+ }
115
+ for (const { col, text, row } of counts) {
116
+ const len = text.length, want = col - Math.floor((len - 1) / 2), line = rows[row];
117
+ const free = (st: number, gap: boolean) => {
118
+ if (st < 0 || st + len > width) return false;
119
+ for (let x = st - (gap ? 1 : 0); x < st + len + (gap ? 1 : 0); x++) if (x >= 0 && x < width && line[x] !== " ") return false;
120
+ return true;
121
+ };
122
+ let at = -1;
123
+ for (const gap of [true, false]) {
124
+ for (let d = 0; d < width && at < 0; d++) for (const st of d ? [want + d, want - d] : [want]) if (at < 0 && free(st, gap)) at = st;
125
+ if (at >= 0) break;
126
+ }
127
+ if (at >= 0) for (let k = 0; k < len; k++) line[at + k] = text[k];
93
128
  }
94
- return s.join("").trimEnd();
129
+ return { rows, height };
95
130
  }
96
131
 
97
- const lifeText = (c: CardData): string => (c.life ? `expires after ${c.life} idle${c.measured ? " (measured)" : ""}` : "");
132
+ /** Colours the cells of a row (one string per column, null = blank) with as few runs as possible. */
133
+ function paint(cells: { ch: string; color: string }[], th: Th): string {
134
+ let out = "", run = "", color = "";
135
+ const flush = () => { if (run) out += color ? th.fg(color, run) : run; run = ""; };
136
+ for (const c of cells) {
137
+ if (c.color !== color) flush(), (color = c.color);
138
+ run += c.ch;
139
+ }
140
+ flush();
141
+ return out.trimEnd();
142
+ }
143
+
144
+ interface Seg { lead: string; word: string; color: string }
145
+ /** "under 5 min: usually still there · 5–10 min: not sure yet · over 10 min: gone", as segments that wrap whole. */
146
+ function captionSegs(c: CardData): Seg[] {
147
+ const s1 = c.s1 ?? CT0, s2 = c.s2 ?? CT1;
148
+ const range = s1 >= 60 && s2 < 3600 ? `${+(s1 / 60).toFixed(1)}–${fmtGap(s2)}` : `${fmtGap(s1)}–${fmtGap(s2)}`; // "5–10 min"
149
+ return [
150
+ { lead: `under ${fmtGap(s1)}: `, word: c.usually ? "usually still there" : "still there", color: "accent" },
151
+ ...(s2 > s1 ? [{ lead: `${range}: `, word: "not sure yet", color: "warning" }] : []),
152
+ { lead: `over ${fmtGap(s2)}: `, word: "gone", color: "muted" },
153
+ ];
154
+ }
155
+ const SEP = " \u00b7 ";
156
+ const segW = (g: Seg) => g.lead.length + g.word.length;
157
+ /** The caption as lines of at most `room` columns, wrapped between segments (a segment wider than the room is cut). */
158
+ function captionLines(segs: Seg[], room: number, th: Th, m: Measure): string[] {
159
+ if (room < 2) return []; // no room for even a cut sentence
160
+ const lines: Seg[][] = [[]];
161
+ for (const g of segs) {
162
+ const cur = lines[lines.length - 1], w = cur.reduce((a, x) => a + segW(x), 0) + SEP.length * cur.length + segW(g);
163
+ if (cur.length && w > room) lines.push([g]);
164
+ else cur.push(g);
165
+ }
166
+ return lines.map((l) => {
167
+ const plain = l.map((g) => g.lead + g.word).join(SEP);
168
+ if (plain.length > room) return th.fg("dim", m.cut(plain, room));
169
+ return l.map((g) => th.fg("dim", g.lead) + th.fg(g.color, g.word)).join(th.fg("dim", SEP));
170
+ });
171
+ }
172
+ const captionPlain = (c: CardData) => captionSegs(c).map((g) => g.lead + g.word).join(SEP);
173
+
174
+ const idle = (c: CardData): string => (c.life ? `expires after ${c.life} idle` : "");
175
+ const NO_DATA = "a chart appears once you come back after a break";
176
+ const hasChart = (c: CardData): boolean => !!c.model && !!c.bins?.length;
98
177
 
99
178
  export function cardText(c: CardData): string[] {
100
179
  const rows: [string, string][] = [];
101
- // One fact the user can act on ("come back within ~5 min and nothing is touched"), drawn and said: no classes, no jargon.
102
- if (c.model && c.alive?.length) rows.push(["model", c.model], ["cache", `${c.alive.map(bar).join("")} ${lifeText(c)}`.trimEnd()], ["away", chartAxis()]);
103
- else if (c.model) rows.push(["cache", `${c.model}${c.life ? ` · ${lifeText(c)}` : ""}`]);
180
+ if (c.model) rows.push(["model", c.model]);
181
+ if (c.model) rows.push(["cache", hasChart(c) ? `each dot is one time you came back · ${captionPlain(c)}` : `${idle(c) || "expires after ~5 min idle"} · ${NO_DATA}`]);
104
182
  const n = (x: number, one: string, many = one + "s") => `${x} ${x === 1 ? one : many}`;
105
183
  rows.push(["session", `${c.folds} zipped ~${c.foldedTokens} · ${n(c.summaries, "summary", "summaries")}${c.summaryUsd > 0 ? ` $${c.summaryUsd.toFixed(2)}` : ""} · ${n(c.recalls, "unzip")}`]);
106
184
  if (c.last) rows.push(["last", c.last]);
@@ -109,47 +187,74 @@ export function cardText(c: CardData): string[] {
109
187
  return rows.map(([k, v]) => `${k.padEnd(LABEL_W)}${v}`);
110
188
  }
111
189
 
190
+ /** The chart rows (without the model row): header, marks above, axis, marks below, axis labels, [legend], caption. */
191
+ function chartRows(c: CardData, width: number, th: Th, m: Measure): string[] {
192
+ const indent = " ", label = (k: string) => th.fg("dim", (indent + k).padEnd(indent.length + LABEL_W)), pad = " ".repeat(IND);
193
+ const bins = c.bins ?? [];
194
+ const na = bins.reduce((a, b) => a + b.alive, 0), nd = bins.reduce((a, b) => a + b.gone, 0);
195
+ const up = stackRows(bins.map((b) => ({ col: colW(Math.sqrt(b.lo * b.hi)), n: b.alive })));
196
+ const down = stackRows(bins.map((b) => ({ col: colW(Math.sqrt(b.lo * b.hi)), n: b.gone })), "\u25cb");
197
+ const legUp = ["\u25cf", ` still there ${na}`], legDown = ["\u25cb", ` gone ${nd}`];
198
+ const legW = (l: string[]) => l[0].length + l[1].length;
199
+ const beside = width >= IND + CHART_W + 1 + 3 + Math.max(legW(legUp), legW(legDown));
200
+ const s1 = c.s1 ?? CT0, s2 = c.s2 ?? CT1;
201
+ const out: string[] = [label("cache") + th.fg("dim", "each dot is one time you came back")];
202
+ const side = (rows: string[][], color: string, leg: string[], first: number) =>
203
+ rows.map((cells, r) => {
204
+ const line = paint(cells.map((ch) => ({ ch, color: ch === " " ? "" : color })), th);
205
+ return pad + line + (beside && r === first ? " " + th.fg(color, leg[0]) + th.fg("muted", leg[1]) : "");
206
+ });
207
+ out.push(...side(up.rows.slice().reverse(), "accent", legUp, up.height - 1)); // outermost row first; the legend sits on the row next to the axis
208
+ out.push(pad + paint(Array.from({ length: CHART_W + 1 }, (_, i) => ({ ch: gapW(i) < s1 ? "\u2501" : gapW(i) < s2 ? "\u2505" : "\u2500", color: gapW(i) < s1 ? "accent" : gapW(i) < s2 ? "warning" : "dim" })), th));
209
+ out.push(...side(down.rows, "dim", legDown, 0));
210
+ const ax = Array<string>(CHART_W + 3).fill(" ");
211
+ for (const [t, lab] of AXIS_TICKS) {
212
+ const col = Math.min(colW(t), CHART_W - lab.length + 1);
213
+ for (let k = 0; k < lab.length; k++) ax[col + k] = lab[k];
214
+ }
215
+ out.push(label("away") + th.fg("dim", ax.join("").trimEnd()));
216
+ if (!beside) out.push(pad + th.fg("accent", legUp[0]) + th.fg("muted", legUp[1]) + " " + th.fg("dim", legDown[0]) + th.fg("muted", legDown[1]));
217
+ for (const l of captionLines(captionSegs(c), width - IND, th, m)) out.push(pad + l);
218
+ return out;
219
+ }
220
+
112
221
  export function renderCard(c: CardData, width: number, th: Th, m: Measure = plainMeasure): string[] {
113
222
  const out = renderState(c.state, c.stateNote ?? "", width, th, m);
114
223
  const indent = " "; // under the words after the monogram
115
- // Too narrow for the chart and its axis: the sentence alone says the same thing.
116
- if (c.alive?.length && width < indent.length + LABEL_W + Math.max(c.alive.length, chartAxis().length)) c = { ...c, alive: undefined };
117
- let under = "";
118
- for (const row of cardText(c)) {
119
- if (under && !row.startsWith("away ")) out.push(under), (under = "");
120
- if (c.alive?.length && row.startsWith("cache ")) {
121
- const runs: string[] = [];
122
- for (let i = 0; i < c.alive.length; ) {
123
- const on = c.alive[i] >= ALIVE;
124
- let j = i;
125
- while (j < c.alive.length && c.alive[j] >= ALIVE === on) j++;
126
- runs.push(th.fg(on ? "accent" : "dim", c.alive.slice(i, j).map(bar).join("")));
127
- i = j;
128
- }
129
- // The sentence goes beside the bars when it fits there whole, else on its own line under the axis: never cut mid-fact.
130
- const say = lifeText(c), beside = m.vw(say) <= width - indent.length - LABEL_W - c.alive.length - 2;
131
- out.push(th.fg("dim", (indent + "cache").padEnd(indent.length + LABEL_W)) + runs.join("") + (say && beside ? " " + th.fg("dim", say) : ""));
132
- if (say && !beside) under = th.fg("dim", m.cut(" ".repeat(indent.length + LABEL_W) + say, width));
133
- continue;
224
+ const head = (k: string) => th.fg("dim", (indent + k).padEnd(indent.length + LABEL_W));
225
+ const row = (k: string, v: string) => {
226
+ const line = m.cut(indent + k.padEnd(LABEL_W) + v, width);
227
+ if (!line) return;
228
+ const kk = line.slice(0, indent.length + LABEL_W);
229
+ out.push(th.fg("dim", kk) + th.fg("dim", line.slice(kk.length)));
230
+ };
231
+ if (c.model) {
232
+ row("model", c.model);
233
+ if (!hasChart(c)) {
234
+ // nothing learned yet: one line, the chart appears with the first return after a break
235
+ const room = width - indent.length - LABEL_W;
236
+ const v = [`${idle(c) || "expires after ~5 min idle"} \u00b7 ${NO_DATA}`, idle(c) || "expires after ~5 min idle"].find((x) => m.vw(x) <= room);
237
+ if (v) out.push(head("cache") + th.fg("dim", v));
238
+ else row("cache", idle(c) || "expires after ~5 min idle");
239
+ } else if (width >= IND + CHART_W + 1) out.push(...chartRows(c, width, th, m));
240
+ else {
241
+ // too narrow for the chart: the caption sentence alone says the same thing
242
+ const lines = captionLines(captionSegs(c), width - IND, th, m);
243
+ lines.forEach((l, i) => out.push((i ? " ".repeat(IND) : head("cache")) + l));
134
244
  }
135
- if (!c.alive?.length && c.model && c.life && row.startsWith("cache ")) {
136
- // Narrow: keep the fact whole, dropping the model name first, then "(measured)".
137
- const head = (indent + "cache").padEnd(indent.length + LABEL_W), room = width - m.vw(head);
138
- const v = [row.slice(LABEL_W), lifeText(c), lifeText({ ...c, measured: false })].find((x) => m.vw(x) <= room);
139
- if (v) { out.push(th.fg("dim", head + v)); continue; }
140
- }
141
- if (row.startsWith("unzip ")) {
245
+ }
246
+ for (const r of cardText({ ...c, model: undefined })) {
247
+ if (r.startsWith("unzip ")) {
142
248
  // Narrow: the bracketed reason goes first, so where the files are stays whole.
143
- const head = (indent + "unzip").padEnd(indent.length + LABEL_W), room = width - m.vw(head);
144
- const v = [row.slice(LABEL_W), row.slice(LABEL_W).replace(/ \([^()]*\)$/, "")].find((x) => m.vw(x) <= room);
145
- if (v) { out.push(th.fg("dim", head + v)); continue; }
249
+ const h = head("unzip"), room = width - m.vw(h);
250
+ const v = [r.slice(LABEL_W), r.slice(LABEL_W).replace(/ \([^()]*\)$/, "")].find((x) => m.vw(x) <= room);
251
+ if (v) { out.push(th.fg("dim", h + v)); continue; }
146
252
  }
147
- const line = m.cut(indent + row, width);
253
+ const line = m.cut(indent + r, width);
148
254
  if (!line) continue;
149
255
  const k = line.slice(0, indent.length + LABEL_W), v = line.slice(indent.length + LABEL_W);
150
- out.push(th.fg("dim", k) + (row.startsWith("session") ? th.fg("text", v) : th.fg("dim", v)));
256
+ out.push(th.fg("dim", k) + (r.startsWith("session") ? th.fg("text", v) : th.fg("dim", v)));
151
257
  }
152
- if (under) out.push(under);
153
258
  return out;
154
259
  }
155
260
 
package/src/util.ts CHANGED
@@ -103,3 +103,8 @@ export function tokensOf(m: Any): number {
103
103
  }
104
104
  return Math.ceil(chars / 4);
105
105
  }
106
+
107
+ /** Custom session entry pi-zip appends right before the prompt of a cold return. A warm plan holds back what that return protected
108
+ * (plan.ts FOLD_HOLD); the entry makes the return a fact of the session, so a restart reads it instead of guessing from timestamps at
109
+ * the TTL's edge. Custom entries never reach the model. */
110
+ export const COLD_CUSTOM = "pi-zip/cold";