pi-codex-compaction 0.1.5 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,35 @@ This project follows the spirit of [Keep a Changelog](https://keepachangelog.com
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.2.1] - 2026-10-09
10
+
11
+ ### Fixed
12
+
13
+ - Resolve Pi AI helpers relative to its root entry to avoid Pi's extension-loader alias rewriting submodule imports into invalid paths. Include Pi AI as a runtime dependency and test loading outside the monorepo.
14
+
15
+ ## [0.2.0] - 2026-10-09
16
+
17
+ ### Added
18
+
19
+ - Default-on automatic server-side compaction for eligible public OpenAI Responses requests, including ChatGPT subscription authentication. Keep `/server-compaction off`, persisted overrides, and optional thresholds.
20
+ - Commit streamed checkpoints through Pi's public turn-end boundary and replay the latest checkpoint plus its exact output suffix, including mid-response checkpoints and tool continuations.
21
+ - Bound checkpoint storage, bind replay to the model and credential, preserve readable fallbacks and normal Pi compaction, and avoid counting ordinary response usage twice.
22
+ - Add real-session regression coverage and an explicitly gated, synthetic-text subscription comparison. The small live Astra trial verified replay but did not demonstrate lower total latency or token use than standard Pi compaction.
23
+ - Verify immediate text streaming and tool continuation without a second Pi compaction when its safety net is enabled. Add flow-timing instrumentation to the gated benchmark; the earlier live trial did not measure streaming gaps.
24
+
25
+ ### Changed
26
+
27
+ - Clarify provider/auth billing boundaries and distinguish automatic public compaction from standalone `/responses/compact` and removed RemoteCompactionV2.
28
+
29
+ - Require Pi 1.1.0 or later for the public turn-end boundary and normalized-transcript APIs.
30
+ - Scope cache-warming interception to the actual automatic-compaction request instead of a global enabled toggle.
31
+ - Update the shared Pi development and contract-test baseline to 1.1.0; require Node.js >=22.19.0 to match the host runtime. Pi remains a host-supplied peer dependency.
32
+
33
+ ### Removed
34
+
35
+ - **Breaking:** Remove legacy `openai-codex` RemoteCompactionV2 support, direct backend transport, beta headers, temporary `pi-ai/compat` serialization, old transport helper exports, legacy TUI confirmations, and direct-compaction event-bus contracts.
36
+ - Remove the unused direct `pi-tui` dependency. Existing legacy session files remain untouched, with readable fallback rather than cross-provider opaque replay.
37
+
9
38
  ## [0.1.5] - 2026-09-17
10
39
 
11
40
  ### Fixed
package/CONTRIBUTING.md CHANGED
@@ -4,57 +4,59 @@ Thanks for your interest in contributing to `pi-codex-compaction`.
4
4
 
5
5
  ## Development setup
6
6
 
7
+ Run from the monorepo root:
8
+
7
9
  ```bash
8
10
  npm install
9
11
  npm run -w packages/pi-codex-compaction check
12
+ npm test -w packages/pi-codex-compaction
13
+ npm run -w packages/pi-codex-compaction pack:dry-run
10
14
  ```
11
15
 
12
- This package is source-distributed: Pi loads the TypeScript extension files directly. There is no build step for runtime use.
16
+ The package is source-distributed. Pi loads TypeScript directly; there is no
17
+ runtime build step.
13
18
 
14
19
  ## Pull request checklist
15
20
 
16
- Before opening a pull request:
17
-
18
- - Run `npm run -w packages/pi-codex-compaction check`.
19
- - Run `npm test -w packages/pi-codex-compaction`.
20
- - Run `npm audit --omit=dev`.
21
- - Run `npm run -w packages/pi-codex-compaction pack:dry-run` and confirm the package contents are intentional.
22
- - Update `README.md` if user-visible behavior changes.
23
- - Update `CHANGELOG.md` for notable changes.
24
- - Keep examples and paths generic; do not commit API keys, tokens, auth headers, local settings, or provider configuration containing secrets.
21
+ - Run package checks and tests.
22
+ - Run `npm run validate` for security-sensitive or cross-package changes.
23
+ - Inspect packed contents; never publish tests, session fixtures, or local settings.
24
+ - Update README, CHANGELOG, and SECURITY when their documented behavior changes.
25
+ - Never commit credentials, auth headers, private history, or opaque checkpoints.
25
26
 
26
27
  ## Coding guidelines
27
28
 
28
- - Keep the Codex provider/API capability check explicit and future-compatible.
29
- - Preserve current-model compaction, context bounds, cancellation, HTTPS, and standard Pi fallback behavior.
30
- - Add a regression test when changing wire parsing, request construction, checkpoint rehydration, or model-switch fallback behavior.
31
-
32
- ## Token-budget regression and smoke checks
33
-
34
- `tests/integration.test.mjs` uses the real Pi compaction hook, serializer, and
35
- session tree with fake credentials and mocked Responses transport. A synthetic
36
- long transcript must produce a remote checkpoint even when its UTF-8 bytes
37
- exceed the numeric token budget. Oversized input must still invoke the standard
38
- compactor and record a safe fallback reason. A later model request must not
39
- contain the diagnostic entry. No private session fixtures or live requests are
40
- needed for these tests.
41
-
42
- The token conversion follows Codex's ordinary JSON-item heuristic:
43
- [byte-to-token estimate](https://github.com/openai/codex/blob/654b0a77d0d2f81aa21f61caf7af4be88fe550bb/codex-rs/utils/string/src/truncate.rs)
44
- and [history sizing](https://github.com/openai/codex/blob/654b0a77d0d2f81aa21f61caf7af4be88fe550bb/codex-rs/core/src/context_manager/history.rs).
45
- This package keeps all serialized opaque/image bytes in its estimate instead of
46
- copying Codex's modality-specific discounts. Never equate bytes and tokens or
47
- remove the independent wire-size ceiling.
48
-
49
- For an offline replay, reconstruct a compaction's discarded history in memory,
50
- use fake authentication, replace `fetch` with a fixture checkpoint response,
51
- and call `createRemoteCompaction`. Print only sizes and the success/fallback
52
- code. Do not save the transcript, request, auth headers, or opaque content.
53
-
54
- For an approved live smoke test, follow the procedure in [README.md](./README.md#reference-and-smoke-test).
55
- Check for `fromHook: true` and `details.kind: "pi-codex-compaction"` on success.
56
- On fallback, inspect only the reason and counters in the custom diagnostic
57
- entry. Never inspect or print `encryptedContent`.
29
+ Use public Pi request, stream, and turn-end APIs. Preserve normal streaming and
30
+ the effective request rather than constructing a separate compaction request.
31
+ Do not restore legacy provider transport or the temporary compat serializer.
32
+
33
+ Automatic mode defaults on only for eligible public OpenAI requests. Preserve
34
+ off overrides, safe checkpoint bounds, cancellation, credential binding, ordinary
35
+ usage accounting, and standard Pi recovery. Never infer model-token counts from
36
+ ciphertext length or equate request count with cost.
37
+
38
+ ## Regression and smoke tests
39
+
40
+ `tests/automatic.test.mjs` exercises real Pi sessions, serializer, tool loops and
41
+ session projection with fake credentials and mocked Responses streams.
42
+ `tests/default-mode.test.mjs` checks default policy, unsupported routes and
43
+ cache-warming decisions. No live requests are needed for ordinary validation.
44
+
45
+ Add integration coverage when changing checkpoint capture, ordering, replay,
46
+ threshold coordination, or message normalization. Keep a stream open to verify
47
+ text delivery before completion. Successful checkpoint adoption must not start
48
+ a second Pi compaction between tool continuations. Missing checkpoints must
49
+ leave the standard compactor available.
50
+
51
+ The gated `tests/benchmark-automatic.mjs` consumes subscription usage and is not
52
+ part of `npm test`. Obtain explicit approval before every live run. Use synthetic
53
+ text only, bound requests and input, and never print raw provider errors,
54
+ credentials, or encrypted content. `tests/AUTOMATIC_BENCHMARK.md` documents the
55
+ existing measurements and their limits.
56
+
57
+ For an approved TUI smoke test, follow the procedure in README. Check saved
58
+ checkpoint replay after reload, branch navigation and mode-off. Do not interpret
59
+ passing protocol tests as proof of long-session latency or cost savings.
58
60
 
59
61
  ## Code of conduct
60
62
 
package/README.md CHANGED
@@ -1,19 +1,19 @@
1
1
  # pi-codex-compaction
2
2
 
3
- Keep long Pi sessions usable on OpenAI Codex models by replacing Pi's local summary request with Codex's provider-side **RemoteCompactionV2** checkpoint when the current model supports the Codex Responses API.
3
+ Keep long Pi sessions flowing with **automatic server-side compaction** inside
4
+ ordinary OpenAI Responses requests. The server can emit an encrypted checkpoint
5
+ and continue inference without Pi stopping for a separate summary request.
4
6
 
5
- ## Features
7
+ - **On by default** for eligible public `openai` GPT-5/GPT-6 requests.
8
+ - Uses your existing ChatGPT subscription login or API-key authentication without
9
+ changing providers, models, credentials, or billing mode.
10
+ - Preserves Pi's effective tools, streamed text, tool continuations, and Fast tier.
11
+ - Saves checkpoints on the session branch for reload and continuation.
12
+ - Keeps `/server-compaction off`, manual `/compact`, and Pi's standard compactor
13
+ as recovery options.
6
14
 
7
- - Uses the current `openai-codex` model for each compaction; it never silently changes the session model.
8
- - Sends only the history Pi is discarding, plus the previous Codex checkpoint, so the incoming/kept user message is not duplicated.
9
- - Retains the normal Codex Responses request envelope, including system instructions, active tool schemas, reasoning settings, prompt-cache fields, and routing fields.
10
- - Persists Codex's opaque encrypted checkpoint and rehydrates it only for supported Codex requests.
11
- - Reuses checkpoints only for the same model, trusted endpoint, Codex account, and authentication mode.
12
- - Bounds input with a Codex-style UTF-8 token estimate and a separate hard byte limit, trims tool output when necessary, retries transient failures, and honors cancellation.
13
- - Falls back to standard Pi compaction on failure or when custom compaction instructions are requested.
14
- - Keeps a bounded readable transcript excerpt so switching models or providers remains usable.
15
-
16
- The current Codex catalog includes `gpt-6-astra`, `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna`. Capability detection follows the provider/API contract (`openai-codex` + `openai-codex-responses`) rather than a brittle model-name list. Use Pi 0.85.1 or later.
15
+ Requires **Pi 1.1.0 or later**. The package name is unchanged, but the legacy
16
+ `openai-codex` provider is no longer supported.
17
17
 
18
18
  ## Installation
19
19
 
@@ -33,109 +33,175 @@ Or load it for one run:
33
33
  pi -e /path/to/pi-mono/packages/pi-codex-compaction
34
34
  ```
35
35
 
36
- ## Behavior
37
-
38
- After a Codex checkpoint is saved, the TUI shows
39
- `[compaction (codex)] Checkpoint saved.` Pi's built-in `[compaction]` heading
40
- remains unchanged. The notice is stored as a TUI-only custom session entry, so it
41
- survives the compaction chat rebuild and reload. It is not added to model context.
42
- Standard Pi compaction does not create a Codex success notice.
43
- Print, JSON, and RPC modes receive no extra notification.
44
-
45
- When Pi starts compaction on a supported Codex model, the extension sends a streamed Responses request whose `input` contains only discardable history, any compatible prior checkpoint, and a `compaction_trigger` item. The normal request envelope is retained because Codex's compaction path is parity-tested against ordinary Responses requests; this includes the effective system prompt, active tool definitions, reasoning level, prompt-cache fields, and routing fields. The request uses the `remote_compaction_v2` beta feature, and the returned opaque checkpoint and bounded provider usage are stored in the Pi compaction entry. Later requests rehydrate the raw checkpoint only when the model, endpoint, account, and authentication mode match; other providers/models receive the bounded textual fallback instead.
46
-
47
- Compaction uses the model active when Pi triggers it. If a session switches from a larger to a smaller model, the remote request is bounded against the new model's context window and tool outputs are reduced before sending. A previous opaque checkpoint is treated as incompatible after a model, endpoint, account, or authentication-mode switch; Pi's readable previous summary is sent instead. If the full request still cannot fit, the extension leaves compaction to Pi's normal implementation.
48
-
49
- If the remote request fails or returns an unexpected response, Pi's standard compaction path runs. Cancellation remains cancelled. Custom compaction instructions also use Pi's standard path because RemoteCompactionV2 has no documented custom-instructions field. The direct checkpoint request is restricted to `https://chatgpt.com`, rejects redirects, limits request/response size, and never decodes or logs `encrypted_content`. No configuration is required.
50
-
51
- ### Size limits
52
-
53
- Input is estimated as `ceil(UTF-8 request bytes / 4)`, with 8,192 tokens reserved
54
- from the active model's context window. This uses Codex's ordinary-item heuristic,
55
- not an exact tokenizer. The complete transformed request is counted, including
56
- system instructions, tool definitions, and routing fields. Opaque checkpoints
57
- and image data remain counted at their serialized size; they are not decoded or
58
- discounted. Non-ASCII text uses UTF-8 bytes, not JavaScript string length.
59
-
60
- The uncompressed request also has an independent **16 MiB hard limit**. Tool
61
- outputs are reduced only when one of these limits is exceeded. User messages,
62
- tool calls, and opaque checkpoints are not removed. If the remaining request
63
- still cannot fit, or the model's context limit is unknown, standard Pi compaction
64
- runs. The estimate can differ from the server's token count; a server rejection
65
- still uses the existing fallback.
66
-
67
- ### Fallback diagnostics
68
-
69
- A fallback on a supported model records a local custom session entry with type
70
- `pi-codex-compaction:fallback:v1`. It contains `version: 1` and a reason:
71
- `custom-instructions`, `auth-unavailable`, `request-unavailable`,
72
- `context-window-unavailable`, `context-limit`, `request-size-limit`, or
73
- `remote-failed`. Size failures also include estimated tokens, token budget,
74
- request bytes, byte limit, and the number of tool outputs reduced.
75
- Unexpected preparation failures use `request-unavailable`; `remote-failed`
76
- is reserved for failures from the transport call.
77
-
78
- No prompt, tool content, encrypted checkpoint, account identifier, credential, or
79
- raw provider error is included. These entries are not sent to the model. Pi
80
- shows a warning only in TUI mode when notifications are available; print, JSON,
81
- and RPC modes get no extra notifications or console output. Unsupported models
82
- and cancelled attempts do not create fallback diagnostics. Diagnostic storage
83
- or notification failure does not stop the standard compactor.
84
-
85
- ## Development
86
-
87
- ### Other Codex extensions
88
-
89
- With `pi-fast` installed, direct compaction requests use the current Fast toggle.
90
- With `pi-codex-tools` installed, `apply_patch` keeps its raw grammar definition,
91
- custom-tool calls, and custom-tool results during compaction. Neither package is
92
- required. No package reads a private Pi tool registry.
93
- With `pi-openai-reasoning` installed, verified Astra requests keep the original
94
- request effort and receive the current effort as a configuration update.
95
- Failed compaction does not change saved reasoning state.
96
-
97
- Pi 0.85.1 does not expose grammar metadata in `getAllTools()`. Two synchronous,
98
- versioned `pi.events` contracts let cooperating extensions supply it:
99
-
100
- - `pi-codex-compaction:tools:v1`: `{ model, tools }`, before provider serialization.
101
- A tool owner can attach its own `constrainedSampling` metadata.
102
- - `pi-codex-compaction:request:v1`: `{ ctx, messages, payload }`, after input
103
- assembly and before size checks. A listener can replace `payload`. This event
104
- is not the general `before_provider_request` chain and does not carry auth.
105
- Size checks use the transformed envelope, including field removals.
106
-
107
- Other extensions' private request changes are not applied automatically.
108
- Unknown third-party grammar metadata needs cooperation through the tools event.
109
- The checkpoint entry includes standard Pi `usage`, including cache reads and
110
- cache writes, as well as the original bounded token counters in `details`.
111
- Costs follow Pi's catalog estimates. They are not a ChatGPT subscription bill.
112
-
113
- ### Reference and smoke test
114
-
115
- The automated display smoke tests exercise Pi's actual compaction-end handler
116
- and chat renderer for manual, threshold, and overflow compaction. They check
117
- redraw, reload, standard compaction, and exclusion from subsequent model requests.
118
- Only provider responses and terminal I/O are replaced with local fixtures.
119
-
120
- Behavior was checked against `openai/codex` commit
121
- `654b0a77d0d2f81aa21f61caf7af4be88fe550bb` (2026-09-11), notably
122
- `core/src/compact_remote_v2{,_attempt}.rs`. No Codex code was copied.
123
- RemoteCompactionV2 is a changing Codex protocol, not the public `/responses/compact`
124
- API. Async tools and mid-turn steering require upstream Pi support.
125
-
126
- For a small live test, load this package and select `openai-codex/gpt-6-astra`.
127
- Send two short messages, run `/compact`, then ask about the first message.
128
- Confirm `[compaction (codex)] Checkpoint saved.` remains visible after the
129
- compaction finishes and after `/reload`. Then run
130
- `/compact Focus on recent work` and confirm the standard-compaction warning
131
- appears without a new Codex success notice.
132
- Repeat with `/fast on` and `pi-codex-tools` loaded. Check that compaction succeeds,
133
- the continuation retains context, and session usage includes compaction tokens.
134
- Use only a temporary file if you test `apply_patch`. Do not generate images.
36
+ Select a GPT-5/GPT-6 model on provider `openai`, API `openai-responses`, at
37
+ `https://api.openai.com/v1`. No enable command is needed. A previously saved
38
+ `off` setting on the current session branch remains respected.
39
+
40
+ Load this package **after extensions that transform provider requests**. Pi
41
+ orders handlers by extension load order. The compatibility guard cannot prevent
42
+ a later extension from adding incompatible fields.
43
+
44
+ ## Controls and thresholds
45
+
46
+ ```text
47
+ /server-compaction status
48
+ /server-compaction off
49
+ /server-compaction on
50
+ /server-compaction on 100000
51
+ ```
52
+
53
+ Commands persist the mode and optional token threshold on the current session
54
+ branch, not in global or project settings. `on` without a number uses the startup
55
+ threshold, or the default when none was supplied.
56
+
57
+ `--server-compaction` defaults to true.
58
+ `--server-compaction-threshold 100000` supplies a startup threshold.
59
+ Saved command settings take precedence.
60
+
61
+ The default threshold is 60% of Pi's local trigger
62
+ (`contextWindow - reserveTokens`), including per-model reserve overrides. This
63
+ is a headroom heuristic, not an OpenAI-prescribed optimal threshold. Explicit
64
+ thresholds must be integers of at least 1,000 and below Pi's local trigger.
65
+ Invalid or too-late thresholds leave ordinary inference unchanged.
66
+
67
+ Very small thresholds can compact repeatedly within one response and increase
68
+ latency and token use. Do not use the 1,000-token protocol-test threshold for
69
+ ordinary long sessions.
70
+
71
+ ## How continuation works
72
+
73
+ The extension adds
74
+ `context_management: [{ type: "compaction", compact_threshold: ... }]`,
75
+ `store: false`, and `stream: true` to eligible ordinary requests. It does not
76
+ construct a separate compaction request or reconstruct tool declarations.
77
+ Pi's actual request retains grammar tools, hidden loadouts, reasoning options,
78
+ Fast settings, and declarations carried through `additional_tools`.
79
+
80
+ Text streams normally. After successful completion, the extension commits a
81
+ Pi compaction entry at `turn_end`, before the next assistant response. It stores
82
+ the latest encrypted checkpoint and the exact provider output following it.
83
+ On subsequent requests, it replaces the readable fallback and the retained
84
+ assistant's serialized items with that checkpoint and output suffix. Later
85
+ tool results and messages remain.
86
+
87
+ The original transcript stays in the session file. Reload and tree navigation
88
+ use the selected branch, not a global cache. Turning the mode off stops requesting
89
+ new checkpoints but still replays a compatible saved checkpoint while the
90
+ extension remains loaded.
91
+
92
+ Successful adoption avoids a separate Pi compaction lifecycle, including between
93
+ tool continuations. It does **not** promise zero server-side delay or concurrent
94
+ inference during the server's compaction pass. It does not set `background: true`.
95
+
96
+ ## Support and recovery
97
+
98
+ Authentication stays with Pi: `/login openai` with **Sign in with ChatGPT**
99
+ uses plan quota; an OpenAI API key uses separate API billing. Neither path
100
+ falls back to the other. This is automatic `context_management` in ordinary
101
+ `/responses`, not the standalone `/responses/compact` API or the removed private
102
+ RemoteCompactionV2 protocol.
103
+
104
+ | Route | Behavior |
105
+ | --- | --- |
106
+ | `openai` + `openai-responses`, official endpoint, GPT-5/GPT-6 | Automatic compaction on by default; server/model availability still applies |
107
+ | Same route with ChatGPT sign-in | Live checkpoint/replay verified with `gpt-6-astra` |
108
+ | Same route with an API key | Mocked contract coverage; no live billing test |
109
+ | Legacy `openai-codex`, custom endpoints, other providers, routed model mismatch | Left unchanged; no compaction or auth hooks for those routes |
110
+
111
+ Virtual-model selections are also ineligible, including when their physical
112
+ dispatch returns to the same public OpenAI model. Pi 1.1.0 request hooks expose
113
+ the selected virtual model, not a complete physical request/auth identity.
114
+ Existing checkpoints use their bounded readable fallback rather than opaque
115
+ replay. Pi's ordinary routed context-limit checks and standard compaction remain
116
+ available; no checkpoint identity is inferred from the payload's model name.
117
+
118
+ GPT-6.1 Sol matches the GPT-6 candidate guard, but the existing live replay
119
+ evidence is Astra-only. Model eligibility is not proof of account availability.
120
+ No default model is changed. See the [provider migration assessment](../pi-openai-reasoning/README.md#provider-migration-assessment)
121
+ for the separate reasoning-update limitation.
122
+
123
+ - Requests with `configuration_update`, `compaction_trigger`, an existing
124
+ `context_management`, stateful continuation, truncation, background processing,
125
+ or enabled multi-agent mode are not opted in.
126
+ - Keep Pi's normal automatic compaction enabled. If no checkpoint is adopted,
127
+ its normal threshold and overflow recovery remain available. Manual `/compact`,
128
+ including custom instructions, uses Pi's standard summarizer.
129
+ - Failed, cancelled, malformed, oversized, or unsupported output does not
130
+ replace history. An unsuccessful automatic attempt pauses new automatic
131
+ requests until `/server-compaction on` or session reload. The extension does
132
+ not issue an extra paid retry or switch credentials; Pi owns its own retries.
133
+ - Adoption supports reasoning, single-text assistant messages, function/custom
134
+ calls, and compaction items. Other output shapes leave history intact.
135
+ - Replay requires the same model, endpoint, auth mode, credential and relevant
136
+ headers. Credential rotation, including OAuth refresh, conservatively uses
137
+ the readable excerpt rather than an unverified checkpoint.
138
+ - Edits or transforms of the retained assistant prevent stale raw replay. The
139
+ fallback is a bounded excerpt, **not a complete summary**. Edits to history
140
+ already compacted cannot retroactively change an opaque checkpoint.
141
+ - Pi cannot measure the opaque checkpoint's token footprint natively. Local
142
+ context estimates are not exact server-context measurements.
143
+ - Cache warming is stopped only when the last ordinary request actually enabled
144
+ automatic compaction. Turning the toggle off does not make an already-cached
145
+ automatic request safe to replay; a new ordinary non-automatic request clears
146
+ that restriction. Unrelated providers retain their normal warming behavior.
147
+ - Usage stays on the ordinary assistant response and is counted once. Bounded
148
+ success diagnostics stay local; they contain timing and size counters, not
149
+ conversation content.
150
+
151
+ ## Migration from the legacy provider
152
+
153
+ Legacy RemoteCompactionV2 requests, `chatgpt.com/backend-api` transport, beta
154
+ headers, the temporary `pi-ai/compat` serializer, old transport helper exports,
155
+ and the `pi-codex-compaction:tools:v1` / `:request:v1` event-bus contracts have
156
+ been removed. Fast and grammar integrations now use the ordinary Pi pipeline.
157
+
158
+ Existing legacy session files are not modified or deleted. Old encrypted
159
+ checkpoints are not migrated or replayed across providers; their readable
160
+ fallback remains available. Use compatible package versions together.
161
+ `pi-fast` and `pi-codex-tools` still support their own legacy-provider use cases;
162
+ this change removes only their obsolete direct-compaction adapters.
163
+
164
+ `pi-openai-reasoning` remains a legacy-provider-only extension. This package
165
+ does not migrate it or enable its reasoning-update feature on the public route.
166
+ Automatic compaction and histories containing `configuration_update` are
167
+ incompatible in the public API.
168
+
169
+ ## Validation and measured results
170
+
171
+ Tests use real Pi sessions with mocked transport to cover effective loadouts,
172
+ streaming, tool continuations, checkpoint ordering, branches, reloads,
173
+ cancellation, credential changes, context edits, default-on behavior, and
174
+ the standard compaction safety net.
175
+
176
+ A small live subscription test with `gpt-6-astra` verified checkpoint adoption,
177
+ replay, and synthetic fact recall. At a deliberately low 1,000-token threshold,
178
+ automatic mode used two requests versus four for standard Pi. The first answer
179
+ completed sooner, but overall elapsed time and reported token use were higher.
180
+ That protocol test did not measure long-session streaming continuity and does
181
+ not establish cost or latency savings.
182
+
183
+ In a checkout, `tests/AUTOMATIC_BENCHMARK.md` contains the measurements, limits,
184
+ and Codex source comparison. `tests/benchmark-automatic.mjs` is an explicitly
185
+ gated live test and is not run by `npm test`. Live testing requires approval to
186
+ consume usage; use synthetic text, not private session history or images.
135
187
 
136
188
  ```bash
137
- npm install
138
189
  npm run -w packages/pi-codex-compaction check
139
190
  npm test -w packages/pi-codex-compaction
140
191
  npm run -w packages/pi-codex-compaction pack:dry-run
141
192
  ```
193
+
194
+ For an approved TUI smoke test, cross the configured threshold in a synthetic
195
+ session, then check facts from earlier turns and a tool continuation. Check replay
196
+ after `/reload`, mode-off, and branch navigation. Successful adoption should
197
+ not start Pi's separate compaction lifecycle.
198
+
199
+ References:
200
+
201
+ - [OpenAI compaction guide](https://developers.openai.com/api/docs/guides/compaction)
202
+ - [ChatGPT subscription inference](https://developers.openai.com/siwc/token-sharing-open-source/models-and-inference)
203
+ - [Subscription restrictions](https://developers.openai.com/siwc/token-sharing-open-source/preview-limitations)
204
+ - [Reasoning-update restrictions](https://developers.openai.com/api/docs/guides/deployment-checklist)
205
+
206
+ Install telemetry is best-effort, once per version. Disable it with
207
+ `PI_OFFLINE=1`, `PI_TELEMETRY=0`, or `enableInstallTelemetry: false`.
package/SECURITY.md CHANGED
@@ -7,46 +7,74 @@ Security fixes are provided for the latest released version of `pi-codex-compact
7
7
  ## Reporting a vulnerability
8
8
 
9
9
  Please do not open a public issue for suspected security vulnerabilities.
10
-
11
- Report privately through [GitHub Security Advisories](https://github.com/jvm/pi-mono/security/advisories/new) or by contacting the repository maintainer through GitHub. Include:
12
-
13
- - a description of the issue;
14
- - steps to reproduce;
15
- - affected versions or commits, if known;
16
- - any suggested mitigation.
10
+ Report privately through
11
+ [GitHub Security Advisories](https://github.com/jvm/pi-mono/security/advisories/new)
12
+ or contact the maintainer through GitHub. Include the affected version,
13
+ description, reproduction steps, and any suggested mitigation.
17
14
 
18
15
  ## Security model
19
16
 
20
- `pi-codex-compaction` is a Pi package. Pi extensions execute with the same permissions as the local user running Pi. Users should review installed Pi packages and only install packages from sources they trust.
21
-
22
- For supported `openai-codex` models, the extension sends the portion of conversation Pi is about to discard to the OpenAI Codex Responses endpoint over HTTPS, using credentials and provider headers resolved by Pi. The request also includes the normal Codex Responses envelope: the effective system prompt, active tool schemas, reasoning settings, and provider-generated cache/routing fields. It stores the provider-issued opaque `encrypted_content` checkpoint in the local Pi session's compaction details so compatible Codex requests can reuse it. The checkpoint is not decoded, printed, or logged.
23
-
24
- The extension bounds request, compaction input, and response size, validates the HTTPS endpoint against the official `chatgpt.com` origin, rejects redirects, validates the account claim shape, binds checkpoints to a hashed account identity/model/endpoint/authentication mode, honors Pi cancellation, and retries only transient transport failures. It falls back to standard Pi compaction on authentication, transport, response, or context-limit failures. A bounded textual transcript excerpt remains in the compaction summary so switching to another model or provider does not leave the session with only an unusable Codex checkpoint. The fallback may contain conversation content already present in the local session and is still subject to the user's normal Pi session-file permissions.
25
-
26
- The extension never logs prompts, conversation contents, credentials, authorization headers, or raw provider responses. Install/update telemetry is best-effort and sends only package/version/runtime metadata; it can be disabled with `PI_OFFLINE=1`, `PI_TELEMETRY=0`, or Pi's `enableInstallTelemetry: false` setting.
27
-
28
- Direct requests allow only the official `/backend-api/codex/responses` endpoint
29
- on the default HTTPS port, without query strings or fragments. Total request
30
- time is limited to five minutes, including retries. Completed streams are closed
31
- without waiting for a server disconnect. A pre-aborted request does no network I/O.
32
- Null auth headers remove matching model headers; beta features are merged.
33
-
34
- Cooperating local extensions can inspect and transform compaction inputs through
35
- the documented event bus before size checks. These events contain no credentials.
36
- They have the same trust level as other installed Pi extensions.
37
-
38
- Context sizing uses `ceil(UTF-8 serialized request bytes / 4)` with an 8,192-token
39
- reserve from the active model's context window. It is an estimate, not a strict
40
- tokenizer bound. The complete transformed envelope is counted, including opaque
41
- content at its serialized size. A separate 16 MiB uncompressed request limit is
42
- enforced before network I/O. The HTTPS, redirect, response-size, checkpoint
43
- compatibility, and cancellation checks remain independent of token estimation.
44
-
45
- Fallback diagnostics store only a fixed reason code and finite non-negative
46
- size counters in `pi-codex-compaction:fallback:v1` custom session entries.
47
- These local records never include credentials, account/model identifiers,
48
- request content, encrypted checkpoints, or raw errors, and do not enter model
49
- context. No external diagnostic telemetry is added. They use the existing
50
- session's permissions and retention policy.
51
-
52
- See [CONTRIBUTING.md](./CONTRIBUTING.md) for development and validation instructions.
17
+ Pi extensions execute with the permissions of the local user. Install only
18
+ trusted packages. This extension enables automatic compaction by default only
19
+ for eligible `openai` GPT-5/GPT-6 models on `openai-responses` at the exact
20
+ official endpoint `https://api.openai.com/v1`.
21
+
22
+ It transforms ordinary Pi requests; it does not make direct HTTP inference
23
+ requests, register a provider, or use ChatGPT's legacy backend endpoints.
24
+ Pi owns TLS, transport limits, cancellation, provider retries, authentication,
25
+ and credential refresh. The extension does not copy credentials between
26
+ providers or fall back from subscription authentication to API-key billing.
27
+
28
+ ## Checkpoints and local data
29
+
30
+ Raw `response.output_item.done` items are copied, bounded to 1,024 items and
31
+ 4 MiB total, and adopted only after matching successful stream completion.
32
+ Added/partial items are never treated as checkpoints. Unknown output forms,
33
+ missing indices, failed streams, and cancellation leave history intact.
34
+ Encrypted checkpoint strings are additionally limited to 2 million characters.
35
+ The complete persisted details are bounded to 4 MiB.
36
+
37
+ The checkpoint and its exact post-checkpoint output suffix are sensitive session
38
+ data. Provider encryption does not make the surrounding transcript or suffix
39
+ public. Neither is logged or sent through install telemetry.
40
+
41
+ A domain-separated HMAC-SHA-256 binds replay to the current credential, auth
42
+ mode, model, endpoint and headers. The provider-issued, high-entropy bearer
43
+ credential and potentially secret headers form the HMAC key material and are
44
+ never stored in checkpoint details. Only non-secret routing/auth-mode metadata
45
+ is used as the HMAC message. This is
46
+ credential binding, not storage or verification of user-chosen passwords.
47
+ Rotation invalidates replay, including OAuth token refresh. The fallback is
48
+ a bounded transcript excerpt, not a complete summary. The original history and
49
+ fallback remain subject to the user's Pi session-file permissions and retention.
50
+
51
+ Replay checks the canonical retained assistant and serialized output
52
+ fingerprints. Omitted, edited, or transformed messages fall back rather than
53
+ restoring stale raw content. As with Pi text summaries, edits to already-compacted
54
+ source history cannot alter an opaque checkpoint. Legacy checkpoint details are
55
+ not migrated or replayed across provider/authentication boundaries.
56
+
57
+ ## Extension composition and recovery
58
+
59
+ Load this extension after request transformers. Its compatibility check cannot
60
+ constrain a later trusted extension that changes the payload. It skips observed
61
+ reasoning configuration updates, stateful continuation, truncation, multi-agent
62
+ requests, and preexisting compaction policies.
63
+
64
+ Cache warming is stopped only for a cached request that actually enabled
65
+ automatic compaction, because replay could produce an uncommitted checkpoint.
66
+ Other providers and skipped requests retain Pi's normal warming behavior.
67
+ Turning automatic mode off still blocks warming the old automatic request until
68
+ a new ordinary request replaces it.
69
+
70
+ The extension does not execute tools, duplicate response usage, or issue an
71
+ automatic paid retry after a failed attempt. Pi's standard compaction remains
72
+ available. Local success diagnostics contain only a version, elapsed
73
+ milliseconds, byte count and output-item count.
74
+
75
+ No project settings are read. Install telemetry is best-effort, once per
76
+ version, with a five-second timeout. It sends only package/version/runtime
77
+ metadata and respects CI, `PI_OFFLINE`, `PI_TELEMETRY`, and
78
+ `enableInstallTelemetry: false`.
79
+
80
+ See [CONTRIBUTING.md](./CONTRIBUTING.md) for validation and live-test precautions.