pi-retrospect 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,15 +1,32 @@
1
1
  # pi-retrospect
2
2
 
3
- A [Pi](https://github.com/earendil-works/pi) package that gives an agent tools for
4
- exploring past Pi sessions and the messages they contain — the read side of harness
5
- self-improvement. An agent that can look back over its own previous sessions can find
6
- the conversation where it hit a given error, recover a decision it made, and audit
7
- what actually happened before repeating it.
3
+ `pi-retrospect` is a [Pi](https://github.com/earendil-works/pi) package for
4
+ learning from previous Pi sessions. It helps agents recover prior context, audit
5
+ completed work, compare approaches, understand how a task was handled, and improve
6
+ the local harness that produced it.
8
7
 
9
- **Status: 0.x, early.** One operation is implemented — `listSessions` (see
10
- [`docs/tool-api.md`](docs/tool-api.md) for its contract), registered as one
11
- codemode-callable tool. Exploring the *messages* inside a session is not built yet;
12
- that product scope stays open until it is discussed.
8
+ Past sessions can reveal repeated failures, ineffective instructions, tool friction,
9
+ delegation problems, and opportunities to improve local skills, prompts, tools, and
10
+ workflows.
11
+
12
+ The package provides two read-only tools for discovering recorded sessions and
13
+ inspecting their entries. It also includes a skill that guides agents through
14
+ retrospective analysis and an optional setting that makes steering corrections easier
15
+ to find later.
16
+
17
+ > **Status:** `pi-retrospect` is an early 0.x package. Its session-discovery and
18
+ > transcript-reading tools are implemented and usable, but broader capabilities such
19
+ > as ranking or modifying history remain out of scope until their behavior is
20
+ > discussed.
21
+
22
+ ## What can it help with?
23
+
24
+ - Recover decisions or context from an earlier session.
25
+ - Understand how a previous task was approached.
26
+ - Audit what an agent or delegated subagent actually did.
27
+ - Compare approaches used across multiple sessions.
28
+ - Find recurring failures, corrections, or misunderstandings.
29
+ - Identify improvements to instructions, skills, prompts, tools, and workflows.
13
30
 
14
31
  ## Install
15
32
 
@@ -33,11 +50,35 @@ so the minimum is stated here rather than in `package.json`.
33
50
 
34
51
  ## Using it
35
52
 
36
- The registered tool is `list_sessions`. It walks the Pi sessions root, reads only
37
- each file's header line, and returns session metadata — id, absolute path, absolute
38
- `cwd`, timestamp, fork lineage (`parentSessionPath`) — with subagent transcripts
39
- nested under the session that launched them, plus a warning for every file it had to
40
- skip.
53
+ Two tools register, and they chain: `list_sessions` names transcript files, `session_entries` opens
54
+ one of them.
55
+
56
+ **`list_sessions`** walks the Pi sessions root, reads only each file's header line, and returns
57
+ session metadata — id, absolute path, absolute `cwd`, timestamp, fork lineage
58
+ (`parentSessionPath`) — with subagent transcripts nested under the session that launched them, plus
59
+ a warning for every file it had to skip.
60
+
61
+ All of its parameters are optional, and all of them act on **top-level** sessions — a
62
+ matching parent always arrives with its complete subagent tree:
63
+
64
+ | Parameter | Default | Meaning |
65
+ | --- | --- | --- |
66
+ | `cwds` | all working directories | Absolute `cwd`s to keep. |
67
+ | `cwdMatch` | `"exact"` | `"sibling-prefix"` also keeps sibling directories whose basename extends the requested one — the shape of git worktrees placed next to the main checkout (a lexical path rule; no git metadata is read). |
68
+ | `includeCurrentSession` | `false` | Keep the session this call runs inside. By default it is dropped, with the transcripts nested under it, before filtering, sorting, and `limit` — a retrospective normally means earlier sessions. Pi names it by its session file, never by id, so a copy of it survives; an ephemeral session has no file and so excludes nothing. |
69
+ | `startTimestamp`, `endTimestamp` | unbounded | Inclusive ISO 8601 bounds, read in the **host timezone**: a bare date is one whole calendar day, and a date-time with no offset is local to the machine running the tool. |
70
+ | `sortBy`, `sortDirection` | `"timestamp"`, `"asc"` | Top-level order only; children always stay in launch order. `"desc"` puts the newest first. |
71
+ | `limit` | none | Cap on returned rows, after filtering and sorting. Headers are still all read. |
72
+
73
+ ```js
74
+ // in a codemode script — the project and its worktrees, ten newest first
75
+ const { sessions } = await tools.list_sessions({
76
+ cwds: ["/Users/me/repos/my-app"],
77
+ cwdMatch: "sibling-prefix",
78
+ sortDirection: "desc",
79
+ limit: 10,
80
+ });
81
+ ```
41
82
 
42
83
  **It is exposed to codemode, not to the model.** The tool registers with
43
84
  `exposure: "codemode"`, so it is never declared in the model's tool list and is not
@@ -45,6 +86,7 @@ activated on registration. Call it from a codemode script:
45
86
 
46
87
  ```js
47
88
  // in a codemode script
89
+ // every session but this one, oldest first
48
90
  const { sessions, warnings } = await tools.list_sessions({});
49
91
  ```
50
92
 
@@ -53,10 +95,151 @@ the tool directly. This is deliberate — the result is structured JSON that a s
53
95
  can filter before it costs context — but it does mean the tool is invisible to a
54
96
  session running without codemode.
55
97
 
98
+ **`session_entries`** takes `sessionPath` plus optional filters, and returns the entries of that file:
99
+ `{ lineNo, id, parentId, timestamp, type, messageRole, text, raw }`, where `raw` is the whole parsed JSON
100
+ line unchanged and `text` is the entry's primary human-readable body — a message's content, a system
101
+ message's content plus its prompt sections, a compaction or branch summary, a `custom_message` content,
102
+ a `context_edit` replacement, a `session_info` name, a usage note, a label, or the command of a `!`
103
+ shell run — or `null` when the entry has no such payload. Assistant thinking, tool calls, and images
104
+ never reach `text`; they are still in `raw`. Line 1 is the session header and is never returned, so
105
+ `lineNo` starts at 2 and a
106
+ malformed line costs a warning without shifting the lines after it. Unknown entry types and unknown
107
+ message roles come back verbatim, and the call cannot leave the sessions root — a relative path, a
108
+ `..` traversal, a symlink that resolves outside it, and a file whose first line is not a session
109
+ header all throw.
110
+
111
+ | Parameter | Default | Meaning |
112
+ | --- | --- | --- |
113
+ | `startLineNo`, `endLineNo` | unbounded | Inclusive physical line bounds. |
114
+ | `ids`, `parentIds` | unfiltered | Exact, case-sensitive sets of entry ids. A `null` field matches nothing, so a version 1 file is never selected. |
115
+ | `types`, `messageRoles` | unfiltered | Exact, case-sensitive sets. `messageRoles` reaches only `type: "message"` rows. |
116
+ | `search` | unfiltered | Literal substring search over `text`: `{ terms: string[], caseSensitive?: boolean }`. A row matches when its non-null `text` contains **any** term. Case-insensitive by default. |
117
+ | `startTimestamp`, `endTimestamp` | unbounded | Inclusive ISO 8601 bounds on each entry's own timestamp, read in the host timezone — the same grammar `list_sessions` uses. |
118
+ | `limit` | none | Cap on returned entries, applied after filtering. |
119
+
120
+ Filters are ANDed, values inside one array are ORed, and order is never configurable: rows come back
121
+ in file order. Filtering narrows the **result**, never the **scan** — `warnings` still describe the
122
+ whole file. To page, pass `startLineNo` one past the last `lineNo` you already read.
123
+
124
+ `search` is the one filter that is not exact. Terms are literal bytes — no pattern, no tokenization,
125
+ no glob — so `["h.llo"]` matches only `h.llo` and `["the"]` matches inside `there`; a term of `""` and
126
+ an empty `terms` array are refused rather than read as "every row". It runs over `text`, never `raw`,
127
+ so thinking, tool calls, images, and the output of a `!` shell run are unreachable (they are still in
128
+ `raw` for a script to filter), and a row whose `text` is null is never a hit. Because a `system` row's
129
+ `text` is that message's rendered prompt, an ordinary word matches harness text — the preamble, the
130
+ tool rules, every `AGENTS.md` — so AND the search with `types` or `messageRoles` when the question is
131
+ about what was said.
132
+
133
+ ```js
134
+ // in a codemode script — project the rows in the script, never hand `raw` to a model
135
+ // the newest *previous* session: this one is dropped before `limit` applies
136
+ const { sessions } = await tools.list_sessions({ sortDirection: "desc", limit: 1 });
137
+ const { entries, warnings } = await tools.session_entries({
138
+ sessionPath: sessions[0].path,
139
+ messageRoles: ["assistant"],
140
+ });
141
+
142
+ if (warnings.length > 0) return { skipped: warnings.length, warnings };
143
+
144
+ return entries.map((entry) => ({
145
+ lineNo: entry.lineNo,
146
+ said: entry.text,
147
+ stopReason: entry.raw.message.stopReason,
148
+ }));
149
+ ```
150
+
151
+ ```js
152
+ // in a codemode script — the rows that mention one error, in the conversation only
153
+ const { entries } = await tools.session_entries({
154
+ sessionPath,
155
+ search: { terms: ["ETIMEDOUT", "connection timed out"] },
156
+ messageRoles: ["user", "assistant", "toolResult"],
157
+ });
158
+
159
+ return entries.map((entry) => ({ lineNo: entry.lineNo, role: entry.messageRole, text: entry.text }));
160
+ ```
161
+
162
+ `raw` is unbounded per row — as large as the entries it keeps (a 2.4 MB session returned 2.46 MB of
163
+ `raw`) — and `text` is bounded only by the entry it was projected from (measured max 51 KB on a tool
164
+ result), so it is a codemode-only tool by design. A `system` row is the largest category `text` carries
165
+ by mean (16.7 KB, max 38.3 KB): it projects that message's rendered prompt, which is one message's own
166
+ state — a session folds several such rows to get the prompt the model actually had.
167
+
168
+ It reads stored history: no compaction, no `context_edit`, no branch selection is applied, so it is not
169
+ the model's context view. In a session file older than version 2, `id` and `parentId` come back `null`
170
+ even where the line stores them — Pi
171
+ replaces every id when it migrates such a file — and one `legacy_version` warning says so; `lineNo`
172
+ is the handle that stays valid, and `raw` keeps what was written.
173
+
174
+ ## Steering-message markers
175
+
176
+ A mid-run correction often means the agent misunderstood the task. Making those corrections findable
177
+ turns one-off friction into evidence of recurring harness problems: an instruction that is not
178
+ landing, a tool that keeps getting misused, or a repo whose `AGENTS.md` needs a clearer rule.
179
+
180
+ ### Enable marking
181
+
182
+ Installing the package registers the two read-only tools without changing input. To opt into marking,
183
+ set `markSteeringMessages` in the user-level `~/.pi/agent/settings.json` or a trusted project's
184
+ `.pi/settings.json`, which Pi merges over the user value:
185
+
186
+ ```json
187
+ {
188
+ "piRetrospect": {
189
+ "markSteeringMessages": true
190
+ }
191
+ }
192
+ ```
193
+
194
+ The default is `false`. After editing the settings file, run `/reload`.
195
+
196
+ ### What gets marked
197
+
198
+ While the setting is enabled, `pi-retrospect` prepends `STEERING: ` to steering input submitted from
199
+ the interactive UI or an RPC client while the agent is streaming. It leaves idle prompts, queued
200
+ follow-ups, extension-generated input, slash-prefixed input, and text that already starts with the
201
+ exact marker unchanged. Attached images are preserved.
202
+
203
+ The prefix is part of the user message sent to the model and stored in the transcript, not separate
204
+ metadata. This changes what the model sees — often usefully, because the correction is explicitly
205
+ labelled — as well as making the message searchable later.
206
+
207
+ ### Find and interpret markers
208
+
209
+ Use a literal, case-sensitive search over user messages, then keep entries where the automatic marker
210
+ appears as a prefix:
211
+
212
+ ```js
213
+ // in a codemode script — the marked steering messages of one session
214
+ const { entries, warnings } = await tools.session_entries({
215
+ sessionPath,
216
+ messageRoles: ["user"],
217
+ search: { terms: ["STEERING: "], caseSensitive: true },
218
+ });
219
+
220
+ return {
221
+ steeringMessages: entries
222
+ .filter(({ text }) => text?.startsWith("STEERING: ") === true)
223
+ .map(({ lineNo, text }) => ({ lineNo, text })),
224
+ warnings,
225
+ };
226
+ ```
227
+
228
+ Treat matches as high-signal candidates, not authoritative metadata. A user can type the same prefix
229
+ manually, and unmarked steering can exist when the setting was disabled or the input belonged to an
230
+ excluded category. Inspect the surrounding conversation before deciding why the user intervened, and
231
+ compare sessions before concluding that a misunderstanding recurs.
232
+
233
+ Only messages submitted while marking is enabled receive the prefix. Existing history is never
234
+ rewritten, and disabling the setting does not remove markers already stored. The full behavior is in
235
+ [`docs/tool-api.md`](docs/tool-api.md#marking-steering-messages).
236
+
56
237
  ## Reference
57
238
 
58
239
  - [`docs/tool-api.md`](docs/tool-api.md) — the contract for the operations this
59
- package exposes, including `listSessions`, its discovery rules, and its guarantees.
240
+ package exposes: `list_sessions` (discovery rules, filters, ordering, guarantees) and
241
+ `session_entries` (sessions-root confinement, line addressing, entry filters, the literal
242
+ `text` search, the `text` projection, `raw`, warning codes).
60
243
  - `test/fixtures/generate.mjs` (source repository, not in the npm tarball) — rebuilds
61
244
  the synthetic session tree the tests run against.
62
245
 
@@ -65,6 +248,40 @@ analysis this contract was derived from — lives in the author's Obsidian vault
65
248
  than in the published package, because it documents Pi's schema (which changes with
66
249
  Pi, not with this package) and quotes counts from one developer's local session store.
67
250
 
251
+ ## Development
252
+
253
+ Run the gate locally with `npm run check` — it cleans `test/tmp/`, runs the Vitest suite,
254
+ then type-checks with `tsc --noEmit`. Do not run two Vitest processes in one checkout at
255
+ once: the start-of-run purge is not concurrency-safe.
256
+
257
+ CI is `.github/workflows/ci.yml`: on pull requests, pushes to `master`, and manual dispatch it
258
+ installs with `npm ci`, runs `npm run check`, and verifies the tarball contents with
259
+ `npm pack --dry-run`. It publishes nothing.
260
+
261
+ ### Releasing
262
+
263
+ Releases are run locally with [release-it](https://github.com/release-it/release-it). Start from a
264
+ clean `master` branch that tracks its upstream and make sure npm is authenticated for this package.
265
+ Preview the interactive flow without changing Git or npm state:
266
+
267
+ ```sh
268
+ npm run release:dry-run
269
+ ```
270
+
271
+ A dry run still performs read-only prerequisite checks such as npm authentication. For a real
272
+ release, run either the interactive version selector or name the SemVer increment explicitly:
273
+
274
+ ```sh
275
+ npm run release
276
+ npm run release -- patch
277
+ ```
278
+
279
+ The release runs `npm run check`, updates `package.json` and `package-lock.json`, creates and pushes
280
+ a `chore: release vX.Y.Z` commit and `vX.Y.Z` tag, and publishes the package to npm. The existing
281
+ `prepublishOnly` guard runs the checks again immediately before publication, so a direct
282
+ `npm publish` remains protected too. This workflow does not create a GitHub Release or maintain a
283
+ changelog.
284
+
68
285
  ## License
69
286
 
70
287
  MIT — see [`LICENSE`](LICENSE).