@henryqw/pi-session-recall 0.1.8 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,10 +1,12 @@
1
1
  # `@henryqw/pi-session-recall`
2
2
 
3
- FTS5 search over past Pi sessions: single `session_search` tool with four arg-inferred modes, zero LLM calls, raw messages only.
3
+ Search past Pi sessions with local FTS5. The `session_search` tool has four modes and makes no model calls.
4
+
5
+ The package also includes `pi-session-pattern-miner`. This skill finds repeated work that may be worth automating.
4
6
 
5
7
  ## Why
6
8
 
7
- - **Created for**: Pi users who need to recover decisions and context from prior sessions without keeping every transcript in the active prompt.
9
+ - **Created for**: Recover decisions and context from prior sessions without keeping every transcript in the active prompt.
8
10
  - **Advantage**: Local FTS5 search gives fast, private recall with zero standing context cost and no model calls.
9
11
 
10
12
  ## Install
@@ -17,32 +19,56 @@ pi install npm:@henryqw/pi-session-recall
17
19
 
18
20
  | Surface | Type | Purpose |
19
21
  | --- | --- | --- |
20
- | `session_search` | tool | Search past sessions or inspect one: discovery (`query`), scroll (`sessionId` + `aroundMessageId`), read (`sessionId`), browse (no args) |
22
+ | `session_search` | tool | Search past sessions or inspect one. |
23
+ | `pi-session-pattern-miner` | skill | Find repeated work and choose the smallest useful automation. |
24
+
25
+ BM25 is a text-ranking method. Hydrated results include messages read from saved session files.
26
+
27
+ | Mode | Call | Result |
28
+ | --- | --- | --- |
29
+ | Discovery | `query` | BM25-ranked top sessions. The top hit is hydrated with a ±5 message window and first/last-3 bookends. Lower hits include the matched anchor message and metadata. `detail:"full"` hydrates all. |
30
+ | Scroll | `sessionId` + `aroundMessageId` | ±`window` messages ([1,20]) around the anchor on its branch. Re-anchor on the last or first message ID to scroll. Across forks, pass the previous response's `branchTip`; `aroundMessageId` only centers the window and must lie on that branch. |
31
+ | Read | `sessionId` | The whole session. Large sessions return head 20 + tail 10. Oversized content is bounded to 50k characters and marked with `contentTruncated`. |
32
+ | Browse | no args | Recent sessions with path, name, cwd, started date, and preview. |
33
+
34
+ In the interactive TUI, the collapsed tool block shows the last five visual lines and the earlier-line count. Press `Ctrl+O` to expand the full bounded response. The model always receives the complete tool result.
21
35
 
22
- **Discovery** — BM25-ranked top sessions; top hit hydrated with a ±5 message window and first/last-3 bookends; lower hits carry the matched anchor message plus metadata (`detail:"full"` hydrates all).
36
+ ### Find work worth automating
23
37
 
24
- **Scroll** — ±`window` messages ([1,20]) around the anchor on its branch; scroll forward/backward by re-anchoring on the last/first message id of the returned window. Across forks, pass the previous response's `branchTip` to select the branch — `aroundMessageId` only centers the window and must lie on that branch.
38
+ Run `/skill:pi-session-pattern-miner` to find repeated workflows in past sessions. It requires evidence from two independent sessions and checks for existing automation. It prefers a fixed script when model judgment is not needed.
25
39
 
26
- **Read** — whole session; head 20 + tail 10 when large, with oversized content bounded to 50k characters and flagged by `contentTruncated`.
40
+ ### Query syntax and indexed text
27
41
 
28
- **Browse** — recent sessions: path, name, cwd, started date, preview.
42
+ - Prefer distinctive identifiers, package names, issue numbers, or uncommon terms. Use quoted phrases only when exact wording is known.
43
+ - The FTS5 trigram index uses AND for multiple words by default. Use `OR` for breadth, quoted phrases for exact matches, and `NOT` to exclude. Wildcards help only stems ≥3 characters.
44
+ - Only user and assistant text is indexed. Thinking blocks and tool output are not searchable.
45
+ - For message text over the 20,000-character indexing budget, only the first and last regions are indexed. The middle is omitted. Phrases and `NEAR` cannot cross those regions, but ordinary AND terms can.
46
+ - `sessionId` must be a `.jsonl` file under the Pi sessions directory.
29
47
 
30
- Query syntax: Prefer distinctive identifiers, package names, issue numbers, or uncommon terms; use quoted phrases only when exact wording is known. FTS5 over a trigram index — multi-word = AND by default, `OR` for breadth, quoted phrases for exact match, `NOT` to exclude. Wildcards only help stems ≥3 chars. Only user/assistant text is indexed; thinking blocks and tool output are not searchable. For message text over the 20,000-character indexing budget, only first/last regions are indexed and the middle is omitted; phrases and `NEAR` cannot cross those regions, but ordinary AND terms can. `sessionId` must be a `.jsonl` file under the Pi sessions directory.
48
+ ### Context and sync
31
49
 
32
- Hits inside the current session's live context are suppressed; compacted-away or inactive-branch history stays discoverable. Forked sessions collapse into their parent when both match.
50
+ Hits inside the current session's live context are suppressed. Compacted-away or inactive-branch history stays discoverable. Forked sessions collapse into their parent when both match.
33
51
 
34
- If the lazy index sync before browse/discovery cannot fully enumerate the session tree, or throws entirely, results are still served from the current index — potentially partially updated and stale: files discovered before the failure may already reflect their new content, while rows for files the walk never reached remain stale — and carry a top-level `syncWarning`: `{kind:"incomplete-walk"}` for a partial walk (indexed-but-unseen paths are never purged in that case), or `{kind:"sync-failed", error}` with the capped failure message. The warning is omitted once a sync completes.
52
+ Before browse or discovery, lazy index sync can fail while walking the session tree. Results still come from the current index and can be partly updated or stale. Files found before failure may have new content, while rows for files the walk did not reach stay stale.
53
+
54
+ - A partial walk returns top-level `syncWarning`: `{kind:"incomplete-walk"}`. Indexed-but-unseen paths are never purged in that case.
55
+ - A total sync failure returns top-level `syncWarning`: `{kind:"sync-failed", error}` with the capped failure message.
56
+ - The warning is omitted after a completed sync.
35
57
 
36
58
  ## State
37
59
 
38
- | Path | Purpose |
39
- | --- | --- |
40
- | `~/.pi/agent/config/pi-session-recall/index.db` | Derived SQLite search index, maintained by the extension. |
60
+ The extension maintains the derived SQLite search index at `~/.pi/agent/config/pi-session-recall/index.db`.
41
61
 
42
62
  ## Deliberate exclusions
43
63
 
44
- Session directories whose encoded path starts with `--tmp-` or `--private-tmp-` (sessions run from `/tmp` or `/private/tmp`) are never indexed. Session files over 32 MiB are excluded from indexing and hydration: discovery cannot newly find them; READ/SCROLL return an explicit size error, while a stale discovery hit retained from before the file grew is returned as metadata with empty messages and that error.
64
+ Session directories whose encoded path starts with `--tmp-` or `--private-tmp-` are never indexed. These sessions run from `/tmp` or `/private/tmp`.
65
+
66
+ Session files over 32 MiB are excluded from indexing and hydration. Discovery cannot newly find them.
67
+
68
+ READ and SCROLL return an explicit size error. A stale discovery hit from before a file grew returns metadata with empty messages and that error.
45
69
 
46
70
  ## Storage & privacy
47
71
 
48
- The SQLite index lives at `~/.pi/agent/config/pi-session-recall/index.db`. It is derived state: delete it and it rebuilds from your session files. Everything stays local — transcripts are read in place and nothing leaves the machine beyond what tool results already show the model.
72
+ This is derived state. Delete it and it rebuilds from your session files.
73
+
74
+ Everything stays local. Transcripts are read in place, and nothing leaves the machine beyond what tool results already show the model.
@@ -1,9 +1,10 @@
1
1
  /**
2
2
  * pi-session-recall entry point: tool registration and mode dispatch.
3
3
  */
4
- import { getAgentDir } from "@earendil-works/pi-coding-agent";
4
+ import { getAgentDir, keyHint, truncateToVisualLines } from "@earendil-works/pi-coding-agent";
5
5
  import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
6
6
  import { StringEnum } from "@earendil-works/pi-ai";
7
+ import { Text, truncateToWidth } from "@earendil-works/pi-tui";
7
8
  import { Type } from "typebox";
8
9
  import { realpathSync } from "node:fs";
9
10
  import { join, sep } from "node:path";
@@ -119,6 +120,21 @@ export default function (pi: ExtensionAPI): void {
119
120
  limit: Type.Optional(Type.Number({ description: "Max results, [1,10], default 3." })),
120
121
  detail: Type.Optional(StringEnum(["adaptive", "full"] as const)),
121
122
  }),
123
+ renderResult(result, { expanded }, theme) {
124
+ const output = result.content.find((part) => part.type === "text")?.text ?? "";
125
+ const styledOutput = theme.fg("toolOutput", output);
126
+ if (expanded) return new Text(`\n${styledOutput}`, 0, 0);
127
+ return {
128
+ render(width: number) {
129
+ const preview = truncateToVisualLines(styledOutput, 5, width);
130
+ if (preview.skippedCount === 0) return ["", ...preview.visualLines];
131
+ const hint = theme.fg("muted", `... (${preview.skippedCount} earlier lines,`) +
132
+ ` ${keyHint("app.tools.expand", "to expand")}${theme.fg("muted", ")")}`;
133
+ return ["", truncateToWidth(hint, width, "..."), ...preview.visualLines];
134
+ },
135
+ invalidate() {},
136
+ };
137
+ },
122
138
  async execute(_toolCallId, rawParams: ToolParams, _signal, _onUpdate, ctx) {
123
139
  try {
124
140
  // LLMs sometimes send numeric ids/queries despite the string schema.
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@henryqw/pi-session-recall",
3
- "version": "0.1.8",
4
- "description": "FTS5 search over past Pi sessions: single tool, four arg-inferred modes (discovery/scroll/read/browse), zero LLM calls.",
3
+ "version": "0.2.1",
4
+ "description": "Local FTS5 search over past Pi sessions plus a skill for turning recurring work into deterministic automation.",
5
5
  "keywords": [
6
6
  "pi-package",
7
7
  "pi",
@@ -15,9 +15,10 @@
15
15
  },
16
16
  "license": "MIT",
17
17
  "files": [
18
+ "LICENSE",
18
19
  "extensions",
19
- "README.md",
20
- "LICENSE"
20
+ "skills",
21
+ "README.md"
21
22
  ],
22
23
  "scripts": {
23
24
  "test": "node --test test/*.test.ts",
@@ -27,6 +28,7 @@
27
28
  "peerDependencies": {
28
29
  "@earendil-works/pi-ai": "^0.84.4",
29
30
  "@earendil-works/pi-coding-agent": "^0.84.4",
31
+ "@earendil-works/pi-tui": "^0.84.4",
30
32
  "typebox": "^1.3.15"
31
33
  },
32
34
  "repository": {
@@ -43,6 +45,9 @@
43
45
  "pi": {
44
46
  "extensions": [
45
47
  "./extensions/session-recall.ts"
48
+ ],
49
+ "skills": [
50
+ "./skills"
46
51
  ]
47
52
  }
48
53
  }
@@ -0,0 +1,55 @@
1
+ ---
2
+ name: pi-session-pattern-miner
3
+ description: Mine repeated work, corrections, and manual procedures from past Pi sessions with session_search, then identify the smallest deterministic script or reusable skill that would prevent repeating them. Use when the user asks what to automate, which workflows recur, or what skill or script should be extracted from prior Pi sessions.
4
+ ---
5
+
6
+ # Pi Session Pattern Miner
7
+
8
+ Find repeated work in past Pi sessions and turn the best-supported pattern into an automation candidate. Use `session_search` for session history; do not scan Pi session files directly. The current model mines the evidence once; the resulting automation should remove model judgment from recurring execution wherever the workflow permits it.
9
+
10
+ ## Scope
11
+
12
+ Use the scope or topic in the request. Otherwise sample recent sessions. If the user specifies the current repository, use the current Git root and session `cwd` metadata to filter where possible; search distinctive repository or package names to find related sessions from other worktrees.
13
+
14
+ State the sampled scope and its limits. `session_search` browse returns at most ten recent sessions, discovery is query-driven, and current live-context matches are suppressed. Its messages contain user/assistant text, not hidden thinking or tool output. Never claim exhaustive coverage or invent commands that are absent from the evidence.
15
+
16
+ ## Mine
17
+
18
+ 1. Call `session_search` with only `limit: 10` to seed the sample.
19
+ 2. Read each in-scope session by `sessionId`. For a truncated session, inspect only relevant gaps with `sessionId`, `aroundMessageId`, and `window: 20`; retain `branchTip` while scrolling a fork.
20
+ 3. Extract candidate episodes: repeated user intents, corrections, manual procedures, avoidable retries, and agent-authored steps that recur. Ignore generic coding work such as inspect/edit/test unless the same concrete procedure repeats.
21
+ 4. Cluster by underlying job, not wording. A pattern needs evidence from at least two independent sessions; do not count forks, retries, or continuations of one task as separate evidence.
22
+ 5. Confirm each candidate with one or more distinctive `query` searches using `limit: 10` and `detail: "full"`. Record session path, date or name, relevant entry id, and a short paraphrase. Do not copy secrets or unnecessary transcript text.
23
+ 6. Inspect the current repository's commands, package scripts, executable scripts, skills, and agent instructions before proposing anything. Reuse or repair an existing automation when it already owns the workflow.
24
+
25
+ Stop with “not enough repeated evidence” when fewer than two independent sessions support a pattern. Do not manufacture a recommendation from one occurrence.
26
+
27
+ ## Choose the owning surface
28
+
29
+ Stop at the first option that fully handles the pattern:
30
+
31
+ 1. **Existing command or skill** — document, fix, or invoke it instead of adding another path.
32
+ 2. **Script** — choose when inputs, decisions, outputs, and failures can be specified without model judgment. Prefer a repository command or standard-library script over a new dependency.
33
+ 3. **Script plus thin skill** — choose when an agent must gather inputs or explain results, but the repeated operation itself can be deterministic. The skill must call the script rather than restate its algorithm.
34
+ 4. **Skill only** — choose only when the reusable work inherently requires judgment, repository inspection, or user decisions.
35
+ 5. **Product change** — choose when the root cause belongs in an extension, API, CI check, or other code rather than agent instructions.
36
+
37
+ A deterministic candidate must define its trigger, inputs, outputs, side effects, failure behavior, and one runnable check. If those cannot be defined from session evidence plus the current codebase, recommend a focused investigation instead of automation.
38
+
39
+ ## Report
40
+
41
+ Return at most three ranked patterns:
42
+
43
+ | Pattern | Independent sessions | Repeated cost or failure | Existing coverage | Smallest automation |
44
+ | --- | ---: | --- | --- | --- |
45
+
46
+ For each pattern, cite the session paths and entry ids, explain why the proposed owning surface fits, and state what model work it removes. Then give the highest-confidence candidate a minimal implementation contract:
47
+
48
+ - **Trigger and inputs**
49
+ - **Deterministic steps**
50
+ - **Output and side effects**
51
+ - **Failure behavior**
52
+ - **Runnable check**
53
+ - **Files to add or change**
54
+
55
+ Do not grade the agent, generate a report site, or modify files unless the user asks to implement a candidate.