kankaku 0.4.5 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/README.md +964 -16
  2. package/package.json +4 -2
  3. package/src/adapters/agent-info.ts +86 -0
  4. package/src/adapters/ancestry.ts +260 -0
  5. package/src/adapters/cached-catalog.ts +131 -0
  6. package/src/adapters/file-inflight-store.ts +32 -0
  7. package/src/adapters/file-modes.ts +35 -0
  8. package/src/adapters/hub-credentials.ts +95 -0
  9. package/src/adapters/jsonl-work-log.ts +19 -2
  10. package/src/adapters/kankaku-command.ts +729 -0
  11. package/src/adapters/kankaku-dir.ts +92 -0
  12. package/src/adapters/lazy-jsonl-work-log.ts +4 -0
  13. package/src/adapters/machine-process-registry.ts +256 -0
  14. package/src/adapters/pi-tracker.ts +432 -266
  15. package/src/adapters/pocketbase-catalog.ts +60 -0
  16. package/src/adapters/pocketbase-client.ts +197 -0
  17. package/src/adapters/pocketbase-sink.ts +224 -0
  18. package/src/adapters/process-identity-memo.ts +102 -0
  19. package/src/adapters/process-identity.ts +162 -0
  20. package/src/adapters/project-config.ts +73 -1
  21. package/src/adapters/report.ts +89 -8
  22. package/src/adapters/session-client.ts +116 -0
  23. package/src/adapters/session-dir.ts +28 -0
  24. package/src/adapters/session-target.ts +262 -0
  25. package/src/adapters/status-bar.ts +86 -0
  26. package/src/adapters/subagent-startup.ts +66 -0
  27. package/src/adapters/sync-runner.ts +301 -0
  28. package/src/adapters/sync-state-store.ts +227 -0
  29. package/src/adapters/target-picker.ts +82 -0
  30. package/src/config.ts +444 -6
  31. package/src/domain/ancestry-match.ts +84 -0
  32. package/src/domain/hub-entry.ts +339 -0
  33. package/src/domain/registry-health.ts +87 -0
  34. package/src/domain/subagent-profile.ts +495 -0
  35. package/src/domain/sync-plan.ts +251 -0
  36. package/src/domain/task-view.ts +307 -23
  37. package/src/domain/work-record.ts +165 -4
  38. package/src/domain/work-target.ts +185 -0
  39. package/src/domain/work-tracker.ts +257 -56
  40. package/src/extension.ts +303 -6
  41. package/src/ports/catalog.ts +31 -0
  42. package/src/ports/process-registry.ts +75 -0
  43. package/src/ports/work-log.ts +9 -0
  44. package/src/ports/work-sink.ts +35 -0
package/README.md CHANGED
@@ -4,6 +4,8 @@ A [pi](https://pi.dev) extension that measures how long an agent actually
4
4
  spends working on each prompt, so the time can later be accounted for
5
5
  (billing, reporting).
6
6
 
7
+ Docs and guide: [kankaku.io](https://kankaku.io).
8
+
7
9
  ## What it measures
8
10
 
9
11
  For every prompt, kankaku tracks the span from `before_agent_start` to
@@ -57,6 +59,12 @@ Each line in `worklog.jsonl` is one JSON object:
57
59
  "model": "anthropic/claude-opus",
58
60
  "client": "acme",
59
61
  "sessionName": "billing sprint",
62
+ "sessionDir": "/abs/custom/session/dir",
63
+ "clientId": "pocketbase-record-id",
64
+ "clientName": "Acme",
65
+ "projectId": "pocketbase-record-id",
66
+ "projectName": "Portal",
67
+ "machine": "laptop",
60
68
  "prompt": "first 200 chars of the first prompt",
61
69
  "startedAt": "2026-09-10T16:00:00.000Z",
62
70
  "settledAt": "2026-09-10T16:04:10.000Z",
@@ -69,7 +77,9 @@ Each line in `worklog.jsonl` is one JSON object:
69
77
  "subagents": [{ "toolCallId": "…", "agent": "sdd-explore", "mode": "task", "taskId": "t1", "ms": 90000 }],
70
78
  "segments": { "review": 62000 },
71
79
  "usage": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0, "cost": 0 },
72
- "status": "completed"
80
+ "status": "completed",
81
+ "roleConfidence": "uncertain",
82
+ "orchestratorRef": { "pid": 4000, "project": "/abs/other-worktree", "startedAt": "2026-09-10T15:59:00.000Z", "dir": "/abs/other-worktree/.kankaku" }
73
83
  }
74
84
  ```
75
85
 
@@ -77,6 +87,46 @@ Each line in `worklog.jsonl` is one JSON object:
77
87
  `stopReason: "aborted"`), or `interrupted` (pi shut down while still
78
88
  running).
79
89
 
90
+ `runs` counts the agent loops inside the record: the first one plus every
91
+ continuation pi ran before settling it (an automatic retry after a provider
92
+ error, overflow recovery, a queued steer or follow-up). Informational only.
93
+
94
+ `trigger` is optional. `"extension"` marks a record no user prompt started:
95
+ an extension woke the agent itself — this is how gentle-pi resumes the
96
+ orchestrator when a background subagent finishes. Its `prompt` is the fixed
97
+ text `(no user prompt — run started by an extension)`. Without it that work
98
+ would not be recorded at all, since pi only announces user prompts.
99
+
100
+ `roleConfidence` and `orchestratorRef` are both optional and normally
101
+ absent — see "Subagents" below. `roleConfidence` is only ever set to
102
+ `"uncertain"`, and only on an `orchestrator`-role record kankaku could not
103
+ positively prove top-level; `orchestratorRef` is only ever set on a
104
+ `subagent`-role record that discovered its tracked ancestor via the
105
+ machine-wide process registry. Its optional `dir` field carries that
106
+ orchestrator's resolved kankaku directory — the real top-level one even
107
+ across a subagent-of-subagent chain — and is what this process's own
108
+ work log and inflight checkpoints were actually routed into when it
109
+ differs from this process's own (see "Subagents" > "Cross-worktree write
110
+ routing"). Neither field, nor `orchestratorRef.dir`, bumps
111
+ `WORK_RECORD_SCHEMA` — a record without them (from an older kankaku build)
112
+ remains valid.
113
+
114
+ `clientId`, `clientName`, `projectId`, `projectName` and `machine` are only
115
+ present once a hub is configured (see "Hub (PocketBase)"); every report and
116
+ export written before this feature, or by a user without a hub, is
117
+ unaffected.
118
+
119
+ `sessionDir` is present only when pi's session manager reports a
120
+ *non-default* session directory (`--session-dir`, or a resumed session
121
+ started that way) — exactly the condition under which pi's own printed "To
122
+ resume this session: ..." line includes `--session-dir`. Most records never
123
+ carry it. `/kankaku doctor` shows it for the current session when set, and
124
+ it is available on a task's orchestrator record (`TaskView.sessionDir`) for
125
+ anything that wants to reconstruct the exact `pi --session-dir <dir>
126
+ --session <id>` resume command locally. When a hub is configured it is also
127
+ sent as `session_dir` on every sync (see "Hub (PocketBase)" > "Sync" >
128
+ "Agent and measurement quality").
129
+
80
130
  ## Task and session views
81
131
 
82
132
  Each `WorkRecord` still measures one pi process's own prompt-to-idle span.
@@ -87,11 +137,14 @@ correct for that, built purely from `pid`/`parentPid`/`startedAt`/`settledAt`
87
137
  already present on every record — no new fields are persisted to
88
138
  `worklog.jsonl`.
89
139
 
90
- - **Task**: one orchestrator record plus every subagent record matched to
91
- it — same `project`, `parentPid === orchestrator.pid`, and the child's
92
- `startedAt` falling inside the orchestrator's `[startedAt, settledAt]`
93
- window. (If a pid is reused across runs and several orchestrator records
94
- match, the child attaches to the latest-starting one.) A task's `wallMs`
140
+ - **Task**: one *confirmed* orchestrator record (see "Subagents" below —
141
+ an orchestrator-role record flagged uncertain never anchors a task) plus
142
+ every subagent record matched to it — `parentPid === orchestrator.pid`
143
+ and the child's `startedAt` falling inside the orchestrator's
144
+ `[startedAt, settledAt]` window; `project` is only a **hint**, preferred
145
+ when it matches but never a hard filter (see "Subagents"). (If a pid is
146
+ reused across runs and several orchestrator records match, a same-project
147
+ candidate is preferred, then the latest-starting one.) A task's `wallMs`
95
148
  is the **union** of the orchestrator's interval and every matched child's
96
149
  interval — never their sum — so parallel background children are not
97
150
  double-counted, and a child that outlives the orchestrator's own settle
@@ -104,10 +157,531 @@ already present on every record — no new fields are persisted to
104
157
  `waitingMs` is the sum of each task's `waitingMs`, and `workMs = wallMs -
105
158
  waitingMs`.
106
159
  - **Orphan subagents**: a subagent record with no matching orchestrator
107
- record (for example, its parent's record was lost, or it belongs to a
108
- different project) is excluded from every task but is not silently
109
- dropped — it stays visible so gaps in the log are noticeable rather than
110
- hidden.
160
+ record (for example, its parent's record was lost, or a cross-worktree
161
+ registry entry had already expired) is excluded from every task but is
162
+ not silently dropped — it stays visible so gaps in the log are noticeable
163
+ rather than hidden. See "Subagents" for how a cross-worktree child is
164
+ usually reunited *before* it ever becomes an orphan.
165
+
166
+ ## Subagents
167
+
168
+ kankaku recognises gentle-pi's `subagent_run` tool as opening a subagent
169
+ span (unchanged from before this section); this describes how it decides,
170
+ for a process that shows no such marker, whether it is a genuine top-level
171
+ session or actually someone's subagent — and how a gentle-pi subagent
172
+ running in a *different git worktree* than its orchestrator still gets
173
+ correctly counted.
174
+
175
+ ### The role/state model
176
+
177
+ Every record still carries the same binary persisted `role`
178
+ (`"orchestrator"` | `"subagent"`, unchanged — see "Record schema"). On top
179
+ of it, kankaku's task/session views and the hub sync apply a four-state
180
+ classification:
181
+
182
+ - **orchestrator** — confirmed top-level: no recognised child-env-marker
183
+ (`GENTLE_PI_AGENTS_CHILD=1`, or an explicit `KANKAKU_ROLE=orchestrator` —
184
+ see "Interactive sessions and `KANKAKU_ROLE`" below) is present, and
185
+ either no live tracked ancestor process was found, or this session is
186
+ itself interactive (see "The registry" and "Interactive sessions" below).
187
+ This is the default for a plain, ordinary `pi` session — unaffected by
188
+ any of this.
189
+ - **subagent (joined)** — a gentle-pi child matched to its orchestrator, as
190
+ described in "Task and session views" above.
191
+ - **subagent (orphan)** — a gentle-pi child that could not be matched to
192
+ any orchestrator (shown separately, never dropped — `orphanSubagents`).
193
+ - **uncertain** — no recognised child-env-marker, a live tracked ancestor
194
+ process *was* found, **and** this process is not itself an interactive
195
+ TUI session: it cannot be proven top-level, so it is never counted as a
196
+ new task locally and never synced to the hub as one, but it is not
197
+ dropped either — `WorkRecord.roleConfidence` is set to `"uncertain"` on
198
+ it, and `/kankaku doctor` (and a one-line hint on the plain `/kankaku`
199
+ summary) surface it so the gap is visible instead of silently wrong.
200
+ This is the fix for a real bug: a subagent mechanism kankaku does not
201
+ specifically recognise (for example, pi's own bundled reference
202
+ `subagent` example, which sets no env marker at all) used to default to
203
+ `"orchestrator"` outright — a phantom top-level task on top of the time
204
+ already measured inside its parent's own tool-call span, billed twice.
205
+ An unrecognised process now degrades to a safe, visible **undercount**
206
+ instead of a silent, unrecoverable **overcount**. An *interactive*
207
+ session is never demoted this way, no matter what its ancestry looks
208
+ like — see "Interactive sessions and `KANKAKU_ROLE`" below for why, and
209
+ for the escape hatch when kankaku still gets it wrong.
210
+
211
+ An `uncertain` classification is recoverable going forward: once the
212
+ mechanism is recognised (for example, by upgrading kankaku, setting
213
+ `KANKAKU_ROLE` explicitly, or — in a later version — registering it via a
214
+ configured tool/env marker), a later `/kankaku sync all` or `backfill`
215
+ picks up the record correctly. It never resolves itself by guessing. A
216
+ record that was *already written* `uncertain`, however, cannot be rewritten
217
+ after the fact — `worklog.jsonl` is append-only and kankaku never edits a
218
+ past line (see AGENTS.md) — so only a run *after* the fix correctly
219
+ anchors a task; there is no migration that goes back and reclassifies old
220
+ lines.
221
+
222
+ ### The registry
223
+
224
+ Every kankaku process writes a small entry to
225
+ `~/.kankaku/run/<pid>.json` at startup — `pid`, `parentPid`, `role`,
226
+ `project`, its resolved (and, for a routed subagent, actually-used —
227
+ see "Cross-worktree write routing" below) `KANKAKU_DIR`, `startedAt`, and
228
+ `processStartId` (below) — independent of any project's own `KANKAKU_DIR`,
229
+ so it survives a project boundary. Both `~/.kankaku/run` and its entry
230
+ files are created owner-only (`0700`/`0600` — an existing looser mode, left
231
+ by an older kankaku build, is tightened on the next write, best-effort);
232
+ they name absolute project paths and session ids. **The registry is a
233
+ startup-time lookup only** — "who is my tracked ancestor, and where does
234
+ it keep its log" — resolved once, at process factory time, and never
235
+ consulted again later as a live pointer (this used to matter: see
236
+ "Cross-worktree write routing" below for why it no longer does). This is
237
+ what powers both of the following:
238
+
239
+ - **Uncertain detection**: a process with no child-env-marker walks its own
240
+ OS ancestor chain (one snapshot, see "Ancestor-chain detection" below)
241
+ looking for *any* live registry entry whose identity it can actually
242
+ **prove** — see "Identity, not just pid" below. Finding one means some
243
+ other tracked kankaku process is an ancestor of this one; combined with
244
+ this process *not* being an interactive TUI session (see "Interactive
245
+ sessions and `KANKAKU_ROLE`" below), it is classified `uncertain` rather
246
+ than defaulting to `orchestrator`.
247
+ - **Cross-worktree write routing** (ADR 0023, rewritten for a real bug —
248
+ see below): a gentle-pi subagent running in a different git worktree than
249
+ its orchestrator walks its ancestor chain, finds its orchestrator's
250
+ registry entry (identity-verified), and resolves it to an
251
+ `orchestratorRef` (`{ pid, project, startedAt, dir }` — `dir` also
252
+ resolves through a subagent-of-subagent chain to the real, top-level
253
+ orchestrator, never a middle hop). When that orchestrator's directory
254
+ differs from this process's own, the child writes its work log **and**
255
+ its inflight crash-recovery checkpoints straight into the orchestrator's
256
+ directory instead of its own cwd-relative one — so parent and child
257
+ records end up in the *same* `worklog.jsonl` from the moment the child's
258
+ first record is appended, not merely discovered there later. The
259
+ orchestrator's later `buildTasks` call joins them with the same
260
+ `pid`/`parentPid`/project-hint keys it always has; the interval-union
261
+ rule itself is still computed in exactly one place (`buildTasks`) — this
262
+ only changes *where the bytes physically live*, never how they are
263
+ joined. If the orchestrator's directory cannot be created or written to
264
+ (gone, or no permission), the child falls back to its own local
265
+ directory instead of losing the record, and `/kankaku doctor` reports the
266
+ fallback so it can be reunited manually; a record is always written to
267
+ **exactly one** log, never both. Because reunification no longer depends
268
+ on any pointer still being alive at read time, it survives the child's
269
+ own exit cleanup removing its registry entry — which, for gentle-pi's
270
+ main case (a blocking `subagent_run` in task mode), has already happened
271
+ by the time the parent regains control. If ancestry could not be
272
+ established at all (or the write genuinely could not go anywhere), the
273
+ child stays a visible orphan instead — undercounted, never lost, and
274
+ never compensated for by summing two independently synced rows: **the
275
+ hub never sums two unions to recover a missing one**, since that would
276
+ double-count the overlap between parent and child. `project` is
277
+ therefore only ever a *hint* for the join (preferred when it matches),
278
+ never a hard filter.
279
+ - The registry is swept opportunistically (when a process writes its own
280
+ entry) — see "Registry cleanup and health" below — so it does not grow
281
+ unbounded and never keeps serving a stale identity.
282
+
283
+ #### Identity, not just pid — the PID-reuse fix
284
+
285
+ Matching an ancestor pid to a registry entry by **pid number alone** is not
286
+ safe: operating systems reuse pids. A kankaku process that dies without
287
+ cleanup (a crash, `kill -9`) can leave its `~/.kankaku/run/<pid>.json`
288
+ entry behind; the OS can later hand that same pid to the user's own
289
+ interactive shell, and every *genuine* top-level pi session launched from
290
+ that shell would then falsely resolve a "tracked ancestor" — silently
291
+ misclassified `uncertain` forever, its task never synced. This inverts the
292
+ whole guarantee this feature exists for, so identity is proven, not
293
+ assumed:
294
+
295
+ - Every registry entry also carries `processStartId`: an approximate,
296
+ self-consistent epoch-ms estimate of that process's actual OS start time.
297
+ **This process's own** `processStartId` (the one it records about
298
+ itself) is derived cheaply and portably — `Date.now() - process.uptime()
299
+ * 1000`, sampled once at factory time — with **no subprocess spawn and no
300
+ `/proc` read at all**, so it is available on every platform, Windows
301
+ included, and never adds startup cost (see "Startup cost" below).
302
+ Verifying *another* process's (an ancestor's) live identity still needs a
303
+ fresh reading of that specific pid from an OS ancestor-chain snapshot: on
304
+ macOS/BSD, `ps -eo pid,ppid,etime` (`[[dd-]hh:]mm:ss` elapsed time,
305
+ forced through the portable `etime` keyword — BSD `ps` has no `etimes`);
306
+ on Linux, `/proc/<pid>/stat`'s `starttime` (clock ticks since boot)
307
+ combined with `/proc/uptime`, assuming the near-universal `USER_HZ=100` —
308
+ a wrong assumption never causes a false match, since the same (possibly
309
+ wrong) constant is used both when an entry is written and whenever it is
310
+ re-verified, and a process's `starttime` ticks never change during its
311
+ life. `process.uptime()`-derived and `ps`/`/proc`-derived readings of the
312
+ *same* process instance agree within the same tolerance (2000ms, which
313
+ also absorbs each source's own second-granularity rounding) — this is
314
+ cross-checked against a real OS reading by
315
+ `scripts/e2e-cross-worktree-real-processes.ts`. Windows has no supported
316
+ source for a *live ancestor's* start time — see "Ancestor-chain
317
+ detection" — so an ancestor still cannot be identity-verified there, even
318
+ though this process's own id is now always available.
319
+ - A match is only trusted when **both** sides prove the same identity: the
320
+ registry entry's own `processStartId` **and** a fresh re-derivation of
321
+ that live pid's start time (from the ancestor's own current snapshot)
322
+ agree within tolerance. A pid with a registry entry but a mismatched — or
323
+ unprovable, on either side — identity is walked past exactly like an
324
+ untracked hop, not treated as a match; if nothing further up the chain is
325
+ provable either, ancestry detection reports "no tracked ancestor," which
326
+ is the same safe fallback as if the registry were empty (this process
327
+ classifies as a confirmed `orchestrator`, never `uncertain`, from an
328
+ unprovable candidate alone).
329
+ - A legacy entry with no `processStartId` at all (written by a kankaku
330
+ build predating this field) is never trusted for identity matching or
331
+ kept around: it reads as stale and is removed by the normal sweep the
332
+ next time any process writes its own entry.
333
+
334
+ #### Registry cleanup and health
335
+
336
+ - Every kankaku process removes its own entry file on a normal exit and on
337
+ `session_shutdown` (best-effort, verifying the on-disk file's `pid` and
338
+ `processStartId` still match its own before unlinking, so it can never
339
+ remove a file it does not verifiably own) — a crash still leaves the
340
+ entry for the next sweep. Immediately before unlinking a *discarded*
341
+ entry, the sweep also re-reads that file and compares it byte-for-byte
342
+ against what it judged stale: if the pid was reused and a fresh entry
343
+ already written to the same path in the meantime, the file is left alone
344
+ instead of destroying a live registration the sweep never actually
345
+ evaluated.
346
+ - The opportunistic sweep (run whenever any process writes its own entry)
347
+ removes: entries for a dead pid; entries whose pid is alive but whose
348
+ recorded identity no longer matches that live process (pid reuse); and,
349
+ as a last resort, entries older than 7 days regardless of
350
+ aliveness/identity. An entry with **no verifiable identity at all**
351
+ (legacy/malformed, no `processStartId`) is *never itself* grounds for
352
+ deletion while its pid is alive and within the age ceiling — such an
353
+ entry is never *used* for ancestor matching either way (matching always
354
+ requires a verifiable `processStartId` on both sides), but deleting it
355
+ outright used to risk un-registering a genuinely live orchestrator whose
356
+ own start-time read happened to fail, at the mercy of an unrelated
357
+ sibling process's sweep. It still gets cleaned up the ordinary way, once
358
+ its pid dies or it ages out. The sweep never removes the entry the
359
+ writing process itself just wrote.
360
+ - `/kankaku doctor` reports registry health: how many entries it currently
361
+ trusts, how many it would discard, and why (dead / stale-reuse /
362
+ over-age).
363
+
364
+ ### Ancestor-chain detection
365
+
366
+ Reading "a live tracked ancestor process" above requires one OS-level
367
+ ancestor-chain snapshot. On Linux this is a set of `/proc/<pid>/stat` reads
368
+ (ppid and start-time ticks together, plus one `/proc/uptime` read); on
369
+ macOS, one `ps -eo pid,ppid,etime` snapshot (ppid and
370
+ elapsed-time-since-start together); a shell-wrapper hop with no registry
371
+ entry of its own is walked past, not stopped at.
372
+
373
+ **Startup cost.** This snapshot is taken at most once per process, at
374
+ extension startup, never on a later hot path — and, since it is the only
375
+ part of startup that ever spawns anything, it is skipped entirely unless
376
+ there is something for it to find: the machine-wide registry is read
377
+ *first*, and the snapshot is only taken when at least one other entry
378
+ exists that could possibly be this process's ancestor. The common case (no
379
+ other kankaku process running on the machine at all) therefore never
380
+ spawns `ps` or reads `/proc` — this process's own identity
381
+ (`processStartId`) is unaffected, since it comes from `process.uptime()`
382
+ instead (see "Identity, not just pid" above).
383
+
384
+ **On a platform or environment where this mechanism cannot run at all** —
385
+ Windows (no supported mechanism in this version), or any platform where a
386
+ fresh attempt still fails (`ps`/`/proc` missing, timing out, or producing
387
+ unreadable output) — ancestor-chain detection degrades gracefully to "no
388
+ ancestor found" (never a spawn attempt beyond the one failed try, never a
389
+ crash). Critically, this does **not** mean every unmarked process there is
390
+ classified `uncertain`: with no way to check, kankaku falls back to the
391
+ same marker-only detection it used before this feature existed
392
+ (`GENTLE_PI_AGENTS_CHILD=1`/`KANKAKU_ROLE=subagent` → subagent, anything
393
+ else → confirmed orchestrator) — the deliberately chosen default, because
394
+ marking *every* genuine top-level session `uncertain` on such a platform
395
+ would drop all of that user's work, which is far worse than the narrow
396
+ overcount risk this guards against elsewhere. The trade-off is visible, not
397
+ silent: `/kankaku doctor` reports ancestor-chain detection as unavailable
398
+ whenever this happens (distinguishing it from "checked, no tracked
399
+ ancestor found" — a separate, always-accurate report never folded into
400
+ `roleConfidence`) and names `KANKAKU_ROLE` as the remedy for a genuine
401
+ subagent system that needs marking explicitly on such a platform — see
402
+ "Interactive sessions and `KANKAKU_ROLE`" below.
403
+
404
+ ### The `/kankaku doctor` diagnostic
405
+
406
+ `/kankaku doctor` reports, with no network call:
407
+
408
+ - How many records are orphaned subagents, and why.
409
+ - How many are `uncertain`, and why.
410
+ - Whether ancestor-chain detection is actually usable right now (see
411
+ above) — and, when it is not, a reminder that an unmarked subagent
412
+ system on this platform/environment may be counted twice, with
413
+ `KANKAKU_ROLE` named as the fix.
414
+ - `KANKAKU_ROLE`, when it decided this process's role, as the deciding
415
+ signal — or, when it did not (a confirmed child marker took precedence,
416
+ or an interactive session's `subagent` override was ignored — see
417
+ "Interactive sessions and `KANKAKU_ROLE`" below), the contradiction and
418
+ the resolved outcome instead.
419
+ - Whether this process is a subagent that could not write to its
420
+ orchestrator's directory and fell back to its own local one (see
421
+ "Cross-worktree write routing" above) — a hint to go reunite that record
422
+ manually, since `worklog.jsonl` can never be rewritten after the fact.
423
+ - Registry health (see "Registry cleanup and health" above).
424
+ - The current session's non-default session directory, when set.
425
+
426
+ The plain `/kankaku` summary also appends a one-line hint (`N uncertain
427
+ record(s) excluded from tasks — run /kankaku doctor`) whenever any exist,
428
+ so an undercount is never silent.
429
+
430
+ ### Interactive sessions and `KANKAKU_ROLE`
431
+
432
+ Every subagent mechanism kankaku recognises today launches its child
433
+ **non-interactively**, over pipes (gentle-pi's `--mode rpc`, pi's own
434
+ bundled `subagent` example's `--mode json -p`, `pi-subagents`) — a human
435
+ never sits in front of one. A process running as an **interactive TUI
436
+ session** (`ctx.mode === "tui"`, pi's own signal for "a real terminal, a
437
+ human is here") is therefore always treated as a genuine top-level session
438
+ and is **never** classified `uncertain`, even when some ancestor in its
439
+ process chain happens to be a tracked pi process (for example, pi launched
440
+ from inside another pi's `bash` tool). Interactivity can only be known once
441
+ pi's own `ExtensionContext` is available, at `session_start` — later than
442
+ this process's binary `role` (orchestrator vs. subagent) is decided, but
443
+ `roleConfidence` is deferred and finalised exactly once, then, and stays
444
+ stable for the rest of the process's life.
445
+
446
+ **`KANKAKU_ROLE=orchestrator` or `KANKAKU_ROLE=subagent`** is an explicit
447
+ escape hatch — validated; any other value is ignored, falling back to
448
+ normal detection. Use it to force a session kankaku still gets wrong: mark
449
+ a genuine subagent system it does not recognise as `subagent` (this is
450
+ also the remedy `/kankaku doctor` names when ancestor-chain detection is
451
+ unavailable on the current platform), or force a session `orchestrator`
452
+ regardless of what its ancestry looks like. It has no effect on a record
453
+ already written — see "The role/state model" above.
454
+
455
+ **Scope it to one invocation. Never export it in a shell rc, tmux config,
456
+ or CI environment file.** `process.env` is inherited by every OS child by
457
+ default: an exported `KANKAKU_ROLE` reaches every `pi` invocation that
458
+ shell/session ever starts, subagents included. Set it only on the one
459
+ command it is meant for:
460
+
461
+ ```
462
+ KANKAKU_ROLE=orchestrator pi ...
463
+ ```
464
+
465
+ **Precedence (rewritten for a real bug — a BLOCKER fix).** `KANKAKU_ROLE`
466
+ no longer overrides every other signal unconditionally:
467
+
468
+ 1. A **confirmed child marker** (`GENTLE_PI_AGENTS_CHILD=1`, set only by
469
+ the subagent runner itself, never something a shell rc/tmux/CI
470
+ environment would export) **always wins**, even over an explicit
471
+ `KANKAKU_ROLE=orchestrator`. Without this, a `KANKAKU_ROLE=orchestrator`
472
+ export that leaked into a shell rc — the natural thing to do after
473
+ hitting a false `uncertain` once — would turn every one of that shell's
474
+ later subagent invocations into a confirmed, independently-billed
475
+ orchestrator: systematic multi-counting, invisible until someone
476
+ compares the hub totals against what actually happened.
477
+ 2. `KANKAKU_ROLE=subagent`, with no confirmed marker, is **ignored for an
478
+ interactive session** (`ctx.mode === "tui"`). No subagent mechanism
479
+ kankaku recognises ever launches its child interactively, so this is
480
+ almost always the *mirror* leak — a globally exported
481
+ `KANKAKU_ROLE=subagent` reaching a genuine top-level terminal session —
482
+ and honouring it would silently drop that session's own work from every
483
+ report and the hub (an orphaned subagent record that never anchors a
484
+ task), with no way to recover it later, since `worklog.jsonl` is
485
+ append-only. Between kankaku's two guiding rules — "undercount is
486
+ recoverable, overcount is not" (which governs the *opposite* risk,
487
+ inventing extra billing, and does not apply to this contradiction) and
488
+ "never silently drop genuine work" — this one is governed by the
489
+ second: the override is ignored, the session is classified
490
+ `orchestrator` (what it structurally must be), and the contradiction is
491
+ surfaced once via `ctx.ui.notify` (a warning) at `session_start` and in
492
+ `/kankaku doctor` — never resolved silently. `KANKAKU_ROLE=orchestrator`
493
+ has no such exception: forcing a session `orchestrator` can never drop
494
+ work, only (rarely) invent a task that should not exist, a risk the
495
+ user accepted by setting it explicitly.
496
+ 3. Otherwise `KANKAKU_ROLE`, when set to a recognised value, decides — as
497
+ before.
498
+
499
+ `/kankaku doctor` reports `KANKAKU_ROLE` as the deciding signal only when
500
+ it actually decided anything: it flags "override present AND child marker
501
+ present" with the resolved outcome (`subagent`, per rule 1) when both are
502
+ set, and reports the resolved `orchestrator` outcome (per rule 2) when a
503
+ `subagent` override was ignored for an interactive session — in neither
504
+ case does it claim the override was the deciding signal.
505
+
506
+ **Non-propagation.** `KANKAKU_ROLE` decides only the process that reads
507
+ it. kankaku strips it from its own `process.env` right after reading it
508
+ (before spawning anything), so a child it spawns — a subagent runner, a
509
+ tool shell — never inherits it, even when this process's own copy came
510
+ from something outside kankaku's control (a shell rc, tmux, CI). This is
511
+ a second, independent layer on top of rule 1 above: rule 1 already
512
+ neutralises a leaked `KANKAKU_ROLE=orchestrator` for any *recognised*
513
+ subagent mechanism (its confirmed marker always wins regardless), but
514
+ stripping means the leak can never reach an *unrecognised* one, or any
515
+ other child process, either.
516
+
517
+ **Captured once per process, survives `/new`/`/resume`/`/fork`/`/reload`.**
518
+ pi re-invokes an extension's factory function in the SAME OS process for
519
+ each of those (it "reloads and rebinds extensions" for the new session);
520
+ kankaku reads `KANKAKU_ROLE` and decides `role` from the very first
521
+ invocation and reuses that exact result for every later one in the same
522
+ process, so a `KANKAKU_ROLE=orchestrator` you set for one `pi` command
523
+ stays honoured across every `/new`/`/resume`/`/fork`/`/reload` you run
524
+ inside that same session, not just the first. This does **not** widen the
525
+ "scope it to one invocation" rule above — it still applies only to the one
526
+ `pi` process you set it on, and is still stripped from that process's own
527
+ `process.env` right after the first read, so it is still never inherited
528
+ by anything that process spawns. It only means "one invocation" is
529
+ honoured for as long as that OS process stays alive, across every reload,
530
+ rather than being silently forgotten the moment pi reloads extensions
531
+ internally.
532
+
533
+ ### Subagent profiles (phase 6b)
534
+
535
+ kankaku recognises a subagent-opening tool call through a `SubagentProfile`
536
+ (one per ecosystem package), not a single hardcoded tool name. Three
537
+ profiles are built in:
538
+
539
+ - **gentle-pi** (first-class): `subagent_run`, joined by explicit `taskId`
540
+ (`result.details.gentleAgents`), confirmed by `GENTLE_PI_AGENTS_CHILD=1`.
541
+ Nothing about gentle-pi changes — every field it already exposed (agent,
542
+ mode, taskId, live status, cross-worktree `cwd`) still does.
543
+ - **pi's bundled reference example**: the `subagent` tool, no env marker at
544
+ all — recognised only through ancestry, always starts `uncertain` until
545
+ the registry/ancestor-chain mechanism above corroborates it.
546
+ - **pi-subagents**: also registers a tool named `subagent`, confirmed by
547
+ `PI_SUBAGENT_DEPTH` (present with any value — its own recursion-depth
548
+ counter, not a fixed sentinel).
549
+
550
+ Two packages registering a tool with the exact same name (`subagent`) is a
551
+ real ambiguity kankaku never guesses through: which ecosystem package
552
+ actually made a given call can only be told apart by its child-env marker
553
+ (present in the *child* process, not visible from the parent's tool-call
554
+ alone), so a call to `subagent` still opens a span (best-effort agent/mode,
555
+ kept only when every candidate profile that reports one agrees), but is
556
+ never attributed to one specific profile unless a marker resolves it.
557
+ **Nothing money- or join-affecting is ever taken from an ambiguous call
558
+ either** — no `usage`, no `taskId` — even when one of the colliding
559
+ profiles would normally forward one, because kankaku cannot tell whether
560
+ that specific call actually came from that profile. `/kankaku doctor`
561
+ reports this as an "ambiguous tool name" line.
562
+
563
+ **`KANKAKU_SUBAGENT_TOOLS`** registers one or more additional tool names as
564
+ subagent-opening spans, comma-separated, parsed exactly like
565
+ `KANKAKU_INTERACTIVE_TOOLS` — always additive to the built-ins, never
566
+ replacing gentle-pi's own recognition.
567
+
568
+ **`KANKAKU_SUBAGENT_CHILD_ENV`** registers one or more child-process env
569
+ markers that confirm a process as this configured tool's subagent,
570
+ `;`-separated `NAME=VALUE` (exact match) or a bare `NAME` (presence-only,
571
+ any non-empty value) — mirrors `KANKAKU_SEGMENTS`'s tolerant parsing:
572
+ malformed entries are skipped, not fatal.
573
+
574
+ ```
575
+ KANKAKU_SUBAGENT_TOOLS=my_subagent_tool
576
+ KANKAKU_SUBAGENT_CHILD_ENV=MY_TOOL_CHILD=1
577
+ ```
578
+
579
+ **What NOT to use as a marker.** A configured marker must be exclusive to
580
+ the child process your subagent tool actually spawns — never an ambient
581
+ variable pi, your shell, npm, or the OS sets on *every* process. kankaku
582
+ rejects an obviously-ambient name outright at load time (case-insensitive):
583
+ `PI_CODING_AGENT` and `AI_AGENT` (pi sets both on every process it runs,
584
+ not just a subagent's child), the generic shell/OS variables `PATH`,
585
+ `HOME`, `USER`, `SHELL`, `PWD`, `CI`, `LANG`, `TMUX`, and anything prefixed
586
+ `PI_`, `TERM`, `LC_`, `NODE_`, `NPM_`, or `KANKAKU_`. A rejected marker
587
+ never reaches the configured profile — it is reported once via
588
+ `ctx.ui.notify` and listed in `/kankaku doctor`, never silently accepted.
589
+ This denylist cannot enumerate every possible ambient variable, though, so
590
+ there is a second, runtime layer: **a configured marker never demotes an
591
+ interactive session**, exactly like `KANKAKU_ROLE=subagent` already does
592
+ not (see "Interactive sessions and `KANKAKU_ROLE`" above) — if a configured
593
+ marker matches on a session that turns out to be interactive, kankaku
594
+ treats it as the orchestrator it structurally must be and warns once
595
+ (escalated to a stronger warning when that session also has no tracked
596
+ ancestor at all, the clearest sign the "marker" is actually ambient). A
597
+ **built-in** marker (`GENTLE_PI_AGENTS_CHILD`, `PI_SUBAGENT_DEPTH`) keeps
598
+ the unconditional precedence it always had — no built-in mechanism kankaku
599
+ recognises ever launches its child interactively, so this exception never
600
+ actually applies to it in practice.
601
+
602
+ **Verify with `/kankaku doctor`.** After configuring
603
+ `KANKAKU_SUBAGENT_CHILD_ENV`, run `/kankaku doctor` from an ordinary
604
+ top-level session: it must **not** report a "configured marker" or
605
+ "rejected marker" line for a ordinary interactive session. If it does, the
606
+ chosen name is either denylisted or ambient enough to trip the interactive
607
+ guard — pick something the third-party tool's own child process sets that
608
+ nothing else on the system would ever set.
609
+
610
+ A confirmed marker from a configured profile that passes both layers above
611
+ still takes the same "always wins over `KANKAKU_ROLE`" precedence gentle-pi's
612
+ own marker already had for a **non-interactive** process — see "Interactive
613
+ sessions and `KANKAKU_ROLE`" above.
614
+
615
+ `/kankaku doctor` reports the active profile set, any configured tools/
616
+ markers, which profile matched each subagent record (or "unmatched" when
617
+ no marker resolved it), any rejected marker names with why, and a
618
+ configured-marker-ignored-for-interactivity contradiction when one occurs.
619
+
620
+ ### In-process subagents (phase 6c)
621
+
622
+ A subagent tool result's `usage` field — pi's own documented convention
623
+ for "a tool making nested LLM calls should return their combined `Usage`
624
+ as `usage`" — is recorded on the span itself (never folded into the
625
+ triggering record's own usage totals at write time any more), and added to
626
+ the *task's* aggregate total by `buildTasks` — the one place per-task
627
+ usage is ever assembled — except when this same task also has a joined
628
+ child record confirmed by the **same** profile: that child's own usage
629
+ already carries this cost through its own confirmed-marker/ancestry join,
630
+ so the span's forwarded figure is excluded instead of counted a second
631
+ time. gentle-pi is unaffected (its result never carries one — cost for its
632
+ children is, and stays, tracked through the registry/ancestry join above).
633
+ A profile whose marker can also produce an ancestry-joined child record
634
+ with its own usage (pi-subagents) never forwards `usage` even when its
635
+ result happens to carry one, to avoid counting the same nested work twice
636
+ by construction; a *configured* profile that declares **both** a marker
637
+ and forwards usage relies on the runtime reconciliation above instead (see
638
+ "Subagent profiles (phase 6b)"). **Usage is never forwarded for an
639
+ ambiguous tool-name match** (2+ profiles registering the same name, e.g.
640
+ `subagent`) — see "Subagent profiles (phase 6b)" above.
641
+
642
+ Real in-process (same-OS-process, no separate `pid`) subagent nesting was
643
+ investigated directly against pi's own source and documented API
644
+ (`docs/extensions.md`) for this release: none of gentle-pi, pi's bundled
645
+ reference example, or pi-subagents actually run a child *inside* the
646
+ parent's process — every one of them spawns a real, separate OS process.
647
+ pi's own in-process mechanism (`ctx.newSession`/`ctx.fork`) replaces one
648
+ session with another *sequentially* in the same process (the old session's
649
+ `session_shutdown` fires, then the new one's `session_start` — never
650
+ concurrently), which is exactly what "kankaku reads/writes a fresh record
651
+ per session_start, same pid" already handles correctly. As a defensive
652
+ guard for the pattern true concurrent nesting *would* leave behind,
653
+ `/kankaku doctor` flags two confirmed-orchestrator records sharing a pid
654
+ with **overlapping** `[startedAt, settledAt]` windows as "likely
655
+ in-process nesting", unioning (never summing) their wall time via the
656
+ same interval-union primitive `buildTasks` itself uses — informational
657
+ only, it never changes a task's own numbers. This has not been observed
658
+ from any real subagent mechanism in this codebase's research; if pi (or an
659
+ extension built on its SDK) grows genuine concurrent in-process nesting in
660
+ the future, this is the signal that would surface it.
661
+
662
+ ### Limitations, honestly
663
+
664
+ - **Windows has no ancestor-chain detection** (an ancestor can never be
665
+ identity-verified there), though this process's own `processStartId` is
666
+ always available regardless of platform — see "Identity, not just pid"
667
+ above. Mark a genuine subagent system explicitly with `KANKAKU_ROLE` on
668
+ such a platform; see "Interactive sessions and `KANKAKU_ROLE`" above.
669
+ - **gentle-pi's child cannot currently read its own task id** — the
670
+ cross-worktree join above relies on ancestry plus the registry, not on an
671
+ explicit shared id, because upstream gentle-pi does not hand the child
672
+ process its task id today. If that changes upstream, a future kankaku
673
+ version can upgrade this join to a higher-confidence explicit-id match.
674
+ - **Ancestor-chain detection only sees the chain as it exists when a
675
+ process looks.** A detached child reparented to init/launchd before that
676
+ point cannot recover its original ancestry this way — the same limitation
677
+ the existing `pid`/`parentPid` capture already has (see AGENTS.md).
678
+ - **A record already written `uncertain` (or already routed to a fallback
679
+ local directory) cannot be rewritten.** `worklog.jsonl` is append-only;
680
+ fixing the underlying cause (upgrading kankaku, setting `KANKAKU_ROLE`,
681
+ restoring access to an orchestrator's directory) only helps a *later*
682
+ run's records, never edits a line already on disk. There is no migration
683
+ planned for this — it follows directly from "never rewrite the log" (see
684
+ AGENTS.md).
111
685
 
112
686
  ## The `/kankaku` command
113
687
 
@@ -128,10 +702,28 @@ Arguments are whitespace-separated and order-insensitive:
128
702
  - `/kankaku client <name>` — set the billing client for the current pi
129
703
  session. `/kankaku client` alone shows the effective client and which
130
704
  source it came from; `/kankaku client --clear` removes the session-level
131
- override. See "Billing labels" below.
705
+ override. See "Billing labels" below. When a hub is configured, `<name>`
706
+ must match a catalog client's code or name (case-insensitive) instead of
707
+ being free text — see "Hub (PocketBase)".
132
708
  - `/kankaku clients` — one line per client (work/waiting/wall time, cost,
133
709
  task count) for today. Add `all` for every day. Tasks with no resolved
134
710
  client are grouped under `(none)`.
711
+ - `/kankaku doctor` — orphan/uncertain subagent record counts and why,
712
+ plus ancestor-detection platform availability. No network call. See
713
+ "Subagents".
714
+
715
+ The following are available only when a hub is configured (see "Hub
716
+ (PocketBase)" below):
717
+
718
+ - `/kankaku target` — show the effective client/project and which source
719
+ produced it. `/kankaku target pick` runs the picker again (works
720
+ mid-session; the new target applies to records settled afterwards).
721
+ `/kankaku target clear` clears the session-level target.
722
+ - `/kankaku catalog refresh` — force a catalog refresh and report the
723
+ client/project counts.
724
+ - `/kankaku projects` — one line per project (work/waiting/wall time, cost,
725
+ task count) for today. Add `all` for every day. Tasks with no resolved
726
+ project are grouped under `(no project)`.
135
727
 
136
728
  Cost figures are the sum of `usage.cost` as priced by pi's model table
137
729
  (per-million-token rates in `models.json`, adjustable with `modelOverrides`).
@@ -167,6 +759,273 @@ orchestrator's) client.
167
759
  `sessionName` is also attached to every record from `pi.getSessionName()`,
168
760
  so reports can show which named session produced a task.
169
761
 
762
+ ## Hub (PocketBase)
763
+
764
+ kankaku can optionally resolve the billing client (and a project) **from a
765
+ PocketBase instance** instead of free text, so `cajamar`/`Cajamar`/`cjamar`
766
+ can no longer become three different clients. This is phase 1 of the hub
767
+ integration (catalog + selection only): nothing is uploaded anywhere.
768
+
769
+ ### Configuration
770
+
771
+ Set `KANKAKU_PB_URL`, `KANKAKU_PB_EMAIL`, `KANKAKU_PB_PASSWORD`, or write
772
+ `~/.kankaku/credentials.json`:
773
+
774
+ ```json
775
+ { "url": "https://pb.example.com", "email": "bot@example.com", "password": "secret" }
776
+ ```
777
+
778
+ Environment variables take precedence over the file, field by field. The
779
+ hub URL must be HTTPS unless it points at `localhost`/`127.0.0.1`/`::1`; a
780
+ plain-HTTP URL for any other host is refused (surfaced once via a
781
+ notification). The project's own `<KANKAKU_DIR>/config.json` is never read
782
+ for credentials — it is project-local and frequently committed.
783
+
784
+ `KANKAKU_MACHINE` optionally names this machine (for a multi-machine setup
785
+ later); it defaults to the OS hostname and is attached to every record as
786
+ `machine` once the hub is configured.
787
+
788
+ **When no hub is configured, kankaku behaves exactly as it does today** —
789
+ this whole feature is additive and every existing behaviour, record shape,
790
+ and report stays unchanged.
791
+
792
+ ### Selection
793
+
794
+ On `session_start`, for the orchestrator role with a UI available:
795
+
796
+ 1. **Session** — restored from the last `kankaku-target` session entry
797
+ (including a remembered "skipped" choice, so a reload does not ask
798
+ again).
799
+ 2. **Project config** — `clientId`/`projectId` in `<KANKAKU_DIR>/config.json`.
800
+ 3. **`repo_paths`** — the current working directory matched against each
801
+ project's `repo_paths` (exact match, or a subdirectory of one; the
802
+ longest match wins).
803
+ 4. Otherwise, a picker: `ctx.ui.select` for the client (active clients,
804
+ sorted by name, plus "— skip —"), then for the project (active projects
805
+ of that client, plus "(no project)" and "— skip —"). Declining at either
806
+ step — "— skip —" or dismissing the dialog — cancels the whole pick and
807
+ is remembered for the session.
808
+
809
+ After a pick, kankaku asks whether to remember it for this repository; a
810
+ "yes" merges `clientId`/`projectId` into `<KANKAKU_DIR>/config.json`.
811
+
812
+ An id from any source that no longer resolves to an active, non-"unassigned"
813
+ catalog entry is treated as absent for that source and resolution falls
814
+ through to the next one, exactly like the legacy client precedence.
815
+
816
+ Once a hub target is active for a run, the legacy `client` label is set to
817
+ the target's client `code` (so every existing report/export keeps grouping
818
+ correctly), and the record additionally carries `clientId`, `clientName`,
819
+ and — when a project is selected — `projectId`/`projectName`. A subagent
820
+ never resolves its own target, exactly like the legacy `client` label — the
821
+ task view exposes it from the orchestrator record only.
822
+
823
+ The status bar shows `💼 <client> · <project>` (or just `💼 <client>` without
824
+ a project) in place of the legacy client label, both idle and during a run.
825
+
826
+ ### Caching and offline behaviour
827
+
828
+ The catalog (clients/projects) is cached machine-wide at
829
+ `~/.kankaku/catalog.json` with a 6-hour TTL. On startup: a fresh cache is
830
+ used as-is; a stale cache is used immediately while a refresh happens in
831
+ the background; when there is no cache at all, one refresh is awaited
832
+ (bounded by the hub client's own request timeout, 3s by default) before
833
+ falling back. If the hub is unreachable and there is no cache, kankaku
834
+ notifies once (`kankaku: hub unreachable, using local labels`) and
835
+ continues exactly as it would without a hub configured. `/kankaku catalog
836
+ refresh` forces a refresh on demand. The cache file is always written
837
+ owner-only (`0600`); if kankaku is the first thing to ever create
838
+ `~/.kankaku` itself (no project has put its own `.kankaku` there), the
839
+ directory is created owner-only (`0700`) too — but an already-existing
840
+ `~/.kankaku` is never chmod'd, since it may be a project's own kankaku
841
+ directory (see "The registry" below for the same rule applied to `run/`).
842
+
843
+ ### Privacy (catalog)
844
+
845
+ The catalog itself (clients/projects) is read-only — nothing about *that*
846
+ data is ever written back. Whether your own work records ever leave the
847
+ machine is a separate, opt-in decision: see "Sync" below.
848
+
849
+ ### Sync
850
+
851
+ Once a hub is configured, kankaku can push consolidated **task** rows (see
852
+ "Task and session views" above) to PocketBase, so a project/task manager
853
+ can report AI time and cost per project. This is an outbox pattern:
854
+ `worklog.jsonl` stays the local source of truth, append-only and never
855
+ rewritten, exactly as without a hub. A separate sync step reads it and
856
+ uploads what is pending — nothing in a pi event handler ever waits on the
857
+ network.
858
+
859
+ **What gets uploaded.** One `task_entries` row per task — never raw
860
+ `WorkRecord`s re-aggregated on the server. The union-of-intervals rule
861
+ (`wallMs`, "Task and session views") is computed exactly once, locally, by
862
+ `buildTasks`; the hub only ever sums already-consolidated rows. When
863
+ `KANKAKU_SYNC_RECORDS` is not `0` (the default), each task's underlying
864
+ `WorkRecord`s are also uploaded as `work_records`, raw per-run detail for
865
+ drilling into a task — these rows overlap each other and must never be
866
+ summed, unlike `task_entries`.
867
+
868
+ **Idempotency and the revisit window.** Every task is upserted by its id
869
+ (the orchestrator record's `id`), never blindly created — safe to
870
+ re-send. A task is not final the moment its orchestrator settles: a
871
+ background subagent can settle *after* it and extend the task's union
872
+ (`wallMs`, cost, subagent count) for a task that may already be in
873
+ PocketBase. So every sync revisits a trailing window behind its own
874
+ watermark — `KANKAKU_SYNC_WINDOW_HOURS`, 24h by default — and re-evaluates
875
+ every task whose `endedAt` falls inside it. A cheap content hash per task
876
+ (`<KANKAKU_DIR>/sync-state.json`) means an unchanged task inside the window
877
+ costs nothing: running `/kankaku sync` twice in a row performs zero writes.
878
+
879
+ The window is anchored to `syncedThrough` (the watermark), never to
880
+ current wall-clock time — see "Limitations" below for what that means for
881
+ a background subagent that settles long after its orchestrator, and after
882
+ the directory has otherwise gone quiet.
883
+
884
+ **Assignment is create-only.** You (or whoever reassigns work in the hub's
885
+ web app) can move a task from one client/project to another directly in
886
+ PocketBase — for example, moving a "Sin determinar" row to its real
887
+ client once you have identified it. A later re-sync of that same task
888
+ **must never undo that**: on create kankaku sends the full row, including
889
+ `client`/`project`/`legacy_client_label`; on every subsequent update it
890
+ sends measurement fields only (`wall_ms`, `cost`, `status`, ...) and never
891
+ touches assignment fields again. If you need kankaku itself to change a
892
+ task's assignment, do it in the web app, not by re-syncing.
893
+
894
+ **Historical ("Sin determinar") records.** A record with no `clientId`, or
895
+ whose `clientId` no longer resolves in the catalog, is routed to the hub's
896
+ "Sin determinar" (unassigned) client, carrying its old free-text `client`
897
+ label (or `clientName`) forward as `legacy_client_label` — the exact
898
+ mechanism that lets you bulk-reassign "everything that said `cjamar`" once,
899
+ in the web app, from the unassigned queue.
900
+
901
+ **Agent and measurement quality.** Every `task_entries` row also carries
902
+ who produced it and how well each figure was measured, so the hub can
903
+ label what it has instead of silently blending incompatible numbers from
904
+ different agents: `agent` (`"pi"`), `agent_version` (pi's own version,
905
+ when it could be determined — never guessed, omitted otherwise), `plugin`
906
+ (`"kankaku"`), `plugin_version` (this package's own version),
907
+ `waiting_quality` (always `"measured"` for kankaku/pi — it always
908
+ instruments waiting time), `cost_quality` (`"measured"` when the task's
909
+ own record or any joined subagent observed a real provider cost figure on
910
+ at least one turn; `"unknown"` when none did, e.g. a subscription/OAuth
911
+ provider that reports no cost — kankaku has no token-price estimator, so
912
+ it never sends `"estimated"`), and `subagent_linkage` (`"not_applicable"`
913
+ when the task opened no subagent spans; `"linked"` when at least as many
914
+ child records were joined as spans were opened; `"unlinked"` otherwise —
915
+ a task-level approximation, since there is no per-span correlation id
916
+ today, see "Subagents" > "Limitations"). These are measurement fields, not
917
+ assignment: sent on every create *and* update, and included in the sync
918
+ content hash, so a background subagent that joins later — improving
919
+ `cost_quality`/`subagent_linkage` without changing any other number —
920
+ still triggers a resync. An older hub predating these fields simply
921
+ ignores them (PocketBase silently drops unrecognized fields on write); no
922
+ capability probing is needed.
923
+
924
+ **Session directory.** `session_dir` carries a task's non-default session
925
+ directory (`TaskView.sessionDir`, see "Record schema") to the hub, so a
926
+ resumable session can be resumed from the web, not just locally via
927
+ `/kankaku doctor`. It is optional — only present when pi reports a
928
+ non-default session directory — and, like the fields above, a measurement
929
+ field: sent on both create and update, and included in the sync content
930
+ hash so a session dir change alone triggers a resync. Like `repo_project`,
931
+ it is an absolute local filesystem path (username, disk layout) — the same
932
+ category of exposure the hub already accepts for `repo_project`, not a new
933
+ one. An older hub predating this field simply ignores it (PocketBase
934
+ silently drops unrecognized fields on write).
935
+
936
+ **Privacy.** `KANKAKU_SYNC_PROMPT` controls whether a task's prompt text
937
+ leaves the machine at all: `none` (default — omitted entirely), `truncated`
938
+ (first 120 chars plus `…`), or `full`.
939
+
940
+ **Commands:**
941
+
942
+ - `/kankaku sync` — push everything pending (new tasks, plus anything
943
+ inside the revisit window that changed).
944
+ - `/kankaku sync all` — a full re-evaluation: every task, not just the
945
+ window. Safe and cheap to run — the content hash still skips anything
946
+ unchanged.
947
+ - `/kankaku sync status` — the current watermark, a locally-computed
948
+ pending count (no network), how many never-synced tasks fall outside the
949
+ current revisit window (needs `sync all` — see "Limitations" below), and
950
+ the last sync error, if any.
951
+ - `/kankaku backfill` — a full sync, reported grouped by
952
+ `legacy_client_label`: how many tasks went to "Sin determinar" and under
953
+ which old label, so you know what to reassign in the web app's
954
+ unassigned queue. This never rewrites `worklog.jsonl` locally — the
955
+ reassignment happens once, in PocketBase, and survives every future sync
956
+ (see "Assignment is create-only" above).
957
+
958
+ **Automatic sync.** Unless `KANKAKU_SYNC_AUTO=0`, kankaku also syncs
959
+ automatically on three triggers (orchestrator role only): fire-and-forget
960
+ (never awaited, errors never surface as a failure of the run that
961
+ triggered them) on `session_start` (after crash recovery) and again after
962
+ `agent_settled`; and, on `session_shutdown`, one **awaited**, time-bounded
963
+ sync — pi awaits its `session_shutdown` handlers with no timeout of its
964
+ own, so this is the one place kankaku's own handler awaits the network, up
965
+ to `shutdownSyncTimeoutMs` (default 3 s). This is what makes the last
966
+ prompt(s) of a session reach the hub when the session ends, rather than
967
+ only on the next session's `session_start`: quitting with an unreachable
968
+ hub costs at most that timeout longer, never more, and cleanup (status
969
+ bar, session-client bookkeeping) still runs even if the sync times out or
970
+ fails. All three triggers share one single-flight guard, so they never
971
+ race each other within a process — a `session_shutdown` sync that arrives
972
+ while one is already in flight awaits that same one rather than starting a
973
+ second — and a lock file (`<KANKAKU_DIR>/sync.lock`, an atomic
974
+ exclusive-create so two racing processes can never both acquire it, stale
975
+ after 5 minutes) keeps two pi processes from syncing the same directory
976
+ concurrently. Subagents never sync. None of the three triggers notify on
977
+ success; on failure (including a shutdown timeout) they notify at most
978
+ once per session (`kankaku: sync failed: ...` / `kankaku: shutdown sync
979
+ timed out`) — check `/kankaku sync status` for the details, including on a
980
+ later run.
981
+
982
+ The automatic path is cheap on every prompt, not just fire-and-forget: it
983
+ skips entirely (no read of `worklog.jsonl`, no network) when the log has
984
+ not changed since the last successful sync, for all three triggers.
985
+ Otherwise, only `agent_settled` — fired once per prompt — is throttled, to
986
+ at most once per `KANKAKU_SYNC_MIN_INTERVAL_MINUTES` (default 5; `0`
987
+ disables the throttle); since right after `agent_settled` the log *has*
988
+ just changed (a record was just appended), this throttle is what actually
989
+ keeps that trigger cheap. `session_start` and `session_shutdown` never
990
+ throttle: a session boundary is worth catching up on regardless of how
991
+ recently the last automatic run happened, so a stuck hub does not stay
992
+ silently unsynced across restarts, and the shutdown sync is already
993
+ bounded by its own timeout. None of this ever applies to a manual
994
+ `/kankaku sync`, `sync all`, or `backfill`.
995
+
996
+ **Network/validation failures.** A network or server (5xx) error stops a
997
+ sync run where it is and does not advance its watermark past the failing
998
+ task — nothing is lost, and the next sync (manual or automatic) picks up
999
+ exactly there. A task that fails **validation** (e.g. a genuinely malformed
1000
+ payload) is recorded with its reason and skipped — not retried on every
1001
+ single run — but is retried automatically the moment its content changes.
1002
+
1003
+ **Limitations:**
1004
+
1005
+ - Sync state (`sync-state.json`) is per repository/machine, not
1006
+ centralized; there is no standalone CLI entry point yet (`npx kankaku
1007
+ sync` outside of pi) — see "Roadmap".
1008
+ - **A late background child, and the revisit window (R3).** A background
1009
+ subagent can settle well after its (possibly cross-worktree)
1010
+ orchestrator process has already exited — its record still writes
1011
+ correctly into the orchestrator's `worklog.jsonl` (see "Subagents" >
1012
+ "Cross-worktree write routing"), but nothing *syncs* it until that
1013
+ directory is next visited: pi opened there again (`session_start`'s
1014
+ auto-sync), or `/kankaku sync`/`sync all` run there manually. Subagents
1015
+ themselves never sync (see "Automatic sync" above). An ordinary
1016
+ incremental sync then picks the late child up wherever the task sits: a
1017
+ task the hub **already holds** is re-synced whenever its content changed,
1018
+ inside the revisit window or not, so a row on the hub never goes stale —
1019
+ including a task that *shrank* because a child moved to another task. The
1020
+ window (`syncedThrough - windowHours`) only bounds how far back work that
1021
+ was **never synced** is looked for; `/kankaku sync status` reports how
1022
+ many such tasks there are, and `/kankaku sync all` (or `backfill`)
1023
+ uploads them.
1024
+
1025
+ **What the stale count means.** Only tasks outside the window that were
1026
+ never synced to this hub. A task the hub already holds never appears
1027
+ here: if it changed it is simply re-synced.
1028
+
170
1029
  ## Tagged segments
171
1030
 
172
1031
  While a run is open, kankaku can also time tool executions that match a
@@ -201,9 +1060,11 @@ it.
201
1060
 
202
1061
  ## Crash recovery
203
1062
 
204
- While a run is open, each pi process periodically writes a checkpoint of
205
- its current record to `<KANKAKU_DIR>/inflight/<pid>.json` (after every
206
- `turn_end` and `tool_execution_end`), and removes it on a normal
1063
+ While a run is open, each pi process writes a checkpoint of its current
1064
+ record to `<KANKAKU_DIR>/inflight/<pid>.json` — first as soon as the run
1065
+ starts (`before_agent_start`), so even a crash on the very first turn still
1066
+ leaves a checkpoint, and then again after every `turn_end` and
1067
+ `tool_execution_end` — and removes it on a normal
207
1068
  `agent_settled`/`session_shutdown`. If the process is killed outright
208
1069
  (`kill -9`, power loss) before it can settle, the checkpoint file survives
209
1070
  it. On the next pi start, `session_start` scans `inflight/` for checkpoints
@@ -213,6 +1074,13 @@ whose owning pid is no longer alive, appends each one to `worklog.jsonl` as
213
1074
  recovered record is the time of its last checkpoint, not the actual crash
214
1075
  time, so `wallMs`/`workMs` are a **lower bound** on the real duration.
215
1076
 
1077
+ The same scan also sweeps `inflight/` for orphaned `.tmp` files: `save`
1078
+ writes to a temp file before renaming it into place, and a process killed
1079
+ between those two steps leaves the temp file behind. A stray `.tmp` file is
1080
+ deleted once its writer pid is no longer alive (or its name cannot be
1081
+ parsed); one still owned by a live writer — including this very process's
1082
+ own in-progress write — is left alone.
1083
+
216
1084
  ## Export
217
1085
 
218
1086
  `/kankaku export [csv|json] [all]` writes one flat row per task (today's
@@ -252,9 +1120,64 @@ Columns (in this order for CSV; the same fields for JSON):
252
1120
  `ask_user_question,ask_user_choice`.
253
1121
  - `KANKAKU_SEGMENTS`: `;`-separated `tag=tool:regex` rules for tagged
254
1122
  segments (see above). Defaults to the single `review` rule.
1123
+ - `KANKAKU_SUBAGENT_TOOLS`: comma-separated list of additional tool names
1124
+ treated as subagent-opening spans, parsed exactly like
1125
+ `KANKAKU_INTERACTIVE_TOOLS`. Always additive to the built-in profiles
1126
+ (gentle-pi, pi's bundled reference example, pi-subagents) — never
1127
+ replaces gentle-pi's own recognition. See "Subagents" > "Subagent
1128
+ profiles (phase 6b)". Unset by default (built-in profiles' tool names
1129
+ only).
1130
+ - `KANKAKU_SUBAGENT_CHILD_ENV`: `;`-separated `NAME=VALUE` (exact match) or
1131
+ bare `NAME` (presence-only) child-process env markers that confirm a
1132
+ process as the configured tool's subagent — parsed like `KANKAKU_SEGMENTS`,
1133
+ malformed entries skipped. The separator is `;`, **not** the `,` that
1134
+ `KANKAKU_SUBAGENT_TOOLS` takes: a name that is not a valid environment
1135
+ variable name (such as `A,B`) is rejected and reported, never silently
1136
+ accepted. A name that looks pi/shell/OS/npm-owned
1137
+ (`PI_CODING_AGENT`, `AI_AGENT`, `PATH`, `HOME`, `USER`, `SHELL`, `PWD`,
1138
+ `CI`, `LANG`, `TMUX`, or a `PI_`/`TERM`/`LC_`/`NODE_`/`NPM_`/`KANKAKU_`
1139
+ prefix, case-insensitive) is rejected outright, and even an accepted
1140
+ marker never demotes an interactive session — see "Subagents" > "Subagent
1141
+ profiles (phase 6b)" for both layers, and verify with `/kankaku doctor`.
1142
+ Unset by default.
1143
+ - `KANKAKU_ROLE`: `orchestrator` or `subagent` — an explicit escape hatch
1144
+ for this process. Any other value is ignored. Scope it to one
1145
+ invocation (`KANKAKU_ROLE=orchestrator pi ...`) — **never export it in
1146
+ a shell rc, tmux config, or CI environment file**: a confirmed child
1147
+ marker always wins over `KANKAKU_ROLE=orchestrator`, `KANKAKU_ROLE=
1148
+ subagent` is ignored for an interactive session, and kankaku strips it
1149
+ from the environment it passes to any child it spawns, but none of that
1150
+ helps if it reaches a session it was never meant for in the first
1151
+ place. See "Subagents" > "Interactive sessions and `KANKAKU_ROLE`" for
1152
+ the full precedence.
255
1153
  - `KANKAKU_CLIENT`: default billing client for this project (see "Billing
256
1154
  labels" above). Lower precedence than the session-level
257
1155
  `/kankaku client` override, higher than `<KANKAKU_DIR>/config.json`.
1156
+ - `KANKAKU_PB_URL`, `KANKAKU_PB_EMAIL`, `KANKAKU_PB_PASSWORD`: hub
1157
+ (PocketBase) credentials (see "Hub (PocketBase)" above). Take precedence,
1158
+ field by field, over `~/.kankaku/credentials.json`.
1159
+ - `KANKAKU_MACHINE`: this machine's display name for the hub, attached to
1160
+ every record as `machine` once the hub is configured. Defaults to the OS
1161
+ hostname.
1162
+ - `KANKAKU_SYNC_PROMPT`: prompt privacy for sync — `none` (default, omitted
1163
+ entirely), `truncated` (first 120 chars + `…`), or `full`. See "Hub
1164
+ (PocketBase)" > "Sync" > "Privacy".
1165
+ - `KANKAKU_SYNC_WINDOW_HOURS`: how far behind the sync watermark to revisit
1166
+ on every run, so a subagent that settles after its orchestrator still
1167
+ reaches its task. Defaults to 24; a non-positive or non-numeric value
1168
+ falls back to the default.
1169
+ - `KANKAKU_SYNC_RECORDS`: `0` disables uploading `work_records` (raw
1170
+ per-`WorkRecord` detail); `task_entries` are always uploaded regardless.
1171
+ Defaults to enabled.
1172
+ - `KANKAKU_SYNC_AUTO`: `0` disables the automatic `session_start`/
1173
+ `agent_settled`/`session_shutdown` sync; `/kankaku sync` still works.
1174
+ Defaults to enabled.
1175
+ - `KANKAKU_SYNC_MIN_INTERVAL_MINUTES`: how often the automatic
1176
+ `agent_settled` sync is allowed to actually run, at most — see
1177
+ "Automatic sync" above. Defaults to 5; `0` disables the throttle. Only
1178
+ ever applies to `agent_settled`: `session_start` and `session_shutdown`
1179
+ are never throttled, and none of this applies to a manual `/kankaku
1180
+ sync`, `sync all`, or `backfill`.
258
1181
 
259
1182
  ## Limitations
260
1183
 
@@ -269,5 +1192,30 @@ Columns (in this order for CSV; the same fields for JSON):
269
1192
 
270
1193
  ## Roadmap
271
1194
 
272
- - Remote sync service: the `id` and `schema` fields are already in place for
273
- a future `synced` cursor that uploads records to a remote store.
1195
+ - Hub sync (phase 2): push consolidated task rows to PocketBase (outbox
1196
+ pattern, idempotent upsert by task id) so a task/project manager can
1197
+ report AI time and cost per project. The catalog/selection layer in "Hub
1198
+ (PocketBase)" above is phase 1; sync itself ("Hub (PocketBase)" > "Sync")
1199
+ is phase 2 — both already shipped.
1200
+ - A standalone CLI entry point (`npx kankaku sync`, for a cron/launchd job
1201
+ outside of any pi session) is deliberately not included yet: Node refuses
1202
+ type stripping for a `.ts` file under `node_modules`, so a bin script
1203
+ needs a build step this package does not have yet. `sync-runner.ts` and
1204
+ its adapters are already decoupled from pi so that build step is the only
1205
+ missing piece.
1206
+ - Linking a `task_entries` row to an existing `tasks` record (phase 3 in the
1207
+ hub's own data model) — kankaku never invents tasks; it would only ever
1208
+ link to one created in the manager.
1209
+ - Generic subagent detection (phase 6): 6a fixed the two correctness bugs
1210
+ described in "Subagents" above (a phantom-orchestrator double count; a
1211
+ gentle-pi cross-worktree child's work going missing). 6b added the
1212
+ `SubagentProfile` abstraction, built-in profiles for pi's bundled
1213
+ reference example and pi-subagents alongside gentle-pi, and
1214
+ `KANKAKU_SUBAGENT_TOOLS`/`KANKAKU_SUBAGENT_CHILD_ENV` for a third-party
1215
+ tool kankaku does not recognise out of the box. 6c added subagent-result
1216
+ `usage` forwarding and the same-pid overlapping-orchestrator guard — see
1217
+ "Subagents" > "Subagent profiles (phase 6b)" / "In-process subagents
1218
+ (phase 6c)" above. Built-in profiles beyond these three
1219
+ (`pi-background-tasks`, `@d3ara1n/pi-subagent`), gentle-pi handing a
1220
+ child its own task id, and the cosmetic `linked_task_id` hub
1221
+ self-relation remain out of scope.